Data and GDPR in AI projects: the questions that need answers
Data protection in AI projects: purpose, processing environments, third-party training, access, logs, retention and deletion before you build.
Mauricio Zaffari
Before an AI project gets approved, four questions reach whoever owns the architecture: where is our data processed, does the provider train on it, who can access it, and how does it get deleted. The questions often show up as objections, and they deserve direct answers, including when the honest answer is "it depends on your contract, and that is exactly where to look".
This post answers the four questions one by one, with one document flow as the running example: a company that processes operational documents with AI support. A test: can your next project meeting answer the four questions in writing?
Where is the data processed?
Direct answer: in every environment the flow touches, and each one needs to be mapped. A single flow can pass through the company database, a cloud service, a contracted language model, and the browser of whoever uses the solution. Each environment has its own rules for access and configuration. In this post's example, an operational document passes through the company database, a cloud service and a contracted language model.
The practical question is where sensitive data is exposed. In some designs, only the necessary part goes to the model and the rest stays in the company environment. In others, the whole record moves. The difference lies less in the model and more in how the flow was built.
flowchart TD
source["Source system"]
company["Company environment"]
log["Event log for operation and audit"]
ai["AI service"]
result["Output"]
review{"Human review of exceptions"}
retention["Retention defined -> deletion"]
source -->|"access control"| company
company --> log
company -->|"contract defines use"| ai
ai --> result
result --> review
review --> retention
For each environment, answer what data arrives, what data leaves, who operates it, and under which conditions. The architecture can include separated environments, encrypted traffic and protected credential management, according to the project requirements. None of that is automatic. It has to be in scope and verifiable.
If the question "where is this data processed?" has no clear answer, the project is not ready to start.
Is the data used to train third-party models?
Direct answer: it depends on the services and contracts adopted, and that is why the answer is settled in the contract and the configuration, not in conversation. Some providers let you turn training use off. Others handle the data differently depending on the contracted plan. In the example, the AI service contract has to define whether the documents can be used for training.
Before deployment, the project defines which providers may process the data, which privacy settings apply, and which uses are allowed, formalized in writing. If a provider does not allow the necessary setting, that is a project constraint, not a detail: it may require switching the service, changing the architecture, or narrowing the data that moves. Better to find out during the diagnostic than in the first audit.
Who can access the data, and what gets logged?
Direct answer: access follows authorized roles, permissions, and integrations, for people and for services, and the log needs to record what operation and audit need to see.
A useful log shows when data was accessed, through which flow, and with what result. A useless log records everything and no one reads it. There is a trade-off here, and we name it: a detailed log helps investigate but stores information that itself needs protection; a minimal log reduces exposure but makes incidents harder to understand. The usual path is to record what is needed to operate and audit, restrict access to the roles that need it, and review that list as the process changes.
In flows with automated decisions, include the human review points in the same map: where a person validates information, stops the flow, or corrects a decision. Without that point, an exception can pass without anyone seeing it.
How long is the data kept, and how is it deleted?
Direct answer: each piece of data gets a destination at the end of the flow, defined before the build. Projects tend to design the entry well and leave the exit undefined, which means data stays somewhere for an indefinite time because no one agreed on when it leaves.
The question to answer is whether the data needs to keep existing after the flow ends, for operation, audit, or a legal obligation. If yes, for how long and with which protection. If no, what is the deletion path and how can it be verified.
In the running example, the company processing operational documents with AI support: the document enters, is read in a separated environment, the extracted fields feed the source system, and the checking record stays for audit. The original document, the extracted fields, and the operation log have different destinations; treating all three as one is the common mistake.
Anonymization enters here too. In some cases you can work with data that does not directly identify people, which reduces exposure. In others, identification is necessary for the process to work. The choice depends on the purpose and needs legal validation, not a generic rule.
What we cannot answer generically
We will not name a controller, a DPO, or fixed retention periods here. Purpose, legal basis, retention, anonymization, deletion, and responsibilities are evaluated according to the processing performed, and the answer changes case by case. The job of engineering is to make the processing visible enough that these decisions can be made with information, not by assumption.
In summary: the four answers
- Where is it processed? In every environment the flow touches, each one mapped and verifiable.
- Does the provider train on it? Settled in the contract and configuration, checked before deployment.
- Who can access it? Roles and permissions, with a log that serves operation and audit.
- How is it deleted? Each piece gets a destination, retention and deletion, defined before the build.
These questions do not block an AI project. They organize it. In the diagnostic, this mapping is usually planned for 2 to 3 weeks, alongside prioritization of the opportunities. When there is a pilot, its build is usually estimated at 4 to 6 weeks after the scope is defined, depending on the data, accesses, and integrations involved.
To review these questions against your context, Request a diagnostic.