GlossaryFloor 4 · The Organisationa mosaic: AI is only one tile among manyFloor 4 · The Organisation
personal data
No. 031 · v2026-08FR: donnée personnellePersonal data is any information that makes it possible to trace back to a person, directly or by cross-referencing several elements: it is not only the name on the letterbox, it is also the floor they live on, the timetable and the car, which together designate someone just as surely.
What it is not
Personal data is not the same thing as sensitive or confidential data: a work address is personal data without being in any way secret, whereas a trade secret is not personal data. Nor is it a category reserved for civil status, since a technical identifier, a voice recording, a photograph or a browsing history fall under it just as much. Finally, replacing a name with a code does not take you out of scope: as long as a reasonable means of restoring the link remains, the data stays personal.
In depth
The decisive criterion
The decisive criterion is identifiability, and it is assessed in the light of the means that can reasonably be brought to bear to trace back to the person. This gives the notion a relative character that is underestimated: one and the same file can be personal for whoever holds the correspondence table and not be personal for a recipient who has no way of reaching it. Indirect identification is the most frequent source of error, because innocuous attributes taken separately become identifying once combined, especially when the population concerned is small. A municipality, a job title and an age bracket are sometimes enough to designate one person and one only, without any name appearing.
The points of exposure
In an AI system, the question arises in more places than the original database. What a person writes in a prompt, what a system retains from one conversation to the next, the documents indexed so as to be retrieved, the technical logs kept in order to understand an incident and the answers produced all form places where the data circulates. Those logs are in fact the most common blind spot, because they are taken for mere machine traces when they attach dated actions to named people. To this is added a question specific to training: when a model has been built on personal data, knowing whether its parameters still carry any has no general answer and is examined case by case. The practical consequence is that personal data has a life cycle inside the system, and that a useful inventory follows that cycle rather than the list of applications.
The stricter regime
Some categories receive stricter treatment, because they expose more: health, opinions, beliefs, origin, orientation, trade union membership, biometric data used to identify. AI adds a difficulty of its own, that of inference: a system can deduce an attribute of this nature from information that does not fall under it, which brings the processing into the strictest regime without anyone having collected anything of the sort. The second trap is claiming anonymity too quickly, often asserted after a simple deletion of names when cross-referencing remains possible. The third is reasoning by tool: an organisation that classifies its rules by software finds itself helpless as soon as a piece of software changes, whereas a classification by category of data outlives the tools.
Relations where the neighbours live
Check 3 questions · click your answer
Level 1 · Recognise
A file contains only customer identifiers and amounts, with no name and no address. Is it personal data?
Level 2 · Distinguish
What is the difference between pseudonymisation and anonymisation?
Level 2 · Distinguish
A system infers a probable state of health from purchasing habits, although no health data has been collected. How should this situation be read?
Try it 1 practice
Concrete things to try where this term comes up, in ten minutes.
Lexigraph, "Personal data", v2026-08, https://www.lexigraph.org/en/personal-data/, CC BY 4.0.