Back to Insights
GDPR ·September 25, 2026 ·7 min read

How do you anonymise documents before they enter a model?

Documents should be anonymised before AI processing by detecting personal information and removing, masking or replacing it — ideally locally, before text reaches an external model or embedding API. Pseudonymisation is often more practical than full anonymisation, but the resulting data remains personal data under the GDPR.

Documents should be anonymised before AI processing by automatically detecting personal and otherwise identifying information and then removing, masking or replacing it. This should happen as early as possible in the document pipeline: ideally locally or on-premises, before any text is sent to an external model or embedding API.

It is also essential to distinguish anonymisation from pseudonymisation. With genuine anonymisation, the information should no longer reasonably allow an individual to be identified. With pseudonymisation, identifying data is replaced, but a mechanism still exists that can reconnect the data to the original person. For many AI and RAG systems, pseudonymisation is often more practical, but the resulting information must still be treated as personal data.

What do anonymisation and pseudonymisation mean for AI and RAG?

In a traditional document management system, a document may remain within a single application. AI introduces a longer processing chain. A file may be retrieved from a document management platform, parsed, split into text chunks, converted into embeddings, stored in a vector database and later supplied to a language model together with a user's question. Personal data can be processed at several points in this chain.

Document anonymisation therefore involves more than removing a name immediately before a chatbot receives the text. The entire data flow must be considered. What information is present in the original document? Which text is sent to the embedding model? What is stored? Which chunks are subsequently provided to the language model as context? What information appears in logs, traces or error messages?

Pseudonymisation usually preserves more useful context. A name such as "John Smith" might, for example, be replaced with "PERSON_014". Every reference to that same person can then use the same identifier. The system can still understand relationships within the document without requiring the real identity.

What is the GDPR distinction between anonymisation and pseudonymisation?

The practical distinction is whether the information remains identifiable. Data is genuinely anonymised only when an individual can no longer reasonably be identified from it, including by combining different attributes. Removing a name alone is therefore usually not sufficient. A combination of job title, employer, town, date and a distinctive event may still make one particular person recognisable.

With pseudonymisation, re-identification remains possible using additional information. For example, a separate table might link "PERSON_014" to the individual's real identity. Such mapping information should be strictly separated and appropriately protected.

This distinction matters in AI projects. A dataset does not automatically become anonymous simply because names have been replaced by numbers. If an organisation can restore the link, or if the remaining attributes are sufficiently distinctive to identify people, the information should still be treated as personal data.

Which techniques can you use?

A robust anonymisation layer will normally combine several techniques.

Named entity recognition (NER): a language-processing or specialised detection model identifies entities such as personal names, organisations, locations and potentially addresses or other contextual information.

Regular expressions and pattern matching: useful for structured information such as email addresses, phone numbers, postal codes, IBANs and certain identifiers.

Redaction: sensitive information is removed completely or replaced with a marker such as [NAME REMOVED].

Tokenisation: a value is replaced by a random or controlled identifier, allowing the same person or entity to remain consistently recognisable within the dataset.

Synthetic replacement: real values are replaced with fictional values with similar characteristics. A real name can, for example, be replaced with a fictional one while preserving the sentence structure.

On-premises processing: detection and transformation are performed inside the organisation's own infrastructure before the document reaches an external API.

No individual technique works equally well for every document. A regular expression may detect an email address very reliably, but it cannot understand that "the managing director from Venlo who was dismissed last week" could refer to one specific individual. NER provides better contextual understanding but may miss patterns that simple rules can identify reliably. Combining methods is therefore usually more effective than relying on a single detector.

How do you build a practical document pipeline?

Start by determining where documents first enter your trusted processing environment. Original files should initially be handled inside a controlled environment rather than being passed directly to an external AI service.

The first technical stage is extraction. Text is retrieved from PDFs, Word documents, emails or other sources. Sensitive-information detection follows. Different detectors can operate alongside one another: pattern matching for structured identifiers and contextual detection for names, roles, places and other entities.

A policy layer then determines what should happen to each category. A telephone number might be removed completely, while a person's name might be consistently replaced by a pseudonym. These decisions are best defined for each use case rather than through one universal filtering rule.

The transformed text should then be checked again. Particularly for sensitive applications, a second detection pass can help identify information that survived the first transformation.

Only after this stage should text be chunked, submitted to an embedding model or provided to a language model as context. This prevents the original personal information from already entering embeddings, logs or external processing systems before the anonymisation layer is applied.

Logging also requires attention. A carefully designed privacy filter achieves little if the application subsequently stores the complete original input in its logs.

What mistakes are commonly made?

A common mistake is assuming that removing names means the data has been anonymised. Identification often occurs through combinations of attributes. A job title, office location, project name and date can collectively be sufficient to identify a person.

Another mistake is filtering only recognisable patterns. Not every personal data element looks like a telephone number or email address. Free text requires contextual detection as well.

Embeddings should not automatically be regarded as anonymous either. An embedding is a numerical representation of information, but that does not mean all privacy risks disappear. If embeddings are derived from personal data, risks related to storage, access, linkage and potential information leakage still need to be assessed.

Problems also arise when anonymisation is performed only immediately before the final model request. By that point, the original information may already have passed through OCR systems, parsers, embedding services, observability tools or application logs.

Finally, excessive anonymisation can make an AI system ineffective. If all names, dates, organisations and relationships are removed, a RAG system may lose precisely the context required to answer a question correctly.

When is full anonymisation not feasible?

Some applications depend fundamentally on personal information. A legal case file, employee record or customer history may lose its meaning if every identifiable detail is removed.

In such situations, it is better not to describe the dataset as anonymous. Instead, consider pseudonymisation, strict access controls, data minimisation or an architecture in which sensitive processing remains entirely within the organisation's own infrastructure.

On-premises AI can be valuable in this scenario. Document extraction, embeddings, retrieval and potentially the language model itself can run within a controlled environment. Personal information therefore does not necessarily have to be sent to an external AI provider.

A hybrid architecture is another option. Sensitive processing and retrieval remain local, while only tightly restricted or pseudonymised context is sent to an external model. The appropriate approach depends on the type of documents, the purpose of the system and how much context is required to generate reliable answers.

Conclusion

Anonymising documents before they enter an AI model is not a single filtering operation but part of the entire data pipeline. Detect personal information before documents are chunked, embedded or transmitted to external services, and combine pattern-based detection with contextual analysis.

It is equally important to distinguish clearly between anonymisation and pseudonymisation. If information remains identifiable, it should still be treated technically and organisationally as personal data. When full anonymisation removes too much useful context, pseudonymisation, data minimisation and on-premises processing are often more realistic alternatives.

Preparing documents safely for AI

Want to know how to make your documents available to an AI system without privacy risk?

We build document pipelines that detect and process personal data before text reaches an external model — fully on-premises where needed. The legal assessment of your specific dataset should then be reviewed by a lawyer.