Anonymisation for text data
Managing narrative detail and semi-automated tools
Text data, including interview transcripts, focus groups, diaries and fieldnotes, often contain rich, contextual accounts of people’s lives. This richness is central to qualitative research, but it also increases the likelihood that individuals or third parties could be identified through narrative detail.
Identification in text data often occurs through context, specificity and coherence across a story. A name may be removed, yet a distinctive life event, job role or geographic reference could still allow someone to infer identity. Effective anonymisation of text data therefore requires careful, contextual judgement.
Direct identifiers in text data are usually explicit and should normally be replaced by pseudonyms or removed prior to sharing.
These may include:
- participant names
- names of family members, colleagues or organisations
- specific employers or institutions
- full addresses or highly specific locations
- contact details.
Replacement should be applied consistently across transcripts and related materials. Pseudonyms or meaningful descriptive tags (e.g. [local authority employee], [colleague 1]) may help preserve narrative coherence.
In qualitative data, disclosure risk frequently arises not from a single identifier but from combinations of contextual detail. This is sometimes referred to as “jigsaw identification”.
Examples include:
- a rare occupation in a small community
- a distinctive career trajectory
- detailed timelines
- references to highly specific locations
- unique personal experiences
- information about identifiable third parties.
Removing names alone is rarely sufficient. Data producers should consider how narrative detail might allow identity to be inferred when combined with external knowledge.
Pseudonymisation and consistent replacement
Names and identifiable references should be replaced consistently throughout the dataset. Care should be taken when using automated search functions to avoid unintended changes or missed variations.
Generalisation
Generalisation reduces specificity while retaining analytical meaning.
For example:
“I moved to Paris and lived on Rue de Rivoli near the Louvre…”
May be generalised to:
“I lived abroad for several years.”
The aim is to remove unnecessary precision without erasing the substantive point being made.
Selective redaction
In some cases, redaction may be necessary where disclosure could cause harm. This may include:
- sensitive health information, particularly where this is declared about someone other than the participant themselves
- potentially defamatory statements
- references to ongoing legal proceedings.
Where redaction is applied, it is good practice to indicate this transparently (e.g. “[section removed]”) and to document decisions in an anonymisation log.
Over-anonymisation can remove too much contextual detail, reducing interpretive value and limiting meaningful reuse. Under-anonymisation may leave sufficient detail for identification. A balanced approach replaces disclosive detail with meaningful descriptors rather than removing content entirely.
The following examples show what ‘over’ and ‘under’ anonymisation might look like. In the example of ‘over’ anonymisation, too much detail has been taken out without replacement of meaningful descriptors. In the ‘under’ anonymisation, too much detail which can lead to a disclosure has been left in the data.
Original: So my first workplace was Arronal, which was about 20 minutes from my home in Norwich. My best colleagues from day one were Andy, Julie and Louise and in fact, I am still very good friends with Julie to this day. She lives in the same parish still with her husband Owen and their son Ryan.
Example A, ‘over’ anonymisation: So my first workplace was X, which was about X minutes from my home in X. My best colleagues from day one were X, X and X and in fact, I am still very good friends with X to this day. X lives in the same parish still with her husband X and their X X.
Example B, ‘under’ anonymisation: So my first workplace was [name], which was about 20 minutes from my home in Norwich. My best colleagues from day one were Andy, Julie and Louise and in fact, I am still very good friends with Julie to this day. She lives in the same parish still with her husband Owen and their son Ryan.
Example C, ‘balanced’ anonymisation: So my first workplace was at [a local company], which was about 20 minutes from my home in [city in Eastern England]. My best colleagues from day one were [colleague 1], [colleague 2], and [colleague 3], and I am still very good friends with [colleague 2] to this day. She lives in the same parish with her husband and their son.
It is good practice to:
- Maintain an anonymisation log recording changes.
- Clearly mark replacements in transcripts.
- Ensure metadata describe what has been modified.
The text anonymisation helper tool (Zip) can help to find disclosive information to remove or mask in text data files. The tool does not anonymise or make changes to data but uses MS Word macros to find and highlight numbers and words starting with capital letters in text. Numbers and capitalised words are often disclosive, e.g. as names, companies, birth dates, addresses, educational institutions and countries.
A range of tools are available to support the automated anonymisation of text data, but they differ in their scope and level of flexibility.
QualiAnon is an open-source, semi-automated tool designed to support the anonymisation and pseudonymisation of qualitative text data, particularly interview transcripts. It allows users to manually control how sensitive information is identified and replaced, supporting data protection compliance while preserving meaningful research context.
Tools such as NLM Scrubber use predefined rules to identify and remove common types of sensitive information and are often applied in highly-structured domains such as health data.
Tools such as Microsoft Presidio, combine rule-based and machine learning approaches to provide more flexible and actively developed solutions. Regardless of the tool used, automated anonymisation should always be supplemented with careful review to ensure accuracy and minimise the risk of missed identifiers or unnecessary data loss.
When using any tools for anonymising or processing textual data it is crucial to ensure they are deployed locally and not reliant on external servers or unrestricted cloud services, which could inadvertently expose sensitive data to third parties. Data producers must take responsibility to verify that any software used for data processing is configured correctly to prevent data from being uploaded to unknown or insecure entities.