Data anonymisation step-by-step
A three-stage framework for assessing and applying privacy controls
Anonymisation is a structured, risk-based process. While specific techniques differ depending on the structure and content of the data, the underlying approach remains consistent.
A simplified high-level approach for preparing data for sharing consists of three core stages.
Begin by identifying information that could directly or indirectly identify participants.
This includes:
- Direct identifiers: information that identifies an individual on its own (e.g. names, addresses, ID numbers, identifiable faces or voices).
- Indirect identifiers: information that may not identify someone alone but could do so when combined with other data (e.g. age, occupation, fine geography, rare characteristics, distinctive life events, network position, precise coordinates).
Identification risk depends not only on what appears in the data, but also on:
- The availability of external information that could be linked, where permissible.
- The size and uniqueness of the data, including the level of geographic or temporal precision.
- The sensitivity of the information and the intended access conditions.
Key questions:
- Could an individual be identified using the information in the data?
- Could combinations of information increase identification risk?
- Could information disclosed cause harm to participants or third parties?
- Does the level of risk align with the planned access conditions?
After assessing risk, apply techniques appropriate to the data type and intended sharing conditions.
- Remove or manage direct identifiers
Direct identifiers should normally be removed unless there is a clear legal and ethical justification for retaining them. Where direct identifiers are retained (for example with explicit permission), access conditions should be carefully considered throughout the research data lifecycle.
- Reduce risk from indirect identifiers
Depending on the data structure, this may involve:
- Aggregation or banding (e.g. age ranges, income bands).
- Generalisation (e.g. specific town to region; detailed job title to broader category).
- Recoding or category reduction.
- Top/bottom coding of extreme values.
- Perturbative methods for microdata such as noise addition or data swapping.
- Selective redaction or contextual editing.
- Reducing spatial precision or masking coordinates.
- Reducing attribute detail or aggregating network structures.
- Editing or masking visual/audio identifiers where appropriate.
Techniques should be selected to balance privacy protection with analytical utility. Over-anonymisation may significantly reduce research value; under-anonymisation may leave unacceptable disclosure risk.
In some cases, more restrictive access arrangements may be more appropriate than extensive data modification.
Key questions:
- Does the anonymisation meaningfully reduce identification risk?
- Has analytical utility been preserved as far as possible?
- Are the techniques proportionate to the sensitivity and intended access level?
Anonymisation should not be assumed to be complete once techniques have been applied. A final review is required to assess whether individuals are identifiable, taking into account all means reasonably likely to be used.
Conduct a systematic review to ensure:
- Anonymisation has been applied consistently across all files.
- No direct identifiers remain.
- Combinations of information do not create unmanaged identification risk.
- Edits have not introduced errors or inconsistencies.
- Contextual narrative detail does not enable “jigsaw identification”.
It is good practice to:
- Maintain an anonymisation log recording all changes.
- Confirm that licensing and access conditions align with residual risk.
Anonymisation does not operate independently of governance. The level of data modification should be proportionate to the access arrangements. For example:
- Data released under open reuse licences require a higher level of anonymisation.
- Data shared under safeguarded or controlled access conditions may permit greater analytical detail, provided appropriate contractual and technical protections are in place.
The key question is whether individuals are identifiable by means reasonably likely to be used in the circumstances, including the legal and technical controls governing access and reuse.
It may be helpful to apply the “motivated intruder” perspective described in guidance from the Information Commissioner’s Office. This involves asking whether a reasonably competent person, with no specialist skills but access to available information, could identify individuals in the data.
If individuals remain identifiable, the data should not be treated as anonymised and must instead be managed as personal data under appropriate access controls and legal safeguards.