Considerations for common indirect identifiers

Data producer support home page

Managing combined variable risks and re-identification threats

In addition to direct identifiers, careful consideration must also be given to indirect identifiers. These are pieces of information that do not identify an individual on their own but may do so when combined with other.

Examples include age, date of birth, gender, ethnicity, postcode, occupation, and geographic location. Although each variable may appear non-identifiable in isolation, their combination can substantially increase the risk of re-identification, particularly when linked with external data.

Consequently, organisations and researchers should assess the cumulative disclosure risk posed by indirect identifiers and apply appropriate safeguards, such as data aggregation, generalisation, suppression, or statistical disclosure control techniques, to reduce the likelihood of re-identification while preserving the utility of the data for analysis.

Age can be presented in a number of formats:

  • full date of birth
  • day of birth
  • month of birth
  • year of birth
  • age (single year)
  • age (in months)
  • age (banded).

Full date of birth should be only made available under very struct conditions as can be used to link to other data sources.

Where single age is available, you might find that small numbers of very young or old respondents are present in the data. This can be problematic as, when used in conjunction with other indirect identifiers, it can uniquely identify someone or be used to link to external data sources.

Top-coding or bottom-coding might be required to group small numbers of very young or old respondents, or anonymisation of other information in the file may make age/date of birth information less disclosive.

Single year of age is often considered one of the most important demographic information, therefore one has to keep in mind that banding and top-coding age can render some numeric analyses impossible (including mean age, for example), hence this needs to be considered when deciding for your course of action. Anonymisation of other information in the file may make age/date of birth information less disclosive and should be considered in conjunction with any aggregation.

Education information can be problematic when they include detailed information which could be unique or only applicable to a small number of the population.

This type of information, when used in combination with other, could lead to the identity of a participant being disclosed. School, college and university names may be disclosive, if analysed in conjunction with course titles or other information, especially where the data cover more than one course with accompanying institutional information.

Education information should also ideally be categorised using a coding frame. Where educational information is detailed, the usefulness of information for secondary analysis should be considered. Educational attainment or qualification level might be more useful than specific institutional information. For example the ONS Census 2021 ‘Highest level of qualification’  codes education level into 8 categories. Removing some level of detail might considerably reduce the risk of identifying an individual.

Education information can be problematic when they include detailed information which could be unique or only applicable to a small number of the population.

This type of information, when used in combination with other, could lead to the identity of a participant being disclosed. School, college and university names may be disclosive, if analysed in conjunction with course titles or other information, especially where the data cover more than one course with accompanying institutional information.

Education information should also ideally be categorised using a coding frame. Where educational information is detailed, the usefulness of information for secondary analysis should be considered. Educational attainment or qualification level might be more useful than specific institutional information. For example the ONS Census 2021 ‘Highest level of qualification’  codes education level into 8 categories. Removing some level of detail might considerably reduce the risk of identifying an individual.

Detailed breakdowns of ethnicity, national identity or religious affiliation can be potentially problematic when used in combination with other information.

Using a standard coding frame is helpful and the ONS provides some guidance on this complex area in Measuring equality: A guide for the collection and classification of ethnic group, national identity and religion data in the UK.

Example of geographic information

Geographical or spatial information present in the data should be considered carefully. Detailed, low level geographic information will increase the ability to potentially identity data subjects. If the geographic information is at a detailed level, it should be ensured that frequencies contain sufficient numbers to avoid identification in combination with other information. For example, one case in a particular small area may mean that the respondent could be identified using a combination of other information.

The research topic of the study must also be borne in mind while checking geographical information. For example in a politics-related study, low-level political geographies may be necessary for analysis, such as Ward or Parliamentary Constituency. Also, Scottish/Welsh/Northern Ireland surveys may include LA or similar level areas, as aggregation to Region/GOR level would mean that only country-level analysis would be possible (Wales, Scotland and NI are each one Region/GOR area).

Access level and geo-references are generally assessed together.

Health information may be disclosive, if for example they contain unusual combinations of conditions or dates of hospital visits or named individual treatment centres. Aggregation, coding or removal of the information affected may be the best solution.

Large households may mean that they are more unique and identifiable when their size is analysed alongside other information Where data are household based, the age and sex structure of the household members and relationships of individuals to each other can increase the risk of identification.

The ONS recommend top-coding household size at 10.

Income and other financial information could be disclosive if they have unique outlying values. Isolated cases of a very high or low income or other financial sums – for example, a large lottery win recorded as unearned income during the time period the data cover – may present a disclosure risk if analysed alongside other information and other publicly-available information (some lottery wins are well-publicised).

Financial information needs to be checked and for example for tabular data it may be advisable to recommend top-/bottom-coding to aggregate unique cases into a group. Banding and top-coding can hamper some statistical analyses where individual numbers are required, so data usability should always be considered.

Example of life events

Similar to date of birth information, exact dates of life events (hospital treatments, court cases, marriages, deaths, etc.) may increase the risk of potential identification. Reducing exact dates to month and year might remove enough detail. Again it depends how crucial the information is to the data usability so increased access conditions might be more appropriate.

Data should always be checked for sensitive information as some information may be potentially harmful if a respondent is identified, such as an unusual or sensitive health condition or status, or details of illegal behaviour.

Some sensitive data will be Special Category data under UK GDPR and this information should be carefully considered alongside all other indirect identifiers in the data in case a disclosure risk is identified. They may need to be edited, but if doing so would harm the usability of the data, access restriction should also be considered.