This site uses necessary cookies

Some of these cookies are essential. Strictly necessary cookies enable core functionality, without which, the website cannot function properly. For more detailed information please see our Cookie Policy.


Website stats

We use Matomo Analytics to understand how our website is used and to improve your experience. This tool gathers limited information about the device you use to access the UK Data Service website. To learn more, please see our Privacy Policy.

Glossary

Data archives

A data archive is a centralised database system that collects, manages, and stores datasets for later use. Similar to a data repository.

Data generalisation (data anonymisation)

The process of replacing detailed or specific data with broader categories (for example, replacing an exact age with an age range). This helps protect privacy by reducing the level of detail that could identify individuals.

Data licensing

Data licensing is a legal arrangement between the creator of the data and the end-user specifying what users can do with the data.

Source: How to FAIR.

Data linkage

Data linkage is the process of joining together records from different sources that pertain to the same entity.

Source: ONSUnderstanding Society.

Data manipulation

Data manipulation is the process of arranging and organising data to make it easier to use, analyse and interpret.

Data masking (data anonymisation)

The process of protecting sensitive data by replacing it with altered or fictitious values. The masked data preserves the format and structure of the original data but cannot be easily traced back to the real information.

Data mining

Data mining is defined as the process of extracting useful information from large data sets through the use of any relevant data analysis techniques developed to help people make better decisions.

Source: SAGE Research Methods .

Data repository

A data repository is a centralised database system that collects, manages, and stores datasets for later use, similar to a data archive.

Data science

A field that draws on statistics, programming, and subject area knowledge to collect, process, and analyse data in order to answer research questions and draw evidence-based conclusions.

Data suppression (data anonymisation)

The process of removing or hiding specific data values in a dataset to reduce the risk of identifying individuals. Suppressed data is typically replaced with blanks or placeholders to prevent sensitive information from being disclosed.

Dataset

Any computer file (or set of files) which is organised under a single title and is capable of being described as a coherent unit.

Derived variable

variable that is created from one or more already existing variables by following some sort of calculation or other data processing technique. For example, each respondent’s estimated annual income from savings and investments could be derived from several reported income variables.

Descriptive statistic

Descriptive statistics are those that describe data. Examples include means, medians, variances, standard deviations, correlation coefficients, etc.

Source: SAGE Research Methods.

Disclosure risk (data anonymisation)

The likelihood that an individual or entity can be identified in a dataset that is intended to be anonymised. This can occur through direct identifiers (such as names) or indirect identifiers (such as combinations of demographic or sensitive attributes). Managing disclosure risk is important for protecting privacy and meeting data protection requirements.

Documentation

Accompanying files that enable users to understand a dataset, exactly how the research was carried out and what the data mean. Usually consisting of data-level documentation i.e. about individual databases or data files and study-level documentation i.e. high-level information on the research context and design, the data collection methods used, any data preparations and manipulations, plus summaries of findings based on the data.

DOI (Digital Object Identifier)

A unique, persistent alpha-numeric string used to identify content like scholarly articles and datasets. It offers a permanent, reliable link to the content’s location online. Unlike URLs, which may break or become outdated, a DOI ensures long-term access, functioning like a “digital passport” or barcode for research.