Data archives
A data archive is a centralised database system that collects, manages, and stores datasets for later use. Similar to a data repository.
Some of these cookies are essential. Strictly necessary cookies enable core functionality, without which, the website cannot function properly. For more detailed information please see our Cookie Policy.
We use Matomo Analytics to understand how our website is used and to improve your experience. This tool gathers limited information about the device you use to access the UK Data Service website. To learn more, please see our Privacy Policy.
A data archive is a centralised database system that collects, manages, and stores datasets for later use. Similar to a data repository.
The process of replacing detailed or specific data with broader categories (for example, replacing an exact age with an age range). This helps protect privacy by reducing the level of detail that could identify individuals.
Data licensing is a legal arrangement between the creator of the data and the end-user specifying what users can do with the data.
Source: How to FAIR.
Data linkage is the process of joining together records from different sources that pertain to the same entity.
Source: ONS; Understanding Society.
Data manipulation is the process of arranging and organising data to make it easier to use, analyse and interpret.
The process of protecting sensitive data by replacing it with altered or fictitious values. The masked data preserves the format and structure of the original data but cannot be easily traced back to the real information.
Data mining is defined as the process of extracting useful information from large data sets through the use of any relevant data analysis techniques developed to help people make better decisions.
Source: SAGE Research Methods .
A data repository is a centralised database system that collects, manages, and stores datasets for later use, similar to a data archive.
A field that draws on statistics, programming, and subject area knowledge to collect, process, and analyse data in order to answer research questions and draw evidence-based conclusions.
The process of removing or hiding specific data values in a dataset to reduce the risk of identifying individuals. Suppressed data is typically replaced with blanks or placeholders to prevent sensitive information from being disclosed.
Any computer file (or set of files) which is organised under a single title and is capable of being described as a coherent unit.
A variable that is created from one or more already existing variables by following some sort of calculation or other data processing technique. For example, each respondent’s estimated annual income from savings and investments could be derived from several reported income variables.
Descriptive statistics are those that describe data. Examples include means, medians, variances, standard deviations, correlation coefficients, etc.
Source: SAGE Research Methods.
The likelihood that an individual or entity can be identified in a dataset that is intended to be anonymised. This can occur through direct identifiers (such as names) or indirect identifiers (such as combinations of demographic or sensitive attributes). Managing disclosure risk is important for protecting privacy and meeting data protection requirements.
Accompanying files that enable users to understand a dataset, exactly how the research was carried out and what the data mean. Usually consisting of data-level documentation i.e. about individual databases or data files and study-level documentation i.e. high-level information on the research context and design, the data collection methods used, any data preparations and manipulations, plus summaries of findings based on the data.
A unique, persistent alpha-numeric string used to identify content like scholarly articles and datasets. It offers a permanent, reliable link to the content’s location online. Unlike URLs, which may break or become outdated, a DOI ensures long-term access, functioning like a “digital passport” or barcode for research.