Data lifecycle
Stages of research data
The data lifecycle model below describes the main stages that research data typically pass through, from initial planning at the start of a project to reuse after publication.
Different organisations and disciplines may use slightly different terms or group stages differently when describing the research data lifecycle. The stages shown here provide a general framework for structuring data management activities and may overlap or be revisited during a project.
Planning occurs before data are created or acquired and establishes the foundation for effective data management throughout a project.
Activities in this stage include thinking ahead about how data will be:
- created or obtained
- organised and stored
- documented for later interpretation
- protected in line with ethical considerations and legal requirements
- managed by clearly assigned roles and responsibilities
- shared or preserved after the project ends.
Planning often involves preparing a Data Management Plan (DMP) or equivalent documentation that captures decisions about data governance, storage, and sharing.
Effective planning reduces risk, saves time later in the project, and helps ensure that data can be reused by others.
The collecting stage is where data are gathered or obtained for use in research.
This may involve:
- Gathering new data through instruments, surveys, interviews, experiments or sensors.
- Acquiring existing data from trusted sources such as data existing in responsible repositories, Trusted Research Environment or administrative systems.
- Ensuring ethical and legal requirements are followed and respected.
- Ensuring that data are captured in formats and structures that support later use.
Effective data collection includes documenting how and where data were collected, the tools used, and any relevant contextual information so that the meaning and provenance of data are clear.
Once data are collected, they typically processed and analysed to produce meaningful results.
This stage includes:
- cleaning and organising raw data
- transforming data into analysis-ready formats
- generating derived variables or indicators
- combining data from multiple sources
- conducting statistical or qualitative analysis
- documenting analytical steps and decisions.
Recording how data have been processed and analysed is critical for transparency and for enabling others to understand, reproduce, or build on your work.
Publishing and sharing makes data available to others beyond the original research team.
This can include:
- Depositing data in responsible repositories.
- Linking data to academic publications.
- Providing metadata and documentation that help others discover and understand the data.
- Choosing appropriate access conditions (for example Open, Safeguarded, or Controlled access – terminology preferred by the UK Data Service), other data service providers may use different terms) depending on legal and ethical requirements.
Sharing data responsibly enables verification of research results, supports new research questions, and increases the value and impact of the original work.
Preservation focuses on ensuring that curated research data collections remain accessible and usable beyond the end of a project. This does not mean keeping all project files. Instead, preservation focuses on retaining the data and supporting materials that have long-term value and enable others to understand, verify and reuse the research.
As a general principle, preserved data should allow another researcher to:
- Understand how the data were created and processed.
- Verify published findings.
- Reuse the data for new research, teaching or policy work.
In social science research, this usually means preserving curated data collections, rather than only selected outputs or publication-specific extracts. Key activities in this stage include:
- Selecting, contacting and negotiating with appropriate responsible repositories for preservation.
- Converting data to sustainable, well-supported formats where needed.
- Capturing and maintaining metadata that describe the data and how they were created.
- Ensuring that documentation and provenance information are preserved alongside the data.
Long-term preservation protects the research investment and enables future reuse by researchers, policymakers, educators and others. Preservation is often supported by making the data available through a responsible repository.
Reusing refers to the use of data for new research, teaching, policy analysis or other purposes after initial publication and preservation.
Data reuse depends on:
- clear and sufficient documentation
- mechanisms for discovery and access
- appropriate rights and licensing to enable lawful use
- confidence in data quality and provenance.
Reuse is most effective when data are prepared as curated collections, with documentation, metadata and access conditions that allow others to understand and use the data appropriately.
When data can be reused effectively, the original research generates additional impact and supports knowledge creation beyond its original context.