Preparing data for sharing and reuse

Data producer support home page

Preparing data for sharing and reuse is the stage where project working files are transformed into stable, well-documented research collections that can be preserved, discovered and accessed by others.

This process builds on earlier data management activities and focuses on technical preparation, preservation decisions, repository deposit, access arrangements and reuse conditions. For social science research, this usually means preparing curated data collections rather than only sharing small extracts or publication-specific subsets.

Preparation for sharing is often iterative. As projects progress and publication plans develop, data producers may refine documentation or create additional derived outputs. Starting preparation early helps reduce delays at deposit, supports compliance with funder requirements and improves the long-term value of research data.

Data cleaning for sharing differs from analytical data cleaning. The goal is not to optimise data for modelling or interpretation, but to improve clarity and usability for external users.

This may include:

  • resolving ambiguous labels or codes
  • removing duplicate records
  • standardising date formats and text encodings
  • checking consistency across linked files.

Any changes made at this stage should be documented so that users can understand how release versions were created.