Third-party rights and restrictions

Data producer support home page

Managing licensing limits, redistribution rules, and secondary data provenance

Many social science research projects rely on data obtained from external sources, including administrative and government records, commercial platforms, licensed databases, web-based collections and previously published research data. While these data may be used legitimately for research purposes, they often come with specific legal, contractual and licensing conditions that affect how they can be shared or redistributed.

Redistribution vs. project-only access

It is important to recognise that permission to access and analyse data does not automatically include permission to publish or redistribute the data. For example, a licence may allow researchers to use data within a project or secure environment but explicitly prohibit onward sharing of the original files.

Before sharing derived or combined datasets, data producers should review the original conditions under which the data were obtained. This includes checking reuse licence terms, data use agreements and any contractual restrictions on redistribution or onward use. These requirements should be considered early in the project and revisited when preparing data for deposit.

Alternative approaches when raw data cannot be shared

Where third-party data cannot be redistributed, data producers should plan alternative approaches that still support transparency, reproducibility and reuse. Depending on the context, this may include sharing derived or aggregated data that do not expose restricted content, creating metadata-only records that describe unavailable source data, depositing code, or, if the licence terms permit, generating synthetic or simulated data that reflect key characteristics without revealing protected information.

Any reuse restrictions or limitations must be clearly documented in repository metadata and user documentation. This ensures that future users understand what data are available, what cannot be accessed, and under what conditions reuse is permitted.

Documenting secondary data sources

Data producers are encouraged to prepare a variable information log that records the provenance and reuse conditions associated with each external data source. This supports responsible data sharing and helps repositories assess compliance with licensing and contractual requirements.

A variable information log typically records:

  • variable names
  • original data sources
  • how the data were collected or obtained
  • brief descriptions of content
  • licensing, contractual or reuse restrictions

This information enables both repositories and secondary users to understand how third-party data were incorporated into the dataset and what limitations apply to reuse.

A sample variable information log template is available here to support consistent documentation of secondary data sources.

The same information log approach can also be applied to other types of data where provenance and reuse conditions need to be clearly documented. For example, this may include administrative data, commercial data, linked datasets, transcript data or image data. In these cases, additional considerations may include documenting data linkage methods, permissions obtained and data processing or transformation steps.

Using an information log supports transparency, responsible data sharing, and helps repositories and secondary users understand how the data were created and what limitations apply to reuse.

Where data cannot be shared, data producers are encouraged, depending on the funder mandated, to prepare a catalogue only record (metadata record) that records the origin (or provenance) of the data and reuse conditions associated with each external data source. Information about metadata records is available on the prepare your data collection for deposit page.