Collections created from secondary sources
Managing provenance, licensing, and onward sharing for complied data
Many projects incorporate variables, datasets or materials obtained from external sources. These may include administrative records, official statistics, commercial datasets, web-scraped material, historical archives or previously published data.
When external data are incorporated into a new collection, or a collection consist of secondary sources only, the research team becomes responsible not only for how those data are used, but also for how their origin, transformation and reuse conditions are communicated.
A data collection that includes secondary material must make clear what originates from the research team and what originates elsewhere.
Users should be able to understand:
- The original source of each externally derived dataset or variable.
- The version or date accessed.
- The method of acquisition (e.g. download, agreement, API, archive access).
- Any cleaning, transformation or harmonisation applied.
- How externally sourced data were integrated into the primary dataset.
Without this clarity, secondary users cannot properly assess data quality, comparability or limitations. Transparency about provenance is not only a matter of good documentation; it is essential for reproducibility and scholarly integrity.
Secondary data may carry specific legal or contractual conditions. These may include:
- data sharing agreements
- copyright restrictions
- open licences with attribution requirements
- restrictions on redistribution or derivative works.
Data that appear publicly accessible are not automatically free for redistribution. For example, website data may be viewable but subject to terms restricting scraping, processing or onward dissemination.
Before preparing a dataset for sharing, data producers should confirm:
- Whether redistribution of the secondary data is permitted.
- Whether derivatives may be shared.
- Whether attribution statements are required.
- Whether downstream users must comply with specific licence terms.
Where redistribution is not permitted, alternative approaches may be required, such as:
- Providing derived indicators rather than raw variables.
- Supplying well documented code which directs users to the original source.
- Catalogue only records (metadata only records).
Clear documentation of reuse conditions protects both the research team and future users.
Where records from different sources have been linked, additional documentation is required. Secondary users should be able to understand:
- The basis for linkage (e.g. deterministic match, probabilistic linkage, unique identifiers).
- Any matching thresholds applied.
- Potential linkage errors or exclusions.
- The impact of linkage decisions on coverage or representativeness.
Linkage decisions can affect analytical conclusions. Transparent documentation enables users to interpret linked datasets responsibly.
Secondary data are often transformed into derived indicators, indices or harmonised classifications. Where this occurs, documentation should explain:
- the rationale for derivation
- the method used
- any assumptions applied
- the implications for interpretation.
Derived variables should not obscure the fact that they originate from externally sourced data.
Where the original sources require citation or attribution, this should be clearly documented within the collection-level documentation and reflected in recommended citation guidance.
Secondary data should be treated as scholarly contributions in their own right.