Data-level documentation
Explaining internal structure, coding schemes, and file relationships
Data-level documentation explains how individual files are structured and how their contents should be interpreted. It makes explicit how information is organised within files, how elements are identified, and how different components relate to one another.
All research data, regardless of format, have structure. That structure may take the form of variables in a table, speakers in a transcript, features in a spatial layer, nodes and edges in a network, or media files linked to transcripts. Data-level documentation clarifies this structure so that it can be understood independently of the original research team.
The specific form of data-level documentation depends on the structure of the collection. However, across all data types, documentation should clarify:
- What each file contains.
- How items, records or features are identified.
- How internal elements are organised.
- How files relate to one another.
- What transformations or modifications have been applied.
- How missing or incomplete material is represented.
The sections below outline how these principles apply to different data structures.
In tabular data, meaning is often embedded in labels, classifications and conventions rather than in the raw values themselves. A numeric value has little meaning unless its coding scheme, units of measurement and applicability are clearly defined. Distinctions such as “not applicable”, “not known” or “not asked” can materially affect analysis, yet may be represented using similar numeric codes.
Data-level documentation for tabular data therefore focuses on clarity at the variable level. It should make explicit how each variable is defined, how it is coded and how it relates to the wider dataset.
Documentation should describe:
- Variable names and descriptive labels.
- Data types.
- Units of measurement.
- Value labels and coding schemes.
- Definitions of missing data and distinctions between types of non-response.
- Applicability or universe information (e.g. filtered questions or subpopulations).
- Weighting variables, where relevant.
- Derived or constructed variables and how they were created.
Where derivations are simple, documentation within the variable label may be sufficient. More complex derivations should be described in sufficient detail to allow the logic to be understood or reproduced. Retaining syntax or command files supports transparency.
For longitudinal datasets, wave identifiers must be clearly defined and any changes in question wording, coding schemes, measurement or sampling across waves should be documented to preserve comparability.
Although some statistical software (e.g. SPSS, Stata) often allows this information to be embedded within the file, embedded metadata should be reviewed for completeness. Supplementary documentation, such as a structured codebook or data dictionary, may be required to ensure clarity.
Text data, such as interview transcripts, fieldnotes, and diaries, require documentation that preserves context, structure and internal coherence. Without clear documentation, secondary users cannot reliably interpret what the material represents or how it was produced.
Each text file should contain sufficient contextual information to stand on its own. This is typically provided through a header or summary section at the beginning of the file. At minimum, this should include:
- a unique identifier
- date of data collection
- context or setting (where appropriate).
Consistency is essential. Identifiers used in transcripts should match those used in related materials.
Where transcripts are involved, clarity of speaker attribution is critical. Speaker tags should be applied consistently throughout the file. If pseudonyms or role-based descriptors are used, they should be used systematically across the collection. Examples can be seen in our model transcription template.
Transcription conventions must be documented. This includes notation for pauses, emphasis, overlapping speech, interruptions or non-verbal expressions. Without an explanation of these conventions, secondary users may misinterpret tone, meaning or analytical intent.
For collections containing multiple interviews, documents or items, a structured data list should accompany the deposit. This list enables users to identify and navigate materials efficiently.
A data list should:
- Assign a unique identifier to each item.
- Provide key contextual descriptors (e.g. interview ID, date, location, topic area).
- Indicate the existence of associated files (e.g. transcript, audio recording, fieldnotes).
- Clearly mark any missing or partial materials.
This structured overview allows users to understand the scope and composition of the collection before engaging with individual files.
You can use our data listing template to develop a data list. You can also see examples of data lists in any of our collections, such as Inequality and the Media in Brexit-Covid-19-Britain, 2020-2021, Growing Risk? The Potential Impact of Plant Disease on Land Use and the United Kingdom Rural Economy, 2007-2011, Black Immigrants to Britain, 1890-1975 or Angels in Marble.
For image, audio and video materials, documentation must explain both the content and the technical characteristics of the files. Meaning may be conveyed through visual composition, sound quality, sequencing or editing decisions. Without contextual and technical information, reuse becomes limited and potentially misleading.
Each file should be identifiable and traceable within the collection. Documentation should therefore ensure that:
- Each file has a unique identifier.
- File names are meaningful and consistent.
- The relationship between different files (e.g. recording and transcript, raw and edited versions) is clear.
For audio-visual materials, the circumstances of recording are often essential to interpretation. Documentation should describe:
- Date of recording.
- General location or setting (as appropriate and proportionate).
- Purpose of the recording.
- Role of participants (e.g. interviewee, observer, facilitator).
- Any relevant situational factors affecting the recording.
Where recordings form part of a sequence (e.g. observational sessions or longitudinal interviews), the order and relationship between files should be documented.
Technical characteristics influence usability and preservation. Documentation should include:
- File format (e.g. WAV, MP4, TIFF).
- Duration (for audio and video).
- Resolution or frame rate (for images and video, where relevant).
- File size.
- Any compression applied.
- Hardware or software used for capture or editing, where this affects quality or interpretation.
Where technical metadata are embedded automatically by recording devices, this should still be verified for accuracy.
Where transcripts, subtitles, captions or logs exist, they should be included and clearly linked to the relevant recording, along with any file relationships (such as groupings or sequencing across the collection). If time-stamping conventions are used, these should be explained.
For observational video or complex audio recordings, shot logs or content summaries may assist navigation and reuse.
[Add data list template for audio-visual data]
For spatial data collections, accurate interpretation depends on clear documentation of how location is represented and structured. Spatial meaning is inseparable from coordinate systems, scale and geometry type. Without this information, spatial files cannot be reliably mapped, compared or analysed.
Documentation at the data level should enable a user to correctly load, map and interpret the data without ambiguity.
Data-level documentation for spatial data should make explicit:
- Coordinate reference system (CRS)
This includes:
- projection
- datum
- units of measurement.
If the CRS is not documented, spatial alignment with other datasets may be incorrect.
If the data have been reprojected, the current CRS should be clearly identified.
- Geometry and resolution
Documentation should describe:
- geometry type (point, line, polygon, raster)
- spatial resolution (e.g. raster cell size)
- scale or level of aggregation
- spatial extent (geographic coverage).
For raster data, cell size and grid structure should be specified.
- Attribute structure
Where spatial data include attribute tables, each variable should be described clearly, including:
- variable name and label
- units of measurement
- coding schemes
- definitions of any derived spatial measures (e.g. calculated area, distance, density).
Users must be able to understand how attribute values relate to spatial features.
For spatial data collections, accurate interpretation depends on clear documentation of how location is represented and structured. Spatial meaning is inseparable from coordinate systems, scale and geometry type. Without this information, spatial files cannot be reliably mapped, compared or analysed.
Documentation at the data level should enable a user to correctly load, map and interpret the data without ambiguity.
Data-level documentation for spatial data should make explicit:
- Coordinate reference system (CRS)
This includes:
- projection
- datum
- units of measurement.
If the CRS is not documented, spatial alignment with other datasets may be incorrect.
If the data have been reprojected, the current CRS should be clearly identified.
- Geometry and resolution
Documentation should describe:
- geometry type (point, line, polygon, raster)
- spatial resolution (e.g. raster cell size)
- scale or level of aggregation
- spatial extent (geographic coverage).
For raster data, cell size and grid structure should be specified.
- Attribute structure
Where spatial data include attribute tables, each variable should be described clearly, including:
- variable name and label
- units of measurement
- coding schemes
- definitions of any derived spatial measures (e.g. calculated area, distance, density).
Users must be able to understand how attribute values relate to spatial features.