Data quality and integrity
Data quality and integrity
Maintaining data quality and integrity during active research is essential for producing reliable results and enabling future reuse. Good quality management helps ensure that data accurately represent what they are intended to measure and that files remain trustworthy throughout the project (data?) lifecycle.
Quality considerations begin at the point where data are created or obtained. During data collection and acquisition, data producers should ensure that records accurately reflect observed values, responses or system outputs. This may involve applying consistent collection protocols, calibrating instruments or sensors, validating incoming data feeds and carefully documenting how data were generated or obtained.
When data are digitised, transcribed or manually entered into systems, consistency becomes particularly important. Standardised procedures, clear coding schemes and well-designed data entry processes help reduce transcription errors and inconsistencies.
As data move through cleaning and preparation stages, quality control typically involves a combination of automated and manual checks. For example, for quantitative data producers may review value ranges and formats, identify missing or inconsistent entries, verify samples against source materials and run basic summaries to detect anomalies. Any corrections or exclusions should be recorded so that changes can be traced and understood later. For qualitative, textual and multimedia data, quality control may involve checking transcription accuracy, ensuring consistent formatting of transcriptions, reviewing annotations and metadata, and confirming that files are complete and correctly labelled. For administrative and linked quality processes, these may also include validating record-linkage outcomes, checking for unexpected structural changes in incoming data, and monitoring completeness and coverage over time.
Across all data types, documenting quality assurance procedures and decisions is essential. This supports transparency within research teams, enables reproducibility, and reduces the effort required to prepare data for later documentation, anonymisation, and sharing.
Many projects create additional derived variables or enriched datasets during processing. These enhancements can increase analytical potential and support future reuse. Examples include harmonised variables across datasets, derived indicators, spatially referenced data or integrated classification variables.
When such additions are made, it is important to document the methods, assumptions and decisions involved. Clear documentation ensures that derived data remain interpretable and trustworthy.
Quality management should be revisited regularly as projects evolve. New data sources, changing workflows or expanding collaborations can introduce additional risks. Periodic review of quality control procedures helps ensure that data remain reliable and usable throughout the active research phase.
Maintaining good quality and integrity practices during the project also reduces the effort required later when preparing data for documentation, anonymisation and sharing.
Data Quality Checklist
Before data collection check:
- Is there an agreed file format and structure?
- Is the location of data file storage secure?
- Is it agreed who has access to data files?
- Is there a clear tracking mechanism for data files and changes?
During data collection check:
- Are there missing or unexpected values?
- Are variables names and labels consistent?
- Are files saved in agreed format and structure?
- Are any files missing or corrupted?
Pre-analysis check
- Are all variables correctly coded?
- Are there duplicates or inconsistencies?
- Do data summaries/descriptive statistics look reasonable?
Pre-deposit check
Please see Prepare your data collection for deposit for an in-depth pre-deposit check.