What are social science data?

Data producer support home page

Defining social science data in research

Social science data can take many different forms. Understanding how data are produced, structured and combined helps data producers and their teams manage them appropriately and select suitable approaches to sharing and access.

Social science data can be understood as evidence about people and society. This includes, for example: data about how we live, work, learn, vote, travel, care for others, spend and save, use services, and online technologies, as well as data on markets, businesses, production, trade and their geographic or environmental context.

These data are used to study social, economic, cultural and behavioural patterns and outcomes across different populations, organisations and contexts.

Social science data can be described in different ways depending on the research context and purpose. The examples below illustrate common perspectives that can help data producers understand the characteristics of their data and the implications for management and sharing.

These are not fixed categories. Many projects will combine elements from more than one group.

Social science data can be understood according to their technical structure. The structure of the data influences how they should be documented, preserved and prepared for deposit.

For example:

  • Tabular data: Organised in rows and columns (e.g. surveys, administrative datasets, experimental data).
  • Text data: Document-based materials such as transcripts, policy documents or corpora.
  • Image, audio and video data: Visual or auditory recordings.
  • Spatial data: Data that include geographic coordinates or mapped features.
  • Relational data: Network or graph-structured data describing relationships between entities.
  • Web and digital trace data: Data generated through social media or online platforms.

Many research projects combine more than one structural type.

Understanding the structural form of your data helps you prepare appropriately for the sharing and archiving of data.

Another way to think about social science data is in terms of how they are generated or acquired.

Primary data

Primary data are generated directly by a research project. Examples may include:

  • survey microdata
  • interviews and focus groups
  • experimental and evaluation data
  • field observations and diaries
  • sensor or device-generated streams.

These data are usually collected for a specific research purpose and designed by the research team.

Secondary data

Secondary data are obtained from external sources and reused for new research purposes. Common examples include:

  • official statistics and census data
  • administrative and government records
  • platform and commercial datasets
  • web-scraped corpora
  • sensor or device-generated streams when these are not produced for the primary purpose of research
  • macro-level indicators such as GDP, CPI and labour market statistics.

When working with secondary data, researchers must consider the original conditions of data collection and any legal, contractual or ethical requirements associated with reuse.

Supporting and derivative objects

Social science research often produces additional materials that are essential for understanding, verifying or reusing data. These materials are often part of the research data collection and should be managed and prepared alongside the data themselves. These may include:

  • code and syntax files
  • data processing workflows
  • synthetic data generated from original data or their metadata
  • teaching and training packages.

These objects form part of the research data record and play an important role in transparency and reproducibility.

Not all materials will be shared in the same way. Access conditions may differ depending on legal, ethical or contractual restrictions, but these materials should still be documented and considered as part of the overall data collection.

Social science data may also be described by the nature of the data, or how the data are interpreted.

Quantitative data

Quantitative data typically consist of numerical and structured information. Examples include survey datasets, administrative microdata, transaction records, indicators and aggregated statistics. These data are often analysed using statistical methods.

Qualitative data

Qualitative data often consist of textual, audio or visual materials. Examples include interview transcripts, focus group recordings, fieldnotes, open-ended survey responses, images and video. These data often require different documentation, anonymisation, and access approaches.

Another way to think about social science data is in terms of how they are generated or produced.

For example, data may be generated through:

  • surveys and questionnaires
  • interviews and focus groups
  • experiments and evaluations
  • observational and field-based research
  • administrative and operational systems
  • sensors and digital devices
  • web-based collection and web scraping
  • reusing existing published data collections
  • transactional and platform systems.

These approaches influence how data are structured, documented, anonymised and shared. However, many research projects combine multiple generation methods, and the same dataset may fall into more than one category.

Another useful perspective is to consider what the data are about.

Examples may include:

  • Data about people and households: Such as demographic, behavioural, health, education and employment data collected through surveys, administrative systems, experiments or digital platforms.
  • Data about organisations and institutions: Including information about organisations, schools, hospitals, charities and public bodies, such as organisational surveys, administrative records, financial or transactions data, and policy documents.
  • Data about places and environments: Including neighbourhood, regional and national statistics, geospatial datasets, environmental sensor data, and macroeconomic indicators such as GDP.

Data about culture, communication and media: including textual corpora, social media content, news archives, images, audio recordings and video data used for content or discourse analysis.

Some types of social science data involve additional legal, ethical or governance responsibilities, particularly where data relate to identifiable individuals or where multiple datasets are combined.

Data about people

A substantial proportion of social science data relate to identifiable individuals.

Where research data include information about people, additional legal and ethical responsibilities apply. In these cases, the UKRI Guidance on Sharing Data About People should be followed in full, alongside the guidance provided on these pages.

UKRI also provides additional guidance on research involving human participants through the Good Research Resource Hub, including resources on topics such as:

These resources should be consulted for those working with UKRI funding but as a good resource where research involves collecting or using data about people.

Linked data

Data are increasingly created by combining information from multiple sources. For example, survey data may be linked with administrative, health or education records.

When working with linked data, data producers must ensure that:

  • Legal and contractual requirements associated with each data source are respected.
  • Data protection and confidentiality risks are assessed.
  • Appropriate access conditions are applied to shared or deposited data.

These considerations should be addressed during project planning and reflected in data sharing decisions.