Data producer support home page

Organising data details to help others correctly find and use it

Metadata is structured information that describes, explains, locates, or otherwise makes it easier to find, use, and manage data over time. Created throughout the research lifecycle, metadata can be embedded within raw data files, recorded in documentation, or published within repository catalogue records.

Metadata exist in several forms and serve different functions. Without clear metadata, datasets become difficult to discover, cite, or interpret. Depending on its role, metadata can help researchers locate a study, understand how different files relate to one another, or ensure data is managed in line with legal and preservation standards.

Descriptive metadata support identification and discovery. They enable data to be found, distinguished from other resources and cited correctly.

Descriptive metadata typically include:

  • title
  • creator(s) and affiliated institution(s)
  • funder information
  • persistent identifier (e.g. DOI)
  • abstract or summary
  • keywords and subject classifications
  • geographic and temporal coverage.

These elements form the basis of repository catalogue records and enable interoperability between systems.

Structural metadata describe how data are organised and how components relate to one another. They provide information necessary to interpret and analyse data correctly.

Examples include:

  • variable names and labels
  • value labels and coding schemes
  • relationships between files in a multi-file dataset
  • definitions of nodes and edges in network data
  • folder or hierarchical structures.

Structural metadata are often embedded within data files. For example, variable labels and missing value definitions in SPSS form part of a dataset’s structural metadata.

Administrative metadata support the management, governance and preservation of data.

They may include:

  • file formats and software versions
  • version history
  • rights and copyright information
  • access restrictions and embargo periods
  • preservation actions.

Administrative metadata ensure that datasets can be managed appropriately and used in accordance with legal and ethical requirements.

Metadata are closely related to documentation but serve a distinct function. Metadata provide structured, often standardised descriptions of data that support discovery, identification and management.

Documentation provides more detailed explanatory material, such as user guides, codebooks and methodological reports, that enable meaningful interpretation and reuse.

Both are necessary to ensure that data remain accessible, understandable and reusable over time.

The Data Documentation Initiative (DDI) is a rich and detailed metadata standard originally designed for describing social, behavioural and economic sciences data. It is used by most social science data archives in the world.

DDI catalogue records contain mandatory and optional metadata elements relating to study description, data file description and variable description. The study description contains information about the context of the data collection, such as bibliographic citation of the study and data, the scope of the study, like topics, geography, time, method of data collection, sampling and processing, data access information, and information on accompanying materials.

The data file description indicates data format, file type, file structure, missing data, weighting variables and software used. Variable descriptions indicate the variable labels and codes.

The UK Data Service uses DDI to structure catalogue records. The use of standardised records in eXtensible Mark-up Language (XML) brings key data documentation together into a single document, creating rich and structured content about the data.

Metadata can be harvested for data sharing through the Open Archives Initiative Protocol for Metadata Harvesting (OAI-PMH).

Our DDI records contain mandatory and optional metadata elements on the following:

  • Study description – information about the context of the data collection, such as bibliographic citation of the study and data, the scope of the study (topics, geography, time), methodology of data collection, sampling and processing, data access information, and information on accompanying materials.
  • Data file description – information on data format, file type, file structure, missing data, weighting variables and software.
  • Variable descriptions.

We collect the initial catalogue record information from our data deposit form, which is completed by the data depositor. We then enhance information from the accompanying documentation to create a conformant metadata record.

Where data producers can provide detailed and meaningful data collection titles, descriptions, keywords, contextual and methodological information in the deposit form, it helps us to create rich resource-discovery metadata for their deposited collections. We assign keywords from our own ELSST Thesaurus.

Depositors are encouraged to provide information about original and subsequent reports and publications or presentations based on our data collections, so these references can be added as further documentation.

We are always interested in capturing case studies of data reuse in our collections to encourage further use of data. We prepare a standard bibliographic citation for each data collection so that users can correctly cite the data sources in research outputs.

Example extract from a UK Data Service DDI catalogue record

<stdyInfo>
<subject>
<keyword>
</keyword>
<topcClas Vocab=”unknown”>Economic processes and indicators – Economics</topcClas>
<topcClas Vocab=”unknown”>Economic systems and development – Economics</topcClas>
<topcClas Vocab=”unknown”>General – Employment and labour</topcClas>
<topcClas Vocab=”unknown”>Elites and leadership – Social stratification and groupings</topcClas>
<topcClas Vocab=”unknown”>Management and organisation – Industry and management</topcClas>
</subject>
<abstract>This project aims to develop knowledge and understanding of the contemporary globalisation of the headhunting industry in Europe and its implications for new forms and geographies of executive search and selection. Europe has become the most complex and sophisticated pan-regional market for executive search, fuelled by free labour mobility within the EU, thereby offering a unique environment in which to study the changing practices of the headhunting industry.</abstract>
<sumDscr>
<nation>UK</nation>
<geogCover>London</geogCover>
<nation>France</nation>
<geogCover>Paris</geogCover>
<nation>The Netherlands</nation>
<geogCover>Amsterdam</geogCover>
<nation>Germany</nation>
<geogCover>Frankfurt</geogCover> <nation>Belgium</nation> <geogCover>Brussels</geogCover> <universe>Executive search consultants, researchers and associations in London, Paris, Frankfurt, Amsterdam and Brussels, 2006-2007</universe>
<anlyUnit>Individuals</anlyUnit>
<anlyUnit>Institutions/organisations</anlyUnit>
<collDate event=”start” date=”01/2006″>January 2006</collDate>
<collDate event=”end” date=”08/2007″>August 2007</collDate>
<timePrd event=”start” date=”01/1980″>January 1980</timePrd>
<timePrd event=”end” date=”08/2007″>August 2007</timePrd>
</sumDscr>
<notes>
</notes>
<method>
<dataColl>
<sources>
<dataSrc rule=”Sources used”>The Executive Grapevine, The Directory of Executive Recruitment, published by The Executive Grapevine International Ltd. Editions consulted: 1980, 1985, 1990, 1994, 2000, 2005</dataSrc>
<dataSrc rule=”Source location and access”>Copies are held at the British Library and the most recent edition is available for private purchase.</dataSrc>
</sources>
<collMode>Face-to-face interview</collMode>
<collMode rule=”Other”>Time series for search firm and office data collated from the Executive Grapevine Directories of International Recruitment</collMode>
<sampProc>Purposive selection/case studies</sampProc>
<timeMeth>Cross-sectional (one-time) study</timeMeth>
</dataColl>
</method>
<othrStdyMat>