Additional considerations
Additional considerations
In addition to establishing ownership and obtaining the necessary permissions, researchers should be aware of a range of considerations that may influence whether and how research data can be shared. These factors can affect the rights associated with research materials, the conditions under which data may be reused or distributed, and the legal obligations that apply in different contexts. Considering these issues alongside institutional policies, funder requirements, and ethical responsibilities helps ensure that research data are shared lawfully, responsibly, and in a manner that respects the rights of all relevant stakeholders. The following sections outline the key additional considerations that researchers should consider when sharing research data.
The duration of copyright depends on the type of work. The table below is highly simplified.
| Type of work | Copyright duration |
|---|---|
| Literary and artistic works | 70 years from the end of the year of the death of the creator. |
| Sound recordings | 50 years from date of creation. |
| Typographical arrangements | 25 years from date of publication. |
| Crown Copyright | 50 years from date of publication or 125 years from date of creation. |
Copyright supports creators by ensuring they receive fair compensation for their work, which encourages more creativity and ensures proper recognition. However, copyright law also includes limitations and exceptions, allowing for certain uses without the creator’s permission. Many countries allow the adaptation of copyrighted books for people with visual impairments or other disabilities, making them more accessible. For example, UK copyright law permits the creation of accessible format copies (such as Braille or subtitled versions) for disabled individuals without breaching copyright, supported by the Marrakesh Treaty.
In the same way, researchers and students are allowed to copy limited portions of copyrighted works (such as books, plays, music, and photos) for non-commercial research or personal study. This exemption is known as “Fair dealing” which is a legal term used to determine whether the use of copyrighted material is lawful or infringes on copyright. There is no fixed definition of fair dealing and it depends on the specific situation and context. Essentially, fair dealing means that using a copyrighted work under certain exceptions must be both reasonable and justifiable.
This only applies to literary, dramatic, musical or artistic work, not to films or recordings. An acknowledgement should give credit to the data source used, the data distributor and the copyright holder.
Fair dealing use could be assessed by asking how a fair-minded and honest person would handle the work. Key factors include:
- Market Impact: If the use substitutes the original work and causes revenue loss, it is likely to be unfair.
- Amount Used: The part taken should be reasonable and necessary, typically involving only part of the work.
The importance of these factors depends on the specific case and type of use.
While exceptions exist under the concept of fair dealing, which permits the use of copyrighted material for non-commercial teaching or research, private study, criticism, or review without infringing copyright provided that the original source and author are properly acknowledged caution must be exercised in the context of data sharing. In the United Kingdom, such exceptions are not typically recognised by responsible data repositories, and the sharing of data or information generally requires explicit copyright permission.
Case study
Background: A social sciences researcher at a UK university, has completed a qualitative research project examining media representations of migration. As part of the research, the researcher collected interview transcripts, compiled annotated excerpts from newspaper articles, captured screenshots of online news content, and developed detailed coding frameworks and analytical notes. In line with the funder’s open research requirements, researcher intends to deposit the dataset with the UK Data Service so that other researchers can access and reuse it.
The Issue: Although much of dataset is the original work, it also incorporates substantial amounts of copyrighted material. This includes verbatim extracts from newspaper articles owned by publishers, screenshots of digital news platforms, and annotated documents containing large portions of reproduced text. Researcher assumes that these inclusions are permissible under the UK concept of fair dealing, particularly as her work is for non-commercial research and includes proper attribution to original sources.
Repository Review: When the dataset was deposited, the repository conducts a review and identifies potential copyright risks. It was advised that while fair dealing may allow the use of copyrighted material within research activities, it does not generally extend to the redistribution or public sharing of that material through a data repository. Because repositories are responsible for ensuring legal compliance, they cannot rely on fair dealing exceptions alone and instead require clear evidence of permission or rights to share all included content.
Resolution: Faced with this challenge, the researcher must reconsider how to proceed with the data deposit. Seeking permission from copyright holders is one option, or she could remove or redact the copyrighted content, replacing it with summaries or descriptions, though this may reduce the depth and transparency of her dataset is the only other option.
Copyright laws vary by country, but global harmonisation has been achieved through treaties like the Berne Convention (1886), administered by the World Intellectual Property Organization (WIPO). A key principle of the Berne Convention is ‘national treatment’, requiring member countries to protect works from other countries from infringement as they would their own. The treaty sets minimum standards for aspects like types of protected works, duration, and exceptions.
There is a detailed overview of copyright legislation in different European countries in the CESSDA Data Management Expert Guide Diversity in copyright – Data Management Expert Guide.
When copyright expires, a work enters the public domain and can be used freely. To determine if a work is in the public domain, consider who created it and when they died, if it has been over 70 years since the creator’s death, the work is likely public domain. However, there are key considerations:
- Public domain status varies by country.
- Reproductions or recordings of public domain works can still have their own copyright.
- Adaptations of public domain works, such as modern films or adaptations, are protected by copyright and require permission to use.
Further guidance Information in the public domain | ICO.
If you use the orphan work for non-commercial purposes, like one of the many uses considered fair use (e.g. criticism, parody, news reporting, classroom teaching, scholarship, or research), you may fall within the fair use. However, when sharing research data, all rights holders must be identified, and copyright permissions obtained. If one or more rights holders cannot be found, the work is classified as an ‘orphan work’.
Generally, orphan works cannot be reproduced without permission, except in specific circumstances. To use an orphan work, a ‘diligent search’ must be conducted to locate the rights holders, consulting appropriate sources for the type of work. This process, which varies slightly by country, should be completed before using the work and is often time-consuming.
Detailed guidance on Orphan works diligent search guidance – GOV.UK.
Social media platforms generate a vast amount of data from user posts, comments, likes, photos, videos, and interactions, offering valuable insights into user behaviour, preferences, and trends. Accessing social media data for research is best done through application programming interfaces (APIs), which provide a direct and authentic connection to platform data.
These APIs, offered by social media platforms or their authorised resellers, deliver machine-readable data, enabling researchers to process large datasets quickly and efficiently in real time and detect patterns at scale. APIs act as a bridge between the platform and the user and define how data can be accessed and used, including technical guidelines and any rules or restrictions.
Social media data are a powerful tool for both commercial and academic use. The terms of use for major social media platforms are similar when it comes to IP rights. Just like books or journals, the content you post on these platforms is protected by copyright.
By agreeing to the platform’s terms and conditions when creating an account, you grant the site a licence to use your content in various ways. This can include allowing researchers to access your data for academic purposes. Therefore, if researchers are using social media data, they must follow the platform’s terms and conditions, as well as the guidelines set by API developers.
For example, X allows researchers to access tweets via its public API. However, challenges arise when researchers attempt to publish or archive data for future use. X’s policy restricts sharing or storing collected data, although it does allow archiving tweet IDs (the unique numbers for each tweet) and user IDs. This enables researchers to recreate datasets in the future, but only if X continues to provide access to historical data. X also imposes restrictions on how tweets can be published. For example, tweets must be quoted in their original form without modifications, which may pose privacy concerns, as anonymisation is not allowed. Therefore, it is essential to check the terms and conditions associated with the platform when it comes to publishing the data for future use.
Case Study: Copyright Complexity in Video Data Sharing
Background: A cultural studies researcher at a UK university, conducts an ethnographic study exploring street performance culture in urban public spaces. The research focuses on how street dancers engage with popular music to create spontaneous performances that attract public audiences and social media attention. As part of the fieldwork, they record high-quality videos of street performers dancing in locations across London and Manchester. These videos capture not only the dancers’ movements but also the background audio of well-known commercial songs played through portable speakers. Following the completion of her project, researcher plans to deposit the video dataset with the UK Data Service to support transparency and future research reuse.
The Issue: The dataset presents multiple layers of copyright and related rights. First, the music captured in the recordings is protected by copyright, typically owned by record labels and music publishers. Second, the dancers themselves may hold performers’ rights in their performances, giving them control over how recordings of their performances are used and shared. Third, while the researcher owns the copyright in the video recordings that were created, their rights do not override those of the music owners or performers. Although the recordings were made in public spaces and for research purposes, the researcher initially assumes that fair dealing for research may allow them to share the videos with appropriate attribution.
Repository Review: When the researcher submits the dataset, the repository identifies significant risks, noting that the videos contain copyrighted music used without permission and identifiable performers whose consent for redistribution may be unclear or insufficient. The repository explains that fair dealing does not typically extend to making such materials publicly available, especially when multiple rights holders are involved. As a result, the dataset cannot be accepted in its current form without evidence of permissions or substantial modification.
Resolution: Seeking permission from music rights holders would be highly impractical due to the number of songs involved and the cost and administrative burden of licensing. Contacting each performer to obtain explicit consent for data sharing is also challenging, particularly where performers were recorded informally in public settings and may not have anticipated their inclusion in a public dataset. At the same time, removing or heavily editing the videos could undermine the integrity of her research, which relies on the interaction between music, movement, and public space. To address these challenges, the researcher adopts a multi-layered strategy. They created summaries to illustrate key points in the dataset with detailed metadata to share the information via UK Data Service.
Key Lessons: This case illustrates the complexity of working with audiovisual data that incorporates multiple layers of copyright and related rights. It highlights that ownership of a recording does not equate to ownership of all content within it, particularly when music and human performances are involved. The case also demonstrates that fair dealing is limited in scope and does not generally support open data sharing through repositories. Researchers working with video data should plan ahead by designing consent processes that address data sharing, considering the use of non-copyrighted or licensed materials, and exploring technical solutions to mitigate legal risks.
Secondary data often provide valuable insights and save considerable time and resources compared to collecting primary data. Under the principle of fair dealing, researchers are permitted to copy and extract limited portions of such data for the purpose of research, provided this use is justified and properly acknowledged. However, when it comes to data sharing, researchers may encounter several challenges.
Copyright and licensing issues, in particular, can significantly restrict the extent to which data can be redistributed or published. Many data sources are protected by copyright or subject to specific licence terms that limit reuse or sharing beyond personal research purposes. The UK Data Service rarely encounters copyright problems relating to sharing data beyond the original research project collecting primary data. To help researchers ensure no copyright breaches arise when wishing to deposit derived data from existing resources we make available a variable information log template (Excel). The template can be adapted depending on the type of analysis carried out.
We advise all researchers to check licences, and terms and conditions for any data sources they use. If any uncertainty arises they should get in touch with the data provider at their earliest convenience.
For further guidance and example case studies please visit our copyright scenarios web page.
Copyright scenarios for data sharing
Scenario 1: Copyright of a media database
Scenario: A researcher has collated articles about the Prime Minister from The Guardian over the past ten years, using the LexisNexis database to source articles. These are then transcribed/copied by the researcher into a database so that content analysis can be applied. The researcher offers a copy of the database with the original transcribed text to the UK Data Service.
Rights issue: Researchers cannot share either of these data sources as they do not have copyright in the original material. The UK Data Service cannot accept these data as to do so would be breach of copyright. The rights holders, in this case The Guardian and LexisNexis, would need to provide consent for archiving.
Scenario 2: Copyright of archived data
Scenario: A researcher uses International Social Survey Programme (ISSP) data obtained from GESIS – Leibniz Institute for the Social Sciences in Germany. These data are available to registered GESIS users. The researcher incorporates some of the ISSP data within a database containing his own research data. Can this database be placed on the researcher’s website?
Rights issue: Although the ISSP data are available for free to all registered GESIS researchers, this does not mean that the data can be published on a website and made available to others. The data can be incorporated into a database and used for personal analysis. Before this dataset is placed on a website, permission must be sought from the data owner.
Scenario 3: Safeguarded data obtained from the UK Data Service
Scenario: A researcher has used data from the National Diet and Nutrition Survey (NDNS), obtained via the UK Data Service. NDNS data are available under safeguarded access and under Crown Copyright. The researcher has processed the NDNS data (filtered, integrated and aggregated data across variables, while maintaining individual records) and used the processed data to model food chain risks. The researcher would like to archive the processed data that were used as input data for the modelling, as well as the modelling code, at the UK Data Service.
Rights Issue: There is joint copyright over the processed data, shared between the researcher and the Crown (holding copyright over the NDNS data). The researcher must declare this joint copyright for the modelling data. Although the NDNS data are available as Crown Copyright, this does not mean that there are no restrictions associated to the use and sharing of data. The data is effectively anonymised and made available as ‘safeguarded’ data. This means use of the data is subject to the UK Data Service End User Licence (PDF), hence permission from the original data creator (Clause 15) is necessary to publish any derived data.
Scenario 4: Open data obtained from the UK Data Service
Scenario: A researcher has used data from the Participation Survey, 2021-2022, obtained via the UK Data Service. This data is available under Crown Copyright and available as an ‘open’ access collection. The researcher creates derived variables for their analysis, and they would like to archive the derived data at the UK Data Service.
Rights issue: There is joint copyright over the processed data, shared between the researcher and the Crown (holding copyright over the original Participation survey data). The researcher must declare this joint copyright. The data collection is published under the Open Government Licence (OGL) hence no further permission from the data creator is required to archive the derived data if acknowledgement is provided as described in the OGL.
Scenario 5: Copyright of the information available online
Scenario: A researcher studies how health issues around obesity are reported in the media in the last ten years. Freely available newspaper websites and library sources are used to obtain articles on this topic. Articles or excerpts are copied into a database and coded according to various criteria for content analysis. Can the researcher use such public data without breaching copyright? Can the database be archived and shared with other researchers?
Rights issue: Even though the articles obtained are freely available online, they might still be subject to copyright. Whilst such information can be used for personal research purposes (fair dealing), the articles cannot be archived, unless permission is obtained from the newspapers; otherwise, this would breach copyright. Terms and conditions of all data used should be checked before the archiving processes begins.
Scenario 6: Copyright of interviews with company directors
Scenario: A researcher has interviewed company directors about their careers and produced audio recordings and near verbatim transcripts herself which they have agreed to be made available as collected for future research. The researcher analyses this material and offers it to a data archive. What are the rights issues surrounding this offer of data?
Rights issues: In this case the company directors hold the copyright in their own recorded words, whilst the researcher holds copyright over the transcribed interviews. Quoting large extracts of the data, either in publications, or by archiving the transcripts, would breach the copyright of the interviewees in their recorded words.
If the researcher wants to publish large extracts of data, or archive transcripts, they need to request permission to do so from the interviewees or request that the interviewee transfers the copyright of the interview content to the researcher, which could be achieved by a Recording Agreement.
Scenario 7: Transcription from a printed work into a spreadsheet
Scenario: A researcher has copied a series of statistical information from a printed work into a spreadsheet. The transcription is a direct copy with minimal alterations. The book is in copyright.
Rights issue: The researcher should technically have cleared copyright before transcription. If the work is for personal use only, this can probably be disregarded, but if the newly constructed dataset is to be archived and disseminated, copyright clearance will need to be gained from the copyright holder.
Scenario 8: Copyright of survey questions
Scenario: A researcher wishes to reuse a set of questions from an existing survey questionnaire, to compare results between the newly proposed survey and the original.
Rights issue: It should be assumed that all survey questions and instruments are copyright protected, with copyright residing with the organisation who commissioned, designed or conducted the survey. Our advice, therefore, is to contact the copyright holder directly for permission to reproduce questionnaire text for any new use.
In our experience, the copyright holder will almost always grant that permission. Some questionnaires contain measurement scales, batteries of questions or classifications. These instruments are copyright to the institution or company that produced them and should not be reproduced without permission. In many cases, the copyright for these instruments is printed on the relevant page of the questionnaire.
Regardless of what the blanket terms and conditions of data use are, we would always advise researchers to try to negotiate the sharing of derived data with the data suppliers at the time of acquisition or purchase. There are cases where the agreement has been successful and has set a precedent for opening up the sharing of such data under a broader licence. We can advise on such negotiations.
Scenario 9: Copyright in academic research
Scenario: A research student wishes to deposit data in an archive that was collected as part of their PhD or an academic staff during their employment with the university.
Rights issue: Although, IP ownership will depend on national law and individual institutions’ policies. Most universities recognise as a general principle that students who are not employees of the university own the IP rights in the works, they produce purely based on knowledge received from lectures and teaching. However, there may be some circumstances where ownership must be shared or assigned to the university or a third party.
On the other hand, many universities or research centres claim ownership of any IP that is generated by academic staff in the course of their employment, and when IP is created using substantial institutional resources.
Best practice is to check the institutional policies to determine who owns the IP in the produced work.