Platform and web-derived data

Data producer support home page

Managing web-scraped data: terms of use, privacy, and code sharing

Online content is now central to much contemporary social science research. Researchers routinely collect and analyse material from social media platforms, online forums, news sites, commercial websites and other digital environments. These data may include posts, comments, images, interaction metrics, hyperlinks, engagement statistics or network structures.

In many cases, such data are publicly accessible and may be lawfully collected for research purposes, provided that appropriate legal and ethical considerations are addressed. When planning to redistribute, archive, or share these data for reuse, it is important to review the relevant terms of use and any applicable legal or ethical requirements to ensure that the data can be shared.

A critical distinction in online research is the difference between using data for analysis and redistributing those data through a repository or archive.

Most online platforms and websites operate under contractual terms that govern how their content may be accessed, processed and reused. Even where content is publicly visible, terms of service may restrict automated harvesting, bulk downloading, republication or redistribution outside the original platform.

Researchers must therefore consider not only whether data collection is lawful, but whether onward sharing is permitted. Public visibility does not imply unrestricted reuse.

A critical distinction in online research is the difference between using data for analysis and redistributing those data through a repository or archive.

Most online platforms and websites operate under contractual terms that govern how their content may be accessed, processed and reused. Even where content is publicly visible, terms of service may restrict automated harvesting, bulk downloading, republication or redistribution outside the original platform.

Researchers must therefore consider not only whether data collection is lawful, but whether onward sharing is permitted. Public visibility does not imply unrestricted reuse.

Online content frequently contains personal data, even where individuals use pseudonyms. Aggregating and redistributing such material may alter its context, visibility and permanence. Content originally shared within a specific platform environment may take on different implications when archived and made discoverable in a research repository.

Where full redistribution of online data is not permitted, careful documentation becomes especially important. Transparency about how data were collected, processed and filtered enables scholarly scrutiny even where raw content cannot be shared.

Documentation for platform or web-derived data should clearly describe:

  • The source websites or platforms.
  • The date or time period of collection.
  • The method of collection (e.g. API access, automated harvesting).
  • Any sampling, filtering or cleaning procedures applied.
  • Any legal, contractual or ethical constraints affecting reuse.

In many cases, sharing well-documented code alongside detailed methodological documentation provides a proportionate way to support reproducibility without breaching contractual or copyright restrictions.