← New search

Other meanings of Research data

Research Methods

Research data

Research data are the factual records (as raw measurements, observations, or computational outputs) collected, observed, or generated to validate and produce original research findings. They underpin the scientific method, enabling reproducibility and secondary analysis across disciplines.

2.5 quintillion
bytes of data created daily (2023 estimate)
global data volume
~50%
of research data are lost within 2 decades
data preservation gap
1,000+
data repositories registered in re3data
repository count
1

Definition and types

Research data encompass both quantitative and qualitative records, including sensor readings, survey responses, interview transcripts, genomic sequences, simulation outputs, and laboratory notebooks. They are distinct from derived results (e.g., published figures) and from metadata, which describe the data's context. The FAIR Guiding Principles (Findable, Accessible, Interoperable, Reusable) have become the international benchmark for data stewardship, adopted by major funders like the European Commission and the U.S. National Science Foundation.1

Data can be classified by origin: observational (e.g., climate records), experimental (e.g., clinical trial results), simulation-based (e.g., climate models), or derived/compiled (e.g., census aggregates). Each type poses distinct challenges for curation and sharing.

2

Lifecycle and management

The data lifecycle spans planning, collection, processing, analysis, preservation, and sharing. Data management plans (DMPs) are now mandatory for many grants, detailing formats, storage, and access policies. Tools like the DMPTool and institutional repositories support researchers in meeting these requirements.2

Version control and documentation are critical; without them, data become unusable even for the original researcher. The Open Science Framework and Zenodo provide free platforms for versioned, citable data sharing, while domain-specific repositories (e.g., GenBank for genetic sequences) enforce community standards.

3

Sharing and open data

Open data policies, driven by funders and journals, aim to increase reproducibility and accelerate discovery. The PLOS and Nature journals require data availability statements, and many mandate deposition in recognized repositories. However, barriers persist: concerns about privacy, intellectual property, and the effort of curation lead many researchers to share only minimal data.

Data citation is evolving; the DataCite DOI system assigns persistent identifiers to datasets, enabling proper credit and tracking of reuse. Despite this, a 2020 study found that fewer than 20% of datasets in repositories were ever cited, highlighting the gap between sharing and actual reuse.

4

Lesser-known aspects

Beyond the mainstream, research data include 'dark data'—information collected but never analyzed, often due to resource constraints. For example, the Human Genome Project generated petabytes of raw sequencing data, much of which remains unexplored. Similarly, historical datasets, such as the Millennium Cohort Study, offer longitudinal insights that were not anticipated at collection time.

Edge cases include 'data fossils'—legacy formats (e.g., floppy disks) that require emulation to read, and 'data rescue' efforts like DataRescue that preserve at-risk federal datasets. The RDA (Research Data Alliance) has developed standards for such challenges, but many remain unresolved, underscoring the field's ongoing evolution.

Glossary

FAIR principles
Guidelines for making data Findable, Accessible, Interoperable, and Reusable.
Metadata
Data about data, providing context such as collection methods and units.
Data management plan
A formal document outlining how data will be handled during and after a project.
Persistent identifier
A stable reference (e.g., DOI) that uniquely identifies a digital object.

This entry focuses on the scientific research context; for other uses, see data (computing) or data (statistics).