A research consortium combines decades of Earth-observation imagery with field measurements and clinical-study datasets. The original projects may be complete, but researchers still need to reproduce published findings, compare new observations against the historical record, and evaluate new AI methods.
The challenge is not simply to keep files. Future researchers need to know what each object is, where it came from, how it was processed, what restrictions apply, and whether it can still be interpreted and verified.
Scientific data must outlive the project that produced it—but the preservation design must fit the research, the data, and the obligations attached to each collection.
The grant ends. The research questions keep changing.
A completed grant may leave behind raw observations, processed products, calibration files, quality masks, code, lab measurements, documentation, and data derived from people. Some of that material supports published conclusions. Some may be valuable only when a later team combines it with new observations or a new method.
If the only copy lives on a project server, an aging disk array, or with one investigator, technology turnover and staff changes can break the chain between a dataset and its scientific meaning. A later AI project may be able to use the bits but be unable to establish which instrument, protocol, model version, or preprocessing step produced them.
Preservation has to retain the context needed to interpret and evaluate the data, not just the payload.
A 50-year Earth-observation record becomes new research material
The USGS Landsat Collection 2 provides Level-1 observations from Landsats 1–9 beginning in 1972, alongside later Level-2 and Level-3 science products. USGS has reprocessed the archive as methods and reference data improved, and publishes collection identifiers, metadata, product documentation, and reprocessing information.
That long record lets researchers compare current conditions with earlier observations, build consistent time series, and develop analyses that were not foreseeable when the first scenes were acquired. The archive is a real example of scientific value accumulating over decades when source observations and processing context remain available.
It is not a claim that ELS stores Landsat data. It illustrates why research teams may want a preservation architecture that can retain source data and versions while allowing future authorized users to locate and retrieve them.
Institutional and funded research: plan for sharing and retention
For NIH-funded or conducted research that generates scientific data, the NIH Data Management and Sharing Policy requires a Data Management and Sharing Plan and compliance with the plan approved by the relevant NIH Institute, Center, or Office. Shared data should be made accessible as soon as possible and no later than the associated publication or the end of the award/support period, whichever comes first.
That policy does not require every dataset to be openly released. NIH recognizes that legal, ethical, privacy, technical, and other factors can limit sharing, and strongly encourages established repositories where practical. Investigators must account for consent, controlled access, tribal data governance where relevant, genomic-data policy, award terms, repository commitments, and institutional rules.
For NSF-supported work, the applicable Proposal and Award Policies and Procedures Guide and award terms set expectations for sharing primary data and supporting materials within a reasonable time, with exceptions for privacy, confidentiality, field-specific concerns, and legitimate interests. NSF has issued supplements to PAPPG 24-1 that may apply based on award date, so the current award terms and policy version need review.
In either case, storage capacity is only one line of a data-management plan. Teams also need named stewardship, repository or access arrangements, metadata, persistent identifiers, retention/disposition decisions, and funding for preservation and sharing activities.
Earth observation: preserve source, products, and processing lineage
A climate or land-use team may combine new satellite scenes with decades of prior imagery, field observations, and derived analytical products. To make comparisons meaningful, preserve the received observations alongside calibrated or analysis-ready products, quality masks, geospatial metadata, collection identifiers, processing versions, and documented transformations.
A generated product should remain connected to the source observations and processing recipe that produced it. If a new algorithm reprocesses an archive, retain enough version and provenance information to distinguish the earlier result from the replacement and to reproduce either one when required.
NASA and USGS have agency-specific scientific data policies and archive practices. Those are not interchangeable with one generic federal retention statute; project requirements can arise from a mission, data center, award, repository, or agency policy. The Landsat archive shows why continuity, reprocessing records, and standardized metadata matter for reuse.
Clinical and human-participant research: protect access as well as longevity
A medical research center may retain study datasets, images, and supporting records for future approved analysis. Requirements depend on the study type and records involved. Consent, IRB protocol, award conditions, sponsor contracts, institutional policy, and state law can all affect what must be retained and who may access it.
The Common Rule at 45 CFR 46.115 sets a minimum retention period for specified IRB records: at least three years, and research-related IRB records for at least three years after completion. It is not a blanket three-year retention period for all scientific datasets.
For covered investigational drug studies, 21 CFR 312.62(c) specifies retention periods for required investigator records tied to approval or discontinuation of the investigation. For covered investigational device studies, 21 CFR 812.140(d) defines retention for specified investigator and sponsor records in relation to investigation completion and regulatory submissions. These rules cover defined record classes and study contexts; other award, sponsor, contract, and institutional terms may extend retention.
HIPAA applies to protected health information held or handled by covered entities and their business associates, not to every research institution or every de-identified dataset. Research use and disclosure may require authorization, an IRB or Privacy Board waiver, a limited-data-set agreement, or another permitted basis. Long-term preservation should maintain those access constraints and agreements rather than turn preservation into open publication.
AI reuse raises the value of provenance
Historical observations can become training inputs, validation sets, or comparison baselines for new AI systems. That reuse can be valuable, but only if the team can determine whether the data is appropriate for the task, what transformations have been applied, and whether access or consent conditions permit the intended use.
Preserve the original dataset, labels, derived features, model and preprocessing versions, evaluation outputs, and links between each derivative and its source. Keep a record of access restrictions and data-use agreements so a future team does not treat technical availability as permission to reuse.
The NIST AI Risk Management Framework is voluntary guidance, not a research-data regulation or certification. It can help teams think about traceability, data quality, governance, and risk when datasets are reused in AI development and evaluation.
An archive is not a repository or a research information system
A scientific archive tier can help retain large source collections and make them retrievable over time. It does not automatically provide a domain repository, DOI registration, searchable scientific catalog, participant-consent enforcement, laboratory information management, validated clinical workflow, or publication-quality metadata.
A useful architecture connects those systems: repository or research-management tools govern discovery, description, sharing, and access; policy defines who may use the data and for what; storage services hold and protect the objects; preservation procedures verify integrity and plan for migration.
For ELS environments, oRain can manage object and media-location awareness across networked libraries. Integrations should use interfaces supported by the research, repository, and storage systems involved; system-specific design and validation remain part of implementation.
Design for revalidation, not just retention
A practical preservation workflow is: receive and identify the data; retain its source and required metadata; record a cryptographic fixity value; create the preservation copy; preserve access restrictions and data-use terms; periodically verify integrity; test retrieval; and document any migration or transformation.
For each dataset, identify a responsible steward, authoritative version, repository or discovery record, allowed uses, retention trigger, review date, restore procedure, and disposition authority. Where long-term reproducibility matters, retain processing code, parameters, calibration references, and software or format information alongside the outputs.
Maintain redundancy appropriate to the risk. An online optical copy can complement performance storage, while an offline copy can add physical separation. Copies in one device or one facility do not provide geographic disaster recovery; the number and placement of copies should reflect project, institution, funder, and risk requirements.
Preserve the evidence behind the next discovery
Research records often become more useful when joined to later observations, new measurements, or better analytical tools. The Landsat archive demonstrates how continuity across decades can support questions that could not be fully formed when the original data was collected.
A strong preservation design keeps observations findable, interpretable, verifiable, and appropriately governed as media, software, teams, and analytical methods change. For clinical data, it also preserves the restrictions that protect participants. For AI reuse, it retains the provenance and version links needed to explain where training and evaluation material came from.
The goal is not simply to keep research data. It is to keep enough of the research record for future teams to understand what they have, establish what they are allowed to do with it, and determine whether a result can be reproduced.
