SCIENCE

Software Heritage Grew From a Bet on Legal Deposit Copies

Software Heritage, unveiled by Inria in 2016 with UNESCO support, archives the world's source code and assigns persistent identifiers. The project grew from a wager that legal deposit, the centuries-old rule requiring publishers to send copies to national libraries, would eventually cover software. Here is how the archive was built and what it changes for reproducible computational science.

A Deposit Law Meets Source Code

Legal deposit is a legal requirement that mandates individuals or organizations submit copies of their publications to a designated repository, typically a national library. The practice preserves a nation's published heritage and ensures long-term access to information. Books, maps, newspapers, and official documents fall under deposit laws in many countries. Software does not.

Inria, the French national research institute for digital science, launched Software Heritage in 2016 as a non-profit multi-stakeholder initiative. UNESCO backed the effort. The framing was deliberate: source code is published heritage, not just a tool for building other things. A program released under an open license is a readable, versioned, authored work. Deposit law simply had not caught up.

The archive needed an identifier. The SoftWare Hash IDentifier, or SWHID, is a persistent identifier that uniquely identifies a piece of software source code and its version. It works like a DOI but is tailored for source code and compatible with versioning systems such as Git. A SWHID can point to a single file, a directory, a commit, or a full release snapshot.

The bet was not that libraries would demand code tomorrow. It was that archiving practice would come first, and law would follow. That sequence matters. Deposit mandates are slow, national, and uneven. A working archive with stable identifiers gives legislators something concrete to point at.

The Procedural Bet on Legal Deposit

National libraries already collect published works through deposit. Extending that duty to software is a legal stretch in most jurisdictions. Deposit law typically names tangible media and registered publishers. Source code lives in repositories, under licenses, and often without a publisher in the traditional sense. The legal category does not fit cleanly.

Rather than wait for statutes, the project chose to crawl public repositories. If the code is already publicly accessible under a license that permits archiving, the archive can ingest it without a deposit mandate. This is the procedural bet: build the collection first, and let legal frameworks catch up to the practice.

The approach has a cost. Crawling public repositories captures what is visible, not what is legally deposited. Proprietary code, internal research software, and code behind access controls stay out unless the holder chooses to contribute. The archive is comprehensive for open source and thin for everything else.

There is a related tension that this site has explored: Git repositories and lab notebooks diverge on reproducibility. A public repository is not the same as a preserved record. Repositories get deleted, force-pushed, or made private. The archive treats the public snapshot as the durable artifact.

Scale brings its own legal questions. The archive ingests millions of repositories, each with its own license and provenance. Some licenses permit redistribution; others are silent. The project navigates this by storing content for preservation and making it available under terms that respect the original license. That balance between preservation and rights is an ongoing negotiation, not a settled rule.

Crawling, Hashing, and Storing the World's Code

Crawlers ingest Git repositories at scale. The archive walks commits, trees, and blobs, pulling objects from public hosts and from repository mirrors. Each object is content-addressed. A file's identity is derived from its content, so the same file appearing in ten thousand repositories is stored once.

Every file is hashed into a persistent SWHID. The identifier encodes what the object is and how it is hashed, which makes it verifiable without trusting the archive. A researcher can recompute the hash from a local copy and confirm the SWHID matches. That property is what turns an archive into a citation mechanism.

Deduplication keeps storage costs manageable. Because identical content collapses to a single stored object, the marginal cost of archiving another repository that shares dependencies with thousands of others is small. The archive stores the graph of relationships, not just the bytes.

Origin and license metadata are preserved per snapshot. The archive records where a repository came from, when it was crawled, and what license the repository declared. That metadata matters for reuse. A SWHID without provenance is a hash; a SWHID with provenance is a citable record.

As of recent public reporting, the archive holds on the order of billions of unique source files and hundreds of millions of repository snapshots, with growth continuing as new hosts are added. That scale is what makes the legal deposit bet plausible: the collection already exists, so the policy debate is about recognition, not construction.

What the Archive Changes for Reproducibility

Citations can point to exact code versions. Instead of citing a project name and hoping the repository still exists, a paper can cite a SWHID that resolves to a specific snapshot. This is the same logic that made DOIs useful for articles, applied to source trees and commits.

Reviewers can verify builds from archived snapshots. If the code is preserved and the build environment is documented, a reviewer can reconstruct the software that produced a result. The archive does not solve environment drift on its own, but it removes the first failure mode: the code is gone.

Software joins data and papers as citable output. This site has argued that climate models adopt a seismology tool for ocean heat, a case where method travels across fields. Citable code makes that travel traceable. A method reused in a new field can be followed back to the version that was actually run.

Gaps remain for proprietary and ephemeral code. Internal pipelines, licensed dependencies, and short-lived scripts often never enter a public archive. A SWHID for the open part of a project does not cover the closed part. Reproducibility is bounded by what was archived, not by what was used.

Lessons for Scientific Computing Practice

Archive code at publication, not after. The moment a paper is accepted is the moment the code is most likely to still exist in the form that produced the result. Waiting for a grant renewal or a lab move increases the chance that the repository is gone.

Prefer persistent identifiers over URLs. A repository URL is a location; a SWHID is an identity. Locations change when hosting moves, accounts are renamed, or projects migrate. Identifiers survive those changes if the archive holds the content.

Document the build environment alongside source. The archive preserves code, not compilers, system libraries, or hardware. A build recipe, a container definition, or a dependency lockfile closes the gap between archived source and a runnable artifact.

Treat legal deposit as a model, not a mandate. The deposit analogy is useful for arguing that code deserves preservation. It is a poor template for compliance, because deposit law varies by country and rarely names software. Voluntary archiving is the practical path today.

Actions for Researchers and Institutions

Deposit code in Software Heritage with a SWHID. When a project is ready to publish, submit the repository and record the resulting identifier in the project's metadata.

Cite the exact snapshot in papers and data sets. Use the SWHID in the references section and in the data availability statement, so a reader can resolve the precise version that was run.

Ask libraries to include software in deposit policies. A library that already collects datasets is a natural home for a software deposit workflow, even if the legal mandate lags.

Budget for long-term code preservation. Storage and curation have costs, and a grant that funds a three-year project rarely funds a thirty-year archive. Line-item the preservation.

Teach archiving as part of methods training. A graduate student who learns to cite a SWHID alongside a dataset is more likely to build the habit into the next project than one who learns it after a failed rebuild.