SCIENCE
Git Repositories and Lab Notebooks Diverge on Reproducibility
Every computational result leaves two paper trails. One is the version-control history: commits, branches, tags, diffs. The other is the lab notebook: dated entries, parameter values, reasoning, dead ends. Reproducibility depends on both records agreeing, and in practice they diverge in predictable ways. This piece explains what each artifact actually captures, where the mismatch bites, and which habits close the gap.
The Reproducibility Gap in Computational Science
Reproducibility means a different researcher, given the same inputs and the same method, obtains results consistent with the original. In computational work that requires three things to survive: the code, the data, and the environment it ran in. None of the three is fully captured by either a repository or a notebook alone.
Version control tracks how code changed over time. A lab notebook tracks what the researcher intended and observed. The first is mechanical and complete about text files. The second is interpretive and selective. Neither was designed as a provenance system, and treating either as one produces the familiar failure mode: a published figure whose generating command nobody can reconstruct.
Surveys of research software practices have repeatedly found that a substantial share of published computational work cannot be rerun from the materials provided, though estimates vary widely with field and definition. The gap is not mainly about secrecy. It is about artifacts that record different things than readers assume.
What Git Histories Actually Record
A git commit records a snapshot of tracked files, an author, a timestamp, and a message. That is the whole payload. The message is free text, and in research repositories it is frequently a single word or a terse fragment. Nothing in the format forces the commit to say why a parameter changed or which run it supported.
Branches carry information too, but of an awkward kind. An abandoned branch often marks an analysis path that was tried and discarded, which is exactly the context a replicator needs. Most repositories delete those branches, or never push them, so the negative results vanish while the final code survives.
Diffs show what changed between two states, not the rationale behind the change. Merge conflicts, when they appear, expose coordination gaps inside a team: two people editing the same analysis script because no one owned it. Those structural signals are visible in the history, but only to someone who reads it as a record of human decisions rather than a file ledger.
One more limit matters. Git was built for source code, and it handles large binary inputs poorly. Model weights, satellite granules, and simulation outputs usually live outside the repository, sometimes with no record of which version of the data fed which run.
The distributed design of git also creates a subtle provenance problem. When a collaborator clones a repository, they receive the full history, but only up to the point of the clone. If they later contribute commits back, those commits may carry a different author identity and timestamp than the original work. The resulting history is a patchwork of local decisions, not a single authoritative timeline. For a replicator, this means the commit graph they see may not reflect the sequence in which experiments were actually run.
Lab Notebooks as Partial Evidence
A well-kept notebook captures intent: the hypothesis, the parameter sweep that was planned, the reason a run was repeated. That is precisely what a commit message omits. The problem is that notebooks are written by hand, often after the fact, and rarely to a schema that a machine can parse.
Entries frequently omit the values that matter most for replication: exact software versions, library dependencies, random seeds, and hardware details such as GPU model or thread count. A notebook may say a simulation was run in March; it will seldom say which build of the solver, compiled with which compiler flags.
Dates drift. An entry written on Monday may describe a run launched the previous Friday, and the notebook date is what gets cited. Cross-referencing a notebook page to a specific commit hash is rare, which means the two records cannot be joined without guesswork.
There is a real trade-off here. Notebooks are fast and low-friction, which is why people keep them. Imposing a rigid template on every entry slows the exploratory work that notebooks exist to support. The cost of that freedom shows up later, at replication time, when the person paying it is someone else.
In some fields, the notebook is not a bound volume but a collection of loose files: text documents, spreadsheets, and scanned pages stored in a shared drive. This fragmentation multiplies the points of failure. A single missing file can break the chain of evidence for an entire experiment, and there is no version history to recover it.
A Worked Example from Climate Modeling
Model intercomparison projects, the coordinated experiments in which many modeling centers run the same protocol, are a demanding test case. Participants must report results for a defined set of experiments, which requires knowing the exact code state that produced each submission.
In practice, the code state is often identified by a version number or a release name rather than a commit hash. If the repository has no tag at that point in its history, the mapping from published run to source code is ambiguous. Later commits, including bug fixes, sit between the tag and the actual run.
Notebooks from these projects record boundary conditions inconsistently. One group writes down the forcing dataset and its version; another notes only the experiment name. Replication attempts then fail on undocumented flags: a switch that changes aerosol treatment, a compile-time option, or a configuration file edited outside the repository.
The scale of these projects adds another layer. A single intercomparison may involve dozens of modeling centers, each with its own naming conventions and repository layout. Even when everyone uses git, the tags and branches are not harmonized. A run labeled "v1.2" at one center might correspond to a completely different set of code changes than "v1.2" at another. Without a shared identifier, cross-center replication becomes a manual detective exercise.
This site has covered related methodological strain in climate model ocean heat estimates and in proxy disagreements on interglacial warming, where small analysis choices move the headline number. The same fragility applies to the code that computes those numbers.
Bridging the Two Records
The most direct bridge is the executable notebook: a document that interleaves narrative, code, and stored output, and that can be rerun end to end. It carries the notebook's explanatory value and the code's executability in one file. The catch is that such documents rot quickly unless their environment is pinned.
Persistent identifiers help on the archiving side. A repository can be deposited with a data repository and receive a DOI that resolves to an immutable snapshot, so a cited version stays fixed even as the working repository moves on. Several major archives offer this for software, and some journals now require it.
Automated provenance capture is the other lever. Tools that log the environment, package versions, and command line at run time produce a record no human has to remember to write. The output is machine-readable, which makes it comparable across runs in a way handwritten entries are not.
Versioned environments, whether container images or lockfiles, turn an implicit dependency set into an explicit one. A related piece on this site describes how archived plate data moved into genomics pipelines, a case where the analysis environment had to be reconstructed years after the observations were made.
In practice, the most robust bridges combine several of these approaches. A DOI-archived snapshot with a lockfile and an automated provenance log covers code, environment, and execution context. The notebook then serves its original purpose: explaining the reasoning that no automated system can infer.
Practical Steps for Reproducible Workflows
Tag every release that supports a publication with a semantic version, and cite the tag alongside the commit hash in the paper's data statement. A tag alone is not enough if the branch moves afterward.
Archive the tagged code and the accompanying notebook in a persistent repository such as Zenodo, which mints a DOI for the deposit, so the cited version cannot silently change.
Record exact package versions in a lockfile, and commit that lockfile with the analysis code rather than generating it at run time on a different machine.
Link each substantive notebook entry to the commit hash of the code it describes, even if the link is a single line at the top of the page. This is the cheapest step and the one most often skipped.
Test the whole thing on a clean machine before submission. A fresh environment, the archived snapshot, and the documented commands will surface missing dependencies faster than any review.
Finally, treat the two records as complementary rather than redundant. The git history answers what changed; the notebook answers why. When they are cross-referenced, replication becomes a matter of following links, not reconstructing lost context.