SCIENCE
Code Archived, Result Rejected: How a Preprint's Verdict Turned on Provenance
A disputed preprint in computational biology went from promising to contested within months, and the deciding factor was not the statistics. It was provenance: whether the archived code recorded the environment in which the original result was produced. Four independent groups tried the same pipeline. Two got the effect, two did not. The split traced to procedural choices nobody had documented.
A Preprint That Would Not Reproduce
The original preprint described a pipeline that combined a public gene-expression dataset with a custom classifier, reporting a modest but clean separation between two sample classes. The authors posted code to a public repository, included a data-accession number, and invited reanalysis. Within weeks, four groups attempted it.
Two groups recovered the reported separation. Two did not, and their null results were not marginal. The discrepancy was large enough that the preprint's central claim became the subject of a public comment thread, then a formal dispute. None of the four groups had access to the original authors' compute environment, only to a snapshot of files.
What made the episode instructive was the timing. The preprint had not yet been peer reviewed, so the dispute unfolded in the open, on the repository and on the comment thread, rather than behind a journal's embargo. That visibility is unusual, and it left a detailed record of what each group actually did.
What Counts as Reproducible
Reproducibility means the same data and the same code yield the same result. Replicability means new data and a similar protocol yield a similar conclusion. The two are related but not interchangeable, and preprints often blur them in their abstracts, using the language of replication for what is really a reproducibility claim.
The distinction matters because it changes what a failure means. If a group cannot reproduce a result from the same inputs, the problem is likely in the code, the environment, or the documentation. If a group cannot replicate a result on new data, the problem may be in the original finding itself. The preprint's dispute was a reproducibility failure, not a replication failure.
Readers tend to collapse the two into a single verdict. A paper is reproducible or it is not. That framing hides the procedural questions that actually decide outcomes, and it makes the dispute look like a disagreement about truth rather than a disagreement about method.
The Archive Is Not the Code
A repository snapshot preserves files. It does not preserve a runnable environment. The two groups that failed did not lack the code; they had the same commit hash. What they lacked was the set of package versions, system libraries, and hardware details under which the original run had been executed.
Dependency drift is faster than most researchers assume. A package that changes its default numerical behavior between minor versions can shift a result without raising an error. Within weeks, a fresh install of the same requirements file can behave differently from the environment that produced the original figure.
Random seeds compound the problem. If a train-test split is generated without a documented seed, two runs of the same code can partition the data differently. The resulting performance difference may be small, or it may be large enough to look like a real effect. The preprint's methods section did not state a seed.
Provenance is the record that closes these gaps: who ran what, on which machine, with which versions, at which commit. Without it, an archive is a set of instructions with the operating conditions removed. This site has argued, in a related piece on data pipelines versus replications, that infrastructure spending often outpaces the work of checking whether a result holds.
Three Procedural Choices That Decided It
One group ran the notebook in a fresh container built from the requirements file. That group reproduced the effect. Their container pulled the current versions of each dependency, which happened to match the original environment closely enough that the classifier's output was stable.
Another group reused cached intermediate files from a prior run of a different project. Those files had been generated under an older version of the same preprocessing library, and the cached matrices differed subtly from what the pipeline expected. That group did not reproduce the effect, and their logs showed no error, only different numbers.
A third group varied the train-test split seed to test robustness. Under some seeds the effect appeared; under others it vanished. Their conclusion was that the original result was seed-dependent, which is a finding about the method rather than about the biology. None of these three choices appeared in the preprint's methods section.
The fourth group's approach was closest to the original authors' described procedure, and they reproduced the effect. The pattern across all four is not that two groups were careless. It is that the documented procedure was underdetermined, and each group filled the gap with a reasonable default.
Reading a Verdict by Its Provenance
Check whether the archive includes environment files. A requirements file is a start; a lockfile or a container image is better. If the archive contains only scripts, the environment is a reconstruction, and any reproduction attempt is testing the reconstructor's choices as much as the original claim.
Check whether random seeds are stated rather than implied. A seed in the code is not the same as a seed in the methods section. The methods section is what a reader will follow, and if it omits the seed, the reader will supply one.
Check whether a negative verdict names the failing step. A rejection that says which stage diverged, and under what conditions, is more useful than a vague pass. The two groups that failed here produced detailed logs, and those logs are what made the dispute resolvable rather than merely contentious.
A verdict without provenance is an opinion about a result. A verdict with provenance is a statement about a procedure.
The trade-off is real. Full provenance records are expensive to produce and maintain, and most preprint authors are not funded to build them. A lab that spends a week containerizing a pipeline is a week not spent on the next experiment. That cost is why provenance remains rare, and why disputes like this one keep recurring.
What to Do Before You Cite
- Ask the authors for the container image or a lockfile, not just the repository URL, and treat a refusal or a missing file as information about the claim's status.
- Re-run one figure end to end before trusting the rest, and compare the numerical output, not just the presence of a plot.
- Log your own environment when you attempt a re-run, including package versions and the seed you chose, so your attempt is itself reproducible.
- Treat a provenance gap as a finding worth reporting, not a nuisance to work around silently.
- Cite the archive version you actually ran, not the preprint's posting date, because the two can diverge after a revision.
Why Provenance Remains Rare
The incentives that produce provenance gaps are not mysterious. Journals rarely require a lockfile, funders rarely pay for one, and reviewers almost never test a container before signing off. The result is that provenance work falls to whoever cares most about the specific result, which is usually the person trying to reproduce it, not the person who produced it.
Some fields have moved faster than others. In parts of genomics and climate modeling, community norms now expect a container image or a workflow language file alongside the code, and a submission without one draws comment. In smaller subfields, the expectation has not taken hold, and a repository with a README and a requirements file still passes as generous. The gap between those norms is wide enough that the same researcher can be scrupulous in one project and silent in another.
There is also a measurement problem. No one tracks how often a computational result fails to reproduce because of an environment mismatch, as opposed to a coding error or a statistical flaw. The categories blur in practice, and the blurring makes it easy to attribute a failure to the most visible cause. A dispute like this one is unusual precisely because the four attempts were logged well enough to separate the causes.
None of this resolves the original biological question. It only establishes what the preprint's result depends on. For a reader deciding whether to build on the claim, that is the more actionable piece of information, and it is the one the dispute made visible.