SCIENCE
Hippocampal Replay Study Shrinks Under Larger Samples
Hippocampal replay is real, but the human effect is roughly half the size early papers claimed. Larger, preregistered studies have pulled the numbers down without killing the phenomenon. What follows is how that happened, who measured what, and what a careful reader should check before trusting the next replay result.
A Finding That Shrank Under Scrutiny
Replay names a specific event: a sequence of place-cell activations from a waking path fires again during rest or sleep, compressed into tens of milliseconds. Rodent work in the 1990s established this in sharp-wave ripples, the 30–100 ms electrophysiological bursts that accompany idle states. The human version arrived later, inferred from EEG–fMRI coupling rather than implanted electrodes, and it arrived with big numbers attached.
Early human reports often described effects around d = 0.8, the kind of signal that feels solid in a sample of 15 to 20 participants. When independent groups ran larger cohorts with preregistered analysis plans, the same comparisons landed nearer d = 0.4. The direction held. The magnitude did not.
The tension is not whether replay exists. It is whether the human measures are precise enough to carry the theoretical weight placed on them. A halved effect size changes power calculations, changes which individual differences are detectable, and changes how confidently anyone can claim a task manipulation boosted or suppressed replay.
How Replay Was First Measured
Wilson and McNaughton's rodent recordings in the mid-1990s gave the field its founding observation: cells that fired together on a track fired together again during subsequent sleep, in the same order but faster. The measurement depended on implanted electrodes, stable single-unit isolation, and animals that would sleep in a recording chamber. Those constraints kept sample sizes small by design.
Human work traded electrodes for scalp EEG and fMRI, which meant trading precision for access. Replay-like reactivation was inferred from transient hippocampal activity coupled with coordinated changes across distributed networks. The inference is indirect. A scalp sensor cannot see a 30–100 ms sequence in CA1 directly, so the analysis pipeline carries much of the argument.
That pipeline became the soft spot. Different groups chose different classifiers, different time windows, different definitions of a replay event. A finding that depends on pipeline choices is a finding that needs independent pipelines to confirm it. For much of the 2000s and 2010s, few independent pipelines existed.
The Replication That Changed the Picture
Between roughly 2023 and 2024, multi-lab efforts began reporting on larger cohorts with analysis plans fixed before data collection. The pattern across these studies was consistent: effects in the predicted direction, smaller than the original reports, and sometimes below the threshold the original authors had treated as meaningful.
Statistical power explains part of the gap. A study with 18 participants and a true effect of d = 0.4 has low odds of detecting anything, so the studies that got published were disproportionately the ones that happened to land high. Publication bias compounds this: null and small-effect results sat in file drawers while the striking ones circulated.
None of this required misconduct. It required ordinary research practice under ordinary incentives, applied to a measurement that was noisier than anyone wanted to admit. The replication did not overturn replay. It recalibrated what a human replay study can credibly detect.
What Shaped the Verdict
Preregistration and open data requirements did most of the work. When analysis choices are fixed in advance and the data are posted, the space for post hoc flexibility narrows. Bayesian reanalyses of the original datasets often found evidence weaker than the frequentist claims suggested, with Bayes factors that supported a smaller effect or remained inconclusive.
Task differences matter too. Spatial replay, where an animal or person retraces a path, and reward-related replay, where sequences favor goal locations, are not the same phenomenon and do not replicate at the same rates. Pooling them inflates apparent consistency. A related piece on this site, Code Archived, Result Rejected, traces how provenance questions can decide a preprint's fate when the analysis pipeline is doing the arguing.
Funding sources were not the lever here. Replay research has little industry money attached. The pressure came from career incentives: a first-author paper in a high-visibility journal rewards a striking effect, and no promotion committee rewards a careful null. That incentive operates regardless of who pays.
Why the Effect Size Matters for Theory
A halved effect size is not a minor correction. It changes which theoretical claims are testable. If the true effect is d = 0.4, then a study needs on the order of 100 participants to have decent power, and most human replay studies have run far fewer. That arithmetic makes many published claims unfalsifiable in practice: they were never powered to detect the effect they now report.
Consider the claim that a specific task manipulation—say, reward timing—boosts replay. With a small sample, a significant result could reflect a true boost or just noise. With a larger sample, the same manipulation might show a boost of, say, 0.2 standard deviations, which may or may not matter for memory consolidation. The theoretical weight placed on the manipulation shrinks along with the number.
This is not an argument for abandoning small studies. It is an argument for reading them as exploratory. A small study can generate a hypothesis; it cannot confirm one. The field's early literature mixed the two roles, and the replication wave is sorting them out.
The Trade-off Between Precision and Access
The choice between rodent and human replay work is not a matter of one being better. It is a trade-off between precision and access. Rodent studies offer direct access to single-unit activity and can resolve replay sequences with millisecond precision. They can test causal manipulations—optogenetic silencing of sharp-wave ripples, for instance—that are impossible in humans. But they cannot tell us about human memory in its natural context, and the species gap leaves open whether the same mechanisms operate.
Human studies, by contrast, can link replay-like signals to behavior, individual differences, and clinical outcomes. They can ask whether replay during rest predicts later recall in a way that matters for education or rehabilitation. But they rely on indirect measures and cannot isolate the precise circuit events that rodent work can. The replication crisis in human replay is partly a consequence of this trade-off: the measures are noisier, so effects are smaller and harder to pin down.
The field is learning to live with this trade-off rather than pretending it away. Some groups now combine both approaches, using rodent findings to constrain human hypotheses and human data to test ecological validity. That kind of cross-species triangulation is slow and expensive, but it may be the only way to build a cumulative science of replay.
Lessons for Circuit Neuroscience
Small samples inflate effects across subfields, not just replay. The same arithmetic applies to any hippocampal measure where individual variability is large and the outcome is a difference score. Replication is not rejection. It is calibration, and calibration is what lets a field build on a result instead of around it.
Anatomical distinctions carry weight that pooled analyses hide. CA1, CA3, and the dentate gyrus play different roles in sequence generation and pattern separation, and a replay measure that averages across them may be measuring a mixture. Rodent work can resolve this with electrodes. Human work mostly cannot.
The working consensus is narrow and defensible: replay exists in rodents and has plausible human analogues; the magnitude of the human effect is uncertain and task-dependent; and the field's early numbers were too large. A related piece, Optogenetics Found Its Funding, shows how a technique's trajectory can be shaped by budgets rather than by replication alone.
Practical Steps for Researchers and Readers
Check the sample size before you trust the effect. A human replay study with fewer than 30 participants and a large reported effect deserves a power calculation before it deserves citation.
Look for preregistration and open data as a pair. One without the other leaves room for the pipeline to absorb the flexibility that preregistration was meant to remove.
Read replication attempts in independent cohorts as the real evidence. A result reproduced by the original lab using the original pipeline is a consistency check, not a replication.
Hedge the claim when you cite it. Replay likely supports memory consolidation; the effect size varies by task and by measure, and the honest citation says so.
Follow multi-lab consortium updates rather than single-lab papers. The aggregate is slower to publish and harder to spin. A related piece, Physicists Import a Spin Glass Tool, shows what happens when a method crosses fields and meets new noise.