SCIENCE

How Evolutionary Biologists Date a Trait Without a Fossil

Most of what makes an organism interesting leaves no fossil. A bird's egg-rejection behavior, a frog's call, a plant's flowering time: none of it mineralizes. So when a paper says a trait evolved roughly 30 million years ago, where does that number come from? It comes from a chain of procedural choices, and understanding those choices is how you decide whether to believe the date.

The problem with trait origins

The fossil record documents hard parts. Bones, shells, teeth. Soft tissue decays long before mineralization can preserve it, and behavior leaves no physical trace at all unless it happens to alter a bone or a burrow. This means the traits biologists most want to date, such as courtship displays or immune strategies, are exactly the ones the rock record cannot speak to.

Indirect evidence becomes necessary. Biologists reconstruct history from the traits of living species, from DNA sequences, and from the shape of the evolutionary tree that connects them. Each of those sources carries its own assumptions, and the assumptions do the heavy lifting.

The result is that a date for a trait origin is a model output, not an observation. That is not a flaw, but it does mean the date is only as good as the model. Reading such a paper well requires knowing which model produced the number.

Building a molecular clock

Émile Zuckerkandl and Linus Pauling proposed in the 1960s that biomolecules accumulate changes at roughly steady rates over time. If that holds, the number of differences between two DNA sequences becomes a measure of how long ago their lineages split. The technique is now standard in evolutionary biology.

Calibration is where the clock meets the rocks. A molecular clock must be anchored to at least one divergence with a known fossil date, and the estimate for everything else scales from that anchor. Use a different calibration point and the whole timeline shifts.

Rates also vary. Some lineages tick faster than others because of generation time, metabolic rate, or DNA repair efficiency. A single global rate applied across a tree can therefore produce dates that are systematically wrong in one direction for one group and the other direction elsewhere.

This is why modern analyses use relaxed clocks, which allow rates to drift across branches. Relaxed clocks are more flexible but also more parameter-rich, and a model with more parameters can fit noise as easily as signal. The trade-off is real and does not resolve cleanly.

Phylogenetic comparative methods

Comparing species as if they were independent data points is a statistical error. Closely related species share traits through common descent, so a sample of fifty songbirds is not fifty independent draws. Phylogenetic comparative methods exist to correct for this shared ancestry.

Joseph Felsenstein's 1985 paper on independent contrasts gave the field a workable correction. The method transforms trait values along the branches of a phylogeny so that the resulting contrasts are statistically independent, which then permits ordinary regression and correlation tests.

Model choice matters. Whether you assume traits evolve by Brownian motion, by an Ornstein-Uhlenbeck process that pulls toward an optimum, or by some early-burst model will change the correlations you recover. A related piece on this site has argued that pipeline choices shift results in brain imaging findings, and the same logic applies here.

A worked example: avian brood parasitism

Some cuckoo species lay their eggs in the nests of other birds. The host then raises the parasite's chick. In response, many hosts have evolved the ability to recognize and reject foreign eggs, and the parasite in turn evolves eggs that mimic the host's. This is an escalating coevolutionary cycle.

To date the escalation, biologists map egg-rejection behavior and egg mimicry onto a host-parasite phylogeny. If rejection appears in a clade of hosts that share a recent common ancestor with a parasitized lineage, the defense can be assigned to that branch. The date of the branch then becomes the date of the trait.

The trouble is ancestral state reconstruction. If a host species rejects eggs and its sister species accepts them, the reconstruction of what the ancestor did depends on the model. Different models can place the origin of rejection on different branches, moving the date by millions of years.

Uncertainty in these reconstructions is often reported as a probability distribution over ancestral states. A paper that reports only the most likely state, without the distribution, has thrown away the part a careful reader needs most.

Where dating methods fail

Rate variation across lineages is the largest single source of error. If a group has experienced a burst of molecular evolution, its branch lengths will be inflated, and any date anchored to those branches will be too old or too young depending on the direction of the correction.

Incomplete taxon sampling biases estimates in a predictable way. Drop the species that lack the trait, and the reconstruction of the ancestral state shifts toward the sampled majority. The missing taxa are often the ones that would have told you the trait is older than you think.

Convergent evolution mimics shared ancestry. If two distant lineages independently evolve the same trait, a naive analysis will treat it as a single origin and place it too far back on the tree. Distinguishing convergence from homology requires dense sampling and careful model comparison.

Fossil calibrations remain sparse for most groups. Many clades have only one or two usable fossils, and those fossils are often fragmentary. A single calibration point carries the entire timeline, which is a lot of weight for one jawbone.

Practical steps for evaluating claims

Check the calibration points first. A dated trait origin is only as trustworthy as the fossils anchoring the clock, so look for how many were used and whether their placements are justified in the supplement.

Ask whether the phylogeny is fully resolved. If the tree has polytomies, the analysis cannot place a trait origin on a specific branch, and any single date is an average over possibilities.

Look for sensitivity analyses. A careful paper will rerun the analysis under different clock models, different tree topologies, and different calibration schemes, then report how much the date moves. This site has noted that ice core dust layers recalibrate radiocarbon dating, and the principle of testing calibration choices against alternatives applies across dating methods.

Demand ancestral state uncertainty. A single reconstructed state is a point estimate, and point estimates without intervals invite overconfidence. If the paper gives a probability distribution, read the tails.

Treat any single-date estimate as provisional. The honest output of these methods is a range, often wide, and the range is the finding. A number without a range is a number you should hold loosely.

How dates get reported

Even when the analysis is sound, the number that reaches a news story has usually been rounded, stripped of its interval, and detached from the model that produced it. A study might report that a trait originated in the Miocene, roughly 5 to 23 million years ago, and the press release may convert that into a single phrase like "10 million years ago." The rounding is not dishonest, but it hides the width of the estimate. When you see a trait date, ask whether the original paper gave a range and whether the range was carried through to the summary.

Another reporting habit is to treat the date as a property of the trait itself rather than of the analysis. A trait does not have an age the way a rock has an age. It has an inferred age under a particular model, tree, and calibration set. Change any of those and the number changes. Papers that present a single date without qualification are making a stronger claim than the methods can support.

What would make these dates firmer

Two developments would narrow the intervals more than any refinement of current models. The first is denser fossil coverage for soft-tissue-adjacent structures. Feathers, eggshells, and stomach contents occasionally preserve, and each new such fossil constrains a node that previously floated. The second is genomic sampling of populations rather than single individuals per species. When a species is represented by one genome, the analysis cannot see the variation that would let it distinguish a recent origin from an old one that has since spread. Population-level sampling turns a point on a tree into a distribution, and the distribution is what the dating methods actually need as input.

Until then, the practical stance is to read trait dates the way you would read a weather forecast: useful, conditional, and best consumed with the uncertainty attached. The methods are not broken. They are doing exactly what they were designed to do, which is to infer history from the only evidence available. The mistake is to forget that the inference is the product, not the history itself.