Skip to content

Article image
The Molecular Clock: Dating Evolutionary Events

May 16, 2026 · Updated: May 25, 2026

Overview

The molecular clock hypothesis proposes that nucleotide or amino acid substitutions accumulate at a relatively constant rate over evolutionary time. If the rate is known, calibrated against the fossil record or geological events, the amount of sequence divergence between two lineages can be converted into an estimate of their divergence time. This insight transformed evolutionary biology by providing a quantitative framework for dating speciation events across the tree of life, including groups with sparse fossil records.

Key Concepts

A strict molecular clock assumes the same substitution rate for all lineages, which is often violated in real data. Relaxed clock models allow rates to vary among lineages, typically drawn from a lognormal or exponential distribution. Calibration involves anchoring nodes with dated fossils or known biogeographic events to convert relative times into absolute ages. Programs such as BEAST and MCMCTree implement Bayesian relaxed clock approaches that jointly estimate topology, rates, and divergence times while accounting for uncertainty in calibration points.

Practical Workflow

A typical molecular clock analysis begins with a sequence alignment of orthologous genes from the taxa of interest. The first decision is clock model selection: a likelihood ratio test comparing the strict clock against a relaxed clock model can determine whether rate heterogeneity among lineages is significant. For Bayesian relaxed clock analyses in BEAST, the input XML file is generated through BEAUti, where the user specifies the substitution model (e.g., GTR+G), the clock model (e.g., uncorrelated lognormal relaxed clock), and the tree prior (e.g., Yule process or birth-death). Calibration is performed by assigning prior distributions to node ages based on fossil evidence, for example, a lognormal prior with an offset equal to the fossil age and a standard deviation reflecting the uncertainty of the calibration. The MCMC chain is run for 50–200 million generations, sampling every 5,000–10,000 steps, and convergence is assessed by ensuring effective sample sizes exceed 200 for all parameters. Tracer software visualizes the posterior distributions of divergence times and checks for stationarity. After discarding burn-in, TreeAnnotator summarizes the posterior tree sample as a maximum clade credibility tree with mean node ages and 95% highest posterior density intervals. The resulting time-calibrated phylogeny reveals when key speciation events occurred, providing a temporal framework for evolutionary and ecological analyses.

Applications

Molecular clocks have revolutionized our understanding of evolutionary timescales. They date the origin of major clades, track the emergence of drug resistance in pathogens, and time viral pandemics. The method depends on reliable DNA sequencing data to measure divergence and is frequently applied in bacterial genetics to estimate the age of pathogenic lineages. In virology, molecular clocks calibrated by known collection dates help trace the emergence of novel strains, complementing viral structure and classification studies.