Genetics and Ancestry

The Fossil Record showed what physical remains can and cannot establish about historical transitions. This chapter turns to a different, complementary body of evidence: what genomes themselves reveal about ancestry. Genetic evidence for common descent does not depend on fossilization at all, which is exactly what makes it a valuable independent check on the fossil-based case built in the previous chapter — and, because genomes are so information-rich, it is also where the strongest technical challenges to common descent live.

By the end of this chapter you should be able to:

Nested Hierarchy

Living things can be classified into groups nested inside larger groups — species inside genera, genera inside families, and so on — such that almost every organism fits cleanly into exactly one place in a single branching pattern, rather than needing to belong simultaneously to several unrelated categories. This nested structure is itself an observation, independent of any particular explanation for why it exists.

Common descent explains the nested pattern directly: each branch point represents a real historical population split, so features that arose before a split are shared by everything descended from it, and features that arose after a split are restricted to that branch alone. Simple similarity between two organisms is comparatively weak evidence on its own, because similar function can produce similar structure independently of shared history. The strongest evidence for a genuinely nested (rather than merely convenient) classification comes from features that are unlikely to be explained by function alone — particularly historically contingent details that a shared ancestor would pass down but that convergent function has no obvious reason to reproduce independently. Chromosome 2 and endogenous retroviral insertions, covered next, are the strongest examples of that kind of evidence available in this guide.

Homology

Homology is similarity that exists because of shared inherited ancestry, as distinct from analogy, similarity that exists because of shared function or shared environment despite independent origins. A bird wing and a bat wing are homologous as forelimbs (both inherited the basic tetrapod limb plan from a shared ancestor) while being only analogous as flight structures (flight itself arose independently in each lineage). This distinction matters because similarity can arise from more than one cause: shared ancestry, common function, convergent evolution, or shared developmental or physical constraints can all produce resembling structures.

Because similarity is not uniquely diagnostic of ancestry, homology claims are strongest when supported by evidence beyond gross anatomical resemblance — underlying developmental pathways, positional relationships to other structures, and, where available, the genomic evidence discussed throughout the rest of this chapter. Anatomical resemblance alone is a starting hypothesis, not a conclusion.

Convergence

Convergent evolution can produce similar traits, and even similar molecular solutions, in unrelated lineages independently, without shared ancestry of the trait itself. This is a real and well-documented phenomenon, and it is the main reason simple similarity is treated as weak evidence throughout this chapter: “these two organisms resemble each other” is not, by itself, a strong argument for common descent, because convergence offers an alternative explanation for many kinds of resemblance.

Convergence is a much stronger alternative explanation for adaptive, functionally driven similarity — a streamlined body shape in unrelated aquatic animals, for instance, where the same physical constraints of moving through water plausibly favor the same solution more than once — than it is for similarity in features with no obvious functional reason to recur independently in the same specific form. That distinction is exactly why the chromosome 2 and endogenous retrovirus evidence below carries more weight than anatomical resemblance: those features are not merely similar, but historically contingent in ways convergence has no obvious mechanism to reproduce at matching genomic locations.

Human Chromosome 2

Humans have one chromosome that corresponds, broadly and end to end, to two separate chromosomes found in other great apes. If two ancestral chromosomes fused, three specific signatures would be expected at the fusion site: correspondence to both ape chromosomes, internal telomere-like sequences at the fusion point (telomeres normally occur only at chromosome ends), and remnants of a second, now-unused centromere. Researchers identified internal head-to-head telomeric repeats on human chromosome 2 surrounded by sequences characteristic of chromosome ends Human chromosome 2 fusion site, and later work identified remnants associated with a degenerated second ancestral centromere Degenerate ancestral centromere on human chromosome 2.

This is strong evidence for an ancestral chromosome-fusion event consistent with shared human-ape genomic history, because it is not merely that human chromosome 2 resembles two ape chromosomes — it is that chromosome 2 carries the specific, functionally unnecessary structural scars a real fusion event would be expected to leave behind, in exactly the location a fusion model predicts.

Human chromosome 2 does not, by itself, prove whale evolution, the fish-to-tetrapod transition, universal common ancestry, or abiogenesis. It is strong evidence for a specific, well-defined claim — a shared human-ape ancestral fusion event — not a general-purpose proof of every broader ancestry claim in this guide. Each broader claim needs its own supporting evidence.

Endogenous Retroviruses

Retroviruses can insert copies of their genetic material into a host's genome; if that insertion happens in a germline cell, it can become a permanently inherited feature of the host lineage's genome, passed down at the same genomic location in every descendant. Humans and other primates share numerous endogenous retroviral insertions at corresponding genomic positions, and comparative studies also identify lineage-specific insertions that follow the expected primate branching pattern Primate endogenous retroviral insertion patterns.

The reasoning is analogous to manuscript genealogy: shared content between two documents could have many explanations, but a shared unusual copying error, in the same location, in multiple copies, is much stronger evidence of a shared exemplar those copies both descend from. A retroviral insertion at one specific position in the genome is exactly this kind of unusual, historically contingent “copying mark” — not something a designer or convergent process would have any obvious functional reason to reproduce at the same genomic coordinate in an unrelated lineage.

Shared Neutral Mutations

A neutral mutation is one with no meaningful effect on fitness — it is neither favored nor disfavored by selection, and its fate is governed largely by chance (genetic drift) rather than adaptive advantage. Two unrelated lineages sharing the exact same neutral mutation, at the exact same genomic position, is comparatively hard to explain by anything other than inheritance from a shared ancestor who carried that mutation, because there is no selective pressure that would make convergent evolution independently reproduce a functionally inconsequential change at that specific site.

This is part of the same broader pattern as the retroviral-insertion evidence above: precisely because a neutral mutation does nothing useful, common function offers no explanation for why it would recur independently. Shared neutral mutations at matching positions are therefore treated as some of the strongest available evidence for shared ancestry, alongside chromosome rearrangements and retroviral insertions — all three are historically contingent markers that common descent explains directly and that convergent, function-driven processes have no obvious reason to reproduce.

Shared Genomic Rearrangements

Genomes are occasionally rearranged by large-scale structural events — fusions, fissions, inversions, and translocations of chromosome segments — rather than only by point mutations. The human chromosome 2 fusion described above is the clearest single example in this guide, but the same logic applies more generally: a specific, structurally unusual rearrangement shared by two lineages, at the same genomic location, is difficult to explain by anything other than descent from a common ancestor in whom that rearrangement first occurred and then was passed down.

Genomic rearrangements are useful evidence for the same reason retroviral insertions and neutral mutations are: they are rare, historically contingent events rather than recurring, predictable outcomes of shared function, so their recurrence at matching locations across lineages is not well explained by convergence.

Synteny

Synteny is the conservation of gene order and arrangement along chromosomes across different species. Genes that sit next to each other, in the same order, on the same relative chromosomal position in two different species are more plausibly explained by both species having inherited that arrangement from a shared ancestor than by two entirely separate origins independently landing on the identical gene order by chance or by function alone.

Synteny is also a practical tool, not only an evidential argument: because related genomes tend to preserve large blocks of conserved gene order, syntenic comparison is one of the standard methods researchers use to identify corresponding chromosomal regions across species — including, in the chromosome 2 case above, confirming which two ape chromosomes the fused human chromosome corresponds to.

Homoplasy

Homoplasy is similarity that is not due to shared inheritance — it covers both convergent evolution (independent origin of similar traits under similar pressures) and cases where two lineages independently acquire what looks like the identical marker by coincidence. Even markers normally treated as strong evidence for shared ancestry, such as retroviral- or transposon-style insertions, can in principle be homoplastic: two independent insertion events could, rarely, land at the same genomic site by chance.

This possibility has actually been tested rather than merely acknowledged. One primate study estimated insertion-site homoplasy for a particular class of transposable-element insertion (LINEs) at about 0.52%, and found no independent insertion events at all among the specific orthologous human-chimpanzee sites it surveyed LINE insertion-site homoplasy in primates. The honest conclusion is two-sided: a single shared insertion is not, strictly, logically irrefutable proof of common ancestry on its own, since a rare coincidental match remains conceivable — but many independently matching insertions across many genomic positions become correspondingly much stronger evidence, because the chance of coincidental homoplasy compounds unfavorably across dozens or hundreds of independent matching sites.

Incomplete Lineage Sorting

An ancestral population is not genetically uniform: it can carry multiple different genetic variants at a given locus simultaneously. If population splits happen in relatively quick succession, different descendant lineages can end up inheriting different ancestral variants by chance, so that the genealogical history of that one locus does not match the overall species branching order. This is incomplete lineage sorting, and it is a well-understood, expected consequence of population genetics operating during rapid successive speciation — not an anomaly requiring a special explanation each time it appears.

The gorilla genome provides a documented real-world example: although the dominant, genome-wide species relationship groups humans and chimpanzees together to the exclusion of gorillas, roughly 30% of the genome locally groups gorilla with either humans or chimpanzees instead Gorilla genome and incomplete lineage sorting, exactly the pattern incomplete lineage sorting predicts when population splits occur close together in time.

Conflicting Gene Trees

Different individual genes can, and regularly do, produce different branching trees from one another, for several distinct reasons covered across this chapter: incomplete lineage sorting, horizontal gene transfer, recombination, hybridization, and gene duplication and loss can each cause one gene's history to diverge from the overall species history. This is real, and it means that the overly simple assumption “every gene must share exactly the same genealogical tree” is false as a strict rule.

Conflicting gene trees do not automatically falsify common descent. They falsify the overly simple assumption that every stretch of DNA must have precisely the same branching history — not the underlying claim that the species themselves share common ancestry. The relevant question for any specific case of gene-tree conflict is whether the discordance fits the quantitative expectations of a known process such as incomplete lineage sorting, or instead systematically contradicts the proposed ancestry in a way none of the known processes can account for.

Hybridization

Distinct lineages that have already diverged can sometimes still interbreed, at least partially, exchanging genetic material (introgression) without fully merging back into one population. Documented cases in this guide include extensive historical hybridization among tropical eel species that nonetheless remain distinguishable lineages over millions of years, and gene flow between Heliconius butterfly lineages that homogenizes most of the genome while a smaller portion retains the specific traits responsible for ecological and reproductive isolation.

Hybridization complicates the picture of ancestry as a single, cleanly branching tree in the same general way incomplete lineage sorting and horizontal gene transfer do: it means some genetic material can cross between branches after they have already split, rather than being inherited strictly along one path. It does not mean the branches themselves are not real, historically distinct lineages — the tropical eel and Heliconius cases both document lineages that remain distinguishable over long timescales despite this genetic exchange, not lineages that dissolve back into a single population.

Horizontal Gene Transfer

Horizontal gene transfer is the movement of genetic material between organisms other than by ordinary parent-to-offspring inheritance — common in microorganisms, which can exchange genes across lineages that are otherwise only distantly related. Because different genes can each have their own independent horizontal-transfer history, Dagan and Martin argued that a strictly bifurcating universal tree may accurately describe only a small fraction of microbial genome histories, an argument often summarized as the “tree of one percent” Tree of one percent critique. Subsequent work has documented extensive lateral gene transfer across prokaryotic evolution more broadly Lateral gene transfer in prokaryotic evolution.

Horizontal gene transfer strongly challenges a simple model in which every gene, in every organism, follows the same clean branching tree — early evolutionary history, especially among microorganisms, may have been considerably more network-like than a single tree diagram can represent.

Horizontal gene transfer does not necessarily challenge shared ancestry itself. A population of ancestral organisms could share descent, exchange genes extensively with each other, and still produce different individual histories for different genes — a better model of early evolution may therefore be a genuinely tree-like population history with extensive network-like gene exchange layered on top of it, rather than either a pure tree or a pure network alone.

Universal Common Ancestry

Universal common ancestry is the claim that all presently known cellular organisms ultimately trace their genetic heritage to a common ancestral population. The strongest evidence for it is deep molecular similarity shared across all three domains of cellular life — Bacteria, Archaea, and Eukarya — including DNA as hereditary material, RNA intermediates, ribosomes, a broadly shared genetic code, ATP-based energy chemistry, homologous proteins, and overlapping metabolic pathways. A comparative analysis of 23 conserved proteins across 45 organisms recovered a broad phylogenetic signal linking all three domains Conserved proteins across domains.

Rather than simply assuming that sequence similarity implies ancestry, Douglas Theobald built explicit statistical models comparing universal common ancestry against separate-ancestry alternatives using universally conserved proteins; under the tested models, universal common ancestry was favored overwhelmingly Theobald statistical test of universal common ancestry. That formal test was itself methodologically challenged: Yonezawa and Hasegawa argued the approach could favor common ancestry partly because aligned protein-coding sequences already contain correlations that need not originate from common ancestry, and demonstrated cases where apparently unrelated sequence families could still cause the method to prefer a common-origin model Critique of Theobald's universal-common-ancestry test. Theobald responded that the counterexample itself introduced correlations through how the coding sequences were aligned Theobald response to methodological critique, and a later independent methodological analysis concluded that an assumption-free formal proof of universal common ancestry had not been achieved — since some common-ancestry signal is already embedded in sequence similarity and alignment choices — while affirming that the broader comparative-genomic evidence for common ancestry remained very strong Analysis of formal tests of universal common ancestry.

This methodological debate shows that one particular statistical method may not constitute an assumption-free proof of universal common ancestry. It does not show that universal common ancestry has been disproven: criticizing a specific proof or methodology is a different claim from disproving the underlying hypothesis, and the broader comparative-genomic case remained strong even in the assessment of the critique's own authors.

LUCA

LUCA stands for the Last Universal Common Ancestor: the most recent population from which all presently known cellular life descends. LUCA does not necessarily mean the first life to exist, a single individual cell, or the only origin of life that ever occurred — each of those is a separate, stronger claim than universal common ancestry itself makes.

LUCA is better understood as a population, a genetic community, or a network of early cells exchanging genes, rather than one individual organism with a single, fully resolved genome. Gene histories at this depth are complicated by horizontal gene transfer, duplication, loss, endosymbiosis, and recombination, so not every gene in a modern organism was necessarily vertically inherited from LUCA in a simple, direct line. Universal common descent also cannot easily detect whether life originated more than once: if life arose independently in multiple lineages and all but one went extinct, all currently known life could still trace to a single common ancestor without that meaning life itself only originated once. Universal common descent, in other words, begins after heritable biological systems already exist; it is a claim about the ancestry of existing life, not an origin-of-life theory, a distinction developed further in Abiogenesis.

Tree versus Network Models

A strictly branching tree, in which every lineage splits cleanly and never rejoins, is a useful simplification for many ancestry patterns — but, as the sections above on incomplete lineage sorting, hybridization, and horizontal gene transfer show, it is not a complete description of how genetic history actually moves through populations. A network model, which allows branches to exchange material and rejoin, can represent horizontal transfer, hybridization, and gene-tree conflict that a pure tree cannot.

Diagram comparing a branching-tree model of the history of life, described as a useful approximation for many ancestry patterns, against a network model with gene exchange, described as able to represent horizontal transfer, hybridization, or conflicting histories.
A strictly branching tree and a network model both describe real features of evolutionary history at different scales; the appropriate model depends on how much horizontal exchange affected the lineages and time depth in question.

Neither model is simply “correct” on its own. Within well-studied groups such as great apes, a tree with a modest, well-quantified amount of local gene-tree discordance (from incomplete lineage sorting) is a good approximation. Among microorganisms, and potentially near the earliest history of life, extensive horizontal gene transfer may make a network a substantially better description. A useful synthesis available in this guide is a genuinely tree-like population history — real, historically distinct lineages that trace to shared ancestral populations — with network-like gene exchange layered on top of it at varying intensity, rather than treating “tree” and “network” as mutually exclusive models competing for the single correct description of all of life's history.

Common Descent versus Common Design

Common descent explains shared biological features as the result of inheritance from shared ancestral populations. An alternative model, common design, interprets the same similarity as reuse by a designer rather than inheritance — the way an engineer might reuse a proven authentication library, database library, or networking module across otherwise unrelated software systems, rather than writing each one from scratch. Biologically, a designer under this model might be expected to reuse DNA, ATP, ribosomes, protein domains, or developmental gene toolkits across separately originated organisms.

Table comparing common descent and common design as explanatory models across four questions: core proposal (lineages share ancestry vs. features reflect intelligent causation), expected patterns (nested hierarchy, homology, shared changes vs. functional unity, reuse, possible design signatures), and each model's explanatory strength.
Common descent and common design are both coherent explanatory models; the comparison that actually distinguishes them is which one generates independent, testable predictions about specific patterns in the data, not which one is merely compatible with the data.

The two models are not simply “evolution versus creationism.” Intelligent Design and common descent can coexist: some prominent ID advocates, including Michael Behe, accept broad common descent while still arguing that the underlying source of biological variation required design input Michael Behe on common descent and design. Under that kind of model, endogenous retroviruses, pseudogenes, chromosome rearrangements, neutral substitutions, and shared synteny could all still represent genuine evidence of ancestry; the design question would concern how certain specific innovations arose, not whether genealogical descent occurred at all. This is why build spec section 25's distinction between creationism and Intelligent Design matters in practice, not just in principle.

What each model predicts, and where each has explanatory strength
Question Common descent predicts Common design predicts
Functionally useful similarity (shared proteins, metabolic pathways) Expected where the function is old and broadly needed, or where descent from a shared ancestor already carrying the function explains its spread. Directly and simply expected — reusing a working solution is exactly what an engineer reusing proven components would do.
Functionally arbitrary shared details (chromosome 2's fusion scars, matching retroviral insertions, matching neutral mutations) Directly and simply expected — these are exactly the historically contingent markers inheritance from a shared ancestor would be expected to leave behind. Not directly predicted; must additionally explain why a designer would repeatedly reuse specifically nonfunctional or damaged historical accidents at matching genomic locations, rather than only the functional components.
Nested, nonoverlapping classification (species fit cleanly into one branching hierarchy) Directly and simply expected as the natural outcome of successive population splits. Not directly predicted; would require an additional, independently motivated assumption about why a designer's reuse pattern should happen to fall into a strictly nested hierarchy rather than a mix-and-match pattern across unrelated groups.
Gene-tree conflict at rates matching known population-genetic processes Directly and quantitatively predicted by incomplete lineage sorting, hybridization, and horizontal transfer models. Not independently predicted by the model itself; would need to be separately accommodated.

Common design can straightforwardly explain functional similarity — reuse of a working solution requires no additional assumptions. It is markedly less parsimonious for the historically contingent, functionally arbitrary similarities this chapter has emphasized throughout: the same broken gene, the same neutral mutation, the same retroviral insertion, the same chromosome rearrangement, repeated in the same genomic location across species. Common ancestry gives those patterns a direct, single-mechanism historical explanation. A common-design model can still accommodate them, but only by additionally explaining why a designer would repeatedly reuse specifically nonfunctional historical details — and, as of the evidence surveyed in this guide, that model has not yet generated independent, testable predictions about which such details should or should not appear that differ from what common descent already predicts. Without such independent predictions, the two models remain difficult to distinguish empirically on this evidence alone, even though they remain logically distinguishable as explanations.

Key Takeaways

  • Nested, historically contingent genomic details — chromosome fusion scars, matching retroviral insertions, shared neutral mutations — are stronger evidence for common descent than simple similarity, because convergent, function-driven processes have no obvious reason to reproduce them independently at matching genomic locations.
  • Homoplasy, incomplete lineage sorting, hybridization, and horizontal gene transfer are all real, well-documented, and individually different reasons a strict single-tree model can fail locally — none of them, on the evidence reviewed here, amounts to a systematic contradiction of shared ancestry.
  • Universal common ancestry is strongly supported by deep molecular conservation across all three domains of life; its formal statistical test has been seriously, though not fatally, contested, and the nature of LUCA and the earliest lineage structure remain genuinely uncertain.
  • Common descent and common design are not simply “evolution versus creationism”; they are separable explanatory models that should be compared on their predictions, and common design has not yet generated independent predictions that outperform common descent's on the historically contingent evidence this chapter emphasizes.

Common Overstatements

Check Your Understanding

Why is the human chromosome 2 fusion considered stronger evidence for common ancestry than the simple observation that human and chimpanzee genomes are similar?

Because it is not merely a similarity claim. Chromosome 2 carries specific, functionally unnecessary structural features — internal telomere-like repeats and remnants of a degenerated second centromere — in exactly the location a real ancestral fusion event predicts. Those features have no function that would make convergent, independent processes likely to reproduce them; a real historical fusion event, passed down by inheritance, is a much more economical explanation.

Does incomplete lineage sorting mean the gorilla genome study disproves the standard human-chimpanzee-gorilla family tree?

No. Incomplete lineage sorting is a well-understood, quantitatively predicted consequence of population genetics when population splits happen close together in time; it was not invented solely to explain the gorilla result. The genome-wide dominant relationship still groups humans and chimpanzees together; the roughly 30% local exception at individual loci is exactly the kind of discordance the model predicts, not evidence against the overall tree.

What would it take for common design to become a genuinely competing explanation with common descent, rather than merely a compatible alternative?

Common design would need to generate independent, testable predictions about which specific historically contingent genomic patterns — which shared neutral mutations, which retroviral insertions, which chromosome rearrangements — should or should not appear, in a way that differs from what common descent already predicts and that can be checked against real data. Without such independent predictions, the two models remain empirically difficult to distinguish on this evidence, even though they remain logically distinct claims.

What We Know

Within well-studied groups, especially humans and other great apes, common descent is very strongly supported by nested, historically contingent genomic evidence — chromosome-fusion signatures and matching retroviral insertions — that carries considerably more evidential weight than anatomical resemblance alone. Universal common ancestry is strongly supported by deep molecular conservation shared across all three domains of cellular life. Homoplasy, incomplete lineage sorting, hybridization, and horizontal gene transfer are all real, documented, and individually explicable complications to a strictly clean branching-tree model.

What Remains Disputed

The precise branching structure of the earliest cellular life, the nature of LUCA as a population versus a single organism, and how much of the deepest part of the genome was vertically inherited versus horizontally exchanged remain genuinely unsettled, even among researchers who accept universal common descent as the best-supported available explanation. Whether common design can be developed into a model that generates independent, testable predictions — rather than remaining a compatible but less economical alternative explanation — is itself an open question, taken up again from the dissenting side in Major Dissenting Arguments.

What Would Move the Debate Forward

Statistical model-comparison methods demonstrably robust to the alignment-correlation critique raised against Theobald's original universal-common-ancestry test would strengthen the formal statistical case. A clearer quantitative accounting of how much of the deep, earliest genome was vertically versus horizontally inherited would narrow one of the genuinely open questions about LUCA. For the common-descent-versus-common-design comparison specifically: a common-design model that generated and then successfully tested independent predictions about which historically contingent genomic patterns should or should not appear would be the clearest way to move that comparison from “compatible alternative” toward “genuinely competing, empirically distinguishable explanation.”

Sources for This Chapter