New Biological Information
The Seven Evolutionary Claims treated speciation, macroevolution, and common descent as claims about lineages branching and diverging. This chapter asks a different, more granular question: when a population's DNA changes, what exactly has been added? “Evolution creates new information” and “mutations only ever destroy information” are both popular slogans, and both are too simple. The word information is used to mean at least four different things in this debate, and conflating them is one of the most common ways the argument goes wrong on either side.
By the end of this chapter you should be able to:
- list at least four distinct things “new biological information” can mean, and explain why they are not interchangeable;
- explain the mechanisms — mutation, gene duplication, recombination, and de novo gene birth — that can plausibly generate each kind;
- state the Hazen functional-information concept and what it does and does not measure; and
- explain the difference between demonstrating that a mechanism can produce an outcome and demonstrating how often it does.
Core distinction for this chapter: a new sequence, a new function, and increased system complexity are different claims. Evidence for one does not automatically establish either of the others.
Ambiguity of the Word “Information”
Outside biology, “information” already carries several meanings — a string of symbols, a reduction in uncertainty, a message with meaning to a receiver, an instruction that produces an effect. Biology inherits all of that ambiguity. When someone asks whether evolution can produce “new information,” the honest first response is: which of the following do you mean?
| Meaning | Example |
|---|---|
| New sequence information | A mutation creates a DNA sequence that did not previously exist |
| More genetic material | A gene or genomic region is duplicated |
| New biological function | A protein acquires an activity its ancestor lacked |
| Greater system complexity | A molecular system acquires additional differentiated, required components |
These are not equivalent, and none implies the others automatically. A new DNA sequence does not automatically imply a useful function. A duplicated gene does not automatically imply a new function. A new biochemical activity does not necessarily imply increased system-level complexity. A useful analysis has to specify which kind of “information” is being discussed before asking whether evolutionary mechanisms can produce it.
New Sequence Information
In the narrowest and least controversial sense, a “new” sequence is simply one that did not exist in that exact form before a mutation created it. This sense of novelty is directly observed: modern pedigree sequencing identifies new DNA sequence variants in offspring that are absent from both parents, and follows some of those variants into further generations Four-generation pedigree mutation study. A large Icelandic trio-sequencing study identified more than 100,000 such de novo mutations and showed that the rate rises with parental age Icelandic de novo mutation study.
What this shows: new DNA sequences that never previously existed in a lineage are produced continuously and can be observed directly at the molecular level.
What this does not show: that the new sequence does anything useful. Most new sequence is functionally neutral or is not tested by selection at all in any given generation. Sequence novelty in this sense is a precondition for the more demanding senses of “new information” discussed below, not a substitute for them.
Increased Genetic Material
A second, distinct sense of “more information” is simply more genetic material — an increase in the total amount of DNA a lineage carries, most commonly through gene or genome duplication (covered in detail below). This sense is also directly documented: a duplicated gene is a countable, sequenceable fact about a genome.
What this does not show: more genetic material is not automatically more functional information. Immediately after a duplication event, both copies do the same job the single ancestral copy did; nothing new has been accomplished functionally until one copy is retained, diverges, and is co-opted for something else, or is lost. Quantity of DNA and functional content are separate variables.
Functional Information
A third sense asks about function directly: has a sequence acquired the ability to do something biologically useful that it, or its ancestor, could not do before? This is the sense most people actually mean when they ask whether evolution can create new information, and it is also the hardest to measure precisely, because “useful” and “function” both require a specified standard before they can be quantified.
New Biological Function
Plainly stated: does a descendant sequence do something biologically relevant that its ancestor did not do, or does it do the same thing detectably better? Under a functional definition of information, the answer for at least some documented cases is yes. Many enzymes have a primary activity alongside weak, accidental promiscuous activities; experimental evolution has shown that mutation and selection can take a weak secondary activity and turn it into a strong, specialized one, often while only modestly affecting the original activity Aharoni et al. promiscuous protein functions. Ancestral protein reconstruction of glucocorticoid receptors has similarly identified historical substitutions that changed hormone specificity — a genuine functional change, not just a sequence change Historical contingency in glucocorticoid receptor evolution.
What this shows: mutation and selection have been directly observed to produce a biologically relevant function a sequence's ancestor lacked, or to improve one substantially. Under a functional definition of information, this reasonably counts as new information.
What this does not show: these cases typically begin with a pre-existing protein fold, an existing active site, and a weak ancestral side activity — not a random, functionless string of amino acids. They do not establish that an advanced molecular function can arise in one step from a completely nonfunctional starting sequence, that every protein function is easily accessible, or that every major biological innovation followed the same pathway. Protein Function and Sequence Space takes up exactly this question — how accessible new functions are, and from what starting points — in full.
System-Level Complexity
A fourth, still more demanding sense concerns the system as a whole: has a molecular system gained additional differentiated, required components, so that it now does more, or does it more reliably, than a simpler ancestral version could? This sense is the one most closely tied to debates over irreducible complexity, taken up directly in Complex Systems and Evolutionary Constraints. It is worth flagging here because it is conceptually the furthest removed from simple sequence novelty: a system can become more complex by a component's losing an ancestral capability rather than by any component gaining one, a counterintuitive process examined in that later chapter's V-ATPase case study.
A new biochemical activity, in the sense discussed above, does not by itself establish increased system-level complexity. A protein can acquire a new activity without any change to how many components a system requires or how interdependent those components are.
Hazen Functional-Information Concept
One attempt to put the “functional information” sense above on a more precise footing is the measure associated with Hazen and colleagues:
I(Ex) = −log2 F(Ex)
where F(Ex) is the fraction of all possible configurations of a system that achieve at least a specified level of function Hazen et al. functional-information framework. In plain language: the rarer a functional configuration is among all the configurations that could exist, the more functional information that configuration is said to carry.
What this measure depends on: the result is not a fixed property of a sequence alone. It depends on the specific function being measured and on the performance threshold chosen to count as “functional.” Change either input and the quantity changes. This is a genuinely useful formal tool, not a settled, single number for any given protein or gene.
The concept reinforces an important distinction that recurs throughout this chapter and the next two:
Random processes can generate novelty; selection can cause functionally advantageous novelty to accumulate.
Selection does not generate mutations. Mutation, recombination, duplication, and related processes generate variants; selection changes their frequencies once they exist.
Mutation
Mutation — a change to an organism's DNA sequence — is the ultimate source of every sense of novelty discussed above; without it, a population's genetic variation could only be reshuffled by recombination, never expanded. As Foundations established, mutation is now directly observable at the molecular level through pedigree sequencing rather than merely inferred Four-generation pedigree mutation study Icelandic de novo mutation study. What mutation alone does not settle is the harder question this chapter is built around: whether a given new sequence does anything, and if so, how much.
Gene Duplication
Gene duplication produces a redundant copy of a gene or genomic region. Conceptually:
Original: A Duplication: A -> A1 + A2 Initially: A1 = old function, A2 = old function
The evolutionary advantage is not that duplication itself creates a new function. It is that redundancy can allow one copy to maintain the ancestral function while the other is freed from some of the selective constraint that kept the single ancestral copy fixed, and can then accumulate changes:
Possible later outcome: A1 -> retains or specializes in old function A2 -> acquires altered or specialized function
Phylogenetic, biochemical, and selection analyses of duplicated anthocyanin-pathway genes in morning glories found functional differentiation after duplication, consistent with escape from adaptive conflict: the ancestral gene performed competing roles, and duplication allowed the descendant copies to specialize Morning-glory duplication and functional specialization.
Duplication alone does not automatically create new functional information. Immediately after A → A + A, the two copies may perform the identical role, and many duplicated genes are eventually deleted, become pseudogenes, retain similar functions, or remain selectively neutral. The relevant evolutionary mechanism is duplication, plus retention, plus divergence, plus selection or drift, which together make functional differentiation possible — not duplication by itself.
Gene duplication is directly observed and can create genetic redundancy from which specialized or novel functions can subsequently evolve. It does not establish that all innovations originate by duplication.
Recombination
Recombination creates new combinations of already-existing genetic variants:
Parent 1: A b Parent 2: a B Recombination: A B
The A B combination may never have existed previously in a single genome, even though neither individual variant, A nor B, is new. Controlled yeast experiments comparing sexual and asexual populations have shown that recombination can combine beneficial mutations, separate beneficial mutations from harmful genetic backgrounds, reduce competition between independently arising beneficial lineages, and increase the overall rate of adaptation Recombination accelerates adaptation in yeast. Other experimental work has directly identified recombination events that combined two beneficial mutations into a fitter genotype Recombination combines beneficial mutations.
Whether recombination creates “new information” depends entirely on which sense from the table above is meant. Recombination generally does not create the underlying alleles being combined — those usually already existed. It cannot supply a needed mutation that does not exist anywhere in the population; it only reorganizes existing variation. But it can create new genotypes, new combinations of mutations, new epistatic interactions, and new phenotypes, and it can therefore generate new combinatorial information even when it generates no new sequence symbols at all.
De Novo Gene Birth
The most demanding claim in this chapter is that an entirely new protein-coding gene can originate from DNA that was previously noncoding:
noncoding DNA
-> occasional transcription
-> open reading frame / translation
-> weak or incidental effect
-> selection
-> functional gene
Comparative genomic studies in yeast have identified lineage-specific transcripts and proposed intermediate “proto-gene” states consistent with a continuum between noncoding DNA and established protein-coding genes Proto-gene model in yeast. Experimental work on candidate human and fly de novo proteins found that these proteins can be expressed, display plausible biophysical properties, and show moderately greater solubility than matched random sequences Biophysical properties of de novo proteins.
Frumkin and Laub generated approximately 108 artificial genes containing random sequences with no homology to any natural gene, screened them for the ability to help E. coli survive growth arrest caused by the toxin MazF, and identified a random-sequence protein that promoted survival by interacting with existing protein-homeostasis pathways Functional random-sequence protein in E. coli. This directly demonstrates that a previously nonexistent protein sequence can possess a selectable biological effect — the strongest available evidence that de novo gene birth is mechanistically possible, not merely a plausible story.
Strongest skeptical qualification: a gene that appears unique to one lineage does not automatically prove de novo origin. Apparent novelty can instead result from very rapid divergence, ancient ancestry that has become unrecognizable, incomplete genome sequencing, annotation errors, or the failure of homology-detection methods to find a real but highly diverged relative. Strong individual de novo cases therefore ideally require the corroborating evidence discussed next: synteny.
The random-sequence experiment establishes possibility, not natural frequency. The screening library was extremely large and the selection conditions were intentionally strong; the experiment does not establish how often anything comparable happens in wild populations. See Capacity versus Frequency below.
Synteny
Synteny is the conservation of gene order along a chromosome across related species. It is the key corroborating evidence that separates a well-supported de novo gene claim from a merely apparent one: if a gene looks unique to one lineage, researchers can check whether the corresponding genomic region can still be identified in close relatives — and, critically, whether the ancestral version of that region lacks the coding structure the descendant gene now has. If the surrounding sequence context lines up (the region is syntenic) but the coding capability does not extend into the relative's genome, that is much stronger evidence for genuine de novo origin than sequence novelty by itself, because it rules out the simplest alternative explanation — that a real, older gene was simply missed by homology searches.
Open question: how many currently accepted de novo gene candidates would survive a rigorous synteny check against a densely sampled set of close relatives? Genome sequencing of closely related species is still incomplete for many lineages, so the population of “confirmed” de novo genes is itself likely to shift as more comparative genomic data becomes available.
Capacity versus Frequency
This distinction closes the chapter because it applies to every mechanism discussed above, not only de novo gene birth. A laboratory or field demonstration that a mechanism can produce an outcome does not establish how often that outcome occurs in nature. The Frumkin and Laub random-sequence experiment Functional random-sequence protein in E. coli used an enormous, deliberately constructed library and strong artificial selection; it shows that selectable effects exist among random sequences, not that such events are common in real populations under real conditions. The same caution applies to gene duplication, recombination, and every functional-novelty case in this chapter: demonstrating capacity is a real scientific result, but it is a different, narrower claim than demonstrating typical frequency.
capacity ≠ frequency
Falsification test: the narrow claim that known mechanisms have demonstrated the capacity to generate new sequence, new function, and increased genetic material would be seriously weakened if controlled mutation-accumulation and selection experiments systematically failed to produce any documented functional novelty across repeated, well-powered trials. That has not happened; the cases in this chapter exist precisely because such experiments have repeatedly succeeded. What would remain genuinely open even so is the separate, harder frequency question above, which no single capacity-demonstrating experiment can settle by itself.
Key Takeaways
- “New biological information” is not one claim; at minimum it means new sequence, more genetic material, new function, and increased system complexity, and evidence for one does not automatically establish the others.
- Mutation, gene duplication, recombination, and de novo gene birth are each independently, directly documented mechanisms capable of producing at least some sense of novelty.
- The Hazen functional-information measure gives a precise formal definition, but its value depends on the function measured and the threshold chosen — it is a tool, not a settled number.
- Synteny is the key corroborating evidence that turns an apparent de novo gene into a well-supported one, by ruling out missed homology as the simpler explanation.
- Demonstrating that a mechanism can generate novelty (capacity) is a different, narrower claim than demonstrating how often it actually does so in nature (frequency).
Common Overstatements
- “Mutations only ever destroy information.” This conflicts with direct evidence of new sequence, new allele combinations, new selectable activity, and altered biochemical specificity documented throughout this chapter.
- “A new protein function has been observed, so complex life is explained.” Demonstrating functional novelty in specific systems does not automatically reconstruct every major historical innovation, and does not by itself establish increased system-level complexity, a separate claim addressed in Complex Systems and Evolutionary Constraints.
- “A gene looks unique to one species, so it must have arisen de novo.” Apparent lineage-specific novelty is also consistent with rapid divergence, incomplete sequencing, or failed homology detection; a strong case requires synteny evidence, not sequence novelty alone.
Check Your Understanding
A student says, “A gene got duplicated, so evolution created new information.” What is missing from this statement?
It does not specify which sense of “information” is meant. Duplication by itself only increases the quantity of genetic material; both copies initially do the same job. Whether that counts as “new information” depends on whether the student means increased genetic material (yes, directly), new function (only if the copies later diverge and one acquires a new role), or increased system complexity (a separate, still more demanding claim). The morning-glory anthocyanin-gene case shows functional divergence actually occurring after duplication, but that is a further step beyond duplication itself.
Why does the Frumkin and Laub random-sequence experiment count as strong evidence for the possibility of de novo gene birth but not for its frequency in nature?
The experiment screened roughly 108 artificial random sequences under intentionally strong selective pressure (survival under toxin-induced growth arrest) and found at least one sequence with a selectable effect. That is a valid demonstration of mechanism — a previously nonexistent sequence can have biological function. It says nothing about how often comparable events occur among the much smaller, unselected “libraries” of noncoding DNA that real populations carry, under real environmental conditions, without a researcher intentionally screening billions of variants for one specific selectable outcome. That is the capacity-versus-frequency gap.
Why does synteny strengthen a de novo gene claim more than simply failing to find a homologous gene in other species?
Failing to find a homolog is ambiguous: it could mean no homolog exists (genuine de novo origin), or it could mean a real homolog exists but was missed because of rapid sequence divergence, incomplete genome assembly, or limits of the homology-search method. Synteny adds an independent line of evidence: if the surrounding genomic neighborhood is recognizably the same region in a relative species, but the coding structure itself is absent from the ancestral version of that region, missed homology becomes a much less likely explanation, because the region itself was findable — only the coding capability was not there ancestrally.
What We Know
Mutation creates novel DNA sequences directly and observably. Duplication creates additional genetic material directly and observably. Recombination creates new allele combinations directly and observably. Mutation and selection have been experimentally shown to improve weak secondary protein functions and to alter biochemical specificity in specific studied cases. Random sequences have been experimentally shown capable of possessing selectable biological effects. Duplicated genes have been shown, in specific documented cases, to become functionally differentiated.
What Remains Disputed
Whether the demonstrated mechanisms in this chapter are sufficient, in combination and at plausible natural rates, to account for every major evolutionary innovation is genuinely disputed and is addressed quantitatively in Complex Systems and Evolutionary Constraints and Major Counterarguments. Individual natural de novo gene candidates also require case-by-case validation; a claimed instance can be contested on synteny or homology-detection grounds even where the general mechanism is well supported.
What Would Move the Debate Forward
Quantitative estimates of the natural rate at which de novo genes, functionally differentiated gene duplicates, and novel promiscuous activities actually arise and fix in wild populations — rather than laboratory demonstrations of capacity alone — would directly close the capacity-versus-frequency gap identified throughout this chapter. Additional syntenic, cross-species confirmation of proposed natural de novo genes would strengthen individual cases beyond what sequence novelty alone can establish.
Sources for This Chapter
- [Primary Research] Four-generation pedigree mutation study
- [Primary Research] Icelandic de novo mutation study
- [Primary Research] Aharoni et al. promiscuous protein functions
- [Primary Research] Historical contingency in glucocorticoid receptor evolution
- [Methods / Conceptual] Hazen et al. functional-information framework
- [Primary Research] Morning-glory duplication and functional specialization
- [Primary Research] Recombination accelerates adaptation in yeast
- [Primary Research] Recombination combines beneficial mutations
- [Primary Research] Proto-gene model in yeast
- [Primary Research] Biophysical properties of de novo proteins
- [Primary Research] Functional random-sequence protein in E. coli