Protein Function and Sequence Space

New Biological Information distinguished sequence novelty from functional novelty and noted that documented cases of new protein function “typically begin with a pre-existing protein fold, an existing active site, and a weak ancestral side activity.” This chapter examines that claim in depth, at the level of individual proteins: how accessible is a new protein function, really, and from what starting points? The honest answer requires holding two well-supported bodies of evidence together at once, rather than picking the one that is more convenient.

By the end of this chapter you should be able to:

Required comparison for this chapter: evidence for functional accessibility versus evidence for strong biochemical constraints. Both bodies of evidence are real. The goal of this chapter is not to declare a winner but to show what each actually measures, so the two literatures can be read together instead of against each other.

Enzyme Promiscuity

Many enzymes possess a primary, well-optimized activity while also exhibiting weak, accidental promiscuous activities — side reactions the enzyme catalyzes poorly, as a byproduct of its main active-site chemistry, without having been “built for” that reaction. Aharoni and colleagues evolved several enzymes in the laboratory and found mutations that dramatically enhanced these promiscuous activities while having comparatively modest effects on the original, primary activity Aharoni et al. promiscuous protein functions. Conceptually:

protein with primary function A
  -> weak accidental activity B
  -> mutation + selection
  -> stronger activity B
  -> specialization

This provides direct evidence that mutation and selection can produce substantial functional change without requiring an entirely nonfunctional protein to become sophisticated in a single step. The starting point already has weak activity; evolution's job is to strengthen and refine what is already there.

Functional Specialization

Functional specialization is what tends to happen next: as a promiscuous side activity is strengthened, the protein (or, after duplication, one of two gene copies) can become progressively better at the new role, sometimes at some cost to the original one. This is the same underlying pattern discussed for duplicated genes in Gene Duplication, viewed here at the level of the protein's biochemistry rather than the gene's copy number: a generalist starting point gradually differentiates into one or more specialists.

Altered Protein Specificity

A stronger test of functional novelty is not just gaining a new side activity but changing which molecule a protein recognizes in the first place. Experiments reconstructing ancestral glucocorticoid receptors identified historical amino-acid substitutions associated with a genuine change in hormone specificity — the ancestral receptor and its modern descendant respond to meaningfully different hormones Historical contingency in glucocorticoid receptor evolution. This is a stronger functional claim than improving an existing weak activity, because the receptor's target itself changed.

Under a functional definition of information, this reasonably counts as new information: a descendant sequence performs a biologically relevant function differently from its ancestor. It demonstrates that specificity itself — not just activity strength — is evolutionarily accessible in at least this documented case.

Ancestral Protein Reconstruction

The glucocorticoid receptor work is possible because of ancestral protein reconstruction: researchers use phylogenetic methods to infer the probable amino-acid sequence of an extinct ancestral protein, chemically synthesize that inferred sequence, and test its actual biochemical behavior in the laboratory, sometimes by introducing individual historical substitutions one at a time to see which step did what Historical contingency in glucocorticoid receptor evolution. This is considerably stronger evidence than a purely verbal evolutionary scenario, because the reconstructed protein's function is empirically tested rather than assumed.

Ancestral sequence reconstruction remains a historical inference, not direct observation of the ancient protein. Researchers reconstruct the most probable ancestral sequence from patterns in living descendants; that inferred sequence is not guaranteed to be identical to the actual historical molecule, though the method's internal consistency and experimental testability make it far more rigorous than an untested narrative.

GFP Mutational Landscape

Here the chapter turns to the constraint side of the evidence. A large mutational study of green fluorescent protein (GFP) systematically tested how single and combined mutations affected fluorescence and found that roughly three-quarters of single-amino-acid mutations reduced fluorescence, about half of variants carrying four mutations completely lost fluorescence, and epistasis (discussed next) made many multi-mutation combinations far less viable than simply adding up their individual effects would predict GFP local fitness landscape.

This is strong, direct evidence that functional proteins often occupy restricted regions of sequence space, and that most nearby mutations make things worse rather than better. It quantifies, rather than merely asserts, the intuition that “most changes to a working protein break it.”

Epistasis

Epistasis means the effect of one mutation depends on which other mutations are already present. A mutation that is harmless, or even beneficial, on one genetic background can be strongly deleterious on another. The GFP landscape study found epistasis making many multi-mutation combinations far less viable than expected from single-mutation effects alone GFP local fitness landscape, and a separate study of mutational pathway accessibility found that epistasis can block some evolutionary routes entirely while leaving others open, depending heavily on the order in which mutations occur Epistatic accessibility of mutational pathways.

Epistasis restricts evolutionary paths. A mutation that would be useful once a prior mutation is already in place may be actively harmful before that prior mutation occurs, meaning not every combination of individually plausible mutations is actually reachable by cumulative selection.

Permissive Mutations

Epistasis has a constructive corollary: some mutations do not themselves change function but make a protein able to tolerate a later mutation that would otherwise be destabilizing. The glucocorticoid receptor work found that the historical transition to new hormone specificity required not only function-changing mutations but also permissive mutations that first made the protein able to tolerate the later functional change, in a particular historical order Historical contingency in glucocorticoid receptor evolution.

This simultaneously demonstrates genuine functional change and strong historical constraint. The transition was real, but it was not available from just any starting sequence at just any time; it depended on a specific permissive mutation occurring first. Alternative evolutionary routes that skipped the permissive step were highly restricted, because later mutations would otherwise have destabilized or damaged the protein.

Local Fitness Constraints

Taken together, the GFP landscape, epistasis, and permissive-mutation findings establish that protein evolution is not unlimited, easy, or unconstrained. Evolutionary paths can be restricted by protein folding, active-site geometry, stability requirements, molecular interactions, epistasis, and genetic background.

Conceptual three-dimensional fitness landscape with multiple peaks and valleys; a highlighted path climbs a nearby local peak toward a higher peak, with a valley shown as a barrier to direct movement, and a note that a path can climb nearby improvements without implying access to every region.
A local fitness constraint means many nearby paths are uphill dead ends or require crossing a valley. It does not by itself establish that every path toward every possible peak is blocked.

What local constraints do not prove: they do not establish a universal law that no protein can ever evolve beyond a fixed amount of novelty. A local fitness constraint shows that many specific pathways, from a specific starting sequence, are inaccessible. It does not establish that every possible path toward a new function, from every possible starting point, is blocked. Complex Systems and Evolutionary Constraints examines exactly when a demonstrated local constraint can, and cannot, be generalized further.

Protein Sequence-Space Size

For a protein of 100 amino acids, 20100 possible amino-acid sequences exist — a number no biological population could ever exhaustively sample. This motivates a serious argument: if useful proteins are extraordinarily rare and functionally isolated within that space, random mutation may have essentially no chance of finding them. Whether that argument succeeds depends on several further questions this chapter now works through: how common functional sequences actually are, whether multiple sequences can perform similar functions, whether functional sequences form connected networks, whether selection can follow viable intermediate steps, and whether evolution actually begins its search from a random sequence or from an already-functional one.

Conceptual map of a vast protein sequence space containing sparse, clustered functional regions connected by a dashed route representing local evolutionary search, distinct from uniform random sampling of the entire space; a note states that rarity estimates depend on the function and assay used.
Functional sequences are not evenly scattered through sequence space; they cluster in identifiable regions, and different experiments measure different things about those regions and the routes connecting them.

Random Sequence-Library Experiments

Keefe and Szostak directly tested how common a selectable function is within a huge library of genuinely random sequences. They constructed a library of approximately 6 × 10 12 proteins with 80 randomized amino acids and screened it for a specific function: the ability to bind ATP Keefe & Szostak random-sequence ATP-binding proteins.

ATP-Binding Experiments

Four unrelated families of ATP-binding proteins were recovered from the random library Keefe & Szostak random-sequence ATP-binding proteins.

This demonstrates that functional sequences exist outside known natural protein families and can occur at experimentally accessible frequencies — not vanishingly rare, but findable within a library that, while enormous, is still minuscule compared to the full 2080 space of possible 80-residue sequences.

The selected function, ATP binding, is substantially simpler than building systems such as ATP synthase, ribosomes, DNA replication machinery, or multi-component signaling networks. The experiment also used an enormous library, repeated rounds of selection, and additional mutagenesis during optimization. It does not establish that highly complex enzymes are common among random protein sequences; it establishes the narrower point that some selectable molecular functions are accessible within random sequence space.

Douglas Axe-Style Protein-Rarity Arguments

Douglas Axe tested sequence tolerance in a beta-lactamase-like enzyme domain using a very different design from Keefe and Szostak's. His approach did not randomly sample all possible sequences; instead, it began from an existing enzyme-related sequence, constrained sampling to the hydropathic pattern (the pattern of water-attracting and water-repelling residues) associated with that fold, randomized clusters of residues within that constraint, and measured retention of low-level biological function, extrapolating from those local experiments to estimate how prevalent sequences capable of sustaining the functional fold might be. The paper estimated that roughly one in 1064 hydropathic-signature-compatible sequences might support a working domain under its model, with further extrapolation to rarer frequencies for the specific function across broader sequence space Douglas D. Axe, Estimating the Prevalence of Protein Sequences Adopting Functional Enzyme Folds — Journal of Molecular Biology.

This strongly supports the claim that a sophisticated enzyme fold and function can impose severe sequence constraints — a large majority of sequences even within the constrained hydropathic class tested did not sustain the function.

The Axe estimate does not directly measure the frequency of any selectable protein function, the accessibility of a function from an already-functional neighboring sequence, the number of alternative folds capable of a useful phenotype, or evolutionary pathways that use promiscuous activities or duplication rather than starting fresh. It is not a direct random sample of all of global protein sequence space; the large rarity estimate depends on extrapolating from the experimental design and the chosen functional threshold.

Functional Accessibility versus Biochemical Constraint

The Axe estimate and the Keefe & Szostak result are often presented in popular debate as though they contradict each other — one says useful sequences are astronomically rare, the other says they were found in a huge random library. They do not directly contradict each other, because they measure different things.

QuestionAxeKeefe & Szostak
Starting materialEnzyme-related fold/signatureLarge random-sequence library
Function measuredRetention of a particular enzyme-domain function/foldATP binding
Functional thresholdSpecific biological activityBinding enrichment
Main questionHow rare is one demanding functional domain under the model?Do random sequences contain selectable binding functions?
Direct global-space measurement?NoSamples a huge but still tiny fraction

Both can be true simultaneously: simple molecular activities can be found among random sequences, while highly constrained modern enzyme functions can occupy much smaller regions of sequence space. Enzyme promiscuity reinforces the reconciliation from a third direction. Aharoni and colleagues' finding that mutation can greatly improve a weak existing secondary activity while causing comparatively small changes to the primary activity Aharoni et al. promiscuous protein functions changes the probability problem from “find a complete modern function from a random sequence” into something closer to “improve a weak existing activity through a local sequence neighborhood” — which can make some functions much more accessible, though it does not show that every major protein innovation had such a precursor available.

Neutral assessment: the empirical record supports two conclusions simultaneously. First, protein sequence space is strongly constrained for many demanding functions — the GFP landscape, epistasis, permissive-mutation, and Axe-style results are real and substantial. Second, functional sequence space is not composed solely of isolated modern proteins; random activities, promiscuous functions, and locally connected viable regions exist in specific studied cases. What is not established is a universal value for the fraction or connectivity of all biologically useful protein functions — the correct question in any specific case is whether a selectable path connects a given ancestral state to a given target function, not which side of this general debate is “right” in the abstract.

Falsification test: the claim that functional protein sequences are accessible to evolutionary search only in a narrow set of already-studied special cases would be seriously weakened by additional random-library selection experiments, targeting more demanding multi-step catalytic functions rather than simple binding, that continued to recover functional sequences at experimentally tractable frequencies. Conversely, the claim that functional accessibility generalizes broadly would be seriously weakened if such experiments, run against harder catalytic targets, consistently failed to find any hits even in enormous libraries.

Key Takeaways

  • Enzyme promiscuity, altered specificity, and ancestral protein reconstruction directly demonstrate that new and changed protein functions are evolutionarily accessible in documented cases.
  • The GFP fitness landscape, epistasis, and permissive-mutation findings directly demonstrate that protein evolution is strongly constrained, with most nearby mutations harmful and many paths blocked or order-dependent.
  • The Axe rarity estimate and the Keefe & Szostak random-library result measure different things — retention of one demanding fold-function versus discovery of a simpler function in random sequence — and are not direct contradictions.
  • Evolutionary theory generally proposes local search from an already-functional starting point, not uniform random sampling of a complete target sequence; neutral-network research supports the plausibility of such local search in specific studied cases without proving it applies universally.
  • The right question in any specific case is whether a selectable path connects a given ancestral state to a given target function — not a single verdict for all of protein evolution.

Common Overstatements

Check Your Understanding

Why are the Axe rarity estimate and the Keefe & Szostak ATP-binding result not a direct contradiction?

They measure different starting points, different functions, and different thresholds. Axe began from an enzyme-related fold and asked how rare sequences retaining a demanding, specific enzymatic function are within that constrained class. Keefe and Szostak began from a fully random-sequence library and asked whether a comparatively simple function (ATP binding) occurs at experimentally findable frequencies. Both results can be correct simultaneously: simple binding functions can be found in random sequence, while retaining a sophisticated, specific enzyme fold and function can be extremely rare. Neither study directly measured the other's question.

What is wrong with calculating the probability of evolution producing a modern protein as 1/20^100?

That calculation models uniform random sampling of one exact target sequence in a single step, which is not the process evolutionary theory proposes. Evolutionary theory proposes local search from an already-functional starting sequence, with selection able to preserve useful intermediates, plus additional mechanisms such as duplication, recombination, promiscuous activity, and co-option that further broaden the accessible search. The 1/20100 calculation would only be the correct model if the evolutionary claim actually required randomly assembling that exact sequence from scratch in one step, which it generally does not.

How does enzyme promiscuity change what “the probability of a new function evolving” actually means?

Without promiscuity, the naive question is the probability of finding a complete modern function starting from an unrelated or random sequence. With promiscuity, many proteins already possess a weak version of a secondary activity as an accidental side effect of their primary chemistry. The relevant probability question becomes: how likely is it that mutation and selection can strengthen an already-existing weak activity through a local sequence neighborhood? That is generally a much more tractable question than finding a function from nothing, though it does not establish that every major innovation had such a promiscuous precursor available.

What We Know

Mutation and selection have experimentally improved weak secondary protein activities and altered protein specificity in documented cases. Protein fitness landscapes are strongly constrained, with most single mutations harmful and epistasis blocking or reordering many evolutionary paths. Selectable functions (ATP binding) occur in large random-sequence libraries at experimentally accessible frequencies. Retention of a specific, demanding enzyme-domain function has been estimated as extremely rare within one constrained sequence class under one specific experimental model.

What Remains Disputed

How far the Axe-style rarity estimate generalizes beyond the one enzyme-domain class it measured, and how far the Keefe & Szostak result generalizes beyond ATP binding to more demanding catalytic functions, both remain genuinely open questions. Whether promiscuity, neutral networks, and local search are sufficient in combination to account for the full range of documented protein diversity — as opposed to the specific, individually studied cases in this chapter — is not settled by any single result presented here.

What Would Move the Debate Forward

Rigorous protein-evolution probability analyses that estimate the probability of some selectable path from a specified ancestral state to a specified target phenotype — accounting for the ancestral sequence, target size, local fitness landscape, mutation supply, neutral-network connectivity, permissible intermediate functions, duplication, and recombination — rather than only the probability of one exact modern sequence from a random draw, would narrow the disagreement identified above. Additional random-library selection experiments targeting more demanding, multi-step catalytic functions (beyond simple binding) would help establish how far the Keefe & Szostak-style result generalizes.

Sources for This Chapter