LEHNINGER PRINCIPLES OF BIOCHEMISTRY - VOLUME 3. INFORMATION PATHWAYS - 2017
PART III. INFORMATION PATHWAYS
Class="center">It is obvious that Harry's [Noller] discoveries do not explain how life began, nor do they answer what came before RNA. However, mounting circumstantial evidence suggests that life existed on our planet long before us, and that is precisely where we came from—that is a fact!
— Gerald Joyce, from a comment in Science, 1992
27. PROTEIN METABOLISM
Proteins are the ultimate product of most informational metabolic pathways. At any given moment, a Cell requires thousands of different proteins. They must be synthesized in accordance with cellular demands, delivered to their proper subcellular localization, and degraded when they are no longer needed.
Unraveling The Mechanism of METABOLISM/35.html">Protein Biosynthesis—a profoundly complex and vital life process—stands as one of the most remarkable triumphs of biochemistry. In eukaryotes, Protein Synthesis involves more than 70 distinct ribosomal proteins, 20 or more Enzymes for Amino Acid Activation, 10 or more auxiliary enzymes and other protein factors for polypeptide initiation, elongation, and termination, roughly 100 additional enzymes for final protein Processing, and 40 or more types of transfer and Ribosomal RNAs. Altogether, nearly 300 different macromolecules participate in polypeptide synthesis. Many of these macromolecules are integrated into the intricate three-dimensional structures of Ribosomes.
The Significance of protein synthesis to The Cell can be grasped by looking at the cellular resources dedicated to this process. Up to 90% of the chemical energy expended by a cell on all biosynthetic reactions can be consumed by protein synthesis. Every bacterial, archaeal, or Introduction/5.html">Eukaryotic Cell contains anywhere from a few to several thousand copies of numerous proteins and RNA molecules. A typical bacterial cell houses 15,000 ribosomes, 100,000 protein factors and enzymes associated with protein synthesis, and 200,000 tRNA molecules, which together account for over 35% of the cell's dry weight.
Despite being exceptionally complex, Protein synthesis proceeds with remarkable speed. A polypeptide consisting of 100 amino acid residues is synthesized in an Escherichia coli cell (at 37 °C) in approximately 5 seconds. The synthesis of thousands of different proteins within the cell is regulated so that their Abundance precisely matches the current metabolic state. To maintain the requisite composition and concentration of proteins, the rates of transport and degradation must keep pace with The rate of synthesis. Through extensive research, we are gradually lifting the veil on the orchestrated molecular dance in which each protein finds its proper place within the cell and is selectively degraded when its services are no longer required.
In the course of studying protein synthesis, we have uncovered a world of catalytic RNA molecules that may have existed before life evolved in its modern form. Researchers have mapped the architecture of bacterial ribosomes, shedding light on the details of cellular protein synthesis with astonishing clarity. What did they find? Proteins are synthesized by giant RNA enzymes!
27.1. The Genetic Code
MODERN CONCEPTS OF protein biosynthesis are rooted in three major scientific breakthroughs. In the early 1950s, Paul Zamecnik and his coworkers conducted a series of experiments to determine where within the cell protein synthesis takes place. They injected radioactively labeled Amino Acids into rats, and at various time intervals post-injection, the animals were sacrificed, their livers removed and homogenized, the homogenate fractionated by centrifugation, and the subcellular fractions examined for the presence of radioactive protein. Hours and days after the injection of labeled amino acids, all subcellular fractions contained labeled proteins. However, within minutes of the injection, labeled proteins were detected exclusively in the fraction containing small ribonucleoprotein particles. These particles, visible in animal Tissues under an Electron microscope, were subsequently identified as the sites of protein synthesis from Amino Acids and were later named ribosomes (Fig. 27-1).

Mahlon Hoagland and Zamecnik made a second pivotal discovery, finding that amino acids were "activated" upon incubation with ATP and the cytosolic fraction of Liver Cells. The amino acids became attached to a heat-stable, soluble RNA of a specific type, which had been discovered and characterized by Robert Holley and was later named Transfer RNA (tRNA); this reaction yielded aminoacyl-tRNA molecules. The enzymes that catalyze this process are known as Aminoacyl-tRNA synthetases.
Fig. 27-1. Ribosomes and The Endoplasmic reticulum. An electron micrograph and a schematic representation of a pancreatic cell fragment, showing the attachment of ribosomes to the outer (cytoplasmic-facing) surface of the endoplasmic reticulum (ER). Ribosomes appear as numerous small dots along the borders of parallel membrane layers.

The third major discovery emerged from the research of Francis Crick. He posed a fundamental question: how can Genetic information encoded in the four-letter language of Nucleic Acids be translated into the 20-letter language of proteins? Small nucleic acid molecules (likely RNA) could act as adaptors—one part of the adaptor molecule binding to a specific amino acid, while another part recognizes The nucleotide sequence encoding that amino acid within the mRNA (Fig. 27-2). This hypothesis was soon successfully tested. The tRNA adaptor molecule translated the nucleotide sequence of the mRNA into the Amino Acid Sequence of a polypeptide. The entire process of mRNA-directed Protein synthesis is commonly referred to simply as translation.
Fig. 27-2. Crick's adaptor hypothesis. We now know that an amino acid is covalently attached to the 3'-end of the tRNA molecule, whereas a specific nucleotide triplet within the tRNA interacts with a corresponding triplet codon in the mRNA via complementary base-pairing Hydrogen Bonds.

Shortly after these discoveries, the principal Stages of Protein Synthesis were elucidated, and The Genetic Code specifying each amino acid was cracked.
The genetic code was deciphered using synthetic mRNAs
By the 1960s, it became apparent that at least three nucleotide residues in DNA are required to code for each amino acid. The four letters of the DNA code (A, T, G, C) grouped in pairs can yield only 42 = 16 different combinations, which is insufficient to encode 20 amino acids. Grouping the DNA code letters in triplets yields 64 combinations (43 = 64).
Several key principles of coding were established through early genetic studies (Figs. 27-3, 27-4). A codon is a nucleotide triplet that specifies a particular amino acid. Translation proceeds such that these nucleotide triplets are read in a continuous sequence. The first codon in the sequence establishes the reading frame, within which each subsequent codon begins exactly three nucleotide residues further down. There are no punctuation marks between codons. The amino acid sequence of a protein is dictated by the linear sequence of adjacent triplets. In principle, any single-stranded DNA or mRNA sequence can possess three reading frames. Each reading frame contains its own sequence of codons (Fig. 27-5), but only one corresponds to a given functional protein. The question remains: which three-letter codons correspond to which amino acids?
Fig. 27-3. Overlapping and non-overlapping genetic codes. In a non-overlapping code, codons (numbered sequentially) are arranged consecutively without intervening NUCLEOTIDES. In an overlapping code, certain nucleotides in the mRNA belong to multiple different codons. In a triplet code with maximal overlap, many nucleotides (e.g., the third from the left (A)) belong to three distinct codons. Note that in an overlapping code, the triplet sequence of the first codon severely restricts the possible sequence options for the subsequent codon. A non-overlapping code provides significantly greater nucleotide Variability in adjacent codons and, consequently, in the possible Amino acid sequences represented by the code. In all organisms known to date, the genetic code is non-overlapping.

In 1961, Marshall Nirenberg and Heinrich Matthaei published a study that marked the first breakthrough in cracking the genetic code. They incubated synthetic polyuridylic acid (poly(U)) with an E. coli extract, GTP, ATP, and a mixture of 20 amino acids in 20 separate test tubes—each tube containing one radioactively labeled amino acid. Because poly(U) mRNA consists of numerous UUU triplets, it can template the synthesis of only a polypeptide composed of a single amino acid encoded by the UUU triplet. Indeed, a radioactive polypeptide was formed in only the single tube containing radioactive phenylalanine. Nirenberg and Matthaei concluded that the UUU triplet encodes phenylalanine. Using a similar approach, it was demonstrated that polycytidylic acid (poly(C)) encodes a polypeptide consisting exclusively of Proline (polyproline), and polyadenylic acid (poly(A)) encodes polylysine. Polyguanylic acid fails to form any Peptides in such an extract because it spontaneously folds into four-stranded structures (see Fig. 8-20) that cannot be bound by ribosomes.

Fig. 27-4. Nonoverlapping triplet code. The principles underlying the Organization OF THE genetic code were established through numerous diverse experiments, including genetic studies involving insertion or deletion Mutations. The insertion or deletion of a single base pair (illustrated using mRNA Transcription) shifts the reading frame of the nonoverlapping code, altering all subsequent amino acids encoded by the mRNA. While a Combination of an insertion and a deletion changes specific amino acids, it can ultimately restore the original amino acid sequence. The insertion or deletion of three nucleotides (not shown) preserves the triplet reading frame, proving that a codon consists of three nucleotides rather than four or five. Codons transcribed from the original Gene are shown in gray, whereas new codons resulting from an insertion or deletion are shown in blue.

The synthetic polynucleotides used in these experiments were generated using polynucleotide phosphorylase (p. 26.2), an enzyme that catalyzes The formation of RNA polymers starting from ADP, UDP, CDP, and GDP. Discovered by Severo Ochoa, this enzyme does not require a template and instead produces macromolecules whose composition depends solely on the relative concentrations of nucleoside 5′-diphosphates in the reaction mixture. In the presence of UDP alone, the enzyme synthesizes poly(U). When added to a mixture of five parts ADP and one part CDP, however, it forms a copolymer in which five-sixths are adenylates and one-sixth are cytidylates. This irregular polymer contains many AAA triplets, fewer AAC, ACA, and CAA triplets, relatively few ACC, CCA, and CAC triplets, and very few CCC triplets (Table 27-1). By utilizing various artificial mRNA molecules synthesized by polynucleotide phosphorylase from different starting mixtures of ADP, GDP, UDP, and CDP, researchers were soon able to identify the triplet compositions encoding nearly all amino acids. Although these experiments revealed the base compositions of the triplets, the exact sequence of bases could not be determined by this method alone.
Fig. 27-5. Reading frames and the genetic code. In a nonoverlapping triplet code, any mRNA molecule can have up to three reading frames, which are distinguished here by color. The triplets and, consequently, the amino acids they encode differ for each reading frame.

Table 27-1. Incorporation of Amino Acids into Polypeptides Directed by Random RNA Copolymers
Amino acid |
Observed incorporation frequency (Lys=100) |
Inferred nucleotide composition* of the corresponding codon |
Expected incorporation frequency based on a normalized frequency for Lys = 100 |
Asparagine |
24 |
A2C |
20 |
Glutamine |
24 |
A2C |
20 |
6 |
AC2 |
4 |
|
100 |
AAA |
100 |
|
Proline |
7 |
AC2, CCC |
4.8 |
26 |
A2C, AC2 |
24 |
Note. The data presented were obtained in one of the early experiments directed at cracking the genetic code. Polypeptides were synthesized using a synthetic RNA template containing only A and C residues in a 5:1 ratio, after which the quantitative and qualitative Amino Acid Composition of the polypeptide was determined. Then, based on the relative frequencies of A and C residues in the synthetic RNA and arbitrarily Setting the frequency of the AAA codon (the most abundant codon) at 100, it was calculated that there should be three isomeric codons of composition A2C, each with a frequency of 20; three isomeric codons of composition AC2, each with a frequency of 4; and a relative frequency for the CCC codon of 0.8. The identity of the CCC codon had been established in previous experiments using poly(C) templates.
* These data provide no information regarding the nucleotide sequence within the codon (with the obvious exception of the AAA and CCC codons).
Conventions.
Throughout our subsequent Structure/133.html">Discussion, we will frequently encounter tRNA molecules. The Specificity of each tRNA for a given amino acid is denoted by a superscript, for example, tRNAAla, whereas aminoacylated tRNA is written with a hyphen: alanyl-tRNAAla or Ala-tRNAAla. ■
In 1964, Nirenberg and Philip Leder made another milestone discovery. Isolated E. coli ribosomes can bind specific aminoacyl-tRNAs in the presence of the corresponding synthetic polynucleotide template. For example, ribosomes incubated with poly(U) and phenylalanyl-tRNAPhe (Phe-tRNAPhe) bind both types of RNA molecules; however, if ribosomes are incubated with poly(U) and certain other aminoacyl-tRNAs, binding does not occur because those aminoacyl-tRNAs fail to recognize the triplets in poly(U) (Table 27-2). Even trinucleotide sequences can specifically bind to their corresponding tRNA molecules, allowing such experiments to be conducted with short synthetic oligonucleotides as well. This method made it possible to determine which aminoacyl-tRNAs bind to approximately 50 of the 64 possible triplet codons. While some codons were bound by no aminoacyl-tRNA molecules and others by more than one, a different approach was required to complete and confirm the genetic code.
Table 27-2. Trinucleotides Capable of Inducing Specific Binding of Aminoacyl-tRNA Molecules to Ribosomes
Relative increase in labeled aminoacyl-tRNA bound to ribosomes* |
|||
Trinucleotides |
Phe-tRNAPhe |
Lys-tRNALys |
Pro-tRNAPro |
UUU |
4.6 |
0 |
0 |
AAA |
0 |
7.7 |
0 |
CCC |
0 |
0 |
3.1 |
* Values represent the fold increase in bound 14C in the presence of the respective trinucleotide compared to a control lacking the trinucleotide.
Around the same time, an auxiliary analytical approach was introduced by H. Gobind Khorana, who devised chemical Methods for synthesizing polyribonucleotides with defined repeating sequences of two to four bases. Polypeptides directed by these mRNA molecules contain repeating sequences of a limited number of amino acids. When combined with data from the random polymer experiments performed by Nirenberg and coworkers, these sequences allowed unambiguous assignment of codon-amino acid correspondences. For example, the (AC)n copolymer contains repeating ACA and CAC codons: ACACACACACACACA. A polypeptide synthesized with such a template contains equal amounts of threonine and histidine. Because histidine codons were already known to contain one A and two C residues (Table 27-1), CAC must encode histidine and ACA must encode threonine.

The results of numerous experiments established the identity of 61 of the 64 possible codons. The remaining three codons proved to be stop codons, which prematurely terminate protein synthesis on synthetic polymer RNA templates (Fig. 27-6). The assignment of every triplet codon (Fig. 27-7) was completed by 1966 and subsequently verified by various independent methods. The cracking of the genetic code is widely regarded as one of the most profound scientific discoveries of the 20th century.
Fig. 27-6. Role of a stop codon in a repeating tetranucleotide. Stop codons (pink) occur after every fourth codon across three different reading frames (distinguished by color). Depending on the initial ribosomal binding site, the synthesis yields either dipeptides or tripeptides.

Fig. 27-7. The mRNA amino acid codon dictionary. Codons are written in the 5′ → 3′ direction. The third base of each codon (in bold) plays a less stringent role in specifying the amino acid than the first two. The three stop codons are shown on a pink Background, and the AUG initiation codon on a green background. All amino acids except Methionine and Tryptophan are specified by more than one codon. As a rule, codons specifying the same amino acid differ only at the third base position.

Codons play a pivotal role in the translation of genetic information, directing the Synthesis of specific proteins. The reading frame is established at the very outset of mRNA Translation and is maintained throughout the entire process as the ribosome moves sequentially from one triplet to the next. If the initial reading frame gains or loses one or two bases, or if translation somehow slips by a nucleotide within the mRNA, the reading frame for all downstream codons is disrupted, typically resulting in a nonsense protein with a severely garbled amino acid sequence.
A few codons perform specialized Functions (Fig. 27-7). The initiation codon AUG is the most prevalent signal for THE START OF polypeptide synthesis in all cells, while internally, it also encodes methionine. Stop codons (UAA, UAG, and UGA), also known as termination codons or nonsense codons, signal the end of polypeptide synthesis and do not encode any amino acids. Certain exceptions to this rule are discussed in Box 27-1.
As described in Section 27.2, the initiation of Protein synthesis in the cell is a complex process dependent on initiation codons and other signals in the mRNA. It is now evident that the experiments by Nirenberg, Khorana, and their colleagues to decipher codon functions would not have succeeded without initiation codons. Fortunately, the experimental conditions were not overly stringent. In The history of biochemistry, revolutionary research breakthroughs have quite frequently been achieved through human diligence and sheer good fortune.
In a random nucleotide sequence, on average, one out of every 20 codons in each reading frame turns out to be a termination codon. A reading frame of 50 or more codons uninterrupted by a stop codon is called an Open Reading Frame (ORF). Long open reading frames typically correspond to protein-coding genes. Complex computer programs exist to analyze nucleotide sequence Databases in search of open reading frames amid a colossal volume of noncoding information. Reading a continuous gene that encodes a typical protein with a Molecular Weight of 60,000 requires an open reading frame of 500 or more codons.
A remarkable property of the genetic code is that a single amino acid can be specified by more than one codon, which is why the code is described as degenerate. This does not mean the code is flawed: although an amino acid may be written using two or more codons, each codon designates only one amino acid. Codon degeneracy is non-uniform. Methionine and tryptophan have one codon each, Three amino acids (Leu, Ser, Arg) are encoded by six codons each, five Amino acids have four codons, isoleucine has three codons, and nine amino acids have two codons (Table 27-3).
Table 27-3. Degeneracy of the genetic code
Amino acid |
Number of codons |
Amino acid |
Number of codons |
Met |
1 |
Туr |
2 |
Тrр |
1 |
llе |
3 |
Asn |
2 |
Ala |
4 |
Asp |
2 |
Gly |
4 |
Cys |
2 |
Pro |
4 |
Gln |
2 |
Thr |
4 |
Glu |
2 |
Val |
4 |
His |
2 |
Arg |
6 |
Lys |
2 |
Leu |
6 |
Phe |
2 |
Ser |
6 |
The genetic code is nearly universal. With the exception of a few minor variations in Mitochondria, certain Bacteria, and some Unicellular Eukaryotes (Box 27-1), amino acid codons are identical across all currently known species of organisms. Humans, Escherichia coli, tobacco, amphibians, and Viruses all employ the same genetic code. This strongly suggests that all living creatures share a common ancestor whose genetic code has been conserved throughout biological evolution. Even the exceptions prove the general rule.
Box 27-1. The Exception That Proves the Rule: Natural Variations of the Genetic Code
In biochemistry, as in other disciplines, exceptions to general rules complicate teaching and frustrate students. At the same time, exceptions teach us that life is complex and inspire remarkable discoveries. Furthermore, understanding The Nature of these exceptions can, in unexpected ways, reinforce the rule itself.
At first glance, the genetic code would seem to offer little room for variation. Even a single amino acid substitution can have a devastating effect on Protein Structure. Nevertheless, variations in the code are indeed observed in certain organisms, and these variations are both fascinating and instructive. The types of variations and their exceptional nature provide compelling confirmation of the shared evolutionary origin of all living things.
For code alterations to occur, modifications must affect The genes of one or more tRNAs in a way that alters the anticodon. As a result of such a change, a specific amino acid is incorporated into the polypeptide sequence whenever a given codon appears in the mRNA sequence—a codon that normally (see Fig. 27-7) does not encode that amino acid. The genetic code is determined by two elements: (1) tRNA anticodons, which dictate where a specific amino acid is positioned within the growing polypeptide, and (2) the specificity of aminoacyl-tRNA synthetases, which match amino acids to their respective tRNA molecules.
Most sudden changes to the code can have catastrophic consequences for cellular proteins; therefore, modifications are most likely to occur where their effects are minimal, such as in small genomes encoding only a handful of proteins. Moreover, the biological consequences of code alterations are relatively minor if they occur within the three stop codons, which typically do not appear inside genes (see Box 27-4 regarding exceptions to this rule). This is precisely the pattern observed in reality.
Among the few known variations of the genetic code, the majority are found in Mitochondrial DNA (mtDNA), which encodes merely 10 to 20 proteins. Mitochondria contain their own tRNA molecules, so variations in their genetic code do not significantly impact the cell's main genome. Most known mitochondrial alterations (and the only code changes discovered in cellular genomes) involve stop codons. These changes affect the synthesis of products for only a subset of genes, and their impact is sometimes minimized by the presence of a large (redundant) number of stop codons within the genes.
In vertebrates, mtDNA contains genes for 13 proteins, two rRNAs, and 22 tRNAs (see Fig. 19-38 in Vol. 2). Rare cases of code alterations and an unusual set of wobble bases in codons allow Mitochondrial Genes to be translated using only 22 tRNAs, in contrast to the standard code, which requires 32 tRNAs. These mitochondrial adjustments can be viewed as genome optimization, since a reduced Genome Size confers Replication advantages on the organelle. For the four codon families in which an amino acid is completely determined by the first two nucleotides, a single tRNA with a U residue at the first (wobble) position of the anticodon suffices: either U pairs in some manner with any of the four possible bases at the third codon position, or a "two-out-of-three" mechanism operates, where pairing at the third position is not required. Another group of tRNAs recognizes codons with either A or G at the third position, while a third group recognizes U or C. Thus, practically all tRNAs recognize either two or four codons.
In the standard code, only Two amino acids correspond to a single codon each: methionine and tryptophan (see Table 27-3). If all mitochondrial tRNAs recognize two codons, one would expect additional codons for Met and Trp to exist in mitochondria. Indeed, one of the most widespread code variations involves a shift in the meaning of the common stop codon UGA, which in this case encodes tryptophan. The tRNATrp molecule inserts a Trp residue both at UGA sites and at the standard Trp codon, specifically ITGG. The second most common change involves the AUA codon, which here designates Met rather than Ile; the normal methionine codon AUG also functions, and a single tRNA recognizes both variants. Known mitochondrial code variations are summarized in Table 1.
Table 1. Known Codon Variants in Mitochondria
Codons* |
|||||
UGA |
AUA |
AGA AGG |
CUN |
CGG |
|
Normal codon assignment |
Stop |
Ile |
Arg |
Leu |
Arg |
Animals |
|||||
Vertebrates |
Trp |
Met |
Stop |
+ |
+ |
Drosophila |
Trp |
Met |
Ser |
+ |
+ |
Saccharomyces cerevisiae |
Trp |
Met |
+ |
Thr |
+ |
Torulopsis glabrata |
Trp |
Met |
+ |
Thr |
? |
Schizosaccharomyces pomhe |
Trp |
+ |
+ |
+ |
+ |
Mitochondrial Molds |
Trp |
+ |
+ |
+ |
+ |
Trypanosomes |
Trp |
+ |
+ |
+ |
+ |
Higher plants |
+ |
+ |
+ |
+ |
Trp |
Clamydomonas reinhardtii |
? |
+ |
+ |
+ |
? |
* N denotes any nucleotide; + indicates that the codon has the same assignment as in the standard code;
? indicates that the codon is absent in the given Mitochondrial Genome.
Examining the rarer code alterations in cellular genomes (distinct from mitochondrial ones), we find that in bacteria, the only known variation is again The Use of UGA to encode Trp residues. This alteration occurs in the simplest free-living cells, the bacterium Mycoplasma capricolum. Among eukaryotes, non-mitochondrial code variations are observed in a few species of ciliated Protozoa, where both stop codons UAA and UAG can encode glutamine. There are also rare but fascinating instances where stop codons are employed to encode amino acids outside the set of 20 standard amino acids (see Box 27-3).
Alterations do not necessarily affect all specific codons within a given genome; a codon does not always encode the exact same amino acid. For example, in most bacteria, including E. coli, the GTTG (Val) codon is sometimes utilized as an initiation codon specifying methionine. This occurs exclusively in genes where the GUG codon is positioned appropriately relative to the translation-initiation mRNA sequences (Section 27.2).
These variations demonstrate that the code is not as universal as once believed, yet its variability is strictly constrained. The variations are clearly derived from the standard code, and no Examples of a radically different code have been discovered thus far. The limited number of code variants reinforces the idea that all life on our planet evolved from a single (minimally variable) genetic code.
Wobbling allows certain tRNA molecules to recognize more than one codon
When multiple distinct codons specify a single amino acid, they typically differ only at the third base (located at the 3' end). For instance, Alanine is encoded by the triplets GCU, GCC, GCA, and GCG. The codons for Most amino acids can be represented by the shorthand formula XYAg or XYUC. It is the first two letters in each codon that dictate the amino acid, and this feature has A number of interesting biological consequences.
Base pairing between Transfer RNAs and mRNA codons is mediated by a three-base sequence in the tRNA known as the anticodon. The first Base of the mRNA codon (in the 5' —> 3' direction) pairs with the third base of the anticodon (Fig. 27-8a). If a tRNA anticodon triplet recognized only a single mRNA codon satisfying Watson-Crick base-pairing rules across all three positions, cells would need distinct tRNA molecules for every single codon. However, this is not the case because the anticodons of certain tRNAs contain the nucleotide inosinate (I), which features the unusual base hypoxanthine (see Fig. 8-5b in Vol. 1). Inosinate can form hydrogen bonds with three different nucleotides (U, C, and A; Fig. 27-8b), although these bonds are considerably weaker than the hydrogen bonds in Watson-Crick Base Pairs (G ≡ C and A = U). In yeast, a tRNAArg molecule possesses the anticodon (5')-ICG, which recognizes three Arginine codons: (5')-CGA, (5')-CGU, and (5')-CGC. The first two bases in these codons are identical (CG) and pair with the corresponding anticodon bases strictly according to Watson-Crick rules, whereas the third base (A, U, or C) forms weak hydrogen bonds with the I residue at the first position of the anticodon.
Fig. 27-8. Codon-anticodon pairing. (a) The RNA molecules have opposite orientations. The tRNA molecule is shown in the traditional cloverleaf configuration. (b) Three variants of codon-anticodon pairing in the presence of inosinate within the tRNA anticodon.

The Study of these and other codon-anticodon interactions led Crick to conclude that the third base in most codons is often bound relatively loosely to the corresponding base of the anticodon; in other words, the third base of such codons and the first base of the corresponding anticodons exhibit a non-strict match, or "wobble." Crick formulated four main postulates, known as the Wobble Hypothesis:
1. The first two bases of the mRNA codon always form canonical pairs (Watson-Crick pairs) with the anticodon bases in the tRNA and primarily determine the corresponding amino acid.
2. The first base of the anticodon (read in the 5' —> 3' direction; this base pairs with the third base of the codon) determines the number of codons recognized by the tRNA molecule. If the first base of the anticodon is C or A, base pairing is specific, and such a tRNA recognizes only one codon. If the first base is U or G, binding is less specific, and the tRNA can recognize two different codons. If the first (wobbling) nucleotide of the anticodon is inosine (I), the tRNA can recognize three different codons—the maximum number for any tRNA. These principles are summarized in Table 27-4.
Table 27-4. How Wobble Anticodon Bases Determine the Number of Codons Recognized by tRNA

Note. X and Y are bases complementary to bases X' and Y', respectively, forming Watson-Crick pairs with them. Wobble base pairs at the 3' position of the codons and the 5' position of the anticodons are highlighted in pink.
3. If an amino acid is specified by multiple codons, and if these codons differ by their first or second base, a distinct tRNA is required for each codon.
4. To translate all 61 codons, a minimum of 32 tRNAs is required (31 for amino acid coding and 1 for initiation).
The wobbling (third) base of the codon contributes to binding specificity, but because it binds relatively weakly to the corresponding anticodon base, it allows the tRNA to rapidly dissociate from the codon complex during protein synthesis. If all three codon bases formed regular Watson-Crick pairs with the three anticodon bases, the tRNA would leave the complex too slowly, which could significantly limit the Rate of protein synthesis. Thus, codon-anticodon interaction maintains a balance between the accuracy and the speed of synthesis.
The genetic code provides insight into how information regarding protein sequence is stored in Nucleic Acids and offers the key to understanding how this information is translated into proteins. We now turn to the MOLECULAR MECHANISMS OF translation.
Sequence Reading Depends on Frameshifting and RNA Editing
Once the reading frame is established during protein synthesis, codons are translated without overlapping or gaps until the ribosomal complex encounters a stop codon. The other two possible reading frames typically contain no useful genetic information; however, a few genes are organized in such a way that ribosomes make an error at a specific stage of mRNA translation, resulting in a frameshifting event from that point onward. This mechanism may allow a single transcript to direct the synthesis of multiple related, yet distinct proteins, or to regulate protein synthesis.
One of the best-studied cases of frameshifting is the translation of mRNA from the overlapping gag and pol GENES OF THE Rous Sarcoma virus (see Fig. 26-35). The reading frame of the pol gene is offset by one base pair to the left relative to the gag reading frame (Fig. 27-9).
Fig. 27-9. Frameshifting in retroviral transcripts. The gag-pol gene overlap region in the RNA of the Rous sarcoma virus is shown.

The pol gene product (Reverse Transcriptase) is produced from the same mRNA used for the Synthesis of the gag gene product (see Fig. 26-34). The resulting long polyprotein, or gag-pol protein, subsequently undergoes proteolytic Cleavage, being shortened into the mature protein—reverse transcriptase. The formation of this polyprotein requires a frameshift within the gene overlap region, as this allows the ribosome to bypass the UAG stop codon at the end of the gag gene (highlighted in pink in Fig. 27-9).
Frameshifting occurs in approximately 5% of translation events on this mRNA, meaning the gag-pol polyprotein (and ultimately reverse transcriptase) is synthesized at a rate 20 times lower than the gag gene product, yet this level is entirely sufficient for efficient viral replication. In some Retroviruses, frameshifting allows the translation of an even longer polyprotein that incorporates the env gene product fused to the gag and pol gene products (see Fig. 26-34). A similar mechanism allows the synthesis of the Շ and y subunits of E. coli DNA polymerase III from a single dnaX gene transcript (see Table 25-2).
Occasionally, mRNA is edited prior to translation. RNA editing may involve the addition, deletion, or substitution of nucleotides, thereby altering the meaning of the transcript. The deletion or insertion of nucleotides is most commonly observed in RNAs from the Mitochondrial and Chloroplast genomes of eukaryotes. This mechanism involves a specialized class of RNA molecules encoded within the same Organelles. The sequences of these RNAs are complementary to the sequences of the mRNAs being modified. These guide RNA molecules (Fig. 27-10) serve as templates during the editing process.
An example of editing via nucleotide insertion is found in the primary transcripts of the genes encoding cytochrome c oxidase subunit II in the mitochondria of certain protists. These transcripts do not fully match the sequence required at the C-terminus of the protein product. During post-transcriptional modification, four uracil residues are inserted into the sequence, which shifts the reading frame. Figure 27-10 illustrates the inserted U residues within a small fragment of the transcript undergoing editing. Note that base pairing between the primary transcript and the guide RNA generates several G=U pairs (blue dots), which are frequent occurrences in RNA molecules.
Fig. 27-10. Editing of the cytochrome c oxidase subunit II gene transcript from Trypanosoma brucei mitochondria. (a) The insertion of four uracil residues (pink) alters the reading frame. (b) Guide RNAs, complementary to the modified product, serve as templates during editing. Note the presence of non-Watson-Crick G=U pairs (blue).

Fig. 27-11. Deamination reactions in RNA editing. (a) The conversion of adenosine nucleotides to inosine nucleotides is catalyzed by ADAR enzymes. (b) The conversion of cytidine to uridine is catalyzed by Enzymes of the APOBEC family.

RNA editing via nucleotide substitution most commonly involves the enzymatic deamination of adenosine or cytidine residues, yielding inosine or uridine, respectively (Fig. 27-11), although substitutions of other bases also occur. During translation, inosine is interpreted as G. Adenosine deamination reactions are catalyzed by the ADAR family of enzymes (adenosine deaminases that act on RNA). Cytidine deamination is driven by the APOBEC family of peptides (apoB mRNA editing catalytic peptide), which also includes related AID deaminases (activation-induced deaminase). These two groups of deaminating enzymes contain homologous catalytic domains that coordinate a zinc atom.
A well-studied example of this type of reaction is the deamination of the apolipoprotein B gene, a component of vertebrate low-density lipoprotein. One form of apolipoprotein B, apoB-100 (Mr = 513,000), is synthesized in the liver, whereas the second form, apoB-48 (Mr = 250,000), is produced in the intestine; the mRNA for both forms is transcribed from the apoB-100 gene. The cytidine deaminase APOBEC, found exclusively in the intestine, binds to the mRNA at the codon for amino acid residue 2153 (CAA = Gln) and converts C to U, thereby generating a UAA stop codon. Consequently, the intestine produces apoB-48 from this modified mRNA—a truncated (N-terminal) version of the apoB-100 protein (Fig. 27-12). This reaction exemplifies tissue-specific synthesis of two distinct proteins from a single gene.
Fig. 27-12. Editing of the apoB-100 gene transcript, a component of LDL. Deamination, occurring exclusively in the intestine, converts a specific cytidine to uridine, resulting in the substitution of a Gln codon for a stop codon and the synthesis of a truncated protein.

The substitution of base A with I, mediated by ADAR proteins, is most frequently observed in primate gene transcripts, with over 90% of these substitutions occurring within short interspersed elements (SINEs) known as Alu elements (see Fig. 24-8). The Human Genome contains over a million Alu elements, each about 300 bp long, accounting for roughly 10% of the entire genome. They tend to concentrate near protein-coding genes, such as in introns and untranslated regions
at the 3' and 5' ends of transcripts. Newly synthesized (unprocessed) human mRNA typically contains between 10 and 20 Alu elements. ADAR enzymes specifically bind to double-stranded RNA to catalyze the A-to-I conversion. The abundance of Alu elements provides ample opportunity for intramolecular base pairing within transcripts, forming the duplexes required for ADAR activity. In some cases, these modifications alter gene coding sequences. Dysfunctional ADAR activity has been implicated in various neurological conditions, including AMYOTROPHIC LATERAL SCLEROSIS (Lou Gehrig's disease), Epilepsy, and clinical depression.
Vertebrate genomes contain numerous SINE elements, with most organisms harboring A wide variety of SINE types. Alu elements are notably predominant only in primates. Detailed analyses of genes and transcripts reveal that A-to-I substitution occurs 30 to 40 times more frequently in humans than in mice, largely due to the abundance of Alu elements. The relatively high frequency of A-to-I substitutions, along with the prevalence of Alternative Splicing (see Fig. 26-22), represents two distinctive features that set primate genomes apart from those of other mammals. It remains unclear whether these changes arose by chance or played a pivotal role in primate evolution and, ultimately, The Emergence of humans.
Summary of Section 27.1 The Genetic Code
■ The specific amino acid sequence of a protein is dictated by the translation of information encoded in mRNA, a process that takes place on ribosomes.
■ Amino acids are specified by mRNA codons, which are nucleotide triplets. Translation requires adapter molecules (tRNAs) that recognize these codons and insert amino acids in the correct sequence into the growing polypeptide chain.
■ Base sequences within codons were deciphered using synthetic mRNAs of known composition and sequence.
■ The AUG codon signals the initiation of translation, whereas UAA, UAG, and UGA codons serve as termination signals for protein synthesis.
■ The genetic code is degenerate: Almost all amino acids are specified by multiple codons.
■ The standard genetic code is universal across all organisms, with only minor variations found in mitochondria and certain single-celled organisms.
■ The third position of each codon is far less specific than the First and Second; nucleotides at this position are referred to as wobble nucleotides.
■ Frameshifting and RNA editing can alter how the genetic code is read during translation.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.