Biological Chemistry - Berezov, T. T., Korovkin, B. F. 1998

Protein Biosynthesis
Translation and general requirements for cell-free protein synthesis
The nature of the genetic code

Structure/149.html">The problem of Protein Synthesis is closely tied to METABOLISM/2.html">THE CONCEPT OF The Genetic Code. The Genetic information encoded in the Introduction/19.html">Primary Structure of DNA is translated into The nucleotide sequence of mRNA while still in The Nucleus. How this information is actually transferred to the protein molecule remained unclear for a long time. The first evidence pointing to a direct linear relationship between Gene Structure and its product—a protein—can be found in the works of C. Yanofsky. Through a series of elegant experiments employing genetic mapping and sequencing techniques, he demonstrated that the order of alterations within the mutant Tryptophan synthase gene of E. coli precisely matches the order of Changes in the Amino Acid Sequence of the resulting enzyme protein.

Eukaryotic Cells possess a specialized mechanism for the precise and efficient Translation of the mRNA sequence into the corresponding amino acid sequence of the synthesized protein. Because mRNA molecules themselves have no inherent affinity for Amino Acids, it was hypothesized that translating the nucleotide sequence of mRNA into The amino acid sequence of Proteins requires a specific intermediary, termed an adaptor (see above). This adaptor molecule must be capable of recognizing both the specific nucleotide sequence of the mRNA and its corresponding amino acid. Equipped with such an adaptor molecule, The Cell can position each amino acid precisely within the growing polypeptide chain in strict accordance with the mRNA nucleotide sequence. Thus, it remains an established principle that the Functional groups of Amino acids are inherently incapable of interacting directly with the template and Messenger RNA.

It was demonstrated that the nucleotide sequence of mRNA contains specific code "words" for each amino acid—the genetic code. Most likely, this code resides in a definite sequence of NUCLEOTIDES within the DNA molecule. The questions of which nucleotides are responsible for incorporating a specific amino acid into a protein molecule, and how many nucleotides dictate this incorporation, remained unresolved until 1961. Theoretical analysis showed that the code cannot consist of a single nucleotide, since in that case only 4 amino acids could be coded. Nor can the code be a doublet; that is, combinations of two nucleotides from a four-letter "alphabet" cannot cover all amino acids, as theoretically only 16 such combinations are possible (42 = 16), whereas proteins are composed of 20 amino acids. To encode all the amino acids of a protein molecule, a triplet code is fully sufficient, yielding 64 possible combinations (43 = 64).

From M. Nirenberg's data, it becomes evident that poly-U—that is, an artificially synthesized RNA containing only uridylic mononucleotide residues—promotes the synthesis of a protein built from residues of a single amino acid, phenylalanine. Based on this, it was concluded that the codon for incorporating phenylalanine into a protein molecule is a triplet composed of three uridylic nucleotides, namely UUU. Shortly thereafter, it was shown that synthetic polycytidylic acid (poly-C) guides The formation of polyproline, and polyadenylic acid (poly-A) guides polylysine; the corresponding triplets CCC and AAA indeed proved to be the codons for Proline and Lysine, respectively.

A pivotal role in unraveling the complete genetic code "dictionary" was played by the approaches developed by H. Khorana for synthesizing polyribonucleotides (artificial mRNAs) with defined repeating triplet sequences (copolymers). These were subsequently used as templates in a cell-free protein-synthesizing system. The resulting Polypeptides contained equal proportions of amino acids that perfectly matched the copolymer template.

Soon afterward, in the laboratories of M. Nirenberg, S. Ochoa, and H. Khorana, utilizing these artificially synthesized mRNAs, evidence was provided not only for the composition but also for the triplet sequence of all codons responsible for incorporating each of the 20 amino acids into a protein molecule. Below is the complete code "dictionary," encompassing all 64 codons:

Class="center">

The genetic code for amino acids is degenerate. This means that a significant majority of amino acids are specified by multiple codons. With the exception of Methionine and tryptophan, virtually all Other Amino Acids have more than one specific codon. The recognition of an mRNA codon by a tRNA anticodon is based on more than just standard base pairing, where each codon base forms a pair with a complementary nitrogenous base on the anticodon. Under strict base pairing alone, each anticodon—and consequently each tRNA molecule—could in principle recognize only a single mRNA codon. However, evidence shows that certain tRNAs can recognize more than one codon. Specifically, Yeast Alanine tRNA has been shown to recognize 3 codons: GCU, GCC, and GCA. As can be seen, the differences pertain solely to the identity of the 3rd nucleotide. In this regard, the "wobble" hypothesis was proposed, suggesting that the pairing of the 3rd base is subject to less stringent constraints, allowing for a somewhat flexible or non-strict correspondence of this nucleotide—likely one of the reasons behind the degeneracy of the genetic code. The degree of code degeneracy varies among amino acids. For instance, while Serine, Arginine, and leucine each have 6 code "words," several other amino acids, such as glutamic acid, Histidine, and Tyrosine, have 2 codons, and tryptophan has only 1. It is therefore quite plausible that The sequence of the first two nucleotides primarily determines the Specificity of each codon, whereas the 3rd nucleotide is apparently less critical. Recently, proponents of a "two-out-of-three" hypothesis have emerged, suggesting that the protein synthesis code may be quasi- or pseudo-doublet in nature.

It turns out that the degeneracy of the genetic code carries clear biological significance, providing the Organism with several distinct advantages. In particular, it contributes to the "refinement" of The Genome, because point Mutations driven by chemical or physical factors can lead to various Amino Acid Substitutions, the most advantageous of which are then selected over the course of evolution.

Another hallmark feature of the genetic code is its continuity—the absence of "punctuation marks," or signals indicating the end of one codon and the beginning of another. In other words, the code is linear, unidirectional, and non-overlapping: ACGUCGACC. This property ensures the synthesis of an exact and highly ordered sequence of amino acid residues in the protein molecule. Otherwise, the nucleotide reading frame would be disrupted, leading to the synthesis of "nonsense" polypeptide chains with altered structures and unpredictable Functions. One must also highlight another crucial feature of the code: its universality across All living organisms, ranging from E. coli to humans. The code has remained largely unchanged over millions of years of evolution.

Among the 64 conceivable codons, 61 are sense codons, meaning they specify a particular amino acid. Meanwhile, three of them—namely UAG, UAA, and UGA—are "nonsense" codons; they were designated as nonsense (or stop) codons because they do not encode any of the 20 amino acids. However, these codons are far from meaningless, as at least two of them perform the vital function of termination signals during polypeptide Synthesis on Ribosomes (signaling the cessation and termination of translation).

Investigations of the genetic code through in vivo experiments have likewise corroborated its universality; however, recent years have revealed certain deviations within animal Mitochondria, including human cells. The cytoplasmic genetic code differs from the mitochondrial code by 4 codons. Two codons—AUG, which typically serves as the initiation codon, also codes for internal methionine residues in the chain, while UGA, a standard nonsense codon, encodes tryptophan in mitochondria. Furthermore, the codons AGA and AGG act as termination signals in mitochondria rather than encoding arginine. As a result, decoding the Mitochondrial Genome requires fewer distinct tRNAs, whereas the cytoplasmic translational machinery maintains a full Complement of tRNAs.



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.