Biochemistry and Molecular Biology - Belyasova N.A. 2002

Molecular Foundations and Mechanisms of Heredity
Organization of the Cellular Genetic Apparatus
Genetic Code

The initial concepts regarding how Genetic information is encoded in genes were formulated by F. Crick in his "sequence hypothesis," which posited that the Amino Acid Sequence in a polypeptide chain is determined by The sequence of elements within the Gene. This hypothesis received experimental confirmation following the cracking of METABOLISM/28.html">The Genetic Code through the experiments of C. Yanofsky. In 1964, Charles Yanofsky demonstrated a correspondence between the relative positions of induced Mutations in the trpA gene of E. coli and the Amino Acid Substitutions in the enzyme encoded by this gene, Tryptophan synthetase. Thus, the collinearity between The Structure of a gene and the polypeptide it encodes was proven.

Nevertheless, the Molecular Basis of this collinearity was far from obvious, given that the entire diversity of Amino Acids in Polypeptides is represented by 20, whereas The Diversity of NUCLEOTIDES in DNA is represented by 4. Consequently, a single nucleotide cannot possibly encode a single amino acid in a peptide.

Experiments conducted by F. Crick and his co-workers studying mutations in the T4 bacteriophage of Escherichia coli led to the Conclusion that each amino acid is encoded by three nucleotides; that is, the genetic code is triplet. This deduction stemmed from the observation that mutations involving insertions or deletions of one or two nucleotides in the T4 genome resulted in The formation of abnormal Proteins with impaired function. Conversely, insertions or deletions of three nucleotides were frequently accompanied by minor changes in protein composition, allowing the latter to retain their activity. Crick and Brenner concluded that the genetic code is read in discrete units of 3 nucleotides. Under such a mechanism, the insertion (or deletion) of a nucleotide triplet should result in the addition (or removal) of just a single amino acid from the corresponding polypeptide. In cases where the insertion (or deletion) of nucleotides occurs in a quantity not a multiple of three, a "frameshift" occurs, causing the entire subsequent amino acid sequence in the protein to change completely.

Thus, the genetic code is triplet, meaning THE POSITION OF each amino acid in a polypeptide is specified by a sequence of three nucleotides called a codon. Since the number of different nucleotides in DNA is four, the number of possible nucleotide triplet variants is given by: 4 ґ 4 ґ 4 = 64. Out of the 64 triplets, 61 encode amino acids—each triplet specifying only one amino acid—while the remaining three codons serve as Translation termination signals (Fig. 1.7). These are referred to as stop codons or nonsense codons because they do not specify any amino acid. Additionally, two coding triplets (most commonly ATG for Met, and occasionally GTG for Val) perform a dual function: they encode the amino acids Methionine or valine and simultaneously serve as start codons where the translation process begins (Fig. 1.7).

Class="center">

Fig. 1.7. STRUCTURE OF THE genetic code. Codons functioning as translation start sites are underlined. Stop codons that terminate the translation process are highlighted. A characteristic feature of the genetic code is the absence of commas, meaning there are no punctuation marks separating one codon from another. Furthermore, the genetic code is non-overlapping within a given reading frame, with the reading frame being established by the first "read" nucleotide (Fig. 1.8). The maximum number of reading frames in a gene is 3, which equals the number of "letters" in the code.

Most cellular organisms utilize only a single reading frame, whereas certain Viruses may employ two or even three.

The encoded message is read in the 5'-to-3' direction of the mRNA, which is a transcript of the DNA "+" strand synthesized in the 5' → 3' direction. The codon closest to the 5' end corresponds to the N-terminal amino acid of the polypeptide chain. Consequently, proteins are synthesized from the N-terminus to the C-terminus (Fig. 1.8).

Another property of the genetic code is its degeneracy. This means that a single amino acid can be encoded by more than one nucleotide triplet. Conversely, the code is unambiguous: each codon specifies one and only one amino acid. This pattern implies that while knowing the DNA nucleotide sequence allows one to easily determine The amino acid sequence of a protein, the reverse is not true—a known amino acid sequence cannot be unambiguously translated back into a DNA nucleotide sequence. As a rule, the degeneracy of the genetic code means that for codons specifying the same amino acid, only the first two nucleotides are strictly recognized, whereas the third may be immaterial.

Fig. 1.8. Reading frames of the genetic code. The number "1" designates the first nucleotides that establish each of the three possible reading frames. A polypeptide is synthesized from the N-terminus (free amino group) to the C-terminus (free carboxyl group). A frameshift leads to an alteration of the amino acid sequence in the peptide molecule.

To explain this phenomenon, Crick proposed the "Wobble Hypothesis," which was subsequently confirmed and is now known as the wobble rule. According to this rule, the pairing between the third nucleotide of an mRNA codon and the first nucleotide of a tRNA anticodon is flexible, largely because the first position of the tRNA anticodon is frequently occupied by a minor nucleotide containing inosine as its nitrogenous base. Inosine can form Hydrogen Bonds with uracil, cytosine, or adenine located at the third position of the codon. This mechanism enables a Cell to maintain fewer than 61 distinct tRNAs, as many tRNAs are capable of recognizing up to three different codons.

The genetic code is universal. This property implies that any mRNA molecule, when translated in The Cell of any Organism, will direct the synthesis of a polypeptide with the identical amino acid sequence. However, this rule has exceptions, particularly concerning Mitochondrial DNA genetic codes. For the most part, the standard "genetic dictionary" is used here as well; however, in mammalian Mitochondria, for instance, the UGA codon in mRNA is "read" as tryptophan, inserting tryptophan into the corresponding position of the peptide, whereas in nuclear mRNA this codon acts as a stop codon (Fig. 1.7) signaling the termination of translation. Conversely, in mammalian mitochondria, the nucleotide triplets AGA and AGG are read as termination signals, whereas in The Nucleus they encode the amino acid Arginine. Other deviations from the universal nuclear genetic code may occur in the mitochondria of other organisms.

The structure of nucleotide triplets correlates with The chemical properties of the amino acids they encode. For example, all codons with uridylate In the second position encode amino acids with hydrophobic side chains: phenylalanine, leucine, isoleucine, valine, and methionine. Excluding termination codons, the presence of adenylate in the second position specifies a polar or charged side chain (Tyrosine, Histidine, glutamine, asparagine, Lysine, glutamic acid, and aspartic acid). Furthermore, the codons for most hydrophobic acids differ by only a single nucleotide (Fig. 1.7). A similar situation is observed for the codons of Serine and Threonine (whose side chains contain a hydroxyl group) or Alanine and Glycine (which possess the least complex side chains). Thus, the genetic code is organized in such a way that even when nucleotides are substituted in the first or second position of certain codons, a structurally related amino acid is incorporated into the polypeptide, thereby minimizing disruptions to the protein's Secondary structure.

The genetic code was deciphered by Nirenberg and Khorana in the early 1960s. In their initial experiments, artificially synthesized homopolynucleotides—such as polyuridylic acid, polycytidylic acid, and others—were introduced as mRNA into a cell-free protein-synthesizing system containing all necessary components. Polypeptides synthesized under these conditions were subjected to Amino acid analysis, establishing that poly(U) mRNA (i.e., UUUUUU...) directs the synthesis of polyphenylalanine, poly(C) directs polyproline, and so on. Thus, it was concluded that the nucleotide triplet UUU encodes the amino acid phenylalanine, while CCC encodes Proline. The final decoding of all 64 codons was achieved by utilizing synthetic polyribonucleotides with known repeating sequences in cell-free translation systems. These regular copolymers were obtained by combining organic synthesis and enzymatic Methods.



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.