BIOTECHNOLOGY - V. H. Herasymenko - 2006
Part I. General Biotechnology
Chapter 3. BASICS OF MOLECULAR BIOLOGY
3.2. PROTEIN BIOSYNTHESIS AND ITS REGULATION
3.2.1. The Genetic Code
The realization that DNA serves as the genetic material stems from the pioneering studies of O. T. Avery and co-workers (1944), who demonstrated Bacterial Transformation using purified DNA extracts from pneumococci, as well as the work of A. D. Hershey and M. Chase, who proved that when a bacterial Cell is infected, only the bacteriophage DNA enters The Cell, while its protein coat remains outside. Through the concerted efforts of scientists—biochemists, biophysicists, geneticists, chemists, and others—the three-dimensional Structure of DNA was deciphered in the early 1950s, following the somewhat earlier determination of The chemical composition of these macromolecules. Around the same time, researchers successfully determined the Amino Acid Sequence of Insulin, a protein consisting of just 51 amino acid residues. Thus, it was conclusively proven that DNA is a long, unbranched polymer composed of four monomeric structures—deoxyribonucleotides—repeated in varying sequences, whose nitrogenous bases are adenine (A), cytosine (C), guanine (G), and thymine (T). The mononucleotides are linked together by covalent phosphodiester bonds extending from the 5'-carbon atom of one deoxyribose residue to the 3'-carbon atom of the adjacent pentose residue, thereby forming a chain with a linear sequence.
X-Ray Diffraction Analysis revealed that DNA forms a double-stranded helix in which the nitrogenous bases are directed inward (like the rungs of a spiral staircase), while the deoxyribose-phosphate backbone is positioned on the outside (acting as the handrails). Optimal packing of the linear monomer sequences within the double-helical polynucleotide structure is achieved through the Interaction of a larger purine base (adenine or guanine)—each formed by the Condensation of six-membered and five-membered heterocycles—with a smaller pyrimidine base (thymine or cytosine), which are six-membered heterocycles.
Model experiments demonstrated that more effective Hydrogen Bonds are formed between guanine (G) and cytosine (C), and between adenine (A) and thymine (T), than in any other combinations of NUCLEOTIDES. The complementary pairing of A with T and G with C in the DNA double helix accounted for previously obtained biochemical findings regarding the stoichiometric equivalence of A to T and G to C; that is, the ratio between the nitrogenous bases in these pairs was 1:1 across all investigated DNAs.
Biochemical Analysis of Proteins that are products of mutant genes demonstrated that The sequence of the four monomeric structures (adenine, guanine, thymine, and cytosine) in DNA is collinear with the twenty Amino Acids in proteins. In other words, The nucleotide sequence in a protein-coding region of DNA corresponds to The amino acid sequence in that protein. Consequently, this state of affairs raised a question that became central to molecular biology: What is the mechanism behind such a biochemically complex transformation as translating a DNA nucleotide sequence into a protein amino acid sequence? The flow of information from DNA to protein can be symbolically represented as follows:
Class="center">![]()
This scheme attracted considerable attention from researchers and subsequently spurred the rapid development of biochemical genetics. The synthesis of RNA molecules is termed Introduction/24.html">DNA METABOLISM/31.html">Transcription. Formed on a DNA template along one of its strands, the RNA copy contains the complete Genetic information of that specific DNA region. RNA retains the capacity to form hydrogen bonds between complementary bases because uracil, which replaces thymine in RNA, pairs with adenine in the exact same manner as thymine. However, transcription differs fundamentally from Replication. Upon completion of its synthesis, the RNA copy is released from the DNA template, whereupon the original DNA double helix is restored. The newly synthesized RNA molecules are single-stranded, shorter than DNA, and correspond in length to the DNA segment required to code for one or several proteins. Certain DNA regions (genes) are utilized for RNA Synthesis thousands of times, whereas others are not transcribed at all.
In Eukaryotic Cells, many of the RNA molecules generated during transcription undergo significant chemical modifications before maturing into Messenger RNA (mRNA) and entering the Cytoplasm. In turn, thousands of copies of the corresponding polypeptide chain can be synthesized on each mRNA molecule in the cytoplasm. Given the exceptionally high rate of Protein Synthesis—for instance, a polypeptide chain consisting of one hundred amino acid residues is assembled in an *Escherichia coli* cell within just 5 seconds—the information contained within a relatively small DNA segment can be rapidly amplified into a vast quantity of a specific protein. Thus, on the template of a single Fibroin-encoding Gene, 104 mRNA molecules are synthesized, each directing The production of 105 molecules of fibroin, the primary component of silk. Overall, over a span of four days, a single silk-producing glandular cell manufactures 109 molecules of fibroin (Alberts et al., 1986).
The rules governing the Translation of Nucleotide sequences within the polynucleotide structure of DNA into the Amino acid sequences of proteins (The Genetic Code) were deciphered in the early 1960s by M. Nirenberg, J. Matthaei, P. Leder, and other investigators. Earlier genetic experiments had established that Amino acids are encoded by nucleotide triplets (codons). With four nitrogenous bases (A, T, G, C), it is possible to generate 64 (43) distinct triplet combinations, which is more than sufficient to encode 20 amino acids. Conversely, if combinations of only two nucleotides (a doublet code) were used, this number would clearly be insufficient to encode the entire set of amino acids. By incubating cell-free extracts from *Escherichia coli* with mixtures of synthetic polyribonucleotides, twenty amino acids (with only one radioactively labeled), researchers successfully determined the complete set of triplets encoding all amino acids. Furthermore, utilizing trinucleotides of known base sequences, they deciphered the specific nucleotide sequences of all codons responsible for binding various aminoacyl-tRNAs.
It is also important to highlight the work of H. G. Khorana, who proposed a method for the Chemical synthesis of poly- and oligonucleotides, and R. W. Holley, who elucidated The structure of tRNA along with its anticodon region.
Until recently, available empirical evidence pointed to the universality of the genetic code—meaning that across all organisms (Viruses, prokaryotes, and eukaryotes), the identical nucleotide triplets code for the same amino acids. In recent years, however, studies of Protein Biosynthesis in Mitochondria have revealed deviations from the universal code (Table 3.3).
Table 3.3.
Discrepancies between the mitochondrial and universal genetic codes (after R. Bohinski, 1987)
Codons |
UGA |
AUA |
AGU |
AGG |
AUU |
Universal code |
Termination |
Ile |
Arg |
Arg |
Ile |
Mitochondrial code |
Trp |
Met and initiation |
Termination |
Termination |
Ile and possibly initiation |
Another hallmark of the genetic code is its degeneracy, which means that a single amino acid may be specified by more than one codon. For example, Arginine, leucine, and Serine are each encoded by six codons; Glycine, Proline, valine, Tyrosine, and Alanine by four; whereas Tryptophan and Methionine are specified by only one codon each. Thanks to this degeneracy, errors occurring during replication and transcription do not always alter the genetic information or disrupt Gene Expression, which carries profound biological significance. In all instances of twofold, threefold, and fourfold degeneracy, the variation is confined strictly to the third nucleotide of the triplet. For instance, if alanine were encoded exclusively by the single triplet GCU, any alteration in its nucleotide sequence during replication or transcription would inevitably result in the substitution of alanine with a different amino acid in the corresponding polypeptide chain, with all the attendant consequences. However, owing to fourfold degeneracy, only substitutions affecting the first two nucleotides of a codon lead to A change in its encoded meaning. Codon Specificity is determined primarily by its first two nucleotides. As for the third nucleotide, located at the 3'-end of the oligonucleotide structure, its specificity is markedly less pronounced (A. Lehninger, 1985; R. Bohinski, 1987).
The fidelity of Polypeptide chain synthesis is achieved through the complementary recognition of the nitrogenous bases of a 5'- to 3'-oriented mRNA codon and the 3'- to 5'-oriented nitrogenous base sequence of the tRNA anticodon.
It should be noted that the number of Aminoacyl-tRNA synthetases—the Enzymes catalyzing the activation of amino acids—corresponds to the number of distinct amino acid types used in protein synthesis. As for tRNAs, their number must be at least 32, since Certain amino acids are capable of interacting with two or even three different tRNAs, which in turn recognize and bind one, two, or even three mRNA codons. Codon-anticodon recognition (where the codon on the mRNA has a 5'→3' orientation, and the anticodon on the tRNA is oriented in the 3'→5' direction) involves deviations from the classical base-pairing rules (A–T, G–C, and A–U in DNA and RNA). This deviation was formulated by F. Crick in his Wobble Hypothesis, the biological implications of which share much in common with The phenomenon of genetic code degeneracy. The wobble hypothesis posits that the third nitrogenous base, located at the 5'-end of the tRNA anticodon, can exhibit spatial flexibility, whereas the first two bases at the 3'-end of the anticodon are held more rigidly. This spatial repositioning of the nitrogenous base at the 5'-end of the anticodon enables it to form non-canonical Base Pairs that deviate from classical interactions (A–T, G–C, and A–U). It has been established that the third position from the 3'-end of the anticodon can be occupied by U, G, or I (inosine, a ribonucleoside whose nitrogenous base is hypoxanthine, formed by the deamination of the 6-amino group of adenine). Table 3.4 outlines the possible base-pairing combinations arising from the wobble phenomenon.
Table 3.4.
Possible base-pairing combinations between the 5'-end of the tRNA anticodon and the 3'-end of the mRNA codon, as determined by the wobble hypothesis
(after R. Bohinski, 1987)
Base at the 5'-end of the tRNA anticodon |
Base at the 3'-end of the mRNA codon |
I G U A* C* |
A, C, or U C or U A or G U G |
* The hypothesis does not imply novel combinations if these bases are located in the anticodon
As can be seen from Table 3.4, when hypoxanthine (as part of inosine) occupies the wobble position in the anticodon, recognition and hydrogen bonding are possible with three pairs: I — A, I — C, and I — U (although this complementary interaction is weaker compared to the conventional G — C and A — U pairs); when G or U occupies the wobble position, the number of possible combinations is limited to two: G — C, G — U and U — A, U — G. No new base-pair combinations arise when adenine and cytosine act as the wobble bases. In this case, bonding occurs According to the classical principle: A — U; C — G (Fig. 3.7).

Fig. 3.7. Scheme of codon-anticodon interactions in the context of the wobble hypothesis
(after R. Bohinski, 1987)
The formation of weak hydrogen bonds during codon-anticodon recognition can be illustrated by one of the arginine tRNAs, whose anticodon (5') I — C — G (3') is capable of interacting with three different arginine codons:

The first two codon bases (C — G) form strong Watson-Crick pairs (indicated by three bars) with the corresponding nitrogenous bases of the anticodon. The nitrogenous bases at the third position of the arginine codons (A, U, C) form weak hydrogen bonds (two bars) with the inosine residue (I) in the anticodon. Based on this and other Examples of codon-anticodon interactions, F. Crick concluded that the bases at the third position of most codons possess a certain degree of freedom when pairing with the corresponding anticodon base, i.e., they are wobble bases in Crick's terminology. The Biological Significance of this phenomenon is that it minimizes potential errors. Due to the relative weakness of the bond formed between the wobble base and the corresponding anticodon base, the tRNA is more readily released from the mRNA complex during protein synthesis. If all three base pairs were involved in a strong Watson-Crick codon-anticodon interaction, the bond strength would become a rate-limiting factor in protein synthesis by slowing down the release of tRNA from the mRNA complex (Lehninger A., 1985; Bohinski R., 1987).
Last update: 11/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.