Fundamentals of Molecular Biology. Part 2: Molecular Genetic Mechanisms - A. N. Ogurtsov 2011
The Role of RNA in Translation
Genetic Code
METABOLISM/28.html">The Genetic Code is a system for recording information about The sequence of Amino Acids in Proteins using the sequence of NUCLEOTIDES (A, T, G, C) in DNA.
Since DNA does not take a direct part in Protein Synthesis, the code is transcribed into the language of RNA (A, U, G, C).
In RNA, thymine is replaced by uracil.
4.2.1. Properties of the genetic code.
1. Triplet nature. Each amino acid is encoded by a sequence of three nucleotides (Table 1).
The code cannot be doublet, since 4 (the number of different nucleotides in DNA) is less than 20 (the number of amino acids).
Class="center">Table 1 - The genetic code
|
First nucleotide |
Second nucleotide |
Third nucleotide |
|||
|
U |
C |
A |
G |
||
|
U |
Phe |
Ser |
Tyr |
Cys |
U |
|
Phe |
Ser |
Tyr |
Cys |
C |
|
|
Leu |
Ser |
STOP |
STOP |
A |
|
|
Leu |
Ser |
STOP |
Trp |
G |
|
|
C |
Leu |
Pro |
His |
Arg |
U |
|
Leu |
Pro |
His |
Arg |
C |
|
|
Leu |
Pro |
Gln |
Arg |
A |
|
|
Leu (Met)* |
Pro |
Gln |
Arg |
G |
|
|
A |
Ile |
Thr |
Asn |
Ser |
U |
|
Ile |
Thr |
Asn |
Ser |
C |
|
|
Ile |
Thr |
Lys |
Arg |
A |
|
|
Met (START) |
Thr |
Lys |
Arg |
G |
|
|
G |
Val |
Ala |
Asp |
Gly |
U |
|
Val |
Ala |
Asp |
Gly |
C |
|
|
Val |
Ala |
Glu |
Gly |
A |
|
|
Val (Met)* |
Ala |
Glu |
Gly |
G |
|
|
* AUG is the most common initiation codon; GUG typically encodes valine, and CUG encodes leucine, but on rare occasions these codons can encode Methionine for the initiation of a protein chain |
|||||
|
Ala - Alanine, Asn - asparagine, Asp - aspartic acid, Gly - Glycine, Gln - glutamine, Glu - glutamic acid, Pro - Proline, Ser - Serine, Tyr - Tyrosine, Cys - Cysteine, Arg - Arginine, Val - valine, His - Histidine, Ile - isoleucine, Leu - leucine, Lys - Lysine, Met - methionine, Thr - Threonine, Trp - Tryptophan, Phe - phenylalanine |
|||||
The code cannot be doublet, since 16 (the number of combinations and permutations of four nucleotides taken two at a time) is less than 20.
The code can be triplet, since 64 (the number of combinations and permutations of four nucleotides taken three at a time) is greater than 20.
2. Degeneracy. Out of the 64 possible codons, 61 specify amino acids, and 3 serve as stop codons. Most amino acids are encoded by more than one codon. Only two—methionine (Met) and tryptophan (Trp)—have a single codon. On the other hand, leucine (Leu), serine (Ser), and arginine (Arg) are each encoded by six codons.
Different codons for the same amino acid are called synonyms.
The code itself is said to be degenerate, meaning that more than one codon specifies a single amino acid.
With the exception of methionine and tryptophan, all Amino acids are encoded by more than one triplet, specifically:
✵ Two amino acids are encoded by a single triplet: 2x1=2.
✵ Nine amino acids are encoded by two triplets: 9x2=18.
✵ One amino acid is encoded by three triplets: 1x3=3.
✵ Five amino acids are encoded by four triplets: 5x4=20.
✵ Three amino acids are encoded by six triplets: 3x6=18.
A total of 2+18+3+20+18=61 triplets encode the 20 amino acids.
3. Presence of punctuation marks. A Gene is a region of DNA that encodes a single polypeptide chain or a single RNA molecule.
At the end of every gene encoding a polypeptide, there is at least one of the three terminating or stop codons: UAA, UAG, UGA. They do not encode amino acids, but instead terminate Translation. Sometimes they are referred to as nonsense codons. The end of a gene invariably contains UAA, UAG, or UGA (or even two nonsense codons in a row).
By convention, punctuation marks also include the AUG codon, which is the first one following the start site. It Functions as a capital letter. In this position, it encodes formylmethionine (in prokaryotes). Any subsequent AUG within the gene simply encodes methionine.
In some bacterial mRNAs, GUG is also used as an initiation codon, and CUG is occasionally used as an initiation codon in eukaryotes.
4. Unambiguity. Each triplet encodes only a single amino acid or acts as a translation terminator. The AUG codon is an exception. In prokaryotes, at the first position (capital letter), it encodes formylmethionine, while at any other position, it encodes methionine.
5. Universality. The genetic code is identical for All living organisms on Earth. This is compelling evidence supporting a common ORIGIN AND EVOLUTION.
6. Compactness, or the absence of punctuation marks within genes. Within a gene, every nucleotide is part of a sense codon.
7. Error tolerance (robustness).
Conservative Mutations are nucleotide substitutions that do not alter the class of the encoded amino acid.
Radical mutations are nucleotide substitutions that result in A change in the class of the encoded amino acid.
Figure 33 shows the circular representation of the genetic code.
Notably, the first two bases are crucial for encoding most amino acids, whereas the third base can be arbitrary. Consequently, base substitutions in the third position of many codons will have no phenotypic effect whatsoever.
Furthermore, mutations that replace a polar residue with a non-polar one are often subtle if the residues share similar physicochemical properties. Even when such mutations manifest, the mutant protein does not completely lose its activity, but only partially. Such mutations likely gave rise to so-called synonymous proteins, which share identical folding and enzymatic activity yet possess different primary structures.
Nine single substitutions can be made within each triplet.
The total number of possible nucleotide substitutions is 61 × 9 = 549.

Figure 33 - Circular representation of the genetic code. The inner circle represents the first Base of the codon, the second circle is the second base, the third circle is the third base, the fourth circle indicates amino acids using three-letter Abbreviations, and the fifth circle denotes the polarity (P) or non-polarity (NP) of the amino acids
Out of the 549 possible nucleotide substitutions:
— 134 substitutions do not change the encoded amino acid.
— The remaining 415 mutations alter the meaning of the codons and are referred to as missense mutations.
Missense mutations:
> 23 nucleotide substitutions result in Translation termination codons. Such mutations are called nonsense mutations.
> 230 substitutions do not change the class of the encoded amino acid.
> 162 substitutions result in a change of The amino acid class, meaning they are radical.
> Out of 183 substitutions at the third nucleotide, 7 lead to translation terminators, and 176 are conservative.
> Out of 183 substitutions at the first nucleotide, 9 lead to terminators, 114 are conservative, and 60 are radical.
> Out of 183 substitutions at the second nucleotide, 7 lead to terminators, 74 are conservative, and 102 are radical.
Thus, the number of conservative substitutions is 176+114+74=364, and the number of radical substitutions is 102+60=162 (Table 2).
Table 2 - Possible point mutations
|
Mutations |
Nonsense |
Conservative |
Radical |
|
Out of 183 substitutions of the 3rd nucleotide |
7 |
176 |
0 |
|
Out of 183 substitutions of the 2nd nucleotide |
7 |
74 |
102 |
|
Out of 183 substitutions of the 1st nucleotide |
9 |
114 |
60 |
|
Total 549 substitutions |
23 |
364 |
162 |
The error tolerance of the genetic code is quantified by The ratio of conservative substitutions to radical substitutions: 364/ 162 = 2.25.
Thus, the genetic code exhibits high error tolerance regarding missense mutations—mutations that alter the meaning of codons. Notably, this resilience is most apparent in missense mutations that result in the replacement of polar residues with nonpolar ones, and vice versa.
8. Non-overlapping nature. In 1956, George Gamow proposed an overlapping code model. According to Gamow's code, every nucleotide starting from the third one in a gene belongs to three codons. Once the genetic code was deciphered, it turned out to be non-overlapping, meaning each nucleotide is part of only a single codon.
These are the eight properties of the genetic code.
4.2.2. Information Capacity of DNA. There are 6 billion people living on Earth. The hereditary information encoding them is contained in 6x109 spermatozoa. According to various estimates, humans have between 30,000 and 50,000 genes. Across all humans, this amounts to ~ (50x103)x(6x109) = 30x1013 genes, or (30x1016) = (3x1017) nucleotide pairs, which make up 1017 codons.
An average book page contains 2,500 characters. The DNA of 6x109 spermatozoa holds an Amount of Information roughly equivalent to 4x1013 book pages. These pages would take up the volume of five Derzhprom buildings. Meanwhile, 6x109 spermatozoa take up half a thimble, and their DNA occupies less than a quarter of a thimble.
4.2.3. Reading Frame. The sequence of codons between the start and stop codons is referred to as the reading frame.
This precise sequence of ribonucleotides in triplets within the mRNA determines both the linear sequence of amino acids in the polypeptide chain and the signals indicating where synthesis should begin and end.
Since the genetic code lacks "commas" and is written in non-overlapping triplet codons, a specific mRNA could theoretically be translated in three different reading frames.
Indeed, it has been demonstrated that some mRNAs contain overlapping information that can be translated in different reading frames, yielding distinct Polypeptides (Figure 34).
The vast majority of mRNAs, however, are read in only a single frame; otherwise, long before a functional protein is formed, translation is aborted by a stop codon inevitably encountered when using either of the two other "incorrect" reading frames.
Another potential cause of mistranslation arises from an accidental frameshift.
This occurs when, for instance, four nucleotides are mistakenly read as a single codon and translated into one amino acid, or conversely, a single nucleotide is skipped, causing subsequent triplet reading to proceed in a completely new reading frame. Such frameshift events are relatively rare, but several dozen Examples have been documented.

Figure 34 - Example of how the genetic code can be translated in different reading frames
4.2.4. Variations of the Genetic Code. The fact that the genetic code is identical across all organisms serves as a key argument supporting the common descent of the entire terrestrial biosphere.
However, the genetic code differs for certain codons in many Cell/35.html">Mitochondria and some Protozoa (Table 3). Most of these deviations involve reading a stop codon as an amino acid rather than substituting one amino acid for another.
Table 3 - Deviations from the universal genetic code
|
Codon |
Universal code |
Unusual code |
Occurrence |
|
UGA |
STOP |
Trp |
Mycoplasma, Spiroplasmas, mitochondria of many species |
|
CUG |
Leu |
Thr |
Yeast mitochondria |
|
UAA, UAG |
STOP |
Gln |
Acetabularia, Tetrahymena, Paramecium, |
|
UGA |
STOP |
Cys |
Euplotes |
Last update: 12/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.