BIOLOGY Volume 3 - A Guide to General Biology - 2004

23. THE CONTINUITY OF LIFE

23.7. The Nature of Genes

Class="center">23.7.1. What Are Genes?

In 1866, Mendel suggested that organismal traits are determined by inheritable units, which he termed "elements." These later became known as genes, and were shown to reside in Chromosomes that pass them down from generation to generation. Thus, Mendel would likely have defined a Gene as a unit of heredity. While this is a perfectly acceptable definition, it tells us nothing about the physical Nature of the gene.

Below, we will attempt to explain the physical nature of the gene based on two Structure/97.html">Definitions.

THE GENE AS A UNIT OF RECOMBINATION. In his work on chromosome mapping in Drosophila (Section 24.2), Morgan postulated that a gene is the shortest segment of a chromosome that can be separated from adjacent segments As a result of Crossing Over. According to this definition, a gene represents a specific chromosomal segment that determines a particular trait of the Organism.

THE GENE AS A UNIT OF FUNCTION. Since genes are known to determine the structural, physiological, and biochemical traits of an organism, it was proposed to define a gene as the smallest segment of a chromosome responsible for the synthesis of a specific product. We now know that genes encode Protein Synthesis. Therefore, a gene can be defined as a segment of DNA encoding a specific protein. This definition can be refined further by calling a gene a segment of DNA encoding a specific polypeptide, since some Proteins consist of more than one polypeptide chain and are thus encoded by more than one gene.

23.7.2. METABOLISM/28.html">The Genetic Code Is a Sequence of Bases

When Watson and Crick proposed the helical structure of DNA in 1953, they also suggested that the Genetic information passed from generation to generation and controlling Cell activity is contained within the DNA molecule in the form of a base sequence. Once it was demonstrated that DNA encodes the synthesis of protein molecules, it became clear that The base sequence in DNA must encode the Amino Acid Sequence in proteins. This relationship between bases and Amino Acids is known as the genetic code. In 1953, A number of unresolved problems remained: it was necessary to prove the existence of such a base code, decipher it, and determine how the base sequence in DNA is translated into The amino acid sequence in a protein molecule.

23.7.3. The Triplet Code

The DNA molecule is built from four types of bases: adenine (A), guanine (G), thymine (T), and cytosine (C) (Section 3.6). Each base forms part of a nucleotide, and NUCLEOTIDES are linked into a polynucleotide chain, denoted by the initial letters of their names. Thus, four "alphabet" letters make it possible to write instructions for synthesizing a potentially infinite number of different protein molecules. There are 20 amino acids from which Proteins are built, and these must be encoded by the bases that make up DNA. If THE POSITION OF a single amino acid in the Introduction/19.html">Primary Structure of a protein were determined by a single base, that protein could contain only four different amino acids. If each amino acid were encoded by two bases, such a code could specify 16 amino acids.

23.3. Make a list of the 16 possible pairwise combinations of the bases A, G, T, and C.

The incorporation of all 20 amino acids into protein molecules can only be provided by a code consisting of three bases. Such a code can yield 64 base combinations, which is more than enough. Therefore, Watson and Crick predicted that the code must be a triplet code.

23.4. Four bases used individually can code for four amino acids; used in pairs, 16 amino acids; used in triplets, 64 amino acids. Derive a mathematical expression that explains this.

It was later proven that the code is indeed triplet, meaning that each amino acid is encoded by three bases.

Evidence for the Triplet Code

Such evidence was provided by Francis Crick in 1961. Using phage T4, he obtained Mutations caused by the addition or loss of bases. The addition or deletion of a base alters the "reading" of the code downstream from the point where the change occurred (Fig. 23.23). The resulting mutation is called a "frameshift" mutation. This mutation prevents The formation of base triplet sequences that could direct the synthesis of protein molecules with the original amino acid sequence. Restoring the original base sequence would only be possible by adding or removing a single base at specific points. Such restoration would prevent the appearance of mutants among the experimental T4 phages. The addition of a single base was termed a (+) mutation, and the removal of a single base a (—) mutation. Mutations of type (+) and (—) restore the correct reading frame. Double mutants (++) or (——) also undergo a frameshift, resulting in mutants that synthesize defective proteins. However, in mutants (+++) or (———), no alteration in Protein synthesis is observed. According to Crick, this is because such mutations do not cause frame shifts, but merely result in the addition or deletion of a single amino acid, which often does not affect protein function. This means that the code is read three bases at a time, i.e., in triplets.

Fig. 23.23. Diagram illustrating the results of base deletion or addition in the triplet code. The addition of base C leads to a frameshift, so that the original message GAT, GAT... becomes TGA, TGA... The deletion of base A causes a frameshift, resulting in the replacement of the original message GAT, GAT... with ATG, ATG... The addition of base C and the deletion of base A at the points indicated in the diagram restore the original message GAT, GAT... (After F. H. C. Crick, The genetic code I, 1962, Scientific American Offprint, N 123, Wm. Saunders and Co.)

These experiments also showed that the triplets are non-overlapping, meaning each base belongs to a single triplet only. None of the bases within a given triplet is part of an adjacent triplet (Fig. 23.24).

Fig. 23.24. Base triplets in an overlapping and non-overlapping code.

23.5. Using repeating sequences of the GTA triplet and the base C, demonstrate that the original triplet sequence can only be restored by adding or removing three bases. (Present your answer in the format shown in Fig. 23.24.)

23.7.4. Deciphering the Code

Once the triplet nature of the code was established, the next challenge was to determine which triplets code for each specific amino acid, in other words—to crack the code. To understand the experimental Procedures used for this purpose, one must have a general grasp of the mechanism by which the triplet code is translated into a protein molecule.

Protein synthesis involves two interacting Types of Nucleic acids: deoxyribonucleic acid (DNA) and ribonucleic acid (RNA). There are three MAIN TYPES OF RNA: Messenger RNA (mRNA), ribosomal RNA (rRNA), and Transfer RNA (tRNA). The DNA base sequence is rewritten (transcribed) into a messenger RNA strand, which then passes from The Nucleus into the Cytoplasm. There, these strands attach to Ribosomes, and the mRNA sequence is translated into an amino acid sequence. Each amino acid is bound to a specific tRNA molecule, which attaches to the complementary base triplet on the mRNA. Brought close together by this process, the Amino acids join to form a polypeptide chain. Thus, protein synthesis requires DNA, mRNA, ribosomes, tRNA, amino acids, ATP as an energy source, as well as various Enzymes and Cofactors that catalyze each stage of the process.

Nirenberg used these data and various techniques developed in the late 1950s to design a series of experiments aimed at cracking the code. The Essence of his experiments was to use mRNA with a predetermined base sequence to determine the amino acid sequence in the polypeptide chain synthesized in the presence of that mRNA. Nirenberg managed to synthesize an mRNA consisting of repeating UUU triplets. This compound, named polyuridylic acid [poly(U)], was used as the code. A cell-free extract of E. coli containing ribosomes, tRNA, ATP, enzymes, and a single labeled amino acid was placed into each of twenty test tubes. Poly(U) was then added to each tube and left for some time to allow polypeptide synthesis to take place. Analysis of the tube contents showed that a polypeptide had formed only in the tube containing the amino acid phenylalanine. This was the first step toward deciphering the genetic code. Nirenberg demonstrated that the base triplet, or codon, UUU within the mRNA determines the position of phenylalanine in the polypeptide chain. Nirenberg and his colleagues then set out to create synthetic polynucleotide molecules corresponding to all 64 possible codons, and by 1964 they had cracked the codes for all 20 amino acids (Table 23.4).

Table 23.4. Base sequences of the triplet code and their corresponding amino acids

Note. The codons shown are base sequences in mRNA rather than DNA. The DNA genetic code contains complementary bases, with T replacing U.

* Codon signifying the termination of Polypeptide chain synthesis; equivalent to a full stop.

23.7.5. CHARACTERISTICS OF THE Genetic Code

Triplets

As already demonstrated, the genetic code is a triplet code: three bases in a DNA molecule code for one amino acid in any protein molecule. The DNA code is first transcribed into a messenger RNA complementary to that DNA. Complementary mRNA triplets are called codons. Each codon is three bases long and codes for a single amino acid. The DNA code for any amino acid can be derived by converting the RNA codons into complementary DNA base triplets according to Table 23.5.

Table 23.5. Complementarity between DNA and RNA bases

DNA bases

RNA bases

A (adenine)

U (uracil)

G (guanine)

C (cytosine)

T (thymine)

A (adenine)

C (cytosine)

G (guanine)

23.6. Write down the base sequence in the mRNA formed on a DNA strand with the following sequence:

ATGTTCGAGTACCGATGTAACG

The code is degenerate

Table 23.4 lists the codons of the genetic code. As can be seen from the table, Some amino acids are coded by more than one codon. Such a code is described as degenerate. An analysis of this code also reveals that for many amino acids, apparently only the first two letters of the codon are significant.

The code contains punctuation marks

Among the codons presented in Table 23.4, three act as "full stops," meaning they signal the end of the encoded message. An example is the UAA triplet. Such codons are sometimes called "nonsense codons"; they do not code for any amino acid. These codons presumably mark the end of a given gene, serving as "stop signals" that halt polypeptide chain synthesis during Translation.

Certain other codons, such as AUG (Methionine), act as "start signals," indicating the initiation of polypeptide chain translation.

The code is universal

One of the remarkable Features of the genetic code is that it appears to be universal. All living organisms share the same 20 Amino Acids and the same five bases (A, G, T, C, and U).

Molecular biology has now advanced to the point where it is possible to determine base sequences for entire genes and complete organisms. The first organism whose complete genetic code was successfully deciphered was a virus—phage φX174. This phage has only 10 genes, and its complete genetic code consists of 5,386 bases. The sequence of these bases was determined by Fred Sanger, the researcher who first discovered the amino acid sequence in a protein. He was awarded a Nobel Prize for each of these foundational discoveries. It is now possible to synthesize entire genes, an achievement with important Applications in Genetic Engineering. At the very beginning of the 21st century, the Human Genome Project is expected to decode the complete human genetic code, estimated to be 3,000 million Base Pairs in length. (The Genome is the entire DNA of a given organism.) Work is currently underway to decode the genomes of E. coli, the fruit fly (Drosophila), a nematode worm, and the laboratory mouse.

Summary

The Main Features of the genetic code are briefly summarized below.

1. The code determining the Incorporation of Amino acids into a polypeptide chain is a base triplet in the DNA polynucleotide chain.

2. The code is universal: the same triplets encode the same amino acids across all organisms. (A few triplet codes in Mitochondrial DNA and certain ancient Bacteria differ from the universal code.)

3. The code is degenerate: a given amino acid can be encoded by more than one triplet.

4. The code is non-overlapping: for instance, an mRNA sequence starting with the nucleotides AUGAGCgca is not read as AUG/UGA/GAG... (overlapping by two bases) or AUG/GAG/GCG... (overlapping by one base). (However, overlapping genes have been discovered in some organisms, such as bacteriophage φX174. These cases are likely very rare and may be explained by DNA economy in genomes with a very small number of genes.)



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.