BIOCHEMISTRY - L. Stryer - 1984

VOLUME 3

Part IV INFORMATION

CHAPTER 26. THE GENETIC CODE AND THE RELATIONSHIP BETWEEN GENES AND PROTEINS

26.10. Gene Base Sequence and Polypeptide Amino Acid Sequence Are Colinear

Let us now turn to the relationship between genes and Proteins. As shown by Seymour Benzer's high-resolution genetic mapping work, genes are unbranched structures. This important finding was consistent with the established fact that DNA is a linear sequence of Base Pairs. Polypeptide chains also possess an unbranched Structure. Consequently, in the early 1960s, the following question arose: is there a linear correspondence between a Gene and its polypeptide product?

Charles Yanofsky's approach to this problem involved using E. coli mutants that produced an altered enzyme molecule. Many mutants affecting the a-chain of Tryptophan synthase were isolated, and the positions of the Mutations on the Genetic Map of the a-chain were determined using recombination experiments with a transducing phage. Some of these mutations mapped close to each other on the genetic map, whereas others were widely separated within the same gene. The next crucial task was to determine THE POSITION OF The amino acid substitution for each of these ten mutants. First, The sequence of 168 Amino Acids of the wild-type a-chain was established. Then, using the fingerprinting method, the Location and Nature of the amino acid substitution in each case were determined. The order of mutations on the genetic map was identical to the order of the corresponding substitutions in the Amino Acid Sequence of the polypeptide product (Fig. 26.5). In other words, the gene encoding the α-chain and its polypeptide product are colinear.

Class="center">Fig. 26.5. Colinearity of the gene and the amino acid sequence of the tryptophan synthase a-chain. The positions of mutations in the DNA (yellow line) were determined by genetic mapping Methods. The Amino Acid Substitutions in the amino acid sequence (blue line) are arranged in the same order as the corresponding mutations

26.11. Some Viral DNA Sequences Encode More Than One Protein

The discovery that the DNA of phage ɸX174 encodes more proteins than its nucleotide count would seem to allow was utterly puzzling. How can this virus encode more than 2,000 amino acid residues if it contains only 5,375 NUCLEOTIDES? The answer came from the complete base Sequence Determination (Section 24.29), which revealed that some genes in the ɸX174 DNA overlap. Specific Regions of the corresponding transcripts are translated in different reading frames, resulting in The production of proteins with distinct Amino acid sequences (Fig. 26.6). For example, the exact same 300 nucleotides encode the protein of gene E and a major portion of the protein of gene D. An even more striking example is provided by the related bacteriophage G4, in the DNA of which certain short regions encode three different proteins. These Viruses utilize overlapping genes to pack more information into small DNA molecules. However, this genetic economy comes at a price: strict constraints are imposed on the amino acid sequences encoded by overlapping genes. Consequently, overlapping genes appear to be used extensively only when The amount of DNA is strictly limited, as in the case of viruses with protein coats of strictly defined dimensions.

Fig. 26.6. Overlapping genes

in ɸX174 phage DNA. Adjacent to The base sequence are two amino acid sequences determined by different reading frames

26.12. Eukaryotic Genes Are Mosaics of Translated and Untranslated DNA Sequences

In Bacteria, polypeptide chains are encoded by a continuous sequence of triplet codons. For many years, it was assumed that The genes of higher organisms were likewise continuous. This view was unexpectedly disproven in 1977 when several laboratories discovered that certain genes have a discontinuous structure. For example, the gene for the Hemoglobin β-chain is interrupted within the region encoding the amino acid sequence by a long noncoding intervening sequence of 550 base pairs and a short sequence of 120 base pairs. Thus, the β-globin gene is divided into three coding sequences:

This remarkable structure was discovered through electron microscopic studies of hybrids formed between β-globin mRNA and a mouse DNA fragment containing the β-globin gene (Fig. 26.7). The double-stranded DNA is partially denatured, allowing the mRNA to hybridize with the complementary DNA strand. The single-stranded DNA region then loops out and appears as a thin line in electron micrographs, in contrast to double-stranded DNA or DNA-RNA hybrid regions, which appear significantly thicker. If the β-globin gene were continuous, a single loop would be visible. However, Cytology/cytology/93.html">ELECTRON MICROGRAPHS OF such hybrids (Fig. 26.8) clearly reveal three loops. This demonstrates that the gene is interrupted by at least one region of DNA that is absent from the corresponding mRNA. Additional data on intervening sequences were obtained by Restriction mapping of the β-globin gene and the reverse METABOLISM/31.html">Transcription product of mRNA. The large discrepancies between these maps showed that the genomic DNA contains untranslated sequences interspersed between the coding sequences. Restriction maps enabled the precise localization of such intervening sequences.

Fig. 26.7. Detection of intervening sequences by Electron Cell/15.html">Microscopy. An mRNA molecule (red line) hybridizes with genomic DNA containing the corresponding gene. A – if the gene is continuous, a single loop of single-stranded DNA (shown in blue) is visible; B – if the gene contains an intervening sequence, two loops of single-stranded DNA (shown in blue) and one loop of double-stranded DNA (blue and yellow) are visible

Fig. 26.8. Electron micrograph of a hybrid between ß-globin mRNA and a genomic DNA fragment containing the ß-globin gene. The thick loops of double-helical DNA represent intervening sequences in the DNA that are absent from the mRNA (as in Fig. 26.7, B). The upper arrow points to the large intervening sequence, and the lower arrow points to the small one

At what stage of Gene Expression are the intervening sequences removed? Newly synthesized RNAs isolated from The Nucleus are significantly longer than the mRNA molecules derived from them. Specifically, the primary transcript of the ß-globin gene contains two untranslated regions. These intervening sequences in the primary 15S transcript are excised, and the coding sequences are simultaneously joined together by the action of a splicing enzyme. This enzyme operates with high precision, yielding mature 9S mRNA (Fig. 26.9). The coding sequences of discontinuous (“split”) genes are called exons, and the intervening sequences are called introns (derived from expressed regions and intervening sequence, respectively).

Fig. 26.9. Transcription of the ß-globin gene and removal of intervening sequences from the primary RNA transcript. The formation of the 5' cap and The addition of the poly(A) tail are discussed in Section 29.22

Another interrupted eukaryotic gene is the chicken Ovalbumin gene, which consists of eight exons separated by seven long introns (Fig. 26.10). Even more remarkable is the conalbumin gene, which contains no fewer than 17 exons. A common property of the expression of these genes is that their exons are arranged in the mRNA and in the DNA in the exact same sequence. Thus, interrupted genes, much like continuous genes, are collinear with their polypeptide products.

Fig. 26.10. STRUCTURE OF THE chicken ovalbumin gene. Introns (non-coding regions) are shown in yellow, and exons in blue.

All avian and mammalian genes mapped to date, with the exception of histone genes (Section 29.13), contain introns. Why do virtually all higher eukaryotic genes contain intervening sequences? One possible answer is that interrupted genes reflect the evolutionary process. Exons may correspond to major structural or functional elements (domains) that joined together to form proteins with novel properties. Another possibility is that the excision of intervening sequences regulates the flow of mRNA from the nucleus to the Cytosol. According to this hypothesis, the splicing of the primary transcript plays a pivotal role, as it determines which proteins The Cell synthesizes. The discovery of interrupted genes in higher organisms has opened up a fascinating field of study that appears to be of great importance for understanding cell growth and differentiation.



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.