LEHNINGER PRINCIPLES OF BIOCHEMISTRY - VOL. 3. INFORMATION PATHWAYS - 2017

CHAPTER III. INFORMATION PATHWAYS

24. GENES AND CHROMOSOMES

The sheer physical size of DNA presents an intriguing biological paradox, as DNA molecules are typically vastly larger than the Cells or Viral Particles that house them (Fig. 24-1). This raises the fundamental question of how such enormous molecules are compactly packaged within The Cell. To answer this, we must transition from the Introduction/11.html">Secondary Structure of DNA (see Chapter 8) to its remarkable tertiary structure, which forms the architectural basis of Chromosomes—the primary repositories of Genetic information. We begin this chapter by examining the fundamental elements of chromosomes, followed by a Discussion of their size and Organization. Next, we delve into DNA topology, exploring various modes of coiling and supercoiling. Finally, we analyze the complex interactions between DNA and Proteins that enable the compact compaction of chromosomes.

Fig. 24-1. The protein capsid of bacteriophage T2 encases the single linear DNA molecule of the phage.

Class="center">Upon osmotic lysis of the bacteriophage particles in distilled Water, the DNA was released from the capsid and dispersed across the water's surface. A T2 bacteriophage particle consists of an icosahedral HEAD and a tail apparatus used to attach to the outer surface of a bacterial cell. The entirety of the DNA visualized in this electron micrograph is normally packaged tightly within the phage head.

24.1. Chromosomal Elements

Cellular DNA contains both genes and intergenic regions; both play vital biological roles. More complex genomes, such as those of eukaryotes, require correspondingly higher levels of chromosomal organization. We begin by examining the diverse types of DNA sequences and structural elements that comprise chromosomes.

Genes are segments of DNA molecules that encode polypeptide and RNA chains

Our understanding of genes has evolved profoundly over the past century. Historically, a Gene was defined as the chromosomal region encoding or determining a single phenotypic trait, such as eye color. In 1940, George Beadle and Edward Tatum proposed a groundbreaking molecular Definition of the gene. By treating spores of the fungus Neurospora crassa with X-rays and other mutagenic agents that alter DNA sequences (Mutations), they isolated mutant fungal strains deficient in specific Enzymes, which in some cases disrupted an entire metabolic pathway. Beadle and Tatum concluded that a gene is a segment of genetic material that specifies or encodes a single enzyme, giving rise to the "one gene-one enzyme" hypothesis. This concept was later broadened to "one gene-one polypeptide," as many genes encode proteins that are not enzymes, and a polypeptide may function as a subunit of a multi-protein complex.

The modern biochemical definition of a gene is even more precise. Genes are defined as all DNA segments that encode the primary sequence of a final functional product, which may be a polypeptide or an RNA molecule possessing structural or catalytic activity. Alongside genes, DNA contains other sequences that perform exclusively regulatory Functions. Regulatory sequences may denote the start or end points of genes, influence METABOLISM/31.html">Transcription rates, or specify origins of Replication or recombination (Chapter 28). Furthermore, many genes can be expressed via alternative pathways, allowing a single DNA segment to serve as a template for multiple distinct products. The complex transcriptional and translational mechanisms governing these processes are detailed in Chapters 26–28.

We can estimate the minimum size of a gene encoding an average-sized protein. As discussed in detail in Chapter 27, each amino acid in a polypeptide chain is specified by a sequence of three NUCLEOTIDES (Fig. 24-2); these triplet sequences (codons) correspond directly to the Amino Acid Sequence of the encoded polypeptide. A polypeptide of 350 amino acid residues (an average length) corresponds to a coding sequence of 1,050 bp. However, Many eukaryotic genes and some prokaryotic genes are interrupted by non-coding DNA segments, making them significantly longer than this simple calculation suggests.

Fig. 24-2 The correspondence among DNA coding regions, mRNA, and The amino acid sequence of a polypeptide chain. Nucleotide triplets in DNA specify the amino acid sequence of a protein via an mRNA intermediary. One strand of the DNA acts as a template for mRNA synthesis, with nucleotide triplets (codons) that are complementary to the DNA triplets. In certain Bacteria and many eukaryotes, the coding sequences are interrupted by non-coding regions known as introns.

How many genes reside within a single chromosome? The chromosome of the prokaryote Escherichia coli, whose genome has been completely sequenced, is a circular DNA molecule (more accurately, a closed loop without a beginning or end) consisting of 4,639,675 bp. This sequence encodes approximately 4,300 protein-coding genes and 157 genes for stable RNA molecules. In comparison, The Human Genome comprises roughly 3.1 billion Base Pairs, corresponding to approximately 29,000 genes distributed across 24 distinct chromosomes.

DNA molecules are vastly larger than the cellular or viral structures that enclose them

Chromosomal DNA molecules are typically orders of magnitude longer than the cells or viral particles that contain them (Fig. 24-1; Table 24-1). This principle holds true across all biological kingdoms and Viruses.

Viruses.

Viruses cannot survive or replicate outside of a living host cell. They are Obligate Intracellular Parasites that hijack the metabolic machinery of the host cell to propagate. Many viral particles consist solely of a genome (typically a single RNA or DNA molecule) enclosed within a protective protein coat.

The genomes of nearly all plant viruses and certain bacterial and animal viruses are composed of RNA. These genomes are generally compact in size. For instance, the genomes of mammalian Retroviruses, such as HIV, contain about 9,000 nucleotides, whereas bacteriophage Qβ has 4,200 nucleotides. Both of these viral genomes consist of single-stranded RNA.

DNA-containing viral genomes are substantially larger (Table 24-1). Many Viral DNA molecules adopt a closed circular conformation during a phase of their life cycle. As a virus replicates within a host cell, specialized forms of viral DNA known as replicative forms may emerge; for example, many linear DNA molecules become circularized, while single-stranded DNAs form dimers. A classic medium-sized DNA virus is bacteriophage λ, which infects E. coli. The replicative form of phage λ DNA inside cells exists as a circular double-stranded helix. This double-stranded DNA contains 48,502 bp, with a contour length of 17.5 µm. The Genome of bacteriophage φX174 also consists of DNA, but is much smaller; within the viral particle, the DNA is a single-stranded circle, whereas the double-stranded replicative form contains 5,386 bp. Although viral genomes are relatively small, their extended DNA length far exceeds the physical dimensions of the viral particles that package them (Table 24-1).

Table 24-1. Dimensions of DNA and Viral Particles for Selected Bacterial Viruses (Bacteriophages)

Virus

Viral DNA Size, bp

Viral DNA Length, nm

Viral Particle Length, nm

φX 174

5,386

1,939

25

T7

39,936

14,377

78

λ

48,502

17,460

190

T4

168,889

60,800

210

Note. DNA sizes are given for the replicative (double-stranded) form. DNA lengths were estimated by assuming a base pair length of 3.4 Å (see Fig. 8-13 in Vol. 1).

Bacteria.

A single E. coli cell contains approximately 100 times more DNA than a bacteriophage $\lambda$ particle. The E. coli bacterium has a single double-stranded circular DNA molecule. It consists of 4,639,675 bp and is approximately 1.7 mm long, which exceeds the length of the E. coli cell itself by about 650 times (Fig. 24-3). In addition to the large circular chromosome within the nucleoid, many bacteria contain one or more small circular DNA molecules that reside freely in the Cytosol. These extrachromosomal elements are called Plasmids (Fig. 24-4; see also p. 439 in Vol. 1). Most plasmids consist of only a few thousand base pairs, while some contain more than 10,000 bp. They carry genetic information and replicate to form daughter plasmids that are distributed into daughter cells when the parent cell divides. Plasmids are found not only in bacteria but also in Yeasts and other Fungi.

Fig. 24-3. The E. coli chromosome (1.7 mm long) shown in linear form; an E. coli cell (2 µm) is shown alongside for comparison.

Fig. 24-4. DNA from a lysed E. coli cell. White arrows indicate circular plasmid molecules. White and black spots are artifacts.

In many cases, plasmids confer no apparent advantages on their host cells, and their sole function is independent replication. However, some plasmids carry genes beneficial to the host. For example, plasmid-borne genes can confer resistance to antibacterial agents. Plasmids carrying the $\beta$-lactamase gene provide resistance to $\beta$-lactam Antibiotics such as penicillin and amoxicillin (see Fig. 6-28 in Vol. 1). Plasmids can be transferred from antibiotic-resistant cells to other Cells of the same or different bacterial species, rendering those cells resistant as well. The intensive use of antibiotics serves as a potent selective pressure promoting the spread of Antibiotic Resistance plasmids (as well as Transposons encoding similar genes) among pathogenic bacteria, leading to The Emergence of multi-drug-resistant bacterial strains. Physicians are increasingly recognizing the dangers of widespread antibiotic use and prescribe them only when strictly necessary. For similar reasons, the extensive use of antibiotics in livestock farming is being restricted.

Eukaryotes.

A Yeast cell, one of the smallest eukaryotes, contains 2.6 times more DNA than an E. coli cell (Table 24-2). Cells of the fruit fly Drosophila, a classic subject of genetic research, contain 35 times more DNA, and human cells contain approximately 700 times more DNA than an E. coli cell. Many plants and amphibians contain even larger amounts of DNA. The genetic material of Eukaryotic cells is organized into chromosomes. The diploid chromosome Complement (2n) varies by Organism (Table 24-2). For instance, a human somatic cell has 46 chromosomes (Fig. 24-5). As shown in Figure 24-5a, each eukaryotic chromosome contains a single very large double-helical DNA molecule. The 24 Human chromosomes (22 matched pairs and two sex chromosomes, X and Y) vary in length by more than 25-fold. Each eukaryotic chromosome contains a specific set of genes.

Table 24-2. DNA, Genes, and Chromosomes of Selected Organisms


Total DNA, bp

Chromosome Numbera

Approximate Gene Number

Escherichia coli (bacterium)

4,639,675

1

4,435

Saccharomyces cerevisiae (yeast)

12,080,000

16b

5,860

Caenorhabditis elegans (nematode)

90,269,800

12c

23,000

Arabidopsis thaliana (plant)

119,186,200

10

33,000

Drosophila melanogaster (fruit fly)

120,367,260

18

20,000

Oryza sativa (rice)

480,000,000

24

57,000

Mus musculus (mouse)

2,634,266,500

40

27,000

Homo sapiens (human)

3,070,128,600

46

29,000

Note. Information is continuously updated; for the most recent data, consult websites dedicated to individual genome projects.

a For all eukaryotes except yeast, the diploid chromosome number is given.

b Haploid number. Wild strains of yeast typically have eight (octaploid) or more sets of such chromosomes.

c For females with two X chromosomes. Males have one X chromosome and no Y, i.e., 11 chromosomes in total.

Fig. 24-5. Eukaryotic chromosomes. (a) A pair of linked and condensed sister chromatids from a human chromosome. Eukaryotic chromosomes exist in this form following replication and during metaphase in mitosis. (b) A complete set of chromosomes from a leukocyte of one of the book's authors. Every normal human somatic cell contains 46 chromosomes.

If the DNA molecules of the human genome (22 chromosomes plus the X and Y, or X and X chromosomes) were placed end-to-end, they would form a sequence about one meter long. Most human cells are diploid, so the total DNA length in such cells is about 2 m. An adult human has approximately 1014 cells; thus, the total length of all DNA molecules combined is 2 × 1011 km. For comparison, the Earth's circumference is 4 × 104 km, and the distance from the Earth to the Sun is 1.5 × 108 km. This illustrates the astonishingly compact packaging of DNA within our cells!

Eukaryotic cells also contain other DNA-containing Organelles: Mitochondria and Chloroplasts. Mitochondrial DNA (mtDNA) molecules are much smaller than nuclear chromosomes. The size of double-stranded circular mtDNA in animal cells is less than 20,000 bp (in human mitochondria, the mtDNA molecule consists of 16,569 bp). Each mitochondrion typically carries two to ten copies of mtDNA molecules; this number can increase to hundreds in certain cells, such as differentiating embryonic cells. In some organisms (e.g., trypanosomes), each mitochondrion contains thousands of copies of mtDNA, with the mitochondria organized into a complex structure called a kinetoplast. The sizes of plant cell mtDNA range from 200,000 to 2,500,000 bp. Chloroplast DNA (cpDNA) molecules are also double-stranded and circular, ranging in length from 120,000 to 160,000 bp. Many hypotheses have been proposed regarding THE ORIGIN OF Mitochondrial and Chloroplast DNA. The currently accepted view is that they represent remnants of ancient bacteria that invaded the Cytoplasm of host cells and evolved into the precursors of these organelles (see Fig. 1-36, Vol. 1). Mitochondrial DNA encodes mitochondrial tRNAs and rRNAs, as well as a few mitochondrial proteins. Over 95% of mitochondrial proteins are encoded by nuclear DNA. Mitochondria and chloroplasts divide along with the cell. The DNA of these organelles replicates prior to and during Cell Division, after which daughter mtDNA molecules are distributed into the organelles of the daughter cells.

Fig. 24-6. A dividing mitochondrion. Some mitochondrial proteins and RNA molecules (not visible in the photograph) are encoded by a single copy of mitochondrial DNA. Mitochondrial DNA (mtDNA) replicates every time the mitochondrion divides, preceding cell division.

Eukaryotic genes and chromosomes have a highly complex organization

Many bacterial species have only a single chromosome, and in almost all cases, each chromosome contains a single copy of every gene. Only a few genes, such as those for rRNAs, are present in multiple copies. Genes and regulatory sequences comprise virtually the entire prokaryotic genome. Furthermore, almost every gene corresponds directly to the amino acid sequence (or RNA sequence) that it encodes (Fig. 24-2).

The Structural and functional Organization of Eukaryotic genes is considerably more complex. The Study of eukaryotic chromosomes, and later the sequencing of complete Eukaryotic Genomes, revealed many surprises. Many, if not most, eukaryotic genes share a striking feature: their nucleotide sequences contain one or more DNA regions that do not encode the amino acid sequence of the polypeptide product. These noncoding insertions disrupt the direct correspondence between The nucleotide sequence of a gene and the amino acid sequence of the encoded polypeptide. These noncoding segments within genes are called introns (or intervening sequences), and the coding segments are called exons. In prokaryotes, very few genes contain introns.

In typical higher eukaryote genes, intron sequences are generally much longer than exon sequences. For example, in the gene encoding the single polypeptide chain of the egg white protein Ovalbumin (Fig. 24-7), introns are significantly longer than exons: seven introns together make up 85% of the gene's DNA. In the gene for the Hemoglobin β-subunit, more than half of the DNA is contained within a single intron. The gene for the Muscle protein titin holds the record for the number of introns, with a staggering 178. Histone genes, by contrast, appear to have no introns. In most cases, the function of introns remains unknown. Overall, only about 1.5% of human DNA is «coding», meaning it carries information for proteins or RNA. However, when taking large introns into account, human DNA is found to consist of about 30% genes.

Fig. 24-7. Introns in two eukaryotic genes. The ovalbumin gene contains seven introns (A through G) separating the coding sequences of eight exons (L, 1 through 7). The hemoglobin β-subunit gene carries two introns and three exons, including one intron that contains more than half of the gene's base pairs.

Since genes make up a relatively small fraction of the human genome, a substantial portion of DNA remains unaccounted for. Figure 24-8 illustrates the types of sequences in the genome using a pie chart. Much of the noncoding DNA exists as repetitive sequences of several types. Most surprisingly, about half of the human genome consists of repetitive sequences of Mobile Genetic Elements—DNA segments ranging from a few hundred to several thousand base pairs in length that move from place to place within the genome. These mobile genetic elements (transposons) are Examples of molecular parasites that efficiently propagate within the host genome. Many transposons carry genes for proteins that catalyze the transposition process, which is described in more detail in Chapters 25 and 26. Some transposons in the human genome are active and move at a low frequency, but the majority are inactive relics modified by mutations over the course of evolution. Although these elements most often do not encode functional proteins or RNA, they have played a crucial role in Human Evolution because the movement of transposons has led to the redistribution of other genomic sequences.

Fig. 24-8. Types of sequences in the human genome. The diagram shows the three Main Components of the genome: transposons (mobile genetic elements), genes, and mixed sequences. Four Major Classes of transposons are known (three are shown in the diagram). Long interspersed nuclear elements (LINEs), ranging from 6 to 8 kbp (1 kbp = 1000 bp) in length, typically contain several genes for proteins that catalyze transposition. There are approximately 850,000 LINEs in the genome. Short interspersed nuclear elements (SINEs) are roughly 100 to 300 bp long. The human genome contains about 1.5 million of them, of which more than 1 million are Alu elements—so named because they typically contain a single restriction site for the restriction endonuclease AluI (see Fig. 9-2 in Vol. 1). The genome also harbors 450,000 copies of retrovirus-like transposons ranging from 1.5 to 11 kbp in length. Although they are «captured» by the genome and cannot move from Cell to Cell, they are evolutionarily related to retroviruses (Chapter 26), which include HIV. Another class of transposons (accounting for <3%, not shown here) is represented by transposon fragments of varying lengths.

Approximately 30% of the genome consists of sequences containing protein genes, but only a small fraction of this DNA is located in exons (coding sequences). Mixed sequences include simple sequence repeats (SSRs) and segmental duplications (SDs); the latter occur in more than one copy at various positions. As noted in Chapter 26, nearly the entire genome is transcribed into RNA sequences, with many RNA regions yet to be characterized. In addition, the genome contains remnants of transposons that have evolved to be so heavily altered that they are difficult to identify.

Roughly another 3% of the human genome consists of highly repetitive sequences called simple sequence DNA or simple sequence repeats (SSRs). These short sequences, typically less than 10 bp in size, are sometimes repeated millions of times within the cell. Simple sequence DNA is also referred to as satellite DNA because, when cell DNA fragments are centrifuged in a cesium chloride density gradient, their unusual base composition often causes them to migrate as separate bands (like satellites) accompanying the rest of the DNA. Simple sequence DNA does not encode proteins or RNA. However, unlike mobile genetic elements, satellite DNA may perform specific functions in human cells, as many simple DNA sequences are localized to two specific regions of eukaryotic chromosomes—the centromere and the telomeres.

The centromere (Fig. 24-9) is a DNA sequence to which proteins attach during cell division, binding the chromosome to the mitotic spindle. This interaction is essential for the uniform and precise segregation of chromosome sets into daughter cells. Centromeres of Saccharomyces cerevisiae have been isolated and studied. The sequences critical for centromere function are about 130 bp long and contain a high proportion of A=T pairs. Centromeres of higher eukaryotes are much longer and, unlike those of yeast, typically contain simple sequence DNA consisting of thousands of tandem copies of one or more short sequences 5 to 10 bp in length in the same orientation. The exact role of simple repeats in centromere function remains to be fully elucidated.

Fig. 24-9. Key Structural Features of a yeast chromosome.

Telomeres (from the Greek telos, meaning end) are sequences at the ends of eukaryotic chromosomes that stabilize them. Telomeres terminate in repeated sequences of the form

(5') (TxGy)n

(3') (АxСy)n

where x and y typically range from 1 to 4 (Table 24-3). The number of telomeric repeats, n, is between 20 and 100 in most Unicellular Eukaryotes, whereas in mammals it generally exceeds 1,500. The ends of linear DNA molecules cannot be fully replicated by the cellular replication machinery in the standard manner (this is likely one of the reasons for the circular structure of bacterial DNA). Telomeric repeats are added to the ends of eukaryotic chromosomes primarily by the enzyme telomerase (see Fig. 26-39).

Table 24-3. Telomeric sequences


Telomeric repeat sequence

Homo sapiens (human)

(TTAGGG)n

Tetrahymena thermophila (ciliate protozoan)

(TTGGGG)n

Saccharomyces cerevisiae (yeast)

((TG)1-3(TG)2-3)n

Arabidopsis thaliana (plant)

(TTTAGGG)n

To study the functional role of eukaryotic chromosome structural elements, artificial chromosomes were constructed (Chapter 9 in Vol. 1). For a stable linear artificial chromosome to function properly, only three components are required: a centromere, telomeres at each end of the sequence, and a replication origin site. Yeast artificial chromosomes (YACs; see Fig. 9-7) were developed as a tool for biotechnological research. Human artificial chromosomes (HACs; see Box 9-2) have been designed for the Treatment of Genetic Disorders via somatic Gene Therapy.

Summary of Section 24.1 Chromosome Elements

■ A gene is a chromosomal segment that carries information for a functional polypeptide or RNA molecule. Along with genes, chromosomes contain diverse regulatory sequences involved in replication, transcription, and other processes.

■ Viral genomic DNA and RNA are typically several orders of magnitude longer than the cells or viral particles that contain them.

■ Many eukaryotic genes (but only a few bacterial and archaeal genes) are interrupted by noncoding sequences—introns. The coding regions separated by introns are called exons.

■ Less than a third of human genomic DNA consists of genes. The remaining DNA is rich in repetitive sequences of various types. Parasitic Nucleic Acids known as transposons make up roughly half of the human genome.

■ Eukaryotic chromosomes contain two important repetitive DNA sequences with specific functions: centromeres (mitotic spindle attachment sites) and telomeres, located at the ends of chromosomes.



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.