Molecular Biology of the Cell - Volume 2 - Alberts B., Bray D., Lewis J., Raff M., Roberts K., Watson J. 1993

Intracellular macromolecular sorting and maintenance of cellular compartments
DNA and proteins associated with chromosomes

During the first 40 years of our century, biologists did not take seriously the suggestion that the DNA contained within Chromosomes carries Genetic information. This was partly due to the then-prevalent misconception that Nucleic Acids consisted merely of simple, regularly repeating tetranucleotide sequences (for example, AGCTAGCTAGCT...). Today we know that DNA is an extremely long, unbranched linear polymer that can contain many millions of NUCLEOTIDES arranged in an irregular, yet far from random, sequence: this nucleotide sequence in the DNA polynucleotide chain serves as a coded repository of hereditary information. The linear four-letter Genetic Code—whose symbol words are triplets of nucleotides (codons that specify Amino Acids, see Section 5.1.6)—makes it possible to store a vast Amount of Information within a very small volume. In DNA, 1 million "letters" (nucleotides) fit into a straight segment measuring 3.4 x 105 nm (0.034 cm) in length and occupy a volume of approximately 106 nm3 (10-15 cm3).

Each DNA molecule is packaged into a separate chromosome, and all the genetic information stored in an Organism's chromosomes is called its genome. The Genome of the bacterium E. coli contains 4.7 x 106 nucleotide pairs, which make up a single DNA molecule (one chromosome). The Human Genome is represented by 6 x 109 nucleotide pairs distributed among 46 chromosomes (22 pairs of autosomes and 2 distinct sex chromosomes), and therefore consists of 24 types of DNA molecules. Diploid organisms, such as you and me, possess two copies of each chromosome type, one inherited from the mother and the other from the father (with the exception of male sex chromosomes, where the Y chromosome is always received from the father and the X chromosome from the mother). Thus, a typical human Cell contains 46 chromosomes and about 6 x 109 DNA nucleotide pairs. Other mammals have genomes of approximately the same size. Theoretically, this amount of DNA could be packed into a cube with a side length of 1.9 µm. By comparison, 6 x 109 letters in such a book would occupy more than a million pages, and its volume would be 1017 times larger.

In this section, we will discuss the relationship between DNA molecules and chromosomes and examine the diverse Proteins that bind to DNA molecules, transforming them into active eukaryotic chromosomes. Some of these proteins control the expression of genetic information by regulating the synthesis of RNA molecules at specific Regions of the genome. Other proteins, primarily Histones, fold each long DNA molecule in such a way that it becomes compact and orderly while still preserving access to the necessary genetic information.

9-3

9.1.1. Each Chromosome Is Formed from a Single Long DNA Molecule [2]

Each individual human chromosome contains from 50 x 106 to 250 x 106 nucleotide pairs. In an uncoiled state, DNA molecules of this size would range from 1.7 to 8.5 cm in length, yet upon the removal of chromosomal proteins, even the slightest mechanical stress causes them to break. Intact DNA molecules can be isolated from certain lower eukaryotic organisms, such as the Yeast Saccharomyces cerevisiae, whose chromosomes are considerably shorter. Using pulsed-field gel Electrophoresis, it has been demonstrated that each yeast chromosome consists of a single linear DNA molecule. These findings are consistent with the results of highly complex measurements based on the degree of supercoiling. Corresponding experiments were performed on Drosophila chromosomes, whose DNA molecules are approximately the same length as Human chromosomes. Data obtained through various Methods lead to the Conclusion that all chromosomes contain only a single DNA molecule.

9-4

9-5

9.1.2. Each DNA Molecule Constituting a Chromosome Must Contain a Centromere, Two Telomeres, and Replication Origins [3]

For a DNA molecule to form an active chromosome, it must possess The ability to replicate, segregate during mitosis, and be maintained through generations of Cells. The application of recombinant DNA techniques to yeast cells made it possible to isolate and identify the elements that "convert" a nucleotide sequence into a chromosome. Two of these three elements were identified through The Study of small circular DNA molecules that replicate autonomously in Cells of the yeast Saccharomyces cerevisiae. It turned out that replication of such a molecule requires a special sequence that acts as a METABOLISM/36.html">DNA replication origin (also referred to as an autonomously replicating sequence). Each yeast chromosome contains several such origins. The second element required for a DNA sequence to function as a chromosome is called a centromere. The centromere connects the DNA molecule containing it to the mitotic spindle during M phase (see Section 13.5.3). Each yeast chromosome has only a single centromere. If a DNA segment acting as a centromere is inserted into a plasmid, Cell Division ensures that each daughter cell receives one of the two copies of the newly replicated plasmid DNA molecule.

The third essential element of a chromosome is a telomere, which must be present at each end of a linear chromosome. If a circular plasmid containing a replication origin and a centromere is cleaved at any site, it will continue to replicate and remain attached to the mitotic spindle, yet it will still be lost in subsequent cell generations. This occurs because replication of the lagging strand requires a DNA sequence ahead of the copied region to serve as a template for An RNA primer (see Fig. 5-43). Because such sequences are lacking for the very last few nucleotides of a linear DNA molecule, its strands become shorter with each new round of replication. In Bacteria and Viruses, the chromosome is circular, and therefore such difficulties at the ends of replication do not arise. Eukaryotic cells, whose chromosomes are linear, have evolved a special telomeric DNA sequence. This is a simple, repeating nucleotide sequence that is periodically extended by a specialized enzyme (see Section 9.3.5).

Class="center">

Fig. 9-4. Function of the three DNA sequence elements required for The formation of stable linear eukaryotic chromosomes. Telomeric sequences prevent the shortening of chromosomes that would otherwise occur with each cycle of DNA replication. Centromeres serve to align DNA molecules on the mitotic spindle during M phase. Replication origins (initiation sites) are necessary for the formation of replication forks in S phase.

Thus, lost telomeric DNA is replenished, enabling the linear chromosome to replicate fully. Figure 9-4 illustrates the Mechanisms of action of the three DNA sequence elements that ensure the stability of a linear chromosome in a yeast cell. Apparently, similar elements are also required to maintain chromosome stability in human cells. However, human DNA replication origins and centromere sequences remain far from fully characterized, and the corresponding yeast sequences have been found not to function in higher eukaryotic cells. On the other hand, recombinant constructs consisting of human and yeast DNA are capable of replicating in yeast cells as artificial chromosomes. Consequently, yeast cells can be used to generate human Genomic Libraries (see Section 5.6.3) in which each DNA clone, propagated as an artificial chromosome, contains up to a million nucleotide pairs of human DNA sequence (Fig. 9-5).

Fig. 9-5. Yeast artificial chromosome (YAC) vector, which allows the cloning of very large DNA molecules. The telomere, centromere, and replication Water/144.html">Origin of the yeast Saccharomyces cerevisiae are designated as TEL, CEN, and ARS, respectively (ARS stands for autonomously replicating sequence. The replication origin enables the plasmid to replicate outside the host cell chromosomes). BamHI and EcoRI are Restriction Endonucleases that cleave the DNA double helix at strictly defined sites. The sequences designated "A" and "B" encode Enzymes that serve as selectable markers for the Selection of transformed yeast cells carrying the artificial chromosome. (Modified from D. T. Burke, G. E. Carle, and M. V. Olson, Science 236: 806-812, 1987.)

9.1.3. The Vast Majority of Chromosomal DNA Does Not Code for Vital Proteins or RNAs [4]

The genomes of higher organisms appear to contain a large excess of DNA. It became clear long ago that the relative amount of DNA in the haploid genomes of various organisms is not directly related to organismal complexity: for example, human cells contain 700 times more DNA than E. coli, while at the same time, the cells of certain amphibians and plants contain 30 times more DNA than human cells (Fig. 9-6). Moreover, the DNA content in the genomes of different amphibian species can vary by a factor of 100.

Population geneticists have attempted to estimate The amount of DNA in higher organisms that codes for cellular proteins or participates in The regulation of genes responsible for the synthesis of such proteins. Their approach was as follows: every Gene is always subject to a small probability of mutation—a random alteration of nucleotides in DNA. The larger the number of genes, the higher the probability that a mutation will occur in at least one of them. Because most Mutations lead to an Impairment of the activity of the gene in which they occur, the mutation rate imposes a limit on the number of essential genes. Taking this argument into account and based on observed mutation rates, one can conclude that no more than a few percent of the mammalian genome is involved in regulating or encoding vital proteins. Additional evidence supporting this conclusion is presented below. Based on these considerations, a very important deduction can be made: although the mammalian genome is in principle large enough (3 x 109 nucleotides) to encode nearly 3 million average-sized proteins, the constraints imposed by The fidelity of DNA replication mean that no organism can have more than 60,000 essential proteins (disregarding THE CONTRIBUTION OF alternative RNA splicing). Thus, from a genetic standpoint, a human is probably only about 10 times more complex than the fruit fly Drosophila, which has roughly 5,000 genes.

Fig. 9-6. The amount of DNA in the haploid genome can vary by a factor of 100,000 between the smallest Prokaryotic Cells and the largest cells of certain plants and amphibians. Note that the Genome Size in humans (3 x 109 nucleotide pairs) is much smaller than that of several other organisms.

Whatever the function of the excess DNA in higher eukaryotic cells may be (see Chapter 10), the data presented in Fig. 9-6 demonstrate that higher eukaryotic cells containing large amounts of extra DNA do not suffer As a result. Indeed, even their essential coding sequences are frequently interrupted by long stretches of noncoding DNA.

9.1.4. Each Gene Is a Complex, Functionally Active Unit Dedicated to the Regulated Synthesis of an RNA Molecule

The primary function of the genome is The production of RNA molecules. Specific regions of the DNA nucleotide sequence are copied to form corresponding RNA sequences that either encode proteins (such as mRNAs) or form "structural" RNAs, such as tRNA or rRNA molecules. Each region of a DNA molecule on which an active RNA molecule is synthesized is called a gene.

Genes within the chromosomes of higher eukaryotes can contain up to 2 million nucleotide pairs, and genes exceeding 100,000 nucleotide pairs in size are quite common (Table 9-1). Meanwhile, coding for an average-sized protein (containing 300 to 400 amino acid residues) requires a mere 1,000 nucleotide pairs. Most of the excess regions consist of long sequences of noncoding DNA interspersed with relatively short segments of coding DNA. The coding sequences are called exons, and the noncoding sequences separating them are called introns. The RNA molecule synthesized from such a gene is termed a primary transcript. For a primary transcript to become an mRNA, it undergoes Processing, during which the noncoding sequences are removed and the coding sequences are "spliced" together into a single continuous molecule (RNA splicing).

Large genes consist of a long series of alternating exons and introns. In addition, each gene contains regulatory DNA sequences that bind regulatory proteins to control Transcription. Many of these regulatory sequences are located upstream (at the 5' end) of the site where RNA transcription begins, but they can also be found within introns, downstream (at the 3' end) of the site where RNA transcription ends, or even within exons. A typical vertebrate chromosome is schematically depicted in Fig. 9-7, which also illustrates one of the many genes it contains.

Table 9-1. Size of Some Human Genes in Kilobases


Gene size 1)

mRNA size

Number of introns

ß-globin

1,5

0,6

2

Insulin

1,7

0,4

2

Protein kinase C

11

1,4

7

Albumin

25

2,1

14

Catalase

34

1,6

12

LDL receptor

45

5,5

17

Factor VIII

186

9

25

Thyroglobulin

300

8,7

36

Dystrophin1)

more than

17

more than


2000


50

1) The specified gene size includes its transcribed region along with adjacent regulatory DNA sequences (According to Victor McKusick).

2) A modified form of this gene causes Duchenne muscular dystrophy.

Fig. 9-7. Organization of genes in a typical vertebrate chromosome. Proteins that bind to DNA in regulatory regions determine gene transcription. Regulatory sequences are typically located at the 5' end of the gene (as shown in the diagram), but they can also reside in introns, exons, and at the 3' end. During the formation of Messenger RNA (mRNA) molecules, intron sequences are removed from primary RNA transcripts. The data presented here regarding the number of genes in the chromosome represent a minimal estimate.

9-6

9.1.5. Comparison of DNA from related organisms reveals conserved and nonconserved regions [5]

As DNA Sequencing Methods improve, the identification of all 3 x 109 nucleotides making up the human genome is becoming increasingly realistic. However, special approaches are required to identify the small fraction of the sequence (less than 10%) that is actually functional. One way to solve this problem is to sequence corresponding regions of the genome of a related species, such as the mouse. Humans and mice are believed to have diverged from a common ancestor about 80 x 106 years ago. This is sufficient time for roughly two out of every three nucleotides to have undergone changes through random mutations. Consequently, the few regions that turn out to be similar in both genomes (conserved sequences) represent precisely those stretches where mutation leads to functional impairment. Organisms carrying such deleterious mutations are presumably eliminated by natural selection. Nonconserved regions correspond to noncoding DNA located between genes and within introns. Such DNA sequences do not exert such a strong influence on function. Conversely, conserved regions contain functionally important exons and regulatory areas. By studying the results of this long-term "experiment" conducted by nature, one can identify the most interesting regions of genomes. Comparison of DNA nucleotide sequences across different species indicates that in vertebrates, more than 90% of the DNA has no essential significance.

9.1.6. Chromosomes contain diverse proteins associated with specific DNA sequences

The information stored in DNA is organized, replicated, and read by a variety of DNA-binding proteins (DNA-binding proteins). Some of these proteins bind relatively nonspecifically along the entire length of the molecule and participate in its packaging without interfering with the functioning of other proteins. These packaging proteins are discussed below (see Section 9.1.17). Other proteins associate with specific short DNA sequences that are often evolutionarily conserved across various genomes (see Fig. 10-34). Such site-specific DNA-binding proteins serve A wide variety of Functions. Some of them presumably participate in folding the long DNA molecule to form distinct domains, others facilitate the initiation of DNA replication, and many control gene transcription. Each cell type in a multicellular organism contains a diverse mixture of such regulatory proteins. Acting in concert, they determine the pattern of expression of various genes (see Section 10.2.8).

The Control of Gene Expression is discussed in Chapter 10. Here we will consider only how a protein binds to DNA. All site-specific DNA-binding proteins described so far recognize their sequence from the outside of the molecule and associate with DNA without disrupting base-pairing. This is possible because parts of each nucleotide pair are located in two separate grooves—the Major and minor grooves (Fig. 9-8, A). In the major groove, each of the four possible pairs (A-T, T-A, G-C, or C-G) is uniquely recognized by the specific arrangement of protruding atoms, whereas the minor groove is less informative in this regard. As will be shown below, amino acids in the binding site of a site-specific protein are positioned to enhance electrostatic and Hydrogen Bonds formed during the interaction between the protein and a specific DNA sequence. As expected, most hydrogen bonds with nucleotide pairs appear to reside in the major groove. A schematic of the interaction between an amino acid side chain and a single base pair in this groove is presented in Fig. 9-8, B.

Fig. 9-8. Recognition of specific Base Pairs within a DNA molecule by DNA-binding proteins. A. The DNA double helix (B-form): major and minor grooves are highlighted in color. The "edges" of each base pair protrude into these grooves, enabling DNA-binding proteins to recognize different nucleotide sequences by forming hydrogen bonds with them. The putative interaction between an Amino Acid and an A-T pair is schematically depicted in B (viewed along the helix axis). The arrangement of hydrogen bonds of donor and acceptor groups in this groove differs for each of the four possible combinations of nucleotide pairs. It should be noted that the B-form of the DNA helix is right-handed (see Fig. 3-4). Site-specific DNA-binding proteins can recognize sequences ranging from 4 to 50 nucleotide pairs in length.

10-3

10-12

9.1.7. Gel mobility shift assays allow the detection of site-specific DNA-binding proteins in cell extracts [6]

The DNA molecule carries a high negative charge and therefore rapidly moves toward the positive electrode in an electric field. During Polyacrylamide gel electrophoresis, DNA molecules are separated by size because smaller molecules pass through the gel pores more easily and quickly. Protein molecules, upon binding to DNA, cause a decrease in the mobility of its molecules in the gel. The larger the bound protein, the slower the protein-DNA complex moves. This phenomenon forms The basis of the gel mobility shift assay. This technique makes it possible to detect even trace amounts of a site-specific DNA-binding protein. Short DNA fragments of known length and sequence (obtained either by DNA Cloning or chemical synthesis) are radiolabeled and mixed with a cell extract; the resulting mixture is loaded onto a polyacrylamide gel and subjected to electrophoresis. If the DNA fragment corresponds to a chromosomal region to which many site-specific proteins bind, autoradiography reveals a series of bands with differing mobilities. The proteins bound to DNA in each gel band can be isolated by subsequent fractionation of The Cell extracts (Fig. 9-9).

Fig. 9-9. Gel retardation. The Principle of the method is schematically illustrated in Fig. 9-9, A. An extract from an antibody-producing cell line is mixed with radioactive DNA fragments containing the sequence from position —131 to +36 (relative to the Transcription initiation site at position +1) of the gene encoding the light chain of the corresponding antibody molecule. The Effect of proteins present in the extract on the mobility of the DNA fragment is determined by polyacrylamide gel electrophoresis followed by autoradiography. Free DNA fragments move rapidly to the bottom of the gel, whereas protein-bound fragments are retarded. The detection of six retarded zones indicates the presence of six different site-specific DNA-binding proteins in the extract (designated C1-C6). In scheme B, extracts were fractionated using standard chromatographic techniques, and each fraction was mixed with a radioactive DNA fragment, loaded onto a single lane of a polyacrylamide gel, and further analyzed as indicated in scheme A. (B, modified from C. Scheidereit, A. Heguy, R.G. Roeder, Cell 51: 783-793, 1987.)

9.1.8. Site-specific DNA-binding Proteins can be isolated and characterized using their affinity for DNA [7]

Site-specific DNA-binding proteins were first discovered in bacteria. Genetic analysis in these microorganisms proved the existence of regulatory proteins such as the lac Operon repressor, bacteriophage lambda repressor, and cro protein. These proteins were isolated by sequential Fractionation of Cell extracts on chromatographic columns, and their specific DNA-binding sites were mapped using footprinting (see Section 4.6.6). Structure/131.html">Similar Methods were used to isolate and characterize the first eukaryotic site-specific DNA-binding proteins: the SV40 virus T-antigen, transcription factor TFIIIA, and the steroid hormone receptor.

Currently, much more sophisticated methods have been developed for isolating site-specific DNA-binding proteins. The procedure typically begins with a gel mobility shift assay. This makes it possible to determine which specific region of a DNA fragment an unknown protein in a cell extract binds to (see Fig. 9-9). Then, a double-stranded oligonucleotide corresponding to this binding site is chemically synthesized and can be used in two ways. In Affinity Chromatography, the oligonucleotide is coupled to an insoluble porous support, such as agarose, and this loaded support is packed into a Column that selectively binds proteins recognizing specific DNA sequences. This relatively straightforward method allows for a 10,000-fold purification.

Most proteins that bind to a specific DNA sequence are present in higher eukaryotic cells at a level of a few thousand copies per cell (corresponding to roughly one molecule out of every 50,000 cellular protein molecules). This amount is sufficient to isolate the protein by affinity chromatography to a degree of purity that permits determination of its amino-terminal Amino Acid Sequence. This makes it possible to synthesize an oligonucleotide probe and use it to identify the corresponding cDNA clone. Armed with the desired cDNA clone, a researcher can determine the complete amino acid sequence of the protein and produce it in unlimited quantities. In some cases, a cDNA clone encoding a site-specific DNA-binding protein is more easily obtained through a simpler approach utilizing a second method, which is even more efficient than DNA affinity chromatography. This method begins with the Construction of a cDNA library in an appropriately chosen vector. An individual bacterial colony (if the expression vector is a plasmid) or a plaque (if the vector is a bacteriophage) will produce a large amount of the protein encoded by the cDNA it contains. To find the rare colony producing the desired protein, an oligonucleotide containing the appropriate binding site is radiolabeled and used as a probe on nitrocellulose filters bearing replicas of individual colonies (see Section 5.6.5). The few colonies that synthesize proteins specifically binding the labeled oligonucleotide are grown separately and tested further to identify the single one producing the target protein.

This highly efficient method was developed relatively recently, and therefore, to date, out of the many hundreds of site-specific DNA-binding proteins believed to exist in higher eukaryotic cells, only a small fraction has been successfully isolated.

9.1.9. Many site-specific DNA-binding proteins share common domains [8]

The Eukaryotic Transcription factor III A (TFIIIA) is required to initiate the synthesis of small ribosomal RNA (5S rRNA); it binds as a monomer to a DNA sequence of about 50 base pairs located almost in the middle of the very small 5S rRNA gene. The amino acid sequence of this protein suggests that it is organized into a series of nine repeating domains, each containing 30 amino acids folded into a single structural unit around a Zn atom that coordinates two Cysteine and two Histidine residues. Other putative gene-regulatory proteins contain fewer domains with a similar structure. Such domains are commonly referred to as "zinc fingers." Proteins containing them function during early Drosophila development, more than five such proteins are involved in yeast mating, and this same class includes the common mammalian transcription factor Sp1 (see Section 10.2.8) and a large group of steroid hormone receptor proteins (see Fig. 10-25).

The three-dimensional structure of this type of protein has not yet been determined, and consequently, it remains unknown how they bind to DNA. Figure 9-10 presents a hypothetical scheme supported by footprinting data.

9.1.10. Symmetrical dimers of DNA-binding proteins often recognize symmetrical nucleotide sequences [8]

Determining the three-dimensional structure of a protein usually requires X-Ray Diffraction Analysis of large crystals, obtaining which is often a challenging task. One of the first regulatory proteins studied by this method is the bacteriophage lambda cro protein. This is a small protein (66 amino acid residues) that lacks zinc fingers; nevertheless, it binds to a cluster of specific DNA sequences, each 17 base pairs long. One of these sequences is shown in Fig. 9-11. The highlighted part of the sequence is symmetrical, meaning it remains unchanged when the DNA helix is rotated by 180°. Many binding sites of site-specific proteins also exhibit Symmetry. This symmetry can be explained based on The structure of the cro protein determined by X-ray diffraction.

Fig. 9-10. The zinc-finger family of site-specific DNA-binding proteins. A. Simplified diagram of the structure of a DNA-binding domain; circles indicate individual amino acids. It is currently believed that the polypeptide chain forming each "finger" has a complex globular structure. B. Diagram of the interaction of four zinc fingers with a DNA sequence. According to this model, each finger recognizes a specific sequence of approximately five base pairs. (Modified from A. Klug, D. Rhodes, Trends in Biochem. Sci. 12: 464-469, 1987.)

Fig. 9-11. Specific DNA sequence recognized by the bacteriophage lambda cro protein. The highlighted nucleotides in this sequence are symmetrically arranged, allowing each half of the dimeric protein to recognize each half of the site.

The cro protein is a symmetrical homodimer that binds to DNA in the manner depicted in Fig. 9-12. Because the twofold axis of symmetry of the protein sequence coincides with the twofold axis of symmetry of the DNA sequence, each of the identical halves of the dimer can form identical bonds with the DNA base pairs it recognizes. Whenever The sequence of a DNA binding site is symmetrical, the protein that recognizes it is likely to be a dimer or a larger symmetrical structure.

9.1.11. The cro protein belongs to the helix-turn-helix family of DNA-binding proteins [9]

The DNA-contacting site of the protein monomer is formed by a sequence of 20 amino acids that make up two $\alpha$-helices separated by a short turn. This helix-turn-helix motif has also been found in A number of other bacterial site-specific DNA-binding proteins whose three-dimensional structures are known (Fig. 9-13). Furthermore, amino acid sequence analysis (revealing significant Homology) indicates that this motif is also present in other proteins involved in gene regulation in bacteria, yeast, and Drosophila.

All the proteins containing the helix-turn-helix structure shown in Fig. 9-13 are symmetrical homodimers. One of the $\alpha$-helices in each monomer, called the recognition helix, lies in the major groove, where protruding amino acid side chains form hydrogen bonds with specific DNA bases. Importantly, the two identical recognition helices in the dimer are separated by precisely one turn of the DNA chain (3.4 nm). This Separation allows each recognition helix to interact in an identical manner with symmetrically positioned base pairs in the binding site (Fig. 9-14).

Fig. 9-12. Binding of the bacteriophage lambda cro protein to DNA. The protein is a symmetrical homodimer that binds to the symmetrical DNA sequence shown in Fig. 9-11. The mode of DNA binding was determined by analyzing molecular models. A. Wire model of the protein molecule combined with a schematic representation of the DNA double helix. B. Space-filling model of the DNA-protein complex shown in A. In this model, each amino acid in the protein chain is depicted as a sphere; colored spheres represent the DNA backbone. (Courtesy of Brian W. Matthews, after W. F. Anderson, D. H. Ohlendorf, Y. Takada, and B. W. Matthews, Nature 290: 754-758, 1981, © 1981, Macmillan Journals Ltd.)

Fig. 9-13. Family of dimeric DNA-binding proteins. These regulatory proteins function in bacterial systems: the lambda repressor and cro protein control bacteriophage lambda gene expression, whereas the catabolite activator protein (CAP) regulates the expression of a set of E. coli genes that can be turned on only in the absence of glucose. In each case, X-ray diffraction analysis revealed the presence of two copies of the recognition helix (brown cylinder) separated by one turn of the DNA helix (3.4 nm).

Fig. 9-14. Contact of amino acid side chains with DNA base pairs during nucleotide sequence recognition by the bacteriophage 434 repressor regulatory protein. The DNA complex with this protein was studied by X-ray diffraction analysis. (Modified from J. E. Anderson, M. Ptashne, and S. C. Harrison, Nature 326: 846-852, 1987.)

The allosteric effector molecule bound by this type of protein can significantly increase or decrease its affinity for DNA by altering the distance between the two recognition helices. Similarly, allosteric effectors can alter the DNA-binding affinity of Other types of gene regulatory proteins. Such changes are crucial for turning genes on and off in response to environmental changes (see Section 10.2.10).

9.1.12. Protein molecules frequently compete or cooperate with each other upon binding to DNA [10]

Most genetic processes depend on interactions between protein molecules that bind simultaneously to adjacent DNA sites. In the simplest case, two site-specific proteins whose binding sites partially or completely overlap compete with each other for a position on the DNA helix (Fig. 9-15, A). For example, a repressor protein can inhibit gene transcription by blocking the binding of an activating protein to DNA. However, Proteins can also assist each other in holding onto DNA more tightly. Such cooperative binding can occur either between two different protein molecules (Fig. 9-15, B) or between two copies of the same type of molecule. In the latter case, the proteins typically bind in an all-or-none fashion and form clusters on the DNA. As the concentration of these proteins increases, their binding to DNA rises sharply (Fig. 9-15, C). Examples of such cooperatively binding proteins include single-strand DNA-binding proteins, the RecA protein (Ch. 5), and histone H1.

The interaction mechanisms of cooperative and competitive binding using two different site-specific DNA-binding proteins as examples will be discussed further in connection with the Regulation of transcription (Ch. 10).

9.1.13. The geometry of the DNA helix depends on The nucleotide sequence [11]

For 20 years following the Discovery of the DNA double helix in 1953, it was assumed that the molecule possessed a uniform structure throughout its entire length, with a helical twist between adjacent base pairs of precisely 36° (10 nucleotide pairs per helical turn). Subsequent experiments revealed that DNA is far more polymorphic, with its conformational variants dictated by the nucleotide sequence. The conformation of the helix exerts a profound influence on its interactions with proteins.

Fig. 9-15. Examples of competitive and cooperative interactions in protein-DNA binding. A. Competition occurs when the Binding of Proteins X and Y to specific DNA sites is mutually exclusive. B. Cooperation takes place when the binding of one protein to DNA increases the affinity of another protein for a nearby site. C. Self-cooperativity leads to the binding of an individual protein as a cluster of molecules, resulting in an "all-or-none" binding to a given region of DNA.

Fig. 9-16. Three forms of the DNA helix, each containing 22 nucleotide pairs. All of these structures are formed by two antiparallel DNA strands held together by complementary base pairing. Each form is shown from the side and from above. The sugar-phosphate backbone and base pairs are highlighted in different shades of gray: dark gray and light gray, respectively. A. B-form DNA, the conformation most frequently encountered in cells. B. A-form DNA, which becomes predominant upon the dehydration of any DNA, regardless of its sequence. C. Z-form DNA: certain sequences adopt this conformation under specific conditions. Both the B-form and A-form are right-handed, whereas the Z-form is left-handed (see Fig. 3-4). (Courtesy of Richard Feldmann.)

Several types of DNA helices exist, and in some cases, structural deviations are quite significant. The most stable form of DNA is the right-handed B-form helix (Fig. 9-16, A). X-ray diffraction studies have demonstrated that short sequence stretches can form a right-handed helix distinct from the B-form, known as A-form DNA. This form is characterized by a greater base tilt, resulting in a helix that is shorter and wider than the B-form (Fig. 9-16, B). This is likely of great functional importance in certain contexts—for instance, when a DNA strand pairs with RNA in the primer region of Okazaki fragments (the extra ribose hydroxyl group prevents RNA-DNA hybrids from adopting the B-form). For the same reasons, RNA-RNA helices also exist in the A-form. The A-form is characteristic of the helical regions in hairpin structures of all single-stranded RNAs and is therefore critically important for cellular viability. The Biological Role of the third DNA helical form is less clear. DNA sequences consisting of alternating Purines and Pyrimidines (GCGCGCGC) readily form a left-handed double helix known as Z-form DNA (Fig. 9-16, C). It is believed that short segments with this structure occur rarely in chromosomes. Nevertheless, one can surmise that they are specifically recognized by proteins and thus may also play a vital role in cell physiology.

Although the A- and Z-forms of DNA—which differ drastically from the stable B-structure—are quite rare, milder distortions of the B-duplex are common and undoubtedly possess major biological significance. In nucleic acids adopting the B-form, the base tilt and the helical twist angle between base pairs depend significantly on the identity of neighboring nucleotides within the sequence. As a result, the atoms of the helix deviate from their ideal positions. Even DNA-binding proteins lacking the ability to specifically recognize particular nucleotide pairs can "sense" such distortions. The Importance of structural variations in the DNA helix for the binding of site-specific proteins is clearly illustrated by the bacteriophage λ repressor. X-ray crystallographic analysis has shown that upon binding to this protein, the DNA closely approaches an ideal B-form helix (Fig. 9-17).

Fig. 9-17. Comparison of the DNA Conformation in the complex with the bacteriophage 434 repressor (right) to an ideal B-form helix (left). DNA phosphates contacting the protein are marked with colored circles. The twist angles between base pairs are indicated for the distorted helix. Deviations from the ideal helix are thought to be essential for tighter protein binding. Binding energy depends in part on "nonspecific" interactions between NH groups and the DNA phosphates. Such binding-promoting contacts can be formed only by a slightly distorted B-form helix: the middle of the minor groove within the binding site must be slightly narrowed, and the helix moderately bent. Because AT pairs must be present to permit this bending and central compaction of the helix, substituting them with GC significantly impairs protein binding. Thus, centrally located base pairs are recognized by the repressor even when the protein does not make direct contact with these bases (see also Fig. 9-14). (Modified from J. E. Anderson, M. Ptashne, and S. C. Harrison, Nature 326: 846-852, 1987.)

Fig. 9-18. Electron micrograph of DNA fragments containing a strongly bent helical segment. DNA fragments isolated from kinetoplast minicircles of the trypanosomatid Crithidia fasciculata contain only about 200 nucleotide pairs, yet some are bent to such an extent that the entire fragment closes into a circle. A normal DNA helix of this length can, on average, bend to form only a quarter of a circle (one uniform right-handed turn). (From J. Griffith, M. Bleyman, C. A. Rauch, P. A. Kitchin, and P. T. Englund, Cell 46: 717-724, 1986.)

9-7

9.1.14. Some DNA sequences are strongly bent [12]

The DNA helix possesses sufficient conformational freedom to accommodate elastic bending and rotational motions. However, because the helix is also relatively rigid, approximately 200 nucleotide pairs are required for the molecule to bend through an angle of 90°. Certain DNA sequences prove to be exceptionally flexible and adopt a curved shape much more readily than others.

Certain DNA sequences have an intrinsic propensity to remain permanently bent. Among these are sequences containing repeats of AAAAANNNNN (where N is any nucleotide) spaced every 10-11 nucleotides (Fig. 9-18). The pronounced bending of such molecules represents an extreme form of helical variation. It remains unclear whether this curvature stems from the cumulative contribution of minor tilts between specific base pairs, a sharp bend at the junction of short regions, or a combination of both mechanisms.

Even in the absence of specific sequences that promote helical bending, Introduction/20.html">DNA Structure can be significantly distorted by the binding of proteins involved in assembling highly condensed protein-DNA complexes.

9.1.15. Proteins can wrap DNA into a tight helix [13]

Many proteins bend DNA upon binding, and when multiple such DNA-Protein Interactions occur, the DNA strand can wrap into a tight helix around the protein core to form a nucleoprotein particle. In bacteria, such nucleoprotein particles are known to form during the binding of initiator proteins to the replication origin (see Section 5.3.9), as well as during the binding of DNA to phage lambda integrase to catalyze Site-Specific Recombination (Fig. 9-19). Evidently, both competitive and cooperative interactions contribute to such complex three-dimensional assemblies. Similar types of interactions are employed in regulating the catalytic activity of nucleoprotein particles, as exemplified by the protein complex containing phage lambda integrase (Fig. 9-20). Because EUKARYOTIC GENE EXPRESSION involves the binding of clusters of gene-regulatory proteins to specific regulatory DNA sequences, analogous Protein Complexes likely participate in controlling DNA Transcription in eukaryotes as well (see Fig. 10-23).

Fig. 9-19. Schematic diagram of the DNA-protein complex formed by phage lambda integrase, the enzyme responsible for inserting bacteriophage DNA into the E. coli host chromosome. This complex catalyzes site-specific recombination by cleaving and rejoining the DNA helices of phage lambda and the bacterium at specific sequences termed attachment sites (see Fig. 5-66). (Modified from E. Richet, P. Abcarian, and H. A. Nash, Cell 52: 9-17, 1988.)

Fig. 9-20. The excision of bacteriophage lambda from the bacterial chromosome is controlled via cooperative and competitive interactions between site-specific DNA-binding proteins. The reaction is catalyzed by phage lambda integrase and is the reverse of the site-specific recombination shown in Fig. 9-19. A. General reaction scheme and some of the protein-binding sites involved (not all sites are shown). Excision requires the breakage and reunion of the DNA double helix at recombination sites 1 and 2, yielding a circular phage lambda chromosome. Int is the phage lambda integrase, Xis is the phage lambda excisionase, and IHF and FIS are proteins produced by the bacterial host cell. B. Activation of excision by the FIS protein; the indicated steps presumably occur at low concentrations of the Int and Xis proteins. As shown, the binding of several proteins sharply bends the DNA. Although the Int protein catalyzes the site-specific recombination reaction itself, its activity is modulated by other proteins. (Courtesy of Arthur Landy.)

Obviously, any complexes involved in regulating specific genes must occur relatively infrequently. The predominant type of nucleoprotein particle present in Eukaryotic Chromosome structure is the nucleosome, which plays a central role in packaging and organizing DNA within the Cell Nucleus.

9.1.16. Histones are the major structural proteins of eukaryotic chromosomes [14]

The best-characterized structural proteins of chromosomes are undoubtedly histones, which are found exclusively in eukaryotic cells. Their intracellular Abundance is so high that in eukaryotes it is customary to divide DNA-binding proteins into two classes: histones and non-histone chromosomal proteins. The complex formed by both classes of proteins with nuclear DNA in eukaryotic cells is known as Chromatin. Histones are present in such staggering quantities (about 60 million molecules of each type per cell, compared to roughly 10,000 per cell for a typical site-specific protein) that their total mass in the chromosome roughly equals that of the DNA.

Histones are relatively small proteins with an exceptionally high content of positively charged amino acids (Lysine and Arginine). Their net positive charge enables them to bind tightly to DNA regardless of its nucleotide composition. Histones are most likely associated with DNA at all times and, consequently, play a crucial role in all processes related to genome function.

Fig. 9-21. Amino acid sequence of histone H4, one of the core histones. Amino acids are designated by single-letter Abbreviations, with positively charged residues highlighted in color. As with other core histones, the extended amino-terminal sequence of the molecule is reversibly modified in the cell by Acetylation of individual lysine residues. The sequence shown corresponds to bovine histone H4. In pea, this histone has an almost identical amino acid sequence, except that one valine residue is replaced by isoleucine, and one lysine residue by arginine.

The five types of histones can be divided into two main groups: (1) core histones and (2) H1 histones. Core Histones are small proteins (102–135 amino acid residues) responsible for nucleosome formation. They include four histones: H2A, H2B, H3, and H4. H3 and H4 form the interior of the nucleosome and are known to be the most evolutionarily conserved proteins identified to date; for instance, the Amino acid sequences of histone H4 in peas and cows differ by only two amino acid residues (Fig. 9-21). Such evolutionary stability implies that nearly every amino acid within these proteins plays a vital role, and a mutation at any position could prove detrimental to the cell.

Histone H1 is larger (approximately 220 amino acid residues) and has proven to be less evolutionarily stable than the core histones. In the yeast Saccharomyces cerevisiae, H1 appears to be entirely absent (see Section 10.3.15).

9-8

9-9

9.1.17. Histone binding to DNA leads to the formation of nucleosomes—particles that serve as the fundamental unit of chromatin [15]

If it were possible to stretch out the DNA strand of each human chromosome, its length would exceed the size of The Nucleus by thousands of times. Histones play a critical role in packing this exceptionally long DNA molecule into a nucleus that is only a few microns in diameter. These proteins are important for another reason as well: DNA can be packaged in various ways, and the specific packaging of a genomic region into chromatin in a given cell can apparently influence The activity of the genes contained within that region (see Section 10.3.8).

In-depth research into Chromatin Structure began with the discovery of its fundamental structural unit, the nucleosome, in 1974. Due to the presence of nucleosomes, partially decondensed chromatin resembles beads on a string in electron micrographs (Fig. 9-22). A nucleosome "bead" can be separated from the long DNA strand by treating the chromatin preparation with DNA-cleaving enzymes. Enzymes that cause the degradation of both DNA and RNA are called Nucleases, whereas enzymes that act exclusively on DNA are termed deoxyribonucleases or DNases. The nuclease used to isolate individual nucleosomes is derived from micrococcal cells (micrococcal nuclease). Brief Treatment with this enzyme cleaves only the DNA regions located between nucleosomes; the remaining DNA is protected by the histones bound to it, causing the entire polymer molecule to break down into double-stranded fragments 146 base pairs in length. In electron micrographs, these DNA-histone complexes appear as disk-shaped particles approximately 11 nm in diameter. Each nucleosome contains a set of eight histone molecules—two molecules of each of the four highly conserved core histones: H2A, H2B, H3, and H4. This histone octamer essentially forms the protein core of the nucleosome around which a segment of double-stranded DNA is wound (Fig. 9-23).

Fig. 9-22. Cytology/cytology/93.html">ELECTRON MICROGRAPHS OF chromatin fibers before and after treatment that leads to the deconsolidation of the native structure and the formation of "beads on a string." (A) Native structure characteristic of basic chromatin fibrils 30 nm in diameter. (B) Decondensed form of the chromatin fiber "beads on a string," shown at the same magnification. A schematic representation of both chromatin forms is given in Fig. 9-38. The electron micrographs were obtained using the modified procedure described in Fig. 9-71. (A, courtesy of Barbara Hamkalo; B, courtesy of Victoria Foe.)

In intact chromatin, DNA extends as a continuous thread from one nucleosome to the next. Each nucleosomal bead is separated from the adjacent one by a linker sequence that varies in length from 0 to 80 nucleotide pairs. On average, nucleosomal particles (the nucleosome core plus the linker sequence) repeat every 200 nucleotides (see Fig. 9-23). Thus, a eukaryotic gene consisting of 10,000 nucleotide pairs is associated with 50 nucleosomes, and each human cell, whose DNA comprises 6 × 109 nucleotide pairs, contains 3 × 107 nucleosomes.

9-10

9.1.18. Some nucleosomes are positioned on DNA in a non-random manner [16]

In vitro experiments with isolated chromatin suggest that under physiological conditions, histone octamers remain fixed in a single position because their tight association with the nucleic acid prevents sliding along the DNA helix. An open question is whether these octamers are positioned randomly on the DNA or not (randomness implies that in one cell a particular DNA sequence is tightly wrapped around a histone octamer, whereas in another cell the same sequence separates adjacent nucleosome beads).

To determine the positions of nucleosomes within cells, one must treat them with an enzyme or reagent that introduces cuts into the DNA and then analyze the protected regions using a method analogous to DNA footprinting (see Section 4.6.6). Although most nucleosomes appear to be positioned randomly, striking examples of non-random positioning are known. For instance, in the yeast Saccharomyces cerevisiae, 15 nucleosomes position themselves in a strictly fixed manner around the centromere DNA (the CEN sequence) (Section 13.5.3). A unique localization site is also occupied by the nucleosome associated with the very small 5S rRNA gene. Furthermore, at least one nucleosome is known to be positioned immediately upstream of the transcription initiation site for ß-globin.

How can this non-random positioning of nucleosomes be explained? It has been demonstrated that in certain cases (e.g., for nucleosomes associated with 5S rRNA genes), a mixture of the four purified core histones reconstitutes nucleosomes in vitro at precisely the same sites they occupy in vivo. The reason for this may be that nucleosomes tend to bind in a way that maximizes the Filling of the AT-rich minor groove of DNA. This preference arises because the DNA double helix is difficult to bend into two tight turns around the histone octamer without significant compression within the minor groove of the DNA helix (Fig. 9-24). As demonstrated with the bacteriophage repressor protein (see Fig. 9-17), a cluster of two or three AT pairs located in the minor groove facilitates this compression. Other, as yet unknown Properties of the DNA sequence must also influence nucleosome positioning.

Fig. 9-23. Structure of nucleosomes. Nucleosomal particles consist of two complete turns of DNA (83 nucleotide pairs per turn) wound around a core of a histone octamer and connected to one another by linker DNA. The nucleosomal particle is isolated from chromatin by limited Hydrolysis of the linker DNA regions using micrococcal nuclease. In each nucleosomal particle, a double-stranded DNA fragment 146 base pairs in length is wound around the histone core. This protein core contains two molecules of each of the histones H2A, H2B, H3, and H4. Histone polypeptide chains range from 102 to 135 amino acid residues in length, and the total mass of the octamer is approximately 100,000 Da. In the decondensed form of chromatin, each "bead" is connected to the neighboring particle by a thread-like stretch of linker DNA.

Fig. 9-24. DNA bending in the nucleosome. The DNA helix makes two full turns around the histone octamer, with 83 nucleotide pairs per turn. The relative proportions of DNA and protein in the diagram are close to scale to illustrate the compression of the minor groove on the inner side of the coil. As noted above (see Fig. 9-17), the narrow minor groove is populated predominantly by AT base pairs.

9.1.19. Specific sites on chromosomes lack nucleosomes [17]

Nucleosomes are absent from certain DNA regions, despite these regions being hundreds of nucleotide pairs long. Such areas can be detected by treating cell nuclei with trace amounts of deoxyribonuclease (DNase I). Using minimal concentrations of the enzyme ensures the degradation of long stretches of nucleosome-free DNA, while the short stretches of linker DNA located between nucleosomes remain intact. Chromatin treated in this manner is cleaved predominantly at regions that evidently lack nucleosomes. Typically, these sites are spaced several thousand nucleotide pairs apart.

The first Evidence for the Biological Significance of nuclease-hypersensitive sites came from experiments with the SV40 virus. In addition to its circular DNA, the viral chromosome contains histones produced by the host cell. This chromosome features a 300-nucleotide-pair region that is devoid of nucleosomes and is rapidly degraded upon exposure to DNase I. This region is located very close to the DNA sequences where both Viral DNA Replication and RNA Synthesis initiate. Several site-specific DNA-binding proteins also localize here, protecting only a small portion of this molecule—which appears to be completely devoid of nucleosomes—from nuclease degradation. Similarly, many DNase-hypersensitive regions of chromatin within the cell are located in the regulatory regions of genes (Fig. 9-25); in cells where these genes are active, such sites are more abundant than in other cells. It is believed that site-specific DNA-binding proteins involved in the regulation of eukaryotic genes are responsible for the displacement of nucleosomes (see Fig. 9-27).

9.1.20. Nucleosomes are typically packed together to form ordered higher-order structures [18]

In living cells, chromatin rarely appears as extended "beads on a string." Instead, nucleosomes associate with one another to form regular structures in which the DNA is even more condensed. Electron microscopic analysis of nuclei lysed directly on the grid reveals that the bulk of chromatin is organized into fibrils approximately 30 nm in diameter (Fig. 9-22A). One possible way of packing nucleosomes within a 30-nm chromatin fibril is illustrated in Fig. 9-26. This model represents an ideal structure. In reality, both the variation in linker lengths (reflecting a specific arrangement of nucleosomes) and the presence of randomly generated nucleosome-free sequences impart distinct properties to different regions of the 30-nm fibril (Fig. 9-27).

Fig. 9-25. Location of nuclease-hypersensitive sites (colored arrows) in the regulatory regions of active genes. Although such chromatin regions are typically located at the 5'-end of a gene—as shown here for the histone gene cluster (H1, H2A, H2B, H3, and H4) in Drosophila—these sites may also be found in other regions (see Fig. 10-40A).

Fig. 9-26. A model proposed to explain the packing of the nucleosomal filament into the 30-nm fibril observed by Electron Microscopy (see Fig. 9-22A). A. Top view. B. Side view. In this type of packing, one molecule of histone H1 is associated with each nucleosome (not shown). Although the attachment site of histone H1 on the nucleosome is defined, the precise arrangement of H1 molecules along this fibril remains unknown (see also Fig. 9-27).

If the chromatin of a typical human chromosome existed entirely as a 30-nm fibril, it would span 0.1 cm when fully stretched, exceeding the nuclear dimensions by 100-fold. Microscopic analysis of intact chromosomes suggests that 30-nm fibrils undergo further folding within cells to form chromatin threads approximately 100 nm thick. The exact arrangement of nucleosomes within such a structure is currently unclear.

9.1.21. Histone H1 proteins help link nucleosomes together [19]

Mammalian cells contain approximately six closely related variants (subtypes) of histone H1 that differ slightly in their amino acid sequences. These molecules are presumably responsible for packing nucleosomes into the 30-nm fibril. H1 molecules possess an evolutionarily conserved globular central domain flanked by protruding amino-terminal and carboxy-terminal "tails" whose amino acid sequences evolve more rapidly. The globular domain of each H1 molecule binds to a unique site on the nucleosome, whereas the tails are thought to span a broader area and contact neighboring sites on histones of adjacent nucleosomes. This action pulls the nucleosomes together into a regular, repeating structure (Fig. 9-28). According to various hypotheses, H1 molecules are located either inside or outside the 30-nm chromatin fibril.

Fig. 9-27. Diagram illustrating the disruption of the regular nucleosome-non-nucleosome chromatin structure by short regions where the DNA is unusually sensitive to DNase I Digestion. At each of these nuclease-hypersensitive sites, the nucleosomes on the DNA are presumably replaced by one or more site-specific DNA-binding proteins.

Fig. 9-28. Diagram illustrating how histone H1 (220 amino acids) might mediate contact between adjacent nucleosomes. The globular domain of H1 binds to each nucleosome near the site where the DNA helix enters and exits the histone octamer. In the presence of histone H1, two full turns of DNA (166 nucleotide pairs) are protected from micrococcal nuclease digestion (see Fig. 9-23). However, neither the three-dimensional structure of histone H1 nor the precise interaction regions of its protruding amino- and carboxy-terminal tails with the nucleosome are yet known.

The histone octamer forming the core of each nucleosome is a symmetrical structure, whereas the single histone H1 molecule bound to each nucleosome lacks symmetry. Consequently, the binding of H1 molecules to chromatin establishes a local polarity (Fig. 9-29).

In vitro experiments demonstrate that when histone H1 binds to DNA, eight or more protein molecules attach simultaneously, providing an example of cooperative binding. It is highly likely that chromatin is organized through cooperative interactions of this type, and that their disruption by regulatory proteins leads to local chromatin decondensation (as occurs in active gene regions). The resulting "active chromatin" domains appear to have an unusually low affinity for histone H1. In this respect, a restricted region of chromatin may behave like a tiny crystal capable of undergoing an all-or-none conformational switch during gene activation processes (Fig. 9-30).

9.1.22. Nucleosomes do not obstruct RNA synthesis [20]

Biochemical evidence indicates that during transcription, the bulk of DNA remains associated with nucleosomes; electron microscopic examination of spread chromatin preparations typically reveals a uniform distribution of nucleosomes in both transcribed and untranscribed regions (Fig. 9-31). The histone octamer appears to be bound so tightly to DNA that it remains permanently attached. Nevertheless, it is difficult to envision how RNA polymerase transcribes nucleosomal DNA without causing temporary alterations in nucleosome structure. As RNA polymerase passes, the DNA wrapped around the nucleosome presumably unwinds from the histone octamer without fully dissociating from it. Alternatively, the entire nucleosome may temporarily open up, splitting the histone octamer into two halves. Another possibility is that the octamer remains intact but transiently "tilts" or shifts to let the polymerase pass (Fig. 9-32). The inherent difficulty of such maneuvers may explain the remarkable evolutionary conservation of histone amino acid sequences. If a nucleosome served merely as a passive spool for winding DNA, one would expect greater sequence Variability among core histones.

Fig. 9-29. Chromatin polarity is conferred by its binding to histone H1. A. In the absence of histone H1, chromatin lacks polarity because each nucleosome is symmetrical. B. In the presence of H1, chromatin becomes polarized. The precise biological role of this polarity remains unknown.

Fig. 9-30. Individual chromatin regions can function as structural units due to the cooperative nature of nucleosome packing mediated by histone H1. This diagram illustrates the abrupt decondensation of such a unit (a 30-nm fibril or more compact structure) triggered by an external regulatory signal. This type of chromatin decondensation may accompany gene activation (see Fig. 9-50).

As discussed above, the majority of cellular nucleosomes are packed into a 30-nm chromatin fibril, which then undergoes further Condensation. It is hard to imagine that chromatin in this state can be transcribed by RNA polymerase without significant alterations in its nucleosomal packaging (Fig. 9-33). Certain experiments suggest that structural changes of this sort do indeed occur (see Fig. 9-50).

Fig. 9-31. A region of chromatin in the form of a nucleosomal filament. Three chromatin threads are shown, with two RNA polymerase molecules actively transcribing DNA on one of them. Most chromatin in the nucleus of higher eukaryotes lacks active genes and is therefore devoid of RNA transcripts. Notably, nucleosomes are present in both transcribed and untranscribed regions and remain associated with the DNA immediately ahead of and directly behind moving RNA polymerase molecules. (Courtesy of Victoria Foe.)

Fig. 9-32. Two plausible models explaining how RNA polymerase can transcribe chromatin without displacing nucleosomes from it. A. Transcription through temporarily unlatched polynucleosomes. B. Transcription through an intact histone octamer.

Fig. 9-33. Diagram showing an RNA polymerase molecule approaching a 30-nm chromatin fiber. The number of depicted nucleosomes corresponds to approximately 7,000 base pairs of DNA, which is comparable to an average-sized human gene (Table 9-1). Somehow, all of this DNA must become accessible to RNA polymerase without slipping off the chromatin histone octamers.

Obviously, this requires chromatin unwinding.

Conclusion

A gene is a nucleotide sequence that serves as a functional unit for generating an RNA molecule. A chromosome consists of a single, incredibly long DNA molecule containing numerous genes. Chromosomal DNA also contains other types of nucleotide sequences necessary for its function: a replication origin and a telomere (which ensure DNA replication), as well as a centromere (which serves to attach DNA to the mitotic spindle). The human haploid genome contains 3 × 109 nucleotide pairs distributed among 22 distinct autosomes and 2 sex chromosomes. Apparently, only a few percent of this DNA encodes proteins.

Eukaryotic DNA is intimately associated with an abundance of histones, which serve to form numerous repeating protein-DNA particles known as nucleosomes. Nucleosomes are typically packed together into a regular structure—a fiber with a diameter of 30 nm. However, in regions of DNA containing genes, the histone octamer forming each nucleosome must compete with a variety of site-specific proteins for DNA binding sites. Such regions, where the nucleosome is replaced by DNA-binding proteins, are generally identified as domains of more active DNA transcription.

A Eukaryotic Cell contains hundreds of diverse site-specific DNA-binding proteins. Each of these recognizes a short DNA sequence through hydrogen bonds with base pairs and via the shape of the helix. When forming protein complexes at specific DNA sites, these proteins engage in cooperative or competitive interactions with one another.



Last update: 12/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.