Molecular Biology of the Cell - Volume 2 - Alberts B., Bray D., Lewis J., Raff M., Roberts K., Watson J. 1993

Control of Gene Expression
Molecular and genetic mechanisms involved in the formation of different cell types

At the beginning of this chapter, we examined various strategies of Gene Expression regulation that enable a multicellular Organism to generate a vast array of distinct Cell types. Control genes exhibit a certain 'economy,' which manifests in the way regulatory Proteins assemble into diverse combinations. This likely explains why the regulatory sequences of most genes in higher eukaryotes associate with numerous proteins. The existence of discrete cell types implies that The activity of regulatory proteins within this combinatorial network is governed by a complex hierarchy of feedback loops, presumably controlled by a small number of master regulatory proteins. However, the underlying logic of this complex gene-regulatory network remains poorly understood. For instance, it is still unclear how the transient ectopic Synthesis of the master regulatory protein Antennapedia in Drosophila transforms an antenna into a leg.

While this section will not provide a definitive answer to that specific question, we will explore how the gene-regulatory network executes switches that are faithfully maintained through subsequent cell generations. Cellular memory is a fundamental prerequisite for the establishment and stability of differentiated Cells.

We will begin our Structure/133.html">Discussion of this problem by examining well-characterized differentiation mechanisms in Bacteria and Yeasts, although some of these mechanisms are known to be absent in higher eukaryotes.

10-16

10.3.1. In Many Bacteria, Specific DNA Rearrangements Result in Gene Activation and Inactivation [20]

As noted above, the differentiation of higher Eukaryotic cells typically occurs without noticeable Changes in DNA sequence (see Section 10.1.2). In many prokaryotes, however, stably inherited patterns of gene regulation are achieved through specialized DNA rearrangements that turn specific genes on or off. Because all Changes in the DNA sequence are faithfully copied during subsequent rounds of Replication, an altered gene activity state is inherited by all descendants of the rearranged cell. Some of these rearrangements are inherently reversible, allowing a gene to switch to a different activity state after a sufficiently long period of time.

Class="center">

Figure 10-28. Alternative METABOLISM/31.html">Transcription of two genes in Salmonella bacteria resulting from Site-Specific Recombination that leads to the inversion of a short DNA segment. The recombination mechanism is described above (see Section 5.4.7) and is activated only rarely (approximately once every 105 cell divisions). Consequently, the alternative production of flagellin H12 (A) or flagellin H1 (B) is generally strictly inherited within each cell clone.

A classic example of this type of cellular differentiation is the genetic rearrangement found in Salmonella bacteria. In this case, a specific DNA segment 1,000 nucleotide pairs long is inverted through a reaction catalyzed by a site-specific recombinase (Figure 10-28). Thus, the DNA at this site can exist in two states. The Effect of this inversion on gene expression is due to the presence, within the 1,000-nucleotide segment, of a promoter responsible for synthesizing a specific surface protein known as flagellin. When this promoter is in one orientation, one type of protein is produced; when it is in the opposite orientation, a different protein is synthesized. Because such switching occurs very rarely, cell clones grow with either one type of flagellin or the other. This phenomenon is known as phase variation. Most likely, this differentiation mechanism helps bacterial populations evade the host Immune Response. If the host produces Antibodies against one type of flagellin, certain bacteria whose flagellin has been altered via gene rearrangement can survive and multiply.

Phase variation affecting one or more phenotypic traits is frequently observed in wild-type bacterial populations. Such 'instability' is typically lost in standard laboratory strains, which is why relatively few of these mechanisms have been studied in detail. Not all of them involve DNA inversion. For example, Neisseria gonorrhoeae, the bacterium that causes Gonorrhea in humans, evades the immune response through heritable variations in cell Surface Properties generated by Gene Conversion (see Section 5.4.6). This mechanism depends on the RecA recombination protein and relies on The transfer of a specific sequence from a silent 'cassette' into an active expression gene (see Section 10.3.2). This process can generate over 100 Variants of the major bacterial surface protein.

10-17

10-18

10.3.2. A Master Regulatory Locus Determines Mating Type in Yeast [21]

Yeast are Unicellular Eukaryotes that can exist in either haploid or diploid states. Diploid cells are formed by mating, in which two haploid cells fuse (see Fig. 13-17). For this to occur, the haploid cells must differ in their mating type (sex). In ordinary baker's yeast, Saccharomyces cerevisiae, There are two mating types, a and α. Cells of the two types are adapted to mate with each other: each secretes diffusing signaling and receptor molecules that allow cells of the opposite mating type to recognize and fuse with them. The resulting diploid cells, designated a/α, possess distinct properties and differ from either parental type. Diploid cells are incapable of mating, but they can form spores that undergo Meiosis to give rise to haploid cells.

The genetic changes responsible for the existence of Three types of yeast cells are based on the action of master regulatory genes. The master regulatory proteins are encoded by a single locus called the mating type locus (MAT). The combinatorial action of these proteins determines The Cell type by regulating the transcription of many genes. The MECHANISM OF ACTION of these master regulatory proteins is well understood (Fig. 10-29) and provides a clear illustration of the combinatorial control principle described above (see Section 10.1.5).

The production of the three master regulatory proteins that determine yeast mating type (the a1, a2, and α1 proteins) is itself tightly regulated, as haploid yeast cells regularly switch their mating type. The molecular mechanism responsible for this developmental switching appears to be unique to yeast: the cell type depends on which of two DNA sequences—either α (encoding the α1 and α2 proteins) or a (encoding the a1 protein)—currently occupies the MAT locus; A change in mating type occurs As a result of replacing the DNA at this locus. Whenever a-type cells convert to α-type cells, the a gene at the MAT locus is excised and replaced by a newly synthesized α gene copied from a silent cassette located elsewhere in The Genome. Because one gene is removed from an Active Site and replaced by another, this mechanism is called the cassette mechanism. The change is reversible because when the a gene is removed from the MAT locus, a silent copy of it remains in the genome. The copies of the silent a and α genes act as cassettes "inserted" into the MAT locus, which can be likened to a "tape player" (Fig. 10-30).

Fig. 10-29. Yeast cell type is determined by regulatory proteins encoded by the MAT locus. Different sets of Genes are transcribed in haploid a cells, haploid α cells, and diploid (a/α) cells. Haploid cells express either the aSG set of genes (a-specific genes) or the αSG set (α-specific genes) plus the hSG set (haploid-specific genes). In diploid cells, none of these sets are expressed. The a1, α2, and α1 proteins encoded by the MAT locus bind (individually or in combinations) to specific DNA sequences in elements located upstream of the promoter, thereby acting as master regulatory proteins. Note that the α1 protein acts as an activator, whereas the α2 protein acts as a repressor. The a1 protein alone has no effect (and consequently, a cell lacking the mating type locus belongs to the a type). However, when the a1 and α2 proteins are synthesized simultaneously, they form a complex that regulates a different set of genes than the α2 protein acting alone. Thus, this simple three-protein system serves as an excellent illustration of THE PRINCIPLE OF combinatorial gene control presented in Fig. 10-7.

Fig. 10-30. The cassette model of mating-type switching in yeast. Cassette switching occurs via a gene conversion process triggered when the HO endonuclease makes a double-strand break at a specific DNA sequence within the MAT locus. The DNA surrounding the break is then degraded and replaced by a copy of a silent cassette specifying the opposite mating type.

10.3.3. The ability to Switch Mating Type Is Inherited Asymmetrically [22, 23]

Mating-type switching is initiated by a site-specific endonuclease (HO endonuclease), which is the product of the HO gene. This enzyme makes a double-stranded break in the DNA of the MAT locus, resulting in the excision and subsequent resynthesis of this region using the silent gene of the opposite mating type as a template (Fig. 10-30). The Transcription of the HO gene, which determines when and where switching occurs, is tightly controlled. Genetic analysis has shown that this control is mediated by at least six regulatory genes (SWI1 through SWI6). Because yeast cells divide asymmetrically during budding, one of the resulting cells is larger (the "mother" cell) than the other (the "daughter" cell). Most mother cells switch their mating type upon further growth, whereas newly formed daughter cells (arising from the bud) do not synthesize the HO gene product and are unable to switch until they themselves become mother cells through subsequent division (Fig. 10-31). The Asymmetry of switching has been linked to the asymmetric Inheritance of the SWI5 protein, which binds to the DNA upstream of the HO gene and is required for its transcription. It is believed that the SWI5 protein (or its active form) is inherited exclusively by the mother cell. While it remains unclear why this protein is absent from the bud, its pattern of inheritance serves as a model for the asymmetric segregation of certain traits observed in higher eukaryotes.

Fig. 10-31. Mating-type switching in haploid S. cerevisiae cells. A. Gene expression during the yeast Cell Cycle. According to this scheme, mating type can switch only in cells that have inherited the SWI5 site-specific DNA-binding protein required for HO gene transcription. Furthermore, switching can occur only during the G1 phase—this is the only window in the cell cycle when the HO endonuclease can be synthesized (following synthesis, the protein rapidly disappears, likely due to intensive proteolysis). Because DNA replication occurs after every switch, a dividing cell always gives rise to two cells of the same mating type. B. Scheme of switching. Cell Division proceeds via budding. After each division, a larger "mother" cell and a smaller "daughter" cell can be distinguished. Newly formed daughter cells are incapable of switching their mating type (indicated by different colors in this figure) until they pass through mitosis as mother cells and enter the next G1 phase. This asymmetry presumably reflects the asymmetric inheritance of the SWI5 protein (designated here as S).

10.3.4. A Silencer Likely "Closes" a Chromatin Domain in Yeast [23, 24]

Had the coding sequences of the two silent mating-type genes lacked promoters, their activation mechanism upon moving to the MAT locus would have likely resembled phase variation. However, mapping the primary Introduction/20.html">DNA Structure revealed that silent genes possess all the regulatory sites necessary for transcription. These genes remain silenced thanks to a DNA sequence located some distance downstream that somehow blocks their expression. While the exact mechanism of such silencers remains unknown, researchers have successfully identified several proteins that mediate their function. This became possible because mutants defective in any of the four SIR (silent information regulator) genes still undergo expression of the silent cassettes. The action of SIR proteins requires a silencer DNA sequence, and all genes located within several kilobase pairs of the silencer are repressed. Normally, this same mechanism prevents the HO endonuclease from cleaving DNA in the region of the silent mating-type genes, ensuring that transfer proceeds exclusively from the silent locus to the MAT locus and never in the reverse direction.

The fact that SIR proteins suppress both transcription and the action of the HO endonuclease indicates that these proteins can alter yeast Chromatin Structure, promoting the "closing" of entire neighboring chromatin domains and rendering them inaccessible to A wide variety of Enzymes. Two other observations suggest that The Mechanism of silencer action is rather unusual. DNA replication is required for repression to occur, and the sequence essential for replication initiation (ARS) is an integral part of the silencer region. A detailed investigation into this novel mechanism of Genetic control may shed light on how chromatin structure influences gene activity in higher eukaryotic cells.

10.3.5. Two bacteriophage proteins that mutually repress each other's synthesis can participate in stable molecular switching [25]

Up to this point, we have discussed certain changes in cell type that occur via a switching mechanism, i.e., through DNA rearrangements. The fact that the Genetic information contained within a single somatic Cell Nucleus can give rise to an entire plant or vertebrate demonstrates that irreversible changes in DNA sequence are unlikely to be the primary mechanism of Cell Differentiation in higher eukaryotes (although such changes do underlie lymphocyte differentiation). Some heritable changes in gene expression observed in higher organisms might rely on a switching mechanism similar to those described in Salmonella and yeast, but there is currently no data to support this hypothesis.

As will be shown below, several mechanisms can account for the heritable regulation of genetic activity. The most prominent among them is the switching mechanism characteristic of bacteriophage lambda. It determines whether phage particles will replicate in the Cytoplasm of E. coli (leading to cell death) or whether the phage genome will integrate into the host cell DNA and replicate automatically alongside it. This switch involves a protein encoded by the bacteriophage genome, which comprises about 50 genes. These genes are transcribed in completely different ways across two stable states. For instance, a virus destined to integrate must synthesize the integrase protein required to insert the phage DNA into the bacterial chromosome while simultaneously suppressing the production of viral proteins responsible for replication (which are lethal to the host cell). Once established, either transcription pattern is stably maintained. As a result, the integrated prophage may remain dormant in the E. coli genome for thousands of cell generations.

Without delving into the intricacies of this complex regulatory system, let us outline its General Properties. At the core of the entire system are two phage proteins: the repressor protein (the cI protein) and the cro protein. Each of these blocks the synthesis of the other by binding to its gene's operator. The presence of one protein or the other, in turn, activates a set of other genes, ultimately leading to the establishment of one of two stable states. In state 1 (the lysogenic state), the lambda repressor dominates, and it is synthesized rather than the cro protein. In state 2 (the lytic state), the cro protein dominates and is synthesized rather than the lambda repressor (Fig. 10-32). In state 1, the majority of the bacteriophage DNA (the prophage) stably integrated into the host cell genome is not transcribed. In state 2, the phage DNA is intensively transcribed, replicated, packaged into new particles, and released upon cell lysis.

The sequence of events unfolding in a bacterial cell following infection by lambda phage is far from random. If the host strain cells are growing healthily, the bacteriophage is likely to follow the lysogenic pathway, allowing its DNA to replicate rapidly along with the host chromosome. Should a cell carrying the prophage become weakened for any reason, the phage transitions from state 1 to state 2, replicating in the Cell Cytoplasm and exiting it rapidly. Information regarding the host cell's status is "delivered" to the phage by other proteins that influence the switching between the repressor protein and the cro protein. Bacteriophage lambda clearly demonstrates how a relatively complex behavioral pattern can be governed by just a few regulatory proteins that mutually control each other's synthesis and activity. Considering the sheer number of proteins contained within a Eukaryotic Cell, the potential avenues for gene regulation are staggering.

Fig. 10-32. The regulatory system governing The behavior of bacteriophage lambda in E. coli host cells. In stable state 1 (lysogenic state), large quantities of the lambda phage repressor protein are synthesized. This regulatory protein turns off the synthesis of several bacteriophage proteins, including the cro protein. As a result, the phage DNA integrates into the E. coli chromosome and is automatically duplicated along with it as the bacterium grows. In stable state 2 (lytic state), large amounts of the cro protein are synthesized. This regulatory protein turns off the synthesis of the lambda phage repressor protein. Consequently, numerous phage proteins are produced, the viral DNA replicates freely within the E. coli cells, and new bacteriophage particles are formed, leading to cell death. Cooperative and competitive interactions between the lambda repressor and the cro protein facilitate an all-or-none switch between these two states.

Fig. 10-33. Two major regulatory proteins promote molecular switching in eukaryotes. Notably, the exact same protein can exert either an activating or a repressing effect on transcription depending on the specific DNA sequence it binds to. Examples of this type of action are believed to occur among DNA-binding proteins that control developmental pathways during early embryonic stages in Drosophila.

10.3.6. Regulatory proteins in eukaryotes can also determine alternative stable states

Genetic analysis of Drosophila development indicates that the fly's body plan is influenced by more than 30 proteins (see Section 16.5). Many of these are master regulatory proteins that bind to enhancers and control the transcription of a wide variety of genes. Some master regulatory proteins stimulate their own transcription while simultaneously repressing analogous master genes expressed elsewhere in the embryo. The paramount role of enhancers in controlling Eukaryotic Transcription allows numerous positive and negative signals to act at a distance on the same gene. This has fostered an extensive network of interacting regulatory proteins (see Fig. 10-73). However, to illustrate the General Principles underlying this type of cellular memory, we can limit ourselves to a simple two-gene system analogous to the lambda phage switch (Fig. 10-33).

The complexity of the regulatory network in higher eukaryotes can be appreciated by examining the DNA sequence of the regulatory region of the gene that encodes the master regulatory protein controlling The Development of many major body parts in Drosophila. By comparing the very long regulatory regions of this gene in two distantly related Drosophila species, researchers identified 20 short, evolutionarily conserved sequences arranged in a specific pattern. Each of these is thought to play a vital role in binding regulatory proteins (Fig. 10-34).

Fig. 10-34. Comparison of a portion of the regulatory region upstream of the engrailed gene in two Drosophila species. The corresponding sequences from Drosophila melanogaster and Drosophila virilis are shown, with regions of up to 90% Homology highlighted in color. Nucleotide positions are indicated at the top of the figure. Loops correspond to sites where nucleotide insertions or deletions occurred during the divergence of these species from a common ancestor approximately 60 million years ago. (Courtesy of Judith A. Kassis and Patrick H. O'Farrell.)

10-19

10.3.7. Cooperatively binding clusters of regulatory Proteins can be transmitted directly from parents to progeny [26]

There are several explanations for stably heritable patterns of gene expression. One is based on the idea that multiple copies of a regulatory protein bind cooperatively to a specific region of chromatin. If this protein cluster remains attached to the DNA during replication, each daughter DNA molecule inherits a portion of it. Because the binding of this protein to DNA is cooperative, the inherited portion of the protein cluster will stimulate the recruitment of additional protein subunits, ultimately ensuring the reconstruction of the entire cluster. Thus, a given functional state of a gene is inherited directly and immediately via tightly bound chromosomal proteins (Fig. 10-35). In principle, such directly inherited protein clusters can maintain individual genes in a permanently active or permanently inactive state.

While conclusive evidence for this mechanism of Genetic regulation is not yet available, several examples suggest that it may be of great significance.

10.3.8. In higher eukaryotic cells, heterochromatin contains specially condensed DNA regions [27]

Regulatory proteins that bind to specific DNA sequences in eukaryotic cells must interact not simply with bare DNA, as bacteria do, but with DNA that is continuously associated with nucleosomes. The necessity of transcribing DNA packaged within chromatin undoubtedly complicates transcriptional control, yet very little is known about how the underlying mechanisms operate. The only certainty is that in eukaryotes, variations in DNA packaging influence gene expression. As noted earlier, the silencer that regulates transcription in yeast somehow "closes off" adjacent chromatin domains, rendering them inaccessible to transcription and endonuclease Cleavage (see Section 10.3.4). However, long before this phenomenon was discovered, studies of higher eukaryotes revealed the existence of much more densely compacted chromatin accompanied by visible structural changes.

Fig. 10-35. A mechanism explaining the direct inheritance of a gene's expression state during DNA replication. According to this hypothetical model, parts of a cooperatively bound regulatory protein cluster are directly transferred from the parental DNA helix to both daughter molecules. The inherited protein cluster promotes the binding of additional copies of the same regulatory proteins to each daughter DNA helix. Because binding is cooperative (see Fig. 9-15), DNA synthesized on a parental strand that lacks attached regulatory proteins will not bind them. If the attached regulatory protein turns off gene transcription, the inactivated state of the gene will be directly inherited, much like X-chromosome inactivation. If the bound regulatory protein turns on gene transcription, the active state of the gene will be inherited.

Chromosome studies using light Microscopy in the 1930s revealed that certain regions fail to decondense during interphase, apparently maintaining the highly condensed state characteristic of metaphase Chromosomes (see Section 9.2.2). These regions were termed heterochromatin, in contrast to the rest of the chromatin in the interphase nucleus, which is called euchromatin. Certain chromosomal segments condense into heterochromatin in all cells of an organism. In human mitotic chromosomes, this constitutive heterochromatin localizes around the centromeres and is easily identified upon specialized staining as darkly colored zones (Fig. 10-36). In some other mammals, constitutive heterochromatin is also found in specific zones along the chromosome arms. During interphase, blocks of constitutive heterochromatin can aggregate to form chromocenters (see Fig. 9.43). In mammals, the number and distribution of such chromocenters depend on the cell type and developmental stage. Most regions of constitutive heterochromatin contain relatively simple tandemly repeated sequences known as satellite DNA. These highly repetitive sequences are not transcribed, and both their function and that of the condensed structures they form during interphase remain elusive.

Certain regions of DNA during interphase condense to form heterochromatin only in specific cells. It is believed that these regions are also transcriptionally inactive, yet they do not consist of simple sequences. The overall amount of such facultative heterochromatin varies noticeably among different cell types: it is minimal in embryonic cells, whereas highly specialized cells contain high levels of it. One can hypothesize that as cells develop, an increasing number of genes become inactive because their DNA adopts a condensed conformation that prevents genes from interacting with activator proteins. Much of our understanding of facultative heterochromatin comes from studying the inactivation of one of the two X chromosomes in female mammalian cells.

10.3.9. Hereditarily Determined X-Chromosome Inactivation [28]

All somatic cells of female mammals possess two X chromosomes, whereas male cells contain one X and one Y chromosome. It is believed that a double dose of gene products encoded on the X chromosome is lethal to the organism; this may explain why a special mechanism evolved in female cells to ensure that one of the two X chromosomes remains permanently inactivated. In mice, such inactivation occurs between the third and sixth days of embryonic development: in each female cell, either one or the other X chromosome condenses with equal probability to form heterochromatin. Such condensed chromosomes can be visualized under a Light Microscope: during interphase, they appear as distinct structural entities known as Barr bodies, which are localized near the nuclear membrane. They replicate in the late S phase, and most of their constituent DNA is not transcribed in either of the daughter cells. Because the inactivated X chromosome is stably inherited, every female organism has a mosaic structure in the sense that it is composed of clonal cell populations: in roughly half of these groups, the maternally inherited X chromosome (Xm) is active, whereas in the other half, the paternally inherited X chromosome (Xp) is active. In other words, cells expressing Xm or, conversely, Xp are arranged in small clusters in the adult organism, reflecting the tendency of sister cells to remain closely associated during embryonic development and growth, while some degree of mixing occurs at earlier developmental stages (Fig. 10-37).

Fig. 10-36. Human chromosomes in metaphase. Special staining techniques reveal regions of constitutive heterochromatin (dark areas). Certain chromosomes are identified by numbers and letters. (Courtesy of James German.)

Fig. 10-37. Schematic illustration of the clonal inheritance of the condensed, inactive X chromosome in female mammals.

The Condensation process that forms X-chromosome heterochromatin tends to spread along the chromosome.

This was demonstrated in experiments using mutant individuals in which one of the X chromosomes was translocated to the end of an autosome (a non-sex, somatic chromosome). In such mutant cells, autosomal regions bordering the inactivated X chromosome frequently condensed into heterochromatin, accompanied by the heritable inactivation of the genes they contained. These findings suggest that X-chromosome inactivation is a cooperative process that can be viewed as a "crystallization" spreading outward from a nucleation center located on the X chromosome. Once chromatin condensation is complete, it is inherited through all subsequent DNA replication cycles via a mechanism analogous to the one shown in Fig. 10-35. The condensed chromosome can become active again during The formation of Germ Cells. Thus, no permanent changes occur in the DNA sequence comprising this chromosome.

10.3.10. Drosophila Genes Can Be Turned Off via Heritable Chromatin Structure Properties [29]

A phenomenon analogous to X-chromosome inactivation in female mammals also occurs in Drosophila. Genetic Methods have proven exceptionally powerful for studying it. In flies carrying chromosomal rearrangements that shift a central heterochromatic region into euchromatin, euchromatic genes positioned close to the heterochromatin become inactivated. This situation is analogous to attaching an autosome to an inactive mammalian X chromosome; inactivation in both cases proceeds similarly, with the zone of inactivation spreading from the chromosomal breakpoint and engulfing one or more genes. The rate of this "spreading effect" varies among cells, but an inactivation zone established in an embryonic cell is stably inherited by all subsequent cell generations (Fig. 10-38).

Studies of this effect in Drosophila have shown that the rate of spreading of the inactivation zone is reduced in flies whose genomes contain additional constitutive heterochromatin. This may be because the cells of such mutants are depleted of proteins required for heterochromatin assembly. Many Gene Mutations have a similar effect. Now that the corresponding genes have been cloned and sequenced, we can hope that the proteins involved in heterochromatin formation in Drosophila will soon be identified.

Regardless of the molecular mechanisms responsible for packaging specific Regions of the eukaryotic genome into heterochromatin, The phenomenon of heterochromatinization represents a regulatory process that fundamentally distinguishes eukaryotic cells from bacterial cells. A key feature of this unique form of regulation is that the memory of a gene's functional status is stored as a heritable chromatin structure rather than relying on a stable feedback loop of self-regulating regulatory proteins that can shift their nuclear localization. It remains unclear whether mechanisms of this type operate exclusively in the inactivation of large chromosomal domains or if they can also function at the level of individual genes or small groups of genes. The data presented below suggest that the expression of individual genes is frequently regulated by nearby control sequences and is not entirely dependent on the overall chromosomal environment.

Fig. 10-38. Position-effect mosaicism in Drosophila. The spread of heterochromatin (indicated in grey) into adjacent euchromatic regions is normally blocked by specialized boundary sequences of unknown identity. However, in flies carrying specific chromosomal translocations, these boundary sequences are absent. A. During early development in such flies, heterochromatin begins to spread into the neighboring chromosomal region, advancing varying distances in different cells. B. This spreading soon halts, but the achieved level of heterochromatin extension is inherited; as a result, large clones of descendant cells are formed in which the exact same genes are condensed into heterochromatin and are consequently inactivated (hence the "mosaic" appearance of some of these flies). This phenomenon shares many features with X-chromosome inactivation in mammals.

10.3.11. Optimal Gene Expression Frequently Requires a Specific Chromosomal Position [30]

Genes relocated to other chromosomal sites are transcribed in a cell-type-dependent manner. Consequently, they must carry the information necessary for their selective expression in appropriate cell types. A case in point is the Drosophila Sgs-3 gene, which requires approximately 600 nucleotide pairs of upstream sequence for proper transcription in Salivary Glands; it has been established that its transcription proceeds at a normal rate at almost any site (except within heterochromatin) on a polytene chromosome. Similarly, transgenic mammals carrying a tissue-specific gene fragment express that gene in the correct cells, provided the fragment is sufficiently long and retains most of the enhancer sequence.

Most introduced genes in transgenic animal cells require significantly more flanking sequence for proper expression than does the Sgs-3 gene. Moreover, varying levels of gene activity are observed among different transgenic individuals. Because such variations depend on the exact genomic integration site within the host animal's chromosome, they are referred to as position effects. Selected results obtained by inserting the mouse $\alpha$-fetoprotein gene into random genomic locations are presented in Table 10-2. In this example, the gene is active exclusively in the Tissues where it is normally expressed, yet its transcription level is typically five- to tenfold lower than normal in tissues where its activity should be high. Conversely, transcription levels can be abnormally high in tissues where the gene's expression is normally low.

If the human $\beta$-globin gene (which exhibits a strong position effect when expressed in transgenic mouse erythrocytes) is linked to a DNA fragment normally located 50,000 nucleotide pairs away from its promoter, the position effect is abolished and transcriptional activity is fully restored.

Table 10-2. Position effects typically accompany the expression of genes transferred into the mouse genome. Gene activity levels vary among independently derived transgenic individuals


Percentage of total mRNA in cells of

Yolk sac

Liver

intestine

Brain

Gene copy number per cell

Endogenous gene

20

5

0.1

0

2

Transgenic individual 1

3.4

1

0.1

0

4

Transgenic individual 2

4.8

30

1.3

0

4

Transgenic individual 3

4.4

13

4.7

0

4

Transgenic individual 4

0.4

0.4

0

0

12

In addition to the $\alpha$-fetoprotein gene, the DNA fragment microinjected into the fertilized mouse egg included a 14,000-nucleotide-pair sequence containing three enhancers that regulate $\alpha$-fetoprotein gene expression.

mRNA synthesis levels for the injected gene were compared with those produced by the normal endogenous gene in the indicated embryonic tissues using Hybridization assays. (From R. E. Hammer et al., Science 235:53–58, 1987.)

Fig. 10-39. The human $\beta$-like globin gene cluster. A. Organization of these genes. A large chromosomal region spanning 100,000 nucleotide pairs is shown. It contains a cluster of nuclease-hypersensitive sites that constitute a "domain control region" (locus control region). In cases of $\gamma\beta$-thalassemia, this control locus and all but one of the globin genes are deleted. B. Developmental changes in the expression of human $\beta$-like globin genes. Each of the globin chains encoded by these genes associates with an $\alpha$-globin chain to form erythrocyte Hemoglobin. In thalassemia patients, $\beta$-globin synthesis is substantially reduced (not shown). (A from F. Grosveld, G. B. van Assendelft, D. R. Greaves, and G. Kollias, Cell 51: 975–985, 1987.)

This fragment, which contains six nuclease-hypersensitive sites (see Section 9.1.19), affects the entire globin gene cluster; it is referred to as the locus control region (LCR). The main Conclusion to be drawn from studies of position effects is that many vertebrate genes require specific sequences located at a distance from them to achieve the proper level of expression. We shall discuss The Role of these sequences shortly.

10.3.12. Local chromatin decondensation may be required for the activation of eukaryotic genes [31]

In humans with a certain form of thalassemia (an inherited form of anemia), large deletions of DNA are found upstream of the ß-globin gene. The deleted region (~100,000 nucleotide pairs) contains several ß-globin-like genes as well as the locus control region identified in transgenic mouse experiments (Fig. 10-39A). Although the ß-globin gene itself is undamaged, its rate of transcription is significantly reduced. Unlike the normal ß-globin gene, when treated with nuclease, this gene exhibits a reaction rate as low as that of the bulk chromatin and consequently lacks The structure of active chromatin. Its normal homolog in the same erythrocyte lacks the deletion; by the time transcription begins at the first gene of this group (the s-globin gene), the entire cluster of ß-globin-like genes (90,000 nucleotide pairs) apparently decondenses into active chromatin (Fig. 10-39A).

Such findings support a two-step model for the induction of gene transcription in higher eukaryotes. In step 1, all chromatin in the region spanning tens of thousands of nucleotide pairs is converted into a relatively decondensed "active" conformation (Fig. 10-40). This step may be triggered by a specific type of regulatory protein that induces a structural change in the neighboring chromatin. This change propagates from the locus control region throughout the entire loop of that chromatin domain. In step 2, regulatory proteins acting on enhancers and promoter-proximal elements regulate the transcription of specific genes located within the exposed active chromatin region. Through this localized control, the human s-globin gene is expressed first in the embryonic yolk sac, followed by the expression of the two y-globin genes in the embryonic liver, and finally the activation of the ß-globin genes around the time of birth (Fig. 10-39B).

10.3.13. Torsional stress generated by DNA Supercoiling allows long-range action [32]

How long-range control is exerted within a chromosome remains unknown. Several mechanisms are currently known for transmitting a signal from one site in a DNA molecule to another located many thousands of NUCLEOTIDES away. In one mechanism, topological changes occur within a closed double-helical DNA loop, resulting in supercoiled DNA. DNA supercoiling has been best studied in small circular molecules, such as Plasmids and certain viral chromosomes (see Fig. 5-71). However, similar events can occur in any DNA segment constrained by two ends that cannot rotate freely (for example, a chromatin loop anchored firmly at its base).

Fig. 10-40. Chromatin changes in active genes. A. Chromatin STRUCTURE OF THE Lysozyme gene in the chicken oviduct, where the gene is active. As shown, the region of decondensed active chromatin (defined by DNase sensitivity) spans 24,000 nucleotide pairs. It contains seven nuclease-hypersensitive sites and represents a specific DNA sequence where nucleosomes are thought to be replaced by other DNA-binding proteins. B. Two stages of eukaryotic gene activation. In stage 1, the structure of a small chromatin region is modified to prepare it for decondensation prior to transcription. In stage 2, regulatory proteins bind to specific sites on the modified chromatin to induce RNA Synthesis. Transcription in prokaryotes appears to begin at stage 2 of the eukaryotic control process. (A: From A.E. Sipple et al., In: Structure and function of Eucaryotic Chromosomes [W. Hennig, ed.], Berlin: Springer-Verlag, 1987.)

A simple method for visualizing the topological stress that leads to DNA supercoiling is illustrated in Fig. 10-41. In a double-helical DNA molecule with two fixed ends, one supercoil is formed to compensate for every 10 unwound (uncoiled) nucleotide pairs. The formation of a supercoil restores the normal helical twist in the remaining regions where the bases remain paired; otherwise (for the scheme in Fig. 10-41B), the number of Base Pairs per turn would have to increase from 10 to 11. The DNA double helix resists such deformations, preferring to relieve the stress by buckling into supercoiled loops.

Bacteria such as E. coli possess a specialized enzyme called DNA gyrase, which uses the energy of ATP Hydrolysis to continuously introduce supercoils into DNA, thereby maintaining constant mechanical tension within looped domains. These are termed negative supercoils: they are twisted in the direction opposite to positive supercoils, which are generated when a segment of the helix is unwound. Because the stress caused by supercoiling is thereby reduced, untwisting the DNA double helix in E. coli is energetically more favorable than untwisting non-supercoiled DNA.

Fig. 10-41. Torsional stress from DNA supercoiling results in molecular supertwisting. A. In a DNA molecule with one free end (or a single-strand nick acting as a swivel), The Double Helix unwinds by one turn for every 10 unpaired nucleotide pairs. B. If rotation is hindered in any way, supercoils are introduced into the DNA by unwinding the helix elsewhere. As a result, one supercoil arises for every 10 unpaired nucleotide pairs. This diagram depicts a positive superhelix (see Fig. 10-42).

Fig. 10-42. Supercoiling of a DNA segment by a protein moving along the DNA double helix. As in Fig. 10-41B, the two ends of the DNA are fixed; additionally, the protein molecule is assumed to be anchored relative to these ends, or frictional forces during movement impede its free rotation. Consequently, movement generates an excess of helical turns that accumulate in the DNA ahead of the protein, while a deficit of turns forms behind it. Experimental evidence indicates that the progression of RNA polymerase molecules along the DNA template is accompanied by similar strains. The stress resulting from positive supercoiling ahead of the molecule makes that DNA region harder to open, yet this same stress should facilitate the displacement of DNA from nucleosomes—a process necessary for polymerase function during chromatin transcription (see Fig. 9-32).

DNA gyrase has not been found in eukaryotic cells; instead, eukaryotic type I and type II DNA topoisomerases relieve supercoiling-induced stress rather than enhance it (see Section 5.3.10). This is why the bulk of DNA in eukaryotic cells is not under torsional stress. Nevertheless, Transcription initiation involves unwinding of the DNA helix (see Fig. 9-65). Moreover, the movement of RNA polymerase (along with other proteins) along the DNA causes positive stress to accumulate ahead of the enzyme and negative stress behind it (Fig. 10-42). As a result of such topological alterations, an event occurring at a single site in the DNA can generate forces acting throughout the entire chromatin loop. It remains unclear whether such effects trigger further downstream events—which would be necessary if topological changes truly played a role in controlling EUKARYOTIC GENE EXPRESSION.

10.3.14. The mechanism of active chromatin formation remains elusive

According to the gene activation model shown in Fig. 10-40, certain regulatory proteins in higher eukaryotes possess Functions that distinguish them from their bacterial counterparts. Rather than merely promoting the binding ("landing") of RNA polymerase (or transcription factors) to a nearby promoter (see Fig. 10-27), some site-specific DNA-binding proteins may participate in decondensing chromatin in a specific chromosomal region, displace a nucleosome from an adjacent enhancer or promoter, and thereby grant access to conventional regulatory proteins. However, it is not entirely certain whether these events actually take place in vivo. It is quite probable that the observed differences in chromatin structure at active genes are a natural consequence of the assembly of transcription factors and/or RNA polymerase at the promoter sequence, rather than a prerequisite for initiating transcription.

Identifying specific DNA sequences acting as the locus control region of the human ß-globin gene would make it possible to isolate the proteins that bind to these sequences and to clone their genes (see Section 9.1.8). At present, we can only speculate on how such regions function. Their sequences may represent very strong enhancers (see Section 10.2.11). Three alternative hypotheses are presented in Fig. 10-43.

The fundamental differences among the models shown in Fig. 10-43 highlight how far we still are from understanding the transition of chromatin from an inactive to an active state. It is unknown how many chromatin Conformations exist and which specific structural features cause some regions to be more condensed than others. Simply based on the function of chromatin, it is impossible to explain why the Amino acid sequences of Histones (especially H3 and H4) are so evolutionarily conserved. Recently, certain chemical properties unique to active chromatin have been identified (see Section 9.2.10), and Monoclonal Antibodies have begun to be used to separate active chromatin nucleosomes (and their associated DNA sequences) from bulk nucleosomes. By applying these methods to chromatin isolated from Transgenic Animals, it should in principle be possible to distinguish regulatory regions responsible for chromatin activation from other segments, which will undoubtedly help clarify how eukaryotic genes are controlled.

10.3.15. New levels of gene control emerge during the evolution of Multicellular Organisms [33]

Yeast cells serve as a convenient model system for studying eukaryotic cells. Indeed, the high degree of functional homology between yeast and human proteins never ceases to astonish researchers. Nevertheless, not all aspects of gene expression control can be studied in yeast. Yeast cells apparently lack histone H1, and nearly all of their chromatin is in an active state. Furthermore, they lack DNA control regions capable of exerting long-range effects on genes. Yeast cells also do not utilize alternative RNA splicing, a topic we shall discuss shortly.

Fig. 10-43. Three hypotheses proposed to explain the long-range action of the locus control region on genetic activity. It is generally assumed that a chromatin loop represents an entire chromosomal domain that is looped out and contains 100,000 or more nucleotide pairs of DNA. In reality, it remains unknown what accounts for this long-range action, and other mechanisms may well be operating. Evidence that chromatin undergoes decondensation independently of DNA Transcription comes from observations of polytene chromosome puffs containing the Sgs-3 gene in Drosophila. Certain mutants unable to produce RNA at this site nevertheless form a puff during larval development (see Fig. 9-48).

Fig. 10-44. A. 5-Methylcytosine is formed by the methylation of cytosine within a DNA strand. In vertebrates, this process is restricted to cytosine residues (C) occurring within CG sequences. B. The synthetic nucleotide containing the 5-azacytosine base (5-aza-C) cannot be methylated. Moreover, when small amounts of 5-aza-C are incorporated into DNA, they inhibit the methylation of normal cytosines.

Invertebrates, such as Drosophila, possess larger genomes than yeasts, and contain histone H1, heterochromatin, and at least some types of active chromatin (see Fig. 9-54). Invertebrates also utilize alternative RNA splicing. However, one level of gene expression control characteristic of vertebrates appears to be absent in invertebrates: they lack a general DNA Methylation-based silencing system.

10-20

10.3.16. During vertebrate cell division, the pattern of DNA methylation is inherited [34]

DNA bases undergo modification in certain instances. For example, we have previously discussed how methylation of A in the GATC sequence leads to replication errors in bacteria (see Section 5.3.8); conversely, methylation of A or C at a specific site protects the bacterium from its own restriction enzymes (see Section 4.6.2). Vertebrate DNA contains 5-methylcytosine (5-methyl-C), which presumably has no effect on base-pairing (Fig. 10-44, A). Methylation is restricted to C residues within CG sequences. Because this sequence is precisely paired with an identical sequence (in reverse orientation) on the opposite DNA strand, the inheritance of an existing DNA methylation pattern is ensured by a simple copying mechanism. An enzyme called maintenance methylase acts only on CG sequences that are paired with already methylated CG sequences. As a result, the pre-existing methylation pattern is automatically inherited during DNA replication (Fig. 10-45).

Fig. 10-45. Faithful inheritance of the DNA methylation pattern. In vertebrate DNA, the majority of cytosine residues within CG sequences are methylated (see Fig. 10-44). Once the DNA is tagged with methyl groups, each methylation site is inherited by the daughter DNA molecules. This means that modifications in the DNA methylation pattern are preserved and passed on to progeny cells.

Some restriction enzymes cleave DNA at sites containing the CG dinucleotide, and their activity requires this sequence to remain unmethylated. For example, HpaII cleaves the CCGG sequence, but fails to do so if the central C is methylated. Thus, the sensitivity of DNA to the restriction enzyme HpaII can serve as an assay for the methylation of specific sites. Under normal conditions, the HpaII methylase protects the bacterium from its own HpaII restriction enzyme. Consequently, this methylase can be used to introduce 5-methyl-C bases into specific CG sequences (namely CCGG) within cloned DNA molecules (since DNA Cloning in E. coli results in the loss of methylation at CG sites). Using this technique, it has been demonstrated that each individual methylated CG sequence is typically preserved through many cell divisions in vertebrate cell cultures, whereas unmethylated CG sequences remain unmethylated.

The automatic inheritance of 5-methyl-C raises a "chicken-and-egg" problem: at what stage does the initial methylation occur in vertebrates? Studies have shown that when unmethylated DNA is introduced into a fertilized mouse egg, nearly all of the CG sites in this DNA become methylated (with one notable exception described below). Thus, the bulk of the mammalian genome can become heavily methylated. Since maintenance methylases are known to be normally incapable of unmethylated DNA, one must assume that a different enzyme is present in the egg: de novo methylase. Having performed its function, it presumably disappears, leaving the preservation of methylated nucleotides in the DNA of developing tissues to the maintenance methylase.

10-5

10-20

10.3.17. DNA methylation in vertebrates helps cells maintain their determined developmental pathway [35]

What is the Biological Role of CG methylation? Experiments using the restriction enzyme HpaII indicate that the DNA of inactive genes is more heavily methylated than that of active genes. Furthermore, it has been shown that an inactive gene whose DNA is methylated loses most of its methyl groups upon activation. Compelling evidence that methylation directly influences gene expression comes from experiments with 5-azacytidine (5-aza-C, see Fig. 10-44, B), which was briefly added to cultured cells. 5-aza-C, a base analog incapable of being methylated, is incorporated into DNA and acts as an inhibitor of maintenance methylase, thereby reducing the overall level of DNA methylation. In cells treated with 5-aza-C, certain previously inactive genes became active while simultaneously acquiring unmethylated C residues. Once activated, the active state of these genes was maintained in the absence of 5-aza-C across many subsequent cell generations. This indicates that initial gene methylation promoted their inactive state.

Based on results from 5-azacytidine experiments, it was hypothesized that DNA methylation might play a primary role in generating diverse cell types. This would imply that cell specialization in vertebrates and invertebrates occurs through different mechanisms. Subsequent research revealed, however, that the role of DNA methylation in cell differentiation is likely less pivotal and more of a secondary, auxiliary nature. Crucial developmental switches are executed by regulatory proteins capable of influencing gene activity independently of methylation. For instance, the X chromosome in females is first condensed and inactivated, and only later do some of its genes become heavily methylated. Conversely, certain genes active in the liver are turned on during development while still fully methylated, with their methylation levels decreasing only at a later stage.

The Significance of DNA methylation for gene expression was largely clarified through transfection experiments. For example, the tissue-specific gene encoding Muscle Actin was isolated in both fully methylated and fully unmethylated forms. When both variants of this gene were introduced into a muscle cell culture, they were transcribed with equal efficiency. However, when the gene was introduced into fibroblasts—where it is not normally expressed—the unmethylated variant was transcribed at a low level, which was nevertheless higher than that of the introduced methylated gene or the endogenous gene present in the fibroblast (which is also methylated). These experiments lead to the conclusion that in vertebrates, DNA methylation serves to lock in developmental pathways chosen by other means.

In certain cases, it is possible to determine with high precision whether DNA sequences transcribed with high efficiency in one cell type can be transcribed in another vertebrate cell type. Such experiments have shown that transcription levels in two different vertebrate cell types can sometimes differ by more than a factor of 106. Unexpressed vertebrate genes are much more tightly repressed transcriptionally than unexpressed bacterial genes, where the largest known differences in transcription rates between expressed and unexpressed genes do not exceed 1000-fold. By contributing to the further downregulation of genes already silenced by other mechanisms, DNA methylation likely accounts for this difference.

10.3.18. CG-rich islands allow the identification of approximately 30,000 housekeeping genes in mammals [36]

During evolution, methylated residues in the genome tend to disappear due to Specific features of DNA Repair Mechanisms. Specifically, the Spontaneous deamination of unmethylated cytosine produces uracil, which is not normally found in DNA and is therefore readily recognized by repair enzymes, excised, and replaced with cytosine. The consequences of spontaneous deamination of a 5-methylcytosine residue cannot be corrected in this manner, because deamination of 5-methyl-C yields thymine, which is indistinguishable from other non-mutant thymine residues naturally present in DNA. Consequently, over evolutionary time, methylated cytosine residues in the genome tend to mutate into thymine.

Since the divergence of vertebrates and invertebrates approximately 400 million years ago, more than three out of every four CG sequences have been lost in this manner, resulting in a dramatic reduction of this dinucleotide in vertebrate genomes. The remaining CG sequences are distributed very unevenly across the genome: there are discrete regions 1,000 to 2,000 nucleotides long (CG islands) where the CG content is 10 to 20 times higher than the genome average. Such islands flank the promoters of so-called housekeeping genes—those genes encoding proteins essential for basic cellular maintenance and thus expressed in all cell types (Fig. 10-46). These genes contrast sharply with tissue-specific genes, which encode proteins unique to strictly defined cell types.

Fig. 10-46. CG islands in three mammalian housekeeping genes. Rectangles indicate the size of each island, which are shown flanking the promoter of each gene. Note that in most mammalian genes, exons (shaded) are shorter than introns. (After A. R. Bird, Trends Genet. 3: 342-347, 1987.)

The distribution of CG islands is readily explained as a byproduct of the Evolution of the CG methylation system, which evolved to suppress inactive GENE EXPRESSION IN vertebrates (Fig. 10-47). In germ cells, all tissue-specific genes (except those specific to the egg and sperm) are inactive and methylated. Over long evolutionary timescales, their methylated CG sites were lost via spontaneous deamination. However, CG sequences within the promoter regions of genes active in germ cells (including all housekeeping genes) remained unmethylated and were consistently repaired following spontaneous deamination. It is believed that these genes are recognized by site-specific DNA-binding proteins present in germ cells, which strip away any methyl groups in the promoter regions of housekeeping genes. Experiments with cloned genes demonstrate that only CG sequences within CG islands remain unmethylated when completely unmethylated DNA is injected into mouse eggs.

It is known that the mammalian genome (approximately 3 × 109 nucleotide pairs) contains about 30,000 CG islands, each approximately 1,000 nucleotide pairs long. Most islands coincide with the 5' end of a transcription unit and thus presumably with the 5' end of a gene. Because DNA surrounding CG islands can be isolated and cloned separately, identifying and characterizing vital housekeeping genes is a relatively straightforward task. Presumably, the remaining tens of thousands of genes are tissue-specific and are not expressed in germ cells. Because the majority of their CG sequences have been lost, these tissue-specific genes are much more difficult to identify in the genome.

Fig. 10-47. Scheme illustrating the marked depletion of CG sequences and the presence of CG islands. Black lines indicate the positions of unmethylated CG dinucleotides in the DNA sequence, while red lines denote methylated CG dinucleotides. Methylated genes are unknown in invertebrates.

10.3.19. Complex gene regulation is required for the formation of a multicellular organism [37]

Embryonic cells have a lot in common with computers: they constantly receive information about their current position and integrate it with previously acquired data to act accordingly at each stage of development. Genetic studies of Drosophila have demonstrated that the formation and Maintenance of the fundamental body plan involve a relatively small number (around 100) of genes encoding master regulatory proteins that interact with one another. In any multicellular organism, the vast majority of genes—both vital and tissue-specific—are likely regulated through complex control cascades originating from The genes of these master regulatory proteins. Given that eukaryotic gene regulation relies heavily on mechanisms vastly different from those in bacteria (such as mechanisms dependent on the direct inheritance of chromatin structure), it is reasonable to expect that these very mechanisms control certain master regulatory genes.

At present, little is known about how the expression of master regulatory genes is controlled in vertebrates, but initial insights are beginning to emerge from Drosophila. For instance, homeotic genes that determine the distinct development of individual fly body segments reside within two complex loci, Antennapedia and Bithorax. The bithorax complex, responsible for the differentiation of two thoracic and eight abdominal segments, contains three transcriptional units designated as Ubx, abd-A, and abd-B. Each transcriptional unit apparently encodes a family of regulatory proteins generated via alternative RNA splicing. It is estimated that the coding regions of this complex comprise fewer than 20,000 nucleotide pairs, whereas the regulatory sequences span approximately 300,000 nucleotide pairs. The regulatory regions appear to consist of a series of enhancers whose linear arrangement along the chromosome mirrors the body segments they affect (Fig. 10-48). These findings, combined with evidence from mutant flies in which one of the enhancers has been translocated from one part of the bithorax complex to another, suggest that gene regulation within the complex is mediated by changes in chromatin structure. According to this model, chromatin domains open up sequentially in an orderly progression toward the posterior end of the body in cells comprising the bithorax complex. This sequential opening allows enhancers localized within these domains to become activated in a strict temporal and spatial order (Fig. 10-49). It is highly likely that the mechanisms governing master regulatory genes are similarly complex.

Fig. 10-48. ORGANIZATION OF THE bithorax complex in Drosophila. This crucial 300,000-nucleotide-pair chromosomal region contains three genes—Ubx, abd-A, and Abd-B—which encode master regulatory proteins controlling the Development of the Thorax and abdomen. Homeotic mutations have helped identify nine additional groups of DNA regulatory sequences. Each of these groups is required for the development of specified parasegments as well as other parasegments located closer to the posterior end of the body. These DNA regulatory sequences are thought to act as enhancers, controlling the expression of a nearby gene, and their order along the DNA molecule corresponds to the body segments they influence (see Fig. 10-49). (From M. Peifer, F. Karch, and W. Bender, Genes Dev. 1: 891–898, 1987.)

Fig. 10-49. Diagram illustrating the precise correspondence between the chromosomal position of each regulatory region in the bithorax complex and the parasegments of the fly body affected by mutations in that region. (A) Control at the level of chromatin structure modification. It is proposed that chromatin is decondensed or progressively activated in increasingly posterior parasegments; thus, in parasegment 5, only the bbx/bx regulatory region is open, whereas all other regulatory regions are exposed in the more posterior parasegments controlled by that complex (parasegment 13, see Fig. 10-48). Only three parasegments are shown in this diagram. The Ubx gene can generate several distinct transcripts (indicated here by squares and circles), the Selection of which is governed by the regulatory regions. Consequently, the bithorax complex produces a unique cocktail of regulatory proteins in each parasegment (B).

Conclusion

Animal and plant organisms employ mechanisms that ensure different genes are transcribed in different cell types. Because many specialized cells retain their unique properties when grown in culture, gene regulation mechanisms must be stable and heritable. Prokaryotes and yeast provide exceptionally tractable model systems for investigating gene regulatory mechanisms. Some of these mechanisms may also play a role in generating specialized cell types in higher eukaryotes. One such mechanism is the competitive interaction between two or more master regulatory proteins, each of which suppresses the synthesis of the other while stimulating its own production.

Studies of the expression of engineered genes, as well as gene fragments integrated into random genomic sites in transgenic animals, have demonstrated that most higher eukaryotic genes are controlled by a combinatorial mix of diffusing regulatory proteins that are unique to each cell type. Furthermore, gene expression in higher eukaryotes can be influenced by transitions of chromatin between more or less condensed states; in vertebrates, DNA methylation suppresses the transcription of inactive genes. Nevertheless, for the majority of genes, these additional levels of control can either be dictated or overridden by diffusing regulatory proteins. It remains unknown how vertebrate cells control the genes encoding the master regulatory proteins that ultimately determine cell fate.



Last update: 12/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.