Molecular Biology of the Cell - Volume 2 - Alberts B., Bray D., Lewis J., Raff M., Roberts K., Watson J. 1993

Control of gene expression
Control of transcription initiation

Just 40 years ago, the idea that a Gene could be switched on or off seemed absurd. The hypothesis that played such a crucial role in our understanding of cellular function was put forward based on studies of E. coli growing on a mixture of glucose and lactose (a disaccharide). When given a choice of carbon sources, the bacterium would preferentially consume all the glucose before beginning to metabolize lactose. This shift to lactose was accompanied by a growth lag, during which The Cell synthesized the enzyme ß-galactosidase to hydrolyze lactose into glucose and galactose. The isolation and characterization of mutant Bacteria with specific defects in regulating this transition sparked a wave of biochemical research that, in 1966, led to the identification and Isolation of the lactose Operon repressor protein.

As a result of biochemical and genetic studies on the lac operon repressor, the bacteriophage lambda repressor, and other bacterial regulatory Proteins, a general model for Prokaryotic METABOLISM/31.html">Transcription regulation was formulated. It was hypothesized that site-specific proteins either inhibit or stimulate gene transcription by binding to DNA near the promoter—the region where RNA polymerase initiates RNA Synthesis. It was believed that Changes in the positioning of these regulatory proteins relative to DNA (binding or dissociation) act as the molecular switch for turning genes on and off.

For many years, it remained unknown whether this model of Genetic control could be applied to Eukaryotic Cells. It is well established that eukaryotic DNA is packaged with histone proteins into nucleosomes, whose mass equals that of the DNA itself. The presence of nucleosomes implies the existence of alternative gene regulation pathways. This is further supported by the observation that in eukaryotes, regulatory proteins often bind to sites located thousands of nucleotide pairs away from the promoter they control. Progress in elucidating the mechanisms of eukaryotic gene regulation was hindered largely because most regulatory proteins are present in extremely low cellular concentrations (approximately one per 50,000 protein molecules). It was not until 1983 that researchers successfully characterized two unusually abundant (and perhaps atypical) regulatory proteins: the large T antigen of the SV40 virus and the TFIIIA transcription factor of the 5S rRNA gene.

Recently, the situation has fundamentally changed. Thanks to advances in Introduction/32.html">Genetic Engineering, a vast array of eukaryotic regulatory proteins has become accessible for biochemical and genetic analysis. Furthermore, bacterial regulatory proteins that act over a distance have also been discovered. This section focuses on the common principles uniting the mechanisms of Gene Transcription Regulation in PROKARYOTES AND EUKARYOTES. Processes that may be unique to eukaryotes will be discussed later in the context of Cell Differentiation (see Section 10.3.8).

10-8

10.2.1. Bacterial Repressor Proteins Bind Near Promoters and Suppress Transcription of Specific Genes [9]

The E. coli chromosome consists of a single circular DNA molecule comprising ~4.7 × 106 nucleotide pairs. In principle, this amount of genetic material is sufficient to encode roughly 4,000 different proteins; however, at any given time, the cell synthesizes only a specific subset of them. The expression of many E. coli genes depends on the intracellular levels of specific metabolites.

Investigations into the Genetic Control of lactose utilization in E. coli led researchers to conclude that a repressor protein exists which, in the absence of lactose in the medium, turns off the synthesis of ß-galactosidase. The isolation, purification, and in vitro application of this protein made it possible to decipher The Mechanism of this regulation. It was found that the repressor inhibits lac operon transcription by binding to a 21-nucleotide DNA sequence known as the operator, which overlaps the adjacent RNA polymerase binding site (the promoter). As long as the repressor remains bound to the operator, access for RNA polymerase to its binding site is blocked, and Transcription of the downstream DNA regions cannot occur (see Fig. 10-14).

Class="center">

Fig. 10-11. Mechanism of derepression in bacteria. A small specific molecule binds to the repressor protein, inducing a conformational change that causes it to dissociate from the DNA, thereby permitting transcription of the adjacent genes. In the example described in the text, allolactose binds to the lactose operon repressor protein, resulting in the synthesis of a single RNA transcript that encodes three proteins involved in lactose metabolism (ß-galactosidase, galactoside permease, and galactoside acetylase). Because the three genes encoding these proteins are clustered adjacently and their regulation is tightly coordinated, this gene cluster is referred to as the lactose operon.

The lactose operon repressor protein ensures that the synthesis rate of ß-galactosidase—the enzyme required for lactose Cleavage—matches the metabolic demands of the cell. The lac repressor, in turn, is controlled by allolactose, a small carbohydrate molecule generated within the cell when lactose is present. When intracellular allolactose reaches a sufficiently high level, it acts as an allosteric regulator, inducing Conformational Changes in the repressor protein. This weakens its affinity for DNA to the point where it detaches, freeing the promoter and allowing RNA polymerase to transcribe the adjacent DNA regions. This process is known as gene derepression. Such a regulatory system enables E. coli to produce the enzyme necessary for lactose breakdown only when required (Fig. 10-11).

Many other Examples of specific bacterial gene repression are now known. In each case, the binding of a repressor protein to a specific DNA sequence results in gene silencing. This binding process is invariably regulated by specific signaling molecules analogous to allolactose. Sometimes, as with the lac operon repressor, the presence of signaling molecules in the cell turns a gene or transcription unit on by diminishing the repressor's affinity for its target DNA sequence. Conversely, a signaling molecule can also be used to turn a gene off via a repressor protein. For example, an allosteric shift triggered by the binding of a signaling molecule can increase, rather than decrease, the repressor's capacity to bind a specific DNA sequence. This mechanism operates in the control of five adjacent genes encoding the Enzymes required for Tryptophan Biosynthesis in E. coli (the trp operon). The synthesis of a single long mRNA molecule encoding these five proteins is governed by a repressor protein that binds to the DNA only when complexed with tryptophan (the signaling molecule that activates this operon) (Fig. 10-12).

Fig. 10-12. The binding of tryptophan to the tryptophan operon repressor protein alters the repressor's conformation. These conformational changes enable the regulatory protein to bind tightly to a specific DNA sequence, thereby blocking the transcription of genes encoding proteins involved in tryptophan biosynthesis (the trp operon). The three-dimensional Structure of this bacterial protein (helix-turn-helix), determined by X-Ray Diffraction Analysis, is shown both with and without bound tryptophan. Tryptophan binding increases the distance between the two recognition helices (colored cylinders) within the dimer, facilitating The formation of symmetrically positioned Hydrogen Bonds, depicted on the schematic as colored rays. (After R. Zhang et al., Nature, 327: 591-597, 1987.)

Fig. 10-13. GENERALIZED SCHEME OF the various mechanisms by which specific regulatory proteins control gene transcription in prokaryotes. A. Negative control; B. Positive control. Note that an inducer Ligand can turn a gene on either by causing the removal of a repressor protein (top left) or by promoting the attachment of an activator protein (bottom right). Similarly, an inhibitor ligand can turn a gene off either by removing an activator protein (top right) or by promoting the binding of a repressor protein to the DNA (bottom left).

In the case of the lac operon, an increase in signaling molecule concentration promotes the release of the repressor from the DNA and activates transcription, whereas in the trp operon, it induces repressor binding to the DNA, thereby suppressing transcription. It should be emphasized, however, that both mechanisms are fundamentally similar: in both instances, transcription proceeds in the absence of the regulatory protein. This type of genetic control is termed negative regulation (Fig. 10-13, A).

10-10

10.2.2. Bacterial Activator Proteins Interact with RNA Polymerase and Promote Transcription initiation [10]

In negative regulation, a receptor protein binds to a site adjacent to the promoter and represses RNA polymerase activity. An alternative mode of gene activity regulation relies on activator proteins that enhance RNA polymerase function. In E. coli, such positive regulation plays a vital role in activating transcription units characterized by relatively weak promoters that poorly bind polymerase on their own. The attachment of an activator protein to a specific DNA sequence near the promoter facilitates the "landing" of RNA polymerase, ultimately increasing the probability of transcription.

Despite differences in their Mechanisms of action, activator and repressor proteins share many common properties. Furthermore, certain bacterial regulatory proteins can function as both Repressors and activators, binding to different sites where they either suppress or activate transcription. Like repressors, activators frequently associate with specific signaling ligands that either increase or decrease their DNA-binding affinity, thereby turning genes on or off accordingly. This form of Transcriptional Regulation is known as positive regulation because the efficiency of RNA synthesis increases in the presence of the regulatory protein (Fig. 10-13, B).

Fig. 10-14. Glucose and lactose concentrations control the initiation of lac operon transcription by acting on the lac operon repressor protein and CAP. The addition of lactose leads to an increase in allolactose concentration, which removes the repressor protein from the DNA (see Fig. 10-11). The addition of glucose causes a drop in cAMP levels; because cAMP is no longer available to bind CAP, this activator protein dissociates from the DNA, thereby switching off the operon. Binding sites and proteins are drawn to scale. As illustrated, CAP is believed to make direct contact with the polymerase, facilitating the initiation of RNA synthesis.

The most thoroughly studied example of an activator protein is the catabolite activator protein (CAP) of E. coli. This protein enables the bacterium to utilize alternative carbon sources when glucose, its preferred carbon source, is absent.

In the case of the lac operon, the interplay between the repressor and CAP dictates the characteristic pattern of Gene Expression. As noted above, in the presence of lactose, allolactose dissociates the repressor from the DNA. However, this alone is insufficient to activate lac operon transcription because its promoter bears only a distant resemblance to the consensus E. coli promoter sequence and consequently exhibits very low affinity for RNA polymerase. Transcription from this promoter requires that RNA polymerase binding be reinforced by the attachment of CAP to a site located immediately upstream of the promoter (Fig. 10-14). A similar situation is characteristic of promoters for many genes involved in sugar metabolism.

The binding of DNA to CAP is regulated by glucose, which ensures that alternative carbon sources are utilized by the cell only in the absence of glucose. Glucose starvation triggers an increase in the intracellular concentration of cAMP, which serves as a key signaling molecule in both bacteria and eukaryotes. When cAMP binds to the CAP protein, it induces a conformational change that enables the protein to bind to specific DNA sequences and activate the transcription of neighboring genes. Conversely, high glucose levels cause cAMP concentrations to drop; cAMP dissociates from the CAP protein, rendering it inactive and incapable of DNA binding. As a result, the cell switches back to metabolizing glucose (see Fig. 10-14).

10-9

10.2.3. Changes in protein phosphorylation can influence gene activity [11]

All early studies on repressor proteins were conducted using bacteria. It was discovered that The activity of these proteins in both the lactose and tryptophan operons is controlled through the Reversible Binding of small, specific molecules. In eukaryotic cells, regulatory proteins are similarly controlled by small signaling molecules, such as cAMP. Rather than acting directly, these molecules exert their effects by modulating protein phosphorylation and dephosphorylation. Although phosphorylation does not play as central a regulatory role in bacteria, they do possess one well-characterized control system dependent on protein phosphorylation. Using this system as a model, we will examine certain aspects of gene regulation that help us understand the much more complex regulatory networks of higher eukaryotes.

In bacteria, related proteins are known to regulate nitrogen and phosphate metabolism, membrane Protein Synthesis, chemotaxis, and sporulation. Here, we focus on Nitrogen metabolism in E. coli, where the synthesis of several proteins is upregulated under nitrogen-limiting conditions. One such protein is Glutamine Synthetase, a key enzyme in nitrogen assimilation that catalyzes the reaction: glutamic acid + ammonia → glutamine.

The activation of nitrogen-metabolizing genes involves two primary regulatory components: the ntrC protein, an activator that can switch on genes only when phosphorylated (ntrC-phosphate), and the ntrB protein, which can either phosphorylate ntrC (via its kinase activity) or dephosphorylate it (via its phosphatase activity). Phosphorylation of ntrC is triggered by a drop in nitrogen levels, which leads to the attachment of uridine monophosphate (UMP) groups to the regulatory subunit of ntrB. This modification enhances the relative kinase activity of the ntrB enzyme.

Cellular nitrogen demand is determined by the $\alpha$-ketoglutarate-to-glutamine ratio; an increase in this ratio stimulates the phosphorylation of ntrC and thereby induces transcription of the glutamine synthetase gene (Fig. 10-15). Similar cascades of protein modification occur in eukaryotic cells, where they likewise result in the phosphorylation or dephosphorylation of regulatory proteins, although the precise details remain much less understood.

Fig. 10-15. Part of the regulatory cascade that activates genes involved in bacterial nitrogen metabolism during nitrogen starvation. The PII protein acts as a regulatory subunit for the ntrB protein. The advantage of such multi-step regulation is that it can elicit substantial metabolic shifts in response to relatively minor fluctuations in nitrogen availability. Side branches of this pathway also reversibly modify Other Enzymes, modulating their catalytic activity in response to changing nitrogen concentrations.

Fig. 10-16. Regulatory region of the bacterial glutamine synthetase gene. Glutamine synthetase catalyzes the reaction: glutamic acid + ammonia → glutamine. The two brightly highlighted sites strongly bind the ntrC protein and are essential for transcription activation.

10-10

10.2.4. DNA flexibility allows regulatory proteins bound to distant sites to influence gene transcription [11, 12]

The two-component regulatory system in bacteria described above shares many features with eukaryotic gene control systems. Specifically, the DNA contains two strong binding sites for the ntrC protein located 100 or more nucleotide pairs upstream of the glutamine synthetase gene promoter (Fig. 10-16). In its phosphorylated form (ntrC-phosphate), ntrC binds to these sites, which then effectively stimulate transcription. Furthermore, this stimulation persists even if these sites are genetically engineered further upstream, at a distance of more than 100 nucleotide pairs from the promoter.

While such "action at a distance" might seem unusual in prokaryotes, it is a widespread phenomenon in eukaryotes. In prokaryotes, there is solid evidence that proteins bound to distant DNA sites operate largely by the same mechanisms as those attached to adjacent regions. For instance, the ntrC protein enhances the ability of RNA polymerase to bind to the promoter and form an open initiation complex (see Fig. 9-65). The intervening DNA forms a loop, allowing the proteins bound at distant sites to interact directly with the polymerase (Fig. 10-17). Such interactions occur quite readily because DNA acts as a tether, ensuring that a protein bound even several thousand nucleotide pairs away makes frequent collisions with the promoter (Fig. 10-18).

10.2.5. Different sigma factors enable bacterial RNA polymerase to recognize distinct promoters [11, 13]

As discussed in Chapter 9, bacteria possess another level of transcriptional control. When examining The structure of bacterial RNA polymerase, we noted that one of its five core polypeptide chains—the sigma ($\sigma$) factor—Functions as an initiation factor. This factor dissociates from the DNA as soon as the polymerase initiates RNA synthesis (see Section 9.4.1). In bacteria, consensus promoter sequences are typically recognized by the main form of RNA polymerase containing the $\sigma^{70}$ subunit. However, minor polymerase variants also exist that incorporate alternative $\sigma$ subunits capable of recognizing different promoter sequences. For example, the $\sigma^{54}$ subunit is the product of the ntrA gene and, along with the ntrB and ntrC proteins, is specifically required for the transcription of the glutamine synthetase gene. It is entirely plausible that environmental cues can switch on specific subsets of genes through the synthesis of dedicated sigma factors.

Fig. 10-17. Schematic model of the enhancer action of the ntrC protein on RNA synthesis at the E. coli glutamine synthetase gene. By binding to DNA sequences upstream of the gene, this regulatory protein increases The rate of transcription by RNA polymerase. Although the protein can bind to these sites even when unphosphorylated, only the phosphorylated form (ntrC-phosphate) can activate transcription. This activation likely requires physical contact with the polymerase, which enhances the enzyme's intrinsic ability to unwind DNA and form an open complex (see Fig. 9-65), as illustrated here (see also Fig. 10-18).

Fig. 10-18. The binding of two proteins to separate sites on a DNA double helix can dramatically increase the probability of their interaction. (A) Protein binding facilitated by a 500-base-pair DNA loop increases the frequency of collisions. Color intensity reflects the probability that the colored protein will be located at any given point in space surrounding the white-circled protein. (B) DNA flexibility is such that The Double Helix undergoes smooth bends of about 90° (bent turns) roughly once every 200 nucleotide pairs. Consequently, if two proteins are separated by only 100 nucleotide pairs, their contact is relatively constrained. In such cases, protein-protein interaction is facilitated if they are positioned on the same face of the DNA helix (which contains 10 NUCLEOTIDES per turn). (C) Effective concentration of the colored protein at the binding site of the white protein as a function of the intervening distance. (Courtesy of Gregory Bellomy, after M.C. Mossing and M.T. Record, Science 233: 889-892, 1986.)

How can we distinguish an RNA polymerase initiation factor, such as $\sigma^{54}$, which affects promoter recognition, from a regulatory protein like ntrC? One approach relies on the fact that $\sigma$ factors bind tightly to RNA polymerase but are incapable of binding specific DNA sequences on their own. Conversely, regulatory proteins bind specifically to DNA rather than to RNA polymerase. Another approach takes advantage of the fact that only one $\sigma$ subunit associates with an RNA polymerase molecule at any given time. By varying The ratio of a known $\sigma$ factor (e.g., $\sigma^{70}$) to the test protein in in vitro transcription assays, one can draw a definitive Conclusion. If $\sigma^{70}$ inhibits transcription in proportion to this ratio (competitive inhibition), the test Protein Functions as a $\sigma$ subunit.

Fig. 10-19. Mechanisms by which proteins control the initiation of RNA transcription in bacterial cells. Proteins can act by either stimulating or blocking various steps required to launch productive RNA synthesis (see Fig. 9-65). Mechanism (D) was first discovered during studies of the E. coli arabinose operon. It represents a combination of mechanisms A and B, and is the only one of the four mechanisms not previously discussed in the text. This variant will be examined below in the context of eukaryotes (see Fig. 10-27).

The major types of mechanisms regulating the initiation of gene transcription in bacteria are summarized in Fig. 10-19. Bacterial systems will also be considered when exploring how molecular switches can operate through competitive interactions between regulatory proteins.

10-11

10.2.6. The requirement for general transcription factors leads to the presence of additional elements in the eukaryotic gene transcription control system [14]

Transcription in eukaryotes is a significantly more complex process than in Prokaryotic Cells. The three eukaryotic RNA polymerases (polymerases I, II, and III) are evolutionarily related to the bacterial enzyme, but contain a greater number of subunits. To date, these enzyme complexes have not been fully characterized. Furthermore, whereas bacterial polymerase recognizes a specific DNA sequence, the eukaryotic enzyme typically relies on a DNA-protein complex formed by transcription factors.

Here and elsewhere in this section, we will primarily focus on RNA polymerase II, which synthesizes all mRNA precursors. A vital component of the DNA-protein complex recognized by this polymerase is the TATA-binding protein, also known as transcription factor IID (TFIID). This protein binds to the consensus sequence known as the "TATA box," TATAWAW. Typically, this sequence is located approximately 25 to 30 Base Pairs upstream of the transcription initiation site. For many protein-coding genes, THE POSITION OF this sequence is crucial both for promoter activity and for the precise Determination of the RNA chain initiation site. Much like regulatory proteins, the TATA-binding protein specifically binds to DNA. Because this protein is required for the transcription of most (and possibly all) genes transcribed by polymerase II, it is considered a core component of the general transcription machinery rather than a sequence-specific regulatory protein. In vitro experimental results indicate that upon binding to DNA at the promoter region, the TATA-binding protein typically remains part of a stable transcription complex that undergoes multiple rounds of transcription by RNA polymerase II (see Fig. 9-69).

It is likely that all the mechanisms employed by bacteria to control RNA polymerase activity are also operational in eukaryotic cells (see Fig. 10-19). However, the formation of a stable transcription complex on DNA involving the TATA-binding protein undoubtedly increases The complexity of gene regulation in eukaryotes. In vitro experiments suggest that a primary function of certain eukaryotic activator proteins is to assist the TATA-binding protein in binding to DNA at the promoter region.

10-11

10.2.7. Most eukaryotic genes are controlled by promoters and enhancers [15]

Genetic engineering techniques have vastly expanded our ability to study eukaryotic genes. Researchers can custom-modify The nucleotide sequence of any DNA fragment, generate numerous copies of that sequence, and attach it to virtually any other gene, thereby creating a novel gene. To assess the function of such a constructed gene, it is either introduced into cultured cells via transfection or, far more labor-intensively, used to generate a transgenic animal in which the gene is stably integrated into one of the Chromosomes.

The application of these methodologies has made it possible to identify the regulatory regions of eukaryotic genes even in the absence of direct data regarding their cognate regulatory binding proteins. It has been found that the DNA segment immediately adjacent to the RNA initiation site is critical for efficient transcription (the promoter-proximal element). This region typically spans about 100 nucleotide pairs and encompasses the TATA box. Even more surprisingly, efficient transcription requires sequences located quite far upstream or downstream from the promoter (enhancers). Much like the ntrC protein-binding site in bacteria, eukaryotic enhancers influence transcription over a distance.

The first enhancer to be studied was a small fragment of the SV40 viral genome that enhances the transcription of viral genes required during Cytology/cytology/16.html">Early stages of infection. In 1981, it was demonstrated that attaching this small viral fragment to the ß-globin gene increases the level of ß-globin transcription more than 100-fold. By altering the position and orientation of this SV40 segment, researchers established the following defining characteristics. An enhancer is a regulatory DNA sequence that:

1) activates transcription from an associated promoter, with transcription initiating at the normal RNA synthesis start site;

2) functions in either orientation (normal or inverted);

3) exerts its effect from a distance of over 100 nucleotide pairs in either direction from the promoter.

It is now well established that The regulation of most higher eukaryotic genes involves both promoter-proximal elements and one or more enhancers located far from the RNA initiation site (see Fig. 10-22A). Although enhancers are typically situated several thousand base pairs away from the promoter they regulate, some can act over distances exceeding 20,000 base pairs. Both enhancers and promoter-proximal elements associate with site-specific DNA-binding proteins that mediate their effects. Some of these regulatory proteins appear to be present in all cell types, whereas others are restricted to specific cell lineages.

10-11

10.2.8. Most enhancers and promoter-proximal elements are sequences that bind proteins involved in combinatorial control [15, 16]

The activity of each enhancer and promoter-proximal element is typically focused on a sequence spanning 100 to 200 base pairs. An element of this size contains binding sites for numerous proteins and can generally be dissected by genetic engineering into smaller DNA fragments that retain partial activity. The function of both types of regulatory sequences depends on the binding of specific proteins and is therefore often restricted to cell types where those proteins are present. Only a few regulatory elements (such as the SV40 enhancer) can function in almost all cell types. Most enhancers exhibit marked cell Specificity and are active primarily in cells expressing the gene with which they are normally associated. A classic example is the chicken ß-globin gene enhancer.

Fig. 10-20. A, Construction of a series of regulatory region mutants via oligonucleotide-directed mutagenesis. In each mutant, a block of four nucleotide pairs is altered. Thus, to analyze 108 nucleotides in the ß-globin enhancer, 27 consecutive mutants must be generated. B, Insertion of mutant ß-globin enhancers into a specialized test plasmid. The oligonucleotide and the cloning site are "stitched" together using DNA ligase. The protein produced by the recombinant gene is the bacterial enzyme chloramphenicol acetyltransferase (CAT), whose activity is easily measured. C, Analysis of mutant enhancers based on their effect on RNA synthesis. RNA synthesis is measured indirectly by The amount of protein produced by the recombinant gene (i.e., CAT activity).

The chicken ß-globin enhancer is located downstream of the ß-globin transcription unit. In successive generations of erythroid cells (and exclusively in them), it forms a nuclease-hypersensitive site, indicating that regulatory proteins are bound to the enhancer in these cells. To identify these proteins, one must determine the precise nucleotide sequence required for enhancer activity. To this end, mutant enhancer sequences were linked to a reporter gene. Because the product of such a reporter gene is easily assayed, The Effect of any enhancer mutation on Transcription can be readily evaluated: each recombinant construct was introduced into chicken erythroid cells, and the efficiency of reporter gene expression was measured (Fig. 10-20). Those nucleotides found to be essential for enhancer activity in this assay can be considered specific protein-binding sites. Using this methodology, it was established that three such proteins are involved (Fig. 10-21). Although the cellular concentration of each protein is very low, knowing their binding sites makes it possible to clone the corresponding DNA coding sequences and thereby produce these regulatory proteins in unlimited quantities (see Section 9.1.7).

The insights gained from studying enhancers and promoter-proximal elements have led to several key Conclusions.

1. Regulatory elements have a modular architecture consisting of a series of distinct nucleotide sequences. Each such motif contains 8–15 nucleotides and binds a corresponding set of regulatory proteins. Some of these proteins are restricted to specific cell types, whereas others are ubiquitous.

2. Certain regulatory proteins act to activate transcription upon binding, whereas another group functions as transcriptional repressors. The overall effect of a regulatory element depends on the combinatorial set of bound proteins, and for any given element, this effect may shift during cellular development. For instance, an enhancer is capable of both activating and repressing transcription (see Fig. 10-22B). Consequently, the original name given to these elements does not fully capture their versatile function.

3. Enhancers and promoter-proximal elements appear to share many of the same binding proteins, suggesting that both classes of control elements operate via common mechanistic principles to influence transcription.

4. Analysis of newly discovered vertebrate gene regulatory elements has revealed that many of the proteins that bind to them were previously characterized as regulators of other genes. This likely reflects the fact that transcription in higher eukaryotes is controlled by a relatively small repertoire of regulatory proteins (Table 10-1). Proteins bound to promoter-proximal elements cooperate with enhancer-bound proteins to execute their function. Their net effect on gene activity is the result of antagonistic activating and repressing influences (Fig. 10-22). It is believed that shifts in the balance between positive and negative regulatory proteins account for the varying levels of ß-globin gene transcription at different stages of chicken erythroid cell development (Fig. 10-23).

Fig. 10-21. Comparison of data on the activity of the mutant enhancer described in Fig. 10-20 with the locations of protein-binding sites in the normal enhancer. The colored bars in the figure correspond to Regions of the normal enhancer sequence that are protected by the indicated proteins. These data were obtained by mixing the normal enhancer sequence with chicken erythrocyte extracts, followed by Analysis of the mixture using Electrophoresis and footprinting assays. Activity is expressed as "percentage of enhancer activity": a value of 100 indicates that the mutant stimulates RNA synthesis to the same extent as the normal enhancer; 0 indicates that RNA synthesis does not exceed transcription in the absence of the enhancer. The results show that the largest contributions to The stimulation of transcription by the enhancer are made by the AP1-like, AP2-like, and Eguyl proteins; however, none of these proteins individually can provide full enhancer activity (Fig. 10-23).

Fig. 10-22. The transcriptional activity of a gene is determined by the combined action of regulatory proteins bound to the upstream promoter element and proteins bound to the enhancer. A. Typical arrangement of the Two Types of regulatory elements relative to the coding sequence. B. Examples of the cooperative action of proteins bound to these two elements. Depending on which proteins have bound, each individual element in this example exhibits either a weak positive (+), strong positive (+ +), weak negative (—), or strong negative (— —) effect, as shown in the figure.

Table 10-1. Some well-characterized mammalian regulatory proteins and their corresponding specific DNA-binding sites.

Protein name

Protein mass, kDa1)

DNA-binding site2)

Notes

jun

(or AP1)

36

TGANTCA ACTNAGT

Proto-oncogene product that can bind to the fos protein, the product of another proto-oncogene. Activity is induced by protein kinase C; related proteins encoded by multiple genes

AP2

48

CCCCAGGC GGGGTCCG

Activity is induced by protein kinase C

ATF (or 43 CREB)

(TorG)(TorA)CGTCA (A C)(A T)GCAGT

Activity is induced by protein kinase A

SP1

85

GGGCGG CCCGCC

Binds to promoters of many housekeeping genes

OTF1

90

ATTTGCAT TAAACGTA

Found in all mammalian cells. One of many proteins that bind to the same octamer sequence

NF1 (or CTF)

52-66

GCCAAT CGGTTA

Multiple proteins encoded by alternatively spliced RNA transcribed from a single gene

SRF

67

GATGCCCATA

CTACGGGTAT

Involved in the transient activation of proto-oncogene fos transcription (and other genes) by growth factors. Binds as a dimer to a symmetrical DNA sequence (only the monomer binding site is shown)

1) The DNA encoding each of these proteins has been cloned, and the Amino Acid Sequence of the corresponding protein has been determined.

2) In many cases, only a portion of the recognized DNA sequence is known. N denotes any nucleotide.

Fig. 10-23. Known regulatory proteins controlling chicken ß-globin gene expression during normal erythrocyte development. This summarizes the data from the experiments described in Figs. 10-20 and 10-21, analogous experiments performed with mutant upstream promoter elements, and the results of Other types of analyses. A. Binding sites of regulatory proteins and their effects. Note that the enhancer for this gene is located downstream of the coding sequence. Proteins not marked with + (activation) or — (repression) apparently do not significantly affect transcription, as mutation of their binding sites does not alter the transcription level (see Fig. 10-21). B. Relative amounts of regulatory proteins at various developmental stages. The numbers 0, 1, and 2 indicate the absence, intermediate level, and high level of transcription, respectively, on a non-linear scale. Current data do not provide a precise explanation of why the gene is turned on between 4 and 9 days of development. (Courtesy of Gary Felsenfeld.)

10-12

10-13

10.2.9. Most Regulatory Proteins Contain Distinct Functional Domains [17]

The mechanism by which regulatory proteins act on genes is not yet fully understood in most cases. It is known only that binding to DNA alone is not sufficient to affect transcription. For example, in the gal4 protein, the DNA-binding and transcriptional activation functions are carried out by different domains. This activator protein turns on the transcription of many Yeast genes involved in The conversion of galactose to glucose when yeast cells are grown on a galactose-containing medium. The gal4 protein binds to a specific 17-base-pair sequence that acts as an enhancer in yeast (UAS, upstream activating sequence). The gal4 protein contains about 900 amino acid residues, whereas binding to this specific 17-base-pair sequence requires only a 73-amino-acid amino-terminal domain, which forms a series of "zinc fingers" (see Section 9.1.9). However, the DNA-binding domain by itself is not capable of activating transcription. Yet, in the absence of this domain, the remainder of the protein sequence is apparently also inactive. Domain-swap experiments have demonstrated that the DNA-binding and transcriptional activation functions reside in separate domains within this protein (Fig. 10-24).

Fig. 10-24. Description of an experiment designed to identify independent DNA-binding and transcription-activating domains within the yeast gal4 activator protein. A functional activator protein can be generated by fusing the carboxy-terminal portion of the gal4 protein with the DNA-binding domain of a bacterial regulatory protein (the lexA protein) using Gene Fusion techniques. The resulting bacterial-yeast hybrid will activate transcription of yeast genes if a specific binding site required for its recognition is inserted upstream of these genes. A. Normal activation of transcription by the gal4 protein. B. The chimeric regulatory protein requires a lexA-binding DNA site to exhibit its activity. Similar experiments have demonstrated the existence of such separate domains in other eukaryotic regulatory proteins as well.

Fig. 10-25. Evolutionarily related regulatory proteins belonging to the steroid hormone receptor family. The short DNA-binding domains in each receptor are highlighted in color. Domain-swap experiments indicate that many hormone-binding, transcription-activating, and DNA-binding domains in these receptors can act as interchangeable modules.

Similar results have been obtained for mammalian regulatory proteins. For example, the action of regulatory proteins belonging to the steroid hormone receptor family has been thoroughly studied. These receptor proteins mediate cellular responses to various lipid-soluble Hormones by activating or repressing the activity of specific genes. These receptor proteins contain a central DNA-binding domain of approximately 100 amino acid residues. As in the case of gal4, this domain features a series of "zinc fingers" and recognizes a specific DNA sequence. In some members of the family, the transcription-activating domain is located at the amino terminus. In addition, all receptors contain a hormone-binding domain at the carboxyl terminus of the protein (Fig. 10-25). Domain-swap experiments have demonstrated their interchangeability. For example, replacing the DNA-binding domain of the glucocorticoid receptor protein with the DNA-binding domain of the estrogen receptor causes

the modified glucocorticoid receptor to bind to an enhancer that normally activates estrogen-responsive genes. As a result, these genes begin to be transcribed in response to glucocorticoids rather than estrogens. Similarly, introducing the hormone-binding domain of the glucocorticoid receptor into another regulatory protein not belonging to this family results in the modified protein requiring glucocorticoids for its activity. Each of the domains in these receptor proteins is encoded by one or more separate exons. Analysis of the DNA sequences encoding various Proteins of the steroid hormone receptor family (Fig. 10-25) suggests that each protein arose through random chromosomal exchanges that brought together domains from different regulatory proteins (see Section 10.5.4).

10-5

10.2.10. Eukaryotic Regulatory Proteins can be Turned On and Off [18]

Like other members of the steroid hormone receptor protein family, the glucocorticoid and estrogen receptors resemble the bacterial CAP protein (see Section 3.3.6) in that they can be activated by the binding of small signal molecules to distinct ligand-binding domains. Other eukaryotic regulatory proteins are likely activated and inactivated through phosphorylation or via binding to other proteins that modulate their activity (see Table 10-1). These and other currently known mechanisms for controlling the activity of gene-regulatory proteins in Eukaryotic cells are illustrated in Fig. 10

26.

10.2.11. The Mechanism of Enhancer Action Remains Incompletely Understood [15, 19]

The mechanisms regulating eukaryotic gene transcription exhibit a high degree of conservation. For example, the yeast GAL4 protein stimulates transcription of mammalian genes in human cell culture if these genes contain the GAL4 UAS sequence. Similarly, if the human estrogen receptor protein is synthesized in yeast cells, the yeast genes associated with the human estrogen-responsive enhancer are turned on upon the addition of this hormone.

Such functional conservation is not mirrored by conservation of The amino acid sequence. Sequencing results for many genes determining regulatory proteins in yeast and humans indicate that the domains responsible for transcriptional activation in these species lack obvious similarity. In both yeast and human cells, a sufficient level of gene activation is also achieved when the transcription-activating domain of a regulatory protein is replaced by a region rich in acidic Amino Acids (glutamic acid and aspartic acid). This fact may suggest that the mechanism of activation is relatively simple.

Fig. 10-26. Various modes of regulatory protein activation in eukaryotic cells. Examples are known for each of these mechanisms. A. The regulatory protein is synthesized only when needed and is rapidly degraded by proteolysis, thus preventing its accumulation. B. Ligand binding. C. Phosphorylation or other covalent modification. D. Complex formation with a separate protein that acts as a transcription-activating domain.

Fig. 10-27. Two models of enhancer action in eukaryotic cells, based on the DNA looping observed in bacteria (see Fig. 10-17). In this example, the enhancer is located upstream of the coding sequence; however, a loop can also form if the enhancer is located downstream of the coding sequence, as in the case of the ß-globin gene (see Fig. 10-23).

One proposed mechanism of enhancer action is based on studies of bacterial systems. It is known that in bacterial cells, the initiation of transcription is facilitated by DNA looping. This is consistent with data showing that enhancers are typically most active when located close to the promoter; their activity gradually drops as the distance increases. Figure 10-27 illustrates two variants of enhancer action involving loop formation. Other hypotheses regarding the MECHANISM OF ACTION of this regulatory element have also been proposed. 1. An enhancer may act over a long distance by activating DNA topoisomerase, which introduces torsional stress into a large DNA loop using the energy of ATP Hydrolysis. 2. An enhancer may influence transcription by acting as a landing site for mobile proteins that bind to DNA and then move along its molecule. 3. An enhancer may bind proteins that facilitate the attachment of a nearby gene to a specific region of The Nucleus where transcription factors are localized.

It will be shown later that certain DNA sites exert a significant effect on transcription even when located 50,000 or more nucleotide pairs away from the promoter (see Section 10.3.12). The mechanism of action of such sites likely differs from that of enhancers.

Summary

In prokaryotes, regulatory proteins typically bind to specific DNA sequences near the transcription initiation site and either repress or activate the transcription of neighboring genes. Owing to The flexibility of DNA, the molecule can form loops, allowing proteins bound at some distance from the promoter to influence RNA polymerase. Action at a distance is widespread in eukaryotic cells, where gene expression is frequently controlled by enhancers located thousands of nucleotide pairs away from the promoter.

In many eukaryotic protein-coding genes transcribed by RNA polymerase II, a sequence known as the "TATA box" is located immediately upstream of the transcription initiation site.

Before RNA synthesis can be initiated, a crucial transcription factor (TATA-binding protein, TFIID) must form a stable transcription complex. The sequences flanking the TATA box form a core promoter element (an element located upstream of the promoter). Both the enhancer and the core promoter element contain a series of short nucleotide sequences that bind to corresponding regulatory proteins. These proteins interact with one another, and this interaction results in the activation or inactivation of genes. Recombinant DNA techniques have demonstrated that regulatory proteins frequently consist of multiple domains, each possessing its own function.



Last update: 12/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.