LEHNINGER PRINCIPLES OF BIOCHEMISTRY - VOL. 3. INFORMATION PATHWAYS - 2017

CHAPTER III. INFORMATION PATHWAYS

Class="center">RNA occurs in The Cell in The Nucleus, in cytoplasmic particles, and as a “soluble” RNA in the cell sap; and it has been shown by many workers that all these three fractions of RNA have different turnovers. In any Structure/133.html">Discussion of The Role of RNA in the cell It is important to bear in mind that RNA is extremely heterogeneous as to its METABOLISM, and presumably, therefore, belongs to more than one type.

— Francis Crick, Symposium of the Society for Experimental Biology, 1958

26. RNA METABOLISM

The expression of Genetic information is typically achieved through The production of an RNA molecule transcribed from a DNA template. Although RNA and DNA sequences may appear at first glance very similar, their chemical structures differ only in the replacement of thymine by uracil in RNA, and a 2'-hydroxyl group on the aldopentose. However, unlike DNA, most RNA molecules function as single-stranded polymers that fold into complex structures, making them far more structurally diverse than DNA (see Chapter 8, Vol. 1). As a result, RNA is uniquely adapted to perform A wide variety of cellular Functions.

RNA is the only known macromolecule that participates in both the storage and transmission of information, as well as in catalysis—a property that led to the hypothesis that it may have been the most crucial molecular precursor during the early evolution of life on our planet. The discovery of catalytic RNAs, or ribozymes, revolutionized the very concept of Enzymes, whose functions were previously thought to be exclusive to Proteins. Nevertheless, proteins remain essential for RNA function. In modern Cells, all Nucleic Acids, including RNA molecules, are associated with proteins. Some of these complexes are highly sophisticated, with the RNA component combining both structural and catalytic roles.

With the exception of the genomic RNAs of certain Viruses, all RNA molecules transfer information that is permanently stored in the form of DNA. During Transcription, an enzyme system converts the genetic information from a segment of double-stranded DNA into an RNA chain whose base sequence is complementary to one of the DNA strands. Three principal types of RNA are produced. Messenger RNA (mRNA) encodes the Amino Acid Sequence of one or more Polypeptides, as specified by a Gene or set of genes. Transfer RNA (tRNA) reads the information encoded in mRNA and delivers the corresponding amino acid to the growing polypeptide chain during Protein Synthesis. Ribosomal RNA (rRNA) is a Structural and functional component of Ribosomes—the complex cellular machinery that carries out protein synthesis. Many additional specialized RNA molecules perform regulatory or catalytic functions, or serve as precursors to the three primary types of RNA outlined above. These RNAs are no longer considered minor variants in the cellular RNA repertoire. Vertebrates possess a much broader array of RNA types than merely the "classic" mRNAs, tRNAs, or rRNAs.

While DNA Replication typically copies an entire chromosome, transcription is far more selective. Only specific genes or groups of Genes are transcribed at any given moment, and certain Regions of the DNA genome are never transcribed at all. The cell restricts the expression of genetic information so that only those gene products required at a particular time are synthesized. Specific regulatory sequences mark the beginning and end of transcribed DNA segments, as well as which of the two DNA duplex strands serves as the template. Transcription regulation is discussed in detail in Chapter 28.

The complete set of RNA molecules produced by a cell under specific conditions is referred to as the cell's transcriptome. Given that only a relatively small fraction of The Human Genome encodes proteins, one might expect that only a minor part of The Genome is transcribed. However, this is not the case. Modern microarray-based transcription analysis Methods have revealed that a significant portion of the human and mammalian genomes is actively transcribed into RNA, with the primary products being not mRNAs, tRNAs, or rRNAs, but a multitude of diverse RNAs with specialized functions. Many of these likely participate in the Introduction/30.html">Regulation of Gene Expression, though the Functions of the majority remain unknown.

This chapter examines the synthesis of RNA on a DNA template, followed by RNA Processing and turnover. We will discuss many of the Specialized Functions of RNA, including catalysis. Interestingly, the substrates for catalytic RNAs are frequently other RNA molecules. We will also explore systems in which an RNA molecule serves as the template and DNA as the product, rather than the reverse. Information pathways thus form a complete cycle, demonstrating that template-directed nucleic acid synthesis follows standard rules regardless of The Nature of the template or product (RNA or DNA). This mutual interconversion of DNA and RNA as information carriers naturally leads to the question of the evolutionary origin of biological information.

26.1. DNA-Dependent RNA Synthesis

We begin our discussion of RNA Synthesis by comparing transcription with DNA replication (Chapter 25). Transcription resembles replication in its fundamental chemical mechanism, its polarity (direction of synthesis), and its requirement for a template. Like replication, transcription proceeds through initiation, elongation, and termination stages, although in the literature describing transcription, initiation is often subdivided into two distinct phases—DNA binding and RNA chain initiation. Transcription differs from replication in that it does not require a primer and typically involves only discrete segments of the DNA molecule. Furthermore, only One DNA strand serves as the template for any given RNA transcript.

Figure 26-1. Reaction mechanism. Transcription in E. coli catalyzed by RNA polymerase. DNA is temporarily unwound to synthesize an RNA chain complementary to one of the two strands of The Double Helix; (a) at any given moment, a region of about 17 bp is maintained in the unwound state. RNA polymerase and the transcription bubble move from left to right along the DNA as RNA synthesis proceeds, as shown in the diagram. The DNA molecule is unwound ahead of the bubble and rewound behind it. Red arrows indicate the direction in which the DNA must rotate to accommodate this process. Following rewinding, the RNA–DNA hybrid dissociates, and the RNA chain is released. RNA polymerase maintains close contact with the DNA ahead of the transcription bubble, as well as with the separated DNA strands and the RNA chain within and immediately behind the bubble. Incoming nucleoside triphosphates (NTPs) enter the polymerase Active Site through a channel in the protein. During elongation, the polymerase covers a DNA segment of approximately 35 bp. (b) The Mechanism of RNA synthesis by RNA polymerase is analogous to that of DNA polymerase (see Fig. 25-5b). The reaction involves two Mg2+ ions coordinated by the phosphate groups of the incoming NTP and three conserved Asp residues (Asp460, Asp462, and Asp464 in the β' subunit of E. coli RNA polymerase), which are conserved across RNA polymerases from all species. One Mg2+ ion facilitates the attack of the 3'-hydroxyl on the α-phosphate of the NTP; the second Mg2+ assists in the departure of the pyrophosphate; and both temporarily stabilize the pentacoordinate Transition State. (c) Transcription induces changes in DNA Supercoiling. The movement of RNA polymerase along the DNA generates positive supercoils (overwound DNA) ahead of the transcription bubble and negative supercoils (underwound DNA) behind it. Topoisomerases rapidly remove positive supercoils and regulate negative supercoiling (Chapter 24).

RNA Is Synthesized by RNA Polymerase

The Discovery of DNA-dependent DNA polymerase spurred the search for an enzyme capable of synthesizing an RNA molecule complementary to a DNA strand. By 1960, four independent research groups had discovered an enzyme in cell extracts capable of forming an RNA polymer from ribonucleoside 5'-triphosphates. Subsequent studies on purified Escherichia coli RNA polymerase helped elucidate the fundamental principles of transcription (Fig. 26-1). DNA-dependent RNA polymerase requires, In addition to a DNA template, all four ribonucleoside 5'-triphosphates (ATP, GTP, UTP, and CTP) as nucleotide precursors, as well as Mg2+ ions. The enzyme also tightly binds a single Zn2+ ion. In chemical essence and mechanism, RNA synthesis strongly resembles DNA Synthesis by DNA polymerase (see Fig. 25-5). RNA polymerase elongates the RNA chain by adding ribonucleotide units to the 3'-hydroxyl end, extending the RNA in the 5' —> 3' direction. The 3'-terminal hydroxyl group acts as a nucleophile, attacking the α-phosphate of the incoming ribonucleoside triphosphate (Fig. 26-1b) and releasing pyrophosphate. The general reaction can be written as follows:

RNA polymerase requires DNA and exhibits maximal activity when bound to double-stranded DNA. As noted above, only one DNA strand acts as the template. The template DNA strand is copied in the 3' —> 5' direction (opposite to the direction of new RNA chain growth), exactly as in DNA replication. Each nucleotide in the nascent RNA is incorporated according to Watson–Crick base-pairing rules: a U residue in the RNA pairs with an A residue in the DNA template, and a G residue pairs with a C. Base-pair geometry plays a crucial role in maintaining pairing fidelity (see Fig. 25-6).

Unlike DNA polymerase, RNA polymerase does not require a primer to initiate synthesis. Initiation occurs when RNA polymerase binds to specific DNA sequences known as promoters (described below). The 5'-triphosphate group of the first residue in the synthesized RNA molecule is not cleaved to release PPi, but is retained throughout the transcription process. During the elongation phase, bases at the growing end of the new RNA chain temporarily pair with bases in the DNA template, forming a short hybrid double-helix of RNA-DNA approximately 8 bp in length (Fig. 26-1a). Soon after this structure forms, the RNA dissociates from the duplex, and the DNA duplex reforms.

For RNA polymerase to synthesize an RNA chain complementary to one of the two DNA strands, the DNA duplex must unwind locally, creating a transcription "bubble." In E. coli cells, RNA polymerase typically maintains an unwound region of about 17 bp, of which 8 bp form the RNA-DNA hybrid. The elongation rate of the RNA transcript by E. coli RNA polymerase ranges from 50 to 90 NUCLEOTIDES per second. Because DNA exists as a double helix, progression of the bubble requires Rotation of the nucleic acid strands. In most DNA molecules, strand rotation is restricted by DNA-binding proteins and other structural barriers. As a result, the moving RNA polymerase generates a series of positive supercoils ahead of the transcription bubble and negative supercoils behind it (Fig. 26-1c). This phenomenon occurs both in vitro and in vivo (in Bacteria). Within the cell, the topological stress associated with transcription is relieved by the action of topoisomerases (Chapter 24).

Conventions.

During transcription, the two complementary DNA strands serve distinct roles. The strand that acts as the template for RNA synthesis is called the template strand. The complementary nontemplate, or coding, strand has a nucleotide sequence identical to that of the RNA transcript produced from the gene, except that the RNA contains U instead of T (Fig. 26-2). For any given gene, either strand of a given chromosome may serve as the coding strand (as illustrated in Fig. 26-3 for a virus). By convention, the regulatory sequences that control transcription (described later in this chapter) are mapped and designated on the coding strand. ■

Figure 26-2. Template and nontemplate (coding) DNA strands. The two complementary DNA strands perform different functions during transcription. The RNA transcript synthesized on the template strand is identical in sequence to the nontemplate, or coding, strand (except for the substitution of T with U).

Fig. 26-3. Information coding in the adenovirus genome. The genetic information of the adenovirus genome (the simplest example) is encoded in a double-stranded DNA molecule 36,000 bp in size, with both strands encoding proteins. Information for most proteins is encoded in the upper strand; by convention, the strand is oriented from left to right, from the 5' end to the 3' end. For these transcripts, the lower strand serves as the template. However, certain proteins are encoded by the lower DNA strand, which is transcribed in the opposite direction (using the upper strand as a template). In reality, the synthesis of adenovirus mRNA molecules is far more complex. Many mRNA molecules shown on the upper strand are initially synthesized as a single long transcript (~25,000 nucleotides), which is subsequently processed into individual mRNA molecules. Adenoviruses cause Upper Respiratory Tract infections in certain vertebrates.

DNA-dependent RNA polymerase from E. coli is a large, complex enzyme consisting of five subunits, α2ββ'ω (termed the core enzyme; Mr = 390,000), and a sixth subunit, σ, which exists in various size (molecular weight) variants. The σ subunit associates temporarily with the core enzyme and directs it to a specific site on the DNA molecule (shown below). Together, these six subunits constitute the RNA polymerase holoenzyme (Fig. 26-4). Thus, the E. coli RNA polymerase holoenzyme exists in several forms depending on the type of σ subunit. Most commonly, this is the σ70 subunit (Mr = 70,000), and we will focus on this particular variant of the RNA polymerase holoenzyme in the subsequent discussion.

Fig. 26-4. STRUCTURE OF THE RNA polymerase holoenzyme from the bacterium Thermus aquaticus (based on PDB ID 1IW7). The overall architecture of the enzyme strongly resembles that of E. coli RNA polymerase (neither RNA nor DNA is shown here). The β subunit is highlighted in dark gray, the β' subunit in light gray; the two α subunits are shown in different shades of red; the ω subunit is yellow; the σ subunit is beige. The left view is oriented identically to Fig. 26-6. Rotating the structure by 180° about the y-axis (right) reveals the small ω subunit.

RNA polymerases lack an independent 3' → 5' proofreading exonuclease activity (unlike many DNA polymerases), and the transcription error rate is higher than that of chromosomal DNA replication—approximately one error per 104–105 incorporated ribonucleotides. Because multiple RNA copies are typically transcribed from a single gene, and all RNA molecules are eventually degraded and turned over, an error in an RNA molecule does not have the severe consequences for the cell that errors in DNA—the permanent information repository—would entail. Many RNA polymerases, including bacterial RNA polymerase and eukaryotic RNA polymerase II (discussed below), stall if an incorrect base is incorporated during transcription and can remove aberrant nucleotides from the 3' end of the transcript via a reaction that is the reverse of the polymerization reaction. However, it remains unclear whether this activity serves as a true proofreading mechanism or what its precise contribution to transcription fidelity is.

RNA synthesis begins at promoters

Initiating RNA synthesis at random points along the DNA molecule would be an unacceptable waste of resources. Therefore, RNA polymerase binds to specific DNA sequences called promoters, which direct transcription of adjacent DNA regions (genes). The DNA sequences recognized by RNA polymerases vary considerably, and numerous studies have been devoted to identifying the sequence elements that dictate promoter function.

Binding of E. coli RNA polymerase occurs within a region extending from approximately 70 bp upstream to 30 bp downstream of the transcription start site. By convention, the DNA Base Pairs corresponding to THE START OF the RNA molecule are numbered positively, while those upstream of the transcription start site are numbered negatively. Thus, the promoter region is located between positions -70 and +30. Comparative Analysis of the most common class of bacterial promoters (recognized by the holoenzyme containing the σ70 subunit) has revealed characteristic sequences centered around the -10 and -35 positions (Fig. 26-5). These sequences are crucial for interaction with the σ70 subunit. Although these sequences are not identical across all bacterial promoters of this class, certain nucleotides occur most frequently at specific positions, defining a consensus sequence (recall the E. coli oriC sequence; see Fig. 25-11). The consensus sequence at the -10 region is (5') TATAAT (3'); the consensus sequence at the -35 region is (5') TTGACA (3'). A third, AT-rich segment known as the UP element (upstream promoter) is located between positions -40 and -60 in the promoters of certain highly expressed genes. The UP element binds to the α subunit of RNA polymerase. The efficiency with which RNA polymerase binds to a promoter and initiates transcription is largely determined by these sequences, the spacing between them, and their distance from the transcription start site.

Fig. 26-5. Typical E. coli promoters recognized by the σ70-containing RNA polymerase holoenzyme. Sequences of the non-template strand are shown in the standard convention (5' → 3' direction). While promoter sequences vary, comparisons reveal notable similarities, particularly at the -10 and -35 positions. The UP element, absent in some E. coli promoters, is found in the P1 promoter of the highly expressed rrnB rRNA gene. UP elements, typically located between -40 and -60, strongly stimulate transcription from promoters containing them. In the rrnB P1 promoter, the UP element spans positions -38 to -59. The consensus sequence for σ70-recognized E. coli promoters is shown second from the top. Spacer regions contain a somewhat variable number of nucleotides (N). Only the first nucleotide of the transcript-encoding sequence (position +1) is indicated.

Box 26-1. BIOCHEMISTRY IN ACTION. RNA Polymerase Leaves Its Footprint on the Promoter

Footprinting is a technique based on principles adapted from DNA Sequencing that allows the identification of DNA sequences bound by a specific protein. A DNA fragment presumed to contain sequences recognized by a DNA-binding protein is isolated, and one end of a single strand is radiolabeled (Fig. 1). Chemical Reagents or enzymes are then used to introduce random breaks into the DNA fragment (averaging about one break per molecule). High-resolution Electrophoresis of the labeled Cleavage products (fragments of varying lengths) generates a "ladder" of radioactive bands. In a separate tube, the same cleavage Procedure is repeated on the DNA fragment in the presence of the DNA-binding protein. The two sets of products are then subjected to electrophoresis in adjacent lanes and compared. A gap (the "footprint") in the ladder of radioactive bands in the protein-bound DNA sample corresponds to the region protected by the DNA-binding protein, thereby revealing the binding site sequences.

Fig. 1. The footprinting technique for identifying a polymerase binding site on a DNA fragment. Separation is performed in the presence (+) and absence (-) of polymerase.

The exact position of the protein-binding site can be determined by direct sequencing (Fig. 8-34, Vol. 1) of the same DNA fragment, with the products resolved on the same gel (not shown). Figure 2 presents the results of RNA polymerase binding to a DNA fragment containing a promoter. The polymerase covers a region 60 to 80 bp in length, which protects both the -10 and -35 regions, among others.

Fig. 2. Mapping the RNA polymerase binding site on the lac promoter (see Fig. 26-5). In this experiment, the 5' end of the non-template strand was radiolabeled. Lane K represents a control in which the labeled DNA was cleaved with a chemical reagent, producing a uniform distribution of bands.

The Functional Significance of the sequences at the -10 and -35 positions is supported by a wealth of independent evidence. Mutations that

affect promoter activity frequently alter bases within these regions. Changes to the consensus sequence also affect RNA polymerase binding efficiency and Transcription initiation. Substituting a single base pair can reduce binding affinity by several orders of magnitude. Thus, the promoter sequence establishes a basal level of expression that can vary dramatically among different E. coli genes. METHODS FOR STUDYING the interaction between RNA polymerase and promoters are discussed in Box 26-1.

The mechanism of transcription initiation and the role of the σ subunit are becoming increasingly clear (Fig. 26-6a). The initiation process consists of two major phases—binding and initiation—each comprising multiple sequential steps. First, guided by the σ factor, the polymerase binds to the promoter, leading to the stepwise formation of a closed complex (in which the bound DNA remains intact) and an open complex (in which the bound DNA is partially unwound around the -10 position). Transcription then initiates within the complex, triggering a conformational shift into an elongation-competent state, followed by promoter clearance (promoter escape). Each of these stages depends on specific promoter sequences. As the polymerase transitions to the elongation phase, the σ subunit dissociates spontaneously. The participation of RNA polymerase in elongation is schematized in Fig. 26-6b. The NusA protein (Mr = 54,430) binds to RNA polymerase, competing with the σ subunit. Upon completion of transcription, NusA dissociates from the enzyme, RNA polymerase detaches from the DNA template, and the σ factor (σ70 or another variant) can rebind to the enzyme to re-initiate transcription—a cycle often referred to as the σ cycle (Fig. 26-7).

Fig. 26-6. Transcription initiation and elongation by E. coli RNA polymerase. (a) Initiation typically involves two main stages: binding and initiation. During binding, initial interactions between RNA polymerase and the promoter yield a closed complex, in which the promoter DNA is tightly bound but not unwound. Subsequently, a 12–15 bp DNA segment spanning from -10 to +2 or +3 is unwound to form an open complex. Additional intermediates and protein conformational changes (not shown) precede The formation of closed and open complexes. Initiation of transcription culminates in promoter clearance. After Synthesis of the first 8 or 9 nucleotides of the nascent RNA, the σ subunit is released, and the polymerase escapes the promoter to carry out RNA elongation. (b) Structure of the E. coli RNA polymerase core enzyme during elongation. Subunit color coding matches Fig. 26-4: β and β' subunits are dark and light gray, respectively; α subunits are red; the ω subunit is not visible from this angle. The σ subunit is absent, having dissociated following initiation. The top panel shows the complete complex with associated DNA and RNA fragments. The catalytic center for transcription lies within the cleft between the β and β' subunits. In the middle panel, the β subunit is omitted to reveal the active site and the DNA–RNA hybrid region. A Mg2+ ion is positioned in the active site (red sphere). In the bottom panel, all protein components are removed to illustrate how the DNA and RNA strands thread through the complex.

Fig. 26-7. The role of σ factors in transcription. Guided by an associated σ subunit, RNA polymerase binds to DNA at the promoter region. Once RNA synthesis is initiated, the σ subunit dissociates and is replaced by the NusA protein. When RNA polymerase reaches a termination sequence, RNA synthesis ceases, NusA is released from the polymerase complex, and RNA polymerase detaches from the DNA. The free core enzyme can then bind any available σ subunit. The type of σ subunit bound dictates which promoter the RNA polymerase will target in the subsequent round of synthesis.

In the DNA of E. coli, there are other classes of promoters that bind RNA polymerase holoenzymes containing alternative σ-subunits (Table 26-1), such as heat-Shock promoters. The products of these genes are produced at much higher concentrations when the cell undergoes severe stress, such as a sudden Temperature increase. RNA polymerase binds to the promoters of these genes only when the σ70 subunit is replaced by the σ32 subunit (Mr = 32,000), which is specific to heat-shock protein promoters (see Fig. 28-3). By utilizing various σ-subunits, the cell can coordinate Gene Expression, allowing it to significantly adapt its physiological state. The specific set of genes expressed depends on the availability of different σ-subunits, which in turn is determined by several factors: the regulated rates of Synthesis and degradation, post-synthetic modifications that switch individual σ-subunits between active and inactive forms, and a specialized class of anti-σ proteins that bind to specific types of σ-subunits, rendering them unavailable for transcription initiation.

Table 26-1. Seven types of σ subunits in Escherichia coli

σ Subunit

Kd (nmol/L)

Number of molecules per cell*

Fraction of this holoenzyme (%)*

Function

σ70

0.26

700

78

Housekeeping genes

σ54

0.30

110

8

Regulation of cellular nitrogen levels

σ38

4.26

<1

0

Stationary-phase genes

σ32

1.24

<10

0

Heat-shock genes

σ28

0.74

370

14

Flagellar and chemotaxis genes

σ24

2.43

<10

0

Extracytoplasmic functions; Participation in the heat-shock response

σ18

1.73

<1

0

Extracytoplasmic functions, including iron-citrate transport

* Approximate number of each type of σ-subunit per cell, as well as the fraction of RNA polymerase holoenzyme complexed with that σ-subunit during the exponential growth phase. These values depend on growth conditions. The fraction of RNA polymerase molecules complexed with each σ-subunit reflects both the Abundance of that subunit type and its affinity for the enzyme.

Transcription is regulated at multiple levels

The cellular demand for the products of any given gene varies depending on the cell's physiological state and developmental stage, and the transcription of each gene is tightly regulated to ensure that the specific product is synthesized in the exact required amount. Regulation can occur at any stage of transcription, including elongation and termination. However, the stages of polymerase binding and transcription initiation, shown in Figure 26-6a, are most frequently subjected to regulation. One example of

multilevel control is the presence of promoters with diverse sequences.

The Binding of Proteins to sequences near and far from the promoter also influences the level of gene expression. Protein binding can activate transcription by facilitating RNA polymerase binding, or repress it by blocking polymerase activity. In E. coli, one such transcription-activating protein is the cAMP receptor protein (CRP), which, in the absence of glucose in the growth medium, enhances the transcription of genes encoding Enzymes for the metabolism of sugars other than glucose. Repressor proteins block the synthesis of RNA from specific genes. For example, the Lac repressor (Chapter 28) blocks the transcription of genes encoding lactose-metabolizing enzymes when lactose is unavailable.

Transcription is the initial stage in the complex and energetically costly process of protein synthesis; consequently, the Regulation of Protein concentration in both bacteria and eukaryotes is frequently exerted at the transcriptional level, particularly during its early stages. Chapter 28 explores the various mechanisms through which this regulation is achieved.

Specific sequences signal the cessation of RNA synthesis

RNA synthesis proceeds with high processivity (see p. 50), because if RNA polymerase were to prematurely release the RNA transcript, it would fail to complete synthesis and be forced to start over. Nevertheless, specific DNA sequences cause pausing and, occasionally, termination of RNA synthesis. Because the mechanism of termination in eukaryotes is not yet fully understood, we will focus on bacteria. E. coli cells utilize at least Two Types of termination signals: one involving the protein factor ρ (rho), and the other ρ-independent.

Most ρ-independent terminators share two key features. First, they contain a sequence whose transcript features complementary regions capable of forming hairpin structures (see Fig. 8-19, vol. 1), located 15–20 nucleotides upstream from the 3′ end of the RNA chain. Second, their template strand contains a conserved tract of three A residues, which are transcribed into U residues near the 3′ end of the hairpin. When the polymerase reaches the termination site containing this structure, it pauses (Fig. 26-8). The Formation of the hairpin within the RNA disrupts several A=U base pairs in the RNA-DNA hybrid and destabilizes critical contacts between the RNA and RNA polymerase, thereby facilitating transcript dissociation.

Fig. 26-8. Model of ρ-independent transcription termination in E. coli cells. RNA polymerase pauses at various sites along the DNA sequence, including terminators. This can lead to one of two outcomes: either the polymerase bypasses the obstacle and continues movement, or the complex undergoes conformational changes and isomerizes. In the latter case, the pairing of intramolecular complementary segments within the newly synthesized transcript can lead to the formation of a hairpin that disrupts the RNA-DNA hybrid and/or alters the interaction between the RNA and the polymerase, inducing isomerization. The AUU hybrid segment at the 3′ end of the nascent transcript is unstable, causing the RNA to completely dissociate from the complex and resulting in transcription termination. This is the typical sequence of events when RNA polymerase encounters terminators. Following isomerization, the complex may resume RNA synthesis at other sites.

ρ-Dependent terminators lack repetitive A residues in the template strand, but typically feature a C-rich sequence known as the rut element (derived from rho utilization). The ρ protein binds to a specific region of the RNA and translocates in the 5′ → 3′ direction until it catches up with the stalled transcription complex at the termination site, where it promotes the release of the RNA transcript. The ρ protein possesses ATP-dependent RNA-DNA helicase activity that drives its movement along the RNA; during termination, the ρ protein hydrolyzes ATP. The detailed mechanism by which this protein facilitates the release of the RNA transcript remains unknown.

Eukaryotic cells contain Three types of RNA polymerases

The mechanism of Transcription in Eukaryotic nuclei is considerably more complex than in bacteria. Eukaryotes possess three distinct RNA polymerases (I, II, and III), which differ in subunit composition yet share several common subunits. Each polymerase fulfills a distinct function and associates with a specific promoter sequence.

RNA polymerase I (Pol I) is responsible for the synthesis of only a single type of RNA—pre-ribosomal RNA (pre-rRNA), which contains the precursors for the 18S, 5.8S, and 28S rRNAs (see Fig. 26-25). Pol I promoter sequences vary substantially among different species. An essential function of RNA polymerase II (Pol II) is the synthesis of mRNA molecules and certain specialized RNA molecules. This enzyme recognizes thousands of promoters with highly diverse sequences. Many Pol II promoters share several common features, including a TATA box (the eukaryotic consensus sequence TATAAR) near position -30 and an initiator (Inr) sequence near the transcription start site at position +1 (Fig. 26-9).

Fig. 26-9. Common promoter sequences recognized by eukaryotic RNA polymerase II. The assembly point for Pol II preinitiation complex proteins is the TATA box. The DNA within the initiator sequence (Inr) is unwound, and the transcription start site typically lies within or immediately adjacent to this sequence. In the consensus Inr sequence shown here, N represents any nucleotide, and Y represents a pyrimidine nucleotide. Numerous additional sequences serve as binding sites for a vast array of proteins that modulate Pol II activity. These sequences are critical for The regulation of Pol II promoters and vary widely in type and number; as a rule, eukaryotic promoters are far more complex than depicted here (see Fig. 15-23, vol. 2). Many of these sequences are located several hundred base pairs upstream of the TATA box, while others may lie thousands of base pairs away. The consensus sequences of eukaryotic Pol II promoters shown here exhibit much greater divergence than those of E. coli promoters (see Fig. 26-5). Many Pol II promoters lack either the TATA box, the Inr element, or both sequences. One or more transcription factors recognize additional sequences surrounding the TATA box or downstream from it (to the right in this diagram).

RNA polymerase III (Pol III) synthesizes tRNA molecules, 5S rRNA, and several other small, specialized RNA molecules. The promoters recognized by Pol III are well characterized. Interestingly, some of the sequences required by polymerase III for regulated transcription initiation are located within the gene itself, whereas others are positioned more conventionally upstream of the transcription start site (Chapter 28).

RNA polymerase II activity requires additional protein factors

RNA polymerase II plays a central role in EUKARYOTIC GENE EXPRESSION and has therefore been intensely investigated. Although this polymerase is considerably more complex than its bacterial counterpart, this complexity masks a striking conservation of structure, function, and mechanism. Yeast Pol II is a giant enzyme composed of 12 subunits. Its largest subunit, Rpb1, exhibits significant Homology with the β′ subunit of bacterial RNA polymerase. Another subunit, Rpb2, is structurally similar to the bacterial β subunit, and two additional subunits, Rpb3 and Rpb11, show some structural homology with the two bacterial α subunits. Pol II must operate within vastly more complex genomes and interact with much more intricately packaged DNA molecules than bacterial polymerases. The added complexity of the eukaryotic polymerase stems from the necessity to navigate a labyrinth of numerous protein factors and engage in intricate Protein-Protein Interactions.

The largest subunit of Pol II possesses a distinctive feature: its C-terminus contains a long amino acid sequence consisting of multiple repeats of the heptapeptide consensus sequence -YSPTSPS-. The yeast enzyme contains 27 such repeats (18 of which perfectly match the consensus sequence), whereas the mouse and human enzymes each contain 52 repeats (21 exact matches). This carboxyl-terminal domain (CTD) is separated from the main body of the enzyme by an unstructured linker sequence. As discussed below, the CTD is essential for the multifaceted functions of Pol II.

Table 26-2. Proteins Required for Initiation of Transcription from Eukaryotic RNA Polymerase II (Pol II) Promoters

Protein

Number of subunits

Subunit Mr

Functions

Pol II initiation

12

10 000-220 000

Catalyzes RNA synthesis

TBP (TATA-binding protein)

1

38 000

Recognizes the TATA box

TFIIA

3

12 000.19 000, 35 000

Stabilizes binding of TFIIB and TBP to the promoter

TFIIB

1

35 000

Binds to TBP; assembles the Pol II-TFIIF complex

TFIIE

2

34 000, 57 000

Recruits TFIIH; possesses ATPase and helicase activities

TFIIF

2

30 000, 74 000

Binds tightly to Pol II; interacts with TFIIB and prevents nonspecific binding of Pol II to DNA

TFIIH

12

35 000-89 000

Unwinds DNA at the promoter (helicase activity); phosphorylates Pol II (in the CTD); recruits nucleotide Excision Repair proteins

Elongationa

ELLб

1

80 000


pTEFb

2

43 000, 124 000

Phosphorylates Pol II (in the CTD)

SII (TFIIS)

1

38 000


Elongin (SIII)

3

15 000, 18 000, 110 000


a The function of all elongation factors is to overcome pausing or premature termination of transcription by the Pol II-TFIIF complex.

б Short for eleven-nineteen Lysine-rich leukemia. The ELL gene is frequently involved in chromosomal translocations in acute myeloid leukemia.

To form an active transcription complex, RNA polymerase II requires a set of auxiliary proteins called transcription factors. The general transcription factors for each Pol II promoter (typically designated TFII with additional identifiers) are remarkably conserved across all eukaryotes (Table 26-2). The transcription process catalyzed by Pol II can be divided into several stages: assembly, initiation, elongation, and termination, with specific proteins participating at each stage (Fig. 26-10). The step-by-step process described below leads to active transcription in vitro. Within the cell, many of these proteins likely exist as larger, preassembled complexes that facilitate assembly at promoters. The factors involved in transcription are listed in Figure 26-10 and Table 26-2.

Figure 26-10. Transcription from RNA polymerase II promoters. (a) Sequential recruitment of TBP (often alongside TFIIA), TFIIB, TFIIF, followed by Pol II, TFIIE, and TFIIH leads to the formation of the closed complex. DNA unwinding within the Inr region occurs inside the complex, driven by the activities of TFIIH and possibly TFIIE, resulting in open complex formation. The C-terminal domain of the largest Pol II subunit is phosphorylated by TFIIH, after which the polymerase escapes the promoter and initiates transcription. Elongation is accompanied by the release of many transcription factors and is stimulated by elongation factors (see Table 26-2). Following termination, Pol II is released, dephosphorylated, and made available for a new round of synthesis. (b) Human TBP (gray) bound to DNA (dark and light blue) (PDB ID 1TGH). (c) Schematic model of transcription elongation catalyzed by the RNA polymerase II core enzyme.

Assembly of RNA Polymerase and Transcription Factors at the Promoter

Formation of the closed complex begins with the binding of the TATA-binding protein (TBP) to the TATA box (Fig. 26-10b). TBP in turn recruits the transcription factor TFIIB, which also contacts the DNA on either side of TBP. Although TFIIA binding is not strictly essential, it can stabilize the TFIIB-TBP complex on the DNA, which is particularly important at nonconsensus promoters where TBP-DNA binding is relatively weak. The TFIIB-TBP complex then joins with a preformed complex of TFIIF and Pol II. The TFIIF factor assists in the precise positioning of Pol II at the promoter, both through direct interactions with TFIIB and by dampening the polymerase's affinity for nonspecific DNA sequences. Finally, TFIIE and TFIIH join the assembly to complete the closed complex. TFIIH possesses DNA helicase activity and initiates unwinding of the DNA near the transcription start site (a process requiring ATP Hydrolysis), thereby generating the open complex. Counting all the subunits of the various factors (excluding TFIIA), this minimal active complex comprises 30 or more polypeptides. Structural studies conducted by Roger Kornberg and his colleagues have provided a detailed view of the RNA polymerase II core enzyme during elongation (Fig. 26-10c).

Initiation of the RNA Chain and Promoter Clearance

During initiation, TFIIH fulfills an additional role. The kinase activity of one of its subunits phosphorylates the CTD heptad repeats of Pol II at multiple sites (Fig. 26-10a). Several other protein Kinases, including CDK9 (cyclin-dependent kinase 9), which is a component of the pTEFb complex (positive transcription elongation factor b), also phosphorylate the CTD, particularly at Serine residues. This triggers conformational changes throughout the complex and commits the enzyme to transcription. CTD phosphorylation is also crucial for the subsequent elongation stage, with the pattern of phosphorylation changing as transcription proceeds. These dynamic modifications alter the interaction network between the transcription complex and various auxiliary enzymes, meaning that different proteins are associated with the complex during initiation versus later stages. Some of these associated proteins participate in transcript processing, as described below.

As the first 60 to 70 nucleotides are synthesized, TFIIE dissociates first, followed by TFIIH, allowing Pol II to transition into the elongation phase.

Elongation, Termination, and Release

During elongation, TFIIF remains tightly associated with Pol II. At this stage, the processivity of the polymerase is significantly enhanced by specific proteins known as elongation factors (Table 26-2). These elongation factors, some of which associate with the phosphorylated CTD, prevent premature transcriptional arrest and coordinate interactions with Protein Complexes involved in the post-transcriptional processing of mRNA molecules. Once synthesis of the RNA transcript is complete, transcription halts, and RNA polymerase II is dephosphorylated, rendering it ready to initiate another transcription cycle (Fig. 26-10a).

Regulation of RNA Polymerase II Activity

Regulation of Pol II transcription is highly intricate, involving the interplay of a diverse array of proteins with the preinitiation complex. Some of these regulatory proteins interact directly with transcription factors, while others target Pol II itself. Transcriptional Regulation is discussed in greater detail in Chapter 28.

The Multifunctional Role of TFIIH

In eukaryotes, Repair of Damaged DNA (see Table 25-5) is much more efficient in actively transcribed genes than in transcriptionally inactive regions, and the template strand is repaired more rapidly than the non-template strand. These striking observations are explained by the fact that TFIIH subunits serve dual roles. In addition to participating in closed complex formation during transcription preinitiation complex assembly (as described above), several TFIIH subunits also function as essential Components of the independent nucleotide excision repair complex (see Fig. 25-26).

When Pol II transcription stalls due to DNA damage, TFIIH can be recruited to the lesion site to assemble the entire excision repair complex. In humans, Genetic Defects in certain TFIIH subunits can cause severe disorders, such as xeroderma pigmentosum (see Box 25-1) and Cockayne syndrome, which are characterized by growth failure, extreme photosensitivity, and neurological abnormalities. ■

Selective Inhibition of DNA-Dependent RNA Polymerases

RNA chain elongation by both bacterial and eukaryotic RNA polymerases is inhibited by the antibiotic actinomycin D (Fig. 26-11).

Figure 26-11. The transcription inhibitors actinomycin D and acridine. (a) The boxed portion of the actinomycin D molecule is planar and intercalates between adjacent G=C base pairs in double-stranded DNA. The two cyclic peptide rings of actinomycin D bind within the minor groove of the double helix. Sar (Sarcosine) is N-methylglycine; meVal is methylvaline. Acridine also intercalates into DNA. (b) Complex of actinomycin D with DNA (PDB ID 1DSC). The DNA backbone is shown in blue, the bases in gray, the intercalating phenoxazone ring of actinomycin (boxed in part a) in orange, and the peptide rings in red. The DNA is significantly distorted (bent) upon actinomycin binding.

The planar ring system of this molecule intercalates into the DNA double helix between adjacent G=C base pairs, inducing a local structural distortion. This physical roadblock prevents the polymerase from translocating along the template. Because actinomycin D inhibits RNA elongation in intact cells as well as in cell-free extracts, it is widely used as an experimental tool to identify cellular processes dependent on RNA synthesis. Acridine inhibits RNA synthesis in a similarly disruptive manner (Fig. 26-11).

Rifampicin suppresses bacterial RNA synthesis by binding tightly to the β subunit of bacterial RNA polymerases, thereby blocking promoter clearance during transcription initiation (Fig. 26-6). It is frequently deployed clinically as an antibiotic.

The death cap mushroom (Amanita phalloides) features a highly efficient defense mechanism against animals. It synthesizes α-amanitin, which halts mRNA production in animal cells by blocking Pol II, and at high concentrations, Pol III as well. Neither Pol I nor bacterial RNA polymerase is sensitive to α-amanitin, but most importantly, the RNA polymerase II of the fungus A. phalloides itself is completely immune to it!

Summary of Section 26.1 DNA-Dependent RNA Synthesis

■ Transcription is catalyzed by DNA-dependent RNA polymerases, which utilize ribonucleoside-5'-triphosphates to synthesize RNA molecules complementary to the template strand of the DNA duplex. Transcription proceeds in several stages: binding of RNA polymerase to the promoter region of the DNA, initiation of transcript synthesis, elongation, and termination.

■ Bacterial RNA polymerase requires a specialized subunit to recognize the promoter. The binding of RNA polymerase to the promoter and the initiation of transcription are tightly coupled, forming The First stage of transcription. Transcription is halted at specific DNA sequences known as terminators.

■ Eukaryotic cells contain three distinct types of RNA polymerases. The binding of RNA polymerase II to its promoters requires protein transcription factors, while elongation factors take part in the elongation phase. The lengthy C-terminal domain of the largest Pol II subunit is phosphorylated during the initiation and elongation stages.



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.