BIOCHEMISTRY AND MOLECULAR BIOLOGY - W. ELLIOTT - 2002

CHAPTER 21. GENE TRANSCRIPTION — THE FIRST STEP IN THE REGULATION OF PROTEIN BIOSYNTHESIS

The information for Cell/13.html">Protein Structure is encoded in the DNA molecule as a sequence of 4 nitrogenous bases. A sequence of 3 bases encodes a single amino acid. With 4 nitrogenous bases, $4 \cdot 4 \cdot 4$ triplet combinations are possible, so there is no limitation on coding capacity. The Gene itself does not directly participate in METABOLISM/35.html">Protein Biosynthesis. In eukaryotes, DNA is enclosed within the nuclear membrane, whereas the protein-synthesizing machinery is located in the Cytoplasm. How then does the gene direct Protein Synthesis? It dispatches copies carrying the encoded information from The Nucleus to the cytoplasm (for example, in E. coli, such a copy enters the cytoplasm directly). Since the copied information is represented as a sequence of nitrogenous bases, the copy must also be a nucleic acid. It is known as ribonucleic acid (RNA), mentioned earlier; considering its function, it is termed Messenger RNA (or template RNA) (abbreviated as mRNA).

Messenger RNA

Introduction/21.html">RNA Structure

RNA is a polynucleotide similar to DNA, but with distinct features of its own.

✵ The sugar in RNA is ribose, which has an OH group at the 2' position, rather than deoxyribose.

Class="center">

✵ mRNA is a single-stranded molecule rather than a duplex of two molecules, from which it can be concluded that mRNA is a copy of only one of the two DNA strands of the gene.

✵ RNA contains 4 nitrogenous bases: A, C, G, and U. T is absent. Both U and T pair with A. T can be viewed as U with a methyl group attached (an explanation of its possible evolutionary origin is given on p. 256).

Were it not for these differences, The structure of single-stranded DNA (see p. 235) could well be mistaken for RNA: the same 3' —> 5' phosphodiester bonds between adjacent NUCLEOTIDES, no nucleotide substituent at the 5'-OH on the 5'-end, and a free 3'-OH group at the 3'-end. The hydroxyl group at the 2'-position makes the RNA molecule chemically more unstable compared to DNA (see p. 235). In a dilute alkali solution, RNA is degraded at room Temperature, whereas the DNA molecule remains stable.

How is mRNA synthesized?

The "building blocks" for mRNA synthesis are ATP, CTP, GTP, and UTP, which are synthesized in all Cells (see Chapter 18). In E. coli, mRNA synthesis is catalyzed by DNA-dependent RNA polymerase (abbreviated as RNA polymerase). Eukaryotic cells contain three such polymerases.

For RNA Synthesis catalyzed by RNA polymerase to occur—that is, for nucleotides to assemble in a specific sequence on a single template strand—the DNA strands must separate. In the region of RNA synthesis, the two DNA strands temporarily part, and reanneal after the polymerase has passed.

The process of RNA synthesis is largely similar to DNA Synthesis: the nitrogenous Base of the incoming ribonucleotide must be complementary to the base on the DNA template (Fig. 21.1).

Fig. 21.1. mRNA copying on the DNA template strand. The second DNA strand is not shown. Note that the Separation of the DNA strands is temporary

RNA polymerase, moving along the template, links nucleotides in the order determined by the DNA template. As with DNA, RNA synthesis is always

carried out in the 5' —> 3' direction. Nucleotides are added to the 3'-OH groups, and therefore chain elongation proceeds in the 5' —> 3' direction. The chemical reaction catalyzed by the polymerase involves The transfer of the $\alpha$-phosphoryl group of nucleoside triphosphates (directly attached to ribose) to the 3'-OH group of the preceding nucleotide, with the release of inorganic pyrophosphate. The latter is hydrolyzed to form two molecules of Pi, rendering this reaction (Fig. 21.2) strongly exergonic.

Fig. 21.2. Reaction catalyzed by RNA polymerase

RNA polymerase is capable of initiating new chains and does not require a primer. In the presence of a template, RNA polymerase can synthesize the entire mRNA molecule from start to finish. This represents another important distinction between Transcription and DNA synthesis, since DNA polymerase always requires a primer (see p. 248).

Some General Properties of mRNA

A typical chromosome contains thousands of different genes. Each mRNA molecule is encoded by a single gene (or a small group of genes in prokaryotes), so a multitude of different mRNAs are constantly produced in The Cell corresponding to the number of currently active genes. Compared to the long DNA molecule, the mRNA molecule is relatively small. In the cytoplasm, numerous mRNAs direct the Synthesis of the many Proteins encoded by their respective genes.

For a given living cell, the DNA molecule is immortal. The half-life of mRNA ranges from 20 minutes to several hours in eukaryotes, and is about two minutes in Bacteria.

Thus, Gene Expression (where "expression" simply means that the gene is active and synthesizing a protein) requires the continuous transcription of mRNA molecules. You could say that a gene constantly "stamps out" copy after copy, with RNA polymerase acting as its "printing press." As soon as a gene stops producing mRNA, this degradation acts as the main "off switch" for protein synthesis.

E. coli cells lack a nuclear membrane, meaning their genes are in direct contact with the cytoplasm. In eukaryotes, a single mRNA molecule always corresponds to a single gene, whereas in prokaryotes, one mRNA can carry information for several proteins at once. Such a molecule is called polycistronic (the genetic term "Cistron" refers to a gene controlling a single function). Clusters of genes transcribed into a single polycistronic mRNA ensure the coordinated synthesis of proteins that function as a single unit—such as the Enzymes of a single metabolic pathway. There are several other significant differences in mRNA production between pro- and eukaryotes, which we will discuss below.

Some Important Terminology

The flow of information during gene expression follows this direction:

The arrows indicate the direction of information transfer.

When DNA is copied into RNA, the information is rewritten (transcribed); therefore, mRNA production is called gene transcription, or simply transcription. The DNA itself is said to be transcribed. The resulting RNA molecules are called transcripts, while the synthesis of proteins based on mRNA is called Translation.

So far, we have discussed mRNA synthesis as the copying of DNA. The DNA template strand is "copied" strictly by complementary base pairing: through Watson-Crick pairing between incoming ribonucleotide bases and the bases of the template. For instance, wherever there is an A in the template, a U is incorporated into the copy, and so on. The DNA strand that serves as the template for mRNA synthesis is called the template strand, and the other is the non-template strand.

In addition, terms such as coding and non-coding strands, as well as sense and antisense strands, are widely used, though different authors often define them in various ways. While there is no universal, standardized terminology, It is important to understand what is meant when a particular term is used. In this book, we assume that the mRNA carries the Amino Acid Sequence of the protein, thereby conveying the gene's message to the translational machinery.

The sequence of nitrogenous bases in mRNA is identical to that of the non-template DNA strand (except for the substitution of T for U), which is why we will refer to the non-template strand as the coding or sense strand. Conversely, the template strand is defined here as the antisense or non-coding strand (Fig. 21.3). In Viruses, the template strand is often called the minus strand or (-)-strand, and the non-coding strand is called the plus strand or (+)-strand.

Fig. 21.3. Schematic diagram illustrating the terms "template" and "non-template" for the information-carrying DNA strand, using mRNA as an example. In viruses, the template strand is often referred to as the (-)-strand, and the non-coding strand as the (+)-strand

What Lies Ahead

So far, we have looked at gene expression by treating the gene simply as a DNA template. However, many other crucial and fascinating questions remain. For instance, how does a cell select which genes to express? How does RNA polymerase "recognize"

where to start copying the DNA within a transcribed gene and where to stop? How is The rate of gene expression regulated? Why might one protein encoded by its gene be produced in massive quantities, while another is made in tiny amounts or not at all? And how can certain genes be expressed only at specific times? To answer these questions, we need to examine the structures of actual genes, or more precisely, their base sequences.

Up to this point, we have focused on the biochemistry of animal cells. However, many metabolic processes in PROKARYOTES AND EUKARYOTES share profound similarities and are largely identical in fundamental principles. Yet when it comes to gene expression, the differences become so substantial that they require separate Discussion. Let us begin with E. coli.

Transcription in E. coli

What Defines a "Gene" in Prokaryotes?

In Chapter 19 (see p. 240), we noted that a gene is a segment of a large DNA molecule. Now we need to refine that definition.

Essentially, a gene Functions solely as a template for the Transcription of RNA Molecules whose base sequence corresponds to the non-template DNA strand. A gene is a segment of DNA transcribed into RNA. A large group of genes encodes ribosomal and Transfer RNAs.

Other genes are responsible for producing mRNA molecules, which in turn encode the Amino acid sequences of specific proteins. Both ends of an mRNA molecule contain regions that are not translated into protein. The 5' end features an untranslated region (UTR) containing encoded signals essential for initiating translation, while the 3' end harbors a sequence signaling Translation termination. Thus, besides the encoded instructions for protein structure, a gene also includes transcribed DNA regions (sites) that are not used to code for the protein itself.

This does not exhaust the list of DNA segments associated with a gene. The DNA region adjacent to the 5' end of the gene is called the promoter. It is vital for gene transcription, although it is not transcribed into RNA itself.

Conversely, the 3' end—known as the terminator region—is required to terminate transcription and is likewise not transcribed into RNA. The Cytology/cytology/92.html">SCHEMATIC STRUCTURE OF a gene is shown in Fig. 21.4. Transcription begins at the start point, which serves as the template for the first nucleotide of the mRNA. The template for the first nucleotide is designated as +1, and the adjacent nucleotide toward the 5' end as -1. The start point is indicated by an arrow pointing in the direction of transcription. Nucleotides located upstream toward the 5' end are referred to as being upstream, while those toward the 3' end are downstream of the start point. In Fig. 21.4, one end of the gene is labeled 5', which brings us to our next question.

Fig. 21.4. Structure of a prokaryotic gene and its mRNA. The 5' end of the gene corresponds to the non-template, or sense strand. In a neighboring gene, the opposite strand might serve as the non-template strand, with transcription proceeding from right to left

What do we mean when we speak of the 5' end of a gene?

The 5' end of a gene is the end containing the promoter. Since RNA synthesis always proceeds in the 5' —> 3' direction, the 5' end of a gene always corresponds to the 3' end of the template strand (see Fig. 21.3). It should be noted that in a physical sense there is no true end: the DNA chain continues into the next gene.

The 5' end of a gene is conventionally placed on the left in chromosome diagrams. A DNA strand that serves as the template for one gene may be the non-template (coding) strand for another. The choice of which strand acts as the template can switch from one gene to the opposite. If the complementary strand is used in this way for a second, adjacent gene, transcription proceeds in the opposite direction, as synthesis still proceeds in the 5' —> 3' direction.

Phases of gene transcription

There are three phases of gene transcription: initiation, elongation, and termination. Of these, initiation is the most complex.

Transcription initiation in E. coli

Promoters contain short sequences of nitrogenous bases known as boxes or elements. While these are standard terms, the base sequences are not literally "boxes" in any physical sense, nor do they have anything to do with atomic elements. A typical E. coli promoter features two boxes: the Pribnow box (named after its discoverer), located near the -10 region, and another box situated further upstream at the -35 region relative to the transcription start point.

These sequences are called consensus sequences because they occur frequently across various genes and exhibit low Variability (although instances lacking some of these sequences are known). Single-stranded sequences are always written in the 5' —> 3' direction, but a gene possesses a second DNA strand. Thus, although we state that the Pribnow box has the sequence TATAAT, the actual situation is as follows:

That is, it is a double-stranded structure recognized by the proteins that regulate transcription.

Proper transcription initiation is critically important because mRNA synthesis must begin at a precisely defined nucleotide on the template. How does RNA polymerase locate the correct site? The Pribnow box and the -35 element play a vital role in this process. E. coli DNA-dependent RNA polymerase is a large protein complex whose core is formed by 4 subunits. While possessing an affinity for any DNA segment to which it can randomly

bind, the enzyme cannot "on its own" recognize the initiation site. Upon associating with another cytoplasmic protein—the sigma protein, or sigma factor—the resulting holoenzyme loses its affinity for random DNA sequences, instead binding tightly to the -35 element and the Pribnow box (even though the DNA nitrogenous bases within the duplex are Watson-Crick base-paired, they can be recognized by proteins in the Major and minor grooves of the DNA; see Fig. 19.4). This properly orients the enzyme. RNA polymerase is now able to separate the DNA strands and initiate RNA synthesis. A "transcription bubble" (about 1.5 helical turns long) is formed at the site of DNA strand separation, making the nitrogenous bases of the template strand accessible for pairing with nucleoside triphosphates. From these, the enzyme forms the first few phosphodiester bonds, completing initiation. At this point, the sigma factor dissociates, and the polymerase moves along the gene, synthesizing mRNA at a rate of approximately 40 nucleotides per second. This continues until the enzyme reaches a terminator. What triggers the release of the sigma factor from the enzyme remains unknown.

DNA unwinding

Gene transcription requires the temporary separation of the DNA double helix. The Pribnow box, characterized by a high A and T content and consequently weaker Hydrogen Bonds, serves as the site for initial strand separation. DNA is unwound ahead of the polymerase and rewound behind it. Negative DNA Supercoiling facilitates this process (see p. 244). Thus, the transiently formed "bubble" travels along the gene together with the polymerase. Over a short region, the newly synthesized mRNA base-pairs with the template strand before detaching from it. Figure 21.6 illustrates the movement of RNA polymerase along the DNA double helix.

Fig. 21.6. Schematic representation of DNA Transcription by E. coli RNA polymerase. The polymerase unwinds a DNA segment approximately 17 Base Pairs long, forming a transcription bubble that moves along the DNA. DNA is unwound ahead of the polymerase and rewound behind it. The synthesized RNA forms a DNA-RNA hybrid double helix approximately 12 base pairs long with the DNA

Transcription termination

At the end of certain prokaryotic genes lie sequences that form hairpin structures in the transcribed RNA. Although the mRNA molecule is single-stranded, its thermodynamic properties permit intramolecular base pairing. The newly synthesized mRNA is attached to the DNA template strand via base pairing. Near the end of the gene, The base sequence in the transcribed mRNA promotes The formation of a hairpin structure (Fig. 21.7). The base sequence in this region favors the formation of stable G-C pairs, each stabilized by 3 hydrogen bonds. Immediately following the hairpin structure in the mRNA transcript are several U residues that hold the RNA and DNA together only weakly. This presumably facilitates mRNA release and transcription termination. The hairpin structure disrupts the binding of the mRNA to the template: once bases preferentially pair internally within the mRNA chain, they can no longer form pairs with the DNA template. The stretch of U residues facilitates the final detachment of the mRNA due to the weakness of A-U hydrogen bonds. Regardless of the exact mechanism, this hairpin structure followed by U residues leads to the termination of transcription.

Many prokaryotic genes employ a second mechanism of transcription termination. This involves the so-called Rho protein, or Rho factor, which is required to detach the mRNA from the DNA-RNA hybrid. The Rho factor attaches to the transcribed mRNA and moves along it behind the RNA polymerase. At the termination site, the polymerase stalls—possibly due to the difficulty of unwinding a G-C rich region of DNA—allowing the Rho factor to catch up to the polymerase. This protein possesses helicase activity specific for the DNA-RNA duplex formed during transcription. Consequently, upon overtaking the polymerase, the Rho factor releases the RNA transcript, thereby terminating transcription. The mRNA directs protein synthesis. In E. coli, this process begins even before the complete mRNA molecule has finished forming, as transcription occurs in direct contact with the cytoplasm.

However, the precise mechanisms controlling gene transcription in E. coli remain an open question.

Rate of transcription initiation in prokaryotes

Constitutively expressed genes are those that are constantly in an "on" state, and their expression rate is not selectively modulated. Such genes encode enzymes and other proteins constantly required for cell viability. Nevertheless, the cellular demand for certain constitutive proteins can sometimes increase dramatically. The primary influence on gene expression rates in bacteria is the rate of mRNA production, which is largely determined by the frequency of transcription initiation for a given gene. This frequency can vary widely among different genes because they possess promoters of varying "strengths." A "strong" promoter initiates transcription frequently, leading to The production of numerous mRNA transcripts and large amounts of the corresponding protein. A "weak" promoter has the opposite effect. Promoter strength is a function of the base sequences in the Pribnow and -35 boxes, the spacing between them, and the bases in the region from +1 to +10. A higher polymerase affinity for these regions and promoter strength may or may not strictly correlate.

Transcription control by alternative sigma factors

Sigma factors are one of the most effective tools for controlling entire blocks of genes in prokaryotes. Under certain conditions, the standard sigma factor is replaced by an alternative one that directs the polymerase to initiate transcription of a different set of genes. Such conditions include: 1) bacterial sporulation, where a new sigma factor is synthesized upon receiving signals of unfavorable environmental conditions, triggering the expression of a series of sporulation-related genes; 2) nitrogen starvation, which induces a specialized factor; and 3) heat Shock (a sudden temperature rise), which prompts E. coli to temporarily increase the synthesis, stability, and activity of another sigma factor whose baseline level is normally very low. This factor directs the transcription of genes encoding "heat shock proteins," which protect the cell from thermal damage.

Lac Operon

How is The regulation of individual E. coli genes carried out?

Bacteria are forced to adapt to constantly changing environmental conditions. Moreover, fierce competition exists among bacterial populations. A typical E. coli cell under optimal conditions divides every 20 minutes. If a strain is able to shorten this time by just 1 minute, it will rapidly outcompete other strains. In such a competitive environment, biochemical inefficiency cannot be tolerated.

The production of enzymes requires resources and energy, and if these enzymes are not currently needed, E. coli will not synthesize them. As noted earlier, constitutive enzymes are produced constantly under all circumstances. Because glucose is the primary sugar, and other sugars can enter the glucose metabolic pathway, enzymes that metabolize glucose are constitutive. The cell "assumes" they will always be needed, so the promoters of the genes encoding them lack an "off switch" and remain permanently "on".

However, E. coli may encounter other sugars, such as lactose, a disaccharide present in milk. If this is the only available carbon source, its utilization will allow the cell to survive. This requires the enzyme β-galactosidase, named so because lactose is a β-galactoside that hydrolyzes into galactose and glucose (Fig. 21.8). Transporting lactose into the cell requires an additional transport protein, β-galactoside permease. A third protein, galactoside transacetylase, protects the cell from non-metabolized, potentially toxic β-galactosides. Normally, these proteins are synthesized within a few minutes. They are unnecessary until the cell encounters lactose, but when lactose becomes the sole energy source, the synthesis of these three proteins is activated almost instantaneously, enabling the cell to utilize lactose as a source of carbon and energy. However, if glucose is present alongside lactose, producing these enzymes is disadvantageous: why synthesize glucose inside the cell when it is already abundant in the environment? In this case, the cell "ignores" the lactose signal and does not produce the aforementioned enzymes. This regulation, which we will discuss next, operates at the level of gene transcription initiation.

Fig. 21.8. Reactions Catalyzed by β-galactosidase. The enzyme involved in lactose metabolism is often referred to as lactase

STRUCTURE OF THE E. coli lac operon

First, a few Definitions: Protein synthesis in response to a chemical signal is called induction, and the chemical itself is the inducer. The suppression of Protein synthesis is termed repression.

As previously noted, eukaryotic genes are singular structures. Many prokaryotic genes, however, are grouped together, and such a group may fall under the transcriptional control of a single promoter. RNA polymerase transcribes the entire group, producing a polycistronic mRNA molecule that encodes multiple proteins. Such a group of genes controlled by a single promoter is called an operon, and the promoter is part of this operon. The genes for β-galactosidase, lactose permease,

and transacetylase belong to such an operon. They are often designated as z, y, and a, respectively. There is also the I gene (where i stands for inducibility), which encodes the lac repressor protein. The repressor can bind to a DNA region called the operator. Finally, There is a DNA fragment capable of binding a specific protein, the cAMP receptor. It is called the catabolite activator protein (CAP) because it is involved in the induction of various Other Enzymes that catabolize substrates. The approximate Organization of these elements is shown in Fig. 21.9.

Fig. 21.9. Structure of the lac operon. The I gene encodes the lac repressor protein

Now let us examine the control mechanism itself. It is worth noting that the lac promoter on its own is a weak promoter, and RNA polymerase barely binds to it to initiate transcription. Without external assistance, transcription from the lac operon occurs only at a low baseline level. Help arrives in the form of the CAP protein, which binds to the template at a site adjacent to the polymerase binding site. CAP binding induces a bend in The Double Helix at the attachment site, thereby enhancing the promoter's strength. However, CAP does not bind until it interacts with cAMP, which acts as an allosteric regulator of this protein. cAMP is produced within the cell only when glucose levels are low. When glucose concentrations are high, cAMP levels in E. coli are low; consequently, neither the CAP protein nor RNA polymerase binds to the operon, and lac operon transcription remains minimal. This is illustrated in Fig. 21.10a.

Is the lac operon always transcribed when glucose levels are low? No, the presence of lactose is also required for transcription. In the absence of lactose, the lac repressor protein is attached to the operator, blocking RNA polymerase, and transcription does not occur (Fig. 21.10b). The lac repressor is an allosteric protein. When lactose is absent from the environment, it exhibits a high affinity for the operator. If lactose is present in the medium, a low basal level of permease activity allows small amounts of the sugar to enter the cell. During lactose Hydrolysis, a fraction of this disaccharide is converted into an isomer, allolactose (see Fig. 21.8). Allolactose binds to the repressor protein, triggering an allosteric change that causes the repressor to dissociate from the operator, thereby unblocking the operon.

Fig. 21.10. Expression of the lac operon. a — At high glucose concentrations, cAMP is absent, preventing the binding of the CAP protein, which is required for RNA polymerase to attach to the promoter; b — at low glucose concentrations but in the absence of lactose, CAP binds and facilitates RNA polymerase attachment to the promoter. However, transcription still does not occur because the lac repressor is bound to the operator, blocking polymerase movement; c — allolactose binds to the repressor, causing it to detach from the operator, and transcription begins (in the presence of lactose, a small amount of its isomer, allolactose, is formed, which serves as the actual inducer). CAP stands for catabolite activator protein

Why is allolactose the inducer rather than lactose itself? This is likely because allolactose, as a byproduct of the β-galactosidase reaction, is produced in sufficient quantities for induction only when lactose concentrations are high. This prevents induction in the presence of mere trace amounts of lactose.

The Use of allolactose (rather than lactose) as an inducer is highly adaptive. If lactose itself—with its low affinity for the repressor protein—were the inducer, induction would be impossible, as the induced β-galactosidase would maintain a low level of lactose regardless of its external concentration. Once all the lactose was depleted, β-galactosidase would begin consuming allolactose (which is also a substrate for this enzyme), thereby providing a mechanism to terminate induction. Thus, the system continuously tests the rate of lactose influx to decide whether to maintain induction.

Once the operator is freed, RNA polymerase is free to move along the operon, synthesizing a polycistronic mRNA (Fig. 21.10c), the translation of which ultimately leads to the production of the three enzymes.

The lactose operon provided the first deciphered mechanism of prokaryotic gene regulation, which has since proven applicable to many metabolic pathways. The Tryptophan operon (trp operon) of E. coli, containing 5 structural genes, is required to produce three enzymes involved in tryptophan synthesis. It is controlled by the trp repressor protein. When bound to tryptophan, this protein blocks the operator, preventing mRNA transcription.

A second mechanism regulating the trp operon comes into play for those RNA polymerase molecules that manage to overcome the repression block when tryptophan concentrations are high. This mechanism is called attenuation.

In prokaryotes, mRNA Translation begins long before the complete molecule is fully synthesized (after a short segment has been made), meaning a translating ribosome closely follows the polymerase (the translation mechanism is detailed in the next chapter). A ribosome is a nucleoprotein particle that moves along the mRNA, adding amino acid residues to the growing polypeptide chain. The operon initially encodes a 14-amino-acid leader peptide. This differs from the signal Peptides attached to certain proteins (see Chapter 22). The attenuation leader peptide is degraded immediately after synthesis. It contains two consecutive tryptophan residues. If tryptophan molecules are abundant, the ribosome attaches them and moves further along the mRNA; if tryptophan is scarce, however, the ribosome stalls at the tryptophan codons in the attenuation region. When tryptophan levels allow the ribosome to translate past this region, nucleotide pairs forming within the partially synthesized mRNA molecule generate a terminator hairpin, bringing transcription to a premature halt. Conversely, if the ribosome stalls in the attenuation region (due to tryptophan starvation), it interferes with base-pairing, allowing the polymerase to complete Transcription of the operon. Thus, transcription levels are regulated by tryptophan availability. The lower the tryptophan concentration, the more polymerase molecules are able to complete full mRNA synthesis.

Attenuation is known to control the operation of six operons responsible for The biosynthesis of specific Amino Acids. Among them, only the trp operon features repressor control. All six operons share a sensitive method for sensing the intracellular level of their respective amino acid, which relies on the presence of multiple codons for that amino acid within the leader peptide. In the Histidine operon, for example, this peptide contains 7 histidine residues.

The control mechanism of the lac operon (and other operons) in bacteria has been studied in detail. It operates with military precision: CAP binds and assists polymerase attachment, the repressor protein binds and blocks transcription, and the inducer unblocks it. Eukaryotic gene transcription control is considerably more complex. However, we must first examine the process of transcription in eukaryotic genes.

Gene transcription and its control in eukaryotes

Key processes underlying mRNA biogenesis in eukaryotes

The building blocks for RNA synthesis in eukaryotes are the same as those in prokaryotes. Therefore, Fig. 21.1, which illustrates RNA synthesis from four ribonucleotides on a DNA template, applies equally to eukaryotic RNA. The DNA-dependent RNA polymerase that transcribes genes encoding eukaryotic proteins is called RNA polymerase II. (RNA polymerases I and III will be discussed in Chapter 22, which covers the transcription of genes for other RNA molecules.)

The initiation of transcription for eukaryotic genes is complex and will be considered in the section on regulation. Let us assume that initiation has already occurred. This will allow us to examine eukaryotic RNA biogenesis, which differs from the prokaryotic process discussed above.

Termination of transcription in eukaryotes

It remains unclear how RNA polymerase II terminates transcription at eukaryotic genes. Eukaryotic genes lack anything resembling prokaryotic termination signals. When RNA polymerase II reaches the 3'-OH end of a gene, it synthesizes the sequence AAUAAA encoded by the template strand; transcription then continues further before the enzyme somehow completes its task. A specific enzyme cleaves the RNA transcript (beyond the aforementioned sequence), and then another template-independent enzyme carries out polyadenylation, adding approximately 200 adenine nucleotides (using ATP) to form a poly(A) tail. Histone mRNAs lack this tail, suggesting it may not be strictly essential for translation. Some experimental data suggest that polyadenylation ensures mRNA stability within the cell.

RNA polymerase II caps the transcribed RNA

Immediately after the initiation of RNA synthesis, the primary RNA transcript of a eukaryotic gene undergoes a 5'-end modification known as capping. The 5' end of the RNA bears a triphosphate group: the first nucleotide triphosphate incorporated into RNA simply accepts a nucleotide at its 3'-OH group, leaving the original triphosphate group intact. The terminal phosphate of this group is removed and replaced by a GMP residue transferred from GTP (Fig. 21.11). A 5'-5' triphosphate linkage is extremely rare. This is followed by methylation of G at the N-7 position and of the 2'-OH group of the second and occasionally the third nucleotide. Because the cap is formed on the first nucleotide of the RNA, the gene start site (nucleotide +1) is often referred to as the cap site. The cap is believed to protect the mRNA end from exonuclease attack and to participate in the initiation of translation.

Fig. 21.11. Structure of the 5' cap in eukaryotic mRNA. During the reaction with GTP, the terminal nucleoside triphosphate of the primary RNA transcript is converted to a diphosphate with the release of pyrophosphate. Methylation reactions then follow (a third methyl group may be added to the 2'-OH of the primary transcript's next nucleotide). The capped primary transcript is subsequently processed into mRNA

This difference in mRNA biogenesis between pro- and eukaryotes is not the only one, as most eukaryotic genes are interrupted genes.

What are interrupted genes?

Except for short 5'- and 3'-untranslated regions, the length of prokaryotic and eukaryotic mRNA is proportional to the size of the proteins they encode (or,

in the case of polycistronic mRNAs, several encoded proteins). While this holds true for bacterial primary mRNA transcripts, the primary transcripts of most eukaryotic genes are much longer (sometimes 10fold) than expected. This is because the protein-coding base sequence of the mRNA molecule is split into several segments separated by intervening regions that do not encode amino acid sequences. The DNA sequences from which these non-coding fragments are transcribed are termed introns, whereas the protein-coding sequences are called exons. A human gene may contain anywhere from 2 to 50 introns (Fig. 21.12, a), ranging in length from 50 to 20,000 base pairs. Exons are typically no longer than 1000 base pairs. The process of removing introns from the primary transcript and joining the exons to form a mature mRNA molecule is called splicing (Fig. 21.12, b). Although introns are generally considered to be devoid of information, some of them—such as enhancers—are known to contain regulatory sequences (see below).

Fig. 21.12. Primary RNA transcript of a eukaryotic gene by RNA polymerase II. (a) Introns after capping and poly(A) tail addition; (b) the removal of introns to yield mature mRNA is called splicing

Mechanism of splicing

Removing introns from the primary RNA transcript and joining exons into an mRNA molecule is no trivial task. The key to this process is a transesterification reaction, in which a phosphodiester bond is transferred to another -OH group. This occurs without hydrolysis or significant energy loss:

If X–Y represents an RNA chain, it will be cleaved.

We now turn to RNA maturation (Processing)—splicing. Exon-intron junctions are "marked" by consensus sequences. All introns begin with GU and end with AG, although the consensus sequences themselves are longer. In the reaction shown above, ROH is actually the 2'-OH group of an adenine nucleotide within the intron chain (Fig. 21.13). This group attacks the 5'-phosphate of the G nucleotide at the splice site, forming a lariat-like structure. This cleaves the chain at the 3' end of exon 1, releasing a 3'-OH group that then attacks the 5' end of exon 2, thereby joining the two exons.

Fig. 21.13. Two-stage mechanism of mRNA splicing. Because splicing proceeds via the transfer of phosphodiester bonds, it does not require energy consumption

In most eukaryotes, nuclear splicing is catalyzed by complex ribonucleoprotein particles called spliceosomes. The RNA molecules of these particles are designated snRNAs (Small nuclear RNAs).

SnRNAs presumably form hybrid molecules by base-pairing with consensus sequences at the splice sites of the primary RNA transcript, thereby localizing the introns to be excised. A mutation in a consensus sequence can lead to aberrant

splicing or the failure of intron removal. In the genetic disorder β-thalassemia, normal amounts of the Hemoglobin β-chain are not produced due to a G-to-A Substitution at the 5' end of the intron, which prevents the primary transcript from being properly processed into mature mRNA.

In the protozoan Tetrahymena, precise splicing occurs without the assistance of proteins, driven solely by the RNA transcript itself. This discovery struck biochemists like a bombshell: an RNA molecule catalyzing specific Chemical Reactions by itself?! Until then, such functions were attributed exclusively to proteins. Consequently, this finding marked the dawn of a new era in RNA biochemistry and provided further support for certain evolutionary theories suggesting that Earth's earliest life forms were RNA-based, with proteins emerging later.

Thus, the primary transcript of a protein-coding eukaryotic gene—synthesized with the participation of RNA polymerase II—is capped, trimmed, polyadenylated, subjected to splicing, and only then exported from The Nucleus as a mature mRNA.

What is the biological status of introns?

The Biological Significance of introns lies in the fact that they facilitate evolution. Much like protein structure studies revealed discrete domains, genetic investigations uncovered the existence of exons: the latter were found to frequently encode distinct Protein domains, although this correspondence is not always absolute. As noted earlier (see p. 47), new proteins are thought to have evolved through "domain shuffling." According to this concept, a single domain can be reused by combining with others, allowing a new family of proteins to be assembled from already existing domains. Dividing gene segments into exons that encode distinct protein domains facilitates their shuffling. The latter provides a higher rate of new protein formation through recombination rather than point Mutations in DNA, which lead to single Amino Acid Substitutions.

The Role of introns in exon shuffling warrants clarification. Exon shuffling can occur As a result of chromosomal rearrangement. The DNA sequence encoding a protein domain must remain intact, and the presence of introns flanking each side of the exon can facilitate successful transposition. Rearrangements within a protein domain-encoding region would likely disrupt its structure and prove non-viable, given that each protein domain (see p. 47) represents an autonomous structure capable of existing as a discrete entity. Since introns do not carry protein-coding sequences, gene migration is likely accompanied by a lower risk of domain damage and, consequently, more successful exon shuffling.

What is THE ORIGIN OF split genes?

Two hypotheses exist. According to the first, the continuous gene of prokaryotes is primitive, and introns appeared later in the course of evolution (the "late intron" model). However, this fails to explain the origin of introns.

The second hypothesis is represented by the exon theory of genes—the "early intron" model. It postulates the primitivity of introns, which may have originally been fused, non-transcribed flanking regions of ancient minigenes that served as the "building blocks" of modern genes. The fact that RNA molecules can possess self-splicing capabilities removes the necessity for a complex splicing mechanism in early evolutionary stages. In accordance with this hypothesis, the continuous genes of prokaryotes are viewed as a later evolutionary outcome that facilitates rapid Cell Division.

However, if introns were inserted haphazardly into ancient continuous genes, certain protein domain-coding regions might have acquired introns by chance, and if favorably positioned on either side of these regions, such introns would have ensured successful exon shuffling. Assuming that such shuffling conferred an evolutionary advantage, many modern genes formed via exon shuffling should automatically demonstrate a correspondence between Protein Domains and exons. In light of this, it is shuffling, rather than anything else, that led to the correspondence between exons and domains. Nevertheless, this correspondence helps clarify the origin of introns.

The investigation of several ancient genes encoding proteins conserved over a long evolutionary period has failed to reveal any correspondence between exon-intron structure and the structural Properties of the proteins. These findings likely argue against the "early intron" model and point to the primitivity of continuous genes. It is worth noting that exon shuffling is not disputed by either hypothesis; it is the subject of a separate discussion unrelated to the origin of introns.

Split genes in prokaryotes could create complex splicing problems because mRNA is translated by Ribosomes from the 5' end even before the synthesis of the 3' end is complete.

Alternative Splicing

There is one well-known advantage conferred by split genes: alternative splicing. In typical splicing, all exons of a primary transcript are joined together to form a mature mRNA. However, numerous instances are known where splicing occurs according to different patterns: one group of exons forms one mRNA, while another group from the same gene transcript forms another, ultimately resulting in the production of distinct proteins. Such a mechanism can be utilized to generate proteins required in different Tissues or in the same tissue at different developmental times.

Mechanism of Eukaryotic Gene Transcription and Its Control

Although the initiation of eukaryotic gene transcription precedes transcription itself, processing, and so forth, we have postponed discussing this topic until now because it is complex and inseparable from The problem of gene activity control.

Unpacking DNA for Transcription

Recall that to fit within the nucleus, the enormous eukaryotic DNA molecule must be densely packaged (see p. 237). Prior to or during transcription, the DNA Structure is selectively unpacked in the transcribed regions. Two lines of evidence point to this selective DNA unpacking. First, in the polytene Chromosomes of insect Salivary Glands, DNA "unpacking" can be observed under a Light Microscope. Since these unique chromosomes were used in classical studies, we shall examine them first.

The salivary gland cells of fruit fly Drosophila larvae are large and possess an unusual structure. The chromosomes replicate approximately 1,000 times without cell division and lie aligned side-by-side, elongated and tightly packed. Under a microscope, this ordered structure appears banded. While individual bands cannot be resolved in a single chromosome, they become distinguishable at magnifications of several thousand times. As larval development proceeds, successive groups of genes are expressed, and the decondensation of specific bands—known as chromosomal puffs—can be observed during DNA transcription (Fig. 21.14).

Fig. 21.14. Schematic representation of chromosomal puffs in transcriptionally active Chromatin from insect salivary gland polytene chromosomes

Second, in vitro experiments have demonstrated that the transcription of eukaryotic genes is associated with the loosening of DNA packaging. It has been shown that transcriptionally active regions of chromatin become more susceptible to The addition of DNase, an enzyme that hydrolyzes DNA. For instance, globin genes become particularly sensitive to DNase action only in the chromatin of globin-synthesizing cells, and exclusively at the moment of active transcription.

The Need for Genetic control in Differentiated Eukaryotes

An E. coli cell possesses about 4,000–5,000 genes, whereas a human cell has 50,000–100,000. In an E. coli cell, most or all genes are expressed during division. Although the lac operon, for example, is active only at low glucose concentrations and in the presence of lactose, every E. coli cell—unlike eukaryotes—is inherently required to transcribe all of its genes.

In eukaryotes, the induction of specific genes occurs through the action of chemical Inducers. In the Liver, for example, phenobarbital causes an increase in the enzymes that metabolize it. Gene repression also occurs: Cholesterol, for instance, represses the first enzyme of its own biosynthetic pathway (see p. 146). For the most part, such effects are not as critical for cell survival as, say, in the case of the lac operon, but they remain critically important.

Animals possess a variety of Hormones and other regulatory agents that elicit the selective modulation of specific protein synthesis by controlling

transcription of their genes. The complexity of such regulatory effects is extremely high, and indeed, gene activity can be controlled by various factors (see Chapter 26). Such complex regulation has not been found in E. coli.

In eukaryotes, different genes are expressed at different times. Some of them, the so-called housekeeping genes, are expressed constitutively and in all tissues. Glycolytic enzymes, proteins required for DNA and protein synthesis, etc., are needed by all cells, so housekeeping genes are expressed in practically every cell, except for highly specialized ones, such as mature erythrocytes. Other proteins are required only in specialized tissues: immature red Blood Cells produce hemoglobin; there are proteins specific to the liver, Muscles, Kidneys, etc. Each cell type in an Organism contains the same set of genes. This means that in a Muscle cell, for example, only certain genes are expressed. However, the issue of expression is not limited to tissue Specificity: during embryonic development, at an early stage of the differentiation process, certain genes are required, while others begin to function later during The Development of individual tissues. It follows that eukaryotic genes, on the one hand, can be expressed at a constant constitutive rate, and on the other hand, are subject to multiple controls by various hormones and other factors.

The problem is that in prokaryotes, a gene can be "switched off" by a repressor or, in the case of the lac operon, by cAMP or lactose. But how is multiple regulation achieved in a eukaryotic gene?

Structure of Type II Eukaryotic Genes

In this section, we will discuss the transcription of protein-coding genes that use RNA polymerase II.

In both eukaryotic and prokaryotic genes, a promoter with a short sequence between nucleotides -3 and +5 exists at the 5' end of the start fragment, in the so-called initiation region (Fig. 21.15). About 80% of genes have a TATA box centered at position -25. Its sequence resembles the Pribnow box in prokaryotic promoters.

Fig. 21.15. Some upstream activating sequences in the promoter region of a eukaryotic gene

At a distance of 100–200 base pairs upstream of the start site lie short specific DNA sequences recognized by special proteins—transcription factors: the CAAT box, the GC box, and the octamer box (comprising 8 base pairs). In various genes, they are located in the region between -100 and -200 nucleotides. However, there are other DNA boxes or elements—enhancers. They bind transcription factors that accelerate gene transcription. Enhancers can be located thousands of base pairs away from the gene, as their function does not depend on orientation in the DNA.

All this makes the definition of a eukaryotic gene promoter less precise compared to a prokaryotic gene. Is a distant enhancer part of the promoter? One useful working definition is given in Lewin's textbook Genes: "A eukaryotic promoter includes the DNA sequences that must be in relatively fixed positions with respect to the start point." Consequently, the promoter includes the initiation region, the TATA box, and the GC and CAAT boxes located upstream of the initiation point, but excludes enhancers.

One of the most difficult aspects to comprehend is the seemingly random Nature of the genetic control mechanism in eukaryotes. This leads us to the remarkable and somewhat disconcerting Conclusion that no single element is strictly obligatory for transcription, and that different genes have various combinations of elements. About 20% of genes, including many housekeeping genes, lack TATA boxes, some lack other elements as well, or they possess multiple copies of the same element. Variations found in three sample genes are shown in Fig. 21.16.

Fig. 21.16. Promoter elements of three eukaryotic genes. Promoters contain various combinations of TATA boxes, CAAT boxes, GC boxes, and other elements

Mechanism of Transcription Initiation in Eukaryotic Genes

Stable Initiation Complex

In bacteria, RNA polymerase recognizes the correct site on the promoter and binds directly to the DNA, a process facilitated in some cases by, for example, CAP. In eukaryotes, the Formation of the basal initiation complex on DNA—required for all type II genes and to which RNA polymerase attaches—requires A number of proteins known as general transcription factors, which are present in all cells. These include: the TATA-binding protein (TBP), which attaches to the TATA box; eight or more TBP-associated factors (TAFs) associated with the TATA box, forming the TFIID complex (transcriptional factor D for polymerase II); RNA polymerase II and other proteins bind to TFIID, completing the assembly of the basal initiation complex. All of this occurs in genes that have a TATA box. In genes without a TATA box, the complex assembles at the initiation region; in doing so, another factor is presumably used to achieve the correct positioning of the polymerase. If Genes are transcribed by polymerases I and III, various other stable complexes assemble to recruit the corresponding polymerase.

However, as already noted, there are various types of genes. Some of them, such as housekeeping genes, are expressed constantly; others, which are tissue-specific, are expressed only in certain cells; still others are inducible (for example, specific genes whose transcription is activated by Steroid Hormones). Finally, there are genes that, conversely, are subject to negative control, and some that are subject to various combinations of the above types of regulation. How are all these types of regulation carried out? Here it is appropriate to consider specific transcription factors that combine with the enhancer and upstream elements.

What Is the Role of Transcription Factors?

All genes in chromatin (which is DNA bound to proteins; see p. 237) are in an "off" or repressed state in the absence of transcription factors, because a nucleosome blocks the initiation region of each promoter. Until the nucleosome is displaced—allowing the assembly of the initiation complex—the gene is not transcribed.

Fig. 21.17. Model of the basal initiation complex for RNA polymerase II. The TFIID complex includes the TATA-binding protein (TBP) and numerous TAF proteins. The shapes, sizes, and positions of the components are arbitrary

Various transcription factors can participate in the activation of the same gene. Chief among these are factors that bind to the GC and CAAT boxes located upstream of the start site. They are present in all cells, and constitutively expressed genes may require only these. How do transcription factors that bind to distant elements participate in complex formation? This can occur through DNA looping. Each factor has at least two binding domains—one for DNA and another for proteins involved in the formation of the transcription complex, as shown in Fig. 21.18.

Fig. 21.18. Hypothetical model showing how a protein binding to a DNA site distant from the start point can participate in the initiation of transcription by RNA polymerase II. The interaction regions between the transcription factor, the basal initiation complex, and RNA polymerase are drawn arbitrarily. Enhancer elements located thousands of base pairs away can participate in complex formation via a looping mechanism (see Fig. 21.19)

How is the selective Control of Gene Expression achieved? This is the function of elements located in various Regions of the DNA. Even at a distance of thousands of base pairs, they can influence gene transcription. For example, the presence of a distant enhancer in DNA increases β-globin gene transcription by hundreds of times.

The β-globin gene is expressed exclusively in immature erythrocytes; its transcription depends on the binding of a transcription factor to the GATA box. This factor is present only in these cells, thereby ensuring expression specificity.

How is the transcription rate of an inducible gene regulated? As an example, let us consider the Activation of a number of genes by steroid hormones.

Target cells contain inactive transcription factors. They are unable to bind independently to the corresponding DNA elements. A hormone entering the cell activates the transcription factor, which subsequently binds to DNA, triggering transcriptional activation. The Mechanism of this factor's activation is described in Chapter 26, where we will explore cell signaling, a primary mechanism of which is the activation of transcription factors. Note that negative control of gene transcription also exists.

The fundamental question—how the interaction between transcription factors and the initiation complex regulates initiation—remains unanswered for now. An important feature of EUKARYOTIC GENE EXPRESSION regulation is its high flexibility; that is, many factors can contribute to Transcriptional Regulation provided an appropriate binding site is available on the DNA molecule (Fig. 21.19).

Fig. 21.19. Multiple control of eukaryotic gene transcription initiation (schematic). The sizes, positions, and interactions of proteins, as well as the combination of regulatory elements, are conventional.

Structure of DNA-binding proteins

From the foregoing in this chapter, the crucial role of proteins that bind to specific DNA regions becomes apparent. There are numerous repressors and transcription factors involved in differentiation, embryonic development, and so on. Their operational efficiency depends on the ability of these compounds to recognize and bind to DNA.

Extensive research has been undertaken to elucidate the structure of proteins interacting with corresponding DNA regions. It turned out that, based on structural characteristics, most DNA-binding Proteins can be grouped into several families, the most important of which are helix-turn-helix motif proteins, homeodomain proteins, leucine zipper proteins, and finally, zinc finger motif proteins.

The structures, or motifs, that gave these names represent only small regions of the proteins characterizing the entire family. For instance, Fig. 21.20 shows only the family-characteristic helix-turn-helix fragment of a DNA-binding protein monomer.

Fig. 21.20 Fragment of a DNA-binding protein and its interaction with DNA. a - Helix-turn-helix fragment; b - binding of this protein to DNA within the major grooves

DNA-binding proteins attach to double-stranded DNA; site-specific binding occurs through interactions between the side chains of amino acid residues and DNA bases. Additional stabilization can be achieved by binding to the sugar-phosphate backbone. Contact interactions between the protein's recognition motifs and nitrogenous bases occur predominantly in the major groove of the double helix, where certain atoms of the nitrogenous bases are exposed. In many cases, such binding is driven by the recognition of a protein α-Helix that nestles into the major groove.

It is important to note that specific protein attachment on the DNA molecule requires the presence of two adjacent sites. These may be identical, formed by a palindromic arrangement of bases (see below), or distinct. Capable of binding to palindromic sites are, for example,

homodimeric proteins, consisting of two identical subunits, while others bind heterodimeric proteins, consisting of different subunits. Proteins with zinc finger fragments may feature multiple DNA-recognizing motifs.

The Essence of variable gene control is that DNA-binding proteins, such as transcription factors and repressors, in many cases interact with DNA only after receiving appropriate "instructions." We have already encountered situations where transcription factor activation is achieved through allosteric modification of a regulatory protein upon Ligand binding (recall the Regulation of the lac operon). In eukaryotes, such an "instruction" typically arrives in the form of an external extracellular signal, such as a hormone or growth factor. Such mechanisms of transcription factor activation are discussed in Chapter 26.

Now let us examine the families of DNA-binding proteins.

Proteins with the helix-turn-helix (HTH) motif

This is the first identified family of DNA-binding proteins. Many proteins of this type have been found in prokaryotes, such as the lambda repressor. E. coli bacteriophage lambda is described on p. 314. We will not discuss the role of the repressor protein here as it is rather complex; suffice it to say that it binds to a specific region of phage lambda DNA and participates in the "decision-making" process regarding whether the infecting virus enters the lytic or lysogenic cycle (for clarification, see pp. 314–315). What is far more important right now is that the mechanism of its binding to DNA has been studied in great detail.

As we have already explained, the helix-turn-helix (HTH) motif represents a small region of the protein through which binding to the operator region of DNA occurs. (This is not a protein domain and cannot exist independently upon isolation; a motif denotes a recognizable structural property.) The motif comprises two α-helices connected by a β-turn. One of the helices (the recognition helix) is located in the major groove of DNA (see Fig. 21.20, a). The operator or DNA-recognizing region is an imperfect palindrome. A palindrome is a phrase or word that reads the same forwards and backwards. For example, the phrase "Madam I'm Adam" is symmetrical with respect to the letter "I". Palindromic DNA-binding regions exhibit dyad Symmetry (where two units act as one or as a group of two). The structure of the operator is shown below.

The two halves of the palindrome represent identical binding sites, although nucleotides indicate that the palindrome is not absolute. The repressor protein is a dimer, and the recognition helices of its two HTH fragments fit these specific nucleotides (see Fig. 21.20, b).

Leucine zipper-containing proteins

These have been discovered among numerous Eukaryotic Transcription factors. In this case, the name denotes not a recognition fragment (as in HTH proteins), but a characteristic structure that causes the dimerization of two subunits (the latter may not necessarily possess identical recognition sites). In a leucine zipper fragment, every seventh amino acid is a leucine. Since an α-helix turn comprises 3.6 amino acid residues, all leucine residues appear on one side of the α-helix, forming a Hydrophobic surface. Two such subunits join through hydrophobic interactions between the leucine side chains (Fig. 21.21, a). Incidentally, the term "zipper" is a misnomer. It was proposed when it was assumed that leucine residues interdigitate like the Teeth of a zipper. In reality, the DNA-binding region of each monomer is a region rich in positively charged amino acid residues—Arginine and Lysine. Figuratively speaking, the dimer attaches to DNA at the handle of "scissors," positioning itself between the two "blades" in adjacent segments of the duplex major groove (Fig. 21.21, b).

Fig. 21.21. Leucine zipper-containing proteins. a - Structure of the leucine zipper fragment. Hydrophobic leucine residues lie opposite each other but do not interdigitate like teeth on a zipper; b - attachment of a leucine zipper protein to DNA (viewed down the DNA axis)

Zinc finger proteins

In this DNA-binding motif, a zinc atom coordinated to 2 histidine residues and 2 Cysteine residues stabilizes a finger-like structure (Fig. 21.22, a), one part of which is a recognition α-helix located within the major groove of DNA (Fig. 21.22, b). Transcription factors may contain multiple such fingers, establishing multiple contacts with DNA (Fig. 21.22, c). A single such gene regulatory protein possesses approximately 30 zinc fingers. Not all zinc fingers bind specifically to DNA; some are nonspecific, yet they help stabilize specific binding. The zinc finger motif is characteristic of many regulatory proteins of various eukaryotic genes, such as intracellular steroid Hormone Receptors.

Homeodomain proteins and embryonic development

Deciphering the mechanisms of Cell Differentiation and embryonic development is one of the paramount tasks of biology. The fertilized egg of a multicellular organism divides and differentiates into various tissues: epidermis, nerves, muscles, liver, etc. Since every cell contains the same set of genes, differentiation involves selective GENE EXPRESSION IN different tissues. Little is known yet about how this process actually occurs. However, several years ago, a number of homeotic genes responsible for proper embryonic development were identified. Their significance was established by studying the effects of mutations in these genes on the Development of the fruit fly Drosophila. Of note in this context is that the products of homeotic genes—all the proteins they encode—share similar domains known as homeodomains. The DNA segments encoding these homeodomains are called homeoboxes. Homeodomains are DNA-recognizing proteins about 60 amino acid residues in length; they resemble the helix-turn-helix motif, with the recognition helix positioned in the major groove. The proteins encoded by homeotic genes bind to specific DNA sites, suggesting that homeotic genes encode transcription factors. Although the target genes are not yet known, it has been established that in some cases these factors act as activators of certain genes and repressors of others. The structure of homeodomains exhibits a high degree of conservation and striking similarity between insects (Drosophila) and vertebrates, including humans.

Fig. 21.22. Zinc finger proteins. a - Structure of a zinc finger; b - model of interaction between a target DNA region and a zinc finger-containing protein, where one side of the finger structure forms an α-helix that binds to the major groove of DNA; c - multiple zinc finger motifs bound to DNA

Let us now clarify the terminology. In the context of transcription, the term "box" is usually associated with a DNA sequence recognized by a protein (for example, the TATA and GC boxes discussed previously). When discussing homeobox genes, however, the word "box" refers to the base sequence encoding the homeodomain. In this case, the term carries a completely different meaning, even though both Examples involve specific DNA regions.

mRNA stability and the control of gene expression

Although transcriptional regulation plays a profound role in gene expression, the stability of individual mRNAs is no less important. In most cases, the Rate of protein synthesis reflects The amount of mRNA produced, except when messenger RNA (mRNA) translation is controlled via mechanisms such as those described for the synthesis of erythrocyte aminolevulinate synthase and globin (see pp. 366 and 367, respectively). The cellular mRNA level depends on the balance between its rates of Synthesis and degradation. At a given rate of mRNA production, its cellular content will be higher the longer its half-life, naturally leading to a higher rate of synthesis for the corresponding protein. Mechanisms influencing the mRNA half-life thus provide a means of regulating gene expression.

The half-life of prokaryotic mRNA is typically limited to 2–3 minutes. Rapid turnover of the messenger allows for a prompt response to environmental changes. In mammals, the half-life of individual mRNAs ranges from 10 minutes to 2 days. One of the most stable mRNAs (globin mRNA) has a half-life of approximately 10 hours. Regulatory proteins are usually encoded by short-lived mRNAs, so Changes in the transcription rate of their genes are rapidly reflected in their expression rate. The half-life of transcription factor mRNAs is typically less than 30 minutes. Thus, the cell contains mRNAs with varying degradation rates, and the stability of individual mRNAs can be modulated.

Despite the potential importance of mRNA stability, the process of its decay is much less understood than synthesis and its regulation. Little is known

either about the enzymes involved in the degradation process, yet some mechanisms determining mRNA lifespan have already been elucidated.

Structures determining mRNA stability and their role in expression regulation

Nearly all mRNAs possess a poly(A) tail attached to the 3' end before the RNA molecule leaves the nucleus (Fig. 21.23, a); histone mRNAs are a notable exception. The poly(A) tail is thought to protect eukaryotic mRNA from rapid degradation. Deadenylation often precedes mRNA destruction. A protein that binds to the poly(A) tail, thereby inhibiting mRNA decay from the 3' end of the molecule, has already been discovered, but the precise mechanism by which poly(A) protects mRNA remains fully understood. It is hypothesized that poly(A) influences mRNA transport, translation, and decay.

Histone mRNAs lack a poly(A) tail, and the stability of these molecules is instead determined by a characteristic stem-loop structure at the 3' end (Fig. 21.23, b). Histone synthesis is required exclusively during the S phase of The Eukaryotic Cell cycle, when DNA Replication and nucleosome assembly take place. Histone genes are transcribed during the S phase, but this process ceases in the G2 phase, causing histone mRNA levels to drop rapidly. The latter is partly due to transcriptional arrest, but additionally, the mRNA half-life plummets from 40 to 10 minutes, a process dependent on the presence of the 3' stem-loop structure. Transferring this stem-loop to globin mRNA creates a hybrid messenger that, when introduced into cultured cells, is destabilized at the end of the S phase just like histone mRNA. Destabilization of histone mRNA requires the presence of free histone monomers, which accumulate shortly after the completion of DNA synthesis (and hence nucleosome assembly) at the end of the S phase. Rapid shutdown of histone synthesis is essential due to the cytotoxicity of free histone monomers. A fourfold decrease in histone messenger half-life means that following the cessation of histone gene transcription, mRNA degradation continues for about 2 hours, whereas without destabilization it would take approximately 9 hours (Fig. 21.24).

Fig. 21.23. Organization of structures present in the 3' untranslated region (UTR) of mammalian mRNAs that influence the half-life of the molecules within the cell. a - Poly(A) tail found in the majority of eukaryotic mRNAs; b - stem-loop structure found in histone mRNAs; c - iron-responsive element (IRE) of transferrin mRNA; d - AU-rich element (AURE) found in A large number of unstable eukaryotic mRNAs

Fig. 21.24. Effect of a fourfold decrease in mRNA half-life on its cellular Abundance following the cessation of synthesis

The synthesis of β-tubulin is regulated by a feedback mechanism that modulates mRNA stability. Chapter 29 describes the role of tubulin aggregation and its impact on microtubule formation. The presence of free tubulin monomers destabilizes its mRNA. For degradation to occur, the mRNA must undergo translation, as the partially synthesized peptide somehow facilitates the destabilization process. The mechanism of tubulin-induced mRNA decay is not yet fully understood, but the regulatory logic of this system is clear.

The synthesis of the transferrin receptor protein provides another example of how mRNA stability regulates protein synthesis. This receptor is responsible for iron uptake into cells (see p. 368). The 3' untranslated region of its mRNA contains a cluster of five stem-loop structures known as iron-responsive elements (IREs) (see Fig. 21.23, c). In the absence of iron, an IRE-binding protein attaches to the IRE and stabilizes the mRNA, thereby increasing receptor synthesis and enhancing cellular iron uptake. In the presence of excess iron, an iron-protein complex forms, preventing the protein from binding to the IRE and presumably rendering this region accessible to endonucleases.

Short-lived mRNAs may contain so-called instability sequences that signal rapid mRNA degradation within the cell. A structural hallmark of many unstable mRNAs is the presence of adenine- and uracil-rich elements (AUREs) in the 3' untranslated regions of the molecules (see Fig. 21.23, d). These elements vary among different mRNAs, but each contains AU sequences at least 9 bases long. They trigger mRNA destabilization, likely via AURE-binding proteins. How AUREs function remains elusive. Their effect on mRNA half-life may be indirect, mediated through influences on translation; however, the presence of AUREs is invariably associated with rapid mRNA deadenylation. A particularly interesting example is the mRNA product of the c-fos gene, which encodes transcription factors containing a leucine zipper. This mRNA contains an AURE in its untranslated region (the terms c-fos, v-fos, and oncogenicity are explained on p. 356). The v-fos gene is a potent oncogene that induces Cancer development; its mRNA lacks the AURE found in c-fos mRNA and exhibits a longer cellular half-life than c-fos mRNA. Deletion of the AURE region converts the normal c-fos gene into an oncogene (for the relationship between regulatory proteins and oncogenicity, see Chapter 26).

The mechanisms governing mRNA stability in the cell are complex. In addition to the points discussed above, a number of hormones are known to modulate the stability of specific messengers. In conclusion, the regulation of mRNA decay represents a vital mechanism for controlling gene expression.

Mitochondrial gene transcription

The Central Role of Mitochondria in Energy production through The oxidation of dietary substrates was already discussed in previous chapters. Mitochondria are self-replicating Organelles of eukaryotic cells. Like Chloroplasts, they possess their own DNA and protein-synthesizing machinery. Mitochondria divide, thereby maintaining a characteristic number of these organelles in different cells: about 1,000 in a rat hepatocyte and 107 in a frog oocyte. Mammalian Mitochondrial DNA is typically a circular duplex containing fewer than 20,000 nucleotides; each mitochondrion may contain 5 or 10 copies of this duplex.

Mitochondrial DNA encodes only a small fraction of the proteins within these organelles; the remaining proteins are encoded by nuclear DNA. These proteins are synthesized in the cytoplasm and transported into the mitochondria (see Chapter 22). Within the mitochondria, both DNA strands are fully transcribed into long linear transcripts (polymerase molecules move in opposite directions, as synthesis always proceeds in the 5' —> 3' direction). These primary transcripts undergo processing (via a mechanism not yet fully understood) to generate mRNA, tRNA, and rRNA. To produce sufficient amounts of rRNA, many shorter transcripts surrounding the rRNA genes are synthesized.

Unlike cytoplasmic protein synthesis (cytoplasmic translation), mitochondrial translation exhibits several features characteristic of prokaryotes. This led to the hypothesis that mitochondria originated from Prokaryotic Cells engulfed by eukaryotic cells, which subsequently became endosymbionts (a similar theory explains the origin of chloroplasts in plant cells). Indeed, The properties of the mitochondrial transcription system represent a blend of prokaryotic and eukaryotic types. In mammalian mitochondria, mRNAs are polyadenylated (a eukaryotic trait) but lack a cap (a prokaryotic trait), and, just like in prokaryotes, Mitochondrial Genes contain no introns.

In trypanosome mitochondria, the resulting mRNAs undergo subsequent editing. RNA transcripts possess additional nucleotides inserted at specific sites to form the correct protein-coding sequence, which were not originally encoded in the DNA. It was later discovered that mRNA editing can also occur in nuclear transcripts. For instance, transcription of the mammalian apolipoprotein B gene (see p. 96) yields two different mRNAs from a single transcript. The "unedited" mRNA encodes apoproduct B100, whereas the smaller apolipoprotein B48 is encoded by the "edited" version, in which a specific C residue in the DNA is converted to U, generating a translational stop codon (see Chapter 22).

Simian virus (whose genes possess eukaryotic characteristics) produces mRNAs that are edited by the insertion of two G residues not encoded in its DNA. It remains unclear whether these variations serve a biological purpose or merely represent an example of evolutionary "patchwork."

Non-Protein-Coding Genes

In addition to genes responsible for protein synthesis, there are genes that encode specialized RNA molecules which do not act as messengers. As you will learn in the next section, proteins are synthesized by nucleoprotein particles called ribosomes, and the assembly of polypeptide chains from amino acids requires numerous small transfer RNAs (tRNAs). Ribosomes contain ribosomal RNA (rRNA) molecules; prokaryotes contain 3 distinct rRNA molecules, whereas eukaryotes contain 4. During transcription, roughly half of the total RNA produced corresponds to rRNAs. Because tRNAs and rRNAs have much longer half-lives than mRNAs, the vast majority of cellular RNA consists of these forms. In E. coli cells, the three rRNAs and several tRNAs are transcribed together as a large precursor molecule, which is subsequently cleaved into the appropriate smaller fragments. Unlike prokaryotes, eukaryotes possess three distinct RNA polymerases.

RNA polymerase II, which transcribes protein-coding genes, has already been described. RNA polymerase I transcribes the bulk of the rRNAs, while RNA polymerase III transcribes tRNAs and smaller rRNAs. Because cells require large amounts of tRNAs and rRNAs, multiple copies of their genes exist. Eukaryotes may harbor hundreds or thousands of rRNA genes organized in tandem HEAD-to-tail arrays located within the nucleolar regions of the nucleus. Genes transcribed by polymerase III are unusual in that their regulatory elements, or boxes, are located not only upstream of the 5' transcription start site of the DNA but also within the transcribed region itself, downstream of the start site.

Questions for Chapter 21

1. What are the key differences between RNA synthesis and DNA synthesis?

2. Using diagrams, describe the components of an E. coli gene and its flanking regions.

3. Describe the process of transcription initiation in an E. coli gene.

4. Describe two mechanisms of transcription termination in an E. coli gene.

5. What factors determine promoter strength in a constitutive E. coli gene?

6. Describe the regulation of the lac operon.

7. Describe the main differences between RNA biogenesis in eukaryotes and prokaryotes.

8. Using diagrams, explain the mechanism of splicing. What is the potential biological significance of introns?

9. How does transcription initiation in eukaryotes differ from that in prokaryotes?

10. What families of transcription factors have been classified based on structural motif similarities?



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.