LEHNINGER PRINCIPLES OF BIOCHEMISTRY - VOL 3. INFORMATION PATHWAYS - 2017
CHAPTER III. INFORMATION PATHWAYS
26. RNA METABOLISM
26.3. RNA-Dependent Synthesis of RNA and DNA
Thus far in our Structure/133.html">Discussion of DNA and RNA Synthesis, we have considered exclusively DNA molecules acting as templates. However, some Enzymes use an RNA template for the synthesis of Nucleic Acids. With the very important exception of Introduction/7.html">RNA-containing Viruses, these enzymes play only a modest role in informational pathways. It is precisely RNA-containing viruses that are the source of most currently known RNA-dependent polymerases.
The existence of RNA Replication required a revision of the Central dogma of molecular biology (Fig. 26-32; compare with the diagram on p. 5). Enzymes involved in RNA replication help shed light on The Nature of self-replicating molecules that may have existed in prebiotic times.
Class="center">Fig. 26-32. Expansion of the central dogma of molecular biology to include RNA-dependent RNA and DNA Synthesis.

Reverse Transcriptase synthesizes DNA from a viral RNA template
Some RNA-containing viruses that infect animal Cells carry an RNA-dependent DNA polymerase, called reverse transcriptase, within their Viral Particles. Upon infection, the enzyme enters the host Cell along with the single-stranded viral RNA genome (~10,000 NUCLEOTIDES). Reverse transcriptase first catalyzes the synthesis of a DNA strand complementary to the viral RNA (Fig. 26-33), then degrades the RNA strand in the RNA-DNA hybrid and replaces it with DNA. The resulting double-stranded DNA is frequently integrated into the host cell genome. Such integrated (and latent) viral genes can be activated and transcribed, and the products of these genes—viral Proteins, as well as the viral RNA genome itself—can be packaged into new viral particles. RNA viruses containing reverse transcriptases are called Retroviruses (from the Latin prefix retro, meaning "backward").
Fig. 26-33. Retroviral infection in a mammalian cell and Integration of the retrovirus into the host chromosome. Viral particles entering the host cell carry viral reverse transcriptase and cellular tRNA (captured from the previous host cell) already base-paired with the viral RNA. The tRNA molecules mediate the immediate conversion of viral RNA into double-stranded DNA with the participation of reverse transcriptase, as described in the text. The double-stranded DNA enters The Nucleus and is integrated into the host genome. Integration is catalyzed by viral integrase. The Mechanism of viral DNA integration into host DNA resembles that of transposon insertion into bacterial Chromosomes (see Fig. 25-45). For example, at the site of integration, a few Base Pairs of host DNA are duplicated, forming short repeats of 4 to 6 bp on either side of the retroviral DNA insert (not shown).

The existence of reverse transcriptases in RNA-containing viruses was predicted in 1962 by Howard Temin, and in 1970 these enzymes were discovered by Temin and independently by David Baltimore. This discovery attracted widespread attention because it challenged the central dogma and demonstrated that Genetic information can flow "backward" from RNA to DNA.

Retroviruses typically carry three genes: gag (from the historical English name, group-associated antigen), pol, and env (Fig. 26-34). The transcript containing gag and pol is translated into a long polyprotein—a polypeptide that is cleaved into six proteins with distinct Functions. The proteins encoded by the gag Gene form the core of the viral particle. The pol gene encodes a protease that cleaves the long polypeptide, an integrase that inserts the viral DNA into the host chromosome, and reverse transcriptase. Many reverse transcriptases consist of two subunits, $\alpha$ and $\beta$. The pol gene encodes the $\beta$ subunit ($M_r$ = 90,000), whereas the $\alpha$ subunit ($M_r$ = 65,000) is a proteolytic fragment of the $\beta$ subunit. The env gene encodes viral envelope proteins. At each end of the linear RNA genome are long terminal repeats (LTRs) consisting of several hundred nucleotides. When transcribed into double-stranded DNA, these sequences facilitate the integration of the viral chromosome into the host DNA and contain promoters for viral Gene Expression.
Fig. 26-34. Structure and products of the integrated retroviral genome. Long terminal repeats (LTRs) contain sequences required for METABOLISM/31.html">Transcription regulation and initiation. The $\Psi$ sequence is necessary for packaging retroviral RNA molecules into mature viral particles. Transcription of retroviral DNA yields a primary transcript comprising the gag, pol, and env genes. Translation (Chapter 27) produces a long polyprotein—a polypeptide encoded by the gag and pol genes, which is cleaved into six proteins with different functions. Splicing of the primary transcript synthesized from the env gene yields an mRNA that is also translated into a polyprotein, which is subsequently cleaved to form viral envelope proteins.

Reverse transcriptases catalyze three distinct reactions: (1) RNA-dependent DNA synthesis, (2) RNA degradation, and (3) DNA-dependent DNA synthesis. Like many DNA and RNA polymerases, reverse transcriptases contain Zn2+. Although any transcriptase is most active with its own viral RNA, all of them can be used experimentally to generate DNA molecules complementary to A wide variety of RNA molecules. Different active sites on the protein are involved in DNA and RNA synthesis and RNA degradation. To initiate DNA synthesis, reverse transcriptase requires a primer: a cellular tRNA captured from the previous host and carried within the viral particle. The 3' end of this tRNA is base-paired with a complementary sequence of the viral RNA. The new DNA chain is synthesized in the 5' $\rightarrow$ 3' direction, as in all reactions involving RNA and DNA polymerases.
Like RNA polymerases, reverse transcriptases lack a 3' $\rightarrow$ 5' proofreading exonucleolytic activity. The error rate made by these enzymes is approximately 1 per 20,000 incorporated nucleotides. Such a high error rate is highly unusual for DNA replication and is likely characteristic of most enzymes that replicate RNA virus genomes. The result is a high mutation rate and rapid viral evolution, leading to the frequent emergence of new pathogenic retroviral strains.
Reverse transcriptases have become essential tools in The Study of DNA-RNA interactions and in DNA cloning Methods. They enable the synthesis of DNA complementary to an mRNA template; synthetic DNA prepared in this manner is called complementary DNA (cDNA) and can be used to clone cellular genes (see Fig. 9-14 in Vol. 1).
Some retroviruses cause Cancer and AIDS
■ The study of retroviruses has greatly advanced our understanding of the molecular nature of cancer. Most retroviruses do not kill host cells; instead, they remain integrated into the cellular DNA and replicate when The Cell divides. Some retroviruses, classified as RNA tumor viruses, carry an oncogene capable of inducing uncontrolled cell growth. The first virus of this type to be studied was the Rous Sarcoma virus (also known as avian sarcoma virus; Fig. 26-35), named after F. Peyton Rous, who investigated tumors in chickens caused by this virus. Following the discovery of oncogenes in retroviruses by Harold Varmus and Michael Bishop, dozens of such genes have been identified.
Fig. 26-35. Genome of the Rous sarcoma virus. The src gene encodes a Tyrosine kinase, belonging to a class of enzymes that influence Cell Division, intercellular interactions, and intracellular signal Transduction (Chapter 12 in Vol. 1). This same gene has been found in the DNA of chickens (the natural host of this virus) and in the genomes of many eukaryotes, including humans. Upon infection by the Rous sarcoma virus, this oncogene becomes overexpressed, promoting uncontrolled cell division and tumorigenesis.
![]()
The HUMAN IMMUNODEFICIENCY VIRUS (HIV), which causes Acquired Immunodeficiency Syndrome (AIDS), is also a retrovirus. Identified in 1983, HIV has an RNA genome with the standard Complement of retroviral genes along with several unusual genes (Fig. 26-36). Unlike many other retroviruses, HIV tends to kill the cells it infects (primarily T lymphocytes) rather than induce tumor formation. This progressively leads to suppression of the host immune system. The reverse transcriptase of HIV makes even more errors than other known reverse transcriptases—10-fold or more—leading to a high rate of viral mutation. With each replication of the HIV genome, one or more errors typically occur, meaning that any two RNA molecules of this virus may differ from one another.
Fig. 26-36. The HIV genome. In addition to typical retroviral genes, the HIV genome contains several small genes with various functions (not all of which have been identified and discussed here). Some of these genes overlap (see p. 175). As a result of Alternative Splicing, this compact genome (9.7 • 103 nucleotides) encodes a large variety of proteins.

Many modern Vaccines effective against viral infections consist of one or more viral envelope proteins, produced using the methods described in Chapter 9 (Vol. 1). These proteins themselves do not cause infection, but they stimulate The Immune System, helping it recognize and combat the virus upon potential exposure (Chapter 5, Vol. 1). Due to the high error rate of reverse transcriptase, the HIV env gene (and the entire genome) mutates very rapidly, making The Development of an effective vaccine extremely difficult. However, because ongoing cycles of cell infection and replication are required for the spread of HIV infection, the most effective current treatments rely on inhibiting viral enzymes. HIV protease is the target of pharmacological agents known as protease inhibitors (see Box 6-3, Vol. 1). Reverse transcriptase is the target of several other drugs widely used in the Treatment of HIV-infected patients (see Box 26-2). ■
Box 26-2. MEDICINE. Combating AIDS with Reverse Transcriptase Inhibitors
The life cycle and STRUCTURE OF THE human immunodeficiency virus (HIV) were elucidated through the study of template-directed nucleic acid Biosynthesis using modern molecular biology techniques. Just a few years after the isolation of HIV, these investigations led to the development of drugs that extend the lives of individuals infected with the virus.
The first drug approved for clinical trials was azidothymidine (AZT), a structural analogue of deoxythymidine. AZT was first synthesized in 1964 by Jerome Horwitz. In 1985, it was discovered that this unsuccessful anticancer agent (originally developed to combat cancer) could be useful in treating AIDS patients. T lymphocytes, the immune system cells most susceptible to HIV infection, uptake AZT and convert it to AZT triphosphate. (Direct administration of AZT triphosphate is ineffective because this compound cannot cross The Plasma Membrane.) HIV reverse transcriptase has a higher affinity for AZT triphosphate than for dTTP, and the binding of AZT triphosphate to the enzyme competitively inhibits dTTP binding. When AZT is incorporated into the 3' end of the growing DNA chain, the absence of a 3'-hydroxyl group halts further synthesis of viral DNA.

AZT triphosphate is non-toxic to T lymphocytes themselves because cellular DNA polymerases have a lower affinity for this compound compared to dTTP. At concentrations of 1–5 µmol, AZT affects HIV reverse transcription without interfering with cellular DNA replication. Unfortunately, AZT proved toxic to Bone Marrow precursor cells of erythrocytes, causing many patients taking AZT to develop anemia. Administration of AZT can extend the lifespan of individuals with chronic HIV infection by approximately one year and significantly delays the onset of AIDS symptoms in patients in the Cytology/cytology/16.html">Early stages of HIV infection. Several other anti-HIV agents, such as dideoxyinosine (DDI), share a similar MECHANISM OF ACTION. More modern therapeutics inactivate HIV protease. Due to the high error rate of reverse transcriptase and the resulting Rapid Evolution of the virus, most effective treatment regimens are based on a combination of drugs that suppress both protease and reverse transcriptase.
Many Transposons, retroviruses, and introns may share a common evolutionary origin
The structure of certain well-characterized transposons from sources as diverse as Yeast and Drosophila closely resembles that of retroviruses; these are sometimes referred to as retrotransposons (Fig. 26-37). Retrotransposons encode an enzyme homologous to retroviral reverse transcriptase, and their coding regions are flanked by LTR sequences. They move within the cellular genome from one position to another via an RNA intermediate, using reverse transcriptase to generate a DNA copy of the RNA, which is subsequently integrated into a new site. Most eukaryotic transposons utilize this mechanism for transposition, distinguishing them from bacterial transposons, which move directly as DNA fragments from one chromosomal Location to another (see Fig. 25-45).
Fig. 26-37. Eukaryotic transposons. The Saccharomyces Ty element and the Drosophila fruit fly copia element are Examples of eukaryotic transposons with a retrovirus-like structure that lack the env gene. The sequences of the Ty element's δ element are the functional equivalent of retroviral LTR sequences. In the copia element, the INT and RT sequences are homologous to the integrase and reverse transcriptase domains of the pol gene.

Retrotransposons have lost the env gene and cannot form viral particles. Therefore, they should be regarded as defective viruses trapped within the cell. A comparison of retroviruses with eukaryotic transposons indicates that reverse transcriptase is an extremely ancient enzyme that emerged before the evolution of Multicellular Organisms.
Interestingly, many group I and group II introns are also Mobile Genetic Elements. Not only are they capable of self-splicing, but they also encode DNA endonucleases that facilitate their mobility. During genetic exchange between Cells of the same species, or upon the introduction of DNA into a cell by parasites or via other means, these endonucleases promote the insertion of the intron into an identical site in another DNA copy of a homologous gene that lacks this intron. This process is known as homing (Fig. 26-38). Homing of group I introns is DNA-based, whereas homing of group II introns proceeds via an RNA intermediate. Group II intron endonucleases also exhibit reverse transcriptase activity. These proteins can form complexes with intron RNA after the introns have been excised from primary transcripts. Because the homing process involves the insertion of intron RNA into DNA followed by reverse transcriptase-mediated cDNA synthesis, the movement of these introns is termed retrohoming. Within a population, each copy of a given gene may eventually acquire an intron. Much less frequently, an intron spontaneously inserts into a novel site within an unrelated gene. If the host cell survives this event, it can mark the beginning of evolutionary changes and lead to the proliferation of the intron at a new locus. The structures of mobile introns and the mechanisms of their propagation support the hypothesis that at least some of them originated as molecular parasites whose evolution traces back to retroviruses and transposons.
Рис. 26-38. Mobile introns: homing and retrohoming. Certain introns contain a gene (red) for an enzyme that promotes homing (some group I introns) or retrohoming (some group II introns). (a) The gene within the excised intron is bound by a ribosome and translated. Group I homing introns encode a site-specific endonuclease termed a homing endonuclease. Group II retrohoming introns encode a protein with dual endonuclease and reverse transcriptase activities. (b) Homing. Allele a of gene X, which contains a group I homing intron, resides in a cell containing allele b of the same gene but lacking the intron. The homing endonuclease generated by allele a cleaves allele b at the position corresponding to the intron in allele a; subsequently, double-strand break repair (recombination with allele a; see Fig. 25-31a) generates a new copy of the intron in allele b. (c) Retrohoming. Allele a of gene Y contains a group II retrohoming intron; allele b lacks the intron. The excised intron inserts into the coding strand of allele b via a reaction that is the reverse of splicing (whereby the intron is excised from the primary transcript; see Fig. 26-15); however, insertion occurs into DNA rather than RNA. The noncoding DNA strand in allele b is cleaved by the intron-encoded endonuclease/reverse transcriptase. This same enzyme uses the inserted RNA as a template to synthesize the complementary DNA strand. The RNA is then degraded by cellular ribonucleases and replaced with DNA.

Telomerase is a specialized reverse transcriptase
Telomeres are specialized structures at the ends of eukaryotic linear chromosomes (see Fig. 24-9); they typically consist of many tandem repeats of a short oligonucleotide sequence. This sequence generally takes the form TxGy in one strand and CyAx in the complementary strand, where x and y range from 1 to 4 (p. 13). Telomeres vary in length from a few dozen base pairs in certain ciliated Protozoa to tens of thousands of base pairs in mammals. The TG strand is longer than its complementary strand, leaving a single-stranded DNA overhang of up to several hundred nucleotides at the 3' end.
The replication of linear chromosome ends by cellular DNA polymerases is a complex process. DNA replication requires both a template and a primer, yet beyond the terminus of a linear DNA molecule, no template exists for the annealing of An RNA primer. Without a specialized mechanism for end replication, chromosomes would progressively shorten with each cell division. The enzyme telomerase solves this problem by adding telomeric repeats to the chromosome ends.
While the existence of such an enzyme may not be surprising, its mechanism of action is quite unusual. Like several Other Enzymes described in this chapter, telomerase is a ribonucleoprotein consisting of both RNA and Protein components. The RNA component is approximately 150 nucleotides long and contains about 1.5 copies of the CyAx telomeric repeat. This region of the RNA serves as a template for the Synthesis of the TxGy telomere strand. Thus, telomerase functions within The Cell as a reverse transcriptase, catalyzing RNA-dependent DNA synthesis at its Active Site. Unlike retroviral reverse transcriptase, telomerase copies only a small stretch of RNA contained within the enzyme complex itself. The 3' end of the chromosome serves as a primer for Telomere Synthesis, which proceeds in the conventional 5' -> 3' direction. After synthesizing one repeat, the enzyme translocates to continue extending the telomere (Fig. 26-39a).
Fig. 26-39. The TG strand and T-loop in telomeres. (a) Telomerase's intrinsic RNA template binds and base-pairs with the DNA TG primer. ① Telomerase adds additional T and G residues to the TG primer, then ② translocates its internal RNA template to ③ resume The addition of T and G bases. The complementary strand is synthesized by cellular DNA polymerases (not shown). (b) Proposed structure of T-loops in telomeres. The single-stranded tail synthesized by telomerase folds back and invades the duplex region, base-pairing with the complementary sequence of the double-stranded segment. The telomere associates with several proteins, including TRF1 and TRF2 (telomere repeat binding factors). (c) Transmission electron micrograph of a T-loop at the end of a chromosome isolated from mouse hepatocytes. The length of the scale bar at the bottom corresponds to 5000 bp.

Following the extension of the TxGy strand by telomerase, the complementary CyAx strand is synthesized by cellular DNA polymerases, initiating from an RNA primer (see Fig. 25-13). In many lower eukaryotes—particularly species whose telomeres span no more than a few hundred base pairs—the single-stranded region is protected by specialized binding proteins. In higher eukaryotes (including mammals), with telomeres thousands of base pairs long, the single-stranded terminus adopts a specialized higher-order structure termed a T-loop (Fig. 26-39b). The single-stranded overhang folds back and invades the duplex DNA of the double-stranded telomeric region, likely via a mechanism analogous to the initiation of homologous genetic recombination (see Fig. 25-33). In mammals, the resulting DNA loop is bound by two proteins, TRF1 and TRF2, with the latter directly facilitating T-loop formation. T-loops protect chromosome 3' ends, rendering them inaccessible to Nucleases and double-strand break repair enzymes.
In protozoa (such as Tetrahymena), the loss of telomerase activity leads to progressive telomere shortening with each cell division, ultimately resulting in cell death. A similar correlation between telomere length and cellular Aging (replicative senescence) is observed in humans. In germline cell lineages, which maintain telomerase activity, telomere length remains constant; in somatic cells that lack telomerase, telomeres shorten over time. An inverse linear relationship exists between telomere length in cultured fibroblasts and the age of the human Donors: the older the individual, the shorter the telomeres in their somatic cells. If telomerase reverse transcriptase is experimentally introduced into human somatic cells in vitro, telomerase activity is restored, and the replicative lifespan of the cells is substantially extended.
Is the progressive shortening of telomeres the key to understanding the aging process? Does our telomere length at birth determine our lifespan? Further research in this fascinating field may yield remarkable breakthroughs.
Some viral RNAs are replicated by an RNA-dependent RNA polymerase
Several E. coli Bacteriophages, including f2, MS2, R17, and Qβ, as well as various Eukaryotic Viruses (such as Influenza virus and Sindbis virus, which causes a form of encephalitis), possess RNA genomes. The chromosomes of these viruses consist of single-stranded RNA molecules that also function as mRNAs for the synthesis of viral proteins; they are replicated in host cells by an RNA-dependent RNA polymerase (RNA replicase). All RNA-containing viruses, with the exception of retroviruses, must encode a protein with RNA-dependent RNA polymerase activity because host cells do not produce this enzyme.
The RNA replicase of most RNA-containing bacteriophages consists of four subunits and has a molecular mass of approximately 210,000. One subunit (Mr = 65,000) is the product of the replicase gene contained within the viral RNA; it contains the active site for replication. The other three subunits are host proteins normally involved in cellular Protein Synthesis: E. coli elongation factors Tu (Mr = 30,000) and Ts (Mr = 45,000) (which deliver aminoacyl-tRNA molecules to Ribosomes) and protein S1 (part of the 30S ribosomal subunit). These three host proteins help RNA replicase bind to the 3'-ends of viral RNA molecules.
RNA replicase from Qβ-infected E. coli cells catalyzes The formation of RNA complementary to the viral RNA via a reaction analogous to that of DNA-dependent RNA polymerases. The synthesis of new RNA strands proceeds in the 5' -> 3' direction; the mechanism is identical to that of all other template-directed nucleic acid synthesis reactions. RNA replicase uses RNA as a template and does not function with DNA. It lacks an independent proofreading endonuclease activity and makes roughly as many errors as RNA polymerase. Unlike DNA and RNA polymerases, RNA replicases are specific for the RNA of their own virus and generally do not replicate host cell quanta. This is why, in host cells containing numerous other RNA molecules, RNA-containing viruses are preferentially replicated.
RNA synthesis opens an important approach to the study of biochemical evolution
Structural complexity and high Organization are what distinguish living organisms from non-living systems, and these same traits are central to the manifestation of life processes. Maintaining life requires certain chemical transformations to occur very rapidly—especially those that harness external sources of energy to synthesize complex and specialized cellular macromolecules. Life depends on powerful and selective catalysts (enzymes) and on information systems capable of reliably storing copies of these enzymes and faithfully reproducing them for transmission from generation to generation. Chromosomes encode not the blueprint of the cell, but the blueprints of the enzymes that build and maintain the cell. The simultaneous need for information and catalysis presents a classic chicken-and-egg puzzle: which came first—information, which requires a specific structure, or enzymes, which are necessary to maintain and transmit information?
The elucidation of the Structural and functional complexity of RNA led Carl Woese, Francis Crick, and Leslie Orgel in the 1960s to conclude that macromolecules could serve as both information carriers and catalysts. The discovery of catalytic RNA molecules elevated this proposition from conjecture to hypothesis and fueled the widespread theory that an "RNA world" may have been pivotal in the transition from prebiotic chemistry to life (see Fig. 1-34, vol. 1). The common ancestor of all living things on our planet—capable of self-replication across generations from the dawn of life to the present day—may have been a self-replicating RNA or another polymer with equivalent chemical characteristics.

How could such a self-replicating polymer have arisen? How could it have survived in an environment lacking an Abundance of precursors for its synthesis? Starting from such a polymer, how could evolution have forged the modern DNA-protein world? Answering such profound questions requires rigorous experimentation, which may ultimately reveal how life originated and evolved on planet Earth.
The presumed origin of purine and pyrimidine bases is supported by experiments designed to test hypotheses regarding The chemical composition of the prebiotic world (pp. 55-56 in vol. 1). Beginning with simple molecules that likely populated the ancient atmosphere of our planet (CH4, NH3, H2O, H2), electrical discharges (such as lightning) drove the formation of more reactive molecules, such as HCN and aldehydes, followed by a suite of Amino Acids and organic acids (see Fig. 1-33, vol. 1). Apparently, once molecules like HCN became abundant, Purines and Pyrimidines began to synthesize. Indeed, after only a few days in a concentrated solution of ammonium cyanide heated in a reflux apparatus, up to 0.5% adenine is formed (Fig. 26-40). Adenine may have been the first and predominant nucleotide component to emerge on Earth. Interestingly, most enzyme Cofactors contain adenosine in their structure, even though it does not participate directly in their function (see Fig. 8-38, vol. 1). This facile synthesis of adenine from cyanide may point to deep evolutionary connections.
Fig. 26-40. Proposed reaction for the synthesis of adenine from ammonium cyanide in the prebiotic atmosphere. Adenine is formed from five cyanide molecules (highlighted in pink).

The "RNA world" hypothesis posits that a nucleotide polymer is capable of self-replication. Can a ribozyme catalyze its own synthesis on a template? The self-splicing intron from Tetrahymena rRNA (Fig. 26-30) catalyzes the reversible attack of a guanosine residue on the phosphodiester bond at the 5' splice site (Fig. 26-41). If the 5' splice site and the internal guide sequence are removed from the intron, the remaining intron can bind RNA strands base-paired to short oligonucleotides. Part of the remaining intact intron effectively acts as a template for aligning and ligating short oligonucleotides. In essence, this is the reverse of the reaction in which guanosine attacks the 5' splice site, resulting in the template-directed synthesis of short RNA sequences.
Fig. 26-41. RNA-dependent synthesis of an RNA polymer from oligonucleotide precursors. (a) In the first step of removing the group I self-splicing intron from the Tetrahymena rRNA precursor, a guanosine residue reversibly attacks the 5' splice site. Only the P1 region is shown in detail, encompassing the internal guide sequence (boxed) and the 5' splice site; the rest of the ribozyme is depicted as a gray-green globule. The complete Secondary structure of the ribozyme is shown in Fig. 26-30. (b) If the P1 region is excised (indicated by the dark green cavity), the ribozyme retains both its three-dimensional structure and its catalytic capacity. A novel RNA molecule introduced in vitro can bind to the ribozyme in the same manner as the internal guide sequence of P1 in (a). This provides a template for subsequent RNA polymerization reactions involving base pairing between the introduced RNA and complementary oligonucleotides. The ribozyme can ligate these oligonucleotides in a reaction the reverse of that shown in (a). Although only one such reaction cycle is illustrated in (b), repeated rounds of binding and catalysis can drive the RNA-dependent synthesis of long RNA polymers.

Box 26-3. PRACTICAL BIOCHEMISTRY. The SELEX Method for Generating RNAs with Specific Properties
SELEX (Systematic Evolution of Ligands by Exponential Enrichment) is used to select aptamers—oligonucleotide sequences that bind tightly to a specific molecular target. Automated application of this technology allows the identification of one or more aptamers with desired binding specificities.
Figure 1 illustrates how the SELEX technology is used to select RNA sequences that bind to ATP. In step ①, a random pool of RNA sequences is subjected to "artificial Selection" by passing it through a resin bearing immobilized ATP. In practice, researchers work with a pool containing approximately 1015 different sequences, which corresponds to the complete set of sequences of 25 nucleotides (425 ≈ 1015). For longer RNA molecules, the initial RNA pool does not necessarily contain all possible sequences. ② RNAs that fail to bind to the Column are discarded; ③ molecules that bind to ATP are eluted with a salt solution and collected. ④ The selected RNAs are amplified using reverse transcriptase to generate a large yield of complementary DNA molecules, after which RNA polymerase is used to synthesize RNAs complementary to those DNA molecules. ⑤ The new RNA pool is subjected to the same selection Procedure, and the entire cycle is repeated 10 or more times. Ultimately, only a few aptamers remain (RNA sequences exhibiting pronounced affinity for ATP).
Fig. 1. The SELEX method.

The structural elements required for ATP binding are shown in Fig. 2. Such molecules bind ATP (and other adenosine nucleotides) with Kd < 50 μM. Figure 3 illustrates the three-dimensional structure of the 36-nucleotide cAMP aptamer complex obtained by SELEX. The schematic in Fig. 2 corresponds precisely to this RNA aptamer.
Fig. 2. SCHEMATIC STRUCTURE OF an RNA aptamer with affinity for ATP. Labeled nucleotides are essential for binding RNA to ATP.

Beyond investigating RNA functions, the SELEX method holds major practical significance for identifying short pharmaceutical RNA molecules. Naturally, aptamers cannot be generated to bind specifically to every potential target. Nonetheless, this technology for the rapid selection and Amplification of specific oligonucleotide sequences from a highly complex pool provides a powerful route to novel therapeutic agents. For example, one can select RNAs that bind tightly to protein receptors protruding from cell surfaces, such as those on tumor cells. By blocking receptor activity or delivering a toxin linked to the aptamer into the tumor cell, one can eliminate tumor cells. The SELEX method has also been used to select DNA aptamers for the detection of anthrax spores, and other exciting Applications of this technology are currently being explored. ▪
Fig. 3. (Adapted from PDB ID 1RAW.) Complex of the RNA aptamer with AMP. In the conserved nucleotide sequences (forming the pocket where AMP binds), the bases are shown in gray, and the bound AMP is shown in red.

Self-replicating macromolecules rapidly consume precursors, which under prebiotic atmospheric conditions were supplied by slow processes. Thus, for the efficient synthesis of precursors as early as the primitive stages of evolution, metabolic pathways were required in which ribozymes likely acted as catalysts. Known ribozymes today have limited functions, and no trace remains of those "ancient" ribozymes whose existence can only be inferred. To further investigate the "RNA world" hypothesis, we must determine whether RNA is capable of catalyzing the reactions that took place in primitive metabolic pathways.
The discovery of RNA molecules with novel catalytic functions was aided by a rapid screening method that makes it possible to identify and isolate sequences within oligonucleotide mixtures possessing specific activities. This technique is known as SELEX technology—systematic evolution of ligands by exponential enrichment (see Box 26-3). The method has been used to generate RNA molecules that bind to amino acids, organic Dyes, nucleotides, cobalamin, and other molecules. Researchers have isolated ribozymes that catalyze the formation of ester and amide bonds, SN2 reactions, the metallation of Porphyrins (the insertion of Metal Ions into molecules), and carbon-carbon bond formation. Through the evolution of enzyme cofactors featuring a nucleotide "handle" to facilitate binding to the ribozyme, the range of accessible chemical processes hypothesized to occur in primitive systems has expanded.
As we will see in the next chapter, the Synthesis of Peptide Bonds is catalyzed by certain natural RNA molecules; therefore, it seems entirely fair to state that the RNA world was transformed thanks to the high catalytic potential of proteins. Protein synthesis could have been a breakthrough in the Evolution of the RNA world, but it might also have accelerated its demise. The Role of information carrier could have shifted from RNA to DNA because DNA is chemically more stable. RNA replicase and reverse transcriptase are likely today's counterparts of enzymes that once played a crucial role in effecting the transition to the modern DNA-based biosystem.
Molecular parasites could also have originated from the RNA world. With the advent of the first inefficient self-replicating molecules, transposition emerged as a potentially viable alternative to replication as a strategy for successful reproduction and survival. Ancient parasitic RNA molecules could simply have integrated into a self-replicating molecule via catalytic transesterification and subsequently undergone passive replication. Natural selection could have driven transposition to become site-specific, targeting sequences that did not interfere with the catalytic activity of the host RNA. Replicators and RNA transposons may have coexisted in a primitive symbiotic relationship, each driving the evolution of the other. Modern introns, retroviruses, and transposons may well be remnants of this ancient parasitic RNA "freeloader" journey. To this day, these elements continue to shape the evolution of their hosts.
Box 26-4. The Expanding RNA World, or Transcripts of Unknown Function
Throughout this book, modern estimates regarding the number of genes in The Human Genome and many other organisms have been cited repeatedly. These estimates assume that scientists recognize a gene if they can "see" it based on current paradigms of DNA, RNA, and proteins. But how valid is this approach?
As noted in Chapter 9, less than 2% of the human genome appears to encode proteins. Even when introns are taken into account, only a small fraction of The Genome is transcribed into RNA, primarily mRNA encoding these proteins. The rest of the genome is sometimes referred to as junk DNA. However, the term "junk" merely reflects a gap in our knowledge, as we are now gradually coming to realize that much of this genomic DNA is fully functional.
In attempts to map the human transcriptome more precisely, scientists have developed novel tools to identify genomic sequences transcribed into RNA with high accuracy. The research results have been unexpected. It turns out that a significantly larger portion of our genome is transcribed than previously thought. Much of the resulting RNA appears to be non-protein-coding. Many types of RNA lack the basic structural elements (such as 3'-poly(A) tails) that characterize mRNA. What, then, is the purpose of this RNA?
Most methods applied to address these questions fall into two major categories: cDNA cloning and microarray analysis. The Construction of cDNA libraries to study the transcribed genes of a eukaryotic genome is described in Chapter 9 (see Fig. 9-14 in Vol. 1). However, classical cDNA preparation methods often result in the cloning of only a portion of a given transcript's sequence. Because reverse transcriptase can stall at regions with specific mRNA secondary structures or simply dissociate, full-length DNAs typically account for no more than 20% of clones in a cDNA library. This complicates The Use of such libraries for mapping transcription start sites (TSS) and studying the regions of genes that encode the N-terminal sequences of proteins. Figure 1 illustrates one of the numerous methods developed to overcome these limitations. Advanced techniques now make it possible to construct cDNA libraries in which full-length clones exceed 95%, significantly enhancing our ability to assess cellular RNA composition. Nonetheless, cDNAs are conventionally generated from RNA transcripts containing poly(A) tails. The application of microarray analysis combined with poly(A)-independent cDNA preparation methods (Fig. 2) has revealed that a substantial fraction of RNA in Eukaryotic cells lacks these characteristic terminal structures.
Figure 1. Strategy for cloning full-length cDNAs. An mRNA pool is isolated from a tissue sample. In some cases, mRNA bound to a specific protein can be obtained by immunoprecipitation of that protein followed by Isolation of the associated mRNA. For example, biotin is covalently attached to the 5' ends of mRNA molecules exploiting the unique Properties of the 5' cap. To initiate reverse Transcription of the mRNA molecules, a poly(dT) primer is used. RNase I digests RNA that is not part of DNA-RNA hybrids, thereby destroying incomplete cDNA-RNA pairs. Full-length cDNA-RNA hybrids are isolated using streptavidin-coated beads (which bind biotin), converted into double-stranded DNAs, and cloned.

The complete picture is not yet fully clear, but several Conclusions can already be drawn. If repetitive sequences (such as transposons), which make up as much as half of the mammalian genome, are excluded, at least 40% (and potentially the vast majority) of the remaining genomic DNA is transcribed into RNA. There appear to be more species of RNA lacking a poly(A) tail than those possessing one. A large proportion of such RNAs is not exported to the Cytoplasm but is instead retained exclusively in the nucleus. Many genomic regions are transcribed from both strands, yielding transcripts that are complementary to one another (referred to as antisense transcripts). Often, antisense RNAs are involved in regulating the RNAs to which they base-pair. Many types of RNA are produced only in one or a few Tissues, with novel transcripts continually uncovered as new tissues are analyzed. Thus, to date, the complete transcriptome of no Organism has been fully mapped. Furthermore, new types of RNA are frequently transcribed from genomic regions (such as those illustrated in Fig. 9-20) that are syntenic across multiple organisms (meaning genes maintain a conserved chromosomal arrangement). This evolutionary conservation points to an essential function for these RNAs.
Among the novel classes of RNA whose existence has come to light only over the past 20 years are snRNA, snoRNA, and microRNA. New transcription start sites (TSS) have been identified, novel classes of RNA molecules are emerging, and new alternative splicing variants have been discovered. All these findings are reshaping our understanding of genes. In the human and mouse genomes, even the number of protein-coding mRNA transcripts may prove to be far greater than previously estimated, and the tally of identified protein genes may likewise increase. Nevertheless, the functions of most newly discovered transcripts remain unknown; they are simply designated as transcripts of unknown function (TUFs). The discovery of this vast realm of novel RNAs raises hopes that we will gain a deeper understanding of Eukaryotic Cell mechanics and perhaps a clearer picture of the primordial RNA world of the distant past.
Figure 2. Microarray Analysis of the transcriptome. (a) Genomic microarrays are synthesized using non-repetitive genomic regions. Adjacent oligonucleotides on the microarray overlap in sequence such that any individual nucleotide (e.g., T, highlighted in blue) may be represented multiple times. (b) A tissue sample is fractionated to separate nuclei from the cytoplasm, and RNA is isolated from each fraction. RNA samples containing a poly(A) tail are separated from those lacking it by passing the RNA through a poly(dT)-conjugated column. Poly(A)-containing RNA is converted into cDNA (Fig. 1). Using random primers, poly(A)-minus RNA is also converted into cDNA. Although not all resulting DNA fragments match the exact length of their template RNAs, the aggregate DNA pool encompasses the majority of the original RNA sequences. The cDNAs are then labeled and used as probes on microarrays. Fluorescent signals on the microarray reveal the transcribed RNA sequences.

The "RNA world" remains a hypothesis with several unresolved questions, yet experimental evidence continues to support a growing body of core tenets within this theory. Future experiments will further broaden our understanding of biological origins. Crucial missing pieces of this puzzle will come from fundamental chemistry, cell biology research, and perhaps even studies on other planets. Meanwhile, the RNA world continues to expand (Box 26-4).
Summary of Section 26.3 RNA-Dependent Synthesis of RNA and DNA
■ RNA-dependent DNA polymerase (reverse transcriptase) was first discovered in retroviruses, which must convert their RNA genome into double-stranded DNA during their life cycle. These enzymes transcribe viral DNA from RNA, a process widely harnessed to generate complementary DNA.
■ Many eukaryotic transposons are related to retroviruses, and their transposition proceeds through an RNA intermediate.
■ Telomerase—the enzyme responsible for synthesizing the telomeric ends of linear chromosomes—is a specialized reverse transcriptase that carries its own internal RNA template.
■ RNA-dependent RNA polymerases, such as bacteriophage RNA replicases, exhibit high Specificity for the viral RNA used as their template.
■ The discovery of catalytic RNA molecules and metabolic pathways for the interconversion of RNA and DNA suggests that a pivotal milestone in evolution was The Emergence of an RNA molecule (or an equivalent polymer) capable of catalyzing its own replication. The biochemical potential of RNA molecules can be explored using SELEX, a technique that allows the rapid selection of RNA sequences with specific binding affinities or distinct catalytic activities.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.