Fundamentals of Biochemistry - Filippovich Yu. B. 1999
Nucleic Acids and Their Metabolism
Structure and Properties of Nucleic Acids
The Chemistry of Nucleic acids has been developing at an exceptionally rapid pace in recent years, driven by a fundamental reassessment of the Biological Role of these compounds. The view of Selection/9.html">Nucleic Acids AS inert Structure/83.html">Structural elements of The Nucleus and Cytoplasm has been abandoned for good, as they play a crucial role in directing the specific synthesis of Biopolymers in humans, animals, plants, and microorganisms.
Nucleic acids are high-molecular-weight compounds with a defined elemental composition that break down upon Hydrolysis into purine and pyrimidine bases, pentose, and phosphoric acid. A particularly characteristic feature of nucleic acids is their high content of P (8–10%) and N (15–16%). They also contain C, H, and O.
Nucleic acids were first isolated by F. Miescher over a century ago (1869) from pus Cell nuclei as a protein complex called nuclein (from Latin nucleus), and the term itself was proposed by A. Kossel in 1889. By the end of the 19th century, R. Altman (1899) had obtained them in a protein-free state from animal Tissues and Yeast, and in 1936 A. N. Belozersky and co-workers isolated them from plant material.
Isolation of nucleic acids. Most nucleic acids in plant, animal, and bacterial Cells are bound to Proteins. Therefore, In addition to disrupting cell walls by homogenizing the biological material, isolating nucleic acids requires breaking the bonds between the nucleic acid and the protein. This is achieved by treating the material with a concentrated salt solution, such as a 10% NaCl solution, which simultaneously extracts the nucleic acids. After removing the solid residue, the nucleic acids are precipitated from a solution chilled to 0° C using ethanol or a trichloroacetic acid solution. The nucleic acid precipitate is separated by centrifugation, thoroughly washed, and dried.
Alternatively, the phenol method is used to isolate nucleic acids, allowing for the recovery of native preparations. To do this, tissue minced in a homogenizer under cooling is mixed with a Water-saturated phenol solution, shaken vigorously for 1 hour, and centrifuged. The Contents of the centrifuge tube separate into four distinctly defined layers differing in consistency and color. The top aqueous layer and the underlying viscous white layer contain the bulk of the nucleic acids. The third layer—a transparent, yellowish, gelatinous phase—contains phenol with dissolved proteins. The fourth and bottommost brown layer contains tissue debris, denatured proteins, and trace amounts of nucleic acids. Because this method effectively removes a significant portion of proteins, the Procedure is known as phenol deproteinization.
Nucleic acid preparations can also be freed from protein by treating their salt extract with a double volume of chloroform containing a small amount of isoamyl alcohol. After thorough mixing for 15 minutes to form a stable emulsion, the mixture is centrifuged, and the upper aqueous phase containing the nucleic acids is separated (denatured protein remains at the interface between the aqueous and chloroform layers). The nucleic acids are then precipitated from the aqueous phase using a double volume of chilled ethanol.
The crude nucleic acid preparations obtained in this manner are further fractionated using more sophisticated techniques, such as chromatography (including Affinity Chromatography), agarose and Sephadex Gel filtration, partition in two-phase polymer systems, ultracentrifugation, and Electrophoresis. Most of these Methods are covered in Chapter II and differ only in minor details when applied to nucleic acids. Ultimately, preparations of individual nucleic acids are obtained.
Extraction from biological sources is no longer the sole method for obtaining nucleic acids; chemical synthesis has also become widespread. The Development of chemical and engineering approaches has led to automated nucleic acid synthesizers characterized by high performance and reliability. For instance, an LKB synthesizer model (Sweden) enables the automated assembly of nucleic acid molecules up to 160 units long, with each cycle taking just ten minutes.
Chemical synthesis of nucleic acids is especially vital for producing genes, their fragments, regulatory regions, and the like. The first Gene synthesis—for Alanine transfer ribonucleic acid—was accomplished in 1972 by H. Khorana. Today, the number of synthesized genes has reached several dozen. In the USSR, researchers synthesized the human interferon a2 gene (M. N. Kolosov et al., 1982), the valine tRNA gene (M. N. Kolosov et al., 1983), the human Calcitonin gene (A. A. Bayev et al., 1985), A number of promoters, and others.
Chemical composition. When heated with perchloric acid, nucleic acids break down into the structural units that make up their giant macromolecules. Other acids, such as HCl, cause severe Degradation of Nucleic acids accompanied by the release of NH3, indicating the destruction of their constituent structural elements.
The structural Building Blocks of nucleic acids include pyrimidine bases, purine bases, CARBOHYDRATES, and phosphoric acid.
Pyrimidine bases are derivatives of the heterocyclic compound pyrimidine:
Class="center">
The structural formula shown here, depicting a six-membered ring with alternating double bonds, is, much like in the case of benzene, merely a historical convention. In reality, as can be deduced by comparing interatomic distances, the pyrimidine molecule contains neither typical double nor typical single bonds; instead, there is an interaction of the n-electrons across all atoms constituting the ring. A measure of this n-electron interaction is the so-called bond order, which characterizes the conjugation strength of n-electrons between two adjacent atoms. For a typical double bond, the n-electron interaction strength—i.e., the bond order—is taken as unity. Due to the delocalization of n-electrons in molecules with conjugated double bonds, such as the pyrimidine molecule under consideration, bond orders take fractional values. The higher the bond order, the greater its propensity for addition reactions.
The nucleic acid composition includes the following pyrimidine derivatives (pyrimidine bases): cytosine, uracil, thymine, 5-methylcytosine, and 5-hydroxymethylcytosine:

Cytosine, uracil, and thymine are present in nucleic acids in significant amounts, whereas 5-methylcytosine and 5-hydroxymethylcytosine are found in trace quantities and by no means universally. Consequently, they are referred to as minor (exotic) bases. By analogy with rare Amino Acids in proteins, they could be described as bases that occur only occasionally within nucleic acids. In recent years, the list of minor pyrimidine bases discovered in nucleic acids has grown (Table 15).
The purine bases of nucleic acids are derivatives of the bicyclic heterocycle purine:

Here, as with the pyrimidine ring, the placement of single and double bonds in the formula is conventional. Both interatomic distances and bond orders in the purine molecule point to a high degree of n-electron conjugation among the C and N atoms that form the purine ring.
Nucleic acid hydrolysates invariably reveal two purine derivatives—adenine and guanine:

In addition, nucleic acids contain A large number of minor purine bases, which are methylated derivatives of adenine and guanine (Table 15):

An important feature of hydroxy derivatives of pyrimidine and purine is their capacity for tautomeric (lactam-lactim) transformations, for example:

Due to this, in particular, pyrimidine bases interact via the 1st nitrogen atom of the lactam form with the carbohydrates found in nucleic acids.
The carbohydrate component of nucleic acids is represented by two very similar right-handed Monosaccharides: ribose and deoxyribose. In the free state, these monosaccharides exist in all possible tautomeric forms arising from ring-chain Tautomerism. Within
nucleic acids, both monosaccharides are in the $\beta$-D-ribofuranose form:

Compared to $\beta$-D-ribose, the second monosaccharide ($\beta$-D-2-deoxyribose) is a compound reduced at the 2nd carbon atom. Since the reduction process involves the removal of a hydroxyl group, the resulting derivative is termed deoxyribose, with the numeral 2 indicating THE POSITION OF the ribose carbon atom where the hydroxyl group is replaced by an H atom.
It has recently been clarified that ribose and deoxyribose are not the only carbohydrates found in nucleic acids: glucose has been discovered in a number of phage DNAs and the DNA of certain Cancer cell types.
The Study of nucleic acid hydrolysis products led to an important Conclusion: the composition of hydrolysis products obtained from nucleic acids isolated from various sources is not uniform. This was first discovered when comparing the Composition of Nucleic acids isolated from the calf Thymus (thymonucleic acid) and yeast (yeast nucleic acid). Subsequently, it was demonstrated that they correspond to two Types of Nucleic acids differing in composition, structure, and Functions. In accordance with The Nature of their carbohydrate component, one of them was named deoxyribonucleic acid (DNA), and the other, ribonucleic acid (RNA). Between DNA and RNA, there are other Similarities and differences in composition (Table 15).
Table 15 Composition of Nucleic Acids
Chemical compound |
DNA |
RNA |
|
Purine bases |
Adenine |
Adenine |
Guanine |
Guanine |
|
Pyrimidine bases |
Cytosine |
Cytosine |
Thymine |
Uracil |
|
Carbohydrates |
Deoxyribose Glucose (sometimes) |
Ribose |
Inorganic substance |
Phosphoric acid |
Phosphoric acid |
Minor bases |
||
Purine |
N6-Methyladenine |
N6-Methyladenine |
1-Methylguanine |
N6-Dimethyladenine |
|
3-Methylguanine |
1-Methyladenine |
|
7-Methylguanine |
2-Methyladenine |
|
N2-Methylguanine |
2-Methylthio-N6-isopentenyladenine |
|
N2-Dimethylguanine |
N2-Methylguanine |
|
N2-Dimethylguanine |
||
1-Methylguanine |
||
7-Methylguanine |
||
Pyrimidine |
5-Methylcytosine 5-Hydroxymethylcytosine Hydroxymethyluracil Uracil |
5-Methylcytosine 5-Hydroxymethylcytosine N4-Methylcytosine 3-Methylcytosine 3-Methyluracil Thymine 5-Methylaminoethyl-2-thiouracil Dihydrouracil |
As can be seen from the data in Table 15, DNA and RNA also differ in the qualitative composition of their pyrimidine bases: the former is characterized by the presence of thymine, and the latter by uracil. The differences in minor purine and pyrimidine bases of DNA and RNA are particularly striking, with RNA containing a much richer set (over 50) of them.
A trend has emerged linking the presence and distribution of minor methylated bases in DNA and RNA to several crucial functions of nucleic acids: their interaction with proteins, including a number of Enzymes; the coding and transmission of information regarding macromolecule Biosynthesis; participation in memory mechanisms and Organism Aging; and The regulation of nucleic acid biosynthesis, among others.
Molecular weight, content, and cellular localization of DNA and RNA; types of DNA and RNA. The Molecular Weight of DNA is determined primarily by hydrodynamic and electron microscopic methods, although it can also be measured via light scattering of DNA solutions and by several other techniques.
The hydrodynamic method is based on the linear dependence of the DNA sedimentation coefficient, determined by ultracentrifugation of DNA solutions, on its molecular weight, which can be established using a calibration curve or calculated by the formula: $0.445 \lg M = 1.819 + \lg(s^{\circ}_{20,\omega} - 2.7)$, where $s^{\circ}_{20,\omega}$ is the sedimentation coefficient extrapolated to infinite dilution ($s^{\circ}$), standard Temperature ($20^{\circ}\text{C}$), and water viscosity ($\omega$).
The electron microscopic METHOD FOR DETERMINING the molecular weight of DNA is based on measuring the length of extended DNA molecules. It is known that a 0.1 nm stretch of the DNA molecule corresponds to a mass of 197 Da. Multiplying this value by the experimentally determined length of the DNA molecule yields its molecular weight.
Isolation of native DNA from Eukaryotic cells is exceptionally difficult, as appropriate methods to avoid shearing the extracted DNA molecules have not yet been developed. Therefore, reliable figures have been obtained only for viral and phage DNAs (Table 16); from these, DNA is extracted more easily by gently removing the protein coat.
As seen from Table 16, the molecular weights of viral and phage DNAs are measured in tens and hundreds of millions of Daltons. It can be assumed that the molecular weight of eukaryotic DNA is significantly higher. This is evidenced by the molecular weight of DNA isolated with necessary precautions from the largest chromosome of the fruit fly, Drosophila. The entire chromosome DNA is represented by a single molecule with $M = 40 \times 10^{9}$.
The DNA content in cells of a given species exhibits remarkable constancy, whereas interspecies differences in this parameter are quite large. The amount of DNA in a cell is measured in picograms ($10^{-12}\text{ g}$) and ranges from 0.01 pg in *Escherichia coli* to several picograms in haploid cells of higher organisms.
Table 16 Molecular Weights of Viral and Phage DNAs
|
Object |
Molecular weight, million Daltons |
|
Hydrodynamic method |
Electron microscopic method |
|
Bacteriophage fd |
1.9 |
— |
Polyoma virus |
— |
3.2 |
Adenovirus |
21 |
24 |
Bacteriophage T7 |
23–28 |
22–28 |
» T5 |
66 |
67 |
» T2 |
123 |
105–119 |
» T4 |
111–131 |
116–152 |
Depending on its localization within The Cell, nuclear, mitochondrial, chloroplast, centriolar, and episomal DNA are distinguished. Nuclear DNA in eukaryotes sharply predominates over the DNA of other subcellular structures. For instance, Mitochondria contain from $0.5 \times 10^{-16}$ to $5 \times 10^{-16}\text{ g}$ of DNA, METABOLISM/14.html">Chloroplasts from $10^{-16}$ to $150 \times 10^{-16}\text{ g}$, and centrioles $2 \times 10^{-16}\text{ g}$, which accounts for a few percent of the nuclear DNA. A similar ratio exists between the DNA content of the bacterial chromosome and episomes—extrachromosomal, autonomously replicating hereditary determinants in microorganisms that ensure The transfer of Genetic information, such as Antibiotic Resistance (otherwise known as R-factors or resistance factors). The existence of extrachromosomal DNA, transport or communication DNA, cytoplasmic membrane DNA, and finely dispersed supercoiled DNA is also under Discussion. By functional purpose, ribosomal DNA (rDNA) and satellite DNA (sDNA) are distinguished.
In addition to intracellular DNA, there is also DNA that makes up Viruses and Phages. Its amount in viruses and phage particles is significantly lower than in bacterial cells (thousandths of a picogram and less).
The molecular weights of RNA are determined by the same methods as those for DNA; however, Polyacrylamide gel electrophoresis is additionally employed, since the migration distance of RNA in the gel is inversely proportional to its molecular weight. Regarding the content and localization of RNA in cells, it is neither uniform nor stable: in cells undergoing intensive Protein Biosynthesis, the RNA content is several times higher than that of DNA (for example, rat Liver contains 4 times more RNA than DNA), but where Protein Synthesis is low, the DNA-to-RNA ratio may be reversed (for example, rat Lungs contain twice as little RNA as DNA).
Based on their functional significance and molecular weights, as well as their localization within the cell contents, RNAs are divided into the following types.
1. Transfer RNAs (tRNAs) are characterized by relatively low molecular weights (17,000–35,000) and are localized in the hyaloplasm of the cell, nuclear sap, and the unstructured part of chloroplasts and mitochondria. They encode Amino Acids and transport them to the cellular ribosomal apparatus during protein biosynthesis.
2. Ribosomal RNAs (rRNAs) are characterized primarily by high molecular weights (550,000–700,000 for RNAs of 30–40S ribosomal subunits; 1.1 ∙ 106 – 1.7 ∙ 106 for RNAs of 50–60S ribosomal subunits; but 40,000 for 5S RNA and ~ 50,000 for 5.8S RNA from Ribosomes); they are localized in ribosomes, serving as their structural foundation and performing diverse functions within them.
3. Messenger, or template, RNAs (mRNAs) possess molecular weights varying over a wide range (from 300,000 to 4 ∙ 106). Arising in the form of high-molecular-weight precursors in the Cell Nucleus or on the DNA of other subcellular particles, mRNAs (in the form of ribonucleoproteins) migrate to the ribosomes; within the latter, they perform a template function during the assembly of polypeptide chains.
4. Viral RNAs are distinguished by diverse and high molecular weights, lying mainly within the range of several million daltons. They constitute Structural components of viral and phage ribonucleoproteins, carrying all the information necessary for viral reproduction in host cells.
The current literature discusses the feasibility of classifying several additional types of RNA into distinct categories: nuclear, chromosomal, mitochondrial, low-molecular-weight regulatory, and antisense.
STRUCTURE OF THE structural elements of nucleic acids. Upon mild Treatment of RNA with an aqueous alkali solution (e.g., 1 N NaOH or KOH for 18 hours at room temperature), they decompose into structural units, each consisting of a purine or pyrimidine base, ribose, and a phosphoric acid residue. This breakdown process occurs even more efficiently under the action of specific biocatalysts, such as Ribonuclease. DNAs, in contrast to RNAs, are resistant to dilute alkaline solutions at normal temperatures. Therefore, the hydrolysis of DNA into structural units—each composed of a purine or pyrimidine base, deoxyribose, and phosphoric acid—can be achieved only in the presence of a special biological catalyst: deoxyribonuclease. Acids, both mineral (HCl, HClO4, etc.) and organic (HCOOH, etc.), are unsuitable for this type of hydrolysis because they degrade both RNA and DNAS into free purine and pyrimidine bases, carbohydrate, and phosphoric acid. The constituent structural units of nucleic acids were first isolated from their hydrolysates in 1908 by P. Levene and J. Mandel, who named them NUCLEOTIDES. Hydrolysis of RNA yields ribonucleotides, whereas hydrolysis of DNA yields deoxyribonucleotides.
Purine or pyrimidine bases, ribose or deoxyribose, and phosphoric acid are linked in nucleotide molecules in an entirely uniform manner. The chemical Structure of Nucleotides along with their full and abbreviated names are as follows:


The names of individual ribo- and deoxyribonucleotides are derived from the characteristic purine or pyrimidine base, while the presence of deoxyribose is indicated by the prefix "deoxy". Since the phosphoric acid residue in natural nucleotides is attached to the ribose (or deoxyribose) residue via the hydroxyl group at the 3rd or 5th carbon atom, Two Types of nucleotides exist:

To indicate the position of the phosphoric acid residue in the nucleotide molecule, the carbohydrate carbon atoms in the ribose residue are numbered. To avoid confusion with the corresponding numbering of atoms in purine or pyrimidine bases, these digits are marked with a prime symbol ("'").
The configuration of one of the nucleotide molecules is shown in Fig. 65. As can be seen from the figure, the projection formula of the cytidine-3'-phosphate molecule (as well as other nucleotides), drawn taking into account the true conformation of the molecule, differs significantly from the conventional projection formula. This must be taken into account when examining the Structure of Nucleic Acids.
Upon the removal of the phosphoric acid residue from a nucleotide, an even simpler compound—a nucleoside—is obtained. This term was first proposed by P. Levene and W. Jacobs in 1909 to designate carbohydrate derivatives of Purines isolated from RNA. Subsequently, it was extended to compounds of carbohydrates with Pyrimidines. The names of nucleosides are formed from the names of the purine or pyrimidine bases and the corresponding endings. In the structural formulas of nucleotides given above, one can easily find the parts corresponding to nucleoside residues, and in the names of nucleotides, distinguish the part denoting the nucleoside residue. Consequently, nucleotides are phosphoric esters of nucleosides. Everything stated above fully applies to deoxynucleotides and deoxynucleosides.
Since nucleoside phosphates are fairly strong acids, they are frequently referred to as cytidylic, uridylic, adenylic, and guanylic acids.
The mononucleotides characteristic of each type of nucleic acid, when combined in quantities of several hundred or sometimes thousands into a single molecule, form massive polynucleotide chains (see below).
Thus, in terms of their chemical structure, nucleic acids are polyribonucleotides (RNA) and polydeoxyribonucleotides (DNA). The joining of nucleotide residues in RNA and DNA molecules is accomplished in the same way: via ester bridges formed between pairs of nucleotides by phosphoric acid residues. The latter are always linked to the 3rd carbon atom of ribose (or deoxyribose) of one nucleotide residue and to the 5th carbon atom of ribose (or deoxyribose) of another, as can be seen from the depictions of DNA and RNAS molecule fragments given above. If one mentally completes The structure of the DNA and RNA fragments depicted above with terminal nucleotides, the 5'-carbon atom of the deoxyribose residue in DNA (or the ribose residue in RNA) will bear a phosphoric acid residue, while the opposite end of the chain at the 3'-carbon atom of the deoxyribose (or ribose) residue will bear a hydroxyl group. These nucleotide residues form the 5' and 3' ends of DNA and RNA molecules. The first, i.e., the 5'-terminal nucleotide, is conventionally considered the beginning of the nucleic acid molecule, and the second, i.e., the 3'-terminal nucleotide, its end.

Fig. 65. Projection formula of the cytidine-3-phosphate molecule:
distances between atoms are given in nanometers; hydrogen atoms are not labeled

Structure of DNA and RNA molecule fragments
A characteristic feature of the combination of nucleotide residues into polynucleotide chains is that the hydroxyl at the 2nd carbon atom of ribose is never involved in The formation of an ester bridge with the phosphoric acid residue during nucleic acid synthesis. Consequently, in both DNA and RNAS, bonds via the phosphoric acid residue exist exclusively from the 3rd to the 5th carbon atom of the carbohydrate. No branching of the chain has been detected in either DNA or RNA.
Depending on the molecular weight of the nucleic acid, the polycondensation coefficient, i.e., the number of nucleotide units linked into a single polynucleotide chain, varies over a wide range. Assuming the average molecular weight of a nucleotide residue to be 330, the polycondensation coefficient for DNA molecules is expressed in tens and hundreds of thousands of units, whereas for RNA it is merely in the tens, hundreds, and in some cases thousands of units.
Nucleotide Composition of DNA and RNA. Tables 17 and 18 present selected Examples of the ratios of purine and pyrimidine bases in the DNA and RNA of certain organisms.
Table 17 Nucleotide composition of DNA
|
Source of nucleic acid |
Molar base ratios, % |
G+A C+T |
G+C A+T |
|||
G |
A |
C |
T |
|||
Animals |
||||||
Mouse |
21.9 |
29.7 |
22.8 |
25.6 |
1.07 |
0.81 |
Octopus |
17.6 |
33.2 |
17.6 |
31.6 |
1.03 |
0.54 |
Plants |
||||||
Wheat |
23.8 |
25.6 |
24.6 |
26.0 |
0.98 |
0.94 |
Onion |
18.4 |
31.8 |
18.2 |
31.3 |
1.02 |
0.58 |
Microorganisms |
||||||
Tubercle bacillus |
34.2 |
16.5 |
33.0 |
16.0 |
1.03 |
2.08 |
Streptococcus |
16.6 |
33.4 |
17.0 |
33.0 |
1.00 |
0.51 |
Table 18. Nucleotide composition of total RNA
|
Source of nucleic acid |
Molar base ratios, % |
G+A |
G+C |
|||
G |
A |
C |
U |
C+U |
A+U |
|
Animals Rat |
32.8 |
18.7 |
29.6 |
18.9 |
1.06 |
1.66 |
Mosquito |
23.5 |
30.6 |
19.7 |
26.1 |
1.18 |
0.76 |
Plants |
||||||
Wheat |
30.8 |
25.2 |
25.4 |
18.6 |
1.27 |
1.28 |
Pine |
27.6 |
25.3 |
23.4 |
23.7 |
1.12 |
1.04 |
Microorganisms |
||||||
Tubercle bacillus |
33.0 |
22.6 |
26.1 |
18.3 |
1.25 |
1.45 |
Staphylococcus |
28.7 |
26.9 |
22.4 |
22.0 |
1.25 |
1.05 |
The letters denote the respective bases: G — guanine, A — adenine, C — cytosine, T — thymine, U — uracil, while the figures are expressed in molar percentages (mol. %), i.e., as a percentage of the total amount of nitrogenous bases found in the nucleic acid, with the content of each base calculated in moles or mole fractions within the nucleic acid preparation.
An analysis of these data led E. Chargaff to formulate a set of rules (Chargaff's rules):
1. In DNA, the molar sum of G and A (purine bases) is equal to the molar sum of C and T (pyrimidine bases). This regularity is not characteristic of RNA, where The ratio of purine to pyrimidine bases varies over a wide range.
2. In DNA molecules, the number of A residues is always equal to the number of T residues. G and C maintain the same relationship. This is not observed in RNA molecules, although in many cases the molar base ratios are close (here, T and U in DNA and RNA molecules are considered equivalent).
3. The ratio of the sum of the molar concentrations of G and C to the sum of the molar concentrations of A and T in DNA (and A and U in RNA)
varies significantly between the two types of nucleic acids. The range of variation for this parameter is particularly broad in DNA.
The first two regularities in the structure of DNA and RNA are associated with Specific features of their Secondary structure. Just as Hydrogen Bonds between —CO- and —NH- groups play a decisive role in stabilizing the helical conformation of protein molecules, hydrogen bonds formed between Base Pairs—A and T (A and U in RNA) and G and C (in both types of nucleic acids)—are of paramount importance in establishing the Introduction/11.html">Secondary structure of the polynucleotide chain. Bases that form pairs linked by Hydrogen bonds are termed complementary.
The Nature of the hydrogen bonds between complementary bases is illustrated in Fig. 66. On the one hand, hydrogen bonds are formed between amino and keto groups occupying specific positions in the purine and pyrimidine base molecules (the 6-amino group of A and the 4-keto group of T or U, as well as the 2-amino and 6-keto groups of G, and the 2-keto and 4-amino groups of C). This constitutes the complementarity, or chemical complementarity, of purine and pyrimidine bases: wherever an amino group is located on a purine base, a keto group is found on the pyrimidine base, and vice versa. On the other hand, hydrogen bonds are formed through interactions involving the N atoms and —NH— groups located at positions 1 and 3 in the purine and pyrimidine rings, respectively, where they again Complement each other. This is precisely why the number of moles of A equals the number of moles of T in DNA, and the number of moles of G equals that of C. Naturally, the total content of purines and pyrimidines is equal.

Fig. 66. Hydrogen bonds between base pairs in nucleic acids:
A — between adenine and thymine; B — between guanine and cytosine; C — between pairs of complementary bases in a double-helical DNA fragment; arrows indicate the antiparallel orientation of the helices shown as hatched ribbons
The third regularity in the ratio of purine to pyrimidine bases is related to the Specificity of DNA and RNA. Therefore, the ratio
is referred to as the specificity coefficient of nucleic acids.
It has been established that DNA exhibits pronounced species specificity. Particularly sharp species differences are characteristic of DNA isolated from
microorganisms. In RNA, species specificity, as expressed by the ratio
, is less pronounced. Tables 17 and 18 list the specificity coefficients for a number of total DNA and RNA samples. Despite the small number of arbitrarily chosen objects, these tables clearly illustrate the regularities formulated above. In general, the specificity coefficient for DNA varies from 0.45 to 2.57 in microorganisms, from 0.58 to 0.94 in higher plants, and from 0.54 to 0.81 in animals. For total RNA, the specificity coefficient is typically greater than unity.
For a long time, the nucleotide composition of DNA and RNA was determined by separating nucleotides, nucleosides, or free nitrogenous bases—obtained via the degradation of the studied nucleic acids—using chromatographic methods followed by spectrophotometric Analysis of the identity and content of each base. However, classical methods for analyzing the nucleotide composition of nucleic acids have recently been superseded by physical methods. For instance, the content of GC pairs in DNA is now determined from the melting temperature of DNA and its buoyant density. In the first case, the GC-pair content is related to the DNA melting temperature (tm) by the following relationship: GC (mol. %) = (tm — 69.3) ∙ 2.44. This equation holds true when DNA is dissolved in standard saline citrate (0.15 M NaCl and 0.015 M sodium citrate per liter, pH 7.0). In the second case, the GC-pair content is directly proportional to the buoyant density of DNA, determined by ultracentrifugation in a CsCl density gradient.
Physical methods have accelerated the accumulation of empirical data on the nucleotide composition of DNA across various biological species and made it possible to append new regularities, discovered by A. N. Belozersky and his disciples, to Chargaff's rules.
It turned out that the Variability of DNA nucleotide composition is very high in evolutionarily ancient taxa and relatively low in younger ones; thus, the nucleotide composition can serve as an indicator of the evolutionary age of a taxon. Furthermore, two types of DNA are widespread in nature: an AT-type DNA predominates in Chordates and invertebrates, higher plants, yeast-like organisms, blue-green Algae, and a number of Bacteria and viruses, whereas a GC-type DNA is predominantly found in non-yeast Fungi, actinomycetes, algae, and certain bacteria and viruses. Finally, the distribution of methylated bases among members of the animal and plant kingdoms is also non-random. For example, 5-methylcytosine (5-mC) is present in significant amounts in the DNA of higher plants (2–10 mol. %) and some animals (up to 3 mol. %), whereas it is scarce in bacterial DNA (less than 0.5 mol. %). Conversely, N6-methyladenine (N6-mA) is found in bacteria and algae, detected in small quantities in plants, and absent in many animals. All of this has laid the groundwork for the further development of chemical Taxonomy.
The revealed regularities in the ratios of purine to pyrimidine bases in DNA and RNA molecules indicated that The sequence of nucleotide residues in nucleic acids is specific in nature. However, these data did not and could not provide insight into the true sequential arrangement of the structural elements that make up nucleic acid molecules, i.e., their Primary Structure.
Primary structure of DNA. Even in the simplest case of bacteriophage fd (see Table 16), whose DNA is represented by a single-stranded polydeoxyribonucleotide with M = 1.9 ∙ 106, it must contain approximately 5,760 nucleotide residues (1,900,000 / 330, where 330 is the average molecular weight of a nucleotide unit) arranged in a strictly defined sequence.
Determining the primary structure of such a huge polydeoxyribonucleotide is extremely difficult because it consists of only four types of nucleotide residues containing A, G, C, and T as nitrogenous bases, which are relatively rarely methylated. In the case of DNA with a higher molecular weight, the situation is even more complex. Nevertheless, thanks to the development of two highly promising DNA Sequencing Methods between 1975 and 1977, decisive progress has been achieved in this regard.
The first method, proposed by F. Sanger and R. Coulson, involves generating a complete set of single-stranded DNA copies of decreasing length using a DNA polymerase reaction (see p. 249), separating them by polyacrylamide gel electrophoresis (the number of observed bands determines the number of nucleotide residues in the DNA or its fragment), and identifying the positions of deoxyadenylic, deoxyguanylic, deoxycytidylic, and deoxythymidylic residues at the 3'-end of each fragment using specific techniques.
The second method, developed by A. Maxam and W. Gilbert, boils down to the chemical modification of purine and pyrimidine bases in DNA or its fragments, the Selective Cleavage of phosphodiester bonds at the sites of modified bases, the Separation of the cleavage products by polyacrylamide gel electrophoresis, and the direct reading of the DNA fragment structure from the resulting autoradiograms (the number of bands on the gel for a DNA fragment cleaved at a specific modified nitrogenous base equals the number of nucleotide residues containing that purine or pyrimidine, and their positions on the gel indicate the Location of this nucleotide residue within the analyzed DNA fragment).
As a result of applying these methods in their original or modified forms, DNA primary structure determination has become widespread in recent years (Table 19).
Table 19 displays only a small fraction of the primary DNA structures known to date. As of January 1, 1985, data were available on the primary structure of 4,175 polynucleotide compounds comprising 3 million nucleotide residues (n.r.). In just 2.5 years, by August 1987, the number of deciphered primary nucleic acid structures increased to 14,020, containing 14,855,147 n.r. By April 1993, GenBank contained sequence data for 111,911 polynucleotides comprising 129,968,355 n.r., including the mapping of 2,353,635 n.r. in the E. coli genome (50% of The Genome). All of this has made it possible to set what once seemed a futuristic goal: determining the complete primary structure of The Human Genome, which consists of 3 billion nucleotide pairs (n.p.). According to estimates, this project will cost 3 billion dollars and take several years. It is already underway in our country (the governmental research program "Human Genome"), as well as in the USA and Japan. Its goal is not only to map the localization of genes (which account for 5–10% of the genome) but also to understand the mechanisms regulating their activity, as well as to outline ways to correct defects that lead to Hereditary diseases. With the currently achieved DNA sequencing speed of up to 280,000 n.p. per day and a projected rate of up to 1 million n.p. per day (utilizing fluorescently labeled DNA, supercomputing technology, and custom-designed oligodeoxynucleotides), achieving this goal is becoming a realistic prospect, especially since the cost of reading a single nucleotide has already dropped from 1 dollar to a few cents. Tangible results are already evident: whereas by 1975 the primary structure of only 25 human genes had been elucidated, by 1985 this number had risen to 900, and by 1992 it reached 3,618. The Fourth Scientific Conference on the Progress of the Human Genome Program (Chernogolovka, March 9–11, 1994) provided an analysis of the contribution made by Russian scientists to its Implementation.
Table 19. Number of nucleotide pairs (n.p.) in DNA of certain genomes and genes with fully established primary structure
Genome |
Number of n.p. |
Gene |
Number of n.p. |
Gram-positive bacterium Bacillus subtilis |
4214810 |
Human embryonic globin |
11376 |
Spirochete Borrelia burgdorferi, strain B31 |
910725 |
7564 |
|
Epstein-Barr virus |
172282 |
2600 |
|
Tobacco chloroplast |
155844 |
Salmonella DNA polymerase ß-subunit1 |
1986 |
Liverwort chloroplast |
121024 |
Phage T4 DNA ligase1 |
1469 |
Adenovirus |
35937 |
Human Insulin |
1430 |
Human and bovine mitochondria |
16659 |
Shigella toxin1 |
1380 |
Cauliflower mosaic virus |
8031 |
Soybean glycinin1 |
944 |
Bacteriophage M13 |
6407 |
Na+, K+-ATPase ß-subunit1 |
912 |
Virus ⊘X-174 |
5386 |
cGMP phosphodiesterase γ-subunit1 |
779 |
Simian virus 40 (SV40) |
5243 |
Human leukocyte interferon1 |
520 |
1 Primary structures sequenced by Russian scientists.
Note: Data on the primary DNA Structure of the Escherichia coli genome (4,700,000 n.p.) and all 16 Chromosomes of baker's yeast (12,000,000 n.p.) have just been published.
It is clear that the primary structure of most DNAs is represented by double-stranded deoxypolynucleotides with a unique sequence of tens and hundreds of thousands of constituent deoxyribonucleotide residues. As an example, the structure of the human leukocyte interferon gene is shown below:

Each of the 520 deoxyribonucleotide residues making up the gene (one strand is shown) is designated by capital letters (A — deoxyadenylic, G — deoxyguanylic, T — deoxythymidylic, C — deoxycytidylic, C* — 5-methyldeoxycytidylic acid). It is on this strand that mRNA is synthesized (see p. 220), serving as a template for The biosynthesis of interferon protein in human leukocytes, which exerts a powerful influence on metabolic processes and consequently possesses immense therapeutic efficacy. When integrated into the DNA of Escherichia coli, the interferon gene ensures the biosynthesis of interferon in a bacterial culture, opening up possibilities for its practical production in the pharmaceutical industry. This has been achieved, in particular, at the Shemyakin-Ovchinnikov Institute of Bioorganic Chemistry of the Russian Academy of Sciences, where a corresponding interferon-producing strain of E. coli was engineered.
Secondary structure of DNA. In the vast majority of cases (with the exception of the single-stranded DNA of certain phages), naturally occurring DNA molecules consist of pairs (see Table 19) of interwound polydeoxyribonucleotide chains, each characterized by a specific yet complementary sequence of nucleotide residues (dashed lines indicate hydrogen bonds between complementary bases):

A photograph of a three-dimensional model of DNA molecules constructed according to this pattern is shown in Fig. 67, A. Also shown are the arrangement scheme and characteristic parameters of the polynucleotide chains (Fig. 67, A), as well as the layout of the complementary bases holding the polydeoxyribonucleotide chains together via hydrogen bonds (Fig. 67, B). This model of the DNA molecule was first proposed by J. Watson and F. Crick in 1953 based on X-Ray Diffraction Analysis of DNA structure. Data obtained over the subsequent 45 years have fully confirmed it, and its authors became Nobel laureates.
Fig. 67 illustrates models of the so-called B-form of DNA. The fact is that, depending on various conditions, DNA can exist in a variety of ordered fibrous-crystalline structures. More than ten of these have been obtained, and four of them—the A-, B-, C-, and T-forms of DNA—have been studied by X-ray diffraction analysis. While adhering to the General structural plan of the DNA molecule as a bispiral polydeoxyribonucleotide, they differ in a number of parameters specific to each of these conformational modifications of DNA (Table 20).

Fig. 67. Structure of the DNA molecule:
A — model of the DNA molecule: the dark and light chains of atoms standing out in the figure represent the outwardly directed pentose-phosphate backbone. Directed toward the interior of the model, perpendicular to its long axis, are the purine and pyrimidine bases joined by hydrogen bonds; atoms belonging to the bases (cross-hatched diagonally) densely fill the central part of the model; B — spatial arrangement of the polydeoxyribonucleotide chains in the DNA molecule; the transverse lines in the diagram denote the planes in which the base pairs connecting the chains to one another lie: 0.34 nm is the distance between adjacent deoxyribonucleotide residues; 1 nm is the radius of the molecule; 3.4 nm is the pitch of The Double Helix; C — arrangement of complementary purine and pyrimidine bases in the DNA molecule: deoxyribose molecules are represented by white pentagons outlined with a double line; phosphate groups are indicated by the bend of the double line connecting the deoxyribose residues; bases extend from each carbohydrate residue in the form of dot-shaded hexagons and pentagons; the short double lines connecting the cross-hatched purine and pyrimidine bases represent hydrogen bonds. All helices are right-handed
The most dramatic conformational change occurs during the transition of the A-form of DNA into the B-form. Specifically, this is accompanied by a sharp shift in the Spatial Structure of ß-D-2-deoxyribose within it, as a result of which the tilt angle of the base planes relative to the helix axis changes by more than 20° (Fig. 68).
The footnote to Table 20 lists the conditions for obtaining crystalline DNA fibers. However, within The Cell as well, varying degrees of Hydration in its compartments or membranes, as well as differences in the Ionic strength of the surrounding environment, create conditions for DNA to exist in various Conformations that undergo mutual transitions. In biological terms, the B-form is most optimal for Replication processes, the A-form for Transcription, and the C-form for DNA packaging within supramolecular Chromatin structures and certain viruses. Thus, the secondary structure of DNA molecules appears to be linked to the execution of informational processes in living nature, namely: the A-form of DNA is associated with information transfer from DNA to RNA, the B-form with information Amplification (replication), and the C-form with information storage.
Table 20. Characteristics of Some conformational states of DNA
Parameter |
A-form |
B-form |
C-form |
T-form |
Number of nucleotide residue pairs per turn |
11 |
10 |
9.3 |
8.0 |
Tilt angle of base planes to the helix axis, ° |
20 |
-2 |
-6 |
-6 |
Rotation angle of bases around the helix axis, ° |
32.7 |
36 |
38.6 |
45.0 |
Distance of complementary pairs from the helix axis, nm |
0.425 |
0.063 |
0.213 |
0.143 |
Distance between nucleotide residues along the helix height, nm |
0.256 |
0.338 |
0.332 |
0.304 |
Angle between planes of complementary bases, ° |
8 |
5 |
5 |
— |
Note. The A-form was obtained from aqueous-salt solutions containing alkali Metal Ions except Li+; X-ray diffraction analysis was performed at 75% relative humidity. The B-form was obtained from aqueous-salt solutions containing Li+; X-ray diffraction analysis was performed at 66% relative humidity. The C-form was obtained under the same conditions as the B-form, but in solutions with a different cation ratio; X-ray diffraction analysis was performed at less than 66% relative humidity. The T-form (D-form) was isolated from phage T2; it contains glucosylated hydroxymethylcytosine residues.
In recent years, data have emerged indicating the possible existence of two fundamentally new DNA forms: the Z-form and the SBS-form. The Z-form of DNA, discovered in 1979 in A. Rich's laboratory, is represented by left-handed polydeoxyribonucleotide chains (Fig. 69) within a bispiral molecule, with the phosphate groups arranged in a zigzag pattern (hence the Z-form). The diameter of the DNA molecule in the Z-form is 1.8 nm (compared to 2.0 nm in the B-form), the number of bases per turn is 12, the distance between turns is 0.34 nm, and the tilt of the bases relative to the helix axis is 7°; the Z-form of DNA has only a single groove (instead of the two grooves found in the A- and B-forms of DNA; see Figs. 67 and 68). It is hypothesized that natural DNA may feature alternating right-handed (A-, B-, and C-forms) and left-handed (Z-form) segments, with the latter potentially serving as "hot spots" where DNA participates in various metabolic processes.
The SBS-form of DNA is characterized by the absence of interwinding of the polydeoxyribonucleotide chains into a bispiral molecule; instead, they lie side by side while maintaining the other parameters of the DNA molecule (from which the name SBS-form is derived).
This form of DNA ensures exceptionally easy base unpairing and separation of DNA strands, which is critically important during DNA biosynthesis.

Fig. 68. Polymorphism of DNA secondary structure:
I — A-form; II — B-form; a — side view; b — top view. In the A-form of DNA, the complementarily paired nitrogenous bases are located 0.425 nm away from the helical axis, whereas in the B-form they lie very close to the axis (see Table 20). Consequently, a cavity is formed inside the molecule in the first case, while in the second case, the interior of the double-stranded deoxypolyribonucleotide is almost completely filled with nitrogenous bases.

Fig. 69. Comparison of the B- and Z-forms of DNA:
Dark spheres indicate internucleotide phosphates, and thick lines connecting them represent the path of the pentose-phosphate backbone. Z-like bends are clearly visible in the left-handed pentose-phosphate chains of Z-DNA, whereas such bends are absent in the right-handed chains of B-DNA. The differences in the number and depth of the Major and minor grooves in both DNA forms are clearly prominent.

Fig. 70. Formation of a four-stranded hairpin from a polynucleotide with an oligopurine-oligopyrimidine sequence: the quaternary helix arises due to complementary interactions between the polypurine (dark) and polypyrimidine (light) regions within the Double helices (lower part of the figure).
Information has recently emerged regarding the existence of DNA fragments in the form of triple and quadruple helices. The first of these has been named H-DNA, as it forms at pH 4.0 and is stabilized by H+ ions; here, the triple complex arises on purine-pyrimidine blocks through the attachment of polypyrimidine strands to them. The Biological Significance of H-DNA formation remains unclear. As for quadruple DNA helices (Fig. 70), they can be regarded as one of the variants of transition into the Tertiary Structure of Deoxyribonucleic Acids whose primary structure is rich in oligopyrimidine and oligopurine sequences.
Two types of forces hold the two polydeoxynucleotide strands together in the double-stranded DNA molecule. First, these are hydrogen bonds between complementary nitrogenous bases facing the interior of the DNA double helix (see Fig. 66). In forming hydrogen bonds, the bases lie in a plane perpendicular to the longitudinal axis of the B-form of DNA, meaning these bonds act in a transverse direction. Second, there are hydrophobic interaction forces between nitrogenous bases stacked along the DNA molecule; such packing of bases in an aqueous environment generates forces that prevent non-polar (hydrophobic) bases from contacting water molecules, causing the bases to draw closer together, while the stack-like packing is reinforced by interplanar interactions between them. Consequently, these are referred to as stacking interactions (i.e., interactions directed along the base stack and, ultimately, along the DNA molecule).
Recently, stacking interactions have been assigned a more significant role than previously thought in maintaining the secondary structure of DNA, whereas hydrogen bonds between complementary base pairs are largely attributed a guiding role in the mutual orientation of bases during the stacking process. Indeed, THE CONTRIBUTION OF hydrophobic interactions to the Maintenance of the secondary structure of double-stranded polynucleotides increases when U is replaced by T, i.e., due to the additionally introduced hydrophobic methyl radical.
In this regard, the stability of double-stranded structures in DNA molecules—as well as the mechanisms of their disruption under The Influence of Physical and Chemical agents in various conditions—appears in a different light. Specifically, water molecules and metal ions bind primarily to the pentose-phosphate backbone of DNA, with cations located predominantly in the minor groove of the B-DNA double helix. The stabilizing EFFECT OF WATER molecules is directed specifically at enhancing stacking interactions, which leads to the stabilization of hydrogen bonds between bases. At the same time, when stacking interactions weaken, water molecules interact competitively with the proton-donor and proton-acceptor centers of the bases, contributing to destabilization and initiating the further unwinding of the double helix. AT base pairs are hydrated twice as strongly as GC pairs, and native conformation is preserved at a water content of no less than 0.6 g per 1 g of DNA. All this highlights the dynamic nature of the secondary structure of DNA and the possibility of conformational and other transitions within it under the Influence of Environmental agents.
Double-stranded structures in DNA molecules arise not only through the interaction of two complementary polydeoxynucleotide chains, but also within a single chain. This occurs when complementary DNA strands contain palindromes—inverted repeat sequences of nucleotide units (from the Greek *palin* — back, *drome* — running). Palindromic structures are characteristic of regions in the DNA molecule where recognition sites for enzymes and regulatory proteins are located. For example, due to palindromes, the DNA strands of *Escherichia coli* self-spiral to form hairpins in the region of the lactose Operon—a structure containing information for the biosynthesis of ß-galactosidase, which ensures lactose degradation when it is present in the growth medium (Fig. 71). It has been suggested that palindromes are a source of quadruple DNA helices, which may serve to form elements of the tertiary structure of DNA.

Fig. 71. Palindromes of *Escherichia coli* DNA (lactose operon region) in linear (I) and hairpin (II) states:
in the linear structure, the bold dot indicates the starting point of hairpin formation; dashed rectangles mark the palindrome zones; arrows of various thicknesses and shapes denote identical palindromes and their direction.
Tertiary structure of DNA. DNA molecules exist in linear and circular forms (Fig. 72). Most Native DNA molecules presumably exist in linear form, but the DNA of certain viruses and phages, as well as the DNA of chloroplasts, mitochondria, centrioles, and bacterial Plasmids, possesses a circular structure. The tertiary structure of both linear and circular DNA forms is characterized by coiling and supercoiling. Dynamic transitions are hypothesized to occur between circular and Linear Forms of DNA, as well as between its relaxed and supercoiled states. In the DNA of certain viruses (such as polyomavirus) and in Mitochondrial DNA, such transformations have been studied in detail (Fig. 72).
The situation is more complex in the case of eukaryotic chromosomal DNA. The Isolation Methods used are accompanied to a greater or lesser extent by its degradation and, most importantly, by its separation from deoxyribonucleoprotein—the form in which it exists within the nuclear apparatus of cells—specifically its protein moiety, which is absolutely essential for maintaining the tertiary structure of DNA. Therefore, The problem of nuclear DNA tertiary structure can only be solved by studying the structure of the nuclear chromatin and chromosomes.
DNA in chromatin and chromosomes also exists in a supercoiled state, with several levels of supercoiling realized here. The first level of the supercoiled state of DNA in chromatin is maintained by Histones, which occupy DNA segments about 200 bp long and form the elementary structural unit of chromatin—the nucleosome (Fig. 73). The chain of nucleosomes, in turn, forms a helix of the second and higher orders, culminating in Condensation into a chromosome (Fig. 74, A). Here (Fig. 74, B), each turn of the helix in a nucleosome accounts for 80 bp (a 6- to 7-fold compaction of DNA); in the solenoid (6 nucleosomes per turn), 1,200 bp (a 40-fold compaction); in each chromatin loop, 60,000 bp (a 680-fold compaction); and in the chromosome, 1.1 ∙ 106 bp (a 1.2 ∙ 104-fold compaction).

Fig. 72. Tertiary structure of the DNA molecule:
a — linear; b — circular; c — supercoiled circular; d — compact coil structure.

Fig. 73. Structure of the nucleosome
The basis of the nucleosome is a protein core around which the double-stranded DNA is wound (1.8 left-handed supercoils, 146 ± 2 bp, pitch — 2.75 nm); the histone tetramer (H3)2 ∙ (H4)2 determines the positioning of the central turn of the DNA superhelix in the groove of the protein core, while two heterodimers (H2a ∙ H2b) participate in packing the remaining part of the DNA superhelix by ~0.5 turns each; these histones are called core histones; histone H1 stabilizes the entry and exit sites of DNA into the nucleosome by interacting with the linker DNA (from 10 to 80 bp) connecting adjacent nucleosomes.
Crucially, the transition of DNA into the supercoiled state and back is mediated by a specific group of enzymes known as topoisomerases—i.e., enzymes that alter spatial structure.
Properties of DNA. DNA substances are white, fibrous in structure, poorly soluble in water in their free state, but readily soluble as alkali metal salts. They are also highly soluble in concentrated salt solutions.
Because DNA molecules are sharply asymmetric, their solutions exhibit high viscosity and birefringence. Possessing a large negative charge, DNA molecules are mobile in an electric field. All DNAs are optically active.
When DNA solutions are heated within the temperature range of 80 to 90° C, nucleic acid "melting" occurs, accompanied by A change in solution viscosity and an increase in ultraviolet Light absorption (at 260 nm). The latter phenomenon is known as the hyperchromic effect.
Chemically, DNAs are quite inert, which is why they were long considered indifferent structural elements of the cell content. However, data gradually accumulated regarding a number of Chemical Reactions characteristic of nucleic acids: they tightly bind polyvalent metal ions, with Cu2+ and Me4+ forming insoluble complexes with DNA. Primarily, polyvalent cations react with the N and O atoms of guanine. Metal ions may also be involved in maintaining the tertiary structure of nucleic acids.
Nucleic acids readily interact with Polyamines, such as spermidine [H2N—(CH2)3—NH—(СН2)4—NН2] and spermine [H2N—(СН2)3—NH—(СН2)4—NH—(СН2)3—NH2], which participate in maintaining the tertiary structure of nucleic acids. An important reaction of DNA is the alkylation of the amino groups of A, C, and G. Of equal importance is the deamination of these nitrogenous bases: G and C are deaminated first. Both alkylation and deamination of DNA form the basis of research in chemical mutagenesis, i.e., altering heredity using chemical agents.

Fig. 74. Levels of DNA Supercoiling in chromatin:
A — dynamics of DNA-protein complex supercoiling; B — degree of DNA compaction during supercoiling (see text on p. 211).
STRUCTURE AND FUNCTIONS of transfer RNAs. Transfer RNAs were first isolated from the so-called "soluble" fraction of the cell, i.e., the supernatant of the cell homogenate. The primary function of this type of ribonucleic acid proved to be The ability to accept amino acids and transfer them to the protein-synthesizing apparatus of the cell—the ribosome. Accordingly, they are designated as transfer Ribonucleic Acids (tRNAs).
The soluble RNA fraction, which accounts for 1% of the cell's dry matter or 10% of total cellular RNA, has a highly complex composition. It includes several dozen individual tRNAs, each of which has been obtained in a homogeneous state.
Since each individual tRNA is capable of transferring a single amino acid during protein synthesis, specific tRNAs are named after the proteinogenic amino acid they accept (e.g., Glycine tRNA, alanine tRNA, Lysine tRNA, etc., or abbreviated as tRNAGly, tRNAAla, tRNALys, etc.). If the same amino acid is accepted by several individual tRNAs, the latter are termed isoaccepting and are numbered (e.g., tRNAVal1, tRNAVal2, etc.). Thus, 4 isoaccepting tRNALeu, 3 tRNAPro, and 2 each of tRNASer, tRNAMet, tRNAAla, tRNALys, tRNATyr, tRNAGly, and tRNAThr have been discovered.
The chemical composition of tRNAs proved distinctive in only one respect: compared to Other types of RNA, they are rich in minor nucleotide residues. It is precisely within tRNAs that the minor nitrogenous bases listed in Table 15 are found. Furthermore, tRNAs contain Nucleosides and Nucleotides of unusual structure, such as pseudouridine and neoguanylic acid:

The list of minor tRNA components is continuously expanding and currently numbers around 50. Minor nucleotide residues make up approximately 10% of all nucleotide units in tRNA, with a significant portion of them (4–5%) represented by pseudouridylic acid. It is hypothesized that minor nucleotide residues protect tRNAs from ribonuclease attack, which is essential since tRNAs function in the soluble fraction of the cell. Additionally, some researchers believe that certain minor components participate in amino acid coding and are important for the enzyme (aminoacyl-tRNA synthetase) to "recognize" the specific tRNA that interacts with a given amino acid during its activation. Regarding the ratio of conventional nitrogenous bases in tRNA, it is characterized by a pronounced predominance of the sum of G and C over that of A and U. Thus, tRNA belongs to the GC type of RNA.
As noted above, the molecular weights of tRNAs range from 17,000 to 35,000. In most cases, they are concentrated within a narrower range of 22,000 to 27,000, meaning a tRNA contains between 70 and 80 nucleotide residues (n.r.).
Because R. Holley and coworkers successfully utilized a ribonuclease specific for guanylic nucleotide residues in their research on the primary structure of yeast tRNAAla (at 0° C, pH 7, in the presence of Mg2+), they were able to obtain several large blocks containing from 10 to 39 n.r. and subsequently decipher the primary structure of each. As a result, the complete primary structure of alanine tRNA from baker's yeast was established for the first time in 1965 (Fig. 75). Over the next 5 years, the primary structure of 16 tRNAs was deciphered; over the following five years, 47; and in another five years, 115 tRNAs and a number of tRNA precursors. Furthermore, the primary structures of tRNAVal1 and tRNAVal2 from yeast were elucidated in our country by Academician A. A. Baev and coworkers in 1967 and 1971, respectively. These data primarily cover tRNAs isolated from bacteria and yeast, but in some cases also pertain to tRNAs obtained from higher organisms (mammals, plants, insects, etc.). As of April 1993, The nucleotide sequence had been elucidated for 2,011 tRNAs isolated from Representatives of the animal and plant kingdoms, as well as from microorganisms.
Mass data obtained from the study of tRNA primary structure have revealed several general patterns. In 75% of cases, tRNA molecules begin with a guanosine-3',5'-diphosphate residue and in all cases terminate with a CCA triplet, with the terminal adenosine residue serving to bind (accept) The amino acid:

The primary structure of the studied tRNAs features homologous blocks that are extremely similar in their nucleotide residue arrangement. This is especially evident in the dihydrouridine and pseudouridine loops (Fig. 75). For instance, in the pseudouridine loop of all investigated tRNAs without exception, the tetranucleotide fragment —GTVΨC— is present; the dihydrouridine loop is a concentration site not only for dihydrouridylic acid residues but also for AG sequences and minor bases.
Minor nucleotide residues are also distributed systematically within the primary structure of tRNAs. This is explained by the fact that only certain positions in tRNA molecules are accessible to the action of methylating enzymes, which modify conventional bases into methylated derivatives at the level of the fully formed molecule.
The secondary structure of tRNA is clear from an examination of Fig. 75. Its characteristic feature is the folding and self-coiling of the polynucleotide chain within strictly fixed complementary regions.
There are 4 or 5 such regions, depending on the presence or absence of an extra loop, which—when present—is always located between the anticodon arm and the pseudouridine loop. As a result of limited double-helical folding, a uniform secondary structure resembling a cloverleaf is formed in all tRNAs (see Fig. 75).

Fig. 75. Structure of tRNA:
A — Primary and secondary structures of alanine tRNA; B — generalized cloverleaf structure of tRNA; Ψ — pseudouridylic acid residue; m1G — N1-methylguanylic acid residue; m22G — N2-dimethylguanylic acid residue; m1I — N1-methylinosinic acid residue; I — inosinic acid residue; U(2Н) — dihydrouridylic acid residue; I — dihydrouridine loop; II — anticodon loop; III — pseudouridine loop; IV — acceptor end; V — extra loop; • — hydrogen-bonded bases; O — unpaired bases; Pur — purine base; Pyr — pyrimidine base
During the heyday of the cloverleaf tRNA model, a strictly defined functional significance was attributed to each distinct part of the molecule. A particularly major role in elucidating the functions of various polynucleotide sequences within the tRNA molecule was played by the severed-molecule method, proposed in the laboratory of Acad. A. A. Baev and vividly termed "molecular surgery" by Acad. V. A. Engelhardt. Its essence boils down to enzymatically or chemically cleaving a limited number of internucleotide bonds in a tRNA molecule, separating the resulting fragments, and reassociating tRNA molecules deficient in one of the fragments. The reassociation of tRNA molecules from a given set of fragments occurs via weak interaction forces and proceeds by THE PRINCIPLE OF self-assembly. Testing the biological activity of such molecules showed that the concepts regarding the unambiguous role of the dihydrouridine, pseudouridine, and other loops and PARTS OF THE tRNA molecule are not entirely correct. For example, not only the dihydrouridine loop but also other Regions of the tRNA molecule are responsible for binding aminoacyl-tRNA synthetase.
Elucidating the functional activity of both tRNA fragments and the entire tRNA molecule is undoubtedly tied to breakthroughs in understanding its tertiary structure. Based on X-ray diffraction data obtained at resolutions ranging from 0.6 to 0.25 nm, structural models have been built for a number of tRNAs. These include yeast, E. coli, and thermophilic bacteria tRNAPhe, tRNAAsp, tRNATrp, E. coli formylmethionine tRNA, and many others. The tertiary structure of yeast tRNAPhe has been mapped in the greatest detail by A. Rich (Fig. 76).

Fig. 76. Tertiary structure of yeast phenylalanine tRNA:
A — Schematic representation. Arabic numerals indicate nucleotide residue numbers. The molecule features an L-shaped conformation, with the anticodon loop and acceptor stem situated at opposite ends of the molecule, mirroring the planar cloverleaf model. Due to mutual interaction and hydrogen bonding between complementary nitrogenous bases, the dihydrouridine and pseudouridine loops are brought into close proximity near the longitudinal axis of the molecule. Double-helical motifs remain clearly defined within the anticodon and acceptor regions. Functional interactions of tRNA with enzymes, the ribosome, and substrates involve nucleotide residues scattered across various parts of the polynucleotide chain but brought together during the folding of the spatial structure. To some extent, this resembles The Emergence of active sites in protein molecules. B — 3D spatial model

Fig. 77. Proposed secondary structure of E. coli 16S rRNA, deduced from maximal base-pairing of complementary residues:
I–IV — domains essential for understanding the principles governing the formation of its tertiary structure (their boundaries are marked by solid and dashed lines)
tRNAs also perform non-canonical functions that have been receiving increasing research attention: acting as signals in viral genomes; participating in non-template (ribosome-independent) transfer of amino acid residues to growing peptidoglycan chains in bacterial cell walls or to existing Polypeptides via aminoacyl-tRNA-protein transferases; serving as primers for Reverse Transcriptase activity (see below), and more.
Structure and functions of ribosomal RNAs. Ribosomal RNAs (rRNAs) constitute the bulk of cellular RNA (accounting for 80–85% of total cellular RNA). All organisms contain Three types of rRNAs, which differ in molecular weight and localization within ribosomes (see Chapter VII), of which they are an obligatory structural component. Two of these rRNAs are high-molecular-weight species, while the third is relatively low-molecular-weight. Additionally, eukaryotic ribosomes contain another low-molecular-weight rRNA.
Depending on the ribosomal class (70S or 80S), the sedimentation coefficients and molecular weights of the two large, high-polymer rRNAs (large and small) vary slightly. The smaller high-polymer rRNA (sedimentation coefficient 16.0–18.0S, M = 0.55 ∙ 106 – 0.79 ∙ 106) is localized in the 30–40S ribosomal subunits, whereas the larger one (sedimentation coefficient 23–29S, M = 1.07 ∙ 106 – 1.6 ∙ 106) resides in the 50–60S subunits.
In absolutely all ribosomes, a low-polymer 5S rRNA with a molecular weight of 40 kDa is present. It is localized within the 50–60S ribosomal subunits. A low-polymer rRNA with a sedimentation coefficient of 5.8S (M ~ 50,000) is characteristic exclusively of eukaryotic ribosomes.
The nucleotide composition of high-polymer rRNAs varies over a fairly wide range, shifting increasingly toward a GC-rich composition as organisms become more complex. rRNAs isolated from mitochondrial ribosomes show a marked dominance of AT pairs, which serves as one of the arguments for classifying mitochondrial RNAs into a distinct group. High-molecular-weight rRNAs contain 2 to 5 times fewer minor bases than tRNAs, and their repertoire is much more limited.
The nucleotide composition of 5S rRNA is peculiar: it completely lacks methylated bases, and pseudouridylic acid is found only in certain cases (e.g., in Yeasts) at about 1.3 mol. %. The ratio of major nucleotides in 5S rRNA is characterized by a distinct GC type.
Recent years have seen major breakthroughs in decoding the primary and secondary structures of rRNAs, most notably for 5S rRNA. The primary structure of 5S rRNA has now been elucidated in over 500 instances. With rare exceptions, all 5S rRNAs studied to date contain precisely 120 nt, with sequence variations among different sources being minor yet specific enough to be utilized in chemical taxonomy. The primary structure of E. coli 5S rRNA is presented below:

The primary structure of several dozen 5.8S rRNAs and their precursors has also been deciphered. The vast majority of them consist of 158 nt.

Fig. 78. Models of the secondary (A) and tertiary (B) structure of 5S rRNA
Significant progress has been made in establishing the primary structure of 16–18S rRNAs. The primary structure of E. coli 16S rRNA was first reported in 1974 by J. Ebel (France). Following this, the primary structures of over one hundred 16–18S rRNAs were determined. The number of nucleotide residues is 1,542 in E. coli 16S rRNA (Fig. 77) and 1,874 in rat liver 18S rRNA. Similar studies have been completed for E. coli 23S rRNA, baker's yeast 25S rRNA, and several others. E. coli 23S rRNA comprises 2,904 nt, while rat liver 28S rRNA contains 4,718 nt. As of April 1993, the ribonucleotide sequence had been mapped for 1,850 types of 16–18S rRNAs (including 100 from archaebacteria, 1,400 from eubacteria and chloroplasts, and 350 from eukaryotes), as well as for 23–28S rRNAs isolated from 150 prokaryotic and eukaryotic representatives.
The secondary structure of rRNA is characterized by the self-folding and base-pairing of the polyribonucleotide chain in both low- and high-polymer rRNAs.
This coiling is driven by interactions between complementary bases (G–C and A–U), resulting in the formation of a variable number of double-helical regions within the rRNA molecules. These regions are quite prominent in 5S and 5.8S rRNAs and somewhat less pronounced in 16–18S rRNAs. The secondary structure of 5S rRNA is illustrated in Fig. 78, A.

Fig. 79. Evidence for the existence of rapidly labeled RNA (mRNA):
solid black line — absorbance at 260 nm in fractions after sucrose density gradient ultracentrifugation of 23S, 16S, and 4S RNA; dashed line — incorporation of 14C-uracil into RNA. The trajectory of the dashed curve demonstrates that precursor incorporation peaks in an RNA fraction with a sedimentation coefficient of approximately 8S, which corresponds to mRNA
The planar double-helical and linear segments of 5S and 16S rRNA molecules, in turn, fold into more compact, higher-order structures, thus forming the spatial tertiary structure of these macromolecules. This structure has been characterized in detail for 5S rRNA (Fig. 78, B), whereas for 16–18S rRNAs it is still far from fully understood. Nevertheless, there is no doubt that extensive compaction of constant and variable domains occurs in the latter as well.
The functional role of all types of rRNA is gradually becoming clear: 16–18S and 23–29S rRNAs serve as the structural framework for the assembly of the ribonucleoprotein strand, which folds spatially to give rise to the 30–40S and 50–60S ribosomal subunits. However, the functions of high-molecular-weight ribosomal RNAs are by no means limited to this: they can interact with mRNA and aminoacyl-tRNA; their regions are apparently recognized by protein factors (see Chapter VII) involved in polypeptide chain assembly within the ribosome; and it is also possible that 16–18S and 23–29S rRNAs, localized in different subunits, interact directly with each other or contact via specific proteins during the formation of 70–80S ribosomes upon subunit association or within the translating ribosome. As for 5S and 5.8S rRNAs, their functions are not yet fully understood. It has been definitively established only that 5S rRNA in bacterial ribosomes is associated with three proteins—L5, L18, and L25 (for ribosomal proteins, see Chapter VII below), in archaebacteria with one or two, and in eukaryotes with only one (L8); these complexes are regarded as a third ribosomal subunit, within which 5S rRNA acts as a mediator between the peptidyl transferase center and the EF-G-binding domains (see also Chapter VII for details).
Structure and functions of informational RNAs. The existence of informational ribonucleic acids (iRNAs), or messenger RNAs (mRNAs, serving to transmit genetic information from DNA to the cell's protein-synthesizing machinery), was predicted by A. N. Belozersky and A. S. Spirin in 1958, based on the correlation between the nucleotide compositions of DNA and RNA.
Two years later, the formation of mRNA in bacteria was proven by F. Gros and coworkers through direct experiments involving pulse-labeling of various RNA species (Fig. 79). This provided a fresh perspective on earlier experiments by E. Volkin and L. Astrachan (1956), who observed the appearance of a rapidly labeled fraction of DNA-like RNA in bacteria after phage infection.
The direct isolation of mRNA was achieved in 1962, when E. Bautz and B. Hall devised an ingenious method for the sorption of mRNA—synthesized upon infection of Escherichia coli with T4 phage—onto phage DNA that was covalently linked to phosphocellulose. Subsequently, affinity chromatography became the leading method for mRNA Isolation and Purification, especially after a substantial polyadenylate tract was discovered within the majority of eukaryotic mRNAs, and Sepharose covalently linked to polyuridylic acid via the Cyanogen bromide method was used as a Column matrix. Due to the complementarity between the polyadenylic sequences of the mRNA and the polyuridylic fragments of the matrix, their interaction results in the specific binding of mRNA on such columns.
Due to the application of various methods and techniques, many mRNAs have been obtained in a highly purified state. These include mRNAs that provide the template biosynthesis for Hemoglobin subunits, ovalbumin, light and heavy chains of IMMUNOGLOBULINS, histones, crystallins (eye lens proteins), Myosin, Silk Fibroin, Avidin, protamines, α-casein, and others. The total mRNA content in cells accounts for 2–3% of the total cellular RNA.
The molecular weights of mRNAs vary within a wide range: from several hundred thousand to several million daltons. For instance, the molecular weight of monocistronic globin mRNA, which serves as a template for the Biosynthesis of Hemoglobin subunits, is 150,000 (its sedimentation coefficient is 9S). The molecular weights of many monocistronic mRNAs, i.e., those encoding the biosynthesis of a single protein, are close to this value. However, in polycistronic mRNAs, which encode the biosynthesis of not one but several functionally related proteins, molecular weights are much higher and reach several million daltons. An example is the polycistronic Histidine operon mRNA isolated from Salmonella. Its molecular weight is 4 ∙ 106. Nevertheless, some monocistronic mRNAs can have molecular weights of the same magnitude, which is characteristic, for example, of silk fibroin mRNA (M = 5 ∙ 106).
The nucleotide composition of mRNA is extremely diverse. Bacterial mRNAs are characterized by a DNA-like nucleotide composition, whereas eukaryotic mRNAs do not exhibit this feature. In some cases, the nucleotide content of mRNA is highly specific. This occurs, in particular, when they encode the Biosynthesis of Proteins with highly asymmetric amino acid compositions. A striking example is fibroin mRNA. Its composition contains 40 mol % guanine and 19 mol % cytosine, which closely corresponds to the presence of 42% glycine in silk fibroin.
The primary structure of a number of mRNAs has been elucidated. For example, human globin mRNA has been found to consist of 576 nt (excluding the polyadenylate fragment), chicken ovalbumin mRNA consists of 1,859 nt, and the outer membrane lipoprotein mRNA of Escherichia coli consists of 332 nt. The primary structures are also known for rabbit and African clawed frog globin mRNAs, a number of vitellogenin mRNAs (encoding the biosynthesis of egg yolk storage protein), several dozen mRNAs encoding insect egg chorion proteins, and so forth.
The structure of mRNA is specific: their composition includes informative regions, i.e., those functioning as templates in protein biosynthesis, as well as non-informative regions. Among the non-informative regions, the previously mentioned polyadenylate fragments located at the 3'-end of the mRNA molecule have been studied. The length of polyadenylate fragments in different mRNAs ranges from 50 to 400 nt. These fragments are absent in histone mRNA molecules. Near the polyadenylate fragment, mRNA molecules contain small repeating sequences about 30 nt long, which also lack a coding function, but may serve as acceptor sites during the interaction of mRNA with the ribosome or individual protein factors. In addition, at the 5'-end of mRNA, There is a nucleotide sequence in which the nitrogenous bases are methylated, and one of the nucleoside residues, 7-methylguanosine, is attached via a triphosphate bridge. This part of the mRNA molecule is also non-informative and is called the cap:

The noted Structural Features of RNA are represented by the following generalized structure:

It is believed that the polyadenylate segment of the mRNA molecule is involved in mRNA maturation (see below), determines the lifespan of mRNA, facilitates The transport of mRNA from the nucleus to the cytoplasm, and participates in mRNA Translation (see Chapter VII). The cap is required to protect mRNA from exonucleases, to bind protein factors during interaction with rRNA in the ribosome, and it plays a signaling role in the attachment of mRNA to the ribosome and participates in translation.
The question of the secondary structure of mRNA is still under investigation. Like other types of RNA, the polynucleotide chain in the mRNA molecule folds back upon itself and forms helical regions. Computer-analyzed data on potential complementary base pairings (A—U and G—C) within the mRNA molecule served as the basis for developing a model of rabbit globin mRNA (Fig. 80). As for the tertiary structure of mRNA, nothing is known about it as yet, other than the fact that it is packaged less compactly than rRNA.

Fig. 80. Computer model of the secondary structure of rabbit ß-globin mRNA.
The coding region (see Fig.) comprises 441 nt; the non-coding region on the cap side consists of 56 nt (from the cap to the initiation codon where assembly of the ß-globin polypeptide chain begins), and on the polyadenylate side—excluding the poly(A) tail itself—it consists of 94 nt (from the termination codon, which signals the completion of ß-globin chain assembly, to THE START OF the polyadenylate segment of mRNA). Approximately 40% of the complementary nucleotides in rabbit ß-globin mRNA, as in other mRNAs, are paired, as indicated by dots; in these regions, the mRNA molecule forms self-complementary Helical structures.
At the dawn of mRNA research, when studies focused primarily on bacterial mRNAs, this type of RNA was characterized by rapid turnover—meaning a short half-life measured in seconds or minutes—a nucleotide composition similar to that of DNA, and molecular weights corresponding to sedimentation constants intermediate between those of transfer (4S) and small ribosomal (16S) RNAs. Studies on eukaryotic mRNAs fundamentally changed these criteria for classifying an RNA as Messenger RNA. Eukaryotic mRNAs proved to be relatively long-lived molecules (ranging from several hours to several days or even weeks), their nucleotide composition differed significantly from the base ratio of bulk DNA, and their sedimentation coefficients varied from 9S to 65S. Consequently, these earlier criteria have been discarded due to their inadequacy for characterizing all mRNAs. Today, only a single criterion remains universally valid: an RNA is considered messenger RNA if the sequence of nucleotide residues within its molecule carries the information required to direct the synthesis of a specific protein. Thus, template activity has emerged as the sole universal defining characteristic of this class of RNA.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.