Genetics - A. V. Sivolob 2008
The Nature of Genetic Material
DNA as a Genetic Text: Genome Organization
As a source of information, a Gene is a region of DNA whose nucleotide sequence encodes information for a specific functional product. Based on the type of this product, all genes (the complete set of genes in a given Organism is called the genotype) can be divided into two groups: genes whose ultimate product is specific functional RNA molecules (RNA genes), and genes whose sequence, in accordance with METABOLISM/28.html">The Genetic Code (see Chapter 2), encodes the Amino Acid Sequence of Proteins (protein-coding genes). RNA genes encode various non-coding RNA molecules (see Chapter 2): rRNA — Ribosomal RNAs (ribosomal components); tRNA — Transfer RNAs (a key element of the Translation system); Small nuclear RNAs; small nucleolar RNAs; microRNAs; RNA molecules that serve as components of certain Enzymes; and Other types of RNA whose Functions are not yet fully understood. Protein-coding genes serve as templates for synthesizing RNA—Messenger RNA, or mRNA—which is subsequently used for Protein Synthesis.
The coding DNA sequence, from which information about The nucleotide sequence of the RNA transcript is transcribed, constitutes the essential core of a gene. However, for Genetic information to be expressed (via RNA and subsequent protein synthesis), regulatory DNA sequences are equally critical. Through their affinity for specific proteins, these sequences serve to turn Transcription on or off as the primary stage of Gene Expression. Thus, a gene can also be defined as a DNA region necessary and sufficient for the complete synthesis of a functional RNA molecule. A DNA segment considered to be a gene must contain a coding sequence with information about the product, as well as a specific set of regulatory sequence elements that govern the initiation or blockage of transcription, the reading path of information, and so forth.
Every Cell of a multicellular organism contains several (sometimes up to several dozen) DNA molecules, and this set is identical across all Cells. This DNA contains more than just genes; at the very least, it includes intervening intergenic regions. The totality of DNA sequences within the cells of a given organism is called The Genome. To date, the complete sequences of over 700 bacterial and about 100 Eukaryotic Genomes have been mapped. The primary difference between them is that in Prokaryotic Genomes, coding sequences account for about 95%, whereas in eukaryotic genomes, the proportion of coding sequences does not exceed 3%. The sizes of certain genomes and estimates of the number of protein-coding genes they contain are listed in Table 1.1.
Class="center">Table 1.1. Genome sizes and number of protein-coding genes in selected organisms
|
Organism |
Number of DNA molecules* |
Number of genes |
|
|
Bacteriophage phiX174 |
5386 |
1 |
10 |
|
Bacterium Escherichia coli |
4,6-106 |
1 |
4100 |
|
Ascomycete Saccharomyces cerevisiae |
1,2-107 |
16 |
6700 |
|
Nematode Caenorhabditis elegans |
108 |
6 |
20000 |
|
Fruit fly Drosophila melanogaster |
1,3-108 |
4 |
14000 |
|
Chicken Gallus gallus |
109 |
33 |
13000 |
|
Mouse Mus musculus |
3,3-109 |
20 |
22000 |
|
Human Homo sapiens |
3,2-109 |
23 |
21000 |
* For eukaryotes, the genome size and number of molecules represent half of the nuclear DNA.
Viral genomes are structured with extreme "economy": the coding regions of genes occupy virtually the entire, relatively small, viral DNA. In the prokaryotic genome, The amount of DNA and the number of genes increase significantly, yet THE PRINCIPLE OF economy in utilizing most sequences for encoding genetic information is preserved. For instance, the Escherichia coli genome is represented by a single circular DNA molecule (known as the bacterial chromosome) with a length of 4.6 million base pairs. About 90% of this DNA corresponds to the coding sequences of ~4.1 thousand protein-coding genes and ~120 non-translated RNA genes.
Eukaryotic genomes contain a substantially greater amount of DNA compared to prokaryotic genomes (see Table 1.1), with the overwhelming majority of this DNA being non-coding sequences. Notably, approximately half of the eukaryotic genome consists of sequences present in many copies (repetitive sequences). Eukaryotic DNA resides within the Cell Nucleus as part of Chromosomes, with each chromosome containing a single giant linear DNA molecule. Repetitive sequences are concentrated, in particular, at the ends of chromosomes (telomeres) and in the regions where chromosomes attach to the spindle fibers during Mitosis and Meiosis (centromeres).
A characteristic feature of eukaryotic genes (unlike prokaryotic ones) is the mosaic Structure of their coding region (Fig. 1.9): the actual coding region consists of a sequence of individual informative segments—exons—interrupted by non-informative introns. Exons frequently correspond to individual Structural domains of multidomain proteins: the evolutionary assembly of a protein from modular building blocks can occur through exon shuffling at the DNA level. Introns are non-informative in the sense that they carry no information about the final product, yet they often harbor vital regulatory regions. Moreover, the introns of certain genes may contain other genes complete with their own introns and exons. During transcription, the RNA molecule is synthesized as a continuous chain (the primary transcript contains both exons and introns). Therefore, an essential step in gene expression is splicing (Chapter 2)—the excision of introns and the joining of exons into a mature transcript, which can then serve as a template for protein synthesis. Furthermore, splicing can follow alternative pathways (Fig. 1.9)—known as Alternative Splicing—leading to the generation of diverse final products, namely different proteins.

Fig. 1.9. Mosaic STRUCTURE OF THE gene's coding region and the scheme of generating various mRNAs (broken lines indicate excised introns) via alternative splicing
The total number of genes in the genomes of higher eukaryotes ranges approximately from 20 to 30 thousand (Table 1.1). As shown in Fig. 1.10, the coding sequences of these genes account for only ~1.5% of the genome. The remainder consists of intergenic DNA (which also houses regulatory regions), introns (~30%), and repetitive sequences, which make up more than half of the genome.

Fig. 1.10. Approximate relative content of various sequence types in the eukaryotic genome
The MAIN TYPES OF repeats present in the genomes of higher eukaryotes are:
✵ genes present in multiple (sometimes up to 1,000) copies. Repeated genes are often clustered, meaning they are located adjacent to one another;
✵ pseudogenes — sequences that are homologous to specific genes but are not expressed. Their origin is driven, for example, by Mutations in a portion of repeated genes: undamaged genes take over the function of the damaged ones, while the latter remain in the genome;
✵ multiple repeats of short sequences (tandem repeats), some of which are distributed throughout the genome, while the majority are concentrated in the telomeric and centromeric regions of chromosomes;
✵ interspersed (dispersed) mobile elements capable of transposition and Replication within the genome. Mobile elements occupy a significant portion of the eukaryotic genome (from 30% to 50%), but are distributed unevenly: there are long stretches where 90% of the DNA consists of mobile elements, as well as regions where interspersed elements are entirely absent. Overall, a negative correlation is observed between gene density and mobile element density. The types of eukaryotic mobile elements will be discussed in detail in Chapter 6.
In addition to The Cell nucleus, DNA is also found in Mitochondria and Chloroplasts, where it constitutes an autonomous cytoplasmic element of the eukaryotic genome, small in comparison to the nuclear genome (see Chapter 6).
Last update: 11/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.