Genetics - A. V. Sivolob 2008
Genetics of Multicellular Eukaryotes
Eukaryotic Genomes
General Structural Features of Eukaryotic Genomes
All genes of a multicellular Organism can be divided into two groups: 1) genes responsible for certain universal Functions that are active in all Cells—known as housekeeping genes; 2) genes that are specifically activated in particular Cell types—known as luxury genes. A rather common feature of the first group of genes is the presence of so-called CpG islands in their regulatory regions—segments with an elevated content of CpG dinucleotides (referring to The nucleotide sequence along The Double Helix). Overall, the content of these dinucleotides in Eukaryotic Genomes is approximately five times lower than expected due to the methylation of cytosine within the CpG context: 5mC (5-methylcytosine) spontaneously converts into thymine, which is a major source of Mutations. Cytosine methylation in regulatory regions serves as one of the mechanisms for Gene repression (see below). Consequently, genes that maintain their activity in most cells contain unmethylated CpG dinucleotides, the levels of which remain high.
The total number of protein-coding genes in the genomes of higher eukaryotes varies roughly from 20,000 to 30,000 (see Table 1.1). The approximate functional distribution of eukaryotic Proteins is illustrated in Fig. 6.1.
Class="center">
Fig. 6.1. Approximate functional distribution of proteins in the eukaryotic proteome
Among eukaryotic genes, 25–50% are unique (represented in The Genome by a single copy), while the rest belong to gene families consisting of multiple, generally non-identical copies. The corresponding (homologous, yet non-identical) proteins form a protein family. Several families (such as protein Kinases, certain METABOLISM/31.html">Transcription factors, and IMMUNOGLOBULINS) comprise hundreds of proteins, whereas most families consist of a few (up to 30) proteins. Genes of such families are often grouped in the genome into clusters—located adjacently on a specific chromosome (e.g., heat Shock gene clusters, globin genes). It should be noted that such a cluster is not an Operon; each gene is regulated as an independent transcription unit. For instance, the β-globin gene cluster contains homologous genes that are activated at specific stages of individual development (Fig. 6.2). However, all genes in the cluster also share a common regulatory region responsible for the overall potentially active state of the cluster in cells destined to synthesize Hemoglobin.

Fig. 6.2. The β-globin cluster on human chromosome 16 (each gene contains introns). The developmental stages at which the respective genes are active are indicated
The β-globin cluster also contains an inactive pseudogene (see Chapter 1). Following gene inactivation (due to mutations affecting Transcription initiation, splicing, etc.), the pseudogene ceases to be subject to Selection, and A large number of mutations accumulate within it. Naturally, pseudogenes emerge most readily within clusters—when multiple copies of a gene exist, the damage of one does not lead to fatal consequences.
Several gene clusters are repeated multiple times within the genome. Among protein-coding genes, this applies to histone genes (see Chapter 1)—the genes for the five histone molecules are invariably grouped into a cluster (each gene forming a separate transcription unit) that is repeated up to 100 times. Another example of repeated clusters is the ribosomal RNA genes (see Chapter 2), though in this case, the entire cluster functions as a single transcription unit.
In addition to repeated genes, the eukaryotic genome contains a significant amount of other repetitive sequences (accounting for up to 50% of the genome). Aside from the pseudogenes already mentioned, the MAIN TYPES OF such repeats are:
1. Tandem repeats. These include multiple repeats of short sequences (6–8 Base Pairs) in telomeres and repeats of the so-called α-satellite DNA in centromeres (repeat lengths vary from 7 base pairs in Drosophila to 100–200 base pairs in mammals, and 171 base pairs in humans). So-called simple sequence repeats (SSRs) are also distributed throughout the genome. Typically, a distinction is made between microsatellites—1–15 base pairs repeated from 10 to several thousand times—and minisatellites—15–500 base pairs repeated up to 100 times. The number of mini- and microsatellite loci reaches tens and hundreds of thousands. The distribution of loci based on repeat count serves as a specific individual trait—much like fingerprints.
2. Segmental duplications—large blocks of 1–200 kb characterized by a high degree of Homology. These duplications are presumably the products of past chromosomal rearrangements and are most frequently found in pericentromeric and subtelomeric regions.
3. Interspersed (mobile) elements capable of movement and Replication within the genome, which constitute the bulk of repetitive DNA. Some of these sequences result from the past Activity of Mobile elements (having lost their ability to transpose). The primary types of eukaryotic mobile elements are:
✵ DNA Transposons—movement occurs via the excision of a DNA segment followed by its insertion elsewhere, which is entirely analogous to corresponding elements in prokaryotes. Transposons contain one or two genes (depending on the species) encoding transposase—the enzyme responsible for transposition, i.e., excising the element from the donor site and integrating it into the target site. Transposase genes may be defective, in which case transposition of such an element relies on a transposase encoded by another DNA transposon.
The coding region of the transposon is flanked by short inverted repeats recognized by the transposase during DNA excision. The target site is a small specific DNA sequence also recognized and cleaved by the transposase, which then catalyzes the Integration of the transposon into the target site (Fig. 6.3). The transposition process leaves a double-stranded break at the site previously occupied by the transposon. In replication-independent (non-replicative) transposition, this break is repaired via non-homologous end joining (see Chapter 1), meaning the transposon simply "jumps" to a new Location. However, transposition quite frequently occurs during replication (replicative transposition)—in this case, the break is repaired via Homologous Recombination (see Fig. 1.26): the sister DNA molecule is used as a template, and the region containing the transposon is restored. Thus, the transposon both jumps to a new location and remains at the donor site, effectively "multiplying."

Fig. 6.3. Mechanism of DNA transposon mobility
✵ LTR Retrotransposons—sequence elements containing Long Terminal Repeats (LTRs) and several genes, notably the Reverse Transcriptase and integrase (transposase analog) genes. As with the next Two Types of mobile elements, movement proceeds through an RNA intermediate. The transposition of an LTR retrotransposon copy resembles the retroviral life cycle (see Fig. 5.9). The First stage is Transcription of the retrotransposon, after which the synthesized mRNA is transported to the Cytoplasm for Translation. Reverse transcriptase, produced by this translation, synthesizes DNA using the mRNA as a template and the 3' end of a tRNA molecule as a primer. The resulting DNA copy of the retrotransposon, complexed with integrase, returns to The Nucleus, where the DNA is integrated into the genome.
✵ LINEs (Long INterspersed Elements) contain several genes, including the reverse transcriptase gene. Following transcription and subsequent Introduction/27.html">Translation of the mRNA in the cytoplasm, the synthesized proteins bind to the mRNA, and this complex returns to the nucleus, where reverse transcription and integration into the genome take place. mRNA synthesis during LINE transcription, as with most other eukaryotic mRNAs, terminates at a polyA signal (see Chapter 2). This signal is weak, allowing the element to insert into the introns of conventional genes without major disruption to Gene Expression, as the Processing machinery often overlooks the weak polyA signal. Consequently, LINE elements are extraordinarily abundant mobile elements in the genomes of higher eukaryotes.
Occasionally, they act not merely as pieces of "selfish DNA" autonomously multiplying within the genome, but perform specific biological functions. For instance, Drosophila lacks a telomerase system, and specific types of LINE elements serve to extend chromosome ends following replication: reverse transcriptase acts as a telomerase, while the mobile element's mRNA serves as the telomerase RNA template (see Chapter 1). Interestingly, the DNA sequences of the telomerase gene and LINE elements show high homology, suggesting that the telomerase system may have evolved from LINE mobile elements.
✵ SINEs (Short INterspersed Elements)—short (100–400 bp) non-coding elements that co-opt the enzymatic machinery of LINEs for their mobilization. This class includes the so-called Alu repeat (named after the specific restriction endonuclease capable of cleaving this sequence element).
Mobile elements are unevenly distributed throughout the genome: there are long stretches consisting of up to 90% mobile elements, alongside regions where interspersed elements are entirely absent. Overall, a negative correlation is observed between gene density and mobile element density. An exception to this rule is the positive correlation between gene density and SINE-type elements.
Last update: 11/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.