Fundamentals of Molecular Biology - V. I. Rezyapkin 2009
Molecular Organization of Genes
A Gene is the elementary unit of hereditary information, occupying a specific locus in The Genome and responsible for performing specific Functions within the Organism.
Genes are fragments of NA that encode polypeptide chains or RNAs, such as rRNA, tRNA, snRNA, and others. Proteins may consist of one or more polypeptide chains, which can be identical or different from one another. If a protein consists of several distinct polypeptide chains, each chain corresponds to a specific gene (Fig. 7.1).
Class="center">
Fig. 7.1. Polypeptide chains of Oligomeric Proteins can be encoded by different genes
Cell/35.html">Mitochondria contain proteins whose polypeptide chains are partly encoded in the nuclear genome and partly in the Mitochondrial Genome. Obviously, the synthesis of polypeptide chains encoded in the nuclear genome occurs in the Cytoplasm, while others are synthesized in the mitochondria. Polypeptide chains synthesized in the cytoplasm are transported into the mitochondria. Within these Organelles, the native oligomeric protein is assembled (Fig. 7.2).

Fig. 7.2. Mitochondrial proteins may have a mixed origin
Genes can be unique—represented by a single copy in the genome—or repetitive, represented by several or many copies. For instance, histone genes are present in some eukaryote species in up to 1,000 copies, and rRNA genes are also highly repetitive. Moreover, the number of rRNA genes in amphibian oocytes can increase threefold through Amplification. Nevertheless, the majority of eukaryotic genes, relative to the haploid genome, are represented by a single copy or a small number of copies. Prokaryotic genes are typically present in the genome as a single copy.
Genes present in a high copy number may be scattered across the genome or located adjacent to one another, forming gene clusters. Histone, rRNA, and globin genes (the proteins that make up Hemoglobin), among others, are organized into clusters. Repetitive genes located sequentially one after another are called tandem genes. A tandem gene cluster is formed by repeating units, which in turn consist (Fig. 7.3) of a METABOLISM/31.html">Transcription unit (the transcribed DNA region) and a non-transcribed spacer (the non-transcribed DNA region).

Fig. 7.3. A tandem gene cluster is formed by repeating units consisting of a transcription unit (1) and a non-transcribed spacer (2)
As noted above, histone genes have a clustered Organization. There are five main Histones: H1, H2A, H2B, H3, and H4. These are related proteins that perform a common function related to maintaining a specific Chromatin Structure. All five histone genes are part of a repeating unit, the sequential arrangement of which forms a cluster. Figure 7.4 shows the repeating unit of the histone gene cluster in Drosophila melanogaster. In other eukaryotic species, the repeating unit of such a cluster is organized differently. In some cases, histone genes lack a clustered organization.
![]()
Fig. 7.4. Repeating unit of the histone gene cluster in Drosophila melanogaster. The arrow indicates the direction of gene transcription
Genes are located on both strands of DNA (Fig. 7.5). Non-coding sequences are usually situated between them.

Fig. 7.5. Genes are located on both DNA strands
At the same time, some genes can overlap, meaning they share a common DNA sequence. In some cases, the end of one gene serves as the beginning of another (Fig. 7.6A); in others, one gene is located entirely within another gene (Fig. 7.6B). Genes located on complementary strands can also share common regions (Fig. 7.7). Gene overlapping has been found, in particular, in Viruses. This feature of Introduction/29.html">Gene Organization allows for DNA conservation, which is especially important in viruses since the size of their NAs is limited by the volume of the viral particle. There are known instances in viruses where three genes overlap within a single DNA region. In some Mitochondrial Genes, the last nucleotide of one gene is the first nucleotide of another. Genes may overlap with or without a frameshift. If genes overlap without a frameshift, the Polypeptides encoded by them will have identical Amino acid sequences in the overlapping region. Conversely, if genes overlap with a frameshift, the encoded polypeptides will not share identical amino acid sequences.

Fig. 7.6. Overlapping genes

Fig. 7.7. Genes located on complementary DNA strands may share common regions
All genes can be divided into two groups:
a) constitutive genes, or "housekeeping genes". These genes are constantly expressed: they function at all stages of ontogenesis and in all Tissues. They include genes for tRNAs, rRNAs, DNA polymerases, RNA polymerases, ribosomal proteins, histones, and genes controlling continuous metabolic processes;
b) inducible genes, or "luxury genes". These genes can be switched on and off. Inducible genes include those controlling the course of ontogenesis as well as genes determining the Structure and function of cellular and whole-organism components. Turning on inducible genes is called induction, and turning them off is called repression.
Genes can undergo Mutations, which are alterations in The nucleotide sequence of a DNA strand. Mutations may occur As a result of nucleotide substitution, insertion, or deletion. They can lead to Changes in the Amino Acid Sequence of the polypeptide encoded by the gene, and consequently, to alterations in the biological CHARACTERISTICS OF THE protein. However, due to the degeneracy of The Genetic Code, not all changes in the nucleotide sequence result in changes to the Primary Structure of the protein or its biological functions.
Mutations can potentially lead to the transformation of genes into pseudogenes. Pseudogenes are gene homologues that are not expressed to form a functionally active product. The causes of Gene Conversion into pseudogenes may include:
a) transcriptional impairment;
b) aberrant RNA Processing;
c) translational impairment;
d) other causes.
Both prokaryotic and eukaryotic genes consist of regulatory and transcribed regions. The regulatory region ensures the initiation of Transcription of the transcribed region. The latter carries information regarding the polypeptide chain or specific types of RNA (tRNA, rRNA, etc.). Transcription termination takes place in the distal region of the gene, known as the terminator.
In prokaryotes, genes encoding Proteins of the same metabolic pathway can be linked and transcribed from a single promoter as a single RNA molecule, the Translation of which yields multiple polypeptides (Fig. 7.8).

Fig. 7.8. In prokaryotes, genes can be transcribed from a single promoter as a single RNA molecule, the translation of which produces multiple polypeptides. P — promoter, T — terminator.
A group of genes sharing a common promoter and terminator is called an Operon. Figure 7.9 illustrates the organization of a polypeptide-encoding prokaryotic operon. The regulatory region of such an operon consists of a promoter—the site where RNA polymerase binds—along with other nucleotide sequences that interact with regulatory proteins to enhance or suppress transcription. The transcribed region comprises sequences encoding polypeptide chains flanked by untranslated polynucleotide sequences. The operon arrangement of genes ensures the coordinated synthesis of proteins involved in a common biological function.

Fig. 7.9. Operon organization
Protein-coding genes in prokaryotes can also exist as single genes consisting of a regulatory and a transcribed region (Fig. 7.10).

Fig. 7.10. Single prokaryotic gene
Prokaryotic rRNA genes possess an operon structure. Interestingly, such operons also include tRNA genes (Fig. 7.11 A). Prokaryotic tRNA genes may likewise be represented as single genes or incorporated into an operon (Fig. 7.11 B).

Fig. 7.11. Operon Organization of Prokaryotic genes. A — operon containing rRNA and tRNA genes; B — operon containing tRNA genes
Eukaryotic genes have a more complex structure. All eukaryotic genes can be divided into three groups. The first group of genes is transcribed by RNA polymerase I, and includes the genes for 18S, 28S, and 5.8S rRNAs. The second group is transcribed by RNA polymerase II, encompassing genes that encode polypeptides and certain snRNAs. The third group is transcribed by RNA polymerase III, which includes genes encoding tRNAs, 5S rRNA, and other snRNAs. The genes of all three groups differ in their promoter organization, which dictates their transcription by specific RNA polymerases. The structure of promoters recognized by different RNA polymerases was discussed earlier in the section "Transcription". The organization and expression of rRNA genes were also covered in that section.
Another characteristic feature of eukaryotic RNA polymerases is their inability to initiate transcription independently. They require assistance from protein factors known as transcription factors to start transcription.
The rate of Eukaryotic Transcription is influenced by a variety of regulatory proteins. Activator proteins enhance transcription, whereas repressor proteins exert an inhibitory effect on Gene Expression.
Next, we will examine the organization of genes transcribed by RNA polymerase II in greater detail. The regulatory region of such genes consists of a promoter—where the RNA polymerase complex and general transcription factors assemble—and regulatory sequences (enhancers and silencers) that bind regulatory proteins. Six general transcription factors are known: TFIIA, TFIIB, TFIID, TFIIE, TFIIF, and TFIIH. They are essential for Transcription initiation. One of the subunits of TFIIH possesses protein kinase activity, catalyzing the phosphorylation (addition of a phosphate group) of RNA polymerase II. Phosphorylation of the enzyme is required for transcription initiation. In addition to general transcription factors, other transcription factors participate in regulating RNA polymerase II activity. The transcription complex also includes proteins that assist RNA polymerase in disrupting nucleosomes.
Eukaryotic primary transcripts are considerably longer than mRNA. This is because the transcribed region of a eukaryotic gene consists of exons and introns. During the transcription of such genes, a precursor is synthesized in which exons are separated by introns. Splicing removes the introns and joins the exons together. The Mechanism of splicing and its variants are discussed in the section "RNA Processing".
Not all eukaryotic genes have an exon-intron organization, and the proportion of genes with introns varies across species. As living organisms become more complex, the proportion of genes consisting of exons and introns increases: for instance, it is 5% in Yeast, 83% in Drosophila, and 94% in mammals.
In turn, the number and size of introns and exons vary (see table). For instance, exon size typically ranges from 100 to 600 bp. Intron size, however, varies over a much wider range, from several tens to tens of thousands of bp. Furthermore, the total length of introns frequently exceeds the total length of exons. For example, in the chicken Ovalbumin gene spanning 7000 bp, exons account for 1872 bp.
Table 7.1
Characteristics of Exons and Introns in Selected Genes
|
Gene |
Organism |
Exon length, bp |
Introns |
|
|
number |
total length, bp |
|||
|
a-Interferon gene |
human |
600 |
0 |
0 |
|
Adenosine deaminase gene |
human |
1500 |
11 |
30000 |
|
Apolipoprotein B gene |
human |
14000 |
28 |
29000 |
|
ß-Globin gene |
mouse |
432 |
2 |
762 |
|
a-Globin gene |
mouse |
463 |
2 |
256 |
|
silkworm |
18000 |
1 |
970 |
|
|
Hypoxanthine-phosphoribosyltransferase gene |
mouse |
1307 |
8 |
32000 |
|
Erythropoietin gene |
human |
582 |
4 |
1562 |
|
Phaseolin gene |
bean |
1263 |
5 |
515 |
Given that introns are removed during RNA splicing, their existence might seem pointless. However, the following evidence refutes this notion. The deletion of introns in certain genes leads to the death of the organism. Another highly interesting fact is that an entire separate gene can reside within an intron, which in turn may contain its own intron.
In Ciliates, the removal of sequences corresponding to introns occurs at the DNA level during macronuclear/micronuclear DNA maturation. Interestingly, the excision of DNA fragments can be accompanied by the shuffling of exon-corresponding regions (Fig. 7.12).

Fig. 7.12. In ciliates, intron-corresponding sequences are excised at the DNA level, which may be accompanied by the shuffling of exon-corresponding regions
Unlike prokaryotic genes, eukaryotic genes are not organized into operons. The Expression of Eukaryotic genes yields monocistronic mRNA, the translation of which produces a single polypeptide.
Last update: 12/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.