BASICS OF MEDICAL BIOLOGY - 2012
Structure of Prokaryotic and Eukaryotic Genes. Structural, Regulatory, tRNA, and rRNA Genes
A gene is the fundamental Structural and functional unit of heredity that determines The Development of a specific trait in a Cell or Organism. A gene is a functional unit of Genetic information; a segment of a DNA molecule (or RNA in some Viruses) that contains the information for the Amino Acid Sequence of a polypeptide (protein) molecule or The nucleotide sequence in a ribonucleic acid molecule (rRNA or tRNA).
One of the fundamental principles of genetics is that all traits and properties of organisms are determined by elementary discrete units of heredity—genes (from Greek genos — race, origin). Genes are inherited, thereby ensuring structural and functional continuity between generations. The existence of discrete hereditary factors was first established by G. Mendel in 1865. The Danish geneticist W. Johannsen named them genes in 1909. However, Johannsen did not associate the localization of genes with the Chromosomes of the Cell Nucleus. Our understanding of the gene underwent a radical transformation As a result of the work of T. Morgan and his students, giving METABOLISM/2.html">THE CONCEPT OF the gene in the chromosome theory of heredity a material foundation.
In the 1940s, G. Beadle and E. Tatum experimentally substantiated the "one gene - one protein" hypothesis through Experiments on the fungus Neurospora crassa. It was subsequently established that many Proteins consist of several polypeptide chains, the synthesis of which is coded by separate genes. The "one gene - one protein" hypothesis evolved into its modern form: "one gene - one polypeptide." At the suggestion of the American physicist and geneticist S. Benzer in 1957, three new sub-concepts of the gene were introduced: recon, muton, and Cistron. Benzer called the smallest functional unit of genetic material that encodes the synthesis of a specific polypeptide a cistron. One cistron encodes one polypeptide. The terms "cistron" and "gene" are practically used as synonyms. A gene has a complex structure and is capable of mutation and recombination. Benzer designated the smallest segment of a cistron whose alteration can cause a mutation as a muton (the unit of mutation), and the smallest segment that is indivisible by recombination as a recon (the unit of recombination). Both the muton and the recon correspond to a single base pair.
Each interphase chromosome contains a single DNA molecule holding a vast number of genes. The Human Genome (DNA) size is 3.5x109 Base Pairs, which is theoretically sufficient to form about 1.5 million genes. However, according to expert estimates, the human genome contains only 35,000 to 40,000 genes. This means that the greater part of nuclear DNA is not translated into Amino acid sequences. According to current data, only 1-2% of human cellular DNA contains information for encoding the body's proteins. Most of the DNA is in a repressed (inactive) state: a significant part of The Genome is utilized during embryonic development, differentiation, and growth and is subsequently not transcribed; another part constitutes introns; and the bulk of DNA consists of numerous non-transcribed repeat sequences (satellite DNA). Thus, Eukaryotic Genomes contain both unique and repetitive DNA nucleotide sequences. Unique sequences are represented in the genome by a single copy. They comprise the bulk of genes and account for half of the human genome. Repetitive sequences are represented by multiple copies (from 2 to 107). Moderately repetitive and highly repetitive sequences are distinguished. Moderately repetitive sequences have no more than 106 copies, alternate with unique nucleotide sequences, and are characterized by the presence of inverted repeats (palindromes). Highly repetitive sequences are short segments ranging from 5 to 500 base pairs in length that are arranged consecutively (in tandem). Such sequences form clusters containing up to 10 million copies and make up the satellite DNA fraction. This fraction is localized in regions of constitutive heterochromatin, predominantly near centromeres and telomeres, is genetically inert (non-transcribed), and accounts for 12% of the human genome.
According to the chromosomal theory of heredity, genes are located in chromosomes in a specific linear order (one after another) and do not overlap. Each gene occupies a specific site—a locus. Exceptions to this rule have been discovered, known as overlapping genes. The phenomenon of gene overlapping ("a gene within a gene") was found in bacteriophage ΦX174 and other viruses. Telomeric and centromeric regions of chromosomes contain no genes.
Based on the Organization of NUCLEOTIDES in DNA, it can be divided into the following fragments:
1) Structural genes carry information about The structure of specific Polypeptides. mRNA is transcribed from these DNA regions.
2) Regulatory genes control and regulate The process of Protein Biosynthesis.
3) Satellite DNA contains A large number of repetitive nucleotide groups that are non-coding and non-transcribed. Single genes interspersed within satellite DNA exert a regulatory or enhancing effect on structural genes.
4) Spacer DNA consists of large, non-transcribed DNA regions whose exact role is not fully understood. Spacer regions are most commonly found between gene clusters.
5) Gene clusters are groups of different structural genes within a specific chromosomal region that share common Functions.
6) Repetitive genes consist of the same gene repeated many times (hundreds of times) consecutively without Separation. An example is the rRNA genes.
Structure of Eukaryotic Genes. The elucidation of eukaryotic gene structure is one of the major scientific breakthroughs of the late 20th century. It has been proven that structural genes encoding proteins have a complex architecture.
Organization of a eukaryotic gene using the human Hemoglobin β-chain gene as an example.

The promoter is a region of the regulatory part of a gene to which the enzyme RNA polymerase binds. It includes a non-specific region known as the TATA box, which contains specific sites: a recognition site, a binding site, and an initiation site.
Downstream of the promoter lies the Transcription regulation factor-binding region, known as the operator (O). Next is the CAP sequence (TTAGGTTAC), where Transcription is initiated and the 5'-initial region of RNA is formed (the Transcription initiation site). Further along is the TAC codon (the Translation initiation site), which corresponds to the AUG codon in the mRNA molecule.
Between these sites lies a DNA region (~50 bp) called the leader sequence, followed by the structural part of the gene. At its beginning is exon (E1) of 50 bp, which encodes the first 30 Amino Acids of the hemoglobin β-chain, followed by intron (I1), consisting of 130 bp, then E2, I2...
The termination region begins with the TAA codon, followed by the trailer and the polyadenylation site (AATAAA), which is necessary for The formation of a poly-A tail of 200–300 adenylic nucleotides in the RNA. Transcription then terminates. Further downstream is a sequence known as the transcription terminator (approximately 1000 bp).
Additionally, at a distance of 600–900 bp from the poly-A site, there is an enhancer (a region that accelerates the reading of genetic information).
Thus, the gene structure encodes information not only about the polypeptide structure, but also about transcription processes, mRNA structure, and translation.
Properties of genes include discreteness, stability, polyallelism, Specificity, and pleiotropy. Discreteness means that the development of different traits is controlled by distinct genes whose chromosomal localizations do not coincide. Stability refers to the fact that, in the absence of Mutations, a gene is transmitted unchanged across generations. Polyallelism (multiple allelism) is the presence of more than two alleles of the same gene within a population. Specificity means each gene governs the development of a specific trait or group of traits. Pleiotropy is the ability of a single gene to control the development of multiple traits (e.g., Marfan Syndrome).
Mobile (migrating) genetic elements (MGEs) are DNA nucleotide sequences capable of moving (transposition) within a given genome or between genomes. They have been found in the genomes of viruses, Bacteria, Yeasts, Drosophila, mice, and other organisms. They are associated with the occurrence of mutations and variations. Thanks to MGEs, organisms can acquire pre-existing genes from other organisms, which plays a significant role in evolution. MGEs were first discovered by the American geneticist B. McClintock (1947) while studying seed coloration in maize, and she named them "controlling elements" (Nobel Prize, 1983). They have no phenotypic expression of their own and manifest themselves through Changes in the genes in whose loci they reside or which they affect from a distance.
It is believed that there are at least two mechanisms of MGE mobility. The first involves the excision of the mobile element from one site in a chromosome and its insertion into another site. In the second mechanism, the mobile element retains its original position while creating a copy of itself, which is subsequently inserted elsewhere.
The biological (genetic) code and its properties represent The system of nucleotide base pairing in a DNA molecule that determines The amino acid sequence in a protein molecule. The Diversity of protein molecules is determined by the composition and arrangement of amino acids within polypeptide chains. The amino acid sequence that defines a protein's properties is encoded in the DNA molecule via the biological code. In 1954, G. Gamow suggested that Amino acids are coded in DNA by combinations of multiple nucleotides. Over the following years, it was experimentally proven that The Genetic Code is triplet-based.
Each amino acid is encoded by a triplet of three nucleotides, known as a codon, located within the structural region of a gene. The number of possible triplet combinations formed from four types of nucleotides is 64 (43). Of these, only 61 triplets encode protein amino acids. This is sufficient to code for the 20 most common Natural Amino Acids found in proteins. Three triplets are nonsense or stop codons (UAA, UAG, UGA); they do not code for amino acids, but instead signal the termination of translation (polypeptide synthesis).
Properties of the genetic code:
1. Triplet nature: a single nucleotide triplet (codon) encodes a single amino acid.
2. Specificity: each specific triplet designates a specific amino acid.
3. Degeneracy (redundancy): Certain amino acids are encoded by multiple triplets. For instance, phenylalanine is encoded by 2 triplets (UUU, UUC), whereas Serine is encoded by 6 triplets (UCU, UCC, UCA, UCG, AGU, AGC).
4. Non-overlapping: each nucleotide is part of only a single triplet (e.g., AAT — GGC — TTA).
5. The genetic code is read sequentially, triplet by triplet, and incorporates nonsense codons.
6. Stop codons (UAA, UAG, UGA), which do not code for amino acids, signal termination (the cessation of reading information from RNA during translation).
7. Colinearity: a polypeptide and the gene encoding it are colinear, meaning the amino acid sequence in the protein molecule corresponds directly to the nucleotide sequence in the gene.
8. Universality: the genetic code is uniform across all living nature. A specific codon in DNA or RNA dictates the exact same amino acid in the protein-synthesizing systems of All living organisms, from viruses to humans. However, certain exceptions exist:
Minor deviations in the genetic code have been identified in certain viruses and Mitochondrial DNA of specific species. For example, the triplet ACT (nonsense codon) encodes the amino acid Tryptophan (read as ACC). In yeasts, the triplet GAT encodes Threonine. In mammals, the codon TAG encodes Methionine instead of isoleucine, while the triplets TCG and TCT function as nonsense codons in the Mitochondria of certain species.
THE GENETIC CODE
First |
Second base |
Third |
|||
base |
U (A) |
C (G) |
A (T) |
G (C) |
base |
U (A) |
Phe |
Ser |
Tyr |
Cys |
U (A) C (G) A (T) G (C) |
Phe |
Ser |
Tyr |
Cys |
||
Leu |
Ser |
— |
— |
||
Leu |
Ser |
— |
Trp |
||
C (G) |
Leu |
Pro |
His |
Arg |
U (A) C (G) A (T) G (C) |
Leu |
Pro |
His |
Arg |
||
Leu |
Pro |
Gln |
Arg |
||
Leu |
Pro |
Gln |
Arg |
||
A (T) |
Ile |
Thr |
Asn |
Ser |
U (A) C (G) A (T) G (C) |
Ile |
Thr |
Asn |
Ser |
||
Ile |
Thr |
Lys |
Arg |
||
Met |
Thr |
Lys |
Arg |
||
G (C) |
Val |
Ala |
Asp |
Gly |
U (A) C (G) A (T) G (C) |
Val |
Ala |
Asp |
Gly |
||
Val |
Ala |
Glu |
Gly |
||
Val |
Ala |
Glu |
Gly |
||
First nucleotides represent RNA codons; DNA codons are given in parentheses.
Last update: 08/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.