Genetics - A. V. Sivolob 2008

Gene Expression
What is a gene?

The term "Gene" was coined by Wilhelm Johannsen in 1909, shortly after the rediscovery of Mendel's Laws of heredity (see the Historical Overview at the end of the textbook). Since then, METABOLISM/2.html">THE CONCEPT OF the gene has undergone significant revision several times. Some of its various Structure/97.html">Definitions were outlined in Chapter 1, and it should be noted that all of them are valid. However, It is important to understand that each of them also has rather significant limitations. The only undisputed fact is that a gene is a segment of DNA (even this statement requires qualification regarding Viruses that contain RNA as their genetic material, see Chapter 5); yet, such a definition says nothing about the properties this segment must possess to be considered a gene.

From its inception, the term was interpreted in the Mendelian spirit, that is, as a discrete and indivisible hereditary unit that is independent of other hereditary units and is responsible for the manifestation of a particular trait. In certain cases (such as the pea gene sgr mentioned in Chapter 1), this interpretation is applicable, although it is understood that independent inheritance of different genes can only be discussed on the condition that the genes reside on different Chromosomes.

It is also clear that a gene is neither discrete nor indivisible: as a segment of DNA, a gene has a length, and this segment can be divided into fragments. During their expression, genes interact at various levels: Transcription activation and the splicing pathway depend on The activity of transcription factor genes and splicing regulators; mRNA Translation depends on the activity of translation regulatory protein genes; protein products of different genes interact, and so on. As a result, the activity of one gene can enhance or suppress the expression of another, and generally, the manifestation of a trait typically requires the activity of multiple genes. Recently, a computer metaphor has gained popularity, according to which The hereditary apparatus (The Genome and its expression system) can be viewed as an "operating system" that controls the Organism, and a gene as a "subroutine" of this system.

After it was established that genes are located on chromosomes, the gene began to be viewed as a chromosomal locus, and the chromosome itself as a linear combination of non-overlapping genes. Often this is indeed the case, but on the other hand, a single locus (a single region of chromosomal DNA) can contain multiple genes—genes may overlap either due to the overlap of reading frames (in some Bacteriophages and certain eukaryotic genes), or due to the arrangement of genes (in Eukaryotic Genomes) within the intron of another gene, or due to the placement of two coding sequences on the same DNA region across two different strands (as shown in Fig. 2.17). Furthermore, the Concept of the gene as a locus is not applicable to mobile elements—DNA segments (often containing one or more genes) that can alter their localization within the genome.

The Development of molecular biology initially led to the understanding that a gene is a segment of a DNA molecule responsible for the synthesis of a protein molecule ("one gene - one protein"). Sometimes this is indeed the case, but today it is already clear that, firstly, genes encoding various non-translated RNAs are no less important. Secondly, a single DNA segment (the set of exons of a eukaryotic gene) often yields multiple protein products through Alternative Splicing. Differences between biological species are often caused not only, and not so much, by differences in the sets of coding sequences (exons), but rather by different combinations of these exons. Moreover, such recombination is possible both at the DNA level and at the level of final transcripts. Therefore, if the gene is interpreted as a hereditary factor, it is not merely a DNA segment containing specific information, but also the system for expressing that information.

The intensive development over the past 10–15 years of a new discipline—Genomics, aimed at determining and analyzing The nucleotide sequences of entire genomes—has led to a tendency to view the gene as an annotated genomic region with specific properties. According to the definition by the Sequence Ontology Consortium, a gene is a defined region of a genomic sequence that corresponds to a unit of heredity and contains regulatory regions and transcribed regions. The phrase "unit of heredity" implies that the gene encodes certain (one or more) functional products (Proteins or non-translated RNA molecules). The "transcribed region" refers to a specific group of exons, joined by introns, that is transcribed as a single unit. In this context, according to sequence annotation rules adopted by modern genomic sequence Databases, primary transcripts undergoing alternative splicing are considered to belong to the same gene, even if the final proteins are different. That is, a gene is a group of co-transcribed exons, or a gene is a genomic region that produces a set of final transcripts sharing at least one common exon. Finally, an important aspect of the provided gene definition is that regulatory elements controlling its activity are conventionally included as part of this elementary unit of organismal heredity.

With the qualification that in prokaryotic systems (in the case of operons) regulatory regions can control a group of genes, this interpretation of the gene remained generally accepted until very recently.

To thoroughly analyze the information recorded in the genome and realized through transcription, the US National Human Genome Research Institute launched the international ENCODE (Encyclopedia of DNA Elements) project four years ago, one of the MAIN OBJECTIVES OF which is a comprehensive Analysis of the human transcriptome. The First stage of this work was recently completed—the functioning of 1% of The Human Genome (approximately 30 million Base Pairs) was analyzed, and the initial results brought A number of surprises (The ENCODE Project Consortium // Nature, 2007, Vol. 447, P. 799-816).

First, the relative amount of transcribed DNA turned out to be unexpectedly high—around 80%. Considering the proportion of the genome accounted for by exons together with introns (see Fig. 1.10), the question arises: do all these transcripts correspond to genes? Presumably, some of these primary transcripts are merely a consequence of non-specific, chaotic RNA polymerase activity. However, a significant portion of transcripts contains sequence elements that are conserved (among mammals and human populations)—about 60% of such elements lie outside previously known protein-coding genes or regulatory regions. It is likely that in many cases these transcripts are previously unknown non-translated RNAs, and their functional significance for the most part remains to be elucidated. In addition, it was found that regulatory regions (promoters, enhancers) are frequently transcribed. Likely, such transcription is simply one of the ways to maintain the regulatory region in an accessible, decondensed Chromatin fiber state.

For known protein-coding genes (399 in the studied region of the genome), the presence of A large number of previously unknown transcription start sites was demonstrated. Approximately half of the genes have alternative start sites located 100 kb away from the previously annotated transcription starts of these genes. Some of these start sites utilize promoters of other genes: a single such site can be shared by two or three genes, and the primary transcript sometimes encompasses multiple gene loci—groups of exons (see Fig. 2.17).

Furthermore, for most protein-coding genes, the analysis of their transcripts reveals the presence of previously unknown exons. Some of these exons are located up to several thousand base pairs away from all other exons of the gene, sometimes falling within another gene. As Fig. 2.17 demonstrates, it is sometimes difficult to assign a given exon to one specific gene or another. The number of various mRNA isoforms arising as a result of both alternative splicing and trans-splicing turned out to be greater than expected.

The results of the ENCODE project indicate a dispersed distribution of regulatory elements throughout the genome: many regulatory elements are located within exons and introns, and they may serve as regulatory elements for an entirely different gene.

Thus, the concept of the gene once again requires some revision. According to one recently proposed definition, a gene is a union of genomic sequences encoding a coherent set of functional products that may partially overlap (Gerstein et al. // Genome Res., 2007, Vol. 17, P. 669–681). The coherence of a set of products implies that, in the case of protein-coding genes, every exon is shared by at least two products of that set.

The main emphasis in this definition is placed on the final products of gene activity—overlaps between intermediate transcripts are ignored. If we proceed from overlaps between primary transcripts (interpreting the gene as a cluster of exons that can be co-transcribed), then, for example, only two genes should be defined in the genomic region of Fig. 2.17 (by combining genes 1, 2, and 4 into one). If, however, we proceed from overlaps between final products, this region contains at least six genes (genes 1–2 and 1–4 form separate groups of exons that partially overlap with genes 1, 2, and 4). That is, a gene does not necessarily consist of adjacent exons; a group of exons belonging to a single gene can be dispersed across a genomic region, and individual exons of the group can simultaneously belong to other genes. Naturally, in the simple case where there is no alternative splicing (or no introns at all), the definition reduces to the classical one: a gene is a DNA segment encoding a protein or RNA molecule.

Furthermore, according to the definition under Discussion, regulatory elements are not Components of the gene (it is proposed to call them "gene-associated elements"): the regulatory system is more complex than a simple one-to-one relationship between regulatory elements and coding regions.

The given definition certainly cannot be considered final. Obviously, the concept of the gene is too complex to be formulated precisely: different definitions focusing on various properties of Different types of genes are, in accordance with the complementarity principle, simultaneously valid.

Selection/41.html">Review Questions and Exercises

1. WHAT IS A codon and an Open Reading Frame? How many codons exist? Which nucleotide positions within a codon are the most and least determining?

2. Characterize the Stages of Protein Synthesis. What role do Ribosomes and tRNAs play in Protein Synthesis?

3. What are the characteristic differences in the Mechanisms of Genetic information expression between PROKARYOTES AND EUKARYOTES?

4. In what direction is information read from DNA during transcription? What are the sense and antisense strands? Which of them serves as the template?

5. What is the difference between cis- and trans-acting elements of the transcription regulation system?

6. Describe the regulation System of the lactose Operon.

7. What is The structure of a bacterial promoter and a eukaryotic RNA polymerase II promoter?

8. What is the specialization of various types of eukaryotic RNA polymerases?

9. State the core principles of transcription regulation involving transcription factors.

10. How are transcription and mRNA Processing coordinated in eukaryotes?

11. Describe the primary mechanisms of transcription activation in eukaryotes.

12. What is RNA Interference?

13. How is the selection of alternative splicing pathways regulated?

14. Define trans-splicing and explain the mechanisms by which it occurs.

15. Provide several definitions of a gene and explain the limitations of each.



Last update: 11/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.