GENETICS - Textbook - A.V. Syvolob - 2008

CHAPTER 2. Gene Expression

WHAT IS A GENE?

The term "Gene" was proposed by Wilhelm Johannsen in 1909, shortly after the rediscovery of Mendel's Laws of inheritance (see the Historical Overview at the end of the textbook). Since that time, METABOLISM/2.html">THE CONCEPT OF the gene has undergone significant revision on several occasions. Some of its various Structure/97.html">Definitions were provided in Chapter 1, and It is worth noting that all of them are valid. However, It is important to understand that each of them also has rather significant limitations. The only undisputed fact is that a gene is a region of DNA (and even this statement requires clarification regarding Viruses that contain RNA as their genetic material, see Chapter 5); yet, such a definition says nothing about the properties that this region must possess to be considered a gene.

From its inception, the term was interpreted in the Mendelian spirit, that is, as a discrete and indivisible hereditary unit that does not depend on other hereditary units and is responsible for the manifestation of a specific trait. In certain cases (the pea gene sgr mentioned in Chapter 1 is an example of this type), such an interpretation can be applied, although it is clear that the independent inheritance of different genes can only be discussed provided that the genes are located in different Chromosomes.

It is also clear that a gene is neither discrete nor indivisible: as a region of DNA, a gene has a certain length, and this region can be divided into fragments. In The process of their expression, genes interact at various levels:

transcriptional activation and the splicing pathway depend on The activity of Transcription factor genes and splicing regulators, mRNA Translation depends on the activity of translation regulatory protein genes, protein products of different genes interact with one another, and so forth. As a result, the activity of one gene can enhance or inhibit the expression of another, and in general, the manifestation of a trait typically requires the activity of multiple genes. Recently, the computer metaphor has gained popularity, according to which The hereditary apparatus (The Genome and its expression system) can be viewed as an "operating system" that controls the Organism, and a gene as a "subroutine" of this system.

After it was established that genes reside in chromosomes, the gene came to be viewed as a chromosomal locus, and the chromosome itself as a linear combination of non-overlapping genes. Often this is indeed the case; on the other hand, a single locus (a single region of chromosomal DNA) can contain multiple genes—genes can overlap either due to the overlap of reading frames (some Bacteriophages and certain eukaryotic genes), or due to the Location of genes (in Eukaryotic Genomes) within the intron of another gene, or due to the arrangement of two coding sequences on the same DNA region across two different strands (as shown in Fig. 2.17). Furthermore, the Concept of the gene as a locus is not applicable to mobile elements—DNA regions (often containing one or more genes) that can alter their localization within the genome.

The Development of molecular biology initially led to the understanding that a gene is a region of a DNA molecule responsible for the synthesis of a protein molecule ("one gene - one protein"). Sometimes this is indeed the case, but today it is already clear that, first, genes encoding various non-coding RNAs are no less important. Second, a single DNA region (the set of exons of a eukaryotic gene) frequently yields multiple protein products through Alternative Splicing. Differences between biological species are often caused not only, and not so much, by differences in the sets of coding sequences (exons), but rather by different combinations of these exons. Moreover, such recombination is possible both at the DNA level and at the level of final transcripts. Therefore, if the gene is interpreted as a hereditary factor, it is not merely a DNA region containing specific information, but also the system for expressing that information.

The intensive development over the past 10–15 years of a new discipline—Genomics, aimed at determining and analyzing The nucleotide sequences of entire genomes—has driven the tendency to view a gene as an annotated genomic region with specific properties. According to the definition by the Sequence Ontology Consortium, a gene is a defined region of a genomic sequence that constitutes a unit of heredity and includes regulatory regions as well as transcribed regions. The phrase "unit of heredity" refers to the fact that a gene encodes certain functional product(s) (Proteins or non-coding RNA molecules). The "transcribed region" refers to a specific group of exons joined by introns that is transcribed as a single entity. According to sequence annotation rules adopted by modern genomic Databases, primary transcripts undergoing alternative splicing are considered to belong to the same gene, even if the final proteins are different. That is, a gene is a group of co-transcribed exons, or a genomic region that yields a set of final transcripts containing at least one common exon. Finally, an important aspect of this gene definition is that regulatory regions controlling its activity are conventionally included as part of this elementary unit of heredity in an organism.

With the caveat that in prokaryotic systems (in the case of operons) regulatory regions can control a group of genes, this interpretation of the gene remained generally accepted until very recently.

To thoroughly analyze the information encoded in the genome and realized through transcription, the U.S. National Human Genome Research Institute launched the international ENCODE (Encyclopedia of DNA Elements) project four years ago, one of the main tasks of which is the comprehensive Analysis of the human transcriptome. The First stage of this work has recently been completed, analyzing the functioning of 1% of The Human Genome (approximately 30 million Base Pairs), and the initial findings brought A number of surprises (The ENCODE Project Consortium // Nature, 2007, Vol. 447, P. 799-816).

First, the relative amount of transcribed DNA turned out to be unexpectedly high—at the level of 80%. Considering the proportion of the genome represented by exons together with introns (see Fig. 1.10), the question arises: do all these transcripts correspond to genes? Presumably, some of these primary transcripts are merely the result of nonspecific, chaotic activity of RNA polymerases. However, a significant fraction of transcripts contain sequence elements that are conserved (among mammals and human populations)—about 60% of such elements are located outside previously known protein-coding genes or regulatory regions. It is likely that in many cases these transcripts represent previously unknown non-coding RNAs, the Functional Significance of most of which remains to be elucidated. In addition, it was found that regulatory regions (promoters, enhancers) are frequently transcribed. Such transcription is presumably just one of the ways to maintain a regulatory region in an accessible, decondensed Chromatin fiber state.

For known protein-coding genes (399 in the studied region of the genome), the presence of A large number of previously unknown transcription start sites was demonstrated. Approximately half of the genes have alternative start sites located up to 100 kb away from the previously annotated transcription starts of these genes. Some of these start sites utilize promoters of other genes: a single such site can be shared by two or three genes, and the primary transcript sometimes encompasses multiple gene loci—groups of exons (see Fig. 2.17).

Furthermore, for most protein-coding genes, the analysis of their transcripts reveals the presence of previously unknown exons. Some of these exons are located up to several thousand base pairs away from all other exons of the gene, sometimes falling within the intron of another gene. As Fig. 2.17 demonstrates, it is sometimes difficult to assign a given exon to one specific gene or another. The number of different mRNA isoforms arising from both alternative splicing and trans-splicing turned out to be greater than expected.

The results of the ENCODE project indicate a dispersed pattern of regulatory element distribution throughout the genome: many regulatory elements are located within exons and introns, and they may serve as elements of the regulatory system for an entirely different gene.

Consequently, the concept of the gene once again requires some revision. According to one recently proposed definition, a gene is a union of genomic sequences encoding a coherent set of functional products that may partially overlap (Gerstein et al. // Genome Res., 2007, Vol. 17, P. 669-681). The coherence of the product set implies that in the case of protein-coding genes, each exon is shared by at least two products of the given set.

The main emphasis in this definition is placed on the final products of gene activity; overlaps between intermediate transcripts are ignored. If we base our approach on overlaps between primary transcripts (interpreting a gene as a cluster of exons that can be co-transcribed), then, for example, only two genes should be defined in the genomic region of Fig. 2.17 (by combining genes 1, 2, and 4 into one). If, however, we base it on overlaps between final products, then at least six genes are present in this region (genes 1-2 and 1-4 form separate groups of exons that partially overlap with genes 1, 2, and 4). That is, a gene does not necessarily consist of adjacent exons; a group of exons belonging to a single gene can be dispersed across a genomic region, and individual exons of the group may simultaneously belong to other genes.

Obviously, in the simple case where there is no alternative splicing (or no introns at all), the definition reduces to the classical one: a gene is a DNA region encoding a protein or RNA molecule.

Moreover, according to the definition under Discussion, regulatory elements are not Components of the gene (it is proposed to call them "gene-associated elements"): the regulatory system is more complex than a simple one-to-one correspondence between regulatory elements and coding regions.

This definition, of course, cannot be considered final. Clearly, the concept of the gene is too complex to be formulated precisely: different definitions focusing on different properties of various types of genes are, in accordance with the complementarity principle, simultaneously valid.



Last update: 07/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.