Introduction to Molecular Biology: From Cells to Atoms - Anthony Rees, Michael Sternberg 2002

Nucleic Acids and Genes
Gene Organization

Class="center">Introduction/introduction.files/image074.jpg" width="642"/>

Fig. 27.1. Organization of genetic material.

A structural Gene is the smallest segment of DNA or RНК encoding the complete Amino Acid Sequence of a particular protein. A Cell of a higher Organism may contain up to 100,000 genes (according to recent data, significantly fewer). However, The amount of DNA it contains would be sufficient to form 10 times as many genes. It remains unclear why a cell needs "extra" DNA, although recent findings on eukaryotic DNA Structure suggest some hypotheses in this regard (see below). Viruses may contain as few as 5–6 genes, whereas the prokaryotic genome accounts for approximately 0.1% of that of higher organisms.

THE PROKARYOTIC CHROMOSOME contains approximately 2,000–3,000 non-overlapping genes distributed along the DNA. Structural genes are divided into three main types: independent genes, METABOLISM/31.html">Transcription units (transcriptons), and operons. In addition, Cells may harbor smaller, autonomously replicating units called Plasmids.

Independent genes are so named because their transcription occurs without the involvement of any regulatory mechanisms governing transcriptional activity, unlike the other two classes of genes. Such genes are said to exhibit a constitutive form of expression—that is, expression without Regulation at the transcriptional level. Every structural gene represents a continuous sequence of codons following one another directly, and the mRNA transcribed from such a gene is always monocistronic (a Cistron is defined as a nucleotide sequence encoding a complete single polypeptide chain).

Spacer DNA is located between genes and is not always transcribed. Sometimes the region of such DNA between adjacent genes (the so-called spacer) contains some information related to the regulation and initiation of transcription, but it may also consist merely of short repeated sequences of redundant DNA whose role remains elusive.

Transcription units (transcriptons) are groups of consecutive genes transcribed together. Typically, these are genes for Proteins or Nucleic Acids that are functionally related. For example, in E. coli, transcription units have been identified that include genes for various rRNAs (each gene in a single copy) and a set of tRNA genes. Fig. 27.1 illustrates one such unit, which comprises two tRNA genes and three genes corresponding to three different rRNAs. Transcription units containing three or even four tRNA genes are quite common, and their arrangement relative to the rRNA genes shows no apparent pattern. For this class of genes, the mRNA molecule represents a transcript of an entire group of genes; hence, such mRNA is termed polycistronic.

Operons are groups of contiguous structural genes controlled by a specific DNA region called the operator. A classic example is the lac Operon (Chap. 28), which consists of three structural genes (Z, Y, and A) and a regulatory DNA region comprising two sequences—the promoter and the operator. Furthermore, another gene, the regulator gene (I), encodes a protein known as a repressor, which regulates the transcriptional activity of the lac operon. Several metabolites essential for cell viability are known whose Biosynthesis or metabolism is controlled by Enzymes encoded by genes organized into an operon (Chap. 28). Operons typically lack spacers.

Plasmids are small circular DNA molecules of varying length. Large plasmids may contain up to 100 genes. Such plasmids frequently (though not always) carry Genetic information that enables them to transfer from one cell to another during a process known as conjugation. Small plasmids may contain only about 10 genes and are incapable of intercellular transfer via conjugation. The number of genes in a plasmid is variable. EXCHANGE OF GENETIC information with The Cell genome or other plasmids occurs through The transfer of specific segments of plasmid DNA capable of movement (transposition) from one DNA molecule to another, which are therefore called Transposons.

Transposons—DNA segments capable of translocating from one molecule to another—frequently carry Antibiotic Resistance genes. Genes residing within a transposon can transfer from plasmids to chromosomal DNA and vice versa. Thus, through plasmid transfer during conjugation, resistance genes can spread rapidly within a bacterial population.

VIRUSES INFECTING PROKARYOTES utilize both DNA and RНК as genetic material. RНК in these viruses is invariably single-stranded (Chap. 5). DNA can be either double-stranded (as in Bacteriophages T2, T4, T5, and T6) or single-stranded (as in bacteriophage φX174). One of the key differences in DNA organization between certain prokaryotic viruses, on the one hand, and Prokaryotic Cells, on the other, is that these viruses, unlike cells, possess overlapping genes.

Gene overlapping occurs when the same nucleotide sequence encodes two or three different proteins. Such genes were first discovered in coliphage (i.e., a bacteriophage infecting E. coli) φX174 by Sanger and coworkers in 1977. By determining The nucleotide sequence of the phage DNA, they found that three genes (designated K, C, and A) occupy the same position within the DNA molecule, but their respective nucleotide sequences are read in different reading frames. Although this utilization of DNA achieves considerable economy of genetic material, it severely restricts sequence Variability, particularly in the regions of initiation and termination codons for different proteins. For instance, the final Base of the first codon shown in the diagram for gene A must be an adenine, which serves as the start codon for gene C, encoding fMet. Similarly, the first base of the third codon in gene C must also be an adenine, as it marks the end of the termination codon for gene A.

Eukaryotic cells utilize exclusively double-stranded DNA as genetic material. Structural genes, whose functioning is closely linked to specific DNA sequences known as regulatory regions, are subdivided into independent genes, repeated genes, and gene clusters. The coding sequences of these genes may be interrupted by non-coding sequences called introns. Additionally, intergenic regions may contain highly repetitive DNA (satellite DNA) as well as spacer DNA, which may be either transcribed or untranscribed.

Independent genes are genes whose transcription, as in prokaryotes, is not linked to that of other genes within a transcription unit. Their activity may, however, be regulated by exogenous substances such as Hormones.

Repeated genes are present in the chromosome as multiple copies of a single gene. The 5S rRNA gene is repeated many hundreds of times, with the repeats arranged in tandem—that is, directly adjacent to one another without intervening spaces. Functionally related 5.8S, 18S, and 28S rRNA genes are also present as numerous repeats, but they are localized within the nucleolar DNA.

Gene clusters are groups of distinct genes with related Functions localized in specific chromosomal regions (loci). Clusters are also frequently present in the chromosome as repeats. For example, the histone gene cluster is repeated 10–20 times in The Human Genome, forming a tandem repeat array.

Introns are DNA segments that interrupt the expressed, i.e., coding, portion of a gene into segments called exons. The phenomenon of split genes was first discovered during studies of adenovirus and confirmed in 1977 during investigations of the mouse globin gene and ribosomal GENES OF THE fruit fly Drosophila melanogaster. A single gene may contain numerous introns; for instance, the chicken Ovalbumin gene contains 8 introns, the total length of which exceeds the sum of all coding sequences within that gene. During transcription, RНК polymerase copies the entire gene. Subsequently, specialized splicing enzymes process the transcript—they excise the introns and ligate the exons together, yielding a mature yet unmodified mRNA. For such Processing to occur, specific nucleotide sequences must be present at the intron-exon boundaries in the DNA. Such a sequence is illustrated in Fig. 27.2; it occurs frequently in The Genome and can serve as a recognition site for splicing enzymes following transcription. Normal maturation and Translation of the mRNA then proceed (Chap. 22).

Fig. 27.2.

Satellite DNA consists of characteristic nucleotide sequences (ranging from 10 to 200 NUCLEOTIDES in length) that are arranged in tandem and repeated hundreds of times. The function of this DNA remains to be elucidated.

VIRUSES INFECTING EUKARYOTES employ diverse forms of gene organization; some viruses are known to possess both overlapping and split genes.

Overlapping genes have been found, for example, in the mammalian virus SV40, whose DNA contains a region (from nucleotide residue 1488 to 1601, measured from THE ORIGIN OF Replication) encoding two proteins: VP1 and VP2. Thus, the expansion of genetic capacity through The Use of multiple reading frames for coding occurs in both prokaryotic and Eukaryotic Viruses, whereas nothing comparable has yet been discovered in cellular genomes. The SV40 virus also contains split genes. One of the late transcription genes encoding the TL antigen (commonly referred to as "large T") is split into two exons. The length of the first exon (T1) is 246 Base Pairs, and that of the second (T2) is 1,879 base pairs. The single intervening intron is 345 base pairs long and is excised from the TL gene transcript during splicing, after which the two coding sequences are joined together.



Last update: 13/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.