Fundamentals of Bioinformatics - Ourtsov A.N. 2013
Information Principles in Biotechnology
Genome Function and Organization
Prokaryotic Genomes
Most Prokaryotic Cells store their genetic material as a single large circular molecule with a characteristic length of 5 million Base Pairs. In addition, they may harbor extrachromosomal DNA elements known as Plasmids.
Protein-coding regions in bacterial genomes lack introns. In many prokaryotic genomes, protein-coding regions are organized into operons (tandemly arranged genes transcribed as a single mRNA molecule) and are placed under common transcriptional control (see [7], sec. 2.3). Many operons in bacterial genomes encode functionally related genes. For example, consecutive genes in the Tryptophan trp Operon of E. coli encode Enzymes that catalyze sequential steps in tryptophan Biosynthesis. In archaea, the functional linkage of genes within operons is less frequently observed.
The Escherichia coli genome. Let us examine the E. coli genome as a paradigm for a typical prokaryotic genome.
E. coli strain K-12 has long served as a model Organism in molecular biology. The genome of strain MG1655, published in 1997 by F. Blattner's research group at the University of Wisconsin, comprises 4,639,221 base pairs within a circular DNA molecule.
Compared to Eukaryotic Genomes, the E. coli genome contains relatively little non-coding DNA distributed throughout the genome. In E. coli, only 11% of the DNA is non-coding. Approximately 89% of the sequence encodes Proteins or structural RNAs. Genome annotation has revealed the presence of:
✵ 4,285 protein-coding genes;
✵ 122 structural RNA genes;
✵ non-coding repetitive sequences;
✵ regulatory elements;
✵ transcriptional and translational control elements;
✵ transposases (enzymes that bind single-stranded DNA and integrate it into the genomic DNA);
✵ prophage remnants;
✵ insertion sequence elements;
✵ insertions of atypical fragments, presumably foreign elements acquired via Horizontal Gene Transfer (see sec. 13.3).
Genome sequence analysis involved the identification and annotation of protein-coding genes and other functional regions. Because E. coli was studied as a model organism long before genome-wide sequencing projects began, many E. coli proteins were already known before sequencing was completed. As many as 1,853 proteins had been described prior to the publication of the genomic sequence.
By analogy with homologues found in sequence Databases, it became possible to predict the Functions of other genes as well. The narrower the functional Specificity of these homologues, the more precisely their roles could be defined. Currently, basic functions can be assigned to approximately two-thirds of all protein-coding genes.
Other genomic regions, such as regulatory sites or Mobile Genetic Elements, were likewise identified based on sequence similarity to known homologues from other organisms.
The distribution of protein-coding genes across the E. coli genome does not follow any rigid rules, neither regarding their chromosomal position nor their orientation. Comparisons among different strains demonstrate that gene locations are not fixed.
The E. coli genome represents a relatively densely packed arrangement of genes. Genes encoding proteins or structural RNAs account for approximately 89% of the sequence.
The average Open Reading Frame (ORF) size is 317 Amino Acids. Even if genes were distributed uniformly, the average intergenic distance would be 130 base pairs; the observed average distance between genes is actually 118 bp. Furthermore, intergenic distances vary considerably. There are some large intergenic regions containing regulatory signals and repetitive sequences. The longest intergenic region (1,730 base pairs) harbors non-coding repetitive sequences.
Approximately three-quarters of METABOLISM/31.html">Transcription units comprise a single gene; the remaining units contain multiple consecutive genes or operons. It has been estimated that the E. coli genome contains 630–700 operons. Operons vary in size, although only a few contain more than 5 genes. Genes within the same operon typically share related functions.
In some cases, the exact same DNA sequence encodes parts of more than one polypeptide chain. For instance, the same gene encodes both the τ and γ subunits of DNA polymerase III. Introduction/27.html">Translation of the full-length gene yields the τ subunit. The γ subunit is homologous to the N-terminal two-thirds of the τ subunit. A ribosomal frameshift at this site leads to premature termination in 50% of translation events, resulting in a 1:1 ratio of synthesized τ and γ subunits.
There are no overlapping genes in which different reading frames encode distinct expressed proteins.
In other instances, identical polypeptide chains are shared among multiple distinct enzymes. A protein that functions autonomously as lipoamide dehydrogenase also serves as a subunit of Pyruvate dehydrogenase, 2-oxoglutarate dehydrogenase, and the Glycine Cleavage system (see sec. 14.5).
With the availability of the completely sequenced genome, it has become possible to examine the entire E. coli proteome experimentally.
The largest Class of proteins consists of enzymes, and their coding sequences account for approximately 30% of all genes.
Many enzymatic functions are distributed across multiple proteins. Some of these sets of functionally similar enzymes are remarkably alike, likely having evolved through Gene Duplication either within E. coli itself or in its ancestors.
Other sets of functionally similar enzymes have very dissimilar sequences and differ significantly in specificity, regulation, or subcellular localization.
Certain Features of the E. coli enzyme repertoire provide the metabolic flexibility that enables it to grow and compete effectively under changing environmental conditions.
E. coli independently synthesizes all PROTEIN AND NUCLEIC acid monomers (Amino Acids and NUCLEOTIDES) as well as Cofactors.
E. coli exhibits notable metabolic flexibility, capable of both anaerobic and aerobic metabolism utilizing various energy storage mechanisms. It can grow on A wide variety of carbon and nitrogen sources. Not all Metabolic pathways are continuously active; rather, they are induced in response to environmental fluctuations.
A diverse array of membrane transport proteins allows the bacterium to utilize numerous types of nutrients.
Even for specific metabolic reactions, a wide variety of distinct enzymes exists, enabling the bacterium to rewire its metabolism when environmental conditions shift. Complex regulatory mechanisms are integrated into the overall Regulation of Protein expression.
Nevertheless, E. coli is not entirely metabolically autonomous; for instance, it cannot fix CO2 or N2.
The genome of Mycoplasma genitalium. As an example of a prokaryotic genome from one of The Simplest Living Organisms, let us examine the genome of Mycoplasma genitalium.
Mycoplasma genitalium is a pathogenic bacterium that causes nongonococcal urethritis. Its genome was sequenced in 1995 through a collaboration between research groups at The Johns Hopkins University and The University of North Carolina. The M. genitalium genome consists of a single DNA molecule containing 580,070 bp, making it the smallest known genome to date. Consequently, M. genitalium serves as a model organism for attempts to construct a minimal self-sustaining living Cell.
The M. genitalium genome has a high coding density. Corresponding proteins have been identified for 468 genes, with 85% of the total sequence being protein-coding. The average length of a coding region is 1,040 bp. As in other Bacteria, the coding regions contain no introns. Further genomic compaction is achieved through gene overlapping, much of which appears to have arisen from the loss of stop codons.
Some of the genes in M. genitalium encode proteins essential for independent Replication (such as those involved in DNA replication, transcription, and translation), as well as ribosomal and Transfer RNAs. Other genes are dedicated exclusively to pathogenicity, encoding adhesins involved in binding to host cells, proteins that evade the host immune system, and numerous transport proteins. As a consequence of its parasitic lifestyle, the genome lacks genes for many metabolic enzymes, including—as some researchers suggest—those required for Amino acid biosynthesis.
Last update: 11/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.