Biotechnology - Y.O. Sazykin 2006
General biotechnology
Genomics and proteomics
General characteristics
The breakthroughs in genetics, molecular biology, and biochemistry in the 1990s led to The Emergence of two new fundamental disciplines: Genomics and Proteomics. The rapid development of these fields drives modern progress across various sectors of biotechnology, including pharmaceutical manufacturing.
The term "genomics" is derived from "genome"—the complete set of genes of an Organism; whereas "proteomics" stems from "proteome"—the Complement of structural and catalytic Proteins within a eukaryotic or Introduction/4.html">Prokaryotic Cell. Both disciplines can be seen as the formalization of The current stage of development in genetics and Protein Chemistry, bringing them closer to the whole cell. Chronologically and methodologically, genomics takes precedence; proteomics is built upon genomics, representing the Study of Living systems at the protein level.
Early 19th-century genetics later became known as Formal genetics, as research was conducted at the "Gene-trait" level (marked by the discovery of G. Mendel's renowned foundational laws). Although the existence of the gene was postulated, its material nature remained unknown. It was not until the 1950s—following the emergence and rapid confirmation of the DNA double helix model and METABOLISM/2.html">THE CONCEPT OF a gene as a segment of DNA—that Molecular Genetics experienced explosive growth, establishing the size of individual genes, functional regions within genes, and so forth. Concurrently, biochemists and geneticists uncovered the template mechanism of Protein Synthesis, wherein The Genetic Code is transferred from DNA to protein.
The objective of genomics is to establish the complete genetic profile of a cell—the number and sequence of its genes, the number and sequence of NUCLEOTIDES within each gene, and the function of each gene in relation to the organism's metabolism or, more broadly, its vital activity. By knowing the number of genes and their nucleotide sequences, genomics reveals The Essence of an organism: its potential capabilities, its specific (and even individual) differences from other organisms, and its expected response to environmental stimuli.
The goal of genomics is to acquire information regarding all potential properties of a cell that may not be active at a given moment, such as "silent genes." Proteomics, on the other hand, characterizes The Cell at a specific point in time by capturing all its resident proteins—essentially taking a "snapshot" of the cell's functional state at the proteome level (i.e., the ensemble of all active enzymatic and structural proteins, as opposed to unexpressed genes). Just as genomics largely advanced alongside sequencing technologies, proteomics relies fundamentally on two-dimensional gel Electrophoresis, which separates proteins by molecular weight in one dimension and by isoelectric point in the other. While this method is not entirely new, significant refinements now allow researchers to monitor hundreds of proteins simultaneously in real time. Metaphorically speaking, this is akin to shooting a "motion picture" of the cell at the molecular (protein) level. Successive frames of this molecular film capture the rise and fall of individual cellular proteins over time—for instance, during transitions between Phases of the Cell Cycle or in response to environmental changes—while also revealing post-translational protein modifications, and so on.
Proteomics also makes it possible to track Protein-Structure/156.html">Protein Interactions, such as signal Transduction from the cell surface to selective Transcription factors in The Nucleus. Consequently, this approach can transform not only the screening technology for immunosuppressants but also The Study of signal transduction inhibitors as a whole. Proteomic Methods provide a more comprehensive, multi-faceted picture of how novel potential antimicrobial agents interact with the cell. Furthermore, studying the dynamics of secondary metabolite enzyme Biosynthesis in Microorganisms through proteomics can elevate such research to a radically higher level.
Returning to the relationship between proteomics and genomics, it must be emphasized that proteomics can be considered an extension of functional genomics specifically. Unlike genomics, proteomics focuses on the products encoded by genes that are actively expressed at a given moment.
The minimal genomes of certain microbial species consist of several hundred genes, whereas The Human Genome approaches one hundred thousand. The size of individual genes typically ranges from roughly one thousand Base Pairs upward. Thus, the total number of base pairs comprising an individual genome measures at least in the hundreds of thousands, and commonly in the many millions.
Consequently, fully decoding an organism's genome requires determining The sequence of several million base pairs (A-T for adenine-thymine, G-C for guanine-cytosine). Performing whole-genome sequencing, as the accepted term goes, is only feasible using advanced high-throughput technologies and specialized equipment.
Today, the combined daily output of numerous laboratories worldwide amounts to sequencing roughly one million base pairs. Storing and utilizing these vast datasets is impossible without specialized Databases, several of which hold international status. Prominent Examples include the databases of The Institute for Genomic Research (USA) and the University of Heidelberg (Germany). These international databases provide crucial insights: whether a gene exists and how widespread it is among pathogens; what product the gene encodes and how that product (typically an enzyme) participates in a specific metabolic cycle; and what exact reaction it catalyzes within that cycle. In other words, the primary test object for screening antimicrobial substances and selective metabolic inhibitors is no longer merely the microbial culture, but the gene itself (or more precisely, its encoded product).
It is important to keep in mind that variations in nucleotide sequences across different genomes do not necessarily indicate interspecies differences; for instance, distinct strains of the same microbial species used as industrial producers exhibit genomic variations. Intraspecies genomic differences can be found across the entire spectrum of living organisms, including humans (in the latter case, individual DNA profile variations provide a powerful new tool for forensic examination).
Like any recently established scientific discipline, genomics is branching into several sub-fields, with databases specializing accordingly. Foremost among these is structural genomics, which aims to identify genes using specialized computer software (searching for open reading frames flanked by start and stop codons). Consequently, The Genome under study is characterized by molecular weight, gene count, and The nucleotide sequence of each gene—within the chromosome for prokaryotes, and within each chromosome for eukaryotes.
Comparative genomics makes it possible—by querying a database and receiving a rapid response—to determine whether a sequenced gene is unique or has already been identified in another laboratory, to assess the degree of Homology between related genes (i.e., Sequence homology within the Open Reading Frame), and to answer questions regarding the evolutionary proximity of organisms and other fundamental biological problems. Moreover, comparative genomics holds great practical potential. For example, when searching for inhibitors of a specific pathogen gene to develop new drugs, it is vital to know whether a gene with an identical or highly similar nucleotide sequence is present in the host organism, as this allows researchers to predict the safety profile of the prospective medication.
Following structural and comparative genomics is functional or metabolic genomics, a fully established scientific discipline in its own right. Its goal is to establish The Link Between the genome and metabolism, between gene clusters and multi-step metabolic pathways, and between individual genes and specific metabolic reactions. A key concept in functional genomics is The Use of "model organisms"—primarily specific microorganisms whose connections between genes and their encoded enzymatic and structural proteins have been thoroughly mapped, such as prokaryotes and lower eukaryotes with fully sequenced genomes and exhaustively studied metabolism. Prime examples of such model microorganisms include Escherichia coli (among prokaryotes) and Saccharomyces cerevisiae (among eukaryotes). Comparing a gene from an organism under study with a homologous gene in a model organism allows researchers to infer the gene's function. Conversely, a lack of homology indicates The Need for targeted investigation into the novel gene's role.
In the context of pharmacology, functional genomics is particularly crucial for determining the so-called "essentiality" of individual genes. "Essentiality" refers to whether a gene is indispensable for cellular viability. When developing antimicrobial drugs, it is precisely these "essential" genes that must serve as targets for antimicrobial agents. Notably, a gene may sometimes acquire essential status only under specific environmental conditions that a pathogenic microorganism encounters.
Given genome sizes and gene counts, it is clear that complete genome sequencing is achieved much faster in microorganisms than in higher eukaryotes. To date, the genomes of several dozen bacterial species, including pathogens, have been fully sequenced. While Genome Size varies among bacterial species, it generally spans a few thousand genes or several million base pairs. For instance, E. coli possesses a little over four thousand genes and, accordingly, more than four million base pairs.
Currently, clinical practice employs roughly two hundred natural and synthetic antibacterial agents. Each has its specific target, typically an enzyme or a ribosomal protein, with the total number of validated targets also hovering around two hundred. Consequently, the vast majority of genes remain unutilized as targets for antibacterial agents. To prove gene "essentiality," researchers employ a selective gene "knockout" technique to test whether the organism survives the Procedure; this approach holds immense promise as a screening technology for antibacterial (or more broadly, antimicrobial) agents.
Traditionally, the primary screening of such agents involves testing their impact on the growth of a microbial test culture. Highly active growth-inhibiting substances (natural or synthetic) selected at this stage undergo further testing—specifically determining their antimicrobial activity spectrum, their in vivo efficacy in laboratory animals, and their toxicity profiles for both the macroorganism as a whole and its individual Organs and Tissues.
Upon successfully completing preclinical trials, the prospect of advancing the drug to clinical trials is considered. This is typically followed by an in-depth investigation into the MECHANISM OF ACTION of the antimicrobial agent at the subcellular and molecular levels—meaning researchers search for its intracellular target, which, following modern terminology, is a macromolecule or macromolecular complex (referred to as a target). Next, they identify the gene encoding this macromolecule, or the genes encoding the macromolecules that comprise the macromolecular complex.
In contrast to Traditional Methods, the novel screening technology leverages information regarding the fully sequenced pathogen genome and the presence of "essential" genes within it. Laboratories dedicated to discovering new antimicrobial drugs pre-select a candidate gene to serve as the target for their assays (or more precisely, the product encoded by that gene acts as the target).
Targeted screening allows researchers to select biologically active compounds with a predetermined mechanism of action based on the chosen gene (unlike the traditional "cell-to-gene" search approach). The first phase of targeted screening begins by isolating this gene (the corresponding DNA fragment) from the genome. Next, using the Polymerase Chain Reaction (PCR), the DNA fragment is amplified (by creating a viral vector that is introduced into a plasmid), multiplying the number of gene copies. Following this, researchers construct:
✵ a cell-free system, in which the generated template is used to synthesize Messenger RNA specific to the gene;
✵ a cell-free ribosomal system, in which this Messenger RNA is translated to produce the protein encoded by the gene.
The cell-free system used for primary screening and evaluating The activity of potential drug candidates—specifically Enzyme Inhibitors (the products of the studied gene)—contains both the enzyme and its substrate. Once the protein (the product of the gene under investigation) has been synthesized, a key question arises: how to determine its function? For instance, what reaction does it catalyze as an enzyme, so that inhibitors can be selected based on the suppression of that reaction? If the protein shows close homology to a protein from a "model" organism, finding a suitable cell-free system (the substrate for the new protein) is straightforward. If the homology is present but not very close, researchers resort to motif analysis—examining short Amino Acid Sequence segments distributed throughout the protein chain that may be conserved between the two proteins.
When such homology is absent, or when homology is detected but the function of the homolog (the reference protein) remains unclear, scientists employ yet another method: they determine which proteins are co-transcribed with it (transcribed into a polycistronic messenger RNA sequence). If the transcript forms part of a polycistronic (gene-equivalent) messenger RNA necessary for a subsequent multi-step metabolic process involving several Enzymes, the search space for the function of the protein embedded within this enzyme group is significantly narrowed.
In general, approaches to determining the Functions of studied gene products are numerous and steadily growing in variety. Ultimately, this makes it possible to conduct high-throughput screening of potential function inhibitors for almost any of the thousands of genes that make up a pathogen's genome, uncovering all possible vulnerabilities of the microbial cell. This goal-oriented approach has given rise to the term "reverse genetics" in literature, denoting research that proceeds not from the cell and its phenotype to the gene, but rather, conversely, from the gene to the cell and its phenotype.
Complete genome sequencing, combined with Genetic Engineering techniques, contributes to pharmacology in yet another way. Pathogenic microorganisms possess genes that are "essential" for the infectious process, but "non-essential" during growth in vitro on artificial nutrient media. In the latter case, they evade the researcher's attention, cannot be identified, and thus cannot serve as targets in drug discovery. These hidden, or metaphorically speaking, in vitro "silent" genes of pathogenic microorganisms are termed ivi genes (in vivo-induced genes), even though they include not only genes encoding toxins, adhesins, and other virulence factors. They also encompass genes for enzymes and transport proteins that enable the pathogenic microbial cell to survive and multiply within host tissues under conditions of scarcity for certain organic substances and inorganic ions.
For example, a microbial cell residing in vivo experiences an iron ion deficiency that never occurs on standard nutrient media. Under these conditions, the cell synthesizes a specialized iron transport system to uptake iron from a low-concentration environment, effectively working against the concentration gradient. The expression of specific genes is required to build such a system. From being silent and "non-essential," they become "essential"—meaning that suppressing their functions with selected inhibitors will arrest the growth (proliferation) of the pathogen precisely in vivo, within the infected organism. This is, in essence, the ultimate goal of researchers developing novel drugs.
Genes that become "essential" for the pathogen specifically in vivo include those encoding optimal pathway components, as well as those compensating for a deficiency in Purines and their precursors.
The foregoing, of course, does not imply that only ivi genes are expressed in the pathogen's cell during an infection. Most genes are expressed both in vivo and in vitro, as their products are perpetually required by the cell. Such genes have been figuratively dubbed "housekeeping genes." They are expressed under any conditions, as the cell simply cannot exist without them.
The ratio between housekeeping and ivi genes varies among different pathogenic Bacteria, but on average, over 90% of genes belong to the former group. Because inhibitors of housekeeping genes are identified through in vitro screening on nutrient media, virtually all clinically used Antibiotics and synthetic antibacterial agents function by inhibiting these very genes.
Of considerable interest are the methods for identifying and isolating ivi genes for subsequent use in cell-free inhibitor screening systems. A prime example is the IVET (In Vivo Expression Technology) method.
The genome of a pathogenic bacterium (in this case, the strain Salmonella typhimurium) is cleaved into hundreds of fragments using a broad set of restriction enzymes. Using genetic engineering techniques, each individual fragment is joined to a promoterless chloramphenicol acetyltransferase gene (where the promoter is the DNA region recognized by RNA polymerase). Such a promoterless gene cannot replicate upon introduction into a cell. However, it could replicate if the attached DNA fragment (in this case, a Salmonella DNA segment) provided a functional promoter. In that event, this promoter would drive the Replication not only of its own DNA, but also of the adjacent promoterless gene. Thus, replication of the chloramphenicol acetyltransferase gene can only occur in the cell by co-opting or "hijacking" a foreign promoter.
In the next step, a promoterless lactose Operon ($lac\text{ Z}$), required for lactose metabolism, is also attached to this chimeric fragment (designated as $x$-cat, where $x$ is the Salmonella genomic fragment and cat is the chloramphenicol acetyltransferase gene). This tripartite construct ($x$-cat-$lac\text{ Z}$) is then inserted into a plasmid. In reality, we are dealing not with a single fragment, but with a pool of them, because the $x$ segment derived from the Salmonella genome depends on the restriction enzyme used and therefore contains various genomic regions. The $x$-cat-$lac\text{ Z}$ fragments differ precisely in their $x$ portion. This yields a library of diverse Plasmids, and following their introduction into $E. coli$, a library of various $E. coli$ strains carrying different segments of the Salmonella genome.
The subsequent step involves introducing each $E. coli$ strain into a laboratory animal (mouse) and administering chloramphenicol. Twenty-four hours later, the bacterial culture is recovered from animal tissues and plated onto solid lactose indicator media. The resulting colonies are visually analyzed. They appear either red (acidifying the pH and fermenting lactose) or white (colorless), with red colonies predominating at over 90%. However, it is the white colonies that are selected and subjected to further study. The underlying rationale is as follows: if a viable cell is recovered from a chloramphenicol-treated animal and forms a colony on solid media, it means that the chloramphenicol acetyltransferase gene was expressed in that cell, producing the enzyme that inactivates (acetylates) the antibiotic. Consequently, the given fragment $x$ contains a promoter-active gene. This gene was expressed within the animal's body, which in turn drove the expression of the downstream chloramphenicol acetyltransferase gene (lacking its own promoter). However, the expressed gene could belong to either the ivi or the housekeeping Class.
Red colonies indicate the expression of a gene encoding an enzyme that breaks down lactose; this alters the indicator's color, staining the colony. We can thus conclude that the fragment $x$ contains a promoter along with a housekeeping gene—meaning it is expressed constantly, both within the animal and on artificial growth media. Such genes are of no interest in this context and can be identified via more traditional approaches (for subsequent inhibitor screening against their encoded products).
If the colony on the lactose indicator medium grows colorless, it signifies that the promoter was inactive on artificial growth media, and the gene within fragment $x$ was not expressed. Presumably, it is required exclusively during the infectious process and belongs to the virulence-associated ivi genes that remain silent in vitro. The IVET method is not the only approach for identifying ivi genes in pathogens; other techniques, such as those utilizing Site-Directed Mutagenesis, also exist.
Interest in ivi genes stems from the fact that they (and their products) have virtually untapped potential as drug targets, offering the prospect of high selectivity and safety in novel therapeutics. Furthermore, genomics enables the Classification of pathogen genes across multiple parameters, which in turn facilitates an increasingly targeted screening of antimicrobial agents.
As an example, consider the results obtained from studying members of the genus Chlamydia (intracellular parasites with a relatively small genome of about one million base pairs) responsible for bronchopulmonary and urogenital infections. First, The genes of this prokaryote were separated into housekeeping and ivi genes. Next, a group of genes affecting host cell apoptosis (programmed cell death affecting a portion of a multicellular organism's cell population—a fundamental biological phenomenon responsible for maintaining an optimal and sufficient cell count) was identified. Finally, genes duplicating the parasite's life-support system were characterized. Notably, several of these genes showed high homology to higher plant genes, with approximately 27% being unique to the genus Chlamydia.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.