Fundamentals of Bioinformatics - Ogurtsov A.N. 2013
Information Principles in Biotechnology
Sequencing of Biological Sequences and Gene Expression
Gene Expression
Gene expression is the process by which hereditary information from a gene (a DNA nucleotide sequence) is converted into a functional product, such as RNA or protein. In effect, during expression, a gene acts as a blueprint for the synthesis of a specific protein.
Gene expression profiles provide key insights into uncovering a gene's biological role. All cellular, tissue, and organ Functions are governed by differential gene expression.
Gene expression analysis is performed to study gene function. Information on which genes are expressed in healthy versus diseased Tissues helps identify both the Complement of Proteins characteristic of normal function and the protein aberrations associated with diseases.
These data are subsequently used to develop novel diagnostic tests for various diseases, as well as new drugs capable of modulating The activity of affected genes or proteins.
Historically, gene expression was studied at the RNA or protein level on a gene-by-gene basis using Northern and Western blot analyses (see [7], sec. 11.4). Today, global expression analysis Methods are available, allowing all genes to be investigated simultaneously.
A straightforward yet relatively costly RNA-level analysis method involves direct sequencing sampling from RNA pools, cDNA libraries, or even sequence Databases.
A more advanced technique known as SAGE (Serial ANALYSIS OF GENE Expression) involves synthesizing very short sequence tags (typically 8–15 NUCLEOTIDES) from each cDNA, which are then concatenated into clusters of several hundred to form tag concatemers for sequencing. A single sequencing reaction can yield information on the relative Abundance of hundreds of different mRNAs. Each SAGE tag uniquely identifies a gene, and by counting the tag frequencies, the relative expression levels of each gene can be determined.
The highest-throughput approach for gene expression analysis is DNA microarray technology.
A DNA microarray is an array of DNA strands (often referred to as features or Cells) arranged on a miniature substrate, such as a nylon filter or Microscope slide (see [7], sec. 13.5).
Each element in such an array consists of a multitude of identical single-stranded DNA molecules representing a specific gene.
The phenomenon of DNA Hybridization enables DNA probes to isolate specific DNA molecules from highly complex mixtures, such as an entire genomic DNA library or cellular RNA. A DNA probe is an individual DNA molecule (or RNA molecule in the case of an RNA probe) covalently linked to a radioactive or fluorescent label.
Microarrays are typically hybridized with a complex RNA probe; such a probe is produced by labeling the pooled mixture of RNA molecules extracted from a specific Cell type. Thus, the COMPOSITION OF THE probe reflects the relative abundance of individual RNA molecules in the source cell.
If non-saturating hybridization is achieved, the signal intensity of each microarray element represents the relative abundance of the corresponding RNA in the probe, thereby allowing the relative expression levels of thousands of genes to be visualized simultaneously.
The most widely used method involves the automated deposition of individual DNA clones onto a substrate (a specially coated microscope slide), for instance, using ink-jet printing (IJP). Such IJP DNA arrays can achieve a density of up to 5,000 features per square centimeter. The features contain DNA molecules (clones from The Genome under study or cDNA molecules) up to 400 bp in length, which are denatured prior to hybridization (Figure 71).
First, DNA clones are amplified and spotted onto the substrate to produce the microarray (Figure 71(a)). Concurrently, reference and sample RNA fragments are reverse-transcribed into DNA fragments and labeled with different fluorescent Dyes that emit in distinct spectral regions (e.g., red and green). These fluorescent probes are then hybridized to the DNA on the microarray (Figure 71(a)). Afterward, the microarray is washed with Water to remove unhybridized probes, laser excitation is used to measure the fluorescence intensity of each dye across all microarray elements (genes), and these data are converted into relative gene expression levels in the test sample compared to the reference.
Class="center">
Figure 71 - Schematic diagram of differential expression measurement using a DNA microarray
Another approach utilizes DNA chips, in which short oligonucleotides are synthesized *in situ* photolithographically during chip manufacturing. Such feature arrays are known as gene chips. They achieve densities of up to 1,000,000 elements per square centimeter, with each element comprising up to 109 single-stranded oligonucleotides 25 nucleotides in length.
Each gene on a gene chip is represented by twenty elements (twenty overlapping oligonucleotides); additionally, twenty mismatch control elements are included to normalize for nonspecific hybridization.
Fluorescent RNA probes are widely used for screening DNA Microarrays because different fluorophores facilitate the convenient labeling of multiple RNA sets.
Fluorescently labeled RNA probes can be simultaneously hybridized onto a single microarray, enabling direct measurement of differential gene expression.
Gene chip hybridization is performed using individual probes on two identical chips, after which signal intensities are measured and compared using a computer.
Raw microarray data consist of images of hybridized arrays. The exact Nature of the image depends on the microarray substrate (the type of matrix used). DNA microarrays can contain many thousands of elements; therefore, data collection and analysis processes must be automated.
Image preprocessing software is typically supplied with the scanner. It allows the boundaries of individual spots to be defined and total signal intensity (power) to be measured based on the brightness of whole spots. Signal intensities must be Background-corrected, and appropriate standards must be included in the array to measure nonspecific hybridization and assess variations in hybridization parameters across different microarrays.
The objective of data Processing is to convert hybridization signals into numerical values that can be used to generate a gene expression matrix. The interpretation of microarray hybridization data is performed by grouping them into clusters according to similar expression profiles.
Grouping is a method of simplifying large datasets by clustering similar data together. To automate microarray Data analysis methods, various software Applications have been developed, for example:
✵ Stanford University Resources - http://Genomics.stanford.edu/
- Stanford Microarray Database (SMD) - http://smd.stanford.edu/
- Microarray Resources: Software and Tools - http://smd.stanford.edu/resources/restech.shtml;
✵ TM4 - http://www.tm4.org/ - Microarray Data Manager (MADAM), TIGR Spotfinder, Microarray Data Analysis System (MIDAS), and Multiexperiment Viewer (MeV), as well as a Minimal Information About a Microarray Experiment (MIAME)-compliant MySQL database;
✵ GeneMaths XT - http://www.applied-maths.com/genemaths-xt;
✵ Eisen Lab - http://rana.lbl.gov/EisenSoftware.htm;
✵ Microarray Image Analysis Software - http://www.statsci.org/micrarra/image.html;
✵ The Gene Ontology - http://www.geneontology.org/GO.tools.microarray.shtml;
✵ BioDiscovery - http://www.biodiscovery.com/ - Nexus Expression - http://www.biodiscovery.com/index/nexus-expression;
✵ BxArrays - http://bioinforx.com/lims/microarray-gene-expression-data-analysis/bxarrays;
✵ Array-Pro Analyzer - http://www.mediacy.com/index.aspx?page=ArrayPro;
✵ Premier Biosoft - http://www.premierbiosoft.com/dnamicroarray/index.html;
✵ J. Craig Venter Institute - http://www.jcvi.org/cms/research/software/;
✵ Bioconductor - http://www.bioconductor.org/.
DNA microarrays are used for the following purposes.
1. Investigation of cellular states and underlying processes. Analyzing differential expression depending on The Cell state helps deciphering the mechanisms of processes such as sporulation or the transition from aerobic to Anaerobic METABOLISM.
2. Disease Diagnosis. A microarray test for Mutations can confirm the diagnosis of a suspected genetic disorder. It enables the detection of late-onset symptoms, such as in Huntington's disease (see section 6.7), and the identification of potentially risky genes in prospective parents (family planning counseling).
3. Genetic early-warning markers. Some diseases are not determined solely by genotype, but their probability of development depends on The behavior of specific genes and can be assessed from their expression profile. Being aware of a predisposition to a particular disease allows individuals in some cases to prevent its onset by modifying their lifestyle.
4. Selection of pharmacological treatments. Identifying the genetic factors that govern the body's response to medications; in some patients, such effects render Treatment ineffective, while in others they even provoke severe allergic reactions.
5. Disease Classification. Distinct gene expression profiles can define, for example, various types of leukemia. Knowing the exact disease type is crucial for selecting optimal therapeutic strategies.
6. Target identification for drug discovery. Proteins showing elevated Transcription levels in specific disease states could serve as potential pharmacological targets (provided that other evidence demonstrates that Upregulated transcription is necessary for maintaining or promoting the pathological state).
7. Pathogen resistance. Comparative analysis of genotypes or expression profiles in antibiotic-susceptible and resistant bacterial strains helps identify proteins involved in the resistance mechanism.
Gene discovery. Recently, substantial funding has been allocated to searching for genes associated with specific diseases. The goal of this research is to develop novel therapies to combat a wide range of common Functional and Structural disorders, such as Cancer, tuberculosis, asthma, etc.
Currently, There are two main strategies for discovering proteins that can serve as molecular targets suitable both for drug development and Gene Therapy.
One approach to detecting disease-related genes is positional cloning. This method involves studying a human population exhibiting cases of the disease in question and identifying the chromosome associated with its development.
Next, the disease is linked to a specific chromosomal region, after which a large segment of the chromosome near this region (locus) is sequenced, yielding a DNA sequence several hundred thousand Base Pairs long. In principle, such a locus may contain numerous genes, although most likely only one of them is actually involved (directly or indirectly) in the disease process.
To improve the efficiency of gene recognition within the locus, various sequence-searching and gene-prediction methods can be employed; however, several genes must invariably be expressed, and further analysis (or testing) will be required to determine precisely which gene is implicated in the disease.
Although genes discovered through this method may be entirely satisfactory from an academic standpoint, they are not necessarily good drug targets (or points of therapeutic intervention).
Another approach to gene discovery—which requires significantly less sequencing effort and relies more heavily on the powerful search capabilities of modern computational systems—is based on identifying genes that are actually expressed in healthy and diseased tissues. It allows for the comparison of expression levels between the two states and, based on this analysis, the most effective selection of a potential target. This process analyzes the mRNA molecules used for the synthesis of these proteins.
Typically, gene detection involves the following elements:
✵ splice sites;
✵ start and stop codons;
✵ branch points;
✵ promoters and transcription terminators;
✵ polyadenylation sites;
✵ ribosome binding sites;
✵ topoisomerase II binding sites;
✵ topoisomerase I Cleavage sites;
✵ binding sites for various transcription factors.
Such localized regions are termed "signals" and are detected using "signal sensors." Conversely, extended sequences and sequences of variable length (such as exons and introns) are termed "content" and are detected via "content sensors."
Neural networks are among the most sophisticated signal sensors currently in use.
A typical content sensor is one that predicts coding regions.
To determine complete gene Structure, several systems have been developed that combine signal and content sensors. Such systems are capable of recognizing more complex interdependencies among gene properties. One of the earliest comprehensive gene-finding programs developed to date is Genelang (http://arete.ibb.waw.pl/PL/html/gene_lang.html). Built on dynamic programming algorithms, this program combines selected exons and other regions or segments with assigned scores to predict an entire gene with the maximum total score.
Genomic research makes extensive use of Hidden Markov Models (HMMs) (see Section 9.3). Markov models are implemented, for example, in the following programs:
✵ Genemark - http://exon.biology.gatech.edu/ and http://opal.biology.gatech.edu/GeneMark/gmhmm2_prok.cgi;
✵ GlimmerM - http://www.cbcb.umd.edu/software/glimmerm/;
✵ Critica - http://www.ttaxus.com/soflware.html;
✵ AMIGene - http://www.genoscope.cns.fr/agc/tools/amigene/Form/form.php;
✵ EasyGene - http://www.cbs.dtu.dk/services/EasyGene/.
In prokaryotes, it is still common practice to identify gene loci through a straightforward search for open reading frames. Naturally, this approach is unsuited for higher eukaryotes.
To distinguish between coding and non-coding regions in higher eukaryotes, researchers employ exon sensors based on statistical models of nucleotide usage frequencies, which evaluate specific dependencies observed within codon structures.
Gene expression profiles. The Human Genome is remarkably complex, comprising approximately three billion pairs of DNA nucleotides. Of this total, a mere 3% constitutes the coding sequence—that is, the portion of the genome that is transcribed into RNA and subsequently translated into protein.
The remainder of the genome consists of regions essential for the compact storage of Chromosomes, their Replication during Cell Division, the Regulation of transcription, and other vital functions.
The bulk of genome sequence analysis focuses on investigating the products of cellular transcription and Translation machinery, namely protein sequences and structures.
Recently, substantial efforts have been directed toward automating mRNA research. This is partly because the in-silico translation of mRNA into protein sequences can be readily implemented algorithmically, but primarily because mRNA molecules represent the fraction of the genome that is expressed in a specific cell type at a particular stage of its development.
Thus, three distinct levels of genomic information can be identified:
1) the chromosomal genome (the genome proper)—Genetic information common to all cells of an Organism;
2) the expressed genome (transcriptome)—the portion of the genome expressed in a cell at a specific stage of development;
3) the proteome—the repertoire of protein molecules whose interactions impart distinct individual characteristics to the cell.
Each level requires different Analytical Methods and explanatory algorithms. At various Stages of development and Levels of biological activity, cells express distinct sets of genes.
This characteristic set of expressed genes is referred to as the cell's expression profile.
By capturing the expression profiles of a given cell, one can reconstruct the landscape of gene expression levels under normal or pathological conditions, as well as the relative expression levels of all genes transcribed within that cell. Furthermore, analyzing these expression profiles enables the discovery of novel genes, thereby complementing other methodologies utilized in comprehensive genome sequencing projects.
Last update: 11/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.