Fundamentals of Bioinformatics - Ogurtsov A.N. 2013
Information Principles in Biotechnology
Genome Function and Organization
Challenges in Gene Analysis
The successful sequencing of complete Organism genomes initially created the illusion that determining The nucleotide sequence of an organism's Chromosomes was sufficient to directly obtain information about the Structure AND Functions of the biological macromolecules encoded in The Genome. Reality proved to be significantly less optimistic. Simply mapping Biological Sequences turned out to be merely the "tip of the iceberg" of the problems that had to be solved before the available sequencing results could be translated into clear Answers regarding the Biological Significance of specific macromolecules, and into recommendations for the industrial application of this sequence bioinformatics to produce novel products.
Let us consider a specific Gene contained within a given genome that encodes an individual protein. Only in the simplest case of Prokaryotic Genomes does this gene represent a nucleotide sequence of a single DNA region. In eukaryotes, the majority of genes are "scattered" across different Regions of the DNA molecule, making the Determination of the structure and localization of such a gene within the genome a non-trivial task.
The DNA sequence is collinear with the protein sequence. In species where the genetic material is represented by double-stranded DNA, genes may reside on either strand.
Bacterial genes are continuous stretches of DNA. Thus, the functional unit of Genetic information in Bacteria is a sequence of 3N NUCLEOTIDES encoding a sequence of N Amino Acids, or a sequence of N nucleotides encoding a structural RNA molecule
(e.g., ribosomal) consisting of N residues. Such an annotated sequence can be stored as a typical entry in one of the genetic sequence archives.
In eukaryotes, The nucleotide sequences encoding the Amino acid sequences of individual Proteins are organized in a much more complex manner. Here, the relationship between the size of a gene and the protein it encodes is entirely different from that in bacteria. Frequently, a single gene is presented as separate segments of genomic DNA.
An exon is a DNA region retained in the Messenger RNA that the ribosome translates into a protein. An intron is an intervening DNA region between two exons. Cellular machinery splices specific segments within RNA transcripts based on signal sequences flanking the exons (see [7], sec. 2.4). Many introns are very long—much longer than exons.
Regulatory mechanisms orchestrate Gene Expression. Genes can be turned on or off (or fine-tuned) in response to varying nutrient concentrations, stress, or complex programs of Tissue and organ development throughout the organism's lifespan.
Numerous control regions of DNA are located near protein-coding regions. They contain sequences that serve as binding sites for DNA-transcribing molecules or sequences that bind regulatory molecules (Inducers, activators, repressors) capable of controlling METABOLISM/31.html">Transcription rates (see [7], secs. 3.1–3.5). Genome Analysis must also include the deciphering of such "administrative" non-coding DNA regions.
In bacterial genomes, adjacent genes that encode multiple proteins catalyzing sequential steps of a single biochemical pathway are grouped into operons. Genes within an Operon are switched on and off together because they are controlled by a common regulatory sequence (see [7], sec. 2.3).
In animals, DNA Methylation mechanisms ensure tissue-specific differential gene expression during development.
The products of certain genes trigger Cell apoptosis, and disruptions in the apoptotic mechanism leading to the uncontrolled growth of tissue Cells are characteristic of certain types of Cancer tumors. Blocking these mechanisms is a primary approach in cancer Treatment.
Thus, reducing genetic information solely to individual coding DNA sequences inevitably leads to a loss of information regarding the highly complex nature of interactions between them and other molecular components that ensure the expression of a given gene at the right time, in the right cell, and with the required intensity, while ignoring the historical and integrative aspects of the genome.
Metaphorically, one can imagine The Human Genome, which contains about three gigabytes of information, as three gigabytes of files on a computer's hard drive. Sequencing this genome is equivalent to acquiring a hard drive image. This image itself will contain not only the files of interest (fragments of which may be scattered across the entire drive), but also remnants of previously deleted files—the "junk" from which the functional gene-files of our genome must be extracted and recovered. Yet that is not all, for to understand the operation of the computer as a whole, simply recovering files from the disk is insufficient; one must also investigate all the technologies and engineering solutions used in creating all those devices whose cooperative work in "maintaining" the hard drive's file system constitutes a properly functioning computer. In this analogy, the computer represents the organism, whose functioning is determined by the "file" genome but is not reducible to it.
The Genetic Code directly links the Amino Acid Sequence of a protein to the nucleotide sequence in DNA. Today, information about The sequence of a novel protein is more often obtained through genome analysis than through direct protein sequencing.
The first problem that arises in such analysis is the reliability of predicting novel proteins from DNA sequences. Computer programs designed to recognize protein-coding sequences within DNA molecules are prone to Three types of errors:
1) the predicted Introduction/19.html">Primary Structure of the protein differs from the true one;
2) an incorrect splicing of the RNA transcript is predicted;
3) the protein is described incompletely.
Aside from these trivial errors, several other problems significantly complicate protein prediction.
First and foremost, the presence of Alternative Splicing in different cells of various Tissues (see [7], sec. 2.5) gives rise to a whole spectrum of related proteins even within a single organism.
Furthermore, even those genetic sequences that appear plausible and seemingly ought to encode proteins may either be defective remnants of billions of years of evolution within the genome or simply non-expressible.
Thus, the general rule in protein Prediction Based on genetic information analysis is as follows: a protein derived from a genomic sequence remains a hypothetical object until its existence is experimentally detected.
The second set of problems is associated with post-translational modifications of proteins. Very often, the product of gene expression is a protein molecule that must subsequently undergo Processing within The Cell. As a result of such processing, a functional native protein molecule may be formed that differs radically from the protein predicted by deciphering the gene sequence.
Above all, the modification of amino acid side chains (methylation, Acetylation, phosphorylation, carboxylation, hydroxylation, glycosylation, etc.), which are predominantly carried out by Enzymes of the Endoplasmic reticulum and the Golgi apparatus, can both alter the Amino Acid Composition of the protein and produce a hybrid compound of a polypeptide with Lipids and Oligosaccharides (see [6], sec. 6.5.3).
Proper glycosylation of synthesized proteins is especially critical in the manufacturing of pharmaceutical products. For instance, human Blood proteins of different groups (O, A, B, AB) differ precisely in their carbohydrate residues attached by specific Glycosyltransferases.
Another important type of post-translational modification involves the proteolytic Cleavage of the initial protein chain. This is precisely how enzymes involved in Digestion, blood clotting, and apoptosis are activated (see [9], sec. 12.1).
A clear example of protein activation via proteolysis is The Biosynthesis of Insulin. Insulin is initially translated as a proinsulin polypeptide (Figure 74) consisting of 84 amino acids. Protein folding occurs in this exact form, followed by The formation of three intramolecular Disulfide Bonds. Only then is the C-segment (consisting of 33 amino acids) excised from proinsulin, leaving a 51-amino-acid functional hormone molecule composed of two polypeptide chains: chain A with 21 Amino Acids and chain B with 30 amino acids.
Post-translational modifications also include The addition of prosthetic groups, such as the covalent attachment of a heme group in cytochrome c.
Notably, the formation of disulfide bridges, which is essential for the functioning of many proteins (such as insulin), cannot be predicted by analyzing either the DNA nucleotide sequence or the corresponding protein amino acid sequence. The same applies to predicting Protein Structure from gene structure, during the expression of which mRNA splicing occurs.
Class="center">
Figure 74 - Primary structure of proinsulin
The third problem encountered in genome analysis is that an organism's genome provides a complete yet static set of characteristics for the potentially possible macromolecular structures of that organism. The actual functioning of a given organism at a given time and under specific conditions is determined at THE MOLECULAR LEVEL by the expression intensity of specific genes and the distribution of these expression products across the organism's tissues and Organs.
To characterize the set of all proteins present in a given organism at a given time, THE CONCEPT OF the organism's proteome is used (see sec. 6.4). An organism's proteome changes over time.
Figure 75 shows the expression intensities of cyclin proteins in various tissues and organs, obtained using DNA Microarrays.

Figure 75 - Expression levels of cyclins in various tissues and organs: 1 - A; 2-B; 3-C; 4-D1; 5-D2; 6-D3; 7-E; 8-F; 9-G; 10-H; 11-I
Cyclins are a family of proteins that act as activators of cyclin-dependent Kinases (CDKs) — Key Enzymes involved in regulating The Eukaryotic Cell cycle. Cyclins got their name because their intracellular concentration fluctuates periodically as cells progress through the Cell Cycle, reaching a peak at specific stages. As seen in Figure 2, at any given time, the concentration of cyclins varies significantly across different tissues and organs. Analyzing such multiple-expression "fingerprints" helps reveal correlations (both positive and negative) in the functioning of various cells within a single organism.
Proteomics is dedicated to the systematic Analysis of Protein synthesis profiles in various organism tissues and the analysis of how these profiles depend on EXTERNAL FACTORS AND the organism's developmental stage (see sec. 6.4).
The proteome reflects the biological activity of the genome in a dynamic state.
Metaphorically speaking, if the genome is the musical score of a symphony, the proteome is the orchestra performing it.
Last update: 11/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.