Fundamentals of Bioinformatics - Ogurtsov A.N. 2013
Foundations of Bioinformatics
Genomes and Proteomes
Genomics
Advances in biology and chemistry have significantly accelerated the decoding of Gene and protein sequences. The advent of Recombinant DNA technology has made it relatively easy to integrate foreign DNA sequences into numerous biological systems. Furthermore, this technology has enabled the rapid, mass production of specific DNA sequences—essential components for the laboratory analysis of Biological Sequences.
Oligonucleotide synthesis technology has empowered researchers to construct custom short DNA fragments from nucleotide sequences.
First, these oligonucleotides can be used to probe extensive cDNA libraries to isolate genes containing the target sequence.
Second, these DNA fragments can serve as primers in polymerase chain reactions (PCR) to amplify or modify known DNA sequences.
Biological sequence analysis is conducted when it is necessary to:
a) identify sequences that encode Proteins determining overall cellular METABOLISM (structural genes);
b) detect sequences that regulate Gene Expression or other cellular processes.
Genomics focuses on the development and Application of Molecular mapping and sequencing Methods, as well as techniques for describing, decoding, and analyzing entire organismal genomes and complete sets of gene products.
An Organism's genome is defined as the total DNA of the haploid chromosome set and each of the extrachromosomal genetic elements contained within a single germline Cell of a multicellular organism. Analysis of complete genomes provides insights into the global Organization, expression, regulation, and evolution of hereditary material (Figure 39).
Class="center">
Figure 39 - Genome Analysis: hierarchical representation
Genomics is divided into structural, functional, and comparative genomics.
Structural genomics deals with the construction of genetic and physical maps, as well as the decoding of complete genomes.
Genetic maps serve as a baseline for constructing higher-resolution physical and sequence maps, and also provide molecular entry points for Gene cloning.
Physical maps provide insight into how clones from Genomic Libraries are distributed across the entire genome. They furnish essential data for positional cloning. Genomic DNA sequences are indispensable for characterizing all gene Functions, including gene expression and regulation.
Functional genomics involves the comprehensive Study of the Structure, expression patterns, interactions, and regulation of RNA molecules and proteins encoded by The Genome. It is a comprehensive functional analysis of genes and non-coding sequences conducted at the whole-genome level.
Comparative genomics examines methods for comparing the complete genomes of various biological species to determine the function of each gene and the evolutionary relationships among the organism hosts.
Decoding the complete genomic DNA sequence of an organism enables the identification of all its genes and, consequently, the determination of its genotype. Specialized experimental methods have been developed to handle, analyze, and characterize the vast number of genes and large quantities of DNA.
Since conventional sequencing methods are applicable only to short DNA stretches (100–1000 Base Pairs), longer sequences can be fragmented and subsequently reassembled to yield the complete sequence of a large DNA segment.
A sequence refers to the order of NUCLEOTIDES within a DNA fragment. Two primary methods are used to obtain a complete sequence:
1) chromosome walking (or primer-mediated walking), which yields The sequence of a large DNA segment step by step;
2) shotgun sequencing, which is much faster but more complex, as it relies on random DNA fragments that must subsequently be assembled using specialized computer software.
Shotgun sequencing is a method used to sequence long DNA strands (see also section 11.2).
The Essence of the method lies in generating a random mass sample of cloned DNA fragments—contigs (derived from contiguous)—of a given organism (i.e., "fragmenting" the genome). These contigs are then sequenced using conventional chain-termination methods (see section 6.3 below). The resulting overlapping random DNA fragments are subsequently assembled into a single continuous sequence using specialized software. However, DNA repeats can present certain challenges during assembly.
Genomic sequence analysis demonstrates that every organism possesses both a specific set of housekeeping genes required for essential metabolic processes (such as reproduction, Glycolysis, ATP synthesis, maintenance of genetic machinery, anabolism, and Catabolism) and a set of informational genes whose products determine the organism's unique traits.
Deciphering the complete genome provides the foundational knowledge required to analyze gene expression and Protein Synthesis, yet this sequencing alone is insufficient to determine an organism's full Complement of proteins.
Genome Size—meaning The amount of Genetic information per cell—and the DNA nucleotide sequence are virtually always constant across all individuals of a given species, yet they vary widely among different species.
Table 5 shows the genome sizes of several organisms. Not all DNA encodes proteins; furthermore, some genes occur in numerous copies. Therefore, the number of genes in a genome cannot be estimated solely from genome size.
Table 5 — Genome Sizes
|
Organism |
Number of base pairs |
Number of genes |
Comment |
|
Phage φX174 |
5386 |
10 |
infects E. coli |
|
Human mitochondrion |
16569 |
37 |
subcellular organelle |
|
Epstein-Barr virus (EBV) |
172282 |
80 |
causes mononucleosis |
|
Mycoplasma pneumoniae |
816394 |
680 |
pathogen causing primary atypical Pneumonia |
|
Rickettsia prowazekii |
1 111 523 |
878 |
bacterium, CAUSATIVE AGENT OF epidemic typhus |
|
Treponema pallidum |
1 138 011 |
1039 |
bacterium, causes Syphilis |
|
Borrelia burgdorferi |
1 471 725 |
1738 |
bacterium, causes Lyme disease |
|
Aquifex aeolicus |
1 551 335 |
1749 |
thermophilic bacterium from hot springs |
|
Thermoplasma acidophilum |
1 564 905 |
1509 |
archaeon lacking a Cell wall |
|
Campylobacter jejuni |
1 641 481 |
1708 |
common cause of food poisoning |
|
Helicobacter pylori |
1667 867 |
1589 |
primary cause of Stomach ulcers |
|
Methanococcus jannaschii |
1 664 970 |
1783 |
thermophilic archaeon |
|
Hemophilus influenzae |
1 830 138 |
1738 |
bacterium, cause of Middle ear infections |
|
Thermotoga maritima |
1 860 725 |
1879 |
marine bacterium |
|
Archaeoglobus fulgidus |
2 178 400 |
2437 |
archaeon |
|
Deinococcus radiodurans |
3 284 156 |
3187 |
radiation-resistant bacterium |
|
Synechocystis |
3 573 470 |
4003 |
cyanobacterium, blue-green alga |
|
Vibrio cholerae |
4 033 460 |
3890 |
causative agent of cholera |
|
Mycobacterium tuberculosis |
4 411 529 |
4275 |
causative agent of tuberculosis |
|
Bacillus subtilis |
4214814 |
4779 |
gram-positive soil bacterium |
|
4 639 221 |
4406 |
coliform bacterium |
|
|
Pseudomonas aeruginosa |
6 264 403 |
5570 |
prokaryote |
|
Saccharomyces cerevisiae |
12,1∙106 |
5885 |
|
|
Caenorhabditis elegans |
95,5∙106 |
19099 |
nematode worm |
|
Arabidopsis thaliana |
1,17∙108 |
25498 |
flowering plant (angiosperm) |
|
Drosophila melanogaster |
1,8∙108 |
13601 |
fruit fly |
|
Fugu rubripes |
3,9∙108 |
30000 |
pufferfish (Fugu fish) |
|
Human |
3,2∙109 |
34000 |
Last update: 11/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.