Fundamentals of Bioinformatics - Ogurtsov A.N. 2013

Foundations of Bioinformatics
Genomes and Proteomes
Genomics

Advances in biology and chemistry have significantly accelerated the decoding of Gene and protein sequences. The advent of Recombinant DNA technology has made it relatively easy to integrate foreign DNA sequences into numerous biological systems. Furthermore, this technology has enabled the rapid, mass production of specific DNA sequences—essential components for the laboratory analysis of Biological Sequences.

Oligonucleotide synthesis technology has empowered researchers to construct custom short DNA fragments from nucleotide sequences.

First, these oligonucleotides can be used to probe extensive cDNA libraries to isolate genes containing the target sequence.

Second, these DNA fragments can serve as primers in polymerase chain reactions (PCR) to amplify or modify known DNA sequences.

Biological sequence analysis is conducted when it is necessary to:

a) identify sequences that encode Proteins determining overall cellular METABOLISM (structural genes);

b) detect sequences that regulate Gene Expression or other cellular processes.

Genomics focuses on the development and Application of Molecular mapping and sequencing Methods, as well as techniques for describing, decoding, and analyzing entire organismal genomes and complete sets of gene products.

An Organism's genome is defined as the total DNA of the haploid chromosome set and each of the extrachromosomal genetic elements contained within a single germline Cell of a multicellular organism. Analysis of complete genomes provides insights into the global Organization, expression, regulation, and evolution of hereditary material (Figure 39).

Class="center">

Figure 39 - Genome Analysis: hierarchical representation

Genomics is divided into structural, functional, and comparative genomics.

Structural genomics deals with the construction of genetic and physical maps, as well as the decoding of complete genomes.

Genetic maps serve as a baseline for constructing higher-resolution physical and sequence maps, and also provide molecular entry points for Gene cloning.

Physical maps provide insight into how clones from Genomic Libraries are distributed across the entire genome. They furnish essential data for positional cloning. Genomic DNA sequences are indispensable for characterizing all gene Functions, including gene expression and regulation.

Functional genomics involves the comprehensive Study of the Structure, expression patterns, interactions, and regulation of RNA molecules and proteins encoded by The Genome. It is a comprehensive functional analysis of genes and non-coding sequences conducted at the whole-genome level.

Comparative genomics examines methods for comparing the complete genomes of various biological species to determine the function of each gene and the evolutionary relationships among the organism hosts.

Decoding the complete genomic DNA sequence of an organism enables the identification of all its genes and, consequently, the determination of its genotype. Specialized experimental methods have been developed to handle, analyze, and characterize the vast number of genes and large quantities of DNA.

Since conventional sequencing methods are applicable only to short DNA stretches (100–1000 Base Pairs), longer sequences can be fragmented and subsequently reassembled to yield the complete sequence of a large DNA segment.

A sequence refers to the order of NUCLEOTIDES within a DNA fragment. Two primary methods are used to obtain a complete sequence:

1) chromosome walking (or primer-mediated walking), which yields The sequence of a large DNA segment step by step;

2) shotgun sequencing, which is much faster but more complex, as it relies on random DNA fragments that must subsequently be assembled using specialized computer software.

Shotgun sequencing is a method used to sequence long DNA strands (see also section 11.2).

The Essence of the method lies in generating a random mass sample of cloned DNA fragments—contigs (derived from contiguous)—of a given organism (i.e., "fragmenting" the genome). These contigs are then sequenced using conventional chain-termination methods (see section 6.3 below). The resulting overlapping random DNA fragments are subsequently assembled into a single continuous sequence using specialized software. However, DNA repeats can present certain challenges during assembly.

Genomic sequence analysis demonstrates that every organism possesses both a specific set of housekeeping genes required for essential metabolic processes (such as reproduction, Glycolysis, ATP synthesis, maintenance of genetic machinery, anabolism, and Catabolism) and a set of informational genes whose products determine the organism's unique traits.

Deciphering the complete genome provides the foundational knowledge required to analyze gene expression and Protein Synthesis, yet this sequencing alone is insufficient to determine an organism's full Complement of proteins.

Genome Size—meaning The amount of Genetic information per cell—and the DNA nucleotide sequence are virtually always constant across all individuals of a given species, yet they vary widely among different species.

Table 5 shows the genome sizes of several organisms. Not all DNA encodes proteins; furthermore, some genes occur in numerous copies. Therefore, the number of genes in a genome cannot be estimated solely from genome size.

Table 5 — Genome Sizes

Organism

Number of base pairs

Number of genes

Comment

Phage φX174

5386

10

infects E. coli

Human

mitochondrion

16569

37

subcellular organelle

Epstein-Barr virus (EBV)

172282

80

causes mononucleosis

Mycoplasma pneumoniae

816394

680

pathogen causing primary atypical Pneumonia

Rickettsia prowazekii

1 111 523

878

bacterium, CAUSATIVE AGENT OF epidemic typhus

Treponema pallidum

1 138 011

1039

bacterium, causes Syphilis

Borrelia burgdorferi

1 471 725

1738

bacterium, causes Lyme disease

Aquifex aeolicus

1 551 335

1749

thermophilic bacterium from hot springs

Thermoplasma acidophilum

1 564 905

1509

archaeon lacking a Cell wall

Campylobacter jejuni

1 641 481

1708

common cause of food poisoning

Helicobacter pylori

1667 867

1589

primary cause of Stomach ulcers

Methanococcus jannaschii

1 664 970

1783

thermophilic archaeon

Hemophilus influenzae

1 830 138

1738

bacterium, cause of Middle ear infections

Thermotoga maritima

1 860 725

1879

marine bacterium

Archaeoglobus fulgidus

2 178 400

2437

archaeon

Deinococcus radiodurans

3 284 156

3187

radiation-resistant bacterium

Synechocystis

3 573 470

4003

cyanobacterium, blue-green alga

Vibrio cholerae

4 033 460

3890

causative agent of cholera

Mycobacterium tuberculosis

4 411 529

4275

causative agent of tuberculosis

Bacillus subtilis

4214814

4779

gram-positive soil bacterium

Escherichia coli

4 639 221

4406

coliform bacterium

Pseudomonas aeruginosa

6 264 403

5570

prokaryote

Saccharomyces cerevisiae

12,1∙106

5885

Yeast

Caenorhabditis elegans

95,5∙106

19099

nematode worm

Arabidopsis thaliana

1,17∙108

25498

flowering plant (angiosperm)

Drosophila melanogaster

1,8∙108

13601

fruit fly

Fugu rubripes

3,9∙108

30000

pufferfish (Fugu fish)

Human

3,2∙109

34000




Last update: 11/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.