Fundamentals of Bioinformatics - Ogurtsov A.N. 2013
Foundations of Bioinformatics
Biological Sequences
Information in Molecular Biology
The information archive (genome) of every Organism contains a detailed blueprint for the future development and functioning of that individual. DNA molecules are long, linear, chain-like polymers that form strings written in a four-letter alphabet (Figure 6).
Class="center">
Figure 6 – Diagram of Introduction/20.html">DNA Structure: a – Van der Waals model, b – schematic representation of chemical bonds in DNA
Even the genomes of microorganisms are very long strings, typically consisting of millions of letters. When writing any nucleic acid as a string of characters, it is standard practice to denote NUCLEOTIDES using lowercase English letters:
![]()
The DNA structure fully governs the mechanisms of Replication and The transfer of information from Gene to protein. Nearly flawless replication is essential for genetic stability. However, a slight inaccuracy in replication—much like the mechanism for importing foreign genetic material—is also necessary; otherwise, asexually reproducing organisms could not evolve. The strands of The Double Helix are antiparallel. Their ends are designated as 3' and 5' based on the positions in the deoxyribose ring. During METABOLISM/31.html">Transcription, DNA is read in the 3' to 5' direction, whereas during Translation, mRNA is read in the 5' to 3' direction (see [7], sections 2.1 and 5.1). Genetic information is expressed through the synthesis of RNA and Proteins. Proteins are the molecules responsible for the vital activity of most biological structures. Our Hair, Muscles, Digestive System, receptors, and Antibodies are all proteins. Like Nucleic Acids, proteins
are long, linear, chain-like polymers composed of monomers
– Amino Acids. The twenty natural (proteinogenic) Amino acids can be classified based on the polarity of their side chains into nonpolar, polar, and charged.
We will use single-letter designations for amino acids represented by uppercase Latin letters as follows:
|
Nonpolar amino acids: |
||
|
G – Glycine (Gly) |
A – Alanine (Ala) |
P – Proline (Pro) |
|
V – valine (Val) |
I – isoleucine (Ile) |
L – leucine (Leu) |
|
F – phenylalanine (Phe) |
M – Methionine (Met) |
|
|
Polar amino acids: |
||
|
S – Serine (Ser) |
C – Cysteine (Cys) |
T – Threonine (Thr) |
|
N – asparagine (Asn) |
Q – glutamine (Gln) |
Y – Tyrosine (Tyr) |
|
W – Tryptophan (Trp) |
||
|
Charged amino acids: |
||
|
D – aspartic acid (Asp) |
K – Lysine (Lys) |
|
|
E – glutamic acid (Glu) |
R – Arginine (Arg). |
|
The Genetic Code is a cipher: triplets of letters from the DNA sequence specify amino acids (Table 2).
Table 2 – The Standard Genetic Code
|
First nucleotide |
Second nucleotide |
Third nucleotide |
|||
|
u |
c |
a |
g |
||
|
u |
Phe |
Ser |
Tyr |
Cys |
u |
|
Phe |
Ser |
Tyr |
Cys |
c |
|
|
Leu |
Ser |
STOP |
STOP |
a |
|
|
Leu |
Ser |
STOP |
Trp |
g |
|
|
c |
Leu |
Pro |
His |
Arg |
u |
|
Leu |
Pro |
His |
Arg |
c |
|
|
Leu |
Pro |
Gln |
Arg |
a |
|
|
Leu |
Pro |
Gln |
Arg |
g |
|
|
a |
Ile |
Thr |
Asn |
Ser |
u |
|
Ile |
Thr |
Asn |
Ser |
c |
|
|
Ile |
Thr |
Lys |
Arg |
a |
|
|
Met (START) |
Thr |
Lys |
Arg |
g |
|
|
g |
Val |
Ala |
Asp |
Gly |
u |
|
Val |
Ala |
Asp |
Gly |
c |
|
|
Val |
Ala |
Glu |
Gly |
a |
|
|
Val |
Ala |
Glu |
Gly |
g |
|
DNA regions encode the Amino acid sequences of proteins. Typically, proteins consist of 200–400 amino acids, which requires 600–1200 DNA nucleotides to encode them. The synthesis of RNA molecules, such as the RNA components of Ribosomes, is also determined by The nucleotide sequence of DNA. However, in most organisms, not all DNA codes for RNA or proteins. Certain DNA sequence segments exist to regulate Transcription and Replication processes, while a large portion of The Genome remains unexplored and its Functions are still unknown. DNA molecules containing the standard four "nucleotide" letters (a, c, g, t) are similar in chemical structure, and the Spatial Structure of the DNA molecule is, to a first approximation, uniform.
Proteins, on the contrary, exhibit a vast diversity of three-dimensional Conformations. These conformations are essential for proteins to perform their diverse Structural and functional roles. The Amino Acid Sequence of a protein—its Primary Structure—determines its three-dimensional structure. For every natural amino acid sequence, There is a unique, stable Native State, known as the tertiary structure, into which the sequence spontaneously folds under physiological conditions (voir [9], section 4.3).
If a purified protein is heated or otherwise subjected to conditions that deviate significantly from natural physiological conditions, it "unfolds" (denatures), forming a disordered, biologically inactive structure. This is precisely why our bodies possess mechanisms to maintain relatively constant internal conditions (see [13], section 2.4).
Conversely, when normal conditions are restored, polypeptide molecules regain their functional tertiary structure, which is indistinguishable from the naturally occurring native structure (see [9], section 4.5).
Spontaneous protein folding—The process of forming their native structure—represents the exact point where Nature makes a giant leap from one-dimensional genetic and peptide sequences to the three-dimensional world in which we all live.
However, the following paradox arises.
On the one hand, translating DNA sequences into amino acid sequences is logically straightforward, as it is dictated by the genetic code. Folding a polypeptide chain into a precisely defined three-dimensional structure, on the other hand, is extremely difficult to describe logically. Conversely, translation requires an extraordinarily complex ribosomal machinery, Transfer RNAs (tRNAs), and associated molecules (see [7], section 5), whereas protein folding occurs spontaneously without external assistance (see [9], section 4.5).
Protein Functions depend on the acquisition of their native tertiary structure. For example, the native structure of an enzyme may feature a cleft (Active Site) on its surface that binds a single small substrate molecule and positions it in close proximity to The amino acid residues of the catalytic center.
Thus, we have the following information-driven dependencies:
✵ The nucleotide sequence of DNA determines the amino acid sequence of a protein.
✵ The amino acid sequence determines the Cell/13.html">Protein Structure.
✵ The protein structure determines its function.
For the most part, bioinformatics is precisely concerned with the analysis of data related to these processes.
Currently, this paradigm does not extend beyond THE MOLECULAR LEVEL of structure and Organization. Consequently, issues such as tissue specialization during development or, more broadly, the IMPACT OF ENVIRONMENTAL conditions on genetic events fall outside its scope.
In certain trivial cases involving simple feedback loops, it is easy to understand the molecular mechanisms by which an increase in Substrate Concentration leads to enhanced activity of The enzyme catalyzing its transformation (see [9], clause 15.1). Considerably more complex are the Developmental Programs of an organism throughout its lifecycle.
These fundamental questions concerning the flow of information and its regulation within the organism are now actively being investigated using bioinformatics Methods.
Last update: 11/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.