Fundamentals of Bioinformatics - Ogurtsov A.N. 2013
Foundations of Bioinformatics
Subject Matter of Bioinformatics
Future Prospects of Bioinformatics Applications
In addition to providing researchers studying Proteins and DNA with a theoretical foundation and a computational-analytical toolkit, bioinformatics has found Applications across numerous fields.
Two distinct analytical approaches have emerged in deciphering the functional meaning of Biological Sequences:
✵ the first approach relies on pattern-recognition Methods to detect sequence similarities, thereby uncovering evolutionarily related structures and Functions;
✵ the second approach utilizes ab initio prediction methods (from first principles) to forecast tertiary structures and ultimately deduce function directly from the primary sequence. Direct prediction of a protein's three-dimensional Structure from its primary Amino Acid Sequence is a primary objective of bioinformatics.
Sequence Homology analysis. One of the driving forces of bioinformatics is the search for similarities among various Biomolecules. Beyond the systematic Organization of data, the identification of protein homologs has direct Practical Applications. Theoretical protein models are typically based on experimentally determined structures of close homologs.
Whenever biochemical or structural data are scarce, studies can be carried out on lower eukaryotes, such as Yeast-like organisms, and the results can be interpolated to homologous molecules in higher organisms, such as humans.
This approach significantly simplifies The problem of understanding complex genomes by directly analyzing simpler organisms and subsequently applying the same principles to more complex ones.
Such a method makes it possible to search for potential drug targets by conducting tests on homologs of essential microbial proteins.
Drug discovery. The bioinformatics-driven approach to drug discovery offers a significant advantage. Bioinformatics can characterize genotypes associated with pathophysiological states, which in principle allows for the identification of corresponding molecular targets. Subsequently, the likely amino acid sequence of the encoded target protein can be deduced from the known nucleotide sequence.
If this approach is adopted, sequence analysis methods could be used to search for homologs in model organisms. Based on sequence similarity, it would be possible to model The structure of a specific protein using experimentally established structures as a template. Finally, computer docking algorithms could design molecules that potentially bind to the target protein, selecting the most promising candidates for biochemical assays to test the biological activity of these molecules on the actual protein.
Hypothetical example. To better understand The Role of computer modeling in molecular medicine, let us imagine a future pandemic scenario caused by The Emergence of a novel biological virus. This virus triggers an epidemic of a severe disease affecting both humans and animals. Scientists in the laboratory will isolate its DNA and determine its sequence.
Next, by computationally screening this new genome against Databases of all genetic material known at the time, it will be possible to characterize the virus and identify its relationship to previously studied Viruses. The analysis will then proceed toward developing Antiviral Therapy. Viruses contain protein molecules that serve as suitable targets for drugs acting on the virus's Structure and function. From the viral DNA sequence, computer programs will calculate the Amino acid sequences of one or more viral proteins critical for viral Replication or assembly. From these amino acid sequences, other programs will compute the structures of these proteins, following the fundamental principle that a protein's amino acid sequence uniquely determines its three-dimensional structure and, consequently, its function.
First, databases will be screened to search for related proteins of known structure. If such proteins are found, the structure prediction problem will be reduced to predicting The Effect of sequence variations on the molecular structure. Homology modeling will be employed to predict the structures of the target proteins.
If no related protein with a known structure is found and the viral protein turns out to be entirely novel, structure prediction will be performed ab initio. Such situations will occur less and less frequently as the database of known structures expands and our ability to establish distant evolutionary relationships improves.
Knowledge of viral protein structures will enable drug development. The protein surfaces feature functional regions (sites) that are sensitive to inhibition. A small molecule complementary in Structure and properties to such a site will be discovered or synthesized to act as an antiviral drug. Alternatively, one or more Antibodies can be engineered and synthesized to neutralize the virus.
This sequence of events in a hypothetical situation is based on principles that are already well-established today.
Many challenges at each of these described stages remain unresolved, which is one reason why this scenario cannot be applied today—for instance, to develop anti-AIDS drugs. Another reason is that viruses "know" how to protect themselves.
Finally, it must be acknowledged that purely experimental approaches to antiviral drug discovery may continue to outperform theoretical ones for many years to come. The most likely outcome is the parallel improvement and mutual complementation of experimental and bioinformatics-based drug design methods.
Modeling. Information technologies enabling high-throughput data screening and comparison can answer a range of questions regarding the evolutionary, biochemical, and biophysical CHARACTERISTICS OF THE biomolecules under study. It has become possible to determine:
a) specific protein folding motifs corresponding to particular phylogenetic groups;
b) commonalities among different protein globule folding patterns observed in individual organisms;
c) the proportion of analogous tertiary structures shared among related organisms;
d) quantitative parameters defining the degree of relatedness derived from conventional evolutionary trees;
e) individual differences in metabolic pathways across various organisms.
Furthermore, based on the fact that protein globule folding features are frequently linked to specific biochemical functions, insights into protein function can be obtained. By analyzing Gene Expression data alongside the structural and Functional Classification of proteins, it is possible to predict the interactome (the complete network of Protein-Protein Interactions) of a given Organism and its evolution over time.
Medicine. Medical applications of bioinformatics are primarily related to gene expression analysis. Typically, researchers record expression data from Cells affected by various diseases and then compare these measurements with normal expression levels. Genes that show altered expression in diseased cells are most likely associated with the condition in question. This helps uncover the underlying causes of the disease and points to potential drug targets.
Armed with such information, researchers can develop compounds that bind to the expressed protein. Subsequent microarray experiments can then be performed to evaluate the pharmacological response to the newly synthesized candidate compound. This approach can also assist in designing tests to detect or predict the toxicity of experimental drugs during clinical trials.
Combining bioinformatics with experimental Genomics will help address A number of pressing challenges, such as: (1) postnatal genotyping to assess an individual's susceptibility or resistance to specific diseases and pathogens; (2) personalized prescription of unique vaccine combinations; and (3) reducing healthcare costs by enhancing Treatment efficacy and preventing disease relapses. Together, these innovations could pave the way for personalized diets and early-stage disease detection.
In addition, medication regimens could be tailored to the individual patient and their specific condition, thereby ensuring the most effective course of treatment with minimal side effects.
Specifically, The Human Genome and 1000 Genomes projects will undoubtedly benefit forensic medicine and the pharmaceutical industry, lead to the discovery of numerous "beneficial" and "harmful" genes, and make an invaluable contribution to our understanding of Human Evolution. Furthermore, they will facilitate The Development of diagnostic methods for diseases and potential complications, help predict genetically determined responses to therapy, and foster personalized treatment approaches, novel drug target discovery methods, and ultimately, Gene Therapy.
Intellectual Property Rights. Intellectual property rights are an integral part of modern business relations. They refer to legal mechanisms for protecting any intangible assets. Examples of intellectual property include patents, copyrights, trademarks, and trade secrets. A patent is an exclusive monopoly granted by the government to an inventor for The Use of their invention over a limited period of time.
The main areas of bioinformatics that require intellectual property protection are as follows:
a) information management and analysis tools (e.g., modeling methods, databases, algorithms, software, etc.);
c) Drug Discovery and development.
The lion's share of new developments in bioinformatics relates to the application of software (including protocols) designed to collect and/or process biological data.
These inventions fall under the broad category of computer science inventions and are subdivided into computer-implemented inventions and those utilizing machine-readable storage media.
All such inventions comprise two components: software and computer hardware.
For example, a similarity-based automated system for recognizing novel groups of nucleotide sequences within a given set of nucleotide sequences may include an input device, memory, and a processor (as hardware components), as well as a dataset or a method utilizing instructions stored in memory and executed by the processor (as the software component of the system). Patent protection is essential for safeguarding computational methods, such as sequence alignment, homology searching, and metabolic pathway modeling.
Genomics involves the isolation and characterization of genes, as well as the assignment of specific functions or roles to their sequences (e.g., the expression of a specific protein or the identification of a gene as a marker for a particular disease). This work entails numerous Laboratory tests and the application of diverse computational methods, which can likewise be protected by intellectual property rights.
Proteomics focuses on the purification and characterization of proteins using technologies such as two-dimensional gel Electrophoresis, multidimensional Chromatography, and mass spectrometry. Applying these techniques to determine protein properties and discover correlations between a protein marker and a specific disease is a highly complex and labor-intensive process requiring significant investment.
Drug discovery methods utilizing Computer-Aided Molecular Modeling, which relies on computers and computational algorithms, can also be classified as intellectual property.
METABOLISM/35.html">Selection/41.html">Review Questions and Exercises
1. Define bioinformatics.
2. What date can be considered the emergence of bioinformatics as a distinct scientific field?
3. What are the Specific features of bioinformatics data?
4. What is sequencing, and what role does it play in bioinformatics?
5. Where is bioinformatics data stored?
6. What three components constitute the Subject Matter of bioinformatics?
7. What are the MAIN OBJECTIVES OF bioinformatics?
8. What are the main challenges facing bioinformatics?
9. In what fields of activity is the subject matter of bioinformatics applied?
10. What role does the analysis of homologous sequences play in deciphering biological information?
11. How does bioinformatics contribute to drug development?
12. Which areas of bioinformatics require intellectual property protection?
Last update: 11/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.