Fundamentals of Bioinformatics - Ogurtsov A.N. 2013

Information Principles in Biotechnology
Protein Analysis and Prediction
The Challenge of Protein Structure Decryption

One of the primary goals of informational biotechnology is to establish a functional relationship between the Amino Acid Sequence and the three-dimensional Structure of a protein. Once this relationship is established, it will become possible to predict Cell/13.html">Protein Structure from its amino acid sequence with high accuracy. Consequently, the core objective of informational biotechnology—the computational design of target biotechnological products—will become feasible and cost-effective. Recent breakthroughs in solving the sequence-to-structure prediction problem have been made possible by novel Methods and information resources.

Compared to the nucleic acid alphabet (4 bases), the protein alphabet (20 Amino Acids) allows for encoding an incomparably larger number of Structural and functional variants. This is primarily because the differences in the Chemical Structure of amino acid residues are much more pronounced than those in NUCLEOTIDES.

Each amino acid residue can influence the overall Physical Properties of a protein, as the parent amino acid may have basic or acidic properties, be hydrophobic or hydrophilic, possess a straight or branched chain structure, or contain an aromatic ring.

Thus, each amino acid in the protein chain contributes a specific quality toward The formation of a defined type of structure within a protein domain (a conformation uniquely determined by The amino acid sequence).

Numerous observations show that when the reaction environment returns to its initial state, a denatured protein spontaneously folds back into its unique three-dimensional native conformation. This fact indicates that nature possesses an algorithm for restoring a protein's structure from its amino acid sequence. Some attempts to decipher this algorithm rely entirely on general physical principles, while others are based on the comparative analysis of known Amino acid sequences and protein structures.

We will only be able to confidently claim that humanity has mastered this natural algorithm once a computer program is created that can successfully predict the three-dimensional structures of various Proteins from their amino acid sequences alone.

Understanding a protein's structure leads to understanding its function and MECHANISM OF ACTION. Currently, There is a wide gap between the number of deciphered sequences and the number of known structures. This gap is known as the "protein sequence-structure gap" and serves as the main driving force behind The Development of Protein Structure Prediction methods. Predicting a structure means determining the relative positions of all atoms in a protein molecule in three-dimensional space using only information about its primary sequence.

Structure prediction is performed using various methods: comparative modeling, Fold Recognition, Secondary structure prediction, ab initio prediction, and knowledge-based prediction. Knowledge-based algorithms attempt to predict protein structure using information gathered from Databases of known structures.

Most ab initio Protein Structure Prediction algorithms—those relying solely on fundamental physical principles—attempt to account for all interatomic interactions within the protein molecule and determine the Free energy inherent to any possible conformation of a given protein. Computationally, the protein structure prediction problem presents itself as finding the global minimum of the free energy function for a given conformation.

So far, this approach has not succeeded for two main reasons.

First and foremost is the scale (absolute magnitude) of the problem. An average protein consists of several hundred amino acids. Each is connected to its neighbors by two flexible bonds, which possess a whole set of stable Conformations. Furthermore, each amino acid features a flexible side chain that can also adopt numerous stable conformations. Combined, these numerous torsional degrees of freedom define an unimaginably vast conformational space that cannot be handled even by the most advanced supercomputers.

The second problem lies in the method used to evaluate the stability of each trial conformation during a computer-based iterative experiment.

A folded protein contains thousands of internal contacts, each making a minuscule contribution to the stabilization of the overall structure.

A multitude of Water molecules are released during protein folding when protein chains tuck their hydrophobic regions into the interior of the globule. This release of water molecules is a major driving force compelling proteins to adopt a globular structure.

On the other hand, Entropy "resists" the formation of intrachain bonds and the stripping of water molecules from the protein strand. From the standpoint of entropy reduction, a rigid protein globule with a single conformation is energetically less favorable than a flexible, unfolded protein chain possessing a vast multitude of conformations.

The energy released upon the formation of intrachain contacts and the liberation of water molecules is expended on folding the chain into a compact shape.

Overall, these two opposing energetic contributions practically balance each other out within the system.

It is precisely this very small difference—which constitutes the stabilization energy—that we must predict when attempting to solve the protein folding problem, by selecting the single tertiary structure that possesses the highest stabilization energy.

However, the magnitude of this energy is calculated as the difference between two large values, each obtained by summing a huge number of individual contributions from the atoms of the protein molecule. Even a minor error in determining the interaction contribution of each atom in the protein results in a cumulative error that exceeds the magnitude of the stabilization energy itself.

Both factors—the enormous conformational space and the cumulative errors in the objective Functions—jointly undermine protein folding predictions.

The most successful approximations employ simplified models, often approximating the protein chain with some form of crystal lattice, in order to reduce the number of coordinates in the conformational space. However, these approximations are still far from predicting three-dimensional structures for the de novo design of Biomolecules.

An alternative to ab initio methods is Homology modeling—an approach that involves reconstructing the overall picture of a protein structure by searching for sequences that form similar structures. The methods encompassed by this approach are empirical, meaning they are based on experimental data.

An analysis of all protein structures in the Protein Data Bank (PDB) revealed that proteins sharing about 30% sequence identity have homologous structures. In these structures, the folding and topology of the protein chain are similar, although local details of specific loops may vary. Homology modeling takes advantage of this observation.

The structure of a new protein can be modeled based on the structure of a known protein with a similar amino acid sequence (provided, of course, that such information exists). In this case, computer modeling is used to predict the structure of unaligned loops and to determine the coordinates of specific amino acids by which the new protein differs from the already studied one.

For proteins exhibiting 60% or greater Sequence homology, such models can be highly accurate.

Within the 30–60% sequence homology range, such models can be useful for predicting general protein structural properties, such as identifying surface-exposed amino acid residues or predicting overall globular shape.

If no suitable homologs are available for a given protein, one must resort to an alternative approach: secondary structure prediction.

Motif (fold) recognition methods make it possible to detect distant evolutionary relationships and distinguish them from random sequence similarities not sharing a common fold. Algorithms developed on this basis search a library of known protein structures to find the one best suited for the query sequence whose structure is to be predicted. Once an alignment is established between the query sequence and the distantly related database sequences, a putative 3D protein structure model can be generated.



Last update: 11/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.