Fundamentals of Bioinformatics - Ogurtsov A.N. 2013

Information Principles in Biotechnology
Protein Analysis and Prediction
Prediction of 3D Protein Structures

The tertiary Structure is the foundation of protein functionality, which requires the precise Spatial Organization of large amino acid assemblies. The tertiary structure is defined as the spatial arrangement of all atoms within a protein molecule, and it is entirely determined by its Primary Structure.

Not all protein sequences fold into stable structures. According to current estimates, only a small fraction of possible amino acid combinations are capable of forming stable structures. Researchers believe there are only about 1,000 ways to fold a protein chain into a stable structure.

There are numerous well-documented Examples where similar, homologous Amino acid sequences fold into comparable three-dimensional structures that differ only in minor details.

For instance, the Amino Acid Sequence of Hemoglobins synthesized by various animals varies significantly. Out of the 140–150 Amino Acids in the Hemoglobin chain, only two are conserved across all variant forms: Histidine, which directly coordinates with the active iron ion in the heme group, and phenylalanine, which is essential for the proper orientation of the heme cofactor. All Other Amino Acids may vary, enabling modifications in hemoglobin Functions (e.g., oxygen affinity can differ by a factor of 100,000) or simply resulting from Genetic Drift. Despite this enormous range of Variability, all hemoglobins adopt a similar three-dimensional functional structure during folding.

Structure prediction is quite reliable for homologous amino acid sequences. For sequences sharing 30–40% identity, There is a high probability that their three-dimensional structures will be similar enough to apply bioinformatics Methods of Protein modeling.

However, biotechnologists face a different kind of challenge. The aforementioned statistics apply only to natural Biomolecules whose structures have been optimized by evolution for efficient folding. Point Mutations in the structure can be fatal to protein functionality and folding, even though the protein may still exhibit a high degree of Sequence Homology to successfully folding Proteins.

Therefore, when attempting to modify an existing protein to achieve a new function, changes should be introduced in small increments, verifying each time whether the modified protein retains its ability to fold functionally.

Organizing protein structures according to their folding patterns is logically very convenient for data representation in the Protein Data Bank (PDB). This principle underpins information retrieval. Several Databases derived from PDB are built on the Classification of protein structures, offering user-friendly tools for structural analysis—such as keyword and sequence searches, navigation among similar structures across different levels of the systematic hierarchy, database scanning for structures resembling a query structure, and links to external resources. Examples of such databases include SCOP (Structural Classification of Proteins) and CATH (Class, Architecture, Topology, Homology).

Cell/13.html">Protein Structure databases and homology search tools for Protein Families:

SCOP

http://scop.mrc-lmb.cam.ac.uk/scop

САТН

http://www.cathdb.info/

InterPro

http://www.ebi.ac.uk/interpro/

Pfam

http://pfam.sanger.ac.uk/

PANTHER

http://www.panthcrdb.org/

TIGRFAMs

http://www.jcvi.org/cgi-bin/tigrfams/index.cgi

iProClass

http://pir.georgetown.edu/iproclass/

ProDom

http://prodom.prabi.fr/prodom/current/html/home.php

SMART

http://smart.embl-heidelberg.de/

PRINTS

http://www.bioinf.manchester.ac.uk/dbbrowser/PRINTS/index.php

PROSITE

http://prosite.expasy.org/

CDD

http://www.ncbi.nlm.nih.gov/Structure/cdd/cdd.shtml

COG

http://www.ncbi.nlm.nih.gov/COG/

Dali

http://ekhidna.biocenter.helsinki.fi/daliserver

VAST

http://www.ncbi.nlm.nih.gov/Structure/VAST/vast.shtml

CE

http://source.rcsb.org/jfatcatserver/ceHome.jsp

CATHEDRAL

http://v3-4.cathdb.info/cgi-bin/CathedralServer.pl

SSM

http://www.ebi.ac.uk/msd-srv/ssm/

FATCAT

http://fateat.burnham.org/fatcat/

AnnoLite

http://www.salilab.org/DBAli/

The SCOP Database. The SCOP (Structural Classification of Proteins) database hierarchically organizes protein structures according to their evolutionary relationships and structural features.

At the lowest level of the SCOP hierarchy are domains extracted from PDB. Sets of domains are grouped into families of homologs, where similarities in structure, sequence, and sometimes function indicate a common evolutionary origin.

Families containing proteins with similar structures and functions, but lacking clear evidence of common ancestry, are grouped into superfamilies.

Superfamilies that share a common folding topology (at least for their core region) are grouped into folds.

Finally, folds are grouped into classes. The main classes in SCOP are a, ß, a+ß, a/ß, and a diverse group of "small proteins" that often feature minimal Secondary structure and are stabilized by Disulfide Bonds or ligands.

Example of classification in the SCOP database for the flavodoxin protein from Clostridium beijerinckii

1. ROOT: SCOP

2. Class: Alpha and beta proteins (a/ß). Mostly parallel ß-sheets (ß-a-ß units)

3. Fold: Flavodoxin-like. 3-layer, a/ß/a; parallel ß-sheets of 5 strands, strand order 21345

4. Superfamily: Flavoproteins.

5. Family: Flavodoxin-related binds FMN

6. Protein: Flavodoxin

7. Organism: Clostridium beijerinckii

Release 1.75 of SCOP, dated February 23, 2009, contained 38,221 structures from PDB divided into 110,800 domains as of November 2012. The distribution of entries across various hierarchical levels is presented in Table 21.

Table 21 - Statistics of protein distribution in the SCOP database

Class

Count

families

superfamilies

folds

All-a proteins

871

507

284

All-ß proteins

742

354

174

a/ß proteins

803

244

147

a+ß proteins

1055

552

376

Multi-domain proteins

89

66

66

Membrane and cell surface proteins

123

110

58

Small proteins

219

129

90

Total

3902

1962

1195

The results of comparative protein structure analysis demonstrate that newly predicted protein structures often feature folds or Conformations similar to those already known. It has also been experimentally established that many different amino acid sequences can adopt the same Spatial Structure. Statistical analysis of these sequences reveals that identical short, regular amino acid combinations recur across various structural conformations.

Experiments comparing amino acid sequences with their three-dimensional folding arrangements have identified more than 500 major structural folding patterns occurring within domains across over 13,000 3D protein structures from the PDB (http://www.pdb.org/). Furthermore, these studies have shown that identical folding conformations can arise from A wide variety of primary amino acid sequences. Thus, numerous amino acid combinations can spontaneously fold into identical 3D conformations, filling available space and forming appropriate contacts with neighboring residues to establish the overall spatial structure.

The volume of investigated and annotated protein spatial structures is now so vast that the probability of a new sequence matching a known folding pattern is high enough to warrant automated, computer-based methods for protein structure analysis and prediction. The goal of Fold Recognition is to determine which of the possible folds best fits a novel sequence. Three-dimensional structure prediction employs Hidden Markov Models (HMMs) (see section 9.3) or threading techniques.

If two proteins exhibit significant sequence similarity, they are very likely to share similar three-dimensional structures. This similarity may span the entire length of the sequences or be restricted to one or more specific regions featuring relatively short regular combinations of monomers (with adjacent or interspersed gaps). It is generally accepted that if a global sequence alignment shows greater than 45% amino acid identity, the respective residues will be largely compatible within the 3D protein structure.

Consequently, if The structure of one aligned protein is known, the STRUCTURE OF THE second protein, along with the spatial positions of identical amino acids, can be predicted with high confidence. When sequence identity falls between 25% and 45%, the respective protein structures are still likely to be similar; however, lower sequence identity correlates with more pronounced deviations in spatial positioning.

Homology modeling. Homology (comparative) modeling should be employed when significant structural similarity exists between a target protein of unknown structure and another protein with a known 3D structure. The two sequences are aligned to identify homologous segments. When multiple similar structures are available, Multiple Sequence Alignment is used. The reliability of comparative modeling prediction increases with the number of homologous structures considered.

Based on the alignment that identifies corresponding amino acid residues, the structure of the protein of interest is predicted by evaluating the structures of its homologs. Several algorithms have been developed for this stage, which are classified into:

1) rigid-body assembly;

2) segment matching;

3) spatial restraint satisfaction.

Rigid-body assembly algorithms construct protein structures from incompressible Van der Waals building blocks—such as amino acid residues, a-helices, ß-sheets, and prosthetic groups—much like crystals are assembled from incompressible atomic or molecular units. These building blocks are identified from homologous structures and added to a framework whose configuration is determined by averaging reference atomic positions within conserved fold Regions of the template and its homologs. The segment matching program calculates coordinates based on the approximate positions of conserved atoms within template structures, utilizing a database of short protein structure segments.

Additionally, both geometric (more precisely, stereochemical and steric) constraints and thermodynamic requirements for Free energy Minimization are accounted for to ensure stable structures.

Steps of the homology modeling algorithm.

1. Align The amino acid sequences of the target protein and the protein(s) with known structure. Experience shows that insertions and deletions typically occur in the loop regions connecting a-helices and ß-sheets.

2. Identify backbone segments containing insertions or deletions. Sticking (joining) these segments to the backbone of the known template protein yields the target protein backbone model. Joining involves removing template regions absent in the target protein and inserting regions present in the target but absent in the template.

3. Replace the side chains of mutated amino acid residues while preserving side-chain conformations for non-mutated residues. Mutated residues tend to retain similar side-chain conformations, a property leveraged in modeling. Furthermore, computational algorithms have been developed to search for optimal side-chain conformations among combinatorial possibilities.

4. Verify the model (both visually and computationally) to detect significant steric clashes between van der Waals spheres of different atoms. Resolve such clashes manually wherever possible.

5. Minimize the Free energy of the resulting structure. While keeping the amino acid sequence and standard folding motifs intact, allow side chains to undergo minor adjustments to reach a favorable low-energy state corresponding to the global energy minimum. In practice, this step offers primarily a cosmetic refinement; energy minimization in such models cannot correct major structural errors introduced during the backbone joining stages.

Essentially, this algorithm constructs a model of a novel protein structure by introducing minimal modifications to an available template structure. Unfortunately, without accounting for additional factors, significantly improving such models is challenging. An empirical rule of thumb states that if two sequences share at least 40–50% identity, the described Procedure yields a model sufficiently accurate for many Applications. At lower sequence similarity levels, neither this procedure nor any other currently available algorithm can produce a detailed, highly accurate model based solely on related template structures.

The structures of most protein families comprise both relatively conserved and more variable regions. The structural core of a family retains its folding topology (is conserved), albeit potentially distorted, whereas peripheral regions may be entirely restructured.

When only a single ancestor structure is available, the conserved region of the target protein can be modeled with reasonable reliability, but modeling the variable region remains unfeasible. Moreover, predicting which regions are variable and which are conserved is far from trivial.

A more favorable scenario arises when several related proteins of known structure serve as "parents" for homology modeling. Their comparative analysis helps identify conserved and variable Structural domains within the family. The observed distribution of structural variability across parent templates establishes the constraint boundaries for the modeling algorithm.

Fold recognition. Searching a sequence database with a sequence or a structure database with a structure are well-defined tasks with established solutions. Mixed tasks (searching a sequence database with a structure, or a structure database with a sequence) are less straightforward and require methods to evaluate the compatibility of a given sequence with a specific folding pattern.

The objective is to establish meaningful correlations between amino acid sequences and structural folding patterns. Proteins sharing the same pattern are expected to adopt similar structures.

Protein threading, also known as fold recognition, is a computational method used to model the three-dimensional structure of proteins for which no homologues are yet known (i.e., absent from the PDB database).

Based on the empirical statistical rule that similar amino acid sequences tend to adopt similar folding structures, a threading program searches the PDB database for proteins whose amino acid sequences in functional chain fragments resemble those of the target protein. Once such matches are identified, it assembles the Spatial structure of the protein under study by combining structural elements from known proteins.

Protein threading relies on two key principles: first, the number of genuinely distinct folds is limited and does not exceed 1,000; second, among all newly solved protein structures deposited in the PDB, 90% exhibit folds analogous to those already present in the database.

The core threading algorithm involves generating A large number of rough models for a given sequence by exploring all possible alignments with sequences of known structure.

Both threading and homology modeling aim to generate a model of a protein's 3D structure by aligning a target sequence with a sequence of known three-dimensional structure. However, while homology modeling seeks to predict the detailed spatial structure of the target protein through Multiple Sequence Alignments with homologous proteins, threading employs numerous distinct pairwise alignments and utilizes relatively coarse models, sometimes without explicitly constructing them.

In homology modeling, one first identifies homologues to the target protein, then constructs an optimal multiple sequence alignment, and finally optimizes the resulting single model.

In contrast, threading evaluates all possible alignment variants against all available proteins, selects cases that show at least a rough match, and combines them to build a model of the protein of interest.

Successful fold recognition via threading requires:

1. A scoring function to evaluate models and select the best candidate.

2. A calibration procedure for the evaluation method to determine how adequately the selected protein model fits the threading objective and whether it is biologically meaningful.

As a rule, automated computer threading does not immediately yield the exact structure of the target protein. While it effectively narrows down the pool of potential folds, the final decision regarding model adequacy rests with the researcher. In any case, threading methods maximize the automation of Protein Structure Prediction, yielding a relatively narrow set of structures that, with high probability, includes one similar to the three-dimensional fold of the target sequence.

Global conformation energy minimization and Molecular Dynamics. A native protein consists of numerous atoms whose interactions drive it toward a state of maximum stability.

Analytical determination of this state is a formidable challenge. Existing interatomic potential functions lack sufficient precision; furthermore, even if an adequate model is successfully built, one faces The problem of optimizing a nonlinear objective function within a vast variable space subject to nonlinear constraints, resulting in a highly complex energy landscape with numerous local minima.

Interactions between atoms can be divided into two main categories:

1. The covalent bond system — strong interactions that keep atoms at a short distance from one another. These are treated as permanent interactions that remain intact during structural rearrangements of protein molecules and are preserved across all conformations.

2. Weaker non-covalent interactions, the magnitude of which depends on chain conformation. Their contribution can be significant in some conformations and negligible in others, depending on how closely and in what manner atoms approach each other during conformational transitions.

A protein's conformation can be defined by specifying the set of constituent atoms, their spatial coordinates, and the network of chemical bonds connecting them (this information can be reliably derived from the protein's amino acid sequence). When estimating Conformational Energy, the following contributions are taken into account:

✵ Bond stretching energy:

where r0 is the equilibrium interatomic distance, and Kr is the bond force constant; both r0 and Kr depend on the type of chemical bond.

✵ Valence angle bending energy:

For any i-th atom forming chemical bonds with two (or more) other atoms j and k, the j—i—k angle is characterized by an equilibrium value θ0 and a force constant Kθ.

✵ Additional terms responsible for stereochemical correctness, penalizing deviations from planarity in specific groups or maintaining the correct chirality at designated centers.

✵ Torsional rotation energy:

For any four consecutively bonded atoms, i — j — k — l, the rotation of atom l relative to atom i around the j - k bond axis is governed by an energy barrier with a periodic potential. Vn is the barrier height of internal rotation, and n represents the number of barriers encountered during a full 360° rotation. Example: the torsion angles φ, ψ, χ, and ω between atoms in a peptide chain (see [9], section 4.2).

✵ Dispersion (Van der Waals) interaction energy:

For each pair of non-bonded atoms i and j, the first term accounts for short-range Pauli repulsion forces, while the second term accounts for long-range attraction forces. The parameters A and B depend on the atom types, and Rij is the distance between atoms i and j.

Hydrogen bond energy:

A hydrogen bond is a weak chemical-electrostatic interaction between two electronegative atoms. The energy of such an interaction depends on both interatomic distance and angle. The interaction energy formula given above does not explicitly reflect the dependence on the bond angle. Alternative formulations may incorporate this angular parameter.

✵ Electrostatic interaction energy:

Qi and Qj are the effective partial charges on atoms i and j; Rij is the distance between them; ε is the dielectric permittivity of the medium. This formula is merely an approximation when applied to media such as proteins, which are neither continuous nor isotropic.

✵ Solvation energy. Interactions with the solvent, Water, and other solution components (such as salts and sugars) significantly influence the Thermodynamics of protein structure. Treating Solvents as continuous media characterized by dielectric permittivity as the primary parameter is only an approximation. With the advancement of computer technology, it has become feasible to simulate proteins within a simulation box containing explicitly defined solvent (water) molecules.

A wide variety of potential functions exist to describe conformational energy. The energy of a given conformation is calculated by summing the interactions of various types across all atoms in the system.

The accuracy of the potential energy function is a necessary, but not sufficient, condition for successful protein structure prediction.

One way to test this premise is to take an experimentally determined protein structure as the starting conformation and attempt to minimize its energy. Typically, the root-mean-square deviation (RMSD) of such a minimized structure from the original is on the order of 0.1 nm. This value can be considered a measure of the force field's resolution.

Another approach involves minimizing the conformational energy of a misfolded protein globule. This helps determine whether the energy minimum of the correctly folded structure lies significantly lower than the local minimum of the misfolded globule.

The results of such calculations demonstrate that conformational energy calculations alone cannot reliably distinguish the native protein conformation from the ensemble of alternative conformations.

Attempts to Predict protein structure via conformational energy minimization have not yet led to a method capable of making predictions based solely on the amino acid sequence.

To overcome both obstacles—(1) the problem of getting trapped in false local minima and (2) the lack of an adequate solvent interaction model—The Molecular Dynamics method was developed.

The protein, along with explicitly defined solvent, is simulated using Newtonian mechanics within the framework of a defined force field.

Although this methodology is not yet mature enough for Ab Initio Structure prediction from an amino acid sequence, molecular dynamics is valuable because it allows for the exploration of large regions of conformational space.

At the same time, the efficiency of molecular dynamics simulations directly depends on the advancement of computer technology; therefore, The Emergence of more powerful processors will inevitably improve this state of affairs.

Currently, molecular dynamics can significantly facilitate the experimental Determination of protein structures using techniques such as X-ray crystallography (where it typically helps) and NMR spectroscopy (where it is invariably helpful). How can molecular dynamics be integrated into the structure determination process? For each conformation, the deviation between the model and actual experimental data can be calculated. In X-ray crystallography, the experimental data consist of the absolute values of the Fourier coefficients of the molecule's electron density. In NMR spectroscopy, experimental data are used to calculate distances between specific amino acid residues. In both cases, however, the experimental data underdetermine the protein structure. To fully resolve the structure, one must find a set of coordinates that minimizes both the deviation from experimental data and the conformational energy of the structure.

Molecular dynamics successfully tackles this task by efficiently sampling conformational space and converging toward the correct structure by minimizing deviations from available experimental data.



Last update: 11/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.