Protein Structure and Function. Application of Bioinformatics Methods - John Rigden 2014

Fold recognition
“Threading”
Empirical potentials

To empirically establish the rules linking protein sequences to their spatial structures, two steps are required: 1) gathering A large number of sequence-Structure pairs, and 2) selecting a set of protein structural properties for analysis. A simple illustration of this approach is The Development of solvation potentials. Any globular protein in its native folded state features a set of residues buried in the (strongly hydrophobic) core and a set of (strongly hydrophilic) residues On the surface facing the surrounding solvent. These are referred to as buried and exposed residues, respectively. It is straightforward to calculate the degree to which a given residue R in a protein of known structure is exposed or buried. One method, albeit a crude one, simply estimates the number of residues located within a certain distance from R (more sophisticated Methods are typically used, e.g., Richmond 1984; Kabsch 1983). Thus, one can compile a list of all residues across all known protein structures along with corresponding data on solvent accessibility (relative to neighbors). Armed with such data, various statistical methods can be applied to uncover correlations between amino acid types and their propensity to reside on the surface or in the interior of a protein. Common approaches include Methods based on Statistical Mechanics or Bayesian statistics (for a comparison with other techniques, see Xia and Levitt 2000). First proposed by Tanaka and Scheraga (1976) and later refined by Sippl (1990), and Miyazawa and Jernigan (1996), these methods are rooted in Boltzmann statistics.

First, it is assumed that the protein structures in the database constitute a statistical ensemble, and that the exposure levels of each residue type across Proteins follow a Boltzmann distribution. Next, the potential of mean force is calculated,

which governs the observed statistical distribution According to the Boltzmann equation. The “energy” associated with a given property p is defined by the equation:

Class="center">Image

where nobs(p) is the observed value of p, and nexp(p) is the “expected” value of p in a reference structure, which is assumed to lack specific interactions or preferences.

Applying this approach typically involves discretizing distances and creating a lookup table for force-field values, without utilizing continuously differentiable molecular mechanics Functions (though exceptions exist). When threading is performed, this lookup table makes it possible to determine the “energy” value for a given structure-sequence combination. Every amino acid model residue will exhibit a certain degree of exposure or burial. Depending on the residue type in the sequence under consideration, probability values—such as finding a valine 30% exposed—can be looked up in the table. The total energy of the model can then be determined simply by summing the energy values across all residues in the model. (Note that summation is permissible thanks to the logarithmic term in the equation.)

A more sophisticated and productive energy function comes into play when considering interacting pairs of Amino Acids. In this case, one can calculate the frequency with which specific amino acid types occur in the vicinity of others—for example, how often a leucine residue is observed within 4 Å of a valine residue. As before, statistical data are aggregated for all possible pairs of the 20 amino acid types. Expected frequencies are then calculated to reflect how often a given amino acid type is observed within a specified distance range.

In reality, typical pairwise interaction potentials widely used in practice are developed at a significantly higher level of detail. For readers well-versed in mathematics, a detailed examination of a widely used pairwise interaction potential is provided below. Others are welcome to skip this part of the section.

Contacts can be classified by distance up to a certain threshold (say, 30 Å) and grouped into respective intervals. These distance intervals can be further subdivided based on the Separation of the contacting residues along the sequence into short-range (e.g., 3 to 9 intervening residues) and long-range (more than 9 intervening residues). Furthermore, even with 50,000 structures in the database, data sparsity can become an issue if such a large number of subdivisions is used. Therefore, a weighting scheme for observations is introduced (the Mijkσ term), which essentially calculates the occurrence rate as if an event had been observed 1/σ times. Now, the energy Ekij for a pair of residues ij separated by k residues and falling within distance interval l is calculated using the formula:

Image

where Mijk is the occurrence frequency for the residue pair ij separated by k residues, σ is the observation weight contribution (typically set to 1/50), and fkij(l) is the relative occurrence frequency of the pair ij separated by k residues within distance interval l:

Image

where is the relative frequency of all pairs separated by k residues within distance interval l:

Image

Here, R represents the number of residue types, and N denotes the number of sequence separation classes. The pairwise potential for a given protein is defined as the sum of energies for residue-residue contacts within the given Separation parameters.

The specifics of potential calculation methodologies can be influenced by various factors. For instance, a force field may be based solely on Cα backbone atom distances, which is often sufficient for preliminary recognition of coarse structural topology. Some researchers increase the number of atomic interaction sites, which generally improves the modeling of Hydrogen Bonds. The application of the Boltzmann equation is not limited to distances alone. Certain studies account for angle-dependent terms, including packing angles between beta strands. Force fields may also weight contributions from residues separated by different distances differently—meaning a researcher might use distinct functions for residues closely spaced in the sequence (i, i + 3) versus those further apart (i, i + n; n > 10), as mentioned previously.

Evidently, the computational efficiency of threading methods is largely constrained by the performance of the energy function. Consequently, numerous past and ongoing studies focus on developing more accurate and, hopefully, more efficient empirical potentials.



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.