Protein Structure and Function: Application of Bioinformatics Methods - John Rigden 2014

Fold Recognition
Alignment Accuracy, Model Quality, and Statistical Significance
Evaluation of Statistical Significance

For the Methods described in this chapter to find Structure/182.html">Practical Application across the broader biological community, reliable error-estimation methods are essential. If a molecular biologist encounters a prediction devoid of a confidence measure, that prediction is practically useless. Whether performing sequence searches, library-based fold searches, or screening against a set of threading models, the output typically takes the form of a scored list. When comparing a sequence against a library of potential models, it is well known that the vast majority of these models will be incorrect. Consequently, most sequence-structure scores can be regarded as Background noise. Statistical metrics can then be applied to determine whether a given score exceeds this background noise, and if so, by how much.

Currently, no universal analytical description exists for the score distribution of threading or Fold Recognition across diverse models and sequences, although it is well established that the distribution of optimal scores is non-normal. For gapped local alignments of two sequences or a sequence against a profile, the distribution of optimal alignment scores can be approximated by an extreme value distribution. Systems such as BLAST, PSI-BLAST, Hidden Markov Models, and numerous sequence-profile and profile-profile methods fit their output score distributions to an extreme value distribution, from which the $p$-value and $E$-value can then be derived.

Some profile-based methods approximate the score distribution using a normal distribution and standardized scores. These values are computed using the mean and standard deviation of the alignment score for the sequence in question against a library of all structural models. Similarly, in many threading methods, the optimal raw score serves as the primary metric for structure-sequence compatibility, and its statistical significance is determined by assuming a normal distribution for the scores of sequences threaded through the library of available models. In the Gibbs sampling threading method (Bryant 1996), The Significance of the optimal score is evaluated by comparison with the distribution of scores obtained by threading a randomized version of the target sequence through the same structural model. The distribution of shuffled scores is assumed to be normal. Recently, many fold recognition systems have abandoned explicit statistical calculations altogether, relying instead on machine learning approaches such as neural networks and support vector machines to predict accuracy scores.

Nevertheless, the most advanced structure prediction systems—those attempting to capture extremely distant Homology relationships—are highly empirical and generally lack robust statistical indicators for probable errors. It is crucial for the reader to recognize that Cell/13.html">Protein Structure Prediction remains an inexact science, and results must therefore be interpreted with caution. In such interpretations, an intimate biological understanding of the Gene or system under study invariably remains the most valuable tool.



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.