Protein Structure and Function. Application of Bioinformatics Methods - John Rigden 2014

Fold Recognition
Alignment Accuracy, Model Quality, and Statistical Significance
Alignment Generation Algorithms and Scoring

As already demonstrated, incorporating evolutionary information about Proteins—such as profiles or Hidden Markov Models—along with predicted Secondary Structure data enhances Homology detection, which typically leads to a corresponding improvement in alignment accuracy (Elofsson 2002). Zachariah et al.

(Zachariah et al. 2005) showed that when generating alignments using dynamic programming, employing a more accurate model for gap opening and extension does not improve homology detection, yet it significantly increases alignment accuracy.

A successful approach was proposed at the recent CASP meeting (Venclovas and Margelevicius 2005). In this Procedure, a series of sequences bridging the sequence space between the target sequence and the template(s) are used to initiate additional PSI-BLAST searches against a non-redundant sequence database. The resulting alignments of the target sequence against the template(s) are then extracted and subjected to consistency analysis. For regions where a single predominant alignment variant emerges, the alignment is considered reliable, whereas regions lacking consistency between the target sequence and the template are classified as unreliable. Thus, alignment accuracy can be improved by searching for consistent alignments. This concept is closely related to the idea used in 3D-Jury, which searches for a consensus solution across structural space. Prasad et al. (2004) applied a similar approach, using five different Methods to generate alignments and identify a consensus among them.

Tress and colleagues (2003) investigated the distribution of residue-residue profile scores along the alignment length. They found that accurately aligned regions can be reliably identified by the presence of adjacent segments with high scoring function values for the residues.

As mentioned earlier, the dynamic programming or HMM approach guarantees the construction of an “optimal” alignment for a given scoring function. However, scoring Functions are far from perfect. Consequently, there may exist A large number of suboptimal alignments with fairly high scores that may actually be structurally more accurate. Similarly, alignment algorithms require specific parameterization for insertion and deletion probabilities, yet these parameters cannot be universal for all proteins. To address this, Yarovinsky et al. (2002) conducted a systematic study of near-optimal alignments by varying alignment parameters and relaxing the strict constraints of the dynamic programming matrix. They found that a constrained search “near” the optimal alignment can reveal alignments that are substantially more accurate than the formally “optimal” one in terms of the scoring function value. As a result, the question of how to reliably select such alignment improvements from a vast pool of alternatives remained open.

Chivian and Baker (2006) attempted to solve this problem by building models based on alignments and evaluating each model using a combination of structural clustering (e.g., via 3D-Jury) and a fine-tuned protein 3D energy function. In addition, Wallner and Elofsson (2006) trained a neural network on residue environments and profile-profile scores from a set of protein models to develop a Model quality prediction algorithm. Finally, McGuffin (2008) utilized several model quality assessment programs alongside structural clustering methods, such as 3D-Jury, as inputs for a neural network-based predictor.



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.