Protein Structure and Function: Application of Bioinformatics Methods - John Rigden 2014

Fold Recognition
Remote Homology Detection Without Alignment
Utilization of Predicted Structural Properties

One of the earliest attempts to elevate Homology recognition beyond simple sequence alignment was undertaken by Bowie et al. (1991). The method is based on the fact that certain structural properties of protein sequences can be predicted in the absence of an exact template. Notably, Secondary Structure—namely, the positions of alpha-helices and beta-strands—can now be predicted with up to 80% accuracy using programs such as PSIPRED (Jones 1999a). Given that structure is more conserved than sequence, a pair of Proteins sharing remote homology will exhibit similar secondary structure patterns even in the absence of any obvious sequence similarity. Furthermore, residue solvent accessibility can be predicted with relatively high accuracy (see, for example, Kim and Park 2003), as can the presence of tight turns such as beta-hairpins (e.g., Kumar et al. 2005).

Class="center">Image

Fig. 2.4. (For the color version of this figure, see the insert.) Schematic representation of The Development of sequence-based Fold Recognition Methods. The left side of each diagram shows the query sequence. The grey box on the right indicates a database of templates with known structures. Arrows point to the Procedure of comparing the query sequence with a specific template: a) Simple comparison of the Amino Acid Sequence with database sequences; b) Alignment incorporating information on the predicted (query protein) and known (template) secondary structure. Wavy lines denote alpha-helices, and diamonds denote beta-strands; c) The query sequence is represented by a profile derived from multiple sequences, such as a PSSM or HMM (colored grid). Each row of the grid corresponds to a homologous sequence, and each Column corresponds to a sequence position; d) The reverse of (c), where the query sequence is searched against a profile library; e) Profile-profile comparison (secondary structure is still represented by a simple three-letter string); f) Same as (e), but with the secondary structure represented as a profile. Each sequence position is associated with specific probabilities for each of the three secondary structure types. Note that in this case, the researcher likely uses the predicted Introduction/11.html">Secondary structure of the templates, despite the actual secondary structure being known. The method has been shown to exhibit high performance (e.g., Bennett-Lovsey et al. 2008)

Such predicted structural properties provide more detailed information on Cell/13.html">Protein Structure, which can subsequently be utilized alongside sequence alignment. When aligning Two Amino Acids from the query and template sequences, compatibility can be calculated based on a mutation matrix such as BLOSUM, alongside terms associated with secondary structure matching and solvent accessibility:

Sij = Seqij + SSij + Solvij,

where Sij is the overall alignment score for residue i in the query sequence and residue j in the template sequence, Seqij is the score obtained for aligning i and j in the BLOSUM matrix, SSij is the score for matching the predicted secondary structure type of residue i with the known secondary structure type of residue j, and Solvij is the score for matching the burial degree of residue i with the known burial degree of residue j. Simple versions of such scoring Functions are presented in Tables 2.16 and 2.1c, where identical states (e.g., matching a helix with a helix) receive a score of +1, and all other combinations receive a score of -1. Often, these functions are elaborated in detail and based on empirical observations of the frequencies with which differing states are found aligned in known homologs. This process is analogous to moving from a simple sequence identity-based scoring matrix to a more sensitive BLOSUM-type matrix.

METABOLISM/2.html">THE CONCEPT OF Combining Sequence and secondary structure information during database searching is schematically illustrated in Fig. 2.4b. Methods based on this principle significantly outperform standard sequence search methods and prove to be orders of magnitude faster computationally than most threading algorithms. This property is crucial when searching large template Databases.

In the Cytology/cytology/16.html">Early stages of the CASP international experiment, threading approaches generally demonstrated the highest performance, closely followed by the hybrid sequence-structure approaches described above. However, threading methods were soon displaced from their leading positions due to two factors: 1) the rapid growth of sequence databases, and 2) the Development of the PSI-BLAST method.



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.