Protein Structure and Function. Application of Bioinformatics Methods - John Rigden 2014
Comparative Protein Structure Modeling
Steps in Comparative Protein Structure Modeling
Searching for Structures Potentially Related to the Target
Comparative modeling typically begins by searching the Protein Data Bank (PDB) (Berman et al. 2007) for Proteins of known Structure using the target sequence as a query. This search is generally performed by comparing the target sequence against each of the structures in the database.
There are two main classes of protein comparison Methods used in fold assignment. Methods of the first Class compare the target sequence independently with each template in the database, employing pairwise sequence comparison (Apostolico and Giancarlo 1998). These sequence search methods (Pearson 2000; Sauder et al. 2000) and fold assignment approaches (Brenner et al. 1998) have been comprehensively evaluated. The most popular programs in this category are FASTA (Pearson 2000) and BLAST (Schaffer et al. 2001). To increase sensitivity in sequence searches, evolutionary information can be incorporated in the form of Multiple Sequence Alignments (Altschul et al. 1997; Henikoff et al. 2000; Krogh et al. 1994; Marti-Renom et al. 2004; Rychlewski et al. 2000). In such methods, the search is first performed for all available database sequences that share a clear relationship with the target and can be readily aligned. The multiple alignment of these sequences constitutes a target sequence profile, which implicitly contains additional information regarding the Location and pattern of evolutionarily conserved protein residues. The most widely known program of this class is PSI-BLAST (Altschul et al. 1997), which applies a heuristic short-Motif Search algorithm. The next step toward higher method sensitivity is the pre-computation of sequence profiles for all known structures, followed by a pairwise Dynamic Programming Algorithm to compare the two profiles. This approach has been implemented in programs such as COACH (Edgar and Sjolander 2004) and FFAS03 (Jaroszewski et al. 1998, 2005), among others. Profile-based hidden Markov models (HMMs) represent another sensitive method for identifying universal conserved motifs within sequences (Karplus et al. 1998). Significant improvements in HMM-based methods have been achieved by incorporating predicted Secondary structure elements (Karchin et al. 2003; Karplus et al. 2005). Another development utilized in this group of methods is the application of Phylogenetic Tree-based HMMs, where different subsets of sequences are selected for HMM profile analysis at each node of the evolutionary tree (Edgar and Sjolander 2003). Template searches can also be facilitated by identifying intermediate sequences that are homologous to both sequences under consideration (John and Sali 2004; Sauder et al. 2000). These more sensitive fold assignment methods are particularly useful for detecting remote structural relationships when the sequence identity between the target and the template falls below 25%. More accurate Sequence profiles and structural alignments can be constructed using consistency-based methods such as T-Coffee (Moretti et al. 2007), PROMAL (and PROMAL3D for structures) (Pei and Grishin 2007; Pei et al. 2008), ProbCons (Do et al. 2005), and others. Further details regarding Multiple Sequence Alignment methods can be found in recent reviews (Edgar and Batzoglou 2006; Notredame 2007).
The second class of methods is based on the pairwise alignment of a protein sequence and a Cell/13.html">Protein Structure, searching for compatibility between the target sequence and spatial profiles from a database, or “threading” through a library of spatial fold types. This class of methods is also referred to as Fold Recognition, threading, or 3D-1D compatibility matching (Bowie et al. 1991; Finkelstein and Reva 1991; Jaroszewski et al. 1998; Jones 1999; Shi et al. 2001; Sippl 1995). These approaches were discussed in detail in Chapter 2 and are particularly useful when sequence profiling is unfeasible due to a scarcity of known sequences that exhibit a clear relationship with the target or potential templates.
Template search methods “outstrip” the requirements of comparative modeling in the sense that they can detect sequences so evolutionarily distant that building reliable comparative models for them is impossible. This is because establishing sequence relationships often relies on short conserved segments, whereas successful comparative modeling requires an overall correct alignment of the entire modeled protein region. This highlights a crucial distinction between fold recognition and comparative modeling: while both approaches rely on templates and aim to generate a Description of the target's Spatial Structure, fold recognition seeks to determine the overall spatial fold of the target sequence (or at least the class of folds to which it belongs), whereas comparative modeling aims to construct an all-atom model of the target sequence.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.