Protein Structure and Function. Application of Bioinformatics Methods - John Rigden 2014

Comparative Protein Structure Modeling
Steps in Comparative Protein Structure Modeling
Sequence-to-Structure Alignment

When constructing a model, all comparative modeling programs rely on a predefined list of structural correspondences between the target and template residues. This list is generated during the alignment of the target and template sequences.

Class="center">ImageStructure/structure.files/image016.jpg" width="351"/>

Fig. 3.1. Comparison of model accuracy (y-axis) built for the same set of 765 target protein sequences. Modeling was performed using either a single template (only the result with the best expectation value is shown, dark gray bars) or multiple templates (light gray bars). Sequence identity (x-axis) was calculated based on the result with the highest expectation value and The sequence of the query protein. Error bars indicate the standard error of the mean.

Such alignment is performed in many template search Methods and can sometimes be used directly as input for modeling. However, particularly in complex cases—such as when sequence identity is below 30% (where sequence identity is defined as the number of identical positions in the alignment normalized by the target sequence length)—this initial alignment is often suboptimal. Search methods are typically optimized to detect distant evolutionary relationships, which in practice is often achieved by comparing local motifs rather than performing full-length optimal alignment. Therefore, once templates have been selected, the method should be used to align the templates against the target sequence. Alignment is relatively straightforward when the sequence identity between the target and template is 40% or higher. Below 40% sequence identity, alignment accuracy becomes the primary factor determining the quality of the resulting model. Misalignment of even a single residue introduces an error of approximately 4 Å in the model.

3.2.3.1. Incorporating Structural Information into Sequence Alignment

Alignments in comparative modeling form a unique category because they always incorporate a known 3D structure—the template. Consequently, alignment quality can be improved by leveraging structural information from the template. Specifically, gaps should be avoided within Secondary structure elements, in buried regions, or between residues that are far apart in space. Such criteria are integrated into several alignment methods (Blake and Cohen 2001; Jennings et al. 2001; Shi et al. 2001).

When multiple structural templates are available, they can first be superimposed to generate a multiple structural alignment, which helps identify structurally conserved residues (Al Lazikani et al. 2001; Petrey et al. 2003; Reddy et al. 2001). The next step involves aligning the target sequence against this multiple structural alignment. The advantage of using multiple structures and sequences is that it provides both structural and evolutionary information about the templates, as well as evolutionary information about the target sequence. Utilizing this information frequently yields target-template alignments of higher quality than those obtained from standard pairwise sequence alignment methods (Jaroszewski et al. 2000; Sauder et al. 2000).

The Multiple Mapping Method (MMM) explicitly utilizes Spatial Structure information (Rai and Fiser 2006; Rai et al. 2006). MMM minimizes alignment errors by selecting and optimally assembling fragments generated by various alignment methods. These fragments constitute a set of alternative alignments used as input data. The fragment Selection criterion relies on a scoring function that determines the preferred position of a target sequence fragment within the structural environment of the template. The scoring function incorporates four components to evaluate the compatibility of alternative variable segments within the protein microenvironment: a) environment-specific Substitution Matrices from FUGUE (Shi et al. 2001); b) the BLOSUM amino acid substitution matrix (S. Henikoff and Henikoff 1992); c) the environmental 3D profile scoring function (HSP2) designed to evaluate the agreement between the predicted Introduction/11.html">Secondary structure of the target sequence and the observed secondary structure and solvent accessibility of the template residues (Luthy et al. 1991); and d) a statistical residue-residue pairwise interaction potential (Rykunov and Fiser 2007). Essentially, MMM performs a constrained inverse threading of short fragments: the goal of this threading is not to determine the correct folding pathway, but rather to select from a variety of available mapped alignment alternatives for sequence segments threaded through the same fold. Such local mappings are determined for the remainder of the model, with the alignments serving as the basis for sequential resolution and establishing evaluation boundaries.



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.