Protein Structure and Function. Application of Bioinformatics Methods - John Rigden 2014
Comparative Protein Structure Modeling
Stages of Comparative Protein Structure Modeling
Template Selection
Once a list of templates has been generated using search algorithms, one or more templates suitable for solving a specific modeling task must be selected. Several factors need to be considered during this Selection process.
Class="center">3.2.2.1. General Considerations in Template Selection
The simplest rule for template selection is to choose the Structure whose sequence shows the highest similarity to the target sequence. The protein family to which both the target and the templates belong may include several subfamilies. Constructing a Multiple Sequence Alignment and a Phylogenetic Tree (Felsenstein 1981) can assist in selecting a template from the subfamily most closely related to the target sequence. One should also take into account the similarity between the template's “environment” and the environment in which the target modeling will be carried out. The term “environment” here is used in a broad sense, encompassing everything that is not strictly the protein itself (e.g., solvent, pH, ligands, and quaternary structure interactions). Whenever possible, it is generally advisable to use a template bound to the same or similar ligands as the sequence under study. The quality of experimentally determined structures is another crucial factor in template selection. Resolution and the reliability factor for crystal structures, as well as the number of restraints per residue for NMR structures, serve as indicators of their accuracy. Thus, if two templates exhibit a similar level of sequence similarity to the target, the template determined at a higher resolution should generally be used. Template selection criteria also depend on the purpose for which the comparative model is being built. For instance, when constructing a protein-Ligand model, the presence of a closely related ligand in the template is likely more important than the resolution.
3.2.2.2. Advantages of Using Multiple Templates
It is not strictly necessary to choose a single template for modeling. In fact, the optimal use of multiple templates increases model accuracy (Femandez-Fuentes et al. 2007a, b; Sanchez and Sali 1997; Venclovas and Margelevicius 2005); however, not all modeling programs support multi-template capabilities. The advantage of combining multiple template structures can be twofold. First, template structures that have minimal overlap with one another can be aligned with different domains of the target, allowing the modeling Procedure to construct a Homology model of the entire target sequence. Second, template structures can be aligned with the exact same region of the target, while the model is built using the template that performs best for that specific part of the protein under investigation.
A more sophisticated approach to selecting suitable templates involves generating and evaluating models for each potential template and/or their combinations. The optimized full-atom models can then be assessed using energy or scoring Functions, such as the standardized PROSA score (Sippl 1995) or VERIFY3D (Eisenberg et al. 1997). These scoring Methods are generally quite accurate and make it possible to select the most reliable models from the pool of generated candidates (Wu et al. 2000). Such a trial-and-error approach can be viewed as a constrained threading process (i.e., threading the target sequence through closely related template structures). However, these approaches are only effective for selecting diverse templates at the global level.
In the recently developed M4T method (“Multiple Mapping Method with Multiple Templates”), the selection and combination of multiple template structures are performed via iterative clustering, which takes into account the “unique” contribution of each template, their sequence similarity to one another and to the target sequence, as well as the experimental resolution (Femandez-Fuentes et al. 2007a, b). The resulting models systematically outperform those built using a single best template.
Another important finding emerging from these studies is that when sequence identity falls below 40%, models built using multiple templates are more accurate than those built using a single template. This trend becomes increasingly pronounced as the considered target-template pairs diverge further. At the same time, the advantage of using multiple templates gradually diminishes when the sequence identity between the target and the template reaches 40% or higher (Fig. 3.1). This indicates that within this range, the average differences between the template and target structures are smaller than the average differences among structures of different templates that share a high degree of similarity with the target (Femandez-Fuentes et al. 2007b).
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.