Protein Structure and Function. Application of Bioinformatics Methods - John Rigden 2014

Comparative Protein Structure Modeling
Efficiency of Comparative Modeling Methods
Errors in Comparative Models

The overall accuracy of comparative models varies widely. At the lower end are low-resolution models, for which the only reliably determined feature is their backbone fold. At the upper end are models whose accuracy is comparable to the medium resolution of crystallographic structures (Baker and Sali 2001). When addressing biological problems, even low-resolution models are often useful, as Functions can be predicted based on coarse structural properties, as shown in subsequent chapters of this book.

Errors encountered in comparative models can be divided into five categories: (1) errors in side-chain packing; (2) distortions or shifts in regions with a correctly aligned template Structure; (3) distortions or shifts in regions that lack corresponding segments in any of the structural templates; (4) distortions or shifts in regions with an incorrectly aligned template structure; and (5) overall structural mismodeling resulting from The Use of an inappropriate template. Addressing each of these challenges requires substantial methodological improvements.

Errors 3 through 5 are relatively rare when modeling sequences with greater than 40% sequence identity to the template. Thus, in such cases, the RMSD for 90% of the main-chain atoms is likely to be about 1 Å (Sanchez and Sali 1998). Within this range of structural similarity, generating a structural alignment is straightforward: gaps are few, and structural differences between Proteins are typically limited to loops and side chains. When sequence identity is 30–40%, structural differences become more pronounced, alignment gaps are more frequent and larger, and incorrect alignments and insertions into the target sequence emerge as major problems. As a result, the main-chain RMSD increases to roughly 1.5 Å for about 80% of the residues. Modeling of the remaining residues is subject to larger errors, as current Methods generally fail to model structural distortions and rigid-body shifts, or to overcome The problem of using incorrect alignments. When sequence identity drops below 30%, the primary challenge lies in identifying close templates and aligning them with the target sequence. Overall, at this level of sequence similarity, one can expect roughly 20% of the residues to be misaligned, leading to flawed modeling with errors exceeding 3 Å. Such misalignments pose a serious obstacle to comparative modeling because, as it turns out, the majority of structurally related protein pairs share less than 30% sequence identity (Rost 1999).

To put the errors in comparative models into perspective, let us consider the structural variations observed experimentally for the same protein. A main-chain atomic accuracy of 1 Å is typical for a low-resolution X-ray structure—around 2.5 Å—with a reliability factor of about 25% (Ohlendorf 1994), as well as for medium-resolution structures obtained by NMR with 10 interproton distance restraints per residue (Fig. 3.4). Similarly, differences between high-resolution X-ray and NMR structures of the same protein generally amount to about 1 Å (Clore et al. 1993). Environmental changes (e.g., oligomeric state, crystal packing, solvent, ligands) can also significantly affect the structure (Faber and Matthews 1990). In general, comparative modeling of sequences with over 40% identity to the template yields models of nearly the same quality as medium-resolution experimental structures, simply because proteins with this level of similarity are highly likely to resemble each other just as much as different experimental determinations of the same protein under various conditions do. Nevertheless, a potential risk in comparative modeling is that certain regions, mainly loops and side chains, may contain more substantial errors.

The performance of comparative modeling methods can sometimes be overstated. This is because the metrics commonly discussed in the literature are backbone deviation values. Nevertheless, isolated errors in specific residues critical for protein function—even with an overall main-chain RMSD of less than 1 Å—can still be large enough to prevent reliable Conclusions regarding the MECHANISM OF ACTION, protein function, and drug design.

Class="center">Image

Fig. 3.4. (For the color version of this figure, see the insert.) Demonstration of the accuracy of structural models obtained by various experimental and computational Methods for the same allergen protein Der p 2. (a) Superposition of ten different NMR structures of Der p 2 (PDB code 1A9V); average RMSD = 0.97 Å. (b) Superposition of X-ray crystallographic structures of two Der p 2 isoforms sharing 87% sequence identity: 2F08 (resolution 2.20 Å) and 1KTJ (resolution 2.15 Å); RMSD = 1.33 Å. (c) Superposition of the NMR and X-ray structures of Der p 2 (1A9V and 1KTJ); RMSD = 2.2 Å. (d) Superposition of the comparative model built for the 1NEP protein using 1KTJ as a template and the 1NEP X-ray structure. The 1NEP and 1KTJ structures share 28% sequence identity and represent a typical challenging modeling target; RMSD = 1.66 Å. All RMSD values refer to the superposition of Ca atoms.



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.