Protein Structure and Function. Application of Bioinformatics Methods - John Rigden 2014

Comparative Protein Structure Modeling
Applications of Comparative Modeling
Comparative Modeling and the Protein Structure Research Project

The Vision of genome research projects will only be fully realized once we identify and understand the Functions of novel encoded Proteins. This understanding will be greatly advanced by structural information for all, or nearly all, proteins. A wealth of structural data will be generated through structural Genomics (Burley et al. 2008; Chance et al. 2002)—the large-scale Determination of protein structures using X-ray crystallography and nuclear magnetic Resonance spectroscopy, effectively combined with accurate, automated, and high-throughput Structure/53.html">Comparative Cell/13.html">Protein Structure Modeling. Given the throughput of modern modeling techniques, it is reasonable to expect experimental outcomes to yield models based on at least 30% sequence identity (Vitkup et al. 2001), corresponding to one experimentally determined structure per sequence family rather than per fold family.

To make the large-scale comparative modeling required for structural genomics feasible, the steps of comparative modeling are integrated into fully automated pipelines, such as the SWISS-MODEL or MODBASE repositories (Kopp and Schwede 2006; Pieper et al. 2006), each containing over a million models. Statistical data in these Databases indicate that models can be generated for the constituent domains of approximately 70% of known protein sequences. This is supported by the fact that nearly 2,000 structures have been deposited in databases by structural genomics centers, whose primary research focus is on novel folds and novel structures. These additions accounted for 73% of the new structural features in the PDB over the past 7 years (Burley et al. 2008).

While the number of proteins today for which at least partial models are available looks impressive, the model is typically built for only a single protein domain. On average, however, proteins comprise two or three domains. For example, the average size of a Yeast Open Reading Frame is 472 amino acid residues, whereas the average domain size in CATH, a database of Protein domains, is 175 residues. The average model size in MODBASE, a database of comparative models, is only slightly higher at 192 residues. Furthermore, two-thirds of the modeling cases are characterized by less than 30% sequence identity between the target and the closest template.



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.