Protein Structure and Function. Application of Bioinformatics Methods - John Rigden 2014
Comparative Protein Structure Modeling
Introduction
Protein Structure Prediction Methods
The investigation of the principles governing the Spatial Structure of natural Proteins can be pursued either through the laws of physics or through The Theory of evolution. Depending on the underlying data utilized for Cell/13.html">Protein Structure prediction, these Methods are generally divided into two main groups (Fiser et al. 2002).
The first group comprises ab initio methods, or template-free modeling approaches, which were discussed in Chapter 1. Structure prediction here is performed solely on The basis of sequence data (Bonneau and Baker 2001; Pillardy et al. 2001). It is assumed that the native protein structure corresponds to the global Free energy minimum attained within the lifespan of the molecule. These methods aim to determine this minimum by exploring a vast conformational space of possible protein structures (Dill and Chan 1997; Sali et al. 1994).
The second group of methods is known as template-based modeling. It includes "threading" techniques, which yield a complete spatial structure description for a target molecule (J. Xu et al. 2007) (see also Chapter 2), as well as comparative modeling (Fiser 2004). This group of methods relies on the sequence similarity shared by the majority of modeled sequences and at least one known structure. Comparative modeling refers to those instances of template-based modeling where not only the fold is identified from a set of available templates, but a full-atom model is also constructed (Marti-Renom et al. 2000). If The structure of at least one protein in a family has been experimentally determined, the structures of other family members can be modeled by sequence alignment against the known structure. Protein structure prediction via comparative modeling is feasible because minor changes in a protein sequence typically lead to minor alterations in its spatial structure (Chothia and Lesk 1986). This predictive capability is further reinforced by the fact that the spatial structures of proteins belonging to the same family are more conserved than their Amino acid sequences (Lesk and Chothia 1980). Thus, if sequence-level similarity between two proteins can be established, structural similarity can usually be inferred as well. Comparative modeling or template-based methods are becoming increasingly widespread due to the relatively limited repertoire of distinct protein folds found in nature (Andreeva et al. 2008; Chothia et al. 2003; Greene et al. 2007).
Both approaches to protein structure prediction have distinct advantages and limitations. In principle, ab initio methods can be applied to model any sequence. However, because protein folding is a complex process and our understanding of it remains incomplete, ab initio methods typically yield low-resolution models. Despite significant progress in ab initio Protein Structure Prediction (R. Das et al. 2007), these techniques are still applicable only to a limited set of sequences roughly up to 100 residues in length. Comparisons of modeling results with reference structures demonstrate that obtaining a complete and accurate representation of the folding pathway for most targets using ab initio techniques remains challenging (Jauch et al. 2007). The advancement of our understanding regarding the accuracy and performance of currently available force fields and sampling techniques has been largely driven by breakthroughs in computational power. To fully harness this potential, several large-scale research projects have recently been launched, which are expected to significantly deepen our understanding of the protein folding process. Notable Examples include Rosetta@home (http://boinc.bakerlab.org/rosetta/), Folding@home (http://folding.stanford.edu/), and the IBM-supported Blue Gene projects.
In the Rosetta@home and Folding@home projects, the investigation of Protein folding and structural modeling is carried out by harnessing the idle computing power of volunteer participants' personal computers, united into a distributed network of a million processors worldwide. To tackle the same research challenges, IBM developed Blue Gene, a supercomputing cluster with a peak performance estimated at 596 teraflops. Currently, various configurations of Blue Gene systems occupy four of the top ten positions on the TOP500 list of the world's most powerful supercomputers as of November 2007 (http://www.research.ibm.com/bluegene/).
In contrast to ab initio methods, comparative protein modeling generates models whose quality is comparable to low-resolution structures obtained by X-ray crystallography or medium-resolution structures derived from NMR spectroscopy. However, the application of comparative modeling is restricted to sequences that can be reliably aligned with known structures. Currently, the probability of finding a closely related protein with a known structure for a sequence randomly selected from a genome ranges from 30% to 80%, depending on the Organism. Approximately 70% of all known sequences contain at least one domain that can be linked to at least one protein of known structure (Pieper et al. 2006). This number exceeds by more than an order of magnitude the total count of experimentally determined protein structures deposited in the PDB (Berman et al. 2007). Comparative modeling methods are increasingly utilized for protein structure determination as the database of experimentally solved structures expands. This trend is further amplified by the Protein Structure Initiative (PSI), which aims to determine at least one representative structure for every protein family (Burley et al. 2008; Vitkup et al. 2001). The initial five-year exploratory phase focused on structural Genomics feasibility and high-throughput pipeline technologies (PSI-1, 2000–2005) was followed by the "production phase" (PSI-2, 2005–2010). It is quite possible that the project's core objectives will be substantially met in less than a decade, thereby enabling comparative modeling to be applied to the vast majority of protein sequences.
As we will see, in practice, template-based modeling invariably incorporates supplementary information that is not strictly derived from the template itself—ranging from general statistical restraints to molecular-mechanical force fields. Driven by improvements in force field accuracy and search algorithms, most successful methods increasingly explore template-independent conformational space (R. Das et al. 2007; Y. Zhang 2007). Similarly, most successful ab initio modeling protocols essentially rely on known structural fragments to build their models (Bystroff and Baker 1998; Zhou et al. 2007). While it is conceptually useful to discuss the two fundamental principles of structural modeling separately, recent methodological trends clearly favor hybrid approaches that combine both paradigms. Ab initio methods can shed light on the dynamics of the protein folding process, whereas effective practical structure modeling almost invariably involves some variation of template-based modeling.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.