Protein Structure and Function. Application of Bioinformatics Methods - John Rigden 2014
Comparative Protein Structure Modeling
Steps in Comparative Protein Structure Modeling
Model Building
When discussing the model-building stage in comparative structural modeling, it is helpful to distinguish between two parts: template-based modeling and template-free modeling. This division is necessary because certain Regions of the target must be built without using templates. These regions correspond to gaps in the template sequence during target-template alignment. Modeling these areas is commonly referred to as the loop modeling problem. Obviously, these loops account for most of the characteristic differences between the target and the template, and therefore they dictate the structural and, consequently, functional divergences. Unlike loops, the remaining part of the target—in particular, the core with its conserved packing—is constructed using structural information from the template. We will first examine some fundamental approaches from this latter category, namely template-based modeling. This approach also represents the first logical step in model construction, as template-based modeling provides the structural framework for any subsequent loop modeling.
Class="center">3.2.4.1. Template-Based Modeling
Rigid-Body Assembly Modeling
The earliest approach in comparative modeling, and one that is still widely used, involves generating a model from a small number of rigid bodies obtained by aligning with Cell/13.html">Protein Structure templates (Blundell et al. 1987; Browne et al. 1969; Greer 1990). This approach relies on the natural partitioning of a protein structure into conserved core regions, highly variable connecting loops, and side chains framing the protein backbone (Topham et al. 1993). This method is implemented in the widely used COMPOSER program (Sutcliffe et al. 1987). Model accuracy can be somewhat improved by using more than one structural template to build the framework, and by weighting and averaging templates based on the sequence similarity between the template and target sequences (Srinivasan and Blundell 1993).
Segment Matching or Coordinate Reconstruction Modeling
Most existing hexapeptide segments of protein structures can be clustered into only 100 structurally distinct classes using a clustering Procedure (Unger et al. 1989). This fact forms the theoretical foundation of the coordinate transformation method. Structural models of Proteins under study can be built using information on atomic positions in template structures, along with the identification and assembly of short all-atom segments. The atomic positions commonly utilized in this method correspond to those of Ca atoms in segments that are conserved during the alignment of the template structure and target sequence. Such all-atom segments can be obtained either by scanning all known protein structures, including those unrelated to the modeled sequence (Claessens et al. 1989; Holm and Sander 1991), or through a conformational search constrained by an energy function (Bruccoleri and Karplus 1990; van Gelder et al. 1994). For instance, in the general segment-matching modeling method (SEGMOD) (Levitt 1992), the positions of certain atoms (typically Ca atoms) are used to search a representative database of all known protein structures. This method can model the main chain, side chains, and gaps. Certain side-chain modeling Methods (Chinea et al. 1995), as well as a group of loop modeling Methods based on detecting suitable fragments in Databases of known structures (T.A. Jones and Thirup 1986), can be regarded as segment-matching and coordinate-reconstruction techniques.
Satisfaction of Spatial Restraints Modeling
In this group of methods, a set of structural restraints is initially established for the target sequence by aligning the target with related proteins. Conceptually, this procedure is similar to the METHOD FOR DETERMINING protein structures based on restraints derived from NMR experiments. Restraints are typically established by assuming that the corresponding distances between residue pairs in the target and template structures are identical during alignment. These Homology-derived restraints are usually supplemented with stereochemical restraints—Bond Lengths, Bond Angles, dihedral angles, and non-bonded atomic contacts—which are defined using molecular mechanics force fields (Brooks et al. 1983). A model is then built by minimizing the violations of all restraints. These violations can be minimized either through metric geometry methods or via optimization in real space. A particularly elegant metric geometry approach is the construction of all-atom models using lower and upper bounds for distances and dihedral angles (Havel and Snow 1991). Further attempts have been made to apply metric geometry in comparative modeling, such as by Aszodi and Taylor (1996); however, more successful, albeit more conservative, real-space modeling approaches predominate in this field of research. This is likely because evolution has also proved highly conservative in preserving the Structural Features of various proteins (Kihara and Skolnick 2003).
Comparative modeling based on the satisfaction of spatial restraints criterion is implemented in the MODELLER computer program (Fiser and Sali 2003a; Sali and Blundell 1993), which is currently the most popular computer modeling software. In The First stage of model building, restraints derived from the alignment of the target with spatial template structures are imposed on the distances and dihedral angles of the target sequence. The Nature of these restraints is determined through statistical analysis of relationships between close protein structures. By scanning alignment databases, tables are generated that quantitatively reflect various inter-protein relationships, such as correlations between two equivalent Ca-Ca distances or main-chain dihedral angles of two closely related proteins (Sali and Blundell 1993). These relationships are then represented as conditional probability density Functions and can be used directly as spatial restraints. For example, the probability of various main-chain dihedral angle values is calculated taking into account the residue type in question, the main-chain conformation and corresponding template residue, and the degree of sequence similarity between the two proteins. An important feature of this method is that the functional form of the spatial restraints is determined empirically from protein structure alignment databases and contains no subjective user assumptions. Finally, a model is generated by optimizing the objective function in Cartesian space. Optimization is performed using the variable target function method (Braun and Go 1985), which in turn employs conjugate gradient and simulated annealing Molecular Dynamics algorithms (Gore et al. 1986).
Similar principles are incorporated into the NEST software package, which allows for homology modeling based on a single template-sequence alignment or using multiple templates. The method also accommodates different structures for various regions of the target (Petrey et al. 2003).
Combination of Alignments and Combination of Structures
It is often difficult to select the best templates or determine a reliable alignment. In such cases, the quality of a comparative model can be improved by repeatedly iterating the Template Selection, alignment, and model-building process using various model assessment methods. This iteration can be continued until further improvements in Model quality are no longer observed (Fiser et al. 2003; Guenther et al. 1997). More recently, such manual, trial-and-error approaches have been automated (Petrey et al. 2003). For instance, an automated method was introduced in which both the alignment and the tentative model are optimized simultaneously (John and Sali 2003). This task is accomplished using a genetic algorithm protocol. Initially, a series of primary alignments is generated, followed by multiple rounds of re-alignment, model building, and evaluation, with the objective of improving the final model score. During this iterative process, new alignments are generated using several Procedures, such as alignment Mutations and crossovers; comparative models corresponding to these new alignments are then built and evaluated against various criteria, partly dependent on atomic statistical potentials. In another method, a genetic algorithm was applied to automatically generated templates and alignments. A relatively simple structure-dependent scoring function was used to evaluate the resulting trial combinations. Despite certain limitations, this method has been shown to be robust against alignment errors and successfully simplifies the task of template selection (Contreras-Moreira et al. 2003).
Another attempt to optimize the target-template alignment procedure is the Roberta server, where alignments are generated via dynamic programming using a scoring function that combines various protein property data, including recent insights into The Importance of specific sequence regions for protein folding. By systematically varying THE CONTRIBUTION OF different protein attributes to the alignment score, very large ensembles of diverse alignments are produced. Various approaches have been employed to select the best models from the ensemble, including combination of alignments, degree of hydrophobic burial, low- and high-resolution Energy Functions, and combinations thereof (Chivian and Baker 2006).
Within the framework of the meta-server approaches described, not only are models generated by multiple methods evaluated and classified, but they are also further combined. These approaches can likewise be viewed as methods for exploring the alignment and conformational space of a given target sequence (Kolinski and Bujnicki 2005).
Another alternative to complex servers is the M4T program. M4T is characterized by a specific internal scoring function that primarily accounts for the structural environment Properties of the templates. Using this function, the best templates are automatically identified, and alternative alignments are explored and optimally spliced (Fernandez-Fuentes et al. 2007b).
Meta-Servers
Recently, meta-server approaches incorporating multiple existing programs have been developed. In meta-servers, models generated by other methods are selected and either used as input data to build new models or evaluated in a consensus search. For example, in FAMS-ACE (Terashi et al. 2007), input data from other servers serve as starting points for refinement and rebuilding, after which Verify3D (Eisenberg et al. 1997) is used to select the most accurate solutions. Another consensus approach is PCONS, a neural-network-based method where a consensus model is generated by combining information on the accuracy and structural similarities of models produced by other methods (Wallner et al. 2007). 3D-JURY employs the same concept, relying on the consistency of model structure similarities as the basis for selection (Ginalski et al. 2003).
3.2.4.2. Template-Free Modeling: Loop and Insertion Modeling
In comparative modeling, target sequences frequently contain insertions of residues absent from template structures, or regions that structurally differ from their template counterparts. Consequently, information regarding these insertion segments cannot be extracted from template structures. Such regions often correspond to surface loops. Loops typically play a crucial role in defining the functional Specificity of a given protein framework by forming functional sites, such as antibody-complementarity-determining regions (Rudolph et al. 2006), Ligand-binding sites (e.g., for ATP (Saraste et al. 1990), calcium (Grabarek 2006), and NAD(P) (Lesk 1995)), DNA-binding sites (Tainer et al. 1995), and enzyme active sites (e.g., Serine-Threonine Kinases (Johnson et al. 1998) or aspartic proteases (Wlodawer et al. 1989)). The accuracy of loop modeling is a primary factor determining the utility of comparative modeling in Applications such as ligand docking or functional annotation (Fig. 3.2). Loop modeling can be regarded as a miniature protein-folding problem, since the correct conformation of a given polypeptide segment must be calculated primarily from The sequence of the segment itself. However, most loops are small in size and lack sufficient information regarding their local folding. (Exceptions occur when A large number of fragments of known conformation matching the loop sequence are available.) On the other hand, the microenvironment of each loop is uniquely determined by the solvent and the protein framing the loop. For a few rare cases, it has been demonstrated that even decapeptides with identical sequences in different proteins can adopt different Conformations (Fernandez-Fuentes and Fiser 2006; Mezei 1998).

Fig. 3.2. (For the color version of this figure, see the insert.)
Examples of loops (highlighted in yellow) responsible for the functional Specificity of protein superfamilies. From left to right: flavodoxin, immunoglobulin, and neuraminidase, belonging to Protein Families with an a+ß barrel, immunoglobulin, and antiparallel ß-barrel fold, respectively
There are two main groups of loop modeling methods: 1) database search approaches, which search a database of all known protein structures for a segment matching the core anchor region (Chothia and Lesk 1987; Jones and Thirup 1986); and 2) conformational search approaches (Bruccoleri and Karplus 1987; Moult and James 1986; Shenkin et al. 1987). There are also methods that combine these two strategies (de Bakker et al. 2003; Deane and Blundell 2001; van Vlijmen and Karplus 1997).
Fragment-Based Loop Modeling
Modeling based on database scanning and fragment searching is an accurate and efficient method for loop modeling, provided the database contains specific loops belonging to the same class as the modeled loop. For example, this method is effective in modeling ß-hairpins (Sibanda et al. 1989) and loops of specific fold types, such as hypervariable regions in the immunoglobulin fold (Chothia et al. 1989). Previously, it seemed unlikely that structural databank research would ever reach a level of development where fragment-based approaches could become an efficient means of loop modeling (Fidelis et al. 1994). This led to a rapid growth in Conformational Search Methods starting in the 2000s. Nevertheless, over the past decade, many corners of the protein structure universe have been successfully explored, leading to the experimental determination of a vast number of new folds, which in turn has profoundly impacted the number of known structural fragments. Recent studies have demonstrated that loop fragments are not only well-represented in current structural databanks—their shorter segments are likely already exhaustively cataloged (Du et al. 2003).
It has been reported that sequence segments of up to ten residues long have a close counterpart (i.e., at least 50% identical) of a known conformation in the PDB database. Despite a sixfold increase in the number of sequences in data banks and a doubling of the PDB since 2002, not a single sequence segment has been found with less than 50% identity to already known sequences, nor has any unique loop conformation appeared in the PDB. This indicates that a set of already familiar short structural segments continues to circulate in the field of protein sequencing. All sequence segments of 10-12 residues have at least one corresponding structural segment with a minimum of 50% identity, which confirms their structural similarity. A very small group of segments mentioned above is an exception (Fernandez-Fuentes and Fiser 2006). Consequently, new attempts have recently been made to classify loop conformations using broader categories, thereby extending the applicability of the database search approach to a wider range of cases (Fernandez-Fuentes et al. 2006a; Michalsky et al. 2003). One of the latest studies describes the advantages of using HMM sequence profiles in loop Classification and prediction (Espadaler et al. 2004). In another recently published paper, a sequence-based conformation prediction is first performed for the loop of interest, followed by structural alignments of the predicted structural fragments against a relatively small number of structural loop templates. These alignments of the loop sequence against the templates are then quantified using an artificial neural network model trained on a set of predictions with known outcomes (Peng and Yang 2007).
ArchPred is arguably the most accurate database-driven loop modeling method. The approach is briefly described in (Fernandez-Fuentes et al. 2006a, b). ArchPred utilizes a hierarchical multidimensional database designed to classify approximately 300,000 loop fragments and Secondary structure elements flanking the loops. In addition to loop lengths and the types of surrounding secondary structures, the database contains four internal coordinates—a distance and Three types of angles—that describe the geometry of the junction regions (Oliva et al. 1997). During the search, fragments are selected from the library based on length matching, surrounding secondary structure types, and satisfaction of geometric constraints for the junctions. The fragments are then embedded into the backbone of the protein under study, and their goodness-of-fit is evaluated by the RMSD of the junction regions and the number of hard-sphere clashes with surrounding protein atoms. At The final stage, the selected loops are scored using a standardized score that incorporates information on sequence similarity and the probable values of the predicted and observed protein main-chain dihedral angles φ/ψ. Confidence thresholds for the standardized score are established for each loop length to identify fragments whose predictions are of higher quality than those obtained via the ab initio method. The software Implementation of the method is a web server that both performs calculations and regularly updates the fragment library. The predicted segments are returned as output data or, optionally, can be augmented with side chains and then subjected to annealing in the environment of the target protein using conjugate gradient Minimization.
Thus, recent reports indicating that loop conformations are now more comprehensively represented in the PDB suggest that database-driven methods are limited by their ability to recognize matching fragments, and the reason for this lies not in a shortage of segments, as previously believed.
Ab initio loop modeling
To overcome the limitations characteristic of database search methods, conformational search techniques have been developed. There is A wide variety of such methods, utilizing different protein representations, target function terms, and optimization or scoring algorithms. Search strategies include the minimum perturbation method (Fine et al. 1986), molecular dynamics simulation (Bruccoleri and Karplus 1987), genetic algorithms (Ring and Cohen 1993), Monte Carlo and simulated annealing methods (Abagyan and Totrov 1994; Collura et al. 1993), simultaneous multiple-copy search (Zheng et al. 1993), self-consistent field optimization (Kœhl and Delarue 1995), and graph theory-based scoring (Samudrala and Moult 1998). Optimization-based loop prediction can be applied both in the simultaneous modeling of multiple loops and for loops interacting with ligands—neither of which had a straightforward solution in database search methods, which aggregate fragments from unrelated structures in diverse environments.
The MODLOOP module of the MODELLER program employs an optimization-based approach (Fiser et al. 2000; Fiser and Sali 2003b). Loop optimization in MODLOOP is based on conjugate gradients and simulated annealing molecular dynamics. The pseudo-energy function contains multiple terms, including certain terms from the CHARMM-22 molecular mechanics force field (Brooks et al. 1983) and spatial constraints based on distance distributions (Melo and Feytmans 1997; Sippl 1990) and dihedral angles in known protein structures. To address comparative modeling challenges, the loop modeling procedure was optimized with subsequent quality assessment. A large number of loops of known structure were used in both native and only approximately correct environments. The performance of The method was later improved by incorporating the CHARMM molecular mechanics force field and the generalized Born solvation potential (Fiser et al. 2002). The inclusion of solvation potentials in the scoring function has been a key focus of several subsequent works (Das and Meirovitch 2003; de Bakker et al. 2003; DePristo et al. 2003; Forrest and Woolf 2003). Enhanced loop prediction accuracy resulted from incorporating an Entropy-like potential into the scoring function—the "colony energy"—developed through the geometric comparison and clustering of selected loop conformations (Fogolari and Tosatto 2005; Xiang et al. 2002). The continuous refinement of scoring functions contributes to improving the quality of loop modeling methods. Recently, two loop modeling procedures utilizing efficient statistical pairwise potentials encoded in DFIRE have been introduced (Soto et al. 2008; Zhang et al. 2004). Another method has been developed for predicting very long loops using the ROSETTA approach, which essentially performs mini-assembly of loop segments (Rohl et al. 2004). The Prime program implements a procedure for generating a large number of loops based on dihedral angles, followed by iterative cycles of clustering, side-chain optimization, and full energy minimization of the selected loop structures using an all-atom molecular mechanics force field (OPLS) with an implicit solvent model (Jacobson et al. 2004).
3.2.4.3. Model Refinement
Comparative models are built using the best available set of constraints. Such a set typically combines distance and angle constraints derived from various structural templates, molecular mechanics force field potentials, and constraints imposed by various statistical potential functions. Due to the large number of available constraints, the problem can be described as overdetermined. The model-building step is relatively straightforward and primarily aims to resolve constraint conflicts. In the case of MODELLER, this goal is achieved through a combination of conjugate gradient minimization and molecular dynamics. Generating a model typically takes a few minutes. Due to the dominance of template-based constraints, it is often difficult to produce a model whose protein main chain resembles the target more closely than the actual template (assuming the alignment is error-free). Further refinement of the model is also a challenging task, as the most accurate constraints and force field potentials have already been used during model construction. Essentially, the same problem arises as in ab initio modeling, since any refinements at this stage must be carried out without template guidance. Various studies and recent reviews show that most refinements reduce model accuracy (Summa and Levitt 2007). There has been only a single molecular mechanics energy function that improved the initial model, though the improvements were very marginal; statistical potentials have also achieved only a very slight increase in performance.
Other promising model refinement methods attempt to rationally restrict the conformational search space around a high-quality initial model. This can be achieved by simply defining the maximum deviation allowed for protein backbone shifts during Model Selection (Kolinski et al. 2001). Recently, another promising approach has emerged, which defines an evolutionary and vibrational harmonic subspace. This reduced subspace consists of a combination of evolutionarily preferred directions—determined by the principal components of structural variations in a homologous family—and topologically preferred directions derived from the analysis of low-frequency normal modes of vibrational dynamics, spanning up to 50 dimensions. This subspace is sufficiently accurate—allowing most protein cores to be represented with 1Å accuracy—and sufficiently reduced to enable efficient optimization methods, such as replica-exchange Monte Carlo simulation (Han et al. 2008; Qian et al. 2004).
3.2.4.4. Modeling of Proteins and Complexes with Additional Experimental Constraints
Some comparative modeling methods can incorporate constraints derived not from The structure of a homologous template, but from alternative sources. For instance, the basis for constraints can include secondary structure packing rules (Cohen et al. 1989), Hydrophobicity analysis (Aszodi and Taylor 1994) and correlated mutation data (Taylor and Hatrick 1994), Empirical potentials of mean force (Sippl 1995), nuclear magnetic Resonance experiments (Sutcliffe et al. 1992) or chemical cross-linking, spin-labeling, and photoaffinity-labeling experiments (Orr et al. 1998), hydrogen-Deuterium Exchange combined with mass spectrometry (Xiao et al. 2006), hydroxyl radical footprinting (Kiselar et al. 2003), fluorescence spectroscopy, Electron Microscopy image reconstruction (Topf et al. 2008), Site-Directed Mutagenesis (Boissel et al. 1993), and so on. Thus, a comparative model, especially in complex cases, can be improved by achieving consistency between the model and available experimental data as well as more General Principles of protein structure.
In the past, comparative modeling was primarily based on template information and statistical constraints developed from known protein structures and sequences. However, with the advancement of large-scale genetic and proteomic methods, the number of experimentally determined constraints available for automated integration into the modeling process is expected to grow. In addition to generating more accurate models, this will significantly facilitate the modeling of Protein Complexes and assemblies.
A systematic approach to modeling large protein complexes using experimental constraints was developed for modeling the nuclear pore complex—the largest known complex cellular protein, consisting of 456 proteins (Alber et al. 2008). This approach utilizes diverse experimental information. For example, stoichiometry was determined using quantitative immunoblotting; hydrodynamic experiments provided information on the shape and excluded volume of each nucleoporin; immune electron microscopy (EM) aided in the rough localization of nucleoporins; complex composition was determined via affinity purification; cryo-EM analysis and bioinformatics revealed the localization of transmembrane segments; and cross-linking experiments yielded data on direct pairwise interactions. All these input data were integrated into a hierarchical workflow that combined comparative modeling, threading, and rigid- and flexible-docking methods. The ultimate goal of data integration is to translate all available experimental information into spatial constraints that can guide the generalized modeling procedure. The process is flexible, allowing the combination of various representations, resolution levels (e.g., atoms, atomistic protein models, Symmetry elements, or whole ensembles), and optimization procedures (Alber et al. 2007a, b, 2008). This and Similar Methods will jointly leverage the results of genome sequencing experiments, functional Genomics, Proteomics, systems biology, and structural biology.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.