Protein Structure and Function. Application of Bioinformatics Methods - John Rigden 2014
Ab initio protein structure prediction
Energy functions
Rational energy functions
Strictly rational ab initio Methods describe atomic interactions based on the laws of quantum mechanics and the Coulomb potential, utilizing only a few fundamental constants such as the electron charge and Planck's constant. Atoms are represented by atom types, where the number of electrons is the only significant parameter for each type (Hagler et al. 1974; Weiner et al. 1984). However, no serious attempts to employ quantum mechanics-based methods have been undertaken thus far, simply because the computational resources required for such calculations vastly exceed those currently available. In the absence of a quantum mechanical Treatment of interactions, the starting point for ab initio protein modeling essentially becomes The Use of force fields operating with A large number of atom types; the chemical and Physical Properties of atoms for each type closely approximate parameters calculated from crystal structures or quantum mechanical theory (Hagler et al. 1974; Weiner et al. 1984). Well-known Examples of such all-atom rational force fields include AMBER (Weiner et al. 1984; Cornell et al. 1995; Duan and Kollman 1998), CHARMM (Brooks et al. 1983; Neria et al. 1996; MacKerell Jr. et al. 1998), OPLS (Jorgensen and Tirado-Rives 1988; Jorgensen et al. 1996), and GROMOS96 (van Gunsteren et al. 1996). The potentials of these force fields include terms associated with Bond Lengths, valence and torsional angles, Structure/103.html">Van der Waals interactions, and Electrostatic Interactions. The main differences among them lie in the choice of atom types and interaction parameters.
Class="center">Table 1.1. List of ab initio modeling algorithms reviewed in this chapter, along with their Energy Functions, Conformational Search Methods, Model Selection schemes, and typical CPU time per target
Algorithm and server URL |
Force field type |
Search method |
Model selection |
CPU time requirements |
AMBER; CHARMM/ OPLS (Brooks et al. 1983; Weiner et al. 1984; Jorgensen and Tirado-Rives 1988; Duan and Kollman 1998; Zagrovic et al. 2002) |
Rational |
Molecular Dynamics (MD) |
Lowest energy |
Years |
UNRES (Liwo et al. 1999, 2005; Oldziej et al. 2005) |
Rational |
Conformational space annealing (CSA) |
Clustering/ Free energy |
Hours |
ASTRO-FOLD (Klepeis and Floudas 2003; Klepeis et al. 2005) |
Rational |
aBB/CSA/MD |
Lowest energy |
Months |
ROSETTA (Simons et al. 1997; Das et al. 2007) http://www.robetta.org |
Rational-empirical |
Monte Carlo (MC) |
Clustering/ free energy |
Months |
TASSER/Chunk-TASSER (Zhang and Skolnick 2004a; Zhou and Skolnick 2007) http://cssb.biology.gatech.edu/ skolnick/webservice/MetaTASSER |
Empirical |
MC |
Clustering/ free energy |
Hours |
I-TASSER (Wu et al. 2007; Zhang 2007) http://zhang. bioinformatics.ku.edu/ITASSER |
Empirical |
MC |
Clustering/ free energy |
Hours |
To study protein folding, classical force fields have frequently been used in combination with molecular dynamics (MD) simulations. However, in terms of Cell/13.html">Protein Structure Prediction, the results have been less than successful. (For the use of MD in elucidating protein function based on known protein structures, see Chapter 10.) Probably the first significant breakthrough in applying MD to ab initio protein folding was the work of Duan and Kollman in 1997. They simulated the villin headpiece (a 36-residue fragment) in an explicit solvent for 6 months on parallel supercomputers. Although a high-resolution STRUCTURE OF THE final folded protein could not be obtained, the best model generated had a backbone ROOT-mean-square deviation from the native structure within 4.5 Å (Duan and Kollman 1998). Pande and colleagues recently simulated the folding of this small protein using Folding@Home, a distributed global computing system (Zagrovic et al. 2002). The deviation from the native structure was 1.7 Å, with a total simulation time of 300 ms, corresponding to about 1,000 years of CPU time. Despite these substantial efforts, MD simulation using all-atom force fields is by no means a standard method for predicting The structure of medium-sized Proteins (about 100–300 residues). Moreover, a systematic Assessment of the reliability and accuracy of the results obtained has not been performed even for small proteins.
Another potential application of rational force fields in MD simulations is the refinement of protein structure “quality.” The goal here is to shift protein models, starting from low-resolution structures, closer to the native protein structure by improving the local packing of side chains and the main peptide chain. When the initial model is close to the native structure, the required conformational changes are relatively small, meaning that the simulation time will be significantly shorter than that required for ab initio protein folding. One of the first successful examples of protein structure refinement using MD was the GCN4 leucine zipper (a 33-residue dimer) (Nilges and Brunger 1991; Vieth et al. 1994). The disordered low-resolution (2–3 Å) structure of the dimer was first assembled using Monte Carlo (MC) simulation and subsequently refined via MD. By applying helical conformational constraints to the dihedral angles, Skolnick and coworkers (Vieth et al. 1994) successfully obtained a refined GCN4 structure with a main-chain root mean square deviation (RMSD) of less than 1 Å. They used the CHARMM force field (Brooks et al. 1983) and the TIP3P Water model (Jorgensen et al. 1983).
Later, Lee et al. (2001), using AMBER 5.0 (Case et al. 1997) and the TIP3P water model (Jorgensen et al. 1983), attempted to refine 360 low-resolution structural models generated by the ROSETTA program (Simons et al. 1997) for 12 small proteins (fewer than 75 residues). However, they concluded that a systematic improvement in structural quality was not achieved (Lee et al. 2001). Fan and Mark (Fan and Mark 2004) attempted to refine 60 models generated by ROSETTA for 11 small proteins (fewer than 85 residues) using GROMACS 3.0 (Lindahl et al. 2001) and an explicit water model (Berendsen et al. 1981). They reported that 11 of the 60 models showed a 10% improvement in RMSD values, whereas 18 of the 60 models exhibited worse RMSD values following the refinement Procedure. Chen and Brooks (Chen and Brooks 2007) used CHARMM22 (MacKerell Jr. et al. 1998) to refine five CASP6 target structures3 (70–144 residues in size) obtained through comparative modeling. In four cases, RMSD reductions of up to 1 Å were achieved. Their work employed an implicit solvent model based on the generalized Born (GB) approximation (Im et al. 2003), which significantly accelerated the computations. Additionally, spatial restraints present in the initial models were applied during the refinement procedure (Chen and Brooks 2007).
3 CASP — Critical Assessment of Structure Prediction. — Translator's Note.
Notable results were obtained by Summa and Levitt (Summa and Levitt 2007), who used various molecular mechanics (MM) potentials—specifically AMBER99 (Wang et al. 2000; Sonn and Pande 2005), OPLS-AA (Kaminski et al. 2001), GROMOS96 (van Gunsteren et al. 1996), and ENCAD (Levitt et al. 1995)—to refine the structures of 75 proteins using in vacuo energy Minimization. They found that empirical atomic contact potentials outperformed MM potentials: when using the former, the structural models of nearly all tested proteins approached their native states, whereas with the latter (with the exception of AMBER99), the structural models essentially drifted away from the native states. The unsatisfactory performance of the MM potentials may be partly attributed to performing the simulations in vacuo without solvation. These findings demonstrate the potential of combining Empirical potentials and physical force fields for protein structure refinement.
The application of rational potentials and associated MD simulations has not yielded the expected results in protein structure prediction. At the same time, rapid search methods (such as Monte Carlo simulations and genetic algorithms) based on rational potentials have proven promising for both protein structure prediction and quality improvement. One example of these methods is the ongoing project by Scheraga and colleagues (Liwo et al. 1999, 2005; Oldziej et al. 2005), who are developing a rational method for protein structure prediction based solely on the thermodynamic hypothesis. Their method combines a coarse-grained UNRES potential with a global optimization algorithm known as conformational space annealing (Oldziej et al. 2005). In the UNRES potential, each amino acid residue is represented by two interacting connected particles: a Cα atom and the center of the residue's side chain. This effectively reduces the number of atoms tenfold, making it feasible to investigate polypeptide chains exceeding 100 residues. The prediction time in this case can be reduced to 2–10 hours. The UNRES energy function (Liwo et al. 1993) contains a term accounting for THE CONTRIBUTION OF all pairwise interactions between system particles, as well as additional terms such as local energy and correlation energy. Low-energy UNRES models are subsequently converted into all-atom representations using the ECEPP/3 force field (Nemethy et al. 1992). Although many parameters of the energy function are calculated using Quantum Mechanical Methods, some are still derived using distribution and correlation functions from PDB data. Consequently, one might question the strictly non-empirical, or ab initio, Nature of the described approach. Nevertheless, among available ab initio modeling methods, this approach is arguably one of the most reliable (in terms of applying full global optimization to a Rational energy function). Since 1998, it has been systematically applied to investigate numerous CASP targets. The most notable prediction successes with this method were achieved for target T061 from CASP3. For the generated model of a 95-residue α-helical protein, the RMSD from the native structure was 4.2 Å. The accuracy of models obtained for this protein by other methods was significantly lower. For the first time, it was clearly demonstrated that the quality of ab initio target models can surpass that of template-based models. In CASP6, the structural Genomics target TM0487 (T0230, 102 residues) was folded using this method with an accuracy of 7.3 Å. Nevertheless, the extremely small number of models generated exclusively via ab initio modeling methods, along with the improved yet still low accuracy of such models, has resulted in a lack of widespread interest from the scientific community, where accurate protein models are in high demand.
Another example of a rational modeling approach is the multi-stage hierarchical algorithm ASTRO-FOLD, proposed by Floudas and colleagues (Klepeis and Floudas 2003; Klepeis et al. 2005). First, Secondary structure elements (α-helices and β-strands) are predicted based on the calculation of the free energy function of overlapping oligopeptides (typically pentapeptides) and all possible contacts between pairs of hydrophobic residues. Free energy terms reflecting the contributions of Entropy, cavity formation, and the polarization and ionization of each oligopeptide are utilized. The calculated propensities for forming specific secondary structures are then converted into upper and lower bounds for the main-chain dihedral angles, as well as distance restraints imposed on Cα atoms. Subsequently, global minimization within the all-atom ECEPP/3 force field generates the final tertiary structure model of the full-length protein. This approach has been successfully applied to predict the structure of a 102-residue α-helical protein using a double-blind method (although an open community assessment to compare the relative performance of this and other methods was not conducted). The Cα atom RMSD of the predicted model from the experimental structure was 4.94 Å. The global optimization method used in this approach combines the α-branch and bound (αBB) method, conformational space annealing (CSA), and MD simulation (Klepeis and Floudas 2003; Klepeis et al. 2005). The relative performance of this method in determining protein structures remains to be evaluated in the future.
Taylor and colleagues (2008) recently proposed a new approach. Protein structural models are constructed by enumerating possible topologies in a coarse-grained representation, taking into account predefined secondary structure Definitions and physical contact constraints between secondary structure elements. Conformations are scored based on structural compactness and element exposure. The highest-scoring conformations are then selected for further refinement (Jonassen et al. 2006). The authors successfully folded a set of five αβ-sandwich proteins of up to 160 residues, achieving an RMSD of 4–6 Å for the top-ranking model relative to the native structure. Once again, although this method is methodologically intriguing, its performance in open blind experiments on proteins with diverse fold types remains to be determined.
In the latest development of ROSETTA (Bradley et al. 2005; Das et al. 2007), a rational atomic potential is employed In the second stage of Monte Carlo structure refinement, which is preceded by low-resolution fragment assembly (Simons et al. 1997). The features of this method are discussed in the next section.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.