Protein Structure and Function. Application of Bioinformatics Methods - John Rigden 2014

ab initio Protein Structure Prediction
Energy Functions
Combination of Empirical Energy Functions and Fragment Assembly

The empirical potential relies on empirical energy terms4 derived from statistical data on known protein structures available in the PDB database. According to Skolnick (2006), these energy terms can be divided into two groups. The first group includes general energy terms and sequence-independent energy terms, such as Hydrogen bond contributions or peptide backbone rigidity (Zhang et al. 2003). The second group contains Amino Acid Composition- or sequence-dependent energy terms, such as residue-residue pair potentials (Skolnick et al. 1997), distance-dependent atomic interaction potentials (Samudrala and Moult 1998; Lu and Skolnick 2001; Zhou and Zhou 2002; Shen and Sali 2006), and a term reflecting the propensity to form specific secondary structures (Zhang et al. 2003, 2006; Zhang and Skolnick 2005a).

4 In established terminology, a term refers to a component of the energy function. — Translator's Note.

While most empirical force fields account for Secondary Structure propensity, reproducing local Cell/13.html">Protein Structure remains quite challenging in simplified modeling. In other words, diverse protein sequences in nature typically feature either helical or extended structural elements depending on subtle variations in the local and global sequence environment, yet force fields capable of properly reproducing such subtle differences have yet to be developed. One way to circumvent this issue is to directly incorporate secondary structure fragments—obtained from sequence analysis or profile alignment—into the assembly of Spatial Models. An additional advantage of this approach is that utilizing excised secondary structure fragments can significantly reduce Entropy during conformational searches.

This section presents two Protein Structure Prediction Methods based on empirical energy Functions. These methods are among the most successful Ab Initio Protein Structure Prediction approaches (Simons et al. 1997; Zhang and Skolnick 2004a).

One of the most widely known concepts in ab initio modeling was originally proposed by Bowie and Eisenberg. They generated protein models by assembling small fragments (predominantly nonamers) extracted from the PDB database (Bowie and Eisenberg 1994). Building upon a similar concept, Baker and coworkers developed the ROSETTA method (Simons et al. 1997), which proved highly successful in free modeling targets during CASP experiments. This led to fragment-assembly approaches becoming extremely popular within the scientific community. In recent versions of ROSETTA (Bradley et al. 2005; Das et al. 2007), the authors initially generated simplified models whose Conformations were represented by the heavy protein backbone and Cß atoms. In the second stage, a set of selected low-resolution models underwent refinement using an all-atom physical-based energy function that included Van der Waals interactions, pair-wise solvation Free energy, and orientation-dependent hydrogen bonding potentials. The flowchart of this two-stage modeling Procedure is shown in Fig. 1.2, and detailed descriptions of the energy functions can be found in the References (Bradley et al. 2005; Das et al. 2007). The conformational search involves A large number of Monte Carlo energy Minimization cycles (Li and Scheraga 1987). A prime example of applying this two-stage protocol is the blind Ab Initio Structure prediction of target T0281 from CASP6 (70 residues), for which the Ca RMSD from the crystallographic structure was 1.6 Å (Bradley et al. 2005). In CASP7, extensive sampling was performed using distributed computing via Rosetta@home, providing approximately 500,000 CPU hours for each target domain. One of the targets, T0283, was generated via template-based modeling, yet the simulation was performed by ROSETTA using the ab initio protocol. The resulting model achieved an RMSD of 1.8 Å for 92 out of 112 residues (Fig. 1.3, left). Despite significant successes, this procedure is computationally demanding, which hinders its routine application.

The notable successes of the ROSETTA algorithm, combined with the limited availability of its energy functions, prompted several research groups to independently develop energy functions inspired by the ROSETTA concept. Derivatives of ROSETTA include Simfold (Fujitsuka et al. 2006) and Profesy (Lee et al. 2004); their energy functions comprise the following terms: van der Waals interaction potentials, protein backbone dihedral angle potentials, hydrophobic interaction potentials, backbone hydrogen bonding potentials, rotamer potentials, pairwise interaction energy terms, ß-strand pairwise interaction potentials, and a term controlling the protein compactness radius. However, the prediction results obtained using these methods were only partially successful compared to ROSETTA.

Class="center">Image

Fig. 1.2. Flowchart of the ROSETTA program protocol

Another successful free modeling approach is the TASSER program developed by Zhang and Skolnick (2004a), which constructs 3D protein models exclusively using empirical methods. The target sequence is first threaded through a set of representative protein structures during a search for possible

folding patterns. Subsequently, close fragments (longer than 5 residues) are extracted from regions aligned during threading and used to reassemble full-length models. Unaligned regions are built using ab initio modeling methods (Zhang et al. 2003). Protein conformation in TASSER is represented by a set of Ca atoms and side-chain centers of mass. The reassembly process is carried out via parallel Monte Carlo simulations. TASSER energy potentials incorporate information on predicted secondary structure propensities, backbone Hydrogen Bonds, various short- and long-range correlations, and hydrophobic interaction energies based on statistical data from PDB structures. The empirical energy potential contributions are optimized using a large set of structural templates (Zhang et al. 2003), achieving a balance among complex interrelationships between various interaction potentials.

Image

Fig. 1.3. (For color version of this figure, see the insert.) Two Examples of successful free modeling from CASP7. T0283 (left) is a comparative modeling target (from Bacillus halodurans) of 112 residues. The model was built using the all-atom ROSETTA method (a hybrid approach combining physical and empirical methods) (Das et al. 2007) based on free modeling. The TM-score is 0.74 (Zhang and Skolnick 2004b); the RMSD value is 1.8 Å for 92 residues (overall RMSD is 13.8 Å due to incorrect orientation of the C-terminal helix). T0382 (right) is a comparative modeling target (from Rhodopseudomonas palustris CGA009) of 123 residues. The model was built using the I-TASSER method (an exclusively empirical approach) (Zhang 2007). The score is 0.66; RMSD is 3.6 Å. Model and crystallographic structures are shown in blue and red, respectively

Several new versions of TASSER exist. One of them is Chunk-TASSER (Zhou and Skolnick 2007), developed by Skolnick's group. Here, target sequences are first divided into subsequences ("chunks"), each containing three consecutive standard secondary structure elements (helices and/or strands). These subsequences are then folded independently. Finally, spatial restraints derived from the subsequence models are used in subsequent TASSER modeling.

Another version, I-TASSER (Wu et al. 2007), refines the positions of TASSER cluster centers of mass through multiple stages of Monte Carlo simulation. Spatial restraints are established based on models generated in the first TASSER cycle and structural templates identified via template-based modeling alignments using PDB data, which are then applied in the second simulation cycle. The modeling aims to eliminate steric clashes and refine topology. The I-TASSER algorithm flowchart is shown in Fig. 1.4. Although the procedure utilizes structural fragments and spatial restraints from templates obtained via threading, the method frequently successfully builds models with correct topology even when the template topologies composing the model are incorrect. In CASP7, out of 19 free and template-based modeling targets, I-TASSER successfully constructed models with correct topology (3–5 Å) for 7 sequences of up to 155 residues. Fig. 1.3 (right) illustrates example T0382 (123 residues), for which initial templates had incorrect topology (over 9 Å), yet the final model deviated from the X-ray crystallography structure by 3.6 Å. Helles recently conducted a comparative study of 18 ab initio prediction algorithms and concluded that I-TASSER is one of the best methods in terms of modeling accuracy and CPU time expenditure per target (Helles 2008).

Image

Fig. 1.4. (For color version of this figure, see the insert.) Flowchart of the I-TASSER protein structure modeling program



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.