Protein Structure and Function. Application of Bioinformatics Methods - John Rigden 2014

Predicting Protein Function from Theoretical Models
Accuracy and Added Value of Model-Based Predictions
Implementation

Understanding the added value of specific structural or Structure/108.html">Surface Properties of models raises the question of whether they are useful for function prediction. In 1998, Fetrow and Skolnick proposed a multi-step Procedure to identify functional protein sites in low- and medium-resolution models (Fetrow and Skolnick 1998). Based on geometry, residue identity, Ca-atom distance, and conformation, active-site residues were transformed into a three-dimensional descriptor termed the Fuzzy Functional Form (FFF). These forms were used to analyze a set of Spatial Models in order to select those containing similar Spatial Motifs. The applicability of The method was validated by identifying new members of the glutaredoxin/thioredoxin disulfide protein family in the genomes of Yeast (Fetrow and Skolnick 1998) and E. coli (Fetrow et al. 1998), whose Functions could not be determined previously by sequence comparison. A major achievement of the FFF method and its analogs is that they made it possible to distinguish pairs of Proteins with similar active sites from those that may share similar folds but not necessarily similar active centers.

A further Development of the FFF technique was the active-site profile method (Cammer et al. 2003), which was successfully combined with experimental techniques to discover novel Serine Hydrolases in yeast (Baxter et al. 2004). A key feature of this approach is that the focus was placed not on residues conserved across the entire family, but rather on key functional residues specifically identified among all proteins with a given function, regardless of their sequence similarity. Thus, the method can be applied to identify and annotate various functional centers, including catalytic centers, regulatory centers, and cofactor-binding sites.

It is also worth mentioning a hybrid approach that combines protein surface analysis with evolutionary Methods, proposed by Pawłowski and Godzik (Pawłowski and Godzik 2001). They generated maps of the protein molecular surface by projecting the distribution of various properties (such as charge or Hydrophobicity) onto a sphere approximating the protein surface. This method allows entire protein molecules to be compared and Conclusions to be drawn about their global functional similarity, for instance, based on a numerical similarity measure between their maps. It was shown that comparing such surface maps improves protein function prediction compared to conventional sequence analysis methods and is capable of reproducing known Examples of functional divergence within diverse protein groups, including the identification of unexpected sets of shared functional properties among seemingly distant paralogs. This method, which now features a web interface (Sasin et al. 2007), has proven to be quite robust and allows The Use of Homology models instead of experimental structures.

Other studies have addressed the question of whether more specific function predictions can be made as accurately for models as for experimental structures. The results of the MetSite method, which integrates sequence and structural information for metal-binding sites, proved encouraging (Sodhi et al. 2004). Although performance on models was lower than on experimental structures, correct predictions of metal-binding sites were achieved for about half of the reliable models generated using mGenTHREADER. Notably, these models contained only backbone atoms, so errors in side-chain positioning did not affect performance. A similar prediction method for DNA-binding propensity was also developed for both experimental structures and models, utilizing protein sequence information, spatial Asymmetry in the distribution of certain residues, and dipole moments (Szilagyi and Skolnick 2006). This method is likewise designed for structures containing only Ca atoms. When analyzing models with an RMSD of less than 6 Å from the native structure, the performance of this method was only slightly lower than that with experimental structures. Consequently, the method can be applied to models of any origin, including ab initio models and those obtained by Fold Recognition methods, for which, however, lower accuracy should be expected.

One of the important Practical Applications OF protein models is the virtual screening of small-molecule compound libraries to find suitable inhibitors for detailed development and lead generation (Jacobson and Sali 2004). Since our book focuses on Protein Functions, such applications will not be discussed in this chapter. Nevertheless, small-molecule docking, identical to that used in pharmaceutical pipelines, is beginning to be employed for protein function prediction. As discussed in Chapter 8, this approach assumes that the compounds ranking best among candidates are presumably the true ligands (Hermann et al. 2007; Song et al. 2007). In this context, mention should be made of studies assessing the suitability of protein models for small-molecule docking compared to experimental structures. Here are two studies, each concluding that models are viable, yet using different criteria for comparison with experimental structures. McGovern and Shoichet (McGovern and Shoichet 2003) compared the enrichment of known ligands versus decoys in docking solutions for the holo- and apo-forms of nine Enzymes and their models. The templates chosen for modeling shared 34–87% sequence identity with the targets overall and up to 45–100% in the Active Site. The best enrichment was achieved for eight holo-forms, two apo-forms, and three models, thereby confirming the superiority of experimental structures. However, almost all models performed better than random Selection of active compounds. Models built on closer templates generally proved more effective, though minor conformational distortions in the active site could impair this trend in individual cases. Later, Oshiro et al. (Oshiro et al. 2004) compared the enrichment of known active compounds in docking results for several experimental structures and models of CDK2 and Factor VIIa. The templates selected for modeling had sequence identities in the active-site neighborhood of 37–77%. Notably, the performance of the models was comparable to that of experimental structures when the identity was above 50%, and dropped significantly otherwise. Summarizing the results of these two studies, one can conclude that using models for docking is justified only when an experimental structure is unavailable. It would be tempting to evaluate the performance of models in function prediction directly from docking results such as those described above.



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.