Protein Structure and Function: Application of Bioinformatics Methods - John Rigden 2014
Integrated Servers for Structure-Based Function Prediction
Introduction
Structure-Function Prediction Methods
As shown in the previous chapters, a vast array of different Methods exist for predicting protein function based on Structure. These methods and their suitability are reviewed in several publications (Kim et al. 2003; Watson et al. 2005; Rigden 2006). None of these methods is flawless, nor can any of them be expected to succeed in every single case. For example, some approaches are tailored exclusively for Enzymes and are completely ineffective if the protein in question is not an enzyme. Other methods rely heavily on finding a match—whether in fold, motif, binding site, and so forth—with a protein of known structure. Consequently, if such a match cannot be found, or if it merely aligns with another hypothetical protein, the efficacy of the method drops to zero.
Consequently, a sensible approach is to apply a large battery of these methods to a Cell/13.html">Protein Structure and see what turns up. This is precisely the strategy employed by the two servers discussed in this chapter: ProKnow from the University of California, Los Angeles (UCLA) (http://proknow.mbi.ucla.edu) and ProFunc from the European Bioinformatics Institute (EBI) (http://www.ebi.ac.uk/profunc). Both servers leverage predictions based on both protein sequence and structure and are heavily automated: you simply upload a PDB file and patiently wait for the results.
Class="center">
Fig. 10.1. Schematic diagram of structure- and sequence-based methods applied to a 3D protein structure uploaded to the ProKnow function prediction server. Sequence-based methods include PSI-BLAST (Altschul et al. 1997) and PROSITE (Hulo et al. 2004). Structure-based methods include Dali database searches (Holm and Sander 1998) and RIGOR spatial motif searches (Kleywegt 1999). To identify interesting functional links, the latter two methods incorporate results from the DIP (Database of Interacting Proteins, Xenarios et al. 2002) and Prolinks (Bowers et al. 2004) Databases to enhance PSI-BLAST output. All results are then synthesized to generate Gene Ontology (GO) functional annotations, which are combined using Bayesian weighting to yield a set of predicted Functions along with corresponding confidence scores.
To illustrate these two methods in action, we use a recently solved 3D structure as an example. This is The structure of a putative acetyltransferase from Vibrio cholerae, determined in 2005 by the Midwest Center for Structural Genomics (MCSG). It was deposited in the PDB on February 28, 2006, under the accession code 2fck (Cuff et al. 2007). At the time of its release, the protein's function was only tentatively known; its sequence shared over 50% identity with ribosomal-protein-Serine acetyltransferase and contained several motifs characteristic of acetyltransferase activity. Once the structure became available, these preliminary functional annotations received strong support, as robust global and local structural similarities emerged with other—more distant—acetyltransferases. The strongest similarities were detected in the putative binding site for coenzyme A (CoA). Some of these similarities are discussed below.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.