Protein Structure and Function. Application of Bioinformatics Methods - John Rigden 2014

Integrated Servers for Function Prediction from Structure
Introduction
The Problem of Function Prediction from Structure

Why is this the case? First, if a protein has an unknown function, it not only means that no experimental data are available regarding its role, but also that standard sequence analysis Methods for functional annotation have failed. These methods—particularly various profiling techniques, such as Hidden Markov Models—have grown quite sophisticated in recent years and are now capable of detecting functional similarities even at very low sequence identity levels. Therefore, if these methods too have failed, we have no choice but to rely exclusively on three-dimensional Structure.

Protein structures contain various clues to their function, which vary in reliability as described in the previous chapters. Chapter 6 showed that, at a global level, a protein fold type can very often provide clues to its function, since certain fold types are tightly linked to specific Functions. Consequently, the first step in determining function from structure is invariably to search for a protein of known function with a similar fold. This can be accomplished using numerous web servers designed for fold comparison, for which several comparative reviews have been published (Sierk and Pearson 2004; Novotny et al. 2004; Carugo 2006). However, keep in mind that structural similarity does not necessarily imply functional similarity. For example, so-called superfolds (Orengo et al. 1994; see also Chapter 6), such as the TIM-barrel family, can encompass representatives with a vast diversity of functions (Nagano et al. 1999; Anantharaman et al. 2003). Conversely, if a protein adopts a novel fold—which some research groups consider a successful outcome—no similar folds will be found at all.

Moving beyond the global picture, important functional clues may reside on the protein surface, particularly in its clefts and pockets (Chapter 7), which can provide the specific local arrangement of residues required for catalysis, DNA recognition, and so forth (Chapter 8). Thus, you might be able to identify, say, a putative ATP-binding site. This constitutes a vital clue to function, but the story does not end there.

Various other complications tend to throw a wrench in the works. First, obtaining the native structure of an entire protein is frequently challenging. In such cases, one may only secure The structure of a portion of the protein—such as a single domain. By itself, this domain may reveal little about the function of the whole protein. Second, even if the STRUCTURE OF THE entire protein is solved, it may represent merely a single component of a multi-protein complex. Once again, the structure tells only part of the story. Even more unruly are so-called moonlighting Proteins, which can actually perform multiple functions depending on the context: cellular localization, environment, and so on (Jeffery 1999). Furthermore, some proteins can alter their function depending on which Alternative Splicing isoform is currently expressed (Stamm et al. 2005).

Another challenge in function prediction lies in the difficulty of assessing the success or failure of a given prediction method, and indeed, even in defining what is meant by 'function.' Function can be described at various levels, ranging from biochemical activity to biological processes and pathways, all the way to the level of Organs or organisms (Shrager 2003). Consequently, a specific protein may be annotated at several different levels of functional Specificity: for instance, a ubiquitin-like domain, a signaling protein, a predicted Serine hydrolase, a putative eukaryotic D-aminoacyl-tRNA deacylase, and so forth. As a result, it is difficult to judge the accuracy of any such description, especially when it involves even more ambiguous terms.

The standard strategy for evaluating function prediction methods is to use the Gene Ontology (GO) (The Gene Ontology Consortium 2000; Camon et al. 2004). This is an open Classification system for the functional annotation of protein sequences. It provides a machine-readable ontology based on a controlled vocabulary of functional descriptors, and many function prediction methods report their results using GO classifications. Although not strictly hierarchical, functional GO descriptors range from entirely non-specific (e.g., enzyme) to highly specific (e.g., 1-pyrroline-4-hydroxy-2-carboxylate deaminase).



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.