Protein Structure and Function: Application of Bioinformatics Methods - John Rigden 2014
Functional Diversity in Packing Motifs and Superfamilies
From Folding to Function
Relationship Between Folding Patterns and Function Prediction
This section discusses The properties of Functions whose existence can be hypothesized without relying on Homology data—that is, functions that have arisen through convergent evolution. Issues related to The Study of functions established on The basis of homology information are covered in Section 6.3 of this chapter.
In general, determining a protein's Structure and folding pattern provides researchers with the opportunity to employ a variety of structure-based function prediction Methods that would be inaccessible if the Cell/13.html">Protein Structure were unknown. Some of these methods are based on the principle that structural information allows for the identification of remote homologies not evident at the sequence level (Lee et al. 2007). Other methods use structural information simply under the assumption that it is functionally relevant to the protein's molecular role, disregarding the evolutionary context of the structural properties. Many of these methods are described in other chapters of the book (see Chapters 7, 8, 10, and 11). This chapter discusses only those cases that are directly relevant to fold information.
Class="center">6.2.2.1. Single-Function Folds
A newly determined protein structure can be used to search for similar folding patterns among already known structures. This involves structure comparison programs that typically evaluate The Significance of detected structural similarities using specific scoring schemes. Some of these programs are publicly available and have recently been benchmarked using a large dataset of known structural similarities compiled by CATH (Kolodny et al. 2005; Redfern et al. 2007). Such programs include DALI (Holm and Sander 1996a), FATCAT (Ye and Godzik 2004), SSM (Krissinel and Henrick 2004), CE (Shindyalov and Bourne 1998), and CATHEDRAL (Redfern et al. 2007). When a new structure is found in a protein of unknown function, the next step in the investigation is to assess whether functional annotation can be transferred based on this structural similarity.
Certain folds are restricted exclusively to homologous Proteins, whereas others can be found in proteins that have evolved convergently—referred to as homologous and analogous folds, respectively (Moult and Melamud 2000). Similarly, some folds are functionally uniform, whereas others characterize proteins with highly diverse functions. It is generally accepted that homologous folds are more functionally uniform than analogous ones (Moult and Melamud 2000). Obviously, if a fold is associated with a unique function X, discovering this fold in a protein of unknown function will directly lead to the annotation of function X. In practice, however, the situation is somewhat more complex, as functionally diverse folds may be erroneously identified as being associated with a single function due to sampling bias.
Regardless, there are documented cases where determining the fold has successfully aided in predicting protein function (Moult and Melamud 2000). For example, the Spatial Structure of the uraC Gene product from Escherichia coli contains a fold closely resembling that of the amidohydrolase family of proteins. Subsequent studies revealed that this protein possesses a catalytic apparatus similar to that found in proteins sharing the same fold (Colovos et al. 1998; Moult and Melamud 2000). There is a growing number of successful protein function predictions driven by Fold Determination via structural Genomics initiatives, the overarching goal of which is to determine the structures of as many proteins as possible (Adams et al. 2007). In most cases, however, a successful function prediction is not the result of simple fold determination alone, but rather stems from combining the fold with some other property, such as a sequence recognition motif or the similarity of functional sites among the studied proteins.
6.2.2.2. Super-sites
Overall, spatial structure data are extremely useful for defining a functional site—a subset of residues that are critical for carrying out a protein's molecular function. Functional sites are primarily represented by binding sites (the set of protein residues that interact with ligands) (Dessailly et al. 2008) or catalytic sites (the set of residues directly involved in the enzymatic reaction) (Porter et al. 2004).
One reason structures are so useful for identifying functional sites is that the latter tend to reside within the most conserved topological Regions of the structures. Moreover, even when there is no compelling evidence of homology between proteins sharing a common fold, their functional sites are often located in the exact same regions of their three-dimensional structures. Such functional sites are termed super-sites. It has been shown that super-sites are abundant in analogous packing (or super-packing, see Section 6.2.3.1)—that is, packing shared by non-homologous proteins (Russell et al. 1998). Fig. 6.1 illustrates a well-known example of a super-site: the catalytic site of proteins characterized by the (ß/a)8 barrel fold. The catalytic residues are always positioned at the C-termini of the ß-strands forming the central parallel ß-sheet, even though the structures of the ß-strands themselves may vary (Nagano et al. 2002).
6.2.2.3. Superfolds
A fold that is widespread across many diverse superfamilies and exhibits pronounced functional diversity is called a superfold (Orengo et al. 1994). Superfold elements are constituents of proteins that perform a vast array of different functions. Striking Examples of superfolds include the TIM barrel (ß/a)8 fold, which is found in representatives of over 25 different superfamilies (Nagano et al. 2002), and the Rossmann fold, which can be found in proteins from 114 CATH superfamilies (CATH v3.1), many of which are functionally diverse. Superfolds account for a very small fraction of known folds, yet they appear to be the products of a large portion of known genomes (Lee et al. 2005). Furthermore, superfolds pose one of the major challenges in structure-based function prediction because proteins sharing similar superfold elements do not necessarily share similar functions. The existence of such folds and their wide distribution among protein molecules compel researchers to exercise caution when using relationships between known folds to predict Protein Functions.

Fig. 6.1. (For the color version of this figure, see the insert.) Super-sites of the TIM barrel (ß/a)8 fold. Schematic representations of four proteins possessing the (ß/a)8 barrel fold from different CATH (and SCOP) superfamilies: (a) dihydropteroate synthase from E. coli (CATH domain ID: 1aj0A00); (b) Tryptophan synthase alpha-subunit from P. furiosus (CATH domain ID: 1geqB00); (c) endo-1,4-beta-xylanase Z from C. thermocellum (CATH domain ID: 1xyzA00); (d) aldose reductase from H. sapiens (CATH domain ID: 2a1rA00). Structures were superimposed using CORA (Orengo 1999). The molecules are oriented identically. Common elements of the four structures are shown in red. The positions of catalytic residues (According to the Catalytic Site Atlas) are shown in green. Despite significant structural differences and the lack of evidence for homology between these proteins, the catalytic centers are invariably located near the C-terminus of the central ß-strands. Three-dimensional structure images were generated using Molscript (Kraulis 1991), and image rendering was performed using Raster3D (Merritt and Bacon 1997).
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.