Protein Structure and Function: Applications of Bioinformatics Methods - John Rigden 2014

Functional Diversity in Packing Elements and Superfamilies
From Fold to Function
Fold Determination

Class="center">6.2.1.1. General provisions

In general terms, a protein fold (or folding type) refers to the spatial arrangement of the main Secondary Structure elements, taking into account their mutual orientation and topological connections. A major issue directly arising from this general definition is the lack of objective rules for defining which main secondary structure elements must be considered when determining a fold (Grishin 2001).

One of the aims of this chapter is to describe how information on protein relationships, such as data on overall fold types, helps to transfer functional annotations from well-characterized Proteins to those with unknown Functions. As will be discussed further in Section 6.3, transferring annotations between proteins is based on the assumption that evolutionarily related (i.e., homologous) proteins generally share functional properties. However, proteins sharing the same fold are not necessarily homologous. It has been recently debated that different proteins may acquire the same fold independently through convergent evolution, given that the number of physically feasible folds is limited (Russell et al. 1997).

For instance, it remains unclear whether all protein superfamilies characterized by a TIM barrel (ß/a)8 fold are evolutionarily related, as definitive evidence is yet to be found (Nagano et al. 2002).

6.2.1.2. Practical Approaches

Several Databases exist in which Protein Classification provides a comprehensive framework of structural relationships. Below is a practical definition of a protein fold, which is widely used in certain databases. As the definition implies, METABOLISM/2.html">THE CONCEPT OF a fold applies to domains rather than full-length proteins, although domain Definitions may vary across different databases.

The CATH database is a hierarchical classification system for protein domain structures (Orengo et al. 1997; Greene et al. 2007). The highest level of classification assigns a protein domain to one of three classes based on its overall secondary structure content. Within these CATH classes, domains are grouped into architectures that describe the relative spatial arrangement of secondary structure elements without regard to their connectivity. Domains of a given architecture are further classified into topology types depending on how their secondary structure elements are connected to one another. This topological level corresponds most closely to the general definition of a fold mentioned above. In practice, assigning a domain to a specific CATH topology is carried out automatically using the SSAP structural alignment program (Orengo and Taylor 1996) and empirically derived cut-off values.

SCOP (Structural Classification of Proteins), much like CATH, is a hierarchical classification system for protein domain structures (Murzin et al. 1995; Andreeva et al. 2008), although the classification levels in these databases differ. As in CATH, the highest hierarchical level in SCOP is the structural class, but SCOP defines four distinct classes, whereas CATH has three. The next classification level is the fold; two Protein domains share the same fold if they have similar secondary structure elements with the same relative orientation and topological connections. This definition aligns well with the topology level in CATH, but in practice, the assignment of individual domains to specific hierarchical levels can differ between the two databases due to a degree of subjectivity in each definition (specifically, which secondary structure elements are deemed major) and the different protocols used for domain classification (automated in CATH and largely manual in SCOP).

An exceptionally objective METHOD FOR DETERMINING folds is implemented in FSSP (Families of Structurally Similar Proteins) (Holm and Sander 1996b). FSSP performed all-against-all pairwise alignments for a representative, non-redundant set of PDB structures using the DALI structural alignment program (Holm and Sander 1993). The resulting numerical scores from these pairwise structural alignments were then used in hierarchical clustering to construct a so-called protein fold tree. Fold families were determined automatically by cutting the resulting tree at various similarity thresholds.

6.2.1.3. Paradigm Shift

Generally speaking, Cell/13.html">Protein Structure is more evolutionarily conserved than sequence, a principle reflected in numerous shared structural characteristics of proteins. As the number of experimentally determined 3D protein structures grew in the mid-1990s, The Need for structural classification systems capable of extracting meaningful insights from structural data became increasingly urgent. This situation led to The Development of the hierarchical protein structure classification systems described above. Common Structural motifs, such as (ß/a)8 barrels or four-helix bundles, are found in proteins whose sequences share no detectable similarity. Recognizing this fact shaped our current understanding of protein folds. Until recently, protein folds were viewed as recurring structural motifs that arbitrarily partition the protein structure space. This framework implies that the fold space is discrete, meaning that: a) each protein is characterized by a unique fold that groups it with similar proteins while separating it from largely unrelated proteins (despite explanations for analogous folds, see Section 6.2.2.1); and b) each fold possesses distinct structural features, forming an isolated structural group that does not overlap with other groups (Kolodny et al. 2006).

However, as the volume of available structural data continues to grow—driven largely by advances in structural Genomics—our view of protein folds is shifting, with the fold space appearing continuous rather than discrete (Harrison et al. 2002). There is a growing consensus that homologous proteins can adopt different folds (Grishin 2001; Kolodny et al. 2006), and some proteins exhibit multiple flexible folding motifs, with the specific motif at any given time dictated by environmental conditions (Andreeva and Murzin 2006). All of this has important implications when considering folds in functional studies. The main argument for using folds in functional analysis is that proteins sharing similar folds often exhibit remote Homology that cannot be detected by other Methods, and homologous proteins generally perform similar functions (Moult and Melamud 2000). It naturally follows that if the relationship between fold and homology is ambiguous, the relationship between fold and function is likely ambiguous as well. Nevertheless, recent findings based on structural data from CATH indicate that most folds are structurally consistent and markedly distinct from one another (Cuff et al., manuscript in preparation). Indeed, as will be demonstrated later, structural similarity can aid in inferring functional similarity among the proteins under study (Martin et al. 1998).



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.