Protein Structure and Function. Application of Bioinformatics Methods - John Rigden 2014
Functional diversity in packing elements and superfamilies
Functional diversity of homologous proteins
Definitions
In general, detecting Homology (relationships between superfamilies) is far more useful for function prediction than merely identifying structural similarity (relationships between fold types). This section explores the relationship between structural homology and functional diversity, demonstrating that even when homology is clearly established, significant challenges remain in transferring functional annotations from one protein to another.
Before explaining how function diverges within superfamilies, it is necessary to clearly define the term superfamily and outline how it is applied in practice. The term family, which is also used throughout this section, is introduced below as well.
Class="center">6.3.1.1. General Concepts
A superfamily is a group of Proteins considered to be evolutionarily related. Relationships between proteins within a superfamily can be established through sequence similarity—determined either by traditional sequence alignment Methods or by more sensitive searches using HMMs (Reid et al. 2007). In the absence of sequence similarity, structural analysis can also reveal remote homology and/or functional resemblance. However, unlike sequence similarity, there are no universally accepted metrics for assessing the statistical significance of structural or functional similarity. Consequently, the thresholds used to define superfamily relationships can be arbitrary and somewhat subjective. Today, several Databases, such as CATH and SCOP, have established standard and widely accepted definitions for superfamilies (see Section 6.3.2.1). Nevertheless, a certain degree of subjectivity persists across all these databases when assigning a protein to a superfamily. This is evidenced by two main factors: first, manual curation is still required to verify family membership, and second, different databases often yield conflicting results for the same domains (Greene et al. 2007; Andreeva et al. 2008). It is worth noting that while both CATH and SCOP currently employ automated protocols for the preliminary Classification of novel protein structures, the final assignment to a superfamily still involves manual curation.
METABOLISM/2.html">THE CONCEPT OF a family is more elusive. Currently, a family is generally understood as a sub-classification of homologous proteins that satisfies specific criteria. For instance, a sequence family with a defined similarity threshold encompasses all proteins sharing at least that level of identity; a functional family comprises homologs that share a common function; an orthologous family includes all orthologs, and so on. Depending on the primary focus of a given database, the definition of a family may vary.
6.3.1.2. Practical approaches
This section focuses exclusively on Structure-based databases.
CATH and Gene3D. In the CATH classification, domains sharing a specific topology (see Section 2.1.2) are assigned to the same homologous superfamily (the H-level, standing for "Homologous") if they are presumed to share a common ancestor. Two domains are deemed homologous if they meet at least two of the following criteria: (a) structural similarity determined using empirically derived thresholds; (b) sequence similarity established via standard sequence comparison and HMM search methods; and (c) functional similarity confirmed by manual analysis. Gene3D extends this classification to proteins of unknown structure by searching sequences against HMM profiles from the CATH library, thereby mapping sequence regions to CATH homologous superfamilies (Yeats et al. 2008). CATH superfamilies are further subdivided into sequence families, each defined by specific sequence identity thresholds. A 35% sequence identity cutoff is used to define non-redundant protein groups (s35 families).
SCOP and Superfamily. In SCOP, superfamily homology is determined based on sequence similarity or through manual comparison of Structural and functional properties (Andreeva et al. 2008). While this manual curation approach provides the research community with an expertly maintained domain structure classification, it is inherently subject to the unavoidable human subjectivity associated with manual curation processes. Domains are grouped into the same SCOP family if a "clear evolutionary relationship" is established between them. In practice, this generally means that Protein domains are placed in the same family if their pairwise sequence identity exceeds 30%. However, domains lacking high sequence identity may still be grouped into the same family if their structural and functional similarities provide unambiguous evidence of common descent. This flexibility AIDS in establishing remote homologous relationships, though it also increases the subjectivity of the process. The Superfamily database enables the Classification of Proteins of unknown structure by leveraging SCOP data to annotate sequences at the family and superfamily levels (Wilson et al. 2007). Similarly to Gene3D, Superfamily utilizes SCOP-based HMM profiles to evaluate sequence matches.
SFLD (Structure-Function Linkage Database) is a recently developed database specifically designed to investigate the relationship between Structure and function in homologous Enzymes. Although it currently contains a relatively small number of superfamilies compared to CATH and SCOP, SFLD provides detailed descriptions of their functional evolution. In the SFLD, enzymes belonging to the same superfamily must not only share homology but also exhibit a hallmark catalytic mechanism mediated by conserved structural elements (Pegg et al. 2006). SFLD families consist of enzymes within a superfamily that catalyze the same overall reaction.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.