Protein Structure and Function: Application of Bioinformatics Methods - John Rigden 2014
Fold Recognition
Introduction
Limits of the Fold Type Space
Let us briefly recall some key facts about The Nature of Proteins. The Protein Data Bank currently houses roughly 50,000 experimentally determined protein structures (Berman et al. 2000). According to the SCOP (Structural Classification of Proteins) database (Murzin et al. 1995), these structures are grouped into only about 1,100 unique fold types (unique topologies) and roughly 1,800 superfamilies (evolutionarily related Protein Families). While the number of proteins with experimentally solved structures grows daily, the discovery rate of new fold types is extremely slow. Moreover, the rate at which new folds are found appears to be plateauing (Fig. 2.2). These findings have led to a widespread consensus that the number of naturally occurring protein folds is finite and relatively small (Marsden et al. 2006). Structural Databases contain hundreds, if not thousands, of Examples demonstrating that highly similar structures can arise from completely different Amino acid sequences. Thus, while it is true that closely related sequences share very similar structures, it is equally true that substantially different sequences can adopt strikingly similar structures.
Consequently, for any sequence chosen from a sequenced genome database, There is a high probability that its structural template has already been characterized by researchers. The primary challenge, therefore, lies in selecting the correct template from the 50,000 available structures and accurately aligning the target sequence to that template. Structure/29.html">Fold Recognition relies on scoring Functions capable of reliably assessing sequence-structure compatibility and achieving precise alignment even when standard Sequence Homology is undetectable.
Class="center">
Fig. 2.2. The diagram illustrates the growth over time in the number of experimentally determined protein structures, alongside the number of distinct fold types included in the SCOP database (Murzin et al. 1995). As the graph shows, while the total number of structures added to SCOP increases rapidly, the number of newly discovered fold types has remained virtually flat since 2004.
Despite the vastness of sequence space—that is, the total number of all possible protein sequences—Cell/13.html">Protein Structure space is apparently much more compact. Whether this constraint stems from Thermodynamics, folding kinetics, or evolutionary Selection is a complex question that lies beyond The Scope of this chapter. Nevertheless, this remarkable limitation has proved immensely valuable for The Development of Protein Structure Prediction Methods.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.