Protein Structure and Function. Application of Bioinformatics Methods - John Rigden 2014
Fold Recognition
Prospects
None of the most successful Methods in recent CASP competitions relied solely on threading. In fact, many top-performing methods do not use threading at all. Some approaches employ Structure/24.html">Empirical potentials either to evaluate candidate models upon completion or in combination with profile-based methods (see, e.g., Jones 1999b; Zhang 2007). The initial dominance of threading-based approaches followed by their subsequent decline in popularity raises several interesting questions. A longstanding debate in structural biology concerns The concepts of Homology and analogy. Obviously, A wide variety of different sequences can share a similar fold. Many researchers attribute this phenomenon to the divergent evolution of a common ancestral sequence under selective pressure to maintain a specific structure. However, some researchers suggest that instances of significant sequence divergence sharing the same fold may represent Examples of convergent evolution—that is, the independent evolution toward the same packing arrangement in the absence of a common ancestor. This phenomenon is akin to the convergent Evolution of the bat and bird wing.
Clear examples of convergent evolution do exist in Proteins, where similar local structural elements have evolved independently on multiple occasions. Perhaps the most famous example is the Ser/His/Asp catalytic triad (Dodson and Wlodawer 1998), found in at least five distinct protein folds that are difficult to consider homologous. Such evidence supporting convergent evolution suggests that threading methods could prove useful where sequence-profile approaches fail. Yet, The Use of threading appears to be gradually fading, being superseded by sequence- and profile-based methods.
There are several potential reasons for this trend. First, the question of whether protein folding occurs for the entire molecule as a single cooperative unit or via local structures that have evolved independently multiple times remains open. Nature may have independently stumbled upon simple folding motifs, such as four-helix bundles, several times over, but such an occurrence is much harder to substantiate definitively for more complex structures. For certain fold types, such as TIM barrels and ß-trefoils previously regarded as classic examples of convergence, a growing body of evidence now points toward homology, driven by improvements in the sensitivity of sequence comparison methods (Copley and Bork 2000; Ponting and Russell 2000).
Second, even if true analogs exist, sequence-based methods can often detect them by exploiting the shared biophysical constraints required for a given fold. These constraints are subsequently reflected in high-quality profiles constructed from numerous distantly related homologous sequences. Third, the CASP competition is widely recognized as the standard benchmark for evaluating Cell/13.html">Protein Structure Prediction quality. An unfortunate side effect of CASP's prominence is that methods capable of identifying analogous fold relationships tend to be overshadowed by techniques capable of accurately aligning protein sequences against closely related homologs from ever-expanding structural Databases. Whenever homologous relationships can be established within structural databases, they invariably yield superior models compared to analogous relationships. In other words, as sequence and structural databases continue to grow, the need to identify analogies diminishes because: a) the scope for deeper searches through sequence space expands, and b) a richer Selection of close structural templates becomes available.
This leads us to the Conclusion that simply determining the structures of a carefully curated, representative subset of proteins (Marsden et al. 2007) would enable the relatively accurate modeling of the vast majority of genomic sequences. From the perspective of de novo fold design or Ab Initio Structure prediction, such outcomes may be considered unsatisfactory. However, for the practical purpose of annotating protein structures across whole genomes, this approach is more than sufficient, provided that a sufficiently large and well-chosen set of structural templates is available.
It remains unclear to what extent improvements in structure prediction stem from the growing size of databases versus algorithmic refinements. Although sequence databases are growing exponentially, The amount of genuinely novel information does not increase at the same pace. The vast majority of sequences added to databases annually are closely similar to those already present. Recent work (Chubb, Kelley, and Sternberg, manuscript in preparation) demonstrates that despite the massive expansion of databases, homology detection using standard tools (such as PSI-BLAST) has plateaued. Consequently, it is difficult to envision any further dramatic leaps in homology detection based solely on sequence database growth. It is likewise unclear how much of the recent progress in ab initio prediction methods can be attributed to the growth of structural databases, which supply the structural fragments essential for fragment-assembly techniques (Zhang and Skolnick 2005).
Sequence and structural databases will undoubtedly continue to expand. Even if The Development of structure prediction algorithms were to halt today, prediction accuracy would likely continue to rise incrementally. Setting aside the challenges of protein design, structure prediction remains a practical tool for reducing the time and experimental cost of determining protein structures.
The pursuit of "solving" the protein folding problem is still regarded as one of the "Holy Grails" of molecular biology. Yet even in the absence of such a definitive solution, it is probable that within a reasonable timeframe we will succeed in generating accurate and useful models for most, if not all, naturally occurring proteins. Regardless of how many years (5, 10, or 50) of ingenuity and meticulous effort are required from experimentalists determining structures and genomes, and modelers extracting biological insights from those results, the endeavor is well worth it. At this point, it is merely a matter of time.
Protein design, however, remains a multifaceted and formidable challenge rooted in a deep mechanistic understanding of the protein folding process. Understanding protein folding means understanding how the "software" of DNA is translated into the "hardware" or machinery of functional proteins. It means grasping the fundamental nature of living systems at a molecular level. Nevertheless, it is entirely possible that no single, elegant solution to the protein folding problem exists. Nature is under no obligation to use an elegant solution—it only needs one that works. Reluctantly, we may have to accept that we must make do with complex predictive machinery. Even so, strong hopes persist that a simple, computationally tractable, and as-yet-undiscovered principle governing protein folding will ultimately be found.
Acknowledgments. I would like to thank Dr. Benjamin Jefferis for his invaluable assistance in preparing the illustrations for this chapter.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.