Protein Structure and Function. Application of Bioinformatics Methods - John Rigden 2014
Fold Recognition
Introduction
Ab Initio Structure Prediction and Homology Modeling
If we ever hope to describe The Structure of a significant proportion of Proteins in nature without resorting to revolutionary experimental Methods, we will need a computational approach that enables structure prediction directly from sequence. After Anfinsen demonstrated in 1961 that denatured Ribonuclease can refold while retaining its enzymatic activity, the idea that the information required for a protein to attain its native conformation is encoded entirely in its sequence gained widespread acceptance. Consequently, over recent decades, Cell/13.html">Protein Structure Prediction has relied on a combination of “ab initio” (or de novo) methods—which use only the Amino Acid Sequence as input alongside the laws of physics (or approximations thereof)—and empirical data. While certain successes have been achieved in this direction, as described in Chapter 1 of this book, these methods are generally characterized either by prohibitive computational costs that hinder Practical Application, or by low throughput and inaccurate results when applied to systems larger than small proteins (fewer than 100 amino acid residues). Although a strictly physical approach might seem the only correct solution to the protein folding problem, structure prediction is of immense practical importance; therefore, we must accept existing limitations and move, at least temporarily, toward more pragmatic solutions. This shift in perspective has led the field of protein structure prediction to pivot away from pure physics toward data mining and machine learning approaches.
It has long been established that similar protein sequences fold into similar structures. Therefore, given a novel protein sequence whose structure is to be determined (hereinafter referred to as the “target sequence”), it is straightforward to check whether homologous sequences of known structure already exist. If such sequences with a high degree of similarity are available, the structure determination process becomes readily achievable through amino acid sequence alignment. By employing a simple scoring system for Amino Acid Substitutions, such as the BLOSUM matrix, combined with a Dynamic Programming Algorithm like the Smith-Waterman algorithm, two sequences can be aligned rapidly and optimally According to the scoring function.
Class="center">
Fig. 2.1. (For the color version of this figure, see the color insert.) Schematic representation of a simplified modeling algorithm based on sequence alignment between a target protein and a template. The alignment between a sequence of known structure (the “template sequence”) and the target sequence is shown. Indels (insertions and deletions) are indicated by blurred lines; amino acid substitutions are shown in red letters. Residues are colored according to their biophysical properties. Thin wavy lines connect corresponding positions in the target and template sequences.
Once the sequence alignment against a known structure (hereafter referred to as the “template”) is established, a rough model can be built simply by copying the spatial coordinates of the template and mutating The amino acid residues to match those of the target sequence according to the alignment (Fig. 2.1).
This initial model can subsequently be refined using a variety of Homology modeling techniques discussed in the relevant chapter of this book. The advantages of this approach are clear: it is computationally efficient, and the accuracy of the resulting model is remarkably high, provided There is a high degree of sequence identity between the target and the template. However, this immediately highlights the fundamental limitation of the method: if no homologous sequence with a known structure is available, the approach yields no results at all.
Thus, research into the protein folding problem has historically branched into two main directions. The first direction, grounded in fundamental physical principles, aims to develop a robust, universal method for predicting structure from sequence, which would also facilitate protein design, The Study of Molecular Dynamics, and numerous other vital Applications. Yet, developing such a method presents severe computational and theoretical challenges and will likely remain out of reach for years to come. The second direction represents a straightforward yet inherently limited heuristic approach—homology modeling—which yields high-quality models, but only in a very restricted number of cases. To bridge the gap between these two contrasting paradigms, a methodology known as “Fold Recognition” (or threading) was developed.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.