Principles of Protein Structure - G. Schultz 1982
Prediction of secondary structure from amino acid sequence
Probabilistic methods
Propensity of two residues to simultaneously adopt a secondary structure
Doublet propensities involve residue-residue interactions. Strict prediction relies on doublet frequencies. Contrary to Finkelstein and Ptitsyn [343], many prediction Methods incorporate doublet information because it reflects interactions between residues that are close in the polypeptide chain. Parry [344] used doublets to predict a-helices and ß-sheets via a purely probabilistic approach. He examined 27 doublets where residues occupied positions i ± 1, i ± 2, ..., i ± 6, assuming that interactions are negligible at greater distances. A total of 10,800 doublets of various types were thus obtained**. For each residue, Parry evaluated the probabilities of three states: a, ß, and coil (a state that is neither a nor ß). Consequently, based on an experimental dataset, he compiled a table of 32,400 = 3 × 10,800 frequencies of occurrence, which were interpreted as propensities. To predict the Introduction/11.html">Secondary Structure of a given residue within a specific Amino Acid Sequence, the table values were evaluated and combined using a standard statistical Procedure for the a-, ß-, and coil-structure propensities of 27 doublet types. This yielded the a-, ß-, and coil-potentials. These potentials were then compared, and the highest one was selected to determine the conformational state. It is worth noting that this method contains no adjustable parameters. Furthermore, it does not predict a- and ß-structures independently; instead, it combines them in a rather artificial manner.
* For a residue at position i, there are seven doublets of the form (i, i + 6), (i — 1, i + 5), ..., (i — 6, i), six doublets of the form (i, i + 5), ..., (i — 5, i)..., and two doublets (i, i + 1), (i — 1, i), totaling 7 + 6 + 5 + 4 + 3 + 2 = 27 different doublets. Each doublet consists of two residues. Given 20 different residue types, 20 × 20 = 400 different combinations are possible. Thus, the total number of doublet types is 27 × 20 × 20 = 10,800.
Invariant potentials and adjustable thresholds. To predict a-, ß-, and rt-Conformations, Robson et al. [345–352] used singlets combined with 16 doublets formed by the target residue and residues at positions i ± 1 to i ± 8. Singlet and doublet propensities were determined from their frequencies of occurrence in a training set and then converted into a-, ß-, and rt-potentials for a given residue position using standard information theory methods. These potentials were compared against three distinct thresholds—one for each type of secondary structure. A separate optimization of all three thresholds was also performed to achieve the best agreement between the predicted and experimentally observed secondary structures in the training set. Predictions for a-, ß-, and rt-structures were carried out independently of one another.
Adjustable potentials and adjustable thresholds. Nagano [228, 353–356] used the same doublets as Parry [344], with the exception that the maximum distance was set to 7 residues instead of 6. This yielded 35 doublets influencing the target residue; consequently, the table of observed frequencies of occurrence (propensities) contained 20 × 20 × 35 × 3 = 42,000 values. In this way, the propensities of all doublet types for a given residue in a specific amino acid sequence were determined, and their linear combinations subsequently yielded the a-, ß-, and rt-potentials.
Similar to Robson et al., Nagano based his predictions on a separate comparison of the potentials against three different thresholds. Each threshold was adjusted until the best fit with the training set was achieved. Moreover, by formulating the potentials as linear combinations containing 3 × 35 = 105 coefficients—three (a, ß, rt) for each of the 35 doublet types—Nagano introduced A large number of additional fitted parameters. All of these parameters were optimized to obtain the best agreement in the training set. Thus, Nagano utilized a greater number of fitted parameters than Parry, as well as Robson et al., who derived their potentials from propensities using standard probabilistic Procedures.
Nagano also expanded the training set by incorporating (with a low weight) information derived from Amino acid sequences of Proteins homologous to those with known structures; in his later work [356], he attempted to account for the frequencies and propensities of triplets. It should be noted that Nagano's method implicitly includes singlets, as certain linear combinations of doublet propensities are equivalent to singlet propensities.
Prediction of structure-disrupting residues. Using a single type of doublet—specifically, the pair (i — 1, i + 1) of residues flanking the target residue i—Kabat and Wu [357–359] attempted to develop a "negative" prediction method capable of identifying residues that disrupt a-helices and ß-sheets. Propensities for a- and ß-disruptions were derived from frequencies of occurrence in the training set. The potential was treated as a characteristic of the corresponding residue propensity. Potential thresholds were optimized for the best match with available experimental data.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.