Principles of Protein Structure - H. Schultz 1982
Prediction of secondary structure from amino acid sequence
Probabilistic methods
Residue propensity for secondary structure
The simplest approach is to consider each residue individually, without regard to its near or distant neighbors. A straightforward, purely statistical approach was applied by Dirks [338, 339]. Using crystallographic protein structures, he determined the frequency of occurrence of each of the 20 residues in a-helices, ß-sheets, and Reverse Turns of the chains (rt). These frequencies were then adopted as characteristics reflecting "residue conformational preferences for a-, ß-, and rt-Conformations." Table 6.2 presents Dirks's data regarding the tendency of residues to occur in ß-structures as an example.
To predict the a- (ß-, rt-) Structure of a residue at each position of the chain, the "a-(ß-, rt-) potential*" was calculated as the weighted average of the a-(ß-, rt-)-tendencies of the $m$ nearest residues along the chain. As soon as the a-(ß-, rt-) potential exceeded a certain threshold value, the a-(ß-, rt-)-conformation was predicted for that residue. Thus, this potential function was designed to yield a yes/no answer. The weighting scheme was fixed. Three threshold values, as well as the value of $m$, were chosen based on the best agreement between the predicted and observed secondary structures in the training set. The optimal values of $m$ are 17, 11, and 3 for a-, ß-, and rt predictions, respectively. The fundamental concept of "residue propensity" and the derived concept of "residue potential" will be used in this chapter precisely in the sense defined above.
* The term "potential" is used here, of course, not in a physical sense, but in a broader one. — Ed. note.
Class="center">Table 6.2 Relative frequencies of residue occurrence in ß-structures and reverse turns rta
|
Propensity for β-structure |
rt-propensity |
||||||
|
Chou and Fasman, 1974 [340] |
Burgess et al., 1974 [31] |
Beachy and Dirks, 1975 [339] |
Lewis et al., 1971 [326] |
Kuntz, 1972 [203] |
Chou and Fasman, 1974 [340] |
||
|
A |
Ala |
0,97 |
0.29 |
0,37 |
0,22 |
(T) |
0,15 |
|
С |
Cys |
1,30 |
0,53 |
0,84 |
0,20 |
0,31 |
|
|
D |
Asp |
0,80 |
0,27 |
0,97 |
0,73 |
т |
0,33 |
|
E |
Glu |
0,26 |
0,26 |
0,53 |
0,08 |
т |
0,12 |
|
F |
Phe |
1,28 |
0,32 |
0,53 |
0,08 |
— |
0,19 |
|
G |
Gly |
0,81 |
0,31 |
0,97 |
0,58 |
т |
0,45 |
|
H |
His |
0,71 |
0,20 |
0,75 |
0,14 |
т |
0,18 |
|
I |
Ile |
1,60 |
0,41 |
0,37 |
0,22 |
— |
0,15 |
|
К |
Lys |
0,74 |
0,27 |
0,75 |
0,27 |
т |
0,27 |
|
L |
Leu |
1,22 |
0,40 |
0,53 |
0,19 |
— |
0,14 |
|
M |
Met |
1,67 |
0,38 |
0,64 |
0,38 |
— |
0,18 |
|
N |
Asn |
0,65 |
0,23 |
0,97 |
0,42 |
т |
0,45 |
|
P |
Pro |
0,62 |
0,34 |
0,97 |
0,46 |
т |
0,41 |
|
Q |
Gln |
1,23 |
0,33 |
0,64 |
0,26 |
т |
0,15 |
|
R |
Arg |
0,90 |
0,36 |
0,84 |
0,28 |
т |
0,27 |
|
S |
Ser |
0,72 |
0,35 |
0,84 |
0,55 |
т |
0,41 |
|
T |
Thr |
1,20 |
0,39 |
0,75 |
0,49 |
т |
0,27 |
|
V |
Val |
1,65 |
0,50 |
0,37 |
0,08 |
— |
0,08 |
|
w |
Trp |
1,19 |
0,23 |
0,97 |
0,43 |
— |
0,30 |
|
Y |
Tyr |
1,29 |
0,43 |
0,84 |
0,46 |
т |
0,33 |
а Observed occurrence frequencies were used as propensity characteristics. In Kuntz's nomenclature, T indicates that the corresponding residue tends to reside in chain turns.
In reverse chain turns, the relative position of the residue must be taken into account. Information on individual residues (singlets) was also used to predict reverse turns. Lewis et al. [326] defined peptide chain reverse turns as segments of four residues $i$, $i+1$, $i+2$, $i+3$ in which the distance between the Ca atoms at positions $i$ and $i+3$ is less than 7 Å, and the chain is not in an a-helical conformation. However, unlike in Dirks's approach, the four positions within the turn were not considered equivalent. The frequencies of occurrence determined for each residue at positions $i$, $i+1$, $i+2$, and $i+3$ were termed the "propensity of a given residue to occur at a specific position within the turn." The rt-potential of a residue quartet was then defined as the product of the corresponding propensities. The potential threshold value was established based on the best agreement with available experimental data (Fig. 6.2).
Similar occurrence frequencies were utilized by Crawford et al. [200] and later by Chou and Fasman [340] to predict reverse chain turns identified According to the criteria of Lewis et al. [326]. Table 6.2 lists the frequencies of occurrence found by Lewis et al. [326] as well as Chou and Fasman [340]. However, for the sake of compactness, the table lists only averaged values of residue occurrence frequencies in reverse chain turns. For certain residues, positional differences are notable: Pro is found almost exclusively at position $i+1$, whereas Trp occurs almost exclusively at position $i+3$.
A simple approach utilizing singlet propensities was also applied by Ptitsyn and Finkelstein for a- and ß-predictions [341, 342]. According to these authors [343], no additional information can be extracted from the occurrences of residue pairs (doublets). However, this proposition remains controversial because the initial data used involved only 9 Proteins and because the Classification of the 20 residues into four categories was rather arbitrary.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.