Principles of Protein Structure - H. Schulz 1982
Prediction of secondary structure from amino acid sequence
Physicochemical methods
Methods based on statistical mechanics
Unlike probabilistic approaches, Physicochemical Methods rely not only on known correlations between Amino acid sequences and structures within a benchmark set of Globular Proteins, but also incorporate other experimental and theoretical data. These methods can be broadly divided into those employing stereochemical analysis and those utilizing Statistical Mechanics. The Appendix discusses several key aspects of polypeptide chain statistics that are essential for understanding statistical mechanical methods and the Physical Principles of chain folding.
Energy calculations
Statistical weights can be calculated based on geometry and Structure/103.html">Van der Waals forces. METABOLISM/2.html">THE CONCEPT OF applying statistical mechanics methods to Cell/13.html">Protein Structure Prediction was extensively developed by Scheraga and coworkers. Kotelchuk and Scheraga [363, 364] based their predictions on a simplified method described in the appendix during the derivation of equation (A.4). They calculated the statistical weights zaR , zaL and zsε for each residue* except Gly and Pro, which were treated separately. In their calculations, they accounted for van der Waals interactions (Chapter 3) between various PARTS OF THE chain. Since the system under consideration consisted of both the chain and the solvent, a term accounting for the solvent Free energy contribution was also included. For each residue, the conformation with the highest z value was adopted as its propensity (Table 6.1). Because no preference for the aL state was detected, only "helical" (≡aR) and "random" (coil ≡ non-aR) propensities emerged. Notably, the estimates obtained from these calculations show good agreement with experimental data for synthetic Polypeptides [328] and empirical data, such as triplets in the benchmark set. Subsequently, the average values of all angles ∅, ψ found in the four table positions are used as the predicted values (∅, ψ) for THE POSITION OF the i-th residue in a chain folded in a manner common to the given group of proteins.
* The statistical weights zaR, zaL and zaε represent the probabilities that a given residue adopts a main-chain conformation located in the aR-, aL-, and ε- (extended chain) regions, respectively. These regions are defined in Fig. 2.3.
obtained from the Analysis of the benchmark set of globular proteins (Table 6.1).
Calculation of helical potentials from propensities can account for cooperativity. Building upon these propensities, Kotelchuk and Scheraga [364] proposed a simple prediction algorithm whereby helices are initiated when four consecutive residues with helical propensity occur in a row; propagation toward the C-terminus continues until the process is halted by two sequentially positioned residues with a propensity for random coiling. This directional propagation was a consequence of the computational scheme, according to which The Influence of a given residue on its neighbor was exerted exclusively toward the C-terminus rather than the N-terminus. Using this algorithm implies abandoning the simple non-interacting residue scheme of equation (A.4). Helix initiation and growth are introduced, thereby implicitly accounting for interactions between neighboring residues.
The concept of helix initiation by four residues and unidirectional helix growth until the appearance of two consecutive residues with a tendency to form a random coil was also utilized by Leberman [365], who slightly modified the methodology to achieve better agreement with the benchmark set.
Empirical Distribution Function
Residue propensities for forming helices, derived from occurrence frequencies and found using empirical Functions, are equivalent. Since the volume of structural data for globular proteins is now quite large, it has become possible to determine the statistical weights zaR and zε using equation (A.4). To this end, Badges et al. [31] expressed equation (A.4) in the following form:
Class="center">![]()
where χ = aR or ß; extremely small values were neglected. The integral in parentheses represents the probability that a given residue will exhibit specific angles — and ψ on the (—, ψ) map (Fig. 2.3). This is nothing more than the Boltzmann distribution averaged over all side-chain orientations. Because the function E(—, ψ, χ) is not known with sufficient accuracy, calculating the probability distribution directly is impossible. However, these probabilities must correlate with the observed frequencies of occurrence of residues of a given type on the (∅, ψ) map (Fig. 2.4). Therefore, the expression in parentheses can be satisfactorily described by a continuous function of — and ψ fitted to the observed frequencies of occurrence. Such empirical functions can then be used to calculate the integral (6.1) over the aR- and ε-regions, yielding the probability or frequency of occurrence of aR and ε-Conformations, zaR and zε. The values of zε obtained in this manner are given in Table 6.2*. However, the statistical-mechanics-based approach to determining a residue's propensity to adopt a particular conformational state is, in reality, close to a purely statistical analysis.
The method described above does not allow for the determination of chain bend propensities because, in the corresponding conformations, the angles (∅, ψ) of the residues do not lie within a defined region but are scattered across the (∅, ψ) map. For this case, Badges et al. [31] adopted the statistical approach of Lewis et al. [326], utilizing an expanded benchmark set.
Nonapeptide propensities were used to determine the potential of a central residue to be incorporated into a helix. In their Secondary structure predictions, Scheraga et al. [31, 366] examined combinations of residue propensities from i - 4 to i + 4 and applied them to calculate the aR-, ε-, and rt-potentials (see Definitions in Section 6.1) of residue i. The reason for considering precisely nine residues stems from the fact that, according to earlier energy calculations [367], a nonapeptide represents the exact fragment that critically determines the conformation of the central residue. Subsequently, for all segments comprising more than four residues, a helical conformation (≡aR) was predicted if the aR-potential was higher than the ε- and rt-potentials and also exceeded a threshold value. This threshold was defined as the weighted average over all aR-potentials of the chain. Predictions for ε and rt were made in the same manner.
The Zimm–Bragg model applied to heteropolymers
The Zimm–Bragg model is applicable to natural proteins. A helix prediction algorithm based on the Zimm–Bragg model was proposed by Lewis et al. [368, 369]; in this approach, equation (A.9) accounts for individual relative statistical weights si for each position i of the chain, in contrast to the assumption of equal si values in equation (A.8) for a homopolymer. The probability Pi of finding a residue at position i in the a-conformation is given by the expression:
![]()
* In Table 6.2, ε-propensities are designated as ß-propensities. Scheraga et al. noted that only ε- and not ß-structures can be predicted, because the backbone conformation does not completely define the Hydrogen Bonds.
by means of which, out of all possible 2N chain conformations, only those that are helical at position i are selected (equation (A.6)). Normalization is performed, as usual, by dividing by Z, which characterizes the distribution of possible conformational states of the polypeptide chain (Section A.1). As an example, let us derive P1 for the case N = 2 (see equation (A.7)):

Residues with low relative statistical weights significantly shorten the average helix length. To estimate the helical potential of a given protein, a single value of the initiation parameter σ = 5 ∙ 10-4 was used (Section A.4). In addition, three different s values were introduced for all residue types: s = 0.385 corresponded to helix-breaking residues (B), s = 1.00 to helix-indifferent residues (I), and s = 1.5 to helix-forming residues (H) (Table 6.1)*. The values of σ and s are derived from the slopes and Temperature transitions of curves describing helix–coil transitions in synthetic polypeptides, using equations (A.18) and (A.20). A helical conformation is predicted for all residue positions i for which Pi is greater than the average value of Pi. This yields continuous potential functions because equation (6.2) incorporates the cooperativity of the Zimm–Bragg model, which dictates that helices must have a characteristic length (Fig. A.1). This prediction method yields helical segments about 10 residues long, which is much shorter than the length expected for a given value of a in homopolymers at s = 1, i.e.,
(equation (A.17)).
Such helix shortening is a consequence of incorporating residues with low s values.
Modified Zimm–Bragg model. Similarly, the Zimm–Bragg model was applied by Ptitsyn et al. [370], who used a single initiation parameter σ = 5 ∙ 10-4 for all residue types and six different si values based on experimental data from synthetic polypeptides. The si values for residues lacking experimental data were chosen According to the method of Lewis et al. [368]. In subsequent works [371–374], stereochemical data were also invoked to determine the s values (Table 6.1). The Zimm–Bragg model was modified by accounting for the influence of neighboring residues (e.g., residues i and i + 3 in an a-helix). Furthermore, the authors introduced individual σ values for each residue type and expanded the algorithm in an attempt to predict up to six different conformations instead of just two (helix and coil).
Zimm–Bragg parameters derived from globular protein and synthetic polypeptide data
* As noted in Section A.2, the relative statistical weights s correspond to propensities. In equation (A.8), s is the propensity for helix formation.
Probabilistic Nature of the Chou–Fasman prediction method; incorporation of cooperativity. The prediction method of Chou and Fasman [201, 340], which is in many respects based on the Zimm–Bragg model, is probabilistic in nature. The authors determined propensities for a-, ß-, and rt-conformations from their occurrence frequencies in the benchmark set. For a given Amino Acid Sequence, they sequentially analyzed initiation sites for a- and ß-conformations by combining a- and ß-propensities in hexa- and pentapeptides, respectively. From each identified initiation site, the chain was then propagated in both directions until residues with low propensities were encountered. To predict reverse turns, the authors employed the scheme of Lewis et al. [326]. The application of the Chou–Fasman method is described in Section 6.3, where further details are provided.
Helix propensities derived from the statistical mechanics of synthetic polypeptides closely align with those based on occurrence frequencies in globular proteins. In their subsequent work, Chou and Fasman [201] compared helix propensities determined from observed frequencies in globular proteins with data derived from the temperatures of the helix–coil transition in synthetic polypeptides, based on the Zimm–Bragg model. As shown in Sec. A.5, the relative statistical weights s—and consequently the helix-forming propensities—can be determined from the transition temperature θ. Chou and Fasman demonstrated that the s values for seven residue types, for which data on synthetic polypeptides were available, agreed within 10% with the helix propensities derived from occurrence frequencies in globular proteins. This correspondence was investigated in greater detail by Suzuki and Robson [352].
Initiation parameters derived from synthetic polymer data correlate well with the occurrence frequencies of residues at helix termini. Chou and Fasman [201] also compared the initiation parameters σ, obtained from helix–coil transitions (Sec. A.5), with the frequencies of residues occurring at the ends of helices in globular proteins. As in the studies by Lifson and Roig [375], averaging over both ends of the helix was performed in this case. In addition, the occurrence frequencies of three residues at each end were considered (rather than a single residue as in the Zimm–Bragg model), since each of the three terminal residues of an α-Helix forms only a single Hydrogen bond. When comparing the found frequencies with the six σ values obtained from The Study of helix–coil transitions, a good correlation was found despite the difference in absolute magnitudes (the correlation coefficient between the Logarithms was +0.75). Because random coil–β-sheet transitions in synthetic polymers are considerably less studied, it has not yet been possible to establish such correlations for β-structures.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.