Practical Protein Chemistry - A. Darbre 1989
Prediction of Peptide and Protein Conformation
Secondary Structure Prediction. Special Calculation Methods
The predictions discussed in this section are based on calculations without energy Minimization. Nonetheless, they undoubtedly fall into the category of Heuristic Methods and serve as a convenient strategy for generating a starting conformation for a globular protein prior to energy minimization. Selecting an appropriate starting conformation is particularly critical due to the multiple-minima problem. At the same time, Secondary Structure prediction cannot directly yield the Tertiary Structure of the object; additional Processing, including energy minimization, is required. Two major considerations support this view. First, secondary structure prediction algorithms disregard the interactions responsible for maintaining a compact tertiary structure, focusing exclusively on those interactions that are most crucial for forming the secondary structure. Because the secondary structure still depends to some extent on these neglected interactions, an exact prediction is unattainable. Second, even a correct secondary structure prediction assumes that each dihedral angle defining the protein backbone geometry falls within a specific range. For a large protein, a minor inaccuracy in a dihedral angle value can translate into massive uncertainty in the spatial coordinates of numerous atoms and, ultimately, in the tertiary STRUCTURE OF THE molecule under study. Nevertheless, A number of researchers perform secondary structure prediction without subsequent energy minimization—an approach so widespread that it warrants special Discussion. It should be noted that one can hardly expect too much from predictive methods when applied in isolation from other data about the system, and the inherent Limitations of the method must always be kept in mind.
If potential energy calculations are not used in secondary structure prediction, what do these methods actually entail? With the exception of those that draw upon experimental data regarding the helix–coil transition in synthetic Polypeptides, the vast majority of approaches are statistical in nature. They utilize the observed frequencies of conformational states for individual amino acid residues, derived from sequence–conformation correlation tables for Proteins of known 3D structure. The simplest example of a predictive approach is the empirical fact that Proline is never found within the helical regions of proteins deposited in the Protein Data Bank (except at the N-terminal position). Therefore, when searching for an energy minimum, instances where a proline residue is present within a helical region of the initial protein conformation are excluded from consideration altogether.
Researchers are continuously striving to identify algorithms and parameters capable of capturing the observed correlations between an Amino Acid Sequence and a Protein secondary structure, as well as predicting the Conformations of proteins not included in the initial training dataset. In early predictive methods, each of the 20 amino acid residues was classified as either helix-forming or helix-disrupting. The reality that no single amino acid residue adopts a strictly unique conformation was addressed by introducing a set of simple rules (for instance, that at least four helix-forming residues are required to nucleate an α-Helix). Furthermore, studies of the helix–coil transition in synthetic polypeptides demonstrated that categorizing amino acid residues strictly as helix-forming or helix-disrupting units is an oversimplification. Consequently, a propensity scale for helix formation was proposed for all 20 amino acid residues [48]. In the search for a satisfactory predictive Procedure that would rely on a minimum of assumptions and be grounded in the most objective Properties of the source database,
information theory methods were also brought into play. The goal of this endeavor was to eliminate the unjustified influence of physical factors and to treat sequence and conformation data as two texts linked by an unknown code. To enhance the objectivity of the Conclusions, information-theoretic representations were supplemented by Bayesian estimation. This approach allows for the identification and subsequent reduction or elimination of subjective bias in the final data. A formal justification for this approach, along with its various Applications, is discussed in [58].
The relative ease of implementing secondary structure prediction methods has led to The Development of numerous algorithms across various laboratories. However, methods employing information theory undoubtedly deserve preference. Rather than focusing on the design of specific predictive algorithms, these frameworks establish the necessary general formalism for finding optimal algorithms. For instance, it has been demonstrated [64] that the popular Chou–Fasman method [6, 7] is entirely consistent with information theory [63] if the contribution from inter-residue interactions is neglected. Moreover, this method serves as an excellent illustration of the distinctive features of information-theoretic approaches in this field. The following core relationship is utilized:
Class="center">![]()
The meaning of the function
[] is explained below in relation (21.18). Expression (21.17) defines the information content of a residue of type R (e.g., Alanine) residing in conformational state S (e.g., an α-helix). Here, f(X, R) is the observed frequency of occurrence (number of events) of residue R in state S = X within the database, while
is the frequency with which residue R adopts other conformations
. The frequencies e(X, R) and
correspond to the "expected frequencies" determined via the chi-squared criterion, for example, e(X, P) = f(X) · f(P) / foбщ. The function # [] expresses the "information content" within these frequencies:

Function values for non-integer arguments can be obtained via interpolation.
In reference [58], observed frequencies decreased by one were used as values of i for i>1; however, although widely adopted, this practice rests on a debatable theoretical premise. According to the formula given above, an unmodified frequency value can serve directly as the function argument without leading to any significant alteration in the result. Summing or subtracting small numbers reflects confidence in the result depending on how the data were collected, but The Effect of such adjustments should be minor.
Any specific application of the information method depends on which contributions are neglected in the expression I(Sj; R1 ..., Rn). Here, j is the index of the residue whose conformation is being predicted, and Ri, ..., Rn is the complete protein sequence for which secondary structure prediction is performed by considering all indices j. In the simplest variant of the method—which yields results comparable in quality to other algorithms [12]—the following type of approximation is used:
![]()
Here, j is the index of the residue for which the conformation is predicted, Sj is the conformation type, and Ri+m is the type of residue located m positions away along The amino acid sequence of the protein. Positive m values correspond to amino acid residues toward the C-terminus, while negative values point toward the N-terminus of the polypeptide chain, starting from the residue with index j.
METABOLISM/18.html">The Influence of residues beyond m ± 8 on the protein secondary structure is generally considered negligible, although these boundaries are somewhat arbitrary. The value m = 0 corresponds to THE CONTRIBUTION OF the residue occupying the central position within the segment. A similar prediction is carried out for every residue in the sequence from j = 1 to j = n, after which, if necessary, the calculations are repeated for any conformational state S. As a result, each residue is assigned the conformational state S that exhibits the highest information level. In practice, a certain empirical constant—dependent on the conformational state type S—is subtracted from the results obtained for each state. This constant can be viewed as additional information derived from circular dichroism data. Knowing the exact secondary structure content is not strictly necessary; rather, it is sufficient to classify the protein into one of several general types, such as helical, β-sheet, etc. The physical rationale behind this procedure is clear, as a protein rich in β-regions will tend to undergo further stabilization via the cooperative formation of Hydrogen Bonds within the sheet structure. It is equally important that long α-helices are more stable than short ones.
Study [12] can be considered typical for such investigations; four conformational states were taken into account: right-handed α-helix, extended chain (potential sheet element), β-turns, and irregular structure. When classified into these four conformational types, ~60% of the residues were correctly predicted as belonging (or not belonging) to a given type. This predictive accuracy has been observed for many proteins, though significant deviations occurred in a few cases.
The implemented algorithm is written as a high-level programming language code, but short sequences could easily be evaluated manually. It was demonstrated that the prediction results obtained in this manner can be utilized to generate the starting conformation of proteins when modeling self-assembly processes via energy minimization.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.