Practical Protein Chemistry - A. Darbre 1989

Prediction of Peptide and Protein Conformation
Basic Assumptions of Calculations
Attempts to Predict the Tertiary Structure of Globular Proteins

The idea of predicting the tertiary Introduction/12.html">Structure of Globular Proteins from their Amino Acid Sequence has long attracted researchers' attention. Early attempts in this field were largely haphazard, and it is only recently that studies aimed at developing systematic approaches to this problem have emerged. Some scientists, however, still maintain that the preparatory stage is far from complete and that the time is not yet ripe for serious, in-depth research.

It is now evident that the primary difficulties stem not from the size of the protein molecule, but rather from the existence of numerous minima on its potential energy surface. Consequently, early studies focused on smoothing these potential surfaces to clearly reveal the pathway leading to the global minimum. Naturally, such smoothing required sacrificing certain details of the protein's structural representation. Methodologically, the foundational work in this area is that of Levitt and Warhel [33], while the pioneering study by Ptitsyn and Rashin [54] on predicting Myoglobin structure based on the presence of individual a-helices should also be acknowledged. The protein self-assembly modeling was performed without computers, and thus the helices were represented as cylinders. Hydrophobic patches were identified on the surfaces of these cylinders, allowing them to interact and form optimal structures. As a result, one of the possible helix packings was found to correspond to the observed native structure.

A more general approach proposed by Levitt and Warhel allows for a simplified representation of not only helical but also other Conformations. Their algorithm was implemented in a computer program, which, however, does not make the method free from several serious drawbacks. For instance, the main protein chain lacked the CO-NH peptide group, the Ca carbon atoms were linked by a virtual bond, and the side chains were replaced by 'enlarged atoms' (large artificial atoms representing entire atomic groups). This model was first applied to a small protein, the pancreatic Trypsin inhibitor [33]. Study [33] achieved a certain level of success that, from today's perspective, is primarily of historical significance. It should be noted that this research was highly praised and sparked lively Discussion [14, 44]. This wide debate helped formulate the general criteria currently applied in Cell/13.html">Protein Structure Prediction, according to which any valid Procedure must: 1) be fully automated; 2) use input parameters applicable to any protein; 3) be quantitatively reproducible; and 4) be free from unwarranted ad hoc adjustments to specific objects. Furthermore, The amino acid sequence should serve as the sole variable input data for the protein. Subsequently, Levitt made a major contribution to The Development of Methods for predicting protein tertiary structure.

The method described above differs sharply from the one proposed earlier [30], in which the spatial folding of a protein chain is performed using a computer graphics terminal. In particular, interactive graphics is employed as a powerful tool for Protein Engineering in Blundell's laboratory; however, he believes that such technical innovations should be used with considerable caution. Indeed, many human-made decisions can be executed somewhat more slowly by a non-interactive program, which renders the final result quite objective and reproducible. Therefore, while saving machine time through human intellectual input is valuable, It is important to ensure that an objective (automatic) solution path also exists.

How can the quality of a protein tertiary structure prediction be evaluated? In the case of Secondary structure, a visual (though not always flawless) Assessment of the result is provided by the number of amino acid residues assigned to the correctly predicted conformation type. For tertiary structure, the positions of residues in the calculated and experimental structures are compared. For example, Levitt and Warhel used the ROOT-mean-square deviation of inter-residue distances between the calculated and experimental structures. In Protein Crystallography, this metric is not identical to the root-mean-square deviation between calculated and observed atomic positions when comparing similar structures. Such a comparison involves the superposition of structures using Translation and rotation operations on one structure relative to the other to minimize the root-mean-square deviation. The main difficulty in this comparison is that THE POSITION OF no single amino acid residue can be considered unconditionally correct or erroneous. Moreover, averaging the deviation (the 'score') over the entire protein molecule masks the differences between well-predicted and poorly predicted regions. Therefore, it is useful to Complement the root-mean-square score with a residue-residue distance matrix [47, 49] and molecular stereo images.

One of the model foldings of the pancreatic inhibitor molecule obtained by Levitt and Warhel had a root-mean-square score of 6 Å, which was considered an encouraging result compared to the native structure. It turned out, however, that a root-mean-square score of 6 Å does not necessarily correspond to a structure close to the native one [14]. For a theoretical fold with a 6 Å score to closely resemble the native structure, A number of additional criteria must be met, which was not the case in the pancreatic inhibitor folding study. In recent years, several papers have argued that the root-mean-square deviation on its own reveals very little, and that a value of 6 Å can correspond to virtually any compact random structure. However, if during the simulation of protein self-Organization the root-mean-square score progressively approaches zero with each iteration, it can be concluded that the protein folding is proceeding successfully. According to data obtained in the author's laboratory from the comparison of homologous protein structures, a similarity score of <3 Å indicates significant structural similarity. At the same time, a score of 4 Å may already correspond to substantially different structures. Given this fact, as well as the absence of any other objective score for the result, it is difficult to draw a reliable Conclusion regarding structural similarity unless the 3 Å threshold is surpassed. When Levitt and Warhel applied their previously proposed approach to parvalbumin, the performance was poor, leaving it unclear whether the method is applicable to proteins other than the pancreatic trypsin inhibitor.

In an alternative approach [28], a root-mean-square score of 4.7–6.5 Å for the trypsin inhibitor and 4.0–6.0 Å for rubredoxin was achieved in a small number of iterations. Instead of energy calculations, an optimization method was used, incorporating A large number of constraints derived from experimental data on protein structure. Whenever a calculated physical quantity deviated from its expected value, a penalty was imposed, the magnitude of which increased with the extent of the deviation. The penalty system for regions lacking experimental data remains somewhat ambiguous. Furthermore, the resulting structure is rather coarse because all amino acid residues were approximated by 'giant atoms.' Nevertheless, the prediction quality achieved by this method was the best as of early 1980, and the algorithm appears to be useful for generating a starting protein conformation prior to rigorous energy Minimization.

In addition to the prediction methods discussed above, one can mention Scheraga's algorithm [44, 73], which utilizes the Monte Carlo Method in the Metropolis modification (Section 21.4.3). As the first step in folding the polypeptide chain, Scheraga's approach involves predicting the secondary structure. Although it is difficult to judge the quality of the results obtained, it appears to be no worse than that achieved by the Levitt–Warhel protein self-organization simulation.



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.