Principles of Protein Structure Determination - G. Schulz 1982
Thermodynamics and Kinetics of Polypeptide Chain Folding
Modeling of the Folding Process
To determine a three-dimensional Structure from an Amino Acid Sequence, it is necessary to understand The Mechanism of chain folding.
The three-dimensional structure of a protein is determined by non-covalent interactions between The amino acid residues of the chain, as well as between these residues and the solvent (Ch. 3). In principle, if all these interactions are taken into account, the native conformation can be calculated from a known Covalent Structure. However, since the native conformation may not correspond to the global energy minimum, calculating the energy of all possible chain Conformations may not yield the correct answer. Most significantly, however, considering all possible chain conformations would require computing time that far exceeds the age of the Earth. Obviously, therefore, such calculations can only be carried out if one does not attempt to cover all static structures, but instead tries to simulate the folding process by following the natural folding pathway of the given chain. If this pathway is unique (analogous to a deep ravine and a ball; Section 8.2), calculations of moderate accuracy will be able to lead to the correct solution of the problem. But if the pathway is not sufficiently well-defined, high calculation accuracy is required.
Folding simulation is based on the Hierarchy of Protein structure levels. Several attempts have been made to simulate the folding process of, for example, apomyoglobin [470], BPTI [30, 404], and parvalbumin [471]. All of them are based on the assumption of a strict hierarchy of Cell/13.html">Protein Structure levels (Section 5.6), presuming that the a-helices of the native structure have already formed in the still incompletely folded chain. This assumption is partly supported by immunological data, which can be used to estimate the a-helix content in the unfolded chain [418, 454], as well as by fairly satisfactory results from a-helix prediction Methods (Section 6.5). Chain folding simulation is performed a posteriori for Proteins whose structures have been determined by X-ray crystallography. Applying such simulation to a protein with an unknown three-dimensional structure will require information about the a-helical Regions of the chain; this information can be partially obtained using prediction methods (Ch. 6).
The folding of apomyoglobin was considered as the assembly of a-helices. The simulation of apomyoglobin folding [470] was carried out without using a computer, since an extremely simple model was employed, in which the nine helices observed in the native structure were predefined (Fig. 8.4, a). This significantly narrowed the considered conformational space. At the initial stage of the folding process, four structural nucleations were identified, each containing three adjacent a-helices in the chain. Within each nucleation, the helices were arranged in various ways. The Free energy of these helical formations was calculated solely based on the hydrophobic interactions (Section 3.5) of large non-polar side chains; it was assumed that none of the charged side chains resided inside the protein. Out of A large number of obtained helical constructs, only the 40 most energetically favorable structures falling within a 4 kcal/mol interval from the lowest free energy form were subjected to subsequent analysis. In the subsequent stages of simulation, other helices were added to these nucleations and the structures were combined until the entire chain was folded. Starting from a certain stage, only conformations whose free energy exceeded the minimum by no more than 4 kcal/mol were considered. This Procedure resulted in 19 different ways of chain folding. As expected, the arrangement of helices with the minimum free energy (Fig. 8.4, b) corresponded to the native structure of Myoglobin (Fig. 8.4, c). However, the differences in calculated energies for this form and the structure with the next lowest free energy were only 3%, which is much smaller than the expected error.
Class="center">
Fig. 8.4. Simulation of apomyoglobin folding [470]. a — a-helices predefined at the beginning of the folding process simulation. At the initial stage, the four nucleations indicated in the figure were considered, each including three a-helices. b — result of the chain folding simulation corresponding to the minimum calculated value of ΔGtot. c — model of the myoglobin molecule according to X-Ray Diffraction data.
When reviewing this method, the following remarks should be made: a) the method is applicable only to class I proteins (Table 5.2), i.e., proteins consisting predominantly of a-helices; b) at the initial calculation stage, the existence of large nucleations comprising one-third of the molecule is assumed, although experimental data on such nucleations are currently lacking; c) it was assumed that adjacent helices along the chain are included in such nucleations, which corresponds to low chain Entropy and the presence of strong correlation over adjacent residues (Section 8.3); d) it was considered that the native structure folds via a direct pathway: all intermediates were discarded as being incapable of rearrangement. The latter assumption contradicts experiments on proteins containing Disulfide Bonds, which have revealed the existence of intermediates with a geometry significantly different from the native one (Fig. 8.1). However, the apomyoglobin folding process occurs much faster, and it can therefore be assumed to be more directed compared to the folding of disulfide-containing proteins (Section 8.2).
In BPTI and parvalbumin, folding of the simplified chain was performed by descending along the energy gradient. A somewhat different method was used to simulate the folding of the BPTI chain [30, 404] and parvalbumin [471]. In this case, the polypeptide chain was considered in more detail than in the apomyoglobin simulation, but the approximation remained crude, as only the rotation angle around the virtual Ca–Ca bond was used as an independent variable (Fig. 7.10). The virtual valence angle at the Ca atom was considered a function of the torsional Ca angle; side chains were represented as rigid spheres (Fig. 7.10).
This simplification reduced the number of independent parameters to 1 per residue.
The simulation of the folding process was initiated from a chain having an extended quasi-random conformation, with the exception of native a-helices, which were considered already formed. Then, a descent along the free energy gradient was carried out. The kinetic energy of the chain was not taken into account during this process. To account for Temperature fluctuations (Section 8.1), it is necessary to "shake" the chain at short time intervals. Since this would take too much computer time, chain "excitation" was performed only when it fell into an energy minimum, in order to drive it out of a potential local minimum. Unlike the apomyoglobin calculations, this method considers both hydrophobic forces and binding energy. In addition, no fixation of any subensembles was performed, and continuous rearrangements in all PARTS OF THE chain were permitted. As a result, it was shown that with a 50% probability, the BPTI chain folds into the native conformation with a ROOT-mean-square deviation of the Ca atom of about 6 Å. However, even the best results led to an incorrect topology of the β-structure, for which they were criticized [801]. It should be noted that this simulation did not take into account the specific formation of disulfide bonds that occurs during BPTI folding (Fig. 8.1). Using this information could prove useful in further research. Nevertheless, experimental data on the renaturation of disulfide-containing proteins show that the folding pathway of such Proteins can be significantly more complex than that of other proteins (see p. 189), and therefore disulfide-containing proteins are probably not the most suitable subjects for The Development of theoretical analysis.
In the case of parvalbumin [471], a class I a-helical protein (Table 5.2), folding simulation was less successful. This can be partly explained by the failure to account for the possibility of Ca2+ ion binding at two sites. Since these sites are located in short loops between helices [59], it can be assumed that they influence the relative orientation of adjacent helices, and thus the folding process.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.