Principles of Protein Structure - G. Schulz 1982

Methods of polypeptide chain folding and association
Structural domains
Correlation between closely situated residues in the sequence

Structural domains are clearly delineated on electron density maps. An analysis of electron density maps derived from X-Ray Diffraction studies has revealed that many Proteins consist of several globular regions that are relatively weakly linked to one another. These regions, which are clearly demarcated on electron density distribution maps, have been given the somewhat ambiguous designation of “structural domains.” Evidently, the definition of a domain is rather loose, leaving room for numerous borderline cases. Among Globular proteins, well-defined domains were first discovered in IMMUNOGLOBULINS, as schematically illustrated in Fig. 4.2, c; in this case, the domains are arranged along the polypeptide chain like beads on a necklace.

Residues that are far apart along the sequence also tend to be distant in the three-dimensional Structure. As can be seen from Fig. 4.2, c, the polypeptide chain can be divided into several sequential segments belonging to successively arranged domains [248]. Thus, residues that are distant from one another in the Primary Structure are also significantly separated in physical space. This principle holds true for all globular proteins composed of well-defined domains (Table 5.2). In other words, domain architecture indicates a high degree of “correlation between sequence-neighboring residues,” meaning that the distance between residues along the chain correlates with their spatial distance in the tertiary structure.

The existence of a correlation between sequence-proximal residues facilitates the identification of structural domains. This approach was quantified for Chymotrypsin, whose subdivision into two domains was not entirely obvious from the electron density map. The chosen criterion was the sum of all reciprocals of the distances between pairs of residues separated by 6–25 residues in the sequence:

Class="center">

The results are presented in Fig. 5.14, a. Two peaks separated by a minimum indicate the presence of two regions in which residues close in the sequence were brought together in space during the folding process. These regions should be classified as domains. A comparison with the three-dimensional structure of chymotrypsin shows that each domain contains one of the ß-structures shown in Fig. 5.17, d.

Fig. 5.14. Correlation between adjacent residues of the polypeptide chain. (a) Measure of correlation between neighboring residues as a function of residue index i in chymotrypsin. The curve is smoothed by averaging values over 10 residues. The two-domain structure, schematically depicted in Fig. 5.17, d, is clearly discernible. (b) Domain architecture of adenylate kinase. The small and large structural domains are connected by two polypeptide chains. The correlation between neighboring residues is stronger in the small domain than in the large one.

Consecutive segments of ß-structures tend to be located close to one another. The aforementioned correlation between sequence-proximal residues is manifest not only at the domain level, but also within the domains themselves. A clear quantitative example is the distance histogram for ß-structural chains shown in Fig. 5.15, a. In particular, in antiparallel ß-sheet chains, where the pure peptide chain forms two or more successive ß-pleats, a pronounced correlation is observed between residues that are close in the sequence. In parallel sheets, this correlation is weaker, yet still biologically significant.

Fig. 5.15. Preferred structures of globular proteins.

(a) Correlation between neighboring residues observed in antiparallel and parallel ß-structures. Shown is the number of contacts between a single strand in the ß-structure (along the chain) and its first, second, etc., adjacent strands [327]. (b) A rope initially held at the upper end is allowed to fall. The folded rope is free of knots and exhibits a clear correlation between its neighboring segments. (c) A polypeptide chain knot; not observed experimentally.

The correlation between sequence-proximal residues can be interpreted using both kinetic and thermodynamic approaches. This correlation presumably arises during the chain-folding process. Since there are vastly more Conformations that bring sequence-distant residues close together than those that bring sequence-proximal residues together, the folding process must heavily rely on interactions between the latter. Otherwise, exploring an excessive number of conformational states at the initial stage would be required, making the Separation of correct states from incorrect ones exceedingly difficult, and the chain-folding process would be highly prone to misfolding.

From a thermodynamic standpoint, the considerations outlined above are as follows. During folding, the conformational Entropy of the chain must decrease (Section 3.5). The initial drop in entropy is minimized if chain folding leads to a conformation close to the average statistical conformation of the chain in solution—that is, if it yields a structure with strong correlations between sequence-proximal residues. With such a modest decrease in entropy, no initial (activation) entropic barrier arises that would be difficult to overcome within a reasonable timeframe. Thus, the chain-folding process can proceed gradually without requiring excessive binding energy or hydrophobic forces (see also Section 8.3).

Taking into account the correlation between sequence-proximal residues, a folding chain can be compared to a rope lowered in the manner shown in Fig. 5.15, b. The resulting pile on the floor is not random, but is characterized by correlations between neighboring segments; it is untangled, and the rope can be easily lifted by one end. This analogy is supported by the absence of “knots” in all known protein structures. The term “knot” is used here in its colloquial sense (Fig. 5.15, b) rather than in a mathematical one (mathematical knots can exist only in closed systems).

Structural domains apparently serve as the units of protein folding. The correlation between sequence-proximal residues provides grounds for viewing domains as segments of the chain that fold independently of one another. If this is true, structural domains can be defined as units of chain folding—that is, more specifically than before. The architecture of chymotrypsin lends a degree of support to this definition [18]. This protein contains 13 Water molecules trapped between the two domains within the molecule (Fig. 5.14, a). Evidently, the domains fold separately, and water molecules are retained during the subsequent association of the domains. This notion is further reinforced by the fact that functional domains (as revealed by biochemical assays) indeed fold independently of one another [76, 77].

It is quite evident that most large Proteins can be subdivided into several structural domains containing 100–150 residues, which correspond to a globule with a diameter of about 25 Å (Table 5.2). From a free-energy perspective, such a limitation on domain size is somewhat unexpected, since a single large globule has a significantly smaller surface-to-volume ratio than several smaller ones. A large globule should form a prominent Hydrophobic core and numerous internal Hydrogen Bonds—both of which are energetically favorable. Size restriction is likely necessary to simplify the folding process by maintaining a sufficiently short length for independently folding units.

In some cases, such as in adenylate kinase [186], There are two regions that can be termed domains, which are articulated by two chains instead of one. As can be seen from Fig. 5.14, b, the correlation between sequence-adjacent residues in these two domains differs markedly. Presumably, the domain with the stronger correlation folds first and then acts as a kind of template that facilitates the folding of the other domain.



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.