Molecular Biology: Protein Structure and Functions - Stepanov V.M. 2005
Tertiary Protein Structure
On the Relationship Between Primary and Spatial Protein Structures
The concept that the Amino Acid Sequence of a protein's polypeptide chain contains all the information necessary and sufficient for The formation of a unique Spatial Structure is one of the fundamental principles of molecular biology. However, the limits to which this principle applies and certain aspects of its Implementation warrant further comment.
First of all, this principle does not imply that any randomly generated amino acid sequence will necessarily fold into a stable, compact structure. After all, the extensive experimental data summarized by this rule pertain to evolutionarily selected primary structures, rather than arbitrary sequences of amino acid residues. At present, providing a general solution to this problem remains difficult, and it is still a subject of debate. It is widely held that a unique correspondence between The amino acid sequence and the spatial structure is a distinctive feature of only a relatively small fraction of all possible primary structures—specifically, those that have been selected over the course of Introduction/18.html">Protein Evolution.
According to an alternative view, the range of sequences capable of folding into a compact globule is much broader. It is clear, however, that far from every amino acid sequence possesses this property. For instance, residues 1–126 of staphylococcal nuclease (a protein comprising 149 amino acid residues in its peptide chain) do not form any stable spatial structure, even though their size is comparable to that of many small Proteins. Pancreatic Ribonuclease (which contains 124 amino acid residues in its polypeptide chain) loses its refolding capacity once four residues are cleaved from the C-terminus. Similarly, β-casein appears non-compact, despite containing elements of Secondary structure. Furthermore, numerous mutants of natural proteins are known in which a single amino acid substitution completely destabilizes the structure. Consequently, certain polypeptide chains are incapable of folding into a compact, unique spatial structure under normal conditions.
Nevertheless, a tendency to form a more or less compact—even if still imperfect—spatial structure seems to characterize a fairly broad range of Amino acid sequences.
At the same time, there is no absolute certainty that a given amino acid sequence strictly corresponds to one and only one spatial structure, representing a single free-energy minimum. In other words, the question boils down to whether two (or more) alternative spatial structures can exist for the exact same amino acid sequence. In most cases studied, refolding restores the native Cell/13.html">Protein Structure; there are only isolated indications that a tertiary structure differing from the native one might form during the refolding of Ovalbumin. Even if this latter observation is correct, it appears to be an exception.
Considering this, we should assume that for The polypeptide chains of naturally occurring proteins, the correspondence between the Primary Structure and the spatial structure is indeed unique in the overwhelming majority of cases. Naturally, this does not mean that a particular protein cannot adopt different Conformations under the Influence of External factors, such as the formation of a complex with a Ligand. Such a situation is entirely realistic and, in fact, typical. The issue under Discussion is different: whether a protein could form two spatially distinct forms under identical conditions.
Another important nuance of the relationship between the primary and tertiary structures of a protein must be kept in mind. While a given amino acid sequence typically dictates a single mode of spatial folding—a unique protein tertiary structure—the reverse is not true. Experience shows that a given mode of polypeptide chain folding corresponds not to just one, but to an entire family of amino acid sequences. In other words, the following relationship holds:
Class="center">family of primary structures ↔ spatial structure
Clearly, there must be certain regularities, rules, and a stereochemical code of sorts that govern the formation of a spatial structure from a given amino acid sequence. Decoding this code is a paramount objective of physical Protein Chemistry, one that has not yet been fully achieved despite several promising approaches.
Nevertheless, certain Features of the stereochemical code can be discussed. First and foremost, it is degenerate; as noted earlier, a single tertiary structure and a single mode of peptide chain folding correspond to a set of amino acid sequences. For instance, Globins—heme-containing oxygen-binding proteins, including Hemoglobins and myoglobins from various animal species, plant leghemoglobin, and so forth—share very similar spatial structures. These peptide chains differ in their Amino Acid Composition and sequence, yet all consist of approximately 140–150 residues (see Chapter 8).
Since all of them fold into a nearly identical spatial architecture of the polypeptide chain, differing only in minor details, it is evident that not all amino acid residues contribute equally to the Formation of the tertiary structure. This Conclusion is supported by numerous Site-Directed Mutagenesis experiments involving targeted Amino Acid Substitutions at specific positions along the polypeptide chain. For instance, replacing the Pro-86 residue in T4 phage Lysozyme with Gly, Ser, Cys, Leu, Asp, Arg, or His resulted in structural changes that, while sometimes extending up to 20 Å, were overall quite small in magnitude. Furthermore, all these mutant proteins were virtually identical to the wild-type protein in thermal stability and generally retained the same spatial folding pattern of the polypeptide chain.
Interestingly, their enzymatic activity changed much more significantly, even though the substitution site was located 24 Å away from the catalytic center. On the other hand, instances where the replacement of a single residue drastically affects protein stability are by no means rare. For example, replacing Ser-87 with Pro in E. coli adenylate kinase has been shown to reduce the α-Helix content from 50% to 39% and render the protein structure unstable at 40°C, leading in vivo to extensive proteolytic degradation and complete loss of enzymatic activity.
One might be tempted to think that a specific subset of amino acid residues plays a decisive role in driving spatial structure formation. However, the number of invariant residues—that is, Amino Acids that invariably occupy the same position across a sufficiently large family of evolutionarily related proteins—is usually small, sometimes accounting for only 20–30% or significantly less. For example, among the coat proteins of plant Viruses related to tobacco mosaic virus, 25 out of 158 residues (16%) are invariant. Similarly, in the carboxypeptidase family, which includes Enzymes from animals and certain microorganisms, only about 42 out of roughly 310 amino acid residues (13%) are conserved. In the globin family, the number of invariant residues is surprisingly low: just six. It would be extremely difficult to imagine that the entire set of secondary structure elements (such as the α-helices in globins) and all the ways they fold into a tertiary structure could be dictated by such a tiny number of invariant amino acid residues, even when accounting for the presence of a fairly large invariant moiety like the heme group.
Consequently, we must look for other common denominators across families of protein primary structures. A fairly consistent marker turns out to be the distribution of hydrophobic and hydrophilic amino acid residues along the sequence. Of particular importance is the obligatory presence of hydrophobic amino acids at specific positions (invariantly hydrophobic residues).
For instance, in mammalian hemoglobins, 33 positions in the primary structure are always occupied by hydrophobic amino acids, making them invariantly hydrophobic. Such a high significance of invariant hydrophobic residues is easy to understand in a first approximation: almost all of them are required to form the Hydrophobic core. Hydrophobic residues are crucial for mediating contacts between secondary structure elements during folding. Finally, the characteristic distribution of hydrophobic and hydrophilic residues is essential for proper Formation of secondary structure elements.
Naturally, invariant hydrophobic residues do not encompass the full set of features that determine how a spatial structure is formed by a given amino acid sequence. Preserving specific Sequence Motifs that dictate turns in the peptide chain is apparently also vital. Glycine residues, for instance, are frequently conserved; their inclusion imparts flexibility to the chain by lifting many of the constraints on the range of φ and ψ dihedral angles. Finally, amino acid residues that may not directly participate in building the spatial structure, yet remain indispensable for the protein's specific function (such as those entering its Active Site), are likewise conserved.
Surprisingly, in some cases, the integrity of the polypeptide chain itself is not strictly required to form the correct spatial structure. For example, ribonuclease S—produced by the Limited proteolysis of pancreatic ribonuclease using subtilisin—contains a single break in the polypeptide chain after Ala-20. Although the peptide fragment 1–20 can be physically separated from the remaining protein core (residues 21–124), mixing them back together reconstitutes active ribonuclease S, whose tertiary structure differs from that of the native enzyme only at the site of the peptide bond Cleavage.
Conversely, that same pancreatic ribonuclease (consisting of 124 amino acid residues in its polypeptide chain) loses its ability to fold once four Amino acids are cleaved from its carboxy-terminus. When a previously reduced fragment 1–120 is oxidized, its Disulfide Bonds form incorrectly, failing to adopt the pattern characteristic of the native protein. However, adding the C-terminal peptide 105–124 rescues the structure and restores the enzyme's activity. As for staphylococcal nuclease (149 amino acid residues), it has been demonstrated that mixing fragments 1–126 and 49–149 allows their mutually complementary regions (such as 1–48 and 49–149) to form a globule, whereas the redundant fragment 49–126 is excluded from the compact structure and can be digested with Trypsin if the renatured enzyme is stabilized by the presence of a substrate analog.
The refolding of a protein from two fragments, which sometimes overlap, can presumably be explained by the ability of the respective primary structure segments to independently form elements of a 'fluctuating' secondary structure. Upon fragment interaction, these elements quickly find the pathway toward stabilizing a much more robust spatial structure. Of course, one cannot expect that any arbitrary fragments generated by random cleavage will successfully refold into a unified tertiary structure; the chain break must occur at a permissive site chosen such that the fragments can, with a reasonable probability, autonomously form secondary and potentially tertiary structural elements. Protein refolding from large polypeptide chain fragments is evidently akin to complementation—the restoration of protein activity upon mixing two inactive polypeptide chains bearing Mutations at different sites.
Thus, the code governing the uniqueness of the relationship between the spatial folding pattern of a polypeptide chain (its tertiary structure) and an entire family of amino acid sequences is undoubtedly degenerate. However, the degree of this degeneracy should not be overstated, because we are dealing with the robustness of a crucial yet non-exhaustive characteristic of protein Specificity: the mode in which the polypeptide chain folds in space. This flexibility allows the same tertiary structure to serve as a scaffold for numerous spatial configurations that differ in the surface distribution of their functional groups and, consequently, exhibit distinct functional capabilities. It is worth noting that a conserved spatial folding pattern among peptide chains belonging to the same family does not imply structural identity; minor variations in detail can still occur, and these may prove functionally significant.
This phenomenon is illustrated by data on mutants of a thoroughly studied protein: lysozyme. In birds of the order Galliformes, a cluster situated near the Active Site of egg-white lysozymes is formed by one of two alternative sets of amino acid residues:
![]()
or
![]()
No other amino acid residues are found at these positions. Moreover, natural lysozymes corresponding to intermediate combinations of residues (such as Ser-40, Ile-55, Ser-91, etc.) are entirely unknown. All possible transition variants generated via site-directed mutagenesis proved qualitatively very similar to both parental variants, displaying comparable specific activities, among other properties. X-Ray Structural Analysis revealed that the protein globule is sufficiently plastic to accommodate minor alterations in side-chain architecture. For instance, when Ser-91 is replaced by Thr, the extra methyl group of Threonine partially fills a pre-existing 'cavity' in the structure while also utilizing the space freed up by the rotational shift of the Ile-55 side chain around the Cα—Cβ bond. At the same time, generalized parameters of protein globule stability, such as thermal Denaturation Temperature or thermal activity retention, do not typically fall between those of the parental enzymes in transition forms, but rather exceed the range defined by the parent proteins. Therefore, while not causing dramatic alterations in overall properties, amino acid substitutions at these positions are far from inconsequential for the enzyme's functional characteristics and cannot be regarded as entirely 'neutral'.
The degree of degeneracy of the stereochemical code linking primary and Tertiary protein structure can be estimated from experiments conducted by M. Ptashne and coworkers. Using Protein Engineering Methods, the DNA-recognizing a-helices were exchanged between two repressor proteins: 434 and 1. The fragments encoding homologous a-helical regions in these proteins are compared below:

Replacing the specified region in the 434 repressor with the sequence from the cro-repressor yielded a stable hybrid protein—repressor*—functionally similar to the cro-repressor. Thus, the transplantation of a large fragment (an a-helix) into a structurally related yet far from identical protein was quite successful, with this a-helix evidently establishing non-covalent interactions with the remaining 434 repressor structure. This result strongly Supports the degeneracy of the relationship between tertiary and primary structures. However, the seemingly symmetrical operation—transplanting the a-helical fragment of the 434 repressor into the cro-repressor—resulted in an unstable protein unable to persist within The Cell. It turned out that a leucine residue present in the cro-protein a-helix hinders its contact with the residual STRUCTURE OF THE 434 protein. Only after substituting this single residue with the isoleucine characteristic of the 434 repressor did The structure of the hybrid protein stabilize, enabling it to exist and function. While not refuting the degeneracy of the stereochemical code, these findings highlight its significant limitations.
Although the code determining the relationship between primary and tertiary protein structure has not yet been fully deciphered, attempts are underway to predict the Spatial structure of a protein from a known primary structure, as well as to solve the inverse problem—designing an amino acid sequence that ensures the formation of a Tertiary Structure of a specified type. Such a solution has so far been achieved only for relatively simple structures. Researchers have successfully engineered artificial amino acid sequences that do not exist in nature and are not close analogs of natural ones, yet are capable of forming a compact spatial structure of a predetermined type. In other words, the first steps toward de novo Protein Synthesis have been taken.
One such protein is a bundle of four a-helices connected by loops. The spatial structures of natural myohemerythritol and cytochrome c' are formed in the same manner. The sequence of the a-helical regions of this protein is Gly—Glu—Leu—Glu—Glu—Leu—Leu—Lys—Lys—Leu—Lys—Glu—Leu—Leu—Lys—Gly, and that of the connecting loops is Pro—Arg—Arg—, while the distribution of secondary structure elements along the peptide chain is as follows: Met—a-helix—loop—a-helix—loop—a-helix—loop—a-helix—COOH.
This protein, whose Gene was synthesized and cloned in bacterial Cells, was successfully isolated. It has been shown to be quite stable; for example, its unfolding at room temperature requires a 7 M guanidine hydrochloride solution.
Another de novo designed protein also consists of an a-helix bundle formed by sequences containing only Serine and leucine. It is capable of inserting into membranes and forming channels that conduct cations across a phospholipid bilayer membrane. Thus, researchers have successfully engineered a structure possessing at least a rudimentary function from scratch. A de novo ß-sheet protein has also been described and synthesized, and efforts to create de novo enzymes are currently underway.
Last update: 13/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.