Structural Biochemistry - Study Guide - E. A. Bessolitsyna 2015
Amino Acids and Proteins
Proteins
Proteins are defined as Polypeptides capable of spontaneously folding and maintaining a specific three-dimensional Structure. There is no sharp threshold or boundary that strictly separates proteins from Peptides. Indeed, a marked ability to form preferential Conformations in solution is observed even in relatively short peptides; furthermore, this capacity is essential for the function of certain peptides (such as Hormones), facilitating their interaction with cellular receptors. Nevertheless, this is merely a precursor to the precise correlation between Amino Acid Sequence and three-dimensional structure that constitutes the fundamental distinguishing feature of a protein.
The stabilization of a Spatial Structure requires a well-developed network of non-covalent interactions, which can only be achieved once the polypeptide chain reaches a certain length. Proteins are known whose polypeptide chain contains only about fifty amino acid residues. These include, for example, the pancreatic Trypsin inhibitor, epidermal growth factor, and certain bacteriophage coat proteins. However, such instances are relatively rare, and proteins most commonly contain 100 to 400 amino acid residues in a single polypeptide chain forming a globular structure.
However, the length of the polypeptide chain can be considerably greater, reaching a thousand residues or more. The so-called polyproteins are also known. They represent an even longer polypeptide chain that sequentially forms several structurally and functionally autonomous globules. Once cleaved by proteinases at specific interdomain linker sites, these globules exist as independent Enzymes. Apparently, the Biosynthesis mechanism itself does not impose significant limitations on the length of a protein's polypeptide chain.
It must be emphasized that the transition from a peptide to a spatially structured, compact protein globule is determined not by the mechanical elongation of the polypeptide chain, but by the specific sequence of amino acid residues. In other words, a randomly assembled polypeptide chain does not necessarily form a compact spatial structure spontaneously. It is highly probable that The ability to self-assemble is characteristic of a limited range of sequences, including those corresponding to natural proteins selected over the course of evolution. In any case, Amino acid sequences are known that do not fold into a compact structure.
The great importance attached to the self-assembly of spatial structure as a hallmark of proteins is due to the fact that it serves as the basis for all protein properties, most notably its biological function. The physical characteristics of a protein as a polymer are entirely determined by its ability to form a compact globule, the features of which also govern the oligomerization of many proteins. The chemical and Functional Properties of proteins depend on specific interactions of functional groups brought into close proximity within its three-dimensional structure. Consequently, The behavior of these groups in proteins differs fundamentally from their reactivity in free Amino Acids and small peptides. Accordingly, a protein can function—i.e., act as an enzyme, structural or transport protein, regulator, toxin, or inhibitor—solely because it possesses a strictly defined spatial architecture.
Just like in peptides, amino acids in a protein are linked together into a polypeptide by peptide bonds, possessing an N-terminus (HEAD) and a C-terminus (tail).
PHYSICOCHEMICAL PROPERTIES OF Proteins
The physicochemical Properties of Proteins and peptides are similar and depend on their constituent amino acids. The only difference is related to the isoelectric point, which is determined primarily by the groups of amino acid side chains, specifically those exposed On the surface. N- and C-terminal groups can be neglected. Thus, The properties of a protein are governed by The amino acid side chains exposed on its surface.
Classification of Proteins
According to their composition, proteins are classified into:
Simple proteins. Consist solely of polypeptide chains.
Conjugated Proteins. Contain a non-protein component known as a prosthetic group. Prosthetic groups may include Lipids (in Lipoproteins), CARBOHYDRATES (in Glycoproteins), phosphates (in Phosphoproteins), Hemes (in Hemoproteins), flavin NUCLEOTIDES (in Flavoproteins), Metal Ions (in Metalloproteins), and others.
According to their molecular shape, proteins are classified into:
1. Globular proteins — feature a dense, compact structure that is roughly spherical in shape. They are typically Water-soluble and readily diffuse.
2. Fibrous proteins — feature elongated, rod-like molecules and are generally insoluble (e.g., keratin, Myosin, Collagen). Fibrous proteins include Muscle contractile proteins such as Actin and myosin, microtubule proteins found in eukaryotic Cilia and flagella, and so on.
As can be seen, these classification approaches are inadequate because a single Class ends up encompassing too many diverse proteins. The most precise and widely used classification is based on function.
According to their function, proteins are classified into:
Enzymes. Virtually all reactions in the living body are enzymatic. To date, more than 2,000 enzymes have been identified.
Transport proteins are divided into two groups: transmembrane transporters and "soluble transporters". Transmembrane transporters include permeases and porins, which facilitate passive transport, and transport ATPases, which mediate Active Transport. Membrane transport systems are characteristic of both unicellular and Multicellular Organisms (since cellular transport is fundamental). In multicellular organisms, alongside membrane transport proteins that mediate exchange between The Cell and its environment, there are specialized proteins responsible for transporting substances through the Circulatory system. For instance, Hemoglobin transports oxygen, while serum albumin transports Fatty acids, hydrophobic amino acids, Steroid Hormones, etc.
Nutritive and storage proteins. Utilizing proteins as a reserve energy source is energetically unfavorable. However, this mechanism is essential for supporting the growth of offspring, as seen in seed storage proteins, egg white (Ovalbumin), milk casein, and others. These proteins serve as a source of amino acids, other molecules, or energy required to build the developing Organism. This is precisely why seeds and eggs contain an abundant supply of uniform protein, providing both a pool of amino acids for GROWTH AND DEVELOPMENT, and energy to fuel these processes, though to a lesser extent.
Contractile and motility proteins. Actin and myosin drive Muscle contraction in higher animals. Myosin also Functions as an ATPase, meaning that ATP Hydrolysis is required to power muscle activity. Tubulin is the protein that forms microtubules, while dynein, a component of microtubules, powers the movement of cilia and flagella.
Structural proteins. Provide mechanical strength and support to Tissues. Collagen and Elastin are the major proteins of Connective Tissue, ensuring the tensile strength of tendons, ligaments, and so on. Hair, Nails, and claws consist almost entirely of the protein keratin. Silk and spider webs are made of the protein Fibroin.
Protective. A highly diverse group of proteins. This category includes Antibodies, which ensure the recognition of foreign proteins and Polysaccharides and trigger defense mechanisms. Fibrinogen and Thrombin facilitate Blood clotting, protecting the body from blood loss in case of mechanical injuries. Snake venoms and microbial toxins protect against other organisms. Enzymes responsible for the transformation of foreign substances can also be assigned to this group.
Regulatory. Repressors (which regulate biosynthesis). This group also includes receptor proteins embedded in The Plasma Membrane, serving to transduce various signals from other Cells and Organs.
Other. Protein Functions are so diverse that some of them are difficult to assign to any specific group. For example, the plasma of certain arctic Fishes contains proteins with antifreeze properties, enabling them to survive sub-zero temperatures.
At the dawn of modern Protein Chemistry, Danish biochemist K. Linderstrøm-Lang proposed considering four Levels of Protein molecule Organization: primary, secondary, tertiary, and quaternary structures. This classification has become firmly established in the literature as it reflects the actual stages in The formation of the spatial architecture of protein molecules.
The Primary Structure refers to The sequence of amino acid residues in a protein molecule. It is encoded by the structural Gene of the given protein and contains all the necessary information for the self-assembly of its spatial structure. The amino acid sequence is formed As a result of mRNA Translation. However, the Introduction/19.html">Primary structure of a mature protein molecule does not always fully coincide with the direct translation product, which typically undergoes more or less significant co- and post-translational modification, or Processing, during which amino acid residues, the length of the polypeptide chain, etc., may change.
Amino acids in a polypeptide are linked by peptide bonds. Thus, the primary structure of a protein is formed by peptide bonds. The traditional representation of a peptide bond does not entirely accurately convey its electronic structure. The double bond С=0 of the carbonyl group is polarized and, in a certain sense, can be represented as a single bond with charge Separation, with the negative charge localized on the oxygen atom. Studies of the crystal structures of compounds containing an amide bond have shown that the C—0 distance in them is approximately 10% greater than that characteristic of the carbonyl group in aldehydes and ketones. Conversely, due to the Displacement of the lone electron pair of nitrogen toward carbon, the С — N bond acquires a certain degree of double-bond character and is shortened again by approximately 10%. According to L. Pauling, the electronic STRUCTURE OF THE peptide bond is represented by two Resonance structures, with structure II, featuring the carbon–nitrogen double bond, accounting for about 40% of the hybrid. As a result, the nitrogen of the peptide bond almost completely loses its basicity. Thus, the Spatial structure of the polypeptide chain can be viewed as a sequence of planar elements—peptide groups—interconnected via С-α-atoms, which act as a kind of hinges. Rotation of the planar peptide groups can occur around two single bonds connecting the nitrogen atom to the C α -atom (N — Со) and the C- α -atom to the carbonyl carbon (С α — С (0)). There are no rigid prohibitions on rotation around these bonds, although certain conformations are preferred.
Figure 68. Primary protein structure. Rectangles indicate the peptide bond
The amino acid sequence in a polypeptide determines their subsequent interaction with one another and, consequently, the higher levels of protein organization. In other words, the Structure and function of a protein molecule are dictated by its amino acid sequence. Therefore, determining the primary structure of a protein is of paramount importance.
Determination of protein Composition
Determining the amino acid sequence of proteins (protein sequencing) provides the most comprehensive information about the primary sequence of a protein, but this complex and expensive method emerged relatively late. Prior to it, there was a simpler METHOD FOR DETERMINING protein composition, which made it possible to ascertain which amino acids and in what quantities comprise the protein.
Figure 69. Determination of the Amino Acid Composition of a protein
This is not a fully exhaustive analysis, but it yields a sufficient Amount of Information. Determining the amino acid composition is essential for characterizing the studied protein, as well as the peptide fragments obtained during primary structure analysis. To cleave the peptide or protein into free amino acids, it is subjected to exhaustive hydrolysis by heating with a constant-boiling (5.7 M) HCl solution at 105°C for 24 hours. Since peptide bonds of isoleucine and valine residues are not fully hydrolyzed under these conditions, 48- and 72-hour hydrolyses are also performed.
Consequently, to obtain more reliable values, the content of valine and isoleucine is estimated by extrapolating the found values to "infinite" time, whereas the content of Serine and Threonine is determined by extrapolation to "zero" hydrolysis time. Some amino acids cannot withstand acid hydrolysis at all and are destroyed almost completely; therefore, their determination requires specially developed techniques. The resulting mixture of amino acids is analyzed using Thin-Layer Chromatography (Figure 69).
This method is based on the differing mobility of molecules on a support. The mixture of molecules is applied to a layer of the support (e.g., paper or a silica gel plate), and the edge is immersed in a solvent (usually a mixture of hydrophobic and hydrophilic Solvents). Driven by capillary forces, the solvent begins to move along the support (upward in ascending chromatography, downward in descending). When the solvent reaches the applied spot of the mixture, the molecules dissolve and move along with the support; the lower the affinity of a molecule for the solvent, the faster it settles on the support, resulting in a smaller distance from the origin (l). Conversely, the higher the affinity of a molecule for the solvent, the greater l is. The distance traveled by the solvent from the origin is designated as L. Molecular mobility in chromatography is an individual parameter for each molecule in each specific solvent system. Mobility, or Rf, is calculated using the formula:
Rf=l/L
There are specialized Rf tables for various solvent systems that can be used to identify molecules in the mixtures and hydrolyzates under study. However, it is mandatory to use mobility tables corresponding to the solvent systems actually used for chromatography; otherwise, the data will be distorted.
The protein hydrolyzate is separated by thin-layer chromatography, after which the chromatogram is visualized (using Qualitative reactions for amino acids, most commonly with ninhydrin), and then Rf is calculated or compared with the migration distances of marker amino acids. This is how the amino acid COMPOSITION OF THE protein is determined. The quantity of each amino acid in the protein is determined by eluting the amino acid from the support and quantifying it within the spot. Having this value, along with The amount of hydrolyzed protein and applied hydrolyzate, one can calculate the quantity of the amino acid in the protein. Automated analyzers, first described by W. Haining, S. Moore, and D. Spackman, are currently employed. Their designs vary considerably, but they are all based on the method described above.
Naturally, however, Methods for determining the amino acid sequence are far more informative.
Determination of the Amino Acid Sequence of Polypeptides by the Edman Method
In determining the amino acid sequence, phenylisothiocyanate (Edman's reagent) is added to the protein; the reaction results in the Cleavage of the N-terminal residue as its phenylthiohydantoin derivative (Edman Degradation). Identification of the phenylthiohydantoin derivatives is carried out using high-pressure liquid chromatography. The remaining protein is treated with Edman's reagent once again, repeating the cycle (Figure 70). This method allows the determination process to be automated. Instruments enable fully automated sequencing of polypeptides containing up to 30–40 residues (in some cases up to 60 or even 80 residues) per run. However, since proteins are quite long, they are cleaved into shorter fragments—ideally 30–40 amino acid residues in length—with the cleavage occurring not randomly, but after specific amino acid residues. Cleavage that fulfills these requirements is achieved using
Figure 70. Scheme of amino acid Sequence Determination by the Edman method
Cyanogen bromide (CNBr), trypsin, and o-iodosobenzoic acid; proteinases that recognize and cleave specific amino acid sequences are also frequently used. The resulting fragments are separated by chromatography, and their sequence is determined.
Furthermore, to determine the primary structure of a protein, researchers increasingly resort to analyzing The nucleotide sequence in the corresponding structural gene or cDNA, which requires significantly less time and generally yields accurate results. However, this approach is not always practical, as it necessitates the isolation and cloning of the gene. In addition, it fails to reveal post-translational modifications, without the identification of which protein structure research cannot be considered complete.
The number of established primary structures is growing rapidly: back in 1965, they could be counted on the fingers of one hand; by 1975, about 600 were known; in 1984, there were 2,500 sequences containing about 0.5 million amino acid residues; and by the end of 1980, these figures had risen to 14,400 and 4 million, respectively.
Proteins that perform the same function in different species (such as hemoglobin) are called homologous. Typically, homologous proteins are polypeptides of identical or nearly identical length. At many positions, the same amino acid residues, known as invariant residues, are always present. At other positions, the amino acids vary from species to species; such residues are called variable. The sum total of similarities in the sequences of homologous proteins is encompassed by METABOLISM/2.html">THE CONCEPT OF Sequence Homology; the existence of such homology implies that the hosts of these proteins share a common evolutionary origin. The number of differing residues between molecules is proportional to the phylogenetic distance between the species. For example, the Cytochromes of horses and Yeast differ by 48 out of 104 residues, those of ducks and chickens by 4 residues, chickens and turkeys are identical, while pigs, cows, and sheep are also identical. Sequence homology studies are used to construct evolutionary trees.
Before moving on to Other types of protein structure, let us recall what we mean by two important terms.
Configuration refers to the Spatial Organization of an organic molecule determined by: 1) the presence of double bonds around which free rotation is impossible, and 2) asymmetric carbon atoms with surrounding substituent groups arranged in a specific sequence. A distinctive feature of configurational isomers is that they cannot be interconverted without breaking a covalent bond. Configurational isomers can be separated.
Conformation is a term used to describe the spatial arrangement of substituent groups in an organic molecule that are capable of changing their position without bond cleavage, owing to free rotation around single C–C bonds. Different conformations cannot be separated from one another.
Secondary structure refers to the spatial arrangement of atoms in the main chain of a protein molecule within its individual segments. According to this definition, any region of a protein possesses a secondary structure. Sometimes, only those elements of secondary structure that are periodic—namely, the α-Helix and the β-sheet—are considered as such. As emphasized in the definition, the concept of secondary structure applies not to the entire protein molecule as a whole, but to individual, more or less extended regions of its polypeptide chain. Secondary structure is maintained by non-covalent Hydrogen Bonds between the partially negative oxygen of the peptide bond carbonyl group and the partially positive hydrogen of the amide group of another peptide bond.
Two Types of secondary structures should be distinguished: regular and irregular. Regular structures include the α-helix, β-sheet, and β-turn, whereas irregular structures include disordered segments.
α-Helix
In the 1950s, L. Pauling and R. Corey, based on structural data from amino acid crystals and simple peptides, examined possible periodic conformations of the polypeptide chain and concluded that the most probable structure is the one they named the α-helix (Figure 71). The Selection of this specific secondary structure was based on the following criteria:
1. Formation of a tightly packed, compact structure without voids or atom overlapping.
2. Maximum saturation of the structure with hydrogen bonds, under the condition that the geometry of the C–O···H–N Hydrogen bond is close to linear.
3. Preservation of the interatomic distances and angles characteristic of amino acids and simple peptides.
Both right-handed and left-handed α-helices can be constructed while meeting these conditions; however, the right-handed α-helix is energetically somewhat more favorable than the left-handed one when the polypeptide chain is formed by L-amino acids. Hydrogen bonds are established between C=O and N–H groups such that the carbonyl group of the $i$-th residue forms a bond with the N–H group of the $(i+4)$-th residue, the carbonyl group of the $(i+1)$-th residue with the N–H group of the $(i+5)$-th residue, and so on. The fact that the NH group, rather than the carbonyl of the $(i+4)$-th residue, is located "above" the carbonyl group of the $i$-th residue corresponds approximately to 3.6 residues per turn of the helix; negative charge is concentrated on the carbonyl oxygen, and positive charge on the nitrogen atom. The capacity for hydrogen bonding between peptide groups in the α-helix is fully utilized. In this sense, the α-helix is saturated and internally closed.
It is incapable of interacting with other secondary structure elements via hydrogen bonding between its peptide groups. Non-covalent contacts involving the side chains of amino acids within the α-helix play an important role in shaping the spatial structure of the protein. Crucially, the side chains of amino acids separated by two or three residues in the polypeptide chain are brought into close proximity on The surface of the α-helix. If hydrophobic amino acids occupy these positions, they form a distinct hydrophobic ridge.
The interaction between such ridges serves as a mechanism for packing α-helices within the Tertiary Structure of a protein. Similarly, hydrophilic surfaces—in particular, charged ones—can also be formed. Therefore, an α-helix can be purely hydrophobic, which is typical for transmembrane regions of proteins, or amphiphilic, meaning it possesses both hydrophobic and hydrophilic sides.
Figure 71. Structure of the α-helix. A — diagram of amino acid interactions in the α-helix; B — structure of the α-helix
The length of α-helical regions in globular proteins is relatively short, typically spanning 5–15 amino acid residues and rarely exceeding 3–4 turns of the helix; in fibrous proteins, they are much more extended. Kinks in the helix are sometimes observed, usually at sites where Proline residues are incorporated, which disrupt the hydrogen bonding network. At these points, the helical axis deviates by 20–30°.
Not every polypeptide will form an α-helix. For instance, if a chain contains several consecutive glutamate or aspartate residues at neutral pH, the polypeptide will not adopt an α-helical conformation due to the mutual repulsion of the negative carboxyl groups. Similarly, an α-helix will not form in the presence of numerous closely spaced Lysine residues carrying a positive charge. Closely spaced amino acids with negatively charged side chains also repel each other, disrupting the α-helix. Histidine, Tryptophan, leucine, and Arginine hinder α-helix formation because of their bulky R-groups. In proline, the nitrogen atom is part of a rigid ring, which precludes rotation around the N–C bond. Furthermore, the nitrogen atom of the peptide bond formed involving proline's amide group lacks a hydrogen atom. Consequently, an intrachain hydrogen bond cannot form. As a result, whenever one or more proline residues are incorporated into the chain, the regular structure is disrupted, leading to the formation of a loop or bend.
β-Structure
Unlike the α-helix, the β-structure is stabilized by interactions between adjacent segments of the polypeptide chain—that is, more distant amino acids that may even belong to different polypeptide chains (Figure 72). These segments can run in the same direction (N-termini interacting with N-termini), forming a parallel β-structure, or in opposite directions (N-terminus interacting with C-terminus), forming an antiparallel β-structure. The NH group of a given residue in segment A forms a hydrogen bond with the carbonyl group of the $i$-th residue in polypeptide chain B, while the carbonyl group of that same amino acid residue in chain A interacts with the NH group of the $(i+2)$-th residue in the parallel chain B. Importantly, the $(i+1)$-th residue of chain B is effectively skipped, and its NH and carbonyl groups do not participate in interaction with chain A. Similarly, far from all C=O and NH groups of chain A participate in interaction with chain B.
Thus, two segments of a parallel β-structure are linked by hydrogen bonds, leading to the formation of large 12-membered rings. It should be emphasized once again that in each of the chains, only half of the groups capable of forming hydrogen bonds participate in interaction with the neighboring chain. Consequently, each of these polypeptide segments can form an identical system, thereby interacting with the next fragment of the polypeptide chain, while its remaining free peptide groups interact with the subsequent fragment, which causes the sheet to expand.
Unlike in the helix, the radicals do not participate in the Formation of the β-sheet; instead, the amino acid side chains project alternately above and below the surface of the β-sheet, resembling "nap on a carpet." It should be noted that the surface of a β-pleated sheet is rarely flat; more often, it is twisted to the left when viewed perpendicularly to the direction of the strands. The angle between adjacent chain segments is approximately 25°. The repetitive occurrence of this motif causes the sheet to twist into a staircase-like structure. "Barrels" have also been described, which form when an extended β-sheet rolls up and the terminal segments of the polypeptide chain close in on each other, establishing a hydrogen-bonding network characteristic of the β-structure. Obviously, the β-structure in such cases loses its local character, sometimes permeating the entire protein globule (thus extending beyond the definition of secondary structure given at the beginning of the chapter). Given this, the term suprasecondary structure is applied to extended β-pleated sheets that play a crucial role in forming the tertiary structure of a protein.
Figure 72. Structure of the β-sheet
β-Turn
Both the α-helix and the β-structure are typically represented in globular proteins by relatively short segments; therefore, a significant portion of a protein's secondary structure consists of various loops that allow the polypeptide chain to reverse direction. The most economical structural element that enables a polypeptide to turn by 180° using only three peptide groups is called a β-turn (Figure 73). The β-turn is stabilized by a single hydrogen bond. Its formation can be hindered by bulky amino acid side chains, making the incorporation of a Glycine residue preferable. Notably, the β-turn almost always resides on the surface of the protein globule, which is why it frequently plays a vital role in its interactions with other molecules, such as IMMUNOGLOBULINS.
Figure 73. Structure of a β-turn
Unordered Chain
Unordered fragments are most frequently found at the C-terminus, though they also occur within the chain itself. These regions lack a fixed, predefined structure, adopting one only as part of the tertiary structure. On the one hand, this structure is relatively unstable; on the other hand, such loops very often participate in the formation of functional protein centers, such as catalytic sites in enzymes. Furthermore, the disorder of an individual fragment is compensated for by complete ordering within the protein's overall tertiary structure.
The Emergence of the term "supersecondary structure" stems from the fact that as data on protein structures accumulated, it became clear that certain elements of secondary structure no longer fit neatly into that category. For instance, a β-sheet of sufficient width can bend to form barrel-like or blade-like structures. Such architectures "permeate" the entire protein globule, yet are still formed by a portion of the polypeptide chain, shaping something akin to a central ("backbone") structure of the protein globule. Consequently, these constructs are classified as secondary structures, but their scale warrants the prefix "super-"
Figure 74. Supersecondary structure of a collagen protein. A — fibril organization; B — bonding structure between lysine residues
A second such example emerged during The Study of collagen structure. In higher animals, connective tissue makes up a significant portion of the organism, with its three main components being collagen, elastin, and Proteoglycans. Collagen and elastin are prime Examples of fibrous proteins, forming the two primary types of connective tissue fibers: collagenous and elastic.
Collagen features a unique helix found in no other proteins (unlike α- and β-helices, which appear at least in isolated segments of globular proteins). Its fibrils consist of repeating polypeptide units of tropocollagen, aligned along the fibril in parallel head-to-tail bundles. Each molecule is composed of three polypeptide chains tightly wound into a triple-stranded rope, resulting in a rod-like molecule measuring 300 nm in length, 1.5 nm in thickness, and with a Molecular Weight of approximately 300,000 Da. The three chains are of equal length, each containing about 1,000 amino acid residues. In some collagens, all three chains are identical; in others, two are identical and the third differs. Each individual chain also forms a helix distinct from α- and β-structures. A rigid, curved conformation is maintained by a high content of proline and hydroxyproline. Additionally, the chains are linked by transverse Hydrogen bonds and unusual covalent cross-links between lysine residues (Figure 74).
Collagen is virtually inextensible due to its exceptionally tight coiling and abundant cross-linking, the number of which increases with age. Elastin, a highly elastic connective tissue protein, contains a high amount of lysine and very little proline, forming a specialized type of helix. It is characterized by alternating helical regions rich in glycine with shorter segments of lysine and Alanine. Four lysine residues enzymatically combine to form a single desmosine molecule.
Tertiary structure refers to the three-dimensional spatial arrangement of all atoms within a protein molecule, excluding interactions between this globule and neighboring globules or subunits. Often, the concept of tertiary structure is narrowed to focus on its most stable element: the specific folding pattern of the polypeptide chain in space characteristic of a given protein. Notably, for large groups of evolutionarily related proteins—which may differ significantly in primary structure and, consequently, in the spatial distribution of all atoms—the polypeptide chain folding pattern remains largely invariant. This demonstrates that the accepted simplification is justified, as it captures the essential features of tertiary structure.
Tertiary structure is the foundation of protein function, requiring precise spatial organization of large ensembles built from numerous amino acid residues and their side chains. These ensembles form active enzymatic sites, binding domains for other biological molecules, effector centers, and so forth. Consequently, disruption of a protein's tertiary structure (Denaturation) inevitably leads to the loss of its functional capacity.
The formation of tertiary structure involves a multitude of covalent and non-covalent bonds between the radicals of amino acid residues and elements of peptide bonds.
Covalent bonds include disulfide bridges, which provide strong stabilization due to their high bond energy. Disulfide Bonds form via The oxidation of Cysteine residues that are brought into close proximity within the protein's spatial structure, converting them into cystine. As previously mentioned, this is the only type of covalent bond that, In addition to the non-covalent interactions discussed above, participates in stabilizing the tertiary structure. Thus, the presence of disulfide bonds is by no means an absolute requirement for protein stability, and many proteins—including quite stable ones—lack them entirely. For example, Ribonuclease is one of the most stable proteins, retaining its function even after ten minutes of boiling. Apparently, the stabilization of protein structures by disulfide bonds is primarily due to the fact that, by remaining intact in the denatured protein, they drastically restrict the number of possible configurations of the unfolded polypeptide chain, thereby lowering its Entropy.
There is a greater variety of non-covalent interactions involved in tertiary structure formation, and consequently, they provide the primary folding and stabilization of the protein molecule.
Chief among these are hydrogen bonds, which represent Electrostatic Interactions between a partially positively charged atom of one group and a partially negatively charged atom of another. These bonds can form between the carbonyl group of one amino acid residue and the amide of another; these groups are elements of the peptide bond and participate in secondary structure formation.
During tertiary structure formation, hydrogen bonds also form between peptide bond elements within unordered loops, supplemented by hydrogen bonds between partially charged side-chain groups of non-ionic hydrophilic amino acids, thereby increasing the overall number of hydrogen bonds. The second type of bond, similar in nature to hydrogen bonds, is the ionic bond. These are electrostatic interactions between ionized groups in the side chains of acidic and basic amino acid residues.
The final group of non-covalent bonds comprises hydrophobic or Van der Waals interactions, which involve the overlapping of hydrophobic amino acid radicals. Much like in membranes, these radicals are essentially shielded from water beneath the hydrophilic fragments of the polypeptide, interacting exclusively with one another.
Given that tertiary structure is The basis of protein function and considering the immense Diversity of proteins, it can be said that The structure of each individual protein is unique.
X-ray crystallography of polymer molecules is used to determine protein tertiary structure, and its principle is as follows. A beam of X-rays is directed at a protein crystal, where the rays collide with atoms and are deflected from the main axis by a specific angle. Passing through the crystal, the deflected rays expose a photographic film; depending on the atom involved in the collision, the reflection angles of the deflected rays will either reinforce or cancel each other out—a phenomenon known as diffraction. Taking diffraction into account, along with the spots on the X-Ray Diffraction pattern and the known deflection angles, researchers determine the positions and types of atoms within the crystal (which is precisely why the protein must be in crystalline form, to hold all atoms fixed). From these derived data, the spatial coordinates of the atoms in the molecule—that is, the tertiary structure—are established. While this process is highly complex, the structures of numerous proteins have now been elucidated. As mentioned earlier, each of these is unique; studying the structure of every known protein is impossible, and even specialists focus strictly on specific groups of proteins. Nevertheless, General Trends in protein structure formation can be identified.
Most protein molecules exist in an aqueous environment, where hydrophilic Amino acids can form interactions with water. If all amino acids in a protein were hydrophilic, interactions with water would prevail over those with Other Amino Acids, making the formation of a stable globule impossible (Figure 75). Consequently, a portion of the amino acids is hydrophobic; these are shielded from water beneath the hydrophilic amino acids, creating a dense, compact hydrophobic core. Short-range van der Waals interactions facilitate this process, causing the molecule to fold into a roughly spherical shape. This geometry offers an optimal surface-area-to-volume ratio, making it energetically favorable. The hydrophobic core, which does not necessarily approach a perfect spherical configuration but remains more or less compact, typically comprises 20–30% of the total amino acid residues. It is particularly rich in bulky residues such as leucine, isoleucine, phenylalanine, and valine. The hydrophobic core is surrounded by an outer hydrophilic shell that maintains contact with water. This shell contains hydrophilic amino acid residues that interact both with each other and with water molecules, forming the protein's Hydration shell of bound water. However, alongside these, a significant number of hydrophobic amino acid side chains—up to half of their total content—are also present. Frequently, distinct hydrophobic patches or "domains" are formed, giving the protein surface a mosaic character that alternates with strictly hydrophilic regions. Such exposed hydrophobic patches can have functional significance: they form hydrophobic binding sites for substrates and other ligands, participate in Protein-Protein Interactions (particularly in stabilizing quaternary structure), and large Hydrophobic surface regions are also characteristic of integral Membrane Proteins.
Figure 75. Evolutionary Development of the protein core and shells. Dark grey indicates the hydrophobic core; light grey represents the hydrophilic surface; black denotes the hydration shell
At first glance, the characteristic spatial organization of proteins—the formation of a hydrophobic core and a mosaic surface containing both hydrophilic and hydrophobic elements—appears to limit the size of the globule, since as volume increases, the strictly hydrophobic core would make up a progressively smaller fraction. To some extent, this is true, but the limitation applies only to the size of a structure organized in this specific manner, not to the molecule as a whole. Indeed, starting at a molecular mass of approximately 14–16 kDa, there is a discernible tendency for protein molecules to form from two (or more) independently folded globules, each possessing its own hydrophobic core. Such globules—domains—are formed by distinct segments of the same polypeptide chain (Figure 76).
Domains in proteins are regions within the tertiary structure that exhibit a certain degree of structural autonomy. This autonomy is often so pronounced that domains can maintain and even form a spatial structure independently of other PARTS OF THE protein molecule. In many cases, domains can be separated by subjecting the protein to Limited proteolysis. Within the biological function of a given protein, domains frequently—though not always—perform their own distinct tasks, in which case the structural autonomy of the domain is complemented by functional autonomy.
For example, the nucleotide-binding domain of dehydrogenases, which features the same polypeptide chain folding pattern regardless of the specific function of a given enzyme, is responsible for interacting with one of the reaction substrates—the coenzyme NAD or NADH. The amino-terminal domains of Blood Coagulation SYSTEM enzymes provide binding to Membrane Lipids and other proteins, while the amino-terminal domains of immunoglobulins form the antigen-binding site. The zinc finger domain, which is a two-chain β-sheet with a β-turn, is additionally stabilized by a zinc ion. The zinc finger serves as a DNA-binding domain in eukaryotes.
A unique sequence of amino acid residues forms a unique set of bonds with the groups of nitrogenous bases exposed in the major groove; depending on the DNA nucleotide sequence, the set of exposed groups will vary, and consequently, the set of amino acid residues in the "finger" will differ. All this ensures binding to strictly defined nucleotide sequences in DNA. In prokaryotes, this function is performed by the "helix-turn-helix" domain: one of the α-helices interacts with DNA following the same principle as in the "zinc finger," while a short loop or turn ensures a 900 rotation of the other helix, which is rich in basic amino acids that interact with the sugar-phosphate backbone of DNA.
Most commonly, domains are formed by the autonomous folding of sequentially arranged Regions of the polypeptide chain, although cases are known where the spatial structure of a domain is formed by two segments of the primary protein structure that are far apart from each other.
Figure 76. Examples of the structure of certain domains. A — Pyruvate kinase domain (β-barrel type), B — "zinc finger"; C — "leucine zipper"
Naturally, the formation of the tertiary structure also involves the interaction of secondary structure elements. Based on the ratio and interaction of these elements, the following main spatial structure types are distinguished:
1. α-Helical proteins, constructed predominantly of α-helices. The simplest structural motif in such proteins is a bundle of four α-helices whose axes are more or less parallel. Myohemerythrin and the H-subunit of ferritin possess precisely this structure. Globins are typical α-helical proteins.
2. β-Proteins, formed predominantly by β-pleated sheets. In most such proteins, the β-sheets are superimposed at a slight angle or are orthogonal. Examples of such proteins include Pepsin and other aspartyl proteinases, superoxide dismutase, and a large family of Metabolite Carrier Proteins (e.g., fatty acids).
3. α/β-Proteins, in which α-helices and segments of β-structure alternate. This typically leads to the formation of a central β-pleated sheet surrounded on both sides by α-helices that shield it from water. This type of spatial structure is very widespread. It includes the NAD-binding domains of dehydrogenases, Triosephosphate isomerase, and about a dozen other proteins structurally related to this enzyme.
4. (α + β)-Proteins, in which α-helices and segments of β-structure do not alternate, but rather group with their own kind such that part of the molecule acquires a purely α-helical spatial folding (type 1, see above) and another part acquires a purely β-type folding (type 2), often with elements of (α/β)-structure. Chicken egg white Lysozyme belongs to this type.
Unique structures also exist.
Denaturation refers to a significant change in the secondary and tertiary structure of a protein—that is, the disruption and disordering of the non-covalent interaction system without affecting its Covalent Structure. Denaturation is generally accompanied by the loss of the protein's functional properties and its inactivation. However, inactivation by itself cannot serve as a reliable criterion for denaturation. Conformational transitions in a protein, in which one cooperative system of non-covalent interactions rearranges into another, should not be classified as denaturation.
The fundamental difference is that in the latter case, both states are ordered, whereas the hallmark of denaturation is precisely the loss of order, which leads to an increase in the entropy of the system. However, cases apparently do occur where a protein undergoes partial denaturation, for example, due to the loss of spatial organization by one of its constituent domains or the disruption of the non-covalent interaction system in some region of supersecondary structure. A qualitative consideration shows that protein denaturation can be caused by A number of factors.
Based on the causative agents, two types of denaturation are distinguished: Physical and Chemical. Temperature is a physical denaturing agent. Thus, an increase in temperature leads to an increased contribution of the entropy factor, causing thermal denaturation, which generally occurs abruptly. The denaturation temperature of proteins varies and depends significantly on other conditions; for instance, many proteins are markedly stabilized by Calcium Ions. Certain proteins are distinguished by their thermostability. This is particularly characteristic of proteins from thermophilic organisms that have adapted to life at elevated temperatures. For example, Thermolysin, a proteolytic enzyme secreted by thermophilic bacilli, retains activity up to 80°C. Denaturation caused by temperature is irreversible, meaning that restoring the optimal temperature does not cause the protein to recover its structure.
Chemical denaturation occurs under The Influence of Reagents that disrupt non-covalent interactions, primarily the hydrogen bonding system. This happens because the tertiary structure of a protein contains several types of bonds and, consequently, several mechanisms.
The first mechanism is competition for hydrogen bonds. The denaturing agent contains partially positive and negative charges that form hydrogen bonds with the protein groups. Instead of bonding with each other, the protein groups form bonds with the agent, and the globule falls apart. This is how high concentrations of urea (6–8 M), guanidine hydrochloride, and guanidine isothiocyanate act.
The second mechanism is the alteration of group charges. This occurs during pH changes: upon A change in the medium's pH, chemical groups are either protonated or deprotonated. As a result of interactions between charged groups being disrupted by the charge change, the protein molecule denatures. Denaturation using acids and alkalis proceeds via this mechanism.
The third mechanism is the destruction of the hydrophobic core. This is how denaturation using organic solvents occurs. Organic solvents are hydrophobic, and when a protein enters them, the surface hydrophilic amino acids can be said to "self-isolate," whereas the hydrophobic amino acids of the core begin to interact with the solvent rather than with each other, effectively turning the molecule inside out. However, if an organic solvent is properly chosen, the hydration shell can be preserved, meaning the bonds and environment of the protein will not be disrupted, and denaturation will not occur. This allows organic solvents to be used for the isolation and Separation of proteins.
The fourth mechanism combines the first and third. A charged group competes for hydrogen bonds, while the hydrophobic part of the molecule destroys the Hydrophobic bonds of the core. Ionic detergents belong to such agents, the most well-known being sodium dodecyl sulfate (SDS); soaps effect denaturation through this same mechanism.
Protein Renaturation
Under certain in vitro conditions, protein denaturation can be reversed by transitioning from an unfolded polypeptide chain to a compact globule with a well-defined spatial structure. This process, termed renaturation, models—albeit imperfectly—the folding of a polypeptide chain into a globule during translation in Protein Biosynthesis. As is well known, the spatial structure of a protein is determined by its primary structure. The familiar one gene–one protein dogma is essentially equivalent to the assertion that a genetically determined amino acid sequence is sufficient to unambiguously predetermine its folding into the tertiary structure characteristic of a given protein. Naturally, in the general case, this Conclusion is valid only for conditions close to those existing within a given cell during The biosynthesis of that protein, which may include a specific pH range, the presence of certain ions (such as calcium ions), and Cofactors intrinsic to that protein, such as Coenzymes, heme, etc.
Provided these conditions are met, the aforementioned balance of factors governing the Stability of the protein globule will be achieved—that is, thermodynamic Prerequisites for the formation of the native structure (protein renaturation) will be established. This principle was confirmed by successful in vitro protein renaturation experiments first conducted by C. Anfinsen and coworkers in the 1960s using pancreatic ribonuclease and lysozyme, and subsequently on a number of other objects. However, chemical denaturation is not reversible for all proteins; many proteins cannot be renaturated, which is most likely due to the inability to fully replicate the required folding conditions, or the presence of natural factors unaccounted for by the researcher.
It has been established that during the formation of a protein's spatial structure in vivo, the polypeptide chain is not left to itself, but instead interacts with a host of specialized proteins termed chaperones, whose function is to ensure the rapid Discovery of the correct spatial structure. A number of chaperone families are known, being particularly abundant among the so-called heat Shock proteins. The latter are named thus because they are synthesized by cells in large quantities in response to conditions unfavorable for the efficient folding of protein tertiary structures—specifically, elevated temperatures. However, they are also produced and function under normal conditions. One family of chaperones comprises the so-called 70 stress proteins of Eukaryotic cells, which have a molecular mass of about 70 kDa. They form complexes with polypeptide chains that have not yet completed folding, preventing their interaction with one another and undesirable aggregation. These proteins consist of two interacting domains. One domain forms a complex with regions of the unfolded polypeptide chain, while the other binds ATP and is capable of cleaving it, acting as a slow-acting ATPase. Upon ATP hydrolysis, the protein transitions into a different conformational state, and its complex with the polypeptide chain dissociates. In other words, the chaperone does not fold correctly, but rather unfolds what is incorrect; the criterion is the amount of energy supplied to the system during ATP hydrolysis: if it was enough for unfolding, then it was folded incorrectly; if not enough, then correctly.
Protein polypeptide chains destined for Transport from the Cytoplasm are maintained in an unfolded state—most suitable for membrane translocation—through complexation with 70 stress proteins. The lifespan of such complexes is presumably determined by the time required for the intramolecular hydrolysis of ATP bound by the 70 stress proteins. It remains unclear whether these proteins directly catalyze the Formation of secondary or tertiary structure itself. It is possible that their role is limited to preventing intermolecular interactions and the aggregation of yet-unfolded polypeptide chains that are unfavorable to this process. Another widely distributed group of proteins involved in tertiary structure formation comprises the so-called chaperonins—proteins encoded by the GroEL and GroES genes in E. coli. GroEL proteins are composed of subunits with a molecular mass of about 60 kDa that form a distinctive quaternary structure built from two stacked rings of seven subunits each. It is hypothesized that a partially folded polypeptide chain rests on the surface of such a ring. Its individual segments, bound to the chaperonin subunits with varying affinities, can be released from the complex and form their characteristic secondary structure without Interference from neighboring segments or other polypeptide chains held in the bound state. Under such conditions, the formation of regular secondary structure elements proceeds as if in a sequential manner. Upon completion of this process, the complex dissociates, which depends, as in the case considered above, on the hydrolysis of ATP bound to the GroEL protein. GroES proteins with a molecular mass of 10 kDa associate with the GroEL protein and somehow regulate its ATPase activity and, consequently, the lifespan of the complex.
The quaternary structure refers to the spatial arrangement of interacting subunits formed by individual polypeptide chains of a protein.
The formation of the quaternary structure involves not the peptide chains themselves, but rather the globules formed by each of these chains independently. Thus, the concept of quaternary structure applies to an ensemble of globules. The interaction between the latter is strong enough for the ensemble to act as a single molecule; at the same time, each of the associated globules—the subunits—maintains considerable autonomy, which is generally much more pronounced than the autonomy of a domain within a tertiary structure. However, cases are known where two or more polypeptides make up a single globule. As a rule, this is the result of limited proteolysis—the localized cleavage of an originally intact polypeptide chain (which formed a globule According to the usual rules of tertiary structure formation) into separate segments. Such proteins, naturally, should not be classified as having a quaternary structure (Figure 77).
An example is Insulin. Its monomer is built from two peptide chains, A and B, containing 21 and 30 amino acid residues, respectively.
In some proteins, the polypeptide chain folds to form several globules (domains), between which a system of non-covalent interactions is established. Subsequent proteolytic cleavage of the chain segments connecting the domains makes them fully autonomous, turning them into subunits that form the quaternary structure. Sometimes the term "aggregate" is used in literature as a synonym for "quaternary structure," which is difficult to agree with, since the latter implies a very high level of organization—the association of subunits into a molecule stabilized by a system of non-covalent interactions. Similarly, there is no reason to classify supramolecular (e.g., multienzyme) complexes or extended structures, such as phage coats or protein crystalline inclusion bodies in certain bacilli, as quaternary structure, although the mechanisms of their formation have much in common with the formation of quaternary structure.
The quaternary structure is the final level of protein molecule organization and is optional—up to half of known proteins lack it. The boundary between proteins with and without a quaternary structure is not always entirely definite. Some proteins dissociate relatively easily into subunits, in which one of the subunits (C — catalytic) is responsible for the actual enzymatic activity and catalyzes the Transfer of phosphate from ATP to the protein, while the other is regulatory (R). In the absence of cyclic AMP, the latter is bound to the C-subunit in an R–C complex and inhibits it. Upon complexation with cAMP, the quaternary structure dissociates, and the C-subunit becomes capable of phosphorylating protein substrates. A prime example of a heteromeric protein is RNA polymerase. In homomeric proteins, the subunits are identical. Similar to them are strictly heteromeric proteins whose subunits are not identical, but are sufficiently similar both in the folding pattern of the polypeptide chain and in function. The frequencies with which proteins composed of two, three, four, etc., subunits occur vary widely. For instance, in a more or less random sample of about 200 proteins with a molecular weight not exceeding 300 kDa, there were 102 dimers, 58 tetramers, and 23 hexamers, whereas there were only 9 trimers, no pentamers, and 3 octamers. Thus, the vast majority of proteins with a quaternary structure are dimers, tetramers, and hexamers, with the latter occurring in proteins with a molecular weight greater than 100 kDa. The geometry of symmetric dimers is obvious. The same applies to trimers, which are quite rare yet occur in nature. In particular, they can play a vital role in the formation of transmembrane channels, since a bundle of three subunits naturally forms an internal channel. In tetrameric proteins, subunits may be positioned at the corners of a square, which is rare, or occupy the vertices of a tetrahedron. The latter type of quaternary structure, in which each subunit interacts with the other three with varying strengths while the molecule as a whole remains extremely compact, is observed especially frequently.
Hexameric proteins are characterized by octahedral packing, while flat hexagonal structures are much less common. Subunits forming a symmetric quaternary structure can be identical, but may also differ while remaining homologous, evolutionarily related proteins that share the same spatial folding pattern of the peptide chain. In such cases, the formation of the quaternary structure has distinctive features. If, for example, a tetrameric protein is formed by structurally homologous, similar subunits a and b, two situations are possible depending on the degree of their divergence. Other sets of subunits are unstable and practically undetectable. A typical example is human Lactate dehydrogenase, built from four subunits. The latter can be of two types: M- (from muscle) subunits, which predominate in smooth Muscles, and H- (from Heart) subunits, predominantly synthesized in the myocardium. For lactate dehydrogenase, the following set of multiple forms is possible: H4, M3H1, M2H2, M1H3, M4. Note that multiple forms whose differences are genetically determined rather than caused by post-translational modifications or protein damage are conventionally called isoforms.
The difference in the primary structure of the subunits affects The ratio of cationic and anionic groups, leading to differences in the charge of the isoforms and making it easy to separate them by Electrophoresis. Such Analysis of the isoenzyme composition of blood lactate dehydrogenase makes it possible to monitor the release into the blood of lactate dehydrogenase from certain organs, particularly The Heart, during The Development of relevant pathologies, such as myocardial necrosis. In the latter case, the content of isoforms enriched in the H-subunit increases. This approach has found application in medical Diagnostics and biochemical genetics.
Considering that quaternary structure is not essential for protein function—many proteins function perfectly well without it—quaternary structure offers certain distinct advantages:
An increase in activity within a single compact structure; in other words, the emergence of quaternary structure can be viewed as an essential evolutionary mechanism for domains, where domains ultimately become autonomous and turn into separate subunits. These complexes allow for the rapid execution of various sequences of metabolic reactions.
The capacity for regulation.
Interaction coupled with activity changes. (cooperativity of protomer conformational changes)
Cooperative conformational changes are also characteristic of individual subunits of Oligomeric Proteins and the protein globule as a whole. Let us consider this using hemoproteins, specifically hemoglobin and Myoglobin, as an example. Myoglobin is a monomeric protein, whereas Hemoglobin consists of four protomers of two types. The primary structure of $\alpha$- and $\beta$-type hemoglobin protomers differs in about half of their residues (out of 141 and 147, respectively). The primary structure of myoglobin differs even more significantly. At the same time, the secondary and tertiary structures of hemoglobin and myoglobin protomers are very similar, and they perform analogous functions based on the ability to reversibly bind oxygen. The prosthetic group of these proteins—heme—is a planar molecule containing four pyrrole rings with an iron atom coordinated at the center.
Heme is bound to the protein moiety (globin) via hydrophobic interactions between the pyrrole rings and hydrophobic amino acid side chains. In addition, there is a coordination bond between the iron atom and the imidazole ring of one of the histidine residues in the globin. An additional coordination bond allows an oxygen molecule to attach to the iron atom, resulting in the formation of oxyhemoglobin or oxymyoglobin, while the oxidation state of the iron remains unchanged.
Figure 77. A — structure of heme; B — tertiary structure of myoglobin; C — Quaternary Structure of hemoglobin
The pyrrole rings of heme lie in a single plane, whereas the iron atom protrudes slightly from this plane. The binding of oxygen "flattens" the heme molecule, causing the iron atom to move into the plane of the pyrrole rings. Because the iron is linked to a histidine residue of the globin, this displacement pulls a segment of the polypeptide chain, thereby altering the protein conformation. In myoglobin, conformational changes are restricted to a single polypeptide chain. In contrast, hemoglobin contains four protomers, each possessing a heme group capable of binding oxygen:
Hb ↔ HbO2 ↔ Hb (O2) 2 ↔ Hb (O2) 3 ↔ Hb (O2) 4
The binding of the first oxygen molecule alters the conformation of the protomer to which it attaches. Because this protomer is linked to the other three, the conformation of the remaining protomers also changes (Figure 78). This phenomenon is termed cooperative conformational transition in protomer subunits. A consequence of the structural shift following the binding of the first oxygen molecule is an increased oxygen affinity in the remaining three protomers. The binding of a second molecule further facilitates oxygen uptake by the remaining two vacant heme groups. Consequently, hemoglobin's affinity for the fourth oxygen molecule is approximately 300 times greater than for the first. As a result, the oxygen saturation curve of hemoglobin plotted against partial pressure is a sigmoid (S-shaped) curve, rather than the hyperbolic curve characteristic of myoglobin.
Hemoglobin binds oxygen from alveolar air, where the partial pressure is close to 100 mm Hg, and hemoglobin saturation here is nearly 100%. Oxygen is released at a partial pressure of about 40 mm Hg, where hemoglobin oxygen saturation drops to roughly 75%, after which the cycle repeats. These curves clearly demonstrate that myoglobin would be unsuited for this transport role, as its oxygen saturation remains near 100% at both partial pressure levels. The physiological function of myoglobin is to serve as an intermediate in Oxygen transport to Mitochondria (the primary oxygen consumers) and to maintain a local oxygen reserve in Muscle tissue (which is particularly vital for diving mammals). Cooperative Conformational Changes in protomers represent a critical regulatory mechanism, notably for allosteric enzymes. Incidentally, this same example highlights two additional features. Heme within hemoglobin and myoglobin exhibits the ability to bind oxygen with high Specificity. This specificity arises from a mechanism distinct from structural complementarity, depending primarily on the chemical reactivity of the iron atom; nevertheless, polypeptide context strongly modulates this specificity. Conversely, heme embedded in other hemoproteins, such as cytochromes, cannot bind oxygen and instead functions as an electron carrier, with the oxidation state of the heme iron changing during electron uptake and transfer. A second notable feature is that deoxyhemoglobin protomers can selectively bind 2,3-diphosphoglycerate present in erythrocytes, with the diphosphoglycerate binding site distinct from the oxygen-binding site. Diphosphoglycerate bears 5 negative charges distributed among 5 oxygen atoms. Its molecule fits snugly into a cavity formed by the two $\beta$-subunits, where it interacts simultaneously with 7 cationic groups of these chains (the $\alpha$-amino groups of Val-1, the imidazole rings of His-2 and His-143 in both chains, and the amino group of Lys-82 in one chain). This binding stabilizes the deoxy conformation and reduces oxygen affinity. In the absence of diphosphoglycerate, HbA is 50% saturated with oxygen at a partial pressure of 12 mm Hg, whereas in its presence, 50% saturation requires 50 mm Hg. This is a classic example of the so-called allosteric ("other site") mechanism of protein activity regulation. The ability of an effector to influence a structurally distant functional center of a protein relies on the plasticity of the quaternary structure in response to relatively subtle structural perturbations. Evolution has fine-tuned this mechanism to match specific environmental pressures. For instance, in high-altitude llamas living under oxygen-deficient conditions, His-2 is replaced by asparagine, which lacks a positive charge, whereas in fetal hemoglobin, His-143 is replaced by serine. In both cases, the outcome is a lower affinity for diphosphoglycerate and, consequently, a higher affinity of hemoglobin for oxygen. In the first instance, this allows adequate blood oxygenation in hypobaric environments; In the second, it enables the fetus to "extract" oxygen from the maternal blood. Domain-structured proteins share structural similarities with oligomeric proteins. They likewise contain largely independent globules known as domains, except that these globules are formed by a single continuous polypeptide chain.
Figure 78. Cooperative interaction of protomers. A — mechanism of oxygen binding by heme; B — oxygen saturation curves for hemoglobin and myoglobin as a function of oxygen partial pressure
Domains may also be connected by two peptide loops. These concepts help clarify the mechanisms underlying point Mutations. Point mutations involve the substitution of a single nucleotide in a gene, altering The Genetic Code and, consequently, substituting a specific amino acid in the polypeptide chain. The functional consequences of such substitutions vary widely and depend upon:
1. Differences in the physicochemical Properties of the R-groups (Hydrophobicity, charge, size, chain rigidity, and the ability to form hydrogen or disulfide bonds). Depending on these factors, significant alterations in secondary, tertiary, or quaternary structure may (or may not) occur.
2. The functional "criticality" of the specific chain segment where the amino acid substitution takes place, with active sites and hinge regions being the most vulnerable.
The substitution of just a single amino acid at position 6 of the hemoglobin polypeptide sequence (glutamate to valine) results in a severe pathology known as Sickle cell anemia. A "sticky" hydrophobic patch appears on the globule surface, causing deoxyhemoglobin A molecules to aggregate into abnormally long, fibrous polymers that distort erythrocytes into a characteristic sickle shape. The Physical Properties of a protein as a polymer are entirely determined by its capacity to fold into a compact globule, the Structural Features of which also govern the oligomerization of many proteins. The chemical and functional properties of proteins rely on specific interactions among functional groups brought into spatial proximity within its three-dimensional structure. Consequently, the behavior of these groups differs fundamentally from their reactivity when present in free amino acids or small peptides. A protein can function—acting as an enzyme, structural or transport protein, regulator, toxin, or inhibitor—solely by virtue of possessing a precisely defined three-dimensional molecular architecture stabilized by hydrophobic radical interactions.
Isoproteins
As noted above, homologous proteins are proteins that perform the same function in different organisms. If several forms of a single protein are present within the same organism and/or among Representatives of the same species that differ in metabolic activity, they are referred to as isofunctional proteins or isoproteins. For example, the blood of a healthy adult human contains 96% HbA, along with 2% each of HbF and HbA2. All of these are tetramers whose protomers are represented by four forms: α, β, γ, and δ. The corresponding structures of the aforementioned Hemoglobins are α2β2, α2γ2, and α2δ2, respectively. Although they exhibit very minor differences in their primary, secondary, and tertiary structures and perform the same biological function, they differ in their oxygen affinity. Specifically, HbF has a higher oxygen affinity than HbA and is capable of abstracting oxygen from it:
HbAО2 + HbF ↔ HbA+ HbFО2.
HbF is characteristic of the Embryonic Stage of Human Development. The existence of these two isoproteins enables The transfer of oxygen from the maternal blood to the embryo across anatomically separate circulatory systems. It is important to distinguish between isoproteins and homologous proteins: the latter belong to organisms of different species.
The multiplicity of protein forms arises from a variety of factors that can be divided into two main categories. The first category comprises genetic-level factors, wherein an organism possesses multiple genes, each encoding a specific subunit of an oligomeric protein. The second category involves post-translational modifications; in this case, differences among subunits encoded by the same gene stem from diverse modifications of initially homogeneous polypeptide chains. Gene multiplicity, in turn, manifests in two forms: 1) multiple alleles at a single locus, and 2) multiple gene loci.
In a diploid genome, every gene locus is duplicated. An individual may be heterozygous for any given locus, meaning it possesses different alleles. If such a locus encodes a specific enzyme subunit, homozygous individuals will synthesize subunits of only one type, whereas heterozygous individuals should produce subunits of two different types. Obviously, allelic multiplicity within a single individual cannot generate an immense diversity of protein forms, as the number of alleles at a diploid locus cannot exceed two. However, within the gene pool as a whole, the total number of alleles can be quite large, leading to significant Variability in enzyme subunit types when comparing different individuals. Differences between subunits encoded by distinct alleles of the same locus are typically minor, often involving single Amino Acid Substitutions driven by point mutations in the DNA nucleotide sequence. When a gene locus is active, both alleles are generally expressed. Consequently, despite fluctuations in total enzyme activity depending on cell type, an individual's isozyme profile remains constant: homozygous individuals always display a single subunit type, while heterozygotes display two. Many proteins are encoded not by a single gene, but by several, in which case each locus corresponds to a distinct enzyme subunit.
Because the expression of each gene locus is regulated independently, certain cells within an organism may synthesize one type of subunit while others produce a different type. Furthermore, the expression of gene loci frequently changes during organismal development, leading to corresponding shifts in the types of subunits synthesized within a given tissue. Thus, gene multiplicity allows the isozyme profile to vary from one tissue to another, as well as within the same tissue throughout growth and development. Such shifts in the isozyme profile would be impossible if isoform diversity were solely the result of different alleles at a single locus. Since all individuals of a given species share the same gene loci, a multiplicity of protein-coding loci in the absence of allelic variation could not account for differences in the isozyme spectrum among individual members of a species.
Protein subunits corresponding to different loci typically diverge significantly from one another—exhibiting numerous amino acid substitutions, deletions, or insertions—which may also be reflected in their molecular weights. To understand how isoproteins arise from this dual level of gene multiplicity, let us return to the example of human hemoglobin. It exists in several well-characterized molecular forms generated through a combination of both gene locus multiplicity and allelic variation.
Hemoglobin is encoded by eight distinct gene loci, each specifying a particular subunit. The polypeptide sequences of these subunits are highly similar despite numerous amino acid replacements. For instance, the α- and β-subunits differ at approximately 50% of their amino acid residues, and their polypeptide lengths also vary slightly (141 residues in the α-subunit versus 146 in the β-, γ-, and δ-subunits). As the organism develops from early embryonic stages through childhood, the composition of hemoglobin undergoes qualitative changes driven by shifts in the transcriptional activity of specific gene loci. For example, newborn hemoglobin has an α2γ2 composition, but after several months it is replaced by the α2β2 hemoglobin as the β locus becomes active and the γ locus is downregulated.
Unusual hemoglobins also exist as a result of rare alleles, and their catalog continues to grow. The most prevalent and thoroughly studied is the sickle cell anemia gene, which represents an allele of the hemoglobin β locus. The polypeptide it encodes differs from the normal β-chain by a single amino acid Substitution at the 6th position, where valine replaces glutamate. Even in geographic regions where the sickle cell allele is common, individual hemoglobin profiles vary depending on whether a person is homozygous for the normal allele, homozygous for the sickle cell gene, or heterozygous. Newly synthesized proteins frequently undergo post-translational modifications, such as glycosylation, limited proteolysis, or Covalent Modification of amino acid side chains. If such modifications affect only a fraction of the subunit population rather than the entirety, both modified and unmodified subunits will coexist within the organism, thereby generating isozymes. A classic example is the microheterogeneity of muscle aldolase.
In vertebrates, aldolase is encoded by three gene loci. Although only locus A is active in Skeletal Muscle, two types of subunits—designated Aα and Aβ—can be isolated from this tissue. It has been shown that the primary translation product is the Aα polypeptide, which slowly converts into Aβ via the deamination of an asparagine residue near the carboxy terminus of the chain. When post-translational modification processes are much more active in some tissues than in others, it results in a tissue-specific distribution of secondary isozymes that superficially mimics the effects of gene multiplicity. Consider the subunits of pyruvate kinase. In mammals, this enzyme comprises three Major Types of subunits, each governed by an independent gene locus. Type L subunits (also referred to as type I) are found in The Liver and erythrocytes; however, the type L isozymes from these two sources differ slightly in molecular size and electrophoretic mobility. It was initially hypothesized that the slightly larger type L' subunit found in erythrocytes originated from a fourth gene locus. Subsequent studies, however, revealed that this difference stems from post-translational modification. In all tissues, the primary translation product of the L locus is the type L' subunit, but in the liver, these subunits are rapidly cleaved by Proteolytic Enzymes to yield the shorter type L subunits, whereas this process proceeds much more slowly in erythrocytes. Pyruvate kinase plays a pivotal role in the Regulation of Glycolysis, and the L and L' subunits differ in their regulatory properties, which presumably carries specific physiological significance. It has been proposed that the term "isozyme" should be restricted strictly to Multiple molecular forms of genetic origin, and not applied when diversity arises from post-translational modifications.
Many enzymes exist as oligomers composed of multiple subunits, with dimers and trimers being particularly common. If a cell contains two or more types of subunits capable of associating with one another in arbitrary combinations, a diverse array of isozymes is generated. In the examples discussed earlier, isoform formation resulted from two gene loci; however, when cells contain three or more distinct subunit types, The complexity of the resulting isoprotein profile increases dramatically. For instance, most vertebrate species possess three lactate dehydrogenase subunit types—A, B, and C. Random tetrameric assembly of these subunits can yield up to 15 isozymes, 12 of which are hybrids. Complexity increases even further when allelic variants are added to multiple loci. In the case of vertebrate lactate dehydrogenase, allelic variants encoding the A and B subunits have been identified. An organism heterozygous for these variants will synthesize five distinct lactate dehydrogenase subunits: A, A', B, B', and C, theoretically permitting the formation of up to 70 tetrameric isozymes.
Fortunately for researchers, the complexity of isozyme profiles is often limited in practice, as not all theoretically possible hybrid variants are actually formed. Most commonly, this occurs because not all gene loci are simultaneously active within a cell, and the expression of certain genes is strictly restricted. For instance, the C-subunit of lactate dehydrogenase is synthesized exclusively in primary spermatocytes, where the A and B subunits are absent. Due to this compartmentalization *in vivo*, hybrid isozymes containing the C-subunit do not occur.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.