Fundamentals of Biochemistry - A. A. Anisimov 1986

Proteins
Structure of the Protein Molecule

2.4.1. Polypeptide Structure of Proteins. The first protein substances were isolated more than 250 years ago, and by the second half of the 18th to the early 19th centuries, protein substances of PLANT AND ANIMAL origin had already been described on numerous occasions.

Currently, the elemental Chemical composition of Proteins is well understood. They typically contain 50—55% C, 21—23% O2, 15—17% Na, about 7% H2, and from 0 to 3% S. In addition, Conjugated Proteins may contain P and certain metals.

The structure of protein molecules is a considerably more complex issue. Profound and, for their time, highly fascinating studies on Cell/13.html">Protein Structure were conducted in the 1880s by the prominent Russian biochemist and one of the founders of national biochemistry, A. Ya. Danilevsky (1838—1923). Investigating the products of Protein Cleavage under METABOLISM/18.html">The Influence of weak alkalis, as well as the Specific features of the biuret reaction typical of proteins and some of their breakdown products, he put forward A number of intriguing ideas regarding the STRUCTURE OF THE protein molecule. According to A. Ya. Danilevsky, a protein molecule consists of structurally quite similar chains in which C and N atoms alternate ("carbon-nitrogen chains," "elementary series"), with Amino Acids attached to them.

A. Ya. Danilevsky was the first in science to point out the polymeric nature of protein structure. Of great interest was his assertion that protein cleavage is of a Hydration nature, i.e., proceeds via Hydrolysis. While these ideas of A. Ya. Danilevsky were undoubtedly still very far from the modern theory of protein structure, they nevertheless represented its inception.

The decisive step toward the polypeptide theory of protein molecular structure was taken by the foremost German organic chemist and biochemist Emil Fischer (1852—1919). By hydrolyzing proteins with Hydrochloric acid and Enzymes and subsequently analyzing the cleavage products, he concluded that a protein molecule is formed by A large number of amino acid residues. To determine how they are linked together, Fischer pursued the artificial synthesis of protein-like molecules from amino acids. Since free amino acids do not interact when their solutions are mixed, he used their halogen acyl derivatives as starting Materials for the synthesis. The synthesis conditions were such that the amino group of one amino acid reacted with the carboxyl group of another, forming a peptide bond —СО—NH— (the presence of this type of bond in the protein molecule had previously been suggested by A. Ya. Danilevsky, who called them biuret bonds). In this way, he succeeded in joining up to 19 amino acid residues into a single molecule. A comparison of The properties of these artificially synthesized compounds with Peptides—products of the incomplete breakdown of natural proteins—revealed a striking similarity. This proved that in natural proteins as well, Amino acids are linked to one another by peptide bonds.

According to modern data, 20 types of amino acids are most frequently found in various proteins. It is precisely for these 20 amino acids that a Genetic Code exists in the form of triplets (triplets of NUCLEOTIDES in DNA). Sometimes Other Amino Acids are also present in proteins; these are formed As a result of Post-translational protein modification and are non-coding (cystine, hydroxyproline, hydroxylysine, and several others). Only $\alpha$-Amino acids have been found in proteins, the vast majority being in the L-configuration.

Amino acids are linked together by a covalent peptide or amide bond. Its formation occurs through the amino group (—NH2) of one Amino Acid and the carboxyl group (—COOH) of another, with the release of a Water molecule.

Class="center">

The trans-peptide bond is the most widespread in nature, whereas the less stable cis-peptide bond is encountered less frequently. The peptide bond possesses a partial double- and partial single-bond character, with mutual transitions between these structures. The lifetime of the single bond is somewhat longer than that of the double bond (6:4); thus, it can be said that the peptide bond is 60% single and 40% double.

As a result of Resonance, a fluctuating, dynamic bond is formed that cannot be described on The basis of a single valence structure. Because rotation around the double bond is hindered, all atoms of the peptide bond turn out to be located approximately in the same plane, i.e., it is planar; only around the nitrogen atom do the bonds partially retain a pyramidal character. To date, all valence angles and Bond Lengths in peptide groups have been established (Fig. 2.7).

Polymers formed by amino acids are referred to as peptides or proteins depending on the number of structural units they comprise. It is conventionally accepted that peptides containing up to 20 amino acid residues are classified as oligopeptides, which include di-, tri-, tetrapeptides, etc. Polypeptides contain from 20 to 50 amino acid residues in their molecule. Peptide chains comprising more than 50 Amino Acids and having a molecular weight exceeding 6,000 are classified as proteins.

The lowest molecular weight protein is the hormone Insulin, which consists of 51 amino acid residues. The number of amino acid units in a protein can reach several hundred or even thousands. The number of protein species in nature is enormous; their diversity is related to the diverse set of amino acids making up the protein and the order of their sequence within the molecule. For instance, just Three amino acids can yield 6 different tripeptides, four can yield 24 tetrapeptides, five can yield 120 pentapeptides, 11 can yield 40 million isomers, and from 20 different amino acids, each occurring only once, an astronomical number of isomers ($2\cdot10^{18}$) can theoretically be formed. However, only a tiny fraction of these possible isomers is realized in living nature.

Fig. 2.7. Interatomic distances (nm) and angles in the peptide bond. All atoms within the frame lie approximately in the same plane

To describe the structure of protein molecules, The concepts of primary, secondary, tertiary, and quaternary structures were introduced. In recent years, two more levels have been added: supersecondary structure and domains, which occupy an intermediate position between secondary and tertiary structures.

2.4.2. Primary Structure. The Introduction/19.html">Primary structure of a protein molecule is understood as The sequence of amino acids in the polypeptide chain (or chains) and the Location of Disulfide Bonds. A polypeptide chain contains a free amino group at one end (N-terminus) and a carboxyl group at the other (C-terminus). The N-terminus is taken as the beginning of the chain, and it is precisely from here that amino acid numbering starts. This corresponds to the direction of Polypeptide chain synthesis on the ribosome, which in turn matches the 5'→3' direction of mRNA (see Section 4.4.3).

The amino group at the N-terminus of a polypeptide chain may sometimes be acetylated, having attached an acetic acid residue (CH3—СО—NH—...), as, for example, in cytochrome c1, Ovalbumin, Lactate dehydrogenase, Actin, and Myosin. Acetylation-blocked N-termini are also characteristic of the coat proteins of many plant Viruses, as well as certain animal and bacterial viruses. Formylation of the $\alpha$-amino group has been found in the bee venom melittin and lamprey Hemoglobin, while methylation has been detected in E. coli ribosomal proteins. In a number of proteins (Hormones, light and heavy chains of IMMUNOGLOBULINS), the N-terminal residue is pyroglutamic acid (pyrrolidonecarboxylic acid), which contains no free amino group.

At the C-terminus, one finds either a free carboxyl group (in most proteins) or an amidated one (certain hormones, bee venom). C-terminal modifications are rarer compared to N-terminal modifications.

The names of individual peptides are formed in accordance with their constituent amino acid residues, starting from the N-terminus. In doing so, the endings of all amino acid names, except for the last one, are changed to "-yl". For example, L-alanyl-L-cysteinyl-L-Methionine. The complete Amino Acid Sequence of proteins is indicated using abbreviated amino acid names. Three-letter and single-letter designations for amino acids are universally adopted (Table 2.3).

Table 2.3. Designations of amino acid residues comprising Peptides and Proteins

Full amino acid name

Abbreviated name


three-letter

single-letter

Glycine

Gly, Gly

G

Alanine

Ala, Ala

A

Valine

Val, Val

V

Leucine

Leu, Leu

L

Isoleucine

Ile, Ile

I

Proline

Pro, Pro

P

Phenylalanine

Phe, Phe

F

Tyrosine

Tyr, Tyr

Y

Tryptophan

Trp, Trp

W

Serine

Ser, Ser

S

Threonine

Thr, Thr

T

Aspartic acid

Asp, Asp

D

Glutamic acid

Glu, Glu

E

Asparagine1

Asn, Asn

N

Glutamine1

Gln, Gln

Q

Cysteine

Cys, Cys

C

Methionine

Met, Met

M

Histidine

His, His

H

Lysine

Lys, Lys

K

Arginine

Arg, Arg

R

1 If it is unknown whether The amino acid or its amide is present in the protein chain, the following designations are used: Asx (Asx) or B, and Glx (Glx) or Z.

Fig. 2.8. Torsion angles of the α-carbon atom in a peptide chain

The primary bond of protein structure is the peptide bond. This bond is quite rigid, which limits its conformational mobility. However, each amino acid residue contains an α-carbon atom, which introduces two single bonds into the residue; rotation is possible around these single bonds. The rotation angles of single bonds are referred to as torsion angles and are denoted as φ (N — Ca) and φ (C — Ca) (Fig. 2.8). The number of possible combinations of torsion angles is large, and many of them are realized in proteins. The exceptions are pairs of angles that are sterically impossible due to the close spatial arrangement of atoms in neighboring peptide groups and the side chains of amino acid residues.

Currently, using computers and molecular models, the entire range of possible values for φ and ψ is examined, and the results of such analysis are presented as plots of φ versus ψ, known as conformational maps or Ramachandran plots.

To date, the amino acid sequence has been determined for approximately 2,000 proteins: Cytochromes, ferredoxins, immunoglobulins, ribosomal proteins, Hemoglobins, and a large number of enzymes. Bovine insulin was the first to be sequenced (F. Sanger, 1951–1953), an achievement for which he was awarded the Nobel Prize. One of the large proteins for which the primary structure has been established is the enzyme aspartate aminotransferase (M 93,000). A single polypeptide chain of this dimeric enzyme consists of 412 amino acid residues. Its structure was elucidated through the collaborative work of two laboratories led by Yu. A. Ovchinnikov and A. E. Braunstein in 1971.

The workflow for determining the Primary Structure of Proteins is as follows. 1. Cleavage of the protein polypeptide chain into smaller fragments at specific sequence positions. 2. Determination of the order of amino acids within the resulting peptide fragments. 3. Determination of the arrangement of peptides with known Amino acid sequences within the protein molecule.

The cleavage of the polypeptide chain into fragments must be preceded by the reduction of disulfide bonds and the modification of cysteine residues. To cleave proteins into fragments, selective hydrolysis is employed using Proteolytic Enzymes (such as Trypsin, Chymotrypsin, etc.) that cleave peptide bonds formed by specific amino acids, or chemical agents that also possess selective activity. Trypsin cleaves peptide bonds whose carbonyl group belongs to lys or arg, whereas chymotrypsin hydrolyzes peptide bonds involving the carboxyl group of aromatic amino acids: tyr, phe, and trp.

Proteolytic enzymes used for selective hydrolysis are usually immobilized beforehand—that is, bound to an insoluble matrix—to facilitate the subsequent Separation of the enzymes from the hydrolysis products.

Among the chemical Reagents used for selective hydrolysis, the most successful are Cyanogen bromide, which cleaves peptide bonds whose carbonyl group belongs to met residues, and N-Bromosuccinimide, which cleaves bonds following trp.

Selective hydrolysis yields a mixture of peptides, each of which must then be isolated in pure form. Peptide fractionation is primarily carried out using high-voltage (up to 10,000 V) Electrophoresis and Chromatography (mainly ion-exchange, partition, or molecular sieve chromatography) or a combination of these Methods, such as the peptide mapping or "fingerprint" method. The latter method involves spotting the peptide mixture onto a sheet of chromatographic paper and performing partition chromatography in one direction, followed by electrophoresis in the perpendicular direction.

To determine the sequence of amino acids in peptides, a number of methods have been developed that allow for the investigation of the amino acid sequence from both the N- and C-termini of the molecule. One of the most common reagents used for N-terminal peptide analysis is phenylisothiocyanate (PITC), first introduced by P. Edman. The reaction of PITC with the NH2 groups of peptides yields a product: a phenylthiocarbamyl derivative of the amino acid, or a PTC derivative. In an acidic environment, this product cyclizes, leading to the cleavage of the peptide bond and the release of a phenylthiohydantoin derivative of the N-terminal amino acid (PTH derivative) from the rest of the peptide. The resulting PTH derivatives are identified using various chromatographic techniques, thereby establishing the first amino acid in the peptide. By repeating the PITC Treatment on the shortened peptide, the next amino acid residue is determined, and so on. The reactions underlying the Edman method are as follows:

THE PRINCIPLE OF the Edman method is utilized in automated sequenators, which allow for the automatic determination of the amino acid sequence in a peptide containing several dozen residues. A specific reagent for free NH2 groups in proteins is 1-fluoro-2,4-dinitrobenzene (DNFB), first used by F. Sanger and applied by him to determine N-terminal amino acids in peptides. The reaction proceeds as follows:

Acid hydrolysis of a dinitrophenylated peptide leads to the cleavage of all peptide bonds within the molecule, releasing the dinitrophenylated N-terminal amino acid, which is then identified chromatographically using standard markers: DNP-amino acids.

To determine the amino acid sequence of peptides from the N-terminus, aminopeptidases can be employed; these enzymes sequentially cleave amino acids possessing a free amino group. Utilizing chromatographic techniques, the released amino acids are identified and their accumulation rates in the hydrolyzate are measured, thereby providing insight into the order of amino acids at the N-terminus of the peptide.

A similar approach is used to decipher the C-terminal amino acid sequence using Carboxypeptidases, which cleave amino acids with a free carboxyl group. Information regarding the C-terminal amino acid can also be obtained by cleaving peptide bonds in proteins via Hydrazinolysis. This method, originally proposed by S. Akabori, is based on the reaction of a polypeptide with anhydrous hydrazine at 100°C.

All amino acid residues are converted into amino acid hydrazides, with the exception of the one carrying a free carboxyl group. The latter is obtained as a free amino acid, which is then isolated and identified chromatographically.

Mass spectrometry is applied to analyze amino acid sequences from both the N- and C-termini in previously modified oligopeptides.

To establish the complete amino acid sequence of a protein, it is necessary to determine the arrangement order of the peptides within the molecule. This is achieved through a logical approach known as the "overlapping peptides" method. It can be applied provided that the primary structures of the peptides obtained after Protein Hydrolysis are known, utilizing at least two Different types of selective cleavage. The ordering of the peptides is then deduced from these "overlapping" fragments.

For instance, peptide a from a tryptic hydrolyzate contains the complete sequence of peptide g and a portion of peptide d from a chymotryptic

hydrolyzate. Consequently, peptide d directly follows peptide g in the protein molecule. The relative positions of the remaining peptides (G, W, etc., denoting abbreviated amino acid names) are established in a similar manner.

The primary structure of a protein can also be determined indirectly via the structure of its encoding Gene by analyzing The nucleotide sequence. Using this approach in combination with the aforementioned Methods for determining the order of amino acids in proteins, a group of Soviet scientists led by Yu. A. Ovchinnikov successfully deciphered the primary structure of RNA polymerase—a large protein consisting of approximately 4,000 amino acids.

Elucidating the primary structure of proteins also involves determining their Amino Acid Composition and locating disulfide bonds. The Qualitative and quantitative Amino acid composition of proteins is established after hydrolyzing them into monomers and using specialized equipment known as automated amino acid analyzers. In these instruments, amino acids are separated on ion-exchange resin columns, sequentially eluted, and quantitatively analyzed using ninhydrin or a more sensitive reagent, fluorescamine, which forms fluorescent products with amino acids.

The localization of disulfide bonds in a protein is determined while elucidating its complete amino acid sequence. To achieve this, all free SH groups in the protein are modified prior to selective hydrolysis. Peptides containing S—S bonds are isolated and treated with reagents that cleave the disulfide bridges. The resulting peptides are then separated, and the amino acid sequence of each is determined using the methods outlined above.

Concluding the Overview of the main steps and methods used to study the primary structure of proteins, it must be emphasized that the initial Procedure is determining the number of polypeptide chains—or protomers—in the protein molecule. This is typically deduced from the number of NH2- terminal residues per protein molecule. If a protein contains more than one polypeptide chain, they are dissociated from one another using, for example, detergents, and then separated and purified via electrophoresis or chromatography.

The primary structure elucidation data currently available for a large number of proteins already allow for certain generalizations. Despite the vast diversity of properties among individual proteins and differences in their primary structures, There is a distinct similarity in quantitative amino acid composition. The majority of proteins are characterized by the presence of all 20 types of amino acids, with glycine, alanine, aspartic acid, and glutamic acid usually being particularly abundant, while tryptophan, histidine, methionine, and arginine occur in smaller amounts. Amino acids with hydrocarbon side chains typically account for 30—40% of the total amino acid content. Notable exceptions to this generalized compositional profile are found only in certain protein groups, such as Histones, protamines, and proteinoids. These proteins may lack certain types of amino acids or be enriched in specific ones (e.g., arginine in histones and protamines, and hydroxyproline in proteinoids).

Analysis of a large number of proteins with known primary structures has shown that Globular proteins lack any universal patterns in amino acid ordering (for instance, one cannot state that amino acid X is invariably followed by amino acid Z across all proteins; each protein exhibits its own unique sequence pattern). At the same time, nature evidently did not utilize the entire colossal Diversity of proteins theoretically possible from the twenty amino acid building blocks. As F. Šorm (Czechoslovakia) points out, even proteins as biologically disparate as the hormone insulin and the enzyme Ribonuclease share identical small molecular fragments: three identical tripeptides, one identical tetrapeptide, and five analogous tripeptides differing by only a single amino acid, where the substituted amino acids are structurally similar (e.g., ala instead of val, or asp instead of glu).

The arrangement of amino acids in proteins, as well as the resulting Spatial Structure built upon it, is genetically determined and fine-tuned to perform specific biological Functions. The amino acid sequences of proteins responsible for the same biological function are often remarkably similar across different species. Such proteins, which are identical or structurally similar and perform the same functions in different organisms, are termed homologous. For instance, all vertebrate insulins consist of two chains (A and B) comprising 51 amino acid residues. In the A-chain of insulins from various animals, individual sequence differences occur primarily at positions 4, 8, 9, 10, 13, 14, 15, and 18. These are designated as variable amino acids, whereas other residues generally remain invariant. In Cytochromes c, 35 out of 104 residues are identical across all species, with the number of Amino Acid Substitutions showing a direct correlation with phylogenetic relatedness. While cytochromes c from vertebrate animals and Yeast differ by 43—48 amino acids, those from humans and monkeys exhibit only a single substitution, and the molecules of this protein are identical in chickens and turkeys.

The Study of the primary structure of homologous proteins isolated from various living species serves as one of the key methods in modern Taxonomy.

Differences in the structure of homologous proteins also provide valuable insight into The Role of individual amino acid residues in molecular function. Residues located in active sites or determining the conformation of the polypeptide chain cannot be altered genetically or through chemical modification without impairing function. For instance, currently known primary structure variations in cytochrome c across various species do not entail significant Changes in the protein's functional properties, as the least mutable regions are those near the heme-binding site as well as those responsible for the spatial folding of the chain.

Most amino acid substitutions, deletions, or insertions typically occur On the surface of the protein globule, since exterior residues are less critical for protein stability. However, these alterations are generally observed in regions of lesser importance for the expression of functional properties. The degree of Variability of an amino acid residue in homologous proteins can indicate its spatial location—whether on the surface or in the interior of the protein.

Research on a number of enzymes has demonstrated that Protein Functions are robust against Certain amino acid substitutions, deletions, and insertions. Proteins contain fragments of the polypeptide chain that can be removed without disrupting function. In most cases, minor substitutions involve replacing one amino acid residue with another of similar properties, such as the interchange of aliphatic hydrophobic residues (Ile for Val or Leu, Met, etc.) or polar residues (Arg for Lys, Glu for Asp, Gln for Asn). Substitution with a residue possessing a different type of side chain is highly critical if it occurs in a region of the molecule responsible for function and conformation.

The tolerance for substitutions in the primary structure of homologous proteins is not a universal rule for all proteins. There are also highly conserved proteins; an example is histone H4, which differs by only two out of 102 amino acid residues between plants and animals. Evidently, every residue in histone H4 is critical for its function, and the evolution of this "ancient" protein was completed at an early stage of organismal existence.

Alongside the divergence in the primary structure of proteins performing identical functions in different organisms, biochemistry also recognizes The phenomenon of convergence, wherein proteins with analogous functions are structurally non-homologous. For example, peptide bond hydrolysis is carried out by various types of enzymes that differ in active center structure and catalytic mechanism. Pancreatic proteinases (trypsin, chymotrypsin, etc.) and subtilisin (a proteinase from Bac. subtilis) possess distinct primary structures and Conformations yet share a similar reaction mechanism.

The Determination of Amino acid sequences in proteins has also revealed that Gene Duplication and fusion occurred during the course of evolution. Protein Differentiation typically originates with the duplication of corresponding genes, and the products of these genes may gradually acquire distinct functions. Striking Examples of protein differentiation include a-lactalbumin and Lysozyme, which share similar sequences and spatial structures.

Thus, studying the primary structure of proteins provides us with valuable insight into Evolutionary Processes, helps establish Phylogenetic relationships among individual species of living organisms, and proves the continuity of the evolutionary course. The latter Conclusion stems from the fact that the amino acid sequences of a specific protein are not always entirely identical across every individual of a given species. Such polymorphism results from continuous Mutations occurring within The Genome of that species. Mutations beneficial to the species become fixed in the progeny, whereas those disrupting protein function lead to Hereditary diseases.

2.4.3. The Role of Weak interactions in The formation of Biopolymer Spatial Structure. The Spatial Organization of macromolecules and cellular structures is achieved primarily through chemical bonds that are significantly weaker than covalent ones. Covalently bonded atoms are capable of additional weak interactions with other atoms both within the same molecule and with atoms of neighboring molecules. Weak interactions participate in shaping the conformation of complex Biopolymers (proteins, NUCLEIC ACIDS) and their spatial structure, as well as determining the degree of Stability of the latter.

Weak interactions, sometimes referred to as secondary bonds, include hydrogen and ionic bonds, Van der Waals forces, and hydrophobic interactions. The strength of a chemical bond can be characterized by The change in standard Gibbs Free energy, ∆G (see Section 1.3.2), that occurs upon its formation. For weak interactions, this value ranges from 4—30 kJ/mol, whereas covalent bonds are extremely strong, with a Free energy of formation of 200—450 kJ/mol. Bond strength correlates with the interatomic distance: the stronger the bond, the shorter this distance. Covalent bonds are the shortest at 0.10—0.18 nm, while in weak interactions the interatomic distance is 0.20—0.45 nm.

Hydrogen Bonds play a crucial role in forming the structure of biological macromolecules. They arise between two electronegative atoms when a hydrogen proton, covalently bonded to one of these atoms, is positioned between them. Electronegative atoms (i.e., those with a high electron-attracting capacity) include O, N, and F, and less frequently Cl and S participate in hydrogen bonding. A hydrogen atom contains a single electron, and when the latter departs to form a covalent bond, The Nucleus is left without electron shells. Such hydrogen—that is, a proton—is naturally not repelled by the electron clouds of neighboring atoms; instead, it is attracted by them, forming a Hydrogen bond.

A prerequisite for hydrogen bond formation is the presence of at least one unshared electron pair on the electronegative atom, toward which the hydrogen atom will be attracted. Electronegative atoms possess a high electron affinity, thereby filling their entire outer shell with electrons (8 electrons) as if becoming overloaded with negative charges. If a pair of unshared electrons emerges in the process, it interacts with the proton. An example of a hydrogen bond is shown below, where δ+ represents an excess of positive or negative charge, and the dashed line indicates hydrogen bonds:

Hydrogen bonds are characterized by a low formation energy (about 20 kJ/mol) and are longer than covalent bonds (0.26—0.31 nm). The hydrogen atom forming the hydrogen bond is not located equidistant from the electronegative atoms, but rather closer to the atom with which it shares a covalent bond.

Hydrogen bonds are directional in nature. In the strongest hydrogen bonds, the hydrogen atom lies along the straight line connecting the donor and acceptor atoms. If a hydrogen bond forms at an angle to the covalent bond, its energy is correspondingly lower. Protein molecules feature Two Types of hydrogen bonds: those between peptide bond groups and those between amino acid side chains.

Hydrogen bonds can be intramolecular or intermolecular. A whole range of liquids (water, organic acids, alcohols, etc.) are associated due to the formation of intermolecular hydrogen bonds. Carbon is incapable of forming hydrogen bonds because its electronegativity is significantly lower than that of O or N, closely matching that of H. Consequently, hydrocarbon chains are hydrophobic and penetrate water with difficulty, as they are unable to disrupt its hydrogen bonds.

Compounds containing —OH, —NH2, —COOH, and groups exhibit a pronounced ability to form Hydrogen bonds and are hydrophilic. Water molecules in the ice state are linked to one another via hydrogen bonds, with every six molecules forming a ring-like six-membered structure. Liquid water also appears to contain ice-like clusters (aggregates and associations) that continuously break down and reform. Water molecules form hydrogen bonds not only with each other but also with the polar groups of dissolved compounds. The polarity of the water molecule and its ability to form hydrogen bonds are paramount in its role as a biological solvent—the primary medium of living Cells.

Individual hydrogen bonds formed in aqueous solution are very weak, which is explained by the competition between water molecules and solute molecules for hydrogen bonding partners. However, when a macromolecule contains a large number of hydrogen bonds, a very high cumulative strength arises. This phenomenon is termed cooperative hydrogen bonding.

Hydrophobic interactions are no less important in stabilizing biopolymer structures and their functioning. Water molecules, striving to form hydrogen bonds with one another, push out hydrophobic groups and molecules present in water, forcing them to cluster together and form aggregates. This process occurs spontaneously, and the free energy of the system decreases because the water layer near The surface of hydrophobic groups possesses a higher degree of Structuring and lower Entropy (a measure of disorder) compared to the bulk water; upon the association of hydrophobic groups, their total surface area diminishes. Notably, no special bonds are formed between the hydrophobic groups or molecules themselves (aside from potential van der Waals attractive forces), which is why researchers generally speak of hydrophobic interactions rather than bonds.

The weakest bonds between molecules are caused by dispersion forces of van der Waals attraction (sometimes called van der Waals bonds). They arise only at sufficiently short distances between molecules and are based on Coulombic electrostatic forces of attraction. Nuclei within the electron shells of atoms are in continuous vibrational motion, making a temporary displacement of electron orbits relative to the nucleus possible, which leads to dipole formation. Although such dipoles exist briefly, the duration is sufficient to establish a coordinated orientation between molecules.

It should be kept in mind that opposing the attractive dispersion forces is the mutual repulsion of the electron shells of non-bonded atoms. On the other hand, since covalent bonds between different types of atoms lead to an asymmetric distribution of valence electrons, most atoms in a molecule carry partial charges. Since the net charge of a neutral molecule is zero, it can be approximately viewed as a set of dipoles or multipoles that give rise to Electrostatic Interactions.

For computational convenience, the three aforementioned non-covalent forces (van der Waals attraction, electron shell repulsion, and electrostatic interactions) are usually combined into a single force field: the van der Waals potential. Such potentials make it possible to determine the possible Contact distances between pairs of atoms, known as van der Waals radii. It is generally accepted that the forces of van der Waals interactions are inversely proportional to the sixth power of the distance between the interacting groups. The energy of these interactions ranges from 4 to 8 kJ/mol, and the potential bond length is 0.33–0.45 nm (the largest among all types of weak interactions). Van der Waals forces are of great biological importance, as they play a role in the self-assembly of biological macromolecules and structures, the interaction between proteins and other biopolymers, and the stabilization of the Tertiary and Quaternary structures of biopolymers.

The interaction of atoms with drastically different properties (such as metals and metalloids) results in the formation of ionic bonds. In this process, one atom (the cation) transfers an electron to another atom, so that the resulting electron pair belongs exclusively to the latter (the anion). Alkali metal atoms give up electrons most readily, whereas halogen atoms exhibit the maximum electron affinity. Ionic bonding is fundamentally based on electrostatic interaction. The average energy of an ionic bond In aqueous solutions is about 21 kJ/mol, and its length ranges from 0.20 to 0.33 nm.

Ionic forces are highly significant in molecular interactions. For instance, the attraction between the —COO- and —NH3+ groups plays a major role in Protein-Protein Interactions. A doubly charged Ca2+ cation can act as a "bridge" linking two carboxyl groups.

A crucial aspect of all ionic interactions in aqueous solutions is ion hydration. Every ion in water is surrounded by water dipoles that are strictly oriented relative to it. Ion hydration exerts a major influence on their interaction in solution and, alongside other factors, determines the strength of acids and bases, the bond strength between metal cations and negatively charged groups, and more.

It is known that in many cases the charge of an ionized organic molecule is neutralized either by inorganic cations (Na+, K+, Mg2+, etc.) or by inorganic anions (Cl-, SO2-4, etc.). In aqueous solutions, such neutralizing ions cannot occupy fixed positions due to their hydration. Therefore, in aqueous solutions, ionic interactions with hydrated inorganic ions generally do not play a significant role in determining the conformation of organic molecules.

A vital feature and, at the same time, an advantage of the aforementioned weak interactions is that their energy (4–30 kJ/mol) does not significantly exceed the kinetic energy of thermal motion (2.5 kJ/mol). This slight excess is quite sufficient for relatively stable secondary bonds to form between molecules at physiological temperatures. However, because the difference between the energy of weak interactions and the kinetic energy of thermal motion is small, their formation and disruption occur without the participation of enzymes.

Due to the wide distribution of the kinetic energy of molecules, there are always molecules at physiological temperatures whose kinetic energy is sufficient to break weak bonds. This imparts a certain lability to intermolecular interactions. Otherwise, The Cell would possess a rigid, crystal-like structure, and the diffusion rate would be extremely low, which is incompatible with the existence of a living cell. In particular, the remarkably high rate of Enzymatic Catalysis is partly due to the fact that enzyme-substrate complexes are formed via weak interactions; consequently, these complexes arise and dissociate rapidly even under the influence of random thermal motion.

2.4.4. Secondary structure. The secondary structure is the ordered spatial arrangement of individual segments of a polypeptide chain, independent of the type and conformation of the amino acid side chains. It is formed by the closure of hydrogen bonds between peptide groups. The secondary structure is mainly represented by regular structures such as the a-helix, pleated sheets (ß-Structure), and the ß-turn. Portions of the polypeptide chain that lack an ordered structure are referred to as amorphous or unstructured regions.

In a-helical regions and regions with a ß-pleated structure, all consecutively arranged peptide units of the polypeptide chain have identical relative orientations, since all torsion angles φ and all angles ψ at Ca are identical. In this case, the polypeptide chain segment has a linear structure formed by Linear groups.

A linear group represents a turn of a helix whose parameters (axial Translation per repeating element, number of elements per turn, radius, etc.) depend on the magnitude of the angles φ and ψ. A helix with fewer than two elements per turn is impossible. Several types of linear groups without steric hindrance have been discovered in proteins; they are stabilized by hydrogen bonds either within a single segment of the polypeptide chain (a helix) or between adjacent segments (a ß-pleated structure). When the torsion angles are close to —60, —45°, the secondary structure is represented by right-handed a-helices. This type of helix, described by L. Pauling, possesses the lowest Free Energy and is the most "favorable" given the constraints imposed by the geometry of the peptide bond and the allowable variations of the φ and ψ angles.

In an a-helix, all hydrogen bonds are approximately parallel to the helix axis and collinear with one another, which corresponds to the minimum free energy; each carbonyl group forms a hydrogen bond with the NH group located four residues away along the chain.

During the formation of a-helices, the maximum possible number of hydrogen bonds is closed, imparting stability to this structure. The a-helix is characterized by the following parameters: number of amino acid residues per helical turn — 3.6; number of atoms in the ring closed by the hydrogen bond — 13; helical formula 3.6131; helix diameter ∽0.5 nm; pitch ∽0.54 nm; translation of a single amino acid residue along the helix axis (residue projection on the axis) ~0.15 nm (Fig. 2.9).

Only right-handed a-helices have been found in natural proteins. The amino acid side chains in an a-helix point outward and are located on opposite sides of its axis. Non-polar amino acid side chains typically group together on one side of the a-helix, forming non-polar patches; this creates conditions for the association of different helical segments.

In addition to the a-helix, Other Types of Helical structures have been described, and some of their parameters are given in Table 2.4.

As seen from Table 2.4, besides the a-helix, other types of helices are found in proteins, such as the 310 and n-helices, which occur rarely and mostly in short segments, forming 1–2 turns at the ends of an a-helix.

Fig. 2.9. Segment of a protein a-helical structure

Table 2.4. Helical configurations of proteins

Name

Formula

Radius, nm

Residue projection on axis, nm

Note

310

310

0.19

0.20

Strained structure, rarely occurs as 1–2 turns at the ends of an a-helix

a-Helix

3.613

0.23

0.15

Universal configuration

n-Helix

4.416

0.28

0.11

Rarely occurs as 1–2 turns at the ends of an a-helix

1 In recent years, a new notation system has been introduced. According to this system, the designation NM is decoded as follows: M is the minimum number of turns spanned by an integer N of residues; under this notation form, the a-helix is designated as 185.

Pleated structures of the polypeptide chain are formed when the torsion angles φ and ψ are close to —120, +135°. It should be noted that although pleated structures are distinguished from helical structures in the description of Protein secondary structure, both are actually helical in nature; it is just that in the case of pleated structures, the helix is highly extended. In pleated chains, the number of residues per "turn" is 2 (in a flat pleated sheet) or 2.3 (in a slightly twisted sheet), the residue projection on the axis is 0.33 nm, and the "helix" radius is 0.1 nm. Pleated segments of the polypeptide chain exhibit cooperative properties—that is, they tend to align adjacent to one another in the protein molecule—and form parallel and antiparallel ß-pleated sheets, or layers, which are stabilized by hydrogen bonds between the pleated segments of the chain (Fig. 2.10, 2.11).

Fig. 2.10. Parallel ß-structure of a protein

Fig. 2.11. Antiparallel protein ß-structure

Fig. 2.12. Antiparallel ß-structure and ß-bend

The antiparallel ß-structure is formed when the folded chain makes a U-turn and runs back along itself, i.e., in the reverse direction; a ß-bend is formed at the turning point (Fig. 2.12). A ß-bend comprises four sequentially arranged amino acid residues. The parallel ß-structure is assembled from segments of the polypeptide chain that run in the same direction. The antiparallel arrangement of chains provides the most favorable conditions for the formation of hydrogen bonds between them mediated by peptide groups. In the case of a parallel arrangement of chains within the ß-pleated sheet structure, the interchain hydrogen bonds are less stable. The side chains (radicals) of the amino acid residues (more precisely, the Ca — Cß bonds) are approximately perpendicular to the plane of the ß-pleated sheets, with the amino acid side chains oriented alternately on opposite sides of this plane.

Pleated sheets can be formed not only by a single polypeptide chain (in which case the hydrogen bonds will be intramolecular), but also by a group of closely associated polypeptide chains within the molecule (with hydrogen bonds forming between the chains). The second type of ß-structure is characteristic of Fibrous proteins such as Silk Fibroin and Hair keratin, which consist of multiple polypeptide chains. In globular proteins, approximately 15% of the amino acid residues of the polypeptide chain typically participate in the Formation of the ß-pleated structure. Most pleated sheets contain fewer than six chains. As a rule, pleated sheets are not flat; they exhibit a characteristic slight left-handed twist.

Certain Regions of the protein chain exhibit an irregular spatial arrangement of amino acid residues, which is likewise stabilized by hydrogen bonds and hydrophobic interactions. Such regions within a protein molecule are referred to as disordered, unstructured, or amorphous.

Short a-helices and ß-structures are formed in segments of varying length along the polypeptide chain, with unstructured regions located between these ordered structural types.

Fig. 2.13. General view of a left-handed double-stranded coiled-coil

2.4.5. Supersecondary Structure and Domains. The a-helical and ß-structural segments in proteins can interact with one another to form larger assemblies. The spatial architecture of such secondary structure assemblies is termed the supersecondary structure of the protein molecule. The Supersecondary structures found in native proteins are the most energetically favorable.

An example of a supersecondary structure is the supercoiled a-helix, in which two a-helices are twisted around each other to form a left-handed superhelix (Fig. 2.13). Short segments of this supersecondary structure are found in globular proteins (Bacteriorhodopsin, hemerythrin), and more frequently and in a highly ordered form, in fibrous proteins. Supercoiling is energetically advantageous because additional non-covalent (Van der Waals) contacts are established between the side chains belonging to different a-helices.

Superhelices can be formed by a-helices arranged either in a parallel or antiparallel orientation.

Another potential element of supersecondary structure is the ßxß-motif (Fig. 2.14), which consists of two parallel ß-sheets connected by a linker in the form of a random coil (ßcß), an a-helix (ßaß), or a ß-structure (ßßß). Two sequentially linked ßaß segments are referred to as a Rossmann fold (the ßaßaß-motif). This type of supersecondary structure is found in the NAD+-binding domain of dehydrogenases.

A supersecondary structure consisting of an antiparallel three-stranded ß-structure (ßßß) is called a ß-meander (or ß-zigzag). It is quite widespread in proteins, such as staphylococcal nuclease, lactate dehydrogenase, T4 lysozyme, and several others.

Fig. 2.14. Protein supersecondary structures. A — ßcß-motif; B — Rossmann fold (two sequentially connected ßaß segments); C — ß-meander: arrows indicate ß-pleated sheets, cylinders represent a-helices, and amorphous regions are shaded in black

The next level of organization, characteristic of large globular proteins, is the domain. Domains are structurally and functionally distinct regions (subregions) of the molecule connected to one another by short segments of the polypeptide chain known as hinge regions. Functional domains may consist of one or more Structural domains. Most likely, functional domains with a molecular weight exceeding 20,000 contain multiple structural domains.

Fig. 2.15. Domains of lobster Muscle glyceraldehyde-3-phosphate dehydrogenase. A — NAD+-binding domain; B — catalytic domain.

numbers indicate amino acid residues, other designations are the same as in Fig. 2.14

Most large globular Proteins can be subdivided into several structural domains containing 100–150 amino acid residues and having a diameter of about 2.5 nm. Structural domains have been identified, for instance, in the enzyme Glutathione reductase, which catalyzes The conversion of oxidized glutathione to its reduced form (the tripeptide Glu-Cys-Gly). This enzyme is a dimer, meaning its molecule is built from two polypeptide chains, or subunits. Each subunit, in turn, consists of three structural domains that perform specific functions during the enzymatic action.

For a number of enzymes, it has been demonstrated that the Active Site is located in a cleft between domains. For example, glyceraldehyde-3-phosphate dehydrogenase features two distinct functional domains: an NAD+-binding domain and a catalytic domain, which together form the active site (Fig. 2.15).

Fig. 2.16. Classification of globular proteins based on the presence and arrangement of a- and ß-structures in their molecules (after Yu. B. Filippovich, 1985):

A — arrangement of a- and ß-structures (indicated by circles and rectangles, respectively) in a-, ß-, (a+ß)-, and a/ß-proteins; the arrows denote the course of the polypeptide chain from the N-terminus to the C-terminus of the molecule;

B — structure of three representative a-, ß-, and a/ß-proteins, where the a-structure is depicted as a helix and the ß-structure as arrows. In myohemerythrin, the hatched circles designate two iron atoms, in erabutoxin the dashed lines indicate four disulfide bridges, and in flavodoxin the group responsible for hydrogen atom transfer is shown at the top as three condensed six-membered rings

In some globular proteins (Serine proteinases, immunoglobulins), the structural domains within the molecule are remarkably similar, suggesting gene duplication. Similar structural domains can also occur in different proteins, such as the NAD+-binding domain in dehydrogenases.

Across all domain-based globular proteins, there is a high degree of affinity between amino acid residues that are close to each other along the chain. This implies that domains fold independently of one another, which greatly simplifies the macromolecular folding process.

Based on the ORGANIZATION OF THE Secondary structure of the polypeptide chain—namely, the number of a-helices and ß-structures and their spatial arrangement within domains—it has been proposed to divide structural domains and proteins into five classes or groups (Fig. 2.16).

The first group (a-proteins) includes proteins dominated by a-helices. This category specifically comprises hemoglobin, Myoglobin, and calcium-binding proteins.

The second group (ß-proteins) contains proteins constructed primarily of antiparallel ß-sheets, such as concanavalin A (a lectin from the jack bean), rubredoxin (a simple iron-sulfur electron-transport protein), and chymotrypsin.

The third group comprises a + ß-proteins, which feature regions built entirely of a-helices alongside regions consisting entirely of ß-sheets (mostly antiparallel). Examples of this protein type include Papain (a proteolytic enzyme from papaya fruit), Thermolysin, insulin, cytochrome b5, ribonuclease, and lysozyme.

The fourth group consists of a/ß proteins, in which a-helices and ß-structures alternate along the chain. Most ß-structures, predominantly parallel, are localized in the central core of the molecule, where they bend in a propeller-like fashion ("twist structures") to form a rigid "scaffolding" that anchors the remaining PARTS OF THE molecule (Fig. 2.17). Proteins of the fourth group include carboxypeptidase, hexokinase, phosphoglycerate kinase, Triosephosphate isomerase, subtilisin, lactate dehydrogenase, and malate dehydrogenase.

The fifth group encompasses domains lacking a clearly defined secondary structure. The Spatial structure of a protein is significantly more conserved than its amino acid sequence. The folding pattern of the polypeptide chain sometimes reveals evolutionary relationships between proteins that cannot be detected by Amino acid analysis alone.

Fig. 2.17. Spatial folding of the flavodoxin polypeptide chain.

The arrows indicate the path of the chain from the N-terminus to the C-terminus within the twist structure region

2.4.6. Tertiary Structure. Tertiary structure characterizes the spatial arrangement of ordered and amorphous regions within the polypeptide chain as a whole, which is achieved through side-chain interactions and depends on their type and conformation. Thus, tertiary structure describes the three-dimensional folding of an entire protein molecule if it consists of a single polypeptide chain. Tertiary structure is directly related to the Shape of Protein molecules, which can vary widely from globular (spherical) to fibrous (thread-like). The shape of a protein molecule is characterized by its axial ratio (The ratio of the molecule's long axis to its short axis). Fibrous, or fibrillar, proteins are those with an axial ratio of 80 or greater; these include silk fibroin, the keratin of hair, horns, and hooves, Connective Tissue Collagen, and several other proteins. Proteins with an axial ratio of less than 80 are classified as globular; the majority of these have an axial ratio of 3–5. Consequently, the tertiary Structure of Globular proteins is characterized by a sufficiently dense packing of the polypeptide chain into a coiled molecule approaching a spherical shape.

Various types of bonds are involved in maintaining and stabilizing the Tertiary Structure of globular proteins (Fig. 2.18): covalent, ionic (or salt bridges), hydrogen bonds, and hydrophobic interactions (listed in order of decreasing bond energy). Hydrophobic interactions, which occur between nonpolar amino acid side chains, play a predominant role in the formation of tertiary structure.

Fig. 2.18. Bonds stabilizing the tertiary structure of a protein molecule:

I — ionic bond, II — hydrogen bond, III — hydrophobic interactions, IV — covalent bond

The covalent bonds involved in maintaining protein tertiary structure are represented by disulfide and peptide bonds, the latter formed by the amino and carboxyl groups of amino acid side chains. Disulfide bonds arise between two closely positioned SH groups of cysteine side chains. The closure of covalent disulfide bonds (—S—S—) results from The oxidation of sulfhydryl (thiol) groups in the presence of oxygen or certain Other reagents. In vivo, these bonds form spontaneously if thiol groups end up adjacent to each other as a result of polypeptide chain folding; in other words, disulfide bonds stabilize the molecular conformation rather than dictating the folding pattern of the polypeptide chain. In ribonuclease and lysozyme, which contain four disulfide bonds, theoretically 44 = 256 variants of S—S bonds are possible, yet only four are realized in the native conformation.

Disulfide bonds are most commonly found in secreted proteins (snake venoms, Peptide Hormones, digestive enzymes, milk proteins, etc.). The presence of a large number of disulfide bridges in fibrous proteins (such as keratin)—which are capable of interconverting with sulfhydryl groups, i.e., temporarily breaking and reforming—partially explains the viscosity and elasticity of these proteins. Disulfide bridges are never formed between adjacent cysteine residues.

Salt bridges, or ionic bonds, occur between oppositely charged protein groups—that is, between amino acid side chains dissociated according to acidic and basic types. The basic groups can include the e-amino group of Lys, the guanidine group of Arg, and the imidazole group of His (one of its nitrogen atoms exhibits basic properties, while the other is acidic). In addition to histidine, acidic groups in proteins can include the ß-carboxyl group of Asp and the y-carboxyl group of Glu.

Ionizable groups and polar amino acid groups (see Section 2.3.3) are typically located on the surface of the protein globule and are less frequently found in the interior. For instance, the surface of chymotrypsin contains 4 Arg and 14 Lys residues, yielding a net positive charge of 18 (+), as well as Asp and 5 Glu residues, which determine a net negative charge of 12 (—). Charged groups on the surface of the protein globule are usually solvited and surrounded by counterions, which enhances Protein solubility in aqueous media. Polar amino acid side chains located inside the protein molecule typically form hydrogen bonds with each other or with the polypeptide backbone.

The presence of charged groups inside the globule is energetically unfavorable, making them a rare occurrence there. If charged groups are nonetheless localized within the interior, they form salt bridges.

Hydrophobic side chains, lacking affinity for water, pack compactly primarily within the interior of the globule, forming hydrophobic cores that stabilize the molecular tertiary structure. Consequently, the dielectric permittivity (ε) inside a protein globule is significantly lower than that of water. The hydrophobic regions (cores) at the center of the protein globule exhibit a high packing density characteristic of many crystals. A lower packing density is observed for surface regions of the molecule, such as active enzyme sites, which aligns with the hypothesis regarding the mobility of active sites. On average, the packing density of a protein molecule is also quite high, demonstrating the efficient use of noncovalent forces in organizing the spatial structure of the protein molecule. A small fraction of nonpolar radicals may reside on the molecular surface and cluster together to form hydrophobic patches. Thus, overall, the surface of the protein globule is mosaic—predominantly hydrophilic, yet interspersed with small nonpolar areas.

Only after acquiring its native tertiary structure does a protein exhibit its specific functional activity. The spatial structure of a protein and the mutual arrangement of its constituent groups determine the BIOLOGICAL FUNCTIONS OF the protein molecule within the Organism. Active sites of enzymes, Ligand-binding sites responsible for integrating a given protein into a multienzyme complex and the self-assembly of supramolecular structures, as well as antigenic determinants of the protein, are formed by spatially converging groups during tertiary structure assembly. For example, the similar biological function of hemoglobin and myoglobin, driven by their ability to reversibly bind oxygen, is explained by the structural Homology of their polypeptide chains. Active sites in monomeric proteins form during the polypeptide chain folding process, whereas in multidomain enzymes, active sites incorporate regions from all structural domains of the given globular protein, positioning the active site right between them.

Fibrillar Proteins primarily perform a structural function in the living organism. They are poorly soluble or insoluble proteins characterized by a high content of non-polar amino acids. Examples include the proteins of connective and contractile Tissues, hair, Skin, and certain proteins found in the cell walls of plants, Algae, and various other sources. The molecules of fibrillar proteins are typically built from several polypeptide strands that exhibit an a-helical structure (a-Keratins, myosin), ß-pleated sheets (ß-keratins, fibroin), or a specialized twisted helical conformation (collagens).

The formation of long, elongated molecules by polypeptide strands fundamentally characterizes the tertiary structure of fibrillar proteins. Collagen is a key component of animal and human connective tissue, where it forms specialized strands known as collagen fibers. These fibers are constructed from fibrils, whose structural unit, in turn, is tropocollagen.

The tropocollagen molecule (M 300,000) consists of three polypeptide strands, each wound into a tight left-handed helix with three amino acid residues per turn. These three chains are gently twisted together into a right-handed superhelix, forming a tropocollagen molecule measuring 1.5 nm in diameter and 300 nm in length. The mutual arrangement of the peptide chains within the molecule is stabilized by hydrogen bonds between the peptide groups of adjacent chains and by covalent bonds involving lysine, hydroxylysine, allysine, and hydroxyallysine (the latter Two amino acids being products of the oxidation of ε-NH2 groups to aldehydes). The interaction of these specific amino acid residues yields Schiff bases and aldols.

Tropocollagen is distinguished by a high content of glycine (one-third of all amino acid residues) and imino acids (proline and hydroxyproline—21%), alongside a substantial amount of alanine (11%). 5-Hydroxylysine is found in small quantities within collagens, though it is rarely encountered in other proteins. Several types of collagens have been described, differing in their set of polypeptide chains and amino acid compositions. For instance, the triple helix of collagen in most vertebrates comprises two a1-chains and a homologous a2-chain, with the a1- and a2-chains showing minor differences in amino acid composition. Collagen fibrils are formed from tropocollagen molecules through end-to-end and side-by-side assembly. In boiling water, collagen partially dissolves to yield a gelatin solution, which sets into a gel upon cooling.

The proteins of hair, horns, skin, and feathers (a-keratins) are composed of 3–7 polypeptide chains containing approximately 100 amino acid residues each, arranged in an a-helical configuration. Interconnected by disulfide bridges, these polypeptide chains twist together to form a left-handed superhelix; these superhelical structures subsequently aggregate into microfibrils with a diameter of about 2 nm. When a-keratins are treated with hot steam, The system of intrachain hydrogen bonds within each polypeptide strand is disrupted, causing the strands to transition into a state of ß-pleated sheets (β-keratin) upon stretching. In ß-keratin, hydrogen bonds form between separate polypeptide strands. The repeating unit of keratins is the sequence: — cys — cys — glu — pro — ser —. Silk fibroin shares a similar spatial structure with ß-keratin; this protein is rich in glycine, serine, and alanine. In fibroin, adjacent chains are antiparallel, meaning their C- and N-termini do not coincide, and interchain S—S bonds are absent. Silk fibroin consists primarily of the following periodically repeating sequence: —gly—ser—gly—ala—gly—ala—. During the formation of the pleated structure, all the side groups of ala and ser project to one side of the chain, while those of gly project to the other. A structure analogous to a-keratin forms the basis of the Muscle Proteins myosin and Tropomyosin.

A fibrous fibrillar protein has been discovered in the supportive structures of diatoms, where it acts alongside SiO2 to help form the diatom Skeleton. Besides hydroxyproline, this protein contains dihydroxyproline, ε-N-trimethylhydroxylysine, and several other rare amino acids. Bromotyrosine and iodotyrosine have been identified in the skeletal proteins of corals, Sponges, and jellyfish. The primary Cell wall of higher plants contains extensin, a fibrillar protein exceptionally rich in hydroxyproline (up to 33%). Extensin exists as a rigidly twisted left-handed helix. This protein is anchored to the hemicellulose of the cell walls via a glycosidic bond linking arabinose or galactose to hydroxyproline.

Resilin, a protein rich in gly and ala, has been found in the exoskeleton of insects. It lacks hydroxyproline and Sulfur-Containing Amino Acids, bearing a strong resemblance to silk fibroin.

X-ray crystallography (XRC) is the most informative method for studying the spatial structure of protein molecules (secondary, supersecondary, and tertiary). It allows researchers to determine the electron density distribution within a crystal and, consequently, the three-dimensional architecture of the molecules comprising it. The results of protein analysis via XRC enable the construction of a spatial model of the molecule: mapping the trajectory of the polypeptide chain in 3D space, identifying regions that form regular structures or chain bends, and revealing the spatial orientation of amino acid residues, their contacts with other residues, and the attachment sites of non-protein moieties (in complex proteins or two-component enzymes). Information regarding the presence of ordered structures in a protein, their type, overall proportion within the chain, and conformational shifts in solution can be obtained using optical methods—such as ORD, CD, and IR spectroscopy—as well as nuclear magnetic resonance (NMR).

Presently, computational and theoretical methods are widely employed to predict the folding patterns of polypeptide chains based on primary structure data. This line of research is prominently represented at the Institute of Protein Research of the USSR Academy of Sciences.

2.4.7 Quaternary Structure. Proteins are said to possess a quaternary structure when their molecules consist of two or more polypeptide chains held together by non-covalent interactions. As a rule, quaternary structure is characteristic of proteins with a relative molecular mass exceeding 50,000–100,000. Proteins exhibiting a quaternary structure are referred to as oligomeric.

Quaternary structure refers to the spatial arrangement of individual polypeptide chains within a protein molecule and The Nature of the bonds between them. Each separate polypeptide chain within a quaternary protein is called a protomer or subunit. Some authors reserve the term subunit exclusively for a portion of the molecule that possesses functional activity; such a functional unit may consist of either a single protomer or multiple protomers. Complex supramolecular Protein Assemblies comprising up to several hundred subunits—such as bacterial flagella or viral capsids—are occasionally classified among proteins with quaternary structure. Such proteins are also frequently termed multimeric. The aggregation of protomers in an indefinite quantity that does not confer new biological properties to the protein is distinguished from Quaternary Structure and termed an aggregated state.

The majority of intracellular proteins are oligomeric, whereas extracellular proteins tend to be monomers of low molecular mass, and Plasma Proteins typically function as large monomers. This correlation between Protein Structure and localization is evidently not coincidental. The monomeric nature of extracellular proteins—such as digestive enzymes or salivary lysozyme—relates to the unpredictable fate of their molecules and the resultant evolutionary advantage of maintaining numerous independent units. The large molecular mass of plasma protein monomers (>60,000) prevents their filtration and excretion by the Kidneys, as well as their facile leakage into the extracellular space. Serum albumin, which contains multiple functional domains, serves as a classic example of such a protein.

The presence of numerous diverse Oligomeric Proteins inside the cell helps lower intracellular osmotic pressure and viscosity; furthermore, oligomeric proteins are effectively regulated by effectors (see Section 12.5). The Biological Significance of protein oligomerization is also tied to the fact that it requires less genetic material to encode if all or some of the subunits within the protein molecule are identical. Such oligomers carry a lower probability of generating defective molecules compared to large monomers of equivalent molecular mass. Defective subunits can be eliminated through dissociation and reassembly cycles. Most commonly, an oligomeric protein molecule comprises two or four subunits, less frequently 6, 8, or more, or an odd number. The mutual orientation of individual subunits in quaternary proteins and supramolecular assemblies varies, governed by subunit geometry and count in a manner that minimizes free energy. Both isologous and heterologous interactions between subunits are recognized within the macromolecule. In heterologous interactions, each subunit of a given type possesses mutually complementary regions a and b (Fig. 2.19). When two such subunits associate, region a remains exposed on one, and region b on the other. Additional identical subunits can attach to these exposed sites, yielding either extended chains, rings, or spirals. Rings formed exclusively via heterologous interactions exhibit cyclic symmetry: rotation around the symmetry axis by a specific angle superimposes each subunit onto the next. Depending on subunit geometry, the number of subunits within a ring can vary.

If the angles generated between two subunits prevent the closure of a ring, a spiral is formed. In this scenario, alongside heterologous contacts of the ab type, supplementary interactions may occur if other complementary sites are available. Various types of helical structures built from protein subunits include the actin filaments of muscle fibers, the virions of certain viruses, and bacterial flagella.

When two identical protein subunits form a pair of identical ab-type bonds, the interaction is termed isologous (Fig. 2.20). Such a structure possesses a twofold axis of symmetry, meaning every point on one subunit can be mapped to the corresponding point on the other subunit by rotation around the symmetry axis by 360°/2 = 180°. When two isologous dimers overlap, additional isologous bonds may form between them, yielding tetrameric structures with dihedral symmetry. Examples of such structures include the enzyme lactate dehydrogenase and the plant agglutinin concanavalin A.

Large protein architectures feature two primary types of contacts: heterologous and isologous. These typically give rise to oligomers of various shapes possessing cubic symmetry (polyhedra resembling tetrahedrons, cubes, or icosahedrons). Structures with cubic symmetry possess more than one axis of symmetry of an order higher than two. Cubic symmetry is characteristic of the capsids of numerous virions, such as bacteriophage φX174, SV40 virus, human papillomavirus, Adenoviruses, and Influenza virus.

The quaternary structure of many proteins and supramolecular complexes is assembled from protomers belonging to two or more distinct types. For instance, hemoglobin is a tetramer built from two similar yet non-identical subunits, a and ß (a2β2), while the enzyme aspartate transcarbamylase comprises 12 subunits (two trimers acting as catalytic subunits and three dimers acting as Regulatory Subunits). The protein coats of certain viruses are likewise constructed from protomers of two or more types. Structures formed from non-identical subunits similarly exhibit a symmetrical arrangement.

Fig. 2.19. Heterologous association of subunits. A — ring; B — spiral

Fig. 2.20. Isologous association of subunits (see text for explanation)

Various types of bonds can form between the subunits constituting a quaternary protein: hydrophobic interactions, ionic bonds, and hydrogen bonds. The Contact surfaces of subunits that form the quaternary structure are complementary, which contributes to the thermodynamic stabilization of the protein molecule. Association Specificity is achieved through the geometric complementarity of contact surface profiles, alongside the optimal alignment of hydrogen bond Donors and acceptors and charged residues. Nonetheless, hydrophobic interactions remain the primary driving force stabilizing the quaternary structure. Hydrophobic regions ("sticky patches") on the contacting surfaces of protomers prevent water from forming intermolecular bonds, thereby immobilizing it. In accordance with The Second Law of Thermodynamics, the tendency of a system containing water molecules to maximize entropy—achieved when water molecules are free to form hydrogen bonds with one another—forces the protomers to close in on themselves so that water is excluded from the interface. These hydrophobic patches coalesce, while the complementary electrostatic charges at specific points reinforce the bonding between protomers and subunits. Since individual protomers within the molecule are linked non-covalently, the question naturally arises as to whether this constitutes a single molecule. Because all protomers of a quaternary structure are robustly bound together (albeit non-covalently) and exclusively perform their cooperative biological function when assembled, the entire system should be regarded as a single molecule.

Electron Microscopy is a valuable technique for analyzing quaternary structure, yielding useful insights provided the protein is sufficiently large, typically possessing a molecular mass exceeding 200,000.

Conclusions regarding the presence of quaternary structure can also be drawn by determining molecular masses under varying conditions: near-native conditions versus in the presence of denaturing chemical agents that dissociate the molecule into its constituent protomers (such as guanidine hydrochloride, sodium dodecyl sulfate, urea, or high concentrations of neutral salts). If identical molecular mass values are obtained in both instances, the protein lacks a quaternary structure and is built from a single polypeptide chain.

In cases where the protein is an oligomer composed of uniform subunits, the molecular mass measured in the presence of chemical agents should represent an exact fraction of the native protein's mass. For example, if a protein contains two protomers, its molecular mass following Denaturation will equal half of the native value, and so forth. The presence of subunits of differing sizes within a protein is determined using Polyacrylamide gel electrophoresis or Gel filtration on molecular sieve columns, such as Sephadex. Under such conditions, treating oligomeric proteins with sodium dodecyl sulfate leads to molecular dissociation into individual subunits, revealing two or more distinct fractions that differ from the original protein in molecular mass.

The quaternary structure plays a vital role in regulating the biological activity of proteins, as it is highly sensitive to external conditions. Minor environmental shifts can alter the relative arrangement of subunits and, consequently, the protein's biological function. This phenomenon is a fundamental mechanism of Metabolic Regulation, given that numerous enzymes and other metabolically active proteins possess a quaternary structure.

2.4.8. Interrelation of Structural Levels, Orderliness, and Relative Dynamics of the Protein Molecule. A detailed examination of globular and fibrillar protein architectures has demonstrated that each individual protein features a unique spatial structure, or conformation. The primary structure—namely, the genetically determined amino acid sequence—plays a leading role in shaping this conformation. This is evidenced, in particular, by the intrinsic ability of polypeptide chains to spontaneously fold in space, adopting the conformation dictated by their primary structure.

A Classification of amino acid residues based on their propensity to form specific regular structures has been proposed. Amino acids such as glu, ala, and leu favor the formation of the a-helix, whereas met, val, and ile are more commonly found within ß-structures, and gly, pro, and asn tend to occur at chain bends (where they act as a-helix destabilizers). When several helix-promoting residues happen to be positioned adjacently along the chain, the formation of this structure is triggered. It is generally accepted that a cluster of six residues, four of which favor helicity, can serve as a nucleation site for helix formation. From this center, the helix propagates in both directions until it encounters a tetrapeptide segment composed of residues that inhibit helical structure. In the formation of the ß-sheet, three out of five amino acid residues that favor the ß-structure act as primers.

The premise that the primary structure dictates native conformation is further corroborated by protein refolding experiments. Reversing denaturing conditions often leads to spontaneous renaturation, thereby restoring both the native conformation and specific biological activity. For instance, C. B. Anfinsen (1975) successfully established renaturation conditions that allowed RNase to revert to its original Native State following denaturation.

The primary structure of protein subunits also dictates their quaternary structure, a fact supported by X-ray crystallography data and the successful reconstitution of biologically active proteins from dissociated subunits. For example, treating the globin moiety of hemoglobin with acetone at an acidic pH causes denaturation, heme dissociation, and subunit separation. Neutralizing the pH leads to the recombination of heme and globin, followed by subunit association to form native hemoglobin.

It is hypothesized that a protein's primary structure determines not only its conformation but also the pathways and sequence of events required to achieve it. Protein molecules presumably contain relatively autonomous "nodes" or "blocks" that fold into their respective conformations independently of one another. These blocks may consist of a-helices and ß-structures that exhibit a high affinity for neighboring amino acid residues along the chain. This facilitates the spatial folding process of the polypeptide chain and accelerates it. During the formation of a protein's spatial structure, intermediate states distinct from the native conformation are also observed. The folding of proteins containing disulfide bonds proceeds significantly slower than that of proteins lacking S—S bonds. These structural blocks subsequently fold into two to three large domains, which then associate and undergo final compaction into the complete molecule. In conclusion, while the amino acid sequence is the primary and decisive factor in forming the protein's conformation, it is not the sole determinant; other limiting and corrective factors exert a certain influence as well. For instance, the Formation of secondary structure in a specific region of a polypeptide chain depends not only on its local amino acid sequence but also on the influence of distant regions.

The specific conformation of a polypeptide chain begins to form while the protein is still being synthesized on the ribosome. This has been proven by experiments involving Ribosomes isolated along with the nascent enzyme molecules they were synthesizing. Although the ribosomes were isolated before translation was complete—meaning the protein molecules were still attached—catalytic activity was already detectable, a state achievable only when the tertiary structure is nearly fully formed. The organization of the native structure is ultimately completed upon the termination of polypeptide synthesis and subsequent post-translational modification mediated by various enzymes, or, in the case of oligomeric proteins, the assembly of the quaternary structure. The resulting unique structure of the protein molecule is exquisitely adapted to execute a specific biological function within the cell, be it catalytic, hormonal, transport, or otherwise.

Native protein conformations have been selected throughout evolution and represent the most energetically favorable states possessing the lowest free energy. Such relatively stable spatial configurations govern the specific biological Functions of Proteins. Native protein molecules with a defined three-dimensional structure exhibit a higher degree of order compared to an extended polypeptide chain and, consequently, should be characterized by a lower entropy than unordered protein structures. Based on this, one might expect that the formation of an unordered protein structure from an ordered one would be thermodynamically more favorable. However, because proteins invariably reside in an aqueous cellular environment, one must consider the entire protein–water system rather than an isolated protein. When a protein adopts an ordered three-dimensional structure, the surrounding water molecules achieve maximal entropy; overall, the entropy of the system (protein plus water) either remains constant or increases during the folding process. This precise thermodynamic balance enables the spontaneous formation of ordered structures endowed with specific biological functions. The formation of protein quaternary structures via subunit association is governed by the same principles as the folding of an individual polypeptide chain into its native conformation.

Nevertheless, the orderliness of protein spatial structures is dynamic rather than static. Amino acid side chains located on the molecular surface, for instance, possess considerable freedom of movement. Spontaneous and reversible conformational transitions occur within protein macromolecules, meaning that each macrostate of a protein corresponds to a multitude of microstates. These microstates arise from both small-scale thermal fluctuations of atomic groups around their equilibrium positions and local transconformations or domain movements. Under physiological conditions, small-scale thermal fluctuations and medium-scale local transconformations predominate. Their probability depends on various functional impacts on the protein, even though its statistical macroconformation may remain virtually unchanged.

During their functioning, protein molecules undergo minor conformational changes (fluctuations). Such fluctuations are observed, for example, when the protein moiety of an enzyme binds a coenzyme to form an enzyme-substrate complex. B. F. Poglazov (1967) discovered that the contraction of the T4 bacteriophage tail sheath is accompanied by a dramatic decrease (up to 50% of the total content) in the number of sheath protein a-helices and a commensurate increase in ß-structures.

The conformation of proteins is significantly influenced by environmental conditions (such as pH, ionic composition, and Temperature). Altering these environmental parameters leads to the charge redistribution of specific sites on the protein molecule, restructuring the network of ionic bonds, hydrogen bonds, and hydrophobic interactions between amino acid side chains—in other words, altering the protein's conformation and biological activity.

2.4.9. Self-Assembly of Supramolecular Protein Structures. The information encoded within the amino acid sequence is also expressed during the self-organization (or self-assembly) of supramolecular structures, which occurs without the participation of a template or Genetic control. Alongside the self-organization of tertiary structure, self-assembly refers to the intrinsic capacity of protein molecules to undergo spontaneous and orderly association with one another and with other biopolymers, resulting in the formation of biologically active structures.

The self-assembly and reconstitution of tobacco mosaic virus (TMV) in vitro was achieved as early as 1955. In this cylindrical virus, an RNA helix forms an axis around which over 2,000 protein subunits are arranged. Raising the pH or treating the virus with detergents causes this protein coat to dissociate into individual polypeptide subunits. Conversely, acidifying the solution induces the self-assembly of Viral Particles, as protein subunits reassociate around the RNA. These reconstituted particles retain their infectivity, demonstrating their resemblance to the original native virus—a conclusion further supported by a comparison of their physical properties.

Subsequently, similar artificial assembly of biological structures from subunits was accomplished using the T4 bacteriophage tail fiber, membrane structures, microtubules, and other objects. The assembly of the T4 phage tail fiber begins with the formation of a "hub" from specialized proteins, which serves as the center of the baseplate. Other proteins then assemble around this hub to form the hexagonal baseplate (Fig. 2.21). Activation of the baseplate via the attachment of specific proteins triggers the assembly of the tail tube from structural units, followed by the sheath. In the next stage, phage HEAD proteins approach the tail, the head is formed, and completed tail fibers subsequently attach to the baseplate. Microtubules—cytoplasmic elements involved in cellular motility and Intracellular Transport—are formed from molecules of the globular glycoprotein tubulin, which arrange themselves into a helix. The self-assembly of microtubules is triggered by the activation of the protein utilizing GTP energy. The assembly and disassembly of these structures, which occur continuously within the cell, are regulated by the intracellular concentration of Calcium Ions.

Fig. 2.21. Successive stages of T4 bacteriophage tail fiber assembly:

a — baseplate with spikes, b — tube, s — sheath subunits

Membrane formation also occurs as a result of the self-assembly of its components: proteins, Lipids, CARBOHYDRATES, and others. The formation of membranes and various cellular Organelles does not require a genetic code; rather, organelles arise from the interaction of molecules possessing specific chemical structures. This concept is substantiated by numerous experiments. For example, adding an integral membrane protein to a suspension of membrane phosphoglycerides leads to the formation of vesicles in which all protein molecules are embedded in The Lipid Bilayer and oriented in a specific manner.

Extensive and fascinating research on the self-assembly of bacterial flagella was conducted by B. F. Poglazov (1967). In experiments by A. S. Spirin (1968), ribosomes were first "disassembled" (using high concentrations of CsCl and other monovalent cation chlorides) and subsequently self-assembled (at a high concentration of Mg2+ and heating to 37°C).

The formation of enzyme-substrate and antigen-antibody complexes can likewise be regarded as a manifestation of self-assembly and self-organization.

During the self-assembly of supramolecular structures, analogously to the formation of native protein conformations, the system tends toward a state of minimal free energy and maximal entropy. The surrounding water molecules increase their entropy by an amount sufficient to compensate for the decrease in entropy resulting from molecular association during self-assembly.

Ionic interactions play a major role in the aggregation of Biomolecules, as they are the longest-range forces (acting up to 0.7 nm) and come into play first. These are subsequently followed by shorter-range bonds (at distances up to 0.2 nm) between the molecules: hydrogen, hydrophobic, and van der Waals bonds. For these forces to arise and operate, steric and spatial complementarity of the interacting surfaces is paramount. In other words, these surfaces must be able to approach one another closely enough for the aforementioned bonds to form. Complementarity in the distribution of opposite charges (to generate electrostatic forces), hydrophobic domains, and hydrogen-bonding groups is also required. Thus, the phenomenon of "recognition" and the presence of steric and chemical complementarity play a critical role in the self-assembly process.

Self-assembly can be viewed as the simplest stage of biological morphogenesis at THE MOLECULAR LEVEL. During the formation of elementary biological structures, self-assembly serves essentially as the sole basis of the process (assisted by a number of secondary factors). More complex biological formations, however, require the additional participation of various regulatory and corrective mechanisms. It is reasonable to surmise that the self-assembly of biological structures was of paramount importance at the dawn of life during the transition from non-living to living matter.



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.