Fundamentals of Biochemistry - Filippovych, Y. B. 1999
Proteins
Structure of the Protein Molecule
Among A large number of diverse hypotheses, only one has stood the test of time: the polypeptide theory of protein molecular Structure, first proposed by E. Fischer in 1902 based on the ideas put forward by A. Ya. Danilevsky regarding The Role of —СО—NН bonds in Cell/13.html">Protein Structure. According to this theory, protein molecules are giant Polypeptides composed of several dozen or sometimes hundreds of amino acid residues that constantly occur in Proteins (see Table 6).
Class="center">
Fig. 22. Biologically active Peptides
For decodings of Abbreviations, see Table 4. Numbers near the names of amino acid residues indicate their position in the molecule, counting from the N-terminus to the C-terminus of the peptide. Arrows originate from Amino Acids involved in The formation of peptide bonds by COOH groups. Adjacent Cysteine residues in cycles are linked by disulfide bridges. The NH2 group enclosed in parentheses denotes an amide group. It is implied that amino acid residues are connected by peptide bonds. Oxytocin and vasopressin are Hormones of the posterior Pituitary Gland. Upon entering the Blood, the former stimulates the contraction of Muscle fibers located around the alveoli of the Mammary Glands and uterine Muscles, while the latter acts primarily on the smooth muscles of Blood Vessels and participates in the Regulation of the body's Water balance at the renal level. Phalloidin is the toxic principle of the death cap mushroom, which causes the death of the Organism at negligible concentrations due to the leakage of K+ Enzymes out of Cells; the bond between the cysteine and Tryptophan residues in the phalloidin molecule is formed through the interaction of the sulfhydryl group and the pyrrole ring of their radicals. Gramicidin is an antibiotic active against many gram-positive Bacteria (pneumococci, streptococci, staphylococci, etc.); it alters the permeability of Introduction/36.html">Biological Membranes to low-molecular-weight compounds and causes cell death. Thyroliberin is a hormone synthesized in the Hypothalamus that ensures the release (hence its alternative name, releasing factor) and possibly enhances the Biosynthesis in the pituitary gland of another hormone—thyrotropin—which, in turn, controls The activity of The Thyroid Gland and the formation of thyroxine therein; this hierarchy in the Regulation of Hormone biosynthesis is a striking example of the participation of peptides in Metabolic Regulation. Met-enkephalin is a pentapeptide produced in Nervous Tissue that alters (alleviates) pain sensations. It binds to the receptors affected by narcotics, such as morphine, and is an endogenous substance with psychotropic action. Somatostatin is a peptide that inhibits the release from the anterior pituitary gland of the Growth Hormone synthesized there—somatotropic hormone. Antamanide is a cyclic decapeptide that forms complexes with Na+, Ca2+, and certain other cations, which serve as prototypes of structures responsible for Ion transport across biological membranes.
The fact that protein molecules are constructed in precisely this manner is evidenced by the following facts:
1. In native proteins, it is possible to detect a very small number of free NH2 and COOH groups, since terminal residues constitute a negligible fraction of the huge polypeptide chain of a protein.
2. Protein Hydrolysis is accompanied by the gradual release of NH2 and COOH groups in a strict 1:1 ratio, i.e., The breakdown of —СО—NH peptide bonds occurs.
3. The interaction of biuret H2N—СО—NH—СО—NH2 with alkali and copper sulfate solutions produces a blue-violet coloration (a reaction for the presence of —СО—NH bonds); all proteins give a strong biuret reaction, meaning their molecules indeed contain a large number of peptide bonds.
4. The polypeptide nature of A number of proteins has been proven by chemical synthesis.
5. X-Ray Diffraction Analysis of protein structure reveals a continuous polypeptide chain with an arrangement of amino acid residues characteristic of the protein under study.
6. The Methods used to determine The sequence of amino acid residues in proteins (see below) unambiguously indicate the sequential (in some cases, selective) Cleavage of peptide bonds specifically within polypeptide chains.
A characteristic feature of polypeptide chain structure is that the C and N atoms in its backbone, composed of monotonously repeating —СО—CH—NH fragments, lie approximately in the same plane, whereas the atoms and radicals of the >CHR groupings are oriented at an angle of 109°28' to this plane. Furthermore, in adjacent amino acid residues, the arrangement of H atoms and radicals is opposite. This is clearly seen from Fig. 23, A and The structure of the polypeptide chain fragment of the Insulin molecule (see p. 54).
Another feature of the polypeptide chain structure lies in the specific Nature of the —СО—NH bond. The distance between the C and N atoms in the peptide bond is 0.1325 nm, which is smaller than the normal distance between the α-carbon atom and the N atom of the same chain, equal to 0.146 nm. At the same time, it exceeds the distance between C and N atoms connected by a double bond (0.127 nm). Thus, the C—N bond in the —СО—NH grouping can be considered intermediate between a single and a double bond due to the conjugation of the carbonyl group π-electrons with the lone pairs of the nitrogen atom. This exerts a definite effect on The properties of polypeptides and proteins: tautomeric rearrangement easily occurs at the sites of peptide bonds, leading to the formation of an enol form of the peptide bond characterized by enhanced reactivity:
![]()
The third peculiar feature of the Peptide bond Structure is that all atoms comprising it lie in a single plane (Fig. 23, B): this configuration is called planar and is maintained by the delocalization of the electronic density of the carbonyl group's double bond onto the bond between carbon and nitrogen (Fig. 23, Б).
Finally, the fourth property of the polypeptide chain is that its backbone, built of —NH—CH—СО fragments, is surrounded by side chains chemically diverse in nature, as can be seen when examining the STRUCTURE OF THE insulin molecule fragment (see p. 55).
Side chains are functionally versatile. They are represented (see Table 4 and Fig. 18) by aliphatic (ala, val, leu, ile), aromatic (phe), and heterocyclic (pro, trp, his) radicals. Many of them bear free amino (lys, arg), carboxyl (asp and glu), hydroxyl and phenolic (ser, thr, tyr), thiol (cys), amide (asn, gln), and other functional groups. By interacting with surrounding solvent molecules (water), the free NH2 and COOH groups ionize, forming cationic and anionic centers of the protein molecule. Depending on their ratio, the protein molecule acquires a net positive or negative charge and is characterized by a specific medium pH value upon reaching the isoelectric point of the protein. The number and ratio of cationic, anionic, and other groups in certain proteins are given in Table 7.
Table 7 Distribution of groups in proteins (number of groups per molecule)
Protein |
Total number of amino acid residues |
Hydrocarbon radicals |
Heterocyclic radicals |
Phenolic radicals |
Hydroxyl groups |
Amino groups |
Carboxyl groups |
Thiol groups |
Other radicals |
Horse Myoglobin, M = 17000 |
143 |
48 |
16 |
2 |
13 |
20 |
29 |
0 |
15 |
Pepsin, M = 35 000 |
341 |
136 |
21 |
18 |
72 |
3 |
71 |
4 |
26 |
Egg albumin, M = 46000 |
387 |
141 |
24 |
9 |
52 |
35 |
84 |
5 |
37 |
Horse Hemoglobin, M = 68000 |
541 |
209 |
63 |
11 |
59 |
52 |
89 |
3 |
55 |
The sharp predominance of carboxyl groups over amino groups in the pepsin molecule requires a high hydrogen ion concentration to bring it into the isoelectric state, whereas an almost equal number of both in the myoglobin molecule corresponds to an isoelectric point close to the pH of a neutral medium (see Table 3).
The structure of radicals exerts a major influence on many other Properties of Proteins and especially on the spatial configuration of the polypeptide chain: salt, ester, disulfide, and other bridges formed between them stabilize the relative arrangement of chain segments during the Formation of the protein's tertiary structure (see below). Radicals also largely determine the range of Chemical Reactions characteristic of protein bodies and, along with other factors, have a significant impact on the functional activity of proteins.
For a long time, concepts regarding the structure of the protein molecule were limited to recognizing the polypeptide chain as its backbone without any detailing of the regularities of amino acid residue sequencing or chain configuration. Only The Development of new Methods for determining the sequence of amino acid residues in the polypeptide chain and the refinement of X-ray diffraction Analysis of Protein crystals made it possible to advance our understanding of both the first (regularities of amino acid alternation) and second (regularities in polypeptide chain configuration) problems.
At present, the primary, secondary, tertiary, and quaternary structures of the protein molecule are sufficiently well understood. METABOLISM/2.html">THE CONCEPT OF the four Levels of Protein molecular structure was first put forward by Linderstrøm-Lang based on methodological considerations. Naturally, all the listed structural levels coexist within a protein molecule, and their unique combination in each specific case determines the overall architecture of the protein particle.

Fig. 23. Standard values of interatomic distances (in nanometers) and Bond Angles in the polypeptide chain (A) and electron density delocalization across the peptide unit (B), leading to the stabilization of its planar configuration (C and D)
Fragment of the polypeptide chain of the insulin molecule (amino acid residues 1–9 of chain B)

Primary Protein Structure. Primary protein structure refers to the linear sequence of amino acid residues in one or more polypeptide chains that make up a protein molecule. Knowing the Primary Structure of a protein allows one to write its complete chemical formula.
Figure 1. Sequence of operations for determining the primary structure of a protein

Given that a protein molecule contains at least several dozen amino acid residues, determining the exact Location of each is a highly challenging task. This can be achieved by carrying out the series of operations outlined in Figure 1.
Cleavage of disulfide Bonds in the protein molecule (Figure 1) is necessary, on the one hand, to prepare for the subsequent fragmentation of the protein's polypeptide chain and, on the other hand, to convert all Cys residues in the protein into cysteic acid residues, which are stable during further analysis:

Selective hydrolysis of the denatured protein is carried out using Proteolytic Enzymes, most commonly either Trypsin (specific for Arg and Lys residues) or Chymotrypsin (specific for Trp, Phe, and Tyr residues):

Proteins can also be degraded using chemical agents, such as Cyanogen bromide for Met residues, 2,4-dinitrofluorobenzene for Cys residues, N-Bromosuccinimide for Tyr or Trp residues, etc.:

As a result, two to three different sets of peptides are obtained from the studied protein, enabling the reconstruction of its primary structure in The final stage of the work.
Peptide fractionation is a multi-step, time-consuming, and laborious yet essential component of determining the primary structure of a protein. This is achieved primarily through high-voltage (up to 10,000 V) paper Electrophoresis—especially during the final fractionation stage—along with Various Forms of Chromatography, including ion-exchange and partition chromatography.
Determining the Amino Acid Sequence of individual peptides is the most critical Procedure in establishing the Primary Structure of Proteins. Currently, this is performed predominantly using either the Edman phenylisothiocyanate method or mass spectrometry.
The Edman phenylisothiocyanate method (1950) is carried out in a specially designed instrument known as a sequenator (from the English word sequence). The schematic diagram of a solid-phase sequenator and its operating principle are shown in Fig. 24. Modern models of sequenators can determine the sequence of amino acid residues not only in short peptides but also directly in The polypeptide chains of proteins, analyzing up to 70 amino acid units in a single molecule.

Fig. 24. Diagram of the sequenator setup (explanations in the text)
The Edman method involves treating the protein or peptide (Fig. 25)—attached via its C-terminal amino acid to an inert support (such as polystyrene or porous Glass) within the sequenator Column—with phenylisothiocyanate (coupling reaction). After washing the column with Solvents (methanol, dichloroethane), the resulting phenylthiocarbamyl peptide is treated with anhydrous trifluoroacetic acid. This releases an anilinothiazolinone containing the N-terminal amino acid (cleavage reaction), while the peptide or protein, shortened by one amino acid residue, remains bound to the support. The anilinothiazolinone is sent to a fraction collector and subsequently converted in an aqueous medium into a phenylthiohydantoin (conversion reaction) for chromatographic identification of the cleaved N-terminal amino acid. Once washed with solvents, the sequenator is ready for the next automated cycle.

Fig. 25. Chemical reactions occurring during Peptide and Protein sequencing (explanations in the text)

Fig. 26. Schematic diagram of a mass spectrometer:
1 — sample injection device; 2 — ionization chamber; 3–5 — filter system; 4 — electromagnet; 6 — detector; 7 — recorder drive; 8 — recorder chart
The mass spectrometric METHOD FOR DETERMINING the primary structure of peptides, currently applicable only to relatively short peptides of about 10 residues, involves amino acid-type fragmentation of acyl peptide esters followed by the Separation of the resulting positively charged fragments. Fragmentation is achieved by electron impact, and fragment separation is performed in a mass spectrometer (Fig. 26). This yields a mass spectrum of the peptide fragments (Fig. 27), the interpretation of which allows Conclusions to be drawn about the primary structure of the peptide. Although interpreting the mass spectrum of peptide fragments is quite complex, knowing the Amino Acid Composition of the peptide and using computers to calculate various combinations of amino acid residues allows researchers to identify the structural variant whose calculated data match the experimental results.

Fig. 27. Mass spectrum of methyl ester of N-decanoyl-leucyl-alanyl-alanyl-Alanine
The figure shows the mass numbers corresponding to The amino acid fragmentation pattern of the peptide; as is clear from the name of the studied compound, the amino group of the N-terminal
amino acid — leucine — is acylated and carries a decanoyl
group, whereas the carboxyl group of the C-terminal amino acid — alanine — is esterified and carries a methyl group. It is precisely this peptide derivative that undergoes amino acid-type fragmentation, i.e., predominantly at peptide bonds with the formation of fragments characterized by the corresponding mass numbers
The concluding stage of determining the sequence of amino acid residues in a protein molecule is the reconstruction of the primary protein structure based on data regarding their arrangement in the peptides obtained via selective hydrolysis. Since the targeted cleavage of peptide bonds during primary structure analysis is performed using at least two agents that ensure bond breakage at different points along the polypeptide chain, the primary structures of both sets of resulting peptides inevitably overlap. This makes it possible to reconstruct the primary protein structure (Fig. 28). If a Sequencer is successfully used to determine the sequence of amino acid units in large fragments of the protein molecule, this significantly simplifies the reconstruction process.
Other methods for determining the amino acid sequence in a protein molecule are primarily based on The Use of enzymes capable of accelerating the successive cleavage of amino acids from either the N- or C-terminus of the polypeptide chain. However, these methods have not yet achieved the widespread use and instrumentation as those discussed above. In recent years, the leading role in elucidating primary protein structures has been assumed by a method that can be provisionally called genetic: it is based on deducing the amino acid sequence of a protein molecule from the structure of the Gene encoding its biosynthesis. This is precisely how the structures of the longest polypeptide chains were deciphered — blood clotting factor VIII (2,332 amino acid residues) and the thyroglobulin subunit (2,750 amino acid residues). Mention should also be made of recently emerging reports on determining primary protein structures using laser photodissociation; given its exceptionally high sensitivity (requiring only 5 nmol of protein with a Molecular Weight of 50 kDa for analysis), it apparently has a great future.
The first protein whose primary structure was successfully deciphered through the work of F. Sanger et al. (1951–1953) was bovine insulin (for which F. Sanger was awarded the Nobel Prize in 1958). Insulin is a protein hormone that regulates Carbohydrate Metabolism. It is synthesized in the Pancreas as a precursor, the structure of which is shown in Fig. 29. Impairments in insulin biosynthesis in humans lead to Diabetes Mellitus.
The insulin subunit (M = 6000) consists of 51 amino acid residues assembled into two polypeptide chains (A and B) connected by two disulfide bridges.
According to L. Croft's directory as of August 1975, the primary structure of 932 peptide-nature compounds had been elucidated, of which at least 800 were proteins. Considering that in subsequent years from several dozen to several hundred primary structures were deciphered annually (76 in 1976; 380 in 1977), the list of proteins with a completely known amino acid sequence reached 2,657 by January 1985. In a number of cases, the primary structure of the same protein isolated from different organisms has been determined. For instance, the primary structure of insulin, myoglobin, and hemoglobin has been studied in 20 species, cytochrome c in 60, Lysozyme in 10, and so on. Below are the primary structures of these proteins characteristic of humans:

Fig. 28. Reconstruction of the primary structure of leghaemoglobin based on data on the primary structure of peptides obtained by its selective hydrolysis with trypsin (T), chymotrypsin (C), and Thermolysin (Tl) — a microbial proteinase specific for hydrophobic amino acid residues, as well as by its cleavage with cyanogen bromide (CNBr)
Leghaemoglobin is an oxygen-binding heme protein of legumes that plays an important role in MOLECULAR Nitrogen Fixation by ROOT nodule bacteria. T-1, T-2, etc. are tryptic peptides; C-1, C-IV, etc. are chymotryptic peptides; Tl-XXX, Tl-XXXI, etc. are thermolytic peptides; CNBr denotes cyanogen bromide peptides. The overlapping Zones of the arrows indicate overlaps of peptides from different hydrolysates, and the points of Selective Cleavage of the polypeptide chain are located at the junctions of the arrows. The primary structure of fragments containing amino acid residues 1 to 24 and 29 to 58 was determined using a sequencer

Fig. 29. Primary structure of bovine proinsulin
The C-peptide is cleaved from proinsulin during its conversion into insulin under the action of a specific enzyme that accelerates the hydrolysis of peptide bonds at Arginine residues at positions 31 and 60

At the Shemyakin Institute of Bioorganic Chemistry, the primary structure analysis has been performed for the following proteins: the ß- and ß'-subunits of DNA-dependent RNA polymerase — proteins with very long polypeptide chains of 1,342 and 1,407 amino acid residues, respectively; pig Heart aspartate aminotransferase — 412; yellow lupine nodule leghaemoglobin — 153; Escherichia coli LIV (leucine-, isoleucine-, valine-) binding protein — 344; neurotoxins from Central Asian cobra venom — 72; Escherichia coli ribosomal proteins L10 and L25 — 164 and 94, respectively; DNA-dependent RNA polymerase a-subunit — 329; insectotoxin from Central Asian scorpion venom — 62; visual rhodopsin — 348; elongation factor G — 701; bovine retina GTP-binding protein y-subunit and cGMP phosphodiesterase y-subunit — 69 and 87, respectively; bovine cerebellar C39 protein — 341; and pig Kidney Na+, K+-ATPase a- and ß-subunits — 1,016 and 302 amino acid residues, respectively.
In our country, the primary structures have been deciphered for: porcine pepsin — 313 residues (V. M. Stepanov et al.); sturin — 27 (A. B. Silaev); actinoxanthin — 107 (P. Reshetov); lactogenic hormone — 198 (N. A. Yudaev); nuclear polyhedrosis virus polyhedron protein — 244 (S. B. Serebryany et al.); Central Asian scorpion venom neurotoxin — 65 (E. Grishin et al.); Cholesterol-hydroxylating cytochrome P-450 from bovine adrenal cortex Mitochondria — 481 (V. A. Chashchin et al.); whale pituitary luteinizing hormone a- and ß-subunits — 96 and 118, respectively (V. S. Karasev et al.); human Brain calmodulin — 148 (Yu. B. Alakhov et al.); guanyl-specific RNase from several Aspergillus species — 102 (S. V. Shlyapnikov et al.); Escherichia coli Eco RV restriction endonuclease — 241 (A. A. Baev et al.); sea anemone neurotoxin III — 48 (T. A. Zykova et al.); Influenza virus neuraminidase, nucleoprotein structural protein, and hemagglutinin — 470, 498, and 566, respectively (A. B. Beklemishev et al.); soybean glycinin B4 polypeptide — 251 (E. S. Zakharova et al.); bovine Liver mitochondrial hepatoredoxin — 117 (V. L. Chashchin et al.); aspergillopepsin A — 320 (V. I. Ostromyslovskaya et al.); baker's Yeast phosphoribosylaminoimidazole-succinocarboxamide synthetase — 306 (N. A. Myasnikov et al.); FOOT-and-Mouth disease virus protein VP1 — 254 (A. M. Onishchenko et al.); mouse oncoprotein p53 — 390 (P. M. Chumakov et al.); human pituitary prolactin — 204 (N. P. Mertvetsov et al.); cellar spider venom insectotoxin — 35 (N. Zh. Sagdiev et al.); influenza virus protein PV2 — 759 (N. A. Petrov et al.); Shigella toxin A- and ß-subunits — 315 and 89, respectively (A. A. Baev et al.); human brain Na+, K+-ATPase a-subunit (form III) — 1,013 (O. I. Makarevich et al.); and cGMP phosphodiesterase ß-subunit — 852 amino acid residues (V. M. Lipkin et al.).
Each of the proteins whose primary structures are given above (with the exception of hemoglobin) is represented by a single polypeptide chain of varying length. The sequence of amino acid residues in the polypeptide chain of an individual protein is unique and specific. In some cases, protein molecules are built from two or more polypeptide chains connected to each other by covalent bonds. An example of this is the insulin molecule (see Fig. 29). However, the majority of protein molecules consist of several polypeptide chains held together by weak interaction forces.
When considering the primary structure of protein bodies, a fundamentally important question arises: do proteins realize all potentially possible combinations of their constituent amino acid residues, or are there certain combinations of amino acid residues characteristic of many, or perhaps even all, protein bodies? Indeed, through the rearrangement of amino acid residues in the polypeptide chain, PROTEINS AND PEPTIDES can yield an enormous number of isomers:
Number of amino acid residues in the molecule |
Possible number of isomers |
Number of amino acid residues in the molecule |
Possible number of isomers |
2 |
2 |
7 |
5040 |
3 |
6 |
8 |
40320 |
4 |
24 |
9 |
362780 |
5 |
120 |
10 |
3362780 |
6 |
720 |
20 |
∽2 ∙ 1018 |
It is clear that as the number of amino acid residues in a protein molecule increases to several hundred, the number of possible isomers can reach astronomical values. However, analyses of Amino acid sequences in proteins conducted in recent years have led to the discovery of a number of regularities indicating that far from all possible primary structures are realized in proteins. First of all, identical (matching) peptide groupings are found in different proteins, and frequently within the same protein. A special role in the structural similarity of proteins belongs to identical tripeptide groupings, although in some cases larger fragments also coincide in their amino acid sequence order. For instance, in 52 ribosomal proteins, identical tripeptide blocks repeat 657 times, tetrapeptide blocks 86 times, pentapeptide blocks 11 times, hexapeptide blocks 3 times, and heptapeptide blocks not a single time. At the same time, polypeptide chains contain analogous peptide groupings that differ from each other by interchangeable amino acid residues, i.e., residues that are structurally or biogenetically close, such as: gly-ser, gly-ala, leu-ile, leu-val, glu-asp, etc. It has been established that in a number of cases, the primary structures of various proteins include 50% or more identical peptide fragments.
Furthermore, significant sequence similarities are characteristic of proteins performing similar biological Functions. This has been demonstrated, for example, for a family of enzymes accelerating the dehydrogenation of diverse substrates, as well as for many Hydrolases, in which the amino acid sequence near the active center is extremely similar and standard. A particularly high degree of structural similarity is exhibited by proteins performing the same function in different species, which is clearly revealed when comparing the primary structures of insulin (Table 8), as well as Cytochromes, Histones, Pituitary Hormones, and other proteins. In cytochromes isolated from blue-green, red, and brown Algae, primary structures share 48–67% identity. Thus, the species Specificity of the primary structure of homologous proteins boils down to a limited number of Amino Acid Substitutions in the polypeptide chain at strictly defined positions.
A wealth of material of this kind has been provided by primary structure analyses of abnormal human Hemoglobins, where the substitution of just a single amino acid in the a-, ß-, y-, or d-chain results each time in a new type of abnormal hemoglobin. More than two hundred such hemoglobins are currently known. They are the source of human diseases categorized as molecular diseases. For example, when the sixth Glu residue in the ß-chain of human hemoglobin is replaced by a Val residue, abnormal hemoglobin S arises. This leads to the formation of sickle-shaped erythrocytes with a shortened lifespan, resulting in a severe hereditary disease — Sickle cell anemia.
Table 8. Differences in the primary structures of A and B chains in insulins of various origins
|
Insulin source |
Amino acid residue numbers |
|||||||||
|
Chain A |
Chain B |
|||||||||
4 |
8 |
9 |
10 |
1 |
2 |
3 |
27 |
29 |
30 |
|
Sperm whale, fin whale, |
glu |
thr |
ser |
ile |
phe |
val |
asn |
thr |
lys |
ala |
pig (taken as the standard |
||||||||||
of comparison) |
||||||||||
Human |
thr |
|||||||||
Bull, dog |
ala |
val |
||||||||
Goat, sheep |
ala |
gly |
val |
|||||||
Elephant |
ala |
gly |
val |
thr |
||||||
Sei whale |
ala |
thr |
||||||||
Horse |
gly |
|||||||||
Rabbit |
ala |
val |
ser |
|||||||
Rat |
asp |
ala |
val |
lys |
lys |
ser |
||||
met |
||||||||||
Mouse |
asp |
lys |
lys |
ser |
||||||
met |
||||||||||
Duck |
glu |
asn |
pro |
ala |
ala |
ser |
thr |
|||
Chicken |
his |
asn |
thr |
ala |
ala |
ser |
||||
Note. In all other positions of chains A and B, the sequence of amino acid residues in insulins of different origins is identical; cod insulin differs sharply in its primary structure from the species listed in the table.
Protein Secondary structure. A strictly linear polypeptide chain is characteristic of an extremely limited number of proteins. One such protein is Silk Fibroin, produced by silkworm caterpillars. Due to the specific conditions under which silk fibers are formed within the caterpillar's powerful muscular press, the filamentous fibroin molecules—which almost completely lack side-chain radicals flanking the main polypeptide chain—orient themselves along the silk-spinning duct and pack tightly within the silk fiber, acquiring a pseudocrystalline structure in certain regions. X-ray diffraction analysis first made it possible to reveal precisely this linear nature of polypeptide chains in silk fibroin; consequently, in the 1930s, the concept of a protein molecule as a fully extended polypeptide chain with an identity period of 0.71 nm emerged (similar to the one shown in Fig. 23).
However, it was later also found by X-ray diffraction that even in Fibrous proteins, let alone globular ones, fully stretched polypeptide chains are very rarely observed. X-ray images consistently indicated the presence of chains folded or twisted in some manner, with a repeating structural periodicity of 0.54 nm. Based on the analysis of X-ray diffraction patterns of stretched and normal fibrous proteins, such as Hair keratin, W. Astbury and F. Bell (1941) first proposed a model for polypeptide chain folding. This model successfully accounted for the experimentally observed 0.54 nm identity period and its variation upon stretching of the protein fiber. The model was constructed taking into account that the valence angles and interatomic distances characteristic of the peptide bond and its immediate environment must remain undisturbed.
The approach of modeling as a means of solving the protein molecular structure problem was continued by L. Pauling and R. Corey (1949–1951). They formulated the criteria for constructing a model that reflects the possible conformational state of a polypeptide chain in a protein molecule.
Based on these criteria, L. Pauling and R. Corey built the α- and β-type helices, while B. Low and R. Baybutte constructed the π-helix.
Since only the CHARACTERISTICS OF THE α-helical chain conformation matched the values observed in X-ray diffraction analysis of protein crystals, it was recognized by L. Pauling and R. Corey as one of the structural elements that genuinely exist in protein molecules (Fig. 30). Today, the Concept of the α-conformation of the polypeptide chain in proteins is universally accepted.
As can be seen from the figure, the α-Helix is characterized by extremely tight packing of the twisted polypeptide chain, such that the entire space within the hypothetical cylinder enclosing the twist is filled. Each turn of the right-handed α-helix contains 3.6 amino acid residues, whose radicals are directed outward and slightly backward (upward in the model), i.e., tilted toward the beginning of the polypeptide chain. The right-handed nature of the helix is easily determined by looking down the axis of the helix from the N-terminal amino acid (THE START OF the molecule): the model clearly shows that the polypeptide chain winds in a clockwise direction. The pitch of the helix (the distance between turns) is 0.54 nm, and the rise angle per turn is 26°. The identity period, i.e., the length of a complete turn segment along its course, is 2.7 nm (18 amino acid residues).
Hydrogen Bonds forming between the —CO— and —NH— groups of the polypeptide backbone located on adjacent turns of the helix play a crucial role in the formation and Maintenance of the α-helical configuration of the polypeptide chain (Fig. 30). Although the energy of these individual bonds is small, their large number results in a significant energetic effect, making the α-helical configuration quite stable and rigid. It is hypothesized that the π-electrons of the —CO— and —NH— groups in the polypeptide chain can interact through hydrogen bonds that bridge adjacent turns of the α-helix. As a result, electron conjugation zones arise in the helical Regions of the protein molecule. These zones can serve for The transfer of electronic excitation energy, which is of paramount importance for carrying out chemical reactions and transforming one form of energy into another. Thus, the Secondary structure of a protein molecule refers to a specific configuration characteristic of one or more polypeptide chains comprising the molecule.
One should not assume that the polypeptide chain is completely helical in every protein. Such cases are very rare. Apparently, each protein is characterized by a specific degree of polypeptide chain helicity:
Protein |
Fraction of Helical conformation, % |
Paramyosin |
100 |
Myoglobin |
75 |
Hemoglobin |
75 |
Porcine serum albumin |
50 |
Chicken egg albumin |
45 |
Hen egg-white lysozyme |
35 |
Tobacco mosaic virus (subunit) |
30 |
Pepsin |
28 |
17 |
|
Chymotrypsinogen |
11 |

Fig. 30. Models and diagram of the α-helix
In the α-helix model (left), the PARTS OF THE helix facing the observer are shaded in black; dashed lines indicate hydrogen bonds between the —CO— and —NH— groups located on adjacent turns of the helix; hydrogen atoms are shown as small circles. In the α-helix diagram (adjacent), all radicals are omitted, and the course of the polypeptide backbone is shown similarly to the model. To the right are images of the α-helix and the π-helix (far right), taking into account the planarity of peptide bonds.
Thus, in protein molecules, helical regions of the polypeptide chain regularly alternate with linear ones.
In native proteins, only right-handed α-helical Conformations of polypeptide chains exist, which is associated with the presence of L-series amino acids exclusively in protein structures (with rare exceptions).
Along with the α-helical structure, a β-structure is also observed in protein molecules. The latter refers to sheet-like structures formed by the combination of polypeptide chain regions in a β-conformation, i.e., in the form of linear peptide fragments. The linear conformation of these fragments is maintained through hydrogen bonds formed between parallel segments of the polypeptide chain brought close together at a distance of 0.272 nm (corresponding to the length of the Hydrogen bond between the —CO— and —NH— groups) (Fig. 31). Such structures are abundant in fibrous proteins, such as silk fibroin. However, β-structures are also systematically present in Globular proteins and often predominate over α-structures.
The Emergence of α- and β-structures in a protein molecule is a consequence of amino acids retaining their inherent ability to form hydrogen bonds even within polypeptide chains. Thus, the crucial property of amino acids to bind to one another via hydrogen bonds during the formation of crystalline preparations is realized as an α-helical conformation or a β-structure within the protein molecule. Consequently, the appearance of these structures can be viewed as a process of crystallization of polypeptide chain segments within the same protein molecule (Fig. 32).
The capacity to form hydrogen bonds—which serve as the driving force behind the emergence of α- and β-structures in a protein molecule—varies among different amino acids. Among them, a group of helix-forming amino acids is distinguished, which includes ala, glu, gln, leu, lys, met, and his. If residues of these listed Amino acids are concentrated in a certain part of the polypeptide chain or prevail in its composition, α-helix formation proceeds very smoothly. Conversely, amino acids such as val, ile, thr, tyr, and phe promote the formation of β-sheets in the polypeptide chain. Gly, ser, asp, asn, and pro are associated with the predominant formation of disordered fragments within its structure. Fig. 33 illustrates one of the Variants of the coexistence of α- and β-structures in the same protein.
Tertiary Structure of a protein. Information on the sequence of amino acid residues in a polypeptide chain (primary structure) and the presence of specialized, sheet-like, and disordered fragments in the protein molecule (secondary structure) does not yet provide a complete picture of the volume, shape, or, moreover, the mutual spatial arrangement of polypeptide chain segments relative to one another. These structural features are revealed by studying its tertiary structure, which is defined as the overall spatial arrangement of one or more polypeptide chains comprising the molecule and linked by covalent bonds. This structure was formerly referred to as the architectonics of protein molecules.
Solving this highly complex problem in Protein Chemistry has been linked to advancements in X-ray diffraction analysis, which increased the resolution of the method to 0.14 nm. Since interatomic distances in organic molecules are 0.1–0.2 nm, such high resolution makes it possible to precisely determine the spatial position of every atom in the polypeptide chain, i.e., to construct an exhaustive model of the protein molecule. In terms of resolution, X-ray diffraction is surpassed only by nuclear magnetic Resonance spectroscopy (0.03–0.09 nm), which is increasingly used to study protein tertiary structure. In addition, Electron Microscopy, circular dichroism, optical rotatory dispersion, and neutron crystallography are employed for this purpose, not to mention the computer-based mass Prediction of Protein tertiary structures from their primary structures.

Fig. 31. Protein β-structure:
A — formation of the β-structure; dashed lines indicate hydrogen bonds between the —CO— and —NH— groups of laterally positioned polypeptide chains; B — flat sheet formed by antiparallel polyglycine molecules; I — fragment of a space-filling model of an individual polyglycine molecule; II — sheet composed of space-filling models of polyglycine molecules; B' — β-sheet composed of two polypeptide chain fragments constructed taking into account the planarity of peptide bonds.

Fig. 32. Models of polypeptide chain folding into an a-helix with the formation of a hydrogen bond system between the turns of the helix:
A — shown as a paper strip with the polypeptide chain structure written on it; hydrogen bonds, which act as the interaction source leading to helix formation, are indicated by dashed lines; B — shown as the backbone of the polypeptide chain, composed of three-dimensional representations of NH, CO, and CHR groups; hydrogen and oxygen atoms involved in hydrogen bond formation are marked with "+" and "-" signs, respectively. The paper strip model can be used in high school chemistry courses and organic chemistry study circles for students
The first protein molecule model—myoglobin (Fig. 33), reflecting its tertiary structure—was created by J. Kendrew et al. (1957). Figure 34 shows schematic representations of the tertiary structures of myoglobin, ribonuclease, lysozyme, and chymotrypsinogen. Despite great difficulties, over the past three and a half decades it has been possible to determine the tertiary structure of nearly three hundred proteins, with more than three dozen of them solved in the USSR. In our country, specifically, the structures of pepsin, leghemoglobin, aspartate aminotransferase, glyceraldehyde-3-phosphate dehydrogenase, plastocyanin, phytohemagglutinin, y-crystallin, pyrophosphatase, a number of ribosomal proteins, actinoxin, lectin, leucine aminopeptidase, a- and ß-interferon, leucine-specific protein, rhodopsin, tubulin, Hydrogenase, methane monooxygenase, and others have been elucidated. Some of these will be characterized in more detail below.
It is believed that the tertiary structure of a protein molecule is determined by its primary structure, since the interaction of amino acid side chains with one another plays the decisive role in maintaining the spatial arrangement of the polypeptide chain characteristic of the tertiary structure. Possible types of bonds between side chains are shown in Fig. 35. Disulfide bridges are considered to be of particular importance in maintaining protein tertiary structure: in a number of proteins (see Figs. 29, 34, and 35), they firmly fix the relative positions of polypeptide chain segments (or chains). Thus, the location of cysteine (and other amino acid) residues in the protein molecule predetermines The Nature of inter-side-chain bonds and, consequently, the tertiary structure. Of course, here too, when disulfide bridges are formed, far from all theoretically possible formation pathways are actually realized, differing sharply from calculated values (according to T. Creighton, for five SS bonds in a protein molecule, the number of combinations reaches 945, for 10 it is 654,729,075, and for 25 it exceeds 5.8 ∙ 1030).

Fig. 33. a- and ß-structures in a fragment of the cytochrome b5 molecule and a-helical regions in the myoglobin molecule:
A — fragments of the cytochrome polypeptide chain (residues 21–25, 28–34, 50–54, and 73–79) forming the ß-sheet are shown by thick lines, and the sheet itself is marked with ß symbols; four a-structures (residues 34–39, 40–50, 54–63, 64–73) are depicted as cylinders; the prosthetic group of cytochrome b5, containing an iron atom, is located between the pairwise combined a-helices; B — eight a-helical regions in the myoglobin molecule (designated by Latin letters) significantly predominate over the bends of the polypeptide chain (designated by combinations like AB, BC, etc.) in length and intersect at various angles; the black disk in the upper left part of the molecule is the heme group with an iron atom at its center

Fig. 34. Tertiary structures of myoglobin (A), lysozyme (B), ribonuclease (C), and chymotrypsinogen (D)
In all cases, the configuration and spatial arrangement of the polypeptide chain backbone are shown. Amino acid residue side chains are not labeled anywhere except for three side chains in the chymotrypsinogen molecule. A — the hatched disk represents the heme; the dashed line indicates its attachment point to the polypeptide chain backbone via a Histidine side chain; a-helices, which make up 75% of the myoglobin polypeptide chain, are clearly visible; numbers indicate the positions of amino acid residues in the polypeptide chain (see also Fig. 33, B); B — disulfide bridges formed by The oxidation of cysteine side chains are designated by —S—S— symbols; it is noticeable that the a-helical configuration is characteristic of no more than 1/3 of the polypeptide chain; a cleft designed to accommodate the substrate cleaved by lysozyme is visible in the upper part of the molecule; C — a-helical configurations are virtually absent; other conventional designations are the same as in A and B; D — clearly defined a-helices are not observed; —S—S— bonds are indicated by dots, and the sequence numbers of the corresponding cysteine residues are also indicated by numbers; histidine side chains (pentagons) and the Serine side chain (a short two-pronged segment), which form the active center of chymotrypsin (after activation of chymotrypsinogen) responsible for the hydrolysis of peptide bonds, are highlighted by hatching in the center of the molecule

Fig. 35. Types of bonds between amino acid residue side chains in a protein molecule:
a — electrostatic interaction; b — hydrogen bonds; c — interaction of nonpolar side chains caused by the expulsion of lyophobic radicals into the "dry zone" by solvent molecules (the so-called "fat droplet" effect); d — Disulfide Bonds. The double curved line represents the polypeptide chain backbone
In addition to covalent bonds, the tertiary structure of a protein molecule is maintained by weak interaction forces (Fig. 35). Examination of the complete chemical structures of certain proteins has shown that their tertiary structures clearly reveal regions where hydrophobic amino acid side chains are concentrated, with the polypeptide chain essentially wrapping around a Hydrophobic core. Moreover, in a number of cases, two or even three hydrophobic cores become isolated within the protein molecule, resulting in a two- or three-core structure. This type of molecular architecture is characteristic of many proteins with catalytic functions (ribonuclease, lysozyme, etc.).
Data on the complete Chemical Structure of several protein molecules served as the starting point for developing The Theory of the domain architecture of protein molecules. A domain is defined as an isolated region of a protein molecule that possesses a certain degree of Structural and functional autonomy. In a number of enzymes, for example, coenzyme-binding domains are segregated. Associated with the theory of domain Organization is the gradually emerging concept in protein chemistry of uniformity, modularity, and standardization in protein tertiary structures, as well as the limited set of spatial foldings of polypeptide chains that actually exist in natural proteins.
Domains are now considered fundamental Structural elements of protein molecules, and the proportion and arrangement pattern of a-helices and ß-sheets are believed to provide more insight into the Evolution of protein molecules and phylogenetic relationships than the comparison of primary structures. The reason for this lies in the fact that Evolutionary Processes involved domain fusion, duplication, and the emergence of pseudosymmetrical domains from repeating subdomains; some of these events are associated with Gene Duplication and other genetic apparatus alterations.
There are many reasons to believe that the tertiary structure of a protein molecule forms entirely automatically. The driving force that folds the protein polypeptide chain into a strictly defined three-dimensional structure is the interaction of amino acid side chains with molecules of the surrounding solvent. In this process, lyophobic side chains are driven into the interior of the protein molecule, forming dry zones there ("fat droplet"), while lyophilic ones orient toward the solvent. At a certain point, an energetically favorable conformation of the molecule as a whole is reached, and the protein molecule becomes stabilized.
The self-organization of a polypeptide chain of a specific protein into its uniquely inherent Spatial Structure—that is, The process of tertiary structure formation—occurs in several stages (Fig. 36).
The conformation of the resulting globule is strongly influenced by such factors as medium pH, Ionic strength of the solution, and the interaction of protein molecules with other substances, which forms The basis of metabolic regulation, in particular the allosteric Regulation of enzyme Activity.
The Development of concepts regarding protein globule self-organization was accompanied not only by the introduction of the domain concept, as mentioned above, but also by a new approach to characterizing the Structural levels of protein bodies: to these were added the aforementioned domain level and suprasectoral/supersecondary structure. The latter refers to the Regularities of the emergence during polypeptide chain folding of elementary structures represented by ß-sheets (ß, ß'-structure), combinations of a-helical regions (a, a'-structure), or both simultaneously (Fig. 37). The Greek key and Greek fret topologies proved to be the most prevalent among Supersecondary structures.
It is significant that polypeptide chain fragments corresponding to domains and even subdomains are capable of independently maintaining a structure close to the native one. Thus, the cyanogen bromide fragment (amino acid residues 121–316) and its subdomain (residues 205–316) of thermolysin (a proteolytic enzyme from a thermophilic bacterium) spontaneously form a stably retained native structure. Meanwhile, the entire process of tertiary structure formation for proteins such as Carbonic anhydrase, a- and ß-lactoglobulin, phosphoglycerate kinase, and lactamase takes a mere 0.2 s. At the same time, factors limiting The rate of polypeptide folding during tertiary structure formation have been identified; these include the cis-trans isomerization of the X–Pro bond (where X is any amino acid), which is accelerated by peptidyl-prolyl cis-trans isomerase. The Mechanism of protein globule self-organization considered here has received elegant confirmation in studies (beginning in 1988) on the synthesis of artificial proteins. More than two dozen of them have been created to date. Based on the concept that a-helices, ß-sheets, and regions of unstructured polypeptide chain are associated with specific amino acid sequences in synthetically obtained polypeptides, researchers have managed to design protein molecules in which elements of secondary and supersecondary structure automatically and spontaneously arise, inevitably occupying pre-calculated positions during the formation of the Spatial structure of such an artificial protein. This is how synthetic proteins created via Protein Engineering have been named conditionally: "Felix" (from "four helices", i.e., consisting of 4 a-helices), "albebetin" (constructed from a-helices and ß-sheets in a 1:2 ratio, i.e., al:bebe), etc. Furthermore, they possessed pre-programmed properties and biological activity.

Fig. 36. Structural transformations of the polypeptide chain during protein globule formation
In stage I, local interactions occur at various sites throughout the polypeptide chain, resulting in the appearance of fluctuating a-helices and ß-sheets. The formation of a-helices is initiated from the polypeptide region where dicarboxylic amino acid residues are concentrated and terminates in the zone of diamino acid residues, whereas the central part of the emerging a-helices is occupied by amino acid residues with hydrophobic side chains. At this stage of protein molecule self-organization, the maximum possible number of a-helical conformations of the polypeptide arises, encompassing the major part of the polypeptide chain. Therefore, the illustrated hypothesis of protein molecule self-organization has been named the redundant helix hypothesis. In stage II, directed approximation of embryonic structures takes place, along with the "collapse" of a-helices and the formation of one or more hydrophobic cores due to contacts between the hydrophobic side chains of the a-helical amino acids. At this moment, a globular structure emerges. Stage III reduces to the compaction of embryonic structures and the transformation of the intermediate, highly helical globule into the native globule. In stage IV, the final tertiary structure of the molecule, characteristic of the given protein, is established

Fig. 37. Suprasequential (supra-secondary) protein structures
At the same time, over the past decade, views on the self-regulation of polypeptide chain folding during the formation of their tertiary structure in vivo have changed significantly. It turned out that the transformations mentioned above (see previous page and Fig. 36) do not occur spontaneously, but under The Influence of special specific proteins designed specifically for this purpose, called chaperones (this is what the English called an elderly lady who protected a young girl from ill-considered contacts when she first made her debut in society under her guidance). Being mostly tubular oligomers composed of 10–90 kDa subunits combined into stacked seven-membered rings, they draw the as-yet-unorganized polypeptide chain inside the oligomer and control the formation of its secondary and suprasequential structures as well as their mutual spatial packing (see Fig. 36), ensuring the emergence of the tertiary structure inherent in the native (functionally significant) globule of a given protein.
However, this concept is also hardly the final word in the complex and ambiguous problem of polypeptide chain folding. Based on experimental and theoretical approaches developed during The Study of the Protein Synthesis code (see below, Chapter VII), it has been suggested (L. B. Mekler and R. G. Idlis; G. I. Chipens) that a universal stereochemical Genetic Code exists, making it possible to understand how three-dimensional protein molecules were formed from linear polypeptide chains. The core of the matter lies in the existence of a code governing the interaction of amino acid residues with each other, which ensures the realization of the information already embedded in the primary structure of the protein (and, naturally, in mRNA and DNA) in the form of secondary and suprasequential structural elements and, ultimately, in the unique tertiary structure of the protein in question.
Quaternary Protein Structure. It has been noted previously that large protein molecules generally consist of subunits with a relatively low molecular weight. Such molecules are called epimolecules (supermolecules) or multimers, and their constituent elements are referred to as subunits or protomers.
The structure characterized by the presence of a specific number of polypeptide chains (subunits) in a protein epimolecule occupying strictly fixed spatial positions, as a result of which the protein exhibits a particular biological activity, is called The quaternary structure.
Quaternary structure should be distinguished from the oligomeric and aggregated states of a protein. A structure characterized by the presence of several polypeptide chains within a protein particle, the number of which varies in a certain proportion, is called oligomeric. It is extremely crucial that, despite the relative constancy of the number of peptide bonds in a protein oligomer and their ordered arrangement, the oligomer does not exhibit biological activity. For example, bovine serum albumin exists as a monomer (M = 68,000), dimer (M = 136,000), trimer (M = 204,000), and tetramer, with the monomers being arranged in an orderly fashion within the di-, tri-, and tetramers when they combine into oligomeric structures. However, this is not accompanied by the emergence of any new qualities compared to those possessed by the monomer of the given protein.
The aggregated state of a protein refers to a structure of protein particles represented by an indefinite and widely varying number of polypeptide chains. Here, too, the aggregation of monomers does not lead to the formation of any special properties in the protein residing in such a structural state. For instance, cytochrome c (and other cytochromes) possess a pronounced ability to aggregate molecules with one another, but this phenomenon is not accompanied by any change in the enzyme's properties.
The Quaternary Structure of several hundred proteins has now been elucidated. In 1965, information on quaternary structure was limited to approximately 20 proteins; in 1970, the first substantial review appeared, encompassing 108 proteins, and in 1976, D. Darnall and I. Klotz published a table containing a list of over 500 proteins along with data on the molecular weights of multimers, protomers, and the number of the latter in the epimolecule.
It turned out that the number of subunits in epimolecules varies over a very wide range: from 2 to 162. Most frequently, multimeric molecules contain 2 or 4 protomers, much less often 6, 8, 10, 12, or 24, and in rare cases, an odd number of them. Quaternary structure is characteristic primarily of proteins with a molecular weight above 50,000–60,000, whereas proteins with smaller molecular weights generally exist as monomers. The critical molecular weight limit of a protein molecule, above which the protein exhibits a quaternary structure in the majority of cases, is considered to be 100,000. As for the molecular weights of the subunits, they span A wide variety of values, from several thousand (e.g., 6,000 for insulin) to 330,000 (for each of the two subunits of thyroglobulin, a thyroid gland protein responsible for The biosynthesis of the hormone thyroxine — see Chapter XII).
A classic example of a protein with a quaternary structure is hemoglobin (Fig. 38). The hemoglobin molecule (M = 68,000) is built from four subunits with M = 17,000 each. The primary, secondary, and tertiary structures of the hemoglobin molecule subunits have been completely elucidated. They were found to be pairwise identical and were designated as β-axis type subunits (or rather, α and β subunits). The α-type subunit is represented by a polypeptide chain of 141 amino acid residues, and the β-type by 146. Their tertiary structures are similar. Four subunits (two α-type and two β-type) combine into a single hemoglobin molecule, occupying the corners of an almost regular tetrahedron (Fig. 38, B). Thus, a nearly spherical molecule is formed with dimensions of 0.50 × 0.55 × 0.64 nm.

Fig. 38. Quaternary structure of protein molecules:
A — models of hemoglobin subunits of type α (left) and β (right). The blocks making up the models characterize the electron density distribution in different parts of the molecule; the black (left) and white (right) lines indicate the backbone path of the polypeptide chain; B — three-dimensional model of the hemoglobin molecule; α-type subunits (light) and β-type subunits (dark) are located at the corners of an almost regular tetrahedron, dark disks represent heme groups; C — model of the tobacco mosaic virus molecule: protein subunits are visible on the outside, the dark helix is the nucleic acid; D — arrangement of subunits in the tobacco mosaic virus molecule (cross-section)
Of fundamental interest to future chemistry and biology teachers is the question of how the structure of hemoglobin is interrelated with its function—The ability to bind, transport, and readily release oxygen. This phenomenon is studied in detail in secondary school. The oxygen molecule itself attaches to Fe2+, which is fixed at the center of the heme molecule (Fig. 39); the heme, in turn, is held within the hydrophobic pocket of each subunit through coordination bonds with the imidazole radicals of histidine located in the distal and proximal parts of the polypeptide chain forming the α- or β-protomer of hemoglobin. The addition of oxygen to Fe2+ proceeds without changing the valence of the latter, utilizing one of its free coordination bonds; upon this, the radius of the Fe2+ atom decreases, and it, together with O2, moves into the plane of the porphyrin ring. It is held there until the hemoglobin molecule is transported to a tissue with a lower O2 content, where the reverse process of oxygen release takes place. Both the binding and release of O2 are accompanied by Conformational Changes in the structure of the α- and β-subunits of hemoglobin and their mutual arrangement within the multimer.

Fig. 39. Structure of the active center and mechanism of oxygen binding by a hemoglobin subunit (explanations in the text)
Fig. 38 also shows a diagram of the structure of a complex protein (nucleoprotein)—the tobacco mosaic virus. Its giant molecule (M = 40,000,000) contains a small amount (about 6%) of RNA, with the rest accounted for by protein. The protein moiety consists of a large number (2,130) of subunits, each with M = 17,500. The tobacco mosaic virus molecule is a hollow rod about 300 nm long and approximately 17 nm thick, with a central opening 4 nm in diameter. Each subunit has dimensions of 2 × 7 nm. The subunits are arranged in a helix, each turn of which is formed by approximately 16 subunits. The nucleic acid molecule follows the helical arrangement of the subunits, running between their rows. Over a hundred turns of protein subunits are arranged along the molecule.
The most striking phenomenon observed when studying the quaternary structure of protein molecules is that the assembly of protomers into a multimer molecule occurs spontaneously. It is hypothesized that each protomer molecule possesses specific regions that interact with corresponding regions in other protomers. When protomers join to form a multimer, ionic bonds arise. Metal Ions and sometimes low-molecular-weight Organic compounds take part in their formation. However, the greatest contribution to maintaining the integrity of multimer structures is made by weak interaction forces—specifically, hydrophobic interactions and hydrogen bonds; the combined effect of both is large enough to ensure the stabilization of the quaternary structure of proteins. Both outside the organism and presumably within cells, multimers are capable of reversibly dissociating into protomers.
It is fundamentally important that the slightest change in the tertiary structure of protomers makes it impossible for them to assemble into multimer molecules, which drastically affects the biological activity of the protein. Since the tertiary structure of a protein is determined by its primary structure and also depends on a number of other factors (environmental pH, salt concentration, etc.), even a minor change in the primary structure of the protein or the standard conditions within The Cell leads to an alteration in the functional activity of proteins. These phenomena form the basis of regulatory processes within the organism.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.