Molecular Biology of the Cell - Volume 1 - Alberts B., Bray D., Lewis J., Raff M., Roberts K., Watson J. 1994
Introduction to Cell Biology
Macromolecules: Structure, Shape, and Informational Functions
Protein Structure
Cells are largely composed of Proteins, which account for more than half of their dry weight (see Table 3-1). Proteins determine the Structure and shape of The Cell; moreover, they serve as instruments of Molecular recognition and catalysis. DNA, although it contains all the information necessary to build a cell, has little direct effect on cellular processes. For example, the Hemoglobin Gene itself does not carry oxygen: this is a property of the protein it encodes. Using computer terminology, DNA and mRNA can be thought of as "software"—instructions received by a cell from its parent cell. Proteins and catalytic RNA molecules constitute the "hardware"—the physical mechanisms that execute the program stored in memory.
DNA and RNA are chains built from NUCLEOTIDES that are chemically very similar to one another. In contrast, protein molecules are assembled from 20 very different Amino Acids, each possessing a distinct chemical identity. This diversity underlies the extraordinary versatility of The chemical properties of different proteins, and evolution has apparently selected proteins rather than RNA molecules as catalysts for most reactions in the cell.
3.3.1. The shape of a protein molecule is determined by its Amino Acid Sequence [21]
In a long polypeptide chain, free rotation of atoms around many bonds is possible, making the backbone of the protein molecule highly flexible. Therefore, any protein molecule can, in principle, adopt an almost infinite number of different shapes (Conformations). However, most polypeptide chains exist in only one of these conformations, which is determined by The amino acid sequence. This is because the amino acid side chains interact with each other and with Water to form weak noncovalent bonds (see Scheme 3-1). In this case, the appropriate side chains are positioned at key locations along the chain, forming strong bonds between them, which makes a specific conformation highly stable.
Class="center">
Figure 3-22. Schematic representation of a protein folding into a globule. Polar amino acid side chains tend to position themselves on the outer surface of the protein, where they can interact with water. Nonpolar amino acid side chains are located on the inside, where they form a hydrophobic "core" hidden from water.
The polypeptide chain of most proteins folds spontaneously to adopt its correct conformation. Upon Treatment with certain agents, a protein can be unfolded, or denatured; when the denaturing agent is removed, the protein usually spontaneously refolds into its original conformation. This indicates that all the information necessary to specify the shape of a protein is contained within the amino acid sequence itself.
One of the most important factors directing the folding of a polypeptide chain is the distribution of polar and nonpolar side chains. Numerous hydrophobic side chains tend to cluster inside the protein molecule, allowing them to avoid contact with the aqueous environment (just as droplets of oil mechanically dispersed in water coalesce). At the same time, all polar groups tend, conversely, to arrange themselves On the surface of the protein molecule, where they can interact with water and other polar groups (Figure 3-22). It is in this way that almost all polar groups that end up inside the protein globule are paired. Thus, Hydrogen Bonds play a major role in the interaction between different regions of a single polypeptide chain in a folded protein molecule; furthermore, they are of exceptional importance for many interactions occurring on The surface of protein molecules (Figure 3-23).
Secreted proteins, or cell-surface proteins, often form additional covalent bonds between different Regions of the same polypeptide chain. For example, The formation of Disulfide Bonds (also called —S—S-bridges) between two Cysteine SH groups that come into proximity in a folded polypeptide chain stabilizes the three-dimensional structure of extracellular proteins (Figure 3-24). These bonds are not required for proper protein folding, as it occurs normally in the presence of reducing agents that prevent the formation of —S—S-bridges. Indeed, —S—S-bridges are rarely (if ever) formed in protein molecules within the Cytosol, where There is a high concentration of agents that reduce SH groups and disrupt such bridges (see Section 8.6.11).

Figure 3-23. Hydrogen bonds (highlighted in color) that can form between amino acids in proteins. Peptide bonds are shown in gray.

Figure 3-24. Formation of a covalent disulfide bond between adjacent cysteine residues of a protein.
The net result of all individual amino acid interactions is that most protein molecules spontaneously adopt a characteristic conformation: usually a compact globular one, but occasionally an elongated fibrous one. The core of the globule is formed by tightly packed, almost crystalline, hydrophobic side chains, while the polar side chains form a complex and irregular outer surface. The Specificity of protein binding to small molecules and other macromolecular surfaces is determined by the arrangement and Chemical properties of the various atoms on this complex surface (see below). Chemically, proteins are the most complex molecules known.
3.3.2. The same folding patterns are repeatedly found in different proteins [22]
Although the amino acid sequence of a polypeptide chain contains all the information necessary for its folding, we still do not know how to read this information to predict the detailed three-dimensional structure of a protein from its sequence. Consequently, the native conformation of a protein can only be determined using the highly laborious method of X-ray crystallography of protein crystals. To date, more than 100 proteins have been fully analyzed using this method. The specific conformation of each is so complex that a detailed description of it would require an entire chapter.
Comparing the three-dimensional structures of different proteins has revealed that, although the conformation of each protein is unique, a few folding patterns are repeatedly found in parts of macromolecules. Two folding patterns are particularly common because they result from the regular formation of hydrogen bonds between the peptide groups themselves, rather than from unique interactions of the side chains. Both patterns were correctly predicted in 1951 using models based on X-Ray Diffraction studies of silk and Hair. Today, these Periodic structures are called the β-sheet and the α-Helix. The β-sheet conformation is found in the silk protein Fibroin, while the α-helix is found in α-keratin—a protein of Cytology/cytology/66.html">Skin and its derivatives (hair, Nails, and feathers).

Figure 3-25. The β-sheet is a common structure in regions of Globular proteins. Shown at the top is a 115-amino-acid domain of an immunoglobulin molecule. It consists of two β-sheets packed together like a sandwich, one of which is highlighted in color. Shown in more detail at the bottom is a perfect antiparallel β-sheet. Note that each peptide group forms hydrogen bonds with neighboring peptide groups. The β-sheets found in globular proteins are usually somewhat less regular than the structure shown here; often, β-sheets are slightly twisted (see Figure 3-27).
The β-sheet structure forms an essential part of the core of most (though not all) globular proteins.
Figure 3-25 shows, as an example, a portion of an antibody molecule; the antiparallel β-sheet of this molecule is formed by the polypeptide chain folding back on itself by 180° several times, so that the direction of each straight segment of the chain is opposite to that of its immediate neighbors. This structure is highly stable, owing to the formation of hydrogen bonds between the peptide groups of adjacent chain segments. Therefore, the antiparallel β-sheet often serves as a framework upon which a globular protein is assembled.
An α-helix is generated when a polypeptide chain twists around itself to form a rigid cylinder, in which every peptide group is hydrogen-bonded to neighboring peptide groups in the chain. Many globular proteins contain short regions of such α-helices (Figure 3-26); transmembrane regions of proteins that span The Lipid Bilayer are also almost always α-helices, due to the constraints imposed by the hydrophobic lipid environment (see Section 6.2.1). In an aqueous environment, an isolated α-helix is usually unstable. However, two identical α-helices with repeating patterns of nonpolar side chains can wrap around each other to form an extremely stable structure. Such long, rodlike structures are found in many Fibrous proteins, notably in the intracellular fibers of α-keratin that give skin its strength. Three-dimensional models of the α-helix and β-sheet of proteins, with and without side chains, are shown in Figure 3-27.

Figure 3-26. The α-helix is another common structure usually formed in specific regions of a protein's polypeptide chain. A. The oxygen-carrying Myoglobin molecule (153 amino acids long) is shown; one of the α-helical regions is highlighted in color. B. A detailed view of a perfect α-helix. As in the β-sheet, each peptide group is hydrogen-bonded to neighboring peptide groups. C. Atoms in an amino acid residue. Note that in B, the amino acid side chains are omitted for simplicity (they are located on the outer surface of the helix). In C, they are designated as R on the α-carbon atom of each amino acid (see also Figure 3-27).

Figure 3-27. Three-dimensional models of the α-helix and β-sheet. On the left, the structures are shown without amino acid side chains; on the right, with side chains. A. α-Helix (part of The myoglobin structure). B. A region of the β-sheet (part of the immunoglobulin domain structure). In the photographs on the left, each chain surface is represented by only one black atom (R groups in Fig. 3-25 and Fig. 3-26); the entire chain surface is shown on the right. (Courtesy of Richard J. Feldmann.)
3.3.3. Protein molecules are characterized by extraordinary diversity [23]
Differences in the Chemical Nature of amino acid side chains account for the remarkable diversity of possible protein conformations. As extreme Examples, let us consider Two Types of proteins secreted by Connective Tissue cells: Collagen and Elastin, which are Extracellular matrix proteins. In collagen, three separate polypeptide chains, rich in Proline and containing Glycine at every third position, are wound around one another to form a triple helix (see Section 14.2.6). These collagen molecules are, in turn, packed into fibrils where adjacent molecules are held together by covalent cross-links between neighboring Lysine residues. This results in the formation of fibers capable of withstanding exceptionally high loads (Fig. 3-28).
The other extreme is elastin, in which relatively loose and unstructured polypeptide chains are covalently cross-linked to form a rubberlike elastic network. This network enables Tissues such as Arteries and Lungs to deform and stretch without damage. As shown in Fig. 3-29, this elasticity is due to the ability of individual molecules to reversibly unfold under a stretching force.

Figure 3-28. A collagen molecule is a triple helix formed by three extended protein chains. Many rodlike collagen molecules cross-linked together form strong, inextensible collagen fibrils (top), which have a tensile strength comparable to that of steel.

Figure 3-29. Elastin consists of polypeptide chains that form stretchable fibers through cross-linking. Upon stretching, each elastin molecule unfolds into a more extended conformation. The striking contrast between the Physical Properties of elastin and collagen is due to the major differences in their Amino acid sequences.

Figure 3-30. Possible Sizes and Shapes of a protein molecule of 300 amino acids. The specific structure is determined by the amino acid sequence (from D. E. Metzler, Biochemistry, New York: Academic Press, 1977; adapted).
It is remarkable that the same chemical structure—a polypeptide chain—can adopt such highly diverse conformations. Examples include rubberlike elastin, collagen resembling a steel cable, and various globular proteins—Enzymes—which differ greatly in the shape of their catalytic surfaces. Figure 3-30 illustrates the highly diverse shapes that a polypeptide chain of 300 Amino acids can, in principle, adopt. The actual conformation, as we have already noted, depends entirely on the amino acid sequence.
3.3.4. Proteins have different LEVELS OF STRUCTURAL Organization [24]
In analyzing Protein structure, it is useful to distinguish several levels of structural organization. The amino acid sequence is called the Introduction/19.html">Primary Structure of the protein. Regular hydrogen bonding along the length of a continuous polypeptide chain leads to the formation of α-helices and β-sheets, which constitute the Secondary structure of the protein. Certain combinations of α-helices and β-sheets packed together form compactly folded globular units, each of which is called a protein domain. Domains usually consist of polypeptide chain segments containing between 50 and 350 amino acids; they appear to be the modular units from which proteins are constructed (see below). Small proteins may contain only a single domain, whereas larger proteins consist of multiple domains connected by relatively open regions of the polypeptide chain. Finally, individual Polypeptides can serve as subunits to form larger molecules, often called protein aggregates or Protein Complexes. In such complexes, the subunits are held together by A large number of weak noncovalent interactions (see Section 3.1.1); in extracellular proteins, these interactions are often stabilized by disulfide bonds.
The three-dimensional structure of a protein can be illustrated in several ways. Consider, for example, an unusually small protein—the pancreatic Trypsin inhibitor, which contains 58 amino acid residues packed into a single domain. This protein can be represented as a stereo pair showing all its non-hydrogen atoms (Fig. 3-31, A), or as a carefully rendered three-dimensional model where many details are omitted (Fig. 3-31, B). The protein can also be depicted more schematically, without side chains and atoms, to focus attention on the path of the polypeptide backbone (Fig. 3-31, C, D, and E). Such schematic representations are essential for revealing The structure of proteins that are typically larger than the trypsin inhibitor, as they allow one to trace the irregular path of the polypeptide chain within each domain (Fig. 3-32).
Figure 3-33 shows how the structure of a large protein can be resolved into different Levels of organization, each built hierarchically from the preceding ones. These levels of increasing structural organization may correspond to the stages of folding of a newly synthesized protein into its final native structure within the cell.

Figure 3-31. Three-dimensional conformation of a small protein, the pancreatic trypsin inhibitor, shown in five commonly used representations. A. A stereo pair showing the positions of all non-hydrogen atoms. The backbone is drawn with a thick line, and the side chains with thin lines. B. A space-filling model showing the Van der Waals radii of all atoms (see Panel 3-1). C. A skeletal wire model composed of segments connecting all α-carbon atoms along the polypeptide backbone. D. A ribbon model representing all regions of regular hydrogen bonding either as helices (α-helices) or as a set of arrows (β-sheets) pointing toward the carboxyl terminus of the chain; this model also shows hydrogen bonds. E. A "sausage" model showing the path of the polypeptide chain without any details. Note that the core of all globular proteins is tightly packed with atoms, and the impression of empty space is due solely to The Nature of models C, D, and E. (B and C courtesy of Richard J. Feldmann; A and D courtesy of Jane Richardson.)

Figure 3-32. Ribbon models of the three-dimensional structure of several Protein domains with different organizations. A. Cytochrome b562, a single-domain protein composed almost entirely of α-helices. B. The NAD-binding domain of Lactate dehydrogenase, consisting of a mixture of α-helices and β-sheets. C. The variable domain of an immunoglobulin light chain, forming a sandwich of two β-sheets. In these drawings, α-helices and connecting loops are colored, while the strands forming β-sheets are shown as gray arrows. Note that the polypeptide chain typically traverses the domain back and forth, making sharp turns only on the surface of the protein molecule. (Courtesy of Jane Richardson.)

Figure 3-33. The three-dimensional structure of a protein can be described in terms of different levels of folding, each built hierarchically from the structures of the preceding level. These levels are illustrated here using a two-domain bacterial catabolite activator protein. When the large domain binds cyclic AMP, a conformational change occurs in the protein, enabling the small domain to bind to a specific DNA sequence. The amino acid sequence is defined as the primary STRUCTURE OF THE protein, and the first level of folding of the polypeptide chain as its secondary structure. As indicated at the bottom of the figure under the brackets, the combination of the second and third levels of folding shown here is usually referred to as the tertiary structure, and the fourth level (subunit assembly) as The quaternary structure of the protein. (Adapted from drawings by Jane Richardson.)

Figure 3-34. Comparison of the amino acid sequences of two members of the Serine protease family. The carboxyl-terminal regions of the two proteins (from amino acid 149 to 245) are shown. Identical Amino acids are connected by colored lines, and the active-site serine residues at position 195 are highlighted. In the regions of The polypeptide chains enclosed in colored boxes, each amino acid of these two enzymes occupies the same position in the three-dimensional structure (see Fig. 3-35). B. Standard single-letter and three-letter Abbreviations for amino acids. (Adapted from J. Greer, Proc. Natl. Acad. Sci. USA 77: 3393-3397, 1980.)
3.3.5. Relatively few of the potential polypeptide chains would be useful
Since all 20 amino acids are chemically distinct and each can, in principle, occupy any position in a polypeptide chain, for a peptide of four amino acids There are 20 ∙ 20 ∙ 20 ∙ 20 = 160000 different possible chains, and for a polypeptide of n amino acids, 20n chains. Thus, more than 10390 different proteins could exist with an average typical length of about 300 amino acids.
We know, however, that only a very small fraction of all possible proteins will adopt a stable spatial conformation. All the rest would have many different conformations with different chemical properties and roughly equal energy. Proteins with such variable properties cannot be useful and, therefore, must be eliminated by natural Selection during evolution.
The remarkably precise fit of modern protein structures to their Functions is ensured by their ability to fold in a unique way. The amino acid sequence not only provides exceptional stability to one of the conformations but also determines the features of this conformation and its chemical properties required to perform a catalytic or structural function in the cell. Proteins are built so precisely that replacing even a few atoms of a single amino acid can disrupt the structure and lead to catastrophic changes in function.
3.3.6. New proteins often arise from minor changes in pre-existing ones [25]
Cells possess genetic mechanisms that ensure Gene Duplication, modification, and recombination during evolution (see Section 10.5.1). Consequently, once a protein with a useful surface property arises, its basic structure can then become part of many other proteins. In modern organisms, different proteins with related functions often share a similar amino acid sequence. Such Protein Families are thought to have arisen through the duplication of a single ancestral gene, followed by the evolutionary accumulation of Mutations that gradually gave rise to related proteins with new functions.

Fig. 3-35. Comparison of the three-dimensional structures of Elastase (A) and Chymotrypsin (B). In these evolutionarily related proteinases, only the amino acids located in the colored regions of the polypeptide chain are identical. Nevertheless, the protein conformations are very similar. The active sites of the enzymes are circled; both active sites contain an activated serine residue (see Fig. 3-47). The chymotrypsin molecule has several (more than two) chain ends because it is formed by the proteolytic Cleavage of chymotrypsinogen, an inactive precursor.
Consider the family of proteolytic (cleaving) enzymes, the Serine proteinases, which include the digestive enzymes chymotrypsin, trypsin, and elastase, as well as many of the clotting factors—proteinases that control Blood Coagulation. Comparing any two enzymes in this family reveals that approximately 40% of the positions in the polypeptide chain are occupied by the same amino acids (Fig. 3-34). An even more striking similarity is revealed when comparing their conformations determined by X-ray crystallography: most of the turns and bends of the polypeptide chains, which are several hundred amino acids long, turn out to be identical (Fig. 3-35).
Nevertheless, different serine proteinases have completely distinct functions. Some of the Amino Acid Substitutions responsible for the differences among enzymes in this group were apparently selected during evolution because they led to changes in substrate specificity and regulatory Properties of the proteins, which in turn generated the full diversity of modern functional properties. Other amino acid substitutions may have been 'neutral'—that is, preserved because they affected neither the structure nor the function of the protein. Since mutation is a random process, harmful substitutions must also have occurred, altering the Spatial Structure of the enzyme enough to inactivate it. These altered variants were lost during evolution because the individual organisms carrying them would have been at a disadvantage and eliminated by natural selection. It is therefore not surprising that cells contain a whole set of structurally related polypeptide chains that share common ancestors but perform different functions.
3.3.7. New proteins often arise from the combination of different polypeptide domains [26]
Once a set of stable protein surfaces has arisen in a cell, new surfaces with different binding specificities can be created by combining two or more individual proteins through noncovalent interactions. Cells characteristically associate globular proteins into larger functional protein aggregates: the Molecular Weight of many protein aggregates reaches 1 million or more, even though the molecular weight of a typical polypeptide chain is only 40,000–50,000 (approximately 300–400 amino acids); only a few polypeptides are three times larger than this average size.
A similar but distinct way of generating new proteins from existing polypeptide chains is the fusion of the corresponding DNA sequences to form a gene encoding a single large polypeptide chain (see Section 10.5.4). Proteins that arose in this way are thought to fold independently in different PARTS OF THE polypeptide chain into separate globular domains. Such a 'multidomain' structure is characteristic of many proteins, and, as would be expected from the evolutionary premises discussed above, functionally important binding sites are often located at the interface between different domains (Fig. 3-36). Fig. 3-37 shows the structure of a specific multidomain protein.

Fig. 3-36. The general principle by which the juxtaposition of two different protein surfaces during evolution leads to The Emergence of proteins containing new binding sites for other molecules. As shown in this figure, Ligand-binding sites are often located at the interface between two protein domains.

Fig. 3-37. Structure of the glycolytic enzyme glyceraldehyde-3-phosphate dehydrogenase. The protein consists of two domains (highlighted in different colors). Regions of α-helices are represented as cylinders, and β-sheets as arrows. The reaction catalyzed by this enzyme is shown in detail in Fig. 2-21. Note that the three substrate-binding sites are located at the interface between the two domains. (Courtesy of Alan J. Wonacott.)

Fig. 3-38. An example of the 'shuffling' of protein sequence blocks, which is widespread in Protein Evolution. Protein regions indicated by colored geometric shapes are evolutionarily related but not identical. A. The bacterial CAP protein consists of two domains; one of them (shaded triangle) binds to a specific DNA sequence, the other binds cAMP (see Fig. 3-33). The DNA-binding domain is related to the DNA-binding domains of many other regulatory gene proteins, including the lac repressor and cro repressor proteins. In addition, two copies of the cAMP-binding domain are found in eukaryotic Kinases regulated by cyclic nucleotide binding. B. Two domains of approximately 40 amino acids are shown, each of which occurs in three large vertebrate proteins. For example, the low-density lipoprotein (LDL) receptor is an 839-amino-acid transmembrane protein responsible for clearing Cholesterol from cells. It contains many domains also found in other proteins, including seven copies of a cysteine-rich domain (light circles) involved in LDL binding, and three copies of the same size (colored circles) whose functions are unknown.
Another way of reusing an amino acid sequence is especially common among long fibrous proteins, such as collagen (see Fig. 3-28). In this case, their structure is formed from multiple internal repeats of an ancestral amino acid sequence. Clearly, bringing amino acid sequences together by combining pre-existing DNA coding sequences is a more efficient strategy for the cell than generating new protein sequences through random DNA mutations.
3.3.8. Structural homologies can help determine the functions of newly discovered proteins [27]
The Development of rapid DNA Sequencing Methods has made it possible to determine the amino acid sequences of many proteins and The nucleotide sequences of their corresponding genes (see Section 4.6.6). A constantly updated 'protein database' is searched by computer to find possible sequence homologies between a newly sequenced protein and those studied previously. Currently, the sequences of only a small fraction of eukaryotic proteins have been determined, and it often turns out that a newly sequenced protein is homologous to an already known protein over some part of its length. It follows that most proteins have apparently evolved from a limited number of ancestral types. As expected, the sequences of many large proteins often show signs of having arisen by combining pre-existing domains in new combinations, a process known as 'domain shuffling' (Fig. 3-38).
Establishing domain homologies can also be useful in another aspect. Determining the three-dimensional structure of a protein is much more difficult than determining its amino acid sequence. However, the configuration of a domain in a newly sequenced protein can be 'guessed' if it is homologous to a domain of a protein whose conformation has previously been determined by X-ray crystallography. It is often possible to determine the structure of a new protein with reasonable accuracy by assuming that the turns and bends of the polypeptide chain in the two proteins will be the same, even if there are differences in the amino acid sequence.

Fig. 3-39. Schematic diagram of dimer formation from identical protein subunits. If the binding site recognizes itself, the dimers will be symmetrical.
These pairs often subsequently associate with other subunits to form tetramers and more complex assemblies.
Such protein comparisons are also important because similar structures often imply similar functions. Years of experimental research can be bypassed by establishing amino acid Sequence Homology with a protein of known function. For example, such sequence homologies first indicated that certain Yeast cell-cycle regulatory genes and some genes that cause cancerous transformation in mammalian cells encode protein kinases. In the same way, many of the proteins controlling morphogenesis in the fruit fly Drosophila were shown to be regulatory gene proteins, and one protein involved in morphogenesis was identified as a serine proteinase.
Each year, this database is updated with new protein sequence data, increasing the likelihood of finding useful homologies. Thus, comparing protein amino acid sequences will become an increasingly important tool in cell biology.
3.3.9. Protein subunits are capable of self-assembly into large cellular structures [28]
The principle that allows protein domains to associate and form new binding sites also operates in the assembly of much larger cellular structures. Supramolecular structures, such as enzyme complexes, Ribosomes, protein fibers, Viruses, and membranes, are not synthesized as single giant molecules held together by covalent bonds, but are assembled through the noncovalent aggregation of macromolecular subunits.
Using subunits to build large structures has several advantages: 1) less Genetic information is required to build a large structure from multiple copies of a smaller subunit; 2) because the subunits are held together by many relatively weak bonds, their assembly and disassembly are easily controlled; 3) assembling a structure from subunits minimizes errors, as a proofreading mechanism during assembly can eliminate defective subunits.
3.3.10. Identical protein subunits can interact to form geometrically regular structures [29]
If a protein has a binding site complementary to some region on its own surface, it will spontaneously aggregate. In the simplest case, a binding site recognizes itself, resulting in a symmetrical dimer. Many enzymes and other proteins form such dimers, which often, in turn, serve as subunits for the assembly of larger aggregates (Fig. 3-39 and Fig. 3-40).
If a protein's binding site is complementary to a different region on its surface, a chain of subunits is formed. With certain mutual orientations of the two binding sites, the chain will close upon itself and growth will stop. This results in a ring of two, three, four, or more subunits (Fig. 3-41). In a more general case, an infinitely long polymer of protein subunits will form. Provided that all subunits are bound to one another in an identical manner, the subunits in such a chain will arrange themselves in a helix (see Fig. 3-3). For example, an Actin filament is a helical structure assembled from identical subunits of the globular protein actin; actin filaments are major Components of the cytosol of most Eukaryotic cells. When mechanical strength is particularly important, supramolecular aggregates are usually constructed from fibrous rather than globular subunits, because fibrous subunits, winding around each other in a helix, have extensive areas for protein-protein interaction (Fig. 3-42, A).

Fig. 3-40. Ribbon model of a dimer formed from two identical protein subunits (monomers). The protein shown is the bacterial CAP protein, previously illustrated in Fig. 3-33 and Fig. 3-38, A. (Courtesy of Jane Richardson.)

Fig. 3-41. Identical subunits interacting with one another can form rings or helices. Helix formation was shown in Fig. 3-3; a ring forms instead of a helix if the subunits fit into each other in a way that halts further chain growth.
Hexagonally packed protein subunits can form flat sheets. Specialized membrane transport proteins sometimes aggregate this way in lipid bilayers (see Section 6.2.8). With a slight change in subunit geometry, the hexagonal sheet rolls up into a hollow tube (Fig. 3-42, B). Such cylindrical tubes form the protein coats of some elongated viruses (Fig. 3-43).
The formation of closed structures—such as rings, tubes, or spherical particles—further stabilizes the entire aggregate by increasing the total number of bonds between protein subunits. Moreover, because such a structure is formed through interdependent cooperative interactions, assembly and disassembly can be driven by relatively small Changes in the subunits themselves. This is most clearly illustrated by the protein coats of many simple viruses, which are shaped like hollow spheres. These shells are often assembled from hundreds of identical protein subunits that enclose and protect the viral nucleic acid (Fig. 3-43). The structure of the coat proteins must be highly flexible to accommodate Different types of intersubunit contacts and to allow subunit repacking when the nucleic acid is released at THE START OF the viral Replication cycle.

Fig. 3-42. Some structures formed by the self-assembly of protein subunits. A. Three common types of helical Protein Assemblies. An actin filament contains about two globular protein subunits per turn, while many other cytoskeletal proteins contain rodlike regions in which two α-helices associate to form a "coiled-coil" structure. In the Collagen helix, three elongated protein chains wind around one another over a long distance to form an extremely strong rodlike structure. B. Hexagonally packed globular protein subunits can form either flat sheets or tubes.

Fig. 3-43. Structure of a spherical virus. In many viruses, identical protein subunits pack to form a spherical shell enclosing the viral genome, which consists of RNA or DNA. For geometric reasons, no more than 60 subunits can pack in a perfectly symmetrical manner. However, if slight deviations from regularity are allowed, more subunits can be used to form a larger capsid. For example, tomato bushy stunt virus (TBSV) is a sphere about 33 nm in diameter. The electron micrograph (A) and diagram (B) show that it consists of more than 60 subunits. The proposed assembly pathway and the three-dimensional structure determined by X-ray crystallography are shown in C. The viral particle consists of 180 identical copies of the capsid protein (each containing 386 amino acids) and an RNA genome of 4500 nucleotides. To form such a large capsid, the protein must be able to pack in three slightly different ways (indicated by different colors). (Drawings by Steve Harrison; electron micrographs courtesy of John Finch.)

Fig. 3-44. Electron micrograph of tobacco mosaic virus (TMV). The virus consists of a single long RNA molecule surrounded by a tightly packed helix of identical protein subunits that form a cylindrical coat. When mixed in a test tube, purified RNA and coat protein spontaneously assemble into fully infectious Viral Particles. (Courtesy of Robley Williams.)
3.3.11. Self-assembling structures can consist of different protein subunits and Nucleic Acids [30]
Many proteinaceous cellular structures, such as viruses and ribosomes, are built from protein subunits and RNA or DNA molecules. The information for the assembly of these complex aggregates is contained within the structure of the macromolecular subunits themselves, and under appropriate conditions, isolated subunits can spontaneously assemble in a test tube into the final structure. The self-assembly of a large macromolecular aggregate from its individual components was first demonstrated for tobacco mosaic virus (TMV). This virus is a long rod in which a protein cylinder surrounds a helical RNA core (Fig. 3-44 and Fig. 3-45). If purified viral RNA and Protein subunits are mixed in solution, they aggregate to form fully active viral particles. The self-assembly process proved to be surprisingly complex, involving the formation of specific intermediate structures—double protein rings that add to the growing viral coat.
Another example of a macromolecular aggregate that can reassemble after dissociation into its individual components is the bacterial ribosome. Bacterial ribosomes consist of approximately 55 different protein molecules and three different RNA molecules (see Section 5.1.8). If all the individual components are incubated together in a test tube under appropriate conditions, they will spontaneously assemble into a ribosome. Most importantly, these reconstituted ribosomes are capable of Protein Synthesis. As expected, ribosome reconstitution occurs in an orderly fashion: first, specific proteins bind to the RNA, then other proteins recognize the resulting complex, and so on, until the complete structure is formed.
It is still not entirely clear how some of the more complex self-assembly processes are regulated. For example, many cellular structures have a precisely defined length that is many times greater than the length of any of their constituent macromolecules. How such precise length determination is achieved remains a puzzle. Fig. 3-46 illustrates three possible mechanisms for this limitation. In the simplest case, a long template of protein or another macromolecule acts as a ruler that determines the size of the final structure. This is the mechanism that determines the length of the TMV particle, where the RNA molecule serves as the template. Similarly, a protein template has been shown to determine the tail length of certain bacterial viruses (Fig. 3-47).

Fig. 3-45. Model of a structural element of tobacco mosaic virus. A single-stranded RNA molecule of 6000 nucleotides is packed into a protein coat consisting of 2130 copies of a specific protein (each of its molecules consists of 158 amino acid residues).

Figure 3-46. Three possible ways in which large protein assemblies can maintain a fixed length: A. Assembly along an elongated protein or other macromolecular scaffold that serves as a "ruler"; B. Adding additional subunits to the polymer structure beyond a certain length requires too much energy, and subunit assembly ceases. C. Vernier assembly. Two sets of rod-like molecules differ in length from the assembled complex, and its growth stops when the ends of these molecules align exactly.

Figure 3-47. Electron micrograph of bacteriophage lambda. The tip of the phage tail attaches to a specific protein on the surface of the bacterial cell, after which the DNA, tightly packed in the viral HEAD, is injected through the tail into the cell. The tail has a precise length, which is determined by the mechanism shown in Figure 3-46, A.
3.3.12. Not all cellular structures are formed by self-assembly [31]
Some cellular structures held together by noncovalent bonds are not capable of self-assembly. For example, Mitochondria, cilia, or myofibrils cannot spontaneously assemble in solution from their macromolecular components, because some of the information required for their assembly is contained in special enzymes and other cellular proteins that function as templates or scaffolds but do not form part of the final structure. Sometimes even small structures lack some of the components required for their assembly. For example, during the formation of certain bacterial viruses, the head, built from identical protein subunits, is assembled on a temporary scaffold made of another protein. This second protein is absent from the final viral particle, and therefore the head cannot spontaneously assemble in its absence. Other examples are known where proteolytic cleavage is an essential and irreversible step in the assembly process. This is how the capsids of some bacterial viruses and even some simple proteins, including the structural protein collagen and the hormone Insulin, are formed (Figure 3-48). Based on these relatively simple examples, one can conclude that the assembly of complex structures such as a mitochondrion or a cilium is controlled in both time and space by other cellular components and, furthermore, involves irreversible maturation steps catalyzed by Proteolytic Enzymes.

Figure 3-48. The polypeptide hormone insulin is synthesized as a precursor protein, proinsulin, which folds into the correct conformation and is then cleaved by a proteolytic enzyme. Consequently, after its disulfide bonds are reduced, insulin cannot spontaneously refold into its original conformation. The excision of a portion of the proinsulin polypeptide chain thus results in the loss of information required for the self-assembly of the molecule.
Summary
The amino acid sequence of a protein molecule determines its three-dimensional structure. The specific conformation of a polypeptide chain is stabilized by noncovalent interactions between its parts. Amino acids with hydrophobic side chains tend to cluster in the interior of the molecule, while the formation of local hydrogen bonds between neighboring peptide groups leads to the formation of α-helices and β-sheets. Many proteins are built modularly from small globular units called domains; small proteins typically consist of a single domain, whereas larger ones contain multiple domains linked together by short lengths of polypeptide chain. During the evolution of new proteins, domains are modified and combined with other domains.
The same forces that determine the three-dimensional structure of proteins are also responsible for the formation of protein assemblies. Proteins with a binding site complementary to their own surface can form dimers, closed rings, spherical particles, or helical polymers. A mixture of many different proteins, sometimes containing structural nucleic acids, can spontaneously assemble in a test tube into large, complex structures. However, not all cellular structures are capable of spontaneous reconstitution after dissociation into their individual components, as the assembly process in many cases involves irreversible steps.
Last update: 12/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.