BIOCHEMISTRY AND MOLECULAR BIOLOGY - W. ELLIOTT - 2002

CHAPTER 2. PROTEIN STRUCTURE

It is fair to say that no other substance possesses such remarkable properties as protein. Nothing in biochemistry can be understood without appreciating the all-percomplying, fundamental role of Proteins in life. Whenever a Cell needs to perform work, it is almost invariably carried out by a specific protein. Life depends on thousands of proteins, exquisitely engineered so that their molecules recognize and interact with other molecules with astonishing precision. Chemical Reactions within The Cell proceed through the binding of Enzymes to substrates, and the Rate of Enzymatic reactions is frequently determined by the selective binding of Allosteric regulators to specific sites on the enzyme. The function of such a highly complex Structure as Muscle relies entirely on Protein-Protein Interactions. Gene Expression is governed by proteins that interact with specific DNA sequences. Similarly, hormonal control is rooted in the selective Interaction of Hormones with protein receptors. Protein-mediated Molecular recognition drives transmembrane transport, including the generation of nerve impulses. Protein-protein interactions safeguard Immunity through The formation of antigen-antibody complexes. And such selective interactions are virtually endless.

This chapter focuses on the Chemical Structure of proteins—giant molecules that perform a vast array of biological Functions.

Introduction/19.html">Primary Structure of Proteins

Proteins are essentially polypeptide chains. Their molecules may consist of multiple chains, but for now we will limit our Discussion to single-chain proteins. A polypeptide chain is composed of A large number of interconnected Amino Acids. Nature utilizes 20 amino acids to build proteins; if we represent them as letters, a polypeptide chain can be envisioned as a word comprising hundreds of such letters. The longest known polypeptide chain contains roughly 5,000 amino acids, yet the majority of protein chains feature fewer than 2,000 amino acid residues. In principle, The amino acid "alphabet" permits the existence of a virtually infinite variety of different proteins, or "words."

This implies that evolutionary Selection is not constrained by a limited number of protein structures. This is likely why proteins were chosen as carriers of such a vast multitude of biological functions. Had the number of possible chains been restricted, evolution could never have selected optimal ones, as the probability of achieving the desired outcome through random chain assembly is exceedingly low.

The structure of an α-amino acid, which exists as a zwitterion in neutral aqueous solutions, is shown below.

Class="center">

Each amino acid, with the exception of Proline, contains a common backbone fragment H2N-CH-COOH, meaning they differ solely in the side chain R attached to the α-carbon atom. All amino acids found in proteins exhibit the L-configuration (with the exception of Glycine, whose molecule lacks an asymmetric atom).

Two amino acids can be joined together by splitting out a molecule of Water (a process achieved through a more complex pathway inside the cell). This yields a dipeptide, the structure of which is shown below (depicted in its un-ionized state for simplicity). The -CO-NH- bond is termed a peptide bond.

The dipeptide possesses terminal amino and carboxyl groups, allowing further amino acids to be successively attached and thereby extending the chain into a polypeptide.

Although there is no strict demarcation, short chains are conventionally called Peptides or oligopeptides (oligo- meaning few), whereas a polypeptide generally refers to a chain of eighty or more amino acids. Proteins may consist of multiple Polypeptides, in which case they are referred to as subunits. These subunits can be held together in a multimeric protein either exclusively by weak bonds or through stable covalent non-peptide bonds.

A brief note On terminology. The end of the chain bearing the terminal NH3+ group is designated as the N-terminus, while the opposite end is the C-terminus. Strictly speaking, the polypeptide backbone stripped of all amino acid side chains is termed the polypeptide backbone.

An amino acid incorporated into a protein is referred to as an amino acid residue, while the amino acid radicals R are termed amino acid side chains, protein side chains, or simply side chains.

The linear arrangement of amino acids within a polypeptide is known as its Amino Acid Sequence. The Procedure used to determine this sequence is called sequencing. Sequencing plays a pivotal role in Protein Chemistry. The amino acid sequence constitutes the primary structure of a protein. Sanger in Cambridge was the first to determine the amino acid sequence of a protein (Insulin), for which he was awarded the Nobel Prize in 1958. Subsequently, Edman developed a purely chemical approach to automate this procedure. Modern automated sequencers require a mere 0.001 µg of protein for the task. In addition, specialized instruments known as amino acid analyzers are employed to determine the Amino Acid Composition of a protein, after the latter has been subjected to complete acid Hydrolysis.

WHAT IS A Native Protein?

Depicting the primary structure of a protein on paper creates the impression that its molecule adopts an extended, thread-like conformation. This is a misconception, as demonstrated by egg albumin—a colorless, water-soluble protein derived from chicken eggs. It can be purified and crystallized into sharp crystals. The ability to crystallize stems from a compact, three-dimensional molecular architecture. Crystallization requires all molecules to possess a uniform shape. In their natural state, polypeptide chains in Globular proteins are folded into a compact spherical structure known as a globule; unlike Fibrous proteins, whose long chains stretch out along a single axis and exhibit a low axial ratio. Although there is no rigid definition of globular proteins, it is evident that their three-dimensional architecture—the tight packing of the polypeptide chain—is typically disrupted by heating. A protein folded into its original, natural conformation is termed native, one with an unfolded, random conformation is called denatured, and The conversion of a native protein into its denatured state is referred to as Denaturation.

When an egg is boiled, the transparent egg white first turns cloudy and subsequently transforms into an insoluble white mass. What is happening here? Upon heating, individual, specifically folded protein chains unfurl and entangle with one another in a random and irreversible manner (Fig. 2.1). This phenomenon indicates that the chain folding in native proteins is largely maintained by weak non-covalent bonds, since covalent bonds would remain intact under moderate heating. Meanwhile, the polypeptide chain itself remains structurally intact under these conditions. Thermal denaturation is a universal property of proteins.

Fig. 2.1. Denaturation of egg albumin (hypothetical scheme)

The biological activity of proteins, such as enzymes, is almost invariably abolished by heating (with only a few known thermostable proteins). For this reason, heat is lethal to the cell.

Now that we have a grasp of the primary structure of proteins, we can proceed to discuss the three-dimensional folding of polypeptide chains.

What are the main Factors Determining the three-dimensional structure of a protein?

As already noted, most proteins fold into compact globules. Such globules are stable in aqueous systems because their polar groups reside On the surface in contact with water, while nonpolar groups are tucked away inside the molecule, minimizing their exposure to water. If any groups buried within the globule are capable of forming ionic and Hydrogen Bonds but effectively lack partners, this destabilizes the packing. Consequently, an ionized group located in a hydrophobic environment will destabilize the protein molecule.

At first glance, The Challenge of forming a functionally active globular protein might seem insurmountable. After all, hundreds of amino acid residues must be arranged into a tightly packed globular structure with polar groups on the exterior and hydrophobic ones on the interior. Furthermore, in the case of enzymes, the surface must feature clefts that perfectly Complement the shapes of substrate molecules, as well as allosteric regulators of protein activity (see Chapter 12).

You now understand why evolution proceeds so slowly on a geological timescale—its primary task, if you think about it, is precisely the design of new proteins. The chance of randomly stumbling upon an amino acid sequence that not only satisfies the conditions for stable chain folding but also fulfills specific functional requirements is infinitesimally small. Therefore, it is hardly surprising to encounter proteins with different functions that are nevertheless so structurally similar as to suggest they share a common ancestor or evolved from one another. It appears that when faced with a particular challenge, evolution prefers not to engineer proteins de novo, but rather to co-opt well-established structures and adapt them for novel purposes.

Before delving into protein folding, we need to learn more about the STRUCTURE OF THE 20 amino acids that serve as their building blocks.

Structure of the 20 Proteinogenic Amino Acids

Studying the structures of the 20 different amino acids becomes much more meaningful when keeping in mind the purpose of their side chains. It is precisely their differences in shape, size, and polarity that allow amino acids to serve as the building blocks evolution uses to meet the diverse and stringent demands of protein architecture. It is easy to picture how this works: a small hydrophobic group is needed here to fill a cavity near a neighboring group, a strong polar group is required there, and a weaker one elsewhere. With 20 different players in the game, the master coach—evolution—has the flexibility to construct complex designs.

Although numerous non-protein Amino acids have been discovered in various organisms, all known life forms build their proteins using the exact same 20 amino acids. F. Crick referred to them as the "magic twenty." Only these are encoded by METABOLISM/28.html">The Genetic Code. (If this doesn't mean much to you yet, bear with us until the following chapters.) Let us begin with aliphatic hydrophobic chains, listed in order of increasing size.

In glycine, the side chain is simply a hydrogen atom. It is the smallest amino acid; its residue within a protein exhibits neither distinctly hydrophobic nor hydrophilic properties. Next come the aliphatic nonpolar side chains in order of increasing Hydrophobicity, with the final three chains being branched.

Methionine stands somewhat apart among hydrophobic aliphatic amino acids. This is not because its side chain contains a sulfur atom, but because its terminal methyl group plays a crucial role in metabolism.

Next, we turn to the truly bulky side chains of aromatic amino acids:

It is worth noting that the hydroxyl group in Tyrosine can participate in hydrogen bonding, which is why tyrosine cannot be unambiguously classified as a purely hydrophobic amino acid.

Hydrophilic Amino Acids

All ionized groups are hydrophilic; therefore, hydrophilic amino acids definitively include those whose side chains contain carboxyl or amino groups. Both of these groups are ionized at physiological pH. Aspartic and glutamic acids are acidic amino acids, Lysine and Arginine are strongly basic, and Histidine is weakly basic. The ring structure within the histidine molecule is called the imidazole ring.

Both acidic amino acids—aspartic and glutamic acids—are also represented in proteins by their respective amides, asparagine and glutamine.

Hydrophilic amino acids also include the hydroxyl-containing residues Serine and Threonine.

Specialized Amino Acids

Cysteine resembles serine, but contains a thiol group (-SH) instead of a hydroxyl group (-OH). Its specific role in proteins is twofold: it allows thiol groups to be introduced into the active sites of proteins, and two cysteine residues can be linked together by a covalent -S-S- bond.

Proline is remarkable in that its residue introduces a kink into the polypeptide chain. Unlike Other Amino Acids, free proline contains an imino group rather than an amino group.

Ionization of Amino Acids

As already mentioned, free amino acids In aqueous solutions exist as zwitterions, in which the α-amino and α-carboxylate groups are ionized, since the former have a pKa of 8–10, and the latter a pKa of 2.

In the peptide chain, all these dissociating groups are blocked (except for the terminal ones) because they participate in the formation of peptide bonds. Therefore, the ionic status of proteins is almost entirely determined by the dissociation of groups in the side chains of aspartic and glutamic acids, lysine, arginine, and histidine.

Charged amino acid residues within a peptide chain are shown below:

The carboxyl groups in the side chains of aspartic and glutamic acids have a pKa of ~4 and are almost completely dissociated at pH 7.4; consequently, their side groups carry a negative charge. The amino groups in the side chains of basic amino acids, such as lysine (pKa 10.5) and arginine (pKa 12.5), are also fully ionized (NH3+) at physiological pH values. The third basic amino acid, histidine, which features an imidazole ring as its side chain, has a pKa of 6.04—quite close to physiological pH levels. Due to this, the imidazole moiety is frequently part of the Active Site of enzymes that catalyze reactions involving hydrogen ion transfer. During these reactions, the imidazole residue can act either as a proton donor or acceptor, depending on its ionic state.

The distribution of charged amino acid residues in a polypeptide chain strongly influences its conformation, since, as is well known, like charges in close proximity repel each other, whereas opposite charges attract. To regulate enzymatic activity within the cell, phosphorylation is frequently employed, whereby a strong negative charge is introduced into the protein molecule, triggering a conformational change.

Levels of Protein structural Organization: from primary to quaternary structure

As you already know, The sequence of amino acids covalently linked into a polypeptide chain is referred to as the primary structure of a protein (Fig. 2.2, a). In itself, the primary structure provides no information about how the polypeptide chain is folded in three-dimensional space.

The polypeptide chain adopts a specific conformation known as the Secondary structure of a protein (see Fig. 2.2, b), which, in turn, folds into a compact entity called the tertiary structure (see Fig. 2.2, c). Such an entity, formed by the primary, secondary, and tertiary structures, may either function as an independent protein or associate as a monomer (subunit) with identical or different monomers to form a complex multimeric protein. The quaternary structure refers to the spatial arrangement of interacting subunits formed by individual polypeptide chains of a protein (see Fig. 2.2, d).

Fig. 2.2. Schematic representation of the primary (a), secondary (b), tertiary (c), and quaternary (d) structures of proteins

Having adopted this terminology, we will now examine the various levels of protein structural organization in greater detail.

Secondary structure of proteins

Here we will focus on the organization not so much of the side chains as of the polypeptide backbone itself. We have already mentioned that in water-soluble globular proteins, polar side groups should ideally be located on the outside of the globule, while hydrophobic groups should reside inside. But what about the polypeptide backbone? No matter how you coil it into a compact shape, a portion of the backbone will inevitably end up embedded in the central, i.e., hydrophobic region. The issue here is that this backbone—the peptide chain proper—contains numerous C=O and N-H groups of peptide bonds, each of which is potentially capable of participating in Hydrogen bond formation. Until this potential bonding occurs (with the release of Free energy), the structure will remain unstable. What can serve as partners for these numerous groups? The simple answer is: analogous groups from the same or a neighboring polypeptide chain.

Let us emphasize once again that we are not yet discussing the packing of amino acid side groups, which becomes significant primarily at the tertiary structure level. For now, we limit our discussion to The problem of how to satisfy the hydrogen-bonding requirements of the C=O and N-H groups in the polypeptide chain. There are two primary types of structures that make this possible: the α-Helix, into which the chain coils like a telephone cord, and the pleated β-sheet, in which extended segments of one or more chains lie side by side. Both of these structures are remarkably stable; they are found

at the periphery of protein globules (where hydrophilic amino acid side groups are located) and in their interior region (where hydrophobic side groups reside).

α-Helix

The term α-helix was introduced by L. Pauling, who discovered that the peptide chain in the protein α-keratin is folded into a right-handed helix. For L-amino acids, a right-handed helix is more stable than a left-handed one. You can visualize a right-handed helix by imagining the trajectory of a point on the edge of the HEAD of a driving screw; the shape of such a helix is shown in Fig. 2.3, a.

Fig. 2.3. Schematic representation of the polypeptide chain folded into an α-helix

a - Approximate arrangement of hydrogen bonds (dashed lines) between C=O and N-H groups (side chains not shown); b - end-on view of the α-helix. The projections of the R side groups are oriented randomly, while the groups themselves are spaced evenly along the chain (3.6 residues per turn, or 100° on a 360° projection: 360°/3.6 = 100°)

It is natural to assume that such a helix contains an integer number of amino acid residues per turn. However, this appealing idea had to be abandoned when Pauling demonstrated that an α-helix actually contains 3.6 amino acid residues per turn. This means that the C=O group of one peptide bond forms a hydrogen bond with the N-H group of another peptide bond located four residues away along the chain. Both the C=O and N-H bonds are directed parallel to the axis of the helix and oppose each other in pairs; this arrangement is optimal for hydrogen bond formation and, consequently, for the stabilization of the α-helix. In cross-section, the α-helix resembles a disk with the amino acid side chains projecting outward (Fig. 2.3, b). The Van der Waals radii of the atoms are such that there is no empty space inside the helix, which contributes to the Stability of the α-helix. Not the entire polypeptide chain in a globular protein is helical. On average, individual helical segments comprise about 10 amino acid residues, but the length of the helices can vary significantly among different proteins.

Amino acids vary in their frequency of occurrence within helical regions. Proline plays a special role here, as its residue acts as a kind of terminator for α-helices. In all other cases, the polypeptide chain can rotate freely around two single bonds per amino acid residue. Note that free rotation about the CO-NH peptide bond itself does not occur because, due to electronic delocalization, its properties are close to those of a double bond, representing a hybrid of two Resonance structures: a) in which the bond between the Carbon and Oxygen atoms is a double bond; b) in which this bond is a single bond.

The peptide group involving the N atom of proline lacks the hydrogen atom necessary for hydrogen bond formation, and more importantly, rotational isomerism is restricted at this site, preventing the peptide chain from adopting a conformation compatible with an α-helix.

Before discussing The Role of α-helices in Protein Structure, we should examine an alternative type of polypeptide chain arrangement —

the β-pleated sheet. Typically, a protein structure represents a combination of both folding types, which occupy different Regions of the polypeptide chain.

The β-Pleated Sheet

This is also a stable structure in which the polar groups of the peptide bond are paired via hydrogen bonds, providing stability even within the Hydrophobic core of a protein globule.

The organizational principle here is remarkably simple. The polypeptide chain is in an extended state (or β-form, named after the protein β-keratin, in which it was first discovered), and its C=O and N-H groups are hydrogen-bonded to identical groups of an adjacent, parallel-oriented polypeptide chain (Fig. 2.4). Both chains may be independent or represent fragments of a single, shared chain.

Fig. 2.4. Structure of the β-pleated sheet

a - Hydrogen bonds between unidirectionally oriented polypeptide chains in a parallel β-pleated sheet. Side chains R attached to -CH- residues are positioned above and below the plane of the page; b - hydrogen bonds between oppositely directed polypeptide chains in an antiparallel β-pleated sheet. Such a sheet can form within a single folded polypeptide chain

When such a structure is formed by several parallel chains, the term sheet is entirely justified. It is called

pleated because the α-carbon atoms of the amino acid residues are located alternately on either side of the central plane of the sheet. The polypeptide chains forming the pleated sheet can run in the same or opposite directions; in the former case, the pleated sheet is termed parallel, and in the latter, antiparallel. The antiparallel β-structure typically arises when the peptide chain reverses direction, forming a so-called hairpin. The turning point is referred to as a β-turn.

Random Coil, or Polypeptide Chain Loops

The random coil is, in fact, neither random nor a coil. This term was given to regions of the polypeptide chain whose conformation can be classified neither as an α-helix nor as a β-pleated sheet. A more accurate term is connecting loops. Their structure is largely determined by interactions among the side chains of their constituent amino acid residues; in any specific protein molecule, this structure is fixed rather than random, contrary to what the name might suggest. Furthermore, the word "coil" is not entirely appropriate here, since we are not dealing with an arbitrary conformation. In connecting loops, not all peptide C=O and N-H groups participate in hydrogen bonding; consequently, these regions of the polypeptide chain are typically located on The surface of the protein globule in areas of contact with water.

Tertiary Structure of Proteins

Thus, evolution utilizes 20 amino acids to construct proteins. The sequence of amino acid units in a polypeptide chain is characteristic and constant for each protein, defining its primary structure. The polypeptide chain can fold into an α-helix or a β-pleated sheet, or exist in less ordered loops, with peptide bonds participating in hydrogen bonding exclusively within α-helices and β-structures. All three of these folding types are distributed along the polypeptide chain, occupying distinct regions and constituting the secondary structure of the protein. The further folding of secondary structures into the compact architecture of a globular protein is termed tertiary structure.

To simplify the depiction of protein molecules, conventional symbols for secondary structures are used. For instance, an α-helix is represented as a cylinder (sometimes

with a helical line drawn inside) or as a helical ribbon (Fig. 2.5, a), while regions of the chain incorporated into β-pleated sheets are depicted as arrows indicating the direction of the peptide chains (Fig. 2.5, b). Sometimes they are shown slightly twisted to reflect the actual weak right-handed twist of the chain in a β-structure. In both cases, the symbol refers to a specific, fixed region of the polypeptide chain. Chain segments that separate helices and pleated sheets fall into the category of loops and are represented by a simple line.

Fig. 2.5. Symbols used to represent α-helix regions (a) and β-pleated sheets (b). The right-hand curvature characteristic of antiparallel sheets is shown on the right. Intermediate loops and disordered chain regions are depicted by a simple line

How Are Proteins Constructed from α-Helices, β-Pleated Sheets, and Connecting Loops?

In principle, any combination of helices, pleated sheets, and loops can be used to build proteins, provided there are no Steric hindrances related to the packing of amino acid side chains caused by their attraction or repulsion. Additionally, protein folding can be influenced by the arrangement of polar side groups of amino acid residues on the surface of the connecting loop. Overall, certain requirements must be met: hydrogen bond formation is maximized, while contacts between hydrophobic groups and polar groups or water are minimized. However, only a few Amino acid sequences out of the countless possibilities can satisfy these requirements.

The three-dimensional architecture of a substantial number of proteins has been determined using X-Ray Diffraction data. This research has revealed that certain Structural motifs are strongly preferred, meaning the number of basic protein designs may be limited.

Three-dimensional protein structures are rarely of interest to biochemists outside this field, yet familiarity with them is often essential, for instance, when discussing the Biochemical Mechanisms of specific processes. Some of these structures are illustrated in Fig. 2.6.

Fig. 2.6. Structures of various proteins: a - Myoglobin, with the heme group highlighted in red at the center; b - staphylococcal nuclease; c - Triosephosphate isomerase (the arrangement of α-helices and β-pleated sheets in the chain is shown below); d - Pyruvate kinase

Myoglobin (see Fig. 2.6, a) consists exclusively of α-helices connected by loops, with a heme group embedded in a deep cleft. The molecule of staphylococcal nuclease—an enzyme that hydrolyzes Nucleic Acids (see Fig. 2.6, b)—exhibits a combination of antiparallel β-pleated sheets (1–4) and three α-helices. Another structural motif is the so-called α/β-barrel, whose central core is formed by β-strands (1–4) arranged like staves in a wooden barrel (except that they are tilted), while the periphery is filled with α-helices. This motif is exemplified by the structure of triosephosphate isomerase (see Fig. 2.6, c). Yet another packing motif, formed by alternating helices and β-structure regions, can be observed in the structure of pyruvate kinase (see Fig. 2.6, d).

What forces maintain the tertiary structure?

We have already discussed that the elements of secondary structure—α-helices and β-pleated sheets—are stabilized by hydrogen bonds. Now let us discuss what holds these elements together into a single tertiary structure. Of course, ionic and hydrogen bonds between the side chains of amino acid residues make a certain contribution, but the primary role is played by hydrophobic forces. These forces drive the localization of hydrophobic chains to the interior of the molecule, thereby reducing their contact area with water and maximizing van der Waals interactions. In most globular proteins, weak interactions alone are sufficient to maintain the tertiary structure.

What is the role of covalent S-S bonds in tertiary structure?

The special functions of cysteine residues were mentioned earlier. One of them is that the thiol groups of these residues participate in the formation of active sites in enzymes, such as the glycolytic enzyme glyceraldehyde-3-phosphate dehydrogenase (see p. 112). However, cysteine has another function as well. The tertiary structure of proteins is maintained primarily by weak interactions between side chains. While this is sufficient to protect proteins from damaging influences within the cell, extracellular proteins—such as Blood insulin, digestive enzymes, and the like—require additional protection against a harsher extracellular environment. This is achieved by pairing spatially close cysteine residues to form covalent Disulfide Bonds, which act like steel staples holding log structures together. Even a small number of such bonds is enough to stabilize the protein structure very effectively; for example, insulin contains only three. Disulfide bonds, or S-S bridges as they are also called, are not disrupted by moderate heating, and proteins containing them often exhibit greater thermal stability.

A simple experiment helps to illustrate how disulfide bonds are formed. If a neutral solution of cysteine is left exposed to the air, the surface of the liquid will become covered with a layer of white, insoluble crystals within a few hours. The chemical reaction taking place can be represented by the following equation:

Disulfide bonds between cysteine residues in proteins are formed in the exact same way.

The stabilizing role of S-S bridges is most vividly demonstrated by keratin, the primary protein of Hair. Bundles of keratin polypeptides are cross-linked by numerous S-S bridges, which impart rigidity to the hair. During a so-called chemical permanent wave, the S-S bonds are cleaved by a reducing agent, the hair is reshaped into a new configuration, and then the S-S bonds are re-formed using a suitable oxidizing agent. The number of S-S bonds may remain the same after these Procedures, but they now form between those cysteine residues that happen to be brought into close proximity by the mechanical alteration of the hair's configuration.

Quaternary Protein Structure

Tertiary structure completes the description of a single protein molecule's architecture. However, there are quite a few proteins whose molecules are complexes formed from multiple protein chains held together by noncovalent bonds. Such complexes are referred to as oligomeric, multimeric, or subunit proteins. Their composition and stoichiometry are constant, which proves that the assembling protein subunits "recognize" one another thanks to complementary patches on their surfaces. Examples of such Oligomeric Proteins include Hemoglobin (see Chapter 27) and allosterically regulated enzymes (see p. 158). The arrangement of subunits within a functionally active protein complex is termed The quaternary structure of the protein (see Fig. 2.2, d).

Membrane Proteins

So far, we have focused on globular proteins, in which the globular structure ensures the water solubility of the molecule as a whole. Many proteins that span Biological Membranes, however, are organized differently: they consist of two surface-exposed "water-soluble" domains connected by a helical segment embedded within the membrane (see p. 61). The drawback of this design is obvious: contacts between the polar groups of the protein and the hydrophobic core of the membrane render the molecule unstable.

Protein conjugates

A vast number of proteins are no more complex than those described above. For normal functioning, they require only a properly folded polypeptide chain and, at most, Cofactors such as Metal Ions or noncovalently bound Coenzymes. However, certain additional components can be covalently attached to proteins. These are called prosthetic groups, the resulting active complex is termed the holoenzyme, and its protein moiety is the apoenzyme. Examples of such structures include Cytochromes, which contain a heme group as a prosthetic group, or dehydrogenases containing flavin adenine dinucleotide (FAD).

Proteins covalently bound to carbohydrate components are classified as Glycoproteins. These include many Plasma Membrane proteins, which have Oligosaccharides attached to the outer surface of their molecules (see p. 61). These carbohydrate residues replace a hydrogen atom in the hydroxyl group of serine or threonine (O-Glycosides), as well as the amide proton in the side chain of asparagine (N-glycosides). The structure of such N-glycosides is shown schematically below:

Secretory proteins and many blood proteins are glycoproteins. The physiological purpose of carbohydrate components in glycoproteins is not always clear. Sometimes they apparently ensure protein stability, while in other cases they determine its half-life. For instance, the degradation of carbohydrate moieties on Serum proteins or Erythrocyte membranes serves as a signal for Liver Cells to engulf and destroy the corresponding proteins. Conversely, glycosylation protects proteins from the action of proteinases. In some instances, carbohydrate components serve recognition functions, as diverse combinations of Monosaccharides in carbohydrate chains allow for the creation of unique "tags." Such tags are utilized, for example, by the Golgi apparatus during protein sorting (see Chapter 22).

In the case of mucin glycoproteins (mucus), which protect Tissues from damage, O-glycosides facilitate the formation of an extended polypeptide conformation necessary to maintain the mesh-like structure of mucins even in highly dilute solutions. Maintaining the protein chain in an extended state appears to be a function of O-glycosides in other contexts as well. In low-density lipoprotein receptors, one can distinguish a membrane-anchoring region and the receptor domain proper, which projects beyond the membrane. Both parts are connected by a long peptide tether decorated with numerous O-glycosidic residues (Fig. 2.7). There is an extensive family of Proteoglycans, in which the carbohydrate components are rich in aminosugars, sulfosugars, and carboxylated sugar derivatives. These substances are thought to play a vital role in the Formation of the Extracellular matrix, though this topic lies beyond The Scope of this book.

Fig. 2.7. Structure of the low-density lipoprotein (LDL) receptor protein. The glycosylated region of the polypeptide chain forms a tether that holds the functionally active protein globule at a distance of more than 10 nm from The Plasma Membrane surface

Protein modules, or domains

In proteins consisting of a single polypeptide chain, the native structure is typically a compact entity whose individual parts cannot exist independently while retaining the structure they possess within the intact globule. However, this is not always the case, particularly when a protein contains more than 200 amino acid residues. The three-dimensional structure of large proteins sometimes reveals not just one, but several more or less independently folded compact regions connected by poorly structured polypeptide stretches. One gets the impression that if these compact regions could be isolated in their Native State, they would retain their folding intact. Indeed, this has been achieved in some cases. Based on this observation, THE CONCEPT OF protein modules, or domains, was introduced, referring to polypeptide fragments whose properties resemble those of independent globular proteins. It is generally assumed that a domain corresponds to a continuous segment of a polypeptide, thereby ruling out a situation where the chain leaves a compact domain and subsequently loops back into it. A domain must be autonomous. An apt analogy here is with the individual movements of a symphony: they are complete musical pieces and can be performed separately, yet they are subordinate to the overarching concept, and only within that framework do their true meaning and interrelationship become clear. What domains actually look like can be seen in Fig. 2.6, d, which shows the structure of the enzyme pyruvate kinase, where three compact regions are readily discernible.

Why are domains of interest?

Domain architecture is frequently associated with the ability to view a protein's functional activity as the sum of distinct elementary processes. Perhaps the most striking example of this is the mammalian enzyme fatty acid synthase, whose single polypeptide chain contains everything required to catalyze seven sequential reactions. Given that in Bacteria each of these reactions is catalyzed by a separate enzyme, it is quite likely that the domains of the synthase once evolved as separate proteins that subsequently merged via Gene Fusion. Many enzymes with similar functions bind at least two substrates. A typical example is provided by nicotinamide adenine dinucleotide dehydrogenases, or (NAD+) dehydrogenases, which catalyze the reaction:

АН2 + NAD+ <-> А + NADH + Н+

All of them bind NAD+, but their oxidizable substrates differ.

It turns out that dehydrogenase molecules contain two domains: one of them binds NAD+ and has a similar structure across all enzymes of this family, whereas the other binds oxidizable substrates AH2 and varies in structure among different dehydrogenases. These data can be interpreted in two ways. One might assume that There is a unique protein structure capable of efficiently binding NAD+, and that evolution independently finds the single correct solution every time (known as convergent evolution). An alternative hypothesis is that evolution repeatedly reuses a successful design of the NAD+-binding module once it has been found, engaging in something akin to shuffling ready-made modules.

The idea of a modular design in enzymes and other proteins allows for the frequent appearance and Rapid Evolution of new functional proteins. This is somewhat reminiscent of assembling an electronic device by combining ready-made, versatile microchips for various purposes, all inserted into a common motherboard. Since we are dealing with microchips rather than triodes or resistors, it is quite feasible to end up with a sensibly functioning circuit. Similarly, A wide variety of functions can emerge in new proteins formed by the association of different domains from pre-existing molecules.

All of this might seem like pure fantasy based on the similarities between functionally related proteins, especially given that there is an alternative interpretation of convergent evolution. Chapter 21 will discuss how domains are sometimes encoded by separate gene fragments called exons.

This provides a certain material basis for considering the "shuffling" of exons—and consequently domains—as a possible mechanism of Protein Evolution. In any case, such a mechanism could operate much faster than random point Mutations.

Proteins of Hair and Connective Tissues

It is remarkable that proteins form The basis of such water-insoluble and durable Materials as horns, hooves, wool, Skin, and tendons. Such proteins must meet various biological requirements. Hair is essentially a long, fairly durable, and insoluble fiber whose structural basis is the protein α-keratin. Tendons, which connect Muscles to bones, must be stronger than steel and devoid of elasticity. These properties are provided by another structural protein, Collagen. Skin, arterial walls, and pulmonary alveoli require entirely different mechanical characteristics. Here, not only strength is needed, but also elasticity and resilience—that is, the ability to reversibly change shape under load. It is hardly surprising that the protein responsible for these properties is called Elastin.

Next, we will examine the structure of keratin, collagen, and elastin. A common feature of these proteins is the involvement of covalent non-peptide bonds in forming their spatial architecture. Recall that in standard proteins, weak interactions make the primary contribution to stabilizing the three-dimensional structure; however, these interactions are clearly insufficient to form the unique structures characteristic of these proteins.

α-Keratins of Hair, Wool, Horns, and Hooves

Hair and wool keratins form so-called Intermediate filaments. They consist of long polypeptide chains with large domains formed by α-helices containing repeating sequences of seven amino acid residues (heptapeptides). This constitutes the central region of the polypeptide chain, whereas the terminal domains lack any preferred regular structure and may vary among different keratins. Two parallel keratin chains form a superhelix in which nonpolar amino acid residues are shielded from water by facing inward. This structure is further stabilized by numerous disulfide bonds formed between cysteine residues of adjacent chains. Superhelical dimers, in turn, associate to form tetramers resembling a four-strand cable. Evidence suggests that tetramers may aggregate within intermediate filaments to form octameric structures.

Structure of Collagens

Collagen is formed extracellularly from a protein secreted by cells called procollagen, which is converted into mature collagen through the action specific enzymes. A procollagen molecule is a triple superhelix formed by three twisted, helical polypeptides (Fig. 2.8).

Fig. 2.8. Models of fibrous proteins

a - Packing of fibrils in collagen fibers (cross-links between lysine residues, which are also present in the triple helix, are highlighted in color); b - structure of cross-links between adjacent lysine residues. The Cytology/cytology/26.html">Structure of Different collagen types depends on their biological functions

Procollagen conversion begins with the Cleavage of terminal peptides. Afterward, the protein—now called tropocollagen—is assembled into collagen fibers. Each of the three polypeptides in the tropocollagen molecule adopts a left-handed helix (recall that the typical α-helix in proteins is right-handed). Approximately one-third of the amino acid residues in tropocollagen are proline, and every third residue is glycine.

During collagen formation, many proline and lysine residues are hydroxylated, turning into hydroxyproline and hydroxylysine, respectively. Notably, these Amino acids are not part of the "magic twenty"; they are incorporated into the protein not via template-directed synthesis, but through post-translational chemical modification of existing amino acid residues. Proline hydroxylation requires ascorbic acid (Vitamin C) as a cofactor, which is necessary to maintain the Fe2+ ion in the active center of the enzyme prolyl hydroxylase in a reduced state. This is why a Vitamin C Deficiency disrupts Connective Tissue formation, leading to the disease known as scurvy.

The architecture of the collagen superhelix is a prime example of how evolution efficiently exploits the possibilities of Protein Engineering. Three spirally intertwined tropocollagen molecules are covalently bound together, creating an extremely robust structure. Such an association is impossible in a standard protein helix because bulky side chains would sterically hinder it. In collagen, the helices are more extended, with three residues per turn. Because every third residue is glycine, which lacks a side chain, the helices come into exceptionally close contact at these points. Additional stabilization is provided by the participation of hydroxylated lysine and proline residues in hydrogen bond formation.

Tropocollagen molecules contain about 1000 amino acid residues. They assemble into collagen fibrils by joining head-to-tail (see Fig. 2.8, a). The gaps in this structure can serve as initial nucleation sites for the deposition of hydroxyapatite crystals, Ca5(OH)(PO4)3, which play a crucial role in bone mineralization. To acquire the required mechanical properties—for instance, in tendons—collagen undergoes enzymatic modification. During this process, lysine residues at the terminal regions of tropocollagen chains are covalently cross-linked (see Fig. 2.8, b). Thus, tendons consist of bundles of parallel-oriented fibrils of mature collagen (i.e., collagen that has completed all stages of post-translational modification). Unlike tendons, collagen fibrils in the skin form a random two-dimensional meshwork. Several types of collagen are known, differing in the structure of their polypeptide chains. These variations are determined by both the content of hydroxyproline and hydroxylysine and the extent of their glycosylation. Genetically determined defects in collagen maturation lead to severe hereditary disorders known as collagen diseases.

Structure of Elastin

In its architecture, elastin is entirely different from collagen or α-keratin. It contains regular α-helices, but due to specific Structural Features of the molecule, it exhibits greater elasticity than rubber. Elastin forms a cross-linked network whose exceptional mechanical properties stem from a unique mode of covalent bonding between lysine side chains: four closely spaced lysine residues form a so-called desmosine structure (Fig. 2.9), which anchors four distinct regions of peptide chains into a single node.

Fig. 2.9. Desmosine cross-link between four polypeptide chains of elastin. Such cross-links arise from the enzymatic modification of four lysine residues

The Mystery of Protein Folding

How do proteins fold into their secondary and tertiary structures? This is the topic of Chapter 21, but it is worth emphasizing here that the amino acid sequence merely determines which three-dimensional structure is possible for a polypeptide chain, rather than dictating how this structure actually forms. The latter remains one of the greatest unsolved problems in modern biology. It is known that a "bad" sequence prevents the formation of a "good" structure, but it is still unknown how a "good" sequence leads to a "good" structure.

During Protein Biosynthesis in the cell, the folding of polypeptide chains into globules takes on the order of a few minutes. If such folding involved randomly sampling all conceivable chain Conformations, it would take millions of years. Moreover, some of these conformations can be metastable, and their transition into the correct conformation would involve overcoming a significant energy barrier. Nature has managed to solve this extremely complex problem.

Questions for Chapter 2

1. What is the primary structure of a protein?

2. What is Protein Denaturation?

3. Draw the structure of an amino acid whose side chain is:

a) a hydrogen atom;

b) an aliphatic hydrophobic group;

c) an aromatic hydrophobic group;

d) acidic;

e) basic.

4. Give approximate pKa values for the Functional groups of the side chains of:

a) acidic amino acids;

b) basic amino acids;

c) histidine.

5. Which amino acid residues determine the net charge of a protein molecule containing all 20 amino acids?

6. Describe the four levels of protein molecular organization.

7. What functional role do α-helices and β-sheets play in protein structure?

8. How do proteins, which are so sensitive to various types of environmental influences, manage to form tendons that are incredibly resistant to tension?

9. What properties of elastin determine its elasticity?



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.