LEHNINGER PRINCIPLES OF BIOCHEMISTRY - VOLUME 1. THE FOUNDATIONS OF BIOCHEMISTRY: STRUCTURE AND CATALYSIS - 2011

PART I. STRUCTURE AND CATALYSIS

3. AMINO ACIDS, PEPTIDES, AND PROTEINS

3.4. Protein Structure: Primary Structure

Protein Purification is merely the starting point for a detailed biochemical analysis of a protein's Structure and function. Why is one protein an enzyme, another a hormone, a third a structural protein, and yet another an antibody? How do they fundamentally differ from a chemical perspective? The most obvious differences are reflected in Cell/13.html">Protein Structure at every level of Organization.

The task of describing and understanding The structure of such large macromolecules as Proteins is addressed at various levels of complexity. Proteins generally exhibit four LEVELS OF STRUCTURAL organization (Fig. 3-23). All covalent bonds (chiefly peptide bonds and disulfide cross-links) linking amino acid residues in a polypeptide chain constitute the Primary Structure of a protein. The most important element of primary structure is the Amino Acid Sequence. Secondary structure refers to particularly stable and recurring local Conformations of the polypeptide chain. Tertiary structure describes all aspects of the three-dimensional folding of a polypeptide. When a protein consists of two or more polypeptide chains, their spatial arrangement is described as quaternary structure. As we explore proteins, we will encounter complex protein aggregates consisting of dozens or even thousands of subunits. The Introduction/19.html">Primary structure of the protein is the subject of this chapter; higher Levels of Protein organization are discussed in Chapter 4.

Class="center">Fig. 3-23. Levels of organization in protein molecules. The primary structure of a protein is its amino acid sequence, linked by peptide bonds and disulfide bridges. The resulting polypeptide can fold into a specific secondary structure, such as an α-Helix. The helix forms part of the Tertiary Structure of the molecule, which in turn may represent just one of several subunits in a multimeric protein, together forming its quaternary structure (the structure of Hemoglobin is shown here).

Differences in primary structure can be particularly informative. Each protein is characterized by a specific number and sequence of amino acid residues. As we will see in Chapter 4, it is the primary structure of a protein that dictates its three-dimensional structure, which in turn governs its function. Our focus in this chapter is on primary structure. First, we will examine the empirical observation that a protein's amino acid sequence and its function are intimately related. Next, we will discuss Methods for determining Amino acid sequences, and finally, we will explore how much information can be gleaned from a protein's primary structure.

Protein Function Depends on Amino Acid Sequence

The bacterium Escherichia coli synthesizes over 3,000 different proteins, whereas humans have about 25,000 genes encoding an even greater variety of proteins (genetic aspects are discussed in Part III of this book). In both cases, each type of protein is characterized by a unique three-dimensional structure required for a specific function. Each protein also possesses a unique amino acid sequence. Intuition suggests that The amino acid sequence must determine the three-dimensional STRUCTURE OF THE protein and, ultimately, its function. But is this actually the case? A brief survey of The Diversity of protein primary structures provides a body of empirical evidence demonstrating a strong correlation between amino acid sequence and biological function.

First, as noted above, proteins with different Functions invariably differ in their amino acid sequences. Furthermore, thousands of human genetic diseases are associated with The production of abnormal proteins. These defects range from the substitution of a single amino acid (as in Sickle-Cell Anemia, see Chapter 5) to the deletion of a large segment of a polypeptide chain (as in most cases of Duchenne muscular dystrophy, where a deletion in the Gene encoding dystrophin leads to the synthesis of a truncated, inactive protein). We now know that altering the primary sequence can alter protein function. Finally, when comparing functionally homologous proteins across different species, these proteins often share a high degree of sequence similarity. For example, ubiquitin, a 76-amino-acid protein involved in regulating the degradation of other proteins, has been shown to have an identical amino acid sequence in organisms as diverse as fruit flies and humans.

Is the amino acid sequence of a given protein entirely invariant? The answer is no; some variation is tolerated. It is estimated that 20–30% of human proteins are polymorphic, meaning that these proteins exist in human populations with slightly differing amino acid sequences. Many of these sequence variations have little to no effect on protein function. Moreover, proteins with similar functions from distantly related species may vary widely in size and amino acid sequence.

While alterations in certain regions of an amino acid sequence may not affect protein function, most proteins contain sequence segments that are critical for their activity, and these regions are highly conserved. The fraction of the sequence critical for function varies among different proteins, which complicates the correlation of primary and tertiary structures with structure and function. Before examining this problem in more detail, we will discuss how the amino acid sequences of proteins are determined.

The Amino Acid Sequences of Millions of Proteins Have Already Been Decoded

Two major breakthroughs in 1953 played a pivotal role in The history of biochemistry. That year, James D. Watson and Francis Crick proposed the double-helix model of DNA and suggested a structural mechanism for its Replication (Chapter 8). Their hypothesis provided the molecular foundation for METABOLISM/2.html">THE CONCEPT OF the gene. At about the same time, Frederick Sanger determined the amino acid sequence of the hormone Insulin (Fig. 3-24), astonishing many scientists who considered decoding the amino acid sequence of a polypeptide chain an insurmountably difficult task. It soon became clear that The nucleotide sequence of DNA and the amino acid sequence of a protein are somehow related. Just a decade after these discoveries, it was established that DNA dictates the amino acid sequence of proteins (Chapter 27). Today, vast numbers of protein sequences can be deduced from DNA sequences compiled in ever-expanding Databases. However, classical protein sequencing remains widely used in modern Protein Chemistry because it can sometimes reveal details—such as post-translational modifications—that cannot be inferred from gene sequences alone.

Fig. 3-24. Amino acid sequence of bovine insulin. The two polypeptide chains of the protein are linked by Disulfide Bonds. Human, porcine, canine, rabbit, and whale insulins share identical A chains, whereas bovine, porcine, canine, goat, and equine insulins have identical B chains.

Chemical protein sequencing now complements a growing array of newer techniques that provide numerous ways to acquire such data, which are vital to all areas of biochemical research.

Frederick Sanger

Short Polypeptides Are Sequenced Using Automated Sequenators

A variety of methods are used to analyze protein primary structure. Techniques have been developed to label and identify the amino-terminal (N-terminal) amino acid residue (Fig. 3-25a). Sanger pioneered The Use of 1-fluoro-2,4-dinitrobenzene for this purpose. Dansyl chloride and dabsyl chloride are also used, yielding derivatives that are easier to detect than dinitrophenyl derivatives. After coupling one of these Reagents to the N-terminal residue, the polypeptide chain is hydrolyzed (with 6 M HCl) into its constituent free Amino acids, and the labeled residue is identified. Because this Procedure destroys the polypeptide, it cannot be used to determine subsequent residues beyond the terminus. However, it can be used to determine the number of distinct polypeptide chains in a protein if they all possess different N-termini. For example, subjecting insulin to this procedure reveals two N-terminal residues, Phe and Gly (Fig. 3-24).

Fig. 3-25. Stages of polypeptide chain sequencing. (a) The initial step in sequencing may involve identifying the N-terminal amino acid residue. The Sanger method for N-terminal residue determination is shown here. (b) The Edman Degradation allows the Determination of the complete amino acid sequence. For short Peptides, this method permits complete sequencing of the entire chain, rendering step (a) often unnecessary. Step (a) is particularly useful for large proteins that are first cleaved into smaller fragments prior to sequencing (Fig. 3-27).

To sequence (from the Latin sequentia — sequence) an entire protein, the method proposed by Pehr Edman is typically used. In the Edman degradation method, only the N-terminal residue is cleaved from the polypeptide chain without affecting all the others (Fig. 3-25, b). To do this, the peptide is treated under mildly alkaline conditions with phenyl isothiocyanate, which attaches to the N-terminal amino acid. The resulting adduct is then cleaved using anhydrous trifluoroacetic acid as an anilinothiazolinone derivative, extracted with organic Solvents, converted into a more stable phenylthiohydantoin derivative in an aqueous acid solution, and identified. Performing the reactions sequentially—first under alkaline and then under acidic conditions—makes it possible to monitor the progress of the entire process. All Reactions Involving the N-terminal amino acid residue have no effect on the rest of the amino acid sequence. After the terminal residue is removed and identified, the newly exposed N-terminal residue can be labeled, removed, and identified again using the same sequence of reactions. This procedure is repeated until all amino acid residues of the peptide chain have been determined. The Edman degradation METHOD FOR DETERMINING amino acid sequences is carried out using an automated instrument called a sequenator, in which the stages of reagent mixing, product Separation, identification, and data recording are fully automated. This method is extraordinarily sensitive. Sometimes it allows the amino acid sequence of a protein to be determined using only a few micrograms of a sample.

The length of a polypeptide chain that can be sequenced via Edman degradation depends on the efficiency of the individual chemical steps. Imagine a protein whose N-terminus begins with the sequence Gly-Pro-Lys, and so on. If Glycine is removed with an efficiency of 97%, this means that 3% of the polypeptide molecules in solution still retain a glycine residue at their N-terminus. In the second cycle, Proline accounts for 94% of the free amino acids (97% of Pro residues are cleaved from 97% of the molecules ending in Pro), glycine accounts for 2.9% (97% of Gly residues are cleaved from 3% of the molecules ending in Gly), and 3% of the molecules retain either glycine (0.1%) or proline (2.9%) at their N-terminus. Thus, in each cycle, the peptides that failed to be cleaved in the previous step introduce an increasingly larger error; eventually, it becomes impossible to determine which amino acid currently resides at the end of the chain. Modern sequenators achieve an efficiency exceeding 99% per cycle, making it possible to determine sequences of 50 or more residues. The primary structure of insulin, which Sanger and his coworkers spent over 10 years working on, can now be determined in 1–2 days by direct sequencing in a protein sequenator. (As we will discuss in Chapter 8, DNA Sequencing is even more efficient.)

Large proteins must be cleaved into fragments prior to sequencing

As the length of a polypeptide increases, the number of errors in determining the amino acid sequence typically rises. To determine the sequences of large polypeptides and proteins, they must be cleaved into smaller fragments, and each fragment must then be sequenced separately. This process involves several stages. First, the protein is cleaved into specific fragments using chemical or enzymatic methods. If the protein contains disulfide bonds, they must be broken. Each fragment is isolated from the mixture and sequenced using the Edman method. Finally, the order of the fragments in the original chain and the locations of any disulfide bonds are determined.

Cleavage of disulfide Bonds. The presence of disulfide bonds hinders the determination of the amino acid sequence. A Cysteine residue (Fig. 3-7) cleaved from the polypeptide chain during the Edman reaction may remain attached to another polypeptide chain via a disulfide bridge. In addition, disulfide bridges interfere with the enzymatic and chemical Cleavage of the polypeptide chain into fragments. Figure 3-26 schematically illustrates two Methods for the irreversible cleavage of disulfide bonds.

Fig. 3-26. Cleavage of disulfide bridges in proteins by two well-known methods. Oxidation of cysteine residues with performic acid yields cysteic acid residues. Following reduction with dithiothreitol or β-mercaptoethanol to form two cysteine residues, further Modification of the reactive -SH groups is necessary to prevent The formation of a new disulfide bridge. Acetylation with iodoacetate is used for this purpose.

Cleavage of the polypeptide chain. There are several methods for splitting a polypeptide chain into fragments. The Hydrolysis of peptide bonds is catalyzed by Enzymes called proteases. Certain proteases specifically cleave only those peptide bonds that link specific amino acid residues (Table 3-7). Using such proteases makes it possible to reproducibly obtain predictable fragments. Certain chemical reagents are also capable of cleaving the peptide bond between specific amino acids.

Table 3-7. Specificity of some methods for polypeptide chain cleavage

Reagent (biological source)*

Cleavage site**

Trypsin (bovine Pancreas)

Lys, Arg (C)

Submandibular gland protease (mouse)

Arg (C)

Chymotrypsin (bovine pancreas)

Phe, Trp, Tyr (C)

Protease V8 (bacterium Staphylococcus aureus)

Asp, Glu (C)

Asp-N-protease (bacterium Pseudomonas fragi)

Asp, Glu (N)

Pepsin (porcine Stomach)

Leu, Phe, Trp, Tyr (N)

Endoproteinase Lys C (bacterium Lysobacter enzymogenes)

Lys (C)

Cyanogen bromide

Met (C)

* All reagents except cyanogen bromide are proteases; all are commercially available.

** The indicated amino acid residues are the recognition sites for the enzyme or reagent, which cleaves the peptide bond on the C- or N-side of the specified residue.

The digestive enzyme trypsin, which belongs to the class of proteases, catalyzes the hydrolysis of only those peptide bonds whose carbonyl group belongs to a Lysine or Arginine residue, regardless of the length of the polypeptide chain. Thus, knowing the total number of Lys and Arg residues in the polypeptide chain—determined from complete hydrolysis of the sample (Fig. 3-27)—one can predict the number of fragments into which the polypeptide chain will be split upon Treatment with trypsin. A polypeptide containing five Lys and/or Arg residues in its sequence should be cleaved into six fragments by the action of trypsin. Each of these fragments, except for one, will have a Lys or Arg residue at its C-terminus. The fragments resulting from cleavage by trypsin, another enzyme, or a chemical reagent are subsequently purified by chromatographic or electrophoretic methods.

Fig. 3-27. Cleavage and sequencing of peptides and determination of the order of peptide fragments. First, the Amino Acid Composition and N-terminal amino acid of the initial peptide are determined. Then, any disulfide bonds present in the molecule that interfere with sequencing are broken. In this polypeptide, there are only two cysteine (C) residues and, accordingly, only one possible Location for a disulfide bond. In polypeptides with three or more cysteine residues, the arrangement of disulfide bridges is determined as described in the text. For the decoding of one-letter and three-letter amino acid Abbreviations, see Table 3-1.

Peptide sequencing. Each fragment obtained by trypsin Digestion is sequenced separately using the Edman method.

Determining the order of fragments in the original polypeptide. At this stage, it is necessary to establish the order of the "tryptic fragments." To do this, a sample of the original polypeptide is subjected to cleavage by another enzyme or reagent that breaks peptide bonds between amino acid residues different from those targeted by trypsin. For example, cyanogen bromide can be used, as it exclusively cleaves peptide bonds whose carbonyl group belongs to a Methionine residue. The resulting fragments are purified again and sequenced.

Next, the amino acid sequences of all available fragments must be examined to find fragments generated by the second cleavage that overlap the gaps between the fragments obtained in the first experiment (Fig. 3-27). The amino acid sequence of these overlapping peptide regions makes it possible to establish the order of the fragments produced by the first cleavage. If the amino acid at the N-terminus of the polypeptide was identified prior to the first cleavage, this information can be used to find the fragment adjacent to the N-terminus. Furthermore, by comparing the two sets of fragments, potential errors in determining the amino acid sequence can be identified. Sometimes a second cleavage of the polypeptide into fragments is not sufficient to find the overlapping sequences for certain fragments. In such cases, a third or even a fourth cleavage method is applied, ultimately yielding a complete set of overlapping sequences for the original chain.

Localization of disulfide bonds. The locations of disulfide bonds in the original molecule are determined after the sequence has been fully elucidated. To do this, the original polypeptide is cleaved again—for example, with trypsin—this time without prior disruption of the disulfide bonds. The resulting fragments are separated by Electrophoresis and compared with the set of fragments obtained during the initial tryptic digestion. If a disulfide bond exists between two fragments, those fragments will be absent from the new set, and a heavier fragment will appear in their place. The two missing peptides correspond to the Regions of the original polypeptide chain connected by the disulfide bond.

Amino acid sequences can be deduced from other data

The method described above is not the only way to establish an amino acid sequence. New mass spectrometric

methods make it possible to determine the sequences of short peptides (20–30 amino acid residues) in just a few minutes (Box 3-2). In addition, with the advancement of rapid DNA Sequencing Methods (Chapter 8), the elucidation of The Genetic Code (Chapter 27), and The Emergence of gene identification techniques (Chapter 9), it has become possible to determine the amino acid sequence of a polypeptide based on the nucleotide sequence of the gene encoding it (Fig. 3-28). DNA and protein sequencing methods are complementary. For a protein whose gene is available, DNA sequencing can yield faster and more accurate results than sequencing the protein itself. The amino acid sequences of most proteins are now determined using this indirect approach. If the gene has not been isolated, protein sequencing must be performed; the resulting data provide more information than DNA sequence analysis alone, which, for instance, reveals nothing about the positions of disulfide bonds in the molecule. Moreover, knowing the amino acid sequence of even a small region of a polypeptide chain can greatly facilitate the Isolation of the corresponding gene (Chapter 9).

Fig. 3-28. Correspondence between DNA and amino acid sequences. Each amino acid is encoded by a specific sequence of three NUCLEOTIDES in the DNA strand. For details on the genetic code, see Chapter 27.

Practical Biochemistry: Mass Spectrometric Methods for Protein Study

Mass spectrometry is one of the most powerful instruments for chemical research. The molecules under analysis (analytes) are first ionized in a vacuum. The resulting charged particles are then introduced into an electric and/or magnetic field, where their motion depends on their mass-to-charge ratio (m/z). This measurable value can be used to determine the molecular weight (M) of the analyte with high precision.

Mass spectrometry has been used as an analytical method for many years, but until relatively recently, it was impossible to apply this technique to macromolecules such as proteins and Nucleic Acids. The challenge is that measuring the m/z ratio takes place in the gas phase, and heating or any other procedure required to vaporize the substance usually leads to its degradation. In 1988, two approaches were introduced that successfully resolved this problem. In one method, the protein is embedded in a light-absorbing matrix. Upon pulsed laser irradiation, the proteins within the matrix are ionized and released into the vacuum chamber. This method, known as matrix-assisted laser desorption/ionization mass spectrometry (MALDI), is successfully used to determine the molecular weights of a wide range of macromolecules. In the second widely used method, macromolecules are transferred directly to the gas phase from solution. The analyte solution is passed through a spray nozzle placed in a strong electric field, converting the solution into extremely fine charged droplets. The solvent instantly evaporates, leaving the intact charged macromolecules in the gas phase. This method is called electrospray ionization mass spectrometry (ESI-MS). Protons attaching to the macromolecules as they pass through the spraying device impart an additional charge. The m/z value is then determined in the vacuum chamber.

Mass spectrometry provides a wealth of information essential for research in Proteomics, enzymology, and protein chemistry as a whole. These experiments require minimal amounts of samples: sufficient protein for analysis can be obtained, for example, by two-dimensional electrophoresis. Precise determination of a protein's molecular weight is a key step in its identification. Once the molecular weight is known, mass spectrometry allows the detection of all changes associated with the binding of Cofactors, Metal Ions, covalent protein modifications, and so forth. Fig. 1 illustrates an example of determining a protein's molecular weight using electrospray mass spectrometry.

Fig. 1. Electrospray mass spectrometry method for determining protein molecular weight. a) A protein solution is sprayed into highly charged ultrafine droplets by passing it through a spray nozzle placed in a strong electric field. The solvent evaporates, and the analyte ions (in this case, protonated) are introduced into the mass spectrometer to determine the m/z ratio. b) The resulting spectrum consists of a series of peaks, each differing from the preceding one (from right to left) by an increment of one in mass and charge. The inset shows the computer-processed result of this spectrum.

Upon transitioning to the gas phase, a protein acquires a certain number of protons (i.e., positive charges) from the solvent molecules. This produces a spectrum of particles with varying mass-to-charge ratios. Each registered peak corresponds to a distinct type of particle, differing from its neighbor by one unit of charge and one unit of mass (one proton). The Molecular Weight of the protein can be calculated from any two adjacent peaks. For a single peak:

where M is the molecular weight of the protein, n2 is the number of charges, and X is the molecular weight of the added group (in this case, a proton). Similarly, for the adjacent peak:

We now have two unknowns (M and n2) and two equations, allowing us to solve first for n2 and subsequently for M:

This method of calculation based on the m/z values of any two peaks from the spectrum (Fig. 1, b) typically yields the protein mass with an error not exceeding 0.01% (in this case, the Analysis of the protein aerolysin k yielded M = 47,342). Using multiple series of peaks, repeated calculations, and averaging the results can further enhance the accuracy of the molecular weight determination M. Computer software can also obtain an exact result from the m/z value of a single peak (Fig. 1, b, inset).

Mass spectrometry is an invaluable tool for rapidly determining the short amino acid sequence of an unknown protein. This technique is called tandem mass spectrometry. First, the test protein solution is treated with a protease or chemical reagent to generate a mixture of shorter peptides. This mixture is then introduced into an instrument consisting of two mass spectrometers arranged in series

(Fig. 2, a, top). In the first mass spectrometer, the mixture of ionized peptides is sorted so that only a single type of cleavage fragment reaches the next compartment. This sample, with every molecule carrying a specific charge, passes through a vacuum chamber situated between the two mass spectrometers. In this so-called "collision cell," the peptide is broken down into even smaller pieces through energetic collisions with inert gas molecules (argon or helium) fed into the chamber. The process is conducted under conditions such that each individual peptide fragment is cleaved into an average of no more than two parts. Most cleavages occur at the peptide bonds. Water molecules do not participate in this reaction since everything takes place in a vacuum chamber, and consequently, reaction products include radicals, such as carbonyl radicals (Fig. 2, a). The charge originally residing on the parent fragment is now concentrated on one of its portions.

Fig. 2. Determination of Amino acid sequence by tandem mass spectrometry. a) The protein solution after proteolytic cleavage is introduced into the mass spectrometer (MS-1). Peptides are sorted, selecting only a single type of fragment for further analysis. This fragment is then subjected to fragmentation in a collision cell located between the two mass spectrometers. In the second mass spectrometer (MS-2), the m/z value for each fragment is determined. Many of the resulting ions are produced by the cleavage of peptide bonds. In the figure, they are designated as b-type and y-type ions depending on whether the charge remains on the C-terminal or N-terminal portion of the fragment. b) A typical spectrum obtained from Processing a peptide consisting of 10 amino acid residues. Labeled peaks correspond to y-type ions. The highest peak (adjacent to the y5" peak) corresponds to a doubly charged ion and does not belong to this series. The peaks differ from one another by a single adjacent amino acid in the peptide chain. In this case, the determined peptide sequence is Phe-Pro-Gly-Gln-(Ile/Leu)-Asn-Ala-Asp-(Ile/Leu)-Arg. Note that there is an ambiguity in identifying leucine and isoleucine, which share the same molecular weight. In our example, the series of peaks belonging to y-type ions is predominant, which greatly facilitates spectrum interpretation. This is due to the presence of an Arg residue at the C-terminus of the peptide, which concentrates most of the positive charge.

The second mass spectrometer measures the m/z ratios of all charged fragments (uncharged particles are not detected). This yields one or more series of peaks. Each peak series (Fig. 2, b) reflects the presence of all charged fragments generated by the cleavage of the same type of bond (albeit in different PARTS OF THE polypeptide molecule) on the same side of the bond (N or C). Adjacent peaks in a single series differ by one amino acid. By examining the molecular mass differences between corresponding peaks, one can determine which specific amino acid comes next in the polypeptide chain. Difficulties arise solely in distinguishing between leucine and isoleucine, as they have identical molecular weights.

During peptide fragmentation, the charge can be retained on either the N-terminal or C-terminal fragment. Furthermore, cleavage may occur at sites other than peptide bonds, so a single experiment typically yields multiple series of peaks. The two most prominent spectra usually correspond to charged fragments resulting from peptide bond cleavage, and the set of C-terminal fragments is readily distinguishable from the set of N-terminal fragments. This is because

the fragmentation process in the collision cell does not generate conventional carboxyl or amino groups at the cleavage sites. The only authentic α-amino and α-carboxyl groups are those located at the very ends of the fragments (Fig. 2, a). Consequently, the two sets of fragments exhibit slight differences in molecular weights. Determining the amino acid sequence using both sets of fragments increases result accuracy.

If the nucleotide sequence of a gene is known, reading just a short fragment of the amino acid sequence is usually sufficient to unambiguously correlate the gene with its protein. Mass spectrometry sequencing cannot replace Edman degradation for sequencing long fragments, but it is ideally suited for proteomics research aimed at cataloging hundreds of cellular proteins that can be separated by two-dimensional electrophoresis.

Current methods of PROTEIN AND NUCLEIC acid analysis have paved the way for The Development of a new discipline: "whole-cell biochemistry." It is now possible to study the complete DNA sequences (genomes) of A wide variety of organisms, from Viruses and Bacteria to Multicellular Organisms (Tables 1-2). New genes are being discovered by the thousands, and some of them encode proteins with yet-unknown functions. To describe the complete Complement of proteins encoded by an Organism's genome, scientists have introduced the term "proteome." As will be discussed further in Chapter 9, the new fields of proteomics and Genomics are complementary areas of research in cellular and NUCLEIC ACID METABOLISM, aimed at building an increasingly complete picture of the biochemistry of Cells and entire organisms.

Small Peptides and Proteins can be synthesized chemically

Many peptides are used in pharmacology, making their synthesis of considerable commercial interest. There are three ways to obtain a peptide: 1) isolation from tissue, which is often hindered by its extremely low concentration; 2) Recombinant DNA technology (Chapter 9); and 3) chemical synthesis. In many cases, direct chemical synthesis methods are preferred. Beyond commercial Applications, chemical synthesis is used to generate specific peptide fragments required for studying Protein Structure and function.

Traditional organic synthesis methods are unsuitable for producing peptides and proteins containing more than 40-50 amino acid residues due to the inherent complexity of protein structures. One of the challenges encountered along this path is the purification of the product following synthesis.

A major contribution to the development of Peptide Synthesis methods was made by Robert Bruce Merrifield in 1962. He proposed synthesizing a peptide by anchoring one of its ends to a solid support. An insoluble polymer (resin) placed in a Column similar to those used in Chromatography can serve as such a support. Amino acids are attached one by one to the support-bound peptide using a standard set of reactions, after which the cycle is repeated with another amino acid (Fig. 3-29). At each stage of the cycle, protecting chemical groups are used to prevent Side Reactions. Today, the technology of chemical peptide synthesis is fully automated. Much like sequencing reactions, this process is limited by the efficiency of each individual reaction cycle (this can be understood by calculating the overall efficiency of a process in which each step proceeds with a yield of 96.0% or 99.8% (Table 3-8)). Incomplete reaction at one stage leads to the formation of an impurity (a shorter peptide) in the next stage. Improvements in chemical methods have made it possible to synthesize proteins consisting of 100 amino acids in good yields in just a few days. A very similar approach is used for the synthesis of nucleic acids (Fig. 8-35). Yet, laboratory synthesis methods still pale in comparison to Biosynthesis in living organisms! In a bacterial cell, an identical 100-amino-acid protein is synthesized with supreme precision in approximately five seconds.

Table 3-8. Effect of individual step yields on the overall yield of peptide synthesis

Number of residues in the final polypeptide

Overall yield of final product (%) depending on the

96.0%

yield at each step

99.8%

11

66

98

21

44

96

31

29

94

51

13

90

100

1.8

82

Robert Bruce Merrifield, 1921–2006

Fig. 3-29. Solid-phase chemical peptide synthesis. The formation of each peptide bond requires reactions 1–4. The protecting fluorenylmethoxycarbonyl group (FMOC, highlighted in blue) prevents side reactions at the $\alpha$-amino group of the amino acid residue (highlighted in red). Chemical synthesis proceeds from the C-terminus to the N-terminus of the peptide chain, i.e., in the opposite direction compared to synthesis in vivo (Chapter 27).

New, efficient methods for ligating (joining) individual peptide fragments make it possible to assemble larger proteins from synthetic peptides. These methods allow the creation of novel proteins in which chemical groups—including those not found in natural proteins—are positioned at strictly defined sites. Such new proteins enable researchers to study the Mechanisms of Enzymatic Catalysis and to engineer proteins with unique chemical properties and predetermined structures. The latter provides conclusive Evidence for the relationship between a protein's primary structure and its three-dimensional structure in solution.

Amino acid sequence serves as a source of vital biochemical information

Knowledge of a protein's amino acid sequence provides insight into its three-dimensional structure, function, cellular localization, and evolution. Crucial information can be obtained by comparing the amino acid sequence of a given protein with those of other known proteins. The Internet provides access to databases containing thousands of sequences known to date. Comparing a novel sequence with this databank frequently reveals both specific features and general patterns.

To this day, we do not know precisely how an amino acid sequence determines three-dimensional structure, nor are we able to unambiguously predict a protein's function solely from its amino acid sequence. However, based on sequence similarity between a given protein and others, it can be assigned to one of the known Protein Families characterized by specific Functional and Structural features. Family members typically share at least 25% sequence identity and possess certain common Structural and functional traits. Nevertheless, some families include members that share only a few identical amino acid residues essential for a specific function. Furthermore, many proteins with diverse functions share similar supramolecular structures (known as domains; see Chapter 4). These domains frequently form specific structures that exhibit unexpectedly high stability or are adapted to particular environments. In addition, structural and functional similarities among protein families provide clues regarding their evolutionary origins.

Certain amino acid sequences act as signals that determine cellular localization, mark proteins for chemical modification, and dictate protein lifespan. Specialized signal sequences, typically located at the N-terminus of a protein, are used to target proteins for Transport Across the cell membrane; other proteins are destined for localization within The Nucleus, at The Cell surface, or in other cellular compartments. Specific sequences serve as binding sites for prosthetic groups, such as sugar residues in Glycoproteins and Lipids in Lipoproteins. Some of these signal sequences are well characterized and easily recognized within the amino acid sequence of a novel protein (Chapter 27).

Consensus sequences and Sequence logos

Consensus sequences can be depicted in several different ways. To illustrate two representation options, we use two Examples of consensus sequences shown in Fig. 1: (a) an ATP-binding structure known as the P-loop (see Box 12–2); (b) a Ca2+-binding structure known as the EF-hand (see Fig. 12–11). The rules presented here are adapted from those used by the PROSITE website (expasy.org/prosite); this software employs the single-letter code to denote amino acid residues.

Fig. 1. Representation of two consensus sequences. (a) P-loop, an ATP-binding structure; (b) EF-hand, a Ca2+-binding structure.

In one method of representing consensus sequences (top portion of panels a and b), each position is separated from the adjacent position by a hyphen. A position that can accommodate any amino acid is designated by x. In cases of ambiguity, all possible amino acids are listed within square brackets. For example, in case (a), [AG] stands for Ala or Gly. If virtually any amino acid except for a few specific ones can occupy a given position, those excluded amino acids are listed within curly braces at that position. For instance, in Fig. 1b, {W} indicates that Trp cannot be present at this position. Repetition of a sequence element is denoted by a number or a range of numbers enclosed in parentheses following that element. Thus, in example (a), x(4) means x-x-x-x, and x(2,4) means x-x, x-x-x, or x-x-x-x. If the region in question represents the N-terminus or C-terminus of the sequence, the notation starts or ends accordingly (which is not the case in these examples). A period is placed at the end of the fragment. Applying these rules to consensus sequence (a), we find that the first position can be occupied by A or G. The next four positions can contain any amino acid, followed obligatorily by G and K, with the final position occupied by S or T.

The Sequence logos system provides a more informative graphical representation of Multiple Sequence Alignments of amino acids (or nucleotides). Each position displays a stack of characters corresponding to the amino acids (or nucleotides) that may occur there.

The total height of the character stack at a given position (in bits) reflects the degree of conservation at that position, while the height of each individual character within the stack indicates the relative frequency of that specific amino acid (or nucleotide). In amino acid sequence logos, the PHYSICOCHEMICAL CHARACTERISTICS OF the amino acids are color-coded: polar amino acids (G, S, T, Y, C, Q, N) are green, basic (K, R, H) are blue, acidic (D, E) are red, and hydrophobic (A, V, L, I, P, W, F, M) are black. The Classification of amino acids in this scheme differs slightly from that presented in Table 3–1 and Fig. 3–5. Amino acids with aromatic side chains are classified as both nonpolar (F, W) and polar (Y). Glycine, which is notoriously difficult to classify, is conventionally considered polar. When one or a small number of Amino acids can occupy a particular position, it is rare for them to occur with equal probability; typically, one or a few amino acids predominate. This mode of representation highlights such predominance, making the conserved sequence within a protein more apparent. However, it may obscure Certain amino acids that could occasionally occupy a specific position, such as Cys, which is sometimes found at position 8 in the EF-hand motif (Fig. 1b).

Key Conventions.

A significant portion of the functional information contained in a protein sequence resides in its consensus sequences. This term denotes corresponding sequences not only in proteins, but also in DNA and RNA. By comparing sets of related amino acid or nucleotide sequences, researchers identify positions where the same amino acids or nucleotides occur most frequently; these are the consensus sequences. Sequence regions that show the highest degree of similarity across a range of organisms often correspond to evolutionarily conserved functional domains. A wide array of mathematical algorithms and software tools accessible via the Internet has been developed for analyzing and searching consensus sequences. Box 3–3 outlines the General Principles for depicting consensus sequences. ■

Protein sequences shed light on the evolution of life on Earth

The simple string of letters in an amino acid sequence does not fully convey the wealth of information actually embedded within it. As the number of known protein sequences grows, increasingly powerful methods are emerging to extract diverse information from them. The analysis of data stored in continuously expanding biological databases—containing gene and protein sequences as well as macromolecular structures—has given rise to bioinformatics as a new field of knowledge. One outcome of this discipline's development is the advent of computer software suites, many of which are freely accessible online for scientists, students, and laypersons alike. A protein's function is intimately linked to its three-dimensional structure, which, in turn, is largely determined by its primary sequence. Thus, the biochemical information reflected in an amino acid sequence is fundamentally limited only by our own understanding of the structural and functional principles of protein organization. The continuously evolving toolkit of bioinformatics allows researchers to identify functional segments in novel proteins and to establish both their sequence and structural relationships with already known proteins. Furthermore, examining protein sequences from an evolutionary perspective enables us to understand how proteins originated and, ultimately, how life evolved on our planet.

The history of molecular evolution is usually associated with the work of Linus Pauling and Emile Zuckerkandl, who in the mid-1960s began actively utilizing nucleotide and amino acid sequences to study Evolutionary Processes. The foundation of this science is deceptively simple: if two organisms are related, the sequences of their genes and proteins should be similar. As organisms diverge evolutionarily, these sequences become increasingly dissimilar. The potential of this approach became evident in the 1970s when Carl Woese used ribosomal RNA sequences to designate archaea as a distinct group separate from bacteria and eukaryotes (Fig. 1–4). Protein sequences frequently provide opportunities to refine existing knowledge. With genome projects underway for a vast array of organisms, from bacteria to humans, the number of available sequences is growing at an astonishing rate. This information can be harnessed to trace the course of evolution. The challenge lies in deciphering these genetic hieroglyphs.

Evolution does not follow a simple, linear path. Complexities arise at every attempt to extract evolutionary information encoded within protein sequences. Throughout the course of evolution, certain amino acid residues holding the greatest importance have remained unchanged in each specific protein

value for its functioning. Residues that do not play such a critical role in protein activity could change over time—that is, one amino acid could replace another—and it is precisely these altered residues that can reveal something about the course of the evolutionary process. However, Amino Acid Substitutions are not always random. In certain regions of the amino acid sequence, only strictly defined substitutions are possible, which is driven by the need to preserve protein function. The amino acid composition of some proteins has changed more drastically than that of others. For these and other reasons, the rate at which proteins evolve varies.

Another factor that hinders tracing the course of evolution is The transfer of genes or groups of genes from one organism to another, known as horizontal (lateral) gene transfer. Transferred genes can be quite similar to those in the original organism, whereas the majority of other genes in these two organisms share only a very distant resemblance. The result of Horizontal Gene Transfer is the rapid spread of Antibiotic Resistance in bacterial populations observed today. Proteins synthesized based on these transferred genes make poor candidates for studying bacterial evolution, because their evolutionary history in the new "host" organism is very short.

The subject of research in molecular evolution is usually families of closely related proteins. Protein families that play a crucial role in cellular metabolism are selected for analysis, as they must have been present in precursor cells as well, which greatly reduces the likelihood of their appearance in cells As a result of horizontal gene transfer. For example, the protein eEF-1α (elongation factor 1α) is involved in Protein Synthesis in all eukaryotes. A similar protein, EF-Tu, is found in bacterial cells. The similarity in sequences and functions indicates that eEF-α1 and EF-Tu are members of a protein family descended from a common ancestor. Proteins belonging to such a family are called homologous proteins, or homologs. If two members of a family, i.e., two homologs, are present in organisms of the same species, they are called paralogs. Homologous proteins from organisms of different species are called orthologs. To study the evolutionary process, it is first necessary to identify suitable families of homologous proteins and then reconstruct the course of evolution based on them.

The identification of homologs is carried out using powerful computer programs that allow for the direct comparison of two or more protein sequences, as well as Searching for Similar sequences in databases. The electronic search process can be visualized as sliding one sequence along the other until a region of sufficient similarity is found. In this sequence alignment process, each position where amino acids from two amino acid sequences match is assigned a specific score, which varies depending on the computer program used; this score serves as an indicator of alignment quality. Certain complexities arise along this path. Sometimes the compared proteins match quite well, say, in two sequence regions separated by less similar regions of varying lengths. As a result, simultaneous alignment of the two matching regions is impossible. To solve this problem, the computer program inserts a gap into one of the sequences so that the two similar regions can overlap simultaneously (Fig. 3-30). Obviously, with a sufficient number of gaps, almost any two sequences can be aligned in some way. To avoid uninformative alignments, the computer program subtracts a specific penalty for each gap, thereby reducing the alignment score. Next, the program selects the optimal alignment that maximizes the number of identical amino acid residues while minimizing the number of gaps.

Fig. 3-30. Alignment of protein sequences using gaps. Short segments of the Hsp70 protein (a widely distributed class of chaperone proteins) from two well-studied bacteria, E. coli and Bacillus subtilis, are shown here. Introducing a gap into the B. subtilis amino acid sequence allows for a better alignment on either side of the gap. Identical amino acid residues are highlighted in yellow.

The arrangement of identical amino acids often fails to determine the evolutionary degree of relatedness between proteins. Studying The chemical properties of amino acid substitutions is much more useful in this regard. Amino acid substitutions within a protein family are often conservative; this means that one amino acid residue is replaced by another with similar chemical properties. For example, in a certain position within one protein of a family, there is Glu, while in another there is Asp; note that both residues carry a negative charge. Logically, such a conservative substitution allows for a more reliable alignment than a non-conservative substitution of that same Asp, for instance, with a hydrophobic Phe residue.

In most cases, for identifying homologs and determining evolutionary relationships, protein sequences (whether obtained through direct amino acid sequencing or deduced from DNA sequences) are much more preferable than non-coding nucleotide sequences (those that do not encode protein or functional RNA sequences). In the case of nucleic acids, which are constructed from only four different residues, a random alignment of non-homologous sequences typically results in a match of at least 25% of positions. Introducing several gaps can often increase this value to 40% or more, resulting in a very high probability of aligning unrelated sequences. The presence of 20 different amino acids in a protein significantly reduces the possibility of such uninformative alignments.

All sequence alignment programs are equipped with methods for testing alignment reliability. In one of the tests, the amino acid sequence of one of the compared proteins is shuffled to generate a random sequence, and the alignment program is then run to compare it with the original sequence. The new alignment is again scored, and the shuffling and alignment processes can be repeated many times. The initial alignment performed before shuffling should have a much higher score than alignments of random sequences. In this case, one can be confident that the aligned sequences truly belong to two homologs. Note that the lack of a significant alignment score does not necessarily mean the absence of an evolutionary relationship between two proteins. As we will see in Chapter 4, studying three-dimensional structures sometimes reveals evolutionary relatedness even when Sequence Homology is not detected.

To study evolutionary relationships, researchers strive to use protein families with similar functions from the widest possible range of organisms. The resulting information can be used to track the course of evolution of these organisms. Based on the analysis of divergence within selected protein families, a researcher can classify organisms into groups according to their evolutionary relationships. These data must be consistent with the results of classical PHYSIOLOGICAL AND BIOCHEMICAL studies of organisms.

Certain regions of a protein sequence may be found in organisms belonging to one taxonomic group but not to others. These regions can be used as marker sequences for the groups in which they are discovered. An example of such a marker sequence is a 12-amino-acid insertion in the N-terminal region of eEF-α1/EF-Tu proteins in all archaea and eukaryotes, but not in bacteria (Fig. 3-31). Marker sequences are one of the clues that help establish the evolutionary relationship between eukaryotes and archaea. Other characteristic sequences allow researchers to establish evolutionary connections between groups of organisms across many taxonomic levels.

Fig. 3-31. Marker sequence of the eEF-α1/EF-Tu protein family. The marker sequence (boxed) is an insertion of 12 amino acid residues located in the N-terminal region. Bases identical across all sequences are highlighted in yellow. The marker sequence exists in both archaea and eukaryotes, although these insertions differ quite significantly between the two groups of organisms. Changes within the marker sequence reflect the significant evolutionary divergence that occurred in this region after it first appeared in the common ancestor of both groups.

By analyzing complete amino acid sequences of proteins, one can construct more refined evolutionary trees containing many species within each taxonomic group. Figure 3-32 shows such a tree for bacteria, constructed based on sequence divergences of the GroEL protein (a protein present in all bacteria and essential for proper folding). The tree can be further refined by analyzing the sequences of many proteins and incorporating data on the unique biochemical and physiological properties of each species. There are numerous methods for constructing an evolutionary tree, each with its own Advantages and disadvantages, as well as many ways to represent evolutionary relationships. The branch tips in the tree shown in Fig. 3-32 correspond to extant species whose names are indicated. The nodes where two lines converge correspond to their extinct ancestor. In most methods of tree representation, including Fig. 3-32, the lengths of the lines between nodes are proportional to the number of amino acid substitutions distinguishing one species from another. Thus, if we consider the points corresponding to two extant species and their common ancestor, the length of each line connecting a branch tip to its branching point corresponds to the number of amino acid substitutions distinguishing each extant species from the ancestor. The sum of the line lengths extending from a branching point to the branch tips reflects the number of amino acid substitutions by which these extant organisms differ from one another. To determine the time intervals required for species divergence, the tree must be calibrated using fossil record data and other available information.

Fig. 3-32. Evolutionary tree constructed based on the Comparison of Amino acid sequences. This figure presents an evolutionary tree of bacteria based on sequence divergence of the GroEL protein family. In addition, the positions of Chloroplasts (chl.) of certain other organism species are indicated on the tree.

As the volume of information in databases grows, it becomes possible to construct evolutionary trees based on the sequences of an increasing number of diverse proteins. Thanks to increasingly sophisticated Analytical Methods, The amount of Genetic information continues to expand, allowing the structure of these trees to be further refined. All such efforts bring us closer to the goal of creating a detailed tree of life that describes the evolution and interrelationships of all organisms on Earth. Research in this field continues to advance (Fig. 3-33); the questions it raises are vital for humanity's understanding of itself and the surrounding world. Molecular evolution promises to be one of the most vibrant branches of science in the twenty-first century.

Fig. 3-33. Consensus tree of life. The tree shown here is based on the analysis of many protein sequences and additional genome features. Branches indicated by dashed lines are still being investigated. Such a tree represents only a fraction of the available information, and only some of the problems have found resolution in it. Every living group today has its own complex evolutionary history.

Summary of Section 3.4 Protein Structure: Primary Structure

■ The diverse Functions of Proteins are determined by their amino acid composition and sequence. Some minor changes in a protein's amino acid sequence may not alter its function.

■ Determining the amino acid sequence of a peptide occurs in several stages: 1) cleaving the polypeptide into smaller fragments using reagents that break specific peptide bonds; 2) determining the amino acid sequence of each fragment via the Edman degradation method; 3) determining the arrangement of the fragments within the original polypeptide chain by generating another set of overlapping fragments. Alternatively, the amino acid sequence can be deduced from the known nucleotide sequence of the corresponding gene.

■ Short PROTEINS AND PEPTIDES (up to 100 residues) can be synthesized chemically. At each stage of synthesis, a single amino acid is added while the entire peptide chain remains anchored to a solid support.

■ The amino acid sequence of a protein is a vital source of information not only regarding protein structure and function, but also about the evolution of life on our planet. Sophisticated methods have been developed to trace the course of evolution based on the analysis of amino acid sequence changes in homologous proteins.



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.