Principles of Biochemistry Volume 1 - A. Lehninger 1985
Biomolecules
Proteins: Covalent Structure and Biological Functions
Determination of the Amino Acid Sequence of Polypeptide Chains
Two major discoveries made in 1953 marked the beginning of a new era in biochemistry. That year, James D. Watson and Francis Crick in Cambridge, England, proposed the structural model of DNA (The Double Helix) and suggested a structural basis for its precise Replication. This proposal essentially (though implicitly) put forward the idea that The sequence of nucleotide units in DNA encodes Genetic information. In the same year, Frederick Sanger, working in the same Cambridge laboratory, deciphered the Amino Acid Sequence of The polypeptide chains of the hormone Insulin. This achievement was of immense significance in its own right, as it had long been believed that determining The amino acid sequence of a polypeptide was an utterly hopeless and insurmountable task. Moreover, Sanger's findings, appearing almost simultaneously with the Watson-Crick hypothesis, strongly hinted at a connection between The nucleotide sequence of DNA and the amino acid sequence of Proteins. Over the next decade, this concept led to the deciphering of all nucleotide codons in DNA and RNA that uniquely determine the Amino acid sequences of protein molecules.
Prior to Sanger's work, which took several years to complete, there was no certainty that all molecules of a given protein were strictly identical in molecular weight and Amino Acid Composition. Today, the amino acid sequences of many hundreds of proteins isolated from various sources are known. The determination of a polypeptide chain's amino acid sequence relies on principles first developed by Sanger, which are still in use today, albeit with various modifications and refinements. Deciphering the amino acid sequence of any polypeptide requires six main stages.
a. Stage 1: Determination of Amino Acid Composition
The first step toward determining the amino acid sequence is the complete Hydrolysis of all peptide bonds in a pure polypeptide. The resulting mixture of Amino Acids is then analyzed using Ion-exchange Chromatography (Sec. 5.18), which reveals which Amino acids are present in the hydrolysate and in what proportions.
б. Stage 2: Identification of Amino- and Carboxyl-Terminal Residues
The next step is to identify the amino acid residue located at the end of the polypeptide chain that bears a free $\alpha$-amino group, i.e., the amino terminus ($\text{NH}_2$-terminus, or N-terminus). For this purpose, Sanger proposed using 1-fluoro-2,4-dinitrobenzene (Sec. 5.22) as a labeling reagent, which attaches to the amino-terminal (N-terminal) residue of the chain, forming a yellow-colored 2,4-dinitrophenyl (DNP) derivative of the polypeptide. Upon acid hydrolysis, all peptide bonds within this DNP-polypeptide derivative are cleaved, yet the covalent bond between the 2,4-dinitrophenyl group and the $\alpha$-amino group of the N-terminal residue remains intact. Consequently, the N-terminal residue appears in the hydrolysate as a 2,4-dinitrophenyl derivative (Fig. 6-6). This derivative is readily separated from unsubstituted free Amino Acids and identified chromatographically by comparison with authentic dinitrophenyl derivatives of various amino acids.
Class="center">
Fig. 6-6. Identification of the amino-terminal residue of a tetrapeptide via The formation of a 2,4-dinitrophenyl derivative. The tetrapeptide reacts with 1-fluoro-2,4-dinitrobenzene (FDNB) to form a 2,4-dinitrophenyl derivative. The latter is then subjected to boiling in the presence of 6 N HCl to cleave all peptide bonds, leaving the amino-terminal amino acid in the form of a 2,4-dinitrophenyl derivative.

Fig. 6-7. Labeling the amino-terminal residue of a tripeptide using dansyl chloride. Following the hydrolytic Cleavage of all peptide bonds, the dansyl derivative of the amino-terminal amino acid can be isolated and identified. Due to the intense fluorescence of dansyl groups, they can be detected in significantly smaller amounts than dinitrophenyl groups, making the dansyl method far more sensitive than the fluorodinitrobenzene Procedure.
Another labeling reagent used to identify the N-terminal residue is dansyl chloride (Fig. 6-7), which reacts with the free $\alpha$-amino group to yield a dansyl derivative. Because this derivative fluoresces intensely, it can be detected and quantified at much lower concentrations than dinitrophenyl derivatives.
The carboxyl-terminal (C-terminal) amino acid residue of a polypeptide chain can also be identified using Structure/129.html">Specific Methods. One such approach involves incubating the polypeptide with the enzyme carboxypeptidase, which selectively hydrolyzes the peptide bond located at the carboxyl end ($\text{COOH}$-terminus, or C-terminus) of the chain. By determining which amino acid is released first from the polypeptide upon Treatment with carboxypeptidase, the C-terminal residue can be identified.
By identifying the N- and C-terminal residues of the polypeptide, we obtain two crucial reference points for determining its amino acid sequence.
c. Stage 3: Cleavage of the Polypeptide Chain into Fragments
Now we take another portion of the test sample containing intact polypeptide chains and cleave them into smaller pieces—short Peptides consisting on average of 10–15 amino acid residues. The purpose of this procedure is to separate the resulting fragments and determine the amino acid sequence within each of them.
Several methods can be used to cleave a polypeptide chain into individual fragments. One widely used method is partial Enzymatic hydrolysis of the polypeptide using the digestive enzyme Trypsin. The catalytic action of this enzyme is highly specific: it hydrolyzes only those peptide bonds involving the carboxyl group of a Lysine or Arginine residue, regardless of the length and amino acid sequence of the polypeptide chain (Table 6-6). The number of smaller peptides produced by trypsin can therefore be predicted from the total number of lysine and arginine residues in the original polypeptide. A polypeptide containing five lysine and/or arginine residues will typically yield six smaller peptides upon trypsin Digestion, and all of these peptides, except one, will have a lysine or arginine residue at their carboxyl terminus. The fragments generated by trypsin are separated either by Column ion-exchange chromatography or by paper Electrophoresis and chromatography; often, two-dimensional chromatographic Separation of the peptides is performed on a sheet of paper, resulting in a peptide map with the peptides distributed across it (Fig. 6-8).
Table 6-6. Specificity of Four Important Methods for Polypeptide Chain Fragmentation
|
Treatment |
Cleavage Sites (residues contributing the carbonyl group of the cleaved peptide bond) |
|
Trypsin |
Lysine Arginine |
|
Phenylalanine |
|
|
Phenylalanine Tryptophan Tyrosine Certain other residues |
|

Fig. 6-8. Peptide map obtained after trypsin digestion of normal human Hemoglobin. Each spot contains one of the peptides. To obtain such a two-dimensional map, the peptide mixture is applied to a square sheet of paper, and electrophoresis is performed in one direction parallel to one side of the square. The paper is then dried, and chromatographic separation of the peptides is carried out in a second direction perpendicular to the first. Neither of these two processes alone can completely resolve the peptides, but their sequential application proves to be a very efficient method for separating complex peptide mixtures.
d. Stage 4: Determination of the Sequence of Peptide Fragments
At this stage, the amino acid sequence of each peptide fragment obtained in Stage 3 is established. This is typically accomplished using the chemical method developed by Pehr Edman. Edman Degradation relies on labeling and cleaving only the N-terminal residue of the peptide while leaving all other peptide bonds intact (Fig. 6-9). Following the identification of the cleaved N-terminal residue, a label is introduced into the next residue—which has now become the N-terminal residue—and it is cleaved in the exact same manner through the same series of reactions. By successively removing residues in this fashion, the entire amino acid sequence of the peptide can be determined using a single sample. Figure 6-9 illustrates how Edman degradation is carried out. First, the peptide reacts with phenylisothiocyanate, which attaches to the free α-amino group of the N-terminal residue. Treatment of the peptide with cold dilute acid causes the release of the N-terminal residue as a phenylthiohydantoin derivative, which can be identified by chromatographic methods. The rest of the polypeptide chain remains undamaged after the removal of the N-terminal residue. The shortened peptide is then subjected to the same series of reactions again, allowing the identification of the new N-terminal residue. By repeating this sequential removal of N-terminal residues, the amino acid sequence of peptides consisting of 10-20 residues can be easily determined.
The amino acid sequence is determined for all the peptides generated by trypsin. This immediately raises a new problem: determining the original order of the tryptic fragments within the initial polypeptide chain.
e. Stage 5: Cleavage of the Original Polypeptide Chain by Another Method
To establish the order of the peptide fragments produced by trypsin, a fresh portion of the original polypeptide preparation is taken and cleaved into smaller fragments by an alternative method that targets peptide bonds resistant to trypsin. In this case, a chemical rather than an enzymatic method is often preferred. Particularly good results are obtained with the reagent cyanogen bromide, which cleaves only those peptide bonds where the carbonyl group belongs to a methionine residue (Table 6-6). Consequently, if a polypeptide contains eight methionine residues, treatment with cyanogen bromide will typically yield nine peptide fragments. The fragments obtained in this manner can be separated by electrophoresis or chromatography. Each of these short peptides is then subjected to Edman degradation, as described for Stage 4, thereby establishing their amino acid sequences.

Fig. 6-9. Scheme for determining the amino acid sequence of a peptide by Edman degradation. The initial tetrapeptide reacts with phenyl isothiocyanate to yield a phenylthiocarbamoyl derivative of the amino-terminal residue. This residue is cleaved from the peptide without breaking other peptide bonds and is obtained as a phenylthiohydantoin derivative, which can be identified by chromatography. The remaining tripeptide is subjected to the same cycle of reactions again, allowing the identification of the second residue. These operations are repeated until all residues have been identified.
Thus, we have obtained two sets of peptide fragments: one after treating the original polypeptide with trypsin, and another after chemical cleavage of the same polypeptide with cyanogen bromide. We also know the amino acid sequence of each peptide belonging to these two sets.

Fig. 6-10. Ordering of peptide fragments based on overlapping segments. In the example shown here, a polypeptide consisting of 16 amino acid residues was fragmented in two different ways after identification of its N- and C-terminal residues. The top portion shows the resulting fragments, and the bottom portion illustrates the reconstruction of the complete polypeptide sequence using overlapping regions.
e. Step 6: Ordering the peptide fragments using overlapping segments
Next, the amino acid sequences of the peptide fragments obtained by the two different cleavage methods are compared to find peptides In the second set whose sequences overlap (i.e., match) with portions of the peptides in the first set. THE PRINCIPLE OF peptide alignment is illustrated in Fig. 6-10. The peptides from the second set with overlapping sequences allow the peptide fragments produced by the first cleavage of the original polypeptide chain to be linked in the correct order. Furthermore, these two sets of fragments help to reveal any potential errors in determining the amino acid sequence of individual fragments.
Sometimes, a second fragmentation of the polypeptide is not sufficient to find overlapping regions for two or more peptides generated by the first cleavage. In such cases, a third or even a fourth cleavage method is employed, ultimately yielding a set of peptides that provide all the overlaps necessary to establish the complete sequence of the original chain. For this purpose, other Proteolytic Enzymes, such as chymotrypsin or pepsin, can be used; however, these enzymes cleave peptide bonds much less selectively than trypsin (Table 6-6).
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.