Biological Chemistry - Berezov T. T., Korovkin B. F. 1998
Protein Chemistry
Structural Organization of Proteins
Methods for Determining the C-Terminal Amino Acid
Enzymatic Methods are frequently employed to determine The Nature of the C-terminal amino acid. Treating a polypeptide with carboxypeptidase, which cleaves the peptide bond at the end containing the free COOH group, releases the C-terminal amino acid, whose identity can then be determined chromatographically.
The chemical method proposed by S. Akabori, which is based on the Hydrazinolysis of the polypeptide, has also been introduced:
Class="center">
Hydrazine induces the Cleavage of peptide bonds sensitive to it and reacts with all Amino Acids except the C-terminal amino acid, as the latter's carboxyl group does not participate in peptide bond formation. This yields a mixture of aminoacyl hydrazides and the free C-terminal amino acid. After treating the entire mixture with DNFB, the free amino acid is separated and identified chromatographically, for which the resulting dinitrophenyl derivatives of aminoacyl hydrazides are first extracted with ethyl acetate.
The C-terminal amino acid can also be identified by treating the polypeptide with a reducing agent, such as sodium borohydride. In its simplest form, this Procedure can be represented as follows:

It is evident that under these conditions, only a single amino acid—specifically the C-terminal one—is converted into an α-amino alcohol, which is readily identified by Chromatography. Thus, the nature of both N- and C-terminal amino acids can be determined using these methods.
The next stage of the work involves determining the order (sequence) of amino acids within the polypeptide chain. To achieve this, selective partial (chemical and enzymatic) Hydrolysis of the polypeptide chain is first carried out to yield shorter peptide fragments, whose Amino Acid Sequence can be precisely determined using the methods described above.
Chemical methods of selective and partial hydrolysis rely on Reagents that induce the selective, highly Specific Cleavage of peptide bonds formed by specific amino acids while leaving other peptide bonds intact. These selective hydrolytic agents include Cyanogen bromide, CNBr (targeting Methionine residues), hydroxylamine (targeting bonds between aspartic acid and Glycine residues), and N-Bromosuccinimide (targeting Tryptophan residues). Because Proteins generally contain fewer methionine residues than Other Amino Acids, Treatment with CNBr is preferred; it yields a small number of Peptides whose Primary Structure is determined using the previously discussed methods, each time starting with the identification of the N- and C-terminal amino acids.
Enzymatic hydrolysis methods are based on the specific action of proteolytic (protein-degrading) Enzymes that cleave peptide bonds formed by particular amino acids. Specifically, Pepsin accelerates the hydrolysis of bonds formed by phenylalanine, Tyrosine, and glutamic acid residues; Trypsin targets Arginine and Lysine; and Chymotrypsin cleaves at tryptophan, tyrosine, and phenylalanine. Several Other Enzymes, such as Papain, subtilisin, pronase, and other bacterial proteinases, are also used for the partial Hydrolysis of Proteins. As a result, the polypeptide chain is cleaved into smaller peptides, sometimes consisting of only a few amino acids, which are separated from one another using combined electrophoretic and chromatographic techniques to produce characteristic peptide maps. Next, The amino acid sequence within each individual peptide is determined. The procedure concludes by reconstructing the Introduction/19.html">Primary structure of the entire polypeptide chain based on the established Amino acid sequences of the individual peptides.
The technique of peptide mapping, aptly dubbed the "fingerprint method," is used to determine the similarities or differences in the primary structure of homologous proteins. The protein is incubated with a proteolytic enzyme. Often, portions of the protein are incubated with both pepsin and trypsin. Due to the hydrolysis of strictly specific peptide bonds, this produces a mixture of short peptides that are easily separated by chromatography in one dimension and Electrophoresis In the second, at a 90° angle to the first (peptide map).
Subsequent tasks involve establishing the amino acid sequence within each isolated peptide (using the phenylthiohydantoin or other methods), comparing the obtained data, and determining the primary STRUCTURE OF THE entire molecule.
The feasibility of applying X-ray crystallography to determine the amino acid sequence in a protein molecule was discussed earlier. It is worth noting a completely novel approach to solving this crucial problem: determining the amino acid sequence of a protein molecule using data on the complementary nucleotide sequence of DNA. This approach is facilitated by both rapid DNA Sequencing Methods and the techniques for isolating and accessing the Gene itself*.
Today, elucidating the Primary Structure of Proteins is merely a matter of time and laboratory equipment. The primary structure of many natural proteins has been fully resolved, most notably Insulin, which contains 51 amino acid residues [Sanger F., 1954]. A larger protein whose primary structure was successfully determined is immunoglobulin, whose four polypeptide chains comprise 1,300 amino acid residues. For this work, J. Edelman and R. Porter were awarded the Nobel Prize (1972).
* The primary structure of numerous Proteins and Enzymes has been described, notably the replicase of phage MS2 (544 amino acid residues), established using this methodological approach.

Fig. 1.14. Structure of proinsulin.
The primary structures have been deciphered for human Myoglobin (153 amino acid residues), the α-chain (141) and β-chain (146) of human Hemoglobin, human Heart Muscle cytochrome c (104), human milk Lysozyme (130), bovine chymotrypsinogen (245), and many other proteins, including enzymes and toxins. Fig. 1.14 illustrates the amino acid sequence of proinsulin. It can be seen that the insulin molecule (highlighted by dark circles), consisting of two chains (A with 21 and B with 30 amino acid residues), is formed from its precursor, proinsulin (84 amino acid residues, represented by a single polypeptide chain), following the cleavage of a 33-amino-acid peptide. The structure of the insulin molecule (51 amino acid residues) can be schematically represented as follows:

Disulfide (—S—S—) bonds are formed between the A and B chains and within the A chain of insulin. The primary structure of more than 18 insulins isolated from various sources has been elucidated. Insulins from the Pancreas of humans, pigs, and sperm whales have been found to be structurally similar. The sole difference in human Insulin is the presence of Threonine at position 30 of the B chain instead of Alanine.
The second protein whose primary structure was deciphered by C. Moore and W. Stein is pancreatic Ribonuclease (Fig. 1.15), which catalyzes the cleavage of RNA. The enzyme consists of 124 amino acid residues with an N-terminal lysine and a C-terminal valine, and disulfide (—S—S—) bonds are formed between Cysteine residues in 4 regions.
The amino acid sequence of the polypeptide chain of lysozyme has been fully elucidated. This enzyme holds significant protective and medical importance because it lyses certain Bacteria by hydrolyzing the core substance of their Cell wall. Hen egg-white lysozyme contains 129 amino acids (Fig. 1.16) with an N-terminal lysine and a C-terminal leucine.
Soviet researchers have established the primary structure of numerous proteins and Polypeptides, including the large RNA polymerase protein (specifically, the sequences of its β and β' subunits, comprising 1,342 and 1,407 amino acid residues, respectively*), elongation factor G from E. coli (701 amino acids) (Y.A. Ovchinnikov et al.), the enzyme aspartate aminotransferase, consisting of 412 amino acid residues (A.E. Braunstein, Y.A. Ovchinnikov et al.), leghemoglobin, the L25 protein from E. coli Ribosomes, neurotoxins from cobra venom (Y.A. Ovchinnikov et al.), pepsinogen and pepsin (V.M. Stepanov et al.), L-lipotropin and bovine lactogenic hormone (N.A. Yudaev, Y.A. Pankov), and others.
* It should be noted that In addition to the aforementioned β and β' subunits, the DNA-dependent RNA polymerase molecule also contains α and α' subunits as well as the σ factor, whose primary structures have not yet been deciphered, although the total number of amino acids (~5,000) making up this giant protein is known.

Fig. 1.15. Primary structure of RNase. Four Disulfide Bonds are highlighted in color.

Fig. 1.16. Primary structure of the lysozyme polypeptide chain (diagram).
Investigations into the primary structure of hemoglobin α- and β-chains have contributed significantly to elucidating the structure of unusual, so-called abnormal Hemoglobins found in the Blood of patients with hemoglobinopathies. In some cases, the progression of the disease, as well as the alteration of the three-dimensional structure of human hemoglobin, is caused by the substitution of just a single amino acid within the β-chain (less frequently, the α-chain) of hemoglobin (see Chapter 2).
An analysis of data on protein primary structure allows the following general Conclusions to be drawn.
1. The primary structure of proteins is unique and genetically determined. Each individual homogeneous protein is characterized by a unique amino acid sequence: a frequency of amino acid substitution leads not only to structural rearrangements, but also to alterations in physicochemical properties and biological Functions.
2. The Stability of the primary structure is maintained primarily by covalent peptide bonds, with the possible involvement of a small number of disulfide bonds.
3. A wide variety of amino acid combinations can be found within the polypeptide chain; repeating sequences are relatively rare in polypeptides.
4. Certain enzymes with similar catalytic properties share identical peptide structures containing invariant regions alongside variable amino acid sequences, particularly within their active sites. This principle of structural similarity is most typical of A number of Proteolytic Enzymes, including trypsin and chymotrypsin, among others (see Chapter 4).
5. The primary structure of a polypeptide chain determines its secondary, tertiary, and quaternary structures, which together dictate the overall spatial conformation of the protein molecule.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.