Protein Chemistry. Structure, Properties, Research Methods - Shendryk A.N. 2022

Protein Structure
Protein Structure
Chemical Fragmentation - Determination of the Amino Acid Sequence of Protein Molecule Fragments

Insulin was the first protein to have its Amino Acid Sequence determined. Devising the experimental strategy and carrying it out took the brilliant biochemist F. Sanger several years, and in 1951 he successfully deciphered the Introduction/19.html">Primary Structure of insulin. The enzyme consists of two chains (containing 21 and 30 amino acid residues, respectively) linked together by Disulfide Bonds.

In 1958, Sanger was awarded the Nobel Prize for his research into the primary structure of insulin. In 1980, he received a second Nobel Prize for establishing the nucleotide base sequence in polynucleotides.

Class="center">Image

SANGER, Frederick — born August 13, 1918.

Nobel Prize in Chemistry, 1958

Nobel Prize in Chemistry, 1980 (shared with Paul Berg and Walter Gilbert)

English biochemist Frederick Sanger was born in Rendcombe, Gloucestershire. His mother was the daughter of a prosperous textile magnate, and his father worked as a physician. From 1932 to 1936, the future scientist attended Bryanston School in Blandford, Dorset, and in 1936 he entered St John's College, Cambridge. Initially, Sanger intended to study medicine, but he soon became captivated by biochemistry. “It seemed to me,” he wrote many years later, “that this was the way to a real understanding of living matter and to providing a more scientific basis for solving many of the problems facing medicine.”

In 1939, Sanger earned his Bachelor of Arts degree from Cambridge University, followed by his Ph.D. in 1943. He joined a research group led by A.C. Chibnall, who had just succeeded Frederick Hopkins as Professor of Biochemistry at Cambridge. At the time, Chibnall was engaged in Protein Chemistry research.

In 1902, Emil Fischer hypothesized that Proteins are composed of Amino Acids linked together by peptide bonds. By the early 1940s, Fischer's hypothesis was widely, though not universally, accepted. Chibnall suggested that Sanger determine the terminal sequence of a peptide chain using chemical Methods.

In 1945, Sanger reported that under mild alkaline conditions, dinitrophenol could attach to the nitrogen atom of an amino acid via a bond stronger than a peptide bond. Consequently, a protein could be broken down into its constituent amino acids by cleaving the peptide bonds, and The amino acid structure could then be determined using Chromatography.

Much of the research in Chibnall's laboratory focused on insulin, one of the few proteins available at the time in pure form and large quantities. Sanger's initial study of insulin revealed that it contained two different N-terminal amino acids, meaning each insulin molecule consists of Two Types of polypeptide chains. Two Cysteine molecules can combine to form cystine through a disulfide bridge either between two polypeptide chains or between different regions of a single chain. In 1949, Sanger announced that he had discovered a way to break these disulfide bridges, thereby providing a method to separate the two chains.

Working with Hans Tuppy, a visiting researcher from Vienna, Sanger developed a plan to determine the amino acid sequence of each polypeptide chain in insulin using various Enzymes (proteases). In 1950, having established The structure of the longer of the two insulin chains, Tuppy left Cambridge. The shorter insulin chain proved much less amenable to chemical analysis, and its complete amino acid sequence was not fully determined until 1953. Sanger continued his work on mapping the disulfide bridges between the two chains, and in 1955 he presented the complete STRUCTURE OF THE insulin molecule. This was the first protein molecule to be studied in such detail. Sanger's work had profound implications for biochemistry and the burgeoning field of molecular biology. His findings conclusively proved that proteins are made of amino acids linked into chains by peptide bonds. Sanger also demonstrated that certain enzymes can cleave peptide chains at specific, predetermined sites. The application of this method subsequently helped biochemists determine the structures of many other proteins.

Image

In 1958, Sanger was awarded the Nobel Prize in Chemistry “for his work on the structure of proteins, especially that of insulin.” In his Nobel lecture, Sanger emphasized the great Practical significance of his work. “The Determination of the insulin structure certainly opens the way to the investigation of other proteins,” he noted. “One may also hope that The Study of proteins will help to reveal the changes that occur in the Organism during disease, and that our efforts may bring great practical benefit to humanity.”

Even before receiving the Nobel Prize, Sanger turned his attention to genetics, influenced in part by his friendship with Francis Crick. One of the most striking things to Sanger about The sequence of groups in insulin was the apparent absence of any obvious principle behind the unique arrangement of amino acids—yet vital physiological activity depended on this seemingly random order. Sanger could not understand how a protein could assemble in such a precise sequence, but it was clear that this order must have underlying origins. In the mid-1950s, Crick (who, along with James D. Watson, had first described the structure of deoxyribonucleic acid, or DNA) explained Sanger's discoveries through the “sequence hypothesis,” which posited that the information determining the amino acid sequence in a protein is carried by genes. It was later established that genes themselves consist of sequences of units, specific groups of which correspond to particular amino acids.

Nucleic Acids—DNA and ribonucleic acid (RNA)—are chains of linked NUCLEOTIDES. METABOLISM/28.html">The Genetic Code for amino acids embedded within them is determined by sequences of three bases. The protein-building process begins when the corresponding region of the DNA molecule, which contains the complete assembly instructions, “unzips” the bonds connecting the bases. Free nucleotides (floating randomly within The Cell) align themselves along the exposed DNA strand, forming a mirror-image chain known as Messenger RNA (mRNA). The completed mRNA strand detaches from the DNA (which then “zips up” again) and moves to cellular structures called Ribosomes, where protein assembly will take place. Portions of a shorter chain are formed by mRNA and then move away to pick up appropriate free nucleotides, which they subsequently bring back to the mRNA for incorporation into the Protein Structure. These short chains are called Transfer RNAs (tRNA). When Sanger began his study of nucleic acids, very little was known about these processes, and absolutely nothing was known about nucleotide sequences.

DNA and RNA sequences present much greater analytical challenges than protein sequences because they are much longer. A typical protein chain may contain up to fifty amino acids, whereas a typical mRNA contains hundreds of nucleotides. The DNA of even a tiny virus consists of thousands of nucleotides. Nevertheless, nucleic acid sequences are somewhat easier to decode than protein sequences due to one fundamental difference: while each position in a protein chain can be occupied by any of 20 different amino acids, there are only 4 “candidates” for each position in a DNA sequence—the nucleotides abbreviated as A, T, C, and G.

In 1958, Robert W. Holley set out to determine the sequence of a tRNA chain. Although these short chains are no longer than 100 nucleotides, The complexity of sequencing made the work drag on until 1965. Holley's work deeply impressed Sanger, but he sought a more powerful sequencing method applicable to mRNA chains, which frequently reach lengths of several hundred nucleotides. In the early 1960s, he and his colleagues developed such a technique. Using enzymes, they selectively cleaved mRNA chains into smaller fragments and determined the sequence of each. By piecing together how these fragments overlapped, they were able to deduce the sequence of the entire chain.

However, this approach required an immense amount of time and patience, and Sanger resolved to develop an analytical method for DNA Sequencing. He succeeded in 1973. His Procedure involved splitting a double-stranded DNA molecule into single strands and then dividing the material into four separate samples. Each sample is used as a template to rebuild the original double-stranded sequence. However, researchers halt the rebuilding process at specific nucleotides for each sample—either by limiting the concentration of a particular free nucleotide or by introducing a modified nucleotide into the chain that prevents further synthesis. As a result, the reconstructed chains are produced in various lengths, but every chain in a given sample ends with the exact same nucleotide. These four samples are then run simultaneously through an ultra-thin polyacrylamide gel, which separates the chains by length because shorter chains travel through the gel faster. The nucleotide sequence of the original DNA strand can then be read directly from the gel by comparing the bands produced by the various samples.

While Sanger and his colleagues were working on this method (known as the dideoxy method, named after the terminating chemical used), American scientists Walter Gilbert and Allan Maxam were developing an alternative sequencing procedure. In their method, DNA fragments of varying lengths are obtained by cleaving the chain at specific bases. This approach bore similarities to the method Sanger had previously used to sequence Protein and RNA chains. Both Sanger's and Gilbert's techniques became vital tools in Genetic Engineering, although Sanger's method proved somewhat more efficient for handling very long sequences. As early as 1978, Sanger and his colleagues demonstrated the power of the dideoxy method by sequencing 5,375 bases in the DNA chain of a bacterial virus. This was the first time a DNA chain of such length had been deciphered in such detail.

In 1980, Sanger and Gilbert were awarded half of the Nobel Prize in Chemistry “for their contributions concerning the determination of base sequences in nucleic acids.” The other half of the prize was awarded to Paul Berg. These three scientists, as B.G. Mlström stated in his presentation speech on behalf of the Royal Swedish Academy of Sciences, “have made it possible to penetrate to an even greater depth in our understanding of the interplay between chemical Structure and Biochemical function of genetic material.”

In 1983, Sanger retired from his position at the Medical Research Council. A modest and private man, he lives in Cambridge with his wife, Margaret Joan Howe, whom he married in 1940. The couple has two sons and a daughter.

Sanger has received numerous awards and honors, including the Corday-Morgan Medal and Prize from the Chemical Society (1951), the Alfred Benzon Prize from the Alfred Benzon Foundation (1966), the Royal Medal of the Royal Society (1969), the Gairdner Foundation Annual Award (1971 and 1979), the Hanbury Memorial Medal of the Pharmaceutical Society of Great Britain (1976), the Copley Medal of the Royal Society (1977), and the Albert Lasker Basic Medical Research Award (1979). Sanger is an honorary member of the American Society of Biological Chemists and the U.S. National Academy of Sciences, and holds honorary degrees from the universities of Leicester, Strasbourg, Cambridge, and Oxford.

Source of information: Nobel Laureates: An Encyclopedia: Trans. from English.—M.: Progress, 1992.

To date, the number of proteins whose amino acid sequence has been mapped runs into the thousands, and this number is growing rapidly. One of the most remarkable achievements in this field was determining the amino acid sequence of a human gamma-globulin antibody molecule, which comprises 1,320 amino acids. Currently, all accumulated data on protein spatial structures are entered into a databank (available on the Internet at http://nist.rcsb.org/pdb/). As more information is gathered, this databank is expected to be used in the future to predict protein shapes and Functions directly from their Amino acid sequences.

The Methods for determining the amino acid sequence of Peptides are based on the Edman Degradation (1950–1956), which involves the stepwise Cleavage of the polypeptide chain from the N-terminus using PITC.

Each cycle of degradation consists of three stages.

1. Formation of the phenylthiocarbamoyl peptide (PTC-peptide)

Image

2. Cleavage of the N-terminal amino acid as an anilinothiazolinone derivative:

Image

3. Isomerization of the thiazolinone into a 3-phenyl-2-thiohydantoin (PTH) derivative followed by its identification. This step proceeds via The formation of a phenylthiocarbamoyl amino acid (PTH-amino acid):

Image

The First stage involves the attachment of PITC to the unprotonated α-amino group of the terminal amino acid (pH 9–9.5, using highly volatile buffer systems). Tertiary or heterocyclic amines (such as triethylamine, dimethylallylamine, or pyridine) can be used as bases.

Side Reactions.

> Oxidative desulfurization of the PTC group caused by atmospheric oxygen (therefore, all steps are carried out under an inert gas atmosphere).

> Hydrolysis of FITC yielding diphenylthiourea and aniline, along with side products that complicate the identification of amino acid phenylthiohydantoins. Therefore, upon completion of the first stage, the residue of FITC and its by-products is extracted with benzene.

The Second Stage involves the cleavage of thiazolinone. This reaction proceeds readily and, as a rule, without the formation of side products. An exception is that C-terminal glutamine residues can convert into pyroglutamic acid residues, thereby blocking the further degradation of the peptide chain:

Image

The second stage is carried out in the presence of anhydrous trifluoroacetic acid.

The Third Stage is isomerization, which proceeds in two steps. First, rapid hydrolysis takes place, converting thiazolinone into the PTC-amino acid, followed by cyclization into phenylthiohydantoin (PTH). This reaction occurs under the following conditions: 1 N HCl, 80∘C, 10 min. Some PTH derivatives undergo partial degradation under these conditions; for instance, the PTH derivatives of Serine and Threonine are subject to dehydration.

In the second step, butyl acetate is added, and the phases are separated. The organic phase contains the PTH, while the aqueous phase contains the residual peptide with a chain length shorter by one residue. The aqueous phase is separated and directed to the next degradation cycle.

Identification of the PTH is performed in the organic phase. For this purpose, paper or Thin-Layer Chromatography, liquid chromatography, Gas-Liquid Chromatography, and other physicochemical analysis methods are employed. Alternatively, the PTH can be hydrolyzed using HI back to the free amino acid for subsequent identification.

In 1967, Edman and Begg developed an instrument (a Sequencer) for the automated determination of the amino acid sequence of short peptides and their fragments. In this sequencer, all reactions are carried out in a cylindrical Glass vessel rotating around its own axis at a constant speed in an inert gas atmosphere (see fig.).

Image

The protein sample is applied as a thin film onto the walls of the vessel, whereas Reagents in solution are delivered to the bottom. Due to centrifugal force, the rotating vessel causes the solution to creep up its walls, washing over the protein and collecting in a special groove. From this groove, dissolved products are removed from the vessel via a dedicated discharge tube. This reaction setup provides a large surface area of contact between the protein and the reagents, ensuring a high reaction rate and the continuous Nature of the sequencing process.

Currently, instrument-making industries in A number of developed countries manufacture automated sequencers based on the rotating-cup principle proposed by Edman and Begg. The Separation and Analysis of the resulting phenylthiohydantoins in these instruments are performed using High-Performance Liquid Chromatography (HPLC). The appearance of one such sequencer is shown in the figure below:

Image



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.