Protein Chemistry. Structure, Properties, Research Methods - Shendryk A.N. 2022
Methods for Experimental Investigation of Protein Structure
X-ray Diffraction Analysis
Specific Features of Protein XRD
Determining a Structure by X-Ray Diffraction analysis (XRD) means reconstructing the original crystal lattice—which produced the diffraction pattern in reciprocal space—back into the real space of the crystal itself. Solving this problem has always posed serious challenges, even for simple molecules. When moving on to biological objects, entirely unconventional or highly specific XRD techniques were required.
The key contribution made by the Cambridge school (Perutz, Kendrew) to the XRD Analysis of Proteins was The Development of the experimental method for phase determination. How was this achieved?
Protein macromolecules form a molecular lattice in which they are packed closely together. The model (see Fig. 4.20) illustrates the packing of Myoglobin molecules within the lattice. The crystal lattice is monoclinic, with two macromolecules per unit Cell. While the intramolecular structure of myoglobin itself is exceptionally complex, the way entire macromolecules pack within the lattice is simple. Gaps inevitably remain between macromolecules of such intricate shapes, and these gaps are filled with an aqueous solution during protein crystallization. In protein crystals, Water of crystallization often accounts for half, or even a greater part, of the total volume. Drying a protein crystal typically disrupts its structural regularity. It is worth recalling that for a long time, protein crystals were not even considered true crystals because they did not yield X-ray diffraction. Bernal and Dorothy Hodgkin demonstrated that the error lay in drying the crystals; by capturing X-ray diffraction patterns of protein crystals kept in their mother liquor, researchers managed to obtain astonishingly detailed patterns, often comprising over 20,000 independent reflections (see Fig. 4.21).
Class="center">
Fig. 4.21 X-ray diffraction pattern of a Hemoglobin crystal
Let us now assume that we introduce a heavy atom with a high atomic number into each protein macromolecule (specifically considering hemoglobin for clarity), and that we do so systematically—that is, via a chemical reaction with a strictly defined functional group. Another essential requirement is that the resulting protein–heavy-atom conjugate must be isomorphous with the native protein, meaning it must crystallize in the same lattice. When these conditions are met, the heavy atoms primarily alter the intensities of individual spots on the X-ray pattern without fundamentally changing the diffraction pattern as a whole. The simplest Examples of heavy atoms well-suited for this purpose are mercury atoms, which exhibit a specific affinity for sulfhydryl groups. By introducing the specific mercurial reagent sodium p-chloromercuribenzoate into the mother liquor used for hemoglobin crystallization, one can obtain a protein in which two sulfhydryl groups per macromolecule have reacted with the mercury compound.
![]()
Thus, exactly four strongly X-ray-scattering mercury atoms will be located at specific points within each unit cell of the crystal. By capturing the diffraction pattern of the mercury derivative crystal and comparing it with corresponding data for the unsubstituted protein, we obtain an X-ray diffraction pattern corresponding to a spatial lattice formed, as it were, solely by mercury atoms positioned at the sites where they attach via chemical interaction with the protein. Naturally, this mercury atom lattice has a simple structure (containing a total of only 4 mercury atoms per unit cell).
Determining the diffraction pattern for the difference "mercury" lattice is by no means straightforward. It cannot be done simply by subtracting the spot intensities of the combined protein-mercury diffraction pattern from those of the pure protein. What adds up are not the intensities (the squares of the amplitude coefficients), but the amplitude coefficients themselves. As mentioned above, the latter are defined by both magnitude and phase. These are complex numbers that add up as vectors in a plane. Consequently, solving the phase problem boils down to simultaneously solving vector equations obtained from several different protein–heavy-atom derivatives. Double derivatives containing two different heavy atoms at two points in the macromolecule are of great assistance here. Furthermore, the presence of a 2nd-order crystal Symmetry axis greatly simplifies structural analysis.
Without delving into the details, we should note that although the unit cell of the "mercury" lattice remains monoclinic, it is nonetheless primitive because it contains only a few atoms per cell. Therefore, its complete X-Ray Structural Analysis is relatively straightforward and free of fundamental difficulties. This makes it possible to retrospectively calculate The values of all amplitude coefficients in full—that is, determining both amplitudes and phases. Knowledge of the entire spatial picture in both direct and reciprocal space for the "mercury" lattice is then used to solve the main problem: finding the phase values of the amplitude coefficients for the protein crystal.
For self-validation, crystallographers prefer to have not just 2, but 4 to 5 or even more isomorphous heavy-atom-substituted protein derivatives. In this case, the reliability and accuracy of the phase determination for each diffraction spot increase significantly, and any random error is readily exposed because it leads to data inconsistency.
From all of the above, one can only roughly imagine the colossal scope of measurements and calculations involved. However, once the phases are determined and Fourier series are summed up, we have everything we need: the values of the electron density $\rho(x,y,z)$ at every point in space within the unit cell. The points where the electron density passes through a maximum give us the coordinates of the atoms.
To construct a spatial molecular model using electron density, it is reconstructed layer by layer. This process is much like creating histological sections of complex tissue: by stacking successive slices on top of one another, we reconstruct the three-dimensional STRUCTURE OF THE tissue. In a similar way, the three-dimensional structure of a molecule is reconstructed from planar sections of equal electron density surfaces (see Fig. 4.22).

Fig. 4.22 Spatial distribution of electron density in myoglobin.
To illustrate METABOLISM/2.html">THE CONCEPT OF isomorphous replacement for solving the phase problem in X-ray structure analysis of complex molecules, Bragg proposed a simple yet highly original experiment. Instead of X-ray diffraction, Bragg suggested observing the diffraction of visible light using a planar grating rather than a spatial lattice. To achieve visible light diffraction, a simple apparatus shown in the figure is used. Light from source A is collected by a condenser at the principal focus of a plano-convex lens. Parallel rays then pass through mask O, which is a piece of black paper with punched holes.


Fig. 4.23 A mask modeling one (A), two (B), and four hexamethylbenzene molecules (left part of the figure). The same masks with one additional hole in the center (right part of the figure)
The arrangement of the holes corresponds to the positions of the atoms in the molecule under study. For example, Fig. 4.23 shows a mask representing The structure of hexamethylbenzene (recall that hydrogen atoms can be disregarded in X-ray structure analysis). After the second lens, the rays are focused in plane F, where the diffraction pattern is formed. It is viewed through a low-power Microscope. That is the entire setup for the visual observation of light diffraction.
Fig. 4.23 (left) shows diffraction patterns obtained from masks modeling one, two, and four regularly arranged hexamethylbenzene molecules on a plane. It can be seen that even a single molecule produces a system of diffraction spots, the overall arrangement of which is preserved for a lattice of multiple molecules. However, each spot splits into several, ultimately resulting in a complex pattern where intermolecular and intramolecular interferences are mixed.
The question arises whether, knowing the diffraction pattern, one can calculate or measure the original pattern of holes or illuminated points that produced it. It turns out that this would be possible and easy if we knew not only the amplitudes (their square being the spot intensity), but also the Phases of the oscillations in the diffraction spots. In this case, we are dealing with a direct-space pattern much more elementary than a protein crystal: it is planar and possesses a center of symmetry. Furthermore, all Interference phases take only two values: 0 and п. In other words, the phase ambiguity of the spots reduces to the ambiguity of the signs of the amplitude coefficients, i.e., "+" or "-". If we knew the signs corresponding to each spot—that is, to each point in reciprocal space—it would not be difficult to reconstruct (by calculating via Fourier series) the structure of the “hexamethylbenzene unit cell” based on the diffraction pattern. Since the signs of the amplitude coefficients are unknown, Fourier series cannot be calculated. Crucially, however, one can bypass calculations by using an analogue approach, applying the same visible light diffraction technique to transform the spot pattern from reciprocal space into direct space. To do this, one must prepare a mask with holes corresponding to the diffraction pattern (Fig. 4.23, left). But this is not enough. Some of the holes must emit a light beam with an opposite phase of oscillation relative to the rest, and the locations of these holes must be known. Knowing these apertures, we cover them with a transparent mica plate that introduces a phase delay of exactly п (the thickness of the plate is chosen to be equal to half the wavelength of light in the given medium). As a result, we obtain the pattern of atomic positions in the hexamethylbenzene "molecules" at the focal plane of our optical system.
Thus, the inverse problem of transitioning from the diffraction pattern to the “lattice structure” can be solved by analogue Methods only if the phases (in this case, the signs) of the interference maxima are known. Since the phases are not known in advance (in practice, we only see and measure spot intensities), we must resort to the method of isomorphous replacement. To do this, we introduce an extra illuminated point into the structure of the “molecule” at a specific Location. For simplicity, in our example, the extra hole in the mask is placed right in the center (Fig. 4.23, right). We obtain the diffraction patterns from the modified “molecules” and compare them with those from the original “molecules”. Since the situation here is much simpler and requires a choice between only two phases—0 and п—a qualitative comparison of the two patterns is quite sufficient and immediately gives a complete answer to the question posed. Indeed, the additional hole is located at the center of symmetry of the “molecule”. Analysis shows that all amplitude coefficients produced by these central holes will have a positive sign. Therefore, if a diffraction spot increases in intensity after adding the central “atom” to the hexamethylbenzene “molecule”, its amplitude coefficient is positive; if it decreases, it is negative.
By comparing the left and right sides of Fig. 4.23, we easily identify the spots with positive and negative signs of the amplitude coefficients. This—albeit with significant simplifications—captures the core physical idea of the isomorphous replacement method. When analyzing protein crystals, the situation is incomparably more complex for the following reasons.
Firstly, the lattice is three-dimensional; secondly, there is no center of symmetry (proteins cannot have a center of symmetry if only because they contain asymmetric carbon atoms). In principle, however, Bragg’s illustration helps to grasp the very Concept of the method.
Returning to the X-ray structural analysis of proteins, let us look at the journey of myoglobin research via X-ray crystallography, which is now of purely historical interest. This work featured several successive stages. The resolution of X-ray structural analysis is, to a certain extent, up to the researcher. Initially, it was decided to limit the study to a coarser picture, neglecting finer details. It was assumed that the analysis would be conducted at a 6 Å resolution. This meant that interferences corresponding to distances in the direct lattice of less than 6 Å were disregarded. In this case, all spots whose distances from the center exceeded a certain threshold were discarded in reciprocal space. In other words, choosing the resolving power allows us to select diffraction maxima sufficiently close to the central spot, thereby limiting the experimental material used. For instance, obtaining a 3D model of myoglobin at 6 Å resolution required using 400 diffraction spots. When moving to a 2 Å resolution, the volume of reciprocal space that must be used for analysis increases by a factor of 33 = 27. Consequently, the number of diffraction spots used for the calculations reached 10,000. When the resolution was increased to 1.5 Å, the volume of reciprocal space and, accordingly, the number of diffraction maxima doubled once again, meaning that 20,000 interference spots were used for the analysis.
The image obtained at a 6 Å resolution did not reveal the positions of individual atoms, but the positions of the Pauling–Corey helices formed by the polypeptide chain were clearly established, given the helix diameter of 10.1 Å. The amino acid side chains appeared here as a featureless amorphous mass filling the spaces between the rod-like helices. Because individual protein groups could not be identified, nothing could be said about the Amino Acid Sequence within the chain. The next step was to increase the resolving power to 2 Å by analyzing 9,600 interferences. Such a titanic undertaking could only be accomplished through significant mechanization, particularly for all numerical calculations. Myoglobin crystals (from sperm whale Muscle) were photographed from 22 different orientations to obtain so-called oscillation photographs. Then, after microphotometering all the spots, the intensities of all individual X-ray patterns were normalized to produce a fully comparable series of 9,600 coefficients. Afterward, the exact same measurement Procedure was performed on four isomorphous protein derivatives containing heavy metal atoms.
Fortunately, it turned out that such derivatives are formed quite easily and in A wide variety. The protein molecule contains numerous active functional groups capable of binding metal atoms or ions both covalently and through chelation/complexation. For reliability, one should choose compounds in which the number of bound metal atoms per protein molecule is small (ideally 1–2). It is essential that the positions of the heavy atoms are strictly fixed at specific points within the macromolecule, though knowing these coordinates in advance is not strictly necessary. Finally, isomorphism between the protein derivative and the native protein is mandatory. As a rule, this requirement is not difficult to meet: introducing one or two heavy metal groups into a protein macromolecule often leaves the packing of macromolecules in the unit cell (lattice) virtually unchanged. In any case, the resulting lattice distortions are negligibly small. All these circumstances help solve the chemical part of the problem—obtaining protein derivatives with heavy element atoms, with the X-ray patterns of the corresponding crystals typically serving as the criterion for successful formation.
For myoglobin, the best Reagents proved to be the following. The first is p-chloromercuribenzenesulfonate.

This substance attaches at one single site, even when added to the crystallization liquid in a 20-fold excess.
The second reagent consists of mercury-ammonia complex ions formed by heating mercuric oxide with aqueous ammonium sulfate. These ions attach to the myoglobin molecule at a single site as well, yet at a completely different location than the first organomercury compound. The third derivative suitable for X-ray structural analysis contained the gold complex ion AuCl4- (added as KAuCl4 salt during myoglobin crystallization). Gold also binds at a specific point on the myoglobin macromolecule. The reaction proceeds very slowly, taking several months. The fourth compound contained a combination of mercury in the form of chloromercuribenzenesulfonate and the mercury-ammonia complex.
For all four derivatives, X-ray photographs were likewise obtained from 22 orientations, and approximately 10,000 diffraction maxima were precisely measured. Next, difference maps corresponding to the heavy metal atom were obtained, the coordinates of the metal atom were calculated, and its amplitude coefficients were determined in both magnitude and phase. Finally, by solving vector equations, the phases of all 9,600 myoglobin interferences were obtained, culminating in the summation of Fourier series for 92 layers, based on which electron density contour maps were constructed. The clarity of this map is such that the secondary and Tertiary Structure of the protein emerged in full detail, and the structure of the Pauling–Corey helical regions was fully confirmed. The helicity of myoglobin turned out to be close to 77%, which agreed well with the estimate made by Benson and Linderstrøm-Lang based on peptide hydrogen isotope exchange. The Nature of the amino acid side chains was established in most, though not all, cases at this approximation level. Only in the subsequent final calculation at 1.5 Å resolution, dealing with 20,000 diffraction spots, did the level of analytical detail prove sufficient to identify all amino acid residues.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.