Fundamentals of Bioinformatics - Ogurtsov A.N. 2013
Information Principles in Biotechnology
Sequencing of Biological Sequences and Gene Expression
Analysis of Protein Expression
Protein sequencing. Modern protein sequencing Methods rely on mass spectrometry, a technique that accurately determines the mass-to-charge ratio of an ion in a vacuum (m/e or m/z), allowing the calculation of the molecular mass.
Cell/13.html">Protein Structure is determined using X-ray crystallography or nuclear magnetic Resonance (NMR) spectroscopy.
X-ray crystallography involves reconstructing atomic positions based on the diffraction pattern produced when X-rays pass through a precisely oriented protein crystal. The scattered X-rays undergo constructive and destructive Interference, creating a regular pattern of signals, or reflections. The result depends on three variables:
1) scattering amplitude;
2) scattering phase (both amplitude and phase depend on the number of electrons in each atom);
3) wavelength of the incident X-rays.
To obtain structural information about a protein via X-ray crystallography, a crystal must first be grown from a solution containing the molecule of interest. Key parameters in this process include the concentration of the biological object, the type and concentration of salt, pH, the type and concentration of surfactant and other additives, as well as Temperature and crystallization rate. Sometimes, protein crystals can only be obtained after testing more than a thousand different crystallization conditions. The resulting crystal is then exposed to an X-ray beam.
The diffraction pattern, which appears as a complex yet sample-specific distribution of spots, is recorded by a detector and processed by a computer into an electron density map that reflects the spatial position and shape of the molecule's components.
The highest-quality biomolecular crystals yield very high-resolution data, allowing molecular features separated by as little as 0.1 nm (1 Angstrom, 1 Å) to be readily distinguished on such maps. Proteins are typically analyzed at lower resolutions in the 1.5–3.0 Å range. At a resolution of 1.5 Å, individual atoms are easily resolved; however, at 3.0 Å, prior knowledge of covalent bonding geometry is required to decipher the map. The three-dimensional electron density distribution is then interpreted in terms of atomic structure. High-resolution data are relatively easy to interpret and can be used with high confidence.
The main drawback of X-ray crystallography is the requirement to grow a crystal. Furthermore, the conformation of molecules within a crystal may not reflect that of free molecules in solution; more precisely, the crystal state corresponds to a single, frozen conformation of the free molecule.
NMR spectroscopy is based on the principle that certain atomic nuclei—including natural isotopes of nitrogen, phosphorus, and hydrogen—behave like tiny magnets and alter their spin magnetic moment in an applied alternating magnetic field. These processes are driven by the absorption of short-wave electromagnetic radiation. Other methods, such as magic-angle spinning NMR spectroscopy and circular dichroism spectroscopy, are also used to determine protein structure.
Protein Secondary structure prediction relies on one of three approaches:
1) empirical statistical Methods based on evaluating parameters of known three-dimensional structures;
2) methods grounded in physicochemical criteria (such as compactness of folding, Hydrophobicity, charge, Hydrogen bond energy, etc.);
3) prediction algorithms that assign a secondary structure to a polypeptide by comparing it with known structures of homologous proteins.
One of the standard empirical-statistical methods is the Chou-Fasman method, which is based on assessing the conformational preferences of Amino Acids observed in non-homologous proteins (Section 14.2). However, despite its widespread use, the reliability of this approach in determining amino acid conformational potentials has proven unsatisfactory. Conversely, prediction algorithms yield significantly higher accuracy in this domain through the analysis of Multiple Sequence Alignments.
Tertiary structure prediction (especially when built upon predicted secondary structures of the molecule) still remains beyond the capabilities of modern computers.
Protein analysis. A standard biochemical method for protein analysis—where proteins are separated based on two independent parameters: (1) isoelectric point (pI) (charge) and (2) molecular weight—is two-dimensional Polyacrylamide gel Electrophoresis (2D-PAGE, or SDS-PAGE—Sodium Dodecyl Sulfate Polyacrylamide Gel Electrophoresis).
Separation in the first dimension is performed using isoelectric focusing in an immobilized pH gradient.
The pH gradient is established by a series of buffers, and the immobilized pH gradient is created by covalently binding buffer groups to the gel, which prevents the buffer itself from migrating during electrophoresis.
Isoelectric focusing involves the forced migration of proteins under METABOLISM/18.html">The Influence of an electric field until the local pH equals the protein's pI.
The isoelectric point of a protein is the pH value at which the net charge of the protein is zero, causing it to remain stationary in an applied electric field.
Upon completion of migration, the gel is equilibrated with the surfactant sodium dodecyl sulfate (SDS) (Figure 72), which binds uniformly to all proteins and imparts a net negative charge to them.
Class="center">
Figure 72 - Diagram of the SDS molecule
This enables separation In the second dimension based on molecular weight.
Following separation in the second dimension, the protein gel is stained with a universal dye to reveal the positions of all protein spots.
Subsequently, reproducible SDS-PAGE runs with similar (tissue) samples can be performed to compare protein expression levels. This method yields a diagnostic protein profile for a given sample.
Figure 73 shows an example of the result obtained from two-dimensional polyacrylamide gel electrophoresis. Each spot corresponds to a specific protein. The proteins in the sample were separated by isoelectric point (horizontally) and by molecular weight (vertically in the figure).

Figure 73 - Two-dimensional SDS-PAGE electropherogram
The stained protein gel is then scanned to generate a digital image, in which individual protein spots are detected and measured, and the signal intensity of each spot is corrected against the surrounding Background. Several algorithms have been developed for this purpose, based on Gaussian approximation or the detection of Laplacian of Gaussian spots. Spots whose Morphology deviates from a standard Gaussian profile can be interpreted using an overlapping-shape model.
A simpler image analysis method is linear chain analysis, in which the software scans columns of pixels in the digital image and registers signal density peaks. This process is repeated for adjacent pixel columns, allowing the algorithm to determine both the spot centers and the total signal intensity of each spot.
Another approach is known as the "watershed transformation". In this method, pixel intensities are represented as a topographic map, enabling the identification of hills and valleys. This method is useful for separating clusters, chains, and small spots overlapping with larger ones (satellite spots or shoulders), as well as for merging regions belonging to the same spot.
The output of software utilizing any of these methods is a spot list.
The 2D-PAGE method can also be applied to analyze differential protein expression. It can be used to identify proteins that are upregulated or downregulated by a specific Treatment or various drugs, to search for proteins associated with particular disease states, or to monitor changes in protein expression occurring during The Development of a cell or an Organism. Once recorded, the protein expression analysis data are organized into a protein expression matrix. The results of 2D-PAGE experiments are stored in 2D-PAGE Databases, which can be accessed at WORLD-2DPAGE List Index to 2-D PAGE databases and services - http://world-2dpage.expasy.org/list/.
Last update: 11/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.