Fundamentals of Bioinformatics - Ogurtsov A.N. 2013

Foundations of Bioinformatics
Genomes and Proteomes
Protein Structure and Information

The Functional Properties of Proteins are determined by their tertiary Structure, which is the specific spatial folding of the protein molecule's polypeptide chain. The transition from one-dimensional Amino acid sequences to three-dimensional molecular structures must be accompanied by corresponding adaptations in the biological information used to describe these structures.

Proteins fulfill a wide range of vital biological roles: structural proteins (such as viral coat proteins, the keratinized outer layer of Skin in humans and animals, and cytoskeletal proteins); catalytic proteins or Enzymes; transport proteins (like Hemoglobin) and informational proteins; regulatory proteins, including Hormones and receptors; proteins that control genetic METABOLISM/31.html">Transcription; and recognition proteins involved in Cell Adhesion, Antibodies, and Other components of The Immune System.

Proteins are relatively large macromolecules. In most cases, only a small fraction of their structure—the functional site—directly performs a specific function. The remainder of the protein globule acts as a structural scaffold, featuring a largely flexible and sometimes even loose, disordered conformation whose primary purpose is to accurately position the catalytic and binding residues in three-dimensional space.

Proteins evolve through alterations caused by Mutations in their encoding genes, which subsequently lead to Changes in the Amino Acid Sequence. The fundamental principle of evolution is that DNA variations alter both protein Structure and function, directly impacting an individual's reproductive success and thereby driving natural Selection.

To date, over 86,000 protein structures have been deposited in the PDB (Figure 1). The vast majority of this structural data has been obtained using X-ray crystallography and nuclear magnetic Resonance (NMR) spectroscopy. Understanding the spatial folding patterns of amino acid chains has made it possible to elucidate the specific functions of individual proteins, such as explaining the catalytic mechanisms of enzymes.

The Introduction/19.html">Primary Structure of a protein refers to the linear sequence of amino acid residues within its polypeptide chain.

The polypeptide chain determines spatial turns; the direction of the chain defines the bending pattern.

Protein Secondary structure refers to the ordered conformation of polypeptide chains stabilized by intramolecular Hydrogen Bonds between the C=O and N-H groups of different Amino Acids. The most stable secondary structures that ensure the maximum number of intramolecular Hydrogen bonds are α-helices and β-sheets (Figure 44).

Supersecondary structures (elementary motifs) are distinguished as thermodynamically or kinetically stable complexes of α-helices and β-sheets.

The tertiary structure refers to the Spatial Organization of all protein α-helices and β-sheets, as well as the three-dimensional distribution of all atoms within the protein molecule.

The quaternary structure of a protein is defined as the assembly of two or more polypeptide chains with tertiary structure into a functionally active oligomeric complex.

In some cases, subunits may evolutionarily merge into a single polypeptide chain, transforming the quaternary structure into a tertiary one. For instance, five distinct enzymes in the bacterium E. coli that catalyze successive steps in The Biosynthesis of aromatic acids correspond to five domains of a single protein in the fungus Aspergillus nidulans.

Class="center">

Figure 44 - Schematic representation of secondary structures: a - a-helix, b - ß-Structure

Sometimes homologous monomers form oligomeric quaternary protein structures in various ways; for example, globins form tetramers in mammalian hemoglobin, whereas in the mollusk Scapharca inaeguivalvis these same globins form dimers.

In addition to the four main LEVELS OF STRUCTURAL organization mentioned above, the following additional levels are distinguished.

✵ Supersecondary structures. Interactions between ß-structures and a-helices frequently recur in proteins; supersecondary structures include a-helix hairpins, ß-hairpins, and ß-a-ß motifs.

✵ Domains. Many proteins comprise several compact units within a single chain that can exist independently and stably. These are called domains. In the structural hierarchy, domains lie between supersecondary and tertiary structures.

✵ Modular proteins. Modular proteins are multidomain proteins that often contain multiple copies of closely related domains. These domains appear in various structural contexts, so that different modular proteins represent a mosaic of such domains.

For instance, Fibronectin, whose linear sequence is represented as (F1)6(F2)2(F3)15(F1)3, is a large extracellular protein involved in cell adhesion and migration, containing 29 domains that include multiple tandem repeats of Three types of domains designated as F1, F2, and F3. Fibronectin domains also appear in other modular proteins. A dedicated website covers modular proteins (Figure 45)

http://www.bork.embl-heidelberg.de/Modules/.

The site also provides diagrams of modular proteins and their nomenclature.

Figure 45 - Modular proteins website

The most general Classification of protein structural families is based on the secondary and Tertiary Structure of the protein.

Within these rather broad categories, proteins exhibit A wide variety of folding patterns (Table 6).

Proteins with similar folds encompass families that share a considerable number of structural, sequential, and functional details due to their evolutionary relationships. However, unrelated proteins frequently adopt similar folding patterns as well.

Protein Structure classification plays a central role in bioinformatics, serving, at the very least, as a bridge between sequence and function.

Table 6 - Protein Classes

Class

Characteristics

a-helix

secondary structure consists almost exclusively of a-helices

ß-structure

secondary structure consists almost exclusively of ß-sheets

a + ß

a-helices and ß-sheets are localized in different Regions of the molecule; ß-a-ß supersecondary structures are absent

a/ß

helices and sheets are assembled from ß-a-ß structural units

a/ß barrel

the line passing through the centers of gravity (strands) of the sheets is nearly straight

Unstructured

contains few or no secondary structure elements

The amino acid sequence (primary structure) of a protein dictates its three-dimensional structure. When a denatured protein is placed in appropriate conditions, such as those found in The Cell Cytosol, it refolds into its native active state—a process known as spontaneous protein folding. While certain proteins require the assistance of specialized proteins called chaperones to fold correctly, these helpers merely accelerate the process rather than direct it (see [9], sec. 4.5).

If the amino acid sequence contains sufficient information to determine its own three-dimensional structure, it should be feasible to develop an algorithm that predicts the Spatial Structure from the sequence. Nevertheless, this is exceptionally challenging. Consequently, to address the fundamental problem of Protein Structure Prediction from sequence, researchers break it down into more manageable tasks:

✵ Secondary structure prediction. Which sequence segments form a-helices or ß-sheet strands?

✵ Fold Recognition. Given a library of known structures alongside their amino acid sequences, and a sequence of unknown structure, can we identify a library structure that is most likely to share a folding pattern with the unknown protein?

✵ Homology modeling. Suppose we are given a protein with a known sequence and an unknown structure, along with homologs of this protein whose structures are known. We can then assume that the target protein will share similarities with the known protein, which can serve as a template for modeling the corresponding structure. The completeness and reliability of the outcome depend primarily on sequence identity. It is generally accepted that if the sequences of two related proteins share 50% or more identical residues in an alignment, they will likely adopt a similar three-dimensional conformation with at least 90% probability (see also sec. 8.5).

Figure 46 illustrates the superposition of the three-dimensional structures of two related proteins: chicken egg-white Lysozyme (LYSC_CHIK) and baboon a-lactalbumin (LALBA PAPCY).

Sequence alignment of these two proteins using ClustalW2 revealed that their sequences are quite similar (37% identical residues between the two sequences), and consequently, their three-dimensional structures are remarkably similar. Either protein could serve as a viable model for the other, given how closely their main-chain (peptide) backbone traces match.

Before the advent of Genetic Engineering, molecular biologists resembled astronomers—they could observe the molecular objects under investigation but were unable to modify them. That is no longer the case. In modern laboratories, NUCLEIC ACIDS and Proteins can be modified at will. We can study them by introducing mutations and observing functional changes, confer novel functions upon old proteins—such as in the design of catalytic antibodies (abzymes) (see [9], sec. 12.3)—and even attempt to engineer entirely new proteins.

Figure 46 - Polypeptide backbone traces of lysozyme (from chicken egg white) and a-lactalbumin (from baboon)

Most rules regarding protein structure have been derived from observations of natural, native proteins. These rules do not necessarily apply to synthetic proteins. In natural proteins, characteristics are governed by the fundamental principles of physical chemistry and the Mechanisms of Protein evolution. While synthetic proteins must obey the laws of physical chemistry, they are not the product of evolution. As a result, Protein Engineering is rapidly emerging as a distinct scientific discipline today.



Last update: 11/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.