Molecular Biology: Protein Structure and Functions - Stepanov V.M. 2005

Primary protein structure
Primary structure as a level of protein organization

Proteins are defined as Polypeptides capable of spontaneously forming and maintaining a specific Spatial Structure. There is no sharp threshold or boundary that strictly separates proteins from Peptides. Indeed, a marked ability to form preferred Conformations in solution is already observed in relatively short peptides; moreover, this capacity is essential for the function of certain peptides (such as Hormones), facilitating their interaction with cellular receptors. Nevertheless, this is merely a precursor to the precise relationship between Amino Acid Sequence and spatial structure that constitutes the fundamental distinguishing feature of a protein.

The stabilization of a spatial structure requires a well-developed network of non-covalent interactions, which can only be achieved once the polypeptide chain reaches a certain length. Proteins are known whose polypeptide chain contains as few as about fifty amino acid residues. These include, for example, the pancreatic Trypsin inhibitor, epidermal growth factor, and certain bacteriophage coat proteins. However, such cases are relatively rare; proteins most commonly contain 100–400 amino acid residues within a single polypeptide chain forming a globular structure.

However, the length of a polypeptide chain can be much greater, reaching a thousand residues or more. So-called polyproteins are also known. These represent even longer polypeptide chains that successively form several structurally and functionally autonomous globules which, upon Cleavage of the polyprotein at 'hinge' regions by proteinases, exist as independent Enzymes. Apparently, the Biosynthesis mechanism itself does not impose significant restrictions on the length of a protein polypeptide chain.

It should be emphasized that the transition from a peptide to a spatially structured, compact protein globule is determined not by the mechanical elongation of the polypeptide chain, but by a specific sequence of amino acid residues. In other words, a randomly constructed polypeptide chain does not necessarily form a compact spatial structure spontaneously. It is highly likely that the capacity for self-assembly is characteristic of a limited range of sequences, including those corresponding to natural proteins selected over the course of evolution. In any case, Amino acid sequences are known that do not fold into a compact structure.

The great importance attached to the self-assembly of spatial structure as a hallmark of proteins is explained by the fact that it serves as the foundation for all protein properties, most notably biological function. The physical characteristics of a protein as a polymer are entirely determined by its ability to form a compact globule, the Specific features of which also govern the oligomerization of many proteins. The chemical and Functional Properties of proteins depend on the specific interactions of functional groups brought into close proximity within its spatial structure. Consequently, The behavior of these groups in proteins differs fundamentally from their reactivity in free Amino Acids and small peptides. Accordingly, a protein can function—i.e., act as an enzyme, structural or transport protein, regulator, toxin, or inhibitor—solely because it possesses a strictly defined spatial architecture.

At the dawn of modern Protein Chemistry, Danish biochemist K. Linderstrøm-Lang proposed categorizing the Organization of a protein molecule into four levels: primary, secondary, tertiary, and quaternary structures. This Classification has become firmly established in literature as it reflects the actual stages of Formation of the spatial architecture of protein molecules.

Primary Structure refers to The sequence of amino acid residues in a protein molecule. It is encoded by the structural Gene of the given protein and contains all the necessary information for the self-organization of its spatial structure. The amino acid sequence is formed As a result of mRNA Translation. However, the Introduction/19.html">Primary structure of a mature protein molecule does not always fully coincide with the immediate product of translation, which typically undergoes more or less extensive co- and post-translational modification, or Processing, during which amino acid residues, polypeptide chain length, etc., may be altered (see Chapter 11).

To determine protein primary structure, researchers increasingly resort to analyzing The nucleotide sequence in the corresponding structural gene or cDNA, which requires significantly less time and generally yields more accurate results. However, this approach is not always practical, as it requires the isolation and cloning of the gene. Furthermore, it fails to detect post-translational modifications, without which The Study of Cell/13.html">Protein Structure cannot be considered complete. Consequently, primary structure Determination Methods based on direct protein analysis retain their importance and continue to be refined. In addition to full primary structure analysis, these methods are used (usually partially) for the initial characterization of newly isolated proteins, comparison of proteins, analysis of their functionally important fragments, localization of chemically modified groups, and the like.

The number of established primary structures is growing rapidly: back in 1965, they were counted in units; in 1975, about 600 were known; in 1984, there were 2,500 sequences containing about 0.5 million amino acid residues; by the end of 1989, these figures had risen to 14,400 and 4 million, respectively.



Last update: 13/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.