BOTANY, VOLUME 1: CELL BIOLOGY. ANATOMY. MORPHOLOGY – 2007

1. MOLECULAR BASIS — THE BUILDING BLOCKS OF CELLS

1.3. Proteins

Proteins, which exhibit immense diversity (sometimes referred to as Polypeptides or simply proteins, from the Greek protos meaning 'first'), are found in all Cells. A vast number of them function as Enzymes, driving METABOLISM as specific biocatalysts (see 6.1.6) or acting as catalysts to ensure the proper folding of newly synthesized protein molecules into their Tertiary and Quaternary structures. These helper proteins are divided into two classes: chaperones and chaperonins (from the French chaperon — chaperone) (see 7.3.1.2; 7.3.1.4). Structural proteins lack enzymatic activity, yet they are characterized by high stability and The ability to form very large, highly ordered complexes, including thread-like or tubular structures such as filaments and microtubules. Both structural Proteins and Enzymes are located not only inside cells but also externally. Finally, receptor proteins serve to specifically recognize signaling molecules, such as Hormones, pheromones, and elicitors, or specific surface structures, such as those on the egg Cell during Fertilization. The binding of ligands via highly specific recognition triggers a cascade of second messenger reactions within The Cell, ultimately leading to a cellular response. Translocation proteins are integral membrane components specialized in recognizing and transporting specific molecules or ions across membranes. Contractile proteins (motor proteins) act as molecular “power machines,” converting chemical energy into mechanical mechanical work. Storage proteins accumulate in large quantities in seeds, as well as in vegetative storage Organs, and in smaller amounts across most cell types. Through proteolysis, storage proteins release Amino Acids necessary for the Synthesis of Other new proteins (see 6.17.4)1.

1 Proteins frequently perform multiple Functions rather than just one. For instance, the structural proteins of microtubules (tubulins) participate in intracellular molecular transport. Filament proteins exhibit ATPase catalytic activity, etc. — Editor's note.

1.3.1. Amino Acids Constituting Proteins

Proteins are polypeptides — heteropolymeric macromolecules composed of linearly linked α-aminocarboxylic acids, commonly referred to simply as amino acids. Fig. 1.11 illustrates the 20 amino acids found in proteins, grouped according to several characteristic features.

The General structural plan shared by all amino acids, shown at the top left in Fig. 1.11, is characterized by a Cα atom bonded to a carboxyl group, an amino group, a hydrogen atom, and a side chain (R group) that varies among different amino acids. In the simplest case, Glycine, R = H. Unlike all Other Amino Acids, glycine is optically inactive because its Cα atom is symmetrically substituted. The remaining 19 amino acids exhibit optical activity and belong to the L-series. Membership in the L-series can be determined by the arrangement of substituents in structural formulas depicted according to Emil Fischer (the so-called Fischer projection): if the most oxidized carbon atom is placed at the top (in this case, the carboxyl group) and the longest carbon chain is arranged vertically, the compound belongs to the L-series if the defining group (here the amino group, –NH2) is positioned on the left (Latin laeve — on the left)1. If this group stands on the right, it is the D-form (Latin dexter — right). Cα exhibits an S-configuration According to the Cahn-Ingold-Prelog nomenclature (with the exception of Cysteine, which is R, and glycine). The R or S configuration of an asynchronously substituted carbon atom is established based on priority rules for four different radicals and constitutes a Classification system entirely independent of the D- or L-nomenclature.

1 The original text gives “laevis, links,” which is incorrect. Laevis means smooth, whereas the adjective laeva existed only in the feminine gender, deriving from the lost laevus — left. — Editor's note.

Class="center">Fig. 1.11. The twenty Proteinogenic Amino Acids.

The spatial arrangement of substituents at the Cα atom is identical for all amino acids except glycine, which is symmetrically substituted; the remaining amino acids belong to the L-series with respect to THE POSITION OF the amino group. For simplicity, the three-letter and one-letter Abbreviations used in Amino Acid Sequence data are provided beneath the standard amino acid names.

Individual Amino acids are linked linearly in proteins via peptide bonds formed between the carboxyl group of one Amino Acid and the amino group of the next. The formation of a peptide bond corresponds to the formation of an acid amide and can formally (!) be viewed as a Condensation reaction accompanied by the release of Water (Fig. 1.12). In reality, polypeptide synthesis (see 7.3.1.2), which takes place in cellular Ribosomes, is considerably more complex. Peptide bonds can be cleaved hydrolytically. Protein Digestion essentially corresponds to their Hydrolysis.

Fig. 1.12. Formation of a peptide bond: A — peptide bond formation can be considered (formally!) as a condensation reaction with the release of water. Peptide bonds involving Cα atoms form the backbone of the molecule, from which side chains (R) project outward. Because the peptide bond partially possesses double-bond characteristics (B), it is rigid and planar, whereas adjacent bonds to Cα atoms can rotate freely.

In all proteins, the atoms participating in peptide bond formation, along with the Cα atoms, are linked into a uniformly structured linear framework. The Diversity of protein structures and properties is determined by The amino acid sequence (i.e., the side chains R) and the resulting structural peculiarities. As shown in Fig. 1.11, side substituents vary in size and polarity, while basic and acidic amino acids also feature groups capable of dissociation. Within physiological pH ranges — from 4 (cell walls, vacuoles) and 7 (Cytoplasm) to 8.5 (chloroplast stroma in the light) — proteins typically carry electrical charges. The pH value at which a protein exhibits no net electrical charge (i.e., positive and negative charges mutually neutralize each other) is termed the isoelectric point. At the isoelectric point, proteins precipitate especially easily because they possess weak Hydration shells (see 1.1).

1.3.2. Protein Structure

1.3.2.1. Primary Structure

The linear sequence of amino acids in proteins constitutes their primary structure. The amino acid sequence is read starting from the amino acid that has a free NH2 group at the Cα atom (the N-terminus) and ending with the amino acid carrying a free carboxyl group (the carboxyl terminus, or C-terminus). The reading direction corresponds to the direction of molecular synthesis.

The number of possible Amino acid sequences is unimaginably large. For instance, if a given protein sequence contains $n$ amino acids, and each position can be occupied by any of the 20 amino acids, the number of possible combinations is $20^n$. Even for a small protein of just 100 amino acids, the number of possible sequences is $20^{100} = 1.26 imes 10^{130}$. Nature is estimated to contain $10^{10}$ to $10^{20}$ different proteins; a single plant synthesizes approximately 20,000 to 60,000 diverse proteins. For comparison, the total number of water molecules in the World Ocean is only about $4 imes 10^{46}$.

A characteristic amino acid sequence is typical for each protein, though this alone is insufficient for fully understanding its function. Nevertheless, based on sequence similarity, related Proteins can be identified, and even evolutionary relationships among organisms can be deduced by comparing the sequences of multiple proteins (or genes) (molecular systematics; see 11.1.3.1).

For example, cytochrome c occurs as a crucial electron carrier in prokaryotes and in the Mitochondria of all eukaryotes. This protein consists of approximately 110 Amino Acids and a single covalently bound heme group. Its amino acid sequence (Fig. 1.13) is known for more than 100 organisms. Comparisons reveal that in certain positions, even unrelated organisms invariably possess the same amino acid; in other positions, similar amino acids are consistently found, while some positions accommodate A wide variety of different amino acids. Highly conserved amino acids are often critical for Protein Structure and/or function. The number of identical or similar amino acids located at equivalent positions is expressed as a percentage when comparing sequences. If sequence matches clearly (!) exceed random expectation (about 5%, noting furthermore that non-homologous proteins are extremely unlikely to share even short stretches of complete identity), the compared sequences are considered homologous, meaning they are phylogenetically related. All proteins sequenced to date can be categorized into fewer than 150 non-homologous families. Furthermore, each of these families contains numerous proteins with distinct functions. The evolution of proteins (and consequently genes) evidently proceeded through the elongation of a limited number of ancestral sequences.

Fig. 1.13. Sequence comparison of cytochrome c (compiled by S. Rensing).

Ten selected amino acid sequences (single-letter code) from a wide variety of organisms are aligned vertically so that corresponding positions lie directly above one another. Complete identity is indicated by dark shading, positions with similar amino acid residues (e.g., I/L/V: isoleucine/leucine/valine) are shaded in gray. Shown are Cytochromes c from human (Homo sapiens, Hom_sa), fruit fly (Drosophila melanogaster, Dro_me), ascomycetes — Saccharomyces cerevisiae (Sac_ce) and Neurospora crassa (Neu_cr), angiosperms — pumpkin (Cucurbita maxima, Cuc_ma), bean (Phaseolus aureus, Pha_au), wheat (Triticum aestivum, Tri_ae), ginkgo tree (Ginkgo biloba, Gin_bi), the green alga Chlamydomonas reinhardtii (Chl_re), and the bacterium Rhodospirillum rubrum (Rho_ru) representing prokaryotes.

Most proteins contain from 100 to 800 amino acids, although both shorter and longer polypeptide chains do occur. When a protein consists of fewer than 30 amino acids, it is referred to as an oligopeptide or simply a peptide. The Molecular Weight of a protein allows for an approximate calculation of the number of amino acids composing it, and vice versa. The average molecular weight of an amino acid residue within a polypeptide chain is taken to be 111 Da. Consequently, polypeptides consisting of 100 to 800 Amino acids have a molecular weight of approximately 11–88 kDa. Polypeptide chains with a mass exceeding 100 kDa (>900 amino acids) are rare. An interesting example is the plant proton ATPase, a protein essential for the energization of the Plasmalemma, which pumps hydrogen ions out of the cell (see Fig. 6.4; 6.5) via ATP hydrolysis. This enzyme is a polypeptide chain of approximately 950 amino acids (with a molecular weight of about 105 kDa).

1.3.2.2. Spatial Structure

The spatial arrangement of a polypeptide chain is determined by its primary structure. Admittedly, the underlying principles governing the folding of a protein molecule into higher-order structures are not yet fully understood. Small segments of a polypeptide chain consisting of 5 to 20 amino acids form local secondary structures, which are stabilized by Hydrogen Bonds between C=O and NH groups of amino acids that are distant in the primary sequence. Because the peptide bond can be regarded in part as a double bond, it is flatter and more rigid, whereas the bonds involving adjacent Cα-atoms can rotate freely (see Fig. 1.12). Chains of alternating peptide bonds and Cα-atoms can therefore adopt several spatial Conformations. The most common elements of Secondary structure are the right-handed α-Helix and the β-sheet; In addition to these, β-turns and random coils are also found. Random coils typically link α-helices and/or β-sheets to one another.

In an α-helix, Hydrogen bonds are formed between the C=O group of one amino acid and the NH group of every fourth subsequent amino acid along the sequence (Fig. 1.14). This gives rise to a right-handed helix, a full turn of which contains 3.6 amino acids. Amino acid residues that do not participate in forming the backbone of peptide bonds and Cα-atoms project outward from the helix. The amino acids Alanine, glutamic acid, leucine, and Methionine are commonly found in α-helical secondary structures, whereas asparagine, Tyrosine, glycine, and especially Proline occur much less frequently.

Fig. 1.14. Secondary structures of polypeptides (after R. Karlson): A—α-helix; B—antiparallel and parallel β-sheets: in the parallel β-sheet, the C=O and NH elements of the peptide bonds lie opposite an identical element, whereas in the antiparallel β-sheet, the C=O group lies opposite the NH group (and vice versa). Cα-atoms are indicated by black dots, R represents amino acid side chains, and dashed lines denote hydrogen bonds. When illustrating tertiary structures (see Fig. 1.15) for better clarity, it is customary to depict secondary structural elements schematically and without amino acid residues R. In such representations, sheet structures are typically shown as arrows pointing from the N-terminus to the C-terminus, and helices as cylinders or helical ribbons.

The β-sheet is formed As a result of hydrogen bonding between C=O and NH Functional groups of peptide bonds in different extended regions of a single polypeptide molecule, i.e., between so-called β-strands. β-Strands can be arranged either parallel or antiparallel, meaning that adjacent β-strands run either in the same direction—both from the N-terminus to the C-terminus—or in opposite directions—one from the N-terminus to the C-terminus and the second from the C-terminus to the N-terminus (Fig. 1.14). Amino acid residues are located within the β-strands, alternating above and below the plane of the strand backbone. Alanine, isoleucine, and aromatic amino acids are commonly found in β-sheets, while acidic and basic amino acids occur less frequently. Neighboring elements of secondary structure, particularly β-strands, are often linked via β-turns consisting of 4 to 8 amino acids, which are likewise stabilized typically by hydrogen bonds. In the region of the β-turn, the polypeptide chain bends sharply, which is why these are also referred to as hairpin turns. It is precisely for this reason that β-turns facilitate the formation of compact protein structures. The amino acids proline and glycine, as well as asparagine and aspartic acid, are commonly located in β-turns.

The folding process of a protein molecule culminates in the formation of a compact three-dimensional tertiary structure—an ensemble of secondary structural elements within the polypeptide chain. Small proteins containing up to 200 amino acids form a single domain. Molecules with a larger number of Amino acids can fold into two or more domains, each of which folds independently of the others. Auxiliary proteins frequently participate in this process—specifically the aforementioned chaperones and chaperonins (see 7.3.1.2; 7.3.1.4). The stabilization of the tertiary structure often proceeds through:

✵ the formation of additional hydrogen bonds;

✵ the formation of Disulfide Bonds;

✵ the establishment of nonpolar interactions, particularly within the interior of the molecule;

✵ other complex modifications, such as glycosylation;

✵ the isomerization of X-Pro peptide bonds, which, in contrast to other peptide bonds that are always trans-isomers (see Fig. 1.12), can occur in both Cis- and trans-configurations (where X is any amino acid and Pro is proline).

Unlike the formation of Hydrogen bonds and nonpolar interactions, protein modification processes are catalyzed by enzymes.

Using X-ray crystallography or nuclear magnetic Resonance (NMR) spectroscopy, researchers have successfully elucidated the Spatial structure of many complex proteins down to the positioning of individual atoms (Fig. 1.15). It turned out that the number of types of three-dimensional protein tertiary structures is limited, allowing for the classification of a finite number of Protein Families: just over 1,100 of them are known. Within a single structural family, proteins with non-homologous amino acid sequences may occur.

Fig. 1.15. Tertiary Structure of Triosephosphate isomerase from baker's Yeast (Saccharomyces cerevisiae). Schematic representation of the enzyme monomer (after L. Stryer).

In the cell, the enzyme exists in its active form as a dimer. Only the backbone conformation of the amino acid chain is shown (cf. Fig. 1.14). The structure consists of 8 parallel β-sheets (see Fig. 1.14) (shaded arrows in the center of the protein molecule) and 8 peripheral α-helices linked together by these turns

Depending on the proportion of different secondary structural elements, either globular or Fibrous proteins are formed. The former structure is characteristic of enzymes, whereas the latter is found in many structural proteins. Many proteins carry non-peptide prosthetic groups (from the Greek prostetos, meaning added). Depending on the type of additional prosthetic group, proteins are classified as Glycoproteins, Lipoproteins, Chromoproteins, Phosphoproteins, or Metalloproteins. The aforementioned cytochrome c is a chromcoprotein, as it incorporates a heme group as its prosthetic group.

The three-dimensional structure of a protein combines stability with dynamics. For instance, very small, typically active sites of an enzyme influence the overall conformation of the molecule. A significant portion of the tertiary structure serves to ensure the precise formation and stabilization of the Active Site structure. Functional conformational changes are characteristic of many proteins. For example, the conformation of receptors and enzymes changes upon Ligand binding. This process is referred to as an induced fit between the conformation of the protein's active site and its substrate. Conformational changes also occur in motor proteins during their working cycle (e.g., Myosin, dynein, kinesin; see 2.2.2.2) or in translocators (carrier proteins) during transport (see 6.1.5; 6.2.3). Reversible chemical modifications of specific amino acids, such as phosphorylation, typically affect protein activity via conformational changes. The same applies to the binding of Allosteric regulators to enzymes (see 6.1.7). Thus, protein Structure and function are governed by a variety of regulatory processes. In terms of their mode of action, proteins can be described as the molecular machines of cells.

The structure and function of most proteins depend on the conditions of the intracellular environment (including pH and Ionic strength). Soluble proteins are heavily hydrated due to the presence of polar and charged amino acid residues on their surface (see 1.1). Conversely, amino acids located in the interior of the protein molecule stabilize its conformation through nonpolar interactions.

Abrupt changes in pH or heating lead to protein Denaturation. This process disrupts the tertiary structure and, under certain conditions, the secondary structure as well. The interaction between different proteins—for instance, resulting from the exposure of nonpolar residues during denaturation—triggers aggregation and ultimately leads to the coagulation (precipitation) of the protein. Such a state is typically irreversible and is referred to as irreversible denaturation.

1.3.2.3. Protein Complexes

Many proteins can only perform their functions in association with molecules of the same or other proteins. Such protein complexes are regarded as quaternary structures, and their constituent subunits are referred to as protomers (from the Greek meros, meaning part). When only a single type of subunit is present, the complex is termed a homooligomeric protein complex; heterooligomeric protein complexes consist of two or more different subunits. Quaternary structures are held together not by covalent bonds, but by other interactions (hydrogen and ionic bonds, hydrophobic interactions). In structural proteins, quaternary structures can attain considerable dimensions: microtubules and Actin filaments reach lengths of several micrometers, whereas their globular subunits (tubulin, actin) have a diameter of only 4 nm.

As an example of a protein complex, Fig. 1.16 illustrates a proteasome. Proteasomes are present in all organisms, and their function is the degradation of regulatory and misfolded proteins. Thus, they serve in protein turnover, i.e., the continuous renewal of cellular proteins through degradation and de novo synthesis (see 7.3.1). The proteasome exhibits a tubular quaternary structure (Fig. 1.16), with the active sites of its constituent proteases located on the inner surface of the tube. As a result, only polypeptides that penetrate the interior of the proteasome are degraded. Other Examples of protein complexes include chaperonins, such as the plastid Hsp60 chaperonin, which consists of 14 identical protomers (see 7.3.1.4, Fig. 7.18).

Fig. 1.16. 20S proteasome as an example of a multimeric enzyme complex (A, B — originals by H. Zühl, C — after H. Zühl, with kind permission). A — high-resolution electron micrograph of the complex; B — reconstructed image of the proteasome from several electron micrographs, showing discernible subunits; C — schematic longitudinal section through the 20S proteasome. This multimeric, barrel-shaped protease complex consists of 4 rings of 7 subunits each, with active sites located inside the β-subunits. The narrow pores of the tubular chambers (diameter <2 nm) allow passage only to unfolded protein molecules. Recognition of the protein targeted for degradation, which is marked with ubiquitin, is mediated by additional proteins at the right or left ends of the 20S complex (not shown here), thereby forming the 26S complex (S — sedimentation units, see Box 2.1)

A multienzyme complex is formed when different enzymes are assembled into a single quaternary structure. Some of these complexes, capable of catalyzing an entire chain of sequential reactions, have extremely high molecular weights exceeding 7 106 Da, such as the Pyruvate dehydrogenase complex, which consists of nearly 100 protomers (see 6.10.3.1). Catalytic proteins are frequently associated with Regulatory Subunits. In general, protomers within quaternary structures can mutually influence one another; for example, the transition of one protomer from an inactive to an active conformation facilitates the corresponding activation of all other protomers (cooperativity, see 6.1.7).

Nucleic Acids are typically associated with protein complexes. For instance, nuclear chromosomal DNA is largely assembled with octameric histone complexes into nucleosomes (see Fig. 2.21), while Ribosomal RNAs combine with numerous different proteins to form ribosomes (see 2.2.4). Finally, many Viruses are also ribonucleoprotein particles (Fig. 1.17).

Fig. 1.17. Viral particle of Turnip Yellow Mosaic Virus (TYMV) shown in negative stain (EM micrograph by P. Klengler, Siemens AG)

The capsid, a protein coat of the virus built from 32 capsomeres, encloses the RNA-containing core. Each capsomere, in turn, consists of 5 or 6 globular protein molecules functioning as quaternary structure protomers



Last update: 07/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.