Peptide Self-Regulation of Living Systems (Facts and Hypotheses) - Shataeva L. K. 2003
Peptides in Aqueous Solutions
Regulatory Peptides as Carriers of Molecular Information
Informational Load of Amino Acid Sequences in Regulatory Peptides
When examining the mechanisms of organismal self-regulation, the signaling functions of its components are of particular interest. As I. I. Schmalhausen wrote, “the connections between an Organism and its environment are not limited to the phenomena of METABOLISM and Energy Exchange. Of great importance is the perception of signals that have no direct significance in metabolism, but determine The behavior of the organism” (Schmalhausen, 1968, p. 268).
The primary level of connection between a living system and the external environment is based on changes in those environmental parameters that can be called energetic: Temperature, pressure, environmental acidity, and the flux of trophic substances. The perception of these changes is driven by the sensitivity of corresponding cellular receptors—through changes in their conformation, Hydration, degree of dissociation of ionogenic groups, and the shifting of metabolic equilibria within The Cell. This is the primary, or energetic, level of interaction between a living System and Its external environment.
The secondary level of connection between an organism and its environment focuses on changes in environmental order: this is the perception of energetically constant environmental parameters undergoing changes in orientation (direction), duration, or sequence. We perceive not only the pitch and intensity of sounds (the energetic level), but also changes in their rhythm (degenerativity) or shifts in the succession of identical sounds. Such perception (termed informational) is not energetic; rather, it is the perception of changes in environmental Entropy against a backdrop of constant energetic interaction with the environment. As a living system and its connections with the external surroundings grow more complex, metabolic fluxes are complemented by informational ones.
A defining feature of Life as a thermodynamic phenomenon is that an open living system perceives changes in environmental order (i.e., a flux of information) and alters its internal parameters (behavioral program) accordingly for self-preservation. In other words, a living system is characterized not only by metabolism and self-Replication, but also by the perception of the information flux necessary for self-regulation.
The forms of environmental order are diverse. Due to certain causes and laws of inanimate nature, out of an infinite number of simultaneously coexisting particles, vibrations, and waves, a narrow frequency band of increased intensity is formed. For instance, against the backdrop of ragged gray-blue clouds and a blue sky, a seven-color rainbow appears (sometimes a double rainbow) with a strict sequence of electromagnetic wavelength intervals. Visual Perception of a rainbow requires only one condition—that the recipient has Color Vision. A similar principle underlies The formation of musical sounds in nature. The Aeolian harp, for example, isolates a single resonant frequency of sound vibrations from the continuous noise of the wind, amplifies it, and a human hears a specific musical note. These are Examples of perceiving order against the Background of the “noise” of propagating sound or electromagnetic waves. The energetic component of these vibrations is important solely for overcoming the sensitivity threshold of the perceiving organ.
At the level of energetic interactions, regulatory functions are, strictly speaking, inherent to all substances involved in cellular metabolism. The concentrations of hydrogen and ammonium ions, Metal Ions, as well as Amino Acids and their derivatives, provide potentiation or inhibition of biochemical reactions as participants or products of these reactions (Krichevskaya et al., 1983). It can be said that biochemical regulation lies within a component concentration range of 10-2–10-7 M.
At THE MOLECULAR LEVEL, The ratio of energetic and informational components of interaction depends on the organizational level of the system. Evidently, among the components of living systems, the highest level of molecular Organization is achieved by Polypeptides, enabling them to act as Inducers and recipients of molecular information. Individual polypeptide fragments (regulatory Peptides) can serve as carriers of molecular information.
The biological activity of regulatory peptides is usually examined from two Perspectives: pharmacokinetics, which studies the duration of their action within the organism, and pharmacology, which studies the relationship between the dose of a substance (typically normalized to the body mass of the experimental animal) and the observed effect. This approach employs a molar concentration scale, which allows for the comparison of peptides with different molecular weights. The range of active concentrations can serve as a criterion for the trophic or informational functions of a peptide.
One of the distinctive features of peptide regulation is that it is observed at extremely low concentrations (on the order of 10-18–10-12 M), at which the direct participation of peptides in biosynthetic reactions cannot be substantial (Ashmarin, Kamenskaya, 1988; Ashmarin et al., 1992; Sazanov, Zaitsev, 1992). From this perspective, the effective dose of the Lys—Pro peptide equal to 10-4 mol per mouse, which affects tissue regeneration (Kohl et al., 1989), is trophic rather than signaling in nature. At the same time, synthetic peptides affecting thymocyte receptor expression at concentrations on the order of 10-13–10-14 M can be considered signaling molecules (Khavinson, Zhukov, 1992).
Cytological experiments and clinical observations indicate that preparations of the cytomedin Class also contain specifically informational tissue-specific oligopeptides, since within their effective concentration range (10-10–10-12 M) they apparently hold no trophic significance (Morozov, Khavinson, 1983; Khavinson, Zhukov, 1992; Kuznik et al., 1999).
Chemical signals transmitted by amino acids or simple Neurotransmitters, regardless of concentration, have a very brief duration. An analogy suggests itself with more highly organized communication systems: neurotransmitters, much like spoken language, are instantaneous signals because they “sound” for a very short time. As the peptide chain lengthens, the stability and lifespan of the signal increase. Information encoded by The sequence of amino acid residues, much like a written message, persists longer in time and can diffuse over greater distances in space (Ashmarin, Karazeeva, 1999).
As a rule, the Amino Acid Sequence of a peptide for which a specific regulatory or modulating function has already been established can be found within one or more significantly larger Proteins. This established fact has several interpretations. On the one hand, there is I. P. Ashmarin’s theory regarding the continuous functional spectrum of regulatory peptides that derive from common precursors (e.g., proopiomelanocortin and preprotachykinin) through stepwise Cleavage into smaller fragments. The precursor itself exhibits no special activity, but its fragments perform various regulatory functions, including neurotransmitter and neurohormonal ones (Ashmarin, Obukhova, 1986; Ashmarin, Kamenskaya, 1988).
On the other hand, V. T. Ivanov's hypothesis is well known, suggesting the existence of a specific peptide pool derived from functionally active high-molecular-weight proteins via tissue-specific Enzymatic Hydrolysis of these proteins down to peptides of a certain size (Ivanov et al., 1997). In particular, studies on Hemoglobin hydrolysis products and peptides present in the culture medium of human erythrocytes have demonstrated that specific segments of the a- and ß-globin chains exhibit biological activity not characteristic of native hemoglobin: they bind to opiate receptors, inhibit glucose-6-phosphate isomerase, inhibit angiotensin-converting Enzymes, and potentiate the action of bradykinin. Some of these peptides, designated as hemorphins, have been detected in the human Cerebellum, Pituitary Gland, and CEREBROSPINAL FLUID, as well as in the bovine Brain and porcine Hypothalamus. The authors of these studies pointed out the similarity between hemorphins and Peptides of the cytomedin class. This viewpoint was further supported by the work of R. V. Petrov on the isolation of regulatory peptides from the supernatant of Bone Marrow cell cultures, known as myelopeptides. Two of these peptides, MP-1 and MP-2, were found to be identical to the conserved fragments of the a- and ß-chains of hemoglobin (Petrov et al., 2000).
Both concepts rely on the notion of specific and programmed cleavage of high-molecular-weight precursor proteins into short peptides that perform signaling functions in maintaining tissue Homeostasis (Yankovsky, Dovnar, 1986).
Similar ideas form the basis for searching for regulatory peptides among food protein fragments, although the exact Nature of the peptide's effect on cellular functions remains undefined (Yamamoto, 1997).
However, these hypotheses regarding THE ORIGIN OF regulatory oligopeptides do not clarify the mechanisms of their regulatory action on cell function. At the same time, there are examples of feedback loops between low-molecular-weight and high-molecular-weight Hormones (regulatory peptides). Specifically, the hypothalamus synthesizes a prohormone containing six repeats of the Glu—His—Pro sequence, which is subsequently hydrolyzed to form the hormone thyroliberin (thyrotropin-releasing hormone): pyroGlu—His—ProNH2. Low-molecular-weight thyroliberin, in turn, regulates the synthesis in the adenohypophysis of thyrotropin, a polypeptide with a Molecular Weight of 28,300 Da (Oxford Dictionary..., 1997). Apparently, it is precisely the participation of oligopeptides in initiating the Transcription process of high-molecular-weight peptides that ensures the self-Regulation of Protein Synthesis.
While investigating the regulatory properties of components isolated from various tissue extracts, we put forward the “Concept of the existence in the organism of cytomedins, which represent a group of informational molecules involved in maintaining the Structural and functional homeostasis of cell populations” (Morozov, Khavinson, 1996).
The conventional approach to identifying an informational site within a Protein Structure relies on immunospecific Methods (such as RIA or FIA). For instance, the region of The amino acid B1-chain of Laminin responsible for epithelial Cell Adhesion was identified as the nonapeptide CDPGYIGSR. Further research demonstrated that this peptide can be shortened to YIGSR while fully retaining its affinity for the laminin receptor and its effect on cell adhesion; however, the removal of either Y or R drastically reduces its specific activity. It is precisely this pentapeptide that serves as the signal for epithelial cell adhesion (Graf et al., 1987).
In addition to the empirical search for information-bearing segments of the peptide chain, a theoretical approach to this problem is currently being developed—namely, statistical analysis of sequence elements in DNA chains and polypeptides (Herzel et al., 1994; Atchley et al., 2000; Weiss et al., 2000). This approach makes it possible to evaluate stable Amino acid sequences (conserved blocks) within a given protein and to calculate the configurational entropy of its chain. Based on the theory that regulatory peptides originate endogenously through the stepwise hydrolysis of protein precursors, it can be hypothesized that these very blocks are responsible for regulatory activity.
The statistical approach to analyzing amino acid sequences differs significantly from the Statistical Mechanics of gas theory. Unlike gas molecules, which can freely change their spatial positions, the elements of a polypeptide chain cannot be rearranged arbitrarily. The fixed arrangement of amino acid residues along the chain corresponds to a specific order, the measure of which is sequence entropy. If a peptide consists solely of residues of a single amino acid (a homopolymer), the probability of finding this amino acid residue at any position in the chain is 1, and the sequence entropy of the chain is 0 (i.e., minimal). This statistical certainty of residue positioning within the chain should not be confused with the certainty of chain conformation.
In polymer Thermodynamics (including polypeptides), statistical calculations employ THE CONCEPT OF “configurational entropy,” which stems from the notion of rotational isomerism in polymers (Volkenstein, 1981). Energy in a synthetic or natural polymer is distributed not among individual units (monomers), but rather among chain segments whose mobility is restricted by pivot points. The greater the number of spatial configurations a polymer chain can adopt through the rotation of its elements, the more degrees of freedom it has and the higher its entropy. As mentioned in Section 1.2.3, the conformational entropy of a polypeptide chain is proportional to its length and increases due to The Diversity of its side groups and their rotations. Any spatial fixation of a specific chain configuration—achieved in polypeptides via cooperative intramolecular bonds—reduces the number of accessible energy states and leads to the next level of Structural organization of the macromolecule: The Emergence of a- or ß-Conformations of the chain (Secondary structure).
We will not delve into the conformational ordering of peptide chains, limiting our Discussion instead to the sequence of amino acid residues fixed by covalent bonds. This is especially true given that, regardless of the spatial configuration a peptide chain assumes, the sequence of amino acids within it (the one-dimensional Primary Structure) remains constant.
Peptide chains are composed of 20 different amino acid residues. Their sequential arrangement along the chain can be random (disordered), in which case the entropy of such a sequence is maximal. As demonstrated in Tables 6 and 7, Amino acids differ in their demand within regulatory peptides and, accordingly, form repeating blocks within their chains. It can be assumed that these blocks serve as sources of specific molecular signals carrying informational significance, since their composition varies among peptides involved in regulating different physiological functions.
As one of the founders of information theory, R. Hartley, noted, in the conventional sense the term “information” is too elastic, and a specific meaning must be established for each field of study (Hartley, 1928). Since we will subsequently encounter The problem of the quantitative measurement of information, it is advisable to clarify certain terms borrowed into molecular biology from mathematical statistics and communication theory.
When investigating molecular systems, the concept of information is intrinsically linked to the concept of entropy. Entropy serves as a quantitative measure of the randomness and disorder of a system, whereas information characterizes the order and determinism of a system required for its replication. It is generally accepted that information is equal to negative entropy (Brillouin, 1966).
As is well known, the actual problem of information measurement arose during The Development of technical systems for information transmission. To transmit information, a system of symbols (dots, dashes, letters, numbers) is used, which, by mutual agreement between the transmitting and receiving parties, have a definite meaning. Thus, the primary task of information transmission is its most accurate reproduction. When establishing a quantitative measure of information I, it was decided to use the relationship
I = n log2 s,
where s is the total number of distinct symbols in the adopted transmission system, and n is the length of the sequence composed of these signals, i.e., the text length. Here, a base-2 logarithm is used, and the bit is adopted as the unit of information measurement (Hartley, 1928).
When applying information theory methods to the analysis of peptide structures, we must provisionally exclude any attempts to interpret the obtained results, assuming that each combination of symbols (amino acid residues) is determined by an internal coding system evolutionarily fixed by living systems for accurate reproduction. Therefore, instead of the terms “word entropy” and “information content”, we will use the term “information charge”, since at the current level of knowledge, the content and meaning of intermolecular interactions are insufficiently understood.
In statistical calculations, a distinction must be made between exact and probabilistic information. Exact information differs from probabilistic information primarily by its irreversibility: an informational “message” and the form of its expression have a beginning, a definite sequence, and an end, i.e., this order is irreversible. For the long-term fixation of exact information, humanity has long used written culture.
Examining the amino acid sequences of regulatory peptides presented in Tables I–IV of the Appendix, we will find much that brings their structure close to The structure of written language. They consist of 20 distinct elements (amino acids), have a beginning (N-terminal amino acid) and an end (C-terminal amino acid), and are unbranched. As early as the 1960s, J. Bernal noted the analogy in the structural hierarchy of proteins and Nucleic Acids with the Construction of Human conversational language (Bernal, 1969).
Later, another researcher, de Duve, noting that peptides differ in the number, nature, and sequence of their amino acid residues, compared them to words of varying lengths written using a 20-letter alphabet. This is particularly evident when single-letter symbols are used in notation. For example, the tetrapeptide Cys—Glu—Leu—Leu becomes “CELL”, and the nonapeptide Ala—Arg—Cys—His—Glu—Thr—Tyr—Pro—Glu turns into “ARCHETYPE”. “Essentially, all English-language literature could be written in the amino acid alphabet if it were not missing the letters B, J, O, U, X, and Z” (de Duve, 1987, p. 43). The problem is to decode the texts of the language in which the elements of living systems, particularly peptides and Cells, “speak” or “correspond” with one another, and to establish the molecular (physicochemical) mechanism by which they exchange messages. Apparently, such decoding requires knowing the subject matter and context associated with the text.
At the time, a statistical method for finding stationary regions in a sequence of coding signs was developed to decode encoded texts. The works of C. Shannon (1963) and N. Wiener (1983) pursued a specific goal: to predict the sequence of letters in a text when, for example, half of them are missing, but the Frequency Characteristics of their pair, trigram, and tetramram combinations are available.
Since the pioneering work of Shannon, the probability Distribution Function in a linear sequence of signs has been used to estimate the information load per sign. As a measure of ordering and informational significance of individual elements (letters) of a text, it became customary to use The change in sequence entropy upon transitioning from a random distribution of elements in a chain to a distribution that accounts for the frequency of repeats of identical letter combinations within a given sequence. We will use this method to study the amino acid sequences of several regulatory polypeptides.
The calculation consists of a series of approximations F0, F1, ..., FN, which Shannon called N-gram entropy, since at each stage it takes into account more definite statistical correlations between N amino acid residues. It should be emphasized that no energetic content is attributed to these correlations, just as we do not consider the bond energy between the letters of a written word. Semantic content, i.e., meaning reflecting a specific physical entity, is also unrelated to the Amount of Information calculated for a given sequence (Shannon, 1951).
The entropy Fn for short sequences is calculated from standard frequency tables of individual signs, digrams, trigrams, and tetramrams. For the convenience of subsequent quantitative comparisons, sequence entropy is calculated in bits, i.e., using base-2 Logarithms.
For an amino acid “alphabet” of 20 letters, i.e., for the 20 encoded amino acids, the zeroth approximation is
Fq = log2 20 = 4.32 bits/residue
under the assumption that all amino acids occur within a protein with equal probability (Shannon, 1963). However, as we shall see, this is not the case in Natural peptides.
By analyzing the probability distribution for individual amino acids and their combinations within each protein, one can estimate the averaged change in entropy per amino acid residue as the difference between the N-gram and (N — 1)-gram entropies. The change in N-gram entropy, ∆FN, upon adding another residue to the block measures The amount of information attributable to the added residue, and is calculated using the formulas (Shannon, 1963):
FN = -Σp(bi) log2р(bi),
∆FN = -Σ p(bi,j) log2p(bi,j) + Σp(bj) log2p (bi),
where bi is a block of (N - 1) residues; p(bi) is the probability of block bi, determined from its frequency of occurrence; j is the amino acid residue following block bi; p(bi, j) is the probability of a block of N residues; ∆FN is the mean entropy value of a single amino acid residue when added to the preceding (N - 1) residues.
Calculations were performed for several regulatory polypeptides with sequence lengths of 165–482 amino acid residues (a. r.). These are significantly shorter than the length (2.9 million a. r.) used in the work of Weiss (Weiss et al., 2000), but sufficient for statistical estimations. Their amino acid sequences are presented in Table V of the Appendix.
The neurospecific calmodulin-binding protein P-57 was isolated from brain membranes. This protein is characterized by the absence of ordered chain regions with α- or β-structure, and nearly half of its amino acid residues are hydrophilic in nature. It presumably constitutes the outer part of the membrane receptor (Wakim et al., 1987).
Troponin T is one of the subunits of the troponin protein, which regulates the contractile activity of Actomyosin in striated Muscle. Unlike the previous protein, three-quarters of the amino acid residues of troponin T are incorporated into α-helical regions; 50% of the amino acid residues possess ionogenic side groups (Leszyk et al., 1987).
Platelet-derived epithelial cell growth factor (PDECGF) differs significantly in molecular mass and Amino Acid Composition from vascular endothelial growth factor (VEGF), although their physiological functions are quite similar (Heldin et al., 1991). Calculations were also performed for the glial cell line-derived neurotrophic factor (GDNF), which stimulates the growth of dopaminergic Neurons (Lin et al., 1993), and for the aquaporin-0 membrane receptor, The properties of which we will discuss in Section 2.2 (Nemeth-Cahalan, 2000). Frequency CHARACTERISTICS OF THE amino acid composition of these proteins were presented in Table 6.
The calculation of FN was performed for the amino acid sequence of each protein by reading it from left to right (from the N- to the C-terminus) and breaking it down into blocks of 2, 3, 4, and 5 residues; similarly to the Procedure used for the total sequence of tissue-specific peptides in Section 1.4.1, the frequency of occurrence of each N-gram was determined (Shataeva et al., 2002).
Table 8 presents the results of FN calculations (N = 1, 2, 3, 4, and 5) for the selected proteins, as well as repeating (stable) blocks of amino acid sequences in their structure. For all specified amino acid sequences, the value of F1 is lower than the theoretical value of 4.32 bits/residue for an equiprobable amino acid content in the peptide. Due to the unequal frequency of Incorporation of Amino acid residues into the chain (selectivity of amino acid insertion into the chain), the Introduction/19.html">Primary structure of a natural peptide turns out to be more informative than a model peptide containing all amino acids in equal proportions.
When listing pairs of amino acid residues contained in the sequence as signs, we effectively transition from a 20-letter alphabet to an alphabet consisting of 400 signs, for which the mean entropy value F0 is 8.63 bits/sign assuming an equiprobable presence of all signs in the sequence. However, most of the studied peptides are considerably shorter, and by no means all of these signs are found in natural polypeptides.
Table 8. Average information load per amino acid residue (a. a.) in regulatory proteins
Number of residues in N-gram |
FN, bit/block |
∆FN, bit/residue |
Repeating blocks |
FN, bit/block |
∆FN, bit/residue |
Repeating blocks |
|
Glial cell line-derived neurotrophic factor (GDNF), n = 211 a.a. |
Protein P-57, n = 239 a. a. |
|||||
1 |
4.08 |
-0.24 |
A, L, D, S, R |
3.51 |
-0.81 |
А, Е, D, G, R |
2 |
7.00 |
2.92 |
DD, KR, RG, RR |
6.51 |
3.0 |
АЕ, АР, ЕА, ED, ЕР |
3 |
7.64 |
0.64 |
AAA, KRL, QAA, TSD |
6.84 |
0.33 |
APA, AED, АЕА |
4 |
7.66 |
0.02 |
LTSD, QAAA |
7.84 |
1.0 |
AEDA, АРАА, DAPA |
5 |
— |
— |
None |
7.79 |
-0.05 |
АРААЕ |
|
Vascular endothelial growth factor (VEGF), n = 165 a. a. |
Platelet-derived endothelial cell growth factor (PDECGF), n = 482 a. a. |
|||||
1 |
4.09 |
— |
Е, С, R, К, Р, Q, F, S |
3.82 |
— |
А, Е, G, L, R, V |
2 |
6.88 |
2.79 |
KP, SC, CR, RC, RQ, GG |
7.26 |
3.44 |
АА, AL, LV, RV |
3 |
7.24 |
0.36 |
KAR, ARQ, CVP, CRP, KPH |
8.68 |
1.42 |
AAL, GVG, LVL, RAL |
4 |
7.35 |
0.11 |
KARQ |
8.91 |
0.23 |
ALVL, LAPА, PADG |
5 |
— |
— |
None |
8.86 |
-0.05 |
RVAAA, VAAAL |
|
Troponin T, n = 284 a. a. |
Aquaporin 0, n = 263 a. a. |
|||||
1 |
3.42 |
— |
A, E, D, G, R |
3.61 |
— |
L, А, G, V, S, F, R, Т, Р |
2 |
6.79 |
3.37 |
AE, EA, ED, EK, ER |
6.76 |
3.15 |
LG, GA, AV, LA, VG, SL |
3 |
7.74 |
0.95 |
AER, AED, EEE, AEE |
7.83 |
1.07 |
ALA, VAL, RAI, AIC, ASL |
4 |
7.96 |
0.22 |
AEDG, EAVE, EREK, RAER |
7.90 |
0.07 |
VALA, RAIC |
5 |
7.96 |
0 |
EAVEE |
— |
— |
None |
It is interesting to compare the statistical data obtained for amino acid residues in a peptide chain with the calculations for letter sequences in English text performed by C. Shannon (1963). The zero-order value F0 for the English alphabet is 4.7 bits/character. The occurrence frequency of these letters and their combinations in text yields F1 = 4.14, F2 = 3.56, F3 = 3.3 bits/character. Frequency tables for English text for N > 3 were not available at that time, but the analogy between the observed patterns is evident.
As N increases, the value of FN accounts for increasingly distant statistical correlations. The limiting value of FN depends on the sequence length and equals 3.32 lg n, where n is the total number of residues in the chain.
The presented data demonstrate a gradual decrease in entropy increment upon adding a single amino acid residue to the preceding block. A decrease in ∆FN is equivalent to an increase in the information load per coding unit when transitioning from a single amino acid residue to a coding unit the size of a di-, tri-, or tetrapeptide. The largest contribution to the decrease in the entropy increment ∆FN (i.e., to the information content) is made by the number of repetitions of the same regular sequence within a given chain. The significant number of dipeptide block repeats compared to quartet repeats does not indicate their elevated information content. Conversely, an analysis of letter and word statistics in English shows that the letters E, H, T, and O occur most frequently, with the highest usage frequency belonging to the definite article "the" and the prepositions "of" and "to"—that is, words that carry no independent meaning but define the relational system among meaningful words (Shannon, 1963).
The repeating peptide blocks in the studied proteins contain neither aromatic side groups nor Histidine, and they differ in their content of ionogenic and nonpolar amino acid residues. Comparing the data in Tables 6 and 8, one can see that the repeating blocks comprise exclusively amino acid residues occupying the top 9 frequency ranks. The most hydrophilic blocks are found in troponin T, while the most hydrophobic ones are part of aquaporin 0.
A comparison of the data in Tables 7 and 8 reveals that regulatory protein P-57 and Neuropeptides belonging to the same functional system share identical dimers and structurally similar tetramers, GEDA and AEDA, in their structures. Recall that 4 amino acid residues represent a spatially complete turn of an α-Helix without the rigidity of this helix (see Section 2.2 above), which allows tetrapeptide blocks to adapt to other macromolecules with helical conformation by adjusting the pitch of their own helix within permissible limits.
If we adhere to the definition that common sense is a train of thought resilient to The impact of chaotic external information, the blocks presented in Table 8 can be termed "sensible," as they have emerged and persisted for millennia among the less ordered Regions of the peptide chain. Their repetition within the protein structure underscores The Significance of their information load. Discussing the reliability enhancement of information channels in the brain and computers, N. Wiener emphasized The Importance of duplicating information transmission mechanisms: "It is scarcely to be supposed that the transmission of an important message is entrusted to a single neural mechanism. Like a computing machine, the brain in all probability acts according to a variant of the famous principle stated by Lewis Carroll in *The Hunting of the Snark*: 'What I tell you three times is true'" (Wiener, 1961).
The increase in FN with the elongation of the analyzed blocks, as shown in Table 8, reflects the stepwise information increment in the sequence as the difference in average entropy values when a block is enlarged by a single amino acid residue. This is not a very precise measure, since the probability distributions for dimers, trimers, and longer blocks are quite broad. A more accurate Assessment of the information increment relative to the elongation of adjacent blocks is possible by calculating the average entropy difference values for each stage of block extension (Eigen and Winkler, 1979).
The presented method for calculating the statistical entropy of a peptide chain is the simplest of those used, as it is computationally straightforward. However, the regulatory peptide system, which determines a continuous spectrum (continuum) of regulatory functions in the organism, is characterized not only by a set of peptides but also by their natural hierarchy. To identify periodic and coherent Signal Sequences in such systems, correlation analysis methods are employed. Recently, the method of nonparametric or rank correlations has been widely used to estimate information in a linear sequence of signs (Klimontovich, 1999).
It should be emphasized, however, that the information measure in all statistical calculation methods is established on The basis of mathematical rather than biochemical considerations. Their use is justified by empirical observations showing that any changes in a molecule's state that increase its anisotropy or packing compactness are accompanied by a decrease in entropy and a corresponding increase in information content. In particular, the standard entropy values of cationic forms of amino acids and oligopeptides in solution are significantly higher than the standard entropy of their zwitterionic forms (Biochemical Microcalorimetry, 1969)—that is, the amount of information contained in a zwitterion molecule is higher than that in the corresponding cation. Denaturation of a protein macromolecule caused by the disruption of its native structure is typically accompanied by an increase in entropy. Conversely, the transition of a polypeptide chain from a globular state to an α-helix (i.e., a transition from disorder to order) is accompanied by a decrease in entropy and a corresponding increase in information content (Lumry and Rajender, 1970; Volkenshtein, 1981). The simplest form of molecular information one can conceive reflects the stability of a particular form of molecular existence. All additional content represents the elaboration of this elementary message.
It is possible that correlation analysis of amino acid sequences will reveal a specific relationship between the structure of information peptide blocks and The Mechanism of peptide regulation. In the future, this will help define the "semantic" content of molecular information signals in terms of physical biochemistry.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.