Fundamentals of Bioinformatics - Ogurtsov A.N. 2013

Foundations of Bioinformatics
The Concept of "Information"
Amount of Information

In 1948, Claude Elwood Shannon (1916–2001), a 28-year-old researcher at the American company Bell Telephone Laboratories, published a foundational paper entitled "A Mathematical Theory of Communication" in the Bell System Technical Journal. The birth of classical (statistical) information theory is generally associated with its appearance.

Around this time, The Development of technical communication systems necessitated the creation of optimal Methods for transmitting information across communication channels. Solving the relevant problems (message encoding and decoding, choosing error-correcting codes, etc.) required, first of all, answering the question of how much information can be transmitted per unit of time using a given set of signals.

And although the classical information theory does not even pose the question "What is information?" and, generally speaking, is practically useless for bioinformatics, it is still useful for general education to become familiar with The amount of information introduced by Shannon using the example of a text message.

Shannon's formula. The amount of information, IN, in a message containing N symbols is equal to

Class="center">

where M is the number of letters in the alphabet; pi is the probability (frequency) of occurrence of the i-th letter in the language in which the message is written; the minus sign in front of the entire right-hand side of the formula is placed so that the amount of information is always positive, despite the fact that log2pi < 0, since pi < 1.

Binary Logarithms in Shannon's formula were chosen for convenience. For example, in a single coin toss, M = 2 ("heads" or "tails"), N = 1, and

Thus, we obtain the minimum amount

of information (I = 1), which is called a "bit" (derived from "binary digit").

Sometimes natural logarithms are used in Shannon's formula. In that case, the unit of information is called a "nat" and is related to the bit by the ratio: 1 bit = 1.44 nats.

Shannon's formula made it possible to determine the channel capacity of communication links, which served as the basis for improving message encoding and decoding methods, selecting error-correcting codes, and, ultimately, developing the foundations of communication theory.

As an example, let us take a certain text that can be viewed as the result of choosing a specific arrangement of letters.

In the general case, when one option is chosen out of n possible ones (occurring with a probability pi = 1,2,...,n), the amount of information is expressed by the formula

If all options are equally probable, i.e., then

In the special case of a message consisting of N letters from a binary alphabet (M = 2), the number of options is: n = 2N and the amount of information is I = N.

This example is useful for clarifying what the word "equally probable" means in the definition of information. Imagine that the text contains symbols that are not contained in the alphabet at all (not "letters").

The a priori probability of such a symbol is considered very small and is ignored during summation because it falls outside the set under consideration.

It should be noted that Shannon's formula reflects the amount of information, but not its value.

Let us illustrate this with an example. The amount of information in a message, as determined by Shannon's formula, does not depend on a particular combination of letters: a message can be made meaningless by rearranging the letters. In this case, the value of the information disappears, while the amount of information remains the same. This example shows that we cannot substitute the definition of information (taking into account all its qualitative aspects) with the Definition of the amount of information.

Let us return once again to Shannon's formula and analyze, for example, the text "Tomorrow there will be a storm". Indeed, the meaningfulness or information content of the text "Tomorrow there will be a storm" is obvious. However, it is enough to keep all the elements (letters) of this message and rearrange them randomly, for example, into "rdea Zvubub trayai", for it to lose all meaning. Yet, meaningless information does not exist. According to Shannon's formula, however, both sentences contain the "same amount of information". What kind of information are we talking about here? Or, generally speaking, can we speak of information in relation to disjoint elements of a message?

Obviously, individual elements of a message can be called "Shannon information" only on the condition that we stop associating information with meaningfulness, that is, with content. But then this contentless entity is hardly worth calling "information", as it imbues the primary term with a meaning foreign to it.

Considering, however, that message elements are actually used to compile meaningful texts containing information, it is more convenient to treat these elements (letters, signals, sounds) as an information container that may or may not contain information, and may be contentless or empty.

Obviously, the capacity of the container does not depend on whether it is filled or what it is filled with. Therefore, the frequency characteristic of message elements (or the amount of information associated with the i-th letter of the alphabet), defined as Hi = -log2 pi, is better called not the "amount of information", but the "capacity of the information container". This, incidentally, agrees well with Shannon's formula, according to which the "amount of information" in a given message does not depend on the order of its constituent letters, but only on their number and frequency characteristics.

Obviously, in Shannon's terms, the amount of information in an intron and an exon of equal length is the same, whereas the exon participates in METABOLISM/35.html">Protein Biosynthesis (makes sense) and the intron does not.

It is worth noting that while the phrase "Tomorrow there will be a storm" is perfectly clear to a Russian reader, it is "Greek" to a foreigner. This indicates that whenever we discuss semantics, we must take into account the semantic affinity between the message and the perceiving system.

Semantics is a branch of linguistics that studies the meaning of language units.

Let us consider an example. Suppose we have a Russian text containing NK Cyrillic letters (the alphabet contains 32 letters). Its English Translation contains NL letters of the Latin alphabet (26 letters). The Russian text is the result of selecting a specific arrangement of Russian letters (the number of possible order permutations is on the order of 32NK). The English translation is a Selection of a specific arrangement of Latin letters predetermined by the Russian text (information reception). The number of possible variations in the English text is on the order of 26NL. The amount of valuable information is the same (provided the meaning is not distorted), whereas the amount of information, in the sense defined by Shannon's formula, differs.

Below, through Examples, we will see that the processes of generation, reception, and Processing of valuable information are accompanied by the "pouring" of information from one container into another.

Thus, during translation, Genetic information is "poured" from nucleotide information encoded in DNA molecules into amino acid protein information. As a rule, the total amount of information changes in this process, but the amount of valuable information is preserved.

Sometimes "information containers" are so different that we can speak of Different types of information. We will also apply this term to pieces of information that share the same meaning and value yet differ greatly in magnitude, meaning they are housed in different containers.

Shannon himself, although he did not distinguish between "information" and "amount of information," sensed that they were not one and the same. "It is very rarely," Shannon wrote, "that one can simultaneously unlock several secrets of nature with the same key. The edifice of our somewhat artificially constructed prosperity can all too easily collapse the moment it turns out that a few magic words, such as information, Entropy, redundancy... cannot solve all unresolved problems."

"Information and Entropy." THE CONCEPT OF "entropy" (from the Greek word meaning "turning" or "transformation") was introduced into physics in 1865 by Rudolf Clausius as a quantitative measure of uncertainty. According to the second law of Thermodynamics, in a closed system, entropy either remains constant (if reversible processes take place in the system) or increases (in non-equilibrium processes), reaching its maximum at equilibrium.

Statistical Mechanics views entropy (denoted by the symbol S) as a measure of the probability that a system is in a given state. Ludwig Boltzmann noted (in 1894) that entropy is associated with the "loss of information," since entropy is accompanied by a decrease in the number of mutually exclusive possible states that remain permissible in a physical system after the macroscopic information pertaining to it has already been recorded.

By analogy with statistical mechanics, Claude Shannon introduced the concept of entropy into information theory as a property of a message source to generate a certain number of output signals per unit of time. Message entropy is a frequency characteristic of a message, expressed by Shannon's formula.

Norbert Wiener wrote: "Just as entropy is a measure of disorganization, the information conveyed by a series of signals is a measure of Organization. Indeed, the information conveyed by a signal can be interpreted essentially as the negation of its entropy and as the negative logarithm of its probability. That is, the more probable a message is, the less information it contains." The measure of uncertainty is the number of binary digits required to record (write down) an arbitrary message from a specific source, or the average length of the code string corresponding to the most economical coding method.

Leon Nicolas Brillouin developed the so-called negentropic principle of information, according to which information is entropy with the opposite sign—negative entropy (negentropy). Brillouin proposed expressing information (I) and entropy (S) in the same units—either informational (bits) or physical (ergs/degree). Unlike entropy, which is viewed as a measure of a system's disorder, negentropy is a measure of its order. Using a probabilistic approach, one can reason as follows. Suppose a physical system has several possible states. According to Brillouin, increasing information about a physical system is equivalent to fixing this system in one specific state, which leads to a decrease in the system's entropy: I + S = const. The more that is known about a system, the lower its entropy. When information about a system is lost, its entropy increases. One can increase information about a system only by increasing the amount of entropy in the environment outside the system, with the condition that always ∆S ≥ I.

In accordance with The Second Law of thermodynamics, the entropy of a closed system cannot decrease over time. Formally, this means that in a closed system (for example, in a text), an increase in entropy can only signify the "forgetting" of information in order for the equality I + S = const to be maintained. At the same time, The Emergence of new information is only possible in an open system, where order parameters become dynamic variables.



Last update: 11/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.