Principles of Protein Structure - H. Schulz 1982
Protein Evolution
Detection of Distant Evolutionary Relationships
Comparison of Amino Acid Sequences
A priori significance serves as an upper bound for The Significance of relationship. The analogy between two distantly related Proteins is not always obvious, and effective criteria are required to detect it. As an example, let us consider the human Hemoglobin a-chain and sperm whale Myoglobin, aligned by Dayhoff [20] while ignoring insertions and deletions. In this case, the length of the aligned chain is 142 residues, 37 of which are identical. The “a priori probability” of finding common residues in any of the 37 positions of these chains is extremely low:
Class="center">![]()
Conversely, the observation of such similarity indicates that it cannot be random and carries high significance. Accordingly, “a priori significance” is defined as the reciprocal of the “a priori probability.” High similarity significance for biological objects points to an evolutionary relationship.
Biological constraints significantly reduce a priori significance. A priori probability is a mathematical concept. To obtain more realistic values, biological constraints must be taken into account. For instance, not all Amino Acids occur with equal frequency (Table 1.1); consequently, the probability of finding a given residue at a given position will not equal 1/20 for each residue. Furthermore, all residues constitute Building Blocks of a well-defined three-dimensional Structure. For example, a charged residue generally cannot be located in the interior of a protein. This raises the probability of finding other residues in that region from 1/20 to a higher value.
Biological constraints are reflected in the relative substitution frequencies. These constraints are best accounted for by analyzing experimental data on natural proteins. An acceptable approach utilizes the relative substitution frequency matrix (Table 1.2). The elements of this matrix represent The ratio of the observed amino acid substitution frequency to the frequency expected by chance for a given amino acid occurrence frequency distribution (Table 1.1). Therefore, these elements reflect, on average, the natural constraints imposed on Amino Acid Substitutions.
When identifying the similarity of two sequences using this matrix, the result for a given sequence position is not obtained as a binary “yes” (match) or “no,” but in a more “comparative” form: the substitution Ile ↔ Ile occurs with a high level of significance, whereas Ile ↔ Val (a frequent substitution), Ile ↔ Thr, and Ile ↔ Arg correspond to progressively decreasing absolute values of the matrix elements (Table 1.2). The elements corresponding to all sequence positions can be summed and compared with random sequences. In doing so, it must be ensured that The amino acid frequencies in the random sequences are identical to those in the proteins under study. This scheme allows the detection of a relationship between completely different sequences containing A large number of probable substitutions. Conversely, sequences with 15% identity and a high number of improbable substitutions should be regarded as unrelated.
This scheme was proposed by McLachlan [598]. He converted the matrix elements m(i, j) from Table 1.2 into integers ranging from 0 to 7 by proportionally scaling all off-diagonal elements down to a 0–5 range. The diagonal elements were found to be 6 and 7, reflecting a only slightly greater significance of residue identity compared to frequently occurring substitutions. To obtain a similarity criterion, the matrix elements mr(i, j) for each position r of the chain containing amino acid i in one chain and amino acid j in the other are summed. In our example (myoglobin and the hemoglobin a-chain), this criterion equals the sum:
![]()
The distribution of all sums M expected for random sequences with the same amino acid occurrence frequencies is then calculated using combinatorial Methods (see [598]). This calculation is significantly facilitated by using small integers as the elements mr(i, j).
Since the primary interest lies not in the probability of obtaining the sums M themselves, but in the probability that a given sum M indicates a relationship, it is necessary to compare the cumulative probability (the sum of all probabilities) of all sums greater than or equal to M with the cumulative probability of all sums less than M. The latter cumulative probability can be set to 1.0, because only high values of M are of interest—i.e., low cumulative probabilities of all sums greater than or equal to M—and the sum of both types of cumulative probabilities equals 1.0. The cumulative probability distribution for all sums greater than or equal to M is shown in Fig. 9.5. When comparing sperm whale myoglobin with the human hemoglobin a-chain, this cumulative probability reaches 2 ∙ 10-9 [598], which is eight orders of magnitude higher than the a priori probability calculated above.

Fig. 9.5. Comparison of Amino acid sequences. a — cumulative probability distribution of all sums greater than or equal to M for random sequences in chains consisting of 142 residues [598]. For the calculation of M, see text. The sum and corresponding cumulative probability for the comparison of sperm whale myoglobin and human a-hemoglobin are indicated by an arrow; b — diagram of the sequence comparison matrix for two sequences I and II [598]. Both segments ABCDEFG and PQRSTUV have a length of 7 and yield a single matrix element upon comparison. The diagonal shows the correlation between the sequences, which reveals insertions in sequence I at position a and in sequence II at position.
McLachlan's method provides a standard measure of significance for the relatedness of two sequences. Because the cumulative probability accounts for biological constraints imposed on amino acid substitutions (albeit in a highly generalized form), it is far more reliable than the a priori probability. Such biologically consistent probability is termed "standard probability," and its reciprocal is known as "standard significance" [387]. The threshold above which standard significance implies evolutionary relatedness is not sharply defined. Clearly, for values greater than 100, such relatedness can be considered a working hypothesis. In principle, a high standard significance can rule out convergent evolution, since convergence is determined by biological constraints, which have already been accounted for. In practice, however, one can never be entirely certain that all constraints are known.
The comparison matrix can assist in the proper alignment of sequences. The standard significance calculated above assumes an exact match without insertions and deletions, which is unlikely given an evolutionary distance as large as that between Myoglobin and hemoglobin. To somehow account for insertions and deletions, it is necessary to shift certain Regions of the sequence. However, such shifts increase the probability of obtaining high M-sum values because the search for a good fit is performed automatically. Therefore, the original probability distribution (Fig. 9.5, a) can no longer be applied.
In this case, the primary task is to localize insertions and deletions. For this purpose, a sequence of finite length (e.g., seven residues) is selected, assuming no shifts (i.e., insertions and deletions) occur within it. Next, all possible pairs of disruptions (p, q) corresponding to position p of the first chain and position q of the second chain are aligned. Each superposition yields a sum M(p, q), as shown in Fig. 9.5, b. This sum (or corresponding designations) is entered into the comparison matrix (Fig. 9.5, b). Within this matrix, lines of high sum values denote the best match, while disruptions in these lines indicate insertions or deletions.
Once the alignment is established using the matrix shown in Fig. 9.5, b, one can calculate the cumulative probabilities—and consequently, the standard significance of relatedness—using the Procedure described above for comparing myoglobin and hemoglobin.
A qualitative criterion of relatedness can be derived from the observed distribution of M-sums. Very often, however, a clear-cut match cannot be found. In such cases, the observed distribution of all M(p, q) values (20,000 when comparing a chain of 100 residues with a chain of 200 residues) can be compared with the distribution expected for random sequences. If high M(p, q) values occur much more frequently than expected, the sequences share many matching regions, and the proteins are related. However, the degree of this relatedness is difficult to quantify because standard significance cannot be calculated.
The summation system eliminates false positive signals and can reveal structural repeats. Because this method places a very high value on amino acid similarities while not overemphasizing their exact identities, it is insensitive to misalignments caused by complex residue labeling, such as Trp-59 and Trp-56 in Cytochromes c and c551, respectively (Section 9.5).
This method can also be applied to comparisons within a single chain to search for repeats that may indicate Gene Duplication. Such repeats may manifest as high-significance lines running parallel to the diagonal.
In special cases where highly ordered repeats are expected—such as in Collagen or Tropomyosin—such periodicities can be detected using one-dimensional Fourier analysis [599]. In this approach, each type of residue is assigned a specific number, for example, 1 through 20 for amino acids ordered as in Table 1.2. This arrangement provides a one-dimensional representation of amino acid substitution probabilities. Using this distribution, the given sequence is converted into a continuous function whose height characterizes The properties of the respective amino acid (very high values, for instance, correspond to aromatic side chains). Subsequent Fourier transformation of this function reveals periodicities, i.e., repeats of amino acids with similar properties.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.