Fundamentals of Bioinformatics - O-gurtsov A.N. 2013
Methods of Bioinformatic Analysis
Substitution Matrices
PAM Matrices
The PAM (Percent Accepted Mutation) matrix reflects the probability of Amino Acid Substitutions during the evolutionary divergence of Amino acid sequences in protein chains. PAM matrices (Tables 8–11) illustrate changes expected over a specific period of evolutionary time, accompanied by a decreasing sequence similarity as genes encoding the same protein diverge with increasing evolutionary time.
The measure of sequence divergence is expressed in PAM units — the percentage of accepted (or fixed) Mutations. Thus, two sequences are separated by a distance of 1 PAM if they are 99% identical (in other words, one point mutation is fixed per 100 amino acid residues).
To derive the PAM matrices, Margaret Dayhoff evaluated amino acid substitutions across a group of evolving Proteins, recording 1,572 substitutions within 71 groups of protein sequences that shared at least 85% similarity. Because such amino acid substitutions are observed in closely related proteins, they represent mutations that do not significantly alter protein function. Consequently, they are termed "accepted" (or "fixed") mutations, as these amino acid substitutions were "accepted" by natural Selection and "fixed" within the population.
Class="center">Table 8 - PAM30 amino acid substitution matrix

Table 9 - PAM70 amino acid substitution matrix

Table 10 - PAM120 amino acid substitution matrix

Table 11 - PAM250 amino acid substitution matrix

Initially, such protein sequences were organized into a Phylogenetic Tree. Next, the number of substitutions of each amino acid for every other amino acid was counted. To make these counts useful for sequence analysis, information on the relative Variability (mutability, or susceptibility to substitution) of each amino acid was required.
Relative mutabilities were estimated by counting the number of substitutions for each amino acid in each group of related sequences and dividing this number by a value termed the mutation exposure of The amino acid. This factor is equal to the product of the frequencies of all substitutions that occurred at 100 random sequence positions within that group. This factor normalizes the data for varying amino acid compositions, mutation frequencies, and sequence lengths. The normalized frequencies were then summed across all sequence groups. According to these calculations, asparagine, Serine, aspartic acid, and glutamic acid were the most mutable Amino Acids, whereas Cysteine and Tryptophan were the least variable.
Based on the amino acid substitution frequencies and mutability values obtained using this method, a 20x20 probabilistic mutation matrix was constructed, reflecting all possible amino acid replacements. Because each amino acid substitution was modeled using a Markov model (see section 9.3)—where the mutation at each position is independent of previous mutations—the changes predicted for more distantly related proteins that have undergone many ($N$) mutations could also be calculated.
According to this model, the 1 PAM matrix can be multiplied by itself $N$ times to yield transition matrices for comparing sequences with progressively lower levels of similarity due to divergence over longer periods of evolutionary history (as $N$ increases).
Tables 8–11 present the PAM30, PAM70, PAM120, and PAM250 matrices. To avoid fractional numbers, all values within the matrices have been multiplied by 10 and rounded.
A distance of 250 PAMs, which corresponds to approximately 20% sequence identity, is considered the threshold level of similarity for which a correct alignment can be reliably obtained based solely on sequence analysis, without incorporating additional information such as the three-dimensional Structure OF THE protein globule. A distance of 250 PAMs implies that during the evolution of a sequence 100 amino acid residues long, 250 mutations occurred at random positions. Therefore, some positions experienced no mutations at all, whereas others underwent 3 or more mutation events.
If natural selection did not operate in nature, the frequencies of all possible amino acid substitutions would depend primarily on the Background frequencies of those amino acids within the sequence. However, the substitution frequencies observed in related proteins (target frequencies) are driven by substitutions that do not cause severe disruptions to protein function.
PAM matrices are typically converted into log-odds matrices.
The odds score (mutation score) represents The ratio of the probabilities of an amino acid substitution under two different hypotheses:
1) the observed mutation rate reflects true evolutionary change at a given site (numerator);
2) the substitution occurred due to a random mutation determined solely by amino acid occurrence frequencies and has no biological significance (denominator).
These odds ratios are converted into Logarithms to yield log-odds scores. As a result, multiplying the odds scores of Two amino acids in an alignment is conveniently replaced by summing their logarithms.
![]()
The values in the Cells of PAM matrices reflect mutation probabilities. For example, in the PAM250 matrix for the V↔M substitution, the mutation score is +2. This means that in the evolutionarily related sequences being compared, this mutation occurs with a probability 1.6 times higher than expected by chance. The value of +2 was obtained after multiplication by 10; therefore, the mutation probability is equal to 100.2 = 1.6.
In amino acid Substitution Matrices, it is assumed that the probability of substituting amino acid A with amino acid B is always equal to the reciprocal probability of replacing B with A, as it is impossible to distinguish between these two events.
Last update: 11/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.