Fundamentals of Bioinformatics - Ogurtsov A.N. 2013

Foundations of Bioinformatics
Biological sequences
Dot matrix of similarity

A dot plot is the simplest graphical representation used to visualize the similarity between two sequences.

A dot matrix of similarity (or identity) is a table or matrix in which the rows correspond to the elements of one sequence and the columns to the elements of another. In its simplest form, the Cells of a dot matrix are left empty if the compared elements are different, and filled if they match. Matching sequence regions appear as diagonal lines running from the top-left to the bottom-right corner.

As an example, let us construct a dot plot showing the matches between the short string PROFESSOR GURTSOV and the long string PROFESSOR ALEXANDER NIKOLAEVICH GURTSOV (Figure 7).

Class="center">

Figure 7 - Dot matrix of similarity between two strings

Letters corresponding to long matching regions are shown in bold, whereas single, isolated matches are not bolded. Clearly, the resulting alignment will take the form

Figure 8 presents a dot matrix showing both global and local matches of the repetitive sequence ABRACADABRACADABRA with itself.

Figure 8 - Dot matrix of matches in a repetitive sequence

The appearance of a dot matrix can clearly reveal the presence of palindromic sequences within the analyzed string.

For instance, DNA Restriction sites for restriction Enzymes are typical palindromes (Figure 9).

Figure 9 - DNA Cleavage by the restriction enzyme EcoRI

Sometimes, the palindromic nature of a DNA sequence is dictated by the requirement for a dimeric protein to interact with this DNA region, where one subunit binds to one arm of the palindrome and the other binds to the opposite arm on the complementary strand—as, for example, in the binding of the glucocorticoid receptor to the hormone response element (HRE) of DNA (Figure 10).

Figure 10 - Palindromic DNA hormone response element (HRE) bound to dimerized steroid Hormone Receptors

The HRE is a palindrome, meaning a DNA region whose two nucleotide strands are identical when read in the 5'→3' direction (see [7], sec. 3.5). For our example, the HRE has the form

Each HRE strand contains a 6-nucleotide sequence AGAACA, known as the core recognition motif. Since the HRE contains two such motifs, two receptors bind to the HRE.

The two 6-nucleotide sequences are separated by three Base Pairs (designated as NNN in Figure 10), which provide sufficient spacing for the receptor homodimer to bind to the HRE. These three base pairs can be of any type, as they do not affect binding affinity to the receptor complex.

Figure 11 shows the characteristic appearance of the dot plot for the palindrome AROSAUPALANALAPUAZORA.

Figure 11 - Dot matrix of similarity (matches) for the palindromic sequence AROSAUPALANALAPUAZORA

Long DNA or RNA regions containing inverted repeats of this type can form hairpin structures. In addition, certain mobile elements isolated from plants contain imperfect palindromic sequences—inverted repeats of non-complementary sequences located on the same strand.

Another example of a palindrome is ttttcgtgagtgcggaggctttt, a fragment of the Wheat Dwarf Virus genome, which causes stunting in wheat.

Dot plots provide a quick way to illustrate the relationship between two sequences, making striking similarities clearly visible. For instance, the dot plot showing the relationship between the mitochondrial ATPase GENES OF THE sea lamprey Petromyzon marinis and the smallspotted catshark Scyliorhinus canicula reveals that sequence similarity is least pronounced at the beginning (Figure 12).

Figure 12 - Dot matrix of matches for ATPase-6 from sea lamprey and catshark

Sometimes, dot plots are drawn in a "traditional" orientation, where the "origin" (THE START OF the sequences) is placed in the bottom-left corner rather than the top-left. Consequently, the direction of the vertical axis changes accordingly (Figure 13).

Figure 13 - Dot matrix of matches between the linear chromosome of S. meliloti and the circular chromosome of A. tumefaciens

Figure 13 suggests that these organisms shared a common ancestor.

Another example of using dot plots to compare The nucleotide sequences of genes encoding human Hemoglobin α and β subunits is shown in Figure 14. The main diagonal of the plot demonstrates significant sequence similarity.

Regions of similarity are often shifted, which causes them to appear on parallel diagonals of the dot matrix. Such shifts occur As a result of insertions or deletions (indels).

The example in Figure 7 demonstrates the result of inserting the string ALEKSANDRNIKOLAEVICH into the string PROFESSOROGURTSOV, or, equivalently, deleting the substring ALEKSANDRNIKOLAEVICH from the string PROFESSORALEKSANDRNIKOLAEVICHOGURTSOV. Both actions result in the displacement of diagonal matches away from the main diagonal.

Figure 14 - Dot matrix of matches for genes encoding human hemoglobin α and β subunits

Shifts in the diagonal line are also noticeable in the nucleotide sequences of the genes encoding human hemoglobin α and β subunits, indicating the presence of insertions or deletions within these hemoglobin genes.

Figure 15 shows the dot matrix of matches between the PAX-6 Proteins of the mouse and the eyeless protein of the fruit fly Drosophila melanogaster.

Figure 15 - Dot matrix of matches for PAX-6 proteins from mouse (vertical axis) and eyeless from the fruit fly Drosophila melanogaster (horizontal axis)

Figure 15 clearly displays three extended regions of similarity. Two of them are located near the beginning of the sequences, while the third is in the middle. Between two of these three regions, the mouse protein sequence contains a longer intervening segment than the fruit fly protein sequence.

Comparing global alignment, which searches for similarity across the entire length of sequences, with local alignment, which focuses solely on isolated regions of similarity in specific parts of sequences, it should be noted that from a biologist's perspective, searching for local similarity can yield more meaningful and precise results than evaluating alignments along the entire length of sequences.

This is because functionally active regions are typically located within relatively short domains that remain conserved regardless of deletions or Mutations occurring elsewhere in the sequence.

The main advantage of the dot matrix method in sequence alignment searches is that it allows the identification of all possible residue matches between two sequences, giving the researcher the opportunity to select the most valuable ones. Subsequently, the sequences of well-aligned regions can be determined using other sequence alignment Methods (such as dynamic programming). The alignments produced by these programs can be compared with the dot plot alignment; such a comparison will reveal whether the longest regions match and whether insertions and deletions are placed in the most appropriate positions.

The accuracy of identifying matching regions can be improved by filtering out random matches found in the dot matrix. This filtering is performed using a sliding window that allows the two sequences to be compared simultaneously.

The identification of sequence alignments using the dot matrix method can be carried out by counting dots on all possible diagonals of the matrix (to statistically determine which diagonals yield the most matches) and then comparing these match scores with the results of random sequence comparisons.

Dot matrix analysis is primarily a method for comparing two sequences to find potential alignments of their elements. Additionally, this approach is used to predict complementary regions within RNA that may participate in The formation of RNA secondary structures (such as hairpins), as well as to search for direct or inverted repeats in protein and DNA sequences.

Thus, for example, repeated regions distributed across the entire length of individual Chromosomes or even an entire set of chromosomes can be detected.

For example, Figure 16 shows a dot plot comparing the genomes of Sorghum bicolor and Oryza sativa.

Figure 16 - Dot matrix of sequence matches between the genomes of Sorghum bicolor and Oryza sativa; Mb - megabases - millions of base pairs

Parallel to the main diagonal running from the top-left to the bottom-right corner, direct matches in identical DNA strands of the genomes are located. Parallel to the diagonal running from the top-right to the bottom-left corner, inverted repeats in complementary DNA strands (inverted repeats between genomes) are situated.

Thus, for instance, both significant direct similarity between chromosome 2 of O. sativa and chromosome 4 of S. bicolor, and the presence of an inverted region can be observed. Meanwhile, for chromosome 1 of S. bicolor and chromosome 3 of O. sativa, only two inverted regions are observed.

Thus, the dot plot method clearly demonstrates any possible sequence alignments in the form of matrix diagonals. Analysis of a dot matrix can easily reveal the presence of insertions or deletions, as well as direct and inverted repeats, which are much harder to find using other, even more automated methods.

A dot matrix does not merely visualize the similarity between two sequences; rather, it demonstrates all possible alignments and displays their relative quality.

Alignment must not alter the "meaning" of the sequences; therefore, the order of characters within a string must be preserved during alignment, and character permutations are not allowed. Consequently, when constructing an alignment starting from the top-left corner of the dot matrix, only Three types of steps are permitted:

1) strictly to the right (→);

2) strictly downward (↓);

3) diagonally from left to right and from top to bottom (↘).

Any path through the dot matrix from the top-left corner to the bottom-right corner constructed using these steps corresponds to one of the possible alignments.

For example, Figure 17 shows three alignment options for the strings OLEKSANDRNIKOLAEVICHOGURTSOV and OLEKSANDROGURTSOV:

Figure 17 - Possible alignment options

Any path through the dot matrix from the top-left to the bottom-right corner traverses a sequence of cells, each of which represents a pair of positions: one from the row and one from the Column that match in the alignment, or indicate a gap in one of the sequences. The path does not necessarily have to pass only through filled positions. Nevertheless, the more filled positions there are along a diagonal segment of the path, the more matching residues there are in the alignment.

If the direction of movement between consecutive cells is diagonal, the pair of successive compared residues appears in the alignment without an insertion between them (they are aligned).

If the movement direction is horizontal, a gap is inserted into the sequence serving as the row index.

If the movement direction is vertical (downward), a gap is inserted into the sequence indexing the columns.

It should be noted that no movement can be made upward or to the left, as this would correspond to comparing multiple residues of one sequence with just a single residue of the other. The mathematical interpretation of the aforementioned method for choosing a path through a dot matrix is based on representing the alignment path as a graph.

A graph is defined as a set of vertices (or nodes) and a set of connections between nodes, which are called edges (or arcs).

A directed graph (or digraph for short) is a (multi)graph in which edges are assigned a direction.

A route in a digraph is defined as an alternating sequence of vertices and arcs (vertices may repeat). The length of a route is the number of arcs it contains.

A path is a route in a digraph with no repeating arcs; a simple path is one with no repeating vertices. If a path exists from one vertex to another, the second vertex is reachable from the first.

Consider two sequences of length m and n. The alignment of these sequences can be represented as a directed graph G with nodes (i, j) ($0 \le i \le m, 0 \le j \le n$) on a grid of size $(m + 1) \times (n + 1)$. A graph edge from node (i, j) to node (i', j') is possible only if $0 \le i' - i \le 1$ and $0 \le j' - j \le 1$.

Figure 18 shows the alignment graph for the sequences X = gtccgtg and Y = atactgg, which contains vertical, horizontal, and diagonal edges.

Figure 18 - Alignment graph for the sequences gtccgtg and atactgg

In a directed graph, the sink is the node toward which all incident edges are directed, whereas the source is the node from which all incident edges originate. In a global alignment graph, the single source is the node (0,0), and the single sink of the alignment graph is the node (m, n).

The path highlighted with bold arrows in the alignment graph (Figure 18) can be represented as

where s is the origin (0,0). The corresponding alignment will take the form

Another way to interpret a path on a dot matrix is through an edit script. It specifies a series of operations that transform the sequence indexing the columns—the horizontally arranged sequence (above the table)—into the sequence indexing the rows—the vertically arranged sequence (to the left of the table).

Each movement indicates the execution of one of the operations: substitution, insertion, or deletion. Upon reaching the end of the path, one sequence is successfully transformed into the other.

Generally speaking, several different sequences of edit operations can transform one sequence into another in the same number of steps, though they may correspond to different alignments.

To quantitatively compare the results of various alignment options, a specific scoring scheme is assigned to the edit operations. These schemes can be very simple, such as (+1 point) for each pair of matching characters (substitution) and a penalty (-1 point) for a mismatch.

Penalties can be fixed, proportional, or linear, consisting of a gap Introduction penalty combined with an additional gap extension penalty.

Scoring schemes vary widely, but in any case, every alignment can be assigned a total score (weight or value), making it possible to quantitatively evaluate the similarity of the aligned sequences (see also section 7.2).

Of course, it is desirable to have a scoring scheme that assigns a higher score to a biologically correct alignment, taking into account the fact that biological molecules possess an evolutionary history, Spatial Structure, biological function, and other features that constrain arbitrary sequence variation. Therefore, alongside alignment algorithms, methods for constructing the scoring system must be carefully designed, which can ultimately be quite complex and multifactorial.

An optimal alignment is defined as one with the maximum score, the highest number of matches, and the fewest differences.

A suboptimal alignment is a conditionally optimal alignment where the highest score falls below the optimal level.

In an optimal alignment, non-identical characters and gaps are positioned to maximize the number of identical or similar characters within the alignment columns.

Optimal alignments help reveal the evolutionary relationships between sequences by providing the best possible information regarding which sequence characters should occupy the same alignment columns and which represent insertions in one sequence (or, correspondingly, deletions in the other). This information is essential for predicting the Functions, structures, and evolutionary relationships of sequences based on their alignment.

It should be emphasized that although The sequence of edit operations is derived from the optimal alignment and may correspond to a real evolutionary pathway, it is impossible to prove that this is indeed the case. The larger the "edit distance," the greater the number of plausible evolutionary pathways between two sequences (see also section 7.2).

Figure 19 displays dot matrix comparisons of the same protein—the sulfhydryl proteinase Papain from papaya (PAPA_CARPA)—with four homologs of increasing evolutionary distance:

1) with a close relative, actinidin from kiwifruit (ACTN_ACTCH, Figure 19(a));

2) with a more distant relative, human procathepsin L (CATL_HUMAN, Figure 19(b));

3) with human cathepsin B (CATB_HUMAN, Figure 19(c));

4) with staphylopain from Staphylococcus aureus (STPA_STAAU, Figure 19(d)).

Figure 19 - Dot plots of sequence alignment of papain, a sulfhydryl proteinase from papaya, with related proteins

As sequences diverge progressively, it becomes increasingly difficult to deduce the correct alignment from a dot plot, and incorporating additional data on protein structures becomes necessary to achieve a biologically meaningful alignment.

METABOLISM/35.html">Selection/41.html">Review Questions and Exercises

1. Write down the single-letter codes for NUCLEOTIDES and Amino Acids.

2. What is The Genetic Code?

3. What is the paradoxical difference between the processes of Translation and protein folding?

4. Which three information-Processing and control systems are the primary focus of bioinformatics?

5. What is Biological Sequence Alignment?

6. What three types of changes occur during the evolutionary divergence of sequences from a common ancestor?

7. What is global sequence alignment?

8. What is local sequence alignment?

9. What is the purpose of motif searching in sequence alignment?

10. What is Multiple Sequence Alignment?

11. WHAT IS A similarity dot plot, and what is it used for?

12. How do inverted and palindromic sequences appear in similarity dot plots?

13. How do insertions and deletions appear in similarity dot plots?

14. Why, from a biological perspective, can searching for local similarity yield more meaningful and accurate results than evaluating alignment across the entire length of sequences?

15. What are the Similarities and differences between a graph and a directed graph (digraph)?

16. What is a path in a directed graph?

17. What are the source and sink of a graph? Which nodes act as sources and sinks in a global alignment graph?

18. What is an optimal alignment, and how does it differ from a suboptimal one?



Last update: 11/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.