Fundamentals of Bioinformatics - Ogartsov, A.N. 2013
Methods of Bioinformatic Analysis
Multiple Sequence Alignment
Visualization of Alignment Results
The alignment of two or more sequences is referred to as Multiple Sequence Alignment.
Group analysis of sequences belonging to Gene families involves establishing relationships among more than two group members, which helps uncover hidden conservative CHARACTERISTICS OF THE family.
The goal of multiple sequence alignment is to provide a concise yet comprehensive summary of sequence Structure data, which can then be used to determine whether these sequences belong to the gene family under study. Compared to pairwise alignment, multiple sequence alignment yields more information about evolutionary conservation. To maximize its informativeness, a multiple sequence alignment should contain a balanced sample of closely and distantly related sequences.
To construct an optimal multiple sequence alignment, as many similar characters as possible are aligned into corresponding columns. Multiple alignment of a group of sequences can provide insight into the most conserved regions characteristic of that group. In Proteins, such regions may be represented by conserved domains—either functionally active or structural.
If The structure of one or more alignment members is known, it is sometimes possible to predict which Amino Acids form similar spatial structures in other protein members of the alignment, or which genes occupy the same positions in the sequences of other nucleic acid members.
Multiple sequence alignment is also used to design probes specific to other group members, or to discover families of similar sequences from the same or different organisms.
To facilitate the analysis of multiple protein alignment results, Different types of amino acid residues are displayed in distinct colors on a computer screen. One possible coloring scheme is presented in Table 16.
Class="center">Table 16 - One of the possible coloring schemes for amino acid residues in the visualization of multiple protein sequence alignments
|
Color |
Residue type |
Amino acids |
|
Yellow |
Small nonpolar residues |
Gly, Ala, Ser, Thr |
|
Green |
Hydrophobic |
Cys, Val, He, Leu, Pro, Phe, Tyr, Met, Trp |
|
Purple |
Polar |
Asn, Gin, His |
|
Red |
Negatively charged |
Asp, Glu |
|
Blue |
Positively charged |
Lys, Arg |
An example of using one of the color palettes is shown in Figure 60.
Annotations or consensus sequences are an important element in visualizing multiple alignment results.
The simplest type of annotation is used by the ClustalW program, where the degree of conservation of amino acid physicochemical properties in a given alignment position (Column) is denoted by symbols: "*" for "identity", ":" for "conservative", and "." for "semi-conservative" Amino Acid Substitutions within that column (Figure 27).

Figure 60 - Representation of a multiple protein sequence alignment
A more informative way to visualize multiple alignment results is through computer generation of a consensus sequence, which displays the amino acids that occur most frequently in the corresponding columns of the alignment. An example of such a consensus is the bottom row ("Consensus") in Figure 60.
To display the degree of conservation, the consensus sequence is rendered as a histogram (Figure 61(a)), often labeled with the most probable amino acids, where the size of each amino acid symbol is proportional to its frequency of occurrence in that specific column of the multiple alignment (Figure 61(b, c)).
To be informative, a multiple sequence alignment must include sequences spanning a range of evolutionary distances.
If all sequences are overly close, the information they carry is redundantly duplicated, making the alignment poorly informative.
Conversely, if all sequences are too distantly related, it becomes difficult to build an accurate alignment (except for proteins with known structures), in which case the reliability of the results and the Conclusions drawn from them is called into question.

Figure 61 - Examples of annotations: a - conservation histogram; b - consensus sequence with amino acid symbols; c - sequence logo
Ideally, a multiple sequence alignment should contain a broad spectrum of proteins with varying levels of similarity, including distantly related instances among a set of closely related homologs.
Last update: 11/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.