Fundamentals of Bioinformatics - Ogurtsov, A.N. 2013

Methods of Bioinformatic Analysis
Phylogenetic Analysis
Phylogeny and Phenetics

Usually, Living organisms are classified into groups based on observable Similarities and differences. If two organisms are closely related, it is generally assumed that they share a common ancestor.

Phylogeny is the description of biological relationships, typically illustrated as a tree. The observed similarities and differences among organisms are used to reconstruct this phylogeny. The science dedicated to studying the evolutionary relationships among organisms is known as phylogenetics.

Phylogenetic Analysis is essentially a method for evaluating evolutionary relationships. The evolutionary history reconstructed through phylogenetic analysis is usually depicted as branching, tree-like diagrams that represent the hypothesized pedigree of ancestral relationships among molecules, organisms, or both.

Statements regarding the phylogeny of various organisms imply their Homology and depend on Classification. Phylogeny establishes the topology of relationships (the pedigree scheme) derived either from the classification based on the similarity of one or more sets of characters, or from a model of Evolutionary Processes. In many cases, phylogenetic relationships derived from different characters are quite robust and even mutually supportive. Against the backdrop of traditional Taxonomy, molecular approaches to determining phylogeny are currently considered the most reliable.

Compared to traditional trees constructed from morphological traits, molecular phylogenies are far more informative due to their broader scope (for instance, flowering plants and mammals can be compared using protein sequences, but not morphological characters); furthermore, the results derived from this type of data are consistent and objective.

Thus, for example, based on the analysis of 16S and 18S ribosomal RNA sequences, Carl Richard Woese reconstructed the general CLASSIFICATION OF LIVING organisms (Figure 20).

Ribosomal RNA (rRNA) is an extremely conservative yet universal molecule present in the Cells of All living organisms (animals, plants, Fungi, Bacteria, parasites, etc.). It exhibits low tolerance to Mutations and evolves very slowly. The well-developed Introduction/11.html">Secondary Structure of rRNA ensures a slow rate of evolutionary change, as double-helical regions require mutually compensating base substitutions (the probability of which is negligible). The tree presented in Figure 20 is compatible with the alignment and Cluster Analysis of these molecules, and the Conclusions drawn from its evaluation do not contradict those obtained from other macromolecular studies.

The goals of phylogenetic research are to uncover the interrelationships among species, populations, individuals, or genes. Interrelationships imply kinship or genealogy, meaning the pattern (model) of descent from a common ancestor. A tree that displays all descendants from a single common ancestor is called a rooted tree.

Phylogenetic analysis of a family of related nucleic acid or protein sequences consists of establishing the possible Pathways of the family's evolution over time.

Currently, in phylogenetic analysis, DNA sequences provide the best measure of similarity between species.

By utilizing either the third ("wobble") codon position (see [7], sec. 4.4), untranslated regions (such as pseudogenes), or The ratio of synonymous to non-synonymous codon substitutions, it is even possible to distinguish selective genetic changes from non-selective ones.

For comparison, it is necessary to find genes that have diverged to an appropriate degree.

Genes that remain nearly identical among the species of interest provide no measure of similarity, while genes that have diverged too extensively cannot be aligned.

Fortunately, genes vary widely in their degree of Variability. The mammalian Mitochondrial Genome (a circular double-stranded DNA molecule approximately 16,000 bp long) provides a set of rapidly changing sequences useful for studying the evolution of closely related species. In contrast, the conserved sequences of Ribosomal RNAs were used by Carl Woese to identify the three major taxonomic domains: Archaea, Bacteria, and Eukaryota (Figure 20).

It must be taken into account that differing rates of sequence evolution across various genes can lead to disparate and even conflicting results in phylogenetic studies. This is especially true when the goal is not merely to reconstruct the topological branching pattern, but to determine the branch lengths of the tree.

In addition, Horizontal Gene Transfer and convergent evolution represent competing phenomena that complicate the inference of phylogenetic relationships.

When analyzing nucleic acid and protein sequences, the most closely related sequences can be identified by their positions on neighboring Branches of the tree. If a gene family can be identified within an Organism or a group of organisms, the Phylogenetic relationships among those genes can help predict which ones are likely to have equivalent Functions.

If the molecular sequences of two Nucleic Acids or Proteins found in different organisms are similar, it indicates that they likely originated from a common ancestral sequence. Sequence alignment reveals which positions in these sequences have been conserved and which have diverged from the ancestral sequence. With absolute confidence that these two sequences share an evolutionary relationship, they can be considered homologous.

An evolutionary tree is a two-dimensional graph that reflects the evolutionary relationships of both organisms and their genes. Individual sequences are treated as taxa, representing phylogenetically distinct units—the branches of the tree. It is important to realize that each node of the tree represents a branching of an organism's (or gene's) evolutionary pathway into two distinct species that are reproductively isolated from one another. When constructing a tree of evolutionary relationships, sequences are depicted as the outer branches, while the branching connections within the tree's crown reflect the degree of relatedness among the various sequences.

The goal of phylogenetic analysis is to uncover all branching relationships within the tree and determine its branch lengths.

In phylogenetic trees, edge lengths represent either a measure of dissimilarity between two species or The amount of time that has elapsed since their divergence (see, for example, Figure 20). The assumption that differences among living species reflect their divergence times is valid only if the rates of divergence are uniform across all branches of the tree.

In general, There are two approaches to constructing a Phylogenetic Tree.

The first approach, phenetics (numerical taxonomy or clustering), has no direct bearing on the historical model of relationship among species. Here, one begins by measuring distances between species and constructs a tree using hierarchical clustering Procedures.

Clustering is defined as the grouping of similar items (or characters), distinguishing classes of objects that are more similar to each other than to objects outside those classes.

Hierarchical clustering is a multi-stage grouping of clusters into higher-level clusters.

The second approach, the cladistic (temporal) one, involves examining potential evolutionary pathways, postulating a hypothetical ancestor for each node, and selecting the optimal tree according to a specific model of evolutionary change.

Phenetics is based on phenotypic similarity, whereas cladistics is grounded in genealogy.

Hierarchical clustering excels at tree construction even in the absence of evolutionary relationships.

A simple clustering Procedure is carried out as follows: given a sample of species where a measure of similarity or dissimilarity has been established for each pair. This measure may depend on physical bodily traits, such as the difference in average adult height between representatives of two species. Alternatively, one can use the number of mismatched bases in Mitochondrial DNA alignments. To construct a tree from a set of dissimilarities, one first selects the two most closely related species and adds a node representing their common ancestor. Then, the two chosen species are replaced by a group containing both, and the distance from this pair to the remaining taxa is replaced by the average of the distances from the two chosen species to the others. We now have a set of pairwise dissimilarities not between individual species, but rather among groups of species.

Each remaining individual species is treated as a set containing only a single element.

This tree-building process is known as UPGMA (Unweighted Pair Group Method with Arithmetic Mean). A Modification of the UPGMA method, introduced by Naruya Saitou and Masatoshi Nei, is called the Neighbor-Joining method, which was designed to account for rate variations in evolution across different branches of the tree.



Last update: 11/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.