Principles of Protein Structure - H. Schultz 1982
Protein Evolution
Protein Differentiation
Construction of a phylogenetic tree for vasotocin, oxytocin, and vasopressin
The process whereby homologous Proteins acquire distinct Functions is known as protein differentiation [473]. In Enzymes, a mutation at one or two amino acid positions can alter the substrate Specificity of the protein [508]. Consequently, a shift in substrate preference from Sx to Sy can drive organismal adaptation by replacing enzyme Ex with Ey. Such adaptation events have been studied experimentally in Bacteria [519–524]. Notable Examples include the evolution of an amidase into a phenylacetamidase [521] and of a ribitol dehydrogenase into a xylitol dehydrogenase [508, 522]. Proteins with novel functions are constantly evolving in nature, as evidenced by The Development of inherited resistance to toxic chemicals in insects and to Antibiotics in bacteria [519].
* Ma = mega-annum = million years.
Protein differentiation typically begins with the duplication of the corresponding Gene. From this point on, the evolutionary Pathways of the different Amino acid sequences diverge in accordance with their functional differences. A classic example [522] of protein differentiation is the existence of multiple globin chains, such as Myoglobin and the human Hemoglobin a-, ß-, y-, ε-, and ζ-chains.
When comparing homologous proteins with different functions, one cannot assume—and indeed frequently does not observe—a constant rate of amino acid substitution fixation. While this does not pose a major obstacle to reconstructing the genealogy of functionally diverse proteins, dating evolutionary milestones based on structural comparisons of such proteins must be approached with caution.
This section examines the genealogical relationships among the Peptide Hormones vasotocin (VT), oxytocin (OT), bovine vasopressin (BV), and porcine vasopressin (PV). Vasotocin is involved in the restoration and Regulation of Water-salt balance in many vertebrates; oxytocin (in mammals) performs only the former function, whereas vasopressin performs only the latter. These hormones, whose amino acid sequences were determined in 1953 by du Vigneaud's group, represent the first studied example of peptide differentiation [145, 526]. Furthermore, their small size makes it possible to trace the major stages in the construction of a Phylogenetic Tree for homologous (poly)Peptides.
The starting point is an ordered matrix of differences.
The first step involves constructing a matrix of random and ordered differences for these hormones, as shown in Fig. 9.3. Upon subsequent refinement, this matrix is converted into a matrix of "minimum base substitutions" (MBS)* in the sequences, which is then re-ordered. Introducing the MBS scale increases The values of the matrix elements because more than 33% of the most frequently occurring Amino Acid Substitutions are incompatible with a single nucleotide substitution (see Fig. 9.11 in [203]). To some extent, this also enhances accuracy, since double and triple nucleotide substitutions can result in more than one fixed amino acid substitution at a given position. In addition, multiple nucleotide substitutions may correspond to conserved residues whose replacement can be fixed only through successive nucleotide exchanges. Thus, the applied MBS scale assigns greater weight to substitutions at strictly conserved positions.
* Objections to using the MBS scale for comparing distantly related proteins are discussed in [509].
It is necessary to establish the topological relationships among the molecular types. Before constructing a phylogenetic tree in our example, we must determine the topological connectivity among the four peptide types. This can be done graphically, as illustrated in Fig. 9.3d, where single base substitutions are marked with crosses. Thus, when moving from one molecular type (e.g., PV) to another (e.g., OT), the accumulated crosses yield an element of the MBS matrix (e.g., 3). If we apply the same approach using The amino acid substitution matrix instead, we fail to obtain a solution for the considered example, even when arranging the molecules as a topological graph (Fig. 9.3d).
A more detailed analysis shows that Structure/149.html">The problem of fitting the difference matrix to a topological graph corresponds to a system of six linear equations (the six matrix elements given in Fig. 9.3c) with only five unknowns (the number of crosses in each branch), which is solvable only in special cases. In our case, we obtain a direct solution using the MBS matrix for one of the three possible topologies. An amino acid difference matrix can also be used if we assume that the mutation in the OT branch occurs at the same site as the mutation in the PV branch—that is, three substitutions occur during the transition from PV to OT, even though the difference matrix shows only two.
Topological relationships can be converted into a phylogenetic tree. Based on existing molecular types, ancestral nodes must be placed at branching points. The branching point is chosen so that the number of substitutions in all branches leading to extant species is as similar as possible. In our example, both branching points α and β are equivalent in this regard (Fig. 9.3), leading to the Selection of BV and VT as the corresponding ancestors. However, when considering actual organisms, the issue is resolved in favor of VT, as it appears in more ancient taxa [145]. Using the tree, one can also establish that the Amino Acid Sequence of bovine vasopressin (BV) is of more ancient origin than that of porcine vasopressin (PV).
Class="center">
Fig. 9.3. Construction of a phylogenetic tree. a — Amino acid sequences of four peptide hormones. Substitutions occur at only two positions. b — Amino acid substitutions, corresponding codons, and minimum base changes (MBC). Codons belonging to MBC-1 are circled. c — Construction of the amino acid difference matrix: disorder to order → transition from amino acid differences to MBC → another type of order. d — Topological relationships among the four peptide hormones. Alternative topologies are given in parentheses. Node points α and β represent intermediate, not necessarily extant species. Crosses denote amino acid substitutions (or MBC values) in the respective branches. The circled cross represents a special case where an amino acid substitution occurs at a residue putatively located in the same position that varies in the PV–α branch. Consequently, this substitution does not correspond to the PV–OT transition in the amino acid difference matrix. e — System of six linear equations corresponding to the topology given in d, where $p_v$ is the number of crosses in the PV–α branch, $c_{onn}$ is the number of crosses in the α–β branch, etc. The system has a solution if the first equation is equated to 3 rather than 2—that is, if an amino acid Substitution at the same position is considered. There is no solution for alternative topologies. f — Phylogenetic tree constructed According to the topology in Fig. d.
It should be noted that this method provides information on the amino acid sequences of ancestors at branching points. This particular step is central to the aforementioned "ancestral sequence method." When comparing more than four (poly)peptides, an exact solution generally cannot be obtained. In such cases—for instance, with cytochrome c (70 specialized proteins) and Serine proteases (20 differentiated proteins, see below)—the ancestral sequence method [513] is used to achieve the best possible fit when constructing the optimal phylogenetic tree from the available data. Obviously, the further back in time a branching point lies, the greater the uncertainty regarding a given ancestral sequence in such a tree.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.