Principles of Protein Structural Organization - G. Schultz 1982
Protein Evolution
Gene Fusion
Gene Multiplication
Up to this point, we have considered the fusion of various genes. However, in many cases, the fusion of identical genes is observed, indicating that Gene Duplication precedes Gene Fusion.
Adjacent multiplex genes encode gene products such as ribosomal RNA and Histones. The repeated iteration (without fusion) of a structural gene leading to specific new genes is a well-known evolutionary process. Adjacent multiplex genes [582, 583] encode gene products, such as ribosomal RNA or histones [487]. In the case of ribosomal RNA, the advantage over a single gene is obvious: the METABOLISM/31.html">Transcription product (disregarding modification reactions) serves as the final product; multiple reproductions of the final gene product are not provided for in the ribosomal machinery. For histones, it has been hypothesized that the transcription of a multiplex gene into histone mRNA yields large amounts of these Proteins, which are required during the Cytology/cytology/16.html">Early stages of Embryogenesis. Numerous copies of histone mRNA likely outcompete Other types of mRNA in Ribosomes [542].
Gene multiplication appears to be important for protein diversification. It is suggested that gene multiplication played a crucial role in the Structure/98.html">Evolution of protein Structure and function [523, 525, 582, 584, 585, 592]. Following the reiteration of genes, one copy retains the original function, whereas another (or others) may evolve to perform a closely related or novel function. The most well-known Examples are human genes encoding the a-, ß-, γ-, δ-, and ζ-chains of Hemoglobins, as well as Myoglobin. There is an ongoing debate (see works [586] and [523]) regarding how many generations of a redundant gene can persist in The Genome and whether inactive (dormant) forms of this extracopy serve as possible intermediates in the evolution of a protein with a new function.
Structural repetition within a protein may also arise from unequal Crossing-over of genes.
Gene multiplication followed by their fusion leads to gene products with two or more identical substructures [587]. However, as the following example demonstrates, other processes can yield the same result. An instance of partial structural duplication has been discovered in the rare a2-chain of human haptoglobin [145, 588]. Since the Amino acid sequences of both parts are identical, and also identical to a large segment of the normal a1-chain, this structural duplication must have occurred quite recently. It is most likely caused by a chromosomal aberration (unequal crossing-over) in the ancestral (human) population. Had this event occurred much earlier—such that Sequence Homology would have been erased by Amino Acid Substitutions, insertions, and deletions—it would be impossible to distinguish between gene duplication followed by fusion, on the one hand, and a chromosomal aberration, on the other. Therefore, all ancient cases of structural repeats are generally classified simply as "gene duplications" without attempting to differentiate between the various mechanisms.
Certain repeats can be detected through the repeating arrangement of Disulfide Bonds within the structure. Structural repeats can frequently be identified by the positioning of S—S bonds, as demonstrated in the case of wheat germ agglutinin (Fig. 7.2, a). The fourfold repeat is confirmed by the presence of a corresponding repeat in the three-dimensional structure [316]. The complex genealogy of serum albumin has been postulated based on internal Amino Acid Sequence homologies [82, 589], which are evident from the arrangement of disulfide bonds (Fig. 7.2, b). Apparently, the following evolutionary stages took place: at the earliest stage, the coding by a gene of a protein comprising approximately 80 Amino Acids and The formation of an S—S bond $ ightarrow$ triplication of this structure $ ightarrow$ deletion of approximately 30 residues $ ightarrow$ duplication of the resulting structure, leading to a 400-amino-acid protein, followed by the duplication of half of this structure.
The Abundance of structural repeats indicates that novel chain-folding motifs were rare yet highly significant innovations. Duplication of The amino acid sequence and chain folding has been found in ferredoxin (Section 5.3). In parvalbumin, sequence repetition is observed primarily for residues located within the Ca2+-binding center, yet the overall repetition of the chain-folding motif is clearly manifested in the tertiary structure [59, 590]. Simple repetition of a chain-folding pattern without any amino acid sequence homology is encountered in rhodanese (Fig. 5.17, a) as well as in Trypsin-like proteins, where these regions appear as two barrels (Fig. 5.17, d). All these examples demonstrate that a new type of chain folding arose extremely rarely, while a new chain topology was repeated and subsequently conserved in each copy.
Gene duplication has been postulated for the NAD-binding domains of four dehydrogenases [91], which feature a repeating Rossmann fold (Fig. 5.12, b). However, the Evolutionary Significance of this fact (Section 9.6) is minor, since such a folding pattern is an energetically favored element of supersecondary structure (Section 5.2). Analogous cases of minor evolutionary significance are represented by structural repeats within each of the Serine protease barrels (Fig. 5.17, d) as well as structural repeats in Triosephosphate isomerase (Fig. 5.17, e).
The hypothesis that all proteins originated from oligopeptides remains unverified. The observed structural repeats gave rise to the hypothesis [591] that all proteins emerged As a result of gene duplications and fusions of shorter segments consisting of approximately 15 residues. However, this hypothesis was not supported by the statistical analysis of amino acid sequences from 50 Globular proteins [587]. The analysis of three-dimensional structures also contributes little to resolving this issue. In this case, short repeats cannot be reliably identified because the number of possible Folding Pathways for such short chain segments is small. Consequently, any structural match has a fairly low significance score (Section 9.6). The range of possibilities increases if one compares the precise dihedral angles of the backbone rather than merely the overall course of the chains. However, these angles are not strictly conserved throughout evolution.
Multiple repetitions of short sequences occur in certain proteins. Short sequence repeats have been discovered in so-called periodic proteins [145, 593], which include Collagen, wool keratin, histones, Tropomyosin, human apolipoprotein A-I, and the antifreeze glycoprotein of the Antarctic fish. In the latter protein, the repeating unit throughout the entire sequence is Ala-Thr-Thr. In some instances, periodicity may reflect Specific features of the corresponding DNA [593]; in other cases, structural features (such as the Formation of the collagen triple helix shown in Fig. 5.6, or the characteristic binding of tropomyosin to filamentous Actin and histone to the DNA double helix) may have arisen under selective pressure.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.