BIOCHEMISTRY: A TEXTBOOK FOR MEDICAL UNIVERSITIES - E. S. Severin - 2004

CHAPTER 4. BIOSYNTHESIS OF NUCLEIC ACIDS AND PROTEINS (TEMPLATE BIOSYNTHESES). FUNDAMENTALS OF MOLECULAR GENETICS

VIII. Mechanisms of Genetic Variation. Protein Polymorphism. Hereditary Diseases

The precise execution of all template biosyntheses—Replication, METABOLISM/31.html">Transcription, and Translation—ensures the faithful copying of The Genome and the transmission of an Organism's phenotypic characteristics across generations, which is The basis of heredity. However, biological evolution and natural Selection are only possible in the presence of genetic variation. It has been established that the genome constantly undergoes various modifications. Despite the efficiency of DNA proofreading and repair mechanisms, a fraction of DNA damage or errors persists. Changes in The sequence of purine or pyrimidine bases within a Gene that escape correction by repair Enzymes are termed "Mutations." Some of these remain confined to the somatic Cells in which they originated, whereas others occur in Germ Cells, are heritable, and may manifest in the offspring's phenotype as a hereditary disease.

Chromosomal rearrangements during Meiosis make a substantial contribution to genetic variation. As noted previously, the fusion of an ovum with a spermatozoon in eukaryotes is accompanied by genetic recombination, during which segments of DNA are exchanged between homologous Chromosomes. This leads to The Emergence of progeny with novel gene combinations.

A gene or portions of a gene can relocate from one chromosomal site to another. These mobile elements or DNA fragments are known as Transposons and Retrotransposons.

Transposons are DNA segments that excise from one chromosomal locus and integrate into another locus within the same or a different chromosome. Retrotransposons do not leave their initial position in the DNA molecule, but they can be copied, and the resulting copies integrate into a new site in a manner similar to transposons. By inserting into genes or regions adjacent to genes, they can induce mutations and alter Gene Expression.

The eukaryotic genome also undergoes modifications upon infection with DNA- or Introduction/7.html">RNA-containing Viruses that integrate their genetic material into the host Cell's DNA.

A. Mutagenesis

Genome alterations can be diverse, affecting DNA regions of varying lengths, ranging from entire chromosomes and genes down to individual NUCLEOTIDES (Table 4-7).

Class="center">Table 4-7. Classification of Mutations

Type of mutation

Nature of mutational changes

Examples of consequences

Genomic

Change in chromosome number

Down syndrome (presence of an extra chromosome 21)

Chromosomal

Total chromosome number remains unchanged. Chromosomal rearrangements are observed, typically visible under a Microscope.

Duchenne muscular dystrophy (X-chromosome deletions)

Gene

Changes affect a single codon or a small segment of a gene and are not detectable cytogenetically

Sickle cell anemia caused by a single nucleotide substitution in the globin β-chain gene

Genomic and chromosomal mutations are the most dramatic and are frequently observed at the somatic cell level. If they occur in germ cells, the consequences for the organism are typically lethal. The mutation rate in germ cells is high. Evidence indicates that chromosomal structural abnormalities are present in 20% of human pregnancies at the Embryonic Stage. In 90% of these cases, this leads to abnormal fetal development and embryo elimination via Spontaneous Abortion. Miscarriages occurring within the first few weeks of Pregnancy are associated with severe Chromosomal aberrations. In 50% of cases, autosomal trisomy is noted—i.e., three chromosomes are present instead of a homologous pair. An example of this pathology is Down syndrome, characterized by three copies of chromosome 21.

Certain Gene Mutations become fixed within a population, are inherited, and drive Evolutionary Processes. Such mutations underlie various hereditary pathologies accompanied either by the cessation of Synthesis of the protein encoded by the damaged gene or by the synthesis of an altered protein.

Gene, or point, mutations are primarily of three types: substitutions, in which one nitrogenous base in DNA is replaced by another; insertions, involving the incorporation of one or more extra nucleotides into the DNA molecule; and deletions (or losses) of one or more nucleotides, resulting in a shortening of the DNA molecule.

1. Substitution-type mutations arise from the replacement of one nitrogenous base by another, which alters one of the codons in the mutant gene. If the coding triplet containing the altered nucleotide results in the incorporation of the same amino acid into the protein as the original (or wild-type) codon due to the degeneracy of The Genetic Code, such a mutation is termed "silent," and the protein product remains unchanged.

When the substitution of a single base leads to the replacement of an amino acid in the mutant protein, it is called a missense mutation. In A number of cases, despite the substitution, the protein retains its biological activity. This is typically because the altered amino acid resides in a functionally unimportant region of the protein and resembles the original amino acid in its Structure and properties. Such a mutation is also considered "silent," and the substitution is deemed conservative (equivalent).


Wild-type triplet

Altered triplet

DNA template

3'-ТАА-5'

3'-GАА-5'

mRNA codon

5'-АUU-3'

5'-СUU-3'

Amino acid

-Ile-

-Leu-

Occasionally, the substituted amino acid is located in a region critical for the functional activity of the protein, and its replacement results in a functionally inactive product. For instance, a point mutation in a Serine codon (Ser is a crucial structural component of the active center in serine proteases such as Trypsin, Chymotrypsin, and several Other Enzymes) leads to a complete loss of activity. If such an enzyme is involved in Major Metabolic Pathways, this "non-equivalent" substitution can be lethal.


Wild-type triplet

Altered triplet

Wild-type triplet

Altered triplet

DNA template

3'-GGТ-5'

3'-GGА-5'

DNA template

3'-АGА-5'

3'-ААА-5'

mRNA codon

5'-ССА-3'

5'-ССU-3'

mRNA codon

5'-UCU-3'

5'-UUU-3'

Amino acid

-Pro-

-Pro-

Amino acid

-Ser-

-Phe-

In some cases, the mutant protein retains The ability to perform its function despite the incorporated amino acid substitution, though it may be less efficient than the wild-type protein. As a result of the mutation, an enzyme may exhibit a higher Km value or a lower Vmax value, and occasionally both simultaneously. Such partially functional Proteins are referred to as mutant proteins with partially impaired function.

Rarely, as a result of a mutation, the protein product of a gene becomes better adapted to perform its function. Such mutations confer advantages in the Struggle for Existence upon the offspring, and a series of such mutations can lead to the emergence of a new species.

The most deleterious effects are caused by mutations that result in The formation of one of the termination codons (nonsense mutations). During Protein Synthesis, ribosome progression will be halted at the mutant mRNA triplet: UAA, UAG, or UGA. The phenotypic expression of nonsense mutations depends on their intragenic Location. The closer the mutation is to the 5'-end of the gene—i.e., to THE START OF transcription—the shorter the resulting protein product, and consequently, the less capable it is of carrying out its biological function.


Wild-type triplet

Altered triplet

DNA template

3'-GТС-5'

3'-АТС-5'

mRNA codon

5'-САG-3'

5'-UАG-3'

Amino acid

-Gln-

Stop codon

2. Nucleotide insertion or deletion mutations

Mutations involving the insertion or deletion (loss) of nucleotides are much more numerous and hazardous to cells.

If a mutation results in an insertion or deletion within a gene of a single nucleotide pair or a segment of a double-stranded DNA molecule with a number of monomers not a multiple of 3, it disrupts the reading of all subsequent codons. This occurs because of a shift in the DNA "reading frame" and a mismatch between the codons in the DNA and the Amino Acids in the final product—the protein (Fig. 4-57).

Fig. 4-57. Nucleotide deletions (A) or insertions (B) cause frameshift mutations.

As seen in Fig. 4-57, disruptions in information decoding begin at the site where the mutation occurred, since this is precisely where the informational "reading frame" shifts. Downstream of the mutation site, the protein product will feature a random Amino Acid Sequence. Frameshift mutations frequently introduce a premature stop codon, causing early termination of Polypeptide chain synthesis and yielding a truncated product lacking biological activity.

Frameshift mutations are induced by matrix synthesis inhibitors known as "intercalators." Their large, flat molecules—resembling conventional nitrogenous bases or Base Pairs—wedge between two adjacent base pairs, effectively introducing an "extra" base into the DNA. During replication of such an altered DNA strand, mispairing with the intercalated molecule can lead to the incorporation of an extra nucleotide into the daughter strand.

Occasionally, albeit very rarely, an oligodeoxynucleotide consisting of 3 nucleotides, or a multiple of 3, is lost or incorporated into DNA. Such mutations are termed in-frame deletions or insertions. In the resulting protein product, one or several amino acids will be omitted or, conversely, added within this region, whereas the rest of The amino acid sequence will match the original molecule. As a rule, these mutations are relatively harmless.

Information regarding various Types of mutations and structural alterations in mutant proteins is summarized in Table 4-8.

Table 4-8. MAIN TYPES OF gene mutations

Types of mutations

Changes in DNA structure

Changes in Protein Structure

Substitution

Silent (no change in codon meaning)

Missense (change in codon meaning)

Nonsense (formation of a stop codon)

Substitution of a single nucleotide in a codon

Protein is unchanged

One amino acid is replaced by another. Peptide chain synthesis is terminated, producing a truncated product

Insertion

In-frame

Frameshift

Insertion of a DNA fragment of 3 nucleotides or a multiple of 3

Insertion of one or more nucleotides not a multiple of 3

The polypeptide chain is extended by one or more amino acids

A peptide with a "random" amino acid sequence is synthesized because the meaning of all codons downstream of the mutation is altered

Deletion

In-frame

Frameshift

Loss of a DNA fragment of 3 nucleotides or a multiple of 3

Loss of one or more nucleotides not a multiple of 3

The protein is shortened by one or more amino acids

A peptide with a "random" amino acid sequence is synthesized because the meaning of all codons downstream of the mutation is altered

3. Mutation frequency

The average mutation frequency in human structural loci (regions where a gene is localized in a chromosome or DNA molecule) is estimated to range from 10-5 to 10-6 per gamete per generation. However, this value can vary significantly among different genes (from 10-4 for genes with high mutation rates to 10-11 for the most stable Regions of the genome). Such substantial variations in mutation frequency are driven by The Nature of the mutational damage, The Mechanism of mutagenesis, the length of the coding region in the mutant gene, and the Functions of the protein encoded by that gene. For instance, the substitution rate of one base for another in the Hemoglobin gene lies within the interval μ = 2.5 x 10-9 — 5 x 10-9 substitutions per gamete per generation. To visualize what these figures mean, let us extrapolate this mutation rate to the entire Human Genome—3 x 109 base pairs. Multiplying the Genome Size by the rate μ, we find that the genome may acquire between 7 and 15 mutations per generation; in other words, each gamete contains this many DNA alterations compared to the parental DNA. Furthermore, because cells in every individual are diploid and formed by the fusion of 2 Gametes, the number of mutations is twice as high.

One might wonder how humanity copes with such a mutational load. In answering this question, It is worth remembering that the protein-coding regions of genes—where changes are most dangerous—account for no more than 10% of the genome. The situation is further mitigated by the fact that far from every mutation in a coding region has a phenotypic manifestation. Many fall into the 3' position of codons and are thus "silent," meaning they do not cause Amino Acid Substitutions due to the degeneracy of the genetic code; others land in domains non-essential for protein function. Only mutations occurring in gametes are passed on to offspring, and their percentage is quite small.

B. Organization OF THE human genome

The genome of Eukaryotic cells (human cells in particular) contains significantly more DNA than the prokaryotic genome. In the bacterium E. coli, the DNA contains 3.8 x 106 nucleotide pairs or about 3,000 genes. Moreover, all of this DNA performs specific functions: it encodes proteins, rRNA, tRNA, or regulates The production of gene products. The total DNA length of a human haploid set of 23 chromosomes is 3.5 x 109 base pairs, which is 1,000 times greater than the size of the prokaryotic genome. This amount of DNA is sufficient to create several million genes. However, according to many independent calculations, the true number of structural genes is around 100,000. Data from the International Consortium and Celera Genomics, obtained by sequencing The Human Genome and published in 2001, estimate the number of protein-coding genes at 31,780 According to the former organization, and 39,114 according to the latter. The nucleotide sequence of human DNA contains protein-coding regions (no more than 2% of the total genome), RNA-coding regions (about 20% of the genome), and repetitive sequences (over 50% of the total genome).

The functions of "excess" (non-coding) DNA are not yet fully understood; it is believed to be involved in gene expression regulation and RNA Processing, perform structural functions, enhance the precision of homologous pairing and chromosome recombination during meiosis, and facilitate successful replication. Much of this DNA arose through the reverse transcription of RNA and the action of mobile elements.

Highly repetitive DNA sequences

Within the fraction of excess DNA, relatively short sequences ranging from 2 to 10 base pairs in length, repeated millions of times, are distinguished. These highly repetitive sequences are termed "satellite DNA." They make up about 10% of the entire human genome, are scattered throughout The Cell's genetic material, and are predominantly localized in the centromeric and telomeric regions of most chromosomes.

Moderately repetitive DNA sequences

Another identified group consists of moderately repetitive DNA sequences, highly heterogeneous in length and copy number, making up over 30% of the human genome. This fraction includes DNA encoding The structure of rRNA, tRNA, and certain mRNAs. Histone genes are present in the genome in several hundred copies and also belong to this class. Moderately repetitive DNA sequences include non-transcribed regions that are crucial for regulating gene expression: promoters and enhancers. This group was found to contain DNA sequences approximately 300 base pairs long that are repeated about a million times (on average every 5,000 base pairs) and dispersed throughout the human genome.

Common representatives of moderately repetitive sequences are the LINE family (long interspersed elements), which comprise DNA sequences from 5,000 to 8,000 base pairs in length and occur in the genome in 850,000 copies.

Unique DNA sequences

These are represented by DNA sequences present in the genome in one or a few copies, which are transcribed to yield mRNAs containing information for various proteins (Fig. 4-58).

Fig. 4-58. Distribution of unique, moderately repetitive, and highly repetitive DNA sequences in a hypothetical human chromosome. Unique sequences (3) are transcribed into mRNA. These genes occur in one or a few copies. rRNA and tRNA genes are represented by multiple copies forming clusters in the genome. Large pre-rRNA genes form the nucleolar organizer. Moderately repetitive sequences (2) are distributed throughout the genome, whereas highly repetitive sequences (1) form clusters in the centromeric regions and chromosome ends—telomeres.

A typical human gene consists of approximately 28,000 nucleotides and contains an average of 8 exons. Its coding sequence spans about 1,340 nucleotide pairs and encodes a protein comprising 447 amino acid residues. The largest known gene is the gene for the Muscle protein dystrophin, which consists of 2.4 x 106 nucleotide pairs, whereas the highest number of exons (234) is found in the gene for titin, a fibrous protein responsible for the passive elasticity of skeletal Muscles.

Unique sequences frequently form multigene families, which are organized as clusters in specific regions of one or more chromosomes. Examples of multigene families include the genes for ribosomal, transfer, and Small nuclear RNAs, as well as those encoding α- and β-globins, tubulins, Myoglobin, Actin, transferrin, and many others.

Along with functionally active genes, multigene families contain pseudogenes—mutationally altered sequences that are either incapable of transcription or produce functionally inactive gene products.

Pseudogenes represent a major structural feature of the human genome. These unique sequences bear a striking structural resemblance to specific genes. Because undamaged, functional genes are preserved alongside pseudogenes within gene families, the viability of the organism remains unimpaired. Pseudogenes have been identified for numerous genes; their copy number ranges from one to several dozen per genome, and they are typically arranged in tandem. In some cases, pseudogenes and their corresponding normal genes are located on different chromosomes.

B. Protein Polymorphism

Because most normal human cells are diploid, they contain two copies of each chromosome—one inherited from the father and the other from the mother. These two copies of the same chromosome are referred to as homologous chromosomes (Fig. 4-59). The DNA of each chromosome contains over a thousand genes. Corresponding genes at the same locus on homologous chromosomes are called alleles. Alleles may be identical, meaning they carry the exact same nucleotide sequence. In this case, an individual possessing such alleles is said to be homozygous for that particular trait. If the alleles differ in their DNA nucleotide sequence, the Inheritance of the gene is described as heterozygous. Under these conditions, the individual will produce two protein products of the gene that differ in their amino acid sequence.

While any single individual possesses only two different alleles of a given gene, a human population may harbor a vast array of allelic variants. As mentioned previously, the Variability of DNA Structure—and consequently The Diversity of alleles—arises from mutation processes and recombination events within the homologous chromosomes of germ cells. If recombinations during meiosis involve the exchange of DNA segments smaller than a gene, this process can generate novel alleles that never existed before. Furthermore, because recombination events occur more frequently than mutations in the coding regions of a gene, they are the primary driver of allelic diversity.

Fig. 4-59. Homologous chromosomes and their corresponding allelic protein products. The diagram illustrates the arrangement of 4 alleles (AA, Bb, CC, dD) on homologous chromosomes. Alleles may be identical, as in the case of genes AA and CC, or they may differ (Bb, Dd). The resulting protein products will be identical for the AA and CC alleles, but will differ in amino acid sequence in the case of the Bb and Dd alleles.

The existence of two or more alleles of a single gene within a population is termed "allelomorphism" or "polymorphism," while the protein products generated through the expression of these gene variants are called "polymorphs." Different alleles occur in a population with varying frequencies. Only those variants whose prevalence in the population reaches at least 1% are classified as polymorphisms.

Over the course of evolution, individual genes undergo Amplification to produce multiple copies, while their structure and chromosomal positioning can shift as a result of mutations and translocations—both within the same chromosome and between different chromosomes. Over time, this leads to the emergence of new genes that encode proteins related to the original molecule yet endowed with distinct properties and located at different gene loci (or sites) on the chromosomes.

Related proteins include isoproteins, which are protein variants that perform the same physiological function and are found within the same species. For instance, among a group of 2,000 human genes encoding transcription factors and transcriptional activators, 900 have been identified as belonging to the zinc-finger protein family. Similarly, there are 46 genes encoding the enzyme glyceraldehyde-3-phosphate dehydrogenase, which catalyzes the single oxidation step in the metabolic pathway of Glucose Catabolism leading to Pyruvate.

Scientists have identified families of related proteins that evolved from a single ancestral gene, or precursor gene. Such families include:

✵ the genes for Myoglobin and hemoglobin protomers;

✵ a group of Proteolytic Enzymes: trypsin, chymotrypsin, Elastase, plasmin, Thrombin, and several other related Proteins and Enzymes.

1. Human Hemoglobins

During the course of evolution, single precursor genes gave rise to the α- and β-globin gene families (Fig. 4-60), located on chromosomes 16 and 11, respectively.

Throughout human ontogeny, Different types of hemoglobins are synthesized to ensure optimal adaptation to changing environments. HbE is the embryonic hemoglobin, synthesized During the first months of embryonic development; HbF is the fetal hemoglobin, which sustains subsequent INTRAUTERINE DEVELOPMENT OF the fetus; and HbA and HbA2 facilitate Oxygen transport in the adult body. These proteins are tetramers composed of Two Types of polypeptide chains: α and β in HbA (2α2β), α and ε in HbE (2α2ε), whereas in the remaining hemoglobins the β-chains are replaced by γ-Polypeptides in HbF (2α2γ) or by δ-chains in HbA2 (2α2δ).

Figure 4-60. Schematic arrangement of the α- and β-globin gene clusters. Two copies of the α-globin gene are found: α1 and α2, each directing the synthesis of an α-globin chain. Within the β-globin gene family, the ε locus is expressed during early embryonic development in the initial months of gestation (α2ε2). The γ genes are expressed throughout intrauterine fetal development (HbF, α2γ2). Adult hemoglobins—HbA (α2β2), which accounts for 96%, and HbA22δ2), accounting for 2–2.5%—are produced via the expression of the β and δ genes. Ψβ is a pseudogene featuring a nucleotide sequence homologous to the β gene, but containing mutations that disrupt its expression.

Hemoglobin polymorphism within human populations is remarkably high. Alongside genes encoding isoproteins that occupy different chromosomal loci, a vast number of hemoglobin A variants have been discovered that represent products of allelic genes. Several HbA variants are summarized in Table 4-9.

Table 4-9. Selected variants of human hemoglobin A

Name

Mutation

Abnormal property

HbI

Amino acid substitutions

α16 Lys -> Glu

None

Torino

α43 Phe -> Val

Impaired heme contact; unstable

Nasharon

α47 Asp -> His

Unstable

Buda

α61 Lys -> Asn

Decreased O2 affinity

Iwate

α87 His -> Tyr

Decreased O2 affinity; readily oxidized to MetHb

Denmark Hill

α95 Pro -> Ala

Impaired α1β2 contact; increased O2 affinity

HbC

β6 Glu -> Lys

None

HbS

β6 Glu -> Val

Decreased solubility and O2 affinity

Baltimore

β16 Glu -> Asn

None

Genova

β28 Leu -> Pro

Increased O2 affinity

Zurich

β63 His -> Arg

Increased O2 affinity; unstable

Koln

β98 Val -> Met

Same as above

Kansas

β102 Asn -> Thr

Impaired α1β2 contact; increased O2 affinity

San Diego

β109 Val -> Met

Impaired α1β1 contact; decreased O2 affinity

Hiroshima

β146 His -> Asp

Sharply increased O2 affinity

Leiden

Amino acid deletions

β6 or β7 -> 0

Unstable

Tochigi

β(56-59) -> 0

Unstable

Green Hill

β(91-95) -> 0

Unstable, increased O2 affinity

Tak

Amino acid insertions

Chain elongated by 10 residues at the C-terminus

Increased O2 affinity

Note. Cited from the textbook by A. Ya. Nikolaev, "Biological Chemistry" (Moscow: Vysshaya Shkola, 1989).

One of the most prominent allelic variants of HbA is HbS, which results from the substitution of a glutamate residue at position 6 of the HbA β-chain with valine (β6 Glu -> Val). Based on the HbA and HbS alleles, all people can be divided into 3 genotypically distinct groups: AA, AS, and SS. The global distribution of the S allele is uneven. Individuals carrying this allele are frequently found in malaria-endemic regions of Africa and Asia (up to 35%). To date, over 300 variants of HbA have been described; on this basis, the entire human population can be categorized into 600 genotypic groups according to the most common alleles.

2. Blood Groups

Another important example of protein polymorphism related to blood transfusion is the existence of three allelic Variants of the glycosyltransferase gene (A, B, and 0) in the human population. This enzyme participates in the synthesis of an oligosaccharide located on the outer surface of The Plasma Membrane, which determines the antigenic properties of erythrocytes. The A and B enzyme variants exhibit different substrate specificities: variant A catalyzes the attachment of N-acetylgalactosamine to the oligosaccharide, whereas variant B catalyzes the attachment of galactose. Variant 0 encodes a protein devoid of enzymatic activity. Consequently, the structures of the Oligosaccharides on the erythrocyte surface differ (Fig. 4-61).

Fig. 4-61. Structure of oligosaccharides determining blood groups. The oligosaccharides differ in their terminal monomers. Oligosaccharide A has N-acetylgalactosamine (GalNAc) at the non-reducing end, oligosaccharide B has galactose (Gal), and oligosaccharide 0 is shortened by one monosaccharide residue. R represents a protein or a lipid, specifically a ceramide.

Antibodies against Antigens A and B are typically present in the blood serum of individuals whose erythrocytes lack the corresponding antigen. In other words, individuals with A antigens on their erythrocyte surface produce antibodies against B antigens (anti-B) in their blood serum, while people with B antigens produce antibodies against A antigens (anti-A). In blood serum, anti-A and anti-B are usually present in high titers and, upon encountering the corresponding antigens, are capable of activating The Complement System enzymes.

Blood transfusions are guided by the rule that the blood of the donor and recipient must not contain antigens and antibodies that react with each other: for instance, a recipient whose blood serum contains anti-A must not be transfused with blood from a donor carrying A antigens on their erythrocytes.

Violation of this rule triggers an Antigen-Antibody Reaction, leading to erythrocyte agglutination (clumping) and their destruction by complement enzymes and phagocytes.

As shown in Table 4-10, heterozygous individuals with blood group AB (IV) have both A and B antigens on their erythrocytes and possess two functional variants of glycosyltransferase (A and B); consequently, no Antibodies Are Formed. These individuals can be considered "universal" recipients who can safely receive erythrocytes from Donors of any blood group. However, people with blood group IV cannot safely receive blood serum from these donors, as it contains antibodies against A and/or B antigens. Conversely, individuals with blood group 0 (I) are homozygous for the inactive variant of the glycosyltransferase 0, and The surface of their erythrocytes lacks antigens. Such people are "universal" donors of packed red Blood Cells, as their erythrocytes can be transfused to individuals with blood groups A, B, 0, or AB. At the same time, the blood serum of these donors contains antibodies against A and B antigens and can only be used for patients with blood group 0 (I).

Table 4-10. Characteristics of blood groups

Erythrocyte antigens

None

А

В

АВ

Genotypes

00

АА or АО

ВВ or В0

АВ

Antibodies in blood serum

Anti-A and anti-B

Anti-B

Anti-A

None

Blood groups

0 (I)

А (II)

В (III)

АВ (IV)

Frequency (%)

45

40

10

5

3. Major Histocompatibility Complex Proteins and Transplant Incompatibility

During the formation of a cellular Immune Response, T lymphocytes recognize a foreign antigen only when it is presented alongside Glycoproteins present on the cell's own membrane. These glycoproteins are called major histocompatibility complex proteins, or MHC proteins (see Section 1). There are two classes of these proteins: class I and class II molecules. Class I MHC proteins are found on virtually all nucleated cells, including cytotoxic T cells, whereas class II MHC proteins are found primarily on cells involved in the immune response, such as antigen-presenting B cells and helper T cells, but not on cytotoxic T cells or macrophages.

The structure of MHC proteins is encoded by a family of genes located on the short arm of chromosome 6, spanning a DNA region of more than 6,000 base pairs. This family consists of a series of closely linked genes responsible for the synthesis of MHC proteins and certain Components of the complement system. The genes of this complex exhibit extremely high polymorphism, with the number of different alleles reaching several millions. MHC proteins are considered the most polymorphic system in humans. The variability of MHC proteins is responsible for transplant incompatibility. Transplant cells possess a set of these proteins that differs from the recipient's MHC proteins (in all cases except for genetically identical twins), which triggers a cellular immune response resulting in the rejection of the transplanted tissue.

Studies have shown that the polymorphism of various proteins is so vast that one can speak of the biochemical individuality and uniqueness of every human being.

G. Hereditary diseases

Each genetic locus is characterized by a specific level of variability, i.e., the presence of different alleles in different individuals. Gene alleles are divided into two groups: normal, or wild-type alleles, in which gene function is unimpaired, and mutant alleles, which disrupt gene function. A "bad" allele encodes the synthesis of a protein whose function is severely impaired and, when inherited homozygously, phenotypically manifests as a hereditary disease. Hereditary diseases are the consequence of mutations that occurred in gametes or a zygote. Such mutations can be primary, if they arose in gametes or during zygote formation, or secondary, if the mutant gene arose earlier and was inherited by subsequent generations.

Primary mutations, as a rule, are not accompanied by the onset of disease because they typically occur in only one of the chromosomes, and the individual who receives such a mutation becomes a heterozygous carrier of the gene defect. In the heterozygous state, the mutant gene often does not manifest as a disease and does not significantly reduce the organism's viability, which facilitates its spread in the population.

In the case of secondary mutations, if both parents are heterozygous carriers of the mutant gene, children homozygous for the defective allele may be born. Under such circumstances, a hereditary disease develops, often accompanied by a very severe clinical course.

According to World Health Organization data, about 2.4% of all newborns worldwide suffer from various hereditary disorders. Approximately 40% of early infant mortality and childhood disability are caused by hereditary pathology.

To date, about 800 genes whose mutations lead to various hereditary diseases have been identified on Human chromosomes. The number of Monogenic Disorders (i.e., those caused by mutations in a specific gene) is even greater, at approximately 950, due to the existence of so-called "allelic series"—groups of diseases that are clinically distinct from one another yet caused by mutations in the very same gene. For example, mutations in the ret gene, which encodes a receptor with Tyrosine kinase activity, can cause four different hereditary disorders.

More than half of the genes in which mutations causing hereditary diseases have been found have been characterized using molecular analysis Methods. The largest group consists of enzymes (31% of the total number), followed by proteins that modulate Protein Functions and participate in the proper folding of polypeptide chains (14%).

On average, about 30 structural genes whose mutations cause hereditary diseases are identified on each chromosome. However, these genes are distributed unevenly across chromosomes. For instance, chromosome 2 contains three times fewer such genes than chromosome 1. The highest number of mutant genes (over 100) has been established on the X chromosome.

Thalassemias are well-studied hereditary disorders associated with impaired synthesis of α or β chains of Hb. The synthesis of α and β chains is normally regulated such that all protomer molecules are utilized to synthesize the α2β2 tetramer. Thalassemias arise as a result of mutations involving substitutions or deletions of one or more nucleotides, or occasionally an entire gene encoding the structure of one of the protomers. These diseases are classified into four types: if one of the chains is not synthesized at all, they are designated as α0 or β0 thalassemias, whereas if the synthesis of either chain is reduced, they are designated as α+ or β+ thalassemias.

α-Thalassemias occur when α-chain synthesis is disrupted. The genome of each individual contains four copies of the α-globin gene (two copies on each chromosome); consequently, several types of α-chain deficiency exist. If one of the four copies is defective, there is no phenotypic manifestation, and such an individual is considered a "silent carrier" of thalassemia. When two gene copies are defective, the mutation carrier exhibits mild symptoms of the disease, and hemolytic anemia develops when three copies are defective. Complete absence of α-chain synthesis (i.e., when all four gene copies are defective) results in intrauterine fetal death, because fetal forms of Hb are not formed, while γ4 tetramers have an extremely high affinity for oxygen and are unable to function as transport proteins.

β-Thalassemias develop as a result of reduced synthesis of Hb β chains, for which There is a single gene on each chromosome. The synthesis of HbA begins after birth. When one gene copy is defective, Hb deficiency is mild and does not require special Treatment. However, when β-chain synthesis is completely shut down, a severe form of anemia develops, and such patients require either periodic blood transfusions or Bone Marrow transplantation.

The reader will encounter many monogenic hereditary disorders in almost all subsequent sections of this textbook. At this point, however, it is worth noting that alongside diseases with a clearly pronounced hereditary nature, there are numerous conditions characterized by familial predisposition. These include such widespread disorders as Diabetes Mellitus, Gout, atherosclerosis, Schizophrenia, and a number of others. Unlike Monogenic Diseases, these conditions are classified as multifactorial. Therefore, research aimed at identifying proteins whose allelic variants are responsible for disease susceptibility represents a critical challenge for both the present and the future.



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.