Molecular Biotechnology: Principles and Applications - Glick, B., & Pasternak, J. 2002
Molecular Biotechnology of Microbiological Systems
Human Molecular Genetics
Construction of Human Chromosome Genetic Maps
Genetic polymorphism Linkage between the ABO locus and the Gene responsible for hereditary nail-Patella syndrome (NPS) was discovered for two main reasons. First, each of the major alleles of the ABO system (IA, IB, IO) can be precisely identified using a simple laboratory test, meaning that the genotypes of all parents and children studied are known with certainty. Second, each ABO allele occurs at a high frequency within the population, resulting in a fairly high probability that the parents will be heterozygous. In the United Kingdom, where the initial studies on ABO-NPS linkage were conducted, the frequencies of the IA, IB, and IO alleles are approximately 0.66, 0.28, and 0.06, respectively.
The term "allele frequency" denotes the proportion of a specific allele among all alleles of a given locus in a population. For example, for a two-allele locus (A1, A2) in a population of 13,000 individuals—where 3,800 have the A1A1 genotype, 6,400 are A1A2, and 2,800 are A2A2—the frequency of the A1 allele is
Class="center">![]()
whereas the frequency of the A2 allele is
![]()
For the majority of loci, the frequency of one allele (≥0.999) vastly exceeds that of the other(s) (≤0.001). Consequently, in large populations, the vast majority (99.8%) of individuals are homozygous for the more common allele, about 0.198% are heterozygous, and 0.001% are homozygous for the rare allele. Under such conditions, it is practically impossible to establish the segregation of alleles at this locus or its linkage with another locus, since most parents will be homozygous for the common allele. Conversely, if the frequencies of two alleles at a given locus are 0.99 and 0.01, approximately 2% of individuals will be heterozygous, and the chances of detecting segregation or linkage increase because the population contains many individuals heterozygous for that locus (Table 20.3). Thus, linkage studies in humans are feasible only for loci with common alleles. When two or more alleles of a given locus occur in a population at a frequency of 0.01 or higher, genetic polymorphism is said to exist, and the locus is termed polymorphic. Because genetic polymorphisms similar to that of the ABO Blood group system are rare, designing chromosome mapping projects requires The Development of Methods that allow large numbers of polymorphic sites to be easily detected.
Table 20.3. Allele and genotype frequencies in a large randomly mating population1)
|
Allele frequencies |
Genotype frequencies |
|||
|
A1 |
A2 |
A1A1 |
A1A2 |
A2A2 |
|
1,0 |
0 |
1,0 |
0 |
0 |
|
0,999 |
0.001 |
0.998001 |
0.001998 |
0,000001 |
|
0,99 |
0,01 |
0.9801 |
0,0198 |
0,0001 |
|
0,90 |
0,10 |
0,81 |
0,18 |
0,01 |
|
0,75 |
0,25 |
0,5625 |
0,3750 |
0,0625 |
|
0,50 |
0,50 |
0,25 |
0,50 |
0,25 |
|
0,25 |
0,75 |
0,0625 |
0,3750 |
0,5625 |
|
0,10 |
0,90 |
0,01 |
0,18 |
0,81 |
|
0,01 |
0,99 |
0,0001 |
0,0198 |
0,9801 |
|
0,001 |
0,999 |
0,000001 |
0,001998 |
0,998001 |
|
0 |
1,0 |
0 |
0 |
1,0 |
1) Allele frequencies corresponding to polymorphism are highlighted in gray.
Restriction Fragment Length Polymorphism
For new alleles to arise, it is sufficient that two homologous genes differ by just a single nucleotide. In many cases, a single-nucleotide substitution leads to significant differences between the altered gene product and the normal protein. However, many single-nucleotide changes do not result in altered gene products; moreover, substitutions can occur within non-coding regions of DNA with no phenotypic consequences whatsoever. Such "silent" substitutions, distributed throughout the length of the chromosome, generate polymorphic sites (marker loci, genetic markers) that can be employed for genetic mapping. First, however, these polymorphic sites must be detected.
In 1980, D. Botstein, R.L. White, M.H. Skolnick, and R.W. Davis established the theoretical framework for identifying single-nucleotide polymorphic sites and utilizing them as markers for constructing human chromosome maps. The underlying methodology is as follows. Restriction Endonucleases (restriction Enzymes) cleave DNA at specific sites. When a single-nucleotide substitution occurs within such a site, the restriction enzyme fails to cleave it, yet still recognizes and cleaves the intact site on the alternative chromosome (Fig. 20.11A). Because one allele contains the recognition site for a given restriction enzyme while the other does not, treating the DNA with this enzyme yields fragments of different lengths. The presence or absence of a polymorphic restriction site can be determined by hybridizing the DNA with a probe strictly specific for a unique chromosomal region.

Fig. 20.11. Use of restriction sites as genetic markers. A. A single base-pair substitution within a restriction site prevents its recognition by the restriction enzyme, thereby blocking DNA Cleavage. Intact and altered restriction sites are indicated by plus (+) and minus (-) signs, respectively. B. A segment of a chromosome containing three sites (A, 1, and B) recognized by the same restriction enzyme. X is the distance between sites A and 1, Y is the distance between sites 1 and B, and X + Y is the distance between sites A and B. Sites A and B are intact in all cases (both +), whereas site 1 can be either intact (+) or altered (-). If site 1 is intact (+), restriction enzyme Digestion produces fragments X and Y. If it is altered (-), a single fragment (X + Y) is generated. Performing Southern blot Hybridization with probe a reveals fragment X if site 1 is intact (+), and fragment (X + Y) if it is altered (-). C. Restriction fragments and fragments detected by Southern hybridization with probe a for each genotype (+/+, +/-, -/-). D. Fragments of a single chromosome generated by restriction digestion and detected by hybridization with probe β or γ.
Suppose, for instance, that a particular chromosomal segment contains three sites recognized by the HindIII restriction enzyme (Fig. 20.11B), with sites A and B being intact in all individuals. This means that the population lacks alternative alleles at these sites, i.e., polymorphism is absent. In contrast, site 1 frequently exhibits a single-nucleotide substitution that renders it resistant to HindIII cleavage. Thus, two Chromosomes in the population differ at this locus: one is cleaved (+), while the other is not (-).
If the distance from site A to site 1 and from site 1 to site B do not exceed 20 kb each, and a single-copy probe is available that hybridizes to the DNA segment between sites A and 1 (Fig. 20.11B), then following Southern blot analysis and agarose gel Electrophoresis of the HindIII-digested DNA fragments, we can distinguish between two scenarios. First, site 1 is cleaved, producing two fragments, and the probe hybridizes to the one flanked by sites A and 1. Second, site 1 is not cleaved, and the probe hybridizes to the DNA fragment flanked by sites A and B (Fig. 20.11B).
Analyzing actual DNA samples is somewhat more complex because chromosomes exist in pairs (Fig. 20.11C). Even so, each genotype (+/+, +/-, -/-) corresponds to a characteristic pattern of fragments generated upon hybridization with the probe. Furthermore, to detect the restriction site at region 1, one can use probes that hybridize to other DNA segments located between sites A and B (Fig. 20.11D). The phenomenon wherein the presence of a common altered restriction site in the population results in a specific set of DNA fragments is termed restriction fragment length polymorphism (RFLP). Polymorphic restriction sites form marker loci on the chromosome where they reside.
The genetic status of each RFLP locus on a single chromosome is referred to as a haplotype. For a single site, There are two possible haplotypes (+ or -); for two different sites, four (++, +-, -+, and --); for n loci, the number of haplotypes equals 2n. Determining the RFLP alleles (or any other polymorphic loci) present on an individual's chromosomes is known as haplotyping (genotyping or DNA typing). RFLP loci are inherited in a Mendelian fashion, allowing their transmission to be tracked within pedigrees. Tracking the inheritance of two or more RFLP loci within a given family makes it possible to detect recombination. Figure 20.12 illustrates the following scenario: the father (I-1) is heterozygous for three different RFLP loci located on the same chromosome, whereas the mother (I-2) lacks restriction sites at all three loci under consideration. The genetic status of the chromosome inherited from the father by each child can be determined through genotyping. Son II-2 received a recombinant chromosome from the father; the other children inherited non-recombinant chromosomes from him. In practice, the DNA of each individual is treated in a separate tube with various restriction enzymes and then hybridized with cloned single-copy DNA fragments used as probes to detect RFLPs. In situ hybridization of a DNA probe with human metaphase chromosome spreads on a Microscope slide makes it possible to determine whether a given probe corresponds to a unique chromosomal region (single-copy DNA). A standard nomenclature has been developed to designate thousands of RFLP loci. For example, the designation D21S18 corresponds to a locus identified using a DNA probe (D) that hybridizes to chromosome 21 (21), is present in a single copy (S), and was registered by the DNA Committee of the International System for Human Linkage Maps (ISLM) under the number 18 (18). Polymorphic marker sites located within known genes are named after the respective gene. For instance, ADH designates a polymorphic locus within the Alcohol dehydrogenase gene (ADH1). There is no standard designation for RFLP alleles. Some laboratories number alleles sequentially (D1S34*1, D1S43*2, etc.), while others name them According to the fragment length (in kb) generated in the presence or absence of the restriction site(s) (D4S56*8, D4S56*12, D4S56*4, etc.).

Fig. 20.12. Detection of segregation and recombination of RFLP loci within a pedigree. Plus (+) and minus (-) signs denote alleles containing intact and altered restriction sites, respectively, for three RFLP loci (A, B, C) located on the same chromosome. The genetic status of I-1 is inferred from the genotypes of his parents, +++/+++ and ---/--- (not shown). Determining the haplotypes of the offspring demonstrates that II-1 received a recombinant chromosome from I-1, whereas II-1 and II-3 received non-recombinant ones. The vertical bar separating the sets of RFLP alleles beneath the symbol for each family member separates homologous chromosomes.
Thousands of RFLP loci have already been identified, significantly expanding the array of alleles available for genetic research. Linkage between an RFLP locus (or loci) and a disease gene can be established by calculating two-point (two-locus) lod scores for haplotyped families displaying cases of the genetic disorder under study. Similarly, linkage between RFLP loci can be detected by analyzing multi-generation pedigree data. However, using RFLPs for mapping has several limitations. These loci are distributed unevenly across chromosomes, maintaining clone probes is inconvenient, and haplotyping large numbers of individuals from multiple families via Southern blotting is quite labor-intensive. Fortunately, The Human Genome harbors a vast Abundance (>100,000) of other polymorphic loci comprising simple repeating units of two, three, or four Base Pairs—known as short tandem repeats (STRs)—which are readily detected using the Polymerase Chain Reaction (PCR).
Short Tandem Repeat Polymorphism
Approximately 100,000 blocks of CA/GT dinucleotide repeats [(CA) ∙ (GT)] (Fig. 20.13), containing anywhere from 1 to 40 repeating CA/GT units, are evenly distributed throughout the human genome. Any such block located at a specific chromosomal site is transmitted from generation to generation while maintaining its repeat number. CA/GT repeats are commonly designated as (CA)n, where n represents the number of CA repeats. The human genome also contains other dinucleotide repeats [e.g., (AT)n, etc.], as well as trinucleotide [(ATG)n, etc.] and tetranucleotide [(ATCG)n, etc.] repeats. Identifying polymorphic STR loci first requires screening a small-insert human genomic library (approximately 1 kb inserts) using an appropriate oligonucleotide probe. Cloned (CA)n repeats are typically identified using a probe composed of 15 CA units. Each positive insert is sequenced to determine the length of the CA repeat and The nucleotide sequences of its flanking regions. To establish whether these flanking sequences are single-copy, in situ hybridization is performed with complementary probes; if the sequences are found to occur only once in The Genome, a pair of complementary primers is synthesized to amplify the CA repeat. Subsequently, PCR testing of DNA obtained from a large cohort of individuals is carried out using this primer pair. The PCR products, whose length is purposefully chosen to be around 200 bp to facilitate electrophoretic Separation, are resolved in a polyacrylamide gel. If the length of the amplified DNA segment is identical across all DNA samples, the repeat is non-polymorphic (Fig. 20.14A); conversely, if PCR products of varying lengths are generated, it indicates polymorphism at that STR (STR polymorphism, STRP) (Fig. 20.14B). CA repeats of varying lengths at a given locus represent distinct alleles (Fig. 20.15). Such alleles frequently occur at frequencies of 0.20 or even higher.
![]()
Fig. 20.13. Dinucleotide tandem repeat (CA)24 containing 24 repeating units.

Fig. 20.14. STR locus typing. A. DNA obtained from different individuals (n = 7) was amplified by PCR using a primer pair (X) flanking a (CA)-(GT) repeat. The size of all resulting PCR products is identical (lanes 1–7), and therefore the STR length is also identical. Based on these data, the STR locus is represented by a single allele. B. Same as in Fig. A, but using a different primer pair (Y) for another STR locus. The formation of two distinct PCR products indicates that this locus is represented by two alleles. Lanes 1–3 correspond to the amplified DNA fragment from individuals homozygous for a single STR allele, lanes 4 and 5 correspond to the amplified DNA fragment from heterozygous individuals carrying two different STR alleles, and lanes 6 and 7 correspond to the amplified DNA fragment from individuals homozygous for another STR allele.
To date, thousands of STRP loci have already been discovered. The same nomenclature rules are used for them as for RFLP loci. At the same time, the names of STRP primers often differ from the names of the loci. Many STRP loci were identified by French researchers with financial support from the French Muscular Dystrophy Association (Association Française contre les Myopathies), which is reflected in the fact that designations for many STRP primer pairs begin with the abbreviation AFM followed by an identification number (AFM349xc5). The designation of a primer pair is often accompanied by the designation of the corresponding locus [AFM349xc5 (D3S2017)].

Fig. 20.15. Two STR alleles. One of them (allele 1) contains the (CA)15 repeat, while the other (allele 2) contains (CA)10. In both cases, the repeats are flanked by identical unique sequences.
Currently, human Genome Mapping relies primarily on STRPs rather than RFLP loci. Unlike RFLP probes, which must be cloned in a vector, purified, and labeled, STRP loci require only nucleotide sequence information for a primer pair, which can be stored in a computer database. Furthermore, STRP loci are evenly distributed throughout the human genome; STR allele frequencies are very high, ensuring high heterozygosity; and the alleles themselves are easily identified following PCR Amplification.
Last update: 11/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.