Molecular Biotechnology: Principles and Applications - Glick, B. R., & Pasternak, J. J. 2002

Molecular Biotechnology of Microbiological Systems
Human Molecular Genetics
Cloning Human Disease Genes

As a rule, a specific human disease Gene cannot be cloned simply by following a predefined set of experimental protocols. The choice of Methods and tools available to a researcher depends on the specific conditions. The search for a disease gene typically begins with the available information about its product. In some cases, the gene product is well known; in others, one can only guess what it might be. Finally, for many Hereditary diseases, The Nature of the gene product is entirely unknown. A dedicated strategy has been developed for each of these scenarios. Overall, there are four main approaches to finding a disease gene: functional cloning, candidate Gene cloning, positional cloning, and positional-candidate cloning. Regardless of the approach used, one can only claim that a given gene is associated with the disease of interest after nucleotide alterations are found in the gene of affected individuals, which are absent in the same gene in healthy subjects.

Identification of Mutations in Human Genes

To detect mutations, a variety of simple and inexpensive approaches have been developed, such as single-strand conformational polymorphism (SSCP) analysis, denaturing gradient gel Electrophoresis (DGGE), heteroduplex analysis (HA), chemical mismatch Cleavage (CMC), and the protein truncation test (PTT).

Among these approaches, SSCP is the most widely used. The Principle of the method is as follows. As many exons as possible (ideally all) of the target gene are separately amplified by PCR using DNA from both affected and healthy individuals as a template. Each primer pair is selected from the sequences flanking the exon or from its terminal regions. In addition, sequencing data are used to select primers for the Amplification of the 5' region preceding the first exon of the gene, the 3' region following the last exon, and regions containing splice sites.

The PCR products of each reaction are denatured, rapidly cooled, and separated by electrophoresis. Due to intrastrand base pairing and the Formation of other bonds, a denatured single-stranded DNA molecule adopts a specific three-dimensional conformation determined by its nucleotide sequence. Because of complementarity, the two strands of a single DNA molecule have different nucleotide sequences and therefore adopt different three-dimensional Conformations, migrating at different rates during gel electrophoresis. As a result, two bands corresponding to the complementary strands are observed after gel Separation. If two DNA molecules representing the same gene segment but derived from different sources differ by a single base pair, the conformations of their single strands are very likely to differ. In other words, each of the four strands will migrate at its own characteristic rate during gel electrophoresis (Fig. 20.22). The SSCP method only allows for the localization of nucleotide changes within a specific exon or region of a gene, rather than determining the exact Nature of the mutation; such information can only be provided by sequencing. SSCP has certain limitations: it detects approximately 90% of single-nucleotide changes in PCR products that are no longer than 200 bp.

Class="center">

Fig. 20.22. Single-strand conformational polymorphism (SSCP) analysis. DNA preparations differing by a single base pair (A:T ↔ G:C) are amplified by PCR using identical primers (P1, P2). The PCR products are denatured and separated by gel electrophoresis in two lanes (1, 2). The migration distance of a single-stranded DNA molecule depends on its conformation, which in turn is determined by its nucleotide sequence. Even if the DNA molecules differ by only a single nucleotide site, the single strands may adopt different conformations and, consequently, the PCR products form four bands in the gel instead of two.

Functional Cloning

Functional cloning of a gene begins with determining the Amino Acid Sequence of a protein with a known function, which allows the reconstruction of The nucleotide sequence of the coding region of the corresponding gene (target gene). Based on this data, oligonucleotide probes are synthesized to screen a cDNA library derived from a tissue where the protein is abundant. If purified mRNA encoding the protein can be obtained, it can serve as a template to synthesize and clone a full-length cDNA. The accuracy of the selected or synthesized cDNA clone is verified by sequencing.

The chromosomal localization of the target gene is determined by in situ Hybridization using the cDNA clone or a genomic clone selected with its aid. For more precise localization, a panel of monospecific Cell hybrids and subsequently a deletion panel can be screened using the cDNA clone or the selected genomic clone.

Next, a cosmid contig spanning the chromosomal region containing the target gene is screened to identify clones that hybridize with the cDNA clone. The selected genomic clones are sequenced, and exons, introns, and the 5' and 3' flanking sequences of the gene are identified using the cDNA sequence data. If a cosmid contig covering the required chromosomal region containing the target gene is not available, a clone with a large insert encompassing this region is isolated using a cDNA or genomic probe. Subclones with smaller inserts are then generated from the large-insert clone and screened with the cDNA probe. Positive clones are sequenced and the target gene is characterized (Fig. 20.23).

Fig. 20.23. Functional cloning. Identification of a gene when The amino acid sequence of its product is known.

Candidate Gene Cloning

Although this approach is not very efficient for mapping human genes, it can be extremely useful in certain cases. THE PRINCIPLE OF the method is as follows. The symptoms of a genetic disorder are analyzed to deduce the type of protein likely to be associated with it. Next, The nucleotide sequences of all currently cloned genes are scanned to select candidate gene(s). Based on The sequence of the candidate gene, a mutation-screening strategy is developed to determine whether the candidate gene is indeed the disease gene of interest (Fig. 20.24). Given that The Human Genome contains a vast number of genes while only a fraction of them have been characterized, it is hardly surprising that the correct gene is not always identified on the first try. However, a negative result is also valuable, as it allows researchers to exclude that gene from the list of candidates responsible for the specific genetic disorder.

Positional Cloning

The positional cloning strategy is applied when nothing is known about the product of the gene responsible for an inherited disease and there are no candidate genes available (Fig. 20.25). In such cases, the chromosomal Location (position) of the disease gene is determined, and the gene is tracked down using various tools and methodologies ("gene hunting"). Positional cloning is greatly facilitated when several affected individuals exhibit a chromosomal rearrangement such as a translocation or a large deletion (> 10 kb). Assuming that this aberration disrupts the gene responsible for the pathological phenotype, linkage analysis can focus on a single specific chromosome region rather than scanning the entire genome using a large set of polymorphic markers. Once the gene has been localized to a specific chromosomal region, its position is refined, and the closest flanking markers are identified using multilocus mapping with additional polymorphic probes. The minimum resolvable distance between mapped marker sites is, at best, about 1 cM, which corresponds to approximately 106 bp. Such a segment typically harbors between 20 and 50 genes. The goal of positional cloning is to determine which of these genes is responsible for the disease in question.

Fig. 20.24. Candidate gene cloning. Identification of a gene based on the Analysis of the symptoms of the associated disease and reasoning regarding which of the already characterized genes might be the target gene.

Fig. 20.25. Positional cloning. Identification of a gene whose product is unknown, using chromosomal mapping and probes specific for tightly linked markers.

Genomic clones that encompass the flanking markers and the intervening DNA segment are selected from a contig spanning the disease gene region. If such a contig is unavailable, genomic DNA libraries are screened using probes specific for tightly linked marker sites to identify clones originating from the region containing the target gene. Between 1986 and 1990, when human gene-hunting techniques were still in their infancy, chromosome jumping (Fig. 20.26) or chromosome walking (Fig. 20.27) methods were used to identify the necessary genomic clones. Following the establishment of Genomic Libraries containing large human DNA fragments and contigs, these strategies became largely obsolete. Regardless of how the required genomic clones are obtained, it is crucial to determine which of the clones or subclones contain exons. A range of Direct and Indirect methods can be employed for this purpose, including CpG island identification, cross-species Southern blotting, hybrid Selection, exon trapping, DNA Sequencing, and computational database searches.

Fig. 20.26. Construction of a library using the chromosome jumping method. Genomic DNA is partially digested with a restriction enzyme that makes infrequent cuts, and fragments of approximately 200 kb are isolated. These fragments are ligated with the supF gene (~7 kb) and circularized. Letters A and B, X and Y, S and T denote sites initially located 200 kb apart from each other, but separated by 7 kb after fragment circularization. The circular molecules, which contain numerous EcoRI sites, are digested with this restriction enzyme, producing, among other products, fragments containing the supF gene and its flanking sequences (A and B; X and Y; S and T). Only fragments of approximately 20 kb in length are selected and cloned into a λ phage vector. Vectors carrying the supF gene are amplified in SupF host Cells. A clone that hybridizes with a probe specific for the original sequence (A, X, S) is identified and then subcloned; the portion of the subclone that does not hybridize with the probe contains the DNA segment (B, Y, T) located 200 kb away from the original sequence.

Fig. 20.27. Chromosome walking. A. Probe 1 is hybridized with a cloned 40-kb DNA fragment. Following subcloning and restriction mapping, the sequence distal to the hybridization site is used to generate probe 2. B. Probe 2 is used to isolate a different clone (distinct from clone 1) from a genomic library, and the sequence distal to the hybridization site of probe 2 is used to generate probe 3. Clones 1 and 2 together span approximately 80 kb (minus the overlapping region corresponding to probe 2 between them). C. Procedures analogous to steps A and B are performed using probe 3. This third clone in the "walk" extends the coverage of the chromosome by another 40 kb. D. Three overlapping DNA fragments span approximately 120 kb of chromosomal DNA. Chromosome walking can proceed in both directions, guided by the restriction map.

Transcribed regions of vertebrate genomes are frequently preceded by nucleotide clusters rich in C and G residues (CpG islands). A group of CpG islands can be identified by the clustering of restriction sites for the endonucleases EagI, BssHII, and SacII on a restriction map. If at least two such sites are separated by 5–10 kb, they are likely located within a CpG island. Although this does not guarantee the presence of an exon at this exact location, it indicates that a transcribed gene resides nearby.

Genomic clones or subclones can be used for Southern blot hybridization with restricted genomic DNA from various vertebrates, such as mouse, rat, rabbit, monkey, cow, chicken, and fish (zoo blot, interspecies blotting, or "Noah's Ark" blotting). Positive cross-hybridization strongly suggests that a given clone contains coding sequences, as many exons have been evolutionarily conserved, whereas repetitive and noncoding DNA sequences, including introns, have undergone significant divergence. A positive zoo blot indicates that the clone contains exon(s), though it does not prove whether it harbors the disease gene of interest. Hybrid selection enables the rapid and highly efficient identification of genomic clones containing exon(s) while simultaneously isolating the corresponding cDNA. This selection can be performed in several ways. Typically, the DNA of a genomic clone from the chromosomal region containing the disease gene is immobilized on a solid support, prehybridized with repetitive DNA sequences, and then hybridized with linear vector molecules carrying inserts from a cDNA library derived from a tissue likely to express the target gene. Unhybridized vector molecules are washed off the filter, while hybridized ones are eluted and amplified by PCR using primers complementary to vector sequences flanking the cDNA insert (Fig. 20.28). If the tissue expressing the target gene is unknown, cDNA libraries from multiple Tissues can be pooled and hybridized with individual genomic clones. The PCR product can then be cloned and tested to determine whether it contains the coding region of the disease gene. This can be accomplished by sequencing the cDNA and performing a computer-assisted sequence comparison with known genes. If a high degree of Homology is found, inferences can be made about the type of protein encoded by the cDNA. If this protein is a likely candidate for the target gene product, the clone(s) are sequenced to identify exons, introns, and 5' and 3' flanking regions. An alternative approach is mutation screening to detect nucleotide differences between the DNA of affected and unaffected individuals. If Structure/154.html">Sequence homology comparisons fail, other genes from the region are sequenced and analyzed. In practice, to save time and resources, multiple genes in the vicinity of the target are initially characterized until the most probable candidate gene is identified, which is then subjected to detailed investigation, including mutation analysis.

Fig. 20.28. Hybrid selection. A cDNA clone is "captured" via hybridization with genomic clone DNA, amplified, cloned, and tested.

Exon trapping (exon capture or exon amplification) is a method used to identify and clone exons residing within subclones derived from genomic clones (Fig. 20.29). The procedure is based on the following principle. Genomic clone DNA is cleaved into 1- to 6-kb fragments, which are then cloned into a specially designed vector. The multiple cloning site (polylinker) of the vector is positioned within an intron flanked by two exons (exon 1 and exon 2). This artificial gene (exon 1–intron–exon 2) is placed under the control of a strong eukaryotic promoter and can replicate in E. coli or mammalian cell culture. Following Introduction (transfection) of the vector without an insert into a mammalian cell, the artificial gene is transcribed, and the intron is removed from the primary transcript. The processed RNA (exon 1–exon 2) can be detected by reverse METABOLISM/31.html">Transcription PCR (RT-PCR). First, mRNA is used as a template to synthesize single-stranded DNA using Reverse Transcriptase. The second strand is synthesized using a primer complementary to a portion of exon 1. A second primer, complementary to a portion of exon 2 of the second-strand DNA, is then added to the reaction mixture, and PCR amplification is performed. The length of the PCR product is determined by gel electrophoresis.

Fig. 20.29. Exon trapping. A. The exon trapping vector contains an artificial gene consisting of a promoter P, two exons separated by an intron containing a polylinker, and a transcription termination site t. Following introduction of the vector into a Eukaryotic Cell, the artificial gene is transcribed, and the intron is excised from the primary transcript. RT-PCR amplification is used to generate a PCR product of defined length containing portions of both exons.

Fig. 20.29. (Continued) B. A human DNA fragment is inserted into the polylinker of the exon trapping vector, and the construct is introduced into a eukaryotic cell (transfection). In this case, the exon contains functional acceptor and donor splice sites. During Processing of the primary transcript, the introns flanking exon A are excised, placing it between exon 1 and exon 2. The length of the RT-PCR product indicates whether the target exon has been "trapped" between exon 1 and exon 2 and, consequently, whether the insert contains that exon.

If a restriction fragment containing a given exon A and its flanking introns is inserted into the vector, the processed transcript following transfection will contain three exons: exon 1–exon A–exon 2. The resulting PCR product will be larger than in control cases where the vector lacks an insert, contains an insert with an exon lacking functional splice sites (donor and acceptor), or contains no exon at all. If the insert contains more than one exon, each with functional splice sites, the processed transcript will include all of them.

If the primers contain restriction sites, the PCR product carrying the "trapped" exon can be cloned and used as a probe to screen a cDNA library. Once the nucleotide sequence of the trapped exon is known, database searches are performed to identify homologous sequences. If there is strong evidence that the trapped exon is part of the disease gene, genomic clones spanning the gene's locus are characterized and sequenced, and DNA samples from affected and unaffected individuals are analyzed for mutations. Because disease-causing mutations are not always evenly distributed across all exons, scanning a larger portion of the putative gene's coding region increases the likelihood of detecting a mutation.

Various computer programs, such as GRAIL (Gene Recognition and Analysis Internet Line), are employed to identify exons. These programs rely on specific sequence features characteristic of exons, such as the expected codon usage of coding regions. Laboratories equipped for high-throughput sequencing can sequence genomic clones spanning the candidate gene region and use computational tools to analyze the data for putative exons. The nucleotide sequence of a predicted exon can then be used to search sequence Databases for homologs or to synthesize an oligonucleotide probe for screening a cDNA library. Ultimately, as with any exon-trapping method, experimental validation is required to prove that the candidate exon is indeed part of the target gene.

Positional cloning projects are notoriously time-consuming. Between 1986 and 1995, this approach led to the discovery of over 50 human disease genes, which was a remarkable achievement. While some gene searches took only 1 to 2 years, isolating the Huntington's disease gene required a decade of collaborative effort by multiple research consortia. Notably, as more genes are cloned and high-resolution transcriptional maps become available, positional cloning is gradually being superseded by positional candidate cloning.

Positional Candidate Cloning

Positional candidate cloning involves determining the chromosomal localization of a disease gene of unknown product and subsequently analyzing current genetic and transcriptional maps to identify coding sequences (genes, intragenic ESTs) residing within the same chromosomal interval (Fig. 20.30). It is highly probable that one of these candidate sequences represents the disease gene. Once a candidate gene is characterized, mutation analysis can be performed. Alternatively, candidate ESTs can be used as probes to isolate a genomic clone, which is then sequenced and subjected to mutation analysis. As physical and transcriptional maps become increasingly detailed, positional candidate cloning is gaining widespread popularity in the search for human disease genes.

Fig. 20.30. Positional candidate cloning. Identification of a disease gene when its product is unknown, but the gene has been mapped to the same chromosomal region as known functional genes and ESTs. Candidate genes and ESTs are selected from this set and evaluated to determine which one corresponds to the target gene.



Last update: 11/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.