Molecular Biotechnology: Principles and Applications - Glick, B. R., & Pasternak, J. J. 2002

Molecular Biotechnology of Microbiological Systems
Human Molecular Genetics
Physical Mapping of the Human Genome

Genetic and RH maps indicate the linear arrangement of marker sites. Distances between these sites are measured in conventional units that reflect either the recombination frequency (cM) or the probability that two sites will be retained in the same radiation hybrid (cP). These units can be converted (with a certain degree of approximation) into actual physical distance units—nucleotide Base Pairs—which are used in physical maps. A physical map of an entire chromosome or a specific region provides a direct representation of Gene locations on DNA, thereby facilitating their identification, characterization, and the systematic sequencing of chromosomal DNA.

Constructing a physical map first requires isolating clones containing overlapping segments from a genomic DNA library. Based on data regarding the overlapping regions and other information about clone positions, it is possible to reconstruct a continuous series of cloned segments for a specific chromosomal region, an entire chromosome, or the whole genome. Ordered sets of contiguous clones (contigs) have been obtained for human DNA using YAC (Yeast artificial chromosome), BAC (bacterial artificial chromosome), PAC (bacteriophage P1 artificial chromosome), and cosmid libraries. Note that the strategy for constructing contigs from large human DNA fragments contained in YAC, BAC, or PAC libraries differs slightly from that used for P1 or cosmid contigs.

Construction of Contigs from YAC, BAC, and PAC Libraries

When constructing physical maps of specific chromosomal regions or entire Chromosomes from Genomic Libraries containing large inserts (YAC, BAC, or PAC libraries), the mapping method based on sequence-tagged sites (STSs) is most appropriate. An STS is a short, single-copy DNA segment (approximately 100–300 bp) that can be detected via PCR using a unique set of primers. Generating an extensive contig that spans a significant chromosomal region requires A large number of STSs spaced 50–100 kb apart. For instance, physically mapping a chromosome approximately 200 million base pairs (Mb) long requires between 1,500 and 3,000 STSs. Constructing a reasonably accurate physical map of the entire Human Genome requires at least 30,000.

Various approaches have been developed to create STSs. In one method, DNA from a purified preparation of a single human chromosome—isolated by flow cytofluorometry—is digested with a restriction endonuclease and cloned into a vector capable of accepting small (<1000 bp) DNA fragments. Inserts from randomly selected clones are then sequenced; clones with inserts shorter than 100 bp, as well as those containing repetitive human DNA sequences, are discarded. The presence of repeats is determined using computer software that compares The nucleotide sequence of the insert against all known human DNA repeat sequences. Next, primer nucleotide sequences are identified for each selected clone. Each STS is then tested to ensure the amplified chromosomal DNA fragment is unique.

PCR screening is used to detect STSs in individual clones from large-insert libraries. Then, based on the distribution of STSs among the clones, computational Methods are applied to determine the most likely set of overlapping clones and the relative positions of the available STSs (Fig. 20.18). The Development of RH mapping has made it easy to order STSs, which simplifies the identification of contig components (Fig. 20.19). Once overlapping clones are identified, their degree of overlap, contig size, and the total length of the spanned DNA are determined through endonuclease mapping using an electrophoretic system capable of resolving DNA fragments longer than 105 bp (such as pulsed-field gel Electrophoresis). Contigs of chromosomal regions spanning from 1 to over 20 Mb—and in some cases entire chromosomes—have already been generated. Eventually, large-fragment DNA contigs covering the entire genome will be obtained.

Construction of Contigs from Cosmid, P1, and λ Libraries

Contigs assembled from small DNA fragments often prove more convenient for genetic research and large-scale sequencing than those built from large fragments. Cosmid libraries are frequently used to construct contigs of specific chromosomal regions or whole chromosomes. Typically, overlapping cosmid clones are identified using genomic fingerprinting. To achieve this, DNA is extracted from each clone and digested with a restriction enzyme. The resulting fragments are labeled, separated by electrophoresis, and visualized via autoradiography. Each clone produces a specific set of fragments—its unique DNA fingerprint; overlapping clones share one or more identical fragments.

For end-labeling DNA fragments—regardless of whether the ends generated by endonuclease Digestion are 5'- or 3'-overhangs or blunt ends—a replacement reaction catalyzed by T4 DNA polymerase can be used. In this approach, DNA polymerase and a single labeled deoxynucleotide are added to the restriction-enzyme-treated cosmid DNA preparation (Fig. 20.20). The 3'-exonuclease activity of the DNA polymerase sequentially cleaves NUCLEOTIDES from the 3' ends. This process continues until a nucleotide complementary to the labeled deoxynucleotide added to the reaction mixture is exposed on the opposite strand. Then, the polymerase activity of the enzyme kicks in, adding the free labeled nucleotide to the 3' end. Since other nucleotides are absent from the reaction mixture, further chain elongation does not occur. Other methods also exist for labeling endonuclease fragments. The truncated 3' end of a fragment can be extended (filled in) using the Klenow fragment, which utilizes the protruding 5' end as a template; this filling-in is driven by deoxynucleotides added to the reaction mixture, one of which carries a label. Additionally, labeled linkers can be attached to the protruding ends of restriction fragments.

Class="center">

Fig. 20.18. STS mapping. A. STSs (1 through 15) detected via PCR screening in clones a–e are indicated by a plus sign (+). The letters L and R denote STSs located at the 5' and 3' ends of the insert (its left and right ends, respectively). B. By identifying overlapping regions, it is possible to construct a contig from five clones and an STS Location map. The available data do not establish the order of certain STSs (numbers in parentheses). Intervals between STSs are shown as equal; in reality, they are unknown.

Fig. 20.19. Mapping using ordered STSs (numbers 20 through 31 above the horizontal line). Dots indicate the STS composition of clones f–j. The STSs are ordered, making it easy to identify overlapping clones.

Identifying overlapping regions requires analyzing a very large number of cosmid clones; therefore, specialized computer programs are used to search for labeled restriction fragments shared by pairs of clones. The DNA fingerprint of each clone is scanned, the information is entered into a computer, and pairwise comparisons are performed. Based on these comparisons, clones are assembled into contigs. The presence of overlapping regions within the generated contig is verified by constructing detailed restriction maps of the inserts. Gaps between contigs are filled by selecting probes from the outermost ends of neighboring contigs and screening a large-insert library to find the missing DNA segments.

Fig. 20.20. End-labeling of double-stranded DNA using T4 DNA polymerase. The 3'-exonuclease activity of DNA polymerase catalyzes the removal of 3'-terminal nucleotides from DNA fragments with blunt ends (A), 3' overhangs (B), or 5' overhangs (C). Cleavage proceeds until a base complementary to the labeled deoxynucleotide introduced into the reaction mixture (dGTP*) is exposed on the opposite strand; then, the polymerase activity of T4 DNA polymerase "switches on," and a free labeled deoxynucleotide is added to the 3' end. The label is incorporated into both ends of the DNA fragments (not shown in the diagram). Any deoxynucleotide can be used for labeling.

Transcriptional Mapping

Clones from a cDNA library represent DNA copies of the transcripts of expressed genes that were present in a specific tissue at the time mRNA was extracted from it. "Anchoring" individual cDNA clones to chromosomal regions lays the groundwork for identifying potential candidate genes for various diseases. If genetic linkage analysis indicates that a disease gene resides in the same chromosomal region as one or more cDNA sequences, researchers can test whether those cDNA clones originate from the disease gene itself.

Anchoring cDNA clones and Other types of expressed nucleotide sequences to specific chromosomal regions is called transcriptional mapping. Various methods are used to construct transcriptional maps. In one approach, individual cDNA clones are partially sequenced, and STSs are generated based on the translated portion of each cDNA insert. This type of STS is known as an expressed STS (eSTS). To determine the chromosomal localization of eSTSs, somatic Cell hybrid lines containing a single human chromosome (monochromosomal hybrids) or fragments of a specific human chromosome (a deletion panel) are used. To do this, the DNA of each monochromosomal hybrid is amplified by PCR using eSTS primers to identify which chromosome harbors the given eSTS. Subsequently, PCR Amplification of DNA from The Cell lines of the deletion panel pinpoints the specific chromosomal region where the eSTS resides (Fig. 20.21). Furthermore, intragenic STSs can be mapped onto existing RH maps. By 1996, approximately 20,000 human intragenic STSs had been mapped across all autosomes and the X chromosome.

In addition to mapping experiments involving cDNA clones in chromosomal regions, The Institute for Genome Research (Rockville, Maryland) and other laboratories have launched a project to partially sequence cDNA clones from libraries representing all human Organs and Tissues. One goal of this program is to build a catalog of short sequences (150–300 nucleotides) for every expressed human gene. These short coding sequences are called Expressed Sequence Tags (ESTs). They make it possible to study the size, diversity, and transcriptional activity of expressed human genes. Moreover, ESTs can serve as a basis for creating STSs, which can then be used to map and screen for genomic clones containing the gene of interest.

Fig. 20.21. Anchoring a marker to a specific chromosomal region using a hybrid cell deletion panel. A. Schematic representation of chromosomal regions present in a monochromosomal cell hybrid (A) and in hybrid cell lines B–J of the deletion panel. Regions (1 through 10) are defined by the deletion boundaries in the chromosomes of the deletion panel. The filled rectangle represents the centromere of each chromosome. B. Results of PCR amplification of hybrid cell line DNA (A–J) using STS markers (STS-a through STS-e). The presence or absence of a PCR product is indicated by a plus or minus sign, respectively. Data on the presence or absence of PCR products from deletion panel cell lines with an STS marker are used to determine the chromosomal region harboring the STS. For instance, STS-d is assigned to region 7 because the corresponding PCR product is generated upon amplification of every deletion panel cell line that contains region 7.

Partial sequencing of cDNA clones and the Processing of resulting data are fully automated. Each new EST is compared against previously sequenced ones; if the sequence is genuinely novel, it is added to the EST database. Additional comparisons are performed to search for homologies between ESTs and known genes or gene families, and to categorize the function of the gene it represents. By 1995, approximately 300,000 ESTs had been identified from 300 cDNA libraries across 37 organs and tissues. Roughly 90,000 ESTs represent distinct expressed human sequences, of which approximately 10,000 correspond to genes with known cellular Functions, while the remaining 80,000 correspond to as-yet- undiscovered genes.



Last update: 11/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.