LEHNINGER PRINCIPLES OF BIOCHEMISTRY - VOL. 1. THE FOUNDATIONS OF BIOCHEMISTRY: STRUCTURE AND CATALYSIS - 2011
PART I. STRUCTURE AND CATALYSIS
9. DNA-BASED INFORMATION TECHNOLOGY
9.2. From Genes to Genomes
Today, the modern scientific discipline of Genomics allows The Study of DNA on a cellular scale, ranging from individual genes to an Organism’s complete genetic Complement—its genome. Genomic Databases are growing rapidly as successive sequencing milestones are reached. Biology in the 21st century is propelled by informational resources that were scarcely imaginable just a few years ago. We now turn to some of the core technologies that drive this progress.
DNA Libraries Provide Specialized Catalogs of Genetic information
A DNA library is a collection of DNA clones gathered together as a source of DNA for sequencing, Gene identification, or functional studies. Libraries can be of several types, depending on the DNA source. One of the largest types is a genomic library, which is constructed by breaking an organism's entire genome into thousands of fragments, all of which are cloned by insertion into a cloning vector.
The first step in generating a genomic library is the partial Digestion of DNA with Restriction Endonucleases, ensuring that any given sequence is represented in fragments of a specific size range—a range compatible with the cloning vector and sufficient to guarantee that virtually all sequences are present within the library's clones. Fragments too large or too small for cloning are removed by centrifugation or Electrophoresis. A cloning vector, such as a BAC or YAC plasmid, is cut with the same restriction endonuclease and ligated to the genomic DNA fragments. The resulting pool of joined DNA is then used to transform bacterial or Yeast Cells, generating a library of Cell types in which each individual type harbors a distinct recombinant DNA molecule. Ideally, all the DNA of the studied genome is represented in the library. Each transformed bacterial or yeast cell grows into a colony, or “clone,” of identical cells, every cell containing the exact same recombinant plasmid—one of many represented throughout the entire library.
Using Hybridization techniques, researchers can order individual clones within a library by identifying clones with overlapping sequences. A set of overlapping clones represents a contiguous catalog of an extended genomic segment and is commonly referred to as a contig (Fig. 9-13). Previously established nucleotide sequences of entire genes can be located within a library via hybridization by determining which specific clones contain The sequence of interest. Furthermore, if a sequence has already been mapped to a chromosome, researchers can pinpoint the chromosomal Location OF THE cloned DNA and any contig of which it is a part. Well-organized libraries can contain thousands of long contigs assigned to specific Chromosomes and ordered accordingly, forming a detailed genomic map. Known sequences within the library (each termed a sequence-tagged site, or STS) can serve as landmark reference points for genome-sequencing projects.
Class="center">Fig. 9-13. Ordering clones in a DNA library. Shown is a segment of a hypothetical organism X chromosome with markers A through Q representing sequence-tagged sites (STSs—stretches of DNA with known nucleotide sequences, including established genes). Below the chromosome is a set of ordered BAC clones numbered 1 through 9. Ordering clones on a genetic map is a multistage process. The presence or absence of an STS in a particular clone can be determined by hybridization—for example, by probing each clone with DNA derived from the STS via PCR. Once all STSs in each BAC clone are identified, the clones (and the STSs themselves, if their exact locations were previously unknown) can be arranged in a definitive order. For instance, comparing clones 3, 4, and 5: marker E (blue) is found in all three clones; F (red) is present in clones 4 and 5, but absent in 3; and G (green) is found only in clone 5. This establishes the order of the sites as E, F, G. The clones partially overlap, determining their sequence to be 3, 4, 5. The final ordered set of clones is termed a contig.

As more and more genomic nucleotide sequences become available, the reliance on Genomic Libraries is gradually decreasing, and researchers are turning toward more specialized libraries designed specifically for studying gene function. One prominent example is a library that includes only those genes that are expressed—that is, transcribed into RNA—in a given organism, cell type, or tissue. Such a library lacks the noncoding sequences that make up a substantial fraction of many Eukaryotic Genomes. To construct it, mRNA is first isolated from an organism or specific cells, and complementary DNAs (cDNAs) are then synthesized from the RNA through a multistage reaction catalyzed by the enzyme Reverse Transcriptase (Fig. 9-14). The resulting double-stranded DNA fragments are subsequently inserted into an appropriate vector and cloned, yielding a population of clones known as a cDNA library. Searching for a specific gene is greatly facilitated by consulting a cDNA library derived from cells in which that gene is known to be expressed. For example, if one wished to clone globin genes, the initial step would be to construct a cDNA library from red Blood cell precursors, in which about half of the mRNA encodes Globins. To assist in mapping large genomes, cDNAs in the library can be partially sequenced at random to generate a useful type of STS called an expressed sequence tag (EST). Ranging from several dozen to hundreds of Base Pairs in length, ESTs can be mapped onto a large genome, thereby providing markers for expressed genes. Hundreds of thousands of ESTs have been mapped to detailed genomic profiles and have served as crucial reference points in deciphering The Human Genome.
Fig. 9-14. Construction of a cDNA library from mRNA. Cellular mRNA consists of transcripts from thousands of genes, making the resulting cDNAs heterogeneous. The double-stranded DNA generated by this method is inserted into a suitable cloning vector. Reverse transcriptase can synthesize DNA using either an RNA or a DNA template (see Fig. 26-33).

A cDNA library can be made even more specialized by cloning a cDNA or cDNA fragment into a vector where the cDNA sequence is fused with the sequence of a marker or reporter gene; the joined genes form a “reporter construct.” Two widely used markers are the green fluorescent protein gene and epitope tags. When a gene of interest is fused with the gene for green fluorescent protein (GFP), it produces a fusion protein that is intensely fluorescent—it literally glows (Fig. 9-15a). GFP from the jellyfish Aequorea victoria features a β-barrel Structure that encloses a fluorophore (see Box 12–3, p. 612). The fluorophore is formed through the rearrangement and oxidation of several amino acid residues in an autocatalytic reaction that requires only molecular oxygen (see Box 12–3, Fig. 3). Consequently, this protein can be readily expressed in active form within virtually any cell. Just a few molecules of the protein are sufficient for microscopic detection, allowing researchers to track its movements within The Cell. Protein Engineering has yielded mutant forms of the protein capable of fluorescing in different colors and possessing enhanced properties such as brightness and stability. Furthermore, several related Proteins have recently been isolated from other species.
Fig. 9-15. Specialized DNA libraries. (a) Cloning a cDNA adjacent to the green fluorescent protein (GFP) gene creates a reporter construct. Both the gene of interest (DNA insert) and the reporter gene undergo RNA METABOLISM/31.html">Transcription, and the resulting mRNA transcript is translated into a fusion protein, with the GFP portion readily visible via Fluorescence Microscopy. The photograph shows a roundworm expressing a GFP fusion protein localized to only four specific sites on its body. Reporter constructs (b) If a cDNA is cloned adjacent to an epitope tag gene, the resulting fusion protein can be immunoprecipitated using Antibodies specific to the epitope. Any interacting proteins bound to the tagged protein are co-precipitated, thereby helping to elucidate Protein-Protein Interactions.

An epitope tag is a short peptide sequence that is bound tightly by a well-characterized monoclonal antibody (Ch. 5). The tagged protein can be specifically precipitated from a total protein extract through this antibody interaction (Fig. 9-15b). If any other proteins are attached to the tagged protein, they will co-precipitate as well, providing valuable insight into intracellular protein-protein interactions. The diversity and utility of specialized DNA libraries continue to grow with each passing year.
The Polymerase Chain Reaction Amplifies Specific DNA Sequences
The Human Genome Project, alongside similar sequencing initiatives across diverse organisms, has provided unprecedented access to genetic information. In turn, this has greatly streamlined the cloning of individual genes for detailed biochemical analysis. If The nucleotide sequence of at least the flanking regions of a DNA segment of interest is known, the copy number of that segment can be vastly increased using the polymerase chain reaction (PCR)—a technique invented by Kary Mullis in 1983. The amplified DNA can be cloned directly or utilized in a wide array of analytical Procedures.
The PCR Procedure is remarkably straightforward. Two synthetic oligonucleotides are prepared that are complementary to sequences on opposite DNA strands immediately flanking the segment to be amplified. These oligonucleotides serve as primers for DNA Replication directed by a DNA polymerase. The 3'-ends of the hybridized primers point toward one another, positioning them to prime DNA Synthesis across the intervening segment (Fig. 9-16). (DNA polymerases synthesize DNA strands from deoxynucleotides using a DNA template, as discussed in detail in Ch. 25.) The isolated DNA containing the target segment is heated briefly to cause Denaturation, and then cooled in the presence of a large molar excess of synthetic oligonucleotide primers. Four types of deoxynucleoside triphosphates are added, and the primer-bound DNA segment is selectively replicated. This cycle of heating, cooling, and replication is repeated 25 or 30 times over the course of a few hours in an automated process, amplifying the primer-bounded DNA segment to the point where it can be readily analyzed or cloned. PCR relies on a heat-stable DNA polymerase, such as Taq polymerase (derived from a bacterium living at 90 °C), which remains active through each heating step and does not need to be replenished. Careful design of the PCR primers—such as incorporating restriction endonuclease Cleavage sites—can further facilitate the subsequent cloning of the amplified DNA (Fig. 9-16b).
Fig. 9-16. Amplification of a DNA segment by the polymerase chain reaction. (a) The PCR procedure consists of three steps. DNA strands are (1) separated by heating, then (2) annealed in the presence of an excess of short synthetic DNA primers (blue) that define the BOUNDARIES OF THE region to be amplified; and (3) new DNA is synthesized by polymerization. These three steps are repeated 25 to 30 times. The thermostable Taq DNA polymerase (from Thermus aquaticus, a bacterium inhabiting hot springs) is not denatured during the heating phases. (b) DNA amplified via PCR can be readily cloned. Primers may contain noncomplementary tails corresponding to restriction endonuclease cleavage sites. Although these portions of the primers do not base-pair with the template DNA, they become incorporated into the amplified DNA during PCR. Cleavage of the amplified fragments at these sites generates "sticky ends" that can be used to ligate the amplified DNA into a cloning vector. Polymerase chain reaction

This technology is exceptionally sensitive: PCR can detect and amplify as little as a single DNA molecule in virtually any sample. Although DNA naturally degrades over time (p. 415), PCR makes it possible to successfully clone DNA from samples more than 40,000 years old. Researchers have applied this technique to clone DNA fragments from the mummified remains of humans and extinct animals, such as the woolly mammoth, thereby opening new avenues in molecular archaeology and molecular paleontology. DNA extracted from archaeological excavation sites has been amplified using PCR and subsequently used to trace ancient human migration patterns. Epidemiologists can employ PCR-derived DNA samples from historical human remains to monitor the evolution of human pathogenic Viruses. Beyond DNA Cloning, PCR is exceptionally powerful in forensic science (Box 9-1). Additionally, PCR can be utilized to detect viral infections before symptoms manifest, as well as for the Prenatal Diagnosis of Genetic Disorders.
PCR is also indispensable for whole-genome sequencing efforts. For instance, mapping expressed regions on a given chromosome frequently involves amplifying ESTs via PCR, followed by hybridization of the amplified DNA to clones from an ordered library. Researchers have discovered numerous other Applications for PCR within the framework of the Human Genome Project, as we will explore next.
A Powerful Weapon in Forensic Science
Traditionally, one of the most accurate ways to determine whether a person was at a crime scene was fingerprinting. With The Development of Recombinant DNA technology, an even more powerful method emerged — DNA fingerprinting (also known as DNA typing or DNA profiling). This method was first described by the English geneticist Alec Jeffreys in 1985.
DNA fingerprinting is based on nucleotide sequence polymorphisms, i.e., minor sequence differences (typically single base-pair changes) among different individuals, averaging about one base pair per 1,000. Each variation from the prototype human genomic sequence (obtained for the first time) appears in a subset of the human population, and every individual has multiple such variations. Some sequence variations occur within restriction enzyme recognition sites. This leads to differences in the sizes of DNA fragments generated when the genomic DNA is cleaved by the corresponding restriction endonuclease. Such variations are called restriction fragment length polymorphisms (RFLPs). Another type of sequence variation frequently used for DNA identification is short tandem repeats (STRs).
The detection of RFLPs relies on a specialized hybridization technique called Southern blotting (Fig. 1). Fragments generated by digesting genomic DNA with restriction endonucleases are separated by size via electrophoresis, denatured by adding alkali to the agarose gel, and then transferred to a nylon membrane to replicate the fragment distribution from the gel. The membrane is immersed in a solution containing a radiolabeled DNA probe. A probe for a sequence that is repeated multiple times in the human genome typically recognizes a few of the thousands of DNA fragments produced by restriction endonuclease cleavage. Autoradiography displays the fragments to which the probe hybridizes, much like in Fig. 1. This method is highly accurate and came into use in forensic science in the late 1980s. However, the analysis requires a large amount of unfragmented DNA (>25 ng). Such quantities of DNA are frequently impossible to recover from a crime scene or disaster site.
Fig. 1. The Southern blotting method for analyzing restriction fragment length polymorphisms. Southern blotting, a technique used in molecular biology for a wide range of applications, is named after its inventor, Edwin Southern. In this forensic example, DNA from semen found on the body of a raped and murdered victim was compared with DNA samples from the victim and two suspects. Each DNA sample was cleaved into fragments and separated in a gel by electrophoresis. A radiolabeled DNA probe complementary to a specific sequence within a target fragment was used to identify the fragments. The sizes of the identified DNA fragments from the victim and the two suspects differed. As can be seen, the DNA banding pattern of one of the suspects is identical to that of the DNA sample recovered from the crime scene.

More sensitive DNA identification Methods utilize the power of the polymerase chain reaction (PCR; see Fig. 9-16), as well as STR analysis. A short tandem repeat is a short DNA sequence repeated multiple times in specific chromosomal regions; most frequently, the sequence consists of four nucleotide pairs. The most useful DNA loci for identification typically contain between 4 and 50 repeats (totaling 16 to 200 base pairs in the case of tetranucleotide repeats) and can vary in length across the human population. Over 20,000 tetranucleotide repeats have been characterized in the human genome. Estimates suggest that the human genome may contain over a million STRs of various types, accounting for about 3% of the total genomic DNA.
Table 1. Characteristics of Loci Used for the CODIS Database
Locus |
Chromosome |
Repeat |
Repeat length (range)* |
Number of alleles studied** |
CSF1PO |
5 |
TAGA |
5-16 |
20 |
FGA |
4 |
СТТТ |
12,2-51,2 |
80 |
ТН01 |
11 |
ТСАТ |
3-14 |
20 |
ТРОХ |
2 |
GAAT |
4-16 |
15 |
VWA |
12 |
[TCTG][TCTA] |
10-25 |
28 |
D3S1358 |
3 |
[TCTG][TCTA] |
8-21 |
24 |
D5S818 |
5 |
AGAT |
7-18 |
15 |
D7S820 |
7 |
GATA |
5-16 |
30 |
D8S1179 |
8 |
[TCTA][TCTG] |
7-20 |
17 |
D13S317 |
13 |
TATC |
5-16 |
17 |
D16S539 |
16 |
GATA |
5-16 |
19 |
D18S51 |
18 |
AGAA |
7-39,2 |
51 |
D21S11 |
21 |
[TCTA][TCTG] |
12-41,2 |
82 |
Amelogenin |
X,Y |
Not applicable |
* Repeat length in the human population. Some alleles may contain partial or imperfect repeats.
** Number of different alleles in the human population studied to date. Thorough analysis of loci across many individuals is a prerequisite for using DNA typing in forensics.
STR analysis employs the polymerase chain reaction. In forensic practice, the advantages of STR analysis over restriction fragment length polymorphism analysis became apparent early on (in the early 1990s) due to the method's much higher sensitivity. The DNA sequences flanking the STRs are unique to each repeat type and are identical (barring extremely rare Mutations) in all humans. PCR uses primers that bind to these sequences, allowing the amplification of the tandem repeat DNA (Fig. 2, a). Thus, the length of the PCR product corresponds to the length of the tandem repeat in the sample. Because every individual inherits one chromosome of a pair from each parent, the lengths of the tandem repeats on the two chromosomes are often different, yielding two signals when an individual's DNA is analyzed. When analyzing STR loci, the banding pattern obtained from DNA fingerprinting is unique to each individual. The PCR method enables researchers to obtain fingerprints from less than 1 ng of DNA, even if it is partially degraded — for example, from a single Hair follicle, a drop of blood, a small amount of semen recovered from a rape victim, or from DNA samples dating back several months or even years.
Fig. 2. PCR analysis of STR loci. a) PCR primers are designed to amplify a DNA fragment containing the repeat. One of the two primers is linked to a fluorescent dye (green circle). An individual's two chromosomes may possess different alleles with varying numbers of repeats at the same locus. Consequently, PCR amplification can yield two products of slightly differing sizes. These products are separated using a very thin polyacrylamide gel in a capillary tube. The resulting fluorescent bands are analyzed instrumentally to yield a series of peaks. By comparing this profile with markers, the size of each PCR product — and therefore the STR length of the corresponding allele — can be determined. b) The combination of PCR products from multiple loci generates a complex pattern, such as the one shown here. (In this case, it is a commercial STR multiplex kit containing sequences from 16 loci, the so-called 16-plex.) Such an analysis requires multiple sets of primer pairs — one set for each locus. To facilitate locus identification, the primers are labeled with different colored Dyes. Furthermore, the PCR primers for each locus are selected so that the size range of the resulting products differs as much as possible from those generated by other primer sets.

For the successful application of short tandem repeat analysis in forensic practice, standards are essential. The first such standard was established in the United Kingdom in 1995. The American standard, designated COmbined DNA Index System (CODIS), was introduced in 1998. The CODIS system utilizes 13 well-characterized STR loci (Table 1); incorporating these controls is mandatory for any DNA identification experiment performed in the United States. The amelogenin gene is also used as a marker. The flanking sequences of this gene, located on the human sex chromosomes (X and Y), differ slightly. Therefore, PCR amplification of the amelogenin gene yields products of different sizes, allowing the biological sex of the DNA donor to be determined. By 2006, the CODIS database contained 2.8 million profiles and was accessible nationwide across all states. As of 2005, this database had been consulted in over 25,000 legal proceedings.
Convenient kits are available that allow the amplification of 16 or more STR loci in a single tube. These kits (Fig. 2, b) contain locus-specific PCR primers. Each primer is synthesized in such a way that it does not cross-hybridize with other primers in the kit and that PCR generates products of distinct sizes, allowing the signals from different loci to be resolved during electrophoresis. Primers are conjugated with dyes to help distinguish the PCR products. The most widely used kit currently contains the 13 CODIS loci, amelogenin, and two additional loci (totaling 16), which are used worldwide in forensic examinations. Such kits are exceptionally useful for personal identification. Given a high-quality DNA profile, the probability of a chance match between two individuals in the entire human population is less than 1:108.
DNA typing is used both to convict and exonerate suspects, as well as in other contexts to establish biological origin with an extremely high degree of certainty. The impact of such procedures on the judicial system will continue to grow as organizations agree on standards and the methods become universally accepted standard practice in forensic laboratories. Even decades-old murder mysteries can be solved: in 1996, DNA fingerprinting helped confirm the identity of the remains of the last Russian Tsar and his family, who were executed in 1918.
Genomic Nucleotide Sequences Serve as the Foundation for Massive Gene Libraries
The Genome is the ultimate source of information about an organism, and no genome interests us more than our own. Less than a decade after practical DNA Sequencing Methods were developed, serious discussions began regarding the feasibility of sequencing all 3 billion base pairs of the human genome. The international Human Genome Project was launched following substantial funding in the late 1980s. Ultimately, major contributions to the project's development came from 20 sequencing centers across six countries: the United States, the United Kingdom, Japan, France, China, and Germany. Overall coordination was provided by the National Center for Human Genome Research at the National Institutes of Health (USA), initially led by James Watson and, after 1992, by Francis Collins. Initially, the task of determining the sequence of 3 • 109 bp of the genome seemed like a titanic undertaking, but it progressively drove technological breakthroughs. The complete human genome sequence was published in April 2003, several years ahead of schedule.
This milestone was the result of carefully planned international collaboration spanning 14 years. Research groups first generated a high-resolution map of the human genome, with clones derived from each chromosome organized into a series of long contigs (Fig. 9-17). Each contig contained sequence-tagged site (STS) landmarks spaced at intervals of less than 100,000 bp. A genome mapped in this fashion could be partitioned among international sequencing centers, with each center sequencing the mapped BAC or YAC clones corresponding to its assigned genomic segments. Because many clones exceeded 100,000 bp in length, and sequencing techniques could resolve only 600 to 750 bp of nucleotide sequence at a time, each clone had to be sequenced piecemeal. The sequencing strategy relied on the whole-genome shotgun approach, in which researchers used powerful new sequencers to determine the sequences of random segments of a given clone and then assembled them using computer algorithms to identify overlaps. The number of sequenced clone clones was statistically monitored to ensure that each entire clone was covered an average of 4 to 6 times. Sequenced DNA fragments were subsequently retrieved from databases to assemble the complete genome. Constructing the genetic map was a laborious endeavor, accompanied by annual progress reports in major journals throughout the 1990s, with the map nearing completion by the end of the decade. The completion of the entire human genome sequencing project was originally scheduled for 2005, but financial and technical investments accelerated the timeline.
Fig. 9-17. Strategy of the Human Genome Project. Clones derived from a genomic library were ordered onto a detailed genetic map, after which individual clones were sequenced using the shotgun method. Commercial sequencing technologies eventually bypassed The Need for a genetic map, enabling whole-genome sequencing directly via shotgun cloning.

Commercial contributions to decoding the human genome were initiated by Celera Corporation, founded in 1997. Led by J. Craig Venter, this group adopted an alternative strategy known as whole-genome shotgun sequencing, which bypassed the construction of a physical genetic map. Instead, DNA segments were randomly sampled from the entire genome and sequenced. The sequenced segments were ordered by computer identification of overlapping sequences, with minimal reliance on the detailed public project map. At the inception of the Human Genome Project, shotgun sequencing on such a scale seemed far from practical. However, advances in computer software and sequencing automation made this approach feasible by 1997. The competition between private and public sequencing efforts significantly shortened the time to completion. The publication of preliminary human genome sequence data in 2001 was followed by two years of finishing work to resolve approximately a thousand gaps and discrepancies, yielding a high-quality, contiguous sequence across the entire genome.

The Human Genome Project marked the culmination of 20th-century biology and heralds a radically different science for the coming century. The human genome is merely one milestone, as the genomes of many other species are being (or have already been) decoded. These include the yeast Saccharomyces cerevisiae (sequencing completed in 1996) and Schizosaccharomyces pombe (2002), the nematode Caenorhabditis elegans (1998), the fruit fly Drosophila melanogaster (2000), the plant Arabidopsis thaliana (2000), the mouse Mus musculus (2002), the zebrafish, and numerous species of Bacteria and archaea (Fig. 9-18). Early genome-sequencing efforts focused primarily on model organisms commonly used in laboratory research. As technology advances, the complete genome sequences of over 1,200 diverse organisms will be known by the time this book goes to press. Today, large-scale efforts are already underway for gene mapping, the discovery of novel proteins and disease-associated genes, alongside many other initiatives.
Fig. 9-18. Genome sequencing timeline. Discussions in the mid-1980s led to the project officially launching only in 1989. Preliminary work, which included compiling a complete genetic map and establishing genomic landmarks, spanned most of the 1990s. Separate projects were also initiated to sequence the genomes of other organisms required for research practice. Currently, the decoding of genomes has been completed for many species of bacteria (e.g., Haemophilus influenzae), yeast (S. cerevisiae), nematodes (C. elegans), insects (D. melanogaster and Apis mellifera), plants (A. thaliana and Oryza sativa L.), rodents (Mus musculus and Rattus norvegicus), primates (Homo sapiens and Pan troglodytes), and certain human sexually transmitted pathogens (e.g., Trichomonas vaginalis). Each genome project has a website that serves as a central data repository.

These studies have generated a database with the potential not only to drive rapid progress in biology but also to reshape humanity's very perception of itself. Our initial impressions of the human genome sequencing ranged from a sense of bewilderment to an appreciation of nature's profound wisdom. We are not as complex as we think we are. Estimates made several decades ago that humans possess approximately 100,000 genes distributed among 3,2 • 109 bp of the genome proved incorrect: we have only 25,000 to 30,000 genes. This is apparently about 1.5 times more than Drosophila (20,000 genes) and slightly more than the nematode (23,000). Although humans appeared relatively recently in the course of evolution, our genome is very ancient. Of the 1,278 Protein Families found in an ancient layer, only 94 were unique to vertebrates. However, while many types of Protein domains are shared with plants, worms, and fruit flies, we utilize them in more complex mechanisms. Alternative pathways of Gene Expression (Chap. 26) make it possible to produce more than one protein from a single gene; this process occurs more frequently in humans and other vertebrates than in bacteria, worms, or any other life forms. This contributes to the greater complexity of proteins generated from our gene complement.
It is now known that only 1.1% to 1.4% of our DNA actually encodes proteins (Fig. 9-19). Over 50% of our genome consists of short repeating sequences, the vast majority of which—about 45% of the genome overall—originate from Transposons, small mobile DNA sequences that act as molecular parasites (Chap. 25). Most transposons have resided in our genome for a long time and have now mutated to the point where they can no longer relocate to new sites. The remainder are still actively moving at a low frequency, rendering the genome dynamic and evolving. Finally, a small number of transposons have been co-opted by their host and appear to perform important cellular Functions.
Fig. 9-19. STRUCTURE OF THE human genome. The diagram shows the proportions of various types of sequences in our genome.

What can we infer from all this information regarding how much one individual differs from another? Within the human population, there are millions of single-base DNA variations called Single Nucleotide Polymorphisms, or SNPs (pronounced "snips"). Every individual differs from another by one base pair per thousand. From these minor genetic variations arises the human diversity we all recognize—differences in hair and eye color, drug allergies, FOOT size, and even (to some unknown degree) behavior. Certain SNPs are associated with specific human populations and can provide vital insights into human Migrations that occurred thousands of years ago and into our more distant evolutionary past.
How does this information help us understand what truly makes us human? Answers to some of these questions can be found by analyzing the genome of our closest relative, the chimpanzee. The human and chimpanzee genomes differ in nucleotide composition by only 1.2%, and the difference in protein-coding genes is even smaller. While this may seem minor, across such large genomes, this difference amounts to about 35 million bp, another 5 million short insertions or deletions, and a fairly large number of larger-scale genomic rearrangements. Figuring out which of these differences are responsible for the phenotypic distinctions between humans and chimpanzees is no simple task. Primate Genome Analysis can provide significant assistance in understanding human biochemical Organization and evolution, but solving this puzzle is still in its infancy.
Impressive as this success is, decoding the human genome in itself is not as daunting as the task that lies ahead—attempting to comprehend all the information stored within the genome. The genomic sequences added monthly to the international database are "roadmaps," parts of which are written in a language we do not yet understand. Nevertheless, they are immensely useful in driving the discovery of new proteins and processes that impact every aspect of biochemistry, as will become clear in the following chapters.
Summary of Section 9.2 From Genes to Genomes
■ The science of genomics is dedicated to the active study of genomes and their contents.
■ Genomic DNA segments can be assembled into libraries—such as gene banks and cDNA libraries—for a vast range of purposes and applications.
■ Polymerase chain reaction (PCR) can be used to amplify individual DNA segments from a DNA library or an entire genome.
■ Through the combined efforts of international research consortia, the genomes of many organisms, including humans, have been fully sequenced, and this information is now accessible in public databases.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.