LEHNINGER PRINCIPLES OF BIOCHEMISTRY - VOL. 1. THE FOUNDATIONS OF BIOCHEMISTRY: STRUCTURE AND CATALYSIS - 2011

CHAPTER I. STRUCTURE AND CATALYSIS

Class="center">Of all the systems in nature, Living matter is the only one that, despite vast transformations, retains within its Structure the largest amount of its own past.

Emil Zuckerkandl and Linus Pauling, article in Journal of Theoretical Biology, 1965

9. DNA-BASED INFORMATION TECHNOLOGY

Let us turn to the technologies that underpin the progress of modern biological sciences, define the current and future frontiers of biochemistry, and illustrate many of its essential principles. Elucidating the laws governing Enzymatic Catalysis, macromolecular structure, cellular METABOLISM, and information transfer makes it possible to investigate increasingly complex biochemical processes. Cell Division, Immunity, Embryogenesis, Vision, taste, oncogenesis, and cognitive function all blend harmoniously in an intricately organized orchestra of molecular and macromolecular interactions that we are now beginning to understand with growing clarity. The ultimate outcome of the biochemical journey that began in the 19th century is an ever-expanding capability to explore and alter living systems.

To understand a complex biochemical process, a biochemist isolates it and investigates its individual components in vitro, then integrates them to obtain a coherent picture of the entire process. The primary source of insight into its molecular essence is The Cell’s own informational archive, its DNA. However, the sheer size of Chromosomes presents a monumental challenge: how to locate and study a single Gene among tens of thousands of genes encompassing billions of Base Pairs in a mammalian genome. Solutions began to emerge in the 1970s.

The achievements of decades of work by thousands of scientists—geneticists, biochemists, cell biologists, and physical chemists—were brought together in the laboratories of Paul Berg, Herbert Boyer, and Stanley Cohen, who developed Methods for detecting, purifying, preparing, and studying small segments of DNA derived from chromosomes vastly larger than themselves. DNA Cloning METHODS paved the way for such modern fields as Genomics and Proteomics, the large-scale study of genes and Proteins in whole Cells and organisms. These novel techniques are revolutionizing basic research, agriculture, medicine, environmental science, forensics, and many other fields, sometimes confronting society with difficult choices and ethical dilemmas.

We begin this chapter with a broad Overview of the fundamental biochemical principles of what is now the classical methodology of DNA cloning. Then, after laying the groundwork for a Discussion of genomics, we will illustrate a range of Applications and capabilities of these technologies, focusing primarily on recent advances in genomics and proteomics.

9.1. DNA Cloning: Basic Concepts

A clone is an identical copy. The term was originally used for cells of a single type isolated and capable of reproduction to yield a population of identical cells. DNA cloning involves isolating a specific gene or DNA segment from a chromosome, attaching it to a small carrier DNA molecule, and then copying this modified DNA thousands or millions of times, either by increasing the number of cells or by generating multiple copies of the cloned DNA within each cell. The result is the selective Amplification of a given gene or DNA segment. Cloning DNA from any Organism involves five principal Procedures.

1. Cutting DNA at precise positions. Site-specific endonucleases (Restriction Endonucleases) provide the necessary molecular "scissors".

2. Selecting a small DNA molecule capable of self-Replication. These DNAs are called cloning vectors (vector meaning a delivery agent); typically Plasmids or viral DNAs.

3. Covalently joining two DNA fragments. The enzyme DNA ligase seals the cloning vector to the DNA to be cloned. Hybrid DNA molecules containing covalently linked segments from two or more sources are called recombinant DNAs.

4. Moving the recombinant DNA from the test tube into a host cell, which provides the enzymatic machinery for DNA replication.

5. Selecting or identifying host cells that contain the recombinant DNA.

The set of techniques used to perform these and similar procedures is called Recombinant DNA technology or, more informally, Introduction/32.html">Genetic Engineering.

Much of our initial discussion will focus on DNA cloning in the bacterium Escherichia coli, the first organism used for recombinant DNA work and still the most common host cell. E. coli offers many advantages: its DNA metabolism (like many of its other biochemical processes) is well understood; numerous naturally occurring cloning vectors associated with E. coli, such as plasmids and Bacteriophages (bacterial Viruses, also called phages), are well characterized; and methods for rapidly transferring DNA from one bacterial cell to another are readily available. We will also examine DNA cloning in other organisms; this topic is discussed more fully below.

Restriction Endonucleases and DNA Ligase Create Recombinant DNA

Crucial to recombinant DNA technology is the array of Enzymes (Table 9-1) made available by decades of research into NUCLEIC ACID METABOLISM. Two classes of enzymes form the foundation of the basic approach for creating and "amplifying" a recombinant DNA molecule (Fig. 9-1). First, restriction endonucleases (restriction enzymes) recognize and cleave DNA at specific nucleotide sequences (recognition sequences or restriction sites), generating a set of smaller fragments. Second, the DNA fragment to be cloned can be inserted into a suitable cloning vector using DNA ligase to join the DNA molecules. The recombinant vector is then introduced into a host cell, which propagates the fragment through multiple rounds of cell division.

Table 9-1. Some Enzymes Used in Recombinant DNA Technology

Enzyme(s)

Function

Type II restriction endonucleases

Cleave DNA at specific base sequences

DNA ligase

Joins two DNA molecules or fragments

DNA polymerase I (E. coli)

Fills in gaps in duplexes by stepwise addition of NUCLEOTIDES to 3'-ends

Reverse Transcriptase

Synthesizes a DNA copy of an RNA molecule

Polynucleotide kinase

Adds a phosphate to the 5'-OH end of a polynucleotide, to label it or permit ligation

Terminal transferase

Adds homopolymer tails to the 3'-OH ends of a linear duplex

Exonuclease III

Removes nucleotide residues from the 3'-ends of a DNA strand

Bacteriophage λ exonuclease

Removes nucleotides from the 5'-ends of a duplex to expose single-stranded 3'-ends

Alkaline phosphatase

Removes terminal phosphates from the 5'- or 3'-end (or both)

Fig. 9-1. Overview of DNA cloning. A cloning vector and eukaryotic chromosomes are cleaved independently by the same restriction endonuclease. The fragments to be cloned are then inserted into the cloning vector. The resulting recombinant DNA (only one recombinant vector is shown) is introduced into a host cell, where it is replicated (cloned). Note: Not to scale; the size of the E. coli chromosome is actually much larger than that of a typical cloning vector (such as a plasmid).

Restriction endonucleases are found in A wide variety of Bacteria. In the early 1960s, Werner Arber discovered that the biological function of restriction enzymes is to recognize and cleave foreign DNA (such as the DNA of infecting viruses)—a process known as DNA Restriction. In the host cell's DNA, The base sequence recognized by its own restriction endonuclease is protected from Cleavage by DNA Methylation, catalyzed by a specific DNA methylase. The restriction endonuclease and the corresponding methylase are sometimes referred to as a restriction-modification system.

Three types of restriction endonucleases are known: types I, II, and III. Type I and III restriction enzymes are typically large, multi-subunit complexes possessing both endonuclease and methylase activity. Type I restriction endonucleases cleave DNA at random sites that can be more than 1,000 base pairs (1,000 bp) away from the recognized nucleotide sequence. Type III restriction endonucleases cleave DNA approximately 25 bp away from the recognized nucleotide sequence. Both types of restriction enzymes move along the DNA via a reaction that requires ATP energy. Type II restriction endonucleases, first isolated by Hamilton Smith in 1970, are more convenient because they do not require ATP and cleave DNA within the recognized nucleotide sequence itself. The exceptional utility of this group of restriction endonucleases was demonstrated by Daniel Nathans, who first used them to develop novel methods for mapping and studying genes and genomes.

Thousands of restriction endonucleases have been discovered in various bacterial species, with one or more such enzymes recognizing over 100 different DNA sequences. The recognized nucleotide sequences are typically 4 to 6 bp in length and are palindromes (see Fig. 8-18). Table 9-2 lists the base sequences recognized by several type II restriction endonucleases. In some cases, the interaction between a restriction endonuclease and its recognized nucleotide sequence has been elucidated in detail at THE MOLECULAR LEVEL; for example, Table 9-2 illustrates the sequences recognized by several type II restriction endonucleases.

Table 9-2. Nucleotide Sequences Recognized by Selected Type II Restriction Endonucleases

Arrows indicate the phosphodiester bonds cleaved by each restriction endonuclease; asterisks indicate bases methylated by the corresponding methylase (where known); N represents any arbitrary base. Note that the name of each enzyme consists of a three-letter italicized abbreviation of the bacterial species from which it was isolated, sometimes followed by a strain designation and Roman numerals (to distinguish different restriction endonucleases isolated from the same bacterial species). For example, BamHI indicates the first (I) restriction endonuclease isolated from strain H of the bacterium Bacillus amyloliquefaciens.

Some restriction endonucleases introduce staggered cuts in the two DNA strands, leaving two to four nucleotides of one strand unpaired at each resulting end. These unpaired single strands are called sticky ends (Fig. 9-2a) because they can hydrogen-bond to each other or to complementary sticky ends of other DNA fragments. Other restriction endonucleases cleave both DNA strands at opposite phosphodiester bonds, leaving no unpaired bases at the ends; such ends are frequently referred to as blunt ends (Fig. 9-2b).

Fig. 9-2. Cleavage of DNA molecules by restriction endonucleases. Restriction endonucleases recognize and cleave only specific nucleotide sequences, leaving either (a) sticky ends (with protruding single-stranded overhangs) or (b) blunt ends. The fragments can be joined to other DNA molecules, such as a cleaved cloning vector (plasmid) shown here. This reaction is facilitated by the annealing of complementary sticky ends. Joining (ligation) is less efficient for DNA fragments with blunt ends than for those with complementary sticky ends, and DNA fragments with different (non-complementary) sticky ends generally do not ligate at all. (c) A synthetic DNA fragment containing sequences recognized by several restriction endonucleases can be inserted into a restriction-cleaved plasmid. An insertion with a single restriction site is called a linker, and one with multiple sites is called a polylinker.

The average size of DNA fragments generated by restriction endonuclease Digestion of genomic DNA depends on the frequency with which the specific restriction site occurs in the DNA molecule, which in turn strongly depends on the size of the recognition sequence. In a random-sequence DNA molecule in which all four bases are present in equal amounts, a 6-bp sequence recognized by a restriction endonuclease such as BamHI would occur on average once every 46 (4,096) bp, assuming the DNA contains 50% G=C pairs. Enzymes that recognize 4-bp sequences would generate shorter fragments from a random-sequence DNA molecule, since a recognition sequence of that size would occur about once every 44 (256) bp. In natural DNA molecules, specific recognition sequences occur less frequently because nucleotide sequences in DNA are not random, and the four nucleotide types are not present in equal amounts. In laboratory experiments, the average size of fragments generated by restriction endonuclease digestion of a large DNA molecule can be increased simply by stopping the reaction before it goes to completion—a result known as partial digestion. Fragment size can also be increased by using a specialized class of enzymes called homing endonucleases (see Fig. 26-38), which recognize and cleave much longer DNA sequences (14 to 20 bp).

Once a DNA molecule has been cleaved into fragments, a specific fragment of a desired size can be isolated by agarose or Polyacrylamide gel Electrophoresis or by HPLC (p. 135). However, in the case of a typical mammalian genome, restriction endonuclease digestion typically yields a vast number of different DNA fragments, making it difficult to isolate a specific fragment by electrophoresis or HPLC. A common intermediate step in cloning a specific gene or DNA segment is the creation of a DNA library (described in Section 9.2).

After isolating the required DNA fragment, DNA ligase can be used to join it to a identically cleaved cloning vector—that is, a vector digested with the same restriction endonuclease. For example, a fragment generated with EcoRI will generally not ligate to a fragment generated with BamHI. As described in more detail in Chapter 25 (see Fig. 25-17), DNA ligase catalyzes The formation of new phosphodiester bonds in a reaction utilizing ATP or a similar cofactor. Base pairing of complementary sticky ends greatly facilitates ligation (Fig. 9-2a). Blunt ends can also be ligated, albeit less efficiently. Researchers can create novel DNA sequences by inserting synthetic DNA fragments (called linkers) between the ends to be ligated. DNA inserts containing multiple restriction endonuclease recognition sequences (commonly used subsequently as insertion sites for additional DNA via cleavage and ligation) are called polylinkers (Fig. 9-2c).

The efficacy of sticky ends in selectively joining two DNA fragments became apparent in the earliest recombinant DNA experiments. Even before restriction endonucleases became widely available, some researchers discovered that sticky ends could be created through the combined action of bacteriophage λ exonuclease and terminal transferase (Table 9-1). The fragments to be joined contained complementary homopolymeric "tails." Peter Lobban and Dale Kaiser used this method in 1971 in the first experiments joining naturally occurring DNA fragments. Shortly thereafter, Similar Methods were employed in Paul Berg's laboratory to join DNA segments of simian virus 40 (SV40) with DNA derived from bacteriophage λ, thereby creating the first recombinant DNA molecule containing segments derived from different species.

Cloning vectors allow the amplification of inserted DNA segments

The principles governing the delivery of recombinant DNA in a clonable form into a host cell and its subsequent amplification can be illustrated using three common cloning vectors typically employed in E. coli experiments—plasmids, bacteriophages, and bacterial artificial chromosomes—along with a vector used for cloning large DNA segments in Yeast.

Plasmids. Plasmids are circular DNA molecules that replicate independently of the host cell chromosome. Bacterial plasmids ranging from 5,000 to 400,000 bp in length occur naturally. They can be introduced into bacterial cells via a process called transformation. Cells (typically E. coli) and plasmid DNA are incubated together at 0 °C in a calcium chloride solution and then subjected to a heat Shock by rapidly shifting the Temperature to 37–40 °C. For reasons not yet fully understood, some cells take up plasmid DNA under these conditions. Certain bacterial species are naturally competent for DNA uptake and do not require calcium chloride Treatment. In an alternative method, cells incubated with plasmid DNA are subjected to a high-voltage electrical pulse. In this approach, termed electroporation, bacterial membranes are rendered transiently permeable to large macromolecules.

Regardless of the method, only a small fraction of cells take up plasmid DNA, necessitating a method to select for successfully transformed ones. A common approach is to use a plasmid containing a gene required for host cell growth under specific conditions, such as an Antibiotic Resistance gene. Only cells transformed with the recombinant plasmid can grow in the presence of the antibiotic, rendering any plasmid-containing cell "selectable" under such growth conditions. Such a gene is called a selectable marker.

By modifying natural plasmids, researchers have developed many different Plasmid Vectors suitable for cloning. The E. coli plasmid pBR322 is a prime example exhibiting properties convenient for a cloning vector (Fig. 9-3).

Fig. 9-3. The plasmid pBR322, designed for E. coli. Note the locations of several important restriction sites—for PstI, EcoRI, BamHI, SalI, and PvuII; the ampicillin and tetracycline resistance genes; and THE ORIGIN OF replication (ori). Created in 1977, it became one of the earliest plasmids specifically engineered for cloning in E. coli.

Important advantages of plasmid pBR322:

1. The origin of replication (ori)—the sequence where replication is initiated by cellular enzymes (Ch. 25). This sequence is required for plasmid duplication and maintenance at 10 to 20 copies per cell.

2. Two antibiotic resistance genes (tetR, ampR), which enable the identification of cells harboring either the intact plasmid or its recombinant derivative (Fig. 9-4).

Fig. 9-4. Use of pBR322 for cloning foreign DNA in E. coli and identifying cells containing it. Plasmid cloning

3. Several unique recognition sequences (PstI, EcoRI, BamHI, SalI, PvuII) that serve as targets for various restriction endonucleases, providing sites where the plasmid can subsequently be cleaved to insert foreign DNA.

4. Small size (4,361 bp), which facilitates its cellular uptake and biochemical manipulation of DNA.

Transformation of typical bacterial cells with purified DNA (itself an inefficient process) becomes progressively less successful as plasmid size increases; consequently, cloning DNA segments larger than 15,000 bp becomes difficult when using plasmids as vectors.

Bacteriophages. Bacteriophage λ features an efficient mechanism for injecting its own DNA (48,502 bp) into a bacterium, making it suitable as a cloning vector for somewhat larger DNA segments (Fig. 9-5). Two key characteristics make it convenient to use:

1. About one-third of the λ genome is nonessential and can be replaced by foreign DNA.

2. DNA is packaged into infectious phage particles only if its length is between 40,000 and 53,000 bp, a restriction that can be exploited to ensure the packaging of recombinant DNA exclusively.

Figure 9-5. Cloning vectors based on bacteriophage λ. Recombinant DNA techniques are used to modify the bacteriophage λ genome—removing genes unnecessary for phage production and replacing them with "stuffer" DNA to make the phage DNA large enough to fit into the phage particles. As shown here, in cloning experiments the stuffer is replaced by foreign DNA. Recombinants are packaged in vitro into viable phage particles only if they contain a foreign DNA fragment of the appropriate size along with the bacteriophage λ DNA end fragments.

Researchers have developed Bacteriophage λ-Based Vectors that can be readily cleaved into three pieces such that two of them contain the essential bacteriophage λ genes and together total only about 30,000 bp in length. The third piece, the "stuffer" DNA, is discarded when the vector is used for cloning. Additional DNA is inserted between the two essential bacteriophage λ segments, generating ligated DNA molecules long enough to produce viable phage particles. This packaging mechanism is used to select for recombinant Viral Particles.

Bacteriophage λ vectors allow the cloning of DNA fragments up to 23,000 bp in length. Once the bacteriophage λ fragments are ligated with foreign DNA fragments of suitable size, the resulting recombinant DNAs can be packaged into phage particles by adding them to crude bacterial cell extracts containing all the proteins required to assemble an intact phage. This process is called in vitro packaging (Fig. 9-5). All viable phage particles will contain a foreign DNA fragment. Subsequent transfer of the recombinant DNA into E. coli cells is extremely efficient.

Bacterial artificial chromosomes (BACs). Bacterial artificial chromosomes are plasmids designed for cloning very large (typically 100,000 to 300,000 bp) DNA segments (Fig. 9-6). They typically contain selectable markers, such as the chloramphenicol resistance gene (CmR), as well as a highly stable origin of replication (ori) that maintains a low plasmid copy number (1–2 copies per cell). BAC vectors can accommodate DNA fragments hundreds of thousands of base pairs in length. These massive circular DNAs are then introduced into host bacteria via electroporation. These procedures employ bacterial strains carrying Mutations that alter The Cell wall structure, facilitating the uptake of large DNA molecules.

Figure 9-6. Bacterial artificial chromosomes (BACs) as cloning vectors. The vector is a relatively simple plasmid equipped with an origin of replication (ori) that controls replication. par genes, derived from plasmids known as F-plasmids, ensure the uniform partitioning of plasmids to daughter cells during cell division. This increases the likelihood that each daughter cell inherits a single copy of the plasmid, even at low copy numbers. A low copy number is advantageous when cloning large DNA segments because it limits unwanted recombination events that could otherwise unpredictably alter large cloned DNAs over time. The BAC vector includes selectable markers. The lacZ gene (required for Synthesis of the enzyme ξta-galactosidase) is positioned within the cloning site such that it is inactivated by the insertion of cloned DNAs. The uptake of recombinant BACs into cells by electroporation is facilitated by using cells with modified (more porous) cell walls. Screening for recombinant DNAs relies on cellular resistance to the antibiotic chloramphenicol (CmR). Growth plates also contain a ξta-galactosidase substrate, which leads to product coloration. Colonies with active ξta-galactosidase—and thus lacking DNA inserts in the BAC vector—turn blue, whereas colonies lacking ξta-galactosidase activity—and therefore carrying the desired DNA inserts—appear white.

Yeast artificial chromosomes (YACs). In genetic engineering, E. coli cells are by no means the only suitable hosts. Yeast is an exceptionally versatile eukaryotic organism for such work. As with E. coli, yeast genetics is well established. The Genome of the most commonly used yeast, Saccharomyces cerevisiae, contains only 14 × 106 bp (by eukaryotic standards, a compact genome less than four times the size of the E. coli chromosome); its entire nucleotide sequence is known. Yeast is also very easy to maintain and cultivate in large quantities in the laboratory. Plasmid vectors have been developed for yeast using the same approaches underlying the E. coli vectors discussed above. Robust methods are now available for transferring DNA both into and out of yeast cells, facilitating The Study of many aspects of Eukaryotic Cell biochemistry. Some recombinant plasmids contain multiple ori sequences and other elements, making them functional in more than one species (e.g., in both yeast and E. coli cells). Plasmids capable of replicating in two or more different species are called shuttle vectors.

Investigations of large genomes and the concomitant need for high-capacity cloning vectors led to The Development of yeast artificial chromosomes (YACs; Fig. 9-7). YAC vectors incorporate all elements essential for maintaining A eukaryotic chromosome in the yeast nucleus: a yeast ori, two selectable markers, and specific sequences (derived from telomeres and centromeres, DNA regions discussed in Chapter 24) required for stability and proper chromosome segregation during cell division. Before the vector is used for cloning, it is amplified as a circular bacterial plasmid. Cleavage with a restriction endonuclease (BamHI in Fig. 9-7) removes the DNA segment between the two telomeric sequences (TEL), leaving telomeres at the ends of the linearized DNA. Cleavage at a second site (EcoRI in Fig. 9-7) splits the vector into two DNA segments, termed vector "arms," each bearing its own selectable marker.

Figure 9-7. Construction of a yeast artificial chromosome (YAC). A YAC vector contains an origin of replication (ori), a centromere (CEN), two telomeres (TEL), and selectable markers (X and Y). Cleavage with the restriction endonucleases BamHI and EcoRI generates two separate DNA "arms"; each possesses one telomeric end and one selectable marker. A large DNA segment (e.g., up to 2 × 106 bp from The Human Genome) is ligated to the two arms to form a yeast artificial chromosome. The YAC is used to transform yeast cells (spheroplasts stripped of their cell walls), which are then selected for markers X and Y; surviving cells replicate the inserted DNA.

Genomic DNA is subjected to partial digestion with restriction endonucleases (EcoRI in Fig. 9-7) to yield fragments of an appropriate size. Genomic fragments are then fractionated by pulsed-field gel electrophoresis (a specialized variant of gel electrophoresis; see Fig. 3-18) which permits the Separation of very large DNA segments. DNA fragments of the proper size (up to 2 × 106 bp) are mixed with the prepared vector arms and ligated. The resulting mixture is then used to transform prepared yeast cells with these very large DNA molecules. Growth on media requiring the expression of both selectable marker genes ensures the Selection of only those yeast cells containing an artificial chromosome with a substantial insert situated between the two vector arms (Fig. 9-7). The stability of YAC clones increases with their size. YAC clones with inserts exceeding 100,000 bp are roughly as stable as native cellular chromosomes, whereas those with inserts smaller than 100,000 bp progressively degrade during mitosis (for instance, yeast clones possessing merely two vector ends ligated together or containing very short inserts are rarely recovered). YACs that lose telomeres at either end rapidly disintegrate.

Specific DNA Sequences Are Identified by Hybridization

DNA hybridization, introduced broadly in Chapter 8 (see Fig. 8-29), is the premier sequence-specific method for detecting individual genes or nucleic acid segments. Numerous variations of the basic technique exist, predominantly utilizing labeled (e.g., radioactive) DNA or RNA fragments (termed probes) that are complementary to the target DNA. In the classical approach for detecting a specific DNA sequence within a genomic library (a collection of DNA clones), a nitrocellulose membrane is pressed onto an Agar plate containing numerous individual bacterial colonies from the library, each harboring a distinct recombinant DNA. A few cells from each colony adhere to the paper, creating a replica. The membrane is subsequently treated with alkali to lyse the cells and denature the DNA; the DNA molecules remain bound to the paper adjacent to the positions of the colonies from which they originated. A radiolabeled DNA probe is added and binds exclusively to its complementary DNA. After washing away unhybridized DNA probes, the hybridized DNA can be visualized by autoradiography (Fig. 9-8).

Figure 9-8. Using hybridization to identify a clone containing a specific DNA segment. A radioactive DNA probe hybridizes to its complementary DNA and is subsequently detected by autoradiography. Once the labeled colonies are identified, the corresponding colonies on the original agar plate can be used in subsequent experiments as a source of the cloned DNA.

Typically, generating the complementary nucleotide strand used as a probe is the rate-limiting step in identifying and cloning a gene. The Nature of the probe depends on what is already known about the gene of interest. Sometimes a suitable probe is a homologous gene cloned from another species. Alternatively, if the protein product of the gene has already been isolated, a probe can be designed and synthesized based on the Amino Acid Sequence (Fig. 9-9). Today, researchers generally retrieve all necessary DNA sequence information from nucleotide sequence Databases, which detail the structures of millions of genes across a wide spectrum of organisms.

Figure 9-9. Probe design for identifying a gene encoding a protein with a known amino acid sequence. Because more than one DNA sequence can encode any given amino acid sequence, The Genetic Code is said to be "degenerate" (as discussed in Chapter 27; each amino acid is specified by a triplet of nucleotides—a codon. Most Amino Acids are specified by two or more codons; see Fig. 27-7). Consequently, the exact DNA sequence corresponding to a known amino acid sequence cannot be predicted outright. A probe is designed to be complementary to a region of the gene exhibiting minimal degeneracy—that is, a region with the fewest possible codons for its amino acids; in the example shown here, no more than two codons. Oligonucleotides are synthesized with selectively randomized sequences, containing one of two possible nucleotides at each position of potential degeneracy (highlighted in pink). The oligonucleotide shown here represents a pool of eight distinct sequences, one of which is perfectly complementary to the gene, while all eight together match at least 17 out of 20 positions.

Expression of cloned genes yields significant amounts of protein

Frequently, the product of the cloned gene is of primary interest rather than the gene itself, especially when the protein holds commercial, therapeutic, or research value. Now that the fundamentals of metabolism and The regulation of DNA, RNA, and proteins in E. coli are becoming increasingly clear, researchers can readily manipulate cells to express cloned genes for the purpose of studying their protein products.

Most eukaryotic genes lack the DNA sequence elements required for their expression in E. coli cells (such as promoters—sequences that direct DNA polymerase to binding sites). Therefore, bacterial transcriptional and translational regulatory sequences must be introduced into the vector DNA at locations appropriate for eukaryotic genes. Promoters, regulatory sequences, and Other Aspects of Gene Expression regulation are discussed in Chapter 28. In some cases, cloned genes are expressed so efficiently that their protein product accounts for more than 10% of total cellular protein; such genes are said to be overexpressed. At these high concentrations, certain foreign proteins can kill the E. coli host, necessitating that gene expression be restricted to a few hours before cell harvest.

Cloning vectors equipped with the Transcription and Translation signals necessary for the regulated expression of a cloned gene are commonly called expression vectors. The rate of cloned gene expression is controlled by replacing the gene's native promoter and regulatory sequences with more efficient and suitable variants provided by the vector. Typically, a well-characterized promoter and its regulatory elements are positioned near several specific restriction sites used in cloning, ensuring that genes inserted at these sites are driven by the regulated promoter element (Fig. 9-10). Some of these vectors also incorporate other characteristic features, such as a bacterial ribosome binding site to enhance the translation of mRNA transcribed from the gene, or a transcription termination sequence.

Fig. 9-10. DNA sequences in a typical E. coli expression vector. The expressed gene is inserted into one of the restriction sites of the polylinker near the promoter (P) such that the end encoding the N-terminal amino acid sequence is closest to the promoter. The promoter drives efficient Transcription of the inserted gene, and a transcription terminator sequence sometimes enhances the yield and Stability of the resulting mRNA. The operator (O) provides regulation via a repressor that binds to it (Chapter 28). The ribosome binding site supplies the signals necessary for efficient Translation of the gene's mRNA. A selectable marker allows for the selection of cells containing recombinant DNA.

Genes can also be cloned and expressed in Eukaryotic cells; various yeast species are commonly used as hosts. A eukaryotic host can sometimes perform post-translational modifications (alterations in Protein Structure occurring after ribosomal synthesis) that may be required for the biological function of the cloned eukaryotic protein.

Alterations in cloned genes lead to The production of modified proteins

Cloning technologies can be employed not only to produce large quantities of protein, but also to synthesize protein products that differ slightly from their natural forms. Specific Amino acids can be replaced individually using Site-Directed Mutagenesis. In this powerful method for investigating protein Structure and function, The amino acid sequence is altered by modifying the DNA sequence within the cloned gene. If suitable restriction sites flank The nucleotide sequence to be modified, researchers can simply excise the DNA segment and replace it with a synthetic analog identical to the original except for the desired change (Fig. 9-11a). In cases where conveniently located restriction sites are absent, modifications can be introduced into a specific DNA sequence using an approach known as oligonucleotide-directed mutagenesis (Fig. 9-11b). A short synthetic DNA strand carrying the specific sequence alteration is hybridized with a single-stranded copy of the cloned gene in a suitable vector. A mismatch of a single base pair in 15–20 does not prevent hybridization if performed at an appropriate temperature. This strand serves as a primer for the synthesis of a chain complementary to the plasmid vector. This double-stranded recombinant plasmid containing a minor mismatch is then used to transform bacteria, in which mismatches are corrected by cellular DNA Repair enzymes (Chapter 25). Approximately half of the repair events will remove and replace the modified base, reverting the gene to its original sequence; the other half will remove and replace the unmodified base, preserving the desired mutation. Transformed cells are thoroughly screened (often by sequencing their plasmid DNA) until a bacterial colony containing the plasmid with the altered sequence is identified.

Fig. 9-11. Two approaches to site-directed mutagenesis. (a) A synthetic DNA segment replaces a DNA fragment removed by restriction endonuclease cleavage. (b) A synthetic oligonucleotide with the desired sequence change at a single site hybridizes with a single-stranded copy of the gene being modified and acts as a primer for the synthesis of double-stranded DNA (containing a single mismatch), which is subsequently used to transform cells. Cellular DNA repair systems correct about 50% of the mismatches, reflecting the desired nucleotide sequence change.

Modifications can also be introduced into more than one base pair. Large regions of a gene can be deleted by excising a segment with restriction endonucleases and ligating the remaining parts to form a smaller gene. Portions of two different genes can be joined to create novel combinations. The product of such a chimeric gene is termed a fusion protein.

Researchers now possess ingenious methods for executing virtually any genetic modification in vitro. Introducing the altered DNA into a cell allows for the Study of the resulting phenotypic changes. Site-directed mutagenesis has dramatically advanced protein research by enabling scientists to introduce specific alterations into a protein's Primary Structure and examine the effects of these changes on conformation, three-dimensional structure, and protein activity.

Terminal sequences provide binding sites for Affinity Chromatography

Affinity chromatography is one of the most effective methods for Protein Purification (see Fig. 3-17c). Unfortunately, for many proteins, suitable ligands that can be immobilized on a solid support for chromatography are not available. The generation of fusion proteins makes it possible to purify virtually any protein using affinity chromatography.

First, a DNA construct is created in which the gene for the protein of interest is joined to a gene encoding a peptide or protein that binds with high affinity and Specificity to a known Ligand. The peptide or protein used for this purpose is attached to either the N- or C-terminus and is referred to as a terminal sequence, or tag. Table 9-3 lists some of the PROTEINS AND Peptides most commonly used as terminal sequences, along with their respective ligands.

Table 9-3. Commonly Used Protein Tags

Tag/peptide

Molecular weight, kDa

Immobilized ligand

Protein A

59

Fc region of IgG

(His)6

0.8

Ni2+

Glutathione S-transferase

26

Glutathione

Maltose-binding protein

41

Maltose

β-Galactosidase

116

o-Aminophenyl-β-D-thiogalactoside (TPEG)

Chitin-binding domain

5.7

Chitin

The experimental Procedure is illustrated using the attachment of glutathione S-transferase (GST) to the target protein. GST is a small enzyme (Mr = 26,000) that binds glutathione with high affinity and specificity (Fig. 9-12). When a construct is engineered in which the GST gene is fused to the gene of the target protein, the resulting chimeric protein acquires The ability to bind glutathione. The chimeric protein is expressed in bacterial or other cells, and a crude cell extract is subsequently prepared. A chromatography Column is packed with a porous support consisting of microscopic beads of a robust polymer (such as cross-linked agarose) onto which the ligand (in this case, glutathione) has been immobilized. All other proteins in the extract pass through the column without binding to the support. GST binds tightly to glutathione, but this interaction is noncovalent, allowing the protein to be easily eluted from the column using a solution containing either a high salt concentration or free glutathione, which competes with the immobilized ligand for GST binding. This approach frequently yields purified protein with high recovery. Sometimes the terminal sequence is partially or completely cleaved from the purified protein using a protease that cleaves the sequence in the region adjacent to the glutathione-binding domain.

Fig. 9-12. Use of fusion constructs for target protein purification. Purification of a protein fused to glutathione S-transferase (GST) is shown as an example. (a) GST is a small enzyme (depicted here as a purple circle) that binds glutathione. (Glutathione is the tripeptide γ-glutamylcysteinylglycine, which contains an unusual peptide bond between the amino group of Cysteine and the side-chain carboxyl carbon of glutamate.) (b) GST is attached to the C-terminus of the target protein using recombinant DNA technology. The chimeric protein is expressed in the host cell and is present in the crude extract following cell lysis. The extract is applied to a chromatography column packed with a support having immobilized glutathione. The chimeric protein binds to glutathione and is retained on the column, whereas other proteins are rapidly washed through. The chimeric protein is subsequently eluted from the column using a solution containing a high salt concentration or free glutathione.

An example of a short terminal sequence with numerous applications is a simple stretch of six or more Histidine residues (6 × His). This sequence binds with high affinity and specificity to nickel ions. A column packed with immobilized nickel ions can be used to effectively separate a histidine-tagged protein from all other proteins in a mixture. Larger terminal sequences, such as maltose-binding protein, can improve the solubility and enhance the stability of target proteins, thereby enabling the isolation of proteins that cannot be purified by other means.

Methods utilizing terminal sequences are convenient and efficient, which accounts for their widespread adoption. However, a certain degree of caution should be exercised when using this approach. The reason is that terminal sequences are not inert. Even very small tags can alter The properties of the proteins to which they are attached and potentially influence experimental results. Protein activity may remain altered even after the removal of the terminal sequence with a protease, for instance, if one or more amino acid residues remain attached to the target protein. Experimental data obtained using this method should always be validated with well-designed controls to account for the effects of the extraneous sequence on the function of the protein under investigation.

Summary of Section 9.1 DNA Cloning: Basic Concepts

DNA Cloning and genetic engineering involve cleaving DNA and joining DNA segments into a novel combination—recombinant DNA.

■ Cloning involves cutting DNA into fragments with enzymes; selecting and optionally modifying the fragment of interest; inserting the DNA fragment into a suitable cloning vector; transferring the vector with the inserted DNA into a host cell for replication; and identifying and selecting the cells containing the DNA fragment.

Key Enzymes in Gene cloning are restriction endonucleases (especially type II enzymes) and DNA ligase.

■ Cloning vectors include plasmids, bacteriophages, and—for the longest DNA inserts—bacterial artificial chromosomes (BACs) and yeast artificial chromosomes (YACs).

■ Cells containing specific DNA sequences can be identified using DNA hybridization methods.

■ Genetic engineering techniques manipulate cells to express and/or modify cloned genes.

■ Proteins or peptides can be genetically engineered to fuse with a target protein, resulting in a chimeric protein. The additional peptide tag can be used to detect the target protein or purify it via affinity chromatography.



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.