LEHNINGER PRINCIPLES OF BIOCHEMISTRY - VOL. 3. INFORMATION PATHWAYS - 2017

PART III. INFORMATION PATHWAYS

25. DNA METABOLISM

25.3. DNA Recombination

The rearrangement of Genetic information within DNA molecules occurs through A wide variety of processes collectively known as genetic recombination. Today, genetic rearrangements find Structure/182.html">Practical Application in modifying an ever-growing number of genomes (ch. 9 in vol. 1).

Three MAIN TYPES OF genetic recombination are distinguished. Homologous genetic recombination (general recombination) involves the exchange of genes between any two DNA molecules (or segments of the same molecule) that share extended regions of nearly identical sequence. The specific sequences do not matter, provided they are sufficiently similar. In Site-Specific Recombination, the exchange is restricted to specific DNA sequences. DNA transposition differs from the other two types in typically involving a short segment of DNA that can move from one site to another within a chromosome. Such "jumping genes" were first discovered by Barbara McClintock in the 1940s in the maize genome. In addition, various other unusual genetic rearrangements exist, whose mechanisms and significance remain to be fully established. In this book, we focus exclusively on these three main types of recombination.

Class="center">

The Functions of genetic recombination systems are as diverse as their mechanisms. They include roles in specialized DNA Repair pathways, specific activities during METABOLISM/36.html">DNA Replication, regulation of certain Gene Expression, facilitation of chromosome segregation during Introduction/5.html">Eukaryotic Cell division, maintenance of genetic diversity, and execution of programmed genetic rearrangements during embryonic development. In most cases, genetic recombination is closely intertwined with other DNA metabolic processes, a relationship we will explore further below.

Homologous genetic recombination serves several functions

In Bacteria, homologous genetic recombination functions in DNA repair and, as noted in Section 25.2, is termed recombinational DNA repair. It typically acts to rescue replication forks stalled at DNA Lesions. Homologous genetic recombination can also occur during conjugation (mating), when chromosomal DNA is transferred from a donor bacterial cell to a recipient cell. Although recombination during conjugation is a relatively rare event in bacterial populations, it contributes to genetic diversity.

In eukaryotes, homologous genetic recombination performs several functions during replication and Cell Division, including the repair of stalled replication forks. Most frequently, recombination occurs during Meiosis, the specialized process by which diploid Germ Cells divide to produce haploid Gametes (sperm or egg cells in animals, or haploid spores in plants); each gamete receives one chromosome from each homologous pair (Fig. 25-31). Meiosis begins with DNA replication in the parent cell, such that at one point each DNA molecule is present in four copies. The cells then undergo two rounds of cell division without an intervening phase of DNA replication, leaving each gamete with a haploid Complement of DNA.

Fig. 25-31. Meiosis in the animal germ line. Chromosomes of a hypothetical diploid germ cell (six chromosomes; three homologous pairs) replicate and remain joined at their centromeres. The replicated double-stranded DNA molecules are called chromatids (sister chromatids). In prophase I, immediately preceding the first meiotic division, the three sets of homologous chromatids align to form tetrads and remain covalently joined at points of homologous connection (chiasmata). Crossing Over occurs at these chiasmata (Fig. 25-32). These temporary associations between homologs ensure the proper segregation of chromosomes in the subsequent stage, as they move toward opposite poles of the dividing cell During the first meiotic division. The products of this division are two daughter cells, each containing three pairs of chromatids. The pairs then align along The Cell equator in preparation for the Separation of chromatids (now referred to as chromosomes). The second meiotic division yields four haploid daughter cells that can function as gametes. Each of these cells contains three chromosomes—half the complement of the precursor diploid cell. The chromosomes have undergone reassortment and recombination.

Following DNA replication in prophase of the first meiotic division, the resulting sister chromatids remain linked by their centromeres. At this stage, each set of four homologous chromosomes (a tetrad) exists as two pairs of chromatids. Genetic information is exchanged between the closely juxtaposed homologous chromosomes via homologous genetic recombination, a process involving breakage and rejoining of DNA sequences (Fig. 25-32). This process, also known as crossing over, can be observed under a Light Microscope. Crossing over holds the two pairs of sister chromatids together at points called chiasmata.

Fig. 25-32. Crossing over. (a) Crossing over frequently results in the EXCHANGE OF GENETIC material. (b) Homologous chromosomes of grasshoppers shown at the prophase I stage of meiosis. Several contact points (chiasmata) are visible between the two homologous chromatid pairs. These chiasmata provide clear evidence of prior Homologous Recombination (crossing over).

Crossing over effectively and physically links all four homologous chromosomes, which is essential for proper chromosome segregation in subsequent meiotic cell divisions. Clearly, crossing over is not a random process; "hotspots" have been identified on many eukaryotic chromosomes. However, treating crossing over as though it can occur at virtually any point along homologous chromosomes serves as a useful approximation in genetic mapping. The probability of homologous recombination occurring between any two points on a chromosome is roughly proportional to the distance between them, allowing the relative positions of genes and the distances between them to be determined.

In summary, homologous recombination serves at least three functions: (1) it is utilized in repairing several types of DNA damage; (2) in Eukaryotic cells, it provides a temporary physical linkage between chromatids that is required for accurate chromosome segregation during the first meiotic division; and (3) it enhances genetic diversity within a population.

Meiotic recombination is initiated by double-strand breaks

A putative pathway for homologous recombination during meiosis is illustrated in Figure 25-33a. This model highlights four key features. First, homologous chromosomes align. Second, an exonuclease acts at a double-strand break in the DNA molecule to generate single-stranded regions with a 3'-hydroxyl group at the terminus (stage ①). Third, these 3' ends invade the intact duplex DNA of the homolog, followed by branch migration (Fig. 25-34) and/or replication to generate a pair of crossover structures (Holliday junctions; Fig. 25-33a, stages ②–④). Fourth, the resolution of these two structures yields two final recombinant products (stage ⑤).

Fig. 25-33. Meiotic recombination. (a) A double-strand break repair model for homologous genetic recombination. The two homologous chromosomes participating in recombination (one shown in blue, the other in red) share sequence similarity. Each of the two genes shown is present on the two chromosomes as different alleles. The Stages of the process are described in the text. (b) A Holliday junction formed in vivo between two bacterial Plasmids, visualized by Electron Microscopy. These structures are named after Robin Holliday, who first predicted their existence in 1964.

In this double-strand break repair model, the 3' ends serve as primers to initiate genetic exchange during recombination. Their base pairing with the complementary strand of the intact homolog creates a region of hybrid DNA containing complementary strands from two different parental DNA molecules (the product of stage ② in Fig. 25-33a). Subsequently, each 3' end can act as a primer for DNA Synthesis. The resulting structures, termed Holliday junctions (Fig. 25-33b), are a hallmark of homologous genetic recombination across all organisms.

While homologous recombination may vary in minor details among species, most of the stages described above are generally conserved in some form. There are two modes of cleaving, or "resolving," the Holliday junction, which can yield either the same gene order following recombination as in the parental chromosomes or a different arrangement (stage ⑤ in Fig. 25-33a). One Cleavage pathway results in flanking regions of hybrid DNA that are recombinant, whereas the other yields non-recombinant flanking regions. Both outcomes are observed in vivo in eukaryotes and prokaryotes.

Homologous recombination, depicted in Fig. 25-33, is a complex process that contributes to genetic diversity. To understand how this occurs, recall that two recombining homologous chromosomes are not necessarily identical. Their linear gene order may be the same, but the base sequences in some genes can differ slightly (representing different alleles). For example, in humans, one chromosome may carry the Hemoglobin A allele (normal hemoglobin), while the other carries the hemoglobin S allele (sickle-cell mutation). The difference may be as little as a single base pair per million. Homologous recombination does not alter the linear order of genes, but it can determine which alleles become linked on the same chromosome and are inherited together by the next generation.

Fig. 25-34. Branch migration. When a template strand pairs with two different complementary strands, a junction forms at the point where three strands meet. This junction "migrates" as base pairing between complementary strands is disrupted and new base pairing is established between one of the complementary strands and an incoming strand. In the absence of a guiding enzyme, the branch can move spontaneously in either direction. Such spontaneous branch migration continues until an non-identical sequence is encountered in one of the complementary strands.

Numerous Enzymes and other Proteins participate in recombination

Enzymes that stimulate various stages of homologous recombination have been isolated from both PROKARYOTES AND EUKARYOTES. In E. coli, the recB, recC, and recD genes encode the heterotrimeric enzyme RecBCD, which exhibits both helicase and nuclease activities. The RecA protein stimulates all key stages of homologous recombination: the pairing of two DNA molecules, The formation of Holliday intermediate structures, and branch migration (as described below). The RuvA and RuvB proteins (derived from "repair of UV damage") form a complex that binds to Holliday intermediates, displaces RecA, and stimulates branch migration at a much higher rate than RecA alone. Nucleases that specifically cleave Holliday structures are often called resolvases; they have been isolated from bacteria and Yeast, with RuvC being one of at least two such nucleases in E. coli.

The RecBCD enzyme binds to linear DNA at a free (cleaved) end and moves along The Double Helix, unwinding and cleaving the DNA in a reaction coupled to ATP Hydrolysis (Fig. 25-35). The RecB and RecD subunits act as helicase motors: RecB moves along one strand in the 3' —> 5' direction, while RecD moves along the other strand in the 5' —> 3' direction. The enzyme's activity changes when it encounters a sequence (5') GCTGGTGG, known as a chi sequence. From this point on, the degradation of the 3'-tailed strand is significantly slowed, whereas the degradation of the 5'-tailed strand is accelerated. As a result, a single-stranded DNA with a free 3' end is generated, which is utilized in subsequent stages of recombination (Fig. 25-33). The E. coli genome contains 1009 dispersed chi sequences, which enhance recombination frequency by 5- to 10-fold over distances of up to 1,000 bp surrounding them. Their influence diminishes with distance from the chi sequence. Sequences that stimulate recombination have also been identified in several other organisms.

Fig. 25-35. Helicase and nuclease activities of the RecBCD enzyme. Starting at a double-stranded end, RecBCD unwinds and cleaves the DNA sequence. The nuclease domain of the RecB subunit is positioned to degrade either One DNA strand or the other. The RecB subunit moves in the 3' —> 5' direction along the single-stranded DNA with a free 3' end, acting as an ATP-dependent helicase. The RecD subunit moves similarly along the second strand in the 5' —> 3' direction. The Active Site of the RecC subunit recognizes and binds chi sequences. The strand containing the chi sequence forms a loop, leaving only the 5'-tailed strand accessible to the RecB nuclease domain. It is hypothesized that RecBCD initiates genetic recombination in E. coli. It also participates in the Repair of Double-stranded breaks at sites of collapsed replication forks.

The RecA protein stands out among DNA metabolism proteins because its active form is an ordered helical filament containing up to several thousand RecA monomers assembled cooperatively on DNA (Fig. 25-36). The RecA filament normally forms on single-stranded DNA, for instance, through the action of the RecBCD enzyme, but it can also assemble on double-stranded DNA containing a single-stranded gap. In this case, the initial RecA monomers bind to the single-stranded gap, and the resulting filament rapidly wraps around the adjacent duplex DNA. The assembly and disassembly of RecA filaments are regulated by other proteins, including RecX, DinI, RecF, RecO, and RecR.

Fig. 25-36. The RecA protein. (a) Nucleoprotein complex of the RecA protein with single-stranded DNA (electron micrograph). The transverse striations indicate that the filament structure is a right-handed helix. (b) Molecular model of a filament composed of 24 RecA subunits. The filament contains six subunits per turn. One subunit is colored red for visibility against the others (from PDB ID 2REB). (c) Following the rate-limiting nucleation of RecA on single-stranded DNA, filaments elongate in the 5' —> 3' direction. Disassembly also occurs in the 5' —> 3' direction from the end opposite to growth. (d) RecF, RecO, and RecR proteins (collectively known as RecFOR) assist in filament assembly. The RecX protein inhibits RecA filament extension. The DinI protein stabilizes RecA by preventing disassembly.

The Role of the RecA filament in recombination can be illustrated using an in vitro DNA strand-exchange model (Fig. 25-37). First, RecA binds a single DNA strand to form a nucleoprotein filament. The RecA filament then captures a homologous double-stranded DNA molecule and pairs it with the bound single strand. Binding (but not hydrolysis) of ATP is required to form the active RecA filament that bridges the two DNA molecules. Subsequently, the two DNA molecules exchange strands, resulting in the formation of hybrid DNA. The exchange occurs at a rate of 6 bp/s in the 5' —> 3' direction relative to the single-stranded DNA within the RecA filament. Three or four DNA strands can participate in this reaction (Fig. 25-37); in the latter case, Holliday structures are formed during the process.

Fig. 25-37. In vitro DNA strand exchange induced by the RecA protein. Strand replacement involves the separation of one strand of a duplex from its complementary strand and its transfer to another complementary strand to form a new DNA duplex (heteroduplex). This process generates a branched intermediate structure. The Formation of the final product depends on branch migration, which is promoted by RecA. The reaction can involve three strands (left) or, during mutual exchange between two homologous duplexes, four strands (right). In the latter case, the intermediate is a Holliday structure. The RecA protein promotes branch migration using the energy of ATP hydrolysis.

When a DNA duplex is incorporated into the RecA filament and aligned alongside the bound single-stranded DNA over a region of hundreds of Base Pairs, one of the duplex strands can switch partners (Fig. 25-38, step ②). Because DNA has a helical structure, continuous strand exchange requires coordinated Rotation of the two adjacent DNA molecules. This is accomplished through unwinding/rewinding (steps ③ and ④), which shifts the branch point along the helix. The final Stages of DNA strand exchange, during which the hybrid DNA formed in the initial step is extended, are driven by ATP hydrolysis. The mechanism coupling these reactions remains unclear.

Fig. 25-38. Model of RecA-mediated DNA strand exchange. A reaction involving three DNA strands is shown. The spheres representing RecA protein molecules are disproportionately small relative to the thickness of the DNA to clearly demonstrate the changes occurring within the DNA. ① The RecA protein forms a helix around a single-stranded DNA. ② A homologous DNA duplex is incorporated into the complex. ③ As rotation shifts the beginning of the triple-stranded region from left to right, one strand of the duplex takes THE PLACE OF the single strand originally bound to the protein filament. The other duplex strand is displaced, and a new duplex is formed within the protein filament. As rotation continues (④ and ⑤), the displaced strand is fully released. In this model, ATP hydrolysis by RecA drives the relative rotation of the two DNA molecules, thereby advancing strand exchange from left to right.

Following the formation of a Holliday structure, a whole group of enzymes is required to complete recombination: topoisomerases, the RuvAB branch migration protein, resolvases and other nucleases, DNA polymerase I or III, and DNA ligase. In E. coli cells, the RuvC protein (Mr = 20,000) cleaves Holliday structures to yield full-length, unbranched chromosomes.

All avenues of DNA metabolism are utilized to repair stalled replication forks

As in all other cells, the level of DNA damage in bacteria remains quite high even under normal growth conditions. Most DNA lesions are rapidly repaired by base Excision Repair, nucleotide excision repair, and other pathways described previously. Nevertheless, nearly every bacterial Replication fork encounters an unrepaired DNA lesion or break on its path from the origin to the terminus of replication (Fig. 25-30). DNA polymerase III cannot bypass many types of DNA damage, forcing it to leave these lesions behind as single-stranded gaps. When the polymerase encounters a single-strand break, a double-strand break is generated. In both situations, recombinational DNA repair is required (Fig. 25-39). Under normal growth conditions, stalled replication forks are reactivated through a complex repair pathway involving recombinational DNA repair, replication restart, and the repair of any bypassed lesions. This process integrates all pathways of DNA metabolism.

Fig. 25-39. Recombinational repair models for rescuing stalled replication forks. A replication fork stalls upon encountering a DNA lesion (left) or a DNA break (right). Recombination enzymes that facilitate DNA strand exchange are required to re-establish the branched DNA Structure at the replication fork. Single-strand gaps are repaired with the assistance of RecF, RecO, and RecR proteins. Double-strand breaks require the RecBCD enzyme. The RecA protein is involved in both processes. Recombination intermediates are processed through The activity of additional enzymes (e.g., RuvA, RuvB, and RuvC proteins are needed to resolve Holliday structures). Damage in the double-stranded DNA is repaired via nucleotide excision repair or other pathways. The replication fork is reassembled with the recruitment of enzymes that catalyze origin-independent replication restart, allowing Chromosome replication to proceed to completion. The entire process requires precise coordination of all bacterial DNA metabolism pathways.

A stalled replication fork can be restarted via at least two complex pathways, both of which require the RecA protein. The repair pathway for lesions containing DNA gaps also utilizes the RecF, RecO, and RecR proteins. Repair of double-strand breaks requires the RecBCD enzyme (Fig. 25-39). Subsequent recombination steps are followed by a process known as replication restart (or origin-independent replication reinitiation), in which a replication fork is re-assembled by a complex of seven proteins (PriA, PriB, PriC, DnaB, DnaC, DnaG, and DnaT). This complex, originally discovered during in vitro DNA replication of phage φX174, is now called the replication restart primosome. DNA polymerase II is also required to rescue the replication fork, although its exact role remains to be established; the activity of DNA polymerase II paves the way for DNA polymerase III to synthesize the extensive stretches required to complete chromosome assembly. Occasionally, replication restarts downstream of the damage site even before the damaged region itself has been repaired.

The repair of stalled forks involves a coordinated transition between replication and recombination. Recombination steps are necessary to fill in DNA gaps or connect a severed strand, thereby re-establishing the branched DNA structure at the replication fork. Lesions left behind the advancing DNA duplex are cleared by base or nucleotide excision repair. Thus, a broad spectrum of enzymes operating across every stage of DNA metabolism participates in rescuing stalled replication forks. This repair mechanism appears to be The primary function of the homologous recombination system in every cell, and defects in recombinational DNA repair play a major role in The Development of human diseases (Box 25-1).

Site-specific recombination leads to precise DNA rearrangements

The homologous genetic recombination we have just discussed can occur between any two homologous sequences. A second type of recombination, site-specific recombination, is fundamentally different in that it is restricted exclusively to specific sequences. Reactions of this type occur in nearly every cell, and their functions vary widely across different organisms. For example, this mechanism regulates the expression of certain genes and drives programmed DNA rearrangements during embryonic development or within the replication cycles of certain viral and plasmid DNAs. Each site-specific recombination system consists of a recombinase enzyme and a short (20 to 200 bp) unique DNA sequence—the recombination site—with which the recombinase interacts. The course and outcome of the reaction are regulated by one or more accessory proteins.

There are two Major Classes of site-specific recombination systems, distinguished by whether their active-site enzymes contain Tyr or Ser residues. The Study of various Tyrosine-class site-specific recombination systems in vitro has helped elucidate the core mechanisms and General Principles of the process (Fig. 25-40, a). Some of these enzymes have been obtained in crystalline form, allowing for a detailed examination of their structure. The isolated recombinase recognizes and binds to each of the two recombination sites on two different DNA molecules or within the same DNA molecule. At each site, one DNA strand is cleaved at a specific point, after which the recombinase becomes covalently attached to the DNA at the cleavage site via a phosphotyrosine linkage (step ①). This transient protein–DNA bond preserves the phosphodiester bond that is lost during DNA cleavage, so subsequent steps do not require energy-rich Cofactors such as ATP. The cleaved DNA strands are joined to new partners to form a Holliday structure, in which new phosphodiester bonds are created at the expense of the protein–DNA linkage (step ②). To complete the reaction, this process must be repeated at a second point in each of the two recombination sites (steps ③ and ④). In systems with an active-site Ser residue, both strands at each recombination site are simultaneously cleaved and rejoined into new pairs without the formation of Holliday structures. In both cases, however, exchange is always reciprocal and precise, and the recombination sites are restored upon completion of the reaction. The recombinase thus acts simultaneously as a site-specific endonuclease and ligase.

Fig. 25-40. The site-specific recombination reaction. (a) The reaction shown here is characteristic of integrase-class site-specific recombinases (named following the discovery of bacteriophage $\lambda$ integrase, the first recombinase to be described). This enzyme features active-site Tyr residues that act as nucleophiles. The recombinase is a tetramer composed of identical subunits that bind to a specific sequence—the recombination site. ① One strand of each DNA molecule is cleaved at specific points within this sequence. The nucleophilic agent is the -OH group of an active-site tyrosine residue, resulting in the formation of a covalent phosphotyrosine linkage between the protein and the DNA. ② The cleaved strands join with new partners to form a Holliday structure. The concluding steps ③ and ④ mirror the first two. The original recombination site sequence is restored following recombination of the DNA flanking the site. These steps proceed with the participation of a multimeric recombinase complex and, occasionally, other proteins not shown here. (b) A space-filling model of the four-subunit Cre recombinase of the integrase class bound to a Holliday structure (depicted as cyan and blue helices). The protein is shown transparently so that the DNA is visible through it (based on PDB ID 3CRX). In the active center of recombinases belonging to the resolvase-invertase family, a Serine residue serves as the nucleophile.

The recombination site sequences recognized by site-specific recombinases are partially asymmetric (non-palindromic); during the reaction, two recombination sites align in the same orientation. The outcome depends on the localization and orientation of the recombination sites (Fig. 25-41). If two sites reside on the same DNA molecule, the reaction either inverts or deletes the intervening DNA, depending on whether the recombination sites have opposite or identical orientations, respectively. If the sites are located on different DNA molecules, the recombination is intermolecular; if one or both DNAs are circular, insertion occurs. Some recombinase systems are specific for just one of these reaction types and act only on sites in a particular orientation.

Fig. 25-41. Outcomes of site-specific recombination. The outcome of site-specific recombination depends on the localization and orientation of the recombination sites (red and green) within the double-stranded DNA molecule. Here, orientation (indicated by arrows) refers to the arrangement of NUCLEOTIDES within the recombination site rather than the 5' -> 3' direction. (a) Recombination sites with opposite orientation on the same DNA molecule, resulting in inversion. (b) Recombination sites with the same orientation located on the same DNA molecule (resulting in deletion) or on two separate DNA molecules (resulting in insertion).

The first site-specific recombination system to be studied in vitro was the one encoded by bacteriophage $\lambda$. When $\lambda$ phage DNA enters an E. coli cell, a complex, regulated sequence of events leads to one of two outcomes. Either the phage DNA replicates and produces more Bacteriophages (which lyse the host cell), or it integrates into the host chromosome (as a prophage) and replicates passively as part of the chromosome over many cell generations. Integration is carried out by a phage-encoded recombinase ($\lambda$ integrase), which acts at recombination sites in the phage and bacterial DNAs—the attP and attB attachment sites, respectively (Fig. 25-42). The role of site-specific recombination in gene expression regulation is discussed in Chapter 28.

Fig. 25-42. Integration and excision of bacteriophage $\lambda$ DNA at a specific chromosomal site. The attachment site in $\lambda$ phage DNA (attP) shares only 15 bp of complete Homology with nucleotides in the bacterial DNA (attB) within the crossover region. The reaction generates two new attachment sites (attR and attL) flanking the integrated phage DNA. The phage recombinase is $\lambda$ integrase (the INT protein). Integration and excision utilize different attachment sites and distinct auxiliary proteins. Excision involves phage-encoded XIS proteins and bacteria-encoded FIS proteins. Both reactions require the bacteria-encoded IHF (integration host factor) protein.

Site-Specific Recombination May Be Required for Complete Chromosomal Replication

Recombinational Repair of DNA in circular bacterial chromosomes sometimes produces detrimental by-products. When a Holliday structure is resolved by a nuclease such as RuvC, followed by the completion of replication, the result can be either two normal monomeric chromosomes or a linked dimeric chromosome (Fig. 25-43). In the latter case, the covalently linked chromosomes cannot segregate into daughter cells during cell division, arresting the Cell Cycle at this stage. E. coli possesses a specialized XerCD site-specific recombination system that converts dimeric chromosomes into monomers, thereby ensuring proper cell division. This site-specific deletion (Fig. 25-41, b) serves as yet another example of the close interconnection between DNA recombination and other DNA metabolic processes.

Fig. 25-43. DNA deletion eliminating the hazardous consequences of recombinational DNA repair. Resolution of the Holliday structure during recombinational DNA repair (cleavage sites marked by red arrows) can lead to the formation of a dimeric chromosome. In E. coli, the specialized XerCD site-specific recombinase converts the dimer into monomers, facilitating chromosome segregation and normal cell division.

Mobile Genetic Elements Move from One DNA Site to Another

We now turn to a third major type of recombination, which involves the movement of mobile elements, or Transposons. Found in virtually all cells, these DNA segments move or "jump" from one Location in a chromosome (the donor site) to another location on the same or a different chromosome (the target site). DNA Sequence homology is not a prerequisite for this movement, termed transposition; the new location of a transposon is determined more or less at random. Because the insertion of a transposon into a gene can be lethal to the cell, transposition is tightly regulated and typically occurs at a very low frequency. Transposons are essentially the simplest molecular parasites, having adapted to replicate passively within host cell chromosomes. In some instances, however, they carry genes beneficial to the host cell and thus maintain a symbiotic relationship with it.

Bacteria possess two major classes of transposons. Insertion sequences (IS elements or simple transposons) contain only the sequences required for transposition along with the genes for the enzymes (transposases) that mediate the process. Complex transposons harbor one or more additional genes besides those strictly necessary for transposition. These extra genes may, for example, confer Antibiotic Resistance, thereby increasing the host cell's chances of survival. Transposition is one of the driving forces behind the spread of antibiotic resistance in pathogenic bacterial populations, which can render certain antibiotic treatments ineffective (p. 10). ■

Bacterial transposons vary in structure, but most feature short terminal repeats that bind transposases. During transposition, a short sequence at the target site (5 to 10 bp) is duplicated to generate an additional short repeat that flanks the inserted transposon on both sides (Fig. 25-44). These repeats are generated through sequence cleavage and the subsequent insertion of the transposon at its new site.

Fig. 25-44. Duplication of target DNA sequences upon transposon insertion. Sequences duplicated following transposon insertion are shown in red. Typically, these sequences comprise only a few base pairs, though they are greatly enlarged here relative to a typical transposon.

Bacteria utilize two primary pathways of transposition. In direct (simple) transposition (Fig. 25-45, left), Cleavage of the DNA on either side of the transposon excises it, allowing it to move to a new location. This leaves a double-stranded break in the donor DNA that must be repaired. A staggered cut is made at the target site (as in Fig. 25-44) into which the transposon is inserted, and the resulting gaps are filled by replication, duplicating the target site sequence. In replicative transposition (Fig. 25-45, right), the entire transposon is replicated, leaving one copy behind in the donor DNA. The intermediate structure, consisting of the donor covalently linked to the target DNA, is termed a cointegrate. A cointegrate contains two complete copies of the transposon in the same orientation. For certain well-characterized transposons, this intermediate cointegrate structure is resolved into separate products via site-specific recombination mediated by specialized recombinases.

Fig. 25-45. The two major modes of transposition: direct (simple) and replicative. ① DNA is first cleaved on each side of the transposon (arrows). ② The newly freed 3'-hydroxyl groups at the transposon ends act as nucleophiles, directly attacking phosphodiester bonds in the target DNA. These phosphodiester bonds in the two DNA strands are staggered (not directly opposite one another). ③ The transposon becomes linked to the target DNA. In direct transposition (left), replication fills the gaps at each end. In replicative transposition (right), the entire transposon is replicated to yield a cointegrate. ④ The cointegrate is often subsequently resolved by a specialized site-specific recombination system. The host DNA left cleaved after direct transposition is either repaired by DNA end-joining or degraded (not shown); the latter outcome can be lethal to the Organism.

Eukaryotes also harbor transposons structurally reminiscent of bacterial transposons, and some of these utilize a similar transposition mechanism. In other cases, however, RNA intermediates participate in transposition. The evolution of such transposons shares much in common with that of certain classes of RNA Viruses. Both are discussed in greater detail in the following chapter.

Immunoglobulin Gene Assembly Proceeds via Recombination

Some DNA rearrangements are developmentally programmed in eukaryotic organisms. A prime example is the assembly of complete immunoglobulin genes from individual segments within vertebrate genomes. Humans, like other mammals, can produce millions of distinct IMMUNOGLOBULINS (Antibodies) with varying specificities, even though The Human Genome contains only about 29,000 genes. This recombination allows the organism to generate an extraordinary diversity of antibodies despite the limited coding capacity of DNA. Investigation into recombination mechanisms has revealed a close connection to DNA transposition; this antibody-diversity system likely originated in the distant past as a result of transposons invading cells.

Let us examine The Mechanism of generating diverse antibodies using human genes encoding immunoglobulins G (IgG) as an example. Immunoglobulins consist of two identical heavy and two identical light polypeptide chains (see Fig. 5-21 in Vol. 1). Each chain contains two regions: a variable region, whose sequence differs significantly from one immunoglobulin to another, and a constant region, which is essentially invariant for a given class of immunoglobulins. There are also two different light-chain families—κ and λ—which differ slightly in the sequences of their constant regions. The Diversity of the variable regions across all Three types of polypeptide chains (heavy, κ light, and λ light) arises through a similar mechanism. The genes for these Polypeptides are split into segments, and The Genome contains clusters with multiple versions of each segment. Joining one version of each segment creates a complete gene.

Figure 25-46 illustrates the Organization of DNA encoding human IgG κ light chains; the mechanism by which a mature κ light chain is formed relies precisely on this DNA organization. In undifferentiated cells, the information for this polypeptide chain is divided into three segments. The V (variable) segment encodes the first 95 amino acid residues of the variable region, the J (joining) segment encodes the remaining 12 residues of the variable region, and the C segment encodes the constant region. The genome contains roughly 300 Variants of the V segment, four variants of the J segment, and a single C segment.

Fig. 25-46. V and J segment recombination of the human IgG κ-light-chain gene. This process generates antibody diversity. Top: The organization of IgG coding sequences in a Bone Marrow stem cell. Recombination involves the deletion of the DNA sequence between a specific V segment and a specific J segment. Following Transcription, the RNA transcript undergoes splicing as described in Chapter 26. Translation yields a light chain that can combine with any of 5,000 possible heavy chains to form a complete antibody.

Following the differentiation of a bone marrow stem cell into a mature B lymphocyte, one V segment and one J segment are joined via a specialized recombination system (Fig. 25-46). During this programmed deletion, all intervening DNA is discarded. There are approximately 300 • 4 = 1,200 possible V-J combinations. The recombination process is not as precise as the site-specific recombination described above, so the junction between the V and J segments can serve as an additional source of sequence Variability. This increases the number of possible variants by about 2.5-fold, enabling cells to generate roughly 2.5 • 1,200 = 3,000 different V-J segment combinations. Finally, following transcription, RNA splicing joins the V-J region to the C region (these processes are detailed in Chapter 26).

The recombination mechanism for Joining V and J segments is shown in Fig. 25-47. Immediately downstream of each V segment and upstream of each J segment lie recombination signal sequences (RSS), to which the RAG1 and RAG2 (recombination activating gene) proteins bind. The RAG proteins catalyze the formation of a double-strand break between the signal sequences and the V (or J) segments to be joined. The V and J segments are subsequently joined by a second protein complex.

Fig. 25-47. Mechanism of immunoglobulin gene rearrangement. RAG1 and RAG2 proteins bind to recombination signal sequences (RSS) and cleave a single DNA strand between the RSS and the V (or J) segments to be subsequently joined. The released 3'-hydroxyl group then acts as a nucleophile, attacking the phosphodiester bond in the opposite strand to create a double-strand break. The resulting hairpins at the V and J segment regions are cleaved, and their ends are covalently joined by a protein complex specialized in non-homologous end joining of double-strand breaks. The double-strand break formation steps catalyzed by RAG1 and RAG2 mirror the stages of transposition reactions.

Heavy-chain and λ-light-chain genes are assembled through analogous processes. Heavy chains possess more gene segments than light chains, allowing for over 5,000 possible combinations. Because any heavy chain can combine with any light chain during immunoglobulin formation, each individual can produce at least 3000 • 5000 = 1.5 • 107 different IgGs. Additional diversity arises from a high frequency of Mutations (by an unknown mechanism) in the V-segment genes during B-lymphocyte differentiation. Each mature B lymphocyte produces antibodies of only a single type, but the repertoire of antibodies produced across different cells is vast.

Did The Immune System indeed originate, in part, from ancient transposons? The Mechanism of double-strand break formation mediated by RAG1 and RAG2 is strikingly identical to several stages of transposition (Fig. 25-47). Furthermore, the excised DNA flanked by RSS ends shares the structure found in most transposons. In in vitro experiments, RAG1 and RAG2 proteins can interact with this excised DNA and integrate it, much like a transposon, into other DNA molecules (this reaction likely occurs quite rarely in B lymphocytes). Although not definitively proven, the immunoglobulin gene rearrangement system may be the evolutionary outcome of an ancient blurring of lines between parasite and host.

Summary of Section 25.3 DNA Recombination

■ DNA sequences undergo rearrangements via recombination reactions that are typically tightly coordinated with DNA replication or repair.

■ Homologous genetic recombination can occur between any two DNA molecules sharing homologous sequences. During meiosis (in eukaryotes), this type of recombination helps ensure proper chromosome segregation and enhances genetic diversity. In both bacteria and eukaryotes, recombination serves to rescue stalled replication forks. Homologous recombination proceeds via a Holliday intermediate structure.

■ Site-specific recombination occurs exclusively at specific sequences and may also proceed through a Holliday junction intermediate. Recombinases cleave DNA at specific sites and rejoin the strands with new partners. This type of recombination is found in virtually all cells; its diverse functions include DNA Integration and expression regulation.

■ Transposons, present in virtually all cells, move within or between chromosomes via recombination. In vertebrates, programmed recombination reminiscent of transposition drives the assembly of immunoglobulin gene segments during B-lymphocyte differentiation.



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.