Fundamentals of Bioinformatics - Ogurtsov A.N. 2013

Foundations of Bioinformatics
Genomes and Proteomes
DNA Sequencing Methods

Several Methods for determining The nucleotide sequence of DNA are known. One such method is referred to as chain-termination sequencing, dideoxy sequencing, or (named after its inventor Frederick Sanger) the Sanger dideoxy method (see [7], p. 11.2) (Figure 41).

Class="center">

Figure 41 - Diagram of the Sanger DNA Sequencing method

The core sequencing reaction involves the following Reagents: a single-stranded DNA template; a primer to initiate polymerization of the synthesized strand; four deoxyribonucleoside triphosphates, dNTPs (dATP, dGTP, dTTP, and dCTP); four dideoxynucleoside triphosphates, ddNTPs (ddATP, ddGTP, ddTTP, and ddCTP); and a DNA polymerase enzyme that incorporates complementary NUCLEOTIDES into the growing DNA strand using the template strand.

Sanger's dideoxy chain-termination DNA sequencing begins with the Denaturation of The Double Helix of the DNA fragment to yield single-stranded templates for in vitro DNA Synthesis.

A synthetic oligodeoxynucleotide is used as a primer for four independent polymerization reactions, each employing a low concentration of one of the four ddNTPs In addition to a high concentration of normal dNTPs (Figure 41).

In each reaction, a ddNTP is randomly incorporated into the growing DNA chain at THE POSITION OF the corresponding dNTP, terminating further polymerization at that position. Each of the four reaction mixtures synthesizes a set of DNA fragments of varying lengths, sharing a common start and terminating at specific bases (of the same type but at different positions in the sequence). The mixture of shortened fragments obtained from each reaction is denatured and analyzed by gel Electrophoresis.

The dideoxy DNA sequencing method is fully automated. Each reaction mixture is labeled with a specific fluorescent tag (either on the primer or on the substrate of one of the nucleotides, such as a dideoxynucleoside triphosphate specific to that mixture), which subsequently allows a scanner to determine the terminal bases of all fragments.

The four reagent mixtures are then combined in a single container, and the DNA fragments are separated by Polyacrylamide gel electrophoresis (PAGE), in which smaller DNA fragments migrate faster than larger ones.

Through this method, the set of DNA fragments is sorted by size. The resolving power of PAGE allows the Separation of polynucleotides differing in length by as little as a single residue. Near the end of the lanes, a scanner (photodetector) reads the fluorescent tag of each passing DNA fragment, and this information is converted into lane-tracking data presented as a chart composed of a group of colored peaks corresponding to specific bases (Figure 42).

Figure 42 - Example of a high-quality lane-tracking chart. Peaks are typically printed in different colors to facilitate visual interpretation. Software such as Phred reads the peaks and assigns them A, C, G, and T nucleotide values

Decoded DNA sequences are stored in Databases. Various databases exist for different DNA sequences, including genomic DNA, complementary cDNA, and recombinant DNA. Genome sequencing is performed using the shotgun method, chromosome walking, or the UTR clone assembly strategy (UTR stands for UnTranslated Region). Many different software tools are used to verify the quality of decoded sequences, for example: Phred, Vector_clip, CrossMatch, RepeatMaster, Phrap, and Staden-Gap4.

The advent of high-throughput automated fluorescently labeled DNA sequencing technology has led to a rapid accumulation of sequence information. This information, in turn, provides the foundation for deriving protein sequence data through computational methods.

Numerous types of research rely on DNA sequence analysis; Examples include Phylogenetic Analysis, Introduction/32.html">Genetic Engineering and restriction mapping, Gene Structure determination via the prediction of introns and exons, and the Analysis of Protein-coding sequences using Open Reading Frames (ORFs), among others.

According to the Central dogma of molecular biology, DNA is transcribed into RNA, which is then translated into protein (Figure 43).

Figure 43 - Diagram of Gene Expression

In eukaryotic systems, exons form part of the final coding sequence, whereas introns, although transcribed, are excised by splicing machinery before the mRNA assumes its final, mature form. DNA sequence databases typically contain information at the level of untranslated genomic sequences, introns and exons, mRNA, cDNA, and Translation products.

Untranslated regions (UTRs) are found in both DNA and RNA. UTRs are segments of the transcribed sequence that flank the coding sequence on both sides and are not translated into protein. The untranslated sequence, particularly the one located at the 3' end of the coding sequence, is highly specific to both the gene itself and the biological species possessing that coding sequence.



Last update: 11/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.