Fundamentals of Bioinformatics - Ogurtsov A.N. 2013
Information Principles in Biotechnology
Biological Sequence Sequencing and Gene Expression
Sequence Determination
A clone is a copied DNA fragment identical to the template from which it was derived. Determining The nucleotide sequence of clones makes it possible to analyze an entire DNA sequence. Upon completing a cloning experiment for a Gene of known sequence, one must verify that the cloned sequence is indeed identical to the published reference.
The original cDNA clone is synthesized using an mRNA template. This clone is then sequenced.
Deciphering the sequences of clones obtained from a physical genome map (Figure 40(b)) involves assembling a whole genome sequenced by the shotgun method. Whole-genome shotgun sequencing is carried out as illustrated in Figure 68.
Class="center">
Figure 68 — Assembly of a scaffold sequenced by the shotgun method
First, individual contigs are constructed based on the analysis of unique overlaps between clone sequence reads (Figure 68(a)). Next, the paired-end reads of the contig ends are sequenced (Figure 68(b)), which correctly orders and orientates the contigs, bridges the gaps between them, and merges them into larger units known as scaffolds (Figure 68(c)).
The adoption of fluorescent sequencing technology has significantly accelerated The rate of DNA sequence data accumulation (Figure 41). A greater number of sequencing reactions can now be performed in the same amount of time, and protocols have become much better suited for automation. When reactions take place in a fluorescent gel, laser-induced fluorescence is directly captured by a computer.
Gel Electrophoresis is typically performed in 36 parallel lanes. The output data is presented as a series of color-coded peaks, with a row of characters denoting the bases positioned directly above them (Figure 42). If the chromatogram-interpreting software cannot determine which base should be assigned to a specific position, a gap character "-" is inserted. In the final sequencing data file, such ambiguous positions are designated by the letter "N".
A clone sample is randomly selected from a library—for instance, 10,000 clones from a library containing 2 million. To initiate 10,000 sequencing reactions and subsequently run them on automated sequencers, a complex automated pipetting and reaction workflow is executed. The resulting data are uploaded to a database for further analysis.
An ideal outcome is a set of 10,000 sequences, each 200–400 NUCLEOTIDES in length, representing a portion of the sequence from each of the 10,000 clones.
In reality, some sequencing reactions fail entirely, some yield insufficient informative data, and others produce data of unacceptable quality. Sequences that successfully navigate this entire process are known as Expressed Sequence Tags (ESTs).
The resulting ESTs are deposited into GenBank, EMBL, and DDBJ, and remain publicly accessible through all these Databases. These same ESTs can also be found in the dbEST (Database of Expressed Sequence Tags) maintained by NCBI.
Last update: 11/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.