Fundamentals of Bioinformatics - Ogurtsov A.N. 2013

Information Principles in Biotechnology
Biological Sequence Sequencing and Gene Expression
Expressed Sequence Tags

An Expressed Sequence Tag (EST) is a short sequenced segment of a clone randomly selected from a cDNA library, used to identify genes expressed in a specific tissue.

We do not always have complete DNA sequence decodings at our disposal; currently, accumulated DNA data mainly consist of individual sequence segments, the majority of which are represented by expressed sequence tags (ESTs).

When analyzing EST sequences, the following points must be taken into consideration:

1) the alphabet of EST sequences consists of five characters (a, c, g, t, u);

2) the sequence may contain phantom indels (insertion/deletion), leading to frameshifts;

3) it is highly probable that the EST sequence will turn out to be a subsequence of some sequence already present in Databases;

4) an EST sequence may not represent a coding sequence of any Gene at all.

The EST sequencing principle is illustrated in Figure 69.

A cDNA library is constructed from Cells of the tissue or Cell line of interest. To do this, mRNA is isolated from the tissue or cell culture. The mRNA is then reverse-transcribed into cDNA—typically using an oligo-(dT) primer, so that one end of the cDNA insert is transcribed from the poly-A tail at the mRNA terminus. The other end of the cDNA generally corresponds to a region of the coding sequence or, if the coding sequence is short, to the 5'-EST region. Finally, the resulting cDNA is cloned using a vector.

Class="center">

Figure 69 - Schematic diagram of Expressed Sequence Tag (EST) construction: 1 - mRNA isolation and reverse METABOLISM/31.html">Transcription into cDNA; 2 - insertion of cDNA into a vector for propagation and cDNA library creation;

3 - Selection of individual clones; 4 - sequencing of the 5'- and 3'-ends of the inserted cDNA; 5 - submission of the EST to the dbEST database

Individual clones are selected from the library, and a single sequence is synthesized from each end of the cDNA insert. This method is known as the RACE method (Rapid Amplification of cDNA Ends).

Thus, each clone is typically represented by a 5'-EST and a 3'-EST. Because EST sequences are short, they usually represent only gene fragments rather than full coding sequences. A typical EST tag ranges from 200 to 500 NUCLEOTIDES in length.

As a rule, The process of EST synthesis is highly automated and typically involves The Use of a fluorescent laser system for reading gel films. Subsequent Analysis of the decoded sequences is performed using computational Methods.

To determine whether such an EST sequence represents a novel gene, a search is performed against a DNA database. If the alignment result shows significant similarity to a sequence in the database, the standard match Classification Procedure will determine whether a genuinely new gene has been discovered.

If, however, the search result does not show significant similarity, we lack sufficient grounds to assume that a new gene has been detected. It may also turn out that the given EST sequence represents a non-coding sequence of a known gene that simply has not yet been deposited in the database.

In many mRNAs (especially in humans), long untranslated regions (UTRs) are located at the 5'- and 3'-ends of the coding sequence (Figure 43). It is highly probable that the given EST sequence was entirely transcribed from one of these non-coding UTR regions. If we are fortunate, some fragments of the untranslated (non-coding) sequence will already be present in the database. If this is the case, the search will yield a direct match, since UTRs are highly conserved and quite specific to coding genes.

In an unfavorable outcome, no match will be found, which points to one of the following two possibilities:

1) the given EST sequence represents a coding sequence for which there is no similar sequence in the database;

2) this EST sequence represents a non-coding sequence that has not yet been deposited in the database.

When interpreting EST sequence analysis, It is important to clearly distinguish between these two situations.

Figure 70 shows a diagram illustrating the correspondence between DNA, cDNA, and ESTs.

Figure 70 - Alignment of fully sequenced cDNA sequences and EST tags with genomic DNA

The dots between the cDNA segments or EST sequences indicate regions in the genomic DNA that do not align with the cDNA or EST sequences; these are intron regions.

The numbers above the cDNA line indicate the coordinates (in nucleotides) of the cDNA sequence, where nucleotide #1 is the closest to the 5' end of the cDNA, and nucleotide #816 is the closest to the 3' end of the cDNA.

Each EST tag represents only a short sequence read from either the 5' or 3' end of the corresponding cDNA.

Thus, EST tags establish the boundaries of transcription units, but provide no information about the Internal Structure of the transcripts unless the sequences of these ESTs span introns (as in the case of the 3'-EST shown in Figure 70).



Last update: 11/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.