Fundamentals of Bioinformatics - Ogurtsov A.N. 2013
Information Principles in Biotechnology
Genome Function and Organization
Gene Localization in the Genome
Computer programs for Genome Analysis identify open reading frames (ORFs). An ORF is a region of DNA that begins with a start codon (atg) (in some prokaryotes, start codons can also be gtg and ctg) and ends with a stop codon. An ORF represents a potential protein-coding region (see Section 11.1).
Two main approaches can be used to identify protein-coding regions in DNA molecules.
The first approach relies on identifying regions similar to known protein-coding regions from other organisms. These regions may encode Amino acid sequences resembling known Proteins or match Expressed Sequence Tags (ESTs). Since ESTs are derived from mRNA, they correspond to genes known to be actively expressed. In this case, sequencing just a few hundred initial NUCLEOTIDES of a cDNA is sufficient to obtain enough information for Gene identification (see Section 11.3).
Metaphorically speaking, identifying a gene via an EST is akin to indexing poems or songs by their opening lines.
The second approach is based on ab initio gene finding and identification, relying solely on sequence data.
Computer-based genome annotation is more accurate and complete for Bacteria than for eukaryotes. Bacterial genes are relatively straightforward to annotate because they are continuous—they lack introns typical of eukaryotes, and intergenic spacers are quite short.
Gene identification is significantly more complex in higher organisms. Exon identification is one of the major challenges, closely tied to another phenomenon: Alternative Splicing.
The ab initio gene prediction Procedure in Eukaryotic Genomes has several specific features.
The initial exon (5') begins at the METABOLISM/31.html">Transcription start site, preceded by a core promoter site such as a TATA box, which is typically located about 30 bp upstream of the gene. The initial exon generally contains no in-frame stop codons and terminates immediately before the gt splicing signal. Occasionally, the exon containing the start codon is preceded by an untranslated region (UTR) exon (see [7], Section 2.4).
Internal exons, much like initial ones, contain no in-frame stop codons. They begin immediately after the ag splice site and terminate right before the gt splice site. The preceding intron contains a so-called branch point and a polypyrimidine tract, which interact with the splicing machinery.
The terminal exon (3') begins immediately after the ag splice site and ends with a stop codon, followed by a polyadenylation site. Sometimes, the exon containing the stop codon is followed by an additional untranslated region (UTR) exon.
All coding sequences differ statistically from non-coding sequences in terms of codon usage bias (see Section 11.1). Empirical studies have shown that hexanucleotide statistics can be used to effectively distinguish between coding and non-coding regions.
Once a set of genes for a given Organism has been established, the specific codon usage statistics can be used to refine the parameters of recognition software, thereby improving the efficiency of identifying other genes within the organism's genome.
Last update: 11/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.