Fundamentals of Molecular Biology. Part 2: Molecular Genetic Mechanisms - A. N. Ogurtsov 2011

Genomics and Proteomics
Gene Identification in Genomic DNA Fragments

The complete genome of an Organism contains the information that determines The Structure of every protein synthesized by the organism's Cells.

For organisms such as Bacteria or Yeast, whose genomes contain very few introns and short intergenic regions, most protein-coding Gene sequences can be identified by computer searches for open reading frames (ORFs) of the required length within The Genome.

An Open Reading Frame is typically defined as a DNA segment longer than 100 codons that begins with a start codon and has the potential to be transcribed and subsequently translated into a polypeptide chain. Since the probability that a random DNA segment lacks a stop codon over 100 codons is extremely low, it is highly likely that most open reading frames encode Proteins.

ORF analysis correctly identifies over 90% of genes in yeast and bacteria.

Of course, this method (1) fails to detect certain very short genes, and (2) flags random long open reading frames that are not actually genes.

Both types of errors can be corrected through more detailed ANALYSIS OF GENE sequences and Genetic Testing of gene Functions.

For example, half of the genes in Saccharomyces cerevisiae discovered via ORF analysis were already known from mutant phenotype studies. The functions of some proteins encoded by the remaining putative genes identified through ORF analysis were established based on their similarity to previously studied proteins in other organisms.

Identifying genes in organisms with more complex genome structures requires more sophisticated algorithms than simply scanning for open reading frames. Figure 115 compares 50 kb genomic fragments from yeast, Drosophila, and humans.

Class="center">

Figure 115 - Gene arrangement in a 50 kb region of yeast, fruit fly, and human genomes

Genes shown above the solid lines in the figure are transcribed from left to right, while those below the lines are transcribed from right to left. Dark regions represent exons, and light gray regions represent introns.

Gene sequences whose functions remain undetermined are designated by special prefixes: Y (for yeast) is used for yeast, CG for Drosophila, and LOC for humans. The remaining genes shown in the figure encode proteins with already established functions.

Since most genes in higher eukaryotes, including humans and Drosophila, consist of numerous relatively short coding segments (exons) separated by noncoding regions (introns), a simple ORF scan is ineffective for gene finding. The best gene-prediction algorithms utilize all available data that can indicate the presence of a gene at a specific genomic Location.

The following types of information are typically used:

1) results of Hybridization with full-length cDNA;

2) comparison with 200–400 base pair cDNA segments known as Expressed Sequence Tags (ESTs);

3) alignment with existing exon and intron models;

4) Sequence Homology with genes from other organisms.

It should be noted that the computerization of database searches for gene information does not eliminate The Need for human participation in research. The computer merely proposes various alternative Answers to a given question, while a human researcher decides on the validity and plausibility of those suggestions.

The convergence of computer science and biology has given rise to new fields such as bioinformatics and computational biology. Computational biologists have identified approximately 35,000 genes in The Human Genome, although for about 10,000 of these putative genes, it remains unknown whether they actually encode proteins or RNA.

In particular, comparing the human and mouse genome sequences has proven to be a highly powerful method for identifying genes in the human genome. Humans and mice are genetically "close relatives" and share many homologous genes. At the same time, many nonfunctional DNA sequences, such as intergenic regions and introns, differ significantly (since they have not been "optimized" by evolution).

Therefore, it is highly probable that the Regions of the human and mouse genomes that exhibit sequence similarity correspond to functional coding DNA segments, such as exons.



Last update: 12/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.