Fundamentals of Molecular Biology. Part 2: Molecular Genetic Mechanisms - A. N. Ogurtsov 2011

Genomics and Proteomics
Determination of the Functions of Novel Genes and Proteins

The Use of recombinant DNA Methods has enabled researchers to identify a vast number of both DNA fragments and complete genomes of humans and several model organisms. This colossal volume of information, which is growing rapidly every year, is stored primarily in two Databases:

1) GenBank at the National Institutes of Health, Bethesda, Maryland, USA http://www.ncbi.nih.gov/Genbank/index.html

2) EMBL Sequence Data Base at the European Molecular Biology Laboratory, Heidelberg, Germany http://www.embl-hcidelberg.de/

These databases are constantly updated with newly sequenced genomes and provide access to all researchers via the Internet. Researchers use these databases to conduct studies in areas such as:

1) determining the Functions of new genes and Proteins by comparison with already studied ones,

2) Comparative Genome Analysis,

3) identifying genes within genomic DNA fragments,

4) determining genome sizes,

5) developing DNA Microarrays,

6) Cluster Analysis of multiple Gene Expression.

Proteins with similar functions often contain comparable Amino acid sequences that correspond to functional domains within the Three-Dimensional Cell/13.html">Protein Structure. By comparing the Amino Acid Sequence encoded by a newly discovered cloned gene with The amino acid sequences of proteins with known functions, a researcher can predict the Functions of the new protein based on the identified similar regions.

Due to the degeneracy of METABOLISM/28.html">The Genetic Code (the same amino acid is encoded by multiple codons), related proteins invariably display much greater similarity in their amino acid sequences than in The nucleotide sequences of the genes that encode them.

One of the computer programs used for such comparisons is called BLAST (Basic Local Alignment search Tool). The BLAST algorithm breaks down the amino acid sequence of the new protein (referred to as the query sequence) into shorter segments and then searches the database for analogs among existing sequences. The comparison program assigns the highest score to identically matched sequences and lower scores to matches based on other parameters such as Hydrophobicity, polarity, amino acid charge, etc.

Once an analog for a given segment is found, the program compares neighboring regions in detail to extend the region of similarity. Upon completing the search, the program outputs a list of potential analogs for the query protein, ranking the list items by the E-value.

The E-value (expectation value) determines the degree of mismatch between two protein sequences. The smaller the E-value, the more similar the two sequences are. An E-value of less than 10-3 is generally considered to indicate that two proteins share a common ancestor.

To illustrate the effectiveness of this approach, let us consider the human NF1 gene. Mutations in this gene lead to the inherited disorder neurofibromatosis type 1, which involves The formation of multiple tumors in the Peripheral Nervous system, manifesting as bumps on the Skin (elephant man's disease syndrome). Following the isolation, sequencing, and cloning of the NF1 gene cDNA, the deduced amino acid sequence of the NF1 protein was compared with other protein sequences in GenBank. It turned out that one of the Regions of the NF1 protein is largely analogous to a fragment of the Yeast Ira protein (Figure 112).

Class="center">

Figure 112 - Comparison of regions of the human NF1 protein and the Ira protein from S. cerevisiae, shown in single-letter code

In Figure 112, identical and chemically similar amino acid pairs are indicated by gray boxes and dots, respectively. Previous studies had established that the Ira protein is a GTPase-accelerating protein (GAP) that modulates the GTPase activity of the monomeric Ras G-protein. The Ira and Ras proteins control Cell Division and differentiation in response to signals from neighboring Cells.

Experimental studies on the function of normal NF1 proteins, carried out using the Expression of cloned wild-type genes, demonstrated that these proteins indeed regulate The activity of Ras proteins, exactly as their Homology with Ira had suggested. Thus, in patients with neurofibromatosis, the expression of the mutant NF1 protein in peripheral nervous system cells leads to disruptions in cell division and the formation of tumors characteristic of this disease.

Even when it is not possible to detect significant similarity with already known proteins, the BLAST algorithm makes it possible to identify short, functionally important amino acid sequences—repeats or motifs—that occur in many proteins. To search for such motifs, The structure of the given protein is compared with a database of known motif structures. Some of the most frequently occurring motifs are listed in Table 6.



Last update: 12/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.