Fundamentals of Molecular Biology. Part 2: Molecular Genetic Mechanisms - A. N. Ogurtsov 2011

Genomics and Proteomics
Cluster Analysis of Multiple Gene Expression

Definitive Conclusions on whether or not common regulators exist for genes exhibiting synchronous changes in expression—and, consequently, whether or not these genes are functionally related—cannot be drawn solely on The basis of DNA microarray experiments.

For instance, many of the detected differences in Yeast Gene Expression upon switching from glucose to ethanol may be an indirect consequence of various physiological shifts occurring in The Cell as it transitions from one environment to another. In other words, Changes in the expression of genes showing a synchronous response in a DNA microarray experiment can be triggered by entirely different causes; hence, these genes may serve completely distinct biological Functions.

This issue can be resolved by combining data from a series of different expression experiments to identify genes that are coregulated under various conditions and synchronously over time.

For example, the simultaneous Study of the expression of 8,600 genes over 24 hours following the exposure of human fibroblasts to an optimal growth serum—yielding over 10,000 fluorescence kinetics measurements of individual DNA microarray Cells—made it possible to determine the relationships among various genes, Structure the obtained data (using computational Processing, of course), and group genes displaying similar temporal expression patterns into "clusters" (Figure 120).

Class="center">

Figure 120 - Cluster analysis of data from multiple gene expression experiments using DNA Microarrays.

Notably, such cluster analysis grouped together sets of genes encoding Proteins involved in general cellular processes, such as Cholesterol Biosynthesis or the Cell Cycle.

In Figure 120, capital English letters indicate gene clusters encoding proteins involved in specific cellular processes: A - cholesterol biosynthesis, B - the cell cycle, C - immediate response, D - signaling functions and angiogenesis (Blood vessel development), and E - wound healing and tissue remodeling.

Because genes exhibiting identical or similar regulatory patterns typically encode functionally related proteins, cluster analysis of combined data from a series of expression microarray experiments serves as an additional tool for determining the functions of novel, recently identified genes.

This approach allows for the integration of any number of diverse experiments. Furthermore, each new experiment enables the iterative refinement of the cluster analysis, thereby detailing the cluster architecture of The Genome.

In Conclusion, a clear manifestation of the recent surge in Genomics and Proteomics research is The Emergence of new specialized scientific journals dedicated to the comparative analysis of genomes and proteins. For instance, beginning in 2006, Elsevier launched the journal "Comparative Biochemistry and Physiology, Part D, Genomics and Proteomics," entirely devoted to research of this kind.

CONCLUSIONS

The functions of a protein that has not yet been isolated in pure form can be predicted based on the similarity of its Amino Acid Sequence to those of proteins whose functions are already known.

The BLAST computer algorithm performs a rapid search against Databases of decrypted protein sequences, identifying regions of similarity between the query (novel) protein and previously studied proteins.

Proteins sharing common functional motifs may not be identified by standard BLAST searches. To detect these short stretches of the protein chain, protein motif (repeat) databases are employed.

Protein Families unite proteins that trace their "Lineage" back to a common ancestral protein. The genes encoding these proteins, which constitute the corresponding gene family, originated from an ancestral gene via duplication followed by divergence during speciation.

Related genes and the proteins they encode that arose through Gene Duplication are termed paralogs, whereas those that emerged via speciation are called orthologs. Orthologous proteins typically share similar functions.

Open reading frames (ORFs) are regions of genomic DNA containing more than 100 codons located between the start and stop codons.

Computer-based searching for open reading frames in complete bacterial and yeast genomes successfully identifies the majority of protein-coding genes. However, due to the more complex gene structure Organization in humans and other higher eukaryotes, additional data must be integrated to identify probable genes within such genomic sequences.

Comprehensive genome analyses of several different organisms have demonstrated that biological complexity is not directly proportional to the number of protein-coding genes.

DNA microarray analysis simultaneously detects the relative expression levels of thousands of genes across different cell types or within a single cell under varying conditions.

Cluster analysis of results from multiple DNA microarray experiments can identify genes that are coregulated under different conditions. Such similarly regulated genes typically encode proteins with biologically similar (or related) functions.

QUESTIONS FOR SELF-Control

1. What gene sequences are referred to as homologous?

2. What gene sequences are termed paralogous, and how do they differ from orthologous gene sequences?

3. What gene sequences are called orthologous, and how do they differ from paralogous gene sequences?

4. WHAT IS A Phylogenetic Tree (cladogram)?

5. What is an Open Reading Frame?

6. What are the main reasons why Genome Size does not correlate with an Organism's biological complexity?

7. How is a DNA microarray structured?

8. What is determined by the cluster analysis method of multiple gene expression?



Last update: 12/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.