Protein Structure and Function. Application of Bioinformatics Methods - John Rigden 2014
Prediction of Protein Function Based on Theoretical Models
Introduction
Iwona A. Cymerman, Daniel J. Rigden, Janusz M. Bujnicki
Currently, Homology modeling is a well-established technique, whereas de novo modeling provides valuable insights for small Proteins with previously unseen folds. Advances in protein function Structure/77.html">Prediction Based on structure—driven largely by structural Genomics initiatives—have led to a broad array of Methods applicable to models of any origin. However, there are significant limitations in model accuracy and, consequently, in the performance of the algorithms that analyze these models to predict function. Nevertheless, this chapter demonstrates how protein function can be illuminated from various Perspectives using different modeling approaches, which often facilitates the planning and interpretation of experimental results. At the same time, important challenges remain in establishing a fruitful dialogue between modelers and experimental biologists, the resolution of which will expand the Practical Applications OF modeling results. Databases containing both protein models and indicators of their accuracy and reliability are likely to become increasingly important in the future.
Iwona A. Cymerman
International Institute of Molecular and Cell Biology,
Trojdena 4, 02-109 Warsaw, Poland
Daniel J. Rigden
School of Biological Sciences, University of Liverpool,
Liverpool L69 7ZB, UK
Janusz M. Bujnicki
Institute of Molecular Biology and Biotechnology,
Faculty of Biology, Adam Mickiewicz
University, Umultowska 89, 61-614 Poznan', Poland e-mail: iamb@genesilico.pl
The rapid progress in computing technology and the growth of data-sharing capabilities observed over the past decade have profoundly influenced the direction and methodology of biological research. This progress has enabled large-scale initiatives such as genome sequencing projects, microarray development, and structural genomics, which in turn have transformed the focus of individual studies. Thus, rather than merely identifying genes and proteins that determine an observed phenotype, scientists frequently concentrate on elucidating the Functions of the vast number of sequences deposited in databases. Clearly, shifting from descriptive approaches to predictive ones necessitates The Development of new methods. The most crucial piece of information about a particular Gene or protein is its associated function. The most common approach to function prediction is based on the observation that proteins with similar sequences often share similar functions. The ever-growing number of available sequences facilitates the alignment and grouping of similar sequences into families. If the function of one family member is known, it is generally assumed that the remaining sequences in the family inherit this function. This assumption raises the fundamental question of whether knowledge of a protein's 3D structure is required to predict its function, or if sequence information alone is sufficient. At first glance, the answer depends on the degree of similarity between the compared protein sequences. It is widely accepted that sequence identity exceeding 30% is a strong indicator that proteins will share a very similar structure, which can be predicted using homology-based methods (Chapter 3) to yield generally accurate models. Below this threshold, however, Protein Structure Prediction requires more sophisticated approaches (such as those described in Chapters 1 and 2) and is considerably less accurate. Because function depends on structure, one might expect sequence similarity to dictate functional similarity; however, this is not necessarily the case, as significant functional variations are observed even among proteins with highly similar sequences and structures. For instance, functional annotations derived from Gene Ontology (GO) data are conserved in only 80% of protein pairs, even when the proteins share 90-100% sequence identity; when identity drops below 30%, this annotation conservation falls below 50%. Some aspects of function are more conserved than others; for example, when enzyme function is classified According to the EC nomenclature, all four digits match with nearly 100% probability at sequence identities above 70%, whereas for sequences with less than 30% identity, the probability of retaining all four EC numbers drops below 50% (Tress et al. 2008).
The conservation of function is a more complex phenomenon than the conservation of structure. Functional redundancy (such as identical functions of two gene copies following duplication) is subject to accelerated evolutionary rates, which nevertheless depend on the biological utility of the protein function encoded by the duplicated gene (Jordan et al. 2004). Thus, duplication events that give rise to paralogous proteins with nearly identical sequences and structures typically lead either to the loss of one copy through inactivating Mutations (i.e., a return to the ancestral state) or to a divergence in the functions of one or both copies, thereby reducing functional redundancy.
On the other hand, orthologous sequences tend to retain identical functions, often despite significant sequence divergence. However, pairwise sequence comparisons do not allow us to distinguish between orthologs and paralogs, making them unsuitable for reliable function annotation. While several methods perform function prediction through evolutionary analysis and successfully differentiate paralogs from orthologs (e.g., FlowerPower (Krishnamurthy et al. 2007)), these approaches require large sets of sequences with relatively uniform divergence rates to reconstruct the putative duplication history. Furthermore, these methods encounter difficulties when sequences lose their shared function despite remaining orthologs.
The analysis of functional conservation can be greatly facilitated by examining a sequence not merely as a linear chain of amino acid residues, but within the context of its Spatial Structure. Because a protein's function is typically mediated by residues that are close in space rather than adjacent in the sequence, functional analysis can often be restricted to specific functional sites. Consequently, functional conservation generally requires only the preservation of the spatial arrangement of key catalytic residues, rather than complete sequence identity. This can be illustrated by a simple example: the loss of just a single residue from the catalytic center has a negligible effect on overall sequence similarity, yet typically results in the complete loss of a specific protein function (for instance, the protein may still bind its substrate but can no longer catalyze the chemical transformation that required the functional group of the missing residue). Therefore, comparing residues within functional sites and analyzing fuzzy properties such as various protein surface features (Chapters 7 and 8) are far better suited for comparative functional analysis than sequence comparisons alone. Naturally, such analyses require knowledge of the 3D STRUCTURE OF THE protein, and in this chapter, we discuss how computational modeling can contribute to obtaining such structures and, specifically, how models enhance our understanding of protein function.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.