Protein Structure and Function: Application of Bioinformatics Methods - John Rigden 2014

Prediction of Protein Function Based on Theoretical Models
Practical Applications
Protein Complexes

For systems biology—which aims to integrate large-scale datasets into a meaningful whole—to ultimately succeed, a thorough understanding of the complex protein-protein interaction networks within The Cell is essential. Consequently, systems biology stands to gain significantly by incorporating comparative modeling predictions into its arsenal of experimental and computational Methods for predicting Protein-Structure/156.html">Protein Interactions (Aloy and Russell 2006). The principle is straightforward: given a known structure of a complex between protein A and protein X, analyzing a potential complex between protein B and protein Y (where B is homologous to A and Y is homologous to X) suggests that this interaction will occur in vivo (Aloy and Russell 2002). Early methods in this field evaluated interface viability by adapting pairwise interaction potentials from threading techniques and analyzing known contacts in the A-X structure following sequence alignments of B with A and Y with X. Today, such analysis can be performed using a variety of web servers, including InterPreTS (Aloy and Russell 2003) and MULTIPROSPECTOR (Lu et al. 2002). In subsequent work, protein complex modeling was performed explicitly, once again utilizing interaction potentials to distinguish between true and false interactions (Davis et al. 2006). A notable, large-scale application of this prediction approach (Davis et al. 2007) targeted human Proteins and those from the genomes of ten pathogens responsible for neglected diseases. The pathogen and host genomes were first scanned for proteins homologous to those with known interactions. When such structural interaction data were unavailable, the process continued using simple sequence similarity scoring Functions. However, this approach yielded only a limited number of predictions due to the strict reliability criteria applied.

More interesting and significant was the explicit modeling of potentially interacting partners based on a protein complex template. The resulting complex models were evaluated using statistical potentials, and those receiving positive scores were advanced to a sophisticated secondary filter. This filter leveraged known information regarding the tissue and intracellular localization and function of the interacting proteins to rule out interactions that could not occur in vivo. Consequently, only host proteins expressed in the Skin, Lymph Nodes, or Lungs were selected as candidate interactors for Mycobacterium leprae proteins. Pathogenic proteins also had to meet specific biological criteria; for instance, M. leprae proteins required appropriate GO annotations (such as pathogenicity markers) or annotations indicating they were extracellular or surface-localized. Following this filtering step, the number of predicted interactions ranged from 0 to 1,501, depending on the pathogen.

Although the authors had a relatively small number of known interactions available to test their methodology, they successfully predicted 4 out of the 33 interactions documented to date. In the remaining cases, no suitable template was available for modeling the interactions, suggesting that template scarcity is the primary reason for the small number of described interactions (Davis et al. 2007). Interestingly, one prediction was experimentally validated: the method successfully predicted the interaction between falcipain-2 and cystatin (PDB code lyvb) based on the earlier structure of cathepsin H in complex with stefin A (PDB code 1nb3) (Fig. 12.2). These two Enzymes share approximately 24% sequence identity, and their inhibitors are only 11% identical. Therefore, achieving a successful prediction despite such low sequence similarity and notable structural differences (Fig. 12.2) powerfully highlights the capability of the method.



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.