Protein Structure and Function: Application of Bioinformatics Methods - John Rigden 2014

Examples: prediction of the function of structures obtained through structural genomics projects
Conclusion

Structural Genomics consortia have solved an enormous number of structures that have greatly enriched our treasury of protein 3D structures. However, due to the high-throughput nature and rapid publication timelines mandated by these projects, a large fraction of these structures lack functional annotation entirely or possess very little of it. Predicting protein function from sequence and Structure has long been something of a Holy Grail for bioinformaticians, inspiring A wide variety of computational Methods over the years. Each approach has its own pros and cons, and numerous research case studies vividly demonstrate how selecting the right method can yield vital biological insights where others fail. Currently, no single method achieves 100% success, so continued development of novel tools for structural analysis and function prediction remains essential for advancing our understanding.

Several attempts to perform large-scale benchmarking of structure-based function prediction have met with mixed success, yet several critical challenges in this field still await resolution. One of the primary hurdles in developing structure-based function prediction methods is the lack of a reliable gold-standard dataset for validation. Consequently, individual developer groups are forced to curate their own test sets, making rigorous comparisons between different methods exceptionally difficult. Another major obstacle to overcome is how to fairly compare structure-based prediction methods with sequence-based ones. Although not always impossible, stripping away information from sequence Databases and profile-based methods derived from them is a daunting task. This is feasible only for retrospective analyses (such as the study by Watson et al., 2007); however, the problem would be solved if all protein function predictions made at the time a structure was first released were systematically archived and later analyzed. Such efforts are already underway using the ProFunc server, where prediction results for structures generated by the MCSG are stored for future analysis and benchmarking. These issues are far from trivial and must be addressed if newly generated datasets are to serve effectively for method evaluation and comparison in the future. To maximize the probability of identifying the correct function, current efforts focus on developing hybrid methods and leveraging all available information. The integration of existing bioinformatics pipelines has emerged as a direct response to the massive backlog of Proteins with little to no functional annotation. One of the toughest Challenges in Protein function prediction lies in synthesizing data across all biological disciplines; as the volume of available knowledge continues to expand, the growing trend toward community annotation projects will likely play a pivotal role in delivering deeper biological insights.

Acknowledgements. The authors would like to thank Roman Laskowski and Vicky Schneider for their valuable comments on this chapter.



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.