Protein Structure and Function. Application of Bioinformatics Methods - John Rigden 2014

Bioinformatics methods for studying the structure and function of disordered proteins
Disorder prediction
Machine learning algorithms

Probably the most advanced Methods for predicting structural disorder are machine learning (ML) algorithms—that is, predictive methods “trained” on specific sequences encoding ordered or disordered structures. Unlike simpler early approaches, machine learning algorithms integrate the consideration of non-trivial amino acid properties and hidden sequence features, which likely accounts for their superior performance. At the same time, accurate predictions are often driven by principles unknown to researchers; in other words, machine learning methods do not provide a deeper understanding of the processes underlying structural disorder.

A classic ML algorithm is PONDR (Predictor of Natural Disordered Regions), which relies on the analysis of local Amino Acid Composition, flexibility, and other sequence properties (Romero et al. 1998). Developed in several variants, it enables the Structure/77.html">Prediction of disorder in terminal protein regions (Li et al. 1999)—regions highly likely to represent characteristic motifs (VL-XT (Iakoucheva et al. 2002))—as well as combinations of short and long disordered regions (VSL2 (Peng et al. 2006)). Because short disordered regions are context-dependent (i.e., their lack of a defined structure is dictated by the structural environment) whereas the disorder of long regions is independent, this combined approach forms The basis of one of the most efficient algorithms for predicting structural disorder.

Another approach, differing in its computational framework, involves the application of support vector machines (SVMs) and is represented by the DISOPRED2 algorithm (Ward et al. 2004). This algorithm searches the feature space for a hyperplane that separates ordered Proteins from disordered ones. The hyperplane can be either linear or non-linear. It accounts for imbalanced Class frequencies of both ordered data (e.g., proteins in the PDB) and disordered data (e.g., proteins in DisProt (Sickmeier et al. 2007)). Sequence profiles generated using PSI-BLAST are also utilized as input data.



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.