Protein Structure and Function: Application of Bioinformatics Methods - John Rigden 2014

Bioinformatics methods for studying the structure and function of disordered proteins
Prediction of IDP functions
Prediction of short recognition motifs in IDPs

A completely different yet important approach is to predict the presence of short Structure/155.html">Sequence Motifs within IDPs/IDRs, which can subsequently be directly linked to specific Functions, such as post-translational modifications or binding to closely interacting partner molecules. As noted above, the functions of IDPs are frequently associated with the presence of short linear motifs involved in Protein-Protein Interactions. Because the Amount of Information contained within these short motifs is limited, specialized Methods have been developed to recognize such protein regions, two of which are described below.

One of these methods—DILIMOT (Discovery of Linear MOTifs) (Neduva and Russell 2006)—capitalizes on the fact that statistical significance can be markedly enhanced by using a set of sequences that share a common functional property for prediction (such as an interaction partner molecule or subcellular localization), driven by the presence of a short motif that is highly likely to be represented in each sequence of the set. Regions of the input sequences that have a low probability of containing instances of linear motifs (such as globular domains, signal Peptides, transmembrane regions, and coiled-coil domains) are excluded from the analysis. Motif searches are then performed among the remaining sequences using a pattern-matching algorithm. The discovered motifs are ranked according to their level of representation redundancy and their level of conservation among homologs in related species. The performance of the method is further improved by comparing Proteins from different biological species, as well as by sequence randomization. Initial application of the method to high-throughput interaction datasets from Yeast, fly, worm, and human sequences resulted in the rediscovery of numerous previously known linear motif Examples, as well as the identification of several novel motifs. Predictions for two putative novel motifs were subsequently confirmed by direct binding assays: the DxxDxxxD motif binds protein phosphatase 2 with a Kd= 22μM; the VxxxRxYS motif binds traslin with a Kd=43μM (Neduva and Russell 2005).

Conceptually similar to DILIMOT is the SlimDisc (Short Linear Motif Discovery) method (Davey et al. 2006). It is based on the premise that Evidence for the presence of a characteristic motif in a protein becomes increasingly compelling the more frequently that motif appears across distinct, unrelated proteins that evolve through convergence. Detecting such motifs is often hindered by sequence similarity in related proteins arising from a common ancestry. To account for this, the search for similar motifs is conducted within a group of proteins sharing a common characteristic property, yet exhibiting low or entirely absent primary sequence similarity. In this context, the shared characteristic property may be the biological function of the proteins, their subcellular localization, or a common partner molecule with which the proteins interact. Motifs discovered using basic pattern-recognition algorithms, such as TEIRESIAS, are considered more significant if they are found in utterly unrelated sequences, and less significant if they evidently share a common evolutionary ancestor. Validation of the SlimDisc method on a benchmark dataset of proteins containing linear motifs (Neduva and Russell 2005) demonstrated a significant improvement in performance.



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.