Protein Structure and Function: Applications of Bioinformatics Methods - John Rigden 2014
Bioinformatics Methods for Studying the Structure and Function of Disordered Proteins
Prediction of IDP Functions
Maintenance of Disorder
To conclude the Structure/133.html">Discussion on predicting IDP Functions from sequences, It is worth mentioning an approach proposed in several studies, where IDPs are clustered based on their Amino Acid Composition features, which are subsequently correlated with functions. While the data obtained from such analyses are too limited for direct functional prediction, they can be useful in guiding future research directions.
METABOLISM/31.html">Transcription factor trans-activation domains are characterized by a pronounced propensity for disorder (Sigler 1996; Minezaki et al. 2006) and can also be classified according to their amino acid composition. Typically, transcription factors are classified based on such compositional features of their trans-activation domains as acidity and a high content of Pro and Gln (Triezenberg 1995). Although these distinctions lack robust statistical backing, the assignment of a transcription factor to a particular category can be justified by the fact that the functions of a given category remain insensitive to Amino Acid Substitutions as long as the defining property of the trans-activation domain is preserved (Hope et al. 1988). Conversely, Mutations that disrupt this characteristic property impair domain function (Gill and Ptashne 1987). Thus, certain properties observed at the level of amino acid composition are closely linked to biological function.
The existence of such general correlations has been investigated directly by clustering IDPs in The amino acid composition space (Vucetic et al. 2003). The premise of this study is that Disorder Prediction Methods trained on one group of Proteins often exhibit poor performance when applied to others, indicating significant differences in sequence properties among disordered proteins. For instance, Dunker and colleagues clustered 145 IDPs using various prediction methods and employed prediction accuracy as a separator for individual proteins. They found that disordered proteins could be divided into three similarly populated groups differing in their amino acid composition, designated as V, C, and S. Group C is enriched in His, Met, and Ala; group S has a lower content of His; and group V is characterized by an elevated proportion of the least flexible Amino Acids (Cys, Phe, Ile, Tyr). Each group exhibits a correlation with specific functions. For example, 9 out of 10 E. coli ribosomal proteins belong to group V. On the other hand, viral genome RNA-binding IDPs, as well as DNA-binding proteins, are virtually absent from group V. IDPs involved in Protein-Protein Interactions predominantly belong to groups V and S. Despite the known limitations of this analysis, it clearly demonstrates that the type of protein disorder encoded in the amino acid composition is related to biological function. More advanced analyses of this kind can be leveraged for functional prediction.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.