Protein Structure and Function: Applications of Bioinformatics Methods - John Rigden 2014
Protein Function Prediction Based on Theoretical Models
Practical Applications
Ab Initio Function Prediction
Until recently, before The Development of more powerful yet resource-intensive algorithms (Bradley et al. 2005), a reasonable goal for template-free (ab initio or de novo) modeling was simply to obtain the correct fold rather than high-accuracy predictions (see Chapter 1). This limited the range of applicable protein function prediction Methods, meaning that most predictions described in the literature rely primarily on protein fold prediction and its correlation with function (as discussed in Chapter 6).
In one of the earliest large-scale Applications of the ROSETTA server, Bonneau and colleagues (Bonneau et al., 2002) generated models for 510 Protein Families from the Pfam Classification with an average length of under 150 residues. These were protein structures unknown at the time, though for some of them, the function was known or hypothesized. In several cases, speculative predictions could be supported by modeling results. For instance, it was hypothesized that PF01938, a TRAM protein domain, might bind Nucleic Acids. This assumption was based largely on the similarity of its ab initio model to structures in SCOP database superfamilies containing various nucleic acid-binding Proteins. We can now judge the accuracy of the model from two unpublished results of the structural Genomics project (lyez and lyvc). An example of function prediction for a completely uncharacterized protein is domain of unknown function 37 (PF01809). Its model matched The Structure of NK-lysin, a hemolytic protein expressed in natural killer Cells. Although the STRUCTURE OF THE PF01809 protein remains unknown, the Pfam database at the time of writing reports unpublished evidence that this protein from Aeromonas hydrophila indeed exhibits hemolytic activity.
Notably, a single ab initio model does not need to match a known fold precisely to provide clues about function; rather, insights are sometimes prompted by the entire broad Class of structures to which a particular model belongs. An example is the model built for the mucin-binding domain (Bumbaca et al. 2007). The preferred model contained a beta-sandwich fold of a type strongly correlated with carbohydrate binding—at the time of publication, half of the carbohydrate-binding domain families with known structures possessed this fold. This was more consistent with the domain binding to the carbohydrate moiety of its target, a heavily glycosylated mucin, than to its protein part. Furthermore, this ab initio model featured three exposed aromatic residues of the type considered characteristic of carbohydrate binding (Quiocho and Vyas 1999).

Fig. 12.2. Prediction of Protein-Protein Interactions based on Cell/13.html">Protein Structure modeling. The approach based on comparative modeling of Protein Complexes (Davis et al. 2008) made it possible, using the structure of cathepsin A in complex with stefin A (PDB code 1nb3; a), to propose a potential interaction between falcipain-2 and cystatin, which was subsequently confirmed crystallographically (PDB code 1yvb; b). In both panels, Enzymes are shown at the top, inhibitors at the bottom.
A recent example demonstrated how a function predicted from an ab initio model fold can be confirmed by other methods (Rigden and Galperin 2008). The SpoVS protein is known to be essential for sporulation in spore-forming Bacteria, yet it is actually far more widely distributed. The phenotypic description of organisms with mutated SpoVS yields little insight into its molecular role. However, the best models generated by the ROSETTA and I-TASSER servers align well with the fold of the Alba protein found in archaeal Chromatin (Fig. 12.3a). Such a fold is closely associated with nucleic acid binding in various contexts and, furthermore, electrostatic potential mapping on the model surfaces revealed a distinct positively charged region characteristic of nucleic acid-binding proteins (Fig. 12.3b; see Chapter 7). Summarizing the results of these methods, it can be suggested that the SpoVS protein is a novel METABOLISM/31.html">Transcription factor involved in controlling the intricate Gene Expression program that occurs during sporulation (Rigden and Galperin 2008).

Fig. 12.3. (For the color version of this figure, see the color insert.) Analysis of SpoVS ab initio models suggested that it possesses nucleic acid-binding function (Rigden and Galperin 2008). a) Models generated by both ROSETTA (shown in gray) and I-TASSER (shown in black) closely resemble the structure of the Alba protein found in archaeal chromatin (PDB code 1nfj; colored by spectrum from blue N-terminus to red C-terminus). b) Electrostatic potential of a putative SpoVS dimer modeled using ROSETTA (blue areas indicate positive potential, red indicates negative potential).

Fig. 12.4. (For the color version of this figure, see the color insert.) Validated structure prediction performed by Malmström and colleagues (2007). The model of the TRS20/YBR254C protein (a) matched the SCOP SNARE superfamily; this match was later confirmed by obtaining the experimental structure (PDB code 1h3q) of a similar protein (b). Structurally similar elements are highlighted in color, while the rest is shown in gray.
A recent large-scale application of ab initio modeling within an approach that also incorporated PSI-BLAST and threading-based structure prediction methods was aimed at analyzing the Yeast genome (Malmstrom et al. 2007). The authors applied a novel strategy utilizing known functional information to facilitate the Selection of correct matches between potential ab initio model structures and SCOP superfamilies. To this end, In addition to structural comparisons, Gene Ontology (GO) term matches between the target protein and the candidate superfamily proteins were evaluated. These complementary sources of information were integrated using Bayesian statistics. The figure (Fig. 12.4) shows an example of predicting the assignment of the TRS20/YBR254C protein to the SCOP SNARE superfamily, which was later confirmed by experimental structure determination. The match between the model and the crystallographic structure is partial and moderate (Fig. 12.4), illustrating the value of GO database information regarding the target protein and the superfamilies of potentially matching structures. In this case, the TRS20/YBR254C target was a subunit of the transport protein particle (TRAPP) complex involved in vesicle tethering and fusion. Its match with the SCOP SNARE superfamily structure was thus reliably corroborated, since Vesicular Transport is one of the primary Functions of Proteins in this superfamily.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.