Protein Structure and Function. Application of Bioinformatics Methods - John Rigden 2014
Examples: prediction of the function of structures obtained in structural genomics projects
Examples of large-scale protein function prediction
Despite the fact that structural Genomics projects have yielded a vast number of structures in recent years, there are surprisingly few Examples of using these structures to predict protein function. Here, we examine various attempts to improve the efficiency of protein function Structure/77.html">Prediction Based on structure, drawing on several structural genomics targets. A Brief Overview of these examples and Methods is provided in Table 11.1.
An Overview of 15 hypothetical Proteins with known spatial structures and the Procedures used to predict their Functions (Teichmann et al. 2001) offers some insight into the reliability of such predictions. These structures, combined with homologous sequence alignments, were used to identify surface clefts; if these clefts were formed by conserved residues, it typically indicated an Active Site. Utilizing information on all potential Cofactors present in the structures, along with available experimental data for the protein in question or related sequences, functional predictions were made with the highest degree of accuracy permitted by the existing data. It turned out that for 15 proteins, detailed predictions were made for a quarter of them, some functional information was obtained for another half, and no information at all could be retrieved for the remaining quarter.
Class="center">Table 11.1. Examples of function prediction for large-scale analysis and their sources.
This table summarizes the cases reviewed in the literature where attempts were made to improve the efficiency of protein function prediction based on structure, utilizing selected structural genomics targets. For each protein discussed in Section 11.3, A brief overview of the analysis is provided, accompanied by an indicator in the corresponding Column showing which structural analysis method proved to be the most informative.
|
Study |
Protein |
Description |
Key analytical method used to determine function |
|||
Fold/ structure (see Chapter 6) |
Surface/ cleft (see Chapter 7) |
Template (see Chapter 8) |
Bound |
|||
Kim et al. (2003) |
MJ0882 |
Identification of potential methyltransferase activity based on fold similarity, which was subsequently confirmed experimentally. |
X |
|||
MJ0577 |
Detection of bound ATP suggested ATP-hydrolyzing activity. |
X |
||||
TM841 |
Detection of bound palmitate indicated the potential for fatty acid binding. |
X |
||||
MJ0226 |
Discovery of a novel fold showing weak similarity to nucleotide-binding proteins and the HAM1 protein. |
X |
||||
MJ0285 |
Multimeric structure forming a hollow sphere with pores, raising questions about its MECHANISM OF ACTION. |
X |
||||
MPN625 |
Detection of two conserved cysteines residing within a potential active-site cleft, similar to that of the 2-Cys peroxiredoxin family. |
X |
||||
Watson et al. (2007) |
BioH |
Novel carboxylesterase. Searching against an enzymatic active-site template revealed a catalytic triad. |
X |
|||
IsdG |
Fold comparison and inverse template methods pointed to monooxygenase activity. |
X |
X |
|||
Adams et al. (2007) |
ChuS |
Three out of four conserved histidines were found to adjoin or face into one of two large clefts, demonstrating unusual heme coordination. |
X |
|||
YgiN |
Fold similarity to the ActVA-Orf6 protein belonging to the monooxygenase family. Discovery of two earlier reports providing additional arguments supporting this similarity. |
X |
||||
YjjX |
Fold similarity to nucleotide-binding proteins. Close examination of the active site revealed significant similarities, suggesting a novel ITP/XTP*-ase activity. |
X |
X |
|||
YhhW |
Fold similarity to a multifunctional family. Local surface similarities pointed to a deep charged pocket near the metal-binding site. Shows significant similarity to quercetin 2,3-dioxygenase. |
X |
X |
|||
z3393 |
General structural comparison and local molecular surface comparison suggest gentisate 1,2-dioxygenase activity. |
X |
X |
|||
* ITP - inosine triphosphate; XTP - xanthosine triphosphate)
In 2003, Kim et al. published an analysis of eight structures, some of which were obtained at the Berkeley Structural Genomics Center. This analysis provided functional and evolutionary insights into the structures, classifying them into five categories:
1. Remote homologs. Here, protein function is inferred based on structural similarity that is not apparent from sequence comparisons. As an example, the authors cite the MJ0882 protein, which was initially assigned as a methyltransferase based on fold similarity, a finding later confirmed experimentally (Huang et al. 2002).
2. Proteins with unexpectedly bound ligands. Here, function is deduced from a serendipitously discovered ligand or cofactor. The first example—the Analysis of the MJ0577 protein from Methanococcus jannaschii—involved the discovery of bound ATP within the structure, suggesting an ATP-hydrolyzing function. A more detailed examination of the ATP-binding pocket in MJ0577 revealed several motifs characteristic of nucleotide-binding proteins; however, their spatial arrangement differed from known analogs, preventing their detection by conventional motif-scanning methods. Subsequent experimental validation confirmed the ATP-hydrolyzing function, but only in the presence of Cell lysate, implying that the protein acts as a molecular switch requiring one or more partner proteins for its operation. The second example, the analysis of the TM841 protein from Thermotoga maritima, showed that it belongs to the large DegV family (Pfam Classification) and the COG1307 protein group, both of which have unknown functions. Determining The structure of TM841 revealed a bound palmitic acid molecule, thereby demonstrating the protein's ability to bind Fatty acids. Comparison with other members of the DegV family and the COG1307 group revealed high conservation within the carboxylic acid-binding region and lower conservation in the tail-binding region, suggesting that various members of these families may selectively bind fatty acids with different tail lengths.
3. "Twilight-zone proteins". In these cases, neither sequence nor structure alone allows for an unambiguous functional assignment. In the presented example, the STRUCTURE OF THE MJ0226 protein displayed a novel fold with weak similarity to nucleotide-binding proteins (Hwang et al. 1999). Biochemical analysis identified its Enzymatic Function as a novel nucleotide triphosphatase. Combined with its slight similarity to the HAM1 protein (Noskov et al. 1996), this led the authors to propose that the protein's role is to prevent Mutations by scavenging non-standard nucleotide triphosphates. This prediction was later validated by complementation experiments (Stepchenkova et al. 2005).
4. Novel molecular function with a known cellular function. Here, the overall function of the protein is known, but the biochemical details of its mechanism are revealed through structural studies. In the first cited example—the analysis of the MJ0285 protein from M. jannaschii—the protein was annotated as a small heat Shock protein induced by intracellular stress. The structure showed that 24 protein molecules form a hollow sphere with eight triangular pores and six square pores (Kim et al. 1998). Based on these findings, a question arose: do partially denatured proteins become trapped inside the sphere or attach to its exterior? The results of biochemical experiments provided strong evidence that partially denatured cellular proteins bind to the outer surface of the sphere during stress, protecting them from aggregation and inactivation. The second example analyzes the MPN625 protein, a member of the OsmC family. This family encompasses quite divergent sequences, yet Multiple Sequence Alignment reveals two conserved cysteines. The crystallographic structure of MPN625 showed that these two cysteines lie within a potential active-site cleft resembling that of the 2-Cys peroxiredoxin family, whose function is to detoxify reactive oxygen species (Schroder et al. 2000). Thus, comparing the two active sites alongside data on cellular function helped elucidate the molecular function of this protein family and explained differences in substrate Specificity.
5. Proteins of unknown function. Two examples are discussed here: the Aq1575 protein from Aquifex aeolicus and the MPN314 protein from Mycoplasma pneumoniae, both of which are hypothetical proteins belonging to Pfam domains of unknown function. Residue conservation data suggest potential active sites for both proteins, but database searches for Sequence Motifs and functions failed to provide any clues regarding their molecular function. Perhaps the most extensive analysis performed to date is that of Watson et al. (2007), who evaluated the performance of ProFunc—a structure-based function prediction server—using structures generated by the Midwest Center for Structural Genomics (MCSG). In that study, all 319 proteins produced by the MCSG During the first phase of the PSI-1 project (NIH/NIGMS Protein Structure Initiative) were classified into proteins with known functions, proteins with putative functions, and proteins of completely unknown function. Only proteins with known functions were evaluated further, as the goal was to assess how successfully ProFunc algorithms could predict those functions. Consequently, 93 proteins with known structures were submitted to the server, and the highest-scoring matches predicted by each method were retrieved and recorded. These results were then cross-referenced with the release dates of each structure to ensure that the calculated scores genuinely reflected how successful the server would have been if used for a priori function prediction. Finally, the top predictions from each method were compared against the known function of each protein to determine prediction accuracy.
The study demonstrated that among all the methods comprising the ProFunc server, Fold Recognition and inverse-template matching were the most successful, correctly predicting function in approximately 60% of cases. A detailed examination revealed that both methods often correctly predicted the function for the same protein, though there were instances where only one method succeeded. This stems from the inherent Nature of the approaches: while fold recognition aims to identify global similarity between compared proteins, inverse-template matching performs local comparisons. One of the primary limitations of this study was its inability to answer whether structure-based or sequence-based prediction is more accurate. However, this remains a general challenge for which no satisfactory solution has yet been described in the literature, largely due to the inherent difficulties in accurately reconstructing the exact historical state of sequence Databases, along with their derived motifs and profiles, at a specific past date.
In addition to a general comparison of methods, Watson et al. (2007) presented several specific examples of function prediction, some of which were experimentally tested. A successful example of structure-based function prediction was the BioH protein from Escherichia coli (Sanishvili et al. 2003). While this protein was known to be involved in biotin synthesis, its biochemical role had not been established. ProFunc structural analysis revealed a highly significant match (RMSD 0.28 Å) between the enzyme's active-site template and the Ser-His-Asp catalytic triad characteristic of lipases. DALI fold comparisons uncovered structural similarity between BioH and numerous proteins with diverse enzymatic functions, despite low sequence identity (15–25%). The closest matches included a bromoperoxidase (EC 1.11.1.10), an aminopeptidase (EC 3.4.11.5), two epoxide Hydrolases (EC 3.3.2.3), two haloalkane dehalogenases (EC 3.8.1.5), and a ligase (EC 4.2.1.39). Only painstaking manual analysis of these Enzymes and a thorough literature review could have revealed with equal clarity that all of them share a Ser-His-Asp catalytic triad in their active sites. In contrast, enzymatic active-site template searching detected these triads instantaneously. Experimental characterization of the BioH protein demonstrated that it functions as a novel carboxylesterase active on short acyl-chain substrates (Sanishvili et al. 2003).
Another example illustrating how functional insights can be gained through structural analysis is the hypothetical IsdG protein from Staphylococcus aureus. ProFunc sequence analysis suggested a wide range of putative functions, including monooxygenase, Cysteine peptidase, oxidoreductase, methyltransferase, epimerase, transporter, and potential RNA-binding activities. However, when the structure was submitted to the MSDfold/SSM service, the top-ranking folds corresponded to hypothetical proteins without functional annotations, while lower-ranking folds matched various Monooxygenases. No significant structural matches were found with Other Enzymes, DNA, or ligands, but inverse-template scanning yielded numerous hits. Again, most of these hits corresponded to proteins of unknown function, but the first meaningful match was a monooxygenase from Streptomyces coelicolor (PDB ID 1lq9). Thus, both fold comparison and inverse-template methods pointed toward monooxygenase activity. Subsequent experimental analysis characterized the protein as a heme-degrading enzyme structurally similar to monooxygenases (Wu et al. 2005). This is a prime example of how structural analysis provided decisive evidence favoring one of many equipotent predictions derived from sequence analysis.
A more recent study by Adams et al. (2007) discusses structure-based functional annotation of hypothetical proteins. The paper presents five examples where a combination of methods and biochemical assays successfully described protein function. The first example is ChuS, a protein with a novel fold for which Operon studies and Gene knockout experiments suggested a role in heme uptake and utilization. The structure was solved for the apo-form; because the fold was entirely novel, initial structure-based function prediction yielded no concrete results, but subsequent biochemical analysis pointed to a hemoxygenase function. Multiple sequence alignment of ChuS with its homologs revealed four conserved cysteines and histidines—specifically, three conserved histidines that structurally adjoined or pointed into one of two large clefts on opposite sides of the protein globule. This observation prompted further structural studies involving the co-crystallization of ChuS with heme and mutagenesis of the conserved histidines. As a result, it was discovered that heme coordination in ChuS occurs differently than in other heme-degrading enzymes, establishing ChuS as the first hemoxygenase identified in E. coli.
The second example examines the YgiN protein. Using the MSDfold (SSM) web server, its structure was compared against representative folds from the SCOP classification, revealing fold similarity to ActVA-Orf6, a monooxygenase from S. coelicolor.
Members of this monooxygenase family are involved in the synthesis of large polyketide compounds during antibiotic Biosynthesis in Gram-positive Bacteria. ActVA-Orf6 acts as a late-stage tailoring enzyme that modifies the antifungal compound dihydrokalafungin to impart specific activity (Sciara et al. 2003). Because The production of such a compound has not been described in E. coli, it was anticipated that the natural substrate for YgiN would differ. Initial attempts to biochemically characterize the enzyme proved fruitless until two earlier reports—actually pertaining to the YgiN protein—were uncovered in the literature. These additional experimental data, combined with previously utilized structural information on various substrates, enabled the authors to hypothesize a role for YgiN in menadione METABOLISM. Subsequent experimental work successfully crystallized the protein in its apo-form as well as in complex with menadione and flavin adenine dinucleotide (Adams and Jia 2006).
The third example is the YjjX protein, for which functional predictions could not be made based on its genomic context or sequence motifs. The structure of this protein reveals a fold similar to A number of nucleotide-binding proteins (including the aforementioned MJ0226 protein from Kim et al. 2003). Detailed examination of the active sites of YjjX and the identified structural matches revealed significant similarities in several conserved and semi-conserved residues. Further biochemical analysis justified classifying YjjX as a novel inosine triphosphatase/xanthosine triphosphatase that acts in E. coli as a housekeeping enzyme during oxidative stress to prevent the accumulation of non-canonical bases and their subsequent incorporation into Nucleic Acids.
The fourth example differs from the previous ones in that it involves functional annotation within a superfamilial context. In this case, the structure of the YhhW protein (previously annotated as belonging to the cupin superfamily) was determined. As expected, it revealed a protein scaffold analogous to that of known cupins; however, sequence analysis strongly suggested a close relationship to pyrins. The vast functional diversity within the cupin superfamily precludes making annotations based solely on global structural similarity, prompting the authors to examine local surface similarities. The discovery of a deep charged pocket near the metal-binding site in YhhW and one of its homologs, h-pyrin, suggested that this served as the active site. Closer inspection of this pocket revealed striking similarity to the pocket in quercetin 2,3-dioxygenase, a finding further reinforced by the successful docking of quercetin into the pyrin homologs. Quercetin 2,3-dioxygenase activity was subsequently confirmed by biochemical assays, marking the first enzymatic function assigned to a member of the pyrin family. This example also illustrates the difficulties encountered when working with large protein superfamilies, where global metrics like fold similarity are often insufficient for Function determination, necessitating more detailed local analysis.
The final example from Adams' work concerns another representative of the cupin superfamily—the product of the z3393 gene from E. coli. Sequence analysis indicates that it is more closely related to gentisate 1,2-dioxygenase than to other cupins. Comparisons of the overall structure and local molecular Surface Properties also support gentisate 1,2-dioxygenase activity. The authors hope that their determined structure of z3393 will facilitate future mechanistic studies of the enzyme and provide a deeper understanding of how the gentisate operon might be linked to pathogenic E. coli strains.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.