Protein Structure and Function. Application of Bioinformatics Methods - John Rigden 2014

Structural Motifs
Application of Molecular Docking in Function Annotation

Ultimately, a protein’s Ligand Specificity and catalytic capabilities depend on the spatial arrangement of atoms within its binding site or active center. While structural motif approaches aim to establish a direct link between spatial patterns and Functions, one can alternatively use Cell/13.html">Protein Structure to screen for likely ligands, and subsequently infer the molecular function of the protein based on the resulting specificity predictions. In computational docking, A large number of low-molecular-weight Organic compounds are individually fitted into the protein's binding site, and the complementarity of the resulting complex is evaluated. Hundreds of thousands of such binding poses may be tested for each compound. The molecules receiving the highest scores are considered the most probable protein ligands and serve as primary candidates for experimental testing.

While such computational Molecular recognition may seem straightforward, in practice it faces numerous challenges. One obstacle is the sheer volume of metabolites that may need to be screened as potential ligands. Another is the difficulty of achieving a fast yet sufficiently accurate assessment of structural complementarity to distinguish true binding ligands from similar, non-binding compounds. Accurate ranking may require accounting for conformational flexibility during the docking process, which further increases the computational burden. Traditionally, docking large compound libraries has been applied to discover lead compounds with potential pharmacological activity. In functional annotation, however, The Need for precise ranking is even greater than in lead discovery, as experimental screening for unknown catalytic activity is far more complex than binding assays, and the ultimate goal is to annotate thousands of Proteins in an automated or semi-automated manner. In contrast, lead discovery typically focuses on one or a few well-characterized targets.

Despite these difficulties, this approach has recently attracted considerable attention (Macchiarulo et al. 2004; Paul et al. 2004; Kalyanaraman et al. 2005; Tyagi and Pleiss 2006; Favia et al. 2008). In two published studies demonstrating function prediction via molecular docking (Hermann et al. 2007; Song et al. 2007), the protein of interest was first identified as a member of a functionally diverse enzyme superfamily, which helped narrow down the search space for potential substrates. However, each study employed different Methods to achieve accurate ranking of the docked molecules.

The first study examined a family of proteins that, based on sequence data, could be assigned to the enolase superfamily, although their precise function remained undetermined (Song et al. 2007). All members of this superfamily catalyze the abstraction of a proton from a carbon atom adjacent to a carboxyl group, yet their specific substrates and overall reactions vary widely (Babbitt and Gerlt 2000). Because no experimentally determined structures were available for this protein family, a Homology model was built for one representative using the most structurally similar protein with a known structure (35% sequence identity)—Alanine-glutamate epimerase. N-succinyl amino acid racemases also localize near an uncharacterized family in sequence similarity space; therefore, dipeptides and N-succinyl Amino Acids were included in the libraries for computational and experimental screening. Despite exhibiting overall greater sequence similarity to epimerases than to racemases, both parallel experimental analysis and in silico docking revealed that the protein catalyzes the racemization of N-succinyl-L-Arginine and -L-Lysine. Furthermore, although the homology model was constructed using the alanine-glutamate epimerase structure as a template, docking and score evaluation incorporating side-chain mobility and physics-based functions successfully confirmed the experimentally observed substrate preferences among N-succinyl amino acids. Subsequent crystallographic structures of the protein with both substrates bound showed remarkable agreement with the docking predictions. Notably, had the side chains been held rigid during the docking process, the identification of these substrates would have been far less successful (Song et al. 2007).

In the second study, a structure determined by a structural Genomics project was assigned to the amidohydrolase superfamily based on its fold and the presence of specific, highly conserved active-site residues (Hermann et al. 2007). However, these residues did not reveal which (if any) of the dozens of Hydrolysis reactions known for this superfamily could be catalyzed by this particular enzyme. For the docking Procedure, the authors selected only metabolites containing hydrolyzable functional groups, such as amide, ester, and phosphoester bonds. Because Enzymes catalyze reactions rather than merely binding substrates, the authors reasoned that using transition-state structures of these molecules in docking would yield superior results. Consequently, transition-state-mimicking structures were generated for each metabolite: amides and esters were converted into tetrahedral geometries, phosphoesters into trigonal bipyramidal geometries, and so forth. Preliminary validation studies demonstrated that utilizing such high-energy forms in docking improves the ranking of known ligands (Hermann et al. 2006).

During this predictive study (Hermann et al. 2007), it was observed that many top-scoring compounds contained an adenine moiety in which the exocyclic nitrogen had been converted into a tetrahedral geometry, mimicking the state it would adopt during deamination. Experimental assays confirmed that the enzyme catalyzes the deamination of three out of four tested adenine-containing metabolites, but not cytosine derivatives, even though the highest sequence similarity was observed with chlorohydrolases and cytosine deaminases. The crystallographic STRUCTURE OF THE enzyme in complex with one of the deamination products revealed the exact interactions predicted by the docking models. The validated functional annotation of this enzyme subsequently allowed for the annotation of several other sequences sharing characteristic active-site residues.

It is worth noting that restricting the reaction space to hydrolysis reactions helped make the problem tractable, as generating transition-state analogues significantly expanded the library of molecules targeted for docking. The scoring function accounted for steric, electrostatic, and desolvation contributions, while protein flexibility was not considered.

Both structural motif methods and molecular docking in functional annotation focus primarily on local structural features rather than global architecture or overall sequence similarity. While computationally more demanding than structural motif matching, docking offers the unique capability of extrapolating to functions unrelated to previously characterized structures, enabling the prediction of "novel" substrates for a given protein.



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.