Protein Structure and Function: Application of Bioinformatics Methods - John Rigden 2014

Spatial Motifs
Overview of Methods
Interpretation of Results

A wide variety of web servers are available for comparing a Cell/13.html">Protein Structure of interest with Databases of Structural motifs (Table 8.1). What Conclusions can be drawn about a protein's function if a specific motif is identified within its structure? Determining whether a match is statistically significant, and if so, what biological implications it carries, requires addressing several key questions.

The putative function largely depends on how the motif was derived. If the motif consists of a set of residues experimentally proven to be essential for a specific function, identifying it in a query protein suggests that the protein may share this function. Similarly, discovering a motif extracted from the structures of multiple Proteins with a similar function (positive Examples) implies that the query protein likely performs the same function, especially if the motif was designed to exclude negative examples. However, the exact function assigned to a structure depends heavily on the criteria used to define the positive examples. For instance, if all positive examples bind adenosine yet catalyze diverse reactions, finding their common motif in a query structure implies adenosine binding, but not necessarily adenosine deaminase activity. If the set of positive examples comprises members of a single SCOP superfamily, identifying the motif in the query protein indicates that it belongs to that same superfamily, though more specific functional predictions cannot be made.

Ideally, the set of positive examples used for motif discovery should be as diverse as possible while sharing a common property, whereas negative examples should closely resemble the positive ones but lack that specific property. In practice, training sets of positive and negative examples are rarely ideal, and some or all derived structural motifs may ultimately reflect common evolutionary origins or coincidental similarities rather than a shared function.

Interpreting motifs extracted from single structures is inherently subjective. As with sequence analysis, the query structure is annotated by analogy to the protein that served as the source of the best-matching motif. While identifying a motif corresponding to a binding site is traditionally interpreted as a prediction of binding Specificity, it is also frequently used (with varying degrees of accuracy) to infer other annotations, such as catalytic activity or membership in a specific family or superfamily. It should be noted that mere spatial proximity to a Ligand does not guarantee that a residue is crucial for binding or catalysis, as it may have arisen there through mutation. Apart from utilizing residues surrounding a ligand, another way to extract a motif from a single structure is to rely on annotations provided in the PDB file (SITE records). Although these records are intended to list residues forming binding sites, the Biological Significance of such motifs remains ambiguous due to a lack of strict criteria for which residues should be included. Furthermore, many PDB files of ligand-binding structures contain no SITE records at all.

Class="center">Table 8.1. Web servers for searching and comparing structural motifs

Name and URL

Server Function

Structural Motif Database11

Catalytic Site Atlas (CSA)

www.ebi.ac.uk/thomton-srv/Databases/cgi-bin/CSS/makeEbiHtml.cgi?file=form.html

Compares the query structure against motif databases using the JESS program and assesses significance

Motifs comprising alpha-carbon, beta-carbon, and functional atoms for 147 well-studied enzyme families; also available for download

fiinClust

www.pdbfim.uniroma2.it/Funclust

Identifies motifs common to three or more structures using Query3D, with a maximum upload limit of 20 structures

No databases

GASPSdb

www.gaspsdb.rbvi.ucsf.edu

Uses the RIGOR6 program to compare the query structure with structural motifs present in SCOP and GO, and evaluates significance

Motifs based on alpha-carbons and side-chain geometric centers: 4,385 motifs representing 272 GO molecular Functions, 3,599 motifs representing 186 SCOP superfamilies and 137 families, and 4,581 motifs representing 376 groups where proteins share both SCOP superfamily Classification and GO molecular function

PAR-3D

www.sunserver.cdfd.org.in:8080/protease/PAR_3D

Compares the query structure against motifs represented as inter-point distances and angular ranges.

Alpha- and beta-carbon motifs for 6 protease classes and 10 glycolytic Enzymes; alpha- and beta-carbon plus side-chain pseudoatom motifs for metal-binding sites consisting of three or four residues.

pdbFun

www.pdbfun.uniroma2.it

Compares a specific query sample with target residue sets using the Query3D program

Over 12 million individual residues, with subsets selectable via boolean combinations of descriptors

PDBSiteScan

www.mgs.bionet.nsc.ru/cgi-bin/mgs/fastprot/pdbsitescan.pl?stage=0

Compares the query structure against all or a subset of motifs in the PDBSite database

36,273 backbone-atom motifs derived from SITE annotations of individual PDB structures or their interfaces with other proteins, RNA, or DNA

PINTS

www.russell.embl.de/pints

Weekly updated results:

www.russell.embl.de/pints-weekly

Compares a query structure to a database motif, a user-defined motif to the Protein Data Bank, or two proteins against each other; assesses significance

Ligand-binding and SITE-annotated motifs consisting of polar residue side-chain points

ProFunc

www.ebi.ac.uk/thomton-srv/databases/profunc

Performs multiple searches including motif searches via JESS: searches the entire query structure against motif databases, or query fragments against full-chain databases; evaluates significance

Active-site motifs from the CSA database, 13,057 ligand-binding and 1,200 DNA-binding motifs from single structures, and 11,750 full chains; residues are represented by side-chain atoms, with smaller residues also including one or more backbone atoms.

ProKnow

www.Proknow.mbi.ucla.edu

Performs multiple searches including motif searching via RIGOR, yielding GO-based functional annotations

10,230 motifs from GO-annotated structures, or 7,819 motifs when electronic annotations are excluded

Protemot

www.protemot.csbb.ntu.edu.tw

Compares the query structure against all ligand-binding motifs or specific subclasses (e.g., enzymes only)

2,362 binding-site motifs based on alpha-carbon atoms, 1,051 of which belong to enzymes

SuMo

www.sumo-pbi1.ibcp.fr

Commercial site:

www.medit.fr/products-page2-med-sumo.html

Compares a query structure, chain, or ligand-binding site against databases of structures or ligand-binding sites alone

34,210 ligand-binding sites In addition to full structures described as spatial patterns of functional groups

a Sourced from publications, web services, or author communications; links may be outdated.

b For non-commercial use only.

Another factor worth considering is the stringency of motif matching, which depends on motif size and representation, as well as the quantitative criteria and numerical thresholds applied during the search. Motif representations based solely on alpha-carbon atoms are less specific than those incorporating side-chain atoms. Structural motifs spanning only a small number of residues are similarly less specific. Strict matching parameters (such as restricting matches to identical residue types or enforcing a low RMSD threshold) may restrict results to closely related homologs, even though relaxed criteria could yield meaningful matches with more evolutionarily distant proteins.

Most Methods employ scoring functions to rank hits and evaluate match quality. For instance, RMSD values indicate the degree of geometric proximity between structural motif points and corresponding atoms in the target structure. While RMSD is a suitable metric for ranking motifs by their similarity to a reference template, it cannot be reliably used to compare motifs of different sizes. Furthermore, certain motifs are inherently more prone to accidental matches simply because they contain a higher proportion of common residues. To address these issues and improve hit ranking, some methods calculate statistical significance (P-value) or expectation values (E-value). However, their applicability comes with certain limitations, as these values depend heavily on foundational assumptions within the statistical model and the data used for its parameterization.

Regardless of how a structural motif was derived, it can be tested against a benchmark set of structures. When this benchmark includes reliable positive and negative examples, performance can be evaluated in terms of sensitivity (The ability to correctly retrieve positive examples) and specificity (the ability to exclude negative examples). When the benchmark consists solely of negative examples, the resulting RMSD distribution can be used to estimate the statistical significance of matches to that motif. The reliability of these derived metrics depends on having a sufficiently large and representative test dataset.

A consensus approach can be highly beneficial, where multiple hits to related motifs or consistent results obtained from different software tools and databases converge on a unified prediction.

Finally, standard biological common sense should always be applied. For example, a high-scoring match to an active-site motif is implausible if the protein lacks a corresponding substrate-binding pocket. Moreover, statistical significance and biological relevance are not synonymous; a biologically significant motif may fail to achieve high statistical significance compared to a structurally similar but functionally irrelevant motif. Therefore, before using motifs to infer the function or other characteristics of a protein, it is crucial to visually inspect the matches and evaluate them against sound biological criteria.



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.