Protein Structure and Function. Application of Bioinformatics Methods - John Rigden 2014
Structural motifs
Specific methods
Motif detection
Class="center">8.3.2.1. Literature
Perhaps the most reliable yet least automatable approach to motif discovery involves mining the published literature for experimental data regarding which residues are critical for protein function. In the case of Structure/135.html">Structural motifs, the focus is placed on residues that provide specific binding or catalytic capability rather than maintaining structural stability, although separating these functional aspects is not always straightforward.
The Catalytic Site Atlas (CSA) (Table 8.1) contains several hundred enzyme families, for each of which it provides a structure annotated with catalytic site residues derived from literature alongside a set of related sequences (Porter et al. 2004). Representative structural templates (structural motifs) based on side-chain functional atoms or a- and ß-carbon atoms are available for A number of families (Torrance et al. 2005). One can search for these motifs within a structure of interest or download the motif database from the CSA website (Table 8.1). The search is performed using the JESS program (Barker and Thornton 2003); matches between chemically similar residue types, such as aspartate and glutamate, are permitted. Statistical significance is evaluated using a formula that incorporates the number of residues in the motif, the number of points per residue, residue prevalence, and parameters empirically fitted from RMSD distributions viewed as exponents of power Functions (Stark et al. 2003). This formula estimates Background RMSD distributions a priori, thus eliminating the need to compare each motif against a random or reference set of structures.
8.3.2.2. Undirected Search
An undirected search entails discovering general patterns within a random set of structures, where “random” implies that the Selection was not based on the presence of any shared features or functions. In practice, there are far too many possible amino acid residue combinations in structures to examine them all, so the search space must be constrained.
Russell performed exhaustive pairwise structure comparisons within a representative set of structures (Russell 1998). The search space was constrained by distance conditions and the exclusion of non-polar residues, disulfide-bonded cysteines, and residues that were insufficiently conserved in sequence alignments. To identify instances of convergent evolution, matches between Proteins with similar folds were ignored. As a result, several metal-binding sites and active-site patterns, including the catalytic triad, were discovered.
The TRILOGY program similarly ignores residues that are insufficiently conserved in sequence alignments (Bradley et al. 2002). Requiring that the patterns are present in at least three different superfamilies According to the SCOR Classification, triples of potentially compatible residues—including conservative substitutions—are identified and merged into broader patterns. However, this program is designed to identify patterns in sequence and structure simultaneously rather than merely structural motifs; residue patterns must be similar in sequence space as well as in three-dimensional space.
Oldfield analyzed a representative set of structures by excluding small non-polar residues, representing all remaining residues as single points, grouping residue triplets into sets of similar types, and sorting the distances between residues in these groups into 0.5 Å-wide bins (Oldfield 2002). In the resulting three-dimensional histogram (each triplet having three inter-residue distances), bins with high population densities represent common patterns of such residue triplets. Wherever possible, these patterns were merged to incorporate more than three residues. This Procedure successfully identified several well-known structural motifs, such as binding sites and catalytic triads. The same study also describes programs for searching motifs within protein structures (Oldfield 2002).
Another study involved all-against-all pairwise comparisons of putative functional sites formed by interior pocket residues that are either located near a Ligand or conserved in sequence alignments (Ausiello et al. 2007). Although aimed at identifying sites potentially critical for protein function, this search remained undirected because the structures were not grouped by any structural or functional criteria. To identify cases of convergent evolution, the authors focused on instances where the matched residues appeared in a different order within their respective sequences. Both known Examples of such permutations and novel ones were uncovered. Multiple-residue matches were found using the Query3D program (Ausiello et al. 2005a), which performs an exhaustive depth-first search using a two-point representation of residues: the a-carbon atom and the geometric center of the side chain. Query3D detects matches of up to ten residue pairs where corresponding residues belong to similar types and motifs superimpose with an RMSD below a threshold value.
8.3.2.3. Individual Structures
Certain structural motif Databases have been constructed using information from individual structures alone. For instance, binding site motifs can be compiled by examining residues located within a specific distance from ligands, Nucleic Acids, or even other protein chains. Often, these studies focus on methodological development rather than database creation, and some also present results from Other types of research, such as searching for literature-described motifs.
The PINTS server (Patterns in Non-homologous Tertiary Structures) (Stark and Russell 2003) (Table 8.1) compares a query structure against a database of structural motifs representing either binding sites—defined as residues located within 3 Å of a ligand—or motifs annotated in the SITE record of a PDB Cell/13.html">Protein Structure file. Alternatively, User-defined motifs can be compared against protein databases (e.g., presented at various SCOP classification levels), or two specific structures can be compared directly. Much like Russell’s earlier work on undirected searching, PINTS performs a depth-first search, ignores non-polar residues, utilizes side-chain atoms, and permits matches between certain similar atom types. Statistical significance is evaluated using the authors' method (Stark et al. 2003) described above for the CSA. The PINTS website also provides weekly comparison results of structures newly deposited in the PDB against motif databases (Stark et al. 2004) (Table 8.1).
In addition to motifs recorded in SITE records, the PDBSite database (Ivanisenko et al. 2005) (Table 8.1) includes interaction sites with other proteins, RNA, and DNA. An interaction site encompasses residues having at least three atoms located within 5 Å of another chain. All sites in the database, or a subset thereof, can be compared with a query structure using the PDBSiteScan program (Ivanisenko et al. 2004) (Table 8.1), which utilizes residue type data, main-chain atom positions, and user-defined threshold values. The resulting superposition of the query structure and the identified structural motifs can be downloaded in PDB format.
The RIGOR program is essentially identical to SPASM, except that it performs the reverse process: comparing a structure against a structural motif database rather than a motif against a structural database (Kleywegt 1999). RIGOR executables and its associated motif databases are available for download from the Uppsala Software Factory website (Table 8.2). The database includes sites surrounding bound ligands, fragments composed of consecutive identical residues, and several other residue groupings. Each ligand-binding site is included twice—once with residue type labeling and once without. A match to an unlabeled motif can indicate that such a motif might be engineered into a query structure using protein design Methods.
Structural motif detection is central to function prediction performed by two servers that integrate results from third-party data. These servers are discussed in detail in Chapter 11 and are briefly mentioned here for completeness. The ProKnow server (Pal and Eisenberg 2005) (Table 8.1) performs multiple sequence- or structure-based searches for a query structure, including the search for single-structure structural motifs using RIGOR. Each database searched by ProKnow contains Gene Ontology (GO) annotations, and the ultimate output is a list of potential GO terms for the query structure along with their Bayesian confidence scores. However, many GO terms are rather general in nature. A second integrative method, the ProFunc server (Laskowski et al. 2005b) (Table 8.1), also performs multiple sequence- and structure-based searches.
The JESS program (Barker and Thornton 2003) is used to search for enzyme active-site templates within the CSA database and to search for ligand- or nucleic-acid-binding residue triplets in a non-redundant subset of the PDB. To achieve more comprehensive coverage of structural space, an “inverse template” search is also performed, wherein a query structure is broken down into structural motifs that are then compared against a representative set of parent PDB structures (Laskowski et al. 2005a). JESS hits are subsequently scored by expanding the residue comparison region into a 10 Å sphere centered on the found motif. Matches are ranked using a scoring function that rewards the superposition of similar residue pairs having similar sequence positions and sequential order. Consequently, the Motif Search becomes more specific yet less local, rendering it less suitable for identifying cases of convergent evolution. Each motif is searched against a sample of PDB structures, and expectation values (E-values) are calculated assuming an extreme value distribution of the scoring function (Laskowski et al. 2005a).
The SuMo server (Jambon et al. 2005) (Table 8.1) compares a query structure either against a database of “complete PDB structures” (all structures with redundant chains removed) or exclusively against ligand-binding sites from the same database. The query structure can be an entire structure, a single chain, or simply a ligand-binding site. The server represents structures as graphs of chemical group triangles, encompassing various hydrogen-bond Donors and acceptors, aromatic rings, and so forth (Jambon et al. 2003). When comparing a pair of structures, SuMo first identifies pairs of similar triangles and then finds consistent sets of pairs, or patches. These patches are further refined by removing pairs of chemical groups that superimpose poorly or differ significantly in their burial depth.
The Protemot web server (Chang et al. 2006) (Table 8.1) compares a structure against a database of motif binding sites, defined as residues where at least one atom lies within 4.5 Å of a ligand. Redundant chains with 60% sequence identity and binding sites for biologically uninteresting ligands were excluded from the database. Users can search all motifs, motifs occurring in Enzymes, or motifs specific to particular enzyme classes. During the search, the query structure is reduced to a-carbon atoms with labeled residue types located near binding pocket cavities. This information is hashed and compared against database element hashes. To match residues, the user sets a similarity threshold. From these rough matches, the top one hundred are refined to account for additional residues, ultimately retaining only those matches that exhibit identical pocket orientations and an RMSD within 1.5 Å. These matches are displayed graphically, though a table of matched residues is not provided; only PDB hit codes are reported.
The pdbFun web server (Ausiello et al. 2005b) (Table 8.1) enables the comparison of query and target residue sets using the Query3D program (Ausiello et al. 2005a) (described in the previous section). Residues can be specified manually one by one, or predefined sets and Boolean combinations of such sets can be utilized. One type of predefined set represents a binding site—i.e., residues located within 3.5 Å of a ligand. Alternatively, active sites from the CATRES database derived from literature data can be employed (Bartlett et al. 2002). A query set may contain residues from a single chain, whereas a target set can encompass up to the entire pdbFun database (~50,000 chains). While specifying query and target sets is extremely user-friendly, it can also be misleading. To ensure rapid searching, RMSD thresholds are set very strictly and cannot be adjusted. However, the Query3D program can be obtained From the Authors for local use (on Unix platforms), in which case the user can customize the thresholds as desired.
8.3.2.4. Positive Examples
Local structural features common to all proteins performing a specific function or belonging to a particular structural class can be treated as structural motifs. This approach leverages diverse positive examples to determine which atoms or residues may be incorporated into a motif, although the motif coordinates themselves may be extracted from a single structure rather than averaged data. Negative examples are omitted during motif derivation, though they are frequently utilized in evaluating these motifs.
Certain ligand-centric studies exploit the rigid portion of a ligand to superimpose binding sites. For instance, the alignment of various adenine mononucleotide binding sites was achieved by superimposing their adenine moieties. One study encompassed all-all comparisons of the 121 adenine mononucleotide complexes available at the time (38 complexes after redundancy filtering) (Kobayashi and Go 1997). For each structure pair, the number of corresponding atom pairs (based on element type and spatial adjacency) near the adenine and the degree of their superposition were evaluated. High similarity was discovered among structures featuring distinct folds: they shared a common structural motif comprising four-residue main-chain segments and three sequence-discontinuous residues (Kobayashi and Go 1997).
A analogous approach was employed to construct consensus binding-site motifs (Nebel et al. 2007). Because a vastly greater number of structures had become available by the time of this study, complexes with adenine mono-, di-, and triphosphates were examined separately. The similarity between a structure pair was evaluated as the fraction of ligand-neighborhood atoms present in both structures. Structures were clustered according to these similarity scores, and outliers were discarded. Within each cluster, only atoms conserved across all pairwise comparisons were retained, and their positions were averaged to generate a structural motif. Finally, highly similar motifs were merged. The resulting 13 motifs, derived from the analysis of 3 to 20 structures, contain between 6 and 71 atoms and in most cases correspond to established structural or functional classifications. Motif coordinates are available as supplementary data to the publication (Nebel et al. 2007).
Consensus templates for porphyrin-binding sites were developed using the Nestor3D program (Nebel 2006). Templates generated by this software can incorporate atoms, functional groups represented as pseudoatoms, and “solvent” (effectively grid points representing the cavity volume). The Nestor3D program also features a graphical user interface and is freely downloadable (Table 8.2). Users must specify a list of PDB files and the atoms suitable for structural superposition; various other parameters can be customized additionally.
Exhaustive comparisons of 3,737 phosphate environments from protein-nucleotide complexes enabled their classification into 476 compact clusters and 10 broader groups (Brakoulias and Jackson 2004). The resulting cluster division largely agrees with classifications based on global protein structure or function. An efficient clique-detection algorithm was employed to identify matching atom sets, eliminating the need to use ligand atoms for structural superposition.
The SOIPPA (Sequence Order-Independent Profile-Profile Alignment) program identifies common local structure patterns through pairwise comparisons (Xie and Bourne 2008). The protein structure is reduced to its a-carbon atoms, each assigned a geometric potential value and a substitution profile derived from automated sequence alignment. The geometric potential of an a-carbon atom is calculated based on its distance to the protein surface and the arrangement of neighboring a-carbon atoms (Xie and Bourne 2007). A potential match between two structures begins with a pair of points sharing similar geometric potentials; adjacent pairs are added if they are consistent in terms of distances and angles relative to the surface normal.
Each pair of a-carbon atoms is assigned a weight based on the similarity of their substitution profiles, followed by the identification of the subgraph with the maximum total weight. The alignment scoring function after atom superposition is a sum over all pairs, incorporating the pair weight, the degree of spatial overlap between the paired atoms, and the angle between the surface normals. Statistical significance is evaluated using a nonparametric distribution model of scoring function values when a pattern is compared against a representative set of structures. The SOIPPA program has been used to compare diverse adenine-binding structures and to retrieve a representative set of structures for superposition with known functional sites; the program proved capable of aligning binding sites and revealing local similarities superior to global sequence or structural comparisons. This work focused more on relationship identification than on motif discovery.
To identify functionally important atoms in structures sharing a common function but evolutionarily unrelated or distantly related, the Common Structural Cliques method was proposed (Milik et al. 2003). Each protein is reduced to a graph comprising only the representative atoms of each side chain. Next, sets of four atoms are extracted and compared to determine common structural cliques—namely, sets of atoms with equivalent types and interatomic distances in both structures. These cliques are then assembled into larger sets of corresponding atoms.
Notably, the resulting structural motifs may exhibit varying weights, even down to zero, for different interatomic distances, enabling the detection of matches even when specific distances vary significantly due to conformational flexibility. For instance, a motif may include atoms located within one of the hinge-connected domains. Low or zero weights assigned to interdomain distances allow for the identification of a motif in structures with distinct hinge Conformations, whereas intradomain distance weights can remain high to define precise geometric relationships within each domain. A limitation of this method is the inability to automatically combine results obtained from pairwise comparisons.
The DRESPAT (Detection of REcurring Sidechain PATtems) program extracts a common motif from a set of positive example structures (Wangikar et al. 2003). Each protein is reduced to a graph of functional atoms (one per residue), excluding residues with nonpolar side chains and disulfide-bonded cysteines. Patterns of three or more residues are then extracted and compared with patterns from other structures consisting of residues of the same type. In addition to functional atoms, a- and ß-carbon atoms are taken into account, while matches with distance deviations between points and/or RMSD values exceeding specified thresholds are filtered out. Other adjustable parameters include the pattern size (defaulting to three to six residues) and the minimum number of input structures that must contain the pattern.
Based on pattern occurrence within a randomly selected set of structures, empirical relationships were derived to calculate the statistical significance of detected motifs according to their size, the total number of structures, and the number of structures required to contain the region. Results were presented for non-redundant sets of 17 SCOP superfamilies. It was found that motifs consisting of at least four residues and derived from sets containing 5 or more structures typically correspond to functional sites. Relying solely on pairwise comparisons of evolutionarily related structures yielded an excessive number of false-positive patterns. The DRESPAT program can be obtained from its developers as C++ source code (Wangikar et al. 2003).
The fiinClust server (Ausiello et al. 2008) (Table 8.1) identifies structural motifs common to multiple input structures, up to a limit of 20. Structures are further filtered based on sequence identity and then pairwise compared using the Query3D program (Ausiello et al. 2005a). Query3D represents each residue using two points: the alpha-carbon and the geometric center of the side chain. In addition to maximum sequence identity, the user can specify whether the thresholds for RMSD and side-chain proximity should be low, medium, or high; whether hydrophobic or buried residues should be excluded; and whether the alignment of similar rather than strictly identical residue types should be permitted. The server reports motifs comprising three or more residues found in three or more input structures.
The PAR-3D (Protein Active Site Residues using 3-Dimensional structural motifs) server (Goyal et al. 2007) (Table 8.1) compares an uploaded structure against motifs representing six protease classes, ten glycolytic pathway enzymes, and metal-containing sites comprising 3 or 4 residues (Goyal and Mande 2008). These motifs, each derived from a training set of structures, are represented as allowable ranges of interatomic distances or other geometric scalars rather than spatial coordinates. Sensitivity and Specificity values for the motifs are available on the website, and the program itself (Perl scripts and associated data files) can also be downloaded from there.
8.3.2.5. Positive and Negative Examples
The primary distinction between the "positive and negative examples" approach and the "positive examples only" approach is that the former explicitly considers structures that do not belong to the class of interest during motif discovery. In other words, motif generation and the evaluation of its specificity are interconnected.
Geometric filtering refines either an existing motif or a list of potentially important residues based on their geometric uniqueness (Chen et al. 2007a). RMSD distributions for motif candidates (subsets of the input list) are obtained by comparing them against a representative set of structures. Geometric filtering does not require Separation into positive and negative examples; instead, it assumes that the low-RMSD tail of the distribution curve represents true positives, while the remainder corresponds to false positives. Among motifs with a given number of residues, the one with the highest median RMSD is selected as the most geometrically unique, as it provides the best separation between the bulk of the distribution and the low-RMSD tail. The main limitation of this approach is that the "correct" residues must be included in the initial motif.
Considering both positive and negative structural examples, the GASPS (Genetic Algorithm Search for Patterns in Structures) algorithm identifies residue patterns that best separate these two groups (Polacco and Babbitt 2006). A preliminary list of residues is not required, and the method is independent of how positive and negative example groups are defined. The core search engine is the SPASM program (Kleywegt 1999), which represents residues as alpha-carbons and side-chain geometric centers and permits the superposition of identical residue types only. To constrain the search space, GASPS considers only the 100 most conserved structural residues, determined from an automatically generated sequence alignment. An initial motif candidate is constructed by randomly selecting one residue and subsequently choosing four additional residues located at a distance from the first.
Each of the 50 initial candidates is evaluated based on how effectively it separates positive and negative structural examples in terms of best-match RMSD values. In each Generation of the genetic algorithm, the 16 highest-scoring motifs serve as parents for 36 new motifs, and after 50 cycles, the top-scoring motif is declared the winner. Motifs may contain between three and ten residues. Sensitive and specific motifs have been generated for various superfamilies (Babbitt and Gerlt 2000) and Serine proteases. Most residues within the motifs were found to be functionally important; however, in some instances, residues lacking a known functional role exhibited comparable predictive value (Polacco and Babbitt 2006).
The GASPSdb server (Polacco and Babbitt, in press) (Table 8.1) compares a query structure against a database of structural motifs previously generated using the GASPS program across several Protein Classification schemes: SCOP superfamilies and families, Gene Ontology (GO) molecular function categories, and proteins belonging to SCOP superfamilies while sharing a GO molecular function. A non-redundant set of structures was utilized. A motif was generated for each structure within the positive example group, with structures from all other groups treated as negative examples. Motifs were constructed exclusively for groups comprising at least 6 non-redundant structures. Searches on the server are powered by the RIGOR program (Kleywegt 1999). Commercial use of this product requires contacting the developers to obtain a license (refer to the Uppsala Software Factory website, Table 8.2). Statistical significance is assessed using a function developed by the authors of the PINTS program (Stark et al. 2003).
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.