Protein Structure and Function: Application of Bioinformatics Methods - John Rigden 2014
Spatial Motifs
Overview of Methods
Identification and Selection of Motifs
The points forming a structural motif are defined as atoms or pseudoatoms derived directly from the atomic coordinates within a Structure. For instance, the geometric center of a side chain represents a pseudoatom whose coordinates equal the arithmetic mean of the coordinates of the side-chain atoms. When describing a motif, up to several points are used from each residue belonging to the motif, and all points are labeled with additional information, such as the atom type, residue type, or physicochemical characteristics.
When screening a structure for matches against a structural motif, quantitative rules are employed to specify which point in the structure can be mapped to which point in the motif, alongside geometric thresholds that determine whether a given set of points exhibits sufficient spatial similarity to be considered a match (a hit). The degree of accuracy of this match also depends on the number of residues and points within the motif. A trade-off exists between match accuracy and the permitted tolerance limit: while it is desirable to account for residue substitution, conformational flexibility, and low-resolution structures, doing so increases the number of biologically meaningless hits and makes it harder to isolate meaningful ones. The inclusion of specific atoms in Structural motifs is designed to highlight local interactions, such as Hydrogen Bonds, whereas The Use of geometric centers of functional groups or side chains is better suited for accommodating flexibility and residue-type variations (Fig. 8.1). Furthermore, representing a symmetrical side chain—such as the aromatic ring in Phe—as a single point eliminates the need to compare them through various orientations (Oldfield 2002).
Motif searching can be computationally demanding, especially considering that thousands of structures may need to be compared against thousands of motifs. Structural motif searches rely on The Development of efficient algorithms, which often incorporate one or more of the following approaches:
1) Geometric hashing. Hashing is a broad term describing the reduction of complex data into a simpler form that can be compared more rapidly. Numerous values, such as distances, angles, and atom types, can be reduced via a specific function to a few numbers or even a single number. Other sets of values that yield the same result correspond to potentially matching substructures. Geometric hashing encodes spatial relationships between points (Fischer et al. 1994), but it may also incorporate Other types of information, such as physicochemical descriptors (Shulman-Peleg et al. 2004). Prior to executing computationally intensive transformation steps and evaluating scoring Functions, individual substructure matches that imply similar transformations (Translation/rotation to align corresponding points) can be clustered into broader groups (Pennec and Ayache 1998). Hashing, or data preprocessing, requires time, but it needs to be performed only once for each structure and can substantially accelerate overall comparisons.
Class="center">
Fig. 8.1. (For the color version of this figure, see the color insert.) Residues of the catalytic center of members of the enolase superfamily, illustrating aspects of motif representation and Specificity. Superimposed side chains of two basic and three acidic residues are shown for each of the following Proteins: mandelate racemase (yellow, PDB 2mnr), enolase (orange-pink, PDB 4enl), and methylaspartate ammonia-lyase (blue, PDB 1kcz). The positions of alpha carbons and side-chain centers of mass are depicted as spheres. The single-letter code displayed next to the alpha carbons denotes different residue types: H — Histidine, K — Lysine, D — aspartic acid, E — glutamic acid. Although The amino acid residues shown in the lower-left corner are highly conserved in type and conformation, the Active Site incorporates the following variations: (1) different (though similar) residue types at three other positions; (2) different side-chain Conformations, illustrated by the two lysines on the right; (3) different positions in the Primary Structure: the basic residue depicted in the top left is C-terminal in enolase, but N-terminal in The sequence of the other two proteins. Using the center of mass of the side chain rather than the positions of functional atoms generally reduces sensitivity to changes in conformation and residue types. Factoring in side-chain atoms or the center of mass (excluding the main chain) diminishes the allowable tolerance for side-chain mobility while, conversely, providing greater specificity regarding the precise spatial arrangement of atoms within the functional site. Incorporating side-chain atoms also decreases sensitivity to functional residue migration events, where a crucial side chain may belong to Amino Acids located at different positions in the primary sequence (Todd et al. 2002). The figure was generated using the UCSF Chimera visualization software (Pettersen et al. 2004) (http://www.cgl.ucsf.edu/chimera)
2) Graph-based Methods. A graph consists of vertices (points) and edges (lines connecting pairs of vertices). A molecular structure or structural motif can be viewed as a labeled graph. For example, atoms can be represented as vertices labeled with their residue type, with edges connecting each pair of vertices, each assigned a corresponding interatomic distance. Using subgraph isomorphism algorithms, smaller graphs and all their edges are searched for within a larger graph. Such algorithms can be applied in conjunction with labeling to identify a set of atoms in structures that match the atom types and interatomic distances of structural motifs (Artymiuk et al. 1994; Spriggs et al. 2003). Permissible deviations in values allow for the comparison of similar rather than identical distances. Graph clique detection (Schmitt et al. 2002) is ultimately a analogous Procedure, but in this case, the graph describes the geometry of both structures simultaneously. Here, the graph vertices represent pairs of atoms or pseudoatoms, one belonging to structure A and the other to structure B (where a "structure" can also be a structural motif). Only identical atom types may form pairs. Two vertices are connected by an edge if the distance between the two atoms in structure A matches the distance between the atoms in structure B within allowed tolerances. A clique is a graph in which every vertex is connected to all other vertices. Thus, clique detection makes it possible to identify a set of atoms in structure A with internal distances fully compatible with those formed by the pairs of atoms in structure B.
3) Depth-first search. All structures are thoroughly examined for the presence of a structural motif. The search space is defined by constraints on the atom or residue types that can be matched, as well as geometric thresholds such as tolerance limits and the upper bound of the ROOT-mean-square deviation (RMSD).
Match extension. Partial or seed matches to the structural motif are identified first, after which attempts are made to expand the match to encompass the entire motif.
All these methods search for motifs within a given static structure. Recent findings demonstrate that combining motif-searching algorithms with MD calculations that sample conformational space (see also Chapter 9) represents a promising, albeit computationally expensive, route to improving search outcomes (Glazer et al. 2008).
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.