Protein Structure and Function: Applications of Bioinformatics Methods - John Rigden 2014
Spatial Motifs
Background and Significance
Elaine C. Meng, Benjamin J. Polacco, Patricia C. Babbitt
Structure/135.html">Structural motifs are patterns of local Cell/13.html">Protein Structure associated with function, typically consisting of amino acid residues in a binding site or catalytic center. Proteins of unknown function can be annotated by comparing them with known structural motifs. A large number of Methods have been developed to identify structural motifs and search for them within structures. These methods vary in the type and amount of input data, the description of motifs and their matching criteria, whether statistical significance is accounted for in the results, and how motif-function relationships are established. Less progress has been made, compared to algorithm development, in creating publicly available structural motif Databases that are both functionally specific and broadly comprehensive of diverse Functions. This challenge stems from the difficulties in developing detailed structure-function classifications; consequently, large-scale automated studies have instead relied on pre-existing structural or functional classifications. Complementary to structural motif identification methods are approaches focused on molecular surface description, global structure (fold type) comparison, prediction of interactions with other macromolecules, and the identification of physiological substrates by docking small molecules from appropriate databases.
Elaine С. Meng, Benjamin J. Polacco, and Patricia C. Babbitt
University of California San Francisco (UCSF)
Department of Pharmaceutical Chemistry,
600 16th Street, San Francisco, CA 94158-2517
Patricia C. Babbitt
UCSF Department of Biopharmaceutical Sciences,
1700 4th Street, San Francisco, CA 94158-2330
e-mai1:babbitt@cgl.ucsf.edu
3D - three-dimensional, structural,
CSA - Catalytic Site Atlas,
DRESPAT - Detection of REcurring Sidechain PATterns,
EC - Enzyme Commission,
FFF - Fuzzy Functional Form,
GASPS - Genetic Algorithm Search for Patterns in Structures,
GO - Gene Ontology,
PAR-3D - Protein Active Site Residues using 3-Dimensional structural motifs,
PDB - Protein Data Bank,
PINTS - Patterns in Non-homologous Tertiary Structures,
S-BLEST - Structure-Based Local Environment Search Tool,
SCOP - Structural Classification of Proteins,
SOIPPA - Sequence Order-Independent Profile-Profile Alignment,
SPASM - SPatial Arrangements of Sidechains and Mainchains,
TESS — TEmplate Search and Superposition, DB — database,
RMSD — ROOT-mean-square deviation,
EC — Enzyme Classification,
MD — Molecular Dynamics
The application of a genomic approach to biology has generated not only vast amounts of sequence and structural data, but also the prospect of obtaining a complete "parts list" for many organisms. However, such a list is of limited use without some understanding of what each part is meant to do. Even with entire genome sequences in hand, not all genes have been identified, and a significant number of identified genes lack any annotated function. Because the number of sequences vastly exceeds the number of available structures, functional assignment (functional annotation) has largely been carried out by large-scale sequence space searching, transferring functional information from any sufficiently similar sequences to the protein of interest (annotation by Homology). Many structural motifs have been identified in specific sets of proteins and linked to some aspect of protein function or structure. However, the reliability and functional Specificity of annotation by homology decrease as sequences become less similar (Devos and Valencia 2001; Rost 2002). By functional specificity, we mean the narrowness of our inference; for example, the term "leucine aminopeptidase" is more specific than "peptidase".
Examining protein structures can reveal important similarities or potential evolutionary relationships that are not apparent from sequence analysis alone. Proteins may diverge during evolution to such an extent that their sequences can no longer be reliably aligned, yet the overall structural similarity, or fold, is still preserved (Chothia and Lesk 1986; Rost 1997). Using fold similarity for annotation by homology (see Chapter 6) shares the same limitations as sequence similarity: on the one hand, the reliability of homology-based annotation decreases with increasing evolutionary distance between related proteins, while on the other hand, proteins with very similar folds may perform entirely different functions (Babbitt and Gerlt 1997; Todd et al. 2001). Therefore, to accurately describe and predict protein function, one must examine structural details, exemplified by structural motifs, which represent patterns of local structure. In the course of evolution, identical proteins may diverge through the accumulation of random neutral changes that do not alter function (neutral drift) while preserving the structural components essential for that function. Ideally, these functionally required structural components are what structural motifs describe, also serving as sensitive and specific markers of function. The presence of a common structural motif may also reflect convergent evolution between different folds, capturing a similar arrangement of side chains associated with a comparable function. A well-known example is the Asp-His-Ser catalytic triad of Serine proteases, which is utilized in a mechanistically similar manner by structurally diverse proteases (see below).
Structural Genomics projects aim to determine the structures of all proteins, acknowledging the critical importance of this information for annotation and other Applications such as drug discovery. The sheer scale of this task can be somewhat mitigated by clustering similar sequences and selecting a representative target for each group, thereby enabling comparative modeling for the remaining sequences. Over the past few years, the total number of structures in the Protein Data Bank (PDB) (Berman et al. 2000) and in structural genomics projects has been growing at an accelerating pace, with the function of many of these structures remaining unknown. This trend suggests that structure-motif methods will become increasingly widespread and useful as a growing number of experimentally solved and modeled structures become available.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.