Protein Structure and Function. Application of Bioinformatics Methods - John Rigden 2014

Functional diversity in packing elements and superfamilies
Function determination

Benoit H. Dessailly, Christine A. Orengo

With the early successes of structural Genomics, an ever-growing number of three-dimensional protein structures with unknown Functions has become available. Nevertheless, the extent to which structural information actually AIDS in understanding function remains a subject of debate. This chapter examines The Significance of Methods that establish structural relationships between Proteins at various levels (typically the folding motif and superfamily) in order to subsequently transfer functional annotations. First, we explore the functional Diversity of proteins that share a common folding motif. Identifying a folding motif can sometimes offer a valuable clue when determining functional properties. However, it has been demonstrated that the functional variety corresponding to a single fold can be remarkably broad; furthermore, for certain highly diverse folds (such as superfolds), the functional insight gained in this manner proves to be very sparse. Next, we analyze the functional diversity among proteins within the same superfamily (homologous proteins), as structural data can assist in identifying Homology even in the absence of sequence similarity. The evolutionary foundations and mechanisms driving the existing functional diversity among related proteins are discussed. Finally, we review useful tools for the integrated Analysis of Protein Structure, function, and evolution.

B. N. Dessailly and C. A. Orengo*

Department of Structural and Molecular Biology, University College London,

London WC1E 6BT, UK

*e-mail: orengo@biochemistry.ucl.ac.uk

Before discussing how identifying relationships at the fold or superfamily level can help in determining protein function, it is necessary to clearly define the term "function" as used in this chapter. This is essential, among other reasons, to outline the specific aspects of function that can be most reliably inferred from structural information.

Function is a relatively vague concept encompassing many different facets of protein activity. Moreover, the dimensions implied by this term vary depending on the specific field of protein science. For instance, a physiologist would likely describe a protein function in terms of its impact on the global phenotype (e.g., "Cell death initiator"), whereas a biochemist typically defines the function of a studied protein based on its characteristic molecular interactions or catalytic activity (e.g., "receptor-interacting Serine/Threonine protein kinase"). Because of these divergent usages of the term, providing a universal and broadly applicable definition of protein function is extremely challenging.

However, formulating a universal definition is not strictly necessary. The Gene Ontology (GO) Consortium has proposed a general framework that makes it possible to define and, more importantly, classify Protein Functions in a standardized way (The Gene Ontology Consortium 2000). In GO, protein function is examined from three distinct Perspectives, with an independent definition provided for each. According to GO, the cellular component describes the biological structures where the protein resides (e.g., Nucleus or ribosome); the biological process encompasses the objectives or pathways in which the protein participates (e.g., METABOLISM, signal Transduction, or Cell Differentiation); and the molecular function of a protein represents the set of activities it can perform (e.g., catalysis or transport).

Information on three-dimensional protein structures is primarily helpful for understanding catalytic mechanisms and predicting potential interactions with other molecules—both of which belong to the category of molecular functions. Consequently, whenever structural-functional relationships are addressed (as in this chapter), the Discussion focuses predominantly on molecular functions.

A number of Databases and annotation systems exist to describe the molecular Functions of Proteins (see Table 6.1); these are immensely useful when investigating structure-function relationships, particularly on an automated basis. Arguably, the oldest scheme for describing protein molecular functions is the Enzyme Commission (EC) numbering system, which classifies enzymatic reactions hierarchically. It employs a four-digit system where each successive level describes reaction features in increasing detail, ranging from the general type of catalytic activity (oxidoreductase, hydrolase, etc.) to the specific molecules acting as substrates (Nomenclature Committee of the IUBMB 1992). To overcome the long-standing Limitations of the EC Classification, two new databases for classifying Enzymes and their reactions have recently been developed: EzCatDB (Enzyme catalytic-mechanism Database) (Nagano 2005) and MACiE (Mechanism, Annotation and Classification in Enzymes) (Holliday et al. 2007). Both databases contain descriptions and classifications of enzymatic reaction mechanisms rather than information on the reactions per se. This stems from the growing acceptance of the notion that reaction-based classifications (such as the EC system) cannot always serve as reliable classifications of the corresponding enzymes (O'Boyle et al. 2007). In addition to these databases, the Catalytic Site Atlas provides detailed information on specific amino acid residues directly involved in catalytic mechanisms for enzymes of known structure (Porter et al. 2004). Several databases offer further descriptions of all protein residues participating in the binding of biologically important molecules, such as substrates and Cofactors (Lopez et al. 2007; Dessailly et al. 2008). Other widely used systems for annotating protein function include KEGG and FUNCAT. KEGG was originally designed to describe metabolic pathways and biological reaction networks, but has since evolved into a broader classification system for biological functions (Kanehisa et al. 2008). FUNCAT (the Functional Catalogue) classifies protein functions and constructs a unique hierarchical tree (Ruepp et al. 2004).

Class="center">Table 6.1. Links and brief descriptions of relevant databases and tools mentioned in the text

Name

URL

Description

CATH

http://cathwww.biochem.ucl.ac.uk/

Protein Structure classification

SCOP

http://scop.mrc-lmb.cam.ac.uk/scop/

Protein structure classification

SFLD

http://sfld.rbvi.ucsf.edu/

Functional Classification of enzyme superfamilies

PROCOGNATE

http://www.ebi.ac.uk/thornton-srv/databases/procognate/index.html

Mapping of domains to their known ligands

Gene Ontology

http://www.geneontology.org

Controlled vocabulary for protein functions

EC

http://www.chem.qmul.ac.uk/iubmb/enzyme/

Classification of enzymatic reactions

EzCatDB

http://mbs.cbrc.jp/EzCatDB/

Enzyme catalytic mechanism database

MACiE

http://www.ebi.ac.uk/thornton-srv/databases/MACiE/

Enzymatic reaction database

KEGG

http://www.genome.jp/kegg/

Integrated representation of genes, gene products, and metabolic pathways

FUNCAT

http://mips.gsf.de/projects/funcat

Protein function annotation scheme

DALI

http://ekhidna.biocenter.helsinki.fi/daliserver

Structural alignment

FATCAT

http://fatcat.burnham.org/

Flexible structural alignment

SSM

http://www.ebi.ac.uk/msd-srv/ssm/

Secondary structure matching

CATHEDRAL

http://www.cathdb.info/cgi-bin/CathedralServer.pl

Algorithm for identifying known folding motifs in protein structures

Like most databases of this type, both KEGG and FUNCAT primarily contain information on the biological processes in which the described proteins participate rather than the specific types of their molecular activity. Nevertheless, both databases provide data that can prove extremely useful when investigating the property referred to in the GO system as molecular function.



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.