Protein Structure and Function: Application of Bioinformatics Methods - John Rigden 2014
Examples: predicting the function of structures obtained from structural genomics projects
Several Special Examples
Although there are relatively few large-scale projects, numerous individual structures of interest have been published by or in collaboration with various structural Genomics consortia, where knowing the Cell/13.html">Protein Structure proved essential for elucidating its function. A prime example is the TT0936 protein from Thermotoga maritima (PDB code 2plm), which was obtained from an external laboratory using clones provided by the Joint Center for Structural Genomics. The protein was originally annotated in Databases as a hypothetical protein belonging to the Pfam amidohydrolase family (PF01979), which includes several deaminases and forms part of the broader amidohydrolase superfamily. Selected as a case study for function prediction using the ProFunc server, TT0936 revealed a potential match with one of the 189 known active-site templates utilized by the server.
A strikingly similar Active Site (with an expectation value of $E = 2.45 \times 10^{-4}$) was identified as the adenosine deaminase active-site template (EC 3.5.4.4) derived from the PDB structure 1a41, which is involved in purine METABOLISM. The overall sequence identity between this structure and TT0936, determined via pairwise alignment using FASTA, was only 24%. However, the structural similarity—calculated by ProFunc as the fraction of residue pairs lying within one or more matched segments relative to the total number of equivalent residues in the alignment—reached an impressive 95%. Within a 10 Å sphere around the active-site template, the local sequence identity rises to 27.7%, which is higher than the structural average and points to greater sequence conservation in the immediate vicinity of the active site. In addition to a strong resemblance to adenosine deaminase, several hits with reverse templates of other deaminases and amidohydrolases were also detected.
The authors of the TT0936 structure published a study predicting its function to be adenine deaminase (Hermann et al. 2007). Their approach involved molecular docking of high-energy metabolic intermediates into the protein structure, based on the premise that docking ground-state substrates or products might be less effective than docking enzyme-stabilized transition-state analogues. The resulting list of potential ligands consisted predominantly of adenine analogues well-suited for C6-deamination. Four of these ligands were tested as substrates, three of which exhibited significant catalytic rate constants. The structure of the complex between TM0936 and the product ($S$-inosylhomocysteine), generated by the deamination of $S$-adenosylhomocysteine, revealed a near-perfect match between this Ligand bound to TM0936 and deoxycoformycin (an inosine analogue) bound to the template structure used in determining the TM0936 structure (PDB code 1a41) (Fig. 11.1).
Interestingly, a fold analysis using MSDfold shows similarities to various amidohydrolases and guanine/cytosine deaminases. The reason this server fails to rank adenosine deaminase as the top hit is that the TM0936 structure features additional embellishments outside the aligned region, causing the structural similarity score to fall below the critical 70% threshold for the number of matching Secondary structure elements (Fig. 11.2). This observation highlights the power of local comparisons and serves as a textbook example of how a function can be accurately pinpointed via an active-site template match that would remain completely undetected by sequence analysis alone. Another compelling example comes from a recent publication by the Midwest Center for Structural Genomics (MCSG), which describes the open (R) and closed (T) states of prephenate dehydratase (PDT) (Tan et al. 2008), shedding light on the Allosteric Regulation of this enzyme by L-phenylalanine and Other Amino Acids. Prephenate dehydratase (EC 4.2.1.51) catalyzes The conversion of prephenate to phenylpyruvate in The Biosynthesis of L-phenylalanine, playing a vital role in this pathway in organisms utilizing the shikimate pathway, which makes the enzyme indispensable for microorganisms. Because humans lack this enzyme, it represents an attractive potential target for antibacterial drug design.
The prephenate dehydratase structures deposited in the PDB (codes 2qmw and 2qmx) originate from two different organisms and represent the first crystallographic snapshots of PDT in relaxed (R) and tense (T) states (from Staphylococcus aureus and Chlorobium tepidum, respectively). Despite a low sequence identity of 27.3%, these Enzymes share a conserved overall architecture and domain Organization: both are tetramers (forming dimers of dimers—see individual dimer views in Fig. 11.3) consisting of a catalytic domain (PDT domain) and a regulatory domain (ACT domain). Based on these PDT structures, the authors proposed that the active site is located within the cleft between the two PDT domains.
This structure-based prediction is strongly supported by sequence analysis and mutagenesis data. Multiple Sequence Alignment and mapping of conserved residues demonstrated that these residues cluster at the bottom of the inter-subdomain cleft. Mutagenesis studies identifying residues critical for PDT activity in E. coli mapped directly to equivalent positions within the cleft between the two PDT subdomains (Zhang et al. 2000). Further mutagenesis data for PDT in Corynebacterium glutamicum confirmed that equivalent residues are involved in substrate binding and/or catalysis (Hsu et al. 2004). Together, these findings confirm that both the cleft and its conserved residues form the Active Site of prephenate dehydratase, with T168 standing out as the most likely key catalytic residue.
Class="center">
Fig. 11.1. (For color version of this figure, see the plate section.) Active-site template match illustrating the spatial overlap between bound $S$-inosylhomocysteine in TM0936 (shown in orange; PDB code 2plm) and deoxycoformycin, an inosine analogue bound to the adenosine deaminase template (shown in purple; PDB code 1a41). Bound zinc atoms are displayed as overlapping spheres matching the colors of their respective ligands. The residues of TM0936 and adenosine deaminase are colored blue and red, respectively.

Fig. 11.2. Stereo view of the superposition of the query structure TM0936 (PDB code 2plm, shown in black) onto the template structure used for active-site matching (PDB code 1a4l, shown in grey). Additional secondary structure elements unique to the query protein are clearly visible on the left side of the image.
The identification of the putative active site was followed by the mapping of the allosteric site. The Location of this L-phenylalanine-binding site in prephenate dehydratase closely mirrors the effector-binding site found in other ACT domain-containing enzymes involved in binding amino acids or other small molecules. The presence of bound L-phenylalanine in the structure allowed the visualization of its interactions with the ACT domains in the Chlorobium tepidum protein. Inspection of the binding residues revealed that most interactions are likely non-specific, which explains why other amino acids, such as Methionine, can also bind at this site and modulate catalytic activity (Liberles et al. 2005).

Fig. 11.3. Prephenate dehydratase structures from the Midwest Center for Structural Genomics highlighting Similarities and differences: (a) R-state structure from Staphylococcus aureus (PDB code 2qmw); (b) T-state structure from Chlorobium tepidum (PDB code 2qmx). The sequence identity between these two enzymes is only 27.3%.

Fig. 11.4. Monomer of prephenate dehydratase from Staphylococcus aureus. Each domain is shown in a distinct color. The putative active site located in the cleft between the two domains is circled.
A structural comparison of PDT from Staphylococcus aureus and Chlorobium tepidum reveals how L-phenylalanine binding induces Conformational Changes in the ACT dimer. Tan et al. (2008) discuss in detail how these structural shifts propagate to the active site, thereby inhibiting enzymatic activity. The authors suggest that L-phenylalanine binding triggers a series of major local and global conformational adjustments that alter the relative orientation of domains across the protein, ultimately restricting access to the active site. Specifically, these rearrangements split a single broad funnel leading to the catalytic center in the middle of the PDT dimer into two smaller channels, which impedes The entry of prephenate and the release of phenylpyruvate. This example underscores how structural genomics analyses can yield insights extending far beyond mere functional prediction.
Another striking example from the Midwest Center for Structural Genomics is the AF0491 protein from A. fulgidus (Savchenko et al. 2005), which is a homologue of the human Shwachman-Diamond syndrome (SDS) protein. SDS is a rare autosomal recessive disorder caused by Mutations in the SBDS Gene on chromosome 7, characterized by exocrine pancreatic dysfunction, skeletal defects, and hematological abnormalities (Boocock et al. 2003). AF0491 serves as an archaeal homologue, and determining its structure revealed a three-domain architecture (Fig. 11.5).
The C-terminal domain adopts a highly prevalent fold, making its function difficult to predict. However, similar domains are known to occur in numerous RNA- and DNA-binding Proteins. The central domain also features a common fold—the winged helix-turn-helix (wHTH) motif. Such domains are frequently involved in DNA binding (Aravind et al. 2005) and are also found in RNA-binding proteins (Schade et al. 1999). In this particular case, however, The surface of AF0491 lacks the positive electrostatic potential typically associated with nucleic acid binding, making this function unlikely. Instead, as the authors suggest, the domain may mediate Protein-Protein Interactions.

Fig. 11.5. Monomer of the A. fulgidus AF0491 protein, a homologue of the human Shwachman-Diamond syndrome (SDS) protein. The three domains are colored light grey, dark grey, and black from the N- to the C-terminus.
The N-terminal domain adopts a novel fold, and it is precisely this domain that harbors the majority of disease-causing mutations identified in SDS patients. Subsequent structural database searches identified the same fold in the Yeast protein YHR087W. Discovering this structural homologue opens the door to experiments that are impossible to perform directly with the human protein. Experimental investigation of structural and sequence homologues of the SDS protein (YHR087W and YLR022C, respectively) points to a link with RNA metabolism. Strains with a deleted YLR022C gene proved non-viable, but TAP-tagged proteins co-purified with numerous ribosomal proteins and factors involved in rRNA Processing. Although YHR087W deletion strains were viable, crossing this knockout with a library of 383 strains carrying deletions in RNA metabolism genes revealed significant synthetic lethality in several combinations. This observation, together with the genetic interactions mapped for YHR087W, strongly Supports a role for this protein in RNA Processing. Although these findings tie the SDS protein to ribosome biogenesis, its exact cellular function remains unknown, and fundamental differences in ribosome biogenesis between Bacteria (Lecompte et al. 2002), eukaryotes, and archaea mean that functional inferences must be drawn with caution. Nonetheless, the SDS protein stands as a prime example of how solving the structure of a bacterial homologue for a human disease protein can lead to the identification of a yeast homologue, ultimately facilitating functional characterization.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.