Protein Structure and Function. Application of Bioinformatics Methods - John Rigden 2014

Prediction of protein function based on their theoretical models
Protein models as a publicly accessible resource
Model databases

To effectively utilize protein structural models, they must be presented to the scientific community in a way that makes them readily accessible to individual researchers. Establishing public repositories could greatly encourage The Use of these models in experimental design. Currently, two MAIN TYPES OF Databases containing spatial Cell/13.html">Protein Structure models are publicly available.

Databases such as MODBASE (Pieper et al. 2006) and the SWISS-MODEL Repository (Kopp and Schwede 2004) (see Table 12.1) contain models generated by fully automated Methods. Other databases, such as PMDB (Protein Model Database) (Castrignano et al. 2006), are designed for storing and managing manually built models.

Regardless of their underlying architecture, all of these databases feature web interfaces for retrieving models of interest and provide supplementary information, including model reliability estimates and protein function annotations.

To ensure that the models are sufficiently accurate for experimental research, the approaches implemented in the SWISS-MODEL Repository and MODBASE rely on comparative modeling, which is currently the most robust method for large-scale projects. Both databases report the sequence identity between the target protein and the template; however, the SWISS-MODEL Repository exclusively houses models with an identity greater than 40%. Target protein sequences are retrieved from the UniProt database (http://www.expasy.uniprot.org/), and both repositories are regularly updated in sync with underlying sequence and structure databases, incorporating newly added or modified target sequences and utilizing novel structures as templates. Both repositories also provide access to the sequence alignments used to generate the models, although their quality assessment approaches differ. The packing quality of models in the SWISS-MODEL Repository is evaluated using the ANOLEA method (Melo and Feytmans 1998) and energy calculations based on the GROMOS96 force field (van Gunsteren 1996), whereas MODBASE assesses model reliability using statistical criteria (Melo et al. 2002). Furthermore, the visualization tools in the SWISS-MODEL Repository offer graphical representations of domain-level annotations and protein function annotations according to InterPro, while MODBASE visualization also highlights potential Ligand-binding sites and single nucleotide polymorphism (SNP) annotation sites.

Class="center">Table 12.1. Web-accessible databases of computationally predicted models

Database

URL http://

SWISS-MODEL Repository (Kopp and Schwede 2004)

swissmodel.expasy.org/repository/

MODBASE (Pieper et al. 2006)

salilab.org/modbase

PMDB Protein Model Database (Castrignano et al. 2006)

www.caspur.it/PMDB

Protein Model Portal

www.proteinmodelportal.org/laboheme.df.ibilce.unesp.br/

DBMODELING (Silveira et al. 2005)

db_modeling/

The collections of models in MODBASE and the SWISS-MODEL Repository, along with data gathered by research centers participating in the PSI project, are accessible through a unified interface provided by the Protein Model Portal. This portal is designed to simultaneously query all integrated databases when searching for pre-computed models for a given Amino Acid Sequence, presenting the results as a list of models hosted in specific databases along with direct access to them.

Models generated with human curation—typically challenging cases requiring non-trivial modeling techniques and careful Selection of the most reliable structure among several alternatives—are hosted in the PMDB database. This database is designed to provide access to models published in scientific literature alongside their supporting experimental data, although currently the majority of entries consist of CASP predictions. PMDB allows individual modelers to contribute their own models and supporting experimental evidence, while users can freely download multiple models for the same target protein or models covering different regions of that protein. Where applicable, links to the SwissProt sequence database (http://www.ebi.ac.uk/swissprot/) are also provided. Once a protein's experimental structure is resolved, a corresponding link to the PDB entry (http://www.rcsb.org/pdb/home/home.do) is added to the database record.

In addition to large-scale model repositories, there are databases focused on specific organisms. A notable example is DBMODELING (Silveira et al. 2005), a database dedicated to supporting drug discovery against various infectious diseases. To date, this database contains comparative models for Proteins encoded in the genomes of only two pathogenic organisms: Mycobacterium tuberculosis and Xylella fastidiosa, the CAUSATIVE AGENT OF citrus variegated chlorosis.



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.