Protein Structure and Function. Application of Bioinformatics Methods - John Rigden 2014

Fold recognition
Detection of remote homology without alignment
Sequence profiles and Hidden Markov Models

While sequence Databases were expanding rapidly in line with global genome sequencing efforts, The Development of technologies aimed at effectively utilizing the generated information was still in its infancy. A simple approach employed by Park et al. (1997) demonstrates how two homologous sequences, having diverged far beyond the point where their Homology can be detected by a simple pairwise comparison, can be linked using a third sequence that serves as a suitable intermediate between the two. This “stepping stone” approach across sequence space, known as intermediate sequence search, held clear potential, and an advanced methodology was subsequently developed in PSI-BLAST (Altschul et al. 1997). Instead of using a fixed 20x20 scoring matrix for every protein and for each residue position within a protein, it became possible to generate an n×20 scoring matrix, or profile, which contained information regarding the specific mutational preferences at each position of a given protein sequence. For this reason, such a profile is often referred to as a position-specific scoring matrix (PSSM).

Following an initial standard BLAST search to detect relatively close homologs, a (pseudo-)Multiple Sequence Alignment of these homologs is performed against the query sequence. This alignment yields statistics on the observed Mutations for each position of the query sequence. These statistics form The basis of a new scoring matrix, which can then be used in subsequent search iterations. This process of searching for homologs, generating a new scoring function, and repeating the search using this new function can be carried out multiple times (typically 5 to 10) and is known as PSI-BLAST (Position-Specific Iterated BLAST). The combination of this efficient iterative approach with information from ever-growing sequence databases has greatly enhanced the detection of extremely distant homologs. At the CASP4 competition, research groups employing this approach (PSI-BLAST or a variation thereof) demonstrated superior performance compared to previously successful groups whose work relied on threading Methods.

The success of the PSI-BLAST approach stems from accounting for the fact that each position in a protein sequence is subject to its own evolutionary pressure. For instance, a Glycine residue at a specific position may be highly conserved if its presence provides a tight turn in the protein chain necessary to maintain topology. Any mutation at such a position could prove lethal due to the potential disruption of proper protein folding. At another position, a glycine residue may experience minimal selective pressure, residing in a highly variable loop region. Consequently, when aligning the query sequence by Structure, the presence of the first glycine residue is mandatory, whereas The Nature of the second residue can vary. It is precisely this consideration of mutational preference, determined inter alia by residue position, that makes the approach vastly more sensitive in detecting distant homology.

One of the most common Applications of PSI-BLAST-generated profiles is searching for the profile of a query sequence against the PDB database, or conversely, searching a query sequence against a database of template profiles. Profiles are not always generated using PSI-BLAST. For example, Hidden Markov Model (HMM)-based profiles are constructed using Multiple Sequence Alignments, yet they contain more information than standard profiles. Specifically, they incorporate information on the positions of typical insertions and deletions, as well as transition probabilities to and from matched states for each position in the chain. Again, this is frequently combined with predicted structural properties, such as Secondary structure. Alternative approaches based on sequence-profile and profile-sequence principles are schematically illustrated in Figs. 2.4c and 2.4d.

Improved profiles and HMMs can be built using structural alignments of remote homologs, as well as by incorporating sequences of unknown structure that can be readily aligned to any of the available structures (Kelley et al. 2000; Tang et al. 2003). However, employing structural alignments to construct higher-quality profiles often yields only marginal improvements in detecting remote homologs or in alignment accuracy. This is likely because sequence alignments derived from structural alignments lack uniqueness, particularly in the presence of large insertions or deletions, or significant structural shifts. These factors can lead to misalignments between the sets of sequences associated with each structure. A solution proposed by Zhou and Zhou (2005) in their successful SP3 method is to generate protein fragments and use them for profile construction.

In recent years, hidden Markov models have been widely adopted by various research groups with notable success. As mentioned earlier, one of the key advantages of HMMs over the relatively simpler profiles generated by PSI-BLAST is the inclusion of additional information regarding gaps and adjacent residues. Nevertheless, for both profiles and HMMs, the quality of the underlying multiple sequence alignment is of paramount importance. The choice of sequences and the quality of the alignment appear to be more critical to profile performance than the statistical methods utilized during the profile-building process. Consequently, many research groups have found it advantageous to gather homologous sequences using PSI-BLAST, while employing a separate, more robust program to construct a more accurate multiple sequence alignment.

As recently demonstrated, profile-profile and HMM-HMM alignments—serving as generalizations of sequence-profile or sequence-HMM comparison methods—exhibit significantly higher performance. Thus, rather than utilizing profiles (or HMMs) exclusively for the target or template sequence, they are generated for both sequences and compared against each other (Fig. 2.4d). Each position within a sequence can be treated as a probability vector. In the case of standard profiles, a 20-dimensional probability vector is used (with one dimension for each amino acid residue). A position in the target sequence resembles a position in the template structure if both positions are subject to similar evolutionary pressures that yield comparable probability vectors. A variety of techniques have recently been developed to compare such vectors (the simplest being the dot product); almost all of them outperform simpler sequence-profile scoring methods (see, for example, Rychlewski et al. 2000; Ohlsen et al. 2004; Soeding 2005; Bennett-Lovsey et al. 2008).

In light of the success achieved by profile-profile methods, numerous research groups have modified their prediction pipelines to incorporate secondary structure profiles. Instead of simply predicting one of three discrete states (alpha-helix, beta-strand, or coil), the probability of each state is calculated and subsequently treated as a vector. Results have confirmed the superior performance of this approach (Tang et al. 2003; Bennett-Lovsey et al. 2008). A schematic Overview of this approach is shown in Fig. 2.4d.

The performance of profile-based prediction methods has steadily advanced due to the expansion and refinement of sequence databases, improvements in profile generation Procedures, and enhanced profile comparison algorithms. As this performance has grown, the relative importance of supplementary predicted structural properties appears to have diminished compared to their initially decisive role in the early methods of Bowie et al. (1991). The most successful secondary structure prediction methods are typically driven by Machine learning algorithms, such as artificial neural networks or support vector machines, trained on windows of PSI-BLAST-generated sequence profiles. The reason why incorporating this information yields only limited gains likely stems from a scarcity of novel, or "independent," data. The input data for secondary structure prediction generally consist of the exact same profiles used for sequence alignment. Therefore, it can be argued that much of the information used in secondary structure prediction is already implicitly encoded in the profile from which it was derived.



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.