Protein Structure and Function: Application of Bioinformatics Methods - John Rigden 2014
Fold Recognition
Detection of Remote Homology without Alignment
Homology Network Traversal
As we saw with PSI-BLAST and intermediate sequence searches, combining a series of relationships among homologs can yield highly productive search Methods. Recent studies have initiated the exploration of this network of relationships at an even higher level of detail. Profile-based approaches attempt to create a single statistical representation for a set of related Proteins—a kind of "averaged" model. Such approaches, however, discard much of the information available within this relationship network. Bateman and Finn (2007) employed a straightforward approach to recover a portion of this information. Their method compares the results of two independent profile search Procedures and asks whether the number of sequences shared by both procedures is greater than expected by chance. If the sequences in question are closely related, their profiles will share A large number of common sequences. Otherwise, the sequences retrieved by their profiles will exhibit only random similarity. This approach is analogous to examining the first-order Structure of a homolog network, i.e., comparing the neighbors of one sequence with the neighbors of another. This simple approach has proven highly effective in detecting Homology (without generating alignments) and significantly outperforms modern profile-comparison methods.
Weston et al. (2004), in their Rankprop algorithm, made deeper use of the global STRUCTURE OF THE homolog network. A key innovation behind the success of the Google search engine is its ability to exploit global structure by inferring it from the local hyperlink topology of the network. Google's PageRank algorithm models The behavior of a random web surfer who clicks on subsequent links at random and occasionally jumps to a completely random page. Web pages are then ranked According to the probability distribution of these random walks. Initially, the Rankprop algorithm utilizes a protein sequence similarity network pre-computed from the entire sequence database. Analogous to a diffusion process, the protein of interest is introduced into the network, after which linkage information (representing connections between this protein sequence and the closely related sequences of other proteins) propagates through the network to immediate neighbors, neighbors of neighbors, and so on. The database proteins are subsequently ranked based on the volume of links they receive originating from the target protein. This approach has been shown to outperform standard sequence-profile search methods and is comparable to profile-profile search methods, despite relying on PSI-BLAST for the initial construction of the similarity network.
Finally, Heger and colleagues (2008) developed the Maxflow algorithm, capable of traversing large homolog networks at the individual residue level. The algorithm searches for consistently aligned pairs of residues within a network of pairwise alignments. What sets this method apart from others is its explicit focus on generating alignments, which is crucial for protein modeling.
All of these novel network-based approaches represent valuable developments for homology detection. A major drawback of this Class of methods is the immense computational power required to construct all-against-all protein similarity networks. It is evident that the performance of these methods will further increase if truly comprehensive networks are built using modern Databases containing approximately 6 million sequences. However, reduced databases containing only sequences with less than 50% identity—and thus being significantly smaller—have been shown in studies to yield performance equal to, if not higher than, complete databases. It is also worth noting that research in homology recognition is likely to receive a significant boost in the near future, driven by emerging techniques rooted in graph theory.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.