Protein Structure and Function. Application of Bioinformatics Methods - John Rigden 2014
Prediction of Protein Function from Surface Properties
Ligand-Protein Interaction
Prediction of Active Site Location
Seemingly, the only truly universal feature of all protein-Ligand binding sites is the presence of pockets that house them. When one molecule is smaller than another, the simplest way to ensure extensive contact is for the pocket to enclose the ligand. Furthermore, in the case of Enzymes, this is also advantageous because the substrate is isolated from the solvent, thereby reducing the high reorganization energy associated with Reactions in Solution (Yadav et al. 1991). There are two main approaches for characterizing protein Structure/108.html">Surface Properties: geometric and energetic; these two approaches are described in the following sections.
Class="center">7.4.2.1. Geometric Definition of Ligand-Binding Sites
The core idea behind the geometric approach is that small molecules preferentially bind to the largest depression on a protein's surface (Laskowski et al. 1996). Various Methods exist for identifying these depressions (Laurie and Jackson 2006), some of which are discussed in this section. The simplest methods initially envelop the Cell/13.html">Protein Structure in a spatial grid. In the Pocket program (Levitt and Banaszak 1992), grid nodes are arranged along the x, y, and z axes, and depressions are defined as empty space surrounded by protein atoms. The LIGSITE program (Hendlich et al. 1997) makes this approach less sensitive to protein orientation by searching for depressions not only along the cubic grid axes, but also along the diagonals. This same technique was later implemented in the Pocket-Finder web server (Laurie and Jackson 2005). In the PASS program (Brady and Stouten 2000), grid nodes are placed at every point around the protein where they can Touch three protein atoms without intersecting them. These nodes cover virtually the entire protein surface and are selected based on the number of protein atoms within a specified distance from them—nodes located within depressions will have more neighboring protein atoms than those situated outside them. Such cycles of node placement and Selection ultimately lead to the Filling of the entire volume of the depressions.
In the SurfNet program (Laskowski 1995), spheres are placed between pairs of atoms in the protein, with the sphere diameter being reduced until overlap with other protein atoms is eliminated (which is not always possible, in which case the sphere is removed). The remaining spheres cluster within the protein cavities. These cavities can be visualized in real time for any PDB structure using the "Clefts" option in PDBsum (Laskowski et al. 2005). In the CASTp program (Binkowski et al. 2003), the outer atomic surface is defined using the so-called Delaunay triangulation. This is a geometric approach that assigns each atom in a molecule a polyhedron of the maximum possible volume. Faces are formed where atoms come into contact; if such a face cannot be constructed in a certain direction, it means the atom has no neighbors in that direction and is a surface atom. Connecting the centers of such atoms yields a surface that bounds a polyhedron with triangular faces. Some of these faces can be pruned using discrete flow theory (see Fig. 7.3), and the resulting central polyhedron defines a cavity within the protein.

Fig. 7.3. Schematic representation of discrete flow theory. A single Delaunay triangle acts as a sink for the flow. In CASTp, this is considered a true pocket
A common Procedure adopted in all these methods is defining the boundary between the pocket and the rest of the protein surface. Additionally, an extra boundary must be established between the pocket and the space surrounding the protein. In some geometric Definitions, these boundaries depend on the protein structure, and the predicted binding site volume tends to increase with protein size. When binding sites are defined using energetic characteristics (see below), their volume remains roughly equal to the ligand size regardless of the overall protein dimensions. This aligns with the assumption that the size of a ligand-binding pocket corresponds to the ligand itself rather than depending on the size of the protein (Laurie and Jackson 2005).
7.4.2.2. Energetic Definition of the Active Site
In addition to the geometric approach, an energetic approach can be used, in which pockets are defined as regions most favorable for interacting with other molecules (Laurie and Jackson 2005). In the Q-SiteFinder program, Proteins are first surrounded by a spatial grid, and a probe methyl (-CH3) group is placed into all grid nodes that do not intersect with the protein to evaluate interactions.

Fig. 7.4. (For the color version of this figure, see the insert.) Prediction of Ligand-binding sites using the Q-SiteFinder program. Shown is an example of the active site on the molecular surface of acetylcholinesterase (PDB code 1EVE), where the binding site for the small drug molecule—an inhibitor of this enzyme, aricept—is accurately identified by the top-ranked prediction (transparent gray surface)
The scoring function for the nodes is calculated as the Van der Waals potential energy of interaction for this probe group at each node. Grid nodes receiving a sufficiently high score are retained for subsequent clustering, and the resulting clusters are then ranked, assuming that pockets with the most favorable energy (typically the highest) represent the ligand-binding sites. It was thus found that 90% of the examined real binding sites appeared among the top three predicted probable pockets (see Fig. 7.4).
7.4.2.3. Theoretical Microscopic Titration Curves
The Electrostatic Properties of a protein surface influence The behavior of ionizable groups in its vicinity. In many cases, ionizable side chains closely involved in chemical catalysis experience an environment that significantly affects their ionization state. For instance, they frequently prove capable of maintaining a specific partially protonated state over an unusually wide pH range. As a result, the resulting pKa values and titration curves deviate in magnitude and shape from the average values for those same residues. Because the electrostatic properties and ionization states of a protein can be computed, catalytic center residues can be predicted based on their anomalous theoretical microscopic titration curves (e.g., Antosiewicz et al. 1994; Elcock, 2001). The application of this method, known as THEMATICS (theoretical microscopic titration curves), has successfully identified catalytic center residues in seven different enzyme structures with a low false-positive prediction rate, while showing no catalytic activity in non-catalytic control proteins. Since then, the method has been improved by developing statistical descriptors for classifying theoretical titration curves (Ko et al. 2005) and incorporating support vector machines to enhance prediction sensitivity (Tong et al. 2008). More recently, it has been demonstrated that the method's performance slightly decreases when analyzing apo-structures instead of holo-structures (which exhibit substantial conformational differences) (Murga et al. 2008). While a key strength of this method is its independence from the presence or absence of homologous sequences, it must be emphasized that it is only suitable for predicting catalytic sites rather than binding sites in general.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.