Fundamentals of Bioinformatics - Ogurtsov A.N. 2013

Information Principles in Biotechnology
Bioinformatics in Pharmacy
Pharmacoinformatics

The term pharmacoinformatics is often used to describe the discipline that combines biology, chemistry, mathematics, and information technology for the purpose of data Processing and analysis in the pharmaceutical industry.

The application of high-throughput screening in drug discovery relies on the availability of diverse chemical libraries (such as those generated by combinatorial chemistry Methods), as they significantly enhance the Prospects of identifying molecules that interact with a specific protein target.

Quantitatively defining chemical diversity is a formidable challenge. Consequently, attempts have been made to address this issue using METABOLISM/2.html">THE CONCEPT OF "chemical space." Essentially, chemical space encompasses chemical compounds with all possible chemical properties concentrated within all potentially active molecular sites. Thus, a library with a high diversity index will provide broad coverage of chemical space, free of gaps and clusters of highly similar molecules.

The Diversity of chemical libraries is typically quantified using metrics based on the comparison of various molecular properties, described by parameters such as atomic arrangement and charge, as well as The ability to form Different types of chemical bonds.

To compare two molecules, one can use the Tanimoto coefficient KT, which reflects the degree of structural similarity between their fragments. The Tanimoto coefficient is calculated using the formula

Class="center">

where $a$ is the number of fragment parameters in compound A; $b$ is the number of fragment parameters in compound B, and $c$ is the number of shared (similar) fragment parameters between these compounds. Consequently, for identical molecules KT = 1, whereas for molecules with no overlapping parameters KT = 0. In a chemical library with an ideal diversity score, the majority of pairwise comparisons would yield a Tanimoto coefficient close to zero.

When little to no information is available regarding the binding Specificity of a target protein, highly diverse, comprehensive chemical libraries can facilitate productive lead discovery. Conversely, if specific data regarding the target's sequence or Structure has been gathered, one can filter general libraries to select focused subsets covering a particular region of chemical space.

For example, if The sequence of a target protein is known, database Homology searches will frequently yield a related protein with a previously resolved structure and documented small-molecule interactions.

In such cases, it is possible to design a chemical library containing a molecular scaffold that preserves the relative spatial arrangement of the features present in the known Ligand, while allowing for modifications through the attachment of diverse functional groups. It is often the case that certain functional groups have already been proven essential for drug binding. Such features are referred to as pharmacophores.

The term pharmacophore was coined by Paul Ehrlich in 1909. Ehrlich defined a pharmacophore as a molecular backbone that carries (-phore) the essential features responsible for the biological activity of a drug (pharmaco-).

In 1977, this definition was modified by Peter Gund: a pharmacophore is a set of structural features in a molecule that are recognized by biological receptors and are responsible for the molecule's biological activity.

The modern IUPAC (International Union of Pure and Applied Chemistry) definition states: a pharmacophore is an ensemble of steric and electronic features that is necessary to ensure optimal supramolecular interactions with a specific biological target and to trigger (or block) its biological response.

Pharmacophoric features are typically understood to mean pharmacophoric centers and the distance intervals between them that are required to manifest a given type of biological activity.

Typical pharmacophoric centers include hydrophobic regions, aromatic rings, Hydrogen bond Donors and acceptors, as well as anionic and cationic centers.

For a more detailed description of a pharmacophore, hydrophobic and excluded volumes are frequently utilized, along with permitted angular orientation tolerances for hydrogen bond vectors and aromatic ring planes.

Pharmacophore searching involves aligning a pharmacophore model with the characteristics of database molecules accommodated within their permissible Conformations.

An example of a pharmacophore is shown in Figure 85.

The Functional groups of the pharmacophore are depicted in Figure 85 as spheres of varying radii. Computer software employs different colors to represent spheres corresponding to hydrophobic and hydrophilic regions, or positively and negatively charged functional groups. Highly directional Hydrogen Bonds are represented by cones. The corresponding lead molecule is fitted into the pharmacophore model.

Pharmacophore screening significantly accelerates the lead discovery process. Before conducting in vitro screening assays, it is advantageous to gather as much information as possible regarding potential drug-target interactions.

One approach to obtaining such data is automated virtual screening of chemical Databases (matching against a molecular target of known structure).

Figure 85 - Pharmacophore model with a lead molecule

In other scenarios, The structure of a compound can be inferred through homology modeling based on a closely related resolved structure, or predicted using threading algorithms.

If the STRUCTURE OF THE target protein is known, fitness-scoring algorithms for potential interacting ligands are applied—these are the so-called "rational," structure-based computer-aided drug design methods.

First, the binding site (receptor) of a low-molecular-weight compound (drug) on the target protein is identified.

Next, molecular docking is employed to perform a molecular-graphics Analysis of the ligand–receptor complex, evaluating whether a potential ligand can fit into the binding cavity of the macromolecule's Active Site, as well as estimating the binding energy and affinity for such a complex. The design of novel ligands can also be accomplished by modifying the structure of the identified compounds.

To date, A large number of docking algorithms have been developed that attempt to fit small molecules into binding sites by analyzing information on spatial constraints and bond energies. The most popular molecular docking programs are available online:

AutoDock

http://autodock. scripps.edu/

FlexX

http://www.biosolveit.de/FlexX/index.html?ct=l

Hex

http://hex.loria.fr/

Dick Vision

http://dockvision.com/

eHiTS

http://www.simbiosys.ca/ehits/index.html

FRED

http://www.eyesopen.com/fred

GOLD

http://www.ccdc.cam.ac.uk/products/life_sciences/gold/

LIGPLOT

http://www.ebi.ac.uk/thornton-srv/software/LIGPLOT/

Pocket-Finder

http://www.modelling.leeds.ac.uk/pocketfmder/

Q-SiteFinder

http://www.modelling.leeds.ac.uk/qsitefmder/

SITUS

http://situs.biomachina.org/index.html

Molecular Doc king Web

http://mgl.scripps.edu/people/gmm/

Figure 86 shows an example of visualizing the docking results for a molecule of the cytostatic drug dactinomycin (actinomycin D, belonging to the antitumor antibiotic group, specifically the actinomycins), which is used for palliative care in certain types of Cancer to alleviate symptoms for the patient.

The docking program treats each potential ligand as a backbone with attached functional groups. First, the algorithm predicts possible docking poses by analyzing the distribution of Van der Waals spheres (restricted to those on the backbone), and then checks individual functional groups for steric compatibility using various combinations of bond rotations. Finally, the algorithm performs the docking and calculates the overall score of the complex.

Figure 86 – Docking of a dactinomycin molecule intercalated between DNA Base Pairs: a – front view; b – side view

Chemical databases can be searched not only for matches to a binding site (searching for interactions between complementary molecules) but also for similarity to a specific ligand. Several algorithms are available for comparing 2D or 3D structures and generating profiles of similar molecules.

Determining the 3D structure of a target (via X-ray crystallography or NMR spectroscopy) is a prerequisite for developing a lead compound that will either bind to it or modulate its activity.

Leads are selected from existing chemical libraries through combinatorial structure docking. Library lead candidates are sequentially docked into the Active Site of the molecular target (by evaluating various modes of complementary binding). This preliminary in silico "fitting" reduces the number of compounds that need to be synthesized and tested in vitro, since databases contain the necessary descriptions of chemical properties and synthetic routes required for simulation modeling.

Another lead discovery approach constructs molecules from a library of functional groups based on an analysis of the binding site on the target's surface. A specialized simulation algorithm thoroughly analyzes the active site of the molecular target and builds a lead molecule piece by piece from individual fragments.

The surface of the molecular target intended to interact with the lead may be adjacent to protein regions with distinct chemical properties, such as hydrophobic patches, hydrogen-bonding areas, or the catalytic active site. Fragments of the hypothetical compound are sequentially placed into these regions. Optimizing the orientation of these fragments helps determine the final Spatial Structure of the lead.

Sometimes an entire molecule fits into the receptor or active site all at once, and the docking program iterates through all possible ways to accommodate the ligand within the receptor pocket. The binding site in a receptor or enzyme molecule contains both hydrogen-bonding and hydrophobic regions.

Initially, the program positions and orients the prototype molecule within the active site in a way that maximizes the number of possible interactions.

It then sequentially adds and adjusts additional functional groups until all necessary bonds between the ligand and the target are formed.

The program simulates the spatial arrangement of the active site elements of the target and then searches databases for chemical structures that match this simulation model.

When experimental data on the 3D structure of a protein are unavailable, various comparative modeling methods are employed in computer-aided design. Functionally important regions within a protein molecule can be identified through comparative analysis of Amino acid sequences from homologous Proteins.

The BLAST program is used to search sequence and structural databases. When building a 3D model of a protein with a given Amino Acid Sequence, the polypeptide chain is first "mapped" onto the coordinates corresponding to The amino acid residues of a homologous protein with a known spatial structure, followed by internal energy Minimization to relieve any potential structural strain.

Molecular Dynamics methods are then used to simulate the motion of individual molecular parts to refine the conformation of flexible loops. The quality of the model is evaluated using a program that compares the positions of amino acid residues against statistical data derived from proteins whose 3D structures have been experimentally resolved.

When analyzing ligand-receptor interactions, the conformational mobility of the ligand molecule must be taken into account. For binding to occur, the ligand must adopt a conformation that is complementary to the Cell/13.html">Protein Structure. Static ligand-receptor models fail to account for conformational flexibility. The conformational space (the set of conformational variants) of the ligand is assessed using molecular dynamics simulations and energy minimization. Docking of various ligand conformations is performed across different positions, and the top-ranking poses are subsequently used for molecular dynamics simulations of the ligand-protein complex.

Simulation results reveal the states under which the receptor most frequently binds to specific conformational Variants of the ligand. Modeling the receptor structure in both the presence and absence of the ligand provides insight into how the protein changes its conformation upon activation triggered by ligand binding.

Dozens of drugs designed using bioinformatics approaches have already successfully passed clinical trials and entered medical practice.

The described drug discovery method—wherein lead compounds are optimized by adding various functional groups to a molecular scaffold and testing each derivative for biological activity—must also incorporate structure-activity analysis. Otherwise, if the modeled molecule possesses numerous open positions available for functionalization, the total number of molecules that would subsequently need to be evaluated in comprehensive screening assays becomes prohibitively large.

Synthesizing and testing all these molecules would demand significant time and experimental effort, even though it is evident that the vast majority of these candidate molecules possess no useful biological activity.

Quantitative Structure-Activity Relationship (QSAR) analysis methods, utilizing two-dimensional (2D) or three-dimensional (3D) ligand representations, Comparative Molecular Field Analysis (CoMFA), Comparative Molecular Similarity Indices Analysis (CoMSIA), and Hologram Quantitative Structure-Activity Relationship (HQSAR), provide spatial mapping of the ligand-binding site, pharmacophore modeling, and virtual screening of potential ligands against chemical databases.

QSAR evaluation enables the Selection of only those molecules with the highest probability of exhibiting beneficial activity, thereby reducing the number of candidate molecules required for subsequent chemical synthesis.

QSAR represents a mathematically expressed relationship that correlates molecular structure with biological activity.

Molecules are treated as sets of molecular properties (parameters) organized in a tabular format. The QSAR software analyzes these data to identify compatible relationships between individual parameters and biological Functions, thereby establishing a set of rules that can be used to score new molecules when assessing their potential activity. QSAR is typically expressed as a linear equation

where A is bioactivity; Pi are parameters (molecular properties) determined for each of the N molecules in the dataset; Ci are coefficients calculated by fitting molecular parameters to their biological functions.

Once a set of leads has been identified, the molecules must be optimized for potency, selectivity, and pharmacokinetic properties. High oral bioavailability (gastrointestinal absorption) is indicated by the presence of the following four characteristics:

1) number of hydrogen bond donors < 5;

2) number of hydrogen bond acceptors < 10;

3) molecular weight < 500;

4) lipophilicity < 5.

Furthermore, drugs targeting the Central Nervous system must possess adequate Blood-Brain barrier (BBB) permeability.



Last update: 11/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.