Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5364_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
02.09.2026
Размер:
21 Мб
Скачать
8 Drug Design in Motion: Concepts and Applications of Classical... 223
engagement in turn can lead to agonistic or antagonistic effects, depending on a particular helix involved in ligand binding [115119]. Moreover, distance analyses can be used in drug resistance studies, where the distance between ligands and the mutated residues can explain how mutations affect ligand binding [95, 120122].
Moreover, distance analyses have led to the discovery of cryptic binding pockets, often overlooked in static structures, which can be targeted to increase the potency and selectivity of designed compounds [123]. Frequently, ligand binding at one site inuences the conformational rearrangements and interactions at a distant site on the protein. By tracking distances between ligand-bound residues and residues at the allosteric site, it is possible to quantify the extent of this communication and predict how ligand binding triggers changes in distant regions. For instance, in nuclear receptor studies distance calculations can be used to spot the changes in cofactor binding or dimerisation interface alterations, induced by ligand binding [105, 124].
4.2.3 Predicting Protein-Ligand Binding Energies Through MD
Trajectory
There are severa l methods employed to predict protein-ligand free binding energies, among those we highlight the free energy calculations, such as free energy pertur­bation (FEP [125]), relative binding energy prediction, thermal integration (TI [126]), and the Born and surface area continuum solvation (MM/GBSA). In general, the FEP and TI methods are among the most rigorous calculations of the binding free energy, with the trade-off on their high time-consumption and the limitations in the calculation of the relative free energy for highly exible systems. As a comprehensive guideline for protein-ligand binding energy prediction, please consult the work from Ruiter and Oostenbrink [127]. We focus on the MM/GBSA calculations, more details on FEP can be found in Chap. 10.
MM/GBSA Energy Calculations Molecular mechanics energies combined with the generalised Born and surface area continuum solvation (MM/GBSA) allow estimating binding free energy ΔG, which determines the strength of protein-ligand binding [128] (Fig. 8.8a illustrates the protocol involved in the calculation). This technique has been effectively utilised to explain experimental results as well as to enhance virtual screening and docking outcomes [129]. This method is computa­tionally more efcient than rigorous alchemical perturbation methods (e.g., free energy perturbation, beyond the scope of this chapter), but still more robust than simple molecular docking based on scoring functions. When derived from the MD trajectory, MM/GBSA offers data on the interactions the ligand employs to anchor itself inside the binding pocket and offers the advantage of better incorporating the explicit solvent dynamicsinuence in the ligand binding, therefore accurately estimating their entropic contribution and allowing the analyses of protein-water­mediated interactions. These calculations may be used in the hit-to-lead approach to guide synthetic chemistsefforts toward improving the weakest feature of the ligand or to indicate which moiety should be modied to boost binding afnity or to
224 E. Shevchenko et al.
Fig. 8.8 MM/GBSA calculations can help to rank and rationalise protein–ligand complexes. (a) schematic representation of MM/GBSA calculations protocol and respective equation for binding energy calculation. Ligand efciency binding energy (normalising binding energy by the number of heavy atoms). MM/GBSA calculations start with the ligand being extracted from the optimised complex and an energy calculation is run on it without minimisation (see cartoon), to get the energy of the ligand as optimised in the binding pocket. Next, energy minimisation is run on the ligand outside of the receptor. Both calculations are being done with the ligand alone in the solution. The energy difference is the ligand strain energy. The same procedure is performed for the receptor counterpart, generating the receptor energy. The differences between those two calculations generate ligand binding energy, which can be decomposed by residues or by properties. In all violin plot graphics (b), the median of the calculated energies is displayed below, together with its standard deviation, and free energy binding calculation (Kcal/mol normalised by the Heavy Atoms Count, HAC), can be decomposed as the average per residue of the protein sequence (c), where one can observe which residues contribute to the binding differences. (Adapted from Rashidian et al. [137])
8 Drug Design in Motion: Concepts and Applications of Classical... 225
evaluate the bioaccumulation behaviours of non-tested compounds and support their risk assessment. In computational chemistry, Binding free Energy (ΔG) is frequently calculated per the thermodynamic cycle (Eq. 8.6)[130].
ΔG
bind,solv
= ΔG
bind,vacuum
þ ΔG
solv,complex
- ΔG
solv,ligand
þ ΔG
solv,receptor
ð8:6Þ
Binding free energy in computational approaches (see Fig. 8.8a for a schematic representation). ΔG only, as well as their complex) while ΔG
In turn, the Solvation free energy (ΔG (ΔG
) and non-polar ( ΔG
pol
indicate calculations on solvent (for ligand- and receptor-
solv
refers to calculations without solvent.
vacuum
) is described as a sum of the polar
solv
) components. When MM/GBSA is applied on the
nonpol
MD trajector y, the calculations are performed for every simulation frame following Eq. 8.7.
ΔG
bind
=<ΔG
complex
iðÞ- ΔG
Binding free energy ap plied for MD trajectory, where <...>
protein
iðÞ- ΔG
ligand
iðÞ>
i
indicates that
i
ð8:7Þ
calculations are conducted over i simulation frames.
Next, each ΔG in Eq. 8.7 is decomposed to the terms described in Eq. 8.8.
ΔG = E
þ E
int
þ E
ele
vdw
þ ΔG
þ ΔGnp- TΔS ð8:8Þ
pol
Decomposition of Binding free energy into physicochemical terms, where T represents temperature and the ΔS the variation in entropy (see [128] for further theory).
In Eq. 8.8, E electrostatic energy, and E this equation mean the following: The ΔG
stands for the molecular mechanic internal energy, E
int
for van der Waals energies. The last three terms of
vdw
and ΔGnpare the polar and nonpolar
pol
for
ele
Solvation free energies, T is the absolute temperature, and ΔS represents the entropy estimate. While in MD, the protein is frequently solvated in a water box, before the MM/GBSA computations, the water molecules are removed and replaced with an implicit representation [131]. The generalised Born (GB, reects term MM/GBSA) or Poisson-Boltzmann (PB, reects MM/PBSA) can be used as implicit representa­tion [132]. For the GB model, the polar solvation energy is described with Equation 9[130].
ΔG
pol
≈ -
121
ij
ΔG
pol
tj
=
1
-
ϵ
ϵ
in
out
The polar solva tion energy for the GB model.
In Eq. 8.9, ε constant of the solvent); β = ε
= 1 (dielectric constant of the solute) and ε
in
; α = 0.571412. The A corresponds to the
in/εout
molecules electrostatic dimensions. The f
1
1 þ βα
GB
functional form (Eq. 8.10) describes
ij
ij
qiq
1
þ
j
GB
f
ij
out
αβ
A
= 80 (dielectric
ð8:9Þ
226 E. Shevchenko et al.
the distance between atomic charges (rij) and effective Born radii (R), which in turn indicates the depth with how each atom is buried in the solvent (Eq. 8.11, references highlight the constant denitions [133, 134]).
1 2
2
r
GB
f
ij
= r
2
ij
þ RiRjexp -
ij
4RiR
j
ð8:10Þ
Distance-Dependent Electrostatic Solvation Function is used to estimate the contribution of each atom pairs electrostatic interaction to the overall energy of the molecular system.
1
r - r
R
i
- 3
=-
4π
i
dS ð8:11Þ
6
r - r
jj
i
Effective Born Radii Contour integral is limited to the dielectric boundary (V) of the molecule, r and ridescribe the molecules position per the surface vector element
dS.
The ΔG
from Eq. 8.8 is estimated as proportional to the moleculessolvent
np
accessible area (SASA) multiplied by the factor γ [135]. Once conducted, MM/GBSA calculations generate plenty of energy properties. These properties, which are broken down into contributions from diff erent components in the energy expression, report energies for the ligand, receptor, and complex structures, as well as energy changes related to strain and binding.
Below we will briey introduce the most used properties that one can derive from MM/GBSA.
Before diving into energy contributions, one important aspect of MM/GBSA should be mentioned: the impact of ligand size on overall energy and its contribu­tions. When comparing ligands with notable differences in molecular weight, one may consider normalising the results by the Heavy Atom Count (HAC, i.e., number of non-hydrogen atoms in the ligand molecule) to account for differences in molec­ular sizes. In general, normalising energy values by HAC also enhances the compa­rability of ligandsbinding afnities, irrespective of their sizes. This approach is based on the assumption that the ligand–protein interactions scale linearly with the number of heavy atoms in the ligand, which has been shown to hold reasonably well for smaller ligands. This normalised binding energy, known as Ligand Efcacy, is a frequently employed parameter in this context. However, HAC normalisation has limitations and should not be the only criterion for assessing the binding afnity of a ligand. Factors like shape and electrostatic complementarity should also be considered.
8 Drug Design in Motion: Concepts and Applications of Classical... 227
While the absolute energy values derived from MM/GBSA calculations should be interpreted with care, the focus should be on energy difference ratios when comparing ligands or ligand-WT and ligand-mutated protein systems, bearing in mind the relevance of the experimental design in including controlsimulations. Comparing ΔG Binding Free Energy values with ligand efcacy values (normalised by HAC) can reveal variations in absolute v alues across specic systems. These variations originate from using different scales (ΔG in kcal/mol and ligand efcacy in kcal/mol per heavy atom count) or disregarding molecular sizes (Fig. 8.8b).
A notable observation one can make is the ratios of the values between ligands remain consistent across different systems. This consistency in the distribution trend is noteworthy and provides an additional layer of validation when analysing MM/GBSA results. Validity is best conveyed by presenting energy calculations as a statist ical distribution (i.e., box plots), rather than single data points (i.e., mean/ average). The derived values from MM/GBSA fall into two major categories: ΔG with its energy components and Ligand strain and its constituents. In this context, ΔG normalised by HAC (ligand efcacy) belongs to the rst category. The most frequently derived energy components from MM/GBSA can be categorised as follows (see Eqs. 8.7 and 8.8 for exemplary denitions):
Coulombic Interaction Energy: Acco unts for the electrostatic interactions
between charged atoms. It includes both attractive (negative) and repulsive
(positive) contributions. Attractive interactions occur between opposite charges,
while repulsive interactions occur between like charges.
Hydrogen Bond Interaction Energy: Quanties the energy of hydrogen bonds
formed between the ligand and protein. These bonds involve a hydrogen atom
bonded to an electronegative atom and a nearby electronegative atom in the
protein, contributing favourably to binding energy.
Packing Interaction Energy. This term considers the steric complementarity
between the ligand and the protein binding site. It accounts for favourable
interactions when hydrophobic or nonpolar groups in the ligand and protein
come into close contac t, creating a well-packed binding interface.
Van der Waals (VdW) Interaction Energy: The van der Waals forces are attractive
forces that arise from uctuations in electron distributions around atoms. This
term represents the energy associated with favourable interactions between non-
polar groups in the ligand and protein. It includes both attractive van der Waals
forces and repulsive steric clashes.
Next, moving to second category, associated with Ligand Strain Energy. Ligand Strain Energy can provide valuable insights into the structural and energetic distor­tions that a ligand undergoes upon binding. It quanties the energetic cost associated with adapting or rearranging the ligands conformation to achieve optimal interac­tions within the binding pocket. In more detail, the Ligand Strain Energy accounts for conformational changes (rotations, bending, or stretching of the ligands bonds), steric interactions or clashes, distortion of ligand electronic structure (adjustments in bond lengths, angles, and dihedrals), and hydrogen bonding and interaction matching. In turn, Ligand Strain Energy is constituted by the sum of their own
228 E. Shevchenko et al.
different terms/components, namely Lipophilic, Coulomb, and Solvation energy. While Coulomb and Solvation may often appear quite similar, they still provide valuable insights into the ligands strain energy. Heres how one can approach reading the values.
Ligand Strain Lipophilic Interaction Energy: represents the strain energy asso-
ciated with the ligands conformation in the context of lipophilic interactions. It
quanties the distortion or deformation of the ligand molecule as it interacts with
lipophilic regions in the binding site. A higher value indicates a greater distortion
or deformation of the ligand molecule as it interacts with lipophilic regions in the
binding site. It highlights the strain caused by the ligands interaction with
hydrophobic regions of the protein.
Ligand Strain Coulombic Interaction energy: refers to the Coulombic energy
term associated with the ligand in the binding site of a protein, quanties the
electrostatic interactions between the ligand and the surrounding environment,
including the protein and solvent molecules. It accounts for the strain energy
arising from the repulsion or attraction of charged particles. While the Coulomb
values may appear similar to Ligand Strain Solvation Energy, they provide
information about the strain resulting specically from electrostatic interactions.
Ligand Strain Solvation Energy: refers to the strain energy of the ligand due to
solvation effects, quanties the strain or distortion of the ligands molecular
structure caused by the presence of solvent molecules. This energy term captures
the energetic cost associated with reorganising the ligands conformation or
adjusting its shape to accommodate solvent molecules.
As mentioned, the predicted binding energy (using MM/GBSA) can rank ligands more accurately than docking scores. The comparison between crystal structures/ complexes and short MD simulations showed that [136] though the timescale impacted the predictions, the shorter evaluated simulations (~5 ns) have not led to better predictions. This is consistent with the observation that larger conformational changes required a higher number of representative frames to establish accurate predictions. We propose to observe the distribution of predicted energy values (Fig. 8.8b), as a control for this relationship between stable poses and predicted energy, where monomodal distributions would indicate reliable binding modes. Our recommendation is to compute the difference between the free energies of the targeted states (ligands of interest) in comparison to references, multiple crystal structures, and/or relevant models, as the quality of prediction can greatly vary among the different targets.
Additionally, the predictions calculated in the article were sensitive to changes in the solute dielectric constant, suggesting that this parameter should be carefully determined according to the characteristics of the protein/ligand binding interface. This concern could be dealt with by running simulations with known binders as controls for determining an optimal setup. Additionally, the decomposition of predicted energy among the individual contributing amino acid resi dues (Fig. 8.8c) provides data that can be used for further hit-to-l ead optimisation or to interpret conformational changes in the protein.
8 Drug Design in Motion: Concepts and Applications of Classical... 229
4.3 Ligand Perspective
4.3.1 Ligand Properties
Plenty of ligand properties or quantitative descriptors are routinely applied for compound ltering before or after virtual screenings, and large-scale dockings can be used for calculation on an MD trajectory. However, the number of properties that are worth calculating as a function of time is quite limited and primarily used for purposes in industrial drug development. The typical examples of such descriptors are polar surface area (PSA) or solve nt-accessible surface area (SASA). PSA is used to examine such parameters as oral absorption or blood–brain barrier permeation [138].
4.3.2 Ligand Root Mean Squar e Deviation
The theoretical background of RMSD calculations is the same for both ligand and protein, although the RMSD results are interpreted somewhat differently when a ligand is involved. For instance, if RMSD values of the ligand can be aligned on the ligand itself, the graph will reproduce the ligands internal uctuations as a function of time (Fig. 8.4a, red). When a shift of values is observed, the ligands conformation changes during the trajectory (Fig. 8.4a, shift in RMSD indicated with a yellow circle). Alternately, the ligand can be aligned on the protein, in which case the plot will demonstrate how the ligand uctuates concerning the protein. An expressive shift in those values may indicate that the ligand is unstable inside the binding site or moving outside of it. Additionally, RMSD calculations can be used to compare the difference between a docking output to a known binding pose from a crystal structure to validate the docking precision.
4.3.3 Ligand Root Mean Squar e Fluctuation
Ligand RMSF shares the theoretical background with the protein RMSF and indi­cates the atom uctuations in the ligand. Similar to ligand RMSD, ligand RMSF can be aligned either on the ligand itself or on the protein. The data in the rst case indicates ligand uctuations, whereas, in the second, it depicts uctuations in correspondence to the binding site. The exibility or conformational alterations of a ligand inside a binding site are often evaluated using the ligand RMSF.
4.3.4 Angles an d Dihedrals
The analysis of ligand torsional angles and dihedrals can suppl ement the under­standing of ligand–receptor interaction and binding dynamics. Torsional angles refer to the rotations around a single bond, while dihedrals dene the angle between two
230 E. Shevchenko et al.
planes dened by four atoms following one after the other in the ligand structure. Practical applications of this analysis include identifying ligand key binding confor­mations, determining preferred binding modes, mapping the binding pocket, and characterising ligand-induced conformational changes in the receptor.
Moreover, it can supplement the identication of specic torsional angles that are responsible for favourable interactions with the target receptor, supporting the rationalisation of compoundsactivity. Further, the discovered reliable torsional angles can be integrated into a pharmacophore or employed directly as lters in virtual screening campaigns to select compounds with a higher likelihood of binding effectively. Furthermore, torsional angles provide information on compound exi­bility that can assist medicinal chemists to design and synthesise analogues with optimised torsional angles to enhance binding afnity and specicity of existing compound series. Ligand angles analysis can be conducted with tools such as CPPTRAJ [139] in Amber [57], MDAnalysis [110] in Python, GROMACS [107, 140], or Simulation Interaction Diagram in Maestro [47]. These software packages offer functionalities for calculating and visualising torsional angles and dihedrals in different formats, such as plots, radar charts, or heatmaps, providing valuable insights into ligand behaviour over time.
5 MD Simulations Together with Other Relevant SBDD
Techniques
5.1 Protein Structure Prediction and Preparation
This step has become an emerging issue in computational chemistry, especially when rening low-resolution crystal structures, a task, which is further amplied by the continuous growth of structural data. Factors like missing hydrogens, tautomer uncertainties, and crystal packing complement the need for structure preparation. Hydrogen bond (H-bond) networks play a critical role in shaping the quality and stability of the future model, while also imp acting binding specicity studies. This is especially relevant on X-ray structures, which often lack hydrogens [141, 142]. Other factors, such as crystal packing can lead to distortions in protein side-chain conformations, thereby inuencing the interpretation of binding pockets [143, 144]. Accurate protein preparation involves such essential components as addressing missing hydrogens and optimising hydrogen bond networks. These steps are integral to various structure preparation protocols offered by software such as GROMACS [107], AutoDock [145], HADDOCK [146, 147], Chimera [148]/Chimera X, and Protein Preparation Wizard [149] (Schrödinger LLC). Finally, applying a chosen force eld during the minimisation process ensures the accurate renement of a protein structure. By attentiv ely handling these aspects, reliable protein-ligand models can be created, preventing misleading outcomes in subse­quent studies.
8 Drug Design in Motion: Concepts and Applications of Classical... 231
In terms of structure prediction, for instance, homology modelling can be a powerful technique for predicting the proteins secondary structures based on the sequence similarity to template structures. Combined with molecular dynamics, this method can enable in silico mutagenesis experiments, studying conformational changes, and allosteric effects.
The process involves model co nstruction, which is based on frequently several experimentally established structures sharing common structural features with the target prote in. If the sequence similarity between the target protein and existing templates is low, one can rely on well-established common features within the protein family of interest. For instance, class A GPCRs with unknown crystal structures can still be modelled using the classical seven transmembrane helices as a template [116].
Homology modelling can be divided into several steps, including template selection, sequence alignment, model construction, and renement, which nalise generating a competent 3D structure of the protein query. The evaluation step is crucial for model quality, and various software and techniques are described in the respective chapter (see Chap. 14).
Since homology modelling can yield different results based on template choice, it is important to construct multiple models using various templates and compare them for accuracy. Several tools and online resources can facilitate homology model ling tasks, with the latter being the latest articial intelligence-based method for predicting protein structures from amino acid sequences.
5.2 Molecular Docking
The docking methodology has been a pioneering computational tool in small molecule drug design since the early 1980s [150]. It serves to predict the orientation of ligands within a proteins binding site. In general, docking involves several stages, such as ligand and protein preparation, initial pose generation, and pose scoring [151]. Scoring functions can be empirical, knowledge-based, or force eld­based [150]. Diverse approaches can be found in molecular docking software. Some require separate preparation of the protein structure and ligands in a relevant force eld (e.g., Glide [152, 153]), while others combine all stages under a unied protocol (e.g., HADDOCK [146], AutoDock Vina [154, 155]). Regardless of the approach, ligand preparation protocols typically include generat ing 3D geometry and predicting tautomer and ionisation states. After ligand poses are generated, the crucial step of results validation takes place.
Validation may involve including known active compounds in the docking set, as they are expected to exhibit suitable binding modes and receive high docking scores when the procedure works accurately. Additi onally, including known negatives or decoys in the ligand dataset helps assess the models precision and reduces false positives, avoiding articial enrichment [156]. While the enrichment factor is valu­able, it has limitations, notably its dependence on the number of known actives in the
232 E. Shevchenko et al.
initial dataset [156, 157]. An alternative approach is using ROC curves (receiver operating characteristics) combined with the area under the curve (AUC) to evaluate docking efciency and determine the level of enrichment [156]. Nevertheless, ROC curvesrestrictions [32] and trade-offs are known [158]. Each docking software follows a unique protocol. One can rapidly screen a substantial compound library, prioritising speed over quality with reduced conformations, just to have an additional thorough step with just the high-ranking candid ates, aiming to balanc e resource.
Lastly, utilising derived docking poses as an initial frame for molecular dynamics (MD) simulations allows to explore dynamic interactions, gaining valuable insights into binding mechanisms and energy components. However, before conducting an MD simulation, it is crucial to assess docking accuracy. This is a result of a myriad of factors, relying on high-quality initial data and model, as well as carefully curating information on the protein family, ligand conformations, and protonation states. Additionally, prior data on active and decoy compounds should be considered. It is also important to acknowledge the reliance on force elds and scoring functions. Integrating docking and MD has proven successful in drug design, enabling the identication of lead compounds, and designing potent inhibitors in various scien­tic disciplines [159].
We cite the work from Chen [160], which critically review pitfalls of docking with further interesting insights in the MD simulation application. They propose that protein–ligand complexes with stable interactions in simulations would lead to a better representation of the binding modes. We also believe that the so-called unstable systems can be equally informative in show ing protein–liga nd interactions that can be changed. Relevant approaches directly integrating short explicit solvent MD simulations during the docking pipeline [161], generating an ensemble of protein conformations. This approach would, to some extent, account for the ligand-induced conformational changes.
Alternatively, the selection of relevant conformations from MD-trajectories can be performed the ability of each conformation to enrich known binder s from know n non-binders [162], which is relevant for targets with established decoy sets (see DUD-E database for readily available sets and the possibility to generate datasets [163]). In this approach the structures/conformations would be ranked by their discrimination power (using ROC and enrichment factors). It is relevant to notice that the increased enrichment seems to be dependent on a limited number of MD structures, when compared to crystal-exclusive docking [164, 165]. This alternative to exclusively use experimentally determined structures, as was the state-of-art for classical ensemble dockings, can be interesting in low-data scenarios. More details on molecular docking and sampling the protein conformational space are available in Chap. 9.