Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5606_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Foreword
- •Acknowledgments
- •Contents
- •1.1 Structure-Based Drug Discovery (SBDD)
- •1.2 Ligand-Based Drug Design (LBDD)
- •1.3 Echoes from the Past, Visions from the Future
- •References
- •1 Introduction
- •2.2 Second Step: Data Curation
- •2.4 Fourth Step: Updating and Maintenance
- •2 Databases and Curation
- •8 Perspectives
- •9 Conclusion
- •References
- •1 Introduction
- •2.1 Making and Matching Protein Models
- •2.2 Simulating Protein Movements
- •2.3 Analyzing Changes in Protein Shape
- •3 Pharmacogenomics in Drug Development
- •4 Case Studies of Genomics-Based Drug Design
- •References
- •1 Historical Background
- •1.1 Timeline
- •2 Methodology Overview
- •2.1 Neural Networks
- •2.1.1 Perceptron
- •2.1.2 Multilayer Neural Networks
- •2.1.3 Types of Neural Networks
- •Feedforward
- •Recurrent Neural Networks
- •LSTM
- •2.2 Deep Learning
- •3 Using Machine Learning
- •3.2 Data Collection
- •3.3 Data Preprocessing
- •3.4 Model Selection
- •3.5 Model Training
- •3.6 Validation
- •3.7 Tuning
- •3.8 Prediction
- •4 Limitations
- •4.1 Bias
- •4.3 Interpretability
- •4.4 Computational Cost
- •4.5 Data Dependency
- •4.6 Robustness
- •5 Applications in Drug Discovery
- •5.2 Lead Discovery
- •5.3 Preclinical and Clinical Development
- •6 Resources and Tools
- •7 Challenges and Perspectives
- •7.1 Future Trends
- •9 Conclusions
- •References
- •1 Historical Background
- •1.1 Applications in Drug Discovery
- •2 Validations and Controls
- •2.1 Internal Validation
- •2.2 External Validation
- •2.3 Relative Cluster Validation
- •3 Challenges and Perspectives
- •4 Conclusions
- •References
- •1 Historical Background
- •2 OECD Principles
- •2.1 A Defined Endpoint
- •2.2 An Unambiguous Algorithm
- •2.5 A Mechanistic Interpretation, if Possible
- •3 Software and Tools
- •4 Validations and Controls
- •4.1 Internal and External Validation
- •4.1.1 Regression Metrics
- •4.2 Applicability Domain
- •4.3 Randomization Tests
- •5 Interpretation
- •6 Practical Advice During QSAR Modeling
- •7 Application
- •8 Challenges and Perspectives
- •References
- •1 Molecular Docking
- •2 Advances in Scoring Functions and Search Algorithms
- •2.2 Critical Characteristics of Search Algorithms
- •2.3 Docking Programs and Scoring Functions
- •3 Calculations Performed During Docking Simulations
- •4 Essential Components for a Good Docking Program
- •5 Limitations of the Docking Technique
- •6 Validation of Docking Results
- •7 Inappropriate Use of Validation Methods in Docking
- •9 Use of Machine Learning in Molecular Docking
- •11 Challenges
- •12 Conclusions
- •References
- •3 System Preparation for MD Simulations
- •3.1 Solvation and Microensemble
- •3.2 Force Fields: General Concept and Relevant Choices
- •3.3 The Concept of Replicas and Timescale
- •4.1.2 Protein Root Mean Square Fluctuation (RMSF)
- •4.1.4 Protein Secondary Structure Analysis
- •4.1.5 Principal component Analysis (PCA)
- •4.1.6 Markov State Modelling
- •4.1.7 Distance Calculations
- •4.1.8 Angle and Plane Calculations
- •4.2.2 Distances and Ligand-Induced Geometry Rearrangements
- •4 Molecular Dynamics Analysis
- •4.1 Protein Perspective
- •4.1.1 Protein Root Mean Square Deviation (RMSD)
- •4.3 Ligand Perspective
- •4.3.1 Ligand Properties
- •4.3.2 Ligand Root Mean Square Deviation
- •4.3.3 Ligand Root Mean Square Fluctuation
- •4.3.4 Angles and Dihedrals
- •5.1 Protein Structure Prediction and Preparation
- •5.2 Molecular Docking
- •6 Concluding Remarks and Outlook
- •Glossary
- •References
- •1 Introduction
- •2.1 MDeNM
- •2.2 Collective Molecular Dynamics (coMD)
- •2.3 ClustENM and ClustENMD
- •3 Ensemble Docking
- •References
- •1 Introduction
- •1.1 Advantages, Disadvantages, Innovations, and Challenges
- •1.2 Recent Advances in Accessible FEP Software Tools
- •1.3 Applications of FEP in Industry and Consortiums
- •2 Expanding the Potential of FEP Calculations
- •2.1 Validating Binding Poses
- •2.2 Dealing with Solvent
- •2.3 FEP and Allostery
- •2.4 FEP and Covalent Ligands
- •2.5 Applications of FEP in Scaffold Hopping
- •2.6 Positional Analogue Scanning
- •2.7 Combinations and Alternative Approaches
- •3 Machine Learning for FEP
- •3.4 Implications for ML in FEP Calculations
- •4 Final Considerations
- •5 First Steps to FEP Simulations
- •References
- •1 Background
- •2 Ultra-Large Screening Libraries and Chemical Spaces
- •3.1 Implications of Dataset Size
- •4 Ligands on the Ultra-Large Scale
- •4.1 Ultra-Large 2D Similarity Searches
- •7 Challenges and Future Perspectives
- •7.1 Hit Triage: An Old Problem on a New Dimension
- •8 Conclusions
- •Appendix
- •References
- •1 Introduction
- •2 Enzymatic Activity Evaluations
- •3 Cytotoxicity Evaluation and Cell Viability
- •4 Antiviral Assays in Experimental Validation
- •6 In Vivo Evaluation of Compounds
- •7 Conclusions
- •References
- •1 Introduction
- •3.1 Data Collection
- •3.2 Data Preprocessing
- •3.4 Model Choice
- •3.5 Model Training
- •3.6 Model Assessment
- •3.7 External Validation
- •3.8 Implementation and Availability
- •3.9 Continuous Update
- •5 Conclusions and Perspectives
- •References
- •1 Experimental Approaches to Obtain Protein Structure
- •1.1 X-Ray Crystallography
- •1.2 Nuclear Magnetic Resonance
- •1.3 Cryo-EM
- •1.4 Hybrid Methods
- •2 Modeling Approaches to Obtain Protein Structure
- •2.1 Homology Modeling
- •2.2 Ab Initio Modeling
- •2.3 New Approaches
- •3 Conformational Diversity of Proteins
- •3.1 Characterization of Protein Conformational States
- •3.2 Experimental Methods to Study Protein Dynamics and Conformations
- •3.4 Molecular Dynamics Simulation
- •3.5 Sampling Strategies
- •4 Remarks and Perspectives
- •References
- •1 Introduction
- •2 Structure-Based Drug Design of HIV Protease Inhibitors
- •2.1 HIV-1 Protease as a Therapeutic Target
- •2.2.1 Saquinavir
- •2.2.2 Indinavir
- •2.3.1 Lopinavir
- •2.3.2 Darunavir
- •6 Conclusions
- •References
- •4 Experimental Methods to Analyze NR Activity
- •4.2 Coregulator-Recruitment
- •5 Concluding Remarks and Outlook
- •References

8 Drug Design in Motion: Concepts and Applications of Classical... 223
engagement in turn can lead to agonistic or antagonistic effects, depending on a
particular helix involved in ligand binding [115–119]. Moreover, distance analyses
can be used in drug resistance studies, where the distance between ligands and the
mutated residues can explain how mutations affect ligand binding [95, 120–122].
Moreover, distance analyses have led to the discovery of cryptic binding pockets,
often overlooked in static structures, which can be targeted to increase the potency
and selectivity of designed compounds [123]. Frequently, ligand binding at one site
influences the conformational rearrangements and interactions at a distant site on the
protein. By tracking distances between ligand-bound residues and residues at the
allosteric site, it is possible to quantify the extent of this communication and predict
how ligand binding triggers changes in distant regions. For instance, in nuclear
receptor studies distance calculations can be used to spot the changes in cofactor
binding or dimerisation interface alterations, induced by ligand binding [105, 124].
4.2.3 Predicting Protein-Ligand Binding Energies Through MD
Trajectory
There are severa l methods employed to predict protein-ligand free binding energies,
among those we highlight the free energy calculations, such as free energy perturbation (FEP [125]), relative binding energy prediction, thermal integration
(TI [126]), and the Born and surface area continuum solvation (MM/GBSA). In
general, the FEP and TI methods are among the most rigorous calculations of the
binding free energy, with the trade-off on their high time-consumption and the
limitations in the calculation of the relative free energy for highly flexible systems.
As a comprehensive guideline for protein-ligand binding energy prediction, please
consult the work from Ruiter and Oostenbrink [127]. We focus on the MM/GBSA
calculations, more details on FEP can be found in Chap. 10.
MM/GBSA Energy Calculations Molecular mechanics energies combined with
the generalised Born and surface area continuum solvation (MM/GBSA) allow
estimating binding free energy ΔG, which determines the strength of protein-ligand
binding [128] (Fig. 8.8a illustrates the protocol involved in the calculation). This
technique has been effectively utilised to explain experimental results as well as to
enhance virtual screening and docking outcomes [129]. This method is computationally more efficient than rigorous alchemical perturbation methods (e.g., free
energy perturbation, beyond the scope of this chapter), but still more robust than
simple molecular docking based on scoring functions. When derived from the MD
trajectory, MM/GBSA offers data on the interactions the ligand employs to anchor
itself inside the binding pocket and offers the advantage of better incorporating the
explicit solvent dynamics’ influence in the ligand binding, therefore accurately
estimating their entropic contribution and allowing the analyses of protein-watermediated interactions. These calculations may be used in the hit-to-lead approach to
guide synthetic chemists’ efforts toward improving the weakest feature of the ligand
or to indicate which moiety should be modified to boost binding affinity or to

224 E. Shevchenko et al.
Fig. 8.8 MM/GBSA calculations can help to rank and rationalise protein–ligand complexes. (a)
schematic representation of MM/GBSA calculations protocol and respective equation for binding
energy calculation. Ligand efficiency binding energy (normalising binding energy by the number of
heavy atoms). MM/GBSA calculations start with the ligand being extracted from the optimised
complex and an energy calculation is run on it without minimisation (see cartoon), to get the energy
of the ligand as optimised in the binding pocket. Next, energy minimisation is run on the ligand
outside of the receptor. Both calculations are being done with the ligand alone in the solution. The
energy difference is the ligand strain energy. The same procedure is performed for the receptor
counterpart, generating the receptor energy. The differences between those two calculations
generate ligand binding energy, which can be decomposed by residues or by properties. In all
violin plot graphics (b), the median of the calculated energies is displayed below, together with its
standard deviation, and free energy binding calculation (Kcal/mol normalised by the Heavy Atoms
Count, HAC), can be decomposed as the average per residue of the protein sequence (c), where one
can observe which residues contribute to the binding differences. (Adapted from Rashidian et al.
[137])

8 Drug Design in Motion: Concepts and Applications of Classical... 225
evaluate the bioaccumulation behaviours of non-tested compounds and support their
risk assessment. In computational chemistry, Binding free Energy (ΔG) is frequently
calculated per the thermodynamic cycle (Eq. 8.6)[130].
ΔG
bind,solv
= ΔG
bind,vacuum
þ ΔG
solv,complex
- ΔG
solv,ligand
þ ΔG
solv,receptor
ð8:6Þ
Binding free energy in computational approaches (see Fig. 8.8a for a schematic
representation). ΔG
only, as well as their complex) while ΔG
In turn, the Solvation free energy (ΔG
(ΔG
) and non-polar ( ΔG
pol
indicate calculations on solvent (for ligand- and receptor-
solv
refers to calculations without solvent.
vacuum
) is described as a sum of the polar
solv
) components. When MM/GBSA is applied on the
nonpol
MD trajector y, the calculations are performed for every simulation frame following
Eq. 8.7.
ΔG
bind
=<ΔG
complex
iðÞ- ΔG
Binding free energy ap plied for MD trajectory, where <...>
protein
iðÞ- ΔG
ligand
iðÞ>
i
indicates that
i
ð8:7Þ
calculations are conducted over i simulation frames.
Next, each ΔG in Eq. 8.7 is decomposed to the terms described in Eq. 8.8.
ΔG = E
þ E
int
þ E
ele
vdw
þ ΔG
þ ΔGnp- TΔS ð8:8Þ
pol
Decomposition of Binding free energy into physicochemical terms, where
T represents temperature and the ΔS the variation in entropy (see [128] for further
theory).
In Eq. 8.8, E
electrostatic energy, and E
this equation mean the following: The ΔG
stands for the molecular mechanic internal energy, E
int
for van der Waals energies. The last three terms of
vdw
and ΔGnpare the polar and nonpolar
pol
for
ele
Solvation free energies, T is the absolute temperature, and ΔS represents the entropy
estimate. While in MD, the protein is frequently solvated in a water box, before the
MM/GBSA computations, the water molecules are removed and replaced with an
implicit representation [131]. The generalised Born (GB, reflects term MM/GBSA)
or Poisson-Boltzmann (PB, reflects MM/PBSA) can be used as implicit representation [132]. For the GB model, the polar solvation energy is described with Equation
9[130].
ΔG
pol
≈ -
121
ij
ΔG
pol
tj
=
1
-
ϵ
ϵ
in
out
The polar solva tion energy for the GB model.
In Eq. 8.9, ε
constant of the solvent); β = ε
= 1 (dielectric constant of the solute) and ε
in
; α = 0.571412. The A corresponds to the
in/εout
molecule’s electrostatic dimensions. The f
1
1 þ βα
GB
functional form (Eq. 8.10) describes
ij
ij
qiq
1
þ
j
GB
f
ij
out
αβ
A
= 80 (dielectric
ð8:9Þ

226 E. Shevchenko et al.
the distance between atomic charges (rij) and effective Born radii (R), which in turn
indicates the depth with how each atom is buried in the solvent (Eq. 8.11, references
highlight the constant definitions [133, 134]).
1
2
2
r
GB
f
ij
= r
2
ij
þ RiRjexp -
ij
4RiR
j
ð8:10Þ
Distance-Dependent Electrostatic Solvation Function is used to estimate the
contribution of each atom pair’s electrostatic interaction to the overall energy of
the molecular system.
1
r - r
R
i
- 3
=-
4π
i
∙ dS ð8:11Þ
6
r - r
jj
i
Effective Born Radii Contour integral is limited to the dielectric boundary (∂V) of
the molecule, r and ridescribe the molecule’s position per the surface vector element
dS.
The ΔG
from Eq. 8.8 is estimated as proportional to the molecules’ solvent
np
accessible area (SASA) multiplied by the factor γ [135]. Once conducted,
MM/GBSA calculations generate plenty of energy properties. These properties,
which are broken down into contributions from diff erent components in the energy
expression, report energies for the ligand, receptor, and complex structures, as well
as energy changes related to strain and binding.
Below we will briefly introduce the most used properties that one can derive from
MM/GBSA.
Before diving into energy contributions, one important aspect of MM/GBSA
should be mentioned: the impact of ligand size on overall energy and its contributions. When comparing ligands with notable differences in molecular weight, one
may consider normalising the results by the Heavy Atom Count (HAC, i.e., number
of non-hydrogen atoms in the ligand molecule) to account for differences in molecular sizes. In general, normalising energy values by HAC also enhances the comparability of ligands’ binding affinities, irrespective of their sizes. This approach is
based on the assumption that the ligand–protein interactions scale linearly with the
number of heavy atoms in the ligand, which has been shown to hold reasonably well
for smaller ligands. This normalised binding energy, known as Ligand Efficacy, is a
frequently employed parameter in this context. However, HAC normalisation has
limitations and should not be the only criterion for assessing the binding affinity of a
ligand. Factors like shape and electrostatic complementarity should also be
considered.

8 Drug Design in Motion: Concepts and Applications of Classical... 227
While the absolute energy values derived from MM/GBSA calculations should
be interpreted with care, the focus should be on energy difference ratios when
comparing ligands or ligand-WT and ligand-mutated protein systems, bearing in
mind the relevance of the experimental design in including “control” simulations.
Comparing ΔG Binding Free Energy values with ligand efficacy values (normalised
by HAC) can reveal variations in absolute v alues across specific systems. These
variations originate from using different scales (ΔG in kcal/mol and ligand efficacy
in kcal/mol per heavy atom count) or disregarding molecular sizes (Fig. 8.8b).
A notable observation one can make is the ratios of the values between ligands
remain consistent across different systems. This consistency in the distribution trend
is noteworthy and provides an additional layer of validation when analysing
MM/GBSA results. Validity is best conveyed by presenting energy calculations as
a statist ical distribution (i.e., box plots), rather than single data points (i.e., mean/
average). The derived values from MM/GBSA fall into two major categories: ΔG
with its energy components and Ligand strain and its constituents. In this context,
ΔG normalised by HAC (ligand efficacy) belongs to the first category. The most
frequently derived energy components from MM/GBSA can be categorised as
follows (see Eqs. 8.7 and 8.8 for exemplary definitions):
• Coulombic Interaction Energy: Acco unts for the electrostatic interactions
between charged atoms. It includes both attractive (negative) and repulsive
(positive) contributions. Attractive interactions occur between opposite charges,
while repulsive interactions occur between like charges.
• Hydrogen Bond Interaction Energy: Quantifies the energy of hydrogen bonds
formed between the ligand and protein. These bonds involve a hydrogen atom
bonded to an electronegative atom and a nearby electronegative atom in the
protein, contributing favourably to binding energy.
• Packing Interaction Energy. This term considers the steric complementarity
between the ligand and the protein binding site. It accounts for favourable
interactions when hydrophobic or nonpolar groups in the ligand and protein
come into close contac t, creating a well-packed binding interface.
• Van der Waals (VdW) Interaction Energy: The van der Waals forces are attractive
forces that arise from fluctuations in electron distributions around atoms. This
term represents the energy associated with favourable interactions between non-
polar groups in the ligand and protein. It includes both attractive van der Waals
forces and repulsive steric clashes.
Next, moving to second category, associated with Ligand Strain Energy. Ligand
Strain Energy can provide valuable insights into the structural and energetic distortions that a ligand undergoes upon binding. It quantifies the energetic cost associated
with adapting or rearranging the ligand’s conformation to achieve optimal interactions within the binding pocket. In more detail, the Ligand Strain Energy accounts
for conformational changes (rotations, bending, or stretching of the ligand’s bonds),
steric interactions or clashes, distortion of ligand electronic structure (adjustments in
bond lengths, angles, and dihedrals), and hydrogen bonding and interaction
matching. In turn, Ligand Strain Energy is constituted by the sum of their own

228 E. Shevchenko et al.
different terms/components, namely Lipophilic, Coulomb, and Solvation energy.
While Coulomb and Solvation may often appear quite similar, they still provide
valuable insights into the ligand’s strain energy. Here’s how one can approach
reading the values.
• Ligand Strain Lipophilic Interaction Energy: represents the strain energy asso-
ciated with the ligand’s conformation in the context of lipophilic interactions. It
quantifies the distortion or deformation of the ligand molecule as it interacts with
lipophilic regions in the binding site. A higher value indicates a greater distortion
or deformation of the ligand molecule as it interacts with lipophilic regions in the
binding site. It highlights the strain caused by the ligand’s interaction with
hydrophobic regions of the protein.
• Ligand Strain Coulombic Interaction energy: refers to the Coulombic energy
term associated with the ligand in the binding site of a protein, quantifies the
electrostatic interactions between the ligand and the surrounding environment,
including the protein and solvent molecules. It accounts for the strain energy
arising from the repulsion or attraction of charged particles. While the Coulomb
values may appear similar to Ligand Strain Solvation Energy, they provide
information about the strain resulting specifically from electrostatic interactions.
• Ligand Strain Solvation Energy: refers to the strain energy of the ligand due to
solvation effects, quantifies the strain or distortion of the ligand’s molecular
structure caused by the presence of solvent molecules. This energy term captures
the energetic cost associated with reorganising the ligand’s conformation or
adjusting its shape to accommodate solvent molecules.
As mentioned, the predicted binding energy (using MM/GBSA) can rank ligands
more accurately than docking scores. The comparison between crystal structures/
complexes and short MD simulations showed that [136] though the timescale
impacted the predictions, the shorter evaluated simulations (~5 ns) have not led to
better predictions. This is consistent with the observation that larger conformational
changes required a higher number of representative frames to establish accurate
predictions. We propose to observe the distribution of predicted energy values
(Fig. 8.8b), as a control for this relationship between stable poses and predicted
energy, where monomodal distributions would indicate reliable binding modes. Our
recommendation is to compute the difference between the free energies of the
targeted states (ligands of interest) in comparison to references, multiple crystal
structures, and/or relevant models, as the quality of prediction can greatly vary
among the different targets.
Additionally, the predictions calculated in the article were sensitive to changes in
the solute dielectric constant, suggesting that this parameter should be carefully
determined according to the characteristics of the protein/ligand binding interface.
This concern could be dealt with by running simulations with known binders as
controls for determining an optimal setup. Additionally, the decomposition of
predicted energy among the individual contributing amino acid resi dues (Fig. 8.8c)
provides data that can be used for further hit-to-l ead optimisation or to interpret
conformational changes in the protein.

8 Drug Design in Motion: Concepts and Applications of Classical... 229
4.3 Ligand Perspective
4.3.1 Ligand Properties
Plenty of ligand properties or quantitative descriptors are routinely applied for
compound filtering before or after virtual screenings, and large-scale dockings can
be used for calculation on an MD trajectory. However, the number of properties that
are worth calculating as a function of time is quite limited and primarily used for
purposes in industrial drug development. The typical examples of such descriptors
are polar surface area (PSA) or solve nt-accessible surface area (SASA). PSA is used
to examine such parameters as oral absorption or blood–brain barrier
permeation [138].
4.3.2 Ligand Root Mean Squar e Deviation
The theoretical background of RMSD calculations is the same for both ligand and
protein, although the RMSD results are interpreted somewhat differently when a
ligand is involved. For instance, if RMSD values of the ligand can be aligned on the
ligand itself, the graph will reproduce the ligand’s internal fluctuations as a function
of time (Fig. 8.4a, red). When a shift of values is observed, the ligand’s conformation
changes during the trajectory (Fig. 8.4a, shift in RMSD indicated with a yellow
circle). Alternately, the ligand can be aligned on the protein, in which case the plot
will demonstrate how the ligand fluctuates concerning the protein. An expressive
shift in those values may indicate that the ligand is unstable inside the binding site or
moving outside of it. Additionally, RMSD calculations can be used to compare the
difference between a docking output to a known binding pose from a crystal
structure to validate the docking precision.
4.3.3 Ligand Root Mean Squar e Fluctuation
Ligand RMSF shares the theoretical background with the protein RMSF and indicates the atom fluctuations in the ligand. Similar to ligand RMSD, ligand RMSF can
be aligned either on the ligand itself or on the protein. The data in the first case
indicates ligand fluctuations, whereas, in the second, it depicts fluctuations in
correspondence to the binding site. The flexibility or conformational alterations of
a ligand inside a binding site are often evaluated using the ligand RMSF.
4.3.4 Angles an d Dihedrals
The analysis of ligand torsional angles and dihedrals can suppl ement the understanding of ligand–receptor interaction and binding dynamics. Torsional angles refer
to the rotations around a single bond, while dihedrals define the angle between two

230 E. Shevchenko et al.
planes defined by four atoms following one after the other in the ligand structure.
Practical applications of this analysis include identifying ligand key binding conformations, determining preferred binding modes, mapping the binding pocket, and
characterising ligand-induced conformational changes in the receptor.
Moreover, it can supplement the identification of specific torsional angles that are
responsible for favourable interactions with the target receptor, supporting the
rationalisation of compounds’ activity. Further, the discovered reliable torsional
angles can be integrated into a pharmacophore or employed directly as filters in
virtual screening campaigns to select compounds with a higher likelihood of binding
effectively. Furthermore, torsional angles provide information on compound flexibility that can assist medicinal chemists to design and synthesise analogues with
optimised torsional angles to enhance binding affinity and specificity of existing
compound series. Ligand angles analysis can be conducted with tools such as
CPPTRAJ [139] in Amber [57], MDAnalysis [110] in Python, GROMACS
[107, 140], or Simulation Interaction Diagram in Maestro [47]. These software
packages offer functionalities for calculating and visualising torsional angles and
dihedrals in different formats, such as plots, radar charts, or heatmaps, providing
valuable insights into ligand behaviour over time.
5 MD Simulations Together with Other Relevant SBDD
Techniques
5.1 Protein Structure Prediction and Preparation
This step has become an emerging issue in computational chemistry, especially
when refining low-resolution crystal structures, a task, which is further amplified by
the continuous growth of structural data. Factors like missing hydrogens, tautomer
uncertainties, and crystal packing complement the need for structure preparation.
Hydrogen bond (H-bond) networks play a critical role in shaping the quality and
stability of the future model, while also imp acting binding specificity studies. This is
especially relevant on X-ray structures, which often lack hydrogens
[141, 142]. Other factors, such as crystal packing can lead to distortions in protein
side-chain conformations, thereby influencing the interpretation of binding pockets
[143, 144]. Accurate protein preparation involves such essential components as
addressing missing hydrogens and optimising hydrogen bond networks. These
steps are integral to various structure preparation protocols offered by software
such as GROMACS [107], AutoDock [145], HADDOCK [146, 147], Chimera
[148]/Chimera X, and Protein Preparation Wizard [149] (Schrödinger LLC). Finally,
applying a chosen force field during the minimisation process ensures the accurate
refinement of a protein structure. By attentiv ely handling these aspects, reliable
protein-ligand models can be created, preventing misleading outcomes in subsequent studies.

8 Drug Design in Motion: Concepts and Applications of Classical... 231
In terms of structure prediction, for instance, homology modelling can be a
powerful technique for predicting the protein’s secondary structures based on the
sequence similarity to template structures. Combined with molecular dynamics, this
method can enable in silico mutagenesis experiments, studying conformational
changes, and allosteric effects.
The process involves model co nstruction, which is based on frequently several
experimentally established structures sharing common structural features with the
target prote in. If the sequence similarity between the target protein and existing
templates is low, one can rely on well-established common features within the
protein family of interest. For instance, class A GPCRs with unknown crystal
structures can still be modelled using the classical seven transmembrane helices as
a template [116].
Homology modelling can be divided into several steps, including template
selection, sequence alignment, model construction, and refinement, which finalise
generating a competent 3D structure of the protein query. The evaluation step is
crucial for model quality, and various software and techniques are described in the
respective chapter (see Chap. 14).
Since homology modelling can yield different results based on template choice, it
is important to construct multiple models using various templates and compare them
for accuracy. Several tools and online resources can facilitate homology model ling
tasks, with the latter being the latest artificial intelligence-based method for
predicting protein structures from amino acid sequences.
5.2 Molecular Docking
The docking methodology has been a pioneering computational tool in small
molecule drug design since the early 1980s [150]. It serves to predict the orientation
of ligands within a protein’s binding site. In general, docking involves several
stages, such as ligand and protein preparation, initial pose generation, and pose
scoring [151]. Scoring functions can be empirical, knowledge-based, or force fieldbased [150]. Diverse approaches can be found in molecular docking software. Some
require separate preparation of the protein structure and ligands in a relevant force
field (e.g., Glide [152, 153]), while others combine all stages under a unified protocol
(e.g., HADDOCK [146], AutoDock Vina [154, 155]). Regardless of the approach,
ligand preparation protocols typically include generat ing 3D geometry and
predicting tautomer and ionisation states. After ligand poses are generated, the
crucial step of results validation takes place.
Validation may involve including known active compounds in the docking set, as
they are expected to exhibit suitable binding modes and receive high docking scores
when the procedure works accurately. Additi onally, including known negatives or
decoys in the ligand dataset helps assess the model’s precision and reduces false
positives, avoiding artificial enrichment [156]. While the enrichment factor is valuable, it has limitations, notably its dependence on the number of known actives in the

232 E. Shevchenko et al.
initial dataset [156, 157]. An alternative approach is using ROC curves (receiver
operating characteristics) combined with the area under the curve (AUC) to evaluate
docking efficiency and determine the level of enrichment [156]. Nevertheless, ROC
curves’ restrictions [32] and trade-offs are known [158]. Each docking software
follows a unique protocol. One can rapidly screen a substantial compound library,
prioritising speed over quality with reduced conformations, just to have an additional
thorough step with just the high-ranking candid ates, aiming to balanc e resource.
Lastly, utilising derived docking poses as an initial frame for molecular dynamics
(MD) simulations allows to explore dynamic interactions, gaining valuable insights
into binding mechanisms and energy components. However, before conducting an
MD simulation, it is crucial to assess docking accuracy. This is a result of a myriad of
factors, relying on high-quality initial data and model, as well as carefully curating
information on the protein family, ligand conformations, and protonation states.
Additionally, prior data on active and decoy compounds should be considered. It is
also important to acknowledge the reliance on force fields and scoring functions.
Integrating docking and MD has proven successful in drug design, enabling the
identification of lead compounds, and designing potent inhibitors in various scientific disciplines [159].
We cite the work from Chen [160], which critically review pitfalls of docking
with further interesting insights in the MD simulation application. They propose that
protein–ligand complexes with stable interactions in simulations would lead to a
better representation of the binding modes. We also believe that the so-called
unstable systems can be equally informative in show ing protein–liga nd interactions
that can be changed. Relevant approaches directly integrating short explicit solvent
MD simulations during the docking pipeline [161], generating an ensemble of
protein conformations. This approach would, to some extent, account for the
ligand-induced conformational changes.
Alternatively, the selection of relevant conformations from MD-trajectories can
be performed the ability of each conformation to enrich known binder s from know n
non-binders [162], which is relevant for targets with established decoy sets (see
DUD-E database for readily available sets and the possibility to generate datasets
[163]). In this approach the structures/conformations would be ranked by their
discrimination power (using ROC and enrichment factors). It is relevant to notice
that the increased enrichment seems to be dependent on a limited number of MD
structures, when compared to crystal-exclusive docking [164, 165]. This alternative
to exclusively use experimentally determined structures, as was the state-of-art for
classical ensemble dockings, can be interesting in low-data scenarios. More details
on molecular docking and sampling the protein conformational space are available in
Chap. 9.
Соседние файлы в папке Библиотека им академика М.И. Перельмана
