Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5364_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Foreword
- •Acknowledgments
- •Contents
- •1.1 Structure-Based Drug Discovery (SBDD)
- •1.2 Ligand-Based Drug Design (LBDD)
- •1.3 Echoes from the Past, Visions from the Future
- •References
- •1 Introduction
- •2.2 Second Step: Data Curation
- •2.4 Fourth Step: Updating and Maintenance
- •2 Databases and Curation
- •8 Perspectives
- •9 Conclusion
- •References
- •1 Introduction
- •2.1 Making and Matching Protein Models
- •2.2 Simulating Protein Movements
- •2.3 Analyzing Changes in Protein Shape
- •3 Pharmacogenomics in Drug Development
- •4 Case Studies of Genomics-Based Drug Design
- •References
- •1 Historical Background
- •1.1 Timeline
- •2 Methodology Overview
- •2.1 Neural Networks
- •2.1.1 Perceptron
- •2.1.2 Multilayer Neural Networks
- •2.1.3 Types of Neural Networks
- •Feedforward
- •Recurrent Neural Networks
- •LSTM
- •2.2 Deep Learning
- •3 Using Machine Learning
- •3.2 Data Collection
- •3.3 Data Preprocessing
- •3.4 Model Selection
- •3.5 Model Training
- •3.6 Validation
- •3.7 Tuning
- •3.8 Prediction
- •4 Limitations
- •4.1 Bias
- •4.3 Interpretability
- •4.4 Computational Cost
- •4.5 Data Dependency
- •4.6 Robustness
- •5 Applications in Drug Discovery
- •5.2 Lead Discovery
- •5.3 Preclinical and Clinical Development
- •6 Resources and Tools
- •7 Challenges and Perspectives
- •7.1 Future Trends
- •9 Conclusions
- •References
- •1 Historical Background
- •1.1 Applications in Drug Discovery
- •2 Validations and Controls
- •2.1 Internal Validation
- •2.2 External Validation
- •2.3 Relative Cluster Validation
- •3 Challenges and Perspectives
- •4 Conclusions
- •References
- •1 Historical Background
- •2 OECD Principles
- •2.1 A Defined Endpoint
- •2.2 An Unambiguous Algorithm
- •2.5 A Mechanistic Interpretation, if Possible
- •3 Software and Tools
- •4 Validations and Controls
- •4.1 Internal and External Validation
- •4.1.1 Regression Metrics
- •4.2 Applicability Domain
- •4.3 Randomization Tests
- •5 Interpretation
- •6 Practical Advice During QSAR Modeling
- •7 Application
- •8 Challenges and Perspectives
- •References
- •1 Molecular Docking
- •2 Advances in Scoring Functions and Search Algorithms
- •2.2 Critical Characteristics of Search Algorithms
- •2.3 Docking Programs and Scoring Functions
- •3 Calculations Performed During Docking Simulations
- •4 Essential Components for a Good Docking Program
- •5 Limitations of the Docking Technique
- •6 Validation of Docking Results
- •7 Inappropriate Use of Validation Methods in Docking
- •9 Use of Machine Learning in Molecular Docking
- •11 Challenges
- •12 Conclusions
- •References
- •3 System Preparation for MD Simulations
- •3.1 Solvation and Microensemble
- •3.2 Force Fields: General Concept and Relevant Choices
- •3.3 The Concept of Replicas and Timescale
- •4.1.2 Protein Root Mean Square Fluctuation (RMSF)
- •4.1.4 Protein Secondary Structure Analysis
- •4.1.5 Principal component Analysis (PCA)
- •4.1.6 Markov State Modelling
- •4.1.7 Distance Calculations
- •4.1.8 Angle and Plane Calculations
- •4.2.2 Distances and Ligand-Induced Geometry Rearrangements
- •4 Molecular Dynamics Analysis
- •4.1 Protein Perspective
- •4.1.1 Protein Root Mean Square Deviation (RMSD)
- •4.3 Ligand Perspective
- •4.3.1 Ligand Properties
- •4.3.2 Ligand Root Mean Square Deviation
- •4.3.3 Ligand Root Mean Square Fluctuation
- •4.3.4 Angles and Dihedrals
- •5.1 Protein Structure Prediction and Preparation
- •5.2 Molecular Docking
- •6 Concluding Remarks and Outlook
- •Glossary
- •References
- •1 Introduction
- •2.1 MDeNM
- •2.2 Collective Molecular Dynamics (coMD)
- •2.3 ClustENM and ClustENMD
- •3 Ensemble Docking
- •References
- •1 Introduction
- •1.1 Advantages, Disadvantages, Innovations, and Challenges
- •1.2 Recent Advances in Accessible FEP Software Tools
- •1.3 Applications of FEP in Industry and Consortiums
- •2 Expanding the Potential of FEP Calculations
- •2.1 Validating Binding Poses
- •2.2 Dealing with Solvent
- •2.3 FEP and Allostery
- •2.4 FEP and Covalent Ligands
- •2.5 Applications of FEP in Scaffold Hopping
- •2.6 Positional Analogue Scanning
- •2.7 Combinations and Alternative Approaches
- •3 Machine Learning for FEP
- •3.4 Implications for ML in FEP Calculations
- •4 Final Considerations
- •5 First Steps to FEP Simulations
- •References
- •1 Background
- •2 Ultra-Large Screening Libraries and Chemical Spaces
- •3.1 Implications of Dataset Size
- •4 Ligands on the Ultra-Large Scale
- •4.1 Ultra-Large 2D Similarity Searches
- •7 Challenges and Future Perspectives
- •7.1 Hit Triage: An Old Problem on a New Dimension
- •8 Conclusions
- •Appendix
- •References
- •1 Introduction
- •2 Enzymatic Activity Evaluations
- •3 Cytotoxicity Evaluation and Cell Viability
- •4 Antiviral Assays in Experimental Validation
- •6 In Vivo Evaluation of Compounds
- •7 Conclusions
- •References
- •1 Introduction
- •3.1 Data Collection
- •3.2 Data Preprocessing
- •3.4 Model Choice
- •3.5 Model Training
- •3.6 Model Assessment
- •3.7 External Validation
- •3.8 Implementation and Availability
- •3.9 Continuous Update
- •5 Conclusions and Perspectives
- •References
- •1 Experimental Approaches to Obtain Protein Structure
- •1.1 X-Ray Crystallography
- •1.2 Nuclear Magnetic Resonance
- •1.3 Cryo-EM
- •1.4 Hybrid Methods
- •2 Modeling Approaches to Obtain Protein Structure
- •2.1 Homology Modeling
- •2.2 Ab Initio Modeling
- •2.3 New Approaches
- •3 Conformational Diversity of Proteins
- •3.1 Characterization of Protein Conformational States
- •3.2 Experimental Methods to Study Protein Dynamics and Conformations
- •3.4 Molecular Dynamics Simulation
- •3.5 Sampling Strategies
- •4 Remarks and Perspectives
- •References
- •1 Introduction
- •2 Structure-Based Drug Design of HIV Protease Inhibitors
- •2.1 HIV-1 Protease as a Therapeutic Target
- •2.2.1 Saquinavir
- •2.2.2 Indinavir
- •2.3.1 Lopinavir
- •2.3.2 Darunavir
- •6 Conclusions
- •References
- •4 Experimental Methods to Analyze NR Activity
- •4.2 Coregulator-Recruitment
- •5 Concluding Remarks and Outlook
- •References

8 Drug Design in Motion: Concepts and Applications of Classical... 203
interest, computational methods can be applied to suggest the probable locations.
One can choose between the methods relying on geometrical properties, such as
POCKET [7], PASS [8], LIGSITE [9], or combined with the physics approach, such
as PocketFinder [10] or SiteMap [11, 12].
There are several options to pursue after the binding site is identified. For
instance, one can establish a pharmacophore model, which offers an additional
opportunity to validate the model before proceeding with pharmacophore-based
virtual screening. Another option is to carry out the docking (see Chap. 9)ofa
substantial chemical library. Once the method is chosen, and the screening database
(see Chap. 2 for more detailed information about molecular databases) is selected
and prepared according to the requirements of the method, the virtual screening
(VS) can be utilised. Next, when the initial hits are obtained, the compound set
should be scored and filtered according to the properties required for project
purposes. The essential pharmacokinetic properties for the hits list validation include
physicochemical parameters (molecular weight, number of heavy atoms, hydrogen
bond donors and acceptors, and rotatable bonds), lipophilicity (Log Po/w), water
solubility (log S), pharmacokinetics, drug-l ikeness, and chemical synthesis accessibility [13]. The cycle between model validation and the top hits choice is frequently
repeated aiming for refinement. Once the initial hits satisfy the criteria, optional
short-scale MD simulations (200–500 ns) can be performed to validate the ligand
stability within the binding pocket, ensuing in selecting a list of top hits for
experimental validation. The relevance of MD simulations in multiple stages of
this pipeline is discussed in the next sections of this chapter.
2 The Early Beginnings of Molecular Dynamics
Simulations
Molecular dynamics (MD) simulations have transformed the landscape of drug
design and development, providing valuable insights into the behaviour of biomolecules at the atomic level. The roots of MD can be traced back to the
mid-twentiethcentury when the exploration of methods to study the dynamics of
molecules and their interactions in silico started to gain popularity. The early
development of MD can be attributed to the pioneering work of Dr. Martin Karplus
and Dr. Michael Levitt, who in the early 1970s, independently laid the foundation
for the theoretical framework of MD simulations [14, 15]. Their seminal papers
explored the concept of using classical mechanics to describe the motion of atoms
and molecules over time. These fundamental contributions ignited interest in the
scientific community and paved the way for further advancements.
In 1975, Levitt and Warshel introduced a simplified representation of protein
structures that reduced the degrees of freedom and the number of interaction centres
by averaging over groups of atoms [16]. They represented each residue using only
two centres, the Cα atom and the centroid of the side chain and assumed interactions

204 E. Shevchenko et al.
to occur only between side chains. The model reduced the dimensionality of the
conformational space and led to smoother side chain representations. They used
effective time-averaged potential functions to account for the fine details and rapidly
changing variables. The simplified model was tested on a bovine pancreatic trypsin
inhibitor (PTI) and successfully simulated the stable conformation of a folded
protein. The authors also proposed a hierarchical approach, starting with simplified
models and gradually incorporating more atoms, to understand and simulate complex biological assembly processes [17, 18]. Their work laid the foundation for MD
simulations in protein folding and biomolecular studies.
The 1980s witnessed significant progress in MD simulations, as computational
power and algorithms improved. Scientists began to apply MD to study the dynamics of small molecules and biomolecules such as proteins and nucleic acids
[19]. Driven by advancements in both hardware and software, MD simulations
became increasingly popular for exploring the structural and dynamic properties of
biological systems.
In the 1990s, MD simulations evolved into a versatile tool for understanding
complex biological phenomena and interactions [20, 21]. Researchers started using
MD to investigate protein-ligand binding, protein folding, and protein–protein
interactions, among other biologically relevant processes [22–25]. These studies
provided valuable insights that complemented experimental findings and aided in
rational drug design.
With the advent of more sophisticated force fields, parallel computing, and
efficient algorithms, MD simulations entered a new era in the twenty-first century
[26–30]. Simulations of larger biomolecular systems, such as membrane proteins
and nucleic acid complexes, became feasible [31]. The combination of MD with
other computational methods, like quantum mechanics/molecular mechanics
(QM/MM) (see Chap. 10), further expanded the capabilities of drug design
approaches [32, 33].
Nowadays, MD simulations have become an indispensable tool in modern drug
discovery and development. They play a crucial role in predicting ligand binding
affinities, exploring the protein–ligand complexes [34, 35], facilitating drug binding
studies [36, 37], understanding protein–protein interactions [38], and discovering
structural states and binding sites [39]. The integration of MD simulations with
experimental techniques accelerates the identification and optimisation of potential
drug candidates.
MD simulations can empower the studying of the structure within molecular
systems at the atomic scale in a dynamic manner. The MD simulation generates a
trajectory that allows observing an atom’s individual movement as a function of
time. The forces between interacting atoms are estimated using a force field, and the
system
’s total energy is computed according to Newton’s law of classical mechanics.
The integration of Newton’s laws of movements produces subsequent configurations
of the developing system during MD simulations, producing trajectories that
describe the locations and velocities of the particles throughout time. The

8 Drug Design in Motion: Concepts and Applications of Classical... 205
applications of MD simulations cover a wide range of possibilities within in silico
drug design. Among others are the docking postprocessing of protein–ligand complexes [34, 35], drug binding studies [36, 37], protein–protein interactions [38], and
the discovery of structural states and binding sites [39].
3 System Preparation for MD Simulations
A basic pipel ine for setting up a system for MD simulation requires the protein of
interest to be prepared by the simulation software and chosen force field. System
preparation follows the structure preparation with additional examination of
crystallisation artefacts, correct bond orders , the assignment of correct protonation/
ionisation states to protein residues and the presence of ions and associated molecules. On occasion, the modelling of missing residues in pdb structures also takes
place here. In the next sessions we discuss, timescales and replicas, followed by
ensemble models, parametrisation, and force fields, followed by solvation and
periodic conditions.
3.1 Solvation and Microensemble
Next, the solvation model is to be chosen. Many types of water models have been
developed regarding the system solvation with water. These models are categorised
based on several important factors, including the number of interaction points the
molecule can make (site), the inclusion of polarisation effects, and flexibility or
rigidity. One of the most common water models is TIP3P, a three-site rigid water
molecule with assigned Lennard-Jones parameters and potential [40].
A thermodynamics ensemble used to characterise the system must be chosen
during the system preparation step. The thermodynamics ensemble is an idealisation
of the model system composed of multiple replicas of this system, all taken into
account simultaneously, and each reflects a potential state in which the actual system
may be [41]. In other words, the thermodynamics ensemble is a statistical approximation that allows to extrapolate the fundamental features of the thermodynamic
system through classical and quantum mechanics [42]. The ensemble classification
is based on how the model system is separated from the outer environment. The first
category is the microcanonical ensemble, which depicts the model system as entirely
isolated from the environment (Fig. 8.2a). The microcanonical ensemble refers to the
acronym NVE, which means that the total number of particles N, volume V, and
energy E remains constant [43]. As a factor that assumes the thermodynamic
interaction of the model system with its surroundings, the temperature cannot be
defined for the NVE ensemble. NVE is used for describing states of a system with
defined total energy. The next ensemble category is termed canonical and provides
the option to determine the system temperature, therefore referring to the acronym

206 E. Shevchenko et al.
Fig. 8.2 Thermodynamic ensembles in molecular dynamics. (a) A microcanonical (NVE) ensemble is an isolated system with a fixed number of particles, total volume, and energy. (b) Canonical
(NVT) ensemble as a system with fixed total volume but with temperature-dependent boundary
energy transfer. (c) Isothermal-Isobaric (NpT) ensemble as a system allowing volume alteration for
the pressure equilibration. Ext external, environmental, in internal, model system value
NVT, where N stands for the total number of particles, V for volume, and T for
temperature. NVT allows the energy transfer between the boundaries of the model
system and the environment while prohibiting substance exchange (Fig. 8.2b)[43].
In other words, this ensemble addresses the model system’s potential states that
are in thermal equilibrium with the environment. This ensemble model can be
visualised as a heat bath with constant temperature, which is several magnitudes
larger than the model system [44]. The difference in size ensures that no amount of
heat produced by the model system will cause the heat bath’s temperature to rise
considerably. As a result of the thermal contact between the model system and the
environment, the system will now trans mit heat to and across the environment until
they reach thermal equilibrium [45]. The last ensemble category is the isothermalisobaric, which as with the canonical ensemble, allows the energy transfer between
the model system and environment, but allows the volume change. The system’s
volume fluctuates to equilibrate the system’s internal pressure with the pressure
applied to the system by its environment (Fig. 8.2c)[45]. The acronym for the
isothermal-isobaric ensemble is NpT, where N stands for the persistent total number
of particles, p for pressure, and T for temperature.
Moreover, NpT is the most used ensemble. However, prior to the start of MD, the
so-called model relaxation step is conducted. The model relaxation is an essential
step which ensures the proper system equilibration and removal of the high strain
degree from the newly constructed system. The model relaxation involves several
phases, where a particular ensemble is combined with specific pressure, temperature,
and timestep until the equilibration is achieved. The placement and selection of ions
and the specification of temperature and pressure are all components of the system
preparation process. It is frequently desir able to have an electrically neutral system

8 Drug Design in Motion: Concepts and Applications of Classical... 207
for MD simulation. However, it is not strictly necessary, for example, with the
Desmond engine [46], which applies a uniform background charge distribution to
neutralise the system in the Ewald summation [47]. Moreover, systems can be set up
in a salt solution rather than a pure solvent or even include other small molecules as
buffers in a mixed-solvent model [48].
Finally, periodic boundary conditions (PBC) are then applied to the system by
selecting the simulation box’s form and cut-off. It is necessary to note that to account
for potential protein movements, the distance from the simulation box’s edge to the
protein must be considerable. If the simulation box size is insufficient, MD-derived
findings can frequently be artef actual or misleading.
Once the box shape and dimensions, the total number of atoms, number of water
molecules, implicit or explicit membranes or solvents, salt concentration, lipid
composition, cofactors, and metals have been determined, the force field should be
chosen to continue with the MD simulation.
3.2 Force Fields: General Concept and Relevant Choices
While in classical physics, force fields evaluate a system’s potential energy, the
fundamental distinction between the two in molecular modelling is that the energy
landscape is depicted as an energy gradient dispersed across the particle positions
[76]. Reconstructing a realistic simulation of a molecular system on an atomistic
level is strongly dependent on force field parameters, which are the core for deriving
meaningful structural information and relative energies from MD simulation.
To date, additive or non-polarisable force fields are the most prevalent in small
molecule drug design, frequently sharing the basic potential energy function and the
parameters comprising the function itself [49]. The name “additive” originates from
Coulomb’s equation for electrost atic interactions, which states that a system’s
potential energy is the sum of all atom–atom individual interactions (Eq. 8.1)
[49]. The potential functions, which are essentially an array of equations
representing the potential energy, and its components, form the basis of the force
field core. The parameters in this set of equations are another component
distinguishing a force field.
In the vast majority of condensed-phase simulations, the total potential energy of
a cohort of molecules is determined as a sum of inter- and intra-molecular interaction
energies between all components of the system (Eq. 8.1)[50, 51].
→
EX
=
a < b
Eab þ
The total potential energy of a set of constituents ab with coordinates X
int
E
a
a
ð8:1Þ
→
Composed by the bonded terms (first term) and nonbonded interactions (second
term), ab refers to a set of molecules and/or ions. This notation applies to two
exemplary particles a and b, which in a real calculation condition would be expanded
to multiple pairs of particles.
.

208 E. Shevchenko et al.
The following terms commonly describe the intramolecular potential energy:
harmonic bond stretching (Eq. 8.2), angle bending (Eq. 8.3), torsional angle definition as Fourier series (Eq. 8.4), Coulomb electrostatics and Lennard-Jones potential
(Eq. 8.5)[51].
E
bond
=ik
ðÞ
b,iri
- r
2
0,i
ð8:2Þ
The first force field term defines bond stretching energies in harmonic (ideal)
conditions, where k
is the bond-stretching constant, which regul ates the rigidity of
b
the bond spring.
E
bend
=ik
ðÞ
ϑ,iϑi
The second force field term defines the angle bending energy, where k
- ϑ
2
0,i
ϑ
ð8:3Þ
is the
angle bending const ant, determining the rigidity of the spring’s angle.
E
torsion
ðÞ
=
1,i
i
2
V
1 - cos 2φ
ðÞ
2,i
i
þ
2
V
1 þ cos 3φ
ðÞ
3,i
i
þ
2
i
ð8:4Þ
V
1 þ cos φ
The third force field term represents the torsional energies as a Fourier series.
2
E
=
nb
i < j
qiqje
r
þ 4ε
ij
12
σ
ij
ij
r
ij
6
σ
ij
-
r
ij
ð8:5Þ
The fourth force field term represents the nonbonde d energy between all atom
pairs, where σ
is the equilibrium distance and rijis the distance between interacting
ij
atoms.
Force fields can be roughly divided into three major classes. One of the defining
characteristics of class one is the utilisation of harmonic movements to depict bond
stretching and angle bending. According to the assumption, the amount of restoring
force is proportional to the displacement from the equilibrium position for class one
force fields [52]. The approximation in class one force fields is referred to as
quadratic because the square of the displacement energy is linearly associated with
the energy of the harmonic oscillator [53]. Moreover, the parametrisation of bond
stretching and the angle bending often approach harmonic behaviour only close to
the equilibrium. The most famous examples of the class one force fields are the
Optimised Potentials for Liquid Simulations (OPLS) [54], AMBER [55, 56],
CHARMM [57], and GROMOS [58]. One of the first techniques to be established
with constant parameter optimisation for the propagation of thermodynamic characteristics in the liquid state applied to small molecules was OPLS, which stands for
optimised potentials for liquid simulations [54]. On the OPLS core, the next generation of force fields was built: OPLS3 [28], OPLS3e
29,
and OPLS4 [30] with the
maintenance of the nonbonded parameters. In general, AMBER refers to a collection

8 Drug Design in Motion: Concepts and Applications of Classical... 209
of force fields that can be split according to the simulated biomolecular system. For
instance, for protein simulations, AMBER suggests ff19SB [59], and its evolution
ff19SB-ILDN (with improved side-chain torsion potentials [60]); for lipids or
complex membrane simulations—LIPID21 [61]; q and for nucleic acids—DNA
OL15 [62], and RNA OL3 [63]. As AMBER, CHARMM also represents a set of
force fields, for instance, the all-atom CHARMM22 [64] and extended atom force
field CHARMM19 [65], which are annually updated (see https://www.
academiccharmm.org/program/versions for a current list, accessed on February
2024). Last but not least, the very popular GROMOS force-fields (not to be confused
with the GROMOS software), which are united atom force fields, i.e., without
explicit aliphatic (non-polar) hydrogens and are considered an “all purposed” set,
with accurate parameters for phosphorylation and others post-translational modifications (see [66] for the most recent version and https://www.gromos.net/#Reif2012,
for current list and modifications, accessed on February 2024).
Class two provides additional anharmonic cubic and quartic components for the
potential energy of bonds, resulting in more detailed geometrical modelling of the
vibrations of bonds. In addition, these force fields include cross-terms that describe
the interactions between nearby located angles and dihedrals. Class two force fields
include MMFF94 (Merck Molecular Force Field [67]), which parameters are primarily derived from quantum calculations rather than experimental data [68].
Another example is UFF [69] (Universal Force Field), whose original application
was somewhat restricted due to the parameters, particularly for metals and inorganic
substances [41]. The UFF initiative seed the idea recent ly followed up, in 2018, by
the Open Force Field consortium (OFF, https://docs.openforcefield.org/en/latest/,
accessed on February 2024). On their own words, they aim to develop automated
and systematic data-driven techniques to parameterise and assess new generations of
more accurate force fields. Despite their efforts, the most recent benchmarking
shows that public force fields (i.e., OpenFF Parsley and Sage, GAFF and CGenFF)
had comparable accuracy, while OPLS3e was significantly more accurate [70].
Class three includes force fields that contain extended parameters applicable to
organic chemistry, such as the Jahn-Teller effect or stereoelectronic effects. For
example, AMOEBA [71] is a polarisable force field that employs atomic-induced
dipole to model polarisation while assuming that averaging polarisation is insufficient [72]. Another example is DRUDE [73], which uses non-polarisable force fields
to leverage atom-to-atom Coulomb electrostatic interactions as its core while integrating polarisation effects via NAMD and a dual-Langevin thermostat
approach [74].
Important to highlight another class of force fields, known as coarse-grained force
fields, which employ a distinct strategy in molecular dynamics simulations. The idea
behind the coarse-grained approach is the reduction of the number of degrees of
freedom within a system. This is achieved by parametrising the most significant
interactions with the force field while representing a particular set of atoms as a
single bead. The definition of the most significant interactions might be intricate
depending on the parametrisation method, hence tabulated potentials are frequently
employed. The purpose of coarse-grained models is to replicate specific

210 E. Shevchenko et al.
characteristics of a given system, which can encompass an atomistic protein model
or experimental data. The properties one intends to replicate in the model determine
the classification of the coarse-grained force fields. For instance, free energy conservation is the focus of the MARTINI [40] force field and the simplex method
[73]. Another example is inverse Monte Carlo with structure-based coarse-grained
modelling, emphasising the radial distribution [75].
In the case no parameters are available for a molecule in that particular force field,
one can resource to reparameterisation or development of new parameters, which
will be force field compatible. In order to ensure reproducibility, though, the new
data would need to be validated against experimental measurements or high-level
QM data (for more information on quantum mechanics, see Chap. 10). It is also
important to provide a description of bonded and nonbonded potential parame ters,
by atom and provide the topology(-ies) used for that determination. Software tools
such as the Antechamber (from the Amber package [76]) or web server interfaces
like LigParGen (for OPLS force-fields [77]) facilitate the topology analyses, writing
the input for the correspondi ng simulation engine [78].
3.3 The Concept of Replicas and Timescale
One of the essential aspects that can be carefully considered when the MD simulation is conducted is the timescale, as different types of protein motions occur at
distinct timescales. For instance, the side-chain rotamer movements can be observed
in the range of ps to μs, followed by the loop motions in the field of ns to μs, with
more significant domain movements starting to be evident in μs + timescale
[79]. Thus, one should consider the reasonable timescale of the MD simulation in
accordance with the movements to be observed to deliver valuable results (Fig. 8.3).
To obtain meaningful and comprehensive results from MD simulations, one
should carefully consider one of the essential aspects of this methodology. The
selection of an appropriate timescale directly influences the range of molecular
motions and events that can be captured within the simulated system. Shorter
timescales are suitable for studying fast processes such as bond vibrations and
side-chain rotamer movements (fs to ns). In contrast, longer timescales are required
to investigate slower and more complex events, such as loop motions (ns to μs),
significant domain movements and large-scale conformational transitions (μs+).
However, excessively long simulations conducted without a clear research purpose,
may not only become computationally expensive but also lead to diminishing returns
in terms of new insights. To strike a balance between accuracy, computational
feasibility and deriving meaningful results, one should align the simulation timescale
requirements with the specific research question to be answered through MD .
The current model for understanding protein folding and dynamics asks us to
imagine the protein dynamics process as a large funnel, in which the surface would
represent the diversity of protein conformation, and its deepness would be proportional to the energy systems. The deeper pockets of this funnel constitute global

8 Drug Design in Motion: Concepts and Applications of Classical... 211
Fig. 8.3 A spectrum of
dynamic processes in
proteins across timescale
occurring in molecular
dynamics simulation.
(Adapted from HenzlerWildman et al. [79],
dynamic personalities of
protein [79])
low-energy protein conformations. This funnel is rather wrinkly, with smaller wells,
which represent the energy minima of different protein conformations. What the MD
simulation does in this model is to navigate these inner surfaces and unbiasedly
sample around the protein conformational space.
Considering the funnel model and the problem of timescale for protein dynamics,
we understand that classical unbiased MD simulations might not be always suitable
to sample molecular events involving a transition between those energy barriers,
such as large conformational changes. Alternatively, given enough justification and
experimental evidence, one can apply enhanced sampling simulations, such as
nonadaptive biasing potential methods (e.g., Gaussian-accelerated MD), adaptive
bias simulations (e.g., metadynamics), replica exchange methods, etc. Examples of
experimental evidence can rise from (TR-)FRET assays or site-direct mutagenesis in
combination with photocrosslinking, showing that two portions of the protein are in
close proximity, multiple diverse crystal structures, showing different conformations
and even highlighting large “missing” portions due to high flexibility.
If enhanced sampling simulations are used, the choices for sampling methodology, parameters, and, most importantly, the convergence criteria (where to stop
sampling and when to consider that the generated conformations are artificial)
should be described. We refer to the recently published guidelines for reporting
MD simulation data [80].
The other crucial point for comprehensive MD simulation is the number of
replicas of the single system to be conducted to avoid false positive conclusions
[81]. The term “replica” refers to the simulations of the identical system, sharing the
same number of atoms, initial structure, and preparation protocol, repeated several

212 E. Shevchenko et al.
times. Given the principle of energy conservation within the ensemble required for
the MD calculations and the almost derisory effects arising from changes in the
floating-point precision or hardware, multiple runs of these syst ems would lead to
very similar thermodynamical and, consequently, conformational properties. In this
sense, the difference between replicas is in the initial velocities generated randomly
according to the Maxwell distribution [81]. As a result, the velocities are unique for
each replica, which leads to the production of different simulation trajectories.
However, even replicas with identical velocities can produce distinct trajectories
for various reasons, including machi ne-specific settings and the specification of the
compiling [81]. In theory, the ergodic principle claims that the velocities have no
impact on infinite dynamic simulations [82 ]. While molecular dynamics are not run
eternally in real life, random velocities are strictly important, especially for shortscale MDs. The random initial velocity assignment ensures that the results of every
simulation are slightly different, even if the other settings are identical. In other
words, the random velocities provide the opportunity to observe real-world phenomena happening with the same system at different time points. Moreover, multiple
replicas ensure a specific movement or interaction observed not in a single simulation but in a statist ically significant number of replicas is not biased or artificially
generated but related to the real-world evidence. The choice between the multiple
simulation replicas over a single but long-scale one is frequently a question that
should be answered before the initiation of the MD according to the aims and
resources of the project.
The work from Knapp et al. [81] performed 100 identically parametrised replicas
of 3 μs each for a small 10 amino acid long peptide, and also 100x replicas of 100 ns
each for the T-cell receptor/MHC system (827 amino acids). Using those simulations
as show-case, they compared randomly chosen subgroups within these replica sets,
being able to estimate the reproducibility and reliability that could be achieved by a
given number of replicas at a given simulation time. They highlight two major
points: i) conclusions drawn from single simulations are not reproducible and ii)
observations drawn from multiple shorter replicas were more reliable than using a
single longer trajectory. Based on their data, a minimum of 5–10 replicas were
required [81]. Their comment that the actual number of replicas needed will depend
on the level of reliability sought and, herein we add that is highly dependent on the
size and intrinsic flexibility of the system.
Our next example applied multiple short sim ulations on the conformational space
of aptamers [83]. They combine the enhanced sample approach of starting from
high-temperature equilibration runs (six in total), leading to higher flexibi lity, with
independent replicas from the output of each of these systems (10 × 100 ns per
output). They analysed the recurrence rate of found conformations and properties
among the different starting points, showing a clear dependence on initial conformation. This highlights the need of using different initial configurations as simulation starting points to avoid being trapped in energy minima. Especially, since
analyses of convergence of those unbiased trajectories might not allow the detection
of slow trans itions between kinetically trapped met astable states.
Соседние файлы в папке Библиотека им академика М.И. Перельмана
