Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5884_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
02.09.2026
Размер:
21 Мб
Скачать
8 Drug Design in Motion: Concepts and Applications of Classical... 203
interest, computational methods can be applied to suggest the probable locations. One can choose between the methods relying on geometrical properties, such as POCKET [7], PASS [8], LIGSITE [9], or combined with the physics approach, such as PocketFinder [10] or SiteMap [11, 12].
There are several options to pursue after the binding site is identied. For instance, one can establish a pharmacophore model, which offers an additional opportunity to validate the model before proceeding with pharmacophore-based virtual screening. Another option is to carry out the docking (see Chap. 9)ofa substantial chemical library. Once the method is chosen, and the screening database (see Chap. 2 for more detailed information about molecular databases) is selected and prepared according to the requirements of the method, the virtual screening (VS) can be utilised. Next, when the initial hits are obtained, the compound set should be scored and ltered according to the properties required for project purposes. The essential pharmacokinetic properties for the hits list validation include physicochemical parameters (molecular weight, number of heavy atoms, hydrogen bond donors and acceptors, and rotatable bonds), lipophilicity (Log Po/w), water solubility (log S), pharmacokinetics, drug-l ikeness, and chemical synthesis accessi­bility [13]. The cycle between model validation and the top hits choice is frequently repeated aiming for renement. Once the initial hits satisfy the criteria, optional short-scale MD simulations (200–500 ns) can be performed to validate the ligand stability within the binding pocket, ensuing in selecting a list of top hits for experimental validation. The relevance of MD simulations in multiple stages of this pipeline is discussed in the next sections of this chapter.
2 The Early Beginnings of Molecular Dynamics
Simulations
Molecular dynamics (MD) simulations have transformed the landscape of drug design and development, providing valuable insights into the behaviour of bio­molecules at the atomic level. The roots of MD can be traced back to the mid-twentiethcentury when the exploration of methods to study the dynamics of molecules and their interactions in silico started to gain popularity. The early development of MD can be attributed to the pioneering work of Dr. Martin Karplus and Dr. Michael Levitt, who in the early 1970s, independently laid the foundation for the theoretical framework of MD simulations [14, 15]. Their seminal papers explored the concept of using classical mechanics to describe the motion of atoms and molecules over time. These fundamental contributions ignited interest in the scientic community and paved the way for further advancements.
In 1975, Levitt and Warshel introduced a simplied representation of protein structures that reduced the degrees of freedom and the number of interaction centres by averaging over groups of atoms [16]. They represented each residue using only two centres, the Cα atom and the centroid of the side chain and assumed interactions
204 E. Shevchenko et al.
to occur only between side chains. The model reduced the dimensionality of the conformational space and led to smoother side chain representations. They used effective time-averaged potential functions to account for the ne details and rapidly changing variables. The simplied model was tested on a bovine pancreatic trypsin inhibitor (PTI) and successfully simulated the stable conformation of a folded protein. The authors also proposed a hierarchical approach, starting with simplied models and gradually incorporating more atoms, to understand and simulate com­plex biological assembly processes [17, 18]. Their work laid the foundation for MD simulations in protein folding and biomolecular studies.
The 1980s witnessed signicant progress in MD simulations, as computational power and algorithms improved. Scientists began to apply MD to study the dynam­ics of small molecules and biomolecules such as proteins and nucleic acids [19]. Driven by advancements in both hardware and software, MD simulations became increasingly popular for exploring the structural and dynamic properties of biological systems.
In the 1990s, MD simulations evolved into a versatile tool for understanding complex biological phenomena and interactions [20, 21]. Researchers started using MD to investigate protein-ligand binding, protein folding, and protein–protein interactions, among other biologically relevant processes [2225]. These studies provided valuable insights that complemented experimental ndings and aided in rational drug design.
With the advent of more sophisticated force elds, parallel computing, and efcient algorithms, MD simulations entered a new era in the twenty-rst century [2630]. Simulations of larger biomolecular systems, such as membrane proteins and nucleic acid complexes, became feasible [31]. The combination of MD with other computational methods, like quantum mechanics/molecular mechanics (QM/MM) (see Chap. 10), further expanded the capabilities of drug design approaches [32, 33].
Nowadays, MD simulations have become an indispensable tool in modern drug discovery and development. They play a crucial role in predicting ligand binding afnities, exploring the protein–ligand complexes [34, 35], facilitating drug binding studies [36, 37], understanding protein–protein interactions [38], and discovering structural states and binding sites [39]. The integration of MD simulations with experimental techniques accelerates the identication and optimisation of potential drug candidates.
MD simulations can empower the studying of the structure within molecular systems at the atomic scale in a dynamic manner. The MD simulation generates a trajectory that allows observing an atoms individual movement as a function of time. The forces between interacting atoms are estimated using a force eld, and the system
s total energy is computed according to Newtons law of classical mechanics. The integration of Newtons laws of movements produces subsequent congurations of the developing system during MD simulations, producing trajectories that describe the locations and velocities of the particles throughout time. The
8 Drug Design in Motion: Concepts and Applications of Classical... 205
applications of MD simulations cover a wide range of possibilities within in silico drug design. Among others are the docking postprocessing of protein–ligand com­plexes [34, 35], drug binding studies [36, 37], protein–protein interactions [38], and the discovery of structural states and binding sites [39].

3 System Preparation for MD Simulations

A basic pipel ine for setting up a system for MD simulation requires the protein of interest to be prepared by the simulation software and chosen force eld. System preparation follows the structure preparation with additional examination of crystallisation artefacts, correct bond orders , the assignment of correct protonation/ ionisation states to protein residues and the presence of ions and associated mole­cules. On occasion, the modelling of missing residues in pdb structures also takes place here. In the next sessions we discuss, timescales and replicas, followed by ensemble models, parametrisation, and force elds, followed by solvation and periodic conditions.
3.1 Solvation and Microensemble
Next, the solvation model is to be chosen. Many types of water models have been developed regarding the system solvation with water. These models are categorised based on several important factors, including the number of interaction points the molecule can make (site), the inclusion of polarisation effects, and exibility or rigidity. One of the most common water models is TIP3P, a three-site rigid water molecule with assigned Lennard-Jones parameters and potential [40].
A thermodynamics ensemble used to characterise the system must be chosen during the system preparation step. The thermodynamics ensemble is an idealisation of the model system composed of multiple replicas of this system, all taken into account simultaneously, and each reects a potential state in which the actual system may be [41]. In other words, the thermodynamics ensemble is a statistical approx­imation that allows to extrapolate the fundamental features of the thermodynamic system through classical and quantum mechanics [42]. The ensemble classication is based on how the model system is separated from the outer environment. The rst category is the microcanonical ensemble, which depicts the model system as entirely isolated from the environment (Fig. 8.2a). The microcanonical ensemble refers to the acronym NVE, which means that the total number of particles N, volume V, and energy E remains constant [43]. As a factor that assumes the thermodynamic interaction of the model system with its surroundings, the temperature cannot be dened for the NVE ensemble. NVE is used for describing states of a system with dened total energy. The next ensemble category is termed canonical and provides the option to determine the system temperature, therefore referring to the acronym
206 E. Shevchenko et al.
Fig. 8.2 Thermodynamic ensembles in molecular dynamics. (a) A microcanonical (NVE) ensem­ble is an isolated system with a xed number of particles, total volume, and energy. (b) Canonical (NVT) ensemble as a system with xed total volume but with temperature-dependent boundary energy transfer. (c) Isothermal-Isobaric (NpT) ensemble as a system allowing volume alteration for the pressure equilibration. Ext external, environmental, in internal, model system value
NVT, where N stands for the total number of particles, V for volume, and T for temperature. NVT allows the energy transfer between the boundaries of the model system and the environment while prohibiting substance exchange (Fig. 8.2b)[43].
In other words, this ensemble addresses the model systems potential states that are in thermal equilibrium with the environment. This ensemble model can be visualised as a heat bath with constant temperature, which is several magnitudes larger than the model system [44]. The difference in size ensures that no amount of heat produced by the model system will cause the heat baths temperature to rise considerably. As a result of the thermal contact between the model system and the environment, the system will now trans mit heat to and across the environment until they reach thermal equilibrium [45]. The last ensemble category is the isothermal­isobaric, which as with the canonical ensemble, allows the energy transfer between the model system and environment, but allows the volume change. The systems volume uctuates to equilibrate the systems internal pressure with the pressure applied to the system by its environment (Fig. 8.2c)[45]. The acronym for the isothermal-isobaric ensemble is NpT, where N stands for the persistent total number of particles, p for pressure, and T for temperature.
Moreover, NpT is the most used ensemble. However, prior to the start of MD, the so-called model relaxation step is conducted. The model relaxation is an essential step which ensures the proper system equilibration and removal of the high strain degree from the newly constructed system. The model relaxation involves several phases, where a particular ensemble is combined with specic pressure, temperature, and timestep until the equilibration is achieved. The placement and selection of ions and the specication of temperature and pressure are all components of the system preparation process. It is frequently desir able to have an electrically neutral system
8 Drug Design in Motion: Concepts and Applications of Classical... 207
for MD simulation. However, it is not strictly necessary, for example, with the Desmond engine [46], which applies a uniform background charge distribution to neutralise the system in the Ewald summation [47]. Moreover, systems can be set up in a salt solution rather than a pure solvent or even include other small molecules as buffers in a mixed-solvent model [48].
Finally, periodic boundary conditions (PBC) are then applied to the system by selecting the simulation boxs form and cut-off. It is necessary to note that to account for potential protein movements, the distance from the simulation boxs edge to the protein must be considerable. If the simulation box size is insufcient, MD-derived ndings can frequently be artef actual or misleading.
Once the box shape and dimensions, the total number of atoms, number of water molecules, implicit or explicit membranes or solvents, salt concentration, lipid composition, cofactors, and metals have been determined, the force eld should be chosen to continue with the MD simulation.
3.2 Force Fields: General Concept and Relevant Choices
While in classical physics, force elds evaluate a systems potential energy, the fundamental distinction between the two in molecular modelling is that the energy landscape is depicted as an energy gradient dispersed across the particle positions [76]. Reconstructing a realistic simulation of a molecular system on an atomistic level is strongly dependent on force eld parameters, which are the core for deriving meaningful structural information and relative energies from MD simulation.
To date, additive or non-polarisable force elds are the most prevalent in small molecule drug design, frequently sharing the basic potential energy function and the parameters comprising the function itself [49]. The name additiveoriginates from Coulombs equation for electrost atic interactions, which states that a systems potential energy is the sum of all atom–atom individual interactions (Eq. 8.1) [49]. The potential functions, which are essentially an array of equations representing the potential energy, and its components, form the basis of the force eld core. The parameters in this set of equations are another component distinguishing a force eld.
In the vast majority of condensed-phase simulations, the total potential energy of a cohort of molecules is determined as a sum of inter- and intra-molecular interaction energies between all components of the system (Eq. 8.1)[50, 51].
EX
=
a < b
Eab þ
The total potential energy of a set of constituents ab with coordinates X
int
E
a
a
ð8:1Þ
Composed by the bonded terms (rst term) and nonbonded interactions (second term), ab refers to a set of molecules and/or ions. This notation applies to two exemplary particles a and b, which in a real calculation condition would be expanded to multiple pairs of particles.
.
208 E. Shevchenko et al.
The following terms commonly describe the intramolecular potential energy: harmonic bond stretching (Eq. 8.2), angle bending (Eq. 8.3), torsional angle deni­tion as Fourier series (Eq. 8.4), Coulomb electrostatics and Lennard-Jones potential (Eq. 8.5)[51].
E
bond
=ik
ðÞ
b,iri
- r
2
0,i
ð8:2Þ
The rst force eld term denes bond stretching energies in harmonic (ideal) conditions, where k
is the bond-stretching constant, which regul ates the rigidity of
b
the bond spring.
E
bend
=ik
ðÞ
ϑ,iϑi
The second force eld term denes the angle bending energy, where k
- ϑ
2
0,i
ϑ
ð8:3Þ
is the
angle bending const ant, determining the rigidity of the springs angle.
E
torsion
ðÞ
=
1,i
i
2
V
1 - cos 2φ
ðÞ
2,i
i
þ
2
V
1 þ cos 3φ
ðÞ
3,i
i
þ
2
i
ð8:4Þ
V
1 þ cos φ
The third force eld term represents the torsional energies as a Fourier series.
2
E
=
nb
i < j
qiqje
r
þ 4ε
ij
12
σ
ij
ij
r
ij
6
σ
ij
-
r
ij
ð8:5Þ
The fourth force eld term represents the nonbonde d energy between all atom pairs, where σ
is the equilibrium distance and rijis the distance between interacting
ij
atoms.
Force elds can be roughly divided into three major classes. One of the dening characteristics of class one is the utilisation of harmonic movements to depict bond stretching and angle bending. According to the assumption, the amount of restoring force is proportional to the displacement from the equilibrium position for class one force elds [52]. The approximation in class one force elds is referred to as quadratic because the square of the displacement energy is linearly associated with the energy of the harmonic oscillator [53]. Moreover, the parametrisation of bond stretching and the angle bending often approach harmonic behaviour only close to the equilibrium. The most famous examples of the class one force elds are the Optimised Potentials for Liquid Simulations (OPLS) [54], AMBER [55, 56], CHARMM [57], and GROMOS [58]. One of the rst techniques to be established with constant parameter optimisation for the propagation of thermodynamic charac­teristics in the liquid state applied to small molecules was OPLS, which stands for optimised potentials for liquid simulations [54]. On the OPLS core, the next gener­ation of force elds was built: OPLS3 [28], OPLS3e
29,
and OPLS4 [30] with the
maintenance of the nonbonded parameters. In general, AMBER refers to a collection
8 Drug Design in Motion: Concepts and Applications of Classical... 209
of force elds that can be split according to the simulated biomolecular system. For instance, for protein simulations, AMBER suggests ff19SB [59], and its evolution ff19SB-ILDN (with improved side-chain torsion potentials [60]); for lipids or complex membrane simulationsLIPID21 [61]; q and for nucleic acidsDNA OL15 [62], and RNA OL3 [63]. As AMBER, CHARMM also represents a set of force elds, for instance, the all-atom CHARMM22 [64] and extended atom force eld CHARMM19 [65], which are annually updated (see https://www.
academiccharmm.org/program/versions for a current list, accessed on February
2024). Last but not least, the very popular GROMOS force-elds (not to be confused with the GROMOS software), which are united atom force elds, i.e., without explicit aliphatic (non-polar) hydrogens and are considered an all purposedset, with accurate parameters for phosphorylation and others post-translational modi­cations (see [66] for the most recent version and https://www.gromos.net/#Reif2012, for current list and modications, accessed on February 2024).
Class two provides additional anharmonic cubic and quartic components for the potential energy of bonds, resulting in more detailed geometrical modelling of the vibrations of bonds. In addition, these force elds include cross-terms that describe the interactions between nearby located angles and dihedrals. Class two force elds include MMFF94 (Merck Molecular Force Field [67]), which parameters are pri­marily derived from quantum calculations rather than experimental data [68].
Another example is UFF [69] (Universal Force Field), whose original application was somewhat restricted due to the parameters, particularly for metals and inorganic substances [41]. The UFF initiative seed the idea recent ly followed up, in 2018, by the Open Force Field consortium (OFF, https://docs.openforceeld.org/en/latest/, accessed on February 2024). On their own words, they aim to develop automated and systematic data-driven techniques to parameterise and assess new generations of more accurate force elds. Despite their efforts, the most recent benchmarking shows that public force elds (i.e., OpenFF Parsley and Sage, GAFF and CGenFF) had comparable accuracy, while OPLS3e was signicantly more accurate [70].
Class three includes force elds that contain extended parameters applicable to organic chemistry, such as the Jahn-Teller effect or stereoelectronic effects. For example, AMOEBA [71] is a polarisable force eld that employs atomic-induced dipole to model polarisation while assuming that averaging polarisation is insuf­cient [72]. Another example is DRUDE [73], which uses non-polarisable force elds to leverage atom-to-atom Coulomb electrostatic interactions as its core while inte­grating polarisation effects via NAMD and a dual-Langevin thermostat approach [74].
Important to highlight another class of force elds, known as coarse-grained force elds, which employ a distinct strategy in molecular dynamics simulations. The idea behind the coarse-grained approach is the reduction of the number of degrees of freedom within a system. This is achieved by parametrising the most signicant interactions with the force eld while representing a particular set of atoms as a single bead. The denition of the most signicant interactions might be intricate depending on the parametrisation method, hence tabulated potentials are frequently employed. The purpose of coarse-grained models is to replicate specic
210 E. Shevchenko et al.
characteristics of a given system, which can encompass an atomistic protein model or experimental data. The properties one intends to replicate in the model determine the classication of the coarse-grained force elds. For instance, free energy con­servation is the focus of the MARTINI [40] force eld and the simplex method [73]. Another example is inverse Monte Carlo with structure-based coarse-grained modelling, emphasising the radial distribution [75].
In the case no parameters are available for a molecule in that particular force eld, one can resource to reparameterisation or development of new parameters, which will be force eld compatible. In order to ensure reproducibility, though, the new data would need to be validated against experimental measurements or high-level QM data (for more information on quantum mechanics, see Chap. 10). It is also important to provide a description of bonded and nonbonded potential parame ters, by atom and provide the topology(-ies) used for that determination. Software tools such as the Antechamber (from the Amber package [76]) or web server interfaces like LigParGen (for OPLS force-elds [77]) facilitate the topology analyses, writing the input for the correspondi ng simulation engine [78].
3.3 The Concept of Replicas and Timescale
One of the essential aspects that can be carefully considered when the MD simula­tion is conducted is the timescale, as different types of protein motions occur at distinct timescales. For instance, the side-chain rotamer movements can be observed in the range of ps to μs, followed by the loop motions in the eld of ns to μs, with more signicant domain movements starting to be evident in μs + timescale [79]. Thus, one should consider the reasonable timescale of the MD simulation in accordance with the movements to be observed to deliver valuable results (Fig. 8.3).
To obtain meaningful and comprehensive results from MD simulations, one should carefully consider one of the essential aspects of this methodology. The selection of an appropriate timescale directly inuences the range of molecular motions and events that can be captured within the simulated system. Shorter timescales are suitable for studying fast processes such as bond vibrations and side-chain rotamer movements (fs to ns). In contrast, longer timescales are required to investigate slower and more complex events, such as loop motions (ns to μs), signicant domain movements and large-scale conformational transitions (μs+). However, excessively long simulations conducted without a clear research purpose, may not only become computationally expensive but also lead to diminishing returns in terms of new insights. To strike a balance between accuracy, computational feasibility and deriving meaningful results, one should align the simulation timescale requirements with the specic research question to be answered through MD .
The current model for understanding protein folding and dynamics asks us to imagine the protein dynamics process as a large funnel, in which the surface would represent the diversity of protein conformation, and its deepness would be propor­tional to the energy systems. The deeper pockets of this funnel constitute global
8 Drug Design in Motion: Concepts and Applications of Classical... 211
Fig. 8.3 A spectrum of dynamic processes in proteins across timescale occurring in molecular dynamics simulation. (Adapted from Henzler­Wildman et al. [79], dynamic personalities of protein [79])
low-energy protein conformations. This funnel is rather wrinkly, with smaller wells, which represent the energy minima of different protein conformations. What the MD simulation does in this model is to navigate these inner surfaces and unbiasedly sample around the protein conformational space.
Considering the funnel model and the problem of timescale for protein dynamics, we understand that classical unbiased MD simulations might not be always suitable to sample molecular events involving a transition between those energy barriers, such as large conformational changes. Alternatively, given enough justication and experimental evidence, one can apply enhanced sampling simulations, such as nonadaptive biasing potential methods (e.g., Gaussian-accelerated MD), adaptive bias simulations (e.g., metadynamics), replica exchange methods, etc. Examples of experimental evidence can rise from (TR-)FRET assays or site-direct mutagenesis in combination with photocrosslinking, showing that two portions of the protein are in close proximity, multiple diverse crystal structures, showing different conformations and even highlighting large missingportions due to high exibility.
If enhanced sampling simulations are used, the choices for sampling methodol­ogy, parameters, and, most importantly, the convergence criteria (where to stop sampling and when to consider that the generated conformations are articial) should be described. We refer to the recently published guidelines for reporting MD simulation data [80].
The other crucial point for comprehensive MD simulation is the number of replicas of the single system to be conducted to avoid false positive conclusions [81]. The term replicarefers to the simulations of the identical system, sharing the same number of atoms, initial structure, and preparation protocol, repeated several
212 E. Shevchenko et al.
times. Given the principle of energy conservation within the ensemble required for the MD calculations and the almost derisory effects arising from changes in the oating-point precision or hardware, multiple runs of these syst ems would lead to very similar thermodynamical and, consequently, conformational properties. In this sense, the difference between replicas is in the initial velocities generated randomly according to the Maxwell distribution [81]. As a result, the velocities are unique for each replica, which leads to the production of different simulation trajectories. However, even replicas with identical velocities can produce distinct trajectories for various reasons, including machi ne-specic settings and the specication of the compiling [81]. In theory, the ergodic principle claims that the velocities have no impact on innite dynamic simulations [82 ]. While molecular dynamics are not run eternally in real life, random velocities are strictly important, especially for short­scale MDs. The random initial velocity assignment ensures that the results of every simulation are slightly different, even if the other settings are identical. In other words, the random velocities provide the opportunity to observe real-world phe­nomena happening with the same system at different time points. Moreover, multiple replicas ensure a specic movement or interaction observed not in a single simula­tion but in a statist ically signicant number of replicas is not biased or articially generated but related to the real-world evidence. The choice between the multiple simulation replicas over a single but long-scale one is frequently a question that should be answered before the initiation of the MD according to the aims and resources of the project.
The work from Knapp et al. [81] performed 100 identically parametrised replicas of 3 μs each for a small 10 amino acid long peptide, and also 100x replicas of 100 ns each for the T-cell receptor/MHC system (827 amino acids). Using those simulations as show-case, they compared randomly chosen subgroups within these replica sets, being able to estimate the reproducibility and reliability that could be achieved by a given number of replicas at a given simulation time. They highlight two major points: i) conclusions drawn from single simulations are not reproducible and ii) observations drawn from multiple shorter replicas were more reliable than using a single longer trajectory. Based on their data, a minimum of 5–10 replicas were required [81]. Their comment that the actual number of replicas needed will depend on the level of reliability sought and, herein we add that is highly dependent on the size and intrinsic exibility of the system.
Our next example applied multiple short sim ulations on the conformational space of aptamers [83]. They combine the enhanced sample approach of starting from high-temperature equilibration runs (six in total), leading to higher exibi lity, with independent replicas from the output of each of these systems (10 × 100 ns per output). They analysed the recurrence rate of found conformations and properties among the different starting points, showing a clear dependence on initial confor­mation. This highlights the need of using different initial congurations as simula­tion starting points to avoid being trapped in energy minima. Especially, since analyses of convergence of those unbiased trajectories might not allow the detection of slow trans itions between kinetically trapped met astable states.