Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5606_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
02.09.2026
Размер:
21 Мб
Скачать
Chapter 9
Conformational Sampling of Proteins: Methods for Simulate Protein Plasticity and Ensemble Docking
Ana Ligia Scott, Simone Queiroz Pantaleão, and Eric Allison Philot
Abstract The understanding of protein function by emphasizing the critical role of
structural dynamics, or protein plasticity, in biological processes is a eld of knowledge in constant evolution. It highlights that proteins are not static entities but exhibit intrinsic, nonrandom movements crucial for their interactions with other molecules. Advances in cryoelectron microscopy (cryo-EM) and crystallography have provided snapshots of proteins in various conformations, revealing the dynamic nature of macromolecules. Different mechanisms of molecular recognition, such as conformational selection and induced t, could be simulated by using hybrid methods that integrate molecular dynamics and normal mode analysis to study protein conformational spaces. These approaches have signicant implications for drug discovery, as they allow for more accurate modeling of protein–ligand inter­actions by considering the inherent exibility of protein structures. Methods such as MDeNM, which integrates normal modes into molecular dynamics simulations, is highlighted for its ability to promote signicant conformational changes, facilitating the exible tting of atomic structures into cryo-EM maps and revealing complex protein motions. Collective molecular dynamics (coMD) and ClustENM are also discussed, emphasizing their roles in conformational sampling and protein–ligand interactions. Ensemble docking strategies, including the recent essential dynamic ensemble docking (EDED) protocol, are reviewed for their potential in drug design, particularly in improvi ng the accuracy of molecular docking by considering the exibility of both ligands and target proteins. These methodologies, combined with advancements in force elds and postdocking techniques, contribute to a more comprehensive understanding and prediction of biomolecular interactions.
A. L. Scott · S. Q. Pantaleão () Center for Mathematics, Computing and Cognition, Federal University of ABC, Santo André, São Paulo, Brazil e-mail: simone.qrzp@gmail.com
E. A. Philot Educational and Research Institute, Molecular Oncology Research Center, Barretos Cancer Hospital, Barretos, São Paulo, Brazil
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2024 V. G. Maltarollo (ed.), Computer-Aided and Machine Learning-Driven Drug Design, Computer-Aided Drug Discovery and Design 3,
https://doi.org/10.1007/978-3-031-76718-0_9
243
244 A. L. Scott et al.
Keywords Protein plasticity · Conformational sampling · Normal modes · Ensemble docking

1 Introduction

For a long time, the idea that the protein structure de nes its function has been well accepted. In recent decades, this concept has been adjusted to include a new element: protein structural dynamics or plasticity. This chapter seeks to discuss the impor­tance of protein local and global movements and thei r relevance for interaction with other macromolecules or small molecules.
Proteins are not static entities; they have nonrandom intrinsic movements that are extremely important for their function, often because of the interaction with different binding partners. Recent developments in cryoelectron microscopy (cryo-EM) and crystallography techniques have revealed multiple snapshots of increasingly large and exible systems [51]. Cryo-EM has increasingly been used in structural biology to characterize not only a single structure but an ensemble of conformations, often relevant to biological function, particularly for supramolecular assembly [8, 48 ]. Thanks to radically new technological advances in both microscope hardware and image processing software, termed the resolution revolution, cryo-EM has dramat­ically improved in the last years and it can now capture information about biomo­lecular structure in multiple states to near-atomic resolution [44]. Macromolecules adopt many conformations in solution depending on their structure and shape, which determine their dynamics and function. These conformations exhibit unique structure-encoded dynamic properties that have a profound inuence on their bio­logical function [5, 58]. Spectacular progress has been achieved in cryogenic electron microscopy (cryo-EM) through technological advances in both microscope hardware and image processing software (the so-called resolution revolution). Cryo­EM can now capture information about biomolecular structures in multiple states to near-atomic resolution [44]. Not only has the number of structures solved by Cryo­EM increased steadily in recent years, but it has also allowed the characterization of ensembles of conformations.
Macromolecules can adopt multiple conformations in solution depending on their structure and shape, which are directly linked to their function. These conformations exhibit unique structure-encoded dynamic properties that have a profound inuence on their biological function [5, 58]. Several studies have indicated that these dynamic properties are primarily determined by the topology of native contacts [36].
In the last decade, a series of studies established the role of structural dynamics, also called intrinsic dynamics, facilitating, if not driving, the interactions, and functioning of biomolecular systems in the cell [5 , 51, 58]. Therefore, it is crucial to explore different conformations to model dynamical processes and determine stable states [71]. Such studies require both experimental and computational efforts. In this context, theoretical methods are becoming essential for understanding how different conformers interconvert to mediate biological function due to the accumu­lated structural information. A large number of studies based on coarse-grained
9 Conformational Sampling of Proteins: Methods for Simulate... 245
Fig. 9.1 Molecular recognition mechanisms
models (such as the Elastic Network Model, ENM) or full atomic normal mode analysis (NMA) have been applied to investigate macromolecular structures and the molecular mecha nisms of their conformational changes [36]. Integrated methods that combine NMA (using ENM or complete atomic models), molecular dynamics, and experimental data from biophysical techniques, such as X-ray crystallography, are useful for studying the structural dynamics of macromolecules.
Related to molecular recognition mechanisms, it is necessary to better understand whether the interaction of proteins with binding partners occurs through possibili­ties: (i) conformational selection; (ii) induced adjustment, or even (iii) a combination of both. Figure 9.1 illustrates these three situations. In the rst, the target visits several conformations energetically acce ssible by intrinsic motions and the ligand chooses the best (Fig. 9.1a). In the induced adjustment, the ligand binding to the target and provokes a conformational change to a lower energy level (Fig. 9.1b). The third consists of the combinations of both: there is a conformational change before the ligand binding to DNA after, as illustrated in Fig. 9.1c.
There are several examples illustrating each of these three cases and it is impor­tant to consider their role in the molecular interaction to improve the drug discovery methods/software. In several cases, it is very important to consider this targets plasticity when we perform in silico studies, such as molecular docking (please, see Chap. 7). In recent years, the importance of considering tting effects in molecular docking calculations was widely recognized in the molecular modeling community. Although small-scale protein side chain movements are now accounted for in many high-end docking strategies, the explicit modeling of large-scale protein movements is still a challenging task. Ensemble-based methods have been introduced taking into account several protein conformations in Drug Design and/or Discovery. Some software couples a technique to sample the target with its docking algorithm.
We can use two distinct strategies to perform an ensemble docking (simulating the conformational selection): (a) use software that integrates sampling and docking algorithms and (b) use a software/method to sample the conformations for the Target
246 A. L. Scott et al.
and/or the ligand as a rst step and after use software to simulate the docking of several conformations. This second strategy allows the use of methods more ef­ciently to generate the conformers for the target and/or ligand. In this case, we can use several different methods: Molecular Dynamics (please, see Chap. 8), Metadynamics, and Hybrid methods that combine Elastic Network Models or Normal Mode all atoms with methods such as Molecular Dynamics, Monte Carlo, and Metadynamics. We should remember that several times it is important to simulate the adjustment induced after the complexes are formed. There are several examples of protein–ligand where this kind of molecular motion is important to make the complexes more stabl e or facilitate the catalyses of substrates.
MDeNM coupled with ensemble docking helped in elucidating the mechanisms guiding the recognition of substrates and inhibitors by the sulfotransferase [21]. In the next section, we discuss some of these hybrid methods and applications to simulate ensemble docking protein–ligand and protein–protein.
2 Hybrid Methods to Sample the Protein Conformational
Space
The protein plasticity, corresponding to the collective and noncollective motions, has become important in understanding the relationship between sequence, structure and function; and consequently, various aspects of biological processes such as [36]:
(i) The interaction of proteins with other macromolecules or with small molecules.
(ii) The molecular mechanisms involved in metabolic cycles.
(iii) The effect of mutations on the structure and function of proteins and others.
These protein movements can be classied according to different aspects. Yang and collaborators classify protein movements [68] into three classes: (1) local movements of protein fragments; (2) domain movements; and (3) movements involving more than one unit.
In another way, Bahars group proposed the following classication [41]:
(i) Evolution of motions in the global (the three motions with lower frequency).
(ii) Low-frequency motions (from mode 4 until 20). (iii) Low-to-intermedi ate frequency (from mode 21 to 60). (iv) High-frequency (modes >60).
This classication is very interesting and useful for some kinds of analysis [41]. But in this text, we will use a generic classication using the nomenclature: collec­tive and noncollective movements.
Collective Movements involve a considerable part of the molecule moving in a coordinated manner, leading to a global conformational change. For example, the ordered collective movements of one domain or more than one domain [6, 62]. In the case of involving more than one domain, we can imagine that this movement is
9 Conformational Sampling of Proteins: Methods for Simulate... 247
composed of several collective (coordinated) movements, independent or not [37]. The collective one can be considered the rst 20 motions.
Noncollective movements involve a small portion of the protein generally the result of a stimulus. As an example, we can mention an allosteric conformational change due to a small ligand binding. Another example that can be cited involves secondary structure transitions that result in local folding or unfolding (change in secondary structure) of the molecule, as can occur in the case of the Prion protein [22, 42].
Despite the great advances in scalable codes, graphics processing units (GPUs), and parallelization of simulation algorithms [51, 53, 56] allow us to simulate increasingly larger systems and therefore longer biological times such as entire bacterial cytoplasm in the submicrosecond range [69]. Still, for most proteins, these time scales cover a small part of the structural picture, and longer simulations are only accessible with special-purpose supercomputers like Anton (https://www.
psc.edu/resources/anton-3/)[20, 63]. In addition to these technical aspects, there is a
fundamental sampling problem.
Several studies are showing that the way conguration spaces are sampled can be more critical than the simulation length. Normal Modes calculation using all-atom models or coarse-grained methods such as Elastic Network Models can predict with impressive accuracy, just from the general shape of a protein, not only the experi­mentally observed conformational changes but also entire sequences of intermedi­ates structures [52]. The hybrid approaches allow for efcient conformational space sampling, considering both slow and fast movements [36]. In recent years, several hybrid methods have been developed to take into account local and global move­ments in conformational sampling such as MDeNM, CoMD, Dynomics, VMOD, and others [12, 36].
These methods have been used for many purposes [15, 23], such as: (i) efciently sampling the conformational space of macromolecules; (ii) rening biomolecular complexes and assemblies; (iii) integrating experimental data for improving the determination of conformational ensembles; and (iv) conducting biomedical appli­cations such as the study of allosteric movements, docking, and the impact of mutations.
Efcient conformational sampling is essential to identify stable structures as well as transition states of macromolecules and to calculate free energy differences [7]. Kaynak et al. [35] provided a comparative analysis of such hybrid methods developed for conformational space efcient sampling and the possible transitions between functional states [35]. They compared four methods: ClustENM [39], its recent extension ClustENMD [35], collective MD (CoMD) [27], and MD with excited NMs (MDeNM) [12]. The rst three hybrid methods combine Elastic Network Models with minimizations and/or molecular dynamics to sample the conformational space of macromolecules or macromolecular complexes. The last method, MDeNM, combines normal mode calculation based on a force eld (either with all atoms or just their alpha carbons) with MD to sample the conformers. In this chapter, the authors show that these methods can be more efcient than classical molecular dynamics, sampling some open, closed and intermediate states of
248 A. L. Scott et al.
Fig. 9.2 Key combinations of methods for protein structural studies of proteins. Hybrid methods may combine elastic network model (ENM)-based or force-eld (FF)-based or anisotropic network models NMA with various kinds of simulations such as Molecular Dynamics (MD) simulations, Monte Carlo and can improve the heterogeneity conformational of proteins
macromolecules. Another hybrid method that considers all-atom normal mode calculation is VMOD. The main methods present in hybrid approaches that use normal modes are summarized in Fig. 9.2.
2.1 MDeNM
The method MD with excited Normal Modes (MDeNM) was developed to promote large conformational changes during molecular dynamics simulations by incorpo­rating normal modes. This technique is very useful for obtaining conformational changes that are rarely achieved by standard MD simulations unless excessively long simulations are carried out [12]. The MDeNM approach consists of a multireplica protocol designed to enhance conformational exploration in a subspace dened by a set of low-frequency normal modes, considering the dockings with the localized motions occurring within the Cartesian space [14]. A proof-of-concept study has shown that MDeNM can be a very efcient tool for performing exible tting of atomic structures into cryo-EM maps [13]. Each replica starts from an equilibrated
9 Conformational Sampling of Proteins: Methods for Simulate... 249
initial structure whose normal modes are calculated, and the system is driven dynamically along a direction obtained by a linear combination of selected lowest­frequency modes. An efcient conformational sampling is essential for identifying stable, transient, and transition states of macromolecules and calculating free energy differences.
A recent MDeNM new version updates the direction used to move the system during simulations, taking into account the stru ctural and energetic constraints imposed by the system itself and the environment, which allows the system to explore new paths, especially considering complex movements [14]. Without the need to recalculate the NMs like other methods (such as ClustENM, ACM, ACM-PCA, NMA-ITS), the advantage of this approach is to describe alternative paths in the conformational landscape using only a few initial modes as input.
The adaptive MDeNM (aMDeNM[59]) has results compatible with several available methods when simulating simple large-scale movements such as a hinge or twisting movements while showing better performances considering complex move­ments involving reorganizations between protein domains [35]. Quiroz et al. com­bined the use of MDeNM with the Markov ModelPyEMMA [57] to study the effects of phosphorylation on the meta-state transition of the Human Dopamine Transporter (hDAT) and obtained relevant results [57].
Duda et al. used MDeNM to elucidate molecular mechanisms that guide the recognition of diverse substrates and inhibitors by SULT1A1. The hybrid method allowed them to explore the conformational space in a broad way for SULT1A1 in complex and to clarify a littl e more the mechanisms of binding of the protein to the substrate and the inhibitor [21 ].
2.2 Collective Molecular Dynamics (coMD)
The coMD method uses an Anisotropic Network Model (ANM) to guide the molecular dynamics for the system. ANM modes are selected with the help of a Monte Carlo (MC)/Metropolis scheme that allows the system to occasionally diverge from the shortest path and bypass energy barriers. This allows movements not to be limited to proceeding along a subset of low-frequency modes, which is quite interesting and may include a stochastic factor in the generation of conformations.
The use of Molecular Dynamics with a complete model of atoms at the step of the ANM allows us to generate coordinates of perfectly compacted side chain atoms and avoid nonphysical distortions in the geometry of the protein: As the ANM is dened exclusively by the contact topology between residues, changes in geometry affect the adaptively generated ANM modes in the next cycle of the simulations. There­fore, the basic approach is to deform the structure collectively along the modes predicted by the anisotropic network model, upon selecting them via a Monte Carlo/ Metropolis algorithm from among the complete pool of all accessible modes. The
250 A. L. Scott et al.
CoMD method was proposed by [27] and implemented in the package Prody developed by Bahars group [71, 72].
2.3 ClustENM and ClustENMD
ClustENM [39] is a fully automated conformational sampling method composed of multiple generations/cycles consisting of each of the following steps: (1) generating conformers by deformation along global NMs calculating using the lattice model anisotropic (ANM) [4], (2) clustering of generated conformers and (3) relaxation of cluster representatives. ANM modes are updated for each parent conformer in each generation. A series of deformations along the ± directions of some global modes (3–5) are performed, occurring a specic root mean square deviation (RMSD) for deformation (Step 1). Representative conformers selected to calculate the next generation of conformers are relaxed by energy minimization (EM) in implicit solvent (Step 3). Kurkcuoglu and Bovin discussed (pre and postdocking sampling of conformational changes using ClustENM and HADDOCK for protein–protein and protein–DNA). They applied ClustENM combined with HADDOCK software to perform pre and postdocking sampling to incorporate the effect of plasticity in protein–protein and protein–DNA complexes [28]. The authors concluded that conformational selection (presampling and docking) has the greatest impact on the quality of the nal models especially in the case of protein–protein complexes. They also observed that for protein–DNA complexes, in contrast, the induced tuning stage of the pipeline (postsampling) signicantly improved the quality of the nal product models.

3 Ensemble Docking

Ensemble docking is a strategy that carries multiple independent molecular docking calculations (please, see Chap. 7), each of them using different target conformations, either from side-chains or larger backbone variation. This ensemble mimics the macromolecule movements, despite using individual rigid conformations. It is possible to discuss Ensemble Docking in two distinct contexts but similar in some aspects: treating problems involving complex of protein–ligand; and protein–protein interaction. These problems are important for the problem of Drug Design and/or Discovery.
Currently, protein–protein dockings also have an important guideline for success, which consists of understanding the modulation of protein–protein interactions, i.e., understanding which are primordial and which are advisory, even acting in different regions, with allosteric mechanisms. Facts like this can be observed through the stabilization of exible loops by extensive sampling of the protein’s conformational space.
9 Conformational Sampling of Proteins: Methods for Simulate... 251
Molecular dynamics have aspects that need to be constantly observed for biolog­ical systems, from adequate force elds to realistic sampling. In the study by [3], the author also mentions that an improvement in the force elds applied with the inclusion of electronic polarization together with quantum mechanics simulations also helps to provide good informat ion about the biological system.
A computer program for coupling exible ligands to exible proteins using a genetic algorithm, FITTED 1.0, was developed by [11]. This tool includes the ProCESS and SMART modules which correspond to the protein/ligand congura­tion steps. Its validation was obtained by molecular docking of biological systems involved with protein inhibitors such as HIV-1 protease, among others.
There are two very distinct strategic approaches for Drug Design, the rst consists of structure-based drug design (SBDD) and the other of ligand-based drug design (LBDD). Both can be applied to analyze the properties and characteristics of bioactive ligands. In the following session we discuss the main differences between their uses and applicati ons in the docking of molecules.
Ensemble docking represents a molecular docking strategy in which the rst step is to generate an ensemble of structural conformations of a biological target, usually obtained by molecular dynamics simulations, in which candidate molecules are subsequently docked. More than the number of coupli ngs/target structures used, what determines the success of the approach is the selection of which structures are the most representative of the system of interest. When information such as the exibility and thermal uctuation of the atoms is taken into account, it is possible to explore conformational substates of these complexes that may correspond to impor­tant specic moments in the shape of the binding pockets. The need for more realistic simulations has led to the development of studies using a combination of techniques, which form the basis of ensemble docking. As reported by [3], the computational design of ligands based on the structure has been strategically improving since 1970. With computational advances, it has become possible to carry out more complete simulations, considering a greater number of structures with diverse states in terms of exibility and the chemical space of the ligands. The author points out that some historical events have marked the eld, such as in 1999 when work began to consider the exibility of ligand pharmacophores. The main key points of the ensemble docking technique are illustrated in Fig. 9.3.
McKay et al. [43] proposed a new proposal for ensemble coupling, known as essential dynamic ensemble docking (EDED), the authors proposed this protocol for agonists of the PAC1 receptor, a class B GPCR, where three million compounds from the ZINC database were screened, with experimental validation of 23 com­pounds for PAC1 targeting. Its differential is that this technique can reduce the number of false negatives in screening and improve precision in the search for bioactive molecules. In this case, an important point is to obtain and select chemi­cally signicant receptor models for ligand docking. The technique consists of parameterizing the ligand with an OPLS3 force eld. The trajectory frames with Principal Components (PCAs) values closest to the centers by RMSD measurement were adopted as receptors for the ensemble docking.
252 A. L. Scott et al.
Fig. 9.3 Key points of the ensemble docking technique
The difference with other techniques is that this approach allowed us to arrive at a minimum of relevant receptors in the region of the binding pocket, as opposed to other techniques which generate many conformations of protein structures but not all of them relevant to the region in which the ligand will act, demonstrating that it is also possible to arrive at good docking results using a homology model.
Authors who generally follow the LBDD line also involve molecular activities, but related to other techniques, such as QSAR models with subsequent steps of molecular dynamics and study of postdocking normal modes. Two dist inct lines can be interesting depending on the focus of each research. In the LBDD approach, the focus on knowledge generation seeks to elucidate the relationship between the structure and physicochemical attributes of molecules with their biological activity. The structure– activity relationship (SAR) considers pharmacophore models of the ligands that identify the main characteristics responsible for the activity, the quan­titative structure–activity relationships (QSAR), providing quantitative predictions of activity, similarity, and combinations of chemical descriptors. In the work of Shim et al. [64], it is mentioned that 2D-QSAR is based on the calculation of physico­chemical, electronic, topological and shape descriptors and is not closely related to the molecular docking technique. These are models that require preprocessing, the values are normalized and correlated, or redundant descriptors can be eliminated using the partial least squares (PLS) technique or principal component analysis (PCA), which require cross-validation, randomization or a set of external tests. The author, [64], also mentions other types of QSARs, now related to molecular docking, as they use 3D descriptors in bioactive conformations, such as in compar­ative molecular eld analysis (CoMFA) and comparative molecular similarity indi­ces (CoMSIA), which depend on the chemical and spatial characteristics of the molecules, such as molecular volume, surface area, solvation ΔG, dipole moments, HOMO and LUMO, based on the molecular interaction eld in a 3D grid around the molecules with polar probes. The molecular docking step is used to align the ligands in their bioactive conformation. Statistical methods such as PLS or PCA are applied and related to biological activities. There is a type of model, known as 4D-QSAR, which takes exibility into account and uses multiple conformations. 5D-QSAR, on the other hand, builds a pseudoreceptor based on information from the ligand to