Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5364_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Foreword
- •Acknowledgments
- •Contents
- •1.1 Structure-Based Drug Discovery (SBDD)
- •1.2 Ligand-Based Drug Design (LBDD)
- •1.3 Echoes from the Past, Visions from the Future
- •References
- •1 Introduction
- •2.2 Second Step: Data Curation
- •2.4 Fourth Step: Updating and Maintenance
- •2 Databases and Curation
- •8 Perspectives
- •9 Conclusion
- •References
- •1 Introduction
- •2.1 Making and Matching Protein Models
- •2.2 Simulating Protein Movements
- •2.3 Analyzing Changes in Protein Shape
- •3 Pharmacogenomics in Drug Development
- •4 Case Studies of Genomics-Based Drug Design
- •References
- •1 Historical Background
- •1.1 Timeline
- •2 Methodology Overview
- •2.1 Neural Networks
- •2.1.1 Perceptron
- •2.1.2 Multilayer Neural Networks
- •2.1.3 Types of Neural Networks
- •Feedforward
- •Recurrent Neural Networks
- •LSTM
- •2.2 Deep Learning
- •3 Using Machine Learning
- •3.2 Data Collection
- •3.3 Data Preprocessing
- •3.4 Model Selection
- •3.5 Model Training
- •3.6 Validation
- •3.7 Tuning
- •3.8 Prediction
- •4 Limitations
- •4.1 Bias
- •4.3 Interpretability
- •4.4 Computational Cost
- •4.5 Data Dependency
- •4.6 Robustness
- •5 Applications in Drug Discovery
- •5.2 Lead Discovery
- •5.3 Preclinical and Clinical Development
- •6 Resources and Tools
- •7 Challenges and Perspectives
- •7.1 Future Trends
- •9 Conclusions
- •References
- •1 Historical Background
- •1.1 Applications in Drug Discovery
- •2 Validations and Controls
- •2.1 Internal Validation
- •2.2 External Validation
- •2.3 Relative Cluster Validation
- •3 Challenges and Perspectives
- •4 Conclusions
- •References
- •1 Historical Background
- •2 OECD Principles
- •2.1 A Defined Endpoint
- •2.2 An Unambiguous Algorithm
- •2.5 A Mechanistic Interpretation, if Possible
- •3 Software and Tools
- •4 Validations and Controls
- •4.1 Internal and External Validation
- •4.1.1 Regression Metrics
- •4.2 Applicability Domain
- •4.3 Randomization Tests
- •5 Interpretation
- •6 Practical Advice During QSAR Modeling
- •7 Application
- •8 Challenges and Perspectives
- •References
- •1 Molecular Docking
- •2 Advances in Scoring Functions and Search Algorithms
- •2.2 Critical Characteristics of Search Algorithms
- •2.3 Docking Programs and Scoring Functions
- •3 Calculations Performed During Docking Simulations
- •4 Essential Components for a Good Docking Program
- •5 Limitations of the Docking Technique
- •6 Validation of Docking Results
- •7 Inappropriate Use of Validation Methods in Docking
- •9 Use of Machine Learning in Molecular Docking
- •11 Challenges
- •12 Conclusions
- •References
- •3 System Preparation for MD Simulations
- •3.1 Solvation and Microensemble
- •3.2 Force Fields: General Concept and Relevant Choices
- •3.3 The Concept of Replicas and Timescale
- •4.1.2 Protein Root Mean Square Fluctuation (RMSF)
- •4.1.4 Protein Secondary Structure Analysis
- •4.1.5 Principal component Analysis (PCA)
- •4.1.6 Markov State Modelling
- •4.1.7 Distance Calculations
- •4.1.8 Angle and Plane Calculations
- •4.2.2 Distances and Ligand-Induced Geometry Rearrangements
- •4 Molecular Dynamics Analysis
- •4.1 Protein Perspective
- •4.1.1 Protein Root Mean Square Deviation (RMSD)
- •4.3 Ligand Perspective
- •4.3.1 Ligand Properties
- •4.3.2 Ligand Root Mean Square Deviation
- •4.3.3 Ligand Root Mean Square Fluctuation
- •4.3.4 Angles and Dihedrals
- •5.1 Protein Structure Prediction and Preparation
- •5.2 Molecular Docking
- •6 Concluding Remarks and Outlook
- •Glossary
- •References
- •1 Introduction
- •2.1 MDeNM
- •2.2 Collective Molecular Dynamics (coMD)
- •2.3 ClustENM and ClustENMD
- •3 Ensemble Docking
- •References
- •1 Introduction
- •1.1 Advantages, Disadvantages, Innovations, and Challenges
- •1.2 Recent Advances in Accessible FEP Software Tools
- •1.3 Applications of FEP in Industry and Consortiums
- •2 Expanding the Potential of FEP Calculations
- •2.1 Validating Binding Poses
- •2.2 Dealing with Solvent
- •2.3 FEP and Allostery
- •2.4 FEP and Covalent Ligands
- •2.5 Applications of FEP in Scaffold Hopping
- •2.6 Positional Analogue Scanning
- •2.7 Combinations and Alternative Approaches
- •3 Machine Learning for FEP
- •3.4 Implications for ML in FEP Calculations
- •4 Final Considerations
- •5 First Steps to FEP Simulations
- •References
- •1 Background
- •2 Ultra-Large Screening Libraries and Chemical Spaces
- •3.1 Implications of Dataset Size
- •4 Ligands on the Ultra-Large Scale
- •4.1 Ultra-Large 2D Similarity Searches
- •7 Challenges and Future Perspectives
- •7.1 Hit Triage: An Old Problem on a New Dimension
- •8 Conclusions
- •Appendix
- •References
- •1 Introduction
- •2 Enzymatic Activity Evaluations
- •3 Cytotoxicity Evaluation and Cell Viability
- •4 Antiviral Assays in Experimental Validation
- •6 In Vivo Evaluation of Compounds
- •7 Conclusions
- •References
- •1 Introduction
- •3.1 Data Collection
- •3.2 Data Preprocessing
- •3.4 Model Choice
- •3.5 Model Training
- •3.6 Model Assessment
- •3.7 External Validation
- •3.8 Implementation and Availability
- •3.9 Continuous Update
- •5 Conclusions and Perspectives
- •References
- •1 Experimental Approaches to Obtain Protein Structure
- •1.1 X-Ray Crystallography
- •1.2 Nuclear Magnetic Resonance
- •1.3 Cryo-EM
- •1.4 Hybrid Methods
- •2 Modeling Approaches to Obtain Protein Structure
- •2.1 Homology Modeling
- •2.2 Ab Initio Modeling
- •2.3 New Approaches
- •3 Conformational Diversity of Proteins
- •3.1 Characterization of Protein Conformational States
- •3.2 Experimental Methods to Study Protein Dynamics and Conformations
- •3.4 Molecular Dynamics Simulation
- •3.5 Sampling Strategies
- •4 Remarks and Perspectives
- •References
- •1 Introduction
- •2 Structure-Based Drug Design of HIV Protease Inhibitors
- •2.1 HIV-1 Protease as a Therapeutic Target
- •2.2.1 Saquinavir
- •2.2.2 Indinavir
- •2.3.1 Lopinavir
- •2.3.2 Darunavir
- •6 Conclusions
- •References
- •4 Experimental Methods to Analyze NR Activity
- •4.2 Coregulator-Recruitment
- •5 Concluding Remarks and Outlook
- •References

9 Conformational Sampling of Proteins: Methods for Simulate... 253
diversify the locations of the grid points in order to simulate a favorable induce d fit
of the system. Compared to 6D-QSAR, this model considers solvation energy terms.
There is even another type, the 7D-QSAR, built with 3D structural data of the
biological target, from X-ray crystallography or NMR, but they cannot be considered
as an LBDD strategy because the starting point is not a specific descriptor of the
ligand.
Pre and postdocking sampling related to conformational changes in protein–
protein systems and protein–DNA systems have some very peculiar aspects and
have contributed to various studies, as illustrated in the work of [38]. Simulating the
dynamics of biological systems is challenging because it is necessary to consider the
extent and direc tion of conformational changes that actually occur, it consists of both
obtaining and selecting a series of conformations that represent the induced adjustment of the proteins studied, for this it is useful to apply combined methods capable
of providing a conformational sampling in which the extent of the deformation is
considered with its energy cost, as in the ClustENM technique—a modeling procedure based on elastic networks—using the HADDOCK software [18, 19, 67]. With
regard to molecular coupling, it is possible to carry out pre and postcoupling steps as
a strategy to produce models of adequate quality for the representability of biological
systems.
In his work, the author [38] demonstrates the applicability of effective joint
techniques, combining the generation and selection of conformational sampling
with multidomain analysis of regions of interest, resulting in more realistic molecular couplings by considering the flexibility of the region, which can be applied to
protein-protein or protein-DNA studies if the technique known as ClustENMHADDOCK is used. So far, this technique is considered to have superior performance compared to traditional rigid coupling for protein-protein systems, but inferior in terms of the flexibility of the multidomain region. As far as the protein-DNA
system is concerned, the simulations prove to be adequate in most cases.
In addition to raising discussion about the gaps in the molecular docking technique, the author [38], proposed an interesting pipeline to meet the challenges of the
area, where he combined simulations before and after complexation, containing the
following steps: (1) energy minimization of the structure subjected to elastic network
modeling, with ClustENM and deformation test to categorize the RMSD deformation; (2) generation of conformers in which clusters with parent structures are
excluded; (3) selection of 20 conformers based on the minimization energy, to
perform the coupling, whereby, the initial structure is also present regardless of the
energy value; (4) use of HADDOCK for the semi-flexible docking of structures;
(5) selection of the best scoring coupled model for the second round of ClustENM
(which is considered post-sampling); (6) verification and parameterization of
ClustENM for each structure; (7) generation of final models, obtained after applying
the technique for two more generations; (8) verification and classification of the
quality of the final models, using the HADDOCK score.
The combined use of techniques that provide simulations of biological systems
containing a longer profile, that is conformations before and after a molecular
docking, has also been applied for large-scale docking (virtual screenings), as

254 A. L. Scott et al.
demonstrated in the work of [10]. A study was conducted with the dihydroorotate
dehydrogenase (DHODH) protein, which is considered a key enzyme in the pyrimidine biosynthesis pathway and is involved in processes related to cancers, and some
autoimmune and viral diseases.
The study by [10] followed important steps that provided good results: (1) selection and preparation of 38 hDHODH structures complexed with inhibitors of
resolution <2 Å, superimposed with the ICM [1], with selection of the representative
set of structures, based on clustering by the root mean square deviation (RMSD) with
the aid of the bio3D R software [24]. The parameterization included the addition of
polar hydrogen atoms corresponding to protonation at pH = 7, using OpenBabel
2.4.0 [50]. (2) Benchmarking of datasets and evaluation metrics, formation of a
set/library using enhanced data (DUD-E) [47], with 67 active compounds and 1933
inactive compounds. The Boltzmann enhanced ROC curve (BEDROC, with α = 20)
was used and enrichment factors (EF) were calculated using the balancer [40].
The software and molecular docking parameters used were AutoDock Vina [66],
LeDock [70], rDOCK [61], and ICM [1]. These tools were chosen because of their
high overall performance, differences in the methods used to predict/score the poses
generated, and availability. The consensus stage for generating various conformations of the ligand is just as important as the richness of generating the various
conformations of the target protein used. According to the author’s data, the success
of the set fitting and consensus scoring lies precisely in the way the sets are selected
and analyzed, with different conformations, not forgetting to normalize the fitting
score. As for the consensus bond pose approach, all possible combinations of
molecular docking were evaluated, and the compound that scored well in each
docking program was considered the final candidate. The authors point out that,
whenever possible, it is recommended to work with pre-validated data and the
selection of suitable software and structures (representative of the system of interest)
to increase the performance of simple molecular docking and virtual screening.
According to [22], a tool that can improve molecular docking results is HTP
SurflexDock, which allows the simulation of the receptor’s implicit flexibility.
Considering a post-processing step that allows the reclassification of favorable
compounds by exploring the confor mational space and free energy of binding by
MM/PBSA (Molecular mechanics Poisson-Boltzmann surface area). This is a suitable tool when information on the receptor’s contact area is incomplete or specific
optimization is required for a given ligand.
Considering a representative set of conformations of the biological target before
molecular docking makes it possible to simulate the implicit flexibility of the
receptor; however, to correct the accuracy of the scoring funct ions of the compounds, a post-processing step is necessary which explores the conformational space
of the active site, generates new poses and estimates the binding free energy of the
protein-ligand interaction, such as the MM/PBSA method, which takes into account
the relative binding free energy obtained from the protein-ligand molecular dynamics in aqueous solvent, the change in potential energy in vacuum, the desolvation
energy of the system, and the entropy of the complex in the gas phase. It is a wellknown method for rescoring molecules.

9 Conformational Sampling of Proteins: Methods for Simulate... 255
With the HTP SurflexDock, it is possible to load a three-dimensional structure of
the receptor and a set of small molecules (ligands) in PDBQT format. The molecules
are classified according to the ΔG calculated for the ensemble, an ensemble with the
original structure of the receptor is then built and three conformations are obtained
from five-nanosecond molecular dynamics using Gromacs 5.1.5 [2] (a relaxation for
other conformations), the trajectory is followed by clustering, and the conformations
with a maximum RMSD of the binding site of 0.10–0.20 nm and grouped by the
GROMOS algorithm [16]. A representative structure from the three most representative groups is included in the new set. AutoDock [49] and ADT scripts [46] are
used, and ten poses are generated for the complexes [45]. HTP SurflexDock presents
a score table with the ΔG of the best pose, and it is possible to apply two postprocessing options: (I) select 10 compounds and expand the conformational space
for a new molecular docking experiment cycle; (II) explore the conformational space
of 30 new poses.
The latest version of DockThor-VS (availabl e at: www.dockthor.lncc.br, last
accessed on July, 2024) provides users with 3D structures for performing molecular
docking of various ligands in selected structures of the wild-type and mutated
proteins: Nsp3; Nsp5; Nsp12; Nsp15; and Spike. Encouraging the importance of
using several conformations of the biological target in virtual screening, as according
to the authors, using a single conformation of the protein would generate an impaired
classification of the ligands in the complexes. A list of some interesting ensemble
docking tools is presented in Table 9.1.
Another interesting study [65], which presents a protocol for generating various
conformations of cyclin-dependent protein kinase 2 (CDK2), used normal mode
analysis (DIMB [54] with harmonic constraints during minimization, followed by
shifting by 25 lower frequency vectors until the mass-weighted mean square deviation (MRMSD) to 2 Å or –2 Å. The authors selected the structures for their
conformational diversity and binding site topology, corresponding to five open
conformations with more space in the protein’s binding site to docking ligands.
An innovative study [21] used hybrid molecular dynamics methods, and excited
normal modes to obtain conformations of sulfotransferases (SULT) for docking
molecular 132 ligands, helping to understand the mechanism of molecule binding
at the site of this protein. Some studies involving proteins can be more structural,
with a focus on understanding the structure itself and which regions are most
important, such as the one conducted by Yang et al., [68], or more applied, such
as the one by [41], where the focus is on demonstrating the impact of the composition of the protein structure. Both are relevant. Dudas et al. performed MD and
MDeNM simulations for the SULT1A1/PAPS as well as MD and docking simulations with the substrate’s estradiol and fulvestrant. The authors demonstrated that
large conformational changes of the PAPS-bound SULT1A1 can occur and would
be sufficient to accommodate large substrates, e.g. fulvestrant, independently of the
co-factor movements. In this work, the structural displacements were successfully
detected by the MDeNM simulations and suggest that a wider range of drugs could
be recognized by PAPS-bound SULT1A1. The Hybrid method called MDeNM
enables an extended sampling of the conformational space by running multiple

256 A. L. Scott et al.
Table 9.1 Main tools of the ensemble docking technique
Tool
ensemble
docking Key points Docking References
HEX Combined strategy between molecular dynam-
HTP
SurflexDock
EDock-ML It uses machine learning to decide whether
MDock Automated docking that can dock ligands into
MDR
SurFlexDock
DockThor-VSIt uses structures of different conformations
ics and molecular docking with the program
HEX, which performs a systematic search of
6 degrees of freedom and classification of orientations by interaction energy
It has two strategies: ensemble coupling with
implicit receptor flexibility simulation and postprocessing step with new scoring of promising
compounds considering the conformational
space or estimating the binding free energy
using the MM/PBSA protocol
compounds are good candidates for docking
with different receptor conformations, considering the receptor’s flexibility
multiple conformations of the receptor, automated, using an ensemble docking algorithm
[47, 48]
Uses for molecular docking a discrete/representative set of receptors contact surfaces
obtained from the clustering of molecular simulation trajectories, for simulation of the intrinsic flexibility of the ligand contact surface, the
classification of ligands can be by the inhibition
constant (Ki) and is also applicable to receptors
originated by homology modeling
complexed with ligands, taking into account the
flexibility of the receptor and different states of
protonation
Proteinprotein
Proteinligand/protein-protein
Protein-ligand Chandak
Protein-ligand Huang et al.
Protein-ligand De Almeida
Protein-ligand Guedes et al.
Ritchie et al.
[60]
Filho et al.
[22]
et al. [9]
[29]
Huang et al.
[30]
Huang et al.
[31]
Huang et al.
[32]
Huang et al.
[33]
Huang et al.
[34]
Filho et al.
[17]
26]
[25,
short MD simulations during which motions described by a subset of low-frequency
Normal Modes are kinetically excited. It was possible to detect “open”-like conformations of SULT1A1 ef ficiently.
Finally, there are several hybrid methods proposed in the literature; some of them
integrate simulation methods such as molecular dynamics, Monte Carlo,

9 Conformational Sampling of Proteins: Methods for Simulate... 257
optimization with elastic network models or normal modes analysis based on force
field as MDeNM, CoMD , ClustENM, CLusteNMD and VMOD [14, 35, 42, 55].
These methods can be combined with the software listed in Table 9.1 which allows
us to obtain conformations of protein structures that are representative of interesting
states for the molecular coupling of bioactive ligands. Using structures from different points, considering the flexibility of a biological target, brings more realism to
computer simulations, and is a good strategy for understanding the mechanism of
action in the target and developing new molecules.
4 Good Practices for Simulations and Sample
the Conformational Space
Molecular simulation techniques play a crucial role in our quest to understand and
predict the properties, structure, and function of molecular systems. They are a
fundamental tool as we aim to enable predictive molecular design. Simulation
methods are useful for studying the structure and dynamics of complex systems
that are too complicated for traditional theoretical approaches, helping to interpret
experimental data in terms of molecular movements. We list some questions that can
be important to plan your simulation on Table.
Questions to plan the simulations
Guiding questions for implementing conformational sampling in drug design
What is the quality of your initial structures or models?
What are the protonation and phosphorylation conditions for your simulation?
How do you scale your simulation with multiple CPUs/GPUs?
Which force field should be used?
Which kind of motions is important for your problem: local motions or global motions?
Do I have computational resource and time to use explicit solvent and all-atoms models?
How can I optimize the parallelization of the simulations?
How representative are the structures obtained when you project them against principal components space?
How diverse is your sampling of the conformational space in terms of structural (RMSD) and
energetic aspects?
References
1. Abagyan, R., Totrov, M., & Kuznetsov, D. (1994). ICM - A new method for protein modeling
and design: Applications to docking and structure prediction from the distorted native conformation. Journal of Computational Chemistry, 15, 488–506.
2. Abraham, M. J., Murtola, T., Schulz, R., et al. (2015). GROMACS: High performance
molecular simulations through multi-level parallelism from laptops to supercomputers.
SoftwareX, 1,19–25.

258 A. L. Scott et al.
3. Amaro, R. E., Baudry, J., Chodera, J., Demir, Ö., McCammon, J. A., Miao, Y., & Smith, J. C.
(2018). Ensemble docking in drug discovery. Biophysical Journal, 114(10), 2271–2278.
4. Atilgan, C. (2018). Computational methods for efficient sampling of protein landscapes and
disclosing allosteric regions. Advances in Protein Chemistry and Structural Biology, 113,33–
63. . Epub 2018 Jul 25. PMID: 30149905.
5. Bahar, I., Lezon, T. R., Yang, L. W., & Eyal, E. (2010). Global dynamics of proteins: bridging
between structure and function. Annual Review of Biophysics, 39,23–42.
6. Bahar, I., Jernigan, R., & Dill, K. (2017). Protein actions: Principles & modeling. Garland
Science, Taylor & Francis Group. ISBN: 9780815341772.
7. Black, K. A., He, S., Jin, R., Miller, D. M., Bolla, J. R., Clarke, O. B., Johnson, P., Windley, M.,
Burns, C. J., Hill, A. P., Laver, D., Robinson, C. V., Smith, B. J., & Gulbis, J. M. (2020). A
constricted opening in Kir channels does not impede potassium conduction. Nature Communi-
cations, 11(1), 3024.
8. Callaway, E. (2020). Revolutionary cryo-EM is taking over structural biology. Nature, 578
(7794), 201.
9. Chandak, T., & Wong, C. F. (2021). EDock-ML: A web server for using ensemble docking with
machine learning to aid drug discovery. Protein Science, 30(5), 1087–1097. . Epub 2021 Mar
25.
10. Chilingaryan, G., Abelyan, N., Sargsyan, A., et al. (2021). Combination of consensus and
ensemble docking strategies for the discovery of human dihydroorotate dehydrogenase inhibitors. Scientific Reports, 11, 11417.
11. Corbeil, C. R., Englebienne, P., & Moitessier, N. (2007). Docking ligands into flexible and
solvated macromolecules. 1. Development and validation of FITTED 1.0. Journal of Chemical
Information and Modeling, 47(2), 435–449.
12. Costa, M. G. S., Batista, P. R., Bisch, P. M., & Perahia, D. (2015). Exploring free energy
landscapes of large conformational changes: Molecular dynamics with excited normal modes.
Journal of Chemical Theory and Computation, 60(5), 2419–2423.
13. Costa, M. G. S., Fagnen, C., Vénien-Bryan, C., & Perahia, D. (2020). A new strategy for atomic
flexible fitting in Cryo-EM maps by molecular dynamics with excited normal modes (MDeNMEMFit). Journal of Chemical Information and Modeling, 60(5), 2419–2423.
14. Costa, M. G. S., Batista, P. R., Gomes, A., Bastos, L. S., Louet, M., Floquet, N., Bisch, P. M., &
Perahia, D. (2023). MDexciteR: Enhanced sampling molecular dynamics by excited normal
modes or principal components obtained from experiments. Journal of Chemical Theory and
Computation.
15. Das, A., Gur, M., Cheng, M. H., Jo, S., Bahar, I., & Roux, B. (2014). Exploring the conformational transitions of biomolecular systems using a simple two-state anisotropic network
model. PLoS Computational Biology, 10(4), e1003521.
16. Daura, X., Gademann, K., Jaun, B., et al. (1999). Peptide folding: When simulation meets
experiment. Angewandte Chemie International Edition, 38(1–2), 236–240.
17. De Almeida Filho, J. L., & Fernandez, J. H. (2020). MDR SurFlexDock: A semi-automatic
webserver for discrete receptor-ensemble docking. In L. Kowada & D. de Oliveira (Eds.),
Advances in bioinformatics and computational biology. BSB 2019 (Lecture Notes in Computer
Science) (Vol. 11347). Springer.
18. De Vries, S. J., van Dijk, M., & Bonvin, A. M. (2010). The HADDOCK web server for datadriven biomolecular docking. Nature Protocols, 5(5), 883–897.
19. Dominguez, C., Boelens, R., & Bonvin, A. M. J. J. (2003). HADDOCK: A protein-protein
docking approach based on biochemical or biophysical information.
Chemical Society, 125(7), 1731–1737.
20. Dror, R. O., Dirks, R. M., Grossman, J. P., Xu, H., & Shaw, D. E. (2012). Biomolecular
simulation: a computational microscope for molecular biology. Annual Review of Biophysics,
41, 429–452.
Journal of the American

9 Conformational Sampling of Proteins: Methods for Simulate... 259
21. Dudas, B., Toth, D., Perahia, D., Nicot, A. B., Balog, E., & Miteva, M. A. (2021). Insights into
the substrate binding mechanism of SULT1A1 through molecular dynamics with excited
normal modes simulations. Scienti fic Reports, 11, 13129.
22. Filho, E. D. (org.), Scott, A. L., Philot, E. A., Lima, A. N. (2019). Métodos computacionais no
estudo de macromoléculas biológicas, 2nd ed. (pp. 69–90). Editora Livraria da Física.
23. Fratev, F., & Sirimulla, S. (2019). An improved free energy perturbation FEP+ sampling
protocol for flexible ligand-binding domains. Scientific Reports, 9(1).
24. Grant, B. J., Rodrigues, A. P. C., ElSawy, K. M., McCammon, J. A., & Caves, L. S. D. (2006).
Bio3d: An R package for the comparative analysis of protein structures. Bioinformatics, 22,
2695–2696.
25. Guedes, I. A., et al. (2021). Drug design and repurposing with DockThor-VS web server
focusing on SARS-CoV-2 therapeutic targets and their non-synonym variants. Scientific
Reports, 11, 5543.
26. Guedes, I. A., Barreto, A. M. S., Marinho, D., Krempser, E., Kuenemann, M. A., Sperandio, O.,
Dardenne, L. E., & Miteva, M. A. (2021). New machine learning and physics-based scoring
functions for drug discovery. Scientific Reports, 11(1), 3198.
27. Gur, M., Zomot, E., & Bahar, I. (2013). Global motions exhibited by proteins in micro- to
milliseconds simulations concur with anisotropic network model predictions. The Journal of
Chemical Physics, 139, 121912.
28. Haliloglu, T., & Bahar, I. (2015). Adaptability of protein structures to enable functional
interactions and evolutionary implications. Current Opinion in Structural Biology, 35,17–23.
29. Huang, S. Y., & Zou, X. Q. (2006). An iterative knowledge-based scoring function to predict
protein-ligand interactions: I. Derivation of interaction potentials. Journal of Computational
Chemistry, 27, 1866–1875.
30. Huang, S. Y., & Zou, X. Q. (2006). An iterative knowledge-based scoring function to predict
protein-ligand interactions: II. Validation of the scoring function. Journal of Computational
Chemistry, 27, 1876–1882.
31. Huang, S. Y., & Zou, X. Q. (2007). Ensemble docking of multiple protein structures: Considering protein structural variations in molecular docking. Proteins, 66, 399–421.
32. Huang, S. Y., & Zou, X. Q. (2007). Efficient molecular docking of NMR structures: Application
to HIV-1 protease. Protein Science: A Publication of the Protein Society, 16,43–51.
33. Huang, S. Y., & Zou, X. Q. (2008). An iterative knowledge-based scoring function for proteinprotein recognition. Proteins, 72, 557–579.
34. Huang, S. Y., & Zou, X. Q. (2013). A non-redundant structure dataset for benchmarking
protein-RNA computational docking. Journal of Computational Chemistry.
35. Kaynak, B. T., Krieger, J. M., Dudas, B., Dahmani, Z. L., Costa, M. G., Balog, E., Scott, A. L.,
Doruker, P., Perahia, D., & Bahar, I. (2022). Sampling of protein conformational space using
hybrid simulations: A critical assessment of recent methods. Frontiers in Molecular Biosci-
ences, 9, 832847.
36. Krieger, J. M., Doruker, P., Scott, A. L., Perahia, D., & Bahar, I. (2020). Towards gaining sight
of multiscale events: Utilizing network models and normal modes in hybrid methods. Current
Opinion in Structural Biology, 64,34
37. Krieger, J. M., Sorzano, C. O. S., Carazo, J. M., & Bahar, I. (2022 Apr 1). Protein dynamics
developments for the large scale and cryoEM: case study of ProDy 2.0. Acta Crystallographica
Section D: Structural Biology, 78(Pt 4), 399–409.
38. Kurkcuoglu, Z., & Bonvin, A. M. J. J. (2020). Pre- and post-docking sampling of conformational changes using ClustENM and HADDOCK for protein-protein and protein-DNA systems.
Proteins, 88(2), 292–306.
39. Kurkcuoglu, Z., Bahar, I., & Doruker, P. (2016). ClustENM: ENM-based sampling of essential
conformational space at full atomic resolution. Journal of Chemical Theory and Computation,
12, 4549–4562.
–41.

260 A. L. Scott et al.
40. Lätti, S., Niinivehmas, S., & Pentikäinen, O. T. (2016). Rocker: Open source, easy-to-use tool
for AUC and enrichment calculations and ROC visualization. Journal of Cheminformatics, 8,
45.
41. Li, H., Chang, Y. Y., Lee, J. Y., Bahar, I., & Yang, L. W. (2017). DynOmics: Dynamics of
structural proteome and beyond. Nucleic Acids Research, 45(W1), W374–W380.
42. Lima, A. N., de Oliveira, R. J., Braz, A. S. K., de Souza Costa, M. G., Perahia, D., & Scott, L. P.
B. (2018). Effects of pH and aggregation in the human prion conversion into scrapie form: A
study using molecular dynamics with excited normal modes. European Biophysics Journal, 47
(5), 583–590.
43. McKay, K., Hamilton, N. B., Remington, J. M., Schneebeli, S. T., & Li, J. (2022). Essential
dynamics ensemble docking for structure-based GPCR drug discovery. Frontiers in Molecular
Biosciences, 9, 879212.
44. Mitra, A. K. (2019). Visualization of biological macromolecules at near-atomic resolution:
Cryo-electron microscopy comes of age. Acta Crystallographica. Section F, Structural Biology
Communications, 75(Pt 1), 3–11.
45. Morris, G. M., Goodsell, D. S., Halliday, R. S., et al. (1998). Automated docking using a
Lamarckian genetic algorithm and an empirical binding free energy function. Journal of
Computational Chemistry, 19(14), 1639–1662.
46. Morris, G. M., Huey, R., Lindstrom, W., et al. (2009). AutoDock4 and AutoDockTools4:
Automated docking with selective receptor flexibility. Journal of Computational Chemistry, 30
(16), 2785–2791.
47. Mysinger, M. M., Carchia, M., Irwin, J., & Shoichet, B. K. (2012). Directory of useful decoys,
enhanced (DUD-E): Better ligands and decoys for better benchmarking. Journal of Medicinal
Chemistry, 55, 6582–6594.
48. Nogales, E. (2016). The development of cryo-EM into a mainstream structural biology technique. Nature Methods, 13,24–27.
49. Norgan, A. P., Coffman, P. K., Kocher, J.-P. A., Katzmann, D. J., & Sosa, C. P. (2011).
Multilevel parallelization of AutoDock 4.2. Journal of Cheminformatics, 3(1), 12.
50. O’Boyle, N. M., et al. (2011). Open babel: An open chemical toolbox. Journal of
Cheminformatics, 3, 33.
51. Orellana, L. (2019). Large-scale conformational changes and protein function: Breaking the in
silico barrier. Frontiers in Molecular Biosciences, 6.
52. Orellana, L., Yoluk, O., Carrillo, O., et al. (2016). Prediction and validation of protein
intermediate states from structurally rich ensembles and coarse-grained simulations. Nature
Communications, 7, 12575.
53. Páll, S., Abraham, M. J., Kutzner, C., Hess, B., & Lindahl, E. (2015). Tackling exascale
software challenges in molecular dynamics simulations with GROMACS. In S. Markidis &
E. Laure (Eds.), Solving software challenges for exascale. EASC 2014. Lecture notes in
computer science (Vol. 8759). Springer.
54. Perahia, D., & Mouawad, L. (1995). Computation of low-frequency normal modes in macromolecules: Improvements to the method of diagonalization in a mixed basis and application to
hemoglobin. Computers & Chemistry, 19(3), 241–246.
55. Philot, E. A., Perahia, D., Braz, A. S., Costa, M. G., & Scott, L. P. (2013). Binding sites and
hydrophobic pockets in Human Thioredoxin 1 determined by normal mode analysis. Journal of
Structural Biology, 184(2), 293–300.
56. Pouya, I., Pronk, S., Lundborg, M., & Lindahl, E. (2017). Copernicus, a hybrid datafl
peer-to-peer scientific computing platform for efficient large-scale ensemble sampling. Future
Generation Computer Systems, 71(-),18–31.
57. Quiroz, R. C. N., Philot, E. A., General, I. J., Perahia, D., & Scott, A. L. (2023). Effect of
phosphorylation on the structural dynamics, thermal stability of human dopamine transporter: A
simulation study using normal modes, molecular dynamics and Markov State Model. Journal of
Molecular Graphics & Modelling, 118, 108359.
ow and

9 Conformational Sampling of Proteins: Methods for Simulate... 261
58. Resende-Lara, P. T., Perahia, D., Scott, A. L., & Braz, A. S. K. (2020). Unveiling functional
motions based on point mutations in biased signaling systems: A normal mode study on nerve
growth factor bound to TrkA. PLoS One., 15(6), e0231542.
59. Resende-Lara, P. T., Costa, M. G. S., Dudas, B., & Perahia, D. (2022). Adaptive collective
motions: A hybrid method to improve conformational sampling with molecular dynamics and
normal modes. biorxiv [Preprint]. Available at: https://www.biorxiv.org/.
60. Ritchie, D. W., & Kemp, G. J. (2000). Protein docking using spherical polar Fourier correlations. Proteins, 39(2), 178–194.
61. Ruiz-Carmona, S., et al. (2014). rDock: A fast, versatile and open source program for docking
ligands to proteins and nucleic acids. PLoS Computational Biology, 10, e1003571.
62. Schay, G., Herényi, L., Fidy, J., & Osváth, S. (2013). Role of domain interactions in the
collective motion of phosphoglycerate kinase. Biophysical Journal, 104(3), 677–682.
63. Shaw, D. E., Dror, R. O., Salmon, J. K., et al. (2009). Millisecond-scale molecular dynamics
simulations on Anton. In Proceedings of the Conference on High Performance Computing
Networking, Storage and Analysis (SC09), Portland, OR, USA (pp. 1–11).
64. Shim, J., & Mackerell, A. D., Jr. (2011). Computational ligand-based rational design: Role of
conformational sampling and force fields in model development. MedChemComm, 2(5), 356–
370.
65. Sperandio, O., Mouawad, L., Pinto, E., Villoutreix, B. O., Perahia, D., & Miteva, M. A. (2010).
How to choose relevant multiple receptor conformations for virtual screening: A test case of
Cdk2 and normal mode analysis. European Biophysics Journal, 39(9), 1365–1372.
66. Trott, O., & Olson, A. J. (2010). AutoDock Vina: Improving the speed and accuracy of docking
with a new scoring function, efficient optimization, and multithreading. Journal of Computa-
tional Chemistry, 31, 455–461.
67. van Zundert, G. C. P., Rodrigues, J. P. G. L. M., Trellet, M., Schmitz, C., Kastritis, P. L.,
Karaca, E., Melquiond, A. S. J., van Dijk, M., de Vries, S. J., & Bonvin, A. M. J. J. (2016). The
HADDOCK2.2 web server: User-friendly integrative modeling of biomolecular complexes.
Journal of Molecular Biology, 428(4), 720–725.
68. Yang, L., Song, G., & Jernigan, R. L. (2007). How well can we understand large-scale protein
motions using normal modes of elastic network models? Biophysical Journal, 93(3), 920–929. .
Epub 2007 May 4.
69. Yu, I., Mori, T., Ando, T., Harada, R., Jung, J., Sugita, Y., & Feig, M. (2016). Biomolecular
interactions modulate macromolecular structure and dynamics in atomistic model of a bacterial
cytoplasm. Elife, 5, e19274.
70. Zhang, N., & Zhao, H. (2016). Enriching screening libraries with bioactive fragment space.
Bioorganic & Medicinal Chemistry Letters, 26, 3594–3597.
71. Zhang, S., Krieger, J. M., Zhang, Y., Kaya, C., Kaynak, B., Mikulska-Ruminska, K., Doruker,
P., Li, H., & Bahar, I. (2021). ProDy 2.0: Increased scale and scope after 10 years of protein
dynamics modelling with Python. Bioinformatics, 37(20), 3657– 3659.
72. Zhang, Y., Krieger, J., Mikulska-Ruminska, K., Kaynak, B., Sorzano, C. O. S., Carazo, J.-M.,
Xing, J., & Bahar, I. (2021). State-dependent sequential allostery exhibited by chaperonin
TRiC/CCT revealed by network analysis of Cryo-EM maps. Progress in Biophysics and
Molecular Biology, 160, 104–120.

Chapter 10
Free Energy Perturbation and Free-Energy
Calculations Applied to Drug Design
Deborah Antunes
, Lucianna Helene Santos
Ana Carolina Ram os Guimarães
, and Ernesto Raul Caffarena
,
Abstract Free energy perturbation (FEP) is a computational technique used to
evaluate ligand-protein binding affinities for computer-aided drug optimization.
FEP has been shown to be a valuable tool in both academic and pharmaceutical
settings for optimizing drug candidates in a rational, efficient, and cost-effective
manner. Recent advancements in algorithms, software tools, hardware capabilities,
and machine learning integration have significantly improved the scope, applicability, and reliability of FEP calculations. In this chapter, we review recent developments in force field parameterization, software platforms, and automated workflows
that have consolidated FEP as an essential methodology for structure-based drug
discovery and have resulted in FEP calculations becoming more accessible to nonspecialists, as well as applicable to a broad range of scenarios. We also describe the
utility of the FEP technique in diverse contexts, including validating the binding
modes and optimizing allosteric and covalent inhibitors. We illustrate its potential
through vignettes and in-depth case studies, demonstrating its integration into
machine-learning frameworks for predicting binding energies based on molecular
structures. Furthermore, this chapter discusses the remaining challenges in sampling
sufficiency and scalability to ultra-large compound libraries as well as emerging
solutions through cloud computing and machine learning.
Keywords Free energy perturbation · Drug design · Bbinding affinity ·
Computational chemistry · Machine learning · Molecular structures
D. Antunes · A. C. R. Guimarães
Oswaldo Cruz Institute, Rio de Janeiro, Brazil
L. H. Santos (
Institut Pasteur de Montevideo, Montevideo, Uruguay
e-mail: lsilva@pasteur.edu.uy
E. R. Caffarena (
Scientific Computing Program, Rio de Janeiro, Brazil
e-mail: ernesto.caffarena@fiocruz.br
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2024
V. G. Maltarollo (ed.), Computer-Aided and Machine Learning-Driven Drug
Design, Computer-Aided Drug Discovery and Design 3,
https://doi.org/10.1007/978-3-031-76718-0_10
✉)
✉)
263
Соседние файлы в папке Библиотека им академика М.И. Перельмана
