Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5364_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Foreword
- •Acknowledgments
- •Contents
- •1.1 Structure-Based Drug Discovery (SBDD)
- •1.2 Ligand-Based Drug Design (LBDD)
- •1.3 Echoes from the Past, Visions from the Future
- •References
- •1 Introduction
- •2.2 Second Step: Data Curation
- •2.4 Fourth Step: Updating and Maintenance
- •2 Databases and Curation
- •8 Perspectives
- •9 Conclusion
- •References
- •1 Introduction
- •2.1 Making and Matching Protein Models
- •2.2 Simulating Protein Movements
- •2.3 Analyzing Changes in Protein Shape
- •3 Pharmacogenomics in Drug Development
- •4 Case Studies of Genomics-Based Drug Design
- •References
- •1 Historical Background
- •1.1 Timeline
- •2 Methodology Overview
- •2.1 Neural Networks
- •2.1.1 Perceptron
- •2.1.2 Multilayer Neural Networks
- •2.1.3 Types of Neural Networks
- •Feedforward
- •Recurrent Neural Networks
- •LSTM
- •2.2 Deep Learning
- •3 Using Machine Learning
- •3.2 Data Collection
- •3.3 Data Preprocessing
- •3.4 Model Selection
- •3.5 Model Training
- •3.6 Validation
- •3.7 Tuning
- •3.8 Prediction
- •4 Limitations
- •4.1 Bias
- •4.3 Interpretability
- •4.4 Computational Cost
- •4.5 Data Dependency
- •4.6 Robustness
- •5 Applications in Drug Discovery
- •5.2 Lead Discovery
- •5.3 Preclinical and Clinical Development
- •6 Resources and Tools
- •7 Challenges and Perspectives
- •7.1 Future Trends
- •9 Conclusions
- •References
- •1 Historical Background
- •1.1 Applications in Drug Discovery
- •2 Validations and Controls
- •2.1 Internal Validation
- •2.2 External Validation
- •2.3 Relative Cluster Validation
- •3 Challenges and Perspectives
- •4 Conclusions
- •References
- •1 Historical Background
- •2 OECD Principles
- •2.1 A Defined Endpoint
- •2.2 An Unambiguous Algorithm
- •2.5 A Mechanistic Interpretation, if Possible
- •3 Software and Tools
- •4 Validations and Controls
- •4.1 Internal and External Validation
- •4.1.1 Regression Metrics
- •4.2 Applicability Domain
- •4.3 Randomization Tests
- •5 Interpretation
- •6 Practical Advice During QSAR Modeling
- •7 Application
- •8 Challenges and Perspectives
- •References
- •1 Molecular Docking
- •2 Advances in Scoring Functions and Search Algorithms
- •2.2 Critical Characteristics of Search Algorithms
- •2.3 Docking Programs and Scoring Functions
- •3 Calculations Performed During Docking Simulations
- •4 Essential Components for a Good Docking Program
- •5 Limitations of the Docking Technique
- •6 Validation of Docking Results
- •7 Inappropriate Use of Validation Methods in Docking
- •9 Use of Machine Learning in Molecular Docking
- •11 Challenges
- •12 Conclusions
- •References
- •3 System Preparation for MD Simulations
- •3.1 Solvation and Microensemble
- •3.2 Force Fields: General Concept and Relevant Choices
- •3.3 The Concept of Replicas and Timescale
- •4.1.2 Protein Root Mean Square Fluctuation (RMSF)
- •4.1.4 Protein Secondary Structure Analysis
- •4.1.5 Principal component Analysis (PCA)
- •4.1.6 Markov State Modelling
- •4.1.7 Distance Calculations
- •4.1.8 Angle and Plane Calculations
- •4.2.2 Distances and Ligand-Induced Geometry Rearrangements
- •4 Molecular Dynamics Analysis
- •4.1 Protein Perspective
- •4.1.1 Protein Root Mean Square Deviation (RMSD)
- •4.3 Ligand Perspective
- •4.3.1 Ligand Properties
- •4.3.2 Ligand Root Mean Square Deviation
- •4.3.3 Ligand Root Mean Square Fluctuation
- •4.3.4 Angles and Dihedrals
- •5.1 Protein Structure Prediction and Preparation
- •5.2 Molecular Docking
- •6 Concluding Remarks and Outlook
- •Glossary
- •References
- •1 Introduction
- •2.1 MDeNM
- •2.2 Collective Molecular Dynamics (coMD)
- •2.3 ClustENM and ClustENMD
- •3 Ensemble Docking
- •References
- •1 Introduction
- •1.1 Advantages, Disadvantages, Innovations, and Challenges
- •1.2 Recent Advances in Accessible FEP Software Tools
- •1.3 Applications of FEP in Industry and Consortiums
- •2 Expanding the Potential of FEP Calculations
- •2.1 Validating Binding Poses
- •2.2 Dealing with Solvent
- •2.3 FEP and Allostery
- •2.4 FEP and Covalent Ligands
- •2.5 Applications of FEP in Scaffold Hopping
- •2.6 Positional Analogue Scanning
- •2.7 Combinations and Alternative Approaches
- •3 Machine Learning for FEP
- •3.4 Implications for ML in FEP Calculations
- •4 Final Considerations
- •5 First Steps to FEP Simulations
- •References
- •1 Background
- •2 Ultra-Large Screening Libraries and Chemical Spaces
- •3.1 Implications of Dataset Size
- •4 Ligands on the Ultra-Large Scale
- •4.1 Ultra-Large 2D Similarity Searches
- •7 Challenges and Future Perspectives
- •7.1 Hit Triage: An Old Problem on a New Dimension
- •8 Conclusions
- •Appendix
- •References
- •1 Introduction
- •2 Enzymatic Activity Evaluations
- •3 Cytotoxicity Evaluation and Cell Viability
- •4 Antiviral Assays in Experimental Validation
- •6 In Vivo Evaluation of Compounds
- •7 Conclusions
- •References
- •1 Introduction
- •3.1 Data Collection
- •3.2 Data Preprocessing
- •3.4 Model Choice
- •3.5 Model Training
- •3.6 Model Assessment
- •3.7 External Validation
- •3.8 Implementation and Availability
- •3.9 Continuous Update
- •5 Conclusions and Perspectives
- •References
- •1 Experimental Approaches to Obtain Protein Structure
- •1.1 X-Ray Crystallography
- •1.2 Nuclear Magnetic Resonance
- •1.3 Cryo-EM
- •1.4 Hybrid Methods
- •2 Modeling Approaches to Obtain Protein Structure
- •2.1 Homology Modeling
- •2.2 Ab Initio Modeling
- •2.3 New Approaches
- •3 Conformational Diversity of Proteins
- •3.1 Characterization of Protein Conformational States
- •3.2 Experimental Methods to Study Protein Dynamics and Conformations
- •3.4 Molecular Dynamics Simulation
- •3.5 Sampling Strategies
- •4 Remarks and Perspectives
- •References
- •1 Introduction
- •2 Structure-Based Drug Design of HIV Protease Inhibitors
- •2.1 HIV-1 Protease as a Therapeutic Target
- •2.2.1 Saquinavir
- •2.2.2 Indinavir
- •2.3.1 Lopinavir
- •2.3.2 Darunavir
- •6 Conclusions
- •References
- •4 Experimental Methods to Analyze NR Activity
- •4.2 Coregulator-Recruitment
- •5 Concluding Remarks and Outlook
- •References

284 D. Antunes et al.
between the ligand end states in λ-space must first be flattened [94, 98]. This is
possible by identifying and incorporating bias potentials into λ-dynamic simulations.
Therefore, determining these biases requires a consi derable amount of simulation
time prior to production sampling, which reduces the efficacy and cost of the
method. In recent years, several new developments have been introduced to expand
the applicability of λ-dynamics for drug discovery, including multisite λ-dynamics
[86, 99], which enables multiple substituents at multiple sites, the use of a biasing
potential replica exchange to enhance transitions between states [100], and an
alternative λ sampling strategy employing Gibbs sampling [101, 102].
3 Machine Learning for FEP
3.1 Introduction to the Use of Machine Learning in Free
Energy Perturbation Calculations
The challenges and limitations discussed in the previous sections pose substantial
obstacles to implementing FEP calculations in p ractical applications, necessitating
the development of novel approaches and methodologies to overcome them. Thus,
ML enhances the precision and effectiveness of FEP calculations. Rec ent advancements in ML methodologies offer new possibilities for overcoming such difficulties
and improving the reliability of FEP predictions to advance drug discovery
[49]. Machine learning is a branch of artificial inte lligence (AI) that specifically
deals with the creation of algorithms and statistical models that allow computers to
carry out tasks by learning from data, rather than relying on explicitly written
instructions. This data-driven approach allows systems to improve their performance
on tasks over time as they are exposed to more data. In the realm of AI, ML
techniques are employed to recognize patterns, make predictions, and inform
decision-making processes across various domains, including natural language
processing, computer vision, and computational chemistry [103].
ML employs statistical models trained on large datasets to capture the complex
patterns. For FEP, relevant training data may include FEP simulation trajectories,
experimentally determined structural and thermodynamic measurements, and quantum mechanical calculations on small molecule subsets. ML algorithms, particularly
those employing deep learning architectures, have demonstrated the capability of
learning complex patterns in data, enabling more accurate predictions of free energy
changes [33]. By training on physics-based simulations and experimental data, ML
models can learn to correct weaknesses in fixed-charge force fields [104], drive
enhanced sampling of binding modes [105], and bypass costly simulations to
directly predict binding affinities from molecular structures [106]. This efficiency
is paramount in the high-throughput screen ing of drug candidates, allowing for the
rapid evaluation of numer ous compounds [107]. ML algorithms assist in the efficient
description of potential energy surfaces and in the accurate estimation of the

10 Free Energy Perturbation and Free-Energy Calculations Applied to Drug Design 285
parameters for these simulations. By leveraging ML, FEP can more effectively
model the molecular interactions and transformations that are essential in drug
discovery, leading to more precise predictions of binding af finities and thermodynamics of molecular processes. The integration of ML into FEP represents a
significant advancement in computational chemistry.
3.2 Machine Learning Methods in Free Energy Perturbation
Calculations
Several ML methods and algorithms can be utilized in the field of computational
chemistry to enhance the FEP calculations. These ML techniques offer distinct
solutions to specific issues and applications, thereby significantly contributing to
the advancement of the FEP calculations. For more details on ML theory and
methods, please see Chap. 4.
One of the prominent ML algorithms used in this context is neural networks
(NNs), especially deep learning architectures such as convolutional neural networks
(CNNs). These networks are adept at predicting free energy changes directly from
molecular structures and learning the complex relationships between molecular
features and free energies. CNNs, in particular, proces s spatial information in
molecular systems, offering predictions of free energy changes without explicit
simulation.
Gaussian Process Regression (GPR) is another ML algorithm that can be used to
construct surrogate models of free energy landscapes. This approach decreases the
amount of computer resources needed by allowing for efficient exploration of the
parameter space and offering estimates of uncertainty for predictions. This is particularly advantageous for optimizing simulation parameters and incorporating various
data sources [108].
Decision trees and random forests are algorithms applied for feature selection,
determining the chemical descriptors that have a substantial impact on free energy.
This knowledge can guide the configuration of the FEP calculations and improve
their comprehension.
Support vector machines (SVMs) have been employed as ML algorithms to
distinguish compounds based on their propensity to exhibit specific free energy
changes. This approach has been utilized in drug design to differentiate between
ligands that bind and those that do not. Moreo ver, SVMs can be trained to predict
changes in the free energy using chemical descriptors [109]. This has proven
effective in identifying compounds that are likely to exhibit the desired free energy
changes.
Reinforcement learning (RL) is a ML algorithm in which an agent learns to make
decisions, such as choosing mutations or adjustments to a ligand, to maximize the
total rewards associated with the desired change s in free energy. RL enhances the
alchemical pathway in FEP calculations, thereby enhancing the effectiveness of
sampling tactics in the simulations [ 103].

286 D. Antunes et al.
In addition to the specific ML algorithms, various machine learn ing methods are
also applied in the contex t of FEP calculations. These methods define how the
overall process of ML is applied to solve the problem.
Bayesian methods, including Bayesian optimization, are applied to refine free
energy estimates, integrating experimental data with simulation results. This
approach is extremely beneficial for selecting optimal simulation parameters.
Transfer learning is utilized to apply models trained on a specific set of FEP data
to distinct, yet interconnected problems, resulting in a substantial reduction in the
computational resources required for new calculations. This approach has great
potential in the field of drug development as it allows for the application of models
that have been trai ned on large datasets to new molecules [110].
Autoencoders, including variational autoencoders (VAEs), are employed to
reduce the dimensionality of FEP calculations. They condense intricate molecular
features into a space with fewer dimensions, simplifying the representation of
molecular systems and assisting in their analysis and interpretation [111].
Each of these ML algorithms and methods improves the accuracy of the FEP
calculations, effectively managing high-dimensional data. Additionally, they reduce
the computing expenses associated with these calculations and provide valuable
insights into the factors that determine changes in free energy. The incorporation of
ML in FEP calculations is a rapidly growing area that has consistently gained
advantages from the progress in both ML and computational chemistry. Therefore,
these methods play a crucial role in transforming the investigation of molecular
systems, allowing for predictions on a wide scale and for complicated systems that
were previously impossible using conventional computational methods.
3.3 Advanced Applications of Machine Learning in Free
Energy Perturbation: A Compilation of Case Studies
The integration of ML with FEP calculations has significantl y advanced the field of
drug discovery, as exemplified by the studies of Konze et al. [112] and Ghanakota
et al. [33]. Both studies focused on optimizing the hit-to-lead process, particularly in
designing potent inhibitors of Cyclin-Dependent Kinase 2 (CDK2), but they
approached the problem with distinct methodologies and techniques.
In 2019, Konze et al. [112] explored the use of PathFinder, a reaction-based
enumeration tool, combined with active learning and FEP simulations. This
approach enabled rapid exploration of synthetically tractable chemical spaces and
optimized the potency of CDK2 inhibitors. This study involved generating large
virtual libraries of lead-like compounds through retrosynthetic analysis and combinatorial synthesis. These libraries were then filtered based on their drug-like properties and docked to the CDK2 binding site. The FEP+ tool was used to predict
binding affinities, with active learning iteratively prioritizing compounds for further
FEP calculations. This methodology allowed the exploration of over 300,000 ideas
and the performance of more than 5000 FEP simulations, ultimately identifying over
100 ligands with predicted IC
values below 100 nM.
50

10 Free Energy Perturbation and Free-Energy Calculations Applied to Drug Design 287
In 2020, Ghanakota et al. [33] also integrated ML with FEP calculations but
focused on cloud-based profiling and generative ML for large-scale chemical exploration. Their approach combines extensive enumeration and FEP profiling with
generative ML, resulting in a higher concentration of potent compounds. The
study adhered to a predetermined drug-like property space using PathFinder rulebased enumeration optimized for a multi-parameter function based on a weighted
sum QSAR approach. Employing the REINVENT technique, the authors trained a
network of chemical components and adjusted it to enhance the utility function. This
process created millions of unique compounds at specified R-group positions.
Different strategies for selec ting compounds for FEP calculations were evaluated,
including random enumeration, ML-model-based selection, generative ML prioritization, and combined strategies. The results demonstrated the efficacy of generative
ML in producing novel and potent chemical compounds efficiently, preserving or
improving important physicochemical properties, such as lipophilic ligand efficiency (LLE).
Both studies highlighted the significant impact of integrating ML with FEP in
enhancing drug discovery processes. They leveraged the synergy between ML and
FEP to improve the accuracy and efficiency of drug discovery. Each approach
emphasizes the rapid exploration of large chemical spaces, with Konze et al. [112]
exploring over 300,000 ideas and Ghanakota et al. [33] generating millions of
compounds. Both studies aimed to discover and optimize potent CDK2 inhibitors,
demonstrating the practical application of these techniques in medicinal chemistry.
Additionally, they employed iterative processes in which ML models are continuously updated based on the initial FEP results, refining predictions, and selections.
Furthermore, each study ensured that the generated compounds adhered to drug-like
property criteria, optimizing them for therapeutic potential.
However, there are key differences between these two approaches. Konze et al.
[112] utilized PathFinder for reaction-based enumeration and active learning, focusing on retrosynthetic analysis and combinatorial synthesis, whereas Ghanakota et al.
[33] employed the REINVENT generative ML technique to create unique compounds. Konze et al. [112] explored 300,000 ideas and performed over 5000 FEP
simulations, whereas Ghanakota et al. [33] generated millions of compounds using
cloud computing GPUs for FEP calculations. In terms of selection strategies, Konze
et al. [112] used active learning for iterative improvement, whereas Ghanakota et al.
[33] evalua ted multiple strategies, including random enumeration, ML-based selection, generative prioritization, and combined approaches.
In summary, both studies illustrated the transformative potential of combining
ML with FEP for drug discovery. Konze et al. [112] and Ghanakota et al. [33]
demonstrated different complementary approaches for optimizing chemical libraries
and enhancing the efficiency and effectiveness of the lead optimization process.
Their study underscores the versatility and power of ML and FEP integration, paving
the way for new advancements in computational chemistry and drug development.
An additional utilization of ML in FEP computations was performed by Willow
et al. [113], who demonstrated a substantial improvement in the efficiency and
precision of FEP calculations through the application of ML. The researchers

288 D. Antunes et al.
utilized Targeted Free Energy Perturbation (TFEP), a technique that employs invertible mapping to ensure overlap in the configuration space and convergence in free
energy estimations. The technique was employed on a flexible bonded deca-alanine
molecule, exploiting harmonic biases with different spring centers. An important
element of this method involves employing real-valued non-volume-preserving (real
NVP) transformations, which are highly compatible with TFEP because of their
reliable invertibility and easy calculation of the transformation Jacobian.
An interesting component of this study was the use of an identity map for the
initial configuration of a real NVP map. This was accomplished by setting transformation and scaling factors to zero. The neural network was trained using the
AMBER force field implemented in JAX to minimize the value of the loss function.
This procedure involved partitioning the data into a training set, which accounted for
80% of the data, and a test set, which accounted for 20%. This division further
improves the accuracy of the model. The study observed that the TFEP method
could accurately replicate the reference free energy differences for most state pairs
with a spacing of Δλ = 1 Å. The precision of the mapping approximations was the
greatest for pairs of states with free energy disparities (ΔF) below 2 kJ/mol. Utilizing
trained mapping with early termination demonstrated more reliability, resulting in a
more rigorous calculation of errors compared to conventional approaches.
In the evolving field of drug discovery, the integration of ML with FEP calculations has made significant strides, as evidenced by three distinct studies. The first
study, focusing on the application of ML algorithms such as DeepLDA and
Autoencoders in metadynamics, emphasizes the enhancement in binding mode and
free energy landscape analysis for drug–target interactions. The second study demonstrated the synergy of cloud-based FEP calculations, synthetically aware enumerations, and generative ML in accelerating the hit-to-lead process, notably in
identifying potent compounds and optimizing drug-like properties. Finally, the
third study illustrated the innovative use of ML in targeted FEP, employing learned
mappings for peptide conformations to improve the accuracy and efficiency of FEP
calculations. Collectively, these studies not only highlight the versatility of ML in
different aspects of FEP but also underline its transformative potential in advancing
computational methods for drug discovery and design.
3.4 Implications for ML in FEP Calculations
As we conclude this section on the use of ML in FEP calculations, it is clear that we
are in the bricks of a new era of drug discovery. The future of this field is marked by
ongoing research efforts aimed at refining these ML models, with a specific focus on
enhancing their predi ctive accuracy. Such advancements are pivotal in ensuring that
these models become inte gral components of drug design pipelines .
The integration of ML with FEP calculations has the potential to significantly
transform drug discovery processes. By utilizing these modern computational techniques, researchers can not only accelerate the rate of discovery but also uncover

10 Free Energy Perturbation and Free-Energy Calculations Applied to Drug Design 289
novel compounds that may avoid identification using traditional methods. This
provides opportunities for investigating chemical spaces and comprehending molecular interactions at an unprecedented level.
One major implication of this integration is the increase in the precision of free
energy predictions. ML can significantly enhance the accuracy of FEP calculations
by providing corrections and reducing systematic errors, resulting in more reliable
predictions of compound binding affinity. Additionally, the utilization of ML allows
for the rapid screening of large libraries of compounds, prioritizing those most
promising for detailed FEP simulations, thereby speeding up the discovery process.
The ability to efficiently explore vast chemical spaces is another critical advantage. ML techniques enable the exploration of chemical spaces that are infeasible
with traditional methods, increasing the likelihood of discovering new active compounds. This capability, combined with the enhanced predictive accuracy, can lead
to the identification of novel compounds with potential therapeutic benefits.
Moreover, the integration of ML into FEP calculations can lead to a significant
reduction in drug development costs. By improving the efficiency of the computational process and reducing the need for extensive experimental validation, ML
helps lower the overall cost of drug development. This cost-effectiveness is crucial
in fields where research and development expenses are exceedingly high.
Looking ahead, personalized medicine is an exciting frontier. The application of
ML in FEP could pave the way for personalized medical strategies where drug
compounds are tailored based on individual genetic profiles. This level of customization promises more effective treatments and fewer side effects, thereby revolutionizing patient care.
Additionally, as processing power conti nues to advance and ML algorithms
become more sophisticated, we can expect further improvements in the efficiency
and precision of the FEP calculations. These advancements will reduce both the time
and cost of drug development, thereby making the entire process more streamlined
and accessible.
Essentially, the integration of ML into FEP calculations is more than just an
enhancement of existing methodologies. This revolutionary advancement fundamentally changes the boundaries of what can be achieved in the field of drug
development. As this dynamic area continues to evolve, it holds great potential for
advancing medicine and providing hope for faster and more efficient treatments for a
wide range of diseases.
4 Final Considerations
FEP has emerged as a highly effective tool for computer-assisted drug discovery
owing to advancements in computational hardware, force fields, and methodologies.
FEP allows for precise prediction of ligand binding affinities through alchemical
transformations between different states. The latest developments have significantly
enhanced the range and reliability of the FEP calculations. Despite the emergence of

290 D. Antunes et al.
a plethora of alternative methodologies, FEP remains highly robust and easily
adaptable to new technologies owing to ongoing improvements in force fields,
sampling techniques, and supporting hardware. FEP continues to be the gold
standard for validating binding free energy predictions as it captures the physical
interactions between molecules through statistical thermodynamic principles. Extensive testing has shown the remarkable accuracy of FEP across a wide range of
systems when properly implemented by using theoretical grounds rather than empirical parameterization to relate free energy differences to binding constants.
The integration of various approaches has made it possible to address complex
molecular perturbations such as core modifications, scaffold hopping, and reversible
covalent inhibitors. These perturbations were previously difficult to access. The use
of automated workflows, enhanced sampling techniques, and multi-GPU acc elerations has reduced computational barriers, enabling extensive virtual screens. Furthermore, combining FEP with synergistic machine learning frameworks has
demonstrated the potential of this methodology. FEP enables high-precision prediction of ligand-binding strengths, empowering medicinal chemists to improve the
potency and selectivity of their compounds rationally. Among the computational
techniques, FEP quantifies the thermodynamic forces that drive molecular recognition and directly guides optimization. Recent advancements in efficiency and automation have made it possible for FEP to assess affinities on a large scale, thereby
accelerating the critical progressive improvement needed in drug development.
Advances in structure-based drug optimization are expected to significantly
accelerate and transform this field. The use of increasingly accurate force fields
that incorporate growing experimental data will play a key role in this transformation. Enhanced sampling protocols, which are now tightly integrated, will enable the
simulation of increasingly complex molecular systems at a lower cost. The use of
hybrid grand canonical methodologies holds great promise in obtaining precise
hydration and binding free energies. With the generation of extensive data from
high-throughput experiments, AI-integrated FEP techniques are expected to achieve
new levels of predictive capability. The use of cloud computing and automated
analysis will also facilitate the earlier application of FEP in the discovery pipeline
and reduce the number of iterations required. Despite progress in this field, the
application of FEP to large chemical libraries remains computationally expensive.
The integration of emerging AI capabilities with efficient GPU hardware implementation will be crucial for overcoming this challenge and facilitating the widespread
adoption of this essential methodology.
Overall, free energy perturbation has achieved remarkable progress, earned
widespread trust and reliability, and has preserved its status as the gold standard.
Ongoing advancements have enhanced precision and broadened its potential to drive
drug discovery. The future of this computational method appears promising and
positioned at the forefront of structure-based design.

10 Free Energy Perturbation and Free-Energy Calculations Applied to Drug Design 291
5 First Steps to FEP Simulations
When beginning the FEP experiments, it is important to have a foundational
understanding of the relevant keywords and concepts discussed in this work.
Although much more information can be covered, the scope of this work is limited.
To get started, prospective users should be able to answer the following questions:
Questions to answer when starting a FEP simulation
What computational resources are available, multiple CPUs or GPUs?
Do you have a license for commercially available software or access to open-source software?
Are you aiming to calculate the binding event of a solvated ligand to a protein target (ABFE) or
the relative free energy of binding between two ligands (RBFE)?
For RBFE calculations, what is the size of your ligand set? Do the ligands in your set have
minimal or substantial modifications? Is it possible to create a perturbation map?
In RBFE, are there transformations between ligands with different formal charges? Does this
involve the total annihilation or emergence of groups within the ligands?
Is the binding site well-defined, or does it undergo significant rearrangements?
Are there conserved water molecules or metal ions within the binding site?
Do you have an established binding mode for the ligand(s)? If so, is it a covalent binding mode?
Acknowledgments This study was financed in part by the Coordenação de Aperfeiçoamento de
Pessoal de Nível Superior—Brasil (CAPES)—Finance Code 001.
The authors also thank the agencies CPNq (Processes: 305524/2022-4 and 308254/2022-8) and
FAPERJ (Processes: E-26/201.155/2021 and E-26/201.462/2021) for their financial support.
References
1. King, E., Aitchison, E., Li, H., & Luo, R. (2021). Recent developments in free energy
calculations for drug discovery. Frontiers in Molecular Biosciences, 8, 712085.
2. Guedes, I. A., Pereira, F. S., & Dardenne, L. E. (2018). Empirical scoring functions for
structure-based virtual screening: Applications, critical aspects, and challenges. Frontiers in
Pharmacology, 9, 411637.
3. Shirts, M. R. (2012). Best practices in free energy calculations for drug design. Computational
Drug Discovery and Design, 425–467.
4. Cavasotto, C. N. (2020). Binding free energy calculation using quantum mechanics aimed for
drug lead optimization. In A. Heifetz (Ed.), Quantum mechanics in drug discovery
(pp. 257–268). New York, NY, Springer US.
5. Chen, H., & Chipot, C. (2022). Enhancing sampling with free-energy calculations. Current
Opinion in Structural Biology, 77, 102497.
6. Zwanzig, R. W. (1954). High-temperature equation of state by a perturbation
method. I. Nonpolar gases. The Journal of Chemical Physics, 22, 1420–1426.
7. Kirkwood, J. G. (1935). Statistical mechanics of fluid mixtures. The Journal of Chemical
Physics, 3, 300–313.
8. de Donder, T. (1927). L’affinité. Gauthier-Villars Paris.
9. York, D. M. (2023). Modern alchemical free energy methods for drug discovery explained.
ACS Physical Chemistry Au, 3, 478–491.

292 D. Antunes et al.
10. Bennett, C. H. (1976). Efficient estimation of free energy differences from Monte Carlo data.
Journal of Computational Physics, 22, 245–268.
11. Postma, J. P., Berendsen, H. J., & Haak, J. R. (1982). Thermodynamics of cavity formation in
water. A molecular dynamics study (pp. 55–67). Royal Society of Chemistry.
12. Jorgensen, W. L., & Ravimohan, C. (1985). Monte Carlo simulation of differences in free
energies of hydration. The Journal of Cchemical Physics, 83, 3050–3054.
13. Jorgensen, W. L., & Tirado-Rives, J. (1988). The OPLS [optimized potentials for liquid
simulations] potential functions for proteins, energy minimizations for crystals of cyclic
peptides and crambin. Journal of the American Chemical Society, 110, 1657–1666.
14. Case, D. A., Cheatham, T. E., III, Darden, T., Gohlke, H., Luo, R., Merz, K. M., Jr., et al.
(2005). The Amber biomolecular simulation programs. Journal of Computational Chemistry,
26, 1668–1688.
15. Brooks, B. R., Bruccoleri, R. E., Olafson, B. D., States, D. J., Swaminathan, S. A., & Karplus,
M. (1983). CHARMM: A program for macromolecular energy, minimization, and dynamics
calculations. Journal of Computational Chemistry, 4, 187–217.
16. Merz, K. M., Jr., & Kollman, P. A. (1989). Free energy perturbation simulations of the
inhibition of thermolysin: Prediction of the free energy of binding of a new inhibitor. Journal
of the American Chemical Society, 111, 5649–5658.
17. Cournia, Z., Allen, B., & Sherman, W. (2017). Relative binding free energy calculations in
drug discovery: Recent advances and practical considerations. Journal of Chemical Informa-
tion and Modeling, 57, 2911–2937.
18. Song, L. F., & Merz, K. M., Jr. (2020). Evolution of alchemical free energy methods in drug
discovery. Journal of Chemical Information and Modeling, 60, 5308–5318.
19. Boresch, S., Tettinger, F., Leitgeb, M., & Karplus, M. (2003). Absolute binding free energies:
A quantitative approach for their calculation. The Journal of Physical Chemistry. B, 107,
9535–9551.
20. Aldeghi, M., Heifetz, A., Bodkin, M. J., Knapp, S., & Biggin, P. C. (2016). Accurate
calculation of the absolute free energy of binding for drug molecules. Chemical Science, 7,
207–218.
21. Chen, W., Cui, D., Jerome, S. V., Michino, M., Lenselink, E. B., Huggins, D., et al. (2023).
Enhancing hit discovery in virtual screening through accurate calculation of absolute proteinligand binding free energies. Journal of Chemical Information and Modeling, 63(10),
3171–3185.
22. Ross, G. A., Lu, C., Scarabelli, G., Albanese, S. K., Houang, E., Abel, R., et al. (2023). The
maximal and current accuracy of rigorous protein-ligand binding free energy calculations.
Communications Chemistry, 6, 222.
23. Sun, S., & Huggins, D. J. (2022). Assessing the effect of forcefield parameter sets on the
accuracy of relative binding free energy calculations. Frontiers in Molecular Biosciences,9.
24. Shirts, M. R., Mobley, D. L., & Chodera, J. D. (2007). Chapter 4 Alchemical free energy
calculations: Ready for prime time? In D. C. Spellmeyer & R. Wheeler (Eds.), Annual reports
in computational chemistry (pp. 41–59). Elsevier.
25. Mondal, D., Florian, J., & Warshel, A. (2019). Exploring the effectiveness of binding free
energy calculations. The Journal of Physical Chemistry. B, 123, 8910
26. Fratev, F., & Sirimulla, S. (2019). An improved free energy perturbation FEP+ sampling
protocol for flexible ligand-binding domains. Scientific Reports, 9, 16829.
27. Rocklin, G. J., Mobley, D. L., & Dill, K. A. (2013). Calculating the sensitivity and robustness
of binding free energy calculations to force field parameters. Journal of Chemical Theory and
Computation, 9, 3072–3083.
28. Wang, L., Chambers, J., & Abel, R. (2019). Protein–ligand binding free energy calculations
with FEP+. In M. Bonomi & C. Camilloni (Eds.), Biomolecular simulations: Methods and
protocols (pp. 201–232). New York, NY, Springer New York.
29. Lu, C., Wu, C., Ghoreishi, D., Chen, W., Wang, L., Damm, W., et al. (2021). OPLS4:
Improving force field accuracy on challenging regimes of chemical space. Journal of Chem-
ical Theory and Computation, 17, 4291–4300.
–8915.

10 Free Energy Perturbation and Free-Energy Calculations Applied to Drug Design 293
30. Sugita, Y., & Okamoto, Y. (1999). Replica-exchange molecular dynamics method for protein
folding. Chemical Physics Letters, 314, 141–151.
31. Wang, L., Friesner, R. A., & Berne, B. J. (2011). Replica exchange with solute scaling: A more
efficient version of replica exchange with solute tempering (REST2). The Journal of Physical
Chemistry. B, 115, 9431–9438.
32. Corey, R. A., Vickery, O. N., Sansom, M. S. P., & Stansfeld, P. J. (2019). Insights into
membrane protein–lipid interactions from free energy calculations. Journal of Chemical
Theory and Computation, 15, 5727–5736.
33. Ghanakota, P., Bos, P. H., Konze, K. D., Staker, J., Marques, G., Marshall, K., et al. (2020).
Combining cloud-based free-energy calculations, synthetically aware enumerations, and goaldirected generative machine learning for rapid large-scale chemical exploration and optimization. Journal of Chemical Information and Modeling, 60, 4311–4325.
34. Gilson, M. K., & Zhou, H.-X. (2007). Calculation of protein-ligand binding affinities. Annual
Review of Biophysics and Biomolecular Structure, 36,21–42.
35. Bieniek, M. K., Wade, A. D., Bhati, A. P., Wan, S., & Coveney, P. V. (2023). TIES 2.0: A
dual-topology open source relative binding free energy builder with web portal. Journal of
Chemical Information and Modeling, 63, 718–724.
36. Senapathi, T., Suruzhon, M., Barnett, C. B., Essex, J., & Naidoo, K. J. (2020). BRIDGE: An
open platform for reproducible high-throughput free energy simulations. Journal of Chemical
Information and Modeling, 60, 5290–5295.
37. Kim, S., Oshima, H., Zhang, H., Kern, N. R., Re, S., Lee, J., et al. (2020). CHARMM-GUI free
energy calculator for absolute and relative ligand solvation and binding free energy simulations. Journal of Chemical Theory and Computation, 16, 7207–7218.
38. Zhang, H., Kim, S., Giese, T. J., Lee, T.-S., Lee, J., York, D. M., et al. (2021). CHARMM-GUI
free energy calculator for practical ligand binding free energy simulations with AMBER.
Journal of Chemical Information and Modeling, 61, 4145–4151.
39. He, X., Liu, S., Lee, T.-S., Ji, B., Man, V. H., York, D. M., et al. (2020). Fast, accurate, and
reliable protocols for routine calculations of protein–ligand binding affinities in drug design
projects using AMBER GPU-TI with ff14SB/GAFF. ACS Omega, 5, 4611–4619.
40. Wang, L., Wu, Y., Deng, Y., Kim, B., Pierce, L., Krilov, G., et al. (2015). Accurate and
reliable prediction of relative ligand binding potency in prospective drug discovery by way of a
modern free-energy calculation protocol and force field. Journal of the American Chemical
Society, 137, 2695–2703.
41. Carvalho Martins, L., Cino, E. A., & Ferreira, R. S. (2021). PyAutoFEP: An automated free
energy perturbation workflow for GROMACS integrating enhanced sampling methods. Jour-
nal of Chemical Theory and Computation, 17, 4262 – 4273.
42. Petrov, D. (2021). Perturbation free-energy toolkit: An automated alchemical topology
builder. Journal of Chemical Information and Modeling, 61, 4382–4390.
43. Muegge, I., & Hu, Y. (2023). Recent advances in alchemical binding free energy calculations
for drug discovery. ACS Medicinal Chemistry Letters, 14, 244–250.
44. Kuhn, B., Tichý, M., Wang, L., Robinson, S., Martin, R. E., Kuglstatter, A., et al. (2017).
Prospective evaluation of free energy calculations for the prioritization of cathepsin L inhibitors. Journal of Medicinal Chemistry, 60, 2485–2497.
45. Abel, R., Mondal, S., Masse, C., Greenwood, J., Harriman, G., Ashwell, M. A., et al. (2017).
Accelerating drug discovery through tight integration of expert molecular design and predictive scoring. Current Opinion in Structural Biology, 43,38–44.
46. Gaieb, Z., Parks, C. D., Chiu, M., Yang, H., Shao, C., Walters, W. P., et al. (2019). D3R grand
challenge 3: Blind prediction of protein–ligand poses and affinity rankings. Journal of
Computer-Aided Molecular Design, 33 ,1–18.
47. Cournia, Z., Allen, B. K., Beuming, T., Pearlman, D. A., Radak, B. K., & Sherman, W. (2020).
Rigorous free energy simulations in virtual screening. Journal of Chemical Information and
Modeling, 60, 4153–4169.
48. Scheen, J., Wu, W., Mey, A. S. J. S., Tosco, P., Mackey, M., & Michel, J. (2020). Hybrid
alchemical free energy/machine-learning methodology for the computation of hydration free
energies. Journal of Chemical Information and Modeling, 60, 5331–5339.
Соседние файлы в папке Библиотека им академика М.И. Перельмана
