Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5884_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Foreword
- •Acknowledgments
- •Contents
- •1.1 Structure-Based Drug Discovery (SBDD)
- •1.2 Ligand-Based Drug Design (LBDD)
- •1.3 Echoes from the Past, Visions from the Future
- •References
- •1 Introduction
- •2.2 Second Step: Data Curation
- •2.4 Fourth Step: Updating and Maintenance
- •2 Databases and Curation
- •8 Perspectives
- •9 Conclusion
- •References
- •1 Introduction
- •2.1 Making and Matching Protein Models
- •2.2 Simulating Protein Movements
- •2.3 Analyzing Changes in Protein Shape
- •3 Pharmacogenomics in Drug Development
- •4 Case Studies of Genomics-Based Drug Design
- •References
- •1 Historical Background
- •1.1 Timeline
- •2 Methodology Overview
- •2.1 Neural Networks
- •2.1.1 Perceptron
- •2.1.2 Multilayer Neural Networks
- •2.1.3 Types of Neural Networks
- •Feedforward
- •Recurrent Neural Networks
- •LSTM
- •2.2 Deep Learning
- •3 Using Machine Learning
- •3.2 Data Collection
- •3.3 Data Preprocessing
- •3.4 Model Selection
- •3.5 Model Training
- •3.6 Validation
- •3.7 Tuning
- •3.8 Prediction
- •4 Limitations
- •4.1 Bias
- •4.3 Interpretability
- •4.4 Computational Cost
- •4.5 Data Dependency
- •4.6 Robustness
- •5 Applications in Drug Discovery
- •5.2 Lead Discovery
- •5.3 Preclinical and Clinical Development
- •6 Resources and Tools
- •7 Challenges and Perspectives
- •7.1 Future Trends
- •9 Conclusions
- •References
- •1 Historical Background
- •1.1 Applications in Drug Discovery
- •2 Validations and Controls
- •2.1 Internal Validation
- •2.2 External Validation
- •2.3 Relative Cluster Validation
- •3 Challenges and Perspectives
- •4 Conclusions
- •References
- •1 Historical Background
- •2 OECD Principles
- •2.1 A Defined Endpoint
- •2.2 An Unambiguous Algorithm
- •2.5 A Mechanistic Interpretation, if Possible
- •3 Software and Tools
- •4 Validations and Controls
- •4.1 Internal and External Validation
- •4.1.1 Regression Metrics
- •4.2 Applicability Domain
- •4.3 Randomization Tests
- •5 Interpretation
- •6 Practical Advice During QSAR Modeling
- •7 Application
- •8 Challenges and Perspectives
- •References
- •1 Molecular Docking
- •2 Advances in Scoring Functions and Search Algorithms
- •2.2 Critical Characteristics of Search Algorithms
- •2.3 Docking Programs and Scoring Functions
- •3 Calculations Performed During Docking Simulations
- •4 Essential Components for a Good Docking Program
- •5 Limitations of the Docking Technique
- •6 Validation of Docking Results
- •7 Inappropriate Use of Validation Methods in Docking
- •9 Use of Machine Learning in Molecular Docking
- •11 Challenges
- •12 Conclusions
- •References
- •3 System Preparation for MD Simulations
- •3.1 Solvation and Microensemble
- •3.2 Force Fields: General Concept and Relevant Choices
- •3.3 The Concept of Replicas and Timescale
- •4.1.2 Protein Root Mean Square Fluctuation (RMSF)
- •4.1.4 Protein Secondary Structure Analysis
- •4.1.5 Principal component Analysis (PCA)
- •4.1.6 Markov State Modelling
- •4.1.7 Distance Calculations
- •4.1.8 Angle and Plane Calculations
- •4.2.2 Distances and Ligand-Induced Geometry Rearrangements
- •4 Molecular Dynamics Analysis
- •4.1 Protein Perspective
- •4.1.1 Protein Root Mean Square Deviation (RMSD)
- •4.3 Ligand Perspective
- •4.3.1 Ligand Properties
- •4.3.2 Ligand Root Mean Square Deviation
- •4.3.3 Ligand Root Mean Square Fluctuation
- •4.3.4 Angles and Dihedrals
- •5.1 Protein Structure Prediction and Preparation
- •5.2 Molecular Docking
- •6 Concluding Remarks and Outlook
- •Glossary
- •References
- •1 Introduction
- •2.1 MDeNM
- •2.2 Collective Molecular Dynamics (coMD)
- •2.3 ClustENM and ClustENMD
- •3 Ensemble Docking
- •References
- •1 Introduction
- •1.1 Advantages, Disadvantages, Innovations, and Challenges
- •1.2 Recent Advances in Accessible FEP Software Tools
- •1.3 Applications of FEP in Industry and Consortiums
- •2 Expanding the Potential of FEP Calculations
- •2.1 Validating Binding Poses
- •2.2 Dealing with Solvent
- •2.3 FEP and Allostery
- •2.4 FEP and Covalent Ligands
- •2.5 Applications of FEP in Scaffold Hopping
- •2.6 Positional Analogue Scanning
- •2.7 Combinations and Alternative Approaches
- •3 Machine Learning for FEP
- •3.4 Implications for ML in FEP Calculations
- •4 Final Considerations
- •5 First Steps to FEP Simulations
- •References
- •1 Background
- •2 Ultra-Large Screening Libraries and Chemical Spaces
- •3.1 Implications of Dataset Size
- •4 Ligands on the Ultra-Large Scale
- •4.1 Ultra-Large 2D Similarity Searches
- •7 Challenges and Future Perspectives
- •7.1 Hit Triage: An Old Problem on a New Dimension
- •8 Conclusions
- •Appendix
- •References
- •1 Introduction
- •2 Enzymatic Activity Evaluations
- •3 Cytotoxicity Evaluation and Cell Viability
- •4 Antiviral Assays in Experimental Validation
- •6 In Vivo Evaluation of Compounds
- •7 Conclusions
- •References
- •1 Introduction
- •3.1 Data Collection
- •3.2 Data Preprocessing
- •3.4 Model Choice
- •3.5 Model Training
- •3.6 Model Assessment
- •3.7 External Validation
- •3.8 Implementation and Availability
- •3.9 Continuous Update
- •5 Conclusions and Perspectives
- •References
- •1 Experimental Approaches to Obtain Protein Structure
- •1.1 X-Ray Crystallography
- •1.2 Nuclear Magnetic Resonance
- •1.3 Cryo-EM
- •1.4 Hybrid Methods
- •2 Modeling Approaches to Obtain Protein Structure
- •2.1 Homology Modeling
- •2.2 Ab Initio Modeling
- •2.3 New Approaches
- •3 Conformational Diversity of Proteins
- •3.1 Characterization of Protein Conformational States
- •3.2 Experimental Methods to Study Protein Dynamics and Conformations
- •3.4 Molecular Dynamics Simulation
- •3.5 Sampling Strategies
- •4 Remarks and Perspectives
- •References
- •1 Introduction
- •2 Structure-Based Drug Design of HIV Protease Inhibitors
- •2.1 HIV-1 Protease as a Therapeutic Target
- •2.2.1 Saquinavir
- •2.2.2 Indinavir
- •2.3.1 Lopinavir
- •2.3.2 Darunavir
- •6 Conclusions
- •References
- •4 Experimental Methods to Analyze NR Activity
- •4.2 Coregulator-Recruitment
- •5 Concluding Remarks and Outlook
- •References

264 D. Antunes et al.
1 Introduction
Drug design is the process of creating new pharmaceutical compounds or optimizing
existing ones to target specific biological molecules or disease-related processes in
the human body. Its main goal is to develop safe and effective medications that
interact with specific molecular targets, such as proteins, enzymes, or receptors, and
modulate their activity to treat or prevent diseases.
The drug discovery process consists of several crucial stages including target
identification and validation, hit identification, hit expansion, hit-to-lead, and lead
optimization. In the first stage, target identification and validation, the focus is on
identifying the target molecule, usually a protein or another molecule involved in the
disease process. The second stage, hit identification, involves testing thousands of
compounds to identify potential candidates for therapeutic treatment. Compounds
are evaluated through high-throughput screening and other methods to determine
their efficacy against the target. In the hit expansion stage, promising compounds
from the hit identification phase undergo further evaluation and testing to expand
their potential as lead candidates. This stage involves refining the selection of
compounds based on their activity, selectivity, and other desirable properties. The
hit-to-lead phase involves optimizing the selected compounds to enhance their
affinity, selectivity, efficacy, metabolic stability, and oral bioavailability. Lead
optimization is the final stage, in which lead compounds undergo further optimization to improve their drug-like properties and enhance their potential as viable drug
candidates. This phase involves iterative cycles of medicinal chemistry to refine the
lead compounds and address issues related to their efficacy, safety, and pharmacokinetic properties.
The utilization of free-energy methods can yield essential thermodynamic and
kinetic data through rigorous computational approaches. However, until a few years
ago, high computing costs, sampling method restrictions, force field limitations, and
a lack of automation prevented the widespread use of free-energy methods.
Over the last decade, considerable emphasis has been placed on enhancing the
efficiency and feasibility of these approaches using workflows that utilize modern
CPU- and GPU-based architectures. Therefore, from both academic and industrial
standpoints, the utilization of free-energy calculation workflows has emerged as a
compelling computational tool to facilitate the progress of drug discovery.
Today, scientists use a variety of methods to calculate drug design binding free
energies. These include alchemical calculations, endpoint methodologies, empirical
scoring functions, knowledge-based potentials, and quantum mechanics. Free
energy perturbation (FEP) and thermodynamic integration (TI) methods are considered the “gold standard” for accurate in silico potency predictions, as they use
alchemical transformations to calculate the relative binding free energies between
two ligands. Second-group methods, such as molecular mechanics PoissonBoltzmann surface area (MM-PBSA) and generalized Born surface area
(MM-GBSA, see Chap. 8), estimate the binding free energy from the difference
between bound and unbound free energies, making them computationally less
expensive than alchemical methods [1].

10 Free Energy Perturbation and Free-Energy Calculations Applied to Drug Design 265
Empirical scoring functions can also be used to evaluate the binding affinity using
weighted terms representing hydrogen bonding, van der Waals interactions, and
desolvation [2]. While they are fast, they are less accurate than other approaches.
Additionally, knowledge-based statistical potentials can be derived from the analysis
of known protein– ligand complex structures, which are computationally efficient but
require structural data [3]. Finally, quantum mechanical methods such as the fragment molecular orbital (FMO) method use quantum mechanical calculations to
evaluate the binding free energy, making them more precise but computationally
intensive than classical force fields [4].
Other techniques utilize enhanced sampling of the free energy landscape, such as
Metadynamics, ABF, and umbrella sampling, to explore the complex energy landscapes of biomolecular systems. These methods are more effective than standard
molecular dynamics simulations for overcoming the limitations imposed by rare
events [5].
Out of these above-mentioned methods, the methodology called Free Energy
Perturbation (FEP) plays a significant role in the computational drug design step,
helping researchers select and optimize potential drug candidates with improved
binding properties and selectivity.
In drug desig n, FEP has emerged as a valuable tool for estimating the binding free
energy of ligands to proteins, which is a critical determinant of their potential as
successful drug candidates. This represents a robust and thermodynamically rigorous
computational approach capable of predicting the binding affinity of small molecules, given that there is a direct relationship between the free energy and affinity of a
ligand, as shown in Eq. (10.1):
ΔG
=-kT ln K
bind
i
ð10:1Þ
where ΔG represents the binding free energy, calculated as the difference in free
energy between the bound and unbound states of the ligand–target complex, k is the
Boltzmann constant, K
K
could also be used), and T is the absolute temperature.
d
is the inhibition constant (although the dissociation constant
i
FEP’s application in drug design extends the exploration and optimization of
potential drug candidates across a vast chemical space, thereby facilitating the
enhancement of multip le properties. Moreover, FEP’s computationally driven
approach can significantly reduce the cost and time associated with experimental
testing by e nabling the screen ing of less promisi ng candidates in silico.
Many theoretical methodologies have significantly paved the way for the development of FEP, particularly prior to the seminal computational breakthroughs.
Although most scientific sources attribute the FEP method to Zwanzig’s[6], previ-
ous works by Kirkwood [7] and De Donder [8] contributed to setting the theoretical
basis by introducing the generalized-extent parameter (λ), which reconciles statistical mechanics with the degree of evolution in a chemical reaction. Zwanzig derived
an expression for the free energy difference between the two states of a system in
terms of the probability distribution of the perturbation energy. The free energy
difference for going from state A to state B is obtained from the Zwanzig equation
(Eq. 1 0.2)

266 D. Antunes et al.
ΔFA→ BðÞ= FB- FA=-kBT ln exp
UB- U
kBT
A
A
, ð10:2Þ
where F denotes the free energy, U represents the internal energy of the system, T is
the temperature, k
is Boltzmann's constant, and the angular brackets denote the
B
average over a simulation run for state A. Briefly, in statistical mechanics, the phase
space encompasses every conceivable arrangement of positions and momenta for
atoms in a simulated system [9]. FEP utilizes a succession of intermediate
overlapping states, referred to as λ-windows, which are determined by a coupling
parameter λ that connects the potential energy functions of states A and B. The
transition from A to B is accomplished by incrementally adjusting λ from 0 to 1 in a
specified number of discrete steps, both forward and backward.
Subsequently, Valleu and Card [10] introduced a stratification strategy that
connects the reference and target states, breaking the total free-energy difference
into the sum of the finite free-energy differences between the intermediate states with
increased overlap. In 1976 [10], Bennett independently developed a method called
the Bennett Acce ptance Ratio (BAR) and improved the efficiency and reliability of
FEP calculations using a weighted average of forward and reverse transformations.
Soon after the beginning of the 1980s, almost 30 years after Zwanzig’s equation
was presented to the scientific community, the first successful attempt to use the FEP
appeared. Postma et al., in 1982 [11], reported FEP calculations on forming a cavity
in an explicit water simulation box. In 1985, Jorgensen et al. [12 ] used alchemical
transformations of alkanes into alcohols to calculate their hydration-free energies
with high precision. This method involves gradually transforming the ligand into the
final compound, and then calculating the free energy change associated with each
transformation step. This work is considered a cobblestone of FEP application in
drug design because it hints at its potential given that the relative solvation-free
energies play a major role in determining the relative binding free energy of two
ligands at a common receptor site.
The introduction of the OPLS force field by Jorgensen and Tirado-Rives in 1988
[13] improved the accuracy and transferability of the molecular simulations. Similarly, other force fields develo ped simultaneously have incorporated the FEP methodology into the calculations [14, 15]. In the early 1990s, Kollman et al. [16]
employed FEP methodology integrated into the AMBER program to investigate
the binding energies of a range of thermolysin inhibitors.
Among the many contr ibutions that consol idate Structure-Based Drug Discovery,
such as the understanding of molecular interactions and the design of novel drug
candidates based on structural information or molecular docking, which cleared the
way for virtual screening, FEP calculations can be considered as one of the early and
fundamental contributions to the field. Extensive reviews have thoroughly detailed
the development timeli ne for a comprehensive historical account of FEP methodology [10, 11].

10 Free Energy Perturbation and Free-Energy Calculations Applied to Drug Design 267
Fig. 10.1 The thermodynamic cycle in which molecule A alchemically transforms into molecule
B at the protein binding site and the solvent (water)
From the beginning of 2000, the FEP technique has demonstrated a persistent
trajectory of evolution and enhancement. Researchers have developed novel methodologies for the computation of ligand-protein binding free energy, thereby innovating approaches that harness FEP for novel drug design. This escalating
computational potency has propelled FEP to become an indispensable instrument
in the realm of drug discovery.
The predominant approach for implementing FEP in drug design entails the
utilization of the alchemical transformation technique, which involves a thermodynamic cycle that establishes a correlation between the bound and free states. In
thermodynamics, the free energy is a state function; consequently, the total variation
throughout the cycle is zero. For example, if two ligands (A and B) are bound to an
identical receptor, the free energy difference between them can be estimated by
transforming one ligand into another through a nonphys ical process, considering the
medium in which they are immersed (Fig. 10.1). The thermodynamic cycle is an
effective method for determining relative binding free energies.
ΔG
- ΔGA= ΔG
B
prot
A → B
- ΔG
wat
A → B
ð10:3Þ
The terms on the left-hand side of Eq. (10.3)reflect the variation in the ΔG values
of interest (ΔΔG), whereas the subsequent term indicates the alchemical transformation used to compute the difference. Calculating the ΔΔG is crucial in drug
design because it offers precise information on how structural modifications or
replacements in compounds affect the energy difference. By determining the sign
of ΔΔG, it is possible to ascertain whether modifications in compounds result in an

268 D. Antunes et al.
increase (negative ΔΔG) or decrease (positive) in the molecule’saffinity for the
receptor. This information is essential for the development of effective drugs.
Undoubtedly, utilizing the aforementioned equation directly can lead to the
emergence of artifacts and noise in simulations. To mitigate this impact, it is
imperative to gradually convert the ligand into its final form, thereby enabling the
calculation of the free energy changes at each subsequent stage of the transformation. Consequently, the sum of these free-energy differences culminates in an
estimate of the binding free energy.
Although the thermodynamic cycle is more commonly used for determining the
relative binding free energies, it can be adapted to estimate the absolute binding free
energies, in which the ligands are annihilated in the binding site and solvent
[19, 20]. In this case, the entire ligand is coupled/decoupled, which results in a
significantly larger perturbation of the system, longer sampling times are necessary
to achieve convergence compared to the sampling required to converge ΔΔG
estimates.
The FEP approach is highly reliable and p roduces accurate results compared with
experimental data [21]. Nevertheless, successful implementation of this approach
requires a systematic and meticulous methodology.
This chapter provides an overview of the latest advancements and challenges in
the application of free-energy perturbation in drug design. It covers topics such as
improved utilities for FEP calculations, including force fields and enhanced
approaches for system preparation, strategies for addressing solvent, allosteric, and
covalent ligand cases in FEP, and the utilization of Machine Learning (ML) in FEP
calculations.
1.1 Advantages, Disadvantages, Innovations, and Challenges
FEP offers numerous advantages for drug design and serves as a reliable technique
for computing the binding free energy between proteins and diverse compounds
when ABFE is calculated, or a congeneric series of compounds for RBFE calculations. Owing to its thermodynamic rigor, the FEP method can yield more precise
predictions than other computational methods [22], aiding in the creation of more
powerful and selective drugs.
However, implementing FEP in realistic systems, such as protein–ligand complexes, poses notable challenges and limitations that require resolution. One prominent hurdle is the substantial demand for computational resources for FEP
calculations, which renders it a costly and time-intensive methodology. Consequently, their widespread application in large-scale virtual screening remains limited
(see Chap. 11 for fundamentals and other aspects of ultra-large virtual screening).
Another challenge arises from the reliance on different force fields, which can lead to
results that are difficult to comprehend and compare [23].

10 Free Energy Perturbation and Free-Energy Calculations Applied to Drug Design 269
It is crucial to note that FEP typically involves nonphysical (“alchemical”)
perturbations [24], gradually transforming one ligand into another through a series
of steps. However, although FEP shows satisfactory performanc e with ligands that
have minimal structural changes, its accuracy decreases when dealing with substantial modifications [25].
For instance, when substantial structural rearrangements are essential, FEP calculations may not always produce outcomes independent of the structure. This is
particularly apparent when intricate changes, such as loop and backbone movements, are crucial for the ligand-binding process [26].
Accurate sampling is a pivotal factor in FEP simulations. However, a notable gap
exists in comprehensive studies aimed at defining an optimal sampling duration,
representing a primary limitation of the FEP calculations [25]. Inadequate equilibration and coexistence of various stable binding conformations further underline the
critical areas that require enhancement. In many cases, FEP calculations may be
contingent upon the structure and may lack reliability, especially when simulations
start from unknown or undefined binding poses. Sensitivity to the choice of force
field parameters and various other considerations further accentuates the potential
limitations of FEP calculations in certain scenarios [27]. For example, choosing an
appropriate initial ligand–protein structure is crucial for achieving accuracy in FEP,
especially in dynamic systems, when exploring an ensemble of structures to identify
the optimal receptor conformation [22]. Furthermore, it is critical to determine the
equilibration timescale for each window to ensure that the free energies converge.
The coupling parameter λ influences the alchemical transformation between states,
and it should be sampled with sufficient intermediate windows to capture the free
energy changes accurately [27].
Recent years have witnessed notable progress and innovations in free-energy
perturbation and its role in drug design. These advancements have been propelled by
enhancements in both computational hardware and FEP methodologies, enabling the
extensive evaluation of the accuracy and dependability of FEP calculations in drug
discovery initiatives. Key developments encompass the establishment of FEP+
technology by Schrödinger Inc., detailed in https://www.sc hrodinger.com/science-
articles/free-energy-methods-fep and [28]. FEP+ integrates cutting-edge force fields,
advanced sampling techniques, and GPU acceleration to facilitate precise and
reliable computations of protein–ligand binding free energies, thereby finding applications in d rug discovery.
Moreover, FEP+ incorporates the latest OPLS force fields, such as OPLS4 and
OPLS5, which offer comprehensive coverage of the chemical space for both drug
discovery and materials science applications. These force fields build upon the
extensive coverage and accuracy achieved in previous OPLS versions by improving
the accuracy of functional groups, which have presented significant modeling
challenges in the past. The new OPLS5, in particular, is a polarizable force field
that improves the relative bindi ng accuracy in the FEP+ and Desmond models by
adding explicit polarization for polarizable atoms, molecular ions, and cation–pi
interactions [29].

270 D. Antunes et al.
An additional advancement is the FEP/REST method [30], which introduced an
efficient λ-hopping protocol meticulously designed to sample local structural
rearrangements, facilitating the assessment of relative protein–ligand binding affinity within manageable simulation durations. REST is particularly helpful in cases
where there are significant binding site rearrangements upon ligand binding, or when
studying a series of diverse ligands.
Additionally, REST2 [31], an advancement in enhanced sampling, was designed
to speed up the traversal of the phase space and accelerate convergence, which is
especially beneficial for flexible binding sites and those exhibiting a relevant degree
of induced fit. REST2 significantly reduces CPU demands compared to regular
replica exchange, substantially enhancing sampling efficiency, particularly when
addressing substantial solute conformational changes in aqueous protein solutions. It
is important to note that the computational load may vary between CPU and GPU
implementation. This innovation has been seamlessly integrated into FEP+, thereby
broadening its scope and effectiveness.
A recent extension of the FEP+ method has enabled handling of challenging
perturbations, including core-hopping transformations, macrocycle modifications,
and optimization of reversible covalent inhibitors [28]. Specifically, in the domain of
macrocyclic drugs, FEP has demonstrated its utility in the design and optimization of
such drugs by capturing their conformat ional flexibility a nd diversity.
Furthermore, the use of FEP to study membrane proteins is promising. FEP aids
in predicting ligand-binding affinity and specificity of membrane proteins by
accounting for intricate environmental factors and interactions present in the membrane environment [32]. Significant progress has also been made in the application
of FEP to reversible covalent inhibitors, a class of drugs that form reversible bonds
with target proteins to modulate their activity and optimize drug function.
Finally, the Machine Learning (ML, see Chap. 4) methods have significantly
impacted the realm of FEP applied to drug design. Notably, a recent study integrated
cloud-based FEP calculations with synthetically aware enumerations and goaldirected generative machine learning to facilitate extensive chemical exploration
and optimization on a large scale [33]. Cloud-based free-energy calculations are a
computational methodology employed to estimate the binding affinity of molecules
to a target protein. This approach harnesses the computational power of cloud
computing, allowing parallel execution of large-scale simulations. Consequently, it
reduces computational expenses and time. Moreover, these calculations can be
synergistically combined with other methodologies such as synthetically aware
enumerations and goal-directed generative machine learning. This combined
approach enables exploration and optimization across vast chemical spaces [33].
1.2 Recent Advances in Accessible FEP Software Tools
FEP has become one of the most promising approaches for accurately predicting
ligand-binding affinities and is a crucial application in computer-aided drug design
and discovery [34]. Through simulation of the transformation between the initial and

10 Free Energy Perturbation and Free-Energy Calculations Applied to Drug Design 271
final end states along a coupling parameter λ, FEP allows for rigorous calculation of
changes in the binding free energy based on statistical thermodynamics
[34]. Although the theory behind FEP has been firmly established, practical applications on a large scale have faced hindrances owing to complexities in system
preparation, simulation protocols, and convergence of computations [17, 18].
Nevertheless, noteworthy developments have been made in recent years to
broaden the horizons and decrease the barriers to utilizing FEP in real-world drug
discovery campaigns. Open-source tool s, such as TIES 2.0 [35], aid in free-energy
calculations through flexible alignment algorithms and automated workflows accessible via web portals. TIES 2.0 offers the capability to perform both FEP calculations
and thermodynamic integration (TI), which is another method for computing free
energy differences. TIES 2.0 specifically applies a dual-topology method to predict
relative binding free energies using TI, as demonstrated using sets of congeneric
ligands [36]. Although the practical aspects of the FEP and TI are similar, the
underlying theory and implementation of the partial derivatives of the Hamiltonian
with respect to the coupling parameter λ differ between the two methods. Platforms
such as BRIDGE [4] enhance the reproducibility and sharing of FEP protocols by
integrating codes from GROMACS and YANK into Docker containers, which
operate seamlessly across various computing environments. The effectiveness of
BRIDGE’s capabilities was demonstrated by the discovery of drug targets, including
cyclin-dependent kinase 2 (CDK2) and ST3Gal-I, where both absolute and relative
FEP were combined [36].
Notably, efforts have been made to simplify the intricate setup process of FEP
simulations, even with numerous variations in ligands or receptors [35]. CHARMMGUI modules have been expanded to generate input files and analysis scripts for FEP
via various molecular dynamics engines such as NAMD, GENESIS, and AMBER
[37–39]. Automated workflow capabilities were verified by testing diverse ligand
solvation and protein binding systems. Notably, the AMBER implementation facilitates a range of force- field combinations and advanced options, such as hydrogen
mass repartitioning, to expedite FEP convergence [39]. Studies, including those
benchmarked with BACE1, have demonstrated the essential role of multiple independent FEP runs in achieving statistically reliable binding free energies despite
minor protocol variations [40]. These tools provide high-throughput FEP and make
the methodology accessible to nonspecialist researchers [37 – 39].
Additional efforts have been dedicated to improving the automation of
GROMACS simulation packages, one of the most widely used molecular dynamics
software programs. The open-source Python tool PyAutoFEP enables adaptable
configuration of alchemical free-energy calculations in GROMACS, enabling perturbation mapping between numerous ligand states, and integration of advanced
sampling techniques such as replica exchange [41]. As demonstrated by a set of
Farnesoid X receptor compounds, PyAutoFEP enables large-scale prediction of
relative binding affinities that are comparable to those of the top-performing
methods. In addition to PyAutoFEP, the SMArt Python package provides automation capabilities via a perturbation topology builder that employs graph theory
algorithms [42]. By identifying the maximum number of common substructures,

272 D. Antunes et al.
SMArt defines an optimal transformation pathway and generates compatible topology files for GROMACS and GROMOS simulation packages. When applied to a set
of lysine post-translational modifications, the perturbation topologies generated by
SMArt yielded consistent free energy differences between simulations performed
with the two simulation packages, except for perturbations involving net charge
changes. The thermodynamic cycle closures obtained from these calculations were
robust, with a value of 0.5 ± 0.3 kJ mol
-1
for GROMOS and 0.2 ± 0.2 kJ mol-1for
GROMACS, indicating that SMArt produces precise and compatible alchemical
transformations for both simulation packages.
These tools offer automated workflows tailored to GROMA CS, making them
more accessible. PyAutoFEP and SMArt effortlessly manage topology generation
and analysis, making alchemical free energy methods easier to apply. The incorporation of improved sampling and multistate capabilities expands the range and
dependability of free-energy predictions and, as open-source platforms, stimulates
additional development and customization. PyAutoFEP and SMArt are prime examples of targeted endeavors to unleash the potential of FEP in practical applications.
Precise prediction of subtle energy differences between structurally similar ligand
conformations or binding modes has become an essential tool for computer-aided
drug optimization, and recent advances have strengthened its potential [43]. Opensource, user-friendly platforms permit the evaluation and exchange of optimal FEP
protocols, whereas the automation of system preparation minimizes workflow barriers on a large scale. Combined with increased computational capabilities and
improved force fields, FEP has immense potential for significantly enhancing and
hastening molecular discovery and design.
1.3 Applications of FEP in Industry and Consortiums
The application of FEP and free-energy calculations has expanded beyond the
academic realm, becoming an integral part of the drug disco very and development
processes within the pharmaceutical industry. These computational techniques have
emerged as powerful tools for predicting the binding affinities of drug candidates for
their target proteins, thereby streamlining the overall drug discovery pipeline.
In the industrial context, FEP is extensively employed in the optimization of lead
compounds. For example, the FEP+ tool developed by Schrödinger has been
successfully integrated into the workflow of numerous pharmaceutical companies,
enabling the accurate prediction of relative binding free energies and guiding
medicinal chemists towards compounds with improved efficacy and selectivity
[40]. Major pharmaceutical companies, including Merck, Novartis, and Pfizer,
have integrated FEP calculations into their drug design pipelines, leveraging these
methods to prioritize compounds for synthesis and biological testing, thereby
enhancing the efficiency of drug development [44 , 45].

10 Free Energy Perturbation and Free-Energy Calculations Applied to Drug Design 273
Furthermore, industry–academia collaborations such as the Drug Design Data
Resource (D3R) initiative have fostered the application and advancement of FEP
methodologies in drug discovery [46]. These collaborative endeavors have led to
significant advancements in computational techniques and their validation using
experimental data, thereby underscoring the growing recognition of the utility of
FEP in the industry.
The adoption of FEP in the pharmaceutical industry is driven by its unique
advantages. Primarily, the cost-effectiveness of FEP is evident in its ability to reduce
the number of compounds that require physical synthesis and experimental testing,
leading to substantial savings. Second, the speed at which FEP can provide reliable
predictions of binding affinities accelerates drug discovery. Finally, the enhanced
precision of FEP in lead optimization aids in the development of more effective and
selective drug candidates [47].
2 Expanding the Potential of FEP Calculations
The FEP methodology can be utilized to obtain a comprehensive understanding of
the intricate biological circumstances that necessitate a thorough analysis of the
underlying context. For instance, this method can be utilized to validate the accuracy
of molecular binding positions by calculating the relative changes in the binding
affinity between a group of molecules through nonphysical modifications. The
accuracy of FEP calculations relies on several parameters, including the structural
integrity of the prote in, precise positioning of the ligand, and affinity range and
appropriateness of the ligands used for FEP calculations. The presence of water
molecules at the binding site can drastically affect the estimation of the Gibbs free
energy (ΔG). Furthermore, the prediction of the affinity is significantly affected by
whether the binding is covalent. The stru ctural attributes of a ligand set can determine the level of easiness or difficulty of using this technology. The need for
computational resources is another important factor to consider when using FEP,
mainly because of the extensive sampling of the relevant configurations of the
system. Hence, although FEP is a powerful instrument for clarifying intricate
biological situations, its utilization requires meticulous deliberation of a particular
context and accessible resources.
Currently, machi ne learning (ML) is incorporated into free energy perturbation
(FEP) computations to enhance precision and effectiveness. In particular, one
approach utilizes ML-derived correction terms to improve the accuracy of FEP
predictions, as demonstrated in Scheen et al. [48], where ML was used to generate
corrections for hydration free-energy calculations, leading to more precise results
than standalone FEP methods. Another significant application involves the use of
active learning (AL), a special case of ML, to prioritize molecules from large
datasets for FEP calculations. Thompson et al. [49] showcased an AL framework
that efficiently identified top-scoring molecules by iteratively sampling and updating
the ML model, thereby optimizing the selection process and reducing the
Соседние файлы в папке Библиотека им академика М.И. Перельмана
