Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5884_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Foreword
- •Acknowledgments
- •Contents
- •1.1 Structure-Based Drug Discovery (SBDD)
- •1.2 Ligand-Based Drug Design (LBDD)
- •1.3 Echoes from the Past, Visions from the Future
- •References
- •1 Introduction
- •2.2 Second Step: Data Curation
- •2.4 Fourth Step: Updating and Maintenance
- •2 Databases and Curation
- •8 Perspectives
- •9 Conclusion
- •References
- •1 Introduction
- •2.1 Making and Matching Protein Models
- •2.2 Simulating Protein Movements
- •2.3 Analyzing Changes in Protein Shape
- •3 Pharmacogenomics in Drug Development
- •4 Case Studies of Genomics-Based Drug Design
- •References
- •1 Historical Background
- •1.1 Timeline
- •2 Methodology Overview
- •2.1 Neural Networks
- •2.1.1 Perceptron
- •2.1.2 Multilayer Neural Networks
- •2.1.3 Types of Neural Networks
- •Feedforward
- •Recurrent Neural Networks
- •LSTM
- •2.2 Deep Learning
- •3 Using Machine Learning
- •3.2 Data Collection
- •3.3 Data Preprocessing
- •3.4 Model Selection
- •3.5 Model Training
- •3.6 Validation
- •3.7 Tuning
- •3.8 Prediction
- •4 Limitations
- •4.1 Bias
- •4.3 Interpretability
- •4.4 Computational Cost
- •4.5 Data Dependency
- •4.6 Robustness
- •5 Applications in Drug Discovery
- •5.2 Lead Discovery
- •5.3 Preclinical and Clinical Development
- •6 Resources and Tools
- •7 Challenges and Perspectives
- •7.1 Future Trends
- •9 Conclusions
- •References
- •1 Historical Background
- •1.1 Applications in Drug Discovery
- •2 Validations and Controls
- •2.1 Internal Validation
- •2.2 External Validation
- •2.3 Relative Cluster Validation
- •3 Challenges and Perspectives
- •4 Conclusions
- •References
- •1 Historical Background
- •2 OECD Principles
- •2.1 A Defined Endpoint
- •2.2 An Unambiguous Algorithm
- •2.5 A Mechanistic Interpretation, if Possible
- •3 Software and Tools
- •4 Validations and Controls
- •4.1 Internal and External Validation
- •4.1.1 Regression Metrics
- •4.2 Applicability Domain
- •4.3 Randomization Tests
- •5 Interpretation
- •6 Practical Advice During QSAR Modeling
- •7 Application
- •8 Challenges and Perspectives
- •References
- •1 Molecular Docking
- •2 Advances in Scoring Functions and Search Algorithms
- •2.2 Critical Characteristics of Search Algorithms
- •2.3 Docking Programs and Scoring Functions
- •3 Calculations Performed During Docking Simulations
- •4 Essential Components for a Good Docking Program
- •5 Limitations of the Docking Technique
- •6 Validation of Docking Results
- •7 Inappropriate Use of Validation Methods in Docking
- •9 Use of Machine Learning in Molecular Docking
- •11 Challenges
- •12 Conclusions
- •References
- •3 System Preparation for MD Simulations
- •3.1 Solvation and Microensemble
- •3.2 Force Fields: General Concept and Relevant Choices
- •3.3 The Concept of Replicas and Timescale
- •4.1.2 Protein Root Mean Square Fluctuation (RMSF)
- •4.1.4 Protein Secondary Structure Analysis
- •4.1.5 Principal component Analysis (PCA)
- •4.1.6 Markov State Modelling
- •4.1.7 Distance Calculations
- •4.1.8 Angle and Plane Calculations
- •4.2.2 Distances and Ligand-Induced Geometry Rearrangements
- •4 Molecular Dynamics Analysis
- •4.1 Protein Perspective
- •4.1.1 Protein Root Mean Square Deviation (RMSD)
- •4.3 Ligand Perspective
- •4.3.1 Ligand Properties
- •4.3.2 Ligand Root Mean Square Deviation
- •4.3.3 Ligand Root Mean Square Fluctuation
- •4.3.4 Angles and Dihedrals
- •5.1 Protein Structure Prediction and Preparation
- •5.2 Molecular Docking
- •6 Concluding Remarks and Outlook
- •Glossary
- •References
- •1 Introduction
- •2.1 MDeNM
- •2.2 Collective Molecular Dynamics (coMD)
- •2.3 ClustENM and ClustENMD
- •3 Ensemble Docking
- •References
- •1 Introduction
- •1.1 Advantages, Disadvantages, Innovations, and Challenges
- •1.2 Recent Advances in Accessible FEP Software Tools
- •1.3 Applications of FEP in Industry and Consortiums
- •2 Expanding the Potential of FEP Calculations
- •2.1 Validating Binding Poses
- •2.2 Dealing with Solvent
- •2.3 FEP and Allostery
- •2.4 FEP and Covalent Ligands
- •2.5 Applications of FEP in Scaffold Hopping
- •2.6 Positional Analogue Scanning
- •2.7 Combinations and Alternative Approaches
- •3 Machine Learning for FEP
- •3.4 Implications for ML in FEP Calculations
- •4 Final Considerations
- •5 First Steps to FEP Simulations
- •References
- •1 Background
- •2 Ultra-Large Screening Libraries and Chemical Spaces
- •3.1 Implications of Dataset Size
- •4 Ligands on the Ultra-Large Scale
- •4.1 Ultra-Large 2D Similarity Searches
- •7 Challenges and Future Perspectives
- •7.1 Hit Triage: An Old Problem on a New Dimension
- •8 Conclusions
- •Appendix
- •References
- •1 Introduction
- •2 Enzymatic Activity Evaluations
- •3 Cytotoxicity Evaluation and Cell Viability
- •4 Antiviral Assays in Experimental Validation
- •6 In Vivo Evaluation of Compounds
- •7 Conclusions
- •References
- •1 Introduction
- •3.1 Data Collection
- •3.2 Data Preprocessing
- •3.4 Model Choice
- •3.5 Model Training
- •3.6 Model Assessment
- •3.7 External Validation
- •3.8 Implementation and Availability
- •3.9 Continuous Update
- •5 Conclusions and Perspectives
- •References
- •1 Experimental Approaches to Obtain Protein Structure
- •1.1 X-Ray Crystallography
- •1.2 Nuclear Magnetic Resonance
- •1.3 Cryo-EM
- •1.4 Hybrid Methods
- •2 Modeling Approaches to Obtain Protein Structure
- •2.1 Homology Modeling
- •2.2 Ab Initio Modeling
- •2.3 New Approaches
- •3 Conformational Diversity of Proteins
- •3.1 Characterization of Protein Conformational States
- •3.2 Experimental Methods to Study Protein Dynamics and Conformations
- •3.4 Molecular Dynamics Simulation
- •3.5 Sampling Strategies
- •4 Remarks and Perspectives
- •References
- •1 Introduction
- •2 Structure-Based Drug Design of HIV Protease Inhibitors
- •2.1 HIV-1 Protease as a Therapeutic Target
- •2.2.1 Saquinavir
- •2.2.2 Indinavir
- •2.3.1 Lopinavir
- •2.3.2 Darunavir
- •6 Conclusions
- •References
- •4 Experimental Methods to Analyze NR Activity
- •4.2 Coregulator-Recruitment
- •5 Concluding Remarks and Outlook
- •References

7 Molecular Docking: State-of-the-Art Scoring Functions and Search Algorithms 173
Fig. 7.3 Diagram representing the search algorithms that can be used to explore ligand flexibility in
a docking simulation
2.3 Docking Programs and Scoring Functions
Scoring functions are mathematical algorithms used in molecular docking simulations to assess and rank different ligand confor mations in the target protein. The
main goal of scoring functions is to estimate the affinity or binding energy between
the ligand and the protein in a given conformation. Table 7.3 lists some of the most
commonly used molecular docking software, the scoring functions implemented in
each, and their potential applications [56 ].
It is important to note that most docking programs offer options to customize
scoring functions or allow users to define their energy terms, depending on the
specific needs of their research. Additionally, the choice of software and scoring
function should be based on the characteristics of the target protein, the ligand of
interest, and the objectives of the docking simulation [71]. Below is an overview of
scoring functions and the calculations involved. Scoring functions are based on an
energy function that models interactions between the ligand and the protein. This
energy function consists of terms describing various molecular effects, such as:
• van der Waals energy: related to London dispersion forces representing attraction
between atoms due to temporary fluctuations in their electron clouds , as well as
steric repulsion forces preventing ligand atoms from overlapping with protein
atoms, ensuring physically feasible binding.
• Hydrogen bonding energy: considers hydrogen bonds formed between the ligand
and critical residues in the protein’s active site.
• Electrostatic energy: related to interactions between electric charges, such as
ionic interactions.

174 R. M. de Angelo et al.
Table 7.3 Most commonly used docking programs and their scoring functions
Program
Schrödinger
Suite (Glide)
[57]
Scoring
Functions Applications Ligand Flexibility
Empirical
functions, datadriven scoring
Drug discovery,
protein–ligand
interaction studies
Exhaustive ligand
conformation
search
Distinctive
Features
High accuracy and
speed, integration
of multiple scoring
methods
AutoDock
[58]
Energy-based
scoring functions, empirical, grid-based
docking
Drug discovery,
protein–ligand
interaction studies
Genetic Algorithm, Simulated
Annealing, Local
Search, Lamarckian Genetic
Open source
Algorithm
Autodock
Vina* [59]
Energy-based
scoring
functions
Drug discovery,
protein–ligand
interaction studies
Genetic Algorithm, Simulated
Annealing, Local
Search, Lamarck-
Open source, userfriendly graphical
interface, GPU
compatibility
ian Genetic
Algorithm
GOLD
(Genetic
Optimization
for Ligand
Docking)
[60]
Energy-based
scoring functions, solvent
potential
Drug discovery,
protein–ligand
interaction studies
Genetic algorithm Flexibility in
defining force
fields, ligand
encapsulation
method, advanced
parameterization
options
FlexX [61] Energy-based
scoring functions, stereochemical
Drug discovery,
protein–ligand
interaction studies
Incremental construction
algorithm
Use of atomistic
potentials learned
by Ma chine
Learning
penalization
flex-Dock
Sur
[62]
Energy-based
scoring functions, ligand
selection
Drug discovery,
protein–ligand
interaction studies
Incremental construction
algorithm
High-precision
ligand selection
methods, combi-
nation of diverse
methods
Dock [63] Force field-
based and
empirical scoring functions
Drug discovery,
protein–ligand
interaction studies
Incremental construction
algorithm
Scores based on
force field and
empirical methods,
flexibility in
developing
docking protocols
Dock4 [64] Physical-,
empirical-, and
knowledgebased scores
Drug discovery,
protein–ligand
interaction studies
Incremental construction
algorithm
Combine grid-
based docking
approaches with
molecular optimi-
zation and custom
docking methods.
Dock6 [65] Interaction
energy, solvation, and
Drug discovery,
protein–ligand
interaction studies
Incremental construction
algorithm
Efficient search
algorithms,
improved scoring
(continued)

7 Molecular Docking: State-of-the-Art Scoring Functions and Search Algorithms 175
Table 7.3 (continued)
Program
LibDock [66] Data-driven
Molecular
Operating
Environment
(MOE) [67]
PLANTS
[68]
*
Smina: This is an enhanced AutoDock Vina version designed for improved performance and
flexibility. Smina retains Vina’s core functionality while introducing optimizations for multi-
threading, energy minimization, and scoring function customization [69]
*
Gnina: Gnina is an extension of Smina that integrates deep learning (in particular, convolutional
neural networks—CNNs) in molecular docking. By leveraging CNNs, Gnina aims to enhance
ligand pose prediction and scoring accuracy and efficiency. It offers advanced features such as atom
typing, symmetry detection, and ensemble docking [70]
*
These implementations represent advancements in the Vina docking framework, incorporating
state-of-the-art methodologies to address the challenges of molecular docking more effectively
Scoring
Functions Applications Ligand Flexibility
penalties for
sterically
unfavorable
scoring functions, atom pair
clustering
Energy-based
scoring
functions
Interaction
energy, solvation, and penalties for
sterically
unfavorable
(particularly useful
for large compound libraries)
Prediction of highaffinity ligands,
protein– ligand
interaction studies
Drug discovery,
protein–ligand
interaction studies
Drug discovery,
protein–ligand
interaction studies
Incremental construction
algorithm
Search algorithms, Genetic
algorithms,
Verlet’s algorithm, Quantum
chemistry
algorithms
Incremental construction
algorithm
Distinctive
Features
methods, and
improvements in
the handling of
flexible ligands. It
is also highly
modular and
extensible.
Use of simplified
force fields, focus
on predicting
pharmacological
ligands
Intuitive graphical
interface, integrated database,
algorithm
integration
Efficient search
algorithms,
improved scoring
methods, and
improvements in
the handling of
flexible ligands
• Desolvation terms: considers how ligand binding affects water molecules
around it.
• Global scoring function: considers an overall score for each ligand conformation
in the protein. This score is an estimate of the total binding energy in that
conformation.
• Entropic terms: scoring functions may also consider entropic terms reflecting
changes in entropy during the binding process. This is important because ligand
binding to the protein can affect the flexibility of the involved molecules.

176 R. M. de Angelo et al.
SurflexDock
AutoDock
Fig. 7.4 The most used molecular docking programs in the last years (2012–2022). The search was
carried out on Web of Science, with the keywords: “software name” and docking (N = 8496; the
search was based on the following topics in the Web of Science = title and abstract)
Dock6
Frigate
FlexX
LibDock
FRED
GOLD
MOE
Glide
• Optimization: to find the best binding conformation, the docking program scans
for different positions and orientations of the ligand in the protein. At each scan
step, the scoring function is calculated to assess the conformation’s quality in
terms of binding energy.
• Ranking: conformations can be ranked based on the calculated scores. The
conformation with the lowest score (most favorable energy) is usually considered
the most likely. It is expected to represent the most realistic interaction of the
ligand complexed with the protein.
• Results: at the end of the docking simulation, the program lists the most prom-
ising ligand conformations in the target protein, ranked by their scores obtained
via the scoring function. This helps researchers identify the most likely and
favorable conformations for interaction.
Each docking program may offer various options for different scoring functions,
with varied weights in terms of energy and entropy [72]. The choice of the scoring
function can significantly affect docking results, and developers often adjust and
validate these functions to optimize program performance on known and experimental datasets. Figure 7.4 presents the main docking programs employed in the last
years (2012–2022).
The four most widely used docking programs in the last years primarily focus on
drug discovery and studying interactions between proteins and ligands. The
AutoDock program is renowned for its versatility and effectiveness in predicting
interactions between ligands (molecules) and target proteins. Moreover, it is one of
the fastest and most widely used open-source molecular docking programs
[73, 74]. The Glide program (by Schrödinger) is known for its computer efficiency

7 Molecular Docking: State-of-the-Art Scoring Functions and Search Algorithms 177
and accuracy in docking simulations, making it a popular choice in the pharmaceutical industry [75]. The GOLD program stands out for its genetic optimization
algorithm, allowing efficient exploration of the conformational space of molecules
during docking. It is widely used in drug candidate discovery and identifying ligand
poses (conformations) that fit well into target proteins [76]. Finally, the MOE
program is a software suite that includes docking and molecular modeling features
and tools for visua lization and analysis of chemical structures. In addition to docking
simulations, MOE is used in molecular design studies, molecular dynamics simulations, and structure analysis [77].
It is worth noting that each docking program has its characteristics and advantages, and the choice between them often depends on the project’s specific needs,
resource availability, and the user’s familiarity with the tool. Accuracy, calculation
speed, accessibility, and compatibility with the system and target proteins should be
considered when selecting a docking program for a specific task.
3 Calculations Performed During Docking Simulations
Molecular docking simulations involve complex calculations to predict how a
molecule (ligand) fits into a target protein. The essential calculations performed
during the docking process include estimating the several energy terms related to the
binding between the ligand and receptor [78]. This energy is the sum of contributions
from various interactions, as presented in Table 7.4.
Molecular docking simulations entail a series of intricate calculations to predict
how a molecule (ligand) interfaces with a target protein. Critical calculations
performed during docking involve estimating several energy terms pertinent to the
binding between the ligand and receptor [78]. This energy is the cumulative result of
various interactions, as outlined in Table 7.4. In addition to considerations regarding
binding energy, another critical aspect to address is the flexibility of structures
involved in the simulation. Many proteins and ligands exhibit high flexibility,
implying significant structural changes during interaction. Consequently, in flexible
docking, computations must account for determining the most stable conformations
of the protein and ligand throughout the binding process while accommodating the
flexibility of specific residues and the ligand. This can render the docking process
impractical due to hardware limitations in scanning the conformational space of the
entire system.
It is worth noting that the calculations mentioned above are carried out iteratively
to ascert ain the most probable binding conformation. The efficacy of docking
simulations hinges on the accuracy of molecular models utilized, the quality of
energy parameters, and efficient search strategies. These simulations are indispensable tools in drug candidate design and discovery, facilitating an understanding of
crucial interactions between ligands and biological targets.

178 R. M. de Angelo et al.
Table 7.4 Energy terms related to the binding between the ligand and receptor
van der Waals Dispersion Energy: represents the attraction between atoms due to fluctuations
Hydrogen
Bonding
Electrostatic Electrostatic Potential: describes the resulting electric field from charges on
Solvation Free Solvation Energy: measures the contribution of solvation to ligand–
Desolvation In this step, consideration is given to how the protein and ligand affect the
Entropic Terms Conformational Entropy: measures the freedom of movement of the ligand and
in electron clouds.
van der Waals Radius: defines the contact distance between atoms.
Lennard-Jones Potential: describes the attractive and repulsive interaction
between atoms.
Steric Repulsion Energy: represents repulsion when atoms come too close.
These properties are crucial for assessing the complementarity between ligands
and proteins in docking simulations.
Hydrogen Bonding Energy: measures the strength of the interaction between
hydrogen atoms and hydrogen bond donors or acceptors.
Hydrogen Bonding Distance: defines the ideal distance between donor and
acceptor hydrogen atoms.
Hydrogen Bonding Angle: evaluates the orientation of hydrogen atoms in a
hydrogen bond.
Free Energy of Hydrogen Bonding: represents the total contribution of hydrogen bonds to the stability of the ligand-protein interaction.
atoms and influences interaction direction in docking.
Coulombic Interaction Energy: measures the contribution of Coulomb forces in
ligand–protein interactions.
Potential Surface Maps: visually represent the distribution of electrostatic
properties around atoms. These properties are essential for understanding
electrostatic interactions in ligand–protein docking.
protein affinity.
Free Desolvation Energy: evaluates the energetic cost of removing solvation
around the ligand and protein during docking.
Ligand Solvation Free Energy: calculates the specific contribution of the ligand
to solvation energy.
Protein Solvation Free Energy: represents the contribution of the protein to the
total solvation energy. These properties aid in understanding interactions with
the aqueous environment during docking.
structure of water molecules around them. This is important because ligand
binding to the protein often involves the displacement of water molecules from
the interface.
protein, affecting system entropy.
Translational Entropy: refers to the variation in entropy due to changes in the
position of molecules in space.
Rotational Entropy: represents the entropy resulting from molecular rotation.
Vibrational Entropy: evaluates the variation in entropy due to molecular
vibrations. These entropic terms influence the stability and affinity of ligand–
protein interactions during docking simulations.
It is important to note that the list of interaction/energy terms provided is not
comprehensive, and docking scores frequently miss capturing all interactions. An
illustrative instance of extending a scoring function to encompass significant missing
interactions and generate more precise docking results can be found here: [79].

7 Molecular Docking: State-of-the-Art Scoring Functions and Search Algorithms 179
4 Essential Components for a Good Docking Program
“A docking protocol requires extensive sampling of the ligand’s conformational
space to position it at the binding site of a protein, generating a large number of
potential orientations and conformations of a ligand, known as poses. A good
positioning algorithm samples ‘all’ possible binding modes, while the scoring
system ranks all solutions and identifies the most probable ‘binding mode’ of the
ligand.” [80].
A well-designed molecular docking program should incorporate essential components to provide accurate and valuable results. The following are the critical
components that contribute to an effective docking progra m [81, 82]:
• Accurate molecular modeling: The program must be able to accurately represent
the three-dimensional structures of the target protein and the ligand. This includes
considering the flexibility of the protein and the ligand.
• Precise energy parameters: The accuracy of binding energy calculations is
crucial. Energy parameters should be well-calibrated and reflect the key molec-
ular interactions of real/experimental systems, such as hydrogen bonds, van der
Waals interactions, and electrostatics.
• Efficient search algorithms: The program should use efficient search algorithms
to explore the conformational space for the most stable conformations of the
protein–ligand complex. Algorithms like Monte Carlo, genetic algorithms, or
optimization methods are most common in docking programs.
• Accurate scoring function: An appropriate scoring function is crucial for evalu-
ating and ranking ligand conformations in the protein. The scoring function
should accurately reflect the binding energy and consider entropic factors.
• Desolvation: Considering desolvation is important to assess how ligand binding
affects water molecules around. This step is critical for accurate models.
• Molecular flexibility: In many cases, target protein and ligand structures exhibit
high degrees of freedom. Thus, the program should handle the flexibility of the
systems, allowing optimization of molecular conformations during the docking
process.
• User-friendly interface: An intuitive user interface is desirable for researchers to
set parameters, visualize results, and interpret data easily.
• Adequate computational resources: Docking calculations can be computer inten-
sive, especially in high-precision and complex simulations. The program should
be able to use suitable hardware or cluster computing systems.

180 R. M. de Angelo et al.
• Validation and benchmarking: A reliable docking program shoul d be validated
and benchmarked on known/experimental datasets to ensure consistent and
accurate results.
• Customization and extensibility: The program should allow customization and
application extension to meet the specific needs of each research project.
• Integration with other software: The ability to export and import data from other
related software, such as molecular viewers and analysis tools, is desirable.
A well-designed molecular docking program should consider all the mentioned
components to provide reliable and valuable results related to drug candidate
discovery and design processes. It is important to note that the choice of the program
should be based on the specific needs of the project and the research context.
5 Limitations of the Docking Technique
The molecular docking technique has significant limitations that need to be acknowledged by the user. They start with the fact that docking may not always accurately
predict a compound’s binding conformation in the active site of a protein. This
happens because the search and scoring functions used may not capture all the
complexities related to the molecular docking [83].
Another point to consider is related to the fact that proteins and ligands have
flexible structures, but most docking programs treat molecules as rigid during
calculations. This can lead to inaccurate results, in particular for the systems with
a high number of degrees of freedom. For this kind of systems, an option would be to
use flexible and/or ensemble docking. Additionally, most scoring functions do not
handle entropic terms well, meaning they do not adequately consider changes in
entropy during the binding process. Ignoring these changes can lead to inaccurate
predictions regarding binding affinity.
Regarding the proper treatment of solvation (how water molecules are around the
protein and ligand), it can be said that this is still a challenge and can significantly
affect docking results. Besides, for particular systems, the addition of structural
water is very important to properly describe the action mechanisms of the protein.
In other cases, for example, a system with a low structural information/quality, the
addition of solvation is not the most important feature to be considered in the initial
analysis. The choice of initial conformations can also be a problem, as the quality of
docking results largely depends on the choice of the initial conformations of the
ligand and protein. If these conformations are not representative, the results may be
distorted [84]. Furthermore, the technique relies on high-quality three-dimensional
structures of the protein and ligand. When these structures are unavailable or of
lower quality, docking simulations may be less reliable.

7 Molecular Docking: State-of-the-Art Scoring Functions and Search Algorithms 181
It is important to note that the results of docking simulations should be experimentally validated, which can be a costly and time-consuming step. Experimental
validation is crucial to determine the accuracy of predictions. In summary, molecular
docking is a valuable tool, but input data must be reliable, and results should be
interpreted cautiously, considering its intrinsic limitations. It is vital to use docking
in conjunction with other experimental and computational techniques for a more
comprehensive assessment of protein–ligand interactions and to increase the reliability of predictions.
6 Validation of Docking Results
The validation of docking results is a critical step in determining the reliability of
predictions generated by these simulations. Various approaches and validation
methods can be employed to assess the performance of docking techniques, for
example [85, 86]:
• Cross-docking validation refers to the simulation (docking) of the ligand of a
particular biological target to the other tridimensional structures of the same
receptor. This practice provides the accuracy of the docking algorithm to predict
the best ligand pose at the binding site of the biological target.
• Benchmarking with experimental data: comparing docking results with experi-
mental data is a crucial form of validation. This may involve comparing predicted
binding affinity with experimentally measured affinity, overlaying predicted
binding conformations with experimental structures, or analyzing predicted inter-
actions against experimentally observed interactions [87, 88].
– The validation and benchmarking of docking programs are crucial to ensure
their reliability, consistency, and accuracy in predicting ligand–protein interactions. Several methods can be employed to achieve this point, including:
– Experimental data comparison: docking results can be compared with exper-
imental data, such as crystallographic structures of ligand–protein complexes.
If the predicted binding poses and binding affinities align well with experimental observations, it indicates the reliability of the docking algorithm. For
example, the Directory of Useful Decoys (DUD) dataset provides a benchmark
for assessing docking performance against experimentally validated ligand–
protein complexes [89].
– Cross-docking studies: in cross-docking studies, ligands are docked into
multiple protein structures to evaluate the program’s ability to discriminate
between binding sites and protein conformations. This assesses the docking
program’s robustness and versatility. The Astex Diverse Set (ADS) is a
commonly used dataset for cross-docking studies due to its diverse protein–
ligand complexes [90].

182 R. M. de Angelo et al.
– Scoring function evaluation: The performance of scoring functions within
docking programs can be assessed by comparing predicted binding affinities
with experimental binding constants or free energies. Statistical metrics such
as root mean square error (RMSE) or Pearson correlation coefficient can
quantify the agreement between predicted and experimental values [91 ].
– Virtual screening campaigns: virtual screening campaigns involve the docking
of large compound libraries against a target protein to identify potential lead
compounds. The enrichment factor, receiver operating characteristic (ROC)
curve, and area under the curve (AUC) are metrics used to evaluate virtual
screening performance. The DUD-E dataset [40] is often used for virtual
screening benchmarking [92].
– Community challenges and collaborative efforts: participation in community
challenges, such as the D3R Grand Challenge or the Critical Assessment of
PRediction of Interactions (CAPRI), allows docking programs to be evaluated
in blind tests against undisclosed protein–ligand complexes. These challenges
provide an independent assessment of program performance in a competitive
setting [93].
– By employing these approaches of validation and benchmarking, docking
programs can demonstrate their reliability and accuracy in predicting ligand–
protein interactions across diverse datasets and scenarios. Additionally, publications in peer-reviewed journals often report detailed validation studies and
benchmarking results for docking programs, providing valuable insights into
their performance characteristics [92].
• Retrospective validation: using historical datasets where binding activity is
known and comparing docking predictions with these known results is the
focus of this approach. This is useful for evaluating the accuracy and performance
of the method under controlled conditions.
• Negative control: this step involves performing docking simulations where the
ligand has no affinity for the target protein. This helps identify false positives, i.e.,
ligands that are erroneously predicted to have affinity for the target protein.
• Statistical analysis: conducting statistical analyses, for example, via Receiver
Operating Characteristic (ROC) and Precision-Recall (PR) curves, to evaluate the
performance of docking simulations in classifying active and inactive ligands.
• Reproducibility tests: it is important to repeat docking simulations with different
configurations or parameters to verify the consistency of results. This helps
determine if the results are robust.
• Validation with analog ligands: in this step, we perform docking simulations with
a congener series of analog ligands and compare predictions based on their
structural similarity. This can be used to assess the method ’ s ability to recognize
structural similarities.
Соседние файлы в папке Библиотека им академика М.И. Перельмана
