Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5884_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
02.09.2026
Размер:
21 Мб
Скачать
7 Molecular Docking: State-of-the-Art Scoring Functions and Search Algorithms 183
Prospective validation: in this step, the goal is to apply the docking method to
systems not yet experimentally studied and compare predictions with subsequent
experimental results.
Blind docking: in this type of validation, the aim is to perform docking simula-
tions without prior knowledge of the location of the binding site on the target
protein. This tests the methods ability to predict the location of known binding
sites and estimate unknown sites.
Comparative validation: here, the focus is on comparing the performance of the
docking method with other protein– ligand afnity prediction methods, such as
force-eld-based docking, molecular dynamics, or approaches employing
machine learning techniques.
It is important to note that there is no single validation method universally applicable in docking studies, as the choice of validation method depends on the context, research objectives, and availa ble data. In many cases, a combination of validation methods is used to obtain a comprehensive assessment of docking technique performance. Additionally, collaboration among researchers and the availability of well-dened experimental datasets play a fundamental role in the successful validation of docking predictions [94].
6.1 Success in Docking Predictions and the Consensus
Technique
Using a consensus employing multiple dockin g algorithms is a common and valu­able strategy in research involving drug discovery and studies of protein–ligand interactions. This approach involves conducting multiple independent docking sim­ulations for a protein–ligand complex and then combining the results to obtain a more reliable and accurate prediction of the binding conformation and afnity of the compound at the receptor site.
Docking simulations can be sensitive to various variables and parameters, and this diversity may inuence individual results. By achieving a consensus of results obtained from different docking algorithms, it is possible to reduce the impact of outliers or inaccuracies, obtaining a more accurate estimate. The consensus takes into account the inherent uncertainty in docking simulations. Di fferent simulations may produce slightly different results, but the consensus provides a robust measure reecting the intrinsic variability of docking algorithms [86, 95].
The consensus process can help identify binding conformations consistent across multiple simulations, increasing condence that these conformations are more likely to occur in reality. By verifying if various simulations produce consistent results, the consensus helps assess the reliability of docking predictions. Conformations or interactions that are consistent in multiple simulations are more likely to be reliable.
184 R. M. de Angelo et al.
The consensus allows researchers to analyze protein–ligand interactions better by considering a variety of conformations, and possible Consensus helps reduce false positives (ligands that appear to bind but are not relevant) and false negatives (ligands that are mistakenly discarded) in dockin g predic tions [86].
In drug discovery research, decisions regarding selecting potential bioactive compounds and designing new molecules can be guided by consensus-based docking results, providing a solid foundation for decision-making with lower error chances. It is important to note that consensus from results via different docking algorithms is not a denitive solution to all chall enges associated with molecular docking, and its application should be carefully planned and validated. Additionally, how to obtain the consensus and the number of simulations to be performed may vary depending on the system of interest and research objectives. However, in many cases, a consensus from various docking programs is a practical approach to improve the reliability of predictions and the quality of results.

7 Inappropriate Use of Validation Methods in Docking

The inappropriate use of validation methods is a signicant concern when applying molecular docking techniques. Validation is essential to determine the reliability of predictions generated in docking simulations. Improper use can lead to erroneous conclusions and excessive condence in results that are not genuinely predictive [95]. Several factors may contribute to less accurate conclusions about docking results, such as:
Lack of experimental data: Initial validation of docking algorithms relies on the
availability of relevant experimental data. Validation becomes compromised if
high-quality experimental data are unavailable for comparison with docking
results.
Validation on biased datasets: Researchers sometimes validate their results on
datasets previously used to develop or tune scoring functions. This can
overestimate the methods accuracy, as training data may be biasedto t the
algorithms used.
Lack of diversity in molecular targets: Some docking studies may focus only on a
specic type of target protein, limiting the generalization of results. Proper
validation should include a variety of molecular targets to determine applicability
in different contexts.
Failure to consider system exibility: If the docking algorithm does not account
for the exibility of the protein and/or ligand, validation may be inaccurate,
especially in highly exible systems.
7 Molecular Docking: State-of-the-Art Scoring Functions and Search Algorithms 185
Neglecting entropic terms: The exclusion of entropic terms in estimating binding
energy (terms describing changes in entropy during the binding process) can lead
to inaccurate validation, as these changes are critical for estimating protein–
ligand afnity [2].
Validation limited to single metrics: Validation often focuses on a single metric,
such as binding afnity. However, other metrics, such as ligand specicity and
correct binding conformation, are also essential and should be considered.
Ignoring solvent inuence: Validation may be inadequate if the solvent is not
adequately treated in docking simulations, as solvation is a critical factor in
protein–ligand interactions.
Failure to conduct repeatability studies: Docking results may vary depending on
execution conditions and parameters. Repeating studies is vital to assess result
robustness.
A comprehensive approach should be adopted to overcome the previously cited issues and appropriately use validation methods in the molecular docking process. This includes selecting independent validation datasets, considering different per­formance metrics, evaluating exibility and solvation, and repeating experiments to ensure result consistency. Collaboration within the scientic community is crucial for sharing data and knowledge and promoting rigorous validation practices [85, 86,
95].
8 Evaluation from an Experts Perspective
Expert evaluation is essential when analyzing results obtained through molecular docking simulations, and this is due to a series of fundamental reasons. The results of these simulations can be incredibly complex, involving various metrics, binding energy data, conformations, and structural information. In this context, the expertise of a specialist is necessary, enabling the precise interpretation of these complex data. Experts can assess these results in light of the biological and chemical context, considering the target proteins function and the specic application. They can discern whether the identied conformations are relevant in the biological and chemical context [81].
Experimental validation is critical to conrming predictions made through docking. In this regard, a specialist can conceive and conduct experiments that conrm or refute predictions, becoming a crucial component in the validation process. Assessing the accuracy of docking simulations is a task that requires technical and scientic knowledge. A specialist can identify potential sources of error, such as the quality of protein and ligand structures, the effectiveness of scoring functions, and the appropriateness of simulated conditions.
186 R. M. de Angelo et al.
In areas such as drug design and discovery, the analysis of docking results often directly affects strategic decisions, such as the selection of drug candidates and the design of potential bioactive compounds. In this context, the guidance of a specialist is of utmost importance. As mentioned earlier, docking simulations have their limitations and uncertainties. An expert can identify and communicate these limita­tions clearly, providing a balanced and realistic view of the conclusions. As proteins and ligands can exhibit exibility, specialized analys is is necessary to understand how this exibility affects docking predictions and consider how dynamic confor­mations may inuence binding. Experts can assess whether docking results align with current knowledge about the target proteins biology, its function in the organism, and its relationship with the overall biological context [82].
In this way, based on their knowledge, an expert can draw solid conclusions and provide recommendations on the next steps of research, such as additional validation experiments, ligand optimization, or structural modications. Expert evaluation is an essential step in trans lating raw data into scientic knowledge, experimental valida­tion, and guiding strategic decisions in research related to drug discovery, molecular biology, and chemistry.

9 Use of Machi ne Learning in Molecular Docking

Machine learning and molecular docking have an increasingly close relationship in the eld of drug discovery and studies of protein–ligand interactions. The applica­tion of machine learning techniques throu ghout molecular docking simulations has the potential to enhance the accuracy and efciency of predictions. Below are listed some ways in which these two areas relate [9698]:
Improvement of scoring functions: Scoring functions used in docking simulations
are crucial for estimating the binding afnity between protein and ligand.
Machine learning techniques can be used to develop more accurate and specic
scoring functions, incorporating a broader range of molecular features and
interactions [30].
Drug candidate selection: Machine learning algorithms can analyze large librar-
ies of chemical compounds and predict which ones are more likely to be prom-
ising ligands for a specic target protein. This can save time and resources in
virtual screening. For example, a study [74] on COVID-19 utilized about 2.178
million unique protein/ligand binding site combinations available in databases to
optimize the search for potent inhibitors, applying ML techniques to facilitate and
guide in silico studies.
Classication of active and inactive ligands: Machine learning techniques can be
applied to classify ligands as active or inactive based on their chemical and
structural properties. This is useful for the virtual screening of compound
7 Molecular Docking: State-of-the-Art Scoring Functions and Search Algorithms 187
libraries. In a study by Salimi et al. [99], a screening pipeline was built using
machine learning algorithms integrated with checks for similarity to approved
drugs to nd new inhibitors for the vascular endothelial growth factor receptor-2
signaling pathway.
Prediction of proteinligand interactions: Machine learning algorithms can pre-
dict specic interactions between proteins and ligands, identifying critical amino
acid residues for binding or preferred binding sites [97].
Modeling molecular exibi lity: Molecular exibility is a signicant challenge in
docking simulations. Machine learning methods can be applied to model the
exibility of proteins and ligands, allowing more realistic predictions. Harmalkar
and Gray [100] studied the difculties in predicting protein interactions and ways
to apply machine learning to optimize these simulations. The study showed that
docking simulations in the CAPRI [101] challenge included a wide variety of
target types, with 11 out of 28 easytargets achieving high-quality structure.
However, for the 17 difculttargets, the intrinsic exibility of biomolecules
remains a challenge, with only 2 achieving high quality.
Enhancement of validation: Machine learning techniques can be employed to
enhance the validation of docking results, identifying more reliable and effective
metrics to assess the methods performance. In this case, a study by Zhang et al.
[102] presented an internal and external dataset used to cross-validate eight
machine learning methods. The results showed that the extremely random tree
model performed better and was adopted as the rst step in virtual screening.
Increase in computational efciency: Machine learning methods can be used to
accelerate the processing of large volumes of data in docking simulations, making
the process more efcient [98].
Discovery of structureactivity relationships: Machine learning can help discover
complex relationships between the chemical structure of ligands and their bio-
logical activity, aiding in compound optimization. For example, in Hermansyahs
study [103], selective DPP-4 inhibitors against DPP-8 and DPP-9 were identied
using an AI-based quantitative structure–activity relationship (QSAR) workow,
enabling faster screening of millions of molecules for the DPP-4 target compared
to other screening methods.
Integration of diverse data: Machine learning techniques enable the integration of
data from various sources, such as high-throughput screening data, structural
biology information, and gene expression data, for a more comprehensive system
analysis and optimization of simulations [ 104].
To assess the evolution in the use of machi ne learning in docking simulations, publications indexed in the Web of Sciencedatabase between 2012 and 2022 were compiled, relating to the use of molecular docking with machine learning (Fig. 7.5).
188 R. M. de Angelo et al.
Fig. 7.5 Number of studies (2012–2022) considering the use of machine learning techniques and docking simulations for the design and discovery of drug candidates. Research conducted on the
Web of Scienceplatform on October 15, 2023, with the combination of the following keywords:machine learningAND molecular docking
A signicant increase, approximately 14 times, in the number of publications using ML techniques to enhance docking studies can be observed.
In summary, integrating machine learning techniques in molecular docking studies can signicantly enhance the accuracy and efciency of predictions, accel­erating the discovery of drug candidates and enabling a deeper understanding of protein–ligand interactions. Collaboration among computer scientists, structural biologists, and chemists is crucial for the success of this approach.
10 Advancements and Improvements in Computational
Resources
Recent advancements in computational resources have led to signicant progress in molecular dockin g, playing a crucial role in drug discovery. A rising technique is structure-based virtual screening, which bene ts from the increasing availability of high-resolution target structures and large-scale virtual compound libraries, such as REAL combinatorial libraries. A REAL combinatorial library collects virtually generated chemical compounds representing a wide structural diversity. The term REALstems from Readily Available for Synthesis,indicating their readiness for laboratory synthesis. These libraries are designed to offer a vast array of molecules that can be synthesized efciently and economically in the laboratory [105].
However, new approaches are necessary to keep pace with these libraries exponential growth. In this regard, based on a modular approach, the V-SYNTHES method emerges as a promising solution. This method conducts
7 Molecular Docking: State-of-the-Art Scoring Functions and Search Algorithms 189
hierarchical screening of a REAL library containing over 11 billion compounds. Initially, V-SYNTHES identies the most promising scaffold–synthon combina­tions as seeds for iterative growth, rening them to select complete molecules with the best docking scores. This approach enables efcient detection of high-scoring compounds in an extensive chemical space while docking only a min imal fraction of the library. This is highly advantageous in terms of computational and economic efciency. Furthermore, the method has been experimentally validated in synthesiz­ing and testing cannabinoid antagonists, demonstrating signicant improvement over conventional virtual screenings [105].
Another notable advancement lies in incorporating molecular strain as an addi­tional parameter in evaluating ligand scores during the docking process. A recent approach utilizes information on relative torsional populations from the Cambridge Structural Database to precalculate these strain energies, resulting in a more accurate and efcient evaluation. Retrospective studies have shown that including these strain energies signicantl y improves success rates by preferentially excluding false high­scoring molecules. This approach, independent of the scoring function used, stands out for its speed and practicality, and it is applicable even in large-scale compound libraries [106].
Moreover, careful selection of small molecule libraries is crucial to this process. Libraries capable of synthesizing billions of compounds, such as the Examine library, are available. Finally, despite the advancements represented by AlphaFold2 in predicting protein structures with high precision, studies have shown that the quality of predicted structures does not always translate into satisfactory perfor­mance during the docking process. Removing low-condence regions from the predicted structure and exibilizing side chains are promising strategies to optimize docking results, underscoring the importance of rened adjustments for successfully applying these models in drug discovery. It is also important to mention that recent approaches have been developed to consider the refolding around the ligand. New models like AlphaFold3 and Neuraplex have showed excellent performance in this area [ 107].

11 Challenges

Improving molecular docking techniques faces various technical and scientic challenges as the pursuit of greater accuracy and applicability is ongoing. Some of the main challenges include enhancing the accuracy of scoring functions and adequately considering the exibility of the systems under consideration. Cur rent functions may not accurately capture all ligand–receptor interactions, especially in highly exible systems or those with weak interactions. Integrating the exibility of both the protein and the ligand in simulations is a complex task. Developing effective methods to handle exibility, including molecular dynamics coupled with docking, is a signicant challenge [2]. Considering entropic terms in scoring functions is another complicating factor but essential for accurately predicting
190 R. M. de Angelo et al.
protein–ligand afnity. Correctly modeling changes in entropy during the binding process is equally challenging. The inuence of solvent on protein–ligand interac­tions must be accurately addressed as well. Developing more precise solvation methods is a considerable challenge. Additionally, experimental validation is crucial for rening docking techniques. However, it is not always easy to conduct due to the complexity of protein–ligand interactions and the availability of high-quality exper­imental data [4].
Extending the applicability of docking to more diverse targets, such as membrane proteins and protein–protein complexes, is another challenge, as these systems can be highly complex and require specic approaches. Integrating the docking process with other computational techniques, such as molecular dynamics, machine learn­ing, and quantum simulations, is an evolving approach to improving the accuracy and comprehensiveness of predictions. It is worth noting that docking simulations on multiple targets are complex and challenging, as they require considering the competition between different ligands at a single or multiple binding sites [71 ].
Other factors that make docking simulations somewhat challenging involve the following factors: validation and prediction of weak molecular interactions, such as those involved in allosteric inhibitors or protein modulators; docking simulations can be computational ly intensive, requiring signicant resources; the need for the development of efcient algorithms to reduce computation time; better understand­ing of the complexities of biology involved in protein–ligand interactions may be essential to enhance docking techniques and make them more realistic; availability of high-quality protein and ligand structures is essential for the quality of docking results; use of high-quality reference experimental datasets to validate and improve docking techniques.
Overcoming the challenges mentioned above and others not described here regarding docking simulations requires multidisciplinary collaboration among com­puter scientists, structural biologists, chemists, and pharmacologists. Continuous research and development of more advanced methods will improve molecular docking techniques and their application in various areas, including drug candidate discovery.

12 Conclusions

Molecular docking techniques play a crucial role in drug discovery and candidate drug design worldwide and have undergone signicant technological advances. Table 7.5 presents the key reasons why molecular docking algorithms are essential in the drug discovery and design process and how various improvements have contributed to the advancement and qua lity of docking analyses.
Molecular docking techniques are vital in global drug discovery and design. They help expedite the entire process, save resources, and enhance accuracy in identifying drug candidates with lower error rates. Ongoing technological advancements in molecular docking furt her broaden its impact and potential in pharmaceutical and biological researches.
7 Molecular Docking: State-of-the-Art Scoring Functions and Search Algorithms 191
Table 7.5 Importance and challenges/advancements in docking algorithms
Fact Importance
Accelerates drug candidate discovery
Resource savings The use of docking techniques helps reduce the number of com-
Expands research scope Molecular docking algorithms make it possible to investigate a
Assists in selectivity studies Docking techniques aid in designing compounds that can selec-
Understanding molecular mechanisms
Personalization of therapies Docking can also be used to develop drugs for personalized ther-
Increased accuracy of scor­ing functions
Inclusion of exibility Improved docking software allows for considering certain exi-
Integration with Machine Learning
Use of supercomputers Supercomputers and high-performance clusters allow the execu-
Quality data related to structural biology
More accurate experimental validation
The design of drug candidates is lengthy, costly, and risky. Molecular docking enables the virtual screening of compounds, expediting the identication of drug candidates with higher success rates. This saves time and resources
pounds to be synthesized and experimentally tested, saving nan­cial and laboratory resources
wide variety of compounds, including those not easily accessible through experimental approaches
tively bind to specic targets, minimizing unwanted side effects
Analyzing protein–ligand interactions through docking contributes to understanding the molecular mechanisms underlying diseases, leading to the development of more effective treatments
apies tailored to individual patient needs
Challenge/Advancement
Scoring functions used in docking simulations have been enhanced, becoming more accurate and representative of ligand– receptor interactions
bility in the protein and the ligand, enhancing prediction accuracy
Integrating machine learning techniques and docking algorithms improves the prediction of protein–ligand interactions, making simulations more effective
tion of more complex docking simulations
The availability of high-quality structural biology data, such as protein structures determined by crystallography or nuclear mag­netic resonance, enhances the accuracy of docking simulations
Technological advances in experimental validation, such as high­resolution mass spectrometry and cryo-electron microscopy, assist in conrming predictions from docking simulations
192 R. M. de Angelo et al.
Questions to Answer When Planning a Docking Experiment
What is the main objective of the docking experiment?
What are the features and limitations of the biological receptors structure?
What are the main molecular properties of the ligands under study?
Have the three-dimensional structures been veried and optimized?
Are there structural or conformational errors in the selected molecules?
Which docking software will be used? Is it suitable for the system under study?
What is the best scoring function to use for the study in question?
Are there any licensing limitations for the tools being used?
What positive and negative controls will be used to validate the docking protocol?
How will the results be analyzed? What criteria will be used to evaluate the estimated binding energy and molecular interactions?
What recent advancements in scoring functions and docking exibility can be leveraged?
Is there a possibility to integrate machine learning techniques to improve the prediction of protein–ligand interactions?
Is high-quality structural biology data available to enhance the accuracy of docking simulations?
What experimental methods will be used to validate the results of the docking simulations?
What are the limitations of molecular docking simulations?

References

1. Fan, J., Fan, J., Fu, A., & Zhang, L. (2019). Progress in molecular docking. Quantitative Biology, 7,83–89.
2. Winkler, D. A. (2020). Ligand entropy is hard but should not be ignored. Journal of Chemical Information and Modeling, 60(10), 4421–4423.
3. Kuntz, I. D., Blaney, J. M., Oatley, S. J., Langridge, R., & Ferrin, T. E. (1982). A geometric approach to macromolecule-ligand interactions. Journal of Molecular Biology, 161(2), 269–288.
4. Stanzione, F., Giangreco, I., & Cole, J. C. (2021). Use of molecular docking computational tools in drug discovery. In Progress in medicinal Chemistry (Capítulo Quatro) (pp. 273–343). Elsevier.
5. Li, J., Fu, A., & Zhang, L. (2019). An overview of scoring functions used for protein– Ligand interactions in molecular docking. Interdisciplinary Sciences, Computational Life Sciences, 11, 320–328.
6. Zheng, L., Meng, J., Jiang, K., Lan, H., Wang, Z., Lin, M., Li, W., Guo, H., Wei, Y., & Um, Y. (2022). Improving protein–ligand docking and screening accuracies by incorporating a scoring function correction term. Briengs in Bioinformatics, 23(3).
7. Shen, C., Hu, Y., Wang, Z., Zhang, X., Pang, J., Wang, G., Zhong, H., Xu, L., Cao, D., & Hou, T. (2021). Beware of the generic machine learning-based scoring functions in structure-based virtual screening. Briengs in Bioinformatics, 22(3), bbaa070.
8. Nguyen, D. D., & Wei, G.-W. (2019). AG-score: Algebraic graph learning score for protein– ligand binding scoring, ranking, docking, and screening. Journal of Chemical Information and Modeling, 59(7), 3291–3304.
9. Britt, H. M., Cragnolini, T., & Thalassinos, K. (2022). Integration of mass spectrometry data for structural biology. Chemical Reviews, 122(8), 7952–7986.