Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5606_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
02.09.2026
Размер:
21 Мб
Скачать
Chapter 7
Molecular Docking: State-of-the-Art Scoring Functions and Search Algorithms
Rafaela M. de Angelo, Daniel S. de Sousa, Aldineia P. da Silva, Laise P. A. Chiari, Albérico B. F. da Silva, and Kathia M. Honorio
Abstract This study addresses the technique of molecular docking in drug design,
emphasizing the importance of advanced scoring functions and search algorithms. It explores the generation of detailed models involving protein targets and ligand molecules and the evaluation of multiple ligand conformations to identify favorable molecular interactions. Additionally, it underscores the prediction of the ligands most stable and energetically favorable conformation to interact with the target protein. Molecular docking plays a fundamental role in the early stages of drug discovery and development, assisting in selecting promising drug candidates and estimating the afnity between candidate molecules and target proteins. The key contributions of this chapter include the exploration of the importance of molecular docking in drug design and the analysis of advanced scoring functions and search algorithms. These topics are essential components for successful molecular docking. In addition, this chapter will discuss the relevance of result validation and the need for accuracy in the predictions based on docking simulations. Furthermore, it evaluates the challenges faced in molecular docking and reects on using machine learning techniques in this context. It provides insights into the interaction between molecules and target proteins, which is essential for understanding the mechanisms of action of bioactive substances. It also emphasizes the acceleration in drug discovery and design, resource management, and the enhancement of understanding of molecular interactions as signicant outcomes. These contributions underscore
R. M. de Angelo University Federal of ABC (UFABC), Center for Natural and Human Sciences (CCNH), Santo André, SP, Brazil
D. S. de Sousa · L. P. A. Chiari · A. B. F. da Silva University of São Paulo (USP), São Carlos Institute of Chemistry (IQSC), São Carlos, SP, Brazil
A. P. da Silva · K. M. Honorio ( University of Sao Paulo (USP), School of Arts, Sciences and Humanities (EACH), Sao Paulo, Sao Paulo, Brazil e-mail: KMHONORIO@USP.BR
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2024 V. G. Maltarollo (ed.), Computer-Aided and Machine Learning-Driven Drug Design, Computer-Aided Drug Discovery and Design 3,
https://doi.org/10.1007/978-3-031-76718-0_7
✉)
163
164 R. M. de Angelo et al.
the importance of molecular docking as a fundamental tool in pharmaceutical research and drug development.
Keywords Molecular docking · Scoring functions · Search algorithms · Drug design

1 Molecular Docking

Molecular docking is a sophisticated computational technique extensively utilized across various scientic domains, including chemistry, pharmacy, and molecular biology. Its primary objective is to simulate the interaction between a molecule and the three-dimensional structure of a specic target. The docking procedure goal is to generate de tailed models of the target macromolecule (e.g., protein) together with the small (ligand) molecule. It often starts from multiple ligand conformations are assessed to pinpoint the position and orientation that yield favorable molecular interactions, which are quantied using a scoring function (SF). These interactions encompass hydrogen bonds, van der Waals forces at the proteins active site, and other relevant interactions crucial for the binding. Regardless of the scoring function employed, these interactions play a signicant role in determining the docking outcomes [1].
In a nutshell, the primary objective of the docking is to predict a favorable target­ligand conformation, maximizing the number of interactions and minimizing the predicted binding energy (derived from the SF). This information is crucial in the search for drug candidates, as it helps select promising bioactive substances, opti­mize structures of existing molecules, and estimate the afnity between the candi­date molecule and the target protein. In summary, the docking technique is crucial in the early stages of drug discovery and development [2].
Kuntz and colleagues pioneeringly introduced the docking methodology in 1982, with the article titled A Geometric Approach to Macromolecule-Ligand Interac­tions.[3] In this work, the authors proposed that molecular recognition between molecules, based on chemical and geometric aspects, could be explored through three-dimensional models of both the ligand and the target. Since then, molecular docking has been widely applied as a rapid approach to estimate how a particular compound binds to a biological target and to predict/estimate the afnity of this interaction.
In this chapter, we will discuss the relevance of docking in drug design and discovery, advances in scoring functions and search algorithms (including machine learning, ML), calculations involved in docking simulations, essential components of effective docking programs, limitations of the technique, and validation of docking results (including the consensus technique). We end the chapte r discussing perspectives from different experts, the ML integration in molecular docking, and giving a perspective on the most recent challenges encountered in the eld.
7 Molecular Docking: State-of-the-Art Scoring Functions and Search Algorithms 165
1.1 Importance/Relevance of Docking in the Drug Design
and Discovery
Traditionally, the search for new drugs starting from the initial hits involves the synthesis and testing of many compounds; this process is error prone and resource intensive. Docking can help to deprioritize compounds with a low likelihood of success, saving valuable time and resources. Docking results can also be used to optimize the structure of existing hit molecules, during the hit-to-lead phase, improving their effectiveness and selectivity to the target protein [2].
Molecular docking provides detailed insights into the protein–ligand interactions at the atomic level. The early identication of hits as drug candidates with a high probability of success can save signicant resources in the long term by avoiding investments in compounds that will not work or h ave severe side effects. It is worth mentioning that the docking technique is also relevant in the era of personalized medicine, where drugs are tailored to individual patient characteristics, helping design medications that precisely t variant or mutant proteins [4].
When one employs molecular docking and faces some of its complexities, it becomes crucial to recognize the pivotal role played by the choice of scoring functions and search algorithms. These elements constitute the essential foundations of docking tools, guiding the investigation of molecular interactions and determining optimal binding conformations. Before delving into recent advancements in scoring functions and search algorithms, it is relevant to briey contextualize these compo­nentsimportance within the molecular docking realm.

2 Advances in Scoring Functions and Search Algorithms

Molecular docking is an ever-evolving eld and state-of-the-art scoring function and search algorithms continue to advance rapidly. It is crucial to highlight signicant trends and recent advancements within these domains, delineated through the presentation of three principal categories of scoring functions: force-eld-based, empirical, and knowledge-based [5]. Figure 7.1 summarizes the main classes of scoring functions used to estimate protein–ligand interactions, categorizing them into three principal topics: types of functions, datasets (concerning already obtained data), and applications for different functions.
In recent years, several studies related to docking simulations have focused on the development and renement of scoring functions [6], such as:
Machine learning-based scoring functions: Machine learning techniques to
develop more accurate scoring functions have gained prominence in recent
years. Deep learning algorithms are being applied to model complex protein–
ligand interactions [7]. Figure 7.2 illustrates some machine learning-based algo-
rithms that can be applied to scoring functions and optimizing protein–ligand
166 R. M. de Angelo et al.
Fig. 7.1 Diagram of categories, datasets, and applications of scoring functions employed during the protein–ligand docking. (1) Datasets can be used to obtain scoring function (SF), which can estimate binding afnities, for example. (2) Regarding function types, SF can be classied into three categories: knowledge-based, force eld-based, and empirical; below we nd the databases that derive the SF. (3) SF can be applied to estimate interactions between a large number of compounds and proteins, optimizing the search for molecules of interest
7 Molecular Docking: State-of-the-Art Scoring Functions and Search Algorithms 167
Fig. 7.2 Machine learning-based algorithms used in scoring functions, which are present in docking programs. Each acronym represents a unique approach in the eld of machine learning, tailored to different types of data and specic tasks. RF (Random Forest); XGB (XGBoost); GBT (Gradient Boosting Trees); NN (Neural Network); CNN (Convolutional Neural Network); GNN (Graph Neural Network)
interaction calculations. The limitations of generic ML scoring functions include
the possibility of not capturing specic nuances of protein–ligand interaction,
resulting in less precise predictions. Additionally, the reliance on broad datasets
can make models less sensitive to the unique features of specic interactions. On
the other hand, the advantages of generic ML scoring functions lie in their ability
to generalize and computational efciency. These functions can be trained on
large datasets to capture general protein–ligand interaction patterns, making them
applicable to various biological systems. Furthermore, using ML algorithms may
enable the discovery of complex relationships between molecule features and
their interactions, which traditional approaches may not readily identify.
Multiscale scoring functions: Scoring functions that consider multiple scales of
interactions, such as local and global interactions, are being developed to improve
the accuracy of binding predictions. An example of global interactions cited in
this work is related to the long-range electrostatic interactions between charged
residues on the protein surface and polar or charged groups at the ligand mole-
cule. These can play a signicant role in determining the overall stability of the
interactions and the binding afnity of the protein–ligand complex [8].
Integration of structural and energetic data: The integration of structural infor-
mation, such as structural biology data, with energetic data in scoring functions is
becoming more common, enhancing the accuracy of docking simulations. While
168 R. M. de Angelo et al.
this integration improves docking accuracy, it is noteworthy that many original
scoring functions like ALP/ChemScore were actually based on energy results
from pharmaceutical companies [911].
Scoring functions and solvation effect: Scoring functions that take into account
the inuence of solvent on protein–ligand interactions are being developed and
rened to improve accuracy in solvated systems [12, 13].
Taking into account the search algorithms, it is important to mention that some of the original algorithms focused on just sampling the ligand conformation, while the more advanced ones expand on the protein–ligand sampling as a single process. So, the search algorithms integrated into molecular docking platforms serve as robust computational tools for exploring the conformational space of ligands for protein active sites. Each algorithm embodies distinct methodologies tailored to optimize ligand-binding interactions. Below, we can cite some algorithms:
Random Forest (RF): An ensemble learning technique predicated on decision trees,
RF excels in classication and regression tasks. Its adaptability renders it condu-
cive to identifying promising ligand poses within docking studies [14]. XG-Boost (XGB): An optimized implementation of gradient boosting, XGB is
procient in handling classication and regression challenges. While powerful,
vigilant parameter regularization is essential to mitigate overtting risks [15]. Gradient Boosted Trees (GBT): Analogous to XGB but with a more generalized
boosting methodology, GBT offers efcacy in classication and regression
endeavors. Prudent parameter tuning is requisite to circumvent overtting
concerns [16]. Neural Networks (NN): NNs, comprising interconnected layers of neurons, exhibit
prowess in capturing intricate data representations. In docking, NNs are instru-
mental in discerning the intricate relationship between ligand structure and
protein binding afnity [17 ]. Convolutional Neural Networks (CNN): Specialized in grid-structured data
processing, such as images, CNNs are adept at extracting local patterns pertinent
to ligand–protein interactions. Their application in molecular docking facilitates
comprehensive structural analysis [18]. Graph Neural Networks (GNN): Tailored for data represented as graphs, GNNs
excel in modeling complex atomistic interactions within molecular systems.
Leveraging their capabilities, GNNs offer versatile representations of molecular
structures, facilitating spatial relationship elucidation [19].
These algorithms, either employed individually or synergistically, augment the precision and efciency of molecular docking endeavors. The selection of an appropriate algorithm hinges upon the specic attributes of the docking problem and the nature of the data under investigation [20].
About the search algorithm s implemented in docking programs, the main char­acteristics of these algorithms include:
7 Molecular Docking: State-of-the-Art Scoring Functions and Search Algorithms 169
Enhancement of global search algorithms: Global search algorithms, such as
genetic algorithms and particle swarm optimization (heuristic), are being
improved to explore ligand conformation spaces more effect ively [21]. It is
worth noting that the diversity of available search functions has increased,
underscoring the importance of selecting the appropriate algorithm for molecular
docking, as it is not always a one-size-ts-all problem [22].
Hardware acceleration: The use of high-performance hardware, such as GPUs
(Graphics Processing Unit) and TPUs (Tensor Processing Unit), accelerates the
search for binding conformations, making the docking process faster and more
efcient [23].
Intelligent local search methods: Local search algorithms are being enhanced
with intelligent strategies to increase the probability of nding high-afnity
binding conformations [2426].
Integration of molecular dynamics: Integrating molecular dynamics with docking
is becoming more common, allowing for a more realistic consideration of molec-
ular exibility during the search [26, 27].
Flexible conformation simulations: Algorithm s that enable simulations of exible
conformation are being developed, allowing ligands and proteins to change their
conformations during the docking process [28, 29].
Rescoring approaches: Processes in which ligands are re-ranked after the initial
search can help improve the accuracy of predictions [4, 5].
Integration of experimental information: Algorithms that integrate docking and
experimental information, such as mass spectrometry or electron microscopy
data, are being develo ped to enhance validation and prediction accuracy [6].
It is important to note that docking simulations continue to benet from multidisciplinary collaboration among computer scientists, structural biologists, chemists, pharmacologists, and others. Furthermore, applying machine learning techniques and using high-performance computing resources signicantly drives the state-of-the-art in these areas. As new technologies and approaches are devel­oped, research in molecular docking continues to advance, providing new insights into drug candidate discovery/planning and understanding protein–ligand interactions.
170 R. M. de Angelo et al.
2.1 Search Algorithms, Scoring Functions, and Machine
Learning
Machine learning algorithms can be used to optimize parameters in molecular docking techniques, enhancing their ability to predict more accurate molecular interactions. In most of the cases, ML-based scoring functions represent a method­ology based on supervised learning, employing structural data with labels reecting experimentally measured binding afnities. Initially, this ML-integration relies on traditional approaches such as Support Vector Machines (SVM), Random Fo rest (RF), and Gradient Boosting Trees (GBT), aiming to improve scoring performance on reference test sets. The inputs for these models included manually designed descriptors such as molecular ngerprints, ligand features, atom pair terms, and force eld terms [30, 31]. Currently, there are other approaches to incorporate machine learning techniques into scoring functions, which can improve simulation performance and/or optimize the types of result analyses. Below, we will describe some machine learning-based scoring functions presented in Table 7.1.
2.2 Critical Characteristics of Search Algorithms
Search algorithms are classied into three main groups, according to the methodol­ogy applied to explore ligand exibility: systematic, deterministic, and stochastic [47]. Table 7.2 presents the critical characteristics of each algorithm.
As expected, an essential factor in the molecular docking process is the choice of the search algorithm that will position the compound in the active site/binding site of the target. Among the primary algorithms, Genetic Algorithms (GA) and Lamarck­ian Genetic Algorithms (LGA) stand out as stochastic methods and explore the conformations of the receptor–ligand complex by allowing for random conforma­tional modications of the ligand and, occasionally, specic protein residues [53]. It is noteworthy to mention the primary reference for the GOLD program, which employs GA considering the ligand exibility and the exibility of some residue side chains.
GA is based on Darwins theory of evolution, considering the ligands degrees of freedom as genes,composing chromosomesreferring to ligand poses. Changes leading to the generation of new conformations are determined by mutations, which randomly alter the genes,and by crossover,allowing the exchange between two genesthat will be evaluated by the scoring function, considering the new gener­ations those with the highest estimated afnity [54].
Still, in line with evolutionary ideas from biology, the LGA algorithm incorpo­rates the same biological concepts but is applied to computational processes. This approach starts with a population of ligand conformations, evaluates their tness based on interaction energy, and selects the ttest ligands as parents for the next generation. The crossover and mutation processes introduce diversity while the
7 Molecular Docking: State-of-the-Art Scoring Functions and Search Algorithms 171
Table 7.1 Description of some machine learning-based scoring functions
Algorithm Name Input feature Dataset Reference
RF-score Counting pairs of protein–ligand
atom types within a predened
PDBbind v2007
[31]
distance cut-off
RF | Random Forest
SFCscore
Descriptors related to ligand­specic interactions and surface
PDBbind v2007
[32]
RF
area
ΔVinaRF20 Vina empirical and surface area
terms
PDBbind v2014
[33]
CSAR dataset
XGB | eXtreme Gra­dient Boosting
ΔVinaXGB Vina empirical factors, surface
area components, ligand stability factors, and bridge water considerations
ΔLinF9XGB A sequence of Gaussian terms
dening protein–ligand interac­tions, surface area components, ligand characteristics, bridge water
PDBbind v2016 CSAR dataset
PDBbind CSAR dataset BindingDB
[34]
[35]
considerations, and pocket attributes
AGL-Score Features of protein–ligand com-
PDBbind [36] plexes based on algebraic graph theory
GBT | Gradi­ent Boosting Tree
ECIF-GBT Counting pairs of protein–ligand
atom types while taking into account the connectivity of each
PDBbind
v2016
[37]
atom
NN | Nearest Neighbors
NNScore 1.0 Descriptors of specic interactions
and ligand-dependent
NNScore 2.0 Vina empirical factors, counts of
protein–ligand atom-type pairs
MOAD
PDBbind
MOAD
PDBbind
[38]
[39]
within a predened distance threshold
AtomNet Local structure-based 3D grid
DUD-E [40] from protein–ligand structures
CNN | Convolutional Neural
Pafnucy Atom property-based 3D grid from
protein–ligand structures
Kdeep Atom type-based 3D grid from
protein–ligand structures
PDBbind
v2016
PDBbind
v2016
[41]
[42]
[43]
Network
OnionNet Element-specic contacts between
protein and ligand atoms without
PDBbind
v2016 the need for rotation, categorized by different distance ranges
PotentialNet Atom node feature and distance
matrix
PDBbind
v2007
[44]
graphDelta [45]
(continued)
172 R. M. de Angelo et al.
Table 7.1 (continued)
Algorithm Name Input feature Dataset Reference
GNN | Graph Neural Network
SIGN Distance matrix of atom nodes and
Table 7.2 Key characteristics of the main categories of search algorithms
Systematic The systematic approach is the only one mentioned earlier that guarantees
convergence to the global minimum. It systematically and exhaustively explores all ligand degrees of freedom, traversing the energy surface. The main problem related to systematic methods is that computational cost grows explosively with the addition of degrees of freedom in the system. This approach becomes prohibitive for larger ligands (i.e., with many degrees of freedom). This approach is not always suitable for applications that rely on more agile methods [48]
Deterministic The deterministic approach uses energy minimization and molecular dynamics
methods. These methods employ the gradient (i.e., the direction of the steepest ascent on a surface) and system movement over time, respectively. Therefore, they always ensure convergence to an energy minimum. However, such a minimum cannot be guaranteed to correspond to the global minimum of the energy surface. These methods are strictly dependent on the initial conditions of the system and often get trapped in local minima with high energy barriers [49]
Stochastic This approach involves heuristic methods in exploring the energy surface. Sto-
chastic methods randomly vary all ligand degrees of freedom (translational, rotational, and conformational) at each step, generating a wide diversity of solutions. These solutions are evaluated based on a probabilistic criterion to decide whether they will be rejected. Examples of this approach include Monte Carlo and Evolutionary Algorithm [50, 51]. The main disadvantage of this methodology is that convergence to the global minimum is not guaranteed, requiring multiple independent runs of the algorithm to maximize the probability of nding an optimal result [52]
Atom node features considering local environment and distance matrix
angle matrix of bond edges
PDBbind
v2018
PDBbind
v2016
[46]
tness of the descendants is re-evaluated. The process is repeated for several generations until stop criteria are met. The method is called Lamarckianbecause it allows information learned in previous generations to be passed on to future generations, leading to convergence in favorable ligand conformations. In this case, an individual (ligand) with ideal interaction characteristics would pass this information to the new individual or population (group of ligands) [54]. It is essential to highlight that the evolution and the widespread use of molecular docking tech­niques have allowed this computational tool to be used to screen new potential ligands against a large variety of biological targets [55]. Figure 7.3 summarizes the search methodologies that can be explored to understand and improve ligand exibility in a dockin g simulation.