Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5435_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
10.10.2026
Размер:
10 Мб
Скачать
☆
[40] Murray, C. W., & Blundell, T. L. Structural biology in fragment-based drug design. Current Opinion in
Structural Biology, 2010, 20(4), 497–507.
[41] Zhang, J., Yang, P. L., & Gray, N. S. Target identification and validation in drug discovery using novel
proteomics and bioinformatics approaches. Expert Opinion on Therapeutic Targets, 2009, 13(1),
1–11.
[42] Jorgensen, W. L., & Zhu, W. Quantum mechanics and virtual screening applications in drug
discovery. Chemical Reviews, 2020, 120(15), 6904–6953.
[43] Baldi, P., & Brunak, S. Bioinformatics: The Machine Learning Approach. 2001, MIT press: Cambridge,
MA. 02142 617-253-5249.
[44] Choi, Y., & Kendrew, S. G. Cloud computing for protein-ligand docking and virtual screening
algorithms. In: Bioinformatics and Biomedical Engineering. 2015, (pp. 677–688). Springer: Cham:
Shanghai, China, 18–20 September 2015.
[45] Goswami, A., & Datta, A. Virtual screening in personalized medicine. In: Himansu Sekhar Behera,
Durga Prasad Mohapatra (eds) Computational Intelligence in Data Mining-Volume 2. 2016,
(pp. 183–198). Springer: Cham.
[46] Rani, I., Goyal, A., & Sharma, M. Computational design of phosphatidylinositol 3-kinase inhibitors.
ASSAY and Drug Development Technologies, 2022, 20(7), 317–337.
[47] Nguyen, T., & Mathias, S. Applications of big data in virtual screening and drug discovery: A review.
Molecules, 2019, 24(23), 4235.
[48] Li, J., & Zhu, F. Multi-target drugs: The trend of drug research and development. PLoS One, 2017,
12(5), e0176656.
[49] Rani, N., Sharma, P., Sharma, V. K., & Kumar, P. Molecular docking approach to identify potential
anticandidal potential of curcumin. Journal of Pharmaceutical Technology, Research and
Management, 2020, 8(2), 67–71. https://doi.org/10.15415/jptrm.2020.82008.
[50] Vamathevan, J., Clark, D., Czodrowski, P., Dunham, I., Ferran, E., Lee, G., & Whitty, A. Applications of
machine learning in drug discovery and development. Nature Reviews Drug Discovery, 2019, 18(6),
463–477.
64 Gagandeep Kaur et al.
https://t.me/med1917
Pawan Kumar and Ajit Kumar
✶
4 State-of-the-art modeling techniques
in performing docking algorithms
and scoring
Abstract: In drug discovery and molecular biology, molecular docking plays a pivotal
role in predicting the binding modes and affinities of small molecules with target pro-
teins. This abstract explores the current state-of-the-art modeling techniques em-
ployed in performing docking algorithms and scoring methods. Docki ng algorithms
are essential tools used to predict the preferred orientation of one molecule to a sec-
ond when bound to each other to form a stable complex. Over the years, various
approaches have been developed, ranging from geometric matching to advanced ma-
chine learning-based methods. These techniques often integrate molecular mechanics,
quantum mechanics, and empirical scoring functions to accurately predict binding
poses. Moreover, scoring functions are critical components in evaluating the affinity
between the ligand and the receptor. Traditional scoring functions are often based on
empirical parameters derived from experimental data. However, recent advance-
ments have witnessed the integration of machine learning models, deep learning ar-
chitectures, and physics-based potentials to enhance scoring accuracy and reliability.
The chapter discusses the significant advancements in docking algorithms, including
flexible docking, induced fit docking, and ensemble docking, which better capture the
dynamic nature of protein-ligand interactions. Additionally, it highlights the emer-
gence of innovative scoring functions, such as free energy-based scoring and machine
learning-driven scoring, which aim to improve the precision of binding affinity pre-
dictions. Furthermore, the chapter addresses the challenges and limitations associated
with current modeling techniques, including computational complexity, scoring func-
tion bias, and the incorporation of protein flexibility. Overall, it provides insights into
the cutting-edge methodologies shaping the landscape of molecular docking and scor-
ing, paving the way for more efficient and accurate drug discovery processes.
Keywords: Molecular docking, Flexible docking, Rigid docking, Scoring function, Bind-
ing energy
✶
Corresponding author: Ajit Kumar, Toxicology and Computational Biology Group,
Centre for Bioinformatics, M. D. University, Rohtak 124001, Haryana, India,
e-mail: akumar.cbt.mdu@gmail.com; ajitkumar.cbinfo@mdurohtak.ac.in
Pawan Kumar, School of Agricultural Biotechnology, Punjab Agricultural University, Ludhiana, Punjab,
India
https://doi.org/10.1515/9783111207117-004
https://t.me/med1917
4.1 Introduction
Molecular docking is a computational technique that has emerged as a crucial tool in
the field of structural bioinformatics, playing a pivotal role in drug discovery to pre-
dict and analyze the binding interactions between small molecules (ligands), typically
drug candidates, and target macromolecules (receptors) such as proteins or nucleic
acids. The fundamental goal of molecular docking is to explore and evaluate the ener-
getically favorable conformations of the ligand within the binding site of the receptor.
The predicted binding free energy (ΔG
bind
) between a ligand and a receptor is mod-
eled in terms of dispersion and repulsion (ΔG
vdw
), hydrogen bond (ΔG
hbond
), desolva-
tion (ΔG
desolv
), electrostatic (ΔG
elec
), torsional free energy (ΔG
tor
), final total internal
energy (ΔG
total
), and unbound system’s energy (ΔG
unb
) [1, 2].
Molecular docking explores the conformational space of both the ligand and the
receptor. The process can be broken down into several key steps, contributing to the
overall accuracy and reliability of the predictions. The first step involves the prepara-
tion of the 3D molecular structure of the ligand with appropriate charges and atom
types. The next step is the retrieval of the 3D molecular structure of the macromole-
cule (molecular modeling if the 3D structure is not available), structure refinement,
and defining the binding site. Once the ligand and the receptor are prepared, the ac-
tual docking simulation takes place. During this phase, various algorithms, scoring
functions, and search strategies (Table 4.1) are employed to explore the vast confor-
mational space between the ligand and the receptor. The docking algorithm generates
and evaluates potential binding geometries, predicting the most energetically favor-
able binding mode based on factors such as van der Waals forces, hydrogen bonding,
and electrostatic interactions [1, 3, 4]. The scoring functions quantify the fitness of a
particular ligand-receptor conformation and based on the data, predict the accuracy
of biologically relevant binding poses from less favorable ones [5, 6].
Molecular docking finds application in various areas of biomedical research, in-
cluding drug discovery, virtual screening, and the study of ligand–protein interactions
(PPI). In drug discovery, it assists researchers in the identification of potential lead
compounds and the optimization of drug candidates with their precise interaction at
the binding pocket [7]. The Virtual high-throughput screening (VHTS) approach lever-
ages molecular docking to rapidly screen large chemical libraries, identifying promis-
ing candidates for further experimental validation [8–10]. T he efficiency of virtual
screening significantly accelerates the drug discovery process, saving time and resour-
ces. Moreover, molecular docking facilitates the exploration of the structure-activity
relationship (SAR), providing valuable insights into how structural modifications im-
pact the binding interactions and pharmacological properties of the compounds.
Despite various applications of molecular docking, there are some challenges and
limitations. The docking accuracy is highly dependent on the molecular structure
quality for both the ligand and the receptor, as inaccurate structure can lead to unre-
liable docking results [11, 12]. Furthermore, the flexibility of both the ligand and the
66 Pawan Kumar and Ajit Kumar
https://t.me/med1917
receptor poses a significant challenge in molecular docking [2]. Capturing different
conformations and dynamics of macromolecule flexibility in a docking study adds a
layer of complexity, which increases the conformational space and requirement of
computational resources. Advanced techniques, such as molecular dynamics simula-
tions and ensemble docking, can address the dynamic nature of biomolecules in the
docking process. Additionally, the docking algorithm and scoring functions used to
evaluate binding affinities may not always capture the complexities of biological sys-
tems accurately [1–6, 12–14].
Table 4.1: Molecular docking program and algorithms.
Name Docking algorithm Docking type
Affinity docking Monte Carlo Flexible–rigid docking
Autodock Genetic algorithm, Lamarckian genetic algorithm, and
Monte Carlo
Flexible–rigid docking
AutoDock Vina Genetic algorithm Flexible–rigid docking
DARWIN Genetic algorithm Flexible–rigid docking
DIVALI Genetic algorithm Flexible–rigid docking
DOCK Fragmentation algorithm, incremental construction,
and matching algorithm
Flexible docking
eHiTS Incremental construction Flexible–rigid docking
FlexX Fragmentation algorithm and incremental construction Flexible–rigid docking
FLOG Matching algorithm Flexible docking
Glide Exhaustive systematic search Flexible docking
GOLD Genetic algorithm Flexible docking
Hammerhead Incremental construction Flexible–rigid docking
ICM Monte Carlo Flexible–rigid docking
LibDock Matching algorithm Flexible–rigid docking
LeDOCK Simulated annealing and genetic algorithm Flexible docking
RDOCK Genetic algorithm, Monte Carlo, and simplex
minimization
Rigid docking
QXP Monte Carlo Flexible–rigid docking
SANDOCK Matching algorithm Flexible docking
SLIDE Incremental construction Flexible–rigid docking
Surflec-Dock Incremental construction Flexible–rigid docking
ZDOCK Geometric complementarity and molecular dynamics Rigid docking
4 State-of-the-art modeling techniques in performing docking algorithms 67
https://t.me/med1917
The simplistic nature of scoring functions may overlook subtle nuances in the binding
interactions, impacting the precision of the predictions. Advanced machine learning
approaches have been employed to optimize scoring functions and enhance the pre-
dictive power of molecular docking algorithms. Various molecular docking algorithms
are used to compute the interaction between the receptor and the ligand.
4.2 Docking algorithms
Molecular docking algorithms aim to explore the conformational space of the ligand
within the binding site of th e recepto r and predict the most energetically favorable
bindingmode[15,16].The“lock-and-key model” was one of the initially proposed
enzymatic–substrate interaction models, which refers to the rigid docking between
receptors and ligands [3, 17]. However, in biological systems, the docking process is
flexible (induced fit model) where receptors and ligands have to change their confor-
mation to fit each other well [18]. Flexible docking methods can consider several pos-
sible conformations of ligand and receptor, but requires a higher computational time.
Hence, rigid docking is fast but is less accurate than flexible docking, where the later
can predict the correct docked position of the ligand. The earlier versions of Molecu-
lar docking programs of DOCK [19], FLOG, FTDOCK, HEX [20] and RDOCK [21], use
rigid dock parameters, while AutoDock4 [22], AutoDock Vina [6], DOCK (latest version)
[23, 24], FlexX [25], Glide (Schrodinger) [26], MS-DOCK, SYSDOC, and Ros ettaLi gand
[27–29] utilize flexible dock parameters. Being fast, the rigid docking method has been
employed for virtual screening of small-molecule database. Docking accuracy can be
increased using better crystallographic structures or using empirical docking algo-
rithms (evolutionary programming, fast shape matching (SM) algorithms, fragment-
based methods (incremental construction algorithm), genetic algorithms (GAs), Monte
Carlo (MC), simulated annealing (SA), and Tabu search (TS).
4.3 Fast shape matching algorithm (SM)
The algorithm maps the pharmacophore between the ligand and the active site of the
receptor after a geometrical overlap between them. The different ligand conforma-
tions are governed by the distance and scoring matrix between the receptor and the
ligand pharmacophore. Chemical properties like hydrogen bond donors and acceptors
are considered for docking scores. The algorithm has the advantage of speed, and
hence used for screening large libraries of small compounds. Rigid docking applica-
tions (earlier version of DOCK and ZDOCK) use the SM algorithm as their search strat-
egies [30, 31], which combine shape complementarity, desolvation, and electrostatics
parameters through a Fast Fourier Transform (FFT) algorithm to count the binding
68 Pawan Kumar and Ajit Kumar
https://t.me/med1917
energy between the receptor and the ligand. SM algorithm is also an integral part of
Flexible dock applications like DOCK [23], EUDOC [32, 33], and SYSDOC where sphere-
matching procedure is combined with the incremental construction method to find
the best docking conformation of the receptor and the ligand [11].
4.4 Incremental construction algorithm
It is a type of molecular docking algorithm that builds a ligand’s conformation within
the binding site of a target macromolecule step by step [25, 34]. First, the ligand is frag-
mented by breaking its rotatable bonds, followed by docking individual fragments into
the active site, and incrementally adding the remaining fragments within the binding
pocket. Thus, different fragments result in different ligand orientations to fill the active
site, which provides a more systematic and controlled exploration of the binding site
[34]. Incremental construction algorithms are particularly useful in situations where a
ligand undergoes conformational changes upon binding or where there is a need for a
more targeted exploration of the binding site. Docking programs Surflex-Dock [35],
DOCK, FlexX [25], Hammerhead [36], SLIDE, and eHiTS [37–39] use the incremental con-
struction algorithm for docking ligands within the active site of the receptor.
4.5 Genetic algorithm (GA)
The basic idea of GA stems from Darwin’s theory of evolution. Degrees of freedom of
the ligand are encoded as binary strings called genes. These genes make up the “chro-
mosome”, which represents the pose of the ligand. Mutation and crossover are two
kinds of genetic operators in GA, causing random changes and excha nges of genes
between two chromosomes, respectively. When the genetic operators affect the genes,
the result is a new ligand structure. New structures will be assessed by a scoring func-
tion, and the ones that survived (i.e., exceeded a threshold) can be used for the next
generation. The GA algorithm requires the approximate size and location of the recep-
tor active site (binding pocket) along with the 3D Cartesian coordinates of the protein
and the ligand [40–42]. GA has been used in AutoDock [43], GOLD [44, 45], DIVALI,
and DARWIN [46] docking programs.
Lamarckian GA (LGA) is also implemented in docking algorithms, which switches
between the genotypic space and the phenotypic space. Mutation and crossover occur
in the genotypic space, while the phenotypic space is determined by the energy func-
tion to be optimized. Energy minimization is performed after genotypic changes have
been made in the phenotypic space, which is conceptually similar to MC minimization
[47]. The phenotypic changes from energy minimization are mapped back onto the
genes (by changing the ligand coordinates in the chromosome). AutoDock uses LGA to
4 State-of-the-art modeling techniques in performing docking algorithms 69
https://t.me/med1917
calculate the binding interaction between the ligand and the binding pocket of the
receptor [8, 43, 48].
4.6 Monte Carlo method
It is a stochastic optimization algorithm that explores the conformational space (pose)
of ligands within the binding site of a target macromolecule. Bond rotation and rigid
body translation are used to generate different conformation spaces of ligands, which
are further tested using energy-based selection criteria [47, 49, 50]. The process is iter-
ated until a predefined quantity of conformation is collected. Earlier versions of Auto-
dock, ICM, QXP, and Affinity docking tools used the MC method for molecular docking
study.
4.7 Simulated annealing (SA) method
In SA, a biomolecular system is simulated by a specific kind of dynamic simulation.
Every docking conformation is carried into a simulation where the temperature is de-
creased gradually during regular intervals of time in each cycle of simulation. It may
give a higher accuracy result when compared with others methods, since it considers
the detailed conformational state and flexibility of both the protein and the ligand in
different thermodynamic states in an interval of time [51].
The accuracy and performance of SA methods can be improved by combination
with other docking algorithms/m ethods. SA may be combined with the MC method,
resulting in the MC-SA protocol. In MC-SA, random changes are made in ligand orien-
tation inside the protein binding site during each SA temperature cycle. The energy of
the current state is compared with previous the state energy and the lowest energy is
chosen to be compared with the next state.
4.8 Tabu search (TS)
It is an iterative procedure designed for obtaining solution of optimization problems.
It was developed and described by Glover and has been used to solve a large variety
of hard o ptimization problems. This procedure can be defined as a Meta-Heuristic
methodology that can move from one solution to another, being able to save in mem-
ory the already visited solutions.
TS algorithm is an extension of local search methods. For molecular docking algo-
rithms, the search space refers to all possible conformations between two molecules.
70 Pawan Kumar and Ajit Kumar
https://t.me/med1917
4.9 Scoring functions
The purpose of the scoring function is to delineate the correct poses from incorrect
poses, or binders from ina ctive compounds in a reasonable computation time [40].
However, scoring functions involve estimating, rather than calculating the binding af-
finity between the protein and ligand, and through these functions, adopting various
assumptions and simplifications [3, 4, 40, 52–54,]. Scoring functions can be divided
into force-field-based, empirical, and knowledge-based scoring functions
4.10 Force-field-based scoring functions
Classical force-field-based scoring functions assess the binding energy by calculating the
sum of the nonbonded (electrostatics and van der Waals) interactions. The electrostatic
and van der Waals terms are calculated by Coulombic and Lennard-Jones potential for-
mulation, respectively. To decrease the computation time, the cutoff distance is used to
handle the nonbonded interactions. The extensions of force-field-based scoring functions
consider the hydrogen bonds, solvation, and entropy contributions. Furthermore, the re-
sults of docking with force-field-based functions can be further refined with other techni-
ques, such as linear interaction energy and free-energy perturbation methods to improve
theaccuracyinpredictingbindingenergies[55,49,50].Dockingprograms,suchasDOCK
[56], GOLD [45], and AutoDock [43], offer such functions.
4.11 Empirical scoring functions
In empirical scoring functions, binding energy decomposes into several energy com-
ponents, like hydrogen bonds, ionic interaction, hydrophobic effect, and binding en-
tropy [57, 58]. Each energy component is multiplied by a coefficient and then summed
up to give a final score.
4.12 Knowledge-based scoring functions
Knowledge-based scoring functions use statistical analysis of ligand–protein complexes
crystal structures to obtain the interatomic contact frequencies and the distances be-
tween them [59]. The scoring is based on the assumption that the more favorable an
interaction, the greater the frequency of occurrence. The score is calculated by favoring
preferred contacts and penalizing repulsive interactions between each atom in the li-
gand and the protein within a given cutoff [60]. PMF, DrugScore, SMoG, and Bleep are
4 State-of-the-art modeling techniques in performing docking algorithms 71
https://t.me/med1917
examples of knowledge-based functions that differ mainly in the training sets size, en-
ergy function, atom types, and distance cutoff.
4.13 Consensus scoring function
Consensus scoring combines different bonded and nonbonded interaction scores to
assess the docking conformation. A pose of ligand is accepted when it scores well
under several different scoring schemes. The consensus scoring function is greatly
used in virtual screening, and improves the prediction of bound conformations and
poses. Consensus score (CScore) is an example of consensus scoring function that com-
bines DOCK, ChemScore, PMF, GOLD, and FlexX scoring functions.
4.14 Docking program
More than 60 docking programs are reported, with about 0.4 million records over
Google Scholar and 60,000 records over PubMed. Here, we discuss the major molecu-
lar docking programs used.
4.15 AutoDock
AutoDock is a widely used molecular docking software, with 20,000 publication search
records over NCBI PubMed central. It was developed by the Olson laboratory at The
Scripps Research Institute to predict the binding modes and binding affinities of small
molecules with target macromolecules, typically proteins. Autodock employs MC-SA,
evolutionary, genetic, and Lamarckian GAs for global optimization of ligand flexibility
while keeping the receptor rigid [43]. The scoring function is based on the AMBER
force field, including van der Waals, hydrogen bonding, electrostatic interactions, con-
formational entropy, and desolvation terms. Recent updates of AutoDock can handle
the flexibility of amino acid in the bindingpocket,alongwithligandflexibility.
Evolvement in computational resources and availability to screen large datasets made
AutoDock the first preference of most researchers across the globe, with higher num-
ber of records over PubMed last year (Figure 4.1).
72 Pawan Kumar and Ajit Kumar
https://t.me/med1917
2010
0
500
Publication search in NCBI PubMed central
1000
1500
2000
2500
3000
3500
4000
4500
5000
Autodock Darwin
Dock
FlexX Glide GOLD Hex Vina zdock
2011 2012 2013 2014 2015 2016
Yea r
2017 2018 2019 2020 2021 2022
Figure 4.1: Publication record of selected molecular docking tools since 2010.
4 State-of-the-art modeling techniques in performing docking algorithms 73
https://t.me/med1917