Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5884_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Foreword
- •Acknowledgments
- •Contents
- •1.1 Structure-Based Drug Discovery (SBDD)
- •1.2 Ligand-Based Drug Design (LBDD)
- •1.3 Echoes from the Past, Visions from the Future
- •References
- •1 Introduction
- •2.2 Second Step: Data Curation
- •2.4 Fourth Step: Updating and Maintenance
- •2 Databases and Curation
- •8 Perspectives
- •9 Conclusion
- •References
- •1 Introduction
- •2.1 Making and Matching Protein Models
- •2.2 Simulating Protein Movements
- •2.3 Analyzing Changes in Protein Shape
- •3 Pharmacogenomics in Drug Development
- •4 Case Studies of Genomics-Based Drug Design
- •References
- •1 Historical Background
- •1.1 Timeline
- •2 Methodology Overview
- •2.1 Neural Networks
- •2.1.1 Perceptron
- •2.1.2 Multilayer Neural Networks
- •2.1.3 Types of Neural Networks
- •Feedforward
- •Recurrent Neural Networks
- •LSTM
- •2.2 Deep Learning
- •3 Using Machine Learning
- •3.2 Data Collection
- •3.3 Data Preprocessing
- •3.4 Model Selection
- •3.5 Model Training
- •3.6 Validation
- •3.7 Tuning
- •3.8 Prediction
- •4 Limitations
- •4.1 Bias
- •4.3 Interpretability
- •4.4 Computational Cost
- •4.5 Data Dependency
- •4.6 Robustness
- •5 Applications in Drug Discovery
- •5.2 Lead Discovery
- •5.3 Preclinical and Clinical Development
- •6 Resources and Tools
- •7 Challenges and Perspectives
- •7.1 Future Trends
- •9 Conclusions
- •References
- •1 Historical Background
- •1.1 Applications in Drug Discovery
- •2 Validations and Controls
- •2.1 Internal Validation
- •2.2 External Validation
- •2.3 Relative Cluster Validation
- •3 Challenges and Perspectives
- •4 Conclusions
- •References
- •1 Historical Background
- •2 OECD Principles
- •2.1 A Defined Endpoint
- •2.2 An Unambiguous Algorithm
- •2.5 A Mechanistic Interpretation, if Possible
- •3 Software and Tools
- •4 Validations and Controls
- •4.1 Internal and External Validation
- •4.1.1 Regression Metrics
- •4.2 Applicability Domain
- •4.3 Randomization Tests
- •5 Interpretation
- •6 Practical Advice During QSAR Modeling
- •7 Application
- •8 Challenges and Perspectives
- •References
- •1 Molecular Docking
- •2 Advances in Scoring Functions and Search Algorithms
- •2.2 Critical Characteristics of Search Algorithms
- •2.3 Docking Programs and Scoring Functions
- •3 Calculations Performed During Docking Simulations
- •4 Essential Components for a Good Docking Program
- •5 Limitations of the Docking Technique
- •6 Validation of Docking Results
- •7 Inappropriate Use of Validation Methods in Docking
- •9 Use of Machine Learning in Molecular Docking
- •11 Challenges
- •12 Conclusions
- •References
- •3 System Preparation for MD Simulations
- •3.1 Solvation and Microensemble
- •3.2 Force Fields: General Concept and Relevant Choices
- •3.3 The Concept of Replicas and Timescale
- •4.1.2 Protein Root Mean Square Fluctuation (RMSF)
- •4.1.4 Protein Secondary Structure Analysis
- •4.1.5 Principal component Analysis (PCA)
- •4.1.6 Markov State Modelling
- •4.1.7 Distance Calculations
- •4.1.8 Angle and Plane Calculations
- •4.2.2 Distances and Ligand-Induced Geometry Rearrangements
- •4 Molecular Dynamics Analysis
- •4.1 Protein Perspective
- •4.1.1 Protein Root Mean Square Deviation (RMSD)
- •4.3 Ligand Perspective
- •4.3.1 Ligand Properties
- •4.3.2 Ligand Root Mean Square Deviation
- •4.3.3 Ligand Root Mean Square Fluctuation
- •4.3.4 Angles and Dihedrals
- •5.1 Protein Structure Prediction and Preparation
- •5.2 Molecular Docking
- •6 Concluding Remarks and Outlook
- •Glossary
- •References
- •1 Introduction
- •2.1 MDeNM
- •2.2 Collective Molecular Dynamics (coMD)
- •2.3 ClustENM and ClustENMD
- •3 Ensemble Docking
- •References
- •1 Introduction
- •1.1 Advantages, Disadvantages, Innovations, and Challenges
- •1.2 Recent Advances in Accessible FEP Software Tools
- •1.3 Applications of FEP in Industry and Consortiums
- •2 Expanding the Potential of FEP Calculations
- •2.1 Validating Binding Poses
- •2.2 Dealing with Solvent
- •2.3 FEP and Allostery
- •2.4 FEP and Covalent Ligands
- •2.5 Applications of FEP in Scaffold Hopping
- •2.6 Positional Analogue Scanning
- •2.7 Combinations and Alternative Approaches
- •3 Machine Learning for FEP
- •3.4 Implications for ML in FEP Calculations
- •4 Final Considerations
- •5 First Steps to FEP Simulations
- •References
- •1 Background
- •2 Ultra-Large Screening Libraries and Chemical Spaces
- •3.1 Implications of Dataset Size
- •4 Ligands on the Ultra-Large Scale
- •4.1 Ultra-Large 2D Similarity Searches
- •7 Challenges and Future Perspectives
- •7.1 Hit Triage: An Old Problem on a New Dimension
- •8 Conclusions
- •Appendix
- •References
- •1 Introduction
- •2 Enzymatic Activity Evaluations
- •3 Cytotoxicity Evaluation and Cell Viability
- •4 Antiviral Assays in Experimental Validation
- •6 In Vivo Evaluation of Compounds
- •7 Conclusions
- •References
- •1 Introduction
- •3.1 Data Collection
- •3.2 Data Preprocessing
- •3.4 Model Choice
- •3.5 Model Training
- •3.6 Model Assessment
- •3.7 External Validation
- •3.8 Implementation and Availability
- •3.9 Continuous Update
- •5 Conclusions and Perspectives
- •References
- •1 Experimental Approaches to Obtain Protein Structure
- •1.1 X-Ray Crystallography
- •1.2 Nuclear Magnetic Resonance
- •1.3 Cryo-EM
- •1.4 Hybrid Methods
- •2 Modeling Approaches to Obtain Protein Structure
- •2.1 Homology Modeling
- •2.2 Ab Initio Modeling
- •2.3 New Approaches
- •3 Conformational Diversity of Proteins
- •3.1 Characterization of Protein Conformational States
- •3.2 Experimental Methods to Study Protein Dynamics and Conformations
- •3.4 Molecular Dynamics Simulation
- •3.5 Sampling Strategies
- •4 Remarks and Perspectives
- •References
- •1 Introduction
- •2 Structure-Based Drug Design of HIV Protease Inhibitors
- •2.1 HIV-1 Protease as a Therapeutic Target
- •2.2.1 Saquinavir
- •2.2.2 Indinavir
- •2.3.1 Lopinavir
- •2.3.2 Darunavir
- •6 Conclusions
- •References
- •4 Experimental Methods to Analyze NR Activity
- •4.2 Coregulator-Recruitment
- •5 Concluding Remarks and Outlook
- •References

11 Ultra-Large-Scale Virtual Screening 315
For example, Gorgulla et al. [7 ] released a 1.4-billion version of Enamine REAL
in pdbqt format for direct docking with tools of the AutoDock family. They also
made a 1.5-billion version of the ZINC15 library available in ready-to-dock format
[5]. Sivula et al. [23] released 3D databases of a 1.56 billion Enamine REAL leadlike library in the proprietary Schrödinger Phase database format for use with
Schrödinger tools. These public endeavors are also summarized in Table 11.11 in
the Appendix.
These effor ts also underline why ZINC stands out as a compound resource: In
contrast to most ultra-large libraries, ZINC offers the direct download of
pregenerated 3D structures for many compounds. For instance, 3D structural data
that can be readily obtained from ZINC20 amounts to around 60 TBs [24, 53].
Finally, the work of Grebner et al. [4] showcases a strategy to speed up 3D
conformer generation: They first created a fragment library from 100 million random
SMILES and precalculated the corresponding torsions. The result was then applied
as a torsion library for OpenEye’s OMEGA and speeded up 3D conformer generation for their ultra-large libraries by about 25%.
5 It Takes Two to Tango: Receptor-Based
Ultra-Large-Scale Screening
Despite the higher throughput of ligand-based methods, the bulk of ultra-large-scale
screening efforts seem to revolve around receptor-based screening methods. Their
independence from prior knowledge, coupled with the explicit consideration of the
relevant protein target as a partner, makes docking-based screening an attractive
choice in many settings. In brute-force approaches, as schematically illustrated in
Fig. 11.3, all compounds in an ultra-large library are considered one after another.
Fig. 11.3 Schematic illustration of a brute-force docking study. In brute-force docking, every
compound in a large compound database is processed one after another. One or several 3D
conformers are generated and docked into the target’s binding site (highlighted in green) to evaluate
the fit and interactions

316 I. Pöhner et al.
Several recent examples of ultra-large brute-force docking efforts highlight the
applicability, limitations, and opportunities of docking on the ultra-large scale.
5.1 Recent Examples of Docking-Based Ultra-Large
Scale-Screening
One of the first studies to leverage ultra-large compound databases in a dockingbased VS campaign was performed by Lyu et al. [17]. They utilized DOCK3.7 to
dock more than 99 million compounds from ZINC15 against the AmpC β-lactamase,
and more than 138 million compounds against the D
their key findings would motivate many future screening efforts on a similar or even
larger scale: The authors demonstrated that hit rates can be expected to improve with
growing library sizes, provided docking enriched true positives over decoys on a
small scale. On the other hand, if enrichment failed on a small scale, performance
deteriorated further when utilizing larger libraries [17]. Thus, every ultra-large
docking study should, whenever possible, include validation procedures such as
re- and crossdocking and the retrospective analysis of known hits. Further details on
docking validation can be found in Chap. 7 and in the excellent ultra-large docking
protocol summary published by Bender et al. [24].
By excluding any relatives of known actives from their initial set of virt ual hits,
Lyu et al. [17] discovered low micromolar AmpC inhibitors with novel chemotypes.
The same strategy also enabled the discovery of D
with previously unseen chemotypes: The study found several new nanomolar
agonists and partial agonists, and low micromolar antagonists. The most potent
compound was a highly subtype-specific full D
4
the potential of ultra-large databases in the discovery of highly potent binders by
VS [17 ].
Following this success, subsets of ZINC15 provided various promising virtual
hits against further targets: For instance, 15 novel MT
scaffolds were identified from a library of around 150 million lead-like compounds
and included highly potent agonists and inverse agonists with up to high picomolar
potencies [18]. Likewise, several novel chemotypes, many of which had nanomolar
affinity and high selectivity for the σ2 receptor, were found by screening a ZINC15
subset of 490 million compounds [19]. DOCK3.7 was also utilized by Gahbauer
et al. [54] to screen up to 330 million lead-like compounds from ZINC15 and an
in-house virtual database against the SARS-CoV-2 nonstructural protein 3 (NSP3).
In the wake of the COVID-19 pandemic, Acharya et al. [55] initiated the
development of a supercomputing pipeline to support drug development efforts.
Their workflow involved ensemble docking with Autodock-GPU and GROMACS
molecular dynamics (MD) simulations. Initially, they focuse d on the small-scale
docking of repurposing databases to eight SARS-CoV-2 targets, followed by long
MD simulations. Additionally, they utilized Autodock-GPU to screen around 1.4
dopamine receptor. One of
4
dopamine receptor modulators
4
agonist at 180 pM, demonstrating
melatonin receptor-targeting
1

11 Ultra-Large-Scale Virtual Screening 317
billion compounds against selected targets on the Summit supercomputer of the Oak
Ridge Leadership Computing Facility. Their work represented one of the first gigascale brute-force docking efforts, and the final dataset was made publicly avail able in
2023: [56] The full billion-scale docking results and machine-learning basedrescoring results for 5 SARS-CoV-2 targets, namely the Spike protein, the main
protease (Mpro) and papain-like protease (PLpro), RNA-dependent RNA polymerase (RdRp), and the nonstructural protein 15 (NSP15). Mpro was also the focus of a
dedicated ultra-large VS effort: Luttens et al. [57] combined the docking of 235 million ZINC15 compounds with a billion-scale fragment-based optimization and
experimental validation, which yielded eight new inhibitors of Mpro.
A dedicated virtual screening platform that leverages massive parallelization
during ligand preparation and docking, VirtualFlow, was recently presented by
Gorgulla et al. [7]. This Open Source platform integrates with several open-source
docking tools like AutoDock Vina [58] and many of its forks (e.g., smina [59]) and
tools with free-for academics licensing like PLANTS [7, 60, 61]. VirtualFlow was
used to screen around 1.3 billion compounds from ZINC15 and Enamine REAL
against the Kelch-like ECH-associated protein 1 (KEAP1) protein [7]. The screen
identified several structurally diverse binders with submicromolar affinity. In
another extensive screening effort, VirtualFlow enabled the billion-scale screening
of 15 potential target proteins of SARS-CoV-2 and two host proteins of critical
relevance to the virus, namely ACE2 and TMPRSS2 [62]. For each of the 45 binding
sites screened, the top one million virtual hits were made publicly available to
facilitate future SARS-CoV-2 drug development efforts.
Sivula et al. [23] focused on a giga-scale benchmarking effort for a machine
learning (ML)-boosted VS approach (see also Sect. 6.1). To generate reference data,
the authors first performed a brute-force docking study of the 1.56 billion Enamine
REAL lead-like library (status March 2021) to two target proteins, the bacterial
periplasmic chaperone SurA and Cyclin-G associated kinase (GAK).
Sadybekov et al. [63] reported a docking-based multi-template VS of 115 million
compounds from Enamine REAL’s 350/3 lead-like library and the REAL diversity
subset against two G protein-coupled receptors (GPCRs). They discovered five
promising submicromolar compounds, four of which were selective antagonists of
cysteinyl leukotriene GPCR CysLT1R, while one compound was a full antagonist of
both screened GPCRs.
Kaplan et al. [64] took a different strategy to find novel selective binders of the
serotonin 5-HT
receptor: The authors noted that although tetrahydropyridine scaf-
2A
folds can be found in several drug compounds derived from natural products, in
typical screening libraries, they are underrepresented in comparison to their aromatic
counterparts. Thus, they first generated a specialized library of 75 million virtual
tetrahydropyridine compounds from commercially available building blocks. This
bespoke library was then screened with DOCK3.7, and chemical novelty was
ensured by Tanimoto-similarity filtering against known actives. The study resulted
in 17 selected virtual hits, which included one novel agonist and three antagonists
with low micromolar activities (24% hit rate) [64].

318 I. Pöhner et al.
Table 11.4 Summary of recent ultra-large-scale screening campaigns performed with conventional brute-force docking approaches
Compounds
screened Target(s) Docking tool
1.56 billion SurA, GAK Glide-HTVS 20-32 cpds/
1.40 billion 5
a
SARS-CoV-2 targets AutoDock-GPU >50 cpds/
Throughput
per core Reference(s)
Sivula et al.
min
[23, 65]
Acharya
min
et al.
[55, 56, 66]
1.30 billion KEAP1 VirtualFlow
(QuickVina, smina,
4–12 cpds/
min
Gorgulla
et al. [7]
AutoDock Vina)
1.01 billion 15 SARS-CoV-2 tar-
gets; host ACE2,
TMPRSS2
VirtualFlow
(QuickVina W,
QuickVina 2)
8 cpds/min Gorgulla
et al. [62]
490 million σ2 receptor DOCK3.8 46 cpds/min Alon et al.
[19]
330 million SARS-CoV-2 NSP3 DOCK3.7 84 cpds/min Gahbauer
et al. [54]
246 million SARS-CoV-2 NSP3 DOCK3.7 65 cpds/min Gahbauer
et al. [54]
235 million SARS-CoV-2 Mpro DOCK3.7 47 cpds/min Luttens et al.
[57]
150 million MT
Melatonin receptor DOCK3.7 56 cpds/min Stein et al.
1
[18]
138 million Dopamine D
receptor DOCK3.7 53 cpds/min Lyu et al.
4
[17]
115 million CysLT1R, CysLT2R ICM-Pro NA Sadybekov
et al. [63]
99 million AmpC β-lactamase DOCK3.7 40 cpds/min Lyu et al.
[17, 24]
75 million Serotonin 5-HT
receptor
2A
DOCK3.7 144 cpds/
min
Kaplan et al.
[64]
The reported number of screened compounds refers to the input prior to any conformer generation.
Throughput per core is approximated from the reported timing information
GAK cyclin G-associated kinase, KEAP1 Kelch-like ECH-associated protein 1, NSP3 nonstructural
protein 3, Mpro main protease, CysLT1R, CysLT2R cysteinyl leukotriene G protein-coupled
receptors, NA not available (authors did not disclose total run-time or speed information)
a
The initial publication investigated 8 targets, while only five were featured in the final giga-scale
dataset release
Table 11.4 summarizes all discussed ultra-large docking-based VS approaches
and their approxi mate throughput.

11 Ultra-Large-Scale Virtual Screening 319
5.2 Dedicated Tools for Ultra-Large Docking Workflows
Many of the presented ultra-large docking campaigns were achieved by “simply”
scaling up conventional docking workflows. One exception is the tool VirtualFlow,
which was specifically developed to address challenges associated with dataintensive docking workflows on the ultra-large scale [7]. Especially for novice
users, VirtualFlow aims to simplify key tasks, such as splitting data into appropriate
batches and scheduling and monitoring computing jobs in an HPC or cloud computing environment. Since workflows supported by VirtualFlow scale linearly, they
straightforwardly scale with the available resources and allow for uncomplicated
cost estimates.
As noted above, VirtualFlow works with various docking tools to enable users to
pick the most well-suited tool for each docking problem. It achieves high throughput
by utilizing up to four levels of parallelization: Each instance of VirtualFlow can run
several jobs, jobs can have several job steps, and each job step can run several
queues with the actual ligand preparation or docking runs—using tools, which are
often also internally parallelized. It supports various common resource managers for
batch jobs on HPCs, e.g., SLURM and Moab/TORQUE/PBS. Given the multiple
layers of parallelization, a dedicated workload balancing system ensures that file I/O
remains unproblematic and each ligand is considered only once [7].
While VirtualFlow was designed to be more of a “Swiss Army knife” easily
deployed in most common HPC and cloud computing environments, a recently
published open-source pipeline takes a different approach: warpDOCK was developed specifically for cloud computing in the Oracle Cloud Infrastructure [67]. The
authors chose QuickVina2 [68] as the primary docking engine, but warpDOCK can
also integrate with AutoDock Vina and its various forks. Its Python-based queue
engine takes care of active monitoring and the preloading of ligands. In a test
docking with equivalent resources, the authors demonstrated warpDOCK to be
about 3.7 times more efficient than VirtualFlow in their cloud computing
environment.
In an ensemble docking study against 21 Staphylococcus aureus D-alanine-Dalanine ligase structures, warpDOCK achieved a throughput of around 12 compounds per minute per physical core [67]. Although the authors utilized a comparatively small ligand library of only 4.75 million ZINC compounds, the ensemble
docking reached the ultra-large scale with around 101 million docking calculations
in total.
The example of warpDOCK highlights how careful optimization for a specific
architecture can maximize throughput. VirtualFlow, on the other hand, guarantees
decent throughput in different architectures and maximizes a user’s choice of the
computing environment. On the ultra-large scale, job scheduling and task distribution over the computing cores represent particular bottlenecks for docking throughput. We highlighted two examples of open-source tools to support users in

320 I. Pöhner et al.
navigating those challenges (see also Table 11.10 in the Appendix for the
corresponding github addresses). With the increasing popularity of ultra-largescale approaches in mind, we anticipate that other, similar tools will likely be
developed.
6 Reducing Computation Needs in Ultra-Large VS
Workflows
As can be extrapolated, for example, from the throughput in Table 11.4, utilizing
brute-force approac hes to ultra-large-scale docking remains a highly time- and
resource-intensive task. Strategies to reduce the required computations were thus
quickly sought after.
One option involves using so-called diversity sets, i.e., subsets of ultra-large
libraries that aim to reflect the chemical diversity of the entire library. For instance,
instead of investing 14 million CPU hours to dock the full ultra-large Enamine and
SAVI libraries, a recent study targeting the STAT3 N-terminal domain relied on
small million-scale Enamine diversity subsets and a bespoke diversity subset of
SAVI with close to three million compounds [69]. Although the target is commonly
considered “nondruggable,” the authors were able to successfully discover novel
ligands in their docking-based VS of diversity sets.
On the other hand, reduction strategies may lead to a loss of promising binders: In
their study of the dopamine D
clustering an ultra-large library and docking only cluster representatives. They found
that screening only cluster representatives would have prevented the discovery of
several of their novel promising binders. This underlines why there remains strong
interest in the ability to screen entire ultra-large libraries.
Two promising solutions have gained particula r traction over the past few years:
The first combines ligand-based QSAR-like approaches with docking by utilizing
machine learning (ML) to construct models that can serve as stand-ins for conventional screening. The second solution relies on the combinatorial nature of many
ultra-large make-on-demand libraries and their step-by-st ep construction from fragments, thereby avoiding the enume ration of most of the library.
receptor, Lyu et al. [17] investigated the prospect of
4
6.1 Accelerating Docking with Machine Learning/Deep
Learning in Hybrid Ligand- and Structure-Based VS
To remove the necessity of brute-force docking a complete ultra-large compound
library, most recently described ML-accelerated docking approaches train ML
models as stand-ins to predict the docking outcome, as illustrated schematically in
Fig. 11.4. Model training relies either on one-shot learning or active learning. In

11 Ultra-Large-Scale Virtual Screening 321
Fig. 11.4 Schematic illustration of a docking study accelerated with machine learning (ML) or
deep learning (DL). A small random subset of the ultra-large compound database is selected,
conformers are generated, and the compounds are docked by conventional means. The docking
outcome is used to train the ML/DL model, and the trained model acts as a surrogate to predict the
docking outcome for the remainder of the ultra-large library
one-shot learning, the model is trained a single time [35]. However, recent
approaches largely rely on several iterative training steps and an active learning
concept. Based on the prediction results of the current iteration, active learning adds
additional training samples for all following steps [35].
A reduced number of brute-force docking calculations is the key to the speed up
of ultra-large-scale screening with ML. While individual approaches differ in details,
their surrogate models commonly ensure that only a small fraction of the library
needs to be docked by conventional means. The concept of using shallow ML
algorithms and QSAR approaches to reduce required dockings is not new [70], but
many recent methods rely on deep learning (DL) implementations or the combination of DL with more classical, shallow ML algorithms (see Table 11.5).
For example, the open-source tool DeepDocking aims to accelerate VS with DL
[71]. DeepDocking first computes molecular descriptors, such as circular fingerprints, for all compounds in a chemical library. Then, a random compound set is
docked by conventional means, and docking scores and descriptors are used to train
a deep neural network (DNN). DeepDocking uses a classification approach, where a
docking score cutoff determines the separation of compounds into the classes “hits”
and “nonhits.” Once trained, the surrogate model predicts the class of each member
in the entire library, and compounds designated as “nonhits” are removed from the
prediction set. DeepDocking relies on active learning: Iteratively, additional random
samples are brute-force docked, and the score threshold to distinguish “hits” from
“nonhits” is progressively adjusted [71]. This stepwise reduction of the considered
compounds in an ultra-large library was able to speed up the screening of 1.36 billion
3D structures from ZINC15 by at least 50 times [38, 71]. DeepDocking was able to
recover up to 90% of the top-scoring virtual hits for 12 different targets docked with
FRED when considering only a small fraction of the library compounds
[38, 71]. The ZINC15 library was additionally screened against the severe acute

322 I. Pöhner et al.
Table 11.5 Summary of discussed ML-/DL-boosted docking implementations
Tool Type Input
Open source
DeepDocking
[71]
HASTEN [79] Regression SMILES D-MPNN Predicted top-scoring
Linear accel-
erated
docking [85]
Lean docking
[84]
MolPAL [77] Regression Fingerprints,
Commercial
Glide active
Learning [81]
This overview lists the type of model, utilized training input, ML/DL architecture, and available
selection metrics. See also Table 11.6 for specific ultra-large-scale application examples of the tools
DNN deep neural network, D-MPNN directed message-passing graph neural network, RF random
forest, GCN graph convolutional neural network, TS Thompson sampling, EI expectation of
improvement, PI probability of improvement, UCB upper confidence bound
Classification Fingerprints,
properties
Regression Fingerprints Linear
Regression 2D ligands Support
SMILES
Regression Fingerprints,
SMILES
ML/DL
architecture Selection strategy
DNN By predicted class
regression
ensemble
Vector
regressor
RF, DNN,
D-MPNN
RF + GCN Predicted top-scoring, most
Predicted top-scoring
Predicted top-scoring
Predicted top-scoring, TS, EI,
PI, UCB
uncertain, most uncertain among
top scoring, random among
top-scoring
respiratory syndrome coronavirus 2 (SARS-CoV-2) main protease, which allowed
for the discovery of novel inhibitors with micromolar IC
values [72]. Following
50
that, an extensive consensus screen of around 40 billion structures sourced from
ZINC15 and the Enamine REAL Space was performed against the same target,
utilizing a DeepDocking pipeline with five different docking tools [34]. This effort
resulted in over 100 additional inhibitors with novel chemotypes. Other application
examples of DeepDocking include the discovery of novel noncovalent inhibitors of
SARS-CoV-2 papain-like protease (PLpro), A
adenosine receptor antagonists, and
2A
inhibitors of a prostate cancer target, the RNA-binding protein Lin28 [73–75]. In
addition to its various use cases as summarized in Table 11.6, detailed instructions
for running a DeepDocking workflow have been published [38], and a graphical user
interface (GUI) was recently released to simplify the adoption of DeepDocking [76].
Another open-source tool, MolPAL (molecular pool-based active learning), relies
on Bayesian Optimization and offers three different ML regression models [77]:
Two rely on molecular fingerprints, namely the more classical random forest (RF),
and a simple feed-forward DNN model. The third option uses SMILES input and the
directed message-passing graph neural network (D-MPNN) Chemprop [78]. The
authors showcase MolPAL’s activ e learning strategy on small-scale docking-based
VS against thymidylate kinase using libraries of 10,000 to around two million
compounds [77]. Additionally, they explore the ultra-large scale with the AmpC
and Dopamine D
receptor docking results published by Lyu et al. [17]. In addition
4

11 Ultra-Large-Scale Virtual Screening 323
Table 11.6 Summary of ultra-large applications of discussed ML-/DL-accelerated docking tools
Compounds
screened Target(s) Tool
>12.3 billion
(38.7 billion)
Mpro DeepDocking
(5 docking
Runtime
[h] CPUs GPUs Reference(s)
456 640 250 Gentile et al.
[34]
tools)
1.56 billion
(3.8 billion)
1.40 billion A
SurA, GAK HASTEN
(Glide-HTVS)
2A
DeepDocking NA NA NA Tang et al.
223 640 10 Sivula et al.
[23]
[74]
1.40 billion PLpro DeepDocking
(Glide-SP)
1.36 billion 12 diverse
targets
DeepDocking
(FRED)
NA NA NA Garland et al.
[73]
NA 60 4 Gentile et al.
[71]
1.30 billion Mpro DeepDocking NA 390 40 Ton et al.
[72]
1.00 billion Lin28 DeepDocking NA NA NA Radaeva
et al. [75]
138 million D
138 million D
4
4
MolPAL
(DOCK3.7)
Glide Active
Learning
NA NA NA Graff et al.
[77]
60,000 8 1 Yang et al.
[81]
(Glide)
99 million AmpC MolPAL
(DOCK3.7)
99 million AmpC Glide Active
Learning
NA NA NA Graff et al.
[77]
NA NA NA Yang et al.
[81]
(Glide)
96 million AmpC Lin. accel.
docking
<164– Marin et al.
[85]
(DOCK3.7)
Compounds screened: Raw total input numbers of utilized library, where reported. 3D structures
after conformer generation are given in parenthesis. Brute-force docking tools used together with
ML-/DL-acceleration tools are likewise reported in parenthesis
Lin. accel. docking: linear accelerated docking. Total runtimes are reported for the highest number
of iterations in active learning approaches and numbers of CPUs and GPUs represent the maximum
(in most cases, more CPUs were utilized for docking and lower numbers for the ML-/DL-step and
more GPUs for prediction than training). GAK cyclin-G associated kinase; A
; PLpro SARSCoV-2 papain-like protease, Mpro SARS-CoV-2 main protease. NA not available
A
2A
adenosine receptor
2A
(authors did not provide total run-time or speed information)
to comparing different ML/DL strategies, Graff et al. [77] study different acquisition
strategies, i.e., how the next compounds for brute-force docking are selected. They
explore the selection based purely on the rank obtained by the docking scores,
referred to as “greedy selection,” Thompson sampling (TS), expectation of improvement (EI), probability of improvement (PI), and upper confidence bound (UCB) (see
Table 11.5). Especially with smaller training dataset sizes, they observe large
variations of recalls depending on the chosen ML archi tecture and acquisition

324 I. Pöhner et al.
metric. Interestingly, the most basic greedy selection typically ranges among the
top-performing strategies. For instance, training the MPNN with only 2.4% of a
100 million-compound dataset allowed for around 89% recall of the top 50,000
compounds. Taking uncertainty into account in the same example improved recalls
to 95% [77]. This work represents a blueprint to show how different architectures,
selection strategies, and training data batch sizes impact the recalls of ML-boosted
strategies.
A tool that combi nes the MPNN Chemprop and a greedy selection algorithm with
built-in support for the proprietary docking tool Glide is the open-source tool
HASTEN (macHine leArning booSTEd dockiNg) [ 78–80]. Like DeepDocking
and MolPAL, HASTEN generates its initial training data by brute-force docking a
random selection of compounds from a large chemical library. The MPNN model
training data consists of compound SMILES and docking scores. Following the
prediction of docking scores for the entire compound library by the regression
model, HASTEN takes an active learning approach: Compounds are ranked by
their predicted scores, and the top-ranked compounds are brute-force docked before
training a new model with all available docking data. The progressive addit ion of
both true and false positives to the training data iteratively refines the model. On the
million scale, HASTEN achieved high recall values across 13 different targets
docked with either FRED or Glide [79]. HASTEN additionally represents the first
tool for ML-accelerated docking that has been rigorously benchmarked on the billion
scale: Sivula et al. [23] confirmed over 90% recall of the true top-scoring 1000 hits
by HASTEN in their comparison with 1.56 billion brute-force docking results for
two different targets. The authors showcased the significant speed-up when using
HASTEN for the billion-scale screening: While their brute-force docking took
approximately 4 months, the time to screen 1.56 billion compounds was reduced
to 10 days or less when using HASTEN with the time-consuming brute-force
docking reduced by at least 99% [23].
The three discussed open-source tools (see Table 11.10 in the Appendix for
source code repositories) demonstrate how both classification and regression models
can reduce the amount of brute-force docking when screening ultra-large chemical
libraries. They highlight the significant speed-up and pinpoint implications of
training dataset sizes and selection metrics in active learning approaches [23, 71,
77, 79].
Unsurprisingly, in addition to the efforts in the academic environment, similar
strategies have also been integrated into commercial screening software. For example, the Glide Active Learning strategy has been available as part of the Schrödinger
Software Suite since 2019 [81]. Glide Active Learning relies on an ensemble of a
graph convolutional NN (GCN) an d a RF model and features various selection
strategies for compounds to dock after the first iteration of randomly selected
training compounds. Typically, only one round of active learning is employed,
which features docking of 0.1% of the most uncertain top 5% according to ensemble
predictions to augment the training data. Other examples of c ommercial tools
include, for example, ICM-Pro Gigascreen (based on a GCN) and OpenEye’s
GigaDock [82, 83 ].
Соседние файлы в папке Библиотека им академика М.И. Перельмана
