Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5364_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
02.09.2026
Размер:
21 Мб
Скачать
11 Ultra-Large-Scale Virtual Screening 315
For example, Gorgulla et al. [7 ] released a 1.4-billion version of Enamine REAL in pdbqt format for direct docking with tools of the AutoDock family. They also made a 1.5-billion version of the ZINC15 library available in ready-to-dock format [5]. Sivula et al. [23] released 3D databases of a 1.56 billion Enamine REAL lead­like library in the proprietary Schrödinger Phase database format for use with Schrödinger tools. These public endeavors are also summarized in Table 11.11 in the Appendix.
These effor ts also underline why ZINC stands out as a compound resource: In contrast to most ultra-large libraries, ZINC offers the direct download of pregenerated 3D structures for many compounds. For instance, 3D structural data that can be readily obtained from ZINC20 amounts to around 60 TBs [24, 53].
Finally, the work of Grebner et al. [4] showcases a strategy to speed up 3D conformer generation: They rst created a fragment library from 100 million random SMILES and precalculated the corresponding torsions. The result was then applied as a torsion library for OpenEyes OMEGA and speeded up 3D conformer genera­tion for their ultra-large libraries by about 25%.
5 It Takes Two to Tango: Receptor-Based
Ultra-Large-Scale Screening
Despite the higher throughput of ligand-based methods, the bulk of ultra-large-scale screening efforts seem to revolve around receptor-based screening methods. Their independence from prior knowledge, coupled with the explicit consideration of the relevant protein target as a partner, makes docking-based screening an attractive choice in many settings. In brute-force approaches, as schematically illustrated in Fig. 11.3, all compounds in an ultra-large library are considered one after another.
Fig. 11.3 Schematic illustration of a brute-force docking study. In brute-force docking, every compound in a large compound database is processed one after another. One or several 3D conformers are generated and docked into the targets binding site (highlighted in green) to evaluate the t and interactions
316 I. Pöhner et al.
Several recent examples of ultra-large brute-force docking efforts highlight the applicability, limitations, and opportunities of docking on the ultra-large scale.
5.1 Recent Examples of Docking-Based Ultra-Large
Scale-Screening
One of the rst studies to leverage ultra-large compound databases in a docking­based VS campaign was performed by Lyu et al. [17]. They utilized DOCK3.7 to dock more than 99 million compounds from ZINC15 against the AmpC β-lactamase, and more than 138 million compounds against the D their key ndings would motivate many future screening efforts on a similar or even larger scale: The authors demonstrated that hit rates can be expected to improve with growing library sizes, provided docking enriched true positives over decoys on a small scale. On the other hand, if enrichment failed on a small scale, performance deteriorated further when utilizing larger libraries [17]. Thus, every ultra-large docking study should, whenever possible, include validation procedures such as re- and crossdocking and the retrospective analysis of known hits. Further details on docking validation can be found in Chap. 7 and in the excellent ultra-large docking protocol summary published by Bender et al. [24].
By excluding any relatives of known actives from their initial set of virt ual hits, Lyu et al. [17] discovered low micromolar AmpC inhibitors with novel chemotypes. The same strategy also enabled the discovery of D with previously unseen chemotypes: The study found several new nanomolar agonists and partial agonists, and low micromolar antagonists. The most potent compound was a highly subtype-specic full D
4
the potential of ultra-large databases in the discovery of highly potent binders by VS [17 ].
Following this success, subsets of ZINC15 provided various promising virtual hits against further targets: For instance, 15 novel MT scaffolds were identied from a library of around 150 million lead-like compounds and included highly potent agonists and inverse agonists with up to high picomolar potencies [18]. Likewise, several novel chemotypes, many of which had nanomolar afnity and high selectivity for the σ2 receptor, were found by screening a ZINC15 subset of 490 million compounds [19]. DOCK3.7 was also utilized by Gahbauer et al. [54] to screen up to 330 million lead-like compounds from ZINC15 and an in-house virtual database against the SARS-CoV-2 nonstructural protein 3 (NSP3).
In the wake of the COVID-19 pandemic, Acharya et al. [55] initiated the development of a supercomputing pipeline to support drug development efforts. Their workow involved ensemble docking with Autodock-GPU and GROMACS molecular dynamics (MD) simulations. Initially, they focuse d on the small-scale docking of repurposing databases to eight SARS-CoV-2 targets, followed by long MD simulations. Additionally, they utilized Autodock-GPU to screen around 1.4
dopamine receptor. One of
4
dopamine receptor modulators
4
agonist at 180 pM, demonstrating
melatonin receptor-targeting
1
11 Ultra-Large-Scale Virtual Screening 317
billion compounds against selected targets on the Summit supercomputer of the Oak Ridge Leadership Computing Facility. Their work represented one of the rst giga­scale brute-force docking efforts, and the nal dataset was made publicly avail able in 2023: [56] The full billion-scale docking results and machine-learning based­rescoring results for 5 SARS-CoV-2 targets, namely the Spike protein, the main protease (Mpro) and papain-like protease (PLpro), RNA-dependent RNA polymer­ase (RdRp), and the nonstructural protein 15 (NSP15). Mpro was also the focus of a dedicated ultra-large VS effort: Luttens et al. [57] combined the docking of 235 mil­lion ZINC15 compounds with a billion-scale fragment-based optimization and experimental validation, which yielded eight new inhibitors of Mpro.
A dedicated virtual screening platform that leverages massive parallelization during ligand preparation and docking, VirtualFlow, was recently presented by Gorgulla et al. [7]. This Open Source platform integrates with several open-source docking tools like AutoDock Vina [58] and many of its forks (e.g., smina [59]) and tools with free-for academics licensing like PLANTS [7, 60, 61]. VirtualFlow was used to screen around 1.3 billion compounds from ZINC15 and Enamine REAL against the Kelch-like ECH-associated protein 1 (KEAP1) protein [7]. The screen identied several structurally diverse binders with submicromolar afnity. In another extensive screening effort, VirtualFlow enabled the billion-scale screening of 15 potential target proteins of SARS-CoV-2 and two host proteins of critical relevance to the virus, namely ACE2 and TMPRSS2 [62]. For each of the 45 binding sites screened, the top one million virtual hits were made publicly available to facilitate future SARS-CoV-2 drug development efforts.
Sivula et al. [23] focused on a giga-scale benchmarking effort for a machine learning (ML)-boosted VS approach (see also Sect. 6.1). To generate reference data, the authors rst performed a brute-force docking study of the 1.56 billion Enamine REAL lead-like library (status March 2021) to two target proteins, the bacterial periplasmic chaperone SurA and Cyclin-G associated kinase (GAK).
Sadybekov et al. [63] reported a docking-based multi-template VS of 115 million compounds from Enamine REALs 350/3 lead-like library and the REAL diversity subset against two G protein-coupled receptors (GPCRs). They discovered ve promising submicromolar compounds, four of which were selective antagonists of cysteinyl leukotriene GPCR CysLT1R, while one compound was a full antagonist of both screened GPCRs.
Kaplan et al. [64] took a different strategy to nd novel selective binders of the serotonin 5-HT
receptor: The authors noted that although tetrahydropyridine scaf-
2A
folds can be found in several drug compounds derived from natural products, in typical screening libraries, they are underrepresented in comparison to their aromatic counterparts. Thus, they rst generated a specialized library of 75 million virtual tetrahydropyridine compounds from commercially available building blocks. This bespoke library was then screened with DOCK3.7, and chemical novelty was ensured by Tanimoto-similarity ltering against known actives. The study resulted in 17 selected virtual hits, which included one novel agonist and three antagonists with low micromolar activities (24% hit rate) [64].
318 I. Pöhner et al.
Table 11.4 Summary of recent ultra-large-scale screening campaigns performed with conven­tional brute-force docking approaches
Compounds screened Target(s) Docking tool
1.56 billion SurA, GAK Glide-HTVS 20-32 cpds/
1.40 billion 5
a
SARS-CoV-2 targets AutoDock-GPU >50 cpds/
Throughput per core Reference(s)
Sivula et al.
min
[23, 65] Acharya
min
et al. [55, 56, 66]
1.30 billion KEAP1 VirtualFlow (QuickVina, smina,
4–12 cpds/ min
Gorgulla et al. [7]
AutoDock Vina)
1.01 billion 15 SARS-CoV-2 tar-
gets; host ACE2, TMPRSS2
VirtualFlow (QuickVina W, QuickVina 2)
8 cpds/min Gorgulla
et al. [62]
490 million σ2 receptor DOCK3.8 46 cpds/min Alon et al.
[19]
330 million SARS-CoV-2 NSP3 DOCK3.7 84 cpds/min Gahbauer
et al. [54]
246 million SARS-CoV-2 NSP3 DOCK3.7 65 cpds/min Gahbauer
et al. [54]
235 million SARS-CoV-2 Mpro DOCK3.7 47 cpds/min Luttens et al.
[57]
150 million MT
Melatonin receptor DOCK3.7 56 cpds/min Stein et al.
1
[18]
138 million Dopamine D
receptor DOCK3.7 53 cpds/min Lyu et al.
4
[17]
115 million CysLT1R, CysLT2R ICM-Pro NA Sadybekov
et al. [63]
99 million AmpC β-lactamase DOCK3.7 40 cpds/min Lyu et al.
[17, 24]
75 million Serotonin 5-HT
receptor
2A
DOCK3.7 144 cpds/
min
Kaplan et al. [64]
The reported number of screened compounds refers to the input prior to any conformer generation. Throughput per core is approximated from the reported timing information GAK cyclin G-associated kinase, KEAP1 Kelch-like ECH-associated protein 1, NSP3 nonstructural protein 3, Mpro main protease, CysLT1R, CysLT2R cysteinyl leukotriene G protein-coupled receptors, NA not available (authors did not disclose total run-time or speed information)
a
The initial publication investigated 8 targets, while only ve were featured in the nal giga-scale
dataset release
Table 11.4 summarizes all discussed ultra-large docking-based VS approaches
and their approxi mate throughput.
11 Ultra-Large-Scale Virtual Screening 319
5.2 Dedicated Tools for Ultra-Large Docking Workows
Many of the presented ultra-large docking campaigns were achieved by simply scaling up conventional docking workows. One exception is the tool VirtualFlow, which was specically developed to address challenges associated with data­intensive docking workows on the ultra-large scale [7]. Especially for novice users, VirtualFlow aims to simplify key tasks, such as splitting data into appropriate batches and scheduling and monitoring computing jobs in an HPC or cloud com­puting environment. Since workows supported by VirtualFlow scale linearly, they straightforwardly scale with the available resources and allow for uncomplicated cost estimates.
As noted above, VirtualFlow works with various docking tools to enable users to pick the most well-suited tool for each docking problem. It achieves high throughput by utilizing up to four levels of parallelization: Each instance of VirtualFlow can run several jobs, jobs can have several job steps, and each job step can run several queues with the actual ligand preparation or docking runsusing tools, which are often also internally parallelized. It supports various common resource managers for batch jobs on HPCs, e.g., SLURM and Moab/TORQUE/PBS. Given the multiple layers of parallelization, a dedicated workload balancing system ensures that le I/O remains unproblematic and each ligand is considered only once [7].
While VirtualFlow was designed to be more of a Swiss Army knifeeasily deployed in most common HPC and cloud computing environments, a recently published open-source pipeline takes a different approach: warpDOCK was devel­oped specically for cloud computing in the Oracle Cloud Infrastructure [67]. The authors chose QuickVina2 [68] as the primary docking engine, but warpDOCK can also integrate with AutoDock Vina and its various forks. Its Python-based queue engine takes care of active monitoring and the preloading of ligands. In a test docking with equivalent resources, the authors demonstrated warpDOCK to be about 3.7 times more efcient than VirtualFlow in their cloud computing environment.
In an ensemble docking study against 21 Staphylococcus aureus D-alanine-D­alanine ligase structures, warpDOCK achieved a throughput of around 12 com­pounds per minute per physical core [67]. Although the authors utilized a compar­atively small ligand library of only 4.75 million ZINC compounds, the ensemble docking reached the ultra-large scale with around 101 million docking calculations in total.
The example of warpDOCK highlights how careful optimization for a specic architecture can maximize throughput. VirtualFlow, on the other hand, guarantees decent throughput in different architectures and maximizes a users choice of the computing environment. On the ultra-large scale, job scheduling and task distribu­tion over the computing cores represent particular bottlenecks for docking through­put. We highlighted two examples of open-source tools to support users in
320 I. Pöhner et al.
navigating those challenges (see also Table 11.10 in the Appendix for the corresponding github addresses). With the increasing popularity of ultra-large­scale approaches in mind, we anticipate that other, similar tools will likely be developed.
6 Reducing Computation Needs in Ultra-Large VS
Workows
As can be extrapolated, for example, from the throughput in Table 11.4, utilizing brute-force approac hes to ultra-large-scale docking remains a highly time- and resource-intensive task. Strategies to reduce the required computations were thus quickly sought after.
One option involves using so-called diversity sets, i.e., subsets of ultra-large libraries that aim to reect the chemical diversity of the entire library. For instance, instead of investing 14 million CPU hours to dock the full ultra-large Enamine and SAVI libraries, a recent study targeting the STAT3 N-terminal domain relied on small million-scale Enamine diversity subsets and a bespoke diversity subset of SAVI with close to three million compounds [69]. Although the target is commonly considered nondruggable,the authors were able to successfully discover novel ligands in their docking-based VS of diversity sets.
On the other hand, reduction strategies may lead to a loss of promising binders: In their study of the dopamine D clustering an ultra-large library and docking only cluster representatives. They found that screening only cluster representatives would have prevented the discovery of several of their novel promising binders. This underlines why there remains strong interest in the ability to screen entire ultra-large libraries.
Two promising solutions have gained particula r traction over the past few years: The rst combines ligand-based QSAR-like approaches with docking by utilizing machine learning (ML) to construct models that can serve as stand-ins for conven­tional screening. The second solution relies on the combinatorial nature of many ultra-large make-on-demand libraries and their step-by-st ep construction from frag­ments, thereby avoiding the enume ration of most of the library.
receptor, Lyu et al. [17] investigated the prospect of
4
6.1 Accelerating Docking with Machine Learning/Deep
Learning in Hybrid Ligand- and Structure-Based VS
To remove the necessity of brute-force docking a complete ultra-large compound library, most recently described ML-accelerated docking approaches train ML models as stand-ins to predict the docking outcome, as illustrated schematically in Fig. 11.4. Model training relies either on one-shot learning or active learning. In
11 Ultra-Large-Scale Virtual Screening 321
Fig. 11.4 Schematic illustration of a docking study accelerated with machine learning (ML) or deep learning (DL). A small random subset of the ultra-large compound database is selected, conformers are generated, and the compounds are docked by conventional means. The docking outcome is used to train the ML/DL model, and the trained model acts as a surrogate to predict the docking outcome for the remainder of the ultra-large library
one-shot learning, the model is trained a single time [35]. However, recent approaches largely rely on several iterative training steps and an active learning concept. Based on the prediction results of the current iteration, active learning adds additional training samples for all following steps [35].
A reduced number of brute-force docking calculations is the key to the speed up of ultra-large-scale screening with ML. While individual approaches differ in details, their surrogate models commonly ensure that only a small fraction of the library needs to be docked by conventional means. The concept of using shallow ML algorithms and QSAR approaches to reduce required dockings is not new [70], but many recent methods rely on deep learning (DL) implementations or the combina­tion of DL with more classical, shallow ML algorithms (see Table 11.5).
For example, the open-source tool DeepDocking aims to accelerate VS with DL [71]. DeepDocking rst computes molecular descriptors, such as circular nger­prints, for all compounds in a chemical library. Then, a random compound set is docked by conventional means, and docking scores and descriptors are used to train a deep neural network (DNN). DeepDocking uses a classication approach, where a docking score cutoff determines the separation of compounds into the classes hits and nonhits.Once trained, the surrogate model predicts the class of each member in the entire library, and compounds designated as nonhitsare removed from the prediction set. DeepDocking relies on active learning: Iteratively, additional random samples are brute-force docked, and the score threshold to distinguish hitsfrom nonhitsis progressively adjusted [71]. This stepwise reduction of the considered compounds in an ultra-large library was able to speed up the screening of 1.36 billion 3D structures from ZINC15 by at least 50 times [38, 71]. DeepDocking was able to recover up to 90% of the top-scoring virtual hits for 12 different targets docked with FRED when considering only a small fraction of the library compounds [38, 71]. The ZINC15 library was additionally screened against the severe acute
322 I. Pöhner et al.
Table 11.5 Summary of discussed ML-/DL-boosted docking implementations
Tool Type Input
Open source
DeepDocking [71]
HASTEN [79] Regression SMILES D-MPNN Predicted top-scoring Linear accel-
erated docking [85]
Lean docking [84]
MolPAL [77] Regression Fingerprints,
Commercial
Glide active Learning [81]
This overview lists the type of model, utilized training input, ML/DL architecture, and available selection metrics. See also Table 11.6 for specic ultra-large-scale application examples of the tools DNN deep neural network, D-MPNN directed message-passing graph neural network, RF random forest, GCN graph convolutional neural network, TS Thompson sampling, EI expectation of improvement, PI probability of improvement, UCB upper condence bound
Classication Fingerprints,
properties
Regression Fingerprints Linear
Regression 2D ligands Support
SMILES
Regression Fingerprints,
SMILES
ML/DL architecture Selection strategy
DNN By predicted class
regression ensemble
Vector regressor
RF, DNN, D-MPNN
RF + GCN Predicted top-scoring, most
Predicted top-scoring
Predicted top-scoring
Predicted top-scoring, TS, EI, PI, UCB
uncertain, most uncertain among top scoring, random among top-scoring
respiratory syndrome coronavirus 2 (SARS-CoV-2) main protease, which allowed for the discovery of novel inhibitors with micromolar IC
values [72]. Following
50
that, an extensive consensus screen of around 40 billion structures sourced from ZINC15 and the Enamine REAL Space was performed against the same target, utilizing a DeepDocking pipeline with ve different docking tools [34]. This effort resulted in over 100 additional inhibitors with novel chemotypes. Other application examples of DeepDocking include the discovery of novel noncovalent inhibitors of SARS-CoV-2 papain-like protease (PLpro), A
adenosine receptor antagonists, and
2A
inhibitors of a prostate cancer target, the RNA-binding protein Lin28 [7375]. In addition to its various use cases as summarized in Table 11.6, detailed instructions for running a DeepDocking workow have been published [38], and a graphical user interface (GUI) was recently released to simplify the adoption of DeepDocking [76].
Another open-source tool, MolPAL (molecular pool-based active learning), relies on Bayesian Optimization and offers three different ML regression models [77]: Two rely on molecular ngerprints, namely the more classical random forest (RF), and a simple feed-forward DNN model. The third option uses SMILES input and the directed message-passing graph neural network (D-MPNN) Chemprop [78]. The authors showcase MolPALs activ e learning strategy on small-scale docking-based VS against thymidylate kinase using libraries of 10,000 to around two million compounds [77]. Additionally, they explore the ultra-large scale with the AmpC and Dopamine D
receptor docking results published by Lyu et al. [17]. In addition
4
11 Ultra-Large-Scale Virtual Screening 323
Table 11.6 Summary of ultra-large applications of discussed ML-/DL-accelerated docking tools
Compounds screened Target(s) Tool
>12.3 billion (38.7 billion)
Mpro DeepDocking
(5 docking
Runtime [h] CPUs GPUs Reference(s)
456 640 250 Gentile et al.
[34]
tools)
1.56 billion (3.8 billion)
1.40 billion A
SurA, GAK HASTEN
(Glide-HTVS)
2A
DeepDocking NA NA NA Tang et al.
223 640 10 Sivula et al.
[23]
[74]
1.40 billion PLpro DeepDocking (Glide-SP)
1.36 billion 12 diverse
targets
DeepDocking (FRED)
NA NA NA Garland et al.
[73]
NA 60 4 Gentile et al.
[71]
1.30 billion Mpro DeepDocking NA 390 40 Ton et al.
[72]
1.00 billion Lin28 DeepDocking NA NA NA Radaeva
et al. [75]
138 million D
138 million D
4
4
MolPAL (DOCK3.7)
Glide Active Learning
NA NA NA Graff et al.
[77]
60,000 8 1 Yang et al.
[81]
(Glide)
99 million AmpC MolPAL
(DOCK3.7)
99 million AmpC Glide Active
Learning
NA NA NA Graff et al.
[77]
NA NA NA Yang et al.
[81]
(Glide)
96 million AmpC Lin. accel.
docking
<164 Marin et al.
[85]
(DOCK3.7)
Compounds screened: Raw total input numbers of utilized library, where reported. 3D structures after conformer generation are given in parenthesis. Brute-force docking tools used together with ML-/DL-acceleration tools are likewise reported in parenthesis Lin. accel. docking: linear accelerated docking. Total runtimes are reported for the highest number of iterations in active learning approaches and numbers of CPUs and GPUs represent the maximum (in most cases, more CPUs were utilized for docking and lower numbers for the ML-/DL-step and more GPUs for prediction than training). GAK cyclin-G associated kinase; A
; PLpro SARSCoV-2 papain-like protease, Mpro SARS-CoV-2 main protease. NA not available
A
2A
adenosine receptor
2A
(authors did not provide total run-time or speed information)
to comparing different ML/DL strategies, Graff et al. [77] study different acquisition strategies, i.e., how the next compounds for brute-force docking are selected. They explore the selection based purely on the rank obtained by the docking scores, referred to as greedy selection,Thompson sampling (TS), expectation of improve­ment (EI), probability of improvement (PI), and upper condence bound (UCB) (see Table 11.5). Especially with smaller training dataset sizes, they observe large variations of recalls depending on the chosen ML archi tecture and acquisition
324 I. Pöhner et al.
metric. Interestingly, the most basic greedy selection typically ranges among the top-performing strategies. For instance, training the MPNN with only 2.4% of a 100 million-compound dataset allowed for around 89% recall of the top 50,000 compounds. Taking uncertainty into account in the same example improved recalls to 95% [77]. This work represents a blueprint to show how different architectures, selection strategies, and training data batch sizes impact the recalls of ML-boosted strategies.
A tool that combi nes the MPNN Chemprop and a greedy selection algorithm with built-in support for the proprietary docking tool Glide is the open-source tool HASTEN (macHine leArning booSTEd dockiNg) [ 7880]. Like DeepDocking and MolPAL, HASTEN generates its initial training data by brute-force docking a random selection of compounds from a large chemical library. The MPNN model training data consists of compound SMILES and docking scores. Following the prediction of docking scores for the entire compound library by the regression model, HASTEN takes an active learning approach: Compounds are ranked by their predicted scores, and the top-ranked compounds are brute-force docked before training a new model with all available docking data. The progressive addit ion of both true and false positives to the training data iteratively renes the model. On the million scale, HASTEN achieved high recall values across 13 different targets docked with either FRED or Glide [79]. HASTEN additionally represents the rst tool for ML-accelerated docking that has been rigorously benchmarked on the billion scale: Sivula et al. [23] conrmed over 90% recall of the true top-scoring 1000 hits by HASTEN in their comparison with 1.56 billion brute-force docking results for two different targets. The authors showcased the signicant speed-up when using HASTEN for the billion-scale screening: While their brute-force docking took approximately 4 months, the time to screen 1.56 billion compounds was reduced to 10 days or less when using HASTEN with the time-consuming brute-force docking reduced by at least 99% [23].
The three discussed open-source tools (see Table 11.10 in the Appendix for source code repositories) demonstrate how both classication and regression models can reduce the amount of brute-force docking when screening ultra-large chemical libraries. They highlight the signicant speed-up and pinpoint implications of training dataset sizes and selection metrics in active learning approaches [23, 71,
77, 79].
Unsurprisingly, in addition to the efforts in the academic environment, similar strategies have also been integrated into commercial screening software. For exam­ple, the Glide Active Learning strategy has been available as part of the Schrödinger Software Suite since 2019 [81]. Glide Active Learning relies on an ensemble of a graph convolutional NN (GCN) an d a RF model and features various selection strategies for compounds to dock after the rst iteration of randomly selected training compounds. Typically, only one round of active learning is employed, which features docking of 0.1% of the most uncertain top 5% according to ensemble predictions to augment the training data. Other examples of c ommercial tools include, for example, ICM-Pro Gigascreen (based on a GCN) and OpenEyes GigaDock [82, 83 ].