Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5884_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
02.09.2026
Размер:
21 Мб
Скачать
294 D. Antunes et al.
49. Thompson, J., Walters, W. P., Feng, J. A., Pabon, N. A., Xu, H., Maser, M., et al. (2022). Optimizing active learning for free energy calculations. Articial Intelligence in the Life Sciences, 2, 100050.
50. Majellaro, M., Jespers, W., Crespo, A., Núñez, M. J., Novio, S., Azuaje, J., et al. (2021). 3,4-dihydropyrimidin-2(1H)-ones as antagonists of the human A2B adenosine receptor: Opti­mization, structure–activity relationship studies, and enantiospecic recognition. Journal of Medicinal Chemistry, 64, 458–480.
51. Shan, Y., Mysore, V. P., Lefer, A. E., Kim, E. T., Sagawa, S., & Shaw, D. E. (2022). How does a small molecule bind at a cryptic binding site? PLOS Computational Biology, 18, e1009817.
52. Kimura, S. R., Hu, H. P., Ruvinsky, A. M., Sherman, W., & Favia, A. D. (2017). Deciphering cryptic binding sites on proteins by mixed-solvent molecular dynamics. Journal of Chemical Information and Modeling, 57, 1388–1401.
53. Goodford, P. J. (1985). A computational procedure for determining energetically favorable binding sites on biologically important macromolecules. Journal of Medicinal Chemistry, 28, 849–857.
54. Alvarez-Garcia, D., & Barril, X. (2014). Molecular simulations with solvent competition quantify water displaceability and provide accurate interaction maps of protein binding sites. Journal of Medicinal Chemistry, 57, 8530–8539.
55. Alvarez-Garcia, D., Schmidtke, P., Cubero, E., & Barril, X. (2022). Extracting atomic contributions to binding free energy using molecular dynamics simulations with mixed solvents (MDmix). Current Drug Discovery Technologies, 19,62–68.
56. Deng, Y., & Roux, B. (2008). Computation of binding free energy with molecular dynamics and grand canonical Monte Carlo simulations. The Journal of Chemical Physics, 128.
57. Ross, G. A., Bodnarchuk, M. S., & Essex, J. W. (2015). Water sites, networks, and free energies with grand canonical Monte Carlo. Journal of the American Chemical Society, 137, 14930–14943.
58. Michel, J., Tirado-Rives, J., & Jorgensen, W. L. (2009). Energetics of displacing water molecules from protein binding sites: Consequences for ligand optimization. Journal of the American Chemical Society, 131, 15403–15411.
59. Ben-Shalom, I. Y., Lin, Z., Radak, B. K., Lin, C., Sherman, W., & Gilson, M. K. (2020). Accounting for the central role of interfacial water in protein–ligand binding free energy calculations. Journal of Chemical Theory and Computation, 16, 7883–7894.
60. Bruce Macdonald, H. E., Cave-Ayland, C., Ross, G. A., & Essex, J. W. (2018). Ligand binding free energies with adaptive water networks: Two-dimensional grand canonical alchemical perturbations. Journal of Chemical Theory and Computation, 14, 6586–6597.
61. Nussinov, R., & Tsai, C.-J. (2012). The different ways through which specicity works in orthosteric and allosteric drugs. Current Pharmaceutical Design, 18, 1311.
62. Liang, L., Liu, H., Xing, G., Deng, C., Hua, Y., Gu, R., et al. (2022). Accurate calculation of absolute free energy of binding for SHP2 allosteric inhibitors using free energy perturbation. Physical Chemistry Chemical Physics, 24, 9904–9920.
63. Zheng, H., Alter, S., & Qu, C.-K. (2009). SHP-2 tyrosine phosphatase in human diseases. International Journal of Clinical and Experimental Medicine, 2, 17.
64. Torrente, E., Fodale, V., Ciammaichella, A., Ferrigno, F., Ontoria, J. M., Ponzi, S., et al. (2023). Discovery of a novel series of imidazopyrazine derivatives as potent SHP2 allosteric inhibitors. ACS Medicinal Chemistry Letters, 14, 156–162.
65. Xiaoli, A., Yuzhen, N., Qiong, Y., Yang, L., Yao, X., & Bing, Z. (2022). Investigating the dynamic binding behavior of PMX53 cooperating with allosteric antagonist NDT9513727 to C5a anaphylatoxin chemotactic receptor 1 through Gaussian accelerated molecular dynamics and free-energy perturbation simulations. ACS Chemical Neuroscience, 13, 3502–3511.
66. Fu, H., Chen, H., Cai, W., Shao, X., & Chipot, C. (2021). BFEE2: Automated, streamlined, and accurate absolute binding free-energy calculations. Journal of Chemical Information and Modeling, 61, 2116–2123.
10 Free Energy Perturbation and Free-Energy Calculations Applied to Drug Design 295
67. Backus, K. M., Correia, B. E., Lum, K. M., Forli, S., Horning, B. D., González-Páez, G. E., et al. (2016). Proteome-wide covalent ligand discovery in native biological systems. Nature, 534, 570–574.
68. Sutanto, F., Konstantinidou, M., & Dömling, A. (2020). Covalent inhibitors: A rational approach to drug discovery. RSC Medicinal Chemistry, 11, 876–884.
69. Chatterjee, P., Botello-Smith, W. M., Zhang, H., Qian, L., Alsamarah, A., Kent, D., et al. (2017). Can relative binding free energy predict selectivity of reversible covalent inhibitors? Journal of the American Chemical Society, 139, 17945–17952.
70. Awoonor-Williams, E., Walsh, A. G., & Rowley, C. N. (2017). Modeling covalent-modier drugs. Biochimica et Biophysica Acta (BBA) - Proteins and Proteomics, 1865, 1664–1675.
71. Bonatto, V., Shamim, A., Rocho, F. D. R., Leitão, A., Luque, F. J., Lameira, J., et al. (2021). Predicting the relative binding afnity for reversible covalent inhibitors by free energy perturbation calculations. Journal of Chemical Information and Modeling, 61, 4733–4744.
72. Zhang, H., Jiang, W., Chatterjee, P., & Luo, Y. (2019). Ranking reversible covalent drugs: From free energy perturbation to fragment docking. Journal of Chemical Information and Modeling, 59, 2093–2102.
73. Lameira, J., Bonatto, V., Cianni, L., dos Reis, R. F., Leitão, A., & Montanari, C. A. (2019). Predicting the afnity of halogenated reversible covalent inhibitors through relative binding free energy. Physical Chemistry Chemical Physics, 21, 24723–24730.
74. Zhao, H. (2007). Scaffold selection and scaffold hopping in lead generation: a medicinal chemistry perspective. Drug Discovery Today, 12, 149–155.
75. Wang, L., Deng, Y., Wu, Y., Kim, B., LeBard, D. N., Wandschneider, D., et al. (2017). Accurate modeling of scaffold hopping transformations in drug discovery. Journal of Chem- ical Theory and Computation, 13,42–54.
76. Wu, D., Zheng, X., Liu, R., Li, Z., Jiang, Z., Zhou, Q., et al. (2022). Free energy perturbation (FEP)-guided scaffold hopping. Acta Pharmaceutica Sinica B, 12, 1351–1362.
77. Jespers, W., Esguerra, M., Åqvist, J., & Gutiérrez-de-Terán, H. (2019). QligFEP: An auto­mated workow for small molecule free energy calculations in Q. Journal of Cheminformatics, 11, 26.
78. Pennington, L. D., Aquila, B. M., Choi, Y., Valiulin, R. A., & Muegge, I. (2020). Positional analogue scanning: An effective strategy for multiparameter optimization in drug design. Journal of Medicinal Chemistry, 63, 8956–8976.
79. Wade, A. D., Rizzi, A., Wang, Y., & Huggins, D. J. (2019). Computational uorine scanning using free-energy perturbation. Journal of Chemical Information and Modeling, 59, 2776–2784.
80. Pérez-Benito, L., Casajuana-Martin, N., Jiménez-Rosés, M., van Vlijmen, H., & Tresadern, G. (2019). Predicting activity cliffs with free-energy perturbation. and Computation, 15, 1884–1895.
81. Hu, Y., & Muegge, I. (2022). In silico positional analogue scanning with Amber GPU-TI. Journal of Chemical Information and Modeling, 62, 4448–4459.
82. De Vivo, M., Masetti, M., Bottegoni, G., & Cavalli, A. (2016). Role of molecular dynamics and related methods in drug discovery. Journal of Medicinal Chemistry, 59, 4035–4061.
83. Jiang, W., Hodoscek, M., & Roux, B. (2009). Computation of absolute hydration and binding free energy with free energy perturbation distributed replica-exchange molecular dynamics. Journal of Chemical Theory and Computation, 5, 2583–2588.
84. Jiang, W., & Roux, B. (2010). Free energy perturbation Hamiltonian replica-exchange molec­ular dynamics (FEP/H-REMD) for absolute ligand binding free energy calculations. Journal of Cchemical Theory and Computation, 6, 2559–2565.
85. Wang, L., Berne, B., & Friesner, R. A. (2012). On achieving high accuracy and reliability in the calculation of relative protein–ligand binding afnities. National Academy of Sciences of
the United States of America, 109, 1937–1942.
Journal of Chemical Theory
296 D. Antunes et al.
86. Raman, E. P., Paul, T. J., Hayes, R. L., & Brooks, C. L. I. (2020). Automated, accurate, and scalable relative protein–ligand binding free-energy calculations using lambda dynamics. Journal of Chemical Theory and Computation, 16, 7895–7914.
87. Wan, S., Bhati, A. P., & Coveney, P. V. (2023). Comparison of equilibrium and nonequilibrium approaches for relative binding free energy predictions. Journal of Chemical Theory and Computation, 19, 7846–7860.
88. Hummer, G. (2007). Nonequilibrium methods for equilibrium free energy calculations. In C. Chipot & A. Pohorille (Eds.), Free energy calculations: Theory and applications in chemistry and biology (pp. 171–198). Springer Berlin Heidelberg.
89. Crooks, G. E. (1998). Nonequilibrium measurements of free energy differences for micro­scopically reversible Markovian systems. Journal of Statistical Physics, 90, 1481–1487.
90. Crooks, G. E. (1999). Entropy production uctuation theorem and the nonequilibrium work relation for free energy differences. Physical Review E, 60, 2721–2726.
91. Gapsys, V., Pérez-Benito, L., Aldeghi, M., Seeliger, D., van Vlijmen, H., Tresadern, G., et al. (2020). Large scale relative protein ligand binding afnities using non-equilibrium alchemy. Chemical Science, 11, 1140–1152.
92. Gapsys, V., Hahn, D. F., Tresadern, G., Mobley, D. L., Rampp, M., & de Groot, B. L. (2022). Pre-exascale computing of protein–ligand binding free energies with open source software for drug design. Journal of Chemical Information and Modeling, 62, 1172–1177.
93. Kong, X., & Brooks, C. L., III. (1996). λ-dynamics: A new approach to free energy calcula­tions. The Journal of Chemical Physics, 105, 2414–2423.
94. Knight, J. L., & Brooks, C. L., III. (2009). λ-Dynamics free energy simulation methods. Journal of Computational Chemistry, 30, 1692–1700.
95. Banba, S., & Brooks, C. L., III. (2000). Free energy screening of small ligands binding to an articial protein cavity. The Journal of Chemical Physics, 113, 3423–3433.
96. Guo, Z., & Brooks, C. L. (1998). Rapid screening of binding afnities: Application of the
λ-dynamics method to a trypsin-inhibitor system. Journal of the American Chemical Society, 120, 1920–1921.
97. Knight, J. L., & Brooks, C. L., III. (2011). Applying efcient implicit nongeometric con­straints in alchemical free energy simulations. Journal of Computational Chemistry, 32, 3423–3432.
98. Robo, M. T., Hayes, R. L., Ding, X., Pulawski, B., & Vilseck, J. Z. (2023). Fast free energy estimates from λ-dynamics with bias-updated Gibbs sampling. Nature Communications,
, 8515.
14
99. Knight, J. L., & Brooks, C. L. I. (2011). Multisite λ dynamics for simulated structure–activity relationship studies. Journal of Chemical Theory and Computation, 7, 2728–2739.
100. Armacost, K. A., Goh, G. B., & Brooks, C. L. I. (2015). Biasing potential replica exchange multisite λ-dynamics for efcient free energy calculations. Journal of Chemical Theory and Computation, 11, 1267–1277.
101. Ding, X., Vilseck, J. Z., Hayes, R. L., & Brooks, C. L. I. (2017). Gibbs Sampler-based
λ-dynamics and Rao–Blackwell estimator for alchemical free energy calculation. Journal of Chemical Theory and Computation, 13, 2501–2510.
102. Vilseck, J. Z., Ding, X., Hayes, R. L., & Brooks, C. L. I. (2021). Generalizing the discrete Gibbs Sampler-based λ-dynamics approach for multisite sampling of many ligands. Journal of Chemical Theory and Computation, 17, 3895–3907.
103. Russell, S. J., & Norvig, P. (2010). Articial intelligence a modern approach.
104. Rufa, D. A., Bruce Macdonald, H. E., Fass, J., Wieder, M., Grinaway, P. B., Roitberg, A. E., et al. (2020). Towards chemical accuracy for alchemical free energy calculations with hybrid physics-based machine learning/molecular mechanics potentials. bioRxiv, 2020.07.29.227959.
105. Rizzi, A., Carloni, P., & Parrinello, M. (2021). Targeted free energy perturbation revisited: Accurate free energies from mapped reference potentials. Journal of Physical Chemistry Letters, 12, 9449–9454.
10 Free Energy Perturbation and Free-Energy Calculations Applied to Drug Design 297
106. Wirnsberger, P., Ballard, A. J., Papamakarios, G., Abercrombie, S., Racanière, S., Pritzel, A., et al. (2020). Targeted free energy estimation via learned mappings. The Journal of Chemical Physics, 153, 144112.
107. Knight, J. L., Leswing, K., Bos, P. H., & Wang, L. (2021). Impacting drug discovery projects with large-scale enumerations, machine learning strategies, and free-energy predictions. In Free energy methods in drug discovery: Current state and future directions (pp. 205–226). American Chemical Society.
108. Deringer, V. L., Bartók, A. P., Bernstein, N., Wilkins, D. M., Ceriotti, M., & Csányi, G. (2021). Gaussian process regression for materials and molecules. Chemical Reviews, 121, 10073–10141.
109. Heikamp, K., & Bajorath, J. (2014). Support vector machines for drug discovery. Expert Opinion on Drug Discovery, 9,93–104.
110. Cai, C., Wang, S., Xu, Y., Zhang, W., Tang, K., Ouyang, Q., et al. (2020). Transfer learning for drug discovery. Journal of Medicinal Chemistry, 63, 8683–8694.
111. Lim, J., Ryu, S., Kim, J. W., & Kim, W. Y. (2018). Molecular generative model based on conditional variational autoencoder for de novo molecular design. Journal of Cheminformatics, 10, 31.
112. Konze, K. D., Bos, P. H., Dahlgren, M. K., Leswing, K., Tubert-Brohman, I., Bortolato, A., et al. (2019). Reaction-based enumeration, active learning, and free energy calculations to rapidly explore synthetically tractable chemical space and optimize potency of cyclin­dependent kinase 2 inhibitors. Journal of Chemical Information and Modeling, 59, 3782–3793.
113. Willow, S. Y., Kang, L., & Minh, D. D. (2023). Learned mappings for targeted free energy perturbation between peptide conformations. arXiv. preprint arXiv:230614010.
Chapter 11
Ultra-Large-Scale Virtual Screening
Ina Pöhner, Toni Sivula, and Antti Poso
Abstract Recently, make-on-demand chemical libraries and spaces began to grow
into the billion scale, resulting in a surge of ultra-large ligand- and receptor-based virtual screening (VS) approaches. This chapter introduces state-of-the-art screening compound resources with up to 290 trillion compounds and provides examples of ultra-large VS campaigns performed between 2019 and 2024. We empha size key considerations for the planning stage of your own ultra-large VS and the choice of the computing environment. Some of the covered highlights include 2D similarity searches in nonenumerated spaces, 3D shape screening, giga-scale brute-force docking, and screening acceleration strategies either leveraging machine learning and deep learning or relying on the reagent-space of recent make-on-demand combinatorial libraries. To scrutinize whether and when bigger is better,we discuss the critical challenge of hit triage on the ultra-large scale and conclude with future perspectives for VS in the face of continuing library growth.

1 Background

The early-stage hit identication stage of drug discovery efforts frequently features in silico methods, which are typically considered cheaper and faster than their in vitro equivalents [1]. In silico screening enables the timely high-throughput evaluation of large numbers of potential candidate molecules. This computational prioritization of initial virtual hits is commonly termed virtual screening (VS). Traditionally, VS campaigns have either been ligand-centric or receptor-based [2].
I. Pöhner () · A. Poso School of Pharmacy, University of Eastern Finland, Kuopio, Finland e-mail: ina.pohner@uef.
T. Sivula School of Pharmacy, University of Eastern Finland, Kuopio, Finland
CSC IT Center for Science Ltd., Life Science Center Keilaniemi, Espoo, Finland
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2024 V. G. Maltarollo (ed.), Computer-Aided and Machine Learning-Driven Drug Design, Computer-Aided Drug Discovery and Design 3,
https://doi.org/10.1007/978-3-031-76718-0_11
299
300 I. Pöhner et al.
Ligand-based VS uses structural information and physicochemical properties of known active compounds and screens a compound library for virtual hits that share some of the activesfeatures. For example, classical quantitative structure–activity/ property relationships (QSAR/QSPR) and ngerprint-based searches rely solely on 2D structure and topology, which allows for their exceptionally high throughput. However, especially in data-sparse scenarios, such models have limitations and often generalize poorly [1, 3, 4]. Approaches that utilize the 3D structural information of the small molecules, such as shape-matching algorithms or pharmacophore searches, are often superior at capturing the intricacies of biological action in our 3D world [3, 4]. They are, however, generally slow er than their 2D-only counterparts [4].
When starting from known active scaffolds, ligand-based approaches are attrac­tive for performing scaffold hopping [5]. However, their inbuilt dependence on known actives (or ligands with the desired properties) is also their biggest drawback: In prospective screens, where such prior knowledge does not exist, ligand-based VS cannot be applied.
If a molecular target, for instance, an enzyme, can be dened, receptor-based VS approaches can be used (instead)even when no prior knowledge about possible binding partners exists. The most common technique to discover virtual hits for a target receptor is molecular docking. Docking rst ts ligand 3D structures from large databases into the receptor binding pocket. It then assigns a docking score, aiming to quantify the formed interactions and rank the ligands [1, 2].
It is a ge neral notion that larger compound library sizes increase the probability of nding a close-to-optimal match with the desired properties during a VS campaign [3, 6, 7]. Thus, unsurprisingly, in step with a growth trend observed for screening compound libraries, ligand- and receptor-based ultra-large VS approaches have seamlessly evolved from their smaller-scale counterparts.
Before we dive deeper into the evolution of ultra-large VS, let us dene what the ultra-large scaleactually is. The term itself is fuzzy and used in the literature for screens of varying sizes, to date typi cally ranging from around 10
7
to 1010processed small molecules. Herein, we dene our lower limit for considering a library or VS approach as ultra-largearbitrarily at 50 million compounds.
Ultra-large libraries themselves are not a new phenomenon: Efforts to enumerate much of the hypothetically possible chemical space gave rise to virtual ultra-large libraries already several years back (see, e.g., references [3, 8, 9]. for a comprehen­sive overview). However, the organic synth esis of a virtual hit compound from such a library was often time-consuming and costly, potentially hampering or preventing the in vitro validation of the VS predictions. A tight project timeframe and budget or the lack of medicinal chemistry facilities wer e therefore arguments in favor of smaller libraries with commercially available compounds [1].
The recent popularity of ultra-large VS is owed to advances in robust parallel organic synthesis and the emergence of ultra-large make-on-demand libraries. These combinatorial libraries rely on predened chemical reactions to link large collections of off-the-shelf building blocks. As a result, numerous make-on-demand compounds with a high probability of synthesis success can be ordered and readily purchasable in-stock libraries have also grown in size. For instance, the popular screening
11 Ultra-Large-Scale Virtual Screening 301
resource ZINC grew from 20 million commercially available compounds in 2012 to more than 37 billion by 2022 [1012]. Likewise, the make-on-demand library Enamine REAL started off on the million scale and to date contains around 48 billion compounds [13, 14 ].
Previously, their high throughput was among the strongest arguments in favor of VS methods. However, as library sizes hit the billion scale, VS throughput became a limiting factor. For instance, the fastest docking methods can typically screen about 30–60 compounds per minute per computing core [3]. Consequently, a docking study of one million compounds on a workstation computer with 8 CPUs takes around 1.5–3 days. Processing 1 billion co mpounds with the same resources would require 4–8 years of nonstop docking for a single target. To tackle ultra-large screening projects, one thus needs to either massively scale up on computational resources or nd strategies to avoid the bulk of the time- and resource-consuming computation steps.
In many cases, the basic operation of individual VS approaches is not altered when they are applied on the ultra-large scale. Therefore, this chapter will assume a basic familiarity of the reader with different VS methods and focus on challenges faced specically on the ultra-large scale and the emerging solutions. For more information on the individual VS methods, such as QSAR, pharmacophores, and docking, the reader is, for examples, referred to Chaps. 6 and 7 of this book.
The chapter will introduce ultra-large screen ing databases and recent examples of ultra-large-scale ligand- and receptor-based screening campaigns with an emphasis on the past 5 years (2019–2024). As most ultra-large-scale screening campaigns utilize high-performance computing (HPC) or cloud computing resources, we will discuss some characteristics of such computing environments, important consider­ations for their use, and tools designed to cope with the growing numbers of screening compounds. As we enter an era where ultra-large screening campaigns will likely become more prominent and strive to govern even large r chemical spaces, we are only beginning to comprehend the associated pitfalls of Big Datain drug discovery. This chapter will also present some arguments in favor of and against ultra-large VS, ag so far unmet needs, and provide a perspective of what lies ahead for in silico screen ing campaigns in the future.

2 Ultra-Large Screening Libraries and Chemical Spaces

As indicated above, ultra-large VS has little to do with technological advances in screening methodology. Rather, it represents a development in response to the emergence of billion-scale make-on-demand chemical libraries. Thus, before diving deeper into the screening process and specic examples, we will look at the culprits that force us to rethink our approaches to VS.
We can distinguish two major types of ultra-large libraries: enumerated chemical libraries, where each compo und is stored explicitly, and nonenumerated spaces [9, 15]. Enumeration of all possible combinations of building blocks is deemed
302 I. Pöhner et al.
inefcient when working with billion-scal e combinatorial libraries. Therefore, nonenumerated spaces rely on storing the building blocks and all chemical reactions that can be used to interconnect them instead of the fully elaborated possible virtual products [15].
One famous example of a freely accessible ultra-large compound library is ZINC. ZINC was initially created as a compound research database focused on docking and offered downloadable precomputed 3D conformers for many of its compounds. In its 2012 release, ZINC already contained around 20 million purchasable compounds from different vendors [11]. By 2015, the library had become an ultra-large resource with around 120 million purchasable drug-like compounds [16]. ZINC15 offered enhanced programmatic access and downloadable ligand 3D coordinates in various formats, such as sdf, mol2, and pdbqt. It is this version of ZINC that would be prominent in many of the recent ultra-large receptor-based VS works (see Sect. 5.1) [1719]. Meanwhile, ZINC continued to grow into the giga-scale: ZINC20, owing to the addition of compounds from make-on-demand suppliers, reached a total size of roughly 1.4 billion compounds, of which around 1.3 billion were purchasable [20]. The latest version ZINC22 added molecules from additional make-on-demand chemical spaces, increasing its tota l size to around 37 billion [12]. Currently, over 4 billion compo unds can be downloaded from ZINC in 3D ready-to-dock formats.
Enamine is one of the vendors that allowed ZI NC to grow into the billion scale. Their REAL (REadily AccessibLe) make-on-demand compounds can be typically obtained with a success rate of at least 80% within less than 1 month [13
15]. Enamine maintains a large collection of over 100,000 in-stock building blocks
and relies on more than 167 robust validated synthesis protocols [13, 14]. REAL compounds represent another popular compound resource of recent ultra-large screening campaigns [2123]. To date (Mar ch 2024), the REA L database contains about 6.75 billion enumerated compounds. Of the several available subsets, some are likewise ultra-large: The Enamine REAL lead-like set comprises 3.93 billion com­pounds, and Enamine also offers a library of nearly 158 million natural product-like compounds. Enamines by far largest collection of compounds is an example of a nonenumerated space: the REAL Space currently contains over 48 billion virtual products, stored primarily as building blocks and connecting chemical reactions (although an enumerated version is available from Enamine upon request [13]).
Another vendor that facilitated ZINCs tremendous growth is WuXi AppTec [24, 25]. Their billion-scale make-on-demand library, GalaXi, likewise relies on predened reaction schemes and a large in-stock building block collection. Cur­rently, the enumerated virtual Gal aXi library contains 16 billion compounds [25].
While working with enumerated ultra-large libraries can be achieved by scaling up conventional ligand- or receptor-based screening approaches, nonenumerated spaces require specialized technologies to search within their building blocks and the connecting reactions to enumerate relev ant virtual products on the y. Approaches for working with nonenumerated spaces will be discussed in detail later in the text (see, e.g., Sects. 4 and 6.2). Sufce it to say that nonenumerated spaces like the Enamine REAL Space or the nonenumerated variant of GalaXi, the GalaXi Space, can be searched with tools like the chemical space navigation
11 Ultra-Large-Scale Virtual Screening 303
Table 11.1 Summary of discussed publicly available/searchable enumerated ultra-large screening libraries and nonenumerated chemical spaces
Library/Space Approx. number of compounds File format(s)
Enumerated
CHIPMUNK 95 million sdf (2D) Enamine REAL
database ER lead-like 3.93 billion CXSMILES ER natural product-
like GalaXi database 16 billion SMILES SAVI 1.75 billion SMILES, sdf (2D) ZINC22 37 billion (2D), 4.5 billion (3D) SMILES, mol2, sdf,
ZINC20 1.4 billion, 1.3 billion purchasable SMILES, mol2, sdf,
ZINC15 750 million (purchasable, 2D), 230 million
Nonenumerated
CHEMriya 12 billion – Enamine REAL
space eXplore 4.9 trillion – GalaXi space 16 billion – KnowledgeSpace 290 trillion
For enumerated libraries, available formats for compound download are reported. Corresponding URLs can be found in Table 11.9 in the Appendix. ER: Enamine REAL; Billion = 10 lion = 10
12
; db2: format for use with the docking tool DOCK
6.75 billion CXSMILES, sdf (2D)
157.7 million CXSMILES
pdbqt, db2
pdbqt, db2
(purchasable, 3D)
48 billion
SMILES, mol2, sdf, pdbqt, db2
9
; tril-
platform inniSee [26]. Some vendors do not offer (part of) thei r enumerated collection for download but instead rely on its integration in chemical space search tools. With inniSee and related approaches, additional billion-scale synthesis-on­demand spaces become searchable, for example, eMoleculeseXplore with 4.9 trillion virtual products or Otavas 12 billion on-demand space CHEMriya [2729].
In light of the gigantic sizes of nonenumerated chemical spaces, the reader may wonder if the ability to search multiple spaces is even required. Interestingly, extensive efforts to compare chemical spaces have revealed surprisingly limited overlap between the different chemical spaces [15, 28].
Libraries and spaces discussed this far constituted commercially available com­pounds and make-on-demand vendor collections (see Table 11.1 for a summary). While the direct purchase of virtual hits can be an attractive option, other efforts cater more to the needs of those who wish to perform the compound synthesis themselves: While still relying on commercially available building blocks, reactions are derived from literature to ensure the synthetic feasibility of the virtual products.
For example, NCIs 1.75 billion enumerated compound library SAVI (Synthet­ically Accessible Virtual Inventory) relies on building blocks from Enamine and
304 I. Pöhner et al.
integrates expert knowledge into the denition of virtual synthesis rules [30]. Com­pounds contained within SAVI can be ranked by their synthetic accessibility, with the most synthesizable class currently encompassing 1.09 billion compounds [30]. Literature-derived chemical reactions and commercially available building blocks are also used to dene KnowledgeSpace, a virtual chemical space created by BioSolveIT, the company behind inniSee. KnowledgeSpace currently com­prises 10
14
virtual molecules and can be searched with inniSee [15, 26, 31].
Many of the recently published ultra-large-scale VS efforts discussed in this Chapter stem from academic context and utilize the mentioned public and commer­cial libraries and spaces. However, ultra-large chemical spaces are also a very prominent sight in industrial drug development: Over the past years, many pharma companies have developed their own proprietary ultra-large chemical spaces, for instance, BICLAIM of Boehringer-Ingelheim and Mercks MASSIV (10 products), or GlaxoSmithKlin es GSK XXL (10
26
virtual products) [9, 15]. There-
20
virtual
fore, tools and protocols to work with and screen ultra-large chemical spaces are of great interest to academic and industrial drug discovery alike.
All libraries and spaces mentioned so far mainly comprise drug-like small molecules. Compounds targeting protein–protein interactions (PPIs) typically exhibit a different property prole that may be underrepresented in libraries focused on drug-likeness. This motivated the author s of CHIPMUNK (CHemically feasible In silico Public Molecular UNiverse Knowledge base) to propose a specialized ultralarge library: They used commercially available building blocks and selected virtual reactions together with targeted property ltering to generate a library of 95 million synthetically accessible compounds with PPI modulator properties [32].
Table 11.1 summarizes the discussed chemical libraries and spaces, and Table 11.9 in the Appendix lists the corresponding URLs to obtain the different compound collections.
3 How Ultra-Large Chemical Libraries Alter VS
Workows
After familiarizing ourselves with the state-of-the-art ultra-large libraries and spaces, we may ponder how working on the ultra-large scale affects VS workows in practice. Before we discuss specic examples of ultra-large ligand-and receptor­based screens, we rst aim to establish important key considerations for ultra-large­scale screening workows and discuss what sets them apart from smaller-scale approaches.