Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5364_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Foreword
- •Acknowledgments
- •Contents
- •1.1 Structure-Based Drug Discovery (SBDD)
- •1.2 Ligand-Based Drug Design (LBDD)
- •1.3 Echoes from the Past, Visions from the Future
- •References
- •1 Introduction
- •2.2 Second Step: Data Curation
- •2.4 Fourth Step: Updating and Maintenance
- •2 Databases and Curation
- •8 Perspectives
- •9 Conclusion
- •References
- •1 Introduction
- •2.1 Making and Matching Protein Models
- •2.2 Simulating Protein Movements
- •2.3 Analyzing Changes in Protein Shape
- •3 Pharmacogenomics in Drug Development
- •4 Case Studies of Genomics-Based Drug Design
- •References
- •1 Historical Background
- •1.1 Timeline
- •2 Methodology Overview
- •2.1 Neural Networks
- •2.1.1 Perceptron
- •2.1.2 Multilayer Neural Networks
- •2.1.3 Types of Neural Networks
- •Feedforward
- •Recurrent Neural Networks
- •LSTM
- •2.2 Deep Learning
- •3 Using Machine Learning
- •3.2 Data Collection
- •3.3 Data Preprocessing
- •3.4 Model Selection
- •3.5 Model Training
- •3.6 Validation
- •3.7 Tuning
- •3.8 Prediction
- •4 Limitations
- •4.1 Bias
- •4.3 Interpretability
- •4.4 Computational Cost
- •4.5 Data Dependency
- •4.6 Robustness
- •5 Applications in Drug Discovery
- •5.2 Lead Discovery
- •5.3 Preclinical and Clinical Development
- •6 Resources and Tools
- •7 Challenges and Perspectives
- •7.1 Future Trends
- •9 Conclusions
- •References
- •1 Historical Background
- •1.1 Applications in Drug Discovery
- •2 Validations and Controls
- •2.1 Internal Validation
- •2.2 External Validation
- •2.3 Relative Cluster Validation
- •3 Challenges and Perspectives
- •4 Conclusions
- •References
- •1 Historical Background
- •2 OECD Principles
- •2.1 A Defined Endpoint
- •2.2 An Unambiguous Algorithm
- •2.5 A Mechanistic Interpretation, if Possible
- •3 Software and Tools
- •4 Validations and Controls
- •4.1 Internal and External Validation
- •4.1.1 Regression Metrics
- •4.2 Applicability Domain
- •4.3 Randomization Tests
- •5 Interpretation
- •6 Practical Advice During QSAR Modeling
- •7 Application
- •8 Challenges and Perspectives
- •References
- •1 Molecular Docking
- •2 Advances in Scoring Functions and Search Algorithms
- •2.2 Critical Characteristics of Search Algorithms
- •2.3 Docking Programs and Scoring Functions
- •3 Calculations Performed During Docking Simulations
- •4 Essential Components for a Good Docking Program
- •5 Limitations of the Docking Technique
- •6 Validation of Docking Results
- •7 Inappropriate Use of Validation Methods in Docking
- •9 Use of Machine Learning in Molecular Docking
- •11 Challenges
- •12 Conclusions
- •References
- •3 System Preparation for MD Simulations
- •3.1 Solvation and Microensemble
- •3.2 Force Fields: General Concept and Relevant Choices
- •3.3 The Concept of Replicas and Timescale
- •4.1.2 Protein Root Mean Square Fluctuation (RMSF)
- •4.1.4 Protein Secondary Structure Analysis
- •4.1.5 Principal component Analysis (PCA)
- •4.1.6 Markov State Modelling
- •4.1.7 Distance Calculations
- •4.1.8 Angle and Plane Calculations
- •4.2.2 Distances and Ligand-Induced Geometry Rearrangements
- •4 Molecular Dynamics Analysis
- •4.1 Protein Perspective
- •4.1.1 Protein Root Mean Square Deviation (RMSD)
- •4.3 Ligand Perspective
- •4.3.1 Ligand Properties
- •4.3.2 Ligand Root Mean Square Deviation
- •4.3.3 Ligand Root Mean Square Fluctuation
- •4.3.4 Angles and Dihedrals
- •5.1 Protein Structure Prediction and Preparation
- •5.2 Molecular Docking
- •6 Concluding Remarks and Outlook
- •Glossary
- •References
- •1 Introduction
- •2.1 MDeNM
- •2.2 Collective Molecular Dynamics (coMD)
- •2.3 ClustENM and ClustENMD
- •3 Ensemble Docking
- •References
- •1 Introduction
- •1.1 Advantages, Disadvantages, Innovations, and Challenges
- •1.2 Recent Advances in Accessible FEP Software Tools
- •1.3 Applications of FEP in Industry and Consortiums
- •2 Expanding the Potential of FEP Calculations
- •2.1 Validating Binding Poses
- •2.2 Dealing with Solvent
- •2.3 FEP and Allostery
- •2.4 FEP and Covalent Ligands
- •2.5 Applications of FEP in Scaffold Hopping
- •2.6 Positional Analogue Scanning
- •2.7 Combinations and Alternative Approaches
- •3 Machine Learning for FEP
- •3.4 Implications for ML in FEP Calculations
- •4 Final Considerations
- •5 First Steps to FEP Simulations
- •References
- •1 Background
- •2 Ultra-Large Screening Libraries and Chemical Spaces
- •3.1 Implications of Dataset Size
- •4 Ligands on the Ultra-Large Scale
- •4.1 Ultra-Large 2D Similarity Searches
- •7 Challenges and Future Perspectives
- •7.1 Hit Triage: An Old Problem on a New Dimension
- •8 Conclusions
- •Appendix
- •References
- •1 Introduction
- •2 Enzymatic Activity Evaluations
- •3 Cytotoxicity Evaluation and Cell Viability
- •4 Antiviral Assays in Experimental Validation
- •6 In Vivo Evaluation of Compounds
- •7 Conclusions
- •References
- •1 Introduction
- •3.1 Data Collection
- •3.2 Data Preprocessing
- •3.4 Model Choice
- •3.5 Model Training
- •3.6 Model Assessment
- •3.7 External Validation
- •3.8 Implementation and Availability
- •3.9 Continuous Update
- •5 Conclusions and Perspectives
- •References
- •1 Experimental Approaches to Obtain Protein Structure
- •1.1 X-Ray Crystallography
- •1.2 Nuclear Magnetic Resonance
- •1.3 Cryo-EM
- •1.4 Hybrid Methods
- •2 Modeling Approaches to Obtain Protein Structure
- •2.1 Homology Modeling
- •2.2 Ab Initio Modeling
- •2.3 New Approaches
- •3 Conformational Diversity of Proteins
- •3.1 Characterization of Protein Conformational States
- •3.2 Experimental Methods to Study Protein Dynamics and Conformations
- •3.4 Molecular Dynamics Simulation
- •3.5 Sampling Strategies
- •4 Remarks and Perspectives
- •References
- •1 Introduction
- •2 Structure-Based Drug Design of HIV Protease Inhibitors
- •2.1 HIV-1 Protease as a Therapeutic Target
- •2.2.1 Saquinavir
- •2.2.2 Indinavir
- •2.3.1 Lopinavir
- •2.3.2 Darunavir
- •6 Conclusions
- •References
- •4 Experimental Methods to Analyze NR Activity
- •4.2 Coregulator-Recruitment
- •5 Concluding Remarks and Outlook
- •References

294 D. Antunes et al.
49. Thompson, J., Walters, W. P., Feng, J. A., Pabon, N. A., Xu, H., Maser, M., et al. (2022).
Optimizing active learning for free energy calculations. Artificial Intelligence in the Life
Sciences, 2, 100050.
50. Majellaro, M., Jespers, W., Crespo, A., Núñez, M. J., Novio, S., Azuaje, J., et al. (2021).
3,4-dihydropyrimidin-2(1H)-ones as antagonists of the human A2B adenosine receptor: Optimization, structure–activity relationship studies, and enantiospecific recognition. Journal of
Medicinal Chemistry, 64, 458–480.
51. Shan, Y., Mysore, V. P., Leffler, A. E., Kim, E. T., Sagawa, S., & Shaw, D. E. (2022). How
does a small molecule bind at a cryptic binding site? PLOS Computational Biology, 18,
e1009817.
52. Kimura, S. R., Hu, H. P., Ruvinsky, A. M., Sherman, W., & Favia, A. D. (2017). Deciphering
cryptic binding sites on proteins by mixed-solvent molecular dynamics. Journal of Chemical
Information and Modeling, 57, 1388–1401.
53. Goodford, P. J. (1985). A computational procedure for determining energetically favorable
binding sites on biologically important macromolecules. Journal of Medicinal Chemistry, 28,
849–857.
54. Alvarez-Garcia, D., & Barril, X. (2014). Molecular simulations with solvent competition
quantify water displaceability and provide accurate interaction maps of protein binding sites.
Journal of Medicinal Chemistry, 57, 8530–8539.
55. Alvarez-Garcia, D., Schmidtke, P., Cubero, E., & Barril, X. (2022). Extracting atomic
contributions to binding free energy using molecular dynamics simulations with mixed
solvents (MDmix). Current Drug Discovery Technologies, 19,62–68.
56. Deng, Y., & Roux, B. (2008). Computation of binding free energy with molecular dynamics
and grand canonical Monte Carlo simulations. The Journal of Chemical Physics, 128.
57. Ross, G. A., Bodnarchuk, M. S., & Essex, J. W. (2015). Water sites, networks, and free
energies with grand canonical Monte Carlo. Journal of the American Chemical Society, 137,
14930–14943.
58. Michel, J., Tirado-Rives, J., & Jorgensen, W. L. (2009). Energetics of displacing water
molecules from protein binding sites: Consequences for ligand optimization. Journal of the
American Chemical Society, 131, 15403–15411.
59. Ben-Shalom, I. Y., Lin, Z., Radak, B. K., Lin, C., Sherman, W., & Gilson, M. K. (2020).
Accounting for the central role of interfacial water in protein–ligand binding free energy
calculations. Journal of Chemical Theory and Computation, 16, 7883–7894.
60. Bruce Macdonald, H. E., Cave-Ayland, C., Ross, G. A., & Essex, J. W. (2018). Ligand
binding free energies with adaptive water networks: Two-dimensional grand canonical
alchemical perturbations. Journal of Chemical Theory and Computation, 14, 6586–6597.
61. Nussinov, R., & Tsai, C.-J. (2012). The different ways through which specificity works in
orthosteric and allosteric drugs. Current Pharmaceutical Design, 18, 1311.
62. Liang, L., Liu, H., Xing, G., Deng, C., Hua, Y., Gu, R., et al. (2022). Accurate calculation of
absolute free energy of binding for SHP2 allosteric inhibitors using free energy perturbation.
Physical Chemistry Chemical Physics, 24, 9904–9920.
63. Zheng, H., Alter, S., & Qu, C.-K. (2009). SHP-2 tyrosine phosphatase in human diseases.
International Journal of Clinical and Experimental Medicine, 2, 17.
64. Torrente, E., Fodale, V., Ciammaichella, A., Ferrigno, F., Ontoria, J. M., Ponzi, S., et al.
(2023). Discovery of a novel series of imidazopyrazine derivatives as potent SHP2 allosteric
inhibitors. ACS Medicinal Chemistry Letters, 14, 156–162.
65. Xiaoli, A., Yuzhen, N., Qiong, Y., Yang, L., Yao, X., & Bing, Z. (2022). Investigating the
dynamic binding behavior of PMX53 cooperating with allosteric antagonist NDT9513727 to
C5a anaphylatoxin chemotactic receptor 1 through Gaussian accelerated molecular dynamics
and free-energy perturbation simulations. ACS Chemical Neuroscience, 13, 3502–3511.
66. Fu, H., Chen, H., Cai, W., Shao, X., & Chipot, C. (2021). BFEE2: Automated, streamlined,
and accurate absolute binding free-energy calculations. Journal of Chemical Information and
Modeling, 61, 2116–2123.

10 Free Energy Perturbation and Free-Energy Calculations Applied to Drug Design 295
67. Backus, K. M., Correia, B. E., Lum, K. M., Forli, S., Horning, B. D., González-Páez, G. E.,
et al. (2016). Proteome-wide covalent ligand discovery in native biological systems. Nature,
534, 570–574.
68. Sutanto, F., Konstantinidou, M., & Dömling, A. (2020). Covalent inhibitors: A rational
approach to drug discovery. RSC Medicinal Chemistry, 11, 876–884.
69. Chatterjee, P., Botello-Smith, W. M., Zhang, H., Qian, L., Alsamarah, A., Kent, D., et al.
(2017). Can relative binding free energy predict selectivity of reversible covalent inhibitors?
Journal of the American Chemical Society, 139, 17945–17952.
70. Awoonor-Williams, E., Walsh, A. G., & Rowley, C. N. (2017). Modeling covalent-modifier
drugs. Biochimica et Biophysica Acta (BBA) - Proteins and Proteomics, 1865, 1664–1675.
71. Bonatto, V., Shamim, A., Rocho, F. D. R., Leitão, A., Luque, F. J., Lameira, J., et al. (2021).
Predicting the relative binding affinity for reversible covalent inhibitors by free energy
perturbation calculations. Journal of Chemical Information and Modeling, 61, 4733–4744.
72. Zhang, H., Jiang, W., Chatterjee, P., & Luo, Y. (2019). Ranking reversible covalent drugs:
From free energy perturbation to fragment docking. Journal of Chemical Information and
Modeling, 59, 2093–2102.
73. Lameira, J., Bonatto, V., Cianni, L., dos Reis, R. F., Leitão, A., & Montanari, C. A. (2019).
Predicting the affinity of halogenated reversible covalent inhibitors through relative binding
free energy. Physical Chemistry Chemical Physics, 21, 24723–24730.
74. Zhao, H. (2007). Scaffold selection and scaffold hopping in lead generation: a medicinal
chemistry perspective. Drug Discovery Today, 12, 149–155.
75. Wang, L., Deng, Y., Wu, Y., Kim, B., LeBard, D. N., Wandschneider, D., et al. (2017).
Accurate modeling of scaffold hopping transformations in drug discovery. Journal of Chem-
ical Theory and Computation, 13,42–54.
76. Wu, D., Zheng, X., Liu, R., Li, Z., Jiang, Z., Zhou, Q., et al. (2022). Free energy perturbation
(FEP)-guided scaffold hopping. Acta Pharmaceutica Sinica B, 12, 1351–1362.
77. Jespers, W., Esguerra, M., Åqvist, J., & Gutiérrez-de-Terán, H. (2019). QligFEP: An automated workflow for small molecule free energy calculations in Q. Journal of
Cheminformatics, 11, 26.
78. Pennington, L. D., Aquila, B. M., Choi, Y., Valiulin, R. A., & Muegge, I. (2020). Positional
analogue scanning: An effective strategy for multiparameter optimization in drug design.
Journal of Medicinal Chemistry, 63, 8956–8976.
79. Wade, A. D., Rizzi, A., Wang, Y., & Huggins, D. J. (2019). Computational fluorine scanning
using free-energy perturbation. Journal of Chemical Information and Modeling, 59,
2776–2784.
80. Pérez-Benito, L., Casajuana-Martin, N., Jiménez-Rosés, M., van Vlijmen, H., & Tresadern,
G. (2019). Predicting activity cliffs with free-energy perturbation.
and Computation, 15, 1884–1895.
81. Hu, Y., & Muegge, I. (2022). In silico positional analogue scanning with Amber GPU-TI.
Journal of Chemical Information and Modeling, 62, 4448–4459.
82. De Vivo, M., Masetti, M., Bottegoni, G., & Cavalli, A. (2016). Role of molecular dynamics
and related methods in drug discovery. Journal of Medicinal Chemistry, 59, 4035–4061.
83. Jiang, W., Hodoscek, M., & Roux, B. (2009). Computation of absolute hydration and binding
free energy with free energy perturbation distributed replica-exchange molecular dynamics.
Journal of Chemical Theory and Computation, 5, 2583–2588.
84. Jiang, W., & Roux, B. (2010). Free energy perturbation Hamiltonian replica-exchange molecular dynamics (FEP/H-REMD) for absolute ligand binding free energy calculations. Journal of
Cchemical Theory and Computation, 6, 2559–2565.
85. Wang, L., Berne, B., & Friesner, R. A. (2012). On achieving high accuracy and reliability in
the calculation of relative protein–ligand binding affinities. National Academy of Sciences of
the United States of America, 109, 1937–1942.
Journal of Chemical Theory

296 D. Antunes et al.
86. Raman, E. P., Paul, T. J., Hayes, R. L., & Brooks, C. L. I. (2020). Automated, accurate, and
scalable relative protein–ligand binding free-energy calculations using lambda dynamics.
Journal of Chemical Theory and Computation, 16, 7895–7914.
87. Wan, S., Bhati, A. P., & Coveney, P. V. (2023). Comparison of equilibrium and
nonequilibrium approaches for relative binding free energy predictions. Journal of Chemical
Theory and Computation, 19, 7846–7860.
88. Hummer, G. (2007). Nonequilibrium methods for equilibrium free energy calculations. In
C. Chipot & A. Pohorille (Eds.), Free energy calculations: Theory and applications in
chemistry and biology (pp. 171–198). Springer Berlin Heidelberg.
89. Crooks, G. E. (1998). Nonequilibrium measurements of free energy differences for microscopically reversible Markovian systems. Journal of Statistical Physics, 90, 1481–1487.
90. Crooks, G. E. (1999). Entropy production fluctuation theorem and the nonequilibrium work
relation for free energy differences. Physical Review E, 60, 2721–2726.
91. Gapsys, V., Pérez-Benito, L., Aldeghi, M., Seeliger, D., van Vlijmen, H., Tresadern, G., et al.
(2020). Large scale relative protein ligand binding affinities using non-equilibrium alchemy.
Chemical Science, 11, 1140–1152.
92. Gapsys, V., Hahn, D. F., Tresadern, G., Mobley, D. L., Rampp, M., & de Groot, B. L. (2022).
Pre-exascale computing of protein–ligand binding free energies with open source software for
drug design. Journal of Chemical Information and Modeling, 62, 1172–1177.
93. Kong, X., & Brooks, C. L., III. (1996). λ-dynamics: A new approach to free energy calculations. The Journal of Chemical Physics, 105, 2414–2423.
94. Knight, J. L., & Brooks, C. L., III. (2009). λ-Dynamics free energy simulation methods.
Journal of Computational Chemistry, 30, 1692–1700.
95. Banba, S., & Brooks, C. L., III. (2000). Free energy screening of small ligands binding to an
artificial protein cavity. The Journal of Chemical Physics, 113, 3423–3433.
96. Guo, Z., & Brooks, C. L. (1998). Rapid screening of binding affinities: Application of the
λ-dynamics method to a trypsin-inhibitor system. Journal of the American Chemical Society,
120, 1920–1921.
97. Knight, J. L., & Brooks, C. L., III. (2011). Applying efficient implicit nongeometric constraints in alchemical free energy simulations. Journal of Computational Chemistry, 32,
3423–3432.
98. Robo, M. T., Hayes, R. L., Ding, X., Pulawski, B., & Vilseck, J. Z. (2023). Fast free energy
estimates from λ-dynamics with bias-updated Gibbs sampling. Nature Communications,
, 8515.
14
99. Knight, J. L., & Brooks, C. L. I. (2011). Multisite λ dynamics for simulated structure–activity
relationship studies. Journal of Chemical Theory and Computation, 7, 2728–2739.
100. Armacost, K. A., Goh, G. B., & Brooks, C. L. I. (2015). Biasing potential replica exchange
multisite λ-dynamics for efficient free energy calculations. Journal of Chemical Theory and
Computation, 11, 1267–1277.
101. Ding, X., Vilseck, J. Z., Hayes, R. L., & Brooks, C. L. I. (2017). Gibbs Sampler-based
λ-dynamics and Rao–Blackwell estimator for alchemical free energy calculation. Journal of
Chemical Theory and Computation, 13, 2501–2510.
102. Vilseck, J. Z., Ding, X., Hayes, R. L., & Brooks, C. L. I. (2021). Generalizing the discrete
Gibbs Sampler-based λ-dynamics approach for multisite sampling of many ligands. Journal of
Chemical Theory and Computation, 17, 3895–3907.
103. Russell, S. J., & Norvig, P. (2010). Artificial intelligence a modern approach.
104. Rufa, D. A., Bruce Macdonald, H. E., Fass, J., Wieder, M., Grinaway, P. B., Roitberg, A. E.,
et al. (2020). Towards chemical accuracy for alchemical free energy calculations with hybrid
physics-based machine learning/molecular mechanics potentials. bioRxiv, 2020.07.29.227959.
105. Rizzi, A., Carloni, P., & Parrinello, M. (2021). Targeted free energy perturbation revisited:
Accurate free energies from mapped reference potentials. Journal of Physical Chemistry
Letters, 12, 9449–9454.

10 Free Energy Perturbation and Free-Energy Calculations Applied to Drug Design 297
106. Wirnsberger, P., Ballard, A. J., Papamakarios, G., Abercrombie, S., Racanière, S., Pritzel, A.,
et al. (2020). Targeted free energy estimation via learned mappings. The Journal of Chemical
Physics, 153, 144112.
107. Knight, J. L., Leswing, K., Bos, P. H., & Wang, L. (2021). Impacting drug discovery projects
with large-scale enumerations, machine learning strategies, and free-energy predictions. In
Free energy methods in drug discovery: Current state and future directions (pp. 205–226).
American Chemical Society.
108. Deringer, V. L., Bartók, A. P., Bernstein, N., Wilkins, D. M., Ceriotti, M., & Csányi,
G. (2021). Gaussian process regression for materials and molecules. Chemical Reviews, 121,
10073–10141.
109. Heikamp, K., & Bajorath, J. (2014). Support vector machines for drug discovery. Expert
Opinion on Drug Discovery, 9,93–104.
110. Cai, C., Wang, S., Xu, Y., Zhang, W., Tang, K., Ouyang, Q., et al. (2020). Transfer learning
for drug discovery. Journal of Medicinal Chemistry, 63, 8683–8694.
111. Lim, J., Ryu, S., Kim, J. W., & Kim, W. Y. (2018). Molecular generative model based on
conditional variational autoencoder for de novo molecular design. Journal of
Cheminformatics, 10, 31.
112. Konze, K. D., Bos, P. H., Dahlgren, M. K., Leswing, K., Tubert-Brohman, I., Bortolato, A.,
et al. (2019). Reaction-based enumeration, active learning, and free energy calculations to
rapidly explore synthetically tractable chemical space and optimize potency of cyclindependent kinase 2 inhibitors. Journal of Chemical Information and Modeling, 59,
3782–3793.
113. Willow, S. Y., Kang, L., & Minh, D. D. (2023). Learned mappings for targeted free energy
perturbation between peptide conformations. arXiv. preprint arXiv:230614010.

Chapter 11
Ultra-Large-Scale Virtual Screening
Ina Pöhner, Toni Sivula, and Antti Poso
Abstract Recently, make-on-demand chemical libraries and spaces began to grow
into the billion scale, resulting in a surge of ultra-large ligand- and receptor-based
virtual screening (VS) approaches. This chapter introduces state-of-the-art screening
compound resources with up to 290 trillion compounds and provides examples of
ultra-large VS campaigns performed between 2019 and 2024. We empha size key
considerations for the planning stage of your own ultra-large VS and the choice of
the computing environment. Some of the covered highlights include 2D similarity
searches in nonenumerated spaces, 3D shape screening, giga-scale brute-force
docking, and screening acceleration strategies either leveraging machine learning
and deep learning or relying on the reagent-space of recent make-on-demand
combinatorial libraries. To scrutinize whether and when “bigger is better,” we
discuss the critical challenge of hit triage on the ultra-large scale and conclude
with future perspectives for VS in the face of continuing library growth.
1 Background
The early-stage hit identification stage of drug discovery efforts frequently features
in silico methods, which are typically considered cheaper and faster than their
in vitro equivalents [1]. In silico screening enables the timely high-throughput
evaluation of large numbers of potential candidate molecules. This computational
prioritization of initial virtual hits is commonly termed virtual screening (VS).
Traditionally, VS campaigns have either been ligand-centric or receptor-based [2].
I. Pöhner (✉) · A. Poso
School of Pharmacy, University of Eastern Finland, Kuopio, Finland
e-mail: ina.pohner@uef.fi
T. Sivula
School of Pharmacy, University of Eastern Finland, Kuopio, Finland
CSC — IT Center for Science Ltd., Life Science Center Keilaniemi, Espoo, Finland
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2024
V. G. Maltarollo (ed.), Computer-Aided and Machine Learning-Driven Drug
Design, Computer-Aided Drug Discovery and Design 3,
https://doi.org/10.1007/978-3-031-76718-0_11
299

300 I. Pöhner et al.
Ligand-based VS uses structural information and physicochemical properties of
known active compounds and screens a compound library for virtual hits that share
some of the actives’ features. For example, classical quantitative structure–activity/
property relationships (QSAR/QSPR) and fingerprint-based searches rely solely on
2D structure and topology, which allows for their exceptionally high throughput.
However, especially in data-sparse scenarios, such models have limitations and often
generalize poorly [1, 3, 4]. Approaches that utilize the 3D structural information of
the small molecules, such as shape-matching algorithms or pharmacophore searches,
are often superior at capturing the intricacies of biological action in our 3D world
[3, 4]. They are, however, generally slow er than their 2D-only counterparts [4].
When starting from known active scaffolds, ligand-based approaches are attractive for performing scaffold hopping [5]. However, their inbuilt dependence on
known actives (or ligands with the desired properties) is also their biggest drawback:
In prospective screens, where such prior knowledge does not exist, ligand-based VS
cannot be applied.
If a molecular target, for instance, an enzyme, can be defined, receptor-based VS
approaches can be used (instead)—even when no prior knowledge about possible
binding partners exists. The most common technique to discover virtual hits for a
target receptor is molecular docking. Docking first fits ligand 3D structures from
large databases into the receptor binding pocket. It then assigns a docking score,
aiming to quantify the formed interactions and rank the ligands [1, 2].
It is a ge neral notion that larger compound library sizes increase the probability of
finding a close-to-optimal match with the desired properties during a VS campaign
[3, 6, 7]. Thus, unsurprisingly, in step with a growth trend observed for screening
compound libraries, ligand- and receptor-based ultra-large VS approaches have
seamlessly evolved from their smaller-scale counterparts.
Before we dive deeper into the evolution of ultra-large VS, let us define what the
“ultra-large scale” actually is. The term itself is fuzzy and used in the literature for
screens of varying sizes, to date typi cally ranging from around 10
7
to 1010processed
small molecules. Herein, we define our lower limit for considering a library or VS
approach as “ultra-large” arbitrarily at 50 million compounds.
Ultra-large libraries themselves are not a new phenomenon: Efforts to enumerate
much of the hypothetically possible chemical space gave rise to virtual ultra-large
libraries already several years back (see, e.g., references [3, 8, 9]. for a comprehensive overview). However, the organic synth esis of a virtual hit compound from such
a library was often time-consuming and costly, potentially hampering or preventing
the in vitro validation of the VS predictions. A tight project timeframe and budget or
the lack of medicinal chemistry facilities wer e therefore arguments in favor of
smaller libraries with commercially available compounds [1].
The recent popularity of ultra-large VS is owed to advances in robust parallel
organic synthesis and the emergence of ultra-large make-on-demand libraries. These
combinatorial libraries rely on predefined chemical reactions to link large collections
of off-the-shelf building blocks. As a result, numerous make-on-demand compounds
with a high probability of synthesis success can be ordered and readily purchasable
in-stock libraries have also grown in size. For instance, the popular screening

11 Ultra-Large-Scale Virtual Screening 301
resource ZINC grew from 20 million commercially available compounds in 2012 to
more than 37 billion by 2022 [10–12]. Likewise, the make-on-demand library
Enamine REAL started off on the million scale and to date contains around 48 billion
compounds [13, 14 ].
Previously, their high throughput was among the strongest arguments in favor of
VS methods. However, as library sizes hit the billion scale, VS throughput became a
limiting factor. For instance, the fastest docking methods can typically screen about
30–60 compounds per minute per computing core [3]. Consequently, a docking
study of one million compounds on a workstation computer with 8 CPUs takes
around 1.5–3 days. Processing 1 billion co mpounds with the same resources would
require 4–8 years of nonstop docking for a single target. To tackle ultra-large
screening projects, one thus needs to either massively scale up on computational
resources or find strategies to avoid the bulk of the time- and resource-consuming
computation steps.
In many cases, the basic operation of individual VS approaches is not altered
when they are applied on the ultra-large scale. Therefore, this chapter will assume a
basic familiarity of the reader with different VS methods and focus on challenges
faced specifically on the ultra-large scale and the emerging solutions. For more
information on the individual VS methods, such as QSAR, pharmacophores, and
docking, the reader is, for examples, referred to Chaps. 6 and 7 of this book.
The chapter will introduce ultra-large screen ing databases and recent examples of
ultra-large-scale ligand- and receptor-based screening campaigns with an emphasis
on the past 5 years (2019–2024). As most ultra-large-scale screening campaigns
utilize high-performance computing (HPC) or cloud computing resources, we will
discuss some characteristics of such computing environments, important considerations for their use, and tools designed to cope with the growing numbers of
screening compounds. As we enter an era where ultra-large screening campaigns
will likely become more prominent and strive to govern even large r chemical spaces,
we are only beginning to comprehend the associated pitfalls of “Big Data” in drug
discovery. This chapter will also present some arguments in favor of and against
ultra-large VS, flag so far unmet needs, and provide a perspective of what lies ahead
for in silico screen ing campaigns in the future.
2 Ultra-Large Screening Libraries and Chemical Spaces
As indicated above, ultra-large VS has little to do with technological advances in
screening methodology. Rather, it represents a development in response to the
emergence of billion-scale make-on-demand chemical libraries. Thus, before diving
deeper into the screening process and specific examples, we will look at the culprits
that force us to rethink our approaches to VS.
We can distinguish two major types of ultra-large libraries: enumerated chemical
libraries, where each compo und is stored explicitly, and nonenumerated spaces
[9, 15]. Enumeration of all possible combinations of building blocks is deemed

302 I. Pöhner et al.
inefficient when working with billion-scal e combinatorial libraries. Therefore,
nonenumerated spaces rely on storing the building blocks and all chemical reactions
that can be used to interconnect them instead of the fully elaborated possible virtual
products [15].
One famous example of a freely accessible ultra-large compound library is ZINC.
ZINC was initially created as a compound research database focused on docking and
offered downloadable precomputed 3D conformers for many of its compounds. In its
2012 release, ZINC already contained around 20 million purchasable compounds
from different vendors [11]. By 2015, the library had become an ultra-large resource
with around 120 million purchasable drug-like compounds [16]. ZINC15 offered
enhanced programmatic access and downloadable ligand 3D coordinates in various
formats, such as sdf, mol2, and pdbqt. It is this version of ZINC that would be
prominent in many of the recent ultra-large receptor-based VS works (see Sect. 5.1)
[17–19]. Meanwhile, ZINC continued to grow into the giga-scale: ZINC20, owing to
the addition of compounds from make-on-demand suppliers, reached a total size of
roughly 1.4 billion compounds, of which around 1.3 billion were purchasable
[20]. The latest version ZINC22 added molecules from additional make-on-demand
chemical spaces, increasing its tota l size to around 37 billion [12]. Currently, over
4 billion compo unds can be downloaded from ZINC in 3D ready-to-dock formats.
Enamine is one of the vendors that allowed ZI NC to grow into the billion scale.
Their REAL (REadily AccessibLe) make-on-demand compounds can be typically
obtained with a success rate of at least 80% within less than 1 month [13–
15]. Enamine maintains a large collection of over 100,000 in-stock building blocks
and relies on more than 167 robust validated synthesis protocols [13, 14]. REAL
compounds represent another popular compound resource of recent ultra-large
screening campaigns [21–23]. To date (Mar ch 2024), the REA L database contains
about 6.75 billion enumerated compounds. Of the several available subsets, some are
likewise ultra-large: The Enamine REAL lead-like set comprises 3.93 billion compounds, and Enamine also offers a library of nearly 158 million natural product-like
compounds. Enamine’s by far largest collection of compounds is an example of a
nonenumerated space: the REAL Space currently contains over 48 billion virtual
products, stored primarily as building blocks and connecting chemical reactions
(although an enumerated version is available from Enamine upon request [13]).
Another vendor that facilitated ZINC’s tremendous growth is WuXi AppTec
[24, 25]. Their billion-scale make-on-demand library, GalaXi, likewise relies on
predefined reaction schemes and a large in-stock building block collection. Currently, the enumerated virtual Gal aXi library contains 16 billion compounds [25].
While working with enumerated ultra-large libraries can be achieved by scaling
up conventional ligand- or receptor-based screening approaches, nonenumerated
spaces require specialized technologies to search within their building blocks and
the connecting reactions to enumerate relev ant virtual products on the fly.
Approaches for working with nonenumerated spaces will be discussed in detail
later in the text (see, e.g., Sects. 4 and 6.2). Suffice it to say that nonenumerated
spaces like the Enamine REAL Space or the nonenumerated variant of GalaXi, the
GalaXi Space, can be searched with tools like the chemical space navigation

11 Ultra-Large-Scale Virtual Screening 303
Table 11.1 Summary of discussed publicly available/searchable enumerated ultra-large screening
libraries and nonenumerated chemical spaces
Library/Space Approx. number of compounds File format(s)
Enumerated
CHIPMUNK 95 million sdf (2D)
Enamine REAL
database
ER lead-like 3.93 billion CXSMILES
ER natural product-
like
GalaXi database 16 billion SMILES
SAVI 1.75 billion SMILES, sdf (2D)
ZINC22 37 billion (2D), 4.5 billion (3D) SMILES, mol2, sdf,
ZINC20 1.4 billion, 1.3 billion purchasable SMILES, mol2, sdf,
ZINC15 750 million (purchasable, 2D), 230 million
Nonenumerated
CHEMriya 12 billion –
Enamine REAL
space
eXplore 4.9 trillion –
GalaXi space 16 billion –
KnowledgeSpace 290 trillion –
For enumerated libraries, available formats for compound download are reported. Corresponding
URLs can be found in Table 11.9 in the Appendix. ER: Enamine REAL; Billion = 10
lion = 10
12
; db2: format for use with the docking tool DOCK
6.75 billion CXSMILES, sdf (2D)
157.7 million CXSMILES
pdbqt, db2
pdbqt, db2
(purchasable, 3D)
48 billion –
SMILES, mol2, sdf,
pdbqt, db2
9
; tril-
platform infiniSee [26]. Some vendors do not offer (part of) thei r enumerated
collection for download but instead rely on its integration in chemical space search
tools. With infiniSee and related approaches, additional billion-scale synthesis-ondemand spaces become searchable, for example, eMolecules’ eXplore with 4.9
trillion virtual products or Otava’s 12 billion on-demand space CHEMriya [27–29].
In light of the gigantic sizes of nonenumerated chemical spaces, the reader may
wonder if the ability to search multiple spaces is even required. Interestingly,
extensive efforts to compare chemical spaces have revealed surprisingly limited
overlap between the different chemical spaces [15, 28].
Libraries and spaces discussed this far constituted commercially available compounds and make-on-demand vendor collections (see Table 11.1 for a summary).
While the direct purchase of virtual hits can be an attractive option, other efforts cater
more to the needs of those who wish to perform the compound synthesis themselves:
While still relying on commercially available building blocks, reactions are derived
from literature to ensure the synthetic feasibility of the virtual products.
For example, NCI’s 1.75 billion enumerated compound library SAVI (Synthetically Accessible Virtual Inventory) relies on building blocks from Enamine and

304 I. Pöhner et al.
integrates expert knowledge into the definition of virtual synthesis rules [30]. Compounds contained within SAVI can be ranked by their synthetic accessibility, with
the most synthesizable class currently encompassing 1.09 billion compounds
[30]. Literature-derived chemical reactions and commercially available building
blocks are also used to define KnowledgeSpace, a virtual chemical space created
by BioSolveIT, the company behind infiniSee. KnowledgeSpace currently comprises 10
14
virtual molecules and can be searched with infiniSee [15, 26, 31].
Many of the recently published ultra-large-scale VS efforts discussed in this
Chapter stem from academic context and utilize the mentioned public and commercial libraries and spaces. However, ultra-large chemical spaces are also a very
prominent sight in industrial drug development: Over the past years, many pharma
companies have developed their own proprietary ultra-large chemical spaces, for
instance, BICLAIM of Boehringer-Ingelheim and Merck’s MASSIV (10
products), or GlaxoSmithKlin e’s GSK XXL (10
26
virtual products) [9, 15]. There-
20
virtual
fore, tools and protocols to work with and screen ultra-large chemical spaces are of
great interest to academic and industrial drug discovery alike.
All libraries and spaces mentioned so far mainly comprise drug-like small
molecules. Compounds targeting protein–protein interactions (PPIs) typically
exhibit a different property profile that may be underrepresented in libraries focused
on drug-likeness. This motivated the author s of CHIPMUNK (CHemically feasible
In silico Public Molecular UNiverse Knowledge base) to propose a specialized
ultralarge library: They used commercially available building blocks and selected
virtual reactions together with targeted property filtering to generate a library of
95 million synthetically accessible compounds with PPI modulator properties [32].
Table 11.1 summarizes the discussed chemical libraries and spaces, and
Table 11.9 in the Appendix lists the corresponding URLs to obtain the different
compound collections.
3 How Ultra-Large Chemical Libraries Alter VS
Workflows
After familiarizing ourselves with the state-of-the-art ultra-large libraries and spaces,
we may ponder how working on the ultra-large scale affects VS workflows in
practice. Before we discuss specific examples of ultra-large ligand-and receptorbased screens, we first aim to establish important key considerations for ultra-largescale screening workflows and discuss what sets them apart from smaller-scale
approaches.
Соседние файлы в папке Библиотека им академика М.И. Перельмана
