Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5364_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Foreword
- •Acknowledgments
- •Contents
- •1.1 Structure-Based Drug Discovery (SBDD)
- •1.2 Ligand-Based Drug Design (LBDD)
- •1.3 Echoes from the Past, Visions from the Future
- •References
- •1 Introduction
- •2.2 Second Step: Data Curation
- •2.4 Fourth Step: Updating and Maintenance
- •2 Databases and Curation
- •8 Perspectives
- •9 Conclusion
- •References
- •1 Introduction
- •2.1 Making and Matching Protein Models
- •2.2 Simulating Protein Movements
- •2.3 Analyzing Changes in Protein Shape
- •3 Pharmacogenomics in Drug Development
- •4 Case Studies of Genomics-Based Drug Design
- •References
- •1 Historical Background
- •1.1 Timeline
- •2 Methodology Overview
- •2.1 Neural Networks
- •2.1.1 Perceptron
- •2.1.2 Multilayer Neural Networks
- •2.1.3 Types of Neural Networks
- •Feedforward
- •Recurrent Neural Networks
- •LSTM
- •2.2 Deep Learning
- •3 Using Machine Learning
- •3.2 Data Collection
- •3.3 Data Preprocessing
- •3.4 Model Selection
- •3.5 Model Training
- •3.6 Validation
- •3.7 Tuning
- •3.8 Prediction
- •4 Limitations
- •4.1 Bias
- •4.3 Interpretability
- •4.4 Computational Cost
- •4.5 Data Dependency
- •4.6 Robustness
- •5 Applications in Drug Discovery
- •5.2 Lead Discovery
- •5.3 Preclinical and Clinical Development
- •6 Resources and Tools
- •7 Challenges and Perspectives
- •7.1 Future Trends
- •9 Conclusions
- •References
- •1 Historical Background
- •1.1 Applications in Drug Discovery
- •2 Validations and Controls
- •2.1 Internal Validation
- •2.2 External Validation
- •2.3 Relative Cluster Validation
- •3 Challenges and Perspectives
- •4 Conclusions
- •References
- •1 Historical Background
- •2 OECD Principles
- •2.1 A Defined Endpoint
- •2.2 An Unambiguous Algorithm
- •2.5 A Mechanistic Interpretation, if Possible
- •3 Software and Tools
- •4 Validations and Controls
- •4.1 Internal and External Validation
- •4.1.1 Regression Metrics
- •4.2 Applicability Domain
- •4.3 Randomization Tests
- •5 Interpretation
- •6 Practical Advice During QSAR Modeling
- •7 Application
- •8 Challenges and Perspectives
- •References
- •1 Molecular Docking
- •2 Advances in Scoring Functions and Search Algorithms
- •2.2 Critical Characteristics of Search Algorithms
- •2.3 Docking Programs and Scoring Functions
- •3 Calculations Performed During Docking Simulations
- •4 Essential Components for a Good Docking Program
- •5 Limitations of the Docking Technique
- •6 Validation of Docking Results
- •7 Inappropriate Use of Validation Methods in Docking
- •9 Use of Machine Learning in Molecular Docking
- •11 Challenges
- •12 Conclusions
- •References
- •3 System Preparation for MD Simulations
- •3.1 Solvation and Microensemble
- •3.2 Force Fields: General Concept and Relevant Choices
- •3.3 The Concept of Replicas and Timescale
- •4.1.2 Protein Root Mean Square Fluctuation (RMSF)
- •4.1.4 Protein Secondary Structure Analysis
- •4.1.5 Principal component Analysis (PCA)
- •4.1.6 Markov State Modelling
- •4.1.7 Distance Calculations
- •4.1.8 Angle and Plane Calculations
- •4.2.2 Distances and Ligand-Induced Geometry Rearrangements
- •4 Molecular Dynamics Analysis
- •4.1 Protein Perspective
- •4.1.1 Protein Root Mean Square Deviation (RMSD)
- •4.3 Ligand Perspective
- •4.3.1 Ligand Properties
- •4.3.2 Ligand Root Mean Square Deviation
- •4.3.3 Ligand Root Mean Square Fluctuation
- •4.3.4 Angles and Dihedrals
- •5.1 Protein Structure Prediction and Preparation
- •5.2 Molecular Docking
- •6 Concluding Remarks and Outlook
- •Glossary
- •References
- •1 Introduction
- •2.1 MDeNM
- •2.2 Collective Molecular Dynamics (coMD)
- •2.3 ClustENM and ClustENMD
- •3 Ensemble Docking
- •References
- •1 Introduction
- •1.1 Advantages, Disadvantages, Innovations, and Challenges
- •1.2 Recent Advances in Accessible FEP Software Tools
- •1.3 Applications of FEP in Industry and Consortiums
- •2 Expanding the Potential of FEP Calculations
- •2.1 Validating Binding Poses
- •2.2 Dealing with Solvent
- •2.3 FEP and Allostery
- •2.4 FEP and Covalent Ligands
- •2.5 Applications of FEP in Scaffold Hopping
- •2.6 Positional Analogue Scanning
- •2.7 Combinations and Alternative Approaches
- •3 Machine Learning for FEP
- •3.4 Implications for ML in FEP Calculations
- •4 Final Considerations
- •5 First Steps to FEP Simulations
- •References
- •1 Background
- •2 Ultra-Large Screening Libraries and Chemical Spaces
- •3.1 Implications of Dataset Size
- •4 Ligands on the Ultra-Large Scale
- •4.1 Ultra-Large 2D Similarity Searches
- •7 Challenges and Future Perspectives
- •7.1 Hit Triage: An Old Problem on a New Dimension
- •8 Conclusions
- •Appendix
- •References
- •1 Introduction
- •2 Enzymatic Activity Evaluations
- •3 Cytotoxicity Evaluation and Cell Viability
- •4 Antiviral Assays in Experimental Validation
- •6 In Vivo Evaluation of Compounds
- •7 Conclusions
- •References
- •1 Introduction
- •3.1 Data Collection
- •3.2 Data Preprocessing
- •3.4 Model Choice
- •3.5 Model Training
- •3.6 Model Assessment
- •3.7 External Validation
- •3.8 Implementation and Availability
- •3.9 Continuous Update
- •5 Conclusions and Perspectives
- •References
- •1 Experimental Approaches to Obtain Protein Structure
- •1.1 X-Ray Crystallography
- •1.2 Nuclear Magnetic Resonance
- •1.3 Cryo-EM
- •1.4 Hybrid Methods
- •2 Modeling Approaches to Obtain Protein Structure
- •2.1 Homology Modeling
- •2.2 Ab Initio Modeling
- •2.3 New Approaches
- •3 Conformational Diversity of Proteins
- •3.1 Characterization of Protein Conformational States
- •3.2 Experimental Methods to Study Protein Dynamics and Conformations
- •3.4 Molecular Dynamics Simulation
- •3.5 Sampling Strategies
- •4 Remarks and Perspectives
- •References
- •1 Introduction
- •2 Structure-Based Drug Design of HIV Protease Inhibitors
- •2.1 HIV-1 Protease as a Therapeutic Target
- •2.2.1 Saquinavir
- •2.2.2 Indinavir
- •2.3.1 Lopinavir
- •2.3.2 Darunavir
- •6 Conclusions
- •References
- •4 Experimental Methods to Analyze NR Activity
- •4.2 Coregulator-Recruitment
- •5 Concluding Remarks and Outlook
- •References

11 Ultra-Large-Scale Virtual Screening 335
Fig. 11.6 Summary of all discussed ultra-large VS campaigns. Studies are grouped by year and
color coded by screening approach. The y-axis indicates library sizes. In the inset on the right-hand
side, sizes of all discussed publicly available chemical libraries and spaces are displayed for
reference. Gray circles mark enumerated, white circles—nonenumerated libraries
reported hit rates (compare Table 11.8; typical hit finding campaigns achieve only
about 1% hit rate according to the literature [89]) should be taken with a grain of salt
and, given the human factor involved, not directly translated into an argument for
pursuing ultra-large-scale screening. Owing to this notion, the debate whether bigger
is better remains ongoing. Arguments can be made both in favor and against ultralarge VS [98]. We are currently only beginning to scratch the surface of ultra-largescale VS, and it is, in our and others’ opinion, premature to draw a conclusion
whether—and, importantly, when—bigger is actually better [1, 98, 99].
Acknowledgements I.P., T.S., and A.P. gratefully acknowledge the support of the Jane and Aatos
Erkko Foundation. We thank Laxman Yetukuri of CSC for his critical review and helpful feedback
on our discussion of HPC systems.

336 I. Pöhner et al.
Appendix
Table 11.9 List of URLs for discussed enumerated ultra-large screening libraries
Database Available from
Enamine REAL database https://enamine.net/compound-collections/real-com
Enamine REAL lead-like, natural
product-like compounds
GalaXi enumerated https://www.labnetwork.com/frontendapp/p/#!/
SAVI https://doi.org/10.35115/37n9-5738
ZINC22 https://cartblanche22.docking.org/
ZINC20 https://zinc20.docking.org/
ZINC15 https://zinc15.docking.org/
CHIPMUNK http://www.ewit.ccb.tu-dortmund.de/ag-koch/
An online version of this Table is also available from https://github.com/ipohner/ultralarge-VS,
where references and table contents will be periodically updated
Table 11.10 Summary of discussed open-source VS tools and their availability
Tool Reference(s) Available from
DeepDocking Gentile et al.
DeepDocking GUI Yaacoub
HASTEN Kalliokoski
Lean docking Berenger
Linear accelerated
docking
MolPAL Graff et al.
Schrödinger GPU similarity (incentive)
SpaceDock Sindt et al.
Thompson sampling Klarich et al.
V-SYNTHES Sadybekov
VirtualFlow Gorgulla
warpDOCK McDougal
An online version of this Table is also available from https://github.com/ipohner/ultralarge-VS,
where references and table contents will be periodically updated
[71]
et al. [76]
[79]
et al. [84]
Marin et al.
[85]
[77]
[90]
[35]
et al. [22]
et al. [7]
et al. [67]
pounds/real-database
https://enamine.net/compound-collections/real-com
pounds/real-database-subsets
library/virtual
https://cactus.nci.nih.gov/download/savi_download/
chipmunk/
https://github.com/jamesgleave/DD_protocol
https://github.com/jamesgleave/DeepDockingGUI
https://github.com/TuomoKalliokoski/HASTEN
Encoder: https://github.com/UnixJunkie/molenc Support Vector regressor: https://github.com/UnixJunkie/
linwrap
https://github.com/marinegor/Linear-accelerated-
docking
https://github.com/coleygroup/molpal
https://github.com/schrodinger/gpusimilarity
https://github.com/litfsindt/LIT-SpaceDock
https://github.com/PatWalters/TS
https://github.com/katritchlab/V-SYNTHES
https://github.com/VirtualFlow
https://github.com/BruningLab/warpDOCK

11 Ultra-Large-Scale Virtual Screening 337
Table 11.11 Summary of available screening libraries with pre-generated 3D conformers and
benchmarking datasets
Pre-generated 3D screening libraries — Description and reference
→ available from
Enamine REAL 1.4 billion compounds in PDBQT format by Gorgulla et al. [7]
→ via https://virtual-flow.org/real-library with login
ZINC15 1.5 billion compounds by Gorgulla et al. [5]
→ via https://virtual-flow.org/virtualflow-version-zinc15-library with login
Enamine REAL lead-like (2021) 3D conformers from 1.56 billion SMILES input in Schrödinger
Phase databases by Sivula et al. [23]
→ https://doi.org/10.23729/2de314bb-59af-452a-955c-c2ff0c5ea57f
Pre-generated 3D conformer tranches of ZINC20 as referenced by Bender et al. [24]
→ https://files.docking.org/3D/
Ultra-large benchmarking datasets — Description and reference
→ available from
1.56 billion SMILES + Glide-HTVS docking scores for SurA and GAK (csv) by Sivula et al. [23]
→ https://doi.org/10.23729/2170dc9c-4905-43c3-aeee-a574d360737f
1.4 billion SMILES + AutoDock-GPU docking scores and re-scoring results for 5 SARS-CoV-2
protein targets (parquet) as described in Rogers et al. [56]
→ https://doi.org/10.13139/OLCF/1783186
138 million SMILES + DOCK docking scores for dopamin D
receptor by Lyu et al. [17]
4
→ https://doi.org/10.6084/m9.figshare.7359401.v3
99 million SMILES + DOCK docking scores for AmpC by Lyu et al. [17]
→ https://doi.org/10.6084/m9.figshare.7359626.v2
138 million D
+ 99 million AmpC Glide docking scores by Yang et al. [81]
4
→ https://s3.amazonaws.com/content.schrodinger.com/Resources/paper_data_share.zip
All URLs reported below were last checked January 31st, 2024. An online version of this Table is
also available from https://github.com/ipohner/ultralarge-VS, where references and table contents
will be periodically updated
References
1. Stumpfe, D., & Bajorath, J. (2020). Current trends, overlooked issues, and unmet challenges in
virtual screening. Journal of Chemical Information and Modeling, 60, 4112–4115.
2. Carpenter, K. A., Cohen, D. S., Jarrell, J. T., & Huang, X. (2018). Deep learning and virtual
drug screening. Future Medicinal Chemistry, 10(21), 2557–2567.
3. Walters, W. P. (2019). Virtual chemical libraries. Journal of Medicinal Chemistry, 62(3),
1116–1124.
4. Grebner, C., Malmerberg, E., Shewmaker, A., Batista, J., Nicholls, A., & Sadowski, J. (2020).
Virtual screening in the cloud: How big is big enough? Journal of Chemical Information and
Modeling, 60(9), 4274–4282.
5. Gorgulla, C., Jayaraj, A., Fackeldey, K., & Arthanari, H. (2022). Emerging frontiers in virtual
drug discovery: From quantum mechanical methods to deep learning approaches. Current
Opinion in Chemical Biology, 69, 102156.
6. Fresnais, L., & Ballester, P. J. (2021). The impact of compound library size on the performance
of scoring functions for structure-based virtual screening. Briefings in Bioinformatics, 22(3),
1–10.
7. Gorgulla, C., Boeszoermenyi, A., Wang, Z.-F., Fischer, P. D., Coote, P. W., Das, K. M. P.,
Malets, Y. S., Radchenko, D. S., Moroz, Y. S., Scott, D. A., Fackeldey, K., Hoffmann, M.,

338 I. Pöhner et al.
Iavniuk, I., Wagner, G., & Arthanari, H. (2020). An open-source drug discovery platform
enables ultra-large virtual screens. Nature, 580, 663–668.
8. Hoffmann, T., & Gastreich, M. (2019). The next level in chemical space navigation: Going far
beyond enumerable compound libraries. Drug Discovery Today, 24(5), 1148–1156.
9. Warr, W. A., Nicklaus, M. C., Nicolaou, C. A., & Rarey, M. (2022). Exploration of ultralarge
compound collections for drug discovery. Journal of Chemical Information and Modeling,
62(9), 2021–2034.
10. ZINC15. https://zinc15.docking.org/. Online; Accessed 18 Jan 2024.
11. Irwin, J. J., Sterling, T., Mysinger, M. M., Bolstad, E. S., & Coleman, R. G. (2012). ZINC: A
free tool to discover chemistry for biology. Journal of Chemical Information and Modeling,
52(7), 1757–1768.
12. Tingle, B. I., Tang, K. G., Castanon, M., Gutierrez, J. J., Khurelbaatar, M., Dandarchuluun, C.,
Moroz, Y. S., & Irwin, J. J. (2023). ZINC-22—A free multi-billion-scale database of tangible
compounds for ligand discovery. Journal of Chemical Information and Modeling, 63(4),
1166–1176.
13. Enamine REAL. https://enamine.net/compound-collections/real-compounds, https://enamine.
13 Mar 2024.
14. Grygorenko, O. O., Radchenko, D. S., Dziuba, I., Chuprina, A., Gubina, K. E., & Moroz, Y. S.
(2020). Generating multibillion chemical space of readily accessible screening compounds.
iScience, 23(11), 101681.
15. Lessel, U., & Lemmen, C. (2019). Modeling the expansion of virtual screening libraries. ACS
Medicinal Chemistry Letters, 10, 1504–1510.
16. Sterling, T., & Irwin, J. J. (2015). ZINC 15—Ligand discovery for everyone. Journal of
Chemical Information and Modeling, 55(11), 2324–2337.
17. Lyu, J., Wang, S., Balius, T. E., Singh, I., Levit, A., Moroz, Y. S., O’Meara, M. J., Che, T.,
Algaa, E., Tolmachova, K., Tolmachev, A. A., Shoichet, B. K., Roth, B. L., & Irwin, J. J.
(2019). Ultra-large library docking for discovering new chemotypes. Nature, 566(7743),
224–229.
18. Stein, R. M., Kang, H. J., McCorvy, J. D., Glatfelter, G. C., Jones, A. J., Che, T., Slocum, S.,
Huang, X. P., Savych, O., Moroz, Y. S., Stauch, B., Johansson, L. C., Cherezov, V., Kenakin,
T., Irwin, J. J., Shoichet, B. K., Roth, B. L., & Dubocovich, M. L. (2020). Virtual discovery of
melatonin receptor ligands to modulate circadian rhythms. Nature, 579(7800), 609–614.
19. Alon, A., Lyu, J., Braz, J. M., Tummino, T. A., Craik, V., O’Meara, M. J., Webb, C. M.,
Radchenko, D. S., Moroz, Y. S., Huang, X.-P., Liu, Y., Roth, B. L., Irwin, J. J., Basbaum, A. I.,
Shoichet, B. K., & Kruse, A. C. (2021). Structures of the σ2 receptor enable docking for
bioactive ligand discovery. Nature, 600, 759–764.
20. Irwin, J. J., Tang, K. G., Young, J., Dandarchuluun, C., Wong, B. R., Khurelbaatar, M., Moroz,
Y. S., Mayfield, J., & Sayle, R. A. (2020). ZINC20— A free ultralarge-scale chemical database
for ligand discovery. Journal of Chemical Information and Modeling, 60(12), 6065–
21. Michino, M., Beautrait, A., Boyles, N. A., Nadupalli, A., Dementiev, A., Sun, S., Ginn, J., Baxt,
L., Suto, R., Bryk, R., Jerome, S. V., Huggins, D. J., & Vendome, J. (2023). Shape-based virtual
screening of a billion-compound library identifies mycobacterial lipoamide dehydrogenase
inhibitors. ACS Bio & Med Chem Au, 3(6), 507–515.
22. Sadybekov, A. A., Sadybekov, A. V., Liu, Y., Iliopoulos-Tsoutsouvas, C., Huang, X. P.,
Pickett, J., Houser, B., Patel, N., Tran, N. K., Tong, F., Zvonok, N., Jain, M. K., Savych, O.,
Radchenko, D. S., Nikas, S. P., Petasis, N. A., Moroz, Y. S., Roth, B. L., Makriyannis, A., &
Katritch, V. (2022). Synthon-based ligand discovery in virtual libraries of over 11 billion
compounds. Nature, 601(7893), 452–459.
23. Sivula, T., Yetukuri, L., Kalliokoski, T., Käsnänen, H., Poso, A., & Pöhner, I. (2023). Machine
learning-boosted docking enables the efficient structure based virtual screening of giga-scale
enumerated chemical libraries. Journal of Chemical Information and Modeling, 63(18),
5773–5783.
6073.

11 Ultra-Large-Scale Virtual Screening 339
24. Bender, B. J., Gahbauer, S., Luttens, A., Lyu, J., Webb, C. M., Stein, R. M., Fink, E. A., Balius,
T. E., Carlsson, J., Irwin, J. J., & Shoichet, B. K. (2021). A practical guide to large-scale
docking. Nature Protocols, 16, 4799–4832.
25. WuXi AppTec — galaXi. https://www.labnetwork.com/frontend-app/p/#!/library/virtual.
Online; Accessed 30 Jan 2024.
26. BioSolveIT infiniSee. https://www.biosolveit.de/infiniSee/. Online; Accessed 28 Feb 2024.
27. eMolecules eXplore. https://www.emolecules.com/explore. Online; Accessed 15 Mar 2024.
28. Neumann, A., Marrison, L., & Klein, R. (2023). Relevance of the trillion-sized chemical space
“eXplore” as a source for drug discovery. ACS Medicinal Chemistry Letters, 14, 466–472.
29. OTAVA Chemicals — CHEMriya On-Demand Chemical Space. https://www.otavachemicals.
com/products/chemriya. Online; Accessed 9 Feb 2024.
30. Patel, H., Ihlenfeldt, W.-D., Judson, P. N., Moroz, Y. S., Pevzner, Y., Peach, M. L., Delannée,
V., Tarasova, N. I., & Nicklaus, M. C. (2020). SAVI, in silico generation of billions of easily
synthesizable compounds through expert-system type rules. Scientific Data, 7, 384.
31. BioSolveIT KnowledgeSpace. https://www.biosolveit.de/chemical-spaces. Online; Accessed
22 Nov 2024.
32. Humbeck, L., Weigang, S., Schäfer, T., Mutzel, P., & Koch, O. (2018). CHIPMUNK: A virtual
synthesizable small-molecule library for medicinal chemistry, exploitable for protein–protein
interaction modulators. ChemMedChem, 13, 532–539.
33. How to solve a jigsaw puzzle with 1.56 billion pieces. CSC Blog Post. https://csc.fi/en/blog/
how-to-solve-a-jigsaw-puzzle-with-1-56-billion-pieces/. Online; Accessed 22 Nov 2024.
34. Gentile, F., Fernandez, M., Ban, F., Ton, A.-T., Mslati, H., Perez, C. F., Leblanc, E., Yaacoub,
J. C., Gleave, J., Stern, A., Wong, B., Jean, F., Strynadka, N., & Cherkasov, A. (2021).
Automated discovery of noncovalent inhibitors of sars-cov2 main protease by consensus deep
docking of 40 billion small molecules. Chemical Science, 12(48), 15960–15974.
35. Klarich, K., Goldman, B., Kramer, T., Riley, P., & Walters, W. P. (2024). Thompson sampling
— An efficient method for searching ultralarge synthesis on demand databases. Journal of
Chemical Information and Modeling, 64(4), 1158–1171.
36. Morris, G. M., Huey, R., Lindstrom, W., Sanner, M. F., Belew, R. K., Goodsell, D. S., & Olson,
A. J. (2009). AutoDock4 and AutoDockTools4: Automated docking with selective receptor
flexibility. Journal of Computational Chemistry, 30(16), 2785–2791.
37. Santos-Martins, D., Solis-Vasquez, L., Tillack, A. F., Sanner, M. F., Koch, A., & Forli,
S. (2021). Accelerating AutoDock4 with GPUs and gradient-based local search. Journal of
Chemical Theory and Computation, 17(2), 1060–1073.
38. Gentile, F., Yaacoub, J. C., Gleave, J., Fernandez, M., Ton, A.-T., Ban, F., Stern, A., &
Cherkasov, A. (2022). Artificial intelligence
libraries with deep docking. Nature Protocols, 17, 672–697.
39. Dalke, A. (2019). The chemfp project. Journal of Cheminformatics, 11, 76.
40. MolSoft ICM-Pro: Gigasearch. https://molsoft.com/giga-search.html. Online; Accessed
26 Feb 2024.
41. Schrödinger LiveDesign. https://newsite.schrodinger.com/platform/products/livedesign/.
Online; Accessed 26 Feb 2024.
42. OpenEye Orion: Molecules as a Service (MaaS). https://www.eyesopen.com/news/openeye-
orion-2020.2-update. Online; Accessed 26 Feb 2024.
43. Schrödinger GPU Similarity. https://github.com/schrodinger/gpusimilarity. Online; Accessed
26 Feb 2024.
44. NextMove Software Ltd. SmallWorld. https://www.nextmovesoftware.com/smallworld.html.
Online; Accessed 25 Mar 2024.
45. NextMove Software Ltd. Arthor. https://nextmovesoftware.com/arthor.html. Online; Accessed
25 Mar 2024.
46. Rarey, M., & Stahl, M. (2001). Similarity searching in large combinatorial chemistry spaces.
Journal of Computer-Aided Molecular Design, 15, 497–520.
–enabled virtual screening of ultra-large chemical

340 I. Pöhner et al.
47. Bellmann, L., Penner, P., & Rarey, M. (2019). Connected subgraph fingerprints: Representing
molecules using exhaustive subgraph enumeration. Journal of Chemical Information and
Modeling, 59, 4625–4635.
48. Bellmann, L., Penner, P., & Rarey, M. (2021). Topological similarity search in large combinatorial fragment spaces. Journal of Chemical Information and Modeling, 61, 238–251.
49. Schmidt, R., Klein, R., & Rarey, M. (2022). Maximum common substructure searching in
combinatorial make-on-demand compound spaces. Journal of Chemical Information and
Modeling, 62, 2133–2150.
50. Glaab, E., Manoharan, G. B., & Abankwa, D. (2021). Pharmacophore model for sars-cov-2
3clpro small-molecule inhibitors and in vitro experimental validation of computationally
screened inhibitors. Journal of Chemical Information and Modeling, 61, 4082–4096.
51. Brüschweiler, S., Fuchs, J. E., Bader, G., McConnell, D. B., Konrat, R., & Mayer, M. (2021). A
step toward NRF2-DNA interaction inhibitors by fragment-based NMR methods.
ChemMedChem, 16, 3576–3587.
52. MolSoft RIDE (Rapid Isostere Discovery Engine). https://molsoft.com/RIDE.html. Online;
Accessed 29 Feb 2024.
53. ZINC20 3D Tranches. https://files.docking.org/3D/. Online; Accessed 31 Jan 2024.
54. Gahbauer, S., Correy, G. J., Schuller, M., Ferla, M. P., Doruk, Y. U., Rachman, M., Wu, T.,
Diolaiti, M., Wang, S., Neitz, R. J., Fearon, D., Radchenko, D. S., Moroz, Y. S., Irwin, J. J.,
Renslo, A. R., Taylor, J. C., Gestwicki, J. E., von Delft, F., Ashworth, A., Ahel, I., Shoichet,
B. K., & Fraser, J. S. (2023). Iterative computational design and crystallographic screening
identifies potent inhibitors targeting the nsp3 macrodomain of sars-cov-2. Proceedings of the
National Academy of Sciences of the United States of America, 120(2), e2212931120.
55. Acharya, A., Agarwal, R., Baker, M. B., Baudry, J., Bhowmik, D., Boehm, S., Byler, K. G.,
Chen, S. Y., Coates, L., Cooper, C. J., Demerdash, O., Daidone, I., Eblen, J. D., Ellingson, S.,
Forli, S., Glaser, J., Gumbart, J. C., Gunnels, J., Hernandez, O., Irle, S., Kneller, D. W.,
Kovalevsky, A., Larkin, J., Lawrence, T. J., LeGrand, S., Liu, S.-H., Mitchell, J. C., Park, G.,
Parks, J. M., Pavlova, A., Petridis, L., Poole, D., Pouchard, L., Ramanathan, A., Rogers, D. M.,
Santos-Martins, D., Scheinberg, A., Sedova, A., Shen, Y., Smith, J. C., Smith, M. D., Soto, C.,
Tsaris, A., Thavappiragasam, M., Tillack, A. F., Vermaas, J. V., Vuong, V. Q., Yin, J., Yoo, S.,
Zahran, M., & Zanetti-Polzi, L. (2020). Supercomputer-based ensemble docking drug discovery
pipeline with application to Covid-19. Journal of Chemical Information and Modeling, 60(12),
5832–5852.
56. Rogers, D. M., Agarwal, R., Vermaas, J. V., Smith, M. D., Rajeshwar, R. T., Cooper, C.,
Sedova, A., Boehm, S., Baker, M., Glaser, J., & Smith, J. C. (2023). SARS-CoV2 billioncompound docking. Scientific Data, 10, 173.
57. Luttens, A., Gullberg, H., Abdurakhmanov, E., Vo, D. D., Akaberi, D., Talibov, V. O.,
Nekhotiaeva, N., Vangeel, L., De Jonghe, S., Jochmans, D., Krambrich, J., Tas, A., Lundgren,
B., Gravenfors, Y., Craig, A. J., Atilaw, Y., Sandström, A., Moodie, L. W. K., Lundkvist, A.,
van Hemert, M. J., Neyts, J., Lennerstrand, J., Kihlberg, J., Sandberg, K., Danielson, U. H., &
Carlsson, J. (2022). Ultralarge virtual screening identifies sars-cov-2 main protease inhibitors
with broad-spectrum activity against coronaviruses. Journal of the American Chemical Society,
144(7), 2905–2920.
58. Trott, O., & Olson, A. J. (2010). Autodock vina: Improving the speed and accuracy of docking
with a new scoring function, efficient optimization, and multithreading. Journal of Computa-
tional Chemistry, 31(2), 455–461.
59. Koes, D. R., Baumgartner, M. P., & Camacho, C. J. (2013). Lessons learned in empirical
scoring with smina from the CSAR 2011 benchmarking exercise. Journal of Chemical Infor-
mation and Modeling, 53(8), 1893–1904.
60. Korb, O., Stützle, T., & Exner, T. E. (2007). An ant colony optimization approach to flexible
protein-ligand docking.
61. Gorgulla, C., Çınaroğlu, S. S., Fischer, P. D., Fackeldey, K., Wagner, G., & Arthanari,
H. (2021). VirtualFlow ants-ultra-large virtual screenings with artificial intelligence driven
Swarm Intelligence, 1(2), 115–134.

11 Ultra-Large-Scale Virtual Screening 341
docking algorithm based on ant colony optimization. International Journal of Molecular
Sciences, 22(11), 5807.
62. Gorgulla, C., Das, K. M. P., Leigh, K. E., Cespugli, M., Fischer, P. D., Wang, Z.-F., Tesseyre,
G., Pandita, S., Shnapir, A., Calderaio, A., Gechev, M., Rose, A., Lewis, N., Hutcheson, C.,
Yaffe, E., Luxenburg, R., Herce, H. D., Durmaz, V., Halazonetis, T. D., Fackeldey, K., Patten,
J., Chuprina, A., Dziuba, I., Plekhova, A., Moroz, Y., Radchenko, D., Tarkhanova, O.,
Yavnyuk, I., Gruber, C., Yust, R., Payne, D., Näär, A. M., Namchuk, M. N., Davey, R. A.,
Wagner, G., Kinney, J., & Arthanari, H. (2021). A multi-pronged approach targeting SARSCoV-2 proteins using ultra-large virtual screening. iScience, 24(2), 102021.
63. Sadybekov, A. A., Brouillette, R. L., Marin, E., Sadybekov, A. V., Luginina, A., Gusach, A.,
Mishin, A., Besserer-Offroy, E., Jean-Michel, L., Borshchevskiy, V., Cherezov, V., Sarret, P.,
& Katrich, V. (2020). Structure-based virtual screening of ultra-large library yields potent
antagonists for a lipid GPCR. Biomolecules, 10(12), 1634.
64. Kaplan, A. L., Confair, D. N., Kim, K., Barros-Álvarez, X., Rodriguiz, R. M., Yang, Y.,
Kweon, O. S., Che, T., McCorvy, J. D., Kamber, D. N., Phelan, J. P., Martins, L. C., Pogorelov,
V. M., DiBerto, J. F., Slocum, S. T., Huang, X. P., Kumar, J. M., Robertson, M. J., Panova, O.,
Seven, A. B., Wetsel, A. Q., Wetsel, W. C., Irwin, J. J., Skiniotis, G., Shoichet, B. K., Roth,
B. L., & Ellman, J. A. (2022). Bespoke library docking for 5-HT
receptor agonists with
2A
antidepressant activity. Nature, 610(7932), 582–591.
65. Pöhner, I., & Sivula, T. (2023). Glide HTVS docking results of Enamine REAL lead-like library
(1.56 billion compounds) for targets SurA and GAK.
66. Rogers, D., Glaser, J., Agarwal, R., Vermaas, J., Smith, M., Parks, J., Cooper, C., Sedova, A.,
Boehm, S., Baker, M., & Smith, J. (2021). SARS-CoV2 docking dataset.
67. McDougal, D. P., Rajapaksha, H., Pederick, J. L., & Bruning, J. B. (2023). warpDOCK: Largescale virtual drug discovery using cloud infrastructure. ACS Omega, 8(32), 29143–29149.
68. Alhossary, A., Handoko, S. D., Mu, Y., & Kwoh, C.-K. (2015). Fast, accurate, and reliable
molecular docking with QuickVina 2. Bioinformatics, 31(13), 2214–2216.
69. Bonilla, P. A., Hoop, C. L., Stefanisko, K., Tarasov, S. G., Sinha, S., Nicklaus, M. C., &
Tarasova, N. I. (2023). Virtual screening of ultra-large chemical libraries identifies cellpermeable small-molecule inhibitors of a "non-druggable" target, STAT3 N-terminal domain.
Frontiers in Oncology, 13, 1144153.
70. Cherkasov, A., Ban, F., Li, Y., Fallahi, M., & Hammond, G. L. (2006). Progressive docking: A
hybrid QSAR/docking approach for accelerating in silico high throughput screening. Journal of
Medicinal Chemistry, 49, 7466–7478.
71. Gentile, F., Agrawal, V., Hsing, M., Ton, A. T., Ban, F., Norinder, U., Gleave, M. E., &
Cherkasov, A. (2020). Deep docking: A deep learning platform for augmentation of structure
based drug discovery. ACS Central Science, 6(6), 939–949.
72. Ton, A.-T., Gentile, F., Hsing, M., Ban, F., & Cherkasov, A. (2020). Rapid identification of
potential inhibitors of SARS-CoV-2 main protease by deep docking of 1.3 billion compounds.
Molecular Informatics, 39(8), 202000028.
73. Garland, O., Ton, A.-T., Moradi, S., Smith, J. R., Kovacic, S., Ng, K., Pandey, M., Ban, F., Lee,
J., Vuckovic, M., Worrall, L. J., Young, R. N., Pantophlet, R., Strynadka, N. C. J., &
Cherkasov, A. (2023). Large-scale virtual screening for the discovery of SARS-CoV-2
papain-like protease (PLpro) non-covalent inhibitors. Journal of Chemical Information and
Modeling, 63(7), 2158–2169.
74. Tang, M., Wen, C., Lin, J., Chen, H., & Ran, T. (2023). Discovery of novel A
R antagonists
2A
through deep learning-based virtual screening. Artificial Intelligence in the Life Sciences, 3,
100058.
75. Radaeva, M., Ho, C.-H., Xie, N., Zhang, S., Lee, J., Liu, L., Lallous, N., Cherkasov, A., &
Dong, X. (2022). Discovery of novel Lin28 inhibitors to suppress cancer cell stemness.
Cancers, 14(22), 5687.
76. Yaacoub, J. C., Gleave, J., Gentile, F., Stern, A., & Cherkasov, A. (2022). DD-GUI: A graphical
user interface for deep learning-accelerated virtual screening of large chemical libraries (Deep
Docking). Bioinformatics, 38(4), 1146–1148.

342 I. Pöhner et al.
77. Graff, D. E., Shakhnovich, E. I., & Coley, C. W. (2021). Accelerating high-throughput virtual
screening through molecular pool-based active learning. Chemical Science, 12(22), 7866–7881.
78. Yang, K., Swanson, K., Jin, W., Coley, C., Eiden, P., Gao, H., Guzman-Perez, A., Hopper, T.,
Kelley, B., Mathea, M., Palmer, A., Settels, V., Jaakkola, T., Jensen, K., & Barzilay, R. (2019).
Analyzing learned molecular representations for property prediction. Journal of Chemical
Information and Modeling, 59(8), 3370–3388.
79. Kalliokoski, T. (2021). Machine learning boosted docking (HASTEN): An open-source tool to
accelerate structure-based virtual screening campaigns. Molecular Informatics, 40(9), 2100089.
80. Friesner, R. A., Banks, J. L., Murphy, R. B., Halgren, T. A., Klicic, J. J., Mainz, D. T., Repasky,
M. P., Knoll, E. H., Shelley, M., Perry, J. K., Shaw, D. E., Francis, P., & Shenkin, P. S. (2004).
Glide: A new approach for rapid, accurate docking and scoring. 1. Method and assessment of
docking accuracy. Journal of Medicinal Chemistry, 47(7), 1739–1749.
81. Yang, Y., Yao, K., Repasky, M. P., Leswing, K., Abel, R., Shoichet, B. K., & Jerome, S. V.
(2021). Efficient exploration of chemical space with docking and deep learning. Journal of
Chemical Theory and Computation, 17(11), 7106–7119.
82. MolSoft ICM-Pro: Gigascreen. https://www.molsoft.com/GigaScreen.html. Online; Accessed
28 Jan 2024.
83. OpenEye Orion: GigaDock. https://www.eyesopen.com/orion/gigadock. Online; Accessed
28 Jan 2024.
84. Berenger, F., Kumar, A., Zhang, K. Y. J., & Yamanishi, Y. (2021). Lean-docking: Exploiting
ligands’ predicted docking scores to accelerate molecular docking. Journal of Chemical
Information and Modeling, 61, 2341–2352.
85. Marin, E., Kovaleva, M., Kadukova, M., Mustafin, K., Khorn, P., Rogachev, A., Mishin, A.,
Guskov, A., & Borshchevskiy, V. (2024). Regression-based active learning for accessible
acceleration of ultra-large library docking. Journal of Chemical Information and Modeling,
64(7), 2612–2623.
86. Kolb, P., Kipouros, C. B., Huang, D., & Caflisch, A. (2008). Structure-based tailoring of
compound libraries for high-throughput screening: Discovery of novel EphB4 kinase inhibitors.
Proteins, 73(1), 11–18.
87. Zhao, H., Dong, J., Lafleur, K., Nevado, C., & Caflisch, A. (2012). Discovery of a novel
chemotype of tyrosine kinase inhibitors by fragment-based docking and molecular dynamics.
ACS Medicinal Chemistry Letters, 3(10), 834–838.
88. Beroza, P., Crawford, J. J., Ganichkin, O., Gendelev, L., Harris, S. F., Klein, R., Miu, A.,
Steinbacher, S., Klingler, F. M., & Lemmen, C. (2022). Chemical space docking enables largescale structure-based virtual screening to discover ROCK1 kinase inhibitors. Nature Commu-
nications, 13(1), 6447.
89. Müller, J., Klein, R., Tarkhanova, O., Gryniukova, A., Borysko, P., Merkl, S., Ruf, M.,
Neumann, A., Gastreich, M., Moroz, Y. S., Klebe, G., & Glinca, S. (2022). Magnet for the
needle in haystack: "crystal structure first" fragment hits unlock active chemical matter using
targeted exploration of vast chemical spaces. Journal of Medicinal Chemistry, 65(23),
15663–15678.
90. Sindt, F., Seyller, A., Eguida, M., & Rognan, D. (2024). Protein structure-based organic
chemistry-driven ligand design from ultralarge chemical spaces. ACS Central Science, 10(3),
615–627.
91. Lyu, J., Irwin, J. J., & Shoichet, B. K. (2023). Modeling the expansion of virtual screening
libraries. Nature Chemical Biology, 19, 712–718.
92. Segler, M. H. S., Kogej, T., Tyrchan, C., & Waller, M. P. (2018). Generating focused molecule
libraries for drug discovery with recurrent neural networks. ACS Central Science, 4 (1),
120–131.
93. Jeon, W., & Kim, D. (2020). Autonomous molecule generation using reinforcement learning
and docking to develop potential novel inhibitors. Scientific Reports, 10(1), 22104.
94. Segler, M. H. S., Preuss, M., & Waller, M. P. (2018). Planning chemical syntheses with deep
neural networks and symbolic AI. Nature, 555, 604–610.

11 Ultra-Large-Scale Virtual Screening 343
95. Popov, K. I., Wellnitz, J., Maxfield, T., & Tropsha, A. (2024). HIt Discovery using docking
ENriched by GEnerative Modeling (HIDDEN GEM): A novel computational workflow for
accelerated virtual screening of ultra-large chemical libraries. Molecular Informatics, 1,
e202300207.
96. Sunkari, Y. K., Siripuram, V. K., Nguyen, T.-L., & Flajolet, M. (2022). High-power screening
(HPS) empowered by DNA-encoded libraries. Trends in Pharmacological Sciences, 43(1),
4–15.
97. Collie, G. W., Clark, M. A., Keefe, A. D., Madin, A., Read, J. A., Rivers, E. L., & Zhang,
Y. (2024). Screening ultra-large encoded compound libraries leads to novel protein–ligand
interactions and high selectivity. Journal of Medicinal Chemistry, 67(2), 864–884.
98. Clark, D. E. (2020). Virtual screening: Is bigger always better? Or can small be beautiful?
Journal of Chemical Information and Modeling, 60, 4120–4123.
99. Kontoyianni, M. (2022). Library size in virtual screening: Is it truly a number’s game? Expert
Opinion on Drug Discovery, 17(11), 1177–1179.

Part II
The Pitfalls Between Experimentation
and Simulation
Соседние файлы в папке Библиотека им академика М.И. Перельмана
