Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5364_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
02.09.2026
Размер:
21 Мб
Скачать
100 D. S. de Sousa et al.
particularly valuable when dealing with the diversity and complexity of biological data, enabling a deeper understanding of causal relationships and underlying patterns [5].
Generative AI This eld is revolutionizing drug discovery through advanced models like deep neural networks. These models efciently explore complex molec­ular structures, predicting drug candidates with unprecedented precision. Their nuanced understanding of biological data uncovers hidden patterns, ushering in a new era of personalized medicine and targeted drug development. The future of drug discovery is shaped by the transformative capabilities of generative AI [131].
As we move toward the future, these converging trends at the intersection of ML and drug discovery promise not only to transform the efciency of the process but also to open new horizons in therapy personalization and the development of innovative treatments. The marriage between biology and articial intelligence is shaping an exciting era where therapeutic innovation is driven by the synergy between the human mind and the analytical capabilities of machines.
8 Guiding Questions for Implementing Machine Learning
in Drug Design
Before we conclude, the following set of questions is designed to guide the effective use of machine learning in drug design. These questions cover various aspects, including the primary objectives, data availability and quality, model selection, performance evaluation, and integration into the drug development process. Addi­tionally, they address challenges related to data preprocessing, model interpretation, and resource requirements. By thoroughly addressing these questions, researchers and practitioners can ensure that their machine learning projects in drug design are well founded, robust, and capable of overcoming specic challenges in this eld. Guiding questions for implementing machine learning in drug design
What is the primary objective of using machine learning in drug design?
Identication of new molecules?
Optimization of existing molecules?
Prediction of efcacy and toxicity?
Target identication/optimization?
What data are available for training machine learning models?
Chemical structures?
Biological assay data?
Genomic data?
What are the quality and quantity of this data?
Is there enough data for effective training?
Is the data well annotated and standardized?
(continued)
4 Machine Learning and Neural Network Methods Applied to Drug Discovery 101
What types of machine learning models will be used?
Supervised learning models (e.g., regression and classication)?
Unsupervised learning models (e.g., and dimensionality reduction)?
Reinforcement learning models?
How will the performance of the models be evaluated?
What performance metrics will be used (e.g., accuracy, recall, and AUC-ROC)?
What validation techniques will be applied (e.g., cross-validation and data split)?
How will imbalanced data be handled, if present?
Resampling techniques (e.g., oversampling and undersampling)?
How will the results of the models be interpreted?
Feature importance analysis?
Model explanations and visualizations?
How will machine learning models be integrated into the drug design process?
Integration with virtual screening platforms?
Integration with in vitro and in vivo experimentation pipelines?
What are the main challenges and limitations expected in the application of machine learning to drug design?
Interpretability issues?
Generalization to new data?
How will the lifecycle of machine learning models be managed?
Model updates and retraining?
Post-implementation performance monitoring?
What computational resources will be needed?
Required hardware and software?
Use of GPUs or TPUs to accelerate training?
How will the impact of machine learning models on the drug development process be measured?
Cost and time reduction?
Increased success rate in clinical trials?

9 Conclusions

In conclusion, the exploration of ML within the realm of AI offers a transformative lens through which we understand, navigate, and revolutionize complex processes like drug discovery. As we traverse the intricate stages of problem denition, model selection, training, and validation, it becomes evident that the success of ML applications hinges on the careful consideration of various factors, including prob­lem specicity and algorithm selection. However, with the power of ML, especi ally within the domain of neural networks and deep learning, we witness a paradigm shift in how we approach challenges in drug discovery.
Yet, this transformative journey is not without its challenges. The omnipresence of biases, the delicate balance between overtting and undertting, the imperative
102 D. S. de Sousa et al.
need for interpretability, computational costs, data dependencies, and the quest for robust models all contr ibute to the nuanced landscape of ML in drug discovery. Each limitation serves as a reminder of the ethical responsibilities, technical intricacies, and economic considerations that accompany the integration of ML into such critical domains.
The application of ML in drug discovery, exemplied through target identica­tion, lead discovery, and preclinical and clinical development, underscores the potential for accelerated advancements in the pharmaceutical industry. The careful orchestration of various ML algorithms, from classiers to regression models, empowers researchers to sift through vast data sets and extract meaningful insights that shape the trajectory of drug development.
As we navigate through the challenges and leverage the emerging trends of transfer learning, explainable AI, and the integration of imaging data and molecular structures, we glimpse into a future where the marriage between biology and articial intelligence paves the way for unprecedented innovation. The continued advancements in deep learning and the collaborative synergy between human expertise and machine analytical capabilities promise not just efciency but also a profound understanding of the intricate interplay between therapeutic compounds and biological targets.
In essence, the union of ML and drug discovery heralds a promising era where scientic exploration is propelled by the fusion of human intellect and the compu­tational prowess of machines. As we stand at the cross roads of science and technol­ogy, the journey forward promises not just accelerated drug development but a deeper, more personalized understanding of therapies, fostering a landscape where innovation thrives at the intersection of AI and the intricacies of life sciences.

References

1. Gelijns, A. C. (2014). Technological innovation: Comparing development of drugs, devices, and procedures in medicine. National Academies Press.
2. Taye, M. M. (2023). Understanding of machine learning with deep learning: Architectures, workow, applications and future directions. Computers, 12(5), 91.
3. Niazi, S. K., & Mariam, Z. (2023). Recent advances in machine-learning-based chemoinformatics: A comprehensive review. International Journal of Molecular Sciences, 24(14), 11488.
4. Vora, L. K., Gholap, A. D., Jetha, K., Thakur, R. R. S., Solanki, H. K., & Chavda, V. P. (2023). Articial intelligence in pharmaceutical technology and drug delivery design. Pharmaceutics, 15(7), 1916.
5. Askr, H., Elgeldawi, E., Aboul Ella, H., Elshaier, Y. A. M. M., Gomaa, M. M., & Hassanien, A. E. (2023). Deep learning in drug discovery: An integrative review and future challenges. Articial Intelligence Review, 56(7), 5975–6037.
6. McCulloch, W. S., & Pitts, W. (1943). A logical calculus of the ideas immanent in nervous activity. The Bulletin of Mathematical Biophysics, 5, 115–133.
7. Weik, M. H. (1961). The ENIAC story. Ordnance, 45(244), 571–575.
8. Turing, A. M. (1950). I.Computing machinery and intelligence. Mind, LIX(236), 433– 460.
4 Machine Learning and Neural Network Methods Applied to Drug Discovery 103
9. McCarthy, J., & Feigenbaum, E. A. (1990). In memoriam: Arthur samuel: Pioneer in machine learning. AI Magazine, 11(3), 10.
10. Samuel, A. L. (1959). Some studies in machine learning using the game of checkers. IBM Journal of Research and Development, 3(3), 210–229.
11. Rosenblatt, F. (1957). The perceptron, a perceiving and recognizing automaton Project Para. Cornell Aeronautical Laboratory.
12. Cover, T., & Hart, P. (1967). Nearest neighbor pattern classication. IEEE Transactions on Information Theory, 13(1), 21–27.
13. E. A. Feigenbaum, B. G. Buchanan, and J. Lederberg, 1970. On generality and problem solving: A case study using the DENDRAL program,
14. Vamathevan, J., et al. (2019). Applications of machine learning in drug discovery and development. Nature Reviews. Drug Discovery, 18(6), 463–477.
15. Dara, S., Dhamercherla, S., Jadav, S. S., Babu, C. H. M., & Ahsan, M. J. (2022). Machine learning in drug discovery: A review. Articial Intelligence Review, 55(3), 1947–1999.
16. Zhang, L., Tan, J., Han, D., & Zhu, H. (2017). From machine learning to deep learning: Progress in machine intelligence for rational drug discovery. Drug Discovery Today, 22(11), 1680–1685.
17. Stephenson, N., et al. (2019). Survey of machine learning techniques in drug discovery. Current Drug Metabolism, 20(3), 185–193.
18. Popova, M., Isayev, O., & Tropsha, A. (2018). Deep reinforcement learning for de novo drug design. Science Advances, 4(7), eaap7885.
19. Ajay, Walters, W. P., & Murcko, M. A. (1998). Can we learn to distinguish between drug­likeand nondrug-likemolecules? Journal of Medicinal Chemistry, 41(18), 3314–3324.
20. Carracedo-Reboredo, P., et al. (2021). A review on machine learning approaches and trends in drug discovery. Computational and Structural Biotechnology Journal, 19, 4538–4558.
21. Irwin, J. J., & Shoichet, B. K. (2005). ZINC- a free database of commercially available compounds for virtual screening. Journal of Chemical Information and Modeling, 45(1), 177–182.
22. Wang, Y., Xiao, J., Suzek, T. O., Zhang, J., Wang, J., & Bryant, S. H. (2009). PubChem: A public information system for analyzing bioactivities of small molecules. Nucleic Acids Research, 37(suppl_2), W623–W633.
23. Wishart, D. S., et al. (2008). DrugBank: A knowledgebase for drugs, drug actions and drug targets. Nucleic Acids Research, 36(suppl_1), D901–D906.
24. Gaulton, A., et al. (2012). ChEMBL: A large-scale bioactivity database for drug discovery.
Nucleic Acids Research, 40
25. Wallach, I., Dzamba, M., & Heifets, A. (2015). AtomNet: a deep convolutional neural network for bioactivity prediction in structure-based drug discovery. arXiv preprint arXiv:1510.02855.
26. Stokes, J. M., et al. (2020). A deep learning approach to antibiotic discovery. Cell, 180(4), 688–702.
27. Burki, T. (2020). A new paradigm for drug development. Lancet Digit Health, 2(5), e226– e227.
28. Whitby, B. (2009). Articial intelligence. The Rosen Publishing Group, Inc.
29. Winston, P. H. (1984). Articial intelligence. Addison-Wesley Longman Publishing, Inc.
30. Ramesh, A. N., Kambhampati, C., Monson, J. R. T., & Drew, P. J. (2004). Articial intelligence in medicine. Annals of the Royal College of Surgeons of England, 86(5), 334.
31. Hunt, E. B. (2014). Articial intelligence. Academic.
32. El Naqa, I., & Murphy, M. J. (2015). What is machine learning? Springer.
33. Rusk, N. (2016). Deep learning. Nature Methods, 13(1), 35.
34. LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. Nature, 521(7553), 436–444.
35. Bonaccorso, G. (2017). Machine learning algorithms. Packt Publishing Ltd.
36. Mahesh, B. (2020). Machine learning algorithms-a review. International Journal of Science and Research (IJSR), 9(1), 381–386.
37. Zhou, Z.-H. (2021). Machine learning. Springer Nature.
(D1), D1100–D1107.
104 D. S. de Sousa et al.
38. Agatonovic-Kustrin, S., & Beresford, R. (2000). Basic concepts of articial neural network (ANN) modeling and its application in pharmaceutical research. Journal of Pharmaceutical and Biomedical Analysis, 22(5), 717–727.
39. Morales, E. F., & Escalante, H. J. (2022). A brief introduction to supervised, unsupervised, and reinforcement learning. In Biosignal processing and classi cation using computational learn- ing and intelligence (pp. 111–129). Elsevier.
40. Rigatti, S. J. (2017). Random forest. Journal of Insurance Medicine, 47(1), 31–39.
41. Lorena, A. C., & De Carvalho, A. C. (2007). Uma introdução às support vector machines. Revista de Informática Teórica e Aplicada, 14(2), 43–67.
42. Siemers, F. M., & Bajorath, J. (2023). Differences in learning characteristics between support vector machine and random forest models for compound classi cation revealed by Shapley value analysis. Scientic Reports, 13(1), 5983.
43. Rosenblatt, F., & Papert, S. (2021). Perceptron, 9.
44. Gallant, S. I. (1990). Perceptron-based learning algorithms. IEEE Transactions on Neural Networks, 1(2), 179–191.
45. Ruck, D. W., Rogers, S. K., & Kabrisky, M. (1990). Feature selection using a multilayer perceptron. Journal of Neural Network Computing, 2(2), 40–48.
46. Popescu, M.-C., Balas, V. E., Perescu-Popescu, L., & Mastorakis, N. (2009). Multilayer perceptron and neural networks. WSEAS Transactions on Circuits and Systems, 8(7), 579–588.
47. Sharma, S., Sharma, S., & Athaiya, A. (2017). Activation functions in neural networks. Towards Data Science, 6(12), 310–316.
48. A. D. Rasamoelina, F. Adjailia, and P. Sinčák, A review of activation function for articial neural network,in 2020 IEEE 18th World Symposium on Applied Machine Intelligence and Informatics (SAMI), IEEE, 2020, pp. 281–286.
49. Leijnen, S., & van Veen, F. (2020). The neural network zoo. Proceedings of the West Mark Educational Association conference, 47(1).
50. Hassoun, M. H. (1995). Fundamentals of articial neural networks. MIT Press.
51. OShea, K., & Nash, R. (2015). An introduction to convolutional neural networks. arXiv preprint arXiv:1511.08458.
52. Chauhan, R., Ghanshala, K. K., & Joshi, R. C. (2018). Convolutional neural network (CNN) for image detection and recognition. In 2018 rst international conference on secure cyber
computing and communication (ICSCCC)
53. Albawi, S., Mohammed, T. A., & Al-Zawi, S. (2017). Understanding of a convolutional neural network. In 2017 international conference on engineering and technology (ICET) (pp. 1–6). IEEE.
54. Grossberg, S. (2013). Recurrent neural networks. Scholarpedia, 8(2), 1888.
55. Medsker, L. R., & Jain, L. C. (2001). Recurrent neural networks. Design and Applications, 5(64–67), 2.
56. Medsker, L., & Jain, L. C. (1999). Recurrent neural networks: Design and applications. CRC Press.
57. Nosouhian, S., Nosouhian, F., & Khoshouei, A. K. (2021). A review of recurrent neural network architecture for sequence learning: Comparison between LSTM and GRU.
58. Sherstinsky, A. (2020). Fundamentals of recurrent neural network (RNN) and long short-term memory (LSTM) network. Physica D, 404, 132306.
59. Dey, R., & Salem, F. M. (2017). Gate-variants of gated recurrent unit (GRU) neural networks. In 2017 IEEE 60th international Midwest symposium on circuits and systems (MWSCAS) (pp. 1597–1600). IEEE.
60. Martín-Guerrero, J. D., & Lamata, L. (2022). Quantum machine learning: A tutorial. Neurocomputing, 470, 457–461.
61. Chao, W.-L. (2011). Machine learning tutorial. Digital Image and Signal Processing.
62. Ball, N. M., & Brunner, R. J. (2010). Data mining and machine learning in astronomy. International Journal of Modern Physics D, 19(07), 1049–1106.
(pp. 278–282). IEEE.
4 Machine Learning and Neural Network Methods Applied to Drug Discovery 105
63. Emmanuel, T., Maupong, T., Mpoeleng, D., Semong, T., Mphago, B., & Tabona, O. (2021). A survey on missing data in machine learning. Journal of Big Data, 8(1), 1–37.
64. Kumar, S. A., et al. (2022). Machine learning and deep learning in data-driven decision making of drug discovery and challenges in high-quality data acquisition in the pharmaceutical industry. Future Medicinal Chemistry, 14(4), 245–270.
65. Ridzuan, F., & Zainon, W. M. N. W. (2019). A review on data cleansing methods for big data. Procedia Computer Science, 161, 731–738.
66. Ali, P. J. M., Faraj, R. H., Koya, E., Ali, P. J. M., & Faraj, R. H. (2014). Data normalization and standardization: A technical report. Machine Learning Technical Report, 1(1), 1–6.
67. Potdar, K., Pardawala, T. S., & Pai, C. D. (2017). A comparative study of categorical variable encoding techniques for neural network classiers. International Journal of Computers and Applications, 175(4), 7–9.
68. Kurita, T. (2019). Principal component analysis (PCA). In Computer vision: A reference guide (pp. 1–4).
69. Wattenberg, M., Viégas, F., & Johnson, I. (2016). How to use t-SNE effectively. Distill, 1(10), e2.
70. Chawla, N. V., Bowyer, K. W., Hall, L. O., & Kegelmeyer, W. P. (2002). SMOTE: Synthetic minority over-sampling technique. Journal of Articial Intelligence Research, 16, 321–357.
71. Barandela, R., Valdovinos, R. M., Sánchez, J. S., & Ferri, F. J. The imbalanced training sample problem: Under or over sampling? In Structural, Syntactic, and Statistical Pattern
Recognition: Joint IAPR International Workshops, SSPR 2004 and SPR 2004, Lisbon, Por­tugal, August 18-20, 2004. Proceedings (pp. 806–814). Springer, 2004.
72. Cai, J., Luo, J., Wang, S., & Yang, S. (2018). Feature selection in machine learning: A new perspective. Neurocomputing, 300,70–79.
73. Raschka, S. (2018). Model evaluation, model selection, and algorithm selection in machine learning. arXiv preprint arXiv:1811.12808.
74. da Silva, A. P., Chiari, L. P. A., Guimaraes, A. R., Honorio, K. M., & da Silva, A. B. F. (2021). Drug design of new 5-HT6R antagonists aided by articial neural networks. Journal of Molecular Graphics & Modelling, 104, 107844.
75. Kim, Y.-J., & Chi, M. (2018). Temporal belief memory: Imputing missing data during RNN training. In Proceedings of the 27th international joint conference on articial intelligence (IJCAI-2018).
76. Christoffersen, P., & Jacobs, K. (2004). The importance of the loss function in option valuation. Journal of Financial Economics, 72(2), 291–318.
77. Wythoff, B. J. (1993). Backpropagation neural networks: A tutorial. Chemometrics and Intelligent Laboratory Systems, 18(2), 115–155.
78. Devarakonda, A., Naumov, M., & Garland, M. (2017). Adabatch: Adaptive batch sizes for training deep neural networks. arXiv preprint arXiv:1712.02029.
79. Humphrey, A., et al. (2022). Machine-learning classi mating F1-score in the absence of ground truth. Monthy Notices of the Royal Astronomical Society Letters, 517(1), L116–L120.
80. Kuhn, M. (2014). Futility analysis in the cross-validation of machine learning models. arXiv preprint arXiv:1405.6974.
81. Roy, K., Das, R. N., Ambure, P., & Aher, R. B. (2016). Be aware of error measures. Further studies on validation of predictive QSAR models. Chemometrics and Intelligent Laboratory Systems, 152,18–33.
82. Rücker, C., Rücker, G., & Meringer, M. (2007). Y-randomization and its variants in QSPR/ QSAR. Journal of Chemical Information and Modeling, 47(6), 2345–2357.
83. Narkhede, S. (2018). Understanding auc-roc curve. Towards Data Science, 26(1), 220–227.
84. Feurer, M., & Hutter, F. (2019). Hyperparameter optimization. In Automated machine learn- ing: Methods, systems, challenges (pp. 3–33).
85. Bergstra, J., Bardenet, R., Bengio, Y., & Kégl, B. (2011). Algorithms for hyper-parameter optimization. In Advances in neural information processing systems (Vol. 24).
cation of astronomical sources: Esti-
106 D. S. de Sousa et al.
86. Kim, Y., & Chung, M. (2019). An approach to hyperparameter optimization for the objective function in machine learning. Electronics (Basel), 8(11), 1267.
87. Kadhim, Z. S., Abdullah, H. S., & Ghathwan, K. I. (2022). Articial neural network hyperparameters optimization: A survey. International Journal of Online & Biomedical Engineering, 18(15), 59.
88. Mazilu, S., & Iria, J. (2011). L1 vs. l2 regularization in text classication when learning from labeled features. In 2011 10th international conference on machine learning and applications and workshops (pp. 166–171). IEEE.
89. Kuhn, M., Johnson, K., Kuhn, M., & Johnson, K. (2013). Over-tting and model tuning. In Applied predictive modeling (pp. 61–92).
90. Probst, P., Wright, M. N., & Boulesteix, A. (2019). Hyperparameters and tuning strategies for random forest. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 9(3), e1301.
91. Ghawi, R., & Pfeffer, J. (2019). Efcient hyperparameter tuning with grid search for text categorization using kNN approach with BM25 similarity. Open Computer Science, 9(1), 160–180.
92. Friedrichs, F., & Igel, C. (2005). Evolutionary tuning of multiple SVM parameters. Neurocomputing, 64, 107–117.
93. Mastorakis, G. (2018). Human-like machine learning: limitations and suggestions. arXiv preprint arXiv:1811.06052.
94. Malik, M. M. (2020). A hierarchy of limitations in machine learning. arXiv preprint arXiv:2002.05193.
95. Kenge, R. (2020). Machine learning, its limitations, and solutions over IT. International Journal of Information Technology Modeling and Computing, 11, 73.
96. Jabbar, H., & Khan, R. Z. (2015). Methods to avoid over- tting and under-tting in supervised machine learning (comparative study). Computer Science, Communication and Instrumenta- tion Devices, 70(10.3850), 978–981.
97. Bilmes, J. (2020). Undertting and overtting in machine learning (UW ECE course notes) (Vol. 5).
98. Ying, X. (2019). An overview of overtting and its solutions. In Journal of physics: Confer- ence series (p. 022022). IOP Publishing.
99. Lavecchia, A. (2015). Machine-learning approaches in drug discovery: Methods and applica­tions. Drug Discovery Today, 20(3), 318–331.
100. Carvalho, D. V., Pereira, E. M., & Cardoso, J. S. (2019). Machine learning interpretability: A survey on methods and metrics. Electronics (Basel), 8(8), 832.
101. Gramegna, A., & Giudici, P. (2021). SHAP and LIME: An evaluation of discriminative power in credit risk. Frontiers in Articial Intelligence, 4, 752558.
102. Justus, D., Brennan, J., Bonner, S., & McGough, A. S. (2018). Predicting the computational cost of deep learning models. In (pp. 3873–3882). IEEE.
103. Miller, T., et al. (2023). Advancements in articial intelligence circuits and systems (AICAS). Electronics (Basel), 13(1), 102.
104. Elbadawi, M., Gaisford, S., & Basit, A. W. (2021). Advanced machine-learning techniques in drug discovery. Drug Discovery Today, 26(3), 769–777.
105. Koscielny, G., et al. (2017). Open targets: A platform for therapeutic target identication and validation. Nucleic Acids Research, 45(D1), D985–D994.
106. Yang, H., et al. (2019). admetSAR 2.0: Web-service for prediction and optimization of chemical ADMET properties. Bioinformatics, 35(6), 1067–1069.
107. Advanced Chemistry Development Inc. (2016). (ACD/Labs) ACD/PERCEPTA Version 2015. Frankfurt am Main. www.acdlabs.com/pka
108. Pires, D. E. V., Blundell, T. L., & Ascher, D. B. (2015). pkCSM: Predicting small-molecule pharmacokinetic and toxicity properties using graph-based signatures. Journal of Medicinal
Chemistry, 58(9), 4066–4072.
2018 IEEE international conference on big data (Big Data)
4 Machine Learning and Neural Network Methods Applied to Drug Discovery 107
109. Schwaller, P., Strobelt, H., Laino, T., & Hoover, B. IBM RXN for Chemistry: Unveiling the grammar of the organic chemistry language.
110. Ramsundar, B. et al., DeepChem: Democratizing deep-learning for drug discovery, quantum chemistry, materials science and biology.
111. Landrum, G. (2013). RDKit: A software suite for cheminformatics, computational chemistry, and predictive modeling. Greg Landrum, 8, 31.
112. Dixon, S. L., Duan, J., Smith, E., Von Bargen, C. D., Sherman, W., & Repasky, M. P. (2016). AutoQSAR: An automated machine learning tool for best-practice quantitative structure– activity relationship modeling. Future Medicinal Chemistry, 8(15), 1825–1839.
113. Yap, C. W. (2011). PaDEL-descriptor: An open source software to calculate molecular descriptors and ngerprints. Journal of Computational Chemistry, 32(7), 1466–1474.
114. Sybyl, X. (2011). Molecular modelling software (Vol. 1). Tripos Certara.
115. Matlab, S. (2012). Matlab. The MathWorks.
116. The UniProt Consortium. (2017). UniProt: The universal protein knowledgebase. Nucleic Acids Research, 45(D1), D158–D169.
117. Burley, S. K., Berman, H. M., Kleywegt, G. J., Markley, J. L., Nakamura, H., & Velankar, S. (2017). Protein Data Bank (PDB): The single global macromolecular structure archive. In Protein crystallography: Methods and protocols (pp. 627–641).
118. Avram, S., et al. (2023). DrugCentral 2023 extends human clinical data and integrates veterinary drugs. Nucleic Acids Research, 51(D1), D1276–D1287.
119. Cai, M.-C., et al. (2015). ADReCS: An ontology database for aiding standardization and hierarchical classication of adverse drug reaction terms. Nucleic Acids Research, 43(D1), D907–D913.
120. Davis, A. P., Wiegers, T. C., Johnson, R. J., Sciaky, D., Wiegers, J., & Mattingly, C. J. (2023). Comparative Toxicogenomics database (CTD): Update 2023. Nucleic Acids Research, 51(D1), D1257–D1262.
121. Piñero, J., et al. (2015). DisGeNET: A discovery platform for the dynamical exploration of human diseases and their genes. Database, 2015, bav028.
122. UK Biobank. (2014). About UK biobank.
123. Szklarczyk, D., et al. (2023). The STRING database in 2023: Protein–protein association networks and functional enrichment analyses for any sequenced genome of interest. Nucleic Acids Research, 51(D1), D638–D646.
124. Zhou, Y., et al. (2024). TTD: Therapeutic target database describing target druggability information. Nucleic Acids Research, 52(D1), D1465–D1477.
125. Gallo, K., Goede, A., Eckert, A., Moahamed, B., Preissner, R., & Gohlke, B.-O. (2021). PROMISCUOUS 2.0: A resource for drug-repositioning. Nucleic Acids Research, 49(D1), D1373–D1380.
126. Zarin, D. A., Tse, T., Williams, R. J., Califf, R. M., & Ide, N. C. (2011). The ClinicalTrials. Gov results databaseUpdate and key issues. New England Journal of Medicine, 364(9), 852–860.
127. Gonoskov, A., Wallin, E., Polovinkin, A., & Meyerov, I. (2019). Employing machine learning for theory validation and identication of experimental conditions in laser-plasma physics. Scientic Reports, 9(1), 7043.
128. Vayena, E., Blasimme, A., & Cohen, I. G. (2018). Machine learning in medicine: Addressing ethical challenges. PLoS Medicine, 15(11), e1002689.
129. Cai, C., et al. (2020). Transfer learning for drug discovery. Journal of Medicinal Chemistry, 63(16), 8683–8694.
130. Ponzoni, I., Páez Prosper, J. A., & Campillo, N. E. (2023). Explainable articial intelligence: A taxonomy and guidelines for its application to drug discovery. Wiley Interdisciplinary Reviews: Computational Molecular Science, 13(6), e1681.
131. Bian, Y., & Xie, X.-Q. (2021). Generative chemistry: Drug discovery with deep learning generative models. Journal of Molecular Modeling, 27,1–18.
Chapter 5
Clustering of Small Molecules
Alan Talevi, Lucas Alberca, and Carolina Bellera
Abstract Clustering of small molecules nds a diversity of applications in chem-
istry and, in particular, in the elds of cheminformatics and drug discovery. It may be used directly as an unsupervised machine-learning strategy to identify existing patterns in a chemical data set or libraries or integrated into supervised machine­learning studies to partition a sample of compounds into representative subsamples (e.g., training and validation data). It may also be applied to select which in silico hits from a virtual screening campaign will be submitted to experimental conrmation, or to dene which hits emerging from a wetscreening campaign will be prioritized for further development or characterization. Here, we review general strategies to validate the output of a clustering algorithm and discuss current challenges and possible future directions in the eld of small molecule clustering.
Keywords Cluster replication · Clustering, External validity · Hierarchical clustering · Internal validity measures · Multi-view clustering · Non-hierarchic al clustering · Relative validity · Small molecules · Stability measures · Subspace clustering · Supervised learning · Unsupervised learn ing · Validation · Validity
A. Talevi () Laboratory of Bioactive Compound Research and Development (LIDeB), Faculty of Exact Sciences, University of La Plata (UNLP), Buenos Aires, Argentina
Consejo Nacional de Investigaciones Cientícas y Técnicas (CONICET), CCT La Plata, La Plata, Buenos Aires, Argentina
Boolzi SA, Buenos Aires, Argentina e-mail: alantalevi@gmail.com
L. Alberca · C. Bellera Laboratory of Bioactive Compound Research and Development (LIDeB), Faculty of Exact Sciences, University of La Plata (UNLP), Buenos Aires, Argentina
Consejo Nacional de Investigaciones Cientícas y Técnicas (CONICET), CCT La Plata, La Plata, Buenos Aires, Argentina
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2024 V. G. Maltarollo (ed.), Computer-Aided and Machine Learning-Driven Drug Design, Computer-Aided Drug Discovery and Design 3,
https://doi.org/10.1007/978-3-031-76718-0_5
109
110 A. Talevi et al.

1 Historical Background

Human beings have developed the ability to detect higher-level organizations and group objects and experiences into functional categories, which is a precondition of sophisticated thought [45]. Categorization provides the basis to recognize and respond instantly and appropriately to objects and situations even if they are not identical to the ones that we have previously met. In other words, categorization is a
way of anticipating behavior based on prior experience with similar objects or scenarios. This is a central proposition to understand the application of the
approaches described in this chapter.
Perceptual grouping is performed at a remarkable speed based on principles of proximity (objects that are close to each other in space are perceived as part of a group), similarity (grouping elements by shape, size, and color), continuity (a relation is naturally perceived between objects that are arranged in continuous lines or curves), and common fate (objects that move at the same speed and/or direction tend to be grouped), among others [9]. The tendency of human beings to group things together is reected in their language: common nouns name people, animals, or things of the same class or species, whereas proper nouns name particular objects individualized from any other of the same species (a particular person, a particular animal, a particular city, etc.). In fact, categories of knowledge that have a common noun in a language are more efciently captured by the eye than those that do not (which is known as conceptual grouping)[31]. Moreover, collective nouns express functional and/or proximity relationships between similar objects (named by the same common noun), for example, a pack of wolves, a crowd, a school of sh, and a choir.
A fundamental challenge in the eld of articial inte lligence is how to teach a machine to perform grouping or categorization as good as (or better than) evolution has taught human beings to do [41]. Let us take, for instance, the molecular structures in Fig. 5.1. Even if the observer has no specic training in chemistry, it is likely that they will perceive the similarities between the two compounds that share the benzodiazepine nucleus. However, when dealing with large-scale chemical data sets or repositories, visually analyzing and comparing every possible pair of molecules to produce a hypothesis on possible clusters that may be derived from it
Fig. 5.1 The human brain performs perceptual grouping at a remarkable speed based on principles such as similarity. It is likely that even people with no training in chemistry will cluster the three compounds that share the benzodiazepine nucleus (alprazolam, diazepam, and lorazepam) apart from the corticosteroid (prednisolone)