Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5606_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Foreword
- •Acknowledgments
- •Contents
- •1.1 Structure-Based Drug Discovery (SBDD)
- •1.2 Ligand-Based Drug Design (LBDD)
- •1.3 Echoes from the Past, Visions from the Future
- •References
- •1 Introduction
- •2.2 Second Step: Data Curation
- •2.4 Fourth Step: Updating and Maintenance
- •2 Databases and Curation
- •8 Perspectives
- •9 Conclusion
- •References
- •1 Introduction
- •2.1 Making and Matching Protein Models
- •2.2 Simulating Protein Movements
- •2.3 Analyzing Changes in Protein Shape
- •3 Pharmacogenomics in Drug Development
- •4 Case Studies of Genomics-Based Drug Design
- •References
- •1 Historical Background
- •1.1 Timeline
- •2 Methodology Overview
- •2.1 Neural Networks
- •2.1.1 Perceptron
- •2.1.2 Multilayer Neural Networks
- •2.1.3 Types of Neural Networks
- •Feedforward
- •Recurrent Neural Networks
- •LSTM
- •2.2 Deep Learning
- •3 Using Machine Learning
- •3.2 Data Collection
- •3.3 Data Preprocessing
- •3.4 Model Selection
- •3.5 Model Training
- •3.6 Validation
- •3.7 Tuning
- •3.8 Prediction
- •4 Limitations
- •4.1 Bias
- •4.3 Interpretability
- •4.4 Computational Cost
- •4.5 Data Dependency
- •4.6 Robustness
- •5 Applications in Drug Discovery
- •5.2 Lead Discovery
- •5.3 Preclinical and Clinical Development
- •6 Resources and Tools
- •7 Challenges and Perspectives
- •7.1 Future Trends
- •9 Conclusions
- •References
- •1 Historical Background
- •1.1 Applications in Drug Discovery
- •2 Validations and Controls
- •2.1 Internal Validation
- •2.2 External Validation
- •2.3 Relative Cluster Validation
- •3 Challenges and Perspectives
- •4 Conclusions
- •References
- •1 Historical Background
- •2 OECD Principles
- •2.1 A Defined Endpoint
- •2.2 An Unambiguous Algorithm
- •2.5 A Mechanistic Interpretation, if Possible
- •3 Software and Tools
- •4 Validations and Controls
- •4.1 Internal and External Validation
- •4.1.1 Regression Metrics
- •4.2 Applicability Domain
- •4.3 Randomization Tests
- •5 Interpretation
- •6 Practical Advice During QSAR Modeling
- •7 Application
- •8 Challenges and Perspectives
- •References
- •1 Molecular Docking
- •2 Advances in Scoring Functions and Search Algorithms
- •2.2 Critical Characteristics of Search Algorithms
- •2.3 Docking Programs and Scoring Functions
- •3 Calculations Performed During Docking Simulations
- •4 Essential Components for a Good Docking Program
- •5 Limitations of the Docking Technique
- •6 Validation of Docking Results
- •7 Inappropriate Use of Validation Methods in Docking
- •9 Use of Machine Learning in Molecular Docking
- •11 Challenges
- •12 Conclusions
- •References
- •3 System Preparation for MD Simulations
- •3.1 Solvation and Microensemble
- •3.2 Force Fields: General Concept and Relevant Choices
- •3.3 The Concept of Replicas and Timescale
- •4.1.2 Protein Root Mean Square Fluctuation (RMSF)
- •4.1.4 Protein Secondary Structure Analysis
- •4.1.5 Principal component Analysis (PCA)
- •4.1.6 Markov State Modelling
- •4.1.7 Distance Calculations
- •4.1.8 Angle and Plane Calculations
- •4.2.2 Distances and Ligand-Induced Geometry Rearrangements
- •4 Molecular Dynamics Analysis
- •4.1 Protein Perspective
- •4.1.1 Protein Root Mean Square Deviation (RMSD)
- •4.3 Ligand Perspective
- •4.3.1 Ligand Properties
- •4.3.2 Ligand Root Mean Square Deviation
- •4.3.3 Ligand Root Mean Square Fluctuation
- •4.3.4 Angles and Dihedrals
- •5.1 Protein Structure Prediction and Preparation
- •5.2 Molecular Docking
- •6 Concluding Remarks and Outlook
- •Glossary
- •References
- •1 Introduction
- •2.1 MDeNM
- •2.2 Collective Molecular Dynamics (coMD)
- •2.3 ClustENM and ClustENMD
- •3 Ensemble Docking
- •References
- •1 Introduction
- •1.1 Advantages, Disadvantages, Innovations, and Challenges
- •1.2 Recent Advances in Accessible FEP Software Tools
- •1.3 Applications of FEP in Industry and Consortiums
- •2 Expanding the Potential of FEP Calculations
- •2.1 Validating Binding Poses
- •2.2 Dealing with Solvent
- •2.3 FEP and Allostery
- •2.4 FEP and Covalent Ligands
- •2.5 Applications of FEP in Scaffold Hopping
- •2.6 Positional Analogue Scanning
- •2.7 Combinations and Alternative Approaches
- •3 Machine Learning for FEP
- •3.4 Implications for ML in FEP Calculations
- •4 Final Considerations
- •5 First Steps to FEP Simulations
- •References
- •1 Background
- •2 Ultra-Large Screening Libraries and Chemical Spaces
- •3.1 Implications of Dataset Size
- •4 Ligands on the Ultra-Large Scale
- •4.1 Ultra-Large 2D Similarity Searches
- •7 Challenges and Future Perspectives
- •7.1 Hit Triage: An Old Problem on a New Dimension
- •8 Conclusions
- •Appendix
- •References
- •1 Introduction
- •2 Enzymatic Activity Evaluations
- •3 Cytotoxicity Evaluation and Cell Viability
- •4 Antiviral Assays in Experimental Validation
- •6 In Vivo Evaluation of Compounds
- •7 Conclusions
- •References
- •1 Introduction
- •3.1 Data Collection
- •3.2 Data Preprocessing
- •3.4 Model Choice
- •3.5 Model Training
- •3.6 Model Assessment
- •3.7 External Validation
- •3.8 Implementation and Availability
- •3.9 Continuous Update
- •5 Conclusions and Perspectives
- •References
- •1 Experimental Approaches to Obtain Protein Structure
- •1.1 X-Ray Crystallography
- •1.2 Nuclear Magnetic Resonance
- •1.3 Cryo-EM
- •1.4 Hybrid Methods
- •2 Modeling Approaches to Obtain Protein Structure
- •2.1 Homology Modeling
- •2.2 Ab Initio Modeling
- •2.3 New Approaches
- •3 Conformational Diversity of Proteins
- •3.1 Characterization of Protein Conformational States
- •3.2 Experimental Methods to Study Protein Dynamics and Conformations
- •3.4 Molecular Dynamics Simulation
- •3.5 Sampling Strategies
- •4 Remarks and Perspectives
- •References
- •1 Introduction
- •2 Structure-Based Drug Design of HIV Protease Inhibitors
- •2.1 HIV-1 Protease as a Therapeutic Target
- •2.2.1 Saquinavir
- •2.2.2 Indinavir
- •2.3.1 Lopinavir
- •2.3.2 Darunavir
- •6 Conclusions
- •References
- •4 Experimental Methods to Analyze NR Activity
- •4.2 Coregulator-Recruitment
- •5 Concluding Remarks and Outlook
- •References

100 D. S. de Sousa et al.
particularly valuable when dealing with the diversity and complexity of biological
data, enabling a deeper understanding of causal relationships and underlying
patterns [5].
Generative AI This field is revolutionizing drug discovery through advanced
models like deep neural networks. These models efficiently explore complex molecular structures, predicting drug candidates with unprecedented precision. Their
nuanced understanding of biological data uncovers hidden patterns, ushering in a
new era of personalized medicine and targeted drug development. The future of drug
discovery is shaped by the transformative capabilities of generative AI [131].
As we move toward the future, these converging trends at the intersection of ML
and drug discovery promise not only to transform the efficiency of the process but
also to open new horizons in therapy personalization and the development of
innovative treatments. The marriage between biology and artificial intelligence is
shaping an exciting era where therapeutic innovation is driven by the synergy
between the human mind and the analytical capabilities of machines.
8 Guiding Questions for Implementing Machine Learning
in Drug Design
Before we conclude, the following set of questions is designed to guide the effective
use of machine learning in drug design. These questions cover various aspects,
including the primary objectives, data availability and quality, model selection,
performance evaluation, and integration into the drug development process. Additionally, they address challenges related to data preprocessing, model interpretation,
and resource requirements. By thoroughly addressing these questions, researchers
and practitioners can ensure that their machine learning projects in drug design are
well founded, robust, and capable of overcoming specific challenges in this field.
Guiding questions for implementing machine learning in drug design
What is the primary objective of using machine learning in drug design?
Identification of new molecules?
Optimization of existing molecules?
Prediction of efficacy and toxicity?
Target identification/optimization?
What data are available for training machine learning models?
Chemical structures?
Biological assay data?
Genomic data?
What are the quality and quantity of this data?
Is there enough data for effective training?
Is the data well annotated and standardized?
(continued)

4 Machine Learning and Neural Network Methods Applied to Drug Discovery 101
What types of machine learning models will be used?
Supervised learning models (e.g., regression and classification)?
Unsupervised learning models (e.g., and dimensionality reduction)?
Reinforcement learning models?
How will the performance of the models be evaluated?
What performance metrics will be used (e.g., accuracy, recall, and AUC-ROC)?
What validation techniques will be applied (e.g., cross-validation and data split)?
How will imbalanced data be handled, if present?
Resampling techniques (e.g., oversampling and undersampling)?
How will the results of the models be interpreted?
Feature importance analysis?
Model explanations and visualizations?
How will machine learning models be integrated into the drug design process?
Integration with virtual screening platforms?
Integration with in vitro and in vivo experimentation pipelines?
What are the main challenges and limitations expected in the application of machine
learning to drug design?
Interpretability issues?
Generalization to new data?
How will the lifecycle of machine learning models be managed?
Model updates and retraining?
Post-implementation performance monitoring?
What computational resources will be needed?
Required hardware and software?
Use of GPUs or TPUs to accelerate training?
How will the impact of machine learning models on the drug development process be
measured?
Cost and time reduction?
Increased success rate in clinical trials?
9 Conclusions
In conclusion, the exploration of ML within the realm of AI offers a transformative
lens through which we understand, navigate, and revolutionize complex processes
like drug discovery. As we traverse the intricate stages of problem definition, model
selection, training, and validation, it becomes evident that the success of ML
applications hinges on the careful consideration of various factors, including problem specificity and algorithm selection. However, with the power of ML, especi ally
within the domain of neural networks and deep learning, we witness a paradigm shift
in how we approach challenges in drug discovery.
Yet, this transformative journey is not without its challenges. The omnipresence
of biases, the delicate balance between overfitting and underfitting, the imperative

102 D. S. de Sousa et al.
need for interpretability, computational costs, data dependencies, and the quest for
robust models all contr ibute to the nuanced landscape of ML in drug discovery. Each
limitation serves as a reminder of the ethical responsibilities, technical intricacies,
and economic considerations that accompany the integration of ML into such critical
domains.
The application of ML in drug discovery, exemplified through target identification, lead discovery, and preclinical and clinical development, underscores the
potential for accelerated advancements in the pharmaceutical industry. The careful
orchestration of various ML algorithms, from classifiers to regression models,
empowers researchers to sift through vast data sets and extract meaningful insights
that shape the trajectory of drug development.
As we navigate through the challenges and leverage the emerging trends of
transfer learning, explainable AI, and the integration of imaging data and molecular
structures, we glimpse into a future where the marriage between biology and
artificial intelligence paves the way for unprecedented innovation. The continued
advancements in deep learning and the collaborative synergy between human
expertise and machine analytical capabilities promise not just efficiency but also a
profound understanding of the intricate interplay between therapeutic compounds
and biological targets.
In essence, the union of ML and drug discovery heralds a promising era where
scientific exploration is propelled by the fusion of human intellect and the computational prowess of machines. As we stand at the cross roads of science and technology, the journey forward promises not just accelerated drug development but a
deeper, more personalized understanding of therapies, fostering a landscape where
innovation thrives at the intersection of AI and the intricacies of life sciences.
References
1. Gelijns, A. C. (2014). Technological innovation: Comparing development of drugs, devices,
and procedures in medicine. National Academies Press.
2. Taye, M. M. (2023). Understanding of machine learning with deep learning: Architectures,
workflow, applications and future directions. Computers, 12(5), 91.
3. Niazi, S. K., & Mariam, Z. (2023). Recent advances in machine-learning-based
chemoinformatics: A comprehensive review. International Journal of Molecular Sciences,
24(14), 11488.
4. Vora, L. K., Gholap, A. D., Jetha, K., Thakur, R. R. S., Solanki, H. K., & Chavda, V. P.
(2023). Artificial intelligence in pharmaceutical technology and drug delivery design.
Pharmaceutics, 15(7), 1916.
5. Askr, H., Elgeldawi, E., Aboul Ella, H., Elshaier, Y. A. M. M., Gomaa, M. M., & Hassanien,
A. E. (2023). Deep learning in drug discovery: An integrative review and future challenges.
Artificial Intelligence Review, 56(7), 5975–6037.
6. McCulloch, W. S., & Pitts, W. (1943). A logical calculus of the ideas immanent in nervous
activity. The Bulletin of Mathematical Biophysics, 5, 115–133.
7. Weik, M. H. (1961). The ENIAC story. Ordnance, 45(244), 571–575.
8. Turing, A. M. (1950). I.—Computing machinery and intelligence. Mind, LIX(236), 433– 460.

4 Machine Learning and Neural Network Methods Applied to Drug Discovery 103
9. McCarthy, J., & Feigenbaum, E. A. (1990). In memoriam: Arthur samuel: Pioneer in machine
learning. AI Magazine, 11(3), 10.
10. Samuel, A. L. (1959). Some studies in machine learning using the game of checkers. IBM
Journal of Research and Development, 3(3), 210–229.
11. Rosenblatt, F. (1957). The perceptron, a perceiving and recognizing automaton Project Para.
Cornell Aeronautical Laboratory.
12. Cover, T., & Hart, P. (1967). Nearest neighbor pattern classification. IEEE Transactions on
Information Theory, 13(1), 21–27.
13. E. A. Feigenbaum, B. G. Buchanan, and J. Lederberg, 1970. “On generality and problem
solving: A case study using the DENDRAL program,”
14. Vamathevan, J., et al. (2019). Applications of machine learning in drug discovery and
development. Nature Reviews. Drug Discovery, 18(6), 463–477.
15. Dara, S., Dhamercherla, S., Jadav, S. S., Babu, C. H. M., & Ahsan, M. J. (2022). Machine
learning in drug discovery: A review. Artificial Intelligence Review, 55(3), 1947–1999.
16. Zhang, L., Tan, J., Han, D., & Zhu, H. (2017). From machine learning to deep learning:
Progress in machine intelligence for rational drug discovery. Drug Discovery Today, 22(11),
1680–1685.
17. Stephenson, N., et al. (2019). Survey of machine learning techniques in drug discovery.
Current Drug Metabolism, 20(3), 185–193.
18. Popova, M., Isayev, O., & Tropsha, A. (2018). Deep reinforcement learning for de novo drug
design. Science Advances, 4(7), eaap7885.
19. Ajay, Walters, W. P., & Murcko, M. A. (1998). Can we learn to distinguish between ‘druglike’ and ‘nondrug-like’ molecules? Journal of Medicinal Chemistry, 41(18), 3314–3324.
20. Carracedo-Reboredo, P., et al. (2021). A review on machine learning approaches and trends in
drug discovery. Computational and Structural Biotechnology Journal, 19, 4538–4558.
21. Irwin, J. J., & Shoichet, B. K. (2005). ZINC- a free database of commercially available
compounds for virtual screening. Journal of Chemical Information and Modeling, 45(1),
177–182.
22. Wang, Y., Xiao, J., Suzek, T. O., Zhang, J., Wang, J., & Bryant, S. H. (2009). PubChem: A
public information system for analyzing bioactivities of small molecules. Nucleic Acids
Research, 37(suppl_2), W623–W633.
23. Wishart, D. S., et al. (2008). DrugBank: A knowledgebase for drugs, drug actions and drug
targets. Nucleic Acids Research, 36(suppl_1), D901–D906.
24. Gaulton, A., et al. (2012). ChEMBL: A large-scale bioactivity database for drug discovery.
Nucleic Acids Research, 40
25. Wallach, I., Dzamba, M., & Heifets, A. (2015). AtomNet: a deep convolutional neural network
for bioactivity prediction in structure-based drug discovery. arXiv preprint arXiv:1510.02855.
26. Stokes, J. M., et al. (2020). A deep learning approach to antibiotic discovery. Cell, 180(4),
688–702.
27. Burki, T. (2020). A new paradigm for drug development. Lancet Digit Health, 2(5), e226–
e227.
28. Whitby, B. (2009). Artificial intelligence. The Rosen Publishing Group, Inc.
29. Winston, P. H. (1984). Artificial intelligence. Addison-Wesley Longman Publishing, Inc.
30. Ramesh, A. N., Kambhampati, C., Monson, J. R. T., & Drew, P. J. (2004). Artificial
intelligence in medicine. Annals of the Royal College of Surgeons of England, 86(5), 334.
31. Hunt, E. B. (2014). Artificial intelligence. Academic.
32. El Naqa, I., & Murphy, M. J. (2015). What is machine learning? Springer.
33. Rusk, N. (2016). Deep learning. Nature Methods, 13(1), 35.
34. LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. Nature, 521(7553), 436–444.
35. Bonaccorso, G. (2017). Machine learning algorithms. Packt Publishing Ltd.
36. Mahesh, B. (2020). Machine learning algorithms-a review. International Journal of Science
and Research (IJSR), 9(1), 381–386.
37. Zhou, Z.-H. (2021). Machine learning. Springer Nature.
(D1), D1100–D1107.

104 D. S. de Sousa et al.
38. Agatonovic-Kustrin, S., & Beresford, R. (2000). Basic concepts of artificial neural network
(ANN) modeling and its application in pharmaceutical research. Journal of Pharmaceutical
and Biomedical Analysis, 22(5), 717–727.
39. Morales, E. F., & Escalante, H. J. (2022). A brief introduction to supervised, unsupervised, and
reinforcement learning. In Biosignal processing and classi fication using computational learn-
ing and intelligence (pp. 111–129). Elsevier.
40. Rigatti, S. J. (2017). Random forest. Journal of Insurance Medicine, 47(1), 31–39.
41. Lorena, A. C., & De Carvalho, A. C. (2007). Uma introdução às support vector machines.
Revista de Informática Teórica e Aplicada, 14(2), 43–67.
42. Siemers, F. M., & Bajorath, J. (2023). Differences in learning characteristics between support
vector machine and random forest models for compound classi fication revealed by Shapley
value analysis. Scientific Reports, 13(1), 5983.
43. Rosenblatt, F., & Papert, S. (2021). Perceptron, 9.
44. Gallant, S. I. (1990). Perceptron-based learning algorithms. IEEE Transactions on Neural
Networks, 1(2), 179–191.
45. Ruck, D. W., Rogers, S. K., & Kabrisky, M. (1990). Feature selection using a multilayer
perceptron. Journal of Neural Network Computing, 2(2), 40–48.
46. Popescu, M.-C., Balas, V. E., Perescu-Popescu, L., & Mastorakis, N. (2009). Multilayer
perceptron and neural networks. WSEAS Transactions on Circuits and Systems, 8(7),
579–588.
47. Sharma, S., Sharma, S., & Athaiya, A. (2017). Activation functions in neural networks.
Towards Data Science, 6(12), 310–316.
48. A. D. Rasamoelina, F. Adjailia, and P. Sinčák, “A review of activation function for artificial
neural network,” in 2020 IEEE 18th World Symposium on Applied Machine Intelligence and
Informatics (SAMI), IEEE, 2020, pp. 281–286.
49. Leijnen, S., & van Veen, F. (2020). The neural network zoo. Proceedings of the West Mark
Educational Association conference, 47(1).
50. Hassoun, M. H. (1995). Fundamentals of artifi cial neural networks. MIT Press.
51. O’Shea, K., & Nash, R. (2015). An introduction to convolutional neural networks. arXiv
preprint arXiv:1511.08458.
52. Chauhan, R., Ghanshala, K. K., & Joshi, R. C. (2018). Convolutional neural network (CNN)
for image detection and recognition. In 2018 first international conference on secure cyber
computing and communication (ICSCCC)
53. Albawi, S., Mohammed, T. A., & Al-Zawi, S. (2017). Understanding of a convolutional neural
network. In 2017 international conference on engineering and technology (ICET)
(pp. 1–6). IEEE.
54. Grossberg, S. (2013). Recurrent neural networks. Scholarpedia, 8(2), 1888.
55. Medsker, L. R., & Jain, L. C. (2001). Recurrent neural networks. Design and Applications,
5(64–67), 2.
56. Medsker, L., & Jain, L. C. (1999). Recurrent neural networks: Design and applications. CRC
Press.
57. Nosouhian, S., Nosouhian, F., & Khoshouei, A. K. (2021). A review of recurrent neural
network architecture for sequence learning: Comparison between LSTM and GRU.
58. Sherstinsky, A. (2020). Fundamentals of recurrent neural network (RNN) and long short-term
memory (LSTM) network. Physica D, 404, 132306.
59. Dey, R., & Salem, F. M. (2017). Gate-variants of gated recurrent unit (GRU) neural networks.
In 2017 IEEE 60th international Midwest symposium on circuits and systems (MWSCAS)
(pp. 1597–1600). IEEE.
60. Martín-Guerrero, J. D., & Lamata, L. (2022). Quantum machine learning: A tutorial.
Neurocomputing, 470, 457–461.
61. Chao, W.-L. (2011). Machine learning tutorial. Digital Image and Signal Processing.
62. Ball, N. M., & Brunner, R. J. (2010). Data mining and machine learning in astronomy.
International Journal of Modern Physics D, 19(07), 1049–1106.
(pp. 278–282). IEEE.

4 Machine Learning and Neural Network Methods Applied to Drug Discovery 105
63. Emmanuel, T., Maupong, T., Mpoeleng, D., Semong, T., Mphago, B., & Tabona, O. (2021). A
survey on missing data in machine learning. Journal of Big Data, 8(1), 1–37.
64. Kumar, S. A., et al. (2022). Machine learning and deep learning in data-driven decision
making of drug discovery and challenges in high-quality data acquisition in the pharmaceutical
industry. Future Medicinal Chemistry, 14(4), 245–270.
65. Ridzuan, F., & Zainon, W. M. N. W. (2019). A review on data cleansing methods for big data.
Procedia Computer Science, 161, 731–738.
66. Ali, P. J. M., Faraj, R. H., Koya, E., Ali, P. J. M., & Faraj, R. H. (2014). Data normalization
and standardization: A technical report. Machine Learning Technical Report, 1(1), 1–6.
67. Potdar, K., Pardawala, T. S., & Pai, C. D. (2017). A comparative study of categorical variable
encoding techniques for neural network classifiers. International Journal of Computers and
Applications, 175(4), 7–9.
68. Kurita, T. (2019). Principal component analysis (PCA). In Computer vision: A reference guide
(pp. 1–4).
69. Wattenberg, M., Viégas, F., & Johnson, I. (2016). How to use t-SNE effectively. Distill,
1(10), e2.
70. Chawla, N. V., Bowyer, K. W., Hall, L. O., & Kegelmeyer, W. P. (2002). SMOTE: Synthetic
minority over-sampling technique. Journal of Artificial Intelligence Research, 16, 321–357.
71. Barandela, R., Valdovinos, R. M., Sánchez, J. S., & Ferri, F. J. The imbalanced training
sample problem: Under or over sampling? In Structural, Syntactic, and Statistical Pattern
Recognition: Joint IAPR International Workshops, SSPR 2004 and SPR 2004, Lisbon, Portugal, August 18-20, 2004. Proceedings (pp. 806–814). Springer, 2004.
72. Cai, J., Luo, J., Wang, S., & Yang, S. (2018). Feature selection in machine learning: A new
perspective. Neurocomputing, 300,70–79.
73. Raschka, S. (2018). Model evaluation, model selection, and algorithm selection in machine
learning. arXiv preprint arXiv:1811.12808.
74. da Silva, A. P., Chiari, L. P. A., Guimaraes, A. R., Honorio, K. M., & da Silva, A. B. F. (2021).
Drug design of new 5-HT6R antagonists aided by artificial neural networks. Journal of
Molecular Graphics & Modelling, 104, 107844.
75. Kim, Y.-J., & Chi, M. (2018). Temporal belief memory: Imputing missing data during RNN
training. In Proceedings of the 27th international joint conference on artificial intelligence
(IJCAI-2018).
76. Christoffersen, P., & Jacobs, K. (2004). The importance of the loss function in option
valuation. Journal of Financial Economics, 72(2), 291–318.
77. Wythoff, B. J. (1993). Backpropagation neural networks: A tutorial. Chemometrics and
Intelligent Laboratory Systems, 18(2), 115–155.
78. Devarakonda, A., Naumov, M., & Garland, M. (2017). Adabatch: Adaptive batch sizes for
training deep neural networks. arXiv preprint arXiv:1712.02029.
79. Humphrey, A., et al. (2022). Machine-learning classifi
mating F1-score in the absence of ground truth. Monthy Notices of the Royal Astronomical
Society Letters, 517(1), L116–L120.
80. Kuhn, M. (2014). Futility analysis in the cross-validation of machine learning models. arXiv
preprint arXiv:1405.6974.
81. Roy, K., Das, R. N., Ambure, P., & Aher, R. B. (2016). Be aware of error measures. Further
studies on validation of predictive QSAR models. Chemometrics and Intelligent Laboratory
Systems, 152,18–33.
82. Rücker, C., Rücker, G., & Meringer, M. (2007). Y-randomization and its variants in QSPR/
QSAR. Journal of Chemical Information and Modeling, 47(6), 2345–2357.
83. Narkhede, S. (2018). Understanding auc-roc curve. Towards Data Science, 26(1), 220–227.
84. Feurer, M., & Hutter, F. (2019). Hyperparameter optimization. In Automated machine learn-
ing: Methods, systems, challenges (pp. 3–33).
85. Bergstra, J., Bardenet, R., Bengio, Y., & Kégl, B. (2011). Algorithms for hyper-parameter
optimization. In Advances in neural information processing systems (Vol. 24).
cation of astronomical sources: Esti-

106 D. S. de Sousa et al.
86. Kim, Y., & Chung, M. (2019). An approach to hyperparameter optimization for the objective
function in machine learning. Electronics (Basel), 8(11), 1267.
87. Kadhim, Z. S., Abdullah, H. S., & Ghathwan, K. I. (2022). Artificial neural network
hyperparameters optimization: A survey. International Journal of Online & Biomedical
Engineering, 18(15), 59.
88. Mazilu, S., & Iria, J. (2011). L1 vs. l2 regularization in text classification when learning from
labeled features. In 2011 10th international conference on machine learning and applications
and workshops (pp. 166–171). IEEE.
89. Kuhn, M., Johnson, K., Kuhn, M., & Johnson, K. (2013). Over-fitting and model tuning. In
Applied predictive modeling (pp. 61–92).
90. Probst, P., Wright, M. N., & Boulesteix, A. (2019). Hyperparameters and tuning strategies for
random forest. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 9(3),
e1301.
91. Ghawi, R., & Pfeffer, J. (2019). Efficient hyperparameter tuning with grid search for text
categorization using kNN approach with BM25 similarity. Open Computer Science, 9(1),
160–180.
92. Friedrichs, F., & Igel, C. (2005). Evolutionary tuning of multiple SVM parameters.
Neurocomputing, 64, 107–117.
93. Mastorakis, G. (2018). Human-like machine learning: limitations and suggestions. arXiv
preprint arXiv:1811.06052.
94. Malik, M. M. (2020). A hierarchy of limitations in machine learning. arXiv preprint
arXiv:2002.05193.
95. Kenge, R. (2020). Machine learning, its limitations, and solutions over IT. International
Journal of Information Technology Modeling and Computing, 11, 73.
96. Jabbar, H., & Khan, R. Z. (2015). Methods to avoid over- fitting and under-fitting in supervised
machine learning (comparative study). Computer Science, Communication and Instrumenta-
tion Devices, 70(10.3850), 978–981.
97. Bilmes, J. (2020). Underfitting and overfitting in machine learning (UW ECE course notes)
(Vol. 5).
98. Ying, X. (2019). An overview of overfitting and its solutions. In Journal of physics: Confer-
ence series (p. 022022). IOP Publishing.
99. Lavecchia, A. (2015). Machine-learning approaches in drug discovery: Methods and applications. Drug Discovery Today, 20(3), 318–331.
100. Carvalho, D. V., Pereira, E. M., & Cardoso, J. S. (2019). Machine learning interpretability: A
survey on methods and metrics. Electronics (Basel), 8(8), 832.
101. Gramegna, A., & Giudici, P. (2021). SHAP and LIME: An evaluation of discriminative power
in credit risk. Frontiers in Artificial Intelligence, 4, 752558.
102. Justus, D., Brennan, J., Bonner, S., & McGough, A. S. (2018). Predicting the computational
cost of deep learning models. In
(pp. 3873–3882). IEEE.
103. Miller, T., et al. (2023). Advancements in artificial intelligence circuits and systems (AICAS).
Electronics (Basel), 13(1), 102.
104. Elbadawi, M., Gaisford, S., & Basit, A. W. (2021). Advanced machine-learning techniques in
drug discovery. Drug Discovery Today, 26(3), 769–777.
105. Koscielny, G., et al. (2017). Open targets: A platform for therapeutic target identification and
validation. Nucleic Acids Research, 45(D1), D985–D994.
106. Yang, H., et al. (2019). admetSAR 2.0: Web-service for prediction and optimization of
chemical ADMET properties. Bioinformatics, 35(6), 1067–1069.
107. Advanced Chemistry Development Inc. (2016). (ACD/Labs) ACD/PERCEPTA Version 2015.
Frankfurt am Main. www.acdlabs.com/pka
108. Pires, D. E. V., Blundell, T. L., & Ascher, D. B. (2015). pkCSM: Predicting small-molecule
pharmacokinetic and toxicity properties using graph-based signatures. Journal of Medicinal
Chemistry, 58(9), 4066–4072.
2018 IEEE international conference on big data (Big Data)

4 Machine Learning and Neural Network Methods Applied to Drug Discovery 107
109. Schwaller, P., Strobelt, H., Laino, T., & Hoover, B. IBM RXN for Chemistry: Unveiling the
grammar of the organic chemistry language.
110. Ramsundar, B. et al., DeepChem: Democratizing deep-learning for drug discovery, quantum
chemistry, materials science and biology.
111. Landrum, G. (2013). RDKit: A software suite for cheminformatics, computational chemistry,
and predictive modeling. Greg Landrum, 8, 31.
112. Dixon, S. L., Duan, J., Smith, E., Von Bargen, C. D., Sherman, W., & Repasky, M. P. (2016).
AutoQSAR: An automated machine learning tool for best-practice quantitative structure–
activity relationship modeling. Future Medicinal Chemistry, 8(15), 1825–1839.
113. Yap, C. W. (2011). PaDEL-descriptor: An open source software to calculate molecular
descriptors and fingerprints. Journal of Computational Chemistry, 32(7), 1466–1474.
114. Sybyl, X. (2011). Molecular modelling software (Vol. 1). Tripos Certara.
115. Matlab, S. (2012). Matlab. The MathWorks.
116. The UniProt Consortium. (2017). UniProt: The universal protein knowledgebase. Nucleic
Acids Research, 45(D1), D158–D169.
117. Burley, S. K., Berman, H. M., Kleywegt, G. J., Markley, J. L., Nakamura, H., & Velankar,
S. (2017). Protein Data Bank (PDB): The single global macromolecular structure archive. In
Protein crystallography: Methods and protocols (pp. 627–641).
118. Avram, S., et al. (2023). DrugCentral 2023 extends human clinical data and integrates
veterinary drugs. Nucleic Acids Research, 51(D1), D1276–D1287.
119. Cai, M.-C., et al. (2015). ADReCS: An ontology database for aiding standardization and
hierarchical classification of adverse drug reaction terms. Nucleic Acids Research, 43(D1),
D907–D913.
120. Davis, A. P., Wiegers, T. C., Johnson, R. J., Sciaky, D., Wiegers, J., & Mattingly, C. J. (2023).
Comparative Toxicogenomics database (CTD): Update 2023. Nucleic Acids Research,
51(D1), D1257–D1262.
121. Piñero, J., et al. (2015). DisGeNET: A discovery platform for the dynamical exploration of
human diseases and their genes. Database, 2015, bav028.
122. UK Biobank. (2014). About UK biobank.
123. Szklarczyk, D., et al. (2023). The STRING database in 2023: Protein–protein association
networks and functional enrichment analyses for any sequenced genome of interest. Nucleic
Acids Research, 51(D1), D638–D646.
124. Zhou, Y., et al. (2024). TTD: Therapeutic target database describing target druggability
information. Nucleic Acids Research, 52(D1), D1465–D1477.
125. Gallo, K., Goede, A., Eckert, A., Moahamed, B., Preissner, R., & Gohlke, B.-O. (2021).
PROMISCUOUS 2.0: A resource for drug-repositioning. Nucleic Acids Research, 49(D1),
D1373–D1380.
126. Zarin, D. A., Tse, T., Williams, R. J., Califf, R. M., & Ide, N. C. (2011). The ClinicalTrials.
Gov results database—Update and key issues. New England Journal of Medicine, 364(9),
852–860.
127. Gonoskov, A., Wallin, E., Polovinkin, A., & Meyerov, I. (2019). Employing machine learning
for theory validation and identification of experimental conditions in laser-plasma physics.
Scientific Reports, 9(1), 7043.
128. Vayena, E., Blasimme, A., & Cohen, I. G. (2018). Machine learning in medicine: Addressing
ethical challenges. PLoS Medicine, 15(11), e1002689.
129. Cai, C., et al. (2020). Transfer learning for drug discovery. Journal of Medicinal Chemistry,
63(16), 8683–8694.
130. Ponzoni, I., Páez Prosper, J. A., & Campillo, N. E. (2023). Explainable artificial intelligence:
A taxonomy and guidelines for its application to drug discovery. Wiley Interdisciplinary
Reviews: Computational Molecular Science, 13(6), e1681.
131. Bian, Y., & Xie, X.-Q. (2021). Generative chemistry: Drug discovery with deep learning
generative models. Journal of Molecular Modeling, 27,1–18.

Chapter 5
Clustering of Small Molecules
Alan Talevi, Lucas Alberca, and Carolina Bellera
Abstract Clustering of small molecules finds a diversity of applications in chem-
istry and, in particular, in the fields of cheminformatics and drug discovery. It may be
used directly as an unsupervised machine-learning strategy to identify existing
patterns in a chemical data set or libraries or integrated into supervised machinelearning studies to partition a sample of compounds into representative subsamples
(e.g., training and validation data). It may also be applied to select which in silico hits
from a virtual screening campaign will be submitted to experimental confirmation, or
to define which hits emerging from a “wet” screening campaign will be prioritized
for further development or characterization. Here, we review general strategies to
validate the output of a clustering algorithm and discuss current challenges and
possible future directions in the field of small molecule clustering.
Keywords Cluster replication · Clustering, External validity · Hierarchical
clustering · Internal validity measures · Multi-view clustering · Non-hierarchic al
clustering · Relative validity · Small molecules · Stability measures · Subspace
clustering · Supervised learning · Unsupervised learn ing · Validation · Validity
A. Talevi (✉)
Laboratory of Bioactive Compound Research and Development (LIDeB), Faculty of Exact
Sciences, University of La Plata (UNLP), Buenos Aires, Argentina
Consejo Nacional de Investigaciones Científicas y Técnicas (CONICET), CCT La Plata, La
Plata, Buenos Aires, Argentina
Boolzi SA, Buenos Aires, Argentina
e-mail: alantalevi@gmail.com
L. Alberca · C. Bellera
Laboratory of Bioactive Compound Research and Development (LIDeB), Faculty of Exact
Sciences, University of La Plata (UNLP), Buenos Aires, Argentina
Consejo Nacional de Investigaciones Científicas y Técnicas (CONICET), CCT La Plata, La
Plata, Buenos Aires, Argentina
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2024
V. G. Maltarollo (ed.), Computer-Aided and Machine Learning-Driven Drug
Design, Computer-Aided Drug Discovery and Design 3,
https://doi.org/10.1007/978-3-031-76718-0_5
109

110 A. Talevi et al.
1 Historical Background
Human beings have developed the ability to detect higher-level organizations and
group objects and experiences into functional categories, which is a precondition of
sophisticated thought [45]. Categorization provides the basis to recognize and
respond instantly and appropriately to objects and situations even if they are not
identical to the ones that we have previously met. In other words, categorization is a
way of anticipating behavior based on prior experience with similar objects or
scenarios. This is a central proposition to understand the application of the
approaches described in this chapter.
Perceptual grouping is performed at a remarkable speed based on principles of
proximity (objects that are close to each other in space are perceived as part of a
group), similarity (grouping elements by shape, size, and color), continuity
(a relation is naturally perceived between objects that are arranged in continuous
lines or curves), and common fate (objects that move at the same speed and/or
direction tend to be grouped), among others [9]. The tendency of human beings to
group things together is reflected in their language: common nouns name people,
animals, or things of the same class or species, whereas proper nouns name particular
objects individualized from any other of the same species (a particular person, a
particular animal, a particular city, etc.). In fact, categories of knowledge that have a
common noun in a language are more efficiently captured by the eye than those that
do not (which is known as conceptual grouping)[31]. Moreover, collective nouns
express functional and/or proximity relationships between similar objects (named by
the same common noun), for example, a pack of wolves, a crowd, a school of fish,
and a choir.
A fundamental challenge in the field of artificial inte lligence is how to teach a
machine to perform grouping or categorization as good as (or better than) evolution
has taught human beings to do [41]. Let us take, for instance, the molecular
structures in Fig. 5.1. Even if the observer has no specific training in chemistry, it
is likely that they will perceive the similarities between the two compounds that
share the benzodiazepine nucleus. However, when dealing with large-scale chemical
data sets or repositories, visually analyzing and comparing every possible pair of
molecules to produce a hypothesis on possible clusters that may be derived from it
Fig. 5.1 The human brain performs perceptual grouping at a remarkable speed based on principles
such as similarity. It is likely that even people with no training in chemistry will cluster the three
compounds that share the benzodiazepine nucleus (alprazolam, diazepam, and lorazepam) apart
from the corticosteroid (prednisolone)
Соседние файлы в папке Библиотека им академика М.И. Перельмана
