Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5364_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Foreword
- •Acknowledgments
- •Contents
- •1.1 Structure-Based Drug Discovery (SBDD)
- •1.2 Ligand-Based Drug Design (LBDD)
- •1.3 Echoes from the Past, Visions from the Future
- •References
- •1 Introduction
- •2.2 Second Step: Data Curation
- •2.4 Fourth Step: Updating and Maintenance
- •2 Databases and Curation
- •8 Perspectives
- •9 Conclusion
- •References
- •1 Introduction
- •2.1 Making and Matching Protein Models
- •2.2 Simulating Protein Movements
- •2.3 Analyzing Changes in Protein Shape
- •3 Pharmacogenomics in Drug Development
- •4 Case Studies of Genomics-Based Drug Design
- •References
- •1 Historical Background
- •1.1 Timeline
- •2 Methodology Overview
- •2.1 Neural Networks
- •2.1.1 Perceptron
- •2.1.2 Multilayer Neural Networks
- •2.1.3 Types of Neural Networks
- •Feedforward
- •Recurrent Neural Networks
- •LSTM
- •2.2 Deep Learning
- •3 Using Machine Learning
- •3.2 Data Collection
- •3.3 Data Preprocessing
- •3.4 Model Selection
- •3.5 Model Training
- •3.6 Validation
- •3.7 Tuning
- •3.8 Prediction
- •4 Limitations
- •4.1 Bias
- •4.3 Interpretability
- •4.4 Computational Cost
- •4.5 Data Dependency
- •4.6 Robustness
- •5 Applications in Drug Discovery
- •5.2 Lead Discovery
- •5.3 Preclinical and Clinical Development
- •6 Resources and Tools
- •7 Challenges and Perspectives
- •7.1 Future Trends
- •9 Conclusions
- •References
- •1 Historical Background
- •1.1 Applications in Drug Discovery
- •2 Validations and Controls
- •2.1 Internal Validation
- •2.2 External Validation
- •2.3 Relative Cluster Validation
- •3 Challenges and Perspectives
- •4 Conclusions
- •References
- •1 Historical Background
- •2 OECD Principles
- •2.1 A Defined Endpoint
- •2.2 An Unambiguous Algorithm
- •2.5 A Mechanistic Interpretation, if Possible
- •3 Software and Tools
- •4 Validations and Controls
- •4.1 Internal and External Validation
- •4.1.1 Regression Metrics
- •4.2 Applicability Domain
- •4.3 Randomization Tests
- •5 Interpretation
- •6 Practical Advice During QSAR Modeling
- •7 Application
- •8 Challenges and Perspectives
- •References
- •1 Molecular Docking
- •2 Advances in Scoring Functions and Search Algorithms
- •2.2 Critical Characteristics of Search Algorithms
- •2.3 Docking Programs and Scoring Functions
- •3 Calculations Performed During Docking Simulations
- •4 Essential Components for a Good Docking Program
- •5 Limitations of the Docking Technique
- •6 Validation of Docking Results
- •7 Inappropriate Use of Validation Methods in Docking
- •9 Use of Machine Learning in Molecular Docking
- •11 Challenges
- •12 Conclusions
- •References
- •3 System Preparation for MD Simulations
- •3.1 Solvation and Microensemble
- •3.2 Force Fields: General Concept and Relevant Choices
- •3.3 The Concept of Replicas and Timescale
- •4.1.2 Protein Root Mean Square Fluctuation (RMSF)
- •4.1.4 Protein Secondary Structure Analysis
- •4.1.5 Principal component Analysis (PCA)
- •4.1.6 Markov State Modelling
- •4.1.7 Distance Calculations
- •4.1.8 Angle and Plane Calculations
- •4.2.2 Distances and Ligand-Induced Geometry Rearrangements
- •4 Molecular Dynamics Analysis
- •4.1 Protein Perspective
- •4.1.1 Protein Root Mean Square Deviation (RMSD)
- •4.3 Ligand Perspective
- •4.3.1 Ligand Properties
- •4.3.2 Ligand Root Mean Square Deviation
- •4.3.3 Ligand Root Mean Square Fluctuation
- •4.3.4 Angles and Dihedrals
- •5.1 Protein Structure Prediction and Preparation
- •5.2 Molecular Docking
- •6 Concluding Remarks and Outlook
- •Glossary
- •References
- •1 Introduction
- •2.1 MDeNM
- •2.2 Collective Molecular Dynamics (coMD)
- •2.3 ClustENM and ClustENMD
- •3 Ensemble Docking
- •References
- •1 Introduction
- •1.1 Advantages, Disadvantages, Innovations, and Challenges
- •1.2 Recent Advances in Accessible FEP Software Tools
- •1.3 Applications of FEP in Industry and Consortiums
- •2 Expanding the Potential of FEP Calculations
- •2.1 Validating Binding Poses
- •2.2 Dealing with Solvent
- •2.3 FEP and Allostery
- •2.4 FEP and Covalent Ligands
- •2.5 Applications of FEP in Scaffold Hopping
- •2.6 Positional Analogue Scanning
- •2.7 Combinations and Alternative Approaches
- •3 Machine Learning for FEP
- •3.4 Implications for ML in FEP Calculations
- •4 Final Considerations
- •5 First Steps to FEP Simulations
- •References
- •1 Background
- •2 Ultra-Large Screening Libraries and Chemical Spaces
- •3.1 Implications of Dataset Size
- •4 Ligands on the Ultra-Large Scale
- •4.1 Ultra-Large 2D Similarity Searches
- •7 Challenges and Future Perspectives
- •7.1 Hit Triage: An Old Problem on a New Dimension
- •8 Conclusions
- •Appendix
- •References
- •1 Introduction
- •2 Enzymatic Activity Evaluations
- •3 Cytotoxicity Evaluation and Cell Viability
- •4 Antiviral Assays in Experimental Validation
- •6 In Vivo Evaluation of Compounds
- •7 Conclusions
- •References
- •1 Introduction
- •3.1 Data Collection
- •3.2 Data Preprocessing
- •3.4 Model Choice
- •3.5 Model Training
- •3.6 Model Assessment
- •3.7 External Validation
- •3.8 Implementation and Availability
- •3.9 Continuous Update
- •5 Conclusions and Perspectives
- •References
- •1 Experimental Approaches to Obtain Protein Structure
- •1.1 X-Ray Crystallography
- •1.2 Nuclear Magnetic Resonance
- •1.3 Cryo-EM
- •1.4 Hybrid Methods
- •2 Modeling Approaches to Obtain Protein Structure
- •2.1 Homology Modeling
- •2.2 Ab Initio Modeling
- •2.3 New Approaches
- •3 Conformational Diversity of Proteins
- •3.1 Characterization of Protein Conformational States
- •3.2 Experimental Methods to Study Protein Dynamics and Conformations
- •3.4 Molecular Dynamics Simulation
- •3.5 Sampling Strategies
- •4 Remarks and Perspectives
- •References
- •1 Introduction
- •2 Structure-Based Drug Design of HIV Protease Inhibitors
- •2.1 HIV-1 Protease as a Therapeutic Target
- •2.2.1 Saquinavir
- •2.2.2 Indinavir
- •2.3.1 Lopinavir
- •2.3.2 Darunavir
- •6 Conclusions
- •References
- •4 Experimental Methods to Analyze NR Activity
- •4.2 Coregulator-Recruitment
- •5 Concluding Remarks and Outlook
- •References

60 G. M. Ferreira et al.
Autonomy, as one of these principles, underscores the importance of securing
free and informed consent from all individuals involved. However, it is crucial to
ensure that autonomy is genuinely free, without defects in consent or omission of
information, which includes addressing issues related to personal vulnerability.
The principles of nonmaleficence and beneficence emphasize a commitment to
improving the quality of life and treatment efficacy while minimizing side effects for
both individual patients and the broader population. This commitment rejects medical approaches based on “trial and error” methods, prioritizing patient well-being
and safety above all else, as highlighted by Brito et al. [50]. The principle of
nonmaleficence, which originates from the Hippocratic oath whose motto is
“primum non nocere,” meaning “first, do not cause damage,” keeps a direct connection with the essential objectives of pharmacogenomics, which strives to reach such
goals, considering that its fundamental purpose is to improve humanity’s quality of
life [50].
In the realm of bioethics, Marcelo Quentin sheds light on the foundational
principle of beneficence, emphasizing its pivotal role in promoting the well-being
of individuals, particularly those who are ill. According to Quentin, beneficence
entails a profound acknowledgment of the moral worth of others, driving healthcare
professionals and society at large to prioritize the prevention of harm and the
enhancement of wellness. This principle compels individuals to meticulously evaluate the potential risks and benefits associated with medical interventions, striving to
maximize the benefits for patients while minimizing any potential harm. Furthermore, beneficence encompasses the complementary principle of “nonmaleficence,”
which underscores the imperative of avoiding actions that may cause harm to the
patient. In essence, Quentin’s perspective underscores the dual nature of ethical
considerations, emphasizing the imperative of balancing benevolence with the
obligation to do no harm [51].
Exploring the principle of justice unveils a complex sphere entangled with social,
political, and economic dynamics, particula rly poignant in the context of developing
countries. Here, tests and treatments linked to genetic data remain prohibitively
expensive, rendering them accessible only to a privileged minority. Remarkably, the
principle of justice emerges as the most intricate in this landscape. The prevailing
economic interests of pharmaceutical giants often prioritize profit maximization over
addressing illnesses affecting marginalized populations, particularly those deemed
unprofitable or lacking potential returns to the system. As noted, treatments for rare
diseases, categorized as “orphan medications,” are sidelined due to their low profitability and lack of investment appeal. Consequently, pharmacogenomics might
inadvertently perpetuate a discriminatory environment, disproportionately excluding
certain social groups, particularly in nations where genetic testing remains financially out of reach for a significant portion of the populace. This predicament poses a
formidable challenge and threatens to undermine the principle of justice in its
entirety.
On the flip side, as underscored by Jorge Alberto Iriart, the emergence of
precision medicine unfolds within the intricate framework of globalized capitalism,
characterized by what Rose terms “economies of vitality.” This concept delineates a

3 A Brief Introduction to Pharmacogenomics and Personalized Medicine in... 61
novel economic domain, known as bioeconomics, wherein biocorporations wield
control over lives, thereby generating value. In this context, the manipulation of
biological data and healthcare interventions becomes instrumental in the accumulation of a distinct form of capital known as biocapital. This perspective sheds light on
the intersection of political and economic forces shaping the trajectory of precision
medicine, highlighting the profound influence of capitalist dynamics on the
healthcare landscape [52].
In simpler terms, pharmacogenomics emerges as a promising solution to fortify
basic and preventive healthcare, potentially slashing costs linked with ineffective
treatments. This perspective holds merit, provided concerns about the affordability
and accessibility of genetic tests for marginalized social groups are adequately
tackled. Nevertheless, within this context, it is cruci al to acknowledge two branches
of the principle of justice: one of utilitarian nature, striving to maximize benefits for
both patients and society and another of egalitarian ethos, dedicated to ensuring
equal value for individuals and equitable opportunities [53].
Based on this premise, compliance with the bioethical principle of justice is
intrinsically connected to the guarantee of egalitarian access and fair opportunities
to all citizens as regarding treatment.
In addition to the bioethical aspect of the matter, the legal aspect is associated to
two fundamental pillars: (i) the patient’s consent regarding the purpose, use, and
disposal of data, and (ii) guaranteed protection of such data against leaks, with
specific and clear guidelines regarding disposal there of. Each country, considering
their legal and regulatory structure in connection with medical secrecy, data protection, and patient’s informed consent, should consider the particularities of such
collections in implementing public policies that regulate the matter.
Genetic data are already collected for miscellaneous purposes worldwide, especially to map genetic diseases at birth, such as the newborn blood spot test. However,
the manner in which such data are collected, stored, and protected is crucial to avoid
breaches of various natures. Thus, the public power is accountable for guaranteeing
safety in the collection, storage, and use of such genetic data.
Hence, pharmacogenomics represents a significant advancement in the healthcare
area. Nevertheless, its ethical and legal application faces complex challenges that go
beyond medicine, involving social, economic, ethical, and legal matters. To guarantee a safe society from the bioethical and legal standpoint, it is essential that law
and bioethics continuously evolve to follow up technological development and
establish solid criteria. This includes the protection of sensitive genetic data, compliance with fundamental bioethical principles, such as autonomy, nonmaleficence,
beneficence, and justice, and the promotion of egalitarian access to treatments.
Moreover, considering the safety and privacy of the genetic data is of the essence
to guarantee that use thereof is consented to and does not breach the fundamental
principles of human dignity. Ultimately, ethics and bioethics play a core role in
providing guidance to such complex matters, promoting a fairer and safer society.

62 G. M. Ferreira et al.
Questions to answer when planning genomics and personalized medicine in drug
design studies?
What is the genomic basis of the disease?
What are the target genes and pathways?
How do genetic variations affect drug response?
What are the biomarkers for disease and drug response?
How can genomics inform drug discovery and development?
What are the ethical, legal, and social implications?
What technologies and methodologies are required?
How can personalized medicine be integrated into clinical practice?
What are the challenges and limitations?
How to ensure collaboration among stakeholders?
References
1. Dere, W. H., & Suto, T. S. (2009). The role of pharmacogenetics and pharmacogenomics in
improving translational medicine. Clinical Cases in Mineral and Bone Metabolism, 6,13–16.
2. Bienfait, K., et al. (2022). Current challenges and opportunities for pharmacogenomics: Perspective of the industry pharmacogenomics working group (I-PWG). Human Genetics, 141,
1165–1173.
3. Ramayanam, N. R., Amarnath, R. N., & Vijayakumar, T. M. (2022). Pharmacogenetic biomarkers and personalized medicine: Upcoming concept in pharmacotherapy. Research Journal
of Pharmacy and Technology, 15, 4289–4292.
4. Castro, K. M., Scheck, A., Xiao, S., & Correia, B. E. (2022). Computational design of vaccine
immunogens. Current Opinion in Biotechnology, 78, 102821.
5. Gourlay, L., Peri, C., Bolognesi, M., & Colombo, G. (2017). Structure and computation in
immunoreagent design: From diagnostics to vaccines. Trends in Biotechnology, 35(12),
1208–1220.
6. Pak, M. A., Markhieva, K. A., Novikova, M. S., Petrov, D. S., Vorobyev, I. S., Maksimova,
E. S., et al. (2023). Using AlphaFold to predict the impact of single mutations on protein
stability and function. PLoS One, 18(3), e0282689.
7. Modeling mutations in protein structures – Feyfant – 2007 – Protein Science – Wiley Online
Library. https://onlinelibrary.wiley.com/doi/10.1110/ps.072855507.
8. Pandurangan, A. P., & Blundell, T. L. (2020). Prediction of impacts of mutations on protein
structure and interactions: SDM, a statistical approach, and mCSM, using machine learning.
Protein Science, 29, 247–257.
9. Zhou, Y., Tremmel, R., Schaeffeler, E., Schwab, M., & Lauschke, V. M. (2022). Challenges
and opportunities associated with rare-variant pharmacogenomics. Trends in Pharmacological
Sciences, 43, 852–865.
10. Schreeck, F., Ahne, G., Tremmel, R., Schaeffeler, E., & Schwab, M. (2022). Pharmacogenomics in pediatric medicine and drug development. Pharmacogenomics, 23, 709–712.
11. Schwab, M., & Schaeffeler, E. (2012). Pharmacogenomics: A key component of personalized
therapy. Genome Medicine, 4,1–3.
12. Moore, T. J., Heyward, J., Anderson, G., & Alexander, G. C. (2020). Variation in the estimated
costs of pivotal clinical benefit trials supporting the US approval of new therapeutic agents,
2015–2017: a cross-sectional study. BMJ Open, 10, e038863.
13. Sun, D., Gao, W., Hu, H., & Zhou, S. (2022). Why 90% of clinical drug development fails and
how to improve it? Acta Pharmaceutica Sinica B, 12, 3049–3062.

3 A Brief Introduction to Pharmacogenomics and Personalized Medicine in... 63
14. Coleman, J. J., & Pontefract, S. K. (2016). Adverse drug reactions. Clinical Medicine, 16,
481–485.
15. Insani, W. N., et al. (2021). Prevalence of adverse drug reactions in the primary care setting:
A systematic review and meta-analysis. PLoS One, 16, e0252161.
16. Ray, S. (2014). Clopidogrel resistance: The way forward. Indian Heart Journal, 66, 530–534.
17. Yin, O., & Vandell, A. (2019). Incorporating pharmacogenomics in drug development. In
Pharmacogenomics (pp. 81–101). Elsevier.
18. Vogel, F. (1959). Moderne Problem Der. Humangenetik, 12.
19. Lander, E. S., et al. (2001). Initial sequencing and analysis of the human genome. Nature, 409,
860–921.
20. Kandi, V., & Vadakedath, S. (2023). Clinical trials and clinical research: A comprehensive
review. Cureus, 15, e35077.
21. Cappuzzo, F., et al. (2005). Epidermal growth factor receptor gene and Protein and Gefitinib
sensitivity in non–small-cell lung cancer. JNCI Journal of the National Cancer Institute, 97,
643–655.
22. Hirsch, F. R., et al. (2006). Molecular predictors of outcome with Gefitinib in a phase III
placebo-controlled study in advanced non–small-cell lung cancer. Journal of Clinical Oncol-
ogy, 24, 5034–5042.
23. Singh, B., Jain, P., Devaraja, K., & Aggarwal, S. (2023). Chapter 3: Pharmacogenomics in drug
discovery and development. In S. A. Ganie, A. Ali, M. U. Rehman, & A. Arafah (Eds.),
Pharmacogenomics (pp. 57–96). Academic.
24. Menden, M. P., et al. (2013). Machine learning prediction of cancer cell sensitivity to drugs
based on genomic and chemical properties. PLoS One, 8, e61318.
25. Lee, K. H., et al. (2016). Genome sequence variability predicts drug precautions and withdrawals from the market. PLoS One, 11, e0162135.
26. Dugger, S. A., Platt, A., & Goldstein, D. B. (2018). Drug development in the era of precision
medicine. Nature Reviews. Drug Discovery, 17, 183–196.
27. Spreafico, R., Soriaga, L. B., Grosse, J., Virgin, H. W., & Telenti, A. (2020). Advances in
genomics for drug development. Genes, 11, 942.
28. Chan, Y.-T., et al. (2022). CRISPR-Cas9 library screening approach for anti-cancer drug
discovery: Overview and perspectives. Theranostics, 12, 3329–3344.
29. Spahn, S., Kleinhenz, F., Shevchenko, E., et al. (2024). The molecular interaction pattern of
lenvatinib enables inhibition of wild-type or kinase-mutated FGFR2-driven
cholangiocarcinoma. Nature Communications, 15, 1287.
30. Emilien, G., Ponchon, M., Caldas, C., & Isacson, O. Impact of genomics on drug discovery and
clinical medicine. QJM.
31. Conejero-Muriel, M., Contreras-Montoya, R., Díaz-Mochón, J. J., de Cienfuegos, L. Á., &
Gavira, J. A. (2015). Protein crystallization in short-peptide supramolecular hydrogels: A
versatile strategy towards biotechnological composite materials. CrystEngComm, 17,
–8078.
8072
32. Patowary, A., et al. (2012). Systematic analysis and functional annotation of variations in the
genome of an Indian individual. Human Mutation, 33, 1133–1140.
33. Tyukavin, A. I., et al. (2021). Biodizine as a civilizational challenge of modern pharmaceuticals.
Pharmacy Formulations, 3, 108–117.
34. Barrot, C.-C., Woillard, J.-B., & Picard, N. (2019). Big data in pharmacogenomics: Current
applications, perspectives and pitfalls. Pharmacogenomics, 20, 609–620.
35. Low, S.-K., Takahashi, A., Mushiroda, T., & Kubo, M. (2014). Genome-wide association
study: A useful tool to identify common genetic variants associated with drug toxicity and
efficacy in cancer pharmacogenomics. Clinical Cancer Research, 20, 2541–2552.
36. Angelbello, A. J., et al. (2018). Using genome sequence to enable the design of medicines and
chemical probes. Chemical Reviews, 118, 1599–1663.
37. Tsourounis, M., Stuart, J., Pignato, W., Toscani, M., & Barone, J. (2015). Current trends in
personalized medicine and companion diagnostics: A summary from the DIA meeting on

64 G. M. Ferreira et al.
personalized medicine and companion diagnostics. Therapeutic Innovation & Regulatory
Science, 49, 530–543.
38. Madian, A. G., Wheeler, H. E., Jones, R. B., & Dolan, M. E. (2012). Relating human genetic
variation to variation in drug responses. Trend in Genetics, 28, 487–495.
39. Cheung, N. Y. C., et al. (2021). Perception of personalized medicine, pharmacogenomics, and
genetic testing among undergraduates in Hong Kong. Human Genomics, 15, 54.
40. Abrahams, E., Ginsburg, G. S., & Silver, M. (2005). The personalized medicine coalition: Goals
and strategies. American Journal of Pharmacogenomics : Genomics-Related Research in Drug
Development and Clinical Practice, 5, 345–355.
41. Prainsack, B., Naue, U. Relocating health governance: personalized medicine in times of
‘global genes’. Personalized Medicine, 3, 349–355.
42. Gershon, E. S., Alliey-Rodriguez, N., & Grennan, K. (2014). Ethical and public policy
challenges for pharmacogenomics. Dialogues in Clinical Neuroscience, 16, 567–574.
43. Mattevi, V. S., & Tagliari, C. F. (2017). Pharmacogenetic considerations in the treatment of
HIV. Pharmacogenomics, 18,85–98.
44. Rodríguez Duque, R., & Miguel Soca, P. E. (2020). Farmacogenómica: principios y
aplicaciones en la práctica médica. Revista Habanera de Ciencias Médicas, e3128.
45. Abbagnano, N. (2007). Dicionário de Filosofia (5th ed). Ed. Martins Fontes.
46. Sgreccia, E. (1996). Manual de bioética: fundamentos e ética biomédica. In Manual de bioética:
fundamentos e ética biomédica (pp. 686–686).
47. Beauchamp, T. L.; Childress, J. F. (2013) Principles of Biomedical Ethics. New York: Oxford
University Press.
48. Ferrer, J. J., & Álvarez, J. C. (2005). Para fundamentar a bioética: teorias e paradigmas
teóricos na bioética contemporânea. Edicoes Loyola.
49. Suarez-Kurtz, G. (2018). Pharmacogenetic testing in oncology: A Brazilian perspective. Clinics
(Sao Paulo), 73, e565s.
50. Ruiz-Hornillos, J., et al. (2021). Bioethical concerns during the COVID-19 pandemic: What did
healthcare ethics committees and institutions state in Spain? Frontiers in Public Health, 9,
737755.
51. Brito, M. (2015). A farmacogenética e a medicina personalizada. Saúde e Tecnologia, 14, 5–10.
52. Ortolan, A., et al. (2021). The genetic contribution to drug response in Spondyloarthritis: A
systematic literature review. Frontiers in Genetics, 12, 703911.
53. Iriart, J. A. B. (2019). Medicina de precisão/medicina personalizada: análise crítica dos
movimentos de transformação da biomedicina no início do século XXI. Cadernos De Saude
Publica, 35, e00153118.

Chapter 4
Machine Learning and Neural Network
Methods Applied to Drug Discovery
Daniel S. de Sousa, Aldineia P. da Silva, Rafaela M. de Angelo,
Laise P. A. Chiari, Kathia M. Honorio, and Albérico B. F. da Silva
Abstract Throughout this chapter, we will explore how machine learning and
neural networks are shaping the evolution of drug design, its fundamental applications, limitations, and the challenges that still need to be overcome. The revolutionary potential of this approach promises to conti nue contributing to the discovery of
new therapies and advancing pharmaceutical science.
Keywords Machine learning · Drug design · Applications · Neural networks
1 Historical Background
The use of machine learning (ML ) in the field of drug discovery represents a
revolutionary approach that has profoundly transformed how scientists and
researchers approach the development of new therapeutic compounds. Throughout
history, the field of drug design has faced numerous challenges in identifying and
optimizing molecules capable of treating diseases effectively, safely, and efficiently.
However, the advent of ML has brought new perspectives and remarkable advances.
Historically, the drug design process has predominantly relied on trial-and-error
approaches, with scientists conducting exhaustive experiments to identify compounds with desirable properties [1]. Nevertheless, the increasing availability of
data and the growing computational processing power have opened new possibilities
for the application of ML algorithms [2]. Notable examples of this evolution include
the use of Quantitative Structure-Activity Relationship (QSAR), which has enabled
D. S. de Sousa · L. P. A. Chiari · A. B. F. da Silva (✉)
São Carlos Institute of Chemistry, University of São Paulo, São Carlos, SP, Brazil
e-mail: alberico@iqsc.usp.br
A. P. da Silva · K. M. Honorio
School of Arts, Sciences and Humanities, University of São Paulo, São Paulo, SP, Brazil
R. M. de Angelo
Center for Natural Sciences and Humanities, Federal University of ABC, Santo André, SP,
Brazil
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2024
V. G. Maltarollo (ed.), Computer-Aided and Machine Learning-Driven Drug
Design, Computer-Aided Drug Discovery and Design 3,
https://doi.org/10.1007/978-3-031-76718-0_4
65

66 D. S. de Sousa et al.
the prediction of biological activities based on the molecular structure, and the
application of deep learning algorithms, such as neural networks, which have
enhanced the ability to identify complex patterns in extensive data sets [3]. For
more details on QSAR, see Chap. 6.
This transformation has significantly impacted applications in drug design. ML is
now widely employed to expedite the identification of drug candidates, optimize the
structure of existing molecules, and predict potential side effects, resulting in a more
efficient and economically advantageous process. Furthermore, the capability to
construct models of interactions between compounds and biological targets has
improved the understanding of drug mechanisms, unveiling new perspectives for
therapeutic development [1–3].
However, despite its notable achievements, the application of ML in drug design
faces critical challenges and limitations. The need for high-quality and reliable data,
which is not always readily available, is one of these barriers. Additionally, ML
models’ interpretability and domain expertise’s incorporation remain ongoing challenges. Ensuring the safety and efficacy of drug candidates identified by ML
algorithms is also essential [4 , 5]. For more details on the application of ML models,
see Chap. 6.
The history of ML and neural network development has been marked by numerous peaks and vall eys, woven with both triumphs and setbacks over its expansive
chronology. Currently, our primary focus does not involve an in-depth exploration
of the historical context. Instead, our goal is centered on encapsulating this narrative
concerning medicinal chemistry, highlighting key events involving ML that have
been impactful in drug discovery and its evolution.
1.1 Timeline
Before we embark on the application of ML into our field of expertise, it is essential
to trace its historical evolution. The initial concepts of ML emerged in 1943, albeit
not in the concrete ML form we know today. Instead, these nascent ideas originated
from an endeavor to understand the workings of neurons. It was the neurophysiologist Warren McCull och and the mathematician Walter Pitts who ventured into
creating a simplified neural network using electrical circuits [6]. Intriguingly, this
development predated the advent of the first electronic digital computer, ENIAC, in
1946 [7]. This highlights that the quest for creating artificial intelligence (AI) has
deep historical roots, evolving alongside the nascent field of computer science.
The initial connection between arti ficial intelligence (AI) and computers was
established in 1950 when the mathematician and computer scientist Alan Turing
introduced a test to evaluate a machine’s capacity to exhibit human-like intelligence,
known as “ Turing test” [8]. However, the term “machine learning“only entered the
lexicon in 1952, coined by computer scientist Arthur Samuel, a pioneer in the field
[9]. Samuel developed a program capable of playing checkers at a championship
level, utilizing an algorithm called “alpha-beta pruning.” Nonetheless, it is worth

4 Machine Learning and Neural Network Methods Applied to Drug Discovery 67
noting that while situated within the realm of artificial intelligence, this program did
not involve the process of learning from data [10].
The first authentic ML algorithm emerged in 1957, courtesy of the psychologist
Frank Rosenblatt. This algorithm marked the prototype of Artificial Neural Network
(ANN) and was named the “perceptron,” a single-layer neural network [11]. Subsequently, a plethora of other methods were conceived, including the Nearest Neighbor Algorithm (1967), paving the way for the diverse methods that are encountered
today [ 12].
With the advent of ML techniques, numerous fields of knowledge have been
embraced and enriched. For instance, the vast domain of chemistry witnessed
significant advancements with the introduction of the Dendral system, developed
by Edward Feigenbaum and Joshua Lederberg in 1965. This innovative system’s
primary objective was to elucidate the chemical structures of compounds by intricately interpreting spectrographic data [13]. In the field of drug discovery, ML made
its initial foray in the 1990s, and its applications have proliferated in the 2000s up to
the present day [3–5, 14–18].
A significant milestone in the history of ML in drug design was the introduction
of the “drug-likeness” concept in 1998 by Ajay et al. Their model was designed to
predict with high accuracy whether a molecule could be categorized as a drug or not,
employing Bayesian neural network algorithms [19]. In the ensuing years, during the
2000s, the creation of new methods, such as Random Forest, and the popularization
of other ML techniques like SVM, decision trees, and Naive Bayes, among others,
along with the increasing availability of data, led to the development of various
strategies to enhance the drug disco very process [14–16, 20].
In the same decade, the first substantial databases were established, including
ZINC and PubChem (2004) [21, 22]. This was followed by the creation of
DrugBank in 2006 [23] and ChEMBL in 2008 [20, 24]. However, signi ficant
advances in the field of drug discovery only materialized in more recent times. In
2015, Atomwise introduced AtomNet, the pioneering deep learning neural network
for structure-based drug design, utilizing three-dimensional representations of chemical interactions. The system identified chemical features like aromaticity, sp
carbons, and hydrogen bonding, akin to how image recognition networks comprehend spatially proximate features. AtomNet subsequently played notable importance
in predicting novel candidate biomolecules for various disease targets, notably
contributing to treatments for the Ebola virus and multiple sclerosis [25]. The
culmination of progress in the field occurred in 2020, marked by two notable events:
the discovery of halicin, an antibiotic [26] through deep learning models, and the
introduction of the first planned machine learning drug candidate for the treatment of
cancer and cardiovascular disease into preclinical testing [27]. Figure 4.1 illustrates
the chronological progression of ML in the field of drug discovery, depicting key
milestones.
3

68 D. S. de Sousa et al.
1965
1957
Fig. 4.1 Timeline with some events within the field of machine learning since its inception
(yellow), spanning from chemistry (green) to drug discovery (purple)
1998
2000s
2006
2004
2008
2015
2020
2 Methodology Overview
ML is a fundamental branch of knowledge within the field of AI that stands out for
its ability to enable computational systems to learn and improve from data, rather
than being explicitly programmed. While AI encompasses a wide spectrum of
techniques and approaches aimed at endowing machines with human-like intelligence and behavior, ML focuses on a system’s capacity to acquire knowledge and
enhance its performance through data analysis, pattern recognition, and continuous
adaptation. This sets ML apart from traditional programming approaches, allowing
autonomous systems and algorithms to make decisions and perform tasks independently based on past experiences [28–32].
Within the realm of ML, Neural Networks (NNs) assume a fundamental role.
These computational models utilize interconnected layers of artificial neurons to
analyze and recognize intricate patterns within data. NNs can be both shallow, with
few layers, or deep, incorporating multiple layers, giving rise to what is known as
Deep Learning (DL). As a specialized branch of ML , DL, places particular emphasis
on employing deep neural networks to address complex tasks, such as image and
speech recognition, natural language processing, and more [30–38]. This hierarchy
of concepts illustrates the progressive depth and specialization that occurs withi n the
broader domain of AI, as illustrated in Fig. 4.2.
In medicinal chemistry, numerous ML algorithms are essential for addressing
diverse needs related to drug discovery. They are integral in tasks ranging from

4 Machine Learning and Neural Network Methods Applied to Drug Discovery 69
Fig. 4.2 Schematic representation of the hierarchy of concepts in the scope of artificial intelligence
to deep learning
target identification to the optimization of clinical trials. Broadly, ML methods can
be categorized into two major groups: supervised learning and unsupervised
learning [14].
In supervised learning methods, algorithms are trained on labeled data sets where
the relationship between inputs and outputs is known. This enables the prediction or
classification of new data based on prior learning. On the other hand, unsupervised
learning methods explore underlying structures in the data, either by grouping them
into clusters or by reducing dimensionality. These methods are important for
uncovering insights and patterns in unlabeled data, facilitating data segmentation
into similar groups, and simplifying the representation of complex data [32, 35, 36,
39].
Furthermore, these methods can be combined to give rise to other approaches,
such as semi-supervised learning and reinforcement learning. These approaches
offer versatile solutions for a wide range of data sets, leveraging the strengths of
both supervised and unsupervised learning paradigms. Semi-supervised learning, for
instance, utilizes a combination of labeled and unlabeled data, capitalizing on the
benefits of limited labeled information and the vastness of unlabeled data. On the
other hand, reinforcement learning introduces a dynamic element by allowing
models to learn through interaction with an environment, receiving feedback in the
form of rewards or penalties to refine their decision-making processes. This amalgamation of methods contributes to the versatility and efficacy of ML applications in
various domains, including drug discovery [14, 30, 32, 39].
The purpose of this chapter is not to delve into the detailed principles of operation
for all methods, but rather to provide an overview of the broader context within the
field of drug discovery, focusing on the commonly used ML methods for this
purpose. Figure 4.3 presents the main ML algorithms employed in various drug
design applications.
Соседние файлы в папке Библиотека им академика М.И. Перельмана
