Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5364_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Foreword
- •Acknowledgments
- •Contents
- •1.1 Structure-Based Drug Discovery (SBDD)
- •1.2 Ligand-Based Drug Design (LBDD)
- •1.3 Echoes from the Past, Visions from the Future
- •References
- •1 Introduction
- •2.2 Second Step: Data Curation
- •2.4 Fourth Step: Updating and Maintenance
- •2 Databases and Curation
- •8 Perspectives
- •9 Conclusion
- •References
- •1 Introduction
- •2.1 Making and Matching Protein Models
- •2.2 Simulating Protein Movements
- •2.3 Analyzing Changes in Protein Shape
- •3 Pharmacogenomics in Drug Development
- •4 Case Studies of Genomics-Based Drug Design
- •References
- •1 Historical Background
- •1.1 Timeline
- •2 Methodology Overview
- •2.1 Neural Networks
- •2.1.1 Perceptron
- •2.1.2 Multilayer Neural Networks
- •2.1.3 Types of Neural Networks
- •Feedforward
- •Recurrent Neural Networks
- •LSTM
- •2.2 Deep Learning
- •3 Using Machine Learning
- •3.2 Data Collection
- •3.3 Data Preprocessing
- •3.4 Model Selection
- •3.5 Model Training
- •3.6 Validation
- •3.7 Tuning
- •3.8 Prediction
- •4 Limitations
- •4.1 Bias
- •4.3 Interpretability
- •4.4 Computational Cost
- •4.5 Data Dependency
- •4.6 Robustness
- •5 Applications in Drug Discovery
- •5.2 Lead Discovery
- •5.3 Preclinical and Clinical Development
- •6 Resources and Tools
- •7 Challenges and Perspectives
- •7.1 Future Trends
- •9 Conclusions
- •References
- •1 Historical Background
- •1.1 Applications in Drug Discovery
- •2 Validations and Controls
- •2.1 Internal Validation
- •2.2 External Validation
- •2.3 Relative Cluster Validation
- •3 Challenges and Perspectives
- •4 Conclusions
- •References
- •1 Historical Background
- •2 OECD Principles
- •2.1 A Defined Endpoint
- •2.2 An Unambiguous Algorithm
- •2.5 A Mechanistic Interpretation, if Possible
- •3 Software and Tools
- •4 Validations and Controls
- •4.1 Internal and External Validation
- •4.1.1 Regression Metrics
- •4.2 Applicability Domain
- •4.3 Randomization Tests
- •5 Interpretation
- •6 Practical Advice During QSAR Modeling
- •7 Application
- •8 Challenges and Perspectives
- •References
- •1 Molecular Docking
- •2 Advances in Scoring Functions and Search Algorithms
- •2.2 Critical Characteristics of Search Algorithms
- •2.3 Docking Programs and Scoring Functions
- •3 Calculations Performed During Docking Simulations
- •4 Essential Components for a Good Docking Program
- •5 Limitations of the Docking Technique
- •6 Validation of Docking Results
- •7 Inappropriate Use of Validation Methods in Docking
- •9 Use of Machine Learning in Molecular Docking
- •11 Challenges
- •12 Conclusions
- •References
- •3 System Preparation for MD Simulations
- •3.1 Solvation and Microensemble
- •3.2 Force Fields: General Concept and Relevant Choices
- •3.3 The Concept of Replicas and Timescale
- •4.1.2 Protein Root Mean Square Fluctuation (RMSF)
- •4.1.4 Protein Secondary Structure Analysis
- •4.1.5 Principal component Analysis (PCA)
- •4.1.6 Markov State Modelling
- •4.1.7 Distance Calculations
- •4.1.8 Angle and Plane Calculations
- •4.2.2 Distances and Ligand-Induced Geometry Rearrangements
- •4 Molecular Dynamics Analysis
- •4.1 Protein Perspective
- •4.1.1 Protein Root Mean Square Deviation (RMSD)
- •4.3 Ligand Perspective
- •4.3.1 Ligand Properties
- •4.3.2 Ligand Root Mean Square Deviation
- •4.3.3 Ligand Root Mean Square Fluctuation
- •4.3.4 Angles and Dihedrals
- •5.1 Protein Structure Prediction and Preparation
- •5.2 Molecular Docking
- •6 Concluding Remarks and Outlook
- •Glossary
- •References
- •1 Introduction
- •2.1 MDeNM
- •2.2 Collective Molecular Dynamics (coMD)
- •2.3 ClustENM and ClustENMD
- •3 Ensemble Docking
- •References
- •1 Introduction
- •1.1 Advantages, Disadvantages, Innovations, and Challenges
- •1.2 Recent Advances in Accessible FEP Software Tools
- •1.3 Applications of FEP in Industry and Consortiums
- •2 Expanding the Potential of FEP Calculations
- •2.1 Validating Binding Poses
- •2.2 Dealing with Solvent
- •2.3 FEP and Allostery
- •2.4 FEP and Covalent Ligands
- •2.5 Applications of FEP in Scaffold Hopping
- •2.6 Positional Analogue Scanning
- •2.7 Combinations and Alternative Approaches
- •3 Machine Learning for FEP
- •3.4 Implications for ML in FEP Calculations
- •4 Final Considerations
- •5 First Steps to FEP Simulations
- •References
- •1 Background
- •2 Ultra-Large Screening Libraries and Chemical Spaces
- •3.1 Implications of Dataset Size
- •4 Ligands on the Ultra-Large Scale
- •4.1 Ultra-Large 2D Similarity Searches
- •7 Challenges and Future Perspectives
- •7.1 Hit Triage: An Old Problem on a New Dimension
- •8 Conclusions
- •Appendix
- •References
- •1 Introduction
- •2 Enzymatic Activity Evaluations
- •3 Cytotoxicity Evaluation and Cell Viability
- •4 Antiviral Assays in Experimental Validation
- •6 In Vivo Evaluation of Compounds
- •7 Conclusions
- •References
- •1 Introduction
- •3.1 Data Collection
- •3.2 Data Preprocessing
- •3.4 Model Choice
- •3.5 Model Training
- •3.6 Model Assessment
- •3.7 External Validation
- •3.8 Implementation and Availability
- •3.9 Continuous Update
- •5 Conclusions and Perspectives
- •References
- •1 Experimental Approaches to Obtain Protein Structure
- •1.1 X-Ray Crystallography
- •1.2 Nuclear Magnetic Resonance
- •1.3 Cryo-EM
- •1.4 Hybrid Methods
- •2 Modeling Approaches to Obtain Protein Structure
- •2.1 Homology Modeling
- •2.2 Ab Initio Modeling
- •2.3 New Approaches
- •3 Conformational Diversity of Proteins
- •3.1 Characterization of Protein Conformational States
- •3.2 Experimental Methods to Study Protein Dynamics and Conformations
- •3.4 Molecular Dynamics Simulation
- •3.5 Sampling Strategies
- •4 Remarks and Perspectives
- •References
- •1 Introduction
- •2 Structure-Based Drug Design of HIV Protease Inhibitors
- •2.1 HIV-1 Protease as a Therapeutic Target
- •2.2.1 Saquinavir
- •2.2.2 Indinavir
- •2.3.1 Lopinavir
- •2.3.2 Darunavir
- •6 Conclusions
- •References
- •4 Experimental Methods to Analyze NR Activity
- •4.2 Coregulator-Recruitment
- •5 Concluding Remarks and Outlook
- •References

90 D. S. de Sousa et al.
4 Limitations
ML is a powerful approach to solving a variety of complex problems, but it is
determined to recognize and understand its limitations. Like any tool, ML has
constraints that can affect its effectiveness in different contexts. The following are
some of the most significant limitations associated with ML.
4.1 Bias
A significant limitation of ML lies in the unavoidable presence of biases. These
biases can be inadvertently incorporated into models due to the nature of the data
used in training. When data sets refle ct historical inequalities or mirror the biases of
their creators, ML algorithms tend to reproduce and, in some cases, amplify these
biases [93, 94].
The source of biases often stems from human decisions during the data collection,
feature selection, and labeling process. Additionally, ML algor ithms can
unintentionally magnify existing biases in the data, resulting in discriminatory outcomes in areas such as drug efficacy predictions, patient stratification, or adverse
effect profiling [93–95].
The complexity of these issues is exacerba ted by the fact that ML algorithms
often operate as black boxes, making it difficult to understand how decisions are
made. This makes it challenging to identify and rectify biases, especially when
results are presented without a clear explanation [93–95].
Mitigating biases in ML models requires a holistic approach that involves
developers’ awareness of ethical challenges, the implementation of more equitable
data collection practices, and the exploration of advanced methods to identify and
correct biases existing in models [95].
4.2 Overfitting and Underfitting
Overfitting and underfitting are two opposing phenomena that can occur when
training ML models. Overfitting occurs when a model is trained too well on the
training data but fails to generalize to new, unseen data. In other words, the model
“learns” the training data so well that it ends up capturing specific patterns from that
set, including noise, instead of learning general patterns that would apply to other
data sets. This often results in lower performance when the model is confronted with
data that were not used during training [89, 93–95].
On the other hand, underfitting occurs when a model is too simple to capture the
complexity of the training data. This means that the model fails to learn the
underlying patterns and cannot adjust adequately to the training data. As a result,
the model performs poorly on both the training data and new data [96–98].

4 Machine Learning and Neural Network Methods Applied to Drug Discovery 91
Finding the right balance between overfitting and underfitting is fundamental for
developing an effective ML model. This often involves adjusting the model’s
complexity using techniques such as regularization, data augmentation, crossvalidation, and hyperparameter tuning. Cross-validation, for example, can help
evaluate the model’s performance on unseen data and fine-tune its behavior to ensure
proper generalization [93–95].
4.3 Interpretability
Interpretability is an essential characteristic in ML models, referring to the ability to
understand and explain the decisions made by the model. As complex algorithms,
such as deep neural networks, gain popularity, the ability to interpret how and why a
model makes a particular prediction becomes increasingly important. The lack of
interpretability in ML models can be a barrier, especially in critical applications such
as target identification, drug design, and pharmacokinetic profile. The inability to
explain model decisions can lead to distrust in the results, hindering the acceptance
and adoption of these technologies in sensitive environments [93– 95 , 99].
Several challenges contribute to the lack of interpretability. Complex models, like
deep neural networks, often operate as black boxes, making it difficult to understand
how a specific input translates into an output. Additionally, interpretability is often
sacrificed in pursuit of performance, especially in more advanced models.
Approaches to improve interpretability include the use of simpler models, such as
decision trees, which can be more easily understood. Additionally, post hoc interpretability methods, such as Lime and SHAP, have been developed to provide
insights into model decisions, even when they are intrinsically complex [100, 101].
Interpretability is not just a technical concern but also an ethical issue. In sectors
where decisions directly impact people’s lives, it is decisive that users can understand and trust the model predictions. Therefore, advancing research and practices in
interpretability is essential to ensure the responsible and ethical use of artificial
intelligence [100, 101].
4.4 Computational Cost
Computational cost is a decisive consi deration in many aspects of computer science
and software development. It refers to the resource-intensive nature of resources
such as the processing time of the Central Processing Unit (CPU), memory, and
energy associated with the execution of operations or algorithms on a computer
system. As problems and data sets addressed by ML algorithms increase in scale and
complexity, computational cost becomes a central concern. Adva nced models, such
as deep neural networks, may require specialized hardware and substantial computing resources for efficient training.

92 D. S. de Sousa et al.
In addition, the utilization of big data, particularly in virtual high-throughput
screening and deep learning applications, especially CNNs, necessitates efficiency to
ensure prompt and effective responses. Strategies for addressing computational costs
encompass algorithm optimization, task parallelization, judicious selection of hardware architectures, and the implementation of techniques such as pruning (eliminating insignificant connections in neural networks) to streamline model
complexity [102].
The evolution of hardware technology, such as Graphics Processing Units
(GPUs) and Tensor Processing Units (TPUs), has played a significant role in
mitigating computational costs in data-intensive tasks like ML model training.
Computational cost is not just a technical consideration but also has economic
implications, especially in cloud environments where resources are often paid for
based on usage. Therefore, computational efficiency becomes necessary for optimizing costs and ensuring the economic viability of ML-based solutions [102, 103].
4.5 Data Dependency
Data dependency is a central concept in ML, emphasizing the critical influence that
input data have on the performance and effectiveness of models. The quality and
representativeness of the data used during model training directly impact the model’s
ability to generalize to new, unseen data. If the training data are not representative of
the real-world domain of the problem or is biased in some way, the model may fail to
make accurate predictions in real-world situations [93–95].
Furthermore, data dependen cy is also related to the need for sufficient data to train
models effectively. Complex models, such as deep neural networks, often require
large data sets to learn meaningful patterns and avoid overfitting. The quality and
diversity of data are also determining for addressing potential biases. If the data
reflect existing inequalities or prejudices, the model may perpetuate or amplify these
biases, leading to unfair or discriminatory outcomes. Data dependency is not limited
to the training phase; it is also relevant during the inference phase when the model
makes predictions or decisions based on input data. If the input data are of low
quality, incorrect, or incomplete, the model’s predictions can be inaccurate or
inappropriate [93–95].
4.6 Robustness
Robustness in ML refers to the ability of a model to make accurate and consistent
predictions across a variety of conditions and situations. However, ML models face
several limitations in terms of robustness. One of the primary limitations is sensitivity to perturbations in input data. Small changes or noise in the data can lead to
significant variations in model predictions, especially in complex models like deep

4 Machine Learning and Neural Network Methods Applied to Drug Discovery 93
neural networks. This makes models more prone to errors when confronted with data
that slightly differ from the training data [93 – 95 ].
Furthermore, the robustness of ML models is often compromised by adversarial
attacks. Adversaries can intentionally manipulate input data subtly to deceive the
model and induce errors. This vulnerability to adversarial attacks is a significant
concern in crit ical applications such as security, health, and finance. Another limitation is related to the dist ribution of data. If the model is trained on a data set that
does not fully represent the real-world application domain, it may struggle to
generalize to new data, resulting in a drop in performance. Additionally, robustness
is also affected by the presence of outliers in the data. Models sensitive to outliers
may have impaired performance when faced with atypical or extreme examples [93–
95].
5 Applications in Drug Discovery
The drug discovery proces s is marked by complex stages, precisely because it deals
with biological systems that possess numerous pieces of information, many of which
are often not completely mapped. This information is now being systematically
measured and mined at unprecedented levels using a plethora of ‘omics techniques
(proteomics, genomics, epigenomics, metabolomics, and transcriptomics). Certainly, these areas have a high volume of data and variables, and therefore, the
application of ML techniques is of great importance [14, 15, 38, 99, 104].
In the drug development stages, one can highlight the identification of targets,
lead disco very, pre-clinical development, and clinical development. ML algorithms
are applied in all these stages, each with a distinct function and interest, with their
applicability defined by the specific characteristics of each algorithm. Figure 4.8
illustrates this relationship in a summarized way [14, 15, 38, 99, 104].
5.1 Target Identification
Target identification is an essential step in the drug discovery process, involving the
identification and validation of specific biological molecules or pathways that play a
key role in a disease. This intricate process requires a comprehensive understanding
of the molecular mechanisms underlying the pathological condition. Researchers
employ various techniques such as genomics, proteomics, and bioinformatics to
pinpoint potential targets, seeking proteins or genes that are aberrantly expressed or
dysregulated in the diseased state. Successful target identification lays the foundation for the development of targeted therapies, allowing for the design of drugs that
precisely intervene in the disease-associated pathways, potentially leading to more
effective and tailored treatments with minimal side effects [14, 15, 20, 99].
Undoubtedly, the process of target identification involves establishing a causal
link between the identified target and the associated disease, necessitating the

94 D. S. de Sousa et al.
Target identification Lead discovery
Identification of
alternative
targets.
Prediction of novel
therapeutic targets
through target-gene
associations
Target druggability
prediction
synthesis plans
creation
Synthesizability
Virtual screening
Molecule
druggability
prediction
prediction
De Novo drug design
Drug
metabolites
prediction
QSAR
Drug
repurposing
Bind affinities
prediction
ADMETox
properties
prediction
Natural product
identification
Identification
of hits
Preclinical and clinical
development
Drug sensitivity
prediction
Metabolic sites
prediction
Patient
stratification
Prediction of tumor
responses
Prediction of
biomarkers
Prediction of
drug
mechanisms
Image analysis
Fig. 4.8 Mainly applications of ML algorithms in the drug discovery process
demonstration that manipulating the target yields a discernible impact on the disease.
This approach demands diverse data sets with varying dimensionalities and magnitudes. In addition to high-reliability literature data, it requires comprehensive omics
data representing both disease and normal states. The application of ML algorithms
further facilitates the treatment, understanding, and elucidation of the causal relationship between the disease and the target [14, 15, 20, 99].
In pursuit of this objective, commonly employed ML algorithms include classifiers such as SVM, Nearest Neighbors, gradient boosting, and Natural Language
Processing (NLP) methods. These classifiers are primarily applied for establishing
associations among target-disease-drug, identifying druggable and nondruggable
targets for specific diseases, and predicting novel therapeutic targets through target–
gene associations. Additionally, regression methods, including ANNs, and RF,
among others, are applied for target identification. These methods are typically
utilized to assess the druggability of targets based on pharmacokinetic properties,
protein structure or sequence, and exploration of novel targets such as those associated with Huntington’s disease and the identification of alternative targets [14, 15].
5.2 Lead Discovery
Lead discovery represents one of the most extensive stages in the drug discovery
process, typically the primary focus of medicinal chemists. This stage encompasses
the discovery of hits, progressing from hit-to-lead to lead optimization. Its focus is
on the identification of promising chemical compounds with the potential to become

4 Machine Learning and Neural Network Methods Applied to Drug Discovery 95
the foundation for new therapeutic agents. This intricate stage involves screening
vast libraries of compounds to pinpoint molecules exhibiting desirable biological
activities against a specific target or pathway associated with a disease. Traditionally,
the lead discovery process involves extensive laboratory assays that are timeconsuming and expensive. The application of ML in this phase optimizes the
process, accelerating the identification of promising drug candidates [14, 15].
In the first phase corresponding to the identification of hits, virtual screening and
drug repurposing stand out. Drug repurposing aims to identify new purposes for
existing drugs. In this process, the use of Deep Neural Networks (DNNs) is common,
but methods such as naive Bayes, SVM, KNN, and RF can also be found. Additionally, these methods can be applied to predict binding affinities in the virtual
screening process based on score functions from docking, without the need to apply
physical concepts and functions (which are more complex). Algorithms based on
deep learning, such as CNNs, shine in handling 3D structures [14, 15].
In the hit-to-lead process, two approaches heavily involve ML algorithms,
namely QSAR and De Novo design. In QSAR, essentially any regression algorithm
can be used for predicting biological activity, showcasing the technique’s versatility
across various methods. In De Novo drug design, RNNs play a significant role in
creating new compounds with desirable biological activity [14, 15]. For more details
on the application of QSAR and ML models, see Chap. 6.
In the final phase corresponding to lead optimization, various methods are used
for predicting ADME properties, toxicity, and ultimately the druggability of biologically active compounds. The most employed methods fall within the realm of deep
learning such as DNNs and CNNs. Conventional ML methods like SVM can also be
found, but their performance is markedly inferior to DL methods. Another determining factor is assessing synthesizability and creating synthesis plans for the
envisioned compounds, a process often performed by reinforcement learning algorithms. Furthermore, ML algorithms find applications in predicting drug metabolites
and metabolic sites, as well as natural product identification [14, 15].
5.3 Preclinical and Clinical Development
The evolution of ML application in clinical and pre-clinical stages is constantly
progressing, and already showing significant advancements. In the pre-clinical
phase, ML algorithms stand out for their ability to identify and predict biomarkers,
predominantly utilizing regression and classification approaches. These biomarkers
are of utmost relevance, enabling a deeper understanding of drug mechanisms and
efficiently differentiating them for specific patients. Commonly used traditional ML
methods for this purpose include SVM and RF. These models are applied in the
initial stages of clinical trials, aiming to reduce both the time and costs associated
with the final stages [14, 15].
In the clinical context, the application of ML model s for biomarker discovery and
prediction of therapeutic responses continues to stand out. The emphasis on

96 D. S. de Sousa et al.
classification approaches is evident, where algorithms are employed to stratify
patients, identify potential indications, and suggest drug mechanisms [14, 15].
Another critical application is the construction of drug sensitivity models, where
supervised learning techniques, such CNNs, are employed. These models are particularly applicable in digital pathology, where image analysis allows the prediction
of tumor responses to different treatments. Drug sensitivity models and their
corresponding biomarkers are validated by independent test data sets, covering
both pre-clinical trials and early clinical stages [14, 15].
6 Resources and Tools
The discovery of new drugs represents a complex challenge, but recent advances in
ML implicated accelerating this process. Var ious software and online platforms have
been developed to integrate ML algorithms in the analysis of vast sets of biological
and chemical data, providing valuab le insights for scientists and researchers. These
tools offer a variety of applications, ranging from target identification to drug
prioritization, prediction of properties, and assistance in generating synthesis pathways. In this dynamic and challenging field, the diversity of available tools is
notable, as detailed in the accompanying Table 4.5 , which provides a comprehensive
overview of various platforms, their availability, software format, and specific
applications.
Beyond tools, databases are fundamental in constructing suitable models for drug
discovery, serving as a critical starting point for the ML process. These repositories
of information are necessary for the collection and organization of relevant data,
providing a robust foundation for the implementation of ML algorithms. Table 4.6
presents the main databases along with their respective descriptions used in the field
of drug discovery for building ML models.
7 Challenges and Perspectives
The integration of ML into drug discovery processes holds immense promise,
offering unprecedented advancements in identifying potential therapeutic compounds and accelerating the development of new drugs. However, this transformative approach is not without its challenges and ethical considerations. Among the
challenges, stand out:
Experimental Validation While ML offers powerful predictive capabilities, it is
not a substitute for traditional experimental methods. Experimental validation
remains indispensable to confirm and interpret ML predictions accurately. The
integration of ML with experimental approaches can synergize the drug discovery
process. ML can efficiently analyze vast data sets and propose potential drug

4 Machine Learning and Neural Network Methods Applied to Drug Discovery 97
Table 4.5 Some key ML-based tools used in the drug discovery process
Name Availability
Open targets [105] Free Web
admetSAR [106] Free Web
ACD/Percepta
[107]
pkCSM [108] Free Web
IBM RXN [109] Free/
DeepChem [110] Free Library Multi-application in drug design
RDKit [111] Free Library Multi-application in drug design
AutoQSAR [112] Commercial Installable QSAR modeling
PaDEL-Descriptor
[113]
SYBYL-X [114] Commercial Installable Molecular descriptors calculation/QSAR
MatLab [115] Commercial Installable QSAR modeling
Commercial Installable Properties
Commercial
Free Installable Molecular descriptors calculation/QSAR
Software
format Application
platform
platform
platform
Web
platform
Target identification and drug
prioritization
Properties
Prediction
Prediction
Properties
Prediction
Synthesis pathway generation
modeling
modeling
candidates, but human researchers must validate and interpret these findings. The
collaborative efforts of ML and human expertise can optimize the drug discovery
process, leading to the accelerated development of new drugs [127]. For more details
on experimental assays, see Chap. 12.
Regulation and Ethics The integration of ML in drug discovery raises ethical
considerations and regulatory challenges. Decision-making processes impacti ng
people’s health, such as drug development choices and clinical trial selections, are
increasingly influenced by AI algorithms. Ensuring fairness, transparency, and
unbiased outcomes becom es paramount. Biases in AI algorithms could lead to
unequal access to medical treatment, contradicting principles of equality and justice.
Additionally, the potential automation of certain tasks in the pharmaceutical industry
may raise concerns about job losses. Ethical AI implementation involves regulatory
compliance, periodic audits for bias, and the development of protocols ensuring data
privacy and security [128].
Generalization The challenge of generalization involves the ability of ML models
to apply learned knowledge to new, unseen data. ML models trained on specific data
sets may struggle to generalize their predictions to diverse patient populations or
different experimental conditions. Overfitting to a particular data set can limit the
model’s adaptability. Techniques such as transfer learning, where models leverage
knowledge gained from one domain to enhance performance in another, can help
address generalization challenges in drug discovery ML [15].

98 D. S. de Sousa et al.
Table 4.6 Main databases used in the drug discovery process
Name Description URL
ChEMBL [24] Information on bioactive molecules with drug-
DrugBank [23] Information on drugs and drug targets https://go.drugbank.com/
PubChem [22] Information on chemical structures, identifiers,
UniProt [116] Information on protein sequences and their
RCSB PDB
[117]
DrugCentral
[118]
ADReCS [119] Adverse Drug Reaction Classification System http://bioinf.xmu.edu.cn/
CTD [120] Comparative Toxicogenomics Database https://ctdbase.org/
DisGeNET
[121]
UK BioBank
[122]
STRING database [123]
ZINC [21] Free database of commercially available com-
Therapeutic Target Database
[124]
Promiscuous 2.0
[125]
ClinicalTrials.
gov [126]
like properties
properties, activities, patents, safety, toxicity
data
functions, many entries from genome sequencing projects
Archive of 3D structure data for large biological
molecules (proteins, DNA, RNA)
Comprehensive information on drugs and their
targets
Discovering gene–disease associations https://www.disgenet.org/
United Kingdom Biobank: Large-scale biomedical database
Protein–protein interaction information https://string-db.org/
pounds for virtual screening
Information on the known and explored therapeutic protein and nucleic acid targets
Database of promiscuous proteins and
compounds
Database of privately and publicly funded clinical studies conducted worldwide
https://www.ebi.ac.uk/
chembl/g/
https://pubchem.ncbi.nlm.
nih.gov/
https://www.uniprot.org/
https://www.rcsb.org/
https://drugcentral.org/
ADReCS/
https://www.ukbiobank.
ac.uk/
https://zinc.docking.org/
https://idrblab.net/ttd/
https://bioinf-applied.
charite.de/promiscuous2/
index.php
https://clinicaltrials.gov/
Data Integration In the dynamic landscape of drug discovery, the integration of
data emerges as a critical hurdle, encompassing both heterogeneous inputs from
“omics” domains and the complexities of homogeneous databases. Robust pipeline
architectures, powered by Extract Transform Load (ETL) tools, are indispensable for
orchestrating the seamless convergence of diverse data sets from public repositories.
This integration process demands not only technical prowess but also the application
of advanced technologies, including ML and big data analytics. Even within homogeneous data sets, challenges persist, spanning testing, logical considerations, crossplatform normalization, and statistical intricacies. ML and big data analytics serve as
invaluable assets in navigating these challenges, providing a cohesive framework for
integrating diverse strands of information. In the pharmaceutical realm, where

4 Machine Learning and Neural Network Methods Applied to Drug Discovery 99
research spans from molecular insights to individual patient data, the challenge of
integrating heterogeneous data becomes particularly pronounced. The need for a
sophisticated level of AI becomes evident, ensuring a comprehensive understanding
of data gathered across varying contex ts and scales. Modern data connectors are
essential for centralizing dissimilar data, offering a strategi c solution to allocate and
manage this diverse information effectively [15].
7.1 Future Trends
At the crossroads of science and technology, drug discovery is undergoing a
significant revolution, thanks to the increasing application of ML techniques. This
revolution not only accelerates the process of identifying new drugs but also provides a deeper and more personalized understanding of therapies. In this context,
several emerging trends act as promising catalysts for the future of drug discovery.
Transfer Learning This approach stands out as a dominant trend in the field,
characterized by the widespread adoption of pre-trained machine learning models.
This methodology involves utilizing models that have been initially trained on
extensive data sets, which may not necessarily pertain to the domains of biology
or medicine. However , these pre-trained models can be fine-tuned and adapted to
specific tasks within drug discovery. By capitalizing on the knowledge acquired in a
diverse range of domains, these models facilitate the accelerated analysis of intricate
biological data. This, in turn, provides a substantial advantage in pinpointing
potential therapeutic targets and identifying promising compounds [129].
Explainable AI (XAI) The reliance on ML-driven drug discoveries crucially
depends on the ability to interpret and understand model decisions. The rise of
Explainable Artificial Intelligence (XAI) is, therefore, a response to this need. As
ML models become more complex, ensuring transparency and interpretability
becomes paramount. XAI techniques provide insights into how models make decisions, allowing researchers to validate and better understand predictions, ensuring
more reliable outcomes [130].
Use of Imaging Data and Molecular Structure The integration of imaging and
molecular structure data marks another innovative frontier in drug discovery.
Advances in imaging techniques, such as high-resolution microscopy and computed
tomography, are now complemented by ML algorithms that can extract crucial
information from these images. Similarly, the analysis of complex molecular structures, aided by ML models, allows for a more refined understanding of interactions
between chemical compounds and biological targets [14].
Deep Learning Advancements Advancements in Deep Learning continue to
shape the landscape of drug discovery. More complex models, such as deep neural
networks, have the ability to learn more abstract hierarchical representations, capturing subtle nuances in data. This capability for extracting complex features is
Соседние файлы в папке Библиотека им академика М.И. Перельмана
