Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5397_Библиотеки_им_академика_М_И_Перельмана
.pdf
392
https://t.me/med1917
K. Nailwal et al.
18.2.2 Machine Learning Life Cycle (Fig.18.4)
The machine learning process starts with the formulation of a problem statement
and then follows a set of predened steps to tackle the problem like “A pharmaceu-
tical company wants to identify potential drug candidates for a specic disease,
optimizing for both efcacy and safety.” ML life-cycle generally contains the fol-
lowing seven major stages:
Gathering Data
This stage involves the identication and gathering of relevant data from different
data sources, which is then combined together to make a dataset, which would
become the base for the training and testing of the ML model, e.g., for the above
problem statement collecting molecular data, biological data, and experiment data
(e.g., assay results in clinical trial data).
Data Preparation
The dataset prepared above needs to be prepared for the further steps. This step
generally involves data exploration. The statistical characteristics, format, and quality of the data are understood.
Data Wrangling
In this step, we convert the raw data into a usable format by performing cleaning
and preprocessing by handling the missing values, normalizing features, and encoding categorical values, e.g., creating informative features such as molecular nger-
prints, protein-drug interaction scores, and biological pathway indicators.
Fig. 18.4 Machine
learning life cycle

18 AI: Catalyst forDrug Discovery andDevelopment
https://t.me/med1917
393
Analyze Data
With the available formatted data, different statistical techniques are deployed in
this stage to explore, summarize, and gain insights and then build a machine learning model. It deploys mathematical techniques such as dimensionality reduction
that selectively pick the features from the raw dataset to serve requirements and
further engineering new features that would assist in the model training. Apart from
this, it involves key activities such as using descriptive statistics, data visualization,
and exploratory data analysis, e.g., choosing appropriate machine learning models,
such as deep neural networks for predicting drug-target interactions or ensemble
methods for compound screening.
Train Model
Filtered relevant data are passed on to the ML model, where it learns patterns and
relationships between input and the target variables, e.g., train the selected model
using labeled data, where the labels represent known drug properties or outcomes
(e.g., drug binding afnity and toxicity).
Test Model
Following the training phase, the model’s performance is assessed using metrics
tailored to the particular task, such as accuracy, precision, recall, and F1 score
for classication or mean-squared error for regression. Cross-validation or
holdout validation is often used to assess how well the model generalizes to
new, unseen data, e.g., assess the model’s performance using relevant evalua-
tion metrics (e.g., ROC-AUC for binary classication tasks) on a validation
dataset.
Deployment
Once the model performs satisfactorily during evaluation, it can be deployed into a
real-world environment where it will make predictions or decisions on new, incoming data, e.g., deploy the trained model into a drug discovery platform or system
where it can predict the properties of new candidate compounds.
The deployed machine learning model, available through the above process, is
used for making predictions or inferences on new data. This task may involve classication, making recommendations, forecasting future values, identifying anomalies, and solving various other problems. As newly classied data or data with
variations emerge, the existing model requires tuning. Therefore, the model is periodically retrained with updated data to adapt to changing conditions and enhance its
predictive capabilities.

394
https://t.me/med1917
K. Nailwal et al.
18.2.3 Machine Learning Algorithms
Machine learning algorithms have been identied to play a very important role in
drug discovery and development. These algorithms are aimed at utilizing leveraging
data-driven approaches to accelerate the identication of potential drug candidates,
optimize their properties, and enhance the efcacy of the drug development procedure. Below are discussed the popular ML models with utility proven in the drug
discovery process.
Random Forests andDecision Trees
Decision trees serve as models for both classication and regression tasks. They
operate by posing a series of sequential questions, leading us down specic paths
within the tree based on our responses (Song and Ying 2015). This model essentially
follows a set of “if this, then that” conditions to ultimately arrive at a particular
outcome. The concept of “tree depth” is essential here, indicating how many questions are asked before reaching a nal classication.
Decision trees offer several advantages, including (Fig.18.5):
Interpretability: They are highly interpretable and allow for intuitive visualiza-
tions, simplifying the understanding of complex decisions.
Overtting: Decision trees can be prone to overtting, especially when they are
very deep, introducing errors due to bias and variance. Their internal workings
are transparent, facilitating result reproducibility.
Versatility: Decision trees are versatile, as they can accommodate both numerical
and categorical data, rendering them suitable for a wide range of applications.
Fig. 18.5 Random forest and decision trees

18 AI: Catalyst forDrug Discovery andDevelopment
https://t.me/med1917
395
Scalability: They perform well with large datasets, making them efcient for data
analysis.
Speed: Decision trees are known for their exceptional speed, enabling rapid
decision- making and analysis. These advantages collectively make decision
trees a valuable tool in data analysis and machine learning. However, they also
have some disadvantages:
Greedy Algorithm: Hunt’s algorithm, used during training, is greedy and may not
always produce globally optimum results.
Overtting: Decision trees can be prone to overtting, especially when they are
very deep, introducing errors due to bias and variance.
Random forests are the remedy to overcome the above-mentioned disadvantages.
A random forest comprises multiple decision trees, and their individual results are
combined to produce a nal outcome. What makes random forests powerful is their
capability to reduce overtting while keeping bias-related errors in check. Random
forests have found utility in many research studies that deal with drug discovery. In
the research conducted by Kapsiani and Howlin (2021), they investigated the utilization of the DrugAge database for the prediction of compounds with anti-aging
properties (Kapsiani and Howlin 2021). To be more specic, they utilized the random forest algorithm to forecast whether a compound would extend the lifespan of
C. elegans (Kapsiani and Howlin 2021). Ahn etal. (2022) and their team developed
a random forest model based on Kullback-Leibler divergence (KLD) feature vectors. They successfully used this model to predict drug-target interactions (DTIs)
for 17 representative targets (Ahn etal. 2022).
Support Vector Machines (SVMs)
Support vector machines (SVMs) are a type of supervised machine learning
algorithm that has found diverse applications in various elds, including drug
discovery and management (Fig.18.6). In drug discovery and management,
SVM nds utility in various tasks, including target identication, virtual screening, quantitative structure- activity relationship (QSAR) modeling, and drug
repositioning (Vilar and Costanzi 2012). SVMs are based on the idea of nding
the best hyperplane that separates the data points of different classes or predicts
continuous output values. In the case of classication, SVM aims to maximize
the margin between two classes by nding the optimal hyperplane, which is a
line or plane (in higher-dimensional space) that separates the data points belonging to different classes. Support vectors are the data points positioned closest to
the hyperplane, and they play a crucial role in dening the location and orientation of the hyperplane. A wider margin between the support vectors and the
hyperplane indicates that the SVM model generalizes better to new, unseen data
(Rokach 2009).

396
https://t.me/med1917
Fig. 18.6 Support vector machine
K. Nailwal et al.
Ensemble Learning
This machine learning approach combines multiple individual models, known as
base models, to construct a more robust and precise predictive model (Fig.18.7). It
combines the predictions of multiple models that would surpass the performance of
any single model. It is based on the concept of “wisdom of crowd” where multiple
perspectives contribute to better decision-making. There are multiple ensemble
techniques such as voting in which each model’s prediction is counted and the class
with maximum predictions in favor is selected as the output of the model. In the
bagging technique, different base models are trained on different data subsets and
then averaged on their predictions; random forest is an example of bagging technique. The boosting technique tends to focus more on difcult examples, thus
improving accuracy by giving more weight to examples that are misclassied by the
previous models. Another method known as stacking combines the predictions generated by multiple base models by training an additional model, called a metamodel, which leverages the predictions made by the base models (Verikas et al.
1999; Zhou etal. 2002; Freund and Schapire 1997).
Reinforcement Learning (RL)
Reinforcement learning (RL) is a subeld of machine learning that focuses on
learning through interaction with an environment. It has demonstrated successful
applications in diverse elds, including game-playing, robotics, and recommendation systems (Han etal. 2023). In recent years, there has been growing interest
in applying RL in drug discovery and development. RL has the potential to accelerate drug discovery by optimizing the selection and design of molecules for testing. Through interaction with a virtual environment that simulates biological
systems, RL algorithms can learn to predict which molecules are most likely to be

18 AI: Catalyst forDrug Discovery andDevelopment
https://t.me/med1917
397
Fig. 18.7 Ensemble learning: classier
effective in treating a particular disease. This can help reduce the time and cost
associated with traditional trial-and-error methods of drug discovery. RL can also
be used to optimize the dosing of drugs in clinical trials. By learning from patient
data, RL algorithms can determine the optimal dosage for a particular patient,
taking into account factors such as age, weight, and medical history (Liu
etal. 2020).
One of the challenges of applying RL in drug discovery and development is the
availability of high-quality data. The effectiveness of RL algorithms is contingent
on the quality and quantity of data accessible for training. In drug discovery, data
may be limited due to the complexity of biological systems and the high cost of
experiments. Despite these challenges, RL has the potential to revolutionize drug
discovery and development by reducing the time and cost associated with traditional methods and by improving the effectiveness of treatments for patients. There
has not been any signicant work in the area of utilizing RL for this purpose.

398
https://t.me/med1917
K. Nailwal et al.
Deep Learning (DL)
Deep learning is a subset of machine learning, and it focuses on the usage of articial neural networks (ANNs). These neural networks aim to mimic the functioning
of the human brain. It kind of mimics the way the human brain processes information and learns from experience. These algorithms can automatically learn to recognize patterns in data, make decisions, and perform tasks without explicit human
programming.
It uses multiple layers of neurons within the ANN.These layers are often
referred to as “deep” networks. They can process and transform the input data
in a hierarchical manner, allowing the network to learn features and representations at varying levels of abstraction, thus enabling the DL models to accurately
recognize and classify complex patterns, such as those found in multimedia.
With the signicant improvement in computing prowess and availability of
large datasets, DL algorithms have gained extreme popularity over the last
decade namely convolutional neural networks (CNNs), recurrent neural networks (RNNs), and transformers. Deep learning plays a crucial role in speeding
up the process of discovering new drugs and contributes to global efforts to
combat infectious diseases by enabling the prediction of drug-target interactions, identication of lead compounds, and optimization of drug candidates.
DL models have been used in the prediction of drug- target interaction (DTI) and
new medication development (Kim etal. 2021). It not only improves the efciency of screening for antimicrobial compounds that can target a wide range of
pathogens but also shows promise in identifying potential drugs against specic
viruses such as severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2).
In fact, deep learning has been successfully used to identify several potential
drugs that could be effective against SARS-CoV-2. Some of these drugs include
atazanavir, remdesivir, kaletra, enalaprilat, venetoclax, posaconazole, and
daclatasvir (Zhang etal. 2021) (Fig.18.8).
Fig. 18.8 Simple deep neural network architecture for drug discovery

18 AI: Catalyst forDrug Discovery andDevelopment
https://t.me/med1917
399
18.3 Transformation ofDrug Discovery withAI
Conventionally, the drug comprises several sequential stages: target identication,
target validation, lead identication, lead optimization, preclinical testing, and clinical trials. The high cost, time, and limited success rate associated with conventional
drug discovery methods for these methods have led researchers to seek innovative
approaches to improve the process. The transition from traditional paper-based systems to electronic laboratory notebooks and regulatory submissions has streamlined
data management, making it easier to store, access, and share information among
researchers, clinicians, and regulatory authorities. With this, the massive ocean of
datasets available to us for the different stages of drug development has eased the
entry of articial intelligence. Researchers can access vast repositories of scientic
literature, clinical trial data, and databases of repurposable drugs. These resources
help identify potential drug targets, explore new therapeutic approaches, and accelerate the discovery and development of novel treatments. Technological advancements have also revolutionized omics proling in drug development. While
genotyping and whole-genome sequencing have been instrumental in understanding genetic variations, new technologies such as microuidics and antibody tagging
have made single-cell technologies widely accessible. This enables researchers to
study the transcriptome (e.g., using RNA-seq) (Heath etal. 2015) and the proteome
(e.g., via mass cytometry) (Spitzer and Nolan 2016) at a single-cell resolution.
Integrating multiple omics modalities (McGinnis etal. 2019) provides a more comprehensive view of the molecular mechanisms underlying diseases and drug
responses. In addition to these advancements, digital technologies have bolstered
the success rate of clinical trials. Electronic patient data capture systems allow realtime data collection, reducing errors and enhancing data quality. Moreover, articial
intelligence and machine learning algorithms can be deployed for pattern recognition in large datasets, predict treatment outcomes, and optimize trial designs. The
transformative power of digital technologies in the drug development process is
undeniable. From data collection and analysis to precision medicine and clinical
trial optimization, these technologies have revolutionized the way drugs are discovered, developed, and delivered to patients, ultimately improving patient outcomes
and advancing health care as a whole.
18.4 AI inTarget Identication andBiomarker Discovery
Identication of therapeutic targets and biomarkers represents a fundamental step
toward the search for novel therapies for many diseases starting from cancer and
neurodegenerative disorders up to infectious diseases (Bodaghi et al. 2023).
Biomarkers are endpoints in drug development that provide insights into disease

400
https://t.me/med1917
K. Nailwal et al.
states, progression, and treatment response. Targets are specic molecular pathways within the human body that can be modulated to achieve a therapeutic endpoint. Historically, this has been an empirical, experimental, and domain-reliant
process that has consumed much time, often with undue uncertainty bedeviling its
outcome. The introduction of articial intelligence (AI) has catalyzed a revolution
in areas of target identication and biomarker discovery. This is where AI as an
important partner comes into action by using advanced machine learning algorithms, and its large- scale analysis potentialities to push drug development and
precision medicine forward in a fast way (Hathout and Metwally 2023). Such a
transformation lies in the ability of AI to process colossal volumes of biological
data, pinpoint complex trends, and suggest potential targets or biomarkers at the
earliest stages. We will go further and investigate how AI has transformed the
process of target identication and biomarker discovery and explore different
techniques used, their applicability, and benets. This is an epic trip that may
yield breakthrough revelations leading to rapid advancement in therapeutic innovation and taking medicine closer to individualized remedies for a multitude of
diseases. Multiple state-of-the-art works have been done by researchers in biomarker discovery. Zhang etal. embarked on the task of developing a machine
learning (ML) algorithm to accurately quantify tumor- stroma ratio (TSR) in
hematoxylin-and-eosin (H&E)-stained whole-slide images (WSI). Subsequently,
they delved into an investigation of its prognostic implications for patients diagnosed with muscle-invasive bladder cancer (MIBC). Their approach involved the
utilization of an optimized cell classier that had been previously constructed,
leveraging the QuPath open-source software and a machine learning algorithm to
achieve a quantitative assessment of TSR (Zheng et al. 2023). Daamen et al.
(2023) have developed a novel iterative machine learning pipeline that employs
gene enrichment proles derived from blood transcriptome data. This pipeline
straties COVID-19 patients according to disease severity and distinguishes
severe COVID-19 cases from other patients experiencing acute hypoxic respiratory failure (Daamen etal. 2023). Nasimina etal. have successfully engineered a
drug sensitivity prediction model tailored specically for the receptor tyrosine
kinase inhibitor sorafenib. Remarkably, their model demonstrated an impressive
prediction accuracy rate exceeding 80% when it came to identifying AXL dependency in patients aficted with leukemia. This achievement not only underscores
the signicance of their groundbreaking research but also holds promise for the
personalized treatment of leukemia by enabling more precise and effective therapeutic choices for individuals with this devastating disease (Nasimian etal. 2023).
Karaglani and colleagues have implemented an in silico pipeline that has been
instrumental in scrutinizing high-throughput methylome datasets. Their primary
objective was to pinpoint distinctive methylation patterns within three prominent
pathological conditions that pose signicant health burdens, namely breast cancer
(BrCa), osteoarthritis (OA), and diabetes mellitus (DM). This innovative approach
promises to enhance our understanding of the epigenetic underpinnings of these
conditions, offering potential insights for improved diagnosis and treatment strategies (Karaglani etal. 2022).

18 AI: Catalyst forDrug Discovery andDevelopment
https://t.me/med1917
401
18.5 Chemoinformatics andQSAR Modeling
Chemoinformatics is the use of computational methods and tools to analyze and interpret chemical data. It involves the storage, retrieval, and analysis of chemical structures and properties, as well as the prediction of chemical behavior and properties.
Quantitative structure-activity relationship (QSAR) modeling is a specic application
of chemoinformatics that involves the development of mathematical models to predict
the biological activity or properties of chemical compounds. QSAR models are based
on the relationship between the structural features of a compound (such as its molecular descriptors) and its biological activity or property. Articial intelligence (AI) and
machine learning (ML) play a crucial role in chemoinformatics and QSAR modeling
(Hathout and Metwally 2023). In the eld of chemoinformatics, articial intelligence
(AI) and machine learning (ML) techniques are employed to examine and model
extensive datasets containing chemical structures and properties. ML algorithms can
learn from these data to recognize patterns and correlations between molecular
descriptors and biological activities or attributes. These models can subsequently be
utilized to predict the activity or property values of novel chemical compounds. AI
and ML approaches have signicantly enhanced QSAR modeling, facilitating the creation of more precise and dependable models. ML algorithms, such as support vector
machines, random forests, and neural networks, are commonly used in QSAR modeling to build predictive models from large datasets. These models can capture complex
relationships between molecular descriptors and activity or property values, allowing
for more accurate predictions. Additionally, AI and ML techniques can also be used
for feature selection, where they automatically identify the most relevant molecular
descriptors to include in the model. This helps in reducing dimensionality and improving model performance. Moreover, AI and ML can aid in virtual screening, which
involves the rapid screening of large chemical databases to identify compounds with
desired properties.
18.6 High-Throughput Screening (HTS) andAutomation
High-throughput screening (HTS) is an essential method in drug discovery and
chemical biology that enables the swift assessment of numerous chemical compounds for their biological activities. In HTS, automated systems and robotics are
employed to conduct experiments on a large scale, allowing for the screening of
thousands or even millions of compounds within a short timeframe. The high-speed
nature of this process generates an immense volume of data, which must be analyzed and interpreted.
The integration of articial intelligence (AI) in HTS and automation has signicantly transformed the drug discovery process. AI algorithms can efciently
process the copious amounts of data produced by HTS, identifying patterns,
relationships, and potential candidates for further exploration. Machine learning
Соседние файлы в папке Библиотека им академика М.И. Перельмана
