Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5397_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
15
Добавлен:
15.09.2026
Размер:
15 Мб
Скачать
☆
392
https://t.me/med1917
K. Nailwal et al.
18.2.2 Machine Learning Life Cycle (Fig.18.4)
The machine learning process starts with the formulation of a problem statement and then follows a set of predened steps to tackle the problem like “A pharmaceu-
tical company wants to identify potential drug candidates for a specic disease, optimizing for both efcacy and safety.” ML life-cycle generally contains the fol-
lowing seven major stages:
Gathering Data
This stage involves the identication and gathering of relevant data from different data sources, which is then combined together to make a dataset, which would become the base for the training and testing of the ML model, e.g., for the above
problem statement collecting molecular data, biological data, and experiment data (e.g., assay results in clinical trial data).
Data Preparation
The dataset prepared above needs to be prepared for the further steps. This step generally involves data exploration. The statistical characteristics, format, and qual­ity of the data are understood.
Data Wrangling
In this step, we convert the raw data into a usable format by performing cleaning and preprocessing by handling the missing values, normalizing features, and encod­ing categorical values, e.g., creating informative features such as molecular nger-
prints, protein-drug interaction scores, and biological pathway indicators.
Fig. 18.4 Machine learning life cycle
18 AI: Catalyst forDrug Discovery andDevelopment
https://t.me/med1917
393
Analyze Data
With the available formatted data, different statistical techniques are deployed in this stage to explore, summarize, and gain insights and then build a machine learn­ing model. It deploys mathematical techniques such as dimensionality reduction that selectively pick the features from the raw dataset to serve requirements and further engineering new features that would assist in the model training. Apart from this, it involves key activities such as using descriptive statistics, data visualization, and exploratory data analysis, e.g., choosing appropriate machine learning models,
such as deep neural networks for predicting drug-target interactions or ensemble methods for compound screening.
Train Model
Filtered relevant data are passed on to the ML model, where it learns patterns and relationships between input and the target variables, e.g., train the selected model using labeled data, where the labels represent known drug properties or outcomes (e.g., drug binding afnity and toxicity).
Test Model
Following the training phase, the model’s performance is assessed using metrics tailored to the particular task, such as accuracy, precision, recall, and F1 score for classication or mean-squared error for regression. Cross-validation or holdout validation is often used to assess how well the model generalizes to new, unseen data, e.g., assess the model’s performance using relevant evalua-
tion metrics (e.g., ROC-AUC for binary classication tasks) on a validation dataset.
Deployment
Once the model performs satisfactorily during evaluation, it can be deployed into a real-world environment where it will make predictions or decisions on new, incom­ing data, e.g., deploy the trained model into a drug discovery platform or system
where it can predict the properties of new candidate compounds.
The deployed machine learning model, available through the above process, is used for making predictions or inferences on new data. This task may involve clas­sication, making recommendations, forecasting future values, identifying anoma­lies, and solving various other problems. As newly classied data or data with variations emerge, the existing model requires tuning. Therefore, the model is peri­odically retrained with updated data to adapt to changing conditions and enhance its predictive capabilities.
394
https://t.me/med1917
K. Nailwal et al.
18.2.3 Machine Learning Algorithms
Machine learning algorithms have been identied to play a very important role in drug discovery and development. These algorithms are aimed at utilizing leveraging data-driven approaches to accelerate the identication of potential drug candidates, optimize their properties, and enhance the efcacy of the drug development proce­dure. Below are discussed the popular ML models with utility proven in the drug discovery process.
Random Forests andDecision Trees
Decision trees serve as models for both classication and regression tasks. They operate by posing a series of sequential questions, leading us down specic paths within the tree based on our responses (Song and Ying 2015). This model essentially follows a set of “if this, then that” conditions to ultimately arrive at a particular outcome. The concept of “tree depth” is essential here, indicating how many ques­tions are asked before reaching a nal classication.
Decision trees offer several advantages, including (Fig.18.5):
Interpretability: They are highly interpretable and allow for intuitive visualiza-
tions, simplifying the understanding of complex decisions. Overtting: Decision trees can be prone to overtting, especially when they are
very deep, introducing errors due to bias and variance. Their internal workings
are transparent, facilitating result reproducibility. Versatility: Decision trees are versatile, as they can accommodate both numerical
and categorical data, rendering them suitable for a wide range of applications.
Fig. 18.5 Random forest and decision trees
18 AI: Catalyst forDrug Discovery andDevelopment
https://t.me/med1917
395
Scalability: They perform well with large datasets, making them efcient for data
analysis. Speed: Decision trees are known for their exceptional speed, enabling rapid
decision- making and analysis. These advantages collectively make decision
trees a valuable tool in data analysis and machine learning. However, they also
have some disadvantages: Greedy Algorithm: Hunt’s algorithm, used during training, is greedy and may not
always produce globally optimum results. Overtting: Decision trees can be prone to overtting, especially when they are
very deep, introducing errors due to bias and variance.
Random forests are the remedy to overcome the above-mentioned disadvantages. A random forest comprises multiple decision trees, and their individual results are combined to produce a nal outcome. What makes random forests powerful is their capability to reduce overtting while keeping bias-related errors in check. Random forests have found utility in many research studies that deal with drug discovery. In the research conducted by Kapsiani and Howlin (2021), they investigated the utili­zation of the DrugAge database for the prediction of compounds with anti-aging properties (Kapsiani and Howlin 2021). To be more specic, they utilized the ran­dom forest algorithm to forecast whether a compound would extend the lifespan of C. elegans (Kapsiani and Howlin 2021). Ahn etal. (2022) and their team developed a random forest model based on Kullback-Leibler divergence (KLD) feature vec­tors. They successfully used this model to predict drug-target interactions (DTIs) for 17 representative targets (Ahn etal. 2022).
Support Vector Machines (SVMs)
Support vector machines (SVMs) are a type of supervised machine learning algorithm that has found diverse applications in various elds, including drug discovery and management (Fig.18.6). In drug discovery and management, SVM nds utility in various tasks, including target identication, virtual screen­ing, quantitative structure- activity relationship (QSAR) modeling, and drug repositioning (Vilar and Costanzi 2012). SVMs are based on the idea of nding the best hyperplane that separates the data points of different classes or predicts continuous output values. In the case of classication, SVM aims to maximize the margin between two classes by nding the optimal hyperplane, which is a line or plane (in higher-dimensional space) that separates the data points belong­ing to different classes. Support vectors are the data points positioned closest to the hyperplane, and they play a crucial role in dening the location and orienta­tion of the hyperplane. A wider margin between the support vectors and the hyperplane indicates that the SVM model generalizes better to new, unseen data (Rokach 2009).
396
https://t.me/med1917
Fig. 18.6 Support vector machine
K. Nailwal et al.
Ensemble Learning
This machine learning approach combines multiple individual models, known as base models, to construct a more robust and precise predictive model (Fig.18.7). It combines the predictions of multiple models that would surpass the performance of any single model. It is based on the concept of “wisdom of crowd” where multiple perspectives contribute to better decision-making. There are multiple ensemble techniques such as voting in which each model’s prediction is counted and the class with maximum predictions in favor is selected as the output of the model. In the bagging technique, different base models are trained on different data subsets and then averaged on their predictions; random forest is an example of bagging tech­nique. The boosting technique tends to focus more on difcult examples, thus improving accuracy by giving more weight to examples that are misclassied by the previous models. Another method known as stacking combines the predictions gen­erated by multiple base models by training an additional model, called a meta­model, which leverages the predictions made by the base models (Verikas et al.
1999; Zhou etal. 2002; Freund and Schapire 1997).
Reinforcement Learning (RL)
Reinforcement learning (RL) is a subeld of machine learning that focuses on learning through interaction with an environment. It has demonstrated successful applications in diverse elds, including game-playing, robotics, and recommen­dation systems (Han etal. 2023). In recent years, there has been growing interest in applying RL in drug discovery and development. RL has the potential to accel­erate drug discovery by optimizing the selection and design of molecules for test­ing. Through interaction with a virtual environment that simulates biological systems, RL algorithms can learn to predict which molecules are most likely to be
18 AI: Catalyst forDrug Discovery andDevelopment
https://t.me/med1917
397
Fig. 18.7 Ensemble learning: classier
effective in treating a particular disease. This can help reduce the time and cost associated with traditional trial-and-error methods of drug discovery. RL can also be used to optimize the dosing of drugs in clinical trials. By learning from patient data, RL algorithms can determine the optimal dosage for a particular patient, taking into account factors such as age, weight, and medical history (Liu etal. 2020).
One of the challenges of applying RL in drug discovery and development is the availability of high-quality data. The effectiveness of RL algorithms is contingent on the quality and quantity of data accessible for training. In drug discovery, data may be limited due to the complexity of biological systems and the high cost of experiments. Despite these challenges, RL has the potential to revolutionize drug discovery and development by reducing the time and cost associated with tradi­tional methods and by improving the effectiveness of treatments for patients. There has not been any signicant work in the area of utilizing RL for this purpose.
398
https://t.me/med1917
K. Nailwal et al.
Deep Learning (DL)
Deep learning is a subset of machine learning, and it focuses on the usage of arti­cial neural networks (ANNs). These neural networks aim to mimic the functioning of the human brain. It kind of mimics the way the human brain processes informa­tion and learns from experience. These algorithms can automatically learn to recog­nize patterns in data, make decisions, and perform tasks without explicit human programming.
It uses multiple layers of neurons within the ANN.These layers are often referred to as “deep” networks. They can process and transform the input data in a hierarchical manner, allowing the network to learn features and representa­tions at varying levels of abstraction, thus enabling the DL models to accurately recognize and classify complex patterns, such as those found in multimedia. With the signicant improvement in computing prowess and availability of large datasets, DL algorithms have gained extreme popularity over the last decade namely convolutional neural networks (CNNs), recurrent neural net­works (RNNs), and transformers. Deep learning plays a crucial role in speeding up the process of discovering new drugs and contributes to global efforts to combat infectious diseases by enabling the prediction of drug-target interac­tions, identication of lead compounds, and optimization of drug candidates. DL models have been used in the prediction of drug- target interaction (DTI) and new medication development (Kim etal. 2021). It not only improves the ef­ciency of screening for antimicrobial compounds that can target a wide range of pathogens but also shows promise in identifying potential drugs against specic viruses such as severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2). In fact, deep learning has been successfully used to identify several potential drugs that could be effective against SARS-CoV-2. Some of these drugs include atazanavir, remdesivir, kaletra, enalaprilat, venetoclax, posaconazole, and daclatasvir (Zhang etal. 2021) (Fig.18.8).
Fig. 18.8 Simple deep neural network architecture for drug discovery
18 AI: Catalyst forDrug Discovery andDevelopment
https://t.me/med1917
399
18.3 Transformation ofDrug Discovery withAI
Conventionally, the drug comprises several sequential stages: target identication, target validation, lead identication, lead optimization, preclinical testing, and clini­cal trials. The high cost, time, and limited success rate associated with conventional drug discovery methods for these methods have led researchers to seek innovative approaches to improve the process. The transition from traditional paper-based sys­tems to electronic laboratory notebooks and regulatory submissions has streamlined data management, making it easier to store, access, and share information among researchers, clinicians, and regulatory authorities. With this, the massive ocean of datasets available to us for the different stages of drug development has eased the entry of articial intelligence. Researchers can access vast repositories of scientic literature, clinical trial data, and databases of repurposable drugs. These resources help identify potential drug targets, explore new therapeutic approaches, and accel­erate the discovery and development of novel treatments. Technological advance­ments have also revolutionized omics proling in drug development. While genotyping and whole-genome sequencing have been instrumental in understand­ing genetic variations, new technologies such as microuidics and antibody tagging have made single-cell technologies widely accessible. This enables researchers to study the transcriptome (e.g., using RNA-seq) (Heath etal. 2015) and the proteome (e.g., via mass cytometry) (Spitzer and Nolan 2016) at a single-cell resolution. Integrating multiple omics modalities (McGinnis etal. 2019) provides a more com­prehensive view of the molecular mechanisms underlying diseases and drug responses. In addition to these advancements, digital technologies have bolstered the success rate of clinical trials. Electronic patient data capture systems allow real­time data collection, reducing errors and enhancing data quality. Moreover, articial intelligence and machine learning algorithms can be deployed for pattern recogni­tion in large datasets, predict treatment outcomes, and optimize trial designs. The transformative power of digital technologies in the drug development process is undeniable. From data collection and analysis to precision medicine and clinical trial optimization, these technologies have revolutionized the way drugs are discov­ered, developed, and delivered to patients, ultimately improving patient outcomes and advancing health care as a whole.
18.4 AI inTarget Identication andBiomarker Discovery
Identication of therapeutic targets and biomarkers represents a fundamental step toward the search for novel therapies for many diseases starting from cancer and neurodegenerative disorders up to infectious diseases (Bodaghi et al. 2023). Biomarkers are endpoints in drug development that provide insights into disease
400
https://t.me/med1917
K. Nailwal et al.
states, progression, and treatment response. Targets are specic molecular path­ways within the human body that can be modulated to achieve a therapeutic end­point. Historically, this has been an empirical, experimental, and domain-reliant process that has consumed much time, often with undue uncertainty bedeviling its outcome. The introduction of articial intelligence (AI) has catalyzed a revolution in areas of target identication and biomarker discovery. This is where AI as an important partner comes into action by using advanced machine learning algo­rithms, and its large- scale analysis potentialities to push drug development and precision medicine forward in a fast way (Hathout and Metwally 2023). Such a transformation lies in the ability of AI to process colossal volumes of biological data, pinpoint complex trends, and suggest potential targets or biomarkers at the earliest stages. We will go further and investigate how AI has transformed the process of target identication and biomarker discovery and explore different techniques used, their applicability, and benets. This is an epic trip that may yield breakthrough revelations leading to rapid advancement in therapeutic inno­vation and taking medicine closer to individualized remedies for a multitude of diseases. Multiple state-of-the-art works have been done by researchers in bio­marker discovery. Zhang etal. embarked on the task of developing a machine learning (ML) algorithm to accurately quantify tumor- stroma ratio (TSR) in hematoxylin-and-eosin (H&E)-stained whole-slide images (WSI). Subsequently, they delved into an investigation of its prognostic implications for patients diag­nosed with muscle-invasive bladder cancer (MIBC). Their approach involved the utilization of an optimized cell classier that had been previously constructed, leveraging the QuPath open-source software and a machine learning algorithm to achieve a quantitative assessment of TSR (Zheng et al. 2023). Daamen et al. (2023) have developed a novel iterative machine learning pipeline that employs gene enrichment proles derived from blood transcriptome data. This pipeline straties COVID-19 patients according to disease severity and distinguishes severe COVID-19 cases from other patients experiencing acute hypoxic respira­tory failure (Daamen etal. 2023). Nasimina etal. have successfully engineered a drug sensitivity prediction model tailored specically for the receptor tyrosine kinase inhibitor sorafenib. Remarkably, their model demonstrated an impressive prediction accuracy rate exceeding 80% when it came to identifying AXL depen­dency in patients aficted with leukemia. This achievement not only underscores the signicance of their groundbreaking research but also holds promise for the personalized treatment of leukemia by enabling more precise and effective thera­peutic choices for individuals with this devastating disease (Nasimian etal. 2023). Karaglani and colleagues have implemented an in silico pipeline that has been instrumental in scrutinizing high-throughput methylome datasets. Their primary objective was to pinpoint distinctive methylation patterns within three prominent pathological conditions that pose signicant health burdens, namely breast cancer (BrCa), osteoarthritis (OA), and diabetes mellitus (DM). This innovative approach promises to enhance our understanding of the epigenetic underpinnings of these conditions, offering potential insights for improved diagnosis and treatment strat­egies (Karaglani etal. 2022).
18 AI: Catalyst forDrug Discovery andDevelopment
https://t.me/med1917
401
18.5 Chemoinformatics andQSAR Modeling
Chemoinformatics is the use of computational methods and tools to analyze and inter­pret chemical data. It involves the storage, retrieval, and analysis of chemical struc­tures and properties, as well as the prediction of chemical behavior and properties. Quantitative structure-activity relationship (QSAR) modeling is a specic application of chemoinformatics that involves the development of mathematical models to predict the biological activity or properties of chemical compounds. QSAR models are based on the relationship between the structural features of a compound (such as its molecu­lar descriptors) and its biological activity or property. Articial intelligence (AI) and machine learning (ML) play a crucial role in chemoinformatics and QSAR modeling (Hathout and Metwally 2023). In the eld of chemoinformatics, articial intelligence (AI) and machine learning (ML) techniques are employed to examine and model extensive datasets containing chemical structures and properties. ML algorithms can learn from these data to recognize patterns and correlations between molecular descriptors and biological activities or attributes. These models can subsequently be utilized to predict the activity or property values of novel chemical compounds. AI and ML approaches have signicantly enhanced QSAR modeling, facilitating the cre­ation of more precise and dependable models. ML algorithms, such as support vector machines, random forests, and neural networks, are commonly used in QSAR model­ing to build predictive models from large datasets. These models can capture complex relationships between molecular descriptors and activity or property values, allowing for more accurate predictions. Additionally, AI and ML techniques can also be used for feature selection, where they automatically identify the most relevant molecular descriptors to include in the model. This helps in reducing dimensionality and improv­ing model performance. Moreover, AI and ML can aid in virtual screening, which involves the rapid screening of large chemical databases to identify compounds with desired properties.
18.6 High-Throughput Screening (HTS) andAutomation
High-throughput screening (HTS) is an essential method in drug discovery and chemical biology that enables the swift assessment of numerous chemical com­pounds for their biological activities. In HTS, automated systems and robotics are employed to conduct experiments on a large scale, allowing for the screening of thousands or even millions of compounds within a short timeframe. The high-speed nature of this process generates an immense volume of data, which must be ana­lyzed and interpreted.
The integration of articial intelligence (AI) in HTS and automation has sig­nicantly transformed the drug discovery process. AI algorithms can efciently process the copious amounts of data produced by HTS, identifying patterns, relationships, and potential candidates for further exploration. Machine learning