Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5443_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
10.10.2026
Размер:
9 Мб
Скачать
☆
15.3 Computational approaches for predicting drug
solubility and permeability using AI
Computational approaches for predicting drug solubility and permeability using AI
have gained significant popularity and have become a transformative force in drug dis-
covery and development. The rise of AI in this field is primarily attributed to the potent
capabilities of ML algorithms, the accessibility of large and diverse datasets, and the
rapid advancements in computational resources as shown in the Fig 15.2 via taking an
example of ML for regression analysis. AI-driven methods offer the potential to revolu-
tionize the drug developme nt process by providing faster, more cost-effective, and
more accurate predictions of crucial pharmacokinetic properties [44].
In drug development, drug solubility and permeability are essential factors that
directly influence a drug candidate’s success or failure [45]. Poor solubility can lead to
inadequate bioavailability, reducing the drug’s therapeutic efficacy. Likewise, low
permeability can hinder the drug’s ability to reach its target site and exert its pharma-
cological effects [46]. Therefore, predicting and optimizing these properties early in
Data
Data Preprocessing
Feature Scaling
Outlier Removal etc.
Data Preprocessing
Regression Model Learning
Hyperparameter tuning etc.
2
RMSE (Root Mean Squared Error
R (R-square) etc
Figure 15.2: Flowchart of machine learning regression
analysis.
15 Computational approaches for predicting drug solubility 353
https://t.me/med1917
the drug discovery process are critical steps in identifying viable drug candidates and
avoiding costly and time-consuming failures later in the development pipeline [47].
AI-based approaches leverage ML algorithms to analyze large datasets containing
information on chemical structures and the corresponding solubility and permeabil-
ity data. By learning from these data, AI models can capture complex relationships
and patterns that may not be evident through traditional statistical approaches. These
AI-driven methods can be broadly categorized into several techniques [48, 49]:
– DL models
– Graph neural networks (GNNs)
– Transfer learning (TL)
– Ensemble methods
– Generative models
– Virtual screening
DL models have emerged as a revolutionary approach in many fields, including drug
discovery, due to their capacity to tackle complex problems by automatically learning
hierarchical representations from large datasets as shown in Fig 15.3. In the context of
drug solubility and permeability prediction, DL techniques have shown great promise
in capturing the intricate relationships between chemical structures and pharmacoki-
netic properties [49].
DL models are built upon networks drawing inspiration from the intricate workings of
the human brain. These models comprise layers of interconnected nodes, also known
as neurons. They are organized into an input layer, one or more layers, and an output
layer. Each neuron receives input data, carries out computations using acquired param-
eters (weights and biases), and generates an output that is then transmitted to the layer.
The learning process involves adjusting these parameters to minimize the difference
between the predicted outputs and the actual target values in the training data [49].
Figure 15.3: Deep learning model.
354 Vimal Arora, Payal Mittal, and Sanjay Kumar Elisetti
https://t.me/med1917
One of the primary reasons DL models are successful in drug solubility and per-
meability prediction is their ability to handle high-dimensional data. Chemical struc-
tures can be represented in various ways, such as molecular fingerprints , 2D or 3D
molecular graphs, or SMILES strings and DL models classify them as insoluble, slightly
soluble, or soluble in water via a d eep neural network trained off of the dataset.
These representations can result in a large number of features or descriptors, making
traditional ML methods less effective due to the “curse of dimensionality.” DL models
excel at automatically extract ing meaningful patterns and features from such high-
dimensional data, making them well-suited for drug discovery tasks [50].
Furthermore, DL models can learn hierarchical representations, which means
they can identify complex patterns and relationships between molecular features. As
a result, these models can capture both local and global characteristics of the chemi-
cal structures, providing a more comprehensive understanding of how specific struc-
tural elements impact drug solubility and permeability [51].
When predicting drug solubility and permeability, DL models take molecular
structures as input and learn from a diverse set of compounds, with known solubility
and permeability data. By analyzing these relationships, the models can generalize to
predict the properties of new, unseen compounds accurately [52].
The success of DL models in drug discovery is not without challenges. Large and
diverse datasets with accurate solubility and permeability data are essential for train-
ing robust models. Additionally, DL models can be computationally expensive and re-
quire significant computational resources, particularly for training large-scale models
with extensive datasets [51].
Despite these challenges, the application of DL models for drug solubility and per-
meability prediction has the potential to significantly accelerate the drug discovery
process. By efficiently a nalyzing and predicting key pharmacokinetic properties,
these models can guide researchers in selecting the most promising drug candidates
for further experimental validation, thereby reducing costs and time in the drug de-
velopment pipeline. As the field of DL continues to advance, so too will its impact on
drug discovery and the development of safe and effective medications.
GNNs have emerged as a powerful class of DL models, specifically designed to work with
graph-structured data. In the context of drug discovery, molecular structures can be rep-
resented as graphs, where atoms and bonds are nodes and edges, respectively. GNNs are
well-suited for analyzing these molecular graphs and have shown great promise in pre-
dicting drug properties, based on the connectivity patterns of atoms and bonds.
Molecular structures are inherently graph-like, and the a rrangement of atoms
and bonds determines a molecule’s properties and behavior. GNNs leverage this
graph structure and the information encoded in it to learn and make predictions
about the molecule’s properties, such as solubility and permeability [52,53].
The GNN architecture is inspired by the concept of message passing in graphs. At
each layer of the GNN, every node (representing an atom) aggregates information from
15 Computational approaches for predicting drug solubility 355
https://t.me/med1917
its neighboring nodes (adjacent atoms) through the edges (bonds). This process allows
the GNN to capture local and global interactions within the molecular graph, effectively
learning from the connectivity patterns and their impact on the molecule’s properties.
GNNs utilize trainable parameters to update the features of each node during the
message passing process. The aggregation of information from neighboring nodes,
along with the node’ s current feature representation, forms a new representation
that encapsulates information from both the node itself and its neighbors. This pro-
cess is repeated through multiple layers, allowing the GNN to iteratively refine and
enrich the node representations as shown in Fig 15.4.
By processing the molecular graph in this manner, GNNs can learn complex patterns and
relationships that are crucial for understanding the structure–property relationships of
drug molecules. For example, GNNs can identify the presence of specific functional groups,
the arrangement of atoms around a ring system, or the orientation of substituents, all of
which can significantly influence a molecule’s pharmacokinetic properties [53].
Moreover, GNNs can handle variable-sized graphs, meaning they can analyze
molecules of different sizes and complexities without requiring fixed-size input repre-
sentations. This adaptability makes GNNs well-suited for analyzing diverse chemical
structures and drug candidates [53].
The ability of GNNs to learn directly from the molecular graph structure makes
them highly valuable for drug discovery tasks. They can predict drug properties, includ-
ing solubility and permeability, with impressive accuracy, especially when combined
with large and diverse datasets. These predictions can provide crucial insights into the
potential pharmacokinetic behavior of drug candidates, enabling researchers to focus
their efforts on the most promising molecules and avoid costly experimental failures.
TL is a method in ML where we utilize the knowledge obtained from one task to enhance
the performance of another related task as represented in Fig 15.5. In the context of drug
Figure 15.4: Graph neural networks.
356 Vimal Arora, Payal Mittal, and Sanjay Kumar Elisetti
https://t.me/med1917
discovery and predicting properties like drug solubility and permeability, TL has gained
prominence as a way to enhance the accuracy and efficiency of predictive models [54].
Traditionally, ML models are trained from scratch on specific datasets for a particular
task. However, in many cases, there might be a scarcity of labeled data for a particu-
lar task, making it challenging to build accurate models. This is where TL comes into
play. Instead of starting from scratch, TL allows a model to start with the knowledge
gained from solving a related task that has more available data, and then fine-tune it
for the target task with a smaller dataset.
In the context of drug discovery and predicting drug properties, TL works as
follows:
1. P retraining: A model is first trained on a large dataset and a related task. For
instance, the model could be trained to predict the binding affinity of molecules
to a specific protein target.
2. Feature extractio n: During pretraining, the model learns to extract features
from the data that are useful for the related task. In drug discovery, these fea-
tures might include molecular fingerprints, chemical substructures, or other rele-
vant descriptors.
3. Fine-tuning: After pretraining, the model is fine-tuned on a smaller dataset for the
target task, which could be predicting drug solubility or permeability. The model’s
parameters are adjusted to make predictions more aligned with the target property.
15.4 Transfer learning offers several advantages
in drug discovery
Data efficiency: Since the model starts with knowledge from a related task, it requires
less data for fine-tuning the target task. This is particularly valuable when there is a
limited amount of labeled data available for the specific property being predicted.
Generalization: TL helps the model generalize better by learning high-level features
that are relevant across tasks. This can lead to improved performance on the target
task even with limited data.
New Model
Pre-trained
Model
Knowledge
Tas k A
Tas k B
Figure 15.5: Transfer learning.
15 Computational approaches for predicting drug solubility 357
https://t.me/med1917
Faster convergence: Models that have been pretrained can converge faster during
fine-tuning since they already have a good initial understanding of the data.
Improved accuracy: By leveraging information learned from a related task, the model
can potentially achieve higher accuracy on the target task than training from scratch.
Despite its advantages, TL also comes with challenges:
Domain shift: If the source task and the target task are too dissimilar, TL might not be
effective, as the learned features may not be relevant.
Task compatibility: The related task should share some underlying features or pat-
terns with the target task for TL to work effectively.
Overfitting: Fine-tuning requires careful balancing to avoid overfitting, especially
when the target dataset is small [55].
In drug discovery, TL can be a valuable approach when predicting drug prope rties
like solubility and permeability. By capitalizing on existing knowledge from related
tasks, TL enhances the efficienc y and accuracy of predictive models, contributing to
faster and more effective drug development processes.
Ensemble methods: are powerful techniques in ML that combine the predictions of
multiple individual models to produce a final prediction. These methods capitalize on
the diversity of individual models’ strengths to create a more accurate and robust
overall prediction. In the context of drug discovery and predicting drug properties
like solubility and permeability, ensemble methods have demonstrated their effective-
ness in improving predictive performance [56].
Ensemble methods operate under the principle that combining the outputs of
multiple models can mitigate the weaknesses of individual models and yield more re-
liable results. There are several types of ensemble methods, each with its own ap-
proach to aggregating predictions: one of the generalized approach of Ensemble
methods is shown in Fig 15.6
Bagging (bootstrap aggregating): Bagging involves training multiple instances of the
same model on different subsets of the training data. These models are then averaged
or combined to produce the final prediction. Random forest, a popular ensemble tech-
nique, is an example of bagging. In drug discovery, ensemble bagging methods can
help mitigate overfitting and enhance model generalization [56].
Boosting: Boosting focuses on sequentially training models, where each model gives
more weight to the data points misclassified by the previous models. This helps the
ensemble progressively improve its performance by prioritizing difficult instances. Al-
gorithms like AdaBoost and Gradient Boosting Trees are commonly used boosting
techniques. In drug property prediction, boosting methods can enhance the model’s
ability to capture complex relationships in the data.
358 Vimal Arora, Payal Mittal, and Sanjay Kumar Elisetti
https://t.me/med1917
Stacking: Stacking involves training multiple models and then using their predictions as
inputs to a higher-level model (meta-model). The meta-model learns to combine these
predictions to make the final prediction. Stacking is particularly effective when individ-
ual models have complementary strengths. In drug discovery, stacking can lead to a
well-rounded prediction by capturing various aspects of molecular properties.
Voting: Voting methods combine predictions by taking a majority vote (for classifica-
tion tasks) or an average (for regression tasks) of individual models’ predictions. Sim-
ple yet effective, this technique is useful when dealing with diverse models that excel
in different scenarios.
Ensemble met hods offer several advantages in drug discovery and predicting drug
properties:
Improved accuracy: By aggregating predictions from multiple models, ensemble
methods can achieve higher accuracy compared to individual models, especially
when the individual models exhibit varying levels of performance.
Robustness: Ensembles are less prone to overfitting and can handle noisy data better,
as the combination of multiple models reduces the impact of outliers or errors in indi-
vidual predictions.
Model diversity: Ensembles can integrate different types of models or models with dif-
ferent parameter settings, capturing a broader range of patterns and relationships in
the data.
Figure 15.6: Ensemble method.
15 Computational approaches for predicting drug solubility 359
https://t.me/med1917
Generalization: Ensemble methods tend to generalize well to new, unseen data, mak-
ing them reliable for predicting drug properties of new compounds.
However, ensemble methods also come with challenges:
Increased complexity: Ensembles require training and maintaining multiple models,
which can lead to increased computational and memory requirements.
Tuning: Finding the right combination of individual models and their weights in the
ensemble can require careful experimentation and tuning.
In drug discovery, ensemble methods are an asset, as they can improve the accuracy
and reliability of predictions for crucial drug properties like solubility and permeabil-
ity. By harnessing the collective wisdom of diverse models, ensemble techniques con-
tribute to more informed decision-making in drug developmentandincreasethe
likelihood of identifying successful drug candidates.
Generative models: These are a class of ML techniques that focus on creating new
data samples that resemble a given dataset. These models have found diverse applica-
tions in fields such as image generation, text synthesis, and, notably, drug discovery.
In the context of drug discovery and predicting properties like drug solubility and
permeability, generative models offer a unique approach to aiding researchers in
identifying novel and promising compounds [57].
Generative models operate by learning the underlying patterns and st ructu res
within a dataset and then generating new data samples that are consistent with those
patterns. These models can generate data that looks remarkably similar to the original
data, making them powerful tools for creating new molecules with desired properties
as illustrated in Fig 15.7.
There are various types of generative models, but two main categories stand out
for drug discovery:
GANs, known as generative adversarial networks, are a type of network that en-
compasses two components; a generator and a discriminator. These components
work in tandem. Compete during the training process.
The generator creates synthetic data samples, and the discriminator tries to distin-
guish between real and generated data. The generator enhances its capacity to generate
data that becomes progressively harder for the discriminator to distinguish from data
through this process. In drug discovery, GANs have been used to generate new molecu-
lar structures with desired properties, including enhanced solubility and permeability.
Variational autoencoders (VAEs) are models that strive to transform data into a
space where each point represents a condensed representation of a data sample. VAEs
have the ability to produce data by selecting points from this hidden space and subse-
quently decoding them into the original data format.VAEshaveproventobevaluablein
the field of drug discovery as they enable the generation of structures that exhibit desir-
able characteristics, all the while ensuring a diverse range of structural variations [57].
360 Vimal Arora, Payal Mittal, and Sanjay Kumar Elisetti
https://t.me/med1917
Generative models offered several advantages in drug discovery and predicting drug
properties:
Novel compound generation: Generative models can suggest entirely new com-
pounds that have not been previously synthesized or evaluated. This can lead to the
discovery of novel drug candidates with improved solubility, permeability, or other
desired properties.
Property optimization: Generative models can be trained to generate compounds
with specific properties by conditioning the model on desired features. This can speed
up finding molecules that match desired pharmacokinetic profiles [58].
Exploration of chemical space: Generative models can explore vast chemical
spaces that are difficult for traditional methods to navigate. This can lead to the dis-
covery of unconventional molecular scaffolds and innovative drug candidates.
However, generative models also pose challenges
Validity and safety: The generated molecules might be chemically unrealistic or even
unsafe. Careful filtering and validation are necessary to ensure that the generated
compounds are synthetically feasible and biologically relevant.
Biased data: If the training data is biased toward certain types of molecules, the gen-
erative model might produce similar biased compounds.
Generative models represent a groundbreaking approach in drug discovery, provid-
ing a tool to explore and exploit the potential of chemical space. By generating novel
compounds with desired properties, these models augment the process of identifying
viable drug candidates and accelerating the drug development pipeline.
C1,C2,C3 are First Level
Classifiers
Figure 15.7: Generative model.
15 Computational approaches for predicting drug solubility 361
https://t.me/med1917
Virtual screening is a computational technique widely used in drug discovery to
identify potential drug candidates from large libraries of compounds. This approach
leverages computer simulations, predictive models, and data analysis to prioritize
molecules that ha ve a high likelihood of being effective against a specific target or
disease. Virtual screening plays a crucial role in accelerating the drug discovery pro-
cess by significantly narrowing down the pool of compounds that need to be experi-
mentally tested [59].
There are two main types of virtual screening: ligand-based and structure-based.
Ligand-based virtual screening: In this approach, virtual screening is guided by the
knowledge of exi sting ligands or molecules that bind to the target of interest. This
technique involves developing predictive models based on the chemical and struc-
tural properties of known ligands. ML algorithms and similarity searching ar e com-
monly used to identify molecules in large databases that resemble the known ligands.
Ligand-based methods are particularly useful when the 3D structure of the target is
unknown or difficult to obtain.
Structure-based virtual screening: In structure-based virtual screening, the 3D struc-
ture of the target protein is used to identify compounds that could potentially bind to
it. This involves docking simulations, where the potential ligands are virtually docked
into the binding site of the protein. The compounds are scored based on their binding
affinity, and those with the highest scores are considered potential drug candidates.
Structure-based methods are valuable when the protein structure is available and can
provide insights into the binding interact ions between the compounds and the tar-
get [59].
The virtual screening process involves several key steps:
Target selection: Identifying the protein target that is associated with a specific dis-
ease or condition is the starting point. The target’s 3D structure (if available) or ligand
information is essential for subsequent steps.
Library preparation: A database of chemical compounds is compiled, ranging from
thousands to millions of molecules. This library includes compounds that are com-
mercially available or virtually generated.
Filtering and scoring: Ligand-based methods involve creating predictive models or using
similarity searching techniques to filter the compound library. Structure-based methods
involve docking simulations to score the potential binding affinities of the compounds.
Ranking and selection: Compounds are ranked based on their predicted or calculated
properties, such as binding affinity, drug-likeness, and pharmacokinetic properties.
Experimental validation: The top-ranked compounds from the virtual screening are se-
lected for experimental testing. These compounds undergo biochemical and biological
assays to confirm their binding to the target and their potential therapeutic effects.
362 Vimal Arora, Payal Mittal, and Sanjay Kumar Elisetti
https://t.me/med1917