Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5435_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
10.10.2026
Размер:
10 Мб
Скачать
☆
Figure 9.11: Computational prediction of the solubility of a drug (salsalate) model in sc-CO
2
was done by Abdelbasset et al. in Nguyen et al. [68].
194 Anchal Sharma et al.
https://t.me/med1917
predict the solubility of decitabine in CO2SCF by employing different ML-based mathe-
matical models. In this chapter, they used three models – linear regression (LR), deci-
sion tree (DT), and GRNN – and used 32 sample points to make solubility models. ADA-
DT (Adaboost algorithm decision tree), ADA-LR (Adaboost algorithm-linear regression),
and ADA-GRNN (generative regression neural network) models showed MAE of 6.54 ×
10
5
,4.66×10
5
, and 8.35 × 10
5
, respectively, and R
2
values of the abovementioned models
were 0.986, 0.983, and 0.911, respectively. The ADA-LR was the major model. At the end
of the study, the above model’s ideal parameters were P = 400, T = 3.38 × 10
2
, and Y =
1.064 × 10
3
[68].
In Abdelbasset et al. [70], a novel predictive model was developed within the frame-
work of DNN with the inclusion of Modred molecular descriptors for the prediction of
aqueous solubility of drugs and drug-like molecules. The DNN model was also com-
pared with modified graph convolutional network (GCN)-based approach. DNN con-
sisted of many layers and each layer contained a particular number of neural units,
which are correlated via nonlinear functions, and the weight and bias functions con-
trolled the connections among these layers but sometimes this type of neural networks
are insufficient for the extraction of molecular information. However, the operations of
DNN are highly convenient for users due to its high transferability among independent
data sets, but it is comparatively less accurate than GCN. On the other hand, modified
GCN model is highly accurate as it also takes into consideration interatomic interactions
inside the molecules which ultimately leads to the complete information extraction
from the 3D molecular structure and does not fully rely on the complete set of molecu-
lar descriptors [69].
Due to the diverse advantages offered by the computational models in various
fields of drug study, especially drug solubility of large sets of compounds, Namazi et al.
[71] used various computational tools to access the solubility of nanocarrier systems
such as COMPASS force field, which calculated the energy field density that plays a very
crucial role in understanding the drug solubility, drug-carrier intermolecular interac-
tions, and melting point, and finally the solubility parameters were determined using
molecular dyn amic simulations. Furthermore, full atomistic simulation with Hilde-
brand parameter for solubility has been used to effectively explain the behavior of non-
polar and non-interacting liquids; group contribution method and MD simulation
readily support select lead excipients and also effectively assess the miscibility of drug
with the excipients, which further paves the way to formulate a stable and efficacious
drug-loaded nanocarrier system. Using MD simulations and Hildebrand parameter, se-
lection of solvent and antisolvent can be done successfully. These models decrease the
burden of experimentally predicting solubility and also prove to be cost-effective [70].
To design the sc-CO
2
-based processes for the development of micro/nanoparticles of
chloroquine and to study its solubility, Gao et al. [72] developed a computational ap-
proach which included thermodynamic models and multilayer perceptron neural net-
work (MLPNN). Thermodynamic models or cubic equation of state (EoS)-based models:
Soave-Redlich-Kowang (SRK-EoS) and Peng-Robinson (PR-EoS) are most widely used to
9 Computational prediction of drug-limited solubility 195
https://t.me/med1917
correlate the solubility of different materials in sc-CO
2
. The equilibrium solubility (y
2
)
can be obtained by the following equation:
y
2
=
P
sub
2
TðÞ
P
’
sat,s
2
TðÞ
’
2
T, P, yðÞ
exp
v
s
2
P − P
sub
2
TðÞ

RT

(9:7)
where ’
sat,s
2
TðÞ

= saturation fugacity coefficient, which can be assumed to be one
due to the very small sublimation pressure obtained for chloroquine, ’
2
T, P, yðÞ= fu-
gacity coefficient of the solute in sc-CO
2
which can be computed through SRK-EoS and
PR-EoS via the following relationship:
RT In ’
i
= −RT In Z +
ð
∞
V
∂P
∂ni

T, V , nj − ni
− RT
"#
dV (9:8)
where Z = compressibility factor, V = molar volume of sc-CO
2
, and n
i
= moles number
of species.
Modified Wilson’s models: This model includes a combinatorial contribution part
based on Flory’s theory and another part based on Gibbs excess energy (G
E
), accord-
ing to the equation:
G
E
RT
= −y
1
In y
1
+ y
2
∧
12
ðÞ− y
2
In y
1
∧
21
+ y
2
ðÞ (9:9)
where ∧
12
and ∧
21
are the dependent-adjustable parameters to the molar volume of sc-
CO
2
(v
1
), the molar volume of solute (v
2
), and the interaction energy (λ) between them.
Now, differentiating the eq. (9.3) and rearranging the obtained functions, the γ
2
can be calculated by the following equation:
Inγ
2
= −In y
2
+ y
1
∧
21
ðÞ− y
1
∧
12
y
1
+ y
2
∧
12
−
∧
21
y
2
+ y
1
∧
21

(9:10)
At infinite dilutions, the above equation can be rewritten as
Inγ
∞
2
= 1 − v
2
ρ exp −
λ
′
12
T
r

− In
1
v
2
ρ
exp −
λ
′
21
T
r

(9:11)
Here, ∧
12
and ∧
21
are written in reduced form.
UNIQUAC model: This model takes into consideration the size and nature of the mole-
cules as well as the intermolecular forces between the solute and solvent molecules.
Also, this model can be used for solutions containing small or large molecules such as
polymers and multilayer perception neural network (MPLNN) or ANN model relies on
recurrent, well-understood, and predictable patterns in the input data to produce logi-
cal and accurate results. The model for this must be built using a dataset of measure-
ments. The determination of input values is one of the ANN modeling’s components.
196 Anchal Sharma et al.
https://t.me/med1917
The ANN model is highly reliable and produces the most accurate results when com-
pared with the experimental data [71].
In an attempt to predict the solubility of salsalate in sc-CO
2
as a green solvent and
to correlate the solubility to input parameters, including temperature and pressure,
Ashwini et al. [73] used three Gaussian process regression (GPR) ML models: simple
(raw) GPR, Ada-boosted GPR, and bagged GPR. When the model outputs’ joint proba-
bility distribution is Gaussian, then the process is called Gaussian process. The major
advantage of GPR lies in the fact that various ML tasks such as uncertainty estimation,
model development and tuning of hyperparameters have been included in GPR raw
that operates under a probabilistic framework, which takes a training dataset D=
[(y
i
x
i
) n = 1, . . . I], including I input data points x
i
€R
d
, as input and y
i
(a noisy scalar) as
the output, which can be acquired as follows
y
i
= fx
i
ðÞ+ ε
i
(9:12)
where f(x
i
) is the latent function values assumed as random variables, whereas the rele-
vant input variables x
i
are used to index them; ɛ
i
indicates the Gaussian noise (variance
σ
2
noise
and mean = 0), i.e., ɛ
i
N (0, σ
2
noise
). The chief benefit of employing the Gaussian
before assumption is that the functions can be defined using a mean function m(x)and
a covariance function cov(x,x’). Let x
✶
reflect a random sample of input variables; ac-
cordingly, the output variable is estimated by the predictive probability distribution
p(y
✶
| X, y, x) with mean and variance:
y ✶
∧
= mx✶ðÞ+ k
T
✶ K + σ
2
n
I

− 1
y − mx✶ðÞðÞ (9:13)
σ
2
y✶
= k
✶
+ σ
2
n
− k
T
✶
K + σ
2
n
I

− 1
k✶ .
(9:14)
where I = identity matrix, k
✶
= some vectors defined by [k✶]
i
= cov (x
i,
x✶), and K=co-
variance matrix of [K]
i.j
= cov (x
i
, x
j
).
To acquire reliable estimations, the covariance and mean function parameters
are redeemed from the provided dataset. The hyper-parameter values can be calcu-
lated by maximizing the log likelihood function of the training group:
Log pyjXðÞ= −
1
2
y
T
K + σ
2
n
I

− 1
y −
1
2
log jK + σ
2
n
I

j −
n
2
log 2πðÞ (9:15)
where n = quantity of training subsets.
Bagging and boosting GPR: These techniques are used to improve recognition and
prediction reliability. The chief advantage of these methods lies in the fact that they
help to diminish the problem of overfitting by integrating and aggregating various
weak learners, thus creating multiple submodel components by intelligently updating
training data, which can be used for developing better learners. Boosted GPR is the
most reliable approach as it creates an ensemble model that results in the construc-
tion of some learners with more accuracy than a single model. Adaboost can be used
9 Computational prediction of drug-limited solubility 197
https://t.me/med1917
as a boosting ensemble as it updates the distribution of the training dataset and also
the new ensemble estimators developed by Adaboost can be forced to focus on more
difficult instances [72].
Computational tools and techniques of pharmacokinetics, molecular docking,
pharmacophore modelling, and molecular dynamics (MD) were employed by Ardes-
tani et al. [74] to have better insight into the therapeutic properties and dynamics of
Cannabis sativa compounds. Cannabinol, cannabichromene, linoelaidic acid, and mor-
phinan-6-one were revealed by employing physics-based binding free energy calcula-
tions (MM-GBSA) as potential drug like molecules, which can be utilized to lead drug
discovery and AD therapy by inhibiting actions of AD target proteins such as AChE,
DCC, MAO-B, and HTR2C. In addition, ADMET profiling of the compounds has been
performed by the toxicity prediction protocol QikProp [73].
According to reports, the majority of recently discovered drugs are not suffi-
ciently soluble in aqueous solutions. Consequently, in order to have a therapeutic ef-
fect, they must be taken in high dosages. As a result, when patients take high doses of
medication, there will be greater negative effects. In 2023a, Huwaimel et al. evaluate
the solubility of lenalidomide sc-CO
2
using multiple tree-based techniques. In the ini-
tial step toward developing the supercritical method for nanonization of APIs, the sol-
ubility is basically determined using some techniques such as gravimetric meth od.
But, for more reliability in the results, decision tree (DT), extra trees (ET), and gradi-
ent boosting (GB) models are used. These models (Figure 9.12) are then optimized
using the SCA algorithm. The models used in this particular research project are
called SCA-DT, SCA-ET, and SCA-GB, and their respective R
2
values are 0.932, 0.951, and
0.997. The SCA-DT model has an RMSE error rate of 0.0948, whereas the SCA-ET model
has an RMSE error rate of 0.0822 and the SCA-GB model has an RMSE error rate of
0.0203. Therefore, the SCA-GB is introduced as the best model in this research for the
prediction of Lenalidomide solubility in the solvent [74].
The solubility of medications must be raised in order to increase their bioavail-
ability, which can be accomplished by nanonizing the pharmaceuticals. Drug solubil-
ity in supercritical CO
2
was estimated usi ng ML analysis. In 2023a, Huwaimel et al.
used a small data set with solubility as the outcome and input characteristics of pres-
sure and temperature. To analyze and model (Figure 9.13) the data, three models –
the multilayer perceptron (MLP), the kernel ridge regression, and GPR – have been
used. Finally, the hyper-parameters of the three models (Figure 9.13) were optimized
using the Bat optimization algorithm (BA), resulting in the tuned models. In the end,
many measures were used to evaluate the models. The GPR model was determined to
be the most efficient based on the R
2
measure. Furthermore, the RMSE criterion pro-
duced an error of 1.96 × 10
–2
, the MAPE criterion produced an error of 2.52 × 10
−2
, and
the MAE value produced an error of 2.90 x 10^-2 [75].
In 2022, Zhang et al. developed an approach to analyze hydrogen solubility across a
large temperature range of 273–433 K using a combination of MD and ML. Various mo-
lecular models were used in ML such as Water models, H
2
models, MD simulation imple-
198 Anchal Sharma et al.
https://t.me/med1917
mentation, Widom test-particle insertion method, and Henry’s constant (Figure 9.14).
Both SPC/E and TIP4P/2005 produced quite similar values for the density of water at am-
bient conditions (298 K and 1 bar) and they compared well with experimental results,
but not at other temperature points. The finite size effects were negligible when the sys-
tem contains sufficiently large, i.e., ~7,500 water molecules, in the simulation cell [76].
In order to improve the solubility of nevirapine (NVP) using solid dispersion tech-
nique and to access drug-polymer interactions with two different polymers, i.e., Hy-
droxypropyl methylcellulose (HPMC K4M) and Eudragit S100 (ES100) by applying
molecular modeling techniques, Huwaimel et al. [78] employed Schrodinger 2021-3 soft-
ware. Hildebrand solubility parameter was used for the solubility parameter calcula-
tions, which is a square root of cohesive energy de nsity, derived from the heat of
vaporization and molar volume of the component . This parameter suggested that
HPMCK4M augmented nevirapine solubility as compared to ES100. Hansen solubility
parameters (HSP) were calculated based on van der Waals forces and electrostatic in-
teractions, and the difference of less than three between these parameters suggests the
probability of higher solubility between the drug and the polymer. The results indicated
a difference of less than three with HPMCK4M, which further indicates its better solu-
Figure 9.12: Evaluation of the solubility of
lenalidomide in supercritical carbon dioxide
using multiple tree-based techniques by
Huwaimel et al. in 2023a.
9 Computational prediction of drug-limited solubility 199
https://t.me/med1917
bility with nevirapine. ANOVA further described the significant difference between
NVP:HPMC, NVP:ES100, and HPMC:ES100 in the case of calculations of both Hildebrand
and Hansen solubility parameters. Hydrogen bond and interaction energy calculations:
Hydrogen bonding is a chief parameter determining drug-polymer interaction and it
has been found that NVP formed relatively stronger and more number of hydrogen
bonds with HPMCK4M than ES100, thus increasing the solubility of NVP with HPMC.
Furthermore, NVP:HPMC showed higher interaction energies than NVP:ES100, thus con-
firming greater miscibility of NVP in HPMC. All the obtained results were found to be in
close association with experimentally predicted solubility data, thus confirming the ac-
curacy of molecular modeling techniques in predicting the solubility of the drug [77].
In Wiercioch and Kirchmair [2], they described the DNN-properties predictor (DNN-
PP) system. In this study DNN-PP, a unique method for predicting small compounds is
related to biological and physicochemical characteristics. To more accurately understand
the representation of the molecules, the well-designed architecture uses two different
blocks of operations. One of the blocks uses stacking attention and combines properties
from both the atom and molecule levels. As a result, the network can collect detailed
information on chemical compounds, such as their skeletal structure and bond and
atom characteristics. Additionally, a group of molecular descriptors are processed by
the second block. Here, they try to increase the generalizability of the molecular charac-
teristics. Although several methods for predicting molecular characteristics have already
been established, our design can collect extensive information about the link between
structure and property. The effectiveness of the suggested DNN-PP was evaluated using
Figure 9.13: Drug solubility in supercritical CO
2
was estimated using ML analysis by Huwaimel et al. in
2023b.
200 Anchal Sharma et al.
https://t.me/med1917
a common molecular ML (Figure 9.15) benchmark. The outcomes indicate that approach
performs better than cutting-edge models. Consider the prediction of hydration-free en-
ergies. DNNPP has the lowest predictive error on the test set, or 0.73, whereas the best
competing technique reports an error of 1.07 on unobserved data. Also, DNN-PP may be
applied in many other disciplines, such as natural language processing [2, 78].
In an attempt to predict the solubility of bisacodyl in isopropanol, methanol and
ethanol, Rao et al. [80] used three computational models to analyze the connection
between solubility and experimental temperature and cosolvent composition [79].
Jouyban-Acree model is given by the equation:
In
WT
= w
1
In x
IT
+ w
2
In x
2T
+
w
1
w
2
T
X
2
i=0
J
i
ðw
1
− w
2
Þ
i
(9:16)
Modified van’t-Hoff Jouyban-Acree model is represented by the equation:
In x
w1T
= D
1
+
D
2
T
+ D
3
w
1
+ D
4
w
1
T
+ D
5
w
2
1
T
+ D
6
w
3
1
T
+ D
7
w
4
1
T
(9:17)
Modified Wilson model is given by the equation:
Figure 9.14: Approach developed by Zhang et al. (2022) to analyze hydrogen solubility across a large
temperature range using a combination of MD and ML.
9 Computational prediction of drug-limited solubility 201
https://t.me/med1917
In x
w
1T
= 1 −
w
1
1 + In x
1
ðÞ½
w
1
+ w
2
λ
12
−
w
2
1 + In x
2
ðÞ½
w
2
+ w
1
λ
21
(9:18)
With respect to the prioritized target-carboxy muconolactone decarboxylase (CMD),
Wiercioch and Kirchmair [2] sought to anticipate the binding potential of two herbal-
based compounds by contrasting them with the binding of co-crystallized substrate. γ-
CMD was chosen as the potential molecular target based on its functional involvement
in metabolic pathways, and the three-dimensional (3D) structure, which was not avail-
able in its original state, was computationally modeled and confirmed. Based on litera-
ture review, databases search and 3D structure data availability, 25 natural compounds
were selected and their drug likeliness, pharmacokinetics, and toxicities were computa-
tionally computed. Utilizing computational modeling, molecular docking, and dynamic
studies, the potential lead candidates among the chosen chemicals were found to be hir-
sutine and thymoquinone. When compared to the interaction of co-crystallized inhibi-
tors, these natural compounds demonstrated considerable binding to CMD. According
to research, plant-based compounds like hirsutine and thymoquinone are likely effec-
tive binders. Future treatment therapies against MDRAb could advance significantly
with a move toward CMD and computational prediction [80].
Figure 9.15: Deep neural network properties predictor (DNN-PP) system demonstrated.
202 Anchal Sharma et al.
https://t.me/med1917
9.6 Conclusion
In this chapter, we extensively discussed the computational prediction of drug-limited sol-
ubility and CYP450-mediated biotransformation. Given their critical impact on a drug’s
efficacy and safety profile, it is imperative to accurately evaluate these factors. Various
computational approaches, such as QSPR models, MD simulations, and ML algorithms,
have demonstrated promising results in estimating drug solubility, allowing researchers
to prioritize drug candidates with optimal solubility profiles. Similarly, CYP450-mediated
biotransformation plays a pivotal role in drug metabolism and clearance, and computa-
tional methods, including molecular docking, pharmacophore modeling, and QSAR mod-
els, are of immense importance in predicting CYP450-mediated biotransformation. These
approaches have provided valuable insights into drug metabolism, aiding in the design of
safer and more effective drugs.
As evidence is the potential of ML models to predict CYP450-mediated toxicity.
Three models that performed best to predict toxicity are NLSD-XGB, CypReact, and
PoSMNA model, with accuracy rates of 0.98, 0.92, and 0.92, respectively. Similarly,
the models that are found to be best for solubility prediction are NuSVR and SCA-GB,
with R
2
value of 0.998 and 0.997, respectively. A key area for future exploration lies
in the integration of the abovementioned models and approaches. By combining ML
algorithms, it is possible to enhance the ac curacy and reliability of predictions. It is
also crucial to deploy more comprehensive and accurate databases for solubility and
biotransformation data, as this would facilitate the development of more robust pre-
diction models and enable the assessment of their generalizability. Additionally, this
would help guide precision medicine approaches.
Figure 9.16: Predict the solubility of bisacodyl by Rao et al. [80].
9 Computational prediction of drug-limited solubility 203
https://t.me/med1917