Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5435_Библиотеки_им_академика_М_И_Перельмана
.pdf
9.4 CYP450-mediated toxicity prediction based
on ML and artificial intelligence
For the first time, the metabolism mediated by CYP450 3A4, 2D6, and 2C9 was predicted
in the current investigation by Dmitriev et al. [43] using four MMRS-based metabolic pre-
diction models. Aliphatic C-hydroxylation, aromatic C-hydroxylation, N-dealkylation, and
O-dealkylation were four significant biotransformations that took place. Under the as-
sumption that pH = 7.4, a set of topological and physicochemical parameters were de-
rived for each SOM to reflect the reactivity of a possible MMRS. To improve the signal-to-
noise ratio, a negative-screening approach was also implemented. The ideal model for
each metabolic reaction was developed by combining a number of FS approaches and
classification algorithms. The results of the internal and external validations showed
that the models had achieved great success. A revolutionary approach can offer a deeper
understanding for drug metabolism during the drug discovery and development process.
Additionally, it can serve as a foundation for the development of a model to forecast
metabolic pathways [44].
In Xu et al. [45], they developed a multiclass model that can predict which of the
seven CYP450 isoforms will be a compound’s main metabolizing isoform if it is a sub-
strate for CYP450-mediated metabolism. If more than one isoform of a substance is
capable of metabolizing it, the model can also propose which other isoforms are the
next most likely, by ranking the probabilities of the seven isoforms. These models will
make it possible to forecast Phase I metabolite profiles more precisely when multiple
Phase I regioselectivity predictions are provided for a chemical based on several
CYP450 isoforms. The avoidance of DDIs is a crucial feature of CYP450-related metabo-
Figure 9.7: Computational model that predicts metabolism-mediated DDIs used by [41].
184 Anchal Sharma et al.
https://t.me/med1917

lism. The model exhibits significant enrichment over randomized models, with a top-1
criterion success rate of 76% and a top-2 criterion success rate of 88% of information
about drug metabolism including metabolites and the enzyme responsible [45].
In [45], they developed a multilabel learning task to predict CYP450 enzyme-
substrate selectivity. They tested four categories of features (physiochemical property
descriptors , extended connectivity fingerprints, mol2vec descriptors, and molecular
access system key fingerprints) and identified the best combination of features. Next,
they applied the seven different models, all of which outperformed the previous
work. The NLSD-XGB model achieved the best performance (Figure 9.8) with an aver-
age top-1 prediction success of 91.1 %, an average top-2 prediction success of 96.2%,
and an average top-3 prediction success of 98.2%, compared to previous work. NLSD-
XGB showed a significant improvement – over 11% on top-1 by the hold-out method
repeated 10 times. The network-based label space division model was first introduced
in drug metabolism, and it performs efficiently for the task [46].
Figure 9.8: A multilabel learning task developed to predict CYP450 enzyme-substrate selectivity by He
et al. [45].
9 Computational prediction of drug-limited solubility 185
https://t.me/med1917

In He et al. [45], they evaluated ADMET, based on the PaDEL-1D&2D descriptors and
PubChem fingerprints. The binary classification models for CYP1A2, 2C9, 2C19, 2D6, and
3A4 were developed using three typical ensemble learning methods (RF, GBDT, and
XGBoost) and two representative deep learning methods (DNN and CNN). The 10-fold
cross-validation was used to evaluate each classifier’s capacity for generalization, and
predictions on the external test set were used to verify each classifier’s actual capacity
for prediction. The outcomes show that for the external test sets, ensemble learning
models typically provide better predictions than deep learning models. The XGBoost
models beat all other models in terms of classification ability, outperforming even
the multitask deep autoencoder neural network model that was previously reported
(88.5%) with an average prediction accuracy of 90.4% for the test sets. The models were
then interpreted and the misclassified molecules were analyzed using the Shapley addi-
tive explanation approach. The key chemical characteristics provided by our models
are consistent with the structural preferences for inhibitors of various CYP450 isoforms,
which may be helpful in early drug discovery for spotting potential DDIs [47].
The five primary CYPs isoenzymes that make up more than 80% of the metabo-
lism of clinical drugs – CYP1A2, CYP2C19, CYP2D6, CYP2C9, and CYP3A4 – are the focus
of the SuperCYPsPred web server. In Shan et al. [47], they categorized that the predic-
tion models for CYPs inhibition are built using tried-and-true ML techniques. The
models were validated by both cross-validation and external validation sets and they
performed well. The web server accepts a 2D chemical structure as input and outputs
the CYP inhibition profiles of the chemical for 10 models using various molecular fin-
gerprints, along with the confidence scores, similar compounds, known CYPs informa-
tion of drugs – published in literature, detailed interaction profiles of individual
cytochromes, including a DDIs table and an overall CYPs prediction radar chart
(http://insilico-cyp.charite.de/SuperCYPsPred/). The web server is free to access and
neither logging in nor registering is necessary [48].
In this study, Wu et al. [48] developed CyProduct, an entirely novel, precise in silico
metabolic prediction suite for human CYP450 metabolism. Three modules make up the
software suite: (1) CypBoM, a module that precisely predicts the reactive bond locations
within the query molecule, (2) CypReact, a previously described module that determines
whether the query compound is a reactant for a specific CYP450 enzyme, and (3) Metabo-
Gen (Metabolite Generator), a knowledge-based module that creates the resulting meta-
bolic byproducts based on predictions from CypBoM. This package’s usage of a novel
approach to handle reactive sites is crucial to its significant performance gain over previ-
ous tools (both commercial and open source). In particular, CyProduct uses the “bond of
metabolism” (BoM), which provides a clear and succinct method of depicting the meta-
bolic reactions in terms of chemical bonds (rather than atoms) as opposed to the usual
SoM representation. They precisely defined this BoM concept and then used this defini-
tion to produce two publicly accessible BoM data sets, EBoMD and EBoMD2. EBoMD con-
tains 2,262 BoMXY out of 26,420 candidate bonds derived from the 679 compounds in
Zaretzki’s CYP450 data set, and EBoMD2 contains 212 BoMXY out of 3,830 bonds from an
186 Anchal Sharma et al.
https://t.me/med1917

additional 98 compounds. Then they developed the nine CYP450 BoM classifiers in the
CypBoM family, each of which can precisely predict the reactive BoMs for any given sub-
strate molecule. They developed CyProduct, a fully operational in silico metabolism pre-
dictor, by combining CypBoM with our previously created substrate predictor (CypReact)
and a new rule-based module called MetaboGen. We show that CyProduct is significantly
more thorough and accurate than any other in silico metabolism predictor currently
available. In terms of forecasting metabolites for each of the nine key CYP450 enzymes,
our empirical findings demonstrate that CyProduct gets very strong Jaccard scores. The
results show that our scores are substantially better than random chance (see cross-
validation results – they are much better than other in silico metabolism prediction soft-
ware programs that are available publicly or at pubs.acs.org/jcim [49]).
There has been a sharp increase in the usage of herbal medicines alongside conven-
tional treatments in the modern world. Estimating the effects of herbs and medications
remains difficult. Based on in vitro and in vivo approaches, herb-drug interactions can
be anticipated. The current in vivo procedure takes more time, whilst the in vitro techni-
ques are quite expensive. As a result, a novel method was developed by Banerjee et al.
[49] to assess the potential for bioactive substances to interact with the CYP3A4 enzyme
utilizing an online tool. A chosen bioactive compound’s cytochrome activity can be accu-
rately predicted by the Herb-CYP450 Enzyme Inhibition Predictor online database in a
very straightforward, user-friendly interference. Since the database was built using rep-
utable sources like PubChem, the accuracy of the data is good [50].
In Tian et al. [50], they used RF and XGBoost to build computational models based
on four different descriptors (2D, CATS, ECFP4, and MACCS) for substrates and inhib-
itors of five significant CYP450 isoenzymes. They took mechanism-specific metabolic
DDIs caused by CYP450 as the pivotal point in this study. The differences in inhibitor
and substrate models’ predictive abilities, when using RF and XGBoost, showed that
models based on datasets with more chemical skeletons and optimal modelling tech-
niques have a wider application domain. As a result, the derived DDI models were
more trustworthy and useful in future applications. The RF and XGBoost models were
combined to create a number of consensus models, which were used to lower the
model uncertainty [51].
In order to anticipate the possible inhibitors of CYP1A2, CYP2C9, CYP2C19, CYP2D6,
and CYP3A4, a unified GCNN model with an attention mechanism was built by Sarvesh
et al. [51]. Overall, the established GCNN model performed well when used to predict
the CYP inhibitory potencies of unidentified drugs. The chemical ligands of CYP were
identified via molecular graph, which is the most direct and straightforward molecular
encoding approach, without manual structure description, in comparison to the predic-
tion methods that are currently in use. The most notable advantage of the GCNN model
over other deep learning techniques is the facilitation of the determination of the im-
portant structural information of inhibitory potency due to the inclusion of the atten-
tion mechanism. It should be noted that 42 important substrate-binding residues that
make up the five CYP isoforms were used to encode them as pseudo sequences. How-
9 Computational prediction of drug-limited solubility 187
https://t.me/med1917

ever, it is possible that the crucial residues for various ligands to attach to various CYP
isoforms are not always the same, and this merits more research. Even if the weighted
random sample technique was used, the test set’s sensitivity values are still far lower
than their specificity values. Therefore, future studies should concentrate on developing
high-quality benchmark datasets and biased sampling methods [52].
In Wang et al. [52], they integrated a new interaction relationship in HeTOP by 14
enzymes, linking a total of 776 medication components. These findings indicate that for
2,493 patients at a hospital, bisoprolol with amiodarone through enzymatic inhibition is
the concurrent prescription that may result in a cytochrome interaction. P-glycoprotein,
cytochrome P-450 CYP2B6, and cytochrome P-450 CYP2D6 were the key enzymes in play.
The EDSaN CDW queries that were sent enabled the simultaneous spotlight on the most
prescribed molecules that may be accountable for cytochromic interactions. In a further
step, it would be intriguing to assess the actual clinical impact by looking for potential
negative impacts of these interactions in the patient files such as say, hemorrhages [53].
This work done by Qiu et al. [53] used theoretical and experimental research to
structurally characterize the important cytotoxic molecule of 4FPP using PES, optimiza-
tion, and vibrational spectra. These experiments confirmed that the 4FPP molecule has a
stable structure in both gas and solution phases. The aqueous phase exhibited lower op-
timization energy (−1,374,719.096 kJ/mol) and higher solvation energy (20.09 kJ/mol) than
other liquid phases. Strong interactions about the title chemical’s structure were re-
vealed by AIM and RDG analyses, while LOL and EFL tests discovered that the com-
pound’s C7-C10-C11 and C8-C4-C5 regions had the strongest delocalized zones. The FMOs
energies and MEP analysis showed that the compound was more reactive in the aqueous
phase and that the ethanol solvent supplied the greatest UV-vis electronic excitation
when the electronic properties of the title compound were evaluated in solutions for
chemical reactivity parameters. Water solvent also exhibited the compound’s highest hy-
perpolarizability (10.4552 1030 e.s.u.). The highest stabilization energy of the 4FPP mole-
cule in the NBO analysis was influenced by the transition of
✶
. According to the docking
analysis’s findings, the named chemical effectively interacts with each of the cancer cell
proteins that were the target. The ROR2 protein’s membrane-proximal epitope on cancer
cells has a maximum binding energy of about 6.80 kcal/mol (PDB code: 6OSH) [54].
9.5 Computational prediction of drug
limited solubility
Large and lipophilic novel chemical entities are still frequently chosen in contempo-
rary drug development programs using high-throughput screening and combinatorial
chemistry. This is true despite the fact that they have poor aqueous solubility [56–58],
that there is a greater understanding of associated issues, and that there are several
188 Anchal Sharma et al.
https://t.me/med1917

mnemonic rules for avoiding substances with low or variable absorption and pharma-
cokinetics [59–61].
More crucially, the solubility of the target compounds cannot be assessed until they
have been synthesized. On the other hand, quick computational predictions of solubility
can be made on huge compound libraries without the need to synthesize the molecules.
Because there is less demand for pricey simulated or aspirated intestinal fluid, the ex-
penses connected with the pharmaceutical profiling cycle are lessened. This gives me-
dicinal chemists solubility profiles on which they can base better-informed decisions.
The use of BDM or, even better, HIF to forecast solubility is therefore highly recom-
mended. There have been many models established for the prediction of intrinsic aque-
ous solubility (S
0
) or the solubility of the neutral substance [61].
Hughes et al. [62] discussed various in silico models for the prediction of water-
based systems especially in media mimicking the human intestinal fluids which di-
rectly aids in drug discovery and development. The relationship between log P (lipo-
philicity) and solubilization ratio (SR) in water-based solvents, including surfactants,
was first described by Mithani et al. as
SR =
SC
bs
SC
aq
(9:1)
where SC
bs
= solubilization capacity of the bile salt and SC
aq
= solubilization capacity
of water.
Furthermore, linear solvation energy relationship (LEFR), which is based on
Abraham descriptors, is also used to predict drug solubility in fasted state simulated
intestinal fluids (FaSSIF) and can be expressed as
logSE = 0.0678 − 0.1857 × E − 0.3963 × S + 0.5571 × A − 0.9423 × B
0
+ 1.1600 × V (9:2)
where E = excess molar refractivity, S = dipolarity/polarizability, A = hydrogen-bonding
acidity, B
0
= hydrogen-bond basicity, V = McGowan characteristic volume, and SE = solu-
bility enhancement.
This model gave satisfactory results with R
2
of 0.81 and mean absolute error
(MAE) of 0.299. Similarly, Fragerberg et al. calculated descriptors from Dragon pack-
age and PLS methodology to predict the SE in FaSSIF which also gave effective results.
Quantitative structure-property relationships (QSPR) is an effective method that pre-
dicts solubility in biorelevant dissolution media and is a comparatively better model
to predict solubility in physiological fluids especially gastro-intestinal fluids, and it
also gives a clear insight into the effect of the composition of these fluids on dissolu-
tion and solubilization after the oral intake of the drug. However, the disadvantage of
this model lies in its inaccuracy which can be attributed due to the deficiency in the
algorithms and descriptors. To deal with this, Palmer and Mitchell in 2014 modeled
the solubility using the thermodynamic cycle and expressed the relationship between
the intrinsic solubility and the change in Gibbs free energy as:
9 Computational prediction of drug-limited solubility 189
https://t.me/med1917

ΔG
sol
= ΔG
sub
+ ΔG
hydr
= −RTIn S
0
V
m
(9:3)
where ΔG
sol
= Gibbs free energy for solution, ΔG
sub
= Gibbs free energy for sublima-
tion, ΔG
hydr
= Gibbs free energy for sublimation, R = molar gas constant, T = tempera-
ture in Kelvin, S
0
= intrinsic solubility (M), and V
m
= molar volume of the crystal.
When octanol is used as an intermediate between the gaseous and hydrated states
in the experimental assessment of hydration process, then eq. (9.3) can be modified as
ΔG
sol
= ΔG
sub
+ ΔG
solv
+ ΔG
tr
= −RTIn S
0
V
m
(9:4)
where ΔG
solv
is Gibbs free energy for solvation of octanol and ΔG
tr
is the Gibbs free
energy for the transfer of the molecule from octanol to water. In case the log P value
of the molecule is determined, then ΔG
tr
can be replaced by 2.303RT log P.QSPR
model represents a new approach for exploring the effect of solid-state versus hydra-
tion effects on the resulting solubility of the compound [48].
In Wang and Hou [63], they described various linear and nonlinear predictive co-
solvency models in combination with Abraham and Hansen parameters for the pre-
diction of solubility of celecoxib in cosolvency systems. Jouyban-Acree model is the
most accurate mathematical cosolvency model, which provides a good description of
solubility’s dependency on temperature and solvent composition and is independent
of the solute and solvent physicochemical properties. This model can be represented
by the equation:
In
WT
= w
1
In x
1
T
+ w
2
In x
2T
+
w
1
w
2
T
X
2
i=0
J
i
ðw
1
− w
2
Þ
i
(9:5)
It is a combined version of Jouyban-Acree-van’t Hoff model with the Abraham and Han-
sen solubility parameters. The modified Wilson model is represented by the equation:
In x
w
1T
= 1 −
w
1
1 + Inðx
1
½Þ
w
1
+ w
2
λ
12
−
w
2
1 + Inðx
2
½Þ
w
2
+ w
1
λ
21
(9:6)
where w
1
and w
2
are the mass fractions of mono-solvents 1 and 2 in the absence of
solute. x
m
, T, x
1T
, and x
2T
, respectively, are the solubilities of the solute in the solvent
mixes, mono-solvents at temperature T, and J
i
parameters are constants.
Combined version of the modified Wilson model with the Abraham and Hansen
solubility parameters
The fundamental benefit of these general models, which might be useful in the
pharmaceutical sector, is that they can be extended to predict solubility values at dif-
ferent temperatures, and require little-to-no experimental data [62].
In Bergström and Larsson [64], they characterized the solubility of sulfanilamide
and sulfacetamide based on artificial neural networks. The solubility data were inter-
preted in terms of the empirical and semiempirical predictive models such as Buchow-
ski-Ksiazczak model (λh-equation), extended Buchowski’smodel(λβ-equation), modified
190 Anchal Sharma et al.
https://t.me/med1917

Apelblat equation, van’t Hoff-Yaws model, nonrandom two liquid (NRTL) model, Wilson
model, and the Weibull two-parameter extrapolation model. On further applying the
model, the lowest values of AICc parameter were found for the Buchowski-Ksiazczak
model (B1) demonstrating the best correlation with experimental solubility data. Re-
markably, λβ-equation (B2) and van’t Hoff-Yaws model (vHY) gave a fitting of almost the
same quality but slightly less accurate when compared to the B1 approach. The results of
the COSMO-RS simulations were substantially flawed, only offering a qualitative estima-
tion of solubility. Consequently, a nonlinear fully predictive model was created using a
ML approach. The set of utilized molecular descriptors characterizes intermolecular in-
teractions in pure solvents and is derived from COSMO-RS calculations. Based on the ef-
fectiveness of the developed ASNN model, it is feasible to validate that the chosen
collection of descriptors may be used for modeling solubility and it contains the most
crucial data required for modeling saturated solutions. This creates a fresh opportunity
for developing solubility models based on sparse descriptors with strong physical signifi-
cance. This is the first effort, to the best of our knowledge, that successfully applies these
chemical descriptors for solubility modeling [64].
In order to predict the solubility of busulfan in supercritical carbon dioxide (sc-
CO
2
) solvent, Rahimpour et al. [65] developed a novel ML method positioned on Neuro
fuzzy system, i.e., ANFIS (adaptive neuro fuzzy inference system). Grid portioning
technique has been used for the development of Fuzzy system, which further adds to
the accuracy in predicting drug solubility. Trimf was used as the membership function
in the fuzzy structure. The data collected from the extensive literature review was
used to train the ANFIS structure, which was further used for the testing of the busul-
fan solubility, and great accuracy has been observed both in the training and testing
steps, as the collected data was in close association with the experimental values. It
was found that the ANFIS model can be used in future for the prediction of drug solu-
bility in supercritical solvents as a function of process parameters such as pressure
and temperature and can thus be employed for the design and optimization of super-
critical drug production [64].
In Cysewski et al. [66], they studied and applied a unique ML modeling technique
to predict the solubility of two APIs, salsalate and decitabine, in a supercritical solvent
(carbon dioxide) model. Due to its advantages over alternative drug processing meth-
ods, this supercritical approach has been chosen. It is true that determining and cor-
relating a drug’s solubility is crucial for the research, design, and problem-solving of
supercritical-based drug manufacture. Salsalate/decitabine solubility data were col-
lected from the literature in a batch, and models were created using the information.
The initial modeling strategy was based on the thermodynamic technique, and it in-
volved deriving, fitting, and correlating five different empirical correlations that are
derived, fitted, and correlated. The models have three fitting parameters that were
acquired by the use of numerical methods. When the thermodynamic models were
rearranged and compared to the experimental data, t hey showed linear behavior,
demonstrating their capacity to extrapolate salsalate solubility data points (Figure 9.9).
9 Computational prediction of drug-limited solubility 191
https://t.me/med1917

A neural network (JMP Pro 15 software) with one hidden layer and a combination of
efficient activation functions was also taken into consideration while developing an
ML approach. The neural model was trained and validated using 32 trials for each
medication, and the coefficient of determination values more than 0.99 was found for
both medicines. Finally, it was discovered that the neural model is better able to accu-
rately predict the solubility data [65].
In Zhu et al. [67], they evaluated the solubility of tolmetin (an anti-inflammatory medi-
cine) in sc-CO
2
. The two characteristics of the input are temperature and pressure, and
the intended output of this modeling is the solubility of the tolmetin. The investigated
models (Figure 9.10) include gradient tree boosting (GBRT), extra tree (ET), and CART
(regression tree). Three final models were produced after their hyper-parameters were
optimized based on various statistical criteria. Based on the MSE parameter, the error
values for the CART, ET, and GBRT models were 9.79E-08, 5.53E-08, and 1.24E-08, respec-
tively. The GBRT model was proposed as the strongest and most accurate model estab-
lished in this research compared to the other two models based on this fact and other
metrics. The created ML models for the prediction of drug solubility in supercritical sol-
vents showed that these models are reliable enough to be taken into consideration. The
maximum error in this model is 2.11E-04, with R
2
score of 0.976 [66].
In Nguyen et al. [68], a study of the computational task was carried out in order to
predict the solubility of a drug model (Figure 9.11), namely salsalate in sc-CO
2
as the
Figure 9.9: Developed a unique machine learning (ML) modeling technique to predict the solubility of two
APIs, salsalate and decitabine by Nguyen et al.
192 Anchal Sharma et al.
https://t.me/med1917

solvent. Several ML data was employed to predict the solubility data as a function of
the input parameters, i.e., temperature (T) and pressure (P). On the provided data, we
employed linear support vector regression (SVR), Nu SVR, and Bayesian ridge regres-
sion (BRR) models with two inputs including X
1
=P(bar)andX
2
= T(K), and only one out-
put, which is the drug solubility (mole fraction unit). The simulation results indicated
that linear SVR, Nu SVR, and BRR have R
2
scores of 0.784, 0.998, and 0.876, respectively.
In addition, they exhibit MAE error rates of 3.71 × 10
−4
,8.41×10
–5
, and 3.04 × 10
–4
,respec-
tively. Another statistic that was considered in the predictions is RMSE, which indicated
error rates of 4.21× 10
–4
,1.39×10
–4
, and 3.60 × 10
−4
for linear SVR, Nu SVR, and BRR, re-
spectively. Indeed, the model of Nu SVR was chosen as the best model based on these
statistical metrics and some visual examination, and it was utilized to discover optimal
values that can be summarized as a vector: X
1
=400,X
2
= 338, Y = 0.00387 [67].
In Abourehab et al. [69], the solubility of decitabine was characterized on the basis
of computational prediction and optimization. The main objective of this study was to
Figure 9.10: Evaluation of the tolmetin (an anti-inflammatory medicine) solubility in sc-CO
2
dioxide by
Abourehab et al.
9 Computational prediction of drug-limited solubility 193
https://t.me/med1917
Соседние файлы в папке Библиотека им академика М.И. Перельмана
