Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5338_Библиотеки_им_академика_М_И_Перельмана
.pdf
3 Data-Driven Molecular Structure Generation for Inverse QSPR/QSAR … 51
https://t.me/med1917
Fig. 3.2 Derivation of the posterior distribution using the Gaussian mixture model combined with
component-wise multivariate linear regression (GMM/cMLR). Using compounds represented as
dots, MLR regression models (lines) and p(x) are modeled. The posterior distribution of x given
the y
value p(x|y = y1) is depicted on the right
1
Each component on the right side of Eq. (3.4) is analytically tractable because p(x)
and p(y|x,z
of the posterior as a mixture of Gaussians from Eq. (
6]. The overview of this approach for B-IQW, termed the GMM/cMLR approach,
[
is illustrated in Fig.
determined using compounds (Fig.
) comprise Gaussian distributions (Eq. (3.1) and Eq. (3.2)). A derivation
k
3.4) is presented in the literature
3.2. In practice, the parameters of cMLR and p(x) are first
3.2 left). This is followed by the derivation of
p(x|y) given a specific y value. According to the density of data distribution and
the distance to the specific y value, the weights of Gaussians in p(x|y) may differ
appreciably from those of p(x) as depicted in Fig.
3.2 right.
GMR is slightly different from the other two approaches. In GMR, x and y are
combined as a vector, and the combined vector is represented as a GMM (Eg.
3.5).
p(x, y) =
k=1
Here, µ
and Δ
δ
kyy
kx
kx
and δ
and μ
are the means of x and y for the k-th component of the Gaussians,
ky
are the covariance and variance of x and y, respectively. δ
ky
represent the relation between x and y for the k-th Gaussian. As in the other two
θkN
µ
μ
Δ
kx
δ
ky
|
M
kx δkxy
kyx δky
kxy
(3.5)
and
approaches, the posterior distribution of x given a specific y value can be analytically
derived. A derivation has been presented in a comprehensive review of machine
19
learning methods [
].
Once p(x|y) is derived, regions for structure generation are heuristically determined according to the density of the posterior; e.g., a hyper-rectangle is determined
by setting upper and lower bounds for each descriptor.

52 T. Miyao a nd K. Funatsu
https://t.me/med1917
3.2.2.2 Generation of the Structure for a Specific Region in Chemical
Space
The Bayesian approach can derive regions in chemical space. The next step is
to generate chemical structures or chemical graphs that satisfy the descriptor
constraints. B-IQWs are distinguishable from F-IQWs in that they incorporate the
descriptor constraints into the generation process and guarantee the exhaustiveness
of the generated chemical structures under certain constraints.
Structure generation in the B-IQW aims to efficiently generate chemical structures satisfying the descriptor constraints. If derived constraints were not used in the
generation process, B-IQW would be identical to F-IQW, such as approaches based
on the genetic algorithm. To incorporate descriptor constraints, we have proposed
using monotonically changing descriptors (MCDs), which are a set of descriptors
whose values monotonically change as an atom or fragment is appended to a chemical structure [
MCDs. The proposed generation algorithm grows chemical structures by attaching
atoms or fragments one by one. When an MCD value of a growing chemical structure exceeds the upper bound of a constraint, the structure is pruned on a generation
tree. A generation tree is a tree whose nodes are generated chemical structures and
whose edges are connections between parent and child nodes. The chemical structure at a child node is generated by appending a fragment to the chemical structure
at its parent node. The concept of structure generation using MCDs is illustrated in
3.3. Details of the structure generation algorithm using a constraint defined with
Fig.
MCDs have been given in the literature [
20]. Most molecular size-dependent descriptors are categorized as
20, 21].
3.2.3 Proof-Of-Concept Study: Bioactive Chemical Structure
Design Through the Proposed B-IQW
We have previously reported several proof-of-concept studies on the B-IQW targeting
the exhaustive generation of chemical structures exhibiting a desired property or
activity. For example, boiling points [
B-IQW applications. Furthermore, potential ligands for human alpha 2A adrenergic
receptors were generated by adopting the B-IQW [
design is briefly presented as a proof-of-concept study of the inverse QSAR problem.
3.2.3.1 Thrombin Inhibitor Data Set and Descriptors
Designing thrombin inhibitors as an inverse QSAR problem is challenging owing
to the relatively large molecular size of the inhibitors (vide infra). The conventional
approach to exhaustive structure generation by atom combination cannot handle large
chemical structures owing to the combinatorial explosion problem. Our structure
] and aqueous solubility [6] were targeted as
8
6, 21]. Herein, thrombin inhibitor

3 Data-Driven Molecular Structure Generation for Inverse QSPR/QSAR … 53
https://t.me/med1917
Fig. 3.3 Generation of chemical structures using monotonically changing descriptors (MCDs).
Transition in the chemical space spanned by the molecular weight (MW) and the number of aromatic
rings (A-rings) from structure 1 to 3 is depicted on the left. The right shows the growing structures
with descriptor values
generation strategy of combining fragments and adding atoms is a solution to this
problem.
From the ChEMBL database ver. 20 [22], 3259 small molecules for thrombin
(CHEMBLID 204) annotated with K
were downloaded. These small molecules
i
were further curated with an in-house script, and the remaining 1705 compounds
were used in the analysis. The K
values were transformed to pKi values as the
i
objective variable (y). The averages of the molecular weight and the number of
heavy atoms for the data set were 475.4 and 33.9, respectively. The range of y was
(1.00, 12.2), i.e., there was a great range of potencies for this target.
Fifty-one MCDs were adopted. Variable selection based on descriptor information provided 27 MCDs. These 27 MCDs included substructure descriptors (e.g., the
number of aromatic rings, the number of CH
R substructures (R: carbon atom), and
3
the number of aromatic ketones), topological descriptors (the first Zagrev index),
the topological polar surface area (TPSA), and topological pharmacophore distance
descriptors. A topological pharmacophore distance descriptor is the summation of
topological distances between specified types of potential pharmacophore point
20
(PPP) [
]. Hydrogen bond acceptors and donors, lipophilic atoms (e.g., carbon
atoms adjacent only to carbon atoms), positively and negatively ionizable atoms, and
aromatic rings were adopted as PPPs. The summation of the topological distances
between a pair of PPPs is expected to represent the pharmacophore information of
].
inhibitors like chemically advanced template search descriptors [
23

54 T. Miyao a nd K. Funatsu
https://t.me/med1917
3.2.3.2 Model Construction and Derivation of P(X|y)
All 1705 compounds represented by the 27 MCDs were used to train GMM and
cMLR models. Eight components in a GMM had the minimum Bayesian information criterion value among different numbers of components. The GMM model was
created using the mclust package [
2
value of 0.656. The density of p(x) and the densities of p(x|y) for several y
an R
values were projected on the generative topographic mapping (GTM) [
tion to the centers of eight Gaussians (Fig.
24] in the programming language R. cMLR reached
25] in addi-
3.4). The GTM created a manifold in the
27-dimensional MCD space so that the nonlinear manifold had an optimal likelihood
for the training compounds. Figure
3.4 clearly shows that the posterior distributions
were different at the centers of the Gaussians and in highly dense regions on the
manifold.
To generate highly potent compounds within the AD of the cMLR model, the
target pK
value was set at 11 and the centers of the eight Gaussians were focused on.
i
The centers of the Gaussians were expected to show high density, such that structure
generation aiming at the centers would produce reasonable chemical structures within
Fig. 3.4 p(x)and p(x|y) distributions on the generative topographic mapping (GTM) manifold.
a The prior density p(x) is shown in grayscale on the GTM manifold, and t he eight centers of
Gaussians are also shown on the manifold. b, c, d The posteriors of p(x|y)given y = 3, 7, and 11
are shown

3 Data-Driven Molecular Structure Generation for Inverse QSPR/QSAR … 55
https://t.me/med1917
Fig. 3.5 Predicted pK
against the density of the
center of each Gaussian.
Dots are centers of
Gaussians and the numbers
identify the Gaussians, as
shown in Fig.
is the logarithm of the
density of p(x|y = 11)
i
3.4.The x-axis
the AD. According to the relation between the predicted pKi and the density at
the centers (Fig.
3.5), the center of Gaussian 1 was selected as a target, and the
coordinates of the center were translated to constraints. The density and predicted
value for the center of Gaussian 1 were 2.84 × 103 and 9.61, respectively. It is
pK
i
noted that, although aiming for a pK
value of 11, the predicted pKi values at the
i
centers of the Gaussians were approximately 9.5 and lower. This was because the
posterior distribution inherits the prior distribution via the Bayes theorem. As the
distribution for the 1705 compounds was unimodal and the average was 6.6, the
pK
i
posterior was biased toward the prior to a certain extent.
3.2.3.3 Structure Generation Based on P(X|y)
From the coordinates of the center of Gaussian 1 (Fig. 3.5), target coordinates or a
grid point were determined by adjusting the coordinates for discrete variables; e.g.,
the number of aromatic rings. MCD constraints were then determined heuristically so
that the constraint hyperrectangle covered the grid point. For example, the coordinates
of the grid point were the number of rings: 2, the number of aromatic rings: 1, TPSA:
112, summation of topological distances between lipophilic points (LL): 125, and
so on. We set constraints (lower–upper bounds) at 2–3 for the number of rings, 1–1
for the number of aromatic rings, 104–120 for the TPSA, 92–158 for LL, and so on.
Chemical structures were generated by combining atom and ring fragments as
long as they satisfied the descriptor constraints. A total of 289 ring fragments were
prepared by dissecting the 1705 compounds into ring systems using the Bemis–
Murcko scaffold definition [
were used as atom fragments. The generation of potentially reactive substructures like
26], and C, N, O, F, Cl, Br, and I with a single valence rule

56 T. Miyao a nd K. Funatsu
https://t.me/med1917
Fig. 3.6 Generated chemical structures for thrombin inhibitors. The blue circles are the coordinates
of the generated chemical structures, and the red square is the target grid point for the generation of
structures on the space spanned by the predicted pK i and the density of p(x|y = 11). Three chemical
structures A, B, and C are shown by their structural formulas, and the corresponding coordinates
are shown on the scatter plot
peroxides was suppressed. Furthermore, we set further constraints to avoid combinatorial explosion, e.g., we constrained the number of fragments to be combined
from 2 to 15 and introduced a probabilistic operation to randomly prune a growing
generation tree with a probability of 0.4. As random pruning was introduced, we
conducted three trials with different random seeds.
The total number of generated chemical structures satisfying the constraints was
11
1739 while 4.25 × 10
structures were built on generation trees. The expected
number of eligible structures without introducing probability generation would be
1.41 × 10
8
and this would require the searching of 6.19 × 10
15
structures. The latter
number is out of our reach and highlights the necessity of avoiding a combinatorial explosion. The densities of p(x|y = 11) and predicted pK
generated chemical structures are shown Fig.
3.6-left. On the scatter plot, the red
values for the 1739
i
square is the target grid point, and the chemical structures residing nearby on the
Pareto frontier are reported in Fig.
3.6-right. These chemical structures show desired
features based on the QSAR and AD models, and they might be useful for further
development of the chemical structures.
3.3 Related Studies and Recent Trends in Research
on the Inverse QSPR/QSAR Problem
Various approaches have been developed since our proposal of adopting the Bayesian
approach for the inverse QSPR problem.

3 Data-Driven Molecular Structure Generation for Inverse QSPR/QSAR … 57
https://t.me/med1917
Ikebata et al. introduced the use of the simplified molecular-input line-entry system
(SMILES)-based chemical language model to generate novel chemical structures by
mutating sentences of chemical structures [
in the formalization of the posterior distribution of the chemical structure given a
desirable probability on y. In their approach, SMILES strings of chemical structures
are mutated using a mutation function, and the mutated strings are then evaluated via
QSAR models to obtain an updated probability of the posterior, like in F-IQWs and
Markov chain Monte Carlo sampling. Sampling based on the updated probabilities of
strings leads to a set of SMILES strings for the next generation. Here, the mutation
function plays a central role in the sampling of chemical structures based on the
posterior distribution. Unfortunately, in their validation scheme, the diversity of the
generated chemical structures in terms of the true compound distribution was not
described, and the performance of the approach was not compared with that of a
genetic algorithm.
The methodological development and potential of GMR in solving the inverse
QSPR problem were intensively investigated by Kaneko et al. for hyperparameter
optimization in GMR [
the posterior [
materials [
A current trend of structure generation for the inverse QSPR problem is to
adopt generative modeling [
of generative models on a massive number of chemical structures is efficient for
the s ampling of sample chemical structures following the distribution of training
chemical structures. Direct sampling from the distribution given a specific y value
using conditional generative models has been proposed [
ative models are effective for reproducing the probability distribution, the B-IQW
explained in this chapter does not lose its value owing to its transparent structure
generation process focusing on a specific region of the chemical space. Moreover,
leveraging advanced computational power, exhaustive chemical structure generation is important to molecular design projects in terms of not missing “interesting”
chemical structures as hypothesis compounds drawn from data.
29]. The method was further developed for the design of thermochromic
30] and the derivation of synthesis conditions in artificial bone design [ 31].
28] and optimal coordinate detection at a maximal density of
32
], mostly deep-learning-based modeling. The training
27]. The novelty of this approach lies
33]. Although these gener-
3.4 Conclusions
In this chapter, we discussed approaches for the inverse QSPR and QSAR problems for molecular design from a methodological point of view. F-IQWs adopt
QSPR models for filtering generated chemical structures like VS, whereas B-IQWs
construct chemical structures from building blocks (fragments) satisfying descriptor
constraints at a generation level.
One way to tackle backward approach is to derive the posterior distribution of
descriptors: x given a specific property value and y for generating chemical structures
within the AD of a model. For that purpose, GMM and cMLR were introduced,
and via the Bayes theorem, the posterior distribution was analytically derived. As a

58 T. Miyao a nd K. Funatsu
https://t.me/med1917
demonstration of the proposed approach, chemical structures of thrombin inhibitors
were designed. In the demonstration, novel chemical structures with desirable and
trustable property values based on a QSPR model were generated. We expect that
this intuitive approach of generating chemical structures will assist experimental
chemists and biologists in enumerating virtual compounds as a data hypothesis.
References
1. Kang P, Liu Z, Abou-Rachid H, Guo H (2020) Machine-Learning Assisted Screening of
Energetic Materials. J Phys Chem A 124:5341–5351.
2647
2. Ding Y, Chen S, Liu H, et al (2023) Discovery of Multitarget Inhibitors against Insect
Chitinolytic Enzymes via Machine Learning-Based Virtual Screening. J Agric Food Chem
71:8769–8777.
3. Wong F, Zheng EJ, Valeri JA, et al (2023) Discovery of a Structural Class of Antibiotics with
Explainable Deep Learning. Nature 626(7997):177–185.
06887-8
4. Brown N, McKay B, Gasteiger J (2006) A Novel Workflow for the Inverse QSPR Problem
Using Multiobjective Optimization. J Comput Aided Mol Des 20:333–341.
1007/S10822-006-9063-1
5. Jensen JH (2019) A Graph-Based Genetic Algorithm and Generative Model/Monte Carlo Tree
Search for the Exploration of Chemical Space. Chem Sci 10:3567–3572.
1039/C8SC05372C
6. Miyao T, Kaneko H, Funatsu K (2016) Inverse QSPR/QSAR Analysis for Chemical Structure
Generation (from y to x). J Chem Inf Model 56:286–299.
5B00628
7. Gantzer P, Creton B, Nieto-Draghi C (2020) Inverse-QSPR for de novo Design: A Review.
Mol Inform 39:1900087.
8. Miyao T, Arakawa M, Funatsu K (2010) Exhaustive Structure Generation for Inverse-QSPR/
QSAR. Mol Inform 29:111–125.
9. Wong WW, Burkowski FJ (2009) A Constructive Approach for Discovering New Drug Leads:
Using a Kernel Methodology for the Inverse-QSAR Problem. J Cheminform 1:1–27.
doi.org/10.1186/1758-2946-1-4
10. Churchwell CJ, Rintoul MD, Martin S, et al (2004) The Signature Molecular Descriptor: 3.
Inverse-Quantitative Structure–Activity Relationship of ICAM-1 Inhibitory Peptides. J Mol
Graph Model 22:263–273.
11. Skvortsova MI, Baskin II, Slovokhotova OL, et al (1993) Inverse Problem in QSAR/QSPR
Studies for the Case of Topological Indices Characterizing Molecular Shape (Kier Indices). J
Chem Inf Comput Sci 33:630–634.
12. Faulon JL, Churchwell CJ, Visco DP (2003) The Signature Molecular Descriptor. 2. Enumerating Molecules from Their Extended Valence Sequences. J Chem Inf Comput Sci 43:721–734.
https://doi.org/10.1021/CI020346O
13. Visco DP, Pophale RS, Rintoul MD, Faulon JL (2002) Developing a Methodology for an Inverse
Quantitative Structure–Activity Relationship Using the Signature Molecular Descriptor. J Mol
Graph Model 20:429–438.
14. Dragos H, Gilles M, Alexandre V (2009) Predicting the Predictability: A Unified Approach
to the Applicability Domain Problem of QSAR Models. J Chem Inf Model 49:1762–1776.
https://doi.org/10.1021/CI9000579
15. Klingspohn W, Mathea M, Ter Laak A, et al (2017) Efficiency of Different Measures for
Defining the Applicability Domain of Classification Models. J Cheminform 9:1–17.
doi.org/10.1186/S13321-017-0230-2
https://doi.org/10.1021/ACS.JAFC.3C00633
https://doi.org/10.1002/MINF.201900087
https://doi.org/10.1002/MINF.200900038
https://doi.org/10.1016/J.JMGM.2003.10.002
https://doi.org/10.1021/CI00014A017
https://doi.org/10.1016/S1093-3263(01)00144-9
https://doi.org/10.1021/ACS.JPCA.0C0
https://doi.org/10.1038/s41586-023-
https://doi.org/10.
https://doi.org/10.
https://doi.org/10.1021/ACS.JCIM.
https://
https://

3 Data-Driven Molecular Structure Generation for Inverse QSPR/QSAR … 59
https://t.me/med1917
16. Gaspar HA, Marcou G, Horvath D, et al (2013) Generative Topographic Mapping-Based Classification Models and Their Applicability Domain: Application to the Biopharmaceutics Drug
Disposition Classification System (BDDCS). J Chem Inf Model 53:3318–3325.
org/10.1021/CI400423C
17. Berenger F, Yamanishi Y (2019) A Distance-Based Boolean Applicability Domain for Classification of High Throughput Screening Data. J Chem Inf Model 59:463–476.
10.1021/ACS.JCIM.8B00499
18. Neath AA, Cavanaugh JE (2012) The Bayesian Information Criterion: Background, Derivation,
and Applications. Wiley Interdiscip Rev Comput Stat 4:199–203.
S.199
19. Stulp F, Sigaud O (2015) Many Regression Algorithms, One Unified Model: A Review. Neural
Netw 69:60–79.
20. Miyao T, Kaneko H, Funatsu K (2016) Ring System-Based Chemical Graph Generation for de
novo Molecular Design. J Comput Aided Mol Des 30:425–446.
822-016-9916-1
21. Miyao T, Kaneko H, Funatsu K (2014) Ring-System-Based Exhaustive Structure Generation
for Inverse-QSPR/QSAR. Mol Inform 33:764–778.
22. Gaulton A, Hersey A, Nowotka ML, et al (2017) The ChEMBL Database in 2017. Nucleic
Acids Res 45:D945–D954.
23. Reutlinger M, Koch CP, Reker D, et al (2013) Chemically Advanced Template Search (CATS)
for Scaffold-Hopping and Prospective Target Prediction for ‘Orphan’ Molecules. Mol Inform
32:133–138.
24. Scrucca L, Fraley C, Murphy TB, Raftery AE (2023) Model-Based Clustering, Classification,
and Density Estimation Using Mclust in R. Chapman & Hall/CRC Press.
25. Bishop CM, Svensén M, Williams CKI (1998) GTM: The Generative Topographic Mapping.
Neural Comput 10:215–234.
26. Bemis GW, Murcko MA (1996) The Properties of Known Drugs. 1. Molecular Frameworks. J
Med Chem 39:2887–2893.
27. Ikebata H, Hongo K, Isomura T, et al (2017) Bayesian Molecular Design with a Chemical
Language Model. J Comput Aided Mol Des 31:379–391.
0008-Z
28. Kaneko H (2021) Extended Gaussian Mixture Regression for Forward and Inverse Analysis.
Chemom Intell Lab Syst 213:104325.
29. Kaneko H (2022) True Gaussian Mixture Regression and Genetic Algorithm-Based Optimization with Constraints for Direct Inverse Analysis. Sci Technol Adv Mater Methods 2:14–22.
https://doi.org/10.1080/27660400.2021.2024101
30. Shimizu N, Kaneko H (2020) Direct Inverse Analysis Based on Gaussian Mixture Regression
for Multiple Objective Variables in Material Design. Mater Des 196:109168.
10.1016/J.MATDES.2020.109168
31. Motojima K, Shiratsuchi R, Suzuki K, et al (2023) Machine Learning Model for Predicting
the Material Properties and Bone Formation Rate and Direct Inverse Analysis of the Model for
New Synthesis Conditions of Bioceramics. Ind Eng Chem Res 62:5898–5906.
10.1021/ACS.IECR.3C00332
32. Sousa T, Correia J, Pereira V, Rocha M (2021) Generative Deep Learning for Targeted
Compound Design. J Chem Inf Model 61:5343–5361.
1496
33. Kang S, Cho K (2019) Conditional Molecular Design with Deep Generative Models. J Chem
Inf Model 59:43–52.
https://doi.org/10.1016/J.NEUNET.2015.05.005
https://doi.org/10.1002/MINF.201400072
https://doi.org/10.1093/NAR/GKW1074
https://doi.org/10.1002/MINF.201200141
https://doi.org/10.1162/089976698300017953
https://doi.org/10.1021/JM9602928
https://doi.org/10.1016/J.CHEMOLAB.2021.104325
https://doi.org/10.1021/ACS.JCIM.0C0
https://doi.org/10.1021/ACS.JCIM.8B00263
https://doi.org/10.1002/WIC
https://doi.org/10.1007/S10
https://doi.org/10.1007/S10822-016-
https://doi.
https://doi.org/
https://doi.org/
https://doi.org/

Chapter 4
https://t.me/med1917
Materials Informatics with Limited Data
Ryo Yoshida
4.1 Introduction
In general, the parameter space for materials research, such as drug developments,
60
is vast. For example, there are approximately 10
ical space of small organic molecules [
1]. On the other hand, the number of small
molecules currently recorded in public databases is on the order of 10
2]. Therefore, a vast, unexplored area remains in the chemical space. Further-
[
candidate molecules in the chem-
8
at most
more, in the development of practical materials, the dimensionality of the parameter space increases with the addition of other design variables such as processing
conditions and the selection of additives, filters, and solvents. The task of materials
informatics (MI) is to identify the unknown parameters that result in the desired
material properties from such a vast parameter space. This is a multi-objective optimization problem. The design of materials, such as drug molecules, is inherently
different from general industrial product design in terms of specificity and diversity
of the parameter space. The parameters take a variety of forms depending on the
problem of interest, including material composition, molecules, crystal structures,
X-ray diffraction spectra, material microstructures, and processing conditions.
The basic workflow of MI consists of forward and inverse prediction tasks
(Fig.
4.1)[3–7
]. The forward problem aims to predict the output
Y for an arbitrary
input X . For example, the input variable is given as a molecule, chemical composition, or crystal structure, and the output is the physical properties or structural
features of the resulting material. In conventional materials research, simulations
based on physical laws, such as first-principles calculations and molecular dynamics
(MD) simulations, have been used for forward prediction tasks. One of the main challenges in MI is to replace computer-intensive calculations, which entail high costs,
R. Yoshida (B)
The Institute of Statistical Mathematics, Research Organization of Information and Systems, 10-3
Midori-Cho, Tachikawa 190-8562, Tokyo, Japan
e-mail: yoshidar@ism.ac.jp
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2024
H. Satoh et al. (eds.), Drug Development Supported by Informatics,
https://doi.org/10.1007/978-981-97-4828-0_4
61
Соседние файлы в папке Библиотека им академика М.И. Перельмана
