Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5580_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
30.08.2026
Размер:
43 Мб
Скачать
3 Data-Driven Molecular Structure Generation for Inverse QSPR/QSAR … 51
https://t.me/med1917
Fig. 3.2 Derivation of the posterior distribution using the Gaussian mixture model combined with component-wise multivariate linear regression (GMM/cMLR). Using compounds represented as dots, MLR regression models (lines) and p(x) are modeled. The posterior distribution of x given the y
value p(x|y = y1) is depicted on the right
1
Each component on the right side of Eq. (3.4) is analytically tractable because p(x) and p(y|x,z of the posterior as a mixture of Gaussians from Eq. (
6]. The overview of this approach for B-IQW, termed the GMM/cMLR approach,
[ is illustrated in Fig. determined using compounds (Fig.
) comprise Gaussian distributions (Eq. (3.1) and Eq. (3.2)). A derivation
k
3.4) is presented in the literature
3.2. In practice, the parameters of cMLR and p(x) are first
3.2 left). This is followed by the derivation of
p(x|y) given a specific y value. According to the density of data distribution and the distance to the specific y value, the weights of Gaussians in p(x|y) may differ appreciably from those of p(x) as depicted in Fig.
3.2 right.
GMR is slightly different from the other two approaches. In GMR, x and y are
combined as a vector, and the combined vector is represented as a GMM (Eg.
3.5).
p(x, y) =
k=1
Here, µ and Δ
δ
kyy
kx
kx
and δ
and μ
are the means of x and y for the k-th component of the Gaussians,
ky
are the covariance and variance of x and y, respectively. δ
ky
represent the relation between x and y for the k-th Gaussian. As in the other two
θkN
µ
μ
Δ
kx
δ
ky
|

M
kx δkxy
kyx δky

kxy
(3.5)
and
approaches, the posterior distribution of x given a specific y value can be analytically derived. A derivation has been presented in a comprehensive review of machine
19
learning methods [
].
Once p(x|y) is derived, regions for structure generation are heuristically deter­mined according to the density of the posterior; e.g., a hyper-rectangle is determined by setting upper and lower bounds for each descriptor.
52 T. Miyao a nd K. Funatsu
https://t.me/med1917
3.2.2.2 Generation of the Structure for a Specific Region in Chemical
Space
The Bayesian approach can derive regions in chemical space. The next step is to generate chemical structures or chemical graphs that satisfy the descriptor constraints. B-IQWs are distinguishable from F-IQWs in that they incorporate the descriptor constraints into the generation process and guarantee the exhaustiveness of the generated chemical structures under certain constraints.
Structure generation in the B-IQW aims to efficiently generate chemical struc­tures satisfying the descriptor constraints. If derived constraints were not used in the generation process, B-IQW would be identical to F-IQW, such as approaches based on the genetic algorithm. To incorporate descriptor constraints, we have proposed using monotonically changing descriptors (MCDs), which are a set of descriptors whose values monotonically change as an atom or fragment is appended to a chem­ical structure [ MCDs. The proposed generation algorithm grows chemical structures by attaching atoms or fragments one by one. When an MCD value of a growing chemical struc­ture exceeds the upper bound of a constraint, the structure is pruned on a generation tree. A generation tree is a tree whose nodes are generated chemical structures and whose edges are connections between parent and child nodes. The chemical struc­ture at a child node is generated by appending a fragment to the chemical structure at its parent node. The concept of structure generation using MCDs is illustrated in
3.3. Details of the structure generation algorithm using a constraint defined with
Fig. MCDs have been given in the literature [
20]. Most molecular size-dependent descriptors are categorized as
20, 21].
3.2.3 Proof-Of-Concept Study: Bioactive Chemical Structure
Design Through the Proposed B-IQW
We have previously reported several proof-of-concept studies on the B-IQW targeting the exhaustive generation of chemical structures exhibiting a desired property or activity. For example, boiling points [ B-IQW applications. Furthermore, potential ligands for human alpha 2A adrenergic receptors were generated by adopting the B-IQW [ design is briefly presented as a proof-of-concept study of the inverse QSAR problem.
3.2.3.1 Thrombin Inhibitor Data Set and Descriptors
Designing thrombin inhibitors as an inverse QSAR problem is challenging owing to the relatively large molecular size of the inhibitors (vide infra). The conventional approach to exhaustive structure generation by atom combination cannot handle large chemical structures owing to the combinatorial explosion problem. Our structure
] and aqueous solubility [6] were targeted as
8
6, 21]. Herein, thrombin inhibitor
3 Data-Driven Molecular Structure Generation for Inverse QSPR/QSAR … 53
https://t.me/med1917
Fig. 3.3 Generation of chemical structures using monotonically changing descriptors (MCDs). Transition in the chemical space spanned by the molecular weight (MW) and the number of aromatic rings (A-rings) from structure 1 to 3 is depicted on the left. The right shows the growing structures with descriptor values
generation strategy of combining fragments and adding atoms is a solution to this problem.
From the ChEMBL database ver. 20 [22], 3259 small molecules for thrombin (CHEMBLID 204) annotated with K
were downloaded. These small molecules
i
were further curated with an in-house script, and the remaining 1705 compounds were used in the analysis. The K
values were transformed to pKi values as the
i
objective variable (y). The averages of the molecular weight and the number of heavy atoms for the data set were 475.4 and 33.9, respectively. The range of y was (1.00, 12.2), i.e., there was a great range of potencies for this target.
Fifty-one MCDs were adopted. Variable selection based on descriptor informa­tion provided 27 MCDs. These 27 MCDs included substructure descriptors (e.g., the number of aromatic rings, the number of CH
R substructures (R: carbon atom), and
3
the number of aromatic ketones), topological descriptors (the first Zagrev index), the topological polar surface area (TPSA), and topological pharmacophore distance descriptors. A topological pharmacophore distance descriptor is the summation of topological distances between specified types of potential pharmacophore point
20
(PPP) [
]. Hydrogen bond acceptors and donors, lipophilic atoms (e.g., carbon atoms adjacent only to carbon atoms), positively and negatively ionizable atoms, and aromatic rings were adopted as PPPs. The summation of the topological distances between a pair of PPPs is expected to represent the pharmacophore information of
].
inhibitors like chemically advanced template search descriptors [
23
54 T. Miyao a nd K. Funatsu
https://t.me/med1917
3.2.3.2 Model Construction and Derivation of P(X|y)
All 1705 compounds represented by the 27 MCDs were used to train GMM and cMLR models. Eight components in a GMM had the minimum Bayesian informa­tion criterion value among different numbers of components. The GMM model was created using the mclust package [
2
value of 0.656. The density of p(x) and the densities of p(x|y) for several y
an R values were projected on the generative topographic mapping (GTM) [ tion to the centers of eight Gaussians (Fig.
24] in the programming language R. cMLR reached
25] in addi-
3.4). The GTM created a manifold in the
27-dimensional MCD space so that the nonlinear manifold had an optimal likelihood for the training compounds. Figure
3.4 clearly shows that the posterior distributions
were different at the centers of the Gaussians and in highly dense regions on the manifold.
To generate highly potent compounds within the AD of the cMLR model, the
target pK
value was set at 11 and the centers of the eight Gaussians were focused on.
i
The centers of the Gaussians were expected to show high density, such that structure generation aiming at the centers would produce reasonable chemical structures within
Fig. 3.4 p(x)and p(x|y) distributions on the generative topographic mapping (GTM) manifold. a The prior density p(x) is shown in grayscale on the GTM manifold, and t he eight centers of
Gaussians are also shown on the manifold. b, c, d The posteriors of p(x|y)given y = 3, 7, and 11 are shown
3 Data-Driven Molecular Structure Generation for Inverse QSPR/QSAR … 55
https://t.me/med1917
Fig. 3.5 Predicted pK against the density of the center of each Gaussian. Dots are centers of Gaussians and the numbers identify the Gaussians, as shown in Fig. is the logarithm of the density of p(x|y = 11)
i
3.4.The x-axis
the AD. According to the relation between the predicted pKi and the density at the centers (Fig.
3.5), the center of Gaussian 1 was selected as a target, and the
coordinates of the center were translated to constraints. The density and predicted
value for the center of Gaussian 1 were 2.84 × 103 and 9.61, respectively. It is
pK
i
noted that, although aiming for a pK
value of 11, the predicted pKi values at the
i
centers of the Gaussians were approximately 9.5 and lower. This was because the posterior distribution inherits the prior distribution via the Bayes theorem. As the
distribution for the 1705 compounds was unimodal and the average was 6.6, the
pK
i
posterior was biased toward the prior to a certain extent.
3.2.3.3 Structure Generation Based on P(X|y)
From the coordinates of the center of Gaussian 1 (Fig. 3.5), target coordinates or a grid point were determined by adjusting the coordinates for discrete variables; e.g., the number of aromatic rings. MCD constraints were then determined heuristically so that the constraint hyperrectangle covered the grid point. For example, the coordinates of the grid point were the number of rings: 2, the number of aromatic rings: 1, TPSA: 112, summation of topological distances between lipophilic points (LL): 125, and so on. We set constraints (lower–upper bounds) at 2–3 for the number of rings, 1–1 for the number of aromatic rings, 104–120 for the TPSA, 92–158 for LL, and so on.
Chemical structures were generated by combining atom and ring fragments as long as they satisfied the descriptor constraints. A total of 289 ring fragments were prepared by dissecting the 1705 compounds into ring systems using the Bemis– Murcko scaffold definition [ were used as atom fragments. The generation of potentially reactive substructures like
26], and C, N, O, F, Cl, Br, and I with a single valence rule
56 T. Miyao a nd K. Funatsu
https://t.me/med1917
Fig. 3.6 Generated chemical structures for thrombin inhibitors. The blue circles are the coordinates of the generated chemical structures, and the red square is the target grid point for the generation of structures on the space spanned by the predicted pK i and the density of p(x|y = 11). Three chemical structures A, B, and C are shown by their structural formulas, and the corresponding coordinates are shown on the scatter plot
peroxides was suppressed. Furthermore, we set further constraints to avoid combi­natorial explosion, e.g., we constrained the number of fragments to be combined from 2 to 15 and introduced a probabilistic operation to randomly prune a growing generation tree with a probability of 0.4. As random pruning was introduced, we conducted three trials with different random seeds.
The total number of generated chemical structures satisfying the constraints was
11
1739 while 4.25 × 10
structures were built on generation trees. The expected
number of eligible structures without introducing probability generation would be
1.41 × 10
8
and this would require the searching of 6.19 × 10
15
structures. The latter number is out of our reach and highlights the necessity of avoiding a combinato­rial explosion. The densities of p(x|y = 11) and predicted pK generated chemical structures are shown Fig.
3.6-left. On the scatter plot, the red
values for the 1739
i
square is the target grid point, and the chemical structures residing nearby on the Pareto frontier are reported in Fig.
3.6-right. These chemical structures show desired
features based on the QSAR and AD models, and they might be useful for further development of the chemical structures.
3.3 Related Studies and Recent Trends in Research
on the Inverse QSPR/QSAR Problem
Various approaches have been developed since our proposal of adopting the Bayesian approach for the inverse QSPR problem.
3 Data-Driven Molecular Structure Generation for Inverse QSPR/QSAR … 57
https://t.me/med1917
Ikebata et al. introduced the use of the simplified molecular-input line-entry system (SMILES)-based chemical language model to generate novel chemical structures by mutating sentences of chemical structures [ in the formalization of the posterior distribution of the chemical structure given a desirable probability on y. In their approach, SMILES strings of chemical structures are mutated using a mutation function, and the mutated strings are then evaluated via QSAR models to obtain an updated probability of the posterior, like in F-IQWs and Markov chain Monte Carlo sampling. Sampling based on the updated probabilities of strings leads to a set of SMILES strings for the next generation. Here, the mutation function plays a central role in the sampling of chemical structures based on the posterior distribution. Unfortunately, in their validation scheme, the diversity of the generated chemical structures in terms of the true compound distribution was not described, and the performance of the approach was not compared with that of a genetic algorithm.
The methodological development and potential of GMR in solving the inverse QSPR problem were intensively investigated by Kaneko et al. for hyperparameter optimization in GMR [ the posterior [ materials [
A current trend of structure generation for the inverse QSPR problem is to adopt generative modeling [ of generative models on a massive number of chemical structures is efficient for the s ampling of sample chemical structures following the distribution of training chemical structures. Direct sampling from the distribution given a specific y value using conditional generative models has been proposed [ ative models are effective for reproducing the probability distribution, the B-IQW explained in this chapter does not lose its value owing to its transparent structure generation process focusing on a specific region of the chemical space. Moreover, leveraging advanced computational power, exhaustive chemical structure genera­tion is important to molecular design projects in terms of not missing “interesting” chemical structures as hypothesis compounds drawn from data.
29]. The method was further developed for the design of thermochromic
30] and the derivation of synthesis conditions in artificial bone design [ 31].
28] and optimal coordinate detection at a maximal density of
32
], mostly deep-learning-based modeling. The training
27]. The novelty of this approach lies
33]. Although these gener-
3.4 Conclusions
In this chapter, we discussed approaches for the inverse QSPR and QSAR prob­lems for molecular design from a methodological point of view. F-IQWs adopt QSPR models for filtering generated chemical structures like VS, whereas B-IQWs construct chemical structures from building blocks (fragments) satisfying descriptor constraints at a generation level.
One way to tackle backward approach is to derive the posterior distribution of descriptors: x given a specific property value and y for generating chemical structures within the AD of a model. For that purpose, GMM and cMLR were introduced, and via the Bayes theorem, the posterior distribution was analytically derived. As a
58 T. Miyao a nd K. Funatsu
https://t.me/med1917
demonstration of the proposed approach, chemical structures of thrombin inhibitors were designed. In the demonstration, novel chemical structures with desirable and trustable property values based on a QSPR model were generated. We expect that this intuitive approach of generating chemical structures will assist experimental chemists and biologists in enumerating virtual compounds as a data hypothesis.
References
1. Kang P, Liu Z, Abou-Rachid H, Guo H (2020) Machine-Learning Assisted Screening of Energetic Materials. J Phys Chem A 124:5341–5351.
2647
2. Ding Y, Chen S, Liu H, et al (2023) Discovery of Multitarget Inhibitors against Insect Chitinolytic Enzymes via Machine Learning-Based Virtual Screening. J Agric Food Chem 71:8769–8777.
3. Wong F, Zheng EJ, Valeri JA, et al (2023) Discovery of a Structural Class of Antibiotics with Explainable Deep Learning. Nature 626(7997):177–185.
06887-8
4. Brown N, McKay B, Gasteiger J (2006) A Novel Workflow for the Inverse QSPR Problem Using Multiobjective Optimization. J Comput Aided Mol Des 20:333–341.
1007/S10822-006-9063-1
5. Jensen JH (2019) A Graph-Based Genetic Algorithm and Generative Model/Monte Carlo Tree Search for the Exploration of Chemical Space. Chem Sci 10:3567–3572.
1039/C8SC05372C
6. Miyao T, Kaneko H, Funatsu K (2016) Inverse QSPR/QSAR Analysis for Chemical Structure Generation (from y to x). J Chem Inf Model 56:286–299.
5B00628
7. Gantzer P, Creton B, Nieto-Draghi C (2020) Inverse-QSPR for de novo Design: A Review. Mol Inform 39:1900087.
8. Miyao T, Arakawa M, Funatsu K (2010) Exhaustive Structure Generation for Inverse-QSPR/ QSAR. Mol Inform 29:111–125.
9. Wong WW, Burkowski FJ (2009) A Constructive Approach for Discovering New Drug Leads: Using a Kernel Methodology for the Inverse-QSAR Problem. J Cheminform 1:1–27.
doi.org/10.1186/1758-2946-1-4
10. Churchwell CJ, Rintoul MD, Martin S, et al (2004) The Signature Molecular Descriptor: 3. Inverse-Quantitative Structure–Activity Relationship of ICAM-1 Inhibitory Peptides. J Mol Graph Model 22:263–273.
11. Skvortsova MI, Baskin II, Slovokhotova OL, et al (1993) Inverse Problem in QSAR/QSPR Studies for the Case of Topological Indices Characterizing Molecular Shape (Kier Indices). J Chem Inf Comput Sci 33:630–634.
12. Faulon JL, Churchwell CJ, Visco DP (2003) The Signature Molecular Descriptor. 2. Enumer­ating Molecules from Their Extended Valence Sequences. J Chem Inf Comput Sci 43:721–734.
https://doi.org/10.1021/CI020346O
13. Visco DP, Pophale RS, Rintoul MD, Faulon JL (2002) Developing a Methodology for an Inverse Quantitative Structure–Activity Relationship Using the Signature Molecular Descriptor. J Mol Graph Model 20:429–438.
14. Dragos H, Gilles M, Alexandre V (2009) Predicting the Predictability: A Unified Approach to the Applicability Domain Problem of QSAR Models. J Chem Inf Model 49:1762–1776.
https://doi.org/10.1021/CI9000579
15. Klingspohn W, Mathea M, Ter Laak A, et al (2017) Efficiency of Different Measures for Defining the Applicability Domain of Classification Models. J Cheminform 9:1–17.
doi.org/10.1186/S13321-017-0230-2
https://doi.org/10.1021/ACS.JAFC.3C00633
https://doi.org/10.1002/MINF.201900087
https://doi.org/10.1002/MINF.200900038
https://doi.org/10.1016/J.JMGM.2003.10.002
https://doi.org/10.1021/CI00014A017
https://doi.org/10.1016/S1093-3263(01)00144-9
https://doi.org/10.1021/ACS.JPCA.0C0
https://doi.org/10.1038/s41586-023-
https://doi.org/10.
https://doi.org/10.
https://doi.org/10.1021/ACS.JCIM.
https://
https://
3 Data-Driven Molecular Structure Generation for Inverse QSPR/QSAR … 59
https://t.me/med1917
16. Gaspar HA, Marcou G, Horvath D, et al (2013) Generative Topographic Mapping-Based Clas­sification Models and Their Applicability Domain: Application to the Biopharmaceutics Drug Disposition Classification System (BDDCS). J Chem Inf Model 53:3318–3325.
org/10.1021/CI400423C
17. Berenger F, Yamanishi Y (2019) A Distance-Based Boolean Applicability Domain for Classi­fication of High Throughput Screening Data. J Chem Inf Model 59:463–476.
10.1021/ACS.JCIM.8B00499
18. Neath AA, Cavanaugh JE (2012) The Bayesian Information Criterion: Background, Derivation, and Applications. Wiley Interdiscip Rev Comput Stat 4:199–203.
S.199
19. Stulp F, Sigaud O (2015) Many Regression Algorithms, One Unified Model: A Review. Neural Netw 69:60–79.
20. Miyao T, Kaneko H, Funatsu K (2016) Ring System-Based Chemical Graph Generation for de novo Molecular Design. J Comput Aided Mol Des 30:425–446.
822-016-9916-1
21. Miyao T, Kaneko H, Funatsu K (2014) Ring-System-Based Exhaustive Structure Generation for Inverse-QSPR/QSAR. Mol Inform 33:764–778.
22. Gaulton A, Hersey A, Nowotka ML, et al (2017) The ChEMBL Database in 2017. Nucleic Acids Res 45:D945–D954.
23. Reutlinger M, Koch CP, Reker D, et al (2013) Chemically Advanced Template Search (CATS) for Scaffold-Hopping and Prospective Target Prediction for ‘Orphan’ Molecules. Mol Inform 32:133–138.
24. Scrucca L, Fraley C, Murphy TB, Raftery AE (2023) Model-Based Clustering, Classification, and Density Estimation Using Mclust in R. Chapman & Hall/CRC Press.
25. Bishop CM, Svensén M, Williams CKI (1998) GTM: The Generative Topographic Mapping. Neural Comput 10:215–234.
26. Bemis GW, Murcko MA (1996) The Properties of Known Drugs. 1. Molecular Frameworks. J Med Chem 39:2887–2893.
27. Ikebata H, Hongo K, Isomura T, et al (2017) Bayesian Molecular Design with a Chemical Language Model. J Comput Aided Mol Des 31:379–391.
0008-Z
28. Kaneko H (2021) Extended Gaussian Mixture Regression for Forward and Inverse Analysis. Chemom Intell Lab Syst 213:104325.
29. Kaneko H (2022) True Gaussian Mixture Regression and Genetic Algorithm-Based Optimiza­tion with Constraints for Direct Inverse Analysis. Sci Technol Adv Mater Methods 2:14–22.
https://doi.org/10.1080/27660400.2021.2024101
30. Shimizu N, Kaneko H (2020) Direct Inverse Analysis Based on Gaussian Mixture Regression for Multiple Objective Variables in Material Design. Mater Des 196:109168.
10.1016/J.MATDES.2020.109168
31. Motojima K, Shiratsuchi R, Suzuki K, et al (2023) Machine Learning Model for Predicting the Material Properties and Bone Formation Rate and Direct Inverse Analysis of the Model for New Synthesis Conditions of Bioceramics. Ind Eng Chem Res 62:5898–5906.
10.1021/ACS.IECR.3C00332
32. Sousa T, Correia J, Pereira V, Rocha M (2021) Generative Deep Learning for Targeted Compound Design. J Chem Inf Model 61:5343–5361.
1496
33. Kang S, Cho K (2019) Conditional Molecular Design with Deep Generative Models. J Chem Inf Model 59:43–52.
https://doi.org/10.1016/J.NEUNET.2015.05.005
https://doi.org/10.1002/MINF.201400072
https://doi.org/10.1093/NAR/GKW1074
https://doi.org/10.1002/MINF.201200141
https://doi.org/10.1162/089976698300017953
https://doi.org/10.1021/JM9602928
https://doi.org/10.1016/J.CHEMOLAB.2021.104325
https://doi.org/10.1021/ACS.JCIM.0C0
https://doi.org/10.1021/ACS.JCIM.8B00263
https://doi.org/10.1002/WIC
https://doi.org/10.1007/S10
https://doi.org/10.1007/S10822-016-
https://doi.
https://doi.org/
https://doi.org/
https://doi.org/
Chapter 4
https://t.me/med1917
Materials Informatics with Limited Data
Ryo Yoshida
4.1 Introduction
In general, the parameter space for materials research, such as drug developments,
60
is vast. For example, there are approximately 10 ical space of small organic molecules [
1]. On the other hand, the number of small
molecules currently recorded in public databases is on the order of 10
2]. Therefore, a vast, unexplored area remains in the chemical space. Further-
[
candidate molecules in the chem-
8
at most
more, in the development of practical materials, the dimensionality of the param­eter space increases with the addition of other design variables such as processing conditions and the selection of additives, filters, and solvents. The task of materials informatics (MI) is to identify the unknown parameters that result in the desired material properties from such a vast parameter space. This is a multi-objective opti­mization problem. The design of materials, such as drug molecules, is inherently different from general industrial product design in terms of specificity and diversity of the parameter space. The parameters take a variety of forms depending on the problem of interest, including material composition, molecules, crystal structures, X-ray diffraction spectra, material microstructures, and processing conditions.
The basic workflow of MI consists of forward and inverse prediction tasks
(Fig.
4.1)[3–7
]. The forward problem aims to predict the output
Y for an arbitrary input X . For example, the input variable is given as a molecule, chemical compo­sition, or crystal structure, and the output is the physical properties or structural features of the resulting material. In conventional materials research, simulations based on physical laws, such as first-principles calculations and molecular dynamics (MD) simulations, have been used for forward prediction tasks. One of the main chal­lenges in MI is to replace computer-intensive calculations, which entail high costs,
R. Yoshida (B) The Institute of Statistical Mathematics, Research Organization of Information and Systems, 10-3 Midori-Cho, Tachikawa 190-8562, Tokyo, Japan e-mail: yoshidar@ism.ac.jp
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2024 H. Satoh et al. (eds.), Drug Development Supported by Informatics,
https://doi.org/10.1007/978-981-97-4828-0_4
61