Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5580_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
30.08.2026
Размер:
43 Мб
Скачать
40 H. Yamamoto
https://t.me/med1917
Table 2.11 Atomic number range of candidate compounds in Group 1
C# H# Br# Cl# F# I# N# O#
Min 2 1 0 0 0 0 0 0
Max 4 7 1 2 7 0 0 1
These are thought to have an anesthetic effect similar to that of diethyl ether, but little data were available to verify this. However, they have reasonable struc­tures similar to those of anesthetics used in practice (Desflurane: CF Sevoflurane: (CF
CHOCH2F, Isoflurane: CF3CHClOCHF2).
3)2
CHFOCHF2,
3
There is no candidate that can satisfy both the index of flammability, 0.75 or less, and the index of atmospheric lifetime, the logarithmic reaction rate constant with OH radicals (OHR) of 13.1 or more. In other words, the difficulty in designing compounds that are not flammable, have a short atmospheric lifetime, and are not likely to contribute to global warming, also became clear. Even the compounds in the original paper are not compatible with flammability and atmospheric lifetime; the range of atoms belonging to Gr. 1 is as shown i n Table
2.11.
Since 53 compounds are candidates for Group 2, the design guidelines expand a bit.
2.6.5 Use of the Usual NN Method
If a program for transfer learning such as the NthNN method is not available, a regular NN will be used shown in Fig.
2.28a.
For example, put the three-solubility data in series, and enter at the end of the table column which solubility index data it is 0,1. In the usual case, the columns log Kow and log S are appended. If the target variable is log Kow, enter 1 in the log Kow column. If the target variable is log S, enter 1 in the log S column. If the data
Fig. 2.28 Learning without Transfer Learning. a Multiple Regression model and NN model. b Correlation between calculated log Kow using NN and calculated log S using NN
2 Screening Methods for Drugs Using Chemoinformatics Methods … 41
https://t.me/med1917
are log BCF, put 0 in both the log Kow and log S columns. Naturally, a multiple regression calculation using this table will result in perfectly parallel straight lines for the three solubility predictions. The difference between the lines is the regression coefficient. This would seemingly yield three different indices, but it does not make sense for a screening study. Because it is not possible to search for compounds that have the same log Kow but are more soluble in water.
The NN method introduces interaction and nonlinearity between the input and middle layers. The actual computed values of log Kow and −log S are slightly different, as shown in Fig. better to use one of them and combine the other indices.
2.28b. However, if it is to be used for reverse design, it is
2.7 Reverse Design of Candidate Compounds
Halogen-containing compounds have many isomers because the hydrogen in the molecule is replaced by a halogen atom. The SMILES structural formulas for those isomers are created automatically. In doing so, I first consider the skeleton. In the case of two carbons, the following three skeletons are possible. C(X)(X)(X)C(X)(X), C(X)(X)(X)OC(X)(X)(X), C(X)(X) = C(X)(X).
A cyclic form of ether compound is also possible but excluded because of its flammability. The (X) is converted to ([H]), (F), (Cl), and (Br).
Symmetric compounds will generate multiple SMILES formulas showing the same structure, but no extra steps are taken. This is because it is easier and safer to generate all the structures and trim them later with physical properties than to write a program that performs extra trimming.
When there are three carbons, the following skeletons are possible.
CCC, C=CC, COCC, C1CC1, C1OCC1, COC=C, and C(C=O)C.
Olefin and ketone ring compounds are also possible, but they are not considered because of their higher flammability. Similarly, when the number of carbons is 4, saturated, olefinic, etheric, and ketone structures are considered. The number of olefins should be limited to one.
To obtain a table of functional groups from the SMILES structural formulas, I used an in-house segmentation program. From this table of functional groups, the calculated values were obtained using the physical property estimation formula created in this study. There were 2124 compounds in the log Kow index range, 2133 compounds in the log S index range, and 798 compounds in the log BCF index range. 258 compounds satisfied all three indices. However, for example, the SMILES formulas in Table three solubility indices are trimmed because they have the same structure. The 13 compounds in Table automatically generated from those with a carbon chain of 3, there were 151,552 compounds. For those with a solubility index of Gr. 1 and a carbon chain of 3, the 55 compounds shown in Table
I will skip the section on C4 compounds.
2.13 are candidates for 2 carbons. When the structures were
2.12 show the same molecule. Those with the same
2.14 were candidates.
42 H. Yamamoto
https://t.me/med1917
Table 2.12 Same SMILES structure generated automatically
Table 2.13 Reverse engineered 2-carbon candidate compounds
SMILES
[H]C([H])([H])C(Cl)(Br)[H]
[H]C([H])([H])C(Br)(Cl)[H]
[H]C([H])([H])C([H])(Br)Cl
[H]C([H])([H])C(Br)([H])Cl
[H]C([H])([H])C([H])(Cl)Br
[H]C([H])([H])C(Cl)([H])Br
[H]C(Cl)(Br)C([H])([H])[H]
[H]C(Br)(Cl)C([H])([H])[H]
ClC([H])(Br)C([H])([H])[H]
ClC(Br)([H])C([H])([H])[H]
BrC([H])(Cl)C([H])([H])[H]
BrC(Cl)([H])C([H])([H])[H]
No SMILES
11 [H]C([H])([H])C(Cl)(Br)[H]
4762 [H]C(F)(Cl)OC(Cl)(Cl)F
7717 BrC(F)([H])OC(F)(F)Cl
3 [H]C([H])([H])C([H])(Br)[H]
5014 [H]C(Cl)(Cl)OC(F)(Cl)F
10 [H]C([H])([H])C(Cl)(Cl)[H]
4715 [H]C(F)(F)OC(Cl)(Br)Cl
4442 [H]C([H])(F)OC(Cl)(Cl)F
4694 [H]C(F)(F)OC(F)(Cl)F
4587 [H]C([H])(Br)OC(Cl)(Br)Cl
5039 [H]C(Cl)(Cl)OC(Br)(Br)Cl
5071 [H]C(Cl)(Br)OC(Br)(Br)[H]
4698 [H]C(F)(F)OC(Cl)(Cl)F
If necessary, further refinement can be based on atomic number range, boiling point, flammability, and atmospheric lifetime. As described above, screening studies do not necessarily require big data on the drug itself. Even if there is little data on the MAC itself, it is possible to propose new candidates in an ordinal order.
2 Screening Methods for Drugs Using Chemoinformatics Methods … 43
https://t.me/med1917
Table 2.14 Reverse designed 3-carbon candidate compounds
SMILES SMILES
[H]C(F)(F)OC([H])([H])C(F)(F)F [H]C([H])(Br)OC(F)(F)C(F)(F)[H]
[H]C(F)(Br)OC([H])(F)C(F)(F)F [H]C(F)(F)OC([H])(Cl)C(F)(Br)F
FC(F)=C(Br)OC(Br)(Br)[H] C1([H])([H])C(F)(F)C1(F)(Cl)
[H]C(F)(Br)OC([H])(Cl)C(F)(Cl)F C1([H])([H])C([H])(F)C1(F)(Cl)
C1([H])([H])C([H])(F)C1([H])(Br) [H]C([H])(F)OC([H])(F)C(F)(Br)[H]
C1([H])([H])C([H])([H])C1([H])(F) C1([H])(F)C([H])(Br)C1(F)(F)
C1([H])([H])C([H])([H])C1(F)(F) [H]C(F)(F)OC([H])(F)C(F)(Br)F
FC(Cl)=C(Br)OC(Cl)(Br)[H] C1([H])(F)C([H])(F)C1([H])(Br)
[H]C([H])(Cl)OC(F)(Br)C(F)(F)[H] [H]C([H])(Cl)OC([H])(Br)C(Cl)(Cl)[H]
[H]C(F)(F)OC(F)(F)C(Br)(Br)F [H]C(F)(F)OC([H])(Br)C(F)(Br)F
[H]C(Cl)=C([H])OC(F)(Cl)[H] [H]C([H])(Cl)OC([H])(F)C(Cl)(Cl)F
ClC(Cl)=C(Br)OC(F)(F)[H] [H]C([H])([H])OC([H])(F)C(F)(F)F
[H]C(F)(F)OC([H])(Cl)C(F)(F)F [H]C([H])=C(Cl)OC(Br)(Br)[H]
ClC(Cl)=C([H])OC(Cl)(Br)[H] [H]C(F)(F)OC(F)(F)C(Cl)(Cl)[H]
[H]C(F)(Br)OC([H])(F)C(Cl)(Br)F [H]C(Cl)=C([H])OC([H])(F)[H]
[H]C([H])(Cl)OC([H])(Br)C(F)(F)[H] [H]C(Br)=C([H])OC(F)(F)[H]
FC(F)=C(Cl)OC(Br)(Br)[H] C1([H])([H])C([H])(F)C1([H])(F)
[H]C(Cl)(Cl)OC([H])(F)C(F)(Cl)F [H]C([H])(Br)OC(F)(F)C(Cl)(Cl)[H]
FC(F)=C([H])OC(F)(Cl)F FC(Cl)=C(Cl)OC([H])(Br)[H]
[H]C([H])([H])OC([H])([H])C(F)(Br)F C1([H])([H])C([H])(F)C1(F)(F)
C1([H])([H])C(F)(Cl)C1(F)(Cl) [H]C(F)(Br)OC([H])(F)C(Cl)(Br)Cl
[H]C([H])(Br)OC([H])([H])C(F)(Br)F [H]C([H])(Br)OC([H])(F)C(F)(Cl)F
[H]C(F)(F)OC([H])(Br)C(F)(F)F [H]C([H])([H])OC([H])(F)C(F)(Br)F
ClC(Cl)=C([H])OC([H])(Br)[H] [H]C([H])([H])OC(Cl)(Cl)C(F)(F)[H]
[H]C([H])([H])OC(F)(F)C(Cl)(Cl)[H] [H]C([H])=C(Br)OC(Br)(Br)[H]
C1([H])(F)C([H])(Br)C1(F)(Cl) [H]C(F)(F)OC([H])(F)C(F)(Cl)F
FC(Cl)=C(Cl)OC(F)(F)[H] C1([H])(Cl)C([H])(Br)C1(F)(Cl)
[H]C([H])([H])OC([H])(Br)C(Br)(Br)Cl
2.8 Useful Tips for Machine Learning
2.8.1 LASSO % Error Analysis
The LC50 index has been tested in Rat and Mouse. As explained earlier, LASSO regression is used to minimize the sum of errors between experimental and calcu­lated values, with the sum of absolute values of regression coefficients also taken
44 H. Yamamoto
https://t.me/med1917
Fig. 2.29 LASSO % error analysis. a LASSO regression results for LC50. b LASSO regression with % error. c Normal error vs. % error
into account as a penalty (Fig.
2.29a). In many cases, convergence calculations are
performed using a Microsoft Excel solver.
So, it is easy to apply the % error as the error. The % error is larger the smaller the experimental value (the higher the toxicity). When the LC50 is 10, a shift of 1 is 10%, while a shift of 1 when the LC50 is 100 is 1%. Therefore, if I introduce the % error into the LASSO regression, I can search for regression coefficients so that the fitting is higher where the LC50 is small. Conversely, the fitting will be poor where the LC50 is large (Fig.
2.29b).
This effect may not seem large. I have calculated the predictions for all compounds in the database. Let’s zoom in on the particularly problematic areas of toxicity, from
3to3(Fig.
2.29c). The compounds that were rated as less toxic in the regular
LASSO regression and more toxic in the % error LASSO regression appear on the left-lower side of the Y = X line.
In compound screening studies, the goal is not to obtain high correlation coeffi­cients. As a way to exclude those with bad (highly toxic) ratings, or at the very least, to give an estimate that calls for caution when handling, LASSO % error regression is a simple and highly effective method. Compounds reverse-designed by a computer often carry no information, so it is a good idea to do the calculations just to be sure. This method is not limited to regions close to the origin but is also a way to increase the accuracy of one’s target region. For example, in the case of yields in synthesis, this method is used when you want to increase the accuracy of a high region.
2 Screening Methods for Drugs Using Chemoinformatics Methods … 45
https://t.me/med1917
2.8.2 LASSO Regression with Upper and Lower Limits
Solubility in water was often expressed as gram of solubility in 100g water. Therefore, highly water-soluble compounds are often not measured above 100g/100g water. When conducting an analysis that includes such data, it is necessary to devise a method for capturing errors when the calculated value of log S is greater than 2. For example, if the regression coefficient for hydroxyl groups is large, some compounds with multiple hydroxyl groups will have calculated values that greatly exceed 2. If the calculated value of 2 or more is set as the upper limit, the calculation error will always be zero no matter how large the regression coefficient of the hydroxyl group is. In such cases, it is effective to use LASSO regression and incorporate the sum of the absolute values of the regression coefficients as a penalty in the evaluation.
References
1. Koblin DD, Laster MJ, Ionescu P, Gong D, Eger E I II, Halsey MJ, Hudlicky T (1999)
Polyhalogenated Methyl Ethyl Ethers: Solubilities and Anesthetic Properties. Anesth Analg
88(5):1161–1167.
2. RDKit: Open-Source Cheminformatics Software. https://www.rdkit.org. Accessed 15 Feb 2024.
3. Weininger, D (1988) Smiles, a chemical language and information system. 1. introduction to
method- ology and encoding rules. J Chem Inf Comp Sci 28:31–36.
0057a005.
4. Open MOPAC. http://openmopac.net. Accessed 15 Feb 2024.
5. Pople JA, Beveridge DL (1970) Approximate Molecular Orbital Theory. McGraw-Hill, New
Yor k .
6. Aoyama T, Ichikawa H (1991) Reconstruction of weight matrices in neural networks : A method
of correlating outputs with inputs. Chem Pharm Bull 39:1222–1228.
39.1222.
7. Kier LB, Hall LH (2000) Intermolecular Accessibility: The Meaning of Molecular Connectivity.
J Chem Inf Comput Sci 40(3):792–795.
8. Microsoft Corporation (2021) Microsoft Excel. https://office.microsoft.com/excel. Accessed 16
Feb 2024.
https://doi.org/10.1097/00000539-199905000-00036.
https://doi.org/10.1021/ci0
https://doi.org/10.1248/cpb.
https://doi.org/10.1021/ci990135s.
Chapter 3
https://t.me/med1917
Data-Driven Molecular Structure Generation for Inverse QSPR/QSAR Problem
Tomoyuki Miyao and Kimito Funatsu
3.1 Introduction
A fundamental goal of chemistry is to find new substances that have desired prop­erties. To that end, data-driven chemistry utilizes quantitative structure–property and structure–activity relationship (QSPR and QSAR) models to estimate property (activity) values for the chemical structures of compounds. A QSPR/QSAR model takes a molecular descriptor as the input and outputs property values through an algo­rithm, as seen in many QSPR/QSAR studies. As there is no conceptual difference between the QSPR and QSAR at a methodological level, in this chapter, we refer to the QSPR when statements are applicable to both the QSPR and QSAR.
For molecular design, QSPR models are frequently used to screen a large compound library to make a focused library. The focused library contains a smaller number of compounds that have desired properties according to the model output. This naïve approach is called virtual screening (VS), and many successful applica-
1
tions of VS have been reported for various targets from energetic materials [
] and antibiotics [3]. One obvious limitation of VS is that it cannot
insecticides [ propose novel compounds that are not included in a screening library. Thus, although VS is an effective method, new ideas are still needed to design novel chemical structures of compounds, and it has been a central issue in chemoinformatics and data-driven chemistry.
The inverse QSPR problem can be defined as a problem of proposing chemical structures that exhibit a desirable (predicted) property based on a QSPR model; i.e.,
2
]to
T. Miyao Data Science Center, and Graduate School of Science and Technology, Nara Institute of Science and Technology, 8916-5 Takayama-Cho, Ikoma 630-0192, Nara, Japan
miyao@dsc.naist.jp
e-mail:
K. Funatsu (B) Data Science Center, Nara Institute of Science and Technology, Ikoma, Nara, Japan e-mail: funatsu@dsc.naist.jp
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2024 H. Satoh et al. (eds.), Drug Development Supported by Informatics,
https://doi.org/10.1007/978-981-97-4828-0_3
47
48 T. Miyao a nd K. Funatsu
https://t.me/med1917
Fig. 3.1 Two approaches to the inverse QSPR problem. a B-IQW generates chemical structures from descriptors selected to show a specific property value. b F-IQW applies QSPR model to predict a property value from a chemical structure for structure modification
inversing the flow of QSPR analysis to propose chemical structures from a property value. Workflows to solve the inverse QSPR problem (inverse QSPR workflows (IQWs)) can be classified into two categories according to the direction of the use of QSPR models: the forward IQWs (F-IQWs, Fig.
3.1a).
Fig.
3.1b) and backward IQWs (B-IQWs,
F-IQWs start with random modification of chemical structures (Fig. 3.1b). These chemical structures are then filtered on the basis of QSPR models and modified for the next iteration of structure evaluation [
4, 5]. This process is repeated until chemical
structures with the desired properties are found or convergence is reached. F-IQWs have been extensively investigated most likely for their simplicity. In contrast, B­IQWs or truly inverse QSPR workflows (Fig.
3.1a) have seldom been reported [6].
In these approaches, a specific value of the objective variable (property value) is set, and the inverse analysis of the QSPR model then determines the descriptor information that leads to the specific property value. Finally, a structure generator generates chemical structures per the descriptor information. Although F-IQWs are simpler and easier to implement as computational systems, B-IQWs have a merit of efficiently generating chemical structures satisfying a specific property value, because the descriptor values specified for structure generation are assured to produce the desired property values by the model. Methods for the inverse QSPR problem and its two-way approaches were previously reviewed [
7].
3.2 B-IQW via a Bayesian Framework
In 2010, Funatsu and Miyao proposed a new B-IQW via a Bayesian framework [8 which generates chemical structures that have desired properties. Before going into the methodological details, we discuss the limitations of conventional B-IQWs in the following section.
],
3 Data-Driven Molecular Structure Generation for Inverse QSPR/QSAR … 49
https://t.me/med1917
3.2.1 Conventional B-IQWs
Conventionally, the methods for the inverse QSPR problem focused on the mathemat­ical retrieval of chemical graphs from a set of descriptor values [ the group of Faulon JL proposed an algorithm for generating molecular structures from a set of values of their developed topological descriptor, i.e., signature descrip-
12]. Targeting HIV-1 protease inhibitors, they successfully proposed a focused
tors [ library of chemical structures having desired activity values [ was first to identify a set of signature values corresponding to a desired activity value predicted using a multivariate linear regression (MLR) model, and then to exhaus­tively generate chemical structures from the signature set with a structure generation algorithm. This mathematically well-defined approach has two limitations in prac­tical QSPR analysis, namely that only the signature descriptor is available and that the applicability domain (AD) of the QSPR model has not been incorporated [
The molecular descriptor or representation plays a central role in QSPR modeling. Various descriptors have been proposed to numericalize chemical structures in a way that is relevant to a target property. It is a common understanding that there is no single descriptor effective in every property prediction, and appropriate descriptors depend on target properties. Thus, the descriptor available in B-IQWs should not be limited to a specific type.
The AD of a QSPR model refers to regions in the chemical space (descriptor space) where the output of the QSPR model is reliable. As all QSPR models are regression models, a valid prediction is assumed only for compounds within the AD. Many approaches have been proposed to define the AD of a QSPR model [ and most approaches share the basic concept that the high-dense region of the training compound distribution should be the AD. In IQW, chemical structures are generated using a QSPR model, and the AD must therefore be incorporated into the analysis workflow irrespective of the forward or backward nature of the analysis.
911]. For example,
13]. Their approach
14].
1417],
3.2.2 Bayesian Approaches for B-IQW
Our B-IQW has two steps like standard methods for the problem, namely the deter­mination of a region of the chemical space where corresponding property values are desirable and the exhaustive generation of structures for the determined region [ In this section, x represents the descriptor (explanatory variable) and y represents the property (objective variable) for simplicity.
8].
50 T. Miyao a nd K. Funatsu
https://t.me/med1917
3.2.2.1 Posterior Distribution of X Given a Specific Y Value: P(X|y)
In general, identifying x corresponding to a y value for a QSPR is not tractable because the QSPR model is a non-injective function; i.e., two descriptors may corre­spond to the same property value. Furthermore, the AD must be incorporated in the inverse analysis. We propose using probability distributions in combination with the Bayesian framework to achieve these two goals simultaneously.
The probability density function of x: p(x) represents the distribution of training data (the prior distribution). In the proposed method, p(x) is expressed as a Gaussian mixture model (GMM) with a weight p(z component of the M Gaussians (eq.
takes a value of 1 when x belongs to the k-th component and a value of zero
z
k
), mean µk, and covariance k for the k-th
k
3.1). The indicator variable (random variable)
otherwise.
M
p(x) =
p(zk )N (x|µk ,∑k ) (3.1)
k=1
As the compound distribution in the descriptor space is usually non-Gaussian and multi-modal, a mixture of Gaussians is expected to well capture the training compound distribution, forming a simple AD. The number of Gaussians, M, and the shape of covariance, criterion [
18]. The parameters in p(x) can be determined using the expectation–maxi-
, can be determined based on the Bayesian information
k
mization algorithm to achieve a maximum of the likelihood function for a training data set.
To capture the nonlinear relationship between x and y while maintaining an analyt­ically tractable form for the posterior density derivation, we have proposed the adop­tion of three regression methods: MLR, component-wise MLR (cMLR), and Gaus­sian mixture regression (GMR) [
6]. cMLR places an MLR model per GMM compo-
nent, allowing pseudo-nonlinear regression models to be constructed. An MLR model is formulated for each component of the Gaussians in Eq.
p(y|x, z
= Ny|a
)
k
T
x + bk ,σ
k
3.2.
2
k
(3.2)
The posterior distribution of x given a specific y value is expressed as Eq. 3.3.
Using the Bayes theorem, the posterior of x is obtained as Eq. 3.4.
p(x|y) =
p(x|y) =
M
p(x|y, zk )
k=1
M
p(x, zk|y) (3.3)
k=1
)p(zk )
p(y|z
k
M
p(y|zl)p(zl)
l=1
(3.4)