Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5580_Библиотеки_им_академика_М_И_Перельмана
.pdf
40 H. Yamamoto
https://t.me/med1917
Table 2.11 Atomic number range of candidate compounds in Group 1
C# H# Br# Cl# F# I# N# O#
Min 2 1 0 0 0 0 0 0
Max 4 7 1 2 7 0 0 1
These are thought to have an anesthetic effect similar to that of diethyl ether,
but little data were available to verify this. However, they have reasonable structures similar to those of anesthetics used in practice (Desflurane: CF
Sevoflurane: (CF
CHOCH2F, Isoflurane: CF3CHClOCHF2).
3)2
CHFOCHF2,
3
There is no candidate that can satisfy both the index of flammability, 0.75 or less,
and the index of atmospheric lifetime, the logarithmic reaction rate constant with
OH radicals (OHR) of −13.1 or more. In other words, the difficulty in designing
compounds that are not flammable, have a short atmospheric lifetime, and are not
likely to contribute to global warming, also became clear. Even the compounds in
the original paper are not compatible with flammability and atmospheric lifetime;
the range of atoms belonging to Gr. 1 is as shown i n Table
2.11.
Since 53 compounds are candidates for Group 2, the design guidelines expand a
bit.
2.6.5 Use of the Usual NN Method
If a program for transfer learning such as the NthNN method is not available, a
regular NN will be used shown in Fig.
2.28a.
For example, put the three-solubility data in series, and enter at the end of the
table column which solubility index data it is 0,1. In the usual case, the columns log
Kow and −log S are appended. If the target variable is log Kow, enter 1 in the log
Kow column. If the target variable is −log S, enter 1 in the log S column. If the data
Fig. 2.28 Learning without Transfer Learning. a Multiple Regression model and NN model.
b Correlation between calculated log Kow using NN and calculated −log S using NN

2 Screening Methods for Drugs Using Chemoinformatics Methods … 41
https://t.me/med1917
are log BCF, put 0 in both the log Kow and log S columns. Naturally, a multiple
regression calculation using this table will result in perfectly parallel straight lines
for the three solubility predictions. The difference between the lines is the regression
coefficient. This would seemingly yield three different indices, but it does not make
sense for a screening study. Because it is not possible to search for compounds that
have the same log Kow but are more soluble in water.
The NN method introduces interaction and nonlinearity between the input and
middle layers. The actual computed values of log Kow and −log S are slightly
different, as shown in Fig.
better to use one of them and combine the other indices.
2.28b. However, if it is to be used for reverse design, it is
2.7 Reverse Design of Candidate Compounds
Halogen-containing compounds have many isomers because the hydrogen in the
molecule is replaced by a halogen atom. The SMILES structural formulas for those
isomers are created automatically. In doing so, I first consider the skeleton. In the
case of two carbons, the following three skeletons are possible. C(X)(X)(X)C(X)(X),
C(X)(X)(X)OC(X)(X)(X), C(X)(X) = C(X)(X).
A cyclic form of ether compound is also possible but excluded because of its
flammability. The (X) is converted to ([H]), (F), (Cl), and (Br).
Symmetric compounds will generate multiple SMILES formulas showing the
same structure, but no extra steps are taken. This is because it is easier and safer to
generate all the structures and trim them later with physical properties than to write
a program that performs extra trimming.
When there are three carbons, the following skeletons are possible.
CCC, C=CC, COCC, C1CC1, C1OCC1, COC=C, and C(C=O)C.
Olefin and ketone ring compounds are also possible, but they are not considered
because of their higher flammability. Similarly, when the number of carbons is 4,
saturated, olefinic, etheric, and ketone structures are considered. The number of
olefins should be limited to one.
To obtain a table of functional groups from the SMILES structural formulas,
I used an in-house segmentation program. From this table of functional groups,
the calculated values were obtained using the physical property estimation formula
created in this study. There were 2124 compounds in the log Kow index range,
2133 compounds in the −log S index range, and 798 compounds in the log BCF
index range. 258 compounds satisfied all three indices. However, for example, the
SMILES formulas in Table
three solubility indices are trimmed because they have the same structure. The 13
compounds in Table
automatically generated from those with a carbon chain of 3, there were 151,552
compounds. For those with a solubility index of Gr. 1 and a carbon chain of 3, the
55 compounds shown in Table
I will skip the section on C4 compounds.
2.13 are candidates for 2 carbons. When the structures were
2.12 show the same molecule. Those with the same
2.14 were candidates.

42 H. Yamamoto
https://t.me/med1917
Table 2.12 Same SMILES
structure generated
automatically
Table 2.13 Reverse
engineered 2-carbon
candidate compounds
SMILES
[H]C([H])([H])C(Cl)(Br)[H]
[H]C([H])([H])C(Br)(Cl)[H]
[H]C([H])([H])C([H])(Br)Cl
[H]C([H])([H])C(Br)([H])Cl
[H]C([H])([H])C([H])(Cl)Br
[H]C([H])([H])C(Cl)([H])Br
[H]C(Cl)(Br)C([H])([H])[H]
[H]C(Br)(Cl)C([H])([H])[H]
ClC([H])(Br)C([H])([H])[H]
ClC(Br)([H])C([H])([H])[H]
BrC([H])(Cl)C([H])([H])[H]
BrC(Cl)([H])C([H])([H])[H]
No SMILES
11 [H]C([H])([H])C(Cl)(Br)[H]
4762 [H]C(F)(Cl)OC(Cl)(Cl)F
7717 BrC(F)([H])OC(F)(F)Cl
3 [H]C([H])([H])C([H])(Br)[H]
5014 [H]C(Cl)(Cl)OC(F)(Cl)F
10 [H]C([H])([H])C(Cl)(Cl)[H]
4715 [H]C(F)(F)OC(Cl)(Br)Cl
4442 [H]C([H])(F)OC(Cl)(Cl)F
4694 [H]C(F)(F)OC(F)(Cl)F
4587 [H]C([H])(Br)OC(Cl)(Br)Cl
5039 [H]C(Cl)(Cl)OC(Br)(Br)Cl
5071 [H]C(Cl)(Br)OC(Br)(Br)[H]
4698 [H]C(F)(F)OC(Cl)(Cl)F
If necessary, further refinement can be based on atomic number range, boiling
point, flammability, and atmospheric lifetime. As described above, screening studies
do not necessarily require big data on the drug itself. Even if there is little data on
the MAC itself, it is possible to propose new candidates in an ordinal order.

2 Screening Methods for Drugs Using Chemoinformatics Methods … 43
https://t.me/med1917
Table 2.14 Reverse designed 3-carbon candidate compounds
SMILES SMILES
[H]C(F)(F)OC([H])([H])C(F)(F)F [H]C([H])(Br)OC(F)(F)C(F)(F)[H]
[H]C(F)(Br)OC([H])(F)C(F)(F)F [H]C(F)(F)OC([H])(Cl)C(F)(Br)F
FC(F)=C(Br)OC(Br)(Br)[H] C1([H])([H])C(F)(F)C1(F)(Cl)
[H]C(F)(Br)OC([H])(Cl)C(F)(Cl)F C1([H])([H])C([H])(F)C1(F)(Cl)
C1([H])([H])C([H])(F)C1([H])(Br) [H]C([H])(F)OC([H])(F)C(F)(Br)[H]
C1([H])([H])C([H])([H])C1([H])(F) C1([H])(F)C([H])(Br)C1(F)(F)
C1([H])([H])C([H])([H])C1(F)(F) [H]C(F)(F)OC([H])(F)C(F)(Br)F
FC(Cl)=C(Br)OC(Cl)(Br)[H] C1([H])(F)C([H])(F)C1([H])(Br)
[H]C([H])(Cl)OC(F)(Br)C(F)(F)[H] [H]C([H])(Cl)OC([H])(Br)C(Cl)(Cl)[H]
[H]C(F)(F)OC(F)(F)C(Br)(Br)F [H]C(F)(F)OC([H])(Br)C(F)(Br)F
[H]C(Cl)=C([H])OC(F)(Cl)[H] [H]C([H])(Cl)OC([H])(F)C(Cl)(Cl)F
ClC(Cl)=C(Br)OC(F)(F)[H] [H]C([H])([H])OC([H])(F)C(F)(F)F
[H]C(F)(F)OC([H])(Cl)C(F)(F)F [H]C([H])=C(Cl)OC(Br)(Br)[H]
ClC(Cl)=C([H])OC(Cl)(Br)[H] [H]C(F)(F)OC(F)(F)C(Cl)(Cl)[H]
[H]C(F)(Br)OC([H])(F)C(Cl)(Br)F [H]C(Cl)=C([H])OC([H])(F)[H]
[H]C([H])(Cl)OC([H])(Br)C(F)(F)[H] [H]C(Br)=C([H])OC(F)(F)[H]
FC(F)=C(Cl)OC(Br)(Br)[H] C1([H])([H])C([H])(F)C1([H])(F)
[H]C(Cl)(Cl)OC([H])(F)C(F)(Cl)F [H]C([H])(Br)OC(F)(F)C(Cl)(Cl)[H]
FC(F)=C([H])OC(F)(Cl)F FC(Cl)=C(Cl)OC([H])(Br)[H]
[H]C([H])([H])OC([H])([H])C(F)(Br)F C1([H])([H])C([H])(F)C1(F)(F)
C1([H])([H])C(F)(Cl)C1(F)(Cl) [H]C(F)(Br)OC([H])(F)C(Cl)(Br)Cl
[H]C([H])(Br)OC([H])([H])C(F)(Br)F [H]C([H])(Br)OC([H])(F)C(F)(Cl)F
[H]C(F)(F)OC([H])(Br)C(F)(F)F [H]C([H])([H])OC([H])(F)C(F)(Br)F
ClC(Cl)=C([H])OC([H])(Br)[H] [H]C([H])([H])OC(Cl)(Cl)C(F)(F)[H]
[H]C([H])([H])OC(F)(F)C(Cl)(Cl)[H] [H]C([H])=C(Br)OC(Br)(Br)[H]
C1([H])(F)C([H])(Br)C1(F)(Cl) [H]C(F)(F)OC([H])(F)C(F)(Cl)F
FC(Cl)=C(Cl)OC(F)(F)[H] C1([H])(Cl)C([H])(Br)C1(F)(Cl)
[H]C([H])([H])OC([H])(Br)C(Br)(Br)Cl
2.8 Useful Tips for Machine Learning
2.8.1 LASSO % Error Analysis
The LC50 index has been tested in Rat and Mouse. As explained earlier, LASSO
regression is used to minimize the sum of errors between experimental and calculated values, with the sum of absolute values of regression coefficients also taken

44 H. Yamamoto
https://t.me/med1917
Fig. 2.29 LASSO % error analysis. a LASSO regression results for LC50. b LASSO regression
with % error. c Normal error vs. % error
into account as a penalty (Fig.
2.29a). In many cases, convergence calculations are
performed using a Microsoft Excel solver.
So, it is easy to apply the % error as the error. The % error is larger the smaller
the experimental value (the higher the toxicity). When the LC50 is 10, a shift of 1 is
10%, while a shift of 1 when the LC50 is 100 is 1%. Therefore, if I introduce the %
error into the LASSO regression, I can search for regression coefficients so that the
fitting is higher where the LC50 is small. Conversely, the fitting will be poor where
the LC50 is large (Fig.
2.29b).
This effect may not seem large. I have calculated the predictions for all compounds
in the database. Let’s zoom in on the particularly problematic areas of toxicity, from
−3to3(Fig.
2.29c). The compounds that were rated as less toxic in the regular
LASSO regression and more toxic in the % error LASSO regression appear on the
left-lower side of the Y = X line.
In compound screening studies, the goal is not to obtain high correlation coefficients. As a way to exclude those with bad (highly toxic) ratings, or at the very least,
to give an estimate that calls for caution when handling, LASSO % error regression
is a simple and highly effective method. Compounds reverse-designed by a computer
often carry no information, so it is a good idea to do the calculations just to be sure.
This method is not limited to regions close to the origin but is also a way to increase
the accuracy of one’s target region. For example, in the case of yields in synthesis,
this method is used when you want to increase the accuracy of a high region.

2 Screening Methods for Drugs Using Chemoinformatics Methods … 45
https://t.me/med1917
2.8.2 LASSO Regression with Upper and Lower Limits
Solubility in water was often expressed as gram of solubility in 100g water. Therefore,
highly water-soluble compounds are often not measured above 100g/100g water.
When conducting an analysis that includes such data, it is necessary to devise a
method for capturing errors when the calculated value of log S is greater than 2. For
example, if the regression coefficient for hydroxyl groups is large, some compounds
with multiple hydroxyl groups will have calculated values that greatly exceed 2. If
the calculated value of 2 or more is set as the upper limit, the calculation error will
always be zero no matter how large the regression coefficient of the hydroxyl group
is. In such cases, it is effective to use LASSO regression and incorporate the sum of
the absolute values of the regression coefficients as a penalty in the evaluation.
References
1. Koblin DD, Laster MJ, Ionescu P, Gong D, Eger E I II, Halsey MJ, Hudlicky T (1999)
Polyhalogenated Methyl Ethyl Ethers: Solubilities and Anesthetic Properties. Anesth Analg
88(5):1161–1167.
2. RDKit: Open-Source Cheminformatics Software. https://www.rdkit.org. Accessed 15 Feb 2024.
3. Weininger, D (1988) Smiles, a chemical language and information system. 1. introduction to
method- ology and encoding rules. J Chem Inf Comp Sci 28:31–36.
0057a005.
4. Open MOPAC. http://openmopac.net. Accessed 15 Feb 2024.
5. Pople JA, Beveridge DL (1970) Approximate Molecular Orbital Theory. McGraw-Hill, New
Yor k .
6. Aoyama T, Ichikawa H (1991) Reconstruction of weight matrices in neural networks : A method
of correlating outputs with inputs. Chem Pharm Bull 39:1222–1228.
39.1222.
7. Kier LB, Hall LH (2000) Intermolecular Accessibility: The Meaning of Molecular Connectivity.
J Chem Inf Comput Sci 40(3):792–795.
8. Microsoft Corporation (2021) Microsoft Excel. https://office.microsoft.com/excel. Accessed 16
Feb 2024.
https://doi.org/10.1097/00000539-199905000-00036.
https://doi.org/10.1021/ci0
https://doi.org/10.1248/cpb.
https://doi.org/10.1021/ci990135s.

Chapter 3
https://t.me/med1917
Data-Driven Molecular Structure
Generation for Inverse QSPR/QSAR
Problem
Tomoyuki Miyao and Kimito Funatsu
3.1 Introduction
A fundamental goal of chemistry is to find new substances that have desired properties. To that end, data-driven chemistry utilizes quantitative structure–property
and structure–activity relationship (QSPR and QSAR) models to estimate property
(activity) values for the chemical structures of compounds. A QSPR/QSAR model
takes a molecular descriptor as the input and outputs property values through an algorithm, as seen in many QSPR/QSAR studies. As there is no conceptual difference
between the QSPR and QSAR at a methodological level, in this chapter, we refer to
the QSPR when statements are applicable to both the QSPR and QSAR.
For molecular design, QSPR models are frequently used to screen a large
compound library to make a focused library. The focused library contains a smaller
number of compounds that have desired properties according to the model output.
This naïve approach is called virtual screening (VS), and many successful applica-
1
tions of VS have been reported for various targets from energetic materials [
] and antibiotics [3]. One obvious limitation of VS is that it cannot
insecticides [
propose novel compounds that are not included in a screening library. Thus, although
VS is an effective method, new ideas are still needed to design novel chemical
structures of compounds, and it has been a central issue in chemoinformatics and
data-driven chemistry.
The inverse QSPR problem can be defined as a problem of proposing chemical
structures that exhibit a desirable (predicted) property based on a QSPR model; i.e.,
2
]to
T. Miyao
Data Science Center, and Graduate School of Science and Technology, Nara Institute of Science
and Technology, 8916-5 Takayama-Cho, Ikoma 630-0192, Nara, Japan
miyao@dsc.naist.jp
e-mail:
K. Funatsu (B)
Data Science Center, Nara Institute of Science and Technology, Ikoma, Nara, Japan
e-mail: funatsu@dsc.naist.jp
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2024
H. Satoh et al. (eds.), Drug Development Supported by Informatics,
https://doi.org/10.1007/978-981-97-4828-0_3
47

48 T. Miyao a nd K. Funatsu
https://t.me/med1917
Fig. 3.1 Two approaches to the inverse QSPR problem. a B-IQW generates chemical structures
from descriptors selected to show a specific property value. b F-IQW applies QSPR model to predict
a property value from a chemical structure for structure modification
inversing the flow of QSPR analysis to propose chemical structures from a property
value. Workflows to solve the inverse QSPR problem (inverse QSPR workflows
(IQWs)) can be classified into two categories according to the direction of the use of
QSPR models: the forward IQWs (F-IQWs, Fig.
3.1a).
Fig.
3.1b) and backward IQWs (B-IQWs,
F-IQWs start with random modification of chemical structures (Fig. 3.1b). These
chemical structures are then filtered on the basis of QSPR models and modified for
the next iteration of structure evaluation [
4, 5]. This process is repeated until chemical
structures with the desired properties are found or convergence is reached. F-IQWs
have been extensively investigated most likely for their simplicity. In contrast, BIQWs or truly inverse QSPR workflows (Fig.
3.1a) have seldom been reported [6].
In these approaches, a specific value of the objective variable (property value) is
set, and the inverse analysis of the QSPR model then determines the descriptor
information that leads to the specific property value. Finally, a structure generator
generates chemical structures per the descriptor information. Although F-IQWs are
simpler and easier to implement as computational systems, B-IQWs have a merit
of efficiently generating chemical structures satisfying a specific property value,
because the descriptor values specified for structure generation are assured to produce
the desired property values by the model. Methods for the inverse QSPR problem
and its two-way approaches were previously reviewed [
7].
3.2 B-IQW via a Bayesian Framework
In 2010, Funatsu and Miyao proposed a new B-IQW via a Bayesian framework [8
which generates chemical structures that have desired properties. Before going into
the methodological details, we discuss the limitations of conventional B-IQWs in
the following section.
],

3 Data-Driven Molecular Structure Generation for Inverse QSPR/QSAR … 49
https://t.me/med1917
3.2.1 Conventional B-IQWs
Conventionally, the methods for the inverse QSPR problem focused on the mathematical retrieval of chemical graphs from a set of descriptor values [
the group of Faulon JL proposed an algorithm for generating molecular structures
from a set of values of their developed topological descriptor, i.e., signature descrip-
12]. Targeting HIV-1 protease inhibitors, they successfully proposed a focused
tors [
library of chemical structures having desired activity values [
was first to identify a set of signature values corresponding to a desired activity value
predicted using a multivariate linear regression (MLR) model, and then to exhaustively generate chemical structures from the signature set with a structure generation
algorithm. This mathematically well-defined approach has two limitations in practical QSPR analysis, namely that only the signature descriptor is available and that
the applicability domain (AD) of the QSPR model has not been incorporated [
The molecular descriptor or representation plays a central role in QSPR modeling.
Various descriptors have been proposed to numericalize chemical structures in a way
that is relevant to a target property. It is a common understanding that there is no
single descriptor effective in every property prediction, and appropriate descriptors
depend on target properties. Thus, the descriptor available in B-IQWs should not be
limited to a specific type.
The AD of a QSPR model refers to regions in the chemical space (descriptor
space) where the output of the QSPR model is reliable. As all QSPR models are
regression models, a valid prediction is assumed only for compounds within the AD.
Many approaches have been proposed to define the AD of a QSPR model [
and most approaches share the basic concept that the high-dense region of the training
compound distribution should be the AD. In IQW, chemical structures are generated
using a QSPR model, and the AD must therefore be incorporated into the analysis
workflow irrespective of the forward or backward nature of the analysis.
9–11]. For example,
13]. Their approach
14].
14–17],
3.2.2 Bayesian Approaches for B-IQW
Our B-IQW has two steps like standard methods for the problem, namely the determination of a region of the chemical space where corresponding property values are
desirable and the exhaustive generation of structures for the determined region [
In this section, x represents the descriptor (explanatory variable) and y represents the
property (objective variable) for simplicity.
8].

50 T. Miyao a nd K. Funatsu
https://t.me/med1917
3.2.2.1 Posterior Distribution of X Given a Specific Y Value: P(X|y)
In general, identifying x corresponding to a y value for a QSPR is not tractable
because the QSPR model is a non-injective function; i.e., two descriptors may correspond to the same property value. Furthermore, the AD must be incorporated in the
inverse analysis. We propose using probability distributions in combination with the
Bayesian framework to achieve these two goals simultaneously.
The probability density function of x: p(x) represents the distribution of training
data (the prior distribution). In the proposed method, p(x) is expressed as a Gaussian
mixture model (GMM) with a weight p(z
component of the M Gaussians (eq.
takes a value of 1 when x belongs to the k-th component and a value of zero
z
k
), mean µk, and covariance ∑k for the k-th
k
3.1). The indicator variable (random variable)
otherwise.
M
p(x) =
p(zk )N (x|µk ,∑k ) (3.1)
k=1
As the compound distribution in the descriptor space is usually non-Gaussian
and multi-modal, a mixture of Gaussians is expected to well capture the training
compound distribution, forming a simple AD. The number of Gaussians, M, and
the shape of covariance, ∑
criterion [
18]. The parameters in p(x) can be determined using the expectation–maxi-
, can be determined based on the Bayesian information
k
mization algorithm to achieve a maximum of the likelihood function for a training
data set.
To capture the nonlinear relationship between x and y while maintaining an analytically tractable form for the posterior density derivation, we have proposed the adoption of three regression methods: MLR, component-wise MLR (cMLR), and Gaussian mixture regression (GMR) [
6]. cMLR places an MLR model per GMM compo-
nent, allowing pseudo-nonlinear regression models to be constructed. An MLR model
is formulated for each component of the Gaussians in Eq.
p(y|x, z
= Ny|a
)
k
T
x + bk ,σ
k
3.2.
2
k
(3.2)
The posterior distribution of x given a specific y value is expressed as Eq. 3.3.
Using the Bayes theorem, the posterior of x is obtained as Eq. 3.4.
p(x|y) =
p(x|y) =
M
p(x|y, zk )
k=1
M
p(x, zk|y) (3.3)
k=1
)p(zk )
p(y|z
k
M
p(y|zl)p(zl)
l=1
(3.4)
Соседние файлы в папке Библиотека им академика М.И. Перельмана
