Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5338_Библиотеки_им_академика_М_И_Перельмана
.pdf
30 H. Yamamoto
https://t.me/med1917
Fig. 2.18 Reaction rate
constants and atmospheric
lifetime with hydroxyl
radicals
Fig. 2.19 Reaction scheme
for hydroxyl radicals
The second is the addition of hydroxyl radicals when there is an unsaturated bond in
the molecule, and the molecular radical collapse. So non-olefin, hydrogen-free CFCs
have a particularly long atmospheric lifetime.
2.5.1.1 Prediction of Hydroxyl Radical Reaction Rates Using
Molecular Orbital Calculation Results
Since this is a physical property related to the reaction, I consider creating a prediction
equation from the results of molecular orbital calculations. As shown in Fig.
the accuracy of the estimated equation is low. If I try to exclude compounds with a
log reaction rate constant of −12.5 or less, the calculated value can be limited to −
14.0 or less. Very few compounds fall within that range.
To improve the prediction accuracy of molecular orbital calculations, I must
use highly accurate basis functions and calculate the transition state of hydrogen
withdrawal. However, that is not suitable for screening studies. This is because the
computational load increases dramatically as the number of hydrogens increases in
molecule.
2.20,

2 Screening Methods for Drugs Using Chemoinformatics Methods … 31
https://t.me/med1917
Fig. 2.20 Reaction rate
constants of hydroxyl
radicals and correlation
using molecular orbital
calculation results
2.5.1.2 Prediction of Hydroxyl Radical Reaction Rates Using LASSO
Regression Method
The functional group contribution method was used to develop an equation that
predicts the reaction rate of hydroxyl radicals. The advantage of using functional
groups is that the effect of the functional group can be clearly seen. The effect on
atmospheric lifetime can be clearly seen by looking at the size of the coefficients.
The usual multiple regression method seeks the regression coefficient that minimizes
the error between the experimental value and the prediction equation. The LASSO
regression method searches for the absolute sum of the regression coefficients as
well as the calculation error to be as small as possible. While the Ridge regression
method searches for the sum of squares of the regression coefficients to be as small
as possible. In actual calculations, it is very difficult to find a balance between the
smallest possible error and the smallest possible regression coefficient.
By performing LASSO regression, the regression coefficients were smaller than
those from the multiple regression method, as shown in Fig.
2.21a.
In this LASSO regression, I also allowed for a sacrifice in the accuracy of the
estimation. The estimation I wish to make here is to exclude compounds with log
reaction rate constants below −12.5 (those with long atmospheric lifetimes). The
usual multiple regression method excluded compounds with a calculated value of −
13.4, as shown in Fig.
2.21b.
In LASSO regression, calculated values below −13.1 are excluded from the
candidates, as shown in Fig.
2.21c.
Compared to the multiple regression method, the correlation is lower for the
LASSO regression method. However, the number of compounds excluded from
the candidate list is 81 for the multiple regression method and 98 for the LASSO
regression method, so there is no problem in using this method for screening.
Functional groups with large regression coefficients increase the reaction rate
constant. Those with unsaturated bonds in the molecule have larger regression coefficients. As shown in Table
2.5, some regression coefficients are larger in the usual
multiple regression method, but smaller in LASSO regression. The difference is
particularly large for functional groups that are used infrequently in the analyzed

32 H. Yamamoto
https://t.me/med1917
Fig. 2.21 Comparison between LASSO- and Multiple Regression methods. a Difference in the
coefficients between the multiple regression and the LASSO regression methods. Calculation of
reaction rate constants for hydroxyl radicals by b multiple regression and c LASSO regression
methods
data set. When designing molecules using regression coefficients, there is the option
of using regression coefficients from LASSO regression, which can significantly
change the results.
This is the difference between the use of estimating equations for the purpose of
constructing property estimating equations and the use of estimating equations in
screening.
2.5.2 Screening from Atmospheric Lifetime Estimates
There are 1865 compounds in the database that have a log reaction rate constant
of −13.1 or greater. The regression coefficients of functional groups suggest
that compounds with short atmospheric lifetimes have more double bonds, but
compounds with double bonds are also known to be flammable. Therefore,
compounds with short atmospheric lifetimes (OHR > −13.1) that are not obviously flammable (logistic FP < 0.75) by the flammability index are reduced to 501
compounds. Furthermore, if the condition that the calculated boiling point falls in
the range of 0–150 °C is added, the number of candidate compounds is reduced to
351 compounds. The range of atoms constituting the candidate molecules is shown
in Table
are excluded from the candidates.
2.6. The range of fluorine is narrower because perfluorinated compounds

2 Screening Methods for Drugs Using Chemoinformatics Methods … 33
https://t.me/med1917
Table 2.5 Functional groups
with large differences in
coefficients between multiple
regression (MR) and Lasso
regression methods
Table 2.6 Atomic number range of candidate compounds
C# H# Br# Cl# F# I# N# O#
Min 1 0 0 0 0 0 0 0
Max 10 7 3 4 18 2 1 2
Label MR LASSO In data
NH 0.97 0.01 1
#C(Hal) 0.89 0.01 1
NH2 0.64 0.24 3
C(Hal)= 0.28 0.36 41
CH(Hal)= 0.37 0.38 26
3_OH 0.63 0.39 7
CH 0.49 0.41 74
COOH 0.58 0.43 26
CH=_R 0.67 0.44 10
OH 0.68 0.58 29
CH2= 0.82 0.86 57
CH= 1.07 1.02 46
C= 1.24 1.02 14
HCO 1.17 1.10 8
#C 1.62 1.38 3
2.6 Solubility Indices
There are no specific target values for solubility indices, only moderate solubility and
partitioning (solubility in water, octanol/water partition ratio, and bioconcentration).
Many formulas have been developed to estimate log Kow and solubility in water
(log S: g/100g water). However, not much data are available for bioconcentration
factor (log BCF). In particular, there are not much experimental data for halogenated
compounds.
Even when the number of data is not large, it is possible to apply various methods
to estimate equations, as in the case of the boiling point estimation. In this section, I
will explain the transfer learning that can be used for physical properties such as log
Kow, log S, and log BCF.

34 H. Yamamoto
https://t.me/med1917
2.6.1 Transfer Learning of Solubility Indices
The basic assumption is that multiple physical property values must be highly correlated to some extent. For example, there is an inverse correlation for log Kow and
log S as shown in Fig.
The log Kow and log BCF are also positively correlated as shown in Fig. 2.22b. log
Kow is an index that increases as the molecule increases in size. log BCF is known
to increase as the molecule increases in size. However, it is known that log BCF
conversely becomes smaller as larger molecules cannot be taken up by the organism.
In addition, some compounds, such as chlorine compounds, are specifically taken up
by the organism. In the case of such correlations, NN methods using transfer learning
NthNN (in-house software) are used for analysis.
When analyzing phenomena such as the presence or absence of flammability as
in Logistic regression using the usual NN method, two neurons are placed in the
output layer. A slight modification of that program would be the QSAR program.
Then, as shown in Fig.
log Kow, −log S, and log BCF. (log S is multiplied by −1 for positive correlation.)
This learning is valid because the three properties are highly correlated. Then, the
middle and output layers learn specifically for each of the three physical properties.
2.22a.
2.23, the input and middle layers are commonly trained with
Fig. 2.22 Correlation between a octanol/water partition ratio and solubility in water and b octanol/
water partition ratio and bio concentration factor
Fig. 2.23 Transfer learning
NthNN
Learning to output only
FG1
FG2
FGn
Learning in Common
log Kow
log S
log BCF

2 Screening Methods for Drugs Using Chemoinformatics Methods … 35
https://t.me/med1917
Fig. 2.24 Correlation between experimental values and transfer learning results. a log Kow. b −
log S. c log BCF
The log BCF has a very small amount of data. Therefore, it is not possible to
establish NN learning from log BCF data alone. Learning of input and middle layers
is established from many log Kow and −log S data. log BCF less data is mainly
used to recognize the relationship between middle and output layers. After the NN
calculations converged, the learning is shown in Fig.
2.24a, b, c.
Once the Learning is complete, I will be able to calculate each solubility index
based on the number of functional groups that make up the molecule. If the accuracy
of the physical property estimation is an issue, then CV is performed. If it is simply
used for screening, examine the predicted values of the compounds that were not
used in the study.
As shown in Fig. 2.25a, for log Kow, the calculated values for the compounds used
for training are distributed from −1 to 11. In the group of compounds for prediction,
there are some compounds with log Kow greater than 11. These were compounds that
are very large n-alkanes with one halogen attached. The log Kow of such compounds
is known to be this large. So, the calculated value of this log Kow is not a problem.
As shown in Fig. 2.25b, solubility in water (−1*log S) is also not a problem.
As for log BCF, the very small amount of training data used, as shown in Fig. 2.25c,
suggests that overlearning occurred.
Although the log BCF value may have lost its original meaning, three dissolution
indices can be obtained once the structure of the compound is determined.
The calculated solubility indices for the 26 compounds with known MACs are
showninTable
2.7. The log KMAC is the logarithm of the MAC multiplied by 1000.

36 H. Yamamoto
https://t.me/med1917
Fig. 2.25 Results of calculations used in training and predictions for compounds not used in
training. a log Kow. b −log S. c log BCF
2.6.2 Self-Organizing Map (SOM) Method Analysis
of Solubility Indices
The SOM method is a mapping method of multidimensional vectors to two dimensions. It is a technique for mapping similar multidimensional vectors to similar
locations on two dimensions. Here, the three-dimensional vector values of log Kow,
−log S, and log BCF calculated by NthNN are mapped to similar two-dimensional
locations, as shown in Fig.
on the two dimensions. This means that the three solubility index vectors are similar.
If screening is done using the average values of log Kow, −log S, and log BCF of
the three compounds with the smallest MACs in the red circles as indicators. I can
search for compounds with small MACs around this SOM area.
2.26. The three smallest MACs are mapped close together
2.6.3 K-Means Analysis of Solubility Indices
If the SOM method is able to classify the compounds, there may be no need to try
other classification methods. However, I do not know if the compound with the lowest

2 Screening Methods for Drugs Using Chemoinformatics Methods … 37
https://t.me/med1917
Table 2.7 Calculated solubility indices for 26 compounds
Compound Log KMAC NthNN log P NthNN −log S NthNN Log BCF
CH3OCF2CHCI
CF2HOCBrHCF
0.43 1.47 0.02 1.97
2
0.72 1.57 0.49 1.95
3
CH3OCF2CBrFH 0.84 1.53 0.57 2.34
CF2HOCCIHCF
1.16 1.74 0.65 1.82
3
CF2HOCF2CBrCIF 1.18 2.04 0.34 1.05
CF2HOCBrCICF
1.18 2.08 1.36 1.15
3
CH3OCF2CCIFH 1.20 1.67 0.53 0.58
CF2HOCF2CCIFH 1.34 1.78 0.55 0.13
CCIF2OCF2CCIFH 1.48 2.55 1.88 0.28
CFH2OCF2CF2H 1.62 1.57 0.45 0.81
CFH2OCFHCF
CCIF2OCCIHCF
CF2HOCFHCF
CF2HOCF2CFCI
CF2HOCCI2CF
CF2HOCH2CF
3
CCIF2OCCI2CF
CCI2FOCF2CCIF
1.67 1.43 0.65 −0.16
3
1.69 2.57 1.97 1.65
3
1.89 1.26 −0.04 1.70
3
1.95 2.46 1.85 2.61
2
1.99 2.60 1.60 2.30
3
2.04 1.51 0.65 1.70
2.12 3.21 2.71 2.46
3
2.26 3.32 2.84 3.39
2
CCIF2OCF2CCI2F 2.27 3.32 2.84 3.39
CCIF2OCH2CF
CCIF2OCFHCF
CF2HOCCIFCF
CF2HOCF2CCIF
CF3OCFHCF
CF2HOCF2CF
3
3
CH2FOCH(CF3)
2.46 2.36 2.00 1.65
3
2.46 2.51 2.41 1.18
3
2.51 2.16 0.70 0.89
3
2.78 2.32 1.50 2.55
2
3.29 1.84 1.28 −0.26
3.75 1.94 0.89 2.35
3.78 2.01 0.61 −0.57
2
MAC is the target indicator. They may be too anesthetic, and they may not wake up
and become highly toxic.
The K-Means method classifies vectors into K types. The method is shown in
2.27. First, classify each compound into K types by some method. Although
Fig.
it is possible to divide the compounds into K types by random numbers, in some
cases, the average values of the groups classified by random numbers may be very
similar. In such cases, the group may end up with no members. It is better to give
some rough initial values for classification to stabilize the results. The next step is to
find the average value of the group. Calculate the distance from each compound to
the mean of the group. Change each compound so that it belongs to the group with
the closest distance. As each group member is changed, the group mean is changed.

38 H. Yamamoto
https://t.me/med1917
Fig. 2.26 SOM analysis using solubility indices
This operation is repeated until the group membership no longer changes. The results
are shown in Table
Classify into 4 types.
Fig. 2.27 Classification method using K-Means method
2.8.
Random numbers to define groups.
New center
Find the
center of
each
member
Repeat
Center
Nearest center from
each member.

2 Screening Methods for Drugs Using Chemoinformatics Methods … 39
https://t.me/med1917
Table 2.8 Classification results using K-Means method
NthNN log Kow NthNN −log S NthNN log BCF Ave. log K MAC
Gr.1 1.58 0.46 1.98 1.55
Gr.2 2.26 1.31 0.78 1.72
Gr.3 2.87 2.22 2.78 2.23
Gr.4 2.48 2.13 1.49 2.20
Gr.5 1.76 0.63 0.22 2.01
Table 2.9 Solubility index
range for compounds in
Group 1
Gr1 NthNN logKow NthNN −log S NthNN log BCF
Min 1.26 −0.04 1.70
Max 1.94 0.89 2.35
Gr. 1 is the group with the highest anesthetic performance (less dose, more effect).
Compounds belonging to Gr.1 have dissolution indices in the following range in
2.9.
Table
2.6.4 Screening Using Solubility Indices
Compounds were screened from the entire database under the condition of solubility
indices entering Gr. 1. 293 compounds were found to match the condition at log
Kow. Similarly, there were 369 compounds for −log S and 246 compounds for log
BCF. However, compounds that fit all three conditions were found in 18 compounds,
7 of which were compounds that were in the original paper. Of the remaining 11
compounds, one was bromine compound (CH
was chlorine compound (CH
CHCl2, CAS: 75-34-3). The anesthetic effect may be
3
3CH2CH2
similar to that of chloroform and bromoform. The rest were ether compounds shown
in Table
Table 2.10 Candidate ether
compounds in Group 1
2.10.
CAS Structure logistic-FP OHR calc
84011-15-4 CF2CH2OCF
461-22-3 CF3CH2CH2OCH
461-24-5 CF3CH2OC2H
32793-56-9 CF3CH(CH3)OCF2H 0.82 −13.32
26885-67-6 CF3CH=CHOCH
50285-05-7 CF3CHFOCH
2356-61-8 CHF2CF2OCF
53997-64-1 CF3CF2OCF2H 0.02 −13.98
Br, CAS: 74-96-4) and one
3
5
3
3
0.41 −13.75
0.99 −12.59
3
0.99 −12.59
1.00 −10.94
3
0.82 −13.26
0.02 −13.98
Соседние файлы в папке Библиотека им академика М.И. Перельмана
