Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5580_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
30.08.2026
Размер:
43 Мб
Скачать
30 H. Yamamoto
https://t.me/med1917
Fig. 2.18 Reaction rate constants and atmospheric lifetime with hydroxyl radicals
Fig. 2.19 Reaction scheme for hydroxyl radicals
The second is the addition of hydroxyl radicals when there is an unsaturated bond in the molecule, and the molecular radical collapse. So non-olefin, hydrogen-free CFCs have a particularly long atmospheric lifetime.
2.5.1.1 Prediction of Hydroxyl Radical Reaction Rates Using
Molecular Orbital Calculation Results
Since this is a physical property related to the reaction, I consider creating a prediction equation from the results of molecular orbital calculations. As shown in Fig. the accuracy of the estimated equation is low. If I try to exclude compounds with a log reaction rate constant of 12.5 or less, the calculated value can be limited to
14.0 or less. Very few compounds fall within that range.
To improve the prediction accuracy of molecular orbital calculations, I must use highly accurate basis functions and calculate the transition state of hydrogen withdrawal. However, that is not suitable for screening studies. This is because the computational load increases dramatically as the number of hydrogens increases in molecule.
2.20,
2 Screening Methods for Drugs Using Chemoinformatics Methods … 31
https://t.me/med1917
Fig. 2.20 Reaction rate constants of hydroxyl radicals and correlation using molecular orbital calculation results
2.5.1.2 Prediction of Hydroxyl Radical Reaction Rates Using LASSO
Regression Method
The functional group contribution method was used to develop an equation that predicts the reaction rate of hydroxyl radicals. The advantage of using functional groups is that the effect of the functional group can be clearly seen. The effect on atmospheric lifetime can be clearly seen by looking at the size of the coefficients. The usual multiple regression method seeks the regression coefficient that minimizes the error between the experimental value and the prediction equation. The LASSO regression method searches for the absolute sum of the regression coefficients as well as the calculation error to be as small as possible. While the Ridge regression method searches for the sum of squares of the regression coefficients to be as small as possible. In actual calculations, it is very difficult to find a balance between the smallest possible error and the smallest possible regression coefficient.
By performing LASSO regression, the regression coefficients were smaller than those from the multiple regression method, as shown in Fig.
2.21a.
In this LASSO regression, I also allowed for a sacrifice in the accuracy of the estimation. The estimation I wish to make here is to exclude compounds with log reaction rate constants below −12.5 (those with long atmospheric lifetimes). The usual multiple regression method excluded compounds with a calculated value of
13.4, as shown in Fig.
2.21b.
In LASSO regression, calculated values below 13.1 are excluded from the candidates, as shown in Fig.
2.21c.
Compared to the multiple regression method, the correlation is lower for the LASSO regression method. However, the number of compounds excluded from the candidate list is 81 for the multiple regression method and 98 for the LASSO regression method, so there is no problem in using this method for screening.
Functional groups with large regression coefficients increase the reaction rate constant. Those with unsaturated bonds in the molecule have larger regression coef­ficients. As shown in Table
2.5, some regression coefficients are larger in the usual
multiple regression method, but smaller in LASSO regression. The difference is particularly large for functional groups that are used infrequently in the analyzed
32 H. Yamamoto
https://t.me/med1917
Fig. 2.21 Comparison between LASSO- and Multiple Regression methods. a Difference in the coefficients between the multiple regression and the LASSO regression methods. Calculation of reaction rate constants for hydroxyl radicals by b multiple regression and c LASSO regression methods
data set. When designing molecules using regression coefficients, there is the option of using regression coefficients from LASSO regression, which can significantly change the results.
This is the difference between the use of estimating equations for the purpose of constructing property estimating equations and the use of estimating equations in screening.
2.5.2 Screening from Atmospheric Lifetime Estimates
There are 1865 compounds in the database that have a log reaction rate constant of 13.1 or greater. The regression coefficients of functional groups suggest that compounds with short atmospheric lifetimes have more double bonds, but compounds with double bonds are also known to be flammable. Therefore, compounds with short atmospheric lifetimes (OHR > 13.1) that are not obvi­ously flammable (logistic FP < 0.75) by the flammability index are reduced to 501 compounds. Furthermore, if the condition that the calculated boiling point falls in the range of 0–150 °C is added, the number of candidate compounds is reduced to 351 compounds. The range of atoms constituting the candidate molecules is shown in Table are excluded from the candidates.
2.6. The range of fluorine is narrower because perfluorinated compounds
2 Screening Methods for Drugs Using Chemoinformatics Methods … 33
https://t.me/med1917
Table 2.5 Functional groups with large differences in coefficients between multiple regression (MR) and Lasso regression methods
Table 2.6 Atomic number range of candidate compounds
C# H# Br# Cl# F# I# N# O#
Min 1 0 0 0 0 0 0 0
Max 10 7 3 4 18 2 1 2
Label MR LASSO In data
NH 0.97 0.01 1
#C(Hal) 0.89 0.01 1
NH2 0.64 0.24 3
C(Hal)= 0.28 0.36 41
CH(Hal)= 0.37 0.38 26
3_OH 0.63 0.39 7
CH 0.49 0.41 74
COOH 0.58 0.43 26
CH=_R 0.67 0.44 10
OH 0.68 0.58 29
CH2= 0.82 0.86 57
CH= 1.07 1.02 46
C= 1.24 1.02 14
HCO 1.17 1.10 8
#C 1.62 1.38 3
2.6 Solubility Indices
There are no specific target values for solubility indices, only moderate solubility and partitioning (solubility in water, octanol/water partition ratio, and bioconcentration). Many formulas have been developed to estimate log Kow and solubility in water (log S: g/100g water). However, not much data are available for bioconcentration factor (log BCF). In particular, there are not much experimental data for halogenated compounds.
Even when the number of data is not large, it is possible to apply various methods to estimate equations, as in the case of the boiling point estimation. In this section, I will explain the transfer learning that can be used for physical properties such as log Kow, log S, and log BCF.
34 H. Yamamoto
https://t.me/med1917
2.6.1 Transfer Learning of Solubility Indices
The basic assumption is that multiple physical property values must be highly corre­lated to some extent. For example, there is an inverse correlation for log Kow and log S as shown in Fig.
The log Kow and log BCF are also positively correlated as shown in Fig. 2.22b. log Kow is an index that increases as the molecule increases in size. log BCF is known to increase as the molecule increases in size. However, it is known that log BCF conversely becomes smaller as larger molecules cannot be taken up by the organism. In addition, some compounds, such as chlorine compounds, are specifically taken up by the organism. In the case of such correlations, NN methods using transfer learning NthNN (in-house software) are used for analysis.
When analyzing phenomena such as the presence or absence of flammability as in Logistic regression using the usual NN method, two neurons are placed in the output layer. A slight modification of that program would be the QSAR program. Then, as shown in Fig. log Kow, log S, and log BCF. (log S is multiplied by 1 for positive correlation.) This learning is valid because the three properties are highly correlated. Then, the middle and output layers learn specifically for each of the three physical properties.
2.22a.
2.23, the input and middle layers are commonly trained with
Fig. 2.22 Correlation between a octanol/water partition ratio and solubility in water and b octanol/ water partition ratio and bio concentration factor
Fig. 2.23 Transfer learning NthNN
Learning to output only
FG1
FG2
FGn
Learning in Common
log Kow
log S
log BCF
2 Screening Methods for Drugs Using Chemoinformatics Methods … 35
https://t.me/med1917
Fig. 2.24 Correlation between experimental values and transfer learning results. a log Kow. b log S. c log BCF
The log BCF has a very small amount of data. Therefore, it is not possible to establish NN learning from log BCF data alone. Learning of input and middle layers is established from many log Kow and log S data. log BCF less data is mainly used to recognize the relationship between middle and output layers. After the NN calculations converged, the learning is shown in Fig.
2.24a, b, c.
Once the Learning is complete, I will be able to calculate each solubility index based on the number of functional groups that make up the molecule. If the accuracy of the physical property estimation is an issue, then CV is performed. If it is simply used for screening, examine the predicted values of the compounds that were not used in the study.
As shown in Fig. 2.25a, for log Kow, the calculated values for the compounds used for training are distributed from 1 to 11. In the group of compounds for prediction, there are some compounds with log Kow greater than 11. These were compounds that are very large n-alkanes with one halogen attached. The log Kow of such compounds is known to be this large. So, the calculated value of this log Kow is not a problem.
As shown in Fig. 2.25b, solubility in water (1*log S) is also not a problem.
As for log BCF, the very small amount of training data used, as shown in Fig. 2.25c, suggests that overlearning occurred.
Although the log BCF value may have lost its original meaning, three dissolution indices can be obtained once the structure of the compound is determined.
The calculated solubility indices for the 26 compounds with known MACs are showninTable
2.7. The log KMAC is the logarithm of the MAC multiplied by 1000.
36 H. Yamamoto
https://t.me/med1917
Fig. 2.25 Results of calculations used in training and predictions for compounds not used in training. a log Kow. b −log S. c log BCF
2.6.2 Self-Organizing Map (SOM) Method Analysis
of Solubility Indices
The SOM method is a mapping method of multidimensional vectors to two dimen­sions. It is a technique for mapping similar multidimensional vectors to similar locations on two dimensions. Here, the three-dimensional vector values of log Kow,
log S, and log BCF calculated by NthNN are mapped to similar two-dimensional locations, as shown in Fig. on the two dimensions. This means that the three solubility index vectors are similar.
If screening is done using the average values of log Kow, log S, and log BCF of the three compounds with the smallest MACs in the red circles as indicators. I can search for compounds with small MACs around this SOM area.
2.26. The three smallest MACs are mapped close together
2.6.3 K-Means Analysis of Solubility Indices
If the SOM method is able to classify the compounds, there may be no need to try other classification methods. However, I do not know if the compound with the lowest
2 Screening Methods for Drugs Using Chemoinformatics Methods … 37
https://t.me/med1917
Table 2.7 Calculated solubility indices for 26 compounds
Compound Log KMAC NthNN log P NthNN −log S NthNN Log BCF
CH3OCF2CHCI
CF2HOCBrHCF
0.43 1.47 0.02 1.97
2
0.72 1.57 0.49 1.95
3
CH3OCF2CBrFH 0.84 1.53 0.57 2.34
CF2HOCCIHCF
1.16 1.74 0.65 1.82
3
CF2HOCF2CBrCIF 1.18 2.04 0.34 1.05
CF2HOCBrCICF
1.18 2.08 1.36 1.15
3
CH3OCF2CCIFH 1.20 1.67 0.53 0.58
CF2HOCF2CCIFH 1.34 1.78 0.55 0.13
CCIF2OCF2CCIFH 1.48 2.55 1.88 0.28
CFH2OCF2CF2H 1.62 1.57 0.45 0.81
CFH2OCFHCF
CCIF2OCCIHCF
CF2HOCFHCF
CF2HOCF2CFCI
CF2HOCCI2CF
CF2HOCH2CF
3
CCIF2OCCI2CF
CCI2FOCF2CCIF
1.67 1.43 0.65 0.16
3
1.69 2.57 1.97 1.65
3
1.89 1.26 0.04 1.70
3
1.95 2.46 1.85 2.61
2
1.99 2.60 1.60 2.30
3
2.04 1.51 0.65 1.70
2.12 3.21 2.71 2.46
3
2.26 3.32 2.84 3.39
2
CCIF2OCF2CCI2F 2.27 3.32 2.84 3.39
CCIF2OCH2CF
CCIF2OCFHCF
CF2HOCCIFCF
CF2HOCF2CCIF
CF3OCFHCF
CF2HOCF2CF
3
3
CH2FOCH(CF3)
2.46 2.36 2.00 1.65
3
2.46 2.51 2.41 1.18
3
2.51 2.16 0.70 0.89
3
2.78 2.32 1.50 2.55
2
3.29 1.84 1.28 0.26
3.75 1.94 0.89 2.35
3.78 2.01 0.61 0.57
2
MAC is the target indicator. They may be too anesthetic, and they may not wake up and become highly toxic.
The K-Means method classifies vectors into K types. The method is shown in
2.27. First, classify each compound into K types by some method. Although
Fig. it is possible to divide the compounds into K types by random numbers, in some cases, the average values of the groups classified by random numbers may be very similar. In such cases, the group may end up with no members. It is better to give some rough initial values for classification to stabilize the results. The next step is to find the average value of the group. Calculate the distance from each compound to the mean of the group. Change each compound so that it belongs to the group with the closest distance. As each group member is changed, the group mean is changed.
38 H. Yamamoto
https://t.me/med1917
Fig. 2.26 SOM analysis using solubility indices
This operation is repeated until the group membership no longer changes. The results are shown in Table
Classify into 4 types.
Fig. 2.27 Classification method using K-Means method
2.8.
Random numbers to define groups.
New center
Find the center of each member
Repeat
Center
Nearest center from each member.
2 Screening Methods for Drugs Using Chemoinformatics Methods … 39
https://t.me/med1917
Table 2.8 Classification results using K-Means method
NthNN log Kow NthNN −log S NthNN log BCF Ave. log K MAC
Gr.1 1.58 0.46 1.98 1.55
Gr.2 2.26 1.31 0.78 1.72
Gr.3 2.87 2.22 2.78 2.23
Gr.4 2.48 2.13 1.49 2.20
Gr.5 1.76 0.63 0.22 2.01
Table 2.9 Solubility index range for compounds in Group 1
Gr1 NthNN logKow NthNN −log S NthNN log BCF
Min 1.26 0.04 1.70
Max 1.94 0.89 2.35
Gr. 1 is the group with the highest anesthetic performance (less dose, more effect). Compounds belonging to Gr.1 have dissolution indices in the following range in
2.9.
Table
2.6.4 Screening Using Solubility Indices
Compounds were screened from the entire database under the condition of solubility indices entering Gr. 1. 293 compounds were found to match the condition at log Kow. Similarly, there were 369 compounds for log S and 246 compounds for log BCF. However, compounds that fit all three conditions were found in 18 compounds, 7 of which were compounds that were in the original paper. Of the remaining 11 compounds, one was bromine compound (CH was chlorine compound (CH
CHCl2, CAS: 75-34-3). The anesthetic effect may be
3
3CH2CH2
similar to that of chloroform and bromoform. The rest were ether compounds shown in Table
Table 2.10 Candidate ether compounds in Group 1
2.10.
CAS Structure logistic-FP OHR calc
84011-15-4 CF2CH2OCF
461-22-3 CF3CH2CH2OCH
461-24-5 CF3CH2OC2H
32793-56-9 CF3CH(CH3)OCF2H 0.82 13.32
26885-67-6 CF3CH=CHOCH
50285-05-7 CF3CHFOCH
2356-61-8 CHF2CF2OCF
53997-64-1 CF3CF2OCF2H 0.02 13.98
Br, CAS: 74-96-4) and one
3
5
3
3
0.41 13.75
0.99 12.59
3
0.99 12.59
1.00 10.94
3
0.82 13.26
0.02 13.98