Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5858_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
30.08.2026
Размер:
43 Мб
Скачать
10 H. Yamamoto
https://t.me/med1917
containing trifluoromethyl groups do not exist in nature. Therefore, real data are very limited.
The following halogenated compounds are known for inhalation anesthetics.
Desflurane CF3CHFOCHF2 boiling point = 23.35 °C
Sevoflurane (Sevoflurane) (CF3)2CHOCH2F boiling point = 58.5 °C
Isoflurane CF3CHClOCHF2 boiling point = 48.5 °C
These anesthetics are supposed to be nonflammable, non-explosive, and fat­soluble, as well as to interfere with the transmission of nerve impulses. They must also be organic liquids with low boiling points that evaporate easily at room temperature. A low blood/gas partition coefficient is important because it allows for faster awak­ening from anesthesia. Environmental impact assessment is also important, since it is released into the environment. The referenced paper [ concentrations (MAC) for 26 halogenated compounds. Thus, the chemoinformatics method is applied to the design of anesthetics, for which there is little available data, and the goal of development is vague, to screen new candidate compounds.
1] lists minimum alveolar
2.2 Analysis Preparation
2.2.1 Target Values of Physical Properties as an Anesthetic
The target values for anesthetics should be set as follows.
Boiling point = between 20 and 70 °C
Nonflammability
Moderate solubility, partitioning (solubility in water, octanol/water partition ratio,
bioaccumulation)
Low global warming potential (GWP), short atmospheric lifetime, ozone deple-
tion potential (ODP) of 0
Low toxicity: LC50
2.2.2 Data Collection
In order for the computer to make machine learning, data is needed. The data source was extracted from the database for HSPiP (Hansen Solubility Parameters in Practice) software maintained by the author for compounds containing halogens. For the sake of simplicity, compounds containing P, S, Si, and B atoms and aromatic compounds were excluded. In total, there were 2484 compounds. The compounds were classified by type, and the list of physical properties is shown in Table more than one type of functional group are counted multiple times, so the actual number of valid data will be less than this.
2.1. Compounds with
2 Screening Methods for Drugs Using Chemoinformatics Methods … 11
https://t.me/med1917
Table 2.1 Collected physical properties and number of records per compound type
Type Tota l BP OHR Flash
point
Olefine 930 610 163 182 164 163 22 16 49 807
Alcohol 97 61 46 54 46 46 6 10 16 82
Ether 564 422 68 138 69 113 9 9 26 547
Ketone 186 134 46 58 46 46 2 5 12 166
Acid 55 30 27 33 27 27 6 2 9 43
Ester 150 114 48 80 48 48 2 2 11 127
Other 909 746 431 438 431 363 39 26 70 763
Tota l 2891 2117 829 983 831 806 86 70 193 2535
OHR: Rate constant for reaction with hydroxyl radicals
log Kow
log S log
BCF
LC50 LD50 MO
calc
2.2.3 Creating Molecular Descriptors
Represent molecular structures in a form that computers can understand. Image recognition technology has advanced rapidly. However, it will take more time before computers can understand molecular structures. In this study, molecular structures are handled by the SMILES molecular structure formula.
The following three methods are likely to be the most common ways to represent molecules.
Group contribution method: It splits a molecule into functional groups that make
it up.
Topological Index: uses bonding information, etc.
Molecular Orbital Calculation: uses indices calculated from molecular orbital
calculations.
The RDKit software [2 lated from the SMILES [
] was used to generate the topological index. This is calcu-
] structural formula. RDKit also creates 3-dimensional
3
structures for molecular orbital calculations.
Molecular orbital calculations were performed from the 3-dimensional structures
] ver. 2012, and the
using the semi-empirical molecular orbital method, MOPAC [ in-house developed CNDO/2 [5
].
4
2.2.4 Set of Functional Groups
For example, for an ether with the basic skeleton CH3OCH2CH3, replacing hydrogen with F, Cl, or Br, a huge number of functional groups are possible, e.g., CH
,CH2Cl, CHCl2, CCl3,CH2Br, CHBr2, CBr3, CHFCl, CHFBr, CHClBr, CF2Cl,
CF
3
,CF2Br, CFBr2, CCl2Br, CClBr2, and CFClBr. Furthermore, if iodine is also
CFCl
2
F, C H F2,
2
12 H. Yamamoto
https://t.me/med1917
Table 2.2 Definition of functional group
CH
3
CHCl
CH
2
CH CF CCl
C
CH2= #CH CH= #C CH(Hal)= #C(Hal)
C= C(Hal)= C(Hal)(Hal)=
OH 2_OH 3_OH
NH
2
O O_R O_EPO
C=O
COOH C#N HCO COO
H F Cl Br I
CH2_R CF2_R CHF_R CCl2_R CHCl_R CFCl_R
CH_R CF_R
C_R
CH=_R
C=_R C(Hal)=_R
N_R
C=O_R
= Double bond, # Triple bond, _R Cyclic functional group, -Ole Attached to an olefin
CF
3
CH2Cl CF2Cl CFCl
2
CF
2
NH N
F-Ole Cl-Ole Br-Ole I-Ole
CHF
2
CHF CCl
CH2F CCl
3
2
2
CHFCl
CHCl CFCl
taken into account, the functional group becomes too large. Therefore, Br and I are treated as atoms. In this section, 66 functional groups are defined as shown in
2.2.
Table
The accuracy can be improved by adding the ring size and, in the case of olefins, cis and trans information. However, here I will simplify and treat only functional groups.
2.2.5 Goal of the Analysis
If the physical properties of all 2484 compounds are known and the target physical properties are clearly known, sorting the table may extract the candidate compounds. However, MAC, an indicator of anesthesia, has data for only 26 compounds at most.
2 Screening Methods for Drugs Using Chemoinformatics Methods … 13
https://t.me/med1917
The solubility index is also not specified as a range, but only as “moderate” hydropho­bicity. The compounds in the database have been registered because some physico­chemical property values existed, but there is a huge number of halogen compounds that have not been registered.
Therefore, using machine learning, candidate compounds are selected from the compounds in the database. In addition, screening of possible candidates by gener­ating new structures is considered. Such machine learning for screening requires ingenuity in the data input method and changes in the program structure. It is also necessary to exclude data that is clearly erroneous and to correctly understand the meaning of descriptor to select the best candidates. This requires knowledge of chemistry.
2.3 Boiling Point Estimation
When used as an inhalation anesthetic, the compound is handled as a gas. If it is only gassed, the temperature can be increased or diluted with air. However, when it is expelled from the body, it must have a certain vapor pressure or awakening from anesthesia will be delayed. Therefore, the boiling point range should be 20–70 °C.
Boiling points have long been estimated by the functional group contribution method. A well-known method is the JOBACK method. By extending the functional group to halogenated compounds, it is easy to construct an equation to estimate the boiling point. Boiling points do not need to be very precise when used for screening. However, the estimation of the boiling point is the basis of the estimation method and will be explained in detail.
2.3.1 Estimation Methods and Evaluation Methods Used
for Boiling Point Estimation
2.3.1.1 Boiling Point Estimation by Multiple Regression
The functional group contribution method (multiple regression method) predicts the target property (boiling point) based on the number of functional groups that make up a molecule. A simultaneous equation is created, and the regression coefficient of each descriptor is obtained. The coefficients can be obtained by solving these simultaneous equations using the Gauss–Seidel method, the Gauss-Jordan method, or the inverse matrix method. This is the basis of data analysis.
14 H. Yamamoto
https://t.me/med1917
2.3.1.2 Boiling Point Estimation by Neural Network (NN) Method
The simple functional group contribution method cannot account for nonlinearities in the interactions and phenomena between functional groups. For example, an increase of one amin group in the molecule raises the boiling point by 73.23 °C. An increase of one carboxylic acid raises the boiling point by 169.09 °C. However, an amino group with an amin group and a carboxyl group on one carbon will be greater than the sum of the increases in each (it may carbonize without a boiling point). NN methods, which easily take into account interactions and nonlinearities among input items, were well studied until the early 2000s. However, they were used less and less due to the overlearning problems described later. Recently, with advances in deep learning, the NN method itself has been used more often. However, its use appears to be slow in the area of chemistry due to the lack of big data and the low quality of data.
2.3.1.3 Boiling Point Estimation Using Reduction of Dimension Method
The principal component analysis (PCA) or PLS method is probably the most common recent research. The characteristic feature of this method is the reduc­tion of multidimensional vectors. Let me explain this reduction of multidimensional vectors. Since the general molecular formula is C 3-dimensional coordinates (Fig.
2.1a).
,this x, y, z can be plotted in
xHyOz
However, it is obvious to those who understand chemistry that the number of hydrogen is equal to the number of carbon multiplied by two and added by two (y = 2x + 2). In other words, by creating a new axis from the x-axis and y-axis, it is possible to reduce the three dimensions to two. All points can be seen to be on one plane (in two dimensions) (Fig.
2.1b). If the effect of this degeneracy is large,
there will be more experimental data than explanatory variables, so there will be less overlearning.
(a)
3
2
O#
1
Fig. 2.1 Reducing the dimension of multidimensional vectors
H#
(b)
3
2
O#
1
1
C#
2
3
4
H#
4
1
2
3
C#
2 Screening Methods for Drugs Using Chemoinformatics Methods … 15
https://t.me/med1917
Applying the PCA method to a table of functional groups (66 dimensions) was not effective. If the molecule becomes more complex, for example, if olefins or cyclic compounds enter, the molecular formula just described becomes C bonds enter, the formula becomes C
xH2x-2Oz
. When halogens enter, the number of
xH2xOz
;iftriple
H’s decreases by the number of halogens entered. Sometimes the number of hydrogen atoms is zero. In this case, the direction of degeneracy in the C–H plane cannot be found.
Similarly, the multiple regression method of variable selection is a kind of dimen­sional degeneracy. Especially when using RDKit’s explanatory variables, it can reduce the dimension of the explanatory variables more efficiently than the PCA method.
2.3.1.4 Evaluation Method of Calculation Results
The accuracy of property estimates is often evaluated by cross-validation (CV) of the calculation results. CV is also called leave-several-out method. CV is especially necessary when creating nonlinear estimation equations such as the NN method. However, it should be recognized that CV is only valid when the compound to be predicted is an interpolation of the compound to be learned. For example, there are only two ester compounds in the LC50 data. If those two compounds are used as the compounds for prediction, the results will be extrapolated and will not predict correctly. This is easy to understand. The difficulty is when what appears to be an interpolation is actually an extrapolation. For example, there was a compound in DB in which an amin group was used twice. There is a compound in which a carboxyl group is also used twice. Then, an amino acid with one amin group and one carboxyl group each would appear to be an interpolation at first glance. However, if the amino group is not studied, it is an extrapolation. CV results are poor, and examining the reasons shows that they are extrapolated. Therefore, the scope of use is often limited to the results. In the case of compound screening, it is meaningless to limit the scope of use. In addition, the number of data i s not large enough to separate data for prediction. Therefore, especially for physical properties for which the number of data is not sufficient, the distribution of predicted values across the entire database is used as an indicator.
2.3.2 Results of Various Boiling Point Estimation Methods
2.3.2.1 Boiling Point Estimation by Multiple Regression
The number of functional groups was taken as the explanatory variable, and the boiling point contribution factor for each functional group was determined by multiple regression method. The number of experimental data is 1891, and the number of functional groups is 66. The results are shown in Fig.
2.2a.
16 H. Yamamoto
https://t.me/med1917
Fig. 2.2 Boiling point calculation by functional group contribution method. Calculation results of a trained data using multiple regression, b prediction data using multiple regression, c trained data
using EBP-NN, and d prediction data using EBP-NN
For boiling points, a new search was performed for compounds with no boiling points in the database, and the boiling points of 90 compounds were obtained to verify the prediction accuracy of the multiple regression equation. The results are shown in Fig.
2.2b.
This level of accuracy is sufficient for screening purposes, so the other boiling point estimation methods in Sect.
2.3.2 can be skipped.
2.3.2.2 Boiling Point Estimation by Error Back-Propagation NN
(EBP-NN) Method
NN methods came to be used for estimating physical properties because of the development of the error back propagation (EBP) method. However, it is also because of the EBP method that it is no longer used. This was a famous saying around the year 2000. The usual EBP-NN method was used to train on the same data set as the multiple regression method. As shown in Fig. converges, and a good-looking estimating equation can be constructed.
However, when the prediction is made using this estimation formula, the results are shown in Fig.
The prediction performance is inferior to that of the multiple regression method. This is due to overlearning, which is characteristic of NN methods. As shown in
2.2c, the machine learning easily
2.2d.
2 Screening Methods for Drugs Using Chemoinformatics Methods … 17
https://t.me/med1917
Fig. 2.3 Overlearning in the NN method
overlearned curve
curve that we want
Fig. 2.3, the NN method finds a curve that connects the data points smoothly. There­fore, the error is very small in the vicinity of the data points, but outside of them, the error deviates significantly.
Therefore, the number of intermediate neurons is reduced, or the training is termi­nated in the middle. However, if the prediction performance is made equal, the convergence of the training data itself becomes poor, and there is no r eason to bother using the NN method. The problem with the error back propagation method is that the nonlinear fitting ability is too high, and the prediction performance is conversely low.
2.3.2.3 Boiling Point Estimation by Reconstruction Learning (RCL)
NN Method
The r econstruction learning (RCL) method developed by Aoyama et al. introduces the forgetting effect during EBP learning [ that small ones become smaller and large ones become larger with respect to the NN’s weight matrix. They call this effect the forgetting effect.
Fig. 2.4 Introducing the forgetting effect into NN training
6] (Fig. 2.4). They introduce learning such
200 Times
1 Reconstruction
18 H. Yamamoto
https://t.me/med1917
When the threshold value of the weight matrix (Wij)isset to ζ , the weight matrix
= 1
= 0
2.1).
ij
W
W
ξ
ij
ij
1 δW

ij
(2.1)
is increased or decreased according to Eq. (
W
= W
ij
sgnW
ij
δW
ij
δW
ij
It can be seen from Fig. 2.5 that the introduction of the forgetting effect in the estimation of the boiling point results in a smaller weight matrix compared to the usual EBP method. As shown in Fig.
2.6a, the accuracy of the learned compounds is
not different from that of the EBP method. Moreover, the prediction performance is sufficiently high compared to the multiple regression method as shown in Fig.
2.6b.
Although there is the problem of time-consuming learning, the effect of just a few
lines added to the program can be significant.
Fig. 2.5 Distribution of weight matrix EBP and RCL
Fig. 2.6 Boiling point calculation by functional group contribution method. Calculation results of a trained data using RCL-NN and b prediction data using RCL-NN
350
300
250
200
150
100
50
0
-0.01 0.01-0.1 0.01-0.5 0.5-1 1-
EBP RCL
2 Screening Methods for Drugs Using Chemoinformatics Methods … 19
https://t.me/med1917
2.3.2.4 Boiling Point Estimation by Hybrid Method
Chemical data tend to have a small number of data and a large number of explanatory variables. Avoiding overlearning becomes an unavoidable problem when using NN methods. The following is a method of estimation that I have used extensively on such occasions. It is a simple and effective method, but I have not seen many examples of its actual use. It can be said that it is a technique belonging to know-how. The advantage of the multiple regression method is that the functional group contribution factors are available, and their chemical implications are clear. The disadvantage is that it cannot follow the interactions and nonlinearities between functional groups. Therefore, the experimental values are divided into linear and nonlinear fractions. The difference obtained by subtracting the multiple regression calculation value from the experimental value is considered as the contribution of the nonlinear portion shown
2.7a. Only this difference is trained by the NN method.
in Fig.
The difference is between 40 and 60 °C, and most interactions are close to zero. Directly, when boiling points are used as teacher data for NN methods, the values are between 200–650 °C, so they are susceptible to overlearning. When an experienced researcher’s brain looks at the structure of a compound and predicts the boiling point, it first considers the average boiling point. They would then predict it by increasing or decreasing nonlinear effects on it. This can be easily tried without changing the program, as shown in Fig.
2.7b, which shows that the formula is almost equivalent
to the NN method.
As shown in Fig. 2.7c, for the data that was not used for training, it was found that the prediction was made without overlearning.
Fig. 2.7 Boiling point prediction with hybrid method. a Concept of hybrid method, which calcu­lates the boiling point using the multiple regression method first. The difference between the exper­imental and calculated values is calculated using NN. Calculation results of b trained data using hybrid method and c prediction data using hybrid method