Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5338_Библиотеки_им_академика_М_И_Перельмана
.pdf
10 H. Yamamoto
https://t.me/med1917
containing trifluoromethyl groups do not exist in nature. Therefore, real data are
very limited.
The following halogenated compounds are known for inhalation anesthetics.
•
Desflurane CF3CHFOCHF2 boiling point = 23.35 °C
•
Sevoflurane (Sevoflurane) (CF3)2CHOCH2F boiling point = 58.5 °C
•
Isoflurane CF3CHClOCHF2 boiling point = 48.5 °C
These anesthetics are supposed to be nonflammable, non-explosive, and fatsoluble, as well as to interfere with the transmission of nerve impulses. They must also
be organic liquids with low boiling points that evaporate easily at room temperature.
A low blood/gas partition coefficient is important because it allows for faster awakening from anesthesia. Environmental impact assessment is also important, since it
is released into the environment. The referenced paper [
concentrations (MAC) for 26 halogenated compounds. Thus, the chemoinformatics
method is applied to the design of anesthetics, for which there is little available data,
and the goal of development is vague, to screen new candidate compounds.
1] lists minimum alveolar
2.2 Analysis Preparation
2.2.1 Target Values of Physical Properties as an Anesthetic
The target values for anesthetics should be set as follows.
•
Boiling point = between 20 and 70 °C
•
Nonflammability
•
Moderate solubility, partitioning (solubility in water, octanol/water partition ratio,
bioaccumulation)
•
Low global warming potential (GWP), short atmospheric lifetime, ozone deple-
tion potential (ODP) of 0
•
Low toxicity: LC50
2.2.2 Data Collection
In order for the computer to make machine learning, data is needed. The data source
was extracted from the database for HSPiP (Hansen Solubility Parameters in Practice)
software maintained by the author for compounds containing halogens. For the sake
of simplicity, compounds containing P, S, Si, and B atoms and aromatic compounds
were excluded. In total, there were 2484 compounds. The compounds were classified
by type, and the list of physical properties is shown in Table
more than one type of functional group are counted multiple times, so the actual
number of valid data will be less than this.
2.1. Compounds with

2 Screening Methods for Drugs Using Chemoinformatics Methods … 11
https://t.me/med1917
Table 2.1 Collected physical properties and number of records per compound type
Type Tota l BP OHR Flash
point
Olefine 930 610 163 182 164 163 22 16 49 807
Alcohol 97 61 46 54 46 46 6 10 16 82
Ether 564 422 68 138 69 113 9 9 26 547
Ketone 186 134 46 58 46 46 2 5 12 166
Acid 55 30 27 33 27 27 6 2 9 43
Ester 150 114 48 80 48 48 2 2 11 127
Other 909 746 431 438 431 363 39 26 70 763
Tota l 2891 2117 829 983 831 806 86 70 193 2535
OHR: Rate constant for reaction with hydroxyl radicals
log
Kow
log S log
BCF
LC50 LD50 MO
calc
2.2.3 Creating Molecular Descriptors
Represent molecular structures in a form that computers can understand. Image
recognition technology has advanced rapidly. However, it will take more time before
computers can understand molecular structures. In this study, molecular structures
are handled by the SMILES molecular structure formula.
The following three methods are likely to be the most common ways to represent
molecules.
•
Group contribution method: It splits a molecule into functional groups that make
it up.
•
Topological Index: uses bonding information, etc.
•
Molecular Orbital Calculation: uses indices calculated from molecular orbital
calculations.
The RDKit software [2
lated from the SMILES [
] was used to generate the topological index. This is calcu-
] structural formula. RDKit also creates 3-dimensional
3
structures for molecular orbital calculations.
Molecular orbital calculations were performed from the 3-dimensional structures
] ver. 2012, and the
using the semi-empirical molecular orbital method, MOPAC [
in-house developed CNDO/2 [5
].
4
2.2.4 Set of Functional Groups
For example, for an ether with the basic skeleton CH3OCH2CH3, replacing hydrogen
with F, Cl, or Br, a huge number of functional groups are possible, e.g., CH
,CH2Cl, CHCl2, CCl3,CH2Br, CHBr2, CBr3, CHFCl, CHFBr, CHClBr, CF2Cl,
CF
3
,CF2Br, CFBr2, CCl2Br, CClBr2, and CFClBr. Furthermore, if iodine is also
CFCl
2
F, C H F2,
2

12 H. Yamamoto
https://t.me/med1917
Table 2.2 Definition of functional group
CH
3
CHCl
CH
2
CH CF CCl
C
CH2= #CH CH= #C CH(Hal)= #C(Hal)
C= C(Hal)= C(Hal)(Hal)=
OH 2_OH 3_OH
NH
2
O O_R O_EPO
C=O
COOH C#N HCO COO
H F Cl Br I
CH2_R CF2_R CHF_R CCl2_R CHCl_R CFCl_R
CH_R CF_R
C_R
CH=_R
C=_R C(Hal)=_R
N_R
C=O_R
= Double bond, # Triple bond, _R Cyclic functional group, -Ole Attached to an olefin
CF
3
CH2Cl CF2Cl CFCl
2
CF
2
NH N
F-Ole Cl-Ole Br-Ole I-Ole
CHF
2
CHF CCl
CH2F CCl
3
2
2
CHFCl
CHCl CFCl
taken into account, the functional group becomes too large. Therefore, Br and I
are treated as atoms. In this section, 66 functional groups are defined as shown in
2.2.
Table
The accuracy can be improved by adding the ring size and, in the case of olefins,
cis and trans information. However, here I will simplify and treat only functional
groups.
2.2.5 Goal of the Analysis
If the physical properties of all 2484 compounds are known and the target physical
properties are clearly known, sorting the table may extract the candidate compounds.
However, MAC, an indicator of anesthesia, has data for only 26 compounds at most.

2 Screening Methods for Drugs Using Chemoinformatics Methods … 13
https://t.me/med1917
The solubility index is also not specified as a range, but only as “moderate” hydrophobicity. The compounds in the database have been registered because some physicochemical property values existed, but there is a huge number of halogen compounds
that have not been registered.
Therefore, using machine learning, candidate compounds are selected from the
compounds in the database. In addition, screening of possible candidates by generating new structures is considered. Such machine learning for screening requires
ingenuity in the data input method and changes in the program structure. It is also
necessary to exclude data that is clearly erroneous and to correctly understand the
meaning of descriptor to select the best candidates. This requires knowledge of
chemistry.
2.3 Boiling Point Estimation
When used as an inhalation anesthetic, the compound is handled as a gas. If it is
only gassed, the temperature can be increased or diluted with air. However, when it
is expelled from the body, it must have a certain vapor pressure or awakening from
anesthesia will be delayed. Therefore, the boiling point range should be 20–70 °C.
Boiling points have long been estimated by the functional group contribution
method. A well-known method is the JOBACK method. By extending the functional
group to halogenated compounds, it is easy to construct an equation to estimate the
boiling point. Boiling points do not need to be very precise when used for screening.
However, the estimation of the boiling point is the basis of the estimation method
and will be explained in detail.
2.3.1 Estimation Methods and Evaluation Methods Used
for Boiling Point Estimation
2.3.1.1 Boiling Point Estimation by Multiple Regression
The functional group contribution method (multiple regression method) predicts the
target property (boiling point) based on the number of functional groups that make
up a molecule. A simultaneous equation is created, and the regression coefficient
of each descriptor is obtained. The coefficients can be obtained by solving these
simultaneous equations using the Gauss–Seidel method, the Gauss-Jordan method,
or the inverse matrix method. This is the basis of data analysis.

14 H. Yamamoto
https://t.me/med1917
2.3.1.2 Boiling Point Estimation by Neural Network (NN) Method
The simple functional group contribution method cannot account for nonlinearities in
the interactions and phenomena between functional groups. For example, an increase
of one amin group in the molecule raises the boiling point by 73.23 °C. An increase
of one carboxylic acid raises the boiling point by 169.09 °C. However, an amino
group with an amin group and a carboxyl group on one carbon will be greater than
the sum of the increases in each (it may carbonize without a boiling point). NN
methods, which easily take into account interactions and nonlinearities among input
items, were well studied until the early 2000s. However, they were used less and less
due to the overlearning problems described later. Recently, with advances in deep
learning, the NN method itself has been used more often. However, its use appears
to be slow in the area of chemistry due to the lack of big data and the low quality of
data.
2.3.1.3 Boiling Point Estimation Using Reduction of Dimension Method
The principal component analysis (PCA) or PLS method is probably the most
common recent research. The characteristic feature of this method is the reduction of multidimensional vectors. Let me explain this reduction of multidimensional
vectors. Since the general molecular formula is C
3-dimensional coordinates (Fig.
2.1a).
,this x, y, z can be plotted in
xHyOz
However, it is obvious to those who understand chemistry that the number of
hydrogen is equal to the number of carbon multiplied by two and added by two (y
= 2x + 2). In other words, by creating a new axis from the x-axis and y-axis, it
is possible to reduce the three dimensions to two. All points can be seen to be on
one plane (in two dimensions) (Fig.
2.1b). If the effect of this degeneracy is large,
there will be more experimental data than explanatory variables, so there will be less
overlearning.
(a)
3
2
O#
1
Fig. 2.1 Reducing the dimension of multidimensional vectors
H#
(b)
3
2
O#
1
1
C#
2
3
4
H#
4
1
2
3
C#

2 Screening Methods for Drugs Using Chemoinformatics Methods … 15
https://t.me/med1917
Applying the PCA method to a table of functional groups (66 dimensions) was not
effective. If the molecule becomes more complex, for example, if olefins or cyclic
compounds enter, the molecular formula just described becomes C
bonds enter, the formula becomes C
xH2x-2Oz
. When halogens enter, the number of
xH2xOz
;iftriple
H’s decreases by the number of halogens entered. Sometimes the number of hydrogen
atoms is zero. In this case, the direction of degeneracy in the C–H plane cannot be
found.
Similarly, the multiple regression method of variable selection is a kind of dimensional degeneracy. Especially when using RDKit’s explanatory variables, it can
reduce the dimension of the explanatory variables more efficiently than the PCA
method.
2.3.1.4 Evaluation Method of Calculation Results
The accuracy of property estimates is often evaluated by cross-validation (CV) of
the calculation results. CV is also called leave-several-out method. CV is especially
necessary when creating nonlinear estimation equations such as the NN method.
However, it should be recognized that CV is only valid when the compound to be
predicted is an interpolation of the compound to be learned. For example, there are
only two ester compounds in the LC50 data. If those two compounds are used as
the compounds for prediction, the results will be extrapolated and will not predict
correctly. This is easy to understand. The difficulty is when what appears to be an
interpolation is actually an extrapolation. For example, there was a compound in DB
in which an amin group was used twice. There is a compound in which a carboxyl
group is also used twice. Then, an amino acid with one amin group and one carboxyl
group each would appear to be an interpolation at first glance. However, if the amino
group is not studied, it is an extrapolation. CV results are poor, and examining the
reasons shows that they are extrapolated. Therefore, the scope of use is often limited
to the results. In the case of compound screening, it is meaningless to limit the
scope of use. In addition, the number of data i s not large enough to separate data
for prediction. Therefore, especially for physical properties for which the number of
data is not sufficient, the distribution of predicted values across the entire database
is used as an indicator.
2.3.2 Results of Various Boiling Point Estimation Methods
2.3.2.1 Boiling Point Estimation by Multiple Regression
The number of functional groups was taken as the explanatory variable, and the
boiling point contribution factor for each functional group was determined by
multiple regression method. The number of experimental data is 1891, and the
number of functional groups is 66. The results are shown in Fig.
2.2a.

16 H. Yamamoto
https://t.me/med1917
Fig. 2.2 Boiling point calculation by functional group contribution method. Calculation results of
a trained data using multiple regression, b prediction data using multiple regression, c trained data
using EBP-NN, and d prediction data using EBP-NN
For boiling points, a new search was performed for compounds with no boiling
points in the database, and the boiling points of 90 compounds were obtained to
verify the prediction accuracy of the multiple regression equation. The results are
shown in Fig.
2.2b.
This level of accuracy is sufficient for screening purposes, so the other boiling
point estimation methods in Sect.
2.3.2 can be skipped.
2.3.2.2 Boiling Point Estimation by Error Back-Propagation NN
(EBP-NN) Method
NN methods came to be used for estimating physical properties because of the
development of the error back propagation (EBP) method. However, it is also because
of the EBP method that it is no longer used. This was a famous saying around the
year 2000. The usual EBP-NN method was used to train on the same data set as
the multiple regression method. As shown in Fig.
converges, and a good-looking estimating equation can be constructed.
However, when the prediction is made using this estimation formula, the results
are shown in Fig.
The prediction performance is inferior to that of the multiple regression method.
This is due to overlearning, which is characteristic of NN methods. As shown in
2.2c, the machine learning easily
2.2d.

2 Screening Methods for Drugs Using Chemoinformatics Methods … 17
https://t.me/med1917
Fig. 2.3 Overlearning in the
NN method
overlearned curve
curve that we want
Fig. 2.3, the NN method finds a curve that connects the data points smoothly. Therefore, the error is very small in the vicinity of the data points, but outside of them, the
error deviates significantly.
Therefore, the number of intermediate neurons is reduced, or the training is terminated in the middle. However, if the prediction performance is made equal, the
convergence of the training data itself becomes poor, and there is no r eason to bother
using the NN method. The problem with the error back propagation method is that
the nonlinear fitting ability is too high, and the prediction performance is conversely
low.
2.3.2.3 Boiling Point Estimation by Reconstruction Learning (RCL)
NN Method
The r econstruction learning (RCL) method developed by Aoyama et al. introduces
the forgetting effect during EBP learning [
that small ones become smaller and large ones become larger with respect to the NN’s
weight matrix. They call this effect the forgetting effect.
Fig. 2.4 Introducing the forgetting effect into NN training
6] (Fig. 2.4). They introduce learning such
200 Times
1 Reconstruction

18 H. Yamamoto
https://t.me/med1917
When the threshold value of the weight matrix (Wij)isset to ζ , the weight matrix
= 1
= 0
2.1).
ij
W
W
ξ
ij
ij
1 − δW
<ξ
>ξ
ij
(2.1)
is increased or decreased according to Eq. (
W
= W
ij
− sgnW
ij
δW
ij
δW
ij
It can be seen from Fig. 2.5 that the introduction of the forgetting effect in the
estimation of the boiling point results in a smaller weight matrix compared to the
usual EBP method. As shown in Fig.
2.6a, the accuracy of the learned compounds is
not different from that of the EBP method. Moreover, the prediction performance is
sufficiently high compared to the multiple regression method as shown in Fig.
2.6b.
Although there is the problem of time-consuming learning, the effect of just a few
lines added to the program can be significant.
Fig. 2.5 Distribution of
weight matrix EBP and RCL
Fig. 2.6 Boiling point calculation by functional group contribution method. Calculation results of
a trained data using RCL-NN and b prediction data using RCL-NN
350
300
250
200
150
100
50
0
-0.01 0.01-0.1 0.01-0.5 0.5-1 1-
EBP RCL

2 Screening Methods for Drugs Using Chemoinformatics Methods … 19
https://t.me/med1917
2.3.2.4 Boiling Point Estimation by Hybrid Method
Chemical data tend to have a small number of data and a large number of explanatory
variables. Avoiding overlearning becomes an unavoidable problem when using NN
methods. The following is a method of estimation that I have used extensively on such
occasions. It is a simple and effective method, but I have not seen many examples
of its actual use. It can be said that it is a technique belonging to know-how. The
advantage of the multiple regression method is that the functional group contribution
factors are available, and their chemical implications are clear. The disadvantage is
that it cannot follow the interactions and nonlinearities between functional groups.
Therefore, the experimental values are divided into linear and nonlinear fractions. The
difference obtained by subtracting the multiple regression calculation value from the
experimental value is considered as the contribution of the nonlinear portion shown
2.7a. Only this difference is trained by the NN method.
in Fig.
The difference is between −40 and 60 °C, and most interactions are close to zero.
Directly, when boiling points are used as teacher data for NN methods, the values are
between 200–650 °C, so they are susceptible to overlearning. When an experienced
researcher’s brain looks at the structure of a compound and predicts the boiling point,
it first considers the average boiling point. They would then predict it by increasing
or decreasing nonlinear effects on it. This can be easily tried without changing the
program, as shown in Fig.
2.7b, which shows that the formula is almost equivalent
to the NN method.
As shown in Fig. 2.7c, for the data that was not used for training, it was found
that the prediction was made without overlearning.
Fig. 2.7 Boiling point prediction with hybrid method. a Concept of hybrid method, which calculates the boiling point using the multiple regression method first. The difference between the experimental and calculated values is calculated using NN. Calculation results of b trained data using
hybrid method and c prediction data using hybrid method
Соседние файлы в папке Библиотека им академика М.И. Перельмана
