Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5580_Библиотеки_им_академика_М_И_Перельмана
.pdf
20 H. Yamamoto
https://t.me/med1917
As nonlinear analysis tools have advanced over time, this devising of input values
will have an even greater effect.
2.3.2.5 Boiling Point Estimation by Topological Index
If molecules are represented by the number of functional groups, the boiling points
of compounds with undefined functional groups cannot be predicted. Therefore,
explanatory factors such as the topological index (TI) calculated by RDKit and the
molecular orbital method are used as data input. The original idea of TI is probably
7
the Molecular Connectivity Index (MCI) [
]. Assign the number of heavy atoms to
be added to the carbon for each carbon–carbon bond in the molecule as shown in
2.8. Multiply this number by −0.5 and take the sum of each bond (Eq. (2.2)).
Fig.
MCI =
√
2
1 ∗ 2
+
√
2
2 ∗ 3
2
+
√
= 2.808 (2.2)
3 ∗ 1
For the same carbon number, the more branches, the smaller the MCI. In general, the
boiling point of a molecule is higher the larger the molecule. For the same molecular
weight, a compound with more branches will have a lower boiling point. This is
because in an unbranched structure, there are many conformational isomers because
of the free rotation between C–C bonds. This means that thermal energy is also
used for internal energy. In contrast, in a branched structure, the rotational barrier is
higher, and the thermal energy is used for translational energy, resulting in boiling
at a lower temperature. For simple saturated hydrocarbons, the boiling point and
molecular weight have the relationship shown in Fig.
2.9a.
This result is obtained even if the number of branches differs and the boiling
point differs, since the molecular weight is constant. In contrast, a plot of MCI
versus boiling point shows a high correlation as shown in Fig.
2.9b.
For saturated hydrocarbon compounds, the MCI can be used to predict the boiling
point, entropy, heat of formation, and heat of evaporation. However, it is difficult to
design a candidate compound that falls within the target boiling point using only the
topological index.
2
1
2
3
1
1
2
2
1
2
2
3
3
3
1
1
Fig. 2.8 Calculation of Molecular Connectivity Index (MCI)

2 Screening Methods for Drugs Using Chemoinformatics Methods … 21
https://t.me/med1917
Fig. 2.9 Relationship between a boiling point and molecular weight, b boiling point and MCI, and
c boiling point and Chi1n
Calculate the Chi1n values of perfluorosaturated alkanes using RDKit. The relationship between the Chi1n value and the boiling point makes a different curve from
that of saturated hydrocarbons, as shown in Fig.
2.9c. Furthermore, the relationship
becomes very complicated when chlorine or bromine is added.
2.3.2.6 Boiling Point Estimation Using RDKit Output
Machine learning with Descriptor Generator output, such as RDKit, is common.
There are some things to be aware of when doing so. The Descriptor Generator
generates many descriptors for a single compound. PCA and PLS methods can be
dimensionally reduced without much concern for multicollinearity. However, we
should not simply use all the explanatory variables calculated by the Descriptor
Generator. It is necessary to examine the value of each Descriptor.
Among the calculated values of RDKit, Kappa3 has some anomalous values as
shown in Fig.
2.10a.
In the TI, a small compound may have no bonds made by the four heavy atoms;
the TI may be zero, but its inverse number would be a very large value. When PCA is
performed with such identifiers included, for example, the t hird and fourth principal
components (Fig.
It is possible to perform principal component regression using such principal
components, but this would lead to a decrease in prediction performance. Since
Kappa3 is not essential, it is reasonable to eliminate that column.
2.10b and c) show anomalous values derived from Kappa3.

22 H. Yamamoto
https://t.me/med1917
Fig. 2.10 Descriptor values of all the DB compounds calculated using Descriptor Generator of
RDKit. a Kappa3. b Third main component of all DB compounds. c Fourth main component of all
DB compounds
2.3.2.7 Boiling Point Estimation Using Variable Selection Multiple
Regression
The Descriptor Generator generates 50–60 descriptors, from which 3–6 descriptors
are selected to create a prediction equation. Although it depends on the number of
data, a modern computer can perform multiple regression calculations for all combinations and select the set of descriptors with the highest coefficient of determination
as the solution. When the number of descriptors exceeds 1000, the exhaustive method
is computationally burdensome, so the genetic algorithm method is used to search
for the optimal solution within the allowable time.
As an example, three of the 39 basic descriptors created by RDKit were selected to
create an equation to predict the boiling point. The variables selected were NumHBD
(number of hydrogen bond donors), log Kow (octanol/water partition ratio), and MR
(molecular refraction).
Compared to the 66 functional group contribution method, higher correlations
were obtained for at most three explanatory variables, as shown in Fig.
2.11a.
The prediction performance is also higher than the functional group contribution
method, as shown in Fig.
2.11b. Predictions can also be obtained for compounds
with undefined functional groups. Compounds with functional groups for which the
coefficients cannot be determined may be calculated to tentatively determine the
coefficients.

2 Screening Methods for Drugs Using Chemoinformatics Methods … 23
https://t.me/med1917
Fig. 2.11 Correlation of boiling point calculations using variable selection multiple regression
methods. a For trained data. b For prediction data
2.3.3 Screening from Boiling Point Estimates
When screening compounds, I include the error amount due to the estimation method
as showninFig.
compounds with experimental boiling points of 20–70 °C. The candidate compounds
are those whose predicted values by the multiple regression method are between 0 °C
and 150 °C.
The number of candidates can be narrowed down by using estimating formulas that
are more accurate than the multiple regression method. However, for the first stage
of screening, the multiple regression method is sufficient. For the 2484 compounds
in the database, 1925 candidates (77%) remain. As shown in Table
the number of carbons, halogens, and oxygen can be narrowed down.
2.12 as a candidate compound. For example, suppose I want to screen
2.3, the limits for
Fig. 2.12 Refinement of
target and calculated boiling
point values
Table 2.3 Refinement by boiling point. Minimum and maximum for each atom
C# H# Br# Cl# F# I# N# O#
Min 1 0 0 0 0 0 0 0
Max 11 18 3 5 22 2 2 3

24 H. Yamamoto
https://t.me/med1917
2.4 Flammability
2.4.1 Flash Point Data
For researchers who do not work with halogenated compounds, flammability is not
much of an issue. With the exception of some phosphorus compounds, all organic
materials burn according to Eq. (
CxHyOzNw +x +
If combustion is viewed as an oxidation reaction, carbon is oxidized to the most
stable form, CO
. Therefore, CO2 does not burn any further, but carbon monoxide
2
(CO) burns one more step according to Eq. (
Similarly, carbon reacts with F2 fluorine to form CF4, which is very stable and
is not oxidized in the presence of oxygen. Therefore, perfluorinated compounds do
not burn. Some nonflammables do not burn as a result of molecular stabilization by
oxidation or fluorination. Apart from that, those containing chlorine, bromine, or
iodine in the molecule may be nonflammable. Combustion can be thought of as a
chain reaction of radical reactions. It is explained that the chain reaction is broken
because chlorine, bromine, and iodine radicals quench the chain reaction.
However, whether it is flammable or not can be very confusing if determined from
physical properties. Experiments to determine flammability are often terminated
when the temperature exceeds 110 °C. This is because the evaluation as a hazardous
material changes when the flash point is above 110 °C. However, “not flammable up
to 110 °C” is sometimes evaluated as not flammable. Sometimes data is lost because
database columns strictly distinguish between numerical and non-numerical values.
The flash point measuring device checks for flammability by emitting sparks from
an ignition source. However, halogen compounds sometimes did not ignite because
the specific gravity of the gas was heavier than air and there was no more oxygen
around the ignition source. There is also t he uncertainty that the gas burned when it
was stirred.
If a flash point exists, it is known to be approximately half the boiling point, as
shown in Fig.
2.13.
Due to this relationship, calculated flash points are sometimes listed for
compounds that are inherently nonflammable.
The logistic regression method is a method to machine-learn presence (1) and
absence (0), such as combustion and mutagenicity. Since it is too complicated to
analyze by logistic regression method, some rough trends will be considered first.
2.3).
z
y
−
+ wO
4
2
→ xCO2 +
2
y
2
H
O + wNO
2
(2.3)
2
2.4).
1
CO +
O
→ CO
2
2
2
(2.4)

2 Screening Methods for Drugs Using Chemoinformatics Methods … 25
https://t.me/med1917
Fig. 2.13 Relationship
between flash point and
boiling point
The 79 compounds in the database were classified as flammable or not depending
on the data source. These compounds were excluded from the analysis and used for
evaluation.
2.4.2 Rough Refinement with Decision Trees
In predicting the activity of a drug design, many studies can be referenced to determine what items to use as explanatory variables. However, in areas where there are
few similar studies, manual work may be required, including the creation of new
indices. Here I define the Halogen Ratio (HR), which is calculated in Eq. (
Halogen Ratio (HR) =
F + Cl + Br + I
C + O + N
2.5).
(2.5)
Consider the general trends for the 671 compounds with flammability determinant
values. First, alcohols, esters, and carboxylic acids are flammable regardless of the
amount of halogen. Next, HR < 1, i.e., combustible if a large amount of hydrogen
remains in the molecule. And HR > 2 makes them nonflammable. The tree structure
is shown in Fig.
A closer look at the flammable compounds with 1 ≤ HR is shown in Fig. 2.15a.
The horizontal axis is the heat of formation divided by the molecular weight. This
value is an indicator of molecular stability.
It can be seen that olefin compounds burn even with compounds with high halogen
substitution ratio of HR > 1.5 or higher.
Compounds that do not burn at HR ≤ 2 are roughly divided into two groups, as
shown in Fig.
2.15b.
Although some actually have more than one type of halogen atom, they are listed in
the order of preference F < Br, I < Cl. Fluorinated compounds stabilize the molecule
per molecular weight and become nonflammable as the fluorine substitution ratio
2.14.

26 H. Yamamoto
https://t.me/med1917
Fig. 2.14 Tree structure of flammability of halogenated compounds
Fig. 2.15 Classification of a flammable compounds with 1 ≤ HR and b nonflammable compounds
with HR ≤ 2
increases. Chlorine, bromine, and iodine do not stabilize the molecule, but act as
radical quenchers and become nonflammable.
In decision tree machine learning, once a table is created with the presence or
absence of combustibility and the explanatory factors that explain it, the decision
tree is automatically created. In the random forest method, several types of decision
trees are created, and conclusions are reached by consensus. If there is a large enough
number of explanatory variables, the answer may be obtained even if the contents
are a black box. However, what kind of explanatory factors can we come up with is
the bottleneck.
•
Anything with a specific functional group (OH, COO, COOH) will always burn.
•
Fluorinated and stabilized (HF/MW < −2.06) do not burn.
•
Chlorine, bromine, and iodine do not burn because they quench radicals.
It still seems difficult to automatically define these things numerically.

2 Screening Methods for Drugs Using Chemoinformatics Methods … 27
https://t.me/med1917
2.4.3 Logistic Regression
I have a qualitative understanding of what items influence the flammability of
halogenated compounds.
HR (Eq. 2.5) evaluates the effect of halogen atoms only from the number of atoms.
However, the difference in atoms cannot be considered equivalent. Also, the effect
of olefin is not quantitatively known.
In logistic regression for flammability determination, multiple regression calculations are first performed. The objective variable is the presence or absence of
flammability (yes:1, no:0), and the explanatory variables are the number of hydrogens, fluorine, chlorine, bromine, iodine, and unsaturated bonds divided by the sum of
the number of carbon, oxygen, and nitrogen atoms. The multiple regression equation
is as in Eq. (
Multiple Regression (MR) = 0.151
The coefficients for hydrogen and olefins are positive, so the more of these,
the more flammable the molecule becomes. Among the halogens, iodine has the
lowest flammability. Plotting the flammability determinant versus flammability on
the horizontal axis (Fig.
I can say that a determination value of 0.8 or higher burns reliably; a value of 0.1
or lower does not burn reliably. However, there will be no judgment as to which of
the cases falls between the two.
The line obtained from the multiple regression calculation is then S-transformed
using the sigmoid function in Eq. (
2.6).
− 0.257
− 0.732
Cl
C + N + O
I
C + N + O
2.16a).
− 0.308
+ 0.455
2.7).
H
C + N + O
C + N + O
Olefine
C + N + O
Br
− 0.313
F
C + N + O
+ 0.732 (2.6)
Fig. 2.16 Flammability determination by a multiple regression method and b logistic regression
method

28 H. Yamamoto
https://t.me/med1917
Y =
1 + EXP(−MR
(
1
))
(2.7)
In this case, the coefficients of multiple regression are searched so that the logarithmic sum of the cross-entropy error of the judgment error after S-shaped transformation is the smallest shown in Fig.
is included in Microsoft Excel [
Logistic Regression = 1.679
− 6.083
− 16.595
Cl
C + N + O
C + N + O
Y'=
1 + EXP(−Logistic Regression
(
2.16b. As a program, you can use Solver, which
].
8
I
C + N + O
− 11.338
+ 8.331
H
C + N + O
C + N + O
− 5.743
Br
Olefine
C + N + O
+ 5.576 (2.8)
1
))
F
(2.9)
As shown in Fig. 2.17a, logistic regression showed that a determination value of
0.75 or higher was flammable (284/285) and a determination value of 0.02 or lower
was not flammable (184/189). No determination was available for 54 compounds.
When HR = (F + Cl + Br + I)/(C + O + N) was used as a criterion, 214 compounds
with 1 ≤ HR ≤2 could no longer be determined. Logistic regression improves greatly.
When collecting data, 79 compounds that were both listed as having a flash point
and not having a flash point were evaluated using the results of logistic regression.
The 30 compounds with a determination value greater than 0.02 and less than or equal
to 0.75 were not determined. However, 12 compounds can be determined under the
condition that they do not burn when the value of heat of formation / MW is less than
−2.06. The remaining 18 compounds cannot be determined, as shown in Fig.
2.17b.
Fig. 2.17 Determination of Flammability. a Results from logistic regression. b Evaluation of
compounds with unknown flammability

2 Screening Methods for Drugs Using Chemoinformatics Methods … 29
https://t.me/med1917
Table 2.4 Atomic number range of candidate compounds
C# H# Br# Cl# F# I# N# O#
Min 0 0 0 0 0 0 0 0
Max 11 15 4 4 22 1 2 3
2.4.4 Screening by Flammability
Of the 2484 compounds in the total database, 1411 compounds were determined to be
flammable. There were 1915 compounds with boiling points of 0–150 °C calculated
by multiple regression from the database. If compounds with a logistic regression
index of 0.75 or higher (flammability) that burn reliably are excluded from them, the
number of candidates is reduced to 881 compounds.
As for the range of atoms, the hydrogen range was greatly reduced as shown in
2.4 due to hydrogen increasing combustion.
Table
2.5 Environmental Assessment
When designing halogen-containing compounds, whether the ODP and GWP are
included in the target is an important indicator. For use in medicines, environmental
assessment may not be necessary because they are not released into the environment
in large quantities. However, if screening is desired, it is more efficient to use multiple
indicators.
2.5.1 Atmospheric Lifetime Estimation
Halogenated compounds, especially chlorofluorocarbon (CFC), have a long atmospheric lifetime. And they emit chlorine radicals in the stratosphere, depleting the
ozone layer. The long atmospheric lifetime also results in high global warming potential. The atmospheric lifetime is plotted against the reaction rate of hydroxyl radicals
as shown in Fig.
It is estimated that hydroxyl radicals account for about 70% of the actual atmospheric lifetime, with hydrolysis and UV degradation accounting for the remaining
30%. However, from a screening point of view, if the logarithmic reaction rate
constant with hydroxyl radicals is −12.5 or higher, the atmospheric lifetime is
within 60 days, then the effects on ODP and GWP are sufficiently small. Therefore, I will estimate the rate constant for the reaction between a halogen compound
and a hydroxyl radical. There are two types of reactions between hydroxyl radicals
and halogen compounds as shown in Fig.
from the molecule to become water and the remaining molecular radical collapse.
2.18. Hydroxyl radicals are abundant in the atmosphere.
2.19. One is the withdrawal of hydrogen
Соседние файлы в папке Библиотека им академика М.И. Перельмана
