Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5858_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
30.08.2026
Размер:
43 Мб
Скачать
20 H. Yamamoto
https://t.me/med1917
As nonlinear analysis tools have advanced over time, this devising of input values will have an even greater effect.
2.3.2.5 Boiling Point Estimation by Topological Index
If molecules are represented by the number of functional groups, the boiling points of compounds with undefined functional groups cannot be predicted. Therefore, explanatory factors such as the topological index (TI) calculated by RDKit and the molecular orbital method are used as data input. The original idea of TI is probably
7
the Molecular Connectivity Index (MCI) [
]. Assign the number of heavy atoms to
be added to the carbon for each carbon–carbon bond in the molecule as shown in
2.8. Multiply this number by 0.5 and take the sum of each bond (Eq. (2.2)).
Fig.
MCI =
2
1 2
+
2
2 3
2
+
= 2.808 (2.2)
3 1
For the same carbon number, the more branches, the smaller the MCI. In general, the boiling point of a molecule is higher the larger the molecule. For the same molecular weight, a compound with more branches will have a lower boiling point. This is because in an unbranched structure, there are many conformational isomers because of the free rotation between C–C bonds. This means that thermal energy is also used for internal energy. In contrast, in a branched structure, the rotational barrier is higher, and the thermal energy is used for translational energy, resulting in boiling at a lower temperature. For simple saturated hydrocarbons, the boiling point and molecular weight have the relationship shown in Fig.
2.9a.
This result is obtained even if the number of branches differs and the boiling point differs, since the molecular weight is constant. In contrast, a plot of MCI versus boiling point shows a high correlation as shown in Fig.
2.9b.
For saturated hydrocarbon compounds, the MCI can be used to predict the boiling point, entropy, heat of formation, and heat of evaporation. However, it is difficult to design a candidate compound that falls within the target boiling point using only the topological index.
2
1
2
3
1
1
2
2
1
2
2
3
3
3
1
1
Fig. 2.8 Calculation of Molecular Connectivity Index (MCI)
2 Screening Methods for Drugs Using Chemoinformatics Methods … 21
https://t.me/med1917
Fig. 2.9 Relationship between a boiling point and molecular weight, b boiling point and MCI, and c boiling point and Chi1n
Calculate the Chi1n values of perfluorosaturated alkanes using RDKit. The rela­tionship between the Chi1n value and the boiling point makes a different curve from that of saturated hydrocarbons, as shown in Fig.
2.9c. Furthermore, the relationship
becomes very complicated when chlorine or bromine is added.
2.3.2.6 Boiling Point Estimation Using RDKit Output
Machine learning with Descriptor Generator output, such as RDKit, is common. There are some things to be aware of when doing so. The Descriptor Generator generates many descriptors for a single compound. PCA and PLS methods can be dimensionally reduced without much concern for multicollinearity. However, we should not simply use all the explanatory variables calculated by the Descriptor Generator. It is necessary to examine the value of each Descriptor.
Among the calculated values of RDKit, Kappa3 has some anomalous values as shown in Fig.
2.10a.
In the TI, a small compound may have no bonds made by the four heavy atoms; the TI may be zero, but its inverse number would be a very large value. When PCA is performed with such identifiers included, for example, the t hird and fourth principal components (Fig.
It is possible to perform principal component regression using such principal components, but this would lead to a decrease in prediction performance. Since Kappa3 is not essential, it is reasonable to eliminate that column.
2.10b and c) show anomalous values derived from Kappa3.
22 H. Yamamoto
https://t.me/med1917
Fig. 2.10 Descriptor values of all the DB compounds calculated using Descriptor Generator of RDKit. a Kappa3. b Third main component of all DB compounds. c Fourth main component of all DB compounds
2.3.2.7 Boiling Point Estimation Using Variable Selection Multiple
Regression
The Descriptor Generator generates 50–60 descriptors, from which 3–6 descriptors are selected to create a prediction equation. Although it depends on the number of data, a modern computer can perform multiple regression calculations for all combi­nations and select the set of descriptors with the highest coefficient of determination as the solution. When the number of descriptors exceeds 1000, the exhaustive method is computationally burdensome, so the genetic algorithm method is used to search for the optimal solution within the allowable time.
As an example, three of the 39 basic descriptors created by RDKit were selected to create an equation to predict the boiling point. The variables selected were NumHBD (number of hydrogen bond donors), log Kow (octanol/water partition ratio), and MR (molecular refraction).
Compared to the 66 functional group contribution method, higher correlations were obtained for at most three explanatory variables, as shown in Fig.
2.11a.
The prediction performance is also higher than the functional group contribution method, as shown in Fig.
2.11b. Predictions can also be obtained for compounds
with undefined functional groups. Compounds with functional groups for which the coefficients cannot be determined may be calculated to tentatively determine the coefficients.
2 Screening Methods for Drugs Using Chemoinformatics Methods … 23
https://t.me/med1917
Fig. 2.11 Correlation of boiling point calculations using variable selection multiple regression methods. a For trained data. b For prediction data
2.3.3 Screening from Boiling Point Estimates
When screening compounds, I include the error amount due to the estimation method as showninFig. compounds with experimental boiling points of 20–70 °C. The candidate compounds are those whose predicted values by the multiple regression method are between 0 °C and 150 °C.
The number of candidates can be narrowed down by using estimating formulas that are more accurate than the multiple regression method. However, for the first stage of screening, the multiple regression method is sufficient. For the 2484 compounds in the database, 1925 candidates (77%) remain. As shown in Table the number of carbons, halogens, and oxygen can be narrowed down.
2.12 as a candidate compound. For example, suppose I want to screen
2.3, the limits for
Fig. 2.12 Refinement of target and calculated boiling point values
Table 2.3 Refinement by boiling point. Minimum and maximum for each atom
C# H# Br# Cl# F# I# N# O#
Min 1 0 0 0 0 0 0 0
Max 11 18 3 5 22 2 2 3
24 H. Yamamoto
https://t.me/med1917
2.4 Flammability
2.4.1 Flash Point Data
For researchers who do not work with halogenated compounds, flammability is not much of an issue. With the exception of some phosphorus compounds, all organic materials burn according to Eq. (
CxHyOzNw +x +
If combustion is viewed as an oxidation reaction, carbon is oxidized to the most stable form, CO
. Therefore, CO2 does not burn any further, but carbon monoxide
2
(CO) burns one more step according to Eq. (
Similarly, carbon reacts with F2 fluorine to form CF4, which is very stable and is not oxidized in the presence of oxygen. Therefore, perfluorinated compounds do not burn. Some nonflammables do not burn as a result of molecular stabilization by oxidation or fluorination. Apart from that, those containing chlorine, bromine, or iodine in the molecule may be nonflammable. Combustion can be thought of as a chain reaction of radical reactions. It is explained that the chain reaction is broken because chlorine, bromine, and iodine radicals quench the chain reaction.
However, whether it is flammable or not can be very confusing if determined from physical properties. Experiments to determine flammability are often terminated when the temperature exceeds 110 °C. This is because the evaluation as a hazardous material changes when the flash point is above 110 °C. However, “not flammable up to 110 °C” is sometimes evaluated as not flammable. Sometimes data is lost because database columns strictly distinguish between numerical and non-numerical values. The flash point measuring device checks for flammability by emitting sparks from an ignition source. However, halogen compounds sometimes did not ignite because the specific gravity of the gas was heavier than air and there was no more oxygen around the ignition source. There is also t he uncertainty that the gas burned when it was stirred.
If a flash point exists, it is known to be approximately half the boiling point, as shown in Fig.
2.13.
Due to this relationship, calculated flash points are sometimes listed for compounds that are inherently nonflammable.
The logistic regression method is a method to machine-learn presence (1) and absence (0), such as combustion and mutagenicity. Since it is too complicated to analyze by logistic regression method, some rough trends will be considered first.
2.3).
z
y
+ wO
4
2
xCO2 +
2
y
2
H
O + wNO
2
(2.3)
2
2.4).
1
CO +
O
CO
2
2
2
(2.4)
2 Screening Methods for Drugs Using Chemoinformatics Methods … 25
https://t.me/med1917
Fig. 2.13 Relationship between flash point and boiling point
The 79 compounds in the database were classified as flammable or not depending on the data source. These compounds were excluded from the analysis and used for evaluation.
2.4.2 Rough Refinement with Decision Trees
In predicting the activity of a drug design, many studies can be referenced to deter­mine what items to use as explanatory variables. However, in areas where there are few similar studies, manual work may be required, including the creation of new indices. Here I define the Halogen Ratio (HR), which is calculated in Eq. (
Halogen Ratio (HR) =
F + Cl + Br + I
C + O + N
2.5).
(2.5)
Consider the general trends for the 671 compounds with flammability determinant values. First, alcohols, esters, and carboxylic acids are flammable regardless of the amount of halogen. Next, HR < 1, i.e., combustible if a large amount of hydrogen remains in the molecule. And HR > 2 makes them nonflammable. The tree structure is shown in Fig.
A closer look at the flammable compounds with 1 HR is shown in Fig. 2.15a. The horizontal axis is the heat of formation divided by the molecular weight. This value is an indicator of molecular stability.
It can be seen that olefin compounds burn even with compounds with high halogen substitution ratio of HR > 1.5 or higher.
Compounds that do not burn at HR 2 are roughly divided into two groups, as shown in Fig.
2.15b.
Although some actually have more than one type of halogen atom, they are listed in the order of preference F < Br, I < Cl. Fluorinated compounds stabilize the molecule per molecular weight and become nonflammable as the fluorine substitution ratio
2.14.
26 H. Yamamoto
https://t.me/med1917
Fig. 2.14 Tree structure of flammability of halogenated compounds
Fig. 2.15 Classification of a flammable compounds with 1 HR and b nonflammable compounds
with HR 2
increases. Chlorine, bromine, and iodine do not stabilize the molecule, but act as radical quenchers and become nonflammable.
In decision tree machine learning, once a table is created with the presence or absence of combustibility and the explanatory factors that explain it, the decision tree is automatically created. In the random forest method, several types of decision trees are created, and conclusions are reached by consensus. If there is a large enough number of explanatory variables, the answer may be obtained even if the contents are a black box. However, what kind of explanatory factors can we come up with is the bottleneck.
Anything with a specific functional group (OH, COO, COOH) will always burn.
Fluorinated and stabilized (HF/MW < 2.06) do not burn.
Chlorine, bromine, and iodine do not burn because they quench radicals.
It still seems difficult to automatically define these things numerically.
2 Screening Methods for Drugs Using Chemoinformatics Methods … 27
https://t.me/med1917
2.4.3 Logistic Regression
I have a qualitative understanding of what items influence the flammability of halogenated compounds.
HR (Eq. 2.5) evaluates the effect of halogen atoms only from the number of atoms. However, the difference in atoms cannot be considered equivalent. Also, the effect of olefin is not quantitatively known.
In logistic regression for flammability determination, multiple regression calcu­lations are first performed. The objective variable is the presence or absence of flammability (yes:1, no:0), and the explanatory variables are the number of hydro­gens, fluorine, chlorine, bromine, iodine, and unsaturated bonds divided by the sum of the number of carbon, oxygen, and nitrogen atoms. The multiple regression equation is as in Eq. (
Multiple Regression (MR) = 0.151
The coefficients for hydrogen and olefins are positive, so the more of these, the more flammable the molecule becomes. Among the halogens, iodine has the lowest flammability. Plotting the flammability determinant versus flammability on the horizontal axis (Fig.
I can say that a determination value of 0.8 or higher burns reliably; a value of 0.1 or lower does not burn reliably. However, there will be no judgment as to which of the cases falls between the two.
The line obtained from the multiple regression calculation is then S-transformed using the sigmoid function in Eq. (
2.6).
0.257
0.732
Cl
C + N + O
I
C + N + O
2.16a).
0.308
+ 0.455
2.7).
H
C + N + O
C + N + O
Olefine
C + N + O
Br
0.313
F
C + N + O
+ 0.732 (2.6)
Fig. 2.16 Flammability determination by a multiple regression method and b logistic regression method
28 H. Yamamoto
https://t.me/med1917
Y =
1 + EXP(−MR
(
1
))
(2.7)
In this case, the coefficients of multiple regression are searched so that the loga­rithmic sum of the cross-entropy error of the judgment error after S-shaped transfor­mation is the smallest shown in Fig. is included in Microsoft Excel [
Logistic Regression = 1.679
6.083
16.595
Cl
C + N + O
C + N + O
Y'=
1 + EXP(−Logistic Regression
(
2.16b. As a program, you can use Solver, which
].
8
I
C + N + O
11.338
+ 8.331
H
C + N + O
C + N + O
5.743
Br
Olefine
C + N + O
+ 5.576 (2.8)
1
))
F
(2.9)
As shown in Fig. 2.17a, logistic regression showed that a determination value of
0.75 or higher was flammable (284/285) and a determination value of 0.02 or lower was not flammable (184/189). No determination was available for 54 compounds. When HR = (F + Cl + Br + I)/(C + O + N) was used as a criterion, 214 compounds with 1 ≤ HR 2 could no longer be determined. Logistic regression improves greatly.
When collecting data, 79 compounds that were both listed as having a flash point and not having a flash point were evaluated using the results of logistic regression. The 30 compounds with a determination value greater than 0.02 and less than or equal to 0.75 were not determined. However, 12 compounds can be determined under the condition that they do not burn when the value of heat of formation / MW is less than
2.06. The remaining 18 compounds cannot be determined, as shown in Fig.
2.17b.
Fig. 2.17 Determination of Flammability. a Results from logistic regression. b Evaluation of compounds with unknown flammability
2 Screening Methods for Drugs Using Chemoinformatics Methods … 29
https://t.me/med1917
Table 2.4 Atomic number range of candidate compounds
C# H# Br# Cl# F# I# N# O#
Min 0 0 0 0 0 0 0 0
Max 11 15 4 4 22 1 2 3
2.4.4 Screening by Flammability
Of the 2484 compounds in the total database, 1411 compounds were determined to be flammable. There were 1915 compounds with boiling points of 0–150 °C calculated by multiple regression from the database. If compounds with a logistic regression index of 0.75 or higher (flammability) that burn reliably are excluded from them, the number of candidates is reduced to 881 compounds.
As for the range of atoms, the hydrogen range was greatly reduced as shown in
2.4 due to hydrogen increasing combustion.
Table
2.5 Environmental Assessment
When designing halogen-containing compounds, whether the ODP and GWP are included in the target is an important indicator. For use in medicines, environmental assessment may not be necessary because they are not released into the environment in large quantities. However, if screening is desired, it is more efficient to use multiple indicators.
2.5.1 Atmospheric Lifetime Estimation
Halogenated compounds, especially chlorofluorocarbon (CFC), have a long atmo­spheric lifetime. And they emit chlorine radicals in the stratosphere, depleting the ozone layer. The long atmospheric lifetime also results in high global warming poten­tial. The atmospheric lifetime is plotted against the reaction rate of hydroxyl radicals as shown in Fig.
It is estimated that hydroxyl radicals account for about 70% of the actual atmo­spheric lifetime, with hydrolysis and UV degradation accounting for the remaining 30%. However, from a screening point of view, if the logarithmic reaction rate constant with hydroxyl radicals is 12.5 or higher, the atmospheric lifetime is within 60 days, then the effects on ODP and GWP are sufficiently small. There­fore, I will estimate the rate constant for the reaction between a halogen compound and a hydroxyl radical. There are two types of reactions between hydroxyl radicals and halogen compounds as shown in Fig. from the molecule to become water and the remaining molecular radical collapse.
2.18. Hydroxyl radicals are abundant in the atmosphere.
2.19. One is the withdrawal of hydrogen