Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5580_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
30.08.2026
Размер:
43 Мб
Скачать
11 Electronic-Structure Informatics for Drug Development 205
https://t.me/med1917
on E-ESI. The applicability of this informatics method has been rationalized by showing its relationship with drug activity parameters.
In addition to comments on the plausible ESI methodologies, the detailed expla­nation of the E-ESI descriptors has been presented. The common feature of the suggested descriptors is that each has a background justifiable with theories in chem­istry and physics. Although there is still arbitrariness in the choices of descriptors, the factors discussed in quantum chemical analyses are considered reasonable as descriptors for “explainable” machine learning models.
Further challenges in the electron-level descriptor development for drug discovery would be autonomic definitions of essentially important machine learning descrip­tors. It will be done not through only versatile applications of multi-scale computer simulations, but through the development of emerging methods for data mining. For these approaches, both ideas on practical computation and theoretical methods are demanded.
It is noted that although domain-specific information in ESI is appreciably different from the others, it should be appropriately and skillfully coupled with other data-centric sciences. There are a variety of concepts, practical ideas, compu­tational methods, and techniques developed so far in chemoinformatics, bioinfor­matics, materials informatics, and computer sciences developing machine learning and artificial intelligence (AI). There are no needs to develop independent infor­matics approaches. Ongoing efforts to develop compound methods for ESI are being made in our laboratory, and the results will be published in future.
Acknowledgements The author expresses sincere thanks to Prof. Johann Gasteiger and Prof. Hiroko Satoh for their careful reading and suggestions on the present manuscript.
Technical Note This manuscript was originally written by the present author and has been linguis­tically updated by referring to the suggestions for rephrasing from ChatGPT4 of OpenAI [ Grammarly [ Japanese language, which is the native language of the present author. After machine-suggested modifications, the final draft was completed under the responsibility of the present author. This note is added to indicate that this is the first trial for the author to publish a paper through these man– machine interactions (MMI). MMI would be a key even in sciences and technologies of molecules and materials. Emerging tools and applications of productive MMI are highly expected for (even unexpected) development of functional molecules and materials.
]. Using the DeepL translator [44], the contents have been checked even in the
43
42
]and
References
1. Sugimoto M (2016) Exploring Functional Molecules and Materials Using Electronic-Structure Informatics. CICSJ Bull 34:112–117 (in Japanese).
2. Sugimoto M, Ideo T, Manggara AB, Yoshida K, Inoue T (2019) An Electronic-Structure Infor­matics Study on Inhibitory Activity of Natural Products against Fatty Acid Synthase. J Comp Aided Chem 20:65–75.
3. Ideo T, Yoshida K, Sugimoto M (2021) Regression Modeling and Virtual Screening of Natural Products Exhibiting Antibacterial Activity. An Application of Electronic-structure Informatics Descriptors. Chem Lett 50:849–852.
https://doi.org/10.2751/jcac.20.65
https://doi.org/10.1246/cl.200966
https://doi.org/10.11546/cicsj.34.112
206 M. Sugimoto
https://t.me/med1917
4. Tateishi Y, Sugimoto, M (2023) Searching for α-Glucosidase Inhibitors Using Electronic­Structure Informatics. J Comp Chem Jpn 22:24–17 (in Japanese).
2023-0034
5. Sugimoto M, Manggara AB, Yoshida K, Inoue T, Ideo T (2020) An Electronic-Structure Infor­matics Study on the Toxicity of Alkylphenols to Tetrahymena Pyriformis. Mol Inf 39:1900121.
https://doi.org/10.1002/minf.201900121
6. Manggara AB, Sugimoto M (2021) Extended Regression Modeling of the Toxicity of Phenol Derivatives to Tetrahymena Pyriformis Using the Electronic-Structure Informatics Descriptor. J Comp Aided Chem 22:17–22.
7. Manggara AB, Ohkawa K, Sugimoto M (2021) Classifying Modes of Toxic Action of Molecules with Electronic-Structure Informatics. Application to Imbalanced Toxicity Data of Phenol Derivatives to Tetrahymena Pyriformis. Chem Lett 50:1887–1891.
210453
8. Tateishi Y, Sugimoto M (2023) Electronic-Structure Informatics for Natural Product Drug Discovery: Discovery of α-Glucosidase Inhibitors. Paper presented at the 8th Autumn School of Chemoinformatics, Nara Kasugano International Forum, Nara, 28–30 November 2023
9. Di L, Kerns EH (2016) Drug-Like Properties: Concepts, Structure Design and Methods from ADME to Toxicity Optimization. 2nd edn. Academic Press, Cambridge
10. Hückel E (1931) Quantentheoretische Beiträge zum Benzolproblem. Z Physik 70:204–286.
https://doi.org/10.1007/BF01339530
11. Ashcroft NW, Mermin ND (1976) Solid State Physics. 1st edn. Cengage Learning.
12. Holstein T (1959) Studies of Polaron Motion: Part I. The Molecular-Crystal Model. Ann Phys 8:325–342.
13. Hubbard J (1963) Electron Correlations in Narrow Energy Bands. Proc Royal Soc A276:238–
257.
14. Mototake YI, Mizukami M, Akai I, Okada M (2019) Bayesian Hamiltonian Selection in X­Ray Photoelectron Spectroscopy. J Phys Soc Jpn 88:034004.
034004
15. Fukui K (1981) The Role of Frontier Orbitals in C hemical Reactions. In: The Nobel Lecture. The Nobel Prize Organization. Accessed 16 February 2024
16. Hoffmann R (1981) Building Bridges Between Inorganic and Organic Chemistry. In: The Nobel Lecture. The Nobel Prize Organization.
hoffman-lecture.pdf. Accessed 16 February 2024
17. Cheng Y, Prusoff WH (1973) Relationship between the Inhibition Constant (KI)and the Concentration of Inhibitor Which Causes 50 Per cent Inhibition (I Biochem Pharmacol 22:3099–3108.
18. Pearson RG (1963) Hard and Soft Acids and Bases. J Am Chem Soc. 85:3533–3539. https://
doi.org/10.1021/ja00905a001
19. Mulliken RS (1934) A New Electroaffinity Scale; Together with Data on Valence States and on Valence Ionization Potentials and Electron Affinities. J Chem Phys 2:782–793.
org/10.1063/1.1749394
20. Parr RG, Weitao Y (1995) Density-Functional Theory of Atoms and Molecules. Oxford University Press, Oxford.
21. Marcus RA (1992) Electron Transfer Reactions in Chemistry: Theory and Experiment (Nobel Lecture) Angew Chem Int Ed Engl 32:1111–1222.
22. Jensen F (2017) Introduction to Computational Chemistry. 3rd edn. John Wiley & Sons, Chichester.
23. Kittle C (1996) Introduction to Solid State Physics. 7th edn. John Wiley & Sons, Hoboken.
24. Adamo C, Jacquemin D (2013) The Calculations of Excited-State Properties with Time­Dependent Density Functional Theory. Chem Soc Rev 42:845–856.
C2CS35394F
25. McQuarrie DA, Simon JD (1997) Physical Chemistry: A Molecular Approach. University Science Books, Sausalito.
https://doi.org/10.1016/0003-4916(59)90002-8
https://doi.org/10.1098/rspa.1963.0204
https://doi.org/10.2751/jcac.22.17
https://www.nobelprize.org/uploads/2018/06/fukui-lecture.pdf.
https://www.nobelprize.org/uploads/2018/06/
https://doi.org/10.1016/0006-2952(73)90196-2
https://doi.org/10.1002/anie.199311113
https://doi.org/10.2477/jccj.
https://doi.org/10.1246/cl.
https://doi.org/10.7566/JPSJ.88.
) of an Enzymatic Reaction.
50
https://doi.
https://doi.org/10.1039/
11 Electronic-Structure Informatics for Drug Development 207
https://t.me/med1917
26. Frisch MJ et al. Gaussian 16, Gaussian, Inc., Wallingford CT, 2016.
27. Wolfram Mathemetica. https://www.wolfram.com/mathematica/ Accessed 16 February 2024
28. Tomasi J, Mennucci B, Cammi, R (2005) Quantum Mechanical Continuum Solvation Models. Chem Rev 105:2999–3094.
29. RDkit. https://www.rdkit.org/ Accessed 16 February 2024
30. Moriwaki H, Tian YS, Kawashita N, Takagi T (2018) Mordred: A Molecular Descriptor Calculator. J Cheminf 10:4.
31. https://www.alvascience.com/alvadesc-descriptors/ Accessed on 16 February 2024
32. Lipinski CA, Lombardo F, Dominy BW, Feeney PJ (2001) Experimental and Computational Approaches to Estimate Solubility and Permeability in Drug Discovery and Development Settings. Adv Drug Deliv Rev 46:3–26.
33. GaussView 6. https://gaussian.com/gaussview6/ Accessed 16 February 2024
34. O’Boyle NM, Banck M, James CA, Morley C, Vandermeersch T, Hutchison GR (2011) Open Babel: An Open Chemical Toolbox. J Cheminf 3:33.
35. Zhao Y, Truhlar, DG (2008) The M06 Suite of Density Functionals for Main Group Thermo­chemistry, Thermochemical Kinetics, Noncovalent Interactions, Excited States, and Transition Elements: Two New Functionals and Systematic Testing of Four M06-Class Functionals and 12 Other Functionals. Theor Chem Acc 120:215–241.
0310-x
36. Sugimoto M, Nakatsuji H (1995) Gauge-Invariant Basis Sets for Magnetic Property Calcula­tions. J Chem Phys 102:285–293.
37. Nakatsuji H, Kanda K, Yonezawa, T (1980) Force in Scf Theories. Chem Phys Lett 75:340–346.
https://doi.org/10.1016/0009-2614(80)80527-6
38. Chawla NV, Bowyer KW, Hall LO, Kegelmeyer WP (2002) SMOTE: Synthetic Minority Over-Sampling Technique. J Artif Intell Res 16:321–357.
39. Sawada R, Iwata M, Umezaki M, Usui Y, Kobayashi T, Kubono T, Hayashi S, Kadowaki M, Yamanishi Y (2018) KampoDB, Database of Predicted Targets and Functional Annotations of Natural Medicines. Sci Rep 8:11216.
40. AutoDock Vina. https://vina.scripps.edu/ Accessed 16 February 2024
41. Rogers D, Hahn M (2010) Extended-Connectivity Fingerprints. J Chem Inf Model 50:742–754.
https://doi.org/10.1021/ci100050t
42. ChatGPT. https://chat.openai.com/ Accessed 16 February 2024
43. Grammarly. https://app.grammarly.com/ Accessed 16 February 2024
44. DeePL. https://www.deepl.com/translator Accessed 16 February 2024
https://doi.org/10.1021/cr9904009
https://doi.org/10.1186/s13321-018-0258-y
https://doi.org/10.1016/S0169-409X(00)00129-0
https://doi.org/10.1186/1758-2946-3-33
https://doi.org/10.1007/s00214-007-
https://doi.org/10.1063/1.469401
https://doi.org/10.1613/jair.953
https://doi.org/10.1038/s41598-018-29516-1
Chapter 12
https://t.me/med1917
Data-Driven Chemistry for Developing Organic Synthesis Routes for Functional Chemicals
Kenji Hori, Shohei Majima, and Toru Yamaguchi
12.1 Introduction
In recent years, drug and material design have been using chemoinformatics to create new compounds with medicinal or useful properties [ synthetic routes for these compounds is conducted in the following order: (i) synthetic chemists create several synthesis routes based on their experience and intuition, (ii) these routes are verified and obtained the target through experiments involving trial and error, (iii) the reaction conditions are optimized on the basis of a plausible reaction mechanism.
There were also many attempts to create synthetic routes using computers. Corey has developed LHASA, an organic synthesis route design system (SRDS), along with the concept of the retrosynthesis [ progress through the efforts of Gasteiger in Germany [ and Funatsu in Japan [ ered that computers could not handle a variety of compounds and that only synthetic chemists could solve the problem. For synthesis targets required multi-step reactions,
6, 7]. Although these enthusiastic efforts, it was often consid-
3]. Since the 1980s, SRDSs made remarkable
1, 2]. The development of
4], Jorgensen in the U.S. [5],
K. Hori ( R&D Center of Functional Materials, Transition State Technology Co. Ltd, Ube 755-0097, Japan e-mail: kenji@tstcl.jp
Faculty of Engineering, Yamaguchi University, Ube 755-8611, Japan
National Institute of Advanced Industrial Science and Technology, Tsukuba 305-8560, Japan
S. Majima Core Technology Department, Shionogi Pharma Co., Ltd, Amagasaki 660-0813, Japan e-mail: shohei.majima@shionogi.co.jp
T. Yamaguchi Division of Computational Chemistry, Transition State Technology Co. Ltd, Ube 755-0097, Japan e-mail: tor@tstcl.jp
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2024 H. Satoh et al. (eds.), Drug Development Supported by Informatics,
https://doi.org/10.1007/978-981-97-4828-0_12
B
)
209
210 K. Hori et al.
https://t.me/med1917
+
+ +
Database
(i) Creation of New Synthesis routes using AIPHOS/TOSP
Fig. 12.1 Overview of data-driven synthesis route development
(ii) Digital verification using transition state database (iii) Ordering the routes
+
(iv) High throughput verification of synthesis routes
SRDS routes often diverge since they do not take into account their synthetic possi­bility. Therefore, the software has not always been accepted by synthetic organic chemists.
It has been recently developed SYNTHIA™ [8] on the basis of many literature data and AiZyntFinder [
9], ASKCOS [10] using machine learning of big data (AI-
SRDS). The possibility of the target synthesis relies only on literature descriptions so that there are no guarantees whether or not the target is obtained. Therefore, trial-and­error approaches are required. This is not much different from the old fashion route developments. Recently, synthetic robots will help to reduce experimental efforts
11].
[
While there have been many attempts to use computers for developing synthesis routes, their potentials have to still be determined by synthetic chemists. We have proposed to change this situation by introducing a method that combines chemoin­formatics and computational chemistry. This method evaluates synthesis routes of target compounds in the order shown in Fig.
12.1.
i. Creation of new synthesis routes using AIPHOS/TOSP [12]. TOSP creates
synthesis routes using transforms, i.e., knowledge-based Information of forma­tion and/or cleavage positions of bonds, types of substituents, etc., involved in name reactions. Therefore, we think that AIPHOS/TOSP has a potential to create synthetic routes that have not been considered before.
ii. Evaluation of synthesis routes by quantum chemical (QC) calculations utilizing
the database (TSDB/QMRDB [ verification is called the digital screening [
iii. Adoption of a small number of routes with high potential for synthesis and their
ranking on the basis of availabilities of reagents and/or experimental easiness, according to the results of the digital screenings.
iv. Experiments to confirm whether or not the target compound can be synthesized
according to the order of synthetic routes ranked.
This procedure, called the data-driven synthetic route development, can signif­icantly shorten the period for synthesis route developments. The introducing data
13], see below) that we have developed. This
14].
12 Data-Driven Chemistry for Developing Organic Synthesis Routes … 211
https://t.me/med1917
science and theoretical chemistry into the synthesis route development has the poten­tial to fundamentally change the way of organic synthetic chemistry. The new proce­dure improves the weaknesses of the traditional method, which requires a lot of trial-and-error experiments. In this paper, we describe the details of (ii) and some results of research on (iv) in a NEDO project named “Development of Synthesis Process Design Technology” launched in 2022, joining one of the NEDO national projects on flow synthesis [ for developing functional chemicals using process informatics, which organically combines computational chemistry, chemoinformatics, and experimental chemistry.
15]. The project aims to significantly shorten the period
12.2 Digital Screenings Using QC Calculations
12.2.1 TOSP as the Standard SRDS of the NEDO Project
As TOSP creates synthesis routes using knowledge from name reactions, it does not depend on big data from chemical journals. The NEDO project is developing AIst-syn (Unpublished result), another SRDS that shares the same concept as TOSP. The mechanism of the synthesis route from TOSP is easily analyzed using QMRDB/ TSDB contents. However, TOSP also cannot escape from the same route divergence problem as other SRDSs, either. In order to use SRDS for the synthesis route devel­opment, it is necessary to evaluate their synthetic possibility and to eliminate the divergence problem.
Computational chemistry is a useful tool for understanding known reactions and has elucidated reaction mechanisms in detail, including the transition state (TS) structures. This property implies that QC calculations can confirm whether the target can be synthesized by new routes that have not previously been examined. This property can be successfully exploited to develop new synthetic routes for target compounds by combining SRDS and QC calculations.
12.2.2 Effects of Digital Screenings
The digital screening is effective in avoiding the divergence problem since it verifies synthetic routes in the reverse order of experiments. In order to explain the reason, let us consider an example shown in Fig.
Obtaining Precursors 1 and 2 are essential to ascertain whether or not Route D produces the target compound. Precursor 1 may be synthesized using one of Routes D1a-D1c and Precursor 2 using Routes D2a-D2e. It may not be necessary to try them all. For the attempts to synthesize the target using other routes such as Route A, we have to synthesize the precursor needed for the route. The effort for the route is wasted if the route does not provide the target.
12.2.
212 K. Hori et al.
https://t.me/med1917
Route A1
Route A2
Route A
Route B
Target
Route C
Route D
+
Precursor 2
Fig. 12.2 Effect of digital screening on synthesis route evaluations
Route A3
Precursor 1
Route D1a
Route D1b
Route D1c
Route D2a
Route D2b
Route D2c
Route D2d
Route D2e
Verification of synthesis routes using the digital screening is performed in the reverse order of the experiments, i.e., at the beginning, Route A~D are evaluated. If the digital screening resulted in judging that Routes A–C do not work but Route D is capable of synthesizing the target compound, there is no need to validate routes A1–A3. The digital screenings result in significantly reducing the number of QC calculations.
If the digital screenings evaluate that Route D1a synthesizes Precursor 1 and Routes D2a and D2e produce Precursor 2, we can use the synthesis routes indicated in red. Therefore, it is likely that only four experiments of Routes D1a, D2a, D2e, and D are enough to synthesize the target.
As demonstrated in this example, the digital screening reduces the number of actual experimental efforts to be performed very much. In addition, the calculated activation free energy allows to roughly set the reaction temperature. The calculation of solvent effects provides insight into the solvent to be used in experiments. The
16, 17
NEDO project employed the QM/MC/FEP method [
], which can estimate
solvent effects with a high degree of accuracy.
In the data-driven synthesis route development, retrievals of TSDB/QMRDB are not used for predictions but for making initials of new calculations for proposed synthesis routes. The digital screenings have to be always performed to obtain new TSs for evaluating possibilities of synthesis routes. This is very different from predic­tions the AI-SRDSs make through models of machine learning with big data. There­fore, the digital screening predicts the possibility for synthesizing the target with high probabilities.
12 Data-Driven Chemistry for Developing Organic Synthesis Routes … 213
https://t.me/med1917
12.3 Development of TSDB/QMRDB and Technologies
for Digital Screenings
12.3.1 TS Motif, the Key to Reducing Computational Times
In organic chemistry, reactions proceeding under the same reaction mechanism, for example, Diels–Alder reaction, Ene reaction, and so on, are grouped together as name reactions. Our experience in analyzing many of those reactions revealed that TS structures of the same name reaction are similar to each other even though substituents of substrates differ greatly.
As an example, a reaction is discussed in which an amide is formed through the condensation reaction of a carboxylic acid and an amine. This reaction forming an amide consists of two elementary reactions, one is the reaction of the condensation reagent and two molecules of carboxylic acid to form an intermediate, and the other is amide formation through the reaction of the intermediate.
The lower left in Fig. 12.3 displays the TS structure (TS_S) of the first elemen­tary reaction with simple substituents (R1 = R3 = CH3), where it takes a six­membered ring consisting of two oxygen, carbon, nitrogen, and two hydrogen atoms.
18
The intrinsic reaction coordinate (IRC) [ relay is involved in this reaction mechanism. The O1-H and C-O2 distances in this structure were calculated to be 1.698 and 2.281 Å, respectively. The six-membered ring in TS_S is retained in TS_C with complex substituents (R1 = 2-CF = C
). The corresponding distances (1.803 and 2.096 Å) of TS_S are not signif-
6H11
icantly different from those in TS_C. The characteristic geometry, in which bond formation and/or dissociation occur is called a “TS motif”. This example indicates the similarity of the TS motifs. In some cases, TS motifs of different name reactions are similar if they proceed under a similar reaction mechanism.
] calculations indicated that the proton
,R3
3C6H5
Fig. 12.3 Similarity of TS motifs
214 K. Hori et al.
https://t.me/med1917
12.3.2 How to Optimize a New TS Structure Using the TS
Motif
The purpose of the digital screenings is to verify new synthesis routes for the target to be feasible. However, it has to be emphasized that the possibility of synthesis routes does not determine on the basis of TSDB/QMRD database searches but assess through new TS calculations and the activation barrier heights. Therefore, we adopted a new method which utilize a TS motif to create an initial structure for optimizing the TS for a given synthesis route. The present method is called the TS motif method, which follows the steps below.
i. The substituents are given at the corresponding positions in the TS motif. In the
example in Fig. reagent are substituted with R1 = 2-CF initial structure for the TS optimization.
ii. The TS optimization is performed using the initial structure with the fixed TS
motif within the red circle (the partial geometry optimization), followed by the TS optimization without fixed parameters. The optimized TS structure is used for performing IRC calculations, followed
iii.
by structure optimizations of reactants and products. This will be confirmed that the obtained TS connects the target compound with the reactant.
We have to complete these calculations within a few hours to one day, the period that experimental chemists without patience can tolerate. As will be discussed later, we are constructing QMRDB, a database gathering TS motifs. We developed a computer cloud system handling the database (Fig. search, specifically performing the TS motif method. This program creates inputs for the Gaussian program [ manages the conformational analysis using the Conflex program [ preliminary data of reaction analyses to QMRDB. We made a web-based manual describing how to use the system and the TS motif method.
Even an organic chemist without experiences in computational chemistry learned how to use the system within a week. Furthermore, he has completed digital screen­ings for the four-step synthesis routes of a drug for only two months as will be given later. Once he is proficient in the system, a similar digital screening could be completed in less than two weeks. This means that synthetic organic chemists can complete the digital screenings of reactions for the target before s tarting experiments.
12.3, the methyl groups of the simple acid and the condensation
and R3 = C6H
3C6H5
12.4) and a program, named TS
19
], submits jobs to and downloads from the cloud system,
to create the
11
20], and registers
12.3.3 Problems in the TS Motif Method
The digital screening is a tool to compare whether one synthetic route is superior to others. The magnitude of activation free energies is one of the factors to determine which reaction is optimal. Therefore, in data-driven synthetic route development, it is
12 Data-Driven Chemistry for Developing Organic Synthesis Routes … 215
https://t.me/med1917
WindowsTerminal
Web browser
Search resemble reactions
i Structure
Fig. 12.4 TSDB cloud system
Reaction Search
TS Coordinates
TS Motif Method
Conformation search
TSDB Cloud system
TSDB
QMRDB
necessary to find the most stable TS structure in the reaction concerned. However, the TS motif method is unlikely to locate the most stable TS conformation for compounds with many substituents [
21].
The simple TS motif of butadiene + ethylene was applied to optimize the TS for the reaction of N-methylpenta-2,4-dienamide + acrolein, and the resulting TS_ 0 is showninFig.
12.5. IRC calculations confirmed t hat the TS connects the reac-
tant with the product, 6-formyl-N-methylcyclohex-2-ene-1-carboxamide. The other conformers are derived from TS_0. TS_2, 3, 4, and 5 were calculated to be less stable by 0.3 to 5.7 kcal/mol than TS_1, the most stable and more stable conformation by
3.2 kcal/mol than TS_0.
Another example is a TS conformation analysis of a relatively large transition metal complex in Scheme shown in Fig.
12.6, this is a large calculation consisting of 555 basis functions. Even
12.1 [22]. Since the molecule has seven benzene rings
for such a large molecule, the calculations including TS search and its conformation
TS_0 TS_1 TS_2 TS_3 TS_4 TS_5
0.0 -3.2 -2.9 0.6 0.3 2.5
Fig. 12.5 Results of TS conformational analysis for Diels–Alder reaction
Gaussian 09 program B3LYP/6-31(d)