Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5338_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
30.08.2026
Размер:
43 Мб
Скачать
7 Language Models in Molecular Discovery 133
https://t.me/med1917
same, which in turn imposes a constraint on the rate of our technological advance­ments. Over the last few years, this conventional approach has been challenged by LLMs. It has been found that scaling up LLMs leads to astonishing performances in few-shot [ dation models” [ perform multiple tasks despite being trained on one large dataset. Essentially, this multi-task learning is achieved by prompting LLMs with task instructions along with the actual query text which has been found to induce exceptional performance in natural language inference and sentence completion [ kicked off new research directions, such as prompt engineering [ learning [
The foundation model paradigm also finds an increasing adoption in chem­istry. There is an increase in task-specific models integrating natural and chemical languages [ advancing through models that combine tasks such as property prediction, reac­tion prediction, and molecule generation either with small task-specific heads (e.g., T5Chem [ dellis et al. [ multitask chemical and natural language model. Despite only 250M parameters, the Multitask Text and Chemistry T5 was shown to outperform ChatGPT [
101
tica [ (natural text
94] and even zero-shot task generalization [95]. Referred to as “foun-
96, 97], these models, with typically billions of parameters, can
95]. These findings have
98] and in-context
94
], in NLP.
59, 81, 82, 99]. Concurrently, multi-tasking in pure CLMs has also been
100]) or via mask infilling (e.g., Regression Transformer [42]). Christofi-
23] were the first to bridge the gap and develop a fully prompt-based
91
] and Galac-
] on a contrived discovery workflow for re-discovering a common herbicide
new molecule synthesis route synthesis execution protocol).
7.4.2 The Coalescence of Chatbots with Chemistry Tools
Given the aforementioned strong task generalization performances of LLMs, building chatbot interfaces around it was a natural next step and thus next to ChatGPT [ many similar tools were launched. Such tools were found to perform well on simplistic chemistry tasks [ interact with chemical data, enabling intuitive access to complex concepts and make valuable suggestions for diverse chemical tasks. Furthermore, AI models specifically developed by computer scientists, e.g., for drug discovery or material science, can be made available through applications powered by LLMs, such as chatbots. This minimizes the access barrier for subject matter experts who would otherwise require the respective programming skills to utilize these AI models. The power of such chat­bots is reached through the coalescence of LLMs and existing chemistry software tools like PubChem [ can unleash the full potential and value of these models by the strongly enhanced usage. An example of how the interaction with such a tool could look like is shown
7.4.
in Fig.
In this example, a user provides a molecule (either as a SMILES string or via a molecule sketcher) and asks to identify the molecule. The chatbot relies on prompt­engineering in order to inform the LLM about all its available tools. The user input
103, 104
102
], RDKit [85], or GT4SD [61]. Together, such applications
], creating potential to reshape how chemists
91],
134 N. Janakarajan et al.
https://t.me/med1917
Fig. 7.4 Screenshot of the LLM-powered chatbot application ChemChat. Embedding the capabil- ities of existing resources such as PubChem [ to execute programming routines in the background and, thus, to answer highly subject-matter specific user requests without the user needing programming skills and without being prone to hallucinations
102], RDKit [85]orGT4SD [61] enables the assistant
is first sent to the LLM which recognizes that one of its supported tools, in this case PubChem, can answer the question. The chatbot then sends a request to the PubChem API and returns a concise description of the molecule. The user subsequently asks to compute various physicochemical properties, including, e.g., the logP partition coefficient [
105] and the drug-likeness (QED) [106
erties is enabled through the GT4SD tool [
61] allowing the chatbot to answer the
]. Calculation of all those prop-
request with certainty. This will trigger a programming routine to accurately format the API request for GT4SD, i.e., composing the SMILES string with the logP or QED endpoint. The computation is then performed asynchronously and a separate call to the post-processing routine formats the LLM-generated string reply and composes the response object for the frontend. This fusion of LLMs with existing tools gives rise to a chatbot assistant for material science and data visualization that can perform simple programming routines without requiring the user to know programming or have access to compute resources.
A conversation involving more complex user queries is shown in Fig. 7.5.After having the initial molecule as an alkaloid, the user requests three similar molecules with a slightly increased logP of 0.5. Here, ChemChat identifies the Regression Transformer [
42] as the available tool to perform substructure-constrained, property-
driven molecule design. Once the routine has been executed and the three candidate SMILES are collected, the text result is post-processed to add more response data
7 Language Models in Molecular Discovery 135
https://t.me/med1917
Fig. 7.5 Screenshot of ChemChat during a molecular design task executed through GT4SD’s Regression Transformer [
42]aswellasproperty[107] and similarity calculation [108, 109]
objects such as molecule visualizations, datasets, or Vega Lite specs for interac­tive visualizations. The user then asks for synthetic accessibility of one of the just­generated compounds and then lets ChemChat compute the Tanimoto similarity of two candidate molecules.
Moreover, for expert knowledge-specific tasks ChemChat also applies Retrieval Augmented Generation (RAG) which has emerged as a solution to LLM-prominent hallucinations, non-transparent reasoning, and outdated or unavailable information. RAG incorporates external databases, identifies and retrieves relevant information to the user input (in part by involving semantic search), and provides it to the context window of the LLM to augment its knowledge. The LLM is then instructed to generate the response to the user input often without considering other inherent knowledge to ensure highest relevance and specificity. External resources accessible by ChemChat include IBM CIRCA (
https://circa.res.ibm.com), a research plat-
form designed for chemistry, biology, and materials enabling information retrieval from about 28 million patents and, thus, offering an excellent resource for domain­specific knowledge (Fig.
7.6). Besides incorporating IBM CIRCA, ChemChat is
also connected to IBM RXN (Sect. 7.3.3) allowing the user to request and analyze forward reaction predictions and modeling of retrosynthetic pathways through the chat interface.
In conclusion, chatbots can facilitate the integration of essentially all major chemo-informatics software in a harmonized and seamless manner. While LLMs are not intrinsically capable of performing complex routines, at least not yet precisely and in a trustworthy manner, the synergy between their natural language abilities
136 N. Janakarajan et al.
https://t.me/med1917
Fig. 7.6 Screenshot of ChemChat during a domain expertise-intensive task for which a RAG process is applied to utilize expert knowledge stored in IBM CIRCA, a research platform and database for chemistry, biology, and materials holding about 28 million patents. Relevant infor­mation is gathered via a multistep process involving semantic search and provided to the context window of the LLM. By respective prompt engineering, the LLM is restricted to not use any other resources to generate a response. Links to the respective reference patents further guide to the IBM CIRCA web interface for additional detailed analysis
with existing chemistry tools has the potential to transform the way chemistry is performed.
Acknowledgements: This work is supported by the EU project Fragment-Screen, grant agreement ID: 101094131.
References
1. OpenAI (2023) Gpt-4 technical report. 2303.08774
2. Wouters OJ, McKee M, Luyten J (2020) Estimated research and development investment needed to bring a new medicine to market, 2009-2018. Jama 323(9):844–853
3. Scannell JW, Blanckley A, Boldon H, Warrington B (2012) Diagnosing the decline in pharmaceutical R&D efficiency. Nat Rev Drug Discov 11(3):191–200
4. Polishchuk PG, Madzhidov TI, Varnek A (2013) Estimation of the size of drug-like chemical space based on gdb-17 data. J Comput Aid Mol Des 27(8):675–679
5. Hargrave-Thomas E, Yu B, Reynisson J (2012) Serendipity in anticancer drug discovery. World Journal of Clinical Oncology 3(1):1
6. Lin Z, Akin H, Rao R, Hie B, Zhu Z, Lu W, Smetanin N, Verkuil R, Kabeli O, Shmueli Y, et al (2023) Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379(6637):1123–1130
7 Language Models in Molecular Discovery 137
https://t.me/med1917
7. Zhavoronkov A, Ivanenkov YA, Aliper A, Veselov MS, Aladinskiy VA, Aladinskaya AV, Terentiev VA, Polykovskiy DA, Kuznetsov MD, Asadulaev A, et al (2019) Deep learning enables rapid identification of potent ddr1 kinase inhibitors. Nat Biotechnol 37(9):1038–1040
8. Das P, Sercu T, Wadhawan K, Padhi I, Gehrmann S, Cipcigan F, Chenthamarakshan V, Strobelt H, Santos CD, Chen PY, et al (2021) Accelerated antimicrobial discovery via deep generative models and molecular dynamics simulations. Nat Biomed Eng 5(6):613–623
9. Park NH, Manica M, Born J, Hedrick JL, Erdmann T, Zubarev DY, Adell-Mill N, Arrechea PL (2023) Artificial intelligence driven design of catalysts and materials for ring opening polymerization using a domain-specific language. Nature Communications 14(1):3686
10. Mikolov T, Chen K, Corrado G, Dean J (2013) Efficient estimation of word representations in vector space. arXiv preprint
11. Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Ł, Polosukhin I (2017) Attention is all you need. Advances in neural information processing systems 30
12. Weininger D (1988) Smiles, a chemical language and information system. 1. introduction to methodology and encoding rules. J Chem Inf Comp Sci 28(1):31–36
13. Gómez-Bombarelli R, Wei JN, Duvenaud D, Hernández-Lobato JM, Sánchez-Lengeling B, Sheberla D, Aguilera-Iparraguirre J, Hirzel TD, Adams RP, Aspuru-Guzik A (2018) Auto­matic chemical design using a data-driven continuous representation of molecules. ACS Central Science 4(2):268–276
14. Grisoni F (2023) Chemical language models for de novo drug design: Challenges and opportunities. Current Opinion in Structural Biology 79:102527
15. Bjerrum EJ (2017) Smiles enumeration as data augmentation for neural network modeling of molecules. arXiv preprint
16. Tetko IV, Karpov P, Bruno E, Kimber TB, Godin G (2019) Augmentation is what you need! In: International Conference on Artificial Neural Networks, Springer, pp 831–835
17. Li X, Fourches D (2020) Inductive transfer learning for molecular activity prediction: Next­gen qsar models with molpmofit. Journal of Cheminformatics 12(1):1–15
18. Arús-Pous J, Johansson SV, Prykhodko O, Bjerrum EJ, Tyrchan C, Reymond JL, Chen H, Engkvist O (2019) Randomized smiles strings improve the quality of molecular generative models. Journal of Cheminformatics 11(1):1–13
19. van Deursen R, Ertl P, Tetko IV, Godin G (2020) Gen: highly efficient smiles explorer using autodidactic generative examination networks. Journal of Cheminformatics 12(1):1–14
20. Schwaller P, Gaudin T, Lanyi D, Bekas C, Laino T (2018) “Found in translation”: predicting outcomes of complex organic chemistry reactions using neural sequence-to-sequence models. Chemical Science 9(28):6091–6098
21. Ucak UV, Ashyrmamatov I, Lee J (2023) Improving the quality of chemical language model outcomes with atom-in-smiles tokenization. Journal of Cheminformatics 15(1):55
22. Li X, Fourches D (2021) Smiles pair encoding: a data-driven substructure tokenization algorithm for deep learning. Journal of Chemical Information and Modeling 61(4):1560–1569
23. Christofidellis D, Giannone G, Born J, Winther O, Laino T, Manica M (2023) Unifying molecular and textual representations via multi-task language modelling. In: International Conference on Machine Learning
24. Krenn M, Häse F, Nigam A, Friederich P, Aspuru-Guzik A (2020) Self-referencing embedded strings (selfies): A 100% robust molecular string representation. Machine Learning: Science and Technology 1(4):045024
25. Heller SR, McNaught A, Pletnev I, Stein S, Tchekhovskoi D (2015) InChl, the IUPAC international chemical identifier. Journal of Cheminformatics 7(1):1–34
26. Handsel J, Matthews B, Knight NJ, Coles SJ (2021) Translating the InChl: adapting neural machine translation to predict iupac names from a chemical identifier. Journal of Cheminformatics 13(1):1–11
27. Born J, Manica M (2021) Trends in deep learning for property-driven drug design. Current Medicinal Chemistry 28(38):7862–7886
28. Segler MH, Kogej T, Tyrchan C, Waller MP (2018) Generating focused molecule libraries for drug discovery with recurrent neural networks. ACS Central Science 4(1):120–131
arXiv:1301.3781
arXiv:1703.07076
138 N. Janakarajan et al.
https://t.me/med1917
29. Flam-Shepherd D, Zhu K, Aspuru-Guzik A (2022) Language models can learn complex molecular distributions. Nature Communications 13(1):3293
30. Polykovskiy D, Zhebrak A, Sanchez-Lengeling B, Golovanov S, Tatanov O, Belyaev S, Kurbanov R, Artamonov A, Aladinskiy V, Veselov M, et al (2020) Molecular sets (moses): a benchmarking platform for molecular generation models. Front Pharmacol 11:1931
31. Joulin A, Mikolov T (2015) Inferring algorithmic patterns with stack-augmented recurrent nets. Advances in Neural Information Processing Systems 28
32. Popova M, Isayev O, Tropsha A (2018) Deep reinforcement learning for de novo drug design. Science Advances 4(7):eaap7885
33. Schilter O, Vaucher A, Schwaller P, Laino T (2023) Designing catalysts with deep generative models and computational data. a case study for Suzuki cross coupling reactions. Digital Discovery 2(3):728–735
34. Lim J, Ryu S, Kim JW, Kim WY (2018) Molecular generative model based on conditional variational autoencoder for de novo molecular design. Journal of Cheminformatics 10(1):1–9
35. Born J, Manica M, Oskooei A, Cadow J, Markert G, Martínez MR (2021) PaccMannRL:De novo generation of hit-like anticancer molecules from transcriptomic data via reinforcement learning. iScience 24(4):102269
36. Born J, Manica M, Cadow J, Markert G, Mill NA, Filipavicius M, Janakarajan N, Cardinale A, Laino T, Martínez MR (2021) Data-driven molecular design for discovery and synthesis of novel ligands: a case study on sars-cov-2. Mach Learn: Sci Technol 2(2):025024
37. Born J, Huynh T, Stroobants A, Cornell WD, Manica M (2021) Active site sequence represen­tations of human kinases outperform full sequence representations for affinity prediction and inhibitor generation: 3d effects in a 1d model. Journal of Chemical Information and Modeling 62(2):240–257
38. Janakarajan N, Born J, Manica M (2022) A fully differentiable set autoencoder. In: Proceed­ings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pp 3061–3071
39. Radford A, Narasimhan K, Salimans T, Sutskever I, et al (2018) Improving language understanding by generative pre-training
40. Bagal V, Aggarwal R, Vinod P, Priyakumar UD (2021) Molgpt: molecular generation using a transformer-decoder model. Journal of Chemical Information and Modeling 62(9):2064–2076
41. Mazuz E, Shtar G, Shapira B, Rokach L (2023) Molecule generation using transformers and policy gradient reinforcement learning. Scientific Reports 13(1):8799
42. Born J, Manica M (2023) Regression transformer enables concurrent sequence regression and generation for molecular language modelling. Nature Machine Intelligence 5(4):432–444
43. Wu Z, Ramsundar B, Feinberg EN, Gomes J, Geniesse C, Pappu AS, Leswing K, Pande V (2018) Moleculenet: a benchmark for molecular machine learning. Chemical Science 9(2):513–530
44. Born J, Markert G, Janakarajan N, Kimber TB, Volkamer A, Martínez MR, Manica M (2023) Chemical representation learning for toxicity prediction. Digital Discovery
45. Bahdanau D, Cho K, Bengio Y (2014) Neural machine translation by jointly learning to align and translate. arXiv preprint https://arxivorg/abs/14090473, arXiv1409.0473
46. Fabian B, Edlich T, Gaspar H, Segler M, Meyers J, Fiscato M, Ahmed M (2020) Molecular representation learning with language models and domain-relevant auxiliary tasks. arXiv preprint
47. Chithrananda S, Grand G, Ramsundar B (2020) Chemberta: large-scale self-supervised
48. Ross J, Belgodere B, Chenthamarakshan V, Padhi I, Mroueh Y, Das P (2022) Large-
49. Maziarka L, Danel T, Mucha S, Rataj K, Tabor J, Jastrzkebski S (2019) Molecule-augmented
arXiv:2011.13230
pretraining for molecular property prediction. arXiv preprint
scale chemical language representations capture molecular structure and properties. Nature Machine Intelligence 4(12):1256–1264
attention transformer. In: Workshop on Graph Representation Learning, Neural Information Processing Systems
arXiv:2010.09885
7 Language Models in Molecular Discovery 139
https://t.me/med1917
50. Maziarka L, Majchrowski D, Danel T, Gainski P, Tabor J, Podolak I, Morkisz P, Jastrzkebski S (2024) Relative molecule self-attention transformer. Journal of Cheminformatics 16(1):3
51. Ovchinnikova K, Born J, Chouvardas P, Rapsomaniki M, Kruithof-de Julio M (2024) Over­coming limitations in current measures of drug response may enable AI-driven precision oncology Abstract npj Precision Oncology 8(1).
52. Born J, Shoshan Y, Huynh T, Cornell WD, Martin EJ, Manica M (2022) On the choice of active site sequences for kinase-ligand affinity prediction. Journal of Chemical Information and Modeling 62(18):4295–4299.
53. Gezelter JD (2015) Open source and open data should be standard practices
54. Wolf T, Debut L, Sanh V, Chaumond J, Delangue C, Moi A, Cistac P, Rault T, Louf R, Funtowicz M, et al (2020) Transformers: State-of-the-art natural language processing. In: Proceedings of the 2020 conference on empirical methods in natural language processing: system demonstrations, pp 38–45
55. Dosovitskiy A, Beyer L, Kolesnikov A, Weissenborn D, Zhai X, Unterthiner T, Dehghani M, Minderer M, Heigold G, Gelly S, Uszkoreit J, Houlsby N (2021) An image is worth 16x16 words: Transformers for image recognition at scale. In: 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3–7, 2021
56. Chen L, Lu K, Rajeswaran A, Lee K, Grover A, Laskin M, Abbeel P, Srinivas A, Mordatch I (2021) Decision transformer: Reinforcement learning via sequence modeling. Advances in Neural Information Processing Systems 34:15084–15097
57. Jumper J, Evans R, Pritzel A, Green T, Figurnov M, Ronneberger O, Tunyasuvunakool K, Bates R, Žídek A, Potapenko A, et al (2021) Highly accurate protein structure prediction with alphafold. Nature 596(7873):583–589
58. Schwaller P, Vaucher AC, Laplaza R, Bunne C, Krause A, Corminboeuf C, Laino T (2022) Machine intelligence for chemical reaction space. Wiley Interdisciplinary Reviews: Computational Molecular Science 12(5):e1604
59. Edwards C, Lai T, Ros K, Honke G, Cho K, Ji H (2022) Translation between molecules and natural language. In: 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022
60. Elnaggar A, Heinzinger M, Dallago C, Rehawi G, Yu W, Jones L, Gibbs T, Feher T, Angerer C, Steinegger M, Bhowmik D, Rost B (2021) Prottrans: Towards cracking the language of life’s code through self-supervised deep learning and high-performance computing. IEEE Transactions on Pattern Analysis and Machine Intelligence pp 1–1,
TPAMI.2021.3095381
61. Manica M, Born J, Cadow J, Christofidellis D, Dave A, Clarke D, Teukam YGN, Giannone G, Hoffman SC, Buchan M, et al (2023) Accelerating material design with the generative toolkit for scientific discovery. npj Computational Materials 9(1):69
62. Huang K, Fu T, Gao W, Zhao Y, Roohani Y, Leskovec J, W CC, Xiao C, Sun J, Zitnik M (2021) Therapeutics data commons: Machine learning datasets and tasks for drug discovery and development. Advances in Neural Information Processing System 35
63. Ramsundar B, Eastman P, Walters P, Pande V, Leswing K, Wu Z (2019) Deep Learning for the Life Sciences. O’Reilly Media,
Microscopy/dp/1492039837
64. von Platen P, Patil S, Lozhkov A, Cuenca P, Lambert N, Rasul K, Davaadorj M, Wolf T (2022) Diffusers: State-of-the-art diffusion models. Accessed: February 2, 2024
65. Zhu Z, Shi C, Zhang Z, Liu S, Xu M, Yuan X, Zhang Y, Chen J, Cai H, Lu J, et al (2022) Torchdrug: A powerful and flexible machine learning platform for drug discovery. arXiv Preprint at
66. Brown N, Fiscato M, Segler MH, Vaucher AC (2019) Guacamol: benchmarking models for de novo molecular design. J Chem Inf Model 59(3):1096–1108
67. Bengio Y, Lahlou S, Deleu T, Hu EJ, Tiwari M, Bengio E (2023) Gflownet foundations. Journal of Machine Learning Research 24(210):1–55
arXiv:2202.08320
https://doi.org/10.1021/acs.jcim.2c00840
https://www.amazon.com/Deep-Learning-Life-Sciences-
https://doi.org/10.1038/s41698-024-00583-0
https://doi.org/10.1109/
https://github.com/huggingface/diffusers
140 N. Janakarajan et al.
https://t.me/med1917
68. Maziarz K, Jackson-Flux H, Cameron P, Sirockin F, Schneider N, Stiefl N, Segler M, Brockschmidt M (2022) Learning to extend molecular scaffolds with structural motif. In: The Tenth International Conference on Learning Representations, ICLR
69. Abid A, Abdalla A, Abid A, Khan D, Alfozan A, Zou J (2019) Gradio: Hassle-free sharing and testing of ml models in the wild. arXiv preprint https://arxivorg/abs/190602569 arXiv1906.02569
70. for Chemistry team IR (2023) rxn4chemistry: Python wrapper for the IBM RXN for Chemistry API.
71. Schwaller P, Laino T, Gaudin T, Bolgar P, Hunter CA, Bekas C, Lee AA (2019) Molecular
72. Pesciullesi G, Schwaller P, Laino T, Reymond JL (2020) Transfer learning enables the
73. Toniato A, Schwaller P, Cardinale A, Geluykens J, Laino T (2021) Unassisted noise reduction
74. Schwaller P, Petraglia R, Zullo V, Nair VH, Haeuselmann RA, Pisoni R, Bekas C, Iuliano
75. Zipoli F, Baldassari C, Manica M, Born J, Laino T (2024) Growing strings in a chemical
76. Probst D, Manica M, Nana Teukam YG, Castrogiovanni A, Paratore F, Laino T (2022)
77. Thakkar A, Vaucher AC, Byekwaso A, Schwaller P, Toniato A, Laino T (2023) Unbiasing
78. Devlin J, Chang MW, Lee K, Toutanova K (2018) Bert: Pre-training of deep bidirectional
79. Schwaller P, Probst D, Vaucher AC, Nair VH, Kreutter D, Laino T, Reymond JL (2021)
80. Schwaller P, Hoover B, Reymond JL, Strobelt H, Laino T (2021) Extraction of organic
81. Vaucher AC, Zipoli F, Geluykens J, Nair VH, Schwaller P, Laino T (2020) Automated extrac-
82. Vaucher AC, Schwaller P, Geluykens J, Nair VH, Iuliano A, Laino T (2021) Inferring
83. Genheden S, Thakkar A, Chadimová V, Reymond JL, Engkvist O, Bjerrum E (2020) Aizyn-
84. Gainski P, Maziarka L, Danel T, Jastrzebski S (2022) Huggingmolecules: An open-source
85. Landrum G (2013) Rdkit documentation. Release 1(1–79):4
86. Lin TS, Coley CW, Mochigase H, Beech HK, Wang W, Wang Z, Woods E, Craig SL,
87. Born J, Shoshan Y, Huynh T, Cornell WD, Martin EJ, Manica M (2022) On the choice of
https://github.com/rxn4chemistry/rxn4chemistry, accessed: February 2, 2024
transformer: a model for uncertainty-calibrated chemical reaction prediction. ACS Central Science 5(9):1572–1583
molecular transformer to predict regio-and stereoselective reactions on carbohydrates. Nature Communications 11(1):4874
of chemical reaction datasets. Nature Machine Intelligence 3(6):485–494
A, Laino T (2020) Predicting retrosynthetic pathways using transformer-based models and a hyper-graph exploration strategy. Chemical Science 11(12):3316–3325
reaction space for searching retrosynthesis pathways Abstract npj Computational Materials 10(1).
https://doi.org/10.1038/s41524-024-01290-x
Biocatalysed synthesis planning using data-driven learning. Nature Communications 13(1):964
retrosynthesis language models with disconnection prompts. ACS Central Science
transformers for language understanding. arXiv preprint
Mapping the space of chemical reactions using attention-based neural networks. Nature Machine Intelligence 3(2):144–152
chemistry grammar from unsupervised learning of chemical reactions. Science Advances 7(15):eabe4166
tion of chemical synthesis actions from experimental procedures. Nature Communications 11(1):3601
experimental procedures from text-based representations of chemical reactions. Nature Communications 12(1):2573
thfinder: a fast, robust and flexible open-source software for retrosynthetic planning. Journal of Cheminformatics 12(1):70
library for transformer-based molecular property prediction (student abstract). In: Proceedings of the AAAI Conference on Artificial Intelligence, vol 36, pp 12949–12950
Johnson JA, Kalow JA, et al (2019) Bigsmiles: a structurally-based line notation for describing macromolecules. ACS Central Science 5(9):1523–1531
active site sequences for kinase-ligand affinity prediction. Journal of Chemical Information and Modeling 62(18):4295–4299
arXiv:1810.04805
7 Language Models in Molecular Discovery 141
https://t.me/med1917
88. Heyndrickx W, Mervin L, Morawietz T, Sturm N, Friedrich L, Zalewski A, Pentina A, Humbeck L, Oldenhof M, Niwayama R, et al (2022) Melloddy: cross pharma federated learning at unprecedented scale unlocks benefits in QSAR without compromising proprietary information
89. Gorgulla C, Boeszoermenyi A, Wang ZF, Fischer PD, Coote PW, Padmanabha Das KM, Malets YS, Radchenko DS, Moroz YS, Scott DA, et al (2020) An open-source drug discovery platform enables ultra-large virtual screens. Nature 580(7805):663–668
90. Ivanenkov YA, Polykovskiy D, Bezrukov D, Zagribelnyy B, Aladinskiy V, Kamya P, Aliper A, Ren F, Zhavoronkov A (2023) Chemistry42: an AI-driven platform for molecular design and optimization. Journal of Chemical Information and Modeling 63(3):695–701
91. OpenAI (2023) Chatgpt. https://chat.openai.com/chat, accessed: August 8, 2023
92. GitHub (2024) Github copilot
93. Christiano PF, Leike J, Brown T, Martic M, Legg S, Amodei D (2017) Deep reinforcement learning from human preferences. Advances in Neural Information Processing Systems 30
94. Brown T, Mann B, Ryder N, Subbiah M, Kaplan JD, Dhariwal P, Neelakantan A, Shyam P, Sastry G, Askell A, et al (2020) Language models are few-shot learners. Advances in Neural Information Processing Systems 33:1877–1901
95. Sanh V, Webson A, Raffel C, Bach SH, Sutawika L, Alyafeai Z, Chaffin A, Stiegler A, Le Scao T, Raja A, et al (2022) Multitask prompted training enables zero-shot task generalization. In: ICLR 2022-Tenth International Conference on Learning Representations
96. Fei N, Lu Z, Gao Y, Yang G, Huo Y, Wen J, Lu H, Song R, Gao X, Xiang T, et al (2022) Towards artificial general intelligence via a multimodal foundation model. Nature Communications 13(1):3094
97. Moor M, Banerjee O, Abad ZSH, Krumholz HM, Leskovec J, Topol EJ, Rajpurkar P (2023) Foundation models for generalist medical artificial intelligence. Nature 616(7956):259–265
98. Wei J, Wang X, Schuurmans D, Bosma M, Xia F, Chi E, Le QV, Zhou D, et al (2022) Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems 35:24824–24837
99. Zeng Z, Yao Y, Liu Z, Sun M (2022) A deep-learning system bridging molecule struc­ture and biomedical text with comprehension comparable to human professionals. Nature Communications 13(1):862
100. Lu J, Zhang Y (2022) Unified deep learning model for multitask reaction predictions with explanation. Journal of Chemical Information and Modeling 62(6):1376–1387
101. Taylor R, Kardas M, Cucurull G, Scialom T, Hartshorn A, Saravia E, Poulton A, Kerkez V, Stojnic R (2022) Galactica: A large language model for science. arXiv Preprint at
09085
102. Kim S, Chen J, Cheng T, Gindulyte A, He J, He S, Li Q, Shoemaker BA, Thiessen PA, Yu B, et al (2019) Pubchem 2019 update: improved access to chemical data. Nucleic Acids Research 47(D1):D1102–D1109
103. White AD, Hocky GM, Gandhi HA, Ansari M, Cox S, Wellawatte GP, Sasmal S, Yang Z, Liu K, Singh Y, et al (2023) Assessment of chemistry knowledge in large language models that generate code. Digital Discovery 2(2):368–376
104. Castro Nascimento CM, Pimentel AS (2023) Do large language models understand chemistry? a conversation with chatgpt. Journal of Chemical Information and Modeling 63(6):1649–1655
105. Wildman SA, Crippen GM (1999) Prediction of physicochemical parameters by atomic contributions. Journal of Chemical Information and Computer Sciences 39(5):868–873
106. Bickerton GR, Paolini GV, Besnard J, Muresan S, Hopkins AL (2012) Quantifying the chemical beauty of drugs. Nat Chem 4(2):90–98
107. Ertl P, Schuffenhauer A (2009) Estimation of synthetic accessibility score of drug-like molecules based on molecular complexity and fragment contributions. Journal of Chem­informatics 1:1–11
108. Tanimoto TT (1957) Ibm internal report. Nov 17:1957
109. Rogers D, Hahn M (2010) Extended-connectivity fingerprints. Journal of Chemical Informa­tion and Modeling 50(5):742–754
arXiv:2211.
Chapter 8
https://t.me/med1917
Transformers and Large Language Models for Chemistry and Drug Discovery
Andres M. Bran and Philippe Schwaller
8.1 Introduction
The capacity to process and accurately model human language has been a persistent pursuit within the machine learning community [ intrinsic to human reasoning capabilities, thus successful language modeling could open the door to numerous applications, enhancing various information processing tasks with the potential to revolutionize several industries [ language processing has witnessed significant advancements in recent years, thanks to improved computing infrastructure, breakthroughs in algorithms, and the prolif­eration of abundant and accessible data [ also play a pivotal role in the domain of chemistry, which serves as the fundamental basis for drug discovery and development. Analogous to human language, under­standing and accurately modeling the language of chemistry is crucial for effective research and development in the pharmaceutical industry. By applying the advance­ments in language modeling and processing from the machine learning community to the domain of chemistry, it is possible to facilitate drug discovery by efficiently analyzing and interpreting vast amounts of chemical data and literature.
Introduced in 2017, the Transformer architecture revolutionized natural language processing [ neural machine translation. In its original architecture, the Transformer consists of an encoder, which encodes a sentence in the source language (e.g., French) and a
5]. The Transformer is a type of neural network initially developed for
15]. The belief is that language is
6, 7]. The field of natural
]. Language and technical terminology
8
A. M. Bran · P. S chwa l l er (B) Laboratory of Artificial Chemical Intelligence (LIAC), ISIC, EPFL, Lausanne, Switzerland e-mail: philippe.schwaller@epfl.ch
National Centre of Competence in Research (NCCR) Catalysis, Ecole Polytechnique Fédérale de Lausanne (EPFL), Lausanne, Switzerland
A. M. Bran e-mail:
andres.marulandabran@epfl.ch
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2024 H. Satoh et al. (eds.), Drug Development Supported by Informatics,
https://doi.org/10.1007/978-981-97-4828-0_8
143