Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5338_Библиотеки_им_академика_М_И_Перельмана
.pdf
7 Language Models in Molecular Discovery 133
https://t.me/med1917
same, which in turn imposes a constraint on the rate of our technological advancements. Over the last few years, this conventional approach has been challenged by
LLMs. It has been found that scaling up LLMs leads to astonishing performances
in few-shot [
dation models” [
perform multiple tasks despite being trained on one large dataset. Essentially, this
multi-task learning is achieved by prompting LLMs with task instructions along
with the actual query text which has been found to induce exceptional performance
in natural language inference and sentence completion [
kicked off new research directions, such as prompt engineering [
learning [
The foundation model paradigm also finds an increasing adoption in chemistry. There is an increase in task-specific models integrating natural and chemical
languages [
advancing through models that combine tasks such as property prediction, reaction prediction, and molecule generation either with small task-specific heads (e.g.,
T5Chem [
dellis et al. [
multitask chemical and natural language model. Despite only 250M parameters, the
Multitask Text and Chemistry T5 was shown to outperform ChatGPT [
101
tica [
(natural text
94] and even zero-shot task generalization [95]. Referred to as “foun-
96, 97], these models, with typically billions of parameters, can
95]. These findings have
98] and in-context
94
], in NLP.
59, 81, 82, 99]. Concurrently, multi-tasking in pure CLMs has also been
100]) or via mask infilling (e.g., Regression Transformer [42]). Christofi-
23] were the first to bridge the gap and develop a fully prompt-based
91
] and Galac-
] on a contrived discovery workflow for re-discovering a common herbicide
→ new molecule → synthesis route → synthesis execution protocol).
7.4.2 The Coalescence of Chatbots with Chemistry Tools
Given the aforementioned strong task generalization performances of LLMs, building
chatbot interfaces around it was a natural next step and thus next to ChatGPT [
many similar tools were launched. Such tools were found to perform well on
simplistic chemistry tasks [
interact with chemical data, enabling intuitive access to complex concepts and make
valuable suggestions for diverse chemical tasks. Furthermore, AI models specifically
developed by computer scientists, e.g., for drug discovery or material science, can
be made available through applications powered by LLMs, such as chatbots. This
minimizes the access barrier for subject matter experts who would otherwise require
the respective programming skills to utilize these AI models. The power of such chatbots is reached through the coalescence of LLMs and existing chemistry software
tools like PubChem [
can unleash the full potential and value of these models by the strongly enhanced
usage. An example of how the interaction with such a tool could look like is shown
7.4.
in Fig.
In this example, a user provides a molecule (either as a SMILES string or via a
molecule sketcher) and asks to identify the molecule. The chatbot relies on promptengineering in order to inform the LLM about all its available tools. The user input
103, 104
102
], RDKit [85], or GT4SD [61]. Together, such applications
], creating potential to reshape how chemists
91],

134 N. Janakarajan et al.
https://t.me/med1917
Fig. 7.4 Screenshot of the LLM-powered chatbot application ChemChat. Embedding the capabil-
ities of existing resources such as PubChem [
to execute programming routines in the background and, thus, to answer highly subject-matter
specific user requests without the user needing programming skills and without being prone to
hallucinations
102], RDKit [85]orGT4SD [61] enables the assistant
is first sent to the LLM which recognizes that one of its supported tools, in this case
PubChem, can answer the question. The chatbot then sends a request to the PubChem
API and returns a concise description of the molecule. The user subsequently asks
to compute various physicochemical properties, including, e.g., the logP partition
coefficient [
105] and the drug-likeness (QED) [106
erties is enabled through the GT4SD tool [
61] allowing the chatbot to answer the
]. Calculation of all those prop-
request with certainty. This will trigger a programming routine to accurately format
the API request for GT4SD, i.e., composing the SMILES string with the logP or QED
endpoint. The computation is then performed asynchronously and a separate call to
the post-processing routine formats the LLM-generated string reply and composes
the response object for the frontend. This fusion of LLMs with existing tools gives
rise to a chatbot assistant for material science and data visualization that can perform
simple programming routines without requiring the user to know programming or
have access to compute resources.
A conversation involving more complex user queries is shown in Fig. 7.5.After
having the initial molecule as an alkaloid, the user requests three similar molecules
with a slightly increased logP of −0.5. Here, ChemChat identifies the Regression
Transformer [
42] as the available tool to perform substructure-constrained, property-
driven molecule design. Once the routine has been executed and the three candidate
SMILES are collected, the text result is post-processed to add more response data

7 Language Models in Molecular Discovery 135
https://t.me/med1917
Fig. 7.5 Screenshot of ChemChat during a molecular design task executed through GT4SD’s
Regression Transformer [
42]aswellasproperty[107] and similarity calculation [108, 109]
objects such as molecule visualizations, datasets, or Vega Lite specs for interactive visualizations. The user then asks for synthetic accessibility of one of the justgenerated compounds and then lets ChemChat compute the Tanimoto similarity of
two candidate molecules.
Moreover, for expert knowledge-specific tasks ChemChat also applies Retrieval
Augmented Generation (RAG) which has emerged as a solution to LLM-prominent
hallucinations, non-transparent reasoning, and outdated or unavailable information.
RAG incorporates external databases, identifies and retrieves relevant information to
the user input (in part by involving semantic search), and provides it to the context
window of the LLM to augment its knowledge. The LLM is then instructed to
generate the response to the user input often without considering other inherent
knowledge to ensure highest relevance and specificity. External resources accessible
by ChemChat include IBM CIRCA (
https://circa.res.ibm.com), a research plat-
form designed for chemistry, biology, and materials enabling information retrieval
from about 28 million patents and, thus, offering an excellent resource for domainspecific knowledge (Fig.
7.6). Besides incorporating IBM CIRCA, ChemChat is
also connected to IBM RXN (Sect. 7.3.3) allowing the user to request and analyze
forward reaction predictions and modeling of retrosynthetic pathways through the
chat interface.
In conclusion, chatbots can facilitate the integration of essentially all major
chemo-informatics software in a harmonized and seamless manner. While LLMs are
not intrinsically capable of performing complex routines, at least not yet precisely
and in a trustworthy manner, the synergy between their natural language abilities

136 N. Janakarajan et al.
https://t.me/med1917
Fig. 7.6 Screenshot of ChemChat during a domain expertise-intensive task for which a RAG
process is applied to utilize expert knowledge stored in IBM CIRCA, a research platform and
database for chemistry, biology, and materials holding about 28 million patents. Relevant information is gathered via a multistep process involving semantic search and provided to the context
window of the LLM. By respective prompt engineering, the LLM is restricted to not use any other
resources to generate a response. Links to the respective reference patents further guide to the IBM
CIRCA web interface for additional detailed analysis
with existing chemistry tools has the potential to transform the way chemistry is
performed.
Acknowledgements: This work is supported by the EU project Fragment-Screen, grant agreement
ID: 101094131.
References
1. OpenAI (2023) Gpt-4 technical report. 2303.08774
2. Wouters OJ, McKee M, Luyten J (2020) Estimated research and development investment
needed to bring a new medicine to market, 2009-2018. Jama 323(9):844–853
3. Scannell JW, Blanckley A, Boldon H, Warrington B (2012) Diagnosing the decline in
pharmaceutical R&D efficiency. Nat Rev Drug Discov 11(3):191–200
4. Polishchuk PG, Madzhidov TI, Varnek A (2013) Estimation of the size of drug-like chemical
space based on gdb-17 data. J Comput Aid Mol Des 27(8):675–679
5. Hargrave-Thomas E, Yu B, Reynisson J (2012) Serendipity in anticancer drug discovery.
World Journal of Clinical Oncology 3(1):1
6. Lin Z, Akin H, Rao R, Hie B, Zhu Z, Lu W, Smetanin N, Verkuil R, Kabeli O, Shmueli Y,
et al (2023) Evolutionary-scale prediction of atomic-level protein structure with a language
model. Science 379(6637):1123–1130

7 Language Models in Molecular Discovery 137
https://t.me/med1917
7. Zhavoronkov A, Ivanenkov YA, Aliper A, Veselov MS, Aladinskiy VA, Aladinskaya AV,
Terentiev VA, Polykovskiy DA, Kuznetsov MD, Asadulaev A, et al (2019) Deep learning
enables rapid identification of potent ddr1 kinase inhibitors. Nat Biotechnol 37(9):1038–1040
8. Das P, Sercu T, Wadhawan K, Padhi I, Gehrmann S, Cipcigan F, Chenthamarakshan V, Strobelt
H, Santos CD, Chen PY, et al (2021) Accelerated antimicrobial discovery via deep generative
models and molecular dynamics simulations. Nat Biomed Eng 5(6):613–623
9. Park NH, Manica M, Born J, Hedrick JL, Erdmann T, Zubarev DY, Adell-Mill N, Arrechea
PL (2023) Artificial intelligence driven design of catalysts and materials for ring opening
polymerization using a domain-specific language. Nature Communications 14(1):3686
10. Mikolov T, Chen K, Corrado G, Dean J (2013) Efficient estimation of word representations
in vector space. arXiv preprint
11. Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Ł, Polosukhin I
(2017) Attention is all you need. Advances in neural information processing systems 30
12. Weininger D (1988) Smiles, a chemical language and information system. 1. introduction to
methodology and encoding rules. J Chem Inf Comp Sci 28(1):31–36
13. Gómez-Bombarelli R, Wei JN, Duvenaud D, Hernández-Lobato JM, Sánchez-Lengeling B,
Sheberla D, Aguilera-Iparraguirre J, Hirzel TD, Adams RP, Aspuru-Guzik A (2018) Automatic chemical design using a data-driven continuous representation of molecules. ACS
Central Science 4(2):268–276
14. Grisoni F (2023) Chemical language models for de novo drug design: Challenges and
opportunities. Current Opinion in Structural Biology 79:102527
15. Bjerrum EJ (2017) Smiles enumeration as data augmentation for neural network modeling of
molecules. arXiv preprint
16. Tetko IV, Karpov P, Bruno E, Kimber TB, Godin G (2019) Augmentation is what you need!
In: International Conference on Artificial Neural Networks, Springer, pp 831–835
17. Li X, Fourches D (2020) Inductive transfer learning for molecular activity prediction: Nextgen qsar models with molpmofit. Journal of Cheminformatics 12(1):1–15
18. Arús-Pous J, Johansson SV, Prykhodko O, Bjerrum EJ, Tyrchan C, Reymond JL, Chen H,
Engkvist O (2019) Randomized smiles strings improve the quality of molecular generative
models. Journal of Cheminformatics 11(1):1–13
19. van Deursen R, Ertl P, Tetko IV, Godin G (2020) Gen: highly efficient smiles explorer using
autodidactic generative examination networks. Journal of Cheminformatics 12(1):1–14
20. Schwaller P, Gaudin T, Lanyi D, Bekas C, Laino T (2018) “Found in translation”: predicting
outcomes of complex organic chemistry reactions using neural sequence-to-sequence models.
Chemical Science 9(28):6091–6098
21. Ucak UV, Ashyrmamatov I, Lee J (2023) Improving the quality of chemical language model
outcomes with atom-in-smiles tokenization. Journal of Cheminformatics 15(1):55
22. Li X, Fourches D (2021) Smiles pair encoding: a data-driven substructure tokenization
algorithm for deep learning. Journal of Chemical Information and Modeling 61(4):1560–1569
23. Christofidellis D, Giannone G, Born J, Winther O, Laino T, Manica M (2023) Unifying
molecular and textual representations via multi-task language modelling. In: International
Conference on Machine Learning
24. Krenn M, Häse F, Nigam A, Friederich P, Aspuru-Guzik A (2020) Self-referencing embedded
strings (selfies): A 100% robust molecular string representation. Machine Learning: Science
and Technology 1(4):045024
25. Heller SR, McNaught A, Pletnev I, Stein S, Tchekhovskoi D (2015) InChl, the IUPAC
international chemical identifier. Journal of Cheminformatics 7(1):1–34
26. Handsel J, Matthews B, Knight NJ, Coles SJ (2021) Translating the InChl: adapting
neural machine translation to predict iupac names from a chemical identifier. Journal of
Cheminformatics 13(1):1–11
27. Born J, Manica M (2021) Trends in deep learning for property-driven drug design. Current
Medicinal Chemistry 28(38):7862–7886
28. Segler MH, Kogej T, Tyrchan C, Waller MP (2018) Generating focused molecule libraries for
drug discovery with recurrent neural networks. ACS Central Science 4(1):120–131
arXiv:1301.3781
arXiv:1703.07076

138 N. Janakarajan et al.
https://t.me/med1917
29. Flam-Shepherd D, Zhu K, Aspuru-Guzik A (2022) Language models can learn complex
molecular distributions. Nature Communications 13(1):3293
30. Polykovskiy D, Zhebrak A, Sanchez-Lengeling B, Golovanov S, Tatanov O, Belyaev S,
Kurbanov R, Artamonov A, Aladinskiy V, Veselov M, et al (2020) Molecular sets (moses): a
benchmarking platform for molecular generation models. Front Pharmacol 11:1931
31. Joulin A, Mikolov T (2015) Inferring algorithmic patterns with stack-augmented recurrent
nets. Advances in Neural Information Processing Systems 28
32. Popova M, Isayev O, Tropsha A (2018) Deep reinforcement learning for de novo drug design.
Science Advances 4(7):eaap7885
33. Schilter O, Vaucher A, Schwaller P, Laino T (2023) Designing catalysts with deep generative
models and computational data. a case study for Suzuki cross coupling reactions. Digital
Discovery 2(3):728–735
34. Lim J, Ryu S, Kim JW, Kim WY (2018) Molecular generative model based on conditional
variational autoencoder for de novo molecular design. Journal of Cheminformatics 10(1):1–9
35. Born J, Manica M, Oskooei A, Cadow J, Markert G, Martínez MR (2021) PaccMannRL:De
novo generation of hit-like anticancer molecules from transcriptomic data via reinforcement
learning. iScience 24(4):102269
36. Born J, Manica M, Cadow J, Markert G, Mill NA, Filipavicius M, Janakarajan N, Cardinale
A, Laino T, Martínez MR (2021) Data-driven molecular design for discovery and synthesis
of novel ligands: a case study on sars-cov-2. Mach Learn: Sci Technol 2(2):025024
37. Born J, Huynh T, Stroobants A, Cornell WD, Manica M (2021) Active site sequence representations of human kinases outperform full sequence representations for affinity prediction and
inhibitor generation: 3d effects in a 1d model. Journal of Chemical Information and Modeling
62(2):240–257
38. Janakarajan N, Born J, Manica M (2022) A fully differentiable set autoencoder. In: Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pp
3061–3071
39. Radford A, Narasimhan K, Salimans T, Sutskever I, et al (2018) Improving language
understanding by generative pre-training
40. Bagal V, Aggarwal R, Vinod P, Priyakumar UD (2021) Molgpt: molecular generation using a
transformer-decoder model. Journal of Chemical Information and Modeling 62(9):2064–2076
41. Mazuz E, Shtar G, Shapira B, Rokach L (2023) Molecule generation using transformers and
policy gradient reinforcement learning. Scientific Reports 13(1):8799
42. Born J, Manica M (2023) Regression transformer enables concurrent sequence regression and
generation for molecular language modelling. Nature Machine Intelligence 5(4):432–444
43. Wu Z, Ramsundar B, Feinberg EN, Gomes J, Geniesse C, Pappu AS, Leswing K, Pande
V (2018) Moleculenet: a benchmark for molecular machine learning. Chemical Science
9(2):513–530
44. Born J, Markert G, Janakarajan N, Kimber TB, Volkamer A, Martínez MR, Manica M (2023)
Chemical representation learning for toxicity prediction. Digital Discovery
45. Bahdanau D, Cho K, Bengio Y (2014) Neural machine translation by jointly learning to align
and translate. arXiv preprint https://arxivorg/abs/14090473, arXiv1409.0473
46. Fabian B, Edlich T, Gaspar H, Segler M, Meyers J, Fiscato M, Ahmed M (2020) Molecular
representation learning with language models and domain-relevant auxiliary tasks. arXiv
preprint
47. Chithrananda S, Grand G, Ramsundar B (2020) Chemberta: large-scale self-supervised
48. Ross J, Belgodere B, Chenthamarakshan V, Padhi I, Mroueh Y, Das P (2022) Large-
49. Maziarka L, Danel T, Mucha S, Rataj K, Tabor J, Jastrzkebski S (2019) Molecule-augmented
arXiv:2011.13230
pretraining for molecular property prediction. arXiv preprint
scale chemical language representations capture molecular structure and properties. Nature
Machine Intelligence 4(12):1256–1264
attention transformer. In: Workshop on Graph Representation Learning, Neural Information
Processing Systems
arXiv:2010.09885

7 Language Models in Molecular Discovery 139
https://t.me/med1917
50. Maziarka L, Majchrowski D, Danel T, Gainski P, Tabor J, Podolak I, Morkisz P, Jastrzkebski
S (2024) Relative molecule self-attention transformer. Journal of Cheminformatics 16(1):3
51. Ovchinnikova K, Born J, Chouvardas P, Rapsomaniki M, Kruithof-de Julio M (2024) Overcoming limitations in current measures of drug response may enable AI-driven precision
oncology Abstract npj Precision Oncology 8(1).
52. Born J, Shoshan Y, Huynh T, Cornell WD, Martin EJ, Manica M (2022) On the choice of
active site sequences for kinase-ligand affinity prediction. Journal of Chemical Information
and Modeling 62(18):4295–4299.
53. Gezelter JD (2015) Open source and open data should be standard practices
54. Wolf T, Debut L, Sanh V, Chaumond J, Delangue C, Moi A, Cistac P, Rault T, Louf R,
Funtowicz M, et al (2020) Transformers: State-of-the-art natural language processing. In:
Proceedings of the 2020 conference on empirical methods in natural language processing:
system demonstrations, pp 38–45
55. Dosovitskiy A, Beyer L, Kolesnikov A, Weissenborn D, Zhai X, Unterthiner T, Dehghani
M, Minderer M, Heigold G, Gelly S, Uszkoreit J, Houlsby N (2021) An image is worth
16x16 words: Transformers for image recognition at scale. In: 9th International Conference
on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3–7, 2021
56. Chen L, Lu K, Rajeswaran A, Lee K, Grover A, Laskin M, Abbeel P, Srinivas A, Mordatch
I (2021) Decision transformer: Reinforcement learning via sequence modeling. Advances in
Neural Information Processing Systems 34:15084–15097
57. Jumper J, Evans R, Pritzel A, Green T, Figurnov M, Ronneberger O, Tunyasuvunakool K,
Bates R, Žídek A, Potapenko A, et al (2021) Highly accurate protein structure prediction with
alphafold. Nature 596(7873):583–589
58. Schwaller P, Vaucher AC, Laplaza R, Bunne C, Krause A, Corminboeuf C, Laino T
(2022) Machine intelligence for chemical reaction space. Wiley Interdisciplinary Reviews:
Computational Molecular Science 12(5):e1604
59. Edwards C, Lai T, Ros K, Honke G, Cho K, Ji H (2022) Translation between molecules and
natural language. In: 2022 Conference on Empirical Methods in Natural Language Processing,
EMNLP 2022
60. Elnaggar A, Heinzinger M, Dallago C, Rehawi G, Yu W, Jones L, Gibbs T, Feher T, Angerer
C, Steinegger M, Bhowmik D, Rost B (2021) Prottrans: Towards cracking the language of
life’s code through self-supervised deep learning and high-performance computing. IEEE
Transactions on Pattern Analysis and Machine Intelligence pp 1–1,
TPAMI.2021.3095381
61. Manica M, Born J, Cadow J, Christofidellis D, Dave A, Clarke D, Teukam YGN, Giannone
G, Hoffman SC, Buchan M, et al (2023) Accelerating material design with the generative
toolkit for scientific discovery. npj Computational Materials 9(1):69
62. Huang K, Fu T, Gao W, Zhao Y, Roohani Y, Leskovec J, W CC, Xiao C, Sun J, Zitnik M
(2021) Therapeutics data commons: Machine learning datasets and tasks for drug discovery
and development. Advances in Neural Information Processing System 35
63. Ramsundar B, Eastman P, Walters P, Pande V, Leswing K, Wu Z (2019) Deep Learning for
the Life Sciences. O’Reilly Media,
Microscopy/dp/1492039837
64. von Platen P, Patil S, Lozhkov A, Cuenca P, Lambert N, Rasul K, Davaadorj M, Wolf T
(2022) Diffusers: State-of-the-art diffusion models.
Accessed: February 2, 2024
65. Zhu Z, Shi C, Zhang Z, Liu S, Xu M, Yuan X, Zhang Y, Chen J, Cai H, Lu J, et al (2022)
Torchdrug: A powerful and flexible machine learning platform for drug discovery. arXiv
Preprint at
66. Brown N, Fiscato M, Segler MH, Vaucher AC (2019) Guacamol: benchmarking models for
de novo molecular design. J Chem Inf Model 59(3):1096–1108
67. Bengio Y, Lahlou S, Deleu T, Hu EJ, Tiwari M, Bengio E (2023) Gflownet foundations.
Journal of Machine Learning Research 24(210):1–55
arXiv:2202.08320
https://doi.org/10.1021/acs.jcim.2c00840
https://www.amazon.com/Deep-Learning-Life-Sciences-
https://doi.org/10.1038/s41698-024-00583-0
https://doi.org/10.1109/
https://github.com/huggingface/diffusers

140 N. Janakarajan et al.
https://t.me/med1917
68. Maziarz K, Jackson-Flux H, Cameron P, Sirockin F, Schneider N, Stiefl N, Segler M,
Brockschmidt M (2022) Learning to extend molecular scaffolds with structural motif. In:
The Tenth International Conference on Learning Representations, ICLR
69. Abid A, Abdalla A, Abid A, Khan D, Alfozan A, Zou J (2019) Gradio: Hassle-free
sharing and testing of ml models in the wild. arXiv preprint https://arxivorg/abs/190602569
arXiv1906.02569
70. for Chemistry team IR (2023) rxn4chemistry: Python wrapper for the IBM RXN for Chemistry
API.
71. Schwaller P, Laino T, Gaudin T, Bolgar P, Hunter CA, Bekas C, Lee AA (2019) Molecular
72. Pesciullesi G, Schwaller P, Laino T, Reymond JL (2020) Transfer learning enables the
73. Toniato A, Schwaller P, Cardinale A, Geluykens J, Laino T (2021) Unassisted noise reduction
74. Schwaller P, Petraglia R, Zullo V, Nair VH, Haeuselmann RA, Pisoni R, Bekas C, Iuliano
75. Zipoli F, Baldassari C, Manica M, Born J, Laino T (2024) Growing strings in a chemical
76. Probst D, Manica M, Nana Teukam YG, Castrogiovanni A, Paratore F, Laino T (2022)
77. Thakkar A, Vaucher AC, Byekwaso A, Schwaller P, Toniato A, Laino T (2023) Unbiasing
78. Devlin J, Chang MW, Lee K, Toutanova K (2018) Bert: Pre-training of deep bidirectional
79. Schwaller P, Probst D, Vaucher AC, Nair VH, Kreutter D, Laino T, Reymond JL (2021)
80. Schwaller P, Hoover B, Reymond JL, Strobelt H, Laino T (2021) Extraction of organic
81. Vaucher AC, Zipoli F, Geluykens J, Nair VH, Schwaller P, Laino T (2020) Automated extrac-
82. Vaucher AC, Schwaller P, Geluykens J, Nair VH, Iuliano A, Laino T (2021) Inferring
83. Genheden S, Thakkar A, Chadimová V, Reymond JL, Engkvist O, Bjerrum E (2020) Aizyn-
84. Gainski P, Maziarka L, Danel T, Jastrzebski S (2022) Huggingmolecules: An open-source
85. Landrum G (2013) Rdkit documentation. Release 1(1–79):4
86. Lin TS, Coley CW, Mochigase H, Beech HK, Wang W, Wang Z, Woods E, Craig SL,
87. Born J, Shoshan Y, Huynh T, Cornell WD, Martin EJ, Manica M (2022) On the choice of
https://github.com/rxn4chemistry/rxn4chemistry, accessed: February 2, 2024
transformer: a model for uncertainty-calibrated chemical reaction prediction. ACS Central
Science 5(9):1572–1583
molecular transformer to predict regio-and stereoselective reactions on carbohydrates. Nature
Communications 11(1):4874
of chemical reaction datasets. Nature Machine Intelligence 3(6):485–494
A, Laino T (2020) Predicting retrosynthetic pathways using transformer-based models and a
hyper-graph exploration strategy. Chemical Science 11(12):3316–3325
reaction space for searching retrosynthesis pathways Abstract npj Computational Materials
10(1).
https://doi.org/10.1038/s41524-024-01290-x
Biocatalysed synthesis planning using data-driven learning. Nature Communications
13(1):964
retrosynthesis language models with disconnection prompts. ACS Central Science
transformers for language understanding. arXiv preprint
Mapping the space of chemical reactions using attention-based neural networks. Nature
Machine Intelligence 3(2):144–152
chemistry grammar from unsupervised learning of chemical reactions. Science Advances
7(15):eabe4166
tion of chemical synthesis actions from experimental procedures. Nature Communications
11(1):3601
experimental procedures from text-based representations of chemical reactions. Nature
Communications 12(1):2573
thfinder: a fast, robust and flexible open-source software for retrosynthetic planning. Journal
of Cheminformatics 12(1):70
library for transformer-based molecular property prediction (student abstract). In: Proceedings
of the AAAI Conference on Artificial Intelligence, vol 36, pp 12949–12950
Johnson JA, Kalow JA, et al (2019) Bigsmiles: a structurally-based line notation for describing
macromolecules. ACS Central Science 5(9):1523–1531
active site sequences for kinase-ligand affinity prediction. Journal of Chemical Information
and Modeling 62(18):4295–4299
arXiv:1810.04805

7 Language Models in Molecular Discovery 141
https://t.me/med1917
88. Heyndrickx W, Mervin L, Morawietz T, Sturm N, Friedrich L, Zalewski A, Pentina A,
Humbeck L, Oldenhof M, Niwayama R, et al (2022) Melloddy: cross pharma federated
learning at unprecedented scale unlocks benefits in QSAR without compromising proprietary
information
89. Gorgulla C, Boeszoermenyi A, Wang ZF, Fischer PD, Coote PW, Padmanabha Das KM,
Malets YS, Radchenko DS, Moroz YS, Scott DA, et al (2020) An open-source drug discovery
platform enables ultra-large virtual screens. Nature 580(7805):663–668
90. Ivanenkov YA, Polykovskiy D, Bezrukov D, Zagribelnyy B, Aladinskiy V, Kamya P, Aliper
A, Ren F, Zhavoronkov A (2023) Chemistry42: an AI-driven platform for molecular design
and optimization. Journal of Chemical Information and Modeling 63(3):695–701
91. OpenAI (2023) Chatgpt. https://chat.openai.com/chat, accessed: August 8, 2023
92. GitHub (2024) Github copilot
93. Christiano PF, Leike J, Brown T, Martic M, Legg S, Amodei D (2017) Deep reinforcement
learning from human preferences. Advances in Neural Information Processing Systems 30
94. Brown T, Mann B, Ryder N, Subbiah M, Kaplan JD, Dhariwal P, Neelakantan A, Shyam P,
Sastry G, Askell A, et al (2020) Language models are few-shot learners. Advances in Neural
Information Processing Systems 33:1877–1901
95. Sanh V, Webson A, Raffel C, Bach SH, Sutawika L, Alyafeai Z, Chaffin A, Stiegler A, Le Scao
T, Raja A, et al (2022) Multitask prompted training enables zero-shot task generalization. In:
ICLR 2022-Tenth International Conference on Learning Representations
96. Fei N, Lu Z, Gao Y, Yang G, Huo Y, Wen J, Lu H, Song R, Gao X, Xiang T, et al (2022) Towards
artificial general intelligence via a multimodal foundation model. Nature Communications
13(1):3094
97. Moor M, Banerjee O, Abad ZSH, Krumholz HM, Leskovec J, Topol EJ, Rajpurkar P (2023)
Foundation models for generalist medical artificial intelligence. Nature 616(7956):259–265
98. Wei J, Wang X, Schuurmans D, Bosma M, Xia F, Chi E, Le QV, Zhou D, et al (2022)
Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural
Information Processing Systems 35:24824–24837
99. Zeng Z, Yao Y, Liu Z, Sun M (2022) A deep-learning system bridging molecule structure and biomedical text with comprehension comparable to human professionals. Nature
Communications 13(1):862
100. Lu J, Zhang Y (2022) Unified deep learning model for multitask reaction predictions with
explanation. Journal of Chemical Information and Modeling 62(6):1376–1387
101. Taylor R, Kardas M, Cucurull G, Scialom T, Hartshorn A, Saravia E, Poulton A, Kerkez V,
Stojnic R (2022) Galactica: A large language model for science. arXiv Preprint at
09085
102. Kim S, Chen J, Cheng T, Gindulyte A, He J, He S, Li Q, Shoemaker BA, Thiessen PA, Yu B,
et al (2019) Pubchem 2019 update: improved access to chemical data. Nucleic Acids Research
47(D1):D1102–D1109
103. White AD, Hocky GM, Gandhi HA, Ansari M, Cox S, Wellawatte GP, Sasmal S, Yang Z, Liu
K, Singh Y, et al (2023) Assessment of chemistry knowledge in large language models that
generate code. Digital Discovery 2(2):368–376
104. Castro Nascimento CM, Pimentel AS (2023) Do large language models understand chemistry?
a conversation with chatgpt. Journal of Chemical Information and Modeling 63(6):1649–1655
105. Wildman SA, Crippen GM (1999) Prediction of physicochemical parameters by atomic
contributions. Journal of Chemical Information and Computer Sciences 39(5):868–873
106. Bickerton GR, Paolini GV, Besnard J, Muresan S, Hopkins AL (2012) Quantifying the
chemical beauty of drugs. Nat Chem 4(2):90–98
107. Ertl P, Schuffenhauer A (2009) Estimation of synthetic accessibility score of drug-like
molecules based on molecular complexity and fragment contributions. Journal of Cheminformatics 1:1–11
108. Tanimoto TT (1957) Ibm internal report. Nov 17:1957
109. Rogers D, Hahn M (2010) Extended-connectivity fingerprints. Journal of Chemical Information and Modeling 50(5):742–754
arXiv:2211.

Chapter 8
https://t.me/med1917
Transformers and Large Language
Models for Chemistry and Drug
Discovery
Andres M. Bran and Philippe Schwaller
8.1 Introduction
The capacity to process and accurately model human language has been a persistent
pursuit within the machine learning community [
intrinsic to human reasoning capabilities, thus successful language modeling could
open the door to numerous applications, enhancing various information processing
tasks with the potential to revolutionize several industries [
language processing has witnessed significant advancements in recent years, thanks
to improved computing infrastructure, breakthroughs in algorithms, and the proliferation of abundant and accessible data [
also play a pivotal role in the domain of chemistry, which serves as the fundamental
basis for drug discovery and development. Analogous to human language, understanding and accurately modeling the language of chemistry is crucial for effective
research and development in the pharmaceutical industry. By applying the advancements in language modeling and processing from the machine learning community
to the domain of chemistry, it is possible to facilitate drug discovery by efficiently
analyzing and interpreting vast amounts of chemical data and literature.
Introduced in 2017, the Transformer architecture revolutionized natural language
processing [
neural machine translation. In its original architecture, the Transformer consists of
an encoder, which encodes a sentence in the source language (e.g., French) and a
5]. The Transformer is a type of neural network initially developed for
1–5]. The belief is that language is
6, 7]. The field of natural
]. Language and technical terminology
8
A. M. Bran · P. S chwa l l er (B)
Laboratory of Artificial Chemical Intelligence (LIAC), ISIC, EPFL, Lausanne, Switzerland
e-mail: philippe.schwaller@epfl.ch
National Centre of Competence in Research (NCCR) Catalysis, Ecole Polytechnique Fédérale de
Lausanne (EPFL), Lausanne, Switzerland
A. M. Bran
e-mail:
andres.marulandabran@epfl.ch
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2024
H. Satoh et al. (eds.), Drug Development Supported by Informatics,
https://doi.org/10.1007/978-981-97-4828-0_8
143
Соседние файлы в папке Библиотека им академика М.И. Перельмана
