Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5858_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
30.08.2026
Размер:
43 Мб
Скачать
154 A. M. Bran and P. Schwaller
https://t.me/med1917
Fig. 8.3 Recent advances in large language models (LLMs) have launched a new era of Trans­formers in chemistry. Their capabilities allow for (a) general task solvers thanks to their flexibility and knowledge transferability [ unlimited modalities, in the form of computational tools [
70, 71]and (b) agent architectures, capable of integrating virtually
47, 72, 73
]
these endeavors aim to make the most of existing knowledge, enabling models to learn as much as possible from the few available data points.
Interestingly, Jablonka et al. [70
] demonstrated that fine-tuning and in-context learning can perform on par with, and in some instances even outperform, these specialized techniques, particularly when data is limited. The high performance of in-context learning, combined with its ease of use and flexibility, makes it one of the most impressive applications of LLMs in chemistry to date. This technique holds the potential to revolutionize the way machine learning is utilized in the scientific field, by rapidly highlighting complex correlations in data.
Another application of regression, which holds significant interest in chemistry and drug discovery cycles, is optimization. This process involves modifying an object until a property of interest reaches a desired value. Typically, this requires a large number of measurements of the desired property, which can be quite costly in chemistry use-cases. Applications of this include yield/selectivity optimization
8 Transformers and Large Language Models for Chemistry and Drug … 155
https://t.me/med1917
in chemical reactions and the generation of molecular candidates with target prop­erties. Bayesian Optimization (BO) has recently been proposed as a solution to such problems in chemistry [ However, BO requires uncertainty-calibrated regression methods, which sets it apart from conventional regression.
In line with the concept of in-context learning, Ramos et al. [71] proposed a system that utilizes GPT models to perform regression while also incorporating uncertainty. This approach enables BO without the need for any feature engineering or fine-tuning. The flexibility of this method allowed the team to perform catalyst and molecular optimization using only the synthesis procedure of the catalyst as input. This work represents a paradigm shift in drug discovery and molecular design. For the first time, it showcases a direct map from the synthesis procedure into property space, effectively overcoming issues like the synthesizability of proposed molecules, a key limitation of structure-based generative models.
8.3.2.2 Molecular Generation
Another fascinating application of the generative capabilities of language models is molecular generation. This area, which is of significant importance in the drug discovery process, has been largely dominated by models that generate molecules in the form of linear string representations, such as SMILES or SELFIES [
86], its successful application is contingent on the ability to specify substances as
graphs and their subsequent conversion to a linear string representation. However, this approach is only suitable for a subset of organic molecules. Other substances, such as macromolecules and materials, necessitate more comprehensive represen­tations. A complete and accurate representation of these substances can only be achieved by specifying atomic positions, boundary conditions, and other factors. This requirement presents a significant challenge and limitation to the current methods of molecular generation. To address these limitations, Flam et al. [ using language models for structure generation, directly generating them with three­dimensional atomic positions. Besides being innovative and valid, the generated structures can be obtained by training models in a variety of formats used for crystals, proteins, and more. This work also demonstrates performance comparable to expert­designed, state-of-the-art algorithms for molecular generation based on graphs, while overcoming the limitations mentioned earlier.
80, 81], particularly in situations where data is small.
82–84
]. While
70, 85,
87
] proposed
8.3.3 Language Model-Powered Agents
Among the most useful emergent abilities of language models are step-by-step reasoning, activated through chain-of-thought (CoT) prompting, and their capacity to effectively use tools [
88]. These capabilities have been the subject of extensive
156 A. M. Bran and P. Schwaller
https://t.me/med1917
research in recent years, and their application has been shown to significantly enhance the performance of LLMs across a variety of tasks. CoT prompting is a technique where language models are instructed to solve a task by following a sequence of reasoning steps, rather than providing an answer in a single response [ language models in this way effectively allows them to perform symbolic operations, much like humans perform arithmetic operations by keeping track of intermediate steps.
The ability to use tools is another significant capability of language models [88]. This allows them to invoke external computational tools, thereby enriching their knowledge through querying search engines, accessing calculators, and so on [ These capabilities have been demonstrated to enhance the performance of large language models in a range of tasks that were previously inaccessible.
The recent advancements and results in revealing and exploiting the capabilities of LLMs suggest the potential for combining some of these capabilities to create more powerful and useful possibilities. This concept has been recently explored, leading to the development of the Modular Reasoning, Knowledge and Language (MRKL [ tool-using capabilities of modern LLMs. By incorporating external tools into a CoT setting, agents of this type have recently been shown to outperform other methods based on large language models.
unimodality issue of LLMs. Under this setting, they become capable of processing different types of input data, making real-time decisions in simulated environments, and even interacting with real-world robotic platforms. The solutions provided by LLMs to tasks also become more grounded in reality, as access to certain tools provides them with real, up-to-date information relevant to the task. This can, to some extent, limit the tendency of these models to generate unrealistic or “hallucinated” responses.
89]) and Reason + Act (ReAct [90]) systems, which combine the CoT and
One direct benefit of effective tool usage is that it partially overcomes the
68]. Instructing
88].
8.3.3.1 Agents in Chemistry: Unleashing the Power from Tools
Despite their strengths as text generators and task solvers, and their remarkable few­shot and zero-shot performance, these models are also well known for their high propensity to generate false and inaccurate content, an issue that extends to easily verifiable matters such as basic arithmetic [ limitations make the direct application of LLMs to chemistry a challenging matter.
The potential applications of large language models in chemistry were first explored in a large-scale collaboration involving researchers from around the world, an effort that resulted in the demonstration of 14 use-cases [ range from wrappers for computational tools, which enhance their accessibility by allowing natural language inputs to modify behaviors, to assistants for reaction opti­mization, and knowledge parsers and synthesizers for scientific question answering, among others. These are just a few of the possibilities that LLMs offer in chemistry,
88] and chemical operations [91]. These
]. The applications
92
8 Transformers and Large Language Models for Chemistry and Drug … 157
https://t.me/med1917
which, when combined with existing chemistry tools and databases, significantly increase the applicability and accessibility of computational applications.
More recently, Bran and Cox et al. [73] extended the concept of LLM-powered agents for chemistry by curating and compiling a set of computational chemistry tools. Their system, ChemCrow, has been shown to be capable of planning and executing tasks in chemistry, effectively streamlining the reasoning process for several common chemical tasks across areas such as drug and materials design and synthesis. The authors demonstrate that this approach has a highly positive effect on LLM’s performance for tasks in chemistry, overcoming hallucinations and grounding their responses with data from reliable sources. Complementary approaches exist,
72
with a sharper focus on cloud lab operability [
The power of platforms like ChemCrow extends beyond merely serving as inde­pendent task solvers. They can be viewed as general chemistry assistants with the ultimate goal of making computational tools more accessible to chemists, thereby accelerating discovery. An additional highlight is the seamless exploitation of tool composability that this allows. It makes it straightforward to enrich the results of one tool with another, or to construct custom tool pipelines, all through simple requests in natural language.
].
8.4 Outlook and Final Remarks
Advancements in neural translation models, and specially with the introduction of the Transformer architecture, have sparked a revolution in machine learning for applica­tions in chemistry and drug development. Analogies between chemical and natural language, and the publication of open databases and benchmarks, have inspired the representation of chemical tasks in the form of text, allowing straightforward application of Transformers to problems in this field.
This revolution has occurred in three stages, differentiated by the specificity of tasks. In a first stage, characterized by task-specificity and use of single-modality models, applications spanned molecule-to-molecule conversion tasks, like reaction outcome prediction and retrosynthetic planning, along with representation learning and downstream tasks like regression and classification. Their excellent performance and relative simplicity made them de-facto models in an array of applications.
In a second stage, researchers attempted to connect multiple additional modali­ties relevant to chemistry, like spectra from experiments, sequences of experimental actions, and even natural language, opening the way for an expanded number of appli­cations involving modalities of any sort, however still task-specific. More recently, powered by vertiginous advancements in training and tuning of large language models, a series of works have been published that leverage a number of capa­bilities from such models. Among others, these contributions showcase applications in regression, classification, molecular generation and reaction optimization, all with unprecedented flexibility and usually improved performance over other methods.
158 A. M. Bran and P. Schwaller
https://t.me/med1917
Another direction explores the integration of virtually unlimited modalities —in the form of tools—into agents powered by LLMs. The power of these agents has been demonstrated through a number of diverse tasks, ranging from molecular generation to automated organic synthesis, in an open-ended, highly customizable fashion.
By leveraging the expressivity and flexibility of natural language, this last wave of applications aims to bridge the gap between the chemical and natural languages. As we continue to explore and harness these capabilities, we can look forward to a future where machine learning plays an even more integral role in accelerating scientific discovery.
References
1. Bahdanau D, Cho K, Bengio Y (2016) Neural Machine Translation by Jointly Learning to Align and Translate.
2. Cho K, van Merrienboer B, Gulcehre C, Bahdanau D, Bougares F, Schwenk H, Bengio Y (2014) Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation.
3. Hochreiter S, Schmidhuber J (1997) Long Short-Term Memory. Neural Comput 9:1735–1780.
https://doi.org/10.1162/neco.1997.9.8.1735
4. Sutskever I, Vinyals O, Le QV (2014) Sequence to Sequence Learning with Neural Networks.
https://doi.org/10.48550/arXiv.1409.3215
5. Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser L, Polosukhin I (2017) Attention Is All You Need.
6. Bommasani R, Hudson DA, Adeli E, Altman R, Arora S, von Arx S, Bernstein MS, Bohg J, Bosselut A, Brunskill E, Brynjolfsson E, Buch S, Card D, Castellon R, Chatterji N, Chen A, Creel K, Davis JQ, Demszky D, Donahue C, Doumbouya M, Durmus E, Ermon S, Etchemendy J, Ethayarajh K, Fei-Fei L, Finn C, Gale T, Gillespie L, Goel K, Goodman N, Grossman S, Guha N, Hashimoto T, Henderson P, Hewitt J, Ho DE, Hong J, Hsu K, Huang J, Icard T, Jain S, Jurafsky D, Kalluri P, Karamcheti S, Keeling G, Khani F, Khattab O, Koh PW, Krass M, Krishna R, Kuditipudi R, Kumar A, Ladhak F, Lee M, Lee T, Leskovec J, Levent I, Li XL, Li X, Ma T, Malik A, Manning CD, Mirchandani S, Mitchell E, Munyikwa Z, Nair S, Narayan A, Narayanan D, Newman B, Nie A, Niebles JC, Nilforoshan H, Nyarko J, Ogut G, Orr L, Papadimitriou I, Park JS, Piech C, Portelance E, Potts C, Raghunathan A, Reich R, Ren H, Rong F, Roohani Y, Ruiz C, Ryan J, Ré C, Sadigh D, Sagawa S, Santhanam K, Shih A, Srinivasan K, Tamkin A, Taori R, Thomas AW, Tramèr F, Wang RE, Wang W, Wu B, Wu J, Wu Y, Xie SM, Yasunaga M, You J, Zaharia M, Zhang M, Zhang T, Zhang X, Zhang Y, Zheng L, Zhou K, Liang P (2022) On the Opportunities and Risks of Foundation Models.
48550/arXiv.2108.07258
7. Bubeck S, Chandrasekaran V, Eldan R, Gehrke J, Horvitz E, Kamar E, Lee P, Lee YT, Li Y, Lundberg S, Nori H, Palangi H, Ribeiro MT, Zhang Y (2023) Sparks of Artificial General Intelligence: Early Experiments with GPT-4.
8. Gao L, Biderman S, Black S, Golding L, Hoppe T, Foster C, Phang J, He H, Thite A, Nabeshima N, Presser S, Leahy C (2020) The Pile: An 800GB Dataset of Diverse Text for Language Modeling.
9. Raffel C, Shazeer N, Roberts A, Lee K, Narang S, Matena M, Zhou Y, Li W, Liu PJ (2020)
doi.org/10.48550/arXiv.1910.10683
10. Devlin J, Chang M-W, Lee K, Toutanova K (2019) BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.
https://doi.org/10.48550/arXiv.2101.00027
https://doi.org/10.48550/arXiv.1409.0473
https://doi.org/10.48550/arXiv.1406.1078
https://doi.org/10.48550/arXiv.1706.03762
https://doi.org/10.
https://doi.org/10.48550/arXiv.2303.12712
https://
https://doi.org/10.48550/arXiv.1810.04805
8 Transformers and Large Language Models for Chemistry and Drug … 159
https://t.me/med1917
11. Reimers N, Gurevych I (2019) Sentence-BERT: Sentence Embeddings using Siamese BERT­Networks.
12. Kryscinski W, Keskar NS, McCann B, Xiong C, Socher R (2019) Neural Text Summariza­tion: A Critical Evaluation. In: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). Association for Computational Linguistics, Hong Kong, China, pp 540–551
13. Liu PJ, Saleh M, Pot E, Goodrich B, Sepassi R, Kaiser L, Shazeer N (2018) Generating Wikipedia by Summarizing Long Sequences.
14. OpenAI (2023) GPT-4 Technical Report. https://doi.org/10.48550/arXiv.2303.08774
15. Jumper J, Evans R, Pritzel A, Green T, Figurnov M, Ronneberger O, Tunyasuvunakool K, Bates R, Žídek A, Potapenko A, Bridgland A, Meyer C, Kohl SAA, Ballard AJ, Cowie A, Romera­Paredes B, Nikolov S, Jain R, Adler J, Back T, Petersen S, Reiman D, Clancy E, Zielinski M, Steinegger M, Pacholska M, Berghammer T, Bodenstein S, Silver D, Vinyals O, Senior AW, Kavukcuoglu K, Kohli P, Hassabis D (2021) Highly Accurate Protein Structure Prediction with AlphaFold. Nature 596:583–589.
16. Akdel M, Pires DEV, Pardo EP, Jänes J, Zalevsky AO, Mészáros B, Bryant P, Good LL, Laskowski RA, Pozzati G, Shenoy A, Zhu W, Kundrotas P, Serra VR, Rodrigues CHM, Dunham AS, Burke D, Borkakoti N, Velankar S, Frost A, Basquin J, Lindorff-Larsen K, Bateman A, Kajava AV, Valencia A, Ovchinnikov S, Durairaj J, Ascher DB, Thornton JM, Davey NE, Stein A, Elofsson A, Croll TI, Beltrao P (2022) A Structural Biology Community Assessment of AlphaFold2 Applications. Nat Struct Mol Biol 29:1056–1067.
022-00849-w
17. Yang Z, Zeng X, Zhao Y, Chen R (2023) AlphaFold2 and its Applications in the Fields of Biology and Medicine. Signal Transduct Target Ther 8:1–14.
023-01381-z
18. Huang K, Fu T, Gao W, Zhao Y, Roohani Y, Leskovec J, Coley C, Xiao C, Sun J, Zitnik M (2021) Therapeutics Data Commons: Machine Learning Datasets and Tasks for Drug Discovery and Development. Proc Neural Inf Process Syst Track Datasets Benchmarks 1
19. Kearnes SM, Maser MR, Wleklinski M, Kast A, Doyle AG, Dreher SD, Hawkins JM, Jensen KF, Coley CW (2021) The Open Reaction Database. J Am Chem Soc 143:18820–18826.
https://doi.org/10.1021/jacs.1c09820
20. Lowe D (2012) Extraction of Chemical Structures and Reactions from the Literature. University of Cambridge.
21. Wu Z, Ramsundar B, Feinberg EN, Gomes J, Geniesse C, Pappu AS, Leswing K, Pande V (2018) MoleculeNet: A Benchmark for Molecular Machine Learning. Chem Sci 9:513–530.
https://doi.org/10.1039/C7SC02664A
22. Cadeddu A, Wylie EK, Jurczak J, Wampler-Doty M, Grzybowski BA (2014) Organic Chemistry as a Language and the Implications of Chemical Linguistics for Structural and Retrosynthetic Analyses. Angew Chem Int Ed 53:8108–8112.
23. Wołos A, Koszelewski D, Roszak R, Szymkuć S, Moskal M, Ostaszewski R, Herrera BT, Maier JM, Brezicki G, Samuel J, Lummiss JAM, McQuade DT, Rogers L, Grzybowski BA (2022) Computer-Designed Repurposing of Chemical Wastes into Drugs. Nature 604:668–676.
https://doi.org/10.1038/s41586-022-04503-9
24. Schwaller P, Hoover B, Reymond J-L, Strobelt H, Laino T (2021) Extraction of Organic Chem­istry Grammar from Unsupervised Learning of Chemical Reactions. Sci Adv 7:eabe4166.
https://doi.org/10.1126/sciadv.abe4166
25. Brammer JC, Blanke G, Kellner C, Hoffmann A, Herres-Pawlis S, Schatzschneider U (2022) TUCAN: A Molecular Identifier and Descriptor Applicable to the Whole Periodic Table from Hydrogen to Oganesson. J Cheminformatics 14:66.
40-5
26. Heller SR, McNaught A, Pletnev I, Stein S, Tchekhovskoi D (2015) InChI, the IUPAC Inter­national Chemical Identifier. J Cheminformatics 7:23.
0068-4
https://doi.org/10.48550/arXiv.1908.10084
https://doi.org/10.48550/arXiv.1801.10198
https://doi.org/10.1038/s41586-021-03819-2
https://doi.org/10.1038/s41594-
https://doi.org/10.1038/s41392-
https://doi.org/10.17863/CAM.16293
https://doi.org/10.1002/anie.201403708
https://doi.org/10.1186/s13321-022-006
https://doi.org/10.1186/s13321-015-
160 A. M. Bran and P. Schwaller
https://t.me/med1917
27. Krenn M, Ai Q, Barthel S, Carson N, Frei A, Frey NC, Friederich P, Gaudin T, Gayle AA, Jablonka KM, Lameiro RF, Lemm D, Lo A, Moosavi SM, Nápoles-Duarte JM, Nigam A, Pollice R, Rajan K, Schatzschneider U, Schwaller P, Skreta M, Smit B, Strieth-Kalthoff F, Sun C, Tom G, von Rudorff GF, Wang A, White A, Young A, Yu R, Aspuru-Guzik A (2022) SELFIES and the Future of Molecular String R epresentations. Patterns 3:100588.
org/10.1016/j.patter.2022.100588
28. Krenn M, Häse F, Nigam A, Friederich P, Aspuru-Guzik A (2020) Self-referencing Embedded Strings (SELFIES): A 100% Robust Molecular String Representation. Mach Learn Sci Technol 1:045024.
29. O’Boyle N, Dalke A (2018) DeepSMILES: An Adaptation of SMILES for Use in Machine­Learning of Chemical Structures.
30. Weininger D (1988) SMILES, a Chemical Language and Information System. 1. Introduction to Methodology and Encoding Rules. J Chem Inf Comput Sci 28:31–36.
1021/ci00057a005
31. Restrepo G (2022) Chemical Space: Limits, Evolution and Modelling of an Object Bigger than our Universal Library. Digit Discov 1:568–585.
32. Gómez-Bombarelli R, Wei JN, Duvenaud D, Hernández-Lobato JM, Sánchez-Lengeling B, Sheberla D, Aguilera-Iparraguirre J, Hirzel TD, Adams RP, Aspuru-Guzik A (2018) Automatic Chemical Design Using a Data-Driven Continuous Representation of Molecules. ACS Cent Sci 4:268–276.
33. Kusner MJ, Paige B, Hernández-Lobato JM (2017) Grammar Variational Autoencoder. https://
doi.org/10.48550/arXiv.1703.01925
34. Öztürk H, Özgür A, Schwaller P, Laino T, Ozkirimli E (2020) Exploring Chemical Space Using Natural Language Processing Methodologies for Drug Discovery. Drug Discov Today 25:689–705.
35. Pesciullesi G, Schwaller P, Laino T, Reymond J-L (2020) Transfer Learning Enables the Molecular Transformer to Predict Regio- and Stereoselective Reactions on Carbohydrates. Nat Commun 11:4874.
36. Schwaller P, Laino T, Gaudin T, Bolgar P, Hunter CA, Bekas C, Lee AA (2019) Molecular Transformer: A Model for Uncertainty-Calibrated Chemical Reaction Prediction. ACS Cent Sci 5:1572–1583.
37. Schwaller P, Petraglia R, Zullo V, Nair VH, Haeuselmann RA, Pisoni R, Bekas C, Iuliano A, Laino T (2020) Predicting Retrosynthetic Pathways using Transformer-Based Models and a Hyper-Graph Exploration Strategy. Chem Sci 11:3316–3325.
704H
38. Ahmad W, Simon E, Chithrananda S, Grand G, Ramsundar B (2022) ChemBERTa-2: Towards Chemical Foundation Models.
39. Chithrananda S, Grand G, Ramsundar B (2020) ChemBERTa: Large-Scale Self-Supervised Pretraining for Molecular Property Prediction.
40. Li J, Jiang X (2021) Mol-BERT: An Effective Molecular Representation with BERT for Molec­ular Property Prediction. Wirel Commun Mob Comput 2021:e7181815.
1155/2021/7181815
41. Schwaller P, Probst D, Vaucher AC, Nair VH, Kreutter D, Laino T, Reymond J-L (2021) Mapping the Space of Chemical Reactions Using Attention-Based Neural Networks. Nat Mach Intell 3:144–152.
42. Vaucher AC, Schwaller P, Geluykens J, Nair VH, Iuliano A, Laino T (2021) Inferring Exper­imental Procedures from Text-Based Representations of Chemical Reactions. Nat Commun 12:2573.
43. Vaucher AC, Zipoli F, Geluykens J, Nair VH, Schwaller P, Laino T (2020) Automated Extraction of Chemical Synthesis Actions from Experimental Procedures. Nat Commun 11:3601.
doi.org/10.1038/s41467-020-17266-6
44. Tetko IV, Karpov P, Van Deursen R, Godin G (2020) State-of-the-art Augmented NLP Trans­former Models for Direct and Single-Step Retrosynthesis. Nat Commun 11:5575.
org/10.1038/s41467-020-19266-y
https://doi.org/10.1088/2632-2153/aba947
https://doi.org/10.26434/chemrxiv.7097960.v1
https://doi.org/10.1039/D2DD00030J
https://doi.org/10.1021/acscentsci.7b00572
https://doi.org/10.1016/j.drudis.2020.01.020
https://doi.org/10.1038/s41467-020-18671-7
https://doi.org/10.1021/acscentsci.9b00576
https://doi.org/10.1039/C9SC05
https://doi.org/10.48550/arXiv.2209.01712
https://doi.org/10.48550/arXiv.2010.09885
https://doi.org/10.1038/s42256-020-00284-w
https://doi.org/10.1038/s41467-021-22951-1
https://doi.
https://doi.org/10.
https://doi.org/10.
https://
https://doi.
8 Transformers and Large Language Models for Chemistry and Drug … 161
https://t.me/med1917
45. Toniato A, C. Vaucher A, Schwaller P, Laino T (2023) Enhancing Diversity in Language Based Models for Single-Step Retrosynthesis. Digit Discov 2:489–501.
D00110A
46. Thakkar A, Vaucher AC, Byekwaso A, Schwaller P, Toniato A, Laino T (2023) Unbiasing Retrosynthesis Language Models with Disconnection Prompts. ACS Cent Sci 9:1488–1498.
https://doi.org/10.1021/acscentsci.3c00372
47. Jablonka KM, Ai Q, Al-Feghali A, Badhwar S, Bocarsly JD, Bran AM, Bringuier S, Brinson LC, Choudhary K, Circi D, Cox S, de Jong WA, Evans ML, Gastellu N, Genzling J, Gil MV, Gupta AK, Hong Z, Imran A, Kruschwitz S, Labarre A, Lála J, Liu T, Ma S, Majumdar S, Merz GW, Moitessier N, Moubarak E, Mouriño B, Pelkie B, Pieler M, Ramos MC, Ranković B, Rodriques SG, Sanders JN, Schwaller P, Schwarting M, Shi J, Smit B, Smith BE, Van Herck J, Völker C, Ward L, Warren S, Weiser B, Zhang S, Zhang X, Zia GA, Scourtas A, Schmidt KJ, Foster I, White AD, Blaiszik B (2023) 14 Examples of How LLMs Can Transform Materials Science and Chemistry: A Reflection on a Large Language Model Hackathon. Digital Discovery 2:1233-1250.
48. Tu Z, Coley CW (2021) Permutation Invariant Graph-To-Sequence Model for Template-Free Retrosynthesis and Reaction Prediction.
49. Mikolov T, Sutskever I, Chen K, Corrado G, Dean J (2013) Distributed Representations of Words and Phrases and their Compositionality. ArXiv.
4546
50. Duvenaud D, Maclaurin D, Aguilera-Iparraguirre J, Gómez-Bombarelli R, Hirzel T, Aspuru-Guzik A, Adams RP (2015) Convolutional Networks on Graphs for Learning Molecular Fingerprints. In Proceedings of the Neural Information Processing Systems (NeurIPS 2015).
844bb94c-Abstract.html
51. Wang S, Guo Y, Wang Y, Sun H, Huang J (2019) SMILES-BERT: Large Scale Unsupervised Pre-Training for Molecular Property Prediction. In: Proceedings of the 10th ACM International Conference on Bioinformatics, Computational Biology and Health Informatics. ACM, Niagara Falls NY USA, pp 429–436
52. Schwaller P, Vaucher AC, Laplaza R, Bunne C, Krause A, Corminboeuf C, Laino T (2022) Machine Intelligence for Chemical Reaction Space. WIREs Comput Mol Sci 12:e1604.
doi.org/10.1002/wcms.1604
53. Neves P, McClure K, Verhoeven J, Dyubankova N, Nugmanov R, Gedich A, Menon S, Shi Z, Wegner JK (2023) Global Reactivity Models are Impactful in Industrial Synthesis Applications. J Cheminformatics 15:20.
54. Schwaller P, Vaucher AC, Laino T, Reymond J-L (2021) Prediction of Chemical Reaction Yields using Deep Learning. Mach Learn Sci Technol 2:015016.
abc81d
55. Ross J, Belgodere B, Chenthamarakshan V, Padhi I, Mroueh Y, Das P (2022) Large-Scale Chemical Language Representations Capture Molecular Structure and Properties
56. Wu F, Radev D, Li SZ (2023) Molformer: Motif-based Transformer on 3D Heterogeneous Molecular Graphs.
57. Rives A, Meier J, Sercu T, Goyal S, Lin Z, Liu J, Guo D, Ott M, Zitnick CL, Ma J, Fergus R (2021) Biological Structure and Function Emerge from Scaling Unsupervised Learning To 250 Million Protein Sequences. Proc Natl Acad Sci 118:e2016239118.
pnas.2016239118
58. Lin Z, Akin H, Rao R, Hie B, Zhu Z, Lu W, Smetanin N, Verkuil R, Kabeli O, Shmueli Y, dos Santos Costa A, Fazel-Zarandi M, Sercu T, Candido S, Rives A (2023) Evolutionary-Scale Prediction of Atomic-Level Protein Structure with a Language Model. Science 379:1123–1130.
https://doi.org/10.1126/science.ade2574
59. Verkuil R, Kabeli O, Du Y, Wicky BIM, Milles LF, Dauparas J, Baker D, Ovchinnikov S, Sercu T, Rives A (2022) Language Models Generalize Beyond Natural Proteins. 2022.12.21.521521.
https://doi.org/10.1101/2022.12.21.521521
https://doi.org/10.1039/D3DD00113J
https://doi.org/10.48550/arXiv.2110.09681
https://proceedings.neurips.cc/paper/2015/hash/f9be311e65d81a9ad8150a60
https://doi.org/10.1186/s13321-023-00685-0
https://doi.org/10.48550/arXiv.2110.01191
https://doi.org/10.1039/D2D
https://doi.org/10.48550/arXiv.1310.
https://
https://doi.org/10.1088/2632-2153/
https://doi.org/10.1073/
162 A. M. Bran and P. Schwaller
https://t.me/med1917
60. Teukam YGN, Dassi LK, Manica M, Probst D, Laino T Language Models can Identify Enzymatic Active Sites in Protein Sequences.
gg-v3
61. Edwards C, Lai T, Ros K, Honke G, Cho K, Ji H (2022) Translation between Molecules and Natural Language.
62. Christofidellis D, Giannone G, Born J, Winther O, Laino T, Manica M (2023) Unifying Molecular and Textual Representations via Multi-task Language Modelling
63. Alberts M, Laino T, Vaucher AC (2023) Leveraging Infrared Spectroscopy for Automated Structure Elucidation. ChemRxiv.
64. Raschka S (2023) Finetuning Large Language Models. https://magazine.sebastianraschka.com/
p/finetuning-large-language-models?utm_campaign=post. Accessed 17 May 2023
65. Zhang R, Han J, Zhou A, Hu X, Yan S, Lu P, Li H, Gao P, Qiao Y (2023) LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.
arXiv.2303.16199
66. Hoffmann J, Borgeaud S, Mensch A, Buchatskaya E, Cai T, Rutherford E, Casas D de L, Hendricks LA, Welbl J, Clark A, Hennigan T, Noland E, Millican K, Driessche G van den, Damoc B, Guy A, Osindero S, Simonyan K, Elsen E, Rae JW, Vinyals O, Sifre L (2022) Training Compute-Optimal Large Language Models.
67. Wei J, Tay Y, Bommasani R, Raffel C, Zoph B, Borgeaud S, Yogatama D, Bosma M, Zhou D, Metzler D, Chi EH, Hashimoto T, Vinyals O, Liang P, Dean J, Fedus W (2022) Emergent Abilities of Large Language Models.
68. Wei J, Wang X, Schuurmans D, Bosma M, Ichter B, Xia F, Chi E, Le Q, Zhou D (2023) Chain­of-Thought Prompting Elicits Reasoning in Large Language Models.
arXiv.2201.11903
69. Ouyang L, Wu J, Jiang X, Almeida D, Wainwright CL, Mishkin P, Zhang C, Agarwal S, Slama K, Ray A, Schulman J, Hilton J, Kelton F, Miller L, Simens M, Askell A, Welinder P, Christiano P, Leike J, Lowe R (2022) Training Language Models to Follow Instructions with Human Feedback.
70. Jablonka KM, Schwaller P, Ortega-Guerrero A, Smit B (2023) Leveraging large language models for predictive chemistry. Nat. Mach. Intell.
88-1
71. Ramos MC, Michtavy SS, Porosoff MD, White AD (2023) Bayesian Optimization of Catalysts With In-context Learning.
72. Boiko DA, MacKnight R, Kline B, Gomes G (2023) Autonomous chemical research with large language models. Nature.
73. Bran AM, Cox S, Schilter O, Baldassari C, White AD, Schwaller P (2024) Augmenting Large­Language Models with Chemistry Tools. Nat. Mach. Intell.
024-00832-8
74. Howard J, Ruder S (2018) Universal Language Model Fine-tuning for Text Classification. In: Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, Melbourne, Australia, pp 328–339
75. Wei J, Bosma M, Zhao V, Guu K, Yu AW, Lester B, Du N, Dai AM, Le QV (2021) Finetuned Language Models are Zero-Shot Learners.
76. Yin X, Chen W, Wu X, Yue H (2017) Fine-tuning and Visualization of Convolutional Neural Networks. In: 2017 12th IEEE Conference on Industrial Electronics and Applications (ICIEA). pp 1310–1315
77. Dai H, Li C, Coley CW, Dai B, Song L (2020) Retrosynthesis Prediction with Conditional Graph Logic Network.
78. Zhang B, Zhang X, Du W, Song Z, Zhang G, Zhang G, Wang Y, Chen X, Jiang J, Luo Y (2022) Chemistry-Informed Molecular Graph as Reaction Descriptor for Machine-Learned Retrosynthesis Planning. Proc Natl Acad Sci 119:e2212711119.
2212711119
79. Jorner K, Turcani L (2022) kjelljorner/morfeus: v0.7.2
https://doi.org/10.48550/arXiv.2204.11817
https://doi.org/10.26434/chemrxiv-2023-5v27f
https://doi.org/10.48550/ARXIV.2206.07682
https://doi.org/10.48550/arXiv.2203.02155
https://doi.org/10.48550/arXiv.2304.05341
https://doi.org/10.1038/s41586-023-06792-0.
https://doi.org/10.48550/arXiv.2001.01408
https://doi.org/10.26434/chemrxiv-2021-m20
https://doi.org/10.48550/
https://doi.org/10.48550/arXiv.2203.15556
https://doi.org/10.48550/
https://doi.org/10.1038/s42256-023-007
https://doi.org/10.1038/s42256-
https://doi.org/10.48550/arXiv.2109.01652
https://doi.org/10.1073/pnas.
8 Transformers and Large Language Models for Chemistry and Drug … 163
https://t.me/med1917
80. Ranković B, Griffiths R-R, Moss HB, Schwaller P (2023) Bayesian Optimisation for Addi- tive Screening and Yield Improvements in Chemical Reactions – Beyond One-Hot Encoding. Digital Discovery, 2024, Advance Article.
81. Shields BJ, Stevens J, Li J, Parasram M, Damani F, Alvarado JIM, Janey JM, Adams RP, Doyle AG (2021) Bayesian Reaction Optimization as a Tool for Chemical Synthesis. Nature 590:89–96.
82. Bagal V, Aggarwal R, Vinod PK, Priyakumar UD (2022) MolGPT: Molecular Generation Using a Transformer-Decoder Model. J Chem Inf Model 62:2064–2076.
jcim.1c00600
83. Rothchild D, Tamkin A, Yu J, Misra U, Gonzalez J (2021) C5T5: Controllable Generation of Organic Molecules with Transformers.
84. Wang W, Wang Y, Zhao H, Sciabola S (2022) A Transformer-based Generative Model for De Novo Molecular Design.
85. Bengio E, Jain M, Korablyov M, Precup D, Bengio Y (2021) Flow Network based Generative Models for Non-Iterative Diverse Candidate Generation.
04399
86. Born J, Manica M (2023) Regression Transformer Enables Concurrent Sequence Regression and Generation For Molecular Language Modelling. Nat Mach Intell 5:432–444.
org/10.1038/s42256-023-00639-z
87. Flam-Shepherd D, Aspuru-Guzik A (2023) Language Models can Generate Molecules, Mate­rials, and Protein Binding Sites Directly in Three Dimensions as XYZ, CIF, and PDB files.
https://doi.org/10.48550/arXiv.2305.05708
88. Schick T, Dwivedi-Yu J, Dessì R, Raileanu R, Lomeli M, Zettlemoyer L, Cancedda N, Scialom T (2023) Toolformer: Language Models Can Teach Themselves to Use Tools.
10.48550/arXiv.2302.04761
89. Karpas E, Abend O, Belinkov Y, Lenz B, Lieber O, Ratner N, Shoham Y, Bata H, Levine Y, Leyton-Brown K, Muhlgay D, Rozen N, Schwartz E, Shachaf G, Shalev-Shwartz S, Shashua A, Tenenholtz M (2022) MRKL Systems: A Modular, Neuro-Symbolic Architecture that Combines Large Language Models, External Knowledge Sources and Discrete Reasoning.
https://doi.org/10.48550/arXiv.2205.00445
90. Yao S, Zhao J, Yu D, Du N, Shafran I, Narasimhan K, Cao Y (2023) ReAct: Synergizing Reasoning and Acting in Language Models.
91. D. White A, M. Hocky G, A. Gandhi H, Ansari M, Cox S, P. Wellawatte G, Sasmal S, Yang Z, Liu K, Singh Y, Ccoa WJP (2023) Assessment of Chemistry Knowledge in Large Language Models that Generate Code. Digit Discov 2:368–376.
92. Jablonka KM, Ai Q, Al-Feghali A, Badhwar S, Bran JDBAM, Bringuier S, Brinson LC, Choud­hary K, Circi D, Cox S, de Jong WA, Evans ML, Gastellu N, Genzling J, Gil MV, Gupta AK, Hong Z, Imran A, Kruschwitz S, Labarre A, Lála J, Liu T, Ma S, Majumdar S, Merz GW, Moitessier N, Moubarak E, Mouriño B, Pelkie B, Pieler M, Ramos MC, Ranković B, Rodriques SG, Sanders JN, Schwaller P, Schwarting M, Shi J, Smit B, Smith BE, Van Heck J, Völker C, Ward L, Warren S, Weiser B, Zhang S, Zhang X, Zia GA, Scourtas A, Schmidt K, Foster I, White AD, Blaiszik B (2023) 14 Examples of How LLMs Can Transform Materials Science and Chemistry: A Reflection on a Large Language Model Hackathon.
ARXIV.2306.06283
93. Su B, Du D, Yang Z, Zhou Y, Li J, Rao A, Sun H, Lu Z, Wen J-R (2022) A Molecular Multimodal Foundation Model Associating Molecule Graphs with Natural Language.
doi.org/10.48550/arXiv.2209.05481
https://doi.org/10.1038/s41586-021-03213-y
https://doi.org/10.48550/arXiv.2210.08749
https://doi.org/10.1039/D3DD00096F
https://doi.org/10.1021/acs.
https://doi.org/10.48550/arXiv.2108.10307
https://doi.org/10.48550/arXiv.2106.
https://doi.
https://doi.org/
https://doi.org/10.48550/arXiv.2210.03629
https://doi.org/10.1039/D2DD00087C
https://doi.org/10.48550/
https://