Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5338_Библиотеки_им_академика_М_И_Перельмана
.pdf
154 A. M. Bran and P. Schwaller
https://t.me/med1917
Fig. 8.3 Recent advances in large language models (LLMs) have launched a new era of Transformers in chemistry. Their capabilities allow for (a) general task solvers thanks to their flexibility
and knowledge transferability [
unlimited modalities, in the form of computational tools [
70, 71]and (b) agent architectures, capable of integrating virtually
47, 72, 73
]
these endeavors aim to make the most of existing knowledge, enabling models to
learn as much as possible from the few available data points.
Interestingly, Jablonka et al. [70
] demonstrated that fine-tuning and in-context
learning can perform on par with, and in some instances even outperform, these
specialized techniques, particularly when data is limited. The high performance of
in-context learning, combined with its ease of use and flexibility, makes it one of the
most impressive applications of LLMs in chemistry to date. This technique holds the
potential to revolutionize the way machine learning is utilized in the scientific field,
by rapidly highlighting complex correlations in data.
Another application of regression, which holds significant interest in chemistry
and drug discovery cycles, is optimization. This process involves modifying an
object until a property of interest reaches a desired value. Typically, this requires
a large number of measurements of the desired property, which can be quite costly
in chemistry use-cases. Applications of this include yield/selectivity optimization

8 Transformers and Large Language Models for Chemistry and Drug … 155
https://t.me/med1917
in chemical reactions and the generation of molecular candidates with target properties. Bayesian Optimization (BO) has recently been proposed as a solution to
such problems in chemistry [
However, BO requires uncertainty-calibrated regression methods, which sets it apart
from conventional regression.
In line with the concept of in-context learning, Ramos et al. [71] proposed a
system that utilizes GPT models to perform regression while also incorporating
uncertainty. This approach enables BO without the need for any feature engineering
or fine-tuning. The flexibility of this method allowed the team to perform catalyst and
molecular optimization using only the synthesis procedure of the catalyst as input.
This work represents a paradigm shift in drug discovery and molecular design. For
the first time, it showcases a direct map from the synthesis procedure into property
space, effectively overcoming issues like the synthesizability of proposed molecules,
a key limitation of structure-based generative models.
8.3.2.2 Molecular Generation
Another fascinating application of the generative capabilities of language models
is molecular generation. This area, which is of significant importance in the drug
discovery process, has been largely dominated by models that generate molecules in
the form of linear string representations, such as SMILES or SELFIES [
86], its successful application is contingent on the ability to specify substances as
graphs and their subsequent conversion to a linear string representation. However,
this approach is only suitable for a subset of organic molecules. Other substances,
such as macromolecules and materials, necessitate more comprehensive representations. A complete and accurate representation of these substances can only be
achieved by specifying atomic positions, boundary conditions, and other factors. This
requirement presents a significant challenge and limitation to the current methods
of molecular generation. To address these limitations, Flam et al. [
using language models for structure generation, directly generating them with threedimensional atomic positions. Besides being innovative and valid, the generated
structures can be obtained by training models in a variety of formats used for crystals,
proteins, and more. This work also demonstrates performance comparable to expertdesigned, state-of-the-art algorithms for molecular generation based on graphs, while
overcoming the limitations mentioned earlier.
80, 81], particularly in situations where data is small.
82–84
]. While
70, 85,
87
] proposed
8.3.3 Language Model-Powered Agents
Among the most useful emergent abilities of language models are step-by-step
reasoning, activated through chain-of-thought (CoT) prompting, and their capacity
to effectively use tools [
88]. These capabilities have been the subject of extensive

156 A. M. Bran and P. Schwaller
https://t.me/med1917
research in recent years, and their application has been shown to significantly enhance
the performance of LLMs across a variety of tasks. CoT prompting is a technique
where language models are instructed to solve a task by following a sequence of
reasoning steps, rather than providing an answer in a single response [
language models in this way effectively allows them to perform symbolic operations,
much like humans perform arithmetic operations by keeping track of intermediate
steps.
The ability to use tools is another significant capability of language models [88].
This allows them to invoke external computational tools, thereby enriching their
knowledge through querying search engines, accessing calculators, and so on [
These capabilities have been demonstrated to enhance the performance of large
language models in a range of tasks that were previously inaccessible.
The recent advancements and results in revealing and exploiting the capabilities
of LLMs suggest the potential for combining some of these capabilities to create
more powerful and useful possibilities. This concept has been recently explored,
leading to the development of the Modular Reasoning, Knowledge and Language
(MRKL [
tool-using capabilities of modern LLMs. By incorporating external tools into a CoT
setting, agents of this type have recently been shown to outperform other methods
based on large language models.
unimodality issue of LLMs. Under this setting, they become capable of processing
different types of input data, making real-time decisions in simulated environments,
and even interacting with real-world robotic platforms. The solutions provided by
LLMs to tasks also become more grounded in reality, as access to certain tools
provides them with real, up-to-date information relevant to the task. This can, to some
extent, limit the tendency of these models to generate unrealistic or “hallucinated”
responses.
89]) and Reason + Act (ReAct [90]) systems, which combine the CoT and
One direct benefit of effective tool usage is that it partially overcomes the
68]. Instructing
88].
8.3.3.1 Agents in Chemistry: Unleashing the Power from Tools
Despite their strengths as text generators and task solvers, and their remarkable fewshot and zero-shot performance, these models are also well known for their high
propensity to generate false and inaccurate content, an issue that extends to easily
verifiable matters such as basic arithmetic [
limitations make the direct application of LLMs to chemistry a challenging matter.
The potential applications of large language models in chemistry were first
explored in a large-scale collaboration involving researchers from around the world,
an effort that resulted in the demonstration of 14 use-cases [
range from wrappers for computational tools, which enhance their accessibility by
allowing natural language inputs to modify behaviors, to assistants for reaction optimization, and knowledge parsers and synthesizers for scientific question answering,
among others. These are just a few of the possibilities that LLMs offer in chemistry,
88] and chemical operations [91]. These
]. The applications
92

8 Transformers and Large Language Models for Chemistry and Drug … 157
https://t.me/med1917
which, when combined with existing chemistry tools and databases, significantly
increase the applicability and accessibility of computational applications.
More recently, Bran and Cox et al. [73] extended the concept of LLM-powered
agents for chemistry by curating and compiling a set of computational chemistry
tools. Their system, ChemCrow, has been shown to be capable of planning and
executing tasks in chemistry, effectively streamlining the reasoning process for
several common chemical tasks across areas such as drug and materials design and
synthesis. The authors demonstrate that this approach has a highly positive effect on
LLM’s performance for tasks in chemistry, overcoming hallucinations and grounding
their responses with data from reliable sources. Complementary approaches exist,
72
with a sharper focus on cloud lab operability [
The power of platforms like ChemCrow extends beyond merely serving as independent task solvers. They can be viewed as general chemistry assistants with the
ultimate goal of making computational tools more accessible to chemists, thereby
accelerating discovery. An additional highlight is the seamless exploitation of tool
composability that this allows. It makes it straightforward to enrich the results of one
tool with another, or to construct custom tool pipelines, all through simple requests
in natural language.
].
8.4 Outlook and Final Remarks
Advancements in neural translation models, and specially with the introduction of the
Transformer architecture, have sparked a revolution in machine learning for applications in chemistry and drug development. Analogies between chemical and natural
language, and the publication of open databases and benchmarks, have inspired
the representation of chemical tasks in the form of text, allowing straightforward
application of Transformers to problems in this field.
This revolution has occurred in three stages, differentiated by the specificity of
tasks. In a first stage, characterized by task-specificity and use of single-modality
models, applications spanned molecule-to-molecule conversion tasks, like reaction
outcome prediction and retrosynthetic planning, along with representation learning
and downstream tasks like regression and classification. Their excellent performance
and relative simplicity made them de-facto models in an array of applications.
In a second stage, researchers attempted to connect multiple additional modalities relevant to chemistry, like spectra from experiments, sequences of experimental
actions, and even natural language, opening the way for an expanded number of applications involving modalities of any sort, however still task-specific. More recently,
powered by vertiginous advancements in training and tuning of large language
models, a series of works have been published that leverage a number of capabilities from such models. Among others, these contributions showcase applications
in regression, classification, molecular generation and reaction optimization, all with
unprecedented flexibility and usually improved performance over other methods.

158 A. M. Bran and P. Schwaller
https://t.me/med1917
Another direction explores the integration of virtually unlimited modalities —in
the form of tools—into agents powered by LLMs. The power of these agents has been
demonstrated through a number of diverse tasks, ranging from molecular generation
to automated organic synthesis, in an open-ended, highly customizable fashion.
By leveraging the expressivity and flexibility of natural language, this last wave
of applications aims to bridge the gap between the chemical and natural languages.
As we continue to explore and harness these capabilities, we can look forward to
a future where machine learning plays an even more integral role in accelerating
scientific discovery.
References
1. Bahdanau D, Cho K, Bengio Y (2016) Neural Machine Translation by Jointly Learning to
Align and Translate.
2. Cho K, van Merrienboer B, Gulcehre C, Bahdanau D, Bougares F, Schwenk H, Bengio Y
(2014) Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine
Translation.
3. Hochreiter S, Schmidhuber J (1997) Long Short-Term Memory. Neural Comput 9:1735–1780.
https://doi.org/10.1162/neco.1997.9.8.1735
4. Sutskever I, Vinyals O, Le QV (2014) Sequence to Sequence Learning with Neural Networks.
https://doi.org/10.48550/arXiv.1409.3215
5. Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser L, Polosukhin I
(2017) Attention Is All You Need.
6. Bommasani R, Hudson DA, Adeli E, Altman R, Arora S, von Arx S, Bernstein MS, Bohg J,
Bosselut A, Brunskill E, Brynjolfsson E, Buch S, Card D, Castellon R, Chatterji N, Chen A,
Creel K, Davis JQ, Demszky D, Donahue C, Doumbouya M, Durmus E, Ermon S, Etchemendy
J, Ethayarajh K, Fei-Fei L, Finn C, Gale T, Gillespie L, Goel K, Goodman N, Grossman S,
Guha N, Hashimoto T, Henderson P, Hewitt J, Ho DE, Hong J, Hsu K, Huang J, Icard T, Jain
S, Jurafsky D, Kalluri P, Karamcheti S, Keeling G, Khani F, Khattab O, Koh PW, Krass M,
Krishna R, Kuditipudi R, Kumar A, Ladhak F, Lee M, Lee T, Leskovec J, Levent I, Li XL, Li
X, Ma T, Malik A, Manning CD, Mirchandani S, Mitchell E, Munyikwa Z, Nair S, Narayan
A, Narayanan D, Newman B, Nie A, Niebles JC, Nilforoshan H, Nyarko J, Ogut G, Orr L,
Papadimitriou I, Park JS, Piech C, Portelance E, Potts C, Raghunathan A, Reich R, Ren H, Rong
F, Roohani Y, Ruiz C, Ryan J, Ré C, Sadigh D, Sagawa S, Santhanam K, Shih A, Srinivasan
K, Tamkin A, Taori R, Thomas AW, Tramèr F, Wang RE, Wang W, Wu B, Wu J, Wu Y, Xie
SM, Yasunaga M, You J, Zaharia M, Zhang M, Zhang T, Zhang X, Zhang Y, Zheng L, Zhou
K, Liang P (2022) On the Opportunities and Risks of Foundation Models.
48550/arXiv.2108.07258
7. Bubeck S, Chandrasekaran V, Eldan R, Gehrke J, Horvitz E, Kamar E, Lee P, Lee YT, Li Y,
Lundberg S, Nori H, Palangi H, Ribeiro MT, Zhang Y (2023) Sparks of Artificial General
Intelligence: Early Experiments with GPT-4.
8. Gao L, Biderman S, Black S, Golding L, Hoppe T, Foster C, Phang J, He H, Thite A, Nabeshima
N, Presser S, Leahy C (2020) The Pile: An 800GB Dataset of Diverse Text for Language
Modeling.
9. Raffel C, Shazeer N, Roberts A, Lee K, Narang S, Matena M, Zhou Y, Li W, Liu PJ (2020)
doi.org/10.48550/arXiv.1910.10683
10. Devlin J, Chang M-W, Lee K, Toutanova K (2019) BERT: Pre-training of Deep Bidirectional
Transformers for Language Understanding.
https://doi.org/10.48550/arXiv.2101.00027
https://doi.org/10.48550/arXiv.1409.0473
https://doi.org/10.48550/arXiv.1406.1078
https://doi.org/10.48550/arXiv.1706.03762
https://doi.org/10.
https://doi.org/10.48550/arXiv.2303.12712
https://
https://doi.org/10.48550/arXiv.1810.04805

8 Transformers and Large Language Models for Chemistry and Drug … 159
https://t.me/med1917
11. Reimers N, Gurevych I (2019) Sentence-BERT: Sentence Embeddings using Siamese BERTNetworks.
12. Kryscinski W, Keskar NS, McCann B, Xiong C, Socher R (2019) Neural Text Summarization: A Critical Evaluation. In: Proceedings of the 2019 Conference on Empirical Methods in
Natural Language Processing and the 9th International Joint Conference on Natural Language
Processing (EMNLP-IJCNLP). Association for Computational Linguistics, Hong Kong, China,
pp 540–551
13. Liu PJ, Saleh M, Pot E, Goodrich B, Sepassi R, Kaiser L, Shazeer N (2018) Generating
Wikipedia by Summarizing Long Sequences.
14. OpenAI (2023) GPT-4 Technical Report. https://doi.org/10.48550/arXiv.2303.08774
15. Jumper J, Evans R, Pritzel A, Green T, Figurnov M, Ronneberger O, Tunyasuvunakool K, Bates
R, Žídek A, Potapenko A, Bridgland A, Meyer C, Kohl SAA, Ballard AJ, Cowie A, RomeraParedes B, Nikolov S, Jain R, Adler J, Back T, Petersen S, Reiman D, Clancy E, Zielinski M,
Steinegger M, Pacholska M, Berghammer T, Bodenstein S, Silver D, Vinyals O, Senior AW,
Kavukcuoglu K, Kohli P, Hassabis D (2021) Highly Accurate Protein Structure Prediction with
AlphaFold. Nature 596:583–589.
16. Akdel M, Pires DEV, Pardo EP, Jänes J, Zalevsky AO, Mészáros B, Bryant P, Good LL,
Laskowski RA, Pozzati G, Shenoy A, Zhu W, Kundrotas P, Serra VR, Rodrigues CHM, Dunham
AS, Burke D, Borkakoti N, Velankar S, Frost A, Basquin J, Lindorff-Larsen K, Bateman A,
Kajava AV, Valencia A, Ovchinnikov S, Durairaj J, Ascher DB, Thornton JM, Davey NE, Stein
A, Elofsson A, Croll TI, Beltrao P (2022) A Structural Biology Community Assessment of
AlphaFold2 Applications. Nat Struct Mol Biol 29:1056–1067.
022-00849-w
17. Yang Z, Zeng X, Zhao Y, Chen R (2023) AlphaFold2 and its Applications in the Fields of
Biology and Medicine. Signal Transduct Target Ther 8:1–14.
023-01381-z
18. Huang K, Fu T, Gao W, Zhao Y, Roohani Y, Leskovec J, Coley C, Xiao C, Sun J, Zitnik M (2021)
Therapeutics Data Commons: Machine Learning Datasets and Tasks for Drug Discovery and
Development. Proc Neural Inf Process Syst Track Datasets Benchmarks 1
19. Kearnes SM, Maser MR, Wleklinski M, Kast A, Doyle AG, Dreher SD, Hawkins JM, Jensen
KF, Coley CW (2021) The Open Reaction Database. J Am Chem Soc 143:18820–18826.
https://doi.org/10.1021/jacs.1c09820
20. Lowe D (2012) Extraction of Chemical Structures and Reactions from the Literature. University
of Cambridge.
21. Wu Z, Ramsundar B, Feinberg EN, Gomes J, Geniesse C, Pappu AS, Leswing K, Pande V
(2018) MoleculeNet: A Benchmark for Molecular Machine Learning. Chem Sci 9:513–530.
https://doi.org/10.1039/C7SC02664A
22. Cadeddu A, Wylie EK, Jurczak J, Wampler-Doty M, Grzybowski BA (2014) Organic Chemistry
as a Language and the Implications of Chemical Linguistics for Structural and Retrosynthetic
Analyses. Angew Chem Int Ed 53:8108–8112.
23. Wołos A, Koszelewski D, Roszak R, Szymkuć S, Moskal M, Ostaszewski R, Herrera BT,
Maier JM, Brezicki G, Samuel J, Lummiss JAM, McQuade DT, Rogers L, Grzybowski BA
(2022) Computer-Designed Repurposing of Chemical Wastes into Drugs. Nature 604:668–676.
https://doi.org/10.1038/s41586-022-04503-9
24. Schwaller P, Hoover B, Reymond J-L, Strobelt H, Laino T (2021) Extraction of Organic Chemistry Grammar from Unsupervised Learning of Chemical Reactions. Sci Adv 7:eabe4166.
https://doi.org/10.1126/sciadv.abe4166
25. Brammer JC, Blanke G, Kellner C, Hoffmann A, Herres-Pawlis S, Schatzschneider U (2022)
TUCAN: A Molecular Identifier and Descriptor Applicable to the Whole Periodic Table from
Hydrogen to Oganesson. J Cheminformatics 14:66.
40-5
26. Heller SR, McNaught A, Pletnev I, Stein S, Tchekhovskoi D (2015) InChI, the IUPAC International Chemical Identifier. J Cheminformatics 7:23.
0068-4
https://doi.org/10.48550/arXiv.1908.10084
https://doi.org/10.48550/arXiv.1801.10198
https://doi.org/10.1038/s41586-021-03819-2
https://doi.org/10.1038/s41594-
https://doi.org/10.1038/s41392-
https://doi.org/10.17863/CAM.16293
https://doi.org/10.1002/anie.201403708
https://doi.org/10.1186/s13321-022-006
https://doi.org/10.1186/s13321-015-

160 A. M. Bran and P. Schwaller
https://t.me/med1917
27. Krenn M, Ai Q, Barthel S, Carson N, Frei A, Frey NC, Friederich P, Gaudin T, Gayle AA,
Jablonka KM, Lameiro RF, Lemm D, Lo A, Moosavi SM, Nápoles-Duarte JM, Nigam A,
Pollice R, Rajan K, Schatzschneider U, Schwaller P, Skreta M, Smit B, Strieth-Kalthoff F,
Sun C, Tom G, von Rudorff GF, Wang A, White A, Young A, Yu R, Aspuru-Guzik A (2022)
SELFIES and the Future of Molecular String R epresentations. Patterns 3:100588.
org/10.1016/j.patter.2022.100588
28. Krenn M, Häse F, Nigam A, Friederich P, Aspuru-Guzik A (2020) Self-referencing Embedded
Strings (SELFIES): A 100% Robust Molecular String Representation. Mach Learn Sci Technol
1:045024.
29. O’Boyle N, Dalke A (2018) DeepSMILES: An Adaptation of SMILES for Use in MachineLearning of Chemical Structures.
30. Weininger D (1988) SMILES, a Chemical Language and Information System. 1. Introduction
to Methodology and Encoding Rules. J Chem Inf Comput Sci 28:31–36.
1021/ci00057a005
31. Restrepo G (2022) Chemical Space: Limits, Evolution and Modelling of an Object Bigger than
our Universal Library. Digit Discov 1:568–585.
32. Gómez-Bombarelli R, Wei JN, Duvenaud D, Hernández-Lobato JM, Sánchez-Lengeling B,
Sheberla D, Aguilera-Iparraguirre J, Hirzel TD, Adams RP, Aspuru-Guzik A (2018) Automatic
Chemical Design Using a Data-Driven Continuous Representation of Molecules. ACS Cent
Sci 4:268–276.
33. Kusner MJ, Paige B, Hernández-Lobato JM (2017) Grammar Variational Autoencoder. https://
doi.org/10.48550/arXiv.1703.01925
34. Öztürk H, Özgür A, Schwaller P, Laino T, Ozkirimli E (2020) Exploring Chemical Space
Using Natural Language Processing Methodologies for Drug Discovery. Drug Discov Today
25:689–705.
35. Pesciullesi G, Schwaller P, Laino T, Reymond J-L (2020) Transfer Learning Enables the
Molecular Transformer to Predict Regio- and Stereoselective Reactions on Carbohydrates.
Nat Commun 11:4874.
36. Schwaller P, Laino T, Gaudin T, Bolgar P, Hunter CA, Bekas C, Lee AA (2019) Molecular
Transformer: A Model for Uncertainty-Calibrated Chemical Reaction Prediction. ACS Cent
Sci 5:1572–1583.
37. Schwaller P, Petraglia R, Zullo V, Nair VH, Haeuselmann RA, Pisoni R, Bekas C, Iuliano A,
Laino T (2020) Predicting Retrosynthetic Pathways using Transformer-Based Models and a
Hyper-Graph Exploration Strategy. Chem Sci 11:3316–3325.
704H
38. Ahmad W, Simon E, Chithrananda S, Grand G, Ramsundar B (2022) ChemBERTa-2: Towards
Chemical Foundation Models.
39. Chithrananda S, Grand G, Ramsundar B (2020) ChemBERTa: Large-Scale Self-Supervised
Pretraining for Molecular Property Prediction.
40. Li J, Jiang X (2021) Mol-BERT: An Effective Molecular Representation with BERT for Molecular Property Prediction. Wirel Commun Mob Comput 2021:e7181815.
1155/2021/7181815
41. Schwaller P, Probst D, Vaucher AC, Nair VH, Kreutter D, Laino T, Reymond J-L (2021)
Mapping the Space of Chemical Reactions Using Attention-Based Neural Networks. Nat Mach
Intell 3:144–152.
42. Vaucher AC, Schwaller P, Geluykens J, Nair VH, Iuliano A, Laino T (2021) Inferring Experimental Procedures from Text-Based Representations of Chemical Reactions. Nat Commun
12:2573.
43. Vaucher AC, Zipoli F, Geluykens J, Nair VH, Schwaller P, Laino T (2020) Automated Extraction
of Chemical Synthesis Actions from Experimental Procedures. Nat Commun 11:3601.
doi.org/10.1038/s41467-020-17266-6
44. Tetko IV, Karpov P, Van Deursen R, Godin G (2020) State-of-the-art Augmented NLP Transformer Models for Direct and Single-Step Retrosynthesis. Nat Commun 11:5575.
org/10.1038/s41467-020-19266-y
https://doi.org/10.1088/2632-2153/aba947
https://doi.org/10.26434/chemrxiv.7097960.v1
https://doi.org/10.1039/D2DD00030J
https://doi.org/10.1021/acscentsci.7b00572
https://doi.org/10.1016/j.drudis.2020.01.020
https://doi.org/10.1038/s41467-020-18671-7
https://doi.org/10.1021/acscentsci.9b00576
https://doi.org/10.1039/C9SC05
https://doi.org/10.48550/arXiv.2209.01712
https://doi.org/10.48550/arXiv.2010.09885
https://doi.org/10.1038/s42256-020-00284-w
https://doi.org/10.1038/s41467-021-22951-1
https://doi.
https://doi.org/10.
https://doi.org/10.
https://
https://doi.

8 Transformers and Large Language Models for Chemistry and Drug … 161
https://t.me/med1917
45. Toniato A, C. Vaucher A, Schwaller P, Laino T (2023) Enhancing Diversity in Language Based
Models for Single-Step Retrosynthesis. Digit Discov 2:489–501.
D00110A
46. Thakkar A, Vaucher AC, Byekwaso A, Schwaller P, Toniato A, Laino T (2023) Unbiasing
Retrosynthesis Language Models with Disconnection Prompts. ACS Cent Sci 9:1488–1498.
https://doi.org/10.1021/acscentsci.3c00372
47. Jablonka KM, Ai Q, Al-Feghali A, Badhwar S, Bocarsly JD, Bran AM, Bringuier S, Brinson
LC, Choudhary K, Circi D, Cox S, de Jong WA, Evans ML, Gastellu N, Genzling J, Gil MV,
Gupta AK, Hong Z, Imran A, Kruschwitz S, Labarre A, Lála J, Liu T, Ma S, Majumdar S,
Merz GW, Moitessier N, Moubarak E, Mouriño B, Pelkie B, Pieler M, Ramos MC, Ranković
B, Rodriques SG, Sanders JN, Schwaller P, Schwarting M, Shi J, Smit B, Smith BE, Van Herck
J, Völker C, Ward L, Warren S, Weiser B, Zhang S, Zhang X, Zia GA, Scourtas A, Schmidt KJ,
Foster I, White AD, Blaiszik B (2023) 14 Examples of How LLMs Can Transform Materials
Science and Chemistry: A Reflection on a Large Language Model Hackathon. Digital Discovery
2:1233-1250.
48. Tu Z, Coley CW (2021) Permutation Invariant Graph-To-Sequence Model for Template-Free
Retrosynthesis and Reaction Prediction.
49. Mikolov T, Sutskever I, Chen K, Corrado G, Dean J (2013) Distributed Representations of
Words and Phrases and their Compositionality. ArXiv.
4546
50. Duvenaud D, Maclaurin D, Aguilera-Iparraguirre J, Gómez-Bombarelli R, Hirzel T,
Aspuru-Guzik A, Adams RP (2015) Convolutional Networks on Graphs for Learning
Molecular Fingerprints. In Proceedings of the Neural Information Processing Systems
(NeurIPS 2015).
844bb94c-Abstract.html
51. Wang S, Guo Y, Wang Y, Sun H, Huang J (2019) SMILES-BERT: Large Scale Unsupervised
Pre-Training for Molecular Property Prediction. In: Proceedings of the 10th ACM International
Conference on Bioinformatics, Computational Biology and Health Informatics. ACM, Niagara
Falls NY USA, pp 429–436
52. Schwaller P, Vaucher AC, Laplaza R, Bunne C, Krause A, Corminboeuf C, Laino T (2022)
Machine Intelligence for Chemical Reaction Space. WIREs Comput Mol Sci 12:e1604.
doi.org/10.1002/wcms.1604
53. Neves P, McClure K, Verhoeven J, Dyubankova N, Nugmanov R, Gedich A, Menon S, Shi Z,
Wegner JK (2023) Global Reactivity Models are Impactful in Industrial Synthesis Applications.
J Cheminformatics 15:20.
54. Schwaller P, Vaucher AC, Laino T, Reymond J-L (2021) Prediction of Chemical Reaction Yields
using Deep Learning. Mach Learn Sci Technol 2:015016.
abc81d
55. Ross J, Belgodere B, Chenthamarakshan V, Padhi I, Mroueh Y, Das P (2022) Large-Scale
Chemical Language Representations Capture Molecular Structure and Properties
56. Wu F, Radev D, Li SZ (2023) Molformer: Motif-based Transformer on 3D Heterogeneous
Molecular Graphs.
57. Rives A, Meier J, Sercu T, Goyal S, Lin Z, Liu J, Guo D, Ott M, Zitnick CL, Ma J, Fergus
R (2021) Biological Structure and Function Emerge from Scaling Unsupervised Learning To
250 Million Protein Sequences. Proc Natl Acad Sci 118:e2016239118.
pnas.2016239118
58. Lin Z, Akin H, Rao R, Hie B, Zhu Z, Lu W, Smetanin N, Verkuil R, Kabeli O, Shmueli Y,
dos Santos Costa A, Fazel-Zarandi M, Sercu T, Candido S, Rives A (2023) Evolutionary-Scale
Prediction of Atomic-Level Protein Structure with a Language Model. Science 379:1123–1130.
https://doi.org/10.1126/science.ade2574
59. Verkuil R, Kabeli O, Du Y, Wicky BIM, Milles LF, Dauparas J, Baker D, Ovchinnikov S, Sercu
T, Rives A (2022) Language Models Generalize Beyond Natural Proteins. 2022.12.21.521521.
https://doi.org/10.1101/2022.12.21.521521
https://doi.org/10.1039/D3DD00113J
https://doi.org/10.48550/arXiv.2110.09681
https://proceedings.neurips.cc/paper/2015/hash/f9be311e65d81a9ad8150a60
https://doi.org/10.1186/s13321-023-00685-0
https://doi.org/10.48550/arXiv.2110.01191
https://doi.org/10.1039/D2D
https://doi.org/10.48550/arXiv.1310.
https://
https://doi.org/10.1088/2632-2153/
https://doi.org/10.1073/

162 A. M. Bran and P. Schwaller
https://t.me/med1917
60. Teukam YGN, Dassi LK, Manica M, Probst D, Laino T Language Models can Identify
Enzymatic Active Sites in Protein Sequences.
gg-v3
61. Edwards C, Lai T, Ros K, Honke G, Cho K, Ji H (2022) Translation between Molecules and
Natural Language.
62. Christofidellis D, Giannone G, Born J, Winther O, Laino T, Manica M (2023) Unifying
Molecular and Textual Representations via Multi-task Language Modelling
63. Alberts M, Laino T, Vaucher AC (2023) Leveraging Infrared Spectroscopy for Automated
Structure Elucidation. ChemRxiv.
64. Raschka S (2023) Finetuning Large Language Models. https://magazine.sebastianraschka.com/
p/finetuning-large-language-models?utm_campaign=post. Accessed 17 May 2023
65. Zhang R, Han J, Zhou A, Hu X, Yan S, Lu P, Li H, Gao P, Qiao Y (2023) LLaMA-Adapter:
Efficient Fine-tuning of Language Models with Zero-init Attention.
arXiv.2303.16199
66. Hoffmann J, Borgeaud S, Mensch A, Buchatskaya E, Cai T, Rutherford E, Casas D de L,
Hendricks LA, Welbl J, Clark A, Hennigan T, Noland E, Millican K, Driessche G van den,
Damoc B, Guy A, Osindero S, Simonyan K, Elsen E, Rae JW, Vinyals O, Sifre L (2022) Training
Compute-Optimal Large Language Models.
67. Wei J, Tay Y, Bommasani R, Raffel C, Zoph B, Borgeaud S, Yogatama D, Bosma M, Zhou
D, Metzler D, Chi EH, Hashimoto T, Vinyals O, Liang P, Dean J, Fedus W (2022) Emergent
Abilities of Large Language Models.
68. Wei J, Wang X, Schuurmans D, Bosma M, Ichter B, Xia F, Chi E, Le Q, Zhou D (2023) Chainof-Thought Prompting Elicits Reasoning in Large Language Models.
arXiv.2201.11903
69. Ouyang L, Wu J, Jiang X, Almeida D, Wainwright CL, Mishkin P, Zhang C, Agarwal S,
Slama K, Ray A, Schulman J, Hilton J, Kelton F, Miller L, Simens M, Askell A, Welinder P,
Christiano P, Leike J, Lowe R (2022) Training Language Models to Follow Instructions with
Human Feedback.
70. Jablonka KM, Schwaller P, Ortega-Guerrero A, Smit B (2023) Leveraging large language
models for predictive chemistry. Nat. Mach. Intell.
88-1
71. Ramos MC, Michtavy SS, Porosoff MD, White AD (2023) Bayesian Optimization of Catalysts
With In-context Learning.
72. Boiko DA, MacKnight R, Kline B, Gomes G (2023) Autonomous chemical research with large
language models. Nature.
73. Bran AM, Cox S, Schilter O, Baldassari C, White AD, Schwaller P (2024) Augmenting LargeLanguage Models with Chemistry Tools. Nat. Mach. Intell.
024-00832-8
74. Howard J, Ruder S (2018) Universal Language Model Fine-tuning for Text Classification. In:
Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics
(Volume 1: Long Papers). Association for Computational Linguistics, Melbourne, Australia,
pp 328–339
75. Wei J, Bosma M, Zhao V, Guu K, Yu AW, Lester B, Du N, Dai AM, Le QV (2021) Finetuned
Language Models are Zero-Shot Learners.
76. Yin X, Chen W, Wu X, Yue H (2017) Fine-tuning and Visualization of Convolutional Neural
Networks. In: 2017 12th IEEE Conference on Industrial Electronics and Applications (ICIEA).
pp 1310–1315
77. Dai H, Li C, Coley CW, Dai B, Song L (2020) Retrosynthesis Prediction with Conditional
Graph Logic Network.
78. Zhang B, Zhang X, Du W, Song Z, Zhang G, Zhang G, Wang Y, Chen X, Jiang J, Luo Y
(2022) Chemistry-Informed Molecular Graph as Reaction Descriptor for Machine-Learned
Retrosynthesis Planning. Proc Natl Acad Sci 119:e2212711119.
2212711119
79. Jorner K, Turcani L (2022) kjelljorner/morfeus: v0.7.2
https://doi.org/10.48550/arXiv.2204.11817
https://doi.org/10.26434/chemrxiv-2023-5v27f
https://doi.org/10.48550/ARXIV.2206.07682
https://doi.org/10.48550/arXiv.2203.02155
https://doi.org/10.48550/arXiv.2304.05341
https://doi.org/10.1038/s41586-023-06792-0.
https://doi.org/10.48550/arXiv.2001.01408
https://doi.org/10.26434/chemrxiv-2021-m20
https://doi.org/10.48550/
https://doi.org/10.48550/arXiv.2203.15556
https://doi.org/10.48550/
https://doi.org/10.1038/s42256-023-007
https://doi.org/10.1038/s42256-
https://doi.org/10.48550/arXiv.2109.01652
https://doi.org/10.1073/pnas.

8 Transformers and Large Language Models for Chemistry and Drug … 163
https://t.me/med1917
80. Ranković B, Griffiths R-R, Moss HB, Schwaller P (2023) Bayesian Optimisation for Addi-
tive Screening and Yield Improvements in Chemical Reactions – Beyond One-Hot Encoding.
Digital Discovery, 2024, Advance Article.
81. Shields BJ, Stevens J, Li J, Parasram M, Damani F, Alvarado JIM, Janey JM, Adams RP,
Doyle AG (2021) Bayesian Reaction Optimization as a Tool for Chemical Synthesis. Nature
590:89–96.
82. Bagal V, Aggarwal R, Vinod PK, Priyakumar UD (2022) MolGPT: Molecular Generation Using
a Transformer-Decoder Model. J Chem Inf Model 62:2064–2076.
jcim.1c00600
83. Rothchild D, Tamkin A, Yu J, Misra U, Gonzalez J (2021) C5T5: Controllable Generation of
Organic Molecules with Transformers.
84. Wang W, Wang Y, Zhao H, Sciabola S (2022) A Transformer-based Generative Model for De
Novo Molecular Design.
85. Bengio E, Jain M, Korablyov M, Precup D, Bengio Y (2021) Flow Network based Generative
Models for Non-Iterative Diverse Candidate Generation.
04399
86. Born J, Manica M (2023) Regression Transformer Enables Concurrent Sequence Regression
and Generation For Molecular Language Modelling. Nat Mach Intell 5:432–444.
org/10.1038/s42256-023-00639-z
87. Flam-Shepherd D, Aspuru-Guzik A (2023) Language Models can Generate Molecules, Materials, and Protein Binding Sites Directly in Three Dimensions as XYZ, CIF, and PDB files.
https://doi.org/10.48550/arXiv.2305.05708
88. Schick T, Dwivedi-Yu J, Dessì R, Raileanu R, Lomeli M, Zettlemoyer L, Cancedda N, Scialom
T (2023) Toolformer: Language Models Can Teach Themselves to Use Tools.
10.48550/arXiv.2302.04761
89. Karpas E, Abend O, Belinkov Y, Lenz B, Lieber O, Ratner N, Shoham Y, Bata H, Levine Y,
Leyton-Brown K, Muhlgay D, Rozen N, Schwartz E, Shachaf G, Shalev-Shwartz S, Shashua
A, Tenenholtz M (2022) MRKL Systems: A Modular, Neuro-Symbolic Architecture that
Combines Large Language Models, External Knowledge Sources and Discrete Reasoning.
https://doi.org/10.48550/arXiv.2205.00445
90. Yao S, Zhao J, Yu D, Du N, Shafran I, Narasimhan K, Cao Y (2023) ReAct: Synergizing
Reasoning and Acting in Language Models.
91. D. White A, M. Hocky G, A. Gandhi H, Ansari M, Cox S, P. Wellawatte G, Sasmal S, Yang Z,
Liu K, Singh Y, Ccoa WJP (2023) Assessment of Chemistry Knowledge in Large Language
Models that Generate Code. Digit Discov 2:368–376.
92. Jablonka KM, Ai Q, Al-Feghali A, Badhwar S, Bran JDBAM, Bringuier S, Brinson LC, Choudhary K, Circi D, Cox S, de Jong WA, Evans ML, Gastellu N, Genzling J, Gil MV, Gupta AK,
Hong Z, Imran A, Kruschwitz S, Labarre A, Lála J, Liu T, Ma S, Majumdar S, Merz GW,
Moitessier N, Moubarak E, Mouriño B, Pelkie B, Pieler M, Ramos MC, Ranković B, Rodriques
SG, Sanders JN, Schwaller P, Schwarting M, Shi J, Smit B, Smith BE, Van Heck J, Völker C,
Ward L, Warren S, Weiser B, Zhang S, Zhang X, Zia GA, Scourtas A, Schmidt K, Foster I,
White AD, Blaiszik B (2023) 14 Examples of How LLMs Can Transform Materials Science
and Chemistry: A Reflection on a Large Language Model Hackathon.
ARXIV.2306.06283
93. Su B, Du D, Yang Z, Zhou Y, Li J, Rao A, Sun H, Lu Z, Wen J-R (2022) A Molecular
Multimodal Foundation Model Associating Molecule Graphs with Natural Language.
doi.org/10.48550/arXiv.2209.05481
https://doi.org/10.1038/s41586-021-03213-y
https://doi.org/10.48550/arXiv.2210.08749
https://doi.org/10.1039/D3DD00096F
https://doi.org/10.1021/acs.
https://doi.org/10.48550/arXiv.2108.10307
https://doi.org/10.48550/arXiv.2106.
https://doi.
https://doi.org/
https://doi.org/10.48550/arXiv.2210.03629
https://doi.org/10.1039/D2DD00087C
https://doi.org/10.48550/
https://
Соседние файлы в папке Библиотека им академика М.И. Перельмана
