Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5629_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
15.09.2026
Размер:
16 Мб
Скачать
☆
5 • Articial Intelligence and Machine Learning 111
Liu, Cynthia, Qiongqiong Zhou, Yingzhu Li, Linda V. Garner, Steve P. Watkins, Linda J. Carter,
Jeffrey Smoot, etal. 2020. “Research and Development on Therapeutic Agents and Vaccines
for COVID‑19 and Related Human Coronavirus Diseases.” ACS Central Science6 (3): 315–31. Liu, Xiaofeng, Richard D. Taylor, Laura Grifn, Shu‑Fen Coker, Ralph Adams, Tom Ceska,
Jiye Shi, Alastair D. G. Lawson, and Terry Baker. 2017. “Computational Design of an
Epitope‑Specic Keap1 Binding Antibody Using Hotspot Residues Grafting and CDR
Loop Swapping.” Scientic Reports 7 (January): 41306. Liu, Yinhan, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike
Lewis, Luke Zettlemoyer, and Veselin Stoyanov. “RoBERTa: A Robustly Optimized BERT
Pretraining Approach.” arXiv. https://doi.org/10.48550/arXiv.1907.11692 Lu, Ruei‑Min, Yu‑Chyi Hwang, I‑ Ju Liu, Chi‑Chiu Lee, Han‑Zen Tsai, Hsin‑Jung Li, and
Han‑Chung Wu. 2020. “Development of Therapeutic Antibodies for the Treatment of
Diseases.” Journal of Biomedical Science27 (1): 1. Maia, Eduardo Habib Bechelane, Letícia Cristina Assis, Tiago Alves de Oliveira, Alisson Marques
da Silva, and Alex Gutterres Taranto. 2020. “Structure‑Based Virtual Screening: From
Classical to Articial Intelligence.” Frontiers in Chemistry 8 (April): 343. Major, Sylvia M., Satoshi Nishizuka, Daisaku Morita, Rick Rowland, Margot Sunshine, Uma
Shankavaram, Frank Washburn, et al. 2006. “AbMiner: A Bioinformatic Resource on
Available Monoclonal Antibodies and Corresponding Gene Identiers for Genomic,
Proteomic, and Immunologic Studies.” BMC Bioinformatics 7 (April): 192. Makowski, Emily K., Lina Wu, Priyanka Gupta, and Peter M. Tessier. 2021. “Discovery‑Stage
Identication of Drug‑like Antibodies Using Emerging Experimental and Computational
Methods.” mAbs 13 (1): 1895540. Mannar, Dhiraj, James W. Saville, Zehua Sun, Xing Zhu, Michelle M. Marti, Shanti S. Srivastava,
Alison M. Berezuk, etal. 2022. “SARS‑CoV‑2 Variants of Concern: Spike Protein Mutational
Analysis and Epitope for Broad Neutralization.” Nature Communications 13 (1): 4696. Marks, Claire, Alissa M. Hummer, Mark Chin, and Charlotte M. Deane. 2021. “Humanization
of Antibodies Using a Machine Learning Approach on Large‑Scale Repertoire Data.”
Bioinformatics, June. https://doi.org/10.1093/bioinformatics/btab434. Marks, Debora S., Lucy J. Colwell, Robert Sheridan, Thomas A. Hopf, Andrea Pagnani, Riccardo
Zecchina, and Chris Sander. 2011. “Protein 3D Structure Computed from Evolutionary
Sequence Variation.” PloS One 6 (12): e28766. Mason, Derek M., Simon Friedensohn, Cédric R. Weber, Christian Jordi, Bastian Wagner, Simon
M. Meng, Roy A. Ehling, etal. 2021. “Optimization of Therapeutic Antibodies by Predicting
Antigen Specicity from Antibody Sequence via Deep Learning.” Nature Biomedical
Engineering 5 (6): 600–612. Miho, Enkelejda, Alexander Yermanos, Cédric R. Weber, Christoph T. Berger, Sai T. Reddy,
and Victor Greiff. 2018. “Computational Strategies for Dissecting the High‑Dimensional
Complexity of Adaptive Immune Repertoires.” Frontiers in Immunology 9 (February): 224. Mikolov, Tomas, Ilya Sutskever, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. “Distributed
Representations of Words and Phrases and Their Compositionality.” arXiv [cs.CL]. arXiv.
https://proceedings.neurips.cc/paper/2013/hash/9aa42b31882ec039965f3c4923ce901b‑Ab
stract.html. Młokosiewicz, Jakub, Piotr Deszyński, Wiktoria Wilman, Igor Jaszczyszyn, Rajkumar Ganesan,
Aleksandr Kovaltsuk, Jinwoo Leem, Jacob D. Galson, and Konrad Krawczyk. 2022.
“AbDiver: A Tool to Explore the Natural Antibody Landscape to Aid Therapeutic Design.”
Bioinformatics 38 (9): 2628–30. Narayanan, Harini, Fabian Dingfelder, Alessandro Butté, Nikolai Lorenzen, Michael Sokolov,
and Paolo Arosio. 2021. “Machine Learning for Biologics: Opportunities for Protein
Engineering, Developability, and Formulation.” Trends in Pharmacological Sciences 42
(3): 151–65.
112 Biopharmaceutical Informatics
Nguyen, Minh N., Mohan R. Pradhan, Chandra Verma, and Pingyu Zhong. 2017. “The Interfacial
Character of Antibody Paratopes: Analysis of Antibody‑Antigen Structures.” Bioinformatics
33 (19): 2971–76. Norman, Richard A., Francesco Ambrosetti, Alexandre M. J. J. Bonvin, Lucy J. Colwell, Sebastian
Kelm, Sandeep Kumar, and Konrad Krawczyk. 2020. “Computational Approaches to
Therapeutic Antibody Design: Established Methods and Emerging Trends.” Briengs in
Bioinformatics 21 (5): 1549–67. North, Benjamin, Andreas Lehmann, and Roland L. Dunbrack Jr. 2011. “A New Clustering of
Antibody CDR Loop Conformations.” Journal of Molecular Biology 406 (2): 228–56. Obrezanova, Olga, Andreas Arnell, Ramón Gómez de la Cuesta, Maud E. Berthelot, Thomas R.
A. Gallagher, Jesús Zurdo, and Yvette Stallwood.2015. “Aggregation Risk Prediction for
Antibodies and Its Application to Biotherapeutic Development.” mAbs 7 (2): 352–63. Olimpieri, Pier Paolo, Anna Chailyan, Anna Tramontano, and Paolo Marcatili. 2013. “Prediction
of Site‑Specic Interactions in Antibody‑Antigen Complexes: The proABC Method and
Server.” Bioinformatics 29 (18): 2285–91. Olsen, Tobias H., Fergus Boyles, and Charlotte M. Deane. 2022a. “Observed Antibody Space:
A Diverse Database of Cleaned, Annotated, and Translated Unpaired and Paired Antibody
Sequences.” Protein Science: A Publication of the Protein Society 31 (1): 141–46. Olsen, Tobias H., Iain H. Moal, and Charlotte M. Deane. 2022b. “AbLang: An Antibody
Language Model for Completing Antibody Sequences.” bioRxiv. https://doi.
org/10.1101/2022.01.20.477061. Ostrovsky‑Berman, Miri, Boaz Frankel, Pazit Polak, and Gur Yaari. 2021. “Immune2vec:
Embedding B/T Cell Receptor Sequences in RN Using Natural Language Processing.”
Frontiers in Immunology 12. https://doi.org/10.3389/mmu.2021.680687. Peng, Hung‑Pin, Kuo Hao Lee, Jhih‑Wei Jian, and An‑Suei Yang. 2014. “Origins of Specicity
and Afnity in Antibody–protein Interactions.” Proceedings of the National Academy of
Sciences 111 (26): E2656–65. Pittala, Srivamshi, and Chris Bailey‑Kellogg. 2020. “Learning Context‑Aware Structural
Representations to Predict Antigen and Antibody Binding Interfaces.” Bioinformatics 36
(13): 3996–4003. Prabakaran, R., Puneet Rawat, Sandeep Kumar, and M. Michael Gromiha. 2021. “ANuPP: A
Versatile Tool to Predict Aggregation Nucleating Regions in Peptides and Proteins.” Journal
of Molecular Biology 433 (11): 166707. Prihoda, David, Jad Maamary, Andrew Waight, Veronica Juan, Laurence Fayadat‑Dilman, Daniel
Svozil, and Danny A. Bitton. 2022. “BioPhi: A Platform for Antibody Design, Humanization,
and Humanness Evaluation Based on Natural Antibody Repertoires and Deep Learning.”
mAbs 14 (1): 2020203. Rawat, Puneet, R. Prabakaran, Sandeep Kumar, and M. Michael Gromiha. 2021a. “Exploring the
Sequence Features Determining Amyloidosis in Human Antibody Light Chains.” Scientic
Reports 11 (1): 13785. Rawat, Puneet, R. Prabakaran, Sandeep Kumar, and M. Michael Gromiha. 2021b. “AbsoluRATE:
An in‑Silico Method to Predict the Aggregation Kinetics of Native Proteins.” Biochimica et
Biophysica Acta: Proteins and Proteomics 1869 (9): 140682. Raybould, Matthew I. J., Aleksandr Kovaltsuk, Claire Marks, and Charlotte M. Deane. 2021.
“CoV‑AbDab: The Coronavirus Antibody Database.” Bioinformatics 37 (5): 734–35. Raybould, Matthew I. J., Claire Marks, Konrad Krawczyk, Bruck Taddese, Jaroslaw Nowak, Alan
P. Lewis, Alexander Bujotzek, Jiye Shi, and Charlotte M. Deane. 2019. “Five Computational
Developability Guidelines for Therapeutic Antibody Proling.” Proceedings of the National
Academy of Sciences of the United States of America 116 (10): 4025–30.
5 • Articial Intelligence and Machine Learning 113
Raybould, Matthew I. J., Claire Marks, Alan P. Lewis, Jiye Shi, Alexander Bujotzek, Bruck
Taddese, and Charlotte M. Deane. 2020. “Thera‑SAbDab: The Therapeutic Structural
Antibody Database.” Nucleic Acids Research 48 (D1): D383–88. Regep, Cristian, Guy Georges, Jiye Shi, Bojana Popovic, and Charlotte M. Deane. 2017. “The H3
Loop of Antibodies Shows Unique Structural Characteristics.” Proteins 85 (7): 1311–18. Renaud, Nicolas, Cunliang Geng, Sonja Georgievska, Francesco Ambrosetti, Lars Ridder, Dario
F. Marzella, Manon F. Réau, Alexandre M. J. J. Bonvin, and Li C. Xue. 2021. “DeepRank:
A Deep Learning Framework for Data Mining 3D Protein‑Protein Interfaces.” Nature
Communications 12 (1): 7068. Repecka, D., V. Jauniskis, and L. Karpus, etal. 2021. “Expanding Functional Protein Sequence
Spaces Using Generative Adversarial Networks.” Nature Machine Intelligence3: 324–333.
https://www.nature.com/articles/s42256‑021‑00310‑5. Richardson, Eve, Jacob D. Galson, Paul Kellam, Dominic F. Kelly, Sarah E. Smith, Anne Palser,
Simon Watson, and Charlotte M. Deane. 2021. “A Computational Method for Immune
Repertoire Mining That Identies Novel Binders from Different Clonotypes, Demonstrated
by Identifying Anti‑Pertussis Toxoid Antibodies.” mAbs 13 (1): 1869406. Rives, Alexander, Joshua Meier, Tom Sercu, Siddharth Goyal, Zeming Lin, Jason Liu, Demi Guo,
etal. 2021. “Biological Structure and Function Emerge from Scaling Unsupervised Learning
to 250 Million Protein Sequences.” Proceedings of the National Academy of Sciences of the
United States of America 118 (15). https://doi.org/10.1073/pnas.2016239118. Robert, Philippe A., Rahmad Akbar, Robert Frank, Milena Pavlović, Michael Widrich, Igor Snapkov,
Andrei Slabodkin, etal. 2021. “Unconstrained Generation of Synthetic Antibody‑Antigen
Structures to Guide Machine Learning Methodology for Real‑World Antibody Specicity
Prediction.” bioRxiv. bioRxiv. https://doi.org/10.1101/2021.07.06.451258. Roberts, Christopher J. 2014. “Therapeutic Protein Aggregation: Mechanisms, Design, and
Control.” Trends in Biotechnology 32 (7): 372–80. Ruffolo, Jeffrey A., Jeffrey J. Gray, and Jeremias Sulam. 2021. “Deciphering Antibody Afnity
Maturation with Language Models and Weakly Supervised Learning.” arXiv [q‑bio.BM].
arXiv. https://arxiv.org/abs/2112.07782. Ruffolo, Jeffrey A., Carlos Guerra, Sai Pooja Mahajan, Jeremias Sulam, and Jeffrey J. Gray.
2020. “Geometric Potentials from Deep Learning Improve Prediction of CDR H3 Loop
Structures.” Bioinformatics 36 (Suppl_1): i268–75. Ruffolo, Jeffrey A., Jeremias Sulam, and Jeffrey J. Gray. 2022. “Antibody Structure Prediction
Using Interpretable Deep Learning.” Patterns (New York, N.Y.) 3 (2): 100406. Saha, Sudipto, Manoj Bhasin, and Gajendra P. S. Raghava. 2005. “Bcipep: A Database of B‑Cell
Epitopes.” BMC Genomics 6 (May): 79. Sang, Zhe, Yufei Xiang, Ivet Bahar, and Yi Shi. 2022. “Llamanade: An Open‑Source Computational
Pipeline for Robust Nanobody Humanization.” Structure 30 (3): 418–29.e3. Schmitz, Samuel, Cinque Soto, James E. Crowe Jr, and Jens Meiler. 2020. “Human‑Likeness of
Antibody Biologics Determined by Back‑Translation and Comparison with Large Antibody
Variable Gene Repertoires.” mAbs 12 (1): 1758291. Schneider, Constantin, Andrew Buchanan, Bruck Taddese, and Charlotte M. Deane. 2021.
“DLAB‑Deep Learning Methods for Structure‑Based Virtual Screening of Antibodies.”
Bioinformatics, September. https://doi.org/10.1093/bioinformatics/btab660. Sharma, Vikas K., Thomas W. Patapoff, Bruce Kabakoff, Satyan Pai, Eric Hilario, Boyan Zhang,
Charlene Li, etal. 2014. “In Silico Selection of Therapeutic Antibodies for Development:
Viscosity, Clearance, and Chemical Stability.” Proceedings of the National Academy of
Sciences of the United States of America 111 (52): 18601–6.
114 Biopharmaceutical Informatics
Sheng, Zizhang, Chaim A. Schramm, Rui Kong, NISC Comparative Sequencing Program, James
C. Mullikin, John R. Mascola, Peter D. Kwong, and Lawrence Shapiro. 2017. “Gene‑Specic
Substitution Proles Describe the Types and Frequencies of Amino Acid Changes during
Antibody Somatic Hypermutation.” Frontiers in Immunology 8 (May): 537. Shin, Jung‑Eun, Adam J. Riesselman, Aaron W. Kollasch, Conor McMahon, Elana Simon, Chris
Sander, Aashish Manglik, Andrew C. Kruse, and Debora S. Marks. 2021. “Protein Design
and Variant Prediction Using Autoregressive Generative Models.” Nature Communications
12 (1): 2403. Singh, Satish Kumar. 2011. “Impact of Product‑Related Factors on Immunogenicity of
Biotherapeutics.” Journal of Pharmaceutical Sciences 100 (2): 354–87. Sircar, Aroop, and Jeffrey J. Gray. 2010. “SnugDock: Paratope Structural Optimization during
Antibody‑Antigen Docking Compensates for Errors in Antibody Homology Models.” PLoS
Computational Biology 6 (1): e1000644. Sirin, Sarah, James R. Apgar, Eric M. Bennett, and Amy E. Keating. 2016. “AB‑Bind: Antibody
Binding Mutational Database for Computational Afnity Predictions.” Protein Science: A
Publication of the Protein Society 25 (2): 393–409. Smialowski, Pawel, Gero Doose, Phillipp Torkler, Stefanie Kaufmann, and Dmitrij Frishman.
2012. “PROSO II‑‑a New Method for Protein Solubility Prediction.” The FEBS Journal 279
(12): 2192–2200. Sormanni, Pietro, Francesco A. Aprile, and Michele Vendruscolo. 2015. “The CamSol Method
of Rational Design of Protein Mutants with Enhanced Solubility.” Journal of Molecular
Biology 427 (2): 478–90. Sormanni, Pietro, Damiano Piovesan, Gabriella T. Heller, Massimiliano Bonomi, Predrag Kukic,
Carlo Camilloni, Monika Fuxreiter, etal. 2017. “Simultaneous Quantication of Protein
Order and Disorder.” Nature Chemical Biology 13 (4): 339–42. Swindells, Mark B., Craig T. Porter, Matthew Couch, Jacob Hurst, K. R. Abhinandan, Jens H.
Nielsen, Gary Macindoe, James Hetherington, and Andrew C. R. Martin. 2017. “abYsis:
Integrated Antibody Sequence and Structure‑Management, Analysis, and Prediction.”
Journal of Molecular Biology 429 (3): 356–64. Tartaglia, Gian Gaetano, Andrea Cavalli, Riccardo Pellarin, and Amedeo Caisch. 2005. “Prediction
of Aggregation Rate and Aggregation‑Prone Segments in Polypeptide Sequences.” Protein
Science: A Publication of the Protein Society 14 (10): 2723–34. Tomar, Dheeraj S., Li, Matthew P. Broulidakis, Nicholas G. Luksha, Christopher T. Burns, Satish
K. Singh, and Sandeep Kumar. 2017. “In‑Silico Prediction of Concentration‑Dependent
Viscosity Curves for Monoclonal Antibody Solutions.” mAbs 9 (3): 476–89. Torjesen, Ingrid. n.d. “Drug Development: The Journey of a Medicine from Lab to Shelf.” The
Pharmaceutical Journal, doi: 10.1211/PJ.2015.20068196 Toseland, Christopher P., Debra J. Clayton, Helen McSparron, Shelley L. Hemsley, Martin J.
Blythe, Kelly Paine, Irini A. Doytchinova, Pingping Guan, Channa K. Hattotuwagama,
and Darren R. Flower. 2005. “AntiJen: A Quantitative Immunology Database Integrating
Functional, Thermodynamic, Kinetic, Biophysical, and Cellular Data.” Immunome Research
1 (1): 4. Tubiana, Jérôme, Dina Schneidman‑Duhovny, and Haim J. Wolfson. 2022. “ScanNet: An
Interpretable Geometric Deep Learning Model for Structure‑Based Protein Binding Site
Prediction.” Nature Methods 19 (6): 730–39. Tunyasuvunakool, Kathryn, Jonas Adler, Zachary Wu, Tim Green, Michal Zielinski, Augustin
Žídek, Alex Bridgland, etal. 2021. “Highly Accurate Protein Structure Prediction for the
Human Proteome.” Nature 596 (7873): 590–96. Vashchenko, Denis, Sam Nguyen, Andre Goncalves, Felipe Leno da Silva, Brenden Petersen,
Thomas Desautels, and Daniel Faissol. 2022. “AbBERT: Learning Antibody Humanness via
Masked Language Modeling.” bioRxiv. https://doi.org/10.1101/2022.08.02.502236.
5 • Articial Intelligence and Machine Learning 115
Vita, Randi, James A. Overton, Jason A. Greenbaum, Julia Ponomarenko, Jason D. Clark, Jason
R. Cantrell, Daniel K. Wheeler, etal. 2015. “The Immune Epitope Database (IEDB) 3.0.”
Nucleic Acids Research 43 (Database issue): D405–12. Vita, Randi, Laura Zarebski, Jason A. Greenbaum, Hussein Emami, Ilka Hoof, Nima Salimi,
Rohini Damle, Alessandro Sette, and Bjoern Peters. 2010. “The Immune Epitope Database
2.0.” Nucleic Acids Research 38 (Database issue): D854–62.
Wang, Chunyan, Wentao Li, Dubravka Drabek, Nisreen M. A. Okba, Rien van Haperen, Albert
D. M. E. Osterhaus, Frank J. M. van Kuppeveld, Bart L. Haagmans, Frank Grosveld,
and Berend‑Jan Bosch. 2020a. “A Human Monoclonal Antibody Blocking SARS‑CoV‑2
Infection.” Nature Communications 11 (1): 2251. Wang, Xiao, Genki Terashi, Charles W. Christoffer, Mengmeng Zhu, and Daisuke Kihara.
2020b. “Protein Docking Model Evaluation by 3D Deep Convolutional Neural Networks.”
Bioinformatics 36 (7): 2113–18. Weitzner, Brian D., Daisuke Kuroda, Nicholas Marze, Jianqing Xu, and Jeffrey J. Gray. 2014.
“Blind Prediction Performance of RosettaAntibody 3.0: Grafting, Relaxation, Kinematic
Loop Modeling, and Full CDR Optimization.” Proteins 82 (8): 1611–23. Wilman, Wiktoria, Sonia Wróbel, Weronika Bielska, Piotr Deszynski, Paweł Dudzic, Igor
Jaszczyszyn, Jędrzej Kaniewski, et al. 2022. “Machine‑Designed Biotherapeutics:
Opportunities, Feasibility and Advantages of Deep Learning in Computational Antibody
Discovery.” Briengs in Bioinformatics 23 (4). https://doi.org/10.1093/bib/bbac267. Wilton, Emily E., Michael P. Opyr, Senthilkumar Kailasam, Ronja F. Kothe, and Hans‑Joachim
Wieden. 2018. “sdAb‑DB: The Single Domain Antibody Database.” ACS Synthetic Biology
7 (11): 2480–84. Wolf, T., L. Debut, V. Sanh, and J. Chaumond. 2020. “Transformers: State‑of‑the‑Art Natural
Language Processing.” Proceedings of the 2020 Conference on Empirical Methods in
Natural Language Processing: System Demonstrations. https://aclanthology.org/2020.
emnlp‑demos.6/?ref=https://codemonkey.link. Wollacott, Andrew M., Chonghua Xue, Qiuyuan Qin, June Hua, Tanggis Bohnuud, Karthik
Viswanathan, and Vijaya B. Kolachalama. 2019. “Quantifying the Nativeness of Antibody
Sequences Using Long Short‑Term Memory Networks.” Protein Engineering, Design &
Selection: PEDS 32 (7): 347–54. Wu, Jiaxiang, Fandi Wu, Biaobin Jiang, Wei Liu, and Peilin Zhao. 2022. “tFold‑Ab: Fast and
Accurate Antibody Structure Prediction without Sequence Homologs.” bioRxiv. https://doi.
org/10.1101/2022.11.10.515918. Wu, Zachary, S. B. Jennifer Kan, Russell D. Lewis, Bruce J. Wittmann, and Frances H.
Arnold.2019. “Machine Learning‑Assisted Directed Protein Evolution with Combinatorial
Libraries.” Proceedings of the National Academy of Sciences of the United States of America
116 (18): 8852–58. Xu, Yingda, Dongdong Wang, Bruce Mason, Tony Rossomando, Ning Li, Dingjiang Liu, Jason
K. Cheung, et al. 2019. “Structure, Heterogeneity and Developability Assessment of
Therapeutic Antibodies.” mAbs 11 (2): 239–64. Zavrtanik, Uroš, and San Hadži. 2019. “A Non‑redundant Data Set of Nanobody‑Antigen Crystal
Structures.” Data in Brief 24 (June): 103754. Zhang, Jie, Yishan Du, Pengfei Zhou, Jinru Ding, Shuai Xia, Qian Wang, Feiyang Chen, etal. 2022.
“Predicting Unseen Antibodies’ Neutralizability via Adaptive Graph Neural Networks.”
Nature Machine Intelligence, November, 1–13.
From Deep Generative Models
6
to Structure-Based Simulations
Computational Approaches for Antibody Design
Daisuke Kuroda

6.1 INTRODUCTION

In the rapidly evolving eld of biotherapeutics, the intersection of computational science and protein engineering has revolutionized the approach to drug discovery and develop‑ ment. Central to this transformation is the art of antibody design, a complex and critical component of modern therapeutic strategies.1 Antibodies, with their unique specicity and versatility, have emerged as a leading modality in the treatment of various diseases, ranging from cancers to autoimmune disorders and infectious diseases. The process of designing these therapeutic antibodies, however, presents a multifaceted challenge, encompassing aspects of protein engineering, bioinformatics, and developability.
The advent of machine learning (ML) and generative models, including large lan‑ guage models (LLMs), has opened new frontiers in deciphering the complex language of proteins, particularly in understanding and predicting the vast diversity of the antibody repertoire.
116
3–5
These computational tools, coupled with advanced molecular simulations,
2
6 • From Deep Generative Models to Structure-Based Simulations 117
enable scientists to explore the structural and functional nuances of antibodies. By lever‑ aging these technologies, the eld of antibody design has transitioned from a largely empirical endeavor, relying on experimental trial and error, to one that is increasingly predictive and rational.
Developability, a key consideration in antibody engineering, involves assessing the suitability of antibody candidates for therapeutic use, focusing on their manufacturabil‑ ity, stability, and efcacy.6 This process has been greatly enhanced by bioinformatics and computational science, which provide insights into the molecular characteristics that govern the behavior of antibodies in biological systems. Through computational approaches, the scope of protein engineering has expanded, allowing for the design of antibodies with improved properties from their amino acid sequences (Figure6.1). The properties amenable to computational enhancement encompass physicochemical attri‑ butes, such as binding afnity and stability, as well as biological and pharmacological aspects, including immunogenicity.
The integration of ML into antibody design is particularly noteworthy. Through the analysis of vast datasets, including sequences and structures of existing antibod‑ ies, enabled by high‑throughput sequencing of immune repertoires, ML algorithms can predict the antigen‑binding afnity and specicity of novel antibody candidates.7 This capability is pivotal in accelerating the drug discovery process, reducing the time and cost associated with the development of new biotherapeutics.
Among the various algorithms in ML, deep learning (DL) has garnered signicant attention due to its unparalleled success in mimicking human‑like decision‑making and
FIGURE 6.1 The computational workow in antibody drug discovery: Seed antibodies are generated through various methods, including animal immunization, synthetic libraries, single-cell analysis, and computational design. The sequences of these seed antibodies can be further optimized using computational design calculations based on either sequences or structures, yielding lead antibody sequences. Structures of the antibody alone, as well as antibody-antigen complexes, can be predicted using computational methods. Additionally, developability assessments of lead antibodies can be conducted in silico. Topics discussed in this chapter are underlined and italicized in red for emphasis.
118 Biopharmaceutical Informatics
learning patterns. This has been especially evident since the early 2010s with break‑ throughs in image and speech recognition.
8,9
Deep learning employs deep neural net‑ works, which comprise several core architectures, such as convolutional neural networks (CNNs),10 recurrent neural networks (RNNs) or long short‑term memory (LSTM),11gen‑ erative adversarial networks (GANs),12 variational autoencoders (VAEs)13, and trans‑ former14models (Figure6.2). Each architecture is distinguished by its specic strengths, weaknesses, and applications. These models autonomously learn features directly from data, thereby eliminating the need for manual feature extraction. The training process, which optimizes weights through backpropagation and gradient descent, renes learn‑ ing over time. Although DL models reduce the necessity for manual feature engineering by learning complex data patterns, data pre‑processing is still crucial in the bioinfor‑ matics workow. This ensures data quality and compatibility, which are essential for the success of DL applications. This aspect becomes particularly critical in antibody design, where the sizes, amino acid compositions, and structural diversity of the functional sites, namely the complementarity‑determining regions (CDRs), vary across antibodies. Such variability complicates the direct comparison of features among antibodies.
Molecular simulations, another vital component of computational antibody engi‑
neering, often rely on structural information of antibodies and their antigens.
15,16
These
simulations play a crucial role in visualizing the dynamic interactions between antibodies
FIGURE6.2 Classication of deep learning architectures: CNN: Designed to process data with a grid-like topology, using convolutional layers to efciently extract and learn spatial hierarchies in data, particularly useful in image analysis. RNN: Designed to handle sequen­tial data, where the output from previous steps is fed back into the network to inform responses at later steps, making them ideal for tasks like language modeling and time-series analysis. LSTM: An advanced type of RNN that is capable of learning long-term depen­dencies in sequential data. VAE: Designed to encode data into a latent space and then reconstruct it, facilitating data generation by sampling from the learned distribution in the latent space. GAN: Consists of two competing networks: a generator that creates data and a discriminator that evaluates it, working together to produce highly realistic data samples. LLM: Designed to understand, generate, and manipulate natural language, often built on architectures like transformers.
6 • From Deep Generative Models to Structure-Based Simulations 119
and their targets, offering valuable insights into their mechanisms of action and potential off‑target effects. This informs the design of more effective and safer antibodies.
The synergy between ML, molecular simulations, and immune repertoire analysis is revolutionizing the eld of antibody design and protein engineering. studies in the literature have utilized these technologies for a variety of antibody‑related prediction tasks, including the prediction of antibody structures, complexes, from large antibody libraries. of these methodologies for de novo generation and optimization of antibody sequences. Therefore, the objective of this chapter is to explore these technological advancements, with a particular emphasis on their impact on antibody sequence design for next‑genera‑ tion biotherapeutics and their inuence on the future of antibody drug discovery.
25–30
their biophysical properties,
41– 45
However, this chapter specically focuses on the use
31– 40
and the identication of lead candidates
17,18
Numerous
19–2 4
antibody‑antigen

6.2 ANTIBODY GENERATION THROUGH DEEP GENERATIVE MODELS

Generating antibody sequences is a critical task in developing therapeutic antibodies and diagnostic tools and conducting basic immunology research (Figure 6.1). Traditional methods like animal immunization and synthetic libraries offer distinct approaches to acquiring antibodies with the desired specicities and afnities. Each method has its strengths and limitations, with the choice often dictated by project‑specic factors such as the nature of antigens, the intended antibody use, ethical considerations, and avail‑ able resources.
The transition from these conventional experimental methods to incorporating deep generative models into antibody design signies a paradigm shift in biotherapeutic development. Originally conceived for processing human languages, LLMs are now ingeniously repurposed to unravel the complex language of proteins.46 This adapta‑ tion not only highlights the versatility of LLMs but also underscores the parallels in pattern recognition and sequence analysis between linguistics and molecular biology. Leveraging their core capabilities in sequence, context, and pattern recognition, LLMs offer a fresh approach to interpreting protein sequences, analogous to parsing sentences in a natural language. This innovative convergence of computational linguistics and molecular biology equips us with potent tools to advance protein engineering and anti‑ body design. This section reviews the considerations involved in computationally gener‑ ating antibodies de novo through LLMs and other DL‑based techniques.
6.2.1 B‑Cell Repertoires in the Era of Articial
Intelligence
ML methods fundamentally rely on data to uncover the intricate patterns of life. This is equally true for LLMs focused on generating antibody sequences, which benet from
120 Biopharmaceutical Informatics
the rich data provided by B‑cell repertoires. These repertoires, now more accessible due to breakthroughs in high‑throughput sequencing technologies, have revolutionized elds such as immunology, vaccine development, and therapeutic antibody discovery.3 Enhancements in sequencing and computational analysis, coupled with a deeper under‑ standing of the immune system, have made these advances possible. The decreasing cost of sequencing and increased throughput capacity have made large‑scale genomic projects and the sequencing of vast cohorts and diverse species more practical, achiev‑ ing previously unimaginable depth.
Integrating B‑cell repertoire sequencing with other “omics” data, like proteomics and transcriptomics, offers a comprehensive view of the immune response.
47,4 8
Such a holistic approach can elucidate the evolution of repertoires in response to disease pro‑ gression, vaccination, or therapeutic interventions. Signature changes in the B‑cell rep‑ ertoire, linked to various diseases, including autoimmune disorders, infectious diseases, and cancers, have been identied.
49–53
These signatures have potential as biomarkers for diagnosis and prognosis. Sequencing efforts pre‑ and post‑vaccination have shed light on vaccine‑induced immunity, informing the design of more effective vaccines.
54–56
Moreover, high‑throughput sequencing has been instrumental in identifying antibodies with therapeutic potential against targets like severe acute respiratory syndrome coro‑ navirus 2 (SARS‑CoV‑2).
57
The development of specialized bioinformatics pipelines and software tools has enhanced the accuracy of sequence assembly, clonotype identication, and lineage trac‑ ing.58 These tools are designed to manage the vast datasets produced, enabling detailed analyses of B‑cell repertoires. By analyzing intricate patterns in repertoire data, com‑ putational algorithms guide both vaccine design and therapeutic antibody design.
4,59
Notable successes include employing ML, trained on high‑throughput sequencing data, to detect changes in B‑cell repertoire patterns in patients with dengue infection60 and relapsing‑remitting multiple sclerosis.61 These studies have demonstrated the capability of ML algorithms to capture the nuances of B‑cell repertoires, a task that would be chal‑ lenging for humans without the aid of ML.
In the realm of articial intelligence, representation learning, a key technique in deep learning, automates the discovery of data representations for feature detection or classication, bypassing the need for manual feature engineering. This approach allows models to identify complex patterns, enhancing performance in tasks such as classi‑ cation and prediction. In protein modeling and design, representation learning is criti‑ cal for encoding protein sequences, structures, or functional features for computational tasks. The signicance of representation stems from its ability to translate the biological essence of a protein into a format that computational models, particularly those rooted in deep learning, can efciently interpret and learn from.62 The choice of represen‑ tation—ranging from direct sequences and structures to more abstract feature‑based representations and embeddings—signicantly inuences model performance in pre‑ dicting protein structure and function and designing novel proteins (Figure6.3).
Effective representation captures essential protein features relevant to biological functions, facilitating accurate predictions and the design of proteins with new prop‑ erties. It is particularly vital in generative models aimed at creating novel protein sequences with specic functions. Here, representation learning navigates the protein