Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5908_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Contents
- •Foreword
- •Preface
- •About the Editors
- •Contributors
- •References
- •2.3.4 Barriers to Automation Adoption
- •2.4 Core Ingredients for Successful Digital Transformation
- •2.1 Introduction
- •2.3.1 Operational Challenges
- •2.3.2 Cultural Challenges
- •2.4.2 Cloud Computing
- •2.5 Case Studies of Successful Digital Transformation
- •2.6 Conclusion
- •References
- •3. Computational Protein Design Strategies for Optimization of Antigen Generation to Drive Antibody Discovery
- •3.1 Introduction
- •3.3 Antigen Generation Strategies
- •3.4 Computational Methods
- •3.4.2 Computational Protein Structure Prediction
- •References
- •4. Bioinformatic Analyses of Antibody Repertoires and Their Roles in Modern Antibody Drug Discovery
- •4.1 Introduction
- •4.6 Summary and Future Directions
- •Acknowledgments
- •References
- •5.1 Introduction
- •5.2 Databases
- •5.2.1 Databases in Machine Learning Approaches
- •5.2.2 Database Types
- •5.3 Applications of Machine Learning in Antibody Discovery and Development
- •5.3.1 Structure Prediction with Deep Learning
- •5.3.3 Developability
- •5.4 Antibody Generation and Design by Language Models
- •5.4.1 Antibody Representations
- •5.4.2 Representation Learning
- •5.4.3 Language Models
- •References
- •6.1 Introduction
- •6.2 Antibody Generation through Deep Generative Models
- •6.3.1 Sampling and Scoring
- •6.5 Conclusions and Perspectives
- •Acknowledgments
- •References
- •7.1 Introduction
- •7.2.3 Computational Approaches to Predict Antibody–Antigen Interaction
- •7.3 Conclusion
- •Competing Interests
- •Acknowledgments
- •References
- •8.2 Common Types of Molecular Simulations for Biomolecules
- •8.2.1 Molecular Dynamics (MD) Simulations
- •8.2.2 Monte Carlo (MC) Simulations
- •8.2.3 Challenges of Molecular Simulations
- •8.3.1 Periodic Boundary Conditions
- •8.4 Uses of Molecular Simulation in Antibody Drug Development
- •8.5 Conclusion
- •References
- •9. Considerations of Developability During the Early Stages of Antibody Drug Discovery and Design
- •9.1 Introduction
- •9.2 Historical Perspective
- •9.3 Clinical Antibody Data Set
- •9.5 Control Antibodies
- •9.7 Assessment of Chemical Liabilities
- •9.8 Conclusions and Future Perspectives
- •Acknowledgments
- •References
- •Abbreviations
- •10.1 Introduction
- •10.4.1 Conclusions and Outlook
- •Acknowledgments
- •References
- •11.8 Conclusions and Future Directions
- •References
- •12.1 Introduction to PK/PD and QSP Modeling
- •12.1.1 PK/PD Modeling
- •12.1.2 QSP Modeling
- •12.2.1 Monoclonal Antibodies (mAbs)
- •12.2.3 Cell Therapies
- •12.2.4 Gene Therapies
- •12.2.5 Vaccines
- •12.2.6 mRNA/siRNA/Oligonucleotide Therapeutics
- •12.4 Case Studies
- •12.5 Conclusions and Future Perspectives
- •References
- •13.1 Introduction
- •13.2 AI/ML: A Game Changer for Antibody Design
- •13.3 Multispecific Antibody Design
- •13.4 Adapting AI to the Design of Multispecific Antibodies
- •13.4.1 Structure Prediction and Modeling
- •13.4.2 Developability Prediction and Optimization
- •13.4.4 In Silico Modeling and Simulation
- •13.5 The Future: Beyond Optimization
- •13.5.1 Market Trends and Commercialization
- •13.5.2 Logic Gates, Biosensors, and De Novo Design
- •13.5.3 Challenges and Opportunities
- •13.6 Conclusion
- •Acknowledgments
- •References
- •Index

5 • Articial Intelligence and Machine Learning 111
Liu, Cynthia, Qiongqiong Zhou, Yingzhu Li, Linda V. Garner, Steve P. Watkins, Linda J. Carter,
Jeffrey Smoot, etal. 2020. “Research and Development on Therapeutic Agents and Vaccines
for COVID‑19 and Related Human Coronavirus Diseases.” ACS Central Science6 (3): 315–31.
Liu, Xiaofeng, Richard D. Taylor, Laura Grifn, Shu‑Fen Coker, Ralph Adams, Tom Ceska,
Jiye Shi, Alastair D. G. Lawson, and Terry Baker. 2017. “Computational Design of an
Epitope‑Specic Keap1 Binding Antibody Using Hotspot Residues Grafting and CDR
Loop Swapping.” Scientic Reports 7 (January): 41306.
Liu, Yinhan, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike
Lewis, Luke Zettlemoyer, and Veselin Stoyanov. “RoBERTa: A Robustly Optimized BERT
Pretraining Approach.” arXiv. https://doi.org/10.48550/arXiv.1907.11692
Lu, Ruei‑Min, Yu‑Chyi Hwang, I‑ Ju Liu, Chi‑Chiu Lee, Han‑Zen Tsai, Hsin‑Jung Li, and
Han‑Chung Wu. 2020. “Development of Therapeutic Antibodies for the Treatment of
Diseases.” Journal of Biomedical Science27 (1): 1.
Maia, Eduardo Habib Bechelane, Letícia Cristina Assis, Tiago Alves de Oliveira, Alisson Marques
da Silva, and Alex Gutterres Taranto. 2020. “Structure‑Based Virtual Screening: From
Classical to Articial Intelligence.” Frontiers in Chemistry 8 (April): 343.
Major, Sylvia M., Satoshi Nishizuka, Daisaku Morita, Rick Rowland, Margot Sunshine, Uma
Shankavaram, Frank Washburn, et al. 2006. “AbMiner: A Bioinformatic Resource on
Available Monoclonal Antibodies and Corresponding Gene Identiers for Genomic,
Proteomic, and Immunologic Studies.” BMC Bioinformatics 7 (April): 192.
Makowski, Emily K., Lina Wu, Priyanka Gupta, and Peter M. Tessier. 2021. “Discovery‑Stage
Identication of Drug‑like Antibodies Using Emerging Experimental and Computational
Methods.” mAbs 13 (1): 1895540.
Mannar, Dhiraj, James W. Saville, Zehua Sun, Xing Zhu, Michelle M. Marti, Shanti S. Srivastava,
Alison M. Berezuk, etal. 2022. “SARS‑CoV‑2 Variants of Concern: Spike Protein Mutational
Analysis and Epitope for Broad Neutralization.” Nature Communications 13 (1): 4696.
Marks, Claire, Alissa M. Hummer, Mark Chin, and Charlotte M. Deane. 2021. “Humanization
of Antibodies Using a Machine Learning Approach on Large‑Scale Repertoire Data.”
Bioinformatics, June. https://doi.org/10.1093/bioinformatics/btab434.
Marks, Debora S., Lucy J. Colwell, Robert Sheridan, Thomas A. Hopf, Andrea Pagnani, Riccardo
Zecchina, and Chris Sander. 2011. “Protein 3D Structure Computed from Evolutionary
Sequence Variation.” PloS One 6 (12): e28766.
Mason, Derek M., Simon Friedensohn, Cédric R. Weber, Christian Jordi, Bastian Wagner, Simon
M. Meng, Roy A. Ehling, etal. 2021. “Optimization of Therapeutic Antibodies by Predicting
Antigen Specicity from Antibody Sequence via Deep Learning.” Nature Biomedical
Engineering 5 (6): 600–612.
Miho, Enkelejda, Alexander Yermanos, Cédric R. Weber, Christoph T. Berger, Sai T. Reddy,
and Victor Greiff. 2018. “Computational Strategies for Dissecting the High‑Dimensional
Complexity of Adaptive Immune Repertoires.” Frontiers in Immunology 9 (February): 224.
Mikolov, Tomas, Ilya Sutskever, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. “Distributed
Representations of Words and Phrases and Their Compositionality.” arXiv [cs.CL]. arXiv.
https://proceedings.neurips.cc/paper/2013/hash/9aa42b31882ec039965f3c4923ce901b‑Ab
stract.html.
Młokosiewicz, Jakub, Piotr Deszyński, Wiktoria Wilman, Igor Jaszczyszyn, Rajkumar Ganesan,
Aleksandr Kovaltsuk, Jinwoo Leem, Jacob D. Galson, and Konrad Krawczyk. 2022.
“AbDiver: A Tool to Explore the Natural Antibody Landscape to Aid Therapeutic Design.”
Bioinformatics 38 (9): 2628–30.
Narayanan, Harini, Fabian Dingfelder, Alessandro Butté, Nikolai Lorenzen, Michael Sokolov,
and Paolo Arosio. 2021. “Machine Learning for Biologics: Opportunities for Protein
Engineering, Developability, and Formulation.” Trends in Pharmacological Sciences 42
(3): 151–65.

112 Biopharmaceutical Informatics
Nguyen, Minh N., Mohan R. Pradhan, Chandra Verma, and Pingyu Zhong. 2017. “The Interfacial
Character of Antibody Paratopes: Analysis of Antibody‑Antigen Structures.” Bioinformatics
33 (19): 2971–76.
Norman, Richard A., Francesco Ambrosetti, Alexandre M. J. J. Bonvin, Lucy J. Colwell, Sebastian
Kelm, Sandeep Kumar, and Konrad Krawczyk. 2020. “Computational Approaches to
Therapeutic Antibody Design: Established Methods and Emerging Trends.” Briengs in
Bioinformatics 21 (5): 1549–67.
North, Benjamin, Andreas Lehmann, and Roland L. Dunbrack Jr. 2011. “A New Clustering of
Antibody CDR Loop Conformations.” Journal of Molecular Biology 406 (2): 228–56.
Obrezanova, Olga, Andreas Arnell, Ramón Gómez de la Cuesta, Maud E. Berthelot, Thomas R.
A. Gallagher, Jesús Zurdo, and Yvette Stallwood.2015. “Aggregation Risk Prediction for
Antibodies and Its Application to Biotherapeutic Development.” mAbs 7 (2): 352–63.
Olimpieri, Pier Paolo, Anna Chailyan, Anna Tramontano, and Paolo Marcatili. 2013. “Prediction
of Site‑Specic Interactions in Antibody‑Antigen Complexes: The proABC Method and
Server.” Bioinformatics 29 (18): 2285–91.
Olsen, Tobias H., Fergus Boyles, and Charlotte M. Deane. 2022a. “Observed Antibody Space:
A Diverse Database of Cleaned, Annotated, and Translated Unpaired and Paired Antibody
Sequences.” Protein Science: A Publication of the Protein Society 31 (1): 141–46.
Olsen, Tobias H., Iain H. Moal, and Charlotte M. Deane. 2022b. “AbLang: An Antibody
Language Model for Completing Antibody Sequences.” bioRxiv. https://doi.
org/10.1101/2022.01.20.477061.
Ostrovsky‑Berman, Miri, Boaz Frankel, Pazit Polak, and Gur Yaari. 2021. “Immune2vec:
Embedding B/T Cell Receptor Sequences in RN Using Natural Language Processing.”
Frontiers in Immunology 12. https://doi.org/10.3389/mmu.2021.680687.
Peng, Hung‑Pin, Kuo Hao Lee, Jhih‑Wei Jian, and An‑Suei Yang. 2014. “Origins of Specicity
and Afnity in Antibody–protein Interactions.” Proceedings of the National Academy of
Sciences 111 (26): E2656–65.
Pittala, Srivamshi, and Chris Bailey‑Kellogg. 2020. “Learning Context‑Aware Structural
Representations to Predict Antigen and Antibody Binding Interfaces.” Bioinformatics 36
(13): 3996–4003.
Prabakaran, R., Puneet Rawat, Sandeep Kumar, and M. Michael Gromiha. 2021. “ANuPP: A
Versatile Tool to Predict Aggregation Nucleating Regions in Peptides and Proteins.” Journal
of Molecular Biology 433 (11): 166707.
Prihoda, David, Jad Maamary, Andrew Waight, Veronica Juan, Laurence Fayadat‑Dilman, Daniel
Svozil, and Danny A. Bitton. 2022. “BioPhi: A Platform for Antibody Design, Humanization,
and Humanness Evaluation Based on Natural Antibody Repertoires and Deep Learning.”
mAbs 14 (1): 2020203.
Rawat, Puneet, R. Prabakaran, Sandeep Kumar, and M. Michael Gromiha. 2021a. “Exploring the
Sequence Features Determining Amyloidosis in Human Antibody Light Chains.” Scientic
Reports 11 (1): 13785.
Rawat, Puneet, R. Prabakaran, Sandeep Kumar, and M. Michael Gromiha. 2021b. “AbsoluRATE:
An in‑Silico Method to Predict the Aggregation Kinetics of Native Proteins.” Biochimica et
Biophysica Acta: Proteins and Proteomics 1869 (9): 140682.
Raybould, Matthew I. J., Aleksandr Kovaltsuk, Claire Marks, and Charlotte M. Deane. 2021.
“CoV‑AbDab: The Coronavirus Antibody Database.” Bioinformatics 37 (5): 734–35.
Raybould, Matthew I. J., Claire Marks, Konrad Krawczyk, Bruck Taddese, Jaroslaw Nowak, Alan
P. Lewis, Alexander Bujotzek, Jiye Shi, and Charlotte M. Deane. 2019. “Five Computational
Developability Guidelines for Therapeutic Antibody Proling.” Proceedings of the National
Academy of Sciences of the United States of America 116 (10): 4025–30.

5 • Articial Intelligence and Machine Learning 113
Raybould, Matthew I. J., Claire Marks, Alan P. Lewis, Jiye Shi, Alexander Bujotzek, Bruck
Taddese, and Charlotte M. Deane. 2020. “Thera‑SAbDab: The Therapeutic Structural
Antibody Database.” Nucleic Acids Research 48 (D1): D383–88.
Regep, Cristian, Guy Georges, Jiye Shi, Bojana Popovic, and Charlotte M. Deane. 2017. “The H3
Loop of Antibodies Shows Unique Structural Characteristics.” Proteins 85 (7): 1311–18.
Renaud, Nicolas, Cunliang Geng, Sonja Georgievska, Francesco Ambrosetti, Lars Ridder, Dario
F. Marzella, Manon F. Réau, Alexandre M. J. J. Bonvin, and Li C. Xue. 2021. “DeepRank:
A Deep Learning Framework for Data Mining 3D Protein‑Protein Interfaces.” Nature
Communications 12 (1): 7068.
Repecka, D., V. Jauniskis, and L. Karpus, etal. 2021. “Expanding Functional Protein Sequence
Spaces Using Generative Adversarial Networks.” Nature Machine Intelligence3: 324–333.
https://www.nature.com/articles/s42256‑021‑00310‑5.
Richardson, Eve, Jacob D. Galson, Paul Kellam, Dominic F. Kelly, Sarah E. Smith, Anne Palser,
Simon Watson, and Charlotte M. Deane. 2021. “A Computational Method for Immune
Repertoire Mining That Identies Novel Binders from Different Clonotypes, Demonstrated
by Identifying Anti‑Pertussis Toxoid Antibodies.” mAbs 13 (1): 1869406.
Rives, Alexander, Joshua Meier, Tom Sercu, Siddharth Goyal, Zeming Lin, Jason Liu, Demi Guo,
etal. 2021. “Biological Structure and Function Emerge from Scaling Unsupervised Learning
to 250 Million Protein Sequences.” Proceedings of the National Academy of Sciences of the
United States of America 118 (15). https://doi.org/10.1073/pnas.2016239118.
Robert, Philippe A., Rahmad Akbar, Robert Frank, Milena Pavlović, Michael Widrich, Igor Snapkov,
Andrei Slabodkin, etal. 2021. “Unconstrained Generation of Synthetic Antibody‑Antigen
Structures to Guide Machine Learning Methodology for Real‑World Antibody Specicity
Prediction.” bioRxiv. bioRxiv. https://doi.org/10.1101/2021.07.06.451258.
Roberts, Christopher J. 2014. “Therapeutic Protein Aggregation: Mechanisms, Design, and
Control.” Trends in Biotechnology 32 (7): 372–80.
Ruffolo, Jeffrey A., Jeffrey J. Gray, and Jeremias Sulam. 2021. “Deciphering Antibody Afnity
Maturation with Language Models and Weakly Supervised Learning.” arXiv [q‑bio.BM].
arXiv. https://arxiv.org/abs/2112.07782.
Ruffolo, Jeffrey A., Carlos Guerra, Sai Pooja Mahajan, Jeremias Sulam, and Jeffrey J. Gray.
2020. “Geometric Potentials from Deep Learning Improve Prediction of CDR H3 Loop
Structures.” Bioinformatics 36 (Suppl_1): i268–75.
Ruffolo, Jeffrey A., Jeremias Sulam, and Jeffrey J. Gray. 2022. “Antibody Structure Prediction
Using Interpretable Deep Learning.” Patterns (New York, N.Y.) 3 (2): 100406.
Saha, Sudipto, Manoj Bhasin, and Gajendra P. S. Raghava. 2005. “Bcipep: A Database of B‑Cell
Epitopes.” BMC Genomics 6 (May): 79.
Sang, Zhe, Yufei Xiang, Ivet Bahar, and Yi Shi. 2022. “Llamanade: An Open‑Source Computational
Pipeline for Robust Nanobody Humanization.” Structure 30 (3): 418–29.e3.
Schmitz, Samuel, Cinque Soto, James E. Crowe Jr, and Jens Meiler. 2020. “Human‑Likeness of
Antibody Biologics Determined by Back‑Translation and Comparison with Large Antibody
Variable Gene Repertoires.” mAbs 12 (1): 1758291.
Schneider, Constantin, Andrew Buchanan, Bruck Taddese, and Charlotte M. Deane. 2021.
“DLAB‑Deep Learning Methods for Structure‑Based Virtual Screening of Antibodies.”
Bioinformatics, September. https://doi.org/10.1093/bioinformatics/btab660.
Sharma, Vikas K., Thomas W. Patapoff, Bruce Kabakoff, Satyan Pai, Eric Hilario, Boyan Zhang,
Charlene Li, etal. 2014. “In Silico Selection of Therapeutic Antibodies for Development:
Viscosity, Clearance, and Chemical Stability.” Proceedings of the National Academy of
Sciences of the United States of America 111 (52): 18601–6.

114 Biopharmaceutical Informatics
Sheng, Zizhang, Chaim A. Schramm, Rui Kong, NISC Comparative Sequencing Program, James
C. Mullikin, John R. Mascola, Peter D. Kwong, and Lawrence Shapiro. 2017. “Gene‑Specic
Substitution Proles Describe the Types and Frequencies of Amino Acid Changes during
Antibody Somatic Hypermutation.” Frontiers in Immunology 8 (May): 537.
Shin, Jung‑Eun, Adam J. Riesselman, Aaron W. Kollasch, Conor McMahon, Elana Simon, Chris
Sander, Aashish Manglik, Andrew C. Kruse, and Debora S. Marks. 2021. “Protein Design
and Variant Prediction Using Autoregressive Generative Models.” Nature Communications
12 (1): 2403.
Singh, Satish Kumar. 2011. “Impact of Product‑Related Factors on Immunogenicity of
Biotherapeutics.” Journal of Pharmaceutical Sciences 100 (2): 354–87.
Sircar, Aroop, and Jeffrey J. Gray. 2010. “SnugDock: Paratope Structural Optimization during
Antibody‑Antigen Docking Compensates for Errors in Antibody Homology Models.” PLoS
Computational Biology 6 (1): e1000644.
Sirin, Sarah, James R. Apgar, Eric M. Bennett, and Amy E. Keating. 2016. “AB‑Bind: Antibody
Binding Mutational Database for Computational Afnity Predictions.” Protein Science: A
Publication of the Protein Society 25 (2): 393–409.
Smialowski, Pawel, Gero Doose, Phillipp Torkler, Stefanie Kaufmann, and Dmitrij Frishman.
2012. “PROSO II‑‑a New Method for Protein Solubility Prediction.” The FEBS Journal 279
(12): 2192–2200.
Sormanni, Pietro, Francesco A. Aprile, and Michele Vendruscolo. 2015. “The CamSol Method
of Rational Design of Protein Mutants with Enhanced Solubility.” Journal of Molecular
Biology 427 (2): 478–90.
Sormanni, Pietro, Damiano Piovesan, Gabriella T. Heller, Massimiliano Bonomi, Predrag Kukic,
Carlo Camilloni, Monika Fuxreiter, etal. 2017. “Simultaneous Quantication of Protein
Order and Disorder.” Nature Chemical Biology 13 (4): 339–42.
Swindells, Mark B., Craig T. Porter, Matthew Couch, Jacob Hurst, K. R. Abhinandan, Jens H.
Nielsen, Gary Macindoe, James Hetherington, and Andrew C. R. Martin. 2017. “abYsis:
Integrated Antibody Sequence and Structure‑Management, Analysis, and Prediction.”
Journal of Molecular Biology 429 (3): 356–64.
Tartaglia, Gian Gaetano, Andrea Cavalli, Riccardo Pellarin, and Amedeo Caisch. 2005. “Prediction
of Aggregation Rate and Aggregation‑Prone Segments in Polypeptide Sequences.” Protein
Science: A Publication of the Protein Society 14 (10): 2723–34.
Tomar, Dheeraj S., Li, Matthew P. Broulidakis, Nicholas G. Luksha, Christopher T. Burns, Satish
K. Singh, and Sandeep Kumar. 2017. “In‑Silico Prediction of Concentration‑Dependent
Viscosity Curves for Monoclonal Antibody Solutions.” mAbs 9 (3): 476–89.
Torjesen, Ingrid. n.d. “Drug Development: The Journey of a Medicine from Lab to Shelf.” The
Pharmaceutical Journal, doi: 10.1211/PJ.2015.20068196
Toseland, Christopher P., Debra J. Clayton, Helen McSparron, Shelley L. Hemsley, Martin J.
Blythe, Kelly Paine, Irini A. Doytchinova, Pingping Guan, Channa K. Hattotuwagama,
and Darren R. Flower. 2005. “AntiJen: A Quantitative Immunology Database Integrating
Functional, Thermodynamic, Kinetic, Biophysical, and Cellular Data.” Immunome Research
1 (1): 4.
Tubiana, Jérôme, Dina Schneidman‑Duhovny, and Haim J. Wolfson. 2022. “ScanNet: An
Interpretable Geometric Deep Learning Model for Structure‑Based Protein Binding Site
Prediction.” Nature Methods 19 (6): 730–39.
Tunyasuvunakool, Kathryn, Jonas Adler, Zachary Wu, Tim Green, Michal Zielinski, Augustin
Žídek, Alex Bridgland, etal. 2021. “Highly Accurate Protein Structure Prediction for the
Human Proteome.” Nature 596 (7873): 590–96.
Vashchenko, Denis, Sam Nguyen, Andre Goncalves, Felipe Leno da Silva, Brenden Petersen,
Thomas Desautels, and Daniel Faissol. 2022. “AbBERT: Learning Antibody Humanness via
Masked Language Modeling.” bioRxiv. https://doi.org/10.1101/2022.08.02.502236.

5 • Articial Intelligence and Machine Learning 115
Vita, Randi, James A. Overton, Jason A. Greenbaum, Julia Ponomarenko, Jason D. Clark, Jason
R. Cantrell, Daniel K. Wheeler, etal. 2015. “The Immune Epitope Database (IEDB) 3.0.”
Nucleic Acids Research 43 (Database issue): D405–12.
Vita, Randi, Laura Zarebski, Jason A. Greenbaum, Hussein Emami, Ilka Hoof, Nima Salimi,
Rohini Damle, Alessandro Sette, and Bjoern Peters. 2010. “The Immune Epitope Database
2.0.” Nucleic Acids Research 38 (Database issue): D854–62.
Wang, Chunyan, Wentao Li, Dubravka Drabek, Nisreen M. A. Okba, Rien van Haperen, Albert
D. M. E. Osterhaus, Frank J. M. van Kuppeveld, Bart L. Haagmans, Frank Grosveld,
and Berend‑Jan Bosch. 2020a. “A Human Monoclonal Antibody Blocking SARS‑CoV‑2
Infection.” Nature Communications 11 (1): 2251.
Wang, Xiao, Genki Terashi, Charles W. Christoffer, Mengmeng Zhu, and Daisuke Kihara.
2020b. “Protein Docking Model Evaluation by 3D Deep Convolutional Neural Networks.”
Bioinformatics 36 (7): 2113–18.
Weitzner, Brian D., Daisuke Kuroda, Nicholas Marze, Jianqing Xu, and Jeffrey J. Gray. 2014.
“Blind Prediction Performance of RosettaAntibody 3.0: Grafting, Relaxation, Kinematic
Loop Modeling, and Full CDR Optimization.” Proteins 82 (8): 1611–23.
Wilman, Wiktoria, Sonia Wróbel, Weronika Bielska, Piotr Deszynski, Paweł Dudzic, Igor
Jaszczyszyn, Jędrzej Kaniewski, et al. 2022. “Machine‑Designed Biotherapeutics:
Opportunities, Feasibility and Advantages of Deep Learning in Computational Antibody
Discovery.” Briengs in Bioinformatics 23 (4). https://doi.org/10.1093/bib/bbac267.
Wilton, Emily E., Michael P. Opyr, Senthilkumar Kailasam, Ronja F. Kothe, and Hans‑Joachim
Wieden. 2018. “sdAb‑DB: The Single Domain Antibody Database.” ACS Synthetic Biology
7 (11): 2480–84.
Wolf, T., L. Debut, V. Sanh, and J. Chaumond. 2020. “Transformers: State‑of‑the‑Art Natural
Language Processing.” Proceedings of the 2020 Conference on Empirical Methods in
Natural Language Processing: System Demonstrations. https://aclanthology.org/2020.
emnlp‑demos.6/?ref=https://codemonkey.link.
Wollacott, Andrew M., Chonghua Xue, Qiuyuan Qin, June Hua, Tanggis Bohnuud, Karthik
Viswanathan, and Vijaya B. Kolachalama. 2019. “Quantifying the Nativeness of Antibody
Sequences Using Long Short‑Term Memory Networks.” Protein Engineering, Design &
Selection: PEDS 32 (7): 347–54.
Wu, Jiaxiang, Fandi Wu, Biaobin Jiang, Wei Liu, and Peilin Zhao. 2022. “tFold‑Ab: Fast and
Accurate Antibody Structure Prediction without Sequence Homologs.” bioRxiv. https://doi.
org/10.1101/2022.11.10.515918.
Wu, Zachary, S. B. Jennifer Kan, Russell D. Lewis, Bruce J. Wittmann, and Frances H.
Arnold.2019. “Machine Learning‑Assisted Directed Protein Evolution with Combinatorial
Libraries.” Proceedings of the National Academy of Sciences of the United States of America
116 (18): 8852–58.
Xu, Yingda, Dongdong Wang, Bruce Mason, Tony Rossomando, Ning Li, Dingjiang Liu, Jason
K. Cheung, et al. 2019. “Structure, Heterogeneity and Developability Assessment of
Therapeutic Antibodies.” mAbs 11 (2): 239–64.
Zavrtanik, Uroš, and San Hadži. 2019. “A Non‑redundant Data Set of Nanobody‑Antigen Crystal
Structures.” Data in Brief 24 (June): 103754.
Zhang, Jie, Yishan Du, Pengfei Zhou, Jinru Ding, Shuai Xia, Qian Wang, Feiyang Chen, etal. 2022.
“Predicting Unseen Antibodies’ Neutralizability via Adaptive Graph Neural Networks.”
Nature Machine Intelligence, November, 1–13.

From Deep
Generative Models
6
to Structure-Based
Simulations
Computational Approaches
for Antibody Design
Daisuke Kuroda
6.1 INTRODUCTION
In the rapidly evolving eld of biotherapeutics, the intersection of computational science
and protein engineering has revolutionized the approach to drug discovery and develop‑
ment. Central to this transformation is the art of antibody design, a complex and critical
component of modern therapeutic strategies.1 Antibodies, with their unique specicity
and versatility, have emerged as a leading modality in the treatment of various diseases,
ranging from cancers to autoimmune disorders and infectious diseases. The process
of designing these therapeutic antibodies, however, presents a multifaceted challenge,
encompassing aspects of protein engineering, bioinformatics, and developability.
The advent of machine learning (ML) and generative models, including large lan‑
guage models (LLMs), has opened new frontiers in deciphering the complex language of
proteins, particularly in understanding and predicting the vast diversity of the antibody
repertoire.
116
3–5
These computational tools, coupled with advanced molecular simulations,
2

6 • From Deep Generative Models to Structure-Based Simulations 117
enable scientists to explore the structural and functional nuances of antibodies. By lever‑
aging these technologies, the eld of antibody design has transitioned from a largely
empirical endeavor, relying on experimental trial and error, to one that is increasingly
predictive and rational.
Developability, a key consideration in antibody engineering, involves assessing the
suitability of antibody candidates for therapeutic use, focusing on their manufacturabil‑
ity, stability, and efcacy.6 This process has been greatly enhanced by bioinformatics
and computational science, which provide insights into the molecular characteristics
that govern the behavior of antibodies in biological systems. Through computational
approaches, the scope of protein engineering has expanded, allowing for the design of
antibodies with improved properties from their amino acid sequences (Figure6.1). The
properties amenable to computational enhancement encompass physicochemical attri‑
butes, such as binding afnity and stability, as well as biological and pharmacological
aspects, including immunogenicity.
The integration of ML into antibody design is particularly noteworthy. Through
the analysis of vast datasets, including sequences and structures of existing antibod‑
ies, enabled by high‑throughput sequencing of immune repertoires, ML algorithms can
predict the antigen‑binding afnity and specicity of novel antibody candidates.7 This
capability is pivotal in accelerating the drug discovery process, reducing the time and
cost associated with the development of new biotherapeutics.
Among the various algorithms in ML, deep learning (DL) has garnered signicant
attention due to its unparalleled success in mimicking human‑like decision‑making and
FIGURE 6.1 The computational workow in antibody drug discovery: Seed antibodies
are generated through various methods, including animal immunization, synthetic libraries,
single-cell analysis, and computational design. The sequences of these seed antibodies can
be further optimized using computational design calculations based on either sequences or
structures, yielding lead antibody sequences. Structures of the antibody alone, as well as
antibody-antigen complexes, can be predicted using computational methods. Additionally,
developability assessments of lead antibodies can be conducted in silico. Topics discussed in
this chapter are underlined and italicized in red for emphasis.

118 Biopharmaceutical Informatics
learning patterns. This has been especially evident since the early 2010s with break‑
throughs in image and speech recognition.
8,9
Deep learning employs deep neural net‑
works, which comprise several core architectures, such as convolutional neural networks
(CNNs),10 recurrent neural networks (RNNs) or long short‑term memory (LSTM),11gen‑
erative adversarial networks (GANs),12 variational autoencoders (VAEs)13, and trans‑
former14models (Figure6.2). Each architecture is distinguished by its specic strengths,
weaknesses, and applications. These models autonomously learn features directly from
data, thereby eliminating the need for manual feature extraction. The training process,
which optimizes weights through backpropagation and gradient descent, renes learn‑
ing over time. Although DL models reduce the necessity for manual feature engineering
by learning complex data patterns, data pre‑processing is still crucial in the bioinfor‑
matics workow. This ensures data quality and compatibility, which are essential for the
success of DL applications. This aspect becomes particularly critical in antibody design,
where the sizes, amino acid compositions, and structural diversity of the functional
sites, namely the complementarity‑determining regions (CDRs), vary across antibodies.
Such variability complicates the direct comparison of features among antibodies.
Molecular simulations, another vital component of computational antibody engi‑
neering, often rely on structural information of antibodies and their antigens.
15,16
These
simulations play a crucial role in visualizing the dynamic interactions between antibodies
FIGURE6.2 Classication of deep learning architectures: CNN: Designed to process data
with a grid-like topology, using convolutional layers to efciently extract and learn spatial
hierarchies in data, particularly useful in image analysis. RNN: Designed to handle sequential data, where the output from previous steps is fed back into the network to inform
responses at later steps, making them ideal for tasks like language modeling and time-series
analysis. LSTM: An advanced type of RNN that is capable of learning long-term dependencies in sequential data. VAE: Designed to encode data into a latent space and then
reconstruct it, facilitating data generation by sampling from the learned distribution in the
latent space. GAN: Consists of two competing networks: a generator that creates data and
a discriminator that evaluates it, working together to produce highly realistic data samples.
LLM: Designed to understand, generate, and manipulate natural language, often built on
architectures like transformers.

6 • From Deep Generative Models to Structure-Based Simulations 119
and their targets, offering valuable insights into their mechanisms of action and potential
off‑target effects. This informs the design of more effective and safer antibodies.
The synergy between ML, molecular simulations, and immune repertoire analysis
is revolutionizing the eld of antibody design and protein engineering.
studies in the literature have utilized these technologies for a variety of antibody‑related
prediction tasks, including the prediction of antibody structures,
complexes,
from large antibody libraries.
of these methodologies for de novo generation and optimization of antibody sequences.
Therefore, the objective of this chapter is to explore these technological advancements,
with a particular emphasis on their impact on antibody sequence design for next‑genera‑
tion biotherapeutics and their inuence on the future of antibody drug discovery.
25–30
their biophysical properties,
41– 45
However, this chapter specically focuses on the use
31– 40
and the identication of lead candidates
17,18
Numerous
19–2 4
antibody‑antigen
6.2 ANTIBODY GENERATION THROUGH DEEP GENERATIVE MODELS
Generating antibody sequences is a critical task in developing therapeutic antibodies and
diagnostic tools and conducting basic immunology research (Figure 6.1). Traditional
methods like animal immunization and synthetic libraries offer distinct approaches to
acquiring antibodies with the desired specicities and afnities. Each method has its
strengths and limitations, with the choice often dictated by project‑specic factors such
as the nature of antigens, the intended antibody use, ethical considerations, and avail‑
able resources.
The transition from these conventional experimental methods to incorporating
deep generative models into antibody design signies a paradigm shift in biotherapeutic
development. Originally conceived for processing human languages, LLMs are now
ingeniously repurposed to unravel the complex language of proteins.46 This adapta‑
tion not only highlights the versatility of LLMs but also underscores the parallels in
pattern recognition and sequence analysis between linguistics and molecular biology.
Leveraging their core capabilities in sequence, context, and pattern recognition, LLMs
offer a fresh approach to interpreting protein sequences, analogous to parsing sentences
in a natural language. This innovative convergence of computational linguistics and
molecular biology equips us with potent tools to advance protein engineering and anti‑
body design. This section reviews the considerations involved in computationally gener‑
ating antibodies de novo through LLMs and other DL‑based techniques.
6.2.1 B‑Cell Repertoires in the Era of Articial
Intelligence
ML methods fundamentally rely on data to uncover the intricate patterns of life. This is
equally true for LLMs focused on generating antibody sequences, which benet from

120 Biopharmaceutical Informatics
the rich data provided by B‑cell repertoires. These repertoires, now more accessible
due to breakthroughs in high‑throughput sequencing technologies, have revolutionized
elds such as immunology, vaccine development, and therapeutic antibody discovery.3
Enhancements in sequencing and computational analysis, coupled with a deeper under‑
standing of the immune system, have made these advances possible. The decreasing
cost of sequencing and increased throughput capacity have made large‑scale genomic
projects and the sequencing of vast cohorts and diverse species more practical, achiev‑
ing previously unimaginable depth.
Integrating B‑cell repertoire sequencing with other “omics” data, like proteomics
and transcriptomics, offers a comprehensive view of the immune response.
47,4 8
Such a
holistic approach can elucidate the evolution of repertoires in response to disease pro‑
gression, vaccination, or therapeutic interventions. Signature changes in the B‑cell rep‑
ertoire, linked to various diseases, including autoimmune disorders, infectious diseases,
and cancers, have been identied.
49–53
These signatures have potential as biomarkers for
diagnosis and prognosis. Sequencing efforts pre‑ and post‑vaccination have shed light
on vaccine‑induced immunity, informing the design of more effective vaccines.
54–56
Moreover, high‑throughput sequencing has been instrumental in identifying antibodies
with therapeutic potential against targets like severe acute respiratory syndrome coro‑
navirus 2 (SARS‑CoV‑2).
57
The development of specialized bioinformatics pipelines and software tools has
enhanced the accuracy of sequence assembly, clonotype identication, and lineage trac‑
ing.58 These tools are designed to manage the vast datasets produced, enabling detailed
analyses of B‑cell repertoires. By analyzing intricate patterns in repertoire data, com‑
putational algorithms guide both vaccine design and therapeutic antibody design.
4,59
Notable successes include employing ML, trained on high‑throughput sequencing data,
to detect changes in B‑cell repertoire patterns in patients with dengue infection60 and
relapsing‑remitting multiple sclerosis.61 These studies have demonstrated the capability
of ML algorithms to capture the nuances of B‑cell repertoires, a task that would be chal‑
lenging for humans without the aid of ML.
In the realm of articial intelligence, representation learning, a key technique in
deep learning, automates the discovery of data representations for feature detection or
classication, bypassing the need for manual feature engineering. This approach allows
models to identify complex patterns, enhancing performance in tasks such as classi‑
cation and prediction. In protein modeling and design, representation learning is criti‑
cal for encoding protein sequences, structures, or functional features for computational
tasks. The signicance of representation stems from its ability to translate the biological
essence of a protein into a format that computational models, particularly those rooted
in deep learning, can efciently interpret and learn from.62 The choice of represen‑
tation—ranging from direct sequences and structures to more abstract feature‑based
representations and embeddings—signicantly inuences model performance in pre‑
dicting protein structure and function and designing novel proteins (Figure6.3).
Effective representation captures essential protein features relevant to biological
functions, facilitating accurate predictions and the design of proteins with new prop‑
erties. It is particularly vital in generative models aimed at creating novel protein
sequences with specic functions. Here, representation learning navigates the protein
Соседние файлы в папке Библиотека им академика М.И. Перельмана
