Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5629_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
15.09.2026
Размер:
16 Мб
Скачать
☆
5 • Articial Intelligence and Machine Learning 101
5.4.3.3 Transformer architectures
Recently, transformer architectures have been gaining popularity in the representation learning area as they achieved state‑of‑the‑art results on a wide range of NLP tasks (Wolf etal. 2020). Using self‑attention mechanisms, transformers are able to learn rela‑ tionships between words in sentences and, therefore, produce distributed representa‑ tions. Such neural networks are often trained with the MLM approach in which one masks or changes part of the input and the model learns to predict the altered part.
5.4.3.4 ProtBERT
To this time, there have been several adaptations of transformer architectures to antibody data. For example, Elnaggar and colleagues (Elnaggar etal., n.d.) showed that embeddings obtained from transformers captured relevant biological information as small‑size mod‑ els trained solely on those representations were able to compete with bigger architectures on various tasks such as protein classication. Their neural networks reached compa‑ rable performance to methods utilizing MSA, which indicates that transformer‑produced embeddings could be a good starting point for various downstream tasks.
5.4.3.5 AbLang
This transformer architecture was also successfully applied in the eld of antibody representation by Olsen and colleagues (Olsen etal. 2022b) who trained the AbLang model on antibody sequences from OAS. The resulting architecture consists of two parts: AbRep and AbHead, which produce representations for sequences and predict the probability of each amino acid on all positions, respectively. Pre‑training was based on the RoBERTa approach (Liu etal. 2019). The team showed the biological information encoded into the vectors by drawing 10,000 naïve and 10,000memory B‑cell sequence representations (Ghraichy etal. 2021) using t‑SNE and compared the results to evo‑ lutionary scale modeling 1b (ESM‑1b) embeddings (Rives etal. 2021). Both models could separate antibody sequences by their V gene families, but AbLang yielded better separation of naïve and memory B cells. The resulting transformer model was capable of restoring missing residues in immunoglobulin sequences, obtaining similar or better results than using IMGT germlines, but without the knowledge of the germlines.
5.4.3.6 AntiBERTa
As proposed by Leem and colleagues (Leem etal. 2022), AntiBERTa (Antibody‑specic Bidirectional Encoder Representation from Transformers) is another example of trans‑ former architecture. The model was pre‑trained using the RoBERTa approach on 57mil‑ lion human BCR sequences from 61 studies available in OAS. The team selected random 1000BCR heavy‑chain sequences (Ghraichy etal. 2021) from naïve and memory B‑cell sequences and showed that on the top of mutational load and V gene used, embeddings carry information about B‑cell type–they were able to partition naïve and memory B‑cell sequences in the representation space. This result was compared to embeddings prepared
102 Biopharmaceutical Informatics
by ProtBERT (Elnaggar etal. 2020), where the separation was not as clear, which indi‑ cates that compared to the general protein transformer model, AntiBERTa representa‑ tions encapsulate more antibody‑specic information. Leem and colleagues noticed that high self‑attention scores presented residue pairs of contacts indicating that the model is capable of understanding the structural information. So, they applied it to the para‑ tope prediction problem, where the model classied each residue from the input sequence. This was achieved by adding a classication head on top of the already existing 12lay‑ ers. The prediction results were compared to Parapred (Liberis etal. 2018) and ProABC (Olimpieri etal. 2013), which demonstrated SoTa results in paratope prediction. The team used the model to produce embeddings of known therapeutic antibodies, showing that it was possible to determine their origin (human, murine, humanized, chimeric) as well as, to a certain degree, the correlation with immunogenicity response scores–ADA. This demonstrates that embeddings learned biologically relevant information, and the learned representations correspond to B‑cell origin, immunogenicity, and structure.
5.4.3.7 AntiBERTy
AntiBERTy, another model based on the BERT architecture, was proposed by Ruffolo and colleagues (Ruffolo etal. 2021). It was trained on 558million sequences from OAS with MLM objective. The team analyzed repertoires from donors with HIV‑1 neutral‑ izing VRC01 antibodies. For each sample, they created a k bearers neighbour (kNN) graph using model embeddings and visualized it in two‑dimensions using Uniform Manifold Approximation and Projection (UMAP). Using these plots, they observed tra‑ jectories from germline sequences and mutated derivatives corresponding to sequence changes in the afnity maturation process. With repertoire data, individual sequences are not labeled. Hence–relying on clonal expansion–the team produced noisy labels, and frequently observed sequences were assumed to be binders. Next, they applied mul‑ tiple instance learning (MIL) to predict whether the sets of sequences contain binding antibodies. They created single‑instance bags of sequences from known VRC01 anti‑ bodies, conrming that the model produces the correct positive predictions. Finally, they annotated each antibody structure with attention, and in most cases (7 out of 10 sequences), attention‑pointed binding residues.
5.4.3.8 AbBERT
Another transformer architecture called AbBERT trained on 20million heavy and light sequences from OAS was published by Vashchenko and colleagues (Vashchenko etal.
2022). The model family is based on ProtBERT, but ne‑tuned on antibodies. Both heavy and light sequences were used to train the models, during which the team anno‑ tated functional regions of input sequences by inserting additional annotation tokens, before and after all CDRs. First, the authors showed that the model is capable of predict‑ ing CDR regions. The team introduces the term “humanness” that is used to measure the similarity between input immunoglobulin sequence and antibodies sampled from people. This score was used to evaluate 600 known therapeutic sequences that have passed various clinical trials and showed that antibodies with a low assigned metric tend
5 • Articial Intelligence and Machine Learning 103
to be immunogenic. Then, the model was applied for in silico antibody optimization in which the baseline anti‑SARS‑CoV‑1 antibody sequence was modied so that the opti‑ mized immunoglobulin was able to bind to another target–SARS‑CoV‑2. Model with AbBERT embeddings on input was used to solve the optimization problem. Finally, they performed in vitro experiments which demonstrated that poorly scored sequences were weakly expressed in the cells. They have also observed a correlation between the model scores and the protein stability metrics calculated using Free Energy Perturbation.
5.4.3.9 BioPhi (Sapiens module)
Sapiens is one of the two BioPhi models aiming at antibody humanization. Similarly, to other SoTa models, it is a transformer‑based model trained toward the MLM goal. Two separate models for light and heavy chains have been created, each having 568,857 parameters and being based on the RoBERTa model. The training dataset consisted of human‑only and unaligned sequences from OAS: 20million heavy and 19million light sequences. Analysis of the averaged attention matrix showed high importance between CDR loops, which are close structurally but apart in sequence–thus proving that the model is able to recognize long‑range interactions.
Antibody humanization works by leveraging the fact that only human mAbs were used for training. Input variable region sequence is processed by the model, giving prob‑ abilities for all 20 amino acids for all positions. The most probable residues for frame‑ works are selected, keeping unchanged CDRs from input, which minimizes the risk of affecting binding properties but making antibodies more similar to human ones. Such a procedure has been performed on 177 antibodies (25 with known parental sequence and 152 humanized mAbs with presumed original sequence), obtaining results comparable to human experts.
5.5 CONCLUSIONS AND FUTURE PERSPECTIVES IN AI FOR
ANTIBODY DISCOVERY
Over the past 40 years, antibodies have rmly established their role as the most important group of biologics. Up until now, the development of currently 100 approved antibody therapeutics relied on a “discovery” process driven by experimental laboratory‑based methods. Thanks to advances in high‑throughput experimental data generation as well as progress in computational model development, it is possible to shift the paradigm from antibody discovery toward “design.”
For designing a novel biologic computationally, one requires two elements. First, one needs to generate biologically or physically plausible sequences and structures. Second, one requires objective functions to gauge whether the molecule has the proper‑ ties expected of it. The sampling of novel molecules has been greatly facilitated by gen‑ erative modeling such as variational auto encoders (VAEs), GANs, and language models
104 Biopharmaceutical Informatics
that learn the representation of antibodies from large‑scale NGS data. The performance of predicting objective antibody features such as antibody‑antigen binding or develop‑ ability still needs to be addressed. Nevertheless, researchers have started to combine the two features, generating molecules that are either biased or ltered for those with better biophysical features.
As such, computational methods are now capable of producing naturally viable starting points that are free from statistically obvious liabilities. Achieving the goal of fully computational antibody design–as opposed to “discovery”–still requires improv‑ ing the prediction of the molecule’s therapeutic features, chiey binding and develop‑ ability. On the binding front, one could hope for a modeling revolution, on par with structure prediction as the two problems bear many parallels. However, on the develop‑ ability front, hoping for such progress is fanciful, mostly due to the lack of data.
Developability is an umbrella term uniting multiple biological assays. Even though many of these are regularly performed at organizations developing biologics, such a plethora of data was scarcely envisaged for training models. Therefore, data are often‑ times not comparable between different runs, projects, and teams since they were generated with a specic therapeutic challenge rather than to develop a generalistic developability prediction method. For this reason, a new paradigm emerges called “pre‑ diction‑rst,” where data are generated specically with model training in mind. Over the short term, they might not contribute to any therapeutic projects, but rather act as a long‑term investment into the development of a foundation for a broadly applicable computational model.
All in all, shifting from discovery to design and from project‑driven data genera‑ tion toward prediction‑rst requires a sizable shift within the organizations responsible for biologics development. Understandably, it is a large diversion of resources from well‑proved experimental methods to the development of innovative methods that still need to be validated. Nevertheless, with the ongoing progress in the development of computational models for antibodies, such a shift has become far more realistic.

REFERENCES

Abanades, Brennan, Guy Georges, Alexander Bujotzek, and Charlotte M. Deane. 2022a.
“ABlooper: Fast Accurate Antibody CDR Loop Structure Prediction with Accuracy Estimation.” Bioinformatics, January. https://doi.org/10.1093/bioinformatics/btac016.
Abanades, Brennan, Wing Ki Wong, Fergus Boyles, Guy Georges, Alexander Bujotzek, and
Charlotte M. Deane. 2022b. “ImmuneBuilder: Deep‑Learning Models for Predicting the Structures of Immune Proteins.” Commun Biol. 6 (1):575. doi: 10.1038/s42003‑023‑04927‑7.
Abhinandan, K. R., and Andrew C. R. Martin. 2007. “Analyzing the ‘Degree of Humanness’ of
Antibody Sequences.” Journal of Molecular Biology 369 (3): 852–62.
Adolf‑Bryfogle, Jared, Oleks Kalyuzhniy, Michael Kubitz, Brian D. Weitzner, Xiaozhen Hu, Yumiko
Adachi, William R. Schief, and Roland L. Dunbrack Jr. 2018. “RosettaAntibodyDesign (RAbD): A General Framework for Computational Antibody Design.” PLoS Computational Biology 14 (4): e1006112.
5 • Articial Intelligence and Machine Learning 105
Adolf‑Bryfogle, Jared, Qifang Xu, Benjamin North, Andreas Lehmann, and Roland L. Dunbrack
Jr.2015. “PyIgClassify: A Database of Antibody CDR Structural Classications.” Nucleic Acids Research 43 (Database issue): D432–38.
Agrawal, Neeraj J., Bernhard Helk, Sandeep Kumar, Neil Mody, Hasige A. Sathish, Hardeep S.
Samra, Patrick M. Buck, Li, and Bernhardt L. Trout. 2016. “Computational Tool for the Early Screening of Monoclonal Antibodies for Their Viscosities.” mAbs 8 (1): 43–48.
Akbar, Rahmad, Philippe A. Robert, Cédric R. Weber, Michael Widrich, Robert Frank, Milena
Pavlović, Lonneke Scheffer, et al. 2022. “In Silico Proof of Principle of Machine Learning‑Based Antibody Design at Unconstrained Scale.” mAbs 14 (1): 2031482.
Akpinaroglu, Deniz, Jeffrey A. Ruffolo, Sai Pooja Mahajan, and Jeffrey J. Gray. 2022.
“Simultaneous Prediction of Antibody Backbone and Side‑Chain Conformations with Deep Learning.” PloS One 17 (6): e0258173.
Almagro, Juan C., Alexey Teplyakov, Jinquan Luo, Raymond W. Sweet, Sreekumar Kodangattil,
Francisco Hernandez‑Guzman, and Gary L. Gilliland.2014. “Second Antibody Modeling Assessment (AMA‑II).” Proteins 82 (8): 1553–62.
Ambrosetti, Francesco, Brian Jiménez‑García, Jorge Roel‑Touris, and Alexandre M. J. J. Bonvin.
2020. “Modeling Antibody‑Antigen Complexes by Information‑Driven Docking.” Structure 28 (1): 119–29.e2.
Amimeur, Tileli, Jeremy M. Shaver, Randal R. Ketchem, J. Alex Taylor, Rutilio H. Clark,
Josh Smith, Danielle Van Citters, etal. 2020. “Designing Feature‑Controlled Humanoid Antibody Discovery Libraries Using Generative Adversarial Networks.” bioRxiv. https:// doi.org/10.1101/2020.04.12.024844.
Anand, Namrata, and Possu Huang. 2018. “Generative Modeling for Protein Structures.”
https://papers.nips.cc/paper_files/paper/2018/hash/afa299a4d1d8c52e75dd8a24c‑ 3ce534f‑Abstract.html.
Anonymous. 2022. “xTrimoABFold: Improving Antibody Structure Prediction without Multiple
Sequence Alignments.” https://arxiv.org/abs/2212.00735.
Asgari, Ehsaneddin, and Mohammad R. K. Mofrad.2015. “Continuous Distributed Representation
of Biological Sequences for Deep Proteomics and Genomics.” PloS One 10 (11): e0141287.
Bachas, Sharrol, Goran Rakocevic, David Spencer, Anand V. Sastry, Robel Haile, John M. Sutton,
George Kasun, et al. 2022. “Antibody Optimization Enabled by Articial Intelligence Predictions of Binding Afnity and Naturalness.” bioRxiv. https://doi.org/10.1101/2022.08.
16.504181.
Baran, Dror, M. Gabriele Pszolla, Gideon D. Lapidoth, Christoffer Norn, Orly Dym, Tamar Unger,
Shira Albeck, Michael D. Tyka, and Sarel J. Fleishman. 2017. “Principles for Computational Design of Binding Antibodies.” Proceedings of the National Academy of Sciences of the United States of America 114 (41): 10900–905.
Brinda, K. V., and Saraswathi Vishveshwara. 2005. “A Network Representation of Protein
Structures: Implications for Protein Stability.” Biophysical Journal 89 (6): 4159–70.
Briney, Bryan, Anne Inderbitzin, Collin Joyce, and Dennis R. Burton. 2019. “Commonality
despite Exceptional Diversity in the Baseline Human Antibody Repertoire.” Nature 566 (7744): 393–97.
Carter, and Lazar. n.d. “Next Generation Antibody Drugs: Pursuit of The ‘high‑Hanging Fruit’.”
Nature Reviews. Drug Discovery. https://www.nature.com/articles/nrd.2017.227.
Chakrabarty, Broto, and Nita Parekh. 2016. “NAPS: Network Analysis of Protein Structures.”
Nucleic Acids Research 44 (W1): W375–82.
Chen, Rong, Li Li, and Zhiping Weng. 2003. “ZDOCK: An Initial‑Stage Protein‑Docking
Algorithm.” Proteins 52 (1): 80–87.
Chen, Xingyao, Thomas Dougherty, Chan Hong, Rachel Schibler, Yi Cong Zhao, Reza Sadeghi, Naim
Matasci, Yi‑Chieh Wu, and Ian Kerman. 2020. “Predicting Antibody Developability from Sequence Using Machine Learning.” bioRxiv. https://doi.org/10.1101/2020.06.18.159798.
106 Biopharmaceutical Informatics
Chennamsetty, Naresh, Vladimir Voynov, Veysel Kayser, Bernhard Helk, and Bernhardt L.
Trout. 2009. “Design of Therapeutic Proteins with Enhanced Stability.” Proceedings of the National Academy of Sciences of the United States of America 106 (29): 11937–42.
Chinery, Lewis, Newton Wahome, Iain Moal, and Charlotte M. Deane. 2022. “Paragraph–Antibody
Paratope Prediction Using Graph Neural Networks with Minimal Feature Vectors.” Bioinformatics. 39 (1):btac732. doi: 10.1093/bioinformatics/btac732. PMID: 36370083.
Chothia, C., and A. M. Lesk. 1987. “Canonical Structures for the Hypervariable Regions of
Immunoglobulins.” Journal of Molecular Biology 196 (4): 901–17.
Clark, Karen, Ilene Karsch‑Mizrachi, David J. Lipman, James Ostell, and Eric W. Sayers. 2016.
“GenBank.” Nucleic Acids Research 44 (D1): D67–72.
Clavero‑Álvarez, Alejandro, Tomas Di Mambro, Sergio Perez‑Gaviro, Mauro Magnani, and
Pierpaolo Bruscolini. 2018. “Humanization of Antibodies Using a Statistical Inference Approach.” Scientic Reports 8 (1): 14820.
Cohen, Tomer, Matan Halfon, and Dina Schneidman‑Duhovny. 2022. “NanoNet: Rapid and
Accurate End‑to‑End Nanobody Modeling by Deep Learning.” Frontiers in Immunology 13 (August): 958584.
Corrie, Brian D., Nishanth Marthandan, Bojan Zimonja, Jerome Jaglale, Yang Zhou, Emily Barr,
Nicole Knoetze, etal. 2018. “iReceptor: A Platform for Querying and Analyzing antibody/ B‑Cell and T‑Cell Receptor Repertoire Data across Federated Repositories.” Immunological Reviews 284 (1): 24–41.
David, Maria Pamela C., Gisela P. Concepcion, and Eduardo A. Padlan. 2010. “Using Simple
Articial Intelligence Methods for Predicting Amyloidogenesis in Antibodies.” BMC Bioinformatics 11 (February): 79.
De Baets, Greet, Joost Van Durme, Rob van der Kant, Joost Schymkowitz, and Frederic Rousseau.
2015. “Solubis: Optimize Your Protein.” Bioinformatics 31 (15): 2580–82.
Del Vecchio, Alice, Andreea Deac, Pietro Liò, and Petar Veličković. 2021. “Neural Message
Passing for Joint Paratope‑Epitope Prediction.” arXiv [q‑bio.QM]. arXiv. https://arxiv.org/ abs/2106.00757.
Deszyński, Piotr, Jakub Młokosiewicz, Adam Volanakis, Igor Jaszczyszyn, Natalie Castellana,
Stefano Bonissone, Rajkumar Ganesan, and Konrad Krawczyk. 2021. “INDI—integrated Nanobody Database for Immunoinformatics.” Nucleic Acids Research 50 (D1): D1273–81.
Detlefsen, Nicki Skafte, Søren Hauberg, and Wouter Boomsma. 2022. “Learning Meaningful
Representations of Protein Sequences.” Nature Communications 13 (1): 1914.
Du, Zongyang, Hong Su, Wenkai Wang, Lisha Ye, Hong Wei, Zhenling Peng, Ivan Anishchenko,
David Baker, and Jianyi Yang. 2021. “The trRosetta Server for Fast and Accurate Protein Structure Prediction.” Nature Protocols 16 (12): 5634–51.
Dunbar, James, Konrad Krawczyk, Jinwoo Leem, Terry Baker, Angelika Fuchs, Guy Georges,
Jiye Shi, and Charlotte M. Deane. 2014. “SAbDab: The Structural Antibody Database.” Nucleic Acids Research 42 (Database issue): D1140–46.
Dyson, Michael R., Edward Masters, Deividas Pazeraitis, Rajika L. Perera, Johanna L. Syrjanen,
Sachin Surade, Nels Thorsteinson, etal. 2020. “Beyond Afnity: Selection of Antibody Variants with Optimal Biophysical Properties and Reduced Immunogenicity from Mammalian Display Libraries.” mAbs 12 (1): 1829335.
Eguchi, Raphael R., Christian A. Choe, and Po‑Ssu Huang. 2022. “Ig‑VAE: Generative Modeling
of Protein Structure by Direct 3D Coordinate Generation.” PLoS Computational Biology 18 (6): e1010271.
Elnaggar, Ahmed, Michael Heinzinger, Christian Dallago, Ghalia Rehawi, Yu Wang, Llion Jones,
Tom Gibbs, etal. n.d. “ProtTrans: Towards Cracking the Language of Life’s Code Through Self‑Supervised Learning.” https://ieeexplore.ieee.org/document/9477085/.
5 • Articial Intelligence and Machine Learning 107
Elnaggar, Ahmed, Michael Heinzinger, Christian Dallago, Ghalia Rihawi, Yu Wang, Llion Jones,
Tom Gibbs, etal. 2020. “ProtTrans: Towards Cracking the Language of Life’s Code through Self‑Supervised Deep Learning and High Performance Computing.” arXiv [cs.LG]. arXiv. https://arxiv.org/abs/2007.06225.
Fathallah, Anas M., Manting Chiang, Anshul Mishra, Sandeep Kumar, Li Xue, C. Russell Middaugh,
and Sathy V. Balu‑Iyer. 2015. “The Effect of Small Oligomeric Protein Aggregates on the Immunogenicity of Intravenous and Subcutaneous Administered Antibodies.” Journal of Pharmaceutical Sciences 104 (11): 3691–3702.
Feng, Jiangyan, Min Jiang, James Shih, and Qing Chai. 2022. “Antibody Apparent Solubility
Prediction from Sequence by Transfer Learning.” iScience25 (10): 105173.
Ferdous, Saba, and Andrew C. R. Martin. 2018. “AbDb: Antibody Structure Database‑a Database
of PDB‑Derived Antibody Structures.” Database: The Journal of Biological Databases and Curation 2018 (January). https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5925428/.
Ferrara, Fortunato, M. Frank Erasmus, Sara D’Angelo, Camila Leal‑Lopes, André A. Teixeira, Alok
Choudhary, William Honnen, etal. 2022. “A Pandemic‑Enabled Comparison of Discovery Platforms Demonstrates a Naïve Antibody Library Can Match the Best Immune‑Sourced Antibodies.” Nature Communications 13 (1): 462.
Fischman, Sharon, and Yanay Ofran. 2018. “Computational Design of Antibodies.” Current
Opinion in Structural Biology 51 (August): 156–62.
Friedensohn, Simon, Daniel Neumeier, Tarik A. Khan, Lucia Csepregi, Cristina Parola, Arthur
R. Gorter de Vries, Lena Erlach, Derek M. Mason, and Sai T. Reddy. 2020. “Convergent Selection in Antibody Repertoires Is Revealed by Deep Learning.” bioRxiv. https://www. biorxiv.org/content/10.1101/2020.02.25.965673v1.
Gainza, P., F. Sverrisson, F. Monti, E. Rodolà, D. Boscaini, M. M. Bronstein, and B. E. Correia. 2020.
“Deciphering Interaction Fingerprints from Protein Molecular Surfaces Using Geometric Deep Learning.” Nature Methods. https://www.nature.com/articles/s41592‑019‑0666‑6.
Gao, Sean H., Kexin Huang, Hua Tu, and Adam S. Adler. 2013. “Monoclonal Antibody Humanness
Score and Its Applications.” BMC Biotechnology 13 (July): 55.
Garofalo, Maura, Luca Piccoli, Margherita Romeo, Maria Monica Barzago, Sara Ravasio,
Mathilde Foglierini, Milos Matkovic, etal. 2021. “Machine Learning Analyses of Antibody Somatic Mutations Predict Immunoglobulin Light Chain Toxicity.” Nature Communications 12 (1): 3532.
Geng, Cunliang, Yong Jung, Nicolas Renaud, Vasant Honavar, Alexandre M. J. J. Bonvin, and Li
C. Xue. 2019. “iScore: A Novel Graph Kernel‑Based Function for Scoring Protein–protein Docking Models.” Bioinformatics 36 (1): 112–21.
Ghraichy, Marie, Valentin von Niederhäusern, Aleksandr Kovaltsuk, Jacob D. Galson, Charlotte
M. Deane, and Johannes Trück. 2021. “Different B Cell Subpopulations Show Distinct Patterns in Their IgH Repertoire Metrics.” eLife 10:e73111. doi: 10.7554/eLife.73111. PMID: 34661527; PMCID: PMC8560093.
Glanville, Jacob, Wenwu Zhai, Jan Berka, Dilduz Telman, Gabriella Huerta, Gautam R. Mehta,
Irene Ni, etal. 2009. “Precise Determination of the Diversity of a Combinatorial Antibody Library Gives Insight into the Human Immunoglobulin Repertoire.” Proceedings of the National Academy of Sciences of the United States of America 106 (48): 20216–21.
Greiff, Victor, Gur Yaari, and Lindsay G. Cowell. 2020. “Mining Adaptive Immune Receptor
Repertoires for Biological and Clinical Information Using Machine Learning.” Current Opinion in Systems Biology 24 (December): 109–19.
Guo, Yicheng, Kevin Chen, Peter D. Kwong, Lawrence Shapiro, and Zizhang Sheng. 2019.
“cAb‑Rep: A Database of Curated Antibody Repertoires for Exploring Antibody Diversity and Predicting Antibody Prevalence.” Frontiers in Immunology 10 (October): 2365.
108 Biopharmaceutical Informatics
Harmalkar, Ameya, Roshan Rao, Jonas Honer, Wibke Deisting, Jonas Anlahr, Anja Hoenig, Julia
Czwikla, etal. 2022. “Towards Generalizable Prediction of Antibody Thermostability Using Machine Learning on Sequence and Structure Features.” bioRxiv. https://pubmed.ncbi.nlm. nih.gov/36683173/.
Hashemi, Atieh, Majid Basafa, and Aidin Behravan. 2022. “Machine Learning Modeling for
Solubility Prediction of Recombinant Antibody Fragment in Four Different E. Coli Strains.” Scientic Reports 12 (1): 5463.
Hebditch, Max, M. Alejandro Carballo‑Amador, Spyros Charonis, Robin Curtis, and Jim
Warwicker. 2017. “Protein–Sol: A Web Tool for Predicting Protein Solubility from Sequence.” Bioinformatics 33 (19): 3098–3100.
Honegger, A., and A. Plückthun. 2001. “Yet Another Numbering Scheme for Immunoglobulin
Variable Domains: An Automatic Modeling and Analysis Tool.” Journal of Molecular Biology 309 (3): 657–70.
Kim, Inyoung, Sang Yoon Byun, Sangyeup Kim, Sangyoon Choi, Jinsung Noh, Junho Chung,
Byung Gee Kim. 2021. “Computational analysis of B cell receptor repertoires in COVID‑19 patients using deep embedded representations of protein sequences.” bioRxiv. https://www. biorxiv.org/content/10.1101/2021.08.02.454701v3.full
Jain, Tushar, Tingwan Sun, Stéphanie Durand, Amy Hall, Nga Rewa Houston, Juergen H.
Nett, Beth Sharkey, et al. 2017. “Biophysical Properties of the Clinical‑Stage Antibody Landscape.” Proceedings of the National Academy of Sciences of the United States of America 114 (5): 944–49.
Jankauskaite, Justina, Brian Jiménez‑García, Justas Dapkunas, Juan Fernández‑Recio, and Iain H.
Moal. 2019. “SKEMPI 2.0: An Updated Benchmark of Changes in Protein‑Protein Binding Energy, Kinetics and Thermodynamics upon Mutation.” Bioinformatics 35 (3): 462–69.
Jarasch, Alexander, Hans Koll, Joerg T. Regula, Martin Bader, Apollon Papadimitriou, and Hubert
Kettenberger. 2015. “Developability Assessment during the Selection of Novel Therapeutic Antibodies.” Journal of Pharmaceutical Sciences 104 (6): 1885–98.
Jones, David T., Tanya Singh, Tomasz Kosciolek, and Stuart Tetchner. 2015. “MetaPSICOV:
Combining Coevolution Methods for Accurate Prediction of Contacts and Long Range Hydrogen Bonding in Proteins.” Bioinformatics 31 (7): 999–1006.
Jumper, John, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger,
Kathryn Tunyasuvunakool, etal. 2021. “Highly Accurate Protein Structure Prediction with AlphaFold.” Nature 596 (7873): 583–89.
Källberg, Morten, Haipeng Wang, Sheng Wang, Jian Peng, Zhiyong Wang, Hui Lu, and Jinbo
Xu. 2012. “Template‑Based Protein Structure Modeling Using the RaptorX Web Server.” Nature Protocols 7 (8): 1511–22.
Kaplon, Hélène, Mrinalini Muralidharan, Zita Schneider, and Janice M. Reichert. 2020.
“Antibodies to Watch in 2020.” mAbs 12 (1): 1703531.
Kaplon, Hélène, and Janice M. Reichert. 2021. “Antibodies to Watch in 2021.” mAbs 13 (1):
1860476.
Kelly‑Scumpia, Kindra M., Philip O. Scumpia, Jason S. Weinstein, Matthew J. Delano, Alex G.
Cuenca, Dina C. Nacionales, James L. Wynn, etal. 2011. “B Cells Enhance Early Innate Immune Responses during Bacterial Sepsis.” The Journal of Experimental Medicine 208 (8): 1673–82.
Kelow, Simon, Bulat Faezov, Qifang Xu, Mitchell Parker, Jared Adolf‑Bryfogle, and Roland L.
Dunbrack. 2022. “A Penultimate Classication of Canonical Antibody CDR Conformations.” bioRxiv. https://doi.org/10.1101/2022.10.12.511988.
Khetan, Rahul, Robin Curtis, Charlotte M. Deane, Johannes Thorling Hadsund, Uddipan Kar,
Konrad Krawczyk, Daisuke Kuroda, etal. 2022. “Current Advances in Biopharmaceutical Informatics: Guidelines, Impact and Challenges in the Computational Developability Assessment of Antibody Therapeutics.” mAbs 14 (1): 2020082.
5 • Articial Intelligence and Machine Learning 109
Kilambi, Krishna Praneeth, and Jeffrey J. Gray. 2017. “Structure‑Based Cross‑Docking Analysis
of Antibody–Antigen Interactions.” Scientic Reports 7 (1): 1–15.
Kim, Jin Hong, and Hyo Jeong Hong. 2012. “Humanization by CDR Grafting and
Specicity‑Determining Residue Grafting.” Methods in Molecular Biology 907: 237–45. Kindt, T. J., and R. A. Goldsby. 2007. “Osborne BA Kuby Immunology.” WH Freeman and Co., NY. Koenig, Patrick, Chingwei V. Lee, Benjamin T. Walters, Vasantharajan Janakiraman, Jeremy
Stinson, Thomas W. Patapoff, and Germaine Fuh. 2017. “Mutational Landscape of
Antibody Variable Domains Reveals a Switch Modulating the Interdomain Conformational
Dynamics and Antigen Binding.” Proceedings of the National Academy of Sciences of the
United States of America 114 (4): E486–95. Kovaltsuk, Aleksandr, Konrad Krawczyk, Jacob D. Galson, Dominic F. Kelly, Charlotte M. Deane,
and Johannes Trück. 2017. “How B‑Cell Receptor Repertoire Sequencing Can Be Enriched
with Structural Antibody Data.” Frontiers in Immunology 8 (December): 1753. Kovaltsuk, Aleksandr, Jinwoo Leem, Sebastian Kelm, James Snowden, Charlotte M. Deane,
and Konrad Krawczyk. 2018. “Observed Antibody Space: A Resource for Data Mining
Next‑Generation Sequencing of Antibody Repertoires.” Journal of Immunology 201 (8):
2502–9. Krawczyk, Konrad, Terry Baker, Jiye Shi, and Charlotte M. Deane. 2013. “Antibody I‑Patch
Prediction of the Antibody Binding Site Improves Rigid Local Antibody–antigen Docking.”
Protein Engineering, Design & Selection: PEDS 26 (10): 621–29. Krawczyk, Konrad, Andrew Buchanan, and Paolo Marcatili. 2021. “Data Mining Patented
Antibody Sequences.” mAbs 13 (1): 1892366. Krawczyk, Konrad, Sebastian Kelm, Aleksandr Kovaltsuk, Jacob D. Galson, Dominic Kelly,
Johannes Trück, Cristian Regep, etal. 2018. “Structurally Mapping Antibody Repertoires.”
Frontiers in Immunology. https://doi.org/10.3389/mmu.2018.01698. Krawczyk, Konrad, Xiaofeng Liu, Terry Baker, Jiye Shi, and Charlotte M. Deane. 2014.
“Improving B‑Cell Epitope Prediction and Its Application to Global Antibody‑Antigen
Docking.” Bioinformatics 30 (16): 2288–94. Krawczyk, Konrad, Matthew I. J. Raybould, Aleksandr Kovaltsuk, and Charlotte M. Deane. 2019.
“Looking for Therapeutic Antibodies in next‑Generation Sequencing Repositories.” mAbs
11 (7): 1197–1205. Kryshtafovych, Andriy, Torsten Schwede, Maya Topf, Krzysztof Fidelis, and John Moult. 2021.
“Critical Assessment of Methods of Protein Structure Prediction (CASP)‑Round XIV.”
Proteins 89 (12): 1607–17. Kumar, M. D. Shaji, K. Abdulla Bava, M. Michael Gromiha, Ponraj Prabakaran, Koji Kitajima,
Hatsuho Uedaira, and Akinori Sarai. 2006. “ProTherm and ProNIT: Thermodynamic
Databases for Proteins and Protein‑Nucleic Acid Interactions.” Nucleic Acids Research 34
(Database issue): D204–6. Kumar, Sandeep, Satish K. Singh, Xiaoling Wang, Bonita Rup, and Davinder Gill. 2011. “Coupling
of Aggregation and Immunogenicity in Biotherapeutics: T‑ and B‑Cell Immune Epitopes
May Contain Aggregation‑Prone Regions.” Pharmaceutical Research 28 (5): 949–61. Kunik, Vered, Shaul Ashkenazi, and Yanay Ofran. 2012. “Paratome: An Online Tool for Systematic
Identication of Antigen‑Binding Regions in Antibodies Based on Sequence or Structure.”
Nucleic Acids Research 40 (Web Server issue): W521–24. Kuriata, Aleksander, Valentin Iglesias, Jordi Pujols, Mateusz Kurcinski, Sebastian Kmiecik, and
Salvador Ventura. 2019. “Aggrescan3D (A3D) 2.0: Prediction and Engineering of Protein
Solubility.” Nucleic Acids Research. https://doi.org/10.1093/nar/gkz321. Kuroda, Daisuke, and Kouhei Tsumoto. 2020. “Engineering Stability, Viscosity, and
Immunogenicity of Antibodies by Computational Design.” Journal of Pharmaceutical
Sciences 109 (5): 1631–51.
110 Biopharmaceutical Informatics
Lai, Pin‑Kuang, Amendra Fernando, Theresa K. Cloutier, Jonathan S. Kingsbury, Yatin Gokarn,
Kevin T. Halloran, Cesar Calero‑Rubio, and Bernhardt L. Trout. 2021. “Machine Learning
Feature Selection for Predicting High Concentration Therapeutic Antibody Aggregation.”
Journal of Pharmaceutical Sciences 110 (4): 1583–91. Lai, Pin‑Kuang, Austin Gallegos, Neil Mody, Hasige A. Sathish, and Bernhardt L. Trout.
2022. “Machine Learning Prediction of Antibody Aggregation and Viscosity for High
Concentration Formulation Development of Protein Therapeutics.” mAbs 14 (1): 2026208. Lauer, Timothy M., Neeraj J. Agrawal, Naresh Chennamsetty, Kamal Egodage, Bernhard Helk, and
Bernhardt L. Trout. 2012. “Developability Index: A Rapid in Silico Tool for the Screening
of Antibody Aggregation Propensity.” Journal of Pharmaceutical Sciences 101 (1): 102–15. Laustsen, Andreas H., Markus‑Frederik Bohn, and Anne Ljungars. 2022. “The Challenges
with Developing Therapeutic Monoclonal Antibodies for Pandemic Application.” Expert
Opinion on Drug Discovery 17 (1): 5–8. Laustsen, Andreas H., Victor Greiff, Aneesh Karatt‑Vellatt, Serge Muyldermans, and Timothy
P. Jenkins. 2021. “Animal Immunization, in Vitro Display Technologies, and Machine
Learning for Antibody Discovery.” Trends in Biotechnology 39 (12): 1263–73. Lazar, Greg A., John R. Desjarlais, Jonathan Jacinto, Sher Karki, and Philip W. Hammond.2007.
“A Molecular Immunology Approach to Antibody Humanization and Functional
Optimization.” Molecular Immunology 44 (8): 1986–98. Lee, Jae Hyeon, Payman Yadollahpour, Andrew Watkins, Nathan C. Frey, Andrew Leaver‑Fay,
Stephen Ra, Kyunghyun Cho, Vladimir Gligorijevic, Aviv Regev, and Richard Bonneau.
2022. “EquiFold: Protein Structure Prediction with a Novel Coarse‑Grained Structure
Representation.” bioRxiv. https://doi.org/10.1101/2022.10.07.511322. Leem, Jinwoo, James Dunbar, Guy Georges, Jiye Shi, and Charlotte M. Deane. 2016.
“ABodyBuilder: Automated Antibody Structure Prediction with Data–driven Accuracy
Estimation.” mAbs 8 (7): 1259–68. Leem, Jinwoo, Laura S. Mitchell, James H. R. Farmery, Justin Barton, and Jacob D. Galson.
2022. “Deciphering the Language of Antibodies Using Self‑Supervised Learning.” Patterns
of Prejudice, May, 18;3(7):100513. doi: 10.1016/j.patter.2022.100513. PMID: 35845836;
PMCID: PMC9278498. Lees, William, Christian E. Busse, Martin Corcoran, Mats Ohlin, Cathrine Scheepers, Frederick
A. Matsen, Gur Yaari, etal. 2020. “OGRDB: A Reference Database of Inferred Immune
Receptor Genes.” Nucleic Acids Research 48 (D1): D964–70. Lefranc, M. P., V. Giudicelli, C. Ginestoux, J. Bodmer, W. Müller, R. Bontrop, M. Lemaitre,
A. Malik, V. Barbié, and D. Chaume. 1999. “IMGT, the International ImMunoGeneTics
Database.” Nucleic Acids Research 27 (1): 209–12. Li, Tong, Robert J. Pantazes, and Costas D. Maranas. 2014. “OptMAVEn‑‑a New Framework
for the de Novo Design of Antibody Variable Region Models Targeting Specic Antigen
Epitopes.” PloS One 9 (8): e105954. Liaw, Chyn, Chun‑Wei Tung, and Shinn‑Ying Ho. 2013. “Prediction and Analysis of Antibody
Amyloidogenesis from Sequences.” PloS One 8 (1): e53235. Liberis, Edgar, Petar Velickovic, Pietro Sormanni, Michele Vendruscolo, and Pietro Liò. 2018.
“Parapred: Antibody Paratope Prediction Using Convolutional and Recurrent Neural
Networks.” Bioinformatics 34 (17): 2944–50. Lim, Yoong Wearn, Adam S. Adler, and David S. Johnson. 2022. “Predicting Antibody Binders
Lima, Wanessa C., Elisabeth Gasteiger, Paolo Marcatili, Paula Duek, Amos Bairoch, and Pierre
Cosson. 2020. “The ABCD Database: A Repository for Chemically Dened Antibodies.”
Nucleic Acids Research 48 (D1): D261–64. Lippow, Shaun M., K. Dane Wittrup, and Bruce Tidor. 2007. “Computational Design of
Antibody‑Afnity Improvement beyond in Vivo Maturation.” Nature Biotechnology 25
(10): 1171–76.