Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5629_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Contents
- •Foreword
- •Preface
- •About the Editors
- •Contributors
- •References
- •2.3.4 Barriers to Automation Adoption
- •2.4 Core Ingredients for Successful Digital Transformation
- •2.1 Introduction
- •2.3.1 Operational Challenges
- •2.3.2 Cultural Challenges
- •2.4.2 Cloud Computing
- •2.5 Case Studies of Successful Digital Transformation
- •2.6 Conclusion
- •References
- •3. Computational Protein Design Strategies for Optimization of Antigen Generation to Drive Antibody Discovery
- •3.1 Introduction
- •3.3 Antigen Generation Strategies
- •3.4 Computational Methods
- •3.4.2 Computational Protein Structure Prediction
- •References
- •4. Bioinformatic Analyses of Antibody Repertoires and Their Roles in Modern Antibody Drug Discovery
- •4.1 Introduction
- •4.6 Summary and Future Directions
- •Acknowledgments
- •References
- •5.1 Introduction
- •5.2 Databases
- •5.2.1 Databases in Machine Learning Approaches
- •5.2.2 Database Types
- •5.3 Applications of Machine Learning in Antibody Discovery and Development
- •5.3.1 Structure Prediction with Deep Learning
- •5.3.3 Developability
- •5.4 Antibody Generation and Design by Language Models
- •5.4.1 Antibody Representations
- •5.4.2 Representation Learning
- •5.4.3 Language Models
- •References
- •6.1 Introduction
- •6.2 Antibody Generation through Deep Generative Models
- •6.3.1 Sampling and Scoring
- •6.5 Conclusions and Perspectives
- •Acknowledgments
- •References
- •7.1 Introduction
- •7.2.3 Computational Approaches to Predict Antibody–Antigen Interaction
- •7.3 Conclusion
- •Competing Interests
- •Acknowledgments
- •References
- •8.2 Common Types of Molecular Simulations for Biomolecules
- •8.2.1 Molecular Dynamics (MD) Simulations
- •8.2.2 Monte Carlo (MC) Simulations
- •8.2.3 Challenges of Molecular Simulations
- •8.3.1 Periodic Boundary Conditions
- •8.4 Uses of Molecular Simulation in Antibody Drug Development
- •8.5 Conclusion
- •References
- •9. Considerations of Developability During the Early Stages of Antibody Drug Discovery and Design
- •9.1 Introduction
- •9.2 Historical Perspective
- •9.3 Clinical Antibody Data Set
- •9.5 Control Antibodies
- •9.7 Assessment of Chemical Liabilities
- •9.8 Conclusions and Future Perspectives
- •Acknowledgments
- •References
- •Abbreviations
- •10.1 Introduction
- •10.4.1 Conclusions and Outlook
- •Acknowledgments
- •References
- •11.8 Conclusions and Future Directions
- •References
- •12.1 Introduction to PK/PD and QSP Modeling
- •12.1.1 PK/PD Modeling
- •12.1.2 QSP Modeling
- •12.2.1 Monoclonal Antibodies (mAbs)
- •12.2.3 Cell Therapies
- •12.2.4 Gene Therapies
- •12.2.5 Vaccines
- •12.2.6 mRNA/siRNA/Oligonucleotide Therapeutics
- •12.4 Case Studies
- •12.5 Conclusions and Future Perspectives
- •References
- •13.1 Introduction
- •13.2 AI/ML: A Game Changer for Antibody Design
- •13.3 Multispecific Antibody Design
- •13.4 Adapting AI to the Design of Multispecific Antibodies
- •13.4.1 Structure Prediction and Modeling
- •13.4.2 Developability Prediction and Optimization
- •13.4.4 In Silico Modeling and Simulation
- •13.5 The Future: Beyond Optimization
- •13.5.1 Market Trends and Commercialization
- •13.5.2 Logic Gates, Biosensors, and De Novo Design
- •13.5.3 Challenges and Opportunities
- •13.6 Conclusion
- •Acknowledgments
- •References
- •Index

5 • Articial Intelligence and Machine Learning 91
5.3.3.1 Immunogenicity
Mice are often used as a source of therapeutic antibodies upon being challenged with a
target antigen. The signicant drawback of this approach is that animal antibodies can
elicit unwanted immune responses in humans. To reduce such risks, researchers need
to engineer antibodies to resemble human antibody molecules without losing activity
as part of the humanization process (Kim and Hong 2012). Computational methods
quantifying the nativeness of human sequences have long played a signicant role in
this process (Table5.3).
Initially, computational humanization was tackled by frequency‑based methods that
quantied the similarity between animal and human sequences–for example, T20 (Gao
etal. 2013) or humanness scores (Abhinandan and Martin 2007). These approaches
were based on a small number of sequences (in thousands), offering a limited ability
to determine the correlation between different residues. The availability of NGS data
increased the antibody sequence samples from thousands to millions, leading to the
creation of improved positional frequencies (Schmitz et al. 2020; Sheng et al. 2017;
Kovaltsuk etal. 2018).
Positional proles may have been enriched with NGS data but do not reect posi‑
tional interdependencies. To quantify positional correlations, researchers developed a
multivariate Gaussian (MG) statistical score based on the OAS data (Clavero‑Álvarez
etal. 2018), later expanded to the LSTM model by Wollacott (Wollacott etal. 2019).
Both models concentrated on predicting what constitutes the human sequence and
brought to light the required element of correlations between positions. The authors
of MG compared their score to therapeutic sequence immunogenicity. It resulted in a
weak correlation (r2 = 0.18), suggesting that sequence identities might not encode the
immunogenicity information alone. This nding is consistent with the industry experi‑
ence around immunogenicity origins. Immunogenicity toward biotherapeutic drugs can
be observed in clinical trials through the generation of anti‑drug antibodies (ADAs) by
patients who receive immunotherapy. Immunogenicity emerges in patients due to mul‑
tiple factors related to drug product quality (formulation, presence of aggregates in the
product, or aggregation of the product in vivo during administration), patient’s specic
disease history, their genetic background, and the humanness of the antibody sequence
(Kumar etal. 2011; Fathallah etal. 2015; Singh 2011).
A much larger scale attempt at employing NGS data for deimmunization was pro‑
posed by Hu‑mab (Marks etal. 2021). The method is based on a random forest model
trained to distinguish human and non‑human sequences of a particular V gene type
from ones from other species. Hu‑mab correctly differentiated human and other animal
sequences in both validation and test sets, noting slightly worse performance on the
light chain, which may have been caused by the greater volume of negative training data
available for variable heavy chain (VH) than variable light chain (VL) models. Another
reason could be the smaller variability of light chains regarding isotypes and CDRs. The
previous LSTM model (Wollacott etal. 2019) could not discriminate between human
and other animal sequences, possibly because LSTM models were only trained on
sequences originating in a single species (human).

92 Biopharmaceutical Informatics
TABLE5.3 Humanization methods
HUMANIZATION
METHOD METHOD DESCRIPTION REFERENCES
T20 Based on thousands of sequences, offers
a limited ability to determine the
correlation between a selected set of
residues
Humanness scores Based on thousands of sequences, one of
the pioneering methods to quantify
nativeness of antibody sequences
Improved positional
frequencies
Multivariate
Gaussian (MG)
statistical score
Multivariate
Gaussian (MG)
statistical score with
the LSTM model
Hu-mab Based on a random forest model trained
BioPhi (Sapiens and
OASis)
AbBERT Transformer-based language model
Llamanade For capturing unique features of
IgReconstruct Based on single nucleotide frequencies, a
AbDiver Positional frequencies from OAS (Młokosiewicz etal.
A short description and reference are presented above for each humanization method.
Based on millions of sequences, quantify
nativeness of antibody sequences by
determining the correlation between
different residues
Based on the OAS data, statistical score
to quantify the nativeness of human
sequence
Based on the OAS data, predicts what
constitutes the human sequence with
usage of LSTM model
to distinguish human and non-human
sequences of a particular V gene type
from ones from other species
Uses language models for capturing the
diversity of natural human antibody
repertoires and humanizes a sequence
trained on up to 20million unpaired
heavy/light chain sequences from the
OAS database
nanobodies, uses large-scale analysis of
nanobodies and IgG
germline gene rearrangement tailored
to the nucleotide frequency observations
made in the repertoire is generated to
estimate the similarity of a target Ab
amino acid sequence to a given
repertoire
(Gao etal. 2013)
(Abhinandan and
Martin 2007)
(Schmitz etal. 2020;
Sheng etal. 2017;
Kovaltsuk etal. 2018)
(Clavero-Álvarez etal.
2018)
(Wollacott etal. 2019)
(Marks etal. 2021)
(Prihoda etal. 2022)
(Vashchenko etal.
2022)
(Sang etal. 2022)
(Schmitz etal. 2020)
2022)

5 • Articial Intelligence and Machine Learning 93
Another computational method that aims to accelerate the process of humaniza‑
tion is BioPhi, an antibody design interface that uses automated methods for captur‑
ing the diversity of natural human antibody repertoires (Prihoda et al. 2022). BioPhi
offers two data‑driven approaches–humanization (Sapiens) and humanness evaluation
methods (OASis). Sapiens is a deep learning humanization approach based on masked
language modeling (MLM). The method is trained with human variable region antibody
sequences from the OAS. Its goal is to recognize and repair masked or mutated positions
in unaligned amino acid sequences. OASis, on the other hand, is a humanness metric
based on peptide search in the OAS, which evaluates the humanness of a sequence by
dividing it into overlapping 9‑mer peptides (inspired by human string content (Lazar
etal. 2007)). Next, the metric compares them against the OAS database to predict their
universality across the human population. Based on the in silico humanization bench‑
mark of 177 antibodies, this solution offers mutational choices that are similar to those
achieved by experimental humanization methods. The primary advantage of BioPhi is
its attention layer and granularity, allowing users to examine the residue dependencies
and mutational effects on the score.
5.3.3.2 Prediction of biophysical properties: Aggregation,
viscosity, solubility, and thermostability
An mAb‑based new drug development relies on the developability properties of these
molecules, such as low aggregation propensity and low viscosity (Xu etal. 2019; Jarasch
etal. 2015; Makowski etal. 2021). Given the limited number of molecules with the avail‑
able sequence and biophysical property data, assessing the stability proles of antibod‑
ies at high concentrations is challenging during early stage discovery. Predictive tools
based on ML may help researchers to evaluate the developability of high‑concentration
antibody formulation early in the discovery process. Researchers have already applied
computational tools to identify drug‑like antibodies characterized by favorable stability
(Makowski etal. 2021).
Viscosity prediction received attention as well. For example, Sharma and colleagues
(Sharma etal. 2014) found viscosity to be highly correlated with Fv net charge and
charge symmetry. At the same time, viscosity was determined to be weakly correlated
with hydrophobicity. The team proposed a linear equation based on these parameters to
calculate the viscosity at 180mg/ml (pH 5.5 and 200 mM arginine‑HCl). Another vis‑
cosity predictive tool is the spatial charge map, calculated using a molecular dynamics
(MD) simulation that explains the exposed surface‑negative charge distribution on the
Fv region (Agrawal etal. 2016). Tomer and colleagues (Tomar etal. 2017) developed
an equation for predicting the concentration‑dependent viscosity curves, which used
charges on the heavy and light chain variable regions, the hinge region, as well as the
hydrophobic surface area of a full‑length antibody. A summary of a comparison of the
viscosity prediction tools described above can be found in Kuroda etal. (Kuroda and
Tsumoto 2020). Lai and colleagues (Lai etal. 2022) recently applied ML to this problem
area. The team developed a model based on 27mAbs to predict antibody viscosity at

94 Biopharmaceutical Informatics
150mg/ml, based on the decision tree classication method including two features: net
charge and high viscosity index (HVI). To predict viscosity at different concentrations,
the team developed a coarse‑grained model combined with hydrodynamic calculations
and HVI‑derived parameters.
Regarding aggregation, researchers can leverage several in silico models for pre‑
dicting the solubility/protein aggregation rates: developability index (Lauer etal. 2012),
Solubis (De Baets etal. 2015), and Camsol (Sormanni etal. 2015). Other models help
to identify aggregation‑prone regions – for example, spatial aggregation propensity
(Chennamsetty etal. 2009), Aggrescan 3D (Kuriata etal. 2019), and ANuPP (Prabakaran
etal. 2021). ML has also been applied to predict protein aggregation kinetics (regres‑
sion) based on sequence features (Tartaglia etal. 2005; Rawat etal. 2021b) and antibody
amyloidogenesis (classication) (Liaw etal. 2013; David etal. 2010; Rawat etal. 2021a;
Garofalo etal. 2021). The latter has limited application in therapeutic development but is
greatly relevant for human diseases (Roberts 2014). Another ML‑based model proposed
by Lai and colleagues (Lai etal. 2021) was trained on 21mAbs to predict therapeu‑
tic antibody aggregation rates at 150mg/ml with the help of structural‑based features
extracted from MD simulations.
Computational approaches and solubility prediction tools open the door to fast
high‑throughput screening without any material requirements (Feng etal. 2022). Current
approaches use molecular descriptors extracted from protein sequence–sequence‑based
predictors (Hebditch et al. 2017; Sormanni etal. 2017) –or from structures– struc‑
ture‑based predictors (Sormanni etal. 2015). The former often fail to consider tertiary
structure information, which differentiates poorly soluble residues that drive protein
folding from the residues exposed to the solvent, potentially eliciting aggregation.
Structure‑based tools, on the other hand, can be applied only when researchers have a
reliable 3D model, limiting the throughput and applications to a massive volume of early
stage mAb candidates. Some computational methods only output a binary classica‑
tion (soluble/insoluble) instead of a numerical value (Hebditch etal. 2017; Smialowski
etal. 2012). Feng and colleagues (Feng et al. 2022) presented a strategy called sol‑
Predict that uses the embeddings from pre‑trained protein language model to predict
the apparent solubility of mAbs in histidine (pH 6.0) buffer. The team used a dataset
of 220 diverse, in‑house mAbs for model training and hyperparameter tuning through
vefold cross‑validation. The solPredict was found to achieve a high correlation with
experimental solubility through an independent test set of 40mAbs, delivering a good
performance for both IgG1 and IgG4 subclasses. Hashemi and colleagues (Hashemi
etal. 2022) modeled a soluble production of the anti‑EpCAM extracellular domain sin‑
gle‑chain variable fragment (scFv). The solution was then optimized as a function of
four numerical factors: post‑induction temperature, post‑induction time, cell density of
induction time, and inducer concentration, as well as one categorical variable using arti‑
cial neural network (ANN) and response surface methodology (RSM). The predicted
value obtained by ANN for the response (106.1mg/L) approached the experimental
result closer than the one obtained by RSM (97.9mg/L).
The application of computational tools to predict stability‑enhancing mutations
has increased in popularity over the past decade. Researchers use the state‑of‑the‑art
approach called protein consensus design, which leverages phylogenetic information from

5 • Articial Intelligence and Machine Learning 95
multiple sequence alignments (MSAs) to output the most frequent 34 amino acids for a
residue position. However, such residues are characterized by improved thermostability
in only 50% of the cases, with poorer performance noted for antibody sequences that
have highly conserved framework regions. ML approaches employed large datasets such
as ProTherm (Kumar etal. 2006), which predicts thermostability by collating mutant 40
effects on protein stability from mesophilic and thermophilic sequences. However, the
publicly available thermostability data exclude mAb and scFv sequences, placing a limit
on their generalization. It is still a challenge for researchers to predict thermally enhanc‑
ing scFv sequences or develop thermostable mutations within approved scFv candidates.
New computational approaches will need to be designed to combine sequence informa‑
tion to predict biophysical attributes (thermostability) and highly accurate, generalizabil‑
ity‑enabling design. Harmalkar and colleagues (Harmalkar etal. 2022) addressed this
by equipping unsupervised and supervised learning approaches over a thermostability
prediction task that has been tuned specically for generated scFv sequences.
Altogether, the prediction of biophysical properties of therapeutic antibodies is
challenging since most of the studies cited perform the predictions within a specic
experimental framework. For instance, “aggregation” can be measured by a variety of
assays (e.g., hydrophobic interaction chromatography (HIC)) that, though correlating
with one another, do not need to give the same reading for the same molecule. This
makes generalizability extremely difcult and rendering benchmarking different meth‑
ods developed with different datasets almost impossible.
5.3.3.3 Deep learning applications
Current efforts in developability prediction using deep learning methods are directed
toward generating novel sequences computationally and selecting for good manufactur‑
ability properties. This can be done by generating the molecules and ltering them using
general‑purpose developability predictors such as therapeutic antibody proler (TAP)
(Mason etal. 2021). Alternatively, a GAN trained on large‑scale natural data can be
ne‑tuned on sequences with specic developability properties (Amimeur etal. 2020).
All in all, there is an interplay between how to efciently sample the antibody sequence/
structure space and how to predict the developability properties well.
Deep learning models aim to learn higher order dependencies missed by posi‑
tion frequency analysis, reducing the chance of generating non‑functional proteins.
However, current limitations of deep learning approaches focus only on CDR regions or
heavy chains and a lack of experimental validation of predicted properties (Khetan etal.
2022). For instance, 74% antigen binding was achieved in a mouse library designed by a
variational autoencoder that generated novel CDR‑H regions (Friedensohn etal. 2020).
However, such an approach ignores non‑CDR region contributions to the paratope, and
the diversity of sequences in this library is unknown.
Other generative approaches, such as GANs, can also be trained on natural human
antibody repertoires and biased via transfer learning to generate sequences predicted
to have the desired biophysical properties (Amimeur etal. 2020). However, researchers
need more insights to better understand how such properties impact the overall developa‑
bility of an antibody. Experimental validation of the predicted properties is also required.

96 Biopharmaceutical Informatics
It has already been conducted for an enzymatically active protein library (Repecka etal.,
2021) and a nanobody library created using a generative deep neural network‑powered
autoregressive model trained on a native llama repertoire (Shin etal. 2021).
Previous studies have demonstrated the use of libraries of mammalian antibod‑
ies for selecting antibodies with optimal biophysical properties, reduced polyreactivity,
and immunogenicity (Dyson etal. 2020). Dyson and colleagues described the use of
a nuclease‑directed integration system to generate antibodies with differing biophysi‑
cal properties only based on their display level on mammalian cells. Other researchers
have demonstrated the use of ML‑guided directed evolution on combinatorial sequences
(Wu etal. 2019). Recently, Chen and colleagues formulated ML pipelines to predict the
developability of a library of 2,400 antibodies using sequence alone (Chen etal. 2020).
Such advances in bioinformatics and in silico methods allowed for the efcient develop‑
ment of commercially viable antibodies. Antibody library variants of an original candi‑
date are created to exhibit better developability than the parent molecule.
Generative ML has been proposed as a means to design antigen‑specic mAbs,
but efforts to conrm this hypothesis have been hampered by the cost and time
required to test large numbers of antibody sequences. To address this challenge, Akbar
and colleagues (Akbar et al. 2022) developed a lattice‑based antibody‑antigen bind‑
ing simulation framework incorporating physiological parameters. They found that a
deep generative model trained only on antibody sequence data can be used to design
conformational epitope‑specic antibodies matching or exceeding the training dataset
in afnity and diversity, showing that lower sequence diversity is necessary for high‑
accuracy generative antibody modeling.
5.4 ANTIBODY GENERATION AND DESIGN BY LANGUAGE MODELS
Designing novel antibodies without resorting to experimental methods has long been
the end goal of computational antibody engineering. The earliest attempts indicated
the possibility of performing exhaustive in silico mutagenesis combined with binding
energy estimation on a cocrystal structure (Lippow etal. 2007). Designing the antibod‑
ies without the cocrystal structure has been performed by variations on sampling the
available structures (Li etal. 2014; Adolf‑Bryfogle etal. 2018). One of the primary chal‑
lenges in this area relates to creating an antibody sequence/structure that is plausible in a
biological sense. Language models, such as Bidirectional encoder representations from
transformers (BERT), have been applied to push antibody manifolds into representation
space, sampling which facilitates novel antibody discovery.
Here, we describe the importance of antibody representation in terms of data inter‑
pretability and downstream model performance. We show how such representations can
be enriched with domain knowledge by both manual feature engineering processes and
representation learning techniques. Finally, we present how language models are used
to produce antibody embedding and improve the antibody discovery process.

5 • Articial Intelligence and Machine Learning 97
5.4.1 Antibody Representations
Data representation is a vital aspect of all ML projects as raw unprocessed data often
cannot be used directly by the model. It is important that feature engineering becomes
an intrinsic step of the ML development pipeline, as good representation signicantly
boosts downstream model performance and stability or reduces the input dimensional‑
ity. In this process, input data are modied and augmented with additional information
and domain knowledge relevant to the problem being solved.
Several approaches to representing proteins can be used in parallel, each capturing
a different perspective. For example, one can represent an antibody as a sequence of
amino acids. Then, with the presence of structural data, it is possible to represent protein
as a structure–for instance, by extracting voxelated atomic maps. Despite being very
easy to interpret, such an approach has a major drawback: it is sensitive to rotations.
That is why graphs are another popular option for representing protein structural data.
The function and the structure of a protein derive from the interaction between all
of its residues. Such a residue interaction network can be abstracted to a graph where
amino acids constitute nodes and relationships between them are the edges. One can
use several denitions to construct such a graph: for example, using Cα atoms as nodes
with connecting edge if Cα atoms are within 7 Å (Chakrabarty and Parekh 2016; Pittala
and Bailey‑Kellogg 2020), analogously for Cβ atoms with distance thresholds between
7 Å and 10 Å (Chakrabarty and Parekh 2016; Pittala and Bailey‑Kellogg 2020), or using
residue center of mass as a node with edge if any atoms are closer than 5 Å. Another
approach introduces edges if the non‑covalent interaction strength between interacting
residues is above some predened threshold (Brinda and Vishveshwara 2005). There
also exist methods using different graph levels (atom‑atom interactions followed by
residue‑residue) at the same time (Tubiana etal. 2022).
Another challenge is posed by representing molecular surfaces of a protein for ML.
Gainza and colleagues (Gainza etal. 2020) applied geometric deep learning to learn
surface ngerprints that are important for biological interactions and tested their model
on three distinct problems: classication of ligand binding pockets, protein binding site
prediction, and protein‑protein complex search based on interaction ngerprints.
In the context of enriching representations, one can augment biological data on
residue, sequence, and structure layers in the feature engineering process. The easiest
residue representation is likely one‑hot encoding where each residue is represented as
a 20‑dimensional unit vector with 1 in place of represented amino acid. This does not
add any domain knowledge to the residue representation but enables the model to treat
all residues equally, as all amino acid representations differ only by rotation, and all
pairwise distances are the same.
Evolutionary relationships can be encapsulated into data. In this case, one can rep‑
resent each residue as a 20‑ (or 21‑) dimensional vector, where each value is taken from
the BLOSUM substitution matrix (Lim etal., 2022). A common practice is to include
the physicochemical properties of residues by adding Vectors of Hydrophobic, Steric,
and Electronic properties to the representation (Liberis etal. 2018).
One can embed residues’ positional information on the sequence level into sequence
representation. To compare protein sequences in a meaningful way, one needs to align them

98 Biopharmaceutical Informatics
so that the proteins from the same family will have corresponding evolutionary‑dependent
residues in the same positions. Upon sequence alignment, inputs have a xed size, which
licenses the use of simple models that can discover patterns and residue relationships from
absolute input positions, which may be important for model interpretability.
Detlefsen and colleagues (Detlefsen et al. 2022) showed that alignment of the
input sequences improved the model’s ability to separate sequences from one protein
family (b‑lactamase) into different phyla. It remains unclear if it is possible to general‑
ize this capability across multiple protein families. For antibodies, sequence alignment
is achieved by numbering methods, which also provide the functional position context
(e.g., CDR, framework) for each residue. At the structural level, one can augment input
data with shape annotations for each residue. For example, ASA or paratope labels.
Most ML models natively employ vector data format of constant dimensionality.
To utilize those models, one needs to transform variable‑in‑length biological sequences
to point in xed‑size dimensional space so that similar input sequences, or functionally
similar proteins, reside close to each other in the vector space.
Such transformation is called embedding, and to dene it, one needs a similarity
measure and the transformation function that transforms input to a point in the embed‑
ded space. Proteins are sequences of amino acids with variable lengths that fold into a
three‑dimensional shape. Therefore, as an example of a similarity measure, one can use
Levenshtein distance between sequences or some shape difference measures.
5.4.2 Representation Learning
As opposed to the manual process of feature engineering, which involves manually
adding domain knowledge to the representation, it is possible to apply representation
learning techniques to render meaningful interpretable data representations. The basic
representation learning approach is PCA, where the representation features are the out‑
put of linear transformations of input data. Representation learning is often used in
combination with transfer learning, especially when there is little data available. That is
because the model can learn to recognize meaningful biological features by making use
of vastly available NGS data.
The biological knowledge embedded into the vector and the meaningfulness of
such learned representations can be then assessed by the performance of downstream
model predictions.
Detlefsen and colleagues (Detlefsen etal. 2022) demonstrated that low‑dimensional
sequence representation space topology shows good correspondence between the ances‑
tral lineage of the proteins that were rendered on latent space. The team also revealed
that the learned representation space is not Euclidean. Hence, arithmetic operations on
vector representations are not a good approach to interpolate points in the latent space
as one might get unpredictable results. They equipped learned representations with a
Riemannian metric, so measuring the distance in the data space was possible. This
demonstrated that the usage of geodesic distances enables more meaningful navigation
over the latent space.

5 • Articial Intelligence and Machine Learning 99
5.4.3 Language Models
Antibody diversity is estimated to be 1018 unique molecules (Briney et al. 2019).
Therefore, with plenty of publicly available NGS datasets, efcient sequence embedding
that captures its biological properties holds the promise of the possibility to navigate
through the whole immunoglobulin sequence space, providing new insights into the
biology of antibodies. Such navigation could be potentially used in the protein engi‑
neering process to interpolate between known sequences in latent space to yield new
proteins functionally similar or to guide mutations to design antibodies with desired
characteristics.
An emerging trend in antibody analysis is the use of natural language processing
(NLP) methods to develop antibody embeddings by making analogies between amino
acid k‑mers, whole antibody sequences, antibody repertoires, and words, sentences, and
documents. In NLP, the most widely used family of models to produce word embed‑
dings is word2vec (Mikolov etal. 2013). Two architectures are used for representation
learning:
• Skip‑Gram, where given an input word, the neural network tries to predict its
context (preceding and following words).
• Continuous Bag of Words architecture, where given context (neighbor words),
the model is tasked with predicting the target word.
Because of the training process, output vectors embed contextual information and are
capable of capturing both syntactic and semantic relationships, so distinct words that
are related by the context are adjacent to each other in the vector space. Such con‑
text‑aware embedding can be called distributed representation. Below, we highlight
some of the language models that are used to analyze antibodies, with a compilation
in Table5.4.
5.4. 3.1 ProtVec
ProtVec (Asgari and Mofrad2015) is a skip‑gram model trained on 546,790 sequences
from Swiss‑Prot to produce100‑dimensional protein representations. The study showed
that the resulting embeddings were able to embed a range of meaningful chemical and
physical properties, e.g., by showing they were able to classify protein families with
93% accuracy.
Inyoung and colleagues (Kim etal 2021) utilized the ProtVec model to represent the
whole BCR repertoire as a 100‑dimensional vector, by selecting the 100most frequent
CDR3 sequences from the BCR repertoire and summing their respective embeddings.
Using this approach, they were able to effectively distinguish between repertoires of
healthy subjects and COVID‑19 patients. Furthermore, they analyzed how repertoire
representation changes over the course of the disease and showed that it is possible to
track disease progression using this approach.

100 Biopharmaceutical Informatics
TABLE5.4 Language models applied to proteins
METHOD DESCRIPTION REFERENCES
ProtVec Word2Vec-based model trained on Swiss-Prot protein
data, produces 100-dimensional embeddings
Immune2vec Word2Vec model trained on BCR data, shown to
produce meaningful embeddings for 3-mers,
sequences, and repertoires
ProtBERT Transformer-based model, used to produce protein
embeddings
AbLang Transformer-based model for restoring missing
residues in antibody sequence data
AntiBERTa Sequence representation utilizing attention-based
language modeling for downstream tasks, used in
paratope prediction, achieving SoTa results
AntiBERTy BERT model used to analyze afnity maturation
process, used in MIL setup, and shown to effectively
classify bags of sequences containing potential
binders
AbBERT ProtBERT-based model ne-tuned on antibodies.
Scores that it produced for sequences correlated with
immunogenicity, protein expression in cells, and
stability metrics
BioPhi
(Sapiens)
A short description and reference are presented above for each method.
Transformer model for antibody humanization (Prihoda etal.
(Asgari and
Mofrad2015)
(Ostrovsky-Berman
etal. 2021)
(Elnaggar etal.,
n.d.) 2020
(Olsen etal.
2022b)
(Leem etal. 2022)
(Ruffolo etal.
2021)
(Vashchenko etal.
2022)
2022)
5.4.3.2 Immune2vec
Immune2vec (Ostrovsky‑Berman et al. 2021) used word2vec modeling to explore
BCR sequences. As opposed to Inyoung and colleagues who employed the pre‑trained
ProtVec neural network, Ostrovsky‑Berman etal. have trained the embedding model
on their own. Using the analogies between biological sequences and natural language
mentioned beforehand, they rst applied the Immune2Vec model to all possible com‑
binations of 8000 3‑mers. The output embeddings were then projected into a 2D plane
using t‑distributed stochastic neighbor embedding (t‑SNE) and colored according to
biophysical properties. This showed that the resulting two‑dimensional space can be
divided into clusters with homogenous biophysical properties. On the sequence level,
they trimmed CDR3 sequences to remove fragments coded by the V gene. Then, they
used their model to produce embeddings and tried to assign one of six V genes to the
resulting vectors. This demonstrated that there is an association between trimmed
CDR3 and V gene, even though the trimmed CDR3 sequence is not coded by it.
Finally, at the repertoire level, they produced embeddings for chronically infected
individuals with hepatitis C virus (HCV) and spontaneous clearers and attempted to
classify the repertoires, achieving roughly 90% accuracy for BCR data and 70% for
T‑cell receptor (TCR) data.
Соседние файлы в папке Библиотека им академика М.И. Перельмана
