Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5629_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
15.09.2026
Размер:
16 Мб
Скачать
☆
5 • Articial Intelligence and Machine Learning 91
5.3.3.1 Immunogenicity
Mice are often used as a source of therapeutic antibodies upon being challenged with a target antigen. The signicant drawback of this approach is that animal antibodies can elicit unwanted immune responses in humans. To reduce such risks, researchers need to engineer antibodies to resemble human antibody molecules without losing activity as part of the humanization process (Kim and Hong 2012). Computational methods quantifying the nativeness of human sequences have long played a signicant role in this process (Table5.3).
Initially, computational humanization was tackled by frequency‑based methods that quantied the similarity between animal and human sequences–for example, T20 (Gao etal. 2013) or humanness scores (Abhinandan and Martin 2007). These approaches were based on a small number of sequences (in thousands), offering a limited ability to determine the correlation between different residues. The availability of NGS data increased the antibody sequence samples from thousands to millions, leading to the creation of improved positional frequencies (Schmitz et al. 2020; Sheng et al. 2017; Kovaltsuk etal. 2018).
Positional proles may have been enriched with NGS data but do not reect posi‑ tional interdependencies. To quantify positional correlations, researchers developed a multivariate Gaussian (MG) statistical score based on the OAS data (Clavero‑Álvarez etal. 2018), later expanded to the LSTM model by Wollacott (Wollacott etal. 2019). Both models concentrated on predicting what constitutes the human sequence and brought to light the required element of correlations between positions. The authors of MG compared their score to therapeutic sequence immunogenicity. It resulted in a weak correlation (r2 = 0.18), suggesting that sequence identities might not encode the immunogenicity information alone. This nding is consistent with the industry experi‑ ence around immunogenicity origins. Immunogenicity toward biotherapeutic drugs can be observed in clinical trials through the generation of anti‑drug antibodies (ADAs) by patients who receive immunotherapy. Immunogenicity emerges in patients due to mul‑ tiple factors related to drug product quality (formulation, presence of aggregates in the product, or aggregation of the product in vivo during administration), patient’s specic disease history, their genetic background, and the humanness of the antibody sequence (Kumar etal. 2011; Fathallah etal. 2015; Singh 2011).
A much larger scale attempt at employing NGS data for deimmunization was pro‑ posed by Hu‑mab (Marks etal. 2021). The method is based on a random forest model trained to distinguish human and non‑human sequences of a particular V gene type from ones from other species. Hu‑mab correctly differentiated human and other animal sequences in both validation and test sets, noting slightly worse performance on the light chain, which may have been caused by the greater volume of negative training data available for variable heavy chain (VH) than variable light chain (VL) models. Another reason could be the smaller variability of light chains regarding isotypes and CDRs. The previous LSTM model (Wollacott etal. 2019) could not discriminate between human and other animal sequences, possibly because LSTM models were only trained on sequences originating in a single species (human).
92 Biopharmaceutical Informatics
TABLE5.3 Humanization methods
HUMANIZATION METHOD METHOD DESCRIPTION REFERENCES
T20 Based on thousands of sequences, offers
a limited ability to determine the correlation between a selected set of residues
Humanness scores Based on thousands of sequences, one of
the pioneering methods to quantify nativeness of antibody sequences
Improved positional
frequencies
Multivariate
Gaussian (MG) statistical score
Multivariate
Gaussian (MG) statistical score with the LSTM model
Hu-mab Based on a random forest model trained
BioPhi (Sapiens and
OASis)
AbBERT Transformer-based language model
Llamanade For capturing unique features of
IgReconstruct Based on single nucleotide frequencies, a
AbDiver Positional frequencies from OAS (Młokosiewicz etal.
A short description and reference are presented above for each humanization method.
Based on millions of sequences, quantify
nativeness of antibody sequences by determining the correlation between different residues
Based on the OAS data, statistical score
to quantify the nativeness of human sequence
Based on the OAS data, predicts what
constitutes the human sequence with usage of LSTM model
to distinguish human and non-human sequences of a particular V gene type from ones from other species
Uses language models for capturing the
diversity of natural human antibody repertoires and humanizes a sequence
trained on up to 20million unpaired heavy/light chain sequences from the OAS database
nanobodies, uses large-scale analysis of nanobodies and IgG
germline gene rearrangement tailored to the nucleotide frequency observations made in the repertoire is generated to estimate the similarity of a target Ab amino acid sequence to a given repertoire
(Gao etal. 2013)
(Abhinandan and
Martin 2007)
(Schmitz etal. 2020;
Sheng etal. 2017; Kovaltsuk etal. 2018)
(Clavero-Álvarez etal.
2018)
(Wollacott etal. 2019)
(Marks etal. 2021)
(Prihoda etal. 2022)
(Vashchenko etal.
2022)
(Sang etal. 2022)
(Schmitz etal. 2020)
2022)
5 • Articial Intelligence and Machine Learning 93
Another computational method that aims to accelerate the process of humaniza‑ tion is BioPhi, an antibody design interface that uses automated methods for captur‑ ing the diversity of natural human antibody repertoires (Prihoda et al. 2022). BioPhi offers two data‑driven approaches–humanization (Sapiens) and humanness evaluation methods (OASis). Sapiens is a deep learning humanization approach based on masked language modeling (MLM). The method is trained with human variable region antibody sequences from the OAS. Its goal is to recognize and repair masked or mutated positions in unaligned amino acid sequences. OASis, on the other hand, is a humanness metric based on peptide search in the OAS, which evaluates the humanness of a sequence by dividing it into overlapping 9‑mer peptides (inspired by human string content (Lazar etal. 2007)). Next, the metric compares them against the OAS database to predict their universality across the human population. Based on the in silico humanization bench‑ mark of 177 antibodies, this solution offers mutational choices that are similar to those achieved by experimental humanization methods. The primary advantage of BioPhi is its attention layer and granularity, allowing users to examine the residue dependencies and mutational effects on the score.
5.3.3.2 Prediction of biophysical properties: Aggregation,
viscosity, solubility, and thermostability
An mAb‑based new drug development relies on the developability properties of these molecules, such as low aggregation propensity and low viscosity (Xu etal. 2019; Jarasch etal. 2015; Makowski etal. 2021). Given the limited number of molecules with the avail‑ able sequence and biophysical property data, assessing the stability proles of antibod‑ ies at high concentrations is challenging during early stage discovery. Predictive tools based on ML may help researchers to evaluate the developability of high‑concentration antibody formulation early in the discovery process. Researchers have already applied computational tools to identify drug‑like antibodies characterized by favorable stability (Makowski etal. 2021).
Viscosity prediction received attention as well. For example, Sharma and colleagues (Sharma etal. 2014) found viscosity to be highly correlated with Fv net charge and charge symmetry. At the same time, viscosity was determined to be weakly correlated with hydrophobicity. The team proposed a linear equation based on these parameters to calculate the viscosity at 180mg/ml (pH 5.5 and 200 mM arginine‑HCl). Another vis‑ cosity predictive tool is the spatial charge map, calculated using a molecular dynamics (MD) simulation that explains the exposed surface‑negative charge distribution on the Fv region (Agrawal etal. 2016). Tomer and colleagues (Tomar etal. 2017) developed an equation for predicting the concentration‑dependent viscosity curves, which used charges on the heavy and light chain variable regions, the hinge region, as well as the hydrophobic surface area of a full‑length antibody. A summary of a comparison of the viscosity prediction tools described above can be found in Kuroda etal. (Kuroda and Tsumoto 2020). Lai and colleagues (Lai etal. 2022) recently applied ML to this problem area. The team developed a model based on 27mAbs to predict antibody viscosity at
94 Biopharmaceutical Informatics
150mg/ml, based on the decision tree classication method including two features: net charge and high viscosity index (HVI). To predict viscosity at different concentrations, the team developed a coarse‑grained model combined with hydrodynamic calculations and HVI‑derived parameters.
Regarding aggregation, researchers can leverage several in silico models for pre‑ dicting the solubility/protein aggregation rates: developability index (Lauer etal. 2012), Solubis (De Baets etal. 2015), and Camsol (Sormanni etal. 2015). Other models help to identify aggregation‑prone regions – for example, spatial aggregation propensity (Chennamsetty etal. 2009), Aggrescan 3D (Kuriata etal. 2019), and ANuPP (Prabakaran etal. 2021). ML has also been applied to predict protein aggregation kinetics (regres‑ sion) based on sequence features (Tartaglia etal. 2005; Rawat etal. 2021b) and antibody amyloidogenesis (classication) (Liaw etal. 2013; David etal. 2010; Rawat etal. 2021a; Garofalo etal. 2021). The latter has limited application in therapeutic development but is greatly relevant for human diseases (Roberts 2014). Another ML‑based model proposed by Lai and colleagues (Lai etal. 2021) was trained on 21mAbs to predict therapeu‑ tic antibody aggregation rates at 150mg/ml with the help of structural‑based features extracted from MD simulations.
Computational approaches and solubility prediction tools open the door to fast high‑throughput screening without any material requirements (Feng etal. 2022). Current approaches use molecular descriptors extracted from protein sequence–sequence‑based predictors (Hebditch et al. 2017; Sormanni etal. 2017) –or from structures– struc‑ ture‑based predictors (Sormanni etal. 2015). The former often fail to consider tertiary structure information, which differentiates poorly soluble residues that drive protein folding from the residues exposed to the solvent, potentially eliciting aggregation. Structure‑based tools, on the other hand, can be applied only when researchers have a reliable 3D model, limiting the throughput and applications to a massive volume of early stage mAb candidates. Some computational methods only output a binary classica‑ tion (soluble/insoluble) instead of a numerical value (Hebditch etal. 2017; Smialowski etal. 2012). Feng and colleagues (Feng et al. 2022) presented a strategy called sol‑ Predict that uses the embeddings from pre‑trained protein language model to predict the apparent solubility of mAbs in histidine (pH 6.0) buffer. The team used a dataset of 220 diverse, in‑house mAbs for model training and hyperparameter tuning through vefold cross‑validation. The solPredict was found to achieve a high correlation with experimental solubility through an independent test set of 40mAbs, delivering a good performance for both IgG1 and IgG4 subclasses. Hashemi and colleagues (Hashemi etal. 2022) modeled a soluble production of the anti‑EpCAM extracellular domain sin‑ gle‑chain variable fragment (scFv). The solution was then optimized as a function of four numerical factors: post‑induction temperature, post‑induction time, cell density of induction time, and inducer concentration, as well as one categorical variable using arti‑ cial neural network (ANN) and response surface methodology (RSM). The predicted value obtained by ANN for the response (106.1mg/L) approached the experimental result closer than the one obtained by RSM (97.9mg/L).
The application of computational tools to predict stability‑enhancing mutations has increased in popularity over the past decade. Researchers use the state‑of‑the‑art approach called protein consensus design, which leverages phylogenetic information from
5 • Articial Intelligence and Machine Learning 95
multiple sequence alignments (MSAs) to output the most frequent 34 amino acids for a residue position. However, such residues are characterized by improved thermostability in only 50% of the cases, with poorer performance noted for antibody sequences that have highly conserved framework regions. ML approaches employed large datasets such as ProTherm (Kumar etal. 2006), which predicts thermostability by collating mutant 40 effects on protein stability from mesophilic and thermophilic sequences. However, the publicly available thermostability data exclude mAb and scFv sequences, placing a limit on their generalization. It is still a challenge for researchers to predict thermally enhanc‑ ing scFv sequences or develop thermostable mutations within approved scFv candidates. New computational approaches will need to be designed to combine sequence informa‑ tion to predict biophysical attributes (thermostability) and highly accurate, generalizabil‑ ity‑enabling design. Harmalkar and colleagues (Harmalkar etal. 2022) addressed this by equipping unsupervised and supervised learning approaches over a thermostability prediction task that has been tuned specically for generated scFv sequences.
Altogether, the prediction of biophysical properties of therapeutic antibodies is challenging since most of the studies cited perform the predictions within a specic experimental framework. For instance, “aggregation” can be measured by a variety of assays (e.g., hydrophobic interaction chromatography (HIC)) that, though correlating with one another, do not need to give the same reading for the same molecule. This makes generalizability extremely difcult and rendering benchmarking different meth‑ ods developed with different datasets almost impossible.
5.3.3.3 Deep learning applications
Current efforts in developability prediction using deep learning methods are directed toward generating novel sequences computationally and selecting for good manufactur‑ ability properties. This can be done by generating the molecules and ltering them using general‑purpose developability predictors such as therapeutic antibody proler (TAP) (Mason etal. 2021). Alternatively, a GAN trained on large‑scale natural data can be ne‑tuned on sequences with specic developability properties (Amimeur etal. 2020). All in all, there is an interplay between how to efciently sample the antibody sequence/ structure space and how to predict the developability properties well.
Deep learning models aim to learn higher order dependencies missed by posi‑ tion frequency analysis, reducing the chance of generating non‑functional proteins. However, current limitations of deep learning approaches focus only on CDR regions or heavy chains and a lack of experimental validation of predicted properties (Khetan etal.
2022). For instance, 74% antigen binding was achieved in a mouse library designed by a variational autoencoder that generated novel CDR‑H regions (Friedensohn etal. 2020). However, such an approach ignores non‑CDR region contributions to the paratope, and the diversity of sequences in this library is unknown.
Other generative approaches, such as GANs, can also be trained on natural human antibody repertoires and biased via transfer learning to generate sequences predicted to have the desired biophysical properties (Amimeur etal. 2020). However, researchers need more insights to better understand how such properties impact the overall developa‑ bility of an antibody. Experimental validation of the predicted properties is also required.
96 Biopharmaceutical Informatics
It has already been conducted for an enzymatically active protein library (Repecka etal.,
2021) and a nanobody library created using a generative deep neural network‑powered autoregressive model trained on a native llama repertoire (Shin etal. 2021).
Previous studies have demonstrated the use of libraries of mammalian antibod‑ ies for selecting antibodies with optimal biophysical properties, reduced polyreactivity, and immunogenicity (Dyson etal. 2020). Dyson and colleagues described the use of a nuclease‑directed integration system to generate antibodies with differing biophysi‑ cal properties only based on their display level on mammalian cells. Other researchers have demonstrated the use of ML‑guided directed evolution on combinatorial sequences (Wu etal. 2019). Recently, Chen and colleagues formulated ML pipelines to predict the developability of a library of 2,400 antibodies using sequence alone (Chen etal. 2020). Such advances in bioinformatics and in silico methods allowed for the efcient develop‑ ment of commercially viable antibodies. Antibody library variants of an original candi‑ date are created to exhibit better developability than the parent molecule.
Generative ML has been proposed as a means to design antigen‑specic mAbs, but efforts to conrm this hypothesis have been hampered by the cost and time required to test large numbers of antibody sequences. To address this challenge, Akbar and colleagues (Akbar et al. 2022) developed a lattice‑based antibody‑antigen bind‑ ing simulation framework incorporating physiological parameters. They found that a deep generative model trained only on antibody sequence data can be used to design conformational epitope‑specic antibodies matching or exceeding the training dataset in afnity and diversity, showing that lower sequence diversity is necessary for high‑ accuracy generative antibody modeling.

5.4 ANTIBODY GENERATION AND DESIGN BY LANGUAGE MODELS

Designing novel antibodies without resorting to experimental methods has long been the end goal of computational antibody engineering. The earliest attempts indicated the possibility of performing exhaustive in silico mutagenesis combined with binding energy estimation on a cocrystal structure (Lippow etal. 2007). Designing the antibod‑ ies without the cocrystal structure has been performed by variations on sampling the available structures (Li etal. 2014; Adolf‑Bryfogle etal. 2018). One of the primary chal‑ lenges in this area relates to creating an antibody sequence/structure that is plausible in a biological sense. Language models, such as Bidirectional encoder representations from transformers (BERT), have been applied to push antibody manifolds into representation space, sampling which facilitates novel antibody discovery.
Here, we describe the importance of antibody representation in terms of data inter‑ pretability and downstream model performance. We show how such representations can be enriched with domain knowledge by both manual feature engineering processes and representation learning techniques. Finally, we present how language models are used to produce antibody embedding and improve the antibody discovery process.
5 • Articial Intelligence and Machine Learning 97

5.4.1 Antibody Representations

Data representation is a vital aspect of all ML projects as raw unprocessed data often cannot be used directly by the model. It is important that feature engineering becomes an intrinsic step of the ML development pipeline, as good representation signicantly boosts downstream model performance and stability or reduces the input dimensional‑ ity. In this process, input data are modied and augmented with additional information and domain knowledge relevant to the problem being solved.
Several approaches to representing proteins can be used in parallel, each capturing a different perspective. For example, one can represent an antibody as a sequence of amino acids. Then, with the presence of structural data, it is possible to represent protein as a structure–for instance, by extracting voxelated atomic maps. Despite being very easy to interpret, such an approach has a major drawback: it is sensitive to rotations. That is why graphs are another popular option for representing protein structural data.
The function and the structure of a protein derive from the interaction between all of its residues. Such a residue interaction network can be abstracted to a graph where amino acids constitute nodes and relationships between them are the edges. One can use several denitions to construct such a graph: for example, using Cα atoms as nodes with connecting edge if Cα atoms are within 7 Å (Chakrabarty and Parekh 2016; Pittala and Bailey‑Kellogg 2020), analogously for Cβ atoms with distance thresholds between 7 Å and 10 Å (Chakrabarty and Parekh 2016; Pittala and Bailey‑Kellogg 2020), or using residue center of mass as a node with edge if any atoms are closer than 5 Å. Another approach introduces edges if the non‑covalent interaction strength between interacting residues is above some predened threshold (Brinda and Vishveshwara 2005). There also exist methods using different graph levels (atom‑atom interactions followed by residue‑residue) at the same time (Tubiana etal. 2022).
Another challenge is posed by representing molecular surfaces of a protein for ML. Gainza and colleagues (Gainza etal. 2020) applied geometric deep learning to learn surface ngerprints that are important for biological interactions and tested their model on three distinct problems: classication of ligand binding pockets, protein binding site prediction, and protein‑protein complex search based on interaction ngerprints.
In the context of enriching representations, one can augment biological data on residue, sequence, and structure layers in the feature engineering process. The easiest residue representation is likely one‑hot encoding where each residue is represented as a 20‑dimensional unit vector with 1 in place of represented amino acid. This does not add any domain knowledge to the residue representation but enables the model to treat all residues equally, as all amino acid representations differ only by rotation, and all pairwise distances are the same.
Evolutionary relationships can be encapsulated into data. In this case, one can rep‑ resent each residue as a 20‑ (or 21‑) dimensional vector, where each value is taken from the BLOSUM substitution matrix (Lim etal., 2022). A common practice is to include the physicochemical properties of residues by adding Vectors of Hydrophobic, Steric, and Electronic properties to the representation (Liberis etal. 2018).
One can embed residues’ positional information on the sequence level into sequence representation. To compare protein sequences in a meaningful way, one needs to align them
98 Biopharmaceutical Informatics
so that the proteins from the same family will have corresponding evolutionary‑dependent residues in the same positions. Upon sequence alignment, inputs have a xed size, which licenses the use of simple models that can discover patterns and residue relationships from absolute input positions, which may be important for model interpretability.
Detlefsen and colleagues (Detlefsen et al. 2022) showed that alignment of the input sequences improved the model’s ability to separate sequences from one protein family (b‑lactamase) into different phyla. It remains unclear if it is possible to general‑ ize this capability across multiple protein families. For antibodies, sequence alignment is achieved by numbering methods, which also provide the functional position context (e.g., CDR, framework) for each residue. At the structural level, one can augment input data with shape annotations for each residue. For example, ASA or paratope labels.
Most ML models natively employ vector data format of constant dimensionality. To utilize those models, one needs to transform variable‑in‑length biological sequences to point in xed‑size dimensional space so that similar input sequences, or functionally similar proteins, reside close to each other in the vector space.
Such transformation is called embedding, and to dene it, one needs a similarity measure and the transformation function that transforms input to a point in the embed‑ ded space. Proteins are sequences of amino acids with variable lengths that fold into a three‑dimensional shape. Therefore, as an example of a similarity measure, one can use Levenshtein distance between sequences or some shape difference measures.

5.4.2 Representation Learning

As opposed to the manual process of feature engineering, which involves manually adding domain knowledge to the representation, it is possible to apply representation learning techniques to render meaningful interpretable data representations. The basic representation learning approach is PCA, where the representation features are the out‑ put of linear transformations of input data. Representation learning is often used in combination with transfer learning, especially when there is little data available. That is because the model can learn to recognize meaningful biological features by making use of vastly available NGS data.
The biological knowledge embedded into the vector and the meaningfulness of such learned representations can be then assessed by the performance of downstream model predictions.
Detlefsen and colleagues (Detlefsen etal. 2022) demonstrated that low‑dimensional sequence representation space topology shows good correspondence between the ances‑ tral lineage of the proteins that were rendered on latent space. The team also revealed that the learned representation space is not Euclidean. Hence, arithmetic operations on vector representations are not a good approach to interpolate points in the latent space as one might get unpredictable results. They equipped learned representations with a Riemannian metric, so measuring the distance in the data space was possible. This demonstrated that the usage of geodesic distances enables more meaningful navigation over the latent space.
5 • Articial Intelligence and Machine Learning 99

5.4.3 Language Models

Antibody diversity is estimated to be 1018 unique molecules (Briney et al. 2019). Therefore, with plenty of publicly available NGS datasets, efcient sequence embedding that captures its biological properties holds the promise of the possibility to navigate through the whole immunoglobulin sequence space, providing new insights into the biology of antibodies. Such navigation could be potentially used in the protein engi‑ neering process to interpolate between known sequences in latent space to yield new proteins functionally similar or to guide mutations to design antibodies with desired characteristics.
An emerging trend in antibody analysis is the use of natural language processing (NLP) methods to develop antibody embeddings by making analogies between amino acid k‑mers, whole antibody sequences, antibody repertoires, and words, sentences, and documents. In NLP, the most widely used family of models to produce word embed‑ dings is word2vec (Mikolov etal. 2013). Two architectures are used for representation learning:
• Skip‑Gram, where given an input word, the neural network tries to predict its context (preceding and following words).
• Continuous Bag of Words architecture, where given context (neighbor words), the model is tasked with predicting the target word.
Because of the training process, output vectors embed contextual information and are capable of capturing both syntactic and semantic relationships, so distinct words that are related by the context are adjacent to each other in the vector space. Such con‑ text‑aware embedding can be called distributed representation. Below, we highlight some of the language models that are used to analyze antibodies, with a compilation in Table5.4.
5.4. 3.1 ProtVec
ProtVec (Asgari and Mofrad2015) is a skip‑gram model trained on 546,790 sequences from Swiss‑Prot to produce100‑dimensional protein representations. The study showed that the resulting embeddings were able to embed a range of meaningful chemical and physical properties, e.g., by showing they were able to classify protein families with 93% accuracy.
Inyoung and colleagues (Kim etal 2021) utilized the ProtVec model to represent the whole BCR repertoire as a 100‑dimensional vector, by selecting the 100most frequent CDR3 sequences from the BCR repertoire and summing their respective embeddings. Using this approach, they were able to effectively distinguish between repertoires of healthy subjects and COVID‑19 patients. Furthermore, they analyzed how repertoire representation changes over the course of the disease and showed that it is possible to track disease progression using this approach.
100 Biopharmaceutical Informatics
TABLE5.4 Language models applied to proteins
METHOD DESCRIPTION REFERENCES
ProtVec Word2Vec-based model trained on Swiss-Prot protein
data, produces 100-dimensional embeddings
Immune2vec Word2Vec model trained on BCR data, shown to
produce meaningful embeddings for 3-mers, sequences, and repertoires
ProtBERT Transformer-based model, used to produce protein
embeddings
AbLang Transformer-based model for restoring missing
residues in antibody sequence data
AntiBERTa Sequence representation utilizing attention-based
language modeling for downstream tasks, used in paratope prediction, achieving SoTa results
AntiBERTy BERT model used to analyze afnity maturation
process, used in MIL setup, and shown to effectively classify bags of sequences containing potential binders
AbBERT ProtBERT-based model ne-tuned on antibodies.
Scores that it produced for sequences correlated with immunogenicity, protein expression in cells, and stability metrics
BioPhi
(Sapiens)
A short description and reference are presented above for each method.
Transformer model for antibody humanization (Prihoda etal.
(Asgari and
Mofrad2015)
(Ostrovsky-Berman
etal. 2021)
(Elnaggar etal.,
n.d.) 2020
(Olsen etal.
2022b)
(Leem etal. 2022)
(Ruffolo etal.
2021)
(Vashchenko etal.
2022)
2022)
5.4.3.2 Immune2vec
Immune2vec (Ostrovsky‑Berman et al. 2021) used word2vec modeling to explore BCR sequences. As opposed to Inyoung and colleagues who employed the pre‑trained ProtVec neural network, Ostrovsky‑Berman etal. have trained the embedding model on their own. Using the analogies between biological sequences and natural language mentioned beforehand, they rst applied the Immune2Vec model to all possible com‑ binations of 8000 3‑mers. The output embeddings were then projected into a 2D plane using t‑distributed stochastic neighbor embedding (t‑SNE) and colored according to biophysical properties. This showed that the resulting two‑dimensional space can be divided into clusters with homogenous biophysical properties. On the sequence level, they trimmed CDR3 sequences to remove fragments coded by the V gene. Then, they used their model to produce embeddings and tried to assign one of six V genes to the resulting vectors. This demonstrated that there is an association between trimmed CDR3 and V gene, even though the trimmed CDR3 sequence is not coded by it. Finally, at the repertoire level, they produced embeddings for chronically infected individuals with hepatitis C virus (HCV) and spontaneous clearers and attempted to classify the repertoires, achieving roughly 90% accuracy for BCR data and 70% for T‑cell receptor (TCR) data.