Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5387_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Contents
- •Foreword
- •Preface
- •About the Editors
- •Contributors
- •References
- •2.3.4 Barriers to Automation Adoption
- •2.4 Core Ingredients for Successful Digital Transformation
- •2.1 Introduction
- •2.3.1 Operational Challenges
- •2.3.2 Cultural Challenges
- •2.4.2 Cloud Computing
- •2.5 Case Studies of Successful Digital Transformation
- •2.6 Conclusion
- •References
- •3. Computational Protein Design Strategies for Optimization of Antigen Generation to Drive Antibody Discovery
- •3.1 Introduction
- •3.3 Antigen Generation Strategies
- •3.4 Computational Methods
- •3.4.2 Computational Protein Structure Prediction
- •References
- •4. Bioinformatic Analyses of Antibody Repertoires and Their Roles in Modern Antibody Drug Discovery
- •4.1 Introduction
- •4.6 Summary and Future Directions
- •Acknowledgments
- •References
- •5.1 Introduction
- •5.2 Databases
- •5.2.1 Databases in Machine Learning Approaches
- •5.2.2 Database Types
- •5.3 Applications of Machine Learning in Antibody Discovery and Development
- •5.3.1 Structure Prediction with Deep Learning
- •5.3.3 Developability
- •5.4 Antibody Generation and Design by Language Models
- •5.4.1 Antibody Representations
- •5.4.2 Representation Learning
- •5.4.3 Language Models
- •References
- •6.1 Introduction
- •6.2 Antibody Generation through Deep Generative Models
- •6.3.1 Sampling and Scoring
- •6.5 Conclusions and Perspectives
- •Acknowledgments
- •References
- •7.1 Introduction
- •7.2.3 Computational Approaches to Predict Antibody–Antigen Interaction
- •7.3 Conclusion
- •Competing Interests
- •Acknowledgments
- •References
- •8.2 Common Types of Molecular Simulations for Biomolecules
- •8.2.1 Molecular Dynamics (MD) Simulations
- •8.2.2 Monte Carlo (MC) Simulations
- •8.2.3 Challenges of Molecular Simulations
- •8.3.1 Periodic Boundary Conditions
- •8.4 Uses of Molecular Simulation in Antibody Drug Development
- •8.5 Conclusion
- •References
- •9. Considerations of Developability During the Early Stages of Antibody Drug Discovery and Design
- •9.1 Introduction
- •9.2 Historical Perspective
- •9.3 Clinical Antibody Data Set
- •9.5 Control Antibodies
- •9.7 Assessment of Chemical Liabilities
- •9.8 Conclusions and Future Perspectives
- •Acknowledgments
- •References
- •Abbreviations
- •10.1 Introduction
- •10.4.1 Conclusions and Outlook
- •Acknowledgments
- •References
- •11.8 Conclusions and Future Directions
- •References
- •12.1 Introduction to PK/PD and QSP Modeling
- •12.1.1 PK/PD Modeling
- •12.1.2 QSP Modeling
- •12.2.1 Monoclonal Antibodies (mAbs)
- •12.2.3 Cell Therapies
- •12.2.4 Gene Therapies
- •12.2.5 Vaccines
- •12.2.6 mRNA/siRNA/Oligonucleotide Therapeutics
- •12.4 Case Studies
- •12.5 Conclusions and Future Perspectives
- •References
- •13.1 Introduction
- •13.2 AI/ML: A Game Changer for Antibody Design
- •13.3 Multispecific Antibody Design
- •13.4 Adapting AI to the Design of Multispecific Antibodies
- •13.4.1 Structure Prediction and Modeling
- •13.4.2 Developability Prediction and Optimization
- •13.4.4 In Silico Modeling and Simulation
- •13.5 The Future: Beyond Optimization
- •13.5.1 Market Trends and Commercialization
- •13.5.2 Logic Gates, Biosensors, and De Novo Design
- •13.5.3 Challenges and Opportunities
- •13.6 Conclusion
- •Acknowledgments
- •References
- •Index

4 • Analyses of Antibody Repertoires 61
but provided only a glimpse into the antibody diversity landscape, is truly remarkable.
NGS technologies for antibodies vary signicantly in their applications and capabili‑
31,32
ties, as summarized in Table4.1 and previously described.
For instance, Roche’s
454 sequencing platform, although now discontinued, was once essential for its ability
to produce long reads up to 700 bp, making it ideal for sequencing the entire vari‑
able (V) region of antibodies in a single run. On the other hand, the Illumina platform,
which produces shorter reads of about 150 nucleotides, is particularly well‑suited for
TABLE4.1 NGS platforms for sequencing antibody repertoires from in vivo immunized
animals and in vitro display libraries.
MAXIMUM
READ
PLATFORM
Illumina
MiSeq
Illumina
HiSeq
454 GS
FLX
Ion
Torrent
PacBio ~ 20 kb ~4million 10–30 hours ~11
The read lengths, number of reads per run, and run times can vary based on the specic
model and conguration of the sequencing kit or chips used for each platform.
LENGTH
2 X 300 bp ~ 25million 4–55 hours ~0.1 Suitable for bulk
2 X 150 bp ~ 1.5 billion ~48 hours ~0.1 Suitable for bulk
~ 700 bp ~ 1million 10–24 hours ~1 Offered longer read
~ 400 bp ~80million 2–7 hours ~1 Suitable for medium
READS PER
RUN RUN TIME
ERROR
(%) REMARKS
sequencing of variable
regions from heavy and
light chains separately.
Low error rate and high
throughput.
sequencing of heavy and
light chains separately.
Low error rate and high
throughput. Principle
platform used for 10X
Genomics sequencing,
enabling barcoding of
single cells for paired
information such as
CDR-H3 and CDR-L3
pairing.
lengths suitable for scFv
libraries with sequencing
of heavy and light chains
separately but
discontinued technology
with limited availability.
length reads, offering a
balance of speed and
detail.
Long reads enable
(raw),
<1 (CCS)
full-length sequencing of
heavy/light chains in
single reads suitable for
scFv or Fab libraries.

62 Biopharmaceutical Informatics
analyzing the heavy chain complementarity determining region 3 (HCDR3), a cru‑
cial segment in determining antibody specicity. Although the HCDR3 sequences are
shorter, Illumina’s high throughput facilitates a more extensive examination. Emerging
platforms like Pacic Biosciences (PacBio) offer the potential to sequence even longer
reads from single DNA or RNA molecules, potentially simplifying the complex library
preparations required for scFv and Fab libraries. However, their application in antibody
repertoire analysis may be constrained by the quality of the sequences they produce.
Meanwhile, platforms like Ion Torrent strike a balance with their moderate read lengths
and faster sequencing capabilities at a lower cost, although their error rate remains a
concern. The continued development and renement of these technologies, including
error correction strategies such as the use of unique molecular identiers (UMIs), is
crucial as they evolve to meet the high standards required for precise and effective
antibody discovery.
Once the raw data from NGS sequencing runs are available, bioinformatics tools
are utilized to process and analyze the data, facilitating the antibody discovery pro‑
cess from understanding diversity to NGS‑based screening and selection of functional
clones. A selected list of bioinformatics software tools for analyzing antibody repertoire
data generated from NGS is presented in Table4.2. Initially, the process involves data
cleaning, which includes removing low‑quality sequences and sequencing adapters to
ensure data integrity. Subsequent steps involve using tools such as IMGT/V‑QUEST
and IgBLAST to identify the variable (V), diversity (D), and joining (J) gene segments,
and to dene the framework regions and complementarity determining regions (CDRs)
essential for antigen‑binding specicity. The selection of these tools often depends on
the volume of data and the user’s expertise. For those less familiar with computational
TABLE4.2 Selected bioinformatics software tools for NGS-generated antibody repertoire
data analysis.
TOOLS DESCRIPTION URL
IgBLAST V(D)J annotation, CDRs assignment, and SHM https://www.ncbi.nlm.
nih.gov/igblast/
IMGT/
HighV-QUEST
MiXCR Raw sequences to clonotypes https://github.com/
PipeBio User-friendly end-to-end workows with
ENPICOM IGX Platform, data visualization and analysis, and
Geneious
Biologics
IgBLAST is freely available for use, while the other tools listed require licensing and involve
costs.
A pioneering international information system in
immunogenetics and immunoinformatics with
tools available for V(D)J annotation, CDRs
assignment, SHM, statistical analysis, and plots
visualization for antibody discovery, NGS analysis,
and antibody engineering
AI-aided antibody development
Powerful tools for antibody sequence annotation
and analysis, statistical analysis, and visualization
https://www.imgt.org/
milaboratory/mixcr/
https://pipebio.com/
https://enpicom.com/
https://www.geneious.
com/biopharma/

4 • Analyses of Antibody Repertoires 63
techniques, platforms like IMGT High‑VQUEST offer user‑friendly interfaces, while
more sophisticated tools like MiXCR provide detailed immunological analyses for
33
advanced users. Furthermore, clustering tools such as CD‑hit
or UCLUST34 are
employed to group similar sequences, aiding in the identication of clonal expansions
and enriching our understanding of immune response diversity. For handling larger
datasets and facilitating collaborative research, cloud‑based commercial software plat‑
forms like PipeBio, ENPICOM, and Geneious Biologics provide comprehensive, inte‑
grated solutions that streamline the workow and enhance data analysis capabilities.
These platforms combine powerful bioinformatics tools with user‑friendly interfaces,
signicantly simplifying the management and analysis of complex data, thereby accel‑
erating antibody research and development.
After performing antibody NGS data processing and analysis, several metrics
and analytical techniques are employed to thoroughly examine the B‑cell repertoire.
Figure4.1 illustrates various NGS‑based antibody repertoire metrics, such as isotype
analysis, V(D)J segment usage frequencies, CDR3 properties, somatic hypermutation
(SHM) analysis, and clonal relationship and lineage analysis, which are depicted using
trees or network graphs. Statistical analysis is also crucial to ensure sequencing depth
is comparable between samples, estimate repertoire diversity, and assess convergence.
By employing this comprehensive approach, researchers can gain deep insights into the
antibody response, including the types of antibodies produced, namely isotypes, the
genes utilized for antibody building blocks comprising V(D)J gene segments, the CDRs,
the extent of antibody diversication through mutations, and the relationships between
different antibody clones. This information is invaluable for understanding the immune
system’s response to pathogens and for developing novel antibody‑based therapies.
FIGURE 4.1 NGS antibody repertoire data analysis employs a multifaceted approach to
dene key repertoire features, serving as essential metrics for NGS-driven antibody discovery. This analysis includes isotype analysis to categorize sequences by constant regions
inuencing effector functions, and V(D)J usage analysis to assess gene segment frequencies, revealing specic germline contributions. Complementarity determining regions (CDRs)
analysis, particularly focusing on CDR3, examines sequence length and amino acid properties crucial for antigen binding. Somatic hypermutation (SHM) analysis explores the diversity of antibody clones, a critical process for generating high-afnity antibodies. Clonal
relationship and lineage analysis track antibody evolution, offering insights into immune
response development. Lastly, with paired sequencing data, chain pairing analysis evaluates the pairing of heavy (VH) and light (VL) chain genes, enhancing our understanding of
antibody specicity.

64 Biopharmaceutical Informatics
In the case of NGS‑based antibody discovery using display libraries, the depth and
breadth of analysis that can be performed on the outputs from selections of these librar‑
ies have dramatically increased with the advent of NGS.
usage frequency, along with other immunogenetic features, such as specic gene lin‑
eages, CDR lengths, and amino acid patterns, can be analyzed to identify trends associ‑
ated with different panning strategies.
11,35
Additionally, data on gene
4.3 NGS‑ENABLED IN VIVO ANTIBODY
DISCOVERY FROM IMMUNIZED ANIMALS
In vivo antibody discovery using a hybridoma‑based approach was established several
decades ago and has subsequently undergone several improvements.
involves immunizing animals with a specic antigen, followed by the laborious process
of fusing B‑cells from the animal’s spleen with myeloma cells, creating immortalized
hybridomas. The generated hybridomas continuously produce antibodies specic to the
immunizing antigen. However, hybridoma technology suffers from several limitations.
Firstly, it captures only a tiny fraction of the vast diversity of antibodies generated by
the immune system. Secondly, the process is highly inefcient and time‑consuming,
requiring extensive screening of individual hybridoma clones to identify those produc‑
ing high‑afnity antibodies. The emergence of NGS technologies has revolutionized in
vivo antibody discovery, offering a powerful and multifaceted approach.39 This revolu‑
tion extends far beyond traditional hybridoma technology by incorporating advanced
technologies that work synergistically with NGS analysis. Particularly, human immu‑
noglobulin transgenic mice offer a more human‑like antibody repertoire for NGS
sequencing. Furthermore, NGS combined with single B‑cell antibody technologies,
antigen‑specic single B‑cell sorting, and recombinant antibody cloning have become
powerful approaches.
from B‑cell antibody genes, preserve the crucial pairing of heavy and light chain genes
for proper antibody function. NGS analysis of these libraries allows for the identica‑
tion of a vast repertoire of diverse and functional antibodies. Figure4.2 outlines the
NGS‑driven in vivo antibody discovery process, starting with B‑cell isolation and
proceeding through sequencing and comprehensive repertoire analysis. This includes
examining V(D)J gene usage, SHM diversity, clonal expansion, and antibody lineage
relationships. The method facilitates the identication of highly diverse, high‑afnity
antibodies that are not detectable by traditional in vivo hybridoma methods.
In one of the use‑cases as adopted by researchers at Genentech, integrating NGS
into hybridoma technology has signicantly advanced the antibody discovery process.42
This approach overcomes traditional limitations related to throughput and storage
capacity by digitizing antibody variable domain sequences for high‑throughput screen‑
ing directly from hybridoma cells. The use of barcoded primers in a 96‑well format
allows for the simultaneous processing of multiple samples, enhancing sequencing out‑
put and efciency substantially. Additionally, robust bioinformatics tools ensure precise
40,41
Additionally, natively paired immune libraries, constructed
36–38
This method

4 • Analyses of Antibody Repertoires 65
FIGURE 4.2 NGS-enabled in vivo antibody discovery process starts with isolating bulk
or single B-cells from a model organism and utilizes NGS for in-depth repertoire analysis.
This analysis investigates V(D)J gene usage, explores SHM diversity critical for high-afnity
antibodies, examines clonal expansion of desirable B-cells, and elucidates antibody lineage
relationships. This comprehensive NGS approach allows the discovery and characterization
of high-afnity antibodies in vivo.
resolution and unambiguous identication of antibody sequences. This method dramati‑
cally reduces the need for physical storage by converting sequences into digital data,
offering a scalable, cost‑effective solution that aligns perfectly with the needs of modern
biomedical research and drug development. Additionally, it facilitates a more dynamic,
rapid, and expansive exploration of potential therapeutic antibodies, demonstrating a
critical advancement in the eld.
In another use‑case study, George Georgiou and colleagues developed a ground‑
breaking method that bypasses the traditional high‑throughput screening processes
for isolating antigen‑specic monoclonal antibodies (mAbs).10 By leveraging NGS
and bioinformatic analysis, this technique directly mines the antibody variable region
(V)‑gene repertoires from bone marrow plasma cells (BMPCs) of immunized mice.
BMPCs, notable for producing most circulating antibodies but are unable to be immor‑
talized, exhibit a highly polarized V‑gene repertoire following immunization. The most
abundant variable heavy (VH) and variable light (VL) genes are paired based on their
relative frequencies, reconstructed through automated gene synthesis, and expressed in
bacterial or mammalian systems as recombinant antibodies. This method signicantly
accelerates the progression from immunization to specic antibody production and
yields antibodies with high antigen specicity and nanomolar afnities, underscoring
its potential as a tool to rapidly respond to emerging infectious diseases and advancing
both immunological research and therapeutic antibody development.
In another case study by GigaGen Inc., the authors leveraged NGS to signicantly
advance the discovery of therapeutic monoclonal antibodies by deeply sequencing yeast
scFv libraries derived from immunized Trianni mice, both before and after uores‑
43
cence‑activated B‑cell sorting (FACS).
This approach allowed the authors to track
the evolution and enrichment of antibody clones, showcasing the technology’s ability to
narrow down thousands of diverse clones to a few with potential therapeutic properties.
The sequence analysis highlighted minimal overlap in antibody sequences enriched by

66 Biopharmaceutical Informatics
different methods, pointing to the inuence of antigen presentation on immune response
specicity. Remarkably, sequences derived from soluble immunogen methods showed
more similarity across FACS methods compared to those from cell/DNA methods,
suggesting differences in immune response elicitation. This study exemplies how
integrating NGS with innovative immunization strategies can streamline the antibody
discovery process, enhance the functional relevance of identied antibodies, and expe‑
dite the development of new therapeutics.
Overall, NGS‑based in vivo antibody discovery offers a revolutionary approach,
merging the strengths of animal immunization with advanced screening methods. This
opens the door for exploring a vast universe of possibilities within the animal‑derived
antibody repertoire.
4.4 NGS‑ENABLED IN VITRO ANTIBODY
DISCOVERY FROM DISPLAY LIBRARIES
In vitro antibody library platforms such as phage and yeast display have become power‑
ful tools for therapeutic antibody discovery, allowing the identication of fully human
antibodies against virtually any desired target.
advantages as they provide: (i) large antibody sequence and structural diversities, typi‑
cally up to 1011 unique clones, (ii) fast and easy selection methods able to apply specic
enrichment pressures on the entire library, such as incubation with solid‑surface immo‑
bilized antigen or FACS, and (iii) rapid identication of binder antibodies by providing
a direct phenotype‑genotype link. Traditional screening methods of in vitro display
libraries and their enrichment outputs, however, can be time and resource intensive.
These methods rely on the isolation of single colonies for various binding or functional
assays and subsequent Sanger sequencing, thus limiting the number of tested antibodies
to typically a few hundred. Additionally, most of the clones picked for characteriza‑
tion in this way represent antibodies present at the highest frequency in the enriched
population. This inevitably means that clones picked at random have a large level of
redundancy, with those high frequency clones being picked multiple times, reducing
even further the number of tested unique antibodies. Furthermore, this dominance of
high frequency clones does not represent only antibodies that have been enriched due to
high‑afnity antigen binding but also those that might have been enriched due to higher
expression or display of the antibodies by the phage system. The presence of such clones
at high frequency will often lead to the failure of the traditional screening methods in
the detection of rarer but more desirable and functional antibodies.
More recently, the combination of NGS and in vitro display technologies has
largely removed the limitations of traditional screening methods and made the use of
these platforms in therapeutic antibody discovery even more powerful.32 NGS allows
deep repertoire proling and mining of antibody libraries and their enrichment outputs,
providing a more holistic view of the population dynamics and the evolution of the
antibody repertoires through consecutive enrichment rounds. As detailed in Figure4.3,
7,4 4–47
These platforms have several
48

4 • Analyses of Antibody Repertoires 67
FIGURE4.3 NGS-enabled in vitro antibody discovery process begins with the panning of
phage-displayed libraries across multiple rounds, typically three rounds (R1, R2, R3), each
targeting a specic antigen. Post-panning, phage particles from each round undergo NGS
and antibody repertoire analysis. This analysis includes evaluating the frequency distribution of sequences from each round, clustering to identify related antibody sequences and
constructing phylogenetic trees to track the evolution and expansion of antibody hits. This
comprehensive approach facilitates the identication and optimization of high-afnity antibodies in vitro.
the NGS‑enabled in vitro discovery process utilizes iterative panning of phage libraries
against a target antigen, followed by NGS analysis, for example, frequency distribution,
sequence clustering, and phylogenetic trees, to identify and optimize high‑afnity anti‑
body candidates. With NGS, not only the high frequency enriched clones but also those
that are rarer with lower frequencies become visible, resulting in a more diverse panel
of identied antibody hits.11 Moreover, as almost all enriched clones are detected, larger
clusters of enriched sequences (families of sequences closely related in sequence space)
can be identied. Therefore, selection of representative clones from these clusters can
be made in a more intelligent manner, addressing, for example, developability criteria
of the selected clones early in the discovery phase.49 Moreover, having access to these
multi‑member sequence clusters presents the possibility of “hit expansion.” This is the
strategy where variants of a known functional antibody sequence are identied in NGS
data with the objective of improving some property of the parental antibody, that is,
afnity, developability, cross‑reactivity, etc.
50
First studies of incorporating NGS in phage display antibody discovery, due to the
limited capacity of the sequencing length of sequencing technologies existing at the
time, focused only on HCDR3 or the heavy chain variable region (VH) sequences.
51,52

68 Biopharmaceutical Informatics
The short‑read sequencing approaches have the advantage of offering a very high
sequencing depth and are very well‑suited for the global analysis of library quality, for
example, looking at the diversity of antibody chains separately. Several groups have also
used these sequencing methods for NGS‑based clone selection, using clone rescue meth‑
ods to recover the full antibody sequence. However, these studies have been limited to
libraries with variability introduced only in specic segments of the variable gene or
to the selection of only a handful of clones for expression and characterization.48 The
Fischer lab, for example, used HCDR3 sequences with high enrichment ratio obtained
from NGS to design oligos and PCR‑amplify the full scFv sequence of six antibodies
that were expressed and tested for binding to the target.51 To perform clone selection at a
larger scale and/or on fully randomized libraries, however, the sequence information of
both VH and VL is needed. The Georgiou lab developed a yeast display method based
on short‑read sequencing but allowing the identication of full‑length antibodies.14 In
this method, the two antibody chains are cloned in opposite orientations separated by
a bidirectional promoter. Short‑read sequencing can hence cover the combination of
CDR3 sequences from both heavy and light chains (HCDR3‑LCDR3). By also sequenc‑
ing separately the full VH and VL genes, they can reconstruct bioinformatically the
entire paired antibody sequence. More recently, new technologies have been developed
which allow long‑read sequencing and the coverage of the full paired sequence from
an scFv or Fab library on a single read. PacBio sequencing, for example, has been suc‑
cessfully used for NGS‑based repertoire analysis and identication of binders of phage
display libraries.
35,53
4.5 NGS‑ENABLED IN SILICO
ANTIBODY DISCOVERY VIA ARTIFICIAL
INTELLIGENCE METHODS
Machine learning (ML) is a domain of articial intelligence (AI) that focuses on
‘learning’ the relationship between input and output variables and can be used for
both regression‑ and classication‑based tasks. Deep learning (DL), a branch of
ML science that has recently gained popularity, leverages interconnected neurons
arranged into various layers to both transform and process input data. Despite the
uptake of DL approaches in cancer screening and disease diagnosis,
commercial antibody discovery pipelines is relatively recent. NGS platforms provide
thousands of paired antibody sequences, which is now sufcient for training deep
learning‑based AI/ML models. Availability of large amounts of antibody sequence
data via NGS has also resulted in the rapid growth of various databases, in the public
domain, aimed at collecting and storing this information. For example, the Observed
Antibody Space (OAS) currently contains over a billion antibody sequences col‑
lected from over 80 studies and includes both paired and unpaired sequences.
databases like SAbDab58 and Thera‑SAbDab59 focus on collating structural data of
54–56
their use in
57
Other

4 • Analyses of Antibody Repertoires 69
public or therapeutic antibodies, respectively. It is the accessibility of well‑annotated
data alongside the almost exponential growth of new algorithms that has driven the
accelerated development of antibody‑specic ML tools.
Antibody 3D structure prediction is the domain which has beneted the most from
the advent of ML, and methods of antibody structures have evolved rapidly. Much of
this evolution is owed to the development of AlphaFold2,60 which has revolutionized
protein structure prediction. However, despite AlphaFold2’s remarkable accuracy, the
method still struggles with antibody structure prediction. AlphaFold2 requires a mul‑
tiple alignment step using various homologues, which is problematic in the context of
antibodies due to overall diversity of the HCDR3 region. This has led to the develop‑
ment of antibody‑specic protein predictors such as ABodyBuilder261 and DeepAb,62
both of which are DL approaches capable of accurately predicting the entire Fv mod‑
elling region,61 whereas ABlooper specializes in CDR loop prediction by leveraging
graph neural networks.63 Protein language models have also shown success in predicting
antibody structure and function by assigning probabilities to predicted amino acids or
‘tokens.’64 IgFold is an example of a structure prediction tool that leverages this type
of approach and utilizes a model trained on embeddings generated from AntiBERTy, a
natural antibody dataset consisting of over 500million sequences.65 Unlike ABlooper
and DeepAb, however, IgFold is able to make use of template structures, which results
in more accurate and faster predictions.66 Lastly, EquiFold attempts to represent anti‑
bodies as geometrical structures, avoiding the need for either protein language model
embeddings or multiple sequence alignments.67 Obtaining accurate predicted structural
representations is crucial for not only understanding the binding potential of candidate
antibodies but also their biophysical properties.68 Features such as solubility, hydro‑
phobicity, shelf‑life, viscosity, and immunogenicity have major roles in determining
whether an antibody is suitable for large‑scale manufacturing and commercialization.69
Indeed, poor developability in post‑Phase 1 antibodies can largely be attributed to a
small number of these features70; therefore, determining developability concerns early
in the development cycle is key. In‑silico screening tools like FreeSASA,71 PROPKA3,72
and TAP73 can leverage predicted structures to identify major developability concerns,
thus addressing these problems much earlier in development when compared with clas‑
sical approaches. Lastly, structure prediction can also help guide selection to a more
diverse set of candidates. Tools such as SPACE2leverage predicted structures to cluster
antibodies based on their predicted epitope of binding,74 thus allowing therapeutic can‑
didates to target a wider range of epitopes.
The modular nature of the methods discussed above ushers in the concept of a more
integrative approach to antibody discovery, where classic approaches are accompanied
and enhanced using various AI/ML tools (Figure4.4). For example, antigen‑specic
antibodies generated by the phage library could be supplemented by exploring the
nearby sequence space using generative models. Generative models offer signicant
benets to researchers, most notably a reduction in resources and time required for can‑
didate selection, and removal of any downstream developability concerns.75 ML models
trained on binder/non‑binder data can then subsequently be used to lter AI‑generated
sequences into those predicted to bind the same target antigen, with major develop‑
ability issues identied and removed. Indeed, a combinatorial approach using a suite

70 Biopharmaceutical Informatics
FIGURE4.4 Illustration depicting how machine learning approaches can augment antibodies identied by traditional discovery workows. The process begins with generating a
list of candidate antibodies through conventional methods, such as mouse immunization
or phage display (blue antibodies). These antibodies are then enhanced using generative
AI tools to explore the nearby sequence space, generating additional variants (red antibodies). Afnity prediction is applied to the candidate list to identify antibodies with improved
binding to the target antigen. Additionally, developability screening is performed to identify
and mitigate potential developability issues in both the original and AI-generated antibody
sequences, enhancing their suitability for therapeutic use.
of ML tools was recently demonstrated in the context of CTLA‑4‑ and PD‑1‑specic
antibodies.76 Heavy and light chain CDR3 sequences were rst used to train a con‑
volutional neural network (CNN) capable of predicting binder or non‑binder. In‑silico
mutagenesis was then carried out to determine antigen‑binding motifs, which in turn
allowed identication of sequences with a higher probability of binding. Lastly, genera‑
tive adversarial networks (GANs) capable of predicting entire HCDR3 sequences were
constructed, and the binding of the AI‑generated sequences was queried using the pre‑
vious CNN.76 Combinatorial platforms using structural data are also emerging, with a
recent example highlighting the power of DL in the context of reprogramming the bind‑
ing capacity of a SARS‑CoV‑2‑specic antibody.77 In this example, an attention‑based
geometric neural network was trained using 3D protein structures and asked to learn the
effect of single‑point CDR3mutations on binding afnity. This model was then applied
to a single SARS‑CoV‑2‑specic antibody to improve binding to the Delta variant of
SARS‑CoV‑2, with the observed increase in binding afnity owed to the removal of
various steric clashes between the antibody and Delta variant as determined by pre‑
dicted 3D structures.
77
Given the combinatorial potential and speed of development of AI/ML tools, a
fully in silico platform capable of identifying a set of potent and diverse candidates is
somewhat inevitable. Generative models can be used for candidate sequence generation,
structure prediction and clustering models can be used for binding afnity and epitope
prediction, and developability concerns can be removed from lead candidates. Despite
this, in vitro experiments will always be required to fully validate candidate antibod‑
25
As such, at least for the immediate future, a synergistic approach that combines
ies.
classic antibody development alongside AI/ML tools that aims to limit the amount of in
vitro work required is the most feasible approach.
Соседние файлы в папке Библиотека им академика М.И. Перельмана
