Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5629_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
15.09.2026
Размер:
16 Мб
Скачать
☆
4 • Analyses of Antibody Repertoires 61
but provided only a glimpse into the antibody diversity landscape, is truly remarkable. NGS technologies for antibodies vary signicantly in their applications and capabili‑
31,32
ties, as summarized in Table4.1 and previously described.
For instance, Roche’s 454 sequencing platform, although now discontinued, was once essential for its ability to produce long reads up to 700 bp, making it ideal for sequencing the entire vari‑ able (V) region of antibodies in a single run. On the other hand, the Illumina platform, which produces shorter reads of about 150 nucleotides, is particularly well‑suited for
TABLE4.1 NGS platforms for sequencing antibody repertoires from in vivo immunized animals and in vitro display libraries.
MAXIMUM
READ
PLATFORM
Illumina
MiSeq
Illumina
HiSeq
454 GS
FLX
Ion
Torrent
PacBio ~ 20 kb ~4million 10–30 hours ~11
The read lengths, number of reads per run, and run times can vary based on the specic model and conguration of the sequencing kit or chips used for each platform.
LENGTH
2 X 300 bp ~ 25million 4–55 hours ~0.1 Suitable for bulk
2 X 150 bp ~ 1.5 billion ~48 hours ~0.1 Suitable for bulk
~ 700 bp ~ 1million 10–24 hours ~1 Offered longer read
~ 400 bp ~80million 2–7 hours ~1 Suitable for medium
READS PER
RUN RUN TIME
ERROR
(%) REMARKS
sequencing of variable regions from heavy and light chains separately. Low error rate and high throughput.
sequencing of heavy and light chains separately. Low error rate and high throughput. Principle platform used for 10X Genomics sequencing, enabling barcoding of single cells for paired information such as CDR-H3 and CDR-L3 pairing.
lengths suitable for scFv libraries with sequencing of heavy and light chains separately but discontinued technology with limited availability.
length reads, offering a balance of speed and detail.
Long reads enable (raw), <1 (CCS)
full-length sequencing of heavy/light chains in single reads suitable for scFv or Fab libraries.
62 Biopharmaceutical Informatics
analyzing the heavy chain complementarity determining region 3 (HCDR3), a cru‑ cial segment in determining antibody specicity. Although the HCDR3 sequences are shorter, Illumina’s high throughput facilitates a more extensive examination. Emerging platforms like Pacic Biosciences (PacBio) offer the potential to sequence even longer reads from single DNA or RNA molecules, potentially simplifying the complex library preparations required for scFv and Fab libraries. However, their application in antibody repertoire analysis may be constrained by the quality of the sequences they produce. Meanwhile, platforms like Ion Torrent strike a balance with their moderate read lengths and faster sequencing capabilities at a lower cost, although their error rate remains a concern. The continued development and renement of these technologies, including error correction strategies such as the use of unique molecular identiers (UMIs), is crucial as they evolve to meet the high standards required for precise and effective antibody discovery.
Once the raw data from NGS sequencing runs are available, bioinformatics tools are utilized to process and analyze the data, facilitating the antibody discovery pro‑ cess from understanding diversity to NGS‑based screening and selection of functional clones. A selected list of bioinformatics software tools for analyzing antibody repertoire data generated from NGS is presented in Table4.2. Initially, the process involves data cleaning, which includes removing low‑quality sequences and sequencing adapters to ensure data integrity. Subsequent steps involve using tools such as IMGT/V‑QUEST and IgBLAST to identify the variable (V), diversity (D), and joining (J) gene segments, and to dene the framework regions and complementarity determining regions (CDRs) essential for antigen‑binding specicity. The selection of these tools often depends on the volume of data and the user’s expertise. For those less familiar with computational
TABLE4.2 Selected bioinformatics software tools for NGS-generated antibody repertoire data analysis.
TOOLS DESCRIPTION URL
IgBLAST V(D)J annotation, CDRs assignment, and SHM https://www.ncbi.nlm.
nih.gov/igblast/
IMGT/
HighV-QUEST
MiXCR Raw sequences to clonotypes https://github.com/
PipeBio User-friendly end-to-end workows with
ENPICOM IGX Platform, data visualization and analysis, and
Geneious
Biologics
IgBLAST is freely available for use, while the other tools listed require licensing and involve costs.
A pioneering international information system in
immunogenetics and immunoinformatics with tools available for V(D)J annotation, CDRs assignment, SHM, statistical analysis, and plots
visualization for antibody discovery, NGS analysis, and antibody engineering
AI-aided antibody development
Powerful tools for antibody sequence annotation
and analysis, statistical analysis, and visualization
https://www.imgt.org/
milaboratory/mixcr/
https://pipebio.com/
https://enpicom.com/
https://www.geneious.
com/biopharma/
4 • Analyses of Antibody Repertoires 63
techniques, platforms like IMGT High‑VQUEST offer user‑friendly interfaces, while more sophisticated tools like MiXCR provide detailed immunological analyses for
33
advanced users. Furthermore, clustering tools such as CD‑hit
or UCLUST34 are employed to group similar sequences, aiding in the identication of clonal expansions and enriching our understanding of immune response diversity. For handling larger datasets and facilitating collaborative research, cloud‑based commercial software plat‑ forms like PipeBio, ENPICOM, and Geneious Biologics provide comprehensive, inte‑ grated solutions that streamline the workow and enhance data analysis capabilities. These platforms combine powerful bioinformatics tools with user‑friendly interfaces, signicantly simplifying the management and analysis of complex data, thereby accel‑ erating antibody research and development.
After performing antibody NGS data processing and analysis, several metrics and analytical techniques are employed to thoroughly examine the B‑cell repertoire. Figure4.1 illustrates various NGS‑based antibody repertoire metrics, such as isotype analysis, V(D)J segment usage frequencies, CDR3 properties, somatic hypermutation (SHM) analysis, and clonal relationship and lineage analysis, which are depicted using trees or network graphs. Statistical analysis is also crucial to ensure sequencing depth is comparable between samples, estimate repertoire diversity, and assess convergence. By employing this comprehensive approach, researchers can gain deep insights into the antibody response, including the types of antibodies produced, namely isotypes, the genes utilized for antibody building blocks comprising V(D)J gene segments, the CDRs, the extent of antibody diversication through mutations, and the relationships between different antibody clones. This information is invaluable for understanding the immune system’s response to pathogens and for developing novel antibody‑based therapies.
FIGURE 4.1 NGS antibody repertoire data analysis employs a multifaceted approach to dene key repertoire features, serving as essential metrics for NGS-driven antibody dis­covery. This analysis includes isotype analysis to categorize sequences by constant regions inuencing effector functions, and V(D)J usage analysis to assess gene segment frequen­cies, revealing specic germline contributions. Complementarity determining regions (CDRs) analysis, particularly focusing on CDR3, examines sequence length and amino acid proper­ties crucial for antigen binding. Somatic hypermutation (SHM) analysis explores the diver­sity of antibody clones, a critical process for generating high-afnity antibodies. Clonal relationship and lineage analysis track antibody evolution, offering insights into immune response development. Lastly, with paired sequencing data, chain pairing analysis evalu­ates the pairing of heavy (VH) and light (VL) chain genes, enhancing our understanding of antibody specicity.
64 Biopharmaceutical Informatics
In the case of NGS‑based antibody discovery using display libraries, the depth and breadth of analysis that can be performed on the outputs from selections of these librar‑ ies have dramatically increased with the advent of NGS. usage frequency, along with other immunogenetic features, such as specic gene lin‑ eages, CDR lengths, and amino acid patterns, can be analyzed to identify trends associ‑ ated with different panning strategies.
11,35
Additionally, data on gene
4.3 NGS‑ENABLED IN VIVO ANTIBODY
DISCOVERY FROM IMMUNIZED ANIMALS
In vivo antibody discovery using a hybridoma‑based approach was established several decades ago and has subsequently undergone several improvements. involves immunizing animals with a specic antigen, followed by the laborious process of fusing B‑cells from the animal’s spleen with myeloma cells, creating immortalized hybridomas. The generated hybridomas continuously produce antibodies specic to the immunizing antigen. However, hybridoma technology suffers from several limitations. Firstly, it captures only a tiny fraction of the vast diversity of antibodies generated by the immune system. Secondly, the process is highly inefcient and time‑consuming, requiring extensive screening of individual hybridoma clones to identify those produc‑ ing high‑afnity antibodies. The emergence of NGS technologies has revolutionized in vivo antibody discovery, offering a powerful and multifaceted approach.39 This revolu‑ tion extends far beyond traditional hybridoma technology by incorporating advanced technologies that work synergistically with NGS analysis. Particularly, human immu‑ noglobulin transgenic mice offer a more human‑like antibody repertoire for NGS sequencing. Furthermore, NGS combined with single B‑cell antibody technologies, antigen‑specic single B‑cell sorting, and recombinant antibody cloning have become powerful approaches. from B‑cell antibody genes, preserve the crucial pairing of heavy and light chain genes for proper antibody function. NGS analysis of these libraries allows for the identica‑ tion of a vast repertoire of diverse and functional antibodies. Figure4.2 outlines the NGS‑driven in vivo antibody discovery process, starting with B‑cell isolation and proceeding through sequencing and comprehensive repertoire analysis. This includes examining V(D)J gene usage, SHM diversity, clonal expansion, and antibody lineage relationships. The method facilitates the identication of highly diverse, high‑afnity antibodies that are not detectable by traditional in vivo hybridoma methods.
In one of the use‑cases as adopted by researchers at Genentech, integrating NGS into hybridoma technology has signicantly advanced the antibody discovery process.42 This approach overcomes traditional limitations related to throughput and storage capacity by digitizing antibody variable domain sequences for high‑throughput screen‑ ing directly from hybridoma cells. The use of barcoded primers in a 96‑well format allows for the simultaneous processing of multiple samples, enhancing sequencing out‑ put and efciency substantially. Additionally, robust bioinformatics tools ensure precise
40,41
Additionally, natively paired immune libraries, constructed
36–38
This method
4 • Analyses of Antibody Repertoires 65
FIGURE 4.2 NGS-enabled in vivo antibody discovery process starts with isolating bulk or single B-cells from a model organism and utilizes NGS for in-depth repertoire analysis. This analysis investigates V(D)J gene usage, explores SHM diversity critical for high-afnity antibodies, examines clonal expansion of desirable B-cells, and elucidates antibody lineage relationships. This comprehensive NGS approach allows the discovery and characterization of high-afnity antibodies in vivo.
resolution and unambiguous identication of antibody sequences. This method dramati‑ cally reduces the need for physical storage by converting sequences into digital data, offering a scalable, cost‑effective solution that aligns perfectly with the needs of modern biomedical research and drug development. Additionally, it facilitates a more dynamic, rapid, and expansive exploration of potential therapeutic antibodies, demonstrating a critical advancement in the eld.
In another use‑case study, George Georgiou and colleagues developed a ground‑ breaking method that bypasses the traditional high‑throughput screening processes for isolating antigen‑specic monoclonal antibodies (mAbs).10 By leveraging NGS and bioinformatic analysis, this technique directly mines the antibody variable region (V)‑gene repertoires from bone marrow plasma cells (BMPCs) of immunized mice. BMPCs, notable for producing most circulating antibodies but are unable to be immor‑ talized, exhibit a highly polarized V‑gene repertoire following immunization. The most abundant variable heavy (VH) and variable light (VL) genes are paired based on their relative frequencies, reconstructed through automated gene synthesis, and expressed in bacterial or mammalian systems as recombinant antibodies. This method signicantly accelerates the progression from immunization to specic antibody production and yields antibodies with high antigen specicity and nanomolar afnities, underscoring its potential as a tool to rapidly respond to emerging infectious diseases and advancing both immunological research and therapeutic antibody development.
In another case study by GigaGen Inc., the authors leveraged NGS to signicantly advance the discovery of therapeutic monoclonal antibodies by deeply sequencing yeast scFv libraries derived from immunized Trianni mice, both before and after uores‑
43
cence‑activated B‑cell sorting (FACS).
This approach allowed the authors to track the evolution and enrichment of antibody clones, showcasing the technology’s ability to narrow down thousands of diverse clones to a few with potential therapeutic properties. The sequence analysis highlighted minimal overlap in antibody sequences enriched by
66 Biopharmaceutical Informatics
different methods, pointing to the inuence of antigen presentation on immune response specicity. Remarkably, sequences derived from soluble immunogen methods showed more similarity across FACS methods compared to those from cell/DNA methods, suggesting differences in immune response elicitation. This study exemplies how integrating NGS with innovative immunization strategies can streamline the antibody discovery process, enhance the functional relevance of identied antibodies, and expe‑ dite the development of new therapeutics.
Overall, NGS‑based in vivo antibody discovery offers a revolutionary approach, merging the strengths of animal immunization with advanced screening methods. This opens the door for exploring a vast universe of possibilities within the animal‑derived antibody repertoire.
4.4 NGS‑ENABLED IN VITRO ANTIBODY DISCOVERY FROM DISPLAY LIBRARIES
In vitro antibody library platforms such as phage and yeast display have become power‑ ful tools for therapeutic antibody discovery, allowing the identication of fully human antibodies against virtually any desired target. advantages as they provide: (i) large antibody sequence and structural diversities, typi‑ cally up to 1011 unique clones, (ii) fast and easy selection methods able to apply specic enrichment pressures on the entire library, such as incubation with solid‑surface immo‑ bilized antigen or FACS, and (iii) rapid identication of binder antibodies by providing a direct phenotype‑genotype link. Traditional screening methods of in vitro display libraries and their enrichment outputs, however, can be time and resource intensive. These methods rely on the isolation of single colonies for various binding or functional assays and subsequent Sanger sequencing, thus limiting the number of tested antibodies to typically a few hundred. Additionally, most of the clones picked for characteriza‑ tion in this way represent antibodies present at the highest frequency in the enriched population. This inevitably means that clones picked at random have a large level of redundancy, with those high frequency clones being picked multiple times, reducing even further the number of tested unique antibodies. Furthermore, this dominance of high frequency clones does not represent only antibodies that have been enriched due to high‑afnity antigen binding but also those that might have been enriched due to higher expression or display of the antibodies by the phage system. The presence of such clones at high frequency will often lead to the failure of the traditional screening methods in the detection of rarer but more desirable and functional antibodies.
More recently, the combination of NGS and in vitro display technologies has largely removed the limitations of traditional screening methods and made the use of these platforms in therapeutic antibody discovery even more powerful.32 NGS allows deep repertoire proling and mining of antibody libraries and their enrichment outputs, providing a more holistic view of the population dynamics and the evolution of the antibody repertoires through consecutive enrichment rounds. As detailed in Figure4.3,
7,4 4–47
These platforms have several
48
4 • Analyses of Antibody Repertoires 67
FIGURE4.3 NGS-enabled in vitro antibody discovery process begins with the panning of phage-displayed libraries across multiple rounds, typically three rounds (R1, R2, R3), each targeting a specic antigen. Post-panning, phage particles from each round undergo NGS and antibody repertoire analysis. This analysis includes evaluating the frequency distribu­tion of sequences from each round, clustering to identify related antibody sequences and constructing phylogenetic trees to track the evolution and expansion of antibody hits. This comprehensive approach facilitates the identication and optimization of high-afnity anti­bodies in vitro.
the NGS‑enabled in vitro discovery process utilizes iterative panning of phage libraries against a target antigen, followed by NGS analysis, for example, frequency distribution, sequence clustering, and phylogenetic trees, to identify and optimize high‑afnity anti‑ body candidates. With NGS, not only the high frequency enriched clones but also those that are rarer with lower frequencies become visible, resulting in a more diverse panel of identied antibody hits.11 Moreover, as almost all enriched clones are detected, larger clusters of enriched sequences (families of sequences closely related in sequence space) can be identied. Therefore, selection of representative clones from these clusters can be made in a more intelligent manner, addressing, for example, developability criteria of the selected clones early in the discovery phase.49 Moreover, having access to these multi‑member sequence clusters presents the possibility of “hit expansion.” This is the strategy where variants of a known functional antibody sequence are identied in NGS data with the objective of improving some property of the parental antibody, that is, afnity, developability, cross‑reactivity, etc.
50
First studies of incorporating NGS in phage display antibody discovery, due to the limited capacity of the sequencing length of sequencing technologies existing at the time, focused only on HCDR3 or the heavy chain variable region (VH) sequences.
51,52
68 Biopharmaceutical Informatics
The short‑read sequencing approaches have the advantage of offering a very high sequencing depth and are very well‑suited for the global analysis of library quality, for example, looking at the diversity of antibody chains separately. Several groups have also used these sequencing methods for NGS‑based clone selection, using clone rescue meth‑ ods to recover the full antibody sequence. However, these studies have been limited to libraries with variability introduced only in specic segments of the variable gene or to the selection of only a handful of clones for expression and characterization.48 The Fischer lab, for example, used HCDR3 sequences with high enrichment ratio obtained from NGS to design oligos and PCR‑amplify the full scFv sequence of six antibodies that were expressed and tested for binding to the target.51 To perform clone selection at a larger scale and/or on fully randomized libraries, however, the sequence information of both VH and VL is needed. The Georgiou lab developed a yeast display method based on short‑read sequencing but allowing the identication of full‑length antibodies.14 In this method, the two antibody chains are cloned in opposite orientations separated by a bidirectional promoter. Short‑read sequencing can hence cover the combination of CDR3 sequences from both heavy and light chains (HCDR3‑LCDR3). By also sequenc‑ ing separately the full VH and VL genes, they can reconstruct bioinformatically the entire paired antibody sequence. More recently, new technologies have been developed which allow long‑read sequencing and the coverage of the full paired sequence from an scFv or Fab library on a single read. PacBio sequencing, for example, has been suc‑ cessfully used for NGS‑based repertoire analysis and identication of binders of phage display libraries.
35,53
4.5 NGS‑ENABLED IN SILICO
ANTIBODY DISCOVERY VIA ARTIFICIAL
INTELLIGENCE METHODS
Machine learning (ML) is a domain of articial intelligence (AI) that focuses on ‘learning’ the relationship between input and output variables and can be used for both regression‑ and classication‑based tasks. Deep learning (DL), a branch of ML science that has recently gained popularity, leverages interconnected neurons arranged into various layers to both transform and process input data. Despite the uptake of DL approaches in cancer screening and disease diagnosis, commercial antibody discovery pipelines is relatively recent. NGS platforms provide thousands of paired antibody sequences, which is now sufcient for training deep learning‑based AI/ML models. Availability of large amounts of antibody sequence data via NGS has also resulted in the rapid growth of various databases, in the public domain, aimed at collecting and storing this information. For example, the Observed Antibody Space (OAS) currently contains over a billion antibody sequences col‑ lected from over 80 studies and includes both paired and unpaired sequences. databases like SAbDab58 and Thera‑SAbDab59 focus on collating structural data of
54–56
their use in
57
Other
4 • Analyses of Antibody Repertoires 69
public or therapeutic antibodies, respectively. It is the accessibility of well‑annotated data alongside the almost exponential growth of new algorithms that has driven the accelerated development of antibody‑specic ML tools.
Antibody 3D structure prediction is the domain which has beneted the most from the advent of ML, and methods of antibody structures have evolved rapidly. Much of this evolution is owed to the development of AlphaFold2,60 which has revolutionized protein structure prediction. However, despite AlphaFold2’s remarkable accuracy, the method still struggles with antibody structure prediction. AlphaFold2 requires a mul‑ tiple alignment step using various homologues, which is problematic in the context of antibodies due to overall diversity of the HCDR3 region. This has led to the develop‑ ment of antibody‑specic protein predictors such as ABodyBuilder261 and DeepAb,62 both of which are DL approaches capable of accurately predicting the entire Fv mod‑ elling region,61 whereas ABlooper specializes in CDR loop prediction by leveraging graph neural networks.63 Protein language models have also shown success in predicting antibody structure and function by assigning probabilities to predicted amino acids or ‘tokens.’64 IgFold is an example of a structure prediction tool that leverages this type of approach and utilizes a model trained on embeddings generated from AntiBERTy, a natural antibody dataset consisting of over 500million sequences.65 Unlike ABlooper and DeepAb, however, IgFold is able to make use of template structures, which results in more accurate and faster predictions.66 Lastly, EquiFold attempts to represent anti‑ bodies as geometrical structures, avoiding the need for either protein language model embeddings or multiple sequence alignments.67 Obtaining accurate predicted structural representations is crucial for not only understanding the binding potential of candidate antibodies but also their biophysical properties.68 Features such as solubility, hydro‑ phobicity, shelf‑life, viscosity, and immunogenicity have major roles in determining whether an antibody is suitable for large‑scale manufacturing and commercialization.69 Indeed, poor developability in post‑Phase 1 antibodies can largely be attributed to a small number of these features70; therefore, determining developability concerns early in the development cycle is key. In‑silico screening tools like FreeSASA,71 PROPKA3,72 and TAP73 can leverage predicted structures to identify major developability concerns, thus addressing these problems much earlier in development when compared with clas‑ sical approaches. Lastly, structure prediction can also help guide selection to a more diverse set of candidates. Tools such as SPACE2leverage predicted structures to cluster antibodies based on their predicted epitope of binding,74 thus allowing therapeutic can‑ didates to target a wider range of epitopes.
The modular nature of the methods discussed above ushers in the concept of a more integrative approach to antibody discovery, where classic approaches are accompanied and enhanced using various AI/ML tools (Figure4.4). For example, antigen‑specic antibodies generated by the phage library could be supplemented by exploring the nearby sequence space using generative models. Generative models offer signicant benets to researchers, most notably a reduction in resources and time required for can‑ didate selection, and removal of any downstream developability concerns.75 ML models trained on binder/non‑binder data can then subsequently be used to lter AI‑generated sequences into those predicted to bind the same target antigen, with major develop‑ ability issues identied and removed. Indeed, a combinatorial approach using a suite
70 Biopharmaceutical Informatics
FIGURE4.4 Illustration depicting how machine learning approaches can augment anti­bodies identied by traditional discovery workows. The process begins with generating a list of candidate antibodies through conventional methods, such as mouse immunization or phage display (blue antibodies). These antibodies are then enhanced using generative AI tools to explore the nearby sequence space, generating additional variants (red antibod­ies). Afnity prediction is applied to the candidate list to identify antibodies with improved binding to the target antigen. Additionally, developability screening is performed to identify and mitigate potential developability issues in both the original and AI-generated antibody sequences, enhancing their suitability for therapeutic use.
of ML tools was recently demonstrated in the context of CTLA‑4‑ and PD‑1‑specic antibodies.76 Heavy and light chain CDR3 sequences were rst used to train a con‑ volutional neural network (CNN) capable of predicting binder or non‑binder. In‑silico mutagenesis was then carried out to determine antigen‑binding motifs, which in turn allowed identication of sequences with a higher probability of binding. Lastly, genera‑ tive adversarial networks (GANs) capable of predicting entire HCDR3 sequences were constructed, and the binding of the AI‑generated sequences was queried using the pre‑ vious CNN.76 Combinatorial platforms using structural data are also emerging, with a recent example highlighting the power of DL in the context of reprogramming the bind‑ ing capacity of a SARS‑CoV‑2‑specic antibody.77 In this example, an attention‑based geometric neural network was trained using 3D protein structures and asked to learn the effect of single‑point CDR3mutations on binding afnity. This model was then applied to a single SARS‑CoV‑2‑specic antibody to improve binding to the Delta variant of SARS‑CoV‑2, with the observed increase in binding afnity owed to the removal of various steric clashes between the antibody and Delta variant as determined by pre‑ dicted 3D structures.
77
Given the combinatorial potential and speed of development of AI/ML tools, a fully in silico platform capable of identifying a set of potent and diverse candidates is somewhat inevitable. Generative models can be used for candidate sequence generation, structure prediction and clustering models can be used for binding afnity and epitope prediction, and developability concerns can be removed from lead candidates. Despite this, in vitro experiments will always be required to fully validate candidate antibod‑
25
As such, at least for the immediate future, a synergistic approach that combines
ies. classic antibody development alongside AI/ML tools that aims to limit the amount of in vitro work required is the most feasible approach.