Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5908_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Contents
- •Foreword
- •Preface
- •About the Editors
- •Contributors
- •References
- •2.3.4 Barriers to Automation Adoption
- •2.4 Core Ingredients for Successful Digital Transformation
- •2.1 Introduction
- •2.3.1 Operational Challenges
- •2.3.2 Cultural Challenges
- •2.4.2 Cloud Computing
- •2.5 Case Studies of Successful Digital Transformation
- •2.6 Conclusion
- •References
- •3. Computational Protein Design Strategies for Optimization of Antigen Generation to Drive Antibody Discovery
- •3.1 Introduction
- •3.3 Antigen Generation Strategies
- •3.4 Computational Methods
- •3.4.2 Computational Protein Structure Prediction
- •References
- •4. Bioinformatic Analyses of Antibody Repertoires and Their Roles in Modern Antibody Drug Discovery
- •4.1 Introduction
- •4.6 Summary and Future Directions
- •Acknowledgments
- •References
- •5.1 Introduction
- •5.2 Databases
- •5.2.1 Databases in Machine Learning Approaches
- •5.2.2 Database Types
- •5.3 Applications of Machine Learning in Antibody Discovery and Development
- •5.3.1 Structure Prediction with Deep Learning
- •5.3.3 Developability
- •5.4 Antibody Generation and Design by Language Models
- •5.4.1 Antibody Representations
- •5.4.2 Representation Learning
- •5.4.3 Language Models
- •References
- •6.1 Introduction
- •6.2 Antibody Generation through Deep Generative Models
- •6.3.1 Sampling and Scoring
- •6.5 Conclusions and Perspectives
- •Acknowledgments
- •References
- •7.1 Introduction
- •7.2.3 Computational Approaches to Predict Antibody–Antigen Interaction
- •7.3 Conclusion
- •Competing Interests
- •Acknowledgments
- •References
- •8.2 Common Types of Molecular Simulations for Biomolecules
- •8.2.1 Molecular Dynamics (MD) Simulations
- •8.2.2 Monte Carlo (MC) Simulations
- •8.2.3 Challenges of Molecular Simulations
- •8.3.1 Periodic Boundary Conditions
- •8.4 Uses of Molecular Simulation in Antibody Drug Development
- •8.5 Conclusion
- •References
- •9. Considerations of Developability During the Early Stages of Antibody Drug Discovery and Design
- •9.1 Introduction
- •9.2 Historical Perspective
- •9.3 Clinical Antibody Data Set
- •9.5 Control Antibodies
- •9.7 Assessment of Chemical Liabilities
- •9.8 Conclusions and Future Perspectives
- •Acknowledgments
- •References
- •Abbreviations
- •10.1 Introduction
- •10.4.1 Conclusions and Outlook
- •Acknowledgments
- •References
- •11.8 Conclusions and Future Directions
- •References
- •12.1 Introduction to PK/PD and QSP Modeling
- •12.1.1 PK/PD Modeling
- •12.1.2 QSP Modeling
- •12.2.1 Monoclonal Antibodies (mAbs)
- •12.2.3 Cell Therapies
- •12.2.4 Gene Therapies
- •12.2.5 Vaccines
- •12.2.6 mRNA/siRNA/Oligonucleotide Therapeutics
- •12.4 Case Studies
- •12.5 Conclusions and Future Perspectives
- •References
- •13.1 Introduction
- •13.2 AI/ML: A Game Changer for Antibody Design
- •13.3 Multispecific Antibody Design
- •13.4 Adapting AI to the Design of Multispecific Antibodies
- •13.4.1 Structure Prediction and Modeling
- •13.4.2 Developability Prediction and Optimization
- •13.4.4 In Silico Modeling and Simulation
- •13.5 The Future: Beyond Optimization
- •13.5.1 Market Trends and Commercialization
- •13.5.2 Logic Gates, Biosensors, and De Novo Design
- •13.5.3 Challenges and Opportunities
- •13.6 Conclusion
- •Acknowledgments
- •References
- •Index

3 • Computational Protein Design Strategies 41
channel
Type I membrane
proteinECD
Ligand gatedion
FIGURE3.1 Structures illustrating diversity and complexity of protein targets of potential
therapeutic interest. Each protein is represented in cartoon format generated in Pymol.
The cytokine, a soluble secreted protein, is shown here bound to a Fab antibody fragment.
Ligand-gated ion channel and GPCR are integral multi-pass membrane proteins. The type I
membrane protein has a single-pass transmembrane region (not shown here) and an extracellular domain (ECD) extending beyond the lipid bilayer.
Cytokine
GPCR
Fab
Lipid
bilayer
bispecic antibodies, radioimmunoconjugates, and single‑domain antibodies, extending
the potential clinical applications [2]. Antibody discovery is also able to isolate antibod‑
ies that bind a wide range of target proteins (Figure3.1), which include soluble mediators
(e.g. TNF‑alpha, IL‑17A), receptors (e.g. epidermal growth factor receptor), viral proteins
(e.g. SARS‑CoV‑2 spike protein), transmembrane proteins (e.g. PD‑1, PD‑L1), and complex
multi‑spanning membrane proteins (e.g. CD20, CGRP receptor) (Figure3.1) [3–8]. These
target proteins, once identied as a valid therapeutic target, are used as antigens to drive the
antibody discovery process to identify potential lead antibodies.
Antibody discovery uses either in vivo or in vitro strategies (Figure3.2) to create
a source of diverse candidate antibodies from which potential lead antibodies can be
isolated [9]. In vivo strategies require the immunization of animal hosts with a target
antigen to generate an immune response. This has typically involved the immunization
of mice (wild type or humanized), followed by the fusion of immune cells with myeloma
cells to generate hybridomas. Alternative hosts such as camelids or chickens may also be
used [10]. In these cases, antibodies may be isolated by the generation of immune dis‑
play libraries or B‑cell screening [11,12]. In vitro strategies involve generation of display
libraries (e.g. phage display), which can be used to select encoded antibodies that bind to
a target antigen presented to the library [13–15].
To support and drive antibody discovery, an essential aspect is the production of the
protein (antigen) targets that are used as immunogens or targets in the application of dis‑
play technologies. This aspect of antibody discovery requires bioinformatic analysis of
the target antigen at the sequence and structural levels, methods to produce high‑quality
target protein antigen, and formulations that enable antibody–drug discovery. This eld
has always relied on the application of databases and computational tools, but there is
signicant potential to enhance this area by the application of articial intelligence (AI)

42 Biopharmaceutical Informatics
Created with BioRender.com
FIGURE3. 2 Strategies used to isolate therapeutic antibodies. Antibodies may be isolated by in vivo or in vitro methods. Immunizations are performed in a chosen host and lymphoid tissues isolated. For mice, hybridomas are made by fusing immune cells with myeloma
followed by screening for antibodies binding to target antigen. Alternatively, B-cells can be
screened directly using various microuidic platforms. Phage display can be performed using
phage libraries generated from immune cells derived from immunized hosts or via entirely
synthetic libraries. The phage are exposed to target antigen (shown here immobilized on a
surface) and binders identied. Antibody sequences can then be produced recombinantly by
transient or stable cell expression followed by purication of the antibody chosen.
and machine learning (ML). In this article, methods for antigen design and production
are discussed, and the potential of computational tools to enhance this area is explored.
3.2 TARGET PROTEIN (ANTIGEN)
CONSIDERATIONS FOR ANTIBODY DISCOVERY
Protein targets that provide potential avenues for therapeutic antibody discovery may
be identied by various routes. Approaches to this include deep understanding of the
underlying biology of diseases and the proteins involved, genomics and proteomic
methods to identify disease‑relevant targets, and target‑agnostic approaches such as
phenotypic screening [16]. More recently, AI‑driven approaches are being used to iden‑
tify potential targets by mining and analysing data from various sources [17]. In the
case of therapeutic antibody discovery, the target protein identied must be accessible
to the antibody, which focuses attention on extracellular proteins secreted from cells or
membrane proteins with accessible potential epitopes on the cell surface.

3 • Computational Protein Design Strategies 43
Once a target protein is selected, a key step in preparing a strategy for antibody dis‑
covery is the bioinformatic analysis of the target protein. This analysis is performed to
dene the optimal antigen design to facilitate antibody discovery according to the selected
antibody lead identication technology. Proteins consist of one or several compact, auton‑
omously folding substructures referred to as domains. Protein targets may also be complex
in nature. For example, a receptor may function as a homo‑ or heterodimer or higher order
complex. Once a drug discovery target is selected, several databases and computational
tools are used to understand the structure–function relationships of proteins and their
expression patterns. Thus, analysis is performed to understand the sequence and structure
of the protein, prevalence of alternative forms such as splice variants, single‑nucleotide
polymorphisms (SNPs), post‑translational modication, domain structure(s), and potential
interaction partners (Figure3.3). A wide range of databases and computational resources
are available for this purpose: Protein features and sequences can be explored using UniProt
[18] and Ensembl [19]; solved protein structures can be obtained from the Protein Data
Bank (PDB) [20]; domains and boundaries can be predicted using tools based on homol‑
ogy, structure, and ab initio methods (reviewed by Wang etal. [21]) and more recent meth‑
ods such as ResDom [22]. Protein haplotypes also need to be considered, as a single amino
acid change in the target protein can impact the binding of an antibody, which would impact
therapeutic efcacy. This can be evaluated using Haplosaurus, a tool available in Ensembl,
which reects real‑world protein sequence variability and prevalence in populations [23].
A useful online resource describing and providing access to general tools in this area
is the Protein Structural Bioinformatics Overview (PreStO) web tool [24]. In the case
of membrane proteins, useful resources are the Orientations of Proteins in Membranes
(OPM) database and the Positioning of Proteins in Membranes (PPM2.0 and PPM3.0)
server [25,26]. This curated web resource provides information on the spatial positions of
membrane proteins, topology, and membrane protein types, while PPM calculates the spa‑
tial position of proteins in membranes. This resource also provides protein images, struc‑
tures, and visualization tools. There are also useful resources aimed at specic membrane
Antigen considerations
• Sequence (Commonvariant)
• Speciesvariants
• Splice forms, SNP’s, Haplotypes
• Post-translational modifications
• Domain structure
• Potentialepitopes
Bioinformaticand Structural Databases
• Uniprot
• Ensembl
• ProteinDatabank(PDB)
Antigen
HER2 ECD
(Domains I-IV)
Fab
Pertuzumab
(DomainII binder)
Fab
Trastuzumab
(DomainIVbinder)
FIGURE 3.3 Structural and sequence considerations for protein antigen generation.
Cryo-EM structure of HER-2–trastuzumab–pertuzumab complex (PDB ID: 60GE) demonstrates that proteins can be targeted by antibodies binding distinct epitopes. For each target
antigen, key considerations that must be addressed to design an effective antigen are listed
along with some key databases.

44 Biopharmaceutical Informatics
Super-4
protein classes, including G protein‑coupled receptors (GPCRs) (GPCRdb) and protein
channels (ChannelsDB2.0) [27–29]. These databases are subject to regular revisions, with
recent releases being linked to AlphaFold.
In addition to understanding the structure and sequence variants of the human target
protein antigen, it is important to analyse sequence identity relationships to species variants of
the target (e.g. mouse, non‑human primate) and if there are any closely related human ortho‑
logues and paralogues. The reason for performing this analysis is to get a good understanding
of aspects of the target which may impact a therapeutic antibody campaign and to understand
and mitigate potential risks of off‑target cross reactivity. It is important to understand if there
are any closely related family members that could result in undesirable toxicities and should be
included in screens to select antibodies that only bind the intended target proteins. An exam‑
ple of this was recently described for the membrane protein Claudin 6 (CLDN6), which is a
potential oncotherapeutic ta rget with high expression levels in solid tumours. CLDN6 is highly
similar to other CLDN family members, with only three amino acid differences in extracel‑
lular loops to CLDN9, which is expressed in healthy tissue (Figure3.4a) [30]. In identifying
an antibody that targets CLDN6, it is necessary to avoid cross reactivity to CLDN9, which
could drive toxicity. For therapeutic antibody discovery, the primary focus is on the human
target–however, it is important to understand the relationship to species variants to under‑
stand if it is feasible to identify an antibody binding both humans and species variants (e.g.
IL-4 recepto
Fab
A
FIGURE3. 4 (a) Structure of Claudin-9 (PDB: 60V2). The transmembrane helical region and
intracellular region are shown as cyan cartoon, while the extracellular loops are shown in
magenta. Red space-lled amino acids show the position of residues that differ between
Claudin-9 and Claudin-6. An antibody binding selectively to Claudin-6 was isolated (see
text), with the lack of Claudin-9 cross-reactivity being driven by a steric block from residue
156. (b). Structure of ‘Stapler’ antibody isolated by phage display against a stabilized ligand–
receptor complex comprising IL-4, IL-4 receptor, and y-chain. The antibody recognizes an
interface formed by the juxtaposition of IL-4 receptor and the γ-chain (PDB: 3BPL).
Claudin-9
B
r
γ-chain

3 • Computational Protein Design Strategies 45
mouse, non‑human primate) to facilitate in vivo translational studies. This is also important
in selecting the antibody discovery strategy–if a protein shares a very high (>95%) sequence
homology with a mouse counterpart, an in vivo antibody discovery campaign may require
the selection of a more divergent species host for immunizations (e.g. chicken) to avoid issues
with tolerance. This is exemplied by the choice of immunisation strategy used for the target
CLDN6, where the high (95%) sequence identity with the mouse homologue led to selec‑
tion of chickens as the host for immunization to identify antibodies binding to CLDN6 [30].
Using this approach, high‑afnity binders against CLDN6 were isolated, and antibodies
were isolated, which showed minimal to no cross reactivity with CLDN9 and 22 other clau‑
din family members.
3.3 ANTIGEN GENERATION STRATEGIES
The discovery of therapeutics based on antibodies and related molecules, such as
single‑domain antibodies (Vhh, nanobodies), requires the selection of an appropriate
discovery platform. A diverse antibody binding panel can be derived from platforms
such as hybridoma, single B‑cell methods, or screening natural or synthetic antibody
libraries via display technologies using phage, yeast, or mammalian display [13,31–33].
In vivo platforms requiring immunization are becoming increasingly diverse, with
options for using laboratory mice, chickens, camelids, rabbits, and an expanding range
of humanized/transgenic hosts such as the ATX‑Gx™ mouse and OmniChicken
[7,34]. In all cases, a critical element is the production of the target protein antigen in
sufcient quantity and quality to enable immunization, in vitro selection of antibodies
binding to antigen, and functional activity assessment. This allows for the identication
of lead antibodies that can be used as therapeutics or further engineered to enhance
properties such as binding afnity and developability.
Protein antigens need to be produced and puried in a format suitable for antibody
discovery via the chosen in vivo or in vi tro platform. The simplest option for antigen genera‑
tion is the use of synthetic peptides designed to represent the surface‑accessible epitopes
[35–38]. This approach directs the antibody response to a very specic epitope contain‑
ing the peptide and has led to marketed biologics against a GPCR target, CCR4 [36].
It is more common for therapeutic antibody discovery to use intact protein antigen, as
this will present both linear and conformational epitopes and is more physiologically
relevant. Protein is typically produced recombinantly by designing expression plasmids
or vectors that allow protein expression and purication from various hosts [39,40].
Selection of an appropriate expression host is somewhat empirical, with no single sys‑
tem being optimal. However, expression hosts that most closely resemble the native
host are more likely to produce high‑quality recombinant protein due to similar folding
and trafcking machinery, cofactors, and post‑translational modication pathways. The
host selection will also depend on the type of protein being expressed. Escherichia coli
remains very attractive as a host for the production of proteins, particularly if the tar‑
get protein does not require post‑translational modication [39]. Integral membrane
®

46 Biopharmaceutical Informatics
proteins will typically require eukaryotic expression systems such as Human embryonic
kidney (HEK) or Chinese hamster ovary (CHO) cells [41]. These expression hosts can
be used as a source of recombinant target protein that can be extracted and puried (e.g.
by detergents or polymers such as styrene maleic acid) or used to generate virus‑like
particles incorporating the target protein [42–44].
Designing appropriate constructs for expression requires an understanding of pro‑
tein structure, the intended use of protein (e.g. immunogen and/or screening), and the
quantity/quality required. In expression construct design, thought must be given to pro‑
moter strength, protein sequence and structure (full length and/or selected domain(s)),
use of signal peptides (native or alternative), and fusion tags, which may be used for
purication or to introduce a specic feature to facilitate screening (e.g. AviTag to allow
biotinylation). Integral membrane proteins require special consideration depending on
their type and structure. For example, type I membrane proteins have a single trans‑
membrane region and may have a large extracellular domain(s), which can be expressed
directly as a well‑folded protein [45]. Multi‑spanning membrane proteins (e.g. GPCRs,
ion channels) are more challenging to produce, require detergent extraction prior to
purication, and often need reconstitution into lipid membranes using formats such as
nanodiscs to maintain function and stability [46–49]. Alternatively, the membrane pro‑
tein can be expressed in a virus‑like particle format, which avoids detergent extraction
[50]. Finally, for membrane protein targets, it is always necessary to have a cell‑based
system expressing the protein in a native membrane lipid environment. This is to enable
conrmation that any antibodies discovered retain binding to the native target protein.
This could be a cell line naturally expressing the target protein or a recombinantly engi‑
neered cell line overexpressing the target protein of interest [51].
3.4 COMPUTATIONAL METHODS
Antibody discovery, whether following an immunization or display strategy, requires
the production of a target antigen to enable both antibody selection/screening and
functional characterization by binding or via biochemical or cell‑based assays. This
typically requires puried protein but may also use cells overexpressing target protein,
Virus‑like particles (VLPs), and immunization strategies employing priming with DNA
or RNA encoding the target of interest, followed by a boost with an alternative source
of antigen [10,12]. The production of recombinant target antigens can be very challeng‑
ing, particularly for multi‑pass membrane proteins such as G‑protein‑coupled receptors,
ion channels, transporters, and tetraspanins [30,52–54]. Although this can be obviated
using DNA/RNA as the immunogen, a source of target protein is essential for screening
and functional validation of antibodies. These methods are currently mostly experimen‑
tal and highly resource‑intensive.
An emerging and accelerating trend in antibody discovery is the impact of improve‑
ments in computational and MI approaches, which hold the promise of revolutioniz‑
ing biologics drug discovery. These developments have the potential to impact antigen

3 • Computational Protein Design Strategies 47
design, antigen production, and antibody discovery/optimization. Developments in
these areas are moving at different paces, with antigen design and antibody discovery/
optimization being more advanced. It is somewhat challenging to separate these two
areas, but several comprehensive reviews have explored the impact of computational
methods and MI on antibody discovery and design [55–59]. In this chapter, the focus is
on antigen design and computational methods that are impacting this eld.
3.4.1 Impact of AI/ML on Target Protein Expression,
Construct Design, and Protein Production
In designing target protein antigens for antibody discovery, a key element of a success‑
ful strategy is the generation of sufcient high‑quality antigens. Experimentally, this
typically requires expression of the protein antigen in a heterologous system, followed
by purication of the protein and formulation. This has often required bespoke design
of expression constructs and an empirical approach to expression leading to mixed or
unpredictable outcomes. This can provide a bottleneck in recombinant antigen pro‑
duction. This is perhaps unsurprising given the complexity of the processes involved;
parameters that need to be considered include choice of expression host, codon usage,
mRNA synthesis, translation initiation, and protein solubility. In the case of secretory
proteins, signal peptides are required for protein translocation and processing via the
secretory pathway in cells. Just taking one of these considerations, mRNA abundance
alone is not sufcient to explain protein expression abundance [60]. There is thus con‑
siderable interest in building models which may make sequence‑to‑expression mod‑
els available that allow better prediction of expression outcomes. Recent progress in
the areas of DNA synthesis, DNA sequencing, automation, and deep learning is now
being explored to develop deep learning models that allow sequence‑to‑expression
predict ion [61].
This eld is advancing, particularly in the area of E. coli protein expression reect‑
ing the relative ease and low costs of performing expression experiments compared to
other hosts. Computation‑based methods are now emerging, which allow improvements
in protein expression, stability, and function and provide tools to ‘tune’ protein expres‑
sion based on more optimal protein designs, introduction of mutations conducive to pro‑
tein expression, or synonymous codon changes to improve the performance of translation
initiation sites [62–64]. Examples of these strategies and tools are ProteinMPNN [62],
MPEPE (mutation predictor for enhanced protein expression) [63], and the TIsigner.com
web service [65]. The TIsigner web service combines TIsigner (translation initiation
coding region designer), SoDoPE (soluble domain for protein expression), and Razor,
a tool for signal peptide analysis. Codon optimization has also been explored using a
deep learning approach called bidirectional long short‑term memory conditional ran‑
dom eld supported by experimental validation by comparison to commercially acces‑
sible algorithms from Genewiz and ThermoFisher [66]. Although these approaches have
focused on bacterial expression, the application of deep learning methods to other hosts,
such as mammalian systems, is to be expected, particularly given the increasing use of
automation to generate sequence‑to‑expression datasets [67].

48 Biopharmaceutical Informatics
One area where we can expect further impact of ML and deep learning method‑
ologies is for more complex targets such as membrane protein sequences, structures,
and expression. The impact of MI in computational modelling of membrane proteins
has been reviewed recently, suggesting a growing number of applications impacting
membrane protein classication, topology identication, and interaction site detec‑
tion [68,69]. Examples of practical applications of ML for multi‑spanning membrane
proteins are emerging. An example of this can be taken from the research area that
enables structural studies of GPCRs. GPCRs play critical roles in cellular signalling and
are major drug targets. However, their inherent instability in non‑native environments
(e.g. when extracted from the membrane using detergents) complicates their structural
analysis and biophysical characterisation. A well established technique to improve GPCR
stability is to introduce point mutations in transmembrane helices to enhance the ther‑
mostability of the GPCR [70]. This has proven useful in solving the structures of these
inherently exible proteins. Experimentally, this is a challenging and laborious process.
Recently, a computational approach coupled with MI, trained on thermostability data
of 1,231mutants, has been used to predict thermostabilizing mutations in advance of
experiments. Using the C5a receptor as an example, a blind prediction located 36% of
thermostable mutants in the top50 prioritized mutants versus 3% in the rst 50 attempts
using systematic alanine scanning [71]. In a further application of MI‑guided engineer‑
ing, Bedbrook etal. trained statistical models enabling the design of highly functional
light‑gated channelrhodopsins (ChR) [72]. Thirty designed ChR variants were made,
which had improved properties and light sensitivity. Ongoing developments in this area
hold promise for facilitating our ability to design and express membrane protein targets
suitable for use as antigens. One note of caution here is that if mutated proteins are used
as antigens, it is important to check their function and ensure that methods are in place
to conrm that any antibodies identied bind to the native, wild‑type protein.
3.4.2 Computational Protein Structure Prediction
Computational structure prediction of target protein antigens has long been part of ther‑
apeutic antibody discovery [73]. Homology modelling has been an important method for
building 3D protein antigen structures based on using protein primary sequence, mul‑
tiple sequence alignments, and knowledge gained from structural similarities to related
proteins with solved structures. This involves a sequential process whereby sequence
alignment is performed, structural templates are selected, protein backbones are built,
and sidechains are added, followed by an optimization step [74,75]. This can be effective
where the sequence identity between a target antigen and a protein homologue is at least
30%. This type of approach is very useful in providing structural information where a
solved crystal structure is unavailable or too challenging to achieve experimentally. The
information provided by the homology modelling can assist with recombinant antigen
design by providing information on domain structure. This has a practical impact on
issues such as where to place purication/epitope tags to facilitate protein production
and use in screening and design of antigens, which may focus antibody generation on a
desired domain/epitope.

3 • Computational Protein Design Strategies 49
An example of the use of antigen homology models for antibody design is demon‑
strated by engineering a functional antibody directed to IL‑17A. In this study, IL‑17A
was modelled in its receptor‑bound conformation using the experimentally solved
structure of IL‑17F in its receptor‑bound conformation. These cytokines have a 61%
sequence identity. Using this structure, a collection of sequence‑unique antibodies was
used to select a scaffold antibody based on in silico docking of each antibody to IL‑17A
using ZDOCK [76]. The selected antibody was used to design a focused yeast display
library, which allowed isolation of a functional IL‑17A binder that inhibited binding of
IL‑17A to its receptor with a nanomolar EC50. In another example, we have used homol‑
ogy modelling combined with protein–protein docking to guide afnity maturation of
an antibody targeting murine CCL20 [77]. The crystal structure of murine CCL20 was
available (PDB ID: 1HA6), and the structure of the antibody (AB1 Fv) was generated
using SabPred [78]. Docking poses were obtained using RDOCK [79] and conrmed
by a combination of knowledge gained from cross‑reactivity proles (AB1 does not
bind human CCL20) and experimental alanine scanning. In silico afnity maturation
was performed by adopting three protocols (Biovia Discovery Studio [80], Schrödinger
Biologics Suite [81], and Rosetta [82]), and two single‑point mutations were identied,
which increased the afnity of AB1 for CCL20.
Recent advances in computational design are offering the prospect of a new vision
for antibody discovery where prospects for in silico design of antibodies are a real‑
istic aspiration [56,83]. One element of this is the reasonable expectation that accu‑
rately predicted structures of most proteins will be available [84]. Table3.1 captures
a selection of computational methods used to predict protein structures. In addition
to computational methods, continuing advances in experimental technologies, par‑
ticularly Cryogenic electron microscopy (Cryo‑EM), are increasing the deposition of
solved structures to the PDB, particularly for membrane proteins, where deposition has
increased from 30–40 unique structures per year to 70–80 [85]. The deep neural net‑
work tools AlphaFold2 [86] and RoseTTAFold [87] have, and continue, to make a sig‑
nicant impact on protein structure availability, enabling high‑accuracy prediction of
protein structures from amino acid sequence. This opens the way to having an accurate
structural model of any given target protein/antigen [86]. The availability of a structure
for a protein antigen is important to assess protein stability and solubility computation‑
ally using tools such as CamSol [88,89]. The structure also allows the design of target
TABLE3.1 Selected computational protein structure prediction tools
PREDICTION TOOL URL REFERENCE
AlphaFold https://alphafold.ebi.ac.uk [86]
I-TASSER https://zhanggroup.org/I-TASSER/ [102]
Modeller https://salilab.org/modeller/ [103]
Phyre2 http://www.sbg.bio.ic.ac.uk/phyre2 [104]
SWISS-MODEL https://swissmodel.expasy.org/ [105]
Rosetta3 https://www.rosettacommons.org/software [106]
C-QUARK https://zhanggroup.org/QUARK/ [107]

50 Biopharmaceutical Informatics
antigens, which contain mutated residues to enhance protein stability or trap proteins
in desirable conformations to focus immune responses to therapeutically important
epitopes. This could, for example, involve the introduction of disulphide bonds to sta‑
bilize conformationally dynamic structures, as performed for the respiratory syncytial
virus (RSV) fusion glycoprotein [90].
In considering the application of computational tools to improve the immunogenic
potential of selected antigens, epitope prediction and the ability to target pre‑selected
epitopes are of key importance, whether in the eld of therapeutic antibody generation
or in vaccine development. In the case of B‑cell epitope prediction, methods can be
described as sequence‑based or structure‑based. Structure‑based methods are generally
considered more reliable, but their use is limited by the availability of solved antibody–
antigen structures. Structural approaches are generally more powerful, as most epitopes
are conformational. Computational epitope prediction methods include EpiPred [91],
MabTope [92], and DiscoTope 3.0 [93]. This creates an interesting issue, however, as it
is apparent that, in principle, any surface‑accessible area on a protein antigen can be a
potential epitope given the correct antibody binding partner [94]. This is observed in
therapeutic antibodies where approved and clinically successful antibodies bind to the
same target but via distinct epitopes (Figure3.3). A further observation from antibody
discovery campaigns is that screening procedures may select the tightest binders, which
often target immunodominant epitopes and may miss other functionally valid epitopes
[95]. The issue then becomes: How can we predict and select the most relevant func‑
tional epitope? In some cases (e.g. targeting a receptor–ligand interaction), we can focus
attention on the protein–protein interaction surfaces of either the receptor or ligand as
likely epitopes of interest. In other cases, multiple non‑overlapping epitopes may be
targeted and can only be resolved by functional testing and assays designed to select the
required mechanism of action. In the following section, some case studies are presented
of antigen design strategies to enhance antibody–drug isolation.
3.5 CASE STUDY EXAMPLES OF ANTIGEN
DESIGN STRATEGIES TO DRIVE DRUG
DISCOVERY AND IMMUNOGEN PERFORMANCE
Rational design of antigens is of considerable interest in driving therapeutic antibody
discovery. This is also relevant in the design of immunogens for vaccination, particu‑
larly as we emerge from the COVID pandemic, which accelerated research for both
vaccines and neutralizing monoclonal antibodies targeting SARS‑CoV‑2 [96]. There
are thus opportunities for learning from the elds of antibody discovery and emerging
vaccine strategies. This focuses attention on antigen design, which optimizes humoral
response, prevents or reduces off‑target antibody responses, and specically targets pre‑
ferred epitopes [97].
In generating antigens for therapeutic antibody campaigns, knowledge of the target
protein structure, interacting partners, feasibility of production, and desired antibody
Соседние файлы в папке Библиотека им академика М.И. Перельмана
