Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5387_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
15.09.2026
Размер:
16 Мб
Скачать
☆
Computational Biophysical Analyses
7
of Antibody Structure-Function Relationships with Emphasis on Therapeutic Antibody-Based Biologics
Puneet Rawat, Eva Smorodina, Divya Sharma, R. Prabakaran, Jack Wade, Rahmad Akbar, Amrinder Singh, Sandeep Kumar, Victor Greiff, and M. Michael Gromiha

7.1 INTRODUCTION

Antibodies or immunoglobulins (IGs) are a key component of the adaptive immune response. They play a pivotal role in recognizing and binding to a foreign molecule, followed by triggering an immune response against the antigen by recruiting other cells and molecules. Antibodies are heavy proteins with an approximate size of 10 nm and a weight of 150 kDa (Reth, 2013). These molecules are roughly Y shape, composed
161
162 Biopharmaceutical Informatics
of two identical heavy and light chains connected by disulde bonds (Glockshuber etal., 1992). The recognition of any foreign molecule called “antigen” mainly involves the variable region of the antibody (upper tips of the Y shape including both heavy and light chains), and the activation of the immune response is regulated by the constant region of the antibody (Segal etal., 1974; Sela‑Culang etal., 2013; Sinclair etal., 1968). Antibodies possess a high level of diversity, allowing them to target a wide range of antigens, particularly in the antibody variable regions. The diversity in the antibody chains is generated through V(D)J recombination, insertion/deletion during the recom‑ bination process, and somatic hypermutation. The variable region of the antibody fur‑ ther consists of four framework regions (FRs) and three complementarity‑determining regions (CDRs) on both light and heavy chains. The CDRs on the antibody predom‑ inantly interact with the antigen (Padlan et al., 1995). However, the highly diverse CDR3 on the heavy chain (CDRH3) contributes signicantly to the antibody–antigen interactions, in most cases (Akbar etal., 2021; Chothia & Lesk, 1987; Xu & Davis,
2000). The interacting residues in the antibody–antigen complex are called “para‑ topes” on the antibody side and “epitopes” on the antigen side (Akbar et al., 2021; Sela‑Culang etal., 2013). Therefore the main features of antibodies can be classied as (i) specicity toward the antigens due to high diversity on the CDRs, (ii) antibody diversity through V(D)J recombination and somatic hypermutations to recognize wide range of antigens, (iii) tolerance toward host proteins/cells, (iv) optimized biophysical properties to work effectively at the biological conditions, (v) immunological memory to build immunity upon reinfection, (vi) trigger Fc mediated effector functions by neu‑ tralization (direct binding to the pathogen to prevent infection), opsonization (activa‑ tion of phagocytic cells), complement activation (activation of downstream cascade to neutralize the pathogens), and antibody‑dependent cellular cytotoxicity (ADCC; lysis of infected cells activated by antibodies) (Lu etal., 2018).
There are two major aspects related to antibody–antigen interaction, namely, “afnity and avidity” (Rudnick & Adams, 2009; Yin etal., 2021). Afnity measures the strength of the epitope binding to an antibody and is often represented by the dis‑ sociation constant KD. The paratope and epitope residues interact through various types of non‑covalent interactions, such as hydrogen bonds, ionic bonds, Van der Waals, and hydrophobic interactions. Avidity measures the overall strength of the antibody–antigen complex. Avidity includes the valency of the protein, the binding afnity of the anti‑ body–antigen complex, as well as the structural arrangement of the antibody(s) (Evans & Thurber, 2022).
Antibodies are also attractive therapeutic candidates due to their biological role in the immune response and high specicity toward the antigen (Lu et al., 2020). Monoclonal antibodies or mAbs are widely used as therapeutic candidates because they are highly specic to particular antigens. Currently, there are approximately 100 Food and Drug Administration (FDA)‑approved mAbs with an estimated market size of ~185.50 billion USD, which is expected to reach more than 500 billion USD by 2030 (Mullard, 2021). Antibody‑based therapeutics development also faces several challenges. For example, antibody repertoire sizes for an individual range from 108 to
10
in humans (Elhanati et al., 2014; Glanville etal., 2009). Although antibody rep‑
10 ertoires may show convergence based on post‑exposure to similar antigens, there is still vast diversity in the repertoires to be analyzed experimentally (Greiff etal., 2017).
7 • Antibody Structure-Function 163
For example, Wardemann et al. used single‑cell cloning strategies and antibody expression to exhibit self‑expression of newly generated B‑cells in the bone marrow (Wardemann etal., 2003). This opened an arena to several immunological insights and furthered the isolation of antibodies to neutralize clinical pathogens like SARS‑CoV, inuenza, Human immunodeciency viruses (HIV), and many others using these B‑cell cloning strategies. However, the main limitation posed by this technique is that it pro‑ vides only a sliver of information on the full antibody repertoire and is usually limited to a subset of antigen binding activity. Further, Ig‑sequencing limitations include iden‑ tication of a suitable source of DNA sequence and quantication errors; ability to dis‑ tinguish which V
genes pair with VL genes in each B‑cell; use of appropriate data and
H
visualization tools to detangle the large amounts of information furnished post‑analysis; cross‑reactivity of antibodies with host protein; and lack of structural information about the antibodies (Brown etal., 2019; Georgiou et al., 2014; Greiff et al., 2015). Once the binding with the target antigen is established, there are several other developabil‑ ity challenges for naturally occurring antibodies to be used as therapeutic antibodies, which include folding stability, aggregation, viscosity, etc. (Ahmed etal., 2021; Akbar etal., 2022; Jain etal., 2017; Młokosiewicz etal., 2022; Pérez etal., 2022; Raybould etal., 2019; Xu etal., 2019). Therefore, experimental validation of potential therapeutic antibody leads is challenging due to high production costs and a lack of large‑scale and high‑throughput methods for developability prediction (Schlander etal., 2021).
With the advent of computational sciences, researchers have probed machine‑learn‑ ing (ML) methods and other informatics technologies to study plausible relationships between an antibody’s structure and function. For example, there are several com‑ putational resources available for antibody structure modeling, in silico screening of epitope/paratope regions, docking of antibodies with antigen, estimation of binding energies, and calculation of biophysical properties of antibodies such as aggregation propensity, solubility, and melting temperature (Ahmed etal., 2021; Chiu etal., 2019a; Harmalkar etal., 2022; Narayanan etal., 2021a; Sankar etal., 2022). However, even in these methods, there are certain limitations wherein predicting models of disordered or post‑translationally modied (e.g., glycosylated) proteins is not possible. The dock‑ ing methods require prior information on epitope regions for better prediction of the antibody–antigen complex, and binding energy prediction methods include inaccura‑ cies in the calculation of absolute binding free energies. Moreover, the sequence‑based approach to study binding prediction is unreliable because the structural conformation of the CDRs can signicantly affect the binding (Akbar etal., 2022).
Conventionally, antibody repurposing and optimization of the therapeutic antibodies have been widely used by in silico researchers due to limited resources and long down‑ stream validation processes (Mason etal., 2021; Rawat etal., 2021; Rodriguez‑Quijada etal., 2020; Wang, Gallolu Kankanamalage, etal., 2021). However, there have been signicant advances in the computational approaches for the design and development of antibodies in recent years (Hummer etal., 2022; Norman etal., 2020; Tiller & Tessier,
2015). In this chapter, we will be focusing on the computational resource developed to curate antibody‑related information, antibody–antigen binding (docking and binding afnity prediction), and biophysical parameters affecting antigen design (Figure7.1). We have also highlighted the role of language models and molecular dynamics (MD) simulations in antibody structure‑function relationship prediction.
164 Biopharmaceutical Informatics
FIGURE 7.1 Overview of the topics related to antibody structure-function considered in the book chapter. There are several antibody-related databases that contain sequence, structure, and other relevant information related to antibodies (e.g., interacting pathogen, residue-level interaction, binding afnity, epitope, developability-related information, and so on). There are several possible approaches for the antibody–antigen structure predic­tion. Docking (faster computation time) and MD simulations (very high computation time) are classical approaches widely used to predict antibody–antigen structures, whereas ML/DL methods and language models are more recently developed approaches showing much better speed and accuracy compared to the classical approaches. Several binding afnity prediction methods are also developed, which can use the sequence/structure infor­mation of the antibody–antigen complex to predict (i) the binding afnity of the complex or (ii) the change in binding afnity upon point mutation at the interaction interface.
7 • Antibody Structure-Function 165
7.2 COMPUTATIONAL RESOURCES FOR
ANTIBODY STRUCTURE AND FUNCTION
7.2.1 Antibody‑Related Online
Resources and Databases
As antibodies become an increasingly interesting topic in the eld of biotherapeutics, the demand for antibody‑specic data in public repositories is growing (Norman etal.,
2020). The online resource and databases on antibodies can be divided into two major categories: (i) primary sequence/structure databases and (ii) derived databases, which contain the secondary information obtained from the antibody sequence/structure or biological activity (such as antibody–antigen interaction, binding afnity/neutralization activity, and epitope/antigen information) (Table7.1).
7.2.1.1 Primary databases
Primary databases contain the sequence/structure of antibodies or immune reper‑ toire sequences from B‑cells. The international ImMunoGeneTics information system (IMGT) is one of the comprehensive resources on IGs or antibodies, T‑cell recep‑ tors (TCRs), and major histocompatibility (MH) of human and other vertebrate spe‑ cies. It consists of sequence databases, genome databases, structure databases, and mAbs’ databases, which are embedded into several web resources and interactive tools (Ehrenmann etal., 2010). Antibody sequences and structures are also curated in the abYsis database (Swindells etal., 2017), providing an inbuilt analysis platform. The AntiBodies Chemically Dened (ABCD) database is a manually curated repository of sequenced antibodies (Lima etal., 2020). The Protein Data Bank (PDB) is a general resource for experimentally determined protein structures, which also include struc‑ tures of antibodies/antibody complexes (Rose etal., 2021). Several antibody‑ specic structure databases were also developed using PDB, which contains PDB struc‑ tures as well as related annotated information. For example, SAbDab (Dunbar etal.,
2014) contains the antibody structures from PDB, which are annotated with several details, including experimental details, antibody nomenclature (e.g., heavy‑light pair‑ ings), curated afnity data, and sequence annotations; abYbank contains sequences (EMBLIG, Kabat, and AbPDBSeq databases) and renumbered experimental structures (Kabat, Chothia, and Martin antibody numbering scheme in the AbDb database) of antibodies from PDB (Ferdous & Martin, 2018). The B‑cell repertoire‑specic data‑ bases include observed antibody space (OAS) (Olsen etal., 2022a), VBASE2 (Retter etal., 2005), cAb‑Rep (Guo etal., 2019), Pan Immune Repertoire Database (PIRD) (Zhang etal., 2020), and VDJbase (Omer etal., 2020). These immune repertoire data‑ bases can be used as benchmarking datasets for humanness and developability param‑ eters of therapeutic antibodies.
TABLE7.1 List of antibody-related sequence, structure, and other specialized databases
SR.NO. DATABASE LINK DESCRIPTION REFERENCE
Sequence/Structure Database
1 IMGT https://www.imgt.org/ Comprehensive sequence/structure/genome/
monoclonal antibody database
2 SAbDab https://opig.stats.ox.ac.uk/
webapps/newsabdab/sabdab/
3 PDB https://www.rcsb.org/ A generalized structure database which also contains
4 abYbank http://www.abybank.org/ Antibody sequence and renumbered structures data Ferdous and
5 cAb-Rep https://cab-rep.c2b2.columbia.edu/ Database of curated antibody repertoires (Guo etal., 2019) 6 OAS http://opig.stats.ox.ac.uk/webapps/
oas/
7 VDJbase https://vdjbase.org/ Database of adaptive immune receptor genes,
8 PIRD https://db.cngb.org/pird/ Database of raw and processed sequences of IGs and
9 VBASE2 http://www.vbase2.org/ Database of human germline variable region
10 ABCD https://web.expasy.org/abcd/ Database is a manually curated depository of
11 abYsis http://www.abysis.org/abysis/ Integrated database of antibody sequence and
Structure database for antibodies Dunbar etal.
antibodies
Annotated immune repertoire database Olsen etal.
genotypes, and haplotypes
T-cell receptors (TCRs) of human and other vertebrate species
sequences
sequenced antibodies
structure data
Ehrenmann etal.
(2010)
(2014)
Rose etal. (2021)
Martin (2018)
(2022a)
Omer etal. (2020)
Zhang etal. (2020)
Retter etal. (2005)
Lima etal. (2020)
Swindells etal.
(2017)
(Continued)
166 Biopharmaceutical Informatics
TABLE7.1 (Continued ) List of antibody-related sequence, structure, and other specialized databases
SR.NO. DATABASE LINK DESCRIPTION REFERENCE
12 iReceptor http://ireceptor.irmacs.sfu.ca/ NGS sequence data on B-cell receptors Corrie etal. (2018)
Specialized Sequence/Structure Database
1 Thera-SAbDab http://opig.stats.ox.ac.uk/webapps/
newsabdab/therasabdab/
2 CoV-Ab-Dab http://opig.stats.ox.ac.uk/webapps/
covabdab/
3 Ab-CoV https://web.iitm.ac.in/bioinfo2/
ab-cov/home 4 IEDB https://www.iedb.org/ Antibody epitope database Vita etal. (2019) 5 bNAber http://bnaber.org/* Database of broadly neutralizing HIV antibodies Eroshkin etal.
6 AgAbDb http://bioinfo.net.in/AgAbDb.htm* Antibody–antigen interaction database Kulkarni-Kale etal.
7 CPAD2.0 https://web.iitm.ac.in/bioinfo2/
cpad2/ 8 AL-Base https://wwwapp.bumc.bu.edu/
BEDAC_ALBase/ 9 AB-Bind https://github.com/sarahsirin/
AB-Bind-Database 10 SKEMPI 2.0 https://life.bsc.es/pid/skempi2/ Database of kinetics and energetics information upon
11 PROXiMATE https://www.iitm.ac.in/bioinfo/
PROXiMATE/
The links that are not active (as checked on Dec 2022) are denoted with “*” sign.
Sequence database for approved or clinical-stage
therapeutic antibodies
Sequence database for coronavirus-related antibodies Raybould etal.
Experimental neutralization prole of
coronavirus-related antibodies
Experimental protein aggregation information which
also includes antibodies
Experimental amyloidogenic antibody light chain
database
Database of experimentally determined changes in
binding free energies
mutation and includes antibody–antigen complexes
A mutant protein–protein interaction kinetics and
thermodynamics database
Raybould etal.
(2020)
(2021)
Rawat etal. (2022)
(2014)
(2014)
Rawat etal. (2020)
Bodi etal. (2009)
Sirin etal. (2016)
Jankauskaitė etal.
(2018)
Jemimah etal.
(2017)
7 • Antibody Structure-Function 167
168 Biopharmaceutical Informatics
7.2.1.2 Specialized sequence/structure databases
Specialized databases contain a variety of secondary information derived from the primary databases. There are several specialized sequence databases such as Thera‑SAbDab (Raybould et al., 2020) for approved and clinical‑stage antibodies, CoV‑AbDab (Raybould etal., 2021) for coronavirus‑related antibodies, broadly neu‑ tralizing antibodies electronic resource (bNAber) for broadly neutralizing HIV anti‑ bodies (Eroshkin etal., 2014), Amyloid Light Chain Database (AL‑Base) for antibody light chains with aggregation capability (Bodi etal., 2009), and so on. AgAbDb is a unique database, which contains the antibody–antigen interaction details at the resi‑ due level along with other parameters such as interaction type and accessible surface area (Kulkarni‑Kale etal., 2014). The AB‑Bind database contains the experimentally determined change in binding free energy values (Sirin etal., 2016). SKEMPI 2.0 and PROXiMATE databases contain changes in thermodynamic parameters and kinetic rate constants upon point mutations for protein–protein interaction, which also includes antibody–antigen interactions (Jankauskaitė et al., 2018; Jemimah et al., 2017). The curated protein aggregation database (CPAD) 2.0 database provides comprehensive experimentally determined information on aggregation‑prone regions and aggregation kinetics for all proteins, including antibodies (Rawat etal., 2020). The Ab‑CoV database is a coronavirus‑specic antibody database that contains the antibodies’ neutralization prole (IC50 and EC50) and binding afnity (KD), as well as computationally predicted changes in stability and binding afnity upon epitope/paratope residue mutation (Rawat etal., 2022). The Immune Epitope Database (IEDB) considers the antigen‑side informa‑ tion and collects epitope information (Vita etal., 2019).
7.2.2 Computational Methods for Investigating Structure‑Function Relationship
7.2.2.1 Docking tools for antibody–antigen
complex structure prediction
Docking is a molecular modeling technique, which predicts the conformation of one molecule (ligand) on the surface of another static and larger molecule (receptor). In the case of antibody–antigen docking, antibodies are usually considered receptors, while antigen is considered a ligand. Molecular docking allows us to see residue‑level interac‑ tions between the receptor and ligand, which is crucial for drug discovery (Pagadala etal., 2017; Pinzi & Rastelli, 2019). Most docking methods generate several docked structures and provide docking scores to rank each pose. These scores can be unrelated to real estimation of the binding strengths of complex structures, but they allow the comparison of different conformations of one binder or different binders between each other within one docking tool. Lower scores usually represent better binder conforma‑ tions. Recent docking methods prefer ensemble of protein structures (usually from MD simulations) to identify the correct pose (Amaro etal., 2018). Docking poses can also be evaluated through an afnity scoring function representing electrostatic and Van der Waals interactions (Pagadala etal., 2017; Rawat etal., 2021). Docking programs
7 • Antibody Structure‑Function 169
typically rely on an estimated binding site to accurately predict the binding interfaces, and they often lack precision in predicting binding energies (Wang etal., 2003). MD simulations provide a more precise estimation of binding energies (Fernández‑Quintero etal., 2022; Kralj etal., 2021; Salmaso & Moro, 2018). Some docking models allow the ligand to be treated as a rigid body object (all atoms and residues are static and immovable) or exible (when some amino acids are allowed to move during the docking procedure). The docking procedure that requires many ligands to bind to one receptor is called virtual screening. It’s a very common approach in the early stages of drug discovery (Schneider etal., 2022). Docking is the widely used approach to investi‑ gate the function or binding of the antibody against an antigen using protein struc‑ tural information (Brooks etal., 2020; Chaves etal., 2020; Guest etal., 2021). ZDOCK (Pierce etal., 2014), Haddock (Dominguez etal., 2003), ClusPro 2.0 (Comeau etal.,
2004), LightDock (Jiménez‑García et al., 2018), and Rosetta (Schoeder etal., 2021)
are highly used for protein–protein docking tasks, including antibody–antigen docking (Ambrosetti etal., 2020). Antibody–antigen‑specic docking models include Antibody i‑Patch and Absolut!. Antibody i‑Patch denes antibody residues which are most likely in contact with antigen (Krawczyk etal., 2013) and Absolut! generates coarse‑grained synthetic antibody–antigen complexes with information about paratope and epitope conformations and afnity (Robert etal., 2022). A list of the most used docking tools for antibody–antigen docking is provided in Table7.2.
ML techniques can outperform classical molecular docking results (Ganea etal.,
2021). One of the models is based on graph neural networks (EquiDock) and predicts
rotations and translations of molecules in rigid docking (Ganea etal., 2021). Another model is a diffusion generative model (DiffDock), which renes random docking poses to reach the best complex conformation via translations, rotations, and torsion angles (Corso etal., 2022). The recently introduced architecture called Hierarchical Equivariant Renement Network (HERN) in “abdockgen” allows not only rigid dock‑ ing but also renes side chains after pose generation (Jin etal., 2022).

7.2.3 Computational Approaches to Predict Antibody–Antigen Interaction

In this section, we have summarized different methods for antibody–antigen binding and afnity prediction tools that do not involve docking approaches. The details regard‑ ing the tools predicting only paratope or epitope regions can be found elsewhere and not discussed here (Akbar etal., 2022; Chinery etal., 2023; Lo etal., 2021). The ini‑ tial methods for antibody–antigen complex prediction were based on amino acid usage. For example, Bepar uses a sliding window of amino acids between the antigen and CDR regions of antibodies to predict the epitope residues using sequence informa‑ tion (Zhao & Li, 2010). EpiPred is another method that predicts the epitope regions for antibodies using conformational matching and specic antibody–antigen scores (Krawczyk etal., 2014). The methods developed later utilized machine learning (ML) or deep learning (DL) approaches for the prediction of epitope–paratope residues. The random forest‑based ML model “PEASE” (Sela‑Culang etal., 2014, 2015) calculates the “residue‑score” for both antibody and antigen to identify interacting epitope–paratope
TABLE7.2 List of molecular docking tools widely used for antibody–antigen complex prediction
DOCKING
SR.NO.
1 ZDOCK https://zdock.umassmed.edu/ Fast Fourier Transform-based protein docking Pierce etal. (2014)
2 Haddock https://wenmr.science.uu.nl/
3 ClusPro 2.0 https://cluspro.org/login.php Fast Fourier Transform-based rigid docking Comeau etal. (2004)
4 LightDock https://lightdock.org/ Docking protocol based on the Glowworm
5 Rosetta
6 PatchDock http://bioinfo3d.cs.tau.ac.il/
7 AbAdapt https://sysimm.org/abadapt/ From sequence to docked structure
8 EquiDock https://github.com/octavian-ganea/
9 DiffDock https://github.com/gcorso/DiffDock A diffusion generative model over the
10 abdockgen https://github.com/wengong-jin/
11 Antibody
12 Absolut! https://github.com/csi-greifab/
TOOL LINK DESCRIPTION REFERENCE
Information-driven exible docking approach Dominguez etal.
(SnugDock)
i-Patch
haddock2.4/
https://www.rosettacommons.org/
software
PatchDock/
equidock_public
abdockgen
http://opig.stats.ox.ac.uk/webapps/
newsabdab/sabpred/antibodyipatch
Absolut
Swarm Optimization (GSO) algorithm
Simulates the induced-t mechanism Schoeder etal.
Geometry-based molecular docking
algorithm
generation pipeline
Pairwise-independent SE(3)-equivariant graph
matching network-based rigid docking
non-Euclidean manifold of ligand poses
Antibody–antigen docking and design via
hierarchical equivariant renement
Contact likelihood score to each residue Krawczyk etal.
Unconstrained lattice-based antibody–
antigen bindings generator
(2003)
Jiménez-García etal.
(2018)
(2021)
Schneidman-Duhovny
etal. (2005)
Davila etal. (2022)
Ganea etal. (2021)
Corso etal. (2022)
Jin etal. (2022)
(2013)
Robert etal. (2022)
170 Biopharmaceutical Informatics