Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5387_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
15.09.2026
Размер:
16 Мб
Скачать
☆
7 • Antibody Structure‑Function 171
residues using 120 properties calculated from the antibody–antigen complexes. It also allows an optional additional step of nding surface patches on antigens with high scores (termed patch scores). Another study combined statistical and ML algorithms to predict the antibody‑specic epitopes. They calculated several geometric and physicochemi‑ cal features related to interacting regions of antibody–antigen complexes. These fea‑ tures were used in Monte Carlo algorithms to generate putative epitope–paratope pairs, used as training datasets in the ML model (Jespersen etal., 2019). A DL‑based frame‑ work, “Paratope and Epitope prediction with graph Convolution Attention Network” (PECAN), uses a protein–protein interaction dataset for training and utilizes transfer learning to predict the binding interface of antibody–antigen complexes. The local resi‑ dues at close spatial proximity of the interfaces were captured using graph convolutions, while an attention layer was employed to encode antibody–antigen pair interactions (Pittala & Bailey‑Kellogg, 2020).
A protein–protein interaction prediction method, “Molecular Surface Interaction Fingerprints” (MaSIF), applies geometric deep learning (GDL), which can incorporate geometric features such as structure and symmetry of the input to improve the quality of the predictions (Gainza etal., 2020). MaSIF also uses antibody–antigen complexes as training data. MaSIF converts the protein surface into a mesh representation, where each vertex contains two geometric features (shape and distance‑dependent curvature) and three chemical features (hydropathy, continuum electrostatics, and location of free elec‑ trons/proton donors). Furthermore, a set of geodesic lters generates an output with a xed dimension. The performance of MaSIF showed an ROC AUC of 0.77 with one geodesic convolutional layer and 0.86 with three layers. A simple version of this GDL model, dif‑ ferentiable molecular surface interaction ngerprinting (dMaSIF), is also now available for large‑scale analysis with simplied surface representation (Sverrisson etal., 2020). Del Vecchio etal. (2021) argued that paratope and epitope prediction require asymmetric treatment. Therefore, they developed a separate paratope model (Para‑EPMP) and epit‑ ope model (Epi‑EPMP) for joint paratope and epitope prediction. Rangel etal. utilized a fragment‑based approach for the combinatorial design of antibody binding loops (CDRs) and grafted them onto antibody scaffolds (Rangel etal., 2021). The designed CDRs were also computationally optimized for solubility and conformational solubility.
7.2.4 In Silico Prediction of Binding Afnity
Understanding protein–protein interactions is crucial in the investigation of biological systems, as these play a critical role in almost all cellular processes (Gromiha, 2020; Jones & Thornton, 1996; Perkins etal., 2010). Binding afnity is dened as the strength of interaction between proteins and/or peptides. However, higher binding afnity does not always lead to the best therapeutic response (Yu etal., 2023). The binding afnity of an interaction is described through the equilibrium dissociation constant KD, or, in thermodynamic terms, the Gibbs free energy ΔG (ΔG = -RT ln KD) (Kastritis & Bonvin,
2013). Experimentally measuring KD values is a time‑consuming and expensive process (Jarmoskaite etal., 2020). Therefore, many computational methods have been developed for predicting the binding afnity (Table7.3). Binding afnity prediction is important because it not only allows to control interactions and develop innovative therapeutics but
TABLE7.3 List of computational resources available for the prediction of the absolute value of binding afnity and change in binding afnity upon mutation(s)
BINDING AFFINITY PREDICTION TOOLS
METHOD FEATURES PERFORMANCE MODEL TYPE TRAINED ON
AB‑AG DATA?
1. Sequence‑Based Prediction Methods
PPA-Pred Sequence-based afnity
prediction using functional information
ISLAND Kernel representation r = 0.44 on
PIPR Pre-trained embeddings r = 0.87 on
RAPPPID Pairs of amino acid sequences r = 0.97 on
TcellMatch Sequence embedding r = 0.63 on 10x
r = 0.90 on 135
complexes selected from structure-based benchmark
structure-based benchmark
SKEMPI
STRING
dataset
Regression Yes (11.1%;
15 out of
135)
Support vector machine
(SVM)
RRCNN Yes (11.02%;
Neural network Yes (2.6%;
Neural network Yes (100%;
Yes (11.1%;
15 out of
135)
781 out of 7,085)
242 out of 9,340)
4,812)
URL/REFERENCE
https://www.iitm.ac.in/
bioinfo/PPA_Pred/ (Yugandhar & Michael Gromiha, 2014)
https://sites.google.com/
view/wajidarshad/software (Abbasi etal., 2020)
https://github.com/
muhaochen/seq_ppi (Chen etal., 2019)
https://github.com/jszym/
rapppid (Szymborski & Emad, 2022)
(Fischer etal., 2020)
172 Biopharmaceutical Informatics
(Continued)
TABLE7.3 (Continued ) List of computational resources available for the prediction of the absolute value of binding afnity and change in binding afnity upon mutation(s)
BINDING AFFINITY PREDICTION TOOLS
METHOD FEATURES PERFORMANCE MODEL TYPE TRAINED ON
AB‑AG DATA?
Structure‑Based Prediction Methods
2.
PRODIGY Inter‑residue contacts and
noninteracting surface
FoldX Empirical function 64% Empirical function No https://foldxsuite.crg.eu/
PPI‑Afnity ProtDcal r = 0.77 on the
CSM‑AB Graph‑based signatures r = 0.64 on
r = 0.73 on
benchmark of 81 protein– protein complexes
PDBbind database
PDBbind, SabDab, RCSB PDB
Linear regression Yes (14.8%;
12 out of 81)
SVM Yes (9.8%; 82
out of 833)
Regression Yes (100%;
472)
URL/REFERENCE
https://wenmr.science.uu.nl/
prodigy/ (Xue etal., 2016)
(Delgado etal., 2019)
https://protdcal.zmb.uni‑due.
de/PPIAfnity* (Romero‑Molina etal.,
2022)
http://biosig.unimelb.edu.au/
csm_ab/prediction (Myung etal., 2022)
(Continued)
7 • Antibody Structure‑Function 173
TABLE7.3 (Continued ) List of computational resources available for the prediction of the absolute value of binding afnity and change in binding afnity upon mutation(s)
BINDING AFFINITY PREDICTION TOOLS
METHOD FEATURES PERFORMANCE MODEL TYPE TRAINED ON
AB‑AG DATA?
Change in Binding Afnity Upon Mutation Prediction Tools
1. Sequence‑Based Prediction Methods
ProAfMuSeq Sequence-based features and
functional class
PANDA Sequence-based r = 0.52 on
SAAMBE-SEQ Sequence-based, physical
properties
2. Structure‑Based Prediction Methods
BeAtMuSic Statistical potentials r = 0.4 on
BindProfX Interface prole score, shape
complementarity, and sequence-based features
r = 0.73 on
mutation data from PROXiMATE
SKEMPI 2.0
r = 0.83 on
SKEMPI 2.0
SKEMPI
r = 0.68 on
SKEMPI
Regression Yes (16.1%;
189 out of 1,173)
Regression Yes (11.02%;
781 out of 7,085)
Gradient-boosting
decision tree
Regression Yes (1.8%; 55
Prole score Yes (1.8%; 55
Yes (11.02%;
781 out of 7,085)
out of 3,047)
out of 3,047)
URL/REFERENCE
https://web.iitm.ac.in/
bioinfo2/proafmuseq/ (Jemimah etal., 2019)
https://github.com/
wajidarshad/panda (Abbasi etal., 2021)
http://compbio.clemson.edu/
saambe_webserver/ indexSEQ.php#started (Li etal., 2021)
http://babylone.ulb.ac.be/
beatmusic/index.php (Dehouck etal., 2013)
https://zhanggroup.org/
BindProfX/ (Xiong etal.,
2017)
174 Biopharmaceutical Informatics
(Continued)
TABLE7.3 (Continued ) List of computational resources available for the prediction of the absolute value of binding afnity and change in binding afnity upon mutation(s)
BINDING AFFINITY PREDICTION TOOLS
METHOD FEATURES PERFORMANCE MODEL TYPE TRAINED ON
AB‑AG DATA?
URL/REFERENCE
SAAMBE Van der Waals, solvation and
Coulomb energy, entropy, hydrophobicity, solvent accessible surface area, hydrogen bonds and interface area
MutaBind Van der Waals energy,
solvation energy, free energy change due to unfolding, solvent accessible surface area
mmCSM-PPI Graph-based signatures and
complementary features
r = 0.82 on
SKEMPI 2.0
r = 0.68 on
SKEMPI
r = 0.75 on
SKEMPI 2.0
XGBoost Yes (11.02%;
781 out of 7,085)
Molecular mechanics
force elds, statistical potentials, and fast side-chain optimization algorithms
Extra trees Yes (11.02%;
Yes (1.8%; 55
out of 3,047)
781 out of 7,085)
http://compbio.clemson.edu/
saambe_webserver/ (Li etal., 2021)
7 • Antibody Structure-Function 175
https://lilab.jysw.suda.edu.cn/
research/mutabind2// (Li etal., 2016)
https://biosig.lab.uq.edu.au/
mmcsm_ppi/ (Rodrigues etal., 2021)
(Continued)
TABLE7.3 (Continued ) List of computational resources available for the prediction of the absolute value of binding afnity and change in binding afnity upon mutation(s)
BINDING AFFINITY PREDICTION TOOLS
METHOD FEATURES PERFORMANCE MODEL TYPE TRAINED ON
AB‑AG DATA?
URL/REFERENCE
176 Biopharmaceutical Informatics
GeoPPI Graph neural network r
TopNetTree CNN, persistent homology r = 0.79 on
PerSpect‑EL Physical properties, persistent
homology
mCSM‑AB Graph‑based signatures r = 0.53 on 29
FoldX Empirical function 64% Empirical function No https://foldxsuite.crg.eu/
The links that are not active (as checked on Dec 2024) are denoted with “*” sign.
= 0.52 on SKEMPI 2.0
SKEMPI 2.0
r = 0.85 on
SKEMPI 2.0
Ab‑Ag complexes
Gradient‑boosting tree Yes (11.02%;
781 out of 7,085)
Gradient‑boosting tree Yes (11.02%;
781 out of 7,085)
CNN+gradient‑boosting
tree
Regression Yes (100%;
Yes (11.02%;
781 out of 7,085)
645)
https://github.com/Liuxg16/
GeoPPI (Liu etal., 2021)
(Wang etal., 2020)
https://github.com/
ExpectozJJ/ PerSpect‑Ensemble‑Learning (Wee & Xia, 2022)
https://biosig.lab.uq.edu.au/
mcsm_ab/prediction (Pires & Ascher, 2016)
(Delgado etal., 2019)
7 • Antibody Structure-Function 177
also for other applications such as protein engineering, computational mutagenesis, and docking (Ben‑Shimon & Eisenstein, 2010; Keskin etal., 2005; Kortemme etal., 2004; Vangone & Bonvin, 2015).
7.2.4.1 Binding afnity prediction methods
There are several sequence‑based binding afnity prediction methods available for protein–protein complexes. Protein‑Protein Afnity Predictor (PPA‑Pred), a tool for predicting the real value of binding afnity from amino acid sequences, is based on a multiple regression model (Yugandhar & Michael Gromiha, 2014). The sequence‑based features include predicted binding site residues and property values of 20 amino acids from the AAindex database (Kawashima, 2000). The training data included antibody– antigen complexes as well and showed a correlation from 0.74 to 0.99 for different classes of complexes. ISLAND (In SiLico protein AfNity preDictor), another sequence‑based tool, combined a kernel representation of protein sequences with the support vector regression to predict the binding afnity. The correlation between the experimental and predicted ΔG was 0.44, and the structure‑based benchmark for protein–protein bind‑ ing afnity data was used (Abbasi etal., 2020). Another tool based on the recurrent convolutional neural network (RCNN) that takes amino acid sequence as the input for prediction of protein–protein binding afnity was developed by Chen etal. (2019). The correlation of 0.87 was obtained from a Siamese residual RCNN with a pre‑trained embedding representation of protein sequences. A similar model known as DPPI was developed by Hashemifar et al., which was based on DL model (Hashemifar et al.,
2018). Another end‑to‑end DL framework that learns both robust local features and contextualized information from sequences is Protein–Protein Interaction Prediction Based on Siamese Residual RCNN (PIPR), which was able to predict the binding afni‑ ties between the interacting proteins (Chen etal., 2019). Moreover, a method named regularized automatic prediction of PPIs using deep learning (RAPPPID) was trained by considering pairs of amino acid sequences of interacting proteins and allowing better distinctiveness between the interacting motifs and other parts of the proteins. Further, Xue etal. developed a method based on pre‑trained embedding; and residual RCNN, structure information, and functions of proteins were used in the pre‑training stage to generate sequence embeddings. However, the performance of the model was poor with a correlation of only 0.26 (Xue etal. 2021). Fischer etal considered the UMI counts in 10x Genomics single‑cell immune proling dataset as binding strength of the TCR‑pMHC complex and developed a model named “TcellMatch” with r2 values of 0.63 (Fischer etal., 2020). Makowski etal. combined high‑throughput experimental methods includ‑ ing deep sequencing and ML to identify therapeutic antibody variants with superior combinations of afnity and non‑specic binding (Makowski etal., 2022). The model is trained on binary datasets for afnity and specicity and does not consider real binding afnity prediction, yet it correlates with continuous afnity values. A similar approach is also used in the pipeline called RESP that is trained on over 3million human B‑cell receptor sequences. The pipeline efciently identies the high‑afnity antibodies but is not designed to predict binding afnity (Parkinson etal., 2023). Bachas et al. (2022) used deep contextual language models trained on high‑throughput afnity data to quantitatively predict binding of unseen antibody sequence variants and included a
178 Biopharmaceutical Informatics
metric to score antibody variants for similarity to natural IGs. It is important to note that sequence‑based methods are unable to perform predictions for different binding poses of the interacting proteins and do not take conformational changes into account (Gromiha etal., 2017).
The structure‑based methods have signicant advantages over sequence‑based pre‑ diction. However, they usually lack the high‑quality structural data. The rst study to relate binding afnities with a set of structures was by Horton and Lewis who used 15 ΔG values from literature as training data and used interface polar and non‑polar groups as features to obtain a linear regression coefcient of r = 0.96, and a mean absolute difference of 0.8 kcal/mol between the calculated and observed ΔG values (Horton & Lewis, 1992). Kastritis etal. (2011) benchmarked the protein–protein binding afnity data for 144 protein–protein complexes (including 19 antibody–antigen complexes) with varying biological functions and observed that the performance was poor on a validation set because of noise in the experimental data (Kastritis & Bonvin, 2011). Faster methods based on empirical functions (empirical, force‑eld‑based potentials, statistical poten‑ tials, and scoring functions used in docking) could be successful on small training sets (Audie & Scarlata, 2007; Horton & Lewis, 1992) but most of them fail to predict bind‑ ing afnity accurately (Rosato, 2010) for large datasets or discriminate between binders and non‑binders (Sacquin‑Mora etal., 2008). Vangone and Bonvin (2015) worked on relating the interfacial contacts (ICs) and noninteracting surface (NIS) residues with the experimental binding afnity and obtained a Pearson correlation of −0.73. They used a training dataset of 81 protein–protein complexes (including ten antibody–antigen com‑ plexes) and developed a method, PROtein binDIng enerGY prediction (PRODIGY), that can predict the binding afnity of protein–protein complexes from their 3D structure with a correlation of 0.73 between experimental and predicted ΔG values (Vangone & Bonvin, 2015; Xue etal., 2016). Vangone and Bonvin (2015) further analyzed 122 com‑ plexes with binding afnity data and observed that structure‑based methods such as free energy perturbation and thermodynamics integration could be very accurate, but due to their computational costs, their application is extremely limited. Apart from regression models, QSAR models were also used for relating structural descriptors with binding afnity of protein–protein complexes using structure‑based benchmark datasets com‑ piled by Kastritis etal. (2011) and Zhou etal. (2013). Using the same dataset Marillet etal. (2016) utilized 12 features which account for enthalpic and entropic changes upon binding and devised protein–protein afnity prediction models. Wang etal. used the knowledge‑based potentials and reformulated the binding afnity based on the Monte Carlo algorithm to obtain a Pearson correlation of 0.7 for the prediction (Wang, Su, etal., 2021). FoldX by Delgado etal. (2019) uses an empirical function for predicting the binding free energy between the protein–protein complexes.
In recent years, ML methods have been developed which are faster and more accu‑ rate for predicting protein–protein binding afnity (Li etal., 2022). PPI‑Afnity is a web‑based tool that predicts the binding afnity using support vector machines and other classic ML models (Romero‑Molina etal., 2022). The ML model showed a per‑ formance of r = 0.77 on the SKEMPI dataset (Romero‑Molina etal., 2022). In addition, a few antibody–antigen‑specic binding afnity prediction methods have also been developed in the past few years. CSM‑AB developed by Myung etal. (2022) is a ML method capable of predicting antibody–antigen binding afnity by modeling interaction
7 • Antibody Structure-Function 179
interfaces as graph‑based signatures. It obtained a correlation of up to 0.64 on a blind test dataset. Yang etal. carried out a ML analysis based on interface and surface areas for antibody–antigen complexes. They constructed different models to predict anti‑ body–antigen binding using area‑based and contacts‑based descriptors through con‑ structing and training different predictive models. They obtained the best correlation of
0.85 (with 33 antibody–antigen complexes) and 0.74 (with 262 antibody–antigen com‑ plexes). Their results showed that the area‑based descriptors are slightly better than the contacts‑based descriptors in terms of predictive power; the new models specic for antibody‒protein antigen binding afnity prediction are superior to the previously used general models for predicting the protein–protein binding afnities; and the per‑ formances of the best area‑based and contacts‑based models are better than the perfor‑ mances of the graph‑based model (i.e., CSM‑AB) specic for antibody–antigen binding afnity prediction (Yang etal., 2023). Recently, Sharma et al. developed a model for predicting the binding afnity for SARS‑CoV‑2 spike protein and neutralizing antibod‑ ies using 29 antibody–antigen complexes. They obtained a correlation of 0.90 for the jack‑knife test on SARS‑CoV‑2 protein data bank structures (Sharma etal., 2022).
7.2.4.2 Change in binding afnity upon
mutation prediction methods
Mutations in a protein cause changes in its structure, function, interactions, and binding afnity, which can lead to disease (Gromiha et al., 2016). Several methods have been developed over the years for the prediction of change in binding afnity upon mutation utilizing sequence, structure, and energy‑based features as well as a combination of them. BeAtMuSiC is a coarse‑grained predictor of the changes in binding free energy induced by point mutations. It is based on a set of statistical potentials derived from known protein structures and integrates the mutation’s effect on: (i) the strength of the interactions at the interface, and (ii) the overall stability of the complex (Dehouck etal., 2013). The cor‑ relation obtained by BeAtMuSiC with 90% of the SKEMPI dataset (including antibody– antigen complexes) was 0.70. Brender etal. used random forest training and combined interface structure prole scores with residue‑level coarse‑grained potentials to develop a composite predictive model. They obtained a correlation of >0.8 between the predicted and observed binding free energy changes upon mutation using the SKEMPI database (Jankauskaitė etal., 2018). The single amino acid mutation‑based change in binding free energy (SAAMBE) method developed by Petukh et al. (2015) took advantage of both sequence and structure‑based methods and utilized structure minimization, statistical energy scoring functions, and modied molecular mechanics energies combined with the Poisson–Boltzmann surface area continuum solvation (MM‑PBSA) for estimat‑ ing the effect of single and multiple mutations on binding afnity. The mmCSM‑PPI, developed by Rodrigues et al. (2021), predicts binding afnity change upon mutation using graph‑based signatures, which describe the distance patterns between atoms on the binding interface (Rodrigues etal., 2021). Another method that depends on graph‑based signatures is mCSM‑AB, which is specic for predicting antibody–antigen afnity changes upon mutation (Pires & Ascher, 2016). Methods such as GeoPPI, TopNetTree, and PerSpect‑EL are based on neural networks and persistent homology (Liu etal., 2021; Wang etal., 2020; Wee & Xia, 2022). All these methods used the SKEMPI dataset
180 Biopharmaceutical Informatics
and showed a correlation of up to 0.85 between predicted and experimental data. FoldX software can also evaluate the effect of mutations on protein stability, interaction, fold‑ ing, and dynamics using structures (Delgado etal., 2019). Jemimah et al. developed ProAfMuSeq, a method that predicts protein–protein binding afnity change upon mutation using sequence‑based features and functional class. It shows a correlation of
0.73 and a mean absolute error (MAE) of 0.86 kcal/mol in cross‑validation (Jemimah etal., 2019). Another sequence‑based predictor is PANDA that could predict a change in protein binding afnity upon mutation with a correlation coefcient of 0.52 (Abbasi etal., 2021). SAAMBE‑SEQ is based on a gradient‑boosting decision tree ML algorithm. It utilized 80 features representing evolutionary information, sequence‑based features, and change of physical properties upon mutation at the mutation site and achieved a Pearson correlation coefcient (PCC) of 0.83 (Li etal., 2021).
7.2.5 Biophysical Parameters Affecting
Antibody Design
There are certain biophysical parameters which need consideration for the optimization of binding afnity, specicity, and developability of an antibody (Khetan etal., 2022). These parameters include variations of protein features such as charge, pH/isoelectric point, hydrophobicity, and CDR length and are calculated for either the whole antibody or a specic region of the antibody. The rst developability guidelines were dened using 137 clinical‑stage antibodies, which empirically dene boundaries of antibody drug‑like behavior (Jain etal., 2017). Further, Raybould etal. presented ve biophysical parameters for antibodies similar to the “Lipinski rule of ves” for orally active drugs, which include CDR length, surface hydrophobicity of the near CDR region, and three charge‑based metrics (patches of positive charge, patches of negative charge near the CDR region, and charge symmetry of the surface exposed residues) (Raybould etal.,
2019). More recently, in silico analyses of the variable regions of 77marketed anti‑ body‑based biotherapeutics have revealed ve non‑redundant physicochemical descrip‑ tors, which represent stability, isoelectric point, and molecular surface characteristics of Fv regions (Ahmed etal., 2021). Another study on the same dataset found that anti‑ body specicity is dependent on the net charge of CDR regions, with positively charged CDRs having a higher risk of low specicity than negatively charged antibodies (Rabia etal., 2018). Negatively charged CDRs were also linked with the poor biophysical prop‑ erties of the antibody. Bashour etal. (2024) conducted a computational assessment of 40 sequence‑based and 46 structure‑based developability parameters (DPs) across more than two million native and human‑engineered single‑chain antibody sequences. Their analysis revealed that structure‑based DPs exhibited lower redundancy compared to sequence‑based DPs, suggesting that sequence DPs are more predictable and oper‑ ate within a more constrained design space. Sharma et al. looked into the viscosity and clearance of antibodies (Sharma etal., 2014). They observed that antibody viscos‑ ity increases with the increase in charge dipole distribution and hydrophobicity, and decreases with net charge. On the other hand, antibody clearance correlates with high hydrophobicity of CDR regions and highly positive/negative net charge. An analysis of a small set of FDA‑approved antibodies revealed that the shift in isoelectric point,