Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5608_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
02.09.2026
Размер:
21 Мб
Скачать
410
J. Bauer et al.
[27]
Reduced conformational stability,
reduced melting temperature,
increased levels of fragments,
aggregates, and particles,
Light-induced, free
radicals,
metal-catalysis
[2833]
immunogenicity, reduced biological
activity, and coloration
and anionic properties, antibody
self-association, aggregation,
[29,
3439]
structural changes, loss of function
biological activity, and
immunogenicity
[4350]
Increased charge heterogeneity [4042]
Aggregation, adsorption, increased
viscosity, decreased solubility, and
monosaccharides
Air-water interface,
freeze-thaw, shear,
[5154]
loss of function
Aggregation, low solubility, high
agitation,
temperature, and
light stresses
[49,
5559]
viscosity, liquid-liquid and liquid-solid
phase separation, off-target binding,
fast antibody clearance
complementarity
determining regions
(CDRs))
Deamidation Asn, Gln (in theCDRs) Basic pH A decreased pI, altered hydrophobic
Category Mechanism Liability Initiator Effect References
Chemical instability Oxidation Trp, Met (in the
Table 14.2 Most common developability challenges in biologics
Chemical modications,
destabilized domain folds
Isomerization Asp Acidic pH Altered conformational exibility,
Glycation Lys Reducing
Changes in the secondary,
Conformational
tertiary, and quaternary
instability
Nonuniform distributions
of hydrophobic and
charged regions on
molecular surface
destabilized domain folds
structural features of the
native fold
of natively folded mAbs
Colloidal stability Reversible self-association
Non-native aggregation Chemical modications,
14 Biopharmaceutical Informatics: A Strategic Vision for Discovering Developable…
all the stages described in this table. In the next section, we discuss the opportunities that are beginning to mature.
411
14.2 In Silico Assessments ofBiologics inResearch
andDevelopment
14.2.1 Antibody Generation
Puried antigens can be used to generate antibodies against them either by immu­nizing animals (typically laboratory mice, humanized/transgenic mice or other ani­mals like chicken, rabbit, or cows), using hybridoma techniques, or screening of natural and/or synthetic antibody libraries via display technologies such as phage or yeast. Promising hits are selected and validated via antigen binding assays. Currently available methods for antibody generation are almost entirely experimental in nature. Depending on the methods used to generate antibodies against a given anti­gen, it commonly takes several months before an initial set of antibody-based bind­ers becomes available for further investigations and for lead identication. However, computational technologies that have been originally developed for small molecule drug discovery can be also transferred to antibody-based drug discovery. Once fully developed and deployed, these in silico methods will open another set of means to generate antibody binders against a target antigen.
A potential computational tool originates from the strategy of pharmacophore­based screening of small molecules. Transferring the idea to biologics, the molecu­lar surface features of a desired antigen epitope must be known to screen a database of antibody structures, and to identify potential hits based on molecular shape and electrostatic complementarity. The hits can be prioritized by modelling antigen: antibody complexes via protein: protein docking. However, protein: protein docking scoring methods and selection of the “correct” poses remain a topic of debate, although considerable progress is being made [60]. Further, rst antibody binding mutational databases have been established that will help to improve computational afnity predictions [61]. Subsequently, the most promising candidates can be selected for experimental conrmation of binding, and computational optimization of the afnity and developability of the antibody. A crucial part of this innovative approach is the database of antibody structures and molecular models required for the pharmacophore search. Importantly, the database does not require the structures of full-length antibodies, but structures and homology-based molecular models of the variable regions only. As of now, Protein Data Bank (PDB) contains thousands of high-resolution crystal structures of Fragment variable (Fv) as well as fragment antigen-binding (Fab) regions. In addition, several homology-based antibody mod­elling methods have been developed in past few years and, although some issues remain, the eld has matured enough to provide highly accurate antibody models capable of supporting pharmacophore like searches. The second potential approach
412
J. Bauer et al.
focuses on the design of computational libraries for phage or yeast display experi­ments. The availability and growth of large and heterogeneous databases of anti­body sequences and structures provide an ideal starting point for the design of computational libraries [6267].
From the perspective of experimental approaches, fully synthetic human anti­body libraries comprising Fabs that have been selected for biophysical characteris­tics favorable to development have already been created [68]. Thereby, a special focus was placed on the selection of molecules with increased chemical, conforma­tional, and colloidal stability [68]. The idea of optimized antibody libraries for the generation of developable antibodies could be hybridized with de novo computa­tional databases of an extremely large number of diverse combinations of humanoid light and heavy chains [69]. Mutations targeted at specic sequence positions (e.g., CDRs) in the antibody sequences could further expand the library either to make it recognize different antigens, or to optimize its binding afnity toward a given anti­gen. Recently, a generative adversarial network was successfully applied to create a diverse library of novel antibodies that mimic somatically hypermutated human rep­ertoire response [70]. This in silico approach further unraveled the residue diversity throughout the variable region [70] that might be helpful for further computational tools such as CDR redesign. This approach uses a highly developable antibody framework that is altered in the original CDRs, i.e., paratope to recognize a novel antigen. In the last years, signicant advances were made in the design of not only thermodynamically stable but also biologically functional antibodies [71].
Remarkably, rst computational methods offer the possibility to design human­ized antibody variable regions against targeted antigen epitopes de novo [72, 73]. Alternatively, it was shown that the design can also start with a structural model of an antigen: antibody (Ag: Ab) complex generated using molecular docking of the structures of the Ag and Ab [73]. In the next step, the afnity of the antigen toward the antibody could be either varied by randomly introducing sequence variations [73] or selectively re-designed via structure-based approaches. Subsequently, inter­facial residues in the Ab and Ag structures which contribute signicantly toward instability of the Ag: Ab complex could be identied via computational alanine (Ala) scanning. In the next step, the identied residue positions could be scanned for mutations that can increase/decrease the stability of Ag: Ab complex and improve or lower afnity of the Ab toward the Ag [74], as per project requirements. Another attractive alternative for the rational antibody design refers to hotspot graft­ing with CDR loop swapping, which only requires information about the interac­tions with the antigen [75].
14.2.2 Hit Selection andLead Identication
After production of antigen-binding antibodies by immunized animals, hybridoma cells, or phage and yeast display techniques, the variable regions of the antibodies are sequenced, and the binders are validated. The wide variety of hits must then be
14 Biopharmaceutical Informatics: A Strategic Vision for Discovering Developable…
413
prioritized, and the most promising lead candidates identied (LI, lead identica­tion). Consequently, extensive resources are required to experimentally test each hit and conrm antigen binding.
Several bioinformatic techniques can support the prioritization and selection of hits for invitro conrmation of antigen binding as well as lead identication. A commonly used strategy is to cluster the hits into bins of high, medium, and low binding afnity, based on the initial estimates, analyze each bin for the diversity of heavy and light chain germlines followed by diversity of the CDRs, and select mul­tiple but few representatives from each germline pair in each bin for experimental testing. Alternatively, one could directly bin the hits based on the germline pairings and CDR diversity and select multiple but few of them based on their estimated antigen binding. In addition to the antigen binding, the aspect of developability can already be considered early at this point of hit selection by employing computa­tional tools. In a simple application, one can score heavy (HC) and light chain (LC) sequences of hits based on presence of potential chemical degradation motifs, aggregation prone regions (APRs), and T-cell immune epitopes present or overlap­ping with the CDRs of the heavy and light chains. Such scoring schemes can be further optimized by adding different weights based on which CDRs contain these motifs and whether they are present in the at the beginning/end or in the middle of the CDRs. In a more structure-based approach, three-dimensional homology mod­els of all or a subset of hits can be analyzed regarding physicochemical descriptors such as pI, charge, dipole moment, and solvent exposed hydrophobic and ionic patches [22, 24]. In subsequent studies, one or few of the best hits are experimen­tally tested thoroughly for biological function, cross-reactivity across species, non­specic binding, and pharmacology indicators such as serum stability. This process culminates into identication of one or more lead candidates.
14.2.3 Lead Humanization andOptimization
Lead optimization (LO) is conducted once one or more lead candidates have been identied and revalidated for function. During LO, the Fv regions may need to be humanized, if necessary, corrected for post-translational modication (PTM) sites, optimized for afnity, and in the best case, for developability (Fig.14.2). Below it is described, how the LO process of therapeutic antibodies can be supported in every aspect by computational biophysics [76].
Humanization aims for an optimal amino acid sequence in the Fv by converting as many nonhuman residues to human germline residues to decrease the likelihood of immunogenicity and antidrug antibodies (ADAs) which can impact drug efcacy and safety [77, 78]. If necessary, mouse residues essential for binding are identied by back mutations to retain Ag binding similar to the mouse/human chimeric candi­date [79]. A better understanding of the structure-function relationship helps to identify tting templates and back-mutations critical to preserve CDR confor­mations [80]. Consequently, computational protein design methods have been
414
Fig. 14.2 Process of lead optimization (LO)
J. Bauer et al.
successfully applied to efciently increase the humanness of antibodies while main­taining their structural stability [81]. In line with that, a retrospective analysis of a humanization campaign suggested that hotspots in both, the FW regions and Vernier Zones, can affect the antigen binding and thermodynamic stability [82]. State-of­the-art software such as MOE from Chemical Computing Group [83] can facilitate CDR grafting via identication of appropriate human FWs and back-mutations by considering large databases of human sequences [63, 84]. Furthermore, the rating and optimization of humanness can be done via in silico calculations [19, 8489]. Besides that, statistical interference approaches have been created to characterize the statistical distribution of human Fv sequences [90]. Furthermore, bioinformatic studies have shed light on subtle structural differences between the lambda (VL) and kappa (VK) isotypes that need to be considered during (re-)engineering [91]. Structure-guided approaches can help improve the biophysical properties of a thera­peutic mAb by switching from a problematic lambda framework (FW) region to a more stable kappa FW [92].
The humanized sequence(s) are then proceeded with liability engineering cam­paigns. The cumulative recommendations based on pre-formulation assessment, forced degradation studies analyzing PTMs and stability, and in silico assessments can be considered along with CDR germline residues in the engineering design plan. At this time, the power of phage display or other screening technologies can be used to screen a large panel of variants (typically 103–109 Escherichia coli expressed Fabs). A panel of nal lead optimized variants (~50–100) might be for­matted as immunoglobulin Gs (IgGs), expressed, and puried at small scales, and characterized by binding and pre-formulation assessments. Complementary, in silico assessments help to investigate intrinsic differences between the candidates. Various predictive in silico tools are available that help to monitor and guide the redesign of the candidate’s individual weak points that mediate chemical,
14 Biopharmaceutical Informatics: A Strategic Vision for Discovering Developable…
conformational, colloidal, and physical instability (see Sect. 14.3). In addition, comparison of the molecular characteristics of the lead candidate against marketed antibodies allows us to estimate the medicine-likeness, which can be further improved by modication of relevant properties such as the extent and magnitude of surface hydrophobicity and charged patches[24].
The resulting recommendations of single or combinations of distinct amino acid exchanges help to design a highly individual engineering strategy of humanization and optimization that is specically tailored to the mAb candidate. Based on the criteria of the research target prole (RTP), the best LO candidates (~3–6) are then selected for large-scale production to supply material for the next phase. These pro­cesses are mostly applied for the optimization of conventional mAbs but can also be extended for multi-specics along with the additional engineering required to opti­mize the second and/or third Fv or scFv domains. In addition, identifying an optimal multi-specic format that combines the individually optimized variable domains is required to nalize the LO candidate(s).
415
14.2.4 Formatting ofConventional
andNext-Generation Antibodies
After optimizing the Fv portions, the engineering of conventional mAbs continues with the formatting of the Fvs in the desired antibody format. In this step, the Fv is combined with the Fc of a desired IgG isotype. In this phase, engineering of the Fc might be required to adapt the receptor-mediated functions of the mAb such as ADCC, ADCP, CDC, and endosomal recycling [93]. Depending upon the therapeu­tic concept, the design of next-generation biotherapeutics as bi- and multi-specic antibodies might require another intermediate formatting step to assess physico­chemical compatibility of individual specicities with one another, assuring that the multi-specic modalities have desirable developability properties. Individual case studies already reported how structure-based engineering can support antibody for­matting. One case study showed that after converting from scFv to IgG, the afnity of a TGFβ1 (Transforming growth factor β1) binder could be successfully restored by structure-guided reengineering of elbow region [94]. Analogous computational­guided approaches have the power to support formatting of more complex next­generation antibodies such as multi-specic biotherapeutics.
In the last step of the discovery process, a few of the top performing lead variants are assessed in pre-formulation studies prior totransfer to development for cell line generation and early developability assessments to test the t of the nale molecule(s) to the platforms of upstream, downstream and formulation development [4]. The research phase is completed by the selection of the nal candidate for the start of development.
416
J. Bauer et al.
14.2.5 In Silico Assessments inDevelopment
The preliminary stages of drug substance and drug product development are known to be resource intense. Consequently, the full development program comprising of cell line development, upstream and downstream manufacturing process, and for­mulation development can in most cases only be conducted for the nal lead candi­date. However, at the time of selection of the nal lead candidate, experimental data is only sparsely available due to limitations on quantity as well as quality of material available. At the same time, the sequence of the nal lead candidate gets locked at the start of development and not even single point mutations are allowed. This deci­sion inherently puts product development in a disadvantaged situation since real­time data of the candidates’ stability are typically not available at the start of development but are essential to meet the regulatory requirements for shelf-life, CQAs, and product heterogeneity. Therefore, there is a particularly strong demand for an early, fast, and reliable prediction of various stability aspects that can be cov­ered by hybrid approaches of invitro and in silico techniques.
14.3 In Silico Tools forDevelopability Assessments
Over the last decade, the scientic community has developed a diverse set of com­putational tools that can be applied in R&D to assess many different developability aspects of biotherapeutics. In discovery, the result obtained from in silico develop­ability assessments can be considered during the selection of hits as well as identi­cation/optimization and engineering of the lead molecule to produce more easily developable antibody drug candidates for drug product development. In the devel­opment phase, the in silico tools can help to estimate the nal t of the candidate to standardized platforms, identify potential developability issues, guide adjustments from the platforms that might be required, and nally interpret the complex results of experimental studies [95]. The following section provides an overview of in silico tools that computationally characterize therapeutic candidates and predict important developability properties.
14.3.1 Structure Prediction
Some risk factors such as chemical modication sites can be even identied at the level of the amino acid sequence, but many other developability factors such as biophysical properties depend on the three-dimensional structure of the biothera­peutic. Therefore, most biopharmaceutical companies conduct experimental studies as X-ray crystallography or nuclear magnetic resonance spectroscopy to solve the molecular structures of the lead candidates alone or in complex with the respective
14 Biopharmaceutical Informatics: A Strategic Vision for Discovering Developable…
417
target. In case no experimental structures are available at the required time in the project, homology modelling is often used to predict the three-dimensional struc­ture of biologic drug candidates using their amino acid sequence. The modelling abilities have been signicantly improved over the last decade by advances in com­putational power, modelling techniques, and databases of sequences (NGS, next­generation sequencing) and structures [96101]. Since the modelling of the variable regions of the antibodies can be performed automatically in a high-throughput man­ner, it is particularly useful toward discovery stages to analyze large sets of candi­dates [102105]. Homology modelling is especially suited for the prediction of mAb structures since the framework regions are highly conserved [106]. Despite technological advances in template-based, fragment-based, and template-free mod­elling, the greatest challenge remains the prediction of diverse CDR canonical classes. Particularly, modelling of the conformations of HCDR3 loops with signi­cantly varying lengths is challenging due to the sequence variability and structure exibility [103, 107111]. Currently, the most accurate loop models seem to be achieved via hybrid strategies that combine the benets of knowledge- and physics­based approaches [107, 112]. For a further improved understanding of the structural characteristics of the dynamic structure of a mAb in solution, additional molecular dynamics (MD) techniques that can capture antibody uctuations can be applied [69, 113, 114]. Currently, modelling the full-length structures for IgG mAbs, and, for next- generation multi-specicantibodies is challenging because the PDB con­tains only a handful of such crystal structures.
DeepMind’s AlphaFold demonstrated the great potential of deep learning for protein structure prediction [100, 101]. Furthermore, the prediction of exible sys­tems suchas interfaces is still complicated due to the complex balancing of polar and nonpolar interactions as well as solvation effects [97]. At the same time, ML techniques have a great potential to revolutionize the eld of template-free predic­tion of protein structures [97]. Particular attention was attracted by a protein-spe­cic fragment library that has been recently generated via deep neural networks [100, 101]. The analysis identied patterns in protein sequence and co-evolutionary couplings, which have been converted in contact maps [100]. Novel algorithms have further been able to engineer de novo high-order assemblies with therapeutic poten­tial using bioinformatics, even so the design from scratch remains a massive under­taking [98, 115, 116].
14.3.2 Biophysical Properties
The developability of drugs is mostly specied by its biophysical properties. Lipinski’s “rule-of-ve” represented a substantial leap for the discovery and devel­opment campaigns of small molecules since it related calculable physicochemical properties with drug characteristics as solubility and permeability [117]. The com­plexity of biological entities hampered the denition of similar guidelines for NBEs (New Biological Entities) for decades, but a biophysical screening of clinical stage
418
antibodies guided the empirical denition of analogous boundaries [22, 24]. Subsequently, the heterogeneous biophysical measures have been successfully cor­related with sequence features, indicating signicant relationships between the sequence of Fv domains and physicochemical properties that dene the developa­bility of antibodies [21, 23]. For instance, general correlates have been established between mAb polyreactivity and the presence of basic amino acids in the CDRs as well as Gln in HCDR2 and 3 [23]. Furthermore, mAb aggregation and self­association have been correlated with CDR length and content of aromatic residues [23]. However, since the particular relationship between specic amino acids and other physicochemical properties of the mAb such as conformational stability strongly depends on the structural microenvironment, key residues and their modi­cations must be individually evaluated under consideration of the structural con­text [118]. Therefore, another study considered Fv models to establish developability guidelines and develop a Therapeutic Antibody Proler (TAP) [24]. This tool assesses the developability of candidates via ve easily calculablemetricsfor the total length of CDRs, positively and negatively charged patchesas well as hydro­phobic patches in the CDRs, and the asymmetry in the surface charges of heavy and light chains [24]. Similar guidelines remain to be developed for multi-specic anti­bodies that often display suboptimal physical properties impeding the development into therapeutics [119].
J. Bauer et al.
14.3.3 Hydrophobicity
The pharmaceutical industry widely recognizes the relevance of hydrophobicity of mAbs for its developability into therapeutics, for which reason Hydrophobic Interaction Chromatography (HIC) is often conducted to compare the apparent hydrophobicity of mAb candidates toassess downstream risks [6, 11, 22]. Even though an accurate prediction of HIC retention times is challenging since it is inu­enced by the solvent (e.g., salt gradients) and other protein characteristics (e.g., charge) [21], HIC retention times could successfully be correlated with sequence and structure features by using diverse methods asQuantitative Structure Property Relationship (QSPR) modeling or machine learning [120122]. To rectify homol­ogy modeling errors and capture the dynamics of the rather exible formation and breakage of hydrophobic patches, it is recommended to consider conformational sampling or MD simulations [121125].
14.3.4 Solution- andColloidal-State Properties
The solution- and colloidal-state properties are amongst the most important aspects to be considered during discovery and development of a Novel Biologic Entity (NBE) drug candidate, but also amongst the hardest to predict due to multiple
14 Biopharmaceutical Informatics: A Strategic Vision for Discovering Developable…
419
inuencing factors as the individual patterning of hydrophobic and charged resi­dues. However, recent advances in understanding the complex principles of protein solubility enabled the development of predictive tools. First computational tools such as SOLpro and PROSO II were trained on tens of thousands of proteins and successfully demonstrated their ability to predict solubility upon expression with an accuracy of ~75% [126, 127]. Later, web-based tools (e.g., Protein-Sol) have been developed that offer an easy access to the prediction of protein solubility from sequence [128]. Complementary to sequence-based approaches, CamSol calculates a residue-specic intrinsic solubility prole under consideration of structural inu­ences, and offers to optimize the solubility of a candidate by screening for suitable mutations [129]. This feature makes it particularly interesting for the design of anti­body libraries and the selection of lead candidates [130]. CamSol was recently extended to predict the aggregation potential of partially unfolded proteins in tem­perature ramps of molecular dynamics (MD) simulations [131]. For early develop­ment activities, such a tool enables us to assess the candidate’s solubility relative to previous molecules without the need for material and laborious experimentalactivi­ties [132]. An alternative software called SODA estimates changes in the protein solubility via the propensity of the sequence to aggregate (via PASTA) [133] and disorder (via ESpritz) [134], and takes further properties as hydrophobicity and secondary structure (via FELLS) [135] into account [136]. Several computational approaches have already successfully guided the rational design of mAbs with increased solution- and colloidal-state properties [137142].
A currently evolving eld focuses on the prediction of mAb specicity, which refers not only to nonspecic binding but also self-interaction. Recently, a method has been described that predicts the overall specicity of antibodies [26]. Individual and combined sets of chemical rules have been dened that recommend limits of certain amino acids exposed to the surface of the variable region.
While mechanistic tools are most valuable for the screening and minimization of APRs in the hits and leadcandidates during discovery [137139, 143149], the kinetic predictors are helpful to estimate the rate of aggregation which is key during the development of liquid formulations that must meet the regulatory requirements for the shelf life of the drug product [150153]. The kinetic models can be trained by ML on large data sets of combined experimental and sequence/structure infor­mation [150, 152]. Desirable predictors could thereby screen different formulations to identify the optimal composition (pH, salt and excipients) for minimal kinetics.
14.3.5 Isoelectric Point (pI)
The isoelectric point (pI) is an important physicochemical property for mAbs, and it has been shown to correlate with specic developability aspects as thermostabil­ity, viscosity and resistance to HMW formation at low pH [4, 154]. Typically, IgG1s with weakly basic isoelectric points between 8 and 8.5 and Fv isoelectric points between 7.5 and 9 typically display the best combinations of strong repulsive