Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5908_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
15.09.2026
Размер:
16 Мб
Скачать
☆
Use of Molecular Simulations to
8
Understand Structural Dynamics of Antibodies
Daniel A. Nissley, Matthew I. J. Raybould, Charlotte M. Deane, and Sandeep Kumar
8.1 WHY RUN MOLECULAR
SIMULATIONS ON ANTIBODIES?
Gaining insight into the structure‑function relationships that underlie antibody‑antigen binding is a critical step in lead development. While static structures provided by com‑ putational predictions
9,10
phy
(MX) provide a wealth of information to direct, for example, mutational studies to optimize antigen binding and limit off‑target interactions, they only ever provide single snapshots of the ensemble of structures that an antibody populates in solution. Furthermore, the MX structures themselves might contain deviations from the structure that would be found in the solution state, as they are obtained under conditions that are dissimilar to the environment they would encounter in a patient or in transport.9
1–8
or experimental methods like macromolecular crystallogra‑
201
202 Biopharmaceutical Informatics
Forexample, almost all MX structures are collected at cryogenic temperatures (~100 K) to minimize radiation damage.11 Cryogenic electron microscopy (cryo‑EM) is another experimental method of imaging biomolecules under more native‑like conditions.12 However, the low throughput of cryo‑EM experiments and a practical resolution thresh‑ old of ~3 Å13 currently limit their application in the drug development pipeline.
Diverse structural questions may be of interest during antibody drug development. Frequently, these are related to the antibody‑antigen binding event: which of the con‑ tacts present in the MX structure remain when it is heated to room or body temperature? Of those contacts that remain, which contribute most strongly to the overall free‑energy change upon binding? Over time, how do the two binding arms reposition with respect to one another in solution? How does a given mutation in a complementarity‑determin‑ ing region (CDR) loop alter the kinetics of binding? Each of these questions requires dynamic information about antibody structures that is difcult or impossible to access from the experimental structures alone.
Other questions may relate to the developability of the antibody, such as predicting its propensity to aggregate or self‑associate, or its response to changes in salt content made during formulation. These solution properties would ideally be assessed at an early development stage before they could become problematic. Complicating the issue, experimental and theoretical studies indicate that factors difcult to predict from either sequence or structure alone, including the anisotropy of electrostatic interactions and the characteristic oligomeric state of the antibody in solution, strongly inuence these properties.
As described in the following sections, the careful application of molecular simula‑ tion techniques can provide suggestions or even answers to these and other common issues that arise throughout the development of an antibody therapeutic. Before we consider these case studies, we rst step through the most common molecular simulation techniques and important considerations when designing a simulation study of antibody biophysics.

8.2 COMMON TYPES OF MOLECULAR SIMULATIONS FOR BIOMOLECULES

8.2.1 Molecular Dynamics (MD) Simulations

For biomolecules, molecular dynamics (MD) simulations are perhaps the most widely used molecular simulation technique today. MD is a mature, broadly used technique as recognized by the 2013 Nobel Prize in Chemistry awarded to Karplus, Levitt, and Warshel for the development of multi‑scale modeling techniques for complex chemical systems, including MD. In this section, we describe in general terms the theory of MD simulations.
The main goal of MD is the prediction of the time evolution of a molecular sys‑
14,15
In our application, the system likely consists of an antibody and perhaps its
tem. antigen in aqueous solution in a patient. We choose to model the system using the rules
8 • Antibody Structural Dynamics 203
T
E EX
()
=
E
bonded
E
non‑bonded
E EE=+
bonded non‑bonded
E EEE=++
bonded bonds angles torsions
E EE=+
non‑bonded VDW el
E
bonded
E
non‑bonded
FE=−∇
of classical (Newtonian) mechanics and represent each particle as a sphere and each bond as a spring. At the beginning of the simulation, each particle in the system is ini‑ tialized with some starting position and velocity. Initial particle velocities are often ran‑ domly selected from a Maxwell‑Boltzmann distribution generated at the user‑dened absolute temperature of the system,
. A mathematical function, termed a forceeld or Hamiltonian, E, determines the forces different particles in the simulation exert on one another. Despite their name, forceelds are typically written as expressions that, given system particle coordinates X, compute the potential energy of the system The specic terms within the forceeld are chosen to match quantum‑mechanical cal‑ culations and, in most cases, experimental data. Typically, potential energy expres‑ sions contain terms that constrain particle motions through both bonded terms, representing constraints on particle motions due to bonds, angles, and torsions, and non‑bonded terms,
, representing non‑covalent interactions like Van der
Waals (VDW) and electrostatics:
.
,
Equations (8.2 and 8.3) provide a typical decomposition of the
and
(8.1)
(8.2)
(8.3)
terms in Equation (8.1) into their component parts. Some typical mathematical func‑ tions used to represent bonds, angles, torsion angles, VDW, and electrostatic interac‑ tions are shown in Figure8.1a–e.
With positions and velocities determined, the forces acting on particles are then
computed as the negative gradient,
, of the potential energy. Having deter‑ mined the forces, Newton’s equations of motion are then used to predict the positions and velocities at some future time of the system. To reduce numerical integration errors, the time step must be small; values on the order of 1 to 10 femtoseconds are typical for MD simulations of biomolecules. By stepping through time and updating the positions and velocities of all particles in the system according to the forceeld, we produce a movie in 3D space of the time evolution of the system. Various meth‑ ods for maintaining the system temperature during the simulation by updating or rescaling particle velocities have been developed.
16–19
Analysis of the resulting “tra‑ jectory” of the system based on the coordinates collected during the run can then be carried out.
Several different types of MD are frequently encountered in the academic liter‑ ature, all of which fall under the umbrella of MD techniques. In Langevin dynam‑ ics,18additional drag and random collision terms are added to Newton’s equations of motion to approximate the inuence of solvent. In Brownian Dynamics,20 which may be considered a simplied form of Langevin dynamics that applies to certain particles in solution, the equations of motion are modied such that there is no average acceleration in the system.
204 Biopharmaceutical Informatics
FIGURE8.1 Common mathematical forms of MD forceeld terms are shown in panels A through E.
87
Each equation is accompanied by a plot generated with sample forceeld
parameters.
(Continued)
8 • Antibody Structural Dynamics 205
E
bond
k
b
r
0
k
kcal
rÅ
0
=
k
θ
0
k
kcal mol
600=°
E
torsion
k
ϕ
k
kcal
mol
=π
n 1, 2,3
{}
=
6
=
kcal mol
i
r
ij
i
1=+
j
1=+
i
1=+
j
1=−
(Continued) (a) In the expression for
and equilibrium bond lengths. The plot shows results for a bond with
and
. (b) Angles are computed as a function of the force constant
2
,
is a force constant and r and
are, respectively, the current
50
=
b
×
Å
mol
and the squared
2
deviation between the equilibrium (
0.1
=
is shown for
, the multiplicity n (equal to the integer number of minima in the period
,
radians, and
0.1
=
θ
, and its phase
tion of force constant
[−π, π]), the torsion angle
ϕ
) and current () bond angles. The energy as a func-
and
. (c)
is expressed as a function of the
. The displayed plots were generated using
. (d) The Lennard-Jones potential or a slightly
modied form is frequently used for VDW interactions; it computes the non-bonded inter­action between two particles as a function of the distance between the particles, r, the collision diameter for the interaction,
displayed for
Å
and
5
=
Coulombic term; the charges of particles the permittivity of vacuum, and are shown for
and
, and the depth of the potential well, . Results are
. (e) Electrostatic interactions may be treated with a
and j are given by i and j, respectively, 0 is
is the current distance between particles i and j. Results
(blue) and for
and
(green). (f) (Top) Cartoon model of an all-atom antibody Fab in a periodic simulation box. (Bottom) All-atom and CG models of the same Fab are shown. In the CG model, each amino acid is reduced to one representative interaction site.

8.2.2 Monte Carlo (MC) Simulations

Monte Carlo (MC) simulation refers to a broad range of simulation tools in which ran‑ dom numbers are used to perform a search across some phase space.18 The goal of an MC simulation is typically the same as an MD simulation: the prediction of a quantity of interest from the system. In the case of biomolecules, we are typically concerned with exploring the conformational and energetic landscape of our system of interest to predict its conformations or energies. In many cases, a forceeld much like those used in MD simulations will be used to assess the energy of molecular conformations.
MC and MD simulations may also be used within the same protocol, such as in the enhanced sampling technique of replica exchange.21 In replica exchange, a set of identi‑ cal replicas of a chemical system are initialized, each at a different temperature. After a short MD simulation, exchanges are attempted between “neighboring” replicas that are at adjacent temperatures. The probability that the exchange is accepted, w, is calculated based on the energy and temperature differences between the two replica conforma‑ tions, and a random number, is accepted and the replicas are swapped between temperatures. By performing many thousands of these exchanges, the replicas undergo a random walk in temperature space. This random walk helps the simulation avoid spending most of its runtime in local potential energy minima and thereby reduces the total time required to fully explore the energy landscape relative to a single‑temperature simulation. of the multitude of MC methods that may be applied to molecular simulations.
, is generated on the interval [0,1]. If
22
This is but one example
, the exchange
206 Biopharmaceutical Informatics
We note that the key difference between MC and MD simulations is that subsequent structures generated by an MC method are not time‑correlated with one another as they are in an MD trajectory.

8.2.3 Challenges of Molecular Simulations

As with all models, molecular simulations have certain limitations that must be consid‑ ered. As described above, the goal of molecular simulation is normally to predict the time evolution or ensemble average of some property for a chemical system of inter‑ est. Due to the use of random numbers when generating particle velocities (as well as numerous other technical factors23), individual MD trajectories, even when initiated from identical starting coordinates, quickly diverge from one another. The use of ran‑ dom numbers in MC simulations likewise leads to divergent behavior between runs. This leads to questions such as: which of these simulations can be considered correct? Are they all reasonable predictions of behavior?
The random nature of the simulations means that each can be thought of as simi‑ lar to one observation from a single‑molecule experiment–each is individually valid, but the average (ensemble) behavior of the system only becomes clear in the limit of many simulations/observations. In practice, this means that multiple simulations are run and the conformations sampled across them are averaged together. If we are trying to compute the average radius of gyration, we compute the radius of gyration of all of the different conformations of the molecule from each simulation and aver‑ age them. Determining if enough statistically independent simulations of sufcient length have been run to collect a representative sample of possible conformations is a key issue in the practical application of molecular simulations. Obtaining sufcient sampling over the different possible conformational states/trajectories of the system to achieve converged results is time consuming and expensive, especially for larger systems.23 Some systems are simply too large to achieve converged results,24mean‑ ing that reliable answers cannot be obtained without reducing the complexity of the calculation. Below we describe methods that have been developed to overcome these sampling problems.
Molecular simulations also require starting structures. In most cases, these are MX structures deposited in the Protein Data Bank25 (PDB). However, many experimental structures have residues with missing side chains or sections in which entire residues could not be resolved, which require careful rebuilding before simulation. Most pro‑ teins solved by MX are small globular proteins or single domains of multi‑domain proteins crystallized in isolation. This means that, in many cases, multiple PDB models must be merged to build a complete model for simulation. With the advent of more accurate structure prediction tools such as AlphaFold protein structure prediction tools,
4–8
the reliance on experimental structures is begin‑ ning to ease. This explosion in predicted structures has opened up exciting new oppor‑ tunities for simulations by providing a wealth of starting structures for previously inaccessible systems.
One shortcoming of classical MD simulations is that they cannot model the for‑
mation and breakage of chemical bonds, meaning that they cannot be used to model
1–3
and various antibody‑specic
8 • Antibody Structural Dynamics 207
enzymatic reactions. Several methods are available to overcome this shortcoming. Multi‑scale modeling methods in which most of the biomolecular system is represented classically and the catalytic region is modeled using quantum‑mechanical methods have been developed,26 and some specialized forceelds can model bond breakage and for‑ mation.27 However, in simulations of antibodies, we tend to be interested in either anti‑ body‑antibody or antibody‑antigen non‑covalent interactions, meaning that the majority of simulations run are classical.
8.3 MODELING PERSPECTIVE: WHY
WE CANNOT SIMULATE EVERYTHING
IN THE REAL SYSTEM
The goal of a molecular simulation is to predict a property of interest for an antibody. In practice, this is achieved by initializing a simulation with some set of particles rep‑ resenting the molecules in the system, dening how they interact with one another, and then by some method sampling the different accessible conformational states of the simulation.
It makes intuitive sense that the most accurate simulation would be one in which every molecule present in the real system and their interactions with one another are represented as accurately as possible. However, the cellular or test tube environment is far too large and complex to be explicitly simulated with current computational power. For example, a 1‑mL aliquot of a 150‑mg/mL full‑length immunoglobulin G 1 (IgG1) monoclonal antibody (mAb) solution contains on the order of 1017 antibody molecules (assuming a molecular weight of 150 kDa28), each of which is composed of ~20,000 atoms. Without even considering the need to add water and cosolutes to our simulation, we can see that the number of atoms exceeds 1021. Complete simulations of a cell‑like environment would be necessarily even more complex, containing each of the thou‑ sands of unique macromolecules and small molecules composing the cellular milieu.
Current computational power limits molecular simulations to an upper bound of 109 particles, though this feat required utilizing 65,000 processors.29 Simplifying assump‑ tions must be applied to reduce the size and complexity of the real system by many orders of magnitude to make up the difference between the ≤109‑atom systems we can simulate and the real system of 1 mL of 150‑mg/mL mAb with >>1021 atoms that we cannot. Some of the most important and frequently employed simplifying assumptions are considered below before we discuss different resolution molecular simulations in detail.

8.3.1 Periodic Boundary Conditions

To avoid needing to explicitly model large volumes of solution, molecular simulations typically apply periodic boundary conditions around a single copy of the biomolecule of interest. These boundary conditions dene an innitely mirrored simulation space, such
208 Biopharmaceutical Informatics
that when a particle’s trajectory causes it to exit the system from one end, it reappears at the other end with the same velocity (Figure8.1f). The volume within the periodic boundary cell is chosen to be sufciently large that the protein cannot interact with itself through the periodic wall during the simulation. Thus, these simulations can be thought of as being run at innite dilution. Limiting the size of the simulation in this way mas‑ sively decreases the number of particles to within the realm of feasibility.
8.3.2 Inclusion versus Exclusion of
Constant Domains
Structurally, antibodies can be broken into the Fab regions at the ends of the two bind‑ ing arms that recognize and bind antigens and the Fc region that is involved in signal‑ ing. Due to the strong emphasis on understanding antibody‑antigen binding, as well as the inherent difculty in crystallizing full‑length antibodies (and larger proteins in general), experimental structures predominantly contain only Fab segments. For example, as of the 1‑Nov‑2022 update, the SAbDab database tures contains 6,029 Fvs, 5,084 Fabs, and 17 full‑length antibodies (all sequence non‑redundant). These numbers indicate that while the entire Fab is present in 84% of structures, intact antibodies are rarely crystallized. Molecular simulation studies are frequently performed using only the Fv and antigen to reduce computational expense. This approximation appears to make intuitive sense for cases in which only the ener‑ getics of antigen/antibody interactions are of interest, as the entire interface is resolved. However, published MD simulations indicate that including the full Fab rather than just the Fv does, in fact, inuence results,32 suggesting that using full Fabs whenever pos‑ sible in simulations should be considered a matter of best practice. Simulations of the properties of solutions of mAbs, however, require a representation of the Fc to account for, at a minimum, the effects of its steric bulk. Simulations of full‑length antibodies have been published, section of the molecule.
33–38
though they are far less common than studies using a reduced
30,31
of antibody struc‑
8.3.3 All‑Atom versus Coarse‑Grain (CG)
Simulations
Simulations in which each atom of the chosen reduced chemical system is explicitly represented are the most common type of molecular simulations performed. All‑atom simulations of this type must use notably short integration time steps of 1 or 2 fs in order to maintain stability during numerical integration. The magnitude of the integration time step is limited by the highest‑frequency vibrations in the system, which for atomic systems described classically are the bond vibrations.39 By constraining covalent bonds containing hydrogen atoms, a time step of up to 2–3 fs may be used. All‑atom methods are considered the standard for accuracy in MD simulations. All‑atom MD simula‑ tions also benet from a plethora of well‑used and documented tools for preparing, running, and analyzing simulations, including Amber, OpenMM,43 and NAMD.44 The popular forceelds like Amber also have many iterations
40
GROMACS,41 CHARMM,42
8 • Antibody Structural Dynamics 209
that are best applied in different situations, making the choice of forceeld an important question. A benet of all‑atom methods is that all of the mainstream protein force‑ elds (e.g., CHARMM, Amber) are transferable, meaning that they can be applied to any protein system without needing to generate custom parameters. As we will see below, reduced‑resolution models frequently include non‑transferable terms that must be parameterized on a case‑by‑case basis. All‑atom simulations are typically run with explicit representations of solvent molecules and ions. Various water models exist, and some are designed to work in concert with specic biomolecular forceelds (for exam‑ ple, the Amber FF14SB protein forceeld45 is optimized to run simulations with the TIP3P water model46).
Reducing the computational expense of all‑atom simulations to allow the simula‑ tion of larger systems for longer timescales while preserving as much of their predictive power as possible is highly desirable. One method to reduce the computational cost of MD simulations is to coarse‑grain (CG) the system by reducing groups of atoms to representative interaction sites
47,4 8
(Figure8.1f). In most cases, this is achieved by start‑ ing with an atomistic structure of the system and using a CG mapping function that determines how groups of atoms are reduced to interaction sites.48 Once a CG mapping has been determined, a forceeld that describes the interactions between CG sites can be used to calculate the forces between them and their time evolution is then predicted with Newton’s equations of motion, just as in all‑atom MD.
Simulations of CG models are faster than all‑atom simulations for several reasons. First, by reducing the number of particles in the system, they signicantly reduce the number of forceeld terms that must be computed at every time step. Returning to our example of a ~20,000‑atom full‑length IgG1mAb, if we CG it to one interaction site per residue, we reduce the number of particles by an order of magnitude to ~1,400 interac‑ tion sites. Reducing the number of degrees of freedom in the system not only reduces the number of forceeld terms to compute but also reduces the roughness of the free‑energy landscape,47 tending to accelerate dynamic processes like protein folding. By increas‑ ing the length scale of the system and reducing the frequency of the highest‑frequency vibrations, CG models can also typically be run with integration timesteps up to 10‑fold larger than all‑atom MD, allowing them to take larger leaps through time and simulate longer timescales faster. Finally, CG models are frequently run without explicit solvent representations, opting instead to represent the solvent implicitly to further reduce the number of particles. Together, these effects mean that CG simulations are often several orders of magnitude faster than all‑atom simulations.
The selection of model resolution is a critical modeling decision that is typically made at the earliest steps of a molecular simulation project. In general, the selection of model resolution should be motivated by the time and length scales of the process of interest. In fact, it is always worth considering whether the time and length scales of the property of interest put it outside of the practical realm of MD simulations. In terms of timescale, all‑atom simulations can routinely access timescales on the order of 10−9 to 10−6 seconds (nano‑ to microsecond), with longer simulations (up to millisecond) possible for small systems or with the heroic application of computational power. For
43,49
example, the benchmarks
for OpenMM v7.7 on a single A100 GPU indicate that for a ~105‑atom system (apolipoprotein A1) one can expect to achieve 429 ns of simulation per day, but with a ~106‑atom system (satellite tobacco mosaic virus), that performance drops to just 32 ns/day on the same hardware. At the far upper limit of the computational
210 Biopharmaceutical Informatics
performance curve, the state‑of‑the‑art Anton 3 supercomputer, which is custom‑built for high‑speed MD simulations, boasts a speed of 100,000 ns/day using 512 nodes to simulate a 106‑atom system.
The speed‑up of a CG simulation relative to an equivalent all‑atom system is not frequently reported in the literature. General estimates suggest CG models tend to be 103 to 104 times more efcient than all‑atom simulations.51 In some cases, however, the improvement can be more drastic; a set of 19 CG models of globular proteins were found to fold on average 4 × 106 times faster than in experiments.52 The specic sources of acceleration in CG simulations are many and differ between models; we direct inter‑ ested readers to the discussion in Refs. 47 and 48.
While CG models enjoy favorable increases in simulation speed, they suffer from a loss of spatial resolution. CG models with one interaction site per amino acid placed at the coordinates of the Cα atom are reduced to a spatial resolution on the order of the average Cα–Cα bond length, 3.8 Å. In many cases when simulating mAbs, resolution is reduced much further, with the mAb represented by perhaps 6, 10, or 12 interac‑ tion sites. These simulations, which may be considered “ultra‑coarse‑grained” as they reduce hundreds of amino acids to single interaction sites, can be used to simulate protein‑protein interactions and predict experimental solution properties. However, a six‑bead model of antibodies is not suitable for the investigation of which residues contribute to antibody‑antigen binding as the individual residues at the interface are not modeled. Given that an innite number of different CG representations may be designed for a given protein, methods to determine the “best” representation have emerged.53 However, in most cases, researchers employ a combination of intuition for the chemical system under investigation and trial and error to achieve a realistic model for the system and parameter of interest. Despite these issues that would appear to limit the accuracy of CG models, when carefully parameterized and used within their limitations, they can accurately predict the dependence of protein properties on osmo‑ lyte concentration,54 the inuence of pH on antibody viscosity,55 and various other properties.
56,57
50

8.4 USES OF MOLECULAR SIMULATION IN ANTIBODY DRUG DEVELOPMENT

8.4.1 Predicting and Understanding Protein‑Protein Interactions in mAb Solutions
Most antibody drugs, such as mAbs, are delivered by sub‑cutaneous injection. The com‑ fort of the patient receiving the injection places certain limits on the volume (1–2 mL58) that can be administered and the viscosity (practical delivery threshold ~20 cP solution, the latter of which determines the injection force required. The injection volume limit and typical dosages dictate mAb solutions with concentrations of >100 mg/mL.58
59
) of the