Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5660_Библиотеки_им_академика_М_И_Перельмана
.pdf
https://t.me/medicina_free

7
https://t.me/medicina_free
PDB-REDO in Computational-Aided Drug Design (CADD)
Ida de Vries, Anastassis Perrakis, and Robbie P. Joosten
Oncode Institute and The Netherlands Cancer Institute, Department of Biochemistry, Plesmanlaan, 121 1066
CX Amsterdam, the Netherlands
PDB-REDO is both a pipeline that aims to optimize crystallographic macromolecular structure models and a databank oering optimized versions of Protein Data
Bank (PDB) models. The automated decision-making system renes, rebuilds, and
validates the models available in the PDB or provided by users, based on their original experimental diraction data. It returns a new structure model with rich metadata on model quality and structural changes. Optimized PDB models are saved to
the PDB-REDO databank, which contains “redone” structure models with their electron density maps and associated validation data. The PDB-REDO databank is a good
resource for structure models in a computer-aided drug design (CADD) project.
201
7. 1 History and Concepts
7.1.1 X-ray Structure Models
The most commonly used technique to obtain macromolecular structure models is
X-raycrystallography (Figure 7.1). In the crystallographic process, the protein, DNA,
RNA, or complex is crystallized and irradiated with X-rays. This results in a set of
2D diraction images, which undergoes a series of complex operations to determine
the intensity and the associated error of diracted X-rays constrained by the symmetry of the crystal. This results into what we will refer to here as “experimental
data,” a list of intensities and their estimated errors. After the experimental data
are available, crystallographers need to retrieve the missing phases of the diracted
X-rays by computational methods that involve prior knowledge about the nature of
the macromolecular structure, often by collecting additional experimental data. The
experimental data and the phase estimates allow the construction of a 3D electron
density map, which is to construct an initial structure model [1]. Next, this atomic
model is rened using renement software [2–4] and validated against targets based
on independent knowledge of the macromolecular structure [5]. Important steps in
this process are dening the parametersfor the renement and judging the quality of
Open Access Databases and Datasets for Drug Discovery, First Edition.
Edited by Antoine Daina, Michael Przewosny, and Vincent Zoete.
© 2024 WILEY-VCH GmbH. Published 2024 by WILEY-VCH GmbH.

202 7 PDB-REDO in Computational-Aided Drug Design (CADD)
https://t.me/medicina_free
Crystal Final modelDiffraction data
Figure 7. 1 Workflow of an X-ray crystallographic experiment to obtain a macromolecular
structure model.
Initial structure
model
Model
rebuilding
Refinement
Validation
the newly obtained model based on the validation. There are various software packages available for the renement and validation [6–8], each with its own strengths
and weaknesses. The rebuilding, renement, and validationsteps should be repeated
because, due to the intricacies of the crystallographic process we briey explained
above, better estimates of the phases are computed and thus better electron density
maps when the atomic model improves. After many iterations, the structure model
cannot be improved any further and is then considered “optimal” [9]. This end-point
remains highly subjective to date [10].
For several decades (50 years at the moment of writing [11, 12]), crystallographers
have uploaded their structure models obtained from crystallography experiments
to the PDB [13]. This databank contains over 180,000 structure models and has
become a key resource for (computational) structural biology, biochemistry, and
drug design with 1.3G downloads in 2020 alone. Experimental techniques and also
the software used to analyze the diraction data continue to develop and improve.
This has resulted in overall more accurate models of macromolecular structures
that were mostly deposited more recently [14]. Usually, crystallographers continue
with other projects and do not update the deposited structure models after PDB
deposition. As a result, especially older structure models are not as accurate as they
could be with the current computational methods. Researchers interested in such
a structure may choose to optimize the structure themselves when experimental
data are available. The latter is not always the case, as only since 2008 the PDB has
made it mandatory to upload the experimental data when depositing a new structure model to this database [5]. Nevertheless, 86% of all X-raydiraction entries have
their experimental data available.As crystallographicskills do not necessarily belong
to the expertise of researchers in CADD, judging the quality of a structure model can
become problematic. For such structural biology research purposes, but also to help
active crystallographers determine better structures, the PDB-REDO procedure and
the associated databank have been developed [15].
7.1.2 PDB-REDO Development
The PDB-REDO databank contains alternative versions of the X-ray crystallographic
structure models deposited in the PDB that were updated with the PDB-REDO
software pipeline using the original experimental data that is also deposited with
most PDB entries. PDB-REDO entries (>155,000) are generally of better quality
than their PDB counterparts in terms of t to the experimental data and molecular
geometry [15]. An additional and crucial benet for CADD approaches is that the

7.1 History and Concepts 203
https://t.me/medicina_free
(methodological) uniformity of the structure models is substantially improved, as
all models are generated using the same pipeline and validated against the same
targets. Both overall model quality and model uniformity have made considerable
steps forward over the course of PDB-REDO development.
7.1.2.1 First Uniformity
The rst version of the PDB-REDO databank contained optimized coordinates of the
structure models as well as the model parameters used in renement for structures
with a resolution of 2.70 Å or better [15]. R-free was used as a model quality indicator and the quality of the model coordinates was veried using WHAT_CHECK [8].
An important aspect of the process was that PDB-REDO optimized in a uniform way
the relative weight of the experimental data and the so-called “geometric restraints.”
The latter consists of a priori expectations of the covalent geometry of amino acids
and other chemical moieties that are found in macromolecular structures and in
molecular simulations terminology can be thought of as a basic “force eld.” At
least at that time, nding the optimal weight between these two factors, “experiment” and “geometry,” has been often up to each user of each software package.
Crucially, PDB-REDO was not only choosing the software package for optimization,
but was also proposing an objective algorithm for determining this weighting factor.
A key in the re-renement process that was done for each entry was the uniform
use of translation, liberation, and screw (TLS) displacement models [16] during the
renement [17]. The TLS models work on groups of atoms that behave as rigid bodies and provide a layer of information to the atomic displacement parameters. These
parameters describe anisotropic movement by adding only 20 model parameters per
group and without changing the characteristics of the input PDB entry in terms of
e.g. ligands, rotamers, and amino acids. This was then followed by general model
renement that netuned the atomicpositions and B-factors. Although the structure
models improved in terms of t to the X-ray and to other model quality indicators,
more gross modeling errors (e.g. side chains out of density) were not handled in this
rst version of the PDB-REDO databank. Manual inspection and adjustments were
still required to resolve these errors [18].
7.1.2.2 Automatic Rebuilding of Protein Backbone and Side Chains
In further development of PDB-REDO, the programs pepip and SideAide were
adapted from the ARP/wARP package [19] and implemented in the pipeline [20].
These two were the rst fully automated tools to systematically check, correct, and
improve the protein with respect to the electron density maps if the electron density
maps indicate that this is needed. As the name implies, pepip systematically checks
for all peptide planes (i.e. the planes consisting of the Cαi, Ci, Oi, N
backbone atoms) outside the cores of α-helices and β-strands whether alternative,
ipped orientations improve the t to the electron density as well as the position of
the two involved residues on the Ramachandran plot. If so, the adjustment is kept,
as it is considered to be an improvement of the model. SideAide was implemented
to optimize amino acid side chain conformations. This tool searches for the rotamer
conformation of a side chain that shows the best t to the electron density map.
, and Cα
i+1
i+1

204 7 PDB-REDO in Computational-Aided Drug Design (CADD)
https://t.me/medicina_free
During this process, the correctly modeled parts of the structure model are kept
rigid, so that no interference occurs. Furthermore, the Cα atoms are allowed to shift,
increasing the sampled search space and thus the detection rate. The best rotamer is
selected for each amino acid, after which renement of the structure model against
the electron density map is performed. Subsequent validation indicates whether
the change in rotamer have indeed improved the structure [20]. Additionally, side
chains of histidine, glutamine, and asparagine are ipped if this improves the local
hydrogen bonding network.
It should be noted that PDB-REDO by default completes missing side chains in
structure models even when the electron density is relatively poor. Although other
methods of dealing with poor side chain density are also used by crystallographers
(e.g. side chain truncation or manipulation of the occupancy of side chain atoms),
the completion of side chains in the most plausible conformation has the advantage that interpretation is still possible even without having the complete structural
model. The positional uncertainty of the added side chain atoms is, to a large extent,
captured in the atomic B-factors.
7.1.2.3 Automated Model Completion Approaches
The addition of missing side-chains was a rst step toward making structure
models more complete. In further PDB-REDO development, Loopwhole was
added to complete loops in protein models. When Loopwhole detects an unmodeled loop in the protein model (based on the deposited sequence), it looks for
homologs of the protein that do contain a modeled loop at that position in the
protein. If “homologous loops” are found, they are transferred to the structure
model of interest by local structural alignment. After real-space optimization,
the loop with the best t to the electron density is retained when it is of sufcient quality in terms of geometry and t to the density map [21]. Besides
auto-completion of the protein part of structures, PDB-REDO also works on
carbohydrates from N-glycosylation. Carbohydrates are often added to proteins
as post-translational modications and are important recognition parameters in
several biological processes (e.g. protein folding). Modeling of the carbohydrates
in protein structure models is often done relatively poorly [22, 23] because it
has received little attention in the past, and interactive tools for handling carbohydrates were not well-developed (that is, handling polysaccharides required
expertise far beyond normal use of the software available) and also because the
task contains particular challenges. The experimental data are commonly less
informative for carbohydrates than for protein and the tools for model building
of carbohydrates are not as well established as those for proteins. Additionally,
many carbohydrates present in structure models are often not the research interest of the depositor [24]. Carbivore was written and added to the PDB-REDO
pipeline to overcome these challenges by automating the extension of existing
carbohydrate trees [25] using, at that time, newly introduced functionality in the
popular model-building software COOT [26, 27]. Carbivore is also able to build
new trees at asparagine residues that are part of the so-called N-glycosylation
sequon Asn-X-Ser/Thr [28]. Additionally, carbohydrates that do not t any known

7.1 History and Concepts 205
https://t.me/medicina_free
glycosylation tree added in vivo, are deleted from the protein structure model
and rebuilt.
7.1.2.4 Systematic Integration of Structural Knowledge
Besides the use of homologous proteins to complete loops in a protein, homolog
structures are also used as an additional source of geometric restraints for the renement of models with low-resolution experimental data [29]. This information is captured by comparing hydrogen bonds in the structure model to equivalent hydrogen
bonds in close homologs (70% sequence identity or better). The mean hydrogenbond
length in the homologs is used to set a distance restraintbetween the hydrogen bonding partners in the model. The standard deviation is used to set the relative weight of
the restraints so that low standard deviations resulting from strong structural conservation and impose tight restrains, while large standard deviations cause loose
restraints. The use of such restraints improves the geometric quality of the structure models [29]. Additionally, the use of such homology-based restraints increases
the consistency of structural homologs, unless there is strong signal in the experimental data for structural dierences. This makes it easier for users to assess the
structural eect of moving a protein from one functional state to another, e.g. by
binding dierent ligands.
The addition of new structural knowledge when rening structure models in
PDB-REDO is not limited to proteins. Other specialized restraints are added for
structural zinc binding sites [30] and for base pairs in nucleic acid structures [31].
7.1.2.5 Overview of PDB-REDO Pipeline
The input for the PDB-REDO pipeline (Figure 7.2) is an atomic coordinate le, the
experimental data le, and the sequence. The pipeline checks if all required data
are provided and determines various parameters that are used for the renement
with REFMAC [2]. Next, several parallel renements of the structure model are
executed with dierent combinations of parameters, and the best parameters for
renement are chosen. Then, the new model and electron density map are used for
structure rebuilding with the tools we described above. The rebuilt model is rened
once more, while ne-tuning the parameters from the previous round to obtain the
nal structure, i.e. the PDB-REDO model. This model is validated on overall structure parameters, e.g. R-free and Ramachandran Z-score. Nucleic acids are validated
against their specic parameters and ligands are validated separately as well. The
parameters, tools and software, the PDB-REDO model, maps, and validation data
are all saved in the PDB-REDO databank for existing PDB entries. As a result of the
Model coordinates
X-ray data
Sequence
Restraints
Parameterisation
Refinement
Rebuilding
Additional
refinement
Model
validation
Figure 7. 2 The fully automated decision-making PDB-REDO pipeline for optimization of
macromolecular structure models that were obtained from X-ray crystallographic
experiments. 3D boxes represent parallel computing.
PDB-REDO
model and data

206 7 PDB-REDO in Computational-Aided Drug Design (CADD)
https://t.me/medicina_free
PDB-REDO pipeline, more accurate descriptions of the structure models supported
by the experimentally obtained electron density are obtained.
In order to optimize structure models in an eective and ecient manner,
all the steps in PDB-REDO are fully automated. The tools in the PDB-REDO
pipeline that enable the functionalities described above are tied together with many
decision-making algorithms that together form a so-called “expert system,” i.e. a
framework that tries to mimic what a human expert would do [32]. Of course, the
complexity of this system comes at a price in terms of speed and therefore many
steps use parallel computing. As a result, a typical PDB-REDO calculation takes less
than 30 minutes on a modern workstation whereas any manual approach might
cost hours or even days.
7. 2 Structure Improvements by PDB-REDO
To showcase the benets of using PDB-REDO models, this section discusses specic
examples that illustrate how updated macromolecular structure models can change
the biological interpretation of a structure model. These cases are, therefore, interesting examples of the benets that PDB-REDO can bring to CADD projects.
7.2.1 Parametrization and Rebuilding Effects on Small Molecule
Ligands
Many macromolecular structures contain (small molecule) ligands, (metal) ions,
or cofactors that are involved in protein function or structural integrity. Also,
small molecules can be used to modulate the protein’s function or to mediate
protein–protein interactions. Such compounds are found in X-ray structure
models and are treated distinctly during parameterization and renement in
the PDB-REDO pipeline [33]. Here, we illustrate how the specic and uniform
treatment of such compounds, as well as the possible consequences of rebuilding in
their proximity, can aect biochemical conclusions and thus CADD projects.
7.2.1.1 Re-refinement Improves Ligand Conformation
Type IIA DNA topoisomerases are ATP-dependent enzymes involved in cell growth
and division by changing the coiling of DNA helices. These enzymes are the
targets for antibiotics and antitumor agents. To obtain structural insights into
the mechanism of ATP hydrolysis that is coupled to the topoisomerase function,
the crystal structure of the human type IIA DNA topoisomerase was determined
with AMP-PNP, a non-hydrolyzable ATP analogue, and with ADP [34]. The authors
describe conformational dierences of the ribose groups in AMP-PNP and ADP
leading to dierences in interactions with the protein, notably through hydrogen
bonding. However, re-renement of both structure models in PDB-REDO leads
to conformational changes in the nucleotide, particularly in the ribose of ADP
(Figure 7.3). The PDB-REDO procedure removed most of the dierences in the
ribose conformations and resulted to highly similar binding modes for ADP and

7.2 Structure Improvements by PDB-REDO 207
https://t.me/medicina_free
(a)
(c)
PDB
PDB
(b)
(d)
PDB-REDO
PDB-REDO
Figure 7. 3 PDB-REDO results of the binding site of AMP-PNP and ADP (carbon atoms in
light blue) in human type IIA DNA topoisomerase (grey, PDB entries 1zxm and 1zxn, chain
A), side chains of SER149 and ASN150 are shown (carbon atoms in grey) and hydrogen
bonds as black dotted lines. Electron density maps (blue) are oversampled at 0.5 for clarity,
contour levels: 2.5σ for 2mFo-DFcmap 1zxm, 2.0σ for 2mFo-DFcmap 1zxn, 3.5σ for
difference density (red and green). (a) Model as deposited in the PDB with AMP-PNP.
(b) Model by PDB-REDO with the improved conformation of AMP-PNP. (c) Model as
deposited in the PDB with negative difference density surrounding the O2′and O3′atoms
of ADP. The ribose has a different conformation from that in AMP-PNP. (d) PDB-REDO model
in which re-refinement has led to a change in ribose conformation removing the difference
between ADP and AMP-PNP. The negative difference density surrounding the O2′and O3
atoms disappeared. Figure and all molecular graphics figures below were made w ith
CCP4mg. Source: Adapted from McNicholas et al. [35].
′
AMP-PNP, contradicting the original interpretation of the authors. These models
show the importance of proper renement parameterization as used in PDB-REDO.
7.2.1.2 Side Chain Rebuilding Improves Ligand Binding Sites
Glycogen synthase-2 is involved in the biosynthesis of glycogen, which is one of
the most important energy sources in eukaryotes. The basal state of this enzyme
in complex with UDP was crystallized, which provided new structural insights
into the activation of glycogen synthase-2 [36]. Inspection of the binding site
in the PDB-deposited model shows that the uridine base is held in place by a

208 7 PDB-REDO in Computational-Aided Drug Design (CADD)
https://t.me/medicina_free
(a)
Figure 7. 4 Side chain rebuilding of Phe480 of Glycogen synthase-2 (carbon atoms in grey)
results in improved π–π stacking with UDP (carbon atoms in light blue) in the crystal
structure of the basal state (PDB entry 3o3c, chain A). Electron density maps (blue) are
oversampled at 0.5 for clarity, contour levels: 1.75σ for 2mFo-DFcmap, 3.50σ for difference
density (red and green). (a) UDP binding site as in PDB model. (b) UDP binding site as in
PDB-REDO model, where π–π interactions are observed with Phe480.
PDB
(b)
PDB-REDO
hydrogen bond to the protein backbone and π–π interactions with Tyr492. The
side-chain rebuilding in PDB-REDO has moved the side-chain of Phe480 such that
its contribution to UDP binding is also apparent. It has additional π–π interactions
with UDP causing the base to be sandwiched between the aromatic side-chains
(Figure 7.4). The change in π–π interactions for UDP is recorded in the ligand
validation data of PDB-REDO (3o3c_ligval.json (see Section 7.3.1 for downloading
details)). In chain A of the protein, the number of π–π interactions increased from
3 to 6, with a π–π strength improvement of 0.69 (2.75 in the PDB model, 3.44 in
the PDB-REDO model) as measured from the knowledge-based potential used
in YASARA [37].
Besides changing the rotamers of amino acid side chains, PDB-REDO also completes residues in which the side-chains were left unmodeled. This is illustrated by
the binding site of BRAF V600E mutant co-crystallized with vem-bisamide. This
kinase is an oncoprotein in the mitogen-activated protein kinase (MAPK) signaling
pathwaythat is found mutated in several forms of cancer. In melanoma for instance,
the V600E pathogenic mutant is observed regularly and is often targeted in drug discovery. Vemurafenib is one of the compounds that resulted from such studies but
can lead to so-called transactivation of wild-type BRAF. By chemical linkage of two
vemurafenib molecules, vem-bisamide was developed to overcome this transactivation by forcing BRAF into an inactive dimeric conformation. One of the key residues
in potency of inhibitors based on linked vemurafenib-moieties is Gln461 [38]. The
interaction with this residue is nicely modeled in chain A of the crystal structure of
BRAF-V600E in complex with vem-bisamide. However, in chain B the side-chain of
Gln461 is not modeled, hiding a key protein–ligand interaction (Figure 7.5). In the
PDB-REDO model, this residue Gln461 has been completed and the interaction with
the ligand is made obvious (Figure 7.5).
7.2.1.3 Histidine Flip and Improved Ligand Parameterization
The crystal structure of the E. coli autoinducer-2 processing protein LsrF
was obtained without and with the ligands ribose-5-phosphate (R5P) or

7.2 Structure Improvements by PDB-REDO 209
https://t.me/medicina_free
(a)
Figure 7. 5 Vem-bisamide (carbon atoms in light blue, partially shown) is symmetrically
bound by two copies of BRAF kinase (grey, PDB entry 5jt2). Electron density maps (blue) are
oversampled at 0.5 for clarity, contour levels: 1.25σ for 2mFo-DFcmap, 3.5σ for difference
density (red and green). (a) In the PDB model, key interacting residue Gln461 (carbon atoms
in grey) is only completely placed in one of the BRAF chains. (b) The missing side chain of
the second Gln461 has been completed by PDB-REDO revealing the full interaction of
VEM-BISAMIDE with BRAF.
PDB
(b)
PDB-REDO
ribulose-5-phosphate. These structures lead to the strong suggestion that LsrF
belongs to class I aldolases, which are involved in maintaining bacterial expression
of specic genes by catalyzing the formation or cleavage of C–C bonds [39]. The
HSSP multiple sequence alignment [40] of LsrF shows that the binding pocket is
highly conserved with His58 fully conserved among species. Therefore, it is most
likely that this histidine is involved in binding of R5P. However, in the crystal
structure of the autoinducer-2 as deposited in the PDB, His58 does not form any
specic interaction with the ligand (Figure 7.6). As a result of the hydrogen bond
optimization module in PDB-REDO, this histidine residue has been ipped to form
(a)
Figure 7. 6 PDB-REDO improves the binding site of R5P (carbon atoms in light blue) in LsrF
(grey, PDB entry 3glc, chain A) by flipping the His58 side chain (carbon atoms in grey).
Additionally, the overall conformation of R5P is improved. Electron density maps (blue) are
oversampled at 0.5 for clarity, contour levels: 1.5σ for 2mFo-DFcmap, 3.5σ for difference
density (red and green). (a) LsrF structure model as deposited in the PDB. (b) PDB-REDO
model in which His58 has flipped and forms a hydrogen bond with the ligand (black dotted
lines). The negative difference density indicates partial occupancy for ribose-5-phosphate.
PDB
(b)
PDB-REDO
Соседние файлы в папке Библиотека им академика М.И. Перельмана
