Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5419_Библиотеки_им_академика_М_И_Перельмана
.pdf
210 7 PDB-REDO in Computational-Aided Drug Design (CADD)
https://t.me/medicina_free
a hydrogen bond with R5P in the PDB-REDO model (Figure 7.6). Additionally,
the PDB-REDO parameterization improves the overall ligand conformational
t to the density of the R5P ligand. Notably, the parameterization used in the
renement made the B-factor of atoms in R5P more similar to the surrounding
protein atoms. This has resulted in negative dierence density indicating that the
ribose-5-phosphate binding site is only partially occupied.
Although PDB-REDO can improve ligand binding sites and ligand geometry automatically, it cannot reinterpret the ligand to the level of providing an alternative
compound. However,De Souza and co-workers showed that PDB-REDO models are
a suitable starting point for manual reinterpretation of ligands [41].
7.2.2 Building of Protein Loops and Ligands into Protein Structure
Models
Loops are the more exible parts of a protein and tend to give weaker diraction
in crystallographic experiments. This results in poorer local quality of electron density maps and therefore loops are harder to model than other secondary structure
elements such as α-helices and β-sheets. Incomplete protein structure models are
deposited to the PDB, mostly for good reasons: when the experimental data do not
support the modeling of explicit atoms, those should not be added to the model.
However, the decision not to model a loop is invariably a personal one, and some
unmodeled loops can be built with reasonable reliability into poor electron density if
some prerequisites are met. PDB-REDO tries to solve this issue in two ways. First, the
renement steps in PDB-REDO typically lead to an improved atomic model, which
in turn leads to better electron density maps than previously available. Second, the
program Loopwhole reduces the vast number of possible loop conformations to the
ones observed in experimental structures of the same or closely related proteins,
which have a high probability of being correct. Combining both techniques allows
PDB-REDO to add thousands of previously missing loops to PDB models. These
more complete models can then enrich downstream functional or mechanistic interpretations of protein structure. Eventually, this could lead to a better description of
potential binding sites to be targeted in drug discovery [21].
7.2.2.1 Loop Building Completes a Binding Site Region
In the crystal structure of galactokinase from Pyrococcus furiosus, PDB-REDO
builds a substantial part of a loop surrounding the ADP binding site. While this
structure was obtained in complex with ADP, magnesium, and galactose [42],
the surroundings of the ADP binding site in the model as deposited in the PDB
were left unmodeled. There is, however, density observed for this region of the
protein (Figure 7.7). The automatic loop-building algorithm in the PDB-REDO
pipeline improved this part of the protein structure by building the missing loop
(residues 48–72) (Figure 7.7) with a good t to the experimental data. Although
new interactions between the built residues with ADP are not observed in the
PDB-REDO model, the description of the protein is more complete, which is
relevant for structure-based CADD.

7.2 Structure Improvements by PDB-REDO 211
https://t.me/medicina_free
(a)
Figure 7. 7 PDB-REDO builds loop surrounding the ADP (carbon atoms in light blue)
binding site of the Pyrococcus furiosus galactokinase (grey, PDB entry 1s4e, chain B).
Electron density maps (blue) are oversampled at 0.5 for clarity, contour levels: 1.0σ for
2mFo-DFcmap, 3.5σ for difference density (red and green). (a) Model as deposited in the
PDB with unmodeled density and difference density surrounding the ADP binding site.
(b) PDB-REDO model with automatically built loop (residues 48–72, salmon pink ribbons)
in the electron density to complete the chain.
PDB
(b)
PDB-REDO
7.2.2.2 Loop Building Results in Improved Binding Sites
Loop rebuilding in PDB-REDO can improve binding sites as showcased for the magnesium binding site of the geranyl diphosphate methyltransferase in complex with
geranyldiphosphate (GPP) and sinefungin. This protein was used to provide insights
in the methyl-group transfer mechanism, which is a common regulatory process in
living organisms [43]. In the original model, magnesium interacts with the phosphate groups of GPP and with Glu89. However, the magnesium ion lacks an additional coordination ligand (Figure 7.8). This missing interaction could be explained
by the unmodeled residues Val44 and Asn45, which are located in the magnesium
(a)
Figure 7. 8 Addition of residues Val44 and Asn45 results in completing magnesium and
GPP (carbon atoms in light blue) binding site in geranyl diphosphate methyltransferase
(grey, PDB entry 4f86, chain H). Electron density maps (blue) are oversampled at 0.5 for
clarity, contour levels: 1.75σ for 2mFo-DFcmap, 3.5σ for difference density (red and green).
(a) Model as deposited in the PDB with unmodeled residues 44 and 45. Magnesium is
coordinated to GPP and Glu89 (carbon atoms in grey). (b) PDB-REDO model with added
residues Val44 and Asn45 (carbon atoms in grey) resulting in more complete magnesium
coordination.
PDB
(b)
PDB-REDO

212 7 PDB-REDO in Computational-Aided Drug Design (CADD)
https://t.me/medicina_free
binding site. The PDB-REDO loop building adds the two missing residues, leading to a proper magnesium binding site. In addition to the interactions with Glu89
and the phosphates of GPP, the magnesium atom is now also coordinated to Asn45
(Figure 7.8).
7.2.2.3 Building new Compounds into Density
The cytochrome P450 monooxygenase enzymes (CYPs) are important in
post-polyketide modications and thereby cause molecular diversity during
metabolism. Insights into the mechanism of the CYP450 proteins provide valuable
information for drug design, given the involvement of such enzymes in many
metabolic processes, both physiological and xenobiotic [44]. Filipin is often used
as probe for cholesterol binding sites in studies regarding this enzyme. One of
such CYP450-lipin complex obtained by X-ray crystallography is the model
of CYP105P1. Besides lipin, the structure also contains SO4ions and glycerol
molecules that were used during the crystallization process [45]. When inspecting
the model of this structure as deposited in the PDB, some well-dened density
regions have no atoms modeled. Although this is not a feature that is employed by
default, PDB-REDO can build missing compounds if provided a list of candidates.
In the case of the CYP450-lipin complex, glycerol molecules and SO4were tted
(data not shown). Furthermore, one of the modeled water networks in the original
structure (Figure 7.9) indicates that it might be replaced by a dierent compound
present in the crystal and was indeed replaced by this compound in the PDB-REDO
model (Figure 7.9).
These results show that, provided the input model has machine-readable metadata that describes other possible compounds, a model can be made more complete
by automatically tting compounds into the density. This becomes particularly
important if a compound directly inuences the binding pose of a ligand of interest.
An example of this issue was described by Dym et al. [46], who showed that
binding position of methylene blue in acetylcholinesterase was shifted outward by
a poly-ethylene glycol (PEG) molecule sitting at the bottom of the binding pocket.
This eect was overlooked in an early experimental structure model because the
(a)
Figure 7. 9 Original modeled water atoms (red spheres) are replaced by a buffer residue
CSX (carbon atoms in light blue) in de PDB-REDO model of CYP105P1 (grey, PDB entry 3aba,
chain A). Electron density maps (blue) are oversampled at 0.5 for clarity, contour levels: 1.0σ
for 2mFo-DFcmap, 3.5σ for difference density (red and green). (a) Water network as present
in the PDB model. (b) PDB-REDO model in which CXS has replaced the water network.
PDB
(b)
PDB-REDO
CXS

7.2 Structure Improvements by PDB-REDO 213
https://t.me/medicina_free
PEG was not tted in the electron density, which caused conicting results in
downstream structural analyses.
Although crystallization additives and a buer were used as examples, this
approach can be used for ligands and fragments in experimental high-throughput
(lead) drug discovery. This is available as a feature in the PDB-REDO software,
which is currently being tested in real-life experimental settings.
7.2.3 Nucleic Acid Improvements by PDB-REDO
PDB-REDO also applies nucleic acid restraints and validation targets for nucleic
acid-containing structure models. These parameters are limited to the most common structural features: Watson–Crick base pairs [31]. Overall nucleic acids are
improved in both protein–nucleic acid complexes, as well as (mostly) nucleic acid
structures such as ribosomal subunits. One of the rst ribosomal structures that
were deposited in the PDB is the 30S subunit that was used to study its interactions
with antibiotics [47]. The antibiotic paromomycin is bound in such a way that it
is surrounded by nucleotides, among which is G1494 (Figure 7.10). PDB-REDO
restraints improve the orientation of this C–G base pair without changing the
interactions of the nucleic acid structure with the ligand (Figure 7.10). The Z
bgG
a metric describing the relative orientation of the bases, is 3.63 in PDB and 2.04
in PDB-REDO, which is closer to an “ideal” C–G base pair with Z
bpG
= 0.0.
The strongest contribution to this improvement comes from reduced shearing
between the bases. Additionally, the negative dierence density on the ligand has
disappeared.
,
(a)
Figure 7.10 G1494-C1407 base pair (carbon atoms in grey) in the 30S ribosomal subunit
(grey, PDB entry 1fjg) in complex with paromomycin (carbon atoms in light blue). Electron
density maps (blue) are oversampled at 0.5 for clarity, contour level: 1.0 σ for 2mFo–DF
map, 3.5 σ for difference density (red and green). Hydrogen bonds in the G–C base pair are
indicated with black dotted lines. (a) Structure model as deposited in the PDB with
suboptimal hydrogen bonds in the G–C base pair (2.7, 3.1, 3.3Å, shear Z-score −8.85 Å).
(b) Structure model as in PDB-REDO with improved base pair geometry (hydrogen bond
distances 2.8, 2.9, 3.0 Å, shear Z-score -4.43) and the negative difference density on the
ligand drastically reduced.
PDB
(b)
PDB-REDO
c

214 7 PDB-REDO in Computational-Aided Drug Design (CADD)
https://t.me/medicina_free
7.2.4 Glycoprotein Structure Model Rebuilding
Glycosylation is a common post-translational modication of proteins. The modications are important in recognition of other proteins, stability, and formation of
protein complexes [48]. In PDB models, the glycosylated parts of proteins are not
always resolved properly, which indicates that there is sucient room for model
improvement. Initially, PDB-REDO worked on carbohydrates by improving model
annotation and thereby helping model renement to improve the atomic coordinates [49], but much more substantial improvements were achieved when automated (re)building of N-glycans was introduced [25]. The latteris illustrated through
the crystal structure of the binding domain of SARS-CoV-2 spike, which contains a
glycosylated Asn53 [50]. However, the electron density map suggests that the glycan tree can be extended (Figure 7.11). The PDB-REDO model indeed contains an
additional NAG (Figure 7.11).
7.2.5 Metal Binding Sites
Description of metal binding sites is challenging in X-ray modeling software, especially at low resolution. The fact that metal binding sites have an enormous variety in geometries and possible interactors is one of the underlying causes. Many
site geometries are too context-sensitive to reliably predict and restrain in model
renement, but there are common structural motifs around e.g. zinc, magnesium,
and sodium, that are amenable to automated model improvement. The PDB-REDO
pipeline includes a tool (platonyzer) that denes geometric restraints for structural
zinc sites (i.e. ZnCysxHisysites as found in zinc ngers and other motifs) and octahedral sodium and magnesium sites. The restraints impose regular geometry during renement, especially when the experimental data are weak. These restraints
can even recover the right conguration of initially extremely poorly modeled zinc
atoms [30]. The latter is reasonably often present in proteins that are targeted in drug
(a)
Figure 7.11 Extension of a glycan tree (carbon atoms in grey) by PDB-REDO in the
SARS-CoV-2 spike receptor-binding domain bound with ACE2 (grey, PDB entry 6m0j, chain
A). Electron density maps (blue) are oversampled at 0.5 for clarity, contour levels: 1.0σ for
2mFo-DFcmap, 3.5σ for difference density (red and green). (a) Model as deposited in the
PDB that shows unmodeled density close to NAG. (b) PDB-REDO model which contains an
extended glycan; no unmodeled density is observed.
PDB
(b)
PDB-REDO

7.2 Structure Improvements by PDB-REDO 215
https://t.me/medicina_free
discovery research. One such protein is the aspartate transcarbamoylase found in
E. Coli [51], which in humans is an interesting drug target for malaria [52] and cancer [53]. The four cysteine side chains that are supposedly ligands for a zinc atom
are apparently not, but instead, form an irrelevant “trisulde” bond in the structure as deposited in the PDB (Figure 7.12). Platonyzer detects that this zinc binding
site is not chemically correct, and generates restraints that will remodel the site during renement. In this process, the model annotation describing incorrect disulde
bridges is removed. As a result, the zinc atom in the transcarbamoylase has the
four cysteine side chains as ligands and thereby shows a correct coordination state
(Figure 7.12). Additionally, the t to the electron density maps has improved and
both the positive and negative dierence density disappeared.
(a)
(c)
Figure 7.12 (a) and (b) Zinc binding site in aspartate transcarbamoylase (grey, PDB entry
1tug, chain D) is improved by the PDB-REDO pipeline. Electron density maps (blue) are
oversampled at 0.5 for clarity, contour levels: 3.00σ for 2mFo− DFcmap, 3.5σ for difference
density (red and green). (a) Binding site as found in the original model wherein the cysteine
side chains form a trisulfide bridge and the zinc atom does not have a valid number of
ligands. (b) Improved zinc binding site in the PDB-REDO model where the cysteine side
chains coordinate to the zinc atom that is now located in the middle of the binding site.
(c) and (d) Magnesium binding site of the human type IIA DNA topoisomerase (grey, PDB
entry 1zxn, chain A) is improved. Electron density maps (blue) are oversampled at 0.5 for
clarity, contour levels: 2.00σ for 2mFo-DFcmap, 3.5σ for difference density (red and green).
(c) Magnesium site as in the original model in which water 960 is too far away to properly
coordinate magnesium. (d) Magnesium binding site as in PDB-REDO model where water
960 has moved inward and describes a relevant magnesium site.
PDB
PDB
(b)
(d)
PDB-REDO
PDB-REDO

216 7 PDB-REDO in Computational-Aided Drug Design (CADD)
https://t.me/medicina_free
Besides structural zinc sites, platonyzer also evaluates magnesium and sodium
binding sites with six-fold coordination. If these sites form (distorted) octahedrons,
angle restraints are generated to clean up the site in renement. Revisiting the DNA
topoisomerase described in Section 2.1.1/Figure 7.3, but focusing on the magnesium site, which incorporates the ADP. The magnesium is coordinated to ADP phosphate groups (ADP O3B at 2.09 Å and ADP O2A at 2.01 Å), Asn91 OD1 at 2.01 Å,
and two water molecules (H2O 902 at 2.03 Å and H2O 958 at 2.24 Å). The sixth ligand, H2O 960, is improbably far away at 2.87 Å (Figure 7.12). In the PDB-REDO
model, this water moves closer to the magnesium (1.99 Å), which makes the site
more chemically realistic (Figure 7.12). This example illustrates the advantages of
the PDB-REDO pipeline comprehensiveness: because both the ADP binding site and
the magnesium coordination are improved, the user gets a clearer perspective on the
structure–function relationship of DNA topoisomerase.
The identication of metal binding sites and metals is still a work in progress.
Especially for sodium and magnesium whose binding sites are hard to distinguish
in X-ray diraction data as these ions have the same number of electrons. The
coordination distances can help, if not directly used as restraints in renement. The
octahedral restraints from platonyzer help to unbias coordination distances, but
the PDB-REDO pipeline does not change modeled ion identities yet. The decision
about the ion identity thus still lies with the user.
Overall, there is a lot of room for improvement in the renement of metal
binding sites, and using metal validation tools such as CheckMyMetal [54] and
MetalPDB [55, 56] is strongly recommended. Nevertheless, the way the PDB-REDO
pipeline takes care of (transition) metal atoms, can provide better structural starting
points for CADD projects.
7.2.6 Limitations of the PDB-REDO Databank
In the examples above, we discussed dierent structural aspects that are addressed
by the very ecient PDB-REDO pipeline, which can have a substantial impact
on the total structure model. Nevertheless, the PDB-REDO comes with some
limitations. The automation of the decision-making pipeline is a “means to an
end” but it is not the solution to all problems. No systematic manual curation is
performed, which means that not all model errors are removed and new errors
may be introduced by PDB-REDO. The software is designed to be conservative
and to make only few mistakes, but with more than 150,000 structure models with
tens of millions of residues in total, problems are unavoidable. The metadata of
each PDB-REDO entry is designed to allow users to make an informed choice on
which models are most suitable for their downstream studies. For detailed studies
involving a few structure models, we recommend that users inspect the models
carefully in the context of the provided electron density maps. For large-scale
studies, ltering models by overall quality indicators such as R-factors and possibly
local density t metrics such as RSCC is important to construct the most suitable
dataset [57].

7.2 Structure Improvements by PDB-REDO 217
https://t.me/medicina_free
Apart from the general problems linked to automation, there are few issues that
need to be addressed separately. Firstly, some are regarding model annotation.
PDB entries carry a lot of information that aects the way they are dealt with in
PDB-REDO. Examples are descriptions of which (macromolecular) compounds
were crystallized, which parts are modeled, residue and atom nomenclature,
alternate conformers, R-factors, space groups, data quality indicators, et cetera.
When these annotations are severely incorrect, the PDB-REDO process will
either fail completely or will give very poor results. During PDB-REDO databank
maintenance, errors like these are analyzed and if they can be solved by model reannotation, update requests are sent to wwPDB annotators. This process developed
into a fruitful collaboration with the wwPDB by which many small and large issues
in structure models have been solved at the source (i.e. the PDB) so that everyone,
not just PDB-REDO users, can benet.
A second issue is that not only new macromolecular structures but also new small
molecule compounds are added to the PDB every week. Renement of such compounds requires restraint targets that are not immediately available in the CCP4
monomer library [58], which is the key restraint source of the PDB-REDO pipeline.
In such cases, restraints are generated “on the y” based on the current atomic coordinates. This can lead to suboptimal restraints, which in turn lead to suboptimal
molecular geometries. We collaborate with CCP4 developers to regularly update the
CCP4 monomer library so that improved restraints become available to PDB-REDO
users [59]. When users notice compounds with poor geometry in the PDB-REDO
databank, brought on by poor restraints, they can request an update of the aected
entry by clicking a link on the PDB-REDO entry page. Alternatively, they are welcome to contact the PDB-REDO developers directly.
A third issue that combines the problems of model annotation and geometric
restraints is related to intermolecular linkages. Although the polymeric linkages
between amino acids, nucleotides, and recently also saccharides [60] are well
standardized by model annotation at the wwPDB and in the CCP4 monomer library
[59], covalent bonds between ligands and the macromolecule are still a signicant
challenge. Generating the correct geometric restraints requires more information
than is currently stored in structure model les. Currently, only which atoms are
bound is stored, but not how they are bound. Changes in chemistry of the parent
compounds, i.e. deleted atoms, changes in atom hybridization, and changes in
bond orders, are not stored. This makes it challenging to generate correct restraints
for covalent linkages without manual intervention [61]. Unfortunately, this phenomenon can be observed in some PDB-REDO entries and therefore users are
advised to be vigilant when dealing with structure models that have such covalent
linkages. Improvements to the standard data model used for macromolecules, the
mmCIF format [62], are required to solve this issue permanently.
A nal issue of note is technical. Due to the many practical limitations of the PDB
le, this most commonly used data format for structure models, is being replaced
by mmCIF. Most notably the size restrictions for models (99,999 atoms in one le,
with 62 chains) and the limited extensibility to capture new metadata have led to
this replacement. The mmCIF le format does not have these limitations and has

218 7 PDB-REDO in Computational-Aided Drug Design (CADD)
https://t.me/medicina_free
been therefore chosen as the current standard for model handling in the PDB [63].
At the same time, the legacy of structural biology software, including software for
CADD does not (fully) support the mmCIF format yet. This is also true for some of
the software in the PDB-REDO pipeline. This means that even though PDB-REDO
can use mmCIF formatted les as input and output, models are still described in PDB
format within the pipeline. Because of this, 400 very large structure models could not
be processed and are currently missing from the PDB-REDO databank. These will be
added once all the software in the PDB-REDO pipeline becomes mmCIF compliant.
7. 3 Access the PDB-REDO Databank and Metadata
7.3.1 Downloading and Inspecting Individual PDB-REDO Entries
All the models that are parsed through the PDB-REDO pipeline are stored in a databank that can be accessed at https://pdb-redo.eu/. The whole databank can be downloaded, but conveniently it is also possible to download a single entry. The latter can
be done either through the website, or users can open the desired structure model
directly in molecular graphics software YASARA [37], COOT [26], or CCP4mg [35].
These software packages have the utility to download and show the PDB-REDO
model and corresponding density maps from their interface [64].
When using the PDB-REDO website, an entry page visualizing the metadata of
the structure model is provided. An example for PDB entry 1lf2 [65] is shown in
Figure 7.13. On top of the entry page, a table (Figure 7.13) provides crystallographic
data such as space group, resolution, and R-factor of the structure model. Also, the
links to download the PDB-REDO data for the structure are provided in this table.
An additional table (Figure 7.13) containing metrics about the validation of the
crystallographic renement and model quality is provided. This table also indicates
signicant improvements or deteriorations of the PDB-REDO model compared
to the model as deposited in the PDB, conveniently marked in green or red,
respectively. Next, a Kleywegt-like plot is provided (Figure 7.13), showing changes
in model geometry before and after PDB-REDO in terms of ϕ- and ψ-angles.
The Ramachandran Z-scores are provided for both the PDB and PDB-REDO
models, as well as details regarding residues in the preferred regions, allowed
regions and outliers. In the subsequent panel (Figure 7.13) boxplots comparing
the Z-score of the Ramachandran plot for both the PDB and PDB-REDO model
with at least 1000 other structures that were obtained at similar resolutions are
shown. This comparison is also provided for R-free and the rotamer quality as box
plots. Additionally, a table (Figure 7.13) containing the signicant model changes
caused by PDB-REDO is provided. This table provides the counts of these changes,
among which the number of rotamers changed, the number of side chains that
were ipped, chiralities that were xed, and also the improvement or deterioration
of t to the density. When a PDB-REDO entry is loaded into COOT, a button list is
provided to users to quickly inspect all the changes in the structure model in 3D.

7.3 Access the PDB-REDO Databank and Metadata 219
https://t.me/medicina_free
(a)
(b)
(d)
(c)
(e)
Figure 7.13 PDB-REDO entry interface for PDB entry 1lf2, which contains Plasmepsin 2, a
potential antimalaria drug target bound to an inhibitor that can be found at https://pdbredo.eu/db/1lf2. (a) Table containing crystallographic data. (b) Table containing the
validation metrics used in PDB-REDO, while comparing to the model as deposited in the
PDB. Green and red boxes indicate significant improvement or deterioration in the
PDB-REDO model, respectively. (c) Kleywegt-like plot displaying changes of ϕ- and
ψ-angles as a result of redoing the structure model. (d) Boxplots illustrating the model
quality of the PDB and PDB-REDO models compared to models of similar resolution.
(e) Table indicating the significant model changes obtained in the PDB-REDO model
compared to the original model.
Соседние файлы в папке Библиотека им академика М.И. Перельмана
