Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5664_Библиотеки_им_академика_М_И_Перельмана
.pdf
7.2 Structure Improvements by PDB-REDO 215
https://t.me/medicina_free
discovery research. One such protein is the aspartate transcarbamoylase found in
E. Coli [51], which in humans is an interesting drug target for malaria [52] and cancer [53]. The four cysteine side chains that are supposedly ligands for a zinc atom
are apparently not, but instead, form an irrelevant “trisulde” bond in the structure as deposited in the PDB (Figure 7.12). Platonyzer detects that this zinc binding
site is not chemically correct, and generates restraints that will remodel the site during renement. In this process, the model annotation describing incorrect disulde
bridges is removed. As a result, the zinc atom in the transcarbamoylase has the
four cysteine side chains as ligands and thereby shows a correct coordination state
(Figure 7.12). Additionally, the t to the electron density maps has improved and
both the positive and negative dierence density disappeared.
(a)
(c)
Figure 7.12 (a) and (b) Zinc binding site in aspartate transcarbamoylase (grey, PDB entry
1tug, chain D) is improved by the PDB-REDO pipeline. Electron density maps (blue) are
oversampled at 0.5 for clarity, contour levels: 3.00σ for 2mF
density (red and green). (a) Binding site as found in the original model wherein the cysteine
side chains form a trisulfide bridge and the zinc a tom does not have a valid number of
ligands. (b) Improved zinc binding site in the PDB-REDO model where the cysteine side
chains coordinate to the zinc atom that is now located in the middle of the binding site.
(c) and (d) Magnesium binding site of the human type IIA DNA topoisomerase (grey, PDB
entry 1zxn, chain A) is improved. Electron density maps (blue) are oversampled at 0.5 for
clarity, contour levels: 2.00σ for 2mF
(c) Magnesium site as in the original model in which water 960 is too far away to properly
coordinate magnesium. (d) Magnesium binding site as in PDB-REDO model where water
960 has moved inward and describes a relevant magnesium site.
PDB
PDB
(b)
(d)
-DFcmap, 3.5σ for difference density (red and green).
o
PDB-REDO
PDB-REDO
− DFcmap, 3.5σ for difference
o

216 7 PDB-REDO in Computational-Aided Drug Design (CADD)
https://t.me/medicina_free
Besides structural zinc sites, platonyzer also evaluates magnesium and sodium
binding sites with six-fold coordination. If these sites form (distorted) octahedrons,
angle restraints are generated to clean up the site in renement. Revisiting the DNA
topoisomerase described in Section 2.1.1/Figure 7.3, but focusing on the magnesium site, which incorporates the ADP. The magnesium is coordinated to ADP phosphate groups (ADP O3B at 2.09 Å and ADP O2A at 2.01 Å), Asn91 OD1 at 2.01 Å,
and two water molecules (H
and, H
O 960, is improbably far away at 2.87 Å (Figure 7.12). In the PDB-REDO
2
O 902 at 2.03 Å and H2O 958 at 2.24 Å). The sixth lig-
2
model, this water moves closer to the magnesium (1.99 Å), which makes the site
more chemically realistic (Figure 7.12). This example illustrates the advantages of
the PDB-REDO pipeline comprehensiveness: because both the ADP binding site and
the magnesium coordination are improved, the user gets a clearer perspective on the
structure–function relationship of DNA topoisomerase.
The identication of metal binding sites and metals is still a work in progress.
Especially for sodium and magnesium whose binding sites are hard to distinguish
in X-ray diraction data as these ions have the same number of electrons. The
coordination distances can help, if not directly used as restraints in renement. The
octahedral restraints from platonyzer help to unbias coordination distances, but
the PDB-REDO pipeline does not change modeled ion identities yet. The decision
about the ion identity thus still lies with the user.
Overall, there is a lot of room for improvement in the renement of metal
binding sites, and using metal validation tools such as CheckMyMetal [54] and
MetalPDB [55, 56] is strongly recommended. Nevertheless, the way the PDB-REDO
pipeline takes care of (transition) metal atoms, can provide better structural starting
points for CADD projects.
7.2.6 Limitations of the PDB-REDO Databank
In the examples above, we discussed dierent structural aspects that are addressed
by the very ecient PDB-REDO pipeline, which can have a substantial impact
on the total structure model. Nevertheless, the PDB-REDO comes with some
limitations. The automation of the decision-making pipeline is a “means to an
end” but it is not the solution to all problems. No systematic manual curation is
performed, which means that not all model errors are removed and new errors
may be introduced by PDB-REDO. The software is designed to be conservative
and to make only few mistakes, but with more than 150,000 structure models with
tens of millions of residues in total, problems are unavoidable. The metadata of
each PDB-REDO entry is designed to allow users to make an informed choice on
which models are most suitable for their downstream studies. For detailed studies
involving a few structure models, we recommend that users inspect the models
carefully in the context of the provided electron density maps. For large-scale
studies, ltering models by overall quality indicators such as R-factors and possibly
local density t metrics such as RSCC is important to construct the most suitable
dataset [57].

7.2 Structure Improvements by PDB-REDO 217
https://t.me/medicina_free
Apart from the general problems linked to automation, there are few issues that
need to be addressed separately. Firstly, some are regarding model annotation.
PDB entries carry a lot of information that aects the way they are dealt with in
PDB-REDO. Examples are descriptions of which (macromolecular) compounds
were crystallized, which parts are modeled, residue and atom nomenclature,
alternate conformers, R-factors, space groups, data quality indicators, et cetera.
When these annotations are severely incorrect, the PDB-REDO process will
either fail completely or will give very poor results. During PDB-REDO databank
maintenance, errors like these are analyzed and if they can be solved by model reannotation, update requests are sent to wwPDB annotators. This process developed
into a fruitful collaboration with the wwPDB by which many small and large issues
in structure models have been solved at the source (i.e. the PDB) so that everyone,
not just PDB-REDO users, can benet.
A second issue is that not only new macromolecular structures but also new small
molecule compounds are added to the PDB every week. Renement of such compounds requires restraint targets that are not immediately available in the CCP4
monomer library [58], which is the key restraint source of the PDB-REDO pipeline.
In such cases, restraints are generated “on the y” based on the current atomic coordinates. This can lead to suboptimal restraints, which in turn lead to suboptimal
molecular geometries. We collaborate with CCP4 developers to regularly update the
CCP4 monomer library so that improved restraints become available to PDB-REDO
users [59]. When users notice compounds with poor geometry in the PDB-REDO
databank, brought on by poor restraints, they can request an update of the aected
entry by clicking a link on the PDB-REDO entry page. Alternatively, they are welcome to contact the PDB-REDO developers directly.
A third issue that combines the problems of model annotation and geometric
restraints is related to intermolecular linkages. Although the polymeric linkages
between amino acids, nucleotides, and recently also saccharides [60] are well
standardized by model annotation at the wwPDB and in the CCP4 monomer library
[59], covalent bonds between ligands and the macromolecule are still a signicant
challenge. Generating the correct geometric restraints requires more information
than is currently stored in structure model les. Currently, only which atoms are
bound is stored, but not how they are bound. Changes in chemistry of the parent
compounds, i.e. deleted atoms, changes in atom hybridization, and changes in
bond orders, are not stored. This makes it challenging to generate correct restraints
for covalent linkages without manual intervention [61]. Unfortunately, this phenomenon can be observed in some PDB-REDO entries and therefore users are
advised to be vigilant when dealing with structure models that have such covalent
linkages. Improvements to the standard data model used for macromolecules, the
mmCIF format [62], are required to solve this issue permanently.
A nal issue of note is technical. Due to the many practical limitations of the PDB
le, this most commonly used data format for structure models, is being replaced
by mmCIF. Most notably the size restrictions for models (99,999 atoms in one le,
with 62 chains) and the limited extensibility to capture new metadata have led to
this replacement. The mmCIF le format does not have these limitations and has

218 7 PDB-REDO in Computational-Aided Drug Design (CADD)
https://t.me/medicina_free
been therefore chosen as the current standard for model handling in the PDB [63].
At the same time, the legacy of structural biology software, including software for
CADD does not (fully) support the mmCIF format yet. This is also true for some of
the software in the PDB-REDO pipeline. This means that even though PDB-REDO
can use mmCIF formatted les as input and output, models are still described in PDB
format within the pipeline. Because of this, 400 very large structure models could not
be processed and are currently missing from the PDB-REDO databank. These will be
added once all the software in the PDB-REDO pipeline becomes mmCIF compliant.
7.3 Access the PDB-REDO Databank and Metadata
7.3.1 Downloading and Inspecting Individual PDB-REDO Entries
All the models that are parsed through the PDB-REDO pipeline are stored in a databank that can be accessed at https://pdb-redo.eu/.The whole databank can be downloaded, but conveniently it is also possible to download a single entry.The latter can
be done either through the website, or users can open the desired structure model
directly in molecular graphics software YASARA [37], COOT [26], or CCP4mg [35].
These software packages have the utility to download and show the PDB-REDO
model and corresponding density maps from their interface [64].
When using the PDB-REDO website, an entry page visualizing the metadata of
the structure model is provided. An example for PDB entry 1lf2 [65] is shown in
Figure 7.13. On top of the entry page, a table (Figure 7.13) provides crystallographic
data such as space group, resolution, and R-factor of the structure model. Also, the
links to download the PDB-REDO data for the structure are provided in this table.
An additional table (Figure 7.13) containing metrics about the validation of the
crystallographic renement and model quality is provided. This table also indicates
signicant improvements or deteriorations of the PDB-REDO model compared
to the model as deposited in the PDB, conveniently marked in green or red,
respectively. Next, a Kleywegt-like plot is provided (Figure 7.13), showing changes
in model geometry before and after PDB-REDO in terms of ϕ-andψ-angles.
The Ramachandran Z-scores are provided for both the PDB and PDB-REDO
models, as well as details regarding residues in the preferred regions, allowed
regions and outliers. In the subsequent panel (Figure 7.13) boxplots comparing
the Z-score of the Ramachandran plot for both the PDB and PDB-REDO model
with at least 1000 other structures that were obtained at similar resolutions are
shown. This comparison is also provided for R-free and the rotamer quality as box
plots. Additionally, a table (Figure 7.13) containing the signicant model changes
caused by PDB-REDO is provided. This table provides the counts of these changes,
among which the number of rotamers changed, the number of side chains that
were ipped, chiralities that were xed, and also the improvement or deterioration
of t to the density. When a PDB-REDO entry is loaded into COOT, a button list is
provided to users to quickly inspect all the changes in the structure model in 3D.

7.3 Access the PDB-REDO Databank and Metadata 219
https://t.me/medicina_free
(a)
(b)
(d)
(c)
(e)
Figure 7.13 PDB-REDO entry interface for PDB entry 1lf2, which contains Plasmepsin 2, a
potential antimalaria drug target bound to an inhibitor that can be found at https://pdbredo.eu/db/1lf2. (a) Table containing crystallographic data. (b) Table containing the
validation metrics used in PDB-REDO, while comparing to the model as deposited in the
PDB. Green and red boxes indicate significant improvement or deterioration in the
PDB-REDO model, respectively. (c) Kleywegt-like plot displaying changes of ϕ- and
ψ-angles as a result of redoing the structure model. (d) Boxplots illustrating the model
quality of the PDB and PDB-REDO models compared to models of similar resolution.
(e) Table indicating the significant model changes obtained in the PDB-REDO model
compared to the original model.

220 7 PDB-REDO in Computational-Aided Drug Design (CADD)
https://t.me/medicina_free
7.3.2 Data Available in PDB-REDO Entries
A PDB-REDO entry comes with the new model coordinates, maps, and validation
details, but also contains all the metadata that describes the model in the original
state, the re-rened model, and the nal (rebuilt and re-rened) model. All les
available for one entry are shown in Table 7.1. This data are uniform throughout the
whole databank, as the same pipeline has been used to generate all entries. Additionally, the data are formatted in ndable, accessible, interoperable and reproducible
(FAIR)le formats. Both these features are advantages of the PDB-REDO databank.
A PDB-REDO entry contains les with the atomic coordinates for both the original model (as deposited in the PDB) and the PDB-REDO model. In addition to the
PDB format for describing atomic coordinates, PDB-REDO also provides mmCIF
coordinate les of the “redone” structure model. Additionally, the electron density
maps for both the original and the redone structure models are provided. Validation data are available for both structure models. The WHAT_CHECK reports [8]
provide comprehensive model validation data and the JSON les contain valid data
for the macromolecular structure model that can be easily used in data mining for
structure selection. If ligands are present in the entry, a ligand validation le is provided that contains the validation metrics for each ligand in both the original and
the PDB-REDO structure models. If homologous structure models are available, the
PDB identiers and chain identiers of the homologs are reported in the available_
homologs.json le. Furthermore, a Dene Secondary Structure of Proteins (DSSP)
analysis of the PDB-REDO model is provided [40, 66] as well as a separate le containing the model change scores that are used at the PDBe entry pages [67]. For
convenient visualization of a PDB-REDO structure model, a COOT [26] script is provided, which also indicates the model changes made by the PDB-REDO pipeline.
Finally, a PDB-REDO entry comes with some descriptive data of the PDB-REDO
parameters (data.json) and, for the sake of provenance tracking, versions of all the
software used during the process (versions.json).
7.3.3 Usage of the Uniform and FAIR Validation Data
Uniform data are an advantage for selecting multiple models for a CADD project,
e.g. for homology modeling or investigation of a protein family. The models are all
generated with the same pipeline and parameters can easily be extracted from the
(meta)data provided. The JSON format respects FAIR data requirements. For each
JSON formatted le, the corresponding schema is provided on the PDB-REDO website, which includes detailed information about the descriptors. The les themselves
are user-friendly regarding data analyses. Desired values are easily extracted, after
which analyses such as calculating dierences or sorting data can be performed
using any computer scripting language. For example, one could extract the number
of cation–π interactions for all ligands found in the PDB-REDO databank. In the ligval.json le, these counts are stored for each ligand in the model as deposited in the
PDB and for the corresponding PDB-REDO model. Once the values for a PDB-ligand
and PDB-REDO ligand are compared, the dierence between the two values can

Tabl e 7.1 All data files for a PDB-REDO entry.
https://t.me/medicina_free
PDB-REDO data files
Files
regarding Description Typical URL
Atomic
coordinates
Electron
density
Valida tion
data
Descriptive
data
The les describe the structure model in the original, the re-rened, and the nal re-built and rened state. Typical URLs of all the les are shown, using the ctitious
PDB entry 9xyz as an example.
Initial model https://pdb-redo.eu/db/PDB identier/PDB identier_0cyc.pdb.gz
Re-rened (only) model https://pdb-redo.eu/db/PDB identier/PDB identier_besttls.pdb.gz
Re-rened and rebuilt structure model https://pdb-redo.eu/db/PDB identier/PDB identier_nal.pdb
Re-rened and rebuilt structure model with total B-factors
(PDB)
Re-rened and rebuilt structure model with total B-factors
(mmCIF)
Map coecients for the original model https://pdb-redo.eu/db/PDB identier/PDB identier_0cyc.mtz.gz
Map coecients for re-rened (only) model https://pdb-redo.eu/db/PDB identier/PDB identier_besttls.mtz.gz
Map coecients for re-rened and rebuilt structure model https://pdb-redo.eu/db/PDB identier/PDB identier_nal.mtz
WHAT_CHECK report for initial model https://pdb-redo.eu/db/PDB identier/wo/pdbout.txt
WHAT_CHECK report for re-rened (only) model https://pdb-redo.eu/db/PDB identier/wc/pdbout.txt
WHAT_CHECK report for re-rened and rebuilt model https://pdb-redo.eu/db/PDB identier/wf/pdbout.txt
Ligand and ligand interaction data for the initial and
re-rened and rebuilt structure model
Validation data for the initial and structure model https://pdb-redo.eu/db/PDB identier/PDB identier_0cyc.json.gz
Validation data for the re-rened and rebuilt structure model https://pdb-redo.eu/db/PDB identier/PDB identier_nal.json
Fitting and geometry scores for reporting at PDBe https://pdb-redo.eu/db/PDB identier/pdbe.json
COOT script to show model changes https://pdb-redo.eu/db/PDB identier/PDB identier_nal.py (Python) and
PDB-REDO statistics for data mining (crystal parameters,
R-factors, validation scores, etc.)
Homologous structures that are available for the structure
model
DSSP analysis of the re-rened and rebuilt structure model https://pdb-redo.eu/db/PDB identier/PDB identier_nal.dssp
Software versions of all programs used https://pdb-redo.eu/db/PDB identier/versions.json
https://pdb-redo.eu/db/PDB identier/PDB identier_nal_tot.pdb
https://pdb-redo.eu/db/PDB identier/PDB identier_nal.cif
https://pdb-redo.eu/db/PDB identier/PDB identier_ligval.json (only for
entries that contain ligands)
https://pdb-redo.eu/db/PDB identier/PDB identier_nal.scm (Scheme)
https://pdb-redo.eu/db/PDB identier/data.json
https://pdb-redo.eu/db/PDB identier/PDB identier_available_homologs.json

222 7 PDB-REDO in Computational-Aided Drug Design (CADD)
https://t.me/medicina_free
(a)
Figure 7.14 Completion of the Arg22 side chain (carbon atoms in grey) in the crystal
structure of the catalytic subunit in protein kinase A (grey, PDB entry 5otg, chain A) results
in cation–π interaction with ligand (carbon atoms in light blue). Electron density maps
(blue) are oversampled at 0.5 for clarity, contour levels: 1.5σ for 2mF
difference density (red and green). (a) Binding site of ligand with non-complete side chain
modeled for Arg22 as found in the model as deposited in the PDB. (b) Binding site of ligand
with rebuilt Arg22 side chain resulting in cation–π interaction with the ligand as seen in
the PDB-REDO model.
PDB
(b)
PDB-REDO
− DFcmap, 3.5σ for
o
be calculated and sorted in descending order to rank the ligands based on gained
cation–π interactions. One of the ligands on top of this list is the AO8 ligand of the
protein kinase A catalytic subunit from Criteculus griseus (PDB-ID 5otg). This ligand
binds in the ATP binding site of the kinase, most likely to inhibit its function [68].
The AO8 ligand shows no cation–π interactions in the original model (Figure 7.14),
whereas in the PDB-REDO model 7, such interactions are reported. This increase
is due to the completed Arg22 side chain that is present in the binding pocket. The
guanidinium cation is interacting with the π–electrons in the phenyl moiety of the
ligand improving the binding mode with an interaction between protein and ligand (Figure 7.14) that was overlooked. For the sake of completeness, note that, the
boronic acid of AO8is not visible in the electron density. As boronic acids are known
to be oxidatively unstable, most likely oxidation occurred during the experimental
process [69].
7.3.4 Creating Datasets from the PDB-REDO Databank
Next to listing close homologs, PDB-REDO has a more in-depth way of creating
datasets for CADD and method development for structural biology. At https://pdbredo.eu/archive users can search the PDB-REDO databank based on all the model
descriptors, validation scores, and crystallographic parameters stored in the databank, via program and property lters. The versions of all the programs used are also
documented and stored to make a specic entry searchable. Complex queries can be
built to perform successive ltering steps on the data. The nal result can be downloaded as JSON le describing the dataset. To facilitate reproducible research, the
description includes a persistent, version-specic identier for each model so that
even when a PDB-REDO model is updated (e.g. because of an algorithmic improvement) the exact model used in the study is referenced. Old versions of PDB-REDO
databank entries are now archived to allow long-term access.

7.4 C onclu s i o n s 223
https://t.me/medicina_free
7.3.5 Submitting Structure Models to the PDB-REDO Pipeline
Besides the PDB-REDO databank, the PDB-REDO server is available, and also
accessible through https://pdb-redo.eu/ [64]. The server provides the user with
the ability to request an update for an existing PDB entry, or to upload their
own structure model and diraction data to the PDB-REDO pipeline. The server
returns a new structure model with rebuilt parts, electron density maps, optimized
renement parameters, models-specic restraints, and a wealth of validation data.
7.4 Conclusions
The PDB-REDO databank contains optimized PDB entries that were obtained from
X-ray crystallography. The underlying PDB-REDO pipeline renes, rebuilds, and
validates the PDB models based on the original experimental data and returns the
optimized structure models. PDB-REDO attempts to correct modeling errors, including the addition of missing side chain atoms, entire protein loops, or missing sugars
in glycosylation sites. Besides model changes, a signicant advantage of PDB-REDO
entries is that they are uniformly treated. This is in sharp contrast to their counterpart entries in the PDB archive, which all reect the idiosyncrasies of the software
and the crystallographers that constructed them. Importantly, metadata and renement, and validation parameters are provided for all entries, in FAIR data formats.
Also, secondary structure information (DSSP analysis), ligand validation, and available homologous structures are provided.Either the entire databank or a single entry
can be downloaded from https://pdb-redo.eu/. Single entries can also be directly
opened in COOT, YASARA, or CCP4mg software packages.
Therefore, the PDB-REDO databank is a convenient starting point for structure
selection in a CADD research project. For example, one of the key residues mutated
in various subtypes of the SARS-CoV-2 spike protein, Asp501, is ipped in the
PDB-REDO model compared to the model as deposited in the PDB. This residue is
located in the receptor binding domain of the spike protein and the ip results in a
hydrogen bond with Tyr41 in ACE2, spotting the importance of this residue in the
protein complex [70]. Another study using structures available from the PDB-REDO
databank involved docking in the ligand-binding pocket of β-lactoglobulin. The
PDB-REDO models were selected for this study because of their improved t to
experimental X-ray data and overall model quality score compared to the models
as deposited in the PDB [71]. Hydrogen bonding between a tryptophan and the
ligand has been characterized in the hydrophobic binding pocket, which can be
helpful in further investigation of this biopolymer involved in the transportation of
hydrophobic nutrients. Similarly, in research describing the conformational exibility of estrogen receptor α, the PDB-REDO structure of the monomer was selected
as this model has been completed with several side-chains that were not present in
the model as deposited in the PDB [72]. Besides using the models available from the
PDB-REDO databank, the PDB-REDO pipeline can also be used during renement
and validation of an X-ray structure model as e.g. done by Min et al. and Musak et al.

224 7 PDB-REDO in Computational-Aided Drug Design (CADD)
https://t.me/medicina_free
in their works for designing inhibitors of programmed cell death-1/programmed
death-ligand 1 [73] and the estrogen receptor [74], respectively.
Here, we have shown how PDB-REDO can contribute to CADD. At the moment
this is limited to structures from X-ray crystallography, which represent the vast
majority of experimental structure models, particularly those suited as drug targets.
Cryo-electron microscopy (Cryo-EM) is an increasingly popular method for providing experimental structure models including those with bound ligands. This opens
new possibilities for drug discovery including CADD. Developing new methods for
Cryo-EM-based drug discovery is an active research eld and will be for the foreseeable future. It is also part of the PDB-REDO research topics.
Acknowledgments and Funding
This work has been supported by iNEXT-Discovery, project number 871037, funded
by the Horizon 2020 program of the European Commission and by EOSC-Life
funding from the European Union’s Horizon 2020 research and innovation program
under grant agreement No. 824087.
List of Abbreviations and Symbols
2D Two dimensional
2mF
-DFcmap The weighted electron density map
o
3D Three dimensional
ÅÅngström
ADP Adenosine diphosphate
AMP-PNP Adenylyl-imidodiphosphate, a non-hydrolyzable ATP analogue
ATP Adenosine triphosphate
Cα C-alpha
CADD Computer-Aided Drug Discovery
Cryo-EM Cryo-Electron Microscopy
CYPs Cytochrome P450 hydroxylase enzymes
CXS CAPS buer
DSSP Dene Secondary Structure of Proteins
GPP Geranyl diphosphate
HSSP Homology derived Secondary Structure of Proteins
MAPK Mitogen-Activated Protein Kinase
NAG N-acetylglucosamine
PDB Protein Data Bank
PEG Poly-ethylene glycol
R5P Ribose-5-phosphate
RSCC Real Space Correlation Coecient
TLS Translation, liberation, and screw
UDP Uridine diphosphate
Z
bgG
Metric describing the relative orientation of nucleic acid bases
Соседние файлы в папке Библиотека им академика М.И. Перельмана
