Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5664_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
15.09.2026
Размер:
11 Мб
Скачать
☆
7.2 Structure Improvements by PDB-REDO 215
https://t.me/medicina_free
discovery research. One such protein is the aspartate transcarbamoylase found in E. Coli [51], which in humans is an interesting drug target for malaria [52] and can­cer [53]. The four cysteine side chains that are supposedly ligands for a zinc atom are apparently not, but instead, form an irrelevant “trisulde” bond in the struc­ture as deposited in the PDB (Figure 7.12). Platonyzer detects that this zinc binding site is not chemically correct, and generates restraints that will remodel the site dur­ing renement. In this process, the model annotation describing incorrect disulde bridges is removed. As a result, the zinc atom in the transcarbamoylase has the four cysteine side chains as ligands and thereby shows a correct coordination state (Figure 7.12). Additionally, the t to the electron density maps has improved and both the positive and negative dierence density disappeared.
(a)
(c)
Figure 7.12 (a) and (b) Zinc binding site in aspartate transcarbamoylase (grey, PDB entry 1tug, chain D) is improved by the PDB-REDO pipeline. Electron density maps (blue) are oversampled at 0.5 for clarity, contour levels: 3.00σ for 2mF density (red and green). (a) Binding site as found in the original model wherein the cysteine side chains form a trisulfide bridge and the zinc a tom does not have a valid number of ligands. (b) Improved zinc binding site in the PDB-REDO model where the cysteine side chains coordinate to the zinc atom that is now located in the middle of the binding site. (c) and (d) Magnesium binding site of the human type IIA DNA topoisomerase (grey, PDB entry 1zxn, chain A) is improved. Electron density maps (blue) are oversampled at 0.5 for clarity, contour levels: 2.00σ for 2mF (c) Magnesium site as in the original model in which water 960 is too far away to properly coordinate magnesium. (d) Magnesium binding site as in PDB-REDO model where water 960 has moved inward and describes a relevant magnesium site.
PDB
PDB
(b)
(d)
-DFcmap, 3.5σ for difference density (red and green).
o
PDB-REDO
PDB-REDO
− DFcmap, 3.5σ for difference
o
216 7 PDB-REDO in Computational-Aided Drug Design (CADD)
https://t.me/medicina_free
Besides structural zinc sites, platonyzer also evaluates magnesium and sodium binding sites with six-fold coordination. If these sites form (distorted) octahedrons, angle restraints are generated to clean up the site in renement. Revisiting the DNA topoisomerase described in Section 2.1.1/Figure 7.3, but focusing on the magne­sium site, which incorporates the ADP. The magnesium is coordinated to ADP phos­phate groups (ADP O3B at 2.09 Å and ADP O2A at 2.01 Å), Asn91 OD1 at 2.01 Å, and two water molecules (H and, H
O 960, is improbably far away at 2.87 Å (Figure 7.12). In the PDB-REDO
2
O 902 at 2.03 Å and H2O 958 at 2.24 Å). The sixth lig-
2
model, this water moves closer to the magnesium (1.99 Å), which makes the site more chemically realistic (Figure 7.12). This example illustrates the advantages of the PDB-REDO pipeline comprehensiveness: because both the ADP binding site and the magnesium coordination are improved, the user gets a clearer perspective on the structure–function relationship of DNA topoisomerase.
The identication of metal binding sites and metals is still a work in progress. Especially for sodium and magnesium whose binding sites are hard to distinguish in X-ray diraction data as these ions have the same number of electrons. The coordination distances can help, if not directly used as restraints in renement. The octahedral restraints from platonyzer help to unbias coordination distances, but the PDB-REDO pipeline does not change modeled ion identities yet. The decision about the ion identity thus still lies with the user.
Overall, there is a lot of room for improvement in the renement of metal binding sites, and using metal validation tools such as CheckMyMetal [54] and MetalPDB [55, 56] is strongly recommended. Nevertheless, the way the PDB-REDO pipeline takes care of (transition) metal atoms, can provide better structural starting points for CADD projects.
7.2.6 Limitations of the PDB-REDO Databank
In the examples above, we discussed dierent structural aspects that are addressed by the very ecient PDB-REDO pipeline, which can have a substantial impact on the total structure model. Nevertheless, the PDB-REDO comes with some limitations. The automation of the decision-making pipeline is a “means to an end” but it is not the solution to all problems. No systematic manual curation is performed, which means that not all model errors are removed and new errors may be introduced by PDB-REDO. The software is designed to be conservative and to make only few mistakes, but with more than 150,000 structure models with tens of millions of residues in total, problems are unavoidable. The metadata of each PDB-REDO entry is designed to allow users to make an informed choice on which models are most suitable for their downstream studies. For detailed studies involving a few structure models, we recommend that users inspect the models carefully in the context of the provided electron density maps. For large-scale studies, ltering models by overall quality indicators such as R-factors and possibly local density t metrics such as RSCC is important to construct the most suitable dataset [57].
7.2 Structure Improvements by PDB-REDO 217
https://t.me/medicina_free
Apart from the general problems linked to automation, there are few issues that need to be addressed separately. Firstly, some are regarding model annotation. PDB entries carry a lot of information that aects the way they are dealt with in PDB-REDO. Examples are descriptions of which (macromolecular) compounds were crystallized, which parts are modeled, residue and atom nomenclature, alternate conformers, R-factors, space groups, data quality indicators, et cetera. When these annotations are severely incorrect, the PDB-REDO process will either fail completely or will give very poor results. During PDB-REDO databank maintenance, errors like these are analyzed and if they can be solved by model rean­notation, update requests are sent to wwPDB annotators. This process developed into a fruitful collaboration with the wwPDB by which many small and large issues in structure models have been solved at the source (i.e. the PDB) so that everyone, not just PDB-REDO users, can benet.
A second issue is that not only new macromolecular structures but also new small molecule compounds are added to the PDB every week. Renement of such com­pounds requires restraint targets that are not immediately available in the CCP4 monomer library [58], which is the key restraint source of the PDB-REDO pipeline. In such cases, restraints are generated “on the y” based on the current atomic coor­dinates. This can lead to suboptimal restraints, which in turn lead to suboptimal molecular geometries. We collaborate with CCP4 developers to regularly update the CCP4 monomer library so that improved restraints become available to PDB-REDO users [59]. When users notice compounds with poor geometry in the PDB-REDO databank, brought on by poor restraints, they can request an update of the aected entry by clicking a link on the PDB-REDO entry page. Alternatively, they are wel­come to contact the PDB-REDO developers directly.
A third issue that combines the problems of model annotation and geometric restraints is related to intermolecular linkages. Although the polymeric linkages between amino acids, nucleotides, and recently also saccharides [60] are well standardized by model annotation at the wwPDB and in the CCP4 monomer library [59], covalent bonds between ligands and the macromolecule are still a signicant challenge. Generating the correct geometric restraints requires more information than is currently stored in structure model les. Currently, only which atoms are bound is stored, but not how they are bound. Changes in chemistry of the parent compounds, i.e. deleted atoms, changes in atom hybridization, and changes in bond orders, are not stored. This makes it challenging to generate correct restraints for covalent linkages without manual intervention [61]. Unfortunately, this phe­nomenon can be observed in some PDB-REDO entries and therefore users are advised to be vigilant when dealing with structure models that have such covalent linkages. Improvements to the standard data model used for macromolecules, the mmCIF format [62], are required to solve this issue permanently.
A nal issue of note is technical. Due to the many practical limitations of the PDB le, this most commonly used data format for structure models, is being replaced by mmCIF. Most notably the size restrictions for models (99,999 atoms in one le, with 62 chains) and the limited extensibility to capture new metadata have led to this replacement. The mmCIF le format does not have these limitations and has
218 7 PDB-REDO in Computational-Aided Drug Design (CADD)
https://t.me/medicina_free
been therefore chosen as the current standard for model handling in the PDB [63]. At the same time, the legacy of structural biology software, including software for CADD does not (fully) support the mmCIF format yet. This is also true for some of the software in the PDB-REDO pipeline. This means that even though PDB-REDO can use mmCIF formatted les as input and output, models are still described in PDB format within the pipeline. Because of this, 400 very large structure models could not be processed and are currently missing from the PDB-REDO databank. These will be added once all the software in the PDB-REDO pipeline becomes mmCIF compliant.
7.3 Access the PDB-REDO Databank and Metadata
7.3.1 Downloading and Inspecting Individual PDB-REDO Entries
All the models that are parsed through the PDB-REDO pipeline are stored in a data­bank that can be accessed at https://pdb-redo.eu/.The whole databank can be down­loaded, but conveniently it is also possible to download a single entry.The latter can be done either through the website, or users can open the desired structure model directly in molecular graphics software YASARA [37], COOT [26], or CCP4mg [35]. These software packages have the utility to download and show the PDB-REDO model and corresponding density maps from their interface [64].
When using the PDB-REDO website, an entry page visualizing the metadata of the structure model is provided. An example for PDB entry 1lf2 [65] is shown in Figure 7.13. On top of the entry page, a table (Figure 7.13) provides crystallographic data such as space group, resolution, and R-factor of the structure model. Also, the links to download the PDB-REDO data for the structure are provided in this table. An additional table (Figure 7.13) containing metrics about the validation of the crystallographic renement and model quality is provided. This table also indicates signicant improvements or deteriorations of the PDB-REDO model compared to the model as deposited in the PDB, conveniently marked in green or red, respectively. Next, a Kleywegt-like plot is provided (Figure 7.13), showing changes in model geometry before and after PDB-REDO in terms of ϕ-andψ-angles. The Ramachandran Z-scores are provided for both the PDB and PDB-REDO models, as well as details regarding residues in the preferred regions, allowed regions and outliers. In the subsequent panel (Figure 7.13) boxplots comparing the Z-score of the Ramachandran plot for both the PDB and PDB-REDO model with at least 1000 other structures that were obtained at similar resolutions are shown. This comparison is also provided for R-free and the rotamer quality as box plots. Additionally, a table (Figure 7.13) containing the signicant model changes caused by PDB-REDO is provided. This table provides the counts of these changes, among which the number of rotamers changed, the number of side chains that were ipped, chiralities that were xed, and also the improvement or deterioration of t to the density. When a PDB-REDO entry is loaded into COOT, a button list is provided to users to quickly inspect all the changes in the structure model in 3D.
7.3 Access the PDB-REDO Databank and Metadata 219
https://t.me/medicina_free
(a)
(b)
(d)
(c)
(e)
Figure 7.13 PDB-REDO entry interface for PDB entry 1lf2, which contains Plasmepsin 2, a potential antimalaria drug target bound to an inhibitor that can be found at https://pdb­redo.eu/db/1lf2. (a) Table containing crystallographic data. (b) Table containing the validation metrics used in PDB-REDO, while comparing to the model as deposited in the PDB. Green and red boxes indicate significant improvement or deterioration in the PDB-REDO model, respectively. (c) Kleywegt-like plot displaying changes of ϕ- and ψ-angles as a result of redoing the structure model. (d) Boxplots illustrating the model quality of the PDB and PDB-REDO models compared to models of similar resolution. (e) Table indicating the significant model changes obtained in the PDB-REDO model compared to the original model.
220 7 PDB-REDO in Computational-Aided Drug Design (CADD)
https://t.me/medicina_free
7.3.2 Data Available in PDB-REDO Entries
A PDB-REDO entry comes with the new model coordinates, maps, and validation details, but also contains all the metadata that describes the model in the original state, the re-rened model, and the nal (rebuilt and re-rened) model. All les available for one entry are shown in Table 7.1. This data are uniform throughout the whole databank, as the same pipeline has been used to generate all entries. Addition­ally, the data are formatted in ndable, accessible, interoperable and reproducible (FAIR)le formats. Both these features are advantages of the PDB-REDO databank.
A PDB-REDO entry contains les with the atomic coordinates for both the origi­nal model (as deposited in the PDB) and the PDB-REDO model. In addition to the PDB format for describing atomic coordinates, PDB-REDO also provides mmCIF coordinate les of the “redone” structure model. Additionally, the electron density maps for both the original and the redone structure models are provided. Valida­tion data are available for both structure models. The WHAT_CHECK reports [8] provide comprehensive model validation data and the JSON les contain valid data for the macromolecular structure model that can be easily used in data mining for structure selection. If ligands are present in the entry, a ligand validation le is pro­vided that contains the validation metrics for each ligand in both the original and the PDB-REDO structure models. If homologous structure models are available, the PDB identiers and chain identiers of the homologs are reported in the available_ homologs.json le. Furthermore, a Dene Secondary Structure of Proteins (DSSP) analysis of the PDB-REDO model is provided [40, 66] as well as a separate le con­taining the model change scores that are used at the PDBe entry pages [67]. For convenient visualization of a PDB-REDO structure model, a COOT [26] script is pro­vided, which also indicates the model changes made by the PDB-REDO pipeline. Finally, a PDB-REDO entry comes with some descriptive data of the PDB-REDO parameters (data.json) and, for the sake of provenance tracking, versions of all the software used during the process (versions.json).
7.3.3 Usage of the Uniform and FAIR Validation Data
Uniform data are an advantage for selecting multiple models for a CADD project, e.g. for homology modeling or investigation of a protein family. The models are all generated with the same pipeline and parameters can easily be extracted from the (meta)data provided. The JSON format respects FAIR data requirements. For each JSON formatted le, the corresponding schema is provided on the PDB-REDO web­site, which includes detailed information about the descriptors. The les themselves are user-friendly regarding data analyses. Desired values are easily extracted, after which analyses such as calculating dierences or sorting data can be performed using any computer scripting language. For example, one could extract the number of cation–π interactions for all ligands found in the PDB-REDO databank. In the lig­val.json le, these counts are stored for each ligand in the model as deposited in the PDB and for the corresponding PDB-REDO model. Once the values for a PDB-ligand and PDB-REDO ligand are compared, the dierence between the two values can
Tabl e 7.1 All data files for a PDB-REDO entry.
https://t.me/medicina_free
PDB-REDO data files
Files regarding Description Typical URL
Atomic coordinates
Electron density
Valida tion data
Descriptive data
The les describe the structure model in the original, the re-rened, and the nal re-built and rened state. Typical URLs of all the les are shown, using the ctitious PDB entry 9xyz as an example.
Initial model https://pdb-redo.eu/db/PDB identier/PDB identier_0cyc.pdb.gz Re-rened (only) model https://pdb-redo.eu/db/PDB identier/PDB identier_besttls.pdb.gz Re-rened and rebuilt structure model https://pdb-redo.eu/db/PDB identier/PDB identier_nal.pdb Re-rened and rebuilt structure model with total B-factors
(PDB) Re-rened and rebuilt structure model with total B-factors
(mmCIF) Map coecients for the original model https://pdb-redo.eu/db/PDB identier/PDB identier_0cyc.mtz.gz Map coecients for re-rened (only) model https://pdb-redo.eu/db/PDB identier/PDB identier_besttls.mtz.gz Map coecients for re-rened and rebuilt structure model https://pdb-redo.eu/db/PDB identier/PDB identier_nal.mtz WHAT_CHECK report for initial model https://pdb-redo.eu/db/PDB identier/wo/pdbout.txt WHAT_CHECK report for re-rened (only) model https://pdb-redo.eu/db/PDB identier/wc/pdbout.txt WHAT_CHECK report for re-rened and rebuilt model https://pdb-redo.eu/db/PDB identier/wf/pdbout.txt Ligand and ligand interaction data for the initial and
re-rened and rebuilt structure model Validation data for the initial and structure model https://pdb-redo.eu/db/PDB identier/PDB identier_0cyc.json.gz Validation data for the re-rened and rebuilt structure model https://pdb-redo.eu/db/PDB identier/PDB identier_nal.json Fitting and geometry scores for reporting at PDBe https://pdb-redo.eu/db/PDB identier/pdbe.json COOT script to show model changes https://pdb-redo.eu/db/PDB identier/PDB identier_nal.py (Python) and
PDB-REDO statistics for data mining (crystal parameters, R-factors, validation scores, etc.)
Homologous structures that are available for the structure model
DSSP analysis of the re-rened and rebuilt structure model https://pdb-redo.eu/db/PDB identier/PDB identier_nal.dssp Software versions of all programs used https://pdb-redo.eu/db/PDB identier/versions.json
https://pdb-redo.eu/db/PDB identier/PDB identier_nal_tot.pdb
https://pdb-redo.eu/db/PDB identier/PDB identier_nal.cif
https://pdb-redo.eu/db/PDB identier/PDB identier_ligval.json (only for entries that contain ligands)
https://pdb-redo.eu/db/PDB identier/PDB identier_nal.scm (Scheme) https://pdb-redo.eu/db/PDB identier/data.json
https://pdb-redo.eu/db/PDB identier/PDB identier_available_homologs.json
222 7 PDB-REDO in Computational-Aided Drug Design (CADD)
https://t.me/medicina_free
(a)
Figure 7.14 Completion of the Arg22 side chain (carbon atoms in grey) in the crystal structure of the catalytic subunit in protein kinase A (grey, PDB entry 5otg, chain A) results in cation–π interaction with ligand (carbon atoms in light blue). Electron density maps (blue) are oversampled at 0.5 for clarity, contour levels: 1.5σ for 2mF difference density (red and green). (a) Binding site of ligand with non-complete side chain modeled for Arg22 as found in the model as deposited in the PDB. (b) Binding site of ligand with rebuilt Arg22 side chain resulting in cation–π interaction with the ligand as seen in the PDB-REDO model.
PDB
(b)
PDB-REDO
− DFcmap, 3.5σ for
o
be calculated and sorted in descending order to rank the ligands based on gained cation–π interactions. One of the ligands on top of this list is the AO8 ligand of the protein kinase A catalytic subunit from Criteculus griseus (PDB-ID 5otg). This ligand binds in the ATP binding site of the kinase, most likely to inhibit its function [68]. The AO8 ligand shows no cation–π interactions in the original model (Figure 7.14), whereas in the PDB-REDO model 7, such interactions are reported. This increase is due to the completed Arg22 side chain that is present in the binding pocket. The guanidinium cation is interacting with the π–electrons in the phenyl moiety of the ligand improving the binding mode with an interaction between protein and lig­and (Figure 7.14) that was overlooked. For the sake of completeness, note that, the boronic acid of AO8is not visible in the electron density. As boronic acids are known to be oxidatively unstable, most likely oxidation occurred during the experimental process [69].
7.3.4 Creating Datasets from the PDB-REDO Databank
Next to listing close homologs, PDB-REDO has a more in-depth way of creating datasets for CADD and method development for structural biology. At https://pdb­redo.eu/archive users can search the PDB-REDO databank based on all the model descriptors, validation scores, and crystallographic parameters stored in the data­bank, via program and property lters. The versions of all the programs used are also documented and stored to make a specic entry searchable. Complex queries can be built to perform successive ltering steps on the data. The nal result can be down­loaded as JSON le describing the dataset. To facilitate reproducible research, the description includes a persistent, version-specic identier for each model so that even when a PDB-REDO model is updated (e.g. because of an algorithmic improve­ment) the exact model used in the study is referenced. Old versions of PDB-REDO databank entries are now archived to allow long-term access.
7.4 C onclu s i o n s 223
https://t.me/medicina_free
7.3.5 Submitting Structure Models to the PDB-REDO Pipeline
Besides the PDB-REDO databank, the PDB-REDO server is available, and also accessible through https://pdb-redo.eu/ [64]. The server provides the user with the ability to request an update for an existing PDB entry, or to upload their own structure model and diraction data to the PDB-REDO pipeline. The server returns a new structure model with rebuilt parts, electron density maps, optimized renement parameters, models-specic restraints, and a wealth of validation data.
7.4 Conclusions
The PDB-REDO databank contains optimized PDB entries that were obtained from X-ray crystallography. The underlying PDB-REDO pipeline renes, rebuilds, and validates the PDB models based on the original experimental data and returns the optimized structure models. PDB-REDO attempts to correct modeling errors, includ­ing the addition of missing side chain atoms, entire protein loops, or missing sugars in glycosylation sites. Besides model changes, a signicant advantage of PDB-REDO entries is that they are uniformly treated. This is in sharp contrast to their counter­part entries in the PDB archive, which all reect the idiosyncrasies of the software and the crystallographers that constructed them. Importantly, metadata and rene­ment, and validation parameters are provided for all entries, in FAIR data formats. Also, secondary structure information (DSSP analysis), ligand validation, and avail­able homologous structures are provided.Either the entire databank or a single entry can be downloaded from https://pdb-redo.eu/. Single entries can also be directly opened in COOT, YASARA, or CCP4mg software packages.
Therefore, the PDB-REDO databank is a convenient starting point for structure selection in a CADD research project. For example, one of the key residues mutated in various subtypes of the SARS-CoV-2 spike protein, Asp501, is ipped in the PDB-REDO model compared to the model as deposited in the PDB. This residue is located in the receptor binding domain of the spike protein and the ip results in a hydrogen bond with Tyr41 in ACE2, spotting the importance of this residue in the protein complex [70]. Another study using structures available from the PDB-REDO databank involved docking in the ligand-binding pocket of β-lactoglobulin. The PDB-REDO models were selected for this study because of their improved t to experimental X-ray data and overall model quality score compared to the models as deposited in the PDB [71]. Hydrogen bonding between a tryptophan and the ligand has been characterized in the hydrophobic binding pocket, which can be helpful in further investigation of this biopolymer involved in the transportation of hydrophobic nutrients. Similarly, in research describing the conformational exi­bility of estrogen receptor α, the PDB-REDO structure of the monomer was selected as this model has been completed with several side-chains that were not present in the model as deposited in the PDB [72]. Besides using the models available from the PDB-REDO databank, the PDB-REDO pipeline can also be used during renement and validation of an X-ray structure model as e.g. done by Min et al. and Musak et al.
224 7 PDB-REDO in Computational-Aided Drug Design (CADD)
https://t.me/medicina_free
in their works for designing inhibitors of programmed cell death-1/programmed death-ligand 1 [73] and the estrogen receptor [74], respectively.
Here, we have shown how PDB-REDO can contribute to CADD. At the moment this is limited to structures from X-ray crystallography, which represent the vast majority of experimental structure models, particularly those suited as drug targets. Cryo-electron microscopy (Cryo-EM) is an increasingly popular method for provid­ing experimental structure models including those with bound ligands. This opens new possibilities for drug discovery including CADD. Developing new methods for Cryo-EM-based drug discovery is an active research eld and will be for the foresee­able future. It is also part of the PDB-REDO research topics.
Acknowledgments and Funding
This work has been supported by iNEXT-Discovery, project number 871037, funded by the Horizon 2020 program of the European Commission and by EOSC-Life funding from the European Union’s Horizon 2020 research and innovation program under grant agreement No. 824087.
List of Abbreviations and Symbols
2D Two dimensional 2mF
-DFcmap The weighted electron density map
o
3D Three dimensional ÅÅngström ADP Adenosine diphosphate AMP-PNP Adenylyl-imidodiphosphate, a non-hydrolyzable ATP analogue ATP Adenosine triphosphate Cα C-alpha CADD Computer-Aided Drug Discovery Cryo-EM Cryo-Electron Microscopy CYPs Cytochrome P450 hydroxylase enzymes CXS CAPS buer DSSP Dene Secondary Structure of Proteins GPP Geranyl diphosphate HSSP Homology derived Secondary Structure of Proteins MAPK Mitogen-Activated Protein Kinase NAG N-acetylglucosamine PDB Protein Data Bank PEG Poly-ethylene glycol R5P Ribose-5-phosphate RSCC Real Space Correlation Coecient TLS Translation, liberation, and screw UDP Uridine diphosphate
Z
bgG
Metric describing the relative orientation of nucleic acid bases