Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5939_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
15.09.2026
Размер:
11 Мб
Скачать
☆
210 7 PDB-REDO in Computational-Aided Drug Design (CADD)
https://t.me/medicina_free
a hydrogen bond with R5P in the PDB-REDO model (Figure 7.6). Additionally, the PDB-REDO parameterization improves the overall ligand conformational t to the density of the R5P ligand. Notably, the parameterization used in the renement made the B-factor of atoms in R5P more similar to the surrounding protein atoms. This has resulted in negative dierence density indicating that the ribose-5-phosphate binding site is only partially occupied.
Although PDB-REDO can improve ligand binding sites and ligand geometry auto­matically, it cannot reinterpret the ligand to the level of providing an alternative compound. However,De Souza and co-workers showed that PDB-REDO models are a suitable starting point for manual reinterpretation of ligands [41].
7.2.2 Building of Protein Loops and Ligands into Protein Structure Models
Loops are the more exible parts of a protein and tend to give weaker diraction in crystallographic experiments. This results in poorer local quality of electron den­sity maps and therefore loops are harder to model than other secondary structure elements such as α-helices and β-sheets. Incomplete protein structure models are deposited to the PDB, mostly for good reasons: when the experimental data do not support the modeling of explicit atoms, those should not be added to the model. However, the decision not to model a loop is invariably a personal one, and some unmodeled loops can be built with reasonable reliability into poor electron density if some prerequisites are met. PDB-REDO tries to solve this issue in two ways. First, the renement steps in PDB-REDO typically lead to an improved atomic model, which in turn leads to better electron density maps than previously available. Second, the program Loopwhole reduces the vast number of possible loop conformations to the ones observed in experimental structures of the same or closely related proteins, which have a high probability of being correct. Combining both techniques allows PDB-REDO to add thousands of previously missing loops to PDB models. These more complete models can then enrich downstream functional or mechanistic inter­pretations of protein structure. Eventually, this could lead to a better description of potential binding sites to be targeted in drug discovery [21].
7.2.2.1 Loop Building Completes a Binding Site Region
In the crystal structure of galactokinase from Pyrococcus furiosus, PDB-REDO builds a substantial part of a loop surrounding the ADP binding site. While this structure was obtained in complex with ADP, magnesium, and galactose [42], the surroundings of the ADP binding site in the model as deposited in the PDB were left unmodeled. There is, however, density observed for this region of the protein (Figure 7.7). The automatic loop-building algorithm in the PDB-REDO pipeline improved this part of the protein structure by building the missing loop (residues 48–72) (Figure 7.7) with a good t to the experimental data. Although new interactions between the built residues with ADP are not observed in the PDB-REDO model, the description of the protein is more complete, which is relevant for structure-based CADD.
7.2 Structure Improvements by PDB-REDO 211
https://t.me/medicina_free
(a)
Figure 7. 7 PDB-REDO builds loop surrounding the ADP (carbon atoms in light blue) binding site of the Pyrococcus furiosus galactokinase (grey, PDB entry 1s4e, chain B). Electron density maps (blue) are oversampled at 0.5 for clarity, contour levels: 1.0σ for 2mFo-DFcmap, 3.5σ for difference density (red and green). (a) Model as deposited in the PDB with unmodeled density and difference density surrounding the ADP binding site. (b) PDB-REDO model with automatically built loop (residues 48–72, salmon pink ribbons) in the electron density to complete the chain.
PDB
(b)
PDB-REDO
7.2.2.2 Loop Building Results in Improved Binding Sites
Loop rebuilding in PDB-REDO can improve binding sites as showcased for the mag­nesium binding site of the geranyl diphosphate methyltransferase in complex with geranyldiphosphate (GPP) and sinefungin. This protein was used to provide insights in the methyl-group transfer mechanism, which is a common regulatory process in living organisms [43]. In the original model, magnesium interacts with the phos­phate groups of GPP and with Glu89. However, the magnesium ion lacks an addi­tional coordination ligand (Figure 7.8). This missing interaction could be explained by the unmodeled residues Val44 and Asn45, which are located in the magnesium
(a)
Figure 7. 8 Addition of residues Val44 and Asn45 results in completing magnesium and GPP (carbon atoms in light blue) binding site in geranyl diphosphate methyltransferase (grey, PDB entry 4f86, chain H). Electron density maps (blue) are oversampled at 0.5 for clarity, contour levels: 1.75σ for 2mFo-DFcmap, 3.5σ for difference density (red and green). (a) Model as deposited in the PDB with unmodeled residues 44 and 45. Magnesium is coordinated to GPP and Glu89 (carbon atoms in grey). (b) PDB-REDO model with added residues Val44 and Asn45 (carbon atoms in grey) resulting in more complete magnesium coordination.
PDB
(b)
PDB-REDO
212 7 PDB-REDO in Computational-Aided Drug Design (CADD)
https://t.me/medicina_free
binding site. The PDB-REDO loop building adds the two missing residues, lead­ing to a proper magnesium binding site. In addition to the interactions with Glu89 and the phosphates of GPP, the magnesium atom is now also coordinated to Asn45 (Figure 7.8).
7.2.2.3 Building new Compounds into Density
The cytochrome P450 monooxygenase enzymes (CYPs) are important in post-polyketide modications and thereby cause molecular diversity during metabolism. Insights into the mechanism of the CYP450 proteins provide valuable information for drug design, given the involvement of such enzymes in many metabolic processes, both physiological and xenobiotic [44]. Filipin is often used as probe for cholesterol binding sites in studies regarding this enzyme. One of such CYP450-lipin complex obtained by X-ray crystallography is the model of CYP105P1. Besides lipin, the structure also contains SO4ions and glycerol molecules that were used during the crystallization process [45]. When inspecting the model of this structure as deposited in the PDB, some well-dened density regions have no atoms modeled. Although this is not a feature that is employed by default, PDB-REDO can build missing compounds if provided a list of candidates. In the case of the CYP450-lipin complex, glycerol molecules and SO4were tted (data not shown). Furthermore, one of the modeled water networks in the original structure (Figure 7.9) indicates that it might be replaced by a dierent compound present in the crystal and was indeed replaced by this compound in the PDB-REDO model (Figure 7.9).
These results show that, provided the input model has machine-readable meta­data that describes other possible compounds, a model can be made more complete by automatically tting compounds into the density. This becomes particularly important if a compound directly inuences the binding pose of a ligand of interest. An example of this issue was described by Dym et al. [46], who showed that binding position of methylene blue in acetylcholinesterase was shifted outward by a poly-ethylene glycol (PEG) molecule sitting at the bottom of the binding pocket. This eect was overlooked in an early experimental structure model because the
(a)
Figure 7. 9 Original modeled water atoms (red spheres) are replaced by a buffer residue CSX (carbon atoms in light blue) in de PDB-REDO model of CYP105P1 (grey, PDB entry 3aba, chain A). Electron density maps (blue) are oversampled at 0.5 for clarity, contour levels: 1.0σ for 2mFo-DFcmap, 3.5σ for difference density (red and green). (a) Water network as present in the PDB model. (b) PDB-REDO model in which CXS has replaced the water network.
PDB
(b)
PDB-REDO
CXS
7.2 Structure Improvements by PDB-REDO 213
https://t.me/medicina_free
PEG was not tted in the electron density, which caused conicting results in downstream structural analyses.
Although crystallization additives and a buer were used as examples, this approach can be used for ligands and fragments in experimental high-throughput (lead) drug discovery. This is available as a feature in the PDB-REDO software, which is currently being tested in real-life experimental settings.
7.2.3 Nucleic Acid Improvements by PDB-REDO
PDB-REDO also applies nucleic acid restraints and validation targets for nucleic acid-containing structure models. These parameters are limited to the most com­mon structural features: Watson–Crick base pairs [31]. Overall nucleic acids are improved in both protein–nucleic acid complexes, as well as (mostly) nucleic acid structures such as ribosomal subunits. One of the rst ribosomal structures that were deposited in the PDB is the 30S subunit that was used to study its interactions with antibiotics [47]. The antibiotic paromomycin is bound in such a way that it is surrounded by nucleotides, among which is G1494 (Figure 7.10). PDB-REDO restraints improve the orientation of this C–G base pair without changing the interactions of the nucleic acid structure with the ligand (Figure 7.10). The Z
bgG
a metric describing the relative orientation of the bases, is 3.63 in PDB and 2.04 in PDB-REDO, which is closer to an “ideal” C–G base pair with Z
bpG
= 0.0. The strongest contribution to this improvement comes from reduced shearing between the bases. Additionally, the negative dierence density on the ligand has disappeared.
,
(a)
Figure 7.10 G1494-C1407 base pair (carbon atoms in grey) in the 30S ribosomal subunit (grey, PDB entry 1fjg) in complex with paromomycin (carbon atoms in light blue). Electron density maps (blue) are oversampled at 0.5 for clarity, contour level: 1.0 σ for 2mFo–DF map, 3.5 σ for difference density (red and green). Hydrogen bonds in the G–C base pair are indicated with black dotted lines. (a) Structure model as deposited in the PDB with suboptimal hydrogen bonds in the G–C base pair (2.7, 3.1, 3.3Å, shear Z-score −8.85 Å). (b) Structure model as in PDB-REDO with improved base pair geometry (hydrogen bond distances 2.8, 2.9, 3.0 Å, shear Z-score -4.43) and the negative difference density on the ligand drastically reduced.
PDB
(b)
PDB-REDO
c
214 7 PDB-REDO in Computational-Aided Drug Design (CADD)
https://t.me/medicina_free
7.2.4 Glycoprotein Structure Model Rebuilding
Glycosylation is a common post-translational modication of proteins. The modi­cations are important in recognition of other proteins, stability, and formation of protein complexes [48]. In PDB models, the glycosylated parts of proteins are not always resolved properly, which indicates that there is sucient room for model improvement. Initially, PDB-REDO worked on carbohydrates by improving model annotation and thereby helping model renement to improve the atomic coordi­nates [49], but much more substantial improvements were achieved when auto­mated (re)building of N-glycans was introduced [25]. The latteris illustrated through the crystal structure of the binding domain of SARS-CoV-2 spike, which contains a glycosylated Asn53 [50]. However, the electron density map suggests that the gly­can tree can be extended (Figure 7.11). The PDB-REDO model indeed contains an additional NAG (Figure 7.11).
7.2.5 Metal Binding Sites
Description of metal binding sites is challenging in X-ray modeling software, espe­cially at low resolution. The fact that metal binding sites have an enormous vari­ety in geometries and possible interactors is one of the underlying causes. Many site geometries are too context-sensitive to reliably predict and restrain in model renement, but there are common structural motifs around e.g. zinc, magnesium, and sodium, that are amenable to automated model improvement. The PDB-REDO pipeline includes a tool (platonyzer) that denes geometric restraints for structural zinc sites (i.e. ZnCysxHisysites as found in zinc ngers and other motifs) and octa­hedral sodium and magnesium sites. The restraints impose regular geometry dur­ing renement, especially when the experimental data are weak. These restraints can even recover the right conguration of initially extremely poorly modeled zinc atoms [30]. The latter is reasonably often present in proteins that are targeted in drug
(a)
Figure 7.11 Extension of a glycan tree (carbon atoms in grey) by PDB-REDO in the SARS-CoV-2 spike receptor-binding domain bound with ACE2 (grey, PDB entry 6m0j, chain A). Electron density maps (blue) are oversampled at 0.5 for clarity, contour levels: 1.0σ for 2mFo-DFcmap, 3.5σ for difference density (red and green). (a) Model as deposited in the PDB that shows unmodeled density close to NAG. (b) PDB-REDO model which contains an extended glycan; no unmodeled density is observed.
PDB
(b)
PDB-REDO
7.2 Structure Improvements by PDB-REDO 215
https://t.me/medicina_free
discovery research. One such protein is the aspartate transcarbamoylase found in E. Coli [51], which in humans is an interesting drug target for malaria [52] and can­cer [53]. The four cysteine side chains that are supposedly ligands for a zinc atom are apparently not, but instead, form an irrelevant “trisulde” bond in the struc­ture as deposited in the PDB (Figure 7.12). Platonyzer detects that this zinc binding site is not chemically correct, and generates restraints that will remodel the site dur­ing renement. In this process, the model annotation describing incorrect disulde bridges is removed. As a result, the zinc atom in the transcarbamoylase has the four cysteine side chains as ligands and thereby shows a correct coordination state (Figure 7.12). Additionally, the t to the electron density maps has improved and both the positive and negative dierence density disappeared.
(a)
(c)
Figure 7.12 (a) and (b) Zinc binding site in aspartate transcarbamoylase (grey, PDB entry 1tug, chain D) is improved by the PDB-REDO pipeline. Electron density maps (blue) are oversampled at 0.5 for clarity, contour levels: 3.00σ for 2mFo− DFcmap, 3.5σ for difference density (red and green). (a) Binding site as found in the original model wherein the cysteine side chains form a trisulfide bridge and the zinc atom does not have a valid number of ligands. (b) Improved zinc binding site in the PDB-REDO model where the cysteine side chains coordinate to the zinc atom that is now located in the middle of the binding site. (c) and (d) Magnesium binding site of the human type IIA DNA topoisomerase (grey, PDB entry 1zxn, chain A) is improved. Electron density maps (blue) are oversampled at 0.5 for clarity, contour levels: 2.00σ for 2mFo-DFcmap, 3.5σ for difference density (red and green). (c) Magnesium site as in the original model in which water 960 is too far away to properly coordinate magnesium. (d) Magnesium binding site as in PDB-REDO model where water 960 has moved inward and describes a relevant magnesium site.
PDB
PDB
(b)
(d)
PDB-REDO
PDB-REDO
216 7 PDB-REDO in Computational-Aided Drug Design (CADD)
https://t.me/medicina_free
Besides structural zinc sites, platonyzer also evaluates magnesium and sodium binding sites with six-fold coordination. If these sites form (distorted) octahedrons, angle restraints are generated to clean up the site in renement. Revisiting the DNA topoisomerase described in Section 2.1.1/Figure 7.3, but focusing on the magne­sium site, which incorporates the ADP. The magnesium is coordinated to ADP phos­phate groups (ADP O3B at 2.09 Å and ADP O2A at 2.01 Å), Asn91 OD1 at 2.01 Å, and two water molecules (H2O 902 at 2.03 Å and H2O 958 at 2.24 Å). The sixth lig­and, H2O 960, is improbably far away at 2.87 Å (Figure 7.12). In the PDB-REDO model, this water moves closer to the magnesium (1.99 Å), which makes the site more chemically realistic (Figure 7.12). This example illustrates the advantages of the PDB-REDO pipeline comprehensiveness: because both the ADP binding site and the magnesium coordination are improved, the user gets a clearer perspective on the structure–function relationship of DNA topoisomerase.
The identication of metal binding sites and metals is still a work in progress. Especially for sodium and magnesium whose binding sites are hard to distinguish in X-ray diraction data as these ions have the same number of electrons. The coordination distances can help, if not directly used as restraints in renement. The octahedral restraints from platonyzer help to unbias coordination distances, but the PDB-REDO pipeline does not change modeled ion identities yet. The decision about the ion identity thus still lies with the user.
Overall, there is a lot of room for improvement in the renement of metal binding sites, and using metal validation tools such as CheckMyMetal [54] and MetalPDB [55, 56] is strongly recommended. Nevertheless, the way the PDB-REDO pipeline takes care of (transition) metal atoms, can provide better structural starting points for CADD projects.
7.2.6 Limitations of the PDB-REDO Databank
In the examples above, we discussed dierent structural aspects that are addressed by the very ecient PDB-REDO pipeline, which can have a substantial impact on the total structure model. Nevertheless, the PDB-REDO comes with some limitations. The automation of the decision-making pipeline is a “means to an end” but it is not the solution to all problems. No systematic manual curation is performed, which means that not all model errors are removed and new errors may be introduced by PDB-REDO. The software is designed to be conservative and to make only few mistakes, but with more than 150,000 structure models with tens of millions of residues in total, problems are unavoidable. The metadata of each PDB-REDO entry is designed to allow users to make an informed choice on which models are most suitable for their downstream studies. For detailed studies involving a few structure models, we recommend that users inspect the models carefully in the context of the provided electron density maps. For large-scale studies, ltering models by overall quality indicators such as R-factors and possibly local density t metrics such as RSCC is important to construct the most suitable dataset [57].
7.2 Structure Improvements by PDB-REDO 217
https://t.me/medicina_free
Apart from the general problems linked to automation, there are few issues that need to be addressed separately. Firstly, some are regarding model annotation. PDB entries carry a lot of information that aects the way they are dealt with in PDB-REDO. Examples are descriptions of which (macromolecular) compounds were crystallized, which parts are modeled, residue and atom nomenclature, alternate conformers, R-factors, space groups, data quality indicators, et cetera. When these annotations are severely incorrect, the PDB-REDO process will either fail completely or will give very poor results. During PDB-REDO databank maintenance, errors like these are analyzed and if they can be solved by model rean­notation, update requests are sent to wwPDB annotators. This process developed into a fruitful collaboration with the wwPDB by which many small and large issues in structure models have been solved at the source (i.e. the PDB) so that everyone, not just PDB-REDO users, can benet.
A second issue is that not only new macromolecular structures but also new small molecule compounds are added to the PDB every week. Renement of such com­pounds requires restraint targets that are not immediately available in the CCP4 monomer library [58], which is the key restraint source of the PDB-REDO pipeline. In such cases, restraints are generated “on the y” based on the current atomic coor­dinates. This can lead to suboptimal restraints, which in turn lead to suboptimal molecular geometries. We collaborate with CCP4 developers to regularly update the CCP4 monomer library so that improved restraints become available to PDB-REDO users [59]. When users notice compounds with poor geometry in the PDB-REDO databank, brought on by poor restraints, they can request an update of the aected entry by clicking a link on the PDB-REDO entry page. Alternatively, they are wel­come to contact the PDB-REDO developers directly.
A third issue that combines the problems of model annotation and geometric restraints is related to intermolecular linkages. Although the polymeric linkages between amino acids, nucleotides, and recently also saccharides [60] are well standardized by model annotation at the wwPDB and in the CCP4 monomer library [59], covalent bonds between ligands and the macromolecule are still a signicant challenge. Generating the correct geometric restraints requires more information than is currently stored in structure model les. Currently, only which atoms are bound is stored, but not how they are bound. Changes in chemistry of the parent compounds, i.e. deleted atoms, changes in atom hybridization, and changes in bond orders, are not stored. This makes it challenging to generate correct restraints for covalent linkages without manual intervention [61]. Unfortunately, this phe­nomenon can be observed in some PDB-REDO entries and therefore users are advised to be vigilant when dealing with structure models that have such covalent linkages. Improvements to the standard data model used for macromolecules, the mmCIF format [62], are required to solve this issue permanently.
A nal issue of note is technical. Due to the many practical limitations of the PDB le, this most commonly used data format for structure models, is being replaced by mmCIF. Most notably the size restrictions for models (99,999 atoms in one le, with 62 chains) and the limited extensibility to capture new metadata have led to this replacement. The mmCIF le format does not have these limitations and has
218 7 PDB-REDO in Computational-Aided Drug Design (CADD)
https://t.me/medicina_free
been therefore chosen as the current standard for model handling in the PDB [63]. At the same time, the legacy of structural biology software, including software for CADD does not (fully) support the mmCIF format yet. This is also true for some of the software in the PDB-REDO pipeline. This means that even though PDB-REDO can use mmCIF formatted les as input and output, models are still described in PDB format within the pipeline. Because of this, 400 very large structure models could not be processed and are currently missing from the PDB-REDO databank. These will be added once all the software in the PDB-REDO pipeline becomes mmCIF compliant.
7. 3 Access the PDB-REDO Databank and Metadata
7.3.1 Downloading and Inspecting Individual PDB-REDO Entries
All the models that are parsed through the PDB-REDO pipeline are stored in a data­bank that can be accessed at https://pdb-redo.eu/. The whole databank can be down­loaded, but conveniently it is also possible to download a single entry. The latter can be done either through the website, or users can open the desired structure model directly in molecular graphics software YASARA [37], COOT [26], or CCP4mg [35]. These software packages have the utility to download and show the PDB-REDO model and corresponding density maps from their interface [64].
When using the PDB-REDO website, an entry page visualizing the metadata of the structure model is provided. An example for PDB entry 1lf2 [65] is shown in Figure 7.13. On top of the entry page, a table (Figure 7.13) provides crystallographic data such as space group, resolution, and R-factor of the structure model. Also, the links to download the PDB-REDO data for the structure are provided in this table. An additional table (Figure 7.13) containing metrics about the validation of the crystallographic renement and model quality is provided. This table also indicates signicant improvements or deteriorations of the PDB-REDO model compared to the model as deposited in the PDB, conveniently marked in green or red, respectively. Next, a Kleywegt-like plot is provided (Figure 7.13), showing changes in model geometry before and after PDB-REDO in terms of ϕ- and ψ-angles. The Ramachandran Z-scores are provided for both the PDB and PDB-REDO models, as well as details regarding residues in the preferred regions, allowed regions and outliers. In the subsequent panel (Figure 7.13) boxplots comparing the Z-score of the Ramachandran plot for both the PDB and PDB-REDO model with at least 1000 other structures that were obtained at similar resolutions are shown. This comparison is also provided for R-free and the rotamer quality as box plots. Additionally, a table (Figure 7.13) containing the signicant model changes caused by PDB-REDO is provided. This table provides the counts of these changes, among which the number of rotamers changed, the number of side chains that were ipped, chiralities that were xed, and also the improvement or deterioration of t to the density. When a PDB-REDO entry is loaded into COOT, a button list is provided to users to quickly inspect all the changes in the structure model in 3D.
7.3 Access the PDB-REDO Databank and Metadata 219
https://t.me/medicina_free
(a)
(b)
(d)
(c)
(e)
Figure 7.13 PDB-REDO entry interface for PDB entry 1lf2, which contains Plasmepsin 2, a potential antimalaria drug target bound to an inhibitor that can be found at https://pdb­redo.eu/db/1lf2. (a) Table containing crystallographic data. (b) Table containing the validation metrics used in PDB-REDO, while comparing to the model as deposited in the PDB. Green and red boxes indicate significant improvement or deterioration in the PDB-REDO model, respectively. (c) Kleywegt-like plot displaying changes of ϕ- and ψ-angles as a result of redoing the structure model. (d) Boxplots illustrating the model quality of the PDB and PDB-REDO models compared to models of similar resolution. (e) Table indicating the significant model changes obtained in the PDB-REDO model compared to the original model.