Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5661_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
15.09.2026
Размер:
11 Мб
Скачать
☆
https://t.me/medicina_free
5
https://t.me/medicina_free
The Protein Data Bank (PDB) and Macromolecular Structure Data Supporting Computer-Aided Drug Design
David Armstrong1, John Berrisford2, Preeti Choudhary1, Lukas Pravda3, James Tolchard4, Mihaly Varadi1, and Sameer Velankar
1
Protein Data Bank in Europe, EMBL-EBI, Wellcome Genome Campus, Hinxton, Cambridge, CB10 1SD, UK
2
AstraZeneca, AstraZeneca Academy House, 136 Hills Rd, Cambridge, CB2 8PA, UK
3
Exscientia, The Schrodinger Building, Oxford Science Park, Oxford, Oxfordshire, United Kingdom, OX4 4GE
4
Centre de RMN à Tr è s Hauts Champs de Lyo n , Claude Bernard University Ly on 1, 5 rue de la Doua, 69100
Villeurbanne, France
1
5.1 Introduction
The Protein Data Bank (PDB) [1], managed by the global Worldwide PDB (wwPDB) consortium [2], is one of the oldest scientic databases in life sciences. The PDB archives structural models of biological macromolecules, derived from experimental data. The wwPDB partners, the Research Collaboratory for Structural Bioinformat­ics (RCSB) [3], Protein Data Bank in Europe (PDBe) [4], Protein Data Bank Japan (PDBj) [5], Electron Microscopy Data Bank (EMDB) [6], and Bio Mag Res Bank (BMRB) [7], manage the PDB archive based on the FAIR principles [8], ensuring structural biology data are Findable, Accessible, Interoperable and Reusable.
The PDB archive contains over 200,000 structures of proteins, DNA, and RNA, and their complexes with small molecules, such as cofactors and inhibitors. This data, along with related metadata and experimental data, are stored in the PDB archive in PDBx/mmCIF formatted les [9]. Since 2019, it has been mandatory for X-ray crys­tallographic structures to be deposited in the PDBx/mmCIF format [10], ensuring maximum capture and validation of metadata for these structures.
The PDB is a vital resource for computer-aided drug design and contains struc­tural information for large numbers of drugs, with 5494 unique ligands in the PDB mapped to entries in DrugBank [11]. Additionally, there are numerous potential drug targets as well as structures containing macromolecular drug targets. This data are of paramount importance in determination of key drug binding sites in macro­molecules [12], and in understanding existing binding modes to support designs of novel compounds to target these binding sites.
In addition to the three-dimensional (3D) coordinate information archived in the PDB, standardization of polymer sequences and chemical identiers in PDB entries facilitates the integration of additional metadata from external resources
141
Open Access Databases and Datasets for Drug Discovery, First Edition. Edited by Antoine Daina, Michael Przewosny, and Vincent Zoete. © 2024 WILEY-VCH GmbH. Published 2024 by WILEY-VCH GmbH.
142 5 The Protein Data Bank (PDB) and Macromolecular Structure Data Supporting Computer
https://t.me/medicina_free
[13, 14]. This additional biological and chemical information is vitally important in understanding the scientic context of the structures in the PDB, which is required to determine functional information.
As part of the European Bioinformatics Institute (EMBL-EBI), the PDBe [4] is uniquely placed to integrate macromolecular structure data from partner resources at EBI and beyond. This allows the integration of relevant biological and chemical data to better demonstrate the functional importance of structures in the PDB. The PDBe-Knowledge Base (PDBe-KB) [15] builds upon this, with partnerships extend­ing across the structural bioinformatics community, allowing the integration of even more data related to macromolecular structure and function. This includes a large number of partners involved in cheminformatics who provide data relating to small molecules and their macromolecular binding sites, including canSAR [16], 3DLi­gandSite [17], P2Rank [18], and many more, detailed at pdbe-kb.org/partners.
The wwPDB, by adhering to the FAIR principles, ensures that all PDB data are pro­vided freely, with no limits upon its use. All PDBe and PDBe-KB tools and resources are also free to use, with scripts and software pipelines made open source wherever possible. In addition to access through the PDBe [4] and PDBe-KB [15] websites, all PDBe data are made available through publicly accessible APIs [19, 20], while much of this data is also available via a distributed PDBe knowledge graph [20].
This chapter will introduce the type of data available in the PDB, how it is orga­nized and curated, and how it can be used to support drug design. It will also give an overview of the tools and resources at PDBe and PDBe-KB, highlighting how these can support understanding of drug binding and function in PDB structures to improve research within this area.
5.2 Small Molecule Data in Protein Data Bank (PDB) Entries
5.2.1 What Data are in the PDB Archive?
The PDB is the single global archive of experimentally determined 3D structure data of biological macromolecules. Each PDB structure must contain a polymeric entity, i.e. protein or nucleic acid, however, information can also be included for all non-polymeric ligands within the structure.
The atomic coordinates of a PDB entry are built with a specic hierarchy, ensuring that each molecule and its constituent parts can be clearly identied. The smallest component of the hierarchy is at the atomic level, with each separate coordinate line in the archive le describing the specic position of an atom in 3D space. Each atom is dened as part of a larger chemical component or residue, dened by a 3-character ID code.
These chemical components can be either individual bound ligands in the struc­ture or individual residues within a polymeric molecule, with each dened by a unique identier (asym or chain ID) and numbering for clear identication. Beyond each specic polymeric chain ID, each unique and individual molecule in the struc­ture is also given a specic “entity” identier, which is used to link all the metadata
5.2 Small Molecule Data in Protein Data Bank (PDB) Entries 143
https://t.me/medicina_free
Figure 5.1 A sample of the atomic coordinates from a PDBx/mmCIF format file (PDB entry 2yi7) is shown on the left. The individual items in the “atom_site” category are listed first, highlighting the data present in each column for the subsequent data. A subset of the atomic coordinates is displayed below, with graphical representation of these coordinates, displayed on the right of the image, only selected atoms are labeled for clarity.
relating to this molecule. More information about ligand molecules in the PDB will be discussed in Section 5.3 (Figure 5.1).
As previously mentioned, the PDB is an archive of experimentally determined structures and, as such, is limited to a subset of accepted experimental struc­ture determination techniques. Depending on the experimental method used, the archive PDB entry le will contain information specic to the technique, while additional experimental data les are also collected to allow assessment of experimental data in the context of the derived coordinate models. The three main methods accepted for PDB depositions are diraction techniques such as X-ray crystallography, cryo-electron microscopy (cryoEM), and nuclear magnetic resonance (NMR). A summary of the PDB deposition data requirements for each of these techniques is given in Table 5.1.
The oldest and most common technique for determining structures in the PDB is X-ray crystallography, which involves the generation of crystalline structures of the biological sample. These crystals are then exposed to a high-powered X-ray beam, often at a synchrotron facility, and the specic diraction pattern from these X-rays is collected by a detector. Before this diraction pattern can be used for structure calculation, rst the phases must be determined, using techniques such as molecular or isomorphous replacement.
Using the intensities and phases of the spots (or reections) in the diraction pattern, crystallographers can calculate a map of electron density for the biologi­cal specimen, into which the atomic model can be built. In addition to coordinates, since 2008, submission of PDB structures solved by X-ray crystallography must also include submission of the structure factor les dening the experimental reec­tions data.
144 5 The Protein Data Bank (PDB) and Macromolecular Structure Data Supporting Computer
https://t.me/medicina_free
Table 5.1 A summary of the main experimental methods accepted for PDB depositions, including information on the type of data deposited in a PDB entry file, the mandatory requirements for deposition, the related experimental data and where it should be deposited, and the raw experimental data and where it is recommended to be deposited.
Experimental method
X-ray diraction
Nuclear Magnetic Resonance (NMR)
Cryo-electron microscopy (cryoEM)
Model solution method
Single model built into experimentally derived electron density map
Multiple models representing a range of conformers that satisfy experimental restraints
Single model built into experimentally derived electric potential map
PDB deposition requirements
Coordinate le (PDBx/mmCIF format)
Structure factor le (mtz or CIF format)
Coordinate le (PDB or PDBx/mmCIF format)
NMR restraints le (STAR or NEF format)
Coordinate le (PDB or PDBx/mmCIF format)
Image for public display at EMDB EMDB map deposition (MRC or CCP4 format)
Experimental data (archive)
Structure factors (PDB)
NMR restraints (BMRB)
NMR chemical shifts (BMRB)
Electric potential map (EMDB)
Raw data (archive)
X-ray diraction image data (SBGrid, IRRMC)
NMR spectral parameters (BMRB)
NMR relaxation data (BMRB)
Electron microscopy images (EMPIAR)
The rst cryoEM structure in the PDB was released in 1991, however, it has taken a long time for the technique to establish itself as a routine method for high-resolution structure determination. The technique involves the imaging of biological molecules and complexes under cryogenic conditions, using a transmission electron micro­scope. The sample is ash-frozen in a thin layer with electrons passing through the sample to a detector, where images of each particle are captured. These images are then sorted and processed computationally to generate a 3D map of the sample, to allow tting of a molecular model.
The use of cryoEM has allowed the determination of larger and more heteroge­neous complexes than was previously possible with X-ray crystallography, however until recently, was not able to reach resolutions sucient to interpret atomic-level details. However, recent advances in the technique, including improved software and hardware including free-electron detectors, have led to cryoEM becoming the fastest-growing experimental method for studying macromolecular structure [21]. This is due to the determination of high-resolution structures in conditions that
5.2 Small Molecule Data in Protein Data Bank (PDB) Entries 145
https://t.me/medicina_free
better represent the native environment of the sample. Since 2016, submission of cryoEM structures to the PDB has also required submission of the experimental maps to the EMDB [6].
The third main technique used for solving structures in the PDB is NMR. This method involves the use of powerful magnetic elds to discern minor dierences in the resonance frequencies, also known as chemical shifts, of atoms in the macro­molecule. These chemical shifts are used to determine interactions among atoms in the molecule and generate atomicrestraints, which can be used to build the structure of the macromolecules.
In solution-state NMR, the biological macromolecules can move freely in the sol­vent, allowing these experiments to capture the range of structural conformations adopted by these molecules. These structures are therefore deposited as multiple models in the PDB, highlighting an ensemble of potential structural conformations adopted by the macromolecules. An alternative technique is solid-state NMR [22], which uses similar underlying principles, however, is used to determine structures of molecular structures within solid or semisolid materials, including macromolecules within biological membranes.
Though the three experimental methods mentioned above account for the vast majority of PDB entries, there are also additional variations in diraction methods that can be used to solve structures in the PDB. Firstly, neutron crystallography [23] is a similar technique to X-ray crystallography, however, it relies upon neutron scat­tering to determine the macromolecular structure. The benet of this technique is that the neutrons interact with atomic nuclei, rather than electrons, which improves observation of hydrogen atoms in the structure. Neutron and X-ray crystallography are often used in conjunction to provide both high-resolution data and to determine positions of hydrogen atoms.
There are also methods that utilize X-raycrystallography toallow high-throughput determination of ligand binding in macromolecules. Fragment screening [24] exper­iments involve the “soaking” of various small molecule fragments into crystals of the sample. Determining how these dierent fragments bind within the pro­tein structure can help to improve the mechanistic understanding of a protein and the identication of suitable drug candidates. Pan-Dataset Density Anal­ysis (PanDDA) experiments [25] can be used to compare multiple fragment screening datasets to identify weak signals of bound small molecules within the noise of the data and can further improve the determination of small molecule binding sites.
5.2.2 Definition of Small Molecules in OneDep
In the PDB archive, ligands are dened as any molecule that is not part of a larger polymeric molecule, i.e. proteins, nucleic acids, or branched carbohydrates [26]. In the case of peptide-like ligands, if these contain at least two standard peptide bonds, then these are classed as polymeric entities. Small molecule ligands in the PDB can have a range of distinct functions, for example, as substrates, products, and inhibitors.
146 5 The Protein Data Bank (PDB) and Macromolecular Structure Data Supporting Computer
https://t.me/medicina_free
Each distinct ligand in the PDB archive is assigned its own unique 3-character ID code, to support easy identication across the full archive [26]. Each of these ligands includes a detailed chemical description, which is contained within the wwPDB Chemical Component Dictionary (CCD), and discussed in more detail in the next section.
5.3 Small Molecule Dictionaries
5.3.1 wwPDB Chemical Component Dictionary (CCD)
The wwPDB maintains a reference dictionary containing chemical descriptions of each unique chemical component present in the whole PDB archive. These components are the building blocks for PDB entries and include amino acids, nucleotides, metal ions, and other nonpeptide small molecules. These chemical component descriptions are stored in the wwPDB’s CCD (https://www.wwpdb.org/ data/ccd) [26].
Each component is given a unique identier and has a CCD denition that describes the molecule. As of July 2022, there are around 37,000 unique denitions in the wwPDB CCD. The unique identier is currently limited to a maximum of three characters. Most denitions contain a three-character unique identier, with one- or two-character identiers mostly used for identication of DNA and RNA bases or for components containing individual elements. Some examples of chemical components include manganese (Mn), adenosine-triphosphate (ATP), and glutamate (GLU). Previously, these identiers were assigned with some meaning to their chemical denition, however, due to the high number of chemical components in the dictionary, these CCD identiers are now randomly assigned for any new ligands.
If the CCD denition is dening a standard amino acid or nucleotide, the IUPAC protein one-letter code [27] is also provided, for example, Glutamate has the unique identier GLU and the one-letter code E. Small modications of standard amino acids or nucleotides, generally, where the modication contains fewer than 10 atoms, have their own CCD denitions, for example, phospho-serine has the identier SEP. These modied amino acids or nucleotides have the same one-letter code as their standard amino acid parent, therefore SEP, which is a modied serine amino acid, also contains the one-letter code S.
Peptide-based chromophores provide a more complex example. In proteins, chromophores typically comprise three amino acids condensed into one molecule. In the CCD component, the one-letter code lists all amino acids that make up the chromophore. For example, the CCD denition for the chromophore CR2, which comprises the amino acids GLY-TYR-GLY and has the corresponding one-letter code GYG.
In addition to providing a unique identier, each CCD denition describes the chemistry of the molecule. This includes the name of each atom in the molecule, the order of bonds between atoms, stereochemistry of chiral atoms, and any charges
5.3 Small Molecule Dictionaries 147
https://t.me/medicina_free
on individual atoms or the whole molecule. There are two sets of 3D coordinates included in the CCD denition: the model coordinates, extracted from the PDB entry from which the CCD was initially generated, and idealized coordinates (from the lowest energy conformation), generated using the expected geometry of the ligand from the CCD denition.
Additional metadata are also included in the CCD denition, including the name of the molecule and any synonyms. If a common name is known for the compound, then it is provided either as the name or as one of the synonyms, otherwise, system­atic names are provided. The systematic name is the default name generated by the OpenEye [28] software at the time of creation of the compound. Chemical descrip­tors in InChI and SMILES formats are provided and are automatically generated by OpenEye during creation of the compound.
Every molecule in the PDB must have a corresponding CCD denition. During deposition and biocuration of PDB entries, each compound in the deposited le is compared to existing CCD denitions using a graph match algorithm [29]. If rene­ment restraints are provided in the uploaded coordinate le, then this information is used for comparing the molecule to entries in the CCD, otherwise, the coordinates are used for searching. If a molecule fails to nd a match in the CCD, then deposi­tors are asked for further details of the molecule so that a new CCD denition can be made. A new CCD denition is then dened during biocuration and added to the CCD. The name of the molecule is either dened by the depositor, if a common name is available, or a systematic name is generated based on the chemistry of the molecule.
5.3.2 The Peptide Reference Dictionary
The PDB archive also contains examples of complex ligands, which comprise combinations of amino acid and amino acid-like components. These peptide-like molecules commonly possess important biological functions, such as antibiotic or inhibitory activity, however, a description of their global chemistry does not t within the classic PDB denitions of polymers and ligands. Therefore, to support the curation of this important class of molecules, their chemical and biological descrip­tions are described in a separate dictionary. The “Peptide Reference Dictionary” (PRD) resource [30] provides a standardized framework for these denitions so that each PRD entity can be described in a way that outlines both its subcomponent composition and global characteristics. These PRD denitions are stored in the wwPDB Biologically Interesting Molecule Reference Dictionary (BIRD) (https:// www.wwpdb.org/data/bird).
The process of dening a new PRD molecule, or matching connected CCDs to an existing PRD, occurs during the biocuration process. Any atomistic information, such as reference coordinates, bonding, and chemical descriptors are then listed in PRD-chemical denition les, whereas subcomponent sequence, CCD-linking, and any related biological details, such as biological function, are stored in PRD-molecular denition les. By using the PRD dictionary to standardize biocura­tion, it allows for the precise interrogation, comparison, and retrieval of global and
148 5 The Protein Data Bank (PDB) and Macromolecular Structure Data Supporting Computer
https://t.me/medicina_free
Figure 5.2 An example of PRD is the antibiotic Vancomycin (PRD_000204). Vancomycin consists of seven nonstandard amino acids (ball and stick, distinguished by color) and two sugar components (cyan and grey in a 3D representation of the SNFG notation). The PRD polymer linking section of the PRD molecular definition file is shown on the right-hand side.
subcomponent-specic data, which would otherwise not be possible. Links to PRD denition les can be found on relevant PDB entry pages and references to key PRD identiers (ID, name, type, and function) are stored in PDB entry PDBx/mmCIF les under the pdbx_molecule_features category (Figure 5.2).
5.4 Additional Ligand Annotations in the PDB Archive
In addition to the standard chemical and biomolecular denitions already described, where possible, the PDB biocuration pipeline generates additional ligand-related information to best describe ligand complexity and intramolecular connectivity and to do so with a consistent and objective methodology [31]. Many software packages rely upon these data for purposes such as molecular renement and accurate molec­ular visualization.
5.4.1 Linkage Information
The most common ligand-associated annotations found in PDB entry les refer to intermolecular bonding. These refer to all nonstandard linkages between and among polymer components and CCD ligands. This information was historically found in PDB-format LINK, SSBOND, and CONECT records, however, the current PDBx/mmCIF archive format provides additional data, including leaving group information, atomic descriptions, and author numbering for up to three chemical component partners per interaction [9]. At present, to ensure standardization across the whole PDB archive, all link records and bonding distances are calculated and validated during the biocuration process [31]. The PDBx/mmCIF dictionary cur­rently supports covalent, hydrogen, covalent–metal, and disulde-bond interaction “types” [32] (Figure 5.3).
5.4 Additional Ligand Annotations in the PDB Archive 149
https://t.me/medicina_free
(a) (b)
Figure 5.3 Tw o examples of PDB linkage information and default visualization styles in the Mol* viewer: (LEFT: 5AUS) covalent interaction between heme c ligand and cysteine side chain (CYS10-HEC), and metal coordination between the iron in heme c and a histidine side chain (HIS14-FE); (RIGHT: 4TPL) a covalent carbohydrate linkage between two N-acetyl-D-glucosamine (NAG) ligands (NAG2-NAG1) and covalent N-Glycosylation between a NAG and an asparagine side chain (NAG1-ASN207).
5.4.2 Carbohydrates
The most recent development concerning PDB ligand denitions relates to the annotation of polymeric carbohydrates [33]. These new “branched poly­mer” descriptions (entity_type item “branched”) are composed of multiple and sequential CCD monosaccharides. The description of polymeric carbohydrates is outlined within the PDBx/mmCIF model le for a given PDB entry, under the new pdbx categories entity_branch, entity_branch_descriptor, entity_branch_link, entity_branch_list, and branch_scheme (Figure 5.4a). Whilst chemically distinct
(a) (b)
Figure 5.4 (a) Example categories from a PDBx/mmCIF PDB model file describing a branched polymer carbohydrate. The example highlights how entity 2 (raffinose (RAF), PRD_900002) is comprised of the subcomponents FRU, GLC, and GLA (fructose, glucose, and galactose) and lists the corresponding atomic linkages. (b) Examples of wwPDB carbohydrate representations by 2D and 3D symbol nomenclature for glycans (SNFG) (top and middle, respectively). Source: (b) These images were taken from the PDBe entry pages for PDB accession 5ofx; https://pdbe.org/5ofx.
6
α α
β
2