Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5938_Библиотеки_им_академика_М_И_Перельмана
.pdf
https://t.me/medicina_free

5
https://t.me/medicina_free
The Protein Data Bank (PDB) and Macromolecular Structure
Data Supporting Computer-Aided Drug Design
David Armstrong1, John Berrisford2, Preeti Choudhary1, Lukas Pravda3, James
Tolchard4, Mihaly Varadi1, and Sameer Velankar
1
Protein Data Bank in Europe, EMBL-EBI, Wellcome Genome Campus, Hinxton, Cambridge, CB10 1SD, UK
2
AstraZeneca, AstraZeneca Academy House, 136 Hills Rd, Cambridge, CB2 8PA, UK
3
Exscientia, The Schrodinger Building, Oxford Science Park, Oxford, Oxfordshire, United Kingdom, OX4 4GE
4
Centre de RMN à Tr è s Hauts Champs de Lyo n , Claude Bernard University Ly on 1, 5 rue de la Doua, 69100
Villeurbanne, France
1
5.1 Introduction
The Protein Data Bank (PDB) [1], managed by the global Worldwide PDB (wwPDB)
consortium [2], is one of the oldest scientic databases in life sciences. The PDB
archives structural models of biological macromolecules, derived from experimental
data. The wwPDB partners, the Research Collaboratory for Structural Bioinformatics (RCSB) [3], Protein Data Bank in Europe (PDBe) [4], Protein Data Bank Japan
(PDBj) [5], Electron Microscopy Data Bank (EMDB) [6], and Bio Mag Res Bank
(BMRB) [7], manage the PDB archive based on the FAIR principles [8], ensuring
structural biology data are Findable, Accessible, Interoperable and Reusable.
The PDB archive contains over 200,000 structures of proteins, DNA, and RNA, and
their complexes with small molecules, such as cofactors and inhibitors. This data,
along with related metadata and experimental data, are stored in the PDB archive in
PDBx/mmCIF formatted les [9]. Since 2019, it has been mandatory for X-ray crystallographic structures to be deposited in the PDBx/mmCIF format [10], ensuring
maximum capture and validation of metadata for these structures.
The PDB is a vital resource for computer-aided drug design and contains structural information for large numbers of drugs, with 5494 unique ligands in the PDB
mapped to entries in DrugBank [11]. Additionally, there are numerous potential
drug targets as well as structures containing macromolecular drug targets. This data
are of paramount importance in determination of key drug binding sites in macromolecules [12], and in understanding existing binding modes to support designs of
novel compounds to target these binding sites.
In addition to the three-dimensional (3D) coordinate information archived in
the PDB, standardization of polymer sequences and chemical identiers in PDB
entries facilitates the integration of additional metadata from external resources
141
Open Access Databases and Datasets for Drug Discovery, First Edition.
Edited by Antoine Daina, Michael Przewosny, and Vincent Zoete.
© 2024 WILEY-VCH GmbH. Published 2024 by WILEY-VCH GmbH.

142 5 The Protein Data Bank (PDB) and Macromolecular Structure Data Supporting Computer
https://t.me/medicina_free
[13, 14]. This additional biological and chemical information is vitally important in
understanding the scientic context of the structures in the PDB, which is required
to determine functional information.
As part of the European Bioinformatics Institute (EMBL-EBI), the PDBe [4] is
uniquely placed to integrate macromolecular structure data from partner resources
at EBI and beyond. This allows the integration of relevant biological and chemical
data to better demonstrate the functional importance of structures in the PDB. The
PDBe-Knowledge Base (PDBe-KB) [15] builds upon this, with partnerships extending across the structural bioinformatics community, allowing the integration of even
more data related to macromolecular structure and function. This includes a large
number of partners involved in cheminformatics who provide data relating to small
molecules and their macromolecular binding sites, including canSAR [16], 3DLigandSite [17], P2Rank [18], and many more, detailed at pdbe-kb.org/partners.
The wwPDB, by adhering to the FAIR principles, ensures that all PDB data are provided freely, with no limits upon its use. All PDBe and PDBe-KB tools and resources
are also free to use, with scripts and software pipelines made open source wherever
possible. In addition to access through the PDBe [4] and PDBe-KB [15] websites, all
PDBe data are made available through publicly accessible APIs [19, 20], while much
of this data is also available via a distributed PDBe knowledge graph [20].
This chapter will introduce the type of data available in the PDB, how it is organized and curated, and how it can be used to support drug design. It will also give
an overview of the tools and resources at PDBe and PDBe-KB, highlighting how
these can support understanding of drug binding and function in PDB structures to
improve research within this area.
5.2 Small Molecule Data in Protein Data Bank (PDB)
Entries
5.2.1 What Data are in the PDB Archive?
The PDB is the single global archive of experimentally determined 3D structure
data of biological macromolecules. Each PDB structure must contain a polymeric
entity, i.e. protein or nucleic acid, however, information can also be included for all
non-polymeric ligands within the structure.
The atomic coordinates of a PDB entry are built with a specic hierarchy, ensuring
that each molecule and its constituent parts can be clearly identied. The smallest
component of the hierarchy is at the atomic level, with each separate coordinate line
in the archive le describing the specic position of an atom in 3D space. Each atom
is dened as part of a larger chemical component or residue, dened by a 3-character
ID code.
These chemical components can be either individual bound ligands in the structure or individual residues within a polymeric molecule, with each dened by a
unique identier (asym or chain ID) and numbering for clear identication. Beyond
each specic polymeric chain ID, each unique and individual molecule in the structure is also given a specic “entity” identier, which is used to link all the metadata

5.2 Small Molecule Data in Protein Data Bank (PDB) Entries 143
https://t.me/medicina_free
Figure 5.1 A sample of the atomic coordinates from a PDBx/mmCIF format file (PDB entry
2yi7) is shown on the left. The individual items in the “atom_site” category are listed first,
highlighting the data present in each column for the subsequent data. A subset of the
atomic coordinates is displayed below, with graphical representation of these coordinates,
displayed on the right of the image, only selected atoms are labeled for clarity.
relating to this molecule. More information about ligand molecules in the PDB will
be discussed in Section 5.3 (Figure 5.1).
As previously mentioned, the PDB is an archive of experimentally determined
structures and, as such, is limited to a subset of accepted experimental structure determination techniques. Depending on the experimental method used,
the archive PDB entry le will contain information specic to the technique,
while additional experimental data les are also collected to allow assessment
of experimental data in the context of the derived coordinate models. The three
main methods accepted for PDB depositions are diraction techniques such as
X-ray crystallography, cryo-electron microscopy (cryoEM), and nuclear magnetic
resonance (NMR). A summary of the PDB deposition data requirements for each of
these techniques is given in Table 5.1.
The oldest and most common technique for determining structures in the PDB is
X-ray crystallography, which involves the generation of crystalline structures of the
biological sample. These crystals are then exposed to a high-powered X-ray beam,
often at a synchrotron facility, and the specic diraction pattern from these X-rays
is collected by a detector. Before this diraction pattern can be used for structure
calculation, rst the phases must be determined, using techniques such as molecular
or isomorphous replacement.
Using the intensities and phases of the spots (or reections) in the diraction
pattern, crystallographers can calculate a map of electron density for the biological specimen, into which the atomic model can be built. In addition to coordinates,
since 2008, submission of PDB structures solved by X-ray crystallography must also
include submission of the structure factor les dening the experimental reections data.

144 5 The Protein Data Bank (PDB) and Macromolecular Structure Data Supporting Computer
https://t.me/medicina_free
Table 5.1 A summary of the main experimental methods accepted for PDB depositions,
including information on the type of data deposited in a PDB entry file, the mandatory
requirements for deposition, the related experimental data and where it should be
deposited, and the raw experimental data and where it is recommended to be deposited.
Experimental
method
X-ray
diraction
Nuclear
Magnetic
Resonance
(NMR)
Cryo-electron
microscopy
(cryoEM)
Model solution
method
Single model
built into
experimentally
derived electron
density map
Multiple models
representing a
range of
conformers that
satisfy
experimental
restraints
Single model
built into
experimentally
derived electric
potential map
PDB deposition
requirements
Coordinate le
(PDBx/mmCIF
format)
Structure factor
le (mtz or CIF
format)
Coordinate le
(PDB or
PDBx/mmCIF
format)
NMR restraints
le (STAR or
NEF format)
Coordinate le
(PDB or
PDBx/mmCIF
format)
Image for public
display
at EMDB
EMDB map
deposition
(MRC or CCP4
format)
Experimental
data (archive)
Structure factors
(PDB)
NMR restraints
(BMRB)
NMR chemical
shifts (BMRB)
Electric
potential map
(EMDB)
Raw data
(archive)
X-ray diraction
image data
(SBGrid,
IRRMC)
NMR spectral
parameters
(BMRB)
NMR relaxation
data (BMRB)
Electron
microscopy
images
(EMPIAR)
The rst cryoEM structure in the PDB was released in 1991, however, it has taken a
long time for the technique to establish itself as a routine method for high-resolution
structure determination. The technique involves the imaging of biological molecules
and complexes under cryogenic conditions, using a transmission electron microscope. The sample is ash-frozen in a thin layer with electrons passing through the
sample to a detector, where images of each particle are captured. These images are
then sorted and processed computationally to generate a 3D map of the sample, to
allow tting of a molecular model.
The use of cryoEM has allowed the determination of larger and more heterogeneous complexes than was previously possible with X-ray crystallography, however
until recently, was not able to reach resolutions sucient to interpret atomic-level
details. However, recent advances in the technique, including improved software
and hardware including free-electron detectors, have led to cryoEM becoming the
fastest-growing experimental method for studying macromolecular structure [21].
This is due to the determination of high-resolution structures in conditions that

5.2 Small Molecule Data in Protein Data Bank (PDB) Entries 145
https://t.me/medicina_free
better represent the native environment of the sample. Since 2016, submission of
cryoEM structures to the PDB has also required submission of the experimental
maps to the EMDB [6].
The third main technique used for solving structures in the PDB is NMR. This
method involves the use of powerful magnetic elds to discern minor dierences in
the resonance frequencies, also known as chemical shifts, of atoms in the macromolecule. These chemical shifts are used to determine interactions among atoms in
the molecule and generate atomicrestraints, which can be used to build the structure
of the macromolecules.
In solution-state NMR, the biological macromolecules can move freely in the solvent, allowing these experiments to capture the range of structural conformations
adopted by these molecules. These structures are therefore deposited as multiple
models in the PDB, highlighting an ensemble of potential structural conformations
adopted by the macromolecules. An alternative technique is solid-state NMR [22],
which uses similar underlying principles, however, is used to determine structures of
molecular structures within solid or semisolid materials, including macromolecules
within biological membranes.
Though the three experimental methods mentioned above account for the vast
majority of PDB entries, there are also additional variations in diraction methods
that can be used to solve structures in the PDB. Firstly, neutron crystallography [23]
is a similar technique to X-ray crystallography, however, it relies upon neutron scattering to determine the macromolecular structure. The benet of this technique is
that the neutrons interact with atomic nuclei, rather than electrons, which improves
observation of hydrogen atoms in the structure. Neutron and X-ray crystallography
are often used in conjunction to provide both high-resolution data and to determine
positions of hydrogen atoms.
There are also methods that utilize X-raycrystallography toallow high-throughput
determination of ligand binding in macromolecules. Fragment screening [24] experiments involve the “soaking” of various small molecule fragments into crystals
of the sample. Determining how these dierent fragments bind within the protein structure can help to improve the mechanistic understanding of a protein
and the identication of suitable drug candidates. Pan-Dataset Density Analysis (PanDDA) experiments [25] can be used to compare multiple fragment
screening datasets to identify weak signals of bound small molecules within the
noise of the data and can further improve the determination of small molecule
binding sites.
5.2.2 Definition of Small Molecules in OneDep
In the PDB archive, ligands are dened as any molecule that is not part of a larger
polymeric molecule, i.e. proteins, nucleic acids, or branched carbohydrates [26].
In the case of peptide-like ligands, if these contain at least two standard peptide
bonds, then these are classed as polymeric entities. Small molecule ligands in the
PDB can have a range of distinct functions, for example, as substrates, products, and
inhibitors.

146 5 The Protein Data Bank (PDB) and Macromolecular Structure Data Supporting Computer
https://t.me/medicina_free
Each distinct ligand in the PDB archive is assigned its own unique 3-character ID
code, to support easy identication across the full archive [26]. Each of these ligands
includes a detailed chemical description, which is contained within the wwPDB
Chemical Component Dictionary (CCD), and discussed in more detail in the next
section.
5.3 Small Molecule Dictionaries
5.3.1 wwPDB Chemical Component Dictionary (CCD)
The wwPDB maintains a reference dictionary containing chemical descriptions
of each unique chemical component present in the whole PDB archive. These
components are the building blocks for PDB entries and include amino acids,
nucleotides, metal ions, and other nonpeptide small molecules. These chemical
component descriptions are stored in the wwPDB’s CCD (https://www.wwpdb.org/
data/ccd) [26].
Each component is given a unique identier and has a CCD denition that
describes the molecule. As of July 2022, there are around 37,000 unique denitions
in the wwPDB CCD. The unique identier is currently limited to a maximum
of three characters. Most denitions contain a three-character unique identier,
with one- or two-character identiers mostly used for identication of DNA and
RNA bases or for components containing individual elements. Some examples of
chemical components include manganese (Mn), adenosine-triphosphate (ATP),
and glutamate (GLU). Previously, these identiers were assigned with some
meaning to their chemical denition, however, due to the high number of chemical
components in the dictionary, these CCD identiers are now randomly assigned for
any new ligands.
If the CCD denition is dening a standard amino acid or nucleotide, the IUPAC
protein one-letter code [27] is also provided, for example, Glutamate has the
unique identier GLU and the one-letter code E. Small modications of standard
amino acids or nucleotides, generally, where the modication contains fewer than
10 atoms, have their own CCD denitions, for example, phospho-serine has the
identier SEP. These modied amino acids or nucleotides have the same one-letter
code as their standard amino acid parent, therefore SEP, which is a modied serine
amino acid, also contains the one-letter code S.
Peptide-based chromophores provide a more complex example. In proteins,
chromophores typically comprise three amino acids condensed into one molecule.
In the CCD component, the one-letter code lists all amino acids that make up the
chromophore. For example, the CCD denition for the chromophore CR2, which
comprises the amino acids GLY-TYR-GLY and has the corresponding one-letter
code GYG.
In addition to providing a unique identier, each CCD denition describes the
chemistry of the molecule. This includes the name of each atom in the molecule,
the order of bonds between atoms, stereochemistry of chiral atoms, and any charges

5.3 Small Molecule Dictionaries 147
https://t.me/medicina_free
on individual atoms or the whole molecule. There are two sets of 3D coordinates
included in the CCD denition: the model coordinates, extracted from the PDB entry
from which the CCD was initially generated, and idealized coordinates (from the
lowest energy conformation), generated using the expected geometry of the ligand
from the CCD denition.
Additional metadata are also included in the CCD denition, including the name
of the molecule and any synonyms. If a common name is known for the compound,
then it is provided either as the name or as one of the synonyms, otherwise, systematic names are provided. The systematic name is the default name generated by the
OpenEye [28] software at the time of creation of the compound. Chemical descriptors in InChI and SMILES formats are provided and are automatically generated by
OpenEye during creation of the compound.
Every molecule in the PDB must have a corresponding CCD denition. During
deposition and biocuration of PDB entries, each compound in the deposited le is
compared to existing CCD denitions using a graph match algorithm [29]. If renement restraints are provided in the uploaded coordinate le, then this information
is used for comparing the molecule to entries in the CCD, otherwise, the coordinates
are used for searching. If a molecule fails to nd a match in the CCD, then depositors are asked for further details of the molecule so that a new CCD denition can
be made. A new CCD denition is then dened during biocuration and added to
the CCD. The name of the molecule is either dened by the depositor, if a common
name is available, or a systematic name is generated based on the chemistry of the
molecule.
5.3.2 The Peptide Reference Dictionary
The PDB archive also contains examples of complex ligands, which comprise
combinations of amino acid and amino acid-like components. These peptide-like
molecules commonly possess important biological functions, such as antibiotic
or inhibitory activity, however, a description of their global chemistry does not t
within the classic PDB denitions of polymers and ligands. Therefore, to support the
curation of this important class of molecules, their chemical and biological descriptions are described in a separate dictionary. The “Peptide Reference Dictionary”
(PRD) resource [30] provides a standardized framework for these denitions so
that each PRD entity can be described in a way that outlines both its subcomponent
composition and global characteristics. These PRD denitions are stored in the
wwPDB Biologically Interesting Molecule Reference Dictionary (BIRD) (https://
www.wwpdb.org/data/bird).
The process of dening a new PRD molecule, or matching connected CCDs to
an existing PRD, occurs during the biocuration process. Any atomistic information,
such as reference coordinates, bonding, and chemical descriptors are then listed
in PRD-chemical denition les, whereas subcomponent sequence, CCD-linking,
and any related biological details, such as biological function, are stored in
PRD-molecular denition les. By using the PRD dictionary to standardize biocuration, it allows for the precise interrogation, comparison, and retrieval of global and

148 5 The Protein Data Bank (PDB) and Macromolecular Structure Data Supporting Computer
https://t.me/medicina_free
Figure 5.2 An example of PRD is the antibiotic Vancomycin (PRD_000204). Vancomycin
consists of seven nonstandard amino acids (ball and stick, distinguished by color) and two
sugar components (cyan and grey in a 3D representation of the SNFG notation). The PRD
polymer linking section of the PRD molecular definition file is shown on the right-hand side.
subcomponent-specic data, which would otherwise not be possible. Links to PRD
denition les can be found on relevant PDB entry pages and references to key PRD
identiers (ID, name, type, and function) are stored in PDB entry PDBx/mmCIF
les under the pdbx_molecule_features category (Figure 5.2).
5.4 Additional Ligand Annotations in the PDB Archive
In addition to the standard chemical and biomolecular denitions already described,
where possible, the PDB biocuration pipeline generates additional ligand-related
information to best describe ligand complexity and intramolecular connectivity and
to do so with a consistent and objective methodology [31]. Many software packages
rely upon these data for purposes such as molecular renement and accurate molecular visualization.
5.4.1 Linkage Information
The most common ligand-associated annotations found in PDB entry les refer
to intermolecular bonding. These refer to all nonstandard linkages between and
among polymer components and CCD ligands. This information was historically
found in PDB-format LINK, SSBOND, and CONECT records, however, the current
PDBx/mmCIF archive format provides additional data, including leaving group
information, atomic descriptions, and author numbering for up to three chemical
component partners per interaction [9]. At present, to ensure standardization across
the whole PDB archive, all link records and bonding distances are calculated and
validated during the biocuration process [31]. The PDBx/mmCIF dictionary currently supports covalent, hydrogen, covalent–metal, and disulde-bond interaction
“types” [32] (Figure 5.3).

5.4 Additional Ligand Annotations in the PDB Archive 149
https://t.me/medicina_free
(a) (b)
Figure 5.3 Tw o examples of PDB linkage information and default visualization styles in
the Mol* viewer: (LEFT: 5AUS) covalent interaction between heme c ligand and cysteine
side chain (CYS10-HEC), and metal coordination between the iron in heme c and a histidine
side chain (HIS14-FE); (RIGHT: 4TPL) a covalent carbohydrate linkage between two
N-acetyl-D-glucosamine (NAG) ligands (NAG2-NAG1) and covalent N-Glycosylation between
a NAG and an asparagine side chain (NAG1-ASN207).
5.4.2 Carbohydrates
The most recent development concerning PDB ligand denitions relates to
the annotation of polymeric carbohydrates [33]. These new “branched polymer” descriptions (entity_type item “branched”) are composed of multiple and
sequential CCD monosaccharides. The description of polymeric carbohydrates is
outlined within the PDBx/mmCIF model le for a given PDB entry, under the
new pdbx categories entity_branch, entity_branch_descriptor, entity_branch_link,
entity_branch_list, and branch_scheme (Figure 5.4a). Whilst chemically distinct
(a) (b)
Figure 5.4 (a) Example categories from a PDBx/mmCIF PDB model file describing a
branched polymer carbohydrate. The example highlights how entity 2 (raffinose (RAF),
PRD_900002) is comprised of the subcomponents FRU, GLC, and GLA (fructose, glucose, and
galactose) and lists the corresponding atomic linkages. (b) Examples of wwPDB
carbohydrate representations by 2D and 3D symbol nomenclature for glycans (SNFG) (top
and middle, respectively). Source: (b) These images were taken from the PDBe entry pages
for PDB accession 5ofx; https://pdbe.org/5ofx.
6
α α
β
2
Соседние файлы в папке Библиотека им академика М.И. Перельмана
