Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5942_Библиотеки_им_академика_М_И_Перельмана
.pdf
5.2 Small Molecule Data in Protein Data Bank (PDB) Entries 143
https://t.me/medicina_free
Figure 5.1 A sample of the atomic coordinates from a PDBx/mmCIF format file (PDB entry
2yi7) is shown on the left. The individual items in the “atom_site” category are listed first,
highlighting the data present in each column for the subsequent data. A subset of the
atomic coordinates is displayed below, with graphical representation of these coordinates,
displayed on the right of the image, only selected atoms are labeled for clarity.
relating to this molecule. More information about ligand molecules in the PDB will
be discussed in Section 5.3 (Figure 5.1).
As previously mentioned, the PDB is an archive of experimentally determined
structures and, as such, is limited to a subset of accepted experimental structure determination techniques. Depending on the experimental method used,
the archive PDB entry le will contain information specic to the technique,
while additional experimental data les are also collected to allow assessment
of experimental data in the context of the derived coordinate models. The three
main methods accepted for PDB depositions are diraction techniques such as
X-ray crystallography, cryo-electron microscopy (cryoEM), and nuclear magnetic
resonance (NMR). A summary of the PDB deposition data requirements for each of
these techniques is given in Table 5.1.
The oldest and most common technique for determining structures in the PDB is
X-ray crystallography, which involves the generation of crystalline structures of the
biological sample. These crystals are then exposed to a high-powered X-ray beam,
often at a synchrotron facility, and the specic diraction pattern from these X-rays
is collected by a detector. Before this diraction pattern can be used for structure
calculation, rst the phases must be determined, using techniques such as molecular
or isomorphous replacement.
Using the intensities and phases of the spots (or reections) in the diraction
pattern, crystallographers can calculate a map of electron density for the biological specimen, into which the atomic model can be built. In addition to coordinates,
since 2008, submission of PDB structures solved by X-ray crystallography must also
include submission of the structure factor les dening the experimental reections data.

144 5 The Protein Data Bank (PDB) and Macromolecular Structure Data Supporting Computer
https://t.me/medicina_free
Table 5.1 A summary of the main experimental methods accepted for PDB depositions,
including information on the type of data deposited in a PDB entry file, the mandatory
requirements for deposition, the related experimental data and where it should be
deposited, and the raw experimental data and where it is recommended to be deposited.
Experimental
method
X-ray
diraction
Nuclear
Magnetic
Resonance
(NMR)
Cryo-electron
microscopy
(cryoEM)
Model solution
method
Single model
built into
experimentally
derived electron
density map
Multiple models
representing a
range of
conformers that
satisfy
experimental
restraints
Single model
built into
experimentally
derived electric
potential map
PDB deposition
requirements
Coordinate le
(PDBx/mmCIF
format)
Structure factor
le (mtz or CIF
format)
Coordinate le
(PDB or
PDBx/mmCIF
format)
NMR restraints
NEF format)
Coordinate le
(PDB or
PDBx/mmCIF
format)
Image for public
display
at EMDB
EMDB map
deposition
(MRC or CCP4
format)
Experimental
data (archive)
Structure factors
(PDB)
NMR restraints
(BMRB)
NMR chemical
shifts (BMRB)
Electric
potential map
(EMDB)
Raw data
(archive)
X-ray diraction
image data
(SBGrid,
IRRMC)
NMR spectral
parameters
(BMRB)
NMR relaxation
data (BMRB)
Electron
microscopy
images
(EMPIAR)
The rst cryoEM structure in the PDB was released in 1991, however, it has taken a
long time for the technique to establish itself as a routine method for high-resolution
structure determination. The technique involves the imaging of biological molecules
and complexes under cryogenic conditions, using a transmission electron microscope. The sample is ash-frozen in a thin layer with electrons passing through the
sample to a detector, where images of each particle are captured. These images are
then sorted and processed computationally to generate a 3D map of the sample, to
allow tting of a molecular model.
The use of cryoEM has allowed the determination of larger and more heterogeneous complexes than was previously possible with X-ray crystallography, however
until recently, was not able to reach resolutions sucient to interpret atomic-level
details. However, recent advances in the technique, including improved software
and hardware including free-electron detectors, have led to cryoEM becoming the
fastest-growing experimental method for studying macromolecular structure [21].
This is due to the determination of high-resolution structures in conditions that

5.2 Small Molecule Data in Protein Data Bank (PDB) Entries 145
https://t.me/medicina_free
better represent the native environment of the sample. Since 2016, submission of
cryoEM structures to the PDB has also required submission of the experimental
maps to the EMDB [6].
The third main technique used for solving structures in the PDB is NMR. This
method involves the use of powerful magnetic elds to discern minor dierences in
the resonance frequencies, also known as chemical shifts, of atoms in the macromolecule. These chemical shifts are used to determine interactions among atoms in
the molecule and generate atomic restraints, which can be used to build the structure
of the macromolecules.
In solution-state NMR, the biological macromolecules can move freely in the solvent, allowing these experiments to capture the range of structural conformations
adopted by these molecules. These structures are therefore deposited as multiple
models in the PDB, highlighting an ensemble of potential structural conformations
adopted by the macromolecules. An alternative technique is solid-state NMR [22],
which uses similar underlying principles, however, is used to determine structures of
molecular structures within solid or semisolid materials, including macromolecules
within biological membranes.
Though the three experimental methods mentioned above account for the vast
majority of PDB entries, there are also additional variations in diraction methods
that can be used to solve structures in the PDB. Firstly, neutron crystallography [23]
is a similar technique to X-ray crystallography, however, it relies upon neutron scattering to determine the macromolecular structure. The benet of this technique is
that the neutrons interact with atomic nuclei, rather than electrons, which improves
observation of hydrogen atoms in the structure. Neutron and X-ray crystallography
are often used in conjunction to provide both high-resolution data and to determine
positions of hydrogen atoms.
There are also methods that utilize X-raycrystallography to allow high-throughput
determination of ligand binding in macromolecules. Fragment screening [24] experiments involve the “soaking” of various small molecule fragments into crystals
of the sample. Determining how these dierent fragments bind within the protein structure can help to improve the mechanistic understanding of a protein
and the identication of suitable drug candidates. Pan-Dataset Density Analysis (PanDDA) experiments [25] can be used to compare multiple fragment
screening datasets to identify weak signals of bound small molecules within the
noise of the data and can further improve the determination of small molecule
binding sites.
5.2.2 Definition of Small Molecules in OneDep
In the PDB archive, ligands are dened as any molecule that is not part of a larger
polymeric molecule, i.e. proteins, nucleic acids, or branched carbohydrates [26].
In the case of peptide-like ligands, if these contain at least two standard peptide
bonds, then these are classed as polymeric entities. Small molecule ligands in the
PDB can have a range of distinct functions, for example, as substrates, products, and
inhibitors.

146 5 The Protein Data Bank (PDB) and Macromolecular Structure Data Supporting Computer
https://t.me/medicina_free
Each distinct ligand in the PDB archive is assigned its own unique 3-character ID
code, to support easy identication across the full archive [26]. Each of these ligands
includes a detailed chemical description, which is contained within the wwPDB
Chemical Component Dictionary (CCD), and discussed in more detail in the next
section.
5.3 Small Molecule Dictionaries
5.3.1 wwPDB Chemical Component Dictionary (CCD)
The wwPDB maintains a reference dictionary containing chemical descriptions
of each unique chemical component present in the whole PDB archive. These
components are the building blocks for PDB entries and include amino acids,
nucleotides, metal ions, and other nonpeptide small molecules. These chemical
component descriptions are stored in the wwPDB’s CCD (https://www.wwpdb.org/
data/ccd) [26].
Each component is given a unique identier and has a CCD denition that
describes the molecule. As of July 2022, there are around 37,000 unique denitions
in the wwPDB CCD. The unique identier is currently limited to a maximum
of three characters. Most denitions contain a three-character unique identier,
with one- or two-character identiers mostly used for identication of DNA and
RNA bases or for components containing individual elements. Some examples of
chemical components include manganese (Mn), adenosine-triphosphate (ATP),
and glutamate (GLU). Previously, these identiers were assigned with some
meaning to their chemical denition, however, due to the high number of chemical
components in the dictionary, these CCD identiers are now randomly assigned for
any new ligands.
If the CCD denition is dening a standard amino acid or nucleotide, the IUPAC
protein one-letter code [27] is also provided, for example, Glutamate has the
unique identier GLU and the one-letter code E. Small modications of standard
amino acids or nucleotides, generally, where the modication contains fewer than
10 atoms, have their own CCD denitions, for example, phospho-serine has the
identier SEP. These modied amino acids or nucleotides have the same one-letter
code as their standard amino acid parent, therefore SEP, which is a modied serine
amino acid, also contains the one-letter code S.
Peptide-based chromophores provide a more complex example. In proteins,
chromophores typically comprise three amino acids condensed into one molecule.
In the CCD component, the one-letter code lists all amino acids that make up the
chromophore. For example, the CCD denition for the chromophore CR2, which
comprises the amino acids GLY-TYR-GLY and has the corresponding one-letter
code GYG.
In addition to providing a unique identier, each CCD denition describes the
chemistry of the molecule. This includes the name of each atom in the molecule,
the order of bonds between atoms, stereochemistry of chiral atoms, and any charges

5.3 Small Molecule Dictionaries 147
https://t.me/medicina_free
on individual atoms or the whole molecule. There are two sets of 3D coordinates
included in the CCD denition: the model coordinates, extracted from the PDB entry
from which the CCD was initially generated, and idealized coordinates (from the
lowest energy conformation), generated using the expected geometry of the ligand
from the CCD denition.
Additional metadata are also included in the CCD denition, including the name
of the molecule and any synonyms. If a common name is known for the compound,
then it is provided either as the name or as one of the synonyms, otherwise, systematic names are provided. The systematic name is the default name generated by the
OpenEye [28] software at the time of creation of the compound. Chemical descriptors in InChI and SMILES formats are provided and are automatically generated by
OpenEye during creation of the compound.
Every molecule in the PDB must have a corresponding CCD denition. During
deposition and biocuration of PDB entries, each compound in the deposited le is
compared to existing CCD denitions using a graph match algorithm [29]. If renement restraints are provided in the uploaded coordinate le, then this information
is used for comparing the molecule to entries in the CCD, otherwise, the coordinates
are used for searching. If a molecule fails to nd a match in the CCD, then depositors are asked for further details of the molecule so that a new CCD denition can
be made. A new CCD denition is then dened during biocuration and added to
the CCD. The name of the molecule is either dened by the depositor, if a common
name is available, or a systematic name is generated based on the chemistry of the
molecule.
5.3.2 The Peptide Reference Dictionary
The PDB archive also contains examples of complex ligands, which comprise
combinations of amino acid and amino acid-like components. These peptide-like
molecules commonly possess important biological functions, such as antibiotic
or inhibitory activity, however, a description of their global chemistry does not t
within the classic PDB denitions of polymers and ligands. Therefore, to support the
curation of this important class of molecules, their chemical and biological descriptions are described in a separate dictionary. The “Peptide Reference Dictionary”
(PRD) resource [30] provides a standardized framework for these denitions so
that each PRD entity can be described in a way that outlines both its subcomponent
composition and global characteristics. These PRD denitions are stored in the
wwPDB Biologically Interesting Molecule Reference Dictionary (BIRD) (https://
www.wwpdb.org/data/bird).
The process of dening a new PRD molecule, or matching connected CCDs to
an existing PRD, occurs during the biocuration process. Any atomistic information,
such as reference coordinates, bonding, and chemical descriptors are then listed
in PRD-chemical denition les, whereas subcomponent sequence, CCD-linking,
and any related biological details, such as biological function, are stored in
PRD-molecular denition les. By using the PRD dictionary to standardize biocuration, it allows for the precise interrogation, comparison, and retrieval of global and

148 5 The Protein Data Bank (PDB) and Macromolecular Structure Data Supporting Computer
https://t.me/medicina_free
Figure 5.2 An example of PRD is the antibiotic Vancomycin (PRD_000204). Vancomycin
consists of seven nonstandard amino acids (ball and stick, distinguished by color) and two
sugar components (cyan and grey in a 3D representation of the SNFG notation). The PRD
polymer linking section of the PRD molecular definition file is shown on the right-hand side.
subcomponent-specic data, which would otherwise not be possible. Links to PRD
denition les can be found on relevant PDB entry pages and references to key PRD
identiers (ID, name, type, and function) are stored in PDB entry PDBx/mmCIF
les under the pdbx_molecule_features category (Figure 5.2).
5.4 Additional Ligand Annotations in the PDB Archive
In addition to the standard chemical and biomolecular denitions already described,
where possible, the PDB biocuration pipeline generates additional ligand-related
information to best describe ligand complexity and intramolecular connectivity and
to do so with a consistent and objective methodology [31]. Many software packages
rely upon these data for purposes such as molecular renement and accurate molecular visualization.
5.4.1 Linkage Information
The most common ligand-associated annotations found in PDB entry les refer
to intermolecular bonding. These refer to all nonstandard linkages between and
among polymer components and CCD ligands. This information was historically
found in PDB-format LINK, SSBOND, and CONECT records, however, the current
PDBx/mmCIF archive format provides additional data, including leaving group
information, atomic descriptions, and author numbering for up to three chemical
component partners per interaction [9]. At present, to ensure standardization across
the whole PDB archive, all link records and bonding distances are calculated and
validated during the biocuration process [31]. The PDBx/mmCIF dictionary currently supports covalent, hydrogen, covalent–metal, and disulde-bond interaction
“types” [32] (Figure 5.3).

5.4 Additional Ligand Annotations in the PDB Archive 149
https://t.me/medicina_free
(a) (b)
Figure 5.3 Two examples of PDB linkage information and default visualization styles in
the Mol* viewer: (LEFT: 5AUS) covalent interaction between heme c ligand and cysteine
side chain (CYS10-HEC), and metal coordination between the iron in heme c and a histidine
side chain (HIS14-FE); (RIGHT: 4TPL) a covalent carbohydrate linkage between two
N-acetyl-D-glucosamine (NAG) ligands (NAG2-NAG1) and covalent N-Glycosylation between
a NAG and an asparagine side chain (NAG1-ASN207).
5.4.2 Carbohydrates
The most recent development concerning PDB ligand denitions relates to
the annotation of polymeric carbohydrates [33]. These new “branched polymer” descriptions (entity_type item “branched”) are composed of multiple and
sequential CCD monosaccharides. The description of polymeric carbohydrates is
outlined within the PDBx/mmCIF model le for a given PDB entry, under the
new pdbx categories entity_branch, entity_branch_descriptor, entity_branch_link,
entity_branch_list, and branch_scheme (Figure 5.4a). Whilst chemically distinct
(a) (b)
Figure 5.4 (a) Example categories from a PDBx/mmCIF PDB model file describing a
branched polymer carbohydrate. The example highlights how entity 2 (raffinose (RAF),
PRD_900002) is comprised of the subcomponents FRU, GLC, and GLA (fructose, glucose, and
galactose) and lists the corresponding atomic linkages. (b) Examples of wwPDB
carbohydrate representations by 2D and 3D symbol nomenclature for glycans (SNFG) (top
and middle, respectively). Source: (b) These images were taken from the PDBe entry pages
for PDB accession 5ofx; https://pdbe.org/5ofx.
6
αα
2
β

150 5 The Protein Data Bank (PDB) and Macromolecular Structure Data Supporting Computer
https://t.me/medicina_free
from the peptide-like PRD molecules, carbohydrates of biological signicance may
also be described by the PRD dictionary les where they hold functional classications such as “nutrients” or “inhibitors.” The annotation of carbohydrate-branched
polymers occurs within the automatic processing of the biocuration pipeline.
Dierent carbohydrate visual representations may now also be found on wwPDB
partner websites, using the community standard SNFG notation [34, 35], both as
2D images on each entry page, and 3D-SNFG [36] units within the Mol* viewer
(Figure 5.4b).
5.5 Validation of Ligands in the Worldwide Protein Data
Bank (wwPDB)
Currently, PDB contains over 34,000 unique small molecules bound to over 133,600
protein structures. Numerous groups have highlighted the need for standardized
validation in structures in the PDB archive [37–44]. Several errors have even led to
the retraction of the respective publication(s) and subsequent obsoletion of the PDB
entry [43]. A Validation Task Force (VTF) was set up to centralize and standardize
validation across the wwPDB archive, which included metrics to judge the reliability
of the ligand-bound macromolecular structures as they play a pivotal role in many
computational methods like ligand docking and structure-based drug design. This
resulted in development of the wwPDB validation reports (VR), rst introduced in
2012.
wwPDB validation pipeline (VP) generates detailed wwPDB VR comprising
an assessment of structure quality using widely accepted standards and criteria
recommended by various method-specic VTFs [45–47]. To enhance ligand validation, in 2015, wwPDB, CCDC (Cambridge Crystallographic Data Center), and
D3R (Drug Design Data Resource) co-organized the ligand validation workshop
(LVW) [48] whose recommendations were recently implemented in VP [14]. VP
is an integral part of the unied OneDep system for structure deposition [13],
validation [49], and biocuration [31]. Before depositing the structure in the wwPDB
archive, the depositor can generate a wwPDB VR using the standalone validation
server (validate.wwpdb.org) or API (wwpdb.org/validation/onedep-validationweb-service-interface) to nd and x any issues highlighted in the VR. At the
time of deposition, a “preliminary” VR is issued to the depositor, followed by the
“ocial” VR after the completion of biocuration process. Once the structure is
released, the VR is made public and is provided as both human-readable PDF les
and machine-readable mmCIF and XML les. In the next section, we will further
discuss various criteria and software used for validation of ligands in the VR.
5.5.1 Various Criteria and Software Used for Validating Ligand
in Validation Reports
wwPDB validation criteria for ligands can be broadly divided into two categories (I)
geometric validation of the atomic coordinates, without considering the associated

5.5 Validation of Ligands in the Worldwide Protein Data Bank (wwPDB) 151
https://t.me/medicina_free
Table 5.2 Summary of the software used for validating ligands in wwPDB validation
reports.
Property Detailed steps involved Software used Software URL
Atomic
coordinates of all
bound ligands
Geometric quality Computes following
2D diagrams of
geometric quality
3D graphical
depiction of the
model t to data
Extracts ligand of
interest (LOI)
reorder and rename
atoms in ligands
according to the
Chemical Component
Dictionary (CCD)
for all ligands:
● Bond-length,
bond-angle, torsion
angle and ring
outliers
Chirality and planarity
outliers
Highlights geometric
analysis provided by
CCDC Mogul
Each model t to
experimental electron
density is shown from
dierent orientations
to approximate a 3D
view
Maxit
(Macromolecular
Exchange and
Input Tool) [50]
Mogul [51]
Validation-pack
[50]
buster-report [43] https://www
buster-report [43] As above
https://sw-tools
.rcsb.org/apps/
MAXIT
https://ccdc.cam.ac
.uk/solutions/csdcore/components/
mogul
https://mmcif
.wwpdb.org/docs/
software-resources
.html
.globalphasing
.com/buster/
manual/
autobuster/
manual/
autoBUSTER9.html
experimental data, and (II) validation of t between the atomic coordinates of the
ligand and the associated experimental data. The summary of various software used
in VR is shown in Table 5.2.
5.5.2 Identification of Ligand of Interest (LOI)
Since 2017, the depositors are required to identify the ligand(s) in the structure that is
either the focus of their study or is considered biologically important. This LOI information is recorded during the deposition of structure in the OneDep system [13].
All its details can be found in the “Entry composition” section of the VR. In the VP,
Maxit [50] extracts the atomic coordinates of ligand(s), reorder and rename its atoms
according to the CCD after which geometric quality of these is assessed further.

152 5 The Protein Data Bank (PDB) and Macromolecular Structure Data Supporting Computer
https://t.me/medicina_free
5.5.3 Geometric and Conformational Validation
Following the recommendation of wwPDB X-ray VTF, ligand geometry is validated
against the Cambridge Structural Database (CSD) of high-quality small molecule
organic structure [45]. Firstly, CSD structures with similar bond order are identied
and then a distribution for the observed values for each bond length, bond angle,
torsion angle, or ring in the ligand is built using the program Mogul [51]. It then
computes the Z-score, which quanties deviation from the observed values. Bond
lengths and bond angles with absolute Z-score >2.0 are agged as “outliers” in the
VR. The normalized Z-scores (RMSZ) of bond lengths/angles are calculated to facilitate the overall assessment of the ligand. RMSZ scores are expected to lie between
0 and 1. For low-resolution structures, geometry should be tightly restrained and
small values are expected. For very high-resolution structures, values approaching 1
may be attained. Values greater than 1 generally indicate over-tting of the data. For
acyclic torsion angles, Mogul computes the local density measure [52], which is the
ratio of incidences in the CSD within 10∘of the torsion angle in question, to the number of total incidences of the torsion angles in the CSD. The torsion angle is agged
as an outlier if the local density measure is <5%. For isolated rings, Mogul compares
the given ring with comparable rings in small molecule structures in the CSD and
calculates an RMSD value based on corresponding constituent torsion angles for
each comparable ring. If both the mean and minimum of these RMSDs are >60∘,
the ring is agged as an outlier. Chirality outliers are calculated by Validation-pack
[14] and are assessed based on chiral volume. If the sign of the computed volume
is incorrect, the handedness is wrong. If the absolute volume is less than 0.7Å
chiral center is as a planar moiety, which is highly likely to be erroneous. All these
geometric outliers are listed in the “Ligand geometry” section of the VR.
Apart from merely identifying the geometric outliers, it is also important to know
the extent of distortion from the Mogul expectation. After the latest update of VP
[14], this information is now clear in 2D colored images generated by Buster-report
[53]. This image depicts Mogul quality analysis of bond lengths, bond angles, torsion angles, and ring geometry [43] and is added to the “Ligand geometry” section
of VR. Color scheme used in these images is coded according to validation results
with green indicating commonly observed values, magenta indicating unusual values, and grey indicating that there was insucient data to derive a validation score
(see Figure 5.5b). Unusual values include model quality and electron density t.
For model quality, individual bond lengths or angles with absolute Z-score >2, the
torsion angle with <5% of local density measure, or RMSD >60∘are considered
unusual and colored in magenta. These images can be helpful for depositors to identify any putative incorrect restraint values used during renement before depositing
the structure.
3
,the
5.5.4 Ligand Fit to Experimental Electron Density Validation
Apart from validating the geometric quality of the ligand, it is also crucial to
inspect local agreement of a structure model with the observed electron density and
Соседние файлы в папке Библиотека им академика М.И. Перельмана
