Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5942_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
15.09.2026
Размер:
11 Мб
Скачать
☆
5.2 Small Molecule Data in Protein Data Bank (PDB) Entries 143
https://t.me/medicina_free
Figure 5.1 A sample of the atomic coordinates from a PDBx/mmCIF format file (PDB entry 2yi7) is shown on the left. The individual items in the “atom_site” category are listed first, highlighting the data present in each column for the subsequent data. A subset of the atomic coordinates is displayed below, with graphical representation of these coordinates, displayed on the right of the image, only selected atoms are labeled for clarity.
relating to this molecule. More information about ligand molecules in the PDB will be discussed in Section 5.3 (Figure 5.1).
As previously mentioned, the PDB is an archive of experimentally determined structures and, as such, is limited to a subset of accepted experimental struc­ture determination techniques. Depending on the experimental method used, the archive PDB entry le will contain information specic to the technique, while additional experimental data les are also collected to allow assessment of experimental data in the context of the derived coordinate models. The three main methods accepted for PDB depositions are diraction techniques such as X-ray crystallography, cryo-electron microscopy (cryoEM), and nuclear magnetic resonance (NMR). A summary of the PDB deposition data requirements for each of these techniques is given in Table 5.1.
The oldest and most common technique for determining structures in the PDB is X-ray crystallography, which involves the generation of crystalline structures of the biological sample. These crystals are then exposed to a high-powered X-ray beam, often at a synchrotron facility, and the specic diraction pattern from these X-rays is collected by a detector. Before this diraction pattern can be used for structure calculation, rst the phases must be determined, using techniques such as molecular or isomorphous replacement.
Using the intensities and phases of the spots (or reections) in the diraction pattern, crystallographers can calculate a map of electron density for the biologi­cal specimen, into which the atomic model can be built. In addition to coordinates, since 2008, submission of PDB structures solved by X-ray crystallography must also include submission of the structure factor les dening the experimental reec­tions data.
144 5 The Protein Data Bank (PDB) and Macromolecular Structure Data Supporting Computer
https://t.me/medicina_free
Table 5.1 A summary of the main experimental methods accepted for PDB depositions, including information on the type of data deposited in a PDB entry file, the mandatory requirements for deposition, the related experimental data and where it should be deposited, and the raw experimental data and where it is recommended to be deposited.
Experimental method
X-ray diraction
Nuclear Magnetic Resonance (NMR)
Cryo-electron microscopy (cryoEM)
Model solution method
Single model built into experimentally derived electron density map
Multiple models representing a range of conformers that satisfy experimental restraints
Single model built into experimentally derived electric potential map
PDB deposition requirements
Coordinate le (PDBx/mmCIF format)
Structure factor le (mtz or CIF format)
Coordinate le (PDB or PDBx/mmCIF format)
NMR restraints
NEF format) Coordinate le
(PDB or PDBx/mmCIF format)
Image for public display at EMDB EMDB map deposition (MRC or CCP4 format)
Experimental data (archive)
Structure factors (PDB)
NMR restraints (BMRB)
NMR chemical shifts (BMRB)
Electric potential map (EMDB)
Raw data (archive)
X-ray diraction image data (SBGrid, IRRMC)
NMR spectral parameters (BMRB)
NMR relaxation data (BMRB)
Electron microscopy images (EMPIAR)
The rst cryoEM structure in the PDB was released in 1991, however, it has taken a long time for the technique to establish itself as a routine method for high-resolution structure determination. The technique involves the imaging of biological molecules and complexes under cryogenic conditions, using a transmission electron micro­scope. The sample is ash-frozen in a thin layer with electrons passing through the sample to a detector, where images of each particle are captured. These images are then sorted and processed computationally to generate a 3D map of the sample, to allow tting of a molecular model.
The use of cryoEM has allowed the determination of larger and more heteroge­neous complexes than was previously possible with X-ray crystallography, however until recently, was not able to reach resolutions sucient to interpret atomic-level details. However, recent advances in the technique, including improved software and hardware including free-electron detectors, have led to cryoEM becoming the fastest-growing experimental method for studying macromolecular structure [21]. This is due to the determination of high-resolution structures in conditions that
5.2 Small Molecule Data in Protein Data Bank (PDB) Entries 145
https://t.me/medicina_free
better represent the native environment of the sample. Since 2016, submission of cryoEM structures to the PDB has also required submission of the experimental maps to the EMDB [6].
The third main technique used for solving structures in the PDB is NMR. This method involves the use of powerful magnetic elds to discern minor dierences in the resonance frequencies, also known as chemical shifts, of atoms in the macro­molecule. These chemical shifts are used to determine interactions among atoms in the molecule and generate atomic restraints, which can be used to build the structure of the macromolecules.
In solution-state NMR, the biological macromolecules can move freely in the sol­vent, allowing these experiments to capture the range of structural conformations adopted by these molecules. These structures are therefore deposited as multiple models in the PDB, highlighting an ensemble of potential structural conformations adopted by the macromolecules. An alternative technique is solid-state NMR [22], which uses similar underlying principles, however, is used to determine structures of molecular structures within solid or semisolid materials, including macromolecules within biological membranes.
Though the three experimental methods mentioned above account for the vast majority of PDB entries, there are also additional variations in diraction methods that can be used to solve structures in the PDB. Firstly, neutron crystallography [23] is a similar technique to X-ray crystallography, however, it relies upon neutron scat­tering to determine the macromolecular structure. The benet of this technique is that the neutrons interact with atomic nuclei, rather than electrons, which improves observation of hydrogen atoms in the structure. Neutron and X-ray crystallography are often used in conjunction to provide both high-resolution data and to determine positions of hydrogen atoms.
There are also methods that utilize X-raycrystallography to allow high-throughput determination of ligand binding in macromolecules. Fragment screening [24] exper­iments involve the “soaking” of various small molecule fragments into crystals of the sample. Determining how these dierent fragments bind within the pro­tein structure can help to improve the mechanistic understanding of a protein and the identication of suitable drug candidates. Pan-Dataset Density Anal­ysis (PanDDA) experiments [25] can be used to compare multiple fragment screening datasets to identify weak signals of bound small molecules within the noise of the data and can further improve the determination of small molecule binding sites.
5.2.2 Definition of Small Molecules in OneDep
In the PDB archive, ligands are dened as any molecule that is not part of a larger polymeric molecule, i.e. proteins, nucleic acids, or branched carbohydrates [26]. In the case of peptide-like ligands, if these contain at least two standard peptide bonds, then these are classed as polymeric entities. Small molecule ligands in the PDB can have a range of distinct functions, for example, as substrates, products, and inhibitors.
146 5 The Protein Data Bank (PDB) and Macromolecular Structure Data Supporting Computer
https://t.me/medicina_free
Each distinct ligand in the PDB archive is assigned its own unique 3-character ID code, to support easy identication across the full archive [26]. Each of these ligands includes a detailed chemical description, which is contained within the wwPDB Chemical Component Dictionary (CCD), and discussed in more detail in the next section.
5.3 Small Molecule Dictionaries
5.3.1 wwPDB Chemical Component Dictionary (CCD)
The wwPDB maintains a reference dictionary containing chemical descriptions of each unique chemical component present in the whole PDB archive. These components are the building blocks for PDB entries and include amino acids, nucleotides, metal ions, and other nonpeptide small molecules. These chemical component descriptions are stored in the wwPDB’s CCD (https://www.wwpdb.org/ data/ccd) [26].
Each component is given a unique identier and has a CCD denition that describes the molecule. As of July 2022, there are around 37,000 unique denitions in the wwPDB CCD. The unique identier is currently limited to a maximum of three characters. Most denitions contain a three-character unique identier, with one- or two-character identiers mostly used for identication of DNA and RNA bases or for components containing individual elements. Some examples of chemical components include manganese (Mn), adenosine-triphosphate (ATP), and glutamate (GLU). Previously, these identiers were assigned with some meaning to their chemical denition, however, due to the high number of chemical components in the dictionary, these CCD identiers are now randomly assigned for any new ligands.
If the CCD denition is dening a standard amino acid or nucleotide, the IUPAC protein one-letter code [27] is also provided, for example, Glutamate has the unique identier GLU and the one-letter code E. Small modications of standard amino acids or nucleotides, generally, where the modication contains fewer than 10 atoms, have their own CCD denitions, for example, phospho-serine has the identier SEP. These modied amino acids or nucleotides have the same one-letter code as their standard amino acid parent, therefore SEP, which is a modied serine amino acid, also contains the one-letter code S.
Peptide-based chromophores provide a more complex example. In proteins, chromophores typically comprise three amino acids condensed into one molecule. In the CCD component, the one-letter code lists all amino acids that make up the chromophore. For example, the CCD denition for the chromophore CR2, which comprises the amino acids GLY-TYR-GLY and has the corresponding one-letter code GYG.
In addition to providing a unique identier, each CCD denition describes the chemistry of the molecule. This includes the name of each atom in the molecule, the order of bonds between atoms, stereochemistry of chiral atoms, and any charges
5.3 Small Molecule Dictionaries 147
https://t.me/medicina_free
on individual atoms or the whole molecule. There are two sets of 3D coordinates included in the CCD denition: the model coordinates, extracted from the PDB entry from which the CCD was initially generated, and idealized coordinates (from the lowest energy conformation), generated using the expected geometry of the ligand from the CCD denition.
Additional metadata are also included in the CCD denition, including the name of the molecule and any synonyms. If a common name is known for the compound, then it is provided either as the name or as one of the synonyms, otherwise, system­atic names are provided. The systematic name is the default name generated by the OpenEye [28] software at the time of creation of the compound. Chemical descrip­tors in InChI and SMILES formats are provided and are automatically generated by OpenEye during creation of the compound.
Every molecule in the PDB must have a corresponding CCD denition. During deposition and biocuration of PDB entries, each compound in the deposited le is compared to existing CCD denitions using a graph match algorithm [29]. If rene­ment restraints are provided in the uploaded coordinate le, then this information is used for comparing the molecule to entries in the CCD, otherwise, the coordinates are used for searching. If a molecule fails to nd a match in the CCD, then deposi­tors are asked for further details of the molecule so that a new CCD denition can be made. A new CCD denition is then dened during biocuration and added to the CCD. The name of the molecule is either dened by the depositor, if a common name is available, or a systematic name is generated based on the chemistry of the molecule.
5.3.2 The Peptide Reference Dictionary
The PDB archive also contains examples of complex ligands, which comprise combinations of amino acid and amino acid-like components. These peptide-like molecules commonly possess important biological functions, such as antibiotic or inhibitory activity, however, a description of their global chemistry does not t within the classic PDB denitions of polymers and ligands. Therefore, to support the curation of this important class of molecules, their chemical and biological descrip­tions are described in a separate dictionary. The “Peptide Reference Dictionary” (PRD) resource [30] provides a standardized framework for these denitions so that each PRD entity can be described in a way that outlines both its subcomponent composition and global characteristics. These PRD denitions are stored in the wwPDB Biologically Interesting Molecule Reference Dictionary (BIRD) (https:// www.wwpdb.org/data/bird).
The process of dening a new PRD molecule, or matching connected CCDs to an existing PRD, occurs during the biocuration process. Any atomistic information, such as reference coordinates, bonding, and chemical descriptors are then listed in PRD-chemical denition les, whereas subcomponent sequence, CCD-linking, and any related biological details, such as biological function, are stored in PRD-molecular denition les. By using the PRD dictionary to standardize biocura­tion, it allows for the precise interrogation, comparison, and retrieval of global and
148 5 The Protein Data Bank (PDB) and Macromolecular Structure Data Supporting Computer
https://t.me/medicina_free
Figure 5.2 An example of PRD is the antibiotic Vancomycin (PRD_000204). Vancomycin consists of seven nonstandard amino acids (ball and stick, distinguished by color) and two sugar components (cyan and grey in a 3D representation of the SNFG notation). The PRD polymer linking section of the PRD molecular definition file is shown on the right-hand side.
subcomponent-specic data, which would otherwise not be possible. Links to PRD denition les can be found on relevant PDB entry pages and references to key PRD identiers (ID, name, type, and function) are stored in PDB entry PDBx/mmCIF les under the pdbx_molecule_features category (Figure 5.2).
5.4 Additional Ligand Annotations in the PDB Archive
In addition to the standard chemical and biomolecular denitions already described, where possible, the PDB biocuration pipeline generates additional ligand-related information to best describe ligand complexity and intramolecular connectivity and to do so with a consistent and objective methodology [31]. Many software packages rely upon these data for purposes such as molecular renement and accurate molec­ular visualization.
5.4.1 Linkage Information
The most common ligand-associated annotations found in PDB entry les refer to intermolecular bonding. These refer to all nonstandard linkages between and among polymer components and CCD ligands. This information was historically found in PDB-format LINK, SSBOND, and CONECT records, however, the current PDBx/mmCIF archive format provides additional data, including leaving group information, atomic descriptions, and author numbering for up to three chemical component partners per interaction [9]. At present, to ensure standardization across the whole PDB archive, all link records and bonding distances are calculated and validated during the biocuration process [31]. The PDBx/mmCIF dictionary cur­rently supports covalent, hydrogen, covalent–metal, and disulde-bond interaction “types” [32] (Figure 5.3).
5.4 Additional Ligand Annotations in the PDB Archive 149
https://t.me/medicina_free
(a) (b)
Figure 5.3 Two examples of PDB linkage information and default visualization styles in the Mol* viewer: (LEFT: 5AUS) covalent interaction between heme c ligand and cysteine side chain (CYS10-HEC), and metal coordination between the iron in heme c and a histidine side chain (HIS14-FE); (RIGHT: 4TPL) a covalent carbohydrate linkage between two N-acetyl-D-glucosamine (NAG) ligands (NAG2-NAG1) and covalent N-Glycosylation between a NAG and an asparagine side chain (NAG1-ASN207).
5.4.2 Carbohydrates
The most recent development concerning PDB ligand denitions relates to the annotation of polymeric carbohydrates [33]. These new “branched poly­mer” descriptions (entity_type item “branched”) are composed of multiple and sequential CCD monosaccharides. The description of polymeric carbohydrates is outlined within the PDBx/mmCIF model le for a given PDB entry, under the new pdbx categories entity_branch, entity_branch_descriptor, entity_branch_link, entity_branch_list, and branch_scheme (Figure 5.4a). Whilst chemically distinct
(a) (b)
Figure 5.4 (a) Example categories from a PDBx/mmCIF PDB model file describing a branched polymer carbohydrate. The example highlights how entity 2 (raffinose (RAF), PRD_900002) is comprised of the subcomponents FRU, GLC, and GLA (fructose, glucose, and galactose) and lists the corresponding atomic linkages. (b) Examples of wwPDB carbohydrate representations by 2D and 3D symbol nomenclature for glycans (SNFG) (top and middle, respectively). Source: (b) These images were taken from the PDBe entry pages for PDB accession 5ofx; https://pdbe.org/5ofx.
6
αα
2
β
150 5 The Protein Data Bank (PDB) and Macromolecular Structure Data Supporting Computer
https://t.me/medicina_free
from the peptide-like PRD molecules, carbohydrates of biological signicance may also be described by the PRD dictionary les where they hold functional classica­tions such as “nutrients” or “inhibitors.” The annotation of carbohydrate-branched polymers occurs within the automatic processing of the biocuration pipeline. Dierent carbohydrate visual representations may now also be found on wwPDB partner websites, using the community standard SNFG notation [34, 35], both as 2D images on each entry page, and 3D-SNFG [36] units within the Mol* viewer (Figure 5.4b).
5.5 Validation of Ligands in the Worldwide Protein Data Bank (wwPDB)
Currently, PDB contains over 34,000 unique small molecules bound to over 133,600 protein structures. Numerous groups have highlighted the need for standardized validation in structures in the PDB archive [37–44]. Several errors have even led to the retraction of the respective publication(s) and subsequent obsoletion of the PDB entry [43]. A Validation Task Force (VTF) was set up to centralize and standardize validation across the wwPDB archive, which included metrics to judge the reliability of the ligand-bound macromolecular structures as they play a pivotal role in many computational methods like ligand docking and structure-based drug design. This resulted in development of the wwPDB validation reports (VR), rst introduced in
2012. wwPDB validation pipeline (VP) generates detailed wwPDB VR comprising
an assessment of structure quality using widely accepted standards and criteria recommended by various method-specic VTFs [45–47]. To enhance ligand vali­dation, in 2015, wwPDB, CCDC (Cambridge Crystallographic Data Center), and D3R (Drug Design Data Resource) co-organized the ligand validation workshop (LVW) [48] whose recommendations were recently implemented in VP [14]. VP is an integral part of the unied OneDep system for structure deposition [13], validation [49], and biocuration [31]. Before depositing the structure in the wwPDB archive, the depositor can generate a wwPDB VR using the standalone validation server (validate.wwpdb.org) or API (wwpdb.org/validation/onedep-validation­web-service-interface) to nd and x any issues highlighted in the VR. At the time of deposition, a “preliminary” VR is issued to the depositor, followed by the “ocial” VR after the completion of biocuration process. Once the structure is released, the VR is made public and is provided as both human-readable PDF les and machine-readable mmCIF and XML les. In the next section, we will further discuss various criteria and software used for validation of ligands in the VR.
5.5.1 Various Criteria and Software Used for Validating Ligand
in Validation Reports
wwPDB validation criteria for ligands can be broadly divided into two categories (I) geometric validation of the atomic coordinates, without considering the associated
5.5 Validation of Ligands in the Worldwide Protein Data Bank (wwPDB) 151
https://t.me/medicina_free
Table 5.2 Summary of the software used for validating ligands in wwPDB validation reports.
Property Detailed steps involved Software used Software URL
Atomic coordinates of all bound ligands
Geometric quality Computes following
2D diagrams of geometric quality
3D graphical depiction of the model t to data
Extracts ligand of interest (LOI)
reorder and rename atoms in ligands according to the Chemical Component Dictionary (CCD)
for all ligands:
● Bond-length,
bond-angle, torsion angle and ring outliers
Chirality and planarity outliers
Highlights geometric analysis provided by CCDC Mogul
Each model t to experimental electron density is shown from dierent orientations to approximate a 3D view
Maxit (Macromolecular Exchange and Input Tool) [50]
Mogul [51]
Validation-pack [50]
buster-report [43] https://www
buster-report [43] As above
https://sw-tools .rcsb.org/apps/ MAXIT
https://ccdc.cam.ac .uk/solutions/csd­core/components/ mogul
https://mmcif .wwpdb.org/docs/ software-resources .html
.globalphasing .com/buster/ manual/ autobuster/ manual/ autoBUSTER9.html
experimental data, and (II) validation of t between the atomic coordinates of the ligand and the associated experimental data. The summary of various software used in VR is shown in Table 5.2.
5.5.2 Identification of Ligand of Interest (LOI)
Since 2017, the depositors are required to identify the ligand(s) in the structure that is either the focus of their study or is considered biologically important. This LOI infor­mation is recorded during the deposition of structure in the OneDep system [13]. All its details can be found in the “Entry composition” section of the VR. In the VP, Maxit [50] extracts the atomic coordinates of ligand(s), reorder and rename its atoms according to the CCD after which geometric quality of these is assessed further.
152 5 The Protein Data Bank (PDB) and Macromolecular Structure Data Supporting Computer
https://t.me/medicina_free
5.5.3 Geometric and Conformational Validation
Following the recommendation of wwPDB X-ray VTF, ligand geometry is validated against the Cambridge Structural Database (CSD) of high-quality small molecule organic structure [45]. Firstly, CSD structures with similar bond order are identied and then a distribution for the observed values for each bond length, bond angle, torsion angle, or ring in the ligand is built using the program Mogul [51]. It then computes the Z-score, which quanties deviation from the observed values. Bond lengths and bond angles with absolute Z-score >2.0 are agged as “outliers” in the VR. The normalized Z-scores (RMSZ) of bond lengths/angles are calculated to facil­itate the overall assessment of the ligand. RMSZ scores are expected to lie between 0 and 1. For low-resolution structures, geometry should be tightly restrained and small values are expected. For very high-resolution structures, values approaching 1 may be attained. Values greater than 1 generally indicate over-tting of the data. For acyclic torsion angles, Mogul computes the local density measure [52], which is the ratio of incidences in the CSD within 10∘of the torsion angle in question, to the num­ber of total incidences of the torsion angles in the CSD. The torsion angle is agged as an outlier if the local density measure is <5%. For isolated rings, Mogul compares the given ring with comparable rings in small molecule structures in the CSD and calculates an RMSD value based on corresponding constituent torsion angles for each comparable ring. If both the mean and minimum of these RMSDs are >60∘, the ring is agged as an outlier. Chirality outliers are calculated by Validation-pack [14] and are assessed based on chiral volume. If the sign of the computed volume is incorrect, the handedness is wrong. If the absolute volume is less than 0.7Å chiral center is as a planar moiety, which is highly likely to be erroneous. All these geometric outliers are listed in the “Ligand geometry” section of the VR.
Apart from merely identifying the geometric outliers, it is also important to know
the extent of distortion from the Mogul expectation. After the latest update of VP [14], this information is now clear in 2D colored images generated by Buster-report [53]. This image depicts Mogul quality analysis of bond lengths, bond angles, tor­sion angles, and ring geometry [43] and is added to the “Ligand geometry” section of VR. Color scheme used in these images is coded according to validation results with green indicating commonly observed values, magenta indicating unusual val­ues, and grey indicating that there was insucient data to derive a validation score (see Figure 5.5b). Unusual values include model quality and electron density t. For model quality, individual bond lengths or angles with absolute Z-score >2, the torsion angle with <5% of local density measure, or RMSD >60∘are considered unusual and colored in magenta. These images can be helpful for depositors to iden­tify any putative incorrect restraint values used during renement before depositing the structure.
3
,the
5.5.4 Ligand Fit to Experimental Electron Density Validation
Apart from validating the geometric quality of the ligand, it is also crucial to inspect local agreement of a structure model with the observed electron density and