Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5419_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
15.09.2026
Размер:
11 Мб
Скачать
☆
https://t.me/medicina_free
9
https://t.me/medicina_free
Mining for Bioactive Molecules in Open Databases
Guillem Macip, Júlia Mestres-Truyol, Pol Garcia-Segura, Bryan Saldivar-Espinoza, Santiago Garcia-Vallvé, and Gerard Pujadas
Departament de Bioquímica i Biotecnologia, Carrer Marcel⋅lí Domingo 1, Universitat Rovira i Virgili, Research group in Cheminformatics & Nutrition, 43007 Tarragona, Catalonia, Spain
9.1 Introduction
An important goal of drug-discovery researchers is to determine which molecules in chemical compound databases have a specic bioactivity. This goal can be achieved either by experimentally testing compound libraries to nd molecules that show the desired bioactivity (a process known as high-throughput screening; HTS) [1, 2] or by computationally predicting the bioactivity of interest in les containing the struc­ture of hundreds of thousands, or even millions, of chemical compounds encoded in the form of electronic archives (a process known as virtual screening; VS) [3, 4]. In VS, the researcher does not need to physically have the chemical compounds since their description in electronic format is sucient. The cost of VS is therefore much lower than the cost of HTS, where the bioactivity of a large number of compounds is measured experimentally.
Like all screening techniques, VS is sometimes compared to looking for a needle in a haystack. However, using VS to nd molecules that exhibit the desired activity is much easier than nding the needle since VS is a targeted search (unlike search­ing for a needle, where there is no privileged direction and all search directions are equally possible). This targeted search can be guided by the characteristics of existing drugs for the same target or by the structure of the macromolecular target itself.
Figure 9.1 shows an example of a typical VS workow that is frequently used in our laboratory [5–7] (other VS pipelines may also be successfully employed [8, 9]). This VS workow consists of several sequential lters (i.e. ADMET/PAINS, protein–ligand docking, pharmacophore, and shape/electrostatic) where the output molecules of one lter are the input molecules of the next lter, and so on. The inverted pyramid shape in Figure 9.1 also shows that, as the VS progresses, the number of available molecules decreases since molecules that do not pass a lter are discarded and not evaluated by the remaining lters. If the VS is designed correctly, it is expected that, as the molecules pass through the various lters, the resulting
271
Open Access Databases and Datasets for Drug Discovery, First Edition. Edited by Antoine Daina, Michael Przewosny, and Vincent Zoete. © 2024 WILEY-VCH GmbH. Published 2024 by WILEY-VCH GmbH.
272 9 Mining for Bioactive Molecules in Open Databases
https://t.me/medicina_free
Figure 9.1 Overview of a typical virtual
INITIAL DATABASE
ADMET/PAINS FILTER
PROTEIN–LIGAND DOCKING
PHARMACOPHORE FILTER
SHAPE/ELECTROSTATIC
FILTER
VS HITS
screening (VS) workflow previously used in our laboratory. The shape of the funnel is related to the number of molecules evaluated at each stage of the VS. Since these are elimination stages in which compounds that do not pass a filter are discarded and therefore not evaluated by the subsequent filters, the number of available molecules becomes smaller and smaller.
molecular sample will be enriched in compounds with the desired bioactivity. In other words, the probability that a randomly chosen molecule is an active one must be lower in the starting database than at the end of the VS.
In this chapter, we describe the main methods used during a VS, determine which open databases are useful during this process, and explain how to validate the accu­racy of the bioactivity predictions made by the VS. All the resources mentioned in this chapter are listed in Table 9.1, which also contains information on how to access them.
9.2 Main Tools for Virtual Screening
VS have traditionally been classied as either Ligand-Based VS (LBVS) or Structure-Based VS (SBVS) depending on whether the 3D structure of the target is used (SBVS) or not used (LBVS) during the VS workow [3]. However, ligand­and structure-based methods can also appear combined in the same VS pipeline [10–12]. Interestingly, if the experimental 3D structure of the target is unavailable but the structure from a homologous protein is known, then a homology model of the target can be built and SBVS methods can be used to nd new active compounds in chemical databases [5, 13]. With the AlphaFold program, [14–16] it is currently possible to predict the structure of almost any protein (even if no experimental 3D structure of homologous proteins of the target of interest is known) with a high degree of accuracy. This means that almost all the technologies used in the VS (SBVS or LBVS) are applicable to any target (though the application of AlphaFold results in the eld of drug discovery is still at a very early stage and rather uncertain [17–19]). The following sections describe the lters in Figure 9.1.
9.2.1 ADMET and PAINS Filtering
To be eective as a drug, a bioactive molecule must be able not only to reach its target in the organism at a sucient concentration but also to remain there in bioac­tive form long enough for the expected biological events to occur [20]. ADME, the
9.2 Main Tools for Virtual Screening 273
https://t.me/medicina_free
Table 9.1 Resources mentioned in this chapter.
Main tools for virtual screening
ADMET and PAINS ltering
SwissADME (http://www.swissadme.ch/) FAF-Drugs4 (https://fafdrugs4.rpbs.univ-paris-diderot.fr/)
Protein–ligand docking Rigid protein–ligand docking
AutoDock Vina (https://vina.scripps.edu/) Induced t protein–ligand docking SLIDE (https://github.com/psa-lab/SLIDE) PELE Web server (https://pele.bsc.es/pele.wt) Ensemble protein–ligand docking DINC-COVID (http://dinc-covid.kavrakilab.org/) Edock-ML (http://edock-ml.umsl.edu/)
Pharmacophore search
Structure-based pharmacophore
Pharmit (http://pharmit.csb.pitt.edu) ZINCPharmer (http://zincpharmer.csb.pitt.edu/)
Ligand-based pharmacophore
PharmaGist (http://bioinfo3d.cs.tau.ac.il/PharmaGist/)
Shape/electrostatic similarity
ESP-Sim (https://github.com/hesther/espsim) ElectroShape (https://ub.cbm.uam.es/chemogenomics/)
Protein-structure databases
PDB (https://www.rcsb.org/) PDB-REDO Databank (https://pdb-redo.eu/) SWISS-MODEL Repository (https://swissmodel.expasy.org/repository) AlphaFold Protein Structure Database (https://alphafold.ebi.ac.uk/)
Validating binding site and ligand coordinates in three-dimensional protein complexes
VHELIBS (https://github.com/URVquimioinformatica-COS/VHELIBS)
Databases for searching new drugs
COCONUT (https://coconut.naturalproducts.net/) GDBs (https://gdb.unibe.ch/downloads/) ZINC20 (https://zinc20.docking.org/)
Databases of active molecules
The Binding Database (https://www.bindingdb.org/) ChEMBL Database (https://www.ebi.ac.uk/chembl/) PubChem (https://pubchem.ncbi.nlm.nih.gov/)
(continued)
274 9 Mining for Bioactive Molecules in Open Databases
https://t.me/medicina_free
Table 9.1 (Continued)
Databases of decoy molecules
DUD-E (http://dude.docking.org/) DEKOIS (http://www.pharmchem.uni-tuebingen.de/dekois/) VDS (http://compbio.cs.toronto.edu/VDS/)
Tools for building custom-based decoy sets
DUD-E tool (http://dude.docking.org/generate) DecoyFinder (https://github.com/URVquimioinformatica-COS/DecoyFinder)
Format descriptions
SDF (https://depth-rst.com/articles/2020/07/13/the-sdle-format/) SMILES (https://www.daylight.com/dayhtml/doc/theory/theory.smiles.html) PDB le format (http://www.wwpdb.org/documentation/le-format-content/format33/
v3.3.html)
acronym for absorption, distribution, metabolism, and excretion [21], describes how pharmacokinetic behavior is simplied into discrete parameters to be addressed in the drug discovery/preclinical phases. All four criteria inuence the levels and kinet­ics of drug exposure to tissues and thus the performance and pharmacological activ­ity of the compound as a drug. When, in addition to ADME properties, the potential or actual toxicity of the compound is taken into account, the acronym ADME-Tox, ADME/T, or ADMET is used instead [22]. PAINS, the acronym for Pan Assay Inter­ference Compounds, on the other hand, refers to compounds that interfere with assay read-out or that can also react nonspecically with numerous targets in such a way that they often generate false positives [23].
Filters to remove from the initial database those molecules that are predicted to be PAINS or that have unfavorable ADMET properties are not part of the VS work­ow itself but part of the initial pretreatment phase of the molecules on which the screening of bioactive compounds will be performed (being a PAINS or having unfa­vorable ADMET properties is an intrinsic characteristic of the molecule itself and independent of the target against which the VS will be performed). Open access servers such as SwissADME [20, 24] and FAFDrugs4 [22, 25] can predict which input molecules (in SMILES format for SwissADME and SDF/SMILES format for FAFDrugs4 [26, 27]) may be PAINS or have bad ADMET properties.
9.2.2 Protein–Ligand Docking
Protein–ligand docking predicts the coordinates of the complex between a protein and a drug from their individual structures [28]. During this process, the protein structure is usually considered rigid (or small conformational changes are allowed in a very localized part of its structure), while the small molecule is considered able to change its initial conformation to adapt to the protein-binding site. However, since some targets show signicant exibility in their binding site, considering them as
9.2 Main Tools for Virtual Screening 275
https://t.me/medicina_free
rigid during protein–ligand docking is a too strict approach. To overcome this limi­tation, several strategies have been developed, including induced t docking [29–33] and ensemble docking [34, 35]. The most important dierence between these two approaches is that, while induced docking attempts to directly simulate the coupled motion of the receptor and ligand, ensemble docking uses a set of dierent target structures (either experimental or from molecular dynamics simulations) to perform calculations. AutoDock Vina [36, 37] is often used for protein–ligand rigid docking whereas SLIDE [32] and Edock-ML [35] are examples of induced t and ensemble protein–ligand docking, respectively.
9.2.3 Pharmacophore Search
A pharmacophore is a 3D abstract description of molecular features (i.e. pharma­cophoric sites) that are necessary for molecular recognition and biological activity of a ligand by its target. A pharmacophore may also contain exclusion volumes that indicate the positions of the binding site that the ligand cannot ll because they are already occupied by the atoms of the target itself [3].
The simplest way to obtain a pharmacophore is from the three-dimensional (3D) structure of a complex between the target of interest and a drug that has the desired activity on this target (referred to as a structure-based pharmacophore).
When several drugs are known to be bioactive with respect to the same target but the 3D structure of the target is not known, a ligand-based pharmacophore can be obtained (assuming that all these ligands also share the target binding site and the binding mode). Here, it is assumed that the combination of conformers (one per ligand) with the maximum number of common pharmacophoric sites corresponds to the way all these ligands bind to the 3D structure of the target (the common sites are also assumed to represent the intermolecular interactions responsible for their recognition and bioactivity with respect to the target). If inactive compounds are also known, they can be used to add exclusion volumes to the ligand-based pharmacophore.
Once a pharmacophore has been obtained, it can be used to search for small molecules that occupy the maximum number of pharmacophore sites and at the same time do not occupy exclusion volumes. Molecules that t the pharmacophore will be candidates for having the same bioactivity for the target as the molecule (or molecules) from which the pharmacophore was obtained.
PharmaGist [38, 39] is a free resource for constructing ligand-based pharma­cophores, while Pharmit (see Figure 9.2) [41–43] and ZINCPharmer [44, 45] are free resources for building structure-based pharmacophores. Moreover, all three resources can be used to screen small molecule databases.
Pharmacophores can also be used to examine the poses resulting from a previ­ous protein–ligand docking step. In this way, it is possible to assess whether the ligands, once oriented within the target binding site, are able to establish the inter­molecular interactions needed to show high activity toward the drug target. In fact, this is another way to exploit the main strength of protein–ligand docking algo­rithms (i.e. their ability to nd the bioactive pose) and obviate their main weakness
276 9 Mining for Bioactive Molecules in Open Databases
https://t.me/medicina_free
Figure 9.2 The Pharmit web interface for a pharmacophore search. A pharmacophore query for SARS-CoV-2 M-pro in complex with a perampanel derivative (PDB 7L10 [40]) is shown on the Pharmit web interface [41].
(i.e. their inability to correctly calculate the anity of each pose for the target). To do this, one would only need to make sure that, for each ligand, the output of the docking program provides a sucient number of poses to ensure that the correct one is among them (regardless of whether the scoring function places it at the top of the predicted anity ranking for the target). This strategy has been successfully used to identify bioactive compounds for dierent targets [10, 11] and is similar to using constraints during protein–ligand docking (though, in our opinion, using pharmacophores allows greater exibility in the VS strategy than using these con­straints because, among other things, pharmacophore site types are more diverse than protein–ligand docking constraints).
9.2.4 Shape/Electrostatic Similarity
Two molecules that, upon binding to a given target, share a 3D shape and an electro­static charge distribution on their surface (even though they have a dierent chem­ical structure) may have similar bioactivity for that target [46, 47]. It is therefore possible to compare the shape and electrostatic surface area of the poses that have passed through the previous lters (these would be the pharmacophore-compliant docking poses; see Figure 9.1) with the experimental poses of drugs that are cocrys­tallized with the target of interest. However, care must rst be taken to superimpose the structure of all these complexes onto the structure of the target used to obtain the docking poses and the pharmacophore. Only in this way can it be guaranteed that the comparison is made between molecules with the same relative orientation in the target binding site. Moreover, if more than one complex between a highly active drug and the target of interest is available, each docking pose must be compared with each experimental pose, and the highest similarity must be chosen. Then, if more than one docking pose is available for the same ligand, only the one that shows the greatest similarity to one of the experimental poses to which it is compared will be kept (while the rest will be discarded) [6, 7, 11]. Finally, poses with a similarity
9.2 Main Tools for Virtual Screening 277
https://t.me/medicina_free
ZINCO3851930H11 at 3BYZ
2D chemical structure
Shape similarity
(ET_shape= 0.560)
Electrostatic surface
similarity
(ET_pb = 0.640)
Figure 9.3 Shape and electrostatic similarity. Shape and electrostatic comparison between the experimental pose of the H11 ligand in the PDB entry 3BYZ [48] and the docked pose of the ZINC03851930 ligand in the 3BYZ binding site.
below a certain threshold will be discarded. This threshold must be selective enough to reduce, as much as possible, the number of false positives during VS validation (even if this means increasing the number of false negatives) [11].
Figure 9.3 shows the comparison between two poses, one of which is experimen­tal (corresponding to the H11 ligand in the 3BYZ structure [48]), and the other corresponds to the docking pose of the ZINC03851930 ligand in the same structure. The shape similarity between the two poses is 0.560 (where 1.0 corresponds to a perfect overlap, i.e. the same shape), and the Poisson–Boltzmann electrostatic similarity is 0.640 (where 1.0 corresponds to an identical electrostatic potential overlap).
ESP-Sim [49] and ElectroShape [47] are examples of free resources for shape/ electrostatic similarity searches.
9.2.5 Protein-Structure Databases
Many tools used during a VS require the 3D structure of the target to which the drug is to bind. These 3D structures may have been obtained experimentally (using tech­niques such as X-ray diraction, nuclear magnetic resonance [NMR], or electron microscopy) or may result from predictions obtained directly from their sequence (using techniques such as homology modeling or protein threading). Several public databases exist from which these 3D structures can be obtained. The most important
278 9 Mining for Bioactive Molecules in Open Databases
https://t.me/medicina_free
ones are the Protein Data Bank (PDB) [50] and the PDB-REDO Databank [51] for experimentally obtained structures and the SWISS-MODEL Repository [52] and the AlphaFoldProtein Structure Database (AlphaFold DB) [15] for computationally pre­dicted structures.
9.2.6 The Protein Data Bank
The PDB, which was created in 1971 at Brookhaven National Laboratory (Upton, NY), originally contained seven structures [53]. Its purpose was to act as an interna­tional repository in which to deposit the coordinates of proteins determined by X-ray diraction and make these data available on request. Since then, it has also included protein structures determined by other techniques (NMR and electron microscopy) and the 3D structures of other biomolecules (nucleic acids, oligosaccharides, and their respective complexes with proteins). A PDB entry, which contains all the information relating to a particular 3D structure deposited in the PDB, is designated by a 4-character alphanumeric identier called a PDB identier or PDB ID, which always begins with a number from 1 to 9 (e.g. 6LU7). There are currently over 193,000 entries in the PDB, of which 168,000 correspond to proteins, 10,600 correspond to protein–nucleic acid complexes, 10,200 to protein–oligosaccharides complexes, 3900 to nucleic acids, and 22 to oligosaccharides. Of these 193,000 entries, 167,000 were determined by X-ray diraction, 13,700 were determined by NMR, and 12,000 were determined by electron microscopy [54]. The coordinates of the PDB structures are contained in text-only les (one le for each PDB entry) in a PDB-specic format [55] that is soon to be replaced by the mmCIF format [56]. All PDB les can be downloaded without restriction from the PDB or PDBe websites [57, 58].
9.2.7 The PDB-REDO Databank
The PDB-REDO databank [59] contains optimized versions of existing PDB entries that simultaneously satisfy two criteria: (i) the entry has been obtained by X-ray diraction; and (ii) the author deposited the corresponding structure factor les along with the coordinates. All PDB entries that satisfy both criteria are treated with a consistent protocol that reduces eects due to dierences in age, software, and depositor, thus making PDB-REDO a great dataset for large-scale structure analysis studies [51, 60]. For example, an evaluation of the goodness of t of the coordinates of 39,820 protein/ligand complexes to the electron density map using the VHELIBS program showed that both binding-site and ligand coordinates were much more reliable if they were obtained from PDB-REDO than if they were obtained from PDB [61]. Specically, while VHELIBS rated 14,304 binding sites and 25,350 lig­ands obtained from PDB-REDO as “Good” (the other categories were “Dubious” and “Bad”), these values decreased by 50% when the data were obtained from PDB [62] and EDS [63] (7671 and 12,419, respectively). It is therefore recommended that, whenever possible, priority should be given to working with structures obtained from PDB-REDO rather than those obtained directly from PDB.
9.2 Main Tools for Virtual Screening 279
https://t.me/medicina_free
9.2.8 The SWISS-MODEL Repository
The purpose of the SWISS-MODEL Repository is to provide access to an up-to-date collection of annotated 3D protein models generated for all UniProtKB sequences [64] for which no experimental structures are available [52]. Since the structures in this repository were obtained by homology modeling (and these types of models can only be obtained by the automated SWISS-MODEL pipeline for those proteins that have a minimum of 30% sequence identity with at least one protein of known 3D structure that acts as a template during model construction), this repository contains only structural models for a fraction of the UniProtKB proteins. Interestingly, in cases where templates corresponding to the same sequence seg­ment exhibit signicant conformational dierences, several models are generated to reect structural diversity.
To inform users about the expected local accuracy, the quality of the various models is evaluated and annotated using the QMEAN tool (only high-quality models – i.e. those with an overall QMEANDisCo score above 0.5 – are imported into the SWISS-MODEL repository) [65]. If the homology-modeled structure for a sequence is unavailable, the user can build it interactively through the SWISS-MODEL workspace [66, 67]. Currently, the repository contains 2,258,794 homology models that can be freely downloaded [68].
9.2.9 The AlphaFold Protein Structure Database
AlphaFold DB was released in July 2021 [15, 69]. This database contains the 3D structures predicted by the AlphaFold program (CASP14 winner) for proteins of experimentally unknown 3D structure [16]. Since, unlike the SWISS-Model, AlphaFold does not need homologous protein structures to make its predic­tions [14], AlphaFold DB contains structural models of proteins that are not in the SWISS-Model Repository. The rst release of this database covered the human proteome and the proteomes of 20 other key organisms (e.g. Arabidopsis thaliana, Escherichia coli, and Oryza sativa), while the second release added the majority of manually curated UniProt entries (e.g. Swiss-Prot) [70]. AlphaFold DB currently contains over 200 million entries that include the human proteome and the proteomes of 47 other key organisms that are important in research and global health [69].
Regarding the quality of the predicted 3D structures, AlphaFold produces a per-residue estimate of their condence on a scale from 0 to 100. This condence estimate, called pLDDT, corresponds to the score predicted by the model on the lDDT-Cα metric. This value is stored in the B-factor elds of the AlphaFold database les in PDB format, but, unlike the B-factor, the higher the value of pLDDT, the better the quality of the prediction. Regions with pLDDT > 90 are thus expected to be modeled with high accuracy, while those with pLDDT between 70 and 90 are expected to be modeled well, and those with pLDDT between 50 and 70 are of low condence and should be treated with caution.
In theory, it would be advisable to use only AlphaFold DB structures during drug discovery or development if their binding sites have a pLDDT value between