Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5660_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
15.09.2026
Размер:
11 Мб
Скачать
☆
60 2 PubChem: A Large-Scale Public Chemical Database for Drug Discovery
https://t.me/medicina_free
searches. Using the PubChem Classication Browser, it is possible to quickly retrieve records that are annotated with a classication or ontological term. In addition, the Identier Exchange Service allows users to convert identiers for a given set of chemical structures into dierent types of identiers for identical or similar chemical structures.
PubChem records are highly interlinked with each other. The Literature Knowl­edge Panels, which are embedded in the Summary page of a PubChem record, help users explore the relationship between entities contained in PubChem (e.g. chem­icals, genes, proteins, and diseases). In addition, for a given compound, PubChem provides a precomputed list of 2D and 3D neighbors. These neighbors help users to predict the molecular properties and biological activities of a compound that does not have much information.
PubChem data are accessible through multiple programmatic interfaces, includ­ing PUG-REST and PUG-View. Bulk data download via FTP is also supported. Especially, the PubChem FTP site hosts PubChemRDF data, which helps users to integrate PubChem data with in-house data or data from other resources across scientic domains.
PubChem contains more than 110 million compounds, and the majority of them are drug-like, satisfying all criteria of Ro5 or violating only one criterion of Ro5. Because data from many public chemical databases (Table 2.1) are integrated within PubChem, a substantial number of compounds in these resources are also contained in PubChem. In addition, 3.6 million compounds have been tested in at least one bioassay in PubChem, and about half of them (1.5 million compounds) have been declared active in at least one bioassay. While the majority of bioactivity data in PubChem are generated from HTS, a substantial amount of data are extracted from literature through text mining and manual curation. PubChem’s bioactivity data can be used to build predictive models for the bioactivities of chemicals.
Acknowledgments
This work was supported by the National Center for Biotechnology Information of the National Library of Medicine (NLM), National Institutes of Health. The authors thank Yolanda L. Jones, National Institutes of Health Library, for editing assistance.
References
1 Eltyeb, S. and Salim, N. (2014). Chemical named entities recognition: a review
on approaches and applications. Journal of Cheminformatics 6: 17.
2 Krallinger, M., Leitner, F., Rabal, O. et al. (2015). CHEMDNER: the drugs and
chemical names extraction challenge. Journal of Cheminformatics 7: S1.
3 Leaman, R., Wei, C.H., and Lu, Z.Y. (2015). tmChem: a high performance
approach for chemical named entity recognition and normalization. Journal of Cheminformatics 7: S3.
References 61
https://t.me/medicina_free
4 Luo, L., Yang, Z.H., Yang, P. et al. (2018). An attention-based BiLSTM-CRF
approach to document-level chemical named entity recognition. Bioinformatics 34 (8): 1381–1388.
5 Krallinger, M., Rabal, O., Lourenco, A. et al. (2017). Information retrieval and
text mining technologies for chemistry. Chemical Reviews 117 (12): 7673–7761.
6 Rajan, K., Brinkhaus, H.O., Zielesny, A. et al. (2020). A review of optical chemi-
cal structure recognition tools. Journal of Cheminformatics 12: 60.
7 Filippov, I.V. and Nicklaus, M.C. (2009). Optical structure recognition software to
recover chemical information: OSRA, an open source solution. Journal of Chemi- cal Information and Modeling 49 (3): 740–743.
8 Rajan, K., Zielesny, A., and Steinbeck, C. (2021). DECIMER 1.0: deep learning
for chemical image recognition using transformers. Journal of Cheminformatics 13: 61.
9 Staker, J., Marshall, K., Abel, R. et al. (2019). Molecular structure extraction from
documents using deep learning. Journal of Chemical Information and Modeling 59 (3): 1017–1029.
10 Oldenhof, M., Arany, A., Moreau, Y. et al. (2020). ChemGrapher: optical graph
recognition of chemical compounds by deep learning. Journal of Chemical Infor- mation and Modeling 60 (10): 4506–4517.
11 Frasconi, P., Gabbrielli, F., Lippi, M. et al. (2014). Markov logic networks for
optical chemical structure recognition. Journal of Chemical Information and Modeling 54 (8): 2380–2390.
12 Kim, S., Chen, J., Cheng, T.J. et al. (2021). PubChem in 2021: new data content
and improved web interfaces. Nucleic Acids Research 49 (D1): D1388–D1395.
13 Kim, S., Chen, J., Cheng, T.J. et al. (2019). PubChem 2019 update: improved
access to chemical data. Nucleic Acids Research 47 (D1): D1102–D1109.
14 Kim, S., Thiessen, P.A., Bolton, E.E. et al. (2016). PubChem Substance and Com-
pound databases. Nucleic Acids Research 44 (D1): D1202–D1213.
15 Kim, S. (2021). Exploring chemical information in pubChem. Current Protocols 1
(8): e217.
16 Kim, S. (2016). Getting the most out of PubChem for virtual screening. Expert
Opinion on Drug Discovery 11 (9): 843–855.
17 Kim, S., Cheng, T.J., He, S.Q. et al. (2022). PubChem Protein, Gene, Path-
way, and Taxonomy data collections: bridging biology and chemistry through target-centric views of pubChem data. Journal of Molecular Biology 434 (11):
167514.
18 Gaulton, A., Hersey, A., Nowotka, M. et al. (2017). The ChEMBL database in
2017. Nucleic Acids Research 45 (D1): D945–D954.
19 Southan, C., Sitzmann, M., and Muresan, S. (2013). Comparing the chemi-
cal structure and protein content of ChEMBL, drugBank, human metabolome database and the therapeutic target database. Molecular Informatics 32 (11-12): 881–897.
20 Gilson, M.K., Liu, T.Q., Baitaluk, M. et al. (2016). BindingDB in 2015: a pub-
lic database for medicinal chemistry, computational chemistry and systems pharmacology. Nucleic Acids Research 44 (D1): D1045–D1053.
62 2 PubChem: A Large-Scale Public Chemical Database for Drug Discovery
https://t.me/medicina_free
21 Sterling, T. and Irwin, J.J. (2015). ZINC 15-Ligand discovery for everyone.
Journal of Chemical Information and Modeling 55 (11): 2324–2337.
22 Harding, S., Armstrong, J., Faccenda, E. et al. (2021). IUPHAR/BPS Guide to
PHARMACOLOGY: expansion for anti-malarials, antibiotics and COVID-19. British Journal of Pharmacology 178 (2): 390–391.
23 Goldwaser, E., Laurent, C., Lagarde, N. et al. (2022). Machine learning-driven
identication of drugs inhibiting cytochrome P450 2C9. PLoS Computational Biology 18 (1): e1009820.
24 Wu, Z.X., Lei, T.L., Shen, C. et al. (2019). ADMET evaluation in drug discov-
ery. 19. reliable prediction of human cytochrome P450 inhibition using articial intelligence approaches. Journal of Chemical Information and Modeling 59 (11): 4587–4601.
25 Li, X., Xu, Y.J., Lai, L.H. et al. (2018). Prediction of human cytochrome P450
inhibition using a multitask deep autoencoder neural network. Molecular Phar- maceutics 15 (10): 4336–4345.
26 Lee, J.H., Basith, S., Cui, M. et al. (2017). In silico prediction of
multiple-category classication model for cytochrome P450 inhibitors and non-inhibitors using machine-learning method. SAR and QSAR in Environmental Research 28 (10): 863–874.
27 Su, B.H., Tu, Y.S., Lin, C. et al. (2015). Rule-based prediction models of
cytochrome P450 inhibition. Journal of Chemical Information and Modeling 55 (7): 1426–1434.
28 Kim, H. and Nam, H. (2020). hERG-Att: self-attention-based deep neural net-
work for predicting hERG blockers. Computational Biology and Chemistry 87:
107286.
29 Ogura, K., Sato, T., Tuki, H. et al. (2019). Support vector machine model for
hERG inhibitory activities based on the integrated hERG database using descrip­tor selection by NSGA-II. Scientic Reports 9: 12220.
30 Shen, M.Y., Su, B.H., Esposito, E.X. et al. (2011). A comprehensive support vec-
tor machine binary hERG classication model based on extensive but biased end point hERG data sets. Chemical Research in Toxicology 24 (6): 934–949.
31 Svensson, F., Norinder, U., and Bender, A. (2017). Modelling compound cytotox-
icity using conformal prediction and PubChem HTS data. Toxicology Research 6 (1): 73–80.
32 Russo, D.P., Strickland, J., Karmaus, A.L. et al. (2019). Nonanimal models for
acute toxicity evaluations: applying data-driven proling and read-across. Envi- ronmental Health Perspectives 127 (4): 047001.
33 Rodríguez-Pérez, R., Miyao, T., Jasial, S. et al. (2018). Prediction of compound
proling matrices using machine learning. ACS Omega 3 (4): 4713–4723.
34 Matlock, M.K., Hughes, T.B., Dahlin, J.L. et al. (2018). Modeling small-molecule
reactivity identies promiscuous bioactive compounds. Journal of Chemical Information and Modeling 58 (8): 1483–1500.
35 Stork, C., Wagner, J., Friedrich, N.O. et al. (2018). Hit dexter: a machine-learning
model for the prediction of frequent hitters. ChemMedChem 13 (6): 564–571.
References 63
https://t.me/medicina_free
36 Su, B.H., Tu, Y.S., Lin, O.A. et al. (2015). Rule-based classication models of
molecular autouorescence. Journal of Chemical Information and Modeling 55 (2): 434–445.
37 Ludwig, M., Dührkop, K., and Böcker, S. (2018). Bayesian networks for mass
spectrometric metabolite identication via molecular ngerprints. Bioinformatics 34 (13): 333–340.
38 Dührkop, K., Shen, H.B., Meusel, M. et al. (2015). Searching molecular struc-
ture databases with tandem mass spectra using CSI:FingerID. Proceedings of the National Academy of Sciences of the United States of America 112 (41): 12580–12585.
39 Qiu, F., Lei, Z.T., and Sumner, L.W. (2018). MetExpert: an expert system to
enhance gas chromatography-mass spectrometry-based metabolite identications. Analytica Chimica Acta 1037: 316–326.
40 Meyer, J.G., Liu, S.C., Miller, I.J. et al. (2019). Learning drug functions from
chemical structures with convolutional neural networks and random forests. Journal of Chemical Information and Modeling 59 (10): 4438–4449.
41 Korkmaz, S. (2020). Deep learning-based imbalanced data classication for drug
discovery. Journal of Chemical Information and Modeling 60 (9): 4180–4190.
42 Ancuceanu, R., Dinu, M., Neaga, I. et al. (2019). Development of QSAR
machine learning-based models to forecast the eect of substances on malignant melanoma cells. Oncology Letters 17 (5): 4188–4196.
43 Ciallella, H.L. and Zhu, H. (2019). Advancing computational toxicology in the
big data era by articial intelligence: data-driven and mechanism-driven model­ing for chemical toxicity. Chemical Research in Toxicology 32 (4): 536–547.
44 Danishuddin, Madhukar, G., Malik, M.Z. et al. (2019). Development and rigorous
validation of antimalarial predictive models using machine learning approaches. SAR and QSAR in Environmental Research 30 (8): 543–560.
45 Lauötter, O., Sturm, N., Bajorath, J. et al. (2019). Combining structural and
bioactivity-based ngerprints improves prediction performance and scaold hopping capability. Journal of Cheminformatics 11 (1): 54.
46 Capuzzi, S.J., Sun, W., Muratov, E.N. et al. (2018). Computer-aided discovery and
characterization of Novel Ebola virus inhibitors. Journal of Medicinal Chemistry 61 (8): 3582–3594.
47 Chen, J.J.F. and Visco, D.P. (2017). Identifying novel factor XIIa inhibitors with
PCA-GA-SVM developed vHTS models. European Journal of Medicinal Chemistry 140: 31–41.
48 Deshmukh, A.L., Chandra, S., Singh, D.K. et al. (2017). Identication of human
ap endonuclease 1 (FEN1) inhibitors using a machine learning based consensus virtual screening. Molecular BioSystems 13 (8): 1630–1639.
49 Bolton, E.E., Kim, S., and Bryant, S.H. (2011). PubChem3D: conformer genera-
tion. Journal of Cheminformatics 3: 4.
50 Kim, S., Bolton, E.E., and Bryant, S.H. (2013). PubChem3D: conformer ensemble
accuracy. Journal of Cheminformatics 5: 1.
51 Berman, H., Henrick, K., and Nakamura, H. (2003). Announcing the worldwide
protein data bank. Nature Structural Biology 10 (12): 980–980.
64 2 PubChem: A Large-Scale Public Chemical Database for Drug Discovery
https://t.me/medicina_free
52 Groom, C.R., Bruno, I.J., Lightfoot, M.P. et al. (2016). The cambridge structural
database. Acta Crystallographica. Section B: Structural Science, Crystal Engineer- ing and Materials 72: 171–179.
53 Hähnke, V.D., Kim, S., and Bolton, E.E. (2018). PubChem chemical structure
standardization. Journal of Cheminformatics 10: 36.
54 Weininger, D. (1990). SMILES .3. DEPICT – graphical depiction of chemical
structures. Journal of Chemical Information and Computer Sciences 30 (3): 237–243.
55 Weininger, D., Weininger, A., and Weininger, J.L. (1989). SMILES .2. Algorithm
for generation of unique smiles notation. Journal of Chemical Information and Computer Sciences 29 (2): 97–101.
56 Weininger, D. (1988). SMILES, a chemical language and information-system .1.
Introduction to methodology and encoding rules. Journal of Chemical Informa- tion and Computer Sciences 28 (1): 31–36.
57 Heller, S.R., McNaught, A., Pletnev, I. et al. (2015). InChI, the IUPAC interna-
tional chemical identier. Journal of Cheminformatics 7: 23.
58 Ihlenfeldt, W.D., Bolton, E.E., and Bryant, S.H. (2009). The PubChem chemical
structure sketcher. Journal of Cheminformatics 1: 20.
59 Grant, J.A., Gallardo, M.A., and Pickup, B.T. (1996). A fast method of molecular
shape comparison: a simple application of a Gaussian description of molecular shape. Journal of Computational Chemistry 17 (14): 1653–1666.
60 Grant, J.A. and Pickup, B.T. (1995). A gaussian description of molecular shape.
The Journal of Physical Chemistry 99 (11): 3503–3510.
61 Rush, T.S., Grant, J.A., Mosyak, L. et al. (2005). A shape-based 3-D scaold hop-
ping method and its application to a bacterial protein-protein interaction. Journal of Medicinal Chemistry 48 (5): 1489–1495.
62 Holliday, J.D., Salim, N., Whittle, M. et al. (2003). Analysis and display of the
size dependence of chemical similarity coecients. Journal of Chemical Informa- tion and Computer Sciences 43 (3): 819–828.
63 Holliday, J.D., Hu, C.Y., and Willett, P. (2002). Grouping of coecients for the
calculation of inter-molecular similarity and dissimilarity using 2D fragment bit-strings. Combinatorial Chemistry & High Throughput Screening 5 (2): 155–166.
64 Chen, X. and Reynolds, C.H. (2002). Performance of similarity measures in 2D
fragment-based similarity searching: comparison of structural descriptors and similarity coecients. Journal of Chemical Information and Computer Sciences 42 (6): 1407–1414.
65 Kim, S., Thiessen, P.A., Cheng, T. et al. (2016). Literature information in Pub-
associations between PubChem records and scientic articles. Journal of
Chem: Cheminformatics 8: 32.
66 Zaslavsky, L., Cheng, T., Gindulyte, A. et al. (2021). Discovering and summariz-
ing relationships between chemicals, genes, proteins, and diseases in PubChem. Frontiers in Research Metrics and Analytics 6: 689059.
67 Bolton, E.E., Kim, S., and Bryant, S.H. (2011). PubChem3D: similar conformers.
Journal of Cheminformatics 3: 13.
References 65
https://t.me/medicina_free
68 Kim, S., Bolton, E.E., and Bryant, S.H. (2016). Similar compounds versus similar
conformers: complementarity between PubChem 2-D and 3-D neighboring sets. Journal of Cheminformatics 8: 62.
69 Kim, S., Thiessen, P.A., Bolton, E.E. et al. (2015). PUG-SOAP and PUG-REST:
web services for programmatic access to chemical information in PubChem. Nucleic Acids Research 43 (W1): W605–W611.
70 Kim, S., Thiessen, P.A., Cheng, T.J. et al. (2018). An update on PUG-REST:
RESTful interface for programmatic access to PubChem. Nucleic Acids Research 46 (W1): W563–W570.
71 Kim, S., Shoemaker, B.A., Bolton, E.E. et al. (2018). Finding potential multitarget
ligands using PubChem. In: Computational Chemogenomics (ed. J.B. Brown), 63–91. Totowa: Humana Press Inc.
72 Kim, S., Thiessen, P.A., Cheng, T.J. et al. (2019). PUG-View: programmatic access
to chemical annotations integrated in PubChem. Journal of Cheminformatics 11: 56.
73 Fu, G., Batchelor, C., Dumontier, M. et al. (2015). PubChemRDF: towards the
semantic annotation of PubChem Compound and Substance databases. Journal of Cheminformatics 7: 34.
74 Boehr, D.L. and Bushman, B. (2018). Preparing for the future: National Library
of Medicine’s®project to add MeSH®RDF URIs to its bibliographic and authority records. Cataloging and Classication Quarterly 56 (2-3): 262–272.
75 Bushman, B., Anderson, D., and Fu, G. (2015). Transforming the medical sub-
ject headings into linked data: creating the authorized version of MeSH in RDF. Journal of Library Metadata 15 (3-4): 157–176.
76 Redaschi, N. and Uniprot Consortium (2009). UniProt in RDF: tackling data
integration and distributed annotation with the semantic web. Nature Precedings https://doi.org/10.1038/npre.2009.3193.1.
77 Kinjo, A.R., Bekker, G.-J., Suzuki, H. et al. (2016). Protein Data Bank Japan
(PDBj): updated user interfaces, resource description framework, analysis tools for large structures. Nucleic Acids Research 45 (D1): D282–D288.
78 Kinjo, A.R., Suzuki, H., Yamashita, R. et al. (2011). Protein Data Bank Japan
(PDBj): maintaining a structural data archive and resource description frame­work format. Nucleic Acids Research 40 (D1): D453–D460.
79 Jupp, S., Malone, J., Bolleman, J. et al. (2014). The EBI RDF platform: linked
open data for the life sciences. Bioinformatics 30 (9): 1338–1339.
80 Erxleben, F., Günther, M., Krötzsch, M. et al. (2014). Introducing Wikidata to the
Linked Data Web, 50–65. Cham: Springer International Publishing.
81 Lipinski, C.A., Lombardo, F., Dominy, B.W. et al. (1997). Experimental and com-
putational approaches to estimate solubility and permeability in drug discovery and development settings. Advanced Drug Delivery Reviews 23 (1-3): 3–25.
82 Ghose, A.K., Viswanadhan, V.N., and Wendoloski, J.J. (1999). A knowledge-based
approach in designing combinatorial or medicinal chemistry libraries for drug discovery. 1. A qualitative and quantitative characterization of known drug databases. Journal of Combinatorial Chemistry 1 (1): 55–68.
66 2 PubChem: A Large-Scale Public Chemical Database for Drug Discovery
https://t.me/medicina_free
83 Veber, D.F., Johnson, S.R., Cheng, H.Y. et al. (2002). Molecular properties
that inuence the oral bioavailability of drug candidates. Journal of Medicinal Chemistry 45 (12): 2615–2623.
84 Walters, W. P. and Namchuk, M. (2003). Designing screens: How to make your
hits a hit. Nature Reviews Drug Discovery 2 (4): 259–266.
85 Bickerton, G.R., Paolini, G.V., Besnard, J. et al. (2012). Quantifying the chemical
beauty of drugs. Nature Chemistry 4 (2): 90–98.
86 Congreve, M., Carr, R., Murray, C. et al. (2003). A rule of three for
fragment-based lead discovery? Drug Discovery Today 8 (19): 876–877.
87 Williams, A.J., Grulke, C.M., Edwards, J. et al. (2017). The CompTox chemistry
dashboard: a community data resource for environmental chemistry. Journal of Cheminformatics 9: 61.
88 Hastings, J., Owen, G., Dekker, A. et al. (2016). ChEBI in 2016: improved ser-
vices and an expanding collection of metabolites. Nucleic Acids Research 44 (D1): D1214–D1219.
89 Armstrong, D.R., Berrisford, J.M., Conroy, M.J. et al. (2019). PDBe: improved
ndability of macromolecular structure data in the PDB. Nucleic Acids Research 48 (D1): D335–D343.
90 Bansal, P. , Morgat, A., Axelsen, K.B. et al. (2021). Rhea, the reaction knowledge-
base in 2022. Nucleic Acids Research 50 (D1): D693–D700.
91 Chambers, J., Davies, M., Gaulton, A. et al. (2014). UniChem: extension of
InChI-based compound mapping to salt, connectivity and stereochemistry layers. Journal of Cheminformatics 6: 43.
92 Chambers, J., Davies, M., Gaulton, A. et al. (2013). UniChem: a unied chemical
structure cross-referencing and identier tracking system. Journal of Cheminfor- matics 5: 3.
3
https://t.me/medicina_free
DrugBank Online: A How-to Guide
Christen M. Klinger1,JordanCox1, Denise So2, Teira Stauth3, Michael Wilson1, Alex Wilson1, and Craig Knox
1
University of Alberta, Edmonton, AB, Canada
2
University of Ottawa, Ottawa, ON, Canada
3
University of Calgary, Calgary, AB, Canada
1
3.1 Introduction
The process of discovering new drugs has undergone substantial changes from the historical paradigm of directly using natural products to high-throughput screening (HTS) to the modern state-of-the-art that melds HTS and computational approaches [1, 2]. Despite these advancements, drug discovery remains challenging; high attri­tion rates and costs remain barriers to entry for smaller biotech companies and aca­demic institutions [3]. In addition, the promise of HTS as a means for overcoming the practical limitations of smaller-scale molecular interaction studies has fallen short. Antibiotics remain an illustrative example where, despite expansive library screen­ing, few truly novel chemical scaolds have been discovered since the 1960s [4].
Modern drug discovery follows a “target-centric” paradigm, wherein both a drug candidate and at least one validated target will be identied prior to proceeding to clinical studies [5]. This not only ensures a conceptual mechanistic rationale for drug action but also facilitates renement of the candidate molecule through the iterative application of structure–activity relationship studies [6, 7]. Regardless of the exact technologies used, the process of identifying both candidate molecules and targets can be conceptually divided into two schemes, depending on the order in which they are sought [2].
In so-called “target-based” discovery [2], the rst step is the identication of one or more putative targets, whose properties may be manipulated through the addition of specic molecules in order to achieve a therapeutic eect. These targets may be pri­oritized based on previous data, by genetic signals linking the target to the condition in question (e.g. in genome-wide association studies), or may represent promising hits identied in large-scale screening eorts [5]. Each putative target might then serve as the basis for the development of one or more candidate molecules.
67
Open Access Databases and Datasets for Drug Discovery, First Edition. Edited by Antoine Daina, Michael Przewosny, and Vincent Zoete. © 2024 WILEY-VCH GmbH. Published 2024 by WILEY-VCH GmbH.
68 3 DrugBank Online: A How-to Guide
https://t.me/medicina_free
In contrast, “phenotypic-based” (sometimes referred to as “molecule-based”) discovery [2] takes advantage of biological screening platforms with a well­characterized phenotypic readout to screen large compound libraries. High-scoring hits are then deconvoluted to identify the putative target(s) prior to more direct target validation [8].
As the end result, a set of candidate molecules and known targets, remains unchanged between these schemes, the choice of which approach to use is largely contextual. More recent advances in the last approximately two decades have also seen an increasing dependency across both schemes for computational tools. The ability to investigate intermolecular interactions between putative candidate molecules and proteins (e.g. by molecular docking) and to simulate these interac­tions in a temporal manner (e.g. by molecular dynamics) have signicantly increased the size of screening libraries from ∼106compounds to virtual libraries of ∼10 compounds [9]. More recently, articial intelligence/machine learning (AI/ML) tools that assist in target prediction, compound identication/design, high-content screening, candidate retrosynthesis, and other areas have been developed [10, 11].
The potential for computational methods to revolutionize drug discovery is undoubtedly exciting, but it brings new challenges with it. Computational approaches require highly structured and accurate data in order to work eectively. However, data within the biomedical space are often unstructured and diuse, being scattered across a collection of publications and databases, each with its own specic schemas and formats.
In this chapter, we present an updated view of the DrugBank database. We explore the nature and kinds of data within DrugBank, with an emphasis on the structured nature of the data. Protocols are provided for the use of the free online data and tools available through our website. Some examples of published research using DrugBank are provided for further context and as starting points for further exploration of the myriad ways in which this resource may assist those working in the drug discovery eld. Lastly, we discuss our current outlook and possible avenues for future development.
13
3.2 DrugBank
3.2.1 Overview of DrugBank
DrugBank maintains a comprehensive database of drugs and related information, which we will explore further in Section 3.2.2. Data within DrugBank are a mixture of data imported from relevant databases, novel content authored by an in-house curation team of pharmacists, doctors, and biomedical experts, and insights gener­ated from proprietary ML-based workows (Figure 3.1). Regardless of the source, data within DrugBank are highly structured, such that they can be used to train ML models or form the basis for novel algorithms and pipelines.
Although commercialized in 2015, DrugBank started as an academic research project and is committed to supporting the research community through access
External datasets such as:
https://t.me/medicina_free
MedDRA, RxNorm,
and Ontologies
Resources such as:
Publications, monographs,
clinical trials, guidelines,
public content
Al Powered
Knowledge Extraction
Advanced Natural Language Processing & Data Synthesis
In-house experts
3.2 DrugBank 69
External
Dataset Linking
Knowledgebase
DrugBank Users
Knowledge
Augmentation
Figure 3.1 The DrugBank knowledgebase. This figure shows a graphical overview of the DrugBank knowledgebase, which powers DrugBank Online and our other product offerings. Proprietary ML tools scour literary resources such as publications, monographs, public data resources, and others to bring relevant content to an in-house curation team (top). This team then vets the content, synthesizes multiple pieces, and authors novel information for inclusion in the knowledgebase. Concurrently, recognized external datasets in the biomedical space, such as ontologies related to drug products and their possible medical effects, are imported and processed for inclusion into the k nowledgebase (left). All of these data together provide the potential for cyclical knowledge augmentation, wherein the combination of information can yield new insights that are “larger than the sum of their parts” (right). The combined information from all of these sources is provided to users in a highly structured and usable format (bottom).
to free datasets and tools. Public-facing content is available as part of DrugBank Online (https://go.drugbank.com/) and will form the bulk of this chapter. Further support for academic (through our Academic + program) and commercial users, including structured data downloads in a variety of formats, will not be discussed in detail here but may be obtained through the “Solutions” tab at the top of this webpage.
3.2.2 DrugBank Datasets
DrugBank contains a wide variety of datasets for use in both clinical and scientic applications, not all of which will be discussed here. The primary purpose of this section is to provide an overview of key datasets within DrugBank, including their