Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5419_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
15.09.2026
Размер:
11 Мб
Скачать
☆
https://t.me/medicina_free
2
https://t.me/medicina_free
PubChem: A Large-Scale Public Chemical Database for Drug Discovery
Sunghwan Kim and Evan E. Bolton
National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, 8600 Rockville Pike, Bethesda, MD 20894, USA
2.1 Introduction
Advances in combinatorial chemistry (CC) and high-throughput screening (HTS) technologies have made it possible to rapidly test the biological activity of mil­lions of chemicals at a very low cost. In addition, computational approaches for named-entity recognition [1–5] and optical structure recognition [6–11] are now commonly used to extract chemical information from various documents, such as scientic articles, patents, and government reports. These technological advances have signicantly increased the amount of chemical data available in the public domain, creating a demand for public information resources that can collect, organize, and disseminate this data. It led to the development of many public databases, such as PubChem [12–17], ChEMBL [18], DrugBank [19], BindingDB [20], ZINC [21], and IUPHAR/BPS Guide to Pharmacology [22].
PubChem (https://pubchem.ncbi.nlm.nih.gov) (Figure 2.1) is a public chemical database at the U.S. National Institutes of Health (NIH) [12–17]. It collects chemical information from hundreds of data sources and disseminates it to the public free of charge. With more than 110 million unique chemical structures (as of 30 January
2022), PubChem is considered as one of the largest chemical databases in the pub­lic domain. Visited by millions of unique users every month [12], PubChem serves a wide range of users, including scientists, chemical safety ocers, patent agents, edu­cators, students, and many others. Especially, PubChem is an important resource for biomedical research communities in the areas of cheminformatics, chemical biology, medicinal chemistry, and drug discovery.Importantly,PubChem data are commonly used to build machine-learning models to predict various chemical properties and biological activities [23–48].
This chapter provides an overview of PubChem, including its data contents rele­vant to drug discovery as well as the tools and services that use these data. The chem­ical space of PubChem is compared with those of other popular chemical databases. Important characteristics of bioactivity data archived in PubChem are discussed.
41
Open Access Databases and Datasets for Drug Discovery, First Edition. Edited by Antoine Daina, Michael Przewosny, and Vincent Zoete. © 2024 WILEY-VCH GmbH. Published 2024 by WILEY-VCH GmbH.
42 2 PubChem: A Large-Scale Public Chemical Database for Drug Discovery
https://t.me/medicina_free
Figure 2.1 PubChem home page (https://pubchem.ncbi.nlm.nih.gov). The user can search PubChem by providing a keyword query in the search box ( can be provided using the PubChem Sketcher ( PubChem record identifiers using the “Upload ID List” button ( classification browser ( have a particular annotation. Clicking the “About” link ( documentation site (called PubChem Docs), where the user can also access additional tools and services.
4
) allows users to get records that belong to a particular class or
2
). It is also possible to provide a list of
1
). A chemical structure query
3
). The PubChem
5
) directs to PubChem’s help
2.2 Data Content and Organization
PubChem provides a wide range of chemical information. It contains computation­ally generated 3D structures of chemicals [49, 50] as well as links to experimentally determined 3D structures of chemicals available at the Protein Data Bank (PDB) [51] and the Cambridge Structural Database (CSD) [52]. Various kinds of molecular properties are available in PubChem and many of them are pertinent to drug discovery (e.g. molecular weight, solubility, octanol–water partition coecient (log P), Caco2 permeability, acid dissociation constant (pKa), carcinogenicity, and mutagenicity). In addition, a large amount of bioactivity data submitted by data depositors are archived in PubChem. Substantial quantities of annotations on approved and investigational drugs are integrated into PubChem, such as drug labeling, indications, target genes and proteins, mechanisms of action, absorption, distribution, metabolism, excretion, and toxicity (ADMET) properties, and clinical trials carried out in the United States, Europe, and Japan. Moreover, PubChem has spectral information, including mass spectrometry (MS), infrared (IR), ultraviolet (UV), and nuclear magnetic resonance (NMR) spectroscopy data. It also provides
2.2 Data Content and Organization 43
https://t.me/medicina_free
information on synthesis and chemical vendors, as well as scientic articles and patent documents that mention chemicals.
The PubChem Data Sources page (https://pubchem.ncbi.nlm.nih.gov/sources) provides an interactive overview of organizations contributing data to PubChem. As of 30 January 2022, the data contained in PubChem are from more than 800 data sources, including U.S. government agencies, international organizations, academic institutions, pharmaceutical companies, chemical vendors, and other chemical biology databases. PubChem plays a dual role as an archive, which stores original chemical data provided by individual sources without any mod­ication, and as a knowledgebase, which provides users with well-organized, high-quality information about chemicals. This dual role is reected in the data organization in PubChem. PubChem data is organized into multiple data col­lections, including Substance, Compound, BioAssay, Gene, Protein, Pathway, Taxonomy, and Patent (Figure 2.2). Substance archives chemical descriptions submitted by individual data providers. Compound contains unique chemical structures extracted from Substance through chemical structure standardization [53]. BioAssay stores biological assay descriptions and test results, provided by assay data providers. The Gene, Protein, Pathway, and Taxonomy collections provide information on chemicals related to a given gene, protein, pathway, and taxon, respectively [17]. The Patent collection contains chemicals mentioned in a given patent document. Among these data collections, Substance and BioAssay are
Knowledgebase
Substance
Depositor-provided
chemical data
Protein
Taxonomy
Chemical data associated with
a protein/gene/pathway/taxon/patent
Figure 2.2 PubChem data collections. While Substance and BioAssay collections (indicated in red boxes) serve as archives, the other data collections (indicated in blue boxes) are knowledgebases.
Compound
Unique chemical
structures
Gene Pathway
Patent
ArchiveArchive
BioAssay
Assay descriptions
& test results
44 2 PubChem: A Large-Scale Public Chemical Database for Drug Discovery
https://t.me/medicina_free
archives that store depositor-provided data, while the other collections serve as knowledgebases.
It is noteworthy that, because of their archival nature, the data in Substance and BioAssay are kept as they were at the time of data submission by the sources. In essence, these data are owned and controlled by the data contributors. When necessary, the data source, not PubChem, may correct or update a record in Substance or BioAssay: the record will be versioned, and both the new and original ones will be retained and accessible. PubChem does, however, facilitate corrections to the data by working with the data submitter, when errors are made known to PubChem.
Each record in the Substance, Compound, and BioAssay collections is assigned a numeric identier called Substance ID (SID), Compound ID (CID), and Assay ID (AID), respectively. Records in the Substance and Compound collections are called substances and compounds. It is worth mentioning that users are often confused with these two PubChem-specic terms. Simply put, while substances are depositor-provide descriptions of chemicals, compounds are unique chemical structures extracted from substances, meaning that a compound may be associated with multiple substances. Currently, PubChem contains 110 million compounds, extracted from 277 million substances (Figure 2.3). Detailed discussion about the substances and compounds is given in our previous paper [14] and blog post (http:// go.usa.gov/x72qw).
300
250
200
150
100
Number of records (milions)
50
0
Figure 2.3 Growth of substance and compound records in PubChem. See the text for the definition of substances and compounds in PubChem. The data underlying this chart were generated on 30 January 2022.
2004
2005
Substances Compounds
2006
2007
2008
2009
2011
2010
2012
2013
2014
2015
2016
2017
2018
2019
2020
2021
2022
Year
2.3 Tools and Services 45
https://t.me/medicina_free
2.3 Tools and Services
PubChem provides various tools and services to assist users in exploiting PubChem data, and they can be accessed through the PubChem homepage or the PubChem Help site (https://pubchemdocs.ncbi.nlm.nih.gov) ( provides a brief overview of some of these tools and services, while more details can be found at the PubChem Help site. In addition, step-by-step instructions on how to explore PubChem data through web browsers are given in our recent protocol paper [15].
2.3.1 PubChem Search
PubChem data can be searched from the PubChem home page (https://pubchem .ncbi.nlm.nih.gov) (Figure 2.1), which also serves as the entry point to various tools and services. A simple keyword search can be initiated by providing the query keyword in the search box (
1
in Figure 2.1). PubChem accepts various
types of keywords, including chemical names, chemical abstract service (CAS) registry numbers, PubChem record identiers (SID, CID, and AID), gene/protein names and symbols, and disease names. When a keyword query is provided, PubChem searches all data collections simultaneously and returns hit records for individual collections (Figure 2.4). It also tries to identify the most relevant record and presents it at the top of the search result page ( hits from a given collection can be viewed by clicking the corresponding tab
2
(
in Figure 2.4). Users can rene this hit list based on some select attributes
using lters ( the search result page (
3
in Figure 2.4). The buttons available on the right column of
4
in Figure 2.4) allow users to perform additional tasks
with the hit records, such as downloading them on a local machine, saving them for later use, or getting other records related to the hits. Clicking one of the returned hits leads to its Summary page, which displays all information available in PubChem for a given record (to be discussed in 2.3.2 Summary pages section).
The Compound collection can also be searched using a chemical structure query. The input structure can be provided using a simplied molecular-input line-entry system (SMILES) [54–56] or International Chemical Identier (InChI) string [57]. Alternatively, it can be drawn using the PubChem Sketcher [58], which is accessible
2
from the PubChem homepage (
in Figure 2.1). When a chemical structure input
is provided, multiple types of structure searches are simultaneously performed, including identity searches, 2-dimensional (2D) and 3-dimensional (3D) similarity searches, and sub- and superstructure searches. The result for each search type can be accessed through the corresponding tab (
1
in Figure 2.5). Users can customize
the structure search by changing the parameters and options used for the search through the Settings button (
2
in Figure 2.5).
It is noteworthy that PubChem supports two types of similarity search, based on ngerprint-based 2D similarity and Gaussian-shape overlay-based 3D similarity methods [59–61]. The 2D similarity between molecules is calculated by using
5
in Figure 2.1). This section
1
in Figure 2.4). The
46 2 PubChem: A Large-Scale Public Chemical Database for Drug Discovery
https://t.me/medicina_free
Figure 2.4 Search result page for a text query (ascorbic acid as an example) (https:// pubchem.ncbi.nlm.nih.gov/#query=ascorbic%20acid). The best hit ( the top, and the hits from each data collection can be accessed by clicking the corresponding tab ( additional tasks can be done through the buttons on the right column of the page (
the PubChem subgraph ngerprints in conjunction with the Tanimoto equation [62–64]:
Tanimoto =
where NAand NBare the respective counts of ngerprint bits set in molecules A and B, and NABis the count of bits set in common. On the other hand, 3D molecular similarity is quantied with the shape-Tanimoto (ST) [15, 59, 60], which evaluates steric shape similarity, and color-Tanimoto (CT) [15], which quanties functional group similarity. They are dened as:
1
) will be presented at
2
). The search result can be refined using the filters (3), and
N
AB
NA+NB−N
AB
4
).
(2.1)
2.3 Tools and Services 47
https://t.me/medicina_free
Figure 2.5 Search result page for a chemical structure query (the SMILES string for ascorbic acid as an example) (https://pubchem.ncbi.nlm.nih.gov/#query=C([C@@H] ([C@@H]1C(=C(C(=O)O1)O)O)O)O). PubChem performs multiple types of structure searches against the Compound collection, and the results for each search type can be accessed through the corresponding tab ( can be adjusted using the Settings button (
3
filters (
), and additional tasks can be done through the buttons on the right column of the
4
page (
).
V
ST =
VAA+VBB−V
CT =
∑
AB
f
V
+
f
AA
1
). The parameters and options used for structure searches
AB
∑
f
V
f
AB
∑
f
V
−
f
BB
2
). The search result can be refined using the
∑
f
V
f
AB
(2.2)
(2.3)
where VAAand VBBare the self-overlap volumes of molecules A and B, respec­tively, and the VABis the overlap volume between them. In Eq (2.3), the index f indicates any of six functional group types (i.e. hydrogen-bond donors and accep­tors, cations, anions, hydrophobes, and rings), represented by ctitious “feature” or “color” atoms. V
f
AA
and V
f
are the self-overlap volumes of A and B for feature
BB
48 2 PubChem: A Large-Scale Public Chemical Database for Drug Discovery
https://t.me/medicina_free
atom type f , respectively, and V
f
is the overlap volume of molecules AandBfor
AB
feature atom type f .
The ST and CT scores can be combined to create a Combo-Tanimoto (ComboT) score, which simultaneously considers both steric shape similarity and functional group similarity:
ComboT = ST + CT (2.4)
Because both ST and CT scores range from 0 to 1, the ComboT score ranges from 0 to 2 (without normalization). The ST, CT, and ComboT scores between molecules can be evaluated in two dierent molecular superpositions: the ST- or shape-optimized superposition and the CT- or feature-optimized superposition. In the shape-optimization, the superposition of two molecules is optimized to have a maximum ST score. In the feature-optimization, both shapes and features of the molecules are simultaneously considered to nd the best superposition.
By default, 2D similarity search returns compounds whose Tanimoto score relative to the query molecule is equal to or greater than 0.90. For 3D similarity searches, compounds with ST ≥ 0.80 and CT ≥ 0.50 are returned. More detailed information on 2D and 3D similarity searches is given in our previous paper [15].
2.3.2 Summary Pages
The Summary page for a given PubChem record displays all information avail­able for that record. For the records in the Substance and BioAssay collections, which are archival in nature, their Summary page shows the current version of depositor-provided data by default. An older version of data can be displayed by selecting the desired version from the dropdown menu (
1
in Figure 2.6). For the
records in the other data collections, which serve as knowledgebases, the Summary page contains not only relevant depositor-provided data but also annotation data collected by the PubChem crew from external authoritative sources. These annotations are regularly updated to provide up-to-date information.
The Summary page for a given record has links to other related records (in the same or dierent collections), providing users with quick access to information about them. In addition, the annotation data are presented with the data sources, allowing users to go to the original data source to check the context of the data and obtain additional information. Users can quickly access the desired information
2
by using the Table of Contents available in the right column (
in Figure 2.6). The
data presented on the Summary page are downloadable using the Download button available at the top of the right column (
3
in Figure 2.6). It is also possible to down-
load the data presented under an individual (sub)section of the Summary page. Because each (sub)section of the Summary page is widgetized, it can be embedded within the user’s web page. More information on PubChem Widgets is available in its help document, available at: https://pubchem.ncbi.nlm.nih.gov/docs/widgets.
2.3 Tools and Services 49
https://t.me/medicina_free
Figure 2.6 Summary page for SID 46505070 (https://pubchem.ncbi.nlm.nih.gov/sub­stance/46505070). The dropdown menu ( substance record. The Table of Contents on the right column ( Summary page. The data presented on this page can be downloaded by using the Download button (
3
) above the Table of Contents.
1
) allows users to view the older version of this
2
) helps navigate the
2.3.3 Literature Knowledge Panel
PubChem contains a great deal of information on scientic articles and patent doc­uments that mention chemicals and their bioactivity data [65]. Users often desire to explore these articles to learn about the relationships among chemicals, genes, pro­teins, and diseases, which is not trivial given the size and scope of PubChem data. To assist users in this task, the Literature Knowledge Panels [66] are presented on the Summary page of a compound, gene, or protein. The Literature Knowledge Panels for a given entity (i.e. a chemical, gene, or protein) display a few of its most rele­vant, nonredundant “neighbors,” which are dened as other entities co-mentioned in scientic articles (e.g. chemicals, genes, proteins, and diseases). The panels also provide a sample of PubMed records co-mentioning the entity and its neighbors.