Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5660_Библиотеки_им_академика_М_И_Перельмана
.pdf
https://t.me/medicina_free

2
https://t.me/medicina_free
PubChem: A Large-Scale Public Chemical Database for Drug
Discovery
Sunghwan Kim and Evan E. Bolton
National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health,
8600 Rockville Pike, Bethesda, MD 20894, USA
2.1 Introduction
Advances in combinatorial chemistry (CC) and high-throughput screening (HTS)
technologies have made it possible to rapidly test the biological activity of millions of chemicals at a very low cost. In addition, computational approaches for
named-entity recognition [1–5] and optical structure recognition [6–11] are now
commonly used to extract chemical information from various documents, such as
scientic articles, patents, and government reports. These technological advances
have signicantly increased the amount of chemical data available in the public
domain, creating a demand for public information resources that can collect,
organize, and disseminate this data. It led to the development of many public
databases, such as PubChem [12–17], ChEMBL [18], DrugBank [19], BindingDB
[20], ZINC [21], and IUPHAR/BPS Guide to Pharmacology [22].
PubChem (https://pubchem.ncbi.nlm.nih.gov) (Figure 2.1) is a public chemical
database at the U.S. National Institutes of Health (NIH) [12–17]. It collects chemical
information from hundreds of data sources and disseminates it to the public free of
charge. With more than 110 million unique chemical structures (as of 30 January
2022), PubChem is considered as one of the largest chemical databases in the public domain. Visited by millions of unique users every month [12], PubChem serves a
wide range of users, including scientists, chemical safety ocers, patent agents, educators, students, and many others. Especially, PubChem is an important resource for
biomedical research communities in the areas of cheminformatics, chemical biology,
medicinal chemistry, and drug discovery.Importantly,PubChem data are commonly
used to build machine-learning models to predict various chemical properties and
biological activities [23–48].
This chapter provides an overview of PubChem, including its data contents relevant to drug discovery as well as the tools and services that use these data. The chemical space of PubChem is compared with those of other popular chemical databases.
Important characteristics of bioactivity data archived in PubChem are discussed.
41
Open Access Databases and Datasets for Drug Discovery, First Edition.
Edited by Antoine Daina, Michael Przewosny, and Vincent Zoete.
© 2024 WILEY-VCH GmbH. Published 2024 by WILEY-VCH GmbH.

42 2 PubChem: A Large-Scale Public Chemical Database for Drug Discovery
https://t.me/medicina_free
Figure 2.1 PubChem home page (https://pubchem.ncbi.nlm.nih.gov). The user can search
PubChem by providing a keyword query in the search box (
can be provided using the PubChem Sketcher (
PubChem record identifiers using the “Upload ID List” button (
classification browser (
have a particular annotation. Clicking the “About” link (
documentation site (called PubChem Docs), where the user can also access additional tools
and services.
4
) allows users to get records that belong to a particular class or
2
). It is also possible to provide a list of
1
). A chemical structure query
3
). The PubChem
5
) directs to PubChem’s help
2.2 Data Content and Organization
PubChem provides a wide range of chemical information. It contains computationally generated 3D structures of chemicals [49, 50] as well as links to experimentally
determined 3D structures of chemicals available at the Protein Data Bank (PDB)
[51] and the Cambridge Structural Database (CSD) [52]. Various kinds of molecular
properties are available in PubChem and many of them are pertinent to drug
discovery (e.g. molecular weight, solubility, octanol–water partition coecient
(log P), Caco2 permeability, acid dissociation constant (pKa), carcinogenicity, and
mutagenicity). In addition, a large amount of bioactivity data submitted by data
depositors are archived in PubChem. Substantial quantities of annotations on
approved and investigational drugs are integrated into PubChem, such as drug
labeling, indications, target genes and proteins, mechanisms of action, absorption,
distribution, metabolism, excretion, and toxicity (ADMET) properties, and clinical
trials carried out in the United States, Europe, and Japan. Moreover, PubChem has
spectral information, including mass spectrometry (MS), infrared (IR), ultraviolet
(UV), and nuclear magnetic resonance (NMR) spectroscopy data. It also provides

2.2 Data Content and Organization 43
https://t.me/medicina_free
information on synthesis and chemical vendors, as well as scientic articles and
patent documents that mention chemicals.
The PubChem Data Sources page (https://pubchem.ncbi.nlm.nih.gov/sources)
provides an interactive overview of organizations contributing data to PubChem.
As of 30 January 2022, the data contained in PubChem are from more than 800
data sources, including U.S. government agencies, international organizations,
academic institutions, pharmaceutical companies, chemical vendors, and other
chemical biology databases. PubChem plays a dual role as an archive, which
stores original chemical data provided by individual sources without any modication, and as a knowledgebase, which provides users with well-organized,
high-quality information about chemicals. This dual role is reected in the data
organization in PubChem. PubChem data is organized into multiple data collections, including Substance, Compound, BioAssay, Gene, Protein, Pathway,
Taxonomy, and Patent (Figure 2.2). Substance archives chemical descriptions
submitted by individual data providers. Compound contains unique chemical
structures extracted from Substance through chemical structure standardization
[53]. BioAssay stores biological assay descriptions and test results, provided by assay
data providers. The Gene, Protein, Pathway, and Taxonomy collections provide
information on chemicals related to a given gene, protein, pathway, and taxon,
respectively [17]. The Patent collection contains chemicals mentioned in a given
patent document. Among these data collections, Substance and BioAssay are
Knowledgebase
Substance
Depositor-provided
chemical data
Protein
Taxonomy
Chemical data associated with
a protein/gene/pathway/taxon/patent
Figure 2.2 PubChem data collections. While Substance and BioAssay collections
(indicated in red boxes) serve as archives, the other data collections (indicated in blue
boxes) are knowledgebases.
Compound
Unique chemical
structures
Gene Pathway
Patent
ArchiveArchive
BioAssay
Assay descriptions
& test results

44 2 PubChem: A Large-Scale Public Chemical Database for Drug Discovery
https://t.me/medicina_free
archives that store depositor-provided data, while the other collections serve as
knowledgebases.
It is noteworthy that, because of their archival nature, the data in Substance
and BioAssay are kept as they were at the time of data submission by the sources.
In essence, these data are owned and controlled by the data contributors. When
necessary, the data source, not PubChem, may correct or update a record in
Substance or BioAssay: the record will be versioned, and both the new and original
ones will be retained and accessible. PubChem does, however, facilitate corrections
to the data by working with the data submitter, when errors are made known to
PubChem.
Each record in the Substance, Compound, and BioAssay collections is assigned
a numeric identier called Substance ID (SID), Compound ID (CID), and Assay
ID (AID), respectively. Records in the Substance and Compound collections are
called substances and compounds. It is worth mentioning that users are often
confused with these two PubChem-specic terms. Simply put, while substances
are depositor-provide descriptions of chemicals, compounds are unique chemical
structures extracted from substances, meaning that a compound may be associated
with multiple substances. Currently, PubChem contains 110 million compounds,
extracted from 277 million substances (Figure 2.3). Detailed discussion about the
substances and compounds is given in our previous paper [14] and blog post (http://
go.usa.gov/x72qw).
300
250
200
150
100
Number of records (milions)
50
0
Figure 2.3 Growth of substance and compound records in PubChem. See the text for the
definition of substances and compounds in PubChem. The data underlying this chart were
generated on 30 January 2022.
2004
2005
Substances
Compounds
2006
2007
2008
2009
2011
2010
2012
2013
2014
2015
2016
2017
2018
2019
2020
2021
2022
Year

2.3 Tools and Services 45
https://t.me/medicina_free
2.3 Tools and Services
PubChem provides various tools and services to assist users in exploiting PubChem
data, and they can be accessed through the PubChem homepage or the PubChem
Help site (https://pubchemdocs.ncbi.nlm.nih.gov) (
provides a brief overview of some of these tools and services, while more details can
be found at the PubChem Help site. In addition, step-by-step instructions on how
to explore PubChem data through web browsers are given in our recent protocol
paper [15].
2.3.1 PubChem Search
PubChem data can be searched from the PubChem home page (https://pubchem
.ncbi.nlm.nih.gov) (Figure 2.1), which also serves as the entry point to various
tools and services. A simple keyword search can be initiated by providing the
query keyword in the search box (
1
in Figure 2.1). PubChem accepts various
types of keywords, including chemical names, chemical abstract service (CAS)
registry numbers, PubChem record identiers (SID, CID, and AID), gene/protein
names and symbols, and disease names. When a keyword query is provided,
PubChem searches all data collections simultaneously and returns hit records
for individual collections (Figure 2.4). It also tries to identify the most relevant
record and presents it at the top of the search result page (
hits from a given collection can be viewed by clicking the corresponding tab
2
(
in Figure 2.4). Users can rene this hit list based on some select attributes
using lters (
the search result page (
3
in Figure 2.4). The buttons available on the right column of
4
in Figure 2.4) allow users to perform additional tasks
with the hit records, such as downloading them on a local machine, saving
them for later use, or getting other records related to the hits. Clicking one of
the returned hits leads to its Summary page, which displays all information
available in PubChem for a given record (to be discussed in 2.3.2 Summary pages
section).
The Compound collection can also be searched using a chemical structure query.
The input structure can be provided using a simplied molecular-input line-entry
system (SMILES) [54–56] or International Chemical Identier (InChI) string [57].
Alternatively, it can be drawn using the PubChem Sketcher [58], which is accessible
2
from the PubChem homepage (
in Figure 2.1). When a chemical structure input
is provided, multiple types of structure searches are simultaneously performed,
including identity searches, 2-dimensional (2D) and 3-dimensional (3D) similarity
searches, and sub- and superstructure searches. The result for each search type can
be accessed through the corresponding tab (
1
in Figure 2.5). Users can customize
the structure search by changing the parameters and options used for the search
through the Settings button (
2
in Figure 2.5).
It is noteworthy that PubChem supports two types of similarity search, based on
ngerprint-based 2D similarity and Gaussian-shape overlay-based 3D similarity
methods [59–61]. The 2D similarity between molecules is calculated by using
5
in Figure 2.1). This section
1
in Figure 2.4). The

46 2 PubChem: A Large-Scale Public Chemical Database for Drug Discovery
https://t.me/medicina_free
Figure 2.4 Search result page for a text query (ascorbic acid as an example) (https://
pubchem.ncbi.nlm.nih.gov/#query=ascorbic%20acid). The best hit (
the top, and the hits from each data collection can be accessed by clicking the
corresponding tab (
additional tasks can be done through the buttons on the right column of the page (
the PubChem subgraph ngerprints in conjunction with the Tanimoto equation
[62–64]:
Tanimoto =
where NAand NBare the respective counts of ngerprint bits set in molecules A
and B, and NABis the count of bits set in common. On the other hand, 3D molecular
similarity is quantied with the shape-Tanimoto (ST) [15, 59, 60], which evaluates
steric shape similarity, and color-Tanimoto (CT) [15], which quanties functional
group similarity. They are dened as:
1
) will be presented at
2
). The search result can be refined using the filters (3), and
N
AB
NA+NB−N
AB
4
).
(2.1)

2.3 Tools and Services 47
https://t.me/medicina_free
Figure 2.5 Search result page for a chemical structure query (the SMILES string for
ascorbic acid as an example) (https://pubchem.ncbi.nlm.nih.gov/#query=C([C@@H]
([C@@H]1C(=C(C(=O)O1)O)O)O)O). PubChem performs multiple types of structure searches
against the Compound collection, and the results for each search type can be accessed
through the corresponding tab (
can be adjusted using the Settings button (
3
filters (
), and additional tasks can be done through the buttons on the right column of the
4
page (
).
V
ST =
VAA+VBB−V
CT =
∑
AB
f
V
+
f
AA
1
). The parameters and options used for structure searches
AB
∑
f
V
f
AB
∑
f
V
−
f
BB
2
). The search result can be refined using the
∑
f
V
f
AB
(2.2)
(2.3)
where VAAand VBBare the self-overlap volumes of molecules A and B, respectively, and the VABis the overlap volume between them. In Eq (2.3), the index f
indicates any of six functional group types (i.e. hydrogen-bond donors and acceptors, cations, anions, hydrophobes, and rings), represented by ctitious “feature”
or “color” atoms. V
f
AA
and V
f
are the self-overlap volumes of A and B for feature
BB

48 2 PubChem: A Large-Scale Public Chemical Database for Drug Discovery
https://t.me/medicina_free
atom type f , respectively, and V
f
is the overlap volume of molecules AandBfor
AB
feature atom type f .
The ST and CT scores can be combined to create a Combo-Tanimoto (ComboT)
score, which simultaneously considers both steric shape similarity and functional
group similarity:
ComboT = ST + CT (2.4)
Because both ST and CT scores range from 0 to 1, the ComboT score ranges
from 0 to 2 (without normalization). The ST, CT, and ComboT scores between
molecules can be evaluated in two dierent molecular superpositions: the ST- or
shape-optimized superposition and the CT- or feature-optimized superposition. In
the shape-optimization, the superposition of two molecules is optimized to have a
maximum ST score. In the feature-optimization, both shapes and features of the
molecules are simultaneously considered to nd the best superposition.
By default, 2D similarity search returns compounds whose Tanimoto score relative
to the query molecule is equal to or greater than 0.90. For 3D similarity searches,
compounds with ST ≥ 0.80 and CT ≥ 0.50 are returned. More detailed information
on 2D and 3D similarity searches is given in our previous paper [15].
2.3.2 Summary Pages
The Summary page for a given PubChem record displays all information available for that record. For the records in the Substance and BioAssay collections,
which are archival in nature, their Summary page shows the current version of
depositor-provided data by default. An older version of data can be displayed by
selecting the desired version from the dropdown menu (
1
in Figure 2.6). For the
records in the other data collections, which serve as knowledgebases, the Summary
page contains not only relevant depositor-provided data but also annotation
data collected by the PubChem crew from external authoritative sources. These
annotations are regularly updated to provide up-to-date information.
The Summary page for a given record has links to other related records (in the
same or dierent collections), providing users with quick access to information
about them. In addition, the annotation data are presented with the data sources,
allowing users to go to the original data source to check the context of the data and
obtain additional information. Users can quickly access the desired information
2
by using the Table of Contents available in the right column (
in Figure 2.6). The
data presented on the Summary page are downloadable using the Download button
available at the top of the right column (
3
in Figure 2.6). It is also possible to down-
load the data presented under an individual (sub)section of the Summary page.
Because each (sub)section of the Summary page is widgetized, it can be embedded
within the user’s web page. More information on PubChem Widgets is available in
its help document, available at: https://pubchem.ncbi.nlm.nih.gov/docs/widgets.

2.3 Tools and Services 49
https://t.me/medicina_free
Figure 2.6 Summary page for SID 46505070 (https://pubchem.ncbi.nlm.nih.gov/substance/46505070). The dropdown menu (
substance record. The Table of Contents on the right column (
Summary page. The data presented on this page can be downloaded by using the
Download button (
3
) above the Table of Contents.
1
) allows users to view the older version of this
2
) helps navigate the
2.3.3 Literature Knowledge Panel
PubChem contains a great deal of information on scientic articles and patent documents that mention chemicals and their bioactivity data [65]. Users often desire to
explore these articles to learn about the relationships among chemicals, genes, proteins, and diseases, which is not trivial given the size and scope of PubChem data. To
assist users in this task, the Literature Knowledge Panels [66] are presented on the
Summary page of a compound, gene, or protein. The Literature Knowledge Panels
for a given entity (i.e. a chemical, gene, or protein) display a few of its most relevant, nonredundant “neighbors,” which are dened as other entities co-mentioned
in scientic articles (e.g. chemicals, genes, proteins, and diseases). The panels also
provide a sample of PubMed records co-mentioning the entity and its neighbors.
Соседние файлы в папке Библиотека им академика М.И. Перельмана
