Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5942_Библиотеки_им_академика_М_И_Перельмана
.pdf
10
https://t.me/medicina_free
Open Access Databases – An Industrial View
Michael Przewosny
Borngasse 43, D-52064 Aachen, Germany
10.1 Academic vs. Industrial Research
Since the spread of the Internet, extensive changes have taken place in many areas,
both professionally and private. Modernization took place in many areas of the economy, new branches of the economy emerged, and communication behavior and the
use of the media changed.
This also led to enormous changes in technical and scientic elds. Until the
mid-1990s, this accumulated knowledge was only available in printed form, mostly
freely accessible in university libraries. Digitization was pushed by institutions and
publishers, not only for commercial purposes but also to create worldwide access to
data and information. Not only are journals and books now available online, but the
inclusion of this content in searchable databases has also been accomplished.
Before these online databases were established, the life of scientists was characterized by time-consuming research in bound data collections. For chemists who were
preparing a research project or a doctoral thesis, it meant disappearing into libraries
for several days to evaluate the current state of knowledge and dening this as the
starting point for their scientic work. The best-known data collections were the
Gmelin for inorganic chemistry and the Beilstein for organic chemistry, in which
structured research was possible [1, 2]. Chemical Abstracts (CA) was also available
across disciplines, provided by the Chemical Abstracts Service (CAS), a subdivision
of the American Chemical Society (ACS) established in 1907. The aim was to bring
together and index all chemistry-related information (journals, books, patents, dissertations, congresses, etc.) in order to make them available [3, 4].
The result of digitization is the commercial online databases Reaxys and SciFinder.
Reaxys (https://www.elsevier.com/solutions/reaxys) is the successor to Crossre,
which provided access to the Beilstein, Gmelin, and Patent Chemistry databases
until 2010. The Windows-based version since then has access to the three databases
and allows all chemistry-related searches for structures, substructures, reactions,
and synthesis planning based on journals and patents. Provider is the publishing
house Elsevier (Table 10.1).
299
Open Access Databases and Datasets for Drug Discovery, First Edition.
Edited by Antoine Daina, Michael Przewosny, and Vincent Zoete.
© 2024 WILEY-VCH GmbH. Published 2024 by WILEY-VCH GmbH.

300 10 Open Access Databases – An Industrial View
https://t.me/medicina_free
Table 10.1 Available databases for virtual screening.
Database Number of Compounds URL
Asinex 91,473 http://www.asinex.com/
BindingDB 520,000 http://bindingdb.org
ChemBridge 1.3million https://www.chembridge.com/
COCONUT 407,270
PubChem 11 million https://pubchem.ncbi.nlm.nih.gov/
Zinc15 230 million
a) Natural products.
b) Purchasable compounds.
a)
b)
https://coconut.naturalproducts.net/
https://zinc15.docking.org/
SciFinder (https://www.cas.org/solutions/cas-scinder-discovery-platform/casscinder) is a database developed by the CAS in which not only chemical but also
biological information can be searched [5, 6].
Reaxys and SciFinder are established for data analysis in academic and industrial
research because of the amount of data and the clear search and lter options as well
as the export of search results in form of Excel lists and chemical structure lists as
sdf-les (structure data le).
Academic and industrial research dier fundamentally in their focus. Research
at universities is free, deals primarily with basic research, and is nanced by state
funding or industrial cooperation. Industrial research, e.g. materials science, pharmaceutical research, and others, have the goal of developing new materials, drugs
or dosage forms and bringing them to the market as commercially viable products
and being renanced through a life cycle process.
In the pharmaceutical industry, the development of a new active ingredient
involves an extensive, time-consuming, and costly process. Not only the indication
but also the selection of a possible chemical or biological agent and its formulation
must be carefully evaluated.
The development of a possible new active ingredient is divided into several development phases (Scheme 10.1):
Much information on the individual development steps can be researched in commercial and publicly accessible databases and collections.
At the beginning of a project in pharmaceutical research, there is an assessment of
both the indication and the possible target. In this Target Assessment (TA), a team
of chemists, biologists, biochemists, and pharmacologists is formed, which, based
on the results of research, decides whether it makes sense to deal with a target or
an indication. A large number of scientic sources are available for obtaining this
information; in addition to journals and patents, databases are the most important
resource. The topicality of the information sought is of great relevance in industry because of its nancial interests and requires both scientic and economically
reliable data, which is a big dierence from university research. In addition, strategic aspects such as contract research organizations (CROs), contract development,

10.1 Academic vs. Industrial Research 301
https://t.me/medicina_free
Target
identication
Hit identication
Preclinical
Target validation HTS
Hit to lead
Lead optimization
Clinical Market access
Scheme 10.1 Overview of the R&D process.
production, manufacturing organizations (CDMOs), in- or out-licensing, outsourcing, oshoring, company takeovers, etc.) must be assessed.
Obtaining this necessary and up-to-date information Competitive Intelligence,
(CI) is time-consuming and costly and can only be provided by the industrial side
with great eort. The alternative is represented by commercial providers who search
for all the necessary information from all available sources (Internet, patent services,
company websites, analyst websites, regulatory institutions, authorities, etc.), compile them and save them in the form of Excel les or SD les. Make les available
for analysis. The best-known and established providers are:
● Adis Insight – https://adisinsight.springer.com/
● Citeline (formerly PharmaProjects) – https://pharmaintelligence.informa.com/
● Clarivate (formerly Cortellis) – https://clarivate.com/cortellis/
● Evaluate – https://www.evaluate.com/
● GlobalData – https://www.globaldata.com/
● Integrity (Clarivate Analytics) – https://integrity.clarivate.com/integrity/xmlxsl
The information provided relates to a variety of aspects and is relevant for deciding
how to proceed:
● Patent status
● Drugs and biologics
● Molecular interactions
● Pharmacology data points
● Discovery and preclinical
● Safety and pharmacovigilance
● Metabolism, pharmacokinetic (PK), and toxicology
● Competitive intelligence (CI)
● Drug reports
● Meeting reports
● Company proles

302 10 Open Access Databases – An Industrial View
https://t.me/medicina_free
● Portfolio and licensing
● Alliances and in-licensing
● Financial data
● Deals
● Benchmarking
● Industry news
● Drug pipeline
● Clinical trials
● Drug approval
● Regulatory aspects
● Generics, biosimilars
● Alerts
The commercial providers make this data and information available and can be
downloaded as reports for any desired search term.
The problem for small and medium-sized biopharmaceutical or pharmaceutical
companies is the limited nancial possibilities to access this information. The establishment of databases at universities and research institutions was promoted over
several years through private- and state-nanced projects in order to generate and
compile general and specic information and make it available for research and
development.
As already mentioned, the TA is the starting point for a new project in which
chemists and biologists collect a large amount of information in order to create a
basis for decision-making.
It is traditionally the task of chemists to create an overview of the patent situation
of biological targets, substances, or pharmaceutical dosage forms, the so-called CI.
The national and international patent oces are state-nanced organizations that
decide on the granting of patents after an examination procedure. All information
on submitted invention disclosures, processing status, and patent granting is freely
accessible on the websites of the patent oces.
● Deutsches Patent- und Markenamt (DPMA) – https://www.dpma.de/
● European Patent Oce (EPO) – https://www.epo.org/
● United States Patent and Trademark Oce (USPTO) – https://www.uspto.gov/
● Japanese Patent Oce (JPO)
● China National Intellectual Property Administration (CNIPA) – https://english
.cnipa.gov.cn/
● Swiss FederalInstitute of Intellectual Property (IGE-IPI)– https://www.ige.ch/en/
It is possible to research a large amount of information in the patent databases
using search masks and to download patents and the associated information in pdf
format for further evaluation.
The EPO website, for example, oers access to more than 130 million patent
documents, an example is the advanced search via https://worldwide.espacenet
.com/advancedSearch?locale=en_EP (Scheme 10.2). A search on the websites of
the regional patent oces is similar.

Enter keywords
https://t.me/medicina_free
Title:
10.1 Academic vs. Industrial Research 303
plastic and bicycle
Title or abstract:
Enter numbers with or without country code
Publication number:
Application number:
Priority number:
Enter one or more dates or date ranges
Publication date:
Enter name of one or more persons/organisations
Applicant(s):
Inventor(s):
hair
WO2008014520
DE201310112935
WO1995US15925
2014-12-31 or 20141231
Institut Pasteur
Smith
Enter one or more classication symbols
CPC
IPC
Scheme 10.2 EPO search mask for an advanced search.
The disadvantage is that many patents are published in their national languages,
which is a problem with Asian patents in particular. Most patents available as pdf
les are scanned image les, and searching in these les is only possible after conversion to readable formats; the alternative is a paid translation. Google Patents (https://
patents.google.com/ (accessed 20 April 2022)) oers a free solution whereby the
patents can be downloaded not only in their national language but also in English
translation as a pdf le, which gives access to the information contained in foreign
patents.
F03G7/10
H03M1/12

304 10 Open Access Databases – An Industrial View
https://t.me/medicina_free
The search for a new drug begins with the identication of a suitable target such as
proteins, signaling pathways, genes, and nucleic acid sequences that are associated
with a disease:
● G-Protein-coupled receptors (GPCRs)
● Ion channels
● Nuclear receptors
● Enzymes
● Transporters
● DNA
● RNA
In order to study the interaction of a possible active substance with a receptor, it
is necessary to nd the binding site of the ligand on a protein. Knowledge of the
three-dimensional (3D) structure of the biological macromolecule is helpful for
understanding the function of a protein. The Protein Data Bank (PDB), founded
in 1971, is the largest structural database for proteins, DNA, and RNA (https://
www.rcsb.org/ (accessed 24 April 2022)). The structures are determined by X-ray
structure analysis and NMR spectroscopy and are freely available. The following
sequences and structures of biological macromolecules can be searched [7].
The following sequences and structures of biological macromolecules can be
searched:
● 189,735 protein structures
● 56,800 structures of human sequences
● 14,225 nucleic acid containing structures
Internationalization took place in 2003 with the establishment of the Worldwide
Protein Data Bank (wwPDB) (http://www.wwpdb.org/ (accessed 24 April 2022)
by Protein Data Bank in Europe (PDBe) (https://www.ebi.ac.uk/pdbe/ (accessed
24 April 2022)), Protein Data Bank Japan (PDBj) (https://pdbj.org/ (accessed 24
April 2022)), and Biological Magnetic Resonance Bank (BMRB) (https://bmrb.io/
(accessed 24 April 2022).
Another protein database with information on peptide sequences, protein
sequences, and functions can be found in UNI-Prot (Universal Protein Resource,
https://beta.uniprot.org/ (accessed 24 April 2022)) [8].
An analog database for 3D structures of nucleic acids is the Nucleic Acid Database
(NDB), founded in 1992, in which sequences, structures, and functions can be
searched freely (http://ndbserver.rutgers.edu/ (accessed 24 April 2022)) [9].
The development of a new database with additional information on sequences,
functions, and interactions of nucleic acids was described in 2018 through a collaboration between Rutgers University NDB and Bowling Green State University
(RNAhub services). The NDB is to be replaced under the name Nucleic Acid Knowledge Base (NAKB) [10].
Another source for information on nucleosides and nucleotides is the DNA Data
Bank of Japan (DDBJ, https://www.ddbj.nig.ac.jp (accessed 24 April 2022)) [11].

10.1 Academic vs. Industrial Research 305
https://t.me/medicina_free
A European archive is The European Nucleotide Archive (ENA, https://www.ebi
.ac.uk/ena/ (accessed 24 April 2022)) in which data can be stored and information
can be searched for [12].
An overview of all existing targets, signaling pathways and ligands can be found
in the DrugBank (https://go.drugbank.com/ (accessed 24 April 2022)) [13].
Another database on binding anities, ligand–target interactions, and pathways
is the Binding Database (bindingDB, http://bindingdb.org/bind/index.jsp (accessed
24 April 2022)). BindingDB contains 249,5891 binding data for 8813 receptors and
10,711,154 small molecules [14].
Not only the structure of the receptors is relevant for development, but the structure of the ligands must also be considered. Fordrugs with chiral centers, knowledge
of the exact structure of stereoisomers (enantiomers, diastereomers, and racemates)
is extremely important for biological function. Biologically active racemates are composed of two enantiomers that can have dierent biological eects. The eutomer
represents the active form, and the distomer represents the less active form. The
absolute conguration is determined by X-ray structure analysis, and the results of
these investigations are publicly available.
Established in 1965, the Cambridge Structural Database (CSD) contains over
one million 3D structures of small organic and organometallic molecules determined by X-ray diraction or neutron diraction (https://www.ccdc.cam.ac.uk/
solutions/csd-core/components/csd/ (accessed 24 April 2022)) [15]. A total of
50,000 new structures are added to CSD every year.
Another database for organic, inorganic, metal–organic compounds, and minerals is the Crystallography Open Database (COD) in which 487,565 structures of
molecules are available (http://www.crystallography.net/cod/ (Accessed 24 April
2022)) [16].
After a successful target validation, it is necessary to establish an appropriate
assay in order to test a large number of compounds in a high-throughput process
(high-throughput screening, HTS) with the aim of identifying biologically active
substances (hit-nding).
The substances that are screened in the substance libraries in the HTS consist of
compounds synthesized in-house, external syntheses from cooperation with CROs,
and substances purchased from commercial suppliers as listed:
● Asinex – https://www.asinex.com/screening-libraries-(all-libraries)
● Charles River – https://www.criver.com/products-services/discovery-services/
screening-and-proling-assays/screening-libraries/compound-screening-
libraries?region=3696
● ChemBridge – https://www.chembridge.com/
● Enamine – https://enamine.net/compound-libraries
● Evotec – https://www.evotec.com/en/execute/drug-discovery-services/hit-
identication
● I.F. Labs – https://iab.com/
● Maybridge – https://www.thermosher.com/de/de/home/industrial/pharma-
biopharma/drug-discovery-development/screening-compounds-libraries-hit-
identication.html

306 10 Open Access Databases – An Industrial View
https://t.me/medicina_free
● SoftFocus® Libraries – https://www.criver.com/products-services/discovery-
services/screening-and-proling-assays/screening-libraries/compoundscreening-libraries/softfocus-subscription-libraries?region=3696
There is a specic assay for each target to determine the inhibitory or activating
properties of a substance.
The ChEMBL or ChEMBLdb database (https://www.ebi.ac.uk/chembl/ (accessed
April 25 2022) provides an overview of the known and available assays. ChEMBL is
a chemical database of biologically active substances with drug-like properties. Currently, 299,151 assays can be accessed and downloaded in the form of report cards,
including references and patents [17].
The world’s largest database of information on chemical compounds and their
physical and biological properties, safety data, and toxicological data is PubChem
(https://pubchem.ncbi.nlm.nih.gov/ (accessed 25 April 2022)) [18]. There are
111 million compounds, 280 million substances, and 295 million bioactivities
available. Biological and toxicological data can be searched at PubChem BioAssays
(https://pubchemdocs.ncbi.nlm.nih.gov/bioassays (accessed 25 April 2022)) [19].
Pharos (https://pharos.nih.gov/ (accessed 25 April 2022)) is a National Institutes
of Health (NIH) sponsored Knowledge Management Center (KMC) database for
Illuminating the Druggable Genome (IDG). It includes 20,412 targets, information
on 13,704 diseases, and 339,220 ligands [20].
Natural products are a reliable source as lead structures for the development of
new drugs. COlleCtion of Open Natural ProdUcTs (COCONUT), https://coconut
.naturalproducts.net/ (accessed 25 April 2022)) is a freely accessible database on natural products [21]. The database lists 407,270 searchable natural substances, which
are also available for download as sd les.
The results of an HTS run are the basis for further procedure in the projects. A
precise analysis of the found active molecules (hits) such as structure, compound
class, and patent status is important in order to develop lead structures. The aim
of lead structure optimization is to improve the pharmacological, pharmacokinetic,
and toxicological properties in order to avoid the risk of possible side eects and to
improve bioavailability.
The most common reason for the occurrence of side eects is the insucient
drug metabolism, the pharmacokinetic prole (DMPK), and the formation of reactive metabolites [22]. When creating a DMPK prole, absorption, the stability and
structure of formed metabolites, and inhibition or activation of cytochrome peroxidase P450 (CYP) are taken into account [23]. Furthermore, the metabolic enzymes
sulfotransferases (SULTs) and UDP-glucuronosyltransferase (UGT) play an important role in the in vivo degradation of drugs. Data and information on the enzymes,
metabolites involved, and their biological and toxicological data can be found in several Open-Access databases.
The Human Metabolome Database (HMDB) (https://hmdb.ca/ (accessed 25 April
2022)) was established by the Human Metabolome Project and is funded by Genome
Canada [24]. HMBD contains 220,945 entries for chemical, biological, biochemical, and molecular biology data and also 8610 enzyme and transporter sequences.

10.1 Academic vs. Industrial Research 307
https://t.me/medicina_free
In addition to chemical structures, analytical data such as NMR and MS are also
available. There is a link to numerous other databases:
● KEGG: Kyoto Encyclopedia of Genes and Genomes https://www.genome.jp/
kegg/
● PubChem – https://pubchem.ncbi.nlm.nih.gov/
● MetaCyc Metabolic Pathway Database – https://metacyc.org/
● ChEBI: Chemical Entities of Biological Interest – https://www.ebi.ac.uk/chebi/
● PDB: Protein Data Bank – https://www.rcsb.org/
● UniProt – https://www.ebi.ac.uk/uniprot/index
● GenBank – www.ncbi.nlm.nih.gov
● DrugBank – https://go.drugbank.com/
● T3DB: Toxin and Toxin Target Database – http://www.t3db.ca/
● SMPDB: Small Molecule Pathway Database – https://www.smpdb.ca/
● FooDB – https://foodb.ca/
The MetaCyc Metabolic Pathway Database (https://metacyc.org/ (accessed 26
April 2022)) contains a collection of 13,698 enzymes, 3006 metabolic pathways and
their primary and secondary metabolites from 3295 dierent organisms [25].
Die Pathway Datenbank HumanCyc (https://humancyc.org/ (accessed 26 April
2022)) ist ein Archiv zu humanen metabolischen Signalwegen, menschlichen
Metaboliten und dem menschlichen Genom [26].
The KEGG Pathway Database (Kyoto Encyclopedia of Genes and Genomes
(accessed 26 April 2022)) is a collection of databases on biological pathways, drugs,
genomes, diseases, and chemicals [27].
SMPDB (The Small Molecule Pathway Database, https://www.smpdb.ca/
(accessed 27 April 2022)) is an interactive database containing 49,827 signaling
pathways, 55,734 substances, and 1576 proteins [28]. It contains more information
on signaling pathways that are not searchable in other signaling pathway databases.
Additional information on ADME-Tox [29] are in the relevant databases
ChEMBL (https://www.ebi.ac.uk/chembl/ (accessed 26 April 2022)), PubChem
(https://pubchem.ncbi.nlm.nih.gov/ (accessed 26 April 2022)), and DrugBank
(https://go.drugbank.com/ (accessed 26 April 2022)) [30].
Further information and links to signaling pathway databases can be found at
Pathway (https://pathbank.org/others#metabolic (accessed 26 April 2022)).
When developing an active ingredient, the toxicological properties must also be
considered. A substance can have toxic eects, but there is also the possibility that
toxic metabolites are formed [31, 32].
TOXNET (TOXicology Data NETwork, https://toxnet.nlm.nih.gov/ (accessed 26
April 2022)) is a collection of databases containing information on active ingredients, chemicals, diseases, environmental data, safety, poisons, and regulations. The
database is operated by the Toxicology and Environmental Health Information Program (TEHIP). TOXNET consists of several databases:
● CCRIS – Chemical Carcinogenesis Research Information System
● CPDB – Carcinogenic Potency Database
● DART® – Developmental and Reproductive Toxicology Database

308 10 Open Access Databases – An Industrial View
https://t.me/medicina_free
● CTD – Comparative Toxicogenomics Database
● GENE-TOX – Genetic Toxicology
● HSDB® – Hazardous Substances Data Bank
● Haz-Map®
● Household Products Database
● IRIS – Integrated Risk Information System
● ITER – International Toxicity Estimates for Risk
● LactMed® – Drugs and Lactation
● TRI – Toxics Release Inventory
● TOXMAP®
● TOXLINE®
During a research and development project, it is not only important to follow
the preclinical advances, but also the clinical development. After a potential drug
candidate clears the hurdle into the clinic, clinical trials take place to test ecacy,
improved ecacy over known therapies, side eects, and safety in humans. The clinical studies are carried out by pharmaceutical companies and commissioned study
centers [33].
The US National Library of Medicine tracks and documents planned, ongoing,
and completed trials that receive public or private funding (https://clinicaltrials.gov/
(accessed 26 June 2022)). To date, 419,313 studies have been documented in 220
countries. On the homepage, it is possible to perform a simple or an advanced search
for tested substance, indication, therapy, and status of the study (Schemes 10.3
and 10.4).
A European database variant, the EU Clinical Trials Register (https://www
.clinicaltrialsregister.eu/ (accessed 26 April 2022)) is also freely accessible. The
EU Clinical Trials Register oers access to 42,312 clinical trials with a EudraCT
protocol.
Scheme 10.3 Search mask for a simple search on https://clinicaltrials.gov/
Соседние файлы в папке Библиотека им академика М.И. Перельмана
