Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5938_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
15.09.2026
Размер:
11 Мб
Скачать
☆
300 10 Open Access Databases – An Industrial View
https://t.me/medicina_free
Table 10.1 Available databases for virtual screening.
Database Number of Compounds URL
Asinex 91,473 http://www.asinex.com/
BindingDB 520,000 http://bindingdb.org
ChemBridge 1.3 million https://www.chembridge.com/
COCONUT 407,270
PubChem 11 million https://pubchem.ncbi.nlm.nih.gov/
Zinc15 230 million
a) Natural products. b) Purchasable compounds.
a)
b)
https://coconut.naturalproducts.net/
https://zinc15.docking.org/
SciFinder (https://www.cas.org/solutions/cas-scinder-discovery-platform/cas­scinder) is a database developed by the CAS in which not only chemical but also biological information can be searched [5, 6].
Reaxys and SciFinder are established for data analysis in academic and industrial research because of the amount of data and the clear search and lter options as well as the export of search results in form of Excel lists and chemical structure lists as sdf-les (structure data le).
Academic and industrial research dier fundamentally in their focus. Research at universities is free, deals primarily with basic research, and is nanced by state funding or industrial cooperation. Industrial research, e.g. materials science, phar­maceutical research, and others, have the goal of developing new materials, drugs or dosage forms and bringing them to the market as commercially viable products and being renanced through a life cycle process.
In the pharmaceutical industry, the development of a new active ingredient involves an extensive, time-consuming, and costly process. Not only the indication but also the selection of a possible chemical or biological agent and its formulation must be carefully evaluated.
The development of a possible new active ingredient is divided into several devel­opment phases (Scheme 10.1):
Much information on the individual development steps can be researched in com­mercial and publicly accessible databases and collections.
At the beginning of a project in pharmaceutical research, there is an assessment of both the indication and the possible target. In this Target Assessment (TA), a team of chemists, biologists, biochemists, and pharmacologists is formed, which, based on the results of research, decides whether it makes sense to deal with a target or an indication. A large number of scientic sources are available for obtaining this information; in addition to journals and patents, databases are the most important resource. The topicality of the information sought is of great relevance in indus­try because of its nancial interests and requires both scientic and economically reliable data, which is a big dierence from university research. In addition, strate­gic aspects such as contract research organizations (CROs), contract development,
10.1 Academic vs. Industrial Research 301
https://t.me/medicina_free
Target
identication
Hit identication
Preclinical
Target validation HTS
Hit to lead
Lead optimization
Clinical Market access
Scheme 10.1 Overview of the R&D process.
production, manufacturing organizations (CDMOs), in- or out-licensing, outsourc­ing, oshoring, company takeovers, etc.) must be assessed.
Obtaining this necessary and up-to-date information Competitive Intelligence, (CI) is time-consuming and costly and can only be provided by the industrial side with great eort. The alternative is represented by commercial providers who search for all the necessary information from all available sources (Internet, patent services, company websites, analyst websites, regulatory institutions, authorities, etc.), com­pile them and save them in the form of Excel les or SD les. Make les available for analysis. The best-known and established providers are:
● Adis Insight – https://adisinsight.springer.com/
● Citeline (formerly PharmaProjects) – https://pharmaintelligence.informa.com/
● Clarivate (formerly Cortellis) – https://clarivate.com/cortellis/
● Evaluate – https://www.evaluate.com/
● GlobalData – https://www.globaldata.com/
● Integrity (Clarivate Analytics) – https://integrity.clarivate.com/integrity/xmlxsl
The information provided relates to a variety of aspects and is relevant for deciding how to proceed:
● Patent status
● Drugs and biologics
● Molecular interactions
● Pharmacology data points
● Discovery and preclinical
● Safety and pharmacovigilance
● Metabolism, pharmacokinetic (PK), and toxicology
● Competitive intelligence (CI)
● Drug reports
● Meeting reports
● Company proles
302 10 Open Access Databases – An Industrial View
https://t.me/medicina_free
● Portfolio and licensing
● Alliances and in-licensing
● Financial data
● Deals
● Benchmarking
● Industry news
● Drug pipeline
● Clinical trials
● Drug approval
● Regulatory aspects
● Generics, biosimilars
● Alerts
The commercial providers make this data and information available and can be
downloaded as reports for any desired search term.
The problem for small and medium-sized biopharmaceutical or pharmaceutical companies is the limited nancial possibilities to access this information. The estab­lishment of databases at universities and research institutions was promoted over several years through private- and state-nanced projects in order to generate and compile general and specic information and make it available for research and development.
As already mentioned, the TA is the starting point for a new project in which chemists and biologists collect a large amount of information in order to create a basis for decision-making.
It is traditionally the task of chemists to create an overview of the patent situation of biological targets, substances, or pharmaceutical dosage forms, the so-called CI. The national and international patent oces are state-nanced organizations that decide on the granting of patents after an examination procedure. All information on submitted invention disclosures, processing status, and patent granting is freely accessible on the websites of the patent oces.
● Deutsches Patent- und Markenamt (DPMA) – https://www.dpma.de/
● European Patent Oce (EPO) – https://www.epo.org/
● United States Patent and Trademark Oce (USPTO) – https://www.uspto.gov/
● Japanese Patent Oce (JPO)
● China National Intellectual Property Administration (CNIPA) – https://english
.cnipa.gov.cn/
● Swiss Federal Institute of Intellectual Property (IGE-IPI)– https://www.ige.ch/en/
It is possible to research a large amount of information in the patent databases using search masks and to download patents and the associated information in pdf format for further evaluation.
The EPO website, for example, oers access to more than 130 million patent documents, an example is the advanced search via https://worldwide.espacenet .com/advancedSearch?locale=en_EP (Scheme 10.2). A search on the websites of the regional patent oces is similar.
Enter keywords
https://t.me/medicina_free
Title:
10.1 Academic vs. Industrial Research 303
plastic and bicycle
Title or abstract:
Enter numbers with or without country code
Publication number:
Application number:
Priority number:
Enter one or more dates or date ranges
Publication date:
Enter name of one or more persons/organisations
Applicant(s):
Inventor(s):
hair
WO2008014520
DE201310112935
WO1995US15925
2014-12-31 or 20141231
Institut Pasteur
Smith
Enter one or more classication symbols
CPC
IPC
Scheme 10.2 EPO search mask for an advanced search.
The disadvantage is that many patents are published in their national languages, which is a problem with Asian patents in particular. Most patents available as pdf les are scanned image les, and searching in these les is only possible after conver­sion to readable formats; the alternative is a paid translation. Google Patents (https:// patents.google.com/ (accessed 20 April 2022)) oers a free solution whereby the patents can be downloaded not only in their national language but also in English translation as a pdf le, which gives access to the information contained in foreign patents.
F03G7/10
H03M1/12
304 10 Open Access Databases – An Industrial View
https://t.me/medicina_free
The search for a new drug begins with the identication of a suitable target such as proteins, signaling pathways, genes, and nucleic acid sequences that are associated with a disease:
● G-Protein-coupled receptors (GPCRs)
● Ion channels
● Nuclear receptors
● Enzymes
● Transporters
● DNA
● RNA
In order to study the interaction of a possible active substance with a receptor, it is necessary to nd the binding site of the ligand on a protein. Knowledge of the three-dimensional (3D) structure of the biological macromolecule is helpful for understanding the function of a protein. The Protein Data Bank (PDB), founded in 1971, is the largest structural database for proteins, DNA, and RNA (https:// www.rcsb.org/ (accessed 24 April 2022)). The structures are determined by X-ray structure analysis and NMR spectroscopy and are freely available. The following sequences and structures of biological macromolecules can be searched [7].
The following sequences and structures of biological macromolecules can be searched:
● 189,735 protein structures
● 56,800 structures of human sequences
● 14,225 nucleic acid containing structures
Internationalization took place in 2003 with the establishment of the Worldwide Protein Data Bank (wwPDB) (http://www.wwpdb.org/ (accessed 24 April 2022) by Protein Data Bank in Europe (PDBe) (https://www.ebi.ac.uk/pdbe/ (accessed 24 April 2022)), Protein Data Bank Japan (PDBj) (https://pdbj.org/ (accessed 24 April 2022)), and Biological Magnetic Resonance Bank (BMRB) (https://bmrb.io/ (accessed 24 April 2022).
Another protein database with information on peptide sequences, protein sequences, and functions can be found in UNI-Prot (Universal Protein Resource, https://beta.uniprot.org/ (accessed 24 April 2022)) [8].
An analog database for 3D structures of nucleic acids is the Nucleic Acid Database (NDB), founded in 1992, in which sequences, structures, and functions can be searched freely (http://ndbserver.rutgers.edu/ (accessed 24 April 2022)) [9].
The development of a new database with additional information on sequences, functions, and interactions of nucleic acids was described in 2018 through a col­laboration between Rutgers University NDB and Bowling Green State University (RNAhub services). The NDB is to be replaced under the name Nucleic Acid Knowl­edge Base (NAKB) [10].
Another source for information on nucleosides and nucleotides is the DNA Data Bank of Japan (DDBJ, https://www.ddbj.nig.ac.jp (accessed 24 April 2022)) [11].
10.1 Academic vs. Industrial Research 305
https://t.me/medicina_free
A European archive is The European Nucleotide Archive (ENA, https://www.ebi .ac.uk/ena/ (accessed 24 April 2022)) in which data can be stored and information can be searched for [12].
An overview of all existing targets, signaling pathways and ligands can be found in the DrugBank (https://go.drugbank.com/ (accessed 24 April 2022)) [13].
Another database on binding anities, ligand–target interactions, and pathways is the Binding Database (bindingDB, http://bindingdb.org/bind/index.jsp (accessed 24 April 2022)). BindingDB contains 249,5891 binding data for 8813 receptors and 10,711,154 small molecules [14].
Not only the structure of the receptors is relevant for development, but the struc­ture of the ligands must also be considered. For drugs with chiral centers, knowledge of the exact structure of stereoisomers (enantiomers, diastereomers, and racemates) is extremely important for biological function. Biologically active racemates are com­posed of two enantiomers that can have dierent biological eects. The eutomer represents the active form, and the distomer represents the less active form. The absolute conguration is determined by X-ray structure analysis, and the results of these investigations are publicly available.
Established in 1965, the Cambridge Structural Database (CSD) contains over one million 3D structures of small organic and organometallic molecules deter­mined by X-ray diraction or neutron diraction (https://www.ccdc.cam.ac.uk/ solutions/csd-core/components/csd/ (accessed 24 April 2022)) [15]. A total of 50,000 new structures are added to CSD every year.
Another database for organic, inorganic, metal–organic compounds, and min­erals is the Crystallography Open Database (COD) in which 487,565 structures of molecules are available (http://www.crystallography.net/cod/ (Accessed 24 April
2022)) [16].
After a successful target validation, it is necessary to establish an appropriate assay in order to test a large number of compounds in a high-throughput process (high-throughput screening, HTS) with the aim of identifying biologically active substances (hit-nding).
The substances that are screened in the substance libraries in the HTS consist of compounds synthesized in-house, external syntheses from cooperation with CROs, and substances purchased from commercial suppliers as listed:
● Asinex – https://www.asinex.com/screening-libraries-(all-libraries)
● Charles River – https://www.criver.com/products-services/discovery-services/
screening-and-proling-assays/screening-libraries/compound-screening-
libraries?region=3696
● ChemBridge – https://www.chembridge.com/
● Enamine – https://enamine.net/compound-libraries
● Evotec – https://www.evotec.com/en/execute/drug-discovery-services/hit-
identication
● I.F. Labs – https://iab.com/
● Maybridge – https://www.thermosher.com/de/de/home/industrial/pharma-
biopharma/drug-discovery-development/screening-compounds-libraries-hit-
identication.html
306 10 Open Access Databases – An Industrial View
https://t.me/medicina_free
● SoftFocus® Libraries – https://www.criver.com/products-services/discovery-
services/screening-and-proling-assays/screening-libraries/compound­screening-libraries/softfocus-subscription-libraries?region=3696
There is a specic assay for each target to determine the inhibitory or activating
properties of a substance.
The ChEMBL or ChEMBLdb database (https://www.ebi.ac.uk/chembl/ (accessed April 25 2022) provides an overview of the known and available assays. ChEMBL is a chemical database of biologically active substances with drug-like properties. Cur­rently, 299,151 assays can be accessed and downloaded in the form of report cards, including references and patents [17].
The world’s largest database of information on chemical compounds and their physical and biological properties, safety data, and toxicological data is PubChem (https://pubchem.ncbi.nlm.nih.gov/ (accessed 25 April 2022)) [18]. There are 111 million compounds, 280 million substances, and 295 million bioactivities available. Biological and toxicological data can be searched at PubChem BioAssays (https://pubchemdocs.ncbi.nlm.nih.gov/bioassays (accessed 25 April 2022)) [19].
Pharos (https://pharos.nih.gov/ (accessed 25 April 2022)) is a National Institutes of Health (NIH) sponsored Knowledge Management Center (KMC) database for Illuminating the Druggable Genome (IDG). It includes 20,412 targets, information on 13,704 diseases, and 339,220 ligands [20].
Natural products are a reliable source as lead structures for the development of new drugs. COlleCtion of Open Natural ProdUcTs (COCONUT), https://coconut .naturalproducts.net/ (accessed 25 April 2022)) is a freely accessible database on nat­ural products [21]. The database lists 407,270 searchable natural substances, which are also available for download as sd les.
The results of an HTS run are the basis for further procedure in the projects. A precise analysis of the found active molecules (hits) such as structure, compound class, and patent status is important in order to develop lead structures. The aim of lead structure optimization is to improve the pharmacological, pharmacokinetic, and toxicological properties in order to avoid the risk of possible side eects and to improve bioavailability.
The most common reason for the occurrence of side eects is the insucient drug metabolism, the pharmacokinetic prole (DMPK), and the formation of reac­tive metabolites [22]. When creating a DMPK prole, absorption, the stability and structure of formed metabolites, and inhibition or activation of cytochrome peroxi­dase P450 (CYP) are taken into account [23]. Furthermore, the metabolic enzymes sulfotransferases (SULTs) and UDP-glucuronosyltransferase (UGT) play an impor­tant role in the in vivo degradation of drugs. Data and information on the enzymes, metabolites involved, and their biological and toxicological data can be found in sev­eral Open-Access databases.
The Human Metabolome Database (HMDB) (https://hmdb.ca/ (accessed 25 April
2022)) was established by the Human Metabolome Project and is funded by Genome Canada [24]. HMBD contains 220,945 entries for chemical, biological, biochemi­cal, and molecular biology data and also 8610 enzyme and transporter sequences.
10.1 Academic vs. Industrial Research 307
https://t.me/medicina_free
In addition to chemical structures, analytical data such as NMR and MS are also available. There is a link to numerous other databases:
● KEGG: Kyoto Encyclopedia of Genes and Genomes https://www.genome.jp/
kegg/
● PubChem – https://pubchem.ncbi.nlm.nih.gov/
● MetaCyc Metabolic Pathway Database – https://metacyc.org/
● ChEBI: Chemical Entities of Biological Interest – https://www.ebi.ac.uk/chebi/
● PDB: Protein Data Bank – https://www.rcsb.org/
● UniProt – https://www.ebi.ac.uk/uniprot/index
● GenBank – www.ncbi.nlm.nih.gov
● DrugBank – https://go.drugbank.com/
● T3DB: Toxin and Toxin Target Database – http://www.t3db.ca/
● SMPDB: Small Molecule Pathway Database – https://www.smpdb.ca/
● FooDB – https://foodb.ca/
The MetaCyc Metabolic Pathway Database (https://metacyc.org/ (accessed 26 April 2022)) contains a collection of 13,698 enzymes, 3006 metabolic pathways and their primary and secondary metabolites from 3295 dierent organisms [25].
Die Pathway Datenbank HumanCyc (https://humancyc.org/ (accessed 26 April
2022)) ist ein Archiv zu humanen metabolischen Signalwegen, menschlichen Metaboliten und dem menschlichen Genom [26].
The KEGG Pathway Database (Kyoto Encyclopedia of Genes and Genomes (accessed 26 April 2022)) is a collection of databases on biological pathways, drugs, genomes, diseases, and chemicals [27].
SMPDB (The Small Molecule Pathway Database, https://www.smpdb.ca/ (accessed 27 April 2022)) is an interactive database containing 49,827 signaling pathways, 55,734 substances, and 1576 proteins [28]. It contains more information on signaling pathways that are not searchable in other signaling pathway databases.
Additional information on ADME-Tox [29] are in the relevant databases ChEMBL (https://www.ebi.ac.uk/chembl/ (accessed 26 April 2022)), PubChem (https://pubchem.ncbi.nlm.nih.gov/ (accessed 26 April 2022)), and DrugBank (https://go.drugbank.com/ (accessed 26 April 2022)) [30].
Further information and links to signaling pathway databases can be found at Pathway (https://pathbank.org/others#metabolic (accessed 26 April 2022)).
When developing an active ingredient, the toxicological properties must also be considered. A substance can have toxic eects, but there is also the possibility that toxic metabolites are formed [31, 32].
TOXNET (TOXicology Data NETwork, https://toxnet.nlm.nih.gov/ (accessed 26 April 2022)) is a collection of databases containing information on active ingredi­ents, chemicals, diseases, environmental data, safety, poisons, and regulations. The database is operated by the Toxicology and Environmental Health Information Pro­gram (TEHIP). TOXNET consists of several databases:
● CCRIS – Chemical Carcinogenesis Research Information System
● CPDB – Carcinogenic Potency Database
● DART® – Developmental and Reproductive Toxicology Database
308 10 Open Access Databases – An Industrial View
https://t.me/medicina_free
● CTD – Comparative Toxicogenomics Database
● GENE-TOX – Genetic Toxicology
● HSDB® – Hazardous Substances Data Bank
● Haz-Map®
● Household Products Database
● IRIS – Integrated Risk Information System
● ITER – International Toxicity Estimates for Risk
● LactMed® – Drugs and Lactation
● TRI – Toxics Release Inventory
● TOXMAP®
● TOXLINE®
During a research and development project, it is not only important to follow the preclinical advances, but also the clinical development. After a potential drug candidate clears the hurdle into the clinic, clinical trials take place to test ecacy, improved ecacy over known therapies, side eects, and safety in humans. The clin­ical studies are carried out by pharmaceutical companies and commissioned study centers [33].
The US National Library of Medicine tracks and documents planned, ongoing, and completed trials that receive public or private funding (https://clinicaltrials.gov/ (accessed 26 June 2022)). To date, 419,313 studies have been documented in 220 countries. On the homepage, it is possible to perform a simple or an advanced search for tested substance, indication, therapy, and status of the study (Schemes 10.3 and 10.4).
A European database variant, the EU Clinical Trials Register (https://www .clinicaltrialsregister.eu/ (accessed 26 April 2022)) is also freely accessible. The EU Clinical Trials Register oers access to 42,312 clinical trials with a EudraCT protocol.
Scheme 10.3 Search mask for a simple search on https://clinicaltrials.gov/
10.1 Academic vs. Industrial Research 309
https://t.me/medicina_free
Scheme 10.4 Search mask for an advanced search pm https://clinicaltrials.gov/ct2/ search/advanced?cond=&term=&cntry=&state=&city=&dist=.
EudraCT (European Union Drug Regulating Authorities Clinical Trials Database, https://eudract.ema.europa.eu/ (accessed 30 April 2022)) is the European database where all clinical trials on medicinal products authorized in the European Union are listed. This database contains information on ongoing clinical trials provided by the trial sponsors.
An overview of research activities and the development status of candidate sub­stances from the competition is possible with a simple internet search; search terms “company name” and “pipeline” deliver current results.
Additional information on all phases and issues of drug development can be obtained from regulatory bodies, such as the European Medicines Agency (EMA), https://www.ema.europa.eu/en (accessed 26 April 2022)), located in Amsterdam (Netherlands) since 2019 (formerly in London (United Kingdom)) and the U.S. Food and Drug Administration (FDA), https://www.fda.gov/ (accessed 26 April
2022)), located in Silver Springs, Maryland (United States). Information on the following topics is available on both websites:
● Submissions
● Registration
● Recent drug approvals
● Manufacturing
● Medication Guides
● Drug applications
● Drug compounding
● Drug safety communications
● Shortages
● Warning letters
● Recalls