Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5364_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
02.09.2026
Размер:
21 Мб
Скачать
2 Molecular Databases 19
DBs, other DBs have emerged with the aim of drug design, such as the Data Repository of Antiviral Peptides and Proteins (DRAVP). This database provides information on the antiviral activity, structure, physicochemical properties, and literature data of the peptides and proteins that make up the repository [34].
The COVID-19 pandemic has also led to the emergence of DBs that have gathered information to optimize research into this new viral disease, including the search for new drugs, such as the Small Molecule Antiviral Compound Collection (SMACC), which includes bioactivity data available in ChEMBL for compounds that have assays for emerging viruses that pose the greatest potential threat to global human health [35].
Furthermore, other databases recently developed to SARS- CoV-2, such as CoV-RDB [36], CORDITE [37], DockCoV2 [38], H2V [39], SARS-CoV-2 3D [40], and ZINCPharmer [41], bring other functionalities: (i) CoV-RDB contains data on the neutralizing susceptibility of SARS-CoV-2 variants to monoclonal anti­bodies, convalescent plasma, and vaccinated plasma; (ii) CORDITE provides drug interactions for SARS-CoV-2 current drug options; (iii) DockCoV2 shows compu­tational representation of molecular docking; (iv) H2V contemplates information how the human body responds to viral infections; (v) SARS-CoV-2 3D provides possible drug targets from the coronavirus proteome; and (vi) the ZINCPharme uses the ZINC database, employs the Pharmer pharmacophore search technology, and also provides tools to construct and rene pharmacophore hypotheses directly from molecular structure.
Arguably, the great challenge of these repositories is the standardization of pro­cedures for the curation of information and compounds, as well as the provision of nancial resources and researchers for their maintenance and sustainability [42]. In this sense, we discussed these matters in the following topics.

2 Databases and Curation

Compound DBs usually are built by universities and research institutes. The devel­opment of a compound DB usually involves major steps (Fig. 2.1): First, searching the literature for the chemical and/or biological information that will constitute the DBs. At this stage, other databases may also be used as a source of chemical and biological information, such as ZINC, PubChem, and HMDB. Examples of chemical information are structure and molecular weight, 2D and 3D structures, SMILES, ClogP, and information on in vitro, in vivo, and ex vivo biological activities. Second, curating the informat ion and compounds, and nally, selecting the management system and creating the DB website.
20 D. Q. de Azevedo et al.
Fig. 2.1 The steps of the development of a molecular database. Step 1search for chemical or biological information in indexed databases: reports the strategies used to nd the information that will make up the DBs. Step 2curation: describes the DB curation processes, automated or manual. Step 3DB management and network visualization: describes different systems to process DBs. Step 4update and maintenance: the last and most important stage is that, in addition to the development of a DB, its maintenance requires various resources, both nancial and human
2.1 First Step: Search for Chemical or Biological
Information in Indexed Databases
The strategy for feeding the compoundsinformation to the DB mainly uses data from indexed DBs. This approach was used in ChEMBL, NuBBE BIOFACQUIM, NPACT, and TCM Database@Taiwan. The information contained in a DB is extensive, as it may comprise 2D and 3D structures of compounds, with or without their biological activity. Antiviral Medicinal Plants and Natural Products DB (avMpNp DB) is a DB developed in Brazil and contains bioactive compounds from biodiversity with antiviral activity. The strategy for building the avMpNp DB consisted rst of an extensive bibliographic search in academic DBs and compound DBs, to systematize information about the bioactive compounds included in the DB [43]. ChEMBL is a DB that was introduced in 2009 as an open-access resource and plays an important role in drug discovery and validation of computational tool s. A large proportion of this bioactivity data in ChEMBL is currently manually extracted from scientic literature [15]. NPASS (Natural Product Activity and Species Source Database) [44], COCUNUT (COlleCtion of Open Natural ProdUcTs) [45], and
DB
,
2 Molecular Databases 21
PHCS (Persian Herbal Constituents Database) [46] are databa ses of natural products that also employ manual information extraction of the literature.
The availability of public chemistr y and bioactivity DBs, along with large-scale data-driven applications, has increased the communitys attention to data curation and integrity issues, such as structure quality, name-to-structure delity, structure– activity mapping, activity data accuracy, assay description sufciency, target assign­ment, author errors, and redundancy. Together, such factors can provide higher condence assertions and therefore more robust applications and models from the available compoundsinformation in DBs.
In particular, the compounds contained in DBs have already led to the develop­ment of drugs in clinical use to treat various diseases and are contributing to the development of compounds in clinical development and basic research. Virtual libraries thus help to increase the success rate of the lead selection process by ensuring the quality, diversity, and consistency of the curated data. The number of public domain compound databases is increasing and the process of building and curating these libraries is critical as the data must be diver se and reliable to enable safe trials. Therefore, the assembly and analysis of compound DBs in terms of their structure and the legitimacy of the structures is crucial [47].
2.2 Second Step: Data Curation
Data curation includes, for example, elimination of salts, adjustment of protonation states, optimization of geometry by energy minimization, and elimination of dupli­cated molecules. This curation process can involve several manual and automated steps and aims to maximize data accessibility and comparability and improve data integrity and ag outliers, ambiguities, and potential errors. Standard protocols are used in manual DB curation processes. Although this step is not easy, it is feasible. It would be advisable for a research group to be responsible for this endeavor, using publicly available tools and scripts or workows available on public repositories such as GitHub, where examples of database curation are freely available [48].
The automated curation process uses platforms with different functionalities. One example is Open Babel, a tool that provides a solution to the proliferation of different
le formats in chemistry. It also contains conformer searching, 2D visualization,ltering, batch conversion, substructure, and similarity searching. For developers, it
can be a programming library for chemical data handling in areas such as organic chemistry, drug design, materials science, and computational chemistry. It is freely available under an open-source license from http://openbabel.org [49].
ChEMBL, for example, is a database that uses both manual and automated curation strategies and implements their validation and standardization proces s using pipelining tools such as Pipeline Pilot [50] or the Konstanz Information Miner (KNIME) analysis platform [51]. These tools also allow for more exibility such as new components that can be added or adapted as needs change. The ChEMBL DB providers routinely include a salt stripping process in their
22 D. Q. de Azevedo et al.
Table 2.2 Some examples of tools of automatized curation: platforms, functionalities, and uses
Platform Advantages
Molecular Operating Environment (MOE)
Konstanz information miner (KNIME)
b
RDKit
Open Babel Open, collaborative project
a
Advantage of KNIME;bAdvantage of RDKit
Integrated computer-aided molecular design platform­small molecules, peptides, biologics Help draw chemical struc­tures and facilitate the storage and interconversion between standard le formats Free and open-source aca­demic version
a
Offers over 300+ connectors to data sources, and integra­tions to all popular machine
a
learning libraries
b
Manipulate molecular struc­tures in Python
a
Freely available
Convert, analyze, or store data from molecular model­ing, chemistry, biochemistry, or related areas
Functionalities for curation (examples) Databases use
Disconnects salts and metals Removes simple components Recalculates states of pro­tonation, determines wedge bonds for bonds from chiral centers Calculates missing chiral parities from existing wedge bond
Removal of characters encoding stereoisomerism in SMILES format (@; \; /) Removal of salts Neutralization of charges Maintenance of compounds containing only the following elements: (H, C, N, O, F, Br, I, Cl, P, S) Creation of compounds in InChI, InChIKey, and canonical SMILES formats from standardized compounds
Removal of salts Adjustment of the proton­ation state of the structures Convert between molecule le formats
NuBBE
DB
BIOFACQUIM
PeruNPDB Super Natural II ChEMBL
SWMD ChemDB
standardization based on a library of pharmaceutically relevant salts [15]. Table 2.2 shows examples of automatized curation, platform, functionalities, and their use.
NuBBE
[31] and BIOFACQUIM [32] use Molecular Operating Environment
DB
(MOE) [52] for automatized curation. This node disconnects salts and metals, removes simple components, recalculates states of protonation, determines wedge bonds for bonds from chiral centers, and calculates missing chiral parities from existing wedge bonds. Using this same software, inorganic compounds can be eliminated, as well as duplicated compounds [53, 54].
Another example of a DB using automated curation strategies is the Seaweed Metabolite Database (SWMD), which comprises compounds, derived from sea­weeds, and it has been curated using the Open Babel software. This strategy allowed the removal of salts and the adjustment of the protonation state of the structures. Duplicate structures were manually removed after detection in SMILES strings using Microsoft Excel 2016, followed by manual inspection [55]. Open Babel is a full-featured open chemical toolbox, designed to translate the many different
2 Molecular Databases 23
representations of chemical data [49]. It allows anyone to search, convert, analyze, or store data from molecular modeling, ch emistry, solid-state materials, biochemistry, or related areas. It provides both ready-to-use programs as well as a complete, extensible programmers toolkit for developing cheminformatics software. In addi­tion, the ChemDB is a small molecule database that also uses Open Babel to convert between molecule le formats [56].
KNIME analysis platform, which includes RDKit, is also used to curate chemical structures in DBs, such as PeruNPDB, the Peruvian Natural Products Database, and Super Natural II [57, 58]. RDKit is an open-source cheminformatic s toolkit written in C++ that is also usable from Java or Python. It includes a collection of standard cheminformatics functionality for molecules, substructure searching, chemical reac­tions, coordinate generation (2D or 3D), ngerpr inting, curation, as well as a high­performance database cartridge for working with molecules using the PostgreSQL DB [57 59].
DataWarrior is a multi-functional and interactive chemical data analysis and visualization tool. It provides interactive options for visualizing and curating data, assessing correlations, and extracting knowledge from large datasets [60]. The DataWarrior tool was used to eliminate duplicate structures in the different studies, such as in the identication of anti-schistosomal, anthelmintic, and antileishmanial compounds [61, 62 ].
DBs also use manual curation, for example, ChEBI (Chemical Entities of Bio­logical Interest), which is a manually curated DB and ontology that organizes small molecule knowledge [63]. Last, PSC-db [64] and avMpNp DB [ 43 ] also employ manual curation using internal scripts.
2.3 Third Step: Database Management and Network
Visualization
Molecular DBs contain a wide variety of data that may be processed by different systems, which, if unrelated, can lead to redundancy and inconsistency for the user. To solve this problem, the data needs to be stored only once on a platform that can be accessed and shared by all the systems involved. This solution results in a more complex software structure that requires the help of a database management system (DBMS) to maintain.
The choice of the DBMS model is fundamental to its development, as it is the basis for structuring a database (data types, relationships, and relevant constraints). The relational model is the most widely used model because of its greater exibility and suitability for design and implementation, and many DB systems today are based on it. These characteristics are due to its structuring of data into relationships. Examples of DBs that are managed using a relational model include Super Natural II [65], ChEBI [63], Viether b [66], ZINC [13], and NPACT [33].
24 D. Q. de Azevedo et al.
Another option for implementing a molecular DB is to use a workow-based management system (WBMS). WBMSs are data management systems that exibly control the execution of a set of tasks, allowing this set of tasks to be modied without changing the system code. A major advantage of using a WBMS is its exibility, changes to the stored information are reected in changes into workows, which makes it much easier to adapt the system and does not require changes in the system code. This is important because different bioactive compounds may have different types of data associated with them, and in a workow-based WBMS, all this heterogeneous data can be easily managed [67].
2.4 Fourth Step: Updating and Maintenance
Usually, the DB function refers to compound repositories. In fact, compound DBs and their chemical datasets are a central part of pharmaceutical companies and private or government research centers. These DBs have been upgraded through the cooperation of chemoinformatics tools and the introduction of the new com­pound. A study that evaluated 52 DBs developed over the last 30 years found that most of them originated in the academic sector, such as in universities and research institutes, where maintenance and upgrading also take place. Private DBs are maintained by industry and thei r data is usually condential [43] (Fig. 2.2).
3 Advantages and Disadvantages of Using Molecular
Databases
Molecular DBs are useful resources in computer-aided drug design and play a central role in many chemoinformatics applications. It is possible to identify poten­tial bioactive compounds with therapeutic activity through several chemoinformatic methodologies. Indeed, molecular DBs can provide access to hundreds, thousands, or even hundreds of thousands of compounds that can be virtually screened to predict which of them have the desired biological activity. Recently, the amount of freely accessible databases has increased, allowing access to an increasing number of compounds [68]. Molecular DBs not only provide access to chemical structures but also contain other useful information for the different drug design stages (Table 2.3).
As illustrated in Table 2.3, molecular databases are useful not only in drug discovery. For instance, SciFinder [78] allows users to check the reported synthesis pathways of a molecule and ChemSpider [24] contains spectroscopic data of the compounds. Other examples are ZINC [13], ChEMBL [15], and ChemSpider [24]. DBs provide information regarding the commercial availability of chemical compounds. Despite the above-mentioned advantages of using molecular DBs during the drug design process, there are still some associated deciencies. The
2 Molecular Databases 25
Fig. 2.2 Example of how to build a DB, CHEMBL. Step 1Search for chemical or biological information in indexed databases: CHEMBL uses manual extraction of data from scientic litera­ture. Step 2Curation: CHEMBL DB uses tools such as KNIME for this purpose. Stage 3DB management and network visualization: Access to the data in ChEMBL, a relational DB, is provided through a user interface, a set of web services, and a range of download formats including XML, JSON, and YAML. Stage 4Updating and maintenance
structural integrity of the molecules is not always assured, and the annotations are not necessarily error-free. For instance, errors in the stereochemistry can be found in DBs, as valence issues and charge imbalances [79] and small structural errors can lead to signicant losses of predictive abilities of quantitative structure–activity relationship (QSAR) models [80]. For more details on QSAR and machine learning predictors, please refer to Chap. 6. This is particularly common for publicly available databases, where it can be challenging to have a dedicated team to curate the chemical content and annotations in the compound DBs. Incomplete or wrong information provided in chemical DBs is not only limited to the chemical structures. Erroneous information regarding the biological activity of a molecule is something else found in chemical databases. For instance, it is not always reported the assay employed to determine the half maximal inhibitory concentration (IC
) or the
50
minimum inhibitory concentration (MIC) [81]. Besides, the assay conditions are not always reported, or the original references are missing [65]. Such incom plete or inaccurate data can lead to less reliable QSAR predictions. Thus, to improve the quality of chemical databases, periodic revisions and error reporting are recommended practices.
26 D. Q. de Azevedo et al.
Table 2.3 Categories into which databases can be divided according to the type of information stored
Database category Content Database
Chemical information
Bioactivity Inhibitor constant (K
Drug Detailed drug data
Natural product Pathways (synthesis and degrada-
Chemical availability
Fragment Physicochemical information
a
Representative examples. Some of the databases have more than one category that is not shown in the table, for example, PubChem (chemical information and chemical availability), ChEMBL (chemical information and drug), ChemSpider (bioactivity and chemical availability), and ZINC (chemical information and bioactivity)
Chemical and crystal structures spectra Reactions and syntheses Thermophysical data
)
i
Dissociation constant (K Half maximal inhibitory concentra­tion (IC Half maximal effective concentration (EC
Comprehensive drug target information
tion) Structures
Available compounds offered by chemical vendors
Binding site preferences
)
50
)
50
)
d
ChemSpider ChEBI Chemical Universe Data­base GDB
PubChem ChEMBL BindingDB ChemBank PDBbind
DrugBank [28]
Universal Natural Product Database MeFSAT Natural Product Atlas
ZINC NCI
FDB-17 Fragment Store PADFrag
a
Reference
[24] [63] [69]
[14] [15] [26] [27] [70]
[71] [72] [73]
[13] [74]
[75] [76] [77]
4 Natural Product Databases for the Search
and Development of New Drugs
Nature is a rich source of bioactive molecules that serve as therapeutic agents. These bioactive molecules are natural products, which can be dened as compounds produced by living beings and can be employed or proposed as therapeutic agents [82]. For instance, of the approved small molecules in the research area of cancer, from 1946 to 1980, 53% of the compounds that became new medicines corresponded to unaltered natural products or natural product derivatives. From 1981 to 2019, 64.9% of the approved small molecules were unaltered natural products or natural products inspired [83]. Moreover, natural products are an abundant source of privileged scaffolds: structures capable of providing useful ligands for more than one receptor [84]. Some examples of privileged scaffolds that come from natural products that are currently used in the design and develop­ment of new drug candidates are the terpenoid, polyketide, phenylpropanoid, and alkaloid structures [85 ]. Regarding natural product DBs, one application is the
2 Molecular Databases 27
Table 2.4 Representative natural product databases
Database
Collection of Open
Natural Products
(COCONUT)
Universal Natural
Product Database
SuperNatural 3.0 449,058 Open access Toxicity
ZINC 80,000 Open access Bioactivities
Dictionary of Natu-
ral Products
SciFinder 300,000 Commercial Reported synthesis routes [78]
Reaxys 200,000 Commercial Bioactivity
TCM@Taiwan 58,000 Open access The largest database of natu-
IMPPAT 10,000 Open access The largest database of natu-
AfroDB 1000 Open access The largest database of natu-
Phyto4Health 3128 Open access Medicinal plants included in
NuBBE
DB
BIOFACQUIM 553 Open access Taxonomic information of
a
Date of search: April 2024
Number of compounds
411,621 Open access Predicted bioactivities [45]
229,000 Open access 3D structures [71]
230,000 Commercial Spectroscopic data [82]
2223 Open access Bioactivity
a
Accessibility Outstanding features References
Vendor information
Vendor information
Toxicity Physicochemical data
ral products from Traditional Chinese Medicine
ral products from traditional medicine in India
ral products from traditional medicine in Africa
the Russian Pharmacopoeia Bioactivities Predicted bioactivities
Predicted spectroscopic data
the producing organism
[65]
[13]
[91]
[92]
[93]
[94]
[95]
[31, 96]
[32]
design of pseudo-natural products, that is, molecules that retain the biological relevance of natural products yet exhibit structures and bioactivities not available in nature or in existing design strategies. Pseudo-natural products may display unexpected bioactivities that differ from the activities of the natural products from which their fragments are derived [8688]. Besides, natural products usually have more structural diversity compared with synthesized small molecules [89].
Natural product DBs can be important tools for computer-aided drug design (CADD), providing access to a large number of diverse chemical structures. These DBs can be divided into commercial and open access (Table 2.4). Between 2000 and 2019, 123 commercial and open-access natural product collections have been published, of which 98 are somewhat accessible, 92 are open access, and only
28 D. Q. de Azevedo et al.
50 contain molecular structures that can be retrieved for a chemoinformatic analysis [90]. For example, the Collection of Open Natural Products (COCONUT) [45] contains more than 411,000 natural product entries collected from 50 open-access natural product DBs. Similarly, the Universal Natural Product Database [71]is another compilation DB with more than 229,000 natural products. It provides 3D structures with stereochemical information and calculated molecular descriptors but is not yet accessible through the link in the original publication. Instead, it is available on another website [81]. Currently, SuperNatural 3 is the bigges t open­access DB, which contains over 449,058 unique natural products and includes information about 2D structures, physicochemical properties, predicted toxicity class, and potential sellers, but it does not yet provide the option to download in bulk.
There are natural product databases comprised of molecules isolated and charac­terized in specic geographical regions. China is one of the regions with more natural product databases published [102 110], because Traditional Chinese Med­icine (TCM) is part of the Chinese public health system. There are two natural product DBs that comprise compounds that are part of the traditional medicine in India (Indian Ayurveda), such as IMPPAT [93] and MedPServer [106]. Regarding African Traditional Medicine, there are different natural product DBs published such as ConMedNP, p-ANAPL library, and others [107111]. There is a database that contains natural products from medicinal Russian plants: Phyto4Health [95]. Latin America is a region that encompasses at least a third of global biodiversity [112]. All the published natural product databases of Latin America and their practical appli­cations in the drug disco very area have been reviewed and discussed elsewher e [113]. Two representative examples of natural product databases from Latin Amer­ica are NuBBE
[31] and BIOFACQUIM [ 32].
DB
5 Functionalization of Databases and Transformation into
a System to Drug Design
The usage of molecular DBs as a drug design tool depends on the capacity of the chemoinformatic software to recognize the molecules. For this purpose, the simpli­ed molecular input line entry system (SMILES) notation is the predominant input notation for the different chemoinformatic software packages [114]. Other notations that can be recognized by the chemoinformatic software overcome some disadvan­tages of the SMILES, such as the International Chemical Identier (InChI) [115] and InChIKey [116]. SMILES arbitrary target specication (SMARTS) notation was developed to specify substructural patterns that allow matching molecules that contain a specied substructural pattern [117]. The different notations used by the chemoinformatic software, as well as their uses, advantages, and disadvantages, have been explained in detail elsewhere [118].