Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5606_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Foreword
- •Acknowledgments
- •Contents
- •1.1 Structure-Based Drug Discovery (SBDD)
- •1.2 Ligand-Based Drug Design (LBDD)
- •1.3 Echoes from the Past, Visions from the Future
- •References
- •1 Introduction
- •2.2 Second Step: Data Curation
- •2.4 Fourth Step: Updating and Maintenance
- •2 Databases and Curation
- •8 Perspectives
- •9 Conclusion
- •References
- •1 Introduction
- •2.1 Making and Matching Protein Models
- •2.2 Simulating Protein Movements
- •2.3 Analyzing Changes in Protein Shape
- •3 Pharmacogenomics in Drug Development
- •4 Case Studies of Genomics-Based Drug Design
- •References
- •1 Historical Background
- •1.1 Timeline
- •2 Methodology Overview
- •2.1 Neural Networks
- •2.1.1 Perceptron
- •2.1.2 Multilayer Neural Networks
- •2.1.3 Types of Neural Networks
- •Feedforward
- •Recurrent Neural Networks
- •LSTM
- •2.2 Deep Learning
- •3 Using Machine Learning
- •3.2 Data Collection
- •3.3 Data Preprocessing
- •3.4 Model Selection
- •3.5 Model Training
- •3.6 Validation
- •3.7 Tuning
- •3.8 Prediction
- •4 Limitations
- •4.1 Bias
- •4.3 Interpretability
- •4.4 Computational Cost
- •4.5 Data Dependency
- •4.6 Robustness
- •5 Applications in Drug Discovery
- •5.2 Lead Discovery
- •5.3 Preclinical and Clinical Development
- •6 Resources and Tools
- •7 Challenges and Perspectives
- •7.1 Future Trends
- •9 Conclusions
- •References
- •1 Historical Background
- •1.1 Applications in Drug Discovery
- •2 Validations and Controls
- •2.1 Internal Validation
- •2.2 External Validation
- •2.3 Relative Cluster Validation
- •3 Challenges and Perspectives
- •4 Conclusions
- •References
- •1 Historical Background
- •2 OECD Principles
- •2.1 A Defined Endpoint
- •2.2 An Unambiguous Algorithm
- •2.5 A Mechanistic Interpretation, if Possible
- •3 Software and Tools
- •4 Validations and Controls
- •4.1 Internal and External Validation
- •4.1.1 Regression Metrics
- •4.2 Applicability Domain
- •4.3 Randomization Tests
- •5 Interpretation
- •6 Practical Advice During QSAR Modeling
- •7 Application
- •8 Challenges and Perspectives
- •References
- •1 Molecular Docking
- •2 Advances in Scoring Functions and Search Algorithms
- •2.2 Critical Characteristics of Search Algorithms
- •2.3 Docking Programs and Scoring Functions
- •3 Calculations Performed During Docking Simulations
- •4 Essential Components for a Good Docking Program
- •5 Limitations of the Docking Technique
- •6 Validation of Docking Results
- •7 Inappropriate Use of Validation Methods in Docking
- •9 Use of Machine Learning in Molecular Docking
- •11 Challenges
- •12 Conclusions
- •References
- •3 System Preparation for MD Simulations
- •3.1 Solvation and Microensemble
- •3.2 Force Fields: General Concept and Relevant Choices
- •3.3 The Concept of Replicas and Timescale
- •4.1.2 Protein Root Mean Square Fluctuation (RMSF)
- •4.1.4 Protein Secondary Structure Analysis
- •4.1.5 Principal component Analysis (PCA)
- •4.1.6 Markov State Modelling
- •4.1.7 Distance Calculations
- •4.1.8 Angle and Plane Calculations
- •4.2.2 Distances and Ligand-Induced Geometry Rearrangements
- •4 Molecular Dynamics Analysis
- •4.1 Protein Perspective
- •4.1.1 Protein Root Mean Square Deviation (RMSD)
- •4.3 Ligand Perspective
- •4.3.1 Ligand Properties
- •4.3.2 Ligand Root Mean Square Deviation
- •4.3.3 Ligand Root Mean Square Fluctuation
- •4.3.4 Angles and Dihedrals
- •5.1 Protein Structure Prediction and Preparation
- •5.2 Molecular Docking
- •6 Concluding Remarks and Outlook
- •Glossary
- •References
- •1 Introduction
- •2.1 MDeNM
- •2.2 Collective Molecular Dynamics (coMD)
- •2.3 ClustENM and ClustENMD
- •3 Ensemble Docking
- •References
- •1 Introduction
- •1.1 Advantages, Disadvantages, Innovations, and Challenges
- •1.2 Recent Advances in Accessible FEP Software Tools
- •1.3 Applications of FEP in Industry and Consortiums
- •2 Expanding the Potential of FEP Calculations
- •2.1 Validating Binding Poses
- •2.2 Dealing with Solvent
- •2.3 FEP and Allostery
- •2.4 FEP and Covalent Ligands
- •2.5 Applications of FEP in Scaffold Hopping
- •2.6 Positional Analogue Scanning
- •2.7 Combinations and Alternative Approaches
- •3 Machine Learning for FEP
- •3.4 Implications for ML in FEP Calculations
- •4 Final Considerations
- •5 First Steps to FEP Simulations
- •References
- •1 Background
- •2 Ultra-Large Screening Libraries and Chemical Spaces
- •3.1 Implications of Dataset Size
- •4 Ligands on the Ultra-Large Scale
- •4.1 Ultra-Large 2D Similarity Searches
- •7 Challenges and Future Perspectives
- •7.1 Hit Triage: An Old Problem on a New Dimension
- •8 Conclusions
- •Appendix
- •References
- •1 Introduction
- •2 Enzymatic Activity Evaluations
- •3 Cytotoxicity Evaluation and Cell Viability
- •4 Antiviral Assays in Experimental Validation
- •6 In Vivo Evaluation of Compounds
- •7 Conclusions
- •References
- •1 Introduction
- •3.1 Data Collection
- •3.2 Data Preprocessing
- •3.4 Model Choice
- •3.5 Model Training
- •3.6 Model Assessment
- •3.7 External Validation
- •3.8 Implementation and Availability
- •3.9 Continuous Update
- •5 Conclusions and Perspectives
- •References
- •1 Experimental Approaches to Obtain Protein Structure
- •1.1 X-Ray Crystallography
- •1.2 Nuclear Magnetic Resonance
- •1.3 Cryo-EM
- •1.4 Hybrid Methods
- •2 Modeling Approaches to Obtain Protein Structure
- •2.1 Homology Modeling
- •2.2 Ab Initio Modeling
- •2.3 New Approaches
- •3 Conformational Diversity of Proteins
- •3.1 Characterization of Protein Conformational States
- •3.2 Experimental Methods to Study Protein Dynamics and Conformations
- •3.4 Molecular Dynamics Simulation
- •3.5 Sampling Strategies
- •4 Remarks and Perspectives
- •References
- •1 Introduction
- •2 Structure-Based Drug Design of HIV Protease Inhibitors
- •2.1 HIV-1 Protease as a Therapeutic Target
- •2.2.1 Saquinavir
- •2.2.2 Indinavir
- •2.3.1 Lopinavir
- •2.3.2 Darunavir
- •6 Conclusions
- •References
- •4 Experimental Methods to Analyze NR Activity
- •4.2 Coregulator-Recruitment
- •5 Concluding Remarks and Outlook
- •References

2 Molecular Databases 19
DBs, other DBs have emerged with the aim of drug design, such as the Data
Repository of Antiviral Peptides and Proteins (DRAVP). This database provides
information on the antiviral activity, structure, physicochemical properties, and
literature data of the peptides and proteins that make up the repository [34].
The COVID-19 pandemic has also led to the emergence of DBs that have
gathered information to optimize research into this new viral disease, including the
search for new drugs, such as the Small Molecule Antiviral Compound Collection
(SMACC), which includes bioactivity data available in ChEMBL for compounds
that have assays for emerging viruses that pose the greatest potential threat to global
human health [35].
Furthermore, other databases recently developed to SARS- CoV-2, such as
CoV-RDB [36], CORDITE [37], DockCoV2 [38], H2V [39], SARS-CoV-2 3D
[40], and ZINCPharmer [41], bring other functionalities: (i) CoV-RDB contains data
on the neutralizing susceptibility of SARS-CoV-2 variants to monoclonal antibodies, convalescent plasma, and vaccinated plasma; (ii) CORDITE provides drug
interactions for SARS-CoV-2 current drug options; (iii) DockCoV2 shows computational representation of molecular docking; (iv) H2V contemplates information
how the human body responds to viral infections; (v) SARS-CoV-2 3D provides
possible drug targets from the coronavirus proteome; and (vi) the ZINCPharme uses
the ZINC database, employs the Pharmer pharmacophore search technology, and
also provides tools to construct and refine pharmacophore hypotheses directly from
molecular structure.
Arguably, the great challenge of these repositories is the standardization of procedures for the curation of information and compounds, as well as the provision of
financial resources and researchers for their maintenance and sustainability [42]. In
this sense, we discussed these matters in the following topics.
2 Databases and Curation
Compound DBs usually are built by universities and research institutes. The development of a compound DB usually involves major steps (Fig. 2.1): First, searching
the literature for the chemical and/or biological information that will constitute the
DBs. At this stage, other databases may also be used as a source of chemical and
biological information, such as ZINC, PubChem, and HMDB. Examples of chemical
information are structure and molecular weight, 2D and 3D structures, SMILES,
ClogP, and information on in vitro, in vivo, and ex vivo biological activities. Second,
curating the informat ion and compounds, and finally, selecting the management
system and creating the DB website.

20 D. Q. de Azevedo et al.
Fig. 2.1 The steps of the development of a molecular database. Step 1—search for chemical or
biological information in indexed databases: reports the strategies used to find the information that
will make up the DBs. Step 2—curation: describes the DB curation processes, automated or manual.
Step 3—DB management and network visualization: describes different systems to process DBs.
Step 4—update and maintenance: the last and most important stage is that, in addition to the
development of a DB, its maintenance requires various resources, both financial and human
2.1 First Step: Search for Chemical or Biological
Information in Indexed Databases
The strategy for feeding the compounds’ information to the DB mainly uses data
from indexed DBs. This approach was used in ChEMBL, NuBBE
BIOFACQUIM, NPACT, and TCM Database@Taiwan. The information contained
in a DB is extensive, as it may comprise 2D and 3D structures of compounds, with or
without their biological activity. Antiviral Medicinal Plants and Natural Products
DB (avMpNp DB) is a DB developed in Brazil and contains bioactive compounds
from biodiversity with antiviral activity. The strategy for building the avMpNp DB
consisted first of an extensive bibliographic search in academic DBs and compound
DBs, to systematize information about the bioactive compounds included in the DB
[43]. ChEMBL is a DB that was introduced in 2009 as an open-access resource and
plays an important role in drug discovery and validation of computational tool s. A
large proportion of this bioactivity data in ChEMBL is currently manually extracted
from scientific literature [15]. NPASS (Natural Product Activity and Species Source
Database) [44], COCUNUT (COlleCtion of Open Natural ProdUcTs) [45], and
DB
,

2 Molecular Databases 21
PHCS (Persian Herbal Constituents Database) [46] are databa ses of natural products
that also employ manual information extraction of the literature.
The availability of public chemistr y and bioactivity DBs, along with large-scale
data-driven applications, has increased the community’s attention to data curation
and integrity issues, such as structure quality, name-to-structure fidelity, structure–
activity mapping, activity data accuracy, assay description sufficiency, target assignment, author errors, and redundancy. Together, such factors can provide higher
confidence assertions and therefore more robust applications and models from the
available compounds’ information in DBs.
In particular, the compounds contained in DBs have already led to the development of drugs in clinical use to treat various diseases and are contributing to the
development of compounds in clinical development and basic research. Virtual
libraries thus help to increase the success rate of the lead selection process by
ensuring the quality, diversity, and consistency of the curated data. The number of
public domain compound databases is increasing and the process of building and
curating these libraries is critical as the data must be diver se and reliable to enable
safe trials. Therefore, the assembly and analysis of compound DBs in terms of their
structure and the legitimacy of the structures is crucial [47].
2.2 Second Step: Data Curation
Data curation includes, for example, elimination of salts, adjustment of protonation
states, optimization of geometry by energy minimization, and elimination of duplicated molecules. This curation process can involve several manual and automated
steps and aims to maximize data accessibility and comparability and improve data
integrity and flag outliers, ambiguities, and potential errors. Standard protocols are
used in manual DB curation processes. Although this step is not easy, it is feasible. It
would be advisable for a research group to be responsible for this endeavor, using
publicly available tools and scripts or workflows available on public repositories
such as GitHub, where examples of database curation are freely available [48].
The automated curation process uses platforms with different functionalities. One
example is Open Babel, a tool that provides a solution to the proliferation of different
file formats in chemistry. It also contains conformer searching, 2D visualization,
filtering, batch conversion, substructure, and similarity searching. For developers, it
can be a programming library for chemical data handling in areas such as organic
chemistry, drug design, materials science, and computational chemistry. It is freely
available under an open-source license from http://openbabel.org [49].
ChEMBL, for example, is a database that uses both manual and automated
curation strategies and implements their validation and standardization proces s
using pipelining tools such as Pipeline Pilot [50] or the Konstanz Information
Miner (KNIME) analysis platform [51]. These tools also allow for more flexibility
such as new components that can be added or adapted as needs change. The
ChEMBL DB providers routinely include a salt stripping process in their

22 D. Q. de Azevedo et al.
Table 2.2 Some examples of tools of automatized curation: platforms, functionalities, and uses
Platform Advantages
Molecular
Operating
Environment
(MOE)
Konstanz
information
miner
(KNIME)
b
RDKit
Open Babel Open, collaborative project
a
Advantage of KNIME;bAdvantage of RDKit
Integrated computer-aided
molecular design platformsmall molecules, peptides,
biologics
Help draw chemical structures and facilitate the storage
and interconversion between
standard file formats
Free and open-source academic version
a
Offers over 300+ connectors
to data sources, and integrations to all popular machine
a
learning libraries
b
Manipulate molecular structures in Python
a
Freely available
Convert, analyze, or store
data from molecular modeling, chemistry, biochemistry,
or related areas
Functionalities for curation
(examples) Databases use
Disconnects salts and metals
Removes simple components
Recalculates states of protonation, determines wedge
bonds for bonds from chiral
centers
Calculates missing chiral
parities from existing wedge
bond
Removal of characters
encoding stereoisomerism in
SMILES format (@; \; /)
Removal of salts
Neutralization of charges
Maintenance of compounds
containing only the following
elements: (H, C, N, O, F,
Br, I, Cl, P, S)
Creation of compounds in
InChI, InChIKey, and
canonical SMILES formats
from standardized
compounds
Removal of salts
Adjustment of the protonation state of the structures
Convert between molecule
file formats
NuBBE
DB
BIOFACQUIM
PeruNPDB
Super Natural
II
ChEMBL
SWMD
ChemDB
standardization based on a library of pharmaceutically relevant salts [15]. Table 2.2
shows examples of automatized curation, platform, functionalities, and their use.
NuBBE
[31] and BIOFACQUIM [32] use Molecular Operating Environment
DB
(MOE) [52] for automatized curation. This node disconnects salts and metals,
removes simple components, recalculates states of protonation, determines wedge
bonds for bonds from chiral centers, and calculates missing chiral parities from
existing wedge bonds. Using this same software, inorganic compounds can be
eliminated, as well as duplicated compounds [53, 54].
Another example of a DB using automated curation strategies is the Seaweed
Metabolite Database (SWMD), which comprises compounds, derived from seaweeds, and it has been curated using the Open Babel software. This strategy allowed
the removal of salts and the adjustment of the protonation state of the structures.
Duplicate structures were manually removed after detection in SMILES strings
using Microsoft Excel 2016, followed by manual inspection [55]. Open Babel is a
full-featured open chemical toolbox, designed to translate the many different

2 Molecular Databases 23
representations of chemical data [49]. It allows anyone to search, convert, analyze, or
store data from molecular modeling, ch emistry, solid-state materials, biochemistry,
or related areas. It provides both ready-to-use programs as well as a complete,
extensible programmer’s toolkit for developing cheminformatics software. In addition, the ChemDB is a small molecule database that also uses Open Babel to convert
between molecule file formats [56].
KNIME analysis platform, which includes RDKit, is also used to curate chemical
structures in DBs, such as PeruNPDB, the Peruvian Natural Products Database, and
Super Natural II [57, 58]. RDKit is an open-source cheminformatic s toolkit written
in C++ that is also usable from Java or Python. It includes a collection of standard
cheminformatics functionality for molecules, substructure searching, chemical reactions, coordinate generation (2D or 3D), fingerpr inting, curation, as well as a highperformance database cartridge for working with molecules using the PostgreSQL
DB [57 –59].
DataWarrior is a multi-functional and interactive chemical data analysis and
visualization tool. It provides interactive options for visualizing and curating data,
assessing correlations, and extracting knowledge from large datasets [60]. The
DataWarrior tool was used to eliminate duplicate structures in the different studies,
such as in the identification of anti-schistosomal, anthelmintic, and antileishmanial
compounds [61, 62 ].
DBs also use manual curation, for example, ChEBI (Chemical Entities of Biological Interest), which is a manually curated DB and ontology that organizes small
molecule knowledge [63]. Last, PSC-db [64] and avMpNp DB [ 43 ] also employ
manual curation using internal scripts.
2.3 Third Step: Database Management and Network
Visualization
Molecular DBs contain a wide variety of data that may be processed by different
systems, which, if unrelated, can lead to redundancy and inconsistency for the user.
To solve this problem, the data needs to be stored only once on a platform that can be
accessed and shared by all the systems involved. This solution results in a more
complex software structure that requires the help of a database management system
(DBMS) to maintain.
The choice of the DBMS model is fundamental to its development, as it is the
basis for structuring a database (data types, relationships, and relevant constraints).
The relational model is the most widely used model because of its greater flexibility
and suitability for design and implementation, and many DB systems today are
based on it. These characteristics are due to its structuring of data into relationships.
Examples of DBs that are managed using a relational model include Super Natural II
[65], ChEBI [63], Viether b [66], ZINC [13], and NPACT [33].

24 D. Q. de Azevedo et al.
Another option for implementing a molecular DB is to use a workflow-based
management system (WBMS). WBMSs are data management systems that flexibly
control the execution of a set of tasks, allowing this set of tasks to be modified
without changing the system code. A major advantage of using a WBMS is its
flexibility, changes to the stored information are reflected in changes into workflows,
which makes it much easier to adapt the system and does not require changes in the
system code. This is important because different bioactive compounds may have
different types of data associated with them, and in a workflow-based WBMS, all
this heterogeneous data can be easily managed [67].
2.4 Fourth Step: Updating and Maintenance
Usually, the DB function refers to compound repositories. In fact, compound DBs
and their chemical datasets are a central part of pharmaceutical companies and
private or government research centers. These DBs have been upgraded through
the cooperation of chemoinformatics tools and the introduction of the new compound. A study that evaluated 52 DBs developed over the last 30 years found that
most of them originated in the academic sector, such as in universities and research
institutes, where maintenance and upgrading also take place. Private DBs are
maintained by industry and thei r data is usually confidential [43] (Fig. 2.2).
3 Advantages and Disadvantages of Using Molecular
Databases
Molecular DBs are useful resources in computer-aided drug design and play a
central role in many chemoinformatics applications. It is possible to identify potential bioactive compounds with therapeutic activity through several chemoinformatic
methodologies. Indeed, molecular DBs can provide access to hundreds, thousands,
or even hundreds of thousands of compounds that can be virtually screened to
predict which of them have the desired biological activity. Recently, the amount of
freely accessible databases has increased, allowing access to an increasing number of
compounds [68]. Molecular DBs not only provide access to chemical structures but
also contain other useful information for the different drug design stages (Table 2.3).
As illustrated in Table 2.3, molecular databases are useful not only in drug
discovery. For instance, SciFinder [78] allows users to check the reported synthesis
pathways of a molecule and ChemSpider [24] contains spectroscopic data of the
compounds. Other examples are ZINC [13], ChEMBL [15], and ChemSpider
[24]. DBs provide information regarding the commercial availability of chemical
compounds. Despite the above-mentioned advantages of using molecular DBs
during the drug design process, there are still some associated deficiencies. The

2 Molecular Databases 25
Fig. 2.2 Example of how to build a DB, CHEMBL. Step 1—Search for chemical or biological
information in indexed databases: CHEMBL uses manual extraction of data from scientific literature. Step 2—Curation: CHEMBL DB uses tools such as KNIME for this purpose. Stage 3—DB
management and network visualization: Access to the data in ChEMBL, a relational DB, is
provided through a user interface, a set of web services, and a range of download formats including
XML, JSON, and YAML. Stage 4—Updating and maintenance
structural integrity of the molecules is not always assured, and the annotations are
not necessarily error-free. For instance, errors in the stereochemistry can be found in
DBs, as valence issues and charge imbalances [79] and small structural errors can
lead to significant losses of predictive abilities of quantitative structure–activity
relationship (QSAR) models [80]. For more details on QSAR and machine learning
predictors, please refer to Chap. 6. This is particularly common for publicly available
databases, where it can be challenging to have a dedicated team to curate the
chemical content and annotations in the compound DBs. Incomplete or wrong
information provided in chemical DBs is not only limited to the chemical structures.
Erroneous information regarding the biological activity of a molecule is something
else found in chemical databases. For instance, it is not always reported the assay
employed to determine the half maximal inhibitory concentration (IC
) or the
50
minimum inhibitory concentration (MIC) [81]. Besides, the assay conditions are
not always reported, or the original references are missing [65]. Such incom plete or
inaccurate data can lead to less reliable QSAR predictions. Thus, to improve the
quality of chemical databases, periodic revisions and error reporting are
recommended practices.

26 D. Q. de Azevedo et al.
Table 2.3 Categories into which databases can be divided according to the type of information
stored
Database
category Content Database
Chemical
information
Bioactivity Inhibitor constant (K
Drug Detailed drug data
Natural product Pathways (synthesis and degrada-
Chemical
availability
Fragment Physicochemical information
a
Representative examples. Some of the databases have more than one category that is not shown in
the table, for example, PubChem (chemical information and chemical availability), ChEMBL
(chemical information and drug), ChemSpider (bioactivity and chemical availability), and ZINC
(chemical information and bioactivity)
Chemical and crystal structures
spectra
Reactions and syntheses
Thermophysical data
)
i
Dissociation constant (K
Half maximal inhibitory concentration (IC
Half maximal effective concentration
(EC
Comprehensive drug target
information
tion)
Structures
Available compounds offered by
chemical vendors
Binding site preferences
)
50
)
50
)
d
ChemSpider
ChEBI
Chemical Universe Database GDB
PubChem
ChEMBL
BindingDB
ChemBank
PDBbind
DrugBank [28]
Universal Natural Product
Database
MeFSAT
Natural Product Atlas
ZINC
NCI
FDB-17
Fragment Store
PADFrag
a
Reference
[24]
[63]
[69]
[14]
[15]
[26]
[27]
[70]
[71]
[72]
[73]
[13]
[74]
[75]
[76]
[77]
4 Natural Product Databases for the Search
and Development of New Drugs
Nature is a rich source of bioactive molecules that serve as therapeutic agents. These
bioactive molecules are natural products, which can be defined as compounds
produced by living beings and can be employed or proposed as therapeutic agents
[82]. For instance, of the approved small molecules in the research area of cancer,
from 1946 to 1980, 53% of the compounds that became new medicines
corresponded to unaltered natural products or natural product derivatives. From
1981 to 2019, 64.9% of the approved small molecules were unaltered natural
products or natural products inspired [83]. Moreover, natural products are an
abundant source of privileged scaffolds: structures capable of providing useful
ligands for more than one receptor [84]. Some examples of privileged scaffolds
that come from natural products that are currently used in the design and development of new drug candidates are the terpenoid, polyketide, phenylpropanoid, and
alkaloid structures [85 ]. Regarding natural product DBs, one application is the

2 Molecular Databases 27
Table 2.4 Representative natural product databases
Database
Collection of Open
Natural Products
(COCONUT)
Universal Natural
Product Database
SuperNatural 3.0 449,058 Open access Toxicity
ZINC ∼80,000 Open access Bioactivities
Dictionary of Natu-
ral Products
SciFinder ∼300,000 Commercial Reported synthesis routes [78]
Reaxys ∼200,000 Commercial Bioactivity
TCM@Taiwan ∼58,000 Open access The largest database of natu-
IMPPAT ∼10,000 Open access The largest database of natu-
AfroDB ∼1000 Open access The largest database of natu-
Phyto4Health 3128 Open access Medicinal plants included in
NuBBE
DB
BIOFACQUIM 553 Open access Taxonomic information of
a
Date of search: April 2024
Number of
compounds
411,621 Open access Predicted bioactivities [45]
∼229,000 Open access 3D structures [71]
∼230,000 Commercial Spectroscopic data [82]
2223 Open access Bioactivity
a
Accessibility Outstanding features References
Vendor information
Vendor information
Toxicity
Physicochemical data
ral products from Traditional
Chinese Medicine
ral products from traditional
medicine in India
ral products from traditional
medicine in Africa
the Russian Pharmacopoeia
Bioactivities
Predicted bioactivities
Predicted spectroscopic data
the producing organism
[65]
[13]
[91]
[92]
[93]
[94]
[95]
[31, 96]
[32]
design of pseudo-natural products, that is, molecules that retain the biological
relevance of natural products yet exhibit structures and bioactivities not available
in nature or in existing design strategies. Pseudo-natural products may display
unexpected bioactivities that differ from the activities of the natural products from
which their fragments are derived [86–88]. Besides, natural products usually have
more structural diversity compared with synthesized small molecules [89].
Natural product DBs can be important tools for computer-aided drug design
(CADD), providing access to a large number of diverse chemical structures. These
DBs can be divided into commercial and open access (Table 2.4). Between 2000 and
2019, 123 commercial and open-access natural product collections have been
published, of which 98 are somewhat accessible, 92 are open access, and only

28 D. Q. de Azevedo et al.
50 contain molecular structures that can be retrieved for a chemoinformatic analysis
[90]. For example, the Collection of Open Natural Products (COCONUT) [45]
contains more than 411,000 natural product entries collected from 50 open-access
natural product DBs. Similarly, the Universal Natural Product Database [71]is
another compilation DB with more than 229,000 natural products. It provides 3D
structures with stereochemical information and calculated molecular descriptors but
is not yet accessible through the link in the original publication. Instead, it is
available on another website [81]. Currently, SuperNatural 3 is the bigges t openaccess DB, which contains over 449,058 unique natural products and includes
information about 2D structures, physicochemical properties, predicted toxicity
class, and potential sellers, but it does not yet provide the option to download in bulk.
There are natural product databases comprised of molecules isolated and characterized in specific geographical regions. China is one of the regions with more
natural product databases published [102 –110], because Traditional Chinese Medicine (TCM) is part of the Chinese public health system. There are two natural
product DBs that comprise compounds that are part of the traditional medicine in
India (Indian Ayurveda), such as IMPPAT [93] and MedPServer [106]. Regarding
African Traditional Medicine, there are different natural product DBs published such
as ConMedNP, p-ANAPL library, and others [107–111]. There is a database that
contains natural products from medicinal Russian plants: Phyto4Health [95]. Latin
America is a region that encompasses at least a third of global biodiversity [112]. All
the published natural product databases of Latin America and their practical applications in the drug disco very area have been reviewed and discussed elsewher e
[113]. Two representative examples of natural product databases from Latin America are NuBBE
[31] and BIOFACQUIM [ 32].
DB
5 Functionalization of Databases and Transformation into
a System to Drug Design
The usage of molecular DBs as a drug design tool depends on the capacity of the
chemoinformatic software to recognize the molecules. For this purpose, the simplified molecular input line entry system (SMILES) notation is the predominant input
notation for the different chemoinformatic software packages [114]. Other notations
that can be recognized by the chemoinformatic software overcome some disadvantages of the SMILES, such as the International Chemical Identifier (InChI) [115] and
InChIKey [116]. SMILES arbitrary target specification (SMARTS) notation was
developed to specify substructural patterns that allow matching molecules that
contain a specified substructural pattern [117]. The different notations used by the
chemoinformatic software, as well as their uses, advantages, and disadvantages,
have been explained in detail elsewhere [118].
Соседние файлы в папке Библиотека им академика М.И. Перельмана
