Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5884_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Foreword
- •Acknowledgments
- •Contents
- •1.1 Structure-Based Drug Discovery (SBDD)
- •1.2 Ligand-Based Drug Design (LBDD)
- •1.3 Echoes from the Past, Visions from the Future
- •References
- •1 Introduction
- •2.2 Second Step: Data Curation
- •2.4 Fourth Step: Updating and Maintenance
- •2 Databases and Curation
- •8 Perspectives
- •9 Conclusion
- •References
- •1 Introduction
- •2.1 Making and Matching Protein Models
- •2.2 Simulating Protein Movements
- •2.3 Analyzing Changes in Protein Shape
- •3 Pharmacogenomics in Drug Development
- •4 Case Studies of Genomics-Based Drug Design
- •References
- •1 Historical Background
- •1.1 Timeline
- •2 Methodology Overview
- •2.1 Neural Networks
- •2.1.1 Perceptron
- •2.1.2 Multilayer Neural Networks
- •2.1.3 Types of Neural Networks
- •Feedforward
- •Recurrent Neural Networks
- •LSTM
- •2.2 Deep Learning
- •3 Using Machine Learning
- •3.2 Data Collection
- •3.3 Data Preprocessing
- •3.4 Model Selection
- •3.5 Model Training
- •3.6 Validation
- •3.7 Tuning
- •3.8 Prediction
- •4 Limitations
- •4.1 Bias
- •4.3 Interpretability
- •4.4 Computational Cost
- •4.5 Data Dependency
- •4.6 Robustness
- •5 Applications in Drug Discovery
- •5.2 Lead Discovery
- •5.3 Preclinical and Clinical Development
- •6 Resources and Tools
- •7 Challenges and Perspectives
- •7.1 Future Trends
- •9 Conclusions
- •References
- •1 Historical Background
- •1.1 Applications in Drug Discovery
- •2 Validations and Controls
- •2.1 Internal Validation
- •2.2 External Validation
- •2.3 Relative Cluster Validation
- •3 Challenges and Perspectives
- •4 Conclusions
- •References
- •1 Historical Background
- •2 OECD Principles
- •2.1 A Defined Endpoint
- •2.2 An Unambiguous Algorithm
- •2.5 A Mechanistic Interpretation, if Possible
- •3 Software and Tools
- •4 Validations and Controls
- •4.1 Internal and External Validation
- •4.1.1 Regression Metrics
- •4.2 Applicability Domain
- •4.3 Randomization Tests
- •5 Interpretation
- •6 Practical Advice During QSAR Modeling
- •7 Application
- •8 Challenges and Perspectives
- •References
- •1 Molecular Docking
- •2 Advances in Scoring Functions and Search Algorithms
- •2.2 Critical Characteristics of Search Algorithms
- •2.3 Docking Programs and Scoring Functions
- •3 Calculations Performed During Docking Simulations
- •4 Essential Components for a Good Docking Program
- •5 Limitations of the Docking Technique
- •6 Validation of Docking Results
- •7 Inappropriate Use of Validation Methods in Docking
- •9 Use of Machine Learning in Molecular Docking
- •11 Challenges
- •12 Conclusions
- •References
- •3 System Preparation for MD Simulations
- •3.1 Solvation and Microensemble
- •3.2 Force Fields: General Concept and Relevant Choices
- •3.3 The Concept of Replicas and Timescale
- •4.1.2 Protein Root Mean Square Fluctuation (RMSF)
- •4.1.4 Protein Secondary Structure Analysis
- •4.1.5 Principal component Analysis (PCA)
- •4.1.6 Markov State Modelling
- •4.1.7 Distance Calculations
- •4.1.8 Angle and Plane Calculations
- •4.2.2 Distances and Ligand-Induced Geometry Rearrangements
- •4 Molecular Dynamics Analysis
- •4.1 Protein Perspective
- •4.1.1 Protein Root Mean Square Deviation (RMSD)
- •4.3 Ligand Perspective
- •4.3.1 Ligand Properties
- •4.3.2 Ligand Root Mean Square Deviation
- •4.3.3 Ligand Root Mean Square Fluctuation
- •4.3.4 Angles and Dihedrals
- •5.1 Protein Structure Prediction and Preparation
- •5.2 Molecular Docking
- •6 Concluding Remarks and Outlook
- •Glossary
- •References
- •1 Introduction
- •2.1 MDeNM
- •2.2 Collective Molecular Dynamics (coMD)
- •2.3 ClustENM and ClustENMD
- •3 Ensemble Docking
- •References
- •1 Introduction
- •1.1 Advantages, Disadvantages, Innovations, and Challenges
- •1.2 Recent Advances in Accessible FEP Software Tools
- •1.3 Applications of FEP in Industry and Consortiums
- •2 Expanding the Potential of FEP Calculations
- •2.1 Validating Binding Poses
- •2.2 Dealing with Solvent
- •2.3 FEP and Allostery
- •2.4 FEP and Covalent Ligands
- •2.5 Applications of FEP in Scaffold Hopping
- •2.6 Positional Analogue Scanning
- •2.7 Combinations and Alternative Approaches
- •3 Machine Learning for FEP
- •3.4 Implications for ML in FEP Calculations
- •4 Final Considerations
- •5 First Steps to FEP Simulations
- •References
- •1 Background
- •2 Ultra-Large Screening Libraries and Chemical Spaces
- •3.1 Implications of Dataset Size
- •4 Ligands on the Ultra-Large Scale
- •4.1 Ultra-Large 2D Similarity Searches
- •7 Challenges and Future Perspectives
- •7.1 Hit Triage: An Old Problem on a New Dimension
- •8 Conclusions
- •Appendix
- •References
- •1 Introduction
- •2 Enzymatic Activity Evaluations
- •3 Cytotoxicity Evaluation and Cell Viability
- •4 Antiviral Assays in Experimental Validation
- •6 In Vivo Evaluation of Compounds
- •7 Conclusions
- •References
- •1 Introduction
- •3.1 Data Collection
- •3.2 Data Preprocessing
- •3.4 Model Choice
- •3.5 Model Training
- •3.6 Model Assessment
- •3.7 External Validation
- •3.8 Implementation and Availability
- •3.9 Continuous Update
- •5 Conclusions and Perspectives
- •References
- •1 Experimental Approaches to Obtain Protein Structure
- •1.1 X-Ray Crystallography
- •1.2 Nuclear Magnetic Resonance
- •1.3 Cryo-EM
- •1.4 Hybrid Methods
- •2 Modeling Approaches to Obtain Protein Structure
- •2.1 Homology Modeling
- •2.2 Ab Initio Modeling
- •2.3 New Approaches
- •3 Conformational Diversity of Proteins
- •3.1 Characterization of Protein Conformational States
- •3.2 Experimental Methods to Study Protein Dynamics and Conformations
- •3.4 Molecular Dynamics Simulation
- •3.5 Sampling Strategies
- •4 Remarks and Perspectives
- •References
- •1 Introduction
- •2 Structure-Based Drug Design of HIV Protease Inhibitors
- •2.1 HIV-1 Protease as a Therapeutic Target
- •2.2.1 Saquinavir
- •2.2.2 Indinavir
- •2.3.1 Lopinavir
- •2.3.2 Darunavir
- •6 Conclusions
- •References
- •4 Experimental Methods to Analyze NR Activity
- •4.2 Coregulator-Recruitment
- •5 Concluding Remarks and Outlook
- •References

2 Molecular Databases 29
Fig. 2.3 Nomenclature and identification of compounds in a database, for example, captopril
Free-access online tools that allow the user to draw a chemical structure and
convert it into either SMILES, InChI, InChIKey, or SMARTS can be found
(Fig. 2.3), such as PubChem Sketcher [119]. Additionally, there are also tools that
automatically recognize the chemical structures inside a .pdf file or image and
transform into a SMILES string. Notwithstanding a manual inspection of the output
molecules is required to detect possible mistakes in the output structures [120]. Molecules can be represented according to thei r 2D or 3D coordinates, and the two most
commonly formats used are .mol2 and .sdf. These files can store the atomic
coordinates of each atom of a molecule and the information of the connectivity
among the atoms. Moreover, both formats allow the storage of physicochemical
information of the molecule, while the .mol2 format allows the storage of the
atomical charges of every atom [121, 122].
SMILES (Simplified Molecular-Input Line-Entry System) uses graph theory to
represent structures. Each atom is represented by its symbol in the periodic table,
brackets are used to indicate branching points, and numerical labels denote ring
junctions. The basic grammar of SMILES also includes isot opic information, configuration on double bonds, and chirality, also known as isomeric SMILES. Later,
IUPAC (International Union of Pure and Applied Chemistry) developed its own
linear notation for representing chemical structures, InChIKey (International Union
of Pure and Applied Chemistry Key), based on the InChI identifier. A compound 2D
structure image can be generated using MarvinSketch [123].

30 D. Q. de Azevedo et al.
Moreover, there are available open-source molecular DBs that can be used during
drug design approaches (Tables 2.3 and 2.4 in this chapter). Once the database is
already built and curated (see Sect. 2 of this chapter), or obtained either open source
or commercial, it is possible to proceed to determine helpful information for the drug
design process, with free-access online tools. One can cite NERDD [124], which
offers access tools for drug discovery. NERDD hosts tools for predicting the sites of
metabolism (FAME) [125 ] and metabolites (GLORY) [126] of small organic molecules, for flagging compounds that are likely to interfere with biological assays (Hit
Dexter) [127], and for identifying natural products and natural product derivatives in
large compound collections (NP-Scout) [128]. NERDD also has a tool that predicts
if a small organic compound is an inhibitor of different human CYP450 isoforms
(CYPlebrity) and another that classifies molecules into skin sensitizers and
nonsensitizers [129]. These online tools also allow the prediction of the absorption,
distribution, metabolism, excretion, and toxicity (ADMET) profile of a single or a
list of molecules, such as ADMETlab [130] and PhaKinPro [131], among at least
18 free web servers capable of predicting ADMET properties, which have been
discussed elsewhere [132].
Additionally, the synth etic accessibility score is another helpful metric that allows
to determine the synthetic feasibility of given a molecule. With the synthetic
accessibility score, it is possible for a user to select from a DB those compounds
with the highest synthetic feasibility [133]. SYBA [134] and GASA [135] are two
different synthetic accessibility scores that can be determined with free-access online
tools. In addition, retrosynthetic planning can also be useful, as in AiZynthFinder,
which recursively breaks down a molecule to its purchasable precursors [136].
Furthermore, one can also capture structural information from molecules. Two
common molecular fingerprints employed to capture structural information are the
Molecular ACCess System (MACCS) keys-166 bits [137] and the Extended Connectivity Finger print (ECFP6) [138]. MACCS keys and ECFP6 fingerprints encode
the structure of the molecules in strings made of bits, that is, every molecule is
represented with a string made by one and zeros. From the fingerprints, the structural
similarity among the molecules can be measured using the Tanimoto coef ficient
(Tc) [139]. Herein, if two molecules share similar structures, then they will likely
have similar bioactivities [140 ]. Thus, a molecule with a certain known biological
activity can be identified using Tc calculations in a molecular database with similar
compound’s structure possibly having a similar bioactivity to the matched molecule.
These fingerprints can be calculated, for example, by both CDK [141] or RDKit
[142] nodes of the free-available software KNIME [51]. RDKit also can be used in
the Python programming language. Last, another useful and free computational tool
to calculate the fingerprints is PaDEL [
143].

2 Molecular Databases 31
6 Chemoinformatics and Computational Tools
for Supporting and Filtering Potential Drug Bioactive
Compounds from Databases
The SMILES strings (see Sect. 5 in this chapter) avail able from the molecular DBs
can also be used to determine physicochemical properties of interest, such as SlogP
[144], molecular weight (MW), topological polar surface area (TPSA) [145], rotatable bonds (Rb), hydrogen bond acceptors (HBA), and hydrogen bond donors
(HBD). The calculation of the physicochemical properties can be performed
using the freely available software DataWarrior [146], KNIME [51] or using Python,
employing the CDK [141 ] and RDKit [142] nodes (KNIME) or RDKit package (Python). One could argue that out of these three options, DataWarrior is the
most user-friendly, as the list of SMILES strings can be imported into the program
just with a copy/paste keyboard shortcut. Once the physicochemical properties have
been calculated, the compounds can be filtered to a chosen numerical range for each
property. There are some drug-likeness guidance parameters, also called rules of
thumb by some authors, which are suggestions that help to choose the potential druglike compounds such as Lipinski’s rule of 5 (Ro5) [147, 148], Veber’s rules [149],
GlaxoSmithKline’s (GSK) 4/400 rule [150], and Pfizer 3/75 rule [151] (Table 2.5).
Compliance with either Lipinski’s, Veber’s, or GSK rules is commonly associated
with good oral bioavailability. These suggestions or rules of thumb are based on the
premise that certain physicochemical properties are directly associated with one or
more parameters of the ADMET profile. Thus, if a group of molecules has a known
ADMET profile, it is expected that other molecules woul d have a similar ADMET
profile whenever they share similar physicochemical properties. Usually, these rules
are helpful in filtering compounds that are expected to have a desirable ADMET
profile [152]. Nonetheless, one should consider that the ADMET profile of a
compound is influenced by many other aspects and not only by some physicochemical properties. Therefore, the fulfillment of one or more of these rules of thumb does
not necessarily assure obtaining the expected desirable value of a parameter of the
ADMET profile [153].
7 Databases in Drug Design and Discovery: Applicability,
Challenges, and Successful Cases
The field of drug design and discovery still faces low efficacy, off-target delivery, as
well as the time spent, and high cost associated [154]. In this scenario, the increase in
biological and chemical data, including in vitro, in vivo, clinical studies, genomics
studies, proteomics studies, metabolomics studies, gene ontology studies, and
molecular pathway data, benefits from different data repositories that have been
developed. For instance, ChemSpider, ChEMBL, ZINC, BindingDB, and PubChem
are the essential DBs for compound synthesis and screening in the drug design and

32 D. Q. de Azevedo et al.
Number of
rotatable
bonds References
Topological
polar surface
area
Sum hydrogen bond
acceptors and donors
Number of
hydrogen bond
donors
Number of
hydrogen bond
acceptors
Molecular
weight LogP
Table 2.5 Rules of thumb or guides associated with drug-likeness
≤ 500 ≤ 5 ≤ 10 ≤ 5[147, 148]
Lipinski’s rule of
<400 <4 [150]
GlaxoSmithKline
Veber’s rules ≤ 12 ≤ 140 ≤ 10 [149]
5
4/400 rule
Pfizer 3/75 rule >3 < 75 [151]

2 Molecular Databases 33
discovery process. ZINC database was the most preferred DB in virtual screening
studies with an average use of 31.2% from 2015 to 2022 [133], selecting compounds
according to their pharmacological and physicochemical properties of pharmaceutical interest [155].
Among the new drug design strategies, the search for information and computational tools is the most widely used due to its great ease of access and extremely low
cost. A DB displays various techniques and information from, for example, in silico
studies, such as molecular docking, molecular dynamics simulation, and
VS. Therefore, a given DB is designed to include the information necessary for
the first stage of drug discovery, including molecular structures, physicochemical
properties, molecular properties of ligands, and their drug-likeness properties [156].
VS is an approach that benefits from virtual DBs containing a large number of
compounds. VS allows the identification of novel hits as well as the prioritization of
compounds for in vitro testing, resulting in a significant reduction in costs and
attrition rates [154]. This pre-selection is done by virtually predicting the biological
activity of interest using different structure-based drug design and/or ligand-based
drug design (SBDD and/or LBDD) approaches. SBDD was the most prominently
used type of VS and it accounted for an average of 57.6%, from 2015 to 2022
[133]. In the early stages of drug design, in silico studies represent a relevant strategy
for reducing costs in this process [28]. DrugBank is one free-access website
containing information on potential drugs and their targets that is usually accessed
for such approaches. Drugbank is a bioinformatics and chemoinformatics resource
that contains detailed information of drugs and their biological targets (29.785
targets). It contains a detailed description of these targets, such as their type,
organism in which they occur, pharmacological function, and specific and general
functions [156].
Despite successful drug discovery approaches accessing different DBs, it has
been reported that between 0.1% and 3.4% of the chemical structures in chemical
DBs are incorrect [157], including errors in molecular DBs regarding compounds’
bioactivity data [158]. An error in the structure representation of the molecules can
lead to inaccurate QSAR predictions, erroneous hazard and risk assessments, and
wrong decisions in the early drug discovery process. In molecular DBs, it has been
identified fundamental errors in stereochemistry, valency issues, and charge imbalances [157]. Unfortunately, still in 2023, an analysis identified a subst antial number
of errors in the identifiers and chemical structures in DBs, such as PubChem,
CompTox Chemicals Dashboard, and European Chemicals Agency (ECHA). Further, the absence of these errors is not guaranteed even in government-funded DBs
[159]. Therefore, it is important to identify and correct those wrong entries and
structures in the molecular DBs.
There have been efforts with the aim of reverting the problem, such as an opensource chemical structure curation pipeline that was published and utilized successfully to standardize the nearly two million compounds in the ChEMBL DB [160]. In
addition, compounds with critical issues were also identified so that they could be
prioritized for manual curation [161]. The solution to these challenges requires
coordinated strategies and a strong finan cial commitment by the government,

34 D. Q. de Azevedo et al.
funders, and institutions. It is estimated that the current global costs of maintaining
public biomedical data repositories are under $300 million annually [162]. To
maintain this, it is important to have a collaboration from universities, pharm aceutical industry, and institutions with the government to raise awareness about the
importance of the situation with the capacity to maintain curate and create
chemical DBs.
8 Perspectives
Based on research about DBs commonly used in drug discovery, we highlight the
relevance of these libraries in this process. Natural product DBs, for example, are
becoming promising accelerators in the development of drugs from bioactive compounds for various therapeutic areas such as cancer, viral, and neglected diseases.
There is a growing trend toward such libraries, i.e., repositories of information on
compounds with specific pharmacological activities. Currently, and in the near
future, these virtual libraries could also enrich the medically relevant chemical
space for inhibitors and potential drug candidates. In addition, the incorporation of
chemoinformatics tools, the automation of curation processes, and the allo cation of
financial and human resources for the maintenance of DBs will contribute significantly to the development of drug design discovery and development projects.
9 Conclusion
The process of developing new bioactive compounds into drugs is long and complex. Scientific and technological advances at the interface of chemistry and biology
have created remarkable opportunities and challenges for research and development.
Computational simulations have been playing an important role in reducing costs
and accelerating the drug design and discovery processes. These advantages,
together with the large number of chemi cal entities available in DBs, represent a
vast area of research for new potential drugs. The importance of DBs in new drug
discovery projects is continuously increasing and is a centerpiece in pharmaceutical
companies as well as in academic and government research centers. Notwithstanding, the quality of the compound DBs is crucial to properly fulfill their role as drug
design tools. Therefore, the curation and maintenance of high-quality data is essential. Compounds discovered in existing DBs have already led to the development of
drugs in clinical use to treat different diseases. This shows the feasibility of accessing
different DBs that can continue to contribute to the field of drug design, discovery,
and development.

2 Molecular Databases 35
Questions to answer when planning an experiment using databases
What computational approaches are being used to accelerate drug discovery
What is a compound database?
What is the relevance of a compound database in drug discovery?
What are the greatest practical challenges to developing compound repositories?
What are the major steps to construct a compound database?
What are the largest compound databases annotated with biological activity?
What is data curation?
Mention examples of software used for data curation.
What are the major deficiencies or shortcomings in databases, particularly in the public domain?
Provide examples of at least five natural product databases publicly available.
What are the main linear notations to store the information of chemical structures in compound
databases?
What are the main rules of thumb or empirical rules used in drug discovery projects?
References
1. Kubinyi, H. (2004). Industrial Pharmacy,7.
2. Faller, B., Ottaviani, G., Ertl, P., Berellini, G., & Collis, A. (2011). Drug Discovery Today,
16, 976.
3. Viegas, C., Bolzani, V. S., & Barreiro, E. J. (2006). Quimica Nova, 29, 326.
4. Yunes, R. A., & Filho, V. C. (Org.). (2007). Química de produtos naturais, novos fármacos e a
moderna farmacognosia. 1 ed (3003 p). Ed. Universidade do Vale do Itajaí.
5. Drews, J. (2000). Drug discovery: A historical perspective. Science, 287, 1960–1964.
6. Rotella, D. P. (2016). The critical role of organic chemistry in drug discovery. ACS Chemical
Neuroscience, 7, 1315–1316.
7. Cummings, J. L., Morstorf, T., & Zhong, K. (2014). Alzheimer’s disease drug development
pipeline: Few candidates, frequent failures. Alzheimer ’ s Research & Therapy, 6, 37.
8. Ridley, R. G. (2002). Medical need, scientific opportunity and the drive for antimalarial drugs.
Nature, 415, 686–693.
9. Klebe, G. (2006). Virtual ligand screening: Strategies, perspectives and limitations. Drug
Discovery Today, 11, 580–594.
10. Song, C. M., Lim, S. J., & Tong, J. C. (2009). Recent advances in computer-aided drug design.
Briefings in Bioinformatics, 10, 579–591.
11. Chan, H. C. S., Shan, H., Dahoun, T., et al. (2019). Advancing drug discovery via artificial
intelligence. Trends in Pharmacological Sciences, 40, 592–604.
12. Roda, C. I. N. (2022). A inteligência artificial na descoberta de novos medicamentos (Doctoral
dissertation).
13. Sterling, T., & Irwin, J. J. (2015). ZINC 15–ligand discovery for everyone. Journal of
Chemical Information and Modeling, 55(11), 2324–2337.
14. Kim, S., Thiessen, P. A., Bolton, E. E., Chen, J., Fu, G., Gindulyte, A., Han, L., He, J., He, S.,
Shoemaker, B. A., Wang, J., Yu, B., Zhang, J., & Bryant, S. H. (2016). PubChem substance
and compound databases. Nucleic Acids Research, 44, D1202–D1213.
15. Papadatos, G., Gaulton, A., Hersey, A., & Overington, J. P. (2015). Activity, assay and target
data curation and quality in the ChEMBL database. Journal of Computer-Aided Molecular
Design, 29, 885–896.

36 D. Q. de Azevedo et al.
16. Benson, D. A., Cavanaugh, M., Clark, K., Karsch-Mizrachi, I., Ostell, J., Pruitt, K. D., &
Sayers, E. W. (2018). GenBank. Nucleic Acids Research, 46(D1), D41–D47.
17. Burley, S. K., Berman, H. M., Kleywegt, G. J., Markley, J. L., Nakamura, H., & Velankar,
S. (2017). Protein data Bank (PDB): The single global macromolecular structure archive.
Protein Crystallography: Methods and Protocols, 627–641.
18. Thul, P. J., & Lindskog, C. (2018). The human protein atlas: A spatial map of the human
proteome. Protein Science, 27(1), 233–244.
19. Wishart, D. S., Guo, A., Oler, E., Wang, F., Anjum, A., Peters, H., et al. (2022). HMDB 5.0:
The human metabolome database for 2022. Nucleic Acids Research, 50(D1), D622–D631.
20. Hummel, J., Selbig, J., Walther, D., & Kopka, J. (2007). The Golm metabolome database: A
database for GC-MS based metabolite profiling. In Metabolomics: A powerful tool in systems
biology (pp. 75–95). Springer.
21. Wang, M., Carver, J. J., Phelan, V. V., Sanchez, L. M., Garg, N., Peng, Y., et al. (2016).
Sharing and community curation of mass spectrometry data with global natural products social
molecular networking. Nature Biotechnology, 34(8), 828–837.
22. Ulrich, E. L., Akutsu, H., Doreleijers, J. F., Harano, Y., Ioannidis, Y. E., Lin, J., et al. (2007).
BioMagResBank. Nucleic Acids Research, 36(suppl_1), D402–D408.
23. Horai, H., Arita, M., Kanaya, S., Nihei, Y., Ikeda, T., Suwa, K., et al. (2010). MassBank: A
public repository for sharing mass spectral data for life sciences. Journal of Mass Spectrom-
etry, 45(7), 703–714.
24. Pence, H. E., & Williams, A. (2010). ChemSpider: An online chemical information resource.
Journal of Chemical Education, 87, 1123–1124.
25. Moda, T. L., Torres, L. G., Carrara, A. E., & Andricopulo, A. D. (2008). PK/DB: Database for
pharmacokinetic properties and predictive in silico ADME models. Bioinformatics, 24(19),
2270–2271.
26. Liu, T., Lin, Y., Wen, X., Jorissen, R. N., & Gilson, M. K. (2007). BindingDB: A
web-accessible database of experimentally determined protein-ligand binding affinities.
Nucleic Acids Research, 35, D198–D201.
27. Seiler, K. P., George, G. A., Happ, M. P., Bodycombe, N. E., Carrinski, H. A., Norton, S.,
Brudz, S., Sullivan, J. P., Muhlich, J., Serrano, M., Ferraiolo, P., Tolliday, N. J., Schreiber,
S. L., & Clemons, P. A. (2008). ChemBank: A small-molecule screening and cheminformatics
resource database. Nucleic Acids Research, 36, D351–D359.
28. Miller, M. A. (2002). Chemical database techniques in drug discovery. Nature Reviews Drug
Discovery, 1(3), 220–227.
29. Medina-Franco, J. L. (2015). Discovery and development of lead compounds from natural
sources using computational approaches. In Evidence-based validation of herbal medicine
(pp. 455–475). Elsevier.
30. Medina-Franco, J. L. (2020). Towards a unified Latin American natural products database:
LANaPD. Future Science OA, 6(8), FSO468.
31. Pilon, A. C., Valli, M., Dametto, A. C., Pinto, M. E. F., Freire, R. T., Castro-Gamboa, I., et al.
(2017). NuBBEDB: An updated database to uncover chemical and biological information
from Brazilian biodiversity.
32. Pilón-Jiménez, B. A., Saldívar-González, F. I., Díaz-Eufracio, B. I., & Medina-Franco, J. L.
(2019). BIOFACQUIM: A Mexican compound database of natural products. Biomolecules, 9.
33. Mangal, M., et al. (2013). NPACT: Naturally occurring plant-based anti-cancer compoundactivity-target database. Nucleic Acids Research, 41(D1), D1124–D1129.
34. Liu, Y., Zhu, Y., Sun, X., Ma, T., Lao, X., & Zheng, H. (2023). DRAVP: A comprehensive
database of antiviral peptides and proteins. Viruses, 15, 820.
35. Martin, H. J., Melo-Filho, C. C., Korn, D., Eastman, R. T., Rai, G., Simeonov, A., et al. (2022).
Small Molecule Antiviral Compound Collection (SMACC): a database to support the discovery of broad-spectrum antiviral drug molecules. bioRxiv.
Scientific Reports, 7(1), 7215.

2 Molecular Databases 37
36. Tzou, P. L., Tao, K., Pond, S. L. K., & Shafer, R. W. (2022). Coronavirus resistance database
(CoV-RDB): SARS-CoV-2 susceptibility to monoclonal antibodies, convalescent plasma, and
plasma from vaccinated persons. PLoS One, 17, e0261045.
37. Martin, R., Loechel, H. F., Welzel, M., Hattab, G., Hauschild, A. C., & Heider, D. (2020).
CORDITE: The curated CORona drug InTERactions database for SARS-CoV-2. Iscience,
23(7).
38. Chen, T. F., Chang, Y. C., et al. (2021). DockCoV2: A drug database against SARS-CoV-2.
Nucleic Acids Research, 49(D1), D1152–D1159.
39. Zhou, N., Bao, J., & Ning, Y. (2021). H2V: A database of human genes and proteins that
respond to SARS-CoV-2, SARS-CoV, and MERS-CoV infection. BMC Bioinformatics,
22, 18.
40. Alsulami, A. F., Thomas, S. E., Jamasb, A. R., Beaudoin, C. A., Moghul, I., Bannerman, B.,
Copoiu, L., Vedithi, S. C., Torres, P., & Blundell, T. L. (2021). SARS-CoV-2 3D database:
Understanding the coronavirus proteome and evaluating possible drug targets. Briefings in
Bioinformatics, 22, 769–780.
41. Koes, D. R., & Camacho, C. J. (2012). ZINCPharmer: Pharmacophore search of the ZINC
database. Nucleic Acids Research, 40, W409–W414.
42. Miranda-Salas, J., Peña-Varas, C., Martínez, I. V., Olmedo, D. A., Zamora, W. J., ChávezFumagalli, M. A., et al. (2023). Trends and challenges in chemoinformatics research in Latin
America. Artificial Intelligence in the Life Sciences, 100077.
43. de Azevedo, D. Q., Campioni, B. M., Pedroz Lima, F. A. L., Medina-Franco, J., Castilho,
R. O., & Maltarollo, V. G. (2024). A critical assessment of bioactive compounds databases.
Future Medicinal Chemistry,1– 23.
44. Zeng, X., Zhang, P., He, W., Qin, C., Chen, S., Tao, L., et al. (2018). NPASS: Natural product
activity and species source database for natural product research, discovery and tool development. Nucleic Acids Research, 46(D1), D1217–D1222.
45. Sorokina, M., Merseburger, P., Rajan, K., Yirik, M. A., & Steinbeck, C. (2021). COCONUT
online: Collection of open natural products database. Journal of Cheminformatics, 13(1),
1–13.
46. Karimi-Jafari, M. H., Firouzi, R., Ashouri, M., & Poursoleiman, A. A. (2022). A database of
chemical compositions of Persian medicinal herbs. https//:chemrxiv.org/engage/chemrxiv/
articledetails/621e71035f1d9a5bb3ad2173. Accessed 20 Mar 2022.
47. Piccirillo, E., & Amaral, A. T. D. (2018). Busca virtual de compostos bioativos: conceitos e
aplicações. Química Nova, 41, 662–677.
48. Chávez-Hernández, A. L., Sánchez-Cruz, N., & Medina-Franco, J. L. (2020). Fragment library
of natural products and compound databases for drug discovery. Biomolecules, 10(11), 1518.
49. O’Boyle, N. M., Banck, M., James, C. A., Morley, C., Vandermeersch, T., & Hutchison, G. R.
(2011). Open babel: An open chemical toolbox. Journal of Cheminformatics, 3(1), 1–14.
50. Warr, W. A. (2012). Scientific workflow systems: Pipeline pilot and KNIME. Journal of
Computer-Aided Molecular Design, 26 (7), 801–804.
51. Berthold, M. R., Cebron, N., Dill, F., Gabriel, T. R., Kötter, T., Meinl, T., Ohl, P., Thiel, K., &
Wiswedel, B. (2009). KNIME—The Konstanz information miner. SIGKDD Exploration
Newsletter, 11, 26.
52. Vilar, S., Cozza, G., & Moro, S. (2008). Medicinal chemistry and the molecular operating
environment (MOE): Application of QSAR and molecular docking to drug discovery. Current
Topics in Medicinal Chemistry, 8(18), 1555–1572.
53. Saldívar-González, F. I., Valli, M., Andricopulo, A. D., da Silva Bolzani, V., & MedinaFranco, J. L. (2018). Chemical space and diversity of the NuBBE database: A
chemoinformatic characterization. Journal of Chemical Information and Modeling, 59(1),
74–85.
54. Durán-Iturbide, N. A., Díaz-Eufracio, B. I., & Medina-Franco, J. L. (2020). In silico ADME/
Tox profiling of natural products: A focus on BIOFACQUIM. ACS Omega, 5(26),
16076–16084.

38 D. Q. de Azevedo et al.
55. Al Sharie, A. H., El-Elimat, T., Al Zu’bi, Y. O., Aleshawi, A. J., & Medina-Franco, J. L.
(2020). Chemical space and diversity of seaweed metabolite database (SWMD): A
cheminformatics study. Journal of Molecular Graphics and Modelling, 100, 107702.
56. O’Boyle, N. M., Banck, M., James, C. A., Morley, C., Vandermeersch, T., & Hutchison, G. R.
(2011). Open Babel: An open chemical toolbox. Journal of Cheminformatics, 3,1–14.
57. Tuerkova, A., & Zdrazil, B. (2020). A ligand-based computational drug repurposing pipeline
using KNIME and programmatic data access: Case studies for rare diseases and COVID-19.
Journal of Cheminformatics, 12(1), 1–20.
58. Banerjee, P., et al. (2015). Super natural II—A database of natural products. Nucleic Acids
Research, 43(D1), D935–D939.
59. Landrum, G. (2013). Rdkit documentation. Release, 1(1–79), 4.
60. Tandi, M., Tripathi, N., Gaur, A., Gopal, B., & Sundriyal, S. (2022). Curation and
cheminformatics analysis of a Ugi-reaction derived library (URDL) of synthetically tractable
small molecules for virtual screening application. Molecular Diversity,1–14.
61. Coghlan, A., Padalino, G., O’Boyle, N. M., Hoffmann, K. F., & Berriman, M. (2022).
Identification of anti-schistosomal, anthelmintic and anti-parasitic compounds curated and
text-mined from the scientifi c literature. Wellcome Open Research, 7.
62. da Paixão, V. G., & da Rocha Pita, S. S. (2020). Novel scaffolds for Leishmania infantum
trypanothione reductase inhibitors derived from Brazilian natural products biodiversity. Anti-
Infective Agents, 18(4), 398–418.
63. Degtyarenko, K., de Matos, P., Ennis, M., Hastings, J., Zbinden, M., McNaught, A.,
Alcántara, R., Darsow, M., Guedj, M., & Ashburner, M. (2008). ChEBI: A database and
ontology for chemical entities of biological interest. Nucleic Acids Research, 36, D344–D350.
64. Valdés-Jiménez, A., Peña-Varas, C., Borrego-Muñoz, P., Arrue, L., Alegría-Arcos, M., NourEldin, H., et al. (2021). PSC-db: A structured and searchable 3d-database for plant secondary
compounds. Molecules, 26(4), 1124.
65. Gallo, K., Kemmler, E., Goede, A., Becker, F., Dunkel, M., Preissner, R., & Banerjee,
P. (2023). SuperNatural 3.0—A database of natural products and natural product-based
derivatives. Nucleic Acids Research, 51(D1), D654–D659.
66. Nguyen-Vo, T., et al. (2018). VIETHERB: A database for Vietnamese herbal species. Journal
of Chemical Information and Modeling, 59(1), 1–9.
67. Silva, T. S. (2018). Desenvolvimento de banco de dados de pacientes submetidos ao
transplante de células-tronco hematopoiéticas. UFRS. Dissertação de Mestrado.
68. Yang, J., Wang, D., Jia, C., Wang, M., Hao, G., & Yang, G. (2019). Freely accessible chemical
database resources of compounds for in silico drug discovery. Current Medicinal Chemistry,
26, 7581–7597.
69. Ruddigkeit, L., van Deursen, R., Blum, L. C., & Reymond, J.-L. (2012). Enumeration of
166 billion organic small molecules in the chemical universe database GDB-17. Journal of
Chemical Information and Modeling, 52, 2864
70. Wang, R., Fang, X., Lu, Y., & Wang, S. (2004). The PDBbind database: Collection of binding
affinities for protein-ligand complexes with known three-dimensional structures. Journal of
Medicinal Chemistry, 47, 2977–2980.
71. Reymond, J. L., & Awale, M. (2012). Exploring chemical space for drug discovery using the
chemical universe database. ACS Chemical Neuroscience, 3(9), 649–657.
72. Vivek-Ananth, R. P., Sahoo, A. K., Kumaravel, K., Mohanraj, K., & Samal, A. (2021).
MeFSAT: A curated natural product database specific to secondary metabolites of medicinal
fungi. RSC Advances, 11, 2596–2607.
73. van Santen, J. A., Poynton, E. F., Iskakova, D., McMann, E., Alsup, T. A., Clark, T. N.,
Fergusson, C. H., Fewer, D. P., Hughes, A. H., McCadden, C. A., Parra, J., Soldatou, S.,
Rudolf, J. D., Janssen, E. M.-L., Duncan, K. R., & Linington, R. G. (2022). The natural
products atlas 2.0: A database of microbially-derived natural products. Nucleic Acids
Research, 50, D1317–D1323.
–2875.
Соседние файлы в папке Библиотека им академика М.И. Перельмана
