Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5884_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
02.09.2026
Размер:
21 Мб
Скачать
1 Echoes from the Past, Visions from the Future: A Journey into... 9
force eld is a computational equation model used to describe the forces between a collection of atoms. The force eld works as a special case of energy functions or interatomic potentials and encompasses not only the calculation form but also the parameter sets used to calculate the atoms potential energy. These models can be used to derive potential energy and forces, and by consequence acceleration and speed, within molecules or between molecules. The main reason why computational methods can be used in drug design is due to the development of fast and reasonably reliable force elds, which are computationally inexpensive when compared to QM calculations.
The work from McCammon et al. [28] highlights the approximate nature of their potential energy function as the main limitation for what was the rst MD simulation ever run. Later generalizations of these models included parameters derived from QM calculations resulting in the rst force elds (this topic is extensively discussed in Chap. 8).
1.1 Structure-Based Drug Discovery (SBDD)
SBDD is a drug desig n method that utilizes 3D structures of molecular targets and focuses on the design and optimization of a ligand that ts accurately inside the binding pocket and results in benecial target-ligand interactions [29]. Structure­based drug discovery is a rapidly growing area due to steadily increasing structural information available, arising not only from genomics and proteomics data but also from AI-augmented systems and novel machine learning approaches [30]. The SBDD process is iterative and frequently requires multiple cycles before an optimal initial hit compound can proceed to experimental validation. A key limitation in the early stages is the creation of misleading and biologically irrelevant models that can ruin the entire pipeline. Relevant biological phenomena to be considered are, for instance, different temperatures, pH, and the possibility that the target explores multimerization, as well as interaction with other proteins and macromolecules such as nucleic acids and/or membranes.
Those functional assemblies frequently lead to conformational changes in the targets substructures, which shift the size and available interactions in the binding site [31, 32]. In addition, understanding the relevant functional changes that a target undergoes upon modulation/activation is crucial to aligning the initial model repre­sentations with real-world experimental observations.
After, the target 3D structure can be obtained from databases (see Chap. 2) or, in case its crystal structure or NMR is yet to be resolved, the 3D structure can be predicted by leveraging other structures with the highest sequence similarity to the target or using machine learning (ML)-based techniques (see Chap. 4). The selection for an appropriate method ranges from homology modelling towards AI-powered approaches, such as AlphaFold, which are chosen according to the uniqueness of the target and the systems complexity.
Subsequently, the selection of the binding site determination needs to be identi­ed or predicted to facilitate further optimization of hit compounds down the line.
10 V. G. Maltarollo et al.
Those sites can be active/catalytic sites or allosterically relevant binding pockets. In some cases, binding sites are well established within a protein family, providing structural features or specic residues that orient the binding pocket generation. As an example, in protein kinases, various substructural elements can be used as landmarks towards the classical ATP binding site such as the gatekeeper, DFG-motif, or G-rich loop. In addition, many structures have co-crystallized ligands available [33].
Whether the binding site is unknown for the target family, computational methods can predict probable locations. One can choose between the methods relying on geometrical properties such as POCKET [34], PASS [35], LIGSITE [36], or combined with the physics approach such as PocketFinder [37] or SiteMap [38, 39]. It is interesting to consider that for less-dened (or transient) binding pockets, there is a chicken-and-egg situation. The presence of a high-afnity binding ligand would allow the experimental determination of that particular target confor­mation with a well-dened pocket, however, a well-dened pocket can be a require­ment for the identication of even initial hits. Computational modelling, utilizing enhanced sampling, and well-designed long timescale MD simulation studies can untwist this loop on a case-by-case basis (see Chaps. 8 and 9).
The identic ation/selection of the binding site opens many potential avenues for the SBDD. For instance, one can establish a pharmacophore model and pharmacophore-based virtual screenings using relevant amino acids, determined by interaction points with ligands, sequence conservation, or prior knowledge. Alternatively, one can carry out structure-based docking (i.e. virtual screening (VS)) of a chemical library, which is often computationally more intensive than the pharmacophore counterpart but requires no prior knowledge. Once the method and the screening database are selected, the VS campaign can be performed.
Next, the initial hit list should be scored and can also be ltered according to the properties required for project purposes. From the ligand perspective, essential pharmacokinetic properties can include physicochemical parameters (MW, numbe r of heavy atoms, hydrogen bond donors and acceptors, and rotatable bonds), lipophilicity (Log P
), water solubility (log S), pharmacokinetics, drug-likeness
o/w
(Lipinski violations), as well as an evaluation of the synthesis accessibility for medicinal chemistry [40]. From the protein-ligand perspective, docking poses are often visually inspected for relevant pharmacophoric points and interactions with key amino acids, and those poses can even be rescored/reranked (please, see Chap. 7). Data from smaller VS campaigns can in addition be utilized as input for ML-boosted docking models that will quickly screen for libraries a hundred times larger than the original data set (see Chap. 11). Once the initial hits satisfy the criteria, optional short-scale MD simulations (200–500 ns) can be performed to validate the ligand stability within the target binding pocket, resulting in selecting top hit compounds for experimental validation. Finally, the cycle between model validation and the selection of top hits is frequently repetitive and requires rene­ment, the determination of which steps can be rened/repeated greatly varies and the
1 Echoes from the Past, Visions from the Future: A Journey into... 11
diversity of shapes in those loops is often just limited by the computational resources.
1.2 Ligand-Based Drug Design (LBDD)
The premise of ligand-based drug design originates from the foundation concept of medicinal chemistry, which asserts that molecules with a high degree of structural similarity possess comparable bioactivity [41]. LBDD approaches utilize primary data on active compounds (approved drugs, published reports, etc.) to predict or generate novel drug-like compounds with similar biological effects. The inactive compounds are also benecial for the predictions, while they can be used to identify undesirable ligand features and validate the accuracy of the computational model. In this sense, the typical workow starts with an initial set of compounds with known potency that proceeds to a chosen computational method for the similarity search. A similarity search may be conducted using a wide range of molecular descriptors or lters. Molecular descriptors are categorized based on the searchs dimensionality. Molecular weight and log P are typical 1D descriptors, whereas topological indica­tions and ngerprints are 2D descriptors. The characteristics covered by 3D descrip­tors encompass a broad range of properties from electrostatic potential to the 3D geometry of ligand moieties. Furthermore, the 3D descriptors are integrated into pharmacophore modelling. According to IUPA C recommendations, a pharmacophore is an ensemble of steric and electronic features that is necessary to ensure the optimal supramolecular interactions with a specic biological target and to trigger (or block) its biological response[42]. The ideal scenario for pharmacophore modelling is when a protein-ligand complex is co-crystallized, and the ligand shows sufcient potency and bioactivity. An additional step using Quan­titative Structure-Activity Relationships (QSAR) (see Chap. 6) or machine learning (see Chap. 4) can be added to the pipeline if a comprehensive data set of active compounds is available. The most popular metric to quantify the similarity of new compounds to the initial set is the Tanimoto coefcient [43, 44].
1.3 Echoes from the Past, Visions from the Future
Considering all those contexts, opinions, and a brief historical background, we can conclude that CADD/CDDD strategies benet from a multitude of methods and approaches. In this book, computational methods will be mainly described in the rst section (General Topics and Methods) with their particularities, physicochemical and mathematical fundamentals, validation procedures, and available tools. This section aims to introduce the methods as well as update them in terms of the state-of­the-art for both beginners and experienced researchers in the drug design eld.
12 V. G. Maltarollo et al.
In addition, the second section of this book (named The pitfalls between experimentation and simulation) provides insights on maybe one of the most important questions of using CADD for targeting to put a compound into clinics: what are the limitations and problems at the interface between experimentation and simulation? In this sense, the readership is stimulated to think about Can we trust this simulation output?, Which biological questions can those methods help to answer?,or“What is the condence level to predict some experimental properties?” considering the errors and condence of predictions as well as experimental errors and variations that we very often forget about.
Lastly, a third section entitled From computer towards the clinicalintroduces several case studies describing examples of CADD employment for bioactive compound design or cases where a posteriori interpretation of the results was made possible by this technology.
As we are moving beyond databases with millions and billions of data points, scientists studying those connections are being daily ooded with new information. We believe it is important to remember that knowledge is not just data accumulation. Building knowledge takes time, effort, and willingness to work in the interfaces of interdisciplinary elds. We hope this book can help you to develop a common language to speak with scientists from your complementary elds, ask the right questions, and take inspiration to build the necessary bridges.
Acknowledgments VGM would like to thank the Fundação de Amparo à Pesquisa do Estado de Minas Gerais - FAPEMIG (grants APQ-01818-21 and RED-00110-23). TK and ES are funded by the Fortune Initiative and from TüCAD2 and CMIF. TüCAD2 and CMIF are supported by the Federal Ministry of Education and Research (BMBF) and the Baden-Württemberg Ministry of Science as part of the Excellence Strategy of the German Federal and State Governments. As well as the German Center for Infection Research (DZIF, TTU06.716).

References

1. Serturner, F. (1817). Ueber das Morphium, eine neue salzfähige Grundlage, und die Mekonsäure, als Hauptbestandtheile des Opiums. Annalen der Physik, 55,56–89.
2. Mahdi, J. G., Mahdi, A. J., Mahdi, A. J., & Bowen, I. D. (2006). The historical analysis of aspirin discovery, its relation to the willow tree and antiproliferative and anticancer potential. Cell Proliferation, 39, 147–155.
3. Riethmiller, S. (2005). From atoxyl to salvarsan: searching for the magic bullet. Chemotherapy, 51, 234–242.
4. Lloyd, N. C., Morgan, H. W., Nicholson, B. K., & Ronimus, R. S. (2005). The composition of Ehrlichs salvarsan: resolution of a century-old debate. Angewandte Chemie, International Edition, 44, 941–944.
5. Bentley, R. (2009). Different roads to discovery: Prontosil (hence sulfa drugs) and Penicillin (hence β-lactams). Journal of Industrial Microbiology & Biotechnology, 36, 775–786.
6. Jack, D. (1991). The 1990 Lilly Prize Lecture. A way of looking at agonism and antagonism: Lessons from salbutamol, salmeterol and other β-adrenoceptor agonists. British Journal of Clinical Pharmacology, 31, 501–514.
1 Echoes from the Past, Visions from the Future: A Journey into... 13
7. Lemke, T. L., Williams, D. A., Roche, V. F., & Zito, S. W. (2013). Foyes principles of medicinal chemistry (p. 1319). Lippincott Williams & Wilkins.
8. Beddell, C. R., Goodford, P. J., Norrington, F. E., Wilkinson, S., & Wootton, R. (1976). Compounds designed to t a site of known structure in human haemoglobin. British Journal of Pharmacology, 57(2), 201.
9. Protein Data Bank. PDB statistics: Overall growth of released structures per year. Available at:
https://rcsb.org/stats/growth/growth-released-structures
10. Burley, S. K., Berman, H. M., Kleywegt, G. J., Markley, J. L., Nakamura, H., & Velankar, S. (2017). Protein Data Bank (PDB): the single global macromolecular structure archive. In Protein crystallography: methods and protocols (pp. 627–641).
11. Protein Data Bank. (1971). Crystallography: Protein Data Bank. Nature: New Biology, 233,
223.
12. Bartusiak, M. (1981, October 5) Designing drugs with computers. Fortune, 47–50.
13. Frye, L., Bhat, S., Akinsanya, K., & Abel, R. (2021). From computer-aided drug discovery to computer-driven drug discovery. Drug Discovery Today: Technologies, 39, 111–117.
14. Fischer, E. (1894). The inuence of conguration on enzyme activity. Berichte der Deutschen Chemischen Gesellschaft, 27, 2984–2993. (Translated from German).
15. Schneider, H.-J. (2003). Introduction to molecular recognition models. In Protein-Ligand interactions (pp. 21–50).
16. Heaven, W. D. (2023, March/April). AI is dreaming up drugs that no one has ever seen. Now weve got to see if they work MIT Technology Review. Available at: https://www.
technologyreview.com/2023/02/15/1067904/ai-automation-drug-development/
17. Metz, C. (2019, February 5). Making new drugs with a dose of articial intelligence. The New York Times. Available at: https://www.nytimes.com/2019/02/05/technology/articial-
intelligence-drug-research-deepmind.html
18. Arnold, C. (2023). Inside the nascent industry of AI-designed drugs. Nature Medicine, 29(6), 1292–1295.
19. Belleau, B. (1970). Rational drug design: mirage or miracle? Canadian Medical Association Journal, 103(8), 850.
20. Bender, A., & Cortés-Ciriano, I. (2021). Articial intelligence in drug discovery: what is realistic, what are illusions? Part 1: Ways to make an impact, and why we are not there yet. Drug Discovery Today, 26(2), 511–524.
21. Prieto-Martinez, F. D., Lopez-Lopez, E., Euridice Juarez-Mercado, K., & Medina-Franco, J. L. (2019). Chapter 2 - Computational drug design methodsCurrent and future perspectives. In K. Roy (Ed.), In Silico drug design (pp. 19–44). Academic Press.
22. Zhao, L., Ciallella, H. L., Aleksunes, L. M., & Zhu, H. (2020). Advancing computer-aided drug discovery (CADD) by big data and data-driven machine learning modeling. Drug Discovery Today, 25(9), 1624
23. Vemula, D., Jayasurya, P., Sushmitha, V., Kumar, Y. N., & Bhandari, V. (2023). CADD, AI and ML in drug discovery: A comprehensive review. European Journal of Pharmaceutical Sciences, 181, 106324.
24. Chandershekar, A., Bhaskar, A., Mekkanti, M. R., & Rinku, M. (2020). A review on computer aided drug design (CAAD) and its implications in drug discovery and development process. International Journal of Health Care and Biological Sciences,27–33.
25. Osakwe, O. (2016). Chapter 5 - The signicance of discovery screening and structure optimi­zation studies. In O. Osakwe & S. A. A. Rizvi (Eds.), Social aspects of drug discovery, development and commercialization (pp. 109–128). Academic Press.
26. dos Santos Nascimento, I. J., & de Moura, R. O. (2023). Ligand and structure-based drug design (LBDD and SBDD): Promising approaches to discover new drugs. Applied Computer-Aided Drug Design: Models and Methods,1.
27. Hassan Baig, M., Ahmad, K., Roy, S., Mohammad Ashraf, J., Adil, M., Haris Siddiqui, M., et al. (2016). Computer aided drug design: Success and limitations. Current Pharmaceutical
Design, 22(5), 572–581.
1638.
14 V. G. Maltarollo et al.
28. McCammon, J. A., Gelin, B. R., & Karplus, M. (1977). Dynamics of folded proteins. Nature, 267(5612), 585–590.
29. Klebe, G. (2013). Protein modeling and structure-based drug design. In G. Klebe (Ed.), Drug design: Methodology, concepts, and mode-of-action (pp. 429–448). Springer.
30. Jumper, J., et al. (2021). Highly accurate protein structure prediction with AlphaFold. Nature, 596, 583–589.
31. Maltarollo, V. G., Shevchenko, E., Lima, I. D. D. M., Cino, E. A., Ferreira, G. M., Poso, A., & Kronenberger, T. (2022). Do go chasing waterfalls: Enoyl reductase (fabi) in complex with inhibitors stabilizes the tetrameric structure and opens water channels. Journal of Chemical Information and Modeling, 62(22), 5746–5761
32. DíazHolguín, A., Rashidian, A., Pijnenburg, D., Monteiro Ferreira, G., Stefela, A., Kaspar, M., ... & Kronenberger, T. (2023). When two become one: Conformational changes in FXR/RXR heterodimers bound to steroidal antagonists. ChemMedChem, 18(4), e202200556
33. Martins, D. M., Fernandes, P. O., Vieira, L. A., Maltarollo, V. G., & Moraes, A. H. (2024). Structureguided drug design targeting abl kinase: How structure and regulation can assist in designing new drugs. ChemBioChem, e202400296.
34. Levitt, D. G., & Banaszak, L. J. (1992). POCKET: A computer graphics method for identifying and displaying protein cavities and their surrounding amino acids. Journal of Molecular Graphics, 10, 229–234.
35. Brady, G. P., & Stouten, P. F. (2000). Fast prediction and visualization of protein binding pockets with PASS. Journal of Computer-Aided Molecular Design, 14, 383–401.
36. Hendlich, M., Rippmann, F., & Barnickel, G. (1997). LIGSITE: automatic and efcient detection of potential small molecule-binding sites in proteins. Journal of Molecular Graphics & Modelling, 15, 359–363, 389.
37. Huang, B., & Schroeder, M. (2006). LIGSITEcsc: Predicting ligand binding sites using the Connolly surface and degree of conservation. BMC Structural Biology, 6, 19.
38. Halgren, T. (2007). New method for fast and accurate bindingsite identication and analysis. Chemical Biology & Drug Design, 69, 146–148.
39. Halgren, T. A. (2009). Identifying and characterizing binding sites and assessing druggability. Journal of Chemical Information and Modeling, 49, 377–389.
40. Opo, F. A. D. M., et al. (2021). Structure based pharmacophore modeling, virtual screening, molecular docking and ADMET approaches for identication of natural anti-cancer agents targeting XIAP protein. Scienti c Reports, 11, 4049.
41. Martin, Y. C., Kofron, J. L., & Traphagen, L. M. (2002). Do structurally similar molecules have similar biological activity? Journal of Medicinal Chemistry, 45, 4350–4358.
42. Wermuth, C. G., Ganellin, C. R., Lindberg, P., & Mitscher, L. A. (1998). Glossary of terms used in medicinal chemistry (IUPAC Recommendations 1998). Pure and Applied Chemistry, 70,
–1143.
1129
43. Bajusz, D., Racz, A., & Heberger, K. (2015). Why is Tanimoto index an appropriate choice for ngerprint-based similarity calculations? Journal of Cheminformatics, 7, 20.
44. Racz, A., Bajusz, D., & Heberger, K. (2018). Life beyond the Tanimoto coefcient: Similarity measures for interaction ngerprints. Journal of Cheminformatics, 10, 48.
Chapter 2
Molecular Databases
Daniela Quadros de Azev edo, Rache l Oliveira Castilho, Alejandro Gómez-García, and José L. Medina-Franco
Abstract Compound databases (DBs) aim to organize the information needed for
the initial stages of drug discovery. The collection and organization of curated information from bibliographic searches in DBs helps the scientic community to develop multidisciplinary research areas. For example, the chemical and biological properties contained in compound libraries, including PubChem, ZINC, BindingDB, and ChEMBL, are widely used in drug discovery projects. The importance of DBs in these projects is continuously increasing, beyond their role as compound reposito­ries. In fact, compound DBs and chemical datasets can be a centerpiece in pharma­ceutical companies, as well as in academic and government research centers. Several research groups have recently used computational methodologies to screen large DBs of compounds before experimental screening and designing their experiments. The number of DB compounds in the public domain, including those for compounds of natural origin, is increasing. This is in line with the growing and synergistic combination of natural product research and chemoinformatics, for example, NuBBE ucts) and BIOFACQUIM (A Mexican Compound Database of Natural Products). Developing a database involves four steps. Step 1Search for chemical or biolog­ical information in indexed databases that report the strategies used to nd the information that will make up the DBs. Step 2Curation: describes the database (DB) curation processes, automated or manual. Step 3DB management and network visualization: describes different systems to process DBs. Step 4Update and maintenance: The last and most important stage is that, in addition to the development of a DB, its maintenance requires various resources, both nancial and human. The quality of the compound DBs is crucial to fulll properly their role
(Nuclei of Bioassays, Ecophysiology and Biosynthesis of Natural Prod-
DB
D. Q. de Azevedo · R. O. Castilho Departamento de Produtos Farmacêuticos, Faculdade de Farmácia, UFMG, Belo Horizonte, Minas Gerais, Brazil
A. Gómez-García · J. L. Medina-Franco ( DIFACQUIM Research Group, Department of Pharmacy, National Autonomous University of Mexico, Mexico City, Mexico e-mail: medinajl@unam.mx
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2024 V. G. Maltarollo (ed.), Computer-Aided and Machine Learning-Driven Drug Design, Computer-Aided Drug Discovery and Design 3,
https://doi.org/10.1007/978-3-031-76718-0_2
✉)
15
16 D. Q. de Azevedo et al.
as drug design tools. For automated curation, some tools can be used, such as KNIME and Open Babel. Currently, and in the near future, these virtual libraries could also enrich the medically relevant chemical space for inhibitors and potential drug candidates. In addition, the incorporation of chemoinformatics tools, the automation of curation processes, and the allocation of nancial and human resources for the maintenance of DBs will contribute signicantly to the develop­ment of drug design discovery and development projects.
Keywords Molecular database · Drug discovery · Chemoinformatic · Virtual screening

1 Introduction

Until the middle of the nineteenth century, the medicines used by humans were mainly of natural origin (animal, vegetable, and/or mineral) [1, 2] and their active principles (drugs) were compounds isolated from natural plant products and second­ary metabolites of microorganisms [ 2 , 3]. Examples include morphine (extracted from the poppy, Papaver somniferum), quinine (extracted from Cinchona species), atropine (extracted from Atropa belladonna), and salicin (extracted from Salix alba). In 1860, Kolbe and Lauteman synthesized salicylic acid, and in 1888, Hofman synthesized acetylsalicylic acid, supporting the arising of the pharmaceutical indus­try began. Consequently, in the twentieth century, the era of synthesis began with the development and production of drugs such as barbital, phenobarbital, epinephrine, amylocaine, and procaine [4]. Arguably, the discovery of penicillin by Alexander Fleming in 1928 sparked a renewed interest in natural products as sources of inspiration for drug development. In this regard, sulfanilide, which was also widely employed during World War, was synthesized in 1930. This led to subsequent breakthroughs, in the so-called golden ageof antibiotics such as actinomycin (1942), streptomycin (1943), chloramphenicol (1947), and neomycin and erythro­mycin A in 1949 and 1950, respectively. Finally, the 1960s included rifamycin, vancomycin, lincomycin, and amphotericin B [5].
The accelerated pace of drug discovery during this period was partly due to advances in biology and chemistry, apart from the serendipitous discovery of drugs, such as penicillin. The resulting effects were signicant in improving the quality of life and increased longevity [6]. Most human diseases have available therapeutic options, such as diabetes, Alzheimers disease, immun ological disorders, human immunodeciency virus (HIV) and its associated acquired immune deciency syndrome (HIV) [7], neglected tropical diseases (NTDs), and rare diseases [8]. How­ever, new drugs are important when considering new diseases, the development of drug resistance, and our increasing understanding of health conditions allowing for the treatment of previously untreatable conditions [9].
In this sense, several strategic approaches have been proposed to increase ef­ciency in the drug discovery and development process, which can be adopted and exploited in a variety of ways. The se approaches include genomics and proteomics
2 Molecular Databases 17
(details in Chap. 3), complementary phenotyping and target-based screening plat­forms, drug repurposing and repositioning, collaborative research between industry and academia, outsourcing, molecular modeling, and articial intelligence (AI) (details in Chap. 4)[6].
Modeling involves the use of in silico simulations to predict various properties of a compound, including pharmacokinetic and pharmacodynamic proles (more details in Chap. 13)[10]. Advances in computer science have enabled the develop- ment of software that allows the simulation of drug-receptor binding processes, as well as a subset of processes known as virtual screening (VS), which increase the efciency of hit identication from small and large existing chemical databases (DBs). VS is a process that generates results that can be validated in vitro and guides the optimization of lead compounds, improving the afnity of the drug to the receptor and its pharmacokinetic properties. VS also facilitates rational drug design by generating new compounds to optimize receptor binding [11]. For more details on virtual screening, refer to Chap. 11.
Moreover, AI is also being increasingly applied to drug design and development. This is possible by the availability of large chemical and biological DBs, which are today an important strategy for the development of accurate predictive models. AI has the potential to revolutionize drug discovery, particularly for unmet clinical needs, by enabling the screen ing of potential compounds with condent identica­tion and validation of biological targets [11]. For more details on AI and machine learning, refer to Chap. 4.
In the period 2020 to 2022, some compounds have been developed exclusively by AI. As most are in clinical trials, they are described by code. DSP-1181 is a potent, long-acting 5-HT
serotonin receptor agonist with indications for the treatment of
1A
obsessive-compulsive disorder; DSP-0038 has the function of treating Alzheimers disease and psychosis; EXS-21546 is an immunological compound with indications for the treatment of various types of tumors [12].
Currently, DBs may combine medicinal chemistry know ledge with important AI approaches and applications, providing not only quality content but also the users ability to interact with this data and integrate it more easily. However, it is worth mentioning that even if these tools are not included in DBs, they help the drug discovery process. ZINC [13] and PubChem [14] are examples of DBs that have the potential to be used to capture data for the generation of predictive models using machine learning. Another database (DB) that includes AI tools is the latest version of ChEMBL (version 32) [15], which incorporates tools for calculating physico­chemical properties of pharmaceutical interest.
Examples of software used in pharmaceutical modeling and AI-driven drug discovery are listed in Table 2.1.
Drug DBs aim to organize the information needed for the initial stages of drug discovery. The collection and organization of curated information from biblio­graphic searches in DBs help the scientic community to develop multidisciplinary research areas such as drug discovery, medicinal chemistry, chemosystematics, ethnopharmacology, and omicsapproaches. For example, many genomic studies, such as human genome mapping, have used GenBank and the DNA Databank of
18 D. Q. de Azevedo et al.
Table 2.1 Examples of pharmaceutical modeling and articial intelligence-guided drug discovery tools
Tool Application in drug discovery Link
AlphaFold AutoDock DeepTox DiscoveryStudio
Glide
PPB2 PotentialNet
DataWarrior
a
Tools that support drug design
b
Tools that support drug design usually included in DBs
a
a
b
a
b
Target modeling https://alphafold.ebi.ac.uk/ Protein ligand-binding modeling http://autodock.scripps.edu/ Predictive toxicology https://deeptox.co
b
Pharmacophore modeling, target identi­cation, lead optimization
Combinatorial chemistry and docking studies
Predictive pharmacology http://ppb2.gdb.tools/
a
Protein-ligand binding and molecular properties modeling
b
Prediction of physicochemical proper­ties of pharmaceutical interest, cheminformatics calculations, multivar­iate data analysis, and interactive visu­alization with dynamic plots
https://www.3dsbiovia.com/prod ucts/collaborative-science/ bioviadiscovery-studio/
https://www.schrodinger.com/ glide
https://pubs.acs.org/doi/ full/10.1021/acscentsci.8b00507
https://openmolecules.org/ datawarrior/
Japan (DDBJ) [16]. Many pharmacological, computational, and proteomic studies have employed the Protein Data Bank (PDB) [17], the Human Proteome Map, and the Peptide Atlas [18]. In addition, several metabolomic studies have relied on the Human Metabolome DB (HMDB) [19], Golm Metabolome DB [20], Global Natural Product Social Molecular Networking (GNPS) [21], Biological Magnetic Resonance Bank (BMRB) [22], and Mass Bank [23]. In addition, the chemical and biological properties contained in compound libraries, including PubChem [14], ChemSpider [24], ZINC [13], PK/DB [25], BindingDB [26], ChemBank [ 27 ], and ChEMBL [15], are widely used in drug discovery projects.
The importance of DBs in new drug discovery projects is continuously increas­ing, beyond their role as compounds repositories. In fact, compound DBs and chemical datasets can be a centerpiece in pharmaceutical companies, as well as in academic and government research centers [28]. Compounds present in DBs have already led to the development of drugs in clinical use to treat different diseases. Several research groups have recently used computational methodologies to screen large DBs of compounds before experimental screening and designing their experiments [29].
The number of DB compounds in the public domain, including those for com­pounds of natural origin, is increasing . This is in line with the growing and synergistic combination of natural product research and chemoinformatics [30], for example, NuBBE
(Nuclei of Bioassays, Ecophysiology and Biosynthesis of
DB
Natural Products) [31], BIOFACQUIM (A Mexican Compound Database of Natural Products) [32], and NPACT (Naturally occurring Plant-based Anticancerous Compound-Activity-Target Database) [33]. In addition to the natural products