Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5884_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
02.09.2026
Размер:
21 Мб
Скачать
152 P. O. Fernandes and V. G. Maltarollo
A virtual screening comprising multiple computational methods was reported by Mushtaq and colleagues [124] to identify compounds with anti-interleukin-2 activ­ity. The authors validated a molecular docking protocol and applied it in combina­tion with pharmacophore ltering starting from a library containing 11.9 million compounds from ZINC and resulting in 24 compounds that had their activities predicted with a properly validated CoMFA model. Then, only nine compounds were submitted to experimental validation and three of them were considered promising hits due to the IL-2 inhibitory effect.
An interesting application of machine-learning-based virtual screening was reported by Barbosa et al. [125] In this work, they generated and validated kNN and Random Forest models to predi ct activity against Trypanosoma cruzi and used the models to screen a natural products database. After the selection of a virtual hit, they identied a plant that produced this hit as a metabolite (Cymbopogon schoenanthus), performed the isolation and identication of the selected compound, a diterpenoid called andrographolide, and tested it. As a result of the experimental validation, this compound showed IC
values of 29.4 and 2.9 μM against
50
trypomastigote and amastigote forms of T. cruzi and a selectivity index equal to 32, comparable with the positive control.
In 2024, Fernandes and colleagues reported a machine learning-based virtual screening [126] or discovering anti bacterial compounds against methicillin­susceptible and resistant strains of Staphylococcus aureus. In this work, they gen­erated descriptor-based QSAR classication models using several diverse machine learning methods for three data sets : one comprised of compounds with activity against susceptible strains of S. aureus; another one comprised of compounds with activity against resistant strains; and a third one comprised of compounds with activity against both strains. This last data set was very important to weigh the consensus selection of hits for experiment al testing since it comprised both modeled activities. In this sense, the hit rate of models generated from this data set was higher than the other two models.
Wong et al. [127] experimentally screened 39,312 compounds against a methicillin-susceptible S. aureus strain (RN4220) as a model for antibacter ial activity and against human liver carcinoma cells (HepG2), human primary skeletal muscle cells (HSkMCs), and human lung broblast cells (IMR-90) as models for cytotoxicity. After, they used Chemprop to train a graph neural network to predict a binary classication task for all properties (activity and toxicities). After training and validation of models, the authors screened two libraries for obtaining compounds with predicted antibacterial activity and no cytotoxic prole. Following this, poten­tial false-positive compounds were removed using PAINS rules and undesired compounds due to reactivity, metabolic instability, and generalized toxicity were also removed by using Brenk structural alerts. Lastly, the authors selected com­pounds with similarity scores equal or lower than 0.5 in comparison with data set compounds. In this sense, they started with approximately 12 million compounds, and, after ltering, they selected 1261 compounds. As a strategy to interpreting the models, the authors used Monte Carlo tree searches to explore the chemical space and understand the smallest portion of a molecule responsible for their classication
6 QSAR and Machine Learning Predictors 153
as active. With this approach, it was possible to highlight important structural features responsible for predictions. Then, using the rational structural predictions they ltered the hits according to analogs that match known antibacterial classes such as quinolone, cephalosporins, and β-lactams, then selecting 9 compounds for experimental validations. Four of the nine tested compounds were actives reaching a 44% success rate.
This last example is not an actual application of QSAR for drug design discovery, but a free platform available for early stages ADME prole prediction [128]. The Pharmacokinetics Proler (PhaKinPro, available at: https://phakinpro .mml.unc.edu/ ) is a web server with QSAR models to predict hepatic stability, microsomal half-life in sub-cellular and tissue, renal clearance, blood–brain barrier (BBB) permeability, central nervous system (CNS) activity, Caco-2 permeability, plasma protein binding, plasma half-life, microsomal intrinsic clearance, and oral bioavailability. Despite the validation metrics of each model, the predictions at the web server provide the condence of prediction, if the compound is inside of AD or not, and the importance of molecular fragments for the predictio n as the interpretation of the model.

8 Challenges and Perspectives

As mentioned, statistical methods and algorithms have a poor ability to handle data from different sources and, consecutively, biological data from different protocols. In that sense, ML methods emerge as promising methods to model data with noise due to the superior ability of generalization. Indeed, this task has naturally evolved since biology and chemistry joined the big data concept with large databases. However, it desired methods to ensure how the data noise affects the quality of the predictions.
Another related issue is better attention and report of imbalanced data sets for classication models and potential gaps in modeled activities for regression models. Very often, authors underreport this aspect of data sets and, of course, biased data sets produce biased models which make biased predictions. The balance between classes introduces a bias in the modeled activity. However, other sources of bias must be avoided and well reported in QSAR modeling protocols, such as structural and physicochemical biases which could be solved (or, at least, used to warn potential users of reported models) with the combination of data set characterization and proper AD denition.
Multi-task models, in other words, models with the ability to predict more than one property ( y) are well established in the literature; however, authors very often report parallel individual models. Maybe, the lack of ready-to-use software with this ability and the need for coding limit the spreading of this modality of modeling. In the same way, transfer learning is a set of methods that transfer knowledge from one trained and validated model to another. Of course, both models must share mecha­nisms or similarities in the modeled property, for example, models to predict the binding afnity of ligands to close homolog and structurally similar enzymes. This
154 P. O. Fernandes and V. G. Maltarollo
concept of transfer learning is extremely useful since it saves enormous amounts of time in training steps for similar task models. But, as multi-task models, transfer learning in QSAR is underreported in comparison to regular standard QSAR models.
Finally, not only the QSAR eld, but all subjects across chemistry should feel the impact of large language models (LLM) because now they have been closing the gap between machines and humans across many knowledge domains [129] and its potential already has been demonstrated to understand complex molecular distribu­tion [130]. The recent interest of the worldwide community on language models such as Generative Pre-trained Transformer 3 (GPT-3) from OpenAI popularized with the ChatGPT API has shed light on the subject. In this sense, LLM has been explored to predict chemical tasks as well as is in the spotlight of novel methods to be explored.
Acknowledgements The authors would like to thank the Fundação de Amparo à Pesquisa do Estado de Minas Gerais - FAPEMIG (grants APQ-01818-21 and RED-00110-23).

References

1. Cros, A. (1863). Action de lalcohol Amylique Sur lorganisme (PhD Thesis, Thesis). Stras­bourg, University of Strasbourg.
2. Mills, E. J. (1884). XXIII. On melting-point and boiling-point as related to chemical compo­sition. The London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science, 17(105), 173–187.
3. Hansch, C., Maloney, P. P., Fujita, T., & Muir, R. M. (1962). Correlation of biological activity of phenoxyacetic acids with Hammett substituent constants and partition coef cients. Nature, 194(4824), 178–180.
4. Hansch, C., Muir, R. M., Fujita, T., Maloney, P. P., Geiger, F., & Streich, M. (1963). The correlation of biological Activity of plant growth regulators and Chloromycetin derivatives with Hammett constants and partition coefcients. Journal of the American Chemical Society, 85(18), 2817–2824.
5. Fujita, T., Iwasa, J., & Hansch, C. (1964). A new substituent constant, π, derived from partition coefcients. Journal of the American Chemical Society, 86(23), 5175–5180.
6. Soares, T. A., Nunes-Alves, A., Mazzolari, A., Ruggiu, F., Wei, G.-W., & Merz, K. (2022). The (Re)-Evolution of Quantitative Structure–Activity Relationship (QSAR) studies propelled by the surge of machine learning methods. Journal of Chemical Information and Modeling, 62(22), 5317–5320.
7. Shi, Y., Yu, M., Liu, J., Yan, F., Luo, Z.-H., & Zhou, Y.-N. (2022). Quantitative structure– property relationship model for predicting the propagation rate coefcient in free-radical polymerization. Macromolecules, 55(21), 9397–9410.
8. Yu, M., Shi, Y., Liu, X., Jia, Q., Wang, Q., Luo, Z.-H., Yan, F., & Zhou, Y.-N. (2023). Quantitative Structure-Property Relationship (QSPR) framework assists in rapid mining of highly thermostable polyimides. Chemical Engineering Journal, 465, 142768.
9. Czub, N., Szlęk, J., Pacławski, A., Klimończyk, K., Puccetti, M., & Mendyk, A. (2023). Articial intelligence-based quantitative structure– property relationship model for predicting human intestinal absorption of compounds with serotonergic activity. Molecular Pharmaceutics, 20(5), 2545–2555.
6 QSAR and Machine Learning Predictors 155
10. Sterling, A. J., Zavitsanou, S., Ford, J., & Duarte, F. (2021). Selectivity in organocatalysis From qualitative to quantitative predictive models. WIREs Computational Molecular Science, 11(5), e1518.
11. Eckhoff, M., Diedrich, J. V., Mücke, M., & Proppe, J. (2024). Quantitative structure–reactivity relationships for synthesis planning: The benzhydrylium case. The Journal of Physical Chemistry. A, 128(1), 343–354.
12. Paradies, J. (2023). Structure-reactivity relationships in borane-based FLP-catalyzed hydro­genations, dehydrogenations, and cycloisomerizations. Accounts of Chemical Research, 56(7), 821–834.
13. Wang, L.-L., Ding, J.-J., Pan, L., Fu, L., Tian, J.-H., Cao, D.-S., Jiang, H., & Ding, X.-Q. (2021). Quantitative structure-toxicity relationship model for acute toxicity of organophos­phates via multiple administration routes in rats and mice. Journal of Hazardous Materials, 401, 123724.
14. Mukherjee, R. K., Kumar, V., & Roy, K. (2022). Ecotoxicological QSTR and QSTTR modeling for the prediction of acute oral toxicity of pesticides against multiple avian species. Environmental Science & Technology, 56(1), 335–348.
15. Rai, M., Paudel, N., Sakhrie, M., Gemmati, D., Khan, I. A., Tisato, V., Kanase, A., Schulz, A., & Singh, A. V. (2023). Perspective on quantitative structure–toxicity relationship (QSTR) models to predict hepatic biotransformation of xenobiotics. Liver, 3(3), 448–462.
16. Fourches, D., Muratov, E., & Tropsha, A. (2010). Trust, but Verify: On the importance of chemical Structure curation in cheminformatics and QSAR modeling research. Journal of Chemical Information and Modeling, 50(7), 1189–1204.
17. Fourches, D., Muratov, E., & Tropsha, A. (2016). Trust, but Verify II: A practical guide to chemogenomics data curation. Journal of Chemical Information and Modeling, 56(7), 1243–1252.
18. McGibbon, M., Shave, S., Dong, J., Gao, Y., Houston, D. R., Xie, J., Yang, Y., Schwaller, P., & Blay, V. (2024). From intuition to AI: Evolution of small molecule representations in drug discovery. Briengs in Bioinformatics, 25(1), bbad422.
19. Khan, A. U. (2016). Descriptors and their selection methods in QSAR analysis: Paradigm for drug design. Drug Discovery Today, 21(8), 1291–1302.
20. Cereto-Massagué, A., Ojeda, M. J., Valls, C., Mulero, M., Garcia-Vallvé, S., & Pujadas, G. (2015). Molecular ngerprint similarity search in virtual screening. Methods, 71,58–63.
21. Verma, J., Khedkar, V. M., & Coutinho, E. C. (2010). 3D-QSAR in drug design - A review. Current Topics in Medicinal Chemistry, 10(1), 95–115.
22. G. Damale, M., N. Harke, S., A. Kalam Khan, F., B. Shinde, D., & N. Sangshetti, J. (2014) Recent advances in multidimensional QSAR (4D-6D): A critical review. Mini Reviews in Medicinal Chemistry, 14(1), 35–55.
23. Polanski, J. (2009). Receptor dependent multidimensional QSAR for modeling drug-receptor interactions. Current Medicinal Chemistry, 16(25), 3243–3257.
24. Tropsha, A. (2010). Best practices for QSAR model development, validation, and exploitation. Molecular Informatics, 29(6–7), 476
25. Gaulton, A., Hersey, A., Nowotka, M., Bento, A. P., Chambers, J., Mendez, D., Mutowo, P., Atkinson, F., Bellis, L. J., Cibrián-Uhalte, E., Davies, M., Dedman, N., Karlsson, A., Magariños, M. P., Overington, J. P., Papadatos, G., Smit, I., & Leach, A. R. (2017). The ChEMBL database in 2017. Nucleic Acids Research, 45(D1), D945–D954.
26. Wang, Y., Xiao, J., Suzek, T. O., Zhang, J., Wang, J., Zhou, Z., Han, L., Karapetyan, K., Dracheva, S., Shoemaker, B. A., Bolton, E., Gindulyte, A., & Bryant, S. H. (2012). PubChems BioAssay database. Nucleic Acids Research, 40(D1), D400–D412.
27. Organisation for Economic Co-operation and Development. OECD principles for the valida- tion, for regulatory purposes, of (Quantitative) structure-activity relationship models. https://
www.oecd.org/chemicalsafety/risk-assessment/37849783.pdf. Accessed 2021-12-07.
28. Di Paolo, T. (1978). Structure-activity relationships of anesthetic ethers using molecular connectivity. Journal of Pharmaceutical Sciences, 67(4), 564–566.
–488.
156 P. O. Fernandes and V. G. Maltarollo
29. Hansch, C., & Klein, T. E. (1986). Molecular graphics and QSAR in the study of enzyme­ligand interactions. On the denition of bioreceptors. Accounts of Chemical Research, 19(12), 392–400.
30. Cramer, R. D., Patterson, D. E., & Bunce, J. D. (1988). Comparative Molecular Field Analysis (CoMFA). 1. Effect of shape on binding of steroids to carrier proteins. Journal of the American Chemical Society, 110(18), 5959–5967.
31. Clark, M., Cramer, R. D., Jones, D. M., Patterson, D. E., & Simeroth, P. E. (1990). Compar­ative Molecular Field Analysis (CoMFA). 2. Toward its use with 3D-structural databases. Tetrahedron Computer Methodology, 3(1), 47–59.
32. Klebe, G., Abraham, U., & Mietzner, T. (1994). Molecular similarity indices in a comparative analysis (CoMSIA) of drug molecules to correlate and predict their biological Activity. Journal of Medicinal Chemistry, 37(24), 4130–4146.
33. Lowis, D. R. (1997). HQSAR: A new, highly predictive QSAR technique. Tripos Technical Notes, 1(5), 17.
34. Seel, M., Turner, D. B., & Willett, P. (1999). Effect of parameter variations on the effective­ness of HQSAR analyses. Quantitative Structure-Activity Relationships, 18(3), 245–252.
35. Abdizadeh, R., Hadizadeh, F., & Abdizadeh, T. (2020). QSAR analysis of Coumarin-based Benzamides as histone deacetylase inhibitors using CoMFA, CoMSIA and HQSAR methods. Journal of Molecular Structure, 1199, 126961.
36. Ding, H., Xing, F., Zou, L., & Zhao, L. (2024). QSAR analysis of VEGFR-2 inhibitors based on machine learning, Topomer CoMFA and molecule docking. BMC Chemistry, 18(1), 59.
37. Edache, E. I., Uzairu, A., Mamza, P. A., Shallangwa, G. A., & Ibrahim, M. T. (2024). Design of some potent non-toxic autoimmune disorder inhibitors Based on 2D-QSAR, CoMFA, molecular docking, and molecular dynamics investigations. Intelligent Pharmacy.
38. Abdizadeh, R., Hadizadeh, F., & Abdizadeh, T. (2020). Molecular modeling studies of anti­Alzheimer agents by QSAR, molecular docking and molecular dynamics simulations tech­niques. Medicinal Chemistry, 16(7), 903–927.
39. Lino, C. I., Gonçalves de Souza, I., Borelli, B. M., Silvério Matos, T. T., Santos Teixeira, I. N., Ramos, J. P., Maria de Souza Fagundes, E., de Oliveira Fernandes, P., Maltarollo, V. G., Johann, S., & de Oliveira, R. B. (2018). Synthesis, molecular modeling studies and evaluation of antifungal activity of a novel series of Thiazole derivatives. European Journal of Medicinal Chemistry, 151, 248–260.
40. Veríssimo, G. C., Menezes Dutra, E. F., Teotonio Dias, A. L., de Oliveira Fernandes, P., Kronenberger, T., Gomes, M. A., & Maltarollo, V. G. (2019). HQSAR and random forest­based QSAR models for Anti-T. Vaginalis activities of nitroimidazoles derivatives. Journal of Molecular Graphics and Modelling, 90, 180–191.
41. Wu, Z., Zhu, M., Kang, Y., Leung, E. L.-H., Lei, T., Shen, C., Jiang, D., Wang, Z., Cao, D., & Hou, T. (2021). Do we need different machine learning algorithms for QSAR modeling? A comprehensive assessment of 16 machine learning algorithms on 14 QSAR data sets. Briengs in Bioinformatics, 22(4), bbaa321.
42. Brown, F. K., Sherer, E. C., Johnson, S. A., Holloway, M. K., & Sherborne, B. S. (2017). The evolution of drug design at Merck Research Laboratories. Journal of Computer-Aided Molec- ular Design, 31(3), 255–266.
43. Loyola-Gonzalez, O. (2019). Black-box vs White-box: Understanding their advantages and weaknesses from a practical point of view. IEEE Access, 7, 154096–154113.
44. Barber, C., Heghes, C., & Johnston, L. (2024). A framework to support the application of the OECD Guidance documents on (Q)SAR model validation and prediction assessment for regulatory decisions. Computational Toxicology, 30, 100305.
45. Organisation for Economic Co-operation and Development. (2007). Guidance document on the validation of (quantitative) structure-activity relationship [(Q) SAR] models. Organisation for Economic Co-operation and Development.
46. Cronin, M. T. D., & Schultz, T. W. (2003). Pitfalls in QSAR. Journal of Molecular Structure: THEOCHEM, 622(1), 39–51.
6 QSAR and Machine Learning Predictors 157
47. Borota, A., Mracec, M., Gruia, A., Rad-Curpăn, R., Ostopovici-Halip, L., & Mracec, M. (2011). A QSAR study using MTD method and dragon descriptors for a series of selective ligands of α2C adrenoceptor. European Journal of Medicinal Chemistry, 46(3), 877– 884.
48. Rodríguez-Pérez, R., & Bajorath, J. (2021). Explainable machine learning for property pre­dictions in compound optimization. Journal of Medicinal Chemistry, 64(24), 17744–17752.
49. Jouan-Rimbaud, D., Massart, D. L., & De Noord, O. E. (1996). Random correlation in variable selection for multivariate calibration with a genetic algorithm. Chemometrics and Intelligent Laboratory Systems, 35(2), 213–220.
50. Hawkins, D. M. (2004). The problem of overtting. Journal of Chemical Information and Computer Sciences, 44(1), 1–12.
51. Topliss, J. G., & Edwards, R. P. (1979). Chance factors in studies of Quantitative Structure­Activity Relationships. Journal of Medicinal Chemistry, 22(10), 1238–1244.
52. Wold, S., & Dunn, W. J. (1983). Multivariate Quantitative Structure-Activity Relationships (QSAR): Conditions for their applicability. Journal of Chemical Information and Computer Sciences, 23(1), 6–13.
53. Clark, M., & Cramer, R. D., III. (1993). The probability of chance correlation using partial least squares (PLS). Quantitative Structure-Activity Relationships, 12 (2), 137–145.
54. Schaper, K.-J., Kunz, B., & Raevsky, O. A. (2003). Analysis of water solubility data on the basis of HYBOT descriptors. QSAR & Combinatorial Science, 22(9–10), 943–958.
55. Mendez, D., Gaulton, A., Bento, A. P., Chambers, J., De Veij, M., Félix, E., Magariños, M. P., Mosquera, J. F., Mutowo, P., Nowotka, M., Gordillo-Marañón, M., Hunter, F., Junco, L., Mugumbate, G., Rodriguez-Lopez, M., Atkinson, F., Bosc, N., Radoux, C. J., Segura-Cabrera, A., Hersey, A., & Leach, A. R. (2019). ChEMBL: Towards direct deposition of bioassay data. Nucleic Acids Research, 47(D1), D930–D940.
56. Kim, S., Chen, J., Cheng, T., Gindulyte, A., He, J., He, S., Li, Q., Shoemaker, B. A., Thiessen, P. A., Yu, B., Zaslavsky, L., Zhang, J., & Bolton, E. E. (2023). PubChem 2023 update. Nucleic Acids Research, 51(D1), D1373–D1380.
57. Gilson, M. K., Liu, T., Baitaluk, M., Nicola, G., Hwang, L., & Chong, J. (2016). BindingDB in 2015: A public database for medicinal chemistry, computational chemistry and systems pharmacology. Nucleic Acids Research, 44(D1), D1045–D1053.
58. Mauri, A., Consonni, V., Pavan, M., Todeschini, R., et al. (2006). Dragon software: An easy approach to molecular descriptor calculations. Match, 56(2), 237–248.
59. Steinbeck, C., Han, Y., Kuhn, S., Horlacher, O., Luttmann, E., & Willighagen, E. (2003). The Chemistry Development Kit (CDK): An open-source Java library for chemo- and bioinfor­matics. Journal of Chemical Information and Computer Sciences, 43(2), 493–500.
60. Tetko, I. V., Gasteiger, J., Todeschini, R., Mauri, A., Livingstone, D., Ertl, P., Palyulin, V. A., Radchenko, E. V., Zerov, N. S., Makarenko, A. S., Tanchuk, V. Y., & Prokopenko, V. V. (2005). Virtual computational chemistry laboratoryDesign and description. Journal of Computer-Aided Molecular Design, 19 (6), 453
61. Hong, H., Xie, Q., Ge, W., Qian, F., Fang, H., Shi, L., Su, Z., Perkins, R., & Tong, W. (2008). Mold2, molecular descriptors from 2D structures for chemoinformatics and toxicoinformatics. Journal of Chemical Information and Modeling, 48(7), 1337–1344.
62. OBoyle, N. M., Morley, C., & Hutchison, G. R. (2008). Pybel: A Ppython wrapper for the OpenBabel cheminformatics toolkit. Chemistry Central Journal, 2(1), 5.
63. Yap, C. W. (2011). PaDEL-descriptor: An open source software to calculate molecular descriptors and ngerprints. Journal of Computational Chemistry, 32(7), 1466–1474.
64. Cao, D.-S., Liang, Y.-Z., Yan, J., Tan, G.-S., Xu, Q.-S., & Liu, S. (2013). PyDPI: Freely available python package for chemoinformatics, bioinformatics, and chemogenomics studies. Journal of Chemical Information and Modeling, 53(11), 3086–3096.
65. Cao, D.-S., Xu, Q.-S., Hu, Q.-N., & Liang, Y.-Z. (2013). ChemoPy: Freely available Python package for computational biology and chemoinformatics. Bioinformatics, 29(8), 1092–1094.
–463.
158 P. O. Fernandes and V. G. Maltarollo
66. Dong, J., Cao, D.-S., Miao, H.-Y., Liu, S., Deng, B.-C., Yun, Y.-H., Wang, N.-N., Lu, A.-P., Zeng, W.-B., & Chen, A. F. (2015). ChemDes: An integrated web-based platform for molecular descriptor and ngerprint computation. Journal of Cheminformatics, 7(1), 60.
67. Cao, D.-S., Xiao, N., Xu, Q.-S., & Chen, A. F. (2015). Rcpi: R/Bioconductor package to generate various descriptors of proteins, compounds and their interactionsc. Bioinformatics, 31(2), 279–281.
68. Dong, J., Yao, Z.-J., Wen, M., Zhu, M.-F., Wang, N.-N., Miao, H.-Y., Lu, A.-P., Zeng, W.-B., & Cao, D.-S. (2016). BioTriangle: A web-accessible platform for generating various molec­ular representations for chemicals, proteins, DNAs/RNAs and their interactions. Journal of Cheminformatics, 8(1), 34.
69. Dong, J., Yao, Z.-J., Zhu, M.-F., Wang, N.-N., Lu, B., Chen, A. F., Lu, A.-P., Miao, H., Zeng, W.-B., & Cao, D.-S. (2017). ChemSAR: An online pipelining platform for molecular SAR modeling. Journal of Cheminformatics, 9(1), 27.
70. Moriwaki, H., Tian, Y.-S., Kawashita, N., & Takagi, T. (2018). Mordred: A molecular descriptor calculator. Journal of Cheminformatics, 10(1), 4.
71. Dong, J., Yao, Z.-J., Zhang, L., Luo, F., Lin, Q., Lu, A.-P., Chen, A. F., & Cao, D.-S. (2018). PyBioMed: A python library for various molecular representations of chemicals, proteins and DNAs and their interactions. Journal of Cheminformatics, 10(1), 16.
72. Mauri, A. (2020). alvaDesc: A tool to calculate and analyze molecular descriptors and ngerprints. In K. Roy (Ed.), Ecotoxicological QSARs (pp. 801820). Springer US.
73. Dong, J., Zhu, M.-F., Yun, Y.-H., Lu, A.-P., Hou, T.-J., & Cao, D.-S. (2021). BioMedR: An R/CRAN package for integrated data analysis pipeline in biomedical study. Briengs in Bioinformatics, 22(1), 474–484.
74. Coley, C. W., Barzilay, R., Green, W. H., Jaakkola, T. S., & Jensen, K. F. (2017). Convolutional embedding of attributed molecular graphs for physical property prediction. Journal of Chemical Information and Modeling, 57(8), 1757–1772.
75. Wang, Y., Wu, S., Duan, Y., & Huang, Y. (2022). A point cloud-based deep learning strategy for protein–ligand binding afnity prediction. Briengs in Bioinformatics, 23(1), bbab474.
76. David, L., Thakkar, A., Mercado, R., & Engkvist, O. (2020). Molecular representations in AI-driven drug discovery: A review and practical guide. Journal of Cheminformatics, 12(1),
56.
77. Heid, E., Greenman, K. P., Chung, Y., Li, S.-C., Graff, D. E., Vermeire, F. H., Wu, H., Green, W. H., & McGill, C. J. (2024). Chemprop: A machine learning package for chemical property prediction. Journal of Chemical Information and Modeling, 64(1), 9–17.
78. Stokes, J. M., Yang, K., Swanson, K., Jin, W., Cubillos-Ruiz, A., Donghia, N. M., MacNair, C. R., French, S., Carfrae, L. A., Bloom-Ackermann, Z., et al. (2020). A deep learning approach to antibiotic discovery. Cell, 180(4), 688–702.
79. Jin, W., Stokes, J. M., Eastman, R. T., Itkin, Z., Zakharov, A. V., Collins, J. J., Jaakkola, T. S., & Barzilay, R. (2021). Deep learning identies synergistic drug combinations for treating COVID-19. Proceedings of the National Academy of Sciences, 118(39), e2105070118.
80. Lim, M. A., Yang, S., Mai, H., & Cheng, A. C. (2022). Exploring deep learning of quantum chemical properties for absorption, distribution, metabolism, and excretion predictions. Jour- nal of Chemical Information and Modeling, 62(24), 6336–6341.
81. Lenselink, E. B., & Stouten, P. F. W. (2021). Multitask machine learning models for predicting Lipophilicity (logP) in the SAMPL7 challenge. Journal of Computer-Aided Molecular Design, 35(8), 901–
82. McGill, C., Forsuelo, M., Guan, Y., & Green, W. H. (2021). Predicting infrared spectra with message passing neural networks. Journal of Chemical Information and Modeling, 61(6), 2594–2609.
83. Masand, V. H., Mahajan, D. T., Nazeruddin, G. M., Hadda, T. B., Rastija, V., & Alfeefy, A. M. (2015). Effect of information leakage and method of splitting (rational and random) on external predictive ability and behavior of different statistical parameters of QSAR model. Medicinal Chemistry Research, 24(3), 1241–1264.
909.
6 QSAR and Machine Learning Predictors 159
84. Martin, T. M., Harten, P., Young, D. M., Muratov, E. N., Golbraikh, A., Zhu, H., & Tropsha, A. (2012). Does rational selection of training and test sets improve the outcome of QSAR modeling? Journal of Chemical Information and Modeling, 52(10), 2570–2578.
85. Puzyn, T., Mostrag-Szlichtyng, A., Gajewicz, A., Skrzyński, M., & Worth, A. P. (2011). Investigating the inuence of data splitting on the predictive ability of QSAR/QSPR models. Structural Chemistry, 22(4), 795–804.
86. Esbensen, K. H., & Geladi, P. (2010). Principles of proper validation: Use and abuse of re-sampling for validation. Journal of Chemometrics, 24(3–4), 168–187.
87. Hawkins, D. M., Basak, S. C., & Mills, D. (2003). Assessing model t by cross-validation. Journal of Chemical Information and Computer Sciences, 43(2), 579–586.
88. Andrada, M. F., Vega-Hissi, E. G., Estrada, M. R., & Garro Martinez, J. C. (2017). Impact assessment of the rational selection of training and test sets on the predictive ability of QSAR models. SAR and QSAR in Environmental Research, 28(12), 1011–1023.
89. Wu, W., Walczak, B., Massart, D. L., Heuerding, S., Erni, F., Last, I. R., & Prebble, K. A. (1996). Articial neural networks in classication of NIR spectral data: Design of the training set. Chemometrics and Intelligent Laboratory Systems, 33(1), 35–46.
90. Kronenberger, T., Windshügel, B., Wrenger, C., Honorio, K. M., & Maltarollo, V. G. (2018). On the relationship of Anthranilic derivatives Structure and the FXR (Farnesoid X receptor) agonist Activity. Journal of Biomolecular Structure and Dynamics, 36(16), 4378–4391.
91. Gomes, R. A., Genesi, G. L., Maltarollo, V. G., & Trossini, G. H. G. (2017). Quantitative structure–activity relationships (HQSAR, CoMFA, and CoMSIA) studies for COX-2 selective inhibitors. Journal of Biomolecular Structure and Dynamics, 35(7), 1436–1445.
92. Veríssimo, G. C., Pantaleão, S. Q., de Fernandes, P. O., Gertrudes, J. C., Kronenberger, T., Honorio, K. M., & Maltarollo, V. G. (2023). MASSA algorithm: An automated rational sampling of training and test subsets for QSAR modeling. Journal of Computer-Aided Molecular Design, 37(12), 735–754.
93. RDKit. https://www.rdkit.org/. Accessed 2022-03-08.
94. Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., & Duchesnay, É. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12(85), 2825–2830.
95. Virtanen, P., Gommers, R., Oliphant, T. E., Haberland, M., Reddy, T., Cournapeau, D., Burovski, E., Peterson, P., Weckesser, W., Bright, J., van der Walt, S. J., Brett, M., Wilson, J., Millman, K. J., Mayorov, N., Nelson, A. R. J., Jones, E., Kern, R., Larson, E., Carey, C. J., Polat, İ., Feng, Y., Moore, E. W., VanderPlas, J., Laxalde, D., Perktold, J., Cimrman, R., Henriksen, I., Quintero, E. A., Harris, C. R., Archibald, A. M., Ribeiro, A. H., Pedregosa, F., & van Mulbregt, P. (2020). SciPy 1.0: Fundamental algorithms for scientic computing in python. Nature Methods, 17(3), 261–272.
96. Sydow, D., Morger, A., Driller, M., & Volkamer, A. (2019). TeachOpenCADD: A teaching platform for computer-aided drug design using open source packages and data. Journal of Cheminformatics, 11(1), 29.
97. Berthold, M. R., Cebron, N., Dill, F., Gabriel, T. R., Kötter, T., Meinl, T., Ohl, P., Sieb, C., Thiel, K., & Wiswedel, B. (2008). KNIME: The Konstanz information miner. In C. Preisach, H. Burkhardt, L. Schmidt-Thieme, & R. Decker (Eds.), Data analysis, machine learning and applications (Studies in classication, data analysis, and knowledge organization) (pp. 319–326). Springer Berlin Heidelberg.
98. Demšar, J., Curk, T., Erjavec, A., Gorup, Č., Hočevar, T., Milutinovič, M., Možina, M., Polajnar, M., Toplak, M., Starič, A., Štajdohar, M., Umek, L., Žagar, L., Žbontar, J., Žitnik, M., & Zupan, B. (2013). Orange: Data mining toolbox in Python. Journal of Machine Learning Research, 14, 2349–2353.
99. Frank, E., Hall, M. A., & Witten, I. H. (2016). The WEKA workbench. Morgan Kaufmann.
160 P. O. Fernandes and V. G. Maltarollo
100. Joshi, R., Zheng, Z., Agarwal, P., Hatmal, M. M., Chang, X., Seidler, P., & Haworth, I. S. (2024). KNIME workows for applications in medicinal and computational chemistry. Arti- cial Intelligence Chemistry, 2(1), 100063.
101. Nantasenamat, C., Worachartcheewan, A., Jamsak, S., Preeyanon, L., Shoombuatong, W., Simeon, S., Mandi, P., Isarankura-Na-Ayudhya, C., & Prachayasittikul, V. (2015). AutoWeka: Toward an automated data mining software for QSAR and QSPR studies. In H. Cartwright (Ed.), Articial neural networks (pp. 119–147). Springer.
102. Ragno, R. (2019). www.3d-Qsar.Com: A web portal that brings 3-D QSAR to all electronic devicesThe Py-CoMFA web application as tool to build models from pre-aligned datasets. Journal of Computer-Aided Molecular Design, 33(9), 855–864.
103. de Silverio, P. S. S. N., de Viana, J. O., & Barbosa, E. G. (2023). 3D-QSARpy: Combining variable selection strategies and machine learning techniques to build QSAR models. Brazil- ian Journal of Pharmaceutical Sciences, 59, e22373.
104. Gramatica, P., Chirico, N., Papa, E., Cassani, S., & Kovarich, S. (2013). QSARINS: A new software for the development, analysis, and validation of QSAR MLR models. Journal of Computational Chemistry, 34(24), 2121–2132.
105. Martins, J. P. A., Barbosa, E. G., Pasqualoto, K. F. M., & Ferreira, M. M. C. (2009). LQTA­QSAR: A new 4D-QSAR methodology. Journal of Chemical Information and Modeling, 49(6), 1428–1436.
106. Teólo, R. F., Martins, J. P. A., & Ferreira, M. M. C. (2009). Sorting variables by using informative vectors as a strategy for feature selection in multivariate regression. Journal of Chemometrics, 23(1), 32–48.
107. Martins, J. P. A., & Ferreira, M. M. C. (2013). QSAR modeling: um novo pacote computacional open source para gerar e validar modelos QSAR. Química Nova, 36, 554–560.
108. Toropova, A. P., & Toropov, A. A. (2014). CORAL software: Prediction of carcinogenicity of drugs by means of the Monte Carlo method. European Journal of Pharmaceutical Sciences, 52,21–25.
109. de Oliveira, D. B., & Gaudio, A. C. (2000). BuildQSAR: A new computer program for QSAR analysis. Quantitative Structure-Activity Relationships, 19(6), 599–601.
110. Veerasamy, R., Rajak, H., Jain, A., Sivadasan, S., Varghese, C. P., & Agrawal, R. K. (2011). Validation of QSAR models-strategies and importance. International Journal of Drug Design & Discovery, 3, 511–519.
111. Pratim Roy, P., Paul, S., Mitra, I., & Roy, K. (2009). On two novel parameters for validation of predictive QSAR models. Molecules, 14(5), 1660–1701.
112. Consonni, V., Ballabio, D., & Todeschini, R. (2009). Comments on the denition of the Q2 parameter for QSAR Validation. Journal of Chemical Information and Modeling, 49(7), 1669–1678.
113. Gramatica, P., & Sangion, A. (2016). A historical excursus on the statistical validation parameters for QSAR Models: A clarication concerning metrics and terminology. Journal of Chemical Information and Modeling, 56(6), 1127
114. Venkatraman, V., Chakravarthy, P. R., & Kihara, D. (2009). Application of 3D Zernike descriptors to shape-based ligand similarity searching. Journal of Cheminformatics, 1(1), 19.
115. Avram, S. I., Crisan, L., Bora, A., Pacureanu, L. M., Avram, S., & Kurunczi, L. (2013). Retrospective group fusion similarity search based on eROCE evaluation metric. Bioorganic & Medicinal Chemistry, 21(5), 1268–1278.
116. Castillo-González, D., Mergny, J.-L., De Rache, A., Pérez-Machado, G., Cabrera-Pérez, M. A., Nicolotti, O., Introcaso, A., Mangiatordi, G. F., Guédin, A., Bourdoncle, A., Garrigues, T., Pallardó, F., Cordeiro, M. N. D. S., Paz-y-Miño, C., Tejera, E., Borges, F., & Cruz­Monteagudo, M. (2015). Harmonization of QSAR best practices and molecular docking provides an efcient virtual screening tool for discovering new G-Quadruplex ligands. Journal
of Chemical Information and Modeling, 55(10), 2094–2110.
1131.
6 QSAR and Machine Learning Predictors 161
117. Hanser, T., Barber, C., Marchaland, J. F., & Werner, S. (2016). Applicability domain: Towards a more formal de nition$. SAR and QSAR in Environmental Research, 27(11), 865–881.
118. Seram, M. S. M., Pantaleão, S. Q., da Silva, E. B., McKerrow, J. H., ODonoghue, A. J., Mota, B. E. F., Honorio, K. M., & Maltarollo, V. G. (2023). The importance of good practices and false hits for QSAR-driven virtual screening real application: A SARS-CoV-2 Main protease (Mpro) case study. Frontiers in Drug Discovery, 3.
119. Rudin, C. (2019). Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence, 1(5), 206–215.
120. Breiman, L. (2001). Random forests. Machine Learning, 45,5–32.
121. Ribeiro, M. T., Singh, S., & Guestrin, C. (2016). Why should I trust you?: Explaining the predictions of any classi er. In Proceedings of the 22nd ACM SIGKDD International Con- ference on Knowledge Discovery and Data Mining (pp. 1135–1144). ACM.
122. Lundberg, S. M., & Lee, S.-I. (2017). A unied approach to interpreting model predictions. In Advances in neural information processing systems (Vol. 30). Curran Associates, Inc.
123. Moreira-Filho, J. T., Neves, B. J., Cajas, R. A., de Moraes, J., & Andrade, C. H. (2023). Articial intelligence-guided approach for efcient virtual screening of hits against Schistosoma Mansoni. Future Medicinal Chemistry, 15(22), 2033–2050.
124. Mushtaq, M., Usmani, S., Jabeen, A., Nur-e-Alam, M., Ahmed, S., Ahmad, A., & Ul-Haq, Z. (2023). Identication of potent anti-immunogenic agents through virtual screening, 3D-QSAR studies, and in vitro experiments. Molecular Diversity.
125. Barbosa, H., Espinoza, G. Z., Amaral, M., de Castro Levatti, E. V., Abiuzi, M. B., Veríssimo, G. C., de Fernandes, P. O., Maltarollo, V. G., Tempone, A. G., Honorio, K. M., & Lago, J. H. G. (2024). Andrographolide: A diterpenoid from Cymbopogon schoenanthus identied as a new hit compound against Trypanosoma cruzi using machine learning and experimental approaches. Journal of Chemical Information and Modeling, 64(7), 2565–2576.
126. Fernandes, P. O., Dias, A. L. T., dos Santos Júnior, V. S., Sá Magalhães Seram, M., Sousa, Y. V., Monteiro, G. C., Coutinho, I. D., Valli, M., Verzola, M. M. S. A., Ottoni, F. M., de Pádua, R. M., Oda, F. B., dos Santos, A. G., Andricopulo, A. D., da Silva Bolzani, V., Mota, B. E. F., Alves, R. J., de Oliveira, R. B., Kronenberger, T., & Maltarollo, V. G. (2024). Machine learning-based virtual screening of antibacterial agents against methicillin­susceptible and resistant staphylococcus aureus. Journal of Chemical Information and Model- ing, 64(6), 1932–1944.
127. Wong, F., Zheng, E. J., Valeri, J. A., Donghia, N. M., Anahtar, M. N., Omori, S., Li, A., Cubillos-Ruiz, A., Krishnan, A., Jin, W., Manson, A. L., Friedrichs, J., Helbig, R., Hajian, B., Fiejtek, D. K., Wagner, F. F., Soutter, H. H., Earl, A. M., Stokes, J. M., Renner, L. D., & Collins, J. J. (2023). Discovery of a structural class of antibiotics with explainable deep learning. Nature, 626(7997), 177–185.
128. Rath, M., Wellnitz, J., Martin, H.-J., Melo-Filho, C., Hochuli, J. E., Silva, G. M., Beasley, J.-M., Travis, M., Sessions, Z. L., Popov, K. I., Zakharov, A. V., Cherkasov, A., Alves, V., Muratov, E. N., & Tropsha, A. (2024). Pharmacokinetics Pro opment, Validation, and implementation as a web tool for triaging compounds with undesired pharmacokinetics proles. Journal of Medicinal Chemistry, 67(8), 6508–6518.
129. White, A. D. (2023). The future of chemistry is language. Nature Reviews Chemistry, 7(7), 457–458.
130. Flam-Shepherd, D., Zhu, K., & Aspuru-Guzik, A. (2022). Language models can learn complex molecular distributions. Nature Communications, 13(1), 3293.
ler (PhaKinPro): Model devel-