Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5364_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
02.09.2026
Размер:
21 Мб
Скачать
5 Clustering of Small Molecules 121
Table 5.1 Resources to perform small molecules
Name Clustering method Pros and cons Reference ChemBioSe-
rver 2.0 (available as web app)
ChemMine tools (avail­able as web app)
Deep-Clus­tering (avail­able as Python code)
Hierarchical and afnity propa­gation clustering
Hierarchical clustering, multidimensional scaling (MDS), and binning clustering
VAE (Variational Auto Encoder) & K-means AE (Auto encoder) & K-means
Pros: Compound ngerprints can be provided by the user or generated on-site using 166-bit MACCS Open Babel ngerprint from .sdf or .mol les. In the case of hierarchical clustering, the user can select among dif­ferent distances, linkage approaches, and thresholds. A tutorial is available on the site. Results are stored for a week Cons: Despite the tutorial and the availability of a supporting paper, the information on the fundamentals of the afnity propagation clustering may be a bit scarce
Pros: Compounds can be input­ted using SMILES notation, as . sdf or using PubChem CIDs. They can also be drawn online. The user can tune different parameters. A tutorial is avail­able online. Depending on the clustering methods, the output can be exported graphically or as tables/.csv le. Past jobs are accessible. Free support avail­able Cons: Even for small sized data sets, jobs run rather slowly. Despite the tutorial and the availability of a supporting paper, the details of the proce­dures are a bit scarce, though most of them use very well­documented functions implemented in R
Pros: The clustering procedure exploits both local and global molecular features. It can be applied to large-scale chemical libraries. The supporting paper describes the clustering proce­dure in detail Cons: Unavailability as web application
[25]
[4]
[21]
(continued)
122 A. Talevi et al.
Table 5.1 (continued)
Name Clustering method Pros and cons Reference ChemmineR
(R package)
hclust (R function)
RDKIT (col­lection of cheminform­atics and machine­learning soft­ware written in C++ and Python)
iRaPCA (available as web app and Python code)
SOMoC (available as a web app and Python code)
Hierarchical clustering:
Maximum Common Sub­structure (MCS) Search. Nonhierarchical clustering:
K-means clustering
Different hierarchical clustering methods (Ward, single, com­plete, and average linkage, among others)
Sphere exclusion and fuzzy clustering
Hybrid clustering method using molecular descriptors and K-means
Non-hierarchical clustering using EState1 ngerprints and UMAP
Pros: ChemmineR is a popular cheminformatics package, with plenty of documentation and an active community forum and blog available through their developers. The developers wel­come community-contributed resources
Pros: This is a well-documented R function to perform hierarchi­cal clustering based on different combinations of distances and linkage methods. In fact, it has been used to build some other clustering resources listed here
Pros: Large community of users. Community forum is available. Well documented Cons: Time-consuming, these methods can be slow to run and are best used on small sets (no more than a few hundred molecules) of small molecules
Pros: The developers have vali­dated their approach across 29 data sets of different sizes, obtaining better metrics than several classic approximations. The user can tune different parameters and input their own molecular descriptor sets to per­form the clustering. Allows iter­ative sub-clustering. Free support is available. The supporting paper describes the clustering procedure in detail Cons: Long computer times for large-scale libraries. Low interpretability
Pros: The developers have vali­dated their approach across 29 data sets of different sizes and obtained better metrics than several classic approximations. The user can tune different parameters. Free support is available. The supporting paper describes the clustering
[12]
[53]
[24]
[38]
[38]
(continued)
5 Clustering of Small Molecules 123
Table 5.1 (continued)
Name Clustering method Pros and cons Reference
procedure in detail. Compatible with large-scale libraries. Cons: Relatively low interpretability
clValid (R package)
NbClust (R package)
cclust (R package)
MASSA (Python code)
Variety of methods for validat­ing the results from a cluster analysis.
Clustering validity indexes can be applied to outputs of two clustering algorithms: K-means and hierarchical agglomerative clustering (HAC), by varying all combinations of a number of clusters, distance measures, and clustering methods
Convex Clustering methods include the K-means algorithm, the On-line Update algorithm (Hard Competitive Learning), and the Neural Gas algorithm (Soft Competitive Learning)
Explores the biological, physi­cochemical, and structural spaces of molecules using PCA, hierarchical clustering analysis and K-modes
[8]
Pros: Provides 30 indexes that determine the number of groups in a data set. The user can simultaneously evaluate several clustering schemes while vary­ing the number of clusters. Nbclust offers functionalities to visualize the results of various indices
Pros: cclust offers a variety of convex clustering methods and provides several clustering indexes. It can handle complex data sets with large numbers of data points and dimensions. Cons: The interpretation of clustering results can be dif­cult, especially for high­dimensional data sets or those with complex structures
Initial evaluation of the algo­rithm in the context of QSAR/ QSPR modeling to split data sets into training and test sets pro­duced models with low variabil­ity and better values for validation metrics, even when the descriptors used in the QSAR/QSPR were different from those used in the separation of training and test sets, Gener­ates graphical representations that can provide further insights into the data
[14]
[63]
[51]
124 A. Talevi et al.
Table 5.2 Questions to answer when planning a clustering study
Have you checked relevant literature on the subject? Is there some prior clue about the structure of the data that can serve as an external validation
criterion? If so, do you trust the data? Can you do anything to minimize data noise and data mislabeling? Would it be convenient for you to include a feature-selection step? What internal cluster validation index(es) will you include? Would you address cluster replicability? Has any bias been reported for the clustering method(s) of choice?
of the authors is that in the real world a ground truth, a ground cluster structure of real data, does not exist (especially if we use clustering as a truly unsupervised approximation, with no preconception or prior labeling of the data points). This view seems in line with the notion of clusterability. As pointed out by Ackerman and Ben-David [2], if a data set is hard to cluster then it does not have a meaningful clustering structure.However, we accept that some representations of the data are more useful for speci c purposes, and their usefulness depends on the value criteria of the user.
This leads us to the problem of variable selec tion, an old debate that has not yet been fully settled. Ketchen et al. [27] describe three fundamental approaches to the selection of relevant features for a clustering experiment: inductive, deductive, and cognitive.
The inductive approach considers as many variables as possible because it is not known a priori which features provide the best differentiation among observations; thus, using as many clustering variables is likely to maximize the chances of discovering signicant differences [28, 34]. In contrast, the deductive and cognitive approaches (pre)select the feature space based on either theoretic knowledge or expert judgment. The (often vast) universe of possible features is thus pruned to a more limited and, in principle, meaningful number based on previous knowledge .
The preceding concepts are captured, in one way or another, by clustering approaches that incorporate dimensionality-reduction and/or feature-selection steps, although the latter have rarely been applied in the eld of small molecule clustering [42]. Small molecules are often represented by high-dimensional vectors (i.e., a large pool of global and/or local molecular features), which, for different reasons (to facilitate visual representation, reduce computational cost, and/or capture orthogonal and signicant feature projections), are often pre-processed using dimen­sionality reduction techniques, such as Principal Component Analysis (PCA) [38]or the Uniform Manifold Approximation and Projection (UMAP) approaches [24, 38]. Hadipour et al. [21] recently reported an open-source deep clustering approach, in which dimensionality-reduction techniques wer e implemented serially. First, the authors resorted to PCA on sets of global and local molecular features, leading to a representation comprising 243 features, which was then subjected to further dimensionality reduction using autoencoders.
5 Clustering of Small Molecules 125
In particular, McKelveys view is well embodied by subspace clustering approx­imations, which attempt to nd relevant cluster structures in different subspaces of the feature space. In this way, large pools of variables could be exploited through smaller subsets, either through systematic exploration (computationally very demanding for high numbers of features) or stoch astic, until combinations of vari­ables useful for the problem in question are found, without a priori selection of variables biased by the researchers prior knowledge (and lack of knowledge!). For example, starting from descriptor sets provided by the user, the iRaPCA approach [38] explores stochastic subsets of molecular descriptors where cohesive and well­separated clusters can be found, as judged by the silhouette coefcient or other validity measures. Each of the so-obtained clusters can be further divided, itera­tively, into sub-clusters. iRaPCA has shown, without iterations, consistent and almost optimal behavior in benchmarking experiments across 29 data sets of variable sizes [38, 39].
On the other hand, multi-view clustering has also attracted much attention lately [13, 19, 57] and is better aligned with the idea of an underlying ground-truth structure of the data (in which case any view of the data would origi nate from an underlying latent space) . However, it has largely been overlooked in the eld of cheminformatics. Different views of the data may be integrated or combined because of their complementarity (this can be of particular interest when the clustering procedure is supervised via external labeling of the data) or based on their consensus (returning to the idea that high-level stability should be veried across different algorithms, features, and entities).

4 Conclusions

The principle of similarity (guilty by association) has been and is of great importance in the eld of cheminformatics and drug discovery, in which molecular similarity (apparent or not) is used to identify new bioactive scaffolds, develop a series of analogs around such scaffolds, and make predi ctive interpolations related to physicochemical and biological properties of interest. However, in several ways, the eld of supervised machine learning has matured more rapidly than its unsupervised counterpart. Not only is supervised learning much more widely used, but rigorous internal and external validation of supervised models is common practice, while validation of small molecule clustering experiments is much more sporadic. Fur­thermore, while supervised machine-learning procedures in the cheminformatics eld have rapidly assimilated the newest algorithms, small molecule clustering still does not fully exploit recent progress such as subspace clustering and multi­view clustering (with exceptions). A good part of the bibliography used in this chapter comes from disciplines as distant and dissimilar as psychology, organiza­tional management or marketing, image recognition, bioinformatics, etc. In addition, of course, to pure mathematics. The progress that clustering theory has experienced in other disciplines has spilled over, rather very slowly, to cheminformatics.
126 A. Talevi et al.
In a chemical universe that is expanding exponentially, as demonstrated by the currently available ultra-large chemical libraries, the development and application of specic clustering methods will allow the detection of unforeseen patterns in the data and will accelerate the exploitation of this virtually innite chemodiversity. It is no longer possible to trust the human ability to detect molecular patterns, not only because the vast number of accessible molecules cannot be encompassed humanly but also because in such richness and chemical diversity, it is likely that relationships that are not apparent and unnoticed to the naked eye will appear, the detection of which requires higher levels of abstraction and the integration of chemical knowl­edge and advanced mathematics. Subspace clustering and multi-view clustering, still largely unexplored in the eld of chemistry, will provide solutions to the problem of automatic or semi-automatic analysis of ultra-large chemical libraries.
Acknowledgments The three authors are members of CONICET and UNLP. They thank these institutions and FONCYT (PICTs 2019-0984, 2019-1075, 2021-0720) for their nancial support.

References

1. Adnan, M., Slavic, G., Martin Gomez, D., Marcenaro, L., & Regazzoni, C. (2023). Systematic and comprehensive review of clustering and multi-target tracking techniques for LiDAR point clouds in autonomous driving applications. Sensors (Basel), 23, 6119.
2. Ackerman, M., & Ben-David, S. (2009). Clusterability: A theoretical study. Proceedings of the Twelfth International Conference on Articial Intelligence and Statistics, PMLR, 5,1–8.
3. Arbelaitz, I., Gurrutxaga, I., Muguerza, J., Pérez, J. M., & Perona, I. (2013). An extensive comparative study of cluster validity indices. Pattern Recognition, 46, 243–256.
4. Backman, T. W., Cao, Y., & Girke, T. (2011). ChemMine tools: An online service for analyzing and clustering small molecules. Nucleic Acids Research, 39, W486–W491.
5. Baker, F. B., & Hubert, L. J. (1975). Measuring the power of hierarchical cluster analysis. Journal of the American Statistical Association, 70,31–38.
6. Böcker, A., Derksen, S., Schmidt, E., Teckentrup, A., & Schneider, G. (2005). A hierarchical clustering approach for large compound libraries. Journal of Chemical Information and Model- ing, 45, 807–815.
7. Breckenridge, J. N. (2000). Validating cluster analysis: Consistent replication and symmetry. Multivariate Behavioral Research, 35, 261–285.
8. Brock, G., Pihur, V., Datta, S., & Datta, S. (2008). clValid: An R package for cluster validation. Journal of Statistical Software, 25,1–22.
9. Brooks, J. L. (2014). Traditional and new principles of perceptual grouping. In J. Wagemans (Ed.), The Oxford handbook of perceptual organization (pp. 57–87). Oxford University Press.
10. Butina, D. (1999). Unsupervised data base clustering based on daylights ngerprint and tanimoto similarity: A fast and automated way to cluster small and large data sets. Journal of Chemical Information and Computer Sciences, 39, 747–750.
11. Calinski, R. B., & Harabasz, J. (1974). A dendrite method for cluster analysis. Communications in Statistics, 3,1–27.
12. Cao, Y., Charisi, A., Cheng, L. C., Jiang, T., & Girke, T. (2008). ChemmineR: A compound mining framework for R. Bioinformatics, 24(15), 1733–1734.
13. Cao, Z., & Xie, X. (2024). Structure learning with consensus label information for multi-view unsupervised feature selection. Expert Systems with Applications, 238, 121893.
5 Clustering of Small Molecules 127
14. Charrad, M., Ghazzali, N., Boiteau, V., & Niknafs, A. (2014). NbClust: An R package for determining the relevant number of clusters in a data set. Journal of Statistical Software, 61, 1–36.
15. Davies, D., & Bouldin, D. W. (1979). A cluster separation measure. IEEE Transactions on Pattern Analysis and Machine Intelligence, 1, 224–227.
16. Domingo-Fernández, D., Gadiya, Y., Mubeen, S., Healey, D., Norman, B. H., & Colluru, V. (2023). Exploring the known chemical space of the plant kingdom: Insights into taxonomic patterns, knowledge gaps, and bioactive regions. Journal of Cheminformatics, 15, 107.
17. Dunn, J. C. (1974). Well-separated clusters and optimal fuzzy partitions. Journal of Cybernet- ics, 4,95–104.
18. Everitt, B. S., Landau, S., Leese, M., & Stahl, D. (2011). Cluster analysis (5th ed., p. 71). Wiley.
19. Guo, J., Sun, Y., Gao, J., Hu, Y., & Yin, N. (2022). Rank consistency induced multiview subspace clustering via low-rank matrix factorization. IEEE Transactions on Neural Networks and Learning Systems, 33, 3157–3170.
20. Gramatica, P. (2013). On the development and validation of QSAR models. Methods in Molecular Biology, 930, 499–526.
21. Hadipour, H., Liu, C., Davis, R., Cardona, S. T., & Hu, P. (2002). Deep clustering of small molecules at large-scale via variational autoencoder embedding and K-means. BMC Bioinfor- matics, 23(Suppl. 4), 132.
22. Handl, J., Knowles, J., & Kell, D. B. (2005). Computational cluster validation in post-genomic data analysis. Bioinformatics, 21, 3201–3212.
23. Hawkins, D. M., Basak, S. C., & Mills, D. (2003). Assessing model t by cross-validation. Journal of Chemical Information and Computer Sciences, 43, 579–586.
24. Hernández-Hernández, S., & Ballester, P. J. (2023). On the best way to cluster NCI-60 molecules. Biomolecules, 13, 498.
25. Karatzas, E., Zamora, J. E., Athanasiadis, E., Dellis, D., Cournia, Z., Spyrou, G. M., Thomas, J. B., & Snow, C. C. (2020). ChemBioServer 2.0: An advanced web server for ltering, clustering and networking of chemical compounds facilitating both drug discovery and repurposing. Bioinformatics, 36(8), 2602–2604.
26. Kaufman, L., & Rousseeuw, P. J. (1990). Partitioning around medoids (program PAM). In Finding groups in data: An introduction to cluster analysis. Wiley.
27. Ketchen, D. J., Thomas, J. B., & Snow, C. C. (1993). Organizational congurations and performance: A comparison of theoretical approaches. Academy of Management Journal, 36, 1278–1313.
28. Ketchen, D. J., & Shook, C. L. (1996). The application of cluster analysis in strategic management research: An analysis and critique. Strategic Management Journal, 17, 441–458.
https://doi.org/10.1002/(SICI)1097-0266(199606)17:6<441::AID-SMJ819>3.0.CO;2-G
29. Krieger, A. M., & Green, P. E. (1999). A cautionary note on using internal cross validation to select the number of clusters. Psychometrika, 64, 341–353.
30. Leonard, J. T., & Roy, K. (2006). On selection of training and test sets for the development of predictive QSAR models. QSAR and Combinatorial Science, 25, 235–251.
31. Lupyan, G. (2008). The conceptual grouping effect: Categories matter (and named categories matter more). Cognition, 108 , 566–577.
32. Mayr, A., Klambauer, G., Unterthiner, T., Steijaert, M., Wegner, J. K., Ceulemans, H., Clevert, D. A., & Hochreiter, S. (2018). Large-scale comparison of machine learning methods for drug target prediction on ChEMBL. Chemical Science, 9, 5441–5451.
33. McClain, J. O., & Rao, V. R. (1975). CLUSTISZ: A program to test for the quality of clustering of a set of objects. Journal of Marketing Research, 12, 456–460.
34. McKelvey, B. (1975). Guidelines for empirical classication of organizations. Administrative Science Quarterly, 20, 509525.
35. Milligan, G. W. (1980). An examination of the effect of six types of error perturbation on fteen clustering algorithms. Psychometrika, 45, 325–342.
36. Milligan, G. W. (1981). A Monte Carlo study of thirty internal criterion measures for cluster analysis. Psychometrika, 46, 187–199.
128 A. Talevi et al.
37. Murtagh, F., & Contreras, P. (2017). Algorithms for hierarchical clustering: An overview. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 2, e1219.
38. Prada Gori, D. N., Llanos, M. A., Bellera, C. L., Talevi, A., & Alberca, L. N. (2022a). iRaPCA and SOMoC: Development and validation of web applications for new approaches for the clustering of small molecules. Journal of Chemical Information and Modeling, 62, 2987–2998.
39. Prada Gori, D. N., Alberca, L. N., Rodriguez, S., Llanos, M. A., Bellera, C. L., & Talev, A. (2022b). LIDeB tools: A Latin American resource of freely available, open-source cheminformatics apps. Articial Intelligence in the Life Sciences, 2, 100049.
40. Punj, G. N., & Stewart, D. W. (1983). Cluster analysis in marketing research: Review and suggestions. Journal of Marketing Research, 20, 134–148.
41. Risser-Maroix, O., Marzouki, A., Djeghim, H., Kurtz, C., & Lomenie, N. (2021, September). Learning an adaptation function to assess image visual similarities. Paper presented at ORASIS 2021, Centre National de la Recherche Scientique. Available from https://hal.science/hal-0333
9731v2/document. Accessed 17 Nov 2023.
42. Rivera-Borroto, O. M., Marrero-Ponce, Y., García-de la Vega, J. M., & Grau-Ábalo, R. C. (2011). Comparison of combinatorial clustering methods on pharmacological data sets represented by machine learning-selected real molecular descriptors. Journal of Chemical Information and Modeling, 51, 3036–3049.
43. Rousseeuw, P. (1987). Silhouettes: A graphical aid to the interpretation and validation of cluster analysis. Journal of Computational and Applied Mathematics, 20,53–65.
44. Schubert, E. (2023). Stop using the elbow criterion for k-means and how to choose the number of clusters instead. ACM SIGKDD Explorations Newsletter, 25,36–42.
45. Seger, C. A., & Miller, E. K. (2010). Category learning in the brain. Annual Review of Neuroscience, 33, 203–219.
46. Sheikholeslami, C., Chatterjee, S., & Zhang, A. (2000). WaveCluster: A multi-resolution clustering approach for very large spatial database. VLDB Journal, 8, 289– 304.
47. Tan, P. N., Steinbach, M., & Kumar, V. (2005). Cluster analysis: Basic concepts and algo­rithms. In Introduction to data mining. Addison-Wesley Longman Publishing.
48. Tichý, M., & Rucki, M. (2009). Validation of QSAR models for legislative purposes. Interdis- ciplinary Toxicology, 2, 184–186.
49. Tropsha, A. (2010). Best practices for QSAR model development, validation, and exploitation. Molecular Informatics, 29, 476–488.
50. Tropsha, A., Gramatica, P., & Gombar, V. K. (2003). The importance of being earnest: Validation is the absolute essential for successful application and interpretation of QSPR models. QSAR and Combinatorial Science, 22,69–77.
51. Veríssimo, G. C., Pantaleão, S. Q., Fernandes, P. O., Gertrudes, J. C., Kronenberger, T., Honorio, K. M., & Maltarollo, V. G. (2023). MASSA Algorithm: An automated rational sampling of training and test subsets for QSAR modeling. Journal of Computer-Aided Molec- ular Design, 37, 735–754.
52. Virshup, A. M., Contreras-García, J., Wipf, P., Yang, W., & Beratan, D. N. (2013). Stochastic voyages into uncharted chemical space produce a representative library of all possible drug-like compounds. Journal of the American Chemical Society, 135, 7296–
53. Voicu, A., Duteanu, N., Voicu, M., Vlad, D., & Dumitrascu, V. (2020). The rcdk and cluster R packages applied to drug candidate selection. Journal of Cheminformatics, 12(1), 3.
54. Yang, Y., Yao, K., Repasky, M. P., Leswing, K., Abel, R., Shoichet, B. K., & Jerome, S. V. (2021). Efcient exploration of chemical space with docking and deep learning. Journal of Chemical Theory and Computation, 17, 7106–7119.
55. Yu, L., He, X., Fang, X., Liu, L., & Liu, J. (2023). Deep learning with geometry-enhanced molecular representation for augmentation of large-scale docking-based virtual screening. Journal of Chemical Information and Modeling, 63, 6501–6514.
56. Zhang, C., Huang, W., Niu, T., Liu, Z., Li, G., & Cao, D. (2023). Review of clustering technology and its application in coordinating vehicle subsystems. Automotive Innovation, 6, 89–115.
7303.
5 Clustering of Small Molecules 129
57. Zhang, C., Fu, H., Hu, Q., Cao, X., Xie, Y., Tao, D., & Xu, D. (2020). Generalized latent multi­view subspace clustering. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41, 86–99.
58. Lopez-Del Rio, A., Nonell-Canals, A., Vidal, D., and Perera-Lluna, A. (2019). Evaluation of cross-validation strategies in sequence-based binding prediction using deep learning. J. Chem. Inf. Model 59, 1645–1657.
59. Harris, C. J., Hill, R. D., Sheppard, D. W., Slater, M. J., and Stouten, P. F. (2011). The design and application of target-focused compound libraries. Comb. Chem. High. Throughput Screen 14, 521–531.
60. MacQueen, J. (1967). Some methods for classication and analysis of multivariate observa­tions,in Proceedings of the 5th Berkeley Symposium on Mathematical Statistics and Proba­bility Volume 1. Nerkeley. Editors L. M. Le Cam, and J. Neyman (University of California Press), 281–297.
61. Frey, T., & Van Groenewoud, H. (1972). A cluster analysis of the D-squared matrix of white spruce stands in Saskatchewan based on the maximum-minimum principle.Journal of Ecology, 60, 873–886.
62. Milligan, G.W., Cooper, M.C. An examination of procedures for determining the number of clusters in a data set. Psychometrika 50, 159–179 (1985).
63. Dimitriadou E (2023) cclust: convex clustering methods and clustering indexes http://CRAN.R­project.org/package=cclust. R package version 0.6-26
Chapter 6
QSAR and Machine Learning Predictors
Philipe Oliveira Fernandes and Vinicius Gonçalves Maltarollo
Abstract This chapter de lves into the fundamental principles and applications of
quantitative structure–activity relationship (QSAR) and machine learning (ML)­based predictors in the realm of drug design and chemical biology. QSAR estab­lishes a quantitative relationship between the chemical structure of molecules and their biological activities or physicochemical properties. The evolution of QSAR from its rst reports to its modern applications was covered comprising the theoret­ical foundations, encompassing descriptors, mathematical models (followed by brief examples of ML applied to this eld), and statistical validation techniques employed in QSAR analysis. Interestingly, many recognized and accepted good practices and validation protocols align with OECD guidelines for QSAR applications for regu­latory purposes. In this sense, notably, the QSAR eld becam e important outside of the academic boundaries. Additionally, this chapte r discusses current challenges and emerging trends in QSAR research, including the incorporation of machine learning algorithms and big data analytics for enhanced predictive accuracy and applicability. Overall, this chapter serves as a comprehensive guide for researchers and practi­tioners in understanding and leveraging QSAR as a pivotal tool in rational drug design and chemical biology.
Keywords QSAR · Machine learning · Molecular descriptors · OECD principles

1 Historical Background

The Quantitative Structure–Activity Relationship (QSAR) eld was introduced when researchers tried to correlate the molecular structure of similar compounds to a specic biological activity. This process can be traced back to 1863 when Cros rst reported a correlation between the toxicity of primary aliphatic alcohols and their water solubility [1], a pioneer work in this eld. Another notable work from the
P. O. Fernandes · V. G. Maltarollo () Departamento de Produtos Farmacêuticos, Faculdade de Farmácia, Universidade Federal de Minas Gerais, Belo Horizonte, Minas Gerais, Brazil
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2024 V. G. Maltarollo (ed.), Computer-Aided and Machine Learning-Driven Drug Design, Computer-Aided Drug Discovery and Design 3,
https://doi.org/10.1007/978-3-031-76718-0_6
131