Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5884_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Foreword
- •Acknowledgments
- •Contents
- •1.1 Structure-Based Drug Discovery (SBDD)
- •1.2 Ligand-Based Drug Design (LBDD)
- •1.3 Echoes from the Past, Visions from the Future
- •References
- •1 Introduction
- •2.2 Second Step: Data Curation
- •2.4 Fourth Step: Updating and Maintenance
- •2 Databases and Curation
- •8 Perspectives
- •9 Conclusion
- •References
- •1 Introduction
- •2.1 Making and Matching Protein Models
- •2.2 Simulating Protein Movements
- •2.3 Analyzing Changes in Protein Shape
- •3 Pharmacogenomics in Drug Development
- •4 Case Studies of Genomics-Based Drug Design
- •References
- •1 Historical Background
- •1.1 Timeline
- •2 Methodology Overview
- •2.1 Neural Networks
- •2.1.1 Perceptron
- •2.1.2 Multilayer Neural Networks
- •2.1.3 Types of Neural Networks
- •Feedforward
- •Recurrent Neural Networks
- •LSTM
- •2.2 Deep Learning
- •3 Using Machine Learning
- •3.2 Data Collection
- •3.3 Data Preprocessing
- •3.4 Model Selection
- •3.5 Model Training
- •3.6 Validation
- •3.7 Tuning
- •3.8 Prediction
- •4 Limitations
- •4.1 Bias
- •4.3 Interpretability
- •4.4 Computational Cost
- •4.5 Data Dependency
- •4.6 Robustness
- •5 Applications in Drug Discovery
- •5.2 Lead Discovery
- •5.3 Preclinical and Clinical Development
- •6 Resources and Tools
- •7 Challenges and Perspectives
- •7.1 Future Trends
- •9 Conclusions
- •References
- •1 Historical Background
- •1.1 Applications in Drug Discovery
- •2 Validations and Controls
- •2.1 Internal Validation
- •2.2 External Validation
- •2.3 Relative Cluster Validation
- •3 Challenges and Perspectives
- •4 Conclusions
- •References
- •1 Historical Background
- •2 OECD Principles
- •2.1 A Defined Endpoint
- •2.2 An Unambiguous Algorithm
- •2.5 A Mechanistic Interpretation, if Possible
- •3 Software and Tools
- •4 Validations and Controls
- •4.1 Internal and External Validation
- •4.1.1 Regression Metrics
- •4.2 Applicability Domain
- •4.3 Randomization Tests
- •5 Interpretation
- •6 Practical Advice During QSAR Modeling
- •7 Application
- •8 Challenges and Perspectives
- •References
- •1 Molecular Docking
- •2 Advances in Scoring Functions and Search Algorithms
- •2.2 Critical Characteristics of Search Algorithms
- •2.3 Docking Programs and Scoring Functions
- •3 Calculations Performed During Docking Simulations
- •4 Essential Components for a Good Docking Program
- •5 Limitations of the Docking Technique
- •6 Validation of Docking Results
- •7 Inappropriate Use of Validation Methods in Docking
- •9 Use of Machine Learning in Molecular Docking
- •11 Challenges
- •12 Conclusions
- •References
- •3 System Preparation for MD Simulations
- •3.1 Solvation and Microensemble
- •3.2 Force Fields: General Concept and Relevant Choices
- •3.3 The Concept of Replicas and Timescale
- •4.1.2 Protein Root Mean Square Fluctuation (RMSF)
- •4.1.4 Protein Secondary Structure Analysis
- •4.1.5 Principal component Analysis (PCA)
- •4.1.6 Markov State Modelling
- •4.1.7 Distance Calculations
- •4.1.8 Angle and Plane Calculations
- •4.2.2 Distances and Ligand-Induced Geometry Rearrangements
- •4 Molecular Dynamics Analysis
- •4.1 Protein Perspective
- •4.1.1 Protein Root Mean Square Deviation (RMSD)
- •4.3 Ligand Perspective
- •4.3.1 Ligand Properties
- •4.3.2 Ligand Root Mean Square Deviation
- •4.3.3 Ligand Root Mean Square Fluctuation
- •4.3.4 Angles and Dihedrals
- •5.1 Protein Structure Prediction and Preparation
- •5.2 Molecular Docking
- •6 Concluding Remarks and Outlook
- •Glossary
- •References
- •1 Introduction
- •2.1 MDeNM
- •2.2 Collective Molecular Dynamics (coMD)
- •2.3 ClustENM and ClustENMD
- •3 Ensemble Docking
- •References
- •1 Introduction
- •1.1 Advantages, Disadvantages, Innovations, and Challenges
- •1.2 Recent Advances in Accessible FEP Software Tools
- •1.3 Applications of FEP in Industry and Consortiums
- •2 Expanding the Potential of FEP Calculations
- •2.1 Validating Binding Poses
- •2.2 Dealing with Solvent
- •2.3 FEP and Allostery
- •2.4 FEP and Covalent Ligands
- •2.5 Applications of FEP in Scaffold Hopping
- •2.6 Positional Analogue Scanning
- •2.7 Combinations and Alternative Approaches
- •3 Machine Learning for FEP
- •3.4 Implications for ML in FEP Calculations
- •4 Final Considerations
- •5 First Steps to FEP Simulations
- •References
- •1 Background
- •2 Ultra-Large Screening Libraries and Chemical Spaces
- •3.1 Implications of Dataset Size
- •4 Ligands on the Ultra-Large Scale
- •4.1 Ultra-Large 2D Similarity Searches
- •7 Challenges and Future Perspectives
- •7.1 Hit Triage: An Old Problem on a New Dimension
- •8 Conclusions
- •Appendix
- •References
- •1 Introduction
- •2 Enzymatic Activity Evaluations
- •3 Cytotoxicity Evaluation and Cell Viability
- •4 Antiviral Assays in Experimental Validation
- •6 In Vivo Evaluation of Compounds
- •7 Conclusions
- •References
- •1 Introduction
- •3.1 Data Collection
- •3.2 Data Preprocessing
- •3.4 Model Choice
- •3.5 Model Training
- •3.6 Model Assessment
- •3.7 External Validation
- •3.8 Implementation and Availability
- •3.9 Continuous Update
- •5 Conclusions and Perspectives
- •References
- •1 Experimental Approaches to Obtain Protein Structure
- •1.1 X-Ray Crystallography
- •1.2 Nuclear Magnetic Resonance
- •1.3 Cryo-EM
- •1.4 Hybrid Methods
- •2 Modeling Approaches to Obtain Protein Structure
- •2.1 Homology Modeling
- •2.2 Ab Initio Modeling
- •2.3 New Approaches
- •3 Conformational Diversity of Proteins
- •3.1 Characterization of Protein Conformational States
- •3.2 Experimental Methods to Study Protein Dynamics and Conformations
- •3.4 Molecular Dynamics Simulation
- •3.5 Sampling Strategies
- •4 Remarks and Perspectives
- •References
- •1 Introduction
- •2 Structure-Based Drug Design of HIV Protease Inhibitors
- •2.1 HIV-1 Protease as a Therapeutic Target
- •2.2.1 Saquinavir
- •2.2.2 Indinavir
- •2.3.1 Lopinavir
- •2.3.2 Darunavir
- •6 Conclusions
- •References
- •4 Experimental Methods to Analyze NR Activity
- •4.2 Coregulator-Recruitment
- •5 Concluding Remarks and Outlook
- •References

5 Clustering of Small Molecules 121
Table 5.1 Resources to perform small molecules
Name Clustering method Pros and cons Reference
ChemBioSe-
rver 2.0
(available as
web app)
ChemMine
tools (available as web
app)
Deep-Clustering (available as
Python code)
Hierarchical and affinity propagation clustering
Hierarchical clustering,
multidimensional scaling
(MDS), and binning clustering
VAE (Variational Auto
Encoder) & K-means
AE (Auto encoder) & K-means
Pros: Compound fingerprints
can be provided by the user or
generated on-site using 166-bit
MACCS Open Babel fingerprint
from .sdf or .mol files. In the
case of hierarchical clustering,
the user can select among different distances, linkage
approaches, and thresholds. A
tutorial is available on the site.
Results are stored for a week
Cons: Despite the tutorial and
the availability of a supporting
paper, the information on the
fundamentals of the affinity
propagation clustering may be a
bit scarce
Pros: Compounds can be inputted using SMILES notation, as .
sdf or using PubChem CIDs.
They can also be drawn online.
The user can tune different
parameters. A tutorial is available online. Depending on the
clustering methods, the output
can be exported graphically or as
tables/.csv file. Past jobs are
accessible. Free support available
Cons: Even for small sized data
sets, jobs run rather slowly.
Despite the tutorial and the
availability of a supporting
paper, the details of the procedures are a bit scarce, though
most of them use very welldocumented functions
implemented in R
Pros: The clustering procedure
exploits both local and global
molecular features. It can be
applied to large-scale chemical
libraries. The supporting paper
describes the clustering procedure in detail
Cons: Unavailability as web
application
[25]
[4]
[21]
(continued)

122 A. Talevi et al.
Table 5.1 (continued)
Name Clustering method Pros and cons Reference
ChemmineR
(R package)
hclust
(R function)
RDKIT (collection of
cheminformatics and
machinelearning software written
in C++ and
Python)
iRaPCA
(available as
web app and
Python code)
SOMoC
(available as
a web app
and Python
code)
Hierarchical clustering:
Maximum Common Substructure (MCS) Search.
Nonhierarchical clustering:
K-means clustering
Different hierarchical clustering
methods (Ward, single, complete, and average linkage,
among others)
Sphere exclusion and fuzzy
clustering
Hybrid clustering method using
molecular descriptors and
K-means
Non-hierarchical clustering
using EState1 fingerprints and
UMAP
Pros: ChemmineR is a popular
cheminformatics package, with
plenty of documentation and an
active community forum and
blog available through their
developers. The developers welcome community-contributed
resources
Pros: This is a well-documented
R function to perform hierarchical clustering based on different
combinations of distances and
linkage methods. In fact, it has
been used to build some other
clustering resources listed here
Pros: Large community of users.
Community forum is available.
Well documented
Cons: Time-consuming, these
methods can be slow to run and
are best used on small sets
(no more than a few hundred
molecules) of small molecules
Pros: The developers have validated their approach across
29 data sets of different sizes,
obtaining better metrics than
several classic approximations.
The user can tune different
parameters and input their own
molecular descriptor sets to perform the clustering. Allows iterative sub-clustering. Free
support is available. The
supporting paper describes the
clustering procedure in detail
Cons: Long computer times for
large-scale libraries. Low
interpretability
Pros: The developers have validated their approach across
29 data sets of different sizes and
obtained better metrics than
several classic approximations.
The user can tune different
parameters. Free support is
available. The supporting paper
describes the clustering
[12]
[53]
[24]
[38]
[38]
(continued)

5 Clustering of Small Molecules 123
Table 5.1 (continued)
Name Clustering method Pros and cons Reference
procedure in detail. Compatible
with large-scale libraries.
Cons: Relatively low
interpretability
clValid
(R package)
NbClust
(R package)
cclust
(R package)
MASSA
(Python
code)
Variety of methods for validating the results from a cluster
analysis.
Clustering validity indexes can
be applied to outputs of two
clustering algorithms: K-means
and hierarchical agglomerative
clustering (HAC), by varying all
combinations of a number of
clusters, distance measures, and
clustering methods
Convex Clustering methods
include the K-means algorithm,
the On-line Update algorithm
(Hard Competitive Learning),
and the Neural Gas algorithm
(Soft Competitive Learning)
Explores the biological, physicochemical, and structural
spaces of molecules using PCA,
hierarchical clustering analysis
and K-modes
– [8]
Pros: Provides 30 indexes that
determine the number of groups
in a data set. The user can
simultaneously evaluate several
clustering schemes while varying the number of clusters.
Nbclust offers functionalities to
visualize the results of various
indices
Pros: cclust offers a variety of
convex clustering methods and
provides several clustering
indexes. It can handle complex
data sets with large numbers of
data points and dimensions.
Cons: The interpretation of
clustering results can be difficult, especially for highdimensional data sets or those
with complex structures
Initial evaluation of the algorithm in the context of QSAR/
QSPR modeling to split data sets
into training and test sets produced models with low variability and better values for
validation metrics, even when
the descriptors used in the
QSAR/QSPR were different
from those used in the separation
of training and test sets, Generates graphical representations
that can provide further insights
into the data
[14]
[63]
[51]

124 A. Talevi et al.
Table 5.2 Questions to answer when planning a clustering study
Have you checked relevant literature on the subject?
Is there some prior clue about the structure of the data that can serve as an external validation
criterion?
If so, do you trust the data? Can you do anything to minimize data noise and data mislabeling?
Would it be convenient for you to include a feature-selection step?
What internal cluster validation index(es) will you include?
Would you address cluster replicability?
Has any bias been reported for the clustering method(s) of choice?
of the authors is that in the real world a ground truth, a ground cluster structure of
real data, does not exist (especially if we use clustering as a truly unsupervised
approximation, with no preconception or prior labeling of the data points). This view
seems in line with the notion of clusterability. As pointed out by Ackerman and
Ben-David [2], “if a data set is hard to cluster then it does not have a meaningful
clustering structure.” However, we accept that some representations of the data are
more useful for speci fic purposes, and their usefulness depends on the value criteria
of the user.
This leads us to the problem of variable selec tion, an old debate that has not yet
been fully settled. Ketchen et al. [27] describe three fundamental approaches to the
selection of relevant features for a clustering experiment: inductive, deductive, and
cognitive.
The inductive approach considers as many variables as possible because it is not
known a priori which features provide the best differentiation among observations;
thus, using as many clustering variables is likely to maximize the chances of
discovering significant differences [28, 34]. In contrast, the deductive and cognitive
approaches (pre)select the feature space based on either theoretic knowledge or
expert judgment. The (often vast) universe of possible features is thus pruned to a
more limited and, in principle, meaningful number based on previous knowledge .
The preceding concepts are captured, in one way or another, by clustering
approaches that incorporate dimensionality-reduction and/or feature-selection
steps, although the latter have rarely been applied in the field of small molecule
clustering [42]. Small molecules are often represented by high-dimensional vectors
(i.e., a large pool of global and/or local molecular features), which, for different
reasons (to facilitate visual representation, reduce computational cost, and/or capture
orthogonal and significant feature projections), are often pre-processed using dimensionality reduction techniques, such as Principal Component Analysis (PCA) [38]or
the Uniform Manifold Approximation and Projection (UMAP) approaches
[24, 38]. Hadipour et al. [21] recently reported an open-source deep clustering
approach, in which dimensionality-reduction techniques wer e implemented serially.
First, the authors resorted to PCA on sets of global and local molecular features,
leading to a representation comprising 243 features, which was then subjected to
further dimensionality reduction using autoencoders.

5 Clustering of Small Molecules 125
In particular, McKelvey’s view is well embodied by subspace clustering approximations, which attempt to find relevant cluster structures in different subspaces of
the feature space. In this way, large pools of variables could be exploited through
smaller subsets, either through systematic exploration (computationally very
demanding for high numbers of features) or stoch astic, until combinations of variables useful for the problem in question are found, without a priori selection of
variables biased by the researcher’s prior knowledge (and lack of knowledge!). For
example, starting from descriptor sets provided by the user, the iRaPCA approach
[38] explores stochastic subsets of molecular descriptors where cohesive and wellseparated clusters can be found, as judged by the silhouette coefficient or other
validity measures. Each of the so-obtained clusters can be further divided, iteratively, into sub-clusters. iRaPCA has shown, without iterations, consistent and
almost optimal behavior in benchmarking experiments across 29 data sets of variable
sizes [38, 39].
On the other hand, multi-view clustering has also attracted much attention lately
[13, 19, 57] and is better aligned with the idea of an underlying ground-truth
structure of the data (in which case any view of the data would origi nate from an
underlying latent space) . However, it has largely been overlooked in the field of
cheminformatics. Different views of the data may be integrated or combined because
of their complementarity (this can be of particular interest when the clustering
procedure is supervised via external labeling of the data) or based on their consensus
(returning to the idea that high-level stability should be verified across different
algorithms, features, and entities).
4 Conclusions
The principle of similarity (“guilty by association”) has been and is of great
importance in the field of cheminformatics and drug discovery, in which molecular
similarity (apparent or not) is used to identify new bioactive scaffolds, develop a
series of analogs around such scaffolds, and make predi ctive interpolations related to
physicochemical and biological properties of interest. However, in several ways, the
field of supervised machine learning has matured more rapidly than its unsupervised
counterpart. Not only is supervised learning much more widely used, but rigorous
internal and external validation of supervised models is common practice, while
validation of small molecule clustering experiments is much more sporadic. Furthermore, while supervised machine-learning procedures in the cheminformatics
field have rapidly assimilated the newest algorithms, small molecule clustering
still does not fully exploit recent progress such as subspace clustering and multiview clustering (with exceptions). A good part of the bibliography used in this
chapter comes from disciplines as distant and dissimilar as psychology, organizational management or marketing, image recognition, bioinformatics, etc. In addition,
of course, to pure mathematics. The progress that clustering theory has experienced
in other disciplines has spilled over, rather very slowly, to cheminformatics.

126 A. Talevi et al.
In a chemical universe that is expanding exponentially, as demonstrated by the
currently available ultra-large chemical libraries, the development and application of
specific clustering methods will allow the detection of unforeseen patterns in the data
and will accelerate the exploitation of this virtually infinite chemodiversity. It is no
longer possible to trust the human ability to detect molecular patterns, not only
because the vast number of accessible molecules cannot be encompassed humanly
but also because in such richness and chemical diversity, it is likely that relationships
that are not apparent and unnoticed to the naked eye will appear, the detection of
which requires higher levels of abstraction and the integration of chemical knowledge and advanced mathematics. Subspace clustering and multi-view clustering, still
largely unexplored in the field of chemistry, will provide solutions to the problem of
automatic or semi-automatic analysis of ultra-large chemical libraries.
Acknowledgments The three authors are members of CONICET and UNLP. They thank these
institutions and FONCYT (PICTs 2019-0984, 2019-1075, 2021-0720) for their financial support.
References
1. Adnan, M., Slavic, G., Martin Gomez, D., Marcenaro, L., & Regazzoni, C. (2023). Systematic
and comprehensive review of clustering and multi-target tracking techniques for LiDAR point
clouds in autonomous driving applications. Sensors (Basel), 23, 6119.
2. Ackerman, M., & Ben-David, S. (2009). Clusterability: A theoretical study. Proceedings of the
Twelfth International Conference on Artificial Intelligence and Statistics, PMLR, 5,1–8.
3. Arbelaitz, I., Gurrutxaga, I., Muguerza, J., Pérez, J. M., & Perona, I. (2013). An extensive
comparative study of cluster validity indices. Pattern Recognition, 46, 243–256.
4. Backman, T. W., Cao, Y., & Girke, T. (2011). ChemMine tools: An online service for analyzing
and clustering small molecules. Nucleic Acids Research, 39, W486–W491.
5. Baker, F. B., & Hubert, L. J. (1975). Measuring the power of hierarchical cluster analysis.
Journal of the American Statistical Association, 70,31–38.
6. Böcker, A., Derksen, S., Schmidt, E., Teckentrup, A., & Schneider, G. (2005). A hierarchical
clustering approach for large compound libraries. Journal of Chemical Information and Model-
ing, 45, 807–815.
7. Breckenridge, J. N. (2000). Validating cluster analysis: Consistent replication and symmetry.
Multivariate Behavioral Research, 35, 261–285.
8. Brock, G., Pihur, V., Datta, S., & Datta, S. (2008). clValid: An R package for cluster validation.
Journal of Statistical Software, 25,1–22.
9. Brooks, J. L. (2014). Traditional and new principles of perceptual grouping. In J. Wagemans
(Ed.), The Oxford handbook of perceptual organization (pp. 57–87). Oxford University Press.
10. Butina, D. (1999). Unsupervised data base clustering based on daylight’s fingerprint and
tanimoto similarity: A fast and automated way to cluster small and large data sets. Journal of
Chemical Information and Computer Sciences, 39, 747–750.
11. Calinski, R. B., & Harabasz, J. (1974). A dendrite method for cluster analysis. Communications
in Statistics, 3,1–27.
12. Cao, Y., Charisi, A., Cheng, L. C., Jiang, T., & Girke, T. (2008). ChemmineR: A compound
mining framework for R. Bioinformatics, 24(15), 1733–1734.
13. Cao, Z., & Xie, X. (2024). Structure learning with consensus label information for multi-view
unsupervised feature selection. Expert Systems with Applications, 238, 121893.

5 Clustering of Small Molecules 127
14. Charrad, M., Ghazzali, N., Boiteau, V., & Niknafs, A. (2014). NbClust: An R package for
determining the relevant number of clusters in a data set. Journal of Statistical Software, 61,
1–36.
15. Davies, D., & Bouldin, D. W. (1979). A cluster separation measure. IEEE Transactions on
Pattern Analysis and Machine Intelligence, 1, 224–227.
16. Domingo-Fernández, D., Gadiya, Y., Mubeen, S., Healey, D., Norman, B. H., & Colluru,
V. (2023). Exploring the known chemical space of the plant kingdom: Insights into taxonomic
patterns, knowledge gaps, and bioactive regions. Journal of Cheminformatics, 15, 107.
17. Dunn, J. C. (1974). Well-separated clusters and optimal fuzzy partitions. Journal of Cybernet-
ics, 4,95–104.
18. Everitt, B. S., Landau, S., Leese, M., & Stahl, D. (2011). Cluster analysis (5th ed., p. 71). Wiley.
19. Guo, J., Sun, Y., Gao, J., Hu, Y., & Yin, N. (2022). Rank consistency induced multiview
subspace clustering via low-rank matrix factorization. IEEE Transactions on Neural Networks
and Learning Systems, 33, 3157–3170.
20. Gramatica, P. (2013). On the development and validation of QSAR models. Methods in
Molecular Biology, 930, 499–526.
21. Hadipour, H., Liu, C., Davis, R., Cardona, S. T., & Hu, P. (2002). Deep clustering of small
molecules at large-scale via variational autoencoder embedding and K-means. BMC Bioinfor-
matics, 23(Suppl. 4), 132.
22. Handl, J., Knowles, J., & Kell, D. B. (2005). Computational cluster validation in post-genomic
data analysis. Bioinformatics, 21, 3201–3212.
23. Hawkins, D. M., Basak, S. C., & Mills, D. (2003). Assessing model fit by cross-validation.
Journal of Chemical Information and Computer Sciences, 43, 579–586.
24. Hernández-Hernández, S., & Ballester, P. J. (2023). On the best way to cluster NCI-60
molecules. Biomolecules, 13, 498.
25. Karatzas, E., Zamora, J. E., Athanasiadis, E., Dellis, D., Cournia, Z., Spyrou, G. M., Thomas,
J. B., & Snow, C. C. (2020). ChemBioServer 2.0: An advanced web server for filtering,
clustering and networking of chemical compounds facilitating both drug discovery and
repurposing. Bioinformatics, 36(8), 2602–2604.
26. Kaufman, L., & Rousseeuw, P. J. (1990). Partitioning around medoids (program PAM). In
Finding groups in data: An introduction to cluster analysis. Wiley.
27. Ketchen, D. J., Thomas, J. B., & Snow, C. C. (1993). Organizational configurations and
performance: A comparison of theoretical approaches. Academy of Management Journal, 36,
1278–1313.
28. Ketchen, D. J., & Shook, C. L. (1996). The application of cluster analysis in strategic
management research: An analysis and critique. Strategic Management Journal, 17, 441–458.
https://doi.org/10.1002/(SICI)1097-0266(199606)17:6<441::AID-SMJ819>3.0.CO;2-G
29. Krieger, A. M., & Green, P. E. (1999). A cautionary note on using internal cross validation to
select the number of clusters. Psychometrika, 64, 341–353.
30. Leonard, J. T., & Roy, K. (2006). On selection of training and test sets for the development of
predictive QSAR models. QSAR and Combinatorial Science, 25, 235–251.
31. Lupyan, G. (2008). The conceptual grouping effect: Categories matter (and named categories
matter more). Cognition, 108 , 566–577.
32. Mayr, A., Klambauer, G., Unterthiner, T., Steijaert, M., Wegner, J. K., Ceulemans, H., Clevert,
D. A., & Hochreiter, S. (2018). Large-scale comparison of machine learning methods for drug
target prediction on ChEMBL. Chemical Science, 9, 5441–5451.
33. McClain, J. O., & Rao, V. R. (1975). CLUSTISZ: A program to test for the quality of clustering
of a set of objects. Journal of Marketing Research, 12, 456–460.
34. McKelvey, B. (1975). Guidelines for empirical classification of organizations. Administrative
Science Quarterly, 20, 509525.
35. Milligan, G. W. (1980). An examination of the effect of six types of error perturbation on fifteen
clustering algorithms. Psychometrika, 45, 325–342.
36. Milligan, G. W. (1981). A Monte Carlo study of thirty internal criterion measures for cluster
analysis. Psychometrika, 46, 187–199.

128 A. Talevi et al.
37. Murtagh, F., & Contreras, P. (2017). Algorithms for hierarchical clustering: An overview. Wiley
Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 2, e1219.
38. Prada Gori, D. N., Llanos, M. A., Bellera, C. L., Talevi, A., & Alberca, L. N. (2022a). iRaPCA
and SOMoC: Development and validation of web applications for new approaches for the
clustering of small molecules. Journal of Chemical Information and Modeling, 62, 2987–2998.
39. Prada Gori, D. N., Alberca, L. N., Rodriguez, S., Llanos, M. A., Bellera, C. L., & Talev,
A. (2022b). LIDeB tools: A Latin American resource of freely available, open-source
cheminformatics apps. Artificial Intelligence in the Life Sciences, 2, 100049.
40. Punj, G. N., & Stewart, D. W. (1983). Cluster analysis in marketing research: Review and
suggestions. Journal of Marketing Research, 20, 134–148.
41. Risser-Maroix, O., Marzouki, A., Djeghim, H., Kurtz, C., & Lomenie, N. (2021, September).
Learning an adaptation function to assess image visual similarities. Paper presented at ORASIS
2021, Centre National de la Recherche Scientifique. Available from https://hal.science/hal-0333
9731v2/document. Accessed 17 Nov 2023.
42. Rivera-Borroto, O. M., Marrero-Ponce, Y., García-de la Vega, J. M., & Grau-Ábalo, R. C.
(2011). Comparison of combinatorial clustering methods on pharmacological data sets
represented by machine learning-selected real molecular descriptors. Journal of Chemical
Information and Modeling, 51, 3036–3049.
43. Rousseeuw, P. (1987). Silhouettes: A graphical aid to the interpretation and validation of cluster
analysis. Journal of Computational and Applied Mathematics, 20,53–65.
44. Schubert, E. (2023). Stop using the elbow criterion for k-means and how to choose the number
of clusters instead. ACM SIGKDD Explorations Newsletter, 25,36–42.
45. Seger, C. A., & Miller, E. K. (2010). Category learning in the brain. Annual Review of
Neuroscience, 33, 203–219.
46. Sheikholeslami, C., Chatterjee, S., & Zhang, A. (2000). WaveCluster: A multi-resolution
clustering approach for very large spatial database. VLDB Journal, 8, 289– 304.
47. Tan, P. N., Steinbach, M., & Kumar, V. (2005). Cluster analysis: Basic concepts and algorithms. In Introduction to data mining. Addison-Wesley Longman Publishing.
48. Tichý, M., & Rucki, M. (2009). Validation of QSAR models for legislative purposes. Interdis-
ciplinary Toxicology, 2, 184–186.
49. Tropsha, A. (2010). Best practices for QSAR model development, validation, and exploitation.
Molecular Informatics, 29, 476–488.
50. Tropsha, A., Gramatica, P., & Gombar, V. K. (2003). The importance of being earnest:
Validation is the absolute essential for successful application and interpretation of QSPR
models. QSAR and Combinatorial Science, 22,69–77.
51. Veríssimo, G. C., Pantaleão, S. Q., Fernandes, P. O., Gertrudes, J. C., Kronenberger, T.,
Honorio, K. M., & Maltarollo, V. G. (2023). MASSA Algorithm: An automated rational
sampling of training and test subsets for QSAR modeling. Journal of Computer-Aided Molec-
ular Design, 37, 735–754.
52. Virshup, A. M., Contreras-García, J., Wipf, P., Yang, W., & Beratan, D. N. (2013). Stochastic
voyages into uncharted chemical space produce a representative library of all possible drug-like
compounds. Journal of the American Chemical Society, 135, 7296–
53. Voicu, A., Duteanu, N., Voicu, M., Vlad, D., & Dumitrascu, V. (2020). The rcdk and cluster R
packages applied to drug candidate selection. Journal of Cheminformatics, 12(1), 3.
54. Yang, Y., Yao, K., Repasky, M. P., Leswing, K., Abel, R., Shoichet, B. K., & Jerome, S. V.
(2021). Efficient exploration of chemical space with docking and deep learning. Journal of
Chemical Theory and Computation, 17, 7106–7119.
55. Yu, L., He, X., Fang, X., Liu, L., & Liu, J. (2023). Deep learning with geometry-enhanced
molecular representation for augmentation of large-scale docking-based virtual screening.
Journal of Chemical Information and Modeling, 63, 6501–6514.
56. Zhang, C., Huang, W., Niu, T., Liu, Z., Li, G., & Cao, D. (2023). Review of clustering
technology and its application in coordinating vehicle subsystems. Automotive Innovation, 6,
89–115.
7303.

5 Clustering of Small Molecules 129
57. Zhang, C., Fu, H., Hu, Q., Cao, X., Xie, Y., Tao, D., & Xu, D. (2020). Generalized latent multiview subspace clustering. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41,
86–99.
58. Lopez-Del Rio, A., Nonell-Canals, A., Vidal, D., and Perera-Lluna, A. (2019). Evaluation of
cross-validation strategies in sequence-based binding prediction using deep learning. J. Chem.
Inf. Model 59, 1645–1657.
59. Harris, C. J., Hill, R. D., Sheppard, D. W., Slater, M. J., and Stouten, P. F. (2011). The design
and application of target-focused compound libraries. Comb. Chem. High. Throughput Screen
14, 521–531.
60. MacQueen, J. (1967). “Some methods for classification and analysis of multivariate observations,” in Proceedings of the 5th Berkeley Symposium on Mathematical Statistics and Probability Volume 1. Nerkeley. Editors L. M. Le Cam, and J. Neyman (University of California
Press), 281–297.
61. Frey, T., & Van Groenewoud, H. (1972). A cluster analysis of the D-squared matrix of white
spruce stands in Saskatchewan based on the maximum-minimum principle.Journal of Ecology,
60, 873–886.
62. Milligan, G.W., Cooper, M.C. An examination of procedures for determining the number of
clusters in a data set. Psychometrika 50, 159–179 (1985).
63. Dimitriadou E (2023) cclust: convex clustering methods and clustering indexes http://CRAN.Rproject.org/package=cclust. R package version 0.6-26

Chapter 6
QSAR and Machine Learning Predictors
Philipe Oliveira Fernandes and Vinicius Gonçalves Maltarollo
Abstract This chapter de lves into the fundamental principles and applications of
quantitative structure–activity relationship (QSAR) and machine learning (ML)based predictors in the realm of drug design and chemical biology. QSAR establishes a quantitative relationship between the chemical structure of molecules and
their biological activities or physicochemical properties. The evolution of QSAR
from its first reports to its modern applications was covered comprising the theoretical foundations, encompassing descriptors, mathematical models (followed by brief
examples of ML applied to this field), and statistical validation techniques employed
in QSAR analysis. Interestingly, many recognized and accepted good practices and
validation protocols align with OECD guidelines for QSAR applications for regulatory purposes. In this sense, notably, the QSAR field becam e important outside of
the academic boundaries. Additionally, this chapte r discusses current challenges and
emerging trends in QSAR research, including the incorporation of machine learning
algorithms and big data analytics for enhanced predictive accuracy and applicability.
Overall, this chapter serves as a comprehensive guide for researchers and practitioners in understanding and leveraging QSAR as a pivotal tool in rational drug
design and chemical biology.
Keywords QSAR · Machine learning · Molecular descriptors · OECD principles
1 Historical Background
The Quantitative Structure–Activity Relationship (QSAR) field was introduced
when researchers tried to correlate the molecular structure of similar compounds to
a specific biological activity. This process can be traced back to 1863 when Cros first
reported a correlation between the toxicity of primary aliphatic alcohols and their
water solubility [1], a pioneer work in this field. Another notable work from the
P. O. Fernandes · V. G. Maltarollo (✉)
Departamento de Produtos Farmacêuticos, Faculdade de Farmácia, Universidade Federal de
Minas Gerais, Belo Horizonte, Minas Gerais, Brazil
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2024
V. G. Maltarollo (ed.), Computer-Aided and Machine Learning-Driven Drug
Design, Computer-Aided Drug Discovery and Design 3,
https://doi.org/10.1007/978-3-031-76718-0_6
131
Соседние файлы в папке Библиотека им академика М.И. Перельмана
