Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5858_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
30.08.2026
Размер:
43 Мб
Скачать
Chapter 10
https://t.me/med1917
Two- and Three-Dimensional Molecular Representations in Ligand-Based Approaches
Tomoyuki Miyao and Kimito Funatsu
10.1 Introduction
Molecular representations, also termed molecular descriptors [1], play a central role in any quantitative analysis of compounds. In the context of drug discovery, molec­ular descriptors are indispensable for virtual screening (VS) and potency prediction. In VS, a small molecule is first converted to a molecular descriptor, which usually forms a vector of real numbers. The vector is then compared with vectors from bioac­tive molecules for similarity searching and is used as input for a machine learning (ML) model for activity prediction. Recent technological advances in deep neural network modeling enable us to predict molecular properties in an end-to-end manner without explicit derivation of molecular descriptors. In these neural network architec­tures, structural formulas are represented as strings [ which are inputs of the model. Furthermore, a set of atomic coordinates with atomic elements [ chemical calculations. Nevertheless, the prediction accuracy of such deep neural network models does not always surpass conventional approaches using molecular descriptors and ML models. For example, a combination of an extended connec­tivity fingerprint (ECFP) [ rable prediction accuracy with graph-neural network models for biological activity prediction of multi-target annotated compound data sets [ from high-throughput screening data from the PubChem database [
6, 7] can also become an input of neural network models like quantum
8
] and random forests showed overall stable and compa-
2, 3] or chemical graphs [4, 5],
9
], which was compiled
10
]. Recently, a
T. Miyao Data Science Center, Graduate School of Science and Technology, Nara Institute of Science and Technology, Takayama-Cho, Ikoma, Nara, Japan e-mail: miyao@dsc.naist.jp
K. Funatsu (B) Data Science Center, Nara Institute of Science and Technology, Takayama-Cho, Ikoma, Nara, Japan e-mail: funatsu@dsc.naist.jp
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2024 H. Satoh et al. (eds.), Drug Development Supported by Informatics,
https://doi.org/10.1007/978-981-97-4828-0_10
175
176 T. Miyao a nd K. Funatsu
https://t.me/med1917
large-scale extensive study using MoleculeNet data sets [11], and a series of opioid­related data sets revealed that graph-neural networks and pre-trained SMILES string­based neural networks did not show overall better performance than conventional ML models [ neural networks [ combination of molecular descriptors and ML models, the prediction process always starts from the chemical structures of molecules. Thus, the choice of the dimension of descriptors, two- or three-dimensional, is crucial for activity or property prediction.
Physically, molecules exist in three-dimensional space; thus, three-dimensional representations are expected to be a prerequisite for any activity/property prediction. Actually, many successful applications using three-dimensional molecular similarity searches have been reported. For example, with a three-dimensional similarity search, Naylor et al. identified a chemical probe for nicotinic acid adenine dinucleotide phosphate (NAADP). The probe is structurally distinct from NAADP thanks to the representation focusing on the pharmacophore and electrostatic surface of NAADP
14]. Nevertheless, in the context of similarity searching, two-dimensional molecular
[ representations are not always inferior to three-dimensional representations, which is supported by many prospective VS applications using two-dimensional descriptors
15].
[
In this chapter, we discuss three topics focusing on representation differences in ligand-based VS: (1) the importance of bioactive conformers, (2) the relative merit of using three-dimensional molecular representation in activity prediction, and (3) potency prediction.
12]. The findings are consistent with a similar study focusing on graph
13]. Irrespective of employing end-to-end deep neural networks or a
10.2 Importance of Conformation in Similarity Searching
Three-dimensional molecular representations depend highly on the molecular conformation. Different conformers can generate different vectors of descriptors. In similarity searching, a set of conformers (conformer ensemble) of active compounds is used as queries for searching for geometrically similar compounds. As expected, the bioactive conformation is superior to t hat of other conformations. However, in ligand-based VS, this premise may not always hold for discerning a small number of active compounds from a large number of inactive ones, as exemplified by low hit rates from high-throughput screening (between 0.1% and 2%) [ conformation is important when analyzing molecular interactions between a ligand and a target macromolecule. Such high-resolution conformations may be different from an ideal conformer for avoiding a large number of false positives. Kirchmair et al. conducted retrospective three-dimensional similarity searching [ directory of useful decoys (DUD-decoys) data sets [ biological targets with 98,266 compounds. In their study, no significant performance difference was observed in various VS metric values, e.g., the area under the receiver operating characteristic curve (AUC-ROC) when comparing query conformers were
16]. The precise
]using the
17
18] consisting of 40 different
10 Two- and Three-Dimensional Molecular Representations … 177
https://t.me/med1917
systematically generated by several conformer generators, including CORINA [19,
20] and OMEGA [21].
We also conducted a similar but different retrospective experiment using similarity searching [
23] was used to score a three-dimensional similarity between a query and a screening
[ compound. ROCS is a tool to align two molecules in the three-dimensional space by translation and rotation based on a metric. As a metric, TanimotoCombo, which consists of shape similarity and pharmacophore equivalence, was used. A set of bioac­tive analogs for a co-crystallized ligand against the same target was collected from the ChEMBL database [ analogs were superimposed on the conformation of the corresponding co-crystallized ligand. Thus, the analog conformations are likely similar to the bound conforma­tion (bioactive-like conformation). Two types of conformations were also generated, energy-minimized and least similar to the bioactive-like conformations. The energy­minimized conformation was the conformation optimized with the MMFF94s force field from several conformers. The least similar to bioactive conformation exhibits the largest root mean square distance (RMSD) against the bioactive-like conformer. As a VS setting, the selected targets were thrombin (THRB), factor Xa (FA10), leukotriene A4 hydrolase (LKHA4), beta-secretase 1 (BACE1), and heat shock protein 90 alpha (HSP90A). The screening active and putative inactive compounds were taken directly from the Enhanced DUD database [
For each X-ray ligand, conformations of 10 randomly selected active analogs were individually used as the search query. The example results are summarized in Fig. ROC among the three types of conformations for these two targets. Moreover, for each target, the least similar conformation outperformed the bioactive-like confor­mation in four of the five X-ray ligands. For the rest of the five targets, FA10 and LKHA4 showed superiority in the bioactive-like conformation; however, the result was not consistent among different X-ray ligands, even for the same target. This result can hardly be rationalized by assuming that the “true” conformation gives the best prediction accuracy. In ligand-based VS, bioactive conformations may not be impor­tant for screening diverse compounds. Going a step further, we also benchmarked virtually generated analogs as a set of queries for VS to test the importance of query compounds. To our surprise, the screening results suggested that ensembles of virtual analog compounds with unknown activity states also improved performance. Addi­tionally, there were only a few advantages of using biological analogs when measured on a three-dimensional similarity metric, i.e., TanimotoCombo [ lution necessary for discerning diverse compounds may differ from the resolution required for conformational recognition or structurally similar analogs. Instead, a fuzzy query to identify the chemotype (compound class) of active compounds may be more important in VS.
22]. TanimotoCombo in the rapid overlay of chemical structures (ROCS)
24] using the matched molecular pair technique [25]. These
26
].
10.1 for THRB and BACE1. No significant difference was observed in AUC-
23
]. Again, the reso-
178 T. Miyao a nd K. Funatsu
https://t.me/med1917
Fig. 10.1 Differences in AUC-ROC between query conformations. For the 10 randomly selected active compounds, the differences in AUC-ROC values between energy-minimized and bioactive­like conformations (green) and between least similar and bioactive-like conformations (orange) are shown as box plots. Two targets: thrombin (THRB), left, and beta-secretase 1 (BACE1), right, are shown
10.3 Three-Dimensional vs Two-Dimensional
Representations in Activity Prediction
10.3.1 Performance Differences in Similarity Searches
The relative merit of three-dimensional molecular representations in activity predic­tions is controversial. In the previous section, we were unable to conclude that a specific type of molecular conformation was preferred for VS trials. Furthermore, it is unclear which of the three-dimensional and two-dimensional molecular represen­tations is superior. In this section, we try to clarify the preference of representations
] was used as a two-
in similarity searches. ECFP with a diameter of four (ECFP4) [ dimensional representation, whereas a ROCS-derived similarity profile [
8
27, 28
]was used as a three-dimensional representation. A similarity profile consists of similarity metric values against a set of query conformers. The three-dimensional arrange­ment of atoms and pharmacophore equivalence of a test compound is indirectly quantified via similarity values against a set of reference compounds. In this study, we employed TanimotoCombo as a similarity metric. Reference compounds were active and inactive compounds in the training data set. Conformations of reference compounds were optimized in the MMFF94s force field. Active compounds were taken from the ChEMBL database for twelve biological targets, and assumed inac­tive compounds were randomly sampled (i.e., 10,000 compounds) from the ZINC database. Conformations of the screening compounds were generated by OMEGA.
To clarify the situation when ECFP4 outperformed the ROCS-derived similarity profile and vice versa, we compiled training data sets by incrementally increasing the diversity based on the molecular scaffolds (cores), as shown in Fig.
10.2. Initially,
active compounds for a target macromolecule were collected, and cores of the
10 Two- and Three-Dimensional Molecular Representations … 179
https://t.me/med1917
compounds were determined using the compound-core relationship (CCR) method
29]. The cores were hierarchically clustered, and a dendrogram was created. The
[ cores were then iteratively sampled from the bottom to the top of t he tree. At each point of sampling a new core, a new training data set was created. Two t est data sets were created during training data set compilation, one consisting of compounds with the most distinct cores (distinct set) and the other with the same cores as in the cumulative training data sets (regular set).
The screening performance for the regular and distinct data sets in AUC-ROC are reported in Fig.
10.3 for adenosine A2a receptor ligands (A), cannabinoid CB2
receptor ligands (B), and delta-opioid receptor ligands (C), as example cases. For these data sets, the AUC-ROC values converged to 1.0 irrespective of the represen­tations when the average similarity for training CPDs reached a minimum, where all cores in the test data set were contained in the training data sets. As long as test compounds were structural analogs to the training compounds, no statistically signif­icant differences were observed between two-dimensional and three-dimensional representations. However, when the training data sets were distinct from test data sets (high averaged pairwise-similarity for training CPDs in Fig.
10.3), a difference
Fig. 10.2 Workflow for preparing the training and test data sets (Reprinted with permission from [
31])
180 T. Miyao a nd K. Funatsu
https://t.me/med1917
Fig. 10.3 AUC-ROC values against the diversity of training data for similarity searching using distinct data sets. For the three data sets, adenosine A2a receptor ligands a, cannabinoid CB2 receptor ligands b, and delta-opioid receptor ligands c, similarity searching using ECFP4 (SS-2D, green, solid line) and similarity profile (SS-3D, orange, dashed line) was conducted. Each line represents mean values with standard deviations shown as shaded regions
between the two types of representations was observed, but the trend was inconsis­tent, i.e., for some targets, the two-dimensional representation was better, whereas for other targets the three-dimensional representation was better.
10.3.2 Performance Difference in Activity Prediction
ML model-based ranking is usually superior to traditional similarity searching in screening performance. Here, two- and three-dimensional molecular representa­tions were tested in combination with a binary classification ML model in terms of screening performance. A support vector machine (SVM) was selected as a clas­sifier because of the stable prediction ability in activity prediction [
30]. The same
descriptors, i.e., ECFP4 and similarity profiles, were used. For ECFP4, as a kernel function in SVM, the Tanimoto kernel was used, and for similarity profiles, the RBF kernel was used for nonlinear modeling. Tanimoto kernel, which is also called as Jaccard coefficient, is a common metric (function) to measure a similarity between two fingerprints. The screened compounds were ranked based on the signed distance from the hyperplane in the direction of active to inactive test compounds.
As for Fig. 10.3, AUC-ROC values using SVM-based approaches for the three biological targets and the distinct data sets are reported in Fig. training data set diversity. As shown in Fig.
10.4, different types of queries were
10.4 against the
also tested (SVM-3D: active and assumed inactive compounds as the query; SVM­3D-Positive: active compounds; SVM-3D-Negative: assumed inactive compounds). Clearly, using SVM had merits in prediction accuracy for two- and three-dimensional representations. Furthermore, all SVMs showed their highest performance when test compounds were analogs to the training compounds. However, in this condition, ECFP4 was slightly better than similarity profiles. Using only active compounds as references in similarity profiles did not work well; rather, diverse negative compounds
10 Two- and Three-Dimensional Molecular Representations … 181
https://t.me/med1917
Fig. 10.4 AUC-ROC values against the diversity of training data for SVM-rankings using distinct data sets. For the three data sets, adenosine A2a receptor ligands a, cannabinoid CB2 receptor ligands b, and delta-opioid receptor ligands c, SVM-ranking-based VS was conducted using ECFP4 (SVM-2D, green, solid line), similarity profile with active and inactive training compounds as the query (SVM-3D, orange, dashed line), similarity profile with active training compounds as the query (SVM-3D-Positive, blue, dotted line) and SVM-3D using only assumed inactive training compounds (SVM-3D-Negative, magenta, dashed/dotted line). Each line represents mean values with standard deviations shown as shaded regions
as references contributed to the high performance in combination with SVM. Inter­estingly, for delta-opioid receptor ligands, the three-dimensional representation was superior when SVM was used, whereas ECFP4 was superior when using simple simi­larity searching. This result indicates that the best molecular representation for VS may depend on the adopted ML model. Thus, simultaneous optimization of represen­tations and ML modeling methods is recommended. Overall, SVM-3D outperformed SVM-2D when the diversity of the training data set was limited and test compounds were not structurally similar to the training compounds. This phenomenon was also confirmed by VS trials using independently compiled diverse and distinct training and test data sets, as shown in [
31].
10.4 Three-Dimensional vs Two-Dimensional
Representations in Potency Prediction
Potency prediction and activity prediction are different tasks. In potency prediction, compounds for model construction are active compounds with potency values. In contrast, many inactive compounds are involved in models for activity prediction. Such a difference was also highlighted in the feature contributions of SVM and support vector regression (SVR) models [ compounds using a two-stage scheme of SVM and SVR models [
To understand prediction performance using two and three-dimensional molec­ular representations, the same type of experiments as in the previous section was conducted: controlling the diversity of training data sets by the iterative sampling of CCR cores, following clustering results [ as modeling methods, including SVR and random forest and neural networks. SVR achieved the most stable performance, so only the results with SVR are presented. As
], and the identification of highly active
30
32
].
33
]. Several ML algorithms were tested
182 T. Miyao a nd K. Funatsu
https://t.me/med1917
molecular representations, similarity profiles and descriptors calculated with molec­ular operating environment (MOE) software were employed as three-dimensional representations, whereas ECFP4 was used as two-dimensional representations. Simi­larity profiles employed three similarity metrics: ShapeTanimoto (shape), ColorTan­imoto (pharmacophore equivalence), and TanimotoCombo (both). Target macro­molecules were ten diverse biological targets, including acetylcholinesterase, kappa opioid receptor and coagulation factor X, and the inhibition constant (K
)was the
i
endpoint, which was used as the objective variable.
Prediction accuracies for the ten activity classes in terms of root mean square error (RMSE) against the diversity of the training data set are reported in Fig.
10.5.The
RMSE values converged to the lowest as the training data set diversity increased. Some outliers appeared for similarity profiles and MOE. Against our expectation, no merit was observed for employing similarity profiles even when the training compounds were structurally different from the test compounds. On the contrary, ECFP4, in combination with the Tanimoto kernel, showed overall stable performance irrespective of the data set diversity.
We further conducted control calculations to predict quantum chemical proper-
ties and lipophilicity (logD) provided by the QM8 data sets [
11]. Similarity profiles
clearly outperformed ECFP4 for the quantum chemical property prediction trials, whereas for logD predictions, ECFP4 was the best among the tested descriptors
33
]. In general, quantum chemical properties depend highly on the conformation of a
[ molecule, unlike logD, which is a more complicated process because it is a thermody­namic property whose values have also been measured from experiments. However, further analyses are needed to draw a clear conclusion.
Based on the results for the three types of prediction trials, QM properties, logD,
and pK
, it can be hypothesized that highly relevant representations of the target
i
property are necessary for numerical value prediction. The employed descriptors in this study might not be sufficient on this point, including ECFP4. Precise prediction
values requires information about the interaction of a small molecule (ligand)
of K
i
with its target macromolecule and the entropic contribution of surroundings, such as water molecules and the conformational change of the macromolecule. Reported K values also contain experimental uncertainty (errors). Even given X-ray complexes of ligands and macromolecules with annotated potency values (K
and IC50), the
i
prediction of potency based on mechanism-oriented descriptors is challenging, as demonstrated by retrospective studies where limited use of interaction information between a ligand and the target macromolecule was observed [
34, 35]. In other words,
the prediction model relies on memorizing activity values for various chemotypes in the training data set. From this point of view, ECFP is still a powerful representation for potency prediction.
i
10 Two- and Three-Dimensional Molecular Representations … 183
https://t.me/med1917
Fig. 10.5 Averaged prediction errors in RMSE against the diversities of the training data sets. (Reprinted with permission from [ b, coagulation factor X c, muscarinic acetylcholine receptor M3 d, neurokinin 1 receptor e, tyrosine­protein kinase ABL f, serotonin 1d (5-HT1d) receptor g, cathepsin S h, calcitonin gene-related peptide type 1 receptor i and apoptosis regulator Bcl-2 j
33
]). The targets are acetylcholinesterase a, kappa opioid receptor
10.5 Conclusions
In this chapter, we discussed molecular representations in the context of VS and potency prediction. As queries in similarity searching, the bioactive-like conforma­tion was not required for enriching bioactive compounds compared with energy­minimized and the least similar to the bioactive-like conformations. Furthermore, virtually generated analogs of an active compound improved performance in a manner similar to using actual bioactive analogs.
In activity prediction, a ROCS-derived similarity profile as a descriptor showed higher prediction accuracy than ECFP4 when the training compounds were struc­turally distinct from test compounds. Thus, three-dimensional representations show some merits in activity prediction trials. However, no significant difference was observed in prediction accuracy between the similarity profile and ECFP4 in the potency prediction. The presented results imply a different nature of activity
184 T. Miyao a nd K. Funatsu
https://t.me/med1917
and potency prediction: distinguishing active compounds from a large number of diverse compounds and estimating the degree of molecular interaction in addition to entropic effects, respectively. Taken together, representations should be selected by considering the objective for analysis.
References
1. Todeschini R, Consonni V (2000) Handbook of Molecular Descriptors. Wiley. https://doi.org/
10.1002/9783527613106
2. Irwin R, Dimitriadis S, He J, Bjerrum EJ (2022) Chemformer: A Pre-trained Transformer for Computational Chemistry. Mach Learn Sci Technol 3:015022.
2153/AC3FFB
3. Zheng S, Yan X, Yang Y, Xu J (2019) Identifying Structure-Property Relationships through SMILES Syntax Analysis with Self-Attention Mechanism. J Chem Inf Model 59:914–923.
https://doi.org/10.1021/acs.jcim.8b00803
4. Xiong Z, Wang D, Liu X, et al (2020) Pushing the Boundaries of Molecular Representation for Drug Discovery with the Graph Attention Mechanism. J Med Chem 63:8749–8760.
doi.org/10.1021/acs.jmedchem.9b00959
5. Karlov DS, Sosnin S, Fedorov MV, Popov P (2020) GraphDelta: MPNN Scoring Function for the Affinity Prediction of Protein-Ligand Complexes. ACS Omega 5:5150–5159.
org/10.1021/acsomega.9b04162
6. Smith JS, Isayev O, Roitberg AE (2017) ANI-1: An Extensible Neural Network Potential with DFT Accuracy at Force Field Computational Cost. Chem Sci 8:3192–3203.
1039/C6SC05720A
7. Schütt KT, Sauceda HE, Kindermans PJ, et al (2018) SchNet—A Deep Learning Architecture for Molecules and Materials. J Chem Phys 148:241722.
8. Rogers D, Hahn M (2010) Extended-Connectivity Fingerprints. J Chem Inf Model 50:742–754.
https://doi.org/10.1021/CI100050T
9. Rodríguez-Pérez R, Miyao T, Jasial S, et al (2018) Prediction of Compound Profiling Matrices Using Machine Learning. ACS Omega 3:4713–4723.
0462
10. Vogt M, Jasial S, Bajorath J (2018) Extracting Compound Profiling Matrices from Screening Data. ACS Omega 3:4706–4712.
11. Wu Z, Ramsundar B, Feinberg EN, et al (2018) MoleculeNet: A Benchmark for Molecular Machine Learning. Chem Sci 9:513–530.
12. Deng J, Yang Z, Wang H, et al (2023) A Systematic Study of Key Elements Underlying Molecular Property Prediction. Nat Commun 14:6395.
41948-6
13. Jiang D, Wu Z, Hsieh CY, et al (2021) Could Graph Neural Networks Learn Better Molecular Representation for Drug Discovery? A Comparison Study of Descriptor-based and Graph-based Models. J Cheminform 13:12.
14. Naylor E, Arredouani A, Vasudevan SR, et al (2009) Identification of a Chemical Probe for NAADP by Virtual Screening. Nat Chem Biol 5:220–226.
io.150
15. Stumpfe D, Bajorath J (2013) Critical Assessment of Virtual Screening for Hit Identification. In: J Bajorath (ed) Chemoinformatics Drug Discovery, Wiley, pp. 113–130.
1002/9781118742785.CH6
16. Lipinski CA (2009) Overview of Hit to Lead: The Medicinal Chemist’s Role from HTS Retest to Lead Optimization Hand Off. In: MM Hayward et al (eds.) Lead-Seeking Approaches. Topics in Medicinal Chemistry, vol 5. Springer, Berlin, Heidelberg, pp 1–24.
7355_2009_4
https://doi.org/10.1021/acsomega.8b00461
https://doi.org/10.1039/C7SC02664A
https://doi.org/10.1186/s13321-020-00479-8
https://doi.org/10.1088/2632-
https://
https://doi.
https://doi.org/10.
https://doi.org/10.1063/1.5019779
https://doi.org/10.1021/acsomega.8b0
https://doi.org/10.1038/s41467-023-
https://doi.org/10.1038/nchemb
https://doi.org/10.
https://doi.org/10.1007/