Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5580_Библиотеки_им_академика_М_И_Перельмана
.pdf
Chapter 10
https://t.me/med1917
Two- and Three-Dimensional Molecular
Representations in Ligand-Based
Approaches
Tomoyuki Miyao and Kimito Funatsu
10.1 Introduction
Molecular representations, also termed molecular descriptors [1], play a central role
in any quantitative analysis of compounds. In the context of drug discovery, molecular descriptors are indispensable for virtual screening (VS) and potency prediction.
In VS, a small molecule is first converted to a molecular descriptor, which usually
forms a vector of real numbers. The vector is then compared with vectors from bioactive molecules for similarity searching and is used as input for a machine learning
(ML) model for activity prediction. Recent technological advances in deep neural
network modeling enable us to predict molecular properties in an end-to-end manner
without explicit derivation of molecular descriptors. In these neural network architectures, structural formulas are represented as strings [
which are inputs of the model. Furthermore, a set of atomic coordinates with atomic
elements [
chemical calculations. Nevertheless, the prediction accuracy of such deep neural
network models does not always surpass conventional approaches using molecular
descriptors and ML models. For example, a combination of an extended connectivity fingerprint (ECFP) [
rable prediction accuracy with graph-neural network models for biological activity
prediction of multi-target annotated compound data sets [
from high-throughput screening data from the PubChem database [
6, 7] can also become an input of neural network models like quantum
8
] and random forests showed overall stable and compa-
2, 3] or chemical graphs [4, 5],
9
], which was compiled
10
]. Recently, a
T. Miyao
Data Science Center, Graduate School of Science and Technology, Nara Institute of Science and
Technology, Takayama-Cho, Ikoma, Nara, Japan
e-mail: miyao@dsc.naist.jp
K. Funatsu (B)
Data Science Center, Nara Institute of Science and Technology, Takayama-Cho, Ikoma, Nara,
Japan
e-mail: funatsu@dsc.naist.jp
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2024
H. Satoh et al. (eds.), Drug Development Supported by Informatics,
https://doi.org/10.1007/978-981-97-4828-0_10
175

176 T. Miyao a nd K. Funatsu
https://t.me/med1917
large-scale extensive study using MoleculeNet data sets [11], and a series of opioidrelated data sets revealed that graph-neural networks and pre-trained SMILES stringbased neural networks did not show overall better performance than conventional
ML models [
neural networks [
combination of molecular descriptors and ML models, the prediction process always
starts from the chemical structures of molecules. Thus, the choice of the dimension of
descriptors, two- or three-dimensional, is crucial for activity or property prediction.
Physically, molecules exist in three-dimensional space; thus, three-dimensional
representations are expected to be a prerequisite for any activity/property prediction.
Actually, many successful applications using three-dimensional molecular similarity
searches have been reported. For example, with a three-dimensional similarity search,
Naylor et al. identified a chemical probe for nicotinic acid adenine dinucleotide
phosphate (NAADP). The probe is structurally distinct from NAADP thanks to the
representation focusing on the pharmacophore and electrostatic surface of NAADP
14]. Nevertheless, in the context of similarity searching, two-dimensional molecular
[
representations are not always inferior to three-dimensional representations, which is
supported by many prospective VS applications using two-dimensional descriptors
15].
[
In this chapter, we discuss three topics focusing on representation differences in
ligand-based VS: (1) the importance of bioactive conformers, (2) the relative merit
of using three-dimensional molecular representation in activity prediction, and (3)
potency prediction.
12]. The findings are consistent with a similar study focusing on graph
13]. Irrespective of employing end-to-end deep neural networks or a
10.2 Importance of Conformation in Similarity Searching
Three-dimensional molecular representations depend highly on the molecular
conformation. Different conformers can generate different vectors of descriptors. In
similarity searching, a set of conformers (conformer ensemble) of active compounds
is used as queries for searching for geometrically similar compounds. As expected,
the bioactive conformation is superior to t hat of other conformations. However, in
ligand-based VS, this premise may not always hold for discerning a small number
of active compounds from a large number of inactive ones, as exemplified by low
hit rates from high-throughput screening (between 0.1% and 2%) [
conformation is important when analyzing molecular interactions between a ligand
and a target macromolecule. Such high-resolution conformations may be different
from an ideal conformer for avoiding a large number of false positives. Kirchmair
et al. conducted retrospective three-dimensional similarity searching [
directory of useful decoys (DUD-decoys) data sets [
biological targets with 98,266 compounds. In their study, no significant performance
difference was observed in various VS metric values, e.g., the area under the receiver
operating characteristic curve (AUC-ROC) when comparing query conformers were
16]. The precise
]using the
17
18] consisting of 40 different

10 Two- and Three-Dimensional Molecular Representations … 177
https://t.me/med1917
systematically generated by several conformer generators, including CORINA [19,
20] and OMEGA [21].
We also conducted a similar but different retrospective experiment using similarity
searching [
23] was used to score a three-dimensional similarity between a query and a screening
[
compound. ROCS is a tool to align two molecules in the three-dimensional space
by translation and rotation based on a metric. As a metric, TanimotoCombo, which
consists of shape similarity and pharmacophore equivalence, was used. A set of bioactive analogs for a co-crystallized ligand against the same target was collected from
the ChEMBL database [
analogs were superimposed on the conformation of the corresponding co-crystallized
ligand. Thus, the analog conformations are likely similar to the bound conformation (bioactive-like conformation). Two types of conformations were also generated,
energy-minimized and least similar to the bioactive-like conformations. The energyminimized conformation was the conformation optimized with the MMFF94s force
field from several conformers. The least similar to bioactive conformation exhibits the
largest root mean square distance (RMSD) against the bioactive-like conformer. As a
VS setting, the selected targets were thrombin (THRB), factor Xa (FA10), leukotriene
A4 hydrolase (LKHA4), beta-secretase 1 (BACE1), and heat shock protein 90 alpha
(HSP90A). The screening active and putative inactive compounds were taken directly
from the Enhanced DUD database [
For each X-ray ligand, conformations of 10 randomly selected active analogs
were individually used as the search query. The example results are summarized in
Fig.
ROC among the three types of conformations for these two targets. Moreover, for
each target, the least similar conformation outperformed the bioactive-like conformation in four of the five X-ray ligands. For the rest of the five targets, FA10 and
LKHA4 showed superiority in the bioactive-like conformation; however, the result
was not consistent among different X-ray ligands, even for the same target. This result
can hardly be rationalized by assuming that the “true” conformation gives the best
prediction accuracy. In ligand-based VS, bioactive conformations may not be important for screening diverse compounds. Going a step further, we also benchmarked
virtually generated analogs as a set of queries for VS to test the importance of query
compounds. To our surprise, the screening results suggested that ensembles of virtual
analog compounds with unknown activity states also improved performance. Additionally, there were only a few advantages of using biological analogs when measured
on a three-dimensional similarity metric, i.e., TanimotoCombo [
lution necessary for discerning diverse compounds may differ from the resolution
required for conformational recognition or structurally similar analogs. Instead, a
fuzzy query to identify the chemotype (compound class) of active compounds may
be more important in VS.
22]. TanimotoCombo in the rapid overlay of chemical structures (ROCS)
24] using the matched molecular pair technique [25]. These
26
].
10.1 for THRB and BACE1. No significant difference was observed in AUC-
23
]. Again, the reso-

178 T. Miyao a nd K. Funatsu
https://t.me/med1917
Fig. 10.1 Differences in AUC-ROC between query conformations. For the 10 randomly selected
active compounds, the differences in AUC-ROC values between energy-minimized and bioactivelike conformations (green) and between least similar and bioactive-like conformations (orange) are
shown as box plots. Two targets: thrombin (THRB), left, and beta-secretase 1 (BACE1), right, are
shown
10.3 Three-Dimensional vs Two-Dimensional
Representations in Activity Prediction
10.3.1 Performance Differences in Similarity Searches
The relative merit of three-dimensional molecular representations in activity predictions is controversial. In the previous section, we were unable to conclude that a
specific type of molecular conformation was preferred for VS trials. Furthermore, it
is unclear which of the three-dimensional and two-dimensional molecular representations is superior. In this section, we try to clarify the preference of representations
] was used as a two-
in similarity searches. ECFP with a diameter of four (ECFP4) [
dimensional representation, whereas a ROCS-derived similarity profile [
8
27, 28
]was
used as a three-dimensional representation. A similarity profile consists of similarity
metric values against a set of query conformers. The three-dimensional arrangement of atoms and pharmacophore equivalence of a test compound is indirectly
quantified via similarity values against a set of reference compounds. In this study,
we employed TanimotoCombo as a similarity metric. Reference compounds were
active and inactive compounds in the training data set. Conformations of reference
compounds were optimized in the MMFF94s force field. Active compounds were
taken from the ChEMBL database for twelve biological targets, and assumed inactive compounds were randomly sampled (i.e., 10,000 compounds) from the ZINC
database. Conformations of the screening compounds were generated by OMEGA.
To clarify the situation when ECFP4 outperformed the ROCS-derived similarity
profile and vice versa, we compiled training data sets by incrementally increasing the
diversity based on the molecular scaffolds (cores), as shown in Fig.
10.2. Initially,
active compounds for a target macromolecule were collected, and cores of the

10 Two- and Three-Dimensional Molecular Representations … 179
https://t.me/med1917
compounds were determined using the compound-core relationship (CCR) method
29]. The cores were hierarchically clustered, and a dendrogram was created. The
[
cores were then iteratively sampled from the bottom to the top of t he tree. At each
point of sampling a new core, a new training data set was created. Two t est data
sets were created during training data set compilation, one consisting of compounds
with the most distinct cores (distinct set) and the other with the same cores as in the
cumulative training data sets (regular set).
The screening performance for the regular and distinct data sets in AUC-ROC
are reported in Fig.
10.3 for adenosine A2a receptor ligands (A), cannabinoid CB2
receptor ligands (B), and delta-opioid receptor ligands (C), as example cases. For
these data sets, the AUC-ROC values converged to 1.0 irrespective of the representations when the average similarity for training CPDs reached a minimum, where
all cores in the test data set were contained in the training data sets. As long as test
compounds were structural analogs to the training compounds, no statistically significant differences were observed between two-dimensional and three-dimensional
representations. However, when the training data sets were distinct from test data
sets (high averaged pairwise-similarity for training CPDs in Fig.
10.3), a difference
Fig. 10.2 Workflow for preparing the training and test data sets (Reprinted with permission from
[
31])

180 T. Miyao a nd K. Funatsu
https://t.me/med1917
Fig. 10.3 AUC-ROC values against the diversity of training data for similarity searching using
distinct data sets. For the three data sets, adenosine A2a receptor ligands a, cannabinoid CB2
receptor ligands b, and delta-opioid receptor ligands c, similarity searching using ECFP4 (SS-2D,
green, solid line) and similarity profile (SS-3D, orange, dashed line) was conducted. Each line
represents mean values with standard deviations shown as shaded regions
between the two types of representations was observed, but the trend was inconsistent, i.e., for some targets, the two-dimensional representation was better, whereas
for other targets the three-dimensional representation was better.
10.3.2 Performance Difference in Activity Prediction
ML model-based ranking is usually superior to traditional similarity searching in
screening performance. Here, two- and three-dimensional molecular representations were tested in combination with a binary classification ML model in terms
of screening performance. A support vector machine (SVM) was selected as a classifier because of the stable prediction ability in activity prediction [
30]. The same
descriptors, i.e., ECFP4 and similarity profiles, were used. For ECFP4, as a kernel
function in SVM, the Tanimoto kernel was used, and for similarity profiles, the RBF
kernel was used for nonlinear modeling. Tanimoto kernel, which is also called as
Jaccard coefficient, is a common metric (function) to measure a similarity between
two fingerprints. The screened compounds were ranked based on the signed distance
from the hyperplane in the direction of active to inactive test compounds.
As for Fig. 10.3, AUC-ROC values using SVM-based approaches for the three
biological targets and the distinct data sets are reported in Fig.
training data set diversity. As shown in Fig.
10.4, different types of queries were
10.4 against the
also tested (SVM-3D: active and assumed inactive compounds as the query; SVM3D-Positive: active compounds; SVM-3D-Negative: assumed inactive compounds).
Clearly, using SVM had merits in prediction accuracy for two- and three-dimensional
representations. Furthermore, all SVMs showed their highest performance when test
compounds were analogs to the training compounds. However, in this condition,
ECFP4 was slightly better than similarity profiles. Using only active compounds as
references in similarity profiles did not work well; rather, diverse negative compounds

10 Two- and Three-Dimensional Molecular Representations … 181
https://t.me/med1917
Fig. 10.4 AUC-ROC values against the diversity of training data for SVM-rankings using distinct
data sets. For the three data sets, adenosine A2a receptor ligands a, cannabinoid CB2 receptor
ligands b, and delta-opioid receptor ligands c, SVM-ranking-based VS was conducted using ECFP4
(SVM-2D, green, solid line), similarity profile with active and inactive training compounds as the
query (SVM-3D, orange, dashed line), similarity profile with active training compounds as the
query (SVM-3D-Positive, blue, dotted line) and SVM-3D using only assumed inactive training
compounds (SVM-3D-Negative, magenta, dashed/dotted line). Each line represents mean values
with standard deviations shown as shaded regions
as references contributed to the high performance in combination with SVM. Interestingly, for delta-opioid receptor ligands, the three-dimensional representation was
superior when SVM was used, whereas ECFP4 was superior when using simple similarity searching. This result indicates that the best molecular representation for VS
may depend on the adopted ML model. Thus, simultaneous optimization of representations and ML modeling methods is recommended. Overall, SVM-3D outperformed
SVM-2D when the diversity of the training data set was limited and test compounds
were not structurally similar to the training compounds. This phenomenon was also
confirmed by VS trials using independently compiled diverse and distinct training
and test data sets, as shown in [
31].
10.4 Three-Dimensional vs Two-Dimensional
Representations in Potency Prediction
Potency prediction and activity prediction are different tasks. In potency prediction,
compounds for model construction are active compounds with potency values. In
contrast, many inactive compounds are involved in models for activity prediction.
Such a difference was also highlighted in the feature contributions of SVM and
support vector regression (SVR) models [
compounds using a two-stage scheme of SVM and SVR models [
To understand prediction performance using two and three-dimensional molecular representations, the same type of experiments as in the previous section was
conducted: controlling the diversity of training data sets by the iterative sampling
of CCR cores, following clustering results [
as modeling methods, including SVR and random forest and neural networks. SVR
achieved the most stable performance, so only the results with SVR are presented. As
], and the identification of highly active
30
32
].
33
]. Several ML algorithms were tested

182 T. Miyao a nd K. Funatsu
https://t.me/med1917
molecular representations, similarity profiles and descriptors calculated with molecular operating environment (MOE) software were employed as three-dimensional
representations, whereas ECFP4 was used as two-dimensional representations. Similarity profiles employed three similarity metrics: ShapeTanimoto (shape), ColorTanimoto (pharmacophore equivalence), and TanimotoCombo (both). Target macromolecules were ten diverse biological targets, including acetylcholinesterase, kappa
opioid receptor and coagulation factor X, and the inhibition constant (K
)was the
i
endpoint, which was used as the objective variable.
Prediction accuracies for the ten activity classes in terms of root mean square error
(RMSE) against the diversity of the training data set are reported in Fig.
10.5.The
RMSE values converged to the lowest as the training data set diversity increased.
Some outliers appeared for similarity profiles and MOE. Against our expectation,
no merit was observed for employing similarity profiles even when the training
compounds were structurally different from the test compounds. On the contrary,
ECFP4, in combination with the Tanimoto kernel, showed overall stable performance
irrespective of the data set diversity.
We further conducted control calculations to predict quantum chemical proper-
ties and lipophilicity (logD) provided by the QM8 data sets [
11]. Similarity profiles
clearly outperformed ECFP4 for the quantum chemical property prediction trials,
whereas for logD predictions, ECFP4 was the best among the tested descriptors
33
]. In general, quantum chemical properties depend highly on the conformation of a
[
molecule, unlike logD, which is a more complicated process because it is a thermodynamic property whose values have also been measured from experiments. However,
further analyses are needed to draw a clear conclusion.
Based on the results for the three types of prediction trials, QM properties, logD,
and pK
, it can be hypothesized that highly relevant representations of the target
i
property are necessary for numerical value prediction. The employed descriptors in
this study might not be sufficient on this point, including ECFP4. Precise prediction
values requires information about the interaction of a small molecule (ligand)
of K
i
with its target macromolecule and the entropic contribution of surroundings, such as
water molecules and the conformational change of the macromolecule. Reported K
values also contain experimental uncertainty (errors). Even given X-ray complexes
of ligands and macromolecules with annotated potency values (K
and IC50), the
i
prediction of potency based on mechanism-oriented descriptors is challenging, as
demonstrated by retrospective studies where limited use of interaction information
between a ligand and the target macromolecule was observed [
34, 35]. In other words,
the prediction model relies on memorizing activity values for various chemotypes in
the training data set. From this point of view, ECFP is still a powerful representation
for potency prediction.
i

10 Two- and Three-Dimensional Molecular Representations … 183
https://t.me/med1917
Fig. 10.5 Averaged prediction errors in RMSE against the diversities of the training data sets.
(Reprinted with permission from [
b, coagulation factor X c, muscarinic acetylcholine receptor M3 d, neurokinin 1 receptor e, tyrosineprotein kinase ABL f, serotonin 1d (5-HT1d) receptor g, cathepsin S h, calcitonin gene-related
peptide type 1 receptor i and apoptosis regulator Bcl-2 j
33
]). The targets are acetylcholinesterase a, kappa opioid receptor
10.5 Conclusions
In this chapter, we discussed molecular representations in the context of VS and
potency prediction. As queries in similarity searching, the bioactive-like conformation was not required for enriching bioactive compounds compared with energyminimized and the least similar to the bioactive-like conformations. Furthermore,
virtually generated analogs of an active compound improved performance in a
manner similar to using actual bioactive analogs.
In activity prediction, a ROCS-derived similarity profile as a descriptor showed
higher prediction accuracy than ECFP4 when the training compounds were structurally distinct from test compounds. Thus, three-dimensional representations show
some merits in activity prediction trials. However, no significant difference was
observed in prediction accuracy between the similarity profile and ECFP4 in
the potency prediction. The presented results imply a different nature of activity

184 T. Miyao a nd K. Funatsu
https://t.me/med1917
and potency prediction: distinguishing active compounds from a large number of
diverse compounds and estimating the degree of molecular interaction in addition to
entropic effects, respectively. Taken together, representations should be selected by
considering the objective for analysis.
References
1. Todeschini R, Consonni V (2000) Handbook of Molecular Descriptors. Wiley. https://doi.org/
10.1002/9783527613106
2. Irwin R, Dimitriadis S, He J, Bjerrum EJ (2022) Chemformer: A Pre-trained Transformer for
Computational Chemistry. Mach Learn Sci Technol 3:015022.
2153/AC3FFB
3. Zheng S, Yan X, Yang Y, Xu J (2019) Identifying Structure-Property Relationships through
SMILES Syntax Analysis with Self-Attention Mechanism. J Chem Inf Model 59:914–923.
https://doi.org/10.1021/acs.jcim.8b00803
4. Xiong Z, Wang D, Liu X, et al (2020) Pushing the Boundaries of Molecular Representation
for Drug Discovery with the Graph Attention Mechanism. J Med Chem 63:8749–8760.
doi.org/10.1021/acs.jmedchem.9b00959
5. Karlov DS, Sosnin S, Fedorov MV, Popov P (2020) GraphDelta: MPNN Scoring Function for
the Affinity Prediction of Protein-Ligand Complexes. ACS Omega 5:5150–5159.
org/10.1021/acsomega.9b04162
6. Smith JS, Isayev O, Roitberg AE (2017) ANI-1: An Extensible Neural Network Potential with
DFT Accuracy at Force Field Computational Cost. Chem Sci 8:3192–3203.
1039/C6SC05720A
7. Schütt KT, Sauceda HE, Kindermans PJ, et al (2018) SchNet—A Deep Learning Architecture
for Molecules and Materials. J Chem Phys 148:241722.
8. Rogers D, Hahn M (2010) Extended-Connectivity Fingerprints. J Chem Inf Model 50:742–754.
https://doi.org/10.1021/CI100050T
9. Rodríguez-Pérez R, Miyao T, Jasial S, et al (2018) Prediction of Compound Profiling Matrices
Using Machine Learning. ACS Omega 3:4713–4723.
0462
10. Vogt M, Jasial S, Bajorath J (2018) Extracting Compound Profiling Matrices from Screening
Data. ACS Omega 3:4706–4712.
11. Wu Z, Ramsundar B, Feinberg EN, et al (2018) MoleculeNet: A Benchmark for Molecular
Machine Learning. Chem Sci 9:513–530.
12. Deng J, Yang Z, Wang H, et al (2023) A Systematic Study of Key Elements Underlying
Molecular Property Prediction. Nat Commun 14:6395.
41948-6
13. Jiang D, Wu Z, Hsieh CY, et al (2021) Could Graph Neural Networks Learn Better Molecular
Representation for Drug Discovery? A Comparison Study of Descriptor-based and Graph-based
Models. J Cheminform 13:12.
14. Naylor E, Arredouani A, Vasudevan SR, et al (2009) Identification of a Chemical Probe for
NAADP by Virtual Screening. Nat Chem Biol 5:220–226.
io.150
15. Stumpfe D, Bajorath J (2013) Critical Assessment of Virtual Screening for Hit Identification.
In: J Bajorath (ed) Chemoinformatics Drug Discovery, Wiley, pp. 113–130.
1002/9781118742785.CH6
16. Lipinski CA (2009) Overview of Hit to Lead: The Medicinal Chemist’s Role from HTS Retest
to Lead Optimization Hand Off. In: MM Hayward et al (eds.) Lead-Seeking Approaches. Topics
in Medicinal Chemistry, vol 5. Springer, Berlin, Heidelberg, pp 1–24.
7355_2009_4
https://doi.org/10.1021/acsomega.8b00461
https://doi.org/10.1039/C7SC02664A
https://doi.org/10.1186/s13321-020-00479-8
https://doi.org/10.1088/2632-
https://
https://doi.
https://doi.org/10.
https://doi.org/10.1063/1.5019779
https://doi.org/10.1021/acsomega.8b0
https://doi.org/10.1038/s41467-023-
https://doi.org/10.1038/nchemb
https://doi.org/10.
https://doi.org/10.1007/
Соседние файлы в папке Библиотека им академика М.И. Перельмана
