Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5587_Библиотеки_им_академика_М_И_Перельмана.pdf

Evolution of Multivariate Image Analysis in QSAR
The earlier studies by Cros in 1863 at the University of Strasbourg indicated that toxicity of alcohols
to mammals increased as the water solubility of the alcohols decreased (Selassie, 2003). Afterwards,
Crum-Brown and Fraser (1869) demonstrated the concept that molecular structure influences the
biologic activity of chemical entities and that alteration in structure produces changes in biological
action. Richet (1893) proposed that toxicity of some alcohols and ethers was inversely related to the
corresponding water solubility and, seven years later, Meyer and Overton independently described a
parameter, which can be considered the precursor of the octanol/water partition coefficient (Meyer,
1899; Overton, 1901). An important point along with this trajectory was the introduction of the
Hammett constant (σ) to relate the ionization constants of meta and para substituted benzoic acids
(Hammett, 1935). This finding opened the exploration of a variety of properties that correlate with
biological outcomes, e.g. the relationship between several properties and the toxicity of homolog
series of compounds (Ferguson, 1939). In 1964, Free and Wilson (1964) developed a mathematical
procedure to encode bioactivities by means of the position of chemical groups in a given molecular
scaffold. Simultaneously, Hansch and Fujita (1964) systematized the technique when introduced logP
(the logarithm of the octanol/water partition coefficient) to correlate with the biological activity of a
series of compounds against a variety of living organisms.
Different QSAR methods appeared since the milestone work by Hansch and Fujita (1964), but
this field gained high appearance again years later, when Cramer and coworkers (1988) developed a
QSAR method based on the tridimensional structure of a congeneric series of molecules. The treatment
of the large amount of data generated was feasible because of the advances in QSAR technologies,
improved with computers and programs. According to the so called CoMFA method (Comparative
Molecular Field Analysis), each molecule of a given data set is put inside a three-dimensional lattice,
whose intersections are probed by steric and electrostatic fields, generating the molecular descriptors.
Contour plots representing structural moieties affecting either positively or negatively the bioactivity
are given, which are possibly the main outcomes of this method. A similar approach was developed
a few years later to account for the molecular similarity indices in a comparative analysis (CoMSIA);
the CoMSIA maps highlight those regions within the area occupied by the ligand skeletons that require
a particular physicochemical property important for activity. Hydrogen bond has been included as
another descriptor in recent CoMSIA models. At least, two steps are required to start building QSAR
models based on three-dimensional methods: conformational screening and three-dimensional alignment. These are not easy tasks, since the energy minimum of an optimized structure does not necessarily reflect the bioactive ligand conformation. Indeed, it has been shown for the herbicide 2,4-D
that different global minima are achieved using different levels of theory for geometry optimization
and conformational search, which are also different from the crystal, bioactive ligand conformation
(Freitas & Ramalho, 2013). Consequently, the contour plots indicating regions with positive, neutral
and negative influence on the bioactivity may not make sense, because they are dependent on the
conformation, which can be either the bioactive one or not. It is worth reinforcing that bioactivity is
almost always dependent on conformation and, therefore, it should be useful in QSAR studies. A recent
example of this is the role of the synclinal conformation of sulfonylureas to explain the inhibition of
acetohydroxyacid synthase (Jana, Delgado, & Medina, 2014); however, the correct conformation is
sometimes difficult to be precisely predicted and selected.
Multidimensional (nD) QSAR methods, such as the 4D formalism in QSAR (Hopfinger et al.,
1997), try to minimize the problem of choosing a single conformation for modeling, by performing
ensemble averaging, the “fourth dimension”. However, the bioactive conformation is not likely to be
86
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use

Evolution of Multivariate Image Analysis in QSAR
represented by an average. Introduction of receptor information in the so-called receptor-dependent
4D-QSAR (RD-4D-QSAR) (Pan, Tseng, & Hopfinger, 2003) and 5D-QSAR (Vedani & Dobler, 2002)
appears to better encode structural properties for further correlation with the dependent variables
(the bioactivities), but the predictive abilities obtained by many nD models have not been improved
significantly in comparison with traditional 3D approaches. Other dimensions have been included
in QSAR to refine the molecular representation, such as multiple solvation scenarios in the 6DQSAR, the real receptor or target-based receptor model data in the 7D-QSAR, and tautomerization
(Vedani, Dobler, & Lill, 2005; Polanski, 2009; Martin, 2009). Nevertheless, classical descriptors
representing connectivity indices, atom counting, molecular weight, molar refractivity, polarizability, hydrophobicity, orbital energies, etc. (which are physicochemical descriptors mostly derived
from 1D and 2D molecular representation) have not shown to be inferior to 3D descriptors, because
the prediction performance is often acceptable and interpretable (Brown & Martin, 1997; Estrada,
Molina, & Perdomo-López, 2001).
The predictive ability of models based on these descriptors has been significantly improved
with the development of mathematical methods to, for instance, select variables and to account for
nonlinearities. Recent examples of feature selection methods in QSAR include ant colony optimization (ACO) and forward stepwise (FS) applied in the modeling of the anti-HIV-1 activities of
3-(3,5-dimethylbenzyl)uracil derivatives (Goodarzi, Freitas, & Jensen, 2009a), as well as genetic
algorithm (GA) and successive projections algorithm (SPA) in the analysis of glycogen synthase
kinase-3β inhibitory activities of a series of compounds (Goodarzi, Freitas, & Jensen, 2009b).
Partial least squares (PLS) regression has been the preferred method to deal with a large amount of
independent variables, while multimode methods, like parallel factor (PARAFAC) and multilinear
PLS (N-PLS), have gained widespread use to deal with N-way arrays (Freitas et al., 2008). Artificial neural networks (ANN) and support vector machines (SVM) have been successfully employed
to account for nonlinearities in QSAR models, improving importantly the predictive accuracy in
comparison to multiple linear regression (Goodarzi, Freitas, & Jensen, 2009b).
In line with this, the multivariate image analysis applied to QSAR (MIA-QSAR) method matches
the characteristics of 2D QSAR approaches, because the two-dimensional representation of molecules
(molecular structure drawings) gives rise to the descriptors used in the modeling. However, 3D
information, such as enantiomers, can be encoded by drawing wedge or hashed bonds at the chiral
center, such as demonstrated in the MIA-QSPR modeling of the electrophoretic enantioseparation
of aromatic amino acids/esters (Goodarzi & Freitas, 2009). Because nothing is expected to be more
direct to represent a chemical structure than itself, such a method has provided satisfactory prediction performance in the modeling of physical, chemical and biological properties. Nevertheless,
the MIA-QSAR is far from perfection, e.g. because of the lack of interpretability for descriptors
representing chemical substituents, like atoms represented by their chemical symbols. Thus, the
MIA-QSAR method has been constantly improved and its reliability has been attested by rigorous
validation protocols, which have been continuously updated. The MIA-QSAR method and its evolution are described in the next sections and, afterwards, an application to a neglected disease is given.
Trypanosomiasis was chosen because a few studies have devoted attention to this disease, which
is quite common in tropical, third world countries and thousands of deaths occur worldwide, especially in the Latin America, due to the Chagas disease (caused by the parasitic protozoan T. cruzi).
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
87

Evolution of Multivariate Image Analysis in QSAR
MULTIVARIATE IMAGE ANALYSIS IN QSAR
Traditional MIA-QSAR
What is the best way to describe a chemical structure, which will be correlated with its biological activity?
Orbital energies, hydrophobicity, electrostatic and steric fields? According to the MIA-QSAR concept,
a chemical structure is best represented by itself, as a drawing. How the structure is drawn is the main
challenge to achieve better QSAR models. Other challenges include the use of different mathematical
techniques for regression and variable selection, as well as the search for ways to analyze the MIA
descriptors in order to indicate which pixel (or set of pixels) is responsible to increased or decreased
bioactivity and why. This chapter is mainly devoted to present the development of the MIA-QSAR
methodology and to show the way in representing chemical structures affects the predictive ability of
the MIA-QSAR method.
Chemical structures in the traditional MIA-QSAR look like wireframes, with lines representing
chemical bonds and chemical symbols to indicate atoms when necessary. Eventually, chiral centers can
be designated as wedge or hashed bonds. These chemical structures can be easily drawn as routinely
performed using a variety of programs, such as ChemDraw and ChemSketch. An example of a chemical
structure drawn with the ChemSketch program is given in Figure 1.
QSAR methods try to search for correlations between an homolog set of compounds and the corresponding bioactivities. Thus, the set of compounds retain a structural similarity (usually the pharmacophoric
group), which should be used in the 2D alignment. Such an alignment is fundamental to make variable
only those moieties not common along with the set of compounds, which are responsible for the variance
in the bioactivities of the molecules. The variable moieties are evident after superposition of all image
files, as in the example of Figure 2. In order to superpose the files for the 2D alignment, each chemical
structure should be drawn systematically, in such a way that common substructures do not vary, while
substituents or variable moieties should be drawn in the same way every time it is repeated in different
chemical structures. Each chemical structure should be pasted in a workspace with x × y dimension (e.g.
500 × 500 pixels size) using, for instance, the Paint application of the Microsoft Windows. Before saving
the images, each one should be completely selected and a given pixel, which is common along with the
set of structures, should be used to manually move (preferably using the mouse) the whole structure to
a given position in the workspace. The 2D alignment is done and the image can be saved e.g. as bitmaps
(.bmp). This procedure is pictographically represented in Figure 3. The file extension and the way in which
Figure 1. The analgesic S-ketamine drawn with the ChemSketch program
88
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use

Evolution of Multivariate Image Analysis in QSAR
substituents are represented (e.g. CH3 or Me to indicate a methyl group) have not shown to statistically
affect the model response (Goodarzi, Freitas, & Ferreira, 2009). The drawings should be readable for
further regression against the bioactivity values. Because the images are composed by pixels, these can
be converted to numbers, according to the RGB (red-green-blue) system of color composition. For the
chemical structure of Figure 1, the black pixels (absence of colors) are zeroed, while the blank spaces
(white pixels as a sum of the red, green and blue colors, valued with 255 each) are represented by 765.
The different coordinates of the pixels composing each image correspond to the chemical variance and
will explain the changes in the bioactivity values. Such a numerical transformation can be performed
using an appropriate program. The following scripts load and convert a bitmap image into a set of numbers, using the Matlab program:
[name,MAP] = imread (‘name.bmp’,’bmp’);
name=double(name);
name=(name(:,:,1)+name(:,:,2)+name(:,:,3));
The n chemical structures of a data set, in which each one is described by a 2D image with x × y
dimension, give a n × x × y 3D array after superposition. This can be regressed against the y block
(the bioactivities column vector) using a multilinear regression method, such as N-PLS. However, it is
common to unfold the 3D array to an X matrix with n × (x × y) dimension and then proceed with the
regression using bilinear partial least squares (traditional PLS). This allows removing columns of the X
matrix with zero variance, i.e. those columns containing pixel data corresponding to common substructures or common blank spaces; the time processing is quite improved, since thousands of descriptors are
usually removed. An example of unfolded binary X matrix, i.e. a descriptors matrix containing 0 (black
pixels forming each chemical structure drawing) and 765 (blank spaces in the drawing = white pixels)
as molecular descriptors, is shown in Figure 3.
Figure 2. Superposition of a congeneric series of compounds illustrating the common substructure and
the variable moieties of thiosemicarbazone derivatives with antitrypanosomal activity
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
89

Evolution of Multivariate Image Analysis in QSAR
Figure 3. Example of superposed workspaces with congeneric structures drawn and systematically
aligned; the bold substructure is similar along with the whole congeneric series of arylsulfonyl compounds as anti HIV-1 compounds (Freitas, 2006). The alignment was performed by selecting the whole
structure and adjusting a pixel in common among all of them (highlighted in red) in a given coordinate
of the workspace. The resulting three-way array was unfolded to the binary X matrix (containing 0
and 765 descriptors as the result of black and white pixels), which can be further regressed against the
bioactivity values using PLS.
Examples of successful MIA-QSAR modeling include the following series of compounds: a) (S)-N[(1-ethyl-2-pyrrolidinyl)methyl]-6-methoxybenzamides as dopamine D
2
= 0.94 and q2 = 0.58 (Freitas, Brown, & Martins, 2005); b) 2-amino-6-arylsulfonylbenzonitriles and
r
their thio and sulfinyl congeners as anti-HIV-1 compounds, with r
(Freitas, 2006); c) 3-phenyl-3-oxo-N-phenylpropanethiamides as antifungal compounds, with r
2
= 0.62 (Freitas & Rittner, 2008); d) sulfonylureas as herbicides, with r2 = 0.92, q2 = 0.62 and r
ad q
receptor subtype inhibitors, with
2
2
= 0.81, q2 = 0.71 and r
2
test
= 0.82
2
= 0.96
2
test
= 0.77 (Bitencourt & Freitas, 2008).
QSAR models should be validated to probe for their reliability in predicting the bioactivities of
potential drug candidates. While the leave-one-out cross-validation procedure is widely used for such
a purpose, external validation has been considered the only way to achieve a reliable QSAR model
(Golbraikh & Tropsha, 2002). Thus, the compound data set should be divided into training and test set
compounds; the latter does not participate in the calibration step, but the activities of the compounds
in this set should be predicted using the multivariate regression parameters. Usually, the statistical pa-
2
rameters required to consider a QSAR model as reliable are based on determination coefficients (r
0.8 for calibration and ≥ 0.5 for external validation and leave-one-out cross-validation (also called q
) ≥
2
).
However, additional criteria have been used to guarantee the reliability and sensitivity of QSAR models
in various scenarios. These are mostly focused on the test set, whose bias concerned to scattering, slope
and intercept of experimental versus predicted bioactivity values should be analyzed (Ojha et al., 2011;
Chirico & Gramatica, 2012). Moreover, an y-randomization test guarantees that an eventual good calibra-
tion result is neither due to chance correlation nor the result of overfitting. Such a test is performed over
the training set and a calibration is carried out using a shuffled y-block and the intact X matrix. A worse
correlation in comparison with the real calibration model is expected if the descriptors really encode
2
the biological properties. To evaluate the statistical difference between r
calculated according to
c
2
r
= r × (r2 - r
P
2
y-rand
1/2
)
; values ≥ 0.5 are acceptable (Mitra, Saha, & Roy, 2010).
and r
2
y-rand
, the cr
2
should be
P
90
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use

Evolution of Multivariate Image Analysis in QSAR
The existing MIA-QSAR models often respect these validation parameters, but more chemically
representative MIA descriptors are expected to improve their quality and interpretability. For instance,
the traditional MIA-QSAR method is based on thin wireframe-like structures; this may cause superposition deviation and imperfection during the 2D alignment. In addition, chemical symbols to indicate
atoms (e.g. Br for bromine and C for carbon) are not chemically interpretable, despite being encoded in
MIA-QSAR models. By the chance, it has been proved that MIA-QSAR really encodes biological data,
since reliable models have been widely obtained, while images of letters of the alphabet do not correlate
with the corresponding sequence number (Cormanich, Nunes, & Freitas, 2012), since this is arbitrary
unlike MIA-QSAR. Even though, new dimensions were introduced in MIA-QSAR to increment quality
and molecular encoding; the aug-MIA-QSAR method emerged.
Augmented MIA-QSAR
The new dimensions introduced in traditional MIA-QSAR to give the aug-MIA-QSAR method are
based on varying colors and spheres with different radii to represent atoms. The molecules can be
drawn with the aid of a program for molecular modeling and can be either optimized or not. Even
though, aug-MIA-QSAR is essentially a 2D method because the images obtained from either 2D or
3D molecular structures are projections in the plane, a workspace with well-defined dimensions in
which a common substructure along with the congeneric series is used for 2D alignment (superposition). The construction of the data matrix is practically the same as for the traditional MIA-QSAR,
but the colors included in the aug-MIA-QSAR approach give rise to a variety of descriptors other than
0 (black pixels) and 765 (white pixels). Thus, while spheres with different sizes can be proportional
to the van der Waals radii of atoms to represent steric effects, the variety of colors obtained from the
sum of red, green and blue components can encode properties that differentiate atoms in the Periodic
Table. For instance, the same series of compounds in Figure 2 was used to generate the structures
to be used in the aug-MIA-QSAR modeling; superposition of these images is given in Figure 4 to
have an idea of the data variance and drawing profile. The chemical structures were built using the
GaussView program (2009) and the colors were kept as the program default.
One of the Periodic properties important in QSAR is the electronegativity, because it is related to
bond polarity and, consequently, dipolar interactions useful to describe the substrate-enzyme interaction. Such information can be indicated by colors in aug-MIA-QSAR, since each atom color can be the
result of the sum of red, green and blue contributions, in such a way that the resulting value (minimum
of zero and maximum of 765) is proportional to the Pauling electronegativity scale. For example, the
electronegativity values for fluorine, chlorine and carbon are 4.0, 3.2 and 2.6, respectively; the corresponding color composition to reflect this trend could be 600 (e.g. 255 red, 255 green, 100 blue)
for fluorine, 480 (e.g. 200 red, 200 green, 80 blue) for chlorine, and 390 (e.g. 190 red, 100 green, 100
blue) for carbon. In this way, the 2D shape, atomic sizes and electronegativity are all encoded in the
aug-MIA-QSAR method, which is illustrated in Figure 5. The comparison among all three image-based
approaches (traditional MIA-QSAR, aug-MIA-QSAR with default atom colors, and aug-MIA-QSAR
with atom colors proportional to electronegativity values) give insight about the limits of encoding
of the traditional MIA-QSAR method, as well as the perspective for the chemical interpretation of
images as molecular descriptors.
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
91

Evolution of Multivariate Image Analysis in QSAR
Figure 4. Superposition of a congeneric series of compounds illustrating the common substructure and
the variable moieties of thiosemicarbazone derivatives with antitrypanosomal activity, according to the
aug-MIA-QSAR concept
Figure 5. Superposition of a congeneric series of compounds illustrating the common substructure and
the variable moieties of thiosemicarbazone derivatives with antitrypanosomal activity, according to the
aug-MIA-QSAR concept using atom colors proportional to the Pauling electronegativity scale
92
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use

Evolution of Multivariate Image Analysis in QSAR
CASE STUDY: TRYPANOSOMIASIS
Despite the expected improvement of augmented MIA descriptors in comparison with traditional MIAQSAR, there are only a few cases comparing the statistical advantages of using aug-MIA-QSAR over
the traditional approach. Thus, this chapter reports a complete comparative analysis of calibration and
validation parameters when using the aforementioned methods. Since QSAR methods are also useful to
predict the bioactivity of non-existing drug candidates, the substructures of existing congeneric compounds were combined to give new drug-like compounds and their biological activities were estimated
using the regression parameters obtained from all three MIA-based methods, in order to give insight
about the numerical difference in pIC
thiosemicarbazones and semicarbazones acting as cruzain inhibitors extracted from the literature (Du et
al., 2002; Lozano et al., 2012), which were divided into 43 training set compounds and 10 test set compounds (Table 1). The test set (ca. 20% of the whole series) was chosen in such a way that compounds
with low, moderate and high activities were homogeneously distributed. The experimental values of
(IC50 is the half maximum inhibitory concentration, in mol L-1) refer to the effectiveness against
pIC
50
the T. cruzi cathepsin L-like protease (cruzain) using purified recombinant protein.
Because a large number of descriptors (generally thousands) is generated in MIA-based matrices,
partial least squares regression is often used in the calibration step, thus requiring the choice of an
optimum number of latent variables (PLS components) to proceed with the construction of the QSAR
model. This can be achieved by analyzing the behavior of the root mean square errors of leave-one-out
cross-validation (RMSECV) with increasing numbers of latent variables. An appropriate number of
latent variables to be used in the model is that in which the RMSECV is minimized or does not vary
significantly. The use of more latent variables than necessary can lead to overfitting, in which unnecessary/undesirable information is calibrated, giving good correlation coefficients in the calibration for the
training set compounds, but not a reliable predictive model for external samples. The traditional MIAQSAR method was applied for the series of compounds of Table 1 (the images corresponding to these
compounds are superposed in Figure 2) and the optimum number of latent variables was 9, according
to the plot of Figure 6.
The model was found to be somewhat overfitted, because the good correlation obtained in the calibra-
2
tion step (r
= 0.967, RMSEC = 0.162) is not significantly different from the results obtained for the PLS
regression between the intact X matrix (the descriptors matrix) and the scrambled y block (the activities
2
column vector), whose r
2
2
and r
r
can be evaluated by the aforementioned cr
y-random
was 0.736 (mean of 10 repetitions). The statistical similarity between
y-random
0.5 (0.473). In addition, despite of giving acceptable correlations in the leave-one-out cross-validation
and external validation (Table 2), the low r
absolute predicted values of pIC
model experiences problems with the slope and/or intercept relative to the optimum line expected for the
experimental vs. predicted pIC
50
in Figure 7. In addition to experimental inaccuracy, statistical limitation, bad image alignment, etc., the
source of such problem can be related to some lack of chemical information captured by the traditional
MIA-QSAR approach. Thus, the QSAR model based on multivariate image analysis can be improved
by introducing other dimensions to account for atomic sizes and other properties. This can be achieved
carrying out aug-MIA-QSAR modeling over the set of compounds with antitrypanosomal activities.
for these new compounds. The series of compounds comprises 53
50
2
parameter, which was found to be inferior to
P
2
value obtained for the test set (0.382) indicates that the
m
does not match well the corresponding experimental data, i.e. the
50
, which should pass through the origin. The correlation plots are given
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
93

Evolution of Multivariate Image Analysis in QSAR
Table 1. Series of thiosemicarbazones and semicarbazones used in the QSAR modeling
Compound R
1
R
2
R
3
X Exp. Pic
1 H H S 7.70
2 H H S 7.70
50
3 H S 7.40
4 H H S 7.30
continued on following page
94
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use

Evolution of Multivariate Image Analysis in QSAR
Table 1. Continued
Compound R
a
5
1
R
2
R
3
X Exp. Pic
H H S 7.30
6 H H S 7.22
7 H S 7.15
50
8 H S 7.10
9 H H S 6.85
continued on following page
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
95
Соседние файлы в папке Библиотека им академика М.И. Перельмана
