Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5865_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
34 Мб
Скачать
Evolution of Multivariate Image Analysis in QSAR
The earlier studies by Cros in 1863 at the University of Strasbourg indicated that toxicity of alcohols to mammals increased as the water solubility of the alcohols decreased (Selassie, 2003). Afterwards, Crum-Brown and Fraser (1869) demonstrated the concept that molecular structure influences the biologic activity of chemical entities and that alteration in structure produces changes in biological action. Richet (1893) proposed that toxicity of some alcohols and ethers was inversely related to the corresponding water solubility and, seven years later, Meyer and Overton independently described a parameter, which can be considered the precursor of the octanol/water partition coefficient (Meyer, 1899; Overton, 1901). An important point along with this trajectory was the introduction of the Hammett constant (σ) to relate the ionization constants of meta and para substituted benzoic acids (Hammett, 1935). This finding opened the exploration of a variety of properties that correlate with biological outcomes, e.g. the relationship between several properties and the toxicity of homolog series of compounds (Ferguson, 1939). In 1964, Free and Wilson (1964) developed a mathematical procedure to encode bioactivities by means of the position of chemical groups in a given molecular scaffold. Simultaneously, Hansch and Fujita (1964) systematized the technique when introduced logP (the logarithm of the octanol/water partition coefficient) to correlate with the biological activity of a series of compounds against a variety of living organisms.
Different QSAR methods appeared since the milestone work by Hansch and Fujita (1964), but this field gained high appearance again years later, when Cramer and coworkers (1988) developed a QSAR method based on the tridimensional structure of a congeneric series of molecules. The treatment of the large amount of data generated was feasible because of the advances in QSAR technologies, improved with computers and programs. According to the so called CoMFA method (Comparative Molecular Field Analysis), each molecule of a given data set is put inside a three-dimensional lattice, whose intersections are probed by steric and electrostatic fields, generating the molecular descriptors. Contour plots representing structural moieties affecting either positively or negatively the bioactivity are given, which are possibly the main outcomes of this method. A similar approach was developed a few years later to account for the molecular similarity indices in a comparative analysis (CoMSIA); the CoMSIA maps highlight those regions within the area occupied by the ligand skeletons that require a particular physicochemical property important for activity. Hydrogen bond has been included as another descriptor in recent CoMSIA models. At least, two steps are required to start building QSAR models based on three-dimensional methods: conformational screening and three-dimensional align­ment. These are not easy tasks, since the energy minimum of an optimized structure does not neces­sarily reflect the bioactive ligand conformation. Indeed, it has been shown for the herbicide 2,4-D that different global minima are achieved using different levels of theory for geometry optimization and conformational search, which are also different from the crystal, bioactive ligand conformation (Freitas & Ramalho, 2013). Consequently, the contour plots indicating regions with positive, neutral and negative influence on the bioactivity may not make sense, because they are dependent on the conformation, which can be either the bioactive one or not. It is worth reinforcing that bioactivity is almost always dependent on conformation and, therefore, it should be useful in QSAR studies. A recent example of this is the role of the synclinal conformation of sulfonylureas to explain the inhibition of acetohydroxyacid synthase (Jana, Delgado, & Medina, 2014); however, the correct conformation is sometimes difficult to be precisely predicted and selected.
Multidimensional (nD) QSAR methods, such as the 4D formalism in QSAR (Hopfinger et al.,
1997), try to minimize the problem of choosing a single conformation for modeling, by performing ensemble averaging, the “fourth dimension”. However, the bioactive conformation is not likely to be
86
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
Evolution of Multivariate Image Analysis in QSAR
represented by an average. Introduction of receptor information in the so-called receptor-dependent 4D-QSAR (RD-4D-QSAR) (Pan, Tseng, & Hopfinger, 2003) and 5D-QSAR (Vedani & Dobler, 2002) appears to better encode structural properties for further correlation with the dependent variables (the bioactivities), but the predictive abilities obtained by many nD models have not been improved significantly in comparison with traditional 3D approaches. Other dimensions have been included in QSAR to refine the molecular representation, such as multiple solvation scenarios in the 6D­QSAR, the real receptor or target-based receptor model data in the 7D-QSAR, and tautomerization (Vedani, Dobler, & Lill, 2005; Polanski, 2009; Martin, 2009). Nevertheless, classical descriptors representing connectivity indices, atom counting, molecular weight, molar refractivity, polarizabil­ity, hydrophobicity, orbital energies, etc. (which are physicochemical descriptors mostly derived from 1D and 2D molecular representation) have not shown to be inferior to 3D descriptors, because the prediction performance is often acceptable and interpretable (Brown & Martin, 1997; Estrada, Molina, & Perdomo-López, 2001).
The predictive ability of models based on these descriptors has been significantly improved with the development of mathematical methods to, for instance, select variables and to account for nonlinearities. Recent examples of feature selection methods in QSAR include ant colony optimi­zation (ACO) and forward stepwise (FS) applied in the modeling of the anti-HIV-1 activities of 3-(3,5-dimethylbenzyl)uracil derivatives (Goodarzi, Freitas, & Jensen, 2009a), as well as genetic algorithm (GA) and successive projections algorithm (SPA) in the analysis of glycogen synthase kinase-3β inhibitory activities of a series of compounds (Goodarzi, Freitas, & Jensen, 2009b). Partial least squares (PLS) regression has been the preferred method to deal with a large amount of independent variables, while multimode methods, like parallel factor (PARAFAC) and multilinear PLS (N-PLS), have gained widespread use to deal with N-way arrays (Freitas et al., 2008). Artifi­cial neural networks (ANN) and support vector machines (SVM) have been successfully employed to account for nonlinearities in QSAR models, improving importantly the predictive accuracy in comparison to multiple linear regression (Goodarzi, Freitas, & Jensen, 2009b).
In line with this, the multivariate image analysis applied to QSAR (MIA-QSAR) method matches the characteristics of 2D QSAR approaches, because the two-dimensional representation of molecules (molecular structure drawings) gives rise to the descriptors used in the modeling. However, 3D information, such as enantiomers, can be encoded by drawing wedge or hashed bonds at the chiral center, such as demonstrated in the MIA-QSPR modeling of the electrophoretic enantioseparation of aromatic amino acids/esters (Goodarzi & Freitas, 2009). Because nothing is expected to be more direct to represent a chemical structure than itself, such a method has provided satisfactory predic­tion performance in the modeling of physical, chemical and biological properties. Nevertheless, the MIA-QSAR is far from perfection, e.g. because of the lack of interpretability for descriptors representing chemical substituents, like atoms represented by their chemical symbols. Thus, the MIA-QSAR method has been constantly improved and its reliability has been attested by rigorous validation protocols, which have been continuously updated. The MIA-QSAR method and its evolu­tion are described in the next sections and, afterwards, an application to a neglected disease is given. Trypanosomiasis was chosen because a few studies have devoted attention to this disease, which is quite common in tropical, third world countries and thousands of deaths occur worldwide, espe­cially in the Latin America, due to the Chagas disease (caused by the parasitic protozoan T. cruzi).
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
87
Evolution of Multivariate Image Analysis in QSAR
MULTIVARIATE IMAGE ANALYSIS IN QSAR
Traditional MIA-QSAR
What is the best way to describe a chemical structure, which will be correlated with its biological activity? Orbital energies, hydrophobicity, electrostatic and steric fields? According to the MIA-QSAR concept, a chemical structure is best represented by itself, as a drawing. How the structure is drawn is the main challenge to achieve better QSAR models. Other challenges include the use of different mathematical techniques for regression and variable selection, as well as the search for ways to analyze the MIA descriptors in order to indicate which pixel (or set of pixels) is responsible to increased or decreased bioactivity and why. This chapter is mainly devoted to present the development of the MIA-QSAR methodology and to show the way in representing chemical structures affects the predictive ability of the MIA-QSAR method.
Chemical structures in the traditional MIA-QSAR look like wireframes, with lines representing chemical bonds and chemical symbols to indicate atoms when necessary. Eventually, chiral centers can be designated as wedge or hashed bonds. These chemical structures can be easily drawn as routinely performed using a variety of programs, such as ChemDraw and ChemSketch. An example of a chemical structure drawn with the ChemSketch program is given in Figure 1.
QSAR methods try to search for correlations between an homolog set of compounds and the correspond­ing bioactivities. Thus, the set of compounds retain a structural similarity (usually the pharmacophoric group), which should be used in the 2D alignment. Such an alignment is fundamental to make variable only those moieties not common along with the set of compounds, which are responsible for the variance in the bioactivities of the molecules. The variable moieties are evident after superposition of all image files, as in the example of Figure 2. In order to superpose the files for the 2D alignment, each chemical structure should be drawn systematically, in such a way that common substructures do not vary, while substituents or variable moieties should be drawn in the same way every time it is repeated in different chemical structures. Each chemical structure should be pasted in a workspace with x × y dimension (e.g. 500 × 500 pixels size) using, for instance, the Paint application of the Microsoft Windows. Before saving the images, each one should be completely selected and a given pixel, which is common along with the set of structures, should be used to manually move (preferably using the mouse) the whole structure to a given position in the workspace. The 2D alignment is done and the image can be saved e.g. as bitmaps (.bmp). This procedure is pictographically represented in Figure 3. The file extension and the way in which
Figure 1. The analgesic S-ketamine drawn with the ChemSketch program
88
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
Evolution of Multivariate Image Analysis in QSAR
substituents are represented (e.g. CH3 or Me to indicate a methyl group) have not shown to statistically affect the model response (Goodarzi, Freitas, & Ferreira, 2009). The drawings should be readable for further regression against the bioactivity values. Because the images are composed by pixels, these can be converted to numbers, according to the RGB (red-green-blue) system of color composition. For the chemical structure of Figure 1, the black pixels (absence of colors) are zeroed, while the blank spaces (white pixels as a sum of the red, green and blue colors, valued with 255 each) are represented by 765. The different coordinates of the pixels composing each image correspond to the chemical variance and will explain the changes in the bioactivity values. Such a numerical transformation can be performed using an appropriate program. The following scripts load and convert a bitmap image into a set of num­bers, using the Matlab program:
[name,MAP] = imread (‘name.bmp’,’bmp’);
name=double(name);
name=(name(:,:,1)+name(:,:,2)+name(:,:,3));
The n chemical structures of a data set, in which each one is described by a 2D image with x × y
dimension, give a n × x × y 3D array after superposition. This can be regressed against the y block (the bioactivities column vector) using a multilinear regression method, such as N-PLS. However, it is common to unfold the 3D array to an X matrix with n × (x × y) dimension and then proceed with the regression using bilinear partial least squares (traditional PLS). This allows removing columns of the X matrix with zero variance, i.e. those columns containing pixel data corresponding to common substruc­tures or common blank spaces; the time processing is quite improved, since thousands of descriptors are usually removed. An example of unfolded binary X matrix, i.e. a descriptors matrix containing 0 (black pixels forming each chemical structure drawing) and 765 (blank spaces in the drawing = white pixels) as molecular descriptors, is shown in Figure 3.
Figure 2. Superposition of a congeneric series of compounds illustrating the common substructure and the variable moieties of thiosemicarbazone derivatives with antitrypanosomal activity
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
89
Evolution of Multivariate Image Analysis in QSAR
Figure 3. Example of superposed workspaces with congeneric structures drawn and systematically aligned; the bold substructure is similar along with the whole congeneric series of arylsulfonyl com­pounds as anti HIV-1 compounds (Freitas, 2006). The alignment was performed by selecting the whole structure and adjusting a pixel in common among all of them (highlighted in red) in a given coordinate of the workspace. The resulting three-way array was unfolded to the binary X matrix (containing 0 and 765 descriptors as the result of black and white pixels), which can be further regressed against the bioactivity values using PLS.
Examples of successful MIA-QSAR modeling include the following series of compounds: a) (S)-N­[(1-ethyl-2-pyrrolidinyl)methyl]-6-methoxybenzamides as dopamine D
2
= 0.94 and q2 = 0.58 (Freitas, Brown, & Martins, 2005); b) 2-amino-6-arylsulfonylbenzonitriles and
r
their thio and sulfinyl congeners as anti-HIV-1 compounds, with r (Freitas, 2006); c) 3-phenyl-3-oxo-N-phenylpropanethiamides as antifungal compounds, with r
2
= 0.62 (Freitas & Rittner, 2008); d) sulfonylureas as herbicides, with r2 = 0.92, q2 = 0.62 and r
ad q
receptor subtype inhibitors, with
2
2
= 0.81, q2 = 0.71 and r
2
test
= 0.82
2
= 0.96
2
test
= 0.77 (Bitencourt & Freitas, 2008).
QSAR models should be validated to probe for their reliability in predicting the bioactivities of potential drug candidates. While the leave-one-out cross-validation procedure is widely used for such a purpose, external validation has been considered the only way to achieve a reliable QSAR model (Golbraikh & Tropsha, 2002). Thus, the compound data set should be divided into training and test set compounds; the latter does not participate in the calibration step, but the activities of the compounds in this set should be predicted using the multivariate regression parameters. Usually, the statistical pa-
2
rameters required to consider a QSAR model as reliable are based on determination coefficients (r
0.8 for calibration and ≥ 0.5 for external validation and leave-one-out cross-validation (also called q
) ≥
2
). However, additional criteria have been used to guarantee the reliability and sensitivity of QSAR models in various scenarios. These are mostly focused on the test set, whose bias concerned to scattering, slope and intercept of experimental versus predicted bioactivity values should be analyzed (Ojha et al., 2011; Chirico & Gramatica, 2012). Moreover, an y-randomization test guarantees that an eventual good calibra- tion result is neither due to chance correlation nor the result of overfitting. Such a test is performed over the training set and a calibration is carried out using a shuffled y-block and the intact X matrix. A worse correlation in comparison with the real calibration model is expected if the descriptors really encode
2
the biological properties. To evaluate the statistical difference between r calculated according to
c
2
r
= r × (r2 - r
P
2
y-rand
1/2
)
; values ≥ 0.5 are acceptable (Mitra, Saha, & Roy, 2010).
and r
2
y-rand
, the cr
2
should be
P
90
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
Evolution of Multivariate Image Analysis in QSAR
The existing MIA-QSAR models often respect these validation parameters, but more chemically representative MIA descriptors are expected to improve their quality and interpretability. For instance, the traditional MIA-QSAR method is based on thin wireframe-like structures; this may cause superpo­sition deviation and imperfection during the 2D alignment. In addition, chemical symbols to indicate atoms (e.g. Br for bromine and C for carbon) are not chemically interpretable, despite being encoded in MIA-QSAR models. By the chance, it has been proved that MIA-QSAR really encodes biological data, since reliable models have been widely obtained, while images of letters of the alphabet do not correlate with the corresponding sequence number (Cormanich, Nunes, & Freitas, 2012), since this is arbitrary unlike MIA-QSAR. Even though, new dimensions were introduced in MIA-QSAR to increment quality and molecular encoding; the aug-MIA-QSAR method emerged.
Augmented MIA-QSAR
The new dimensions introduced in traditional MIA-QSAR to give the aug-MIA-QSAR method are based on varying colors and spheres with different radii to represent atoms. The molecules can be drawn with the aid of a program for molecular modeling and can be either optimized or not. Even though, aug-MIA-QSAR is essentially a 2D method because the images obtained from either 2D or 3D molecular structures are projections in the plane, a workspace with well-defined dimensions in which a common substructure along with the congeneric series is used for 2D alignment (superposi­tion). The construction of the data matrix is practically the same as for the traditional MIA-QSAR, but the colors included in the aug-MIA-QSAR approach give rise to a variety of descriptors other than 0 (black pixels) and 765 (white pixels). Thus, while spheres with different sizes can be proportional to the van der Waals radii of atoms to represent steric effects, the variety of colors obtained from the sum of red, green and blue components can encode properties that differentiate atoms in the Periodic Table. For instance, the same series of compounds in Figure 2 was used to generate the structures to be used in the aug-MIA-QSAR modeling; superposition of these images is given in Figure 4 to have an idea of the data variance and drawing profile. The chemical structures were built using the GaussView program (2009) and the colors were kept as the program default.
One of the Periodic properties important in QSAR is the electronegativity, because it is related to bond polarity and, consequently, dipolar interactions useful to describe the substrate-enzyme interac­tion. Such information can be indicated by colors in aug-MIA-QSAR, since each atom color can be the result of the sum of red, green and blue contributions, in such a way that the resulting value (minimum of zero and maximum of 765) is proportional to the Pauling electronegativity scale. For example, the electronegativity values for fluorine, chlorine and carbon are 4.0, 3.2 and 2.6, respectively; the cor­responding color composition to reflect this trend could be 600 (e.g. 255 red, 255 green, 100 blue) for fluorine, 480 (e.g. 200 red, 200 green, 80 blue) for chlorine, and 390 (e.g. 190 red, 100 green, 100 blue) for carbon. In this way, the 2D shape, atomic sizes and electronegativity are all encoded in the aug-MIA-QSAR method, which is illustrated in Figure 5. The comparison among all three image-based approaches (traditional MIA-QSAR, aug-MIA-QSAR with default atom colors, and aug-MIA-QSAR with atom colors proportional to electronegativity values) give insight about the limits of encoding of the traditional MIA-QSAR method, as well as the perspective for the chemical interpretation of images as molecular descriptors.
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
91
Evolution of Multivariate Image Analysis in QSAR
Figure 4. Superposition of a congeneric series of compounds illustrating the common substructure and the variable moieties of thiosemicarbazone derivatives with antitrypanosomal activity, according to the aug-MIA-QSAR concept
Figure 5. Superposition of a congeneric series of compounds illustrating the common substructure and the variable moieties of thiosemicarbazone derivatives with antitrypanosomal activity, according to the aug-MIA-QSAR concept using atom colors proportional to the Pauling electronegativity scale
92
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
Evolution of Multivariate Image Analysis in QSAR
CASE STUDY: TRYPANOSOMIASIS
Despite the expected improvement of augmented MIA descriptors in comparison with traditional MIA­QSAR, there are only a few cases comparing the statistical advantages of using aug-MIA-QSAR over the traditional approach. Thus, this chapter reports a complete comparative analysis of calibration and validation parameters when using the aforementioned methods. Since QSAR methods are also useful to predict the bioactivity of non-existing drug candidates, the substructures of existing congeneric com­pounds were combined to give new drug-like compounds and their biological activities were estimated using the regression parameters obtained from all three MIA-based methods, in order to give insight about the numerical difference in pIC thiosemicarbazones and semicarbazones acting as cruzain inhibitors extracted from the literature (Du et al., 2002; Lozano et al., 2012), which were divided into 43 training set compounds and 10 test set com­pounds (Table 1). The test set (ca. 20% of the whole series) was chosen in such a way that compounds with low, moderate and high activities were homogeneously distributed. The experimental values of
(IC50 is the half maximum inhibitory concentration, in mol L-1) refer to the effectiveness against
pIC
50
the T. cruzi cathepsin L-like protease (cruzain) using purified recombinant protein.
Because a large number of descriptors (generally thousands) is generated in MIA-based matrices, partial least squares regression is often used in the calibration step, thus requiring the choice of an optimum number of latent variables (PLS components) to proceed with the construction of the QSAR model. This can be achieved by analyzing the behavior of the root mean square errors of leave-one-out cross-validation (RMSECV) with increasing numbers of latent variables. An appropriate number of latent variables to be used in the model is that in which the RMSECV is minimized or does not vary significantly. The use of more latent variables than necessary can lead to overfitting, in which unneces­sary/undesirable information is calibrated, giving good correlation coefficients in the calibration for the training set compounds, but not a reliable predictive model for external samples. The traditional MIA­QSAR method was applied for the series of compounds of Table 1 (the images corresponding to these compounds are superposed in Figure 2) and the optimum number of latent variables was 9, according to the plot of Figure 6.
The model was found to be somewhat overfitted, because the good correlation obtained in the calibra-
2
tion step (r
= 0.967, RMSEC = 0.162) is not significantly different from the results obtained for the PLS
regression between the intact X matrix (the descriptors matrix) and the scrambled y block (the activities
2
column vector), whose r
2
2
and r
r
can be evaluated by the aforementioned cr
y-random
was 0.736 (mean of 10 repetitions). The statistical similarity between
y-random
0.5 (0.473). In addition, despite of giving acceptable correlations in the leave-one-out cross-validation and external validation (Table 2), the low r absolute predicted values of pIC model experiences problems with the slope and/or intercept relative to the optimum line expected for the experimental vs. predicted pIC
50
in Figure 7. In addition to experimental inaccuracy, statistical limitation, bad image alignment, etc., the source of such problem can be related to some lack of chemical information captured by the traditional MIA-QSAR approach. Thus, the QSAR model based on multivariate image analysis can be improved by introducing other dimensions to account for atomic sizes and other properties. This can be achieved carrying out aug-MIA-QSAR modeling over the set of compounds with antitrypanosomal activities.
for these new compounds. The series of compounds comprises 53
50
2
parameter, which was found to be inferior to
P
2
value obtained for the test set (0.382) indicates that the
m
does not match well the corresponding experimental data, i.e. the
50
, which should pass through the origin. The correlation plots are given
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
93
Evolution of Multivariate Image Analysis in QSAR
Table 1. Series of thiosemicarbazones and semicarbazones used in the QSAR modeling
Compound R
1
R
2
R
3
X Exp. Pic
1 H H S 7.70
2 H H S 7.70
50
3 H S 7.40
4 H H S 7.30
continued on following page
94
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
Evolution of Multivariate Image Analysis in QSAR
Table 1. Continued
Compound R
a
5
1
R
2
R
3
X Exp. Pic
H H S 7.30
6 H H S 7.22
7 H S 7.15
50
8 H S 7.10
9 H H S 6.85
continued on following page
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
95