Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5345_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
34 Мб
Скачать
Evolution of Multivariate Image Analysis in QSAR
The aug-MIA-QSAR method was applied to model the bioactivities of the same class of compounds evaluated using traditional MIA-QSAR, in order to get insight about the statistical advantages of using this augmented methodology. The chemical structures were built using the GaussView program, in which default colors for each atom are those represented in Figure 4 for the superposed images (gray for carbon, white for hydrogen, light blue for fluorine and so on). These colors are arbitrary and do not have any correspondence with atomic properties, i.e. the RGB additive color model gives 426 for carbon (resulting in the gray color as a composition of 142 red, 142 green and 142 blue), 612 for hydrogen (resulting in the white color as a composi­tion of 204 red, 204 green and 204 blue), 688 for fluorine (resulting in the light blue color as a composition of 178 red, 255 green and 255 blue), and so on. Thus, different colors for different atoms serve to distinguish the spheres with different sizes in the images, but they do not have chemical meaning. On the other hand, spheres with sizes proportional to the van der Waals radii describe the atomic volume and corresponding properties, such as steric effect. The effect of these changes over the traditional MIA-QSAR model was investigated.
The same procedure described to build the traditional MIA-QSAR model was applied to construct the aug-MIA-QSAR model. Again, 9 PLS components were chosen to be used in the modeling; the RMSECV appears to stabilize from 7 latent variables (Figure 8), but 9 were used to reproduce the conditions applied in the traditional MIA-QSAR model. The calibration gave a high correlation between experimental and fitted pIC
values, while the results did not show to be overfitted, since the cr
50
0.5 (Table 2). In addition, the leave-one-out cross-validation was improved, with q
2
value was superior to
P
2
of 0.763, which is quite satisfactory. However, external validation has been claimed as the only way to establish a reliable QSAR model (Golbraikh & Tropsha, 2002) and, despite the good correlation between experimental and predicted pIC
values for this set of compounds (Figure 9), it did not pass the r
50
2
test, whose value was
m
marginally inferior to 0.5. Thus, the aug-MIA-QSAR model also requires improvement to certify that predictions for new drug candidates are really reliable.
Figure 8. Plot of number of latent variables vs. RMSECV for the aug-MIA-QSAR model
106
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
Evolution of Multivariate Image Analysis in QSAR
Figure 9. Plot of experimental vs. fitted and predicted pIC
for the series of thiosemicarbazones and
50
semicarbazones, according to the aug-MIA-QSAR method (default colors)
A new dimension can be introduced in aug-MIA-QSAR if colors in the chemical images are used to encode atomic properties. In addition to steric effects, other important interactions ruling enzyme­substrate interactions, responsible for the activity of drug molecules, are the dipolar interactions. Only atoms involved in polar bonds are capable of interacting through dipolar interactions. A polar bond is formed by atoms with different electronegativities and, therefore, the Pauling electronegativity can be used as a descriptor in the so called aug-MIA-QSAR
method. According to this approach, different
color
atoms are colored according to the respective electronegativity values. Because colors can be numerically managed using the RGB additive color model, the atom colors can be proportional to the corresponding electronegativity values, e.g. in Figure 5, where the fluorine with electronegativity 4.0 is approximately white (760 as a sum of 255 red, 255 green and 250 blue), the carbon with electronegativity 2.5 is ap­proximately yellow (475 as a sum of 255 red, 225 green and 0 blue), the hydrogen with electronegativity
2.1 is approximately gray (399 as a sum of 133 red, 133 green and 133 blue), and so on.
Similarly to the previous aug-MIA-QSAR, the aug-MIA-QSAR
could be more parsimonious than
color
the traditional MIA-QSAR model, since the RMSECV decayed on going from 1 to 5 latent variables, from which it is stabilized (Figure 10). However, 9 PLS components were used in the aug-MIA-QSAR
color
modeling for comparison with the other models. Here, the calibration results continued statistically significant, with low RMSEC, high r
2
and not overfitted (cr
2
> 0.5). Nevertheless, the main advantages
P
of this new model is related to its validation, because the leave-one-one cross-validation results suggest
2
a high predictability (q
= 0.756), which is confirmed by the external validation, in which r
proved to 0.714, while the absolute experimental and predicted pIC
values matches very well to each
50
2
was im-
test
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
107
Evolution of Multivariate Image Analysis in QSAR
Figure 10. Plot of number of latent variables vs. RMSECV for the aug-MIA-QSAR model (colors pro­portional to the atom electronegativities)
other (r
2
> 0.5). The correlations are illustrated in Figure 11. Now, the aug-MIA-QSAR
m
model for
color
the series of thiosemicarbazones and semicarbazones as antitrypanosomal compounds satisfies the most important validation parameters required to attest the reliability of a QSAR model, suggesting that the chemical properties correlating with biological activities are suitably encoded by the aug-MIA
descrip-
color
tors and that dipolar interactions play a key role for the action mechanism of this series of compounds. Consequently, the aug-MIA-QSAR
model is ready to estimate the bioactivities of new drug candidates.
color
A QSAR analysis based on physicochemical descriptors for the same series of compounds has been performed, but using a different set of test compounds. The best PLS model using 8 selected descrip­tors gave r
2
= 0.85, q2 = 0.78 and r
2
= 0.55 (Lozano et al., 2012). Thus, the MIA-based models are
test
consistent with the classical QSAR using physicochemical descriptors, despite a deeper comparison cannot be done because the lack of additional validation in the later.
A useful strategy to obtain new compounds containing a given scaffold is to combine the molecular substructures of highly active compounds in a congeneric series. This can be achieved using the data set of Table 1. This procedure can result in highly active substances with improved properties, such as lower toxicity and high activity toward strains that are resistant against the available chemotherapy. For example, compounds 1 and 2 of Table 1 are the most active ones in that series, but they differ from each other by the substituents at the benzene ring in R drug-like compound can be designed with R
or 3-CF3,4-Cl or 3,4-Cl2,5-CF3 or 3,5-(CF3)2,4-Cl; all these possibilities have the Cl and CF3 groups
CF
3
containing the benzene ring with 3-Cl,5-CF3 or 4-Cl,5-
1
(3,4-Cl2 in 1 and 3,5-(CF3)2 in 2). Thus, a new
1
calibrated in that ring positions. Accordingly, these compounds (A-E, Figure 12) were proposed and their bioactivities were subsequently estimated using the regression parameters of the MIA-QSAR and aug-MIA-QSAR models.
108
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
Evolution of Multivariate Image Analysis in QSAR
Figure 11. Plot of experimental vs. fitted and predicted pIC semicarbazones, according to the aug-MIA-QSAR
method (colors proportional to the atom electro-
color
negativities)
for the series of thiosemicarbazones and
50
According to the traditional MIA-QSAR model, compounds A and C were the most active pur­poses (Table 3), but not more active than the experimentally available compounds 1 and 2. However, the most predictive aug-MIA-QSAR models indicated that C and D are more active than 1 and 2. Consequently, these proposed compounds can drive the synthesis of new antitrypanosomal drugs. In addition, these outcomes allow the chemical interpretation about structural requirements to achieve successful purposes: thiosemicarbazones containing a tri-substituted phenyl ring in the R
group are
1
promising, while chlorine at position 3 is preferred rather than the trifluoromethyl group.
These results can be further validated by comparing the predicted pIC
values with the docking
50
scores obtained from receptor-based methods. Molecular docking predicts the preferred orientation of a molecule (substrate) to another (enzyme) when bound to each other to form a stable complex. This binding affinity is usually described in terms of scoring functions, which are expected to be correlated with the bioactivity values. Thus, structure- and receptor-based approaches are comple­mentary. A comparison between docking scores and the bioactivity data obtained through several MIA-QSAR models (Antunes et al., 2008; Goodarzi et al., 2010; Deeb et al., 2012; Guimarães et al.,
2014) has shown that MIA descriptors indeed carry biochemical information. Such a comparison has not performed hitherto using aug-MIA descriptors, but the expected response is optimist, since the predictive performance of aug-MIA-QSAR is better than in traditional MIA-QSAR, at least for the case study of this chapter.
109
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
Evolution of Multivariate Image Analysis in QSAR
Figure 12. Proposed thiosemicarbazones obtained from the combination of the substructures of com­pounds 1 and 2
Table 3. Predicted pIC50 for the proposed thiosemicarbazones A-E
Compound Traditional MIA-QSAR Aug-MIA-QSAR Aug-MIA-QSAR A 7.37 7.55 7.48 B 6.73 7.38 7.44 C 7.60 7.98 8.01 D 7.07 7.76 7.76 E 6.95 7.30 7.31
110
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
color
Evolution of Multivariate Image Analysis in QSAR
Additional drug-like profile can be estimated by numerous programs available for this purpose, which have also complemented many pharmacodynamic data obtained by MIA-QSAR (Silva et al., 2012; Gui­marães et al., 2014). The estimation of pharmacokinetic data, namely ADMET (absorption, distribution, metabolism, excretion and toxicity) can be directly and indirectly calculated. The most widely known indirect method to obtain some information about drug-likeness of a given drug candidate is to use the Lipinski’s rule of five (Lipinski et al., 1997). According to this rule, most drug-like molecules have logP ≤ 5, molecular weight ≤ 500, number of hydrogen bond acceptors ≤ 10, and number of hydrogen bond donors ≤ 5. Molecules violating more than one of these rules might have problems with bioavail­ability. The rule is called “rule of 5”, because the border values are 5, 500, 2×5, and 5. For example, while hydrophobicity (expressed in terms of logP) affects drug absorption, bioavailability, hydrophobic drug-receptor interactions, metabolism of molecules, as well as their toxicity, the molecular volume determines transport characteristics of molecules. Obviously, this is an approach, since many current therapies are based on drug molecules violating more than one of these rules, such as Ritonavir, which is used for the treatment of AIDS, but violates 3 rules according to calculations using the Molinspiration
-1
program (www.molinspiration.com), i.e. logP > 5, molecular weight of 721 g mol
and 11 hydrogen bond acceptors. Other programs estimate directly some pharmacokinetic parameters, such as human intestinal absorption and the blood-brain barrier penetration. These calculations are mostly based on models comprising big libraries of compounds having such parameters available.
In fact, the combination of a QSAR technique (to predict the bioactivity of a designed compound) with docking studies (to understand the substrate-enzyme interaction mechanism) and ADMET evalua­tion, is a powerful way to achieve a successful drug candidate in a relatively low cost manner and more quickly than performing experimental compound screening and tests. Thus, multivariate image analysis methods can play an important role on this course.
FUTURE RESEARCH DIRECTIONS
The MIA-based QSAR methods are relatively new and, consequently, many performance tests and im­provements are still to be performed, while the scientific community and the whole society demand more and better medicines to treat increasingly numerous and damaging diseases. This fact makes necessary strategies in medicinal chemistry to provide suitable chemotherapies in a fast, reliable and safe manner. In order to satisfy these requirements, new computational methods are always welcome to be developed, as are mathematical tools to capture the maximum information possible from molecular descriptors and additional validation techniques to guarantee that QSAR models are really predictive and accurate. Thus, future research directions related to MIA-QSAR are concerned to improve the biochemical interpretability of MIA descriptors, to mathematically manage descriptors in order to capture the MIA information not explored using conventional tools and to apply more precautionary approaches for validation purposes.
There is a well-known lack of interpretability in traditional MIA-QSAR. Partially, this can be attributed to the binary-like descriptors and the big three-way data array obtained after images superposition. This means that the three-way array should be unfolded to a two-way array (a matrix) in order to be analyzed by bilinear multivariate methods, such as PLS (for regression) and PCA (for pattern recognition). The unfolding makes difficult to identify which descriptor corresponds to which pixel coordinate in the 2D space forming each image. In addition, in order to make the data management more practicable, in terms of computational cost, columns with zero variance are usually removed; this procedure makes calculation
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
111
Evolution of Multivariate Image Analysis in QSAR
faster, but the information about which number (765 for white pixels and 0 for black pixels) in the data matrix corresponds to which pixel location in the image gets completely lost. In addition, back-folding becomes impossible. This drawback can be at least partially overcome using aug-MIA-QSAR, because the descriptor number corresponding to a given pixel in a chemical image can be identified (since each atom has a different color) and its contribution for the model can be investigated using loading analy­sis. In aug-MIA-QSAR
, this information is expected to be still more chemically meaningful. This
color
is inspired in the 3D contour maps obtained from CoMFA analysis, where different colors are given to surfaces indicating positive, negative or neutral effect of certain descriptors (e.g. steric and electrostatic) on the bioactivity values. Color is also used to code the direction and magnitude of these differential interactions and, in CoMFA, the numerical data used to construct these “coefficient contour” maps, the QSAR coefficients and the data table, are available on request from the authors (Cramer, Patterson, & Bunce, 1988).
Even the more chemically meaningful aug-MIA-QSAR
method (compared to traditional MIA-
color
QSAR and aug-MIA-QSAR) may not offer descriptors linearly correlated with the biactivity values, or then some information may not have any significance at all; in the worst-case scenario, biased descriptors (e.g. from spurious data, bad alignment, etc.) could confuse the model. Thus, variable selection and dif­ferent regression methods can be used to overcome these setbacks. Genetic algorithm is one of the most widely applied methods for feature selection. It has widespread use for dimension reduction and have been reported in the optimization of a number of different and traditionally difficult problems, including image processing, design of complex networks (e.g., computers and integrated circuits), classifications, parameters for neural nets, job scheduling, robotics, and parameter fitting (Davis, 1991). Genetic algorithm (GA) is a search that mimics the process of natural selection and it is based on the Darwin’s evolutionary theory; a population of candidate solutions (compounds) to an optimization problem is evolved toward better solutions. Each candidate solution has a set of properties (descriptors) which can be mutated and altered. The evolution usually starts from a population of randomly generated compounds, and is an itera­tive process, with the population in each iteration called a generation. In each generation, the fitness of every compound in the population is evaluated; the fitness is usually the value of the objective function in the optimization problem being solved. The more fit individuals are stochastically selected from the current population, and each individual’s genome is modified (recombined and possibly randomly mutated) to form a new generation. The new generation of candidate solutions is then used in the next iteration of the algorithm. Commonly, the algorithm terminates when either a maximum number of generations has been produced, or a satisfactory fitness level has been reached for the population (Whitley, 1994). A variety of other feature selection methods is available, from those based on ordered predictors selection (OPS), which consists of obtaining an informative value, then decreasing the ordering of variables by this value, and lastly investigating the ordered variables (Teófilo, Martins, & Ferreira, 2009), to those inspired in ant colony optimization (Jensen, 2006). The use of such methods can be particularly useful if applied to aug-MIA descriptors, since the selected, colored descriptors could be identified and then analyzed on the basis of their contribution for the activity values.
Bilinear partial least squares (PLS) is the main regression method used in multivariate QSAR. In
MIA-QSAR, unfolding is done so that pixels become a single row and, thus, an image that is originally I by J pixels for K compounds is reshaped to form a two-way array that is I×J by K. An X matrix is then built, where each row contains the variables (the pixels) describing each molecule, and is subsequently decomposed into a score vector s and a weight vector w. The score vector is determined to have the prop- erty of maximum covariance with the dependent variable y. The score vectors then replace the original
112
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
Evolution of Multivariate Image Analysis in QSAR
variables as regressors. However, the three-way array can be directly regressed against the y column vector using multi-way regression methods, such as multilinear partial least squares (N-PLS) (Bro, 1996) and parallel factor (PARAFAC) analysis (Bro, 1997). Actually, these methods have not shown much improved results if compared to PLS when applied to traditional MIA descriptors (Freitas et al., 2008). On the other hand, regression methods that account for nonlinearities, particularly artificial neural net­works (ANN) and support-vector machines (SVM), have really enhanced the predictive performance of physicochemical-based QSAR models when compared to multiple linear regression (Goodarzi, Freitas, & Jensen, 2009). Least-squares support-vector machines (LS-SVM) combined with feature selection methods have been applied only in a few cases to MIA descriptors (Cormanich, Goodarzi, & Freitas,
2009); thus, research on this field is a promising perspective to achieve better (aug)MIA-based models where the relationship between structure and activity is not completely linear.
A QSAR model should be reliable enough to allow the prediction of external samples in order to lead the synthesis and development of new drug candidates. Such a confidence can be assessed using the maximum number possible of validation protocols. There are several examples of traditional MIA-QSAR models in the literature submitted to validation tests, such as leave-one-out cross-validation, external validation, y-randomization tests, etc. The results have shown to be satisfactory. The leave-one-out cross-validation (LOOCV) is the most widely used procedure for validation purposes; LOOCV involves using a single observation from the original set of compounds as the validation data, and the remaining observations as the training data. This is repeated such that each observation in the set of compounds is used once as the validation data. Alternatively, more than a single observation can be used in each observation, in the so called leave-many-out cross-validation. The main result is the determination coef-
2
ficient q
obtained from the correlation between experimental versus predicted data. However, the use of LOOCV alone for validation purposes has been criticized (Golbraikh & Tropsha, 2002), while external validation has been claimed as the only way to establish a reliable QSAR model.
An external validation test is performed to a series of compounds different from those pertaining to the training set, which is used for calibration purposes. The regression parameters obtained from the calibration process are used to predict the dependent variables of the external set of compounds. The main
2
result is the determination coefficient r
2
predicted data. In general, r
above 0.5 are considered acceptable, but it is not a sufficient condition
test
obtained from the correlation between experimental versus
test
to attest the reliability of a QSAR model. Therefore, additional statistical parameters applied over the external sample of compounds have been developed to prevent misinterpretation of an eventual high determination coefficient of external validation. There are different sources of problems in the external validation that can lead to equivocal consideration about the suitability of a QSAR model, namely data scattering and deviation of the intercept and slope of the line obtained from the linear regression of actual versus predicted data, in comparison with an ideal scenario (Figure 13).
Highly scattered data could result in a regression line almost coincident with the one obtained from
2
an ideal scenario, but the problem can be captured by the low r
2
errors (RMSE). However, there are cases in which r
is high if the regression line is not forced to
test
value and high root mean square
test
pass through the origin, but the QSAR model is not really predictive. For instance, external validation data in which the predicted values present a systematic error compared with the actual ones will have
2
problem with the line intercept. In another case, a correlation giving high r
value could be obtained
test
from a relationship in which the regression line passes through the origin, but this will have problem with the slope of the curve if the absolute values of actual and predicted data are not closely related. These and other problems can be analyzed using a variety of statistical parameters, like the modified r
2
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
113
Evolution of Multivariate Image Analysis in QSAR
Figure 13. Examples of possible problems in external validation compared with an ideal scenario
2
), Δr
m
2
and corrected-penalized r2 (cr
m
(r al., 2013; Ojha et al., 2011; Mitra, Saha, & Roy, 2010). While r
2
) of Equations 1-3 (Roy et al., 2009; Roy et al., 2012; Roy et
P
2
> 0.5 guarantees that not only a good
m
correlation between the experimental and predicted values is obtained for the test set, but also that the absolute experimental and predicted pEC
2
the statistical difference between r
2
r
= r2 [1 - (r2 - r
m
wherein r
2
and r
1/2
2
)
] (1)
0
2
correspond to the squared correlation coefficient values between measured and pre-
0
and r
values are congruent, the cr
50
2
(values above 0.5 are considered acceptable).
rand
2
parameter gives insight about
P
dicted values for the test set with and without intercept, respectively. Change of the axes gives the value
2
2
, which is used in Equation 2 to obtain Δr
of r
m
2
Δr
= |r
m
cr2
= r (r2 - r
P
2
m
2
- r
| (2)
m
2
1/2
)
(3)
rand
m
2
r
< 0.2 is acceptable).
m
k and k’ indicate the slopes of the observed versus predicted data (Equations 4 and 5) and are useful
for further validation. Ideally, both k and k’ should be ≥ 0.88 and ≤ 1.15.
k = Σ(y
obs
× y
pred
)/Σ(y
)2 (4)
pred
k’ = Σ(y
obs
pred
)/Σ(y
)2 (5)
obs
× y
114
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
Evolution of Multivariate Image Analysis in QSAR
More precautionary criteria have been recently established by Chirico and Gramatica (2012), based on statistical parameters obtained from the calibration, cross-validation and external validation sets. Overall, the authors suggest that, in addition to the most commonly used validation criteria to different kinds of bias in data, the concordance correlation coefficient (CCC) is an important parameter to verify how small the differences are between experimental data and external data set predictions, independently of their range. Because aug-MIA-QSAR is a relatively new method, further research to probe for its adequacy to these trends in validation is a perspective in this field.
CONCLUSION
Multivariate image analysis applied in quantitative structure-activity relationship studies emerged as a necessity to design new drug candidates in a reliable and easy way. This can be demonstrated by a variety of examples in the literature using traditional MIA-QSAR in the prediction of bioactivities of proposed compounds. This chapter described the procedure used to build a MIA-QSAR, as well as its evolution from 2005 to the current augmented versions (aug-MIA-QSAR), which encode chemical information better than the traditional MIA-QSAR. The molecular representation was improved by introducing spheres with different sizes and colors corresponding to the atoms in molecules. While the sphere sizes are expected to account for biological effects depending on the atomic (or group of atoms) volume, like steric effects, the colors could encode electrostatic interactions, especially if sphere colors are numerically proportional (according to the RGB color model) to the electronegativity Pauling scale. This improve­ment was checked for a series of thiosemicarbazones and semicarbazones as antitrypanosomal agents; the so called aug-MIA-QSAR based models. However, despite promising, such method is quite recent and then lacks additional tests (development of models for several classes of compounds and comparisons), improvements concerned to data analysis and interpretability, and application of state-of-the-art validation protocols. Traditional MIA-QSAR is recommended rather than the augmented approaches when the chemical sampling are not trivial to be constructed (drawn) using the available programs (e.g. very complex tridimensional geometries), such as GaussView, or when atomic sizes and electronegativity are not expected to play a determinant role to improve the model quality, e.g. for hydrocarbon chains only as substituents in a con­generic series. Similarly, aug-MIA-QSAR if electronegativity is indeed an effective descriptor and if the series contains structural dependence with atomic electronegativity. For instance, a series of compounds containing chlorine and nitrogen in differ­ent positions of an aromatic ring as substituents does not need to be modeled with aug-MIA-QSAR since both atoms have Pauling electronegativity 3.0.
model was found to be more predictive and reliable than other MIA-
color
is expected to give better outcomes than aug-MIA-QSAR
color
color
,
REFERENCES
Antunes, J. E., Freitas, M. P., da Cunha, E. F. F., Ramalho, T. C., & Rittner, R. (2008). In silico prediction of novel phosphodiesterase type-5 inhibitors derived from Sildenafil, Vardenafil and Tadalafil. Bioor- ganic & Medicinal Chemistry, 16(16), 7599–7606. doi:10.1016/j.bmc.2008.07.022 PMID:18656371
Bitencourt, M., & Freitas, M. P. (2008). MIA-QSAR evaluation of a series of sulfonylurea herbicides. Pest Management Science, 64(8), 800–807. doi:10.1002/ps.1565 PMID:18338340
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
115