Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5345_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
34 Мб
Скачать
An Introduction to the Basic Concepts in QSAR-Aided Drug Design
The basic concepts of partial least squares (PLS) was proposed by Swedish statistician Herman Wold (Jores-Kong & Wold, 1982; Wold, 1966) to generate linear models from those with non-linear parameters by projecting the predicted variables and the observable variables to a new space. A PLS model, an exten­sion of MLR, not only is used in descriptor selection step but also is used for QSAR model generation as well. It consists of three components including the latent variables, measurement, and weight components in which the weights are used to estimate the values of latent variables, so that the resulting values capture most of the variance of the independent variables that is useful for predicting the dependent variables. The overall goal is to use the descriptors to predict the responses (biological activities) in the population. This is achieved indirectly by extracting latent variables from sampled descriptors and responses, respectively. The extracted transformed descriptors are used to predict the responses. The PLS methodology has become popular in most QSAR studies related to CoMFA (comparative molecular field analysis) and CoMSIA (comparative molecular similarity index analysis), however its basic purpose was not for discriminative and classification problems (Barker & Rayens, 2003). A PLS regression can overcome the problems of standard regression when the matrix of predictors has more variables than observations, and when there is multicollinearity among descriptors (Olah, Bologa, & Oprea, 2004). Various combinations of PLS method with non-linear mathematical methods such as genetic function approximation (GA), factor analysis (FA), and orthogonal signal correction (OSC) have also been proposed to give better performance in QSAR analyses. The FA-PLS hybrid method employs FA for selecting the initial descriptors to find the relation­ships among variables in order to reduce the variables to be selected for the next step, i.e., PLS regression. A leave-one-out (LOO) method is mostly used for choosing the optimum number of components for PLS. Many examples related to the application of FA-PLS in QSAR studies have been presented, such as those used on HIV protease inhibitor mannitol derivatives, (phenylpiperazinyl-alkyl) oxindoles as selective 5-HT1A antagonists and cytochrome 3A4 inhibitory activity of structurally diverse compounds (Adhikari, Maiti, & Jha, 2010; Leonard & Roy, 2006; Roy & Pratim Roy, 2009).
GA-PLS is a combination of two calculation methods, namely, genetic algorithm (GA) (Maccari et al., 2006; Raichurkar, Shah, & Kulkarni, 2011; Williams et al., 2010) and PLS methods in which GA selects appropriate basis functions of a data model and PLS regression is used for setting the weights of the basis functions’ relative contributions in the final model which in turn makes the model suitable in construction of larger QSAR equations.
Principal component analysis (PCA), a statistical procedure, was firstly invented by Karl Pearson (Pearson, 1901) and was later independently developed by Harold Hotelling (Hotelling, 1933) that uses orthogonal transformation to convert a set of observations of correlated variables into a set of values of linearly uncorrelated variables known as principal components for generating the predictive models. In this transformation, the first principal component has the largest possible variance, and each succeeding component has the highest variance such that it is orthogonal to the preceding components.
The independent property of principal components is guaranteed if the data set is normally distributed and the number of them is less than or equal to the number of original descriptors (Bagheri, Omidikia, & Kompany-Zareh, 2013). However, the structure of PCA is similar to that of LDA and FA on constructing the linear combinations of variables which best fit the data (Martinez & Kak, 2001).
Goodarzi et al. built their QSAR model based on PCA-ANFIS, a hybrid methodology consisting of the principal component analysis and the adaptive neuro-fuzzy inference systems for predicting the anti-HIV reverse transcriptase activities of tetrahydroimidazo [4,5,1-jk][1,4] benzodiazepine derivatives (Goodarzi & Freitas, 2010). In another study, PCA was employed to generate a robust model to predict the Src kinase inhibitory activity of 4-anilino-3-quinolinecarbonitriles (Sun, Zheng, Wei, Chen, & Ji, 2009).
16
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
An Introduction to the Basic Concepts in QSAR-Aided Drug Design
Non-Linear Methods
The techniques discussed above were based on linear methods or a combination of them to develop the QSAR models. The linear methods may have a strong tendency to overfit the data once the large number of descriptors is used where the number of molecules is limited. In such cases, the non-linear nature of the problem requires the use of non-linear methods for the prediction of the activities based on the struc­tural features of molecules. Modern QSAR models are based on the variations of the Hansch analyses that predicts a biological property as a statistical correlation with steric, electronic, and hydrophobic indices by using more powerful non-linear models, such as artificial neural networks (ANNs), support vector machines (SVMs), or other machine learning algorithms. The structural descriptors are numeri­cal representations of some important molecular features, such as empirical indices (Hammett and Taft substituent constants), physical properties (octanol-water partition coefficient, dipole moment, aqueous solubility), counts of substructures or substituents, graph descriptors (Dury, Latour, Leherte, Barberis, & Vercauteren, 2001; Ivanciuc, 2003), topological (Fernandez-Blanco, Aguiar-Pulido, Munteanu, & Dorado, 2013; Fernandez-Lozano et al., 2013; Ivanciuc, 2013) or electrotopological (Sapre, Pancholi, Gupta, & Sapre, 2008; Wang, Liu, Shan, & Shi, 2010; Wang, Jiang, Pan, Cao, & Cui, 2009) and con­nectivity indices (Farahani, 2013; Mozrzymas, 2013), quantum chemical descriptors (atomic charges, HOMO and LUMO energies) (Angulo & Antolin, 2007; Hemmateenejad, Yousefinejad, & Mehdipour,
2011), and molecular interaction field descriptors (steric, electrostatic, and hydrophobic descriptors) (Baskin & Zhokhova, 2013; Gussregen et al., 2012). Some of these molecular parameters were described previously in Molecular descriptors section.
To find valuable drug candidates, many computational chemistry tools have been employed, among which artificial neural network (ANN) has a special place in modeling approaches because it provides solution to the complex problems with high accuracies.
The simulation of brain’s neural function by variety of artificial neural networks (ANNs) was initially proposed by McCulloch and Pitts (Mc & Pitts, 1948; Pitts & Mc, 1947), and Rosenblatt (Rosenblatt,
1962). The basic structure of an ANN consists of input, hidden, and output layers, commonly structured in a multilayer feedforward (MLF) network usually trained by the backpropagation algorithm to model the non-linear relationships between input variables and the output properties by adjusting their weights and biases. The system uses the mean squared error (MSE) through the training process. A special type of feedforward ANN is the radial basis function (RBF) neural network with only three layers, including input layer, a hidden layer and a linear output layer (Moody & Darken, 1989) which can predict wide range of properties, such as protein folding (Abbasi, Ghatee, & Shiri, 2013), protein interaction (Chen, Xu, Yang, Zhao, & He, 2012), and beta barrel membrane protein detection (Chen, Ou, Lee, & Gromiha, 2011; Ou, Gromiha, Chen, & Suwa, 2008).
Graph machine neural networks, which work based on the graph representation of the molecular structures, have been implemented in variety of computational chemistry tools with the capability of performing QSAR analyses (Baskin, Ait, Halberstam, Palyulin, & Zefirov, 2002; Chupakhin, Marcou, Baskin, Varnek, & Rognan, 2013) such as ChemNet (Kireeva, 1995), and MolNet (Ivanciuc, 1999).
Genetic algorithm (GA) is another type of biologically inspired algorithm initially proposed and developed by Holland (Holland, 1975) and Goldberg (Goldberg, 1989). GA is based on the Darwinian evolution of a population of individuals (i.e., chromosomes) to solve the high dimensional non-linear problems in which each chromosome is a solution to the problem and the chromosomes can be encoded as binary, continuous or a hybrid of both and their qualities are assessed by their fitness values. GA and
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
17
An Introduction to the Basic Concepts in QSAR-Aided Drug Design
its variants have been employed in many cheminformatics and computational chemistry applications. The followings are some demonstrated examples of its use; 3D-QSAR studies on novel phosphodiesterase-4 inhibitors, pharmacophore identification for benzotriazine derivatives, identification of novel inhibitors for cancer targets, acetylcholinesterase, HIV-1 protease, and ion-channels, as well as developing anti­protozoan compounds (Fernandez, Caballero, Fernandez, & Sarai, 2011; Pourbasheer, Riahi, Ganjali, & Norouzi, 2010; Sahin, Saripinar, Yanmaz, & Gecen, 2011). Evolutionary algorithms such as GA are also successfully applied to descriptor selection and global optimization of QSAR models (Perez-Castillo et al., 2012; Sahin, et al., 2011).
Ant colony optimization (ACO), an agent/colony based algorithm developed by Dorigo et al. (Dorigo, Maniezzo, & Colorni, 1996), simulates the foraging behavior of ants. The main aim of the ACO algo­rithm (i.e., the ant population) is to search and find the shortest path to the food resource by means of a chemical substance called pheromone accumulated on the paths that are more regularly explored by ants. The larger the amount of pheromone, the larger the probability for a path to be the shortest one to reach the food resource. The amount of pheromone in each path will be updated according to the amount of pheromone deposited by an ant. Like the GA, ACO method and its modified versions have also many applications in optimization problems such as chemistry and drug design where the target solution space is huge. There are several hybrid implementations including ACO-MLR (Goodarzi, Jensen, & Vander Heyden, 2012), modified ACO-neuro fuzzy interference for serotonin (5-HT7) receptor inhibitors (Jalali­Heravi & Asadollahi-Baboli, 2009), and ACO-random forest (Patil, Raj, Shingade, Kulkarni, & Jayara­man, 2009) for generating the QSAR models in which the main object is to cluster the new compound in large datasets of chemical compounds based on their similarities measured by molecular descriptors. The classical ACO can only be used for optimization problems related to discrete variables, whereas many QSAR studies need to optimize the continuous variables that has been addressed as an extension to ACO proposed by He and co-workers (He, Chen, & Zhao, 2006). Izrailev and Agrafiotis proposed two ACO approaches to find the best regression trees where each ant was an indicator for a regression tree, and to select required features by using ANN (i.e., ANTSELECT) on three QSAR datasets (Izrailev & Agrafiotis, 2001; Izrailev & Agrafiotis, 2002).
Particle swarm optimization (PSO), known as swarm intelligence (SI) proposed by Kennedy and Eberhart (Kennedy & Eberhart, 1995), is used for solving the non-linear optimization problems by simulating the social behaviors such as swarming, herding or flocking of various spices such as bees, birds, or fish started from a random position with random velocity in a search space that can be applied to both binary or continuous variables which is capable of fast convergence by using a small popula­tion and small number of iterations. The new positions in PSO algorithm are determined by either the particle’s experience obtained from its previous movement or the best particle’s memory in the swarm. That is why, in several cheminformatics studies, PSO and its modifications have been known as efficient replacements of GA and ACO for global optimization. Moreover, PSO has been successfully applied in many non-linear problems such as QSAR (Cheng, Zhang, & Zhou, 2011), protein motif discovery (Har­din & Rouchka, 2005), enzyme-inhibitor docking (Goodarzi, Saeys, Deeb, et al., 2013; Lu, Shen, Jiang, Shen, & Yu, 2004), protein-ligand geometry (i.e., Tribe-PSO in AutoDock) (Goodsell, 2009; Goodsell, Morris, & Olson, 1996; Morris, Goodsell, Huey, & Olson, 1996; Tiwari, Saxena, Pant, & Srivastava,
2012), and finding the global minimum geometry of chemical compounds. Based on the literature, the performance of PSO and its variants is promising, and they are popular function optimization methods in many machine learning algorithms.
18
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
F
F
An Introduction to the Basic Concepts in QSAR-Aided Drug Design
STATISTICAL TREATMENTS AND MODEL VALIDATION
Data analysis and processing is usually accompanied by evaluation procedure. Evaluation of the devel­oped models constitutes fundamental and inevitable step in QSAR studies and are paid particular atten­tion. Any QSAR model should be evaluated prior to its use for data interpretation as well as prediction of biological activities. There are different tools to assess the quality, validity, and reliability of the generated QSAR models. The foremost judgment on a QSAR model is based on the statistical param-
2 2
R r
eters of corresponding model. Squared correlation coefficient ( ported parameters in QSAR models that assesses goodness-of-fit and describes the degree of the varia­tion in the variables explained by the regression equation ranging from zero to unity. The ideal and perfect model has an value of
R2demonstrates no explanatory ability of the model. The other parameter is residual standard
R2 value of 1 explaining all the response (i.e., biological activity) whereas zero
deviation (denoted by s) indicative of how well a model can predict the observed activities. Therefore, the minimum requirements for a good QSAR model is to have high
The fisher statistics (
) is another QSAR statistical parameter for assessing the significance of the models by carrying out analysis of variance (ANOVA). It is described as the ratio of the explained mean square to the residual mean square. This term shows how well a regression equation fits to the training data set in a specified confidence interval. A significant QSAR model should have a high should be kept in mind that the best amount of statistical parameters mentioned above are not sufficient criteria for model reliability and validity and can not be considered as the indicators of predictive capa­bility of the QSAR model. They just indicate how a model is capable to reproduce the biological activ­ity (in this case) for the training data set.
As stated earlier there are several strategies for the validation of QSAR models such as,
, ) is one the most commonly re-
R2(close to unity) and low s values.
value. It
1. Cross validation
2. External validation, and
3. Y-scrambling methods.
There are two types of cross validation (CV) methods namely, leave-one-out (LOO) and Leave-many/ group/several-out (LMO) which evaluate the predictivity. They are known as “internal cross validation”. In these methods one or different portion of compounds are removed in such a way that each data point is taken away once from the entire data set followed by iteratively fitting the model to the remaining data points and prediction of the biological activity of the left out molecules. Comparing the observed vs
predicted values of the biological activities leads to calculation of the “cross validating routinely called
Q2 or q2. q2 is calculated according to the following equation (Golbraikh & Tropsha,
R2”, which is
2002; Tropsha, Gramatica, & Gombar, 2003):
2
q 1
= −
N
( )
=
i 1
N
( )
=
i 1
y y
y y
i i
i
2
= −
2
PRESS
1
SSD
(7)
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
19
An Introduction to the Basic Concepts in QSAR-Aided Drug Design
SSD
y
Here
ity
, over the entire data set. PRESS , the predictive residual sum of squares, is the sum of squared
differences between the actual activity
is the sum of squared deviations for each actual activity value yi from the average activ-
yi and the predicted activity
y
. N is the number of data points.
i
LOO is usually performed for small data set (e.g., 25 compounds), while for big data set, the LMO is preferred. Generally, LMO is considered as a stronger validation approach compared to LOO method due to the fact that in each cycle of LMO method constant perturbation is performed by eliminating
certain number of data points while in LOO method by increasing the number of data points just
of entire data set is removed. Consequently, smaller perturbation in data structure leads to obtaining values too close to
R2 (Gramatica, 2007; Wold & Eriksson, 1995). Although, q2determines goodness-
N
1
th
q
of-prediction, however, it should not be regarded as merely criterion of the robustness and predictive ability of the model. Some QSAR practitioners consider the values of q20 5> . are good enough for validity of a model, but in reality, greater values for
2
q
are essential conditions for validity of a model but not adequate. Golbraikh and Tropsha demonstrated that in addition to internal validation methods, further validation methods should be used for establishing a reliable QSAR model (Golbraikh & Trop­sha, 2002). More discussion on this issue can be found elsewhere (Dearden, Cronin, & Kaiser, 2009; Gramatica, 2007; Tropsha, et al., 2003). Therefore, internal cross validation gives only an approximation of the accuracy of a QSAR model. To be more rigorous, the “external validation” approach must be employed for establishing a truly predictive QSAR model. To this end, a validation data set (i.e., data points not used in the model development procedure) is chosen and the model is used to predict their biological activities. The predictions are compared with actual activities and the agreement among them is evaluated. The higher the correlation between the predicted and the observed activities for the valida­tion data set, the more robust is the model. Therefore, the entire available data set is split into training and test data sets. The important requirement is that both of data sets should span in such a manner that all ranges in terms of biological activity and structure are adequately covered and the number of test data set should not be less than five (Golbraikh & Tropsha, 2002; Gramatica, 2007; Tropsha, et al.,
2003). Generalizability of the developed QSAR model for new chemical entities is attained where the external validation step is carried out at the model development stage.
The last but not least validation strategy for evaluating robustness of the QSAR model is using “Y­scrambling” method (also called Y-randomization test). This technique is based on randomization of the responses (i.e. biological activities) in intact independent variables (i.e. molecular descriptors) followed by full data analysis using the scrambled activities. This process is performed repetitively for at least ten
trials. The resultant statistics (i.e.,
R2 and q2) are expected to be very low compared to that of the
original data. Otherwise, the model is not reliable and applicable (Golbraikh & Tropsha, 2002; Tropsha, et al., 2003; Wold & Eriksson, 1995).
For a predictive QSAR model, the following statistical criteria are also recommended (Golbraikh & Tropsha, 2002; Tropsha, 2010; Tropsha, et al., 2003):
2
q .20 5>
1.
2.
R20 6> .
2
R
2
0
.
0 1
<
2
k
( )
R R
3.
0 85 1 15. .
4.
20
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
An Introduction to the Basic Concepts in QSAR-Aided Drug Design
2
q
Here,
is cross validated correlation coefficient, R2 is squared correlation coefficient between
predicted and observed activities, the correlation coefficients between predicted and observed activities
2
R
when the regression line passes the origin, and k shows the corresponding slope. All of the criteria
0
are calculated for test set compounds except the first one (i.e.,
q2).
For developing a highly predictive model, serious concern should be paid during the evaluation step
and a trustworthy QSAR model should encompass carefully the above outlined principles.
3D-QSAR STUDIES
The experimentally measured biological activity of the compounds arises from the interaction(s) be­tween a molecule and its target. Therefore, determining the stereoelectronic properties accounting for ligand-target recognition process and the subsequent interaction(s) is the main goal of a three dimen­sional quantitative-structure activity relationship (3D-QSAR) study. This technique is introduced as an integration of classical QSAR with molecular modeling approaches. From computational point of view 3D-QSAR is more complex than classical QSAR and provides depiction of possible important interac­tions between a series of compounds and target of interest in a quantitative fashion. In other words, to fulfill these interactions, compounds should have a minimum structural requirement which is usually referred to as “pharmacophore”. A pharmacophore abstracts the molecules into the functional groups capable of interacting with target structure. The structural descriptors used for 3D-QSAR studies are three dimensions in nature originated from measured interaction energies between a probe atom/group around a series of molecules.
The attractive advantage of 3D-QSAR relies on ability of predicting the interaction(s) even having
no information about the active site of target that can be a basis for ligand-based drug design (LBDD).
There are different steps for a 3D-QSAR approach. The first step is collecting the biological data and selecting the training set. The biological activities should have a wide range of activity as possible. The next step is conformational analysis of the molecules. In this step, the bioactive conformers are determined. Bioactive conformers refer to the conformation of the molecules that bind to the receptor. This kind of information is achievable both by experimental and theoretical techniques. The useful experimental data can be provided by X-ray crystallography and NMR spectroscopy (Verma, et al.,
2010). In the case of theoretical approaches where the 3D structure of receptor is unknown, the com­parative modeling can be applied to develop most probable model structure of the target for evaluation of ligand-receptor interactions. In order to determine the bioactive conformations of the ligands some criteria are considered for example, the energies of all conformations should be near the lowest energy and all training set molecules should be in similar conformations where the pharmacophoric components are aligned (Young, 2009). Bioactive conformers are aligned/superimposed over each other. Then, the molecular interaction fields (MIF), which will be explained in more details below, are calculated to be used as molecular descriptors. Different multivariate statistical techniques such as partial least square (PLS) and principal component analysis (PCA) are utilized for model generation by extracting the es­sential information from large pool of descriptors for the purpose of correlating with biological activity. Then, favorable and unfavorable regions for biological activities are visualized. The developed model is evaluated using cross validation methods and can be interpreted for subsequent important applications such as activity prediction for new compounds.
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
21
An Introduction to the Basic Concepts in QSAR-Aided Drug Design
Generally, for classification of 3D-QSAR methods one may follow two approaches: alignment­dependent methods and alignment-independent methods. In the former, the calculation of descriptors is strongly influenced by the alignment of the molecules guided by experimental data or computational methods whereas in the latter case, the superposition of molecules is not required.
The most well-known and prototype of alignment-dependent methods is comparative molecular field analysis (CoMFA) which is introduced in 1988 (Cramer, Patterson, & Bunce, 1988). It is based on the calculation of electrostatic and steric interaction energies (i.e. MIFs) between a probe and a series of molecules (Dudek, et al., 2006). MIF explains the interaction energy variation between a selected probe and a set of molecules that are previously overlaid over each other and positioned at the center of a 3D grid box (Cruciani, 2006) (shown in Figure 6). A probe is a small molecule like water or a specific fragment such as methyl placed in a lattice of grid points. Each grid point (intersection of the lattice) defines a point in space and becomes a descriptor variable in 3D-QSAR. In the case of CoMFA, they are specified as electrostatic or steric descriptors. Several force fields at different level of approximation are utilized for determination of interaction energies. A force field represents a mathematical equation containing standard bond lengths, bond angels, dihedral angles, interatomic distances and the other parameters. The standard (6-12) Lennard-Jones function and Coulomb’s law are employed in the force fields of CoMFA for computation of interaction energies (Kerdawy, Gussregen, Matter, Hennemann, & Clark, 2013; Verma, et al., 2010).
Comparative molecular similarity indices analysis (CoMSIA) is the other 3D-QSAR alignment­dependent method based on similarity indices of the molecules. It is too similar to CoMFA. Steric, electrostatic, hydrophobic and hydrogen bonding properties are considered for calculating descriptors using Gaussian-type potential function (Klebe, Abraham, & Mietzner, 1994). One of the advantages of CoMSIA, is that it provides the accurate information in grid points positioned within the molecule (Dudek, et al., 2006; Munteanu et al., 2010; Verma, et al., 2010). Both the CoMFA and the CoMSIA
7
are implemented in SYBYL
program, Tripos1, Inc.
Figure 6. Molecular interaction field calculation based on interaction energy between the probe mol­ecule (i.e., circles) and the molecule in a 3D grid box; each grid point represents a descriptor variable in 3D QSAR analysis.
22
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
An Introduction to the Basic Concepts in QSAR-Aided Drug Design
The GRID28 approach, introduced by Goodford in 1985 (Goodford, 1985), is another useful method in 3D-QSAR studies. The methodology proceeds by calculation of total interaction energies between different chemical probes and data set. GRID determines energetically favorable binding sites on known molecular structures as well as surface properties. In addition to steric and electrostatic potentials, hydrogen bonding and hydrophobic potentials are also included. The major advantage of this method is that GRID potentials have been developed based on experimentally determined ligand-protein com­plexes. Moreover, the availability of wide variety of probes provides the modeling of different types
29
of interactions (Cruciani, 2006; Verma, et al., 2010). GRID is supplied by Molecular Discovery
, Ltd.
The above mentioned 3D-QSAR methods rely on alignment step, however, there is other approach where the 3D-QSAR models are developed free from the alignment of studied compounds. Alignment­independent methods eliminate the superpositioning process which can be an advantage compared to the alignment-dependent methods. One of the examples of alignment-independent 3D-QSAR methods is comparative molecular moment analysis (CoMMA). In this method second-order moments of mo­lecular mass and charge distribution are considered for generating the molecular similarity descriptors (Silverman & Platt, 1996). These kinds of descriptors are insensitive to the molecular alignment and are sensitive to the molecular conformation (Dudek, et al., 2006; Verma, et al., 2010).
GRid-INdependent Descriptors (GRIND) comprise a class of alignment-independent descriptors that were developed by Pastor et al. in 2000 (Pastor, Cruciani, McLay, Pickett, & Clementi, 2000). In GRIND method, the descriptors are obtained from an MIF collection calculated by probing of grid with chemical probes as before and then, the most favorable energies of interactions (hot spots) are selected. This is followed by encoding the relative position of specified regions into a group of values characterizing hot spots positioned at distinct distance ranges (Cruciani, 2006; Dudek, et al., 2006;
30
Duran, Zamora, & Pastor, 2009). GRIND method is implemented in Pentacle
29
Discovery Ltd
.
program, Molecular
These various methods outlined above are just few examples of two approaches (namely alignment dependent and alignment independent methods). More information about the other different approaches can be found in the literature (Akamatsu, 2002; Cruciani, 2006; Datar, Khedkar, Malde, & Coutinho, 2006; Dudek, et al., 2006; Gohlke & Klebe, 2002; Lushington, Guo, & Wang, 2007; Polanski, Giele­ciak, & Bak, 2002; Polanski & Walczak, 2000; Verma, et al., 2010; Walters & Hinds, 1994).
The application of multivariate chemometric techniques is inevitable because of the huge number of generated descriptors, as the traditional linear regression analysis can not be used in these cases. Partial least squares (PLS) and principal component analysis (PCA) are extensively used for data analysis, by correlation of 3D-descriptors (dependent variables) with related biological activities (independent variables) assuming a linear relationship among them. Whenever, the relationship is non-linear, chemometric methods such as artificial neural network is preferred.
31
An advanced PLS method called GOLPE (Generating optimal linear PLS)
was introduced in 1993, with the aim of developing highly reliable models with improved predictivity. In this method, fractional factorial design (FFD) is primarily used for combination of variables and in this process only variables which increase the predictive power of the models are selected for final PLS and clas­sified according to their predictivity contribution (Baroni et al., 1993).
For evaluation procedure cross validation methods are applied as explained in Statistical treatment and model validation section. Leave-one-out and leave-group-out are used for internally validating the model although for further validation, external validation is strongly recommended.
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
23
An Introduction to the Basic Concepts in QSAR-Aided Drug Design
Multi-Dimensional QSAR
The most pronounced recent advance in QSAR methodology was the expansion of its dimensionality. These multi-dimensional methods are steadily establishing their own place in QSAR studies by incor­porating extra dimensions, such as physical characteristics, to overcome the obstacles of the existing QSAR methods. Multi-dimensional QSAR approaches are extended forms of 3D-QSAR methods in which additional multivariate molecular descriptors derived from 3D molecular representations, are included. Table 1 shows different multi-dimensional QSAR methodologies and their properties. Hopefinger and his coworkers were the pioneers of 4D-QSAR development (Hopfinger et al., 1997). In 4D approach, the “fourth dimension” also called “ensemble sampling” is determined by including conformational flexibility and the freedom of alignment by ensemble sampling the spatial features such as conforma­tions, orientations, and protonation states for each molecule (Andrade, Pasqualoto, Ferreira, & Hopfin­ger, 2010; de Melo & Ferreira, 2012; Martins, Barbosa, Pasqualoto, & Ferreira, 2009; Polanski, 2009; Vedani, Dobler, Dollinger, et al., 2005). In this method, the bias associated with the choice of bioactive conformation is reduced (Vedani, Dobler, Dollinger, et al., 2005). 4D-QSAR analysis can be used as receptor independent (RI) and receptor dependent (RD) methods. In RI-4D-QSAR method the maxi­mum structural information is obtained through pharmacophoric feature of substituents on the ligands’ scaffold. When the geometry of receptor is known, RD-4D-QSAR analysis is utilized for mapping the ligand-receptor interaction (Damale, Harke, Kalam Khan, Shinde, & Sangshetti, 2014; Polanski, 2009). An example of 4D-QSAR technique, which has been developed based on CoMFA method is “LQTA­QSAR” (LQTA, Laborato´rio de Quimiometria Teo´rica e Aplicada). In this method, conformational ensemble profile is generated for each compound followed by obtaining the 3D descriptors through calculation of intermolecular interaction energy at each grid point considering probes and all aligned conformations obtained from molecular dynamics (MD) simulation (Damale et al., 2014; de Melo & Ferreira, 2012; Martins et al., 2009). Further progress in adding dimensionality to the QSAR techniques was brought about by considering the changes in the conformation of the receptor upon ligand binding (referred to as induced fit effect). Induced fit effect is considered as adaptation of the binding pocket to the individual ligand topology (Damale et al., 2014; Vedani, Descloux, Spreafico, & Ernst, 2007; Vedani & Dobler, 2002; Vedani, Dobler, Dollinger, et al., 2005). This topology alteration in the binding site through induced fit effect may influence the properties of the binding pocket such as hydrophilic, hydrophobic, and dielectric characteristics (Vedani et al., 2007). Addition of another dimension leads to development of 6D-QSAR which takes into account the solvation function as well as those already covered up to the fifth dimension. In 6D-QSAR, different solvation models are incorporated by mapping the surface area with different solvent properties (Damale et al., 2014; Vedani, Dobler, Dollinger, et al.,
2005). By adding another dimension to 6D-QSA, a new paradigm known as 7D-QSAR is developed by including the real or virtual target-based receptor models (Polanski, 2009).
There are numerous examples based on newly developed multi-dimensional QSAR analyses. In a study by de Melo and Ferreira, inhibitory activity for 85 HIV-1 integrase strand transferase was inves­tigated using LQTA-QSAR method. In this work, a 4D-QSAR model was generated based on confor­mational ensemble profile using GROMACS molecular dynamic package. The resulted model showed good predictive ability and could be used for designing novel HIV-1 integrase strand transfer inhibitors (de Melo & Ferreira, 2012). In another study, Vera-DiVaio and co-workers, developed a successful 4D­QSAR model to predict the toxicity of a new series of N-phenylpyrazole benzylidene-carbohydrazide
2
with R study based on 5D-QSAR analysis was performed by Vedani et al. In this study, novel designed ligands
and Q2 values of 0.85 and 0.79, respectively (Vera-Divaio et al., 2009). A receptor-modeling
24
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
An Introduction to the Basic Concepts in QSAR-Aided Drug Design
Table 1. Dimensionality of QSAR approaches (Polanski, 2009; Vedani, Dobler, & Lill, 2006)
Dimension Descriptors
0D-QSAR Atom and molecular counts, molecular weight, sum of atomic properties 1D-QSAR Fragment counts such as electronic constraints, hydrophobic constraints and steric constraints 2D-QSAR Topological descriptors such as constitutional, topology, total polar surface area, electrostatic and quantum
3D-QSAR Geometrical, atomic coordinates, or energy grid descriptors and static ligand representation 4D-QSAR Like 3D-QSAR + multiple 3D ligand representations (ensemble of conformers, orientations, protonation states,
5D-QSAR Like 4D-QSAR + induced- fit 6D-QSAR Like 5D-QSAR + solvation scenarios 7D-QSAR Real receptor or virtual target-based receptor model
chemical, geometrical and molecular fingerprints property of the chemical compound
tautomers and stereoisomers)
were proposed to inhibit chemokine receptor-3 (CCR3) at the low-nanomolar range using induced fit scenarios (Vedani, Dobler, & Lill, 2005). Dobler et al. developed a virtual laboratory on the Internet for the toxicity prediction of drugs on five receptor types (i.e., aryl hydrocarbon (Ah), 5HT GABA
, and estrogen receptor) using a 5D-QSAR method (Dobler, Lill, & Vedani, 2003).
A
, cannabinoid,
2A
In an investigation by Peristera et al, binding mode and affinity of anabolic steroids toward mineralo­corticoid receptor was studied using 6D-QSAR (Peristera, Spreafico, Smiesko, Ernst, & Vedani, 2009). The generated model was resulted from flexible docking of inhibitors to the crystal structure of the recep-
2
tor with cross validated and predictive R
of 0.81 and 0.66, respectively. The model was claimed useful for in silico identification of compounds with endocrine-disrupting potential. The similar examples can be found in literature (Ducki, Mackenzie, Lawrence, & Snyder, 2005; Oberdorf, Schmidt, & Wunsch, 2010; Vedani et al., 2007; Vedani, Dobler, & Lill, 2005).
Whether or not increasing the dimensionality in QSAR studies will lead to generation of highly predic­tive QSAR models is a debatable issue. There are examples in which the 2D-QSAR can perform as good or better than 3D-QSAR methods in predicting the biological activity (Dastmalchi, Hamzeh-Mivehroud, & Asadpour-Zeynali, 2012; Jimenez Villalobos, Gaitan Ibarra, & Montalvo Acosta, 2013). In contrast, there are also evidences that show the increasing the dimensionality improves the predictive power of QSAR models (Katritzkya, Slavova, Dobcheva, & Karelson, 2007; Vedani, Dobler, & Lill, 2005). Fol­lowing two individual studies present further examples on the performance evaluation of different multi­dimensional QSAR analyses. The inhibitory activities of 36 indole amide hydroxamic acids on histone deacetylases were analyzed based on 2D- (with over 800 descriptors) and 3D-QSAR (using steric and electrostatic fields) methods (Katritzkya et al., 2007), and the overall results (Table 2) showed that the performance of the QSAR model depends mostly on the type of the chosen variable selection method and the descriptors. It may be hard to predict which specific QSAR method from dimensionality point of view will perform the best. However, choosing the best method for a QSAR study can be achieved by trial-and-error or surveying the available literature. Damale et al. carried out 3D to 6D QSAR analyses
2
using 106 estrogen receptor ligands, and concluded that the cross-validated and predictive R
values increase as the dimensionality of the QSAR method increases (Damale et al., 2014). Additionally, one may conclude that if the relevant method is selected, the best QSAR model can be achieved by increas­ing the dimension (Myint & Xie, 2010). The multi-dimensional methods (i.e., 4D-6D approaches) are new in QSAR studies and have been used just in a limited number of investigations. Therefore, it is too early to conclude that “the higher the dimensionality, the higher the performance”.
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
25