Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5865_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
34 Мб
Скачать
QSAR Models towards Cholinesterase Inhibitors for the Treatment of Alzheimer’s Disease
Alkaloids: Galanthamine (Reminyl) is a tertiary amine alkaloid and a selective competitive AChEI.
It is 50 times more effective against hAChE than BuChE at therapeutic doses. Galanthamine has been approved in several countries for the treatment of AD (Sramek, Frackiewicz, & Cutler,
2000). It enhances the response of nicotinic receptors to ACh, which causes increasing ACh re­lease and other neurotransmitters. Furthermore, it increases the bioavailability of ACh to inhibit AChE. Galanthamine derivatives are synthesized, which include P11012 and P11149 (Figure 12) and which are 10-fold more potent and 6-fold more selective than galanthamine.
(-)-Huperzine A (Figure 13) is another alkaloid isolated from herb Huperzia serrata. Currently it is available not as a drug, but as a dietary supplement. It is a very potent, selective and long-acting AChEI with mild adverse effects include sleeping, nausea, and vomiting. Tests have shown that huperzine A also have neuroprotective properties. It has been a lead compound to design derivative such as the 10-methyl substituted that has 8-fold more potent than (-) -Huperzine (Figure 13) (Xiao, Yang, & Tang, 1999).
Natural Non-Alkaloid
At present, attempts have been made to discover a new non-alkaloidal AChEI to avoid the side effects that have been recorded with alkaloids. Most of the other inhibitors found are terpenes (Hostettmann, Borloz, Urbain, & Marston, 2006). Synergistic effects between monoterpenes were established. The oil or extract of two species of Salvia (Salvia lavandulifolia and Salvia officinalis) has shown improvement of cognitive function in animals, and for both healthy and early stages of AD patients. The roots of Withania somnifera have been used for 4000 years in Ayurvedic medicine for the treatment of cognitive
Figure 12. Chemical structure of galanthamine, P11012 and P11149
Figure 13. Chemical structure of (-) -huperzine a
366
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
QSAR Models towards Cholinesterase Inhibitors for the Treatment of Alzheimer’s Disease
Figure 14. Chemical structure of withaferin a, scopoletin, and coumarin106
decline associated with age. The triterpene withaferin A (Figure 14) has shown promising results in animal models. Among the coumarins scopoletin, isolated from Vaccinium oldhami, shows interesting activity against AChE enzyme (Houghton, Ren, & Howes, 2006; Fallarero, Oinonen, Gupta, Blom, Galkin, Mohan, & Vuorela, 2008).
Butyrylcholinesterase (BuChE) Enzyme
Cholinergic hypothesis have been proven to be the most successful therapeutic approach adopted for symp­tomatic relief on AD, till today by the effective use of cholinesterase inhibitors. According to this hypothesis many of the cognitive, functional, and behavioral symptoms faced by AD patients are a result of reduced levels of cholinergic neurotransmission (Holzgrabe, Kapkova, Alptüzün, Scheiber, & Kugelmann, 2007). Hence, the involvement of cholinergic neurons causes decrease levels of ACh within synapses. AChE levels also decline, to compensate for the loss of ACh, while activity of other cholinesterase enzyme BuChE (EC
3.1.1.8) increases. A significant amount of ACh is metabolized by this enzyme as the disease progresses. Another promising approach is the development of dual inhibitors for AChE and BuChE (Fallarero, Oi­nonen, Gupta, Blom, & Galkin, 2008). BuChE activity seems to correlate with AChE activity in AD and a cognitive improvement could be reached (Decker, Kraus, & Heilmann, 2008). Also, many AChEIs also inhibit BuChE, because both AChE and BuChE enzymes are found in the CNS. Moreover, AChE and Bu­ChE share 65% amino acid sequence homology even though being encoded by different genes on human chromosomes (La Du, 1994). It has been observed that in advanced AD, BuChE hydrolyses the already exhausted ACh levels, thus exacerbating the cholinergic imbalance (Greig, Utsuki, & Lahiri, 2005). Fur­ther, the low-activity BuChE in AD patients directly correlates with better cognitive function of the brain.
AN OVERVIEW OF DUAL BINDING SITE ACHEIS
The particular architecture of the enzyme AChE, in which the CS and PAS were close enough as to be simultaneously spanned by a single not too large molecule, had driven the rational design of novel classes of a dual binding site AChEIs (Inestrosa, Alvarez, Perez, Moreno, Vicente, Linker, Casanueva, Soto, & Garrido, 1996). The sole dual binding site AChEI approved for the treatment of AD is donepezil, and at
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
367
QSAR Models towards Cholinesterase Inhibitors for the Treatment of Alzheimer’s Disease
100 μM it can inhibit 22% of the AChE-induced aggregation of Aβ (Bartolini, Bertucci, Cavrini, & Andrisano, 2003). In 1996, the first family of a dual binding site AChEIs was developed by the group of Pang and Carlier, in which the two component units are connected through a heptameth­ylene chain (Pang, Quiram, Jelacic, Hong, & Brimijoin, 1996). Indeed, the most potent of Pang’s compounds, the so-called bis(7)- tacrine (1, Figure 15), turned out to be 150-fold more potent than the parent compound tacrine as an inhibitor of rat brain AChE. It is also 250-fold more selective towards AChE than BuChE. Bis(7)-tacrine has emerged as a very promising anti-Alzheimer drug candidate for its wide range of neuroprotective effects against a variety of neuronal injury models (Du, & Carlier, 2004).
The success of the dual binding site strategy is evidenced by the large increase in AChE inhibi­tory potency of these dimers or hybrids relative to the parent compounds from which they have been designed. Campiani et al. has reported that replacement of the central methylene group of bis(7)-tacrine by a protonatable methylamino group and capable of providing additional specific interactions with this mid-gorge recognition site leads to the dramatic increase in potency (Savini, Gaeta, Fattorusso, Catalanotti, Campiani, Chiasserini, Pellerano, Novellino, McKissic, & Saxena,
2003). Analogously, the group of Valenti and Recanatini has developed a very interesting AChEI, AP2238 (2, Figure 15), consisting of benzylamino and coumarin moieties as the dual binding units, respectively, connected by a p-phenylene linker and able to interact with the aromatic residues at the enzyme gorge. In addition to its high AChE inhibitory potency, AP2238 exhibits a significant β-amyloid anti-aggregating action (Piazzi, Rampa, Bisi, Gobbi, Belluti, Cavalli, Bartolini, Andri­sano, Valenti, & Recanatini, 2003). In order to improve the moderate potency of AP2238 toward both human AChE (hAChE) and hAChE-induced Aβ aggregation, Piazzi et al. has recently reported the synthesis of a large series of AP2238 derivatives (Piazzi, Cavalli, Belluti, Bisi, Gobbi, Rizzo, Bartolini, Andrisano, Recanatini, & Rampa, 2007). The suitability of the o-methoxy-substitution in AP2238 derivatives could be analogously related to its positive effect on the basicity of the nitrogen atom of the N-methyl-substituted o-methoxybenzylamino moiety and consequently on its extent of protonation.
The derivative of AP2248 leads to a higher potency after replacing the methyl group on the nitrogen atom of the benzylamino moiety by an ethyl group AP2243 (3, Figure 15). Melchiorre and Bolognesi developed a new dual binding site AChEI (4, Figure 15), composed of a unit of ta­crine as the active site interacting moiety and the phenantridinium motif of propidium as the PAS interacting moiety, connecting both units through a triamine linker (Bolognesi, Andrisano, Barto­lini, Banzi, & Melchiorre, 2005). This hybridization resulted in a 20000- and 300-fold increase in hAChE inhibitory activity relative to the parent compounds propidium and tacrine, respectively. Neuropharma, developed a new family (5, Figure 15), of tacrine-based dual binding site AChEIs which are the most potent inhibitors of AChE-induced Aβ aggregation and among the most potent inhibitors of hAChE so far reported. Greenblatt et al. designed and reported the bifunctional de­rivatives (6, Figure 15) of the alkaloid galanthamine that interacted with both the active site and PAS of the AChE (Greenblatt, Guillou, Guénard, Argaman, Botti, Badet, Thal, Silman, & Suss­man, 2004). Bermúdez-Lugo et al. reviewed several computational methods for designing AChEIs, such as docking approaches, molecular dynamics studies, quantum mechanical studies, electronic properties, hindrance effects, partition coefficients (Log P) and molecular electrostatic potentials surfaces, among other physicochemical methods that exhibit QSARs. (Bermúdez-Lugo, Rosales­Hernández, Deeb, Trujillo-Ferrara, & Correa-Basurto, (2011).
368
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
QSAR Models towards Cholinesterase Inhibitors for the Treatment of Alzheimer’s Disease
Figure 15. Different dual binding site AChE inhibitors with Aβ anti-aggregating effects
Quantitative Structure-Activity Relationship
Quantitative structure-activity relationship (QSAR) and quantitative structure-property relationship (QSPR) are based on the assumption that the structure of a compound (i.e. its geometric, steric and elec­tronic properties) must contain features responsible for its physical, chemical, and biological properties (Hansch, Hoekman, & Gao, 1996). More than a century ago, Crum-Brown and Fraser (1868) expressed the idea that the physiological action of a substance in a certain biological system (Φ) was a function (f) of its chemical constitution C:
Biological Activity (Φ) = f Physicochemical Properties(C)
QSAR/QSPR attempts to find out the quantitative correlation between physicochemical parameters of compounds (independent variable) and their pharmacodynamic, pharmacokinetic, toxicological prop­erties or biological activities (dependent variables) (Yap, Li, Ji, & Chen, 2007). Determining a QSAR generally proceed as follows:
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
369
QSAR Models towards Cholinesterase Inhibitors for the Treatment of Alzheimer’s Disease
1. Define a quantitative measure of activity (e.g., the amount of ligand needed to produce an interfer-
ence with the functioning of the target).
2. Express the ligand in some quantitative manner; that is, select a collection of numbers that char-
acterize the ligand. These numbers are called ‘molecular descriptors’ or, ‘descriptors.’
3. Determine a functional relationship between activity and the selected descriptors; that is, the search
for a mathematical function, f, that has the property that “activity = f (descriptors)” to a suitably high level of accuracy.
4. Use the know activity values, molecular descriptors and determined functional relationship to
predict the activity of new candidate compound.
QSAR methods are used to generalize experimental data in order to design or optimize new biologically active compounds that are more potent, less toxic, more selective, or satisfy other relevant criteria (Winkler,
2002). Once, a reliable QSAR model is created, it is possible to predict the activity of new compounds, and to learn which structural properties play an important role in the modeled biological response. It can also identify and describe important structural features of the compounds that are relevant to variations in molecular activities. The most widely used method for determining the functional relationship is the statistical technique of regression or least squares. Various multi-dimensional descriptors, especially 2D and 3D descriptors of molecular properties (particularly steric and electrostatic values) have been defined that can be correlated with the biological activity using statistical or machine learning techniques (Li, Yap, Ung, Xue, Li, Han, Lin, & Chen, 2007). An indispensable prerequisite for QSAR/QSPR model generation is the set of compounds with reliable measured bioactivity, also known as ‘training set molecules’. The model is validated by various methods and final model is used to predict the activity of new compounds in the series. A number of automated computer programs are available for the fast generation of various descriptors and QSAR models (Roy, 2007). Apart from drug discovery research the QSAR is being ap­plied in many other disciplines. In fact a QSAR model can be built to predict any type of physical property (Rücker, Meringer, & Kerber, 2004; Sutter, & Jurs, 1996) or biological activity/toxicity, of a given set of observations and molecular structures (Novak, & Rajagopal, 2002; Gupta, 2007).
HISTORY OF QSAR
Historically, at the turn of the 20th century, Meyer (1899) and Overton (1895) independently suggested that the narcotic (depressant) action of a group of organic compounds paralleled their olive oil/water partition coefficients. In following years on the physical organic front, the seminal work of Hammett gave rise to the “σ−ρ” culture (Hammett, 1935) in the delineation of substituent effects on organic reac­tions, while Taft devised a way for separating polar, steric, and resonance effects and introducing the first steric parameter, ES (Taft Jr, 1952).
In 1962, Hansch and Fujita published their first 2D-QSAR study on the plant growth regulators and their dependency on Hammett constants and hydrophobicity (Hansch, Maloney, Fujita, & Muir, 1962). Using the octanol/water system, a whole series of partition coefficients were measured, and thus a new hydrophobic scale was introduced (Hansch, & Fujita, 1964). The contributions of Hammett and Taft together laid the basis for the development of the QSAR paradigm by Hansch and Fujita, which combined the hydrophobic constants with Hammett’s electronic constants to yield the linear Hansch equation and its many extended forms.
370
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
QSAR Models towards Cholinesterase Inhibitors for the Treatment of Alzheimer’s Disease
Types of QSAR Models
QSAR is one of the basic tools of modern drug design science. It is founded on the systematic use of mathematical models from the multivariate point of view. Further, it has an increasing role in environ­mental sciences. QSAR models exist at the intersection of chemistry, statistics and biology (toxicology) studies. Data used in QSAR evaluations are obtained either from the literature or generated specifically for QSAR analysis and should be accurate and precise. It can consist of congeneric series of compounds or assure structural diversity even within a chemical class. This diversity has allowed the generalization of more robust QSARs, applied in an extended way. A structure– activity model is defined and limited by the nature and quality of the data used in model development and should be applied only within the model’s applicability domain.
Nowadays, several multidimensional QSAR approaches (3D-QSAR, 4D-QSAR, 5D-QSAR and 6D­QSAR) are also available depending upon the type of descriptors used (Lill, 2007). It is important to note that all QSAR approaches aims to correlate structural/compounds data with the biological activity or some other properties (Kubinyi, 1997; Dudek, Arodz, & Galvez, 2006). The 3D-QSAR methodol­ogy is much more computationally complex than the 2D-QSAR approach. In general, it involves several steps to obtain different numerical descriptors of the compound structure. First, the conformation of the compound has to be determined either from experimental data or molecular mechanics and then the resulting conformations are refined by minimizing the energy (Akamatsu, 2002). Next, the conformers in dataset have to be uniformly aligned in space. Finally, the space with immersed conformer is probed computationally for various descriptors. Examples of 3D-QSAR are Comparative Molecular Field Analysis (CoMFA) (Cramer Iii, Patterson, & Bunce, 1998), Comparative Molecular Similarity Indices (CoMSIA) (Klebe, Abraham, & Mietzner, 1994) and Self Organizing Molecular Field Analysis (SOMFA).
Some methods independent of the compound alignment have also been developed. Alignment-inde­pendent 3D-QSAR descriptors are 3D descriptors that are invariant to molecular rotation and translation in space. Thus, no superposition of compounds is required. The Comparative Molecular Moment Analysis (CoMMA) (Silverman, & Daniel, 1996) uses second-order moments (moments relate to the center of the mass and center of the dipole) of the mass distribution and charge distributions. The VolSurf approach is based on probing the grid around the molecule with specific probes, e.g. the hydrophobic interactions or hydrogen bond acceptor or donor groups (Crivori, Cruciani, Carrupt, & Testa, 2000). 3D-QSAR models are widely applicable, but often suitable for any series of congeneric compounds.
The broad family of chemical descriptor(s) includes empirical, quantum chemical or non-empirical parameters used in the 2D-QSAR approach share a common property of being independent from the 3D orientation of the compound. Empirical descriptors may be measured or estimated and include physico­chemical properties. Non-empirical descriptors can be based on individual atoms, substituents, or the whole molecule, they are typical structural features. They can be based on topology or graph theory and, as such, they are developed both from the knowledge of 2D structure, or from the 3D structural conformations of a compound. Some of the different types of descriptors and their description are pre­sented in Table 1.
Different software calculates wide sets of different theoretical descriptors, from SMILES, 2D-graphs to 3D- coordinates etc. Some of the software includes: ADAPT (Stuper, & Jurs, 1976), OASIS (Mek­enyan, & Bonchev, 1986), CODESSA (Katritzky, Petrukhin, Yang, & Karelson, 2005), MolConZ, and the DRAGON (Todeschini, Consonni, Mauri, & Pavan, 2003). Thus, descriptor selection must be performed by applying mathematical approaches with the final and crucial aim to maximize, as an optimization
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
371
QSAR Models towards Cholinesterase Inhibitors for the Treatment of Alzheimer’s Disease
Table 1. Types and description of various descriptors employed for obtaining QSAR models
Type Descriptors Description
Electronic Sum of partial charges, sum of formal charges,
dipole moment, HOMO, LUMO Spatial Radius of gyration, shadow indices, area, density, PMI, Vm Structural Molecular weight, number of rotatable bonds, number of hydrogen bond acceptor, number of
hydrogen bond donors Thermodynamics Log of partition coefficient, log of partition coefficient atom-type value, desolvation free energy
of octanol, heat of formation, molar refractivity Topological Wiener index, Zagreb index, Kier and Hall molecular connectivity index, Balaban index, Hosoya
index E-State Indices Electrotopological-state indices
parameter, the predictive power of the QSAR model, as the real utility of any model is considered its pre­dictivity. For the potency modelling, the most widely used mathematical technique is multiple regression analysis (MRA). Regression analysis is a simple approach that leads to a result that is easy to understand. For the modelling of categories, a wide range of classification methods exists, including: discriminant analysis (DA), SIMCA (Soft Independent Modeling of Class Analogy), k-NN (k-Nearest Neighbours), CART (Classification and Regression Tree), Artificial Neural Network (ANN), Support Vector Machine (SVM), etc. In these techniques, the term “quantitative” is referring to the numerical value of the variables (descriptors) necessary to classify the chemicals in the qualitative classes. Nowadays, genetic function algorithm (GFA) has gained great popularity in QSAR research. GFA and genetic PLS (G/PLS) methods developed by Rogers and Hopfinger (1994), employed to select the relevant descriptors.
Linear and Non-Linear Statistical Regression Techniques
GFA: GFA is genetics based linear method of variable selection that combines Holland’s ge-
netic algorithm (GA) with Friedman’s (1991) multivariate adaptive regression splines (MARS). It works in the following way: First, a predefined number of equations (set at 100 by default) are generated randomly, then pairs of ‘‘parent’’ equations are chosen randomly from this set of 100 equations and the ‘‘crossover’’ operations are performed at random. After some preliminary ob­servations on the initial runs, the number of GFA crossovers was set to 5000 for the present study in order to obtain a reasonable convergence. The goodness of each progeny equation is assessed by Friedman’s lack of fit (LOF) scoring, assigned by GFA, which resists over the fitting and estimates the appropriate number of variables in the equation. The LOF is given by the equation,
LOF LSE c dp m= +
/ /1
{
( )
where LSE is the least-squares error, c is the number of basis functions in the model, d is the smoothing parameter, p is the number of descriptors, and m is the number of observations in the training set. The smoothing parameter that controls the scoring bias between the equations of different size was set at
2
(1)
}
372
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
Y f
( )
QSAR Models towards Cholinesterase Inhibitors for the Treatment of Alzheimer’s Disease
a default value of 1.0. A new term was added to the offspring models with a probability of 50%. Only the linear equation terms were used for model building. The best equation out of 100 equations was taken based on the statistical parameters such as regression coefficient, adjusted regression coefficient, regression coefficient cross validation and F-test values.
G/PLS: G/PLS methodology combines GFA, and is used for the selection of descriptors to be
included in a model. In the partial least squares (PLS) regression fitting technique (Roy, Roy,
2008) that weights the relative contribution of each descriptor in the final model. PLS in particu­lar produces statistically significant solutions by identifying a linear combination of the original physico-chemical descriptors that best correlates with the biological response (Dunn Iii, Scott, & Glen, 1989). The main advantage of the G/PLS method is that it constructs QSAR equations from which most variables have been eliminated; thus, PLS avoids over fitting. Application of G/ PLS thus allows the construction of larger QSAR equations, while still avoiding over fitting and eliminating most variables.
SVM: Support vector machine (SVM), first developed by Vapnik and Cortes (1995) is a novel
machine-learning non-linear method for classification problems. The main advantage of SVM is that it was built on the structure risk minimization principle, which has been shown to be superior to the traditional empirical risk minimization principle. After the introduction of ε-insensitive loss function, SVM has been able to solve non-linear regression estimation and has demonstrated much success in QSAR studies (Yuan, Zhang, & Luo, 2009). In brief, the correlation between the structure of the training set and its biological activity, can be represented as Yi = f (x we can represent f (x
) as a linear function of
i
). Simply,
i
y w x b
= + (2)
i i i
where,
wiis the coefficient vector of the linear function and b corresponds to its coefficient.
In SVM, the basic idea is to map the data x into a higher-dimensional feature space F through a non-
linear mapping ϕ by using kernel function and then performed linear regression in this space. In the high dimensional feature space, SVM approximates the set of data with a linear function:
n
= =
where
x w x b
( )
φ x
is the high-dimensional feature space after kernel transformation, which is nonlinearly
i
φ
+
( )
i i
=
1
i
(3)
mapped from the input space x, while wiand b are coefficients. A radial basis function (RBF) kernel is frequently used kernel function in QSAR studies, out of various available kernel functions for nonlinear
transformation of the input data. The most important aim of SVM is minimization of the regularized risk function, by which the coefficients w and b are estimated. The purpose of this risk function is to find a function that has at most ε prediction errors in all the training data points and at the same time is as flat as possible. The regularized risk function is defined as:
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
373
QSAR Models towards Cholinesterase Inhibitors for the Treatment of Alzheimer’s Disease
N
2
{ }
N
( )
22
N
R C C
( ) ( , ( , )) w= +
L y f x w y f x w for y f x wε ε ε, , , , ,
( )
1 1
( )
=
i
=
1
L y f x w
ε (4)
i i
( )
2
( )
0 otherwise
N
The first term of the equation
mated by the ε insensitive loss function
1
C
L y f x w
1
=
i
L y f x wε , ,
ε( , ( , ))
i i
( )
. The term,
is the empirical error (risk), which is esti-
1
w is known as the regularized
term and is used as a measurement of function flatness. The parameter ε is known as the tube size, and it is equivalent to the approximation accuracy placed on the training data points, and C is the regulariza­tion constant determining the trade-off between the training error and the regularized term. The mini­mization of regularized risk function is constrained optimization problem which can be reformulated into dual problem formalism by using Lagrange multipliers. The performance of SVM regression is based on the kernel selection, and the optimized value of C and ε. Parameter tuning must be performed because, which values of C and ε will work the best for a particular data set is not known previously.
ANN: Artificial neural network (ANN) method is an interconnected feed forward network, which
is motivated by the way biological nervous systems, viz. the brain process information, and is a popular tool in function learning (Gasteiger & Zupan, 1993). It has the ability to learn complicated nonlinear function (dataset) with good efficacy. The theory of the ANN and its application in QSAR/QSPR studies were described extensively in many reviews (Zupan, 1994). An ANN tech­nique in general has three interconnected layers: input layer, a hidden layer, and an output layer. The input layer neurons receive the data from molecular descriptors input files, which is subse­quently passed on to the nodes of the hidden layer for further processing. The signals are then relayed onto the output layer, which sends information directly to the outside world, to a second­ary computer process or to other devices such as a mechanical control system. The connections between the nodes of each layer are assigned by randomized weight value. Supervised learning technique was used in back-propagation ANN, the network is trained by minimizing the squared error of the network’s output. This is achieved by back propagating the desired output to the trans­fer function (sigmoid function) for adjustment of weight by using a gradient descent algorithm.
QUALITY OF FIT AND QSAR MODEL VALIDATION METHODS
The success of a QSAR model is to measure the quality of fit on the available training data set. The most common statistical qualities of the equations were measured by the parameters such as explained
2
variance (R freedom (Wold, & Eriksson, 1995). While for G/PLS equations, least-squares error (LSE) was taken as a statistical measure, lack-of-fit (LOF) was noted for the GFA derived equations. The expression for R
2
and the F-test are defined as mentioned below:
R
a
374
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
), squared correlation coefficient (R2) and variance ratio (F-test) at specified degrees of
a
2
,
N d
1
− −
r
( )
QSAR Models towards Cholinesterase Inhibitors for the Treatment of Alzheimer’s Disease
n
Y Y
( )
2
R
2
R
=
a
F
=
(( ) / ( ))
1
i
1= −
( )
N R d
1 1
− −
=
n
Y Y
( )
− −
R N d
i Training
1
i
=
2
1 1
2
R d
( / )
2
In the above Equations (5-7),
2
ˆ
i i
(5)
2
(6)
. (7)
ˆ
Y , i =1,…,n are the values calculated by the QSAR model for the
dependent variable corresponding to object ‘i’, when this compound has been included in the training data set. Whereas,
Y
is the averaged value of the dependent variable for the training set, N is the
Training
number of compounds in the training set, and d is the number of descriptors in the QSAR equation.
As suggested by Tropsha, Gramatica, & Gombar, (2003) the predictive ability of a 2D-QSAR model should be tested on an external set (test set) of data that has not been taken into account during the process of developing the model. Hence, the best models were used to predict the enzyme inhibition potency value of the test set compounds. The prediction qualities of the models were judged by statisti-
2
R
cal parameters like predictive correlation coefficient actual and predicted values, with (
2
) and without (r
0
2
, squared correlation coefficient between
pred
) intercept of the test set compounds. The predic-
tive value is calculated as follows:
Y
( )
R
2
pred
1= −
pred Test
Y Y
( )
( ) ( )
Test Training
In the above equation, Y the test set compounds, and
was previously shown that the use of
characteristics. Thus, an additional parameter
(Test) ( )
pred (Test)
Y
Training
( )
. (8)
2
and Y
indicate predicted and actual activity values, respectively, of
(Test)
indicates the mean activity value of the training set compounds. It
2
R
and r2 might not be sufficient to indicate the external validation
pred
2
r
defined as r r r r
m
2 2 202
m
1=
* which penalizes a
2
Y
model for large differences between the actual and the predicted values was also calculated (Roy, 2007).
QSAR Model Cross-Validation
Cross-validation is a popular technique which is used to explore the predictive ability of the QSAR models. These models were rigorously evaluated using leave one out (LOO) or leave many out (LMO) approach, Y-randomization test and test set predictions. In the LOO or LMO approach, one compound
375
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use