Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5587_Библиотеки_им_академика_М_И_Перельмана.pdf

QSAR Models towards Cholinesterase Inhibitors for the Treatment of Alzheimer’s Disease
• Alkaloids: Galanthamine (Reminyl) is a tertiary amine alkaloid and a selective competitive AChEI.
It is 50 times more effective against hAChE than BuChE at therapeutic doses. Galanthamine
has been approved in several countries for the treatment of AD (Sramek, Frackiewicz, & Cutler,
2000). It enhances the response of nicotinic receptors to ACh, which causes increasing ACh release and other neurotransmitters. Furthermore, it increases the bioavailability of ACh to inhibit
AChE. Galanthamine derivatives are synthesized, which include P11012 and P11149 (Figure 12)
and which are 10-fold more potent and 6-fold more selective than galanthamine.
(-)-Huperzine A (Figure 13) is another alkaloid isolated from herb Huperzia serrata. Currently it is
available not as a drug, but as a dietary supplement. It is a very potent, selective and long-acting AChEI
with mild adverse effects include sleeping, nausea, and vomiting. Tests have shown that huperzine A also
have neuroprotective properties. It has been a lead compound to design derivative such as the 10-methyl
substituted that has 8-fold more potent than (-) -Huperzine (Figure 13) (Xiao, Yang, & Tang, 1999).
Natural Non-Alkaloid
At present, attempts have been made to discover a new non-alkaloidal AChEI to avoid the side effects
that have been recorded with alkaloids. Most of the other inhibitors found are terpenes (Hostettmann,
Borloz, Urbain, & Marston, 2006). Synergistic effects between monoterpenes were established. The oil
or extract of two species of Salvia (Salvia lavandulifolia and Salvia officinalis) has shown improvement
of cognitive function in animals, and for both healthy and early stages of AD patients. The roots of
Withania somnifera have been used for 4000 years in Ayurvedic medicine for the treatment of cognitive
Figure 12. Chemical structure of galanthamine, P11012 and P11149
Figure 13. Chemical structure of (-) -huperzine a
366
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use

QSAR Models towards Cholinesterase Inhibitors for the Treatment of Alzheimer’s Disease
Figure 14. Chemical structure of withaferin a, scopoletin, and coumarin106
decline associated with age. The triterpene withaferin A (Figure 14) has shown promising results in
animal models. Among the coumarins scopoletin, isolated from Vaccinium oldhami, shows interesting
activity against AChE enzyme (Houghton, Ren, & Howes, 2006; Fallarero, Oinonen, Gupta, Blom,
Galkin, Mohan, & Vuorela, 2008).
Butyrylcholinesterase (BuChE) Enzyme
Cholinergic hypothesis have been proven to be the most successful therapeutic approach adopted for symptomatic relief on AD, till today by the effective use of cholinesterase inhibitors. According to this hypothesis
many of the cognitive, functional, and behavioral symptoms faced by AD patients are a result of reduced
levels of cholinergic neurotransmission (Holzgrabe, Kapkova, Alptüzün, Scheiber, & Kugelmann, 2007).
Hence, the involvement of cholinergic neurons causes decrease levels of ACh within synapses. AChE levels
also decline, to compensate for the loss of ACh, while activity of other cholinesterase enzyme BuChE (EC
3.1.1.8) increases. A significant amount of ACh is metabolized by this enzyme as the disease progresses.
Another promising approach is the development of dual inhibitors for AChE and BuChE (Fallarero, Oinonen, Gupta, Blom, & Galkin, 2008). BuChE activity seems to correlate with AChE activity in AD and
a cognitive improvement could be reached (Decker, Kraus, & Heilmann, 2008). Also, many AChEIs also
inhibit BuChE, because both AChE and BuChE enzymes are found in the CNS. Moreover, AChE and BuChE share 65% amino acid sequence homology even though being encoded by different genes on human
chromosomes (La Du, 1994). It has been observed that in advanced AD, BuChE hydrolyses the already
exhausted ACh levels, thus exacerbating the cholinergic imbalance (Greig, Utsuki, & Lahiri, 2005). Further, the low-activity BuChE in AD patients directly correlates with better cognitive function of the brain.
AN OVERVIEW OF DUAL BINDING SITE ACHEIS
The particular architecture of the enzyme AChE, in which the CS and PAS were close enough as to be
simultaneously spanned by a single not too large molecule, had driven the rational design of novel classes
of a dual binding site AChEIs (Inestrosa, Alvarez, Perez, Moreno, Vicente, Linker, Casanueva, Soto, &
Garrido, 1996). The sole dual binding site AChEI approved for the treatment of AD is donepezil, and at
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
367

QSAR Models towards Cholinesterase Inhibitors for the Treatment of Alzheimer’s Disease
100 μM it can inhibit 22% of the AChE-induced aggregation of Aβ (Bartolini, Bertucci, Cavrini,
& Andrisano, 2003). In 1996, the first family of a dual binding site AChEIs was developed by the
group of Pang and Carlier, in which the two component units are connected through a heptamethylene chain (Pang, Quiram, Jelacic, Hong, & Brimijoin, 1996). Indeed, the most potent of Pang’s
compounds, the so-called bis(7)- tacrine (1, Figure 15), turned out to be 150-fold more potent than
the parent compound tacrine as an inhibitor of rat brain AChE. It is also 250-fold more selective
towards AChE than BuChE. Bis(7)-tacrine has emerged as a very promising anti-Alzheimer drug
candidate for its wide range of neuroprotective effects against a variety of neuronal injury models
(Du, & Carlier, 2004).
The success of the dual binding site strategy is evidenced by the large increase in AChE inhibitory potency of these dimers or hybrids relative to the parent compounds from which they have
been designed. Campiani et al. has reported that replacement of the central methylene group of
bis(7)-tacrine by a protonatable methylamino group and capable of providing additional specific
interactions with this mid-gorge recognition site leads to the dramatic increase in potency (Savini,
Gaeta, Fattorusso, Catalanotti, Campiani, Chiasserini, Pellerano, Novellino, McKissic, & Saxena,
2003). Analogously, the group of Valenti and Recanatini has developed a very interesting AChEI,
AP2238 (2, Figure 15), consisting of benzylamino and coumarin moieties as the dual binding units,
respectively, connected by a p-phenylene linker and able to interact with the aromatic residues at
the enzyme gorge. In addition to its high AChE inhibitory potency, AP2238 exhibits a significant
β-amyloid anti-aggregating action (Piazzi, Rampa, Bisi, Gobbi, Belluti, Cavalli, Bartolini, Andrisano, Valenti, & Recanatini, 2003). In order to improve the moderate potency of AP2238 toward
both human AChE (hAChE) and hAChE-induced Aβ aggregation, Piazzi et al. has recently reported
the synthesis of a large series of AP2238 derivatives (Piazzi, Cavalli, Belluti, Bisi, Gobbi, Rizzo,
Bartolini, Andrisano, Recanatini, & Rampa, 2007). The suitability of the o-methoxy-substitution
in AP2238 derivatives could be analogously related to its positive effect on the basicity of the
nitrogen atom of the N-methyl-substituted o-methoxybenzylamino moiety and consequently on its
extent of protonation.
The derivative of AP2248 leads to a higher potency after replacing the methyl group on the
nitrogen atom of the benzylamino moiety by an ethyl group AP2243 (3, Figure 15). Melchiorre
and Bolognesi developed a new dual binding site AChEI (4, Figure 15), composed of a unit of tacrine as the active site interacting moiety and the phenantridinium motif of propidium as the PAS
interacting moiety, connecting both units through a triamine linker (Bolognesi, Andrisano, Bartolini, Banzi, & Melchiorre, 2005). This hybridization resulted in a 20000- and 300-fold increase in
hAChE inhibitory activity relative to the parent compounds propidium and tacrine, respectively.
Neuropharma, developed a new family (5, Figure 15), of tacrine-based dual binding site AChEIs
which are the most potent inhibitors of AChE-induced Aβ aggregation and among the most potent
inhibitors of hAChE so far reported. Greenblatt et al. designed and reported the bifunctional derivatives (6, Figure 15) of the alkaloid galanthamine that interacted with both the active site and
PAS of the AChE (Greenblatt, Guillou, Guénard, Argaman, Botti, Badet, Thal, Silman, & Sussman, 2004). Bermúdez-Lugo et al. reviewed several computational methods for designing AChEIs,
such as docking approaches, molecular dynamics studies, quantum mechanical studies, electronic
properties, hindrance effects, partition coefficients (Log P) and molecular electrostatic potentials
surfaces, among other physicochemical methods that exhibit QSARs. (Bermúdez-Lugo, RosalesHernández, Deeb, Trujillo-Ferrara, & Correa-Basurto, (2011).
368
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use

QSAR Models towards Cholinesterase Inhibitors for the Treatment of Alzheimer’s Disease
Figure 15. Different dual binding site AChE inhibitors with Aβ anti-aggregating effects
Quantitative Structure-Activity Relationship
Quantitative structure-activity relationship (QSAR) and quantitative structure-property relationship
(QSPR) are based on the assumption that the structure of a compound (i.e. its geometric, steric and electronic properties) must contain features responsible for its physical, chemical, and biological properties
(Hansch, Hoekman, & Gao, 1996). More than a century ago, Crum-Brown and Fraser (1868) expressed
the idea that the physiological action of a substance in a certain biological system (Φ) was a function
(f) of its chemical constitution C:
Biological Activity (Φ) = f Physicochemical Properties(C)
QSAR/QSPR attempts to find out the quantitative correlation between physicochemical parameters of
compounds (independent variable) and their pharmacodynamic, pharmacokinetic, toxicological properties or biological activities (dependent variables) (Yap, Li, Ji, & Chen, 2007). Determining a QSAR
generally proceed as follows:
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
369

QSAR Models towards Cholinesterase Inhibitors for the Treatment of Alzheimer’s Disease
1. Define a quantitative measure of activity (e.g., the amount of ligand needed to produce an interfer-
ence with the functioning of the target).
2. Express the ligand in some quantitative manner; that is, select a collection of numbers that char-
acterize the ligand. These numbers are called ‘molecular descriptors’ or, ‘descriptors.’
3. Determine a functional relationship between activity and the selected descriptors; that is, the search
for a mathematical function, f, that has the property that “activity = f (descriptors)” to a suitably
high level of accuracy.
4. Use the know activity values, molecular descriptors and determined functional relationship to
predict the activity of new candidate compound.
QSAR methods are used to generalize experimental data in order to design or optimize new biologically
active compounds that are more potent, less toxic, more selective, or satisfy other relevant criteria (Winkler,
2002). Once, a reliable QSAR model is created, it is possible to predict the activity of new compounds,
and to learn which structural properties play an important role in the modeled biological response. It can
also identify and describe important structural features of the compounds that are relevant to variations
in molecular activities. The most widely used method for determining the functional relationship is the
statistical technique of regression or least squares. Various multi-dimensional descriptors, especially 2D
and 3D descriptors of molecular properties (particularly steric and electrostatic values) have been defined
that can be correlated with the biological activity using statistical or machine learning techniques (Li, Yap,
Ung, Xue, Li, Han, Lin, & Chen, 2007). An indispensable prerequisite for QSAR/QSPR model generation
is the set of compounds with reliable measured bioactivity, also known as ‘training set molecules’. The
model is validated by various methods and final model is used to predict the activity of new compounds
in the series. A number of automated computer programs are available for the fast generation of various
descriptors and QSAR models (Roy, 2007). Apart from drug discovery research the QSAR is being applied in many other disciplines. In fact a QSAR model can be built to predict any type of physical property
(Rücker, Meringer, & Kerber, 2004; Sutter, & Jurs, 1996) or biological activity/toxicity, of a given set of
observations and molecular structures (Novak, & Rajagopal, 2002; Gupta, 2007).
HISTORY OF QSAR
Historically, at the turn of the 20th century, Meyer (1899) and Overton (1895) independently suggested
that the narcotic (depressant) action of a group of organic compounds paralleled their olive oil/water
partition coefficients. In following years on the physical organic front, the seminal work of Hammett
gave rise to the “σ−ρ” culture (Hammett, 1935) in the delineation of substituent effects on organic reactions, while Taft devised a way for separating polar, steric, and resonance effects and introducing the
first steric parameter, ES (Taft Jr, 1952).
In 1962, Hansch and Fujita published their first 2D-QSAR study on the plant growth regulators and
their dependency on Hammett constants and hydrophobicity (Hansch, Maloney, Fujita, & Muir, 1962).
Using the octanol/water system, a whole series of partition coefficients were measured, and thus a new
hydrophobic scale was introduced (Hansch, & Fujita, 1964). The contributions of Hammett and Taft
together laid the basis for the development of the QSAR paradigm by Hansch and Fujita, which combined
the hydrophobic constants with Hammett’s electronic constants to yield the linear Hansch equation and
its many extended forms.
370
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use

QSAR Models towards Cholinesterase Inhibitors for the Treatment of Alzheimer’s Disease
Types of QSAR Models
QSAR is one of the basic tools of modern drug design science. It is founded on the systematic use of
mathematical models from the multivariate point of view. Further, it has an increasing role in environmental sciences. QSAR models exist at the intersection of chemistry, statistics and biology (toxicology)
studies. Data used in QSAR evaluations are obtained either from the literature or generated specifically
for QSAR analysis and should be accurate and precise. It can consist of congeneric series of compounds
or assure structural diversity even within a chemical class. This diversity has allowed the generalization
of more robust QSARs, applied in an extended way. A structure– activity model is defined and limited
by the nature and quality of the data used in model development and should be applied only within the
model’s applicability domain.
Nowadays, several multidimensional QSAR approaches (3D-QSAR, 4D-QSAR, 5D-QSAR and 6DQSAR) are also available depending upon the type of descriptors used (Lill, 2007). It is important to
note that all QSAR approaches aims to correlate structural/compounds data with the biological activity
or some other properties (Kubinyi, 1997; Dudek, Arodz, & Galvez, 2006). The 3D-QSAR methodology is much more computationally complex than the 2D-QSAR approach. In general, it involves several
steps to obtain different numerical descriptors of the compound structure. First, the conformation of
the compound has to be determined either from experimental data or molecular mechanics and then the
resulting conformations are refined by minimizing the energy (Akamatsu, 2002). Next, the conformers
in dataset have to be uniformly aligned in space. Finally, the space with immersed conformer is probed
computationally for various descriptors. Examples of 3D-QSAR are Comparative Molecular Field Analysis
(CoMFA) (Cramer Iii, Patterson, & Bunce, 1998), Comparative Molecular Similarity Indices (CoMSIA)
(Klebe, Abraham, & Mietzner, 1994) and Self Organizing Molecular Field Analysis (SOMFA).
Some methods independent of the compound alignment have also been developed. Alignment-independent 3D-QSAR descriptors are 3D descriptors that are invariant to molecular rotation and translation
in space. Thus, no superposition of compounds is required. The Comparative Molecular Moment Analysis
(CoMMA) (Silverman, & Daniel, 1996) uses second-order moments (moments relate to the center of the
mass and center of the dipole) of the mass distribution and charge distributions. The VolSurf approach is
based on probing the grid around the molecule with specific probes, e.g. the hydrophobic interactions or
hydrogen bond acceptor or donor groups (Crivori, Cruciani, Carrupt, & Testa, 2000). 3D-QSAR models
are widely applicable, but often suitable for any series of congeneric compounds.
The broad family of chemical descriptor(s) includes empirical, quantum chemical or non-empirical
parameters used in the 2D-QSAR approach share a common property of being independent from the 3D
orientation of the compound. Empirical descriptors may be measured or estimated and include physicochemical properties. Non-empirical descriptors can be based on individual atoms, substituents, or the
whole molecule, they are typical structural features. They can be based on topology or graph theory
and, as such, they are developed both from the knowledge of 2D structure, or from the 3D structural
conformations of a compound. Some of the different types of descriptors and their description are presented in Table 1.
Different software calculates wide sets of different theoretical descriptors, from SMILES, 2D-graphs
to 3D- coordinates etc. Some of the software includes: ADAPT (Stuper, & Jurs, 1976), OASIS (Mekenyan, & Bonchev, 1986), CODESSA (Katritzky, Petrukhin, Yang, & Karelson, 2005), MolConZ, and the
DRAGON (Todeschini, Consonni, Mauri, & Pavan, 2003). Thus, descriptor selection must be performed
by applying mathematical approaches with the final and crucial aim to maximize, as an optimization
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
371

QSAR Models towards Cholinesterase Inhibitors for the Treatment of Alzheimer’s Disease
Table 1. Types and description of various descriptors employed for obtaining QSAR models
Type Descriptors Description
Electronic Sum of partial charges, sum of formal charges,
dipole moment, HOMO, LUMO
Spatial Radius of gyration, shadow indices, area, density, PMI, Vm
Structural Molecular weight, number of rotatable bonds, number of hydrogen bond acceptor, number of
hydrogen bond donors
Thermodynamics Log of partition coefficient, log of partition coefficient atom-type value, desolvation free energy
of octanol, heat of formation, molar refractivity
Topological Wiener index, Zagreb index, Kier and Hall molecular connectivity index, Balaban index, Hosoya
index
E-State Indices Electrotopological-state indices
parameter, the predictive power of the QSAR model, as the real utility of any model is considered its predictivity. For the potency modelling, the most widely used mathematical technique is multiple regression
analysis (MRA). Regression analysis is a simple approach that leads to a result that is easy to understand.
For the modelling of categories, a wide range of classification methods exists, including: discriminant
analysis (DA), SIMCA (Soft Independent Modeling of Class Analogy), k-NN (k-Nearest Neighbours),
CART (Classification and Regression Tree), Artificial Neural Network (ANN), Support Vector Machine
(SVM), etc. In these techniques, the term “quantitative” is referring to the numerical value of the variables
(descriptors) necessary to classify the chemicals in the qualitative classes. Nowadays, genetic function
algorithm (GFA) has gained great popularity in QSAR research. GFA and genetic PLS (G/PLS) methods
developed by Rogers and Hopfinger (1994), employed to select the relevant descriptors.
Linear and Non-Linear Statistical Regression Techniques
• GFA: GFA is genetics based linear method of variable selection that combines Holland’s ge-
netic algorithm (GA) with Friedman’s (1991) multivariate adaptive regression splines (MARS).
It works in the following way: First, a predefined number of equations (set at 100 by default) are
generated randomly, then pairs of ‘‘parent’’ equations are chosen randomly from this set of 100
equations and the ‘‘crossover’’ operations are performed at random. After some preliminary observations on the initial runs, the number of GFA crossovers was set to 5000 for the present study
in order to obtain a reasonable convergence. The goodness of each progeny equation is assessed by
Friedman’s lack of fit (LOF) scoring, assigned by GFA, which resists over the fitting and estimates
the appropriate number of variables in the equation. The LOF is given by the equation,
LOF LSE c dp m= − +
/ /1
{
( )
where LSE is the least-squares error, c is the number of basis functions in the model, d is the smoothing
parameter, p is the number of descriptors, and m is the number of observations in the training set. The
smoothing parameter that controls the scoring bias between the equations of different size was set at
2
(1)
}
372
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use

Y f
( )
QSAR Models towards Cholinesterase Inhibitors for the Treatment of Alzheimer’s Disease
a default value of 1.0. A new term was added to the offspring models with a probability of 50%. Only
the linear equation terms were used for model building. The best equation out of 100 equations was
taken based on the statistical parameters such as regression coefficient, adjusted regression coefficient,
regression coefficient cross validation and F-test values.
• G/PLS: G/PLS methodology combines GFA, and is used for the selection of descriptors to be
included in a model. In the partial least squares (PLS) regression fitting technique (Roy, Roy,
2008) that weights the relative contribution of each descriptor in the final model. PLS in particular produces statistically significant solutions by identifying a linear combination of the original
physico-chemical descriptors that best correlates with the biological response (Dunn Iii, Scott,
& Glen, 1989). The main advantage of the G/PLS method is that it constructs QSAR equations
from which most variables have been eliminated; thus, PLS avoids over fitting. Application of G/
PLS thus allows the construction of larger QSAR equations, while still avoiding over fitting and
eliminating most variables.
• SVM: Support vector machine (SVM), first developed by Vapnik and Cortes (1995) is a novel
machine-learning non-linear method for classification problems. The main advantage of SVM is
that it was built on the structure risk minimization principle, which has been shown to be superior
to the traditional empirical risk minimization principle. After the introduction of ε-insensitive
loss function, SVM has been able to solve non-linear regression estimation and has demonstrated
much success in QSAR studies (Yuan, Zhang, & Luo, 2009). In brief, the correlation between the
structure of the training set and its biological activity, can be represented as Yi = f (x
we can represent f (x
) as a linear function of
i
). Simply,
i
y w x b
= + (2)
i i i
where,
wiis the coefficient vector of the linear function and b corresponds to its coefficient.
In SVM, the basic idea is to map the data x into a higher-dimensional feature space F through a non-
linear mapping ϕ by using kernel function and then performed linear regression in this space. In the high
dimensional feature space, SVM approximates the set of data with a linear function:
n
= =
where
x w x b
( )
∑
φ x
is the high-dimensional feature space after kernel transformation, which is nonlinearly
i
φ
+
( )
i i
=
1
i
(3)
mapped from the input space x, while wiand b are coefficients. A radial basis function (RBF) kernel is
frequently used kernel function in QSAR studies, out of various available kernel functions for nonlinear
transformation of the input data. The most important aim of SVM is minimization of the regularized
risk function, by which the coefficients w and b are estimated. The purpose of this risk function is to
find a function that has at most ε prediction errors in all the training data points and at the same time is
as flat as possible. The regularized risk function is defined as:
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
373

QSAR Models towards Cholinesterase Inhibitors for the Treatment of Alzheimer’s Disease
N
2
{ }
N
( )
22
N
R C C
( ) ( , ( , )) w= +
L y f x w y f x w for y f x wε ε ε, , , , ,
( )
1 1
( )
∑
=
i
= −
1
L y f x w
ε (4)
i i
( )
− −
2
≥
( )
0 otherwise
N
The first term of the equation
mated by the ε insensitive loss function
1
C
L y f x w
∑
1
=
i
L y f x wε , ,
ε( , ( , ))
i i
( )
. The term,
is the empirical error (risk), which is esti-
1
w is known as the regularized
term and is used as a measurement of function flatness. The parameter ε is known as the tube size, and
it is equivalent to the approximation accuracy placed on the training data points, and C is the regularization constant determining the trade-off between the training error and the regularized term. The minimization of regularized risk function is constrained optimization problem which can be reformulated
into dual problem formalism by using Lagrange multipliers. The performance of SVM regression is
based on the kernel selection, and the optimized value of C and ε. Parameter tuning must be performed
because, which values of C and ε will work the best for a particular data set is not known previously.
• ANN: Artificial neural network (ANN) method is an interconnected feed forward network, which
is motivated by the way biological nervous systems, viz. the brain process information, and is a
popular tool in function learning (Gasteiger & Zupan, 1993). It has the ability to learn complicated
nonlinear function (dataset) with good efficacy. The theory of the ANN and its application in
QSAR/QSPR studies were described extensively in many reviews (Zupan, 1994). An ANN technique in general has three interconnected layers: input layer, a hidden layer, and an output layer.
The input layer neurons receive the data from molecular descriptors input files, which is subsequently passed on to the nodes of the hidden layer for further processing. The signals are then
relayed onto the output layer, which sends information directly to the outside world, to a secondary computer process or to other devices such as a mechanical control system. The connections
between the nodes of each layer are assigned by randomized weight value. Supervised learning
technique was used in back-propagation ANN, the network is trained by minimizing the squared
error of the network’s output. This is achieved by back propagating the desired output to the transfer function (sigmoid function) for adjustment of weight by using a gradient descent algorithm.
QUALITY OF FIT AND QSAR MODEL VALIDATION METHODS
The success of a QSAR model is to measure the quality of fit on the available training data set. The
most common statistical qualities of the equations were measured by the parameters such as explained
2
variance (R
freedom (Wold, & Eriksson, 1995). While for G/PLS equations, least-squares error (LSE) was taken as
a statistical measure, lack-of-fit (LOF) was noted for the GFA derived equations. The expression for R
2
and the F-test are defined as mentioned below:
R
a
374
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
), squared correlation coefficient (R2) and variance ratio (F-test) at specified degrees of
a
2
,

N d
1
− −
r
( )
QSAR Models towards Cholinesterase Inhibitors for the Treatment of Alzheimer’s Disease
n
Y Y
( )
2
R
2
R
=
a
F
=
(( ) / ( ))
∑
1
i
1= −
( )
N R d
1 1
− − −
=
n
Y Y
−
( )
∑
− − −
R N d
i Training
1
i
=
2
1 1
2
R d
( / )
2
In the above Equations (5-7),
2
ˆ
−
i i
(5)
2
(6)
. (7)
ˆ
Y , i =1,…,n are the values calculated by the QSAR model for the
dependent variable corresponding to object ‘i’, when this compound has been included in the training
data set. Whereas,
Y
is the averaged value of the dependent variable for the training set, N is the
Training
number of compounds in the training set, and d is the number of descriptors in the QSAR equation.
As suggested by Tropsha, Gramatica, & Gombar, (2003) the predictive ability of a 2D-QSAR
model should be tested on an external set (test set) of data that has not been taken into account during
the process of developing the model. Hence, the best models were used to predict the enzyme inhibition
potency value of the test set compounds. The prediction qualities of the models were judged by statisti-
2
R
cal parameters like predictive correlation coefficient
actual and predicted values, with (
2
) and without (r
0
2
, squared correlation coefficient between
pred
) intercept of the test set compounds. The predic-
tive value is calculated as follows:
Y
( )
∑
R
2
pred
1= −
pred Test
Y Y
( )
∑
( ) ( )
Test Training
In the above equation, Y
the test set compounds, and
was previously shown that the use of
characteristics. Thus, an additional parameter
−
(Test) ( )
−
pred (Test)
Y
Training
( )
. (8)
2
and Y
indicate predicted and actual activity values, respectively, of
(Test)
indicates the mean activity value of the training set compounds. It
2
R
and r2 might not be sufficient to indicate the external validation
pred
2
r
defined as r r r r
m
2 2 202
m
1= − −
* which penalizes a
2
Y
model for large differences between the actual and the predicted values was also calculated (Roy, 2007).
QSAR Model Cross-Validation
Cross-validation is a popular technique which is used to explore the predictive ability of the QSAR
models. These models were rigorously evaluated using leave one out (LOO) or leave many out (LMO)
approach, Y-randomization test and test set predictions. In the LOO or LMO approach, one compound
375
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
Соседние файлы в папке Библиотека им академика М.И. Перельмана
