Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5587_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
34 Мб
Скачать
An Introduction to the Basic Concepts in QSAR-Aided Drug Design
Figure 2. Different names and chemical representation as line notations for structure diagram of phe­nylalanine
Figure 3. Molfile format for phenylalanine as an illustrative example of connection table style repre­sentation of molecules
6
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
An Introduction to the Basic Concepts in QSAR-Aided Drug Design
Figure 4. Molecular graph representation for phenylalanine
Although 2D representation of molecules provides chemical information, but in reality, molecules are 3D objects, and their atoms occupy distinct positions in a 3D space identified by the Cartesian co­ordinates. Some structural properties are obtained from 3D structures. The 3D structures are generated through empirical and numerical methods. X-ray crystallography, 2D NMR, IR, electron diffraction, and microwave spectroscopy are examples for experimental 3D structure determination. In numerical methods, quantum mechanics (QM)/molecular mechanics (MM) calculations are used for representing conformers of molecules. The 3D structure of a molecule is used for many different purposes, such as calculation of molecular descriptors, performing docking experiments, identification of pharmacophore, and building pharmacophore based 3D queries for similarity search in databases.
MOLECULAR DESCRIPTORS
Molecular descriptors (or sometimes called molecular parameters), the quantitative depictions of molecular structures, are crucial part of any QSAR study and encode certain structural features of molecules. The nature of descriptors varies according to the molecular properties. One may roughly classify descriptors as physico­chemical parameters (hydrophobic, electronic and steric properties), geometrical descriptors (molecular surface area calculation), constitutional descriptors (frequency of occurrence of substructure), electronic descriptors (molecular orbital calculation), and topological descriptors (connectivity index), however, as it is obvious they are to some extent interrelated. From dimensionality point of view, molecular descriptors are classified as one-dimensional (1D) representing bulk properties of compounds, two-dimensional (2D) topological and charge indices, three-dimensional (3D) descriptors indicating conformational aspects of a molecule. Param­eters are determined for a fragment of a molecule (called fragment-based parameters/substituent constants) or calculated for a molecule as a whole (global descriptors) (Dutt & Madan, 2012; Estrada & Molina, 2001). In the following section these parameters are described briefly.
Fragment-Based Descriptors
Early fragment-based parameters were originated from physical organic chemistry studies where Hammett equation was developed based on rate and equilibrium constant of reaction of meta- and para- substituted benzoic acid (Hammett, 1937). These parameters refer to a part of molecule. In other words, they are resulted from the various functional groups on a fixed core structure of a homologous series. Therefore, different substitutions on molecules have different impact on biological activities. For each substituent, a constant is defined as substitution constant based on the corresponding properties. Electronic, hydro­phobic, and steric aspects of molecules are important properties that can be used for determination of different substituent constants in a series of compounds.
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
7
An Introduction to the Basic Concepts in QSAR-Aided Drug Design
σ
π
π
v
Electron density distribution in a molecule can be influenced by different substituents leading to change in biological activity of the compounds. Electronic constants indicate the polarity of the sub­stituents. A good example of electronic constant is Hammett substituent constant (
), as stated above. It measures the substituent’s electron withdrawing/donating ability once substituted on benzoic acid defined by the following equation:
K
= log( ) (1)
σ
X
K
X
H
In this equation,
KX and KH indicate the acidic dissociation constants of the substituted and unsub-
stituted benzoic acid in aqueous solution. Electron withdrawing substituents will have positive values and vice versa.
Since, the lipophilicity of the molecules is of great importance in receptor interaction and pharma­cokinetic behavior of drug candidates, therefore, the other substituent parameter derived using a similar concept to that of Hammett constant is hydrophobic constant (
). This parameter shows hydrophobic-
ity of a substituent relative to hydrogen. It can be represented as follows:
π
= log logΡ Ρ (2)
X RX RH
where philicity. Positive value of
RX and RH denote the substituted and parent molecules, respectively; and Ρ shows the lipo-
for a substituent shows that this substituent is more hydrophobic than
hydrogen whereas the less hydrophobic substituent is characterized by negative value.
Sometimes, the size and shapes of substituents are dominant factors in biological activity of com­pounds. Quantification of steric substituent constants is accomplished by various steric parameters. Taft’s
steric parameter ( It is based on acid catalyzed ester hydrolysis of
Es) was the first parameter describing the steric effect in QSAR studies (Taft, 1952).
RCOOR ' vs. CH COOR3' shown in the following
equation:
k
E
= log( ) (3)
s
R
A
k
Me
where substituent parameter was Charton’s steric constant (
kR and kMe are the rate constants of hydrolysis for RCOOR ' and CH COOR3' . The other steric
) based on minimum van der Waals radius of a substituent. STERIMOL parameters were the other important steric features developed by Verloop (Verloop, 1987). These types of parameters are defined by spatial arrangement of a substituent in a molecule and obtained from distance-based measurement of the van der Waals and covalent radii along with standard bond angles and lengths. For more comprehensive information see (Todeschini & Con­sonni, 2000; Verloop, 1987).
There are numerous examples in QSAR studies which are based on the fragment-based descriptors. In 2009, Du et al. developed fragment-based (FB) QSAR in which molecular scaffold of studied com­pounds was divided into several fragments with respect to substituents. Then, the biological activities
8
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
π
An Introduction to the Basic Concepts in QSAR-Aided Drug Design
were correlated with physicochemical properties of the molecular fragments using linear free energy equation by obtaining two sets of coefficients, one for weighting factors and the other for physicochemi­cal properties of molecular fragments (Du et al., 2009).
Katritzky and colleagues proposed models for rate of skin permeability for a set of diverse compounds. In this study, substructural molecular fragments (SMF) were generated by splitting the molecular graphs to subgraphs for calculating their contribution for property of interest. Two different types of fragments were utilized including “sequences” and “augmented” atoms. The former one is related to sequences of atoms and bonds while the latter one represented a selected atom with its environment containing neighboring atoms and bonds with taking into the account the hybridization of the augmented atom (Katritzky, Dobchev, et al., 2006). The same approach was conducted for prediction of blood/tissue–air partition coefficient and blood–brain penetration for a set of organic compounds and drug molecules, respectively, using structural fragmental descriptors (Katritzky et al., 2005; Katritzky, Kuanar, et al., 2006).
Whole Molecule Descriptors
As mentioned earlier, another way for calculating the molecular descriptors is achieved by measuring or calculating different properties for a molecule and relating them to its whole structure, and not just to a fragment of it. Some of the important whole molecular descriptors are described briefly below.
Lipophilicity Parameters
Lipophilicity or hydrophobicity is one of the most attractive parameters in QSAR studies. This physi­cochemical property plays pivotal role in drug solubility, absorption (membrane permeation), distribu­tion, and receptor interactions. It is commonly obtained from the logarithm of the partition coefficient of the compound between n-octanol and water represented by LogP. The most routine experimental methods for measuring the LogP are shake-flask method and liquid chromatography method. It can also
5
be calculated by different methodologies implemented in software such as DRAGON
, and ACD/Labs2 suite of programs. It is an additive constitutive parameter as proposed by Hansch (Hansch, et al., 1962), who demonstrated this additive nature by summing up the
substituents for the fragments of a molecule to calculate its lipophilicity in the form of LogP value. The lipophilicity contribution of different groups and substituents in a molecule has also been represented by Rekker’s equation (Rekker & Mannhold,
1992):
log Ρ =∑a f
In this equation,
(4)
i i
ai and fi indicate the number of occurrences of the molecular fragment and lipo-
philicity contribution, respectively.
Topological Descriptors
As the name implies, these descriptors/indices are derived from molecular topology i.e. molecular connectivity. They are conformationally independent and do not need the coordinates of atoms. Useful information about size, shape, degree of branching, presence of heteroatoms, and multiple bonds in a
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
9
An Introduction to the Basic Concepts in QSAR-Aided Drug Design
m
molecule are provided by these types of descriptors (Gozalbes, Doucet, & Derouin, 2002; Hall & Kier, 2001; Katritzky, Lobanov, & Karelson, 1995; Leach & Gillet, 2007; Randic, 2001). The Wiener index was the first topological descriptor developed which counts the number of bonds between pairs of atoms and sums the corresponding distance between all pairs (Wiener, 1947). The well known molecular con­nectivity indices are Randic index, and Kier & Hall index. The Randic index (RI) shows the branching index of a molecule. It is based on hydrogen suppressed graph representation of a molecule and defines a “degree” of adjacent nodes linked together by edges leading to calculation of bond connectivity value as follows:
RI
=
where
allbounds
1
(5)
1 2
mn
( )
and n denote the degree of atoms connected by bonds. Figure 5 shows an example of a con-
nectivity index calculation for 3,4-dimethylhexane.
In the theory of Kier and Hall, the Randic index was modified and the number of hydrogens associated with atoms as well as valence electrons (sigma, pi and lone pair) were introduced (Kier & Hall, 1986). Subsequently, they designed the Kappa indices which are related to the molecular shape. These types of descriptors are not dependent on molecular geometry and obtained from possible “extreme shapes” for that number of atoms. For more detailed information see (Hall & Kier, 1991; Leach & Gillet, 2007).
Electrostatic Descriptors
A wide variety of descriptors can elucidate the electronic feature of the molecules which can be an important factor in determining the biological activity by affecting the strength of intermolecular inter­actions such as ion-ion, dipole-dipole, ion-dipole interactions as well as hydrogen bonding. Generally, these parameters reflect the polarity and ionization property of the molecule. The ability and tendency
of a molecule for ionization can be exhibited via ionization constants (e.g.,
pKa) which can provide
Figure 5. Illustrative example of connectivity index calculation for 3,4-dimethylhexane
10
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
n
MW
XY, YZ
An Introduction to the Basic Concepts in QSAR-Aided Drug Design
some useful information about the extent of molecular ionization. This in turn influences the absorption, distribution, and excretion of the molecule. Another measure of polarity can be interpreted in terms of molar refractivity and molecular polarizability parameters which are highly interrelated (these param­eters can be classified as steric parameters as well). Molar refractivity (
2
MR
=
n
n
MW
1
2
+
2
(6)
d
MR ) can be obtained by:
In this formula,
spectively. The positive value of
denotes the refractivity index,
MR in a QSAR model implies binding of a molecule to a polar surface
and d are molecular weight and density, re-
whereas a negative value suggests an steric hindrance for binding. Refractivity index, itself accounts for polarizability of the molecule.
Quantum-chemical descriptors can most likely be classified in the category of electronic descriptors
by providing valuable information about internal electronic feature of molecules. They are obtained by molecular orbital (MO) calculations. The most commonly used types of these descriptors are
E
, energies of highest occupied and lowest unoccupied MOs, respectively. The HOMO energy shows
LUMO
E
HOMO
and
the ionization ability of the molecule and losing a pair of electrons to an electrophile while the LUMO energy is related to the magnitude of electron affinity of a molecule by accepting a pair of electrons from a nucleophile. Therefore, these descriptors determine the position and possibility of electrophilic and nu­cleophilic attraction of a molecule in drug-receptor interaction (Dastmalchi, Hamzeh-Mivehroud, Ghafou­rian, & Hamzeiy, 2008; Katritzky, et al., 1995). Dipole moment is the other parameter calculated from quantum mechanical calculations that estimates the polarity of the molecule and indicates the polar-type interactions between a molecule and its target (Karelson, Lobanov, & Katritzky, 1996).
Geometric Descriptors
Biological effect of a compound is the consequence of interaction between the molecule and its target. For a desired interaction, the degree of complementarity of compound and receptor is essential. Three­dimensional feature of a molecule structure can be represented by geometric descriptors. The atomic coordinates enced by molecular conformation. Therefore, size, shape, molecular conformation, and spatial arrange­ment of the atoms are important properties determined in geometric descriptors. Molecular volume, molecular surface area, shadow indices, and shape parameters are few examples of such descriptors (Katritzky & Gordeeva, 1993; Katritzky, et al., 1995).
Since the size of molecules is an important determining factor in drug-receptor interaction, molecu­lar volume is regarded as a size related parameter which can be obtained from van der Waals volumes (Dudek, Arodz, & Galvez, 2006; Higo & Go, 1989) and used in QSAR studies. Molecular surface area is defined as amount of surface that is exposed to the solvent and is useful in prediction of drug transport, distribution, and absorption. Shadow parameters are calculated by projection of the molecular surface to its principal plane axes (i.e. Consonni, 2000). The ratio of longest to shortest three dimensions of a rectangle having the shadows of molecule assuming van der Waals radii and standard bond lengths is defined as shape parameter (η) (Katritzky & Gordeeva, 1993; Todeschini & Consonni, 2000).
( , , )x y z of molecules are required for obtaining these descriptors and hence, they are influ-
, and XZ) using van der Waals radii for atoms (Todeschini &
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
11
An Introduction to the Basic Concepts in QSAR-Aided Drug Design
Constitutional Descriptors
These types of descriptors give mainly simple description of molecular composition without considering topologic, geometric, and electronic aspects of molecule. Molecular weight, number of atoms, atoms of different identity, bonds, functional groups, rings, and H-bond donors/acceptors are few examples. They
5
are easily calculated in different programs like DRAGON
(Todeschini & Consonni, 2000). Although very simple, but when the feature being predicted shows significant correlation, these kinds of descrip­tors can appear in QSAR models.
DESCRIPTOR SELECTION METHODS
As discussed previously, the first stage for any QSAR study is to collect the data related to the molecular physicochemical properties. The collected data will include both the low-quality and high-quality data which can lead to low and high quality QSAR, respectively. However, the molecular descriptor selection from many existing types of molecular features is regarded as an important task in QSAR studies. For the purpose of removing low-quality or uninformative descriptors from the pool of molecular descriptors, many algorithms have been proposed to reduce the number of descriptors each representing a particular molecular feature. By omitting the redundant data, the next stage of QSAR study which is the building of QSAR model, can be carried out without interfering of noisy data and useless biases. Then, the pre- and post-processing techniques should be used to do the feature reduction and selection steps. The basic techniques for feature reduction in QSAR model building are partial least square (PLS) (Esposito Vinzi, 2010) and principle component analysis (PCA) (Ferraty & Romain, 2011; Jolliffe, 2002). The latter is a mathematical tool commonly used in cheminformatics to reduce descriptors from a high dimensional data space to a lower dimensional dataset known as orthogonal components by obtaining the transformed features which is generally a linear combina­tion of the original descriptors. The resulted components are then sorted in descending order that finally, the component with the highest variance is selected as the first principal component to explain the most of the characteristics of the original data. In contrast to PCA, partial least squares, as a statistical method, constructs the components by maximizing the covariance between the response variable and a linear combination of predicting descriptors. However, the size of the reduced feature space achieved using PLS and the selected components are critical for classifiers’ generalization problems (Li & Zeng, 2009). Recently, some of QSAR studies have been conducted based on PLS or PCA as a fundamental method along with other algorithms to improve the quality of the descriptor selection step. For example, ordered predictor selection (OPS)-PLS has been used in a QSAR study for the selection of molecular descriptors important for antitrypanosomal activity of a series of thiosemicarbazones (Lozano et al., 2012).
The successive projection algorithm (SPA) is a variable selection tool consisting of two or more steps that can be used along with the linear discriminant analysis (LDA) (Goodarzi, Saeys, Araujo, Galvao, & Vander Heyden, 2013; Zhang, Rivard, & Rogge, 2008). By performing the SPA method, first, it identifies the most outstanding features and then the next distinct feature will be selected iteratively using the orthogonal projection to minimize the cost function by calculating the Mahalanobis distance values of molecules’ features accord­ing to their true classification and misclassification. Gram-Schmidt orthogonalization (GSO) is also another pre-processing step for solving the collinearity problem to omit less informative descriptors by constructing a regression model between an starting descriptor and other descriptors to maximize the residuals as shown for 107 compounds with anti-HIV activity (Bagheri, Omidikia, & Kompany-Zareh, 2013).
12
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
An Introduction to the Basic Concepts in QSAR-Aided Drug Design
Genetic algorithms (GA) known as evolutionary computation method inspired by gene duplication comprises five steps that is conducted iteratively to reduce the number of descriptors efficiently by par­allel searching in a large scale data. These steps include the initialization (setting the population size), selection (a set of selection algorithms to select the features or descriptors that can be reproduced for the next generation), crossover (using numerous cross over algorithms such as 1-point, 2-point and uniform types for combination of selected descriptors to generate the offsprings based on probability P
), muta-
c
tion (this operator is generally applied after the cross over operator to mutate the offspring based on the probability P
), and stopping criteria (the fitness function is needed to stop the iterative operation and
m
converge the population globally) (Aguiar-Pulido et al., 2013). Finally, in the bit based GA, a “bit” trans­lation of the existing features in QSAR study can be defined as active and inactive represented by “1” and “0” in a chromosome, respectively. GA has been implemented successfully in many QSAR studies (Goodarzi, Saeys, Deeb, Pieters, & Vander Heyden, 2013; Noorizadeh, Sajjadifar, & Farmany, 2013).
Simulated annealing (SA) (Oliveira Junior, 2012) is a probabilistic metaheuristic method for optimiz­ing functions implemented in QSAR studies. This method is also based on evolutionary computations, which globally and iteratively searches in a large scale pool of descriptors to find the optimum descriptors with the least error function. Moreover, based on the probability values of Boltzmann distribution, the worst descriptor can also be selected which makes it hard for SA to be trapped in the local minima of error function. Ghosh and his colleagues have used SA effectively for descriptor selection in quinoxaline derivatives acting as anti-tuberculars (Ghosh & Bagchi, 2009).
Sequential feature forward selection (SFFS) (Sun, Todorovic, & Goodison, 2010) is a kind of clas­sification tool which searches locally through the descriptors pool by using a sequential decision process. The SFFS method is used to assess the trained QSAR model based on the selected descriptors achiev­ing a minimum error. At the first step, it is supposed that the best and the most informative descriptor is determined and selected. Then, the remaining descriptors will be used to evaluate the trained QSAR model based on the selected descriptors and the resulted minimum error. Finally, the descriptors with less error will be selected. As an example, a ligand-based virtual high-throughput screening with the
6
PubChem
database has been carried out using the SFFS for assessing the problem-specific descriptor
optimization protocols (Butkiewicz et al., 2013).
Sequential backward feature elimination (SBFE) (Sun, et al., 2010) is a type of sequential feature selection which acts in the reverse order to that of SFFS method. Despite the sequential forward feature selection which is based on adding features with minimum error, the sequential backward feature elimi­nation is based on removing descriptors with maximum errors which is more time consuming. It selects the worst and the most non-informative descriptors for being eliminated from the pool of descriptors. A novel method to overcome this disadvantage has also been presented by Guyon et al. which uses a recursive feature elimination based on support vector machines (SVMs) (Guyon, Weston, & Barnhill,
2002). Li et al. (Li, Cong, Yang, Xue, & Chen, 2013) have optimized their descriptors subset by recur­sive feature elimination method in an in silico study carried out on the spleen tyrosine kinase inhibitors.
Gravitational search algorithm (GSA) inspired from the Newton’s gravity force, makes two massive objects move toward each other as any particle in the space absorbs the other particle (Bababadani & Mousavi, 2013). In this study, the binary based GSA is proposed which begins by generating a ran­domized binary agent as described in GA (i.e., an agent is a string of predefined variables which is a representative of absence or existence of descriptors and has a mass value calculated based on a fitness evaluation function). The GSA method has been employed for descriptor selection for anticancer potency modeling of a set of imidazo[4,5-b]pyridine derivatives in which the results showed that GSA outper-
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
13
An Introduction to the Basic Concepts in QSAR-Aided Drug Design
forms the GA (Bababadani & Mousavi, 2013). However, this technique suffers from over-fitting problem and should be stopped to achieve the generalization property by using an algorithm such as regularization method.
The replacement method (RM) and its variations, with the property of not using the full search of descriptors, have been proved to achieve the optimal descriptors with a small number of linear regressions (Mercader, Duchowicz, Fernandez, & Castro, 2011; M. Sun et al., 2009). Moreover, the enhanced modified versions of RM resemble the simulated annealing techniques (Mercader, Duchowicz, Fernández, & Castro, 2008). The overall aim of RM is to choose an optimum set of descriptors with the minimum standard deviation which enables RM to find the global minimum. The RM method starts by randomly choosing an initial descriptor set and keeps the set with the minimum standard deviation. Then, all the descriptors with large standard deviation values will be chosen and substituted by remaining descriptors with small standard deviation. The process will be continued until the descriptors with minimum standard deviation are retained. However, in modified and enhanced editions of RM, this type of replacement is not performed (Sun, Zheng, Wei, Chen, Cai, et al., 2009). Some of QSAR studies related to the RM and its variants can be found in the literature (Mercader, Duchowicz, Fernandez, & Castro, 2010; Mercader, et al., 2011; Mercader, et al., 2008).
Cluster analysis is also another statistical method that separates the descriptors into individual classes using clustering algorithms such as k-means clustering method (Gonzalez, Teran, Saiz-Urra, & Teijeira, 2008). Most of QSAR studies are using the Fisher’s discriminant ratio as their separation criterion (Lin, Li, & Tsai, 2004). After the criterion value is applied, the values of descriptors will be scaled in to a specified range of values (i.e., in most of the cases between 0 and 1). Then, the Fisher’s discriminant ratio will be computed for all descriptors to choose the descriptors with the highest Fisher’s discriminant ratio which will then be used to compute the cross-correlation coef­ficient between the descriptors with largest ratio and remaining descriptors. Cluster analysis method has been applied in some QSAR studies. For example, Prakash and Khan have used this method for identification of anticancer leads against human breast cancer (Prakash & Khan, 2013). QSAR studies on Pseudomonas aeruginosa deacetylase LpxC inhibitors, and HIV-1 integrase inhibitors are other examples (Kadam & Roy, 2006; Yuan & Parrill, 2005)
Ant colony optimization (ACO) algorithm, an stochastic optimization method, is generally based on the behavior of ants in selecting a path for finding the food which has the most pheromone with high probability rather than traveling randomly (Bonabeau, Dorigo, & Theraulaz, 2000). Some modified versions of ACO have also been proposed by authors to be applied to QSAR studies by simulating the ant colony systems. However, ACO is a time consuming process and like the linear methods suffers from trapping in local optima (Gonzalez, et al., 2008). The descriptor selection based on the ACO method is generally carried out in a binary format (i.e., 1 and 0) in which for each variable, the existence and the absence of a descriptor is presented by 1 and 0, respectively. The amount of pheromone on each path will be updated according to an updating rule. Finally, ant selects a descriptor as its path based on the pheromone amount. The ACO process is iteratively done until a minimum error criterion is reached. However, the reported modified ACO techniques have advantages over the standard ACO on convergence speed and using fewer numbers of parameters (Gonzalez, et al., 2008). Modified ant colony optimization algorithm and its hybrid along with mutiple linear regression (MLR) have been used for variable selection in QSAR studies of cyclooxygenase inhibitors and tyrosine kinase inhibitor, respectively (Shamsipur, Zare-Shahabadi, Hemmateenejad, & Akhond, 2009; Shen, Jiang, Tao, Shen, & Yu, 2005; Shi, Shen, Kong, & Ye, 2007).
14
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
An Introduction to the Basic Concepts in QSAR-Aided Drug Design
QSAR MODEL BUILDING METHODS
When in the preceding step, various sets of proper descriptors have been selected based on their corre­sponding molecular nature, different QSAR models can be generated using linear or non-linear methods (Aguiar-Pulido, et al., 2013; Doo Ho, Sung Kwang, Bum Tae, & Kyoung Tai, 2001). Although QSAR models are used for the prediction of the desired properties, however the predictions may sometimes fail to pass validation criteria even if a good training set has been chosen (Aguiar-Pulido, et al., 2013). QSAR modeling has become one of the most interesting research subjects because it may reduce the problems associated with the expensive synthetic methods, biological assays, and failure of drug developments due to poor ADMET profiles (Verma, Khedkar, & Coutinho, 2010). For this purpose, researchers have developed many QSAR software (Liew & Yap, 2012) ranging from structure drawing (e.g., ChemDraw
8
ACD/ChemSketch
13
smi23d tor
23
R
), descriptor calculation (e.g., ADRIANA Code14, Dragon5, Molconn-Z15 and PaDEL-Descrip-
16
), and modeling (e.g., KNIME17, RapidMiner18, WEKA19, Orange20, TANAGRA21, MATLAB22 and
) to multipurpose molecular modeling software packages including SYBYL24, Discovery Studio25,
MOE (Molecular Operating Environment)
, OpenBabel9), 3D structure generation (e.g, CORINA10, Concord11, Frog12, and
26
, and CODESSA27 developed by commercial vendors and academia. The QSAR modeling methodologies are divided into linear and non-linear methods on datasets with varying size, although some methods may be preferred for small datasets over the large datasets and vice versa (Aguiar-Pulido, et al., 2013).
7
,
Linear Methods
Using the linear methods implies that the function of QSAR model is comprised of linear combination of the molecular descriptors by using linear algorithms such as multiple linear regression (MLR), partial least square (PLS) or linear discriminant analysis (LDA) (Aguiar-Pulido, et al., 2013). By using MLR method, as one of the earliest methods used in building QSAR models, the linear function of biological activity is predicted based on a model with multiple independent descriptors by minimizing the squared
deviations of the errors of the regression equation or maximizing the available parameters (Afantitis et al., 2006). To reduce the correlations between the independent descrip­tors and/or eliminate the chance correlation, the number of independent descriptors are considered to be less than one-fifth the number of compounds in the training set, otherwise, multiple linear regression method cannot be applied (Livingstone & Salt, 2005). An improved version of MLR known as Best MLR (BMLR) is used in various QSAR studies (Du, Wang, Hu, & Yao, 2008; Du, Zhang, Wang, Yao, & Hu, 2008) which overcomes most of the shortcomings of MLR method and makes it working well on training set less than five times the number of molecular descriptors (Liu & Long, 2009). Another ad­vanced MLR method for linear QSAR model construction is the Heuristic Method (HM) (Katritzky, Lobanov, & Karelson, 2007) which can also be used to select descriptors for a non-linear model (Xia et al., 2009). Some advantages of HM are its high speed, no restrictions on the size of data, and good es­timation of quality of correlations (Liu & Long, 2009). The overall descriptor selection in HM is com­prised of ensuring about the availability of descriptor values in each structure and discarding the ones which are not available and/or significant for a specific structure, then calculating the pair correlation matrix for descriptors in order to keep only one of the highly correlated pairs of descriptors, and at the end gives estimates for the error of regression, cross-validated correlation coefficient, the F-test statistics, and the standard deviation used for evaluating the quality of correlation (Katritzky, et al., 2007).
R2value by adjusting each of the
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
15