Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5865_Библиотеки_им_академика_М_И_Перельмана.pdf

Quantitative Structure-Activity/Property/Toxicity Relationships
Vela, A., & Gázquez, J. L. (1990). A relationship between the static dipole polarizability, the global
softness, and the Fukui function. Journal of the American Chemical Society, 112(4), 1490–1492.
doi:10.1021/ja00160a029
Vijayaraj, R., Subramanian, V., & Chattaraj, P. K. (2009). Comparison of global reactivity descriptors
calculated using various density functionals: A QSAR perspective. Journal of Chemical Theory and
Computation, 5(10), 2744–2753. doi:10.1021/ct900347f
Voskresensky, O. N., & Levitsky, A. P. (2002). QSAR aspects of flavonoids as a plentiful source of new
drugs. Current Medicinal Chemistry, 9(14), 1367–1383. doi:10.2174/0929867023369790 PMID:12132993
Wan, J., Zhang, L., Yang, G., & Zhan, C.-G. (2004). Quantitative structure-activity relationship for cyclic
imide derivatives of protoporphyrinogen oxidase inhibitors: A study of quantum chemical descriptors from
density functional theory. Journal of Chemical Information and Computer Sciences, 44(6), 2099–2105.
doi:10.1021/ci049793p PMID:15554680
Wiener, H. (1947). Structural determination of paraffin boiling points. Journal of the American Chemi-
cal Society, 69(1), 17–20. doi:10.1021/ja01193a005 PMID:20291038
Wolff, M. E. (1955). Therapeutic agents. In Burger’s medicinal chemistry and drug discovery (Vol. 4).
New York: John Wiley & Sons.
Yan, D., Jiang, X., Xu, S., Wang, L., Bian, Y., & Yu, G. (2008). Quantitative structure-toxicity relationship study of lethal concentration to tadpole (Bufo vulgaris formosus) for organophosphorous pesticides.
Chemosphere, 71(10), 1809–1815. doi:10.1016/j.chemosphere.2008.02.033 PMID:18395243
Yang, G.-F., & Huang, X. (2006). Development of quantitative structure-activity relationships
and its application in rational drug design. Current Pharmaceutical Design, 12(35), 4601–4611.
doi:10.2174/138161206779010431 PMID:17168765
Yang, W., & Mortier, W. J. (1986). The use of global and local molecular parameters for the analysis
of the gas-phase basicity of amines. Journal of the American Chemical Society, 108(19), 5708–5711.
doi:10.1021/ja00279a008 PMID:22175316
Yang, W., & Parr, R. G. (1985). Hardness, softness, and the fukui function in the electronic theory of
metals and catalysis. Proceedings of the National Academy of Sciences of the United States of America,
82(20), 6723–6726. doi:10.1073/pnas.82.20.6723 PMID:3863123
Yang, W., Parr, R. G., & Pucci, R. (1984). Electron density, Kohn− Sham frontier orbitals, and Fukui
functions. The Journal of Chemical Physics, 81(6), 2862–2863. doi:10.1063/1.447964
Zhang, L., Hao, G.-F., Tan, Y., Xi, Z., Huang, M.-Z., & Yang, G.-F. (2009). Bioactive conformation
analysis of cyclic imides as protoporphyrinogen oxidase inhibitor by combining DFT calculations,
QSAR and molecular dynamic simulations. Bioorganic & Medicinal Chemistry, 17(14), 4935–4942.
doi:10.1016/j.bmc.2009.06.003 PMID:19540767
Zhang, S. G., Lei, W., Xia, M. Z., & Wang, F. Y. (2005). QSAR study on N-containing corrosion inhibitors: Quantum chemical approach assisted by topological index. Journal of Molecular Structure
THEOCHEM, 732(1-3), 173–182. doi:10.1016/j.theochem.2005.02.091
176
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use

Quantitative Structure-Activity/Property/Toxicity Relationships
KEY TERMS AND DEFINITIONS
Aliphatic Compounds: These are non-aromatic compounds made by carbon and hydrogen.
Electrophilicity: It is a measure of ‘electrophilic power’ of a molecule. It can be estimated as the
electronegativity squared divided by twice the hardness.
Energies of Frontier Molecular Orbitals: They are the energies of the highest occupied molecular
orbital and lowest unoccupied molecular orbital, respectively.
Group Philicity: It is a condensed philicity summed over a group of relevant atoms.
Net Atomic Charge: It is a non-integer charge at each atom in a molecule in elementary charge units.
Net Electrophilicity: It is the electron-accepting power of a molecule relative to its own electron-
donating power.
Polyaromatic Hydrocarbons: These are compounds made by only carbon and hydrogen and contain
multiple aromatic rings.
Polychlorinated Biphenyls: These are compounds containing two phenyl rings and several chlorine
atoms are attached with phenyl rings.
Tetrahymena Pyriformis: These are free-living ciliate protozoa, generally found in freashwater ponds.
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
177

Quantitative Structure-Activity/Property/Toxicity Relationships
APPENDIX: LIST OF SYMBOLS AND ABBREVIATIONS
SAR: Structure-Activity Relationship
QSAR: Quantitative Structure-Activity Relationship
QSTR: Quantitative Structure-Toxicity Relationship
DFT: Density Functional Theory
CDFT: Conceptual Density Functional Theory
LSFER: Linear Solvation Free Energy Relationship
CoMFA: Comparative Molecular Field Analysis
QSPR: Quantitative Structure-Property Relationship
PCA: Principal Component Analysis
η: Hardness
χ: Electronegativity
ω: Electrophilicity
f ( )r : Fukui function
f
: Condensed Fukui function
k
s(r): Local softness
η(r): Local hardness
ω(r): Philicity
IP: Ionization Potential
EA: Electron Affinity
μ: Chemical Potential
E: Energy
N: Number of Electrons
v(r): External Potential
±
: Net Electrophilicity
Δω
HIV-1: Human Immunodficiency Virus type 1
NCp7: Nucleocapsid Protein p7
W: Electrical Power
V: Voltage
R: Electrical Resistance
α: Polarizability
ξ: Magnetizability
ρ(r): Electron Density
: Atomic Charge at kth site
q
k
α
: Group Philicity
ω
g
PMH: Principle of Maximum Hardness
MEP: Minimum Electrophilicity Principle
MPP: Minimum Polarizability Principle
MMP: Minimum Magnetizability Principle
: 50% Lethal Dose
LD
50
: 50% Lethal Concentration
LC
50
178
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use

Quantitative Structure-Activity/Property/Toxicity Relationships
Z: Atomic Number
: Number of Nonhydrogenic Atoms
N
NH
R or r: Coefficient of correlation
2
R
CV
or r
2
: Variance of Leave-one-out Cross-Validation
CV
PAH: Polyaromatic Hydrocarbons
TCDD: Tetrachlorodibenzo-p-Dioxin
Ah: Arylhydrocarbon
PCDF: Polychlorinated Dibenzofurans
CA: Chloroanilines
PCB: Polychlorinated Biphenyl
RBA: Relative Binding Affinity
TeBG: Testosterone-Binding Globulins
MAPT Myotrophic to Androgenic Potency in Temporal Activity.
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
179

180
Chapter 5
Importance of Applicability
Domain of QSAR Models
Kunal Roy
Jadavpur University, India
Supratik Kar
Jadavpur University, India
ABSTRACT
Quantitative Structure-Activity Relationship (QSAR) models have manifold applications in drug discovery,
environmental fate modeling, risk assessment, and property prediction of chemicals and pharmaceuticals.
One of the principles recommended by the Organization of Economic Co-operation and Development
(OECD) for model validation requires defining the Applicability Domain (AD) for QSAR models, which
allows one to estimate the uncertainty in the prediction of a compound based on how similar it is to the
training compounds, which are used in the model development. The AD is a significant tool to build a
reliable QSAR model, which is generally limited in use to query chemicals structurally similar to the
training compounds. Thus, characterization of interpolation space is significant in defining the AD. An
attempt is made in this chapter to address the important concepts and methodology of the AD as well as
criteria for estimating AD through training set interpolation in the descriptor space.
INTRODUCTION
Quantitative structure-activity relationship (QSAR) modelling has become an important tool in the new
drug candidate design, environmental fate modeling, toxicity and property prediction of chemicals and
pharmaceuticals since they offer an economical and time-effective alternative to the medium throughput
in vitro and low throughput in vivo assays (Perkins et al., 2003; Selassie, 2003; Walker et al., 2003). A
QSAR model is a simple mathematical equation that is evaluated from a set of molecules with known
activities/properties/toxicities using computational approaches. Thereafter, the developed predictive
QSAR models are also applied by regulatory agencies to evaluate physical, chemical, and biological
properties of individual chemical entities using applications specific for decision-making frameworks in
risk and safety assessments (Kar & Roy, 2010). The QSAR modeling hypothesis also supports the 3Rs
DOI: 10.4018/978-1-4666-8136-1.ch005
Copyright © 2015, IGI Global. Copying or distributing in print or electronic forms without written permission of IGI Global is prohibited.
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use

Importance of Applicability Domain of QSAR Models
(replacement, refinement and reduction in animals in research) paradigm due to an increased pressure
from social and economic background to trim down the use of animal testing as an important alternative
method for future prediction of untested chemical entities (Benigni & Giuliani, 2003).
Robust validation of QSAR models plays a key step for the selection of a predictive model that may
be considered for future prediction of new molecules. As numerous numbers of researches have been
directed to the design of new molecules with the utilization of QSAR technique, validation of a QSAR
model has been certified as the most considerable stride (Carlsen et al., 2009) for assessing the quality
of data, applicability and mechanistic interpretability of the developed model. Thus, a huge number of
investigations are currently directed towards the introduction of more appropriate validation approaches
for more accurate and predictive QSAR model development. One important objective of QSAR modeling
is to predict activity/property/toxicity of new chemical entities falling within the applicability domain of
the developed models. The reliability of any QSAR model depends on the confident predictions of these
new compounds based on applicability domain (AD) of the modeled compounds and therein arises the
importance of the AD study (Ferreira, 2001).
In order to establish the scientific validity of a QSAR model and to facilitate its acceptance for regulatory
purposes, the Organization for Economic Cooperation and Development (OECD) in its joint meeting (OECD,
2007) has agreed to five principles that should be followed during the construction of QSAR models. The
OECD Principle 3 which defines the need of an AD expresses the fact that QSARs are unavoidably associated
with restrictions in terms of the types of chemical structures, physicochemical properties and mechanisms of
action for which the models can generate reliable predictions. The applicability domain of a QSAR model has
been defined as the response and chemical structure space, characterized by the properties of the molecules in
the training set. The developed QSAR model can predict a new compound truly only when it falls within the
applicability domain of the developed model (Netzeva et al., 2005). Thus, to identify the interpolation (true
prediction) or extrapolation (less reliable prediction) of query compounds is an important task for a QSAR
model developer using the information of applicability domain (Netzeva et al., 2005).
Viewing the importance of AD in QSAR model validation, we herein focus to get an overview of different traditional as well as relatively new AD approaches used to judge the quality of the QSAR models.
This book chapter will be helpful for the QSAR learners in order to have a clear idea on the principles
of available AD approaches useful for judging the predictive quality of QSAR models.
BACKGROUND
A QSAR model is essentially valued in terms of its predictability, indicating how well it is able to predict
the endpoint values of the compounds which are not used to develop the correlation. The models that
have been suitably validated internally and externally can be considered reliable for both scientific and
regulatory purposes (Golbraikh & Tropsha, 2002; Tong et al., 2004]. QSAR models should be validated
according to the OECD principles for reliable prediction. A meeting of QSAR experts held in Setúbal,
Portugal in March 2002 formulated guidelines for the validation of QSAR models, in particular for
regulatory purposes (Jaworska et al., 2003). These principles were agreed by OECD member countries,
th
QSAR and regulatory communities at the 37
Party on Chemicals, Pesticides and Biotechnology in November 2004. These principles are best possible summary of the most important points that are required to be addressed to find consistent, reliable
and reproducible QSAR models (OECD, 2004). Five OECD principles to develop reliable models are:
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
Joint Meeting of the Chemicals Committee and Working
181

Importance of Applicability Domain of QSAR Models
Principle 1: A defined endpoint;
Principle 2: An unambiguous algorithm;
Principle 3: A defined domain of applicability;
Principle 4: Appropriate measures of goodness-of fit, robustness and predictivity;
Principle 5: A mechanistic interpretation, if possible (OECD, 2004).
In order to provide guidance on the interpretation of the mentioned principles, the OECD has also
provided a checklist (OECD, 2004). Thus, the current challenge in the process of development of a QSAR
model is no longer in developing a model that is statistically sound to predict the activity within the
training set, but in developing a model with the capability to predict accurately the activity of untested
chemicals. In spite of high fitted accuracy and apparent mechanistic appeal, some published QSAR
models fail rigorous validation tests, and, thus, may lack practical utility as reliable screening tools
(Tropsha et al., 2003]. In summary, predictive power is one of the most important validation features
of QSAR models. It can be defined as the ability of a model to predict accurately the target activity/
property/toxicity of compounds that are not used for model development.
In this context, QSAR model predictions are most reliable if they come from the model’s applicability
domain (AD) which is broadly defined under OECD principle 3. The OECD includes AD assessment as
one of the QSAR acceptance criteria for regulatory purposes (OECD, 2004; OECD, 2007). The Setubal
Workshop report (Jaworska et al., 2003) presented the following regulation for AD assessment: “The
applicability domain of a (Q)SAR is the physico-chemical, structural, or biological space, knowledge
or information on which the training set of the model has been developed, and for which it is applicable
to make predictions for new compounds. The applicability domain of a (Q)SAR should be described in
terms of the most relevant parameters, i.e., usually those that are descriptors of the model. Ideally the
(Q)SAR should only be used to make predictions within that domain by interpolation not extrapolation.
This depiction is useful for explaining the instinctive meaning of the “applicability domain” approach.
Two general principles to AD estimation have been proposed till date (Nikolova-Jeliazkova & Jaworska, 2005). The first one estimates the interpolation region based on the training set in the model’s
descriptor space. The second approach relies on similarity analysis according to the premise that a QSAR
prediction is reliable if the compound is “similar” to the compounds in the training set. It is interesting
to point out that “similarity” is a relative concept as different concepts of similarity are relevant to different endpoints. The similarity approach to AD assessment should be based on measurable key features
relevant to the endpoint modelled precisely (Bender & Glen, 2004; Nikolova & Jaworska, 2004).
APPLICABILITY DOMAIN (AD)
The AD (Eriksson et al., 2003; Gramatica, 2007; Netzeva et al., 2005, Tetko et al., 2006) is a theoretical
region in chemical space encompassing both the model descriptors and modeled response. In the development of a QSAR model, the domain of applicability of molecules plays a crucial role for estimating the
uncertainty in the prediction of a particular compound based on how similar it is to the compounds used
to build the model. Thus, the prediction of a modeled response using QSAR is valid only if the compound
being predicted falls within the domain of applicability of the model since it is impossible to predict an entire
universe of chemicals using a single QSAR model. Again, AD can be described as the physico-chemical,
182
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use

Importance of Applicability Domain of QSAR Models
structural or biological space, knowledge or information based on which the training set of the model is
developed, and the model is applicable to make predictions for new compounds within the specific domain
(Weaver & Paul Gleeson, 2008).
The selection method of the training and test sets has a significant effect on the applicability domain of
constructed QSAR model. Thus, while splitting a dataset for external validation, the training set molecules
should be selected in such a way that they span the entire chemical space for all the dataset molecules.
In order to achieve successful predictions, a QSAR model should always be used for compounds within
its applicability domain. Therefore, one can arguably agree that if a new compound falls outside of the
AD of the training set molecules, its prediction is not reliable.
APPLICABILTY DOMAIN APPROACHES
There are various AD approaches that have been used by different groups of QSAR researchers worldwide. The most common approaches for estimating interpolation regions in a multivariate space include
the following ones (Jaworska et al., 2005; Netzeva et al., 2005):
1. Ranges in the descriptor space.
2. Geometrical methods.
3. Distance-based methods.
4. Probability density distribution.
5. Range of the response variable.
The first four approaches are based on the methodology used for interpolation space characterization in
the model descriptor space. On the contrary, the last one depends solely on response space of the training
set molecules. A compound can be identified as out of the domain of applicability in a simple way, if:
1. At least one descriptor is out of range for the ranges approach, and
2. The distance between the chemical and the center of the training data set exceeds the threshold for
distance approaches.
The threshold for all kinds of distance methods is the largest distance between the training set data
points and the center of the training data set.
Along with the common approaches, Stanforth et al., (2007) proposed a cluster-based approach to
evaluate the AD of any QSAR model. The method applies an intelligent version of the k-means clustering algorithm for modeling the training set as a compilation of clusters in the descriptors space. On the
other hand, the test compounds of individual clusters are assigned a fuzzy membership from which an
overall distance may be calculated to identify the AD of individual test compounds. A classificationbased process involves calculation of regression residuals for identification of good and bad classes
suggested by Guha & Jurs (2005).
All approaches are discussed thoroughly in this section enlightening their main theories to define the
interpolation space as well as the threshold criteria for AD used in QSAR modeling. A summary of hypotheses
of data distribution and interpolation space of most commonly used AD methods is discussed in Table 1.
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
183

Importance of Applicability Domain of QSAR Models
1. RANGES IN THE DESCRIPTOR SPACE APPROACH
1.1 Fixed or Probabilistic Boundaries/Bounding Box
• Theory: The ranges of the individual descriptors are considered with a uniform distribution de-
fining an n-dimensional hyper-rectangle developed on the basis of the maximum and minimum
values of each descriptor used to build the model with sides parallel to the coordinate axes.
• Criteria: This method is simply based on boundary created by individual descriptor ranges of the
training set compounds from which the QSAR model is developed. As mentioned earlier, this is
based on the highest and lowest values of X variables (descriptors to develop the QSAR model)
and Y variable (response for which the QSAR equation is created) of the training set. Any test set
compounds, which are not present in any of these particular ranges, are considered out of the AD
and their predictions are less reliable (Jaworska et al., 2005; Seber, 1984).
• Drawbacks:
◦ For non-uniformly distributed data, the approach encloses considerable empty space,
◦ Empty regions in the interpolation space cannot be recognized, as only the descriptor ranges
are considered, and
◦ Correlation between descriptors cannot be taken into account (Jaworska et al., 2005).
For the ‘bounding box’ method, the domain of applicability is the smallest axis-aligned rectangular box
containing all the data points presented in Figure 1.
1.2 Principal Component Ranges/PCA Bounding Box
• Theory: Principal components convert the original data into a new orthogonal coordinate system
by the rotation of axes and enable to correct for correlations among descriptors. Newly formed
axes are defined as Principal Components (PCs) representing the maximum variance of the
total dataset. The points between the lowest value and the highest value of each PC define an
M-dimensional (M is the number of significant components) hyper-rectangle with sides parallel to
the PCs (Nikolova-Jeliazkova & Jaworska, 2005).
• Criteria: This hyper-rectangle AD also includes empty spaces depending on uniformity of data
distribution (Nikolova-Jeliazkova & Jaworska, 2005). Combining the bounding box method with
principal component analysis (PCA) can overcome the setback of correlation between descriptors
Figure 1. Bounding box plot
184
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use

d x y x y( , ) = −
( )
( )
Importance of Applicability Domain of QSAR Models
Table 1. Summary of hypothesis of most commonly used AD methods
Method Formula/Equation Hypothesis on Data Distribution Shape of the AD Space
Range based
Geometric
(Convex hull)
Distancebased
Mahalanobis
distance/
Hotelling
2
/leverage
T
- Uniform and smallest axis-aligned
−1
T
D x x x
( , )µ µ µ= −
M
( )
∑
( )
Uniform Boundary created by
convex region containing all the data
points
a) Normal, b) arbitrary variances, and c)
−
arbitrary correlation.
Mean μ, Covariance matrix
p x D x
1
=
N
π
2
( )
2
/
∑
exp ,
1
2
individual descriptor
of the training set
compounds
Convex
Ellipses (hyper-ellipses)
∑
1
−
µ
( )
M
2
Euclidean
distance
City block
Probability density
distribution
n
d x y x y
( , ) ( )= −
n=total number of observations for
which a pair wise comparison is
performed (xi and yi),
i is the number of descriptors
d x y x y
( , ) = −
ϕπ=
( )
where, Ф (xi and xj) is the potential
induced on xj by xi and width of the
curve is defined by smoothing
parameter s. The cut off value
associated with Gaussian potential
functions, namely fp, can be
calculated by:
f f q j f f
= + −
p i j j
where,q p
percentile value of probability
density, n is the number of
compounds in the training set and j is
the nearest integer value of q.
1
s
2
= ×
n
=∑1
i
.exp
∑
=
i
1
( )
n
i i
2
s x x
2
100
i i
−
1
−
( )
i j
−
+1
, p is the
2
a) Normal, b) equal variances, and c)
uncorrelated variables
Mean μ, unit covariance matrix
p x D x
Multivariate uniform Rectangle
1. Parametric methods use standard
distributions such as Gaussian and
Poisson distributions.
2
2. Non-parametric techniques permit the
estimation of probability density solely
from data.
1
=
2
π
( )
N
2
/
1
2
−
exp ,
( )
E
2
µ
Spheres (hyper-spheres)
Captures actual data
distribution and no
reference data point are
considered. For each
data point, the method
calculates the probability
of belonging to the set.
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
185
Соседние файлы в папке Библиотека им академика М.И. Перельмана
