Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5587_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
34 Мб
Скачать
Quantitative Structure-Activity/Property/Toxicity Relationships
Vela, A., & Gázquez, J. L. (1990). A relationship between the static dipole polarizability, the global softness, and the Fukui function. Journal of the American Chemical Society, 112(4), 1490–1492. doi:10.1021/ja00160a029
Vijayaraj, R., Subramanian, V., & Chattaraj, P. K. (2009). Comparison of global reactivity descriptors calculated using various density functionals: A QSAR perspective. Journal of Chemical Theory and Computation, 5(10), 2744–2753. doi:10.1021/ct900347f
Voskresensky, O. N., & Levitsky, A. P. (2002). QSAR aspects of flavonoids as a plentiful source of new drugs. Current Medicinal Chemistry, 9(14), 1367–1383. doi:10.2174/0929867023369790 PMID:12132993
Wan, J., Zhang, L., Yang, G., & Zhan, C.-G. (2004). Quantitative structure-activity relationship for cyclic imide derivatives of protoporphyrinogen oxidase inhibitors: A study of quantum chemical descriptors from density functional theory. Journal of Chemical Information and Computer Sciences, 44(6), 2099–2105. doi:10.1021/ci049793p PMID:15554680
Wiener, H. (1947). Structural determination of paraffin boiling points. Journal of the American Chemi- cal Society, 69(1), 17–20. doi:10.1021/ja01193a005 PMID:20291038
Wolff, M. E. (1955). Therapeutic agents. In Burger’s medicinal chemistry and drug discovery (Vol. 4). New York: John Wiley & Sons.
Yan, D., Jiang, X., Xu, S., Wang, L., Bian, Y., & Yu, G. (2008). Quantitative structure-toxicity relation­ship study of lethal concentration to tadpole (Bufo vulgaris formosus) for organophosphorous pesticides. Chemosphere, 71(10), 1809–1815. doi:10.1016/j.chemosphere.2008.02.033 PMID:18395243
Yang, G.-F., & Huang, X. (2006). Development of quantitative structure-activity relationships and its application in rational drug design. Current Pharmaceutical Design, 12(35), 4601–4611. doi:10.2174/138161206779010431 PMID:17168765
Yang, W., & Mortier, W. J. (1986). The use of global and local molecular parameters for the analysis of the gas-phase basicity of amines. Journal of the American Chemical Society, 108(19), 5708–5711. doi:10.1021/ja00279a008 PMID:22175316
Yang, W., & Parr, R. G. (1985). Hardness, softness, and the fukui function in the electronic theory of metals and catalysis. Proceedings of the National Academy of Sciences of the United States of America, 82(20), 6723–6726. doi:10.1073/pnas.82.20.6723 PMID:3863123
Yang, W., Parr, R. G., & Pucci, R. (1984). Electron density, Kohn− Sham frontier orbitals, and Fukui functions. The Journal of Chemical Physics, 81(6), 2862–2863. doi:10.1063/1.447964
Zhang, L., Hao, G.-F., Tan, Y., Xi, Z., Huang, M.-Z., & Yang, G.-F. (2009). Bioactive conformation analysis of cyclic imides as protoporphyrinogen oxidase inhibitor by combining DFT calculations, QSAR and molecular dynamic simulations. Bioorganic & Medicinal Chemistry, 17(14), 4935–4942. doi:10.1016/j.bmc.2009.06.003 PMID:19540767
Zhang, S. G., Lei, W., Xia, M. Z., & Wang, F. Y. (2005). QSAR study on N-containing corrosion in­hibitors: Quantum chemical approach assisted by topological index. Journal of Molecular Structure THEOCHEM, 732(1-3), 173–182. doi:10.1016/j.theochem.2005.02.091
176
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
Quantitative Structure-Activity/Property/Toxicity Relationships
KEY TERMS AND DEFINITIONS
Aliphatic Compounds: These are non-aromatic compounds made by carbon and hydrogen.
Electrophilicity: It is a measure of ‘electrophilic power’ of a molecule. It can be estimated as the
electronegativity squared divided by twice the hardness.
Energies of Frontier Molecular Orbitals: They are the energies of the highest occupied molecular orbital and lowest unoccupied molecular orbital, respectively.
Group Philicity: It is a condensed philicity summed over a group of relevant atoms.
Net Atomic Charge: It is a non-integer charge at each atom in a molecule in elementary charge units.
Net Electrophilicity: It is the electron-accepting power of a molecule relative to its own electron-
donating power.
Polyaromatic Hydrocarbons: These are compounds made by only carbon and hydrogen and contain multiple aromatic rings.
Polychlorinated Biphenyls: These are compounds containing two phenyl rings and several chlorine atoms are attached with phenyl rings.
Tetrahymena Pyriformis: These are free-living ciliate protozoa, generally found in freashwater ponds.
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
177
Quantitative Structure-Activity/Property/Toxicity Relationships
APPENDIX: LIST OF SYMBOLS AND ABBREVIATIONS
SAR: Structure-Activity Relationship QSAR: Quantitative Structure-Activity Relationship QSTR: Quantitative Structure-Toxicity Relationship DFT: Density Functional Theory CDFT: Conceptual Density Functional Theory LSFER: Linear Solvation Free Energy Relationship CoMFA: Comparative Molecular Field Analysis QSPR: Quantitative Structure-Property Relationship PCA: Principal Component Analysis
η: Hardness χ: Electronegativity ω: Electrophilicity
f ( )r : Fukui function
f
: Condensed Fukui function
k
s(r): Local softness η(r): Local hardness ω(r): Philicity IP: Ionization Potential EA: Electron Affinity μ: Chemical Potential E: Energy N: Number of Electrons v(r): External Potential
±
: Net Electrophilicity
Δω
HIV-1: Human Immunodficiency Virus type 1 NCp7: Nucleocapsid Protein p7
W: Electrical Power V: Voltage R: Electrical Resistance α: Polarizability ξ: Magnetizability ρ(r): Electron Density
: Atomic Charge at kth site
q
k
α
: Group Philicity
ω
g
PMH: Principle of Maximum Hardness MEP: Minimum Electrophilicity Principle MPP: Minimum Polarizability Principle MMP: Minimum Magnetizability Principle
: 50% Lethal Dose
LD
50
: 50% Lethal Concentration
LC
50
178
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
Quantitative Structure-Activity/Property/Toxicity Relationships
Z: Atomic Number
: Number of Nonhydrogenic Atoms
N
NH
R or r: Coefficient of correlation
2
R
CV
or r
2
: Variance of Leave-one-out Cross-Validation
CV
PAH: Polyaromatic Hydrocarbons TCDD: Tetrachlorodibenzo-p-Dioxin Ah: Arylhydrocarbon PCDF: Polychlorinated Dibenzofurans CA: Chloroanilines PCB: Polychlorinated Biphenyl RBA: Relative Binding Affinity TeBG: Testosterone-Binding Globulins MAPT Myotrophic to Androgenic Potency in Temporal Activity.
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
179
180
Chapter 5
Importance of Applicability
Domain of QSAR Models
Kunal Roy
Jadavpur University, India
Supratik Kar
Jadavpur University, India
ABSTRACT
Quantitative Structure-Activity Relationship (QSAR) models have manifold applications in drug discovery, environmental fate modeling, risk assessment, and property prediction of chemicals and pharmaceuticals. One of the principles recommended by the Organization of Economic Co-operation and Development (OECD) for model validation requires defining the Applicability Domain (AD) for QSAR models, which allows one to estimate the uncertainty in the prediction of a compound based on how similar it is to the training compounds, which are used in the model development. The AD is a significant tool to build a reliable QSAR model, which is generally limited in use to query chemicals structurally similar to the training compounds. Thus, characterization of interpolation space is significant in defining the AD. An attempt is made in this chapter to address the important concepts and methodology of the AD as well as criteria for estimating AD through training set interpolation in the descriptor space.
INTRODUCTION
Quantitative structure-activity relationship (QSAR) modelling has become an important tool in the new drug candidate design, environmental fate modeling, toxicity and property prediction of chemicals and pharmaceuticals since they offer an economical and time-effective alternative to the medium throughput in vitro and low throughput in vivo assays (Perkins et al., 2003; Selassie, 2003; Walker et al., 2003). A QSAR model is a simple mathematical equation that is evaluated from a set of molecules with known activities/properties/toxicities using computational approaches. Thereafter, the developed predictive QSAR models are also applied by regulatory agencies to evaluate physical, chemical, and biological properties of individual chemical entities using applications specific for decision-making frameworks in risk and safety assessments (Kar & Roy, 2010). The QSAR modeling hypothesis also supports the 3Rs
DOI: 10.4018/978-1-4666-8136-1.ch005
Copyright © 2015, IGI Global. Copying or distributing in print or electronic forms without written permission of IGI Global is prohibited.
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
Importance of Applicability Domain of QSAR Models
(replacement, refinement and reduction in animals in research) paradigm due to an increased pressure from social and economic background to trim down the use of animal testing as an important alternative method for future prediction of untested chemical entities (Benigni & Giuliani, 2003).
Robust validation of QSAR models plays a key step for the selection of a predictive model that may be considered for future prediction of new molecules. As numerous numbers of researches have been directed to the design of new molecules with the utilization of QSAR technique, validation of a QSAR model has been certified as the most considerable stride (Carlsen et al., 2009) for assessing the quality of data, applicability and mechanistic interpretability of the developed model. Thus, a huge number of investigations are currently directed towards the introduction of more appropriate validation approaches for more accurate and predictive QSAR model development. One important objective of QSAR modeling is to predict activity/property/toxicity of new chemical entities falling within the applicability domain of the developed models. The reliability of any QSAR model depends on the confident predictions of these new compounds based on applicability domain (AD) of the modeled compounds and therein arises the importance of the AD study (Ferreira, 2001).
In order to establish the scientific validity of a QSAR model and to facilitate its acceptance for regulatory purposes, the Organization for Economic Cooperation and Development (OECD) in its joint meeting (OECD,
2007) has agreed to five principles that should be followed during the construction of QSAR models. The OECD Principle 3 which defines the need of an AD expresses the fact that QSARs are unavoidably associated with restrictions in terms of the types of chemical structures, physicochemical properties and mechanisms of action for which the models can generate reliable predictions. The applicability domain of a QSAR model has been defined as the response and chemical structure space, characterized by the properties of the molecules in the training set. The developed QSAR model can predict a new compound truly only when it falls within the applicability domain of the developed model (Netzeva et al., 2005). Thus, to identify the interpolation (true prediction) or extrapolation (less reliable prediction) of query compounds is an important task for a QSAR model developer using the information of applicability domain (Netzeva et al., 2005).
Viewing the importance of AD in QSAR model validation, we herein focus to get an overview of dif­ferent traditional as well as relatively new AD approaches used to judge the quality of the QSAR models. This book chapter will be helpful for the QSAR learners in order to have a clear idea on the principles of available AD approaches useful for judging the predictive quality of QSAR models.
BACKGROUND
A QSAR model is essentially valued in terms of its predictability, indicating how well it is able to predict the endpoint values of the compounds which are not used to develop the correlation. The models that have been suitably validated internally and externally can be considered reliable for both scientific and regulatory purposes (Golbraikh & Tropsha, 2002; Tong et al., 2004]. QSAR models should be validated according to the OECD principles for reliable prediction. A meeting of QSAR experts held in Setúbal, Portugal in March 2002 formulated guidelines for the validation of QSAR models, in particular for regulatory purposes (Jaworska et al., 2003). These principles were agreed by OECD member countries,
th
QSAR and regulatory communities at the 37 Party on Chemicals, Pesticides and Biotechnology in November 2004. These principles are best pos­sible summary of the most important points that are required to be addressed to find consistent, reliable and reproducible QSAR models (OECD, 2004). Five OECD principles to develop reliable models are:
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
Joint Meeting of the Chemicals Committee and Working
181
Importance of Applicability Domain of QSAR Models
Principle 1: A defined endpoint; Principle 2: An unambiguous algorithm; Principle 3: A defined domain of applicability; Principle 4: Appropriate measures of goodness-of fit, robustness and predictivity; Principle 5: A mechanistic interpretation, if possible (OECD, 2004).
In order to provide guidance on the interpretation of the mentioned principles, the OECD has also provided a checklist (OECD, 2004). Thus, the current challenge in the process of development of a QSAR model is no longer in developing a model that is statistically sound to predict the activity within the training set, but in developing a model with the capability to predict accurately the activity of untested chemicals. In spite of high fitted accuracy and apparent mechanistic appeal, some published QSAR models fail rigorous validation tests, and, thus, may lack practical utility as reliable screening tools (Tropsha et al., 2003]. In summary, predictive power is one of the most important validation features of QSAR models. It can be defined as the ability of a model to predict accurately the target activity/ property/toxicity of compounds that are not used for model development.
In this context, QSAR model predictions are most reliable if they come from the model’s applicability domain (AD) which is broadly defined under OECD principle 3. The OECD includes AD assessment as one of the QSAR acceptance criteria for regulatory purposes (OECD, 2004; OECD, 2007). The Setubal Workshop report (Jaworska et al., 2003) presented the following regulation for AD assessment: “The applicability domain of a (Q)SAR is the physico-chemical, structural, or biological space, knowledge or information on which the training set of the model has been developed, and for which it is applicable to make predictions for new compounds. The applicability domain of a (Q)SAR should be described in terms of the most relevant parameters, i.e., usually those that are descriptors of the model. Ideally the (Q)SAR should only be used to make predictions within that domain by interpolation not extrapolation. This depiction is useful for explaining the instinctive meaning of the “applicability domain” approach.
Two general principles to AD estimation have been proposed till date (Nikolova-Jeliazkova & Ja­worska, 2005). The first one estimates the interpolation region based on the training set in the model’s descriptor space. The second approach relies on similarity analysis according to the premise that a QSAR prediction is reliable if the compound is “similar” to the compounds in the training set. It is interesting to point out that “similarity” is a relative concept as different concepts of similarity are relevant to dif­ferent endpoints. The similarity approach to AD assessment should be based on measurable key features relevant to the endpoint modelled precisely (Bender & Glen, 2004; Nikolova & Jaworska, 2004).
APPLICABILITY DOMAIN (AD)
The AD (Eriksson et al., 2003; Gramatica, 2007; Netzeva et al., 2005, Tetko et al., 2006) is a theoretical region in chemical space encompassing both the model descriptors and modeled response. In the develop­ment of a QSAR model, the domain of applicability of molecules plays a crucial role for estimating the uncertainty in the prediction of a particular compound based on how similar it is to the compounds used to build the model. Thus, the prediction of a modeled response using QSAR is valid only if the compound being predicted falls within the domain of applicability of the model since it is impossible to predict an entire universe of chemicals using a single QSAR model. Again, AD can be described as the physico-chemical,
182
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
Importance of Applicability Domain of QSAR Models
structural or biological space, knowledge or information based on which the training set of the model is developed, and the model is applicable to make predictions for new compounds within the specific domain (Weaver & Paul Gleeson, 2008).
The selection method of the training and test sets has a significant effect on the applicability domain of constructed QSAR model. Thus, while splitting a dataset for external validation, the training set molecules should be selected in such a way that they span the entire chemical space for all the dataset molecules. In order to achieve successful predictions, a QSAR model should always be used for compounds within its applicability domain. Therefore, one can arguably agree that if a new compound falls outside of the AD of the training set molecules, its prediction is not reliable.
APPLICABILTY DOMAIN APPROACHES
There are various AD approaches that have been used by different groups of QSAR researchers world­wide. The most common approaches for estimating interpolation regions in a multivariate space include the following ones (Jaworska et al., 2005; Netzeva et al., 2005):
1. Ranges in the descriptor space.
2. Geometrical methods.
3. Distance-based methods.
4. Probability density distribution.
5. Range of the response variable.
The first four approaches are based on the methodology used for interpolation space characterization in the model descriptor space. On the contrary, the last one depends solely on response space of the training set molecules. A compound can be identified as out of the domain of applicability in a simple way, if:
1. At least one descriptor is out of range for the ranges approach, and
2. The distance between the chemical and the center of the training data set exceeds the threshold for
distance approaches.
The threshold for all kinds of distance methods is the largest distance between the training set data points and the center of the training data set.
Along with the common approaches, Stanforth et al., (2007) proposed a cluster-based approach to evaluate the AD of any QSAR model. The method applies an intelligent version of the k-means cluster­ing algorithm for modeling the training set as a compilation of clusters in the descriptors space. On the other hand, the test compounds of individual clusters are assigned a fuzzy membership from which an overall distance may be calculated to identify the AD of individual test compounds. A classification­based process involves calculation of regression residuals for identification of good and bad classes suggested by Guha & Jurs (2005).
All approaches are discussed thoroughly in this section enlightening their main theories to define the interpolation space as well as the threshold criteria for AD used in QSAR modeling. A summary of hypotheses of data distribution and interpolation space of most commonly used AD methods is discussed in Table 1.
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
183
Importance of Applicability Domain of QSAR Models
1. RANGES IN THE DESCRIPTOR SPACE APPROACH
1.1 Fixed or Probabilistic Boundaries/Bounding Box
Theory: The ranges of the individual descriptors are considered with a uniform distribution de-
fining an n-dimensional hyper-rectangle developed on the basis of the maximum and minimum values of each descriptor used to build the model with sides parallel to the coordinate axes.
Criteria: This method is simply based on boundary created by individual descriptor ranges of the
training set compounds from which the QSAR model is developed. As mentioned earlier, this is based on the highest and lowest values of X variables (descriptors to develop the QSAR model) and Y variable (response for which the QSAR equation is created) of the training set. Any test set compounds, which are not present in any of these particular ranges, are considered out of the AD and their predictions are less reliable (Jaworska et al., 2005; Seber, 1984).
Drawbacks:
For non-uniformly distributed data, the approach encloses considerable empty space, Empty regions in the interpolation space cannot be recognized, as only the descriptor ranges
are considered, and
Correlation between descriptors cannot be taken into account (Jaworska et al., 2005).
For the ‘bounding box’ method, the domain of applicability is the smallest axis-aligned rectangular box containing all the data points presented in Figure 1.
1.2 Principal Component Ranges/PCA Bounding Box
Theory: Principal components convert the original data into a new orthogonal coordinate system
by the rotation of axes and enable to correct for correlations among descriptors. Newly formed axes are defined as Principal Components (PCs) representing the maximum variance of the total dataset. The points between the lowest value and the highest value of each PC define an M-dimensional (M is the number of significant components) hyper-rectangle with sides parallel to the PCs (Nikolova-Jeliazkova & Jaworska, 2005).
Criteria: This hyper-rectangle AD also includes empty spaces depending on uniformity of data
distribution (Nikolova-Jeliazkova & Jaworska, 2005). Combining the bounding box method with principal component analysis (PCA) can overcome the setback of correlation between descriptors
Figure 1. Bounding box plot
184
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
d x y x y( , ) =
( )
( )
Importance of Applicability Domain of QSAR Models
Table 1. Summary of hypothesis of most commonly used AD methods
Method Formula/Equation Hypothesis on Data Distribution Shape of the AD Space
Range based
Geometric (Convex hull)
Distance­based
Mahalanobis distance/ Hotelling
2
/leverage
T
- Uniform and smallest axis-aligned
1
T
D x x x
( , )µ µ µ=
M
( )
( )
Uniform Boundary created by
convex region containing all the data points
a) Normal, b) arbitrary variances, and c)
arbitrary correlation.
Mean μ, Covariance matrix
p x D x
1
=
N
π
2
( )
2
/
exp ,
1 2
individual descriptor of the training set compounds
Convex
Ellipses (hyper-ellipses)
1
µ
( )
M
2
Euclidean distance
City block
Probability density distribution
n
d x y x y
( , ) ( )=
n=total number of observations for which a pair wise comparison is performed (xi and yi), i is the number of descriptors
d x y x y
( , ) =
ϕπ=
( )
where, Ф (xi and xj) is the potential induced on xj by xi and width of the curve is defined by smoothing parameter s. The cut off value associated with Gaussian potential functions, namely fp, can be calculated by:
f f q j f f
= +
p i j j
where,q p percentile value of probability
density, n is the number of compounds in the training set and j is the nearest integer value of q.
1
s
2
= ×
n
=∑1
i
.exp
=
i
1
    
( )
n
i i
2
s x x
2
100
i i
1
( )
i j
+1
, p is the
2
a) Normal, b) equal variances, and c) uncorrelated variables Mean μ, unit covariance matrix
p x D x
Multivariate uniform Rectangle
1. Parametric methods use standard
 
distributions such as Gaussian and
Poisson distributions.
2
 
2. Non-parametric techniques permit the
estimation of probability density solely from data.
1
=
2
π
( )
N
2
/
1
2
exp ,
( )
E
2
µ
Spheres (hyper-spheres)
Captures actual data distribution and no reference data point are considered. For each data point, the method calculates the probability of belonging to the set.
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
185