Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5587_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
34 Мб
Скачать
Importance of Applicability Domain of QSAR Models
but the issue regarding empty regions within the interpolation space still remains. It is interest­ing to point out that the empty space is smaller than the hyper-rectangle in the original descrip­tor ranges. The most important criterion of this approach remains in the selection of appropriate number of components (Jaworska et al., 2005, Worth et al., 2005). The PCA bounding box has been used to analyze the AD of the KOWWIN model.
The AD in principal component space is presented in Figure 2. The QSAR model training set is repre­sented by the biggest circle. The predictions of query molecules within the training space are considered reliable. Query compounds which are located outside the training model space would be expected to be less reliably predicted.
1.3 TOPKAT Optimal Prediction Space
Theory: Variation of PCA is implemented by the Optimum Prediction Space (OPS) from
TOPKAT OPS (2000). In the PCA approach, instead of the standardized mean value, the data are centered around the mean of individual parameter range ([x
max–xmin
orthogonal coordinate system which is known as OPS coordinate system. Here, the basic process is same, extracting eigenvalues and eigenvectors from the covariance matrix of the transformed data. The OPS boundary is defined by the minimum and maximum values of the generated data points on each axis of the OPS coordinate system.
Criteria: The Property Sensitive object Similarity (PSS) is implemented in the TOPKAT as a
heuristic solution to replicate the data set’s dense and spare regions, and includes the response variable (y). Accuracy of prediction between the training set and queried points is assessed by the PSS. Similarity search method is used to evaluate the performance of TOPKAT in predicting the effects of a chemical that is structurally similar to the training structure.
]/2). Thus, it establishes a new
2. GEOMETRICAL METHODS
2.1 Convex Hull
Theory: This approach estimates the direct coverage of an n-dimensional set using the convex
hull calculation (Preparata & Shamos, 1991). Convex hull calculation is a computational geom­etry problem which is performed based on complex but efficient algorithms. The approach recog­nizes the boundary of the dataset considering the degree of data distribution.
Criterion: Interpolation space is defined by the smallest convex area containing the entire train-
ing set.
186
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
Importance of Applicability Domain of QSAR Models
Figure 2. A graphical illustration of the AD in principal component (PC) space
Drawbacks:
Implementing a Convex Hull can be challenging with increasing data complexity, this means
that an increase in dimensions contributes to the order of complexity. Convex hull calcula­tion is efficient for two and three dimensions. The complexity swiftly amplifies in higher di­mensions. For n points and d dimensions, the complexity is of order O which can be defined as O = [n
[d/2]+1
The approach only analyses the set boundaries without considering the actual data distribu-
tion, and
This approach cannot identify the potential internal empty regions within the interpolation
space (Jaworska et al., 2005).
For the ‘convex hull’ method, the domain of applicability is the smallest axis-aligned convex region containing all the data points presented in Figure 3.
3. DISTANCE-BASED METHODS
Distance-based methods are mainly founded on the ‘distance-to-centroid’ principle. The ellipsoidal region is centered on the dataset grand mean and has its principal axes based on the eigenvectors of the dataset variance/covariance matrix. In Figure 4, a graphical illustration of ‘distance-to-centroid’ is presented.
There are various types of distance-based methods which can be used for the determination of AD of QSAR models. We have tried to represent the most commonly used distance-based AD approaches in the following sections.
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
187
Importance of Applicability Domain of QSAR Models
Figure 3. Convex hull plot
Figure 4. Distance-to-centroid plot
3.1 Leverage Approach
Theory: The leverage (h) of a compound in the original variable space is calculated based on the
HAT matrix as:
H = (X
In Equation (1), H is an [n x n] matrix that orthogonally projects vectors into the space spanned by the columns of X (Eriksson et al., 2003; Gramatica, 2007). The AD of the model is defined as a squared area within the ±3 band for standardized residuals (σ) and the leverage threshold is defined as h*=3(p+1)/n, where p is the number of variables and n is the number of compounds. The leverage values (h) are cal­culated for each compound and plotted vs. cross-validated standardized residuals (σ) (Y-axis) referred to as the Willam’s plot.
leverage points fit the model well (having small residuals), they are called good high leverage points or good influence points. Those points stabilize the model and make it more accurate. On the contrary, high leverage points, which do not fit the model (having large residuals), are called bad high leverage points or bad influence points.
T(XTX)–1
X). (1)
The leverage approach assumes normal data distribution. It is interesting to point out that when high
Criteria: Graphically, the Williams plot (mostly used for the leverage approach) verifies the pres-
ence of response outliers and training set chemicals that are structurally very influential in deter­mining model parameters. The data predicted for high leverage chemicals in the prediction set are extrapolated and could be less reliable. Here, we have tried to simplify the concept with a pictorial representation of Williams plot in Figure 5.
It is interesting to point out that this figure has been drawn based on arbitrary data to show the every possibility of AD assessment in training and test set compounds. Assessing Figure 5, one can draw the following conclusion:
188
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
Importance of Applicability Domain of QSAR Models
1. Compound number 6 of the training set possesses standardized cross-validated residual value greater than -3σ unit, but it lies within the critical HAT (h*) value of 0.58. This compound may have flawed observed value and it can be identified as an outlier due to a wrong value of response variable. Therefore, compound number 6 is a response outlier.
2. Compounds with serial numbers 9 and 25 belonging to the training set are characterized by high leverage values, but they lie within the fixed ±3σ limit of the ordinate. The mentioned chemicals are not response outliers (not Y outliers), but influence the applicability domain to span over larger area and hence can be marked as influential chemicals or X outliers.
3. Compound number 35 belonging to the test set is wrongly predicted having standardized cross­validated residual value greater than -3σ unit) as well as completely outside of the AD as defined by HAT vertical line (higher leverage value than the critical value).
4. On the contrary, one of the test set compounds (41) belongs to AD as it lies within the critical HAT value but the prediction for this compound is not precise (greater than 3σ unit).
5. The test set compound with serial number 27 is within the ±3σ limit of the ordinate, but possess higher leverage value (h > h*). As we can see from Figure 5, this compound is surrounded by other two influential training set chemicals (serial numbers 9 and 25), therefore its prediction can be assumed to be somewhat reliable or dependable though it lies outside the AD.
Figure 5. Williams plot
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
189
Importance of Applicability Domain of QSAR Models
3.2 Euclidean Distance
Theory: This is one of the major distance based AD assessment approaches. Like other distance-
based approaches, it calculates the distance from every other point to a particular point in the data set. The distance scores are calculated by the Euclidean distance norm. A distance score,
, for two different compounds Xi and Xj can be measured by the Euclidean distance norm. The
d
ij
Euclidean distance can be expressed by the following equation:
m
d x x
=
ij ik jk
( )
1
=
k
2
. (2)
The mean distances of one sample to the residual ones are calculated as follows:
n
d
ij
j
=∑1
d
=
i
, (3)
n
1
where, i=1,2,….,n.
The mean distances are then normalized within the interval of zero to one. It is applicable only for
statistically independent descriptors (Jaworska et al., 2004).
Criteria: Compounds with distance values adequately higher than those of the most active
probes are considered to be outside the domain of applicability. The mean normalized dis­tances are measured for both training and test set compounds. The boundary region created by normalized mean distance scores of the training set are considered as the zone of applicability domain for test set compounds. If the test set compounds are inside the domain/area covered by the training set compounds, it means that these compounds are inside the applicability domain, otherwise not.
An example of Euclidean plot is presented in Figure 6. Based on the assessment of the Euclidean plot, one can conclude that test set compounds 19 and 23 are outside of the applicability domain created by the training set compound from which the QSAR model has been developed.
3.3 Mahalanobis Distance
Theory: The Mahalanobis distance approach considers the distance of an observation from the mean
values of the independent variables but not taking the impact on the predicted value. It offers one of the unique and simple approaches for identification of outliers. Mahalanobis distance is unique because it automatically takes into account the correlation between descriptor axes (Hair et al., 2005).
Criterion: The threshold value does not have any rule of thumb. However, observations with
values much higher than those of the remaining ones may be considered to be outside the domain.
190
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
Importance of Applicability Domain of QSAR Models
Figure 6. Euclidean distance plot for AD study of an arbitrary QSAR model
3.4 City Block Distance
City-block distance is simply the summed difference across dimensions and is calculated from the fol­lowing equation:
n
d x y x y
( , ) =
=∑1
i
It examines the absolute differences between coordinates of a pair of objects (x
. (4)
i i
and yi). City-block
i
distance assumes a triangular distribution. The method is particularly useful for the discrete type of descriptors. It is used only for training sets which are uniformly distributed with respect to count-based descriptors (or fragment mapping counts) (Jaworska et al., 2004).
3.5 Hotelling T2 Test
Theory: The Hotelling T2 method is a multivariate student’s t test and proportional to leverage and
Mahalanobis distance approach. It assumes a normal data distribution like the leverage approach (Hair et al., 2005). The method is used to evaluate the statistical impact of the difference on the
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
191
Importance of Applicability Domain of QSAR Models
means of two or more variables between two groups. Hotelling T2 corrects for collinear descrip-
2
tors through the use of the covariance matrix. Hotelling T tion from the center of a set of X observations. A tolerance volume is derived for Hotelling T
measures the distance of an observa-
2
.
Criterion: Based on the t value, the significant compounds within the domain are determined.
3.6 K-Nearest Neighbors Approach
Theory: The approach is based on similarity search for a new chemical entity with respect
to the space created by the training set compounds. The similarity is identified by finding the distance of a query chemical from nearest training compound or its distances from k-nearest neighbors in the training set. Thus, similarity to the training set molecules is significant for this approach in order to associate a query chemical with reliable prediction (Sheridan et al.,
2004). Like other methods, descriptors are used for calculate the similarity between training and test molecules.
Criterion: If the calculated distance values of test set compounds or query molecules are within
the user defined threshold set by the training set molecules, then the prediction of these com­pounds are considered to be reliable.
How One Can Find the Similarity to the Molecules of the Training Set
1. Considering all the descriptors, the mean similarity of a test set molecule to the most similar k
members of the training set can be found. A variation of this is to calculate this similarity using only the subset of descriptors in the training set.
2. The similarity of the test set molecule to the descriptor centroid of the training set.
How to Decide the Neighboring Test or Query Compounds in the Training Set
1. The simplest method is to pick a similarity cutoff above which the test set molecule and a training
set molecule can be considered “neighbors” and count the molecules that exceed that cutoff.
2. Another complicated approach is explained by Sheridan et al. (2004) which demonstrated that “A
more complicated approach is to define a falloff function f(Similarity) that defines how much of a neighbor a training set molecule is as a function of similarity to the test set molecule. The number of neighbors is the sum of f (Similarity) over all members of the training set. For instance, we use a linear function where a molecule is 1.0 of a neighbor at similarity 1.0 and 0.0 of a neighbor at a similarity 0.6; anything between is linearly interpolated.”
A k-nearest neighbors plot is presented in Figure 7, where the distance of a point to the domain is taken as the average distance to the k- nearest data points.
3.7 DModX (Distance to the Model in X-Space)
Theory: This is another important distance based AD finding approach for QSAR models. This
approach was developed by Wold et al., (2001) and usually applied for partial least squares (PLS) models. The basic theory lies in the residuals of Y and X which are of diagnostic value for the quality
192
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
Importance of Applicability Domain of QSAR Models
Figure 7. K-nearest neighbors plot
of the model. As there are number of X-residuals, one needs a summary for each observation. This is accomplished by the residual standard deviation (SD) of the X-residuals of the corresponding row of the residual matrix E. As this SD is proportional to the distance between the data point and the model plane in X-space, it is also usually called DModX (distance to the model in X-space). Here, X is the matrix of predictor variables, of size (N*K), Y is the matrix of response variables, of size (N*M) and E is the (N*K) matrix of X-residuals, N is number of objects (cases, observations), k is the index of X-variables (k = 1, 2, . . ., K) and m is the index of Y-variables (m = 1, 2, . . ., M).
Criteria: A DModX value larger than around 2.5 times the overall SD of the X residuals (corre-
sponding to an F-value of 6.25) indicates that the observation is outside the applicability domain of the model. In the DModX plot, the threshold line is attributed as D-critical line and this plot can be easily drawn in SIMCA-P software for any PLS model (SIMCA-P 10.0, 2002).
An example of DModX plot is presented in Figure 8 where the plot is constructed at 95% confidence level by SIMCA-P. The DModX values of 26 test compounds are within the critical value of 1.986. Chemical 10 along with the compound numbers 3, 25, 53 and 73 are outside the evaluated critical value indicating that these compounds are outside of the applicability domain of the developed model.
3.8 Tanimoto Similarity
The Tanimoto index is a measure of similarity between two molecules based on the number of com­mon molecular fragments (Tetko et al., 2008). In order to calculate the Tanimoto similarity, all unique fragments of a particular length in two compounds are calculated. The Tanimoto similarity between the compounds J and I is defined as:
N
x x
( )
J i K i
, ,
=∑1
TANIMOTO J K
,
( )
=
N
x x x x x
( )
, , , , ,
J i J i K i K i J i
i
where, N is the number of unique fragments in both the compounds, x i-th fragment in the compounds J and K. Based on equation 3, the distance between two compounds J
and K is “1– TANIMOTO(J, K)” and the distance of a compound to the model is the minimum distance between the investigated compound and compounds from the training set of the model.
i
+
N
( )
i
N
.
( )
===
111
i
(5)
..,x
K i
and x
J, i
are the counts of the
K, i
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
193
Figure 8. DModX plot for PLS based QSAR model
Importance of Applicability Domain of QSAR Models
3.9 Standard Deviation of the Ensemble Predictions (STD)
In order to estimate the uncertainty of the model, standard deviation of the predictions obtained from an ensemble of models can be utilized. The basic principle behind this idea is very simple and it depicts that if different models give significantly different predictions for a particular molecule, then the prediction for this compound is more likely to be unreliable. The sample standard deviation can be ideally used as an estimator of model uncertainty.
As an example, consider that Y(J) = {y set of N trained models, the corresponding distance to model STD can be defined by the following equation:
y J y
( )
d J stdev Y J
( ) =
STD
Based on the type of an ensemble used to evaluate the standard deviation, the STD approach can be classified into several subtypes. (a) Consensus STD or CONS-STD for models developed based on dif­ferent machine learning techniques, (b) Associative neural networks or ASNN-STD for an ensemble of neural network models, and (c) BAGGING-STD for an ensemble of models formed utilizing the bagging technique. The distance based method STD has been verified to provide outstanding results for discrimi­nation of highly accurate predictions for regression models (Manallack et al., 2003; Tetko et al., 2008).
( )
( )
=
(J), i=1..N} is a set of predictions for a compound J given by a
i
2
( )
i
N
1
. (6)
194
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use

Importance of Applicability Domain of QSAR Models
A graphical illustration of the STD approach has been given in Figure 9. Here, reliable prediction
has a low prediction spread and unreliable prediction has higher prediction spread.
3.10 Correlation of Prediction Vectors (CORREL)
The basic theory of this approach is based on the correlation of vectors of ensemble’s predictions for the target compound and compounds from the training set (Tetko, 2008). Similar to the STD method, this measure is relevant only for ensembles of models. More specifically, CORREL measure for the target compound J is calculated according to the following expression:
 
and
( ) ( )
i
y J
( )
 
. (7)
define the vectors of ensemble’s predictions for the training set
d J corr y T y J
CORREL
= −
11max ,
( )
In Equation (7),
i n
=
...
y T
  
( )
i
compound Ti and the target compound J, corr is Spearman rank correlation coefficient between the two vectors and N is the number of compounds in the training set. A low value of CORREL (that indicates the high Spearman correlation coefficient) specifies that for target compound J, there is a compound T from training set for which predictions of the ensemble of models are strongly correlated. If a compound T has the same descriptors as the target compound J, then predictions of models will be identical for both molecules and thus resulted CORREL(J) will be 0. Compounds having high correlation coefficient values are considered to be “closer to the model”. The Spearman correlation outperformed several al-
Figure 9. STD graph for AD estimation
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
195