Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_117_библиотеки_им_акад_М_И_Перельмана

.pdf
Скачиваний:
2
Добавлен:
15.09.2026
Размер:
15 Мб
Скачать
☆
Explainable AI forColorectal Cancer Classication
209
feature extension method partly uses customized normalization methods to trans­form each scalar data point in the dataset into a vector. Variations of the sqrt-sum normalization method are also computed by applying random arithmetic computa­tions to it. The experiments were written in Python and executed on a computer with 16,280MB GPU RAM.
3.1 Datasets
This study utilized two publicly accessible datasets of CRC microbiome, which were collected from fecal samples. The rst dataset consists of shotgun sequence­curated samples and is based on the work done in Ref. [35]. The methodology for accessing and manipulating the curated microbiome data was sourced from Ref. [36]. Both the metadata, which includes demographic information, and the data le, which contains microbial counts per sample, were obtained. The dataset comprised of 884 samples and 2031 features, representing a collection of samples obtained from experiments conducted across six distinct countries, namely the United States, Germany, Austria, Italy, Canada, and France. The dataset was expanded to include a total of 2033 features by incorporating microbial count data along with two demo­graphic features, namely BMI and age. Data augmentation was performed due to the close association of both attributes with the development of CRC [37]. The clas­sication target label utilized in the study was based on the demographic attribute of study condition. Table1 presents the distribution of samples with respect to colorec- tal cancer (CRC) and healthy controls across various countries.
Table 1 shows that there were 368 CRC samples and 516 controls in a dataset of 884 combined samples. The disease attribute from metadata was used to lter the dataset. All samples that had an undened value in the disease eld were excluded from the dataset, which reduced the total number of samples to 796. The second dataset consists of 16S rRNA sequence data samples based on a study in Ref. [38]. The dataset, which has a total of 490 samples and 336 features, consists of 172, 198, and 120 that are normal, adenoma, and cancer samples, respectively. Part of the
Table 1 Composition of colorectal cancer-based microbiome samples from six different countries
Country
The United States 77 87 164 Germany 76 10 86 Austria 46 108 154 France 106 206 312 Italy 61 79 140 Canada 2 26 28 Total 368 516 884
Samples
TotalCRC Non-CRC
210
M
M
+
√
µ
σ
ij ij
M
,,
=+σ
M. Mulenga et al.
samples were collected in Toronto, Canada, while the rest were collected in Boston, Houston, and Ann Arbor in the United States.
3.2 Proposed Methods
This chapter proposes two methods, namely square-root sum and the combined method. In tandem with the XAI approach, the performance of the proposed model is described using underlying properties of the data by underscoring how the vari­ous method adjustments in the experiments affect the underlying data properties. As such the rst customized data normalization method proposed in this chapter is related to a method called Pareto scaling [34], which is computed as
ij i
,
=
ij
,
where M, σ,μ, i, and j represent respective feature value, standard deviation, mean, row number, and column number, respectively, in the dataset.
Pareto scaling stems from the concept called Pareto optimality, which seeks to make an individual better off without making another individual of the same group worse off [34]. The method improves the representation of features that have lower values while reducing the effect of noise in data. This method keeps the structure of the data almost intact. Though the method is able to reduce the effect of outliers in the dataset, it is not robust to dominant features [19]. As such, in this chapter we propose a method that attempts to reduce the effect of both outliers and dominant features. The method transforms a dataset by adding the standard deviation of the raw data to each of its data points and then computes the square root of the new data points as
(1)
For the sake of simplicity, we will refer to the parameter ω as sqrt-sum. Adding the standard deviation to a feature value in the sqrt-sum method is meant to shift the feature space in order to reduce the effect of outliers. Computation of the square root is meant to marginally reduce the inuence of dominant features, while slightly reducing the effect of values that have a lower effect. It is evident that the proposed method does not use Pareto scaling in its entirety, but only applies part of the method. The sqrt-sum method is only applied on nonzero values to avoid introduc­ing an unintended effect on the dataset since zeros in sequence data do not only represent counts, but they also represent missing values. The occurrence of missing values is a prevalent phenomenon in datasets pertaining to health-related research [39].
The second method called feature extension seeks to improve the performance of a DNN model by adjusting feature relevance in the dataset. Several normalization
(2)
f xxxxxx
()
=
()
x x1=
x2=ω
x x4=σ
Explainable AI forColorectal Cancer Classication
211
methods comprising sqrt-sum and its variants obtained by applying random compu­tations used to generate additional features in the dataset as
,,,
12345
,
(3)
where x represents the input scalar parameter and x1, x2, x3,x4, andx5 represent the output vector parameters computed based on the sqrt-sum method and its variants, which are also called the customized normalization methods.
Therefore, the output of the feature extension method represents individual out­puts of various normalization functions. The output of the customized normaliza­tion methods is also combined with features of the raw input dataset such that the rst parameter of output vector computed is equal to the original input scalar value and was computed as
(4)
To compute the second vector parameter, use the sqrt-sum method as
(5)
As mentioned above, the rest of the parameters are random variations of the sqrt- sum methods and that of the original input scalar value shown in Eqs. (2) and (4), respectively. Therefore, the third parameter of the vector, which we will simply refer to as absolute value of the difference (abs-diff), was computed using the parametermax(represents the highest value in the raw dataset) as
max2xx−
∣
=
3
(6)
The fourth parameter of the vector, which we will refer to as square root of the product (sqrt-prod), was computed as the square root of the product of the input value and the standard deviation of the raw data as
The fth parameter of the vector, which we will refer to as absolute shift (abs- shift), was calculated using the absolute value of the difference between the input parameter and the second parameter of the vector as
While sqrt-sum, sqrt-prod computed according to Eqs. (2) and (7) are purpose­fully designed to reduce the effect of dominant features, the parameters abs-diff and abs-shift are random computations that are only meant to produce noisy replicates of other methods. Knights etal. [23] and Lo and Marculescu [24] used the concept
x2xx−=∣
5
2
(7)
(8)
212
of noisy replicates to improve classication accuracy of ML models. The main objective of using noisy replicates in our method is to increase the number of fea­tures based on computations that have outputs that lie within the same feature space.
To evaluate the proposed normalization methods, comparisons were performed against un-normalized data on both datasets and four state-of-the-art methods, namely ZSN, MMADN, min–max value, and sigmoid normalization. ZSN is a mean and standard deviation-based normalization method that reduces the effect of outliers and dominate features. MMADN is similar to ZSN but it also uses the median to rescale the data. Hence, MMADN is more robust to outliers and domi­nant features due to the insensitivity of median values to such features. Min–max value-based methods linearly rescale data according to predened upper and lower bounds and are useful for preserving relationships between the original input and output data.
The sigmoid method applies nonlinear transformations on a dataset to reduce the effect of outliers [19]. Data distributions before and after application of proposed and current methods were compared on both datasets 1 and 2. The effect of the methods on feature importance was also compared across all methods on the two datasets. L1 regularization was used to computer feature importance with respect to the above-mentioned methods. The regularization method was adopted based on the observation made by the work in Ref. [40], which established that L1 regularization performed better than L2 regularization in a related experiment that included the two datasets used in this chapter.
Two DNN models were used to evaluate the effect of the proposed and current normalization methods on each dataset. The DNN models were adopted after ne­tuning. Classication of un-normalized data was also performed on both dataset as a basis of comparison of both the proposed and current methods. The DNN model used on dataset 1 had 8 nodes in each of the ve hidden layers and a sigmoid activa­tion function in the output layer. While both models used root mean square propaga­tion (RMSprop) [41] optimization algorithm, the rst model had its learning rate set to 0.0008. The DNN model used on dataset 2 had two hidden layers with 4096 nodes each and a learning rate set to 2e-5. RMSprop was used because of its ability to converge fast and it is easy to ne-tune [42].
M. Mulenga et al.
4 Results andDiscussions
The preceding segment expounded on the theoretical framework underpinning the suggested technique for expanding features, as well as its practical execution. The methodology for calculating supplementary features of the resulting dataset was outlined. The functions sqrt-sum, abs-diff, abs-shift, and sqrt-prod were each used to generate a synthetic feature for every input feature in the dataset. The new fea­tures were then combined with the original features of the input dataset to make an extended dataset. Thus, the proposed method transforms scalar data points in the dataset to vector representations.
1.0
0.
0.
0.
0.
un-normalized
Explainable AI forColorectal Cancer Classication
213
The proposed methods were evaluated on two datasets against existing methods, namely min–max, MMADN, ZSN, and sigmoid normalization methods. Customized methods such as sqrt-sum, abs-diff, abs-shift, and sqrt-prod, which were used to generate values in the vector representation of the extended dataset, are also used in some comparisons for the purpose of benchmarking. The previous section also detailed the DNN models that were used to evaluate the performance of the methods on each dataset. In this section, results of the proposed methods in comparison to the other methods are given. In particular, this section analyzes and discusses the effects of the methods on the properties of the underlying datasets and how they correlate with the classication performance of the DNN models.
To evaluate the performance of the feature extension method, which is also being referred to as the combined method, and other methods on the underlying properties (skewness, variance, and outliers) of the datasets, box and whisker plots were used. The box and whisker plots reveal the properties of un-normalized dataset, as well as its properties after the proposed and other current methods were applied. The prop­erties of dataset 1 before and after normalization are shown in Fig.1.
Based on Fig.1, it can be observed that while reducing the skewness of data, both the sqrt-sum and combined normalization methods are also able to reduce vari­ance more than the conventional normalization methods. However, the proposed methods unlike min–max and sigmoid did not diminish the outliers but instead increased their number. Similar to Pareto scaling, the sqrt-sum and combined method was able to push samples in a narrow range, which explains the observed reduction in variability. However, the slight increase in the number of outliers shows
8
0.6
4
2
0
Fig. 1 Box and whisker plots for a comparison of proposed and conventional normalization meth­ods on dataset 1
min-max
sqrt-sum
zsn
sigmoid
mmadn
combined
214
1.0
0.
0.
0.
0.
un-normalized
M. Mulenga et al.
that some samples did not scale well enough and ended up further way from the mean. The methods seemed to behave differently on dataset 2, where a comparison of both the proposed and current method is shown in Fig.2.
Figure 2 shows that both the sqrt-sum and combined normalization methods were able to eliminate outliers and readjust the underlying data to a distribution that is a nearly normal. Though the methods can be associated to a higher variance than the un-normalized samples, this behavior can also be observed in min–max and MMADN normalization methods. While a normal distribution can improve the per­formance of a statistical model such as a ML model, high variance negatively affects its performance. Feature dominance results associated with the sqrt-sum, combined, and the four conventional methods are listed in Table2.
From Table2 it can be observed that the combined method produced the highest number of dominant features in datasets 1 and 2, while the sqrt-sum method per­formed relatively well as compared to the conventional methods. A different DNN algorithm on each dataset was then used to evaluate the performance of the pro­posed and conventional methods. The classication results of the DNN model that was used on dataset 1 are shown in Fig.3.
Based on Fig.3, it can be observed that while some conventional methods per­form worse than un-normalized data, sqrt-sum normalization method has the best minimum, mean, and maximum AUC scores. While the performance of the DNN model on un-normalized data in terms of the minimum, mean, and maximum AUC scores are 0.8108, 0.8382, and 0.8593, respectively, the performance of the sqrt- sum method based on the same statistical values is 0.8929, 0.9126, and 0.9302,
8
0.6
4
2
0
Fig. 2 Box and whisker plots for a comparison of proposed and current normalization methods on dataset 2
min-max
sqrt-sum
zsn
sigmoid
mmadn
combined
AUC
Normalization methods
Explainable AI forColorectal Cancer Classication
215
Table 2 A comparison of the effect of proposed and conventional normalization methods on feature dominance
Dominant features
Dataset Method
Total featuresCount Percentage
1 Un-normalized 716 35.22 2033
Sqrt-sum 830 40.83 2033 Min–max 716 35.22 2033 ZSN 777 38.22 2033 Sigmoid 794 39.06 2033 MMADN 817 40.19 2033 Combined 3978 39.13 10,165
2 Un-normalized 137 40.90 335
Sqrt-sum 129 38.51 335 Min–max 137 40.90 335 ZSN 124 37.01 335 MMADN 131 39.10 335 Sigmoid 1 0.30 335 Combined 664 39.64 1675
0.92
0.90
0.88
0.86
0.84
0.82
0.80
0.78
zsn
min_max
sigmoid
combined
un-normalized
sqrt-sum
abs-shift
mmadn
Fig. 3 AUC scores of the un-normalized data, proposed and conventional methods on dataset 1
sqrt-prod
abs-diff
216
M. Mulenga et al.
respectively, which represents an approximate performance improvement of 7.4%. The combined normalization method follows in second place with minimum, mean, and maximum AUC scores at 0.8824, 0.9018, and 0.9170, respectively, representing about 6.4% performance improvement. The standard deviation AUC scores of the methods on dataset 1 were also compared as shown in Fig.4.
In Fig.4, it is observed that while the model on un-normalized data has a stan­dard deviation of 0.0137, the combined method had the best value of 0.0079, fol­lowed by 0.0087 of the sqrt-sum method. The observed performance improvements of the sqrt-sum and combined methods are correlated with the respective reduction in variance and increase in the number of relevant features shown in Fig.1 and Table2, respectively. The AUC scores obtained using the second DNN model on dataset 2, across un-normalized data, proposed and conventional normalization methods are shown in Fig.5.
From Fig.5, it can be observed that the sqrt-sum method has the highest mini­mum, mean, and maximum AUC values of 0.7298, 0.7583, 0.7808, respectively, which is about 9.1% higher than the performance of un-normalized data. In this case also, the combined method is in second place with 0.7011, 0.7395, and 0.7646 mini­mum, mean, and maximum AUC values, respectively, which is approximately 7.2% better than the un-normalized data. However, Fig.5 also shows that the difference between the combined and min–max method is very small. The model exhibits a signicantly superior performance on dataset 1 in comparison to dataset 2. One plausible explanation for this variation is that the latter dataset is approximately 50% smaller than the former. It is possible that the DNN employed on dataset 2 was not comprehensively optimized. Furthermore, the amalgamation of adenoma speci­mens and healthy specimens for the purpose of facile categorization within a binary classication framework could potentially compromise the integrity of the data. Fig.6 displays the standard deviation of the AUC scores for dataset 2, as obtained
Fig. 4 Standard deviation AUC scores of un-normalized data, proposed and current methods on dataset 1
AUC
Normalization methods
Explainable AI forColorectal Cancer Classication
0.75
0.70
0.65
0.60
0.55
0.50
zsn
mmadn
min_max
Fig. 5 AUC scores of the un-normalized data, proposed and conventional methods on dataset 2
sigmiod
un-normalized
abs-shift
combined
sqrt-prod
abs-diff
sqrt-sum
217
Fig. 6 Standard deviation of AUC scores of un-normalized data, proposed and conventional meth­ods on dataset 2
from the un-normalized data, as well as from the proposed and conventional methods.
From Fig.6, it can be observed that while the standard deviation of un- normalized data is 0.0244, the sqrt-sum normalization method has the best value at 0.0114 and is followed by the combined method at 0.0133. The good performance of the DNN model on dataset 2in relation to sqrt-sum normalization correlates to the normal
218
M. Mulenga et al.
distribution and elimination of outliers as observed in the samples associated with the method.
An additional evaluation was carried out to ascertain the inuence of the current and proposed normalization methods on the sensitivity and specicity of DNN models 1 and 2, utilizing datasets 1 and 2, respectively. The sensitivity scores of the deep neural network (DNN) models for each normalization method applied to data­sets 1 and 2 are presented in Table3.
From Table3, it can be observed that based on dataset 1, the sqrt-sum has the highest mean sensitivity value of 83%, which is about 6.4% better than un­normalized data. The combined method follows in second place with a sensitivity value of 82.3%, which is approximately 5.9% higher than un-normalized data. While sensitivity on dataset 2 was generally lower, the combined method had the highest mean value of 42.4%, which is approximately 10.3% better than un­normalized method. The sqrt-sum method, which was the second highest, had a mean sensitivity value of 36.1%, which was 4% better than the un-normalized data­set. In addition to the reasons given earlier, the performance of the model on dataset 2 was drastically lower than on dataset 1, partly because dataset 2 was highly imbal­anced with fewer CRC than non-CRC cases. Specicity scores of the two DNN models associated with the un-normalized dataset, proposed and conventional nor­malization methods on both datasets 1 and 2 are listed in Table4.
Table 4 shows that on dataset 1, the sqrt-sum method has the best performance of
86.4%, which is 4.5% higher than the un-normalized method. For the rst time, a closely related method to sqrt-sum, namely sqrt-prod, had the second highest speci­city value on the same dataset of approximately 85.7%, which is about 3.8% better than the un-normalized data. Again, on dataset 2, the sqrt-sum method has the high­est specicity value of approximately 94.8%, which is about 2.4% better than the un-normalized dataset. In this case also, sqrt-prod has the second highest specicity value of 93.7%, which is about 1.3% higher than the un-normalized dataset.
Table 3 Sensitivity of DNN models 1 and 2 for the proposed and current methods on datasets 1 and 2, respectively
Un normalized 0.7662 0.0198 0.7295 0.8187 0.3208 0.0361 0.2583 0.4417 Min–max 0.7788 0.0159 0.7534 0.8137 0.3117 0.0184 0.2750 0.3417 MMADN 0.7916 0.0152 0.7650 0.8197 0.1417 0.1119 0.0583 0.7167 Sigmoid 0.7293 0.0330 0.6585 0.7890 0.3258 0.0341 0.2667 0.4000 ZSN 0.7551 0.0221 0.7163 0.7978 0.3167 0.0297 0.2583 0.3833 Sqrt-sum 0.8300 0.0141 0.8050 0.8630 0.3608 0.0216 0.3083 0.4083 Abs-diff 0.8101 0.0188 0.7678 0.8512 0.3306 0.0333 0.2583 0.4083 Sqrt-prod 0.7978 0.0184 0.7589 0.8517 0.3614 0.0230 0.3167 0.4083 Abs-shift 0.7671 0.0204 0.7363 0.8154 0.2511 0.0189 0.2167 0.2833 Combined 0.8253 0.0158 0.7951 0.8607 0.4236 0.0338 0.3500 0.4917
Dataset 1 Dataset 2 Mean Std Min Max Mean Std Min Max