Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_117_библиотеки_им_акад_М_И_Перельмана
.pdf
Explainable AI forColorectal Cancer Classication
https://t.me/med1917
209
feature extension method partly uses customized normalization methods to transform each scalar data point in the dataset into a vector. Variations of the sqrt-sum
normalization method are also computed by applying random arithmetic computations to it. The experiments were written in Python and executed on a computer with
16,280MB GPU RAM.
3.1 Datasets
This study utilized two publicly accessible datasets of CRC microbiome, which
were collected from fecal samples. The rst dataset consists of shotgun sequencecurated samples and is based on the work done in Ref. [35]. The methodology for
accessing and manipulating the curated microbiome data was sourced from Ref.
[36]. Both the metadata, which includes demographic information, and the data le,
which contains microbial counts per sample, were obtained. The dataset comprised
of 884 samples and 2031 features, representing a collection of samples obtained
from experiments conducted across six distinct countries, namely the United States,
Germany, Austria, Italy, Canada, and France. The dataset was expanded to include
a total of 2033 features by incorporating microbial count data along with two demographic features, namely BMI and age. Data augmentation was performed due to
the close association of both attributes with the development of CRC [37]. The classication target label utilized in the study was based on the demographic attribute of
study condition. Table1 presents the distribution of samples with respect to colorec-
tal cancer (CRC) and healthy controls across various countries.
Table 1 shows that there were 368 CRC samples and 516 controls in a dataset of
884 combined samples. The disease attribute from metadata was used to lter the
dataset. All samples that had an undened value in the disease eld were excluded
from the dataset, which reduced the total number of samples to 796. The second
dataset consists of 16S rRNA sequence data samples based on a study in Ref. [38].
The dataset, which has a total of 490 samples and 336 features, consists of 172, 198,
and 120 that are normal, adenoma, and cancer samples, respectively. Part of the
Table 1 Composition of colorectal cancer-based microbiome samples from six different countries
Country
The United States 77 87 164
Germany 76 10 86
Austria 46 108 154
France 106 206 312
Italy 61 79 140
Canada 2 26 28
Total 368 516 884
Samples
TotalCRC Non-CRC

210
M
M
+
√
µ
σ
ij ij
M
,,
=+σ
https://t.me/med1917
M. Mulenga et al.
samples were collected in Toronto, Canada, while the rest were collected in Boston,
Houston, and Ann Arbor in the United States.
3.2 Proposed Methods
This chapter proposes two methods, namely square-root sum and the combined
method. In tandem with the XAI approach, the performance of the proposed model
is described using underlying properties of the data by underscoring how the various method adjustments in the experiments affect the underlying data properties. As
such the rst customized data normalization method proposed in this chapter is
related to a method called Pareto scaling [34], which is computed as
ij i
,
=
ij
,
where M, σ,μ, i, and j represent respective feature value, standard deviation, mean,
row number, and column number, respectively, in the dataset.
Pareto scaling stems from the concept called Pareto optimality, which seeks to
make an individual better off without making another individual of the same group
worse off [34]. The method improves the representation of features that have lower
values while reducing the effect of noise in data. This method keeps the structure of
the data almost intact. Though the method is able to reduce the effect of outliers in
the dataset, it is not robust to dominant features [19]. As such, in this chapter we
propose a method that attempts to reduce the effect of both outliers and dominant
features. The method transforms a dataset by adding the standard deviation of the
raw data to each of its data points and then computes the square root of the new data
points as
(1)
For the sake of simplicity, we will refer to the parameter ω as sqrt-sum. Adding
the standard deviation to a feature value in the sqrt-sum method is meant to shift the
feature space in order to reduce the effect of outliers. Computation of the square
root is meant to marginally reduce the inuence of dominant features, while slightly
reducing the effect of values that have a lower effect. It is evident that the proposed
method does not use Pareto scaling in its entirety, but only applies part of the
method. The sqrt-sum method is only applied on nonzero values to avoid introducing an unintended effect on the dataset since zeros in sequence data do not only
represent counts, but they also represent missing values. The occurrence of missing
values is a prevalent phenomenon in datasets pertaining to health-related
research [39].
The second method called feature extension seeks to improve the performance of
a DNN model by adjusting feature relevance in the dataset. Several normalization
(2)

f xxxxxx
()
=
()
x x1=
x2=ω
x x4=σ
Explainable AI forColorectal Cancer Classication
https://t.me/med1917
211
methods comprising sqrt-sum and its variants obtained by applying random computations used to generate additional features in the dataset as
,,,
12345
,
(3)
where x represents the input scalar parameter and x1, x2, x3,x4, andx5 represent the
output vector parameters computed based on the sqrt-sum method and its variants,
which are also called the customized normalization methods.
Therefore, the output of the feature extension method represents individual outputs of various normalization functions. The output of the customized normalization methods is also combined with features of the raw input dataset such that the
rst parameter of output vector computed is equal to the original input scalar value
and was computed as
(4)
To compute the second vector parameter, use the sqrt-sum method as
(5)
As mentioned above, the rest of the parameters are random variations of the sqrt-
sum methods and that of the original input scalar value shown in Eqs. (2) and (4),
respectively. Therefore, the third parameter of the vector, which we will simply
refer to as absolute value of the difference (abs-diff), was computed using the
parametermax(represents the highest value in the raw dataset) as
max2xx−
∣
=
3
(6)
The fourth parameter of the vector, which we will refer to as square root of the
product (sqrt-prod), was computed as the square root of the product of the input
value and the standard deviation of the raw data as
The fth parameter of the vector, which we will refer to as absolute shift (abs-
shift), was calculated using the absolute value of the difference between the input
parameter and the second parameter of the vector as
While sqrt-sum, sqrt-prod computed according to Eqs. (2) and (7) are purposefully designed to reduce the effect of dominant features, the parameters abs-diff and
abs-shift are random computations that are only meant to produce noisy replicates
of other methods. Knights etal. [23] and Lo and Marculescu [24] used the concept
x2xx−=∣
5
2
(7)
(8)

212
https://t.me/med1917
of noisy replicates to improve classication accuracy of ML models. The main
objective of using noisy replicates in our method is to increase the number of features based on computations that have outputs that lie within the same feature space.
To evaluate the proposed normalization methods, comparisons were performed
against un-normalized data on both datasets and four state-of-the-art methods,
namely ZSN, MMADN, min–max value, and sigmoid normalization. ZSN is a
mean and standard deviation-based normalization method that reduces the effect of
outliers and dominate features. MMADN is similar to ZSN but it also uses the
median to rescale the data. Hence, MMADN is more robust to outliers and dominant features due to the insensitivity of median values to such features. Min–max
value-based methods linearly rescale data according to predened upper and lower
bounds and are useful for preserving relationships between the original input and
output data.
The sigmoid method applies nonlinear transformations on a dataset to reduce the
effect of outliers [19]. Data distributions before and after application of proposed
and current methods were compared on both datasets 1 and 2. The effect of the
methods on feature importance was also compared across all methods on the two
datasets. L1 regularization was used to computer feature importance with respect to
the above-mentioned methods. The regularization method was adopted based on the
observation made by the work in Ref. [40], which established that L1 regularization
performed better than L2 regularization in a related experiment that included the
two datasets used in this chapter.
Two DNN models were used to evaluate the effect of the proposed and current
normalization methods on each dataset. The DNN models were adopted after netuning. Classication of un-normalized data was also performed on both dataset as
a basis of comparison of both the proposed and current methods. The DNN model
used on dataset 1 had 8 nodes in each of the ve hidden layers and a sigmoid activation function in the output layer. While both models used root mean square propagation (RMSprop) [41] optimization algorithm, the rst model had its learning rate set
to 0.0008. The DNN model used on dataset 2 had two hidden layers with 4096
nodes each and a learning rate set to 2e-5. RMSprop was used because of its ability
to converge fast and it is easy to ne-tune [42].
M. Mulenga et al.
4 Results andDiscussions
The preceding segment expounded on the theoretical framework underpinning the
suggested technique for expanding features, as well as its practical execution. The
methodology for calculating supplementary features of the resulting dataset was
outlined. The functions sqrt-sum, abs-diff, abs-shift, and sqrt-prod were each used
to generate a synthetic feature for every input feature in the dataset. The new features were then combined with the original features of the input dataset to make an
extended dataset. Thus, the proposed method transforms scalar data points in the
dataset to vector representations.

1.0
0.
0.
0.
0.
un-normalized
Explainable AI forColorectal Cancer Classication
https://t.me/med1917
213
The proposed methods were evaluated on two datasets against existing methods,
namely min–max, MMADN, ZSN, and sigmoid normalization methods. Customized
methods such as sqrt-sum, abs-diff, abs-shift, and sqrt-prod, which were used to
generate values in the vector representation of the extended dataset, are also used in
some comparisons for the purpose of benchmarking. The previous section also
detailed the DNN models that were used to evaluate the performance of the methods
on each dataset. In this section, results of the proposed methods in comparison to
the other methods are given. In particular, this section analyzes and discusses the
effects of the methods on the properties of the underlying datasets and how they
correlate with the classication performance of the DNN models.
To evaluate the performance of the feature extension method, which is also being
referred to as the combined method, and other methods on the underlying properties
(skewness, variance, and outliers) of the datasets, box and whisker plots were used.
The box and whisker plots reveal the properties of un-normalized dataset, as well as
its properties after the proposed and other current methods were applied. The properties of dataset 1 before and after normalization are shown in Fig.1.
Based on Fig.1, it can be observed that while reducing the skewness of data,
both the sqrt-sum and combined normalization methods are also able to reduce variance more than the conventional normalization methods. However, the proposed
methods unlike min–max and sigmoid did not diminish the outliers but instead
increased their number. Similar to Pareto scaling, the sqrt-sum and combined
method was able to push samples in a narrow range, which explains the observed
reduction in variability. However, the slight increase in the number of outliers shows
8
0.6
4
2
0
Fig. 1 Box and whisker plots for a comparison of proposed and conventional normalization methods on dataset 1
min-max
sqrt-sum
zsn
sigmoid
mmadn
combined

214
1.0
0.
0.
0.
0.
un-normalized
https://t.me/med1917
M. Mulenga et al.
that some samples did not scale well enough and ended up further way from the
mean. The methods seemed to behave differently on dataset 2, where a comparison
of both the proposed and current method is shown in Fig.2.
Figure 2 shows that both the sqrt-sum and combined normalization methods
were able to eliminate outliers and readjust the underlying data to a distribution that
is a nearly normal. Though the methods can be associated to a higher variance than
the un-normalized samples, this behavior can also be observed in min–max and
MMADN normalization methods. While a normal distribution can improve the performance of a statistical model such as a ML model, high variance negatively affects
its performance. Feature dominance results associated with the sqrt-sum, combined,
and the four conventional methods are listed in Table2.
From Table2 it can be observed that the combined method produced the highest
number of dominant features in datasets 1 and 2, while the sqrt-sum method performed relatively well as compared to the conventional methods. A different DNN
algorithm on each dataset was then used to evaluate the performance of the proposed and conventional methods. The classication results of the DNN model that
was used on dataset 1 are shown in Fig.3.
Based on Fig.3, it can be observed that while some conventional methods perform worse than un-normalized data, sqrt-sum normalization method has the best
minimum, mean, and maximum AUC scores. While the performance of the DNN
model on un-normalized data in terms of the minimum, mean, and maximum AUC
scores are 0.8108, 0.8382, and 0.8593, respectively, the performance of the sqrt-
sum method based on the same statistical values is 0.8929, 0.9126, and 0.9302,
8
0.6
4
2
0
Fig. 2 Box and whisker plots for a comparison of proposed and current normalization methods on
dataset 2
min-max
sqrt-sum
zsn
sigmoid
mmadn
combined

AUC
Normalization methods
Explainable AI forColorectal Cancer Classication
https://t.me/med1917
215
Table 2 A comparison of the effect of proposed and conventional normalization methods on
feature dominance
Dominant features
Dataset Method
Total featuresCount Percentage
1 Un-normalized 716 35.22 2033
Sqrt-sum 830 40.83 2033
Min–max 716 35.22 2033
ZSN 777 38.22 2033
Sigmoid 794 39.06 2033
MMADN 817 40.19 2033
Combined 3978 39.13 10,165
2 Un-normalized 137 40.90 335
Sqrt-sum 129 38.51 335
Min–max 137 40.90 335
ZSN 124 37.01 335
MMADN 131 39.10 335
Sigmoid 1 0.30 335
Combined 664 39.64 1675
0.92
0.90
0.88
0.86
0.84
0.82
0.80
0.78
zsn
min_max
sigmoid
combined
un-normalized
sqrt-sum
abs-shift
mmadn
Fig. 3 AUC scores of the un-normalized data, proposed and conventional methods on dataset 1
sqrt-prod
abs-diff

216
https://t.me/med1917
M. Mulenga et al.
respectively, which represents an approximate performance improvement of 7.4%.
The combined normalization method follows in second place with minimum, mean,
and maximum AUC scores at 0.8824, 0.9018, and 0.9170, respectively, representing
about 6.4% performance improvement. The standard deviation AUC scores of the
methods on dataset 1 were also compared as shown in Fig.4.
In Fig.4, it is observed that while the model on un-normalized data has a standard deviation of 0.0137, the combined method had the best value of 0.0079, followed by 0.0087 of the sqrt-sum method. The observed performance improvements
of the sqrt-sum and combined methods are correlated with the respective reduction
in variance and increase in the number of relevant features shown in Fig.1 and
Table2, respectively. The AUC scores obtained using the second DNN model on
dataset 2, across un-normalized data, proposed and conventional normalization
methods are shown in Fig.5.
From Fig.5, it can be observed that the sqrt-sum method has the highest minimum, mean, and maximum AUC values of 0.7298, 0.7583, 0.7808, respectively,
which is about 9.1% higher than the performance of un-normalized data. In this case
also, the combined method is in second place with 0.7011, 0.7395, and 0.7646 minimum, mean, and maximum AUC values, respectively, which is approximately 7.2%
better than the un-normalized data. However, Fig.5 also shows that the difference
between the combined and min–max method is very small. The model exhibits a
signicantly superior performance on dataset 1 in comparison to dataset 2. One
plausible explanation for this variation is that the latter dataset is approximately
50% smaller than the former. It is possible that the DNN employed on dataset 2 was
not comprehensively optimized. Furthermore, the amalgamation of adenoma specimens and healthy specimens for the purpose of facile categorization within a binary
classication framework could potentially compromise the integrity of the data.
Fig.6 displays the standard deviation of the AUC scores for dataset 2, as obtained
Fig. 4 Standard deviation AUC scores of un-normalized data, proposed and current methods on
dataset 1

AUC
Normalization methods
Explainable AI forColorectal Cancer Classication
https://t.me/med1917
0.75
0.70
0.65
0.60
0.55
0.50
zsn
mmadn
min_max
Fig. 5 AUC scores of the un-normalized data, proposed and conventional methods on dataset 2
sigmiod
un-normalized
abs-shift
combined
sqrt-prod
abs-diff
sqrt-sum
217
Fig. 6 Standard deviation of AUC scores of un-normalized data, proposed and conventional methods on dataset 2
from the un-normalized data, as well as from the proposed and conventional
methods.
From Fig.6, it can be observed that while the standard deviation of un- normalized
data is 0.0244, the sqrt-sum normalization method has the best value at 0.0114 and
is followed by the combined method at 0.0133. The good performance of the DNN
model on dataset 2in relation to sqrt-sum normalization correlates to the normal

218
https://t.me/med1917
M. Mulenga et al.
distribution and elimination of outliers as observed in the samples associated with
the method.
An additional evaluation was carried out to ascertain the inuence of the current
and proposed normalization methods on the sensitivity and specicity of DNN
models 1 and 2, utilizing datasets 1 and 2, respectively. The sensitivity scores of the
deep neural network (DNN) models for each normalization method applied to datasets 1 and 2 are presented in Table3.
From Table3, it can be observed that based on dataset 1, the sqrt-sum has the
highest mean sensitivity value of 83%, which is about 6.4% better than unnormalized data. The combined method follows in second place with a sensitivity
value of 82.3%, which is approximately 5.9% higher than un-normalized data.
While sensitivity on dataset 2 was generally lower, the combined method had the
highest mean value of 42.4%, which is approximately 10.3% better than unnormalized method. The sqrt-sum method, which was the second highest, had a
mean sensitivity value of 36.1%, which was 4% better than the un-normalized dataset. In addition to the reasons given earlier, the performance of the model on dataset
2 was drastically lower than on dataset 1, partly because dataset 2 was highly imbalanced with fewer CRC than non-CRC cases. Specicity scores of the two DNN
models associated with the un-normalized dataset, proposed and conventional normalization methods on both datasets 1 and 2 are listed in Table4.
Table 4 shows that on dataset 1, the sqrt-sum method has the best performance of
86.4%, which is 4.5% higher than the un-normalized method. For the rst time, a
closely related method to sqrt-sum, namely sqrt-prod, had the second highest specicity value on the same dataset of approximately 85.7%, which is about 3.8% better
than the un-normalized data. Again, on dataset 2, the sqrt-sum method has the highest specicity value of approximately 94.8%, which is about 2.4% better than the
un-normalized dataset. In this case also, sqrt-prod has the second highest specicity
value of 93.7%, which is about 1.3% higher than the un-normalized dataset.
Table 3 Sensitivity of DNN models 1 and 2 for the proposed and current methods on datasets 1
and 2, respectively
Un normalized 0.7662 0.0198 0.7295 0.8187 0.3208 0.0361 0.2583 0.4417
Min–max 0.7788 0.0159 0.7534 0.8137 0.3117 0.0184 0.2750 0.3417
MMADN 0.7916 0.0152 0.7650 0.8197 0.1417 0.1119 0.0583 0.7167
Sigmoid 0.7293 0.0330 0.6585 0.7890 0.3258 0.0341 0.2667 0.4000
ZSN 0.7551 0.0221 0.7163 0.7978 0.3167 0.0297 0.2583 0.3833
Sqrt-sum 0.8300 0.0141 0.8050 0.8630 0.3608 0.0216 0.3083 0.4083
Abs-diff 0.8101 0.0188 0.7678 0.8512 0.3306 0.0333 0.2583 0.4083
Sqrt-prod 0.7978 0.0184 0.7589 0.8517 0.3614 0.0230 0.3167 0.4083
Abs-shift 0.7671 0.0204 0.7363 0.8154 0.2511 0.0189 0.2167 0.2833
Combined 0.8253 0.0158 0.7951 0.8607 0.4236 0.0338 0.3500 0.4917
Dataset 1 Dataset 2
Mean Std Min Max Mean Std Min Max
Соседние файлы в папке Библиотека им академика М.И. Перельмана
