Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_117_библиотеки_им_акад_М_И_Перельмана
.pdf
()=() ()
()()
′′
,,
ÆÆ
n
j
n
ii
()
=
1
x
i
′
y
xy
c
xy
∂
1
F
k
c
SFM
k
=
∑
Explainable AI inDisease Diagnosis
https://t.me/med1917
95
Partial Dependence Plot (PDP) [26] PDP is a model-agnostic XAI method that
generates global explanations. It is suitable for evaluating the mean marginal effect
of the input attributes on the predictions. The PDP visualises the effect of the
selected feature values on the outcome by controlling all other features. These
effects are analysed over different values, such as the range between the minimum
and maximum values. A graph plots the average values of the obtained results. The
partial dependence function is given in Eq. (4).
ÆÆÆ
f xExfxx fxxPx
iiii ii i
′′
=∫
d
(4)
After simplifying, we obtain
=
()
i
1
∑
f x
The xi is the selected feature whose importance is being measured, and
fxx
j
,
′
(5)
are the
rest of the features in the dataset. PDP is purposely used to analyse any variable’s
effect on the prediction. However, these variables are kept singular or in a combination of two for simplication.
Gradient Class Activation Map (Grad CAM) [23] Deep Learning, mainly CNN,
performs adequately in computer vision tasks. Explanations for CNN predictions
can help to simulate human intelligence and understand the model better. Grad
CAM utilises weights and feature maps in the CNN model to nd the activated part
of an image that contributes to the corresponding label.
Grad CAM replaces the fully connected layer of the CNN model with global
average pooling.
c
F
=
∑∑
k
N
k
M
∂
(6)
Here, yc, N, andM denote the score obtained for the class ‘c’ before the softmax
function, the number of pixels and the feature map, respectively. The obtained
is
the output of the global average pooling of the gradients. It also represents the
weight (importance) of kth feature map for the class ‘c’.
Sc is obtained by a weighted combination of feature maps and applying the ReLU
function to get the positively inuencing features on the target class ‘c’.
The next section describes the applications of XAI in disease diagnosis.
c
kck
(7)

96
https://t.me/med1917
P. Bedi et al.
4 XAI inDisease Diagnosis
AI has produced advanced healthcare systems and automated various health-related
tasks such as cancer prediction [28], medical chatbots, arrhythmia detection [28],
early diagnosis of fatal blood disease [16], treatment of rare diseases and drug
repurposing [31], personalised treatment [32], digital storage and management of
medical records, and many other tasks. XAI enhances healthcare tasks and understandability by providing a rationale for the AI-generated recommendations.
Disease diagnosis is one of the most popular research areas in the medicinal
domain. Some of the XAI methods applied within the disease diagnosis systems are
given below:
Grad CAM in Arrhythmia Detection An electrocardiogram (ECG) is one of the
tests performed to detect arrhythmia. XAI can be used to explain arrhythmia based
on ECG.Y. Y.Jo etal. [28] introduced an Explainable Deep Learning Model (XDM)
for arrhythmia classication using a neural network-backed ensemble tree approach.
The XDM has six deep learning feature modules to detect arrhythmia characteristics such as irregular R wave patterns, regular PR intervals, the correlation between
P and R wave etc., and an XDM ensemble approach for arrhythmic classication.
The output of each feature module gave the probability of each arrhythmic characteristic, and the nal ensemble module was the multilayer perceptron for the
classication. The study [28] used the Seojong ECG dataset1 which consists of data
from Mediplex Seojong Hospital and Seojong General Hospital. The arrhythmia
was classied into nine classes based on ECG signals, namely, ventricular tachycardia (VT), junctional rhythm (JR), normal sinus rhythm (NSR), pacemaker rhythm
(PM), supraventricular tachycardia (SVT), atrial brillation (AF), complete atrioventricular block (CAVB), second-degree atrioventricular block Mobitz type II
(2AVB-T2), and second-degree atrioventricular block Mobitz type I with
Wenckebach phenomenon (2AVB-T1). The aggregated precision, recall, and
F1-score for the nine classes were found to be 0.936, 0.793, and 0.852, respectively.
The Area Under the Receiver Operating Characteristic (AUROC) curve values of
PM, NSR, AF, JR, SVT, VT, CAVB, 2AVB_T2, and 2AVB_T1 were found to be
0.988, 0.994, 0.976, 0.993, 0.992, 0.921, 0.976, 0.983, and 0.961, respectively.
Gradient-class activation map (sensitivity map) and the guided gradient backpropagation methods were used for the visualisation of XDM interpretations. First-order
gradients of the classier probabilities were used to create the sensitivity map,
which illustrates the sensitive regions responsible for the decision.
PDP and Accumulated Proles in Cancer Prediction Cancer is a deadly disease
which is responsible for uncountable deaths worldwide. Early detection of the disease is the only solution to stop it from further spreading in the body as it becomes
incurable in later stages. AI can predict the disease based on various types of patient
1
Dataset available at request to corresponding author.

Explainable AI inDisease Diagnosis
https://t.me/med1917
97
data such as genomic, proteomic, or clinical data. XAI provides additional support
to AI-based disease diagnosis models by identifying the subset of input features
responsible for cancer from the given patient input feature set and presenting it in a
human-readable form. Z.U. Ahmed etal. [17] proposed an ensemble approach to
explore and analyse the effect of various risk factors on lung and bronchus cancer
(LBC) mortality rate. Twelve risk factors including smoking, sulphur dioxide, and
poverty were considered in the study [17]. These risk factors were extracted from
different sources. National Vital Statistics System at the National Center for Health
Statistics of the Centers for Disease Control and Prevention2 was used to extract the
LBC mortality rate, County Health Ranking3 and Institute for Health Metrics and
Evaluation4 were used to extract cigarette smoking prevalence data, and from US
Census Bureau’s Small Area Income and Poverty Estimates programme5 poverty
data was extracted. The study [17] trained three spatial regression models (Spatial
lag, Spatial error, and Geographically Weighted OLS Regression (GW-OLS)), ve
base learner models (Generalised Linear Model (GLM), Random Forest (RF),
Gradient Boosting Machine (GBM), Extreme Gradient Boosting (XGBoost), Deep
Neural Network (DNN)), and two stack-ensemble models (best base learners and all
base learners) to analyse the risk factors of LBC.The performance of models was
compared using Mean Absolute Error (MAE) and Root Mean Squared Error
(RMSE) metrics. The MAE with Spatial lag, Spatial error, GW-OLS, GLM, RF,
GBM, XGBoost, DNN, ensemble of best base learners and ensemble of all base
learners on testing data were obtained as 9.02, 6.53, 6.20, 6.49, 6.16, 6.20, 6.41,
7.00, 6.08, and 6.06, respectively. The RMSE with Spatial lag, Spatial error,
GW-OLS, GLM, RF, GBM, XGBoost, DNN, ensemble of best base learners and
ensemble of all base learners on testing data were obtained as 11.46, 8.35, 8.09,
8.31, 8.03, 8.06, 8.35, 9.03, 7.95, and 7.74, respectively. The study showed that the
stack-ensemble approaches were more accurate. However, the stack-ensembled
approaches are complex in design as they stack multiple models to generate predictions. The authors in [17] used Permutation-based feature importance XAI methods
to visualise the effect of input on the outcome. These methods included partial
dependence (PD), local dependence (LD), and accumulated local (AL) proles.
XAI methods were also visually represented with plots. The study using XAI methods showed that smoking is the most critical risk factor responsible for lung and
bronchus cancer.
LIME in Early Diagnosis of Fatal Blood Disease Leukaemia is one of the fatal
blood diseases which takes numerous lives annually. It has a high mortality rate
because it is a fast-acting disease. Therefore, the diagnosis in the early stages can
increase the patient’s survival chances. AI-based systems can help detect leukaemia
2
https://www.cdc.gov/nchs/index.html
3
https://www.countyhealthrankings.org/explore-health-rankings
4
https://www.healthdata.org/
5
https://www.census.gov/programs-surveys/saipe.html

98
https://t.me/med1917
P. Bedi et al.
from the patient’s lab results, and XAI can help doctors to analyse the predictions
generated by AI-based systems. In [16], the authors developed an AI system which
uses Transfer Learning to detect the presence/absence of leukaemia from the microscopic images of blood cells. The dataset in [16] was taken from the Kaggle repository named Leukemia Classication,6 containing 15,135 images. Pre-trained CNN
models, ResNet101V2, VGG19, InceptionResNetV2, and InceptionV3, were used
in the diagnosis because of their excellent performance in computer vision tasks.
Accuracy and F1-score were measured to compare the performance of CNN models. InceptionResNetV2 achieved the best results with 0.8002 accuracy and 0.7980
F1-score with testing data whereas InceptionV3 performed best with 0.9665 accuracy and 0.9665 F1-score with validation data. LIME was applied to the predictions
generated by the InceptionV3 model. It highlighted the parts of the image responsible for the model’s predictions.
Multi-Scale CAM in Glaucoma Diagnosis A clinically interpretable deep
learning- based framework, namely EAMNet, was proposed in [33] to diagnose
glaucoma disease. The framework consisted of a CNN network, Multi-layer Average
Pooling, and Evidence Activation Mapping to diagnose glaucoma from retinal fundus images and highlight the affected regions in the images. Distinct regions contributing to the predicted diagnosis were highlighted using the convolution feature
maps that were obtained during the training of the ResBlock CNN model. The
model concatenated the feature maps at different levels by global averaging to get
an activation map. This map emphasises the areas (such as notch on the neuroretinal rim, defect or bleeding on optic disc) on the fundus images that provide
concrete evidence for diagnosing glaucoma. The experiments in [33] were performed with the ORIGA public dataset [34]. EAMNet achieved a 0.88 AUROC
value for diagnosis of glaucoma disease.
SHAP in Parkinson’s Disease XAI methods are not restricted to generating explanations for black-box models only. They have also been used in feature selection.
SHAP was used for feature selection in the Parkinson’s disease diagnosis in [35].
Parkinson’s disease is a neurodegenerative disease in which the patient’s data is collected using force sensors on their feet or voice presentation. It thus, produces a high
feature dimensionality for the disease diagnosis. High-dimensional features may
produce undesirable outcomes such as redundant features, reduce the computational
efciency or generate biased results. Hence, in [35], the authors selected the important features from the high-dimensional feature set of Parkinson’s disease using
Fscore, Anova-F, Mutual Information (MI), and SHAP value. The dataset for the
study [35] was taken from the UCI repository, Parkinson’s Disease Classication.7
The dataset contains 188 patients’ records. A comparative analysis of four feature
selection methods, Fscore, Anova-F, MI, and SHAP, was performed on this dataset.
6
https://www.kaggle.com/datasets/andrewmvd/leukemia-classication
7
https://archive.ics.uci.edu/dataset/470/parkinson+s+disease+classication

Explainable AI inDisease Diagnosis
https://t.me/med1917
99
The most important features extracted by each feature selection method were given
as input to the classication algorithm. Four ML algorithms, RF, deep forest, Light
GBM, and XGBoost, were used for the classication of Parkinson’s disease. Among
the four feature selection methods, SHAP showed the best performance in
Parkinson’s disease diagnosis with four ML models. SHAP-gcForest achieved the
highest accuracy and F1-score of 91.78% and 0.945, respectively. SHAP-XGBoost,
SHAP-LightGBM, and SHAP-RF produced an accuracy of 90.04%, 91.62%,
90.34% and F1-score of 0.936, 0.945, 0.939, respectively.
The following section shows the implementation of XAI post-hoc methods that
generate explanations with ML/DL models.
5 Case Studies ofXAI Post-hoc Methods
This section demonstrates the implementation of XAI post-hoc methods by generating disease diagnosis explanations on the predictions produced by the ML/DL models. This has been exemplied below using two case studies carried out on the
datasets of COVID-19 and Breast cancer diseases.
Case Study 1: Explanations to Support COVID-19 Diagnosis Using Chest
Radiography Images
This case study explains the step-wise implementation of three XAI post-hoc methods (LIME, SHAP, and Grad-CAM) applied to the chest radiography dataset. The
chest radiography dataset was retrieved from the Kaggle repository.8 Model-based
predictions for COVID-19 diagnosis must be performed as the rst step to generate
explanations using XAI post-hoc methods. This is briey explained below.
Model-Based Prediction The chest radiography dataset had 13,808 images, out of
which 3616 were of patients diagnosed with COVID-19, and the rest of the 10,192
images were of patients without the disease. The size of each image was
299×299×3 pixels. However, the size of the images was reduced from 299×299×3
pixels to 100×100×3 pixels using the opencv9 python library to minimise space
utilisation and maximise the processing time. The dataset was divided into training
and testing samples in a 75:25 ratio. A CNN architecture-based, InceptionV3 model
[36] pre-trained on a large ImageNet [37] dataset was used to generate predictions.
The pre-trained weights of the InceptionV3 model were leveraged for transfer learning [38] and to build the model to diagnose COVID-19 disease. To do so, the last
layer was replaced with two dense layers of 1024 and 2 neurons, respectively. This
model was trained on the training samples to generate predictions on the test samples. The model predicted 90% of cases (normal or COVID) correctly. Figure3
shows the confusion matrix obtained from the test data.
8
https://www.kaggle.com/datasets/tawsifurrahman/covid19-radiography-database
9
https://docs.opencv.org/4.x/

100
https://t.me/med1917
P. Bedi et al.
Explanation Generation Three post-hoc XAI methods, LIME, SHAP, and GradCAM, were implemented to generate explanations for the InceptionV3 predictions.
The three subsections given below describe the implementation of the three XAI
methods to generate explanations on the chest radiography images.
Explanations Using LIME LIME highlights regions of interest (ROI) of an image
that either positively or negatively contributes to the model prediction. An image
explainer object, LimeImageExplainer was built using ‘lime’ python library, to
highlight ROI in chest image. The trained model and a chest image were passed as
input parameters to the explainer object of LIME.The explainer object has a hyperparameter, ‘sample_size’ which denes the number of neighbour samples used to
approximate the model predictions. This was set to 500. The images with highlighted ROI as generated by LIME are shown in Fig.4. The green area depicts a
positive contribution to the predicted class and the red area depicts a negative contribution. The lung features that contribute to the correct prediction are shown in
Fig.4a. These features were extracted from the chest radiography image using the
InceptionV3 model. Figure 4b shows an incorrectly predicted image where the
patient not having COVID-19 (normal image) was misdiagnosed as COVID-19
positive. It is so because the model could not perform feature extraction correctly,
Fig. 3 Confusion matrix of COVID-19 diagnosis using InceptionV3 model

Explainable AI inDisease Diagnosis
https://t.me/med1917
101
and the sternum part contributed to the prediction instead of the lungs. Thus, in such
a case, the XAI helped to cross-check the model prediction.
Explanations Using SHAP The SHAP method uses a trained model and an image
to highlight the pixels contributing to the prediction (COVID-19 diagnosis or normal healthy patient). These contributions are determined by computing Shapley
values for each pixel. Shapley values are generated by a shap-explainer object using
a partition method. In Fig.5, the rst image is the original chest image, and the other
two images (normal and COVID) are the SHAP-generated images. The Shapley
images contain light and deep shades of red and blue colour pixels. The colour
becomes more intense while moving away from (Shapley value) zero and becomes
lighter towards zero, as shown in the SHAP scale in Fig.5. The larger area covered
by the positive or negative Shapley values shows a positive or negative contribution
to the predicted outcome. More red pixels (positive Shapley values) than blue pixels
in Fig.5 (normal image) show that more pixels are contributing to ‘normal’ (not
having COVID-19) prediction. Similarly, in the COVID-19 image of Fig.5, there
are more blue pixels (negative Shapley values), showing that the image is COVID-19
negative.
Explanations Using Grad-CAM Grad-CAM method takes a pre-trained model,
an input chest image, and a feature map (the output of a pre-trained model layer) to
visualise the reasons for the model-based predictions. This implementation uses
disease diagnosis (having COVID-19 or not) prediction generated by the InceptionV3
model. Grad-CAM generates visual explanations in the form of a heatmap. The
feature map is superimposed over the input chest image to generate a heatmap. It is
Fig. 4 LIME explanation. (a) Correct feature selection contributing to right prediction. (b)
Incorrect feature selection contributing to wrong prediction

102
https://t.me/med1917
Fig. 5 Shapley values generated by SHAP method for normal and COVID classes
P. Bedi et al.
done so because a feature map contains the most important features (parts) of an
input image extracted after applying convolutions. Hence the superimposed heatmap highlights the important portions of an input image, which are responsible for
the resulting prediction. It uses VIBGYOR colours to highlight the contributing
features in the image. The heatmap shows the contribution of each pixel to the
model-based prediction where pixels in red and violet are the most and least contributing features, respectively. Figure 6 depicts the heatmap obtained with the
InceptionV3 model and the feature map of InceptionV3’s ‘mixed10’ layer. In
Fig.6a, InceptionV3 correctly predicted the X-ray image as Normal (not having
COVID-19). The pixels of the neck area in the chest X-ray image, represented in
green colour, show a lesser contribution to the diagnosis prediction. The lower part
(lungs), red in colour, shows a more signicant contribution in detecting the absence
of COVID-19. In Fig.6b, the actual (true) diagnosis was COVID-19 but InceptionV3
wrongly predicted the image as normal. It happened because the model did not
extract the features correctly, which is evident from the Grad-CAM produced image.
The image contains a large blue area (full right lung and half of left lung), representing a lesser contribution to the model-generated prediction.
Case Study 2: Explanations to Support Cancer Classication (Benign and
Malignant) Using Wisconsin’s Diagnostic Dataset
This case study discusses the implementation of three XAI post-hoc methods,
namely, LIME, PDP, and SHAP applied to the Wisconsin Dataset.10 It is an openly
accessible dataset archived in the UCI Machine Learning repository. The dataset
contains characteristics of breast mass nuclei obtained from digitised breast images.
The characteristics include mean texture, radius error, worst area, worst texture, etc.
10
https://archive.ics.uci.edu/ml/datasets/breast+cancer+wisconsin+(diagnostic)

Explainable AI inDisease Diagnosis
https://t.me/med1917
Fig. 6 Grad-CAM explanations. (a) Correct feature selection contributing to right prediction. (b)
Incorrect feature selection contributing to wrong prediction
103
The explanations of XAI methods were produced based on the model-generated
predictions. Therefore, rstly the model-based prediction is discussed briey.
Model-Based Prediction The dataset tabulated 569 instances and 32 characteristics. The characteristics had two types of values, numerical and categorical. Strings
represented some categorical values. For example, the characteristic ‘type of cancerous cell’ had two values, ‘B’ and ‘M’, where ‘B’ represented Benign and ‘M’
indicated Malignant. Such categorical values were mapped to numerical labels
using ‘scikit-learn’ python library. Next, the dataset was split into training and testing samples in a 75:25 ratio. An ML algorithm, Random Forest, was used to train a
model that generates predictions for cancerous cell classication. Random Forest
has a hyperparameter ‘n_estimater’, which determines the number of distinct forests that can be formed during training. It was set to 100in our experiment. The
trained model predicted the cancerous cells as benign and malignant on test data.
The trained model successfully detected the type of cancer (benign and malignant)
in 95% of cases. The confusion matrix by the Random Forest model is given
in Fig.7.
Explanation Generation Three post-hoc XAI methods, LIME, SHAP, and PDP,
were used to support the model-based prediction with explanations.
Explanations Using LIME The LIME explanations with the Wisconsin dataset
are presented in Fig.8. LIME provides the optimal range of each feature for malignant and benign classes. As demonstrated in Fig.8a, the input instance is more
likely to be malignant if the worst area is less than 516.45 and the radius error is less

104
https://t.me/med1917
P. Bedi et al.
than or equal to 0.23. The worst area and radius error for the testing instance are
515.80 and 0.18, respectively, as shown in the table in Fig.8a. This makes the testing instance to be malignant. Figure8b shows that the model incorrectly predicted
the cancer as malignant when it was benign. The optimal range for worst concave
points is greater than 0.06 and less than 0.10. But as shown, in Fig.8b, the value of
the worst concave points was 0.13 making the prediction benign. Similarly, other
features like worst smoothness and worst concavity also contributed to the wrong
prediction.
Explanations Using SHAP SHAP describes the impact of each feature on the
model prediction in the form of a summary plot. The summary plot was generated
by Kernel Explainer using ‘shap’ python library. The Explainer calculates Shapley
value for each feature corresponding to each instance. A single importance value is
calculated for each feature by averaging the Shapley values over all instances. The
importance value of each feature tells their signicance in model prediction. The
summary plot illustrates the importance values for each feature in a graph, the worst
radius and worst concave points have the greatest impact in detecting benign and
malignant cancerous cells as shown in the summary plot in Fig.9. In comparison,
the fractal dimension error and worst smoothness have the least weightage in cancer
prediction.
Fig. 7 Confusion matrix for breast cancer prediction using random forest
Соседние файлы в папке Библиотека им академика М.И. Перельмана
