Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_117_библиотеки_им_акад_М_И_Перельмана

.pdf
Скачиваний:
2
Добавлен:
15.09.2026
Размер:
15 Мб
Скачать
☆
()=() ()
 
 
()()
′′
,,
ÆÆ
n
j
n
ii
()
=
1
x
i
′
y
xy
c
xy
∂
1
F
k
c
SFM
k
=
∑
Explainable AI inDisease Diagnosis
95
Partial Dependence Plot (PDP) [26] PDP is a model-agnostic XAI method that generates global explanations. It is suitable for evaluating the mean marginal effect of the input attributes on the predictions. The PDP visualises the effect of the selected feature values on the outcome by controlling all other features. These effects are analysed over different values, such as the range between the minimum and maximum values. A graph plots the average values of the obtained results. The partial dependence function is given in Eq. (4).
ÆÆÆ
f xExfxx fxxPx
iiii ii i
′′
=∫
d
(4)
After simplifying, we obtain
=
()
i
1
∑
f x
The xi is the selected feature whose importance is being measured, and
fxx
j
,
′
(5)
are the rest of the features in the dataset. PDP is purposely used to analyse any variable’s effect on the prediction. However, these variables are kept singular or in a combina­tion of two for simplication.
Gradient Class Activation Map (Grad CAM) [23] Deep Learning, mainly CNN, performs adequately in computer vision tasks. Explanations for CNN predictions can help to simulate human intelligence and understand the model better. Grad CAM utilises weights and feature maps in the CNN model to nd the activated part of an image that contributes to the corresponding label.
Grad CAM replaces the fully connected layer of the CNN model with global
average pooling.
c
F
=
∑∑
k
N
k
M
∂
(6)
Here, yc, N, andM denote the score obtained for the class ‘c’ before the softmax
function, the number of pixels and the feature map, respectively. The obtained
is the output of the global average pooling of the gradients. It also represents the weight (importance) of kth feature map for the class ‘c’.
Sc is obtained by a weighted combination of feature maps and applying the ReLU
function to get the positively inuencing features on the target class ‘c’.
The next section describes the applications of XAI in disease diagnosis.
c
kck
(7)
96
P. Bedi et al.
4 XAI inDisease Diagnosis
AI has produced advanced healthcare systems and automated various health-related tasks such as cancer prediction [28], medical chatbots, arrhythmia detection [28], early diagnosis of fatal blood disease [16], treatment of rare diseases and drug repurposing [31], personalised treatment [32], digital storage and management of medical records, and many other tasks. XAI enhances healthcare tasks and under­standability by providing a rationale for the AI-generated recommendations.
Disease diagnosis is one of the most popular research areas in the medicinal domain. Some of the XAI methods applied within the disease diagnosis systems are given below:
Grad CAM in Arrhythmia Detection An electrocardiogram (ECG) is one of the tests performed to detect arrhythmia. XAI can be used to explain arrhythmia based on ECG.Y. Y.Jo etal. [28] introduced an Explainable Deep Learning Model (XDM) for arrhythmia classication using a neural network-backed ensemble tree approach. The XDM has six deep learning feature modules to detect arrhythmia characteris­tics such as irregular R wave patterns, regular PR intervals, the correlation between P and R wave etc., and an XDM ensemble approach for arrhythmic classication. The output of each feature module gave the probability of each arrhythmic charac­teristic, and the nal ensemble module was the multilayer perceptron for the classication. The study [28] used the Seojong ECG dataset1 which consists of data from Mediplex Seojong Hospital and Seojong General Hospital. The arrhythmia was classied into nine classes based on ECG signals, namely, ventricular tachycar­dia (VT), junctional rhythm (JR), normal sinus rhythm (NSR), pacemaker rhythm (PM), supraventricular tachycardia (SVT), atrial brillation (AF), complete atrio­ventricular block (CAVB), second-degree atrioventricular block Mobitz type II (2AVB-T2), and second-degree atrioventricular block Mobitz type I with Wenckebach phenomenon (2AVB-T1). The aggregated precision, recall, and F1-score for the nine classes were found to be 0.936, 0.793, and 0.852, respectively. The Area Under the Receiver Operating Characteristic (AUROC) curve values of PM, NSR, AF, JR, SVT, VT, CAVB, 2AVB_T2, and 2AVB_T1 were found to be
0.988, 0.994, 0.976, 0.993, 0.992, 0.921, 0.976, 0.983, and 0.961, respectively. Gradient-class activation map (sensitivity map) and the guided gradient backpropa­gation methods were used for the visualisation of XDM interpretations. First-order gradients of the classier probabilities were used to create the sensitivity map, which illustrates the sensitive regions responsible for the decision.
PDP and Accumulated Proles in Cancer Prediction Cancer is a deadly disease which is responsible for uncountable deaths worldwide. Early detection of the dis­ease is the only solution to stop it from further spreading in the body as it becomes incurable in later stages. AI can predict the disease based on various types of patient
1
Dataset available at request to corresponding author.
Explainable AI inDisease Diagnosis
97
data such as genomic, proteomic, or clinical data. XAI provides additional support to AI-based disease diagnosis models by identifying the subset of input features responsible for cancer from the given patient input feature set and presenting it in a human-readable form. Z.U. Ahmed etal. [17] proposed an ensemble approach to explore and analyse the effect of various risk factors on lung and bronchus cancer (LBC) mortality rate. Twelve risk factors including smoking, sulphur dioxide, and poverty were considered in the study [17]. These risk factors were extracted from different sources. National Vital Statistics System at the National Center for Health Statistics of the Centers for Disease Control and Prevention2 was used to extract the LBC mortality rate, County Health Ranking3 and Institute for Health Metrics and Evaluation4 were used to extract cigarette smoking prevalence data, and from US Census Bureau’s Small Area Income and Poverty Estimates programme5 poverty data was extracted. The study [17] trained three spatial regression models (Spatial lag, Spatial error, and Geographically Weighted OLS Regression (GW-OLS)), ve base learner models (Generalised Linear Model (GLM), Random Forest (RF), Gradient Boosting Machine (GBM), Extreme Gradient Boosting (XGBoost), Deep Neural Network (DNN)), and two stack-ensemble models (best base learners and all base learners) to analyse the risk factors of LBC.The performance of models was compared using Mean Absolute Error (MAE) and Root Mean Squared Error (RMSE) metrics. The MAE with Spatial lag, Spatial error, GW-OLS, GLM, RF, GBM, XGBoost, DNN, ensemble of best base learners and ensemble of all base learners on testing data were obtained as 9.02, 6.53, 6.20, 6.49, 6.16, 6.20, 6.41,
7.00, 6.08, and 6.06, respectively. The RMSE with Spatial lag, Spatial error, GW-OLS, GLM, RF, GBM, XGBoost, DNN, ensemble of best base learners and ensemble of all base learners on testing data were obtained as 11.46, 8.35, 8.09,
8.31, 8.03, 8.06, 8.35, 9.03, 7.95, and 7.74, respectively. The study showed that the stack-ensemble approaches were more accurate. However, the stack-ensembled approaches are complex in design as they stack multiple models to generate predic­tions. The authors in [17] used Permutation-based feature importance XAI methods to visualise the effect of input on the outcome. These methods included partial dependence (PD), local dependence (LD), and accumulated local (AL) proles. XAI methods were also visually represented with plots. The study using XAI meth­ods showed that smoking is the most critical risk factor responsible for lung and bronchus cancer.
LIME in Early Diagnosis of Fatal Blood Disease Leukaemia is one of the fatal blood diseases which takes numerous lives annually. It has a high mortality rate because it is a fast-acting disease. Therefore, the diagnosis in the early stages can increase the patient’s survival chances. AI-based systems can help detect leukaemia
2
https://www.cdc.gov/nchs/index.html
3
https://www.countyhealthrankings.org/explore-health-rankings
4
https://www.healthdata.org/
5
https://www.census.gov/programs-surveys/saipe.html
98
P. Bedi et al.
from the patient’s lab results, and XAI can help doctors to analyse the predictions generated by AI-based systems. In [16], the authors developed an AI system which uses Transfer Learning to detect the presence/absence of leukaemia from the micro­scopic images of blood cells. The dataset in [16] was taken from the Kaggle reposi­tory named Leukemia Classication,6 containing 15,135 images. Pre-trained CNN models, ResNet101V2, VGG19, InceptionResNetV2, and InceptionV3, were used in the diagnosis because of their excellent performance in computer vision tasks. Accuracy and F1-score were measured to compare the performance of CNN mod­els. InceptionResNetV2 achieved the best results with 0.8002 accuracy and 0.7980 F1-score with testing data whereas InceptionV3 performed best with 0.9665 accu­racy and 0.9665 F1-score with validation data. LIME was applied to the predictions generated by the InceptionV3 model. It highlighted the parts of the image respon­sible for the model’s predictions.
Multi-Scale CAM in Glaucoma Diagnosis A clinically interpretable deep learning- based framework, namely EAMNet, was proposed in [33] to diagnose glaucoma disease. The framework consisted of a CNN network, Multi-layer Average Pooling, and Evidence Activation Mapping to diagnose glaucoma from retinal fun­dus images and highlight the affected regions in the images. Distinct regions con­tributing to the predicted diagnosis were highlighted using the convolution feature maps that were obtained during the training of the ResBlock CNN model. The model concatenated the feature maps at different levels by global averaging to get an activation map. This map emphasises the areas (such as notch on the neuro­retinal rim, defect or bleeding on optic disc) on the fundus images that provide concrete evidence for diagnosing glaucoma. The experiments in [33] were per­formed with the ORIGA public dataset [34]. EAMNet achieved a 0.88 AUROC value for diagnosis of glaucoma disease.
SHAP in Parkinson’s Disease XAI methods are not restricted to generating expla­nations for black-box models only. They have also been used in feature selection. SHAP was used for feature selection in the Parkinson’s disease diagnosis in [35]. Parkinson’s disease is a neurodegenerative disease in which the patient’s data is col­lected using force sensors on their feet or voice presentation. It thus, produces a high feature dimensionality for the disease diagnosis. High-dimensional features may produce undesirable outcomes such as redundant features, reduce the computational efciency or generate biased results. Hence, in [35], the authors selected the impor­tant features from the high-dimensional feature set of Parkinson’s disease using Fscore, Anova-F, Mutual Information (MI), and SHAP value. The dataset for the study [35] was taken from the UCI repository, Parkinson’s Disease Classication.7 The dataset contains 188 patients’ records. A comparative analysis of four feature selection methods, Fscore, Anova-F, MI, and SHAP, was performed on this dataset.
6
https://www.kaggle.com/datasets/andrewmvd/leukemia-classication
7
https://archive.ics.uci.edu/dataset/470/parkinson+s+disease+classication
Explainable AI inDisease Diagnosis
99
The most important features extracted by each feature selection method were given as input to the classication algorithm. Four ML algorithms, RF, deep forest, Light GBM, and XGBoost, were used for the classication of Parkinson’s disease. Among the four feature selection methods, SHAP showed the best performance in Parkinson’s disease diagnosis with four ML models. SHAP-gcForest achieved the highest accuracy and F1-score of 91.78% and 0.945, respectively. SHAP-XGBoost, SHAP-LightGBM, and SHAP-RF produced an accuracy of 90.04%, 91.62%,
90.34% and F1-score of 0.936, 0.945, 0.939, respectively.
The following section shows the implementation of XAI post-hoc methods that generate explanations with ML/DL models.
5 Case Studies ofXAI Post-hoc Methods
This section demonstrates the implementation of XAI post-hoc methods by generat­ing disease diagnosis explanations on the predictions produced by the ML/DL mod­els. This has been exemplied below using two case studies carried out on the datasets of COVID-19 and Breast cancer diseases.
Case Study 1: Explanations to Support COVID-19 Diagnosis Using Chest Radiography Images
This case study explains the step-wise implementation of three XAI post-hoc meth­ods (LIME, SHAP, and Grad-CAM) applied to the chest radiography dataset. The chest radiography dataset was retrieved from the Kaggle repository.8 Model-based predictions for COVID-19 diagnosis must be performed as the rst step to generate explanations using XAI post-hoc methods. This is briey explained below.
Model-Based Prediction The chest radiography dataset had 13,808 images, out of which 3616 were of patients diagnosed with COVID-19, and the rest of the 10,192 images were of patients without the disease. The size of each image was 299×299×3 pixels. However, the size of the images was reduced from 299×299×3 pixels to 100×100×3 pixels using the opencv9 python library to minimise space utilisation and maximise the processing time. The dataset was divided into training and testing samples in a 75:25 ratio. A CNN architecture-based, InceptionV3 model [36] pre-trained on a large ImageNet [37] dataset was used to generate predictions. The pre-trained weights of the InceptionV3 model were leveraged for transfer learn­ing [38] and to build the model to diagnose COVID-19 disease. To do so, the last layer was replaced with two dense layers of 1024 and 2 neurons, respectively. This model was trained on the training samples to generate predictions on the test sam­ples. The model predicted 90% of cases (normal or COVID) correctly. Figure3 shows the confusion matrix obtained from the test data.
8
https://www.kaggle.com/datasets/tawsifurrahman/covid19-radiography-database
9
https://docs.opencv.org/4.x/
100
P. Bedi et al.
Explanation Generation Three post-hoc XAI methods, LIME, SHAP, and Grad­CAM, were implemented to generate explanations for the InceptionV3 predictions. The three subsections given below describe the implementation of the three XAI methods to generate explanations on the chest radiography images.
Explanations Using LIME LIME highlights regions of interest (ROI) of an image that either positively or negatively contributes to the model prediction. An image explainer object, LimeImageExplainer was built using ‘lime’ python library, to highlight ROI in chest image. The trained model and a chest image were passed as input parameters to the explainer object of LIME.The explainer object has a hyper­parameter, ‘sample_size’ which denes the number of neighbour samples used to approximate the model predictions. This was set to 500. The images with high­lighted ROI as generated by LIME are shown in Fig.4. The green area depicts a positive contribution to the predicted class and the red area depicts a negative con­tribution. The lung features that contribute to the correct prediction are shown in Fig.4a. These features were extracted from the chest radiography image using the InceptionV3 model. Figure 4b shows an incorrectly predicted image where the patient not having COVID-19 (normal image) was misdiagnosed as COVID-19 positive. It is so because the model could not perform feature extraction correctly,
Fig. 3 Confusion matrix of COVID-19 diagnosis using InceptionV3 model
Explainable AI inDisease Diagnosis
101
and the sternum part contributed to the prediction instead of the lungs. Thus, in such a case, the XAI helped to cross-check the model prediction.
Explanations Using SHAP The SHAP method uses a trained model and an image to highlight the pixels contributing to the prediction (COVID-19 diagnosis or nor­mal healthy patient). These contributions are determined by computing Shapley values for each pixel. Shapley values are generated by a shap-explainer object using a partition method. In Fig.5, the rst image is the original chest image, and the other two images (normal and COVID) are the SHAP-generated images. The Shapley images contain light and deep shades of red and blue colour pixels. The colour becomes more intense while moving away from (Shapley value) zero and becomes lighter towards zero, as shown in the SHAP scale in Fig.5. The larger area covered by the positive or negative Shapley values shows a positive or negative contribution to the predicted outcome. More red pixels (positive Shapley values) than blue pixels in Fig.5 (normal image) show that more pixels are contributing to ‘normal’ (not having COVID-19) prediction. Similarly, in the COVID-19 image of Fig.5, there are more blue pixels (negative Shapley values), showing that the image is COVID-19 negative.
Explanations Using Grad-CAM Grad-CAM method takes a pre-trained model, an input chest image, and a feature map (the output of a pre-trained model layer) to visualise the reasons for the model-based predictions. This implementation uses disease diagnosis (having COVID-19 or not) prediction generated by the InceptionV3 model. Grad-CAM generates visual explanations in the form of a heatmap. The feature map is superimposed over the input chest image to generate a heatmap. It is
Fig. 4 LIME explanation. (a) Correct feature selection contributing to right prediction. (b) Incorrect feature selection contributing to wrong prediction
102
Fig. 5 Shapley values generated by SHAP method for normal and COVID classes
P. Bedi et al.
done so because a feature map contains the most important features (parts) of an input image extracted after applying convolutions. Hence the superimposed heat­map highlights the important portions of an input image, which are responsible for the resulting prediction. It uses VIBGYOR colours to highlight the contributing features in the image. The heatmap shows the contribution of each pixel to the model-based prediction where pixels in red and violet are the most and least con­tributing features, respectively. Figure 6 depicts the heatmap obtained with the InceptionV3 model and the feature map of InceptionV3’s ‘mixed10’ layer. In Fig.6a, InceptionV3 correctly predicted the X-ray image as Normal (not having COVID-19). The pixels of the neck area in the chest X-ray image, represented in green colour, show a lesser contribution to the diagnosis prediction. The lower part (lungs), red in colour, shows a more signicant contribution in detecting the absence of COVID-19. In Fig.6b, the actual (true) diagnosis was COVID-19 but InceptionV3 wrongly predicted the image as normal. It happened because the model did not extract the features correctly, which is evident from the Grad-CAM produced image. The image contains a large blue area (full right lung and half of left lung), represent­ing a lesser contribution to the model-generated prediction.
Case Study 2: Explanations to Support Cancer Classication (Benign and Malignant) Using Wisconsin’s Diagnostic Dataset
This case study discusses the implementation of three XAI post-hoc methods, namely, LIME, PDP, and SHAP applied to the Wisconsin Dataset.10 It is an openly accessible dataset archived in the UCI Machine Learning repository. The dataset contains characteristics of breast mass nuclei obtained from digitised breast images. The characteristics include mean texture, radius error, worst area, worst texture, etc.
10
https://archive.ics.uci.edu/ml/datasets/breast+cancer+wisconsin+(diagnostic)
Explainable AI inDisease Diagnosis
Fig. 6 Grad-CAM explanations. (a) Correct feature selection contributing to right prediction. (b) Incorrect feature selection contributing to wrong prediction
103
The explanations of XAI methods were produced based on the model-generated predictions. Therefore, rstly the model-based prediction is discussed briey.
Model-Based Prediction The dataset tabulated 569 instances and 32 characteris­tics. The characteristics had two types of values, numerical and categorical. Strings represented some categorical values. For example, the characteristic ‘type of can­cerous cell’ had two values, ‘B’ and ‘M’, where ‘B’ represented Benign and ‘M’ indicated Malignant. Such categorical values were mapped to numerical labels using ‘scikit-learn’ python library. Next, the dataset was split into training and test­ing samples in a 75:25 ratio. An ML algorithm, Random Forest, was used to train a model that generates predictions for cancerous cell classication. Random Forest has a hyperparameter ‘n_estimater’, which determines the number of distinct for­ests that can be formed during training. It was set to 100in our experiment. The trained model predicted the cancerous cells as benign and malignant on test data. The trained model successfully detected the type of cancer (benign and malignant) in 95% of cases. The confusion matrix by the Random Forest model is given in Fig.7.
Explanation Generation Three post-hoc XAI methods, LIME, SHAP, and PDP, were used to support the model-based prediction with explanations.
Explanations Using LIME The LIME explanations with the Wisconsin dataset are presented in Fig.8. LIME provides the optimal range of each feature for malig­nant and benign classes. As demonstrated in Fig.8a, the input instance is more likely to be malignant if the worst area is less than 516.45 and the radius error is less
104
P. Bedi et al.
than or equal to 0.23. The worst area and radius error for the testing instance are
515.80 and 0.18, respectively, as shown in the table in Fig.8a. This makes the test­ing instance to be malignant. Figure8b shows that the model incorrectly predicted the cancer as malignant when it was benign. The optimal range for worst concave points is greater than 0.06 and less than 0.10. But as shown, in Fig.8b, the value of the worst concave points was 0.13 making the prediction benign. Similarly, other features like worst smoothness and worst concavity also contributed to the wrong prediction.
Explanations Using SHAP SHAP describes the impact of each feature on the model prediction in the form of a summary plot. The summary plot was generated by Kernel Explainer using ‘shap’ python library. The Explainer calculates Shapley value for each feature corresponding to each instance. A single importance value is calculated for each feature by averaging the Shapley values over all instances. The importance value of each feature tells their signicance in model prediction. The summary plot illustrates the importance values for each feature in a graph, the worst radius and worst concave points have the greatest impact in detecting benign and malignant cancerous cells as shown in the summary plot in Fig.9. In comparison, the fractal dimension error and worst smoothness have the least weightage in cancer prediction.
Fig. 7 Confusion matrix for breast cancer prediction using random forest