Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_117_библиотеки_им_акад_М_И_Перельмана

.pdf
Скачиваний:
2
Добавлен:
15.09.2026
Размер:
15 Мб
Скачать
☆
198
S. B. Khan et al.
The discriminative ability of the models was further assessed using ROC curves. As illustrated in the ROC curve comparison in Fig.17, the proposed ResNeXt model exhibits a higher area under the curve (AUC) than the traditional CNN model. The higher AUC value signies that the ResNeXt model can better differentiate between positive and negative classes, reinforcing its superior performance.
Apart from quantitative metrics, interpretability is a crucial aspect of medical applications. The proposed ResNeXt model excels in this regard, thanks to applying XAI techniques, particularly Grad-CAM.The Grad-CAM visualizations for both models were generated, and the results differed. The Grad-CAM heatmap for the proposed ResNeXt model highlighted specic regions within the kidney CT scan images that inuenced its predictions. In contrast, the traditional CNN needed more harvestability, making it challenging for medical professionals to understand the reasoning behind its decisions.
Our suggested ResNeXt model has been found to be superior to conventional CNN methods, as demonstrated through statistical evaluations, ROC curve analy­ses, and Grad-CAM visual representations. This framework not only achieves higher accuracy, precision, recall, and F1 score metrics, but also has advanced capa­bility in diagnosing kidney irregularities. In addition, XAI techniques have been integrated into our approach to enhance its interpretability, providing medical prac­titioners with a clearer understanding of the model’s reasoning. Our methodology is
Fig. 17 ROC curve for proposed and CNN model comparison
Enhancing Diagnosis ofKidney Ailments fromCT Scan withExplainable AI
199
at the forefront of solutions in diagnosing and treating kidney anomalies, offering heightened diagnostic precision and clarity in decision-making. This ResNeXt model is poised to make groundbreaking strides in kidney care, elevating patient outcomes. Furthermore, its efcacy on the IoMT platform signies the potential of delivering top-tier diagnostic solutions to areas with constrained resources, thus democratizing access to quality healthcare worldwide.
5 Conclusion
In this chapter, our research endeavors are centered around enhancing the interpret­ability of deep learning models, often characterized as “black boxes,” particularly in detecting kidney-related abnormalities. Our efforts have yielded substantial advancements by integrating AI Shapley values and Grad-CAM for visualization, coupled with the ResNeXt and XAI models, aimed at identifying kidney cysts, stones, and tumors. By harnessing AI Shapley values, we have bestowed the model newfound transparency, enabling healthcare professionals to delve deeper into the factors underpinning its predictions. This level of transparency not only bolsters the model’s reliability but also furnishes clinicians with invaluable insights that can signicantly enhance patient care. Notably, our proposed framework has demon­strated an exceptional level of accuracy, achieving an impressive rate of 99.52% through the utilization of K= tenfold stratied sampling. This amalgamation of heightened accuracy and transparency stands poised to revolutionize the eld of kidney abnormality diagnosis, ultimately translating into improved patient out­comes. With a transparent model, healthcare practitioners can make more informed decisions, leading to heightened diagnostic precision and increased condence in administering treatments.
Our research is an example of the successful integration of ResNeXt and XAI models in medical informatics contexts, paving the way for further development and expansion of deep learning methodologies in the medical domain. Additionally, the framework’s prociency when applied on an IoMT platform shows its feasibil­ity in resource-constrained regions, making sophisticated diagnostic tools more accessible to a broader audience and democratizing access to state-of-the-art healthcare.
References
1. Holzinger A, Langs G, Denk H, Zatloukal K, Müller H (2019) Causability and explainabil­ity of articial intelligence in medicine. Wiley Interdiscip Rev Data Mining Knowl Discov 9(4):e1312
2. Al’Aref SJ, Anchouche K, Singh G, Slomka PJ, Kolli KK, Kumar A, Dey D etal (2020) Clinical applications of machine learning in cardiovascular disease and its relevance to cardiac imaging. Eur Heart J 40(24):1975–1986
200
3. Lundervold AS, Lundervold A (2019) An overview of deep learning in medical imaging focus­ing on MRI.Z Med Phys 29(2):102–127
4. Yoon H, Kim E, Gao Y, Kim HJ, Li Z, Lee J, Nam HG etal (2020) Quantitative criteria for assessing the spatial pattern of lobular carcinoma in situ. Breast Cancer Res 22(1):1–10
5. Azizi S, Bayat S, Yan P, Tahmasebi A, Kwak JT, Xu S, Turkbey B etal (2020) Deep transfer learning for characterizing choline and spermine on prostate cancer treatment response. Med Image Anal 61:101652
6. Goldenberg SL, Nir G, Salcudean SE (2021) A new era: articial intelligence and machine learning in prostate cancer. Nat Rev Urol 18(6):327–340
7. Chen C, Qin C, Qiu H, Tarroni G, Duan J, Bai W, Rueckert D (2020) Deep learning for cardiac image segmentation: a review. Front Cardiovasc Med 7:25
8. Ribeiro ÁH, Ribeiro MH, Paixão GM, Oliveira DM, Gomes PR, Canazart JA, Meira Jr W etal (2021) CardioNet: a large-scale dataset and a deep learning model to predict clinical outcomes in patients undergoing SARS-CoV-2 RT-PCR tests. medRxiv
9. Zhou Y, Shi W, Chen L, Gao X, Tang H (2020) Deep learning in medical ultrasound analysis: a review. Engineering 6(4):427–441
10. Boers T, Slump CH, Keuning J, Maass AH (2021) Imaging biomarkers for the diagnosis and prognosis of patients with heart failure: a systematic review. Eur J Heart Fail 23(3):313–326
11. Saadi RA, Dashtipour K, Hussain A, Zhang L, Ali AS, A. (2020) Explainable deep learning for predicting response to deep brain stimulation in patients with Parkinson’s disease. Expert Syst Appl 150:113263
12. Kim H, Lee G, Park H (2021) Deep learning-based gait analysis for patients with Parkinson’s disease using explainable articial intelligence. Front Aging Neurosci 13:648801
13. Selvaraju RR, Cogswell M, Das A, Vedantam R, Parikh D, Batra D (2017) Grad-CAM: visual explanations from deep networks via gradient-based localization. In: 2017 IEEE International Conference on Computer Vision (ICCV). IEEE, Piscataway, NJ, pp618–626
14. Bychkov D, Linder N, Turkki R, Nordling S, Kovanen PE, Verrill C, Lundin J et al (2018) Deep learning based tissue analysis predicts outcome in colorectal cancer. Sci Rep 8(1):1–11
15. Schlemper J, Oktay O, Schaap M, Heinrich M, Kainz B, Glocker B, Rueckert D (2019) Attention-gated networks for improving deep learning-based segmentation of brain tumors. In: Brainlesion: glioma, multiple sclerosis, stroke, and traumatic brain injuries. Springer, Cham, pp74–83
16. Hannun AY, Rajpurkar P, Haghpanahi M, Tison GH, Bourn C, Turakhia MP, Ng AY (2019) Cardiologist-level arrhythmia detection and classication in ambulatory electrocardiograms using a deep neural network. Nat Med 25(1):65–69
17. Attia ZI, Kapa S, Lopez-Jimenez F, McKie PM, Ladewig DJ, Satam G, Noseworthy PA etal (2019) Screening for cardiac contractile dysfunction using an articial intelligence–enabled electrocardiogram. Nat Med 25(1):70–74
18. Roy Y, Banville H, Albuquerque I, Gramfort A, Falk TH, Faubert J (2019) Deep learning­based electroencephalography analysis: a systematic review. J Neural Eng 16(5):051001
19. Ribeiro DC, Cardoso JS, Silva CA (2020) Interpretable multiple sclerosis lesion segmentation from magnetic resonance imaging using deep learning. Med Image Anal 65:101788
20. Rajaraman S, Candemir S, Kim I, Thoma G, Antani S (2018) Visualization and interpretation of convolutional neural network predictions in detecting pneumonia in pediatric chest radio­graphs. Appl Sci 8(10):1715
21. Rahman T, Chowdhury ME, Khandakar A, Kadir MA, Masud M, Islam K, Mahbub ZB etal (2020) Explainable machine learning model for pneumonia detection from chest X-ray images. Sensors 20(21):6240
22. Poplin R, Varadarajan AV, Blumer K, Liu Y, McConnell MV, Corrado GS, Webster DR etal (2018) Prediction of cardiovascular risk factors from retinal fundus photographs via deep learning. Nat Biomed Eng 2(3):158–164
23. Ronneberger O, Fischer P, Brox T (2015) U-net: convolutional networks for biomedical image segmentation. In: International conference on medical image computing and computer-assisted intervention. Springer, Cham, pp234–241
S. B. Khan et al.
Enhancing Diagnosis ofKidney Ailments fromCT Scan withExplainable AI
24. Islam MN, Hasan M, Hossain M, Alam M, Rabiul G, Uddin MZ, Soylu A (2022) Vision trans­former and explainable transfer learning models for auto detection of kidney cyst, stone and tumor from CT-radiography. Sci Rep 12(1):1–4
25. Xie S, Girshick R, Dollár P, Tu Z, He K (2017) Aggregated residual transformations for deep neural networks. In: Proceedings of the IEEE conference on computer vision and pattern rec­ognition, pp1492–1500
26. Sze V, Chen YH, Yang TJ, Emer JS (2017) Efcient processing of deep neural networks: a tutorial and survey. Proc IEEE 105(12):2295–2329
27. Lundberg SM, Lee SI (2017) A unied approach to interpreting model predictions. In: Advances in neural information processing systems, pp4765–4774
28. Strumbelj E, Kononenko I (2010) An efcient explanation of individual classications using game theory. J Mach Learn Res 11(Aug):1–18
201
Explainable AI forColorectal Cancer
Classication
MwengeMulenga, ManjeevanSeera, SameemAbdulKareem, andAznulQalidMdSabri
Abstract Colorectal cancer (CRC) ranks second highest in global mortality among
nonsex-related cancers. Conventional machine learning (ML) algorithms applied to microbiome-based CRC detection often yield suboptimal accuracy. Conversely, deep neural network (DNN)-based methods encounter limitations due to scarce labeled samples, data imbalance, and dominant features. The lack of interpretability in articial intelligence models further hinders their adoption in healthcare. This chapter proposes an explainable DNN model for improved CRC detection utilizing stool-based microbiome data. The model employs a square root-based normaliza­tion method and a feature extension approach, incorporating customized normaliza­tion techniques to enhance prediction performance. These methods effectively address outliers, dominant features, and dimensionality challenges. The square root-based method mitigates the effect of outliers and feature dominance, while the feature extension technique expands the dataset’s feature space, potentially improv­ing feature relevance across samples. Leveraging automatic feature selection by the DNN algorithm, the model performs classication using a subset of available fea­tures. Evaluation on publicly available datasets demonstrates the efcacy of the proposed methods, with the square root-based method achieving area under the curve scores of 91.3% and 75.8% on datasets 1 and 2, respectively. The feature extension-based method achieves AUC scores of 90.2% and 74% on the respective datasets.
M. Mulenga Business Studies Division, National Institute of Public Administration, Lusaka, Zambia
M. Seera (*) School of Business, Monash University Malaysia, Selangor, Malaysia e-mail: manjeevansingh.seera@monash.edu
S. A. Kareem · A. Q. M. Sabri Faculty of Computer Science and Information Technology, University of Malaya, Kuala Lumpur, Malaysia
Ltd. 2024 R. Aluvalu et al. (eds.), Explainable AI in Health Informatics, Computational Intelligence Methods and Applications,
https://doi.org/10.1007/978-981-97-3705-5_10
203© The Author(s), under exclusive license to Springer Nature Singapore Pte
204
Keywords Colorectal cancer · Microbiome data · Deep neural networks · Normalization techniques · Feature dominance
M. Mulenga et al.
1 Introduction
Colorectal cancer (CRC) ranks second highest among nonsex-related cancers worldwide in terms of mortality rate [1]. Early detection of CRC can reduce the death rate by approximately 90% [2]. The analysis of the microbiome in stool sam­ples has gained signicant interest as a noninvasive method for CRC detection [3,
4]. The advent of Next Generation Sequencing technologies has resulted in the gen-
eration of massive, high-dimensional, and heterogeneous omics data [5]. However, analyzing such complex microbiome data using statistical methods poses signi­cant challenges. Traditional machine learning (ML) algorithms have limitations in selecting relevant features from a large number of microbial taxa [6]. Although recent works have explored deep neural network (DNN)-based methods to handle high-dimensional features, they are prone to overtting [7]. Research indicates that the use of DNNs does not provide signicant performance improvements over tra­ditional ML methods [1, 8]. Nevertheless, other studies have achieved impressive results in DNN-based classication of microbiota [9–11].
The classication performance of ML algorithms on microbiome data is nega­tively affected by its sparsity, dominant features, and skewed nature [12–14]. Consequently, preprocessing techniques such as data normalization play a vital role in microbiome data classication. However, the normalization of microbiome data remains a challenge [15] and represents an ongoing area of research. Conventional data normalization methods such as minimum–maximum (min–max), Z-score nor­malization (ZSN), and median and median absolute deviation normalization (MMADN) are considered unsuitable for normalizing gene sequence-based data [16]. This is because gene sequence-based data, such as microbiome data, exhibit compositional characteristics, with a constant sum and restricted to nonnegative values [17]. To address this limitation, normalization methods specically tailored to gene sequence-based data have been developed, including trimmed mean of M-values (TMM), relative log expression (RLE), and rarefying, aiming to enhance data classication [18]. However, these methods also have limitations, such as the inability to handle the excessive number of zeros commonly found in sequence data [13].
A study conducted by Pereira etal. [16] evaluated the performance of nine gene sequence-based data normalization methods in identifying differentially abundant genes in microbiome data. The study found that TMM and RLE yielded the best performance, exhibiting high true positive rates (TPR), low false positive rates (FPR), and low false discovery rates. However, these methods may not perform optimally on noncompositional data, which arises when demographic data is com­bined with operational taxonomic units (OTUs) to create a single dataset. In such cases, conventional data normalization methods appear to be more appropriate.
Explainable AI forColorectal Cancer Classication
Furthermore, while recently Singh and Singh [19] demonstrated that no single method surpasses others across all datasets, Mulenga etal. [20] showed that extend­ing the feature space of a dataset using conventional normalization methods enhances the performance of a DNN model by adjusting feature importance throughout the dataset.
Moreover, the lack of interpretability or explainability in articial intelligence (AI) models poses a signicant challenge in healthcare applications [21]. In the context of cancer detection, interpretability is crucial for understanding the reason­ing behind an AI model’s decision-making process and building trust in its accuracy [22]. While AI models have demonstrated high accuracy rates, their lack of inter­pretability raises concerns regarding their reliability and safety in healthcare appli­cations. To tackle this issue, researchers have begun exploring the use of explainable AI (XAI) methods in cancer detection, aiming to provide insights into the decision­making process of AI models.
This chapter extends the method proposed in the aforementioned study [20] and demonstrates that customized normalization methods can be employed for feature extension to enhance the classication performance of a DNN model. Customized normalization methods play a signicant role in the implementation of dynamic data normalization, overcoming the limitations of static normalization, which exhibits poor generalization performance due to its dependence on underlying data properties [19]. The contributions of this chapter are as follows:
• A method called square root-sum (sqrt-sum) that transforms a dataset by com-
puting the sum of each data entry and the standard deviation of the dataset, fol-
lowed by taking the square root of the sum.
• A method that utilizes custom normalization methods to convert a two-
dimensional (2D) dataset into a three-dimensional (3D) dataset by representing
each scalar data point as a vector.
205
The chapter is organized as follows: Sect. 2 provides a review of related works, followed by the description of the proposed method in Sect. 3. Results are presented in Sect. 4, and subsequent discussions are provided in Sect. 4. Finally, Sect. 5 con­cludes the chapter and outlines potential future research directions.
2 Related Works
Data imbalance in microbiome samples is an active research area. Knights etal. [23] conducted a study to identify microorganism groups that change with respect to variations in the host’s physiology or disease. They found out that replicating train­ing data by adding noise can lead to improvements in the predictive power of mod­els. Although the use of data augmentation on the training set helped to reduce overtting and consequently improved the models’ prediction accuracy, the decrease in error was not very signicant. Lo and Marculescu [24] proposed a method that uses neural networks (NN) to classify phenotypes of a host based on metagenomic
206
M. Mulenga et al.
data. To address overtting, they used a new data augmentation technique and a dropout technique. The proposed model had a comparatively high classication accuracy on both synthetic and real data. However, the method was not tested on pooled datasets, and hence did not address the issue of variability across CRC datasets.
Another signicant area of microbiome samples classication is data normaliza­tion, a preprocessing technique that identies and removes systematic variability [16]. While the effectiveness of a normalization method depends on the characteris­tics of the target dataset, datasets used in ML tasks differ in terms of underlying features. Therefore, there has been a substantial amount of research on the applica­tion of data normalization methods on various types of datasets. Manor and Borenstein [15] proposed a normalization method that applied ML methods on single-copy genes to correct marked biases within and across human microbiome­based samples, and to obtain measures that have biological meanings, which are also accurate. Though the method corrects spurious variations across samples and produces accurate downstream comparative analysis with meaningful abundance measures, it is limited to identifying bacterial and archaeal organisms only and does not cover fungal and viral organisms. Gloor etal. [17] demonstrated that current methods for compositional data-based analysis can be easily adopted for analysis of high throughput sequence data. Their method produced good results because it accounted for the compositional nature of microbiome data. However, the study was only based on 16S rRNA data and did not consider shotgun sequence data. Kaul etal. [25] conducted a study that addresses the challenges associated with sparsity in microbiome data. They proposed a method that identies three types of zero val­ues found in microbiome data and conducted hypothesis testing on relative taxa abundance in more than one experimental group. Although the method was able to improve the false positive discovery rate, experiments were based on simulated data that may not reproduce the same level of performance on real data.
Peng etal. [12] proposed a zero-inated beta regression method for the identi­cation of features that are differentially abundant for multiple phenotype classica­tion. The method used cumulative sum normalization and outperformed other methods with signicantly higher area under the curve (AUC) scores on simulation data. Though the method accounted for the sparse and compositional nature of metagenomic data that improved its performance, comparisons were based on simu­lated data that may limit the method’s ability to generalize when using real data. Similarly, Douglas etal. [26] conducted a study to determine if multi-omics can differentially classify the state of Crohn’s disease (CD) and its treatment outcome. The method controlled for inter-sample variation due to microbiome genome size by normalizing Kyoto Encyclopedia of Genes and Genomes (KEGG) abundances within each sample using universal single-copy gene abundance. However, model generalization was not attained in the study owing to technical limitations encoun­tered across individual studies.
Zyprych-Walczak etal. [27] proposed a method for selecting optimal normaliza­tion procedure for any dataset by computing bias and variance values of control genes, specicity and sensitivity of a method, classication errors, and diagnostic
Explainable AI forColorectal Cancer Classication
207
plots. Though automation of data normalization selection process is important, datasets used in the study were not sufcient to conclusively establish an optimum procedure for automatic selection of normalization methods. Pereira etal. [16] com­pared nine normalization methods for analyzing metagenomic data and associated high performance to trimmed mean m-value (TMM) and relative log expression (RLE). Though the study used a data-driven method for the evaluation of metage­nomic data, it did not consider phonotype classication, which is an important area in metagenomic based studies.
McKnight etal. [28] investigated potential problems with normalization meth­ods that are used in gene sequence data and observed that TMM and other similar transformation methods contrary to rarefying and proportions do not ensure equal­ity in the number of reads across samples. The authors claimed that the methods used to reduce the effect of dominant features while amplifying rare features could be misleading in terms of community differences. Though the study considered variance standardization, differential abundance testing, and investigated abundance across community levels, it was based on one simulated dataset and a single real dataset that may not be enough for drawing conclusions about the robustness of methods. Weiss etal. [29] investigated how challenges associated with microbiome data affect its normalization procedures and differential abundance testing, and observed that among normalization techniques, rarefying was the only approach that was not frequently confounded by library size, which obscures biological inter­pretation of results. Interpretation of results is as important as accuracy in medical application of automated detection of diseases [30]. Unlike other studies that were solely based on simulated data, the method investigated data normalization based on both real and simulated data. However, rarefying tends to have reduced sensitivity due to the elimination of part of the dataset.
Metagenomic analysis was used by Guo etal. [31] to investigate the phyloge­netic and functional traits of anammox communities in three microbial aggregates, and ZSN was used to preprocess data to reduce the impact of outliers and dominant features. ZSN, however, does not take the compositional nature of microbiome data into account, which may affect classication performance of a ML algorithm. Korpela etal. [32] investigated the use of gut microbiome signatures of obese indi­viduals to predict their host and microbiome response to dietary interventions. The study used min–max normalization in addition to log transformation in order to preprocess input data. Despite the use of log transformation, which is a suitable technique for skewed datasets such as microbiome data, the use of min–max nor­malization is not suitable for microbiome-based datasets due to the presence of dominant features in the data. Singh and Singh [19] investigated how 14 selected data normalization methods. Which included min–max, ZSN, sigmoid, and MMADN, impact the classication performance of ML algorithms based on full feature set, feature selection, and feature weighting methods. The study covered a wide range of normalization methods and the comparisons performed were quite elaborate and showed that no single data normalization method is superior to others. However, the study did not consider metagenomic data, which has slightly different properties than the ones covered in their study.
208
With a few exceptions, most works on microbiome-based data does not apply conventional data normalization methods since the data is considered to be compo­sitional, rendering conventional methods unsuitable for normalization [17]. Microbiome data consists of operational taxonomic unit (OTU) data and accompa­nying metadata. OTUs account for the part of the microbiome data that is composi­tional, while the metadata that have elds such as age, weight, and other demographic attributes represents the part that has absolute values. Therefore, a study that com­bines the metadata and the OTU data in an analysis produces a dataset that appears to be noncompositional, and hence normalization techniques that are normally used for noncompositional data can be used. The study proposes a method that uses a dataset and combines two demographic attributes such as age and biomass index (BMI) with OTU data in order to improve DNN prediction of CRC based on the microbiome in stool samples.
Furthermore, the researchers are realizing the importance of XAI in CRC detec­tion [21]. For example, Zhang etal. [22] proposed a general method for modifying conventional convolutional neural networks (CNNs) to improve their interpretabil­ity. Although CNN models are associated with very high performance, their use in microbiome is still limited. Le etal. [10] proposed an interpretable neural networks algorithm that has both an encoder and decoder for predicting gut metabolites obtained from the gut microbiome. While the model was highly interpretable, the dataset was very small. Carrieri etal. [33] used XAI to reveal changes in skin micro­biome composition caused by phenotypic differences. Although interpretability is attained in the model, it lacks details on disease-based classication.
Based on the need to treat a microbiome dataset as noncompositional data and the limitations discussed in the related methods, our study, while adopting an XAI approach, proposes the sqrt-sum and a feature extension method to improve the classication of microbiome data. As a priory, customized normalization methods such as the sqrt-sum and other related functions are generated based on existing methods such as Pareto scaling [34]. The new methods are targeted at reducing feature dominance and outliers by combining a technique that uses standard devia­tion with the one that computes square roots of nonzero data points in order to transform a dataset. The customized normalization methods are then used to gener­ate additional features in the dataset and potentially improve the feature importance. The transformed dataset is then subjected to a DNN for classication, which per­forms an embedded form of feature selection on the potentially improved fea­ture space.
M. Mulenga et al.
3 Methods
This chapter proposes a square root-based customized normalization method called
sqrt-sum and a feature extension method for microbiome data classication. The sqrt-sum method transforms samples by adding the standard deviation of the dataset
to nonzero items in that dataset, then computes the square root of the sum. The