Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_117_библиотеки_им_акад_М_И_Перельмана
.pdf
146
https://t.me/med1917
problem. Explanation categories can be through simplication, hyperparameters,
local explanation, or graphical representation.
R. Aluvalu et al.
3.2 Category ofExplainability inAI Methods
The rule-based machine learning method is used to identify relationships among
the data parameters, learn inferences from the data, and frame rules on the dataset.
The dataset consists of rules to create rules and updates the rules based on the evolution of the data. The testing of the data by applying the rules generated from learning. The accuracy is measured by testing the sample set from the dataset. It is mostly
used for classication problems. Local explanation is used to learn the specic
problem-based or individual prediction for the machine learning model. Inuencer
function is another approach to explaining the AI results, and mathematical and
statistical methods are used to analyze the effector variable in the model. The change
in hyperparameters and the effect on the result are analyzed.
A linear approximation function is used to nd the relation between the input
variable and the result, based on analysis, past history, and observed data. Linear
approximation is mainly used in medical applications. The difference between the
obtained curve and the estimated curve cannot be explained by the linear approximate method but can be measured by the minus or pulse area of the difference. This
approach is said to be a black box explanation method. Sensitivity is another
approach that measures how well a model is designed based on the true positive
rate. The rate of correctly identied instances from the overall dataset. The confusion matrix is generated based on classication prediction. It shows a truly positive
rate and a truly negative rate. It has a false positive rate. Sensitivity is the true positive rate divided by the total true results (positive and negative).
Counterfactuals is an ML language that interprets the model, one of the emerging concepts in XAI and works on the principle of situations or other possibilities
that have not occurred. The counterfactual will let you know about alternative
options or another possibility.
3.3 XAI Techniques andCase Studies
3.3.1 SHAP
SHAP (SHapley Additive exPlanations) is a highly versatile machine learning technique that can provide a comprehensive and coherent approach to explaining the
output of any machine learning model, including complex models such as deep
learning models. Drawing on the principles of game theory and the concept of
Shapley values, which are used to apportion payouts among a group of players
based on their individual contributions to a collaborative game, SHAP is designed

Explainable AI forBig Data Control
https://t.me/med1917
147
to calculate the contribution of each feature in a model’s output, analogous to how
Shapley values determine the contribution of each player in a cooperative game. As
a result, SHAP allows us to grasp the signicance of each feature in the decisionmaking process of the model [32].
SHAP offers numerous practical applications, including feature selection, model
debugging, and model comparison. Furthermore, it can be employed to provide
explanations for individual predictions and to identify feature interactions within
the model. In essence, SHAP can be an incredibly powerful tool for understanding
and interpreting machine learning models. Case study example: A credit card company wants to predict whether a customer will default on their credit card payments.
They use a machine learning model to make predictions, but they also want to
understand how the model is making its decisions. They decide to use SHAP to
explain the importance of each feature in the model. After training a random forest
classier on their dataset and using SHAP to compute the feature importance, they
nd that the most important features are the customer’s payment history, credit utilization, and income. They then use SHAP to explain the predictions for a specic
customer and nd that the customer is at high risk of defaulting due to late payments, high credit utilization, and low income. This information can be used to take
appropriate actions, such as offering the customer a lower credit limit or a payment plan.
Algorithm for SHAP:
1. Preprocess the dataset to remove any missing values, outliers, or irrelevant
features.
2. Split the dataset into training and testing sets.
3. Train a random forest classier on the training set.
4. Use SHAP to compute the feature importance for the trained model.
5. Identify the most important features based on the SHAP feature’s importance.
6. Use SHAP to explain the predictions for a specic customer by computing the
contributions of each feature to the predicted outcome.
7. Use the explanations to take appropriate actions, such as offering the customer a
lower credit limit or a payment plan.
3.3.2 Anchors
Anchors are a machine learning technique used for explaining individual predictions of black box models, such as deep neural networks. The approach is based on
creating “anchors,” which are simple and interpretable rules that act as proxies for
the predictions of the black box model. Anchors are created using a process called
anchor generation, which identies the most inuential features for a given prediction and creates a rule based on these features. Anchors have been applied in a
variety of domains, including healthcare, nance, and image recognition. It provides a way to interpret the predictions of complex machine learning models, and
can be used to build trust and understanding of the model’s decision-making process.

148
https://t.me/med1917
R. Aluvalu et al.
Case Study Example A company has developed a deep neural network model for
classifying images of animals into different categories (e.g., dog, cat, bird). However,
the model is a black box, and the company wants to understand how the model is
making its predictions in order to improve the model and build trust with customers.
Dataset The company has a dataset of 10,000 labeled images of animals, with each
image labeled as one of ve categories (dog, cat, bird, sh, or reptile). The images
are in different resolutions and formats.
Solution The company decides to use anchors to explain individual predictions of
the black box model. They follow these steps: They sample a set of images from the
dataset and use the black box model to predict the category of each image. For each
predicted category, they use the anchor generation process to generate a set of
anchors that explain the prediction. For example, an anchor for the “dog” category
might be “If the image contains a four-legged animal with fur, then predict the dog
category.” They evaluate the quality of the anchors by checking how well they cover
the samples in the neighborhood of the prediction. They aim to achieve high coverage while keeping the number of anchors small.
They use the anchors to explain individual predictions of the black box model by
checking if the input image satises the rule of the corresponding anchor. If it does,
they explain the prediction based on the corresponding category label. If not, they
try another anchor or leave the explanation as “unavailable.”
Algorithm:
1. Select an image to be explained.
2. Sample a set of “neighbor” images that are similar to the image in terms of the
distribution of pixel values.
3. Use the black box model to predict the category of the image.
4. For each predicted category, generate a set of candidate anchors based on the
features of the input image, such as “If the image contains a four-legged animal
with fur, then predict the dog category.”
5. Use a search algorithm to nd the smallest set of candidate anchors that covers
at least a user-specied proportion of the neighbor images.
6. Evaluate the quality of the anchors by measuring their coverage and
compactness.
7. Use the anchors to explain the prediction by checking if the input image satises
the rule of the corresponding anchor. If it does, explain the prediction based on
the corresponding category label. If not, try another anchor or leave the explanation as “unavailable.”

Explainable AI forBig Data Control
https://t.me/med1917
149
3.3.3 LIME
LIME (Local Interpretable Model-agnostic Explanations) is a machine learning
technique used to explain individual predictions of any black box model. It aims to
provide local- and human-interpretable explanations for the predictions made by
complex models, such as deep learning models. LIME works by creating a local
model around the instance being explained and then approximating the behavior of
the black-box model in the neighborhood of that instance. LIME has a wide range
of applications, including image recognition, natural language processing, and
fraud detection. It can be used to explain individual predictions, detect bias in models, and provide transparency in decision-making systems. For example, let’s say
you work for a medical diagnostics company and have developed a deep learning
model that can diagnose certain types of cancer. You want to ensure that your model
is transparent and explainable so that doctors can understand the reasons behind the
model’s predictions. You decide to use LIME to explain individual predictions made
by the model.
Algorithm:
1. Select an instance to be explained.
2. Generate a set of perturbed instances by randomly masking out features from the
original instance.
3. Generate a set of interpretable features for each perturbed instance using an
interpretable model, such as a linear regression model.
4. Calculate the weights for each interpretable feature using the distance between
the perturbed instances and the original instance.
5. Use the feature importance weights to create an explanation for the prediction
made by the black box model.
3.3.4 ICE
Individual Conditional Expectation (ICE) is a machine learning technique used to
visualize the behavior of a single instance in relation to a specic feature. ICE plots
show how the model’s prediction changes as the value of the feature changes for a
particular instance. Applications of ICE include model explanation, feature engineering, and anomaly detection. It can help us understand how a model makes predictions and identify areas where the model can be improved. It can also be used to
detect outliers or anomalies in the data. A common use case for ICE is in the eld
of healthcare, where models are used to predict patient outcomes based on various
features such as age, gender, medical history, and treatments received. ICE plots can
help clinicians understand how the model is making predictions for individual
patients and identify areas where the model may need to be adjusted for accuracy.
Algorithm to plot:
1. Choose the instance you want to examine and select the feature you want to plot.

150
https://t.me/med1917
R. Aluvalu et al.
2. Create a new dataset by duplicating the instance and varying the selected feature
across a range of values.
3. Use the model to predict the output for each instance in the new dataset.
4. Plot the selected feature against the predicted output for each instance.
3.3.5 PDP
PDP, or partial dependence plot, is a popular machine learning technique that allows
us to visualize the relationship between a particular feature and the output of a
machine learning model. It works by holding all other features in the model constant and plotting the effect of changing the specic feature of interest on the model’s output.
Applications of PDP include understanding the impact of individual features on
the model’s predictions, identifying important features, and detecting feature interactions. It is commonly used in elds such as healthcare, nance, and marketing.
Let’s consider an example of using PDP in the healthcare industry. Suppose we have
a dataset of medical records of patients, including their demographic information,
lifestyle habits, and medical history. We want to build a machine learning model
that predicts the likelihood of a patient developing a certain disease, such as diabetes. We can use PDP to understand the impact of individual features on the model’s
predictions. For example, we can create a PDP for the patient’s age and observe how
the model’s prediction changes as the patient’s age increases or decreases. This can
help us understand whether age is an important factor in predicting the risk of diabetes and whether the model is sensitive to changes in age.
Algorithm for PDP:
1. Train a machine learning model on your dataset, such as a random forest or a
gradient boosting model.
2. Choose a feature of interest that you want to investigate using PDP.
3. Select a range of values for the chosen feature.
4. For each value in the range, create a new dataset by setting the feature of interest
to the selected value and keeping all other features constant.
5. Use the trained machine learning model to predict the output for each dataset
created in Step 4.
6. Compute the average prediction for each value of the feature of interest.
7. Plot the feature values against the average predictions to create a PDP.
3.3.6 InTrees
Decision trees are highly interpretable and are widely used in explainable AI, particularly in applications where transparency and interpretability are critical, such as
healthcare and nance. They can be used to diagnose medical conditions based on
symptoms or predict the likelihood of nancial fraud. Decision trees can also be

Explainable AI forBig Data Control
https://t.me/med1917
151
combined with other decision trees to create more complex models such as random
forests and gradient boosting trees, which improve their predictive power while
maintaining interpretability. Another advantage of decision trees is their ability to
handle both categorical and continuous features, making them a versatile tool for
many applications.
A case study example of decision trees in explainable AI is their use in diagnosing breast cancer. In this application, a decision tree is trained on a dataset of patient
records containing various features such as age, tumor size, and hormone receptor
status. The decision tree is then used to predict whether a patient is likely to have
malignant or benign breast cancer based on their individual features. By examining
the decision path of the tree, doctors can understand how the model arrived at its
prediction and use this information to make informed decisions about treatment.
Algorithm:
1. Initialize the tree with the root node.
2. For each feature, calculate its impurity or information gain using a measure such
as Gini index or entropy.
3. Choose the feature with the highest information gain and split the data at that
feature.
4. Repeat Steps 2 and 3 for each child node until a stopping criterion is met, such
as a minimum number of samples or a maximum depth.
5. Assign a class label or numerical value to each leaf node based on the majority
class or average value of the samples in that node.
6. Repeat the process for each tree in an ensemble model, such as a random forest.
3.3.7 Explainable Boosting Model
Explainable boosting algo will help to improve the accuracy and select the nal best
model by combining more than one model. Each model has one dimensional for a
specic feature. All multimodels for multifeatures are interpreted for effective outcomes. The prediction of errors is used to adjust the weights in the models. EBM
model uses the data and plots the graphs based on the features as input and SHAP
and GAM are combined to interpret the results.
3.3.8 XAI Amortized andNon-Amortized Acceleration Approaches
Amortized model acceleration is divided again into a predictive model, a generative
model, and a reinforcement agent. Non-amortized algorithms are data centric and
model centric acceleration methods. The accelerations are used to explain the model
and amortizations are used to replace the training patterns based on the result of post
hoc local explanations. Feature candidacy reduction models are used to remove the
features and make feature selections, which are precisions in the predicational
model. Remove the less -weighted features from the training model. Feature

152
https://t.me/med1917
R. Aluvalu et al.
searching optimization methods are used to design the formula to optimize the
problem. Adjusting the values of the variables for efcient analysis. Model centric
acceleration is adjusted to improvise the outcome by proposing new model design.
Optimization driven acceleration is used to replace the model as per task oriented.
Reformulate the training attribute to improve the speed and performance.
4 Conclusion
XAI is used in advanced automated applications. The outcomes of AI applications
in some cases are unpredictable and the reasons for results obtained are unknown.
The XAI approach helps to open up the process ow from the input layer to output
extraction. In deep neural networks, XAI helps to explain the model, feature selections, weight adjustment, and so on. The chapter discussed the optimizations of the
features, feature selection, and so on to make it easy to identify the important factors
affecting the result of the model. The case studies for each algorithm clearly remove
the black box algorithms into a white box to improve and maintain the transparency
of AI algorithms.
References
1. Ioannidis JP (2023) Systematic reviews for basic scientists: a different beast. Physiol Rev
103(1):1–5
2. Grajdura S, Niemeier D (2023) State of programming and data science preparation in civil
engineering undergraduate curricula. J Civil Eng Educ 149(2):04022010
3. Prashanth MS, Reddy PVP, Swapna M (2023) AI enabled chat bot for COVID’19. In:
Proceedings of the 14th International conference on soft computing and pattern recognition
(SoCPaR 2022). Springer Nature Switzerland, Cham, pp700–708
4. Serey J, Quezada L, Alfaro M, Fuertes G, Vargas M, Ternero R, Sabattin J, Duran C, Gutierrez
S (2021) Articial intelligence methodologies for data management. Symmetry 13(11):2040
5. Wang C, Yin L (2023) Dening urban big data in urban planning: literature review. J Urban
Plann Dev 149(1):04022044
6. Talaoui Y, Kohtamäki M, Ranta M, Paroutis S (2023) Recovering the divide: a review of the
big data analytics—strategy relationship. Long Range Plann 56:102290
7. Zheng D, Hu D (2023) Recognition method of news dissemination pattern based on computeraided technology in the era of internet of things, p110
8. Manimozhi N, Suganya R, Pandian PS, Suguna G, Devi RS, Ramya S (2023) A compressive
review models for big data analytics relies on articial intelligence. Int Res J Mod Eng Technol
Sci 5:415
9. Zhang B, Zhu J, Su H (2023) Toward the third-generation articial intelligence. Sci China Inf
Sci 66(2):1–19
10. Yu KH, Beam AL, Kohane IS (2018) Articial intelligence in healthcare. Nat Biomed Eng
2(10):719–731
11. Chennam KK, Uma Maheshwari V, Aluvalu R (2021) Maintaining IoT healthcare records using
cloud storage. In: IoT and IoE driven smart cities. Springer International, Cham, pp215–233

Explainable AI forBig Data Control
https://t.me/med1917
12. Dwivedi R, Dave D, Naik H, Singhal S, Omer R, Patel P, Qian B, Wen Z, Shah T, Morgan G,
Ranjan R (2023) Explainable AI (XAI): core ideas, techniques, and solutions. ACM Comput
Surv 55(9):1–33
13. Mendes C, Rios TN (2023) Explainable articial intelligence and cybersecurity: a systematic
literature review. arXiv preprint arXiv:2303.01259
14. Ahmad H, Hanandeh R, Alazzawi F, Al-Daradkah A, ElDmrat A, Ghaith Y, Darawsheh S
(2023) The effects of big data, articial intelligence, and business intelligence on e-learning
and business performance: evidence from Jordanian telecommunication rms. Int J Data Netw
Sci 7(1):35–40
15. Saraladevi B, Pazhaniraja N, Paul PV, Basha MS, Dhavachelvan P (2015) Big data and
Hadoop—a study in security perspective. Procedia Comput Sci 50:596–601
16. Santos MY, e Sá JO, Andrade C, Lima FV, Costa E, Costa C, Martinho B, Galvão J (2017) A
big data system supporting Bosch Braga industry 4.0 strategy. Int J Inf Manag 37(6):750–760
17. Quille RVE, de Almeida FV, Ohara MY, Corrêa PLP, de Freitas LG, Alves-Souza SN, de
Almeida JR Jr, Davis M, Prakash G (2023) Architecture of a data portal for publishing and
delivering open data for atmospheric measurement. Int J Environ Res Public Health 20(7):5374
18. Fernández-Gómez AM, Gutiérrez-Avilés D, Troncoso A, Martínez-Álvarez F (2023) A
new apache spark-based framework for big data streaming forecasting in IoT networks. J
Supercomput 79:1–23
19. Janković S, Mladenović S, Zdravković S, Vesković S, Uzelac A.Mongodb databases in big
data applications in transportation industry
20. Bohar B, Fazekas D, Madgwick M, Csabai L, Olbei M, Korcsmáros T, Szalay-Beko M (2023)
Sherlock: an open-source data platform to store, analyse and integrate big data for computational biologists. F1000Research 10(409):409
21. Hao M, Wang X (2023) Telemetry data processing and analysis platform for ight test based
on Flink. J Phys Conf Ser 2480(1):012020
22. Barroso-Moreno C, Rayon-Rumayor L, García-Vera AB (2023) Big data and business intelligence on Twitter and Instagram for digital inclusion. Comunicar 31(74):49–60
23. Tall AM, Zou CC (2023) A framework for attribute-based access control in processing big data
with multiple sensitivities. Appl Sci 13(2):1183
24. Ullah F, Salam A, Abrar M, Amin F (2023) Brain tumor segmentation using a patch-based
convolutional neural network: a big data analysis approach. Mathematics 11(7):1635
25. Li Z, Pi X, Park Y (2023) S/C: speeding up data materialization with bounded memory. arXiv
preprint arXiv:2303.09774
26. Javed AR, Khan HU, Alomari MKB, Sarwar MU, Asim M, Almadhor AS, Khan MZ
(2023) Toward explainable AI-empowered cognitive health assessment. Front Public Health
11:1024195
27. Guo W (2020) Partially explainable big data driven deep reinforcement learning for green 5G
UAV. In: ICC 2020–2020 IEEE international conference on communications (ICC). IEEE,
Piscataway, NJ, pp1–7
28. Hasanpour Zaryabi E, Moradi L, Kalantar B, Ueda N, Halin AA (2022) Unboxing the black
box of attention mechanisms in remote sensing big data using XAI. Remote Sens (Basel)
14(24):6254
29. Baldominos A, Albacete E, Saez Y, Isasi P (2014) A scalable machine learning online service
for big data real-time analysis. In: 2014 IEEE symposium on computational intelligence in big
data (CIBD). IEEE, Piscataway, NJ, pp1–8
30. Birjali M, Beni-Hssane A, Erritali M (2018) Evaluation of high-level query languages based
on MapReduce in big data. J Big Data 5:1–21
31. Mohammed HH, Doğdu E, Choupani R, Zarbega TS (2023) Distributed query processing and
reasoning over linked big data. In: The recent advances in Transdisciplinary Data Science:
First Southwest Data Science Conference, SDSC 2022, Waco, TX, USA, March 25–26, 2022,
Revised selected papers. Springer Nature Switzerland, Cham, pp158–170
32. Chennam KK, Mudrakola S, Maheswari VU, Aluvalu R, Rao KG (2022) Black box models for
eXplainable articial intelligence. In: Explainable AI: foundations, methodologies and applications. Springer International, Cham, pp1–24
153

Patient Data Analytics Using XAI: Existing
https://t.me/med1917
Tools andCase Studies
SrinivasJagirdar, VijayaKumarVakulabharanam,
ShyamaChandraPrasad G, andAnithaBejugama
Abstract Articial intelligence (AI)-based systems have found extensive applica-
tion within the healthcare sector. In healthcare, AI systems predominantly offer recommendations to physicians based on the analysis of patient health data. Typically,
a doctor reviews diagnostic reports, assesses patient symptoms, and subsequently
arrives at a diagnosis grounded in a comprehensible rationale. In contrast, when an
AI system endeavors to emulate a doctor’s decision-making process, it can invite
criticism due to its adherence to a black box approach when making critical determinations about patient well-being. This approach raises the potential for queries
encompassing medical-legal, ethical, and societal dimensions concerning the guidance provided by the AI model. Consequently, there exists an imperative for the
integration of explainable AI (XAI) within patient data analytics (PDA). XAI serves
as an essential requirement within this context, unveiling the decision-making procedures previously veiled within the opaque construct of deep learning’s black box
model. This chapter casts a spotlight on the pivotal role of XAI within medical
systems. It accomplishes this by delving into the essence of XAI, elucidating its
various categories, exploring the algorithms harnessed to unveil concealed information within black box systems, and addressing the challenges inherent to
XAI.Furthermore, the chapter offers guidance to its readers on constructing intelligible deep learning models tailored for patient data analytics.
S. Jagirdar (*) · S. C. Prasad G
Department of IT, Matrusri Engineering College, Hyderabad, Telangana, India
e-mail: drjsrinivas@matrusri.edu.in; prof.shyam@matrusri.edu.in
V. K. Vakulabharanam
CSE, Anurag University, Hyderabad, Telangana, India
e-mail: dean_rd@anurag.edu.in
A. Bejugama
Department of IT, Bhoj Reddy Engineering College, Hyderabad, Telangana, India
e-mail: anitha.bejugama@slvedu.in
Ltd. 2024
R. Aluvalu et al. (eds.), Explainable AI in Health Informatics, Computational
Intelligence Methods and Applications,
https://doi.org/10.1007/978-981-97-3705-5_8
155© The Author(s), under exclusive license to Springer Nature Singapore Pte

156
https://t.me/med1917
Keywords Patient data · Disease detection · Machine learning · Patient data
analytics · Black box
J. Srinivas et al.
1 An Introduction toExplainable AI
1.1 What Is XAI?
In recent years, eXplainable Articial Intelligence (XAI) has gained signicant
popularity due to its promise of imbuing AI-based systems with traits such as reliability, compliance, comprehensiveness, and accuracy [1]. As human reliance on AI
for day-to-day activities grows, the transparency and explainability of AI systems in
their decision-making processes become increasingly crucial. A notable instance
from 2018 involves a well-known company discontinuing its AI-based recruitment
software due to its demonstrated bias against women in certain job roles [2]. Many
qualied women were unfairly rejected without receiving adequate explanations.
Such instances of biased decisions have the potential to arise in AI systems that
manage critical applications, such as aircraft autopiloting, defense operations, vaccine or drug manufacturing, nancial services, healthcare, and patient diagnosis.
The remarkable proliferation of AI in these domains can be attributed to factors
such as the availability of extensive datasets, the widespread expansion of GPUs,
and the utilization of advanced algorithms [3]. Presently, a signicant portion of AI
systems employs a “black box” methodology, leaving users without insight into the
decision-making mechanisms of the AI system. Given the potential hazards tied to
AI systems, legislative bodies are endeavoring to regulate AI by introducing new
laws such as the “Algorithmic Accountability Act of 2019,” establishing rights to
explanation and establishing ethical principles for articial intelligence [4–6]. XAI
can be dened as the practice of crafting AI systems that possess an inherent capacity for explanation. This practice revolves around unveiling the inner workings that
remain obscured within black box-centric AI systems [7]. The distinction between
conventional AI and XAI is vividly illustrated in Fig.1.
XAI nds its multifaceted utility across diverse domains, encompassing pivotal
areas such as e-commerce, healthcare, biomedicine, education, agriculture, nance,
business, robotics, automotive technology, surveillance, travel, and entertainment
[8]. The comprehensive reach of XAI is graphically depicted in Fig.2.
The rest of the chapter is organized as follows. Section 1.2 gives clarity on the
signicance of XAI in patient data analytics (PDA). The challenges that XAI has to
face for successful adaptation to PDA are discussed in Sect. 1.3. Methods available
for implementing XAI for PDA are introduced in Sect. 1.4. The next section of the
chapter proposes a generic architecture of any XAI model for PDA.In Sect. 1.6 the
various tool kits and frameworks available for implementing XAI in PDA are
explained. Section 1.7 introduces ve case studies in PDA where XAI can
Соседние файлы в папке Библиотека им академика М.И. Перельмана
