Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_117_библиотеки_им_акад_М_И_Перельмана
.pdf
136
https://t.me/med1917
R. Aluvalu et al.
1 Introduction
Data is a source and is important in the information industry. Data science helps to
improve the lifestyle of people. Inference, knowledge, and information are made
from the data. Data is used to monitor the quality system and resolve issues. Data
helps to analyze the performance of the updated strategy and nd dependence features and relationships among variables. Data evidence will help justify the statement and identify the positive aspects. Data helps to set benchmarks, monitor, and
measure performance. The data are collected, stored, and analyzed, and outcomes
are generated [1]. Primary sources for data collection are surveys, observations, and
experimental results, which record live data. Secondary sources for data collection
are digital libraries, government records, businesses, trading, organizations, and
university research [2].
The big data layers are shown in Fig.1.
Articial intelligence (AI) methodologies such as natural language processing
(NLP), expert advisor (EA), speech recognition (SP), computer vision (CV), searching, sorting, machine learning (ML), facial recognition, text analysis, intelligent
recommendation system, AI chatbot, prediction, and forecasting advanced AI methods are used to solve real-time problems [3]. The AI technology implemented in big
data uses reasoning, robotics, general intelligence, NLP, machine learning, and
automated learning are used for business analysts.
Fig. 1 Different layers of big data and big control layer

Explainable AI forBig Data Control
https://t.me/med1917
Programming languages used to implement AI methods are Java, R programming, Python coding, and C++. Articial intelligence uses large amounts of data
and analysis to create rules. Reasoning is another approach to selecting the appropriate algorithm to solve the problem and extract the expected results. Existing algorithms are updated or modied to extract more tting results. The AI advantages are
task-oriented, data processing in less time, reduced manpower by automation, and
consistency in results. The disadvantages of AI are that more intelligent people are
required to remove human intervention [4].
Big data has a large amount of data, is complex due to a large variety of structures and types of data, and requires processing fast. Conventional methods are not
sufcient to process huge amounts of data with a variety of values, veracity, and
velocity [5]. General data analytics methods are regression analysis, ratio analysis,
trend analysis, quantity analytics, qualitative analytics, predictive data analytics,
and descriptive data analytics. The above methods are used to process data and
visualize the data [6]. Big data processing consists of data integration, analysis,
mining, machine learning, optimization, visualization, domain knowledge, process
modeling, and business process management [7].
Big data analytics is needed to process the data phasewise to automatically analyze the data, upgrade the preparation of the data, visualize the data, use predictive
methods, and perform advanced analysis. AI methods are used to analyze the patterns of the data, descriptive analysis, predictive analysis, and predictive analysis
[8]. AI theories to process data are machine learning methods to extract insights
from the data using operational research, statistical analysis, and a neural network
model to nd connectivity among the data and undened data. Deep learning methods with complex computation are applied to data to extract patterns and insights.
Natural language processing is used for computerized understanding and analysis of
human speech [9].
137
1.1 Big Data andAI Applications
Big data requires well-structured and well-designed applications to identify the patterns and know the consistency of the data. The huge database consists of unrelated
data, duplicate data, and unrelated or unwanted data that need to be removed before
loading or processing the data. Advanced AI methods are used to identify the above
data and remove or clean it. Big data and AI create future applications in various
domains, such as health care screening for diabetic retinopathy, which is a method
that records the retinal images on the cloud and compares the differences. The mental assessment is based on text message communications to check on mental health
care. Predicting the patient’s individual health conditions before and after surgery
by analyzing the vitals. The nurses gain experience by clarifying the health suggestion and addressing health issues all over the day (24/7), instead of utilizing the free
time for admin work and nurse station processing work. Medical transcription tasks
can be automated by converting voice data into text. Predicting the room vacancy

138
https://t.me/med1917
R. Aluvalu et al.
based on the status and condition of the patient will reduce the waiting time. AI is
used to process patients’ health details, status, and supply of medicines and food at
the hospital. It will process 4000 people’s data in a 45-bed hospital for 1year
[10, 11].
1.2 Explainable Articial Intelligence (XAI)
Articial intelligence is used to solve problems and automate processes to extract
results. Different AI and ML methods are used to classify and predict outcomes.
The AI is a black box, where the AI -designed model is not explained, and the reason for the classication of classes and prediction of the outcome is not explained.
XAI is a white box, where the model design is explained and reasonable for the
outcome [12]. Explainable AI is used in different layers such as explainable data,
explainable predictors, and explainable algorithms. Explainable AI when loading
data explain what data and why specic data are used to train the model. Explainable
AI is used to extract the features and represent the weights of the features to predict
the outcome. The selection of features from all the variables is made, and the reason
for the selected features is given [13] (Fig.2).
Fig. 2 Explainable articial intelligence goals

Explainable AI forBig Data Control
https://t.me/med1917
139
1.3 Big Data Control Challenges
Big data challenges are enormous data storage, data processing on huge content and
variety of data, control on quality of the data, scale up and scale down the data quantity of data (real-time insights). Environmental management of storage devices, data
validations, security, big data analyst are common challenges in managing the big
data. Advanced AI methods help to address the challenges.
1.3.1 Automated Building andManagement ofData Storages
Security system in managing the physical data storage devices by monitoring the
digital survival cameras, security system by authorization and access control,
AI-based re alarm system, light control over low maintained locations, smoke
detectors, and control system by automated closing emergency doors and open ventilators are important safety and security maintained through AI technology. XAI is
used to explain the behavior of the model decision based on model-specic or datacentric designs.
1.3.2 Big Data Processing Using Algorithms
Big data processing uses MapReduce and Hadoop methods to process the parallel
programming, pattern process identication, detect the predictions, and make decisions. Collect the big data through external sources, analyze the big data through
classication, contextualize, categorize, process, and distribute data. The data mining algorithms are used to extract the data, and machine learning methods are used
to build the model and predict the new scenarios. Data visualization methods are
used to generate the reports and perform the online application processing the data.
Huge data is collected from IoT sensors, mobile application data, social media and
multimedia, nancial transactions, online purchases, and business transactions. Big
data clustering sample -based techniques, dimensional reduction, parallel clustering, and MapReduction approaches are used to cluster the data. Big data classication methods are traditional learning methods. Decision tree, regression methods,
SVM, ANN, optimization methods, deep learning, active learning, transfer learning,
kernel-based learning, representation learning, and parallel learning are some of the
advanced classication methods used in big data.
1.3.3 Big Data Visualizations
Big data visualization is not visually friendly to understand in graphs. The variables
are plotted nearby and cannot be differentiated easily based on the user. The data
loss is when insights are plotted publicly. Big data visualization tools such as

140
https://t.me/med1917
Matlab, Tableau, Microsoft Power BI, and Zoho help to compress huge data in a
simple form, detect the outlier from the graph, understand the trend of the data,
understand the pattern of data, and communicate effectively..
R. Aluvalu et al.
2 Literature Study
Articial intelligence in big data helps to perform data analytics. The advanced AI
methods for business performance on big data to increase the performance of the
business and data control means maintaining data accuracy, transparency in transactions, speed in data operations, creativity in processing, and learning the data from
training data [14]. Business intelligence always has a great impact on maintaining
the data warehouse, extracting the data, and managing processes in business.
Business intelligence will learn customer behavior and reduce expenditures, product prices, and processing fees [15]. The literature survey about big data at different
layers and technologies is shown in Table1.
2.1 Big Data Control Using Articial Intelligence
The advanced methods of articial intelligence are used in different layers of big
data. AI storage systems are used to process the data in the cloud or on big storage
systems. The performance of the storage device is measured in terms of scalability
in storage capacity change requirements, reliability, and very fast access (Fig.3).
AI Support Features on Storage Devices
1. Repaid in access from data centers to cloud edges.
2. Performance in demand and overcome the bottleneck problem.
3. Cost effectiveness in maintaining the device and the purchase cost.
4. Authentication of the storage device in terms of authentication, role-based
permissions.
5. Dynamic routing and load balance.
6. Graphical processing unit (GPU) accelerated devices.
7. High-performance computing clusters (to build advanced AI storage devices).
8. Solid state hard drives create an intermediate layer to access fast.
9. Non-volatile memory express (NVMe) for superfast access to the data.
AI data ingestion is the process of transforming data from multiple sources into
centralized data centers. The centralized data centers are analyzed and use the real
data or transfer content batchwise. AI can address the data ingestion challenges.
Transform the old data into new formats. Data transformation delays need to be
redirected to another application. Data quality is also clean from error data and
automated management in throughput [29].

Explainable AI forBig Data Control
https://t.me/med1917
Big data tools used in different layers and methodology
Table 1
Literature survey on the big data control in big data layers
Author Tools Big data layer Big data control method
Santos etal.
[16]
Quille etal.
[17]
FernándezGómez etal.
[18]
Janković etal.
[19]
Bohar etal.
[20]
Hao and
Wang [21]
Hadoop
Zoho Analytics
Atlas.ti
Storm
CouchDB
StatsiQ
Talend
Spark
Hybrid Data Collection
Airbyte
Hevo
Amazon Kinesis
Apache Flume
Apache Gobblin
Apache Kafka
Apache NiFi,
Dropbase, Integrate.io
HibariDB
MongoDB
Cloudera
Hadoop
Apache Cassandra
Apache HBase
Terrastore, Matillion
Pentaho
MapReduce
Data Cleaner
Open Rene
Talend
Google Cloud Platform
Sisense, Hadoop
SparkSQL
Presto
Amazon RedShift
Data sources Compatible le system
Data
collection
layer
Data ingestion
layer
Data storage
layer
Data
processing
layer
Data query
layer
Insight dashboard
Cloud driven
Platform
Speed data access
Duplicate data maintain
Data integration
Real-time data
Extracting loading
Data ow, data pipeline
Data Strems from sources
Load data into HDFS
Load large data from multiple
sources
Data export and import
Automated ow data
Transform online into ofine
Storage of big data
Scalable platform and portable
NoSQL database
Storing key/value based
scalability, elasticity,
consistency, ETL tool
Data integration, extraction
Data preparing
Advance reporting algo
Data transformation
Data validation, cleaning,
searching, translation
Sorting, clustering
ELT operations
Analyze data
Summarization
Ad hoc query
Interactive query
141
(continued)

142
https://t.me/med1917
Table 1 (continued)
Literature survey on the big data control in big data layers
Author Tools Big data layer Big data control method
Barroso-
Moreno etal.
[22]
Tall and Zou
[23]
Ullah etal.
[24]
Li etal. [25] DICE Monitoring Platform,
Javed etal.
[26]
Guo [27] Double Dueling Deep
Hasanpour
Zaryabi etal.
[28]
Hive, KNIME Analytics
Datameer, Apache Spark,
Apache Storm
Flink, High Performance
Computing Cluster (HPCC)
RapidMiner, Weka, Orange,
Neural Designer
Tableau
Microsoft power BI
Data-driven documents
Apache Knox 2.0., Apache
Ranger
Cyral Monitoring
Attn-LSTM Data
Q-learning Neural Network
Prioritized Experience
Replay (PER) and xed
Q-targets
XAI Model are (Layer
Gradient X Activation and
(Layer Gradient X
Activation)
Data analytics
layer
Data
visualization
layer
Data security
layer
Data
monitoring
layer
reasoning
Data analysis Control on UAV -based 5G
Data
processing
Analyze dataset, data mining
Predicting the future, analytic
platform
Large scale Sql, batch
processing
Machine learning, stream
processing
Real-time stream processing,
data analytics supercomputer,
data analytics neural network,
exploratory data analysis
Predictive modeling
Real-time analysis, data
collaboration
Data blending, document object
model
Strategies, interactive
visualization
Encryption, user access control
Physical layer, centralized key
management
Monitoring end-to-end ow
Code level detail
Processing backend
Measure metric of data
Empowering explainable AI
Identify user empowering
through deductive and inductive
approach, derivation and
contactization
system
Aerospace remote sensor, image
analysis
R. Aluvalu et al.
Data processing layers use articial intelligence for big data analytics in big storage systems. Advanced machine learning algorithms, data mining algorithms, and
clustering algorithms are used in the processing layer. The machine learning algorithms are processed in stream mode or batch mode. Machine learning is used for
prediction, advanced segmentation, recommendation systems, neural networks, and
Markova charts. The random forest, linear regression, logistic regression, SVM,

Explainable AI forBig Data Control
https://t.me/med1917
Fig. 3 AI storages advantage on big data
143
algorithms, etc. are used for data processing. The batch processing methods are
model building, clustering, and processing using tools such as Weka, R, and
Mahaout [30].
The “query processor layer” in big data is used to extract the data and extract the
summaries or dashboard data. Reports are generated. The query processing unit
tries to process the input query from the user in different phases. The basic query
processing steps are
1. Query Parser.
2. Query Translator.
3. Query Optimizer.
4. Execute the Query.
5. Evaluate the result.
The advanced query processing methods are used in big data using high-level
query-map-MapReduce-Language languages such as BigSQL, Hive, Pig, and JAQL
[31]. The query is processed by the MapReduce function on the Hadoop distributed
le system. Big data technologies for semantic web data are processed by standard
protocols such as OWL, RDF, and RDFS. The products of RDF are Fedx,
PigSPARQL, SPLENDID, and SPARKQLGX [32].
The data visualization layer is used to represent big data. The data representation
can be one dimension, two dimensions, or three dimensions. The advantages of
using data visualization methods help to make decisions and represent the data
without losing accuracy. Pie charts, bar charts, histograms, Gantt charts, heat maps,

144
https://t.me/med1917
R. Aluvalu et al.
waterfall charts, area charts, and scatter plots are all available. Data visualization
tools are tableau, data wrapper, chart blocks, fusion chart, chartist, and Grafana.
3 Methodology
Explained articial intelligence (XAI) is used to create trust in machine decisions.
General AI methods are Blackbox approaches, which means the output generated
from machine learning is not addressed with reason. The basic principles of
explained AI are explainability, transparency in operations, and interpretability
(Fig.4).
XAI in machine learning models is used to replace the old learning process with
a new learning process with an explainable model and an explainable interface,
which can answer the basic questions of what, why, and how the results were
obtained. It can predict when models will succeed or fail. XAI can be used in the
pre-modeling and post-modeling phases of a machine learning model. The efcient
data preprocessing is directly related to the performance of the model and results.
Pre-modeling phase XAI is used in data analysis and data transformation. The XAI
is used in post-modeling based on model specic and model-agnostic criteria.
The major goals of XAI are represented in Fig.5. The positive aspect of the
model is that it provides trustworthiness, condence, causality, transferability, and
accessibility to the stakeholder or users.
Types of XAI such as interpretable XAI, transparent XAI, and interactable XAI
are broad categories of explainable articial intelligence. Interpretable XAI has the
ability to obtain the causes and also the impact of the machine learning algorithms.
Fig. 4 Basic principle of explainable AI

Explainable AI forBig Data Control
https://t.me/med1917
Fig. 5 Major goals of XAI
145
A simple example is overeating can cause indigestion problems. Transparent XAI is
a model designed openly, detailed in process, and communicable and understandable to all stakeholders. Interactable XAI is query-based learning that converts the
black box to the open box model. Generally, AI algorithms are transparent and
opaque in their design models.
The transparent XAI used to map the reasons for results and conclusions is
explainable. The sample algorithms are classication, regression, and segmentation,
with models such as linear regression, logistic regression, decision trees, K-near
neighbors, rule-based learners, generative additive models, and Bayesian models.
The opaque models are the AI methods that do not give any detailed explanation of
the model or predict the same outcome for all the trials. The process of evaluation is
not known to the stakeholders. The sample models are a random forest, a support
vector machine, and a multilayer neural network.
3.1 Post Hoc Explainability (PHE) inXAI
The AI algorithms with a black box approach are complex for stockholders to manage their businesses, as the obtained results are unable to predict the reason. The
trained model of AI is given as input to the PHE-XAI module, which tries to understand the inner layer functionality and its relation to the outputs. The decision rules
can be generated and formulated into some representation for better understanding.
Some of the explanations are in logical rules, feature ranks, heatmaps, and natural
language.
The PHE-XAI generates explanations for both transparent and opaque AI algorithms. The explanation scope can be specic to the dataset, problem, or a generalized explanation for the model designed can be applied to any type of dataset and
Соседние файлы в папке Библиотека им академика М.И. Перельмана
