Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_117_библиотеки_им_акад_М_И_Перельмана

.pdf
Скачиваний:
2
Добавлен:
15.09.2026
Размер:
15 Мб
Скачать
☆
136
R. Aluvalu et al.
1 Introduction
Data is a source and is important in the information industry. Data science helps to improve the lifestyle of people. Inference, knowledge, and information are made from the data. Data is used to monitor the quality system and resolve issues. Data helps to analyze the performance of the updated strategy and nd dependence fea­tures and relationships among variables. Data evidence will help justify the state­ment and identify the positive aspects. Data helps to set benchmarks, monitor, and measure performance. The data are collected, stored, and analyzed, and outcomes are generated [1]. Primary sources for data collection are surveys, observations, and experimental results, which record live data. Secondary sources for data collection are digital libraries, government records, businesses, trading, organizations, and university research [2].
The big data layers are shown in Fig.1.
Articial intelligence (AI) methodologies such as natural language processing (NLP), expert advisor (EA), speech recognition (SP), computer vision (CV), search­ing, sorting, machine learning (ML), facial recognition, text analysis, intelligent recommendation system, AI chatbot, prediction, and forecasting advanced AI meth­ods are used to solve real-time problems [3]. The AI technology implemented in big data uses reasoning, robotics, general intelligence, NLP, machine learning, and automated learning are used for business analysts.
Fig. 1 Different layers of big data and big control layer
Explainable AI forBig Data Control
Programming languages used to implement AI methods are Java, R program­ming, Python coding, and C++. Articial intelligence uses large amounts of data and analysis to create rules. Reasoning is another approach to selecting the appro­priate algorithm to solve the problem and extract the expected results. Existing algo­rithms are updated or modied to extract more tting results. The AI advantages are task-oriented, data processing in less time, reduced manpower by automation, and consistency in results. The disadvantages of AI are that more intelligent people are required to remove human intervention [4].
Big data has a large amount of data, is complex due to a large variety of struc­tures and types of data, and requires processing fast. Conventional methods are not sufcient to process huge amounts of data with a variety of values, veracity, and velocity [5]. General data analytics methods are regression analysis, ratio analysis, trend analysis, quantity analytics, qualitative analytics, predictive data analytics, and descriptive data analytics. The above methods are used to process data and visualize the data [6]. Big data processing consists of data integration, analysis, mining, machine learning, optimization, visualization, domain knowledge, process modeling, and business process management [7].
Big data analytics is needed to process the data phasewise to automatically ana­lyze the data, upgrade the preparation of the data, visualize the data, use predictive methods, and perform advanced analysis. AI methods are used to analyze the pat­terns of the data, descriptive analysis, predictive analysis, and predictive analysis [8]. AI theories to process data are machine learning methods to extract insights from the data using operational research, statistical analysis, and a neural network model to nd connectivity among the data and undened data. Deep learning meth­ods with complex computation are applied to data to extract patterns and insights. Natural language processing is used for computerized understanding and analysis of human speech [9].
137
1.1 Big Data andAI Applications
Big data requires well-structured and well-designed applications to identify the pat­terns and know the consistency of the data. The huge database consists of unrelated data, duplicate data, and unrelated or unwanted data that need to be removed before loading or processing the data. Advanced AI methods are used to identify the above data and remove or clean it. Big data and AI create future applications in various domains, such as health care screening for diabetic retinopathy, which is a method that records the retinal images on the cloud and compares the differences. The men­tal assessment is based on text message communications to check on mental health care. Predicting the patient’s individual health conditions before and after surgery by analyzing the vitals. The nurses gain experience by clarifying the health sugges­tion and addressing health issues all over the day (24/7), instead of utilizing the free time for admin work and nurse station processing work. Medical transcription tasks can be automated by converting voice data into text. Predicting the room vacancy
138
R. Aluvalu et al.
based on the status and condition of the patient will reduce the waiting time. AI is used to process patients’ health details, status, and supply of medicines and food at the hospital. It will process 4000 people’s data in a 45-bed hospital for 1year [10, 11].
1.2 Explainable Articial Intelligence (XAI)
Articial intelligence is used to solve problems and automate processes to extract results. Different AI and ML methods are used to classify and predict outcomes. The AI is a black box, where the AI -designed model is not explained, and the rea­son for the classication of classes and prediction of the outcome is not explained. XAI is a white box, where the model design is explained and reasonable for the outcome [12]. Explainable AI is used in different layers such as explainable data, explainable predictors, and explainable algorithms. Explainable AI when loading data explain what data and why specic data are used to train the model. Explainable AI is used to extract the features and represent the weights of the features to predict the outcome. The selection of features from all the variables is made, and the reason for the selected features is given [13] (Fig.2).
Fig. 2 Explainable articial intelligence goals
Explainable AI forBig Data Control
139
1.3 Big Data Control Challenges
Big data challenges are enormous data storage, data processing on huge content and variety of data, control on quality of the data, scale up and scale down the data quan­tity of data (real-time insights). Environmental management of storage devices, data validations, security, big data analyst are common challenges in managing the big data. Advanced AI methods help to address the challenges.
1.3.1 Automated Building andManagement ofData Storages
Security system in managing the physical data storage devices by monitoring the digital survival cameras, security system by authorization and access control, AI-based re alarm system, light control over low maintained locations, smoke detectors, and control system by automated closing emergency doors and open ven­tilators are important safety and security maintained through AI technology. XAI is used to explain the behavior of the model decision based on model-specic or data­centric designs.
1.3.2 Big Data Processing Using Algorithms
Big data processing uses MapReduce and Hadoop methods to process the parallel programming, pattern process identication, detect the predictions, and make deci­sions. Collect the big data through external sources, analyze the big data through classication, contextualize, categorize, process, and distribute data. The data min­ing algorithms are used to extract the data, and machine learning methods are used to build the model and predict the new scenarios. Data visualization methods are used to generate the reports and perform the online application processing the data. Huge data is collected from IoT sensors, mobile application data, social media and multimedia, nancial transactions, online purchases, and business transactions. Big data clustering sample -based techniques, dimensional reduction, parallel cluster­ing, and MapReduction approaches are used to cluster the data. Big data classica­tion methods are traditional learning methods. Decision tree, regression methods, SVM, ANN, optimization methods, deep learning, active learning, transfer learning, kernel-based learning, representation learning, and parallel learning are some of the advanced classication methods used in big data.
1.3.3 Big Data Visualizations
Big data visualization is not visually friendly to understand in graphs. The variables are plotted nearby and cannot be differentiated easily based on the user. The data loss is when insights are plotted publicly. Big data visualization tools such as
140
Matlab, Tableau, Microsoft Power BI, and Zoho help to compress huge data in a simple form, detect the outlier from the graph, understand the trend of the data, understand the pattern of data, and communicate effectively..
R. Aluvalu et al.
2 Literature Study
Articial intelligence in big data helps to perform data analytics. The advanced AI methods for business performance on big data to increase the performance of the business and data control means maintaining data accuracy, transparency in transac­tions, speed in data operations, creativity in processing, and learning the data from training data [14]. Business intelligence always has a great impact on maintaining the data warehouse, extracting the data, and managing processes in business. Business intelligence will learn customer behavior and reduce expenditures, prod­uct prices, and processing fees [15]. The literature survey about big data at different layers and technologies is shown in Table1.
2.1 Big Data Control Using Articial Intelligence
The advanced methods of articial intelligence are used in different layers of big data. AI storage systems are used to process the data in the cloud or on big storage systems. The performance of the storage device is measured in terms of scalability in storage capacity change requirements, reliability, and very fast access (Fig.3).
AI Support Features on Storage Devices
1. Repaid in access from data centers to cloud edges.
2. Performance in demand and overcome the bottleneck problem.
3. Cost effectiveness in maintaining the device and the purchase cost.
4. Authentication of the storage device in terms of authentication, role-based
permissions.
5. Dynamic routing and load balance.
6. Graphical processing unit (GPU) accelerated devices.
7. High-performance computing clusters (to build advanced AI storage devices).
8. Solid state hard drives create an intermediate layer to access fast.
9. Non-volatile memory express (NVMe) for superfast access to the data.
AI data ingestion is the process of transforming data from multiple sources into centralized data centers. The centralized data centers are analyzed and use the real data or transfer content batchwise. AI can address the data ingestion challenges. Transform the old data into new formats. Data transformation delays need to be redirected to another application. Data quality is also clean from error data and automated management in throughput [29].
Explainable AI forBig Data Control
Big data tools used in different layers and methodology
Table 1
Literature survey on the big data control in big data layers Author Tools Big data layer Big data control method
Santos etal. [16]
Quille etal. [17]
Fernández­Gómez etal. [18]
Janković etal. [19]
Bohar etal. [20]
Hao and Wang [21]
Hadoop Zoho Analytics Atlas.ti Storm CouchDB StatsiQ
Talend Spark Hybrid Data Collection
Airbyte Hevo Amazon Kinesis Apache Flume Apache Gobblin Apache Kafka Apache NiFi, Dropbase, Integrate.io
HibariDB MongoDB Cloudera Hadoop Apache Cassandra Apache HBase Terrastore, Matillion
Pentaho MapReduce Data Cleaner Open Rene Talend
Google Cloud Platform Sisense, Hadoop SparkSQL Presto Amazon RedShift
Data sources Compatible le system
Data collection layer
Data ingestion layer
Data storage layer
Data processing layer
Data query layer
Insight dashboard Cloud driven Platform Speed data access Duplicate data maintain
Data integration Real-time data
Extracting loading Data ow, data pipeline Data Strems from sources Load data into HDFS Load large data from multiple sources Data export and import Automated ow data Transform online into ofine
Storage of big data Scalable platform and portable NoSQL database Storing key/value based scalability, elasticity, consistency, ETL tool
Data integration, extraction Data preparing Advance reporting algo Data transformation Data validation, cleaning, searching, translation Sorting, clustering
ELT operations Analyze data Summarization Ad hoc query Interactive query
141
(continued)
142
Table 1 (continued)
Literature survey on the big data control in big data layers Author Tools Big data layer Big data control method Barroso-
Moreno etal. [22]
Tall and Zou [23]
Ullah etal. [24]
Li etal. [25] DICE Monitoring Platform,
Javed etal. [26]
Guo [27] Double Dueling Deep
Hasanpour Zaryabi etal. [28]
Hive, KNIME Analytics Datameer, Apache Spark, Apache Storm Flink, High Performance Computing Cluster (HPCC) RapidMiner, Weka, Orange, Neural Designer
Tableau Microsoft power BI Data-driven documents
Apache Knox 2.0., Apache Ranger
Cyral Monitoring
Attn-LSTM Data
Q-learning Neural Network Prioritized Experience Replay (PER) and xed Q-targets
XAI Model are (Layer Gradient X Activation and (Layer Gradient X Activation)
Data analytics layer
Data visualization layer
Data security layer
Data monitoring layer
reasoning
Data analysis Control on UAV -based 5G
Data processing
Analyze dataset, data mining Predicting the future, analytic platform Large scale Sql, batch processing Machine learning, stream processing Real-time stream processing, data analytics supercomputer, data analytics neural network, exploratory data analysis Predictive modeling
Real-time analysis, data collaboration Data blending, document object model Strategies, interactive visualization
Encryption, user access control Physical layer, centralized key management
Monitoring end-to-end ow Code level detail Processing backend Measure metric of data
Empowering explainable AI Identify user empowering through deductive and inductive approach, derivation and contactization
system
Aerospace remote sensor, image analysis
R. Aluvalu et al.
Data processing layers use articial intelligence for big data analytics in big stor­age systems. Advanced machine learning algorithms, data mining algorithms, and clustering algorithms are used in the processing layer. The machine learning algo­rithms are processed in stream mode or batch mode. Machine learning is used for prediction, advanced segmentation, recommendation systems, neural networks, and Markova charts. The random forest, linear regression, logistic regression, SVM,
Explainable AI forBig Data Control
Fig. 3 AI storages advantage on big data
143
algorithms, etc. are used for data processing. The batch processing methods are model building, clustering, and processing using tools such as Weka, R, and Mahaout [30].
The “query processor layer” in big data is used to extract the data and extract the summaries or dashboard data. Reports are generated. The query processing unit tries to process the input query from the user in different phases. The basic query processing steps are
1. Query Parser.
2. Query Translator.
3. Query Optimizer.
4. Execute the Query.
5. Evaluate the result.
The advanced query processing methods are used in big data using high-level query-map-MapReduce-Language languages such as BigSQL, Hive, Pig, and JAQL [31]. The query is processed by the MapReduce function on the Hadoop distributed le system. Big data technologies for semantic web data are processed by standard protocols such as OWL, RDF, and RDFS. The products of RDF are Fedx, PigSPARQL, SPLENDID, and SPARKQLGX [32].
The data visualization layer is used to represent big data. The data representation can be one dimension, two dimensions, or three dimensions. The advantages of using data visualization methods help to make decisions and represent the data without losing accuracy. Pie charts, bar charts, histograms, Gantt charts, heat maps,
144
R. Aluvalu et al.
waterfall charts, area charts, and scatter plots are all available. Data visualization tools are tableau, data wrapper, chart blocks, fusion chart, chartist, and Grafana.
3 Methodology
Explained articial intelligence (XAI) is used to create trust in machine decisions. General AI methods are Blackbox approaches, which means the output generated from machine learning is not addressed with reason. The basic principles of explained AI are explainability, transparency in operations, and interpretability (Fig.4).
XAI in machine learning models is used to replace the old learning process with a new learning process with an explainable model and an explainable interface, which can answer the basic questions of what, why, and how the results were obtained. It can predict when models will succeed or fail. XAI can be used in the pre-modeling and post-modeling phases of a machine learning model. The efcient data preprocessing is directly related to the performance of the model and results. Pre-modeling phase XAI is used in data analysis and data transformation. The XAI is used in post-modeling based on model specic and model-agnostic criteria.
The major goals of XAI are represented in Fig.5. The positive aspect of the model is that it provides trustworthiness, condence, causality, transferability, and accessibility to the stakeholder or users.
Types of XAI such as interpretable XAI, transparent XAI, and interactable XAI are broad categories of explainable articial intelligence. Interpretable XAI has the ability to obtain the causes and also the impact of the machine learning algorithms.
Fig. 4 Basic principle of explainable AI
Explainable AI forBig Data Control
Fig. 5 Major goals of XAI
145
A simple example is overeating can cause indigestion problems. Transparent XAI is a model designed openly, detailed in process, and communicable and understand­able to all stakeholders. Interactable XAI is query-based learning that converts the black box to the open box model. Generally, AI algorithms are transparent and opaque in their design models.
The transparent XAI used to map the reasons for results and conclusions is explainable. The sample algorithms are classication, regression, and segmentation, with models such as linear regression, logistic regression, decision trees, K-near neighbors, rule-based learners, generative additive models, and Bayesian models. The opaque models are the AI methods that do not give any detailed explanation of the model or predict the same outcome for all the trials. The process of evaluation is not known to the stakeholders. The sample models are a random forest, a support vector machine, and a multilayer neural network.
3.1 Post Hoc Explainability (PHE) inXAI
The AI algorithms with a black box approach are complex for stockholders to man­age their businesses, as the obtained results are unable to predict the reason. The trained model of AI is given as input to the PHE-XAI module, which tries to under­stand the inner layer functionality and its relation to the outputs. The decision rules can be generated and formulated into some representation for better understanding. Some of the explanations are in logical rules, feature ranks, heatmaps, and natural language.
The PHE-XAI generates explanations for both transparent and opaque AI algo­rithms. The explanation scope can be specic to the dataset, problem, or a general­ized explanation for the model designed can be applied to any type of dataset and