Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_145_библиотеки_им_акад_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
36 Мб
Скачать
Brain Computer Interaction Disease Prediction using Machine Learning 143
https://t.me/med1917
the mental and emotional state throughout sessions. Internal non-stationary causes include fatigue and concentration levels as well. Noise and non-stationary causes are major challenges for BCI technology. They involve unnecessary signals triggered by changes in electrode placement as well as noise from the environment. The acquired signals represent a combination of moving objects, such as electrical activity generated by signals created by eye movements and skeletal muscles electromyogram (EMG) and blinking Electrooculogram (EOG), making it difficult to discern the underlying pattern.
Small Training Sets
The training procedure is dominated by usability problems and limited training sets. Although the subjects find the training sessions to be time-consuming and exhausting, they provide the user with the requisite experience to cope with the device and learn to monitor his or her neurophysiologic signals. As a result, developing a BCI requires a significant balance between the technical difficulty of understanding the brain's signals and training requirements for effective usage.
CONCLUSION
Brain-computer interface is a cutting-edge technology that uses brain impulses to control an external device, bypassing the regular neuromuscular system, to complete a task. BCI technology is a relatively new innovation in neuroscience that has drawn researchers from a variety of fields, including entertainment and gaming, security, and marketing. Invasive and non-invasive recording equipment are classified into two groups. Invasive surgery is frequently required for essential paralyzed conditions due to its higher accuracy rates obtained either geographically or temporally. Besides its advantages over the invasive category, which we have previously discussed, the non-invasive category has also been widely adopted in a variety of applications. Other concerns and obstacles that arise as a result of using brain signals have also been examined as the result of the answers provided by algorithms at BCI system processing. The interaction possibilities of new technologies are fully appreciated by HCI researchers today. We also look at the different challenges that must be overcome in order for BCI to be a successful and widely adopted technology. The paper's primary goal is to bring this new emergent technology to the forefront. The BCI research community hopes to work more closely together in the future and accepts that other research areas and society in general will value and direct BCI research, than other medical applications.
REFERENCES
[1] S. Ram, W. Zhang, M. Williams, and Y. Pengetnze, "Predicting asthma-related emergency department
visits using big data", IEEE J. Biomed. Health Inform., vol. 19, no. 4, pp. 1216-1223, 2015.
144 Disease Prediction using Machine Learning M. Kiruthiga Devi
https://t.me/med1917
[http://dx.doi.org/10.1109/JBHI.2015.2404829] [PMID: 25706935]
[2] C.H. Lee, J.C.Y. Chen, and V.S. Tseng, "A novel data mining mechanism considering bio-signal and
environmental data with applications on asthma monitoring", Comput. Methods Programs Biomed., vol. 101, no. 1, pp. 44-61, 2011. [http://dx.doi.org/10.1016/j.cmpb.2010.04.016] [PMID: 20554074]
[3] W.C. Moore, D.A. Meyers, S.E. Wenzel, W.G. Teague, H. Li, X. Li, R. D’Agostino Jr, M. Castro, D.
Curran-Everett, A.M. Fitzpatrick, B. Gaston, N.N. Jarjour, R. Sorkness, W.J. Calhoun, K.F. Chung, S.A.A. Comhair, R.A. Dweik, E. Israel, S.P. Peters, W.W. Busse, S.C. Erzurum, and E.R. Bleecker, "Identification of asthma phenotypes using cluster analysis in the Severe Asthma Research Program", Am. J. Respir. Crit. Care Med., vol. 181, no. 4, pp. 315-323, 2010. [http://dx.doi.org/10.1164/rccm.200906-0896OC] [PMID: 19892860]
[4] C.K.W. Lai, R. Beasley, J. Crane, S. Foliaki, J. Shah, and S. Weiland, "Global variation in the
prevalence and severity of asthma symptoms: Phase Three of the International Study of Asthma and Allergies in Childhood (ISAAC)", Thorax, vol. 64, no. 6, pp. 476-483, 2009. [http://dx.doi.org/10.1136/thx.2008.106609] [PMID: 19237391]
[5] C. Shivade, P. Raghavan, E. Fosler-Lussier, P.J. Embi, N. Elhadad, S.B. Johnson, and A.M. Lai, "A
review of approaches to identifying patient phenotype cohorts using electronic health records", J. Am. Med. Inform. Assoc., vol. 21, no. 2, pp. 221-230, 2014. [http://dx.doi.org/10.1136/amiajnl-2013-001935] [PMID: 24201027]
[6] W. Wu, E. Bleecker, W. Moore, W.W. Busse, M. Castro, K.F. Chung, W.J. Calhoun, S. Erzurum, B.
Gaston, E. Israel, D. Curran-Everett, and S.E. Wenzel, "Unsupervised phenotyping of severe asthma research program participants using expanded lung data", J. Allergy Clin. Immunol., vol. 133, no. 5, pp. 1280-1288, 2014. [http://dx.doi.org/10.1016/j.jaci.2013.11.042] [PMID: 24589344]
[7] M. Zedan, G. Attia, M.M. Zedan, A. Osman, N. Abo-Elkheir, N. Maysara, T. Barakat, and N. Gamil,
"Clinical asthma phenotypes and therapeutic responses", ISRN Pediatr., vol. 2013, pp. 1-7, 2013. [http://dx.doi.org/10.1155/2013/824781] [PMID: 23606983]
[8] M.C.F. Prosperi, U.M. Sahiner, D. Belgrave, C. Sackesen, I.E. Buchan, A. Simpson, T.S. Yavuz, O.
Kalayci, and A. Custovic, "Challenges in identifying asthma subgroups using unsupervised statistical learning techniques", Am. J. Respir. Crit. Care Med., vol. 188, no. 11, pp. 1303-1312, 2013. [http://dx.doi.org/10.1164/rccm.201304-0694OC] [PMID: 24180417]
[9] M. Zolnoori, M. H. F. Zarandi, and M. Moin, "Application of intelligent systems in asthma disease:
designing a fuzzy rule-based system for evaluating level of asthma exacerbation", J Med Syst, vol. 36, no. 4, pp. 2071-2083, 2012. [http://dx.doi.org/10.1007/s10916-011-9671-8]
[10] V. Siroux, X. Basagaña, A. Boudier, I. Pin, J. Garcia-Aymerich, A. Vesin, R. Slama, D. Jarvis, J.M.
Anto, F. Kauffmann, and J. Sunyer, "Identifying adult asthma phenotypes using a clustering approach", Eur. Respir. J., vol. 38, no. 2, pp. 310-317, 2011. [http://dx.doi.org/10.1183/09031936.00120810] [PMID: 21233270]
[11] C. Chakraborty, T. Mitra, A. Mukherjee, and A.K. Ray, "CAIDSA: Computer-aided intelligent
diagnostic system for bronchial asthma", Expert Syst. Appl., vol. 36, no. 3, pp. 4958-4966, 2009. [http://dx.doi.org/10.1016/j.eswa.2008.06.025]
[12] M.P. Pushpalatha, and M.R. Pooja, "A predictive model for the effective prognosis of Asthma using
Asthma severity indicators", 2017 International Conference on Computer Communication and Informatics (ICCCI), pp. 1-6, 2017. [http://dx.doi.org/10.1109/ICCCI.2017.8117717]
[13] S. Schmidt, G. Li, and Y.-P. P. Chen, "Medical knowledge discovery from a regional asthma dataset",
Advanced Intelligent Computing Theories and Applications. With Aspects of Artificial Intelligence, pp. 888-895, 2008.
Brain Computer Interaction Disease Prediction using Machine Learning 145
https://t.me/med1917
[http://dx.doi.org/10.1007/978-3-540-85984-0_107]
[14] M.R. Pooja, and M.P. Pushpalatha, "A hybrid decision support system for the identification of
asthmatic subjects in a cross-sectional study", 2015 International Conference on Emerging Research in Electronics, Computer Science and Technology (ICERECT), pp. 288-293, 2015. [http://dx.doi.org/10.1109/ERECT.2015.7499028]
[15] K. Farion, W. Michalowski, and S. Wilk, Developing a Decision Model for Asthma Exacerbations:
Combining Rough Sets and Expert-Driven Selection of Clinical Attributes. Rough Sets and Current Trends in Computing, 2006, pp. 428-437. [http://dx.doi.org/10.1007/11908029_45]
[16] M.R. Pooja, and M.P. Pushpalatha, "A comparative performance evaluation of hybrid and ensemble
machine learning models for prediction of asthma morbidity", J. Health Med. Inform., vol. 10, no. 330,
2019. [http://dx.doi.org/10.4172/2157-7420.1000330]
[17] P. Mr, and P. Mp, "Analysis of a panel of cytokines in BAL fluids to differentiate controlled and
uncontrolled asthmatics using machine learning model", J. Respirat. Res., vol. 5, no. 1, pp. 142-145,
2019. [http://dx.doi.org/10.17554/j.issn.2412-2424.2019.05.44] [PMID: 31286968]
[18] E.M.S. Mäkikyrö, M.S. Jaakkola, and J.J.K. Jaakkola, "Subtypes of asthma based on asthma control
and severity: A latent class analysis", Respir. Res., vol. 18, no. 1, pp. 24-32, 2017. [http://dx.doi.org/10.1186/s12931-017-0508-y] [PMID: 28114991]
[19] M.R. Pooja, and M.P. Pushpalatha, "Cluster analysis to characterize the patterns of complementary
and alternative medicines usage in asthma controls", Open Public Health J., vol. 13, p. 1, 2020. [http://dx.doi.org/10.2174/1874944502013010227]
[20] G. Rani, A. Misra, V.S. Dhaka, D. Buddhi, R.K. Sharma, E. Zumpano, E. Vocaturo, "A multi-modal
bone suppression, lung segmentation, and classification approach for accurate COVID-19 detection using chest radiographs", Intell. Sys. Appl., vol. 16, 2022. [http://dx.doi.org/10.1016/j.iswa.2022.200148]
[21] Available at: https://towardsdatascience.com/applied-deep-learning-part-4-convolutional-neuralnet-
works-584bc134c1e2 [http://dx.doi.org/]
[22] Y. Dalin, et al. "A synchronized hybrid brain-computer interface system for simultaneous detection
and classification of fusion EEG signals", Complexity, vol. 2020, 2020.
[23] M. Rashid, B.S. Bari, N. Sulaiman, et al. "A hybrid environment control system combining EMG and
SSVEP signal based on brain-computer interface technology", SN Appl. Sci., vol. 3, no. 782, 2021. [http://dx.doi.org/10.1007/s42452-021-04762-7]
[24] M.F. Mridha, S.C. Das, M.M. Kabir, A.A. Lima, M.R. Islam, Y. Watanobe, "Brain-computer
interface: Advancement and challenges", Sensors (Basel), vol. 21, no. 17, pp. 5746, 2021. [http://dx.doi.org/10.3390/s21175746] [PMID: 34502636] [PMCID: PMC8433803]
[25] P. Yuan, X. Gao, B. Allison, Y. Wang, G. Bin, and S. Gao, "A study of the existing problems of
estimating the information transfer rate in online brain–computer interfaces", J. Neural. Eng., vol 10, no. 2, 2013. [http://dx.doi.org/10.1088/1741-2560/10/2/026014]
146 Disease Prediction using Machine Learning, 2024, 146-158
https://t.me/med1917
CHAPTER 9
Mining Standardized EHR Data: Exploration, Issues, and Solution
Shivani Batra
1
KIET Group of Institutions, Delhi-NCR, Ghaziabad, Uttar Pradesh, India
2
GD Goenka University, Gurugram, India
Abstract: Medical database is among the most crucial databases in terms of their applicability to human life. Many researchers are in search of knowledge that is abstracted within the data. Data mining is popular in today's world as it gives access to knowledge that is otherwise unavailable. The concealed knowledge which is offered as a result of it can help the individual to make better decisions. Data mining tools in health have great potential. These solutions may be divided into four categories: therapeutic efficacy assessment, patient care, customer service, and embezzlement monitoring. The authors discovered that giving decision assistance in the medical sector with an emphasis on electronic health records (EHRs) can save lives. Though offering decision assistance in EHRs using data mining is valuable, it needs consistency. As a result, the authors intend to use data mining methods on standardised EHRs to create a decision support system. This paper presents the state-of-the-art data mining approaches and their application in the healthcare sector. It provides an integrated summary and a comparison detail of the existing literature. This chapter surveys several issues that need to be handled before employing data mining on EHRs and further proposes a solution for dealing with these problems. The problems such as multiple origins, multiple formats, missing data, distinguished users, data granularity, flexibility, and sparseness need immediate attention from researchers. Resolving these problems is important to build an efficient standardized EHRs database.
1,*
, Vinay Kumar1, Neha Kohli2 and Vaishali Arya
2
Keywords: Data Mining, EAV Model, Health Data, Standardized EHRs,
Sparseness.
INTRODUCTION
Data mining (DM) is the method of choosing, examining, and modelling massive datasets in order to uncover interesting patterns or correlations that offer the expert a valid and meaningful conclusion [1 - 3].
*
Corresponding author Shivani Batra: KIET Group of Institutions, Delhi-NCR, Ghaziabad, Uttar Pradesh, India;
E-mail: ms.shivani.batra@gmail.com
Geeta Rani, Vijaypal Singh Dhaka & Pradeep Kumar Tiwari (Eds.)
All rights reserved-© 2024 Bentham Science Publishers
Mining Standardized EHR Data Disease Prediction using Machine Learning 147
https://t.me/med1917
DM is the process of extracting useful information from massive datasets maintained in a variety of places, including data stores, spreadsheets, and cloud platforms. This information is useful in a variety of fields, including corporate strategy, biomedical investigation, and policy decisions. DM is a crucial aspect of information extraction since it examines massive amounts of facts and provides us with previously unrecognized, concealed, and usable information. DM has been successfully employed in a variety of industries, including weather forecasting, healthcare, transit, education, finance, and governance. When employed in a given business, DM offers several benefits such as prediction, clinical diagnosis, and categorization. Aside from such benefits, DM has its own drawbacks, such as the potential for confidentiality breaches. For example, if the miner has access to all of the data's details, he may exploit some of the data's secret information.
Although DM may be used in a variety of fields, its application in the health industry has the potential to assist humanity. Data includes a lot of information that might be concealed. Many academics are working to uncover these previously unknown sections of medical information [4 - 6]. Any platform's most precious asset is data. In the health sector, a lack of vital information can be dangerous to health. The discipline of DM is rapidly expanding, but applying it to medical records is much more difficult than applying it to other data due to the existence of several unique traits.
The remainder of the paper is laid out as follows. Section 2 delves into the complexities of the medical industry, with a focus on EHRs. Section 3 focuses on how DM can be used in EHRs. Section 4 looks at the issues of applying DM to EHRs, and Section 5 offers a remedy. Sections 6 and 7 respectively show related work and the conclusion.
COMPLEXITY IN EHRS
Paper-based patient data have a number of drawbacks, including restoration, productivity, and integrity. EHRs address all of the drawbacks of paper-based health data and make them available to users at a single click.
An Integrated Care EHR [7] is defined as: “a repository of information regarding the health of a subject of care in computer processable form, stored and transmitted securely, and accessible by multiple authorized users. It has a commonly agreed logical information model which is independent of EHR systems. Its primary purpose is the support of continuing, efficient and quality integrated healthcare and it contains information which is retrospective, concurrent and prospective”.
148 Disease Prediction using Machine Learning Batra et al.
https://t.me/med1917
The importance of standardisation for health clients cannot be overstated. Examining the available standards in the health industry is crucial. Various standard bodies are attempting to make interoperable EHRs a reality. Prominent organizations include Health Level 7 (HL7) [8], European Committee of Standardization Technical Committee 251 (CEN TC251) [9], International Standard Organization (ISO) [10, 11], and openEHR [12].
Medical advances are a data-intensive discipline that accumulates massive amounts of complicated and varied information. The health domain is very vast. Considering the complexity of the EHR domain, we identified that it consists of EHR_patients, EHR_concepts, EHR_relationships, terminology concepts and terminology relationships. There are a lot of EHR concepts in a clinical concept. Medical terminologies are a useful tool for standardising taxonomy and incorporating interpretation into EHRs [13]. Professionals frequently use jargon and a plethora of identifiers to define diagnoses. Doctors specialise in a variety of fields. For medical terms, there are several standardised labelling standards. Each specifies its own collection of terminology as well as the connections between them. The SNOMED-CT (Systemized Nomenclature of Medicine-Clinical Terms) coding system, for example, covers about 300,000 medical notions and 7 million connections. Amongst the most significant application fields for DM is healthcare [14]. Apart from having a complex structure, medical data is very sensitive in terms of ethical and legal issues. Besides all of this, a key issue is the occurrence of missing values in the data, which might cause a researcher to receive erroneous findings.
IMPLEMENTING DM ON EHRS
In today's technological age, the volume of health data is growing. Manually processing such a big volume of data is difficult and might result in numerous mistakes. Furthermore, it is likely that a regular person will not be able to acquire hidden knowledge. DM is critical in giving solutions to all of these issues. Many research works are going on to mine electronic health records such as finding the effects of drugs, detecting health insurance fraud, and finding patterns in symptoms to predict future diseases of patients.
All research works which are going on or have already had done provide a good accuracy measure but lack standardization in terms of the local schema used by the researchers. Standardization is necessary since it gives individuals perfect control over any asset, regardless of the criterion. Because of the wide range of storage and retrieval systems, standardisation is critical in the health sector. To construct the data structure, each company has its own set of regulations. This study's major aim is to discover gaps in the use of DM techniques to standardised
Mining Standardized EHR Data Disease Prediction using Machine Learning 149
https://t.me/med1917
EHRs, as well as potential solutions to these gaps. Our main focus is standardization because if it is achieved, the benefits of the standardized EHRs will be provided worldwide. DM can be used to predict the future or to find the interrelationship among observed or unobserved variables. As a result, DM algorithms are divided into three groups based on the roles:
1. Classification: The technique of anticipating the category of a new item from a collection of preset classes is known as classification [15]. The data is separated into a testing set, and a training set. The entire framework is designed by using a training set, and its correctness is determined by verifying the categories to which the test set of data samples are classified.
2. Clustering: Clustering [15] is a strategy for unlabeled data in which the categories are undefined. It operates by estimating the distance and similarities between the two items. Items will correspond to the same category if the resemblance metric is in the predefined threshold; alternatively, they would not. Clustering is a term used to describe a procedure in which items belonging to the same category are grouped with each other to establish a cluster.
3. Association: The goal of association [15] is to figure out how distinct data properties are associated with one another. One of the most well-known association techniques is the Apriori method [15]. Finding fascinating trends necessitates two parameters: support and confidence. The effectiveness of association methods is the criterion for assessment, not correctness since they locate each conceivable relationship among the characteristics of the data.
The type of DM work to be undertaken is determined by the result desired. During the present investigation of implementing DM on EHRs databases, the author examined several DM methodologies based on various criteria necessary. The comparison of the DM techniques categories is shown in Table 1.
Table 1. Comparison of various data mining algorithm types.
Characteristic Classification Clustering Association
Modelling With Supervision Without Supervision Without Supervision
Forecasting Predictive Descriptive Descriptive
Output Predicting Class Output
Performance Measure Accuracy Accuracy Efficiency
Efficiency Dependency Training Data Precision Adopted Threshold Value Support and Confidence
Segregating Unlabeled
Classes
Recognizing Features’
Correlation
150 Disease Prediction using Machine Learning Batra et al.
https://t.me/med1917
Next, the information gathered by using a DM algorithm may be important for delivering effective decision assistance. The procedure is demonstrated in Fig. (1). Decisions making in healthcare is directly related to people's lives. Thus, it is mandatory to minimize human errors and biases in decision-making. Humans are capable of making mistakes, but machines are more dependable. It will be highly advantageous for everyone to build a system that can assist in decision-making. The details of the problems identified and the proposed solution as demonstrated in Fig. (1) are presented in Section 4 and 5 respectively by the authors.
Fig. (1). Application of data mining to EHRs.
All parties such as physicians, corporate leaders, and sufferers engaged in the medical field may tremendously benefit from DM applications [1]. DM techniques used on EHRs will aid in the retrieval of previously unknown information. However, due to the intricacy of health data and the existence of distinctive attributes in health data, mining health records is more difficult than mining other sets of data. Data mining can provide decision support in many areas but nothing else can be more useful if we can take the right decision about human life at the right time. Mining health records is undoubtedly significant, although it's not as simple as mining any other sort of data due to the existence of certain distinctive traits described by Cios and Moore [16] as well as difficulties outlined by us in Section 4. Medical data is very special in terms of its contents [13].
Mining Standardized EHR Data Disease Prediction using Machine Learning 151
https://t.me/med1917
Due to the sheer peculiarity of health data, Cios and Moore [14] offer a full overview of all difficulties (heterogeneity, regulatory and political issues, analytical bounds, and protected privilege) that a DM researcher must be aware of before beginning his investigation. These are described in detail below:
1. Heterogeneity: Medical data might include scans, unstructured and semi­structured data, visual information, audible data, infographics, and figures. Due to this nature of data, high-capacity storage is required. Moreover new mining tools are also required to analyze this type of data so that it can be interpreted easily by the user.
2. Regulatory and Political Issues: Dealing with people to gather medical data necessitates dealing with a variety of difficulties, including information privacy, the threat of litigation, safety and confidentiality of data, anticipated benefits, and administration.
3. Analytical Boundaries: Statistics deal with different types of numbers whereas data mining deals with different types of data. Applying statistics operation requires the medical data to be numeric.
4. Protected Privilege: Medicine is an important part of everyone’s life. Special attention has to be given while dealing with mining medical data since the result of carelessness can be very hazardous. All ethics should be maintained by the person who is accessing the data for any legal purpose and no access rights should be given for illegal purposes.
CHALLENGES IN MINING STANDARDIZED EHRS
Based on a thorough examination of DM in the health sector, the present study looked into the following issues when implementing DM on EHR data.
1. Multiple Origins: Every organization holds information on an individual basis. When implemented in the associative data of all organizations, DM techniques would be beneficial. As a result, data ought to be present in a single location, however, data protection would be a major concern in such a situation.
2. Multiple Formats: Each institution has its own data structures for storing patient information. As a result, data must be collected using a single standard format so that DM tools may be used to uncover information. Also, there is a need to develop automatic tools for converting this heterogeneous medical data to a uniform format.
3. Missing Data: There are times when data has missing values or is contaminated by noise. When DM technologies are used on inaccurate data, the
152 Disease Prediction using Machine Learning Batra et al.
https://t.me/med1917
outcomes will be wrong as well. For example, a particular hospital is capturing information by blood tests, X-rays, as well as CT scans. But, another hospital does not provide a facility for CT scan. This leads to a null value in place of a CT scan for some of the patients. Such null values must be handled before decision­making.
4. Distinguished Users: Users are divided into three categories: professional, semi-skilled, and beginner. A professional user such as a physician frequently requests all the information and reasoning to interpret things. A semi-skilled operator such as a nurse is primarily interested in the service's entire information. The average beginner such as a sufferer is simply interested in the outcomes. In this way, the desired outcome is different for different groups of people.
5. Data Granularity: The amount of granularity is determined by the user's requirements. Varying levels of granularity may be required by professional, semi-skilled, or beginner-level clients. Whenever a client checks his vital sign record, he may need to determine if his heartbeat is regular, excessive, or lower, but if a professional checks his heart rate records, he needs to examine the details rate at which his heart is beating.
6. Flexibility: The healthcare era is rapidly approaching. As a result, modern treatment ideas or diagnosis terminology are established, or established health conceptions are coupled with novel diagnostic terminology. To address this issue, we need flexible data structures that can support these changes without disrupting the entire system.
7. Sparseness: Several times, health data is gathered using surveys that ask for personal details that the individual may not wish to share. In this case, the feature containing the specific value has a null value. Thus, this leads to storage wastage.
Several agencies, including HL7, openEHR, and ISO13606, are constantly working to fix these challenges, but we discovered that DM on standardized EHRs has received very little investigation.
SOLUTION FOR MINING STANDARDIZED EHRS DATABASE
Traditionally databases are stored under a relational model which maintains one column per attribute but the relational model is not able to handle the problems defined in the previous section; hence a new data model is required for storing EHR such that a space is utilized in a very efficient way. The EAV model is capable enough to store EHR data in a very space-efficient way. The EAV model [17] facilitates data storage in a column-oriented storage with triplets in every tuple corresponding to Entity, Attribute, and Value. The 'Entity' component