Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5203_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
30.08.2026
Размер:
51 Мб
Скачать
3 The Basic Principles andPrecautions ofDrug Therapy
351
3.5 Role ofArticial Intelligence Technology inDrug
Therapy ofEpilepsy

3.5.1 Introduction

Epilepsy is a prevalent chronic neurological disorder worldwide, affecting people of all ages and ethnic backgrounds. The lifetime risk of epilepsy is approximately 760 per 100,000 individuals. More than 50 million people live with epilepsy, accounting for 13 million disability-adjusted life years (DALYs) globally [96]. The disease burden is signicant at both the individual and societal levels. Recurrent seizures can lead to physical and psychosocial disabilities, while comorbidities such as learning disabilities and psychiatric disorders complicate epilepsy management [97]. Premature mortality can be attributed to status epilepticus and Sudden Unexpected Death in Epilepsy (SUDEP). Epilepsy is associated with social stigma, unemployment, and marital problems across different cultural backgrounds [98100]. Substantial healthcare resources are allocated to epilepsy care, with an average annual cost of over US$ 1000 per person in high-income countries and an estimated global healthcare cost of US$ 119 billion per year [101].
Currently, pharmacological therapies remain the primary treatment modality for epilepsy. ASMs control seizures via various putative mechanisms of action [102], which include sodium channel blockers (e.g., phenytoin and carbamazepine), GABA agonists (e.g., benzodiazepines and barbiturates), calcium channel blockers (e.g., ethosuximide), as well as agents targeting molecular substrates such as synap­tic vesicle protein 2A (SV2A) receptors (e.g., levetiracetam and brivaracetam) and AMPA receptors (e.g., perampanel). The number of ASMs available has signi­cantly increased in the past two decades, with more than 20in the market currently [103]. Newer ASMs generally exhibit improved adverse event proles and fewer drug interactions. However, the effectiveness of these new ASMs in seizure control does not surpass their older counterparts. Recent studies have shown that the pro­portion of people with drug-resistant epilepsy (DRE) remains steady at around 30% of all people living with epilepsy, despite the expanded pharmacological options [104]. This implies the ASMs are not effective in reversing the course of the disease.
Currently, the pharmacological management of epilepsy is largely empirical. Once the diagnosis of epilepsy is conrmed, a patient is typically prescribed an ASM.The choice of drug depends on the clinician’s discretion, considering factors such as epilepsy classication, patient demographics (e.g., childbearing potential), comorbidities, and potential drug interactions. The decision-making process may vary depending on the individual clinician’s experience. If the initial ASM fails to control seizures, a second ASM may be added to, or substitute for the rst. According to the criteria of International League Against Epilepsy (ILAE), a diagnosis of DRE
352
Q. Wang et al.
can be made if a patient do not achieve sustained (12months or more) seizure free­dom after adequate trials of two appropriately ASMs [105]. Currently, there are no reliable biomarkers available to predict DRE at the time of epilepsy diagnosis. Patients often undergo a trial-and-error process until the right ASM may be selected, which can take months or even years to achieve the goal of seizure freedom or con­rm the diagnosis of DRE [106].
On the other hand, the concept of Personalized Medicine or Precision Medicine has gained signicant attention in recent years [107]. In the conventional treatment paradigm, various guidelines on selection of ASMs are based on clinical trials con­ducted with stringently selected patient groups. These studies draw conclusions by using large patient cohorts to ensure sufcient statistical power and minimize con­founding factors. However, this evidence-based practice has been criticized as a “one-size-ts-all” approach that neglects the specic characteristics of individuals. In contrast, Personalized Medicine aims to customize therapy on the individual level. By gathering personal proles, including demographics, genomics, radio­graphic, and physiological features, the most suitable therapy can be tailored to each patient. In the eld of epilepsy, Personalized Medicine promises to be an achievable goal with advances in technology [108, 109]. For example, the development of pharmacogenomic panels may allow the prediction of responsiveness and potential adverse effects of ASMs in an individual patient [110]. Neuroimaging processing can detect subtle structural substrates of epileptogenicity, thereby enhancing the success rate of surgical treatment [111]. Automated analysis of EEG signals can be used to detect seizures and predict the prognosis of epilepsy [112]. All these tech­niques involve a huge amount of data that overwhelms manual analytical capacity.
3.5.2 Articial Intelligence andMachine Learning
Articial intelligence (AI) refers to the capability of a machine to imitate intelligent human thinking and learning. AI has a wide application in medicine. It has the advantage of precision, efciency, and perpetuity. Hence, investment in AI in health care industry may save cost by reducing medical errors, streamline workow and reducing manpower [113]. Novel applications emerge continuously; some notable examples are given below.
Robotics powered by AI has revolutionized the surgical eld. Robotic-assisted surgery refers to a specialized robotic system to assist or replace the surgeon in certain procedures to enhance precision and minimize complication risk. These robotic systems even allow telesurgery, in which surgeons can operate on patients remotely [114, 115].
In the journey of management, a patient may have multiple consultations with various specialists or receive treatment from different institutes. Reconstruction of a patient’s natural history from electronic medical records from different sources is now possible with Natural Language Processing (NLP) [116]. NLP is a subeld of AI that focuses on how computers could understand and generate human language.
3 The Basic Principles andPrecautions ofDrug Therapy
353
An automatic medical image reporting system involves both computer vision and NLP.Computer vision algorithms process and analyze medical images. These algorithms employ various techniques such as image segmentation, object detec­tion, and image classication to identify anatomical structures and abnormalities in the images. The ndings of the input image are expressed as a structured and coher­ent textual report with the help of NLP.
The concept of applying AI in decision-making or problem-solving processes can be traced back to the mid-twentieth century. Clinical Decision Support Systems (CDSS) are software systems designed to facilitate daily clinical decisions [117]. With the incorporation of AI, the capabilities of CDSS can be enhanced. An exem­plary example is the MYCIN system developed in the 1970s, which aimed to facili­tate the management of bacterial infections [118]. It consisted of three major components: the knowledge base, inference engine, and user interface. The knowl­edge base contained facts provided by human experts, while the user interface facili­tated communication between the computer software and the user. The data input into the inference engine by the user was then interpreted using a preset logic ow. The MYCIN system employed a backward chaining technique, which is a top-to- bottom approach that tracks backward to prove the truth of facts. Other examples of CDSS include an alert warning system in the computerized prescription program, where a warning signal appears if a patient has a potential drug allergy to the prescribed medi­cation. Over the past two decades, there have been spectacular advances in computer technology, including the emergence of high-performance processors, massive stor­age, the internet, and cloud computing, nurturing the next generation of AI.
Machine Learning (ML) is a branch of AI that may help to solve many practical problems. ML can be broadly divided into supervised learning and unsupervised learning [119]. In supervised learning, a human-labeled training dataset is used to train the ML algorithm [120]. The training dataset consists of labeled input features matched with corresponding categorized outputs. Examples of supervised learning models include regression, random forest (RF), support vector machine (SVM), k-nearest neighbor (kNN), and Gradient Boosting Tree (GBT) [121123]. Supervised learning has wide-ranging applications, such as classication, NLP, and image rec­ognition. Classication predicts categorical outputs; NLP is used for analyzing text data (e.g., automated translators); and image recognition is used for recognizing images (e.g., face recognition systems). Unsupervised learning is usually used in hidden pattern recognition and association. Examples of unsupervised learning techniques include K-Means Clustering, Principal Component Analysis (PCA), and Autoencoders. Deep learning is a more advanced type of ML.It does not require manual input in feature selection; rather, the model could extract features from raw data. The algorithm structure is in the form of a multilayered articial neural net­work. The inputs go through articial neuron layers in a non-linear way and map into outputs. The deep learning model can also improve its performance as the size of the training data grows.
To develop an ML model, several key steps are involved, including data prepro­cessing, algorithm selection, model training, and performance evaluation. In the data preprocessing stage, data is collected from various sources, and any missing or
354
Q. Wang et al.
irrelevant data is removed. Sometimes, features or variables in a dataset are trans­formed, such as conversion from a continuous scale to a categorical value, for the convenience of model tting.
Next, suitable algorithms are selected based on the specic problem to be solved. In some cases, multiple models are developed and compared to select the one with the best performance.
Once the algorithm is selected, the dataset is divided into a training set and a vali­dation set. The training set is used to teach the supervised learning model to make meaningful predictions. After training is completed, the model will then be vali­dated to assess its performance and generalizability. Various validation techniques can be employed, such as k-fold cross-validation and leave-one-out cross-validation (LOOCV). In k-fold cross-validation, the whole dataset is divided into equal por­tions or folds of number k. The model is then trained on k-1 folds, and one-fold is left out for validation. The model is trained for k times, so that each of the k folds is used for validation [124]. In LOOCV, each observation is selected as a validation set while the rest of the observations (N-1) are considered as the training set. The choice of validation approach depends on the size of the dataset and the complexity of the computation involved.
Several metrics are commonly used to assess the performance of ML models. For example, in a classication task, performance metrics usually include accuracy, precision, recall, F-score, positive predictive value (PPV), negative predictive value (NPV), and area under the receiver operating characteristic curve (AUC). The AUC varies from zero to one, where one signies a perfect classier and 0.5 suggests a random classier that is not better than chance.
It is valuable to understand the relative contribution of each feature in predicting within an ML model. This provides insights into the relationships between features and outcomes and helps improve the model. Established methods, such as model coefcient and permutation importance, can be used to assess the relative impor­tance of different features in a predictive model. Model coefcient reects the weight of a feature in a linear model, with the higher coefcient value indicating greater importance. Permutation importance is a model-agnostic approach that mea­sures the impact of shufing the value of a feature on the model’s performance. By evaluating the drop in performance after shufing each feature, one can determine its relative importance.
ML may be applied to solve problems in the development and clinical use of ASMs. In this chapter, we will focus on two main clinical applications: ASM response and DRE prediction.
3.5.3 Prediction ofASM Response
Devinsky etal. developed an RF model to predict the ASM with the highest likeli­hood of seizure control for individual patients utilizing claims data [125]. About 6years of medical claims data were retrieved from a nationwide comprehensive
3 The Basic Principles andPrecautions ofDrug Therapy
355
claim database in the US.Adults with epilepsy were identied based on International Classication of Diseases, Ninth Revision (ICD-9) coding and pharmacy claims for ASMs. The researchers recruited cases with any one of the three conditions: (1) initiation of rst ever ASM; (2) addition of ASM to an existing regimen; (3) substi­tution of an ASM by another ASM.The date of these major changes in regimen was dened as the index date. Patient’s features and the treatment outcomes were extracted from 12months before and after the index date, respectively. ASM substi­tution or addition of ASM to an existing regimen or complete withdrawal of the ASM was presumed to be treatment failure due to either inefcacy or intolerability.
A huge number of features, about 5000, covering demographics, comorbidities, therapeutic mechanism of ASMs, were available in each case. To avoid overtting, a selection process with statistical criteria was implemented to lter out irrelevant features. The nal training dataset included 34,990 patients, and the validation data­set comprised 8292 patients.
The trained RF model selected the ASM with the highest probability of a suc­cessful outcome as output. The machine-selected ASM was compared to the ASM chosen by physicians. The RF model achieved an AUC of 0.72. The ASM selected by the model matched with the physician’s choice in only 13% of the cases. In this group, the treatment outcomes were found to be better with a lower chance of alter­ing the ASM regimen. Other secondary outcomes, such as health service utilization, also favored ASM selection by the ML model.
This study used a claim database as the data source for developing an ML model. This kind of administrative database has certain advantages. It contains a large num­ber of cases and is readily available. If the database has a broad coverage, the cohort obtained could minimize selection bias.
However, the claim database also has its limitations. Clinical details could not be directly retrieved. Clinical conditions have to be presumed by surrogate markers. In Devinsky’s study, any change in ASM was supposed to be treatment failure, but it could not distinguish whether it was due to lack of efcacy or medication intoler­ance. Moreover, the database could not differentiate whether the ASMs are pre­scribed for seizure prophylaxis or other purposes, for instance, valproate could be used as a mood stabilizer and prophylaxis of migraine; gabapentin and pregabalin are commonly used to treat neuropathic pain. The diagnosis dened by coding may be subject to the pitfalls of either over-coding or under-coding. Commonly used clinical outcomes, such as seizure freedom or 50% seizure reduction, could not be obtained directly from the claim database.
Ouyang et al. designed an SVM model to identify biomarkers to predict the therapeutic outcomes of ASMs [126]. Quantitative EEG (QEEG) analysis of patients’ background EEG was utilized as a biomarker to assess therapeutic ef­cacy. A total of 20 children living with epilepsy, 11 of whom responded to ASMs and 9 had DRE, were enrolled. EEG recordings were obtained from each patient at different time points before and after initiation or change of the ASM regimen. The EEG recordings were then analyzed by the QEEG technique. Discriminative fea­tures of the EEG were selected and fed into the SVM model to classify the treatment
356
Q. Wang et al.
outcomes. A ten-fold cross-validation approach was adopted for internal validation. Finally, six EEG feature descriptors were found crucial in differentiating treatment­effective (seizure reduction >50%) and ineffective groups. The performance metric of the model yielded an accuracy of 83%.
This study integrated clinical and electrophysiological features to enhance the predictive capabilities of the ML model. However, there are drawbacks in the study that undermine its generalizability. First, the number of patients was small. Also, the recruited patients had different seizure subtypes, including both focal and general­ized seizures. The EEG features of these seizure subtypes are known to have great discrepancies. On the other hand, the ASMs used by the patients could also affect the EEG tracing. In this study, the ineffective group more commonly had patients using ASM combinations, while the effective group more commonly used mono­therapy. The effects of multiple ASMs or their drug interactions may cause bias in QEEG analysis when compared to a single ASM.Further study with a larger and more homogeneous patient population is needed to verify the results.
Zhang etal. attempted to estimate the therapeutic outcomes of specic ASMs by using ML [127]. An SVM classier was developed to predict seizure freedom in epilepsy patients receiving levetiracetam based on both clinical and QEEG features. The study recruited 46 patients receiving levetiracetam from a single center in China. Twenty-two patients achieved seizure freedom while 24 did not. Eleven clin­ical features were selected as input, including (1) age; (2) duration of epilepsy; (3) family history of epilepsy; (4) seizure subtype (generalized, focal, or unknown onset); (5) seizure frequency in the 12months before levetiracetam; 6) psychiatric comorbidities; (7) seizure circadian rhythm; (8) temporal lobe epilepsy; (9) time between levetiracetam initiation to the last seizure before levetiracetam; (10) inter­ictal spikes in EEG; (11) positive MRI ndings. For the EEG features, Sample Entropy of alpha, beta, theta, and delta frequencies was used to describe the data [128]. Sample entropy is a measurement used in signal processing and time series analysis to quantify the regularity or predictability of a signal. Four frequency bands of alpha, beta, theta, and delta were reconstructed and then used to calculate Sample Entropy for each patient in every channel.
To prevent overtting, only strictly selective bands were extracted and incorpo­rated into the input dataset. Consequently, four bands from four different channels were selected: beta bands from EEG lead Fp2, alpha bands from F4, theta bands from C3, and beta bands from F8. The 11 clinical and 4 QEEG features (Sample Entropy of alpha, beta, theta, and delta) of all patients formed the dataset and were then used to train and validate the SVM model. The algorithm was validated by ve­fold cross-validation, hold-out validation, and jack-knife validation. Mean Impact Value (MIV) then assessed to determine the contribution of each feature to the model [129]. The results showed that the beta band from Fp2 played a crucial role in the classier model. Clinical features of temporal lobe epilepsy, family history, circadian rhythm of seizure, MRI ndings, comorbidity, time between levetiracetam initiation and last seizure, and duration of epilepsy had descending importance. As for QEEG features, beta band from Fp2 has a signicant contribution. Interestingly,
3 The Basic Principles andPrecautions ofDrug Therapy
357
age, seizure type, interictal spikes, and seizure frequency before levetiracetam did not have a notable impact on the model.
The combined model of both clinical and EEG features yielded an AUC of
0.95 in the training dataset and an AUC of 0.96 in the validation dataset. The researchers tried to compare the performance of the SVM model with input of clini­cal features only, EEG features only, and combination of clinical and EEG features. The model with combined input demonstrated the best results.
The superior performance of the combined input integrating clinical and EEG inputs made clinical sense, as multiple modalities probably enhanced the accuracy of the classier’s prediction. However, the small number of recruited patients is a main drawback of the study. The specic bands of the four EEG channels showed a signicant difference between seizure-free and non-seizure-free groups, but the authors did not postulate the electrophysiological mechanisms behind their nd­ings. The high AUC values in both training and validation may imply the model has high generalizability. It is worthwhile to further evaluate the model in external vali­dation by using an independent cohort.
Yao etal. constructed ML models to predict the outcomes of ASM in new-onset epilepsy [130]. A total of 287 patients with newly diagnosed epilepsy were recruited from a single center. The cohort had at least three years of follow-up. Fourteen demographic and clinical features were selected as input features of the model, including: (1) sex; (2) living area; (3) occupation; (4) education level; (5) age at seizure onset; (6) risk factors of epilepsy (including brain trauma, stroke intracranial infection, perinatal hypoxia, brain tumor); (7) family history of epi­lepsy; (8) seizure types (focal, generalized, or unknown onset); (9) number of seizures before treatment; (10) history of febrile seizure; (11) duration of previous treatment; (12) presence of multiple seizure types (more than one seizure type including focal aware, focal impaired awareness, focal to bilateral tonic–clonic, generalized onset, and unknown onset seizures); (13) EEG ndings; (14) MRI ndings.
In their treatment protocol, monotherapy of ASM was selected by the attending epileptologist. One of seven ASMs were selected: carbamazepine, oxcarbazepine, valproate, lamotrigine, levetiracetam, topiramate, and gabapentin. Outcomes from the rst ASM were classied into three groups: early remission (less than six months after ASM commencement), late remission (more than six months after ASM com­mencement), and no remission (ongoing seizures after one year of ASM treatment). In the dataset, there were 207 remission cases, of which 140 were early remission and 67 were late remission, along with 80 cases of no remission.
The researchers compared the performance of ve supervised ML models: Decision Tree, RF, SVM, Extreme Gradient Boosting (XGBoost), and Logistic Regression. Two experiments were conducted. In the rst experiment, all 287 patients’ data were used as a training dataset, aiming to predict remission versus no remission status. In the second experiment, data from 207 patients in the remission group were used as a training dataset, aiming to predict early versus late remission. A ve-fold cross-validation was carried out for internal validation.
358
Q. Wang et al.
Among the ve models, the XGBoost algorithm achieved the best metrics in predicting remission versus no remission, with an F1 score of 0.95 and an AUC of
0.98. The best predictor for remission versus no remission was the number of sei­zures before treatment. The XGBoost model also performed the best in predicting early remission versus later remission, with an F1 score of 0.84 and an AUC of 0.92. The presence of multiple seizure types had the most signicant impact in predicting early versus late remission.
This study compares different ML models instead of preselection of a single model, therefore enhancing transparency in model selection. Another notable point is that the input features cover investigational results in addition to other clinical features from medical history and demographics. Both the EEG and MRI ndings are incorporated as categorical features in the dataset. The XGBoost model achieved impressively high F1 and AUC values in predicting early remission versus late remission and remission versus no remission. One of the contributing factors was the application of grid search to optimize the hyperparameter set in hyperparameter tuning. Hyperparameters are parameters that are not learned during the training process, but are set before training and determine how the model learns and general­izes from the data. Hyperparameter optimization is the process of nding the best set of hyperparameters for an ML model. Of course, grid search also has its draw­backs, such as the requirement of additional computational resources and inability to deal with a large number of hyperparameters or extensive ranges of hyperparam­eter values.
Two related studies, conducted by Petrovski etal. and Shazadi etal., aimed to predict response to ASM in newly diagnosed epilepsy patients by exploiting phar­macogenomic markers [131, 132].
Petrovski etal. developed a multi-single nucleotide polymorphisms (SNPs) clas­sication model to predict the treatment outcome of the rst ASM in patients with newly diagnosed epilepsy [131]. A total of 115 patients with newly diagnosed epi­lepsy from two hospitals in Australia were enrolled. Treatment response was assessed at the end of the follow-up period of one year. Drug responders were dened as seizure freedom with the rst ASM, while non-responders were dened as those with recurrent seizures while on the initial ASM.
Genotyping for 4041 SNPs in 279 candidate genes was performed. These candi­date genes were chosen based on their roles in the pathogenesis of epilepsy or drug metabolism, as well as their high expression levels in the brain. The list of SNPs was further narrowed down, and ve sophistically selected SNPs were nally included in the ML model development. The algorithm of kNN was employed as a classier.
A ve-fold cross-validation approach was employed for internal validation. In addition, the model was further externally validated by two independent validation cohorts. The rst validation cohort comprised 63 newly diagnosed epilepsy patients from the same population, and the second cohort comprised 108 community-treated chronic epilepsy patients. The performance of the multi-SNP model was further compared with models based on individual SNAPs alone in each validation cohort.
The performance metrics showed promising results. In the cross-validation, the model achieved an accuracy of 84%. In the validation using the 63 newly diagnosed
3 The Basic Principles andPrecautions ofDrug Therapy
359
epilepsy patients, the model yielded a sensitivity of 91% and specicity of 53%. For the chronically treated epilepsy validation cohort, the multi-SNP classier model achieved a sensitivity of 81% and specicity of 50%. The multi-SNP models con­sistently outperformed the single-SNP models in both independent validation cohorts.
Shazadi etal. further investigated Petrovski’s multi-SNP model by externally validating it with larger independent cohorts [132]. They utilized two cohorts from the UK: the Glasgow cohort, consisting of 281 new-onset epilepsy patients, and a subset of patients from the SANAD study, which included 491 patients with genetic sampling [133]. The genotyping process was the same as in Petrovski’s study, focusing on the ve specic SNPs. The kNN algorithm retained identical input fea­tures as the original model. The UK cohorts were used to validate the multi-SNP model in three ways: (1) Petrovski’s training dataset served as the training dataset, and each of the two UK cohorts as the test datasets; (2) retraining the kNN model using the SANAD cohort as both training and test datasets. A random 70:30 train­test split was adopted. The parameters of kNN were identical to those employed in the original Petrovski’s model; (3) testing the performance of Petrovski’s kNN model using an LOOCV approach in the UK cohorts.
The results showed that the multi-SNAPs model could reasonably predict the treatment responses in carbamazepine or valproate in both Glasgow and SANAD cohorts (Glasgow: odds ratio [OR]=3.1, 95% condence interval [CI]=1.4–6.6, p=0.018; SANAD: OR=2.8, 95% CI: 1.3–6.1, p=0.048). However, the model could not reliably predict lamotrigine and other ASM treatment responses in both UK cohorts.
These two studies tried to predict the response to ASM using genetic approach. The hypothesis that multiple SNPs could have a better prediction than a single SNP is sensible, as response to ASM is probably a trait inuenced by multiple genetic variants. The possible reason for the failed external validation of the ML model on the two UK cohorts may be because of different prescription patterns between the Australian and UK cohorts. Carbamazepine and valproate were more commonly used for the new-onset epilepsy in the Australian cohorts, while lamotrigine or other ASMs were used more commonly in the UK cohorts [132]. This highlights the pref­erence for prescription in different institutions or countries could be a source of bias in the dataset. It would be optimal to collect the developmental dataset from a diverse range of institutions and countries.
De Jong etal. combined both clinical and genetic data to develop an ML to pre­dict the response to brivaracetam [134]. The data sources consisted of two patient cohorts from two randomized placebo-controlled clinical trials of adjunctive brivar­acetam. The rst cohort consisted of 235 patients, while the second cohort consisted of the other 47 patients. The data from the larger cohort was used for the develop­mental dataset for both training and internal validation, while the smaller group served as the retrospective validation dataset. The treatment response was dened as a>50% reduction in seizure frequency after 12weeks. Both the clinical and genetic features were incorporated into the ML models. Whole-genome sequencing (WGS) data of the patient cohorts were available. The huge amount of genetic data from
360
Q. Wang et al.
WGS was robustly narrowed down to a workable list of genetic features that related to mechanisms of epilepsy and drug action. Finally, four data modalities were employed, including: (1) mutational load scores; (2) polygenic risk score; (3) SV2A structural variance (the molecular target of brivaracetam); and (4) clinical data.
As some of the input features may be associated with placebo response, the researchers further validated the models with an extra placebo dataset. A cohort of 235 patients who were assigned to the placebo arm in the same clinical trial was used as the placebo dataset. However, genetic data was not available in this placebo cohort. The researchers assessed the placebo response in the model trained only on clinical data instead.
The researchers compared ve different ML models, including sparse multi­block partial least squares discriminant analysis, multimodal neural network, elastic net classier, GBT, and logistic regression-based stacking. The ML models under­went internal cross-validation and were further validated externally using the vali­dation dataset.
Among the tested models, the GBT achieved the best performance parameters, with an AUC of 0.76 (95% CI: 0.76±0.014) in the training dataset and 0.75 (95% CI: 0.75±0.15) in the validation dataset. Analysis of the placebo cohort ensured that the model trained on the clinical data modality was not associated with a pla­cebo effect.
The researchers demonstrated that the integration of genetic and clinical data modalities improved the model’s performance (AUC of 0.76) compared to using clinical data alone (AUC of 0.71). Genetic factors included the presence of struc­tural variants overlapping the SV2A gene and the mutational load in a gene set representing microtubule minus end binding. The only clinical factor that predicted brivaracetam responsiveness was prior levetiracetam use. Failure in levetiracetam use predicted a poor brivaracetam response.
This study has demonstrated that a meaningful model could be developed with a relatively small cohort. Such a drug response prediction model focusing on a single ASM could facilitate enrichment trials. Eligible subjects with a high probability of non-responsiveness could be screened out using an ML model before being recruited into a drug trial. This measure could effectively reduce the sample size by increas­ing the responder rate, making it easier to achieve statistically signicant differ­ences between the treatment and placebo groups in randomized controlled trials. It would also shorten the drug trial process and reduce the cost of drug development. However, the short duration of follow-up (12weeks) and small absolute difference in seizure frequency in the corresponding clinical trials make the reliability of the developmental dataset questionable. On the other hand, the contribution of genetic data to the nal model was surprisingly small, as reected by the small increase from AUC 0.71 of the model with clinical features alone to AUC 0.76 of the model with a combination of clinical plus genetic features [135]. The cost-effectiveness of incorporating genetic data into the ML model is challenged.
Hakeem etal. have developed and validated a deep learning model to predict the response of the rst ASM for new-onset epilepsy [136]. Five cohorts from Scotland,