Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_6027_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
29 Мб
Скачать
106 S. Viswanathan
findings (Picton et al., 2000; Keil et al., 2014 , 2022) (see Chap. 36: Writing Up Your EEG Researchand Chap.
In this chapter, we focus on the role of statist ics for EEG data analysis. Statistics provides a sophisticated toolkit to address a broad class of challenges (related to uncertainty). However, selecting the right tools to apply requires an understanding of the specic challenges to be solved. We provide an overview of the demands that EEG data present for statistical analysis. This is intended to help readers craft a statistical analysis strategy that is customized to their own research.
37: How to Evaluate an EEG Research Paper).

9.2 Why Is Statistics Needed in EEG Research?

When conducting EEG research, a rst (and often overlooked) issue to addres s is why statistical analysis is needed for a study. In this section, we discuss how uncertainties in an EEG experiment can make it necessary to apply statistical analysis.
The term statistics derives from the Italian phrase ragione di stato (i.e., science of
state), referring to the use of quantitative data for statecraf t (Ostasiewicz, 2014).
the In
its modern usage, statistics refers to a general mathematical approach (and related numerical procedures) to address the role of uncertainty in empirical studies. The relationship between uncertainty and statistics, as described nicely by the psychol­ogist S.S. Stevens (Stevens, noisy
, uncertain, and difcult, it is only natural that statistics should ourish. [] At the other extreme, if accurate measurement were achieved in every inquiry, many of the needs for statistics would vanish.
EEG measurement involves several uncertainties. A careful assessment of these
ainties can be valuable before we consider possible statistical solutions.
uncert
EEG measurements are used to detect and record uctuations in electrical poten-
tials
on the scalp that are generated by the brains neuroelectric activity (also see Chap. 2: What is EEG?and Chap. 20: EEG Source Analysis). An EEG record
ing thus provides a structured set of numerical values (i.e., data) that are associated with a persons brain activity. Due to this association, EEG recordings are a valuable dependent measure for experimental studies of brain function. This use of EEG presents a set of typical uncertainties, as described below.
Experimental EEG studies are often conducted to resolve unknowns about how
brain implements a particular capability or function (e.g., visual attention, motor
the control). Although experiments vary greatly in their specics, the rationale is often as follows:
1968): In those disciplines where measurement is
1. Experimental
manipulate the cognitive function of interest while EEG is being measured. These manipulation strategies can take various forms, such as instructed tasks, external stimulation, pharmacological interventions, and so on. It can also involve distinct groups of individuals with relevant characteristics, for example, clinical patient
conditions: Experimenters use controlled strategies to modulate/
9 Applying Statistics in Your EEG Research 107
groups or groups de ned by age. With these manipulations, the objective is to measure a persons EEG under distinct conditions or treatments (e.g., condition 1: resting with eyes open, condition 2: resting with eyes closed).
2. Neural activity: Each experimental condition is assumed to produce neural
vity related to the cognitive/mental function of interest.
acti
3. EEG recording: The condition-specic neural activity is, in turn, assumed to
produce
condition-specic scalp potentials that are recorded by EEG
measurement.
To summarize, the assumed relationships are: (1) each experimental condition
uences neural activity, and (2) this neural activity inuences the recorded EEG.
in Therefore, comparing the effects of the experimental conditions (the independent variables) on the EEG data (the dependent variables) can be informative about the neural processes of interest.
Detecting these effects quantitatively in the EEG data presents two critical
ainties: effect uncertainty and measurement uncertainty.
uncert
Effect Uncertainty Experiments are conducted to understand unknowns about brain
function. Therefore, the effect of the experimental conditions can be uncertain until the EEG data are analyzed. For example, in a particular study, presenting visual images of human faces (condition 1) might be expected to evoke EEG responses that are absent for images of houses (condition 2). The researchers might hypothesize that the perceptual processing of face stimuli engages a distinctive brain network that is inactive for house stimuli (Kanwisher,
e potentially incorrect assumptions about the brain. Therefore, the presence/
involv
2010). However, this hypothesis might
absence of these effects in the actual data can provide novel information to either conrm or revise prior scientic assumptions. This uncertaint y is compounded by a second source of uncertaintyEEG measurements can be noisy.
Measurement Uncertainty Apart from signals related to neural activity, the mea­sured
EEG includes noise and random variation from various other sources (see
Chap. 15: Getting Clean EEG Data: Artifacts and How to Prevent Them). The net
for the measured EEG can be summarized simplistically as: Measured
effect value = True value + Error. EEG measures a weak bioelectric signal in the microvolt range. Therefore, a major source of measurement error is electrical interference from the measuring environment (e.g., power line noise) and other bioelectric phenomena (e.g., muscle activity, eye movements). Several additional factors can produce inter-recording differences both within and between individuals, which can range from individual variability in scalp-to-electrode contact (e.g., related to hair thickness, skin properties) to neural activity itself (e.g., related to individual neuroanatomy/physiology) (Hommelsen et al.,
In summary,
experiments are conducted to resolve effect uncertainty. Evaluating
2022).
the presence/absence of an experimental effect (i.e., the true difference between experimental conditions) and its properties is part of scientic research. However, this objective is confounded by measurement uncertainty from various sources. Specically, in an EEG recording, the relative contribution of relevant brain activity
108 S. Viswanathan
(i.e., True value) and measurement variability (i.e., Error) can be uncertain, espe­cially when the True value is unknown (i.e., effect uncertainty). Therefore, addressing these overlapping uncertainties is critical.
Statistics provides a quantitative framework to navigate these uncertainties,
namely, by estimating the true effect despite the presence of measurement error.
Applying statistics does not eliminate the sources of the above uncertainties.
Therefor
e, practical steps to reduce or even eliminate measurement error should always be a high priority during EEG acquisition and data analysis. This priority has an inuence on when statistics is applied during EEG data analysis, as discussed next.

9.3 When Is Statistics Applied During EEG Data Analysis?

Figure 9.1 is a schematic of a typical organization of EEG data analysis and where statistics has a role. This is meant as a coarse overview as specics can differ from study to study. We describe the steps in detail below.

9.3.1 Raw EEG Data

In a typical EEG study, recordings are acquired from a sample of individuals using a particular experimental protocol. These individuals are selected from one or more well-dened populations (discussed in more detail below). The raw EEG recording
Fig. 9.1 Schematic of a typical EEG data analysis organization for hypothesis testing. Raw data are obtained from individual participants who are sampled from the population (left). They undergo an individual-level or rst-level analysis, which consists of a preprocessing step followed by signal denition steps (see text for details). The processed data from individual participants are then pooled together for group-level or second-level analysis. The group data are then used to evaluate statistical hypotheses by applying statistical tests. The outcomes are numerical measures evaluating these hypotheses
9 Applying Statistics in Your EEG Research 109
from each person usually consists of the following: (1) a structured set of continuous time-series of voltages from multiple electrode locations on the persons scalp (i.e., channel × time) often also with auxiliary data from peripheral physiology sensors (such as electrooculography (EOG) and electromyography (EMG)); (2) markers with information about the persons state and relevant events during the recording (e.g., related to instructions, stimuli, behaviors
The raw data cannot be interpreted as is and requires further analysis. This analysis usually has two sequential stages (see Fig. 9.1). The rst involves data analyses performed independently on individual datasets (referred to as individual­level analysis or rst-level analysis). This is followed by an analysis performed on data pooled from all participants (referred to as group-level or second-level analysis) with the application of statistics.
).

9.3.2 Individual-Level (First-Level) Analysis

The individual-level analysis has two sequential (but interrelated) steps.
Preprocessing The raw data are often acquired with parameter settings optimized
accurate measurement and digitization (e.g., choice of sampling rate, lters).
for Furthermore, the raw data can include extraneous (i.e., non-brain) information, also known as artifacts. Several of these artifacts have distinctive signatures (e.g., related to blinking, muscle contractions). Preprocessing refers to the various transformations applied to the raw data to convert it into a form suitable for analysis. This includes the removal of artifacts of known origin and the correction of biases. The reader can consult Chap. preprocessing and artifact handling.
17 (EEG Pre-processing and Artifact Handling) for more details on
Signal Denition
specic to each experimental condition. This might involve linking the clean data to condition-specic details, for example, using the recorded markers (see Chap. 7:
Basic Time Concepts in EEG Practice). These data are processed further to dene
the EEG signal of interest using one or more approaches, for example, event-related potential analysis (Chap. freque
ncy analysis (Chap. 18: Introduction to EEG Oscillations and Spectral Analysi ysis (Lachaux et al., can include the application of statistics, for example, to model intertrial variability (Pernet et al., Homme
different experimental conditions. Importantly, this activity has a lower contribution of measurement error as compared to the raw data.
s), source analysis (Chap. 20: EEG Source Analysis), connectivity anal-
lsen et al., 2022).
The outcome
The clean preprocessed data is then used to extract EEG activity
19: Event-Related Potentials), spectral and time-
1999) (Chap. 20 : EEG Source Analysis), and so on. This step
2011) or when applying machine learning (King & Dehaene, 2014;
of this stage is a summary of each individuals EEG activity in the
110 S. Viswanathan

9.3.3 Group-Level (Second-Level) Analysis

The individual data from the rst-level analysis are pooled together into groups for second-level or group analysis. This grouping is based on the individuals popula­tion membership (e.g., patient group, healthy control group) and the goal of the statistical analysis.
Unlike the rst-level analysis, where the focus is on the available data, the focus
of
the second-level analysis is on statistical hypotheses (i.e., predictions about the data). The main objective of this stage is to evaluate hypotheses on the group data by the application of statistical methods, discussed in more detail below.
9.3.3.1 Statistical Hypotheses
A statistical hypothesis is a quantitative prediction of the effect of selected indepen­dent variables on selected dependent variables in the experiment. This hypothesis is needed for statistical calculations. An example of a hypothesis might be: The mean
beta power at channel Cz is higher in the rest condition than in the motor imagery condition.
Statistical hypotheses are formulated as general statements relevant to the
tions of interest rather than specic to the sample of participants in the
popula study. Apart from being a numerical prediction, a statistical hypothesis in the EEG context involves the selection of a relevant dependent variable. The above example hypothesis is a prediction for one frequency band (i.e., beta band [13–33 Hz] from a large range of available oscillatory frequencies) at one particular channel (i.e., Cz of the severa l recorded channels). However, EEG data are typically multivariate, namely, involving more than one dependent variable. EEG is recorded continuously over time from multiple scalp locations. Therefore, even a 2-second segment of an EEG recording with 64-channels and 500 Hz sampling rate is associated with 64,000 voltage values (Channel (64) × Time (2 s × 500 Hz)).
Therefore, EEG analyses can benet from statistical hypotheses that include a
formulated rationale for variable selection (Picton et al.,
well-
Gaspe
lin, 2017). These hypotheses can be derived from broader scientic hypoth-
eses,
neuroanatomical considerations (see Chaps. 3: Basic Anatomy: Central Ner­vous Systemand 4: Basic Anatomy: Peripheral Nervous System) and details of the
experiment design. Without variable selection, the large number of dependent variables can dramatically increase effect uncertainty, which can require complex analyses to address (Pernet et al., 2015; Groppe et al., 2011; Maris & Oostenveld,
2007).
The maj iment design and analys is planning. Multiple alternative (or competing) hypotheses can be dened for the same combination of independent and dependent variables based on competing scientic considerations.
or hypotheses are dened before the data are acquired and guide exper-
2000; Luck &
9 Applying Statistics in Your EEG Research 111

9.3.4 Application of Statistical Inference

Statistical inference procedures are applied to evaluate each statistical hypothesis with the actual corresponding data. This is the step where the confounding effect of measurement uncertainty is numerically addressed.
A large variety of statistical procedures (or tests) could be used here as they are
specic to EEG, for example, t-test, Analysis of Variance (ANOVA), and
not general linear model (GLM). For details on specic statistical tests and how to choose a particular test, we recommend the reader consult a standard statistical textbook for further information. In general, before performing a test, the datas consistency with the tests assumptions is evaluated. Typical requirements are that the data are normally distributed, and that there are no extreme outliers. Major violations of a tests assumptions require a re-evaluation of how to proceed, for example, re-evaluating the data quality, or the application of a better-suited test.
This step can be practically implemented with any standard statistical package (for
example, R (https://www.r-project.org/), SPSS (https://www.ibm.com/products/
spss), JASP (https://jasp-stats.org/), NumPy (https://numpy.org/). Furthermore, this
step
requires organizing the data into a suitable format for the statistical tests, often
referred to as data wrangling(Wickham,
exity from simple tables for t-test (Table 9.1) to more involved formats
compl when
using, for instance, a multiple-factor ANOVA or a linear mixed model. The difculties of data wrangling should not be underestimated since errors during this step can directly impact the results. To reduce errors, many statistical programs are integrated with programming environments that enable data wrangling to be conve­niently and reproducibly scripted.
Running a statistical procedure produces multiple test-dependent numerical out-
such as p-values, effect sizes, and condence intervals.
puts,
2014). These formats can vary in
Table 9.1 Illustrative example of a numerical table prepared for a within-subject t-test to evaluate a statistical hypothesis about the mean difference in ERP magnitudes between two conditions (e.g., left, right) at channel Pz at time + 200 ms
Voltage condition = left
(Pz,
Participant 1 0.3285 -0.2143 0.5428 2 1.3332 0.0573 -1.3905 . . 30 -0.6386 0.9211 -1.5597
+200 ms)
Condition = right (Pz, +200 ms)
Difference: left - right
+200 ms)
(Pz,
112 S. Viswanathan

9.3.5 Interpretation and Inference

Despite being numerical input-to-output calculations, statistical procedures involve several complex assumptions about the data and uncertainty (Wasserstein & Lazar,
2016; Wasserstein et al., 2019) (see Fig. 9.2 and next section). Therefore, the
ical outputs of the statistical tests are used along with other considerations
numer (e.g., data plots, contextual information) to evaluate whether the hypotheses are supported by the data or suggest rejection. This reasoning is typically reported in the results section of a scientic article. Furthermore, this step can trigger the denition of further post hoc hypotheses and statistical testing.
In summary, the application of statistical calculations is performed near the end of
EEG
data analysis and relies on the previous steps. The rst-level analysis helps reduce the extent of measurement error in the analyzed data. While the rst-level analysis is strongly focused on data processing, at the group-level, the main focus is
Fig. 9.2 Schematic of statistical inference for a single variable. The hypothesis space (upper plane) is the space of possible statistical hypotheses where H1 and H2 are examples (black dots). For each hypothesis, a probability is assigned to every possible dataset that could be obtained under the effects of random chance if that hypothesis were true. Each dataset is represented by a sample statistic value, and the space of all possible sample statistic values is shown in the middle plane. The shading indicates the probability assigned to each sample statistic (dataset) by a hypothesis. The alpha threshold for each hypothesis is shown as a dotted line. The colored star indicates a sample statistic obtained from the measured sample (bottom plane). It is assigned a different probability by different hypotheses (probability < alpha for H1 but > alpha for H2). The measured sample (bottom plane) is a subset of the population. The sample mean ( sample size (N) are used to calculate the statistic for the sample
xÞ, sample standard deviation (s), and
9 Applying Statistics in Your EEG Research 113
on statistical hypotheses, which are crucial inputs to this stage. Furthermore, the outputs at the second-level are numerical evaluations of these hypotheses, rather than data transformations. Therefore, applying statistics in EEG analysis benets from well-specied hypotheses.
9.4 How Does Statistics Account for Measurement
Uncertainty?
The relative ease of running statistical procedures using modern software tools can underestimate the mathematical concepts and assumptions behind them (Gigerenzer,
2004, 2018; Nieuwenhuis et al., 2011; Wasserstein et al., 2019). A general under-
stand
ing of how statistical methods address EEG-specic uncertainty can be valu-
able to the application of statistics.
Figure 9.2 provides a simplied Venn diagram of a single dependent variable V in the EEG recording. As described earlier, the measured value of a variable (V
) and error (e.g., related to; measurement). For simplicity, V
(V
true
Both V
and the error are unknown.
true
) is assumed to be a mixture of its true value
meas
schematic of statistical evaluation
= V
meas
true
+ error.

9.4.1 Hypotheses (Upper Plane of Fig. 9.2)

A rst step in statistical analysis (e.g., with a t-test) is to make a guess about the possible value of mean V
, namely, a statistical hypothesis. A hypothesis (shown
true
as a black dot, upper plane in Fig. 9.2) is one of many possible alternatives. For examp
le, hypothesis H1 might be that mean V
be that mean V
< 0. By evaluating different hypotheses (i.e., guesses) against the
true
measured data, the objective is to arrive at better estimates of mean V
> 0, while an alternative H2 might
true
, which is
true
unknown and might never be exactly knowable.
The hypothesis is dened as mean V
rather than per individual. The mean is a
true
measure of the central tendency across a collection of individualsit can be evaluated on the entire population as well as on samples of different sizes N.

9.4.2 Population and Sample Data (Bottom Plane of Fig. 9.2)

The bottom plane shows the values of V data obtained from the rst-level analysis. Importantly, the value of N (i.e., the sample size) and the specic individuals included in the analyzed sample are decided well before the group-level analysis.
of N individuals in the sample. This is
meas
114 S. Viswanathan
A critical idea in statistical inference is that properties observed in a sample are informative about properties of the population. Without this assumption, statistical inference would be unable to reason beyond the sample. For this to be possible, the sample is assumed to be representative of the population, namely, sharing properties of the entire population. This can be illustrated by considering an urn lled with 1000 colored marbles (the population), where 70% are colored red and the remaining 30% are green. If the marbles are thoroughly mixed in the urn, then a randomly selected sample of N = 20 marbles is representative of the full population as it would be expected to have color proportions similar to the entire urn, even if somewhat higher or lower. However, if all the red marbles were concentrated at the top of the urn and all the green marbles were concentrated at the bottom of the urn, then a sample from the urn (even if selected randomly) would be biased, leading to erroneous conclusions about the population. Along similar lines, the selection of a representative sample is a critical requirement during experiment design.
Although the population is a mathematical assumption, it is the set of individuals that
we seek to meaningfully learn about in a research study. Therefore, dening a suitable population is critical for the value of the study. The population is often dened using a combination of inclusion and exclusion criteria, for example, age, demographics, and clinical prole.

9.4.3 Sample Statistic (Middle Plane of Fig. 9.2)

How is the hypothesis linked to the acquired data? The key idea is to use hypothesis H to generate a model of all possible ways in which random variation can inuence mean V
, if the hypothesis H were actually true.
meas
In Fig. 9.2, this is shown by the arrow from H to the middle plane. Each point in this plane is a sample statistic value, e.g., t-statistic. For a sample of size N, a sample can be described by a unique t-value (Eq. 9.1)
x - μ0
p
sð Þ= N
ð9:1 Þ
where
t =
x
is the sample mean, s indicates the sample standard deviation and μ0 is the statistical hypothesis. This plane represents the statistical values for all possible samples (of size N ), actually measured as well as hypothetically possible.
The model
for H assigns a probability value to every possible sample statistic value. This is shown by the gradations in gray color, where dark gray indicates a higher probability and lighter colors indicate lower probability. In parametric statis­tics, this model is derived mathematically, i.e., independent of the actual data. This probability is often denoted as P(d|H), namely, the probability of obtaining data d given that the hypothesis H is true.
9 Applying Statistics in Your EEG Research 115
Different hypotheses assign different probabilities to the sample statistic values. Therefore, the same sample statistic can be assigned a probability p according to hypothesis H1 and a different probability p
= P(d|H2) according to
H2
= P(d|H1)
H1
hypothesis H2. Hence, d might have a higher probability if H1 were true but a lower probability if H2 were true. However, this does not imply that H 2 is more probable than H1 since P(H1|d ) is not equal to P(d|H1). The reason is due to Bayes Theorem (van de Schoot et al., 2021), which we do not discuss further here.
The key step is to evaluate the actual sample data (colored star). Specically, if
is true, is the data consistent with this hypothesis? One approach has been to set a
H
threshold probability below which the data is deemed to be inconsistent with H (shown as a dotted line). This threshold is denoted as α (alpha). It is set by the researcher and is typically 0.05. So, an extreme p-value less than 0.05 is interpreted as having a low consistency with hypothesis H, if H is true .
With this sketch of the key concepts, we see that statistical inference can be a
rful aid to reasoning but also involves many interacting assumptions. Meeting
powe these assumptions involves decisions made before data acquisition (e.g., sample size, sample selection ). It can also be seen that this procedure cannot tell us what is truebut evaluates the consistency of data if a particular hypothesis were to be true. Therefore, how these procedures are applied and their interpretation require careful attention to the EEG context. Finally, these considerations highlight that a decision about a hypothesis with statistical inference (e.g., p < 0.05) is incomplete without the details of the evaluation (Curran-Everett & Benos,
2004, 2007; Lang & Altman,
2015).

9.5 Conclusion

In this chapter, we have provided you with an overview of the uncertainties that arise in EEG measurement that motivate the application of statistics. A careful diagnosis of these uncertainties in your study can help you select a suitable statistical approach. Before applying statistical analyses, you should consider minimizing the contribu­tion of measurement error during data acquisition and other analysis stages. Finally, a major recommendation is to formulate well-specied statistical hypotheses. Well­specied hypotheses with a rationale for variable selection can benet statistical analysis and enable a clear interpretation of the results.

References

Curran-Everett, D., & Benos, D. J. (2004). Guidelines for reporting statistics in journals published
by the American Physiological Society. American Journal of Physiology. Regulatory, Integra-
tive and Comparative Physiology, 287, R247–R249. Curran-Everett, D.,
by the American Physiological Society: The sequel. Advances in Physiology Education, 31,
295–298.
& Benos, D. J. (2007). Guidelines for reporting statistics in journals published