Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_612_Библиотеки_им_академика_М_И_Перельмана
.pdf
Standardized Mean Difference
The effect on a continuous outcome can also be expressed in units of standard
deviation (standardized mean difference, SMD = mean difference/pooled SD).
SMD enables the effects between various measures or studies to be
standardized. This helps interpreting the difference when the outcome
measure is unfamiliar or we need to pool different outcomes in a metaanalysis. For example, consider two imaginary studies assessing the effect of
wrist arthroscopy compared with a nonoperative strategy. The first study finds
a mean difference of 10 points on the Disabilities of the Arm, Shoulder, and
Hand (DASH) score with an SD of 20, and the second study finds a 6-point
difference on Patient Rated Wrist Evaluation (PRWE) score with an SD of 12.
The SMD of both studies is 0.5, and the results may be pooled together in a
meta-analysis, even though they are considering two completely different
outcome measures. By convention, 0.2 SMD is considered small, 0.5 SMD
medium, and 0.8 SMD a large effect.5 To better understand the concept of
SMD, consider the following example: The average height of adult women is
165 cm with an SD of 6 cm. An effect size of 0.2 SMD would correspond to a
height difference of 1.2 cm, while a moderate effect size of 0.5 SMD would
correspond to a height difference of 3 cm.
Interaction Effects and Prognostic Factors
Interaction occurs when the effect of an exposure on an outcome differs
depending on the level of a third variable. For example, a study that finds the
efficacy of collagenase (compared with placebo) is different in women
compared with men, sex is said to have an interaction with treatment, ie, sex is
an effect modifier. In this case we have subgroups with different effects. An
effect modifier can also be a continuous variable, like age. If age was an effect
modifier of collagenase, the effect is different depending on the age of the
patient.
On the other hand, a prognostic factor is associated with the outcome, but it
does not modify the effect size of the exposure. For example, depression could
be associated with poor outcomes after surgery. However, depressed
individuals may also have poor outcomes without surgery and the effect of
surgery (difference between surgery and no surgery) is similar in patients with
depression and those without depression. In this example, depression itself
does not affect the absolute magnitude of surgery compared with no surgery.
CONFOUNDING, BIAS, AND IMPRECISION
Confounding
https://t.me/med1917

Confounding is present when a third variable is influencing both the exposure
and the outcome (Figure 16.4). In an RCT such as the collagenase study,
1
participants were randomly assigned to a treatment modality. This effectively
prevents confounding because random allocation ensures that no other
variable (confounder) can systematically influence the treatment they received.
However, the relationship between independent variables (such as
demographics, risk factors or intervention) and dependent variables (such as
outcome) is not so clear in observational studies where patients are not
recruited and assigned to treatment in a random manner. Patients may differ
considerably with respect to baseline demographics, disease severity, and risk
factors. Some of these factors may have influenced the treatment choice.
FIGURE 16.4 Directed acyclic graphs illustrating confounding,
effect modification, and covariates.
Consider a hypothetical observational study comparing two cohorts of
people with DD. One was treated with surgery and another with collagenase
https://t.me/med1917

injection. The study showed that collagenase injections achieved smaller
improvements in final ROM when compared with surgery. But what if there was
an extrinsic factor, like patient age, influencing why people had surgery or
collagenase? An older patient is likely to present with more severe joint
contractures that result in worse outcomes and might be more inclined to
choose an injection over surgery because of greater perioperative risks. In this
hypothetical situation, age is a “common prior cause” of both the dependent
(outcome) and independent variables (treatment) and is a confounding variable
(Figure 16.4). If we stratify the cohort by age (eg, <60 and ≥60 years) or
examine a subset of the cohort (eg, patients less than 60 years), there may no
longer be a difference between surgery and collagenase. Not all confounders
are known or measured and therefore cannot be accounted for in this manner.
Bias
Bias relates to the internal validity of the study and is present when the
measured effect deviates from the true effect owing to systematic error(s) in
the study design (Figure 16.5). The typical sources of bias in clinical trials
include nonrandom allocation of the treatment (groups are systematically
different at baseline, known as selection bias) or nonblinded assessment of the
outcomes (detection bias). The precise amount of bias cannot be quantified in
a single study, and once bias has been introduced, it cannot be reliably
removed. The reader must judge the credibility of the results based on the
reported methodology.
https://t.me/med1917

FIGURE 16.5. Imprecision and bias illustrated with a target board
and forest plot. The bars present 95% confidence intervals of the
effect. Biased studies have systematic error, which results in a
flawed effect size (consistently higher or lower than the true effect).
Imprecision is caused by uncertainty (large variations denoted by
wide confidence intervals) because of too little information.
Imprecision
Imprecision refers to a situation where the amount of information is insufficient
and thus the uncertainty around the effect is wide (Figure 16.5). This
uncertainty can be quantified using confidence intervals (see later) and may
occur because the sample just contains too little information to draw firm
conclusions about the population effect. Unlike bias, imprecision can be
remedied by collecting more data.
INFERENTIAL STATISTICS
Inferential statistics are mathematical approaches that use data from a sample
to make assumptions about the whole population. It is impossible to study an
entire population, and we usually examine a subset of representative patients
so that useful information may be obtained in a timely and cost-effective
manner. However, all study results are subject to random variation. Statistical
models account for this “randomness” in a sample and express the uncertainty
regarding the magnitude of the true population level effect. For example, Hurst
et al randomly allocated 204 participants to receive collagenase injection and
104 participants to placebo for Dupuytren contracture.1 After the treatment, the
https://t.me/med1917

success rate was compared between the groups. Hypothetically, had the
authors randomized the whole population (all people with DD), they could have
calculated RD without using statistical methods. The two main statistical
concepts that are used to make inferences about the underlying population are
the frequentist and Bayesian approaches.6 Inferences can also be made in
observational studies, but the findings should be considered associations and
not causal effects because of the presence of potential confounders.
Hypothesis Testing and the P-value
Hypothesis testing lies at the heart of science. The baseline assumption of an
experiment (trial) is that there is no effect (eg, difference between two groups
that are subject to different conditions, such as treatment), and the investigator
seeks to disprove this. In the study by Hurst et al, the null hypothesis was that
collagenase has no effect and that the success rates between collagenase and
placebo are similar.1 To test this, the researchers took a random sample from
the population and randomized them to receive either collagenase or placebo.
The sample of 308 people in the study is just one of many possible samples
that researchers could have drawn from the entire population of patients with
DD. Imagine that the authors repeated the study 10,000 times—each sample
would produce a slightly different result with a mean and SD. The distribution
of the means from these multiple trials is called the sampling distribution of
means and we can use this to calculate the P-value (Figure 16.6).
https://t.me/med1917

FIGURE 16.6. Sampling distribution for between-group difference in
finger ROM in two hypothetical randomized controlled trials. The
probability distribution on the left was obtained by simulating 10,000
trials under null hypothesis each with a sample of 60 individuals.
Because there is no difference between groups when null hypothesis
is true, the mean differences disperse around 0. However, owing to
random variation, we may observe differences of 15° in either
direction. In a hypothetical study that finds a mean difference of 7°,
we can calculate the area under curve at both tails (two-tailed test)
and determine the probability of observing 7° or a more extreme
effect if the null hypothesis was true. This probability is the P-value.
The graph on the right was simulated with 10,000 trials with a
sample of 100 individuals. There is less variation around the mean
because of a larger sample. The probability of finding a difference of
≥7° is less probable (area is smaller) compared with the smaller
sample.
A P-value less than 0.05 means that the chance of seeing the observed
outcome if the null hypothesis was true is less than 5% and we may reject the
tested null hypothesis, acknowledging that there is a 5% chance that we are
wrong. This value of 0.05 is an arbitrary one but we have accepted it as a
conventional threshold. Thankfully, we do not need to repeat a trial 10,000
times, and knowing the mean, SD, and data distribution provides this P-value
using a test statistic (see later discussion).
False Positive (Type I Error)
The use of probabilities immediately indicates that we can make wrong
conclusions. A false-positive finding, also known as type I error (α), occurs
https://t.me/med1917

when we reject the null hypothesis, even though there is no true effect at the
population level. A P-value of 0.05 means this situation would occur once in
every 20 tests when the null hypothesis is true, and this concept is important
when multiple hypotheses are being tested simultaneously. Imagine an
observational study that tests the association between 10 risk factors and the
outcome at three time points, resulting in 30 tests. If the null hypothesis were
true in all cases, the researchers nevertheless have a good chance (79%) of
finding at least one spurious P-value <.05. The researchers might be tempted
to report only the significant results because it is easier to publish “positive”
studies and this kind of data dredging results in spurious “significant” findings.
This is one of the reasons why researchers are encouraged to publish the
study protocols a priori, before collecting and analyzing the data. Various
techniques such as Bonferroni correction, Tukey method, and Holm(Bonferroni) method may be used to mathematically account for multiple
testing.
False Negative (Type II Error) and Power
A false-negative rate (concluding “no effect” when there is an effect) is known
as a type II error.7 A type II error rate (β) closely relates to the sample size. The
complement of the type II error rate, the probability of detecting an effect when
it is present, is called power (1-β). The power increases (and the type II error
rate decreases) with increasing sample size. A power calculation can be made
based on the predetermined acceptable levels of type I error (α), power (1-β),
and the estimated effect size to determine an appropriate sample size before a
study is initiated. Power analysis is not meant to assess the power of a study
when the results are already known and the study reports no significant effect.
Instead, it is far more productive to assess the precision of the effect estimates
by assessing the confidence intervals.
Confidence Intervals
A sampling distribution of effects from the 10,000 trials (Figure 16.6) displays
how the effect varies between samples. The range of values within which 95%
of the effect will be found is called 95% confidence interval (CI). Instead of
repeating trials, we can calculate 95% CIs for the observed effect. These
values can be considered as the range of effects on population level that are
compatible with the observed effect in the sample.
P-Values or Confidence Intervals?
There has been growing interest to abandon P-values altogether.8 Binary
interpretation (statistically significant vs nonsignificant) is seldom useful. As
described earlier, a P-value >.05 should not be interpreted as no effect
because it simply describes the probability of the observed effect occurring if
https://t.me/med1917

the null hypothesis was true. Second, a P-value is not a measure of effect size
and does not convey any relevance to clinical decision making. A small Pvalue does not signify a large or more clinically meaningful effect, and trivial
treatment effects may produce small P-values simply because of large sample
sizes. CIs are a more useful way to communicate and interpret the results,
because they convey statistical significance, magnitude, and precision all at
once. CIs can also be compared with clinically minimal important differences to
assess whether the effect is relevant to most patients (Figure 16.7).
https://t.me/med1917

FIGURE 16.7. A forest plot showing six hypothetical RCTs
comparing treatment A and B. The bars represent the 95%
confidence intervals (CI) of the treatment effect in each study.
Studies 1 to 3 found a statistically significant effect (CIs do not
overlap 0) while studies 4 to 6 did not. Instead of looking at statistical
significance, we can interpret the results based on the 95% CIs and
whether they exclude clinically minimal important difference. Studies
2, 4, and 6 are inconclusive whether treatment A is clinically superior
to B (the effect estimates are imprecise), although the null
hypothesis was rejected with study 2. Studies 1, 3, and 5 yield
precise estimates. Studies 3 and 5 exclude clinically relevant
benefits while study 1 suggests a relevant benefit.
WHICH TEST SHOULD WE USE?
The choice of appropriate test for each null hypothesis depends on the type of
variables and whether the observations are independent or paired (Table 16.4).
The observations are paired if they are obtained from the same individual,
such as a crossover trial in which each participant receives several
interventions, and we compare the effects of these interventions. On the other
https://t.me/med1917

hand, observations are considered independent when we compare two
separate cohorts of individuals, such as in RCTs.
TABLE 16.4. COMMONLY USED PARAMETRIC AND
NONPARAMETRIC TESTS
Parametric Test
Nonparametric
Test
Purpose
Two sample T-test Mann-Whitney U
test
Compare two independent
samples
Paired T-test Wilcoxon test Compare two dependent
samples
One-way ANOVA Kruskal-Wallis
test
Compare more than two
samples
Repeated
measures ANOVA
Friedman test Multiple measurements/factors
Pearson
correlation
Spearman rank
correlation
Quantify the correlation
between two variables
Parametric Tests
Parametric tests compare means of continuous variables that are normally
distributed. They can be used for both independent and dependent
measurements.
T-tests can be used to compare the means of two groups. Student T-test and
Welch T-test are appropriate when the measurements are independent, and
paired T-tests are used when the observations are dependent. In the T-test, a
test statistic known as the T-value is calculated by dividing the observed mean
difference between the groups by the pooled standard error (SE). This T-value
can then be compared with the sampling distribution of T-values to determine
the probability of observing such a T-value or a more extreme value if the null
hypothesis was true. The probability can be calculated in both tails (two-tailed
test) or one tail (one-tailed test) of the distribution. One-tailed tests are typically
used only in noninferiority studies where the researchers want to find evidence
that a new intervention is not inferior to the reference intervention. Two-tailed
tests are used in situations where random variation is equally likely to cause
extreme observations in both directions. This statistical approach is
fundamental to all statistical tests.
https://t.me/med1917
Соседние файлы в папке Библиотека им академика М.И. Перельмана
