Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_612_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
30.08.2026
Размер:
67 Мб
Скачать
Standardized Mean Difference
The effect on a continuous outcome can also be expressed in units of standard deviation (standardized mean difference, SMD = mean difference/pooled SD). SMD enables the effects between various measures or studies to be standardized. This helps interpreting the difference when the outcome measure is unfamiliar or we need to pool different outcomes in a meta­analysis. For example, consider two imaginary studies assessing the effect of wrist arthroscopy compared with a nonoperative strategy. The first study finds a mean difference of 10 points on the Disabilities of the Arm, Shoulder, and Hand (DASH) score with an SD of 20, and the second study finds a 6-point difference on Patient Rated Wrist Evaluation (PRWE) score with an SD of 12. The SMD of both studies is 0.5, and the results may be pooled together in a meta-analysis, even though they are considering two completely different outcome measures. By convention, 0.2 SMD is considered small, 0.5 SMD medium, and 0.8 SMD a large effect.5 To better understand the concept of SMD, consider the following example: The average height of adult women is 165 cm with an SD of 6 cm. An effect size of 0.2 SMD would correspond to a height difference of 1.2 cm, while a moderate effect size of 0.5 SMD would correspond to a height difference of 3 cm.
Interaction Effects and Prognostic Factors
Interaction occurs when the effect of an exposure on an outcome differs depending on the level of a third variable. For example, a study that finds the efficacy of collagenase (compared with placebo) is different in women compared with men, sex is said to have an interaction with treatment, ie, sex is an effect modifier. In this case we have subgroups with different effects. An effect modifier can also be a continuous variable, like age. If age was an effect modifier of collagenase, the effect is different depending on the age of the patient.
On the other hand, a prognostic factor is associated with the outcome, but it does not modify the effect size of the exposure. For example, depression could be associated with poor outcomes after surgery. However, depressed individuals may also have poor outcomes without surgery and the effect of surgery (difference between surgery and no surgery) is similar in patients with depression and those without depression. In this example, depression itself does not affect the absolute magnitude of surgery compared with no surgery.
CONFOUNDING, BIAS, AND IMPRECISION
Confounding
https://t.me/med1917
Confounding is present when a third variable is influencing both the exposure and the outcome (Figure 16.4). In an RCT such as the collagenase study,
1
participants were randomly assigned to a treatment modality. This effectively prevents confounding because random allocation ensures that no other variable (confounder) can systematically influence the treatment they received. However, the relationship between independent variables (such as demographics, risk factors or intervention) and dependent variables (such as outcome) is not so clear in observational studies where patients are not recruited and assigned to treatment in a random manner. Patients may differ considerably with respect to baseline demographics, disease severity, and risk factors. Some of these factors may have influenced the treatment choice.
FIGURE 16.4 Directed acyclic graphs illustrating confounding,
effect modification, and covariates.
Consider a hypothetical observational study comparing two cohorts of people with DD. One was treated with surgery and another with collagenase
https://t.me/med1917
injection. The study showed that collagenase injections achieved smaller improvements in final ROM when compared with surgery. But what if there was an extrinsic factor, like patient age, influencing why people had surgery or collagenase? An older patient is likely to present with more severe joint contractures that result in worse outcomes and might be more inclined to choose an injection over surgery because of greater perioperative risks. In this hypothetical situation, age is a “common prior cause” of both the dependent (outcome) and independent variables (treatment) and is a confounding variable (Figure 16.4). If we stratify the cohort by age (eg, <60 and ≥60 years) or examine a subset of the cohort (eg, patients less than 60 years), there may no longer be a difference between surgery and collagenase. Not all confounders are known or measured and therefore cannot be accounted for in this manner.
Bias
Bias relates to the internal validity of the study and is present when the measured effect deviates from the true effect owing to systematic error(s) in the study design (Figure 16.5). The typical sources of bias in clinical trials include nonrandom allocation of the treatment (groups are systematically different at baseline, known as selection bias) or nonblinded assessment of the outcomes (detection bias). The precise amount of bias cannot be quantified in a single study, and once bias has been introduced, it cannot be reliably removed. The reader must judge the credibility of the results based on the reported methodology.
https://t.me/med1917
FIGURE 16.5. Imprecision and bias illustrated with a target board
and forest plot. The bars present 95% confidence intervals of the effect. Biased studies have systematic error, which results in a flawed effect size (consistently higher or lower than the true effect). Imprecision is caused by uncertainty (large variations denoted by wide confidence intervals) because of too little information.
Imprecision
Imprecision refers to a situation where the amount of information is insufficient and thus the uncertainty around the effect is wide (Figure 16.5). This uncertainty can be quantified using confidence intervals (see later) and may occur because the sample just contains too little information to draw firm conclusions about the population effect. Unlike bias, imprecision can be remedied by collecting more data.
INFERENTIAL STATISTICS
Inferential statistics are mathematical approaches that use data from a sample to make assumptions about the whole population. It is impossible to study an entire population, and we usually examine a subset of representative patients so that useful information may be obtained in a timely and cost-effective manner. However, all study results are subject to random variation. Statistical models account for this “randomness” in a sample and express the uncertainty regarding the magnitude of the true population level effect. For example, Hurst et al randomly allocated 204 participants to receive collagenase injection and 104 participants to placebo for Dupuytren contracture.1 After the treatment, the
https://t.me/med1917
success rate was compared between the groups. Hypothetically, had the authors randomized the whole population (all people with DD), they could have calculated RD without using statistical methods. The two main statistical concepts that are used to make inferences about the underlying population are the frequentist and Bayesian approaches.6 Inferences can also be made in observational studies, but the findings should be considered associations and not causal effects because of the presence of potential confounders.
Hypothesis Testing and the P-value
Hypothesis testing lies at the heart of science. The baseline assumption of an experiment (trial) is that there is no effect (eg, difference between two groups that are subject to different conditions, such as treatment), and the investigator seeks to disprove this. In the study by Hurst et al, the null hypothesis was that collagenase has no effect and that the success rates between collagenase and placebo are similar.1 To test this, the researchers took a random sample from the population and randomized them to receive either collagenase or placebo. The sample of 308 people in the study is just one of many possible samples that researchers could have drawn from the entire population of patients with DD. Imagine that the authors repeated the study 10,000 times—each sample would produce a slightly different result with a mean and SD. The distribution of the means from these multiple trials is called the sampling distribution of means and we can use this to calculate the P-value (Figure 16.6).
https://t.me/med1917
FIGURE 16.6. Sampling distribution for between-group difference in
finger ROM in two hypothetical randomized controlled trials. The probability distribution on the left was obtained by simulating 10,000 trials under null hypothesis each with a sample of 60 individuals. Because there is no difference between groups when null hypothesis is true, the mean differences disperse around 0. However, owing to random variation, we may observe differences of 15° in either direction. In a hypothetical study that finds a mean difference of 7°, we can calculate the area under curve at both tails (two-tailed test) and determine the probability of observing 7° or a more extreme effect if the null hypothesis was true. This probability is the P-value. The graph on the right was simulated with 10,000 trials with a sample of 100 individuals. There is less variation around the mean because of a larger sample. The probability of finding a difference of ≥7° is less probable (area is smaller) compared with the smaller sample.
A P-value less than 0.05 means that the chance of seeing the observed outcome if the null hypothesis was true is less than 5% and we may reject the tested null hypothesis, acknowledging that there is a 5% chance that we are wrong. This value of 0.05 is an arbitrary one but we have accepted it as a conventional threshold. Thankfully, we do not need to repeat a trial 10,000 times, and knowing the mean, SD, and data distribution provides this P-value using a test statistic (see later discussion).
False Positive (Type I Error)
The use of probabilities immediately indicates that we can make wrong conclusions. A false-positive finding, also known as type I error (α), occurs
https://t.me/med1917
when we reject the null hypothesis, even though there is no true effect at the population level. A P-value of 0.05 means this situation would occur once in every 20 tests when the null hypothesis is true, and this concept is important when multiple hypotheses are being tested simultaneously. Imagine an observational study that tests the association between 10 risk factors and the outcome at three time points, resulting in 30 tests. If the null hypothesis were true in all cases, the researchers nevertheless have a good chance (79%) of finding at least one spurious P-value <.05. The researchers might be tempted to report only the significant results because it is easier to publish “positive” studies and this kind of data dredging results in spurious “significant” findings. This is one of the reasons why researchers are encouraged to publish the study protocols a priori, before collecting and analyzing the data. Various techniques such as Bonferroni correction, Tukey method, and Holm(­Bonferroni) method may be used to mathematically account for multiple testing.
False Negative (Type II Error) and Power
A false-negative rate (concluding “no effect” when there is an effect) is known as a type II error.7 A type II error rate (β) closely relates to the sample size. The complement of the type II error rate, the probability of detecting an effect when it is present, is called power (1-β). The power increases (and the type II error rate decreases) with increasing sample size. A power calculation can be made based on the predetermined acceptable levels of type I error (α), power (1-β), and the estimated effect size to determine an appropriate sample size before a study is initiated. Power analysis is not meant to assess the power of a study when the results are already known and the study reports no significant effect. Instead, it is far more productive to assess the precision of the effect estimates by assessing the confidence intervals.
Confidence Intervals
A sampling distribution of effects from the 10,000 trials (Figure 16.6) displays how the effect varies between samples. The range of values within which 95% of the effect will be found is called 95% confidence interval (CI). Instead of repeating trials, we can calculate 95% CIs for the observed effect. These values can be considered as the range of effects on population level that are compatible with the observed effect in the sample.
P-Values or Confidence Intervals?
There has been growing interest to abandon P-values altogether.8 Binary interpretation (statistically significant vs nonsignificant) is seldom useful. As described earlier, a P-value >.05 should not be interpreted as no effect because it simply describes the probability of the observed effect occurring if
https://t.me/med1917
the null hypothesis was true. Second, a P-value is not a measure of effect size and does not convey any relevance to clinical decision making. A small P­value does not signify a large or more clinically meaningful effect, and trivial treatment effects may produce small P-values simply because of large sample sizes. CIs are a more useful way to communicate and interpret the results, because they convey statistical significance, magnitude, and precision all at once. CIs can also be compared with clinically minimal important differences to assess whether the effect is relevant to most patients (Figure 16.7).
https://t.me/med1917
FIGURE 16.7. A forest plot showing six hypothetical RCTs
comparing treatment A and B. The bars represent the 95% confidence intervals (CI) of the treatment effect in each study. Studies 1 to 3 found a statistically significant effect (CIs do not overlap 0) while studies 4 to 6 did not. Instead of looking at statistical significance, we can interpret the results based on the 95% CIs and whether they exclude clinically minimal important difference. Studies 2, 4, and 6 are inconclusive whether treatment A is clinically superior to B (the effect estimates are imprecise), although the null hypothesis was rejected with study 2. Studies 1, 3, and 5 yield precise estimates. Studies 3 and 5 exclude clinically relevant benefits while study 1 suggests a relevant benefit.
WHICH TEST SHOULD WE USE?
The choice of appropriate test for each null hypothesis depends on the type of variables and whether the observations are independent or paired (Table 16.4). The observations are paired if they are obtained from the same individual, such as a crossover trial in which each participant receives several interventions, and we compare the effects of these interventions. On the other
https://t.me/med1917
hand, observations are considered independent when we compare two separate cohorts of individuals, such as in RCTs.
TABLE 16.4. COMMONLY USED PARAMETRIC AND
NONPARAMETRIC TESTS
Parametric Test
Nonparametric Test
Purpose
Two sample T-test Mann-Whitney U
test
Compare two independent samples
Paired T-test Wilcoxon test Compare two dependent
samples
One-way ANOVA Kruskal-Wallis
test
Compare more than two samples
Repeated measures ANOVA
Friedman test Multiple measurements/factors
Pearson correlation
Spearman rank correlation
Quantify the correlation between two variables
Parametric Tests
Parametric tests compare means of continuous variables that are normally distributed. They can be used for both independent and dependent measurements.
T-tests can be used to compare the means of two groups. Student T-test and Welch T-test are appropriate when the measurements are independent, and paired T-tests are used when the observations are dependent. In the T-test, a test statistic known as the T-value is calculated by dividing the observed mean difference between the groups by the pooled standard error (SE). This T-value can then be compared with the sampling distribution of T-values to determine the probability of observing such a T-value or a more extreme value if the null hypothesis was true. The probability can be calculated in both tails (two-tailed test) or one tail (one-tailed test) of the distribution. One-tailed tests are typically used only in noninferiority studies where the researchers want to find evidence that a new intervention is not inferior to the reference intervention. Two-tailed tests are used in situations where random variation is equally likely to cause extreme observations in both directions. This statistical approach is fundamental to all statistical tests.
https://t.me/med1917