Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_2767_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
25 Мб
Скачать
11.7 Investigations ofsources ofheterogeneity
https://t.me/medicina_free
uncommon in systematic reviews because data needed for such comparisons are often not reported in the included studies. In this Cochrane Review the subgroups were not formally compared using meta- regression due to the small number of studies and dif­ferent test positivity thresholds used in the studies.
Meta- analytical results from heterogeneity analyses are often presented in a table and graphically in SROC plots displaying summary points or SROC curves for each cat­egory. The investigation of heterogeneity presented in Table11.7.a compared studies of type 1 rapid diagnostic tests for Plasmodium falciparum malaria performed in Africa and studies performed in Asia.
Care should be exercised in interpreting the results of heterogeneity investigations. There are several points that should be considered.
First, subgroup findings are more credible if they have a scientific rationale. Ideally, selection of characteristics for investigation should be motivated by biological, clinical and methodological hypotheses supported by evidence from other sources. Subgroup analyses based on characteristics that are implausible or irrelevant are unlikely to be useful and should be avoided.
Second, exploratory heterogeneity investigations are less trustworthy than those that were pre-
specified in the protocol. Review authors are expected to report whether het­erogeneity investigations were pre- specified or data driven. Exploratory analyses are often data driven and prompted by observations made in informal data analyses, rather than by carefully crafted research hypotheses, supported by prior evidence. Pre­specification of investigations of heterogeneity in systematic reviews is often difficult, as review authors can be aware of some of the results in the included studies before they start the review.
Third, the chance of obtaining spurious significant results increases with the number of statistical comparisons that are undertaken. There is no formal rule on the maximum number of investigations that can be undertaken. The total number of hypotheses investigated must be kept in mind when interpreting the significance of the results. Adjustments to P values using rules for multiple testing are not encouraged, as they will be overly conservative due to the inevitable correlations between the factors investigated.
Fourth, heterogeneity investigations based on small numbers of studies are unlikely to produce useful findings and should be avoided. The statistical power of a compari­son depends on the number of studies as well as on the precision of the test accuracy estimates from each study. When the characteristic is unevenly distributed across
Table11.7.a Investigation of heterogeneity between studies of type 1 rapid diagnostic tests for
Plasmodium falciparum malaria
Continent Studies Patients Malaria
Africa 39 21,958 7445 94.0 (91.2 to 96.0) 93.1 (89.7 to 95.3) P = 0.03
Asia 24 15,810 4060 96.4 (93.7 to 97.9) 96.6 (94.0 to 98.1)
CI, confidence interval. P value obtained from a likelihood ratio test comparing models with and without the covariate. Source: Adapted from Table7in Abba 2011.
cases
Summary sensitivity % (95% CI)
Summary specificity % (95% CI)
P value
339
11 Presenting findings
https://t.me/medicina_free
groups, it is possible that important differences may be missed. Since most compari­sons will not be made within study participants, or based on randomization, statisti­cally significant differences may not reflect the actual cause of the difference.
Fifth, only characteristics that were reported at study level or in subgroups defined within studies can be investigated. This limits the ability to detect relevant associations with test accuracy on an individual level. For example, if accuracy varies with age but study groups are similar in mean age, no association will be detected in a systematic review that relies on aggregate data. This problem is known as aggregation or ecologi­cal bias: the failure of ecological (aggregate)- level associations to properly reflect individual- level associations.
Sixth, it should be remembered that most subgroup comparisons are observational, and suffer from the same limitations as other comparisons in observational research. It may not be appropriate to make a causal interpretation of observed differences, although causal effects may have the greatest relevance for clinical practice. Confounding has to be considered: a difference between subgroups may be influenced by other factors that are also associated with differences in test accuracy. For example, if reference standards varied over time, but there were also changes over time in the composition of the study groups, it will not be possible to identify which, if either, is the cause of observed differences in test accuracy. Multivariable analysis, investigating multiple sources of heterogeneity in parallel, is usually infeasible due to the limited number of studies available (Takwoingi 2020).
Many review authors discover sooner or later that their plans for investigating hetero­geneity are infeasible, either because of there being too few studies, or because studies do not report the desired information. Cochrane Reviews of diagnostic test accuracy contain a dedicated section to describe how the review differs from the protocol, where these issues can be described.
11.8 Re- expressing summary estimates numerically
11.8.1 Frequencies
Sensitivity and specificity are typically presented as proportions or percentages, and so are positive and negative predictive values. Presenting probabilities as frequencies has been shown to help readers understand their relative magnitude (Hoffrage 1998, Evans 2000, Zhelev 2013), and this approach is encouraged both in the ‘Summary of main results’ section of the review and in the ‘Summary of findings’ table.
A natural frequency description expresses a proportion as the number of individuals out of a group (typically 10, 100 or 1000) in whom an event or outcome is observed. A 95% sensitivity can be presented as a group of 100 persons with the target condition, of which 95 test positive.
As with conditional probabilities, one should be explicit about the group to which natural frequencies refer. For example, they may refer to all those tested, those with or without the target condition, or those with positive or with negative index test results.
Although sensitivity and specificity do not provide information on the absolute impact of a test at a particular prevalence of the target condition, expressing them as natural frequencies may help readers to interpret them. In addition, natural frequencies
340
11.8 Re- expressing summary estimates numerically
https://t.me/medicina_free
explicitly illustrate that sensitivity provides information on the false negatives and specificity on the false positives.
For example:
For a test with a sensitivity of 90%: the index test will detect 90 out of every 100with the target condition, but 10will be missed (i.e. will be false negatives).
For a test with a specificity of 80%: of every 100individuals without the target condi­tion, 20will be wrongly diagnosed as having it (i.e. will be false positives).
11.8.2 Predictive values
There is a considerable body of empirical evidence demonstrating that sensitivity and specificity can be difficult to interpret for many clinicians and patients (Steurer 2002, Puhan 2005). Some researchers have argued that probabilities conditional on index test results (predictive values) rather than actual disease status (sensitivity and specificity) may be more intuitive to decision makers (Reid 1998).
Historically the use of predictive values has been discouraged, because sensitivity and specificity were assumed to be statistically independent of the proportion of par­ticipants with the target condition in the study group. If there are more study partici­pants with the target condition, and sensitivity and specificity do not change, positive predictive values will be higher and negative predictive values lower. We now know that this is a simplification, as we are becoming increasingly aware of the variation in sensi­tivity and specificity caused by differences in the setting, previous testing and severity of disease (spectrum of disease) (Leeflang 2012). Review authors should therefore be mindful of the transferability of any accuracy measures, regardless of the type of sum­mary statistic used.
Meta- analysis of predictive values is possible and sometimes necessary (Leeflang
2012), but it is not recommended to perform two meta- analyses in parallel: one of sen­sitivity and specificity and a second, separate meta- analysis of positive and negative predictive values. If review authors wish to use predictive values as a means of express­ing test accuracy from a meta- analysis producing summary estimates of sensitivity and specificity, they should compute predictive values based on these estimates for a repre­sentative proportion of those with the target condition.
Predictive values are most simply obtained from summary estimates of sensitivity and specificity by creating an illustrative 2×2 table and computing predictive values directly (the simple equations to do this are in Chapter4). This exercise can be done on paper, using a spreadsheet or a calculator tool (see Figure11.8.a).
To compute predictive values manually or using a calculator, one needs a fictional group size (say, 1000), the proportion with the target condition (‘prevalence’ in the cal­culator), and the summary estimates of sensitivity and specificity of the test– i.e. the boxes in green in Figure11.8.a. For example, a test that has sensitivity of 0.9 and speci­ficity of 0.8 yields the values in Figure11.8.a if 0.25 (25%) have the target condition. This computes the positive predictive value to be 0.60 and the negative predictive value to be 0.96. The same computations can be done using Bayes equation, as shown in Box11.8.a.
Natural frequencies can be used to describe the absolute impact of a test in a popula­tion with a given prevalence (25% in Figure11.8.a and Box11.8.a).
341
Figure11.8.a Using a calculator to convert sensitivity and specificity to positive and negative predictive values at a prevalence of25%. D+,
https://t.me/medicina_free
disease positive; D–, disease negative; FP, false positive; FN, false negative; LR+, positive likelihood ratio; LR–, negative likelihood ratio; NPV, negative predictive value; PPV, positive predictive value; TN, true negative; TP, true positive.
11.8 Re- expressing summary estimates numerically
sensitivity prevalence
https://t.me/medicina_free
Box 11.8.a Calculation ofpredictive values using Bayes equation andestimates ofsensitivity, specificity andprevalence
sensitivity prevalence 1 specificity 1 prevalence
0.9 0.25
0.9 0.25 1 0.8 1 0.25
specificity 1 prevalence
1 sensitivity prevalence specificity 1 prevalence
0.8 1 0.25
1 0.9 0.25 0.8 1 0.25
NPV, negative predictive value; PPV, positive predictive value.
For a test with a positive predictive value of 60%: 60 out of every 100 positive index
test results will have the target condition, but 40will not (i.e. will be false positives). In
a group in which 25% have the target condition, this will result in 150 false positive
test results for every 1000 people tested.
For a test with a negative predictive value of 96%: 96 out of every 100negative index
test results will not have the target condition, but 4will (i.e. will be false negatives). In
a group in which 25% have the target condition, this will result in 25 false negative
test results for every 1000 people tested.
Although predictive values may be intuitive summary metrics, choosing the propor­tion of those with the target condition (prevalence) is not straightforward. The term ‘prevalence’ is often used to refer to the proportion of those with the target condition, for lack of a better term, but the proportion of those with the target condition in the population being tested will rarely be equal to the prevalence of the target condition in the population. In testing for SARS- CoV- 2 infection, for example, the proportion with the SARS- CoV- 2 virus in symptomatic persons undergoing testing will be substantially higher than the population prevalence. Similarly, the proportion with the virus in asymptomatic contacts of persons who recently tested positive will also be higher than the population prevalence, but lower than the proportion in symptomatic persons.
To select a proportion for the conversion to predictive values, the median value for the proportion of those with the target condition might be used, if that median is calcu­lated from studies that relied on consecutive or random sampling of participants in the intended- use setting. Because of their design, studies that separately recruited partici­pants with the target condition and healthy controls should be excluded from calculat­ing the median proportion. Alternatively, review authors may consider computing predictive values across a range of plausible values for the intended- use setting. In some circumstances, estimates of disease prevalence may be more reliably obtained
343
11 Presenting findings
https://t.me/medicina_free
from other data sources, such as disease registries that correspond to the settings in which the studies in the meta- analysis were conducted.
The proportion of those with the target condition selected for presenting predictive values usually falls in the range of corresponding proportions observed in the studies included in the meta- analysis. The selection of a value outside that range would be an extrapolation, and should be done with caution.
11.8.3 Likelihood ratios
The use of likelihood ratios to express test performance (see Chapter4) has been pro­moted by some as a metric that facilitates Bayesian probability updating: the calcula­tion of post­evidence that likelihood ratios improve diagnostic decision- making in groups of clini­cians is lacking.
Parallel or separate meta- analysis of likelihood ratios is not recommended (see Chapter 9, Section 9.3). Summary estimates of the positive and negative likelihood ratios can be calculated from the summary estimates of sensitivity and specificity. These summary positive and negative likelihood ratios (and their 95% confidence inter­vals) can be obtained as additional estimates, as indicated in Chapter9, Section9.3, and illustrated in the software code in the appendices of Chapter10.
Ideally, the confidence intervals for likelihood ratios should be used to calculate con­fidence intervals for predictive values because, unlike sensitivity and specificity, the positive likelihood ratio jointly considers the number of true and false positives and the negative likelihood ratio jointly considers the number of true and false negatives. Using likelihood ratio outputs from SAS and Stata, the lower and upper confidence limits of the positive likelihood ratio can be converted into lower and upper confidence limits of the positive predictive value, at a stated proportion with the target condition. Likewise, confidence limits of the negative likelihood ratio can be converted into confidence lim­its of the negative predictive value.
The simplest approach for deriving predictive values from likelihood ratios is to use a calculator tool, similar to the one for deriving point estimates of the predictive values (see Figure11.8.a). Alternatively, predictive values can be calculated using the equa­tions in Box11.8.b.
Only the uncertainty in the summary estimates of sensitivity and specificity as expres­sions of test accuracy is captured in these 95% confidence intervals, not the uncertainty in the proportion with the target condition. This uncertainty could be explored by com­puting point estimates and 95% confidence intervals for predictive values across a range of plausible values for the proportion with the target condition (see Chapter12).
test probabilities for specific pre- test probabilities (Straus 2019). Actual
11.9 Presenting findings when meta- analysis cannot beperformed
Meta- analysis is not always possible due to few studies, convergence issues with the analysis (see Chapter10, Section10.6), clinical heterogeneity or other factors. Such sit­uations are common. Of the 135 Cochrane Reviews of diagnostic test accuracy pub­lished up to 31July 2020, 32 (24%) did not include a meta- analysis. In the 32 reviews,
344
11.9 Presenting findings when meta- analysis cannot beperformed
https://t.me/medicina_free
Box 11.8.b Calculation ofpredictive values using post- test odds andlikelihood ratios
Pre- test odds = pre- test probability/(1 − pre-test probability) Post-
test odds of disease given positive test result = pre- test odds of disease × LR+ test odds of disease given negative test result = pre- test odds of disease × LR–
Post­Post-
test probability = Post- test odds/(1 + post- test odds) test probability of disease given a positive test result = PPV
Post­Post-
test probability of disease given a negative test result = 1– NPV
Repeating the example shown in Figure11.8.a, a test that has a positive likelihood ratio (LR+) of 4.5 and negative likelihood (LR–) of 0.125will give a positive predictive value (PPV) of 0.60 and negative predictive value (NPV) of 0.96 using the previous equations as follows.
test odds = 0.25/(1– 0.25) = 1/3
Pre­Post- test odds of disease given positive test result = (1/3) × 4.5 = 1.5 Post- test odds of disease given negative test result = (1/3) × 0.125 = 0.0417
test probability of disease given a positive test result = 1.5/(1 + 1.5) = 0.60
Post­Post- test probability of disease given a negative test result = 0.0417/(1 + 0.0417) = 0.04, therefore NPV = 1– 0.04 = 0.96
the number of included studies ranged between 1 and 33; the median was 4 (inter­quartile range 3 to 9). A frequent reason for not performing meta- analysis was consid­erable between- study variation in clinical and methodological characteristics. For example, Chan (2019) did not perform meta- analysis due to risk- of- bias concerns and heterogeneity.
Tables, forest plots and/or SROC plots are valuable tools for presenting findings in reviews without meta- analysis. Plots provide a visual summary across studies and are more readily accessible to readers than listing many results from individual studies in the text. When there are few included studies, an SROC plot showing point estimates of sensitivity and specificity and their confidence intervals can be an effective display if study points and confidence interval lines do not overlap considerably (see Figure11.9.a).
Review authors may also want to present and describe the results of individual stud­ies. For example, Crawford (2016) included four studies and described the results of each study in the text and outlined key findings from the studies in the ‘Summary of findings’ table. Describing the results of individual studies will be impractical when there are many studies. Instead, review authors could provide a narrative summary that includes the ranges of reported estimates of sensitivity and specificity and the body of evidence (number of studies, number of participants, number with the target condition).
Presenting findings in the absence of a meta- analysis becomes even more challeng­ing when there are multiple index tests, target conditions or reference standards. In such situations, review authors should carefully consider how best to present the find­ings, using text, tables and figures as appropriate.
For example, a Cochrane Review reported that the 33included studies evaluated a myriad of index tests against different reference standards (Hanchard 2013). In addi­tion, there were five categories of the target condition. Altogether, there were 170 com­binations of target condition and index test. None of these combinations was assessed in a similar manner by more than two studies. The review authors concluded that
345
11 Presenting findings
1
Sensitivity (95% CI)
0
https://t.me/medicina_free
0.9
0.8
0.7
0.6
0.5
0.4
0.3
0.2
0.1
0
1 0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1
Legend
POD:3 DFA > 600 IU/L POD:3 to 5 DFA > 3 times serum amylase POD:4 DFA > 647 U/L
Specificity (95% CI)
POD:5 DFA > 3 times serum amylase POD:5 DFA > 4000 U/L
346
Figure11.9.a SROC plot with point estimates of sensitivity and specificity and 95% confidence
intervals. The Cochrane Review assessed amylase in drain fluid for detecting pancreatic leak in post- pancreatic resection. In the five included studies, drain fluid amylase was measured on different days and at different thresholds. A meta- analysis was not performed. The numbers following POD (postoperative day) indicate the number of the postoperative day. The numbers or text following DFA (drain fluid amylase) indicate the threshold. Source: Adapted from Davidson 2017
meta- analysis was inappropriate because of substantial clinical heterogeneity and a limited number of studies for each combination. To keep the number of forest plots to a minimum and to enhance readability, the review authors presented estimates of sen­sitivity and specificity on forest plots grouped according to target condition. A narrative summary with a similar structure accompanied the forest plots.
11.10 Chapter information
Authors: Jonathan J. Deeks (Institute of Applied Health Research, University of Birmingham, UK), Patrick M. Bossuyt (Department of Epidemiology and Data Science,
University of Amsterdam, The Netherlands), Mariska M. Leeflang (Department of Epidemiology and Data Science, University of Amsterdam, The Netherlands), Yemisi Takwoingi (Institute of Applied Health Research, University of Birmingham, UK).
11.11 References
https://t.me/medicina_free
Sources of support: Jonathan J. Deeks is a UK National Institute for Health Research (NIHR) Senior Investigator Emeritus. Yemisi Takwoingi is funded by a UK National Institute for Health Research (NIHR) Postdoctoral Fellowship. Jonathan J. Deeks and Yemisi Takwoingi are supported by the NIHR Birmingham Biomedical Research Centre at the University Hospitals Birmingham NHS Foundation Trust and the University of Birmingham. The views expressed are those of the authors and not necessarily those of the NHS, the NIHR or the Department of Health and Social Care. The authors declare no sources of support for writing this chapter.
Declarations of interest: Jonathan J. Deeks, Yemisi Takwoingi and Mariska M. Leeflang are members of Cochrane’s Diagnostic Test Accuracy Editorial Team. Yemisi Takwoingi and Mariska M. Leeflang are co-
convenors of the Cochrane Screening and Diagnostic Tests Methods Group. The authors declare no other potential conflicts of interest rele­vant to the topic of this chapter.
Acknowledgements: The authors would like to thank Karen R. Steingart, Marta Roque and Daniël Korevaar for helpful peer review comments.
11.11 References
Abba K, Deeks JJ, Olliaro P, Naing CM, Jackson SM, Takwoingi Y, Donegan S, Garner P. Rapid
diagnostic tests for diagnosing uncomplicated P. falciparum malaria in endemic countries. Cochrane Database of Systematic Reviews 2011; 7: CD008122.
Alldred SK, Takwoingi Y, Guo B, Pennant M, Deeks JJ, Neilson JP, Alfirevic Z. First trimester
ultrasound tests alone or in combination with first trimester serum tests for Down’s syndrome screening. Cochrane Database of Systematic Reviews 2017; 3: CD012600.
Best LM, Takwoingi Y, Siddique S, Selladurai A, Gandhi A, Low B, Yaghoobi M, Gurusamy KS.
Non- invasive diagnostic tests for Helicobacter pylori infection. Cochrane Database of Systematic Reviews 2018; 3: CD012080.
Chan CC, Fage BA, Burton JK, Smailagic N, Gill SS, Herrmann N, Nikolaou V, Quinn TJ,
Storr AH, Seitz DP. Mini- Cog for the diagnosis of Alzheimer’s disease dementia and
Noel­other dementias within a secondary care setting. Cochrane Database of Systematic Reviews 2019; 9: CD011414.
Chan KK, Joo DA, McRae AD, Takwoingi Y, Premji ZA, Lang E, Wakai A. Chest ultrasonography
versus supine chest radiography for diagnosis of pneumothorax in trauma patients in the emergency department. Cochrane Database of Systematic Reviews 2020; 7: CD013031.
Crawford F, Andras A, Welch K, Sheares K, Keeling D, Chappell FM. D-
the diagnosis of pulmonary embolism. Cochrane Database of Systematic Reviews 2016; 8:CD010864.
Davidson TB, Yaghoobi M, Davidson BR, Gurusamy KS. Amylase in drain fluid for the
diagnosis of pancreatic leak in post- pancreatic resection. Cochrane Database of Systematic Reviews 2017; 4: CD012009.
Deeks JJ, Higgins JPT, Altman DG. Chapter10: Analysing data and undertaking meta-
analyses. In: Higgins JPT, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, Welch VA, editors. Cochrane Handbook for Systematic Reviews of Interventions version 6.0. Cochrane, 2019.
dimer test for excluding
347
11 Presenting findings
https://t.me/medicina_free
Evans JS, Handley SJ, Perham N, Over DE, Thompson VA. Frequency versus probability
formats in statistical word problems. Cognition 2000; 77: 197–213.
Hanchard NC, Lenza M, Handoll HH, Takwoingi Y. Physical tests for shoulder impingements
and local lesions of bursa, tendon or labrum that may accompany impingement. Cochrane Database of Systematic Reviews 2013; 4: CD007427.
Hoffrage U, Gigerenzer G. Using natural frequencies to improve diagnostic inferences.
Academic Medicine 1998; 73: 538–540.
Kohli M, Schiller I, Dendukuri N, Yao M, Dheda K, Denkinger CM, Schumacher SG, Steingart
KR. Xpert MTB/RIF Ultra and Xpert MTB/RIF assays for extrapulmonary tuberculosis and rifampicin resistance in adults. Cochrane Database of Systematic Reviews 2021; 1:CD012768.
Leeflang MM, Deeks JJ, Rutjes AW, Reitsma JB, Bossuyt PM. Bivariate meta-
analysis of predictive values of diagnostic tests can be an alternative to bivariate meta- analysis of sensitivity and specificity. Journal of Clinical Epidemiology 2012; 65: 1088–1097.
Puhan MA, Steurer J, Bachmann LM, ter Riet G. A randomized trial of ways to describe test
accuracy: the effect on physicians’ post-
test probability estimates. Annals of Internal
Medicine 2005; 143: 184–189.
Reid MC, Lane DA, Feinstein AR. Academic calculations versus clinical judgments: practic-
ing physicians’ use of quantitative measures of test accuracy. American Journal of Medicine 1998; 104: 374–380.
Steurer J, Fischer JE, Bachmann LM, Koller M, ter Riet G. Communicating accuracy of tests
to general practitioners: a controlled study. BMJ 2002; 324: 824–826.
Straus SE, Glasziou P, Richardson WS, Haynes RB. Evidence- based medicine: how to practice
and teach EBM. 5th ed. Philadelphia (PA): Elsevier; 2019.
Takwoingi Y, Partlett C, Riley RD, Hyde C, Deeks JJ. Methods and reporting of systematic
reviews of comparative accuracy were deficient: a methodological survey and proposed guidance. Journal of Clinical Epidemiology 2020; 121: 1–14.
Zhang S, Smailagic N, Hyde C, Noel- Storr AH, Takwoingi Y, McShane R, Feng J. (11)C- PIB- PET
for the early diagnosis of Alzheimer’s disease dementia and other dementias in people with mild cognitive impairment (MCI). Cochrane Database of Systematic Reviews 2014; 7:CD010386.
Zhelev Z, Garside R, Hyde C. A qualitative study into the difficulties experienced by
healthcare decision makers when reading a Cochrane diagnostic test accuracy review. Systematic Reviews 2013; 2: 32.
348