Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_2767_Библиотеки_им_академика_М_И_Перельмана
.pdf
11.7 Investigations ofsources ofheterogeneity
https://t.me/medicina_free
uncommon in systematic reviews because data needed for such comparisons are often
not reported in the included studies. In this Cochrane Review the subgroups were not
formally compared using meta- regression due to the small number of studies and different test positivity thresholds used in the studies.
Meta- analytical results from heterogeneity analyses are often presented in a table
and graphically in SROC plots displaying summary points or SROC curves for each category. The investigation of heterogeneity presented in Table11.7.a compared studies of
type 1 rapid diagnostic tests for Plasmodium falciparum malaria performed in Africa
and studies performed in Asia.
Care should be exercised in interpreting the results of heterogeneity investigations.
There are several points that should be considered.
First, subgroup findings are more credible if they have a scientific rationale. Ideally,
selection of characteristics for investigation should be motivated by biological, clinical
and methodological hypotheses supported by evidence from other sources. Subgroup
analyses based on characteristics that are implausible or irrelevant are unlikely to be
useful and should be avoided.
Second, exploratory heterogeneity investigations are less trustworthy than those that
were pre-
specified in the protocol. Review authors are expected to report whether heterogeneity investigations were pre- specified or data driven. Exploratory analyses are
often data driven and prompted by observations made in informal data analyses, rather
than by carefully crafted research hypotheses, supported by prior evidence. Prespecification of investigations of heterogeneity in systematic reviews is often difficult,
as review authors can be aware of some of the results in the included studies before
they start the review.
Third, the chance of obtaining spurious significant results increases with the number
of statistical comparisons that are undertaken. There is no formal rule on the maximum
number of investigations that can be undertaken. The total number of hypotheses
investigated must be kept in mind when interpreting the significance of the results.
Adjustments to P values using rules for multiple testing are not encouraged, as they
will be overly conservative due to the inevitable correlations between the factors
investigated.
Fourth, heterogeneity investigations based on small numbers of studies are unlikely
to produce useful findings and should be avoided. The statistical power of a comparison depends on the number of studies as well as on the precision of the test accuracy
estimates from each study. When the characteristic is unevenly distributed across
Table11.7.a Investigation of heterogeneity between studies of type 1 rapid diagnostic tests for
Plasmodium falciparum malaria
Continent Studies Patients Malaria
Africa 39 21,958 7445 94.0 (91.2 to 96.0) 93.1 (89.7 to 95.3) P = 0.03
Asia 24 15,810 4060 96.4 (93.7 to 97.9) 96.6 (94.0 to 98.1)
CI, confidence interval. P value obtained from a likelihood ratio test comparing models with and without the
covariate.
Source: Adapted from Table7in Abba 2011.
cases
Summary sensitivity %
(95% CI)
Summary specificity %
(95% CI)
P value
339

11 Presenting findings
https://t.me/medicina_free
groups, it is possible that important differences may be missed. Since most comparisons will not be made within study participants, or based on randomization, statistically significant differences may not reflect the actual cause of the difference.
Fifth, only characteristics that were reported at study level or in subgroups defined
within studies can be investigated. This limits the ability to detect relevant associations
with test accuracy on an individual level. For example, if accuracy varies with age but
study groups are similar in mean age, no association will be detected in a systematic
review that relies on aggregate data. This problem is known as aggregation or ecological bias: the failure of ecological (aggregate)- level associations to properly reflect
individual- level associations.
Sixth, it should be remembered that most subgroup comparisons are observational,
and suffer from the same limitations as other comparisons in observational research.
It may not be appropriate to make a causal interpretation of observed differences,
although causal effects may have the greatest relevance for clinical practice.
Confounding has to be considered: a difference between subgroups may be influenced
by other factors that are also associated with differences in test accuracy. For example,
if reference standards varied over time, but there were also changes over time in the
composition of the study groups, it will not be possible to identify which, if either, is the
cause of observed differences in test accuracy. Multivariable analysis, investigating
multiple sources of heterogeneity in parallel, is usually infeasible due to the limited
number of studies available (Takwoingi 2020).
Many review authors discover sooner or later that their plans for investigating heterogeneity are infeasible, either because of there being too few studies, or because studies
do not report the desired information. Cochrane Reviews of diagnostic test accuracy
contain a dedicated section to describe how the review differs from the protocol, where
these issues can be described.
11.8 Re- expressing summary estimates numerically
11.8.1 Frequencies
Sensitivity and specificity are typically presented as proportions or percentages, and so
are positive and negative predictive values. Presenting probabilities as frequencies has
been shown to help readers understand their relative magnitude (Hoffrage 1998, Evans
2000, Zhelev 2013), and this approach is encouraged both in the ‘Summary of main
results’ section of the review and in the ‘Summary of findings’ table.
A natural frequency description expresses a proportion as the number of individuals
out of a group (typically 10, 100 or 1000) in whom an event or outcome is observed. A
95% sensitivity can be presented as a group of 100 persons with the target condition, of
which 95 test positive.
As with conditional probabilities, one should be explicit about the group to which
natural frequencies refer. For example, they may refer to all those tested, those with or
without the target condition, or those with positive or with negative index test results.
Although sensitivity and specificity do not provide information on the absolute
impact of a test at a particular prevalence of the target condition, expressing them as
natural frequencies may help readers to interpret them. In addition, natural frequencies
340

11.8 Re- expressing summary estimates numerically
https://t.me/medicina_free
explicitly illustrate that sensitivity provides information on the false negatives and
specificity on the false positives.
For example:
●
For a test with a sensitivity of 90%: the index test will detect 90 out of every 100with
the target condition, but 10will be missed (i.e. will be false negatives).
●
For a test with a specificity of 80%: of every 100individuals without the target condition, 20will be wrongly diagnosed as having it (i.e. will be false positives).
11.8.2 Predictive values
There is a considerable body of empirical evidence demonstrating that sensitivity and
specificity can be difficult to interpret for many clinicians and patients (Steurer 2002,
Puhan 2005). Some researchers have argued that probabilities conditional on index test
results (predictive values) rather than actual disease status (sensitivity and specificity)
may be more intuitive to decision makers (Reid 1998).
Historically the use of predictive values has been discouraged, because sensitivity
and specificity were assumed to be statistically independent of the proportion of participants with the target condition in the study group. If there are more study participants with the target condition, and sensitivity and specificity do not change, positive
predictive values will be higher and negative predictive values lower. We now know that
this is a simplification, as we are becoming increasingly aware of the variation in sensitivity and specificity caused by differences in the setting, previous testing and severity
of disease (spectrum of disease) (Leeflang 2012). Review authors should therefore be
mindful of the transferability of any accuracy measures, regardless of the type of summary statistic used.
Meta- analysis of predictive values is possible and sometimes necessary (Leeflang
2012), but it is not recommended to perform two meta- analyses in parallel: one of sensitivity and specificity and a second, separate meta- analysis of positive and negative
predictive values. If review authors wish to use predictive values as a means of expressing test accuracy from a meta- analysis producing summary estimates of sensitivity and
specificity, they should compute predictive values based on these estimates for a representative proportion of those with the target condition.
Predictive values are most simply obtained from summary estimates of sensitivity
and specificity by creating an illustrative 2×2 table and computing predictive values
directly (the simple equations to do this are in Chapter4). This exercise can be done on
paper, using a spreadsheet or a calculator tool (see Figure11.8.a).
To compute predictive values manually or using a calculator, one needs a fictional
group size (say, 1000), the proportion with the target condition (‘prevalence’ in the calculator), and the summary estimates of sensitivity and specificity of the test– i.e. the
boxes in green in Figure11.8.a. For example, a test that has sensitivity of 0.9 and specificity of 0.8 yields the values in Figure11.8.a if 0.25 (25%) have the target condition. This
computes the positive predictive value to be 0.60 and the negative predictive value to
be 0.96. The same computations can be done using Bayes equation, as shown in
Box11.8.a.
Natural frequencies can be used to describe the absolute impact of a test in a population with a given prevalence (25% in Figure11.8.a and Box11.8.a).
341

Figure11.8.a Using a calculator to convert sensitivity and specificity to positive and negative predictive values at a prevalence of25%. D+,
https://t.me/medicina_free
disease positive; D–, disease negative; FP, false positive; FN, false negative; LR+, positive likelihood ratio; LR–, negative likelihood ratio; NPV, negative
predictive value; PPV, positive predictive value; TN, true negative; TP, true positive.

11.8 Re- expressing summary estimates numerically
sensitivity prevalence
https://t.me/medicina_free
Box 11.8.a Calculation ofpredictive values using Bayes equation andestimates
ofsensitivity, specificity andprevalence
sensitivity prevalence 1 specificity 1 prevalence
0.9 0.25
0.9 0.25 1 0.8 1 0.25
specificity 1 prevalence
1 sensitivity prevalence specificity 1 prevalence
0.8 1 0.25
1 0.9 0.25 0.8 1 0.25
NPV, negative predictive value; PPV, positive predictive value.
●
For a test with a positive predictive value of 60%: 60 out of every 100 positive index
test results will have the target condition, but 40will not (i.e. will be false positives). In
a group in which 25% have the target condition, this will result in 150 false positive
test results for every 1000 people tested.
●
For a test with a negative predictive value of 96%: 96 out of every 100negative index
test results will not have the target condition, but 4will (i.e. will be false negatives). In
a group in which 25% have the target condition, this will result in 25 false negative
test results for every 1000 people tested.
Although predictive values may be intuitive summary metrics, choosing the proportion of those with the target condition (prevalence) is not straightforward. The term
‘prevalence’ is often used to refer to the proportion of those with the target condition,
for lack of a better term, but the proportion of those with the target condition in the
population being tested will rarely be equal to the prevalence of the target condition in
the population. In testing for SARS- CoV- 2 infection, for example, the proportion with
the SARS- CoV- 2 virus in symptomatic persons undergoing testing will be substantially
higher than the population prevalence. Similarly, the proportion with the virus in
asymptomatic contacts of persons who recently tested positive will also be higher than
the population prevalence, but lower than the proportion in symptomatic persons.
To select a proportion for the conversion to predictive values, the median value for
the proportion of those with the target condition might be used, if that median is calculated from studies that relied on consecutive or random sampling of participants in the
intended- use setting. Because of their design, studies that separately recruited participants with the target condition and healthy controls should be excluded from calculating the median proportion. Alternatively, review authors may consider computing
predictive values across a range of plausible values for the intended- use setting. In
some circumstances, estimates of disease prevalence may be more reliably obtained
343

11 Presenting findings
https://t.me/medicina_free
from other data sources, such as disease registries that correspond to the settings in
which the studies in the meta- analysis were conducted.
The proportion of those with the target condition selected for presenting predictive
values usually falls in the range of corresponding proportions observed in the studies
included in the meta- analysis. The selection of a value outside that range would be an
extrapolation, and should be done with caution.
11.8.3 Likelihood ratios
The use of likelihood ratios to express test performance (see Chapter4) has been promoted by some as a metric that facilitates Bayesian probability updating: the calculation of postevidence that likelihood ratios improve diagnostic decision- making in groups of clinicians is lacking.
Parallel or separate meta- analysis of likelihood ratios is not recommended (see
Chapter 9, Section 9.3). Summary estimates of the positive and negative likelihood
ratios can be calculated from the summary estimates of sensitivity and specificity.
These summary positive and negative likelihood ratios (and their 95% confidence intervals) can be obtained as additional estimates, as indicated in Chapter9, Section9.3,
and illustrated in the software code in the appendices of Chapter10.
Ideally, the confidence intervals for likelihood ratios should be used to calculate confidence intervals for predictive values because, unlike sensitivity and specificity, the
positive likelihood ratio jointly considers the number of true and false positives and the
negative likelihood ratio jointly considers the number of true and false negatives. Using
likelihood ratio outputs from SAS and Stata, the lower and upper confidence limits of
the positive likelihood ratio can be converted into lower and upper confidence limits of
the positive predictive value, at a stated proportion with the target condition. Likewise,
confidence limits of the negative likelihood ratio can be converted into confidence limits of the negative predictive value.
The simplest approach for deriving predictive values from likelihood ratios is to use a
calculator tool, similar to the one for deriving point estimates of the predictive values
(see Figure11.8.a). Alternatively, predictive values can be calculated using the equations in Box11.8.b.
Only the uncertainty in the summary estimates of sensitivity and specificity as expressions of test accuracy is captured in these 95% confidence intervals, not the uncertainty
in the proportion with the target condition. This uncertainty could be explored by computing point estimates and 95% confidence intervals for predictive values across a
range of plausible values for the proportion with the target condition (see Chapter12).
test probabilities for specific pre- test probabilities (Straus 2019). Actual
11.9 Presenting findings when meta- analysis cannot
beperformed
Meta- analysis is not always possible due to few studies, convergence issues with the
analysis (see Chapter10, Section10.6), clinical heterogeneity or other factors. Such situations are common. Of the 135 Cochrane Reviews of diagnostic test accuracy published up to 31July 2020, 32 (24%) did not include a meta- analysis. In the 32 reviews,
344

11.9 Presenting findings when meta- analysis cannot beperformed
https://t.me/medicina_free
Box 11.8.b Calculation ofpredictive values using post- test odds andlikelihood ratios
Pre- test odds = pre- test probability/(1 − pre-test probability)
Post-
test odds of disease given positive test result = pre- test odds of disease × LR+
test odds of disease given negative test result = pre- test odds of disease × LR–
PostPost-
test probability = Post- test odds/(1 + post- test odds)
test probability of disease given a positive test result = PPV
PostPost-
test probability of disease given a negative test result = 1– NPV
Repeating the example shown in Figure11.8.a, a test that has a positive likelihood ratio
(LR+) of 4.5 and negative likelihood (LR–) of 0.125will give a positive predictive value (PPV)
of 0.60 and negative predictive value (NPV) of 0.96 using the previous equations as follows.
test odds = 0.25/(1– 0.25) = 1/3
PrePost- test odds of disease given positive test result = (1/3) × 4.5 = 1.5
Post- test odds of disease given negative test result = (1/3) × 0.125 = 0.0417
test probability of disease given a positive test result = 1.5/(1 + 1.5) = 0.60
PostPost- test probability of disease given a negative test result = 0.0417/(1 + 0.0417) = 0.04,
therefore NPV = 1– 0.04 = 0.96
the number of included studies ranged between 1 and 33; the median was 4 (interquartile range 3 to 9). A frequent reason for not performing meta- analysis was considerable between- study variation in clinical and methodological characteristics. For
example, Chan (2019) did not perform meta- analysis due to risk- of- bias concerns and
heterogeneity.
Tables, forest plots and/or SROC plots are valuable tools for presenting findings in
reviews without meta- analysis. Plots provide a visual summary across studies and are
more readily accessible to readers than listing many results from individual studies in
the text. When there are few included studies, an SROC plot showing point estimates of
sensitivity and specificity and their confidence intervals can be an effective display if
study points and confidence interval lines do not overlap considerably (see Figure11.9.a).
Review authors may also want to present and describe the results of individual studies. For example, Crawford (2016) included four studies and described the results of each
study in the text and outlined key findings from the studies in the ‘Summary of findings’
table. Describing the results of individual studies will be impractical when there are
many studies. Instead, review authors could provide a narrative summary that includes
the ranges of reported estimates of sensitivity and specificity and the body of evidence
(number of studies, number of participants, number with the target condition).
Presenting findings in the absence of a meta- analysis becomes even more challenging when there are multiple index tests, target conditions or reference standards. In
such situations, review authors should carefully consider how best to present the findings, using text, tables and figures as appropriate.
For example, a Cochrane Review reported that the 33included studies evaluated a
myriad of index tests against different reference standards (Hanchard 2013). In addition, there were five categories of the target condition. Altogether, there were 170 combinations of target condition and index test. None of these combinations was assessed
in a similar manner by more than two studies. The review authors concluded that
345

11 Presenting findings
1
Sensitivity (95% CI)
0
https://t.me/medicina_free
0.9
0.8
0.7
0.6
0.5
0.4
0.3
0.2
0.1
0
1 0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1
Legend
POD:3 DFA > 600 IU/L
POD:3 to 5 DFA > 3 times serum amylase
POD:4 DFA > 647 U/L
Specificity (95% CI)
POD:5 DFA > 3 times serum amylase
POD:5 DFA > 4000 U/L
346
Figure11.9.a SROC plot with point estimates of sensitivity and specificity and 95% confidence
intervals. The Cochrane Review assessed amylase in drain fluid for detecting pancreatic leak in
post- pancreatic resection. In the five included studies, drain fluid amylase was measured on different
days and at different thresholds. A meta- analysis was not performed. The numbers following POD
(postoperative day) indicate the number of the postoperative day. The numbers or text following DFA
(drain fluid amylase) indicate the threshold. Source: Adapted from Davidson 2017
meta- analysis was inappropriate because of substantial clinical heterogeneity and a
limited number of studies for each combination. To keep the number of forest plots to
a minimum and to enhance readability, the review authors presented estimates of sensitivity and specificity on forest plots grouped according to target condition. A narrative
summary with a similar structure accompanied the forest plots.
11.10 Chapter information
Authors: Jonathan J. Deeks (Institute of Applied Health Research, University of
Birmingham, UK), Patrick M. Bossuyt (Department of Epidemiology and Data Science,
University of Amsterdam, The Netherlands), Mariska M. Leeflang (Department of
Epidemiology and Data Science, University of Amsterdam, The Netherlands), Yemisi
Takwoingi (Institute of Applied Health Research, University of Birmingham, UK).

11.11 References
https://t.me/medicina_free
Sources of support: Jonathan J. Deeks is a UK National Institute for Health Research
(NIHR) Senior Investigator Emeritus. Yemisi Takwoingi is funded by a UK National
Institute for Health Research (NIHR) Postdoctoral Fellowship. Jonathan J. Deeks and
Yemisi Takwoingi are supported by the NIHR Birmingham Biomedical Research Centre
at the University Hospitals Birmingham NHS Foundation Trust and the University of
Birmingham. The views expressed are those of the authors and not necessarily those of
the NHS, the NIHR or the Department of Health and Social Care. The authors declare no
sources of support for writing this chapter.
Declarations of interest: Jonathan J. Deeks, Yemisi Takwoingi and Mariska M. Leeflang
are members of Cochrane’s Diagnostic Test Accuracy Editorial Team. Yemisi Takwoingi
and Mariska M. Leeflang are co-
convenors of the Cochrane Screening and Diagnostic
Tests Methods Group. The authors declare no other potential conflicts of interest relevant to the topic of this chapter.
Acknowledgements: The authors would like to thank Karen R. Steingart, Marta Roque
and Daniël Korevaar for helpful peer review comments.
11.11 References
Abba K, Deeks JJ, Olliaro P, Naing CM, Jackson SM, Takwoingi Y, Donegan S, Garner P. Rapid
diagnostic tests for diagnosing uncomplicated P. falciparum malaria in endemic
countries. Cochrane Database of Systematic Reviews 2011; 7: CD008122.
Alldred SK, Takwoingi Y, Guo B, Pennant M, Deeks JJ, Neilson JP, Alfirevic Z. First trimester
ultrasound tests alone or in combination with first trimester serum tests for Down’s
syndrome screening. Cochrane Database of Systematic Reviews 2017; 3: CD012600.
Best LM, Takwoingi Y, Siddique S, Selladurai A, Gandhi A, Low B, Yaghoobi M, Gurusamy KS.
Non- invasive diagnostic tests for Helicobacter pylori infection. Cochrane Database of
Systematic Reviews 2018; 3: CD012080.
Chan CC, Fage BA, Burton JK, Smailagic N, Gill SS, Herrmann N, Nikolaou V, Quinn TJ,
Storr AH, Seitz DP. Mini- Cog for the diagnosis of Alzheimer’s disease dementia and
Noelother dementias within a secondary care setting. Cochrane Database of Systematic
Reviews 2019; 9: CD011414.
Chan KK, Joo DA, McRae AD, Takwoingi Y, Premji ZA, Lang E, Wakai A. Chest ultrasonography
versus supine chest radiography for diagnosis of pneumothorax in trauma patients in the
emergency department. Cochrane Database of Systematic Reviews 2020; 7: CD013031.
Crawford F, Andras A, Welch K, Sheares K, Keeling D, Chappell FM. D-
the diagnosis of pulmonary embolism. Cochrane Database of Systematic Reviews 2016;
8:CD010864.
Davidson TB, Yaghoobi M, Davidson BR, Gurusamy KS. Amylase in drain fluid for the
diagnosis of pancreatic leak in post- pancreatic resection. Cochrane Database of
Systematic Reviews 2017; 4: CD012009.
Deeks JJ, Higgins JPT, Altman DG. Chapter10: Analysing data and undertaking meta-
analyses. In: Higgins JPT, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, Welch VA,
editors. Cochrane Handbook for Systematic Reviews of Interventions version 6.0.
Cochrane, 2019.
dimer test for excluding
347

11 Presenting findings
https://t.me/medicina_free
Evans JS, Handley SJ, Perham N, Over DE, Thompson VA. Frequency versus probability
formats in statistical word problems. Cognition 2000; 77: 197–213.
Hanchard NC, Lenza M, Handoll HH, Takwoingi Y. Physical tests for shoulder impingements
and local lesions of bursa, tendon or labrum that may accompany impingement.
Cochrane Database of Systematic Reviews 2013; 4: CD007427.
Hoffrage U, Gigerenzer G. Using natural frequencies to improve diagnostic inferences.
Academic Medicine 1998; 73: 538–540.
Kohli M, Schiller I, Dendukuri N, Yao M, Dheda K, Denkinger CM, Schumacher SG, Steingart
KR. Xpert MTB/RIF Ultra and Xpert MTB/RIF assays for extrapulmonary tuberculosis and
rifampicin resistance in adults. Cochrane Database of Systematic Reviews 2021;
1:CD012768.
Leeflang MM, Deeks JJ, Rutjes AW, Reitsma JB, Bossuyt PM. Bivariate meta-
analysis of
predictive values of diagnostic tests can be an alternative to bivariate meta- analysis of
sensitivity and specificity. Journal of Clinical Epidemiology 2012; 65: 1088–1097.
Puhan MA, Steurer J, Bachmann LM, ter Riet G. A randomized trial of ways to describe test
accuracy: the effect on physicians’ post-
test probability estimates. Annals of Internal
Medicine 2005; 143: 184–189.
Reid MC, Lane DA, Feinstein AR. Academic calculations versus clinical judgments: practic-
ing physicians’ use of quantitative measures of test accuracy. American Journal of
Medicine 1998; 104: 374–380.
Steurer J, Fischer JE, Bachmann LM, Koller M, ter Riet G. Communicating accuracy of tests
to general practitioners: a controlled study. BMJ 2002; 324: 824–826.
Straus SE, Glasziou P, Richardson WS, Haynes RB. Evidence- based medicine: how to practice
and teach EBM. 5th ed. Philadelphia (PA): Elsevier; 2019.
Takwoingi Y, Partlett C, Riley RD, Hyde C, Deeks JJ. Methods and reporting of systematic
reviews of comparative accuracy were deficient: a methodological survey and proposed
guidance. Journal of Clinical Epidemiology 2020; 121: 1–14.
Zhang S, Smailagic N, Hyde C, Noel- Storr AH, Takwoingi Y, McShane R, Feng J. (11)C- PIB- PET
for the early diagnosis of Alzheimer’s disease dementia and other dementias in people
with mild cognitive impairment (MCI). Cochrane Database of Systematic Reviews 2014;
7:CD010386.
Zhelev Z, Garside R, Hyde C. A qualitative study into the difficulties experienced by
healthcare decision makers when reading a Cochrane diagnostic test accuracy review.
Systematic Reviews 2013; 2: 32.
348
Соседние файлы в папке Библиотека им академика М.И. Перельмана
