Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_2767_Библиотеки_им_академика_М_И_Перельмана
.pdf
11.5 Individual andsummary estimates oftest accuracy
https://t.me/medicina_free
Details for the included studies may best be reported in a ‘Characteristics of included
studies’ table. Such a table may present at a glance key details for each study. Chapter7
details the characteristics that should be extracted from each study to populate a
‘Characteristics of included studies’ table. Careful planning of the data extraction form
can help ensure that items are recorded in a manner that allows their direct incorporation in the table. This saves time, particularly in reviews with substantial numbers of
included studies.
11.4 Methodological quality ofincluded studies
This section describes the methodological quality of studies included in the systematic
review. For most systematic reviews of test accuracy, the QUADAS- 2 tool will be used to
evaluate methodological quality in terms of risk of bias and concerns regarding applicability (see Chapter8).
The results of the assessment are usually presented in a table or figure. A figure may
show the results for individual studies or across studies. For reviews with multiple target conditions or index tests, figures may also be presented separately for each target
condition or test. If there are several tables or figures, review authors should carefully
consider which ones to include in the main text and which ones to include in an
appendix.
An informative description of the assessments of risk of bias would indicate, for each
domain, whether studies were considered at risk of producing biased results, and why
they were classified as such. Readers would also like to know to what extent an
included study helped to answer the review question, i.e. whether there were concerns
regarding applicability. Here also, it is informative to indicate whether such concerns
arose, and why.
If studies were potentially eligible but excluded after an evaluation of their risk of bias
or applicability, they should not be reported here. The characteristics of studies
excluded because of risk of bias or concerns regarding applicability should be reported
elsewhere, in a separate table or supplementary file.
If the review addresses comparisons of test accuracy between two or more index
tests, the results of the assessment with the QUADAS- C tool will be presented next. The
QUADAS- C results are typically specific to a comparison of two tests within the review.
If there are two or more such comparisons, separate QUADAS- C assessments need to be
reported.
11.5 Individual andsummary estimates oftest accuracy
For each objective, the review should present the relevant evidence and the findings,
to enable conclusions to be drawn. For primary objectives this typically requires
presentation of results from individual studies as well as from meta- analyses, if such
analyses were performed. Meta- analyses may produce summary estimates of sensitivity and specificity or SROC curves (see Chapter9, Section9.2.2 and Section9.3). If no
meta- analysis could be performed, a different approach to summarizing the results of
individual studies will be adopted (see Section11.9).
329

11 Presenting findings
https://t.me/medicina_free
11.5.1 Presenting results fromincluded studies
It will be informative to present test accuracy results from the individual studies. Such
overviews typically include the study identifier, year of publication, total number of participants, number of participants with the target condition, number of true and false positives, number of true and false negatives, estimates of test accuracy, as defined in the
protocol, and expressions of statistical imprecision (typically as 95% confidence intervals). These overviews may be presented in tables or, more succinctly, on coupled forest
plots of sensitivity and specificity (see Chapter9, Section9.2.1 and example Figure9.2.a).
For pairwise comparisons of accuracy, individual studies will present multiple pairs of
estimates of sensitivity and specificity. In that case, it is helpful to structure the table or
graph in such a way that the link between these pairs is obvious. This can be done by
presenting a linked SROC plot showing the multiple estimates from each study (one for
each test), connected by a line (see Chapter9, Figure9.4.g, for an example).
11.5.2 Presenting summary estimates ofsensitivity andspecificity
Review authors should carefully consider the summary measures that are relevant and
appropriate to present for a given analysis, rather than compute and report all possible
measures of test accuracy. Many systematic reviews of test accuracy present summary
estimates of sensitivity and specificity.
Other summary estimates, like positive and negative predictive values, may also be
produced. The selection of appropriate summary estimates will also depend on the
design of the primary studies. For example, if one reference standard is always used in
index test positives and a different one in index test negatives, it will be more meaningful to present summary estimates of the negative and predictive values. The use of different reference standards is known to introduce bias in estimates of sensitivity and
specificity.
A review may include multiple meta- analyses and corresponding summary estimates.
These estimates can be summarized in one or more tables as appropriate. The table(s)
should include the number of studies, participants and participants with and without
the target condition on which each analysis is based (see an example in Table11.5.a).
When summary estimates of sensitivity and specificity are reported as an output of
meta- analysis, they need to be interpreted as expressions of a central tendency of estimates across the included studies. There may be considerable uncertainty in these
summary estimates and substantial variability in accuracy across studies.
11.5.3 Presenting SROC curves
Including a 2×2 table of the number of true positives, false positives, false negatives and
true negatives from each study regardless of threshold value allows estimation of an
SROC curve. If an SROC curve is assumed to be symmetrical (i.e. the shape parameter in
the HSROC model is assumed to be zero), then accuracy does not depend on threshold.
In that case, the SROC curve can be described by a constant diagnostic odds ratio (DOR)
(see Chapter4, Section 4.5.7). A DOR of 1 represents an uninformative test (the upward
diagonal in Figure11.5.a). SROC curves closer to the top left- hand corner of the SROC
plot represent tests with a higher DOR, hence more discriminatory power.
330

11.5 Individual andsummary estimates oftest accuracy
DOR = 361
Sensitivity
0
Specificity
https://t.me/medicina_free
Table11.5.a Accuracy ofXpert MTB/RIF fordetection ofextrapulmonary tuberculosis
Type of
specimen
Reference
standard
Number of
studies
(participants)
Number
with TB
(%)
Summary sensitivity %
(95% credible interval)
Summary specificity %
(95% credible interval)
CSF Culture 30 (3395) 571 (16.8) 71.1 (62.8 to 79.1) 96.9 (95.4 to 98.0)
CSF Composite 14 (2203) 862 (39.1) 42.3 (32.1 to 52.8) 99.8 (99.3 to 100.0)
Pleural fluid Culture 25 (3065) 644 (21.0) 49.5 (39.8 to 59.9) 98.9 (97.6 to 99.7)
Pleural fluid Composite 10 (1024) 616 (60.1) 18.9 (11.5 to 27.9) 99.3 (98.1 to 99.8)
Lymph node
Culture 14 (1588) 627 (39.5) 88.9 (82.7 to 93.6) 86.2 (78.0 to 92.3)
aspirate
Lymph node
Composite 4 (679) 377 (55.5) 81.6 (61.9 to 93.3) 96.4 (91.3 to 98.6)
aspirate
CSF, cerebrospinal fluid; TB, tuberculosis.
Source: Adapted from Table2in Kohli 2021.
When interpreting DORs, review authors should note that the same DOR may be
achieved by different combinations of sensitivity and specificity (see Chapter 4,
Section4.5.7). In addition, single measures of test accuracy, like the DOR, are often not
clinically informative, because of a lack of information on errors in those with the target
condition (false negatives) and those without the condition (false positives). The relative
DOR = 81
0.0 0.2 0.4 0.6 0.8 1. 0
Figure11.5.a Symmetrical summary receiver operating characteristic (SROC) curves and diagnostic
odds ratios (DORs)
DOR = 16
DOR = 5
DOR = 2
DOR = 1 (uninformative test)
Line of symmetry
0.20.4 0.6 0.81. 0
0.
331

11 Presenting findings
https://t.me/medicina_free
magnitude of such errors is essential for judging the extent and likely impact of consequences of positive and negative test results.
DORs are probably most useful in meta- analysis when making comparisons between
tests or between subgroups (see Chapter9, Section 9.4.6.4, and Section 11.6.2). For
these comparisons, if two curves are symmetrical or have the same shape, then the
relative accuracy of the two curves can be summarized using the ratio of the DORs or
relative DOR (RDOR).
It may be challenging to describe an SROC curve that is not symmetrical, either verbally or numerically. As explained in Chapter10 (Section10.3.1), the anticipated sensitivity at a given value of specificity (or the other way round) can be computed from the
SROC curve to illustrate test performance in such cases. If review authors choose to
report such sensitivity and specificity pairs from an SROC curve, then the most informative estimates will be points on the curve that lie within the range of the sensitivity and
specificity estimates reported by the included studies. For example, these could be values that represent the median and interquartile range from the studies included in the
meta-
analysis (see Chapter10, Box 10.3.b).
Alternatively, if minimizing false positives– therefore maximizing specificity – in a
particular testing context is desired, the sensitivity of the test could be reported at the
minimally acceptable specificity (for example, a specificity of 95%).
Whatever the approach taken to select sensitivity and specificity pairs, the values and
corresponding estimates can be presented in a table, as shown in the example in
Table11.5.b. Likelihood ratios can also be estimated at such fixed values.
11.5.4 Describing uncertainty insummary statistics
The summary sensitivity and specificity are estimates based on the included studies
and there will always be statistical uncertainty about their true value. Review authors
should express the degree of statistical uncertainty associated with summary estimates
of test accuracy, irrespective of the metric used. Confidence intervals or credible intervals should therefore be reported alongside the point estimates of summary sensitivity
and specificity in text and tables. For summary points on SROC plots, confidence regions
should also be presented (see Chapter9, Figure9.4.a).
Table11.5.b Sensitivity andlikelihood ratios for11C- PIB- PET at fixed values ofspecificity
forAlzheimer’s dementia
Statistic Fixed value
of specificity %
Lower quartile 56 96 (88 to 99) 2.19 (2.09 to 2.29) 0.07 (0.02 to 0.23)
Median 58 96 (87 to 99) 2.29 (2.17 to 2.41) 0.07 (0.02 to 0.24)
Upper quartile 81 89 (68 to 97) 4.66 (4.03 to 5.39) 0.14 (0.05 to 0.44)
CI, confidence interval. Nine studies (n: 112with dementia; 162without dementia)
Source: Adapted from Table4in Zhang 2014.
332
Estimated
sensitivity %
(95% CI)
Positive
likelihood ratio
(95% CI)
Negative
likelihood ratio
(95% CI)

11.6 Comparisons oftest accuracy
https://t.me/medicina_free
11.5.5 Describing heterogeneity insummary statistics
Statistical heterogeneity (referred to simply as heterogeneity) exists whenever estimates of test accuracy vary between studies, more than would be expected from withinstudy sampling error (chance) alone. This is extremely common in systematic reviews
of test accuracy. Prediction intervals and prediction regions give an indication of heterogeneity. Prediction regions plotted around summary points in ROC space represent
the area where the sensitivity and specificity estimates from a future test accuracy
study are expected to lie (see Chapter9, Section9.3.2, Section9.4.1 and Section9.4.2).
Systematic reviews of test accuracy studies can include prediction regions with coverage probabilities of 50%, 90% or 95%. The 95% prediction regions often cover large
areas of ROC space because of the presence of substantial heterogeneity common in
many systematic reviews of test accuracy. While the confidence region depicts uncertainty in the summary estimates of sensitivity and specificity due to withinability, the prediction region depicts the uncertainty due to both within- and
between- study variability (see Chapter9, Figure9.4.a). As such, when heterogeneity is
high, the 95% prediction region will be much larger than the 95% confidence region.
As stated in Chapter9, estimation of a prediction interval or region relies on the assumption of normal distributions for the effects across studies. This may be very problematic
when the number of studies is small and can lead to spuriously large (or small) regions
(Deeks 2019). In addition, if the variance parameters and the correlation (or covariance)
between logit sensitivity and specificity in a bivariate model cannot be reliably estimated,
a prediction region will be misleading. Therefore, when there are few studies (e.g. fewer
than 10) in the meta- analysis of a single test, review authors should carefully consider the
appropriateness of presenting a prediction region around the summary point.
study vari-
11.6 Comparisons oftest accuracy
Review authors should consider two issues in reviews that compare the accuracy of multiple tests: (1) the statistical measures that can be used; and (2) the strength of the evidence for the comparison. The strength primarily relates to whether the meta- analysis is
based on within- study (direct) or between- study (indirect) comparisons of tests (see
Chapter9, Section9.1.4.3). This will be considered in Chapter12 (Section12.1.3.5 and
Section12.6.2.4). The appropriate statistical measures are not affected by this issue.
When summarizing findings from a comparison of two tests, review authors should
focus on describing (1) the magnitude and direction of the difference between tests;
(2)the uncertainty in the estimates; and (3) the degree of heterogeneity. Presentation of
test comparisons can be facilitated by summaries of test accuracy in ROC space, which
allow readers to compare test performance in one SROC plot. This may be in the form of
summary estimates of sensitivity and specificity or SROC curves (shape and relative
position). In addition, within- study test comparisons can be annotated to distinguish
them from between- study comparisons.
11.6.1 Comparing tests using summary points
The magnitude and direction of the difference between tests can be summarized
byreporting point estimates of the summary sensitivity and specificity for each test
333

11 Presenting findings
https://t.me/medicina_free
Table11.6.a Accuracy ofchest ultrasonography andchest radiography fordiagnosis
ofpneumothorax
Test Studies Participants (with
pneumothorax)
CUS 9 1271 (410) 0.91 (0.85 to 0.94) 0.99 (0.97 to 1.00)
CXR 9 1271 (410) 0.47 (0.31 to 0.63) 1.00 (0.97 to 1.00)
Absolute
difference
CUS, chest ultrasonography; CXR, chest radiography. The P values are from Wald tests.
Source: Adapted from Table2in Chan 2020.
Summary
sensitivity (95% CI)
0.44 (0.27 to 0.61)
P < 0.001
Summary
specificity (95% CI)
−0.007 (−0.018 to 0.005)
P = 0.26
and by measures of absolute or relative differences in sensitivity and specificity
(seeTable11.6.a). Focusing on the size and the uncertainty in the estimated difference
in summary sensitivity and specificity between tests illustrates the potential impact of
using one test or the other.
For the example in Table 11.6.a, the summary sensitivities of 0.91 for chest ultrasonography and 0.47 for chest radiography indicate that chest ultrasonography can be
expected to correctly detect 44more patients out of every 100 with pneumothorax,
compared to chest radiography. The 95% confidence interval translates into a difference that lies between 27 and 61more patients out of every 100with pneumothorax.
Asimilar approach can be used if predictive values are being compared.
11.6.2 Comparing tests using SROC curves
When SROC curves of two or more tests are compared, it matters whether or not the
SROC curves for the tests have the same shape. If they do, the value of the RDOR will be
constant all the way along the curve. In those situations, the RDOR is a valid reflection
of the difference in accuracy between tests. In that case, it will be informative to report
the summary DOR for each test and estimates of the RDOR, with expressions of statistical uncertainty and heterogeneity.
Table11.6.b presents the results of an indirect comparison of the accuracy of urea
breath test, serology and stool antigen test for detecting a Helicobacter pylori infection.
It is based on an HSROC model with symmetrical shape for the curves (Best 2018). The
RDOR of 3.22 (95% CI 1.24 to 8.37, P = 0.017) indicates that the DOR for the urea breath 13C test is about three times that of serology, and that we are 95% confident that this
ratio lies between 1.24 and 8.37 times higher.
A higher DOR means that, for any selected positivity threshold for the inferior test, we
can find a positivity threshold for the superior test that produces higher sensitivity for
the same specificity, higher specificity for the same sensitivity, or both higher sensitivity
and specificity. We cannot, however, identify the positivity thresholds on the SROC
curve that correspond to specific sensitivity–specificity pairs.
Where tests have been compared using SROC curves, it may nevertheless be informative to report selected sensitivity–specificity pairs on each of the curves, to facilitate test
comparisons. For example, the sensitivity of each test at the same fixed specificity could
be reported, or the other way round. Review authors should be cautious when choosing
334

Table11.6.b Comparison of the accuracy of non- invasive tests for Helicobacter pylori infection
https://t.me/medicina_free
Index test Studies; participants
(with H. pylori)
DOR (95% CI) Relative diagnostic odds ratios (95% CI), P value
13
Urea breath test-
C Urea breath test- 14C Serology
Urea breath test- 13C 34; 3139 (1526) 153 (73.7 to 316)
14
Urea breath test-
C 21; 1810 (1018) 105 (74.0 to 150) 1.45 (0.65 to 3.26)
P = 0.36
Serology 34; 4242 (2477) 47.4 (25.5 to 88.1) 3.22 (1.24 to 8.37)
P = 0.017
Stool antigen test 29; 2988 (1311) 45.1 (24.2 to 84.1) 3.39 (1.30 to 8.83)
P = 0.013
CI, confidence interval. The indirect comparison included all studies that evaluated at least one of the four tests, i.e. all available data. The RDOR is the diagnostic odds
ratio (DOR) of the test in the column relative to the DOR of the test in the row. If the RDOR is greater than one, then the test in the column is more accurate than the test
in the row. The P values are from Wald tests.
Source: Adapted from Table2in Best 2018.
2.22 (1.09 to 4.51)
P = 0.028
2.33 (1.14 to 4.76)
P = 0.020
1.05 (0.44 to 2.53)
P = 0.91

11 Presenting findings
https://t.me/medicina_free
pairs for comparison. To be valid, the pairs should be based on the results in the review:
they should lie within the range of observed sensitivity and specificity estimates from
the included studies.
If the SROC curves for two or more tests differ in shape, the ratio of the corresponding
DOR will not be constant along the entire length of the curve. In these situations, comparisons are challenging and interpretation of meta- analysis results needs to be done carefully,
similarly considering the observed range of sensitivity and specificity estimates in individual studies. Here also, selecting pairs on the fitted curves may assist in interpretation,
provided that these lie within the range of estimates reported by the included studies.
11.6.3 Interpretation ofconfidence intervals fordifferences intest accuracy
It has been shown that some readers wrongly apply heuristics about ratio measures, as
used in systematic reviews of interventions (Zhelev 2013). In comparing interventions,
they know that a confidence interval including unity indicates no significant difference
in outcomes between the interventions. Readers should not interpret confidence intervals for sensitivity and specificity in a similar way. A confidence interval for sensitivity
and specificity that includes unity (1.00, or 100%) does not reflect poor test performance but rather the opposite: the possibility of perfect performance. Review authors
should therefore consider supplementing numerical presentation of statistical uncertainty in confidence intervals with verbal explanations. A natural frequency presentation format may additionally help (see Box11.6.a and Section11.8), particularly if a
systematic review of test accuracy includes both estimates of the accuracy of single
tests and a comparison of test accuracy.
When interpreting confidence intervals associated with comparisons, confidence
intervals that do not overlap can be assumed to represent statistically significant differences. Confidence intervals that overlap may or may not reflect statistically significant
differences; P values will be required to draw conclusions about statistical significance.
Interpretation of a test comparison is illustrated in Box11.6.a.
11.7 Investigations ofsources ofheterogeneity
Investigations of heterogeneity aim to assess whether test accuracy tends to vary with
identifiable features, such as the characteristics of the participants, settings, tests, reference standards or others. As explained in Chapter 9 and Chapter 10, hierarchical
meta- regression models can estimate differences in accuracy between subgroups
(orthe association of accuracy with a continuous measure) and formally test differences
and associations for statistical significance.
Comparisons of subgroups based solely on studies that report data for subgroups
within studies provide more valid evidence of sources of heterogeneity than comparisons between studies. For example, a Cochrane Review on Down syndrome screening
explored the effect of advanced maternal age (< 35 years versus ≥ 35 years) on test
performance. Of the 69 studies for one of the index tests assessed in the review, 5 provided data to compare the performance of the test between women younger than
35years and those 35 years or more within the same study (Alldred 2017). The 5 studies
all showed a higher sensitivity and higher false positive fraction for the ≥ 35 years
compared to the < 35 years subgroup, as shown in Figure11.7.a. Such analyses are
336

Box 11.6.a Interpretation ofconfidence intervals andP values forcomparisons oftest performance
https://t.me/medicina_free
A Cochrane Review compared the accuracy of rapid diagnostic tests (RDTs) for detecting Plasmodium falciparum malaria (Abba 2011). RDTs
use different types of antibody or antibody combinations to detect Plasmodium antigens. Two RDT types, type 1 and type 4, were compared
in one of the meta-
RDT type Summary sensitivity % (95% CI) Summary specificity % (95% CI)
Type 1: HRP- 2 antibody- based tests 94.8 (93.0 to 96.1) 95.2 (93.2 to 96.7)
Type 4: pLDH antibody- based tests 91.5 (84.7 to 95.3) 98.6 (96.9 to 99.5)
Relative difference (type 1/type 4) 0.96 (0.91 to 1.02)
CI, confidence interval. P values were obtained from likelihood ratio tests.
Source: Adapted from Table6in Abba 2011.
analyses in the review.
P = 0.20
1.04 (1.02 to 1.06)
P < 0.001
The results can be interpreted as follows:
●
The summary sensitivity of type 1 is 94.8%. We are 95% confident that the true value of sensitivity lies between 93.0% and 96.1%.
●
The summary specificity of type 1 is 95.2%. We are 95% confident that the true value of specificity lies between 93.2% and 96.7%.
●
The summary sensitivity of type 4 is 91.5%. We are 95% confident that the true value of sensitivity lies between 84.7% and 95.3%.
●
The summary specificity of type 4 is 98.6%. We are 95% confident that the true value of specificity lies between 96.9% and 99.5%.
●
Relative difference in sensitivity of type 4 and type 1 RDTs: the relative sensitivity of 0.96indicates that the sensitivity of type 4 RDTs is
4% lower than the sensitivity of type 1 RDTs, in terms of the corresponding summary estimates. We are 95% confident that this difference lies between 9% lower and 2% higher. This difference is not statistically significant (P = 0.20).
●
Relative difference in specificity of type 1 and type 4 RDTs: the relative specificity of 1.04indicates that the specificity of type 4 RDTs is
4% higher than that of type 1 RDTs. We are 95% confident that this difference lies between 2% higher and 6% higher. This difference is
statistically significant (P < 0.001).

St
udy
Hadlo
Krantz 2000
Mar
Schielen 2006
W
St
Hadlo
Krantz 2000
Mar
Schielen 2006
W
NT, PAPP-A, free βhCG and maternal age-maternal age < 35 years
NT
1
0 0.2 0.4 0.6 0.8 10 0.2 0.4 0.6 0.8 1
https://t.me/medicina_free
w 2005
chini 2010
apner 2003
, PAPP-A, free βhCG and maternal age-maternal age ≥ 35 years
TP
14
FP
FN
TN
Sensitivity (95% Cl)
165
3
8042
7
169
1
3589
2
35
1
1200
1
27
1
1704
8
151
4
3933
0.82 [0.57, 0.96]
0.88 [0.47, 1.00]
0.67 [0.09, 0.99]
0.50 [0.01, 0.99]
0.67 [0.35, 0.90]
Specificity (95% Cl)
0.98 [0.98, 0.98]
0.96 [0.95, 0.96]
0.97 [0.96, 0.98]
0.98 [0.98, 0.99]
0.96 [0.96, 0.97]
Sensitivity (95% Cl)
0 0.2 0.4 0.6 0.8 10 0.2 0.4 0.6 0.8
Specificity (95% Cl)
udy
w 2005
chini 2010
apner 2003
Figure11.7.a Forest plot of the NT, PAPP- A, free ßhCG and maternal age test strategy by maternal age group (< 35 years versus ≥ 35 years). βhCG, beta human
chorionic gonadotrophin; FN, false negative; FP, false positive; NT, nuchal translucency; PAPPtrue positive. Source: Adapted from Figure6in Alldred 2017
TP
15
23
14
44
FP
FN
TN
Sensitivity (95% CI)
209
0
1988
289
2
1729
4
39
1
261
163
5
2118
619
5
3452
1.00 [0.78, 1.00]
0.92 [0.74, 0.99]
0.80 [0.28, 0.99]
0.74 [0.49, 0.91]
0.90 [0.78, 0.97]
Specificity (95% Cl)
0.90 [0.89, 0.92]
0.86 [0.84, 0.87]
0.87 [0.83, 0.91l
0.93 [0.92, 0.94]
0.85 [0.84, 0.86]
Sensitivity (95% Cl) Specificity (95% Cl)
A, pregnancy- associated plasma protein A; TN, true negative; TP,
Соседние файлы в папке Библиотека им академика М.И. Перельмана
