Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_2767_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
25 Мб
Скачать
11.5 Individual andsummary estimates oftest accuracy
https://t.me/medicina_free
Details for the included studies may best be reported in a ‘Characteristics of included studies’ table. Such a table may present at a glance key details for each study. Chapter7 details the characteristics that should be extracted from each study to populate a ‘Characteristics of included studies’ table. Careful planning of the data extraction form can help ensure that items are recorded in a manner that allows their direct incorpora­tion in the table. This saves time, particularly in reviews with substantial numbers of included studies.
11.4 Methodological quality ofincluded studies
This section describes the methodological quality of studies included in the systematic review. For most systematic reviews of test accuracy, the QUADAS- 2 tool will be used to evaluate methodological quality in terms of risk of bias and concerns regarding applica­bility (see Chapter8).
The results of the assessment are usually presented in a table or figure. A figure may show the results for individual studies or across studies. For reviews with multiple tar­get conditions or index tests, figures may also be presented separately for each target condition or test. If there are several tables or figures, review authors should carefully consider which ones to include in the main text and which ones to include in an appendix.
An informative description of the assessments of risk of bias would indicate, for each domain, whether studies were considered at risk of producing biased results, and why they were classified as such. Readers would also like to know to what extent an included study helped to answer the review question, i.e. whether there were concerns regarding applicability. Here also, it is informative to indicate whether such concerns arose, and why.
If studies were potentially eligible but excluded after an evaluation of their risk of bias or applicability, they should not be reported here. The characteristics of studies excluded because of risk of bias or concerns regarding applicability should be reported elsewhere, in a separate table or supplementary file.
If the review addresses comparisons of test accuracy between two or more index tests, the results of the assessment with the QUADAS- C tool will be presented next. The QUADAS- C results are typically specific to a comparison of two tests within the review. If there are two or more such comparisons, separate QUADAS- C assessments need to be reported.
11.5 Individual andsummary estimates oftest accuracy
For each objective, the review should present the relevant evidence and the findings, to enable conclusions to be drawn. For primary objectives this typically requires presentation of results from individual studies as well as from meta- analyses, if such analyses were performed. Meta- analyses may produce summary estimates of sensitiv­ity and specificity or SROC curves (see Chapter9, Section9.2.2 and Section9.3). If no meta- analysis could be performed, a different approach to summarizing the results of individual studies will be adopted (see Section11.9).
329
11 Presenting findings
https://t.me/medicina_free
11.5.1 Presenting results fromincluded studies
It will be informative to present test accuracy results from the individual studies. Such overviews typically include the study identifier, year of publication, total number of par­ticipants, number of participants with the target condition, number of true and false posi­tives, number of true and false negatives, estimates of test accuracy, as defined in the protocol, and expressions of statistical imprecision (typically as 95% confidence inter­vals). These overviews may be presented in tables or, more succinctly, on coupled forest plots of sensitivity and specificity (see Chapter9, Section9.2.1 and example Figure9.2.a).
For pairwise comparisons of accuracy, individual studies will present multiple pairs of estimates of sensitivity and specificity. In that case, it is helpful to structure the table or graph in such a way that the link between these pairs is obvious. This can be done by presenting a linked SROC plot showing the multiple estimates from each study (one for each test), connected by a line (see Chapter9, Figure9.4.g, for an example).
11.5.2 Presenting summary estimates ofsensitivity andspecificity
Review authors should carefully consider the summary measures that are relevant and appropriate to present for a given analysis, rather than compute and report all possible measures of test accuracy. Many systematic reviews of test accuracy present summary estimates of sensitivity and specificity.
Other summary estimates, like positive and negative predictive values, may also be produced. The selection of appropriate summary estimates will also depend on the design of the primary studies. For example, if one reference standard is always used in index test positives and a different one in index test negatives, it will be more meaning­ful to present summary estimates of the negative and predictive values. The use of dif­ferent reference standards is known to introduce bias in estimates of sensitivity and specificity.
A review may include multiple meta- analyses and corresponding summary estimates. These estimates can be summarized in one or more tables as appropriate. The table(s) should include the number of studies, participants and participants with and without the target condition on which each analysis is based (see an example in Table11.5.a).
When summary estimates of sensitivity and specificity are reported as an output of meta- analysis, they need to be interpreted as expressions of a central tendency of esti­mates across the included studies. There may be considerable uncertainty in these summary estimates and substantial variability in accuracy across studies.
11.5.3 Presenting SROC curves
Including a 2×2 table of the number of true positives, false positives, false negatives and true negatives from each study regardless of threshold value allows estimation of an SROC curve. If an SROC curve is assumed to be symmetrical (i.e. the shape parameter in the HSROC model is assumed to be zero), then accuracy does not depend on threshold. In that case, the SROC curve can be described by a constant diagnostic odds ratio (DOR) (see Chapter4, Section 4.5.7). A DOR of 1 represents an uninformative test (the upward diagonal in Figure11.5.a). SROC curves closer to the top left- hand corner of the SROC plot represent tests with a higher DOR, hence more discriminatory power.
330
11.5 Individual andsummary estimates oftest accuracy
DOR = 361
Sensitivity
0
Specificity
https://t.me/medicina_free
Table11.5.a Accuracy ofXpert MTB/RIF fordetection ofextrapulmonary tuberculosis
Type of specimen
Reference standard
Number of studies (participants)
Number with TB (%)
Summary sensitivity % (95% credible interval)
Summary specificity % (95% credible interval)
CSF Culture 30 (3395) 571 (16.8) 71.1 (62.8 to 79.1) 96.9 (95.4 to 98.0)
CSF Composite 14 (2203) 862 (39.1) 42.3 (32.1 to 52.8) 99.8 (99.3 to 100.0)
Pleural fluid Culture 25 (3065) 644 (21.0) 49.5 (39.8 to 59.9) 98.9 (97.6 to 99.7)
Pleural fluid Composite 10 (1024) 616 (60.1) 18.9 (11.5 to 27.9) 99.3 (98.1 to 99.8)
Lymph node
Culture 14 (1588) 627 (39.5) 88.9 (82.7 to 93.6) 86.2 (78.0 to 92.3)
aspirate
Lymph node
Composite 4 (679) 377 (55.5) 81.6 (61.9 to 93.3) 96.4 (91.3 to 98.6)
aspirate
CSF, cerebrospinal fluid; TB, tuberculosis. Source: Adapted from Table2in Kohli 2021.
When interpreting DORs, review authors should note that the same DOR may be achieved by different combinations of sensitivity and specificity (see Chapter 4, Section4.5.7). In addition, single measures of test accuracy, like the DOR, are often not clinically informative, because of a lack of information on errors in those with the target condition (false negatives) and those without the condition (false positives). The relative
DOR = 81
0.0 0.2 0.4 0.6 0.8 1. 0
Figure11.5.a Symmetrical summary receiver operating characteristic (SROC) curves and diagnostic
odds ratios (DORs)
DOR = 16
DOR = 5
DOR = 2
DOR = 1 (uninformative test)
Line of symmetry
0.20.4 0.6 0.81. 0
0.
331
11 Presenting findings
https://t.me/medicina_free
magnitude of such errors is essential for judging the extent and likely impact of conse­quences of positive and negative test results.
DORs are probably most useful in meta- analysis when making comparisons between tests or between subgroups (see Chapter9, Section 9.4.6.4, and Section 11.6.2). For these comparisons, if two curves are symmetrical or have the same shape, then the relative accuracy of the two curves can be summarized using the ratio of the DORs or relative DOR (RDOR).
It may be challenging to describe an SROC curve that is not symmetrical, either ver­bally or numerically. As explained in Chapter10 (Section10.3.1), the anticipated sensi­tivity at a given value of specificity (or the other way round) can be computed from the SROC curve to illustrate test performance in such cases. If review authors choose to report such sensitivity and specificity pairs from an SROC curve, then the most informa­tive estimates will be points on the curve that lie within the range of the sensitivity and specificity estimates reported by the included studies. For example, these could be val­ues that represent the median and interquartile range from the studies included in the meta-
analysis (see Chapter10, Box 10.3.b).
Alternatively, if minimizing false positives– therefore maximizing specificity – in a particular testing context is desired, the sensitivity of the test could be reported at the minimally acceptable specificity (for example, a specificity of 95%).
Whatever the approach taken to select sensitivity and specificity pairs, the values and corresponding estimates can be presented in a table, as shown in the example in Table11.5.b. Likelihood ratios can also be estimated at such fixed values.
11.5.4 Describing uncertainty insummary statistics
The summary sensitivity and specificity are estimates based on the included studies and there will always be statistical uncertainty about their true value. Review authors should express the degree of statistical uncertainty associated with summary estimates of test accuracy, irrespective of the metric used. Confidence intervals or credible inter­vals should therefore be reported alongside the point estimates of summary sensitivity and specificity in text and tables. For summary points on SROC plots, confidence regions should also be presented (see Chapter9, Figure9.4.a).
Table11.5.b Sensitivity andlikelihood ratios for11C- PIB- PET at fixed values ofspecificity
forAlzheimer’s dementia
Statistic Fixed value
of specificity %
Lower quartile 56 96 (88 to 99) 2.19 (2.09 to 2.29) 0.07 (0.02 to 0.23)
Median 58 96 (87 to 99) 2.29 (2.17 to 2.41) 0.07 (0.02 to 0.24)
Upper quartile 81 89 (68 to 97) 4.66 (4.03 to 5.39) 0.14 (0.05 to 0.44)
CI, confidence interval. Nine studies (n: 112with dementia; 162without dementia) Source: Adapted from Table4in Zhang 2014.
332
Estimated sensitivity % (95% CI)
Positive likelihood ratio (95% CI)
Negative likelihood ratio (95% CI)
11.6 Comparisons oftest accuracy
https://t.me/medicina_free
11.5.5 Describing heterogeneity insummary statistics
Statistical heterogeneity (referred to simply as heterogeneity) exists whenever esti­mates of test accuracy vary between studies, more than would be expected from within­study sampling error (chance) alone. This is extremely common in systematic reviews of test accuracy. Prediction intervals and prediction regions give an indication of het­erogeneity. Prediction regions plotted around summary points in ROC space represent the area where the sensitivity and specificity estimates from a future test accuracy study are expected to lie (see Chapter9, Section9.3.2, Section9.4.1 and Section9.4.2).
Systematic reviews of test accuracy studies can include prediction regions with cover­age probabilities of 50%, 90% or 95%. The 95% prediction regions often cover large areas of ROC space because of the presence of substantial heterogeneity common in many systematic reviews of test accuracy. While the confidence region depicts uncer­tainty in the summary estimates of sensitivity and specificity due to within­ability, the prediction region depicts the uncertainty due to both within- and between- study variability (see Chapter9, Figure9.4.a). As such, when heterogeneity is high, the 95% prediction region will be much larger than the 95% confidence region.
As stated in Chapter9, estimation of a prediction interval or region relies on the assump­tion of normal distributions for the effects across studies. This may be very problematic when the number of studies is small and can lead to spuriously large (or small) regions (Deeks 2019). In addition, if the variance parameters and the correlation (or covariance) between logit sensitivity and specificity in a bivariate model cannot be reliably estimated, a prediction region will be misleading. Therefore, when there are few studies (e.g. fewer than 10) in the meta- analysis of a single test, review authors should carefully consider the appropriateness of presenting a prediction region around the summary point.
study vari-
11.6 Comparisons oftest accuracy
Review authors should consider two issues in reviews that compare the accuracy of mul­tiple tests: (1) the statistical measures that can be used; and (2) the strength of the evi­dence for the comparison. The strength primarily relates to whether the meta- analysis is based on within- study (direct) or between- study (indirect) comparisons of tests (see Chapter9, Section9.1.4.3). This will be considered in Chapter12 (Section12.1.3.5 and Section12.6.2.4). The appropriate statistical measures are not affected by this issue.
When summarizing findings from a comparison of two tests, review authors should focus on describing (1) the magnitude and direction of the difference between tests; (2)the uncertainty in the estimates; and (3) the degree of heterogeneity. Presentation of test comparisons can be facilitated by summaries of test accuracy in ROC space, which allow readers to compare test performance in one SROC plot. This may be in the form of summary estimates of sensitivity and specificity or SROC curves (shape and relative position). In addition, within- study test comparisons can be annotated to distinguish them from between- study comparisons.
11.6.1 Comparing tests using summary points
The magnitude and direction of the difference between tests can be summarized byreporting point estimates of the summary sensitivity and specificity for each test
333
11 Presenting findings
https://t.me/medicina_free
Table11.6.a Accuracy ofchest ultrasonography andchest radiography fordiagnosis
ofpneumothorax
Test Studies Participants (with
pneumothorax)
CUS 9 1271 (410) 0.91 (0.85 to 0.94) 0.99 (0.97 to 1.00)
CXR 9 1271 (410) 0.47 (0.31 to 0.63) 1.00 (0.97 to 1.00)
Absolute difference
CUS, chest ultrasonography; CXR, chest radiography. The P values are from Wald tests. Source: Adapted from Table2in Chan 2020.
Summary sensitivity (95% CI)
0.44 (0.27 to 0.61) P < 0.001
Summary specificity (95% CI)
−0.007 (−0.018 to 0.005) P = 0.26
and by measures of absolute or relative differences in sensitivity and specificity (seeTable11.6.a). Focusing on the size and the uncertainty in the estimated difference in summary sensitivity and specificity between tests illustrates the potential impact of using one test or the other.
For the example in Table 11.6.a, the summary sensitivities of 0.91 for chest ultra­sonography and 0.47 for chest radiography indicate that chest ultrasonography can be expected to correctly detect 44more patients out of every 100 with pneumothorax, compared to chest radiography. The 95% confidence interval translates into a differ­ence that lies between 27 and 61more patients out of every 100with pneumothorax. Asimilar approach can be used if predictive values are being compared.
11.6.2 Comparing tests using SROC curves
When SROC curves of two or more tests are compared, it matters whether or not the SROC curves for the tests have the same shape. If they do, the value of the RDOR will be constant all the way along the curve. In those situations, the RDOR is a valid reflection of the difference in accuracy between tests. In that case, it will be informative to report the summary DOR for each test and estimates of the RDOR, with expressions of statisti­cal uncertainty and heterogeneity.
Table11.6.b presents the results of an indirect comparison of the accuracy of urea breath test, serology and stool antigen test for detecting a Helicobacter pylori infection. It is based on an HSROC model with symmetrical shape for the curves (Best 2018). The RDOR of 3.22 (95% CI 1.24 to 8.37, P = 0.017) indicates that the DOR for the urea breath­ 13C test is about three times that of serology, and that we are 95% confident that this ratio lies between 1.24 and 8.37 times higher.
A higher DOR means that, for any selected positivity threshold for the inferior test, we can find a positivity threshold for the superior test that produces higher sensitivity for the same specificity, higher specificity for the same sensitivity, or both higher sensitivity and specificity. We cannot, however, identify the positivity thresholds on the SROC curve that correspond to specific sensitivity–specificity pairs.
Where tests have been compared using SROC curves, it may nevertheless be informa­tive to report selected sensitivity–specificity pairs on each of the curves, to facilitate test comparisons. For example, the sensitivity of each test at the same fixed specificity could be reported, or the other way round. Review authors should be cautious when choosing
334
Table11.6.b Comparison of the accuracy of non- invasive tests for Helicobacter pylori infection
https://t.me/medicina_free
Index test Studies; participants
(with H. pylori)
DOR (95% CI) Relative diagnostic odds ratios (95% CI), P value
13
Urea breath test-
C Urea breath test- 14C Serology
Urea breath test- 13C 34; 3139 (1526) 153 (73.7 to 316)
14
Urea breath test-
C 21; 1810 (1018) 105 (74.0 to 150) 1.45 (0.65 to 3.26)
P = 0.36
Serology 34; 4242 (2477) 47.4 (25.5 to 88.1) 3.22 (1.24 to 8.37)
P = 0.017
Stool antigen test 29; 2988 (1311) 45.1 (24.2 to 84.1) 3.39 (1.30 to 8.83)
P = 0.013
CI, confidence interval. The indirect comparison included all studies that evaluated at least one of the four tests, i.e. all available data. The RDOR is the diagnostic odds ratio (DOR) of the test in the column relative to the DOR of the test in the row. If the RDOR is greater than one, then the test in the column is more accurate than the test in the row. The P values are from Wald tests. Source: Adapted from Table2in Best 2018.
2.22 (1.09 to 4.51) P = 0.028
2.33 (1.14 to 4.76) P = 0.020
1.05 (0.44 to 2.53) P = 0.91
11 Presenting findings
https://t.me/medicina_free
pairs for comparison. To be valid, the pairs should be based on the results in the review: they should lie within the range of observed sensitivity and specificity estimates from the included studies.
If the SROC curves for two or more tests differ in shape, the ratio of the corresponding DOR will not be constant along the entire length of the curve. In these situations, compari­sons are challenging and interpretation of meta- analysis results needs to be done carefully, similarly considering the observed range of sensitivity and specificity estimates in individ­ual studies. Here also, selecting pairs on the fitted curves may assist in interpretation, provided that these lie within the range of estimates reported by the included studies.
11.6.3 Interpretation ofconfidence intervals fordifferences intest accuracy
It has been shown that some readers wrongly apply heuristics about ratio measures, as used in systematic reviews of interventions (Zhelev 2013). In comparing interventions, they know that a confidence interval including unity indicates no significant difference in outcomes between the interventions. Readers should not interpret confidence inter­vals for sensitivity and specificity in a similar way. A confidence interval for sensitivity and specificity that includes unity (1.00, or 100%) does not reflect poor test perfor­mance but rather the opposite: the possibility of perfect performance. Review authors should therefore consider supplementing numerical presentation of statistical uncer­tainty in confidence intervals with verbal explanations. A natural frequency presenta­tion format may additionally help (see Box11.6.a and Section11.8), particularly if a systematic review of test accuracy includes both estimates of the accuracy of single tests and a comparison of test accuracy.
When interpreting confidence intervals associated with comparisons, confidence intervals that do not overlap can be assumed to represent statistically significant differ­ences. Confidence intervals that overlap may or may not reflect statistically significant differences; P values will be required to draw conclusions about statistical significance. Interpretation of a test comparison is illustrated in Box11.6.a.
11.7 Investigations ofsources ofheterogeneity
Investigations of heterogeneity aim to assess whether test accuracy tends to vary with identifiable features, such as the characteristics of the participants, settings, tests, ref­erence standards or others. As explained in Chapter 9 and Chapter 10, hierarchical meta- regression models can estimate differences in accuracy between subgroups (orthe association of accuracy with a continuous measure) and formally test differences and associations for statistical significance.
Comparisons of subgroups based solely on studies that report data for subgroups within studies provide more valid evidence of sources of heterogeneity than compari­sons between studies. For example, a Cochrane Review on Down syndrome screening explored the effect of advanced maternal age (< 35 years versus 35 years) on test performance. Of the 69 studies for one of the index tests assessed in the review, 5 pro­vided data to compare the performance of the test between women younger than 35years and those 35 years or more within the same study (Alldred 2017). The 5 studies all showed a higher sensitivity and higher false positive fraction for the 35 years compared to the < 35 years subgroup, as shown in Figure11.7.a. Such analyses are
336
Box 11.6.a Interpretation ofconfidence intervals andP values forcomparisons oftest performance
https://t.me/medicina_free
A Cochrane Review compared the accuracy of rapid diagnostic tests (RDTs) for detecting Plasmodium falciparum malaria (Abba 2011). RDTs use different types of antibody or antibody combinations to detect Plasmodium antigens. Two RDT types, type 1 and type 4, were compared in one of the meta-
RDT type Summary sensitivity % (95% CI) Summary specificity % (95% CI)
Type 1: HRP- 2 antibody- based tests 94.8 (93.0 to 96.1) 95.2 (93.2 to 96.7)
Type 4: pLDH antibody- based tests 91.5 (84.7 to 95.3) 98.6 (96.9 to 99.5)
Relative difference (type 1/type 4) 0.96 (0.91 to 1.02)
CI, confidence interval. P values were obtained from likelihood ratio tests. Source: Adapted from Table6in Abba 2011.
analyses in the review.
P = 0.20
1.04 (1.02 to 1.06) P < 0.001
The results can be interpreted as follows:
The summary sensitivity of type 1 is 94.8%. We are 95% confident that the true value of sensitivity lies between 93.0% and 96.1%.
The summary specificity of type 1 is 95.2%. We are 95% confident that the true value of specificity lies between 93.2% and 96.7%.
The summary sensitivity of type 4 is 91.5%. We are 95% confident that the true value of sensitivity lies between 84.7% and 95.3%.
The summary specificity of type 4 is 98.6%. We are 95% confident that the true value of specificity lies between 96.9% and 99.5%.
Relative difference in sensitivity of type 4 and type 1 RDTs: the relative sensitivity of 0.96indicates that the sensitivity of type 4 RDTs is 4% lower than the sensitivity of type 1 RDTs, in terms of the corresponding summary estimates. We are 95% confident that this differ­ence lies between 9% lower and 2% higher. This difference is not statistically significant (P = 0.20).
Relative difference in specificity of type 1 and type 4 RDTs: the relative specificity of 1.04indicates that the specificity of type 4 RDTs is 4% higher than that of type 1 RDTs. We are 95% confident that this difference lies between 2% higher and 6% higher. This difference is statistically significant (P < 0.001).
St
udy
Hadlo Krantz 2000 Mar Schielen 2006 W
St
Hadlo Krantz 2000 Mar Schielen 2006 W
NT, PAPP-A, free βhCG and maternal age-maternal age < 35 years
NT
1
0 0.2 0.4 0.6 0.8 10 0.2 0.4 0.6 0.8 1
https://t.me/medicina_free
w 2005
chini 2010
apner 2003
, PAPP-A, free βhCG and maternal age-maternal age ≥ 35 years
TP
14
FP
FN
TN
Sensitivity (95% Cl)
165
3
8042
7
169
1
3589
2
35
1
1200
1
27
1
1704
8
151
4
3933
0.82 [0.57, 0.96]
0.88 [0.47, 1.00]
0.67 [0.09, 0.99]
0.50 [0.01, 0.99]
0.67 [0.35, 0.90]
Specificity (95% Cl)
0.98 [0.98, 0.98]
0.96 [0.95, 0.96]
0.97 [0.96, 0.98]
0.98 [0.98, 0.99]
0.96 [0.96, 0.97]
Sensitivity (95% Cl)
0 0.2 0.4 0.6 0.8 10 0.2 0.4 0.6 0.8
Specificity (95% Cl)
udy
w 2005
chini 2010
apner 2003
Figure11.7.a Forest plot of the NT, PAPP- A, free ßhCG and maternal age test strategy by maternal age group (< 35 years versus ≥ 35 years). βhCG, beta human
chorionic gonadotrophin; FN, false negative; FP, false positive; NT, nuchal translucency; PAPP­true positive. Source: Adapted from Figure6in Alldred 2017
TP
15 23
14 44
FP
FN
TN
Sensitivity (95% CI)
209
0
1988
289
2
1729
4
39
1
261
163
5
2118
619
5
3452
1.00 [0.78, 1.00]
0.92 [0.74, 0.99]
0.80 [0.28, 0.99]
0.74 [0.49, 0.91]
0.90 [0.78, 0.97]
Specificity (95% Cl)
0.90 [0.89, 0.92]
0.86 [0.84, 0.87]
0.87 [0.83, 0.91l
0.93 [0.92, 0.94]
0.85 [0.84, 0.86]
Sensitivity (95% Cl) Specificity (95% Cl)
A, pregnancy- associated plasma protein A; TN, true negative; TP,