Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_2767_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
25 Мб
Скачать
12
https://t.me/medicina_free
Drawing conclusions
Mariska M. Leeflang, Karen R. Steingart, Rob J. Scholten and Clare Davenport
KEY POINTS
Key issues threatening the strength of the evidence in a review are risk of bias, concerns
regarding applicability, heterogeneity, imprecision, and completeness of the body of evidence. When using the GRADE approach for assessing the certainty of the evidence, these key
issues can be translated to the five GRADE domains: risk of bias, indirectness, incon­sistency, imprecision, and publication bias. These key issues alsoshould be addressed in the Discussion section and Authors’ conclusion section of the review, together with an explanation of what the results practically mean. A ‘Summary of findings’ table can present the findings of the review in a clear, trans-
parent and structured format, as well as key information regarding the overall strength or certainty of the evidence. When discussing the implications of the review’s findings for practice, the potential
consequences of testing should beconsidered. Such a discussion should take into account the fact that evidence of consequences is typically not documented in the studies included in the review.
12.1 Introduction
The purpose of Cochrane Reviews is to facilitate healthcare decision- making by patients and the general public, by clinicians or other healthcare workers, administrators and policy makers. Such people will rely on the ‘Summary of findings’ tables, Discussion and Authors’ conclusions to make sense of the information in the review and to help them to interpret the results and the strength of the available evidence. Owing to the importance
This chapter should be cited as: Leeflang MM, Steingart KR, Scholten RJ, Davenport C. Chapter12: Drawing conclusions. In: Deeks JJ, Bossuyt PM, Leeflang MM, Takwoingi Y, editors. Cochrane Handbook for Systematic Reviews of Diagnostic Test Accuracy. 1st edition. Chichester (UK): John Wiley & Sons, 2023: 349–376.
Cochrane Handbook for Systematic Reviews of Diagnostic Test Accuracy, First Edition. Edited by Jonathan J. Deeks, Patrick M. Bossuyt, Mariska M. Leeflang and Yemisi Takwoingi. © 2023 The Cochrane Collaboration. Published 2023 by John Wiley & Sons Ltd.
349
12 Drawing conclusions
https://t.me/medicina_free
of the discussion and conclusions sections, review authors should take great care that these sections accurately reflect the data and information contained in the review.
For systematic reviews of test accuracy, the key results are usually summary estimates of sensitivity and specificity from meta- analysis, or summary estimates of differences in accuracy between tests. The risk of bias and concerns regarding applicability of the accu­racy estimates from an individual study should be assessed in the methodological qual­ity assessment stage of a systematic review, as explained in Chapter5. The included studies contribute to the overall body of evidence for a specific review question.
How much confidence we have in the overall body of evidence is referred to as the strength of the evidence, or certainty in the evidence. Several alternative terms, such as quality, have been used. Throughout this chapter we use ‘strength of the evidence’ as the broader term, to distinguish it from methodological quality, which covers risk of bias and concerns regarding applicability. Review authors who use the GRADE (Grading of Recommendations Assessment, Development and Evaluation) framework will come across the term ‘certainty of the evidence’. In this Handbook, we only use the term ‘certainty of the evidence’ when we refer specifically to GRADE guidance.
In addition to the strength of the evidence, the contribution of test accuracy to evidence- based decision- making needs to be made explicit. Accuracy results usually do not provide readers with clear answers about whether to buy, reimburse, implement or order tests. Such decisions usually need more evidence concerning the potential conse­quences of index test- positive results and index test- negative results, and other ways in which tests have an impact on patients. The Discussion section of a systematic review of test accuracy should at least alert readers to this and indicate where additional information might be found.
Above all, readers should weigh the results and their implications against the risk of bias and concerns regarding applicability of the body of evidence from which they stem, to know how confident they can be that the results are valid and applicable to the review question. In addition, one should take into account how large and complete thebody of evidence is, and likely heterogeneity in accuracy. The following sections inaCochrane Review of diagnostic test accuracy facilitate interpretation of the reviewfindings.
The ‘Summary of findings’ table presents a summary of the review findings in a clear, transparent, and structured format, as well as key information about the strength of the evidence.
The Discussion section usually starts with a summary of the main findings from the review, placed in the context of other research and knowledge. This section should be followed by an explanation of the strengths and weaknesses of the review, and a section about the applicability of findings to the review question.
Finally, the Authors’ conclusions section should explain the implications of the review findings for practice and for research. In this chapter we provide suggestions on how to approach each of these sections.
12.2 ‘Summary offindings’ tables
Similar to ‘Summary of findings’ tables in systematic reviews of interventions, it is important that the main findings of a systematic review of test accuracy are presented in a transparent and simple tabular format. The ‘Summary of findings’ tables should
350
12.2 ‘Summary offindings’ tables
https://t.me/medicina_free
provide key information on the accuracy of the index test(s) under consideration (and the difference in accuracy when index tests are being compared), and important limita­tions arising from the assessment of the strength of the evidence.
A ‘Summary of findings’ table appears at the beginning of a Cochrane Review, before the Background section. Cochrane Reviews of diagnostic test accuracy should have at least one ‘Summary of findings’ table representing the primary review question. Some reviews may include more than one ‘Summary of findings’ table if, for example, the review addresses more than one primary objective (Kohli 2021).
In this section we outline the key features that should be included in a ‘Summary of findings’ table. One method to create a ‘Summary of findings’ table is using the GRADE approach for assessing the certainty in the evidence (templates available through the software package GRADEpro GDT (GRADEpro 2020)). However, unlike for Cochrane Reviews of interventions, this is not a requirement for Cochrane Reviews of diagnostic test accuracy, and review authors may prefer to summarize using their own structure.
The following essential features should be included in a ‘Summary of findings’ table, irrespective of which approach is used.
1)
The review question and its components, i.e. population, (prior tests), setting, index
test(s) and reference standard(s), should be described in full at the head of the table.
2) A brief description of how these components were addressed by the included stud-
ies, to facilitate the assessment of whether the included studies are applicable to the review question.
3) The results for each index test should, at a minimum, include:
a) the number of included studies; b) the number of participants, in sufficient detail to calculate the numbers of partici-
pants with and without the target condition;
c) the accuracy of the index test(s). This is usually reported as summary estimates of
sensitivity and specificity. For comparative accuracy reviews, the estimates of absolute or relative differences in accuracy may also be reported. In situations where it is not considered meaningful to provide summary estimates, review authors may want to provide the range of reported estimates of sensitivity and specificity. With multiple test positivity thresholds and estimation of a summary receiver operating characteristic (ROC) curve, it may not be meaningful to provide a summary sensitivity and specificity. In these cases, a summary diagnostic odds ratio or summary estimates of sensitivity at certain values of specificity (or the other way round) may be reported (see Chapter9, Section 9.4 and Chapter10, Section10.3); and
d) the statistical uncertainty around any summary measure of test accuracy used
(e.g. 95% confidence interval or 95% credible interval).
4) There should be an explanation of what the results mean when applied to a hypo-
thetical cohort of people who will in practice undergo the index test(s). This means that the absolute numbers of true and false positive and negative test results should be stated (with accompanying confidence intervals), so that the reader gets a sense of what the practical implications may be of using the index test(s). How to derive these numbers is explained in Chapter11.
5) There should be a clear statement about the strength or certainty of the evidence,
including risk of bias, concerns regarding applicability, and between- study variability.
351
12 Drawing conclusions
https://t.me/medicina_free
Beyond these essential features, review authors may identify other aspects of the results to include in the ‘Summary of findings’ table, such as variation in results by cut­off, prevalence or any other important potential source of heterogeneity. If desirable, other accuracy measures may also be stated, such as predictive values or likelihood ratios. However, review authors should be aware that using different metrics to report the same information may be confusing for readers.
Although not essential, review authors could also explain the potential consequences of test results, for example that people with a false positive test result may undergo further, unnecessary testing.
Explanations may be provided about the results presented in the table as comments or as footnotes. These may include, for example, reasons for downgrading the certainty of evidence when the GRADE approach is used. In some reviews, review authors may be concerned that the ‘Summary of findings’ table for the test cannot be safely interpreted in isolation from the original data presented in the main body of the review, particularly where there is substantial heterogeneity and where prevalence estimates used to derive absolute numbers differ from the proportion with the target condition in included studies.
Here we present four examples of ‘Summary of findings’ tables that illustrate the points outlined. Table12.2.a shows a review of a single index test using a structure created by the review authors. The authors did not follow the GRADE approach and did not rate the certainty of the evidence as high, moderate, low or very low. They narratively described the strength of the evidence.
Table12.2.b shows a ‘Summary of findings’ table for a review without a meta- analysis, addressing multiplicity in the target condition. The review authors did not use the GRADE approach and the explanation of the strength of the evidence could have been more explicit and detailed. However, this ‘Summary of findings’ table is an example of how to present the results in the absence of a meta- analysis.
Table 12.2.c shows a review of multiple index tests using the GRADE approach (this format can also be applied to a single index test). Table12.2.d shows a review com­paring two index tests using the GRADE approach. A further explanation of the GRADE approach can be found in Section12.4. Review authors should be explicit about whether they used the GRADE approach or not.
12.3 Assessing thestrength ofthe evidence
12.3.1 Key issues toconsider when assessing thestrength ofthe evidence
In this section, we first introduce the key issues that should be considered when assess­ing the strength of the overall body of evidence in a systematic review of test accuracy. In section12.4, we explain the GRADE approach to assessing the certainty in the evidence.
Most systematic reviews will contain one or multiple meta- analyses and provide sum­mary sensitivity and specificity. However, some reviews may have estimated a sum­mary ROC curve and presented a summary diagnostic odds ratio. In other reviews the body of evidence may have been insufficient or too heterogeneous to justify a meta­analysis. In these cases, the assessment of the strength of the evidence may differ. Where applicable, guidance is provided for these situations.
352
Table12.2.a ‘Summary of findings’ table: What is the diagnostic accuracy of serum galactomannan for invasive aspergillosis in immunocompromised patients?
https://t.me/medicina_free
Population: immunocompromised patients, mostly haematology patients; applicable to the review question Prior testing: varied, mostly physical examination and history (fever, neutropenia); applicable to the review question Setting: mostly inpatients in haematology or cancer departments; applicable to the review question Index test: Platelia Aspergillus test, which measures galactomannan, an Aspergillus antigen Importance: depends on the time- gain the test may provide Reference standard: a composite reference standard of clinical and microbiological criteria; the reference standard classifies the patients in four groups: no– possible–
probable– proven invasive aspergillosis (IA). In this ‘Summary of findings’ table, sensitivity refers to the proven or probable IA patients and specificity to the patients with possible or no IA Studies: 29 cross- sectional studies; two- group designs and studies excluding patients with possible IA were not included; studies had to report cut- off values that were used and some studies reported more than one cut- off
What do the results mean?
Cut­value
Summary
off
sensitivity (95% CI)
Summary specificity (95% CI)
No. of participants (studies)
Median proportion with target condition (interquartile range)
 With a prevalence of 11% patients will have IA Strength of the evidence
*
, 11 out of 100
0.5 0.78 (0.70 to
0.85)
0.85 (0.78 to
0.91)
394 proven or probable IA 3549 possible or no IA (27)
11% (6.5% to 16%)
2 (95% CI 2 to 3) IA patients will be missed, but will be tested again. 13 (8 to 20) out of 89 patients without IA will be unnecessarily referred for CT scanning
Risk of bias was unclear for most domains in most studies, due to poor reporting Three studies had concerns regarding applicability of the included patients Low numbers of diseased patients (1to 20) Much between-
study variability, with sensitivity
ranging from 0% to 100%
1.0 0.71 (0.63 to
0.78)
1.5 0.63 (0.49 to
0.77)
0.90 (0.86 to
0.93)
0.93 (0.89 to
0.97)
145 proven or probable IA 1246 possible or no IA (8)
209 proven or probable IA 2412 possible or no IA (15)
13% (4.2% to 31%)
7.4% (4.3% to 16%)
3 (95% CI 2 to 4) IA patients will be missed, but will be tested again 9 (6 to 12) out of 89 patients without IA will be unnecessarily referred for CT scanning
4 (95% CI 2 to 5) IA patients will be missed, but will be tested again 6 (2 to 10) out of 89 patients without IA will be unnecessarily referred for CT scanning
Risk of bias was unclear for most studies, due to poor reporting No concerns regarding applicability for any of the QUADAS-
2 domains
Low numbers of diseased patients (1to 34)
Low numbers of diseased patients (1to 17), except one study (98 IA patients) One study had high risk of bias in the patient domain and three studies in the reference standard domain One study had high concerns regarding applicability of the patients
*
Median proportion with target condition over all studies was 11% (range 0.8% to 56%). CI, confidence interval; CT, computed tomography. Comment: The results in this table should not be interpreted in isolation from the results of the individual included studies contributing to each summary test accuracy measure. These are reported in the main body of the text of the review. Source: Adapted from Leeflang 2015.
Table12.2.b ‘Summary of findings’ table: What is the diagnostic accuracy of physical tests for various causes of shoulder impingements in people whose
https://t.me/medicina_free
symptoms, history or both suggest impingement?
Setting: most people with shoulder pain symptomatic of impingements and related pathologies are diagnosed and managed in the primary care setting; this is applicable to the review question
Index tests: physical tests used singly or in combination to identify shoulder impingement and related pathologies Importance: accurate diagnosis using readily applied, convenient, low- cost physical tests would enable appropriate and well- timed management of these common
causes of shoulder pain Reference standard: while a definitive reference standard is lacking, surgery, whether open or arthroscopic, is generally regarded as the best available; Non- invasive contenders include ultrasound and magnetic resonance imaging (MRI) Studies: 33 studies including 4002 shoulders in 3852 patients; These incorporated numerous standard, modified or combinations of index tests and 14novel index tests
Methodological quality: methodological quality was generally poor; All but two studies failed to meet the criteria for having a representative spectrum of patients Data analysis: the studies assessed 170 target condition/index test combinations, with only six instances of any index test being performed and interpreted similarly
in two studies; Meta- analysis of the latter was considered to be inappropriate, however
Target condition
*
Subacromial and internal impingement
Shoulders/
Subcategory of target condition, if applicable Studies
patients
Subacromial impingement 5 361/356 13
Tests or variants evaluated
Subacromial versus internal impingement 1 110/110 1
Internal impingement 0 0 0
LHB tendinopathy or tears 3 660/557 10
Multiple, undifferentiated target conditions LHB/labral pathology; LHB/SLAP lesions; SA-
SDbursitis/
4 201/200 10 bursal- side degeneration of supraspinatus;and SIS/rotator cuff tendinitis ortear
LHB, long head of biceps; SA- SD, subacromial- subdeltoid bursa; SIS, subacromial impingement syndrome; SLAP lesions, Superior Labrum Anterior to Posterior lesions.
*
A selection of target conditions as presented in the original review is presented here. Note that methodological quality assessment was performed using QUADAS, hence the reason there is no explicit mention of risk of bias and applicability. Source: Adapted from Hanchard 2013.
Table12.2.c ‘Summary of findings’ table: What is the diagnostic accuracy of rapid diagnostic tests (RDTs) for detecting Plasmodium vivax malaria parasitae-
https://t.me/medicina_free
mia in people living in malaria-
Population: people presenting with symptoms of uncomplicated malaria Prior testing: none Setting: ambulatory healthcare settings in P vivax- endemic areas Index tests: immunochromatography- based RDTs for P vivax malaria that meet the World Health Organization (WHO) malaria RDT performance criteria (WHO
2017a); This table presents the results for the CareStart Malaria Pf/Pv Combo test
Reference standards: conventional microscopy, polymerase chain reaction Target condition: P vivax malaria Importance: accurate and fast diagnosis of P vivax from other malaria species allows appropriate treatment to be provided quickly Study design: all cross- sectional studies Findings: 10 studies of six different RDT brands; Only two brands (CareStart Malaria Pf/Pv Combo test and Falcivax Device Rapid test) were evaluated against the
same reference standard by more than one study Limitations: a small number of studies were included in the analyses and meta- analyses were only possible for two RDT brands; Studies often did not report how patients were selected, blinding of the RDT results to the reference standard, and the storage conditions and lot testing of RDTs
endemic areas who present to ambulatory healthcare facilities with symptoms suggestive of malaria?
Outcome
Number of studies
Numbers in a cohort of 1000 patients tested (95%confidence interval (CI))
Number of patients
a
Prevalence of20%
Certainty of the evidence (GRADE)Prevalence of 0.5% Prevalence of 5%
CareStart Malaria Pf/Pv Combo test against microscopy: summary sensitivity (95% CI) = 99% (94% to 100%) and summary specificity (95% CI) = 99%
(99% to 100%), summary positive likelihood ratio (95% CI) = 141.1 (68.2 to 292.0) and summary negative likelihood ratio (95% CI) = 0.01 (0.00 to 0.06)
True positives
(patients with P vivax malaria)
False negatives
(patients incorrectly classified as not having
4 251 5
(5 to 10)
0 (0 to 0)
50 (47 to 50)
0 (0 to 3)
198 (188 to 200)
2 (0 to 12)
⨁⨁⨁◯
MODERATE
1
P vivax malaria)
True negatives
(patients without P vivax malaria)
False positives
(patients incorrectly classified as having P
2147 985
(980 to 995)
10 (0 to 10)
941 (941 to 950)
9 (0 to 9)
792 (792 to 800)
8 (0 to 8)
⨁⨁⨁◯
MODERATE
1
vivax malaria)
(Continued )
Table12.2.c (Continued)
https://t.me/medicina_free
Outcome
Number of studies
Numbers in a cohort of 1000 patients tested (95%confidence interval (CI))
Number of patients
a
Prevalence of20%
Certainty of the evidence (GRADE)Prevalence of 0.5% Prevalence of 5%
Falcivax Device Rapid test against microscopy: summary sensitivity (95% CI) = 77% (53% to 91%) and summary specificity (95% CI) = 99% (98% to 100%),
summary positive likelihood ratio (95% CI) = 120.3 (43.1 to 335.9) and summary negative likelihood ratio (95% CI) = 0.23 (0.10 to 0.53)
True positives
2
(patients with P vivax malaria)
False negatives
(patients incorrectly classified as not having
89 4
(3 to 5)
1 (0 to 2)
50 (47 to 50)
11 (4 to 23)
198 (188 to 200)
46 (18 to 94)
⨁⨁◯◯
1,2
LOW
P vivax malaria)
True negatives
(patients without P vivax malaria)
False positives
(patients incorrectly classified as having P
2621
985 (975 to 995) 10 (0 to 10)
941 (931 to 950) 9 (0 to 19)
792 (784 to 800) 8 (0 to 16)
⨁⨁⨁◯
MODERATE
1
vivax malaria)
a
Median values were chosen from ranges of prevalence considered to be moderate, low and very low transmission settings for P vivax (WHO 2017b).
1
Downgraded for risk of bias by one.
2
Downgraded for imprecision by one due to wide confidence intervals. GRADE Certainty of the evidence High: we are very confident that the true effect lies close to that of the estimate of the effect. Moderate: we are moderately confident in the effect estimate: the true effect is likely to be close to the estimate of the effect, but there is a possibility that it is substantially different. Low: our confidence in the effect estimate is limited: the true effect may be substantially different from the estimate of the effect. Very low: we have very little confidence in the effect estimate: the true effect is likely to be substantially different from the estimate of effect. Source: Adapted from Agarwal 2020.
Table12.2.d ‘Summary of findings’ table: What is the diagnostic accuracy of Xpert Ultra versus Xpert MTB/RIF for the detection of pulmonary tuberculosis in
https://t.me/medicina_free
adults with presumptive pulmonary tuberculosis?
Population: adults with presumptive pulmonary tuberculosis; participants were unselected, meaning they were not enrolled in a study based on prior testing with microscopy examination (smear results) or a history of tuberculosis
Role: an initial test Setting: primary care facilities and local hospitals Index tests: Xpert Ultra and Xpert MTB/RIF on sputum Threshold for index tests: an automated binary result is provided Reference standards: solid or liquid culture Studies: 7 cross- sectional and cohort studies that directly compared the accuracy of Xpert Ultra and Xpert MTB/RIF were included Xpert Ultra summary sensitivity 90.9% (95% credible interval (CrI) 86.2 to 94.7) and summary specificity 95.6% (95% CrI 93.0 to 97.4) Xpert MTB/RIF summary sensitivity 84.7% (95% CrI 78.6 to 89.9) and summary specificity 98.4% (95% CrI 97.0 to 99.3)
Test result
True positives (TP) 23
False negatives (FN) 2
Number of results per 1000 patients tested (95% CrI)
Prevalence 2.5% Prevalence 10% Prevalence 30%
Xpert MTB/
Xpert Ultra Xpert MTB/RIF Xpert Ultra
RIF Xpert Ultra
(22 to 24)21(20 to 22)91(86 to 95)85(79 to 90)
4 (3 to 5)
#
6more TP in Xpert Ultra 19more TP in Xpert Ultra
9 (5 to 14)15(10 to 21)27(16 to 41)46(30 to 64)
2more TP in Xpert Ultra
(1 to 3)
*
273 (259 to 284)
Xpert MTB/ RIF
254 (236 to 270)
2 fewer FN in Xpert Ultra 6 fewer TN in Xpert Ultra 19 fewer TP in Xpert Ultra
Certainty of the Number of participants
evidence
(GRADE)
983 ⊕⊕⊕⊕
High
(Continued )
Table12.2.d (Continued)
https://t.me/medicina_free
*
669 (651 to 682)
Xpert MTB/ RIF
689 (679 to 695)
Certainty of the Number of participants
evidence
(GRADE)
1852 ⊕⊕⊕⊕
High
Test result
Xpert Ultra Xpert MTB/RIF Xpert Ultra
True negatives (TN) 932
(907 to 950)
Number of results per 1000 patients tested (95% CrI)
Prevalence 2.5% Prevalence 10% Prevalence 30%
Xpert MTB/ RIF Xpert Ultra
959 (946 to 968)
860 (837 to 877)
886 (873 to 894)
27 fewer TN in Xpert Ultra 26 fewer TN in Xpert Ultra 20 fewer TN in Xpert Ultra
False positives (FP) 43
(25 to 68)16(7 to 29)
40 (23 to 63)14(6 to 27)31(18 to 49)11(5 to 21)
27more FP in Xpert Ultra 26more FP in Xpert Ultra 20more FP in Xpert Ultra
*
95% credible limits were estimated based on those around the point estimates for summary sensitivity and specificity; 95% confidence intervals were estimated for true positives, false negatives, true negatives and false positives. Prevalence estimates were suggested by the World Health Organization Global Tuberculosis Programme. The median proportion with tuberculosis in the included studies was 30.1% (range 12.8% to 72.2%).
#
These differences have been derived using GRADEpro, which does not provide credible or confidence intervals for the difference between tests. GRADE Certainty of the evidence High: we are very confident that the true effect lies close to that of the estimate of the effect. Moderate: we are moderately confident in the effect estimate: the true effect is likely to be close to the estimate of the effect, but there is a possibility that it is substan- tially different. Low: our confidence in the effect estimate is limited: the true effect may be substantially different from the estimate of the effect. Very low: we have very little confidence in the effect estimate: the true effect is likely to be substantially different from the estimate of effect. Source: Adapted from Zifodya 2021.