Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_2767_Библиотеки_им_академика_М_И_Перельмана
.pdf
12
https://t.me/medicina_free
Drawing conclusions
Mariska M. Leeflang, Karen R. Steingart, Rob J. Scholten and Clare Davenport
KEY POINTS
Key issues threatening the strength of the evidence in a review are risk of bias, concerns
•
regarding applicability, heterogeneity, imprecision, and completeness of the body of
evidence.
When using the GRADE approach for assessing the certainty of the evidence, these key
•
issues can be translated to the five GRADE domains: risk of bias, indirectness, inconsistency, imprecision, and publication bias. These key issues alsoshould be addressed
in the Discussion section and Authors’ conclusion section of the review, together with
an explanation of what the results practically mean.
A ‘Summary of findings’ table can present the findings of the review in a clear, trans-
•
parent and structured format, as well as key information regarding the overall strength
or certainty of the evidence.
When discussing the implications of the review’s findings for practice, the potential
•
consequences of testing should beconsidered. Such a discussion should take into
account the fact that evidence of consequences is typically not documented in the
studies included in the review.
12.1 Introduction
The purpose of Cochrane Reviews is to facilitate healthcare decision- making by patients
and the general public, by clinicians or other healthcare workers, administrators and
policy makers. Such people will rely on the ‘Summary of findings’ tables, Discussion and
Authors’ conclusions to make sense of the information in the review and to help them to
interpret the results and the strength of the available evidence. Owing to the importance
This chapter should be cited as: Leeflang MM, Steingart KR, Scholten RJ, Davenport C. Chapter12: Drawing
conclusions. In: Deeks JJ, Bossuyt PM, Leeflang MM, Takwoingi Y, editors. Cochrane Handbook for Systematic
Reviews of Diagnostic Test Accuracy. 1st edition. Chichester (UK): John Wiley & Sons, 2023: 349–376.
Cochrane Handbook for Systematic Reviews of Diagnostic Test Accuracy, First Edition. Edited by
Jonathan J. Deeks, Patrick M. Bossuyt, Mariska M. Leeflang and Yemisi Takwoingi.
© 2023 The Cochrane Collaboration. Published 2023 by John Wiley & Sons Ltd.
349

12 Drawing conclusions
https://t.me/medicina_free
of the discussion and conclusions sections, review authors should take great care that
these sections accurately reflect the data and information contained in the review.
For systematic reviews of test accuracy, the key results are usually summary estimates
of sensitivity and specificity from meta- analysis, or summary estimates of differences in
accuracy between tests. The risk of bias and concerns regarding applicability of the accuracy estimates from an individual study should be assessed in the methodological quality assessment stage of a systematic review, as explained in Chapter5. The included
studies contribute to the overall body of evidence for a specific review question.
How much confidence we have in the overall body of evidence is referred to as the
strength of the evidence, or certainty in the evidence. Several alternative terms, such as
quality, have been used. Throughout this chapter we use ‘strength of the evidence’ as
the broader term, to distinguish it from methodological quality, which covers risk of
bias and concerns regarding applicability. Review authors who use the GRADE (Grading
of Recommendations Assessment, Development and Evaluation) framework will come
across the term ‘certainty of the evidence’. In this Handbook, we only use the term
‘certainty of the evidence’ when we refer specifically to GRADE guidance.
In addition to the strength of the evidence, the contribution of test accuracy to
evidence- based decision- making needs to be made explicit. Accuracy results usually do
not provide readers with clear answers about whether to buy, reimburse, implement or
order tests. Such decisions usually need more evidence concerning the potential consequences of index test- positive results and index test- negative results, and other ways in
which tests have an impact on patients. The Discussion section of a systematic review
of test accuracy should at least alert readers to this and indicate where additional
information might be found.
Above all, readers should weigh the results and their implications against the risk of
bias and concerns regarding applicability of the body of evidence from which they stem,
to know how confident they can be that the results are valid and applicable to the review
question. In addition, one should take into account how large and complete thebody of
evidence is, and likely heterogeneity in accuracy. The following sections inaCochrane
Review of diagnostic test accuracy facilitate interpretation of the reviewfindings.
The ‘Summary of findings’ table presents a summary of the review findings in a
clear, transparent, and structured format, as well as key information about the strength
of the evidence.
The Discussion section usually starts with a summary of the main findings from the
review, placed in the context of other research and knowledge. This section should be
followed by an explanation of the strengths and weaknesses of the review, and a
section about the applicability of findings to the review question.
Finally, the Authors’ conclusions section should explain the implications of the
review findings for practice and for research. In this chapter we provide suggestions on
how to approach each of these sections.
12.2 ‘Summary offindings’ tables
Similar to ‘Summary of findings’ tables in systematic reviews of interventions, it is
important that the main findings of a systematic review of test accuracy are presented
in a transparent and simple tabular format. The ‘Summary of findings’ tables should
350

12.2 ‘Summary offindings’ tables
https://t.me/medicina_free
provide key information on the accuracy of the index test(s) under consideration (and
the difference in accuracy when index tests are being compared), and important limitations arising from the assessment of the strength of the evidence.
A ‘Summary of findings’ table appears at the beginning of a Cochrane Review, before
the Background section. Cochrane Reviews of diagnostic test accuracy should have at
least one ‘Summary of findings’ table representing the primary review question. Some
reviews may include more than one ‘Summary of findings’ table if, for example, the
review addresses more than one primary objective (Kohli 2021).
In this section we outline the key features that should be included in a ‘Summary of
findings’ table. One method to create a ‘Summary of findings’ table is using the GRADE
approach for assessing the certainty in the evidence (templates available through the
software package GRADEpro GDT (GRADEpro 2020)). However, unlike for Cochrane
Reviews of interventions, this is not a requirement for Cochrane Reviews of diagnostic
test accuracy, and review authors may prefer to summarize using their own structure.
The following essential features should be included in a ‘Summary of findings’ table,
irrespective of which approach is used.
1)
The review question and its components, i.e. population, (prior tests), setting, index
test(s) and reference standard(s), should be described in full at the head of the table.
2) A brief description of how these components were addressed by the included stud-
ies, to facilitate the assessment of whether the included studies are applicable to the
review question.
3) The results for each index test should, at a minimum, include:
a) the number of included studies;
b) the number of participants, in sufficient detail to calculate the numbers of partici-
pants with and without the target condition;
c) the accuracy of the index test(s). This is usually reported as summary estimates of
sensitivity and specificity. For comparative accuracy reviews, the estimates of
absolute or relative differences in accuracy may also be reported. In situations
where it is not considered meaningful to provide summary estimates, review
authors may want to provide the range of reported estimates of sensitivity and
specificity. With multiple test positivity thresholds and estimation of a summary
receiver operating characteristic (ROC) curve, it may not be meaningful to provide
a summary sensitivity and specificity. In these cases, a summary diagnostic odds
ratio or summary estimates of sensitivity at certain values of specificity (or the
other way round) may be reported (see Chapter9, Section 9.4 and Chapter10,
Section10.3); and
d) the statistical uncertainty around any summary measure of test accuracy used
(e.g. 95% confidence interval or 95% credible interval).
4) There should be an explanation of what the results mean when applied to a hypo-
thetical cohort of people who will in practice undergo the index test(s). This means
that the absolute numbers of true and false positive and negative test results should
be stated (with accompanying confidence intervals), so that the reader gets a sense
of what the practical implications may be of using the index test(s). How to derive
these numbers is explained in Chapter11.
5) There should be a clear statement about the strength or certainty of the evidence,
including risk of bias, concerns regarding applicability, and between- study variability.
351

12 Drawing conclusions
https://t.me/medicina_free
Beyond these essential features, review authors may identify other aspects of the
results to include in the ‘Summary of findings’ table, such as variation in results by cutoff, prevalence or any other important potential source of heterogeneity. If desirable,
other accuracy measures may also be stated, such as predictive values or likelihood
ratios. However, review authors should be aware that using different metrics to report
the same information may be confusing for readers.
Although not essential, review authors could also explain the potential consequences
of test results, for example that people with a false positive test result may undergo
further, unnecessary testing.
Explanations may be provided about the results presented in the table as comments
or as footnotes. These may include, for example, reasons for downgrading the certainty
of evidence when the GRADE approach is used. In some reviews, review authors may be
concerned that the ‘Summary of findings’ table for the test cannot be safely interpreted
in isolation from the original data presented in the main body of the review, particularly
where there is substantial heterogeneity and where prevalence estimates used to derive
absolute numbers differ from the proportion with the target condition in included
studies.
Here we present four examples of ‘Summary of findings’ tables that illustrate the
points outlined. Table12.2.a shows a review of a single index test using a structure
created by the review authors. The authors did not follow the GRADE approach and
did not rate the certainty of the evidence as high, moderate, low or very low. They
narratively described the strength of the evidence.
Table12.2.b shows a ‘Summary of findings’ table for a review without a meta- analysis,
addressing multiplicity in the target condition. The review authors did not use the
GRADE approach and the explanation of the strength of the evidence could have been
more explicit and detailed. However, this ‘Summary of findings’ table is an example of
how to present the results in the absence of a meta- analysis.
Table 12.2.c shows a review of multiple index tests using the GRADE approach
(this format can also be applied to a single index test). Table12.2.d shows a review comparing two index tests using the GRADE approach. A further explanation of the GRADE
approach can be found in Section12.4. Review authors should be explicit about whether
they used the GRADE approach or not.
12.3 Assessing thestrength ofthe evidence
12.3.1 Key issues toconsider when assessing thestrength ofthe evidence
In this section, we first introduce the key issues that should be considered when assessing the strength of the overall body of evidence in a systematic review of test accuracy. In
section12.4, we explain the GRADE approach to assessing the certainty in the evidence.
Most systematic reviews will contain one or multiple meta- analyses and provide summary sensitivity and specificity. However, some reviews may have estimated a summary ROC curve and presented a summary diagnostic odds ratio. In other reviews the
body of evidence may have been insufficient or too heterogeneous to justify a metaanalysis. In these cases, the assessment of the strength of the evidence may differ.
Where applicable, guidance is provided for these situations.
352

Table12.2.a ‘Summary of findings’ table: What is the diagnostic accuracy of serum galactomannan for invasive aspergillosis in immunocompromised patients?
https://t.me/medicina_free
Population: immunocompromised patients, mostly haematology patients; applicable to the review question
Prior testing: varied, mostly physical examination and history (fever, neutropenia); applicable to the review question
Setting: mostly inpatients in haematology or cancer departments; applicable to the review question
Index test: Platelia Aspergillus test, which measures galactomannan, an Aspergillus antigen
Importance: depends on the time- gain the test may provide
Reference standard: a composite reference standard of clinical and microbiological criteria; the reference standard classifies the patients in four groups: no– possible–
probable– proven invasive aspergillosis (IA). In this ‘Summary of findings’ table, sensitivity refers to the proven or probable IA patients and specificity to the patients with
possible or no IA
Studies: 29 cross- sectional studies; two- group designs and studies excluding patients with possible IA were not included; studies had to report cut- off values that were used
and some studies reported more than one cut- off
What do the results mean?
Cutvalue
Summary
off
sensitivity
(95% CI)
Summary
specificity
(95% CI)
No. of
participants
(studies)
Median proportion
with target condition
(interquartile range)
With a prevalence of 11%
patients will have IA Strength of the evidence
*
, 11 out of 100
0.5 0.78
(0.70 to
0.85)
0.85
(0.78 to
0.91)
394 proven or
probable IA
3549 possible
or no IA
(27)
11%
(6.5% to 16%)
2 (95% CI 2 to 3) IA patients will be
missed, but will be tested again.
13 (8 to 20) out of 89 patients without
IA will be
unnecessarily referred for CT scanning
Risk of bias was unclear for most domains in most
studies, due to poor reporting
Three studies had concerns regarding applicability
of the included patients
Low numbers of diseased patients (1to 20)
Much between-
study variability, with sensitivity
ranging from 0% to 100%
1.0 0.71
(0.63 to
0.78)
1.5 0.63
(0.49 to
0.77)
0.90
(0.86 to
0.93)
0.93
(0.89 to
0.97)
145 proven or
probable IA
1246 possible
or no IA
(8)
209 proven or
probable IA
2412 possible
or no IA
(15)
13%
(4.2% to 31%)
7.4%
(4.3% to 16%)
3 (95% CI 2 to 4) IA patients will be
missed, but will be tested again
9 (6 to 12) out of 89 patients without IA
will be
unnecessarily referred for CT scanning
4 (95% CI 2 to 5) IA patients will be
missed, but will be tested again
6 (2 to 10) out of 89 patients without IA
will be
unnecessarily referred for CT scanning
Risk of bias was unclear for most studies, due to
poor reporting
No concerns regarding applicability for any of the
QUADAS-
2 domains
Low numbers of diseased patients (1to 34)
Low numbers of diseased patients (1to 17), except
one study (98 IA patients)
One study had high risk of bias in the patient
domain and three studies in the reference standard
domain One study had high concerns regarding
applicability of the patients
*
Median proportion with target condition over all studies was 11% (range 0.8% to 56%). CI, confidence interval; CT, computed tomography.
Comment: The results in this table should not be interpreted in isolation from the results of the individual included studies contributing to each summary test accuracy measure.
These are reported in the main body of the text of the review.
Source: Adapted from Leeflang 2015.

Table12.2.b ‘Summary of findings’ table: What is the diagnostic accuracy of physical tests for various causes of shoulder impingements in people whose
https://t.me/medicina_free
symptoms, history or both suggest impingement?
Setting: most people with shoulder pain symptomatic of impingements and related pathologies are diagnosed and managed in the primary care setting; this is
applicable to the review question
Index tests: physical tests used singly or in combination to identify shoulder impingement and related pathologies
Importance: accurate diagnosis using readily applied, convenient, low- cost physical tests would enable appropriate and well- timed management of these common
causes of shoulder pain
Reference standard: while a definitive reference standard is lacking, surgery, whether open or arthroscopic, is generally regarded as the best available; Non- invasive
contenders include ultrasound and magnetic resonance imaging (MRI)
Studies: 33 studies including 4002 shoulders in 3852 patients; These incorporated numerous standard, modified or combinations of index tests and 14novel index
tests
Methodological quality: methodological quality was generally poor; All but two studies failed to meet the criteria for having a representative spectrum of patients
Data analysis: the studies assessed 170 target condition/index test combinations, with only six instances of any index test being performed and interpreted similarly
in two studies; Meta- analysis of the latter was considered to be inappropriate, however
Target condition
*
Subacromial and internal impingement
Shoulders/
Subcategory of target condition, if applicable Studies
patients
Subacromial impingement 5 361/356 13
Tests or variants
evaluated
Subacromial versus internal impingement 1 110/110 1
Internal impingement 0 0 0
LHB tendinopathy or tears 3 660/557 10
Multiple, undifferentiated target conditions LHB/labral pathology; LHB/SLAP lesions; SA-
SDbursitis/
4 201/200 10
bursal- side degeneration of supraspinatus;and SIS/rotator
cuff tendinitis ortear
LHB, long head of biceps; SA- SD, subacromial- subdeltoid bursa; SIS, subacromial impingement syndrome; SLAP lesions, Superior Labrum Anterior to Posterior lesions.
*
A selection of target conditions as presented in the original review is presented here.
Note that methodological quality assessment was performed using QUADAS, hence the reason there is no explicit mention of risk of bias and applicability.
Source: Adapted from Hanchard 2013.

Table12.2.c ‘Summary of findings’ table: What is the diagnostic accuracy of rapid diagnostic tests (RDTs) for detecting Plasmodium vivax malaria parasitae-
https://t.me/medicina_free
mia in people living in malaria-
Population: people presenting with symptoms of uncomplicated malaria
Prior testing: none
Setting: ambulatory healthcare settings in P vivax- endemic areas
Index tests: immunochromatography- based RDTs for P vivax malaria that meet the World Health Organization (WHO) malaria RDT performance criteria (WHO
2017a); This table presents the results for the CareStart Malaria Pf/Pv Combo test
Reference standards: conventional microscopy, polymerase chain reaction
Target condition: P vivax malaria
Importance: accurate and fast diagnosis of P vivax from other malaria species allows appropriate treatment to be provided quickly
Study design: all cross- sectional studies
Findings: 10 studies of six different RDT brands; Only two brands (CareStart Malaria Pf/Pv Combo test and Falcivax Device Rapid test) were evaluated against the
same reference standard by more than one study
Limitations: a small number of studies were included in the analyses and meta- analyses were only possible for two RDT brands; Studies often did not report how
patients were selected, blinding of the RDT results to the reference standard, and the storage conditions and lot testing of RDTs
endemic areas who present to ambulatory healthcare facilities with symptoms suggestive of malaria?
Outcome
Number of
studies
Numbers in a cohort of 1000 patients tested
(95%confidence interval (CI))
Number of
patients
a
Prevalence
of20%
Certainty of the
evidence (GRADE)Prevalence of 0.5% Prevalence of 5%
CareStart Malaria Pf/Pv Combo test against microscopy: summary sensitivity (95% CI) = 99% (94% to 100%) and summary specificity (95% CI) = 99%
(99% to 100%), summary positive likelihood ratio (95% CI) = 141.1 (68.2 to 292.0) and summary negative likelihood ratio (95% CI) = 0.01 (0.00 to 0.06)
True positives
(patients with P vivax malaria)
False negatives
(patients incorrectly classified as not having
4 251 5
(5 to 10)
0
(0 to 0)
50
(47 to 50)
0
(0 to 3)
198
(188 to 200)
2
(0 to 12)
⨁⨁⨁◯
MODERATE
1
P vivax malaria)
True negatives
(patients without P vivax malaria)
False positives
(patients incorrectly classified as having P
2147 985
(980 to 995)
10
(0 to 10)
941
(941 to 950)
9
(0 to 9)
792
(792 to 800)
8
(0 to 8)
⨁⨁⨁◯
MODERATE
1
vivax malaria)
(Continued )

Table12.2.c (Continued)
https://t.me/medicina_free
Outcome
Number of
studies
Numbers in a cohort of 1000 patients tested
(95%confidence interval (CI))
Number of
patients
a
Prevalence
of20%
Certainty of the
evidence (GRADE)Prevalence of 0.5% Prevalence of 5%
Falcivax Device Rapid test against microscopy: summary sensitivity (95% CI) = 77% (53% to 91%) and summary specificity (95% CI) = 99% (98% to 100%),
summary positive likelihood ratio (95% CI) = 120.3 (43.1 to 335.9) and summary negative likelihood ratio (95% CI) = 0.23 (0.10 to 0.53)
True positives
2
(patients with P vivax malaria)
False negatives
(patients incorrectly classified as not having
89 4
(3 to 5)
1
(0 to 2)
50
(47 to 50)
11
(4 to 23)
198
(188 to 200)
46
(18 to 94)
⨁⨁◯◯
1,2
LOW
P vivax malaria)
True negatives
(patients without P vivax malaria)
False positives
(patients incorrectly classified as having P
2621
985
(975 to 995)
10
(0 to 10)
941
(931 to 950)
9
(0 to 19)
792
(784 to 800)
8
(0 to 16)
⨁⨁⨁◯
MODERATE
1
vivax malaria)
a
Median values were chosen from ranges of prevalence considered to be moderate, low and very low transmission settings for P vivax (WHO 2017b).
1
Downgraded for risk of bias by one.
2
Downgraded for imprecision by one due to wide confidence intervals.
GRADE Certainty of the evidence
High: we are very confident that the true effect lies close to that of the estimate of the effect.
Moderate: we are moderately confident in the effect estimate: the true effect is likely to be close to the estimate of the effect, but there is a possibility that it is
substantially different.
Low: our confidence in the effect estimate is limited: the true effect may be substantially different from the estimate of the effect.
Very low: we have very little confidence in the effect estimate: the true effect is likely to be substantially different from the estimate of effect.
Source: Adapted from Agarwal 2020.

Table12.2.d ‘Summary of findings’ table: What is the diagnostic accuracy of Xpert Ultra versus Xpert MTB/RIF for the detection of pulmonary tuberculosis in
https://t.me/medicina_free
adults with presumptive pulmonary tuberculosis?
Population: adults with presumptive pulmonary tuberculosis; participants were unselected, meaning they were not enrolled in a study based on prior testing with
microscopy examination (smear results) or a history of tuberculosis
Role: an initial test
Setting: primary care facilities and local hospitals
Index tests: Xpert Ultra and Xpert MTB/RIF on sputum
Threshold for index tests: an automated binary result is provided
Reference standards: solid or liquid culture
Studies: 7 cross- sectional and cohort studies that directly compared the accuracy of Xpert Ultra and Xpert MTB/RIF were included
Xpert Ultra summary sensitivity 90.9% (95% credible interval (CrI) 86.2 to 94.7) and summary specificity 95.6% (95% CrI 93.0 to 97.4)
Xpert MTB/RIF summary sensitivity 84.7% (95% CrI 78.6 to 89.9) and summary specificity 98.4% (95% CrI 97.0 to 99.3)
Test result
True positives (TP) 23
False negatives (FN) 2
Number of results per 1000 patients tested (95% CrI)
Prevalence 2.5% Prevalence 10% Prevalence 30%
Xpert MTB/
Xpert Ultra Xpert MTB/RIF Xpert Ultra
RIF Xpert Ultra
(22 to 24)21(20 to 22)91(86 to 95)85(79 to 90)
4
(3 to 5)
#
6more TP in Xpert Ultra 19more TP in Xpert Ultra
9
(5 to 14)15(10 to 21)27(16 to 41)46(30 to 64)
2more TP in Xpert Ultra
(1 to 3)
*
273
(259 to 284)
Xpert MTB/
RIF
254
(236 to 270)
2 fewer FN in Xpert Ultra 6 fewer TN in Xpert Ultra 19 fewer TP in Xpert Ultra
Certainty of the
Number of
participants
evidence
(GRADE)
983 ⊕⊕⊕⊕
High
(Continued )

Table12.2.d (Continued)
https://t.me/medicina_free
*
669
(651 to 682)
Xpert MTB/
RIF
689
(679 to 695)
Certainty of the
Number of
participants
evidence
(GRADE)
1852 ⊕⊕⊕⊕
High
Test result
Xpert Ultra Xpert MTB/RIF Xpert Ultra
True negatives (TN) 932
(907 to 950)
Number of results per 1000 patients tested (95% CrI)
Prevalence 2.5% Prevalence 10% Prevalence 30%
Xpert MTB/
RIF Xpert Ultra
959
(946 to 968)
860
(837 to 877)
886
(873 to 894)
27 fewer TN in Xpert Ultra 26 fewer TN in Xpert Ultra 20 fewer TN in Xpert Ultra
False positives (FP) 43
(25 to 68)16(7 to 29)
40
(23 to 63)14(6 to 27)31(18 to 49)11(5 to 21)
27more FP in Xpert Ultra 26more FP in Xpert Ultra 20more FP in Xpert Ultra
*
95% credible limits were estimated based on those around the point estimates for summary sensitivity and specificity; 95% confidence intervals were estimated for true
positives, false negatives, true negatives and false positives. Prevalence estimates were suggested by the World Health Organization Global Tuberculosis Programme. The
median proportion with tuberculosis in the included studies was 30.1% (range 12.8% to 72.2%).
#
These differences have been derived using GRADEpro, which does not provide credible or confidence intervals for the difference between tests.
GRADE Certainty of the evidence
High: we are very confident that the true effect lies close to that of the estimate of the effect.
Moderate: we are moderately confident in the effect estimate: the true effect is likely to be close to the estimate of the effect, but there is a possibility that it is substan-
tially different.
Low: our confidence in the effect estimate is limited: the true effect may be substantially different from the estimate of the effect.
Very low: we have very little confidence in the effect estimate: the true effect is likely to be substantially different from the estimate of effect.
Source: Adapted from Zifodya 2021.
Соседние файлы в папке Библиотека им академика М.И. Перельмана
