Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_2767_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
25 Мб
Скачать
12.3 Assessing thestrength ofthe evidence
https://t.me/medicina_free
12.3.1.1 How valid are thesummary estimates?
Validity here refers to low risk of bias. The QUADAS- 2 tool and QUADAS- C (its extension for comparative accuracy studies) are recommended for assessing risk of bias and applicability of test accuracy studies (see Chapter8). High risk of bias caused by partici­pant selection, conduct of the index test, conduct and nature of the reference standard or flaws in the flow of participants, including timing of tests or missing values, may have a negative impact on the findings of a systematic review.
Just stating that the findings may come from studies with a high risk of bias is insuf­ficient; review authors should judge to what extent included studies with a high risk of bias may lead to bias in the summary estimates or affect the findings of the systematic review otherwise. For example, one or two studies at high risk of bias in a large number of included studies will have a different impact on estimates compared to one or two studies at high risk of bias in a review containing a small number of included studies. One way to gain some insight into the effect of studies with high or unclear risk of bias is to perform sensitivity analyses by excluding such studies (see Chapter9, Section9.4.9). If the results of the sensitivity analyses indicate that such studies greatly influence the findings of the review and thus question the robustness of the findings, then this should be explicitly noted.
12.3.1.2
How applicable are thesummary estimates?
A study can be at low risk of bias but, because of the included participants, test or refer­ence standard, may nevertheless not apply well to the review question. Concerns regard­ing the applicability of the individual studies to the review question can be assessed using QUADAS- 2. Review authors should judge to what extent studies with some con­cerns regarding applicability may affect the strength of the results in the whole body of evidence. The results of sensitivity analyses may support these judgements.
A specific concern regarding applicability may occur in comparative accuracy studies. As stated in Chapter5, studies that directly compare two tests could have recruited a study group that is unrepresentative of the patient population in whom the tests will be used in clinical practice. For example, a systematic review assessed whether exercise electrocardiography (ECG) may be replaced with computed tomography (CT) coronary angiography to detect coronary stenosis in patients with stable angina pectoris sus­pected of coronary disease (Nielsen 2014). The review authors only included studies that directly compared the two index tests. The exercise ECG test is usually done in patients with less severe coronary disease, whereas CT coronary angiography is used in more advanced cases. Therefore, only including studies of patients who received both tests in the past decreases the confidence in the applicability of the resulting compara­tive accuracy estimates, even though the risk of bias for estimates of accuracy gener­ated by studies that directly compared the two index tests may be low.
12.3.1.3 How heterogeneous are theindividual study estimates?
Variability in the study results may be explained by chance variation, if small and fewstudies have been included. Alternatively, it may be explained by the recruitment from different study populations, differences in the use and/or definition of the index tests, the use of different thresholds for test positivity, or variation in study methods. In Chapter 9, heterogeneity is defined as variation that goes beyond what may be
359
12 Drawing conclusions
https://t.me/medicina_free
expected by chance. Chapter9 also explains how potential sources of heterogeneity can be investigated.
In the Discussion section of a review and in the ‘Summary of findings’ tables, review authors should comment not only on the effect of possible sources of heterogeneity, but also on the consequences of heterogeneity for the interpretation of the findings.
The accuracy of a test is not a fixed property; it is likely to vary across populations and settings. Hence, heterogeneity is expected in systematic reviews of diagnostic accuracy. Whether heterogeneity exists may be difficult to assess in systematic reviews of a single test. If the sensitivity and specificity estimates of all studies in a review lie within a rela­tively narrow space in the ROC space, then one could claim little or no heterogeneity. When the estimates of sensitivity and specificity vary between 0% and 100%, the varia­tion is likely to go beyond chance variation. However, the situations in between are more difficult to assess, especially when the study results are also imprecise and the role of chance may be considerable.
One way to communicate heterogeneity is to report a prediction region. As explained in Chapter9, a 95% prediction region represents the region within which one has 95% confidence that the true sensitivity and specificity of any future study (resembling the studies included in the review) should lie (Harbord 2007). The greater the between­study heterogeneity, the larger this region will be.
Another way to communicate heterogeneity may be to focus on a minimally required estimate for sensitivity, specificity or another outcome measure. One could then explain to what extent the studies provide point estimates above this minimally required estimate, and how the confidence intervals of the studies relate to this point.
Little evidence and guidance have been developed to assess and communicate inconsistency when summary receiver operating characteristic (SROC) curves and odds ratios are presented. In those situations, review authors may comment on the SROC plots and how the study estimates are visually distinct from the summary curve.
Assessment of heterogeneity in comparative studies may be focused on the difference in sensitivity or specificity between the tests evaluated, in case of binary test results. For example, if some of the included studies indicate that test A has a much higher sensitiv­ity than test B, while other studies report that test B has a higher sensitivity, then these studies contradict one another, which may indicate heterogeneity.
12.3.1.4 How precise are thesummary estimates?
If only one small study was found that assessed the accuracy of an index test, then the estimates of sensitivity and specificity will be very imprecise. If a systematic review contains many large studies with a considerable number of both people with and those without the target condition, then the summary estimates will be a more precise reflection of the sensitivity and specificity over all the included studies.
Review authors are therefore expected to report:
the number of included studies;
the number of included participants with the target condition, which informs estima-
tion of sensitivity;
the number of included participants without the target condition, which informs
estimation of specificity; and
the statistical uncertainty around the summary estimates (e.g. 95% confidence interval).
360
12.3 Assessing thestrength ofthe evidence
https://t.me/medicina_free
These may be used to provide a judgement about the precision of the estimates.
The example of only one small study versus many large studies illustrates two extremes. The judgement of imprecision is subjective, but one could pre- specify a required width of the confidence interval (e.g. maximum 10 percentage points wide) to support this judgement.
Another way to address imprecision is to check whether the confidence interval crosses a pre- specified lower or upper bound of sensitivity or specificity. For example, if the test needs to be at least 70% sensitive and the lower 95% confidence interval is 65%, then one could decide that the point estimate was not sufficiently precise to judge whether the test is at least 70% sensitive. Such a pre- specification requires judgement about the potential for a test to have clinical utility based on the role of the test in the clinical pathway.
If no meta-
analysis can be done, review authors may report the individual study estimates of test accuracy, including information about the imprecision of these esti­mates (e.g. confidence intervals). These findings may be summarized by providing a range of estimates across included studies in the ‘Summary of findings’ table. In those situations, it should be made clear that a summary estimate cannot be provided.
12.3.1.5 How complete is thebody ofevidence?
One of the most prominent threats to the validity of a meta- analysis is the omission of certain results from the available literature. Even with the most comprehensive search strategy, reporting bias may occur. This could result from selective reporting of results within a study, from selective publication of studies, or from excluding reports in spe­cific languages.
The most obvious form of this omission manifests as publication bias: studies with promising or favourable results are more likely to be published and thus included in a systematic review, which may bias the summary estimates towards more favourable numbers (Simes1986).
Chapter9, Section9.5.3, addressed investigation and handling of publication bias. Although an appropriate test for detecting funnel plot asymmetry in systematic reviews of test accuracy that may indicate publication bias exists, it has limitations (Deeks
2005). Therefore, review authors are encouraged to explore how mechanisms underpin­ning reporting bias may operate in their clinical area and to inform the reader about the potential for reporting bias whenever possible. For example, review authors with knowledge about clinical testing pathways may be able to comment on the potential for selective reporting of results in studies that describe patients undergoing ‘multiple index tests’ as part of their routine care, but report the results of only some or one of these tests.
Some decisions made in the review methods may also lead to more favourable results than are justified. An example is language bias, caused by a restriction to reports in English and the phenomenon that larger and more positive studies may be more fre­quently published in English- language journals. Another example is missing studies due to flawed search strategies.
More information about the effects of choices in the search strategy and about publi­cation bias can be found in Chapter6. Review authors should judge to what extent any of these effects may be present and, if so, what their potential effect on the summary
361
12 Drawing conclusions
https://t.me/medicina_free
estimates would be. Expressing concerns about missing evidence and explaining why review authors think this may be problematic is more informative than merely stating that ‘there is publication bias’.
12.3.1.6 Were index test comparisons made between or within primary studies?
Many comparative accuracy reviews include studies assessing the accuracy of only one of the index tests in the comparison. Estimates of accuracy from these single index test accuracy studies may be at low risk of bias, with low concerns regarding applicability, but estimates of comparative accuracy will then be based on an indirect comparison (see Chapter9, Section 9.1.4.3) and may be susceptible to bias. Indirect comparisons are at higher risk of bias than direct comparisons made using studies that have com­pared tests head to head.
With an indirect comparison, an observed difference in the accuracy of the index tests may be confounded by the test settings, or by the difference in participants included in the respective studies for each of the tests. If all exercise ECG studies in the example in Section12.3.1.2 were done in patients with less severe coronary disease and all CT angi­ography studies were done in patients with more severe coronary disease, then the observed difference in accuracy between the two tests may be due to differences in disease severity rather than a genuine difference in accuracy.
12.4 GRADE approach for assessing the certainty of evidence
The GRADE Working Group developed a comprehensive and transparent system for grading the certainty of evidence and subsequently for grading the strength of recom­mendations following from evidence. As more recent GRADE guidance uses the term ‘certainty of the evidence’ instead of ‘quality of the evidence’ or ‘strength of the evi­dence’, we also use the term ‘certainty’ in this section. In systematic reviews, the cer­tainty in evidence reflects the extent to which we are confident that estimates are close to the truth. GRADE publications of particular relevance to systematic reviews of test accuracy include Schünemann (2008), Schünemann (2020a) and Schünemann (2020b).
The GRADE approach assesses certainty of evidence using the following five domains that are described in detail in Section12.4.1.
Risk of bias
Indirectness (applicability)
Inconsistency (heterogeneity)
Imprecision
Publication bias
These domains overlap and correspond with the key features presented in Section12.3.1, but may be used and named slightly differently within the GRADE frame­work. The certainty of the evidence starts as high when there are appropriate test accu­racy studies that recruit a group of participants with an uncertain diagnosis and who are representative of the target population. If a reason is found for downgrading, sys­tematic review authors should use their judgement to classify the reason as either seri­ous (e.g. downgraded by one level for serious imprecision) or very serious (e.g.
362
12.4 GRADE approach for assessing the certainty of evidence
https://t.me/medicina_free
downgraded by two levels for very serious imprecision). The overall certainty of the evi­dence for a given outcome can then range from high to very low. Review authors should be transparent about their judgements byexplicitly stating the reason for downgrading in footnotes (labelled as ‘Explanations’ in the GRADE ‘Summary of findings’ table), so that the reader can understand the reason for the decision. See Table 12.2.c and Table12.2.d for examples of ‘Summary of findings’ tables using the GRADE framework.
In addition, the GRADE approach for systematic reviews of interventions uses three reasons to upgrade the certainty of the evidence from a lower to a higher level of cer­tainty. These reasons are (1) a large effect; (2) any plausible confounding that would reduce the effects found; and (3) a dose- response gradient. However, these reasons are difficult to translate to systematic reviews of test accuracy. As clear guidance for upgrad­ing diagnostic accuracy evidence is lacking, we do not advise upgrading.
Review authors using GRADE (or any other formal method) to assess the certainty of the evidence should make this explicit in the Methods section of the review. The descrip­tion of the application of GRADE should include how judgements were made and whether the software package GRADEpro GDT was used to build the ‘Summary of find­ings’ tables (GRADEpro2020).
The GRADE guidance for diagnostic accuracy questions is mainly based on summary estimates of the sensitivity and specificity of a single test, or the summary estimates of the difference in sensitivity and specificity of two tests (Yang 2021).
12.4.1 GRADE domains for assessing certainty of evidence for test accuracy
Using the GRADE approach, the five domains are assessed separately for participants with the target condition (sensitivity estimates, true positive and false negative) and those without the target condition (specificity estimates, true negative and false posi­tive). The five domains are described in detail in the following sections (Schünemann 2008, Schünemann 2020a, Schünemann 2020b).
12.4.1.1 Risk ofbias
Review authors may downgrade the certainty of evidence by one or two levels depend­ing on the number (percentage) of included studies considered to have a high or unclear risk of bias for one or more of the QUADAS- 2 domains (see Section 12.3.1.1 and Chapter 8). Moving from risk- of- bias judgements of individual studies to an overall judgement about downgrading the body of evidence, however, can be challenging and relies on subjective judgement.
In a systematic review on tests for identifying people with glaucoma, the review authors downgraded the certainty of the evidence for four of the five tests they assessed due to risk of bias. For example, the certainty of the evidence for the sensitivity and specificity of the Oblique flashlight test was downgraded one level because 40% of studies had a high risk of bias in one or more QUADAS- 2 domains.
12.4.1.2 Indirectness (applicability)
Indirectness can be regarded as a synonym for applicability and can be assessed by judging whether concerns regarding applicability, as identified with QUADAS- 2, justify downgrading the certainty of evidence.
363
12 Drawing conclusions
https://t.me/medicina_free
Review authors may downgrade the certainty of evidence for indirectness if there are important differences between the participants in the studies and the population stated in the review question, with respect to either prior testing, spectrum of disease or comorbidities, or settings.
They may also downgrade if there are important differences between the character­istics of the index test(s) and/or the expertise of the people applying the test(s) in the studies compared to the characteristics of the test(s) and/or users in real- world settings.
They may downgrade in case of comparative questions if the index tests were not directly compared in each of the studies included in the test comparison (i.e. indirect comparisons).
A systematic review of tests for plague downgraded the certainty of the evidence for specificity one level for indirectness: there was a high concern about applicability due to exclusion of people who received antibiotics prior to sample collection. The index test was performed in a central laboratory, which may not reflect the field conditions in case of an outbreak (Jullien 2020).
12.4.1.3
Inconsistency (heterogeneity)
Inconsistency can be caused by identifiable clinical heterogeneity, or it may remain unexplained. As GRADE recommends downgrading for unexplained inconsistency in sensitivity and specificity estimates, review authors should state whether they carried out pre- specified analyses to investigate potential sources of heterogeneity and con­sider downgrading when they cannot explain inconsistency in the accuracy estimates. Downgrading is not necessary if the inconsistency can be explained, for example due to differences in setting or population or index test execution. If that is the case, the ‘Summary of findings’ tables should report the results per subgroup.
For example, in a systematic review about Xpert MTB/RIF Ultra and Xpert MTB/RIF assays for extrapulmonary tuberculosis and rifampicin resistance in adults (Kohli 2021), the review authors performed sensitivity analyses based on selected QUADAS- 2items and found that they could not explain the heterogeneity. The review authors then downgraded one level for inconsistency.
Questions to consider are ‘Are the individual point estimates in the forest plots more or less the same?’ and, more importantly, ‘How much do confidence intervals overlap?’ A scatter plot of sensitivity and specificity (i.e. an SROC plot) is helpful for visual assess­ment of heterogeneity (‘Are the sensitivity–specificity pairs clustered closely together or are they spread all over the ROC space?’). In addition, the size of the 95% prediction region (if presented) can assist in this assessment, particularly when there are many studies.
Downgrading for inconsistency may also be warranted when the differences between two tests differ too much. What is regarded to be ‘too much’ should then be specified. For example, one could decide– based on expected consequences of testing– that test A is chosen over test B when the sensitivity of test A is at least 10 percentage points higher than test B. When the review then includes studies that show much larger differ­ences, studies showing no difference and studies with a difference the other way (i.e.test B being more sensitive than test A), this may be a reason to downgrade for inconsistency (Hultcrantz 2020).
364
12.5 Summary ofmain results inthe Discussion section
https://t.me/medicina_free
12.4.1.4 Imprecision
Efforts to provide guidance on how to operationalize the assessments of imprecision for diagnostic test accuracy are ongoing. As per GRADE for systematic reviews of interven­tions and the explanation in Section12.3.1.1, one could look at the width of the 95% confidence intervals of the summary estimates of sensitivity and specificity and assess whether they cross a certain clinically acceptable lower limit (for which we would down­grade) or not. This clinically acceptable lower limit should depend on the clinical con­text and the potential consequences of testing, as explained in the clinical pathway (see Chapter5, Section5.3.1).
The systematic review of tests for plague, mentioned earlier, downgraded the cer­tainty of the evidence for specificity one level for imprecision. The review authors based this judgement on the 95% confidence intervals around the proportion of false positives and the proportion of false negatives: the lower limit of these confidence intervals would lead to a different decision than the upper limit of the confidence intervals (Jullien 2020).
12.4.1.5
Publication bias may occur when the body of evidence is incomplete. It can lower cer­tainty of evidence, mainly because studies with favourable results tend to be published more often than those with less favourable results. Although test accuracy studies with promising results about the performance of tests seem to be published more rapidly compared to those reporting lower estimates, there is no strong evidence that specific results, either favourable or unfavourable, end up less frequently in meta- analyses (Korevaar 2016). Still, review authors should be aware of the possibility of publication bias (see Section12.3.1.5).
studies, in particular if they support a pre- existing hypothesis and were funded by a body with a vested interest in a specific test (Schünemann 2020b). In some situations, there may be direct evidence that results of particular studies have been withheld or that studies are reporting partial findings, for example reporting only the sensitivity or specificity of the test, or reporting only particular subgroups of patients. All these exam­ples are publication and reporting practices that should be reported in the review. Alternatively, if few concerns were identified, review authors may judge that publica­tion bias was undetected.
2021) included studies with for- profit interest and small studies, which were thought to be an indication of potential publication bias. However, the review authors did notdowngrade for publication bias, as the literature search was thought to be compre­hensive and the review authors contacted researchers of primary studies to identify unpublished studies.
Publication bias
Downgrading may be considered when published evidence is limited to a few small
The systematic review about Xpert MTB/RIF Ultra and Xpert MTB/RIF assays (Kohli
12.5 Summary ofmain results inthe Discussion section
The Discussion section should begin with a restatement of the clinical question or questions that the review is attempting to answer. It should then give a summary of the results that provide answers to these questions. As a starting point, the number of
365
12 Drawing conclusions
https://t.me/medicina_free
included studies, total number of participants with and without the target condition, results of the risk of bias and applicability assessments, and consistency of findings should be summarized. This summary should be consistent with what was reported inthe ‘Summary of findings’ table and may be an elaboration or explanation of the information in the table.
If meta- analysis was performed, appropriate summary statistics from estimation of summary points or summary curves should be reported (see Chapter11). Although sen­sitivity and specificity are the most frequent measures in test accuracy studies, there is some evidence that predictive values are better understood (Whiting 2015). We there­fore recommend that review authors present potential consequences of testing by applying the summary estimates to a hypothetical cohort of persons to be tested.
If meta­mary of the review findings, if possible, combined with the range of estimates of test accuracy from the included studies. Review authors should also summarize the rele­vance of the findings from investigations of heterogeneity, if such analyses were feasi­ble. This section should be consistent with and link to the ‘Summary of findings’ table(s), including the assessment of the certainty of the body of evidence, as outlined in Chapter13, Section 13.3, or using the GRADE approach, as outlined in Section12.4. Many of the issues discussed in previous sections may be briefly but clearly mentioned in the Discussion section to provide a complete picture of the evidence.
analysis was not possible, review authors should provide a narrative sum-
12.6 Strengths andweaknesses ofthe review
In this section we focus on a narrative of the strengths and weaknesses of the primary studies, the strengths and weaknesses of the systematic review methods and the con­sequences these may have for the interpretation of the review’s results and conclu­sions. If review authors are aware of key issues that potentially limit or bias the results of their review, such as those considered when assessing the certainty of evidence for the ‘Summary of findings’ table, then these should be pointed out to readers. Review authors could consider using the following subheadings, when appropriate: ‘Strengths and weaknesses of the included studies’ (or ‘Certainty of the evidence’) and ‘Strengths and weaknesses of the review process’.
Review authors should discuss the strengths and weaknesses of the review with regard to accuracy estimates, not the strengths and weaknesses of the evidence with regard to policy- making decisions, which would rely on other considerations, such as the acceptability or feasibility of the test, the impact of the test on people- important outcomes, and resource requirements (costs).
12.6.1 Strengths andweaknesses ofincluded studies
This section narratively summarizes assessments made about the strength or certainty of the evidence. It is important to highlight the strength of the evidence, including its potential limitations, not only in the ‘Summary of findings’ tables but also in the main text of the review.
Relative strengths of the included studies may be highlighted here, without overstat­ing the implications. For example, it may be good to know that a review contained many
366
12.6 Strengths andweaknesses ofthe review
https://t.me/medicina_free
large studies with very similar results. For comparative questions, it may be important to highlight that results were obtained from fully paired (within- participant or head- to­head) comparative accuracy studies in the relevant population, and if the superiority of one test over another was consistent across the included studies.
Relative weaknesses of the included studies should at least be summarized with ref­erence to each of the four QUADAS- 2 domains (patient selection, index test(s), reference standard, and flow and timing), highlighting items particularly relevant to the review question, as reflected by the tailoring of QUADAS- 2 or QUADAS- C to the review topic. A detailed discussion of the types of bias that might occur in test accuracy studies, includ­ing comparative test accuracy studies, can be found in Chapter8.
Review authors should be mindful that readers of systematic reviews of test accuracy are likely to be less familiar with the types of bias that are encountered in test accuracy research (Zhelev 2013) and that descriptions of the mechanisms underlying important potential sources of bias may facilitate understanding.
Assessing the impact of bias on estimates of accuracy can be challenging, especially as an overall quality score is discouraged (Whiting 2005). Assessment should include consideration of the relative importance of the four QUADAS-
2 domains to the review topic and the proportion and size of studies at risk of bias. The assessment of the strength of the evidence can be helpful for such a summary, as described in Section12.4.
12.6.2 Strengths andweaknesses ofthe review
Limitations of the review focus on the review process and may include shortcomings in the search strategy, selection of studies, data extraction and analyses. This section should summarize the potential implications of these limitations on the strength of the conclusions. Review authors may not have been able to conduct their review as origi­nally intended in the protocol. For example, paucity of studies may have precluded planned investigations of heterogeneity. Limitations of the search strategy should be addressed as a shortcoming and the potential for bias caused by failure to retrieve or to translate reports should be discussed.
12.6.2.1 Strengths andweaknesses dueto thesearch andselection process
Some decisions made in the retrieval process may lead to omissions in the body of evi­dence that may be relevant for its interpretation; for example, excluding certain lan­guages or only searching in a limited number of databases. The potential consequences of these choices should be addressed in this section.
Lack of consensus in the selection of included studies is another potential limitation of the review process. If there has been substantial disagreement between review authors about the inclusion of studies, there is a risk of including less appropriate stud­ies or of excluding studies from which it is more difficult to extract data. Although the effects of these shortcomings may be limited, they are still potential sources of bias.
12.6.2.2 Strengths andweaknesses dueto methodological quality assessment anddata extraction
Information about the characteristics of participants, setting, study design and other key information may be missing in reports of primary studies and, therefore, limit
367
12 Drawing conclusions
https://t.me/medicina_free
assessment of the methodological quality of included studies. The potential impact of ‘unclear’ assessments of risk of bias will depend on how these ‘unclear’ assessments have been judged or interpreted when assessing risk of bias or concerns regarding applicability. For example, if all ‘unclear’ assessments were ignored when judging risk of bias, then the risk of bias for the whole body of evidence may be judged to be overly optimistic.
Although contacting study authors is not a requirement of the review process, the successful identification of additional data can be regarded as a strength of the review process.
12.6.2.3
Weaknesses dueto thereview analyses
Although meta- analysis in general may result in more precise estimates than analysis of individual primary studies, a small number of included studies or small number of par­ticipants may still jeopardize the precision of the results of the review. This especially holds when substantial heterogeneity is observed, and when sources of heterogeneity cannot be explored, letalone explained. In the assessment of the strength of the evi­dence, this is taken into account when considering imprecision and inconsistency.
With respect to heterogeneity and spectrum effects, review authors should make a priori hypotheses about possible differences in accuracy between subgroups. The Discussion section is the place to put these differences into context. Review authors may be able to discuss how consistent results are across different clinical settings and in different groups of individuals. This allows readers to judge to what extent summary estimates of test accuracy can be applied to different clinical settings.
If review authors find apparent differences in accuracy between subgroups, they should decide whether or not these effects are credible and relevant to readers. These differences should be considered carefully, as variation in results due to chance may play a role. It may be important to consider whether the variable was defined a priori; whether the difference in subgroups was seen within studies rather than between stud­ies; and whether there is additional evidence to support the relevance of findings in a specific subgroup. Cochrane Reviews are written for an international audience, and the discussion should not be limited to the applicability of results to a single healthcare organization or country.
12.6.2.4 Direct andindirect comparisons
Reviews with a comparative question may base their conclusions on direct compari­sons (the studies included in the review all directly compare the accuracy of different index tests) or on indirect comparisons (primary studies evaluated one of the index tests) or on both. As explained in Section12.3.1.6, direct comparisons are ideal because they are less prone to bias due to confounding.
When both direct and indirect comparisons are included in the same review, it should be clear that they are separate analyses and review authors should discuss any differ­ences in results. In addition to potential for bias associated with indirect comparisons (Takwoingi 2013), review authors should acknowledge that estimates of test accuracy derived from indirect comparisons are typically based on a greater number of studies than direct comparisons, leading to more statistical power to detect differences between tests, if such differences exist, and to generate more precise estimates.
368