Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_2767_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
25 Мб
Скачать
9.1 Introduction
https://t.me/medicina_free
tests, or randomizing patients to undergo one of the tests. A within- study comparison of test performance between two independent but comparable groups of patients could also be considered for inclusion in such an analysis. Direct comparison has the advan­tage that each study acts as its own control, thereby dealing with potential confounding that may occur due to study characteristics. However, the number of comparative studies available may not be sufficient to allow such an analysis.
9.1.5 Planning theanalysis
Undertaking meta- analyses for a Cochrane Review of diagnostic test accuracy involves first developing an analysis plan. Some of these decisions can be made at protocol stage (see Chapter1), others only after the data have been extracted from the reports. The established and commonly used statistical methods for test accuracy meta-
analysis use accuracy data for one threshold per study per index test. The following sections and examples used in Section9.4 follow this approach. More advanced methods for dealing with data from multiple thresholds per study are discussed in Section9.4.5 and dealt with in more detail in Chapter10.
The planning stages can be organized as follows.
Clearly stating the main questions that need answering, specifying which tests require estimates of test accuracy, and which tests should be compared with each other.
Detailed planning of the way in which comparisons will be made, identifying the different tests or groups of tests that can be compared, the multiple and pairwise comparisons that will be made, and the studies and data that will be included in each analysis. Review authors need to decide whether comparative analyses should include all studies, or be restricted to studies that evaluate both tests. Covariates for any heterogeneity analyses similarly need to be specified and coded.
From these decisions, a list of the planned main analyses, test comparisons and het­erogeneity analyses will be produced. The quantity of data that are available for each analysis should be determined to guide the choice of analysis method, and to assess whether adequate data are available for planned heterogeneity analyses.
Results can be plotted on forest plots and ROC plots to report study- specific esti­mates of sensitivity and specificity and the variability between these estimates across studies.
A strategy needs to be developed to deal with the mixed reporting of thresholds that may occur across studies. A key issue is deciding whether an analysis should be restricted to studies that share a common threshold value (which allows estimation of the summary sensitivity and specificity of a test at that threshold) or to include all studies regardless of threshold value (which allows estimation of SROC curves but compromises the interpretation of summary points). This will depend on the thresh­olds at which tests were evaluated in the primary studies, and knowledge of how the tests are applied in clinical practice. Consideration should also be given to whether a more complex analysis is feasible that uses multiple thresholds per study (seeSection9.4.5 and Chapter10, Section10.7).
If this analysis plan has been created using RevMan, the data should be exported from RevMan to the chosen statistics package, and appropriate models fitted. Results should be collated and tabulated as required, and parameter estimates copied back
207
9 Understanding meta- analysis
https://t.me/medicina_free
into the RevMan graphics function to produce final graphical output showing summary points or SROC curves as appropriate.
9.2 Graphical andtabular presentation
A Cochrane Review of diagnostic test accuracy uses two main forms of graphical display: coupled forest plots and SROC plots. Review authors can create these figures within RevMan for each analysis that is specified.
9.2.1 Coupled forest plots
Forest plots for diagnostic test accuracy report the number of true positives and false negatives in participants with the target condition (diseased), and true negatives and false positives in participants who do not have the target condition (non- diseased) in each study, and the estimated sensitivity and specificity, together with confidence intervals (CIs). The plots are known as coupled forest plots as they contain two graphi­cal sections: one depicting sensitivity and one specificity (Figure9.2.a). The order of the studies can be sorted, for example by values of sensitivity, or grouped by test type or covariate values. While it is possible to observe heterogeneity in sensitivity and specific­ity individually on such plots, it is not easy to visualize whether there are threshold- like relationships. Summary statistics computed from meta- analyses are rarely added to coupled forest plots. In Cochrane Reviews of diagnostic test accuracy, an archive of cou­pled forest plots for all the tests for which data were entered into RevMan is published with the review to make the 2×2 tables of the index test result for those with the target condition (referred to as the diseased) and without the target condition (referred to as non- diseased) readily accessible.
9.2.2 Summary ROC plots
An SROC plot is a scatterplot of the results of individual studies in ROC space where each study is plotted as a single (specificity, sensitivity) point. The size of the symbol (e.g. rectangle, ellipse) used to mark each point can be controlled to depict the preci­sion of the estimate (typically scaled according to their sample sizes). Using sample- size weights, the height of the symbol is proportional to the number of diseased (and hence the precision of the sensitivity estimate) and the width is proportional to the number of non- diseased (and hence the precision of the specificity estimate) (seeFigure9.2.b).
Both the within- study sampling variability and the heterogeneity between studies contribute to the total variability between studies. Even if the summary plots scale the symbols to indicate the precision of the estimates from individual studies, it is difficult to distinguish visually between these two sources of variability.
Two types of meta- analytical summary can be added to an SROC plot: SROC curves and summary points. Confidence regions for the summary points can be included, as can prediction regions that give an indication of between- study heterogeneity (see later Figure9.4.a).
Studies can also be plotted using different symbols and/or colours to indicate differ­ent subgroups for investigations of heterogeneity or for test comparisons.
208
Study
A
otsuka 2005
Bas 2003 Bizzar Bombar Choi 2005 C De Du Fe Gar Gir Goldbach-Mansk Gr Gr Hitchon 200 J K K Kw Lee 2003 Lopez-Ho Nell 2005 Nielen 200 Quinn 2006 Rantapaa-D Raza 2005 Sa S Schellek Soder Suzuki 2003 Va va va Vincent 2002 Vit Zeng 200
1
TP FP FN TN Generation Sensitivity SensitivitySpecificity Specificity
https://t.me/medicina_free
o 2001
dieri 2004
orrea 2004
Rycke 2004
bucquoi 2004
rnandez-Suarez 2005
cia-Berrocal 2005
elli 2004
einer 2005 ootenboer-Mignot 2004
ansen 2003
amali 2005 umagai 2004
ok 2005
raux 2003
auerland 2005
lin 2004
llbracht 2004 n Gaalen 2005 n Venroooij 2004
tecoq 2004
Figure9.2.a Coupled forest plot of the sensitivity and specificity of Anti- CCP for the diagnosis of rheumatoid arthritis
y 2000
4
yos 2004
5
ahlqvist 2003
ens 2000
3
115
17
16
110
24
40
5
23
0
236
20
74
11
89
4
90
2
31
0
69
8
25
2
43
1
70
5
167
8
26
8
110
3 26 64 71 68 38 42
149 147
47 24 40
171
72
481 190
82
865 139
69 90
1
14
2
14
3
2
7
10
7
3
11 26 14
7
2 23 12 13 79
7
5
7
86 58
88
29 50 22 18 10 63 17 98 15
148
20 15 58 35
60
109
35 20 18 46 60 77
68
105
71 252 101 107 101
73 215 227
7
39 231
8
130 142 129
75
38
40 120 228
88
15 118
56 293
66 132
0
73
96 114 106 375
79 146 443 298
9
51 185 408 301
2218
464 133 313
CCP2 CCP1 CCP1 CCP2 CCP2 CCP2 CCP2 CCP2 CCP2 CCP2 CCP2 CCP1 CCP2 CCP2 CCP2 CCP1 CCP2 CCP2 CCP2 CCP2 CCP2 CCP2 CCP2 CCP2 CCP2 CCP2 CCP1 CCP2 CCP1 CCP2 CCP2 CCP2 CCP2 CCP2 CCP1 CCP2 CCP1
0.88 [0.81, 0.93]
0.56 [0.49, 0.63]
0.41 [0.31, 0.51]
0.77 [0.58, 0.90]
0.73 [0.68, 0.78]
0.90 [0.82, 0.96]
0.75 [0.67, 0.83]
0.64 [0.56, 0.72]
0.58 [0.44, 0.72]
0.79 [0.69, 0.87]
0.71 [0.54, 0.85]
0.41 [0.31, 0.51]
0.80 [0.71, 0.88]
0.63 [0.57 0.69]
0.63 [0.47, 0.78]
0.43 [0.37, 0.49]
0.57 [0.41, 0.71]
0.81 [0.71, 0.89]
0.55 [0.46, 0.64]
0.66 [0.56, 0.75]
1.00 [0.91, 1.00]
0.41 [0.32, 0.51]
0.58 [0.51, 0.64]
0.81 [0.74, 0.86]
0.70 [0.58, 0.81]
0.57 [0.41, 0.72]
0.47 [0.36, 0.58]
0.74 [0.68, 0.80]
0.48 [0.40, 0.57]
0.44 [0.20, 0.70]
0.88 [0.85, 0.90]
0.64 [0.59, 0.70]
0.54 [0.45, 0.62]
0.77 [0.75, 0.80]
0.58 [0.51, 0.64]
0.39 [0.32, 0.47]
0.47 [0.40, 0.54]
0.81 [0.71, 0.89]
0.90 [0.85, 0.93]
0.98 [0.95, 0.99]
1.00 [0.91, 1.00]
0.92 [0.88, 0.95]
0.92 [0.86, 0.96]
0.97 [0.93, 0.99]
0.98 [0.95, 1.00]
1.00 [0.95, 1.00]
0.83 [0.69, 0.92]
0.95 [0.84, 0.99]
0.99 [0.95, 1.00]
0.98 [0.95, 0.99]
0.92 [0.84, 0.96]
0.65 [0.43, 0.84]
0.98 [0.93, 0.99]
0.98 [0.91, 1.00]
0.95 [0.92, 0.97]
0.97 [0.90, 1.00]
0.90 [0.84, 0.95]
0.96 [0.89, 0.99]
0.98 [0.93, 1. 00]
0.94 [0.88, 0.98]
0.91 [0.85, 0.96]
0.98 [0.96, 0.99]
0.96 [0.90, 0.99]
0.93 [0.88, 0.96]
0.94 [0.92, 0.96]
0.96 [0.93, 0.98]
0.96 [0.87, 1.00]
0.89 [0.84, 0.93]
0.97 [0.95, 0.99]
0.96 [0.93, 0.98]
0.97 [0.96, 0.97]
0.99 [0.97, 0.99]
0.96 [0.92, 0.99]
0.98 [0.96, 0.99]
0 0.2 0.4 0.6 0.8 10 0.2 0.4 0.6 0.8
9 Understanding meta- analysis
1
Sensitivity
Specificity
0
https://t.me/medicina_free
0.9
0.8
0.7
0.6
0.5
0.4
0.3
0.2
0.1
0
1 0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1
Figure9.2.b SROC plot of Anti- CCP for the diagnosis of rheumatoid arthritis
9.2.3 Linked SROC plots
Linked SROC plots are used in analyses of paired test comparisons, where both tests have been evaluated within a study. The points are plotted as for a routine SROC plot, but the two estimates (one for each test) from each study are joined by a line. It is thus possible to get a sense of the difference in accuracy between tests within each study, and to assess visually the degree of consistency in this difference across studies. Summary estimates of sensitivity and specificity for each test, as well as SROC curves obtained from meta- analysis, can be added to these plots (see later Figure9.4.g for an example plot).
9.2.3.1
Example 1: Anti- CCP forthe diagnosis ofrheumatoid arthritis– descriptive
plots
These data are taken from a review (Nishimura 2007) of anti- cyclic citrullinated peptide antibody (anti- CCP). The reference standard was based on the 1987 revised American College of Rheumatology (ACR) criteria for clinical diagnosis. The meta- analysis included 37 studies and their sensitivities and specificities are shown on the forest plot (Figure9.2.a); these study- specific estimates are also shown in a scatterplot in ROC space.
The forest plot shows the studies in alphabetical order. The figure gives the numbers for the 2×2 table (true positives (TP), false positives (FP), false negatives (FN), true nega­tives (TN)) for each study that will form the basis for statistical analyses. Study- specific estimates of sensitivity and specificity are shown, with their 95% confidence intervals. These estimates (and confidence intervals) are also represented graphically. The most
210
9.3 Meta- analytical summaries
https://t.me/medicina_free
striking features of this figure are the consistently high specificity and the greater uncer­tainty (indicated by the confidence interval width) and variability (indicated by the scat­ter of point estimates) in sensitivity than specificity. The studies can be ordered in different ways (e.g. in increasing order of sensitivity) to provide a visual representation of any association between sensitivity and specificity. Ordering studies by year of pub­lication may also be useful to identify trends in test accuracy resulting from changes in methodology and/or selection criteria over time (Cohen 2016). The figure also includes information on a covariate, the CCP generation, which may be associated with hetero­geneity in test accuracy. (This will be explored in Section9.4.6.2.)
The SROC plot shown in Figure9.2.b also illustrates the greater variability in esti­mated sensitivity than specificity across studies. Covariate information (e.g. generation of CCP) could be used to distinguish between studies in different subgroups (e.g. CCP1 vs CCP2). (See Section9.4.6.3 for further exploration of these data.)
Formal statistical analyses are required to obtain a summary estimate of test accuracy and to explore heterogeneity. These will be covered in Section9.4.6.3. Before proceeding to these statistical analyses, the review author should decide whether it is appropriate to focus on a summary point(s) or a summary curve(s) in the statistical analyses that follow. This will be influenced by the threshold(s) used by the included studies to define a posi­tive test result (see Section9.4).
9.2.4 Tables ofresults
Review authors need to construct additional tables to report results from their meta­analytical models. Review authors might consider creating tables to report the following.
The numbers of studies and individuals available for each of the key analyses.
Summary estimates of diagnostic accuracy for each test.
Statistics for the comparative accuracy and tests of statistical significance for the
pairwise comparisons between tests (a half- matrix display of all possible pairwise
comparisons may be useful). Separate tables based on direct (within- study) compari-
sons and uncontrolled comparisons may be needed (see Section9.4.7).
Results of investigations of heterogeneity, including estimates of test accuracy in sub-
groups, summary statistics of comparative accuracy, and tests of statistical signifi-
cance (see Section9.4.6).
Results of sensitivity analyses (see Section9.4.9).
This list is not exhaustive, and review authors should explore and identify the best ways of communicating the results of their analyses.
Cochrane Reviews of diagnostic test accuracy also include ‘Summary of findings’ tables, which are described in Chapter12.
9.3 Meta- analytical summaries
Test accuracy meta- analysis aims to compute the average accuracy of a test, compare estimates of accuracy between tests, and investigate the heterogeneity between studies. Review authors should choose which summary statistics to use. In Cochrane
211
9 Understanding meta- analysis
https://t.me/medicina_free
Reviews that use one 2×2 table per study (per test), the choice is between estimating average values of sensitivity and specificity for a test at a common threshold (referred to as the average or summary operating point or summary point), or estimating the underlying ROC curve for a test across many thresholds (referred to as the summary ROC curve or SROC curve). In most scenarios estimation of a summary point is pre­ferred, as it facilitates interpretation of the results in terms of the numbers of true positives, true negatives, false positives and false negatives (see Chapter 11 and Chapter12). However, there are situations when estimation of an SROC curve is the only analysis possible, or there are benefits in terms of precision of estimates and comparisons.
9.3.1 Should Iestimate anSROC curve or asummary point?
In a systematic review of test accuracy studies, it is likely that the collected data on test performance will have been measured at a mixture of different positivity thresholds. Athreshold could, for instance, represent a cut- point on a quantitative scale or a judge­ment on an ordinal scale. A key principle underlying the choice of statistical summary in meta- analysis of test accuracy is that the sensitivity and specificity of a test will vary as the positivity thresholds vary, as graphically depicted using an ROC curve. It is important to note that both hierarchical models recommended for meta- analysis for Cochrane Reviews of diagnostic test accuracy account for correlation between sensi­tivity and specificity observed across studies that is due to the functional relationship between sensitivity and specificity as the criterion for test positivity (threshold) varies. This correlation is accounted for in the analysis regardless of whether an SROC curve or a summary point is the output of choice. Separate meta- analysis of sensitivity and specificity estimates fails to account for the trade- off between sensitivity and specific­ity, which may lead to under- estimates of test accuracy (Deeks2001). Separate meta­analysis of likelihood ratios also ignores correlations between positive and negative likelihood ratios, and theoretically can produce estimates that are impossible (Zwinderman 2008).
While for some tests there is consensus about what value the positivity threshold should take, often tests are evaluated at different thresholds in different studies. Presentation of results at multiple thresholds within a single study is also possible, with some studies presenting estimates of ROC curves that depict the accuracy of the test at all possible thresholds. In addition, selective reporting of thresholds identified to opti­mize test accuracy can introduce bias, because sampling variability may influence which threshold is identified as ‘optimal’ when it is selected based on the observed data (Leeflang 2008).
Review authors need to decide whether they will use all the studies available and estimate an SROC curve, or select only studies that use a common threshold to esti­mate a summary point (an analysis that could be repeated across a set of different com­mon thresholds). Estimating summary sensitivity and specificity by combining studies that mix thresholds will produce an estimate that relates to some notional unspecified average of the thresholds that occur in the included studies. This must be avoided, because it is clinically uninterpretable and not generalizable. More complex methods do allow use of all data to estimate both curves and summary points, but are not currently often used due to challenges in their implementation (see Section9.4.5).
212
9.3 Meta- analytical summaries
https://t.me/medicina_free
In some contexts, test positivity is based on an ordinal rating scale rather than an explicit numerical threshold requiring a judgement rather than measurement, which is likely to lead to variability between individual readers in how they interpret the thresh­olds on the rating scale. For a study- level meta- analysis of test accuracy, it is generally reasonable to assume a common threshold across studies if the same point on the rat­ing scale was applied. At the study level, the estimated sensitivity and specificity repre­sent average estimates of accuracy across readers in that study. Even when it is possible to define a common threshold on the basis of a numerical value or a point on a rating scale, it must be acknowledged that some variability will remain in the actual threshold between studies through calibration differences between equipment, differences between groups of raters or observers in terms of training and experience, as well as variation in the implementation of tests. The consequence of such variability will be additional heterogeneity in test results observed at the common threshold. The sum­mary sensitivity and specificity point will reflect the average observed accuracy across studies, while the prediction region (see Section9.3.2) will reflect the heterogeneity between studies, as illustrated in the example in Section9.4.2.
Thus, the three main strategies used to handle mixed and variable thresholds in an analysis are as follows.
Estimating summary sensitivity and specificity of the test for a common threshold, or
at each of several different common thresholds if there are sufficient data. Each study
can contribute to one or more analyses depending on what thresholds it reports.
Studies that do not report at any of the selected thresholds are excluded.
Estimating the underlying ROC curve that describes how sensitivity and specificity
trade off with each other as thresholds vary. In this case, one threshold per study is
selected to be included in the analysis. A range of thresholds across studies is needed
to inform the shape of the SROC curve.
Simultaneous analysis of multiple (one or more) estimates of sensitivity and specific-
ity reported by each study. This approach allows estimation of an SROC curve and
summary points on the curve for particular thresholds of interest.
For analyses that will include one threshold per study, the choice of analytical approach will be influenced by the variation of thresholds in the available studies. For example, ifthere is little consistency in the thresholds used, meta- analyses that are restricted tocommon thresholds will contain very few data, and estimating an SROC curve may be preferred. If there is little variation in threshold between studies, attempting to fitan SROC curve will be difficult, as the points are likely to be too tightly clustered inROC space and there is little information in the data to estimate the shape of the underlying curve.
It can be reasonable to estimate both SROC curves and summary points in a review, as they may complement each other in providing clinically useful summaries and pow­erful ways of detecting effects. For example, separate analyses of test data at different thresholds may be used to provide clinically informative estimates of sensitivity and specificity based on the studies that provide data at each threshold. Including all stud­ies to estimate how SROC curves depend on covariates or test type will be the most powerful way to test hypotheses and investigate heterogeneity when thresholds vary across studies.
213
9 Understanding meta- analysis
https://t.me/medicina_free
When some studies provide data for multiple thresholds, more complex models that estimate an SROC curve and summary points on that curve may be applied (see Section9.4.5 and Chapter10, Section10.7). However, the potential gain relative to the more straightforward approaches will be limited if only a small proportion of studies provide such data.
9.3.2 Heterogeneity
Heterogeneity is to be expected in meta- analyses of diagnostic test accuracy. A conse­quence of this is that meta- analyses of test accuracy studies tend to focus on computing average rather than typical effects. In systematic reviews of interventions it is some­times noted that the estimates of the effect of the intervention in the different studies are very similar, the differences between them being small enough to be explicable by chance. In such situations it is appropriate to use a fixed-
effect approach meta- analysis, which estimates the underlying common effect (and is interpreted as the actual effect of the intervention). In systematic reviews of test accuracy, large differences are com­monly noted between studies, too big to be explained by chance, indicating that actual test accuracy varies between studies: there is heterogeneity in test accuracy. Random­effects meta- analysis methods are recommended when effects are heterogeneous. These methods focus on providing an estimate of the average accuracy of the test and describing the variability in accuracy between studies. In Cochrane Reviews of diagnos­tic test accuracy, heterogeneity is presumed to exist and random- effects models are fitted by default, only simplified to fixed- effect models where there are too few studies to estimate between- study variability, or analysis and forest or SROC plots demonstrate that fixed- effect models are appropriate (see Section9.4.8).
Univariate tests for heterogeneity in sensitivity and specificity and estimates of the I2 statistic (Higgins 2003) are not recommended for systematic reviews of test accu­racy, as they do not account for heterogeneity explained by phenomena such as posi­tivity threshold effects, whereby an increasing threshold for defining test positivity will decrease sensitivity and increase specificity. Problems with the I2statistic that could lead to misleading conclusions have also been highlighted in the literature (Rücker 2008, Wetterslev 2009, Zhou 2014); other measures such as the variance parameter of the random- effects model (Rücker 2008) and also multivariate and test accuracy- specific I2 statistics (Zhou 2014) have been proposed but are not used routinely.
The numerical estimates of the random- effects terms in the hierarchical models do quantify the amount of heterogeneity observed, but they are not easily interpreted as they represent variation in parameters expressed on log odds scales. Graphical displays provide a more straightforward means of assessing the magnitude of the observed het­erogeneity. For instance, if variation in threshold occurs in a meta- analysis, what mat­ters is the degree to which the observed study results in an SROC plot lie close to the SROC curve, not how scattered they are in ROC space.
For a summary point estimate, the inclusion of a prediction region around the point provides a visual assessment of heterogeneity. The region takes account of the variance of the random effects for logit(sensitivity) and logit(specificity), as well as the correla­tion between them. The region provides a visual summary of the spread of the true underlying test accuracy across the studies included in the random- effects model
214
9.4 Fitting hierarchical models
https://t.me/medicina_free
(seethe example in Section9.4.2). Hence, a 95% prediction region represents the region within which one has 95% confidence that the truesensitivity and specificity of any future study should lie (Harbord 2007). Estimation of a prediction interval or region relies on the assumption of normal distributions for the effects across studies. This may be very problematic when the number of studies is small and can lead to spuriously large (or small) regions (Deeks 2019).
9.4 Fitting hierarchical models
In this section, the bivariate model (Reitsma 2005) and the hierarchical SROC (HSROC) model of Rutter and Gatsonis (Rutter 2001) are described and discussed for the meta­analysis of studies each of which contributes one threshold to the analysis. These hier­archical models include random study effects that account for the unexplained heterogeneity between studies that is typical in systematic reviews of test accuracy. They supersede the earlier, more limited fixed- effect SROC approach of Moses and Littenberg (Littenberg 1993, Moses 1993).
Both the bivariate and Rutter and Gatsonis HSROC models involve statistical distribu­tions at two levels. At the lower level (level 1), they model the cell counts in the 2×2 tables extracted from each study using binomial distributions and logistic (log odds) transformations of proportions, thereby taking account of random sampling variability within studies. At the higher level (level 2), random study effects are assumed to account for heterogeneity in test accuracy between studies beyond that accounted for by sam­pling variability at the lower level.
The two models are mathematically equivalent when no covariates are fitted (Harbord 2007, Arends 2008), but differ in their parametrizations. The bivariate parametrization models sensitivity, specificity and the correlation between them directly, whereas the Rutter and Gatsonis HSROC parametrization models a function of sensitivity and speci­ficity to define an SROC curve. Given their shared statistical properties, in the absence of covariates in the models SROC curves can be computed from bivariate models and summary points from HSROC models.
If review authors are using RevMan, the parameter estimates from either model can be input to produce an appropriate graphical display in ROC space of the summary esti­mates of test accuracy for a particular analysis, usually superimposed on the scatter­plot of study specific estimates. The options are:
the SROC curve; or
the summary point (i.e. summary values for sensitivity and specificity);
a confidence region around the summary point; and
a prediction region around the summary point.
Although it is common to present 95% confidence and 95% prediction regions, RevMan offers options to use other values. The 95% prediction region around a summary point illustrates the extent of statistical heterogeneity by depicting a region within which, assuming the model is correct, we have 95% confidence that the true sensitivity and specificity of any future study should lie (Harbord 2007). There is also an option of plot­ting a 50% prediction region, in which the central half of the true values of future studies would lie (akin to an interquartile range).
215
9 Understanding meta- analysis
AA
AB B
https://t.me/medicina_free
Additional estimates can be derived from the models. Summary estimates for the positive and negative likelihood ratios and the diagnostic odds ratio, with correspond­ing confidence intervals, can be computed at the summary point or at any point on the SROC curve. From the SROC curve the average sensitivity at a given value of specificity (or the other way round) can also be computed (see Chapter10). Not all of these possi­ble summary measures will be relevant or appropriate for a given analysis.
The motivation for choosing one of these two alternative parameterizations becomes clear when covariates are added to explore heterogeneity in test accuracy or to com­pare tests. Ultimately, the choice of method will be determined by the focus one wishes to adopt, and which of the two models addresses the research question given the nature of the available data (see Section9.3.1).
Both models require the use of external statistical software. The results given for the examples included in this chapter have been estimated using frequentist methods. Although this is the most commonly used approach, Bayesian estimation can also be used, as illustrated in Chapter10. Publication-
ready graphical output can be created in
RevMan using the model parameter estimates to add model summaries to SROC plots.
Alternative specifications for summary curves based on functions of the bivariate model parameters have been proposed (Arends 2008, Chappell 2009). This chapter will focus on the Rutter and Gatsonis model, as it is the most established of the HSROC specifications.
9.4.1 Bivariate model
The bivariate method models sensitivity and specificity directly. The model can be regarded as having two levels corresponding to variation within and between studies. At the lower level, the within- study variability for both sensitivity and specificity is assumed to follow a binomial distribution. For sensitivity (denoted by A), the number testing positive yAi~B(nAi, πAi), where nAi and πAi respectively represent the total number of dis­eased individuals tested and the probability of a positive test result in that group in study i. Similarly, for specificity (denoted by B), the number testing negative yBi~B(nBi, πBi), where nBi and πBi respectively represent the total number of non- diseased individuals tested and the probability of a negative test result in that group in study i. The sensitiv­ity–specificity pair for each study must be modelled jointly within the study at the lower level of the analysis, because they are linked by shared study characteristics including the positivity threshold. At the higher (between- study) level, the logit- transformed sen­sitivities are assumed to have a normal distribution with mean μA and variance σ while the logit- transformed specificities have a normal distribution with mean μB and variance σ
2
. Their correlation is included by modelling both at once by a single bivari-
B
ate normal distribution
Ai
Bi
A
N
with
B
B
where σAB is the covariance between logit(sensitivity) and logit(specificity). The model may also be parametrized using the correlation ρAB= σAB/(σAσB), which may be more interpretable than the covariance. The bivariate model therefore has five parameters
216
2
,
A