Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_2767_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
25 Мб
Скачать
9.4 Fitting hierarchical models
https://t.me/medicina_free
2
when no covariates are included: μA, μB, σ
2
, σ
and ρAB. (Note: we follow Harbord (2007)
A
B
in using μ where Reitsma (Reitsma 2005) used θ in order to avoid confusion with the notation from that of the HSROC model that follows.)
The inclusion of a correlation parameter in the model allows for the expected trade­off in sensitivity and specificity. Where variation between studies arises through a trade­off resulting from variation in thresholds between studies, this correlation is expected to be negative. However, the correlation may be positive if there are other sources of heterogeneity.
Reitsma and colleagues originally proposed fitting these models by approximating the binomial within- study distributions by normal distributions (Reitsma 2005). Although this allows the model to be fitted in a slightly larger range of software (e.g. the MIXED procedure in SAS), Chu (2006) later demonstrated that the approximation can perform poorly and recommended that software be used that can explicitly model the binomial within-
study distributions, as is done in the examples that follow and the sam-
ple programs provided in Chapter10.
Meta- analysis of predictive values is possible using the bivariate method (Leeflang
2012), but this is not recommended as it is known that predictive values depend on the proportion of participants with the target condition, which is likely to vary between studies. Hence, the average predictive values will relate to use of the test at some aver­age prevalence. Meta- analysis of predictive values may be appropriate when either sensitivity and specificity is not estimable due to partial verification (i.e. verification of only test positives or test negatives) or when there is differential verification.
9.4.2 Example 1 continued: anti- CCP forthe diagnosis ofrheumatoid arthritis
We now undertake the first stage of a formal statistical analysis of the data from a review of anti- cyclic citrullinated peptide antibody (anti- CCP) (Nishimura 2007). If it can be presumed that the anti- CCP test is deemed positive if any anti- CCP antibody is detected and that detection can be considered a common threshold, it is appropriate to focus on summary estimates for sensitivity and specificity.
As noted in the descriptive analyses of these data (Section9.2.3), there appears to be variability across studies in estimates of test accuracy. This variability appears to be higher for sensitivity than specificity, which could arise either through heterogeneity or through estimates of sensitivity being based on smaller samples than estimates of specificity. Using the bivariate model to estimate a summary point based on the data for all studies, the parameter estimates from the bivariate model are shown in Table9.4.a.
The parameter estimates can be input to RevMan to produce the summary point, 95% confidence region and 95% prediction region shown in Figure9.4.a, superimposed on the individual study estimates. Computation of confidence and prediction regions also requires the standard error of the estimates for mean logit(sensitivity), mean logit(specificity) and the covariance between these estimates, which are 0.1275, 0.1459 and −0.00741, respectively. Note that the covariance shown in Table9.4.a is the covari­ance between observed estimates of logit(sensitivity) and logit(specificity) across stud­ies, not the covariance between the estimates for mean logit(sensitivity) and mean logit(specificity). Thelatter can be extracted as demonstrated in Chapter10.
217
9 Understanding meta- analysis
1
Sensitivity
Specificity
0
https://t.me/medicina_free
Table9.4.a Bivariate model parameter estimates foraccuracy ofanti- CCP forthe diagnosis of
rheumatoid arthritis
Label Parameter Estimate Standard error
Mean logit(sensitivity)
Mean logit(specificity) μ
Variance of random effects for logit(sensitivity) σ
Variance of random effects for logit(specificity) σ
Covariance of logit(sensitivity) and logit(specificity) σ
0.9
0.8
0.7
0.6
0.5
0.4
0.3
0.2
μ
A
B
2
A
2
B
AB
0.6534 0.1275
3.1090 0.1459
0.5426 0.1463
0.5717 0.1873
−0.2704 0.1199
0.1
0
1 0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1
Figure9.4.a Summary sensitivity and specificity of anti- CCP for the diagnosis of rheumatoid arthritis
The variance coefficients indicate similar heterogeneity in sensitivities and specifici­ties on the logit scale. The magnitude of the heterogeneity in sensitivities and specifici­ties (untransformed) is evident in the size of the prediction region on the SROC plot, with the variability in specificity being constrained on this scale because the underlying specificity is high. The summary estimates of sensitivity and specificity are shown by the solid black dot. The sensitivity and specificity at this point can be computed by inverse transformation of the logit estimates to give a sensitivity and specificity of 0.66 and 0.96, respectively. Confidence intervals can be computed by inverse transformation of the confidence intervals computed on the logit scale. (See Chapter 10 for further details.)
218
9.4 Fitting hierarchical models
sd
exp
https://t.me/medicina_free
The plot shows a potential outlier, with a sensitivity of 0.63 and specificity of 0.65. A sensitivity analysis can be performed by omitting this study to assess its influence on the summary estimates (see Section9.4.9).
9.4.3 The Rutter andGatsonis HSROC model
The HSROC model proposed by Rutter and Gatsonis (Rutter 1995, Rutter 2001) is based on a latent scale logistic regression model (McCullagh1980, Tosteson 1988). The HSROC model assumes that there is an underlying ROC curve in each study with parameters α and β that characterize the accuracy and asymmetry of the curve.
Accuracy, defined in terms of the lnDOR (natural logarithm of the diagnostic odds ratio), determines the position of the summary curve relative to the top left corner of the ROC axes. Each study contributes data at a single threshold to the analysis. The 2×2 table for each study then arises from dichotomizing at a positivity threshold denoted by θ. Theparameters α and θ are assumed to vary between studies: both are assumed to have normal distributions, as in conventional random-
effects meta- analysis.
The HSROC model can be regarded as having two levels corresponding to variation within and between studies. Atthe lower level, the number of diseased individuals who test positive is denoted by yi1 for the ith study, and the corresponding number of non- diseased who test positive is denoted by yi2. For each study (i), the number test­ing positive in each disease group (j) is assumed to follow a binomial distribution such that yij~B(nij, πij), j=1, 2, where nij and πij, respectively, represent the total number tested and the probability of a positive test result. The number testing positive in each diseased and non- diseased pair is analysed jointly within each study at the lower level in the analysis.
The model takes the form
where disij represents the ‘true’ disease status (coded as −0.5 for the non- diseased and
0.5 for the diseased), therebytaking into account the within- study variability at the lower level. Using the terminology for this model, θi represents a proxy for positivity threshold calculated as the mean of the log odds of a positive test result for the dis­eased and the log odds of a positive test result for the non- diseased groups in study i. αi (the lnDOR for study i) represents a measure of diagnostic accuracy in the ith study that incorporates both sensitivity and specificity for that study. The shape (scale) parameter (β) provides for asymmetry in the SROC curve by allowing accuracy to vary with thresh­old. Since each study contributes only one estimate of sensitivity and specificity at a single threshold, it is necessary to assume that the shape of the true underlying ROC curve in each study is the same, and hence β is fitted as a fixed effect.
The threshold and diagnostic accuracy for each study are specified as random effects and are assumed to be independent (uncorrelated) and normally distributed. The accuracy parameter has mean Λ (capital lambda) and variance , while the positivity (threshold) parameter has mean Θ (capital theta) and variance . The shape parameter (β) is estimated using data from the studies considered jointly, assuming normally distributed random effects for test accuracy. When no covariates are included, the HSROC model has five parameters: Λ, Θ, β, and .
ogit di
ij iiij ij
is
219
9 Understanding meta- analysis
ee
1
Sensitivity
Specificity
0
https://t.me/medicina_free
An SROC curve can be constructed from the HSROC model by choosing a range of values of 1–specificity and using the estimated average location parameter (Λ) and scale parameter (β) to compute the corresponding values for sensitivity. The average sensitivity at a chosen false positive fraction (1–specificity) is given by
ensitivity logitspecificityexp
When β=0, test accuracy can be summarized by Λ, which represents a common aver­age accuracy (lnDOR) across all thresholds, and the resulting summary curve will be symmetrical.
9.4.4 Example 2: Rheumatoid factor asa marker forrheumatoid arthritis
In this example we will investigate the diagnostic performance of rheumatoid factor (RF) as a marker for rheumatoid arthritis (RA). The 50 studies included in the analysis are taken from the same review as Example 1 (Nishimura 2007). The reference standard was again based on the 1987 revised American College of Rheumatology (ACR) criteria or clinical diagnosis.
The threshold for test positivity for RF varied between studies and ranged from 3 to 100 U/mL. The variability in threshold used to define test positivity between studies is reflected in the variability in study- specific estimates of sensitivity and specificity in the SROC plot shown in Figure9.4.b. Because of the variation in threshold across studies, asummary point does not have a useful interpretation. An SROC curve is appropriate to summarize these data and can be estimated by fitting an HSROC model. The
0.9
0.8
0.7
Figure9.4.b SROC plot and SROC curve for accuracy of rheumatoid factor
220
0.6
0.5
0.4
0.3
0.2
0.1
0
1 0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1
9.4 Fitting hierarchical models
https://t.me/medicina_free
Table9.4.b HSROC parameter estimates forrheumatoid factor model
Label Parameter Estimate Standard error
Mean accuracy
Mean threshold Θ −0.4370 0.1469
Shape of SROC curve β 0.2267 0.1624
Variance of random effects for accuracy
Variance of random effects for threshold
Λ
2.6016 0.1862
1.3014 0.3046
0.5423 0.1237
parameter estimates shown in Table9.4.b are extracted from the output of example programs for these data provided in Chapter10.
The parameter estimates can be input to RevMan to draw the summary curve as shown in Figure9.4.b; 2.6016 estimates the mean of the random effects for accuracy (Λ), −0.4370 estimates the mean of the random effects for threshold (Θ), 0.2267 esti­mates the shape parameter (β), 1.3014 estimates the variance of the random effects for accuracy, and 0.5423 estimates the variance of the random effects for threshold. All of these estimates are on the logit scale. The resulting curve shows the expected trade- off between sensitivity and specificity across thresholds. It is recommended that the fitted curve is displayed across the range of the observed study sensitivities and specificities, as shown here. The average sensitivity at a chosen specificity, or the other way round, can be computed from the fitted curve (see Chapter10).
When interpreting the results of the analysis, it is important to note that RF constitutes part of the ACR criteria. Hence, there is risk of bias in the estimated curve since the index test is incorporated in the reference standard. This could result in an over- estimation of the diagnostic accuracy of RF, and consequently a distorted picture of the value of using RF as a first test for resolving uncertainty in a suspected case of rheumatoid arthritis.
9.4.5 Data reported at multiple thresholds per study
The key approaches that are described and illustrated in this chapter focus on methods of analysis that use data from one threshold per study. To estimate a summary point, we use data from a common threshold across studies. To estimate an SROC curve, data from a range of thresholds are required to inform the shape of the underlying curve across studies.
Some studies may report sensitivity and specificity at more than one threshold. Using data from all available thresholds has the potential to provide more accurate estima­tion of SROC curves and estimates of average sensitivity and specificity values at stated thresholds. The potential gain will increase as the number of studies that report multi­ple thresholds increases.
A range of methods has been proposed for the analysis of such data. An early approach by Dukic (2003) summarizes study- specific ROC curves to obtain an SROC curve within a Bayesian framework. However, single- threshold studies cannot be included and cer­tain assumptions made in fitting the model can lead to sensitivities and specificities that are not monotonic (Hamza 2009). An alternative multivariate random- effects meta­analysis approach was proposed (Hamza 2009) that is applicable when there are
221
9 Understanding meta- analysis
https://t.me/medicina_free
specified thresholds of interest and all studies report sensitivity and specificity at these thresholds. This method can, in principle, be applied when studies do not provide data for all thresholds, but model convergence may be problematic in those circumstances. A subsequent approach based on survival analysis methods (Putter 2010) also requires that studies provide sensitivity and specificity at the specified thresholds of interest.
Recent approaches, including the methods of Steinhauser (2016) and Jones (2019), provide a more rigorous and robust approach for dealing with multiple thresholds per study. These methods can accommodate a variable number of thresholds and a range of different thresholds across studies. Studies that provide data for only one threshold can be included, thereby reducing possible bias resulting from the inclusion of an unrepresentative group of studies. Both methods allow for correlation between sensi­tivity and specificity across thresholds and also heterogeneity between studies through the inclusion of random study effects. An SROC curve is estimated and summary sensi­tivity and specificity can be computed at specified thresholds.
The Steinhauser method models the distribution of test results as a function of the continuous (or ordinal) thresholds for test positivity within both the diseased and non­diseased groups. Linear mixed- effects modelling is used to estimate the distribution parameters for a known underlying parametric distribution (usually assumed to be nor­mal logistic) of the test results in the two groups. The more recent method of Jones (2019) fits multinomial distributions to tables of categorized test results in each of the diseased and non- diseased groups. The number of categories (equal to the number of thresholds + 1) is allowed to vary across studies. The Jones model assumes an underly­ing logistic distribution for some transformation (for example, the natural logarithm) of the underlying continuous results in each group. The approach is flexible in that the best- fitting transformation from the set of Box- Cox transformations can be estimated from the data. Covariates can also be included in the model to investigate sources of heterogeneity. The methods by Steinhauser and by Jones are described in more detail in Chapter10, Section10.4, and illustrated using an example.
9.4.6 Investigating heterogeneity
In systematic reviews of test accuracy it is usual to observe variability in test accuracy between studies that is considerably greater than would be expected from within- study sampling error alone. This is reflected in the model specifications for the bivariate and HSROC models, which both allow for random study effects. For the bivariate model, the summary estimates of sensitivity and specificity represent an average operating point across studies. Similarly, the estimated SROC curve represents an underlying ROC curve across studies.
Some of this heterogeneity in test accuracy between studies is likely to arise because of differences in patient characteristics, test methods, study design and other factors. Exploratory analyses can be conducted to investigate whether such study characteris­tics appear to be associated with test accuracy using symbols and colours in SROC plots to distinguish between studies belonging to different subgroups.
Statistically, it is generally more efficient to make use of all of the data available across studies when investigating heterogeneity by adding study- level covariates to a hierar­chical model to identify factors associated with diagnostic test accuracy. This meta­regression approach also allows statistical inferences to be made. It is usually assumed
222
9.4 Fitting hierarchical models
vZ
https://t.me/medicina_free
that each covariate has a fixed effect when added to the model. This approach is also applicable to test comparisons, as discussed in Section9.4.7.
The bivariate and HSROC models differ in how study- level covariates are included. The bivariate method focuses on the estimation of summary estimates of sensitivity and specificity, and estimating how these values vary with study- level covariates. The HSROC approach, by contrast, focuses on the estimation of the SROC curve as the basis for assessing test accuracy, and investigating how the position and shape of the curve may vary with study- level covariates.
Both models allow the use of categorical and continuous covariates. In practice, covariates relating to study characteristics are usually categorical and indicator varia­bles are created as in standard regression modelling. For continuous covariates, par­ticular care should be taken to check that the assumption of linear associations is valid. For the bivariate model, this refers to association with logit(sensitivity) and/or logit(specificity). For the HSROC model, this refers to association with the accuracy parameter (lnDOR) and/or the threshold parameter.
The uses and limitations of investigating heterogeneity using subgroup analysis and
regression in Chapter10, Section10.11.5 of the Cochrane Handbook for Systematic
meta­Reviews of Interventions (Deeks 2019) apply equally to diagnostic studies.
9.4.6.1 Criteria formodel selection
Irrespective of which model is used, review authors should specify what modelling strategy will be used for adding or removing covariates and what criterion will be used to decide whether or not a covariate should be included in a model.
The decision as to whether a covariate should be retained in the model may be based in part on statistical tests. Commonly used software for fitting these models will provide P values for each estimate in the model based on Wald statistics. A P value based on the likelihood ratio Chi2 statistic is generally more reliable than the Wald statistic, especially for small sample sizes (Agresti2007). The likelihood ratio Chi2 statistic is computed as the change in the −2Log likelihood when a covariate is added (or removed) from a model, with the degrees of freedom equal to the difference in the number of parameters fitted in these models. The effect of adding (or removing) covariates on measures of model fit such as Akaike’s information criterion (AIC) or the Bayesian information criterion (BIC) can also be used. The deviance information criterion (DIC) is commonly used for models fitted by Markov chain Monte Carlo (MCMC) simulation. (See Chapter10 for further details.)
Statistical tests can also be used to assess whether allowing for variance of the ran­dom effects to vary by test in a comparison of two or more index tests provides a better­fitting model (see Chapter10).
9.4.6.2 Heterogeneity andregression analysis using thebivariate model
The bivariate model allows covariates to affect summary sensitivity or summary speci­ficity, or both. Using the notation of Harbord (2007), and assuming that we have a single study- level covariate Z that may affect both sensitivity and specificity, then the model can be extended as follows:
Ai
Bi
AAi
N
vZ
BBi
223
9 Understanding meta- analysis
https://t.me/medicina_free
As before, ∑ represents the covariance matrix for the random effects for logit(sensitivity) and logit(specificity). If the covariate does explain some of the heterogeneity in sensitivity and/or specificity, then we would expect the estimated variance for one or both random effects to be reduced. The estimated covariance (correlation) parameter may also change.
Assuming that we have a binary study- level covariate (Z) coded as 0 or 1 to represent the two groups of studies, then μA estimates the logit(sensitivity) at the summary point for the referent group (Z = 0), and μA+vA estimates the logit(sensitivity) at the summary point for the other group (Z = 1). Hence, exp(vA) estimates the odds ratio for sensitivity in group1 relative to the referent group. The average sensitivity is estimated as exp(μA)/ (1+exp(μA)) for the referent group of studies, and as exp(μA+vA)/(1+exp(μA+vA)) for the other group. Comparisons of specificity between the two groups of studies follow the same approach as described earlier based on μ
and vB. The fit of the model, with and
B
without the additional parameters vA and vB, can be used to assess whether the covariate is associated with sensitivity and/or specificity. This joint test will have 2 degrees of freedom if Z is binary. Separate tests of statistical significance of the covariate with sen­sitivity and specificity can also be conducted, first to assess whether vAdiffers from 0 (astatistically significant result indicates that there is evidence that sensitivity differs between the two groups of studies) and secondly whether vBdiffers from 0 (a statistically significant result indicates that there is evidence that specificity differs between the two groups of studies). See also Section9.4.6.1 relating to criteria for model selection.
The standard error of a new estimate that is a function of the model parameter estimates can be obtained using the delta method, on the assumption that the error distribution of the new estimate is approximately normal (Oehlert1992). The delta method is imple­mented in standard statistical software packages such as SAS and Stata (see Chapter10).
The bivariate model is easily extended to allow for more than one covariate. However, this may not be feasible in practice unless the number of studies is large. Typically only one source of heterogeneity can be investigated at a time. Also, it is important to note that a covariate may only be associated with sensitivity and not specificity, or the other way round. It is not required that the same covariates are fitted for both sensitivity and specificity, although this may commonly be the case. Where a covariate (or covariates) is allowed to affect both the sensitivity and the specificity, thebivariate model is equiv­alent to an HSROC model in which the covariate (or covariates) is allowed to affect both the accuracy and the positivity threshold but not the shape parameter. However, using the estimates from the bivariate model to test for the effect of covariates on the shape and position of the SROC curve is not straightforward. Using the HSROC model para­metrization allows this to be done in a more direct and straightforward manner.
It is usually assumed that the variance of the random effects (and their correlation in the case of the bivariate model) is not associated with the covariate. This is probably a reasonable assumption in most analyses investigating heterogeneity in test accuracy for a single index test. However, for analyses that compare different index tests, this assumption is less likely to hold (see Section9.4.7).
9.4.6.3 Example 1 continued: Investigation ofheterogeneity indiagnostic performance ofanti- CCP
The studies included in the review to assess the diagnostic performance of anti- CCP used two different generations of the assay: first generation (CCP1, 8 studies) and
224
9.4 Fitting hierarchical models
AB
A
B
https://t.me/medicina_free
second generation (CCP2, 29 studies). A binary covariate for test version (i.e. CCP1 or CCP2) with separate coefficients for sensitivity and specificity was added to the model. The covariate was coded as 0 for CCP1 (the referent group) and 1 for CCP2. Allowing both sensitivity and specificity to vary by test version in the model resulted in a –2Log likelihood of 533.4, a reduction of 12.2 compared with the model that contained no covariates. Hence, there is statistical evidence (Chi2 = 12.2, 2 df, P = 0.002) that test accu­racy is associated with the version of test used, but further investigation is required to ascertain whether this association is for sensitivity, specificity or both. The P values in Table9.4.c are based on Wald statistics for each parameter estimate, adjusted for the other variables in the model. Based on these, there is strong evidence that sensitivity is associated with test version (P = 0.0005), but not specificity (P = 0.21).
The variances of the random effects for logit(sensitivity) and logit(specificity), and
their covariance (
and σAB, respectively), are assumed to be common for both gen­erations of CCP. For the referent group (CCP1in this case) thesummary estimates for logit(sensitivity), logit(specificity) and the corresponding standard errors are denoted by μA and μB, respectively). The covariate parameter estimates (νA and νB) give the change in logit(sensitivity) and logit(specificity) for CCP2 relative to CCP1 – thus the logit(sensitivity) and logit(specificity) estimates for CCP2 are obtained by adding the covariate parameter estimates to those of the referent test (μA+νA and μB+νB, respec­tively). The standard errors for the logit(sensitivity) and logit(specificity) for CCP2 are most easily computed by refitting the model and defining CCP2 to be the referent group (covariate coded as 0 for CCP2 and coded as 1 for CCP1). The model parameter esti­mates can be input to RevMan to display the summary points and corresponding confi­dence regions shown in Figure9.4.c.
The summary point estimates and corresponding 95% confidence regions shown in the figure are consistent with the conclusion that the sensitivity varies by test type, but not specificity. Based on the model parameter estimates in Table9.4.c, the summary estimates of specificity were 0.97 (95% CI 0.95 to 0.98) for CCP1 and 0.95 (95% CI 0.94 to
0.97) for CCP2. The summary estimates of sensitivity were 0.48 (95% CI 0.37 to 0.58) for CCP1 and 0.70 (95% CI 0.65 to 0.75) for CCP2. These results indicate an improvement in
Table9.4.c Bivariate parameter estimates forcomparison ofthe accuracy ofCCP1 andCCP2 forthe
diagnosis ofrheumatoid arthritis
Parameter Estimate Standard error P value
μ
A
μ
B
σ
AB
ν
A
ν
B
*These Wald statistics are ignored as they test whether μ
sensitivity = 50% and specificity = 50%, respectively.
−0.09653 0.2203
3.4467 0.2982 < 0.0001*
0.3598 0.1022 0.001
0.5399 0.1802 0.005
−0.1969 0.09836 0.05
0.9626 0.2513 0.0005
−0.4302 0.3377 0.21
= 0 and μB = 0, equivalent to testing hypotheses that
A
0.66
*
225
9 Understanding meta- analysis
1
Sensitivity
0
https://t.me/medicina_free
0.9
0.8
0.7
0.6
0.5
0.4
0.3
0.2
0.1
0
1 0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1
Legend
Generation: CCP1 Generation: CCP2
Figure9.4.c Summary estimates of sensitivity and specificity for CCP1 and CCP2with corresponding
95% confidence regions
Specificity
sensitivity (P < 0.001), without loss of specificity (P = 0.21) for CCP2 compared with CCP1. (The model could be simplified by removing the covariate for specificity. The resulting estimate for specificity would then be assumed to be the same for CCP1 and CCP2.) For final presentation of the results we would wish to estimate the difference in sensitivity and specificity together with 95% confidence intervals. This is covered in Chapter10.
Comparing the output from this model with that of the model with no covariates (see Section 9.4.2), it is clear that the variances of the random effects are smaller, particularly for sensitivity. Also, checks of the distributions of the random effects (notshown here) show that adjusting for the use of first- or second- generation anti- CCP tests results in distributions that more closely follow a normal distribution.
In these analyses, the variances of the random effects were assumed to be the same for logit(sensitivity) and logit(specificity) for both generations of the test. This assump­tion can be investigated by fitting additional models, as shown in Chapter10. For the example here, allowing the variances to differ did not make a substantive change to theestimates or their interpretation. Chapter10 also demonstrates how to compute the difference in sensitivity and also the difference in specificity, with corresponding confidence intervals, using the bivariate model from the example in Section9.4.7.3 for illustration.
For analyses based on a small number of studies (unlike the example shown here), confidence intervals around the estimated difference in sensitivity and/or specificity may be large. Hence, it is important not to rely purely on tests of statistical significance to conclude that there is a lack of evidence of a difference, as the confidence interval for
226