Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_2767_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
25 Мб
Скачать
9.4 Fitting hierarchical models
ii
sZ
ii
NZ
ii
NZ
(a) R par curves crossing (𝛿 ≠ 0)
cross (𝛿 = 0)
1. 0
Sensitivity
0
1. 0
https://t.me/medicina_free
the estimated difference may cover a clinically important effect. In that situation, the results would be regarded as inconclusive.
9.4.6.4 Heterogeneity andregression analysis using theRutter andGatsonis HSROC model
The HSROC model allows covariates to be added to explore heterogeneity in test posi­tivity (threshold), position of the curve (accuracy) and shape of the curve. A covariate may be associated with some but not all three model parameters.
Assuming that we have a binary study- level covariate (Z) coded as 0 or 1 to represent the two groups of studies, then the HSROC model can be extended to estimate the log odds of a positive test for study i and disease group j as follows
t
ij ii iiij
ZZdi
disexp
j
where each of γ, ξ and δ is assumed to be a fixed effect. Hence, the distributions of the random effects for threshold and accuracy are now given by
and
, respectively. The shape parameter for the summary curves for the two groups is estimated as β for the referent group of studies (Z = 0) and β+δ for the other group (Z = 1). If the covariate does explain some of the heterogeneity in threshold and/or accuracy, then we would expect that the estimated variance for one or both ran­dom effects would be reduced.
The first step would be to investigate the shape of the summary curve. If δ≠0, then the shape of the summary curve differs for the two groups of studies, which means that the relative accuracy of the test for the two groups of studies will vary with threshold (Figure9.4.d (a)). This represents the most complex scenario, and the model would not generally be simplified any further. The power to detect differences in shape will be low
0.8
0.6
0.4
group 1
group 0
Sensitivity
group 1
group 0
0.8
0.6
0.4
0.2
0.0
0.0 0.2 0.4 0.6 0.8 1.0
elative accuracy depends on
ticular specificity values, the
Figure9.4.d SROC curves with and without a difference in shape
1-specificity
(b) Group 1 dominates across all specificity values, the curves do not
0.2
0.0
0.0 0.2 0.4 0.6 0.8 1. 1-specificity
227
9 Understanding meta- analysis
https://t.me/medicina_free
when the number of studies in either group is limited. Also, it is important when inves­tigating shape to consider the effect of outlying and potentially influential studies (see Section9.4.9). When there is statistical evidence that the curves differ in shape, a plot of the estimated curves for the two groups will aid in interpretation. Focusing on the region of the plot that covers the observed data, it is then possible to compare the esti­mated curves. Where one curve consistently lies above another in the region of primary interest, there is evidence of superior accuracy even though the separation between the curves will vary across thresholds. If the curves cross, then the interpretation of which curve shows superior accuracy will depend on threshold.
If, based on statistical evidence, similarity of curve shapes and investigation of poten­tially influential studies, it can be assumed that δ=0, then the covariate can be removed for shape. The estimated SROC curves for the two groups will then have the same shape, even though they may not be symmetrical (Figure9.4.d (b)), and the relative diagnostic accuracy of the two curves can be summarized using the relative diagnostic odds ratio (RDOR=exp(ξ)). The RDOR will be constant across all possible values of θ. If the model can be simplified further and both curves can be assumed to be symmetrical, i.e. β=0, the RDOR again provides a measure of relative accuracy as already described, but in addition the DOR in each group will be constant across thresholds.
If the curves can be assumed to have the same shape (either both asymmetrical or both symmetrical), then the
question is whether the covariate is associated with accu­racy, i.e. the position of the curve. If there is evidence that ξ≠0, then the RDOR gives an estimate of the overall relative diagnostic accuracy. This would correspond to a clear separation between the SROC curves for the two groups. Alternatively, ξ=0 implies that there is no separation between the curves and no association between the covariate and accuracy.
If ξ can be assumed to be 0, then the model can be further simplified by removing the covariate for accuracy, which will result in a single summary curve (assuming that the shape of the curve is the same for the two groups of studies). An association between the covariate and the threshold parameter (i.e. γ≠0) would indicate that the underlying test positivity rate for the two groups of studies differs. Such an association is often difficult to interpret unless the curves can be assumed to have the same shape and accuracy.
The RDOR is useful for the statistical comparison of two curves that have the same shape because the RDOR is constant across all values of the threshold parameter θ. However, it does not have a straightforward interpretation when the shapes of the curves differ. In that case, the estimated RDOR will represent the relative accuracy of the points on the curves where they intersect the diagonal line in ROC given by sensitiv­ity = specificity. One approach that can be used to aid interpretation is to compute theestimated sensitivities at a chosen value of specificity (or the other way round) to compare the curves at selected points that are of clinical importance, as illustrated in Chapter10.
9.4.6.5 Example 2 continued: Investigating heterogeneity indiagnostic accuracy ofrheumatoid factor (RF)
We will now investigate whether the laboratory technique used to measure RF is associ­ated with diagnostic performance. Of the 50 studies, 15 used nephelometry (N), 16latex agglutination (LA), 16 ELISA, one study used RA hemagglutination and 2 did not report
228
9.4 Fitting hierarchical models
q
https://t.me/medicina_free
the method used. The analysis is restricted to studies that used N, LA or ELISA. The HSROC model was again used because of the variation in threshold used for test posi­tivity across studies. In keeping with standard methods used in regression analysis, two covariates are defined to distinguish between the three techniques (indicator variables for N and ELISA to compare these groups of studies with LA, the referent group). These covariates were included in the model to assess whether accuracy, threshold or the shape of the SROC curve varied with technique. The variances of the random effects for threshold and accuracy are assumed to be common to all three techniques. (See Chapter10 for example programs and output.)
The - 2Log likelihood for the most complex model that included covariates for shape, accuracy and threshold parameters was 752.9. The increase in the - 2Log likelihood was negligible (an increase to 753.1) when the covariates for shape were removed from the
2
model (Chi
= 753.1 − 752.9 = 0.2, 2 df, P = 0.90). Parameter estimates for the model that assumes a common shape are given in Table9.4.d, and the corresponding estimated HSROC curves shown in Figure9.4.e. The estimates of alpha (Λ), theta (Θ) and beta (β) can be input to RevMan to obtain the summary curve for the referent group (LA).The threshold and accuracy parameter estimates for ELISA are given by Θ + γ Λ+ξ
, respectively, and for N are given by Θ+γN and Λ+ξN, respectively, on the logit
ELISA
ELISA
and
scale. The standard errors of these estimates for ELISA and N are most easily obtained by refitting the model twice, first defining ELISA as the referent group and then defining N as the referent group.
From Figure 9.4.e, it appears that LA may be less accurate than the other two
methods; however, removal of thecovariate for accuracy (i.e. coefficients ξ
ELISA
and ξN assumed to be 0) has a negligible effect on the fit of the model (Chi2 = 753.7 − 753.1 = 0.6, 2 df, P = 0.74), indicating no statistical evidence of a difference in diagnostic accuracy of RF according to technique. This is consistent with the P values based on Wald statistics in Table9.4.d. Hence, it is reasonable to fit a single SROC for RF, as there is no statistical evidence of a difference in test accuracy for the three techniques (see Chapter10 for model estimates for the single common summary curve). These results indicate that it
Table9.4.d HSROC parameter estimates tocompare RF techniques
Parameter Estimate Standard error P value
Λ
Θ −0.5490 0.2137 0.014
β 0.1995 0.1702 0.25
ξ
ELISA
ξ
N
γ
ELISA
γ
N
*This P value is ignored as it tests the uninformative hypothesis that the mean of the pseudo threshold
parameter is zero.
2.4552 0.3245 < 0.0001
*
1.2865 0.3109 0.0002
0.4786 0.1139 0.0001
0.2483 0.4408 0.58
0.3328 0.4439 0.46
−0.1962 0.2614 0.46
0.4960 0.2627 0.065
229
9 Understanding meta- analysis
1
Sensitivity
0
https://t.me/medicina_free
0.9
0.8
0.7
0.6
0.5
0.4
0.3
0.2
0.1
0
1 0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1
Legend
Method: ELISA Method: LA Method: Nephelometry
Figure9.4.e SROC curves to compare accuracy of RF techniques
Specificity
may be reasonable to treat the techniques as equivalent markers for RF (they have similar DOR, but the trade- off between sensitivity and specificity may still differ).
9.4.7 Comparing index tests
Many diagnostic reviews aim to compare the diagnostic accuracy of two alternative index tests that may be used to detect the same condition. In this section, the focus will be on the comparison of two index tests, but the approach can be extended to allow for more than two tests.
Two approaches are generally adopted for test comparisons. The first approach uses test accuracy data from all eligible studies that have evaluated one or both tests. The second approach restricts the analysis to studies that have either evaluated both tests in the same individuals, or have randomized individuals to undergo one or other of the two tests. The second approach has advantages because the comparison is less likely to be biased by confounding and hence these results should be relied upon where possi­ble (Takwoingi 2013). However, the number of studies that report such direct compari­sons is often very limited, which means that such an analysis may not be feasible or may only be considered as a sensitivity analysis (see Section9.4.9).
9.4.7.1 Test comparisons based onall available studies
Often, many of the available studies evaluate only one of the tests of interest (Takwoingi
2013). By using all studies that have evaluated at least one of the tests, we maximize the number of studies in the analysis. However, the studies are likely to be heterogeneous
230
9.4 Fitting hierarchical models
https://t.me/medicina_free
in terms of design and patient characteristics that are associated with test accuracy, and hence confounding may be an issue. In preliminary exploratory analyses this can be dealt with by comparing the tests within subgroups of studies that are homogene­ous with respect to important potential confounders such as study design or spectrum of disease. The value and feasibility of such exploratory analyses will be affected by the number of available studies and missing or inconsistent reporting across studies of information on potential confounders.
The statistical methods described in this section follow directly from the description of hierarchical models in Section9.4.6 and how they can be used to investigate hetero­geneity in test accuracy. For the comparison of two index tests, the type of test is repre­sented by a binary covariate, which is used to identify the test that gave rise to each 2×2 table included in the analysis. Confounders can potentially be adjusted for; however, this may be difficult to do in practice because the number of studies is often small and/ or data on important confounders may be poorly recorded or incomplete.
Both the bivariate model and the Rutter and Gatsonis HSROC model can be used to investigate the relative accuracy of two index tests. However, as noted previously, the choice of approach will be influenced by the nature of the available data. The interpre­tation of the results will depend on which approach is used.
Test comparisons using thebivariate model
9.4.7.2
If, for each index test, the available studies have used a consistent threshold on a continuous or ordinal scale to define test positivity, then the bivariate model provides an appropriate framework for test comparisons. It may also be reasonable to assume acon­sistent threshold when a test comprises a ‘test kit’ that produces positive and negative results (such as a coloured line appearing on a device). By adopting the same strategy described earlier (Section9.4.6.2), a binary covariate for test type can be included in the model to investigate whether sensitivity and/or specificity differs between two tests.
Care must be taken with the interpretation of the results of such a model, particularly if the common threshold for test positivity for either test is applied to a continuous or ordinal scale. Any inferences made about the relative diagnostic accuracy of the two tests is only valid at the chosen threshold for each of the two tests and cannot be extrapolated to other possible thresholds. Where other thresholds are reported, the analysis can be repeated using the available data to investigate the relative diagnostic accuracy of the tests at those alternative thresholds. However, such additional investi­gations should be restricted to thresholds reported by sufficient studies to allow a meaningful analysis.
Because we are analysing test accuracy data for two alternative index tests, it may not be reasonable to assume that the variances of the random effects for logit(sensitivity) and logit(specificity) are the same for the two tests. The bivariate model can be extended to allow the variance of the random effects for each test to depend on the covariate for test type (see Chapter10). This will also affect the estimated correlation between them. Statistically, estimation of the variances of the random effects for logit(sensitivity) and logit(specificity) and correlation between them is subject to a higher level of uncer­tainty than for the main parameters of interest. However, when preliminary plots of the study- level estimates of sensitivity and specificity in ROC space show marked differ­ences in heterogeneity between studies for the two tests, it is advisable to assess
231
9 Understanding meta- analysis
https://t.me/medicina_free
whether the assumption of equal variances of random effects for the two tests is rea­sonable. This can be done by comparing estimates for test accuracy between the alter­native models to assess whether conclusions about the relative sensitivity and/or specificity of the tests are robust to assumptions about the variances of the random effects. Such an investigation may not be feasible if the number of studies is small.
It is usual for most of the studies in such an indirect analysis of test comparisons to have evaluated only one of the tests, but some studies may have evaluated both. If the proportion of studies that have evaluated both is very small, then treating the results ofthe two tests in a study as if they were obtained from different studies is unlikely toaffect the results. Although this is often done in practice, such an approach is not rec­ommended if the proportion of studies evaluating both tests is not small, because it is likely to result in inappropriate standard errors for the test comparison parameters for sensitivity and specificity. In that case the paired sensitivity/specificity data for bothtests from each study should be at the lower level of the hierarchical analysis, anda binary covariate for test type included to identify which 2×2 table corresponds to each test.
9.4.7.3
Example 3: CT versus MRI forthe diagnosis ofcoronary artery disease
Schuetz (2010) evaluated the diagnostic performance of multi- slice computed tomog­raphy (CT) and magnetic resonance imaging (MRI) for the diagnosis of coronary artery disease (CAD). The review included prospective studies that evaluated either CT or MRI (or both), used conventional coronary angiography (CAG) as the reference standard, and used the same index test threshold for clinically significant coronary artery stenosis (a diameter reduction of 50% or greater). A total of 103 studies provided a 2×2 table for one or both tests and were included in the meta- analysis: 84 studies evaluated only CT, 14 evaluated only MRI and 5 studies evaluated both CT and MRI. (See Chapter10 for data and example programs.)
Because study selection was based on a common threshold for clinically significant coronary artery stenosis, the bivariate model was used for data synthesis and test com­parison. In the first stage of the analysis, we base our test comparison on all studies that evaluated at least one test. The approach follows closely the method illustrated in Section 9.4.6.3 for exploring heterogeneity using the bivariate model and the same notation has been used.
A binary covariate is added to the model, which is coded as 0 if the 2×2 table is for MRI (the referent group) and coded as 1 if the 2×2 table is for CT. The 5 studies that evaluated both tests contribute a 2×2 table for each test, hence there are 19 studies included for MRI and 89 studies included for CT. Allowing both sensitivity and specificity to vary by type of test resulted in a –2Log likelihood of 953.0, a reduction of 42.5 compared with the model that contained no covariates. Hence, there is statistical evidence (Chi 2 df, P < 0.001) that sensitivity and/or specificity are associated with test type. Removing the covariate for sensitivity from the model (Chi2 = 976.7 − 953.0 = 23.7, 1 df, P < 0.001) shows strong statistical evidence of a difference in sensitivity between the two tests. Similarly, removing the covariate for specificity from the model (Chi2 = 976.2 − 953.0 =
23.2, 1 df, P < 0.001) shows strong statistical evidence of a difference in specificity between the two tests. These results are consistent with the P values shown in Table9.4.e that are based on Wald statistics.
Table 9.4.e gives the estimated mean logit(sensitivity) and mean logit(specificity) (μ
and μB, respectively) for the referent category (MRI), the common variances of the
A
232
2
= 42.5,
9.4 Fitting hierarchical models
A
B
1
Sensitivity
0
https://t.me/medicina_free
Table9.4.e Bivariate model estimates forcomparison ofCT andMRI
Parameter Estimate Standard error P value
μ
A
μ
B
2.1771 0.2457
0.8754 0.2111 < 0.0001*
0.8749 0.2293 0.0002
0.8447 0.1696 < 0.0001
σ
AB
ν
A
ν
B
0.1803 0.1384 0.20
1.3033 0.2625 < 0.0001
1.0415 0.2154 < 0.0001
*These Wald statistics are ignored as they test whether μ
that sensitivity = 50% and specificity = 50%, respectively.
0.9
0.8
0.7
0.6
0.5
0.4
0.3
< 0.0001
= 0 and μB = 0, equivalent to testing hypotheses
A
*
0.2
0.1
Figure9.4.f Summary estimates of accuracy of CT and MRI for the diagnosis of coronary artery
disease with corresponding 95% confidence regions
random effects (A and B, respectively), their covariance (σAB) and the difference in mean logit(sensitivity) and mean logit(specificity) between CT and MRI (νA and νB, respectively). Estimates of the standard errors of mean logit(sensitivity) and mean logit(specificity) for CT are required for entry into RevMan and can be obtained easily from refitting the model with CT defined as the referent group.
Figure 9.4.f shows the SROC plot with summary points for MRI and CT and their
95%confidence regions superimposed as shown. The squares represent CT and the
0
1 0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1
Legend
CT MRI
Specificity
233
9 Understanding meta- analysis
https://t.me/medicina_free
diamonds represent MRI. Because of the large number of studies for CT, the summary point and region are difficult to see. The figure could be redrawn without the individual study points to show just the summary points and regions, or the size of the study points can be reduced to improve visibility.
Inverse logit transformation of the model estimates and their confidence limits pro­vide the relevant estimates in ROC space: summary estimates for sensitivity are 0.90 (95% CI 0.84 to 0.93) for MRI and 0.97 (95% CI 0.96 to 0.98) for CT. The summary esti­mates for specificity are 0.71 (95% CI 0.61 to 0.78) for MRI and 0.87 (95% CI 0.84 to 0.90) for CT. The P values for νA and νB show strong statistical evidence of an association between sensitivity and test type and also between specificity and test type.
Based on this analysis, there is strong evidence that CT has higher sensitivity and specificity than MRI for detecting clinically significant coronary artery stenosis, defined as a diameter reduction of 50% or more. Chapter10 demonstrates how to compute the difference in sensitivity and difference in specificity, with corresponding confidence intervals, for the two tests.
Further descriptive analyses and modelling may be undertaken to assess whether the assumption of equal variances for the random effects for the two tests has a substantive effect on the results presented here (see Chapter10).
9.4.7.4 Test comparisons using theRutter andGatsonis HSROC model
Comparisons of summary estimates of sensitivity (or specificity) of alternative tests can be misleading if the included studies have used different thresholds to define testpositivity. In this situation, comparisons based on SROC curves provide a more informative approach.
The hierarchical modelling strategy used to investigate heterogeneity described ear­lier for the Rutter and Gatsonis HSROC model (Section9.4.6.4) can be used for compari­sons of test accuracy when there is variability in threshold between studies. The type of test is represented by a binary covariate that is used to identify the test that gave rise to each 2×2 table included in the analysis. This covariate then allows the review authorto investigate whether test type is associated with the shape and position of theSROC curve. Interpretation of the results follows directly from the discussion of the interpretation of investigations of heterogeneity in Section9.4.6.
Statistically, estimation of the variances of the random effects for threshold and accu­racy is subject to a higher level of uncertainty than for the main model parameters of interest. If preliminary plots of the study-
level estimates of sensitivity and specificity in ROC space show marked differences in heterogeneity between studies for the two tests, it is advisable to assess whether the assumption of equal variances of the random effects for the two tests is reasonable (see Chapter10). This is usually done by compar­ing the fit of the alternative models (i.e. where variances do, or do not, depend on the covariate for test type). A comparison of the main estimates of interest between the alternative models is also useful to assess whether conclusions about the relative shapeand accuracy of the summary curves for the two tests are robust to assumptions about the variances of the random effects. Again, such an investigation will not be feasible if the number of studies is small.
As noted for the bivariate model, it is usual for most of the included studies to have
evaluated only one of the tests, but some studies will have evaluated both. If the
234
9.4 Fitting hierarchical models
1
Sensitivity
0
https://t.me/medicina_free
proportion of studies that have evaluated both is very small, then treating the results of the two tests in a study as if they were obtained from different studies is unlikely to affect the results. However, more accurate standard errors will be obtained for the test comparison parameters if the data for both tests are modelled within the study at the lower level in the analysis.
9.4.7.5 Test comparison based onstudies that directly compare tests
As noted in Section9.1.3 and Section9.3.2, heterogeneity in the estimated accuracy of a diagnostic test across studies is likely to occur. This could confound the comparison of two tests if different studies are used to estimate the diagnostic accuracy of each test. Ideally, the comparison should be based on studies that have made a direct compari­son of the tests of interest, either by applying both tests to each study participant, or by randomizing each individual to receive one of the tests (Takwoingi 2013). A common reference standard should be applied to both tests. If there are sufficient studies of thistype on which to base a test comparison, the results are less prone to bias than an analysis based on all available studies that have evaluated one or both tests.
A preliminary graphical analysis can be conducted by plotting the estimated sensitiv­ity and specificity for both tests for each study in ROC space. The two points contributed by each study (one for each test) are joined by a line to highlight the relative test accu­racy within each study to illustrate the pairing of test accuracy estimates at the study level (see Figure9.4.g).
The rationale described earlier for choosing between the bivariate model and the HSROC model when making test comparisons is also applicable here, and the same
0.9
0.8
0.7
0.6
Figure9.4.g SROC plot of direct comparisons of CT vs MRI for the diagnosis of coronary artery
disease
0.5
0.4
0.3
0.2
0.1
0
1 0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1
Legend
CT MRI
Specificity
235
9 Understanding meta- analysis
https://t.me/medicina_free
issues relating to interpretation apply. The only major difference is that the analysis does not include any studies that have evaluated only one of the tests.
Because each study contributes a 2×2 table for each of the two tests to be compared, the data for the two tests should be analysed within the study at the lower level in the analysis. The random effect for that study is assumed to be common to both 2×2 tables for that study. A binary covariate for test type is included to identify which 2×2 table corresponds to each test. Entering a separate 2×2 table for each test (within each study) for analysis in a hierarchical model effectively assumes that the data arise from a randomized or independent groups design. This represents a conservative approach that is often necessitated by the lack of information on paired results at the individual level for truly ‘paired’ studies that have applied both tests to the same individual. Atpresent, it is not common practice for researchers to publish a cross-
classification of test results within both the diseased and non- diseased groups, or to provide individual patient data for analysis that would provide full information on the pairing of test results within an individual.
Meta- analytical models that account for pairing of test results within an individual within each study have been developed as an extension of the bivariate model. The method proposed by Trikalinos (2014) allows for comparisons of two tests based on studies that use a paired design and report fully cross- classified data. For studies that only report sensitivity and specificity for each test, counts for the cross- classification are imputed based on the associations observed in studies that report the full cross­classification. The approach of Dimou (2016) also allows for the inclusion of studies that report results for only one of the tests in the comparison. These methods require further evaluation before they are recommended for routine use. However, as suggested by Trikalinos (2014), they may be useful as a sensitivity analysis.
Network meta- analysis models have also been developed that use data from both direct and indirect comparisons of multiple tests, e.g. Ma (2018), Nyaga (2018), Lian (2019) and Owen (2018). Rücker (2018) provides a very useful overview of this topic. Veroniki (2022) also provides an overview and an empirical evaluation. However, fur­ther evaluation of these methods for dealing with complex correlational structures is required before they can be considered in Cochrane Reviews.
Example 3 continued: CT versus MRI forthe diagnosis ofcoronary artery
9.4.7.6 disease
The meta- analysis by Schuetz included five studies that made a direct comparison of CT and MRI. Basing the analysis on these five studies (10 2×2 tables) has the advantage that the results should be less prone to bias. However, the number of studies in the analysis is dramatically reduced, which reduces the precision of the summary esti­mates. As we will see in this example, simplifying assumptions may also be required to fit complex hierarchical models to these data. We will again apply the bivariate model for these data.
The SROC plot (Figure9.4.g) shows the data for the five paired studies, with squares used to denote CT and diamonds used to denote MRI. A line is used to join the results for CT and MRI within each study. Examining this plot, we can see that sensitivity for CT is lower than for MRI in one study, equivalent in one study and higher in the other three studies. Specificity is higher for CT than for MRI in three studies and lower in the other two.
236