Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_2767_Библиотеки_им_академика_М_И_Перельмана
.pdf
9.4 Fitting hierarchical models
ii
sZ
ii
NZ
ii
NZ
(a) R
par
curves crossing (𝛿 ≠ 0)
cross (𝛿 = 0)
1. 0
Sensitivity
0
1. 0
https://t.me/medicina_free
the estimated difference may cover a clinically important effect. In that situation, the
results would be regarded as inconclusive.
9.4.6.4 Heterogeneity andregression analysis using theRutter andGatsonis
HSROC model
The HSROC model allows covariates to be added to explore heterogeneity in test positivity (threshold), position of the curve (accuracy) and shape of the curve. A covariate
may be associated with some but not all three model parameters.
Assuming that we have a binary study- level covariate (Z) coded as 0 or 1 to represent
the two groups of studies, then the HSROC model can be extended to estimate the log
odds of a positive test for study i and disease group j as follows
t
ij ii iiij
ZZdi
disexp
j
where each of γ, ξ and δ is assumed to be a fixed effect. Hence, the distributions of the
random effects for threshold and accuracy are now given by
and
, respectively. The shape parameter for the summary curves for the
two groups is estimated as β for the referent group of studies (Z = 0) and β+δ for the
other group (Z = 1). If the covariate does explain some of the heterogeneity in threshold
and/or accuracy, then we would expect that the estimated variance for one or both random effects would be reduced.
The first step would be to investigate the shape of the summary curve. If δ≠0, then
the shape of the summary curve differs for the two groups of studies, which means that
the relative accuracy of the test for the two groups of studies will vary with threshold
(Figure9.4.d (a)). This represents the most complex scenario, and the model would not
generally be simplified any further. The power to detect differences in shape will be low
0.8
0.6
0.4
group 1
group 0
Sensitivity
group 1
group 0
0.8
0.6
0.4
0.2
0.0
0.0 0.2 0.4 0.6 0.8 1.0
elative accuracy depends on
ticular specificity values, the
Figure9.4.d SROC curves with and without a difference in shape
1-specificity
(b) Group 1 dominates across all
specificity values, the curves do not
0.2
0.0
0.0 0.2 0.4 0.6 0.8 1.
1-specificity
227

9 Understanding meta- analysis
https://t.me/medicina_free
when the number of studies in either group is limited. Also, it is important when investigating shape to consider the effect of outlying and potentially influential studies (see
Section9.4.9). When there is statistical evidence that the curves differ in shape, a plot of
the estimated curves for the two groups will aid in interpretation. Focusing on the
region of the plot that covers the observed data, it is then possible to compare the estimated curves. Where one curve consistently lies above another in the region of primary
interest, there is evidence of superior accuracy even though the separation between the
curves will vary across thresholds. If the curves cross, then the interpretation of which
curve shows superior accuracy will depend on threshold.
If, based on statistical evidence, similarity of curve shapes and investigation of potentially influential studies, it can be assumed that δ=0, then the covariate can be removed
for shape. The estimated SROC curves for the two groups will then have the same shape,
even though they may not be symmetrical (Figure9.4.d (b)), and the relative diagnostic
accuracy of the two curves can be summarized using the relative diagnostic odds ratio
(RDOR=exp(ξ)). The RDOR will be constant across all possible values of θ. If the model
can be simplified further and both curves can be assumed to be symmetrical, i.e. β=0,
the RDOR again provides a measure of relative accuracy as already described, but in
addition the DOR in each group will be constant across thresholds.
If the curves can be assumed to have the same shape (either both asymmetrical or
both symmetrical), then the
question is whether the covariate is associated with accuracy, i.e. the position of the curve. If there is evidence that ξ≠0, then the RDOR gives an
estimate of the overall relative diagnostic accuracy. This would correspond to a clear
separation between the SROC curves for the two groups. Alternatively, ξ=0 implies that
there is no separation between the curves and no association between the covariate
and accuracy.
If ξ can be assumed to be 0, then the model can be further simplified by removing the
covariate for accuracy, which will result in a single summary curve (assuming that the
shape of the curve is the same for the two groups of studies). An association between the
covariate and the threshold parameter (i.e. γ≠0) would indicate that the underlying test
positivity rate for the two groups of studies differs. Such an association is often difficult
to interpret unless the curves can be assumed to have the same shape and accuracy.
The RDOR is useful for the statistical comparison of two curves that have the same
shape because the RDOR is constant across all values of the threshold parameter θ.
However, it does not have a straightforward interpretation when the shapes of the
curves differ. In that case, the estimated RDOR will represent the relative accuracy of
the points on the curves where they intersect the diagonal line in ROC given by sensitivity = specificity. One approach that can be used to aid interpretation is to compute
theestimated sensitivities at a chosen value of specificity (or the other way round) to
compare the curves at selected points that are of clinical importance, as illustrated in
Chapter10.
9.4.6.5 Example 2 continued: Investigating heterogeneity indiagnostic accuracy
ofrheumatoid factor (RF)
We will now investigate whether the laboratory technique used to measure RF is associated with diagnostic performance. Of the 50 studies, 15 used nephelometry (N), 16latex
agglutination (LA), 16 ELISA, one study used RA hemagglutination and 2 did not report
228

9.4 Fitting hierarchical models
q
https://t.me/medicina_free
the method used. The analysis is restricted to studies that used N, LA or ELISA. The
HSROC model was again used because of the variation in threshold used for test positivity across studies. In keeping with standard methods used in regression analysis, two
covariates are defined to distinguish between the three techniques (indicator variables
for N and ELISA to compare these groups of studies with LA, the referent group). These
covariates were included in the model to assess whether accuracy, threshold or the
shape of the SROC curve varied with technique. The variances of the random effects for
threshold and accuracy are assumed to be common to all three techniques. (See
Chapter10 for example programs and output.)
The - 2Log likelihood for the most complex model that included covariates for shape,
accuracy and threshold parameters was 752.9. The increase in the - 2Log likelihood was
negligible (an increase to 753.1) when the covariates for shape were removed from the
2
model (Chi
= 753.1 − 752.9 = 0.2, 2 df, P = 0.90). Parameter estimates for the model that
assumes a common shape are given in Table9.4.d, and the corresponding estimated
HSROC curves shown in Figure9.4.e. The estimates of alpha (Λ), theta (Θ) and beta (β)
can be input to RevMan to obtain the summary curve for the referent group (LA).The
threshold and accuracy parameter estimates for ELISA are given by Θ + γ
Λ+ξ
, respectively, and for N are given by Θ+γN and Λ+ξN, respectively, on the logit
ELISA
ELISA
and
scale. The standard errors of these estimates for ELISA and N are most easily obtained
by refitting the model twice, first defining ELISA as the referent group and then defining
N as the referent group.
From Figure 9.4.e, it appears that LA may be less accurate than the other two
methods; however, removal of thecovariate for accuracy (i.e. coefficients ξ
ELISA
and ξN
assumed to be 0) has a negligible effect on the fit of the model (Chi2 = 753.7 − 753.1 = 0.6,
2 df, P = 0.74), indicating no statistical evidence of a difference in diagnostic accuracy of
RF according to technique. This is consistent with the P values based on Wald statistics
in Table9.4.d. Hence, it is reasonable to fit a single SROC for RF, as there is no statistical
evidence of a difference in test accuracy for the three techniques (see Chapter10 for
model estimates for the single common summary curve). These results indicate that it
Table9.4.d HSROC parameter estimates tocompare RF techniques
Parameter Estimate Standard error P value
Λ
Θ −0.5490 0.2137 0.014
β 0.1995 0.1702 0.25
ξ
ELISA
ξ
N
γ
ELISA
γ
N
*This P value is ignored as it tests the uninformative hypothesis that the mean of the pseudo threshold
parameter is zero.
2.4552 0.3245 < 0.0001
*
1.2865 0.3109 0.0002
0.4786 0.1139 0.0001
0.2483 0.4408 0.58
0.3328 0.4439 0.46
−0.1962 0.2614 0.46
0.4960 0.2627 0.065
229

9 Understanding meta- analysis
1
Sensitivity
0
https://t.me/medicina_free
0.9
0.8
0.7
0.6
0.5
0.4
0.3
0.2
0.1
0
1 0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1
Legend
Method: ELISA
Method: LA
Method: Nephelometry
Figure9.4.e SROC curves to compare accuracy of RF techniques
Specificity
may be reasonable to treat the techniques as equivalent markers for RF (they have
similar DOR, but the trade- off between sensitivity and specificity may still differ).
9.4.7 Comparing index tests
Many diagnostic reviews aim to compare the diagnostic accuracy of two alternative
index tests that may be used to detect the same condition. In this section, the focus will
be on the comparison of two index tests, but the approach can be extended to allow for
more than two tests.
Two approaches are generally adopted for test comparisons. The first approach uses
test accuracy data from all eligible studies that have evaluated one or both tests. The
second approach restricts the analysis to studies that have either evaluated both tests
in the same individuals, or have randomized individuals to undergo one or other of the
two tests. The second approach has advantages because the comparison is less likely to
be biased by confounding and hence these results should be relied upon where possible (Takwoingi 2013). However, the number of studies that report such direct comparisons is often very limited, which means that such an analysis may not be feasible or
may only be considered as a sensitivity analysis (see Section9.4.9).
9.4.7.1 Test comparisons based onall available studies
Often, many of the available studies evaluate only one of the tests of interest (Takwoingi
2013). By using all studies that have evaluated at least one of the tests, we maximize the
number of studies in the analysis. However, the studies are likely to be heterogeneous
230

9.4 Fitting hierarchical models
https://t.me/medicina_free
in terms of design and patient characteristics that are associated with test accuracy,
and hence confounding may be an issue. In preliminary exploratory analyses this can
be dealt with by comparing the tests within subgroups of studies that are homogeneous with respect to important potential confounders such as study design or spectrum
of disease. The value and feasibility of such exploratory analyses will be affected by the
number of available studies and missing or inconsistent reporting across studies of
information on potential confounders.
The statistical methods described in this section follow directly from the description
of hierarchical models in Section9.4.6 and how they can be used to investigate heterogeneity in test accuracy. For the comparison of two index tests, the type of test is represented by a binary covariate, which is used to identify the test that gave rise to each 2×2
table included in the analysis. Confounders can potentially be adjusted for; however,
this may be difficult to do in practice because the number of studies is often small and/
or data on important confounders may be poorly recorded or incomplete.
Both the bivariate model and the Rutter and Gatsonis HSROC model can be used to
investigate the relative accuracy of two index tests. However, as noted previously, the
choice of approach will be influenced by the nature of the available data. The interpretation of the results will depend on which approach is used.
Test comparisons using thebivariate model
9.4.7.2
If, for each index test, the available studies have used a consistent threshold on a
continuous or ordinal scale to define test positivity, then the bivariate model provides an
appropriate framework for test comparisons. It may also be reasonable to assume aconsistent threshold when a test comprises a ‘test kit’ that produces positive and negative
results (such as a coloured line appearing on a device). By adopting the same strategy
described earlier (Section9.4.6.2), a binary covariate for test type can be included in the
model to investigate whether sensitivity and/or specificity differs between two tests.
Care must be taken with the interpretation of the results of such a model, particularly
if the common threshold for test positivity for either test is applied to a continuous or
ordinal scale. Any inferences made about the relative diagnostic accuracy of the two
tests is only valid at the chosen threshold for each of the two tests and cannot be
extrapolated to other possible thresholds. Where other thresholds are reported, the
analysis can be repeated using the available data to investigate the relative diagnostic
accuracy of the tests at those alternative thresholds. However, such additional investigations should be restricted to thresholds reported by sufficient studies to allow a
meaningful analysis.
Because we are analysing test accuracy data for two alternative index tests, it may not
be reasonable to assume that the variances of the random effects for logit(sensitivity)
and logit(specificity) are the same for the two tests. The bivariate model can be extended
to allow the variance of the random effects for each test to depend on the covariate for
test type (see Chapter10). This will also affect the estimated correlation between them.
Statistically, estimation of the variances of the random effects for logit(sensitivity) and
logit(specificity) and correlation between them is subject to a higher level of uncertainty than for the main parameters of interest. However, when preliminary plots of the
study- level estimates of sensitivity and specificity in ROC space show marked differences in heterogeneity between studies for the two tests, it is advisable to assess
231

9 Understanding meta- analysis
https://t.me/medicina_free
whether the assumption of equal variances of random effects for the two tests is reasonable. This can be done by comparing estimates for test accuracy between the alternative models to assess whether conclusions about the relative sensitivity and/or
specificity of the tests are robust to assumptions about the variances of the random
effects. Such an investigation may not be feasible if the number of studies is small.
It is usual for most of the studies in such an indirect analysis of test comparisons to
have evaluated only one of the tests, but some studies may have evaluated both. If the
proportion of studies that have evaluated both is very small, then treating the results
ofthe two tests in a study as if they were obtained from different studies is unlikely
toaffect the results. Although this is often done in practice, such an approach is not recommended if the proportion of studies evaluating both tests is not small, because it is
likely to result in inappropriate standard errors for the test comparison parameters for
sensitivity and specificity. In that case the paired sensitivity/specificity data for bothtests
from each study should be at the lower level of the hierarchical analysis, anda binary
covariate for test type included to identify which 2×2 table corresponds to each test.
9.4.7.3
Example 3: CT versus MRI forthe diagnosis ofcoronary artery disease
Schuetz (2010) evaluated the diagnostic performance of multi- slice computed tomography (CT) and magnetic resonance imaging (MRI) for the diagnosis of coronary artery
disease (CAD). The review included prospective studies that evaluated either CT or MRI
(or both), used conventional coronary angiography (CAG) as the reference standard,
and used the same index test threshold for clinically significant coronary artery stenosis
(a diameter reduction of 50% or greater). A total of 103 studies provided a 2×2 table for
one or both tests and were included in the meta- analysis: 84 studies evaluated only CT,
14 evaluated only MRI and 5 studies evaluated both CT and MRI. (See Chapter10 for
data and example programs.)
Because study selection was based on a common threshold for clinically significant
coronary artery stenosis, the bivariate model was used for data synthesis and test comparison. In the first stage of the analysis, we base our test comparison on all studies that
evaluated at least one test. The approach follows closely the method illustrated in
Section 9.4.6.3 for exploring heterogeneity using the bivariate model and the same
notation has been used.
A binary covariate is added to the model, which is coded as 0 if the 2×2 table is for MRI
(the referent group) and coded as 1 if the 2×2 table is for CT. The 5 studies that evaluated
both tests contribute a 2×2 table for each test, hence there are 19 studies included for
MRI and 89 studies included for CT. Allowing both sensitivity and specificity to vary by
type of test resulted in a –2Log likelihood of 953.0, a reduction of 42.5 compared with
the model that contained no covariates. Hence, there is statistical evidence (Chi
2 df, P < 0.001) that sensitivity and/or specificity are associated with test type. Removing
the covariate for sensitivity from the model (Chi2 = 976.7 − 953.0 = 23.7, 1 df, P < 0.001)
shows strong statistical evidence of a difference in sensitivity between the two tests.
Similarly, removing the covariate for specificity from the model (Chi2 = 976.2 − 953.0 =
23.2, 1 df, P < 0.001) shows strong statistical evidence of a difference in specificity
between the two tests. These results are consistent with the P values shown in
Table9.4.e that are based on Wald statistics.
Table 9.4.e gives the estimated mean logit(sensitivity) and mean logit(specificity)
(μ
and μB, respectively) for the referent category (MRI), the common variances of the
A
232
2
= 42.5,

9.4 Fitting hierarchical models
A
B
1
Sensitivity
0
https://t.me/medicina_free
Table9.4.e Bivariate model estimates forcomparison ofCT andMRI
Parameter Estimate Standard error P value
μ
A
μ
B
2.1771 0.2457
0.8754 0.2111 < 0.0001*
0.8749 0.2293 0.0002
0.8447 0.1696 < 0.0001
σ
AB
ν
A
ν
B
0.1803 0.1384 0.20
1.3033 0.2625 < 0.0001
1.0415 0.2154 < 0.0001
*These Wald statistics are ignored as they test whether μ
that sensitivity = 50% and specificity = 50%, respectively.
0.9
0.8
0.7
0.6
0.5
0.4
0.3
< 0.0001
= 0 and μB = 0, equivalent to testing hypotheses
A
*
0.2
0.1
Figure9.4.f Summary estimates of accuracy of CT and MRI for the diagnosis of coronary artery
disease with corresponding 95% confidence regions
random effects (A and B, respectively), their covariance (σAB) and the difference in
mean logit(sensitivity) and mean logit(specificity) between CT and MRI (νA and νB,
respectively). Estimates of the standard errors of mean logit(sensitivity) and mean
logit(specificity) for CT are required for entry into RevMan and can be obtained easily
from refitting the model with CT defined as the referent group.
Figure 9.4.f shows the SROC plot with summary points for MRI and CT and their
95%confidence regions superimposed as shown. The squares represent CT and the
0
1 0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1
Legend
CT
MRI
Specificity
233

9 Understanding meta- analysis
https://t.me/medicina_free
diamonds represent MRI. Because of the large number of studies for CT, the summary
point and region are difficult to see. The figure could be redrawn without the individual
study points to show just the summary points and regions, or the size of the study
points can be reduced to improve visibility.
Inverse logit transformation of the model estimates and their confidence limits provide the relevant estimates in ROC space: summary estimates for sensitivity are 0.90
(95% CI 0.84 to 0.93) for MRI and 0.97 (95% CI 0.96 to 0.98) for CT. The summary estimates for specificity are 0.71 (95% CI 0.61 to 0.78) for MRI and 0.87 (95% CI 0.84 to 0.90)
for CT. The P values for νA and νB show strong statistical evidence of an association
between sensitivity and test type and also between specificity and test type.
Based on this analysis, there is strong evidence that CT has higher sensitivity and
specificity than MRI for detecting clinically significant coronary artery stenosis, defined
as a diameter reduction of 50% or more. Chapter10 demonstrates how to compute the
difference in sensitivity and difference in specificity, with corresponding confidence
intervals, for the two tests.
Further descriptive analyses and modelling may be undertaken to assess whether the
assumption of equal variances for the random effects for the two tests has a substantive
effect on the results presented here (see Chapter10).
9.4.7.4 Test comparisons using theRutter andGatsonis HSROC model
Comparisons of summary estimates of sensitivity (or specificity) of alternative tests
can be misleading if the included studies have used different thresholds to define
testpositivity. In this situation, comparisons based on SROC curves provide a more
informative approach.
The hierarchical modelling strategy used to investigate heterogeneity described earlier for the Rutter and Gatsonis HSROC model (Section9.4.6.4) can be used for comparisons of test accuracy when there is variability in threshold between studies. The type of
test is represented by a binary covariate that is used to identify the test that gave rise
to each 2×2 table included in the analysis. This covariate then allows the review
authorto investigate whether test type is associated with the shape and position of
theSROC curve. Interpretation of the results follows directly from the discussion of the
interpretation of investigations of heterogeneity in Section9.4.6.
Statistically, estimation of the variances of the random effects for threshold and accuracy is subject to a higher level of uncertainty than for the main model parameters of
interest. If preliminary plots of the study-
level estimates of sensitivity and specificity in
ROC space show marked differences in heterogeneity between studies for the two tests,
it is advisable to assess whether the assumption of equal variances of the random
effects for the two tests is reasonable (see Chapter10). This is usually done by comparing the fit of the alternative models (i.e. where variances do, or do not, depend on the
covariate for test type). A comparison of the main estimates of interest between the
alternative models is also useful to assess whether conclusions about the relative
shapeand accuracy of the summary curves for the two tests are robust to assumptions
about the variances of the random effects. Again, such an investigation will not be
feasible if the number of studies is small.
As noted for the bivariate model, it is usual for most of the included studies to have
evaluated only one of the tests, but some studies will have evaluated both. If the
234

9.4 Fitting hierarchical models
1
Sensitivity
0
https://t.me/medicina_free
proportion of studies that have evaluated both is very small, then treating the results of
the two tests in a study as if they were obtained from different studies is unlikely to
affect the results. However, more accurate standard errors will be obtained for the test
comparison parameters if the data for both tests are modelled within the study at the
lower level in the analysis.
9.4.7.5 Test comparison based onstudies that directly compare tests
As noted in Section9.1.3 and Section9.3.2, heterogeneity in the estimated accuracy of
a diagnostic test across studies is likely to occur. This could confound the comparison of
two tests if different studies are used to estimate the diagnostic accuracy of each test.
Ideally, the comparison should be based on studies that have made a direct comparison of the tests of interest, either by applying both tests to each study participant, or by
randomizing each individual to receive one of the tests (Takwoingi 2013). A common
reference standard should be applied to both tests. If there are sufficient studies of
thistype on which to base a test comparison, the results are less prone to bias than an
analysis based on all available studies that have evaluated one or both tests.
A preliminary graphical analysis can be conducted by plotting the estimated sensitivity and specificity for both tests for each study in ROC space. The two points contributed
by each study (one for each test) are joined by a line to highlight the relative test accuracy within each study to illustrate the pairing of test accuracy estimates at the study
level (see Figure9.4.g).
The rationale described earlier for choosing between the bivariate model and the
HSROC model when making test comparisons is also applicable here, and the same
0.9
0.8
0.7
0.6
Figure9.4.g SROC plot of direct comparisons of CT vs MRI for the diagnosis of coronary artery
disease
0.5
0.4
0.3
0.2
0.1
0
1 0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1
Legend
CT
MRI
Specificity
235

9 Understanding meta- analysis
https://t.me/medicina_free
issues relating to interpretation apply. The only major difference is that the analysis
does not include any studies that have evaluated only one of the tests.
Because each study contributes a 2×2 table for each of the two tests to be compared,
the data for the two tests should be analysed within the study at the lower level in the
analysis. The random effect for that study is assumed to be common to both 2×2 tables
for that study. A binary covariate for test type is included to identify which 2×2 table
corresponds to each test. Entering a separate 2×2 table for each test (within each study)
for analysis in a hierarchical model effectively assumes that the data arise from a
randomized or independent groups design. This represents a conservative approach
that is often necessitated by the lack of information on paired results at the individual
level for truly ‘paired’ studies that have applied both tests to the same individual.
Atpresent, it is not common practice for researchers to publish a cross-
classification of
test results within both the diseased and non- diseased groups, or to provide individual
patient data for analysis that would provide full information on the pairing of test
results within an individual.
Meta- analytical models that account for pairing of test results within an individual
within each study have been developed as an extension of the bivariate model. The
method proposed by Trikalinos (2014) allows for comparisons of two tests based on
studies that use a paired design and report fully cross- classified data. For studies that
only report sensitivity and specificity for each test, counts for the cross- classification
are imputed based on the associations observed in studies that report the full crossclassification. The approach of Dimou (2016) also allows for the inclusion of studies that
report results for only one of the tests in the comparison. These methods require further
evaluation before they are recommended for routine use. However, as suggested by
Trikalinos (2014), they may be useful as a sensitivity analysis.
Network meta- analysis models have also been developed that use data from both
direct and indirect comparisons of multiple tests, e.g. Ma (2018), Nyaga (2018), Lian
(2019) and Owen (2018). Rücker (2018) provides a very useful overview of this topic.
Veroniki (2022) also provides an overview and an empirical evaluation. However, further evaluation of these methods for dealing with complex correlational structures is
required before they can be considered in Cochrane Reviews.
Example 3 continued: CT versus MRI forthe diagnosis ofcoronary artery
9.4.7.6
disease
The meta- analysis by Schuetz included five studies that made a direct comparison of
CT and MRI. Basing the analysis on these five studies (10 2×2 tables) has the advantage
that the results should be less prone to bias. However, the number of studies in the
analysis is dramatically reduced, which reduces the precision of the summary estimates. As we will see in this example, simplifying assumptions may also be required to
fit complex hierarchical models to these data. We will again apply the bivariate model
for these data.
The SROC plot (Figure9.4.g) shows the data for the five paired studies, with squares
used to denote CT and diamonds used to denote MRI. A line is used to join the results for
CT and MRI within each study. Examining this plot, we can see that sensitivity for CT is
lower than for MRI in one study, equivalent in one study and higher in the other three
studies. Specificity is higher for CT than for MRI in three studies and lower in the other two.
236
Соседние файлы в папке Библиотека им академика М.И. Перельмана
