Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_2767_Библиотеки_им_академика_М_И_Перельмана
.pdf
Table4.5.b Evaluation ofthe Innova lateral flow antigen test todetect SARS- CoV- 2infection inLiverpool, UK
https://t.me/medicina_free
RT- qPCR reference standard
Positive Negative Void Total
Innova
LFT
result
Measure of test accuracy Equation Estimate 95% CI
Sensitivity TP/(TP+FN) 0.400 (0.285 to 0.524)
Specificity TN/(TN+FP) 0.999 (0.998 to 1.000)
Positive predictive value TP/(TP+FP) 0.903 (0.742 to 0.980)
Negative predictive value TN/(TN+FN) 0.992 (0.990 to 0.994)
Positive likelihood ratio sensitivity/(1−specificity) 725 (226 to 2328)
Negative likelihood ratio (1−sensitivity)/specificity 0.600 (0.496 to 0.727)
Diagnostic odds ratio LR+/LR− 1207 (346 to 6311)
Data taken from García- Fiñana (2021).
If the control line of an LFT or RT- qPCR test fails to appear within 30minutes, the result is recorded as void. All void results were excluded from the main analysis.
In a secondary analysis, the study authors included void LFT results as test negatives. CI, confidence interval; LFT, lateral flow test; LR, likelihood ratio; RT- qPCR,
quantitative reverse transcription polymerase chain reaction.
Positive 28 (TP: true positives) 3 (FP: false positives) 2 33
Negative 42 (FN: false negatives) 5431 (TN: true negatives) 341 5814
Void 4 18 0 22
Total 74 5452 343 5869

4.5 Analysis ofa primary test accuracy study
11
sensitivity
specificity
https://t.me/medicina_free
ratio of that test result. This is a simple result of the definition of conditional probabilities.
For a test that is informative, the post- test probability should be higher than the
pre- test probability if the test result is positive, while the post- test probability should
be lower than the pre- test probability if the test result is negative. Considerations
about the use of likelihood ratios in systematic reviews of test accuracy are explained
in Chapter11.
The values of sensitivity and specificity are sometimes combined in a single measure
known as Youden’s index, defined as sensitivity + specificity − 1. Youden’s index is not a
proportion and cannot be interpreted as a probability. Instead, it is a general index of
test accuracy. Values close to 1indicate high accuracy; a value of 0means that the test
cannot discriminate between those with and without the target condition. When the
index is equal to 0 the proportion of true positives in those with the target condition is
equal to the proportion of false positives in those without the target condition. This
means that the test is not informative: a positive result is as likely in those with the target condition as in those without.
Youden’s index combines sensitivity and specificity and, as a result, differences in
the implications of misclassification between those with and without the target condition are lost. This makes it less informative for medical decision-
making. However,
it is not uncommon to encounter primary studies that claim to identify the ‘optimal’
threshold based on maximizing Youden’s index. Such computations are only optimal
under the unlikely assumption that the consequences of false negatives and false
positives are the same, which often is not appreciated by those quoting these
values.
Some test accuracy studies report the overall accuracy, computed as the total of true
positives and true negatives divided by the total sample size. Again, this measure does
not differentiate between the implications of false positive and false negative test
errors and is therefore not recommended for any form of decision- making.
The diagnostic odds ratio (DOR) summarizes the diagnostic accuracy of the index
test as a single number, which describes how many times higher the odds are of
obtaining a positive test result in someone with the target condition, selected at random, than in someone without the target condition, also selected at random (Glas
2003). Alternatively, it can be defined as the odds of the target condition in test positives versus the odds of having the target condition in test negatives. The DOR is
formally defined as
The DOR is also equal to LR+/LR−. It is estimated from Table4.5.a as (ad)/(bc).
The same DOR may be achieved by different combinations of sensitivity and specificity, as shown in Figure4.5.a, where the numbers circled in red indicate sensitivity–
specificity combinations that have the same diagnostic odds ratio of 9. For example,
a DOR of 9 could be achieved by a specificity of 90% and a sensitivity of 50%, or by a
sensitivity of 90% and a specificity of 50%. The fact that it summarizes test accuracy
sensitivity
1
1
specificity
sensitivity specificity
sensitivity specificity
63

4 Understanding test accuracy measures
99
https://t.me/medicina_free
Sensitivity
Specificity 50% 60% 70% 80% 90% 95% 99%
50% 122491
60% 224614 29 149
70% 245921 44 231
80% 469163676 396
90% 914213681 171 891
95% 19 29 44 76 171 361 1881
99% 99 149 231 396 891 1881 9801
Figure4.5.a Diagnostic odds ratios achieved at different values of sensitivity and specificity
9
in a single number makes it popular with some authors. Like Youden’s index and
overall accuracy, it is uninformative about the nature of misclassifications. However,
as will be explained in Chapter9, the DOR has a particular role in meta- analysis models
of test accuracy.
4.6 Positivity thresholds
In studies of the accuracy of tests with ordinal and continuous results, positive and
negative test results are defined based on a threshold for test positivity and change if
the threshold is altered. This dependence on threshold is a fundamental aspect of test
accuracy evaluation. In the case of test sensitivity and specificity, the dependence
induces a trade- off between the two quantities, one value increasing while the other
decreases as the threshold for positivity is changed.
This is illustrated in the panels in Figure4.6.a, which show the same hypothetical
distributions of test results for those with and without the target condition on a continuous scale. The shaded areas show how the false
false positive fraction (green) change as the positivity threshold varies. The panels
vary in the numerical value of the threshold used to define test positive. At each
threshold, the sensitivity of the test reflects theproportion of the area under the
curve to the right of the threshold: those with the target condition. Similarly, the
specificity reflects the proportion of the area under the curve to the left of the
threshold.
As the threshold decreases from panel (a) to panel (e), the proportion of those with
the target condition who are above the threshold– and hence have a positive test–
increases from 69% to 99%. These values stand for the sensitivity of the test. At the
same time, the corresponding proportion of those without the target condition who are
below the threshold– and hence have a negative test result– decreases from 99% to
69%. These values stand for the specificity of the test.
negative fraction (red) and the
64

4.6 Positivity thresholds
(a) (b)
(c)
(e)
Specif
Ta
%
Test result
Specif
Ta
%
Specif
Ta
https://t.me/medicina_free
icity = 99% Sensitivity = 69%
rget condition
absent
TN FN FP TP
0 20 40 60 80 100 120 140 16 0
icity = 93% Sensitivity = 93%
rget condition
absent
020406080100 120140 160
Test result
TN FN FP TP
Test result
Ta rget condition
present
Ta rget condition
present
Specificity = 98% Sensitivity = 84
Ta rget condition
absent
TN FN FP TP
0 20 40 60 80 100 120 140 160
Test result
(d)
Specificity = 84% Sensitivity = 98
Ta rget condition
absent
TN FN FP TP
020406080100 120140 160
Test result
Ta rget condition
present
Ta rget condition
present
icity = 69% Sensitivity = 99%
rget condition
absent
TN FN FP TP
020406080100 120140 160
Figure4.6.a Relationship between sensitivity, specificity and the positivity threshold
Ta rget condition
present
65

4 Understanding test accuracy measures
https://t.me/medicina_free
4.7 Receiver operating characteristic curves
Primary studies that evaluate a test at several thresholds sometimes present their
findings as receiver operating characteristic (ROC) curves. The ROC curve of a test is
thegraph of the values of sensitivity and specificity that are obtained by varying the
positivity threshold across all possible values. The graph plots sensitivity (true- positive
fraction) against 1–specificity (false positive fraction).
The curve for any test moves from the point where sensitivity and 1–specificity are
both 1 (the upper right corner), which occurs when the threshold classifies all participants as test positive (there are no false negatives and all without the target condition
are false positives), to a point where sensitivity and 1−specificity are both 0 (the lower
left corner), which occurs when all participants are classified as test negative (giving no
false positives and all with the target condition are false negatives). The shape and position of the curve between these two fixed points depend on the discriminatory ability of
the test across all possible thresholds. The closer the curve lies to the top leftcorner (where sensitivity and specificity are both 1), the more discriminatory the test.
In practice, the ROC curve is estimated from a finite sample of test results and hence
will not necessarily be a smooth curve, as shown in Figure4.7.a. Note that the horizontal axis for each ROC plot in Figure4.7.a is labelled in terms of specificity decreasing
from 1.0 to 0.0. This style of labelling is used in Cochrane Reviews and is equivalent to
the usual labelling (1−specificity or the false positive fraction ranging from 0.0 to 1.0).
The position of the ROC curve depends on the degree of overlap of the distributions
of the test results in those with and without the target condition. Where a test clearly
discriminates between those with and without the target condition such that there is no
or little overlap of distributions, the ROC curve will indicate that high sensitivity is
achieved with a high specificity, i.e. the curve approaches the upper left- hand corner of
the graph where sensitivity is 1 and specificity is 1 (Figure4.7.a panel (a)). If the distributions of test results in the two subgroups overlap completely, the test would be completely uninformative, and its ROC curve would be the upward diagonal of the square
(Figure4.7.a panel (c)).
The ROC curves shown in Figure4.7.a panel (a) to panel (c) are all symmetrical about
the downward diagonal of the square, where sensitivity equals 1–specificity. It is also
possible for ROC curves to be asymmetrical, as in Figure4.7.a panel (d). Asymmetrical
curves typically occur when the distribution of the test measurement in those with the
target condition has more, or less, variability than the distribution in those without the
target condition. Increased variability might occur, for example, if the target condition
causes a biomarker both to rise and become more erratic; reduced variability might
occur if the target condition lowers biomarker values to a bounding level, such as the
lower level of detection.
The comparison of tests based on their ROC curves takes into consideration their
accuracy across a range of thresholds and is aided by single summary measures. Several
of these have been proposed in the literature. Most used among them is the area under
the curve (AUC), which equals 1 for a perfect test and 0.5 for a completely uninformative
test. The AUC is also sometimes referred to as the c- statistic.
The AUC can be interpreted as a probability: it reflects the chances that, in a pair of
one with and one without the target condition, selected at random, the one with the
target condition will have a test result that is more compatible with the target condition
66
hand

(a)
(b
(c
(d
Ta
absent
0
0
0
0
Specificity
Ta
absent
Ta
absent
Ta
absent
https://t.me/medicina_free
rget condition
Ta rget condition
present
Sensitivity
0 20 40 60 80 100 120 140
Test measurement
)
rget condition
020406080100 120140
Test measurement
Ta rget condition
present
)
rget condition
Ta rget condition
present
0.0 0.2 0.4 0.6 0.8 1. 0
Specificity
Sensitivity
0.0 0.2 0.4 0.6 0.8 1. 0
Specificity
Sensitivity
0.
0.20.40.60.81. 0
0.20.40.60.81. 0
0.
020406080100 120140
Test measurement
)
rget condition
020406080100 120140
Test measurement
Figure4.7.a Examples of ROC curves
Ta rget condition
present
0.0 0.2 0.4 0.6 0.8 1. 0
Specificity
Sensitivity
0.0 0.2 0.4 0.6 0.8 1. 0
0.
0.20.40.60.81. 0
0.20.40.60.81. 0
0.
67

4 Understanding test accuracy measures
11
D
y
00
D
y
10
D
y
01
D
y
11
D
y
00
D
y
10
D
y
01
D
y
A1
D
B1
D
A0
D
B0
D
A0
D
B0
D
A1
D
B1
D
A1 1 A 0 0
DD
An n An n
B1 1 B0 0
DD
Bnn Bn n
https://t.me/medicina_free
than the other. If higher test results point to having the target condition, the AUC represents the probability that in that pair the one with the target condition will have a higher
result than the one without the target condition.
The AUC can also be interpreted as an average sensitivity for the test, taken over all
specificity values, or, equivalently, as the average specificity over all sensitivity values.
Other summaries include partial areas under the curve, values of sensitivity corresponding to selected values of specificity (and the other way round) and operating
points defined according to specified criteria (such as the Q* value, which is the point
on the curve where sensitivity and specificity are equal).
4.8 Analysis ofa comparative accuracy study
As explained in Chapter3, robust comparative studies of diagnostic test accuracy use
either a within- subject paired design, in which all patients undergo all tests together
with a reference standard, or, more rarely, a between- subject randomized design, in
which all patients undergo the reference standard test but are randomly assigned to
have only one of the index tests (Takwoingi 2013). These designs have consequences for
how the accuracy of two or more tests can be characterized.
For studies that use randomized comparisons, a separate 2×2 table will be created for
the results in each arm of the study, and computation of test accuracy measures proceeds as before.
Point estimates will be the same if a paired design is analysed as a randomized design,
but standard errors for comparisons will be inaccurate, leading to confidence intervals
that are likely to be too large. Appropriate computations in studies using a paired design
are more complex (Hayen 2010).
Table4.8.a shows the joint classification of the results of two index tests from a study
that used a paired design.
For two tests labelled A and B in Table4.8.a,
target condition for whom both tests are positive;
thetarget condition for whom both tests are negative;
target condition for whom test A is positive but test B is negative; and
of patients with the target condition for whom test B is positive but test A is negative.
For those without the target condition, the counts
in an equivalent manner.
For tests A and B,
numbers of true negatives;
and
are the numbers of true positives;
and
are the numbers of false negatives; and
are the numbers of false positives. If n1 and n0 represent the number of individuals
with and without the target condition, respectively, then the sensitivity and specificity
of the two tests can be estimated from the marginal frequencies as follows:
is the number of patients with the
is the number of patients with
is the number of patients with the
is the number
,
,
and
can be interpreted
and
are the
and
Table4.8.b shows the results of two interferon gamma release assays (IGRAs)– T- SPOT.
TB and a second- generation IGRA– for diagnosis of active tuberculosis cross- classified
68

4.8 Analysis ofa comparative accuracy study
11
D
y
01
D
y
B1
D
11
D
y
01
D
y
B1
D
10
D
y
00
D
y
B0
D
10
D
y
00
D
y
B0
D
A1DA0
D
A1
D
A0
D
https://t.me/medicina_free
Table4.8.a Joint classification ofpaired index tests andreference standard results
Target condition present Target condition absent
Test A Test A
Positive Negative Total Positive Negative Total
Test B Positive
Negative
Total
n
1
n
Adapted from Takwoingi (2016).
Table4.8.b Comparison ofT- SPOT.TB andsecond- generation IGRA fordiagnosis ofactive
tuberculosis
Active TB present Active TB absent
generation IGRA 2nd- generation IGRA
2nd-
Positive Negative Total Positive Negative Total
T- SPOT.TB Positive 253 0 253 51 0 51
Negative 16 33 49 19 296 315
Total 269 33 302 70 296 366
Test accuracy measure Estimate (95% CI)
Sensitivity of 2nd- generation IGRA 269/302 = 89.1% (85.0 to 92.43)
Sensitivity of T-
Absolute difference in sensitivity
(sensitivity of 2ndsensitivity of T-
Relative sensitivity (sensitivity of
2nd-
generation IGRA/sensitivity of
SPOT.TB 253/302 = 83.8% (79.1 to 87.7)
5.30 (2.44 to 8.16) percentage points
generation IGRA−
SPOT.TB)
1.06 (1.03 to 1.10)
T- SPOT.TB)
Specificity of 2nd- generation IGRA 296/366 = 80.9% (76.5 to 84.8)
Specificity of T- SPOT.TB 315/366 = 86.1% (82.1 to 89.4)
Absolute difference in specificity
−5.19 (−7.74 to −2.65) percentage points
(specificity of 2nd- generation IGRA−
specificity of T- SPOT.TB)
Relative specificity (specificity of
0.93 (0.91 to 0.97)
2nd- generation IGRA/specificity of
T- SPOT.TB)
0
Data taken from Whitworth (2019).
Borderline test results excluded. CI, confidence interval; IGRA, interferon gamma release assay; TB, tuberculosis.
69

4 Understanding test accuracy measures
Absolute difference in sensitivity
Absolute difference in specificity .
sensitivity A sensitivity B
Relative sensitivity /
Relative specificity / .
sensitivity A sensitivity B
11
11
sensitivity A sensitivity B
https://t.me/medicina_free
among those with and those without the target condition. Although such a table of the
joint classification of the results of two tests against those of the reference standard is
potentially useful, many studies do not present results in this format, but rather give a
separate 2×2 table of the results of each index test against the reference standard, as
would be expected for a randomized design.
Absolute differences comparing the sensitivity and specificity of two tests can be estimated while relative comparisons can be estimated for most of the measures described
in Section5.5. (Full details of computations including confidence intervals are given in
Hayen (2010)).
If the newer or experimental test is labelled as test A and the older test, or standard
practice, as test B, then the absolute difference in sensitivity and specificity can be
written as
specificity A specificity B
These absolute differences describe the change in the proportion of those with the target condition who will be additionally detected using test A instead of test B (difference
in sensitivity) and the absolute reduction in the numbers without the target condition
given false positive results (difference in specificity).
The relative probabilities can be written as
specificity A specificity B
It is more difficult to define a clear interpretation of the magnitude of relative sensitivity
and specificity measures than of absolute differences. Ratios of 11- specificity values describe multiplicative increases in the proportions misclassified
(false negative and false positive fractions, respectively) and give a distinct perspective
on the magnitude of differences. For example, if test A and test B have sensitivities of
99% to 90%, respectively, the ratio of sensitivities is 1.1 (a 10% increase), whereas the
ratio of the equivalent false
negative fractions of 1% with 10% is a 10- fold increase in
false negatives.
Results may also be presented as odds ratios:
sensitivity A sensitivity B
specificity A specificity B
specificity A specificity B
These odds ratios are not equivalent to the DOR. Here the ratios compare the
samemeasure (sensitivity or specificity) for one test to that of another test, while
the DOR compares two groups (target condition present or absent) for one test
70
sensitivity and

4.9 Chapter information
https://t.me/medicina_free
(Takwoingi2016). The DORs of two tests can be compared and expressed as the relative
DOR (RDOR).
If logistic regression models are used to compare the sensitivity and specificity of two
or more tests, odds ratios are the natural output. However, odds ratios do not have an
intuitive interpretation. Absolute differences and relative probabilities are more familiar to researchers and are straightforward to interpret; therefore, they are preferred
measures of comparative accuracy.
The approximate relationship between odds ratios and relative risks that is exploited
in epidemiology when events are rare (Altman 1998) is invalid in test research, because
events (i.e. true positives or true negatives) are common.
Table4.8.b shows the results of the comparison of the accuracy of the two IGRAs, in
terms of absolute and relative differences in sensitivity and specificity. The relative sensitivity of 1.06 (95% confidence interval (CI) 1.03 to 1.10) indicates that the sensitivity of
the secondThe 95% CI indicates that we are 95% confident that the sensitivity could be about 3%
to 10% higher for the second- generation IGRA compared to T- SPOT.TB. The relative
specificity of 0.94 (95% CI 0.91 to 0.97) indicates a 6% reduction in the specificity of the
second- generation IGRA relative to that of T- SPOT.TB, and we are 95% confident that
the decrease in sensitivity lies between 9% and 3%.
generation IGRA is about 1.06 times or 6% higher than that of T- SPOT.TB.
4.9 Chapter information
Authors: Jonathan J. Deeks (Institute of Applied Health Research, University of
Birmingham, UK); Yemisi Takwoingi (Institute of Applied Health Research, University
of Birmingham, UK); Petra Macaskill (Sydney School of Public Health, University of
Sydney, Australia); Patrick M. Bossuyt (Department of Epidemiology and Data Science,
University of Amsterdam, The Netherlands).
Sources of support: Jonathan J. Deeks is a UK National Institute for Health Research
(NIHR) Senior Investigator Emeritus. Yemisi Takwoingi is funded by a UK National
Institute for Health Research (NIHR) Postdoctoral Fellowship. Jonathan J. Deeks and
Yemisi Takwoingi are supported by the NIHR Birmingham Biomedical Research Centre
at the University Hospitals Birmingham NHS Foundation Trust and the University of
Birmingham. The views expressed are those of the authors and not necessarily those of
the NHS, the NIHR or the Department of Health and Social Care. The authors declare no
other sources of support for writing this chapter.
Declarations of interest: Jonathan J. Deeks, Yemisi Takwoingi and Petra Macaskill are
members of Cochrane’s Diagnostic Test Accuracy Editorial Team. Yemisi Takwoingi and
Petra Macaskill are coMethods Group. The authors declare no other potential conflicts of interest relevant to
the topic of this chapter.
Acknowledgements: The authors would like to thank Marta Roque and Gianni Virgili
for helpful peer review comments.
convenors of the Cochrane Screening and Diagnostic Tests
71
Соседние файлы в папке Библиотека им академика М.И. Перельмана
