Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_2946_Библиотеки_им_академика_М_И_Перельмана
.pdf
Measuring theaccuracy ofclinical findings 65
To measure the specificity of a test, perform the index test and the gold standard test in a source population. In terms
of the 2
× 2 table (see Table5.1), the specificity of a test is:
5.2.3 Measures ofdiscordance between index test anddisease state
Two types of test results that do not reflect the true state of the patient: false negative results and false positive results.
Unlike sensitivity and specificity, these two quantities do not have their own names. Sometimes, “false-
negative rate”
and “false-
positive rate” are used, but these names can be interpreted in several ways.
False- negative results
The frequency of a false- negative result in conditional probability notation is the following:
pp
TDnegative test result diseaseor
||
which means “the probability that a negative test result will occur if the patient has the target condition.” Patients
with the target condition can only experience a positive result or a negative result, so
pT DpTD
||1
Therefore, p[T−|D+]=1–p[T+|D+]
To obtain the value for 1 − sensitivity, simply subtract the test sensitivity from 1. To measure it directly, perform the
index test and the gold standard test in a source population. In terms of the 2 × 2 table
False- positive results
pp
TDpositive test result no diseaseor
||
which means “the probability that a positive test result will occur if the patient does not have the target condition.”
Patients with the target condition can only experience a positive result or a negative result, so
pT DpTD
||1
Therefore, p[T+|D−]=1–p[T−| D−]
Definition
A negative test result in a patient who has the target condition.
1
sensitivity
No of diseased patients with negative test
No
.
.. of diseased patients
1
sensitivity
FN
TP FN
Definition
False- positive: A positive test result in a patient who does not have the target condition. The frequency of a false-
positive result in conditional probability notation is the following:
Specificity
TN
FP TN
https://t.me/medicina_free

66 Medical decision making
To measure 1 − specificity, simply subtract the test specificity from 1. To measure it directly, perform the index test
and the gold standard test in a source population. In terms of the 2
× 2 table:
5.2.4 Predictive value
Predictive value is another way to describe the results of performing the index test and the gold standard test in terms
of the 2
× 2 table. Unlike the measures of test performance, for which the columns of the 2 × 2 table contain the key
measures, the calculation of predictive value uses the rows. Predictive value is, in effect, a post-
test probability. The two
types of predictive value are positive predictive value and negative predictive value.
Positive predictive value
In terms of the 2 by 2 table,
Negative predictive value
P
V
Number of nondiseased patients with negative test
Numbe
rr of patients with negative test
In terms of the cells in the 2 × 2 table,
P
V
TN
TN FN
Differences between predictive value and post- test probability
Predictive value and post- test probability appear to be very similar. Both answer this all- important question, as a result
of this test result, what is the likelihood that my patient has this disease? However, posterior probability and predictive value
are different in important ways. Posterior probability is far more useful than predictive value, but, unfortunately,
authors of medical articles often use “predictive value” when the correct usage is "posterior probability.”
1
specificity
No of diseased patients with positive test
No
.
.. of nondiseased patients
1
specificity
FP
FP TN
Definition
Positive Predictive Value (PV+): The fraction of patients with a positive test who also have the target condition.
P
V
TP
TP FP
Definition
Negative Predictive Value (PV−): The fraction of patients with a negative test result who do not have the target
condition.
P
V
Number of diseased patients with a positive test
Number
of patients with positive test
https://t.me/medicina_free

Measuring theaccuracy ofclinical findings 67
Predictive value
• Defined as the proportion of patients with a test result who have the target condition.
• Calculated from a 2 by 2 table of results in a specific population of patients.
• Refers to a test as used in a particular population; strictly speaking, it does not apply to any other population
because it depends on the prevalence of the target condition in the specified population. Applying a published
predictive value for a disease to a population with a different prevalence of the target condition will cause error
except when the prevalence of the disease is the same in both populations. The larger the difference in the preva-
lence of the target condition, the greater the error.
• There is no mathematical relationship with which to calculate the predictive value in another population.
Post- test probability
• Defined as the probability of the target condition after using Bayes’ theorem to take new information, such as a test
result, into account.
• A mathematical relationship (Bayes’ theorem) makes it possible to calculate the post- test probability for any
patient, regardless of the source population.
• According to Bayes’ theorem, the post- test probability depends on the pre- test probability of the target condition,
which is often a quantitative representation of a subjective opinion about the likelihood of an event.
• Uses sensitivity, specificity, and likelihood ratios obtained from published research. Assumes that the sensitivity
and specificity applies to the patient at hand.
Thus, the predictive value is an observable number obtained from a defined population. It usually will not apply to another
population.
An important caution: A reader might mistakenly believe that predictive value is a measure of test performance. It is not.
Imagine a population consisting of patients who have undergone a diagnostic test. From this population, it is possible to
calculate a predictive value, which will reflect the sensitivity and specificity of the test due to their influence on thenumber
of true-
positive and false- positive test results, which in turn affect the predictive value. The prevalence of the target condi-
tion also affects predictive value through its effect on the number of patients that can have a true- positive result (those with
the target condition) or a false-
positive result (those who do not have the target condition). Therefore, predictive value is a
consequence of test performance and pre- test probability, not an independent measure of test performance.
The next several sections will illustrate the pitfalls of using predictive value as a surrogate for posterior probability.
5.3 How tomeasure diagnostic test performance: ahypothetical example
5.3.1 Description ofthe study
You must evaluate a new radionuclide scanning test for splenomegaly. The purpose of the test is to detect spleen
enlargement that is too slight to be detected by palpating the abdomen. The scan uses radioactively labeled macroag-
gregated iron particles. When injected into the bloodstream, these particles attach preferentially to splenic mac-
rophages, which ingest and then digest them. You decide to measure the accuracy of this test by comparing the size of
the radionuclide scan image of the spleen to the weight of the spleen, which is a good gold standard test for spleno-
megaly. In order to weigh the spleen, someone must remove it. Therefore, you perform the spleen scan on all patients
who are to undergo elective splenectomy.
A description of your study of the radionuclide spleen scan should contain the following information:
• Index test: Macroaggregated iron scan of the spleen.
• Gold standard test: Weighing the spleen after surgical removal.
• Definition of abnormal index test: Spleen silhouette as seen on the scan is more than 1.5 times normal size.
• Definition of disease (splenomegaly): Splenic weight >250 grams.
• Source populations:
• Patients whose clinicians cannot agree on the size of the spleen; some clinicians can feel it, but others cannot.
• Patients whose spleen is not palpable, but they have a disease that is often accompanied by splenomegaly.
• Verified sample: Patients who are about to undergo splenectomy because their spleen is so large that it interferes
with the survival of the formed elements of the blood.
A description of the flow of patients through the study appears in Figure5.5 (see next page).
5.3.2 Description ofresults
The results of using the scan in 50 study patients appear in Table5.2 (see next page).
https://t.me/medicina_free

68 Medical decision making
As shown in Table5.2, 35 of the patients (70 percent) have splenomegaly. Of these 35 patients, 20have a positive
radioactive iron scan. Of the 15 patients with normal size spleens, 10have a normal scan. Therefore,
pT D
|; .sensitivity
20
35
057
pT D
|; .specificity
10
15
067
Predictive value:
P
V
true positives
total positives
20
25
080.
P
V
true negatives
total negatives
10
25
040.
5.3.3 An important limitation ofthe spleen scan study
The study of the spleen scan will be of little value to anyone because of a major error in design, one that occurs often
and is the major reason why studies of test performance can mislead the unwary reader. Everyone who reads a study
of test performance must be alert to its occurrence.
The error in the study design is the choice of the verified sample.
Recall these definitions:
50 patients suspected of
having splenomegaly
35 have
splenomegaly
15 do not have
splenomegaly
20
true-
positive
results
15
false-
negative
results
5
false-
positive
results
10
true-
negative
results
Test + Test –
Test –
Test
+
Figure 5.5 Flow of patients through an illustrative study of a hypothetical test for splenomegaly.
Definition
Source Population: The patients whose findings lead the doctor to order the index test. (Also known as the
“ clinically relevant population.”)
Table5.2 Results ofspleen scan study: veried sample consists ofpatients withpalpable spleens.
Results of the spleen scan
Number of patients
TotalsSpleen weighs >250 grams Spleen weighs <250 grams
Positive True positive
N=
20
False positive
N=5
N=25
Negative False negative
N=15
True negative
N=10
N=25
Totals N=35 N= 15 N=50
https://t.me/medicina_free

Measuring theaccuracy ofclinical findings 69
When the verified sample and the source population are the same, the reader can be sure that the measurements
of sensitivity and specificity apply to patients who receive the test in clinical practice. When, as in the spleen scan
study, the verified sample differs from the source population, the reader cannot confidently use the findings in
patient care.
Source population: Patients whose doctors cannot agree on the size of the spleen; some can feel it, but others cannot;
patients whose spleen is not palpable, but they have a disease that is often accompanied by splenomegaly.
Verified sample: Patients who are about to undergo splenectomy because their spleen is so large that it interferes with
the survival of the formed elements of the blood.
In fact, many patients in the verified sample have such large spleens that they do not need a spleen scan at all!
So far, in this section, we have learned this important paradox of technology assessment:
The patients in a study of test performance often differ from the patients who usually get the test.
The difference between the source population and the verified sample leads to two errors, which are easy to confuse
with one another.
• The predictive value of the spleen scan in the verified sample is not necessarily the same as the post- test probability in the
source population. The prevalence of an enlarged spleen was 70% in the verified sample, whose spleens were so
large that they required surgery to remove them. The prevalence of splenomegaly would likely be considerably
lower in the source population, patients who got a spleen scan because of uncertainty about the size of their
spleens. Therefore, the predictive value in the verified sample is not necessarily the same as the post-
test probabil-
ity in the source population.
• The measurements of test performance in the verified sample may not apply to the source population. The large spleens in
the verified sample would be easier to detect with the scan, increasing the sensitivity of the scan as measured
inthe verified sample. Many patients in the source population would have spleens that would be enlarged but not
large enough to be detected by palpating the abdomen.
Both errors are important. We address the first in the next section, which is titled Pitfalls of Predictive Value. We
focus on avoiding the second error in Section 5.6 of the chapter, which is about how to avoid biased estimates of test
performance.
5.4 Pitfalls ofpredictive value
Applying the predictive value in one population to another risks error because differences in the prevalence of the
disease in the two populations will affect the predictive value of a test. Here, we use the sensitivity (0.57) and specific-
ity (0.67) of the spleen scan in the verified population to calculate the predictive value of the test in a primary care
source population in which the prevalence of splenomegaly is only 20%. Figure5.6 shows the results of using the scan
in 1000 patients from this population (see next page).
Using the numbers from Figure5.6, the predictive value of the test in the hypothetical source population is:
PV
114
114 264
030.
PV
536
536 86
086.
The predictive value in the two populations differ, as seen in Table5.3 (see next page).
Definition
Verified Sample: Patients who receive the index test and the gold standard test (usually a subset of the source
population).
https://t.me/medicina_free

70 Medical decision making
The positive predictive value of the spleen scan is much lower (and the negative predictive value much higher) in
the source population than in the verified sample. Therefore, if two populations have a different prevalence of the target
condition, their predictive values will differ.
As discussed in Chapter4, the predictive value is a number derived from a 2
× 2 table of results from a study of a test.
The predictive value obtained in one population usually will not apply to another population unless the prevalence of
the target condition in the two populations is the same.
A final point about using the term “predictive value:” the words “predictive value” do not precisely describe the
meaning of the term. A person who is unsure of the meaning of “predictive value” would have to look it up. In
contrast,
consider the precision of the language of probability: “probability of the disease if the test is positive.” More words are
required, but their meaning is self-
explanatory. In this book, we will not use predictive value when we mean post- test
probability.
5.5 How toperform ahigh quality study ofdiagnostic test performance
We now return to a crucial question
Do measurements of the sensitivity and specificity of a test apply to the patients that I care for?
Readers of studies of test performance will find themselves asking this question many times and will often be disap-
pointed. A skeptical frame of mind is the best way to read accounts of original research on diagnostic test performance.
This section will itemize what to look for.
This book focuses on the use of a diagnostic test to help resolve diagnostic uncertainty. For that purpose, the clinician
needs to know the state of the patient at the moment of diagnostic uncertainty. Therefore, the principal use of the index
diagnostic test is to characterize the patient at a moment in time. The measurement of the sensitivity and specificity of
a diagnostic test is an example of a cross- sectional design, in which the interval between the index test and gold standard
test is as short as possible.
5.5.1 The features ofa high- quality prospective study ofa diagnostic test
Forming the study population
• Pre-specify the inclusion criteria. They should reflect the reasons that clinicians order the index test.
• Enroll study patients before they have the index test.
1000 patients suspected of
having splenomegaly
200 have
splenomegaly
800 do not have
splenomegaly
114
true-
positive
results
86
false-
negative
results
264
false-
positive
results
536
true-
negative
results
Test +
Test
+T
est –
Test –
Figure 5.6 Results of applying the test for splenomegaly to a hypothetical population of primary care patients with a 20% pre- test probability of
having splenomegaly.
Table 5.3 Predictive value ina veried sample population andthe source population.
Prevalence of splenomegaly in
population
Positive predictive
value
Negative
predictive value
Verified sample population 0.70 0.80 0.40
Source population 0.20 0.30 0.86
https://t.me/medicina_free

Measuring theaccuracy ofclinical findings 71
• Enroll enough patients to ensure reasonably precise measurements of sensitivity and specificity.
• When comparing two index tests, enroll enough patients to permit a statistically valid noninferiority interpretation
of the measured difference in their sensitivity and specificity.
Study conduct
• Ensure that every patient who has the index test also has the gold standard test.
• Obtain a standard set of clinical features from each patient and use them, if necessary, to compare those who have
the index test and the gold standard test with those who have only the index test.
• Record the time interval between having the index test and the gold standard test. Note any patient whose clinical
status changes during the interval.
Interpreting the Index Test and Gold Standard Test
• Choose a gold standard test that accurately measures the patient’s true state.
• Describe the criteria for labeling test results as positive or negative in the study protocol.
• If the test is a visual image, two clinicians must interpret the image independently, discuss any differences, and
form a consensus interpretation.
• Conceal the clinical features of the patient from the study clinicians who interpret the index test or the gold standard
test.
• Conceal the results of the index test from the study clinicians who interpret the gold standard.
• Conceal the results of the gold standard test from the study clinicians who interpret the index test.
The next parts of this section discuss the rationale underlying these rules of research conduct, focusing on (1) insur-
ing that measurements of sensitivity and specificity apply to usual practice; and (2) assuring accurate, unbiased
interpretation of images.
5.5.2 Study characteristics that help ensure that theresults apply tousual practice
• The criteria for including patients in the study should reflect test- ordering decisions in usual practice
Recruitment and enrollment of study patients should maximize external validity by taking place in typical patient
care settings. For example, the study of a test that is typically ordered by specialists should take place in a specialty
clinic setting.
The setting of care can influence the study results. For example, many primary care physicians order a stress electro-
cardiogram (ECG) when they suspect coronary artery disease in a patient with chest pain. They refer some patients to
cardiologists, often those with a test result that is alarming or confusing or because of a poor response to initial treat-
ment. Study patients enrolled in primary care may differ from those drawn from a cardiologist’s practice in ways that
can influence the results of a test.
What to look for in a study: Read the description of the practice settings of the source population. Is it typical of your practice? How
were patients identified for enrollment in the study? Who ordered the index test?
• The study is prospective: Patients are enrolled before doing the index test.
The ideal study enrolls patients in the clinical setting, not at the point of testing. Some patients may not show up for
the index test, and it is important to know how they differ from those that do. In the ideal study, patients enroll early
in the process of investigating their chief complaint. Data collection should occur according to a study protocol that
enforces the discipline needed for good clinical research. A published report should include the criteria for excluding
patients from the study cohort, as well as the criteria for including them.
Retrospective assembly of the study cohort is a much weaker research design. A retrospective study refers to a
defined period in the past. Within this time frame, someone identifies all patients who had the index test and who
among them also had the gold standard test. This approach invites biased results. First, a positive index test is often
the reason for referring a patient for the gold standard test. Second, a positive index test is often sufficient to make a
firm diagnosis. So, some patients do not get the gold standard test. Third, patients with a positive index test, whether
referred or not, are not typical of the entire population who had the index test.
What to look for in a study: A description of the enrollment process. Were patients enrolled prospectively in the clinical setting where
the test- ordering decisions occurred? Were inclusion and exclusion criteria stated clearly?
• The study protocol describes a standard set of pertinent clinical features to obtain from each patient.
A list of the frequency of clinical features in the source population helps clinicians to decide how well the measure-
ments of test performance apply to patients in their practice. A full description of the source population should include
https://t.me/medicina_free

72 Medical decision making
race/ethnicity, gender, age, reasons for testing, duration and severity of the present illness, and past illnesses. By compar-
ing the frequency of the items from the source population and the verified sample, the reader can identify possible selec-
tion biases that might limit the generalizability (external validity) of the study measurements of sensitivity and
specificity.
What to look for in a study: A table listing key clinical features and their frequency in the source population and the verified sample.
Reasons for failure to get the gold standard test.
• Every patient who has the index test also has the gold standard test.
The most important source of error in measuring test performance is due to differences between the verified sample
and the source population that comprises the patients that undergo the test in usual practice. In the example of the
hypothetical spleen scan, the source population was very different than the verified sample. This extreme example
helps to understand the concept of spectrum bias, which we discuss in the next section.
What to look for in a study: A table comparing the clinical characteristics of those who had the gold standard test and those that did
not have it.
5.5.3 Study characteristics that insure unbiased, reproducible interpretation ofthe index test
andthegold standard test
Some index test results are expressed as numbers. The translation of a test result expressed as a number (e.g., the serum
troponin level) into a positive or negative result requires knowing the cut point for the test as applied to a specific
disease (e.g., acute MI).
In the study of the accuracy of diagnostic tests, imaging tests may be the index test (e.g., CT coronary arteriography)
or the gold standard test (e.g., invasive coronary arteriography) or both, as in this sentence. With diagnostic imaging
tests (e.g., a chest radiograph), the classification of the test result requires judgment about interpreting the image as
positive (lung mass present) or negative (lung mass absent). This section addresses the interpretation of imaging tests.
In a study of the sensitivity and specificity of an imaging test, at least two specialists interpret the images indepen-
dently and adjudicate any disagreements.
• If the test is a visual image, two clinicians should interpret the image independently, discuss any differences,
and form a consensus interpretation.
Many index tests and gold standard tests require someone to look at an image of a disease process, place the pattern
into a category, and assign the correct label to the category. Many studies have shown that clinicians often disagree
about the interpretation of a visual image. These interpretive errors can lead to incorrect numbers in the cells of the
2 × 2 table that describes the performance of a test.
The best way to avoid errors in interpretation is to have several people interpret the same image without knowing
the others’ interpretation. If they disagree, the usual procedure is to discuss the disagreement and come to a consensus
interpretation or to ask a third person to participate in forming a consensus opinion.
What to look for in a study: With many tests (e.g., imaging tests), deciding what to call the result involves judgment (i.e., the result
is not a number). When the index test, the gold standard test, or both is such a test, look for evidence that several people independently
evaluated the test and settled any disagreement by consulting each other. The study should report how often two readers disagreed.
The best measure is the kappa statistic, which accounts for agreement by chance alone (any kappa statistic above +0.5 is good agree-
ment; kappa equal to zero signifies agreement due only to chance).
• The study protocol describes the criteria for labeling test results as positive or negative.
The raw data of a study of test performance are four numbers in a 2
× 2 table: the number of true- positive, false-
positive, true- negative, and false- negative results. The source of these numbers are the results of the gold standard test
and the index test. One essential step is being sure that the people who interpret the gold standard test and index test
mean the same thing each time they classify a test result.
Suppose the index test is an imaging test, such as a chest radiograph. A radiologist looks at this image and sees a
pattern (e.g., “patchy honeycomb appearance”). This pattern has certain characteristics that predict the underlying
disease. The leaders of the study want to be sure that the radiologist looks for this pattern and uses the same name for
it each time it is present and never when it is absent. They list the key features of a “positive” finding and provide
criteria for using them to decide when the finding is present.
https://t.me/medicina_free

Measuring theaccuracy ofclinical findings 73
Studies of test performance try to maximize the internal validity of the judgments about whether a test result is
positive or negative. Having two clinicians make the final decision independently is standard research procedure.
They confer, and, if they disagree, ask a third individual to break the tie. Some studies assemble a group of expert
radiologists, train them to read the imaging test according to the study protocol, and monitor for consistency in making
the final interpretation. Some studies save the images and perform the final review at the end of the study.
What to look for in a study: do the authors define the criteria for deciding how to label a result, either with a name or by calling it
“positive” or “negative?” Do they check for consistency in applying diagnostic criteria to images? Do their diagnostic criteria
correspond to the system used in your hospital? Did at least two experts read each image independently?
• The index test and gold standard test are both performed within a short time period.
A cross-
sectional study to measure the performance of a test describes the patient at one point in time. Its intent is to
see how well the index test reflects the actual state of the patient at that time. Any change in the patient’s condition
between the index test and the gold standard test could bias the measures of sensitivity and specificity. The gold stand-
ard test must be performed before the patient’s target condition has changed.
What do look for in a study: look for the average time elapsed between doing the index test and doing the gold standard test. Is it likely
that the target condition may have improved or worsened since the index test?
• The clinician should interpret the index test without knowing the clinical features of the patient. Likewise, the
interpretation of the gold standard test should be concealed when interpreting the index text (and conversely
for the interpretation of the gold standard test).
Doing otherwise can result in two biases: Test-
review bias and Diagnosis- review bias.
An example of Test- review bias. The physician interprets the index test: an exercise ECG result that is on the
borderline between normal and abnormal. Whether to label the exercise ECG results as positive or negative is a close
call. Knowing that the patient had an abnormal coronary arteriogram (the gold standard test) or had many risk fac-
tors for coronary artery disease may influence the clinician toward calling the difficult-to-interpret exercise ECG
abnormal.
An example of diagnosis- review bias. The clinician is interpreting the gold standard test: a coronary arteriogram.
The results must be classified as positive or negative, but it is not clear which interpretation to make. Knowing that the
patient had very abnormal results on an exercise ECG or had many risk factors for coronary artery disease may influ-
ence the clinician toward calling the coronary arteriogram abnormal.
Test-
review bias and diagnosis- review bias have similar effects on measured test performance: they increase the
likelihood that the index test and the gold standard studies will agree, increasing the measured sensitivity and specific-
ity of the index test.
What to look for in a study: Look at the study protocol. It should say that the clinicians who interpreted the index test were blinded
to the results of the gold standard test and vice versa. Those who interpreted one of the tests should not have had any information
about the other test or the clinical characteristics of the patient.
• The number of enrolled patients, both with the target condition and free of the target condition, is sufficient
for precise measurements of the sensitivity and specificity of the test.
The 95% confidence interval for a proportion such as the sensitivity of a test is given by the following relationship:
where p is the proportion in question such as p[T+|D+], the sensitivity of a test. N is the number of patients with the
target condition.
Figure5.7 shows the relationship between the number of diseased patients used to calculate the sensitivity of a test
and the half- width of the 95% confidence interval (see next page). In this example, above 100 patients with the target
condition, the width of the confidence interval changes very little.
951
96
1
%.
confidence interval
pp
N
https://t.me/medicina_free

74 Medical decision making
What to look for in a study: Evidence that, given the measured sensitivity in the study, the number of patients with the target condition is
large enough for the half-width of the 95% confidence interval for the measured sensitivity of the test to be on the flat of the curve as depicted
in Figure 5.7. The same applies for the number of study patients who do not have the target condition and the measured specificity of the test.
A well- reported study of the sensitivity and specificity of a test will address all of the key characteristics of a well-
designed and carefully executed study. The STARD (Standards for Reporting of Diagnostic Accuracy) statement
describes a 25- item checklist of the important features. QUADAS- 2 is the corresponding statement for systematic
reviews of diagnostic test performance (see Bibliography).
5.6 Spectrum bias inthe measurement oftest performance
Spectrum bias affects the measurement of test performance in two phases of the uptake of a test into day- to- day medi-
cal practice. In the first phase, the test is unproven, clinicians do not order the test, and it is difficult to find participants
for a study of the test. In the second phase, clinicians have too much confidence in the test results and will not refer
patients with a negative test result to undergo the gold standard test. The next section describes the mechanism of
spectrum bias in studies of test performance. Later sections describe the effects of spectrum bias on test sensitivity and
specificity.
5.6.1 The first phase oftest evaluation: testing the“sickest ofthe sick” andthe “wellest ofthe well”
The first studies of a diagnostic test are attempts to learn if the test is accurate enough to justify further study. The first
concern is to be sure that the test will be positive for severe disease. A second concern is to be sure that the test is nega-
tive in most healthy people. The easiest way to accomplish these limited goals is to study the test in populations at the
extreme ends of the spectrum of disease severity.
“The sickest of the sick:” One verified sample is the very sickest patients who, because they have advanced disease, are
ideal for learning if the test can detect disease at all. Test sensitivity is apt to be high because advanced disease is
typically extensive and easy to detect. In a broader spectrum of patients, sensitivity usually falls.
Definition
Spectrum bias: The effect of differences in the spectrum of disease in different study populations on measurements
of sensitivity and specificity.
0.2000
0.1800
0.1600
0.1400
0.1200
0.1000
95% confidence half-interval
0.0800
0.0600
0.0400
0.0200
0.0000
Number of patients with target condition
10
20
30
40
50
60
70
80
90
100
110
120
130
140
150
160
170
180
190
200
Figure 5.7 Relationship between 95% confidence interval for the sensitivity of a test and the number of patients used in a study to measure the
sensitivity of a test. In this calculation, the sensitivity of the test is 0.90.
https://t.me/medicina_free
Соседние файлы в папке Библиотека им академика М.И. Перельмана
