Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_2946_Библиотеки_им_академика_М_И_Перельмана
.pdf
Interpreting new information: Bayes’ theorem 55
In contrast to the earlier case, here the post- test probability for the first test is the pre- test probability for the second
test, the CT angiogram. In both cases, the question is whether the sensitivity and specificity of the second test are the
same whether the first test is positive or negative. If the sensitivity and specificity of the CT angiogram are the same
irrespective of the results of the exercise ECG, the post-
test probability after the CT angiogram is conditionally independ-
ent of the exercise ECG results.
Researchers typically measure the sensitivity and specificity of a test in the entire population of patients referred for
the test irrespective of the results of earlier tests, such as an exercise ECG. When the results are used to interpret the
scan in a patient with a positive exercise ECG, the clinician is assuming that the sensitivity and specificity of the scan
are conditionally independent of the results of the exercise ECG.
In reality, clinicians assume conditional independence of the sensitivity and specificity of the tests in a sequence.
Ideally, all study patients should have the first test in a sequence, the second test, and then a definitive test for the target
condition. With this study design, it is possible to calculate the sensitivity and specificity of the second test in those
with a positive result on the first test and in those with a negative result on the first test.
4.8 Using Bayes’ theorem when many diseases are under consideration
This section contains advanced material and can be skipped without loss of continuity.
In this chapter, we have represented the denominator of Bayes’ theorem by the sum of two quantities: (p[D+] × test
sensitivity) and ([1−p[D+] × (1- test specificity), where p[D+] is the probability of the target condition and (1−p[D+])
represents the probability of all other conditions. In effect, we are saying that the contribution of other diseases to the
denominator of Bayes’ theorem is a variable (1−p[D+]) times a constant (1-
test specificity). As p[D+] increases, the con-
tribution of other diseases decreases.
Suppose that two diseases D
1
and D
2
are under serious consideration and that X is a symptom or test result associ-
ated with both diseases. Assuming that only one of them is actually present, Bayes’ theorem would look like this:
pD
S
pX DpD
pX DpDpXD pD pX
1
11
11 22
|
|
|||
[neitheer norneither norDDpDD
12
12
]
More generally, assume that there are n possible diseases and that the patient has only one of these possibilities.
pD
S
pX DpD
pX DpDpXD pD
nn
1
11
11
|
|
||
In this section, we derive Bayes’ theorem in its most general form, one in which every “other disease” contributes to
the denominator of Bayes’ theorem individually (rather than being lumped under the term “no disease”).
Suppose a clinician is considering three diseases (A, B, and C) as the possible cause of a patient’s symptoms (the
probability of other diseases is vanishingly small). The clinician’s goal is to calculate the probability of disease A given
the presence of a clinical finding.
In this instance, we know the prevalence of diseases A, B, and C, and we also know the frequency of finding X in
these diseases (e.g. p[X|A]). Instead of representing two of the diseases (e.g., B and C) as “no disease,” we can list
them separately and perhaps gain additional precision in estimating the conditional probability that disease A is
present.
The prior probabilities of these diseases are p[A], p[B], and p[C], where p[A] + p[B] + p[C]=1. The clinician makes
observation X. The relationships between observation X and the three diseases are given by the following table
(Table4.5, see next page) of conditional probabilities.
Definition
CONDITIONAL INDEPENDENCE: Two tests are conditionally independent if the sensitivity and specificity of
one test do not depend upon the result of the other test.
https://t.me/medicina_free

56 Medical decision making
The probability that finding X occurs in a patient with one of these diseases (e.g., p[X and A]) arises from the definition
of conditionalprobability:
pXA
pX A
pA
|
and
pX
ApAp
XA
and
|
The probability of finding X occurring in all patients suspected of having diseases A, B, or C is the sum of the prob-
ability of its occurrence in each disease:
pX pA pX ApBpXB pC pX C
|||
The probability of disease A given that finding X has occurred– p[A|X]— follows from the definition of conditional
probability.
pA
X
pX A
pX
|
and
Substituting the expressions for p[X and A] and p[X] into the expression for p[A|X], we obtain the following equa-
tion, which is what the clinician is interested in:
pA
X
pA pX A
pA pX ApBpXB pC pX C
|
|
|||
From this 3- disease case, it is a small step to express Bayes’ theorem in its most general form, with n diseases and
subject to the assumption that the patient has only one of the possible diseases. In this equation, X refers to a clinical
finding, the subscript i refers to a specific disease and n refers to the total number of diseases under consideration.
pD
X
pD pX D
pD pX D
i
n
ii
1
11
1
|
|
|
(
Table 4.5 Probability ofa symptom given three diseases.
Disease
Probability of finding X given
disease A, B, or C
A p[X|A]
B p[X|B]
C p[X|C]
Summary
1. A compelling reason to use probability to express uncertainty is to speak the language of Bayes’ theorem, and
thereby to calculate post-
test probability.
2. Bayes’ theorem takes two equivalent but different forms.
• The algebraic form of Bayes’ theorem is especially useful when several diseases are being considered. Down-
side: calculating a post- test probability requires a lot of facility with mental arithmetic. Most people need a
calculator.
• With the odds ratio form of Bayes’ theorem, calculating the post- test odds requires multiplying two numbers:
the pre- test odds and the LR of the test. Downside: moving back and forth between probability and odds
requires some facility with mental arithmetic.
3. Perhaps the most important idea in this book: the interpretation of a test result depends on the pre- test probability of
the disease.
https://t.me/medicina_free

Interpreting new information: Bayes’ theorem 57
Bibliography
Diamond, G.A. and Forrester, J.S. (1979) Analysis of probability as an aid in the clinical diagnosis of coronary- artery disease. New England
Journal of Medicine, 300, 1350–58.
The authors estimated the pre- test probability of coronary artery disease using age, sex, and chest pain history and calculated post- test
probabilities for 4 tests.
Gorry, G.A. and Barnett, G.O. (1968) Sequential diagnosis by computer. Journal of the American Medical Association, 205, 849–54.
A description of the sequential use of Bayes’ theorem to interpret several clinical findings.
Gorry, G.A., Pauker, S.G., and Schwartz, W.B. (1978) The diagnostic importance of the normal finding. The New England Journal of Medicine,
298, 486–9.
A clear discussion of how a negative test result can be evidence against one disease and evidence for another.
Raiffa, H. (1968) Decision Analysis: Introductory Lectures on Choices Under Uncertainty, Addison-
Wesley Publishing Co. Inc., Reading, MA.
Chapter2 of this classic book contains a derivation of Bayes’ theorem.
Rifkin, R.O. and Hood, W.B. (1977) Bayesian analysis of electrocardiographic stress testing. The New England Journal of Medicine, 297, 681–6.
Probabilistic reasoning in test selection has become particularly well- accepted in cardiology practice. This article was very influential.
Weiner, D.A., Ryan, T.J., McCabe, C.H. et al. (1979) Exercise stress testing: correlation among history of angina, ST- segment response and
prevalence of coronary artery disease in the Coronary Artery Surgery Study (CASS). The New England Journal of Medicine, 301, 230–5.
This study illustrates how measuring sensitivity and specificity in subgroups of patients can yield new insights. This study showed that
in men the test performance of the exercise ECG depends on the patient’s history.
4. Bayes’ theorem provides other insights about interpreting test results and deciding when a test is likely to be
useful.
• Rare is the test that will reduce the probability to near- zero disease when the pre- test probability is quite high.
•
Still rarer is the test that will raise the probability of the disease to nearly 1.0when the pre- test probability is quite
low.
•
If you screen for occult disease, take particular care to avoid tests that have a low specificity.
https://t.me/medicina_free

58
Medical Decision Making, Third Edition. Harold C. Sox, Michael C. Higgins, Douglas K. Owens, and Gillian Sanders Schmidler.
© 2024 John Wiley & Sons Ltd. Published 2024 by John Wiley & Sons Ltd.
CHAPTER5
Measuring theaccuracy ofclinical findings
Clinicians must rely on imperfect knowledge to make decisions. Probability is a system for dealing with the uncer-
tainty created by imperfect information, but it is only part of the story. The other part is the imperfect knowledge itself
and how to measure it, which is the topic of this chapter.
When taking a patient’s history, clinicians interpret the answer to a question as evidence for or against a diagnosis.
They must decide if the answer favors the diagnosis or not and whether it is strong or weak evidence. Interpretive
errors occur when:
• The clinician assumes that a negative response to a question is evidence that the hypothesized disease is absent.
Afalse- negative result occurs when a patient has the target condition but answers no to the question.
• The clinician assumes that a positive response is evidence that the hypothesized disease is present. A false- positive
result occurs when a patient has the finding but does not have the target condition.
In Chapter4, we learned that the probability of a disease when a clinical finding is present depends on its sensitivity
and specificity and the patient’s pre-
test probability. This chapter is about measuring sensitivity and specificity and
evaluating published reports of diagnostic test performance.
5.1 A language fordescribing test results
This part of the chapter is about a language for describing test results. We first describe the distribution of test results
in patients who have the target condition and patients who do not have the target condition. Building on this knowl-
edge, we then consider several ways to describe test results.
Most test results are expressed as continuous variables. For example, the serum concentration of troponin I and
troponin T (protein constituents of heart muscle cells) is a measure of heart muscle damage. The serum troponin level
ranges from less than 100 units/ml to greater than 4000 units/ml, depending on the amount of damaged muscle.
Inapatient with a suspected myocardial infarction (MI), each point on this scale of troponin levels corresponds to
aprobability that the patient has had a recent MI. To understand this claim, first consider the results of a test in a
5.1
A language fordescribing test results 58
5.2
The measurement ofdiagnostic test performance 62
5.3
How tomeasure diagnostic test performance: ahypothetical example 67
5.4
Pitfalls ofpredictive value 69
5.5
How toperform ahigh quality study ofdiagnostic test performance 70
5.6
Spectrum bias inthe measurement oftest performance 74
5.7
When tobe concerned about inaccurate measures oftest performance 79
5.8
Test results asa continuous variable: theROC curve 81
5.9
Combining data fromstudies oftest performance: thesystematic review andmeta- analysis 87
A.5.1
Appendix: derivation ofthe method forusing anROC curve tochoose thedefinition ofan
abnormal test result 89
Bibliography 91
https://t.me/medicina_free

Measuring theaccuracy ofclinical findings 59
No. of patients
Serum troponin concentration
Individuals
with MI
Individuals
without MI
P[MI] = 0
P
[MI] = 1.0
Overlap Region
0<p[MI]<1.0
Figure 5.3 Distribution of test results in individuals with the target condition (dashed line) and individuals who do not have the target condition
(solid line).
population of healthy individuals. Figure5.1 shows the distribution of values around the central value for a hypothetical
test.
Figure5.1 shows a symmetric distribution of values. The curve represents a normal distribution, which has many
important statistical properties. The most important of these are the mean of the distribution, which is the unweighted
average of all individuals’ results, and the standard deviation, which is a measure of the degree of spread of the results
around the mean of the distribution. The normal distribution has another important feature: the test results for 68% of
the population fall within one standard deviation of the mean value. The test results for 95% of the population fall
within two standard deviations of the mean (Figure5.2).
The distribution of test results in patients with the target condition often overlaps the values in individuals who do
not have the target condition (Figure5.3).
No. of patients
Serum concentration
Figure 5.1 Results of a hypothetical test in a healthy population.
No. of patients
Serum concentration
1 SD
2 SD
Figure 5.2 One and two standard deviations of a normal distribution.
https://t.me/medicina_free

60 Medical decision making
Notation: When describing the performance of a test (sensitivity, specificity, likelihood ratio), always specify the target
condition to which these measures apply.
The interpretation of a test result depends on knowing the shape of the distribution curves in persons with and
without the target condition and where the curves intersect. Very low test result values indicate “target condition
absent,” and very high values indicate “target condition present.” Uncertainty is not a problem with these extreme
values. Uncertainty is a problem for test result values in the range where the two distribution curves overlap because
the patient could either have the target condition or not have it. Interpreting values in the overlap region requires
Bayes’ theorem.
5.1.1 Defining atest result
• A test result defined as a dichotomous variable
• Below and above the upper limit of normal
• Normal vs. abnormal
• Positive vs. negative
• A test result defined as a continuous variable
A test result as a dichotomous variable: A test result is a gateway to taking an action. As discussed in Chapter3, test
results are typically expressed as a dichotomous variable: positive or negative. The reason for this practice is that these
words connote action. To simplify, “positive test” implies “treat as if the patient has the target condition,” and “
negative
test” implies “treat as if the patient does not have the target condition.”
Imagine that the larger normal distribution curve in Figure5.3 shows the distribution of serum troponin values in
patients with chest pain who do not have an MI, and the smaller normal curve shows those who do have an MI. The
horizontal axis is the concentration of serum troponin. Decision making is easy for very low or very high levels of
troponin. The probability of an MI is zero in the lower of these two ranges of values and 1.0in the upper range.
The diagnosis is more difficult in the region of troponin values where the two normal distribution curves overlap.
As the troponin value increases, more and more patients do have an MI, and so the proportion of patients with an MI
increases. Treatment for an MI benefits those with an MI and may harm those who do not have an MI. Therefore, each
troponin value within the overlap region has a different ratio of benefits and harms from treating the patient as if they
had an MI. Any value of serum troponin within the overlap region could be the cut point at which to treat the patient
as if they had an MI. Picking the cut point value of serum troponin that maximizes net benefit– the treatment threshold
value– is central to decision making and a recurring theme of this chapter.
The language commonly used to describe test performance uses terms that place test results into two categories. As
described in Chapter4, the terms are sensitivity and specificity, which are, respectively, the frequency of a “positive” test
in patients with the target condition and the frequency of a “negative” test in someone who does not have the target
condition. If the test result is expressed as a continuous variable (like serum troponin), we will take a different action
depending on the troponin value at the cut point test result that separates a positive result from a negative result.
Therefore, deciding on this cut point will require careful thinking about the consequences of treating or not treating the
target condition. We discuss three possibilities for choosing the cut point:
Using theupper limit ofnormal asthe cut point
Clinical laboratories usually report the patient’s value and, as a guide to interpretation, the value that corresponds to
the “upper limit of normal.” The “upper limit of normal” is usually a value two standard deviations above the mean
value. Is “above the upper limit of normal” a good way to describe a test result?
The words “upper limit of normal” sound as if they mean the highest result in someone without the target condition;
this point would correspond to the right-
hand end of the distribution labeled “individuals without disease” in
Figure5.3. While everyone with a serum troponin value above that point on the horizontal axis would have an MI, the
serum troponin for many patients with an MI would be below the cut point value.
As described in the preceding section, the purpose of defining a cut point is to link the test result to action, such as
“treat” or “do not treat.” If the upper limit of normal is used as the cut point, many patients with the target condition
would have results below the cut point and would not get treated for a MI. Lowering the cut point a little would result
in treating more patients with a MI but also treating some who do not have a MI, which might cause more harm than
the added benefit for those with a MI. In short, the placing of the cut point for starting treatment involves balancing
benefit to the sick with harm to the well.
The “upper limit of normal” is a poor choice for the cut point for defining a test result as positive or negative.
It accounts for the shape of the distribution of test results in normal people but not in patients who have an MI.
https://t.me/medicina_free

Measuring theaccuracy ofclinical findings 61
Thisisimportant because the intersection of the distributions for patients with and without an MI defines the overlap
region, in which the balance of harms and benefits differs at each value of the test result.
Another reason for avoiding the “upper limit of normal:” the label encourages simplistic reasoning. The unwary
might assume that a value above the upper limit of normal means the disease is present and a value below it means
that the disease is absent. Figure5.3 says otherwise.
The best cut point value for dividing “positive” from “negative” results will depend on the clinical situation. As we
shall see at the end of this chapter, a higher cut point is best if the pre-
test probability is low, and a lower cut point is
best if the patient’s pre-
test probability is high.
Using normal andabnormal asthe cut point
Clinicians often hear the following: “That chest radiograph is abnormal.” Or, “this MRI scan is perfectly normal.” Or, “that
tuberculin skin test is abnormal.” Here, the clinician has not categorized a test result in reference to a cut point along a
continuous scale of test result values. Instead, the point of reference seems to be the normal condition. A central pur-
pose of this book is to show how decision making can become more patient-
centered by measuring uncertainty and
personal preferences. This task is more difficult when the boundary that defines normal and abnormal– and thus
action or no action– is ill-
defined and personal. So, we will not use the terminology of normal and abnormal to help
develop the ideas in this book.
Using thecut point todefine positive andnegative
In practice, most of the ideas that define the discipline of medical decision making depend on expressing numerical
results as a dichotomous variable. Perhaps this usage reflects the nature of decision making: it is the act of choosing
one action and not choosing another.
We follow the conventions of the scientists who conceived the ideas in this book and define the two categories of test
results as “positive” and “negative.” These terms are abstract and do not have a physical counterpart. The cut point,
which divides the range of possible values into “positive” and “negative” regions, is a critical concept for deciding
what a numerical test result means to the patient.
We define a “positive” test result as one that is more extreme than the cut point. It increases the probability that the
patient has the target condition. A “negative” result is less extreme than the cut point, and it decreases the probability
of the target condition.
A test result asa continuous variable
Dividing the range of values into positive and negative regions has an important shortcoming. It disregards the infor-
mation contained in the magnitude of the numerical result. In general, a larger value of the serum troponin means a
greater likelihood that the patient has an MI, and a very low value means a smaller likelihood. A result that moves the
probability of the target condition closer to 1.0– one that reduces uncertainty to a minimum– can be useful. However,
as we will see in later chapters, crossing a threshold probability for taking action should be sufficient to take that
action. When reducing uncertainty any further will not change management, the resources required are not well spent.
Referring to Figure5.4 (see next page), within the range of test results that occur in both patients with the target
condition and patients who do not have the target condition (the overlap region), a test result is consistent with either
having the target condition or not having it. If the overlap region contains the cutoff value, it defines a positive and
negative test result. Patients with the target condition can have a negative result (false-
negative), and patients who do
not have the target condition can have a positive result (false- positive) (Figure5.4).
In Figure5.4 (see next page), we first introduced the problem of balancing harms and benefits of treatment versus
no treatment when the diagnosis is uncertain.
Within the overlap region of the two distributions of individuals, the problem is distinguishing between persons who have the disease
of concern and persons who do not. The cut point divides the diseased population into those with positive results, who are treated and
benefit, and those with negative results, who are not treated and suffer harm. The cut point also divides the nondiseased population.
Definition
Cut point: The test result that divides the spectrum of test results into a test-positive region and a test-negative region.
https://t.me/medicina_free

62 Medical decision making
A small number of them are treated and suffer some harm, and most are not treated and do fine. The task is to find the cut point that
best balances the potential benefits and harms of treatment for the individual patient, who may or may not have the disease of concern.
We will revisit this problem in the appendix to this chapter and in Chapter13.
In this part of the chapter, we have learned how the cut point test result leads to a language for describing misleading
test results (e.g., false- positives and false- negatives and their truthful counterparts– true- positives and true- negatives).
These terms apply to any information that might affect our certainty about the patient’s true state: the history, physical
examination, or a diagnostic test result. Using a cut point deprives us of the information conveyed by extreme test
results, but it is in keeping with a focus on getting just the information we need to decide.
5.2 The measurement ofdiagnostic test performance
The key measure of any diagnostic information is its ability to discriminate between the target condition and all other
conditions. What do we mean by “discriminate?” A test that perfectly discriminates between the target condition and
all other conditions is positive in all patients with the target condition and negative in all patients who do not have it.
Most tests fall far short of this ideal performance.
5.2.1 How tomeasure test performance
Bayes’ theorem defines the measures of test performance. The odds ratio form of Bayes’ theorem (Chapter4) states that:
This form of Bayes’ theorem reminds us that the function of diagnostic information is to change the probability of disease.
The likelihood ratio is the most meaningful measure of diagnostic information because it directly shows the effect on
the probability of disease. In other words,
To know the performance of a test, know its likelihood ratio.
Posttest odds pretest odds likelihood ratio
Number of patients
Serum concentration
Cutoff value
to define
abnormal result
Diseased
individuals
False-positive result
False-negative result
Normal
individuals
Figure 5.4 The figure shows the distribution of results in patients with and without the target condition. The cut point divides both distributions into
regions that define false- negative results and false- positive results.
Definition
TEST PERFORMANCE: A test’s ability to detect and discriminate.
https://t.me/medicina_free

Measuring theaccuracy ofclinical findings 63
Recall the definition of the likelihood ratio, as it emerges from the derivation of the odds ratio form of Bayes’
theorem:
where p[R|D+] is the probability of a test result in patients with the target condition and p[R|D−] is the probability of
the result in people who do not have the target condition.
This line of argument leads us to the following method for measuring test performance:
1. To measure the performance of a test for a disease, first perform the test in patients who are known to have the
target condition and in patients who are known to be free of it (but might have other diseases).
2. Then, calculate the frequency of a test result in patients with the disease and in patients who do not have the
disease.
This simple prescription is hard to achieve to practice. Many studies of test performance have had serious flaws.
These flaws can lead to inaccurate test performance measurements that can lead to incorrect interpretation of test
results and potentially to mistakes in patient care.
We will first learn how to measure test performance. Then, we will learn how to evaluate articles about test perfor-
mance and identify studies whose measurements of test performance we can rely upon.
Recall that the first step in measuring the performance of a test is to determine the frequency of a test result in
patients with the target condition and in patients known to be free of it. Deciding if the patient has the disease requires
doing another test, which is usually called the “gold standard” test (also “diagnostic reference standard”).
In the ideal study of a test, each patient in a source population containing N patients undergoes both the index test
and the gold standard procedure. If the test results are expressed as a dichotomous variable (positive or negative), a 2
by 2 table is a convenient way to display the results of the study. The name “2 × 2” refers to the array of the four cells
that contain the primary results, labeled TP, FP, FN, and TN in Table5.1 (see next page).
Definition
“Gold Standard” Test: The procedure which defines the true state of the patient in a study of test performance
(also known as “diagnostic reference standard”).
Definition
Index Test: The test whose performance is being measured.
Definition
Source Population: The patients whose findings lead a clinician to order the index test. (Also known as the
“
clinically relevant population.”)
Definition
Verified Sample: Patients who receive the gold standard test to verify their disease status (ideally, identical to the
source population but too often, a sample, or subset, of the source population).
Likelihood ratio
pRD
pRD
|
|
https://t.me/medicina_free

64 Medical decision making
If the index test and the patient’s true state were perfectly concordant, we would have that elusive animal, the perfect
test, one with no false positive results and no false negative results. Misleading test results nearly always occur with
most tests. Measures of the degree of concordance, such as the sensitivity and specificity of a test, are part of the basic
vocabulary of medicine.
5.2.2 Measures ofconcordance between index test anddisease state
Two types of test results reflect the true state of the patient: true positive results and true negative results. The relative
proportions of these two types of results completely characterize the performance of a test.
The Sensitivity of a diagnostic test
In conditional probability notation, the sensitivity of a test result is:
p[positive test result|disease] or p[+|D+] which means “the probability of a positive test result if the patient has the
target condition.”
S
ensitivity
Number of diseased patients with positive test
=
NNumber of diseased patients
To measure the sensitivity of a test, perform the index test and the gold standard test in a source population.
In terms of the 2 × 2 table, the sensitivity of a test, is:
The specificity ofa diagnostic test
In conditional probability notation, the specificity of a test result is:
ppTDnegative test result disease is absent or
which means “the probability of a negative test result if the patient does not have the target condition.”
Table5.1 Test results dened.
Results of the index test
Results of the gold standard test
TotalsPositive Negative
Positive True-
positive (TP) False- positive (FP) TP + FP
Negative False- negative (FN) True- negative (TN) FN + TN
Totals TP + FN FP + TN N=TP+FN+FP+TN
Definition
Sensitivity: Probability of a positive test in a patient with the target condition.
S
ensitivity
TP
TP FN
Definition
Specificity: The probability that a patient who does not have the target condition has a negative test.
S
pecificity
Number of nondiseased patients with negative t
=
eest
Number of nondiseased patients
https://t.me/medicina_free
Соседние файлы в папке Библиотека им академика М.И. Перельмана
