Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / @xirurgi_2025 / @xirurgi_2025 - 1270 - файл
.pdf
Evidence-based History and Examination 13
(1) Detecting pneumonia: In patients with acute respiratory
(2)
chest pain, “dysphagia” is reported in 4% of patients found
to have coronary disease and in 20% of patients with another
cause of chest pain. Therefore,
Infinity
Zero
No change
e
https://t.me/medicina_free
Likelihood Ratios
When a symptom or sign is present (or absent), how do we know
how useful that finding is in making a diagnosis? Likelihood ratios (LRs) are diagnostic weights. The likelihood ratio is the probability of the finding in someone with the disease over the
probability of the finding in someone without the disease. Thus, if
a finding is equally likely in people with and without the disease,
the likelihood ratio is 1 (i.e., unhelpful). Each finding from the
history and physical examination is associated with a unique LR,
a number whose values ranges from zero to infinity. An LR greater
than 1.0 increases the probability of disease, and the higher the
value of the LR, the greater the increase in probability. An LR of
less than 1.0 decreases the probability of disease, and the lower
the value of the LR, the greater the reduction in probability (see
Figures 2.2 and 2.3).
One simple method of interpreting LRs is to memorise the
association between three LR values – 2, 5, and 10 – and the first
three multiples of 15 – 15, 30, and 45. A finding with an LR of 2
increases the absolute probability by around 15% (that is, the clinician adds 15% to the pre-test probability); a finding with an LR
of 5 increases the probability by around 30%, and one with an LR
of 10 increases the probability by around 45%.
For those LRs less than 1.0, the clinician simply inverts the 2, 5,
and 10 ‘rule’ (that is 0.5, 0.2, and 0.1). A finding with an LR of 0.5
decreases the probability by around 15%; one with an LR of 0.2
decreases the probability by around 30%, and one with an LR of
0.1 decreases the probability by around 45%. Provided clinicians
round off final probabilities greater than 100% to 100%, and those
less than 0% to 0%, this method suffices for the purposes of
clinical reasoning.
Box 2.11 summarises the absolute changes in probability for
the most used LRs. Findings with LRs greater than 3 or less than
0.3 are most helpful because these values identify findings that
either increase or decrease probability by 20–25% or more.
Can LRs Be Combined?
LRs can be combined only if the two findings are independent of
one another (independence implies the LR for the first finding is
complaints, “percussion dullness” is found in 18% of patients
with pneumonia and in 6% of patients with another cause of
respiratory distress. Therefore,
for percussion dullness
LR
in detecting pneumonia
Detecting coronary artery disease: In patients with chronic
18
== 3.0
6
10
Increase
probability
Decrease
probability
Figure 2.3 Likelihood ratios: diagnostic weights. Clinicians should classify
LRs into three groups: those with values greater than 1.0 increase probability; those with values less than 1.0 decrease probability; and those with
values near 1.0 change probability very little or not at all.
Box 2.11 Likelihood ratios and bedside estimates
Likelihood ratio Approximate change in probability*
0.1 −45%
0.2 −30%
0.3 −25%
0.5 −15%
1 No change
2 +15%
3 +20%
4 +25%
5 +30%
6 +35%
7
8 +40%
9
10 +45%
*These changes describe absolute increases or decreases in probability.
From McGee (2002). J Gen Intern Med; 17: 646–9.
5
2
1
0.5
0.2
0.1
+45%
+30%
+15%
No chang
–15%
–30%
–45%
for dysphagia
LR
in detecting coronary
artery disease
Figure 2.2 Likelihood ratios: examples. From McGee (further resources).
4
==
20
0.2
the same whether or not the second finding is present). For
example, typical angina (an LR of 5.8) and hyperlipidaemia (an LR
of 2.2) are likely to be independent because the accuracy of a history
of typical angina is unlikely to be affected by the presence or

14 ABC of Clinical Reasoning
https://t.me/medicina_free
absence of hyperlipidaemia. To combine findings, the clinician can
simply multiply the two individual LRs (5.8 × 2.2); the resulting
product (12.7 or a+50% probability) becomes the LR for combined
‘typical angina and hyperlipidaemia’. Alternatively, the clinician
could first apply typical angina (LR of 5.8 or a+35% probability),
then hyperlipidaemia (LR of 2.2 or a+15% probability) to obtain
the increment in probability for the combined findings (35%+15%
or a+50% probability).
Clinicians should not combine the LRs of more than two
individual findings unless clinical studies have proven that the
findings are independent. If there is any possibility that the
individual findings are dependent on each other, their LRs should
not be combined (for example, typical angina and ‘duration of
pain< 5 minutes’ should not be combined, because pain lasting
less than 10 minutes after rest or nitro-glycerine is a criterion for
stable typical angina).
The Limitations of LRs
Statistical calculations are appropriate only when the clinical problem
is defined by a diagnostic (or reference) standard, such as laboratory
testing or clinical imaging (Figure 2.4). Examples, and their reference
standards, are pneumonia (chest radiographs), ascites (ultrasonography), coronary artery disease (coronary angiography), anaemia
(full blood count), and hyperthyroidism (thyroid function tests). In
each of these disorders, the evidence-based approach compares findings from the history or examination to the accepted reference standard and identifies the findings most accurately predicting the results
of that standard. Since many clinical problems lack reference standards, evidence-based reasoning using LRs is not always applicable.
For these problems, empiric observation based on the clinician’s
prior knowledge and experience of similar patients – what the clinician sees, feels, and hears at the bedside – remains the sole diagnostic
standard and LRs cannot be used.
Although LRs describe how the probability changes, they cannot
determine the pre-test probability of a disease. For example, the LR for
the physical finding ‘fluid wave’ in detecting ascites in patients with
abdominal distension is 5.0 (a+30% probability). If the clinician
works in a hepatology practice in which 60% of all patients with
abdominal distension have ascites (that is a pre-test probability of
60%) the finding of a fluid wave is diagnostic (that is 60%+30% or
WHAT IS THE
DIAGNOSTIC STANDARD?
a 90% probability of ascites). On the other hand, if the clinician
works in a community practice where only 20% of patients with
abdominal distension have ascites (the other 80% have increased
abdominal fat or gas), the presence of the fluid wave is less conclusive (20%+30% or a 50% probability of ascites). Proper application
of evidence-based medicine here requires intimate knowledge of the
types of diseases found in one’s own practice.
The Future of the History and Physical
Examination
Increasingly, researchers are comparing clinical findings to diagnostic standards to reveal LRs for a wide variety of clinical disorders. This is through diagnostic accuracy studies reported to the
STARD criteria [13]. These include:
•
Both the test (clinical symptom, sign, or laboratory test) and
diagnostic standard are clearly defined
•
All enrolled patients have symptoms suggestive of the diagnosis
under study
•
Determination of the test result is blinded from determination
of the diagnostic standard
•
The study presents enough information to allow calculation of
LRs and their confidence intervals.
Clinicians applying this approach can focus on findings with
greatest diagnostic accuracy. Nonetheless, this does have limitations. Even when a problem has been studied, conclusions often
rest on relatively few patients. Whether diagnostic accuracy
depends on clinical technique is largely unaddressed, although
the few studies on this subject show diagnostic accuracy with students as observers is the same as with specialists, provided the
finding is well-defined. Finally, most literature on the subject
focusses on individual findings, although it is well known that
expert clinicians typically combine many findings simultaneously
when diagnosing disease.
Point of care ultrasound is increasingly being used in acute care
settings as an extension of the physical examination (e.g., to
estimate volume status, or differentiate fluid from consolidation in
the lungs). However, the same caveats for all diagnostic tests apply
(see Chapter 3) – the history and physical examination remains
fundamental in establishing the clinical probability of disease and
ultrasound ‘findings’ need to be interpreted in light of this. Point of
care ultrasound has several limitations and should be seen as a
decision aid pending more definitive investigations.
Clinical imaging or laboratory
Pneumonia
Ascites
Coronary artery disease
Anemia
Hyperthyroidism
Evidence-based reasoning
can be used
Figure 2.4 Can evidence-based reasoning be used?
Empiric observation
Cellulitis
Parkinson disease
Trochanteric bursitis
Pericarditis
Serotonin syndrome
Evidence-based reasoning
does
not apply
Developing Skills in Teaching
It is challenging for busy clinicians to be experts in clinical communication and in teaching evidence-based history and physical
examination. This has contributed to a decline in bedside teaching
since the 1960s. It is however both a patient and student-centred
activity. The scope for evidence-based history and examination is
exciting, with potential to improve patient safety. Role-modelling
of reflective practice by bedside teachers can assist learners in
developing resilience and dealing with the uncertainty of clinical
practice. Careful planning and engagement with patients can help
develop clinical teachers. Box 2.12 lists some tips for teaching

Evidence-based History and Examination 15
https://t.me/medicina_free
Box 2.12 Tips for teaching evidence-based physical
examination
Practice teaching concepts of diagnostic accuracy
•
• Practice estimating the pre-test probability of disease
• Practice teaching methods to estimate post-test probabilities
• Know where to find evidence-based physical examination data
and prepare to use it
• Prepare an answer to the question, ‘Can the likelihood ratios of
multiple findings be combined?’
• Answer the common question, ‘Why should we examine patients
if it is so unhelpful?’
• Teach the basics of evidence-based physical examination and
prepare students for bedside teaching
• Orientate the patient to the purpose of the teaching and explicitly
discuss evidence-based physical examination
• Encourage students to commit to their own description of
findings
• Encourage students to commit to a next step in management
•
Facilitate deliberate practice and give feedback to learn evidence-
based physical examination
• Acknowledge uncertainty and follow up on unresolved issues
Adapted from Mookherjee S, Hunt S, Chou CL. (2015). Twelve tips
for teaching evidence-based physical examination. Medical Teacher;
37(6): 543–550.
evidence-based physical examination. Seeing variations in demonstration of the physical examination is a source of discomfort
for students, particularly around assessment. Reasons for variation in technique should be discussed with learners to help them
manage their uncertainty and to apply these critical skills.
Summary
Practicing evidence-based history and examination is challenging
but rewarding. Since the history and examination is so critical to
the patient’s care, a robust evidence base is essential, and merits
increased research. An initial step for learners is establishing the
importance of the history and examination not only for initial
formulation of the patient’s problem list and differential diagnosis, but the correct interpretation of any subsequent investigations. Clinical teachers should be supported in developing their
own confidence and skills in teaching evidence-based history and
examination.
tion, including this revision.Thanks also to Lucille Middleton,
Graduate Entry Medicine student at the University of Nottingham,
UK, for contributing Box 2.1 to this chapter.
References
1. Hampton JR, Harrison MJ, Mitchell JR et al. (1975). Relative contributions of history-taking, physical examination, and laboratory investigation to diagnosis and management of medical outpatients. British Medical
Journal; 31; 2(5969): 486–489.
2.
Silverman J, Kurtz SM and Draper J. Skills for communicating with
patients, 3
Silverman J. The consultation. In: Cooper N and Frain J (Eds). ABC of
3.
Clinical Communication. Wiley-Blackwell, 2018.
4. Frain J and Abdalla M. Teaching clinical communication. In: Cooper N
and Frain J (Eds). ABC of Clinical Communication. Wiley-Blackwell,
2018.
5.
Kilian A, Upton LA and Sheagren JN. (2020). Reorganizing the history of
present illness to improve verbal case presenting and clinical diagnostic
reasoning skills of medical students: the all-inclusive history of present
illness. Journal of Medical Education and Curricular Development; 7:
2382120520928996.
Swap CJ and Nagurney JT. (2005). Value and limitations of chest pain
6.
history in the evaluation of patients with suspected acute coronary syndromes. JAMA; 294(20): 2623–2629.
7. Elieson SW and Papa FJ. (1994). The effect of various knowledge formats
on diagnostic performance. Academic Medicine; 69(10 Suppl): S81–S83.
8. Thomas KE, Hasbun R, Jekel J and Quagliarello VJ. (2002). The diagnostic
accuracy of Kernig’s sign, Brudzinski’s sign, and nuchal rigidity in adults
with suspected meningitis. Clinical Infectious Diseases; 35(1): 46–52.
Yusuf S, Hawken S, Ounpuu S et al. (2004). Effect of potentially modifi-
9.
able risk factors associated with myocardial infarction in 52 countries
(the INTERHEART study): case-control study. Lancet; 364: 937–952.
10. Paley L, Zornitzki T, Cohen J et al. (2011). Utility of clinical examination
in the diagnosis of emergency department patients admitted to the
department of medicine of an academic hospital. Archives of Internal
Medicine; 171(15): 1393–1400.
11. Verghese A, Charlton B, Kassirer J et al. (2015). Inadequacies of physical
examination as a cause of medical errors and adverse events: a collection
of vignettes. The American Journal of Medicine; 128(12): 1322–1324.
12. Holboe ES. (2004). Faculty and the observation of trainees’ clinical skills:
problems and opportunities. Academic Medicine; 79: 16–22.
Cohen JF, Korevaar DA, Altman DG et al. (2015). STARD guidelines for
13.
reporting diagnostic accuracy studies: explanation and elaboration. BMJ
Open; 6: e012799. doi:10.1136/bmjopen-2016-012799.
rd
Ed. CRC Press, 2013.
Acknowledgements
Thanks are due to my co-author of the first edition of this chapter,
Steven McGee, Emeritus Professor of Medicine, University of
Washington, Seattle, USA, whose work and contribution to evidence-based history and examination continues to be an inspira-
Further Resources
1. McGee S. Evidence-based physical diagnosis, 5th Ed. Elsevier/Saunders,
2021.
Talley N and O’Connor S. Clinical examination, 9
2.
th
Ed. Elsevier, 2021.

https://t.me/medicina_free

CHAPTER 3
https://t.me/medicina_free
Choosing and Interpreting
Diagnostic Tests
Nicola Cooper
OVERVIEW
• Test results are affected by a number of factors which the clinician
has to take into account
There is no such thing as a perfect test
•
• The interpretation of new information depends on what you
believed beforehand, based on your assessment of the patient
• Predictive values combine information about sensitivity, specificity,
and prevalence and indicate how likely a test result is to be correct
• Thresholds provide a useful way of thinking about whether a test
should be performed at all
Introduction
The history and physical examination provide a differential diagnosis and/or problem list. This is refined further using diagnostic
tests. The appropriate selection of tests depends of the quality of
the history and physical examination. Test results then have to be
interpreted in light of the patient’s history and examination findings because test results are affected by a number of factors (see
Box 1.1):
•
How ‘normal’ is defined
•
Factors other than disease that influence test results
•
Operating characteristics
•
Sensitivity and specificity
•
Prevalence of disease in a population
Unfortunately, commonly used measures of test accuracy, such as
sensitivity and specificity, are poorly understood. A systematic
review of 24 studies found that most qualified healthcare professionals were poor at providing definitions of sensitivity and specificity, and were poor at estimating the post-test probability of
disease [1]. This chapter aims to introduce key concepts and provide further resources for this important area of clinical reasoning.
Box 1.1 Tests are affected by a number of factors
Factor Explanation
How ‘normal’ is
defined
Factors other than
disease that influence
test results
Operating
characteristics
Sensitivity and
specificity
(see Box 3.3)
Prevalence of disease
in a population
(see Box 3.5)
‘Normal’ can refer to values within the
•
reference range for the population to which
the patient belongs
•
It can also refer to a value below or above a
pre-determined cut-off point designed to
maximise true positives and minimise false
positives
It can also be an ‘abnormal’ result that is
•
actually normal for the particular context in
question
These are biological and/or laboratory factors
•
that make test results ‘abnormal’ when they
are not, or vary when there has not been a
true change
• This refers to the method of performing the
test itself which, if not optimal, can affect its
accuracy
•
The sensitivity of a test refers to its ability to
correctly identify patients with the disease
• The specificity of a test refers to its ability to
correctly identify patients without the
disease
•
Sensitivity and specificity are characteristics
relating to the accuracy of a test relative to a
reference standard
The prevalence of disease in a population
•
can significantly alter the predictive value of
a test
The positive predictive value is the
•
proportion of people with a positive test
result who truly have the disease
•
The negative predictive is the proportion of
people with a negative test result who do
not have the disease
ABC of Clinical Reasoning, Second Edition. Edited by Nicola Cooper and John Frain.
© 2023 John Wiley & Sons Ltd. Published 2023 by John Wiley & Sons Ltd.

18 ABC of Clinical Reasoning
(standard deviations from the mean)
https://t.me/medicina_free
How Normal Is Defined
Many diagnostic test results are expressed as continuous variables
on a numerical scale and many quantitative measurements in
human populations have a Gaussian (normal) distribution. The
‘normal’ range is defined as those values that encompass 95% of
the healthy population, or two standard deviations from the mean.
This means that 2.5% of the healthy population will have values
above, and 2.5% of the population will have values below, the
normal range. For this reason, it is more appropriate to use the
term ‘reference range’ (see Figure 3.1). Diagnostic test results in
people with a disease also have a Gaussian distribution but with a
different mean and reference range. In some diseases there is no
overlap between results from the abnormal and normal population,
but in some diseases there is. In the latter, the greater the difference
between the result and the reference range of the normal
population, the higher the chance that the person has the disease.
Arbitrarily dividing a range of values into ‘normal’ and
‘abnormal’ has disadvantages – it does not take into account the
magnitude of the result. For example, a highly sensitive troponin
T result in a patient with chest pain is more likely to indicate myocardial injury when the value is very high, as opposed to slightly
raised. Some test results have a binary classification (‘normal’ vs
‘abnormal’), for example, an exercise electrocardiogram (ECG)
looking for signs of ischaemic heart disease. However, in deciding
where the cut-off point between ‘normal’ and ‘abnormal’ should
be, there is a trade-off between sensitivity (true positives) and
specificity (true negatives). The optimum cut-off point is calculated using receiver operating characteristic (ROC) analysis,
which is described in more detail later.
In medicine there are some situations when a normal result is
abnormal, and an abnormal result is normal. For example in a
clinically severe asthma attack when one expects the PaCO
low, a normal PaCO
on an arterial blood gas is not normal at all
2
and indicates life-threatening asthma. On the other hand, a raised
d-dimer is normal in pregnancy. So what is ‘normal’ and
‘abnormal’ has to be interpreted in light of the clinical picture.
Clinicians who use diagnostic tests should have a good working
knowledge of the tests they use in their everyday practice, and
how they should be interpreted in light of the patient’s history and
examination findings.
Normal population
Number of people
Reference range
–3 –2 –1 1 2 3
Figure 3.1 Normal distribution.
Mean
Test Result
to be
2
Factors Other than Disease Which
Influence Test Results
There are a number of factors other than disease which influence
test results. They include:
•
Age
•
Sex
•
Ethnicity
•
Pregnancy
•
Body position
•
Chance
•
Spurious (in vitro) results
•
Lab error
•
Critical difference values
For example, normal values for paediatric blood results can be significantly different to those of adults. Old people often have a
normal white cell count in the presence of infection, and can have a
significantly reduced glomerular filtration rate with a normal creatinine. Men have slightly different reference ranges to women (e.g.,
for haemoglobin) and healthy black adults may have an ‘abnormal’
12-lead electrocardiogram (due to early repolarisation) that can
resemble serious disease, but is in fact a ‘normal variant’ [2].
Pregnancy significantly alters many test results due to the
physiological changes that occur, particularly in the third trimester. A large foetus splints the diaphragm and compresses the
lungs causing supine hypoxaemia as well as a respiratory alkalosis
(important facts to remember when considering the possibility of
pulmonary embolism in a pregnant woman). Circulating volume
increases by 50% in late pregnancy causing a flow murmur, tachycardia, and a rightward axis on the 12-lead electrocardiogram.
Kidneys also swell as a result, and renal ultrasound shows increased
size and dilatation.
Body position is important in some tests, for example, lung
function and tests where the patient has to lie in a certain position
to get optimal images. Finally, a test result may be abnormal by
chance (e.g., the patient is an outlier on the normal curve); the
result may be spurious (e.g., hyperkalaemia caused by haemolysis
or some haematological conditions); or may be due to lab error
(e.g., as a result of a technical or human error). It is always worth
pausing before acting when a very unexpected test result crops up.
Lab results also vary in the same person at different times. The critical difference, also known as the reference change value, is the smallest
difference between sequential laboratory results in the same patient
which is likely to indicate a true change. Let’s imagine a person has
their cholesterol measured every single day. The result will not be
identical every time. The reason for this is natural biological variation but also lab variation. The combination of the two is the
critical difference – the amount by which the test can vary before
it can be considered a true change. This is calculated using
knowledge of normal intra-individual variation and lab variation
for different tests. The critical difference is different for different
lab tests. Some calculated critical difference values for common
biochemistry results are shown in Box 3.2. For a person having
their serum cholesterol monitored, the critical difference is 17%.
An initial value of 5.2 mmol/L can therefore vary between
4.3mmol/L and 6.1 mmol/L without being a true change.

Choosing and Interpreting Diagnostic Tests 19
https://t.me/medicina_free
Box 3.2 Calculated critical difference (CD) for some common
biochemistry results
Test CD as %
Albumin
Alkaline phosphatase
Aspartate aminotransferase (AST)
Bilirubin
Calcium
Cholesterol
Glucose
Total protein
TSH
Urea
Uric acid
Data from Professor Trefor Higgins, Department of Laboratory Medicine
and Pathology, University of Alberta.
11.2
37.1
27.7
47.5
6.1
17.0
9.9
11.2
63.0
28.9
25.2
Operating Characteristics
Before ordering a test, it is important to be aware of certain
operating characteristics of the test. This refers to the method of
actually performing the test itself. For example, measuring lung
function requires that the patient be able to hear, understand, and
co-operate with instructions, as well as hold their breath. Exercise
electrocardiograms require patients to be able to walk briskly and
cannot be accurately interpreted in people who have left bundle
branch block.
Some tests are highly operator dependent – in other words, the
skill of the operator influences the results and the report provided.
Ultrasound is the best example of this, as dynamic images have to be
skilfully interpreted by the sonographer. For radiology investigations in general, the interpretation of results can be highly influenced by the patient’s body habitus or clinical state. In ultrasound,
for example, morbid obesity can make getting good views difficult,
and in people of all sizes, intra-abdominal organs can be obscured
by bowel gas. In computed tomography, images can be severely
degraded by movement artefact, or interpretation can be affected by
whether or not contrast was used, and whether it circulated as anticipated to get optimal images. If a report says, ‘Limited views due to
… but within these limitations, no abnormality detected’ consider
whether it is in fact a non-diagnostic scan, rather than a ‘normal’
scan. It is also important that radiologists, as well as other clinicians
such as physiologists, are provided with a clear clinical question and
key information in the history, past medical/surgical history, and
physical examination. This is so that ‘abnormalities’ or incidental
findings can be interpreted in light of the clinical context.
Sensitivity and Specificity
The sensitivity of a test refers to its ability to correctly identify
patients with the disease. The specificity of a test refers to its ability
to correctly identify patients without the disease. Even a test with
Box 3.3 Sensitivity and specificity
Disease No disease
Positive test A
(True positive)
Negative test C
(False negative)
The sensitivity of a test refers to its ability to correctly identify
patients with the disease, i.e. A/(A+C) × 100.
The specificity of a test refers to its ability to correctly identify
patients without the disease, i.e. D/(D+B) × 100.
B
(False positive)
D
(True negative)
a high sensitivity, for example 95%, will miss 5% of people with
the disease. Unfortunately, there is no such thing as a perfect test.
Test results consist of ‘true positives’ and ‘false positives’; ‘true
negatives’ and ‘false negatives’. Box 3.3 illustrates this. Tests differ
in their sensitivity and specificity for detecting certain diseases, so
clinicians need to have a sound working knowledge of the accuracy of the tests they use on a day-to-day basis.
A very sensitive test will detect most disease but generate
abnormal findings in healthy people. We see this with the aptly
named high sensitivity troponin T. On the other hand, a very
specific test may miss significant disease but is likely to establish
the diagnosis beyond doubt when the result is positive. You may
have heard of the acronyms ‘SNOUT’ and ‘SPIN’. SNOUT stands
for ‘sensitive test when negative rules out the disease’ and SPIN
stands for ‘specific test when positive rules in the disease’.
However, SNOUT and SPIN are misleading. This is because the
diagnostic power of any test is determined by both its sensitivity
and specificity, as well as the prevalence of disease in the
population – more of that later. The trade-off between sensitivity
and specificity is explored in what is termed a ‘ROC analysis’.
ROC Analysis
ROC stands for ‘receiver operating characteristic’ – so called
because it was developed by radar engineers during World War II
for discriminating enemy objects in the battlefield. It is also known
as the ‘relative operating characteristic’ because it compares two
operating characteristics (true positive results and false positive
results) at various settings (see Figure 3.2). It is used in medicine to
select the best cut-off point for a test in a way that maximises true
positives while minimising false positives. ROC analysis is conducted in a research setting whenever investigators measure the
ability of a test to detect a diagnosis in a population with the disease and exclude the diagnosis in those without it. (Of course, the
results of the analysis also depend on what study population was
used – if the same performance is expected in practice, the test
must be used in a similar population). For example, if we define an
exercise electrocardiogram as ‘abnormal’ when there is at least 0.5
mm of ST depression, we could pick up every case of ischaemic
heart disease but generate many false positives. On the other hand,
if we define an exercise electrocardiogram as ‘abnormal’ when
there is at least 2 mm of ST depression, we could detect most cases
of clinically important ischaemic heart disease but with far fewer
false positives, which is far more practical.

20 ABC of Clinical Reasoning
1.0
1.0
False positive rate
True positive rate
Prior Probability
Posterior Probability
1.00
1.00
//
https://t.me/medicina_free
Perfect test
Good test
Moderate test
0.75
+test
Positive
shift
0.5
0 0.5
Figure 3.2 Receiver operating characteristic (ROC) curve. The curve is
generated by adjusting the cut-off values defining ‘normal’ and ‘abnormal’,
calculating the effect on sensitivity and specificity, and then plotting these
against each other. The closer the curve gets to the top left-hand corner, the
more useful the test is. The dotted line represents a test with no discriminant
value.
Test with no value
Conditional Probability
Conditional probability is the probability that something is true
given that something else is true. Bayes’ Theorem (named after
English clergyman Thomas Bayes 1702–61) is a mathematical
way to describe this. It estimates the post-test probability using
information about pre-test probability and the sensitivity and
specificity of the test.
Figure 3.3 illustrates Bayes’ Theorem and more detailed explanations can be found in the further resources. ‘Bayesian reasoning’ is
the term sometimes used for clinical reasoning using probabilities.
Test results shift our thinking, but sometimes by not very
much. The probability that someone actually has a disease
depends on the clinical (pre-test) probability, a judgement based
on the patient’s background, history and examination findings,
and the sensitivity and specificity of the test. Imagine an elderly
woman has been brought to the emergency department after
falling and hurting her left hip. On examination, the left hip is
extremely painful to move and she cannot weight bear. Both
antero-posterior and lateral X-rays of the left hip are normal (see
Figure 3.4). Is there a fracture? Sox and colleagues (see further
resources) state a fundamental assertion, which they describe as a
profound and subtle principle of clinical medicine: the interpreta-
tion of new information depends on what you believed beforehand.
As a simple rule of thumb, in a high clinical probability patient, a
normal test result does not necessarily exclude the disease, but in
a low clinical probability patient, a normal test result does exclude
the disease. Let’s go back to our elderly woman who has fallen.
The sensitivity of plan X-rays of the hip performed in the
emergency department for suspected hip fracture is 95%. That
means 5% of fractures (or 1 in 20) are missed. In an elderly
woman, likely to have osteoporosis, whose left hip is extremely
painful to move and she cannot weight bear, a normal X-ray does
0.50
Negative
shift
0.25
0
0
Figure 3.3 How a test results shift our thinking using Bayes’ Theorem. The
sensitivity of a troponin test is 95% and the specificity is 80%. If we imagine
a patient with chest pain and our pre-test or prior probability is 50% (i.e.,
we are sitting on the fence) a positive or a negative result would significantly
shift our thinking about whether the patient is having a heart attack. But if
our prior probability was very low (e.g., 10%) a negative test result would
shift our thinking by very little and a positive test result would not by itself
be conclusive (dotted line). Bayes’ Theorem is a method for interpreting
evidence in the context of previous knowledge. It has wide applications and
constitutes a mathematical foundation for reasoning. In clinical practice,
doctors do not use algebra to work out pre- and post-test probabilities,
however an understanding of the principles of Bayesian reasoning is
important because the ability to accurately estimate probability is important
in clinical reasoning. Bayes’ Theorem:
PDisR
/
where P[Dis/R+] is the chance of having the disease given a positive test
result; and P is probability, Dis is disease, and R+is a positive test result.
Figure from BrushJE. Probability: Uncertainty Quantified. In: The Science of
the Art of Medicine, 2015. Reproduced with permission of Dementi
Milestone Publishing.
Figure 3.4 Is there a fracture?
PR Dis PDis PR noDis Pno
0.25 0.50 0.75
P R Dis P Dis
/
DDis
–test

Choosing and Interpreting Diagnostic Tests 21
https://t.me/medicina_free
not necessarily exclude a fracture. But if the examination of the
hip was normal and she could walk easily, a normal X-ray would
be enough to satisfy the clinician that there is probably no fracture. The same test result is interpreted completely differently
when the clinical (pre-test) probability changes.
The example above illustrates that when the clinical probability
and the test result are discordant, we may need to think more care-
fully. For example, CT pulmonary angiography (CTPA) in the
diagnosis of pulmonary embolism (PE) has a specificity of 98%
and a sensitivity of 94%. When patients with an intermediate or
high clinical probability of PE have a positive CTPA, the result can
be trusted. Likewise, when patients with a low clinical probability
of PE have a negative CTPA, the result can also be trusted. But
what if a high clinical probability patient has a negative CTPA, or
a low clinical probability patient has a positive CTPA – what then?
One study found that around 40% of CTPA results were false in
these situations [3]. This is why further imaging (e.g., V/Q
SPECT
may be indicated in high clinical probability patients. It is also why
formal clinical probability assessment, D-dimer testing, and CTPA
which includes imaging of the lower limbs is used in combination
before safely withholding anticoagulation in patients being investigated for possible PE. There are many other examples in medicine where clinical probability really matters in accurately, and
safely, interpreting a diagnostic test result.
The lesson from these examples is that tests, even good tests,
can be wrong.
Tests give us test probabilities, not real probabilities. Tests have
to be interpreted in light of the clinical probability and estimating
clinical probability requires knowledge – formal and experiential
knowledge of basic science, epidemiology, clinical skills, and
clinical medicine.
Prevalence of Disease in a Population
Box 3.4 What is the chance a person found to have a positive
result actually has the disease?
Many doctors give an answer of 95%, but the actual answer is
illustrated in the table below:
Disease No disease Total
Actual 1 999 1000
Positive test 1 50 51
Negative test 0 949 949
If we sent 1000 tests to the lab, we would get 51 positive results – 1
true positive and 50 false positives. This chance of having a positive
result and actually having the disease is 1 out of 51 – or 2%. This
example illustrates the importance of understanding prevalence.
)
Box 3.5 Predictive values
Disease No disease
Positive test A
(True positive)
Negative test C
(False negative)
The positive predictive value – ‘What is the chance that a person
with a positive test truly has the disease?’ – is A/(A+B) × 100.
The negative predictive value – ‘What is the chance that a person
with a negative test does not have the disease?’ – is D/(D+C) × 100.
Positive and negative predictive values are influenced by the
prevalence of the disease in the population being tested. Using a
test in a population with higher prevalence increases positive
predictive value (and decreases negative predictive value).
B
(False positive)
D
(True negative)
Now let’s get more complicated! Consider this problem that was
given to a group of Harvard doctors: if a test to detect a disease
whose prevalence if 1:1000 has a false positive rate of 5%, what is
the chance that a person found to have a positive result actually
has the disease, assuming you know nothing about the person’s
symptoms or signs? (Assume no false negatives.) Just under half
replied with the answer 95%. Now look at Box 3.4 for the answer.
Sensitivity and specificity are characteristics relating to the
accuracy of a test relative to a reference standard. They are an
assessment of the test. But as a clinicians we are interested in the
question, ‘What are the chances that a person with a positive
result actually has the disease?’ In other words, we want to assess
people. Predictive values do just that – by combining sensitivity,
specificity, and prevalence of the disease in a population to answer
this question (see Box 3.5). Just considering test accuracy can be
misleading when the number of ‘positives’ and ‘negatives’ in different groups varies greatly.
In predictive analytics, a confusion matrix (yes, it’s real name)
is a 2×2 table that reports the number of true positives, false positives, true negatives, and false negatives using information about
the prevalence of disease in the population. This allows more
detailed analysis than simply observing the proportion of correct
classifications (or test accuracy).
John Brush, in his book The Science of the Art of Medicine (see
further resources) uses this next example to illustrate. We know
from angiography results and post-mortem studies the actual
prevalence of coronary artery disease in different patient groups.
Young women with non-cardiac sounding chest pain have a low
prevalence of ischaemic heart disease (1%). On the other hand,
older men with typical symptoms of angina have a high prevalence ischaemic heart disease (94%). If we sent a patient from
each of these groups for an imaging stress test, which has a sensitivity of 90% and a specificity of 85%, and both tests came back
positive, how would we interpret the results? In other words, what
is the positive predictive value of the test in these two different
scenarios? Aside from the fact that we should consider whether to
request this test at all in patients with such extreme pre-test probabilities, Box 3.6 shows the results we would get if we tested 100
patients just like each of them.
This example demonstrates the flaws in believing that a positive
result on a highly sensitive test indicates the presence of a condition
and that a negative result on a highly specific test indicates the
absence of a condition. Prevalence matters. In deciding the clinical
(pre-test) probability of disease, novices tend to focus on the patient’s
history and physical examination findings. A more accurate way of

22 ABC of Clinical Reasoning
https://t.me/medicina_free
Box 3.6 Confusion matrix showing results of an imaging
stress test in a) a 35-year-old woman with non-cardiac
sounding chest pain and b) a 65-year-old man with typical
symptoms of angina
a)
IHD No IHD
Actual/total 1 99
Positive test 0.9
True positive (sensitivity,
or 90% of 1)
Negative test 0.1 84.1
Positive predictive value=0.9 / (0.9+14.9) × 100=5.7%
b)
IHD No IHD
Actual/total 94 6
Positive test 84.6
True positive (sensitivity,
or 90% of 94)
Negative test 9.4 5.1
Positive predictive value=84.6 / (84.6+0.9) × 100=99%
An imaging stress test has a sensitivity of 90% and a specificity of
85%. Although both patients had some kind of chest pain and both
were sent for the same test, how we interpret a positive result is
completely different for each one because the prevalence of disease
in the group to which the patient belongs is so different (see Box
3.5 for predictive values).
14.9
True negative (specificity,
or 85% of 99)
0.9
True negative
(specificity, or 85% of 6)
estimating pre-test probability is to first ask yourself, ‘Who is my
patient?’ – in other words, the prevalence of disease in the group to
which the patient belongs – then add in information from the history and physical examination findings to come up with an estimate
of pre-test probability: low, intermediate, or high. Then use this
estimate to choose and interpret diagnostic tests. See Box 3.7 for an
example that illustrates this.
Thresholds
An important consideration in the diagnostic process is whether
to do a test at all. If a test will make no difference to the probability
or outcome of a disease, should the test be done? Tests (when they
are selected rationally, that is) are most helpful when they change
the management of a patient’s condition.
It is also not necessary to know the true state of the patient
before deciding whether to act. The therapeutic threshold combines factors such as test characteristics, risks of the test, the risks
and benefits of treatment, as well as the potential penalty for
being wrong. The point at which the factors are all evenly weighed
is the threshold. If a test or treatment for a disease is effective and
Box 3.7 Estimating clinical (pre-test) probability
A 30-year-old woman complained of a constant, dull left-sided
headache. On examination she was tender over her left temple. A
junior doctor remembered learning about temporal arteritis and
requested an erythrocyte sedimentation rate (ESR), a test for
temporal arteritis. The result was abnormal. The junior doctor
diagnosed temporal arteritis and started steroids.
The problem with this story is that temporal arteritis almost
exclusively affects people aged 50 years or more. So even with this
history, the pre-test probability of temporal arteritis is close to zero
in this patient, which affects the predictive value of the test, and
thus the interpretation of the result.
low risk then one would have a lower threshold for going ahead.
On the other hand, if a test or treatment is less effective or high
risk, one requires greater confidence in the diagnosis and potential benefits of treatment first.
Summary
Tests do not make a diagnosis, clinicians do. Tests give us test
probabilities not real probabilities. A working knowledge of factors
other than disease that influence test results, operating characteristics, and how accurate the test is for the disease in question is
important. Assessing clinical (pre-test) probability is vital, without
this you cannot interpret any test result. Pre-test probability is
derived from knowledge of the prevalence of the disease in the
group to which the patient belongs and information from the individual’s history and physical examination findings. Positive predictive values and negative predictive values are the proportion of
people with a positive (or negative) test result who have (or do not
have) a disease. They can be thought of as the post-test probability
of a disease. Finally, thresholds provide a useful way of thinking
about whether a test should be performed at all.
References
1. Whiting PF, Davenport C, Jameson C et al. (2015). How well do health professionals interpret diagnostic information? A systematic review. BMJ
Open; 5: e008155 (accessed April 2022).
2. Walsh B, Macfarlane PW, Prutkin JM and Smith SW. (2019). Distinctive
ECG patterns in healthy black adults. Journal of Electrocardiology; 56:
15–23.
3. Stein PD, Fowler SE, Goodman LR et al. (2006). Multidetector computed
tomography for acute pulmonary embolism. The New England Journal of
Medicine; 354: 2317–2327.
Further Resources
1. Sox HC, Higgins MC and Owens DK. Medical decision making, 2nd Ed.
Oxford: Wiley-Blackwell, 2013.
2. Brush JE. The Science of the Art of Medicine. Dementi Milestone Publishing,
2015.
3. Stone JV. Bayes’ Rule. A tutorial introduction to Bayesian analysis. Sebtel
Press, 2013.
Соседние файлы в папке @xirurgi_2025
