Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / @xirurgi_2025 / @xirurgi_2025 - 1270 - файл

.pdf
Скачиваний:
0
Добавлен:
29.08.2026
Размер:
3 Мб
Скачать
Evidence-based History and Examination 13
(1) Detecting pneumonia: In patients with acute respiratory
(2) chest pain, “dysphagia” is reported in 4% of patients found to have coronary disease and in 20% of patients with another cause of chest pain. Therefore,
Infinity
Zero
No change
e
https://t.me/medicina_free
Likelihood Ratios
When a symptom or sign is present (or absent), how do we know how useful that finding is in making a diagnosis? Likelihood ra­tios (LRs) are diagnostic weights. The likelihood ratio is the prob­ability of the finding in someone with the disease over the probability of the finding in someone without the disease. Thus, if a finding is equally likely in people with and without the disease, the likelihood ratio is 1 (i.e., unhelpful). Each finding from the history and physical examination is associated with a unique LR, a number whose values ranges from zero to infinity. An LR greater than 1.0 increases the probability of disease, and the higher the value of the LR, the greater the increase in probability. An LR of less than 1.0 decreases the probability of disease, and the lower the value of the LR, the greater the reduction in probability (see Figures 2.2 and 2.3).
One simple method of interpreting LRs is to memorise the association between three LR values – 2, 5, and 10 – and the first three multiples of 15 – 15, 30, and 45. A finding with an LR of 2 increases the absolute probability by around 15% (that is, the cli­nician adds 15% to the pre-test probability); a finding with an LR of 5 increases the probability by around 30%, and one with an LR of 10 increases the probability by around 45%.
For those LRs less than 1.0, the clinician simply inverts the 2, 5, and 10 ‘rule’ (that is 0.5, 0.2, and 0.1). A finding with an LR of 0.5 decreases the probability by around 15%; one with an LR of 0.2 decreases the probability by around 30%, and one with an LR of
0.1 decreases the probability by around 45%. Provided clinicians round off final probabilities greater than 100% to 100%, and those less than 0% to 0%, this method suffices for the purposes of clinical reasoning.
Box 2.11 summarises the absolute changes in probability for the most used LRs. Findings with LRs greater than 3 or less than
0.3 are most helpful because these values identify findings that either increase or decrease probability by 20–25% or more.
Can LRs Be Combined?
LRs can be combined only if the two findings are independent of one another (independence implies the LR for the first finding is
complaints, “percussion dullness” is found in 18% of patients with pneumonia and in 6% of patients with another cause of respiratory distress. Therefore,
for percussion dullness
LR
in detecting pneumonia
Detecting coronary artery disease: In patients with chronic
18
== 3.0
6
10
Increase
probability
Decrease
probability
Figure 2.3 Likelihood ratios: diagnostic weights. Clinicians should classify LRs into three groups: those with values greater than 1.0 increase proba­bility; those with values less than 1.0 decrease probability; and those with values near 1.0 change probability very little or not at all.
Box 2.11 Likelihood ratios and bedside estimates
Likelihood ratio Approximate change in probability*
0.1 −45%
0.2 −30%
0.3 −25%
0.5 −15%
1 No change
2 +15%
3 +20%
4 +25%
5 +30%
6 +35%
7
8 +40%
9
10 +45%
*These changes describe absolute increases or decreases in probability.
From McGee (2002). J Gen Intern Med; 17: 646–9.
5
2
1
0.5
0.2
0.1
+45%
+30%
+15%
No chang
–15%
–30%
–45%
for dysphagia
LR
in detecting coronary
artery disease
Figure 2.2 Likelihood ratios: examples. From McGee (further resources).
4
==
20
0.2
the same whether or not the second finding is present). For example, typical angina (an LR of 5.8) and hyperlipidaemia (an LR of 2.2) are likely to be independent because the accuracy of a history of typical angina is unlikely to be affected by the presence or
14 ABC of Clinical Reasoning
https://t.me/medicina_free
absence of hyperlipidaemia. To combine findings, the clinician can simply multiply the two individual LRs (5.8 × 2.2); the resulting product (12.7 or a+50% probability) becomes the LR for combined ‘typical angina and hyperlipidaemia’. Alternatively, the clinician could first apply typical angina (LR of 5.8 or a+35% probability), then hyperlipidaemia (LR of 2.2 or a+15% probability) to obtain the increment in probability for the combined findings (35%+15% or a+50% probability).
Clinicians should not combine the LRs of more than two individual findings unless clinical studies have proven that the findings are independent. If there is any possibility that the individual findings are dependent on each other, their LRs should not be combined (for example, typical angina and ‘duration of pain< 5 minutes’ should not be combined, because pain lasting less than 10 minutes after rest or nitro-glycerine is a criterion for stable typical angina).
The Limitations of LRs
Statistical calculations are appropriate only when the clinical problem is defined by a diagnostic (or reference) standard, such as laboratory testing or clinical imaging (Figure 2.4). Examples, and their reference standards, are pneumonia (chest radiographs), ascites (ultrasonog­raphy), coronary artery disease (coronary angiography), anaemia (full blood count), and hyperthyroidism (thyroid function tests). In each of these disorders, the evidence-based approach compares find­ings from the history or examination to the accepted reference stan­dard and identifies the findings most accurately predicting the results of that standard. Since many clinical problems lack reference stan­dards, evidence-based reasoning using LRs is not always applicable. For these problems, empiric observation based on the clinician’s prior knowledge and experience of similar patients – what the clini­cian sees, feels, and hears at the bedside – remains the sole diagnostic standard and LRs cannot be used.
Although LRs describe how the probability changes, they cannot determine the pre-test probability of a disease. For example, the LR for the physical finding ‘fluid wave’ in detecting ascites in patients with abdominal distension is 5.0 (a+30% probability). If the clinician works in a hepatology practice in which 60% of all patients with abdominal distension have ascites (that is a pre-test probability of 60%) the finding of a fluid wave is diagnostic (that is 60%+30% or
WHAT IS THE
DIAGNOSTIC STANDARD?
a 90% probability of ascites). On the other hand, if the clinician works in a community practice where only 20% of patients with abdominal distension have ascites (the other 80% have increased abdominal fat or gas), the presence of the fluid wave is less conclu­sive (20%+30% or a 50% probability of ascites). Proper application of evidence-based medicine here requires intimate knowledge of the types of diseases found in one’s own practice.
The Future of the History and Physical Examination
Increasingly, researchers are comparing clinical findings to diag­nostic standards to reveal LRs for a wide variety of clinical disor­ders. This is through diagnostic accuracy studies reported to the STARD criteria [13]. These include:
Both the test (clinical symptom, sign, or laboratory test) and diagnostic standard are clearly defined
All enrolled patients have symptoms suggestive of the diagnosis under study
Determination of the test result is blinded from determination of the diagnostic standard
The study presents enough information to allow calculation of
LRs and their confidence intervals. Clinicians applying this approach can focus on findings with greatest diagnostic accuracy. Nonetheless, this does have limita­tions. Even when a problem has been studied, conclusions often rest on relatively few patients. Whether diagnostic accuracy depends on clinical technique is largely unaddressed, although the few studies on this subject show diagnostic accuracy with stu­dents as observers is the same as with specialists, provided the finding is well-defined. Finally, most literature on the subject focusses on individual findings, although it is well known that expert clinicians typically combine many findings simultaneously when diagnosing disease.
Point of care ultrasound is increasingly being used in acute care settings as an extension of the physical examination (e.g., to estimate volume status, or differentiate fluid from consolidation in the lungs). However, the same caveats for all diagnostic tests apply (see Chapter 3) – the history and physical examination remains fundamental in establishing the clinical probability of disease and ultrasound ‘findings’ need to be interpreted in light of this. Point of care ultrasound has several limitations and should be seen as a decision aid pending more definitive investigations.
Clinical imaging or laboratory
Pneumonia Ascites Coronary artery disease Anemia Hyperthyroidism
Evidence-based reasoning
can be used
Figure 2.4 Can evidence-based reasoning be used?
Empiric observation
Cellulitis Parkinson disease Trochanteric bursitis Pericarditis Serotonin syndrome
Evidence-based reasoning
does
not apply
Developing Skills in Teaching
It is challenging for busy clinicians to be experts in clinical com­munication and in teaching evidence-based history and physical examination. This has contributed to a decline in bedside teaching since the 1960s. It is however both a patient and student-centred activity. The scope for evidence-based history and examination is exciting, with potential to improve patient safety. Role-modelling of reflective practice by bedside teachers can assist learners in developing resilience and dealing with the uncertainty of clinical practice. Careful planning and engagement with patients can help develop clinical teachers. Box 2.12 lists some tips for teaching
Evidence-based History and Examination 15
https://t.me/medicina_free
Box 2.12 Tips for teaching evidence-based physical examination
Practice teaching concepts of diagnostic accuracy
• Practice estimating the pre-test probability of disease
• Practice teaching methods to estimate post-test probabilities
• Know where to find evidence-based physical examination data and prepare to use it
• Prepare an answer to the question, ‘Can the likelihood ratios of multiple findings be combined?’
• Answer the common question, ‘Why should we examine patients if it is so unhelpful?’
• Teach the basics of evidence-based physical examination and prepare students for bedside teaching
• Orientate the patient to the purpose of the teaching and explicitly discuss evidence-based physical examination
• Encourage students to commit to their own description of findings
• Encourage students to commit to a next step in management
Facilitate deliberate practice and give feedback to learn evidence-
based physical examination
• Acknowledge uncertainty and follow up on unresolved issues
Adapted from Mookherjee S, Hunt S, Chou CL. (2015). Twelve tips for teaching evidence-based physical examination. Medical Teacher; 37(6): 543–550.
evidence-based physical examination. Seeing variations in dem­onstration of the physical examination is a source of discomfort for students, particularly around assessment. Reasons for varia­tion in technique should be discussed with learners to help them manage their uncertainty and to apply these critical skills.
Summary
Practicing evidence-based history and examination is challenging but rewarding. Since the history and examination is so critical to the patient’s care, a robust evidence base is essential, and merits increased research. An initial step for learners is establishing the importance of the history and examination not only for initial formulation of the patient’s problem list and differential diag­nosis, but the correct interpretation of any subsequent investiga­tions. Clinical teachers should be supported in developing their own confidence and skills in teaching evidence-based history and examination.
tion, including this revision.Thanks also to Lucille Middleton, Graduate Entry Medicine student at the University of Nottingham, UK, for contributing Box 2.1 to this chapter.
References
1. Hampton JR, Harrison MJ, Mitchell JR et al. (1975). Relative contribu­tions of history-taking, physical examination, and laboratory investiga­tion to diagnosis and management of medical outpatients. British Medical Journal; 31; 2(5969): 486–489.
2.
Silverman J, Kurtz SM and Draper J. Skills for communicating with
patients, 3
Silverman J. The consultation. In: Cooper N and Frain J (Eds). ABC of
3. Clinical Communication. Wiley-Blackwell, 2018.
4. Frain J and Abdalla M. Teaching clinical communication. In: Cooper N and Frain J (Eds). ABC of Clinical Communication. Wiley-Blackwell,
2018.
5.
Kilian A, Upton LA and Sheagren JN. (2020). Reorganizing the history of
present illness to improve verbal case presenting and clinical diagnostic reasoning skills of medical students: the all-inclusive history of present illness. Journal of Medical Education and Curricular Development; 7:
2382120520928996.
Swap CJ and Nagurney JT. (2005). Value and limitations of chest pain
6. history in the evaluation of patients with suspected acute coronary syn­dromes. JAMA; 294(20): 2623–2629.
7. Elieson SW and Papa FJ. (1994). The effect of various knowledge formats on diagnostic performance. Academic Medicine; 69(10 Suppl): S81–S83.
8. Thomas KE, Hasbun R, Jekel J and Quagliarello VJ. (2002). The diagnostic accuracy of Kernig’s sign, Brudzinski’s sign, and nuchal rigidity in adults with suspected meningitis. Clinical Infectious Diseases; 35(1): 46–52.
Yusuf S, Hawken S, Ounpuu S et al. (2004). Effect of potentially modifi-
9. able risk factors associated with myocardial infarction in 52 countries (the INTERHEART study): case-control study. Lancet; 364: 937–952.
10. Paley L, Zornitzki T, Cohen J et al. (2011). Utility of clinical examination in the diagnosis of emergency department patients admitted to the department of medicine of an academic hospital. Archives of Internal Medicine; 171(15): 1393–1400.
11. Verghese A, Charlton B, Kassirer J et al. (2015). Inadequacies of physical examination as a cause of medical errors and adverse events: a collection of vignettes. The American Journal of Medicine; 128(12): 1322–1324.
12. Holboe ES. (2004). Faculty and the observation of trainees’ clinical skills: problems and opportunities. Academic Medicine; 79: 16–22.
Cohen JF, Korevaar DA, Altman DG et al. (2015). STARD guidelines for
13. reporting diagnostic accuracy studies: explanation and elaboration. BMJ Open; 6: e012799. doi:10.1136/bmjopen-2016-012799.
rd
Ed. CRC Press, 2013.
Acknowledgements
Thanks are due to my co-author of the first edition of this chapter, Steven McGee, Emeritus Professor of Medicine, University of Washington, Seattle, USA, whose work and contribution to evi­dence-based history and examination continues to be an inspira-
Further Resources
1. McGee S. Evidence-based physical diagnosis, 5th Ed. Elsevier/Saunders,
2021.
Talley N and O’Connor S. Clinical examination, 9
2.
th
Ed. Elsevier, 2021.
https://t.me/medicina_free
CHAPTER 3
https://t.me/medicina_free
Choosing and Interpreting Diagnostic Tests
Nicola Cooper
OVERVIEW
• Test results are affected by a number of factors which the clinician has to take into account
There is no such thing as a perfect test
• The interpretation of new information depends on what you believed beforehand, based on your assessment of the patient
• Predictive values combine information about sensitivity, specificity, and prevalence and indicate how likely a test result is to be correct
• Thresholds provide a useful way of thinking about whether a test should be performed at all
Introduction
The history and physical examination provide a differential diag­nosis and/or problem list. This is refined further using diagnostic tests. The appropriate selection of tests depends of the quality of the history and physical examination. Test results then have to be interpreted in light of the patient’s history and examination find­ings because test results are affected by a number of factors (see Box 1.1):
How ‘normal’ is defined
Factors other than disease that influence test results
Operating characteristics
Sensitivity and specificity
Prevalence of disease in a population Unfortunately, commonly used measures of test accuracy, such as sensitivity and specificity, are poorly understood. A systematic review of 24 studies found that most qualified healthcare profes­sionals were poor at providing definitions of sensitivity and spec­ificity, and were poor at estimating the post-test probability of disease [1]. This chapter aims to introduce key concepts and pro­vide further resources for this important area of clinical reasoning.
Box 1.1 Tests are affected by a number of factors
Factor Explanation
How ‘normal’ is defined
Factors other than disease that influence test results
Operating characteristics
Sensitivity and specificity (see Box 3.3)
Prevalence of disease in a population (see Box 3.5)
‘Normal’ can refer to values within the
• reference range for the population to which the patient belongs
It can also refer to a value below or above a
pre-determined cut-off point designed to maximise true positives and minimise false positives
It can also be an ‘abnormal’ result that is
• actually normal for the particular context in question
These are biological and/or laboratory factors
• that make test results ‘abnormal’ when they are not, or vary when there has not been a true change
• This refers to the method of performing the test itself which, if not optimal, can affect its accuracy
The sensitivity of a test refers to its ability to
correctly identify patients with the disease
• The specificity of a test refers to its ability to correctly identify patients without the disease
Sensitivity and specificity are characteristics
relating to the accuracy of a test relative to a reference standard
The prevalence of disease in a population
• can significantly alter the predictive value of a test
The positive predictive value is the
• proportion of people with a positive test result who truly have the disease
The negative predictive is the proportion of
people with a negative test result who do not have the disease
ABC of Clinical Reasoning, Second Edition. Edited by Nicola Cooper and John Frain. © 2023 John Wiley & Sons Ltd. Published 2023 by John Wiley & Sons Ltd.
18 ABC of Clinical Reasoning
(standard deviations from the mean)
https://t.me/medicina_free
How Normal Is Defined
Many diagnostic test results are expressed as continuous variables on a numerical scale and many quantitative measurements in human populations have a Gaussian (normal) distribution. The ‘normal’ range is defined as those values that encompass 95% of the healthy population, or two standard deviations from the mean. This means that 2.5% of the healthy population will have values above, and 2.5% of the population will have values below, the normal range. For this reason, it is more appropriate to use the term ‘reference range’ (see Figure 3.1). Diagnostic test results in people with a disease also have a Gaussian distribution but with a different mean and reference range. In some diseases there is no overlap between results from the abnormal and normal population, but in some diseases there is. In the latter, the greater the difference between the result and the reference range of the normal population, the higher the chance that the person has the disease.
Arbitrarily dividing a range of values into ‘normal’ and ‘abnormal’ has disadvantages – it does not take into account the magnitude of the result. For example, a highly sensitive troponin T result in a patient with chest pain is more likely to indicate myo­cardial injury when the value is very high, as opposed to slightly raised. Some test results have a binary classification (‘normal’ vs ‘abnormal’), for example, an exercise electrocardiogram (ECG) looking for signs of ischaemic heart disease. However, in deciding where the cut-off point between ‘normal’ and ‘abnormal’ should be, there is a trade-off between sensitivity (true positives) and specificity (true negatives). The optimum cut-off point is calcu­lated using receiver operating characteristic (ROC) analysis, which is described in more detail later.
In medicine there are some situations when a normal result is abnormal, and an abnormal result is normal. For example in a clinically severe asthma attack when one expects the PaCO low, a normal PaCO
on an arterial blood gas is not normal at all
2
and indicates life-threatening asthma. On the other hand, a raised d-dimer is normal in pregnancy. So what is ‘normal’ and ‘abnormal’ has to be interpreted in light of the clinical picture. Clinicians who use diagnostic tests should have a good working knowledge of the tests they use in their everyday practice, and how they should be interpreted in light of the patient’s history and examination findings.
Normal population
Number of people
Reference range
–3 –2 –1 1 2 3
Figure 3.1 Normal distribution.
Mean
Test Result
to be
2
Factors Other than Disease Which Influence Test Results
There are a number of factors other than disease which influence test results. They include:
Age
Sex
Ethnicity
Pregnancy
Body position
Chance
Spurious (in vitro) results
Lab error
Critical difference values For example, normal values for paediatric blood results can be sig­nificantly different to those of adults. Old people often have a normal white cell count in the presence of infection, and can have a significantly reduced glomerular filtration rate with a normal creat­inine. Men have slightly different reference ranges to women (e.g., for haemoglobin) and healthy black adults may have an ‘abnormal’ 12-lead electrocardiogram (due to early repolarisation) that can resemble serious disease, but is in fact a ‘normal variant’ [2].
Pregnancy significantly alters many test results due to the physiological changes that occur, particularly in the third tri­mester. A large foetus splints the diaphragm and compresses the lungs causing supine hypoxaemia as well as a respiratory alkalosis (important facts to remember when considering the possibility of pulmonary embolism in a pregnant woman). Circulating volume increases by 50% in late pregnancy causing a flow murmur, tachy­cardia, and a rightward axis on the 12-lead electrocardiogram. Kidneys also swell as a result, and renal ultrasound shows increased size and dilatation.
Body position is important in some tests, for example, lung function and tests where the patient has to lie in a certain position to get optimal images. Finally, a test result may be abnormal by chance (e.g., the patient is an outlier on the normal curve); the result may be spurious (e.g., hyperkalaemia caused by haemolysis or some haematological conditions); or may be due to lab error (e.g., as a result of a technical or human error). It is always worth pausing before acting when a very unexpected test result crops up.
Lab results also vary in the same person at different times. The criti­cal difference, also known as the reference change value, is the smallest difference between sequential laboratory results in the same patient which is likely to indicate a true change. Let’s imagine a person has their cholesterol measured every single day. The result will not be identical every time. The reason for this is natural biological var­iation but also lab variation. The combination of the two is the critical difference – the amount by which the test can vary before it can be considered a true change. This is calculated using knowledge of normal intra-individual variation and lab variation for different tests. The critical difference is different for different lab tests. Some calculated critical difference values for common biochemistry results are shown in Box 3.2. For a person having their serum cholesterol monitored, the critical difference is 17%. An initial value of 5.2 mmol/L can therefore vary between
4.3mmol/L and 6.1 mmol/L without being a true change.
Choosing and Interpreting Diagnostic Tests 19
https://t.me/medicina_free
Box 3.2 Calculated critical difference (CD) for some common biochemistry results
Test CD as %
Albumin
Alkaline phosphatase
Aspartate aminotransferase (AST)
Bilirubin
Calcium
Cholesterol
Glucose
Total protein
TSH
Urea
Uric acid
Data from Professor Trefor Higgins, Department of Laboratory Medicine and Pathology, University of Alberta.
11.2
37.1
27.7
47.5
6.1
17.0
9.9
11.2
63.0
28.9
25.2
Operating Characteristics
Before ordering a test, it is important to be aware of certain operating characteristics of the test. This refers to the method of actually performing the test itself. For example, measuring lung function requires that the patient be able to hear, understand, and co-operate with instructions, as well as hold their breath. Exercise electrocardiograms require patients to be able to walk briskly and cannot be accurately interpreted in people who have left bundle branch block.
Some tests are highly operator dependent – in other words, the skill of the operator influences the results and the report provided. Ultrasound is the best example of this, as dynamic images have to be skilfully interpreted by the sonographer. For radiology investiga­tions in general, the interpretation of results can be highly influ­enced by the patient’s body habitus or clinical state. In ultrasound, for example, morbid obesity can make getting good views difficult, and in people of all sizes, intra-abdominal organs can be obscured by bowel gas. In computed tomography, images can be severely degraded by movement artefact, or interpretation can be affected by whether or not contrast was used, and whether it circulated as antic­ipated to get optimal images. If a report says, ‘Limited views due to … but within these limitations, no abnormality detected’ consider whether it is in fact a non-diagnostic scan, rather than a ‘normal’ scan. It is also important that radiologists, as well as other clinicians such as physiologists, are provided with a clear clinical question and key information in the history, past medical/surgical history, and physical examination. This is so that ‘abnormalities’ or incidental findings can be interpreted in light of the clinical context.
Sensitivity and Specificity
The sensitivity of a test refers to its ability to correctly identify patients with the disease. The specificity of a test refers to its ability to correctly identify patients without the disease. Even a test with
Box 3.3 Sensitivity and specificity
Disease No disease
Positive test A
(True positive)
Negative test C
(False negative)
The sensitivity of a test refers to its ability to correctly identify patients with the disease, i.e. A/(A+C) × 100.
The specificity of a test refers to its ability to correctly identify
patients without the disease, i.e. D/(D+B) × 100.
B (False positive)
D (True negative)
a high sensitivity, for example 95%, will miss 5% of people with the disease. Unfortunately, there is no such thing as a perfect test. Test results consist of ‘true positives’ and ‘false positives’; ‘true negatives’ and ‘false negatives’. Box 3.3 illustrates this. Tests differ in their sensitivity and specificity for detecting certain diseases, so clinicians need to have a sound working knowledge of the accu­racy of the tests they use on a day-to-day basis.
A very sensitive test will detect most disease but generate abnormal findings in healthy people. We see this with the aptly named high sensitivity troponin T. On the other hand, a very specific test may miss significant disease but is likely to establish the diagnosis beyond doubt when the result is positive. You may have heard of the acronyms ‘SNOUT’ and ‘SPIN’. SNOUT stands for ‘sensitive test when negative rules out the disease’ and SPIN stands for ‘specific test when positive rules in the disease’. However, SNOUT and SPIN are misleading. This is because the diagnostic power of any test is determined by both its sensitivity and specificity, as well as the prevalence of disease in the population – more of that later. The trade-off between sensitivity and specificity is explored in what is termed a ‘ROC analysis’.
ROC Analysis
ROC stands for ‘receiver operating characteristic’ – so called because it was developed by radar engineers during World War II for discriminating enemy objects in the battlefield. It is also known as the ‘relative operating characteristic’ because it compares two operating characteristics (true positive results and false positive results) at various settings (see Figure 3.2). It is used in medicine to select the best cut-off point for a test in a way that maximises true positives while minimising false positives. ROC analysis is con­ducted in a research setting whenever investigators measure the ability of a test to detect a diagnosis in a population with the dis­ease and exclude the diagnosis in those without it. (Of course, the results of the analysis also depend on what study population was used – if the same performance is expected in practice, the test must be used in a similar population). For example, if we define an exercise electrocardiogram as ‘abnormal’ when there is at least 0.5 mm of ST depression, we could pick up every case of ischaemic heart disease but generate many false positives. On the other hand, if we define an exercise electrocardiogram as ‘abnormal’ when there is at least 2 mm of ST depression, we could detect most cases of clinically important ischaemic heart disease but with far fewer false positives, which is far more practical.
20 ABC of Clinical Reasoning
1.0
1.0
False positive rate
True positive rate
Prior Probability
Posterior Probability
1.00
1.00
//



https://t.me/medicina_free
Perfect test
Good test
Moderate test
0.75
+test
Positive
shift
0.5
0 0.5
Figure 3.2 Receiver operating characteristic (ROC) curve. The curve is generated by adjusting the cut-off values defining ‘normal’ and ‘abnormal’, calculating the effect on sensitivity and specificity, and then plotting these against each other. The closer the curve gets to the top left-hand corner, the more useful the test is. The dotted line represents a test with no discriminant value.
Test with no value
Conditional Probability
Conditional probability is the probability that something is true given that something else is true. Bayes’ Theorem (named after English clergyman Thomas Bayes 1702–61) is a mathematical way to describe this. It estimates the post-test probability using information about pre-test probability and the sensitivity and specificity of the test.
Figure 3.3 illustrates Bayes’ Theorem and more detailed explana­tions can be found in the further resources. ‘Bayesian reasoning’ is the term sometimes used for clinical reasoning using probabilities.
Test results shift our thinking, but sometimes by not very much. The probability that someone actually has a disease depends on the clinical (pre-test) probability, a judgement based on the patient’s background, history and examination findings, and the sensitivity and specificity of the test. Imagine an elderly woman has been brought to the emergency department after falling and hurting her left hip. On examination, the left hip is extremely painful to move and she cannot weight bear. Both antero-posterior and lateral X-rays of the left hip are normal (see Figure 3.4). Is there a fracture? Sox and colleagues (see further resources) state a fundamental assertion, which they describe as a profound and subtle principle of clinical medicine: the interpreta- tion of new information depends on what you believed beforehand. As a simple rule of thumb, in a high clinical probability patient, a normal test result does not necessarily exclude the disease, but in a low clinical probability patient, a normal test result does exclude the disease. Let’s go back to our elderly woman who has fallen. The sensitivity of plan X-rays of the hip performed in the emergency department for suspected hip fracture is 95%. That means 5% of fractures (or 1 in 20) are missed. In an elderly woman, likely to have osteoporosis, whose left hip is extremely painful to move and she cannot weight bear, a normal X-ray does
0.50
Negative
shift
0.25
0
0
Figure 3.3 How a test results shift our thinking using Bayes’ Theorem. The sensitivity of a troponin test is 95% and the specificity is 80%. If we imagine a patient with chest pain and our pre-test or prior probability is 50% (i.e., we are sitting on the fence) a positive or a negative result would significantly shift our thinking about whether the patient is having a heart attack. But if our prior probability was very low (e.g., 10%) a negative test result would shift our thinking by very little and a positive test result would not by itself be conclusive (dotted line). Bayes’ Theorem is a method for interpreting evidence in the context of previous knowledge. It has wide applications and constitutes a mathematical foundation for reasoning. In clinical practice, doctors do not use algebra to work out pre- and post-test probabilities, however an understanding of the principles of Bayesian reasoning is important because the ability to accurately estimate probability is important in clinical reasoning. Bayes’ Theorem:
PDisR
/


where P[Dis/R+] is the chance of having the disease given a positive test result; and P is probability, Dis is disease, and R+is a positive test result. Figure from BrushJE. Probability: Uncertainty Quantified. In: The Science of the Art of Medicine, 2015. Reproduced with permission of Dementi Milestone Publishing.
Figure 3.4 Is there a fracture?

PR Dis PDis PR noDis Pno

0.25 0.50 0.75
P R Dis P Dis

/
DDis
–test
Choosing and Interpreting Diagnostic Tests 21
https://t.me/medicina_free
not necessarily exclude a fracture. But if the examination of the hip was normal and she could walk easily, a normal X-ray would be enough to satisfy the clinician that there is probably no frac­ture. The same test result is interpreted completely differently when the clinical (pre-test) probability changes.
The example above illustrates that when the clinical probability and the test result are discordant, we may need to think more care- fully. For example, CT pulmonary angiography (CTPA) in the diagnosis of pulmonary embolism (PE) has a specificity of 98% and a sensitivity of 94%. When patients with an intermediate or high clinical probability of PE have a positive CTPA, the result can be trusted. Likewise, when patients with a low clinical probability of PE have a negative CTPA, the result can also be trusted. But what if a high clinical probability patient has a negative CTPA, or a low clinical probability patient has a positive CTPA – what then? One study found that around 40% of CTPA results were false in these situations [3]. This is why further imaging (e.g., V/Q
SPECT
may be indicated in high clinical probability patients. It is also why formal clinical probability assessment, D-dimer testing, and CTPA which includes imaging of the lower limbs is used in combination before safely withholding anticoagulation in patients being inves­tigated for possible PE. There are many other examples in medi­cine where clinical probability really matters in accurately, and safely, interpreting a diagnostic test result.
The lesson from these examples is that tests, even good tests, can be wrong.
Tests give us test probabilities, not real probabilities. Tests have to be interpreted in light of the clinical probability and estimating clinical probability requires knowledge – formal and experiential knowledge of basic science, epidemiology, clinical skills, and clinical medicine.
Prevalence of Disease in a Population
Box 3.4 What is the chance a person found to have a positive result actually has the disease?
Many doctors give an answer of 95%, but the actual answer is illustrated in the table below:
Disease No disease Total
Actual 1 999 1000
Positive test 1 50 51
Negative test 0 949 949
If we sent 1000 tests to the lab, we would get 51 positive results – 1 true positive and 50 false positives. This chance of having a positive result and actually having the disease is 1 out of 51 – or 2%. This example illustrates the importance of understanding prevalence.
)
Box 3.5 Predictive values
Disease No disease
Positive test A
(True positive)
Negative test C
(False negative)
The positive predictive value – ‘What is the chance that a person with a positive test truly has the disease?’ – is A/(A+B) × 100.
The negative predictive value – ‘What is the chance that a person
with a negative test does not have the disease?’ – is D/(D+C) × 100.
Positive and negative predictive values are influenced by the prevalence of the disease in the population being tested. Using a test in a population with higher prevalence increases positive predictive value (and decreases negative predictive value).
B (False positive)
D (True negative)
Now let’s get more complicated! Consider this problem that was given to a group of Harvard doctors: if a test to detect a disease whose prevalence if 1:1000 has a false positive rate of 5%, what is the chance that a person found to have a positive result actually has the disease, assuming you know nothing about the person’s symptoms or signs? (Assume no false negatives.) Just under half replied with the answer 95%. Now look at Box 3.4 for the answer.
Sensitivity and specificity are characteristics relating to the accuracy of a test relative to a reference standard. They are an assessment of the test. But as a clinicians we are interested in the question, ‘What are the chances that a person with a positive result actually has the disease?’ In other words, we want to assess people. Predictive values do just that – by combining sensitivity, specificity, and prevalence of the disease in a population to answer this question (see Box 3.5). Just considering test accuracy can be misleading when the number of ‘positives’ and ‘negatives’ in dif­ferent groups varies greatly.
In predictive analytics, a confusion matrix (yes, it’s real name) is a 2×2 table that reports the number of true positives, false pos­itives, true negatives, and false negatives using information about the prevalence of disease in the population. This allows more detailed analysis than simply observing the proportion of correct classifications (or test accuracy).
John Brush, in his book The Science of the Art of Medicine (see further resources) uses this next example to illustrate. We know from angiography results and post-mortem studies the actual prevalence of coronary artery disease in different patient groups. Young women with non-cardiac sounding chest pain have a low prevalence of ischaemic heart disease (1%). On the other hand, older men with typical symptoms of angina have a high preva­lence ischaemic heart disease (94%). If we sent a patient from each of these groups for an imaging stress test, which has a sensi­tivity of 90% and a specificity of 85%, and both tests came back positive, how would we interpret the results? In other words, what is the positive predictive value of the test in these two different scenarios? Aside from the fact that we should consider whether to request this test at all in patients with such extreme pre-test prob­abilities, Box 3.6 shows the results we would get if we tested 100 patients just like each of them.
This example demonstrates the flaws in believing that a positive result on a highly sensitive test indicates the presence of a condition and that a negative result on a highly specific test indicates the absence of a condition. Prevalence matters. In deciding the clinical (pre-test) probability of disease, novices tend to focus on the patient’s history and physical examination findings. A more accurate way of
22 ABC of Clinical Reasoning
https://t.me/medicina_free
Box 3.6 Confusion matrix showing results of an imaging stress test in a) a 35-year-old woman with non-cardiac sounding chest pain and b) a 65-year-old man with typical symptoms of angina
a)
IHD No IHD
Actual/total 1 99
Positive test 0.9
True positive (sensitivity, or 90% of 1)
Negative test 0.1 84.1
Positive predictive value=0.9 / (0.9+14.9) × 100=5.7%
b)
IHD No IHD
Actual/total 94 6
Positive test 84.6
True positive (sensitivity, or 90% of 94)
Negative test 9.4 5.1
Positive predictive value=84.6 / (84.6+0.9) × 100=99%
An imaging stress test has a sensitivity of 90% and a specificity of 85%. Although both patients had some kind of chest pain and both were sent for the same test, how we interpret a positive result is completely different for each one because the prevalence of disease in the group to which the patient belongs is so different (see Box
3.5 for predictive values).
14.9
True negative (specificity, or 85% of 99)
0.9
True negative (specificity, or 85% of 6)
estimating pre-test probability is to first ask yourself, ‘Who is my patient?’ – in other words, the prevalence of disease in the group to which the patient belongs – then add in information from the his­tory and physical examination findings to come up with an estimate of pre-test probability: low, intermediate, or high. Then use this estimate to choose and interpret diagnostic tests. See Box 3.7 for an example that illustrates this.
Thresholds
An important consideration in the diagnostic process is whether to do a test at all. If a test will make no difference to the probability or outcome of a disease, should the test be done? Tests (when they are selected rationally, that is) are most helpful when they change the management of a patient’s condition.
It is also not necessary to know the true state of the patient before deciding whether to act. The therapeutic threshold com­bines factors such as test characteristics, risks of the test, the risks and benefits of treatment, as well as the potential penalty for being wrong. The point at which the factors are all evenly weighed is the threshold. If a test or treatment for a disease is effective and
Box 3.7 Estimating clinical (pre-test) probability
A 30-year-old woman complained of a constant, dull left-sided headache. On examination she was tender over her left temple. A junior doctor remembered learning about temporal arteritis and requested an erythrocyte sedimentation rate (ESR), a test for temporal arteritis. The result was abnormal. The junior doctor diagnosed temporal arteritis and started steroids.
The problem with this story is that temporal arteritis almost exclusively affects people aged 50 years or more. So even with this history, the pre-test probability of temporal arteritis is close to zero in this patient, which affects the predictive value of the test, and thus the interpretation of the result.
low risk then one would have a lower threshold for going ahead. On the other hand, if a test or treatment is less effective or high risk, one requires greater confidence in the diagnosis and poten­tial benefits of treatment first.
Summary
Tests do not make a diagnosis, clinicians do. Tests give us test probabilities not real probabilities. A working knowledge of factors other than disease that influence test results, operating character­istics, and how accurate the test is for the disease in question is important. Assessing clinical (pre-test) probability is vital, without this you cannot interpret any test result. Pre-test probability is derived from knowledge of the prevalence of the disease in the group to which the patient belongs and information from the indi­vidual’s history and physical examination findings. Positive pre­dictive values and negative predictive values are the proportion of people with a positive (or negative) test result who have (or do not have) a disease. They can be thought of as the post-test probability of a disease. Finally, thresholds provide a useful way of thinking about whether a test should be performed at all.
References
1. Whiting PF, Davenport C, Jameson C et al. (2015). How well do health pro­fessionals interpret diagnostic information? A systematic review. BMJ Open; 5: e008155 (accessed April 2022).
2. Walsh B, Macfarlane PW, Prutkin JM and Smith SW. (2019). Distinctive ECG patterns in healthy black adults. Journal of Electrocardiology; 56: 15–23.
3. Stein PD, Fowler SE, Goodman LR et al. (2006). Multidetector computed tomography for acute pulmonary embolism. The New England Journal of Medicine; 354: 2317–2327.
Further Resources
1. Sox HC, Higgins MC and Owens DK. Medical decision making, 2nd Ed. Oxford: Wiley-Blackwell, 2013.
2. Brush JE. The Science of the Art of Medicine. Dementi Milestone Publishing,
2015.
3. Stone JV. Bayes’ Rule. A tutorial introduction to Bayesian analysis. Sebtel Press, 2013.