Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_2767_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
25 Мб
Скачать
4 Understanding test accuracy measures
https://t.me/medicina_free
4.10 References
Altman DG, Deeks JJ, Sackett DL. Odds ratios should be avoided when events are common.
BMJ 1998; 317: 1318.
Brown LD, Cai TT, DasGupta A. Interval estimation for a binomial proportion. Statistical
Science 2001; 16: 101–117.
Clopper C, Pearson ES. The use of confidence or fiducial limits illustrated in the case of the
binomial. Biometrika 1934; 26: 404–413.
García-
Fiñana M, Hughes DM, Cheyne CP, Burnside G, Stockbridge M, Fowler TA, Fowler VL, Wilcox MH, Semple MG, Buchan I. Performance of the Innova SARS­lateral flow test in the Liverpool asymptomatic testing pilot: population based cohort study. BMJ 2021; 374: n1637.
Glas AS, Lijmer JG, Prins MH, Bonsel GJ, Bossuyt PM. The diagnostic odds ratio: a single
indicator of test performance. Journal of Clinical Epidemiology 2003; 56: 1129–1135.
Graham BL, Steenbruggen I, Miller MR, Barjaktarevic IZ, Cooper BG, Hall GL, Hallstrand TS,
Kaminsky DA, McCarthy K, McCormack MC, Oropez CE, Rosenfeld M, Stanojevic S, Swanney MP, Thompson BR. Standardization of Spirometry 2019 Update. An Official American Thoracic Society and European Respiratory Society Technical Statement. American Journal of Respiratory and Critical Care Medicine 2019; 200: e70–e88.
Hayen A, Macaskill P, Irwig L, Bossuyt P. Appropriate statistical methods are required to
assess diagnostic tests for replacement, add­Epidemiology 2010; 63: 883–891.
Leeflang MM, Deeks JJ, Rutjes AW, Reitsma JB, Bossuyt PM. Bivariate meta-
predictive values of diagnostic tests can be an alternative to bivariate meta- analysis of sensitivity and specificity. Journal of Clinical Epidemiology 2012; 65: 1088–1097.
Shinkins B, Thompson M, Mallett S, Perera R. Diagnostic accuracy studies: how to report
and analyse inconclusive test results. BMJ 2013; 346: f2778.
Simel DL, Feussner JR, DeLong ER, Matchar DB. Intermediate, indeterminate, and
uninterpretable diagnostic test results. Medical Decision Making 1987; 7: 107–114.
Takwoingi Y, Leeflang MM, Deeks JJ. Empirical evidence of the importance of comparative
studies of diagnostic test accuracy. Annals of Internal Medicine 2013; 158: 544–554.
Takwoingi Y. Meta-
medical tests [PhD]. Birmingham (UK): University of Birmingham, 2016.
Wang LW, Fahim MA, Hayen A, Mitchell RL, Baines L, Lord S, Craig JC, Webster AC. Cardiac
testing for coronary artery disease in potential kidney transplant recipients. Cochrane Database of Systematic Reviews 2011; 12: CD008691.
Whitworth HS, Badhan A, Boakye AA, Takwoingi Y, Rees- Roberts M, Partlett C, Lambie H,
Innes J, Cooke G, Lipman M, Conlon C, Macallan D, Chua F, Post FA, Wiselka M, Woltmann G, Deeks JJ, Kon OM, Lalvani A, Abdoyeku D, Davidson R, Dedicoat M, Kunst H, Loebingher MR, Lynn W, Nathani N, O’Connell R, Pozniak A, Menzies S. Clinical utility of existing and second- generation interferon- gamma release assays for diagnostic evaluation of tuberculosis: an observational cohort study. Lancet Infectious Diseases 2019; 19: 193–202.
Wilson EB. Probable inference, the law of succession, and statistical inference. Journal of
the American Statistical Association 1927; 22: 209–212.
analytic approaches for summarising and comparing the accuracy of
on, and triage. Journal of Clinical
CoV- 2 antigen rapid
analysis of
72
Part Three
https://t.me/medicina_free
Methods andpresentation ofsystematic reviews oftestaccuracy
5
https://t.me/medicina_free
Defining thereview question
Mariska M. Leeflang, Clare Davenport and Patrick M. Bossuyt
KEY POINTS
Systematic reviews of test accuracy should aim to address clinically relevant questions
for which knowing a test’s sensitivity and specificity, or other accuracy measures, is important. The objective of these systematic reviews is to collate evidence about the accuracy of
a single test or to compare the accuracy of two or more tests for detecting the same target condition. Other objectives may be to study the differences in accuracy related to test character-
istics, such as the type of assays, procedures or positivity thresholds, or to the setting where the tests can be used. The review question should contain information about the Population (including
where and when they will be tested), Index tests (including competing tests) and Target condition. Identifying the clinical pathway in which the index test(s) will be used helps to refine
the review question. The eligibility criteria for studies should match the elements of the review question and
include the acceptable reference standard(s) used to establish the presence or absence of the target condition.
5.1 Introduction
In Chapter3 we discussed different study designs and terminology for designing a test accuracy study. In this chapter, we focus on how to formulate a review question for a systematic review of test accuracy, the objectives of the review and the eligibility crite­ria. These steps should consider the aim of the review: answering relevant questions forwhich knowing a test’s sensitivity and specificity (or other accuracy measures) is
This chapter should be cited as: Leeflang MM, Davenport C, Bossuyt PM. Chapter5: Defining the review question. In: Deeks JJ, Bossuyt PM, Leeflang MM, Takwoingi Y, editors. Cochrane Handbook for Systematic Reviews of Diagnostic Test Accuracy. 1st edition. Chichester (UK): John Wiley & Sons, 2023:75–96.
Cochrane Handbook for Systematic Reviews of Diagnostic Test Accuracy, First Edition. Edited by Jonathan J. Deeks, Patrick M. Bossuyt, Mariska M. Leeflang and Yemisi Takwoingi. © 2023 The Cochrane Collaboration. Published 2023 by John Wiley & Sons Ltd.
75
5 Defining thereview question
https://t.me/medicina_free
important. That means that the objectives and the study eligibility criteria should address key issues such as the intended use of the test, the population to be tested, the role of the index test(s) in the clinical pathway, the relevant target condition and the acceptable reference standard(s).
Test accuracy is not a fixed property of a test: accuracy describes the performance of a test in specific circumstances. The accuracy of a test may therefore vary with the intended use (e.g. screening versus diagnosis), population (e.g. children versus adults), setting (rural health centre in a low- income country versus urban hospital), prior tests (e.g. only signs and symptoms, or also an X- ray before computed tomography (CT) scanning), level of training (novice versus expert readers) and many more elements. Since test accuracy is variable, review authors should be explicit in specifying the cir­cumstances in the review question. The guiding principle is that the review question should flow from the clinical problem: the intended use of the index test, specified in sufficient detail.
In this chapter, we explain the aims of systematic reviews of test accuracy and their scope. We illustrate ways to derive the review question and we discuss setting the eligibility criteria.
5.2 Aims ofsystematic reviews oftest accuracy
Systematic reviews of test accuracy studies can be used to guide policy and decision­making about the use of a test in current clinical practice. To serve this purpose, system­atic reviews of test accuracy should aim to address clinically relevant questions, for which knowing a test’s sensitivity and specificity, or other accuracy measures, is important. This includes comparisons of two or more tests, where decisions about selecting one or another test are guided by the differences in accuracy between the tests.
Some reviews will focus on whether a test is sufficiently sensitive and sufficiently spe­cific to be used in a well- specified setting. This requires criteria that quantify the clinical performance that a new test must attain to achieve the desired health outcomes. These criteria will be guided by an exploration of the existing testing strategy, how the index test would fit in an alternative testing strategy, and the consequences of management actions based on the test results. This will require value judgements. Methods to define minimally acceptable criteria for accuracy have been suggested in the literature (Pepe 2016, Lord 2019).
Reviews could also explore the amount of variation in test accuracy that results from different test types (e.g. immunofluorescence versus enzyme- linked immunosorbent assays in autoimmune disease) or different samples (e.g. cervical swabs, vaginal swabs, or urine for chlamydia testing), which could help in selecting from the available test options.
In general, the aim of systematic reviews of test accuracy will be to inform decisions about the use of tests in defined populations based on the estimated test accuracy in that population. As discussed in Chapter2, evaluating test accuracy is necessary but not sufficient to assess whether tests will benefit patients. Recommendations about testing will always require consideration of the downstream consequences of testing and test results on patient outcomes.
76
5.3 Identifying theclinical problem
https://t.me/medicina_free
For these decision- making purposes, it is important that the review includes those studies that best answer the question and have a low risk of bias. Therefore, the eligibility criteria should ideally address both clinical relevance and methodological quality.
5.2.1 Investigations ofheterogeneity
Heterogeneity is expected in systematic reviews of test accuracy and is typically higher than in intervention reviews. This may be partly due to calculation of point estimates of test accuracy, which are proportions akin to absolute risks in treatment or control groups in randomized controlled trials (RCTs). This could potentially make the esti­mates more heterogeneous than estimates of relative measures, in addition to differ­ences in patient characteristics and other more diverse in design than RCTs, which may lead to heterogeneity.
Another explanation for heterogeneity is the diversity in target populations and disease severity in the study participants. Where intervention studies usually include participants with the same condition, test accuracy studies include participants who have the target condition and those who do not have the target condition. Furthermore, if a test is used in practice in a wide variety of situations and healthcare settings, then a systematic review about the accuracy of this test may also include studies done in different healthcare settings and thus studies that have included participants from different target populations. There is evidence that the accuracy of a test varies with the target population, and this leads to heterogeneity.
There may also be variation in the index test itself, for example because there may be variability in how the test is used and criteria for test positivity. Studies may also use different reference standards to verify the same target condition. These differences in the index test and reference standard may not always be clearly reported.
As heterogeneity is to be expected in systematic reviews of test accuracy, a relevant objective may therefore be to investigate potential sources of heterogeneity. In Chapter9 and Chapter10, we explain how these investigations can be done. Review authors need to realize that the value of these investigations depends on the number of studies included, how well the studies are reported and how consistently the variables are reported.
factors. Test accuracy studies may also be
5.3 Identifying theclinical problem
The review question will guide the rest of the review process. It helps to define what studies to search for and what search terms to use, and guides the definition of the eli­gibility criteria. The review question also drives the tailoring of the signalling questions for assessing risk of bias and the applicability of the results from the studies included in thereview.
5.3.1 Role ofa new test
Test accuracy expresses the ability of a test to identify correctly persons with and without the target condition. For most target conditions, there will already be a testing
77
5 Defining thereview question
https://t.me/medicina_free
strategy in place. The role of the index test should therefore be positioned relative to this existing strategy.
This also implies that clinical decisions about the use of a test require a comparison, explicit or implicit, between a testing strategy with and one without the index test. Few primary test accuracy studies will undertake a direct comparison between such strategies.
In general, three roles can be defined for a new test relative to an existing test: (1) to select patients for whom follow- up testing may be useful (triaging); (2) to increase the accuracy of a testing strategy, by adding an extra test to the existing strategy (add- on); and (3) to replace one or more tests in the existing strategy with the (new) index test (replacement) (Bossuyt 2006). These three roles will be discussed here.
In a situation where the current diagnostic testing strategy poses too high a burden for too many people, a triage test may select those persons who do not need to go through an existing, more extensive testing strategy (see Box5.3.a). The accuracy of the testing strategy with the triage test should then be compared against the accuracy of the testing strategy without the triage test. For many triage tests, existing studies will only provide the accuracy of the triage test in isolation. In that case, the relevant ques­tion is: what accuracy is needed for this test to fulfil the role of triage test? If the aim of the triage test is to rule out disease, a sufficiently high sensitivity of the index test will be required, as the damage due to missed diagnoses may be high. A sufficient level of specificity is also needed, to justify using the triage test for reducing the number of peo­ple undergoing further testing. Box 5.3.a presents an example of a triage test. Alternatively the main aim of a triage test may be to rule in disease. In colorectal cancer screening, for example, a faecal occult blood test (FOBT) is often used as the first screen­ing test, and FOBT positives are invited to undergo colonoscopy. The sensitivity of FOBT for cancer and its precursors is imperfect, but since colonoscopy is an invasive and costly
procedure, for which there is limited capacity, using FOBT as a triage test is
preferred.
Box 5.3.a Example ofa triage question
Urine biomarkers totriage women who may need toundergo surgery forendometriosis
Endometriosis causes painful periods, chronic lower abdominal pain and difficulty con­ceiving. The most reliable way to diagnose it is to perform laparoscopic surgery and visu­alize the endometrial deposits inside the abdomen. Sending all women with complaints suggestive of endometriosis for surgical diagnosis immediately is undesirable. A non­invasive and easy- to- perform test that could identify women who do not need laparos­copy would minimize the diagnostic burden. Several urine tests are available to serve this purpose. This review investigated whether the proposed urine tests are sufficiently accurate for triaging these women. According to the authors of the review, a sensitivity of equal to or greater than 95% is required.
Source: Adapted from Lui 2015.
78
5.3 Identifying theclinical problem
https://t.me/medicina_free
Another clinical need may arise in situations where the current testing strategy leads to a number of false positives that is considered too high. An extra test may then be added to the current testing strategy, to reduce the number of false positive test results. In such cases, an add- on test may be considered in those testing positive with the current strategy. In the example in Box5.3.b, resectability was the target condition. Additional imaging was considered as an add- on staging test in patients (testing positive) qualifying for resection after conventional staging, to reduce the number of unnecessary surgical procedures in irresectable patients. Add- on tests do not necessarily have to reduce the number of false positives in an existing testing strategy. Alternatively, it is possible that too many patients may be missed by the existing testing strategy, leading to the wish to reduce the number of false negatives. In screening for breast cancer, for example, mammography is the standard proce­dure. For certain high-
risk women, however, additional screening with magnetic resonance imaging is encouraged in some of the guidelines to reduce the number of cancers missed.
More straightforward review questions are those where the new test will replace an existing test, or where there is a choice between a number of tests for the same target condition, in the same setting. For example, many infectious diseases require micros­copy, sophisticated molecular tests or culture to identify the relevant pathogen. In emergency settings and low- resource settings, rapid tests and point- of- care tests may be used to replace these more expensive tests, which also have longer turnaround times (Box5.3.c).
Box 5.3.b Example ofan add- on question
Different imaging modalities asadd- on tests following CT scanning forassessing resectability inpancreatic and periampullary cancer
Periampullary cancer includes cancer of the head and neck of the pancreas, cancer of the distal end of the bile duct, cancer of the ampulla of Vater and cancer of the second part of the duodenum. If the tumour expands over an anatomical area that is too large, then it can no longer be surgically removed. Before starting such a surgical operation, the surgeons need to know if a resection of the tumour is possible. CT scanning is the preferred method to assess resectability, but it underestimates the extent to which periampullary cancer has expanded, leading to the decision to operate while the tumour is no longer resectable (false positive test result). Here, the review question is whether other imaging methods such as magnetic resonance imaging (MRI), positron emission tomography (PET), PET- CT and endoscopic ultrasound (EUS) can be used as an add- on test after CT scanning. The expectation is that when these other methods are used in patients who were thought to have a resectable tumour (CT positive), they may indicate further which patients’ tumours are indeed resectable (true positives) and which tumours turn out to be not resectable (false positives) based on the add- on imaging test.
Source: Adapted from Tamburrino 2016.
79
5 Defining thereview question
https://t.me/medicina_free
Box 5.3.c Example ofa replacement question
Rapid tests toreplace blood culture fordiagnosing enteric fever
Enteric fever may be caused by the bacteria Salmonella typhi (typhus) or Salmonella paratyphi A (paratyphus). The currently preferred tests are time consuming and either
invasive (bone marrow culture) or relatively insensitive (blood culture orthe Widal test, a serological measurement). The review authors wanted to assess which rapid tests would be World Health Organization (WHO)–recommended main diagnostic test for enteric fever.
Source: Adapted from Wijedoru 2017.
sufficiently accurate to replace blood culture in the daily clinical setting as the
Sometimes it is not clear whether a test will replace an existing test or be placed before (triage) or after (add- on) an existing test strategy. Or the new test opens up a completely new test–treatment pathway. In these situations, the systematic review may be used in a more exploratory way and the clinical pathway may not (yet) be clear. Possible new pathways may be used as a starting point for defining the review questions.
5.3.2 Defining theclinical pathway
In specifying the clinical question, it may be helpful to draw the existing and alternative testing strategies. Cochrane Reviews of test accuracy studies contain a section called ‘Clinical pathway’ within the Background, in which the review author is invited to explain how and where the test under evaluation is to be used, relative to other tests for the same target condition in the same population.
Drawing a clinical pathway will provide the review team with the information needed to formulate the review question and the study eligibility criteria. It will also help to specify the actions following testing, as well as the respective consequences of these actions for patients, and may therefore help with the interpretation of the review results (Gopalakrishna 2013). This way, describing the clinical pathway facilitates thinking about the effects of the test on patient- important outcomes, providing a useful descrip­tion of context that is relevant for the methods and results of the review.
Review authors are asked to provide a description, either graphical or narrative, of the current pathway that patients follow just before and after testing (Gopalakrishna
2016). A description of the clinical pathway should contain the following elements: (1) the setting and patient groups to be tested, including relevant prior testing; (2) the index test and any comparator index tests; and (3) subsequent steps after testing, driven by the test result, such as further testing or treatment.
Setting and population
Here, we will start with defining the clinical setting, the purpose of testing and the peo­ple who will be tested with the index test. For a systematic review on triage tools for severe neck injuries in children after an accident (Slaar 2017), the review authors took the emergency department as the starting point (Figure 5.3.a). They distinguished
80
Severe blunt trauma
https://t.me/medicina_free
Clinical
decision rule
(e.g. NEXUS)
5.3 Identifying theclinical problem
Low risk of cervical
spine injury
No imaging
No abnormalitiesFracture or
No further imaging
Figure5.3.a Clinical pathway for neck injuries in children (<18 years). The setting (emergency
department) is not shown in the figure but is stated in the text. Source: Adapted from Slaar 2017.
High risk of cervical
spine injury
Plain radiography of
cervical spine
inconclusive image
Computed tomography
of the cervical spine
Neurological symptoms
or very high risk
Consider computed
tomography of the
cervical spine directly
between children under 8 years old versus children between 8 and 18, as each age group has different injury patterns.
The purpose of the testing may differ. For example, a molecular rapid test for tubercu­losis may be intended as a screening test in persons without any symptoms or pro­posed as a diagnostic test for evaluating symptomatic patients. The purpose of testing should be specified explicitly, as well as the intended use population (asymptomatic people versus symptomatic people). The performance of rapid tests for tuberculosis may differ in people with HIV versus people without HIV, in adults versus children, or in patients pre- selected after imaging versus all- comers. Figure5.3.b shows the clinical pathway for a molecular test for tuberculosis in a screening situation.
The reference standard in evaluations of test accuracy may or may not be part of cur­rent clinical practice. The clinical pathway may therefore not be the same as the research study design for the included studies. For example, in current clinical practice only some patients may receive the reference standard (often only index test positives), whereas the study design to estimate sensitivity and specificity requires that all partici­pants (index test positive and index test negative) receive a reference standard. In Figure5.3.a, children with severe neck trauma will first be assessed using a clinical rule (e.g. the National Emergency X- Radiography Utilization Study (NEXUS) rule) and if that indicates cervical spine injury, the child will be referred for radiography first and, if nec­essary, CT scans afterwards. The clinical rules are the index tests. Radiographs and CT (either separately or combined) could be the reference standard, but have a different role in the pathway.
81
5 Defining thereview question
https://t.me/medicina_free
Systematic screening is initiated by national TB programmes and
other national and subnational public health agencies and partners.
People to be screened may or may not have TB symptoms.
All undergo testing with Xpert
MTBC negative
No further action*
Figure5.3.b Clinical pathway for a molecular WHO- recommended rapid diagnostic test to
screen for pulmonary tuberculosis, irrespective of signs and symptoms. symptoms, evaluate clinically for TB in accordance with national guidelines. DST, drug susceptibility testing; INH, isoniazid; MDR- TB, multidrug- resistant TB; MTBC, Mycobacterium tuberculosis complex; RIF, rifampicin; TB, tuberculosis; Xpert, Xpert MTB/RIF or Xpert MTB/RIF Ultra. Source: Adapted from Shapiro 2021.
MTBC positive, RIF resistance
Treat with first­line regimen in accordance with national guidelines. Consider DST for INH if high risk of INH resistance
not detected
INH resistant Treat with INH­resistance TB regimen. Obtain additional DST
MTBC positive, RIF resistance
detected
Evaluate the patient for MDR-TB risk factors. If high risk, treat with MDR-TB regimen in accordance with national guidelines. Obtain additional DST
*If an individual has
Index test(s)
The next element in specifying the clinical pathway is identifying the tests of interest: one or more index tests. For the molecular WHO- recommended rapid diagnostic test for screening for tuberculosis example, the review authors were interested in Xpert MTB/ RIF and Xpert MTB/RIF Ultra, the next generation of the test. For the review about triage tests for neck injuries in children, the authors focused on tests that may be used to decide whether children should be referred for further imaging. These index tests are usually combinations of findings at physical examination or history taking. Two such tests were already in use in clinical practice (NEXUS and DTM) and so the focus was on these two tools.
Subsequent steps for test positives and test negatives
After defining the tests of interest, the steps after the index test results need to be speci­fied. What will happen if theindex test result is positive? What will happen if the index test is negative? Is there a third option, such as a middle ‘unclear’ or ‘intermediate’ test result? What if the test fails? These steps may sometimes be difficult to define, as there may be variability in practice patterns.
In the example of a blood test for fungal disease in immunocompromised patients, the next step after a positive test is usually referral for high- resolution CT scanning. Some hospitals, however, may decide upon treatment based on the blood test, or patients may be referred for radiographs before undergoing CT scanning. Where rele­vant, the clinical pathway should contain information on all alternative next steps following index test results.
82