Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_2767_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
25 Мб
Скачать
2 Evaluating medical tests
https://t.me/medicina_free
Evaluating the accuracy of a staging test is similar to evaluating a diagnostic test, in that results of the new test are compared with results of a reference standard. However, staging tests often classify patients into more than two categories, which makes the application of accuracy measures such as sensitivity and specificity difficult, as these are designed for binary classifications of data.
Staging tests also frequently function as tests of prognosis, in that patients with severe disease often have poorer outcomes. However, longitudinal studies are required to evaluate the relationship between disease stage and future likely outcomes (the prognostic value of the test) and to assess the accuracy of a staging test.
2.6.5 Prognosis
Prognostic tests are used to predict what the future state of a disease is likely to be. Knowledge of likely future health states may lead to modification of a patient’s treat­ment, or enable the patient to make appropriate plans.
Tests for prognosis are similar to those for predisposition, in that they aim to pre­dict what the future holds for a patient and their evaluation requires longitudinal follow- up. Unlike predisposition studies, prognosis studies examine the relationship between test results and a health outcome, such as death or recovery, rather than the onset of disease. Prognosis may be investigated in a sample of all those who have the disease, while predisposition is investigated in healthy people who do not yet have the disease.
Prognosis studies may be done either prospectively or retrospectively, but due to their longitudinal nature they are not immediately suited for evaluation within the test accuracy framework. However, some clinical scenarios where outcomes are observed over a standard time period, such as predicting pregnancy within a year or the outcome of pregnancy, have strong commonalities with test accuracy studies.
2.6.6 Treatment selection
Stratified or precision medicine focuses on identifying subgroups of patients who will benefit from particular interventions, often through pairing a new test or biomarker with an intervention. The greatest advances are being made where the molecular understanding of a disease leads to development of both a biomarker to detect the molecular target and a targeted therapy. Typically, this involves the use of molecular pathology tests such as genomics, proteomics and metabolomics, which have had their greatest successes in cancer.
Benefit from treatment is defined as a difference in outcomes: a better outcome after treatment compared with the counterfactual outcome when not being treated. Such a difference in outcome can never be directly observed, and a reference standard for ben­efit may be difficult to define. Hence, the evaluation of treatment selection tests rarely fits the test accuracy paradigm.
Evaluation of test–treatment combinations is best achieved through randomized trials. Ideally, randomized trials randomly allocate participants to treatment or a comparator, and investigate whether the observed treatment effect differs between patients who are biomarker positive and those who are biomarker negative, via a test of
30
2.6 Other purposes ofmedical testing
https://t.me/medicina_free
the interaction. Such evaluations are more suited to the format of Cochrane Reviews of interventions than those of test accuracy.
2.6.7 Treatment efficacy
Medical tests may be used in both clinical practice and research to predict whether a treatment has worked or not, particularly to get an early indication of a response to a new treatment when long- term follow- up would otherwise be required. In oncology, such tests are called predictive tests, to distinguish them from prognostic tests, which would predict the future health state under no or under conventional treatment.
Such trials often evaluate the presence or absence of a response after treatment in studies in which all participants receive treatment and compare the proportion of responders between biomarker positives and biomarker negatives. Though the results can be expressed in terms of accuracy, the longitudinal nature of such treatment stud­ies makes them unfit for the test accuracy paradigm.
2.6.8 Therapeutic monitoring
Therapeutic monitoring involves using a test repeatedly in a group of patients to see whether their disease is being appropriately controlled by a therapy. For example, patients who are prescribed anti- hypertensive drugs will have their blood pressure reg­ularly measured to assess whether the drug and its dosage are appropriate. If their blood pressure remains high, the dose may be increased, or further anti- hypertensive agents may be added. If their blood pressure is too low, the drug dosage may be reduced.
To inform treatment decisions reliably, monitoring tests need to have good test accu­racy to correctly classify individuals. However, monitoring tests must also respond to changes in patients’ health state quickly, and their clinical value will also depend on how often they are used and whether the associated changes in management actually improve patient health. Evaluation of whether using a monitoring strategy leads to patient benefit is also best assessed using a randomized trial, as for screening interven­tions, and is amenable to systematic review using the Cochrane methods for reviewing intervention studies.
Surveillance forprogression or recurrence
2.6.9
Tests are also used repeatedly in patients who have a progressive disease or have recently undergone an intervention to check for disease progression or recurrence, which may prompt changes in patient management. This is sometimes called surveil­lance, a form of monitoring.
Surveillance has many similarities to screening, but screening is performed in asymp­tomatic people while surveillance is planned in individuals known to have a disease or known to have undergone an intervention.
Surveillance, like monitoring, requires tests that are accurate but also respond to changes in patient status. Thus, systematic reviews of the accuracy of tests for surveil­lance are informative, but whether or not surveillance testing benefits patients is also best assessed in randomized trials.
31
2 Evaluating medical tests
https://t.me/medicina_free
2.7 Chapter information
Authors: Jonathan J. Deeks (Institute of Applied Health Research, University of Birmingham, UK); Patrick M. Bossuyt (Department of Epidemiology and Data Science, University of Amsterdam, The Netherlands).
Sources of support: Jonathan J. Deeks is a UK National Institute for Health Research (NIHR) Senior Investigator Emeritus. Jonathan J. Deeks is supported by the NIHR Birmingham Biomedical Research Centre at the University Hospitals Birmingham NHS Foundation Trust and the University of Birmingham. The views expressed are those of the authors and not necessarily those of the NHS, the NIHR or the Department of Health and Social Care. The authors declare no other sources of support for writing this chapter.
Declarations of interest: Jonathan J. Deeks is a member of Cochrane’s Diagnostic Test Accuracy Editorial Team. The authors declare no other potential conflicts of interest relevant to the topic of this chapter.
Acknowledgements: The authors would like to thank Jenny Doust and Matthew McInnes for helpful peer review comments.
2.8 References
Bossuyt PM, Lijmer JG, Mol BW. Randomised comparisons of medical tests: sometimes
invalid, not always efficient. Lancet 2000; 356: 1844–1847.
Chang SM, Matchar DB, Smetana GW, Umscheid CA. Methods guide for medical test
reviews. Rockville, MD: Agency for Healthcare Research and Quality; 2012. Available at effectivehealthcare.ahrq.gov/sites/default/files/pdf/methods­overview- 2012.pdf.
European Parliament and Council of the European Union. Regulation (EU) 2017/746 of the
European Parliament and of theCouncil of 5 April 2017. Available at eur­legal- content/EN/TXT/PDF/?uri=CELEX:32017R0746&from=EN.
FDA- NIH Biomarker Working Group. BEST (Biomarkers, EndpointS, and other Tools)
Resource. Silver Spring (MD): Food and Drug Administration (US). Co- published by National Institutes of Health (US), Bethesda (MD), 2016. Available at www.ncbi.nlm.nih. gov/books/NBK326791.
Ferrante di Ruffano L, Davenport C, Eisinga A, Hyde C, Deeks JJ. A capture- recapture
analysis demonstrated that randomized controlled trials evaluating the impact of diagnostic tests on patient outcomes are rare. Journal of Clinical Epidemiology 2012a; 65:282–287.
Ferrante di Ruffano L, Hyde CJ, McCaffery KJ, Bossuyt PMM, Deeks JJ. Assessing the value
of diagnostic tests: a framework for designing and evaluating trials. BMJ 2012b; 344:e686.
Ferrante di Ruffano L, Dinnes J, Taylor- Phillips S, Davenport C, Hyde C, Deeks JJ. Research
waste in diagnostic trials: a methods review evaluating the reporting of test- treatment interventions. BMC Medical Research Methodology 2017a; 17: 32.
32
guidance- tests_
lex.europa.eu/
2.8 References
https://t.me/medicina_free
Ferrante di Ruffano L, Dinnes J, Sitch AJ, Hyde C, Deeks JJ. Test- treatment RCTs are
susceptible to bias: a review of the methodological quality of randomized trials that evaluate diagnostic tests. BMC Medical Research Methodology 2017b; 17: 35–35.
Gøtzsche PC, Jørgensen KJ. Screening for breast cancer with mammography. Cochrane
Database of Systematic Reviews 2013; 6: CD001877.
Hewitson P, Glasziou P, Irwig L, Towler B, Watson E. Screening for colorectal cancer using
the faecal occult blood test, Hemoccult. Cochrane Database of Systematic Reviews 2007; 1: CD001216.
Horvath AR, Lord SJ, StJohn A, Sandberg S, Cobbaert CM, Lorenz S, Monaghan PJ,
Verhagen-
Kamerbeek WD, Ebert C, Bossuyt PM. From biomarkers to medical tests: the
changing landscape of test evaluation. Clinica Chimica Acta 2014; 427: 49–57.
Ilic D, Neuberger MM, Djulbegovic M, Dahm P. Screening for prostate cancer. Cochrane
Database of Systematic Reviews 2013;1: CD004720.
Lord SJ, Irwig L, Simes RJ. When is measuring sensitivity and specificity sufficient to
evaluate a diagnostic test, and when do we need randomized trials? Annals of Internal Medicine 2006; 144: 850–855.
Manser R, Lethaby A, Irving LB, Stone C, Byrnes G, Abramson MJ, Campbell D. Screening for
lung cancer. Cochrane Database of Systematic Reviews 2013; 6: CD001991.
National Institute for Health and Clinical Excellence. Diagnostics assessment programme
manual. Manchester (UK): National Institute for Health and Clinical Excellence, 2011. Available at www.nice.org.uk/Media/Default/About/what-
we- do/NICE- guidance/
NICE- diagnostics- guidance/Diagnostics- assessment- programme- manual.pdf.
Odaga J, Sinclair D, Lokong JA, Donegan S, Hopkins H, Garner P. Rapid diagnostic tests
versus clinical diagnosis for managing people with fever in malaria endemic settings. Cochrane Database of Systematic Reviews 2014; 4: CD008998.
Schünemann HJ, Oxman AD, Brozek J, Glasziou P, Jaeschke R, Vist GE, Williams JW Jr, Kunz
R, Craig J, Montori VM, Bossuyt P, Guyatt GH. Grading quality of evidence and strength of recommendations for diagnostic tests and strategies. BMJ 2008; 336: 1106–1110.
33
3
https://t.me/medicina_free
Understanding thedesign oftest accuracy studies
Patrick M. Bossuyt
KEY POINTS
In the basic design for a test accuracy study, a single group of participants suspected
of the target condition undergo the index test, the test under evaluation and the reference standard. The target condition can be a disease, a disease stage or any pathophysiological
condition that has clinical consequences. The reference standard is usually the best available clinical method for finding out
whether patients have the target condition. Instead of including a single group of participants, accuracy studies can also include two
or more groups of participants, such as healthy controls or patients with a previously obtained diagnosis, without performing a reference standard in all participants. Test accuracy studies can rely on more than one reference standard. The reference
standards can be a single test or procedure, a rule based on multiple procedures, apanel­Comparative accuracy studies evaluate two or more index tests. The most valid designs
are fully paired designs and randomized designs.
based adjudication or can be based on latent class modelling.
3.1 Introduction
Test accuracy studies evaluate the performance of a medical test: to what extent the test is able to distinguish patients with the target condition from those without. These studies do so by comparing the results of one or more index tests with the classification obtained with the reference standard, which typically is the best available clinical method for identifying patients with the target condition.
This chapter should be cited as: Bossuyt PM. Chapter3: Understanding the design of test accuracy studies. In: Deeks JJ, Bossuyt PM, Leeflang MM, Takwoingi Y, editors. Cochrane Handbook for Systematic Reviews of Diagnostic Test Accuracy. 1st edition. Chichester (UK): John Wiley & Sons, 2023:35–52.
Cochrane Handbook for Systematic Reviews of Diagnostic Test Accuracy, First Edition. Edited by Jonathan J. Deeks, Patrick M. Bossuyt, Mariska M. Leeflang and Yemisi Takwoingi. © 2023 The Cochrane Collaboration. Published 2023 by John Wiley & Sons Ltd.
35
3 Understanding thedesign oftest accuracy studies
https://t.me/medicina_free
Estimates of test accuracy can be obtained from several types of studies. Before starting a review, review authors should understand what kind of primary study would ideally fit the review question. This helps in defining the search strategy and the eligibility criteria, and is essential for preparing evaluations of the risk of bias in included studies.
This chapter provides an overview of the most common designs used to assess test accuracy. It starts with a description of the basic study design for evaluating the accu­racy of a single test, followed by a description of variations of this basic study design. It then identifies various types of reference standards. Another section describes designs for studies to evaluate and compare the accuracy of two or more tests.
3.2 The basic design fora test accuracy study
Systematic reviews of test accuracy studies are often performed to support clinical decisions and recommendations about the use of specific tests, in a well- defined con­text (Bossuyt 2020). To be informative, the studies included in such a review should match the intended use setting.
In the basic design for a test accuracy study, a single group of consecutive partici­pants is included, all suspected of having the target condition. This study group should be representative for what we will refer to as the target population, or intended use population. This population consists of all patients in whom the test is currently used or will be used. For a new test, this will be the total group of patients in whom the test is considered in the future. For an existing test, the target population consists of all patients for whom the test is currently used. Ideally, the study eligibility criteria define the target population, while the study exclusion criteria specify those in the target population that cannot become a member of the study group.
In addition to sampling the study group from the target population, researchers should also recruit the study group in the intended use setting. The accuracy of many tests is known to vary across settings, and performance of the tests often differs in general practice compared to more specialized settings. For this reason, a test accu­racystudy should preferably be performed in the setting in which the test is used, or will be used.
Most medical tests intend to measure a specific quantity: the measurand. D- dimer assays, for example, intend to measure D- dimer, a fibrin degradation product: a small protein fragment present in the blood after a blood clot is degraded by fibrinolysis.
Very often, the purpose of testing is not so much the measurement itself, but a spe­cific clinical purpose, such as making a diagnosis (see Chapter2). Measuring D- dimer, for example, is used by clinicians to assist in the diagnosis of patients with suspected deep venous thrombosis or suspected pulmonary embolism.
More generally, we refer to the target condition: the condition to be detected with the test. Very often the target condition is a specific disease. Since many diseases vary in severity, it can also be a specific disease stage, as not all stages may be clinically conse­quential. In vascular medicine, for example, there have been discussions about the clinical relevance of subsegmental pulmonary embolism, and whether such emboli should be included in the definition of the target condition when evaluating tests for patients with suspected pulmonary embolism (Baumgartner 2020).
36
3.2 The basic design fora test accuracy study
https://t.me/medicina_free
The target condition can also be a precursor of disease. An example is adenoma in thecolon, which can develop into colon cancer. Tests to screen for colorectal cancer therefore aim to detect not only early cancer, but also other forms of neoplasia. The definition of what should be considered ‘advanced neoplasia’ has changed over time (Young 2016).
The target condition can also be a specific pathophysiological condition, such as hyperlipidaemia: abnormally elevated levels of any or all lipids or lipoproteins in the blood (Lord 2011).
In almost all cases, the target condition is a condition with clinical consequences. People with the target condition could benefit from treatment or from other actions, such as further testing. Detecting the target condition would therefore benefit those being tested, while missing the target condition could cause potential harm to those being tested.
All study group participants in a test accuracy study undergo the test that is being evaluated: the index test. This is the test for which estimates of accuracy will be reported. In a test accuracy study, this can be a single test, two or more tests or a test strategy, consisting of a specific sequence of medical tests.
Very shortly thereafter, or at the same time, all study group participants undergo a second procedure: the reference standard. The reference standard is usually the best available clinical method for finding out whether patients have the target condition. The reference standard can also be a single test or a combination of tests.
Like the target condition, what should be the reference standard is not always fixed; it can vary over time and may vary with changes in the definition of the target condition. See Section3.5 for more information on the different types of reference standards.
In the analysis of a test accuracy study, the index test result of each study participant is then compared with the corresponding result, for the same participant, obtained with the reference standard. Once this cross-
classification is done for every study group member, the comparisons are aggregated and expressed as an estimate of the accuracy of the index test, in that specific setting, for the target population. Accuracy expresses the clinical performance of medical tests: how well this test is able to identify those with the target condition in all who undergo testing. Chapter4 discusses the various meas­ures of test accuracy in more detail.
Figure3.2.a shows a schema for the test accuracy study described in Box 3.2.a. The index test, a D- dimer test, is presented in light blue in Figure3.2.a, the reference standard in orange. The same colours will be used in the other examples in this chapter. In this example, CT imaging is used as the reference standard. In another study for a different target condition, CT imaging may be the test under evaluation: one of the index tests, as shown in Figure3.6.a and Figure3.6.b.
In principle, questions about the accuracy of a medical test require a cross- sectional study design: we want to know how well the test performs in identifying patients at the time of testing, not whether the patients will develop the target condition in the near future (that would be a prognostic question), or whether the patients had a pulmonary embolism in the past (which may be relevant for answering questions about causality).
Since the question about test accuracy is essentially a cross- sectional one, it does notreally matter whether the index test is performed first or the reference standard is performed first. Figure3.2.b shows a different schema for the study summarized in Figure3.2.a. In this case, the blood sample is taken shortly after CT imaging, instead
37
3 Understanding thedesign oftest accuracy studies
https://t.me/medicina_free
ED patients with
suspected pulmonary embolism
Simplify D-dimer
CT-angiography
Cross-classification
Figure3.2.a An example of the basic design for a test accuracy study
ofbefore CT imaging. CT imaging (in orange) is still the reference standard and D- dimer (in blue) the index test. The diagnostic accuracy estimates generated in this study will not be fundamentally different from those obtained with the study in Figure 3.2.a, provided the time interval between D- dimer testing and imaging is sufficiently short inboth.
Because of the relatively recent interest in study designs for test accuracy studies, there is no agreed terminology to describe the different types of studies. Many terms borrowed from epidemiology are ill- fitted to describe this type of study. The study in Example1 (see Box 3.2.a) is sometimes described as a cohort study, but we should keep in mind that, unlike cohort studies in epidemiology, the design of the study in Example1
Box 3.2.a Example 1: Basic design of a test accuracy study
A research group wants to evaluate the diagnostic accuracy of the Simplify D- dimer in correctly identifying patients with pulmonary embolism presenting at the Emergency Department (ED) of a US hospital. The Simplify D- dimer is a point- of- care test that requires a drop of whole blood from the patient undergoing the test. The researchers registered the study in a clinical trials registry before it started and obtained Institutional Review Board approval for the protocol. During the four years it took to complete the study, one of a group of emergency physicians identified ED patients with suspected pulmonary embolism based on presenting symptoms such as dyspnoea, chest pain or syncope, or physical signs such as a rapid pulse or low pulse oximetry reading that could not readily be explained by another disease process. Eligible patients were invited to participate in the study and asked for informed consent. A data collection form was then completed, after which blood was drawn by a qualified respiratory therapist and the D- dimer test was performed. Immediately thereafter, patients underwent computed tomography (CT) angiography of the chest, the reference standard. The radiologist reading the CT images was unaware of the D- dimer results. After the study was completed, all D- dimer results were cross classified with the imaging results, to obtain estimates of the diagnostic accuracy of this Simplify D- dimer test.
38
3.3 Multiple groups ofparticipants
https://t.me/medicina_free
ED patients with
suspected pulmonary embolism
CT-angiography
Simplify D-dimer
Cross-classification
Figure3.2.b A variation on the basic design for a test accuracy study
is not longitudinal but cross- sectional. The variation in Figure3.2.b is sometimes called a case- control study (since the reference standard is performed shortly before the index test), but that term is misleading, and ‘reverse flow’ study design has been suggested as an alternative description (Rutjes 2005).
When the target condition is relatively rare, as in cancer screening, performing the index tests in all participants with the target condition and in all participants without the target condition may be inefficient: the second group may be much larger, generat­ing a precision beyond what is needed for decision- making. In that case, performing the index test in all participants with the target condition and in a random subset of those without the target condition may be a more efficient alternative. This presents a varia­tion on the study design in Figure3.2.b.
An example is a study of DNA methylation- based detection, where a large group of patients had a CT scan for suspected lung cancer, with pathology in case lesions were detected (reference standard (Liu 2020)). The investigators recruited from this single group, after the CT results were known, 74 patients with pathological confirmation of non- small cell lung cancer and 27 patients with a non- cancer diagnosis. The investigators then reported the DNA methylation results in each group, but both these groups were selected from a single study group representing the target population: patients with suspected lung cancer. Even though only 101 participants were included in the analysis, there was a single set of study eligibility criteria, and every participant in this single study group had the same reference standard.
The basic design has a single group of participants, a single index test and a single reference standard. In the following sections we discuss variations on this basic design for each of these three components.
3.3 Multiple groups ofparticipants
In the basic study design, there is a single set of eligibility criteria: one set of inclu­sion criteria that define the target population, the population from whom the study group is sampled and to whom the estimates of test accuracy are assumed to apply.
39
3 Understanding thedesign oftest accuracy studies
https://t.me/medicina_free
Stroke patients
VR-kinematics VR-kinematics
Classification
Figure3.3.a A two- group (or two- gate) test accuracy study
Healthy controls
In Example1, these were patients presenting to the ED with suspected pulmonary embolism.
In having a single set of eligibility criteria, the basic test accuracy study design is simi­lar to that of randomized trials, which also have a single set of inclusion and exclusion criteria. Anyone in the target population either has the target condition or not (pulmo­nary embolism in this case). In principle, the eligibility criteria are defined in such a way that anyone in the target population could have been included in the study group. The study group is therefore a truly random subset of the target population, and measures of test accuracy are assumed to apply to the target population.
A variation on the basic study design, which is commonly observed in test accuracy studies, is the study that uses two or more sets of eligibility criteria (Figure3.3.a). In the study in Example2 (see Box 3.3.a), the authors reported on the discriminative validity of VR- based kinematics by presenting estimates of sensitivity and specificity, which are common measures of test accuracy. They had two sets of eligibility criteria: one for stroke patients and a second set for the healthy controls. Each group was defined in a different way and participants in each group were recruited in a different way. Because of the two distinct groups, we refer to these studies as two- group studies or two- gate studies. Rutjes and colleagues have suggested the term “two- gate studies” for such
Box 3.3.a Example 2: Study with two groups of participants
A group of researchers wants to evaluate kinematic analysis using a virtual reality (VR) environment for evaluating motor function in stroke patients. They relied on two groups of participants:
stroke patients with mild to moderate impairment; and
healthy controls, recruited through a website.
The report estimates sensitivity and specificity. Sensitivity refers to the proportion of individuals with stroke diagnosis correctly identified as having stroke by the cut- off val­ues for the kinematic variables. Specificity refers to the proportion of healthy controls who are correctly identified as controls by the cut- off values for the kinematic variables. This example is based on a study by Hussain and colleagues (Hussain 2018).
40