Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_2767_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
25 Мб
Скачать
7.3 What data tocollect
https://t.me/medicina_free
7.3.7.3 Challenges defining reference standard positive andnegative: strategies when there are more than twocategories
There are a number of situations where accuracy data may not be presented as a simple 2×2 contingency table, including the case of multiple reference categories where the target condition is not classed as present or absent but, for example, as present, possible or absent.
An example is provided in a review of galactomannan for the diagnosis of invasive asper­gillosis (Leeflang 2015), where a composite reference standard classifies the target condi­tion as proven, probable, possible or absent (Table7.3.e). In order to calculate sensitivity and specificity, the reference standard results have to be dichotomized into two categories, representing those with the target condition and those without the target condition. For this review, the reference standard classification was guided by the anticipated treatment decision for each group; the ‘proven’ and ‘probable’ categories would likely be treated with antimycotics and were considered to have the target condition, and ‘possible’ or ‘absent’ would not be treated and were considered not to have the target condition.
There is a further challenge if some studies in a review do not report their results according to all possible reference standard categories. Continuing the aspergillosis example, studies could report results in three categories, e.g. ‘proven’, ‘probable or pos­sible’ and ‘absent’, or in some other format, in which case a decision has to be made regarding how to dichotomize results to allow extraction and how to deal with any vari­ation in the definition of the target condition in subsequent analysis.
Any 2×2 tables that have been derived by combining reference standard categories should be clearly identified on the data extraction form, and the original data presented by the study authors should also be extracted.
Challenges defining index test positive andnegative: inconclusive results
7.3.7.4
The other situation where accuracy data may not be presented as a simple 2×2 contin­gency table is if not all index test results are classified as positive or negative but, for example, as positive, indeterminate or negative, or when a continuous test result is categorized into separate ranges of results with the middle group considered ‘borderline’ (seeChapter4).
Table7.3.e Estimating 2×2 contingency table data when there are multiple reference standard
categories, invasive aspergillosis example
Reported categories of invasive aspergillosis
Proven Probable Possible None Total
Galactomannan positive 1 3  8  2 14
Galactomannan negative 0 2 13 54 69
Estimation of 2×2 data
Galactomannan positive 4 10 14
Galactomannan negative 2 67 69
Column totals 6 77 83
People with invasive aspergillosis
People without invasive aspergillosis
145
7 Collecting data
https://t.me/medicina_free
In order to calculate sensitivity and specificity, inconclusive results are usually con­sidered as either positive or negative results. This decision should be guided by what the clinical decision would be for people with inconclusive results (Shinkins 2013), and
the review authors are not content experts. If some intervention is likely to be initiated, then inconclusive results should be considered positive, but if invasive further investi­gations are indicated, for example, then inconclusive results might be considered nega­tive unless there is strong clinical evidence to support their interpretation as positive. If the clinical pathway for inconclusive results is not clear, or is not definitive, then data can be extracted considering inconclusive results as positive and as negative and sensi­tivity analysis can be used to explore the effect on accuracy.
In the example of an interferon-
gamma release assay for diagnosing active tuberculo­sis in Table7.3.f, study authors excluded borderline results in the main analysis but included them as index test positive in a sensitivity analysis. It is important that the decision is applied consistently across all studies and is agreed within the review author team at the outset. In most cases borderline test results should not be excluded, as that would either over- estimate sensitivity or over- estimate specificity compared to the situ­ation of borderline results being considered negative or positive, as shown in Table7.3.f.
Table7.3.f Estimating 2×2 contingency table data when there are multiple index test categories,
example ofinterferon- gamma release assay (IGRA) fordiagnosis ofactive tuberculosis (TB)
(1) Borderline results excluded People with active TB People with no active TB Row totals
IGRA positive 253 51 304
IGRA negative 58 319 377
IGRA borderline (17) (16) (33)
Column totals 311 370 681
Accuracy measures Sensitivity
253/311 = 81%
(2) Borderline results considered positive
IGRA positive + borderline 270  67 337
IGRA negative  58 319 377
Column totals 328 386 714
Accuracy measures Sensitivity
270/328 = 82%
(3) Borderline results considered negative
IGRA positive 253  51 304
IGRA negative + borderline  75 335 410
Column totals 328 386 714
Accuracy measures Sensitivity
253/328 = 77%
Adapted from Whitworth 2019.
Specificity 319/370 = 86%
Specificity 319/386 = 83%
Specificity 335/386 = 87%
146
7.3 What data tocollect
https://t.me/medicina_free
Any 2×2 tables that have been calculated by combining index test categories should be clearly identified on the data extraction form, and the original data presented by the study authors should also be extracted.
In some reports, the study authors may have excluded intermediate results from the calculation of sensitivity and specificity, or sensitivity may have been calculated based on one cut- off and specificity based on a second cut- off. In those cases it is necessary to reconstruct the 2×2 tables, including all tested study participants with a test result, using a single cut- off for all, and calculate sensitivity and specificity based on that reconstructed 2×2 table, not on the reported results.
Challenges defining index test positive andnegative: test failures
7.3.7.5
Failed index test procedures are distinct from inconclusive results in that there is no valid test result, i.e. they represent some kind of failure of the index test (Shinkins 2013). Test failure could be due to the test not being conducted to a sufficient standard (e.g. inadequate sampling or some technical error with the test) or could be for clinical rea­sons (e.g. concurrent infection masking true results or lack of fasting affecting a blood glucose result). In this case, repeating the test under more optimal conditions is likely to lead to a positive or negative result.
Unlike inconclusive results, it is reasonable for test failures to be excluded from analy­sis if they cannot be considered as either positive or negative. The proportion of test failures should always be extracted and reported in a review, as an indicator of how often the test is likely to fail (Shinkins 2013).
7.3.7.6 Challenges defining index test positive and negative: dealing with multiple thresholds and extracting data from ROC curves or other graphics
Studies of index tests that report results on a continuous scale can report accuracy data at multiple thresholds that may not all be relevant to the review question. Where this situation is anticipated, review authors should set out a strategy in the protocol for selecting the datasets that should be extracted. Possible strategies include selecting:
the ‘standard’ threshold currently used in practice;
the threshold recommended by the manufacturer;
the most commonly reported threshold;
a randomly selected threshold; or
all thresholds.
Data can be extracted from a ROC curve if the thresholds associated with each point on the ROC curve are reported, and if the number with and without the target condition is known. This is not a very precise method of identifying sensitivity and specificity and should be used with caution.
Dot plots can be used to represent the results of biomarker tests. The plots can show the test result for each participant according to the value of the test result (y- axis) and the presence of the target condition or other differential diagnosis on the x- axis. The number of participants with a test result above or below the cut- off value of choice can be counted for each category of participants. Dot plots are often used to represent bio­marker results from studies using multi- gate designs, as depicted in Figure7.3.a.
147
7 Collecting data
4
OD
donors
infection
CoV
CoV
CoV-2
https://t.me/medicina_free
3
450
2
1
0
Non-CoV
Blood
Figure7.3.a Example of a dot plot for estimation of sensitivity and specificity. Source: Okba
2020 / U.S Department of Health and Human Services / Public domain
HCoV MERS-
SARS-
SARS-
Any 2×2 tables that have been calculated by extracting data from graphics should be
clearly identified on the data extraction form.
7.3.7.7 Extracting data fromfigures withsoftware
Numerous tools for extracting data from figures are available, many of which are free. Those available at the time of writing include Plot Digitizer, WebPlotDigitizer, Engauge, Dexter, ycasd and GetData Graph Digitizer. The software works by taking an image of a figure and then digitizing the data points off the figure using the axes and scales set by the users. The numbers exported can be used for systematic reviews, although additional calculations may be needed to obtain or validate accuracy measures.
It has been demonstrated that software is more convenient and accurate than visual estimation or use of a ruler (Gross 2014, Jelicic Kadic 2016). Review authors should con­sider using software for extracting numerical data from figures when the data are not available elsewhere.
7.3.7.8 Corrections formissing data: adjusting forpartial verification bias
Partial verification, whereby some participants eligible for a test accuracy study do not receive the reference standard, can occur for a number of reasons, for example due to the invasive nature of the reference standard, or because there is no practical way of confirming the absence of the target condition in an index test negative participant. It is also an approach used when the condition is rare, only sampling a random subset of those who test negative. Study authors may use statistical approaches to correct for the bias introduced by missing data (de Groot 2011), such as inverse sampling probability weighting. Direct inclusion of the original 2×2 table in a systematic review where cor­rections have been made will re-
introduce the bias that has been corrected for by the analysis. Rather, a new 2×2 table should be created that provides equal estimates of sensitivity and specificity with the same uncertainty as from the weighted analysis.
7.3.7.9 Multiple index tests fromthe same study
Studies that report the use of different index tests in different participants can be considered as different studies, with study naming as suggested in Section7.2.1. More
148
7.3 What data tocollect
https://t.me/medicina_free
often, studies evaluate and report accuracy data for multiple eligible index tests in the same participants.
The same data extraction considerations apply to each index test and threshold. Any differences in the selection criteria for participants undergoing each test should be documented and any differences in the number of participants undergoing each index test, or any resulting differences in the number with the target condition, should be accounted for in the 2×2 tables.
Studies that report accuracy data for multiple eligible index tests could also present data for combinations of different test results. For example, a study of anti- CCP and rheumatoid factor for the diagnosis of rheumatoid arthritis could present results for either test positive or for both tests positive. A worked example is provided in Table7.3.g.
The choice of 2×2 tables to extract should always be guided by the review question and should be specified prior to commencing data extraction.
Table7.3.g Calculating 2×2 contingency table data froma cross- tabulation oftwo index tests against
areference standard
Results of the anti- CCP assays in relation to the presence or absence of RF (reported data Luis Caro- Oleas 2007)
RA patients
(n = 124)
CCP assay
Anti-
QUANTA Lite™ CCP2
Anti- CCP2 positive 67 54.0 3 1.9 70
RF positive 56 45.1 3 1.9 59
RF negative 11  8.9 0 0.0 11
CCP2negative 57 46.0 155 98.1 207
Anti-
RF positive 14 11.3 4 2.5 18
RF negative 43 34.7 151 95.6 194
Estimation of 2×2 data for anti- CCP2
anti- CCP2 positive 56 + 11 = 67 3 + 0 = 3  70
anti- CCP2negative 14 + 43 = 57 4 + 151 = 155 212
Column totals 124 158 282
Estimation of 2×2 data for RF
RF positive 56 + 14 = 70 3 + 4 = 7  77
RF negative 11 + 43 = 54 0 + 151 = 151 205
Column totals 124 158 282
n % n % n
Other groups
(n = 158) Total
(Continued)
149
7 Collecting data
https://t.me/medicina_free
Table7.3.g (Continued)
Estimation of 2×2 data for both tests positive
Both tests positive 56 3  59
Either test negative 11 + 14 + 43 = 68 0 + 4 + 151 = 155 223
Column totals 124 158 282
Estimation of 2×2 data for either test positive
Either test positive 56 + 11 + 14 = 81 3 + 0 + 4 = 7  88
Both tests negative 43 151 194
Column totals 124 158 282
RA, rheumatoid arthritis; RF, rheumatoid factor.
Most software for conducting meta- analysis or plotting data allows only one 2×2 table from each study to be included in a single analysis. Where studies compare different tests in the same participants and review authors wish to include all 2×2 tables for each test in the same analysis, special consideration is needed. For example, a review of commercially available point- of- care biomarker assays for an infectious disease such as tuberculosis or malaria may include studies that compare tests produced by different manufacturers in the same study participants. Studies that evaluate imaging tests often report separate 2×2 tables for different observer interpretations of the same images. Unless it is decided a priori that results for only one observer will be reported, for exam­ple for the most experienced observer, it is usual to extract all available 2×2 tables, clearly labelled by observer experience.
7.3.7.10 Subgroups ofpatients
Accuracy data for subgroups of study participants should be extracted only if there is a stated a priori interest in that subgroup. There is no obligation to extract all possible 2×2 contingency tables from any individual study.
7.3.7.11 Individual patient data
Some reviews (Hooper 2015, Manzotti 2019) present individual patient data (IPD). For example, in case of continuous tests it may be helpful to have all data from all partici­pants, so that ROC plots of individual studies can be reconstructed and alternative thresholds applied. Alternatively, IPD could allow review authors to restrict data to par­ticipants meeting the review question, for example in regard to age or underlying health conditions. If a review looks at a test strategy that combines results from several tests, IPD may allow results for strategies to be obtained where they have not been investi­gated in the original report. If review authors plan to use IPD, this should be specified in the protocol, along with a description of the methods to be used to elicit those data from study authors or, where available, to extract from study reports. The success, or otherwise, of contacting study authors should be documented in the review.
150
7.3 What data tocollect
https://t.me/medicina_free
7.3.7.12 Extracting covariates
Covariates that are relevant to the analysis of study data, or that could be displayed on forest plots of sensitivity and specificity, could be related to study design, participant, index test or reference standard characteristics. Detailed information on each of these should be extracted as described in sections 7.3.2 to 7.3.5. For analysis or display pur­poses, however, a separate field for each covariate can be useful, especially as covari­ates can apply to all 2×2data extracted from an individual study (‘study’ level) or can vary between different 2×2 tables extracted from the same study (‘test’ level).
The decision as to whether a covariate applies at study level or at test level will vary between reviews. For example, a covariate related to participant age would apply at study level if all studies reported data either for children, adults or mixed age groups, or would apply at test level if individual studies reported data for a mixed age group as well as for either adults or children separately. Similarly, if accuracy results are expected to vary according to the experience of the assessor interpreting the test, then 2×2data could be presented for all assessors combined and separately according to the experi­ence level of the assessor, for example high, intermediate or low.
Covariates can be either categorical (with a finite number of categories) or continuous (with an infinite number of values between any two values) variables. Categorical vari­ables are easier for the analysis of test accuracy data, but categorization of continuous variables before data analysis is discouraged, as information will be lost. Categorical covariates are required in order to display results per subgroup on a SROC plot, but either categorical or continuous covariates can be displayed on a forest plot of sensitivi­ties and specificities.
Covariates that will be investigated in heterogeneity analyses should be pre-
specified whenever possible, and any categories to be used should be standardized. In the same way that the number of covariates defined for a review should be guided by the expected number of data sets for the review, the number of categories defined for each covariate should be kept to a minimum in order to maximize the number of data sets in each cat­egory and thereby increase the power of the heterogeneity investigation (see Chapter9, Section9.4.6).
7.3.8 Other information tocollect
Collection of information about the harmful effects of testing may be desirable depend­ing on the nature of the test. In test accuracy studies, adverse events can be collected either systematically or non- systematically. Systematic collection refers to collecting adverse events in the same manner for each participant using defined methods such as a questionnaire or a laboratory test. Non- systematic collection refers to collection of information on adverse events using methods such as open- ended questions (e.g. ‘Have you noticed any symptoms since your last visit?’) or reported by participants spontaneously. In either case, adverse events may be selectively reported based on their severity, and whether the participant suspected that the effect may have been caused by the test, which could lead to bias in the available data.
Further comments by the study authors, for example any explanations they provide for unexpected findings, may be noted; however, it is not necessary to report these in the data extraction. References to other studies that are cited in the study report may be useful, although review authors should be aware of the possibility of citation
151
7 Collecting data
https://t.me/medicina_free
bias (see Chapter6). Documentation of any correspondence with the study authors is important for review transparency.
7.4 Data collection tools
7.4.1 Rationale fordata collection forms
Data collection for systematic reviews should be performed using structured data col­lection forms. These can be paper forms, electronic forms (e.g. Excel or Google Form), or commercially available (e.g. Covidence or EPPI- Reviewer) or custom- built data sys­tems that allow online form building, data entry by several users, data sharing and effi­cient data management (Li 2015). At the time of writing, most commercially available data systems do not have standard extraction form templates for reviews of test accu­racy and, although some modification is possible, 2×2 contingency table data cannot be extracted. All different means of data collection require data collection forms.
The data collection form is a bridge between what is reported by the original investiga­tors (e.g. in journal articles, abstracts, personal correspondence) and what is ultimately reported by the review authors. The data collection form serves several important func­tions (Meade 1997). First, the form is linked directly to the review question and criteria for assessing eligibility of studies, and provides a clear summary of these that can be used to identify and structure the data to be extracted from study reports. Second, the data collection form is the historical record of the provenance of the data used in the review, as well as the multitude of decisions (and changes to decisions) that occur throughout the review process. Third, the form is the source of data for inclusion in an analysis and provides support for judgements related to quality assessment.
Given the important functions of data collection forms, ample time and thought should be invested in their design. Because each review is different, data collection forms will vary across reviews. However, there are many similarities in the types of information that are important. Thus, forms can be adapted from one review to the next. Although we use the term ‘data collection form’ in the singular, in practice it may be a series of forms used for different purposes; for example, a separate form could be used to assess the eligibility of studies for inclusion in the review to assist in the quick identification of studies to be excluded from or included in the review.
Considerations inselecting data collection tools
7.4.2
The choice of data collection tool is largely dependent on review authors’ preferences, the size of the review and resources available to the review author team. Potential advantages and considerations of selecting one data collection tool over another are outlined in Table7.4.a (Li 2015). A significant advantage that data systems have is in data management (Chapter1, Section1.2.4) and re- use, and in their accessibility to multiple author review teams. They make review updates more efficient and facilitate methodo­logical research across reviews. Numerous ‘meta- epidemiological’ studies have been carried out using Cochrane Review data, resulting in methodological advances that would not have been possible if thousands of studies had not all been described using the same data structures in the same system.
152
Table7.4.a Considerations inselecting data collection tools
https://t.me/medicina_free
Paper forms Electronic forms Data systems
7.4 Data collection tools
Examples Forms developed
Suitable review type and team sizes
Resource needs
Advantages Do not rely on
using word processing software
Small-
scale reviews (<10included studies)
Small team with 2 to 3data extractors in the same physical location
Low Low to medium (most if
access to computer and network or internet connectivity
Can record notes and explanations easily
Require minimal software skills
Microsoft Access
Microsoft Excel
Google Forms
Forms developed using word processing software (Microsoft Word, Google Docs, Zoho Writer, etc.)
Small- to medium- scale reviews (10 to 20 studies)
Small to moderate­sized team with 4 to 6data extractors
not all institutions provide access and free word processing software is increasingly available online)
Allow extracted data to be processed electronically for editing and analysis
Allow electronic data storage, sharing and collation
Easy to expand or edit forms as required
Can automate data comparison with additional programming
Can copy data to analysis software without manual re-
entry, reducing
errors
Covidence
Reviewer
EPPI-
Systematic Review Data Repository (SRDR)
DistillerSR (Evidence Partners)
REDCap
For small- , medium- and especially large- scale reviews (>20 studies), as well as reviews that need constant updating
All team sizes, especially large teams (i.e.>6data extractors)
Low to medium (open­tools such as SRDR, or tools for which authors have institutional licences or, for Cochrane Reviews, tools for which Cochrane has obtained licences on behalf of review authors)
High (commercial data systems with no access via an institutional licence)
Specifically designed for data collection for systematic reviews
Allow online data storage, linking and sharing
Easy to expand or edit forms as required
Can be integrated with title/ abstract, full- text screening and other functions
Can link data items to locations in the report to facilitate checking
Can readily automate data comparison between independent data collection for the same study
Allow easy monitoring of progress and performance of the author team
access
(Continued)
153
7 Collecting data
https://t.me/medicina_free
Table7.4.a (Continued)
Paper forms Electronic forms Data systems
Disadvantages Inefficient and
potentially unreliable because data must be entered into software for analysis and reporting
Susceptible to errors
Data collected by multiple authors must be manually collated
Difficult to amend as the review progresses
If the papers are lost, all data will need to be recreated
Require familiarity with software packages to design and use forms
Carry risk of introducing mistakes in data entering or copy- pasting between versions
Possible data security issue if files are not encrypted or securely backed up (especially relevant for IPD)
Facilitate coordination among data collectors such as allocation of studies for collection and monitoring team progress
Allow simultaneous data entry by multiple authors
Can export data directly to analysis software
Can import included and excluded studies, with reasons for exclusion, directly into RevMan (Cochrane Reviews only)
In some cases, improve public accessibility through open data sharing
front and considerable
Up­investment of resources to set up and adapt the form for studies of test accuracy
Requires training of data extractors
Cannot extract 2×2 contingency tables for accuracy data
Structured templates may not be as flexible as electronic forms
Cost of commercial data systems
Require familiarity with data systems
Susceptible to changes in software versions
7.4.3 Design ofa data collection form
Regardless of whether data are collected using a paper or electronic form or a data sys­tem, the key to successful data collection is to construct easy- to- use forms and collect sufficient and unambiguous data that faithfully represent the source in a structured and organized manner (Li 2015). In most cases, a document format should be devel­oped for the form before building an electronic form or a data system. This can be dis­tributed to others, including programmers and data analysts, and used as a guide for
154