Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_2767_Библиотеки_им_академика_М_И_Перельмана
.pdf
7.3 What data tocollect
https://t.me/medicina_free
7.3.7.3 Challenges defining reference standard positive andnegative: strategies
when there are more than twocategories
There are a number of situations where accuracy data may not be presented as a simple
2×2 contingency table, including the case of multiple reference categories where the
target condition is not classed as present or absent but, for example, as present, possible
or absent.
An example is provided in a review of galactomannan for the diagnosis of invasive aspergillosis (Leeflang 2015), where a composite reference standard classifies the target condition as proven, probable, possible or absent (Table7.3.e). In order to calculate sensitivity
and specificity, the reference standard results have to be dichotomized into two categories,
representing those with the target condition and those without the target condition. For
this review, the reference standard classification was guided by the anticipated treatment
decision for each group; the ‘proven’ and ‘probable’ categories would likely be treated with
antimycotics and were considered to have the target condition, and ‘possible’ or ‘absent’
would not be treated and were considered not to have the target condition.
There is a further challenge if some studies in a review do not report their results
according to all possible reference standard categories. Continuing the aspergillosis
example, studies could report results in three categories, e.g. ‘proven’, ‘probable or possible’ and ‘absent’, or in some other format, in which case a decision has to be made
regarding how to dichotomize results to allow extraction and how to deal with any variation in the definition of the target condition in subsequent analysis.
Any 2×2 tables that have been derived by combining reference standard categories
should be clearly identified on the data extraction form, and the original data presented
by the study authors should also be extracted.
Challenges defining index test positive andnegative: inconclusive results
7.3.7.4
The other situation where accuracy data may not be presented as a simple 2×2 contingency table is if not all index test results are classified as positive or negative but, for
example, as positive, indeterminate or negative, or when a continuous test result is
categorized into separate ranges of results with the middle group considered ‘borderline’
(seeChapter4).
Table7.3.e Estimating 2×2 contingency table data when there are multiple reference standard
categories, invasive aspergillosis example
Reported categories of invasive aspergillosis
Proven Probable Possible None Total
Galactomannan positive 1 3 8 2 14
Galactomannan negative 0 2 13 54 69
Estimation of 2×2 data
Galactomannan positive 4 10 14
Galactomannan negative 2 67 69
Column totals 6 77 83
People with invasive
aspergillosis
People without invasive
aspergillosis
145

7 Collecting data
https://t.me/medicina_free
In order to calculate sensitivity and specificity, inconclusive results are usually considered as either positive or negative results. This decision should be guided by what
the clinical decision would be for people with inconclusive results (Shinkins 2013), and
the review authors are not content experts. If some intervention is likely to be initiated,
then inconclusive results should be considered positive, but if invasive further investigations are indicated, for example, then inconclusive results might be considered negative unless there is strong clinical evidence to support their interpretation as positive. If
the clinical pathway for inconclusive results is not clear, or is not definitive, then data
can be extracted considering inconclusive results as positive and as negative and sensitivity analysis can be used to explore the effect on accuracy.
In the example of an interferon-
gamma release assay for diagnosing active tuberculosis in Table7.3.f, study authors excluded borderline results in the main analysis but
included them as index test positive in a sensitivity analysis. It is important that the
decision is applied consistently across all studies and is agreed within the review author
team at the outset. In most cases borderline test results should not be excluded, as that
would either over- estimate sensitivity or over- estimate specificity compared to the situation of borderline results being considered negative or positive, as shown in Table7.3.f.
Table7.3.f Estimating 2×2 contingency table data when there are multiple index test categories,
example ofinterferon- gamma release assay (IGRA) fordiagnosis ofactive tuberculosis (TB)
(1) Borderline results excluded People with active TB People with no active TB Row totals
IGRA positive 253 51 304
IGRA negative 58 319 377
IGRA borderline (17) (16) (33)
Column totals 311 370 681
Accuracy measures Sensitivity
253/311 = 81%
(2) Borderline results considered positive
IGRA positive + borderline 270 67 337
IGRA negative 58 319 377
Column totals 328 386 714
Accuracy measures Sensitivity
270/328 = 82%
(3) Borderline results considered negative
IGRA positive 253 51 304
IGRA negative + borderline 75 335 410
Column totals 328 386 714
Accuracy measures Sensitivity
253/328 = 77%
Adapted from Whitworth 2019.
Specificity
319/370 = 86%
Specificity
319/386 = 83%
Specificity
335/386 = 87%
146

7.3 What data tocollect
https://t.me/medicina_free
Any 2×2 tables that have been calculated by combining index test categories should
be clearly identified on the data extraction form, and the original data presented by the
study authors should also be extracted.
In some reports, the study authors may have excluded intermediate results from the
calculation of sensitivity and specificity, or sensitivity may have been calculated based
on one cut- off and specificity based on a second cut- off. In those cases it is necessary to
reconstruct the 2×2 tables, including all tested study participants with a test result,
using a single cut- off for all, and calculate sensitivity and specificity based on that
reconstructed 2×2 table, not on the reported results.
Challenges defining index test positive andnegative: test failures
7.3.7.5
Failed index test procedures are distinct from inconclusive results in that there is no
valid test result, i.e. they represent some kind of failure of the index test (Shinkins 2013).
Test failure could be due to the test not being conducted to a sufficient standard (e.g.
inadequate sampling or some technical error with the test) or could be for clinical reasons (e.g. concurrent infection masking true results or lack of fasting affecting a blood
glucose result). In this case, repeating the test under more optimal conditions is likely
to lead to a positive or negative result.
Unlike inconclusive results, it is reasonable for test failures to be excluded from analysis if they cannot be considered as either positive or negative. The proportion of test
failures should always be extracted and reported in a review, as an indicator of how
often the test is likely to fail (Shinkins 2013).
7.3.7.6 Challenges defining index test positive and negative: dealing with multiple
thresholds and extracting data from ROC curves or other graphics
Studies of index tests that report results on a continuous scale can report accuracy data
at multiple thresholds that may not all be relevant to the review question. Where this
situation is anticipated, review authors should set out a strategy in the protocol for
selecting the datasets that should be extracted. Possible strategies include selecting:
●
the ‘standard’ threshold currently used in practice;
●
the threshold recommended by the manufacturer;
●
the most commonly reported threshold;
●
a randomly selected threshold; or
●
all thresholds.
Data can be extracted from a ROC curve if the thresholds associated with each point on
the ROC curve are reported, and if the number with and without the target condition is
known. This is not a very precise method of identifying sensitivity and specificity and
should be used with caution.
Dot plots can be used to represent the results of biomarker tests. The plots can show
the test result for each participant according to the value of the test result (y- axis) and
the presence of the target condition or other differential diagnosis on the x- axis. The
number of participants with a test result above or below the cut- off value of choice can
be counted for each category of participants. Dot plots are often used to represent biomarker results from studies using multi- gate designs, as depicted in Figure7.3.a.
147

7 Collecting data
4
OD
donors
infection
CoV
CoV
CoV-2
https://t.me/medicina_free
3
450
2
1
0
Non-CoV
Blood
Figure7.3.a Example of a dot plot for estimation of sensitivity and specificity. Source: Okba
2020 / U.S Department of Health and Human Services / Public domain
HCoV MERS-
SARS-
SARS-
Any 2×2 tables that have been calculated by extracting data from graphics should be
clearly identified on the data extraction form.
7.3.7.7 Extracting data fromfigures withsoftware
Numerous tools for extracting data from figures are available, many of which are free.
Those available at the time of writing include Plot Digitizer, WebPlotDigitizer, Engauge,
Dexter, ycasd and GetData Graph Digitizer. The software works by taking an image of a
figure and then digitizing the data points off the figure using the axes and scales set by
the users. The numbers exported can be used for systematic reviews, although additional
calculations may be needed to obtain or validate accuracy measures.
It has been demonstrated that software is more convenient and accurate than visual
estimation or use of a ruler (Gross 2014, Jelicic Kadic 2016). Review authors should consider using software for extracting numerical data from figures when the data are not
available elsewhere.
7.3.7.8 Corrections formissing data: adjusting forpartial verification bias
Partial verification, whereby some participants eligible for a test accuracy study do not
receive the reference standard, can occur for a number of reasons, for example due to
the invasive nature of the reference standard, or because there is no practical way of
confirming the absence of the target condition in an index test negative participant. It is
also an approach used when the condition is rare, only sampling a random subset of
those who test negative. Study authors may use statistical approaches to correct for the
bias introduced by missing data (de Groot 2011), such as inverse sampling probability
weighting. Direct inclusion of the original 2×2 table in a systematic review where corrections have been made will re-
introduce the bias that has been corrected for by the
analysis. Rather, a new 2×2 table should be created that provides equal estimates of
sensitivity and specificity with the same uncertainty as from the weighted analysis.
7.3.7.9 Multiple index tests fromthe same study
Studies that report the use of different index tests in different participants can be
considered as different studies, with study naming as suggested in Section7.2.1. More
148

7.3 What data tocollect
https://t.me/medicina_free
often, studies evaluate and report accuracy data for multiple eligible index tests in the
same participants.
The same data extraction considerations apply to each index test and threshold. Any
differences in the selection criteria for participants undergoing each test should be
documented and any differences in the number of participants undergoing each
index test, or any resulting differences in the number with the target condition, should
be accounted for in the 2×2 tables.
Studies that report accuracy data for multiple eligible index tests could also present
data for combinations of different test results. For example, a study of anti- CCP and
rheumatoid factor for the diagnosis of rheumatoid arthritis could present results for
either test positive or for both tests positive. A worked example is provided in Table7.3.g.
The choice of 2×2 tables to extract should always be guided by the review question
and should be specified prior to commencing data extraction.
Table7.3.g Calculating 2×2 contingency table data froma cross- tabulation oftwo index tests against
areference standard
Results of the anti- CCP assays in relation to the presence or absence of RF (reported data Luis
Caro- Oleas 2007)
RA patients
(n = 124)
CCP assay
Anti-
QUANTA Lite™ CCP2
Anti- CCP2 positive 67 54.0 3 1.9 70
RF positive 56 45.1 3 1.9 59
RF negative 11 8.9 0 0.0 11
CCP2negative 57 46.0 155 98.1 207
Anti-
RF positive 14 11.3 4 2.5 18
RF negative 43 34.7 151 95.6 194
Estimation of 2×2 data for anti- CCP2
anti- CCP2 positive 56 + 11 = 67 3 + 0 = 3 70
anti- CCP2negative 14 + 43 = 57 4 + 151 = 155 212
Column totals 124 158 282
Estimation of 2×2 data for RF
RF positive 56 + 14 = 70 3 + 4 = 7 77
RF negative 11 + 43 = 54 0 + 151 = 151 205
Column totals 124 158 282
n % n % n
Other groups
(n = 158) Total
(Continued)
149

7 Collecting data
https://t.me/medicina_free
Table7.3.g (Continued)
Estimation of 2×2 data for both tests positive
Both tests positive 56 3 59
Either test negative 11 + 14 + 43 = 68 0 + 4 + 151 = 155 223
Column totals 124 158 282
Estimation of 2×2 data for either test positive
Either test positive 56 + 11 + 14 = 81 3 + 0 + 4 = 7 88
Both tests negative 43 151 194
Column totals 124 158 282
RA, rheumatoid arthritis; RF, rheumatoid factor.
Most software for conducting meta- analysis or plotting data allows only one 2×2 table
from each study to be included in a single analysis. Where studies compare different
tests in the same participants and review authors wish to include all 2×2 tables for each
test in the same analysis, special consideration is needed. For example, a review of
commercially available point- of- care biomarker assays for an infectious disease such as
tuberculosis or malaria may include studies that compare tests produced by different
manufacturers in the same study participants. Studies that evaluate imaging tests often
report separate 2×2 tables for different observer interpretations of the same images.
Unless it is decided a priori that results for only one observer will be reported, for example for the most experienced observer, it is usual to extract all available 2×2 tables,
clearly labelled by observer experience.
7.3.7.10 Subgroups ofpatients
Accuracy data for subgroups of study participants should be extracted only if there is a
stated a priori interest in that subgroup. There is no obligation to extract all possible
2×2 contingency tables from any individual study.
7.3.7.11 Individual patient data
Some reviews (Hooper 2015, Manzotti 2019) present individual patient data (IPD). For
example, in case of continuous tests it may be helpful to have all data from all participants, so that ROC plots of individual studies can be reconstructed and alternative
thresholds applied. Alternatively, IPD could allow review authors to restrict data to participants meeting the review question, for example in regard to age or underlying health
conditions. If a review looks at a test strategy that combines results from several tests,
IPD may allow results for strategies to be obtained where they have not been investigated in the original report. If review authors plan to use IPD, this should be specified in
the protocol, along with a description of the methods to be used to elicit those data
from study authors or, where available, to extract from study reports. The success, or
otherwise, of contacting study authors should be documented in the review.
150

7.3 What data tocollect
https://t.me/medicina_free
7.3.7.12 Extracting covariates
Covariates that are relevant to the analysis of study data, or that could be displayed on
forest plots of sensitivity and specificity, could be related to study design, participant,
index test or reference standard characteristics. Detailed information on each of these
should be extracted as described in sections 7.3.2 to 7.3.5. For analysis or display purposes, however, a separate field for each covariate can be useful, especially as covariates can apply to all 2×2data extracted from an individual study (‘study’ level) or can
vary between different 2×2 tables extracted from the same study (‘test’ level).
The decision as to whether a covariate applies at study level or at test level will vary
between reviews. For example, a covariate related to participant age would apply at
study level if all studies reported data either for children, adults or mixed age groups, or
would apply at test level if individual studies reported data for a mixed age group as
well as for either adults or children separately. Similarly, if accuracy results are expected
to vary according to the experience of the assessor interpreting the test, then 2×2data
could be presented for all assessors combined and separately according to the experience level of the assessor, for example high, intermediate or low.
Covariates can be either categorical (with a finite number of categories) or continuous
(with an infinite number of values between any two values) variables. Categorical variables are easier for the analysis of test accuracy data, but categorization of continuous
variables before data analysis is discouraged, as information will be lost. Categorical
covariates are required in order to display results per subgroup on a SROC plot, but
either categorical or continuous covariates can be displayed on a forest plot of sensitivities and specificities.
Covariates that will be investigated in heterogeneity analyses should be pre-
specified
whenever possible, and any categories to be used should be standardized. In the same
way that the number of covariates defined for a review should be guided by the expected
number of data sets for the review, the number of categories defined for each covariate
should be kept to a minimum in order to maximize the number of data sets in each category and thereby increase the power of the heterogeneity investigation (see Chapter9,
Section9.4.6).
7.3.8 Other information tocollect
Collection of information about the harmful effects of testing may be desirable depending on the nature of the test. In test accuracy studies, adverse events can be collected
either systematically or non- systematically. Systematic collection refers to collecting
adverse events in the same manner for each participant using defined methods such as
a questionnaire or a laboratory test. Non- systematic collection refers to collection of
information on adverse events using methods such as open- ended questions (e.g.
‘Have you noticed any symptoms since your last visit?’) or reported by participants
spontaneously. In either case, adverse events may be selectively reported based on
their severity, and whether the participant suspected that the effect may have been
caused by the test, which could lead to bias in the available data.
Further comments by the study authors, for example any explanations they provide
for unexpected findings, may be noted; however, it is not necessary to report these
in the data extraction. References to other studies that are cited in the study report
may be useful, although review authors should be aware of the possibility of citation
151

7 Collecting data
https://t.me/medicina_free
bias (see Chapter6). Documentation of any correspondence with the study authors
is important for review transparency.
7.4 Data collection tools
7.4.1 Rationale fordata collection forms
Data collection for systematic reviews should be performed using structured data collection forms. These can be paper forms, electronic forms (e.g. Excel or Google Form),
or commercially available (e.g. Covidence or EPPI- Reviewer) or custom- built data systems that allow online form building, data entry by several users, data sharing and efficient data management (Li 2015). At the time of writing, most commercially available
data systems do not have standard extraction form templates for reviews of test accuracy and, although some modification is possible, 2×2 contingency table data cannot
be extracted. All different means of data collection require data collection forms.
The data collection form is a bridge between what is reported by the original investigators (e.g. in journal articles, abstracts, personal correspondence) and what is ultimately
reported by the review authors. The data collection form serves several important functions (Meade 1997). First, the form is linked directly to the review question and criteria for
assessing eligibility of studies, and provides a clear summary of these that can be used
to identify and structure the data to be extracted from study reports. Second, the data
collection form is the historical record of the provenance of the data used in the review,
as well as the multitude of decisions (and changes to decisions) that occur throughout
the review process. Third, the form is the source of data for inclusion in an analysis and
provides support for judgements related to quality assessment.
Given the important functions of data collection forms, ample time and thought
should be invested in their design. Because each review is different, data collection
forms will vary across reviews. However, there are many similarities in the types of
information that are important. Thus, forms can be adapted from one review to the
next. Although we use the term ‘data collection form’ in the singular, in practice it may
be a series of forms used for different purposes; for example, a separate form could be
used to assess the eligibility of studies for inclusion in the review to assist in the quick
identification of studies to be excluded from or included in the review.
Considerations inselecting data collection tools
7.4.2
The choice of data collection tool is largely dependent on review authors’ preferences,
the size of the review and resources available to the review author team. Potential
advantages and considerations of selecting one data collection tool over another are
outlined in Table7.4.a (Li 2015). A significant advantage that data systems have is in data
management (Chapter1, Section1.2.4) and re- use, and in their accessibility to multiple
author review teams. They make review updates more efficient and facilitate methodological research across reviews. Numerous ‘meta- epidemiological’ studies have been
carried out using Cochrane Review data, resulting in methodological advances that
would not have been possible if thousands of studies had not all been described using
the same data structures in the same system.
152

Table7.4.a Considerations inselecting data collection tools
https://t.me/medicina_free
Paper forms Electronic forms Data systems
7.4 Data collection tools
Examples Forms developed
Suitable
review type
and team sizes
Resource
needs
Advantages Do not rely on
using word
processing
software
Small-
scale reviews
(<10included
studies)
Small team with 2
to 3data extractors
in the same
physical location
Low Low to medium (most if
access to computer
and network or
internet
connectivity
Can record notes
and explanations
easily
Require minimal
software skills
Microsoft Access
Microsoft Excel
Google Forms
Forms developed using
word processing
software (Microsoft
Word, Google Docs,
Zoho Writer, etc.)
Small- to medium- scale
reviews (10 to 20
studies)
Small to moderatesized team with 4 to
6data extractors
not all institutions
provide access and free
word processing
software is increasingly
available online)
Allow extracted data to
be processed
electronically for
editing and analysis
Allow electronic data
storage, sharing and
collation
Easy to expand or edit
forms as required
Can automate data
comparison with
additional
programming
Can copy data to
analysis software
without manual
re-
entry, reducing
errors
Covidence
Reviewer
EPPI-
Systematic Review Data
Repository (SRDR)
DistillerSR (Evidence Partners)
REDCap
For small- , medium- and
especially large- scale reviews
(>20 studies), as well as reviews
that need constant updating
All team sizes, especially large
teams (i.e.>6data extractors)
Low to medium (opentools such as SRDR, or tools for
which authors have institutional
licences or, for Cochrane
Reviews, tools for which
Cochrane has obtained licences
on behalf of review authors)
High (commercial data systems
with no access via an
institutional licence)
Specifically designed for data
collection for systematic reviews
Allow online data storage, linking
and sharing
Easy to expand or edit forms as
required
Can be integrated with title/
abstract, full- text screening and
other functions
Can link data items to locations
in the report to facilitate
checking
Can readily automate data
comparison between
independent data collection for
the same study
Allow easy monitoring of
progress and performance of the
author team
access
(Continued)
153

7 Collecting data
https://t.me/medicina_free
Table7.4.a (Continued)
Paper forms Electronic forms Data systems
Disadvantages Inefficient and
potentially
unreliable because
data must be
entered into
software for
analysis and
reporting
Susceptible to
errors
Data collected by
multiple authors
must be manually
collated
Difficult to amend
as the review
progresses
If the papers are
lost, all data will
need to be
recreated
Require familiarity with
software packages to
design and use forms
Carry risk of
introducing mistakes in
data entering or
copy- pasting between
versions
Possible data security
issue if files are not
encrypted or securely
backed up (especially
relevant for IPD)
Facilitate coordination among
data collectors such as allocation
of studies for collection and
monitoring team progress
Allow simultaneous data entry by
multiple authors
Can export data directly to
analysis software
Can import included and
excluded studies, with reasons
for exclusion, directly into
RevMan (Cochrane Reviews only)
In some cases, improve public
accessibility through open data
sharing
front and considerable
Upinvestment of resources to set up
and adapt the form for studies of
test accuracy
Requires training of data
extractors
Cannot extract 2×2 contingency
tables for accuracy data
Structured templates may not be
as flexible as electronic forms
Cost of commercial data systems
Require familiarity with data
systems
Susceptible to changes in
software versions
7.4.3 Design ofa data collection form
Regardless of whether data are collected using a paper or electronic form or a data system, the key to successful data collection is to construct easy- to- use forms and collect
sufficient and unambiguous data that faithfully represent the source in a structured
and organized manner (Li 2015). In most cases, a document format should be developed for the form before building an electronic form or a data system. This can be distributed to others, including programmers and data analysts, and used as a guide for
154
Соседние файлы в папке Библиотека им академика М.И. Перельмана
