Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_2823_Библиотеки_им_академика_М_И_Перельмана
.pdf
reproductive
disorders
Respiratory
disorders
8 9 4 4
Total 70 60 30 40
a
Items related to infectious and neoplastic diseases are
included in the affected organ system.
When items are sampled from fixed content categories, the
analysis becomes more complicated. In this situation, there
are variance components for persons (p); content categories
(c); items nested in content categories (i:c); persons by
content categories (p × c); and persons by items nested in
content categories (p × i:c). In this case, the c component will
not contribute measurement error because this structure is
fixed across test forms. Similarly, because the categories are
fixed, the p × c variance component will contribute to
universe variance. The p × i:c component will contribute to
error, and when comparisons are made across forms the i:c
will contribute to measurement error. The impact of this
latter component again will be mitigated to the extent that
test forms are constructed or equated to be statistically
equivalent. The stratification process typically will yield a
smaller standard error and larger generalizability coefficient
than analysis without stratification; this is one reason that
coefficient alpha is referred to as a lower bound estimate of
reliability. It should be noted, however, that in practice the
difference in coefficients typically is modest.
The error variance estimates produced using generalizability
https://t.me/med1917

theory provide a basis for estimating the standard error of
measurement for the test; these are useful for providing
confidence intervals around scores. Generalizability
coefficients also may be produced as the ratio of the universe
score variance divided by the sum of the universe score
variance and the error variance. Although these indices are
commonly reported, caution is required because they will be
sensitive to the specific sample of examinees used in the
estimation. Consider, for example, estimation of such an
index for one of the Steps of the United States Medical
Licensing Examination (USMLE); if the coefficient is
estimated based on the relatively homogeneous group of US
graduates taking the test for the first time, it may be several
points lower than if it is estimated based on all examinees
completing the test. In contrast, the standard error of
measurement tends to be more stable across groups, making
it a more interpretable and useful index of precision.
33
Example 2: Performance Assessment
The logic of the argument described in the previous example
holds in the context of performance assessments. To draw
conclusions from analyses based on a single administration
of the assessment, the rules employed in test construction
must guarantee that there will not be systematic differences
in test forms. The logistic realities of test delivery may make
this more difficult when the “items” are people, but clearly
generalization across test forms will be threatened if the
tasks (e.g., standardized patients) on one form are
systematically different than those on another. Important
differences could include changes in the types of problems
https://t.me/med1917

portrayed as well as changes in the level of experience and
training of the patients.
The generalizability of standardized tests comprising
multiple-choice items is relatively easy to evaluate, and even
the simpler classical test theory models provide adequate
tools for most situations. The complexity of performance
assessments, however, makes evaluation of the
generalizability of scores a more difficult matter. Consider a
test in which examinees rotate through a set of stations and
at each station they interact with a patient and complete a
post-encounter note. The notes then are scored by a group of
raters. When examinees complete the same set of stations
and notes are rated by the same set of raters, variance
components can be estimated for persons, stations, raters,
persons by stations, persons by raters, stations by raters, and
persons by stations by raters.**
The evaluator will need to
determine how much each component contributes to
measurement error in the specific context. Interaction terms
that include the person and station effect almost always will
contribute to measurement error—regardless of the intended
score interpretation—because the generalization argument is
about the extent to which the score from this test form is
comparable to the score from a similarly constructed test
form. By contrast, generalization over raters may or may not
be important. If the test is administered in a context in which
the same group of raters rates all examinees, and if there is
no intention to draw inferences about how the examinees
may have performed with other raters (which is rarely true),
then raters can be considered a fixed facet in the design. In
this case, the rater and station-by-rater variance components
will not contribute to measurement error and the person-by-
https://t.me/med1917

rater component will contribute to universe score variance.
In most circumstances, however, users of test scores will
want to draw inferences that extend beyond the group of
raters scoring an examinee’s performance, and these variance
components are best viewed as contributing to measurement
error (often substantially if the typical examinee is scored by
a small number of raters).
At this point, it should be clear that when a facet in the
design is considered fixed, the scores will have a smaller
error variance and a higher level of generalizability. The
evaluator may be tempted to try to increase the
generalizability of scores by considering facets fixed. This
strategy is without merit; it gives a promising answer to the
wrong question.
Example 3: Workplace-Based Assessment
When examinees are observed in a practice setting, the
generalization portion of the validity argument may be
problematic. Although there may be explicit rules controlling
the sampling of observations, the logistics of conducting a
WBA could make it likely that the environmental factors and
patient characteristics are more similar from one observation
to another within versus between examinees. This may lead
to an overly optimistic report on the generalizability of
scores. In this setting, the scores will be influenced by the
rater effect as well as an effect for the specific patient or task
that provides the context for the observation. Depending on
the design used to assign raters, it may be difficult to
accurately estimate a rater effect. It also may be difficult or
https://t.me/med1917

even impossible to fully differentiate between variance
associated with the difficulty of the patient’s presentation or
other characteristics of the task and the residual variance.
It usually is the case that the generalizability of scores will
decrease as the type of assessment changes from a highly
structured format—such as a professionally developed
multiple-choice test—to a performance assessment or
workplace-based observational assessment. There are two
reasons for this. First, it is possible to sample from the
domain of interest more widely and efficiently with
multiple-choice items because it takes relatively little time to
respond to them and they are inexpensive to score. Second,
both the sampling of content and the scoring can be more
highly standardized with multiple-choice assessments so
that the contribution of these factors to measurement error
can be markedly reduced. At the same time, multiple-choice
tests can assess only a limited range of skills, and it is
important for the assessment format to have a good match to
the skills to be assessed to avoid construct
underrepresentation.
The potential to sample more widely reduces the impact of
the examinee-by-item interaction as well as the effect of any
higher-order interaction terms (including residual variance).
There is a widely held view that the examinee-by-item
interaction term in the typical person-by-item design
represents “content specificity,” or the tendency for
physician knowledge to be highly problem specific. The
pervasive nature of the effect is well documented: the
examinee-by-item (or case) interaction term is routinely the
largest single source of error variance. It is, however, less
https://t.me/med1917

clear whether this term represents content specificity or other
sources of uncontrolled variability in the design. There is
relatively little research investigating how consistently
examinees respond to the same items or cases on different
occasions. To the extent that the effect of interest actually is
content specificity, examinees completing the same multiplechoice items or completing the same performance task on
multiple occasions would receive highly consistent scores.
There is some evidence from outside the domain of
assessment in medical education to suggest that scores may
not be highly reproducible across occasions. Similarly, there
is evidence that the generalizability of test scores can be
improved by building test forms to consistently sample from
fixed content categories; however, the absolute magnitude of
this improvement generally is small.
As noted previously, a second reason for the lower
generalizability of scores resulting from assessments using
performance tasks or workplace observation is that the
conditions of observation and scoring are more difficult to
standardize. This argues for increasing the structure of the
assessment, but this process requires careful thought. The
decision to implement a less structured assessment instead of
one that is more highly structured (e.g., a clinical rather than
a multiple-choice examination) is based on the perceived
need to more directly assess the construct of interest. The
problem lies in the fact that changing the scoring procedure
may increase the standardization of the assessment by
altering what is being assessed; the focus of the assessment
therefore may shift in the direction of competencies that are
more easily quantified and away from its original intent.
This is not to argue against making every effort to structure
https://t.me/med1917

the assessment; the key is to structure the assessment with a
careful eye on the intended interpretation of the scores.
Inevitably, it will be necessary to strike a balance between
the generalizability of scores and the extent to which one can
extrapolate from those scores to the actual competencies of
interest.
Additional Thoughts on Reliability
In the previous sections we have discussed generalization
(reliability) issues for three testing contexts. There are,
however, perspectives that have not been considered, some
of which relate to the purpose of the test. For example, the
same testing format (e.g., multiple-choice items, performance
assessment) may be used in different ways, such as
providing formative feedback to support learning or
informing summative decisions about an individual.
Similarly, a test may be used for low-stakes or high-stakes
purposes. Generally, it is reasonable to argue that highstakes decisions should be based on precise measurements; a
summative decision that may lead to granting/denying an
individual a license to practice requires a precise
measurement to protect both the public and the test taker.
This does not imply that low-stakes summative assessments
should be held to lower standards of precision. The required
level of precision is dependent on the inference/use that is to
be made of the test results. Medical professionals study and
train to develop expertise. Decades of research on expertise
have demonstrated that development of expertise is
facilitated by accurate, timely, and specific feedback.34 This
https://t.me/med1917

would seem to argue for the importance of precision in
formative feedback. Additionally, the need for reliability
may be increased—rather than decreased—when decisions
are made about relative strengths within rather than between
individuals. One reason for this is that the correlations
between competencies often are high. This, in part, explains
why so much attention continues to be given to how and
when we should report subscores,
35
but the problem has
been recognized for nearly a century.36 There is an extensive
literature on the issues of measuring differences between
correlated competencies and a similar literature on
measuring change for individuals.37 Both argue that the
associated problems are complex and the need for precision
may be substantially increased in these contexts.
In addition, if the primary purpose of an assessment system
is formative—to identify an individual’s areas of strength
and weakness and to provide feedback that motivates
trainees to remediate weaknesses, driving future learning
forward
38
—the criteria for evaluating the utility of the
assessment system logically should include its success in
improving learning outcomes over time.39 If assessment
results are to be used for both formative and summative
purposes—which often is true for assessments given during
training, and more recently, for continuing certification40—
then the reliability and validity considerations discussed in
this chapter also must be addressed. The next section
examines the extrapolation phase of the validity argument.
Extrapolation
https://t.me/med1917

The extrapolation phase of the validity argument focuses on
examining a link between the scores collected as part of the
assessment and performance in the real-world context of
interest. Assessors rarely are interested in knowing about a
learner’s ability to answer multiple-choice questions or, for
that matter, the learner’s ability to interact with standardized
patients. Instead, the interest is in competencies such as
knowledge base, problem-solving skills, clinical judgment,
and ability to communicate effectively. Assessment scores
provide indirect evidence about how the learner is likely to
perform in the context of interest; the extrapolation phase of
the validity argument is concerned with that evidence.
This is the most difficult stage of the validity argument
because the evidence is by nature inferential and the analytic
framework is less well developed than that for the
generalization argument. The extrapolation stage of the
argument is every bit as vital as the generalization phase. A
highly reliable score that measures the wrong characteristic
is of little value. However, it is equally important to
remember that the appearance that an examination measures
the competency of interest is not a substitute for actual
evidence supporing that assertion. Such “face validity” may
support the political acceptability and, perhaps, the legal
viability of an assessment,41 but it does not contribute
meaningfully to the validity argument.
As noted in the introduction to this chapter, the validity of
inferences from test scores cannot be reduced to a correlation
with a criterion measure because completely valid criteria
are rarely (if ever) available. Nonetheless, information about
the relationship between test scores and other relevant
https://t.me/med1917

measures will contribute to the argument. Similarly,
evidence about the content of the examination will be of
interest. Beyond these two types of supportive evidence, the
extrapolation argument must be guided by the quote from
Cronbach that began this chapter: “A proposition deserves
some degree of trust only when it has survived serious
attempts to falsify it.”42 The evaluator will be called upon to
assess both the extent to which scores are influenced by
sources of variability that are not related to the competency
of interest and the extent to which scores fail to reflect
important aspects of the competency of interest; these two
threats to validity are referred to as construct-irrelevant
variance and construct underrepresentation.
The assessment format itself may be one potential source of
construct-irrelevant variance. If a computerized test requires
an examinee to manage a patient in a simulated patient-care
environment (as does the computer-based case simulations
component of USMLE Step 3), facility with the user interface
may impact performance. Not only must examinees identify
the next step in management; they must take that step in a
potentially unfamiliar simulated world. Similarly, if the test
format requires examinees to interact with a simulated
electronic health record, scores may reflect the extent to
which the health record is similar to the system they use in
training or practice. The impact of such factors on test scores
would be considered construct-irrelevant variance.
One way to evaluate the extent to which construct-irrelevant
format effects are impacting scores is with the multitraitmultimethod matrix. For example, the ability to interview
and examine a patient and to describe the critical features of
https://t.me/med1917
Соседние файлы в папке Библиотека им академика М.И. Перельмана
