Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_112_библиотеки_им_акад_М_И_Перельмана
.pdf
the case could be evaluated based on typed responses and
oral presentation by the examinee. These same response
formats could be used to evaluate apparently distinct
competencies. If scores across competencies within response
format are more highly correlated than scores across formats
within competencies, this would be a matter of concern.
One particularly problematic source of construct-irrelevant
variance is systematic bias. Random errors—of the type we
typically consider when assessing the generalizability of
scores—tend to average out across items or judges.
Systematic effects can be more problematic than random
effects because they do not tend to sum to zero. This creates
systematic error or bias. Aspects of the test format or
administration can create this type of effect. For example,
when time limits impact test scores, the effect is likely to vary
across examinees. Examinees who need more time to
respond to one item are more likely to need more time to
respond to other items, and the effect accumulates as they
move through the test. Of particular concern is that these
effects might differentially impact distinct groups of
examinees, such as non-native English speakers. Systematic
effects also can impact scores when raters (e.g., standardized
patients) allow characteristics such as examinee race,
ethnicity, gender, or native language to impact scoring.
Recently, considerable attention has been given to the
potential for scoring systems driven by artificial intelligence
(AI) to display bias because the data samples used to train
the system do not represent the overall population
variability.
43
https://t.me/med1917

Example 1: A Multiple-Choice Examination
Tests of this sort typically assess a defined domain of
interest. Extrapolation of test scores to performance in
practice (or readiness for advancement in training) requires
that the content of the test is matched appropriately to the
demands of practice. Evidence for the content validity of the
test will follow from the procedures used to define the
domain and sample from it in assembling test forms. A job
(or practice) analysis may be used to collect information
about the requirements of practice, and additional studies
may include collecting expert judgments about the relevance
of items on actual test forms.
44
Criterion-related evidence is conceptually central to the
extrapolation stage of the validity argument. Certainly, it
would be desirable to demonstrate that scores from a
licensing examination were directly related to the learner’s
subsequent delivery of safe and effective treatment in
practice. While some researchers have been successful in
collecting this type of evidence, in general results of this sort
have been limited. One reason for this is the lack of valid
measures of the criterion of interest. For example, numerous
studies have shown that learners with better performance on
licensing examinations have a lower probability of being
sanctioned by state medical boards.45 These results generally
are supportive of the use of these tests as part of licensure,
but the tests are designed to measure medical knowledge or
clinical judgment and the criterion measure is at best a very
approximate measure of these competencies. Other
researchers
46,47
have shown a relationship between test scores
https://t.me/med1917

and patient outcomes or adherence to practice guidelines,
but in general these criteria have been limited in scope and
are somewhat removed from the actual competency the test
is designed to measure.
48,49
In the case of licensing
examinations, another limiting factor is the fact that
examinees who fail are not able to practice, making it
impossible to collect criterion measures. This is not to
suggest that studies based on such criterion measures should
not be pursued, but in the end, a more compelling argument
may rest on less direct evidence demonstrating that the
content of the examination reasonably represents the
construct of interest and that scores are not unduly
influenced by sources of construct-irrelevant variance.
Because logistic constraints necessitate administering highstakes multiple-choice examinations within structured time
limits, one potentially important source of constructirrelevant variance with such tests is the impact of these time
limits on outcomes (often referred to as speededness). It may
be (and often is) the case that the ability to respond quickly is
not a part of the construct of interest and is not consistent
with the intended score interpretations.†† The effects of
speededness are another example of a potential source of
construct-irrelevant variance.
50
Example 2: Performance Assessment
The primary attraction of performance-based assessment
formats is that they have the potential to more directly
measure constructs of interest; weakening the generalization
argument may be considered acceptable because the
https://t.me/med1917

extrapolation argument is strengthened. However, even
though simulation formats may be of high fidelity, there are
likely to be aspects that are artificial. There has been
relatively little research into the degree to which interactions
with standardized patients differ from interactions with real
patients, but it seems highly likely that differences exist.
Even when standardized patients appear to be
indistinguishable from actual patients, factors such as the
choice of scoring approach may impact the extent to which
the scores can be extrapolated to the performance of interest
in practice. Checklists, for example, may fail to capture more
subtle interviewing skills that facilitate information
gathering. Similarly, knowledge that the interaction is being
scored based on a checklist may alter an examinee’s
approach to interviewing in order to maximize score points.
The previous comments are intended to highlight the fact
that the appearance of similarity between the assessment
setting and the practice setting is not in and of itself validity
evidence. Using an assessment task that closely
approximates the practice setting has the potential to limit
the effects of construct-irrelevant variance and construct
underrepresentation, but this similarity does not ensure that
the score appropriately represents the competency of
interest.
Example 3: Workplace-Based Assessment
As with performance-based assessment formats such as
those using standardized patients, direct observation is an
attractive assessment approach because it has the potential to
https://t.me/med1917

strengthen the extrapolation stage of the validity argument.
Because observations are done in the practice setting,
differences between the features of the assessment and those
of practice may be minimized or eliminated. This
characteristic may facilitate construction of an assessment
that directly relates to real-world performance, but again it
does not in and of itself make the argument for
extrapolation. The act of observing may alter the
environment. More importantly, the scoring algorithm will
shape what is observed and how that observation is
transformed into a score. Because it is the score and not the
setting that is of interest, collecting observations in the
practice setting does not ensure the elimination of constructirrelevant variance or construct underrepresentation.
In the absence of highly structured scoring algorithms and/or
careful training, assessments based on direct observation
may be particularly susceptible to halo effects
51
and other
sources of construct-irrelevant variance. Unfortunately, in an
effort to more clearly define the behaviors to be assessed and
avoid such effects, the focus of the assessment may shift from
the construct of interest to a set of more easily defined
behaviors. In an effort to avoid the effect of constructirrelevant variance, the scores may suffer from construct
underrepresentation. For example, the complex concept of
physician–patient communication may be reduced to a set of
descriptions, such as “asks open-ended questions” and
“makes eye contact.” This may leave out important aspects
of the competency such as tone of voice or expressing
compassion. When this happens, the knowledge, skills,
behaviors, and attitudes that are included in the assessed
https://t.me/med1917

competency will be a limited subset of those in the intended
competency. Because of these issues, it is important to
remember that even when the real-life behaviors of interest
are directly observed, the resulting scores will be a function
of the specifics of the instrument used to record the
observation.
Decision/Interpretation
The decision stage of the validity argument provides support
for the decision rules and theory-based interpretations that
are applied to test scores. The most common decision rules
will be simple pass/fail classifications based on a single cut
score, but conjunctive or partially compensatory rules are not
uncommon. Arguments supporting the reasonableness of
these rules will be needed if the score interpretations
associated with the resulting classification decisions are to be
considered credible.
Similarly, score interpretations based on psychological
theories about cognition, judgment, or decision-making will
only be as credible as the theories themselves. For example, if
a score is used to classify practitioners as experts or novices
based on their patterns of data collection in reaching a
diagnosis, the theory of expert judgment supporting scoring
would be critical; if the theory were shown to be flawed,
score interpretations would by extension be suspect.
Example 1: A Multiple-Choice Examination
When performance on multiple-choice tests is used to make a
https://t.me/med1917

decision about eligibility for licensure or certification, the
appropriateness of the cut score will be a critical part of any
validity argument supporting the interpretation that failing
candidates are likely to lack some competency that is
necessary for safe and competent practice. That said, it must
be remembered that standard setting decisions are policy
judgments; they are not scientifically verifiable. Given this
reality, Kane has argued that appropriate evidence to
support the use of a cut score will demonstrate that the
procedure used to establish the standard was appropriate.
52
Information about the choice of procedure, selection of
judges, and implementation of the procedure will be central.
The credibility of the decision rule is central to score
interpretation for high-stakes standardized tests, but this
does not reduce the potential importance of theory-based
assumptions. For example, the use of multiple-choice items
may be based on the theoretical assumption that the
knowledge and judgment required to respond to such items
represent a necessary prerequisite for decision-making in
practice. While high scores may not provide assurance of
good performance in practice (because many other factors
can have an influence), low scores on a well-designed test
may indicate sufficiently serious knowledge deficits that are
unlikely to allow for good performance in practice. These
connections represent a theory about clinical decisionmaking: if the theory is shown to be flawed, the validity of
the associated scores similarly will be undermined.
Example 2: Performance Assessment
https://t.me/med1917

Performance assessments (including standardized patient
based examinations) sometimes are used to make
classification decisions in medical schools or postgraduate
education; and in these situations, failing examinees may be
required to complete remedial training. When this is done,
the assessment takes on the characteristics of a placement
test because the test scores result in placing learners in either
the standard education track or a remedial program.
Evidence to support the decision rule(s) used in this setting
might include results demonstrating that learners classified
as requiring remediation will show differential improvement
when exposed to the remediation program. Alternatively,
evidence could be collected to demonstrate that learners so
identified have a significantly greater chance of succeeding
in future training if they complete the remediation program.
Scoring procedures for performance assessments also may be
based, either implicitly or explicitly, on theoretical
assumptions about how information is to be aggregated in
drawing conclusions about competence. Decisions will need
to be made about the relative value of thoroughness and
efficiency. Similarly, decisions may need to be made about
the importance of physical examination maneuvers. If the
practitioner will confirm both negative and positive results
with a diagnostic test, the theoretical basis for drawing
conclusions about the learner’s diagnostic ability based on
their use of a nondiscriminating physical examination
maneuver would be questionable at best. These comments
are not intended to advocate for or against specific
approaches to scoring such examinations; they are intended
to highlight the fact that the structure of the scoring
procedure ultimately rests on a theoretical view of the
https://t.me/med1917

diagnostic process, and the strength of that model limits the
extent to which scores can be interpreted with respect to the
learner’s diagnostic competency.
Example 3: Workplace-Based Assessment
As with the formats discussed previously, assessments based
on direct observation will depend on theoretical
assumptions. Assumptions about the nature of the construct
being assessed will dictate the choice of process as opposed
to product or outcome measures. Similarly, theories relating
to expert-novice differences or cognitive theories about the
nature of the medical diagnosis process—and, more broadly,
medical decision-making—may influence the data that are
collected, the way those data are aggregated, and the way
the resulting scores are interpreted.
WBAs often are the basis for feedback to a learner. In this
case, it may be that no explicit decision is made based on the
scores. In other instances, however, promotional or other
high-stakes decisions may be made based at least in part on
the results of these assessments. In this circumstance, an
implicit (if not explicit) cut score must exist. The argument
for this use of the scores will require evidence to support the
reasonableness of the cut score or, more broadly, the decision
process. The fact that the implicit cut score may be built into
the definition of the score scale rather than the result of a
separate standard-setting process does not reduce the
importance of this part of the validity argument. Such a
circumstance might exist when the observer must rate a
performance as “adequate” or “inadequate.” The definition
https://t.me/med1917

of the rating may make establishing a cut score unnecessary,
but evidence that the definition of “adequate” corresponds to
a skill level that is appropriate to support a specific decision
still is needed.
Consequential Validity and Program
Evaluation
In this chapter, the validity argument has been presented as
the accumulation of scientific evidence to evaluate the
credibility of intended score interpretations. We would,
however, be remiss not to discuss the broader understanding
of validity that has existed within the educational
measurement community for the past 50 years.
Cronbach,
10
Messick,8 and Kane7 all emphasized the
importance of what has come to be known as consequential
validity. There has been debate over the years about whether
the consequences associated with a testing program belong
within the definition of validity. In the end, it may not matter
whether evaluation of the consequences of a testing program
are considered part of validity. What matters is the clear
understanding that just as the test developer has a
responsibility to evaluate the evidence that supports
inferences made based on test scores, the test developer has a
responsibility to evaluate the intended and unintended
consequences of the testing program. Whether implemented
within the classroom or on a national or international level,
testing has consequences; programmatic review of the
positive and negative consequences therefore is an important
https://t.me/med1917
Соседние файлы в папке Библиотека им академика М.И. Перельмана
