Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_112_библиотеки_им_акад_М_И_Перельмана
.pdf
Scoring
1. Were the observations made or stimulus
materials administered under standardized
conditions?
2. Were the scores recorded accurately?
3. Were the scoring algorithms applied
correctly?
4. Were appropriate security procedures
implemented?
Generalization1. What are the sources of measurement
error that contribute to the observed scores
on the assessment?
2. How similar would scores be across
replications of the measurement procedure?
3. How similar would classification
decisions be across replications of the
measurement procedure?
4. To what extent are test forms constructed
using a systematic process?
Extrapolation 1. To what extent do the scores correspond
to real-world competencies of interest?
2. Are there factors that interfere with
assessment of the competencies of interest?
3. Do scores predict real-world outcomes of
interest?
4. Are there artificial aspects of the testing
conditions that impact the scores?
Decision 1. Was the standard established through
implementation of a defensible and
properly implemented procedure?
2. Do examinees identified for remediation
https://t.me/med1917

more from a remediation program than
would those who were not identified?
Scoring
The scoring component of the validity argument must
provide evidence that assessment data have been collected
appropriately and scored accurately. This will include
consideration of a variety of types of evidence, such as the
extent to which the stated conditions of standardization have
been implemented and the accuracy of the scoring process.
As is true for each component of the validity argument, the
specifics of the evidence that will be relevant to the scoring
aspect of the argument will vary with the characteristics of
the assessment. Again, it is important to remember that any
argument for the validity of a score interpretation is only as
strong as the weakest link in that argument!
Example 1: A Multiple-Choice Examination
Standardized tests have been developed to provide the
strongest possible evidence for the scoring component of the
validity argument. The multiple-choice format was
developed in 1915 explicitly to support objective scoring.
21
Adherence to the conditions of standardization ensures that
the data are collected in the same manner for all examinees.
Factors such as the time allowed for the examination, the
seating, the lighting, and the quality of the stimulus
materials are controlled. To the extent that administration
procedures require documentation of the violation of these
conditions and annotation of score reports, the score user
https://t.me/med1917

will have confidence in the conditions under which the test
responses have been collected. Similarly, professionally
administered and scored tests routinely will have quality
control steps built into the scoring process. Key validation—
statistical analyses of examinee responses designed to verify
that the keyed answer is correct—provides evidence that the
scoring rules have been applied accurately. This step
includes activities such as (1) examining the proportion of
examinees receiving credit for each item; and (2) comparing
the probability of a correct response for examinees at
different score levels.
A fundamental consideration for high-stakes
†
tests is
security. In high-stakes settings, examinees may be
motivated to cheat and may attempt to do so in any number
of ways. When items are reused from one administration to
another, it is possible for examinees testing earlier to steal
(i.e., remember, copy, photograph) items and make them
available to those testing on a later date. When computerized
examinations are administered on a continuous basis, this
threat to validity may be increased. Evidence about the size
of the item pool and the frequency with which items are
reused will support the user’s confidence that prior exposure
has not threatened the integrity of the score. For computerbased tests, encryption of test items at all times except when
they are displayed on the screen may provide additional
confidence in the security of the test material. Tests that are
administered nationally or internationally are particularly
likely to be targeted by individuals or groups interested in
breaching security, but the same issues apply to tests
developed and administered within a single medical school
https://t.me/med1917

if items (or entire test forms) are used on multiple occasions.
Although these steps appear to fall under the heading of test
development or test administration, they also are critical
pieces of evidence to support the validity argument for
standardized written examinations. Because they are a
critical part of the validity argument, it is important that the
steps in administration and scoring are verified and
documented; it is not reasonable to assume that this is a
given, even with professionally developed assessments.
Example 2: Performance Assessment
The reproducibility of the stimulus material and scoring
procedures is, as previously noted, a strength of
standardized tests comprising multiple-choice items.
Relatively little effort is required to be satisfied that two
examinees assigned to the same test form but sitting at
different computers are seeing the same items and that those
items are being scored in the same way. The same is not
necessarily true for performance assessments such as
standardized patient–based assessments or other formats
that require humans to present and/or score the assessment.
Adding the human element creates the possibility that two
standardized patients trained to portray the same scenario
may perform in a less-than-standardized manner; the same
standardized patient may not portray the same scenario in
the same way on two different occasions. The scoring phase
of the validity argument will need to include evidence that
standardized patients are trained to an acceptable standard,
and it also will require evidence that standardized patients
https://t.me/med1917

are monitored over time to ensure both inter- and intrapatient consistency. Similar issues arise with scoring for
these tests; whether the scores are produced by a
standardized patient or content expert, it will be necessary to
assess the accuracy of the process. Again, this aspect of
testing must be verified before testing begins and must
continue to be monitored over time. It also is important to
remember that collecting evidence of a high level of rater
agreement during a small-scale pilot administration should
not replace collecting the same evidence once the test is
being administered operationally.
In addition to verifying that the overall error rate is low, both
in standardized patient portrayal and in the scoring process,
it will be important to provide evidence of a lack of
significant relationship between examinee characteristics and
standardized patients’ portrayal or scoring. For example, the
gender or ethnicity of a learner should have no impact on the
way the scenario is portrayed and scored. If data suggest that
learners of otherwise equal competence are likely to receive
better scores if they are, for example, male rather than
female, this would be a serious threat to valid score
interpretation. This type of effect is more serious than
random error in portrayal or scoring because random errors
tend to average out across encounters; systematic effects do
not.
Security issues also may be important with performance
assessments. If the test is used to make important decisions,
learners may attempt to improve their scores by gaining
prior access to test information. In most circumstances,
performance assessments (and standardized patient–based
https://t.me/med1917

tests in paticular) are administered on multiple occasions.
This creates the opportunity for learners who have
completed the examination to share information with others
who will test in the future; in most situations, prior
knowledge about the specific tasks that will appear on a test
should be expected to influence scores.22 This threat to
validity is analogous to the problem associated with the
reuse of material on tests comprising multiple-choice items,
but in the case of performance assessments it is much more
difficult to produce large banks of test “items.” When tests
are administered during a relatively short period,
sequestering examinees to prevent the sharing of
information may provide evidence that this threat to validity
has been controlled. With standardized patient–based tests,
an additional threat to security exists in that standardized
patients (or examiners, if they are used in test
administration) themselves may share information with
learners before or during the test administration.
Example 3: Workplace-Based Assessment
Assessment of a clinician or trainee through direct
observation involves taking another step away from the
completely standardized stimulus material of the multiplechoice examination and the partially standardized conditions
that exist in performance assessment; it takes place in the
largely uncontrolled conditions of the authentic clinical
environment. To support the interpretations of scores
produced in these settings, it will be necessary to produce
evidence that different assessors working in different
settings are, in fact, assessing the same construct in the same
https://t.me/med1917

way. One way to provide such evidence would be to
carefully define the characteristics of performance that will
be rated. A combination of careful specification of what is
being rated and thorough training of assessors may provide
reasonable support for the assertion that individuals are
being assessed on the same construct. The disadvantage of
carefully defining the assessment content is that it may
restrict what can be assessed to those aspects of the construct
that can be defined easily; this will have an impact on the
potential to extrapolate from the scores to the construct of
interest. The alternative may be that each assessor defines the
construct in their own way, but this approach clearly leaves
the scoring aspect of the validity argument seriously
weakened.
Even with careful attention to activities such as specifying
the rating elements and assessor training, it will be important
to collect evidence demonstrating that assessors are, in fact,
assessing the same constructs. In a paper reporting on a
WBA system implemented at 15 US residency programs, the
authors describe several activities intended to strengthen the
scoring component of the validity argument. In the early
stages of instrument development, cognitive interviews were
conducted with assessors in the roles that would be
participating in the actual assessment. This allowed for
collecting important information about how the items were
being interpreted by different groups of assessors and led to
targeted item revisions focused on decreasing the variability
in individual interpretation and completion of items. In
addition, as part of the instrument development process the
researchers often would include a brief introduction to the
item prior to the specific rating question. This was done in an
https://t.me/med1917

attempt to provide a clear foundation for responding to the
item and, by extension, to decrease the extent to which
individual assessors responded to items based on their own
interpretations of what question was being asked. The
following provides an example of one of these items. There is
an initial description of expectations for the person being
assessed:
As a member or leader of a clinical team, a resident is
counted on to keep interprofessional team members aware
of: (1) patients' current status and potential for deterioration;
and (2) any changes in status (e.g., physical exam, labs) that
occur.
And the scorable item follows:
Thinking about situations in which the resident would have been
expected to inform the team about patient status, indicate the
degree to which the resident kept you informed (rating scale:
continuum between did not keep me informed and kept me
fully informed).
Providing specific examples to help assessors focus on the
same construct when responding to assessment items is just
one approach to addressing scoring-related considerations in
WBA settings. Relevant considerations will differ based on
the assessment context and the inferences that are to be made
based on test scores. As such, a critical first step in the
process is to ensure an understanding of the various factors
that can impact the scoring component of the validity
argument. This will help to ensure that the necessary validity
https://t.me/med1917

evidence can be generated and collected.
As with other forms of assessment that do not lend
themselves to machine scoring, careful training of assessors
will be important with WBAs. That said, evidence to support
the scoring component of the validity argument likely will
include an evaluation of the accuracy of the scores rather
than simply documentation of careful training.
Generalization
This stage of the argument focuses on the question of
whether an examinee would receive the same score if the
assessment were repeated. We would not waste our time
standing on a scale if we discovered that the dial reported a
wildly different weight when we stepped off and back on.
Technically speaking, the generalization stage of the
argument is focused on the relationship between the
observed scores and the associated universe scores (or true
scores). Both universe scores and true scores are
conceptualizations. The universe score represents the score
that an examinee would receive if it were possible for that
examinee to respond to all items representing the universe of
acceptable observations (i.e., if the examinee responded to all
items in the domain). The true score is a closely related
concept representing the mean score that the examinee
would receive if they completed an unlimited number of
randomly equivalent (parallel) forms of the test. (The
observed score is the actual score that is recorded when an
examinee completes a specific test form.) The details of these
definitions and related theories are beyond the scope of this
https://t.me/med1917

chapter; the interested reader is referred to Gulliksen1 and
Lord and Novick
23
for a detailed discussion of classical test
theory and to Cronbach and associates24 and Brennan25 for
discussions of generalizability theory.
Two kinds of evidence are required for this stage of the
argument. First, it is necessary to show that the sample of
items or observations made of the examinee are
representative of the domain to which the score is to be
generalized. Second, it is necessary to demonstrate that the
sampling is sufficiently extensive to prevent the observed
scores from being unduly influenced by sampling error. The
extent to which the sample is representative will depend on
the procedures used for test construction (a
blueprint/sampling plan for data collection for WBAs); the
adequacy of the sampling can be examined directly through
a well-developed set of theory-based statistical procedures.
The samples will be representative to the extent that data
collection follows specified rules. In some cases, random
selection from a specified domain will be appropriate; in
others, stratified sampling will be preferred. In some
contexts, rules for the range of conditions under which
observations may (or must) be made will replace the
sampling of stimulus material.
The most developed aspect of test theory by far relates to
evaluation of reliability; conceptually this methodology is
designed to assess the relationship between observed scores
and true scores or universe scores. The most common index
of this relationship is the reliability coefficient; this coefficient
represents the correlation between the observed test scores
https://t.me/med1917
Соседние файлы в папке Библиотека им академика М.И. Перельмана
