Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_2823_Библиотеки_им_академика_М_И_Перельмана
.pdf
of issues that are central to validity and reliability as these
concepts pertain to assessment in medical education.
Historical Context
Practically speaking, the history of test theory as we know it
began around the turn of the 20th century with Charles
Spearman. Spearman’s interest was in the study of
intelligence rather than assessment, and to support his work
he developed the field of correlational psychology. Most of
the basic equations from classical test theory were developed
by Spearman to aid his research on the presence of a
common (g) factor shared by most—if not all—tests of
mental proficiency.
1–4
This groundwork laid the foundation for a science of testing
that expanded explosively during World War I. The US
Army had a monumental personnel problem: tens of
thousands of recruits had to be placed in jobs. Testing
provided a potentially effective and efficient means of
determining appropriate job placements.5 This effort
established the practice of psychological testing in the United
States; not surprisingly, the science of testing was used in an
effort to boost educational and industrial efficiency after the
war. In both military and industrial contexts, the question of
interest was “How well do these tests predict performance
on the job?” Evidence to justify the use of the test naturally
conformed to the approach established by Spearman and
took the form of a correlation between the test scores and an
independent assessment of job performance.
https://t.me/med1917

The dramatic proliferation of placement testing did much to
define the view of validity during the period from 1920
through 1950. Correlational evidence, referred to as criterion
validity, was the standard during this period; in his 1951
chapter in the first edition of Educational Measurement,
Edward Cureton defined validity “in terms of the correlation
between the actual test scores and the ‘true’ criterion score.”
6
As a practical matter, criterion validity has obvious utility. In
placement testing, it has clear relevance to the interpretation
of the score and it provides an objective basis for comparing
multiple assessments available for a given purpose.
However, the strength of this approach is less apparent for
applications beyond placement testing. One problem is that
an obvious and practical criterion may not be available; no
clear and objective external criterion is likely to exist for an
achievement test. And if such a criterion is identified, the test
developer would need to provide validity evidence to
support its use.
7
Questions about the appropriateness of criterion validity as a
primary evaluation of assessments of academic achievement
led to the development of procedures for assessing content
validity. The purpose of such evidence is to establish that the
content of the test reasonably represents the domain of
interest. This type of evidence clearly is necessary, but it is
not sufficient to establish the validity of interpretations for
an achievement test. As Messick pointed out, evidence that
the test content is relevant to the domain(s) of interest
provides no direct support for inferences that are made
based on the test scores.
8
https://t.me/med1917

During the period after World War II, interest in personality
testing and the development of testing instruments to
standardize the process pushed researchers to continue
considering the types of evidence required to support the use
of these new instruments. Neither criterion nor content
validity models provided a particularly good fit to these
tests. It was in this context that Cronbach and Meehl
introduced the idea of construct validity.9 In the second
edition of Educational Measurement,10 Cronbach made the
following comment when describing the underlying
rationale for this new conceptualization of validity:
The rationale for construct validation (Cronbach and Meehl,
1955) developed out of personality testing. For a measure of,
for example, ego strength, there is no uniquely pertinent
criterion to predict, nor is there a domain of content to
sample. Rather, there is a theory that sketches out the
presumed nature of the trait. If the test score is a valid
manifestation of ego strength, so conceived, its relations to
other variables conform to the theoretical expectations.
This approach to validation greatly expanded the types of
evidence that could be considered when evaluating an
assessment. For example, in the context of achievement
testing, construct validation might argue for collecting
evidence to demonstrate that learners with advanced
training in the topic area outperform those with less training.
The 1950s brought two other important changes to the
conceptualization of validity. First, Campbell and Fiske
introduced the multitrait-multimethod matrix: a way to
organize validity evidence about the relative strengths of the
https://t.me/med1917

relationships between traits measured by a single method
and measures of the same trait using different methods.11 In
the context of personality testing, examples of traits may
have included extraversion and aggression; methods may
have included individual examiner-administered
assessments and group-administered paper-and-pencil
assessments. Campbell and Fiske’s matrix provided an
empirical means of assessing the extent to which scores are
impacted by otherwise irrelevant characteristics of the
assessment method or format (signaled by relatively higher
correlations between different traits measured by the same
method compared with the same trait measured by different
methods). This method effect relates to what was later to
become known as construct-irrelevant variance, a concept
that will be addressed in more detail later in this chapter.
The second important change in the conceptualization of
validity came when Loevinger focused attention on the
proposed interpretation of test scores.12 This represented an
important shift in perspective from consideration of the
relationship between the construct the test was designed to
measure and the test score to consideration of the
correspondence between what is measured by the test and
the proposed interpretations of the test score.
By the time of publication of the third edition of Educational
Measurement, Messick had developed a unified theory of
validity.8 Rather than being defined as “the correlation
between the actual test scores and the ‘true’ criterion score,”
6
validity now was viewed as the “. . . degree to which
empirical evidence and theoretical rationales support the
adequacy and appropriateness of interpretations and actions
https://t.me/med1917

based on test scores.”8 Messick’s model built on the
contributions of his predecessors; following Cronbach and
Meehl,9 Cronbach,10 and Loevinger,12 he emphasized the
need to specify the intended meaning and use of the test
score before validation. Consistent with Cronbach and Meehl
and Campbell and Fiske,11 Messick emphasized the
importance of considering alternative hypotheses such as the
impact of construct-irrelevant variance. Additionally, like his
predecessors, Messick argued that the process of validation
would involve an extended program of research.
*
Messick’s formulation is consistent with previous validity
frameworks, although his conceptualization introduces a
change in emphasis. In particular, he placed increased
emphasis on evaluating the consequences of the testing
program. He believed that both the actual and potential
social consequences of a test must be evaluated. Considering
as an example a test for medical licensure, at a minimum this
requirement leads to examination of consequences such as
the test’s impact on what instructors choose to teach and
what learners choose to learn. More broadly, consequential
validity would require consideration of the test’s impact on
the availability of medical practitioners both for the
community at large and, perhaps, specifically for
underserved communities. Messick’s view of consequential
validity went beyond these considerations; he additionally
required consideration of the impact that such an
examination might have on the entrance of minority
candidates into the profession. This broad definition of
consequential validity emphasizes the importance of test
developers and administrators accepting responsibility for
https://t.me/med1917

their actions. The definition takes the validation process
beyond the scientific evaluation of the assessment into the
arena of social and political values.
By 1999, the role of consequences within the sphere of
validity was sufficiently well established that it was included
as one of the five sources of validity evidence referenced in
the Standards for Educational and Psychological Testing.13 They
are presented here to provide continuity with previous
secondary sources describing Messick’s concept of validity
theory. It is, however, important to remember that both
Messick’s unified theory of validity and the Standards
emphasize that these are not different types of validity.
Rather, they are different sources of evidence, each of which
may be more or less important in providing support for a
specific score interpretation.
The history of validity theory should make it clear that the
definition of validity has expanded over time. The emphasis
also has changed as the focus of testing has changed.
Criterion validity (evidence based on relationships to other
variables) has not been replaced; this type of evidence
remains essential in evaluating admissions and employment
tests. Similarly, content validity represents an important
source of evidence in support of tests of achievement. The
history of validity is a history of both an expansion in
meaning and a shift in emphasis.
More recently, Kane has introduced an additional shift in
perspective by representing validity as an argument in
support of the proposed interpretations and uses of a test
score.
7,14
As with previous phases in the evolution of validity
https://t.me/med1917

theory, Kane’s view does not deny the importance of the
evidence and perspectives that have been discussed during
the past half century; it provides a shift in perspective rather
than a rejection of the basic arguments. That shift in
perspective does have one important characteristic: it
highlights the fact that the collection of evidence in support
of the interpretations of test scores must form a structured
and coherent argument that leads from the test
administration to the interpretation. That structured
argument is only as strong as its weakest component.
Kane’s View of Validity
Implicit in the interpretation of a test score is a series of
assertions and assumptions that support that interpretation.
For example, the interpretation of a passing score on a
medical licensing examination requires the assumption that
the test was administered under standardized conditions
and that the examinee did not have prior access to the test
material. If the examinee cheated, no interpretation can be
made about the score regardless of other characteristics of
the test. Interpretation of the test score requires assumptions
about the precision of the score; if the test score is not
reproducible, there is no basis for making an interpretation.
Interpretation of the score assumes that the test measures
some relevant aspect of the overall set of knowledge, skills,
and abilities required for the practice of medicine. It also
assumes that the cut score (pass/fail standard) has been
established in a way that supports the interpretation. If any
one of these assumptions is unfounded, the strength of the
others may be of little relevance.
https://t.me/med1917

Kane provides a structure for this validity argument that
outlines four links in the inferential chain from the test
administration to the final decision or interpretation.
7,14
He
labels these four components scoring, generalization,
extrapolation, and decision. Support for the scoring component
of the overall argument includes evidence that the test was
administered properly, examinee behavior was captured
correctly, and scoring rules were appropriate and were
applied accurately and consistently. The generalization
component of the argument requires evidence that the
observations were appropriately sampled from the universe
of test items, clinical encounters, et cetera. Generalization
also requires evidence that the sample of observations was
large enough to produce scores with an acceptable level of
precision. Broadly speaking, this stage in the argument asks
the question: Is the test reliable? In this context, generalization
refers to generalizing from the sample of behavior that was
part of the test (the observed score) to the test-taker’s true score
or universe score. The extrapolation component of the
argument requires evidence that the observations
represented by the test score were relevant to the target
competency or construct measured by the test. This requires
a demonstration that the observations were relevant to the
interpretation and that the scores were not unduly
influenced by sources of variance that are irrelevant to the
intended interpretation. Extrapolation also requires that the
target competency that the test is intended to measure is
reasonably well represented by the test score. The decision
component of the argument requires evidence in support of
any theoretical framework necessary for score interpretation
or evidence in support of decision rules. For tests with a cut
https://t.me/med1917

score, this evidence would include support for the procedure
used to establish that cut score. Again, the score user can
have confidence in an interpretation only if there is evidence
for each component of the overall argument. The types of
evidence required will vary with the purpose and
characteristics of the assessment.
These last two sentences are critically important and warrant
special comment because, since the last edition of this
volume, there has been considerable interest within medical
education in longitudinal assessment, formative assessment,
and other forms of assessment for learning. There also have
been numerous publications advocating the importance of
including a broad range of assessment that allows for
“triangulation.”
15,16
These publications sometimes have
drawn a distinction between tests designed to differentiate
among individuals and those designed to identify strengths
and weaknesses within an individual. They also have
described these changes as leading to a “postpsychometric”
era. To avoid confusion, we want to be clear that we believe
richer sources of evidence are valuable regardless of the
purpose of the assessment. At the same time, declaring that
something represents a rich source of evidence does not
make it so. Sources of feedback should be taken seriously
when (and only when) there is evidence to support the
validity of that feedback. This includes assessments that
result in numeric scores and those that produce written or
narrative feedback. Formative assessment of clinical
competence is intended to support the development of
expertise. That development requires practice; it also
requires accurate and timely feedback. We will return to this
https://t.me/med1917

issue throughout the chapter.
The next sections of the chapter will further explicate the
four components of the validity argument as described by
Kane. Within each section, content relevant to the particular
component will be provided for three types of assessments
that span a range of the types of assessments currently used
in medical education: a multiple-choice examination, a
performance assessment, and a workplace-based assessment
(WBA). Multiple-choice examinations are ubiquitous in
medical education from selection to medical school and
through classroom assessment to credentialing. Performance
assessments have a long history in medical education, with
objective structured clinical examinations (OSCEs) and
standardized-patient-based examinations being in common
use within medical schools and having a history as part of
licensure assessment.17 WBAs are becoming an increasingly
important part of assessment within medical education,
particularly during residency training.
18–20
Table 2.1 provides
examples of the kinds of questions that arise at each stage of
the validity argument. The questions are provided as
examples and are not intended to represent an exhaustive
list. As you read what follows, we encourage you to think
about the assessments that are presented, extrapolate to
other assessments, and add your own questions to this list.
Table 2.1
Examples of Questions Supporting Each of the Four
Components of Kane’s Argument-Based Approach to
Validity
Component Questions
https://t.me/med1917
Соседние файлы в папке Библиотека им академика М.И. Перельмана
