Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_2823_Библиотеки_им_академика_М_И_Перельмана
.pdf
from two equivalent forms of the test. The square root of this
value represents the correlation between observed scores and
true scores on the test. In the classical test theory framework,
the reliability coefficient also is directly related to the
standard error of measurement, which represents the
distribution of observed scores around a given true score.
Numerous approaches have been developed to estimate the
relationship between observed scores and true scores; the
usefulness of these procedures will depend on how one
conceptualizes the meaning of “a replication of the
measurement procedure.”26 Because the specific set of items
and the specific time and date on which the test was
administered rarely are central to how scores are to be
interpreted, it is generally desirable to view “replication” as
including measurements with different test forms on
different occasions. This common condition makes
correspondence between scores achieved on two forms of a
test on different occasions a single standard for assessing
replicability. The value of this standard rests on the
assumption that the characteristic to be measured has not
changed between administrations. This includes change in
the narrow sense of learning as well as change in relevant
conditions of observation such as motivation, fatigue, and
familiarity with the test format.
When it is unlikely that relevant conditions of testing remain
constant across occasions, it may be more appropriate to
conceptualize a replication such that occasion is held
constant. In the practical sense in which the test is
administered twice on the same day, replication on the same
occasion is open not only to the effects of fatigue but also to
https://t.me/med1917

the effects of practice leading to increased familiarity with
the format. In the literal sense of replication on the same
occasion (in which two forms are administered
simultaneously), these effects are absent but actual
replication is not possible; only a conceptual or theoretical
replication can exist.
For a test comprising multiple-choice items, the definition of
replication should include consideration of occasion and the
selection of items. For more complex testing formats, the
definition of replication similarly will be more complex.
Consider, for example, an essay examination in which the
stimulus will be standardized but a replication may involve
a different set of essay prompts to which an examinee will
respond on a different occasion. Additionally, the responses
may be scored by a different set of judges and the judges
may evaluate the material on different occasions. The
definition of a replication therefore will depend on which
features are considered fixed and which are considered
random. In this context, the definition of fixed and random
variables is guided by the desired interpretation of scores. If
score interpretation assumes that judgments were made by a
specific group of experts and that all examinees were
evaluated by the same experts, judges will be considered a
fixed facet in the design.‡ More commonly, it is appropriate
to view judges as having been sampled from a larger group
of similarly qualified and acceptable judges such that the
judges should be viewed as a random facet. Similarly, if
score interpretations assume a specific set of test items or
other stimulus materials, this facet is fixed; alternatively, if
the stimuli are viewed as sampled from a larger domain,
https://t.me/med1917

items must be viewed as a random facet. Random facets will
vary from one replication to the next; fixed facets will not.
Again, it is uncommon for tasks to be fixed, but in certain
situations, such as tests of specific procedural skills, this may
be the case.
The appropriate methodology for examining the strength of
the relationship between observed scores and true scores or
estimating the standard error of measurement will depend
on the complexity of the data collection design. When
practicality allows, actually repeating the measurement
procedure will provide a sound basis for assessing the
relationship of interest. The correlation between scores
produced across replications will provide an appropriate
estimate of the reliability of the test. Again, the square root of
this value will represent the correlation between observed
and true scores, and the well-known formula
provides an estimate of the standard error of measurement
in this situation. (In this formula, σx represents the standard
deviation of the observed scores, σe represents the standard
error, and r
xx′
represents the reliability of the test.)
In many circumstances, replication will not be practical; for
example, candidates for licensure cannot be called upon to
retest under the same high-stakes conditions after they have
completed and passed an examination. Numerous
procedures are available to evaluate the reliability
https://t.me/med1917

(reproducibility, precision) of test scores based on a single
examination administration. Nearly a century ago, Spearman
and Brown introduced the first of these procedures based on
the correlation between split halves (e.g., even- and oddnumbered items) of an examination.
4
,27
Kuder-Richardson
Formula 20 (KR-20)
28
and coefficient alpha29 estimate a value
equal to the average of all possible split halves. These
procedures provide estimates of reliability based on the
strength of relationship between items in a single test form.
They work on the assumption that the strength of
relationship (covariance) between item n and item m (n ≠ m)
on a single test form will provide a good approximation of
the strength of relationship between item “n” on test form 1
and item “m” on test form 2.
Coefficient alpha and the Kuder-Richardson formulas are
useful tools for collecting evidence about the generalization
of test scores. Unfortunately, they have become a kind of
knee-jerk reaction to the question of score reliability. Too
often, researchers appear to view estimating reliability as a
requirement that allows them to report a coefficient that a
journal editor will demand rather than as an opportunity to
better understand the characteristics of their assessment.
When applying these procedures, two important
considerations arise. First and foremost, the evaluator must
ask the question about what is meant by a replication in that
specific context. The described estimation procedures are
appropriate when generalization is viewed in terms of
replication across items (or test forms) with all other
conditions of measurement held constant. In the relatively
simple context of tests based on multiple-choice items, this
https://t.me/med1917

approach generally will underestimate the standard error of
measurement that would be observed for replications across
both test forms and testing occasions. For more complex
testing formats, interpretation of the results from these
procedures will be more difficult and often much more
problematic.
A second important consideration in applying these
procedures is related to the assumptions used in their
derivation. The central assumption in interpreting coefficient
alpha (or KR-20) is that, on average, the strength of
relationship between any two items on a single test form is
equal to that between any two items on different forms of the
test. When this assumption is violated, the results may
substantially misrepresent the actual reliability of the
assessment. Typically, this violation will result in an
overestimate of reliability. Consider, for example, the case in
which a passage describing a clinical scenario is followed by
several questions. It is common that the strength of
relationship between questions associated with the same
passage will be greater than the strength of relationship
between items from different passages. Because scenarios
typically will be different from one test form to the next, the
average relationship between items across test forms will be
best approximated by the relationship between items from
different scenarios on a single test.
Another example of a situation in which these procedures
may be misapplied occurs in assessments that require
multiple raters to assess an examinee’s performance on the
same task. For example, consider the circumstance in which
raters work in pairs to evaluate an examinee’s interaction
https://t.me/med1917

with a real patient. The assessment requires that each
examinee interact with five patients, and each interaction is
evaluated by a different pair of raters. If the raters score
separately, the examinee will receive 10 scores. If all
examinees have interacted with the same five patients and
have been scored by the single set of raters assigned to that
patient, the evaluator may be tempted to calculate coefficient
alpha based on this set of 10 scores. Again, however, because
the strength of relationship between scores from raters
evaluating performance with the same patient typically will
be greater than that between pairs evaluating performance
with different patients, this approach will not appropriately
approximate the strength of relationship between scores
from different tests. In this instance, the error in estimation
may be substantial (e.g., the estimated standard error may be
50% of the correct value) and may grossly overstate the
reproducibility of the scores.
Lee Cronbach’s coefficient alpha paper29 may be the most
cited paper in the history of educational measurement. As
we mentioned, for many researchers it has become a kind of
knee-jerk reaction to the measurement of reliability.
Unfortunately, the interpretation of this coefficient also has
been made mechanically, without consideration of the
context. Too often a value of .80 has been viewed as a cut
score to determine whether a test is reliable. Although there
may be instances in which this is a reasonable target, the
precision that is needed to make a specific inference based on
a test score will depend on the inference itself.
Classical test theory divides observed scores into two
components: true score and error. Because an examinee’s
https://t.me/med1917

true score is defined as uncorrelated with error, it follows
that observed-score variance is composed of true-score
variance and error variance. Generalizability theory expands
this framework to divide (partition) the overall variance into
multiple components. Consider as an example the simple
testing situation in which examinees respond to essay
prompts and the essays are scored by raters. To study the
generalizability of the results, a researcher collected data for
a group of examinees; all examinees responded to the same
prompts, and all responses were scored by the same raters.
In the framework of generalizability theory, essay prompts
and raters become distinguishable sources of error variance.
As in the classical test theory framework, it is possible to take
data from a single administration, estimate the reliability (or
generalizability) of the test, and project the expected
reliability of the test with differing numbers of essay
prompts. However, because generalizability theory provides
a means of making explicit the error contributed by
variability both in essay prompts and raters, this framework
makes it possible to further project how the reliability of the
test would change if the number of raters assessing
performance on each prompt also is varied.
Example 1: A Multiple-Choice Examination
The focus of the generalization stage of the validity argument
is on the extent to which scores will be comparable across
replications of the assessment procedure. In the context of
standardized multiple-choice–based assessments, the
interpretation of scores typically will require that they are
comparable across multiple test forms. For example, a
https://t.me/med1917

licensing or certifying examination would lose credibility if
the test forms to which examinees were assigned led to
widely varying scores.
Viewed from a generalizability theory framework, this part
of the argument will require several types of evidence. First,
it will be necessary to demonstrate that the sampling
procedure used for test construction supports the creation of
comparable test forms. The simplest case of the construction
of multiple forms would be based on random selection of
items from an available pool of acceptable items. This is
conceptually simple, but it is unusual for standardized tests.
A more common approach would be to select items to meet
the constraints of a table of specifications or test “blueprint.”
In this case, items may be randomly selected from each of a
number of content categories (see Table 2.2
for a hypothetical
200-item multiple-choice test in internal medicine). When
different item formats are included on the test, the table of
specifications may indicate the number of items that should
be included from each combination of format and category.
A common variation on this theme is to write items for a new
form of the test to meet the specifications of the previous
form. When systematic differences exist in the test
construction process across forms, estimation of the
correlation between scores on multiple forms based on
generalizability analysis of a single form will be
inappropriate. When systematic test construction procedures
are used, multiple-choice–based tests typically will have a
reasonably simple data collection design, examinees
typically will be the focus of the measurement procedure
(referred to as the object of measurement in generalizability
theory terminology), and the sampling of items will
https://t.me/med1917

represent a potential source of measurement error. With this
simple design, three sources of variance (referred to as
variance components) can be estimated: a person variance
component, which is conceptually equivalent to true-score
variance in classical test theory; an item variance component,
which represents the variability in item difficulty; and a
person-by-item variance component, which represents
residual variance not explained by the other two effects. The
person-by-item variance component divided by the number
of items will represent the error variance when comparisons
are being made between examinees who have completed the
same test form; when comparisons are being made between
examinees who have completed different test forms, the
definition of error variance is more complicated. If the test
forms are constructed through a process that approximates
random sampling from an undifferentiated item pool and
there is no formal procedure to adjust the scores for difficulty
differences, the appropriate error variance will be the sum of
the item variance component and the person-by-item
variance component both divided by the number of items.
When statistical equating§ procedures are used, the impact of
the item variance component will be reduced; because
equating is not likely to be error free, the error variance
estimate based on the person-by-item variance component
alone will represent a lower bound of the error variance
when forms are equated.
Table 2.2
Sample Blueprint for a 200-Item Multiple-Choice Test in Internal
Medicine
Number of Questions per Clinical Task
https://t.me/med1917

Disease
Category/Organ
System
a
Making
a
Diagnosis
Making
Therapeutic
Decisions
Preventing
Disease
Using
Diagnostic
Studies
Cardiovascular
disorders
10 9 5 6
Dermatologic
disorders
4 2 2 2
Endocrine
and metabolic
disorders
7 6 3 4
Gynecologic
disorders
3 3 2 2
Hematologic
disorders
3 3 1 3
Immunologic
disorders
3 3 2 2
Mental
disorders
4 3 1 2
Musculoskeletal
disorders
8 6 2 4
Neurologic
disorders
6 4 2 3
Nutritional
and digestive
disorders
8 9 4 4
Renal,
urinary, and
male
6 3 2 4
https://t.me/med1917
Соседние файлы в папке Библиотека им академика М.И. Перельмана
