Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_112_библиотеки_им_акад_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
05.09.2026
Размер:
18 Мб
Скачать
from two equivalent forms of the test. The square root of this value represents the correlation between observed scores and true scores on the test. In the classical test theory framework, the reliability coefficient also is directly related to the standard error of measurement, which represents the distribution of observed scores around a given true score.
Numerous approaches have been developed to estimate the relationship between observed scores and true scores; the usefulness of these procedures will depend on how one conceptualizes the meaning of “a replication of the measurement procedure.”26 Because the specific set of items and the specific time and date on which the test was administered rarely are central to how scores are to be interpreted, it is generally desirable to view “replication” as including measurements with different test forms on different occasions. This common condition makes correspondence between scores achieved on two forms of a test on different occasions a single standard for assessing replicability. The value of this standard rests on the assumption that the characteristic to be measured has not changed between administrations. This includes change in the narrow sense of learning as well as change in relevant conditions of observation such as motivation, fatigue, and familiarity with the test format.
When it is unlikely that relevant conditions of testing remain constant across occasions, it may be more appropriate to conceptualize a replication such that occasion is held constant. In the practical sense in which the test is administered twice on the same day, replication on the same occasion is open not only to the effects of fatigue but also to
https://t.me/med1917
the effects of practice leading to increased familiarity with the format. In the literal sense of replication on the same occasion (in which two forms are administered simultaneously), these effects are absent but actual replication is not possible; only a conceptual or theoretical replication can exist.
For a test comprising multiple-choice items, the definition of replication should include consideration of occasion and the selection of items. For more complex testing formats, the definition of replication similarly will be more complex. Consider, for example, an essay examination in which the stimulus will be standardized but a replication may involve a different set of essay prompts to which an examinee will respond on a different occasion. Additionally, the responses may be scored by a different set of judges and the judges may evaluate the material on different occasions. The definition of a replication therefore will depend on which features are considered fixed and which are considered random. In this context, the definition of fixed and random variables is guided by the desired interpretation of scores. If score interpretation assumes that judgments were made by a specific group of experts and that all examinees were evaluated by the same experts, judges will be considered a fixed facet in the design.‡ More commonly, it is appropriate to view judges as having been sampled from a larger group of similarly qualified and acceptable judges such that the judges should be viewed as a random facet. Similarly, if score interpretations assume a specific set of test items or other stimulus materials, this facet is fixed; alternatively, if the stimuli are viewed as sampled from a larger domain,
https://t.me/med1917
items must be viewed as a random facet. Random facets will vary from one replication to the next; fixed facets will not. Again, it is uncommon for tasks to be fixed, but in certain situations, such as tests of specific procedural skills, this may be the case.
The appropriate methodology for examining the strength of the relationship between observed scores and true scores or estimating the standard error of measurement will depend on the complexity of the data collection design. When practicality allows, actually repeating the measurement procedure will provide a sound basis for assessing the relationship of interest. The correlation between scores produced across replications will provide an appropriate estimate of the reliability of the test. Again, the square root of this value will represent the correlation between observed and true scores, and the well-known formula
provides an estimate of the standard error of measurement in this situation. (In this formula, σx represents the standard
deviation of the observed scores, σe represents the standard error, and r
xx
represents the reliability of the test.)
In many circumstances, replication will not be practical; for example, candidates for licensure cannot be called upon to retest under the same high-stakes conditions after they have completed and passed an examination. Numerous procedures are available to evaluate the reliability
https://t.me/med1917
(reproducibility, precision) of test scores based on a single examination administration. Nearly a century ago, Spearman and Brown introduced the first of these procedures based on the correlation between split halves (e.g., even- and odd­numbered items) of an examination.
4
,27
Kuder-Richardson
Formula 20 (KR-20)
28
and coefficient alpha29 estimate a value equal to the average of all possible split halves. These procedures provide estimates of reliability based on the strength of relationship between items in a single test form. They work on the assumption that the strength of relationship (covariance) between item n and item m (n ≠ m) on a single test form will provide a good approximation of the strength of relationship between item “n” on test form 1 and item “m” on test form 2.
Coefficient alpha and the Kuder-Richardson formulas are useful tools for collecting evidence about the generalization of test scores. Unfortunately, they have become a kind of knee-jerk reaction to the question of score reliability. Too often, researchers appear to view estimating reliability as a requirement that allows them to report a coefficient that a journal editor will demand rather than as an opportunity to better understand the characteristics of their assessment. When applying these procedures, two important considerations arise. First and foremost, the evaluator must ask the question about what is meant by a replication in that specific context. The described estimation procedures are appropriate when generalization is viewed in terms of replication across items (or test forms) with all other conditions of measurement held constant. In the relatively simple context of tests based on multiple-choice items, this
https://t.me/med1917
approach generally will underestimate the standard error of measurement that would be observed for replications across both test forms and testing occasions. For more complex testing formats, interpretation of the results from these procedures will be more difficult and often much more problematic.
A second important consideration in applying these procedures is related to the assumptions used in their derivation. The central assumption in interpreting coefficient alpha (or KR-20) is that, on average, the strength of relationship between any two items on a single test form is equal to that between any two items on different forms of the test. When this assumption is violated, the results may substantially misrepresent the actual reliability of the assessment. Typically, this violation will result in an overestimate of reliability. Consider, for example, the case in which a passage describing a clinical scenario is followed by several questions. It is common that the strength of relationship between questions associated with the same passage will be greater than the strength of relationship between items from different passages. Because scenarios typically will be different from one test form to the next, the average relationship between items across test forms will be best approximated by the relationship between items from different scenarios on a single test.
Another example of a situation in which these procedures may be misapplied occurs in assessments that require multiple raters to assess an examinee’s performance on the same task. For example, consider the circumstance in which raters work in pairs to evaluate an examinee’s interaction
https://t.me/med1917
with a real patient. The assessment requires that each examinee interact with five patients, and each interaction is evaluated by a different pair of raters. If the raters score separately, the examinee will receive 10 scores. If all examinees have interacted with the same five patients and have been scored by the single set of raters assigned to that patient, the evaluator may be tempted to calculate coefficient alpha based on this set of 10 scores. Again, however, because the strength of relationship between scores from raters evaluating performance with the same patient typically will be greater than that between pairs evaluating performance with different patients, this approach will not appropriately approximate the strength of relationship between scores from different tests. In this instance, the error in estimation may be substantial (e.g., the estimated standard error may be 50% of the correct value) and may grossly overstate the reproducibility of the scores.
Lee Cronbach’s coefficient alpha paper29 may be the most cited paper in the history of educational measurement. As we mentioned, for many researchers it has become a kind of knee-jerk reaction to the measurement of reliability. Unfortunately, the interpretation of this coefficient also has been made mechanically, without consideration of the context. Too often a value of .80 has been viewed as a cut score to determine whether a test is reliable. Although there may be instances in which this is a reasonable target, the precision that is needed to make a specific inference based on a test score will depend on the inference itself.
Classical test theory divides observed scores into two components: true score and error. Because an examinee’s
https://t.me/med1917
true score is defined as uncorrelated with error, it follows that observed-score variance is composed of true-score variance and error variance. Generalizability theory expands this framework to divide (partition) the overall variance into multiple components. Consider as an example the simple testing situation in which examinees respond to essay prompts and the essays are scored by raters. To study the generalizability of the results, a researcher collected data for a group of examinees; all examinees responded to the same prompts, and all responses were scored by the same raters. In the framework of generalizability theory, essay prompts and raters become distinguishable sources of error variance. As in the classical test theory framework, it is possible to take data from a single administration, estimate the reliability (or generalizability) of the test, and project the expected reliability of the test with differing numbers of essay prompts. However, because generalizability theory provides a means of making explicit the error contributed by variability both in essay prompts and raters, this framework makes it possible to further project how the reliability of the test would change if the number of raters assessing performance on each prompt also is varied.
Example 1: A Multiple-Choice Examination
The focus of the generalization stage of the validity argument is on the extent to which scores will be comparable across replications of the assessment procedure. In the context of standardized multiple-choice–based assessments, the interpretation of scores typically will require that they are comparable across multiple test forms. For example, a
https://t.me/med1917
licensing or certifying examination would lose credibility if the test forms to which examinees were assigned led to widely varying scores.
Viewed from a generalizability theory framework, this part of the argument will require several types of evidence. First, it will be necessary to demonstrate that the sampling procedure used for test construction supports the creation of comparable test forms. The simplest case of the construction of multiple forms would be based on random selection of items from an available pool of acceptable items. This is conceptually simple, but it is unusual for standardized tests. A more common approach would be to select items to meet the constraints of a table of specifications or test “blueprint.” In this case, items may be randomly selected from each of a number of content categories (see Table 2.2
for a hypothetical 200-item multiple-choice test in internal medicine). When different item formats are included on the test, the table of specifications may indicate the number of items that should be included from each combination of format and category. A common variation on this theme is to write items for a new form of the test to meet the specifications of the previous form. When systematic differences exist in the test construction process across forms, estimation of the correlation between scores on multiple forms based on generalizability analysis of a single form will be inappropriate. When systematic test construction procedures are used, multiple-choice–based tests typically will have a reasonably simple data collection design, examinees typically will be the focus of the measurement procedure (referred to as the object of measurement in generalizability theory terminology), and the sampling of items will
https://t.me/med1917
represent a potential source of measurement error. With this simple design, three sources of variance (referred to as variance components) can be estimated: a person variance component, which is conceptually equivalent to true-score variance in classical test theory; an item variance component, which represents the variability in item difficulty; and a person-by-item variance component, which represents residual variance not explained by the other two effects. The person-by-item variance component divided by the number of items will represent the error variance when comparisons are being made between examinees who have completed the same test form; when comparisons are being made between examinees who have completed different test forms, the definition of error variance is more complicated. If the test forms are constructed through a process that approximates random sampling from an undifferentiated item pool and there is no formal procedure to adjust the scores for difficulty differences, the appropriate error variance will be the sum of the item variance component and the person-by-item variance component both divided by the number of items. When statistical equating§ procedures are used, the impact of the item variance component will be reduced; because equating is not likely to be error free, the error variance estimate based on the person-by-item variance component alone will represent a lower bound of the error variance when forms are equated.
Table 2.2
Sample Blueprint for a 200-Item Multiple-Choice Test in Internal
Medicine
Number of Questions per Clinical Task
https://t.me/med1917
Disease
Category/Organ
System
a
Making
a
Diagnosis
Making Therapeutic Decisions
Preventing
Disease
Using
Diagnostic
Studies
Cardiovascular disorders
10 9 5 6
Dermatologic disorders
4 2 2 2
Endocrine and metabolic disorders
7 6 3 4
Gynecologic disorders
3 3 2 2
Hematologic disorders
3 3 1 3
Immunologic disorders
3 3 2 2
Mental disorders
4 3 1 2
Musculoskeletal disorders
8 6 2 4
Neurologic disorders
6 4 2 3
Nutritional and digestive disorders
8 9 4 4
Renal, urinary, and male
6 3 2 4
https://t.me/med1917