Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_112_библиотеки_им_акад_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
05.09.2026
Размер:
18 Мб
Скачать
reproductive disorders
Respiratory disorders
8 9 4 4
Total 70 60 30 40
a
Items related to infectious and neoplastic diseases are
included in the affected organ system.
When items are sampled from fixed content categories, the analysis becomes more complicated. In this situation, there are variance components for persons (p); content categories (c); items nested in content categories (i:c); persons by content categories (p × c); and persons by items nested in content categories (p × i:c). In this case, the c component will not contribute measurement error because this structure is fixed across test forms. Similarly, because the categories are fixed, the p × c variance component will contribute to universe variance. The p × i:c component will contribute to error, and when comparisons are made across forms the i:c will contribute to measurement error. The impact of this latter component again will be mitigated to the extent that test forms are constructed or equated to be statistically equivalent. The stratification process typically will yield a smaller standard error and larger generalizability coefficient than analysis without stratification; this is one reason that coefficient alpha is referred to as a lower bound estimate of reliability. It should be noted, however, that in practice the difference in coefficients typically is modest.
The error variance estimates produced using generalizability
https://t.me/med1917
theory provide a basis for estimating the standard error of measurement for the test; these are useful for providing confidence intervals around scores. Generalizability coefficients also may be produced as the ratio of the universe score variance divided by the sum of the universe score variance and the error variance. Although these indices are commonly reported, caution is required because they will be sensitive to the specific sample of examinees used in the estimation. Consider, for example, estimation of such an index for one of the Steps of the United States Medical Licensing Examination (USMLE); if the coefficient is estimated based on the relatively homogeneous group of US graduates taking the test for the first time, it may be several points lower than if it is estimated based on all examinees completing the test. In contrast, the standard error of measurement tends to be more stable across groups, making it a more interpretable and useful index of precision.
33
Example 2: Performance Assessment
The logic of the argument described in the previous example holds in the context of performance assessments. To draw conclusions from analyses based on a single administration of the assessment, the rules employed in test construction must guarantee that there will not be systematic differences in test forms. The logistic realities of test delivery may make this more difficult when the “items” are people, but clearly generalization across test forms will be threatened if the tasks (e.g., standardized patients) on one form are systematically different than those on another. Important differences could include changes in the types of problems
https://t.me/med1917
portrayed as well as changes in the level of experience and training of the patients.
The generalizability of standardized tests comprising multiple-choice items is relatively easy to evaluate, and even the simpler classical test theory models provide adequate tools for most situations. The complexity of performance assessments, however, makes evaluation of the generalizability of scores a more difficult matter. Consider a test in which examinees rotate through a set of stations and at each station they interact with a patient and complete a post-encounter note. The notes then are scored by a group of raters. When examinees complete the same set of stations and notes are rated by the same set of raters, variance components can be estimated for persons, stations, raters, persons by stations, persons by raters, stations by raters, and persons by stations by raters.**
The evaluator will need to determine how much each component contributes to measurement error in the specific context. Interaction terms that include the person and station effect almost always will contribute to measurement error—regardless of the intended score interpretation—because the generalization argument is about the extent to which the score from this test form is comparable to the score from a similarly constructed test form. By contrast, generalization over raters may or may not be important. If the test is administered in a context in which the same group of raters rates all examinees, and if there is no intention to draw inferences about how the examinees may have performed with other raters (which is rarely true), then raters can be considered a fixed facet in the design. In this case, the rater and station-by-rater variance components will not contribute to measurement error and the person-by-
https://t.me/med1917
rater component will contribute to universe score variance. In most circumstances, however, users of test scores will want to draw inferences that extend beyond the group of raters scoring an examinee’s performance, and these variance components are best viewed as contributing to measurement error (often substantially if the typical examinee is scored by a small number of raters).
At this point, it should be clear that when a facet in the design is considered fixed, the scores will have a smaller error variance and a higher level of generalizability. The evaluator may be tempted to try to increase the generalizability of scores by considering facets fixed. This strategy is without merit; it gives a promising answer to the wrong question.
Example 3: Workplace-Based Assessment
When examinees are observed in a practice setting, the generalization portion of the validity argument may be problematic. Although there may be explicit rules controlling the sampling of observations, the logistics of conducting a WBA could make it likely that the environmental factors and patient characteristics are more similar from one observation to another within versus between examinees. This may lead to an overly optimistic report on the generalizability of scores. In this setting, the scores will be influenced by the rater effect as well as an effect for the specific patient or task that provides the context for the observation. Depending on the design used to assign raters, it may be difficult to accurately estimate a rater effect. It also may be difficult or
https://t.me/med1917
even impossible to fully differentiate between variance associated with the difficulty of the patient’s presentation or other characteristics of the task and the residual variance.
It usually is the case that the generalizability of scores will decrease as the type of assessment changes from a highly structured format—such as a professionally developed multiple-choice test—to a performance assessment or workplace-based observational assessment. There are two reasons for this. First, it is possible to sample from the domain of interest more widely and efficiently with multiple-choice items because it takes relatively little time to respond to them and they are inexpensive to score. Second, both the sampling of content and the scoring can be more highly standardized with multiple-choice assessments so that the contribution of these factors to measurement error can be markedly reduced. At the same time, multiple-choice tests can assess only a limited range of skills, and it is important for the assessment format to have a good match to the skills to be assessed to avoid construct underrepresentation.
The potential to sample more widely reduces the impact of the examinee-by-item interaction as well as the effect of any higher-order interaction terms (including residual variance). There is a widely held view that the examinee-by-item interaction term in the typical person-by-item design represents “content specificity,” or the tendency for physician knowledge to be highly problem specific. The pervasive nature of the effect is well documented: the examinee-by-item (or case) interaction term is routinely the largest single source of error variance. It is, however, less
https://t.me/med1917
clear whether this term represents content specificity or other sources of uncontrolled variability in the design. There is relatively little research investigating how consistently examinees respond to the same items or cases on different occasions. To the extent that the effect of interest actually is content specificity, examinees completing the same multiple­choice items or completing the same performance task on multiple occasions would receive highly consistent scores. There is some evidence from outside the domain of assessment in medical education to suggest that scores may not be highly reproducible across occasions. Similarly, there is evidence that the generalizability of test scores can be improved by building test forms to consistently sample from fixed content categories; however, the absolute magnitude of this improvement generally is small.
As noted previously, a second reason for the lower generalizability of scores resulting from assessments using performance tasks or workplace observation is that the conditions of observation and scoring are more difficult to standardize. This argues for increasing the structure of the assessment, but this process requires careful thought. The decision to implement a less structured assessment instead of one that is more highly structured (e.g., a clinical rather than a multiple-choice examination) is based on the perceived need to more directly assess the construct of interest. The problem lies in the fact that changing the scoring procedure may increase the standardization of the assessment by altering what is being assessed; the focus of the assessment therefore may shift in the direction of competencies that are more easily quantified and away from its original intent. This is not to argue against making every effort to structure
https://t.me/med1917
the assessment; the key is to structure the assessment with a careful eye on the intended interpretation of the scores. Inevitably, it will be necessary to strike a balance between the generalizability of scores and the extent to which one can extrapolate from those scores to the actual competencies of interest.
Additional Thoughts on Reliability
In the previous sections we have discussed generalization (reliability) issues for three testing contexts. There are, however, perspectives that have not been considered, some of which relate to the purpose of the test. For example, the same testing format (e.g., multiple-choice items, performance assessment) may be used in different ways, such as providing formative feedback to support learning or informing summative decisions about an individual. Similarly, a test may be used for low-stakes or high-stakes purposes. Generally, it is reasonable to argue that high­stakes decisions should be based on precise measurements; a summative decision that may lead to granting/denying an individual a license to practice requires a precise measurement to protect both the public and the test taker.
This does not imply that low-stakes summative assessments should be held to lower standards of precision. The required level of precision is dependent on the inference/use that is to be made of the test results. Medical professionals study and train to develop expertise. Decades of research on expertise have demonstrated that development of expertise is facilitated by accurate, timely, and specific feedback.34 This
https://t.me/med1917
would seem to argue for the importance of precision in formative feedback. Additionally, the need for reliability may be increased—rather than decreased—when decisions are made about relative strengths within rather than between individuals. One reason for this is that the correlations between competencies often are high. This, in part, explains why so much attention continues to be given to how and when we should report subscores,
35
but the problem has been recognized for nearly a century.36 There is an extensive literature on the issues of measuring differences between correlated competencies and a similar literature on measuring change for individuals.37 Both argue that the associated problems are complex and the need for precision may be substantially increased in these contexts.
In addition, if the primary purpose of an assessment system is formative—to identify an individual’s areas of strength and weakness and to provide feedback that motivates trainees to remediate weaknesses, driving future learning forward
38
—the criteria for evaluating the utility of the assessment system logically should include its success in improving learning outcomes over time.39 If assessment results are to be used for both formative and summative purposes—which often is true for assessments given during training, and more recently, for continuing certification40— then the reliability and validity considerations discussed in this chapter also must be addressed. The next section examines the extrapolation phase of the validity argument.
Extrapolation
https://t.me/med1917
The extrapolation phase of the validity argument focuses on examining a link between the scores collected as part of the assessment and performance in the real-world context of interest. Assessors rarely are interested in knowing about a learner’s ability to answer multiple-choice questions or, for that matter, the learner’s ability to interact with standardized patients. Instead, the interest is in competencies such as knowledge base, problem-solving skills, clinical judgment, and ability to communicate effectively. Assessment scores provide indirect evidence about how the learner is likely to perform in the context of interest; the extrapolation phase of the validity argument is concerned with that evidence.
This is the most difficult stage of the validity argument because the evidence is by nature inferential and the analytic framework is less well developed than that for the generalization argument. The extrapolation stage of the argument is every bit as vital as the generalization phase. A highly reliable score that measures the wrong characteristic is of little value. However, it is equally important to remember that the appearance that an examination measures the competency of interest is not a substitute for actual evidence supporing that assertion. Such “face validity” may support the political acceptability and, perhaps, the legal viability of an assessment,41 but it does not contribute meaningfully to the validity argument.
As noted in the introduction to this chapter, the validity of inferences from test scores cannot be reduced to a correlation with a criterion measure because completely valid criteria are rarely (if ever) available. Nonetheless, information about the relationship between test scores and other relevant
https://t.me/med1917
measures will contribute to the argument. Similarly, evidence about the content of the examination will be of interest. Beyond these two types of supportive evidence, the extrapolation argument must be guided by the quote from Cronbach that began this chapter: “A proposition deserves some degree of trust only when it has survived serious attempts to falsify it.”42 The evaluator will be called upon to assess both the extent to which scores are influenced by sources of variability that are not related to the competency of interest and the extent to which scores fail to reflect important aspects of the competency of interest; these two threats to validity are referred to as construct-irrelevant variance and construct underrepresentation.
The assessment format itself may be one potential source of construct-irrelevant variance. If a computerized test requires an examinee to manage a patient in a simulated patient-care environment (as does the computer-based case simulations component of USMLE Step 3), facility with the user interface may impact performance. Not only must examinees identify the next step in management; they must take that step in a potentially unfamiliar simulated world. Similarly, if the test format requires examinees to interact with a simulated electronic health record, scores may reflect the extent to which the health record is similar to the system they use in training or practice. The impact of such factors on test scores would be considered construct-irrelevant variance.
One way to evaluate the extent to which construct-irrelevant format effects are impacting scores is with the multitrait­multimethod matrix. For example, the ability to interview and examine a patient and to describe the critical features of
https://t.me/med1917