Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_2823_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
05.09.2026
Размер:
18 Мб
Скачать
the case could be evaluated based on typed responses and oral presentation by the examinee. These same response formats could be used to evaluate apparently distinct competencies. If scores across competencies within response format are more highly correlated than scores across formats within competencies, this would be a matter of concern.
One particularly problematic source of construct-irrelevant variance is systematic bias. Random errors—of the type we typically consider when assessing the generalizability of scores—tend to average out across items or judges. Systematic effects can be more problematic than random effects because they do not tend to sum to zero. This creates systematic error or bias. Aspects of the test format or administration can create this type of effect. For example, when time limits impact test scores, the effect is likely to vary across examinees. Examinees who need more time to respond to one item are more likely to need more time to respond to other items, and the effect accumulates as they move through the test. Of particular concern is that these effects might differentially impact distinct groups of examinees, such as non-native English speakers. Systematic effects also can impact scores when raters (e.g., standardized patients) allow characteristics such as examinee race, ethnicity, gender, or native language to impact scoring. Recently, considerable attention has been given to the potential for scoring systems driven by artificial intelligence (AI) to display bias because the data samples used to train the system do not represent the overall population variability.
43
https://t.me/med1917
Example 1: A Multiple-Choice Examination
Tests of this sort typically assess a defined domain of interest. Extrapolation of test scores to performance in practice (or readiness for advancement in training) requires that the content of the test is matched appropriately to the demands of practice. Evidence for the content validity of the test will follow from the procedures used to define the domain and sample from it in assembling test forms. A job (or practice) analysis may be used to collect information about the requirements of practice, and additional studies may include collecting expert judgments about the relevance of items on actual test forms.
44
Criterion-related evidence is conceptually central to the extrapolation stage of the validity argument. Certainly, it would be desirable to demonstrate that scores from a licensing examination were directly related to the learner’s subsequent delivery of safe and effective treatment in practice. While some researchers have been successful in collecting this type of evidence, in general results of this sort have been limited. One reason for this is the lack of valid measures of the criterion of interest. For example, numerous studies have shown that learners with better performance on licensing examinations have a lower probability of being sanctioned by state medical boards.45 These results generally are supportive of the use of these tests as part of licensure, but the tests are designed to measure medical knowledge or clinical judgment and the criterion measure is at best a very approximate measure of these competencies. Other researchers
46,47
have shown a relationship between test scores
https://t.me/med1917
and patient outcomes or adherence to practice guidelines, but in general these criteria have been limited in scope and are somewhat removed from the actual competency the test
is designed to measure.
48,49
In the case of licensing examinations, another limiting factor is the fact that examinees who fail are not able to practice, making it impossible to collect criterion measures. This is not to suggest that studies based on such criterion measures should not be pursued, but in the end, a more compelling argument may rest on less direct evidence demonstrating that the content of the examination reasonably represents the construct of interest and that scores are not unduly influenced by sources of construct-irrelevant variance.
Because logistic constraints necessitate administering high­stakes multiple-choice examinations within structured time limits, one potentially important source of construct­irrelevant variance with such tests is the impact of these time limits on outcomes (often referred to as speededness). It may be (and often is) the case that the ability to respond quickly is not a part of the construct of interest and is not consistent with the intended score interpretations.†† The effects of speededness are another example of a potential source of construct-irrelevant variance.
50
Example 2: Performance Assessment
The primary attraction of performance-based assessment formats is that they have the potential to more directly measure constructs of interest; weakening the generalization argument may be considered acceptable because the
https://t.me/med1917
extrapolation argument is strengthened. However, even though simulation formats may be of high fidelity, there are likely to be aspects that are artificial. There has been relatively little research into the degree to which interactions with standardized patients differ from interactions with real patients, but it seems highly likely that differences exist. Even when standardized patients appear to be indistinguishable from actual patients, factors such as the choice of scoring approach may impact the extent to which the scores can be extrapolated to the performance of interest in practice. Checklists, for example, may fail to capture more subtle interviewing skills that facilitate information gathering. Similarly, knowledge that the interaction is being scored based on a checklist may alter an examinee’s approach to interviewing in order to maximize score points.
The previous comments are intended to highlight the fact that the appearance of similarity between the assessment setting and the practice setting is not in and of itself validity evidence. Using an assessment task that closely approximates the practice setting has the potential to limit the effects of construct-irrelevant variance and construct underrepresentation, but this similarity does not ensure that the score appropriately represents the competency of interest.
Example 3: Workplace-Based Assessment
As with performance-based assessment formats such as those using standardized patients, direct observation is an attractive assessment approach because it has the potential to
https://t.me/med1917
strengthen the extrapolation stage of the validity argument. Because observations are done in the practice setting, differences between the features of the assessment and those of practice may be minimized or eliminated. This characteristic may facilitate construction of an assessment that directly relates to real-world performance, but again it does not in and of itself make the argument for extrapolation. The act of observing may alter the environment. More importantly, the scoring algorithm will shape what is observed and how that observation is transformed into a score. Because it is the score and not the setting that is of interest, collecting observations in the practice setting does not ensure the elimination of construct­irrelevant variance or construct underrepresentation.
In the absence of highly structured scoring algorithms and/or careful training, assessments based on direct observation may be particularly susceptible to halo effects
51
and other sources of construct-irrelevant variance. Unfortunately, in an effort to more clearly define the behaviors to be assessed and avoid such effects, the focus of the assessment may shift from the construct of interest to a set of more easily defined behaviors. In an effort to avoid the effect of construct­irrelevant variance, the scores may suffer from construct underrepresentation. For example, the complex concept of physician–patient communication may be reduced to a set of descriptions, such as “asks open-ended questions” and “makes eye contact.” This may leave out important aspects of the competency such as tone of voice or expressing compassion. When this happens, the knowledge, skills, behaviors, and attitudes that are included in the assessed
https://t.me/med1917
competency will be a limited subset of those in the intended competency. Because of these issues, it is important to remember that even when the real-life behaviors of interest are directly observed, the resulting scores will be a function of the specifics of the instrument used to record the observation.
Decision/Interpretation
The decision stage of the validity argument provides support for the decision rules and theory-based interpretations that are applied to test scores. The most common decision rules will be simple pass/fail classifications based on a single cut score, but conjunctive or partially compensatory rules are not uncommon. Arguments supporting the reasonableness of these rules will be needed if the score interpretations associated with the resulting classification decisions are to be considered credible.
Similarly, score interpretations based on psychological theories about cognition, judgment, or decision-making will only be as credible as the theories themselves. For example, if a score is used to classify practitioners as experts or novices based on their patterns of data collection in reaching a diagnosis, the theory of expert judgment supporting scoring would be critical; if the theory were shown to be flawed, score interpretations would by extension be suspect.
Example 1: A Multiple-Choice Examination
When performance on multiple-choice tests is used to make a
https://t.me/med1917
decision about eligibility for licensure or certification, the appropriateness of the cut score will be a critical part of any validity argument supporting the interpretation that failing candidates are likely to lack some competency that is necessary for safe and competent practice. That said, it must be remembered that standard setting decisions are policy judgments; they are not scientifically verifiable. Given this reality, Kane has argued that appropriate evidence to support the use of a cut score will demonstrate that the
procedure used to establish the standard was appropriate.
52
Information about the choice of procedure, selection of judges, and implementation of the procedure will be central.
The credibility of the decision rule is central to score interpretation for high-stakes standardized tests, but this does not reduce the potential importance of theory-based assumptions. For example, the use of multiple-choice items may be based on the theoretical assumption that the knowledge and judgment required to respond to such items represent a necessary prerequisite for decision-making in practice. While high scores may not provide assurance of good performance in practice (because many other factors can have an influence), low scores on a well-designed test may indicate sufficiently serious knowledge deficits that are unlikely to allow for good performance in practice. These connections represent a theory about clinical decision­making: if the theory is shown to be flawed, the validity of the associated scores similarly will be undermined.
Example 2: Performance Assessment
https://t.me/med1917
Performance assessments (including standardized patient based examinations) sometimes are used to make classification decisions in medical schools or postgraduate education; and in these situations, failing examinees may be required to complete remedial training. When this is done, the assessment takes on the characteristics of a placement test because the test scores result in placing learners in either the standard education track or a remedial program. Evidence to support the decision rule(s) used in this setting might include results demonstrating that learners classified as requiring remediation will show differential improvement when exposed to the remediation program. Alternatively, evidence could be collected to demonstrate that learners so identified have a significantly greater chance of succeeding in future training if they complete the remediation program.
Scoring procedures for performance assessments also may be based, either implicitly or explicitly, on theoretical assumptions about how information is to be aggregated in drawing conclusions about competence. Decisions will need to be made about the relative value of thoroughness and efficiency. Similarly, decisions may need to be made about the importance of physical examination maneuvers. If the practitioner will confirm both negative and positive results with a diagnostic test, the theoretical basis for drawing conclusions about the learner’s diagnostic ability based on their use of a nondiscriminating physical examination maneuver would be questionable at best. These comments are not intended to advocate for or against specific approaches to scoring such examinations; they are intended to highlight the fact that the structure of the scoring procedure ultimately rests on a theoretical view of the
https://t.me/med1917
diagnostic process, and the strength of that model limits the extent to which scores can be interpreted with respect to the learner’s diagnostic competency.
Example 3: Workplace-Based Assessment
As with the formats discussed previously, assessments based on direct observation will depend on theoretical assumptions. Assumptions about the nature of the construct being assessed will dictate the choice of process as opposed to product or outcome measures. Similarly, theories relating to expert-novice differences or cognitive theories about the nature of the medical diagnosis process—and, more broadly, medical decision-making—may influence the data that are collected, the way those data are aggregated, and the way the resulting scores are interpreted.
WBAs often are the basis for feedback to a learner. In this case, it may be that no explicit decision is made based on the scores. In other instances, however, promotional or other high-stakes decisions may be made based at least in part on the results of these assessments. In this circumstance, an implicit (if not explicit) cut score must exist. The argument for this use of the scores will require evidence to support the reasonableness of the cut score or, more broadly, the decision process. The fact that the implicit cut score may be built into the definition of the score scale rather than the result of a separate standard-setting process does not reduce the importance of this part of the validity argument. Such a circumstance might exist when the observer must rate a performance as “adequate” or “inadequate.” The definition
https://t.me/med1917
of the rating may make establishing a cut score unnecessary, but evidence that the definition of “adequate” corresponds to a skill level that is appropriate to support a specific decision still is needed.
Consequential Validity and Program Evaluation
In this chapter, the validity argument has been presented as the accumulation of scientific evidence to evaluate the credibility of intended score interpretations. We would, however, be remiss not to discuss the broader understanding of validity that has existed within the educational measurement community for the past 50 years.
Cronbach,
10
Messick,8 and Kane7 all emphasized the importance of what has come to be known as consequential validity. There has been debate over the years about whether the consequences associated with a testing program belong within the definition of validity. In the end, it may not matter whether evaluation of the consequences of a testing program are considered part of validity. What matters is the clear understanding that just as the test developer has a responsibility to evaluate the evidence that supports inferences made based on test scores, the test developer has a responsibility to evaluate the intended and unintended consequences of the testing program. Whether implemented within the classroom or on a national or international level, testing has consequences; programmatic review of the positive and negative consequences therefore is an important
https://t.me/med1917