Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_112_библиотеки_им_акад_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
05.09.2026
Размер:
18 Мб
Скачать
Scoring
1. Were the observations made or stimulus materials administered under standardized conditions?
2. Were the scores recorded accurately?
3. Were the scoring algorithms applied correctly?
4. Were appropriate security procedures implemented?
Generalization1. What are the sources of measurement
error that contribute to the observed scores on the assessment?
2. How similar would scores be across replications of the measurement procedure?
3. How similar would classification decisions be across replications of the measurement procedure?
4. To what extent are test forms constructed using a systematic process?
Extrapolation 1. To what extent do the scores correspond
to real-world competencies of interest?
2. Are there factors that interfere with assessment of the competencies of interest?
3. Do scores predict real-world outcomes of interest?
4. Are there artificial aspects of the testing conditions that impact the scores?
Decision 1. Was the standard established through
implementation of a defensible and properly implemented procedure?
2. Do examinees identified for remediation
https://t.me/med1917
more from a remediation program than would those who were not identified?
Scoring
The scoring component of the validity argument must provide evidence that assessment data have been collected appropriately and scored accurately. This will include consideration of a variety of types of evidence, such as the extent to which the stated conditions of standardization have been implemented and the accuracy of the scoring process. As is true for each component of the validity argument, the specifics of the evidence that will be relevant to the scoring aspect of the argument will vary with the characteristics of the assessment. Again, it is important to remember that any argument for the validity of a score interpretation is only as strong as the weakest link in that argument!
Example 1: A Multiple-Choice Examination
Standardized tests have been developed to provide the strongest possible evidence for the scoring component of the validity argument. The multiple-choice format was developed in 1915 explicitly to support objective scoring.
21
Adherence to the conditions of standardization ensures that the data are collected in the same manner for all examinees. Factors such as the time allowed for the examination, the seating, the lighting, and the quality of the stimulus materials are controlled. To the extent that administration procedures require documentation of the violation of these conditions and annotation of score reports, the score user
https://t.me/med1917
will have confidence in the conditions under which the test responses have been collected. Similarly, professionally administered and scored tests routinely will have quality control steps built into the scoring process. Key validation— statistical analyses of examinee responses designed to verify that the keyed answer is correct—provides evidence that the scoring rules have been applied accurately. This step includes activities such as (1) examining the proportion of examinees receiving credit for each item; and (2) comparing the probability of a correct response for examinees at different score levels.
A fundamental consideration for high-stakes
tests is security. In high-stakes settings, examinees may be motivated to cheat and may attempt to do so in any number of ways. When items are reused from one administration to another, it is possible for examinees testing earlier to steal (i.e., remember, copy, photograph) items and make them available to those testing on a later date. When computerized examinations are administered on a continuous basis, this threat to validity may be increased. Evidence about the size of the item pool and the frequency with which items are reused will support the user’s confidence that prior exposure has not threatened the integrity of the score. For computer­based tests, encryption of test items at all times except when they are displayed on the screen may provide additional confidence in the security of the test material. Tests that are administered nationally or internationally are particularly likely to be targeted by individuals or groups interested in breaching security, but the same issues apply to tests developed and administered within a single medical school
https://t.me/med1917
if items (or entire test forms) are used on multiple occasions.
Although these steps appear to fall under the heading of test development or test administration, they also are critical pieces of evidence to support the validity argument for standardized written examinations. Because they are a critical part of the validity argument, it is important that the steps in administration and scoring are verified and documented; it is not reasonable to assume that this is a given, even with professionally developed assessments.
Example 2: Performance Assessment
The reproducibility of the stimulus material and scoring procedures is, as previously noted, a strength of standardized tests comprising multiple-choice items. Relatively little effort is required to be satisfied that two examinees assigned to the same test form but sitting at different computers are seeing the same items and that those items are being scored in the same way. The same is not necessarily true for performance assessments such as standardized patient–based assessments or other formats that require humans to present and/or score the assessment. Adding the human element creates the possibility that two standardized patients trained to portray the same scenario may perform in a less-than-standardized manner; the same standardized patient may not portray the same scenario in the same way on two different occasions. The scoring phase of the validity argument will need to include evidence that standardized patients are trained to an acceptable standard, and it also will require evidence that standardized patients
https://t.me/med1917
are monitored over time to ensure both inter- and intra­patient consistency. Similar issues arise with scoring for these tests; whether the scores are produced by a standardized patient or content expert, it will be necessary to assess the accuracy of the process. Again, this aspect of testing must be verified before testing begins and must continue to be monitored over time. It also is important to remember that collecting evidence of a high level of rater agreement during a small-scale pilot administration should not replace collecting the same evidence once the test is being administered operationally.
In addition to verifying that the overall error rate is low, both in standardized patient portrayal and in the scoring process, it will be important to provide evidence of a lack of significant relationship between examinee characteristics and standardized patients’ portrayal or scoring. For example, the gender or ethnicity of a learner should have no impact on the way the scenario is portrayed and scored. If data suggest that learners of otherwise equal competence are likely to receive better scores if they are, for example, male rather than female, this would be a serious threat to valid score interpretation. This type of effect is more serious than random error in portrayal or scoring because random errors tend to average out across encounters; systematic effects do not.
Security issues also may be important with performance assessments. If the test is used to make important decisions, learners may attempt to improve their scores by gaining prior access to test information. In most circumstances, performance assessments (and standardized patient–based
https://t.me/med1917
tests in paticular) are administered on multiple occasions. This creates the opportunity for learners who have completed the examination to share information with others who will test in the future; in most situations, prior knowledge about the specific tasks that will appear on a test should be expected to influence scores.22 This threat to validity is analogous to the problem associated with the reuse of material on tests comprising multiple-choice items, but in the case of performance assessments it is much more difficult to produce large banks of test “items.” When tests are administered during a relatively short period, sequestering examinees to prevent the sharing of information may provide evidence that this threat to validity has been controlled. With standardized patient–based tests, an additional threat to security exists in that standardized patients (or examiners, if they are used in test administration) themselves may share information with learners before or during the test administration.
Example 3: Workplace-Based Assessment
Assessment of a clinician or trainee through direct observation involves taking another step away from the completely standardized stimulus material of the multiple­choice examination and the partially standardized conditions that exist in performance assessment; it takes place in the largely uncontrolled conditions of the authentic clinical environment. To support the interpretations of scores produced in these settings, it will be necessary to produce evidence that different assessors working in different settings are, in fact, assessing the same construct in the same
https://t.me/med1917
way. One way to provide such evidence would be to carefully define the characteristics of performance that will be rated. A combination of careful specification of what is being rated and thorough training of assessors may provide reasonable support for the assertion that individuals are being assessed on the same construct. The disadvantage of carefully defining the assessment content is that it may restrict what can be assessed to those aspects of the construct that can be defined easily; this will have an impact on the potential to extrapolate from the scores to the construct of interest. The alternative may be that each assessor defines the construct in their own way, but this approach clearly leaves the scoring aspect of the validity argument seriously weakened.
Even with careful attention to activities such as specifying the rating elements and assessor training, it will be important to collect evidence demonstrating that assessors are, in fact, assessing the same constructs. In a paper reporting on a WBA system implemented at 15 US residency programs, the authors describe several activities intended to strengthen the scoring component of the validity argument. In the early stages of instrument development, cognitive interviews were conducted with assessors in the roles that would be participating in the actual assessment. This allowed for collecting important information about how the items were being interpreted by different groups of assessors and led to targeted item revisions focused on decreasing the variability in individual interpretation and completion of items. In addition, as part of the instrument development process the researchers often would include a brief introduction to the item prior to the specific rating question. This was done in an
https://t.me/med1917
attempt to provide a clear foundation for responding to the item and, by extension, to decrease the extent to which individual assessors responded to items based on their own interpretations of what question was being asked. The following provides an example of one of these items. There is an initial description of expectations for the person being assessed:
As a member or leader of a clinical team, a resident is counted on to keep interprofessional team members aware of: (1) patients' current status and potential for deterioration; and (2) any changes in status (e.g., physical exam, labs) that occur.
And the scorable item follows:
Thinking about situations in which the resident would have been expected to inform the team about patient status, indicate the degree to which the resident kept you informed (rating scale: continuum between did not keep me informed and kept me
fully informed).
Providing specific examples to help assessors focus on the same construct when responding to assessment items is just one approach to addressing scoring-related considerations in WBA settings. Relevant considerations will differ based on the assessment context and the inferences that are to be made based on test scores. As such, a critical first step in the process is to ensure an understanding of the various factors that can impact the scoring component of the validity argument. This will help to ensure that the necessary validity
https://t.me/med1917
evidence can be generated and collected.
As with other forms of assessment that do not lend themselves to machine scoring, careful training of assessors will be important with WBAs. That said, evidence to support the scoring component of the validity argument likely will include an evaluation of the accuracy of the scores rather than simply documentation of careful training.
Generalization
This stage of the argument focuses on the question of whether an examinee would receive the same score if the assessment were repeated. We would not waste our time standing on a scale if we discovered that the dial reported a wildly different weight when we stepped off and back on. Technically speaking, the generalization stage of the argument is focused on the relationship between the observed scores and the associated universe scores (or true scores). Both universe scores and true scores are conceptualizations. The universe score represents the score that an examinee would receive if it were possible for that examinee to respond to all items representing the universe of acceptable observations (i.e., if the examinee responded to all items in the domain). The true score is a closely related concept representing the mean score that the examinee would receive if they completed an unlimited number of randomly equivalent (parallel) forms of the test. (The observed score is the actual score that is recorded when an examinee completes a specific test form.) The details of these definitions and related theories are beyond the scope of this
https://t.me/med1917
chapter; the interested reader is referred to Gulliksen1 and Lord and Novick
23
for a detailed discussion of classical test theory and to Cronbach and associates24 and Brennan25 for discussions of generalizability theory.
Two kinds of evidence are required for this stage of the argument. First, it is necessary to show that the sample of items or observations made of the examinee are representative of the domain to which the score is to be generalized. Second, it is necessary to demonstrate that the sampling is sufficiently extensive to prevent the observed scores from being unduly influenced by sampling error. The extent to which the sample is representative will depend on the procedures used for test construction (a blueprint/sampling plan for data collection for WBAs); the adequacy of the sampling can be examined directly through a well-developed set of theory-based statistical procedures.
The samples will be representative to the extent that data collection follows specified rules. In some cases, random selection from a specified domain will be appropriate; in others, stratified sampling will be preferred. In some contexts, rules for the range of conditions under which observations may (or must) be made will replace the sampling of stimulus material.
The most developed aspect of test theory by far relates to evaluation of reliability; conceptually this methodology is designed to assess the relationship between observed scores and true scores or universe scores. The most common index of this relationship is the reliability coefficient; this coefficient represents the correlation between the observed test scores
https://t.me/med1917