Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_2823_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
05.09.2026
Размер:
18 Мб
Скачать
of issues that are central to validity and reliability as these concepts pertain to assessment in medical education.
Historical Context
Practically speaking, the history of test theory as we know it began around the turn of the 20th century with Charles Spearman. Spearman’s interest was in the study of intelligence rather than assessment, and to support his work he developed the field of correlational psychology. Most of the basic equations from classical test theory were developed by Spearman to aid his research on the presence of a common (g) factor shared by most—if not all—tests of mental proficiency.
1–4
This groundwork laid the foundation for a science of testing that expanded explosively during World War I. The US Army had a monumental personnel problem: tens of thousands of recruits had to be placed in jobs. Testing provided a potentially effective and efficient means of determining appropriate job placements.5 This effort established the practice of psychological testing in the United States; not surprisingly, the science of testing was used in an effort to boost educational and industrial efficiency after the war. In both military and industrial contexts, the question of interest was “How well do these tests predict performance on the job?” Evidence to justify the use of the test naturally conformed to the approach established by Spearman and took the form of a correlation between the test scores and an independent assessment of job performance.
https://t.me/med1917
The dramatic proliferation of placement testing did much to define the view of validity during the period from 1920 through 1950. Correlational evidence, referred to as criterion validity, was the standard during this period; in his 1951 chapter in the first edition of Educational Measurement, Edward Cureton defined validity “in terms of the correlation between the actual test scores and the ‘true’ criterion score.”
6
As a practical matter, criterion validity has obvious utility. In placement testing, it has clear relevance to the interpretation of the score and it provides an objective basis for comparing multiple assessments available for a given purpose. However, the strength of this approach is less apparent for applications beyond placement testing. One problem is that an obvious and practical criterion may not be available; no clear and objective external criterion is likely to exist for an achievement test. And if such a criterion is identified, the test developer would need to provide validity evidence to support its use.
7
Questions about the appropriateness of criterion validity as a primary evaluation of assessments of academic achievement led to the development of procedures for assessing content validity. The purpose of such evidence is to establish that the content of the test reasonably represents the domain of interest. This type of evidence clearly is necessary, but it is not sufficient to establish the validity of interpretations for an achievement test. As Messick pointed out, evidence that the test content is relevant to the domain(s) of interest provides no direct support for inferences that are made based on the test scores.
8
https://t.me/med1917
During the period after World War II, interest in personality testing and the development of testing instruments to standardize the process pushed researchers to continue considering the types of evidence required to support the use of these new instruments. Neither criterion nor content validity models provided a particularly good fit to these tests. It was in this context that Cronbach and Meehl introduced the idea of construct validity.9 In the second edition of Educational Measurement,10 Cronbach made the following comment when describing the underlying rationale for this new conceptualization of validity:
The rationale for construct validation (Cronbach and Meehl,
1955) developed out of personality testing. For a measure of, for example, ego strength, there is no uniquely pertinent criterion to predict, nor is there a domain of content to sample. Rather, there is a theory that sketches out the presumed nature of the trait. If the test score is a valid manifestation of ego strength, so conceived, its relations to other variables conform to the theoretical expectations.
This approach to validation greatly expanded the types of evidence that could be considered when evaluating an assessment. For example, in the context of achievement testing, construct validation might argue for collecting evidence to demonstrate that learners with advanced training in the topic area outperform those with less training.
The 1950s brought two other important changes to the conceptualization of validity. First, Campbell and Fiske introduced the multitrait-multimethod matrix: a way to organize validity evidence about the relative strengths of the
https://t.me/med1917
relationships between traits measured by a single method and measures of the same trait using different methods.11 In the context of personality testing, examples of traits may have included extraversion and aggression; methods may have included individual examiner-administered assessments and group-administered paper-and-pencil assessments. Campbell and Fiske’s matrix provided an empirical means of assessing the extent to which scores are impacted by otherwise irrelevant characteristics of the assessment method or format (signaled by relatively higher correlations between different traits measured by the same method compared with the same trait measured by different methods). This method effect relates to what was later to become known as construct-irrelevant variance, a concept that will be addressed in more detail later in this chapter. The second important change in the conceptualization of validity came when Loevinger focused attention on the proposed interpretation of test scores.12 This represented an important shift in perspective from consideration of the relationship between the construct the test was designed to measure and the test score to consideration of the correspondence between what is measured by the test and the proposed interpretations of the test score.
By the time of publication of the third edition of Educational Measurement, Messick had developed a unified theory of validity.8 Rather than being defined as “the correlation between the actual test scores and the ‘true’ criterion score,”
6
validity now was viewed as the “. . . degree to which empirical evidence and theoretical rationales support the adequacy and appropriateness of interpretations and actions
https://t.me/med1917
based on test scores.”8 Messick’s model built on the contributions of his predecessors; following Cronbach and Meehl,9 Cronbach,10 and Loevinger,12 he emphasized the need to specify the intended meaning and use of the test score before validation. Consistent with Cronbach and Meehl and Campbell and Fiske,11 Messick emphasized the importance of considering alternative hypotheses such as the impact of construct-irrelevant variance. Additionally, like his predecessors, Messick argued that the process of validation would involve an extended program of research.
*
Messick’s formulation is consistent with previous validity frameworks, although his conceptualization introduces a change in emphasis. In particular, he placed increased emphasis on evaluating the consequences of the testing program. He believed that both the actual and potential social consequences of a test must be evaluated. Considering as an example a test for medical licensure, at a minimum this requirement leads to examination of consequences such as the test’s impact on what instructors choose to teach and what learners choose to learn. More broadly, consequential validity would require consideration of the test’s impact on the availability of medical practitioners both for the community at large and, perhaps, specifically for underserved communities. Messick’s view of consequential validity went beyond these considerations; he additionally required consideration of the impact that such an examination might have on the entrance of minority candidates into the profession. This broad definition of consequential validity emphasizes the importance of test developers and administrators accepting responsibility for
https://t.me/med1917
their actions. The definition takes the validation process beyond the scientific evaluation of the assessment into the arena of social and political values.
By 1999, the role of consequences within the sphere of validity was sufficiently well established that it was included as one of the five sources of validity evidence referenced in the Standards for Educational and Psychological Testing.13 They are presented here to provide continuity with previous secondary sources describing Messick’s concept of validity theory. It is, however, important to remember that both Messick’s unified theory of validity and the Standards emphasize that these are not different types of validity. Rather, they are different sources of evidence, each of which may be more or less important in providing support for a specific score interpretation.
The history of validity theory should make it clear that the definition of validity has expanded over time. The emphasis also has changed as the focus of testing has changed. Criterion validity (evidence based on relationships to other variables) has not been replaced; this type of evidence remains essential in evaluating admissions and employment tests. Similarly, content validity represents an important source of evidence in support of tests of achievement. The history of validity is a history of both an expansion in meaning and a shift in emphasis.
More recently, Kane has introduced an additional shift in perspective by representing validity as an argument in support of the proposed interpretations and uses of a test score.
7,14
As with previous phases in the evolution of validity
https://t.me/med1917
theory, Kane’s view does not deny the importance of the evidence and perspectives that have been discussed during the past half century; it provides a shift in perspective rather than a rejection of the basic arguments. That shift in perspective does have one important characteristic: it highlights the fact that the collection of evidence in support of the interpretations of test scores must form a structured and coherent argument that leads from the test administration to the interpretation. That structured argument is only as strong as its weakest component.
Kane’s View of Validity
Implicit in the interpretation of a test score is a series of assertions and assumptions that support that interpretation. For example, the interpretation of a passing score on a medical licensing examination requires the assumption that the test was administered under standardized conditions and that the examinee did not have prior access to the test material. If the examinee cheated, no interpretation can be made about the score regardless of other characteristics of the test. Interpretation of the test score requires assumptions about the precision of the score; if the test score is not reproducible, there is no basis for making an interpretation. Interpretation of the score assumes that the test measures some relevant aspect of the overall set of knowledge, skills, and abilities required for the practice of medicine. It also assumes that the cut score (pass/fail standard) has been established in a way that supports the interpretation. If any one of these assumptions is unfounded, the strength of the others may be of little relevance.
https://t.me/med1917
Kane provides a structure for this validity argument that outlines four links in the inferential chain from the test administration to the final decision or interpretation.
7,14
He labels these four components scoring, generalization, extrapolation, and decision. Support for the scoring component of the overall argument includes evidence that the test was administered properly, examinee behavior was captured correctly, and scoring rules were appropriate and were applied accurately and consistently. The generalization component of the argument requires evidence that the observations were appropriately sampled from the universe of test items, clinical encounters, et cetera. Generalization also requires evidence that the sample of observations was large enough to produce scores with an acceptable level of precision. Broadly speaking, this stage in the argument asks the question: Is the test reliable? In this context, generalization refers to generalizing from the sample of behavior that was part of the test (the observed score) to the test-taker’s true score or universe score. The extrapolation component of the argument requires evidence that the observations represented by the test score were relevant to the target competency or construct measured by the test. This requires a demonstration that the observations were relevant to the interpretation and that the scores were not unduly influenced by sources of variance that are irrelevant to the intended interpretation. Extrapolation also requires that the target competency that the test is intended to measure is reasonably well represented by the test score. The decision component of the argument requires evidence in support of any theoretical framework necessary for score interpretation or evidence in support of decision rules. For tests with a cut
https://t.me/med1917
score, this evidence would include support for the procedure used to establish that cut score. Again, the score user can have confidence in an interpretation only if there is evidence for each component of the overall argument. The types of evidence required will vary with the purpose and characteristics of the assessment.
These last two sentences are critically important and warrant special comment because, since the last edition of this volume, there has been considerable interest within medical education in longitudinal assessment, formative assessment, and other forms of assessment for learning. There also have been numerous publications advocating the importance of including a broad range of assessment that allows for “triangulation.”
15,16
These publications sometimes have drawn a distinction between tests designed to differentiate among individuals and those designed to identify strengths and weaknesses within an individual. They also have described these changes as leading to a “postpsychometric” era. To avoid confusion, we want to be clear that we believe richer sources of evidence are valuable regardless of the purpose of the assessment. At the same time, declaring that something represents a rich source of evidence does not make it so. Sources of feedback should be taken seriously when (and only when) there is evidence to support the validity of that feedback. This includes assessments that result in numeric scores and those that produce written or narrative feedback. Formative assessment of clinical competence is intended to support the development of expertise. That development requires practice; it also requires accurate and timely feedback. We will return to this
https://t.me/med1917
issue throughout the chapter.
The next sections of the chapter will further explicate the four components of the validity argument as described by Kane. Within each section, content relevant to the particular component will be provided for three types of assessments that span a range of the types of assessments currently used in medical education: a multiple-choice examination, a performance assessment, and a workplace-based assessment (WBA). Multiple-choice examinations are ubiquitous in medical education from selection to medical school and through classroom assessment to credentialing. Performance assessments have a long history in medical education, with objective structured clinical examinations (OSCEs) and standardized-patient-based examinations being in common use within medical schools and having a history as part of licensure assessment.17 WBAs are becoming an increasingly important part of assessment within medical education, particularly during residency training.
18–20
Table 2.1 provides examples of the kinds of questions that arise at each stage of the validity argument. The questions are provided as examples and are not intended to represent an exhaustive list. As you read what follows, we encourage you to think about the assessments that are presented, extrapolate to other assessments, and add your own questions to this list.
Table 2.1
Examples of Questions Supporting Each of the Four
Components of Kane’s Argument-Based Approach to
Validity
Component Questions
https://t.me/med1917