Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_2823_Библиотеки_им_академика_М_И_Перельмана
.pdf
responsibility. The importance of evaluating the intended
consequences should be obvious. If a licensing examination
is implemented with the explicit purpose of protecting the
health of the public from individuals who lack the
knowledge, skills, and attitudes necessary to provide safe
and competent care, it is reasonable to expect evidence
indicating that the program is contributing to that outcome.
Similarly, if a test is a part of a program of formative
assessment intended to accelerate learning, good intentions
are not sufficient. Evidence is needed to support the view
that the assessment program leads to improved educational
outcomes.
Reasonable program evaluation for a testing program must
provide evidence about the extent to which the program is
having the intended outcome; it also must examine the
extent to which the program is producing unintended
consequences. In the case of a licensing examination, one
possible unintended consequence is the impact that the test
content has on the curriculum of training programs (e.g.,
medical schools). Typically, a licensing examination will
assess only a subset of the knowledge and skills needed for
practice. If this leads to a narrowing of the curriculum, a
program intended to ensure the competency of healthcare
providers actually may have a negative impact on the
training system. (It is worth noting that the same issue may
arise with progress testing in medical schools.) Good
intentions do not ensure uniformly good outcomes.
Test content, format, and scoring procedures all may have
unintended consequences. Consider a test of clinical skills in
which an individual’s competence in taking a patient history
https://t.me/med1917

is scored using a checklist that awards points for asking
specific relevant questions of the patient. The checklist score
might provide a useful surrogate for a more nuanced and
complex approach to collecting evidence for competence in
clinical reasoning. It also might motivate students to collect a
patient history using a “shotgun” approach: asking a broad
range of questions to maximize the score rather than
focusing on a more appropriate history-taking approach that
supports inferences about competency in diagnostic
reasoning.
There is, of course, an important difference between
consequential validity and the rest of the validity evidence
we have discussed to this point: the evaluation of
consequences requires value-based judgments. The intended
consequences of the testing program will—at least in the
eyes of those responsible for the program—have a positive
valence. Unintended consequences could be happy surprises
that increase the value of the program, or they could
represent unanticipated costs. Both the valence and the
associated magnitude may be a function of the policy driving
the program and the values of those evaluating the program.
Consider, as a hypothetical example, a test impacting
acceptance to, advancement in, or graduation from medical
school. This hypothetical test has ample evidence to support
the inferences as described in the previous pages, and yet use
of the test has led to a reduction in the number of minority
candidates entering practice. The argument-based approach
to validity described in this chapter would need to include
evidence that the differential performance was the result of
actual differences in the competencies the test was intended
to measure and not construct-irrelevant variance in the test
https://t.me/med1917

scores. The decision about when and how to use the scores
would require value-based judgments, and these values may
differ for different stakeholder groups.
Conclusion
In this chapter we have conceptualized validity as a
systematic argument in support of score interpretations.
Reliability has been viewed as a component of the overall
validity argument. The details of the specific examples
should be viewed as unimportant, and whether a specific
piece of evidence is seen as part of the extrapolation stage or
the decision stage of the argument is secondary. The central
issue is that the overall argument must be complete and coherent.
There is no such thing as a valid test; the validity argument
must focus on intended interpretations of test scores. To
construct such an argument, researchers and users of the test
scores must systematically and self-critically collect a wide
array of evidence that provides clear insight into the
credibility of those interpretations.
Placement of this chapter at the beginning of the volume is
intentional, as the included concepts are critical for any
serious consideration of measurement. Reliability and
validity are at once nuanced yet simple concepts. Reliability
is an evaluation of the stability or precision of a measure. In
this regard, an individual’s weight, A1C level, and
competence in clinical reasoning share one essential thing in
common: we would not give credence to a measure of any of
these characteristics if we believed that we would get widely
disparate results if we stepped off and back on the scale,
https://t.me/med1917

repeated the reading on the same blood sample, or
implemented an equivalent form of our assessment of
clinical reasoning. Similarly, validity might be reduced to the
question, does the test measure what it was intended to
measure? Clearly, if the answer is no, we should not use the
measure. Unfortunately, the answer to that question is not
always obvious—face validity can be deceiving. Simply
speaking, the centrality of reliability, and validity more
broadly, cannot be overstated.
It also is important to note that the importance of these
concepts is not linked to the specifics of the measurement.
These issues are critical in high-stakes and low-stakes testing,
assessments that produce numeric scores and assessments
that result in text-based feedback, tests intended to measure
individual differences and tests intended to measure an
individual’s relative strengths, tests that break down
complex constructs into components and tests that
holistically evaluate performance, tests that are used for
formative purposes, tests that are used for summative
purposes, and tests that are used for both.
As we have described in this chapter, reliability and validity
are the foundation of the psychometric requirement that the
interpretations we make based on test results are supported
by evidence. The place of evidence in the interpretation of
test results—evidence-based measurement—is no different
than the place of evidence in evidence-based medicine. In
both cases, collecting the evidence is hard work, but it is
essential work.
https://t.me/med1917

Annotated Bibliography
1. Kane MT. Validating the interpretations and uses of test
scores. J Educ Measure. 2013 Mar 14;50:1-73.
doi:10.1111/jedm.12000. This publication provides a current,
in-depth discussion of Michael Kane’s approach to validity
and validation, describing the validation process as a
structured argument in support of the intended
interpretations made based on test scores.
2. Clauser BE, Margolis MJ, Case SM. Testing for licensure
and certification in the professions. In: Brennan RL, ed.
Educational Measurement. 4th ed. American Council on
Education/Praeger; 2006:701–731. This chapter provides an
overview of assessment methods commonly used in
licensure and certification examinations in the professions. It
includes an expanded discussion of the evolution of validity
and validation over the past century, and discusses
additional applications of Kane’s framework to assessment
methods commonly used in the professions.
3. Cook DA, Brydges R, Ginsburg S, Hatala R. A
contemporary approach to validity arguments: a practical
guide to Kane’s framework. Med Educ. 2015 Jun;49(6):560-
575. doi:10.1111/medu.12678. This paper provides a very
readable introduction to Kane’s validity framework for both
quantitative and qualitative assessment methods commonly
used in medical education. As a part of the discussion, it
highlights some of the parallels between validation work and
evaluation of diagnostic studies in medicine.
https://t.me/med1917

4. Cook DA, Zendejas B, Hamstra SJ, Hatala R, Brydges R.
What counts as validity evidence? Examples and prevalence
in a systematic review of simulation-based assessment. Adv
Health Sci Educ Theory Pract. 2014 May;19(2):233-250.
doi:10.1007/s10459-013-9458-4.
5. Cook DA, Brydges R, Zendejas B, Hamstra SJ, Hatala R.
Technology-enhanced simulation to assess health
professionals: a systematic review of validity evidence,
research methods, and reporting quality. Acad Med. 2013
Jun;88(6):872-883. doi:10.1097/ACM.0b013e31828ffdcf. Using
the frameworks proposed by Messick and Kane, these
systematic reviews summarize sources of validity evidence
from studies of technology-enhanced simulation-based
assessments, identifying methodological and reporting
shortcomings and recommending directions for
improvement in future research.
6. Hatala R, Cook DA, Brydges R, Hawkins R. Constructing a
validity argument for the Objective Structured Assessment of
Technical Skills (OSATS): a systematic review of validity
evidence. Adv Health Sci Educ Theory Pract. 2015
Dec;20(5):1149-1175. doi:10.1007/s10459-015-9593-1. This
systematic review uses Kane’s framework to analyze the
validity argument for the objective structured assessment of
technical skills (OSATS). They found that, in general, validity
evidence supports the use of OSATS for formative feedback,
but more research is required to support use of OSATS for
making higher-stakes decisions.
7. Hawkins RE, Margolis MJ, Durning SJ, Norcini JJ.
Constructing a validity argument for the mini-clinical
https://t.me/med1917

evaluation exercise: a review of the research. Acad Med. 2010
Sep;85(9):1453-1461. doi:10.1097/ACM.0b013e3181eac3e6.
This systematic review of research conducted from 1995 to
2009 uses Kane’s validity framework to evaluate validity
evidence related to the mini clinical evaluation exercise
(mini-CEX). It concludes that scoring-related issues (e.g.,
leniency error and high interitem correlations) limit the
utility of the mini-CEX for providing feedback to trainees,
though evidence related to the generalization and
extrapolation components is generally supportive of the
validity of mini-CEX score interpretations.
References
1. Gulliksen H. Theory of Mental Tests. John Wiley & Sons;
1950.
2. Spearman C. Proof of the measurement of association
between two things. Am J Psychol. 1904;15:72–101.
3. Spearman C. “General intelligence” objectively
determined and measured. Am J Psychol. 1904;15:201–292.
4. Spearman C. Correlation calculated with faulty data. Br J
Psychol. 1910;3:271–295.
5. Yoakum CS, Yerkes RM. Mental Tests in the American
Army. Sidgwick & Jackson; 1920.
6. Cureton EE. Validity. In: Lindquist EF, ed. Educational
Measurement. American Council on Education; 1951:621–
https://t.me/med1917

694.
7. Kane MT. Validating the interpretations and uses of test
scores. J Educ Measure. 2013 Mar 14;50:1-73.
doi:10.1111/jedm.12000
.
8. Messick S. Validity. In: Linn RL, ed. Educational
Measurement. 3rd ed. American Council on
Education/Macmillan; 1989:13–103.
9. Cronbach LJ, Meehl PE. Construct validity in
psychological tests. Psych Bull. 1955 Jul;52(4):281–302.
doi:10.1037/h0040957.
10. Cronbach LJ. Test validation. In: Thorndike RL, ed.
Educational Measurement. 2nd ed. American Council on
Education; 1971:443–507.
11. Campbell DT, Fiske DW. Convergent and divergent
validation by the multitrait-multimethod matrix. Psych Bull.
1959 Mar;56(2):81–105.
12. Loevinger J. Objective tests as instruments of
psychological theory. Psych Rep. 1957 Jun;3:635–694.
doi:10.2466/pr0.1957.3.3.635.
13. American Educational Research Association, American
Psychological Association, National Council on
Measurement in Education. Standards for Educational and
Psychological Testing. American Educational Research
Association; 1999.
14. Kane M. An argument-based approach to validation.
https://t.me/med1917

Psych Bull. 1992;112(3):527–535. doi:10.1037/0033-
2909.112.3.527.
15. Hodges B. Assessment in the post-psychometric era:
learning to love the subjective and collective. Med Teach.
2013 Jul;35(7):564–568. doi:10.3109/0142159X.2013.789134.
16. Schuwirth LWT, van der Vleuten CPM. A history of
assessment in medical education. Adv Health Sci Educ. 2020
Dec;25(5):1045–1056. doi:10.1007/s10459-020-10003-0.
17. Swanson DB, van der Vleuten CP. Assessment of clinical
skills with standardized patients: state of the art revisited.
Teach Learn Med. 2013;25 Suppl 1:S17-S25.
doi:10.1080/10401334.2013.842916.
18. Norcini J, Burch V. Workplace-based assessment as an
educational tool: AMEE Guide No. 31. Med Teach. 2007
Nov;29(9-10):855-871. doi:10.1080/01421590701775453.
19. Hicks PJ, Margolis MJ, Carraccio C, et al. A novel
workplace-based assessment for competency-based decisions
and learner feedback. Med Teach. 2018 Nov;40(11):1143-1150.
doi:10.1080/0142159X.2018.1461204.
20. Hicks PJ, Margolis MJ, Poynter SE, et al. The Pediatrics
Milestones Assessment Pilot: development of workplacebased assessment content, instruments, and processes. Acad
Med. 2016 May;91(5):701-709.
doi:10.1097/ACM.0000000000001057.
21. Kelly FJ. The Kansas Silent Reading Test. The Kansas
Printing Plant; 1915.
https://t.me/med1917

22. Swanson DB, Clauser BE, Case SM. Clinical skills
assessment with standardized patients in high-stakes tests: a
framework for thinking about score precision, equating, and
security. Adv Health Sci Educ Theory Pract. 1999;4(1):67–106.
doi:10.1023/A:1009862220473.
23. Lord FM, Novick MR. Statistical Theories of Mental Test
Scores. Addison-Wesley; 1968.
24. Cronbach LJ, Gleser GC, Nanda H, Rajaratnam N. The
Dependability of Behavioral Measurements: Theory of
Generalizability for Scores and Profiles. John Wiley & Sons;
1972.
25. Brennan RL. Generalizability Theory. Springer-Verlag;
2001.
26. Brennan RL. An essay on the history and future of
reliability from the perspective of replications. J Educ Meas.
2001;38(4):295–317. doi:10.1111/j.1745-3984.2001.tb01129.x.
27. Brown W. Some experimental results in the correlation of
mental abilities. Br J Psych. 1910;3:296–322.
28. Kuder GF, Richardson MW. The theory of estimation of
test reliability. Psychometrika. 1937;2:151–160.
doi:10.1007/BF02288391.
29. Cronbach LJ. Coefficient alpha and the internal structure
of tests. Psychometrika. 1951;16:297–334.
doi:10.1007/BF02310555.
30. Angoff WH. Scales, norms, and equivalent scores. In:
https://t.me/med1917
Соседние файлы в папке Библиотека им академика М.И. Перельмана
