Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_2823_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
05.09.2026
Размер:
18 Мб
Скачать
responsibility. The importance of evaluating the intended consequences should be obvious. If a licensing examination is implemented with the explicit purpose of protecting the health of the public from individuals who lack the knowledge, skills, and attitudes necessary to provide safe and competent care, it is reasonable to expect evidence indicating that the program is contributing to that outcome. Similarly, if a test is a part of a program of formative assessment intended to accelerate learning, good intentions are not sufficient. Evidence is needed to support the view that the assessment program leads to improved educational outcomes.
Reasonable program evaluation for a testing program must provide evidence about the extent to which the program is having the intended outcome; it also must examine the extent to which the program is producing unintended consequences. In the case of a licensing examination, one possible unintended consequence is the impact that the test content has on the curriculum of training programs (e.g., medical schools). Typically, a licensing examination will assess only a subset of the knowledge and skills needed for practice. If this leads to a narrowing of the curriculum, a program intended to ensure the competency of healthcare providers actually may have a negative impact on the training system. (It is worth noting that the same issue may arise with progress testing in medical schools.) Good intentions do not ensure uniformly good outcomes.
Test content, format, and scoring procedures all may have unintended consequences. Consider a test of clinical skills in which an individual’s competence in taking a patient history
https://t.me/med1917
is scored using a checklist that awards points for asking specific relevant questions of the patient. The checklist score might provide a useful surrogate for a more nuanced and complex approach to collecting evidence for competence in clinical reasoning. It also might motivate students to collect a patient history using a “shotgun” approach: asking a broad range of questions to maximize the score rather than focusing on a more appropriate history-taking approach that supports inferences about competency in diagnostic reasoning.
There is, of course, an important difference between consequential validity and the rest of the validity evidence we have discussed to this point: the evaluation of consequences requires value-based judgments. The intended consequences of the testing program will—at least in the eyes of those responsible for the program—have a positive valence. Unintended consequences could be happy surprises that increase the value of the program, or they could represent unanticipated costs. Both the valence and the associated magnitude may be a function of the policy driving the program and the values of those evaluating the program. Consider, as a hypothetical example, a test impacting acceptance to, advancement in, or graduation from medical school. This hypothetical test has ample evidence to support the inferences as described in the previous pages, and yet use of the test has led to a reduction in the number of minority candidates entering practice. The argument-based approach to validity described in this chapter would need to include evidence that the differential performance was the result of actual differences in the competencies the test was intended to measure and not construct-irrelevant variance in the test
https://t.me/med1917
scores. The decision about when and how to use the scores would require value-based judgments, and these values may differ for different stakeholder groups.
Conclusion
In this chapter we have conceptualized validity as a systematic argument in support of score interpretations. Reliability has been viewed as a component of the overall validity argument. The details of the specific examples should be viewed as unimportant, and whether a specific piece of evidence is seen as part of the extrapolation stage or the decision stage of the argument is secondary. The central issue is that the overall argument must be complete and coherent. There is no such thing as a valid test; the validity argument must focus on intended interpretations of test scores. To construct such an argument, researchers and users of the test scores must systematically and self-critically collect a wide array of evidence that provides clear insight into the credibility of those interpretations.
Placement of this chapter at the beginning of the volume is intentional, as the included concepts are critical for any serious consideration of measurement. Reliability and validity are at once nuanced yet simple concepts. Reliability is an evaluation of the stability or precision of a measure. In this regard, an individual’s weight, A1C level, and competence in clinical reasoning share one essential thing in common: we would not give credence to a measure of any of these characteristics if we believed that we would get widely disparate results if we stepped off and back on the scale,
https://t.me/med1917
repeated the reading on the same blood sample, or implemented an equivalent form of our assessment of clinical reasoning. Similarly, validity might be reduced to the question, does the test measure what it was intended to measure? Clearly, if the answer is no, we should not use the measure. Unfortunately, the answer to that question is not always obvious—face validity can be deceiving. Simply speaking, the centrality of reliability, and validity more broadly, cannot be overstated.
It also is important to note that the importance of these concepts is not linked to the specifics of the measurement. These issues are critical in high-stakes and low-stakes testing, assessments that produce numeric scores and assessments that result in text-based feedback, tests intended to measure individual differences and tests intended to measure an individual’s relative strengths, tests that break down complex constructs into components and tests that holistically evaluate performance, tests that are used for formative purposes, tests that are used for summative purposes, and tests that are used for both.
As we have described in this chapter, reliability and validity are the foundation of the psychometric requirement that the interpretations we make based on test results are supported by evidence. The place of evidence in the interpretation of test results—evidence-based measurement—is no different than the place of evidence in evidence-based medicine. In both cases, collecting the evidence is hard work, but it is essential work.
https://t.me/med1917
Annotated Bibliography
1. Kane MT. Validating the interpretations and uses of test scores. J Educ Measure. 2013 Mar 14;50:1-73. doi:10.1111/jedm.12000. This publication provides a current, in-depth discussion of Michael Kane’s approach to validity and validation, describing the validation process as a structured argument in support of the intended interpretations made based on test scores.
2. Clauser BE, Margolis MJ, Case SM. Testing for licensure and certification in the professions. In: Brennan RL, ed. Educational Measurement. 4th ed. American Council on Education/Praeger; 2006:701–731. This chapter provides an overview of assessment methods commonly used in licensure and certification examinations in the professions. It includes an expanded discussion of the evolution of validity and validation over the past century, and discusses additional applications of Kane’s framework to assessment methods commonly used in the professions.
3. Cook DA, Brydges R, Ginsburg S, Hatala R. A contemporary approach to validity arguments: a practical guide to Kane’s framework. Med Educ. 2015 Jun;49(6):560-
575. doi:10.1111/medu.12678. This paper provides a very readable introduction to Kane’s validity framework for both quantitative and qualitative assessment methods commonly used in medical education. As a part of the discussion, it highlights some of the parallels between validation work and evaluation of diagnostic studies in medicine.
https://t.me/med1917
4. Cook DA, Zendejas B, Hamstra SJ, Hatala R, Brydges R. What counts as validity evidence? Examples and prevalence in a systematic review of simulation-based assessment. Adv Health Sci Educ Theory Pract. 2014 May;19(2):233-250. doi:10.1007/s10459-013-9458-4.
5. Cook DA, Brydges R, Zendejas B, Hamstra SJ, Hatala R. Technology-enhanced simulation to assess health professionals: a systematic review of validity evidence, research methods, and reporting quality. Acad Med. 2013 Jun;88(6):872-883. doi:10.1097/ACM.0b013e31828ffdcf. Using the frameworks proposed by Messick and Kane, these systematic reviews summarize sources of validity evidence from studies of technology-enhanced simulation-based assessments, identifying methodological and reporting shortcomings and recommending directions for improvement in future research.
6. Hatala R, Cook DA, Brydges R, Hawkins R. Constructing a validity argument for the Objective Structured Assessment of Technical Skills (OSATS): a systematic review of validity evidence. Adv Health Sci Educ Theory Pract. 2015 Dec;20(5):1149-1175. doi:10.1007/s10459-015-9593-1. This systematic review uses Kane’s framework to analyze the validity argument for the objective structured assessment of technical skills (OSATS). They found that, in general, validity evidence supports the use of OSATS for formative feedback, but more research is required to support use of OSATS for making higher-stakes decisions.
7. Hawkins RE, Margolis MJ, Durning SJ, Norcini JJ. Constructing a validity argument for the mini-clinical
https://t.me/med1917
evaluation exercise: a review of the research. Acad Med. 2010 Sep;85(9):1453-1461. doi:10.1097/ACM.0b013e3181eac3e6. This systematic review of research conducted from 1995 to 2009 uses Kane’s validity framework to evaluate validity evidence related to the mini clinical evaluation exercise (mini-CEX). It concludes that scoring-related issues (e.g., leniency error and high interitem correlations) limit the utility of the mini-CEX for providing feedback to trainees, though evidence related to the generalization and extrapolation components is generally supportive of the validity of mini-CEX score interpretations.
References
1. Gulliksen H. Theory of Mental Tests. John Wiley & Sons;
1950.
2. Spearman C. Proof of the measurement of association between two things. Am J Psychol. 1904;15:72–101.
3. Spearman C. “General intelligence” objectively determined and measured. Am J Psychol. 1904;15:201–292.
4. Spearman C. Correlation calculated with faulty data. Br J Psychol. 1910;3:271–295.
5. Yoakum CS, Yerkes RM. Mental Tests in the American Army. Sidgwick & Jackson; 1920.
6. Cureton EE. Validity. In: Lindquist EF, ed. Educational Measurement. American Council on Education; 1951:621–
https://t.me/med1917
694.
7. Kane MT. Validating the interpretations and uses of test scores. J Educ Measure. 2013 Mar 14;50:1-73. doi:10.1111/jedm.12000
.
8. Messick S. Validity. In: Linn RL, ed. Educational Measurement. 3rd ed. American Council on Education/Macmillan; 1989:13–103.
9. Cronbach LJ, Meehl PE. Construct validity in psychological tests. Psych Bull. 1955 Jul;52(4):281–302. doi:10.1037/h0040957.
10. Cronbach LJ. Test validation. In: Thorndike RL, ed. Educational Measurement. 2nd ed. American Council on Education; 1971:443–507.
11. Campbell DT, Fiske DW. Convergent and divergent validation by the multitrait-multimethod matrix. Psych Bull. 1959 Mar;56(2):81–105.
12. Loevinger J. Objective tests as instruments of psychological theory. Psych Rep. 1957 Jun;3:635–694. doi:10.2466/pr0.1957.3.3.635.
13. American Educational Research Association, American Psychological Association, National Council on Measurement in Education. Standards for Educational and Psychological Testing. American Educational Research Association; 1999.
14. Kane M. An argument-based approach to validation.
https://t.me/med1917
Psych Bull. 1992;112(3):527–535. doi:10.1037/0033-
2909.112.3.527.
15. Hodges B. Assessment in the post-psychometric era: learning to love the subjective and collective. Med Teach. 2013 Jul;35(7):564–568. doi:10.3109/0142159X.2013.789134.
16. Schuwirth LWT, van der Vleuten CPM. A history of assessment in medical education. Adv Health Sci Educ. 2020 Dec;25(5):1045–1056. doi:10.1007/s10459-020-10003-0.
17. Swanson DB, van der Vleuten CP. Assessment of clinical skills with standardized patients: state of the art revisited. Teach Learn Med. 2013;25 Suppl 1:S17-S25. doi:10.1080/10401334.2013.842916.
18. Norcini J, Burch V. Workplace-based assessment as an educational tool: AMEE Guide No. 31. Med Teach. 2007 Nov;29(9-10):855-871. doi:10.1080/01421590701775453.
19. Hicks PJ, Margolis MJ, Carraccio C, et al. A novel workplace-based assessment for competency-based decisions and learner feedback. Med Teach. 2018 Nov;40(11):1143-1150. doi:10.1080/0142159X.2018.1461204.
20. Hicks PJ, Margolis MJ, Poynter SE, et al. The Pediatrics Milestones Assessment Pilot: development of workplace­based assessment content, instruments, and processes. Acad Med. 2016 May;91(5):701-709. doi:10.1097/ACM.0000000000001057.
21. Kelly FJ. The Kansas Silent Reading Test. The Kansas Printing Plant; 1915.
https://t.me/med1917
22. Swanson DB, Clauser BE, Case SM. Clinical skills assessment with standardized patients in high-stakes tests: a framework for thinking about score precision, equating, and security. Adv Health Sci Educ Theory Pract. 1999;4(1):67–106. doi:10.1023/A:1009862220473.
23. Lord FM, Novick MR. Statistical Theories of Mental Test Scores. Addison-Wesley; 1968.
24. Cronbach LJ, Gleser GC, Nanda H, Rajaratnam N. The Dependability of Behavioral Measurements: Theory of Generalizability for Scores and Profiles. John Wiley & Sons;
1972.
25. Brennan RL. Generalizability Theory. Springer-Verlag;
2001.
26. Brennan RL. An essay on the history and future of reliability from the perspective of replications. J Educ Meas. 2001;38(4):295–317. doi:10.1111/j.1745-3984.2001.tb01129.x.
27. Brown W. Some experimental results in the correlation of mental abilities. Br J Psych. 1910;3:296–322.
28. Kuder GF, Richardson MW. The theory of estimation of test reliability. Psychometrika. 1937;2:151–160. doi:10.1007/BF02288391.
29. Cronbach LJ. Coefficient alpha and the internal structure of tests. Psychometrika. 1951;16:297–334. doi:10.1007/BF02310555.
30. Angoff WH. Scales, norms, and equivalent scores. In:
https://t.me/med1917