Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_2823_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
05.09.2026
Размер:
18 Мб
Скачать
Thorndike RL, ed. Educational Measurement. 2nd ed. American Council on Education; 1971:508–600.
31. Hambleton RK, Swaminathan H. Item Response Theory: Principles and Applications. Kluwer-Nijhoff Publishing;
1985.
32. Kolen MJ, Brennan RL. Test Equating, Scaling, and Linking. Methods and Practices. 3rd ed. Springer; 2014.
33. Cronbach LJ. My current thoughts on coefficient alpha and successor procedures. Educ Psychol Measure. 2004;64(3):391-418. doi:10.1177/0013164404266386.
34. Ericsson A, Smith J. Prospects and limitations of the empirical study of expertise: an introduction. In: Ericsson A, Smith J. Toward a General Theory of Expertise: Prospects and Limits. Cambridge University Press; 1991.
35. Zapata-Rivera D. Score Reporting Research and Applications. Routledge; 2018.
36. Kelley TL. Interpretation of Educational Measurements. World Book; 1927.
37. Lord FM. Elementary models for measuring change. In: Harris CW, ed. Problems in Measuring Change. Wisconsin Press; 1963.
38. Norcini J, Anderson B, Bollela V, et al. Criteria for good assessment: consensus statement and recommendations from the Ottawa 2010 Conference. Med Teach. 2011;33(3):206-214. doi:10.3109/0142159X.2011.551559.
https://t.me/med1917
39. Norcini J, Anderson MB, Bollela V, et al. Consensus framework for good assessment. Med Teach. 2018 Nov;40(11):1102-1109. doi:10.1080/0142159X.2018.1500016.
40. Price D, Swanson DB, Irons M, Hawkins RE. Longitudinal assessments in continuing specialty certification and lifelong learning. Med Teach. 2018 Sep;40(9):917-919. doi:10.1080/0142159X.2018.1471202.
41. Clauser BE, Margolis MJ, Case SM. Testing for licensure and certification in the professions. In: Brennan RL, ed. Educational Measurement. 4th ed. American Council on Education/Praeger; 2006:701–731.
42. Cronbach LJ. Validity on parole: how can we go straight? New directions for testing and measurement: measuring achievement over a decade. Proceedings of the 1979 ETS Invitational Conference. Jossey-Bass; 1980:99–108.
43. Norori N, Hu Q, Aellen FM, Faraci FD, Tzovara A. Addressing bias in big data and AI for health care: A call for open science. Patterns (N Y). 2021 Oct 8;2(10):100347. doi:10.1016/j.patter.2021.100347
44. Cuddy MM, Dillon GF, Clauser BE, et al. Assessing the validity of the USMLE Step 2 Clinical Knowledge Examination through an evaluation of its clinical relevance. Acad Med. 2004 Oct;79(10 Suppl):S43–S45. doi:10.1097/00001888-200410001-00013.
45. Cuddy MM, Young A, Gelman A, et al. Exploring the relationship among USMLE performance and disciplinary
https://t.me/med1917
action in practice: a validity study of score inferences from a licensure examination. Acad Med. 2017 Dec;92(12):1780-1785. doi:10.1097/ACM.0000000000001747.
46. Norcini JJ, Lipner RS, Kimball HR. Certifying examination performance and patient outcomes following acute myocardial infarction. Med Educ. 2002 Sep;36(9):853-
859. doi:10.1046/j.1365-2923.2002.01293.x.
47. Norcini JJ, Boulet JR, Opalek A, Dauphinee WD. The relationship between licensing examination performance and the outcomes of care by international medical school graduates. Acad Med. 2014 Aug;89(8):1157-1162. doi:10.1097/ACM.0000000000000310.
48. Tamblyn R, Abrahamowicz M, Dauphinee WD, et al. Association between licensure examination scores and practice in primary care. JAMA. 2002 Dec 18;288(23):3019–
3026. doi:10.1001/jama.288.23.3019.
49. Swanson DB, Roberts TE. Trends in national licensing examinations in medicine. Med Educ. 2016 Jan;50(1):101-114. doi:10.1111/medu.12810.
50. Harik P, Feinberg RA, Clauser BE. How examinees use time: examples from a medical licensing examination. In: Margolis MJ, Feinberg RA, eds. Integrating Timing Considerations to Improve Standardized Testing Practices. Routledge; 2020: 73–89. doi:10.4324/9781351064781-6
51. Margolis MJ, Clauser BE, Cuddy MM, et al. Use of the mini-CEX to rate examinee performance on a multiple-
https://t.me/med1917
station clinical skills examination: a validity study. Acad Med. 2006 Oct;81(10 Suppl):S56–S60. doi:10.1097/01.ACM.0000236514.53194.f4.
52. Kane M. Validating the performance standards associated with passing scores. Rev Educ Res. 1994;64:425–461.
* Throughout this chapter we argue that it is important to collect a range of evidence to evaluate the credibility of interpretations that are to be made based on test scores. We share the view of Cronbach, Messick, and Kane that this is likely to require a program of research. At the same time, it is clear that issues of practicality come into play. Although a test that contributes to a grade in a single class or clerkship may raise the same validity issues as a national licensing examination, the resources available to evaluate validity will be far greater in the latter context than in the former. In the case of the licensing examination, an extensive program of research certainly will be appropriate; in the case of a classroom test, the evaluation may be much more limited. That said, when educators introduce novel testing formats, they have a significant responsibility to provide empirical justification for the associated score interpretations and uses.
†High-stakes testing refers to situations in which the outcome of the test has important consequences for the examinee. In medical education, admissions tests and tests for licensing and certification have very high stakes. Tests that result in grades or pass/fail decisions also can be considered high stakes. A self-assessment would be considered low stakes.
https://t.me/med1917
‡In generalizability theory terminology, sources of variability —such as the sampling of items or judges—are referred to as facets. Facets are similar to factors used in analysis of variance.
§A wide variety of procedures are in use for putting scores from different forms of the same test on a common scale. The simplest of these approaches is to administer the two test forms to the same group or to randomly equivalent groups of examinees and set the mean (or mean and standard deviation) for the two forms to be equal.30 More sophisticated approaches include item response theory31 and equipercentile equating.32 Each approach is designed to minimize the differences in difficulty across test forms that are reflected in the item variance component.
**This is a relatively infrequent occurrence for large-scale OSCEs. Even if all examinees rotate through the same set of stations, multiple “circuits” with different standardized patients portraying case roles and different raters grading performance are commonly used, and this adversely affects precision.
21
††In some areas of practice (e.g., emergency medicine, trauma surgery), speed of response may be critical in practice. In these areas it may seem attractive to include speed of response as part of the assessment. When this is done, it will be appropriate to collect evidence linking response speed on the test to response speed in practice.
https://t.me/med1917
3
Programmatic Assessment Using Systems Thinking
Eric S. Holmboe, MD
https://t.me/med1917
Chapter Outline
Overview of Chapter Overview of Programmatic Assessment
Three Overarching Principles to Improve Programmatic
Assessment Importance of Groups in Programmatic Assessment Importance of Longitudinal Design Thinking in
Programmatic Assessment
Specific Opportunities to Improve Programmatic Assessment
Gain Clarity Around the “Why” of Assessment Ensure Comprehensive and Fair Coverage of All Core
Competencies Rebalance the Focus of Programmatic Assessment to
Emphasize More Work-Based Assessments Embrace Narrative Assessment to Make Entrustment
Decisions Take Seriously and Address Unwarranted Variation Use Learning Analytics and Big Data Explicitly Define Assessors’ Roles and Responsibilities in
the Assessment System Address Bias in Assessment Recognize the Importance of Coproduction With Learners
Putting It All Together: Implementation Science and Programmatic Assessment Conclusion Acknowledgments References
https://t.me/med1917
Overview of Chapter
The introduction of competency-based medical education (CBME) has catalyzed advancements in assessment as highlighted throughout this textbook. Yet much work remains to be done to realize the full promise of CBME. Programmatic assessment provides the information needed for feedback, to support coaching, determine appropriate supervision levels, support the creation of individualized learning plans, inform progress decisions, and most importantly, help ensure that patients and families receive high­quality, safe care in the training environment and the future practice setting of graduates. Many training programs have not fully implemented programmatic assessment that includes assessment of all the key core competencies needed for effective clinical practice, such as a systems view of professionalism, interprofessional teamwork, quality improvement and patient safety, care coordination, systems thinking, and evidence-based practice (EBP), to name a few.
It is important to distinguish programmatic assessment from a simple, nonintegrated program of assessment. Programmatic assessment requires a systematic, longitudinal approach (i.e., systems thinking) in its design and execution. A key feature of programmatic assessment is the systematic approach to how data are synthesized and combined for purposes of judgment, feedback, and support of professional development. In many typical programs of assessment, data are often combined only when they have the same format to produce a score for a single competency domain or category of assessment (e.g., one objective structured clinical exam [OSCE] station with another OSCE station to produce a mean score or average outcome, or results from in-training multiple-choice exams).
Programmatic assessment seeks to synthesize, combine, and triangulate information so that it meaningfully contributes to a more
https://t.me/med1917
holistic understanding of where the learner is in their journey and what they and the program can do with the learner to best support their longitudinal professional development. The term programmatic explicitly incorporates systems and developmental thinking. Finally, I want to briefly distinguish programmatic assessment from program evaluation (see Chapter 18
). Program evaluation involves analyzing the performance of the educational system, training program, and all its interacting parts needed for changes in curriculum, assessment practices, and the learning environment; an educational system is made up of more than just its individual learners. Programmatic assessment of individual learners is thus one critical component of program evaluation. Many other sources of information will be important in formulating judgments about program quality and informing change and improvement strategies in the training program (see Chapter 18).
The chapters that follow will address specific assessment approaches needed to assess all the competencies (i.e., abilities) required for mastery in clinical practice. This chapter will guide the reader in how to improve assessment practices and assemble all the assessment components, or “parts,” into an integrated and effective programmatic assessment approach to produce accurate entrustment decisions supported by strong validity evidence, and addresses the pernicious effects of bias. Although we have discussed important potential differences in terminology, “programs of assessment” will be interchangeable in this chapter with “programmatic assessment” with the understanding that both refer to a systematic, integrated, and longitudinal approach.
Overview of Programmatic Assessment
https://t.me/med1917
The growing recognition of serious deficiencies in healthcare in the late 20th century led to an examination of medical education’s role in healthcare system performance. One result of this examination was to pressure the medical education enterprise to shift the focus and design of medical education programs to be outcomes based. However, uptake of an outcomes-based approach has varied around the world based on local needs and culture, and this is an important consideration for anyone implementing programmatic assessment.
Competency-based models are the primary approach to implementing outcomes-based medical education. Competencies, defined as “observable abilities of a health professional, integrating multiple components such as knowledge, skills, values and attitudes,” are the predominant framework used to define learners’
educational outcomes (see Chapter 1).1 CBME depends on effective programmatic assessment to achieve the desired outcome goals of training. Using a systems lens, programmatic assessment can be defined as a group of integrated and related assessment activities (or methods) that are managed in a coordinated manner. These interdependent activities have a common goal or success “vision” under integrated management and are embedded in systems. Systems thinking acknowledges that the components and activities of an assessment program are interdependent and must work together to accomplish the shared aims of a training program.
Effective programmatic assessment, using a systems perspective, most importantly involves a group of people who work together on a
regular basis to perform assessment and provide feedback to a population of trainees over a period of time, and share:
2
• Educational goals and outcomes
• Linked individual learner assessments and program evaluation processes (see Chapter 18)
• Information about learner performance to support professional
https://t.me/med1917