Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_2823_Библиотеки_им_академика_М_И_Перельмана
.pdf
Thorndike RL, ed. Educational Measurement. 2nd ed.
American Council on Education; 1971:508–600.
31. Hambleton RK, Swaminathan H. Item Response Theory:
Principles and Applications. Kluwer-Nijhoff Publishing;
1985.
32. Kolen MJ, Brennan RL. Test Equating, Scaling, and
Linking. Methods and Practices. 3rd ed. Springer; 2014.
33. Cronbach LJ. My current thoughts on coefficient alpha
and successor procedures. Educ Psychol Measure.
2004;64(3):391-418. doi:10.1177/0013164404266386.
34. Ericsson A, Smith J. Prospects and limitations of the
empirical study of expertise: an introduction. In: Ericsson A,
Smith J. Toward a General Theory of Expertise: Prospects
and Limits. Cambridge University Press; 1991.
35. Zapata-Rivera D. Score Reporting Research and
Applications. Routledge; 2018.
36. Kelley TL. Interpretation of Educational Measurements.
World Book; 1927.
37. Lord FM. Elementary models for measuring change. In:
Harris CW, ed. Problems in Measuring Change. Wisconsin
Press; 1963.
38. Norcini J, Anderson B, Bollela V, et al. Criteria for good
assessment: consensus statement and recommendations from
the Ottawa 2010 Conference. Med Teach. 2011;33(3):206-214.
doi:10.3109/0142159X.2011.551559.
https://t.me/med1917

39. Norcini J, Anderson MB, Bollela V, et al. Consensus
framework for good assessment. Med Teach. 2018
Nov;40(11):1102-1109. doi:10.1080/0142159X.2018.1500016.
40. Price D, Swanson DB, Irons M, Hawkins RE.
Longitudinal assessments in continuing specialty
certification and lifelong learning. Med Teach. 2018
Sep;40(9):917-919. doi:10.1080/0142159X.2018.1471202.
41. Clauser BE, Margolis MJ, Case SM. Testing for licensure
and certification in the professions. In: Brennan RL, ed.
Educational Measurement. 4th ed. American Council on
Education/Praeger; 2006:701–731.
42. Cronbach LJ. Validity on parole: how can we go straight?
New directions for testing and measurement: measuring
achievement over a decade. Proceedings of the 1979 ETS
Invitational Conference. Jossey-Bass; 1980:99–108.
43. Norori N, Hu Q, Aellen FM, Faraci FD, Tzovara A.
Addressing bias in big data and AI for health care: A call for
open science. Patterns (N Y). 2021 Oct 8;2(10):100347.
doi:10.1016/j.patter.2021.100347
44. Cuddy MM, Dillon GF, Clauser BE, et al. Assessing the
validity of the USMLE Step 2 Clinical Knowledge
Examination through an evaluation of its clinical relevance.
Acad Med. 2004 Oct;79(10 Suppl):S43–S45.
doi:10.1097/00001888-200410001-00013.
45. Cuddy MM, Young A, Gelman A, et al. Exploring the
relationship among USMLE performance and disciplinary
https://t.me/med1917

action in practice: a validity study of score inferences from a
licensure examination. Acad Med. 2017 Dec;92(12):1780-1785.
doi:10.1097/ACM.0000000000001747.
46. Norcini JJ, Lipner RS, Kimball HR. Certifying
examination performance and patient outcomes following
acute myocardial infarction. Med Educ. 2002 Sep;36(9):853-
859. doi:10.1046/j.1365-2923.2002.01293.x.
47. Norcini JJ, Boulet JR, Opalek A, Dauphinee WD. The
relationship between licensing examination performance and
the outcomes of care by international medical school
graduates. Acad Med. 2014 Aug;89(8):1157-1162.
doi:10.1097/ACM.0000000000000310.
48. Tamblyn R, Abrahamowicz M, Dauphinee WD, et al.
Association between licensure examination scores and
practice in primary care. JAMA. 2002 Dec 18;288(23):3019–
3026. doi:10.1001/jama.288.23.3019.
49. Swanson DB, Roberts TE. Trends in national licensing
examinations in medicine. Med Educ. 2016 Jan;50(1):101-114.
doi:10.1111/medu.12810.
50. Harik P, Feinberg RA, Clauser BE. How examinees use
time: examples from a medical licensing examination. In:
Margolis MJ, Feinberg RA, eds. Integrating Timing
Considerations to Improve Standardized Testing Practices.
Routledge; 2020: 73–89. doi:10.4324/9781351064781-6
51. Margolis MJ, Clauser BE, Cuddy MM, et al. Use of the
mini-CEX to rate examinee performance on a multiple-
https://t.me/med1917

station clinical skills examination: a validity study. Acad
Med. 2006 Oct;81(10 Suppl):S56–S60.
doi:10.1097/01.ACM.0000236514.53194.f4.
52. Kane M. Validating the performance standards associated
with passing scores. Rev Educ Res. 1994;64:425–461.
* Throughout this chapter we argue that it is important to
collect a range of evidence to evaluate the credibility of
interpretations that are to be made based on test scores. We
share the view of Cronbach, Messick, and Kane that this is
likely to require a program of research. At the same time, it is
clear that issues of practicality come into play. Although a
test that contributes to a grade in a single class or clerkship
may raise the same validity issues as a national licensing
examination, the resources available to evaluate validity will
be far greater in the latter context than in the former. In the
case of the licensing examination, an extensive program of
research certainly will be appropriate; in the case of a
classroom test, the evaluation may be much more limited.
That said, when educators introduce novel testing formats,
they have a significant responsibility to provide empirical
justification for the associated score interpretations and uses.
†High-stakes testing refers to situations in which the
outcome of the test has important consequences for the
examinee. In medical education, admissions tests and tests
for licensing and certification have very high stakes. Tests
that result in grades or pass/fail decisions also can be
considered high stakes. A self-assessment would be
considered low stakes.
https://t.me/med1917

‡In generalizability theory terminology, sources of variability
—such as the sampling of items or judges—are referred to as
facets. Facets are similar to factors used in analysis of
variance.
§A wide variety of procedures are in use for putting scores
from different forms of the same test on a common scale. The
simplest of these approaches is to administer the two test
forms to the same group or to randomly equivalent groups
of examinees and set the mean (or mean and standard
deviation) for the two forms to be equal.30 More
sophisticated approaches include item response theory31 and
equipercentile equating.32 Each approach is designed to
minimize the differences in difficulty across test forms that
are reflected in the item variance component.
**This is a relatively infrequent occurrence for large-scale
OSCEs. Even if all examinees rotate through the same set of
stations, multiple “circuits” with different standardized
patients portraying case roles and different raters grading
performance are commonly used, and this adversely affects
precision.
21
††In some areas of practice (e.g., emergency medicine,
trauma surgery), speed of response may be critical in
practice. In these areas it may seem attractive to include
speed of response as part of the assessment. When this is
done, it will be appropriate to collect evidence linking
response speed on the test to response speed in practice.
https://t.me/med1917

3
Programmatic Assessment Using
Systems Thinking
Eric S. Holmboe, MD
https://t.me/med1917

Chapter Outline
Overview of Chapter
Overview of Programmatic Assessment
Three Overarching Principles to Improve Programmatic
Assessment
Importance of Groups in Programmatic Assessment
Importance of Longitudinal Design Thinking in
Programmatic Assessment
Specific Opportunities to Improve Programmatic
Assessment
Gain Clarity Around the “Why” of Assessment
Ensure Comprehensive and Fair Coverage of All Core
Competencies
Rebalance the Focus of Programmatic Assessment to
Emphasize More Work-Based Assessments
Embrace Narrative Assessment to Make Entrustment
Decisions
Take Seriously and Address Unwarranted Variation
Use Learning Analytics and Big Data
Explicitly Define Assessors’ Roles and Responsibilities in
the Assessment System
Address Bias in Assessment
Recognize the Importance of Coproduction With Learners
Putting It All Together: Implementation Science and
Programmatic Assessment
Conclusion
Acknowledgments
References
https://t.me/med1917

Overview of Chapter
The introduction of competency-based medical education (CBME)
has catalyzed advancements in assessment as highlighted
throughout this textbook. Yet much work remains to be done to
realize the full promise of CBME. Programmatic assessment provides
the information needed for feedback, to support coaching, determine
appropriate supervision levels, support the creation of
individualized learning plans, inform progress decisions, and most
importantly, help ensure that patients and families receive highquality, safe care in the training environment and the future practice
setting of graduates. Many training programs have not fully
implemented programmatic assessment that includes assessment of
all the key core competencies needed for effective clinical practice,
such as a systems view of professionalism, interprofessional
teamwork, quality improvement and patient safety, care
coordination, systems thinking, and evidence-based practice (EBP),
to name a few.
It is important to distinguish programmatic assessment from a
simple, nonintegrated program of assessment. Programmatic
assessment requires a systematic, longitudinal approach (i.e., systems
thinking) in its design and execution. A key feature of programmatic
assessment is the systematic approach to how data are synthesized
and combined for purposes of judgment, feedback, and support of
professional development. In many typical programs of assessment,
data are often combined only when they have the same format to
produce a score for a single competency domain or category of
assessment (e.g., one objective structured clinical exam [OSCE]
station with another OSCE station to produce a mean score or
average outcome, or results from in-training multiple-choice exams).
Programmatic assessment seeks to synthesize, combine, and
triangulate information so that it meaningfully contributes to a more
https://t.me/med1917

holistic understanding of where the learner is in their journey and
what they and the program can do with the learner to best support
their longitudinal professional development. The term programmatic
explicitly incorporates systems and developmental thinking. Finally,
I want to briefly distinguish programmatic assessment from program
evaluation (see Chapter 18
). Program evaluation involves analyzing
the performance of the educational system, training program, and all
its interacting parts needed for changes in curriculum, assessment
practices, and the learning environment; an educational system is
made up of more than just its individual learners. Programmatic
assessment of individual learners is thus one critical component of
program evaluation. Many other sources of information will be
important in formulating judgments about program quality and
informing change and improvement strategies in the training
program (see Chapter 18).
The chapters that follow will address specific assessment
approaches needed to assess all the competencies (i.e., abilities)
required for mastery in clinical practice. This chapter will guide the
reader in how to improve assessment practices and assemble all the
assessment components, or “parts,” into an integrated and effective
programmatic assessment approach to produce accurate entrustment
decisions supported by strong validity evidence, and addresses the
pernicious effects of bias. Although we have discussed important
potential differences in terminology, “programs of assessment” will
be interchangeable in this chapter with “programmatic assessment”
with the understanding that both refer to a systematic, integrated,
and longitudinal approach.
Overview of Programmatic
Assessment
https://t.me/med1917

The growing recognition of serious deficiencies in healthcare in the
late 20th century led to an examination of medical education’s role in
healthcare system performance. One result of this examination was
to pressure the medical education enterprise to shift the focus and
design of medical education programs to be outcomes based.
However, uptake of an outcomes-based approach has varied around
the world based on local needs and culture, and this is an important
consideration for anyone implementing programmatic assessment.
Competency-based models are the primary approach to
implementing outcomes-based medical education. Competencies,
defined as “observable abilities of a health professional, integrating
multiple components such as knowledge, skills, values and
attitudes,” are the predominant framework used to define learners’
educational outcomes (see Chapter 1).1 CBME depends on effective
programmatic assessment to achieve the desired outcome goals of
training. Using a systems lens, programmatic assessment can be
defined as a group of integrated and related assessment activities (or
methods) that are managed in a coordinated manner. These
interdependent activities have a common goal or success “vision”
under integrated management and are embedded in systems. Systems
thinking acknowledges that the components and activities of an
assessment program are interdependent and must work together to
accomplish the shared aims of a training program.
Effective programmatic assessment, using a systems perspective,
most importantly involves a group of people who work together on a
regular basis to perform assessment and provide feedback to a population of
trainees over a period of time, and share:
2
• Educational goals and outcomes
• Linked individual learner assessments and program evaluation
processes (see Chapter 18)
• Information about learner performance to support professional
https://t.me/med1917
Соседние файлы в папке Библиотека им академика М.И. Перельмана
