Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_112_библиотеки_им_акад_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
05.09.2026
Размер:
18 Мб
Скачать
without narrative comments.
GME, Graduate medical education; ILP, Individualized learning plan; MCQs, multiple-choice questions; PBLI, Practice-based learning and improvement; SAQs, short-answer questions; SPs, Standardized patients.
To ensure comprehensive coverage, programs should complete an assessment map that includes where, when, and how the competency is being taught, experienced, and assessed (and must include more than one tool or method per competency). Assessment maps are a good way to ensure that all core competencies are being properly assessed (see Appendix 3.1 for an example using the US competency framework). The map also can help identify competencies receiving insufficient attention. In many locales around the world, that includes quality improvement and patient safety, interprofessional teamwork, care coordination such as daily handoffs and patient discharge, interpersonal communication skills, EBP, and
systems thinking.
43,44
Providing high-quality healthcare requires integrated contributions of all members of the interprofessional care team. Indeed, most care is provided by teams, not individuals acting in silos. However, the predominant lens for most assessment is attribution, not contribution. Attribution in assessment, meaning the results of an assessment can be clearly and predominantly attributed to the learner, may infer that the quality or safety of a clinical care process or outcome was only the direct result of a particular learner’s clinical skills. Contribution, in contrast, implies that the quality or safety of clinical care is the result of the interdependent contributions from all healthcare professions involved with the patient’s care. This creates a challenging situation where assessment is viewed as unfair if a specific patient care outcome cannot be predominantly attributed to the learner. Moving forward, assessment programs will have to pay
https://t.me/med1917
greater attention to how the learner, in collaboration with the team, contributed to the patient’s care and outcomes.45 Even in
competencies such as clinical reasoning, group involvement using distributed and situated cognition can lead to reductions in
diagnostic and therapeutic errors (see Chapter 8).46 Therefore, going forward, it will be important to discern not only the individual learner’s contributions to care, but also how that learner contributes to the effectiveness of their teams.
Rebalance the Focus of Programmatic Assessment to Emphasize More Work-Based Assessments
As noted earlier, assessment programs overemphasize high-stakes assessments, such as clerkship grades, single-event tests, and end-of­rotation summative faculty evaluations, that can undermine timely formative assessments that better support professional development. As a learner progresses along the continuum, assessment programs should shift the balance of assessments to WBAs that trained groups (e.g., CCCs) can integrate and synthesize into developmental judgments, such as Accreditation Council for Graduate Medical Education (ACGME) Milestones, the Royal College of Physicians and Surgeons of Canada’s (RCPSC’s) Entrustable Professional Activities (EPAs), and the Association of American Medical Colleges’ core
Entrustable Professional Activities (EPAs).
47–51
There are in fact multiple specialties that have created EPAs worldwide that can be a good place to start in designing what and where assessment programs should target their efforts and ensure that the EPAs sufficiently cover all competencies (see Chapter 1).
Higher-stakes WBA and non-WBA assessments (e.g., clerkship
grades or licensing examinations and objective structured clinical
https://t.me/med1917
examinations) will continue to have a role in holistic programmatic assessment. Use of these higher-stakes assessments should be fit-for­purpose and support professional development. Govaerts and colleagues recommended using “both/and” polarity thinking when designing local and national assessment systems.
7
For example, assessment programs must properly balance standardization and authenticity, quantitative and qualitative data, and WBA and non­WBA assessments. Van der Vleuten also cautioned against making firm distinctions between formative and summative, describing a formative-summative spectrum of stakes based on assessment purpose.
35
For example, a single, direct observation–based assessment should primarily focus on gathering accurate information for feedback, for coaching, and to provide a data point to the program for aggregation and synthesis. In contrast, a CCC making a higher-stakes graduation decision must use multiple, aggregated, and longitudinally generated assessments.
Embrace Narrative Assessment to Make Entrustment Decisions
Entrustment is quickly becoming a predominant assessment construct in medical education. As noted in Chapter 1
, ten Cate and colleagues extended Miller’s pyramid, placing “entrusted with future care” as the ultimate assessment decision. Entrustment has been operationalized as EPAs, defined as the routine professional-life activities of physicians based on their specialty and subspecialty, where entrustable means “a practitioner has demonstrated the necessary knowledge, skills, and attitudes to be trusted> to perform
this activity unsupervised”49 (see Chapter 1). In transitioning to EPAs, medical educators have created entrustment scales. Entrustment rating scale anchors are defined by either the amount of supervision the rater believes the learner will require in future
https://t.me/med1917
clinical encounters or the level of supervision and clinical care the rater (usually clinical faculty) contributed to the patient’s care during learner–patient interactions or procedures. Faculty report more satisfaction using entrustment scales because the scales align more closely with how they think about the learner’s ability, trustworthiness, and supervision needs
52
(see Chapter 4).
Despite faculty satisfaction with entrustment-based scales, they are no panacea for what ails more effective use of WBA. For decades, numbers, especially rating scales, have dominated faculty assessment approaches. However, numeric rating scales are nothing more than a
code that requires translation by the rater (and learner).
53
Simply stopping the assessment process at assigning a rating is inadequate if the data informing the rating cannot be effectively communicated to the learner in a way that the learner finds credible, actionable, trustworthy, and free of bias.
While studies have found evidence for improved reliability of and
satisfaction with entrustment scales, two recent studies have questioned their validity and accuracy.
54,55
For example, Schumacher and colleagues found no to little correlation between supervisors’ entrustment ratings and the quality of care pediatric residents
delivered in an emergency department.56 Kogan and colleagues found that faculty entrustment ratings on scripted videos of learners at different entrustment levels were highly variable, with greater variation in ratings occurring at the lowest level of ability in medical
interviewing and counseling.57 In sum, there is nothing magical about scales. While they enable statistical and psychometric analysis, these analyses still depend on the quality of the ratings provided on the scale. If the codes are not an accurate and valid reflection of the learner’s performance, the utility of the rating scale output (results) is greatly diminished. This is not meant to diminish the meaningful
(i.e., warranted) variation seen between faculty raters.58 As discussed
https://t.me/med1917
extensively in Chapter 5, raters bring variable levels of ability in assessment and clinical skills. When variable interpretations of performance by faculty are grounded in strong educational and clinical science, that can be a good thing for both the learner and patient. When ratings and interpretations are not, both the learner and patient can be adversely affected.
Additionally, an overemphasis on scales can undermine the importance of narrative assessments, which capture rich descriptions of what happened in an encounter, particularly when the assessment is grounded in evidence-based clinical and educational practice. For example, faculty can provide feedback on medical interviewing by referencing specific behaviors (e.g., agenda setting, active listening with silence, etc.) instead of the generic “good bedside manner” comment. Ginsburg and colleagues found that narrative comments can possess high levels of reliability. They noted that using written comments to discriminate between residents can be extremely reliable even after only several reports are collected. This suggests a way to identify residents early on who may require attention. These findings contribute evidence to support the validity argument for
using qualitative data for assessment.”
59,60
Prentice and colleagues in Australia found that flagging procedures (i.e., identifying learners at risk) using qualitative data was useful in identifying trainees in need
of interventions to help them.60 In addition, assessment programs can leverage robust and rigorous qualitative research methodologies designed to analyze qualitative data as part of group process such as
CCCs.61 The specific qualitative techniques and their correlates in assessment programs are provided in Table 3.3.
Table 3.3
A Model for Programmatic Assessment Fit for Purpose
Strategies to establish Criteria Potential assessment
https://t.me/med1917
trustworthiness strategy Credibility Prolonged
engagement
Train assessors People who know that the learner (coach, peers) best provides information for assessment Incorporate intermittent feedback cycles in the assessment procedure
Triangulation Involve many
assessors and different credible groups Use multiple sources of assessment within or across methods Organize a sequential judgment procedure where conflicting information necessitates the gathering of more information
Peer examination (sometimes called peer debriefing)
Assessors talk about benchmarking, the assessment process, and results before and at the halfway point during an activity Separate assessors’ multiple roles by
https://t.me/med1917
removing summative assessment decisions from the coaching role
Member checking
Incorporate the learner’s point of view in the assessment procedure (using coproduction) Incorporate longitudinal, intermittent feedback cycles
Structural coherence
Assessment committee (e.g., clinical competency committee) discusses inconsistencies in the assessment data and looks for signs of bias
Transferability Time
sampling
Sample broadly over different contexts and patients over time
Thick description (or dense description)
Assessment instruments facilitate inclusion of qualitative, narrative information Give narrative information a lot of weight in the assessment procedure
Dependability Stepwise Sample broadly over
https://t.me/med1917
replication different assessors
and over time
Dependability/confirmability Audit Document the
different steps in the assessment process (a formal assessment plan approved by an examination board, especially for high­stakes assessments; provide overviews of the results per phase) Quality assessment of procedures with external auditor Learners can appeal the assessment decision
From van der Vleuten CPM, Schuwirth LWT, Driessen EW, et al. A model for programmatic assessment fit for purpose. Med Teach. 2012;34(3):205-214. doi:10.3109/0142159X.2012.652239
.
Take Seriously and Address Unwarranted Variation
Unwarranted variation is an underappreciated problem in programmatic assessment.58 Some variation in assessments is good,
such as when a faculty member leverages a particular strength or aspect of effective practice (i.e., warranted variation). However, assessments driven by idiosyncrasies or bias that is no evidence based represent unwarranted variation (see Chapter 5). Unwarranted variation can be harmful, can contribute to suboptimal and variable educational outcomes, and, by extension, risks graduates delivering suboptimal healthcare. Bias is a particularly unwelcome form of unwarranted variation. Faculty are often the primary sources of
https://t.me/med1917
unwarranted variation, which we will address shortly, but poor group process is another source of unwarranted variation in programmatic assessment. There is abundant evidence that unwarranted variation at the program and institutional levels also can affect educational outcomes and the quality of care graduates
deliver.
62–69
Learners are nested within training programs (medical schools, residencies, fellowships, and other postgraduate training programs) that are nested within institutions that are nested within communities. These nested relationships are interdependent, and this interdependence impacts clinical and educational outcomes and can amplify competency deficiencies and structural biases with potentially profound effects. Warm and colleagues argued that optimizing assessment requires all participants within the system to clearly identify its assessment purpose, develop a deep and nuanced understanding of the variation occurring within the training program, and then use this information for improvement
interventions.
30
Programs can analyze a faculty member’s rating patterns and provide feedback on suboptimal and counterproductive assessment behaviors, such as straight-lining (i.e., all assessment items are given the same rating), leniency, stringency, halo rating errors, and bias (see Chapter 5), as part of the assessment program’s ongoing evaluation and continuous improvement efforts (see Chapter 18). All participants, including learners, should use all available data to identify and address the sources and causes of unwarranted variation in the training program, including the assessment
program.
30
Feedback loops across the continuum are also needed to support continuous quality improvement of assessment programs and reduce unwarranted variation and bias. As one example, UME programs
https://t.me/med1917
should assess how their graduates perform in residency (postgraduate) training as feedback about the effectiveness of the medical school program. Similarly, GME programs will need to leverage clinical performance measures in early practice as feedback on the effectiveness of residency and/or fellowship programs. More work is needed on how to access, interpret, and carefully use such data on graduates for continuous improvement of training programs. Use of clinical performance measures should be a part of program evaluation (see Chapter 18) with a clear-eyed understanding of the limitations of such measures. Undergraduate and postgraduate training programs are not factories all producing the same product. Therefore abundant caution and humility is needed to ensure that the focus of the assessment program is directed toward using meaningful measurement-type data fit for purpose. Patients being cared for by recent graduates deserve high-quality care, and where we have information that can provide useful insights to improve care we should use it. Several recent studies examining the relationship between Milestone ratings and early clinical practice provide some optimism for how these feedback loops could be created leveraging different clinical data sets despite the challenges in using clinical data
(see Chapter 11).
70–73
Regardless, all training programs, especially postgraduate training programs, need to use more clinical performance data as part of their assessment programs (see Chapter 11). As noted earlier, attribution of clinical performance, such as care for patients with chronic conditions or procedural complications, often cannot be directly or fully attributed to the actions and behaviors of the learner. Yet the learner contributed to the quality of care provided and the learner should leverage quality and safety measures for formative purposes as an important way to reduce unwarranted variation in clinical care
provided by training programs.
56,74
https://t.me/med1917