Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_112_библиотеки_им_акад_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
05.09.2026
Размер:
18 Мб
Скачать
not eliminate poor accuracy, particularly when assessing learners with poor performance.99 Even with entrustment ratings, raters are the largest source of variability.
94,100
In sum, the optimal choice of scale anchors is still not settled although the trend is moving toward entrustment scales. What is clear, however, is that regardless of which assessment tool and scale anchors are selected, the assessor will account for much of the variance in ratings. Therefore faculty training is essential to successfully implement WBA.
Narrative Comments
Narrative assessment is becoming an increasingly important component of WBA.
101,102
Many WBA instruments have space for narrative comments. WBA that includes a space for narrative comments or prompts verbal feedback provides learners with rich information about how they can improve (more so than numbers). Narrative comments can both explain and elaborate checklist scores, contextualize the assessment, and convey unique content. Work by Ginsberg and colleagues has described the reliability and validity of
narrative assessments.
103
Learners too describe how narrative
feedback is more helpful than receiving a rating.
104,105
Fig. 5.5 shows an example of a narrative assessment form that also includes an entrustment rating. We recommend that narrative assessments are made before selecting a numeric rating so that the rating can be informed by the narrative (Fig. 5.6; see Chapter 3 for more detail on managing bias).
https://t.me/med1917
https://t.me/med1917
FIG. 5.5 Revised rater assessment form.
https://t.me/med1917
FIG. 5.6 Rethinking the assessment process.
Issues With Reliability and Validity
Despite the importance of WBA, the quality of assessment remains challenged by the reliability and validity issues described in the following subsections.
Accuracy
A significant problem with WBA is poor accuracy of faculty ratings. For example, faculty do not detect up to 68% of errors committed by a resident scripted to depict marginal
performance on a training video.
80,99,106
Use of specific checklists that prompt faculty to look for certain skills can increase error detection accuracy nearly twofold but do not
produce more accurate overall ratings of competence.
80
Nearly 70% of faculty may still rate a resident depicting marginal performance as satisfactory or superior.80 Kalet and
colleagues examined the reliability and validity of faculty observations using videos of student performance on an objective structured clinical examination (OSCE) evaluating
interviewing skills.
107
Faculty inconsistently identified the use of open-ended questions and empathy, and the positive predictive value of faculty ratings for “adequate” interviewing skills was only 12%. Another study found that faculty members could not reliably evaluate 32% of the physical examination skills assessed and had the most difficulty with examination of the head, neck, and
abdomen.
108
https://t.me/med1917
Variable Assessment Standards
There is an increasing body of literature exploring rater cognition: how faculty observe, synthesize their observations,
and make assessment judgments.
109,110
The rater cognition literature has shown that faculty use variable standards against which they compare the learner. Furthermore, assessors often develop their assessment criteria experientially and idiosyncratically and interpret scale
anchors variably.
111,112
Different faculty subsequently focus
on different aspects of performance and have variable definitions of quality and competence.
111–113
When assessors use different standards to assess learners, they rate learners differently, decreasing interrater reliability.
Common standards are normative, self, and use of gestalt. Some faculty use a normative standard and compare the learner to learners at a similar level of training. However, faculty are often uncertain about what skills are expected at a
given stage of training.
111,112,114
More commonly, faculty use
themselves as the standard when evaluating a learner.
111
This can be problematic given the robust literature describing the shortcomings practicing physicians can have in their clinical
skills.
28,31–33,38–41,115
Medical educators with clinical skill deficiencies may be less likely to detect those skill deficits in learners. This is supported by a small study showing that faculty who were more complete in their history taking and had better interpersonal skills, as rated by a standardized patient, were more likely to be stringent when evaluating
residents.
116
An important limitation to the study was the use of standardized patients to assess faculty skills since standardized patient assessments often reward completeness
https://t.me/med1917
over the use of sophisticated illness scripts and efficiency.
14
Nevertheless, another study found that assessors often feel they lack content expertise in the skills they are being asked
to assess.
114
Another standard used in WBA is the use of gestalt. Many faculty say that their assessment decisions are based on gestalt and they are unable to articulate how they
arrived at their assessment decision(s).
111
While gestalt assessments can be accurate, they are problematic when faculty cannot deconstruct their gestalt to provide the learner with specific, constructive feedback. In contrast to the common use of normative standards, self as standard, and gestalt, faculty rarely use a criterion-referenced standard. When faculty use a criterion-referenced standard, they compare what the learner has done to the evidence-based
best practice components for that particular skill.
111
An example of using a criterion-referenced approach would be assessing which of nine shared decision-making behaviors a resident performed when counseling a patient about starting a new medication.
Role of Inference
Another source of WBA error is when faculty make inferences during observation. Inference occurs when observers derive what seems to be logical conclusions from premises that are assumed to be true rather than assessing
observable behaviors.
111
,117
For example, imagine a resident standing with their arms crossed breaking bad news to a patient. Rather than note that the resident had their arms crossed (the observable behavior), some faculty may conclude the resident does not know how to break bad news (inference). Other faculty may believe the resident lacks
https://t.me/med1917
empathy and humanism or might assume the resident is uncomfortable having never broken bad news before (also inference). Assessors make inferences about learners’ knowledge, skills (competence), and attitudes (work ethic, emotions, intentions, personality).
111
,117
Assessors often do not recognize when they are making inferences and rarely validate the accuracy of their inferences.
111
Unchecked inferences risk “distorting” accurate assessment of the learner because the assessor’s inferences cannot be observed and measured; this leads to greater interassessor variability and ultimately faulty assessment.
Impact of Educational Culture
While some faculty embrace their roles and responsibilities as assessors and coaches and do not shy away from giving low
ratings,
111
others modify their assessments to avoid
unpleasant repercussions (both for themselves and their learners). This is another source of interrater variability.
111
Medicine has a “failure to fail” culture. As a result, some faculty inflate ratings to avoid discussing a marginal rating
with a learner.
118,119
Other faculty may inflate assessments to
be seen as a popular, likable teacher.
111
Some faculty avoid
stringent assessments so that they will not have to defend their assessments with institutional leaders.
111,120,121
Rater Limitations and Bias
Limitations in human cognition may also explain variability in WBA.85 For example, contrast effects occur when the
scores faculty give to learners are influenced by their recent observations of other learners. Faculty may rate a learner
https://t.me/med1917
with borderline skills lower if they are observed after a learner with excellent skills; conversely, the same learner with borderline skills may be rated higher if the faculty
observed them after a learner who is less skilled.
122,123
These
biases affect both numerical ratings and narrative feedback.
124,125
In addition to contrast effects and recency bias, there is an urgent need to confront and address the persistent and pernicious effects of bias in assessment. Explicit bias refers to conscious beliefs and attitudes one possesses about another person or group. Implicit bias refers to an individual’s “prejudicial attitudes toward and stereotypical beliefs about a
particular social group or members therein.”
126
These individual attitudes are often subconscious. A growing body of literature describes the factors contributing to, and the harmful effects of, implicit bias and educational inequity on learners who are underrepresented in medicine (URiM) or otherwise at risk for marginalization in assessment practices. Multiple studies have found that learners from groups historically URiM receive lower assessment ratings from
faculty.
127–129
In addition, gender-based differences have
repeatedly been demonstrated in faculty assessments of medical trainees.
130–135
These differences have been observed in quantitative measurements where women residents receive lower competency ratings despite no differences in
performance.
136–138
Narrative assessments often include
linguistic differences and describe personality traits that differ by gender and URiM status.
134,135,139–141
Conflicting priorities, time constraints, and burnout can lower the threshold for faculty to unconsciously apply their activated biases when they assess learners. Fig. 5.6 provides
https://t.me/med1917
suggestions on where raters can use self-awareness and self­assessment for bias in the observation and rating process.
Conceptual Model Summarizing Rater Cognition
Fig. 5.7 is a conceptual model of the factors that influence
faculty’s WBA.
111
This model highlights the complexity of direct observation and WBA. Letter labels direct you to the corresponding part of the figure. When faculty observe, they bring their attitudes and emotions about observation and feedback, clinical competence (i.e., their own clinical skills in the domain they are assessing the learner), educational competence (i.e., their skills in observation, assessment, and feedback as well as their biases), and traits (e.g., age, gender, clinical, and teaching experience) (A). Faculty observe learner–patient interactions through two lenses (B). While observing, they make inferences and use variable frames of reference (normative, self, criterion-referenced, gestalt). After the observation, faculty must interpret and synthesize their observations (C) to select a numerical rating or create a narrative assessment (D). Sometimes faculty modify their assessments in anticipation of the feedback they need to give (E). For example, faculty may select a higher rating to avoid discussing a low rating with the learner. Experiential modifiers can further influence faculty’s observations and their interpretation and synthesis of observations and feedback (F). For example, prior experiences working with a learner (e.g., learner’s prior performance, receptivity to feedback, etc.) can influence faculty assessment and feedback. Finally, observations are influenced by the clinical system (G) (e.g., familiarity with the patient, patient complexity, system
https://t.me/med1917
factors such as multiple patients waiting to be seen) and the institution’s educational culture (e.g., assessment culture, culture of oversight/supervision) (H).
FIG. 5.7 Conceptual model: process of direct observation (see
text for details).
Improving the Quality of Direct Observation Assessments
In this section of the chapter, we describe the need for faculty development to address some of the aforementioned factors that negatively influence WBA. Before doing so, we recognize that some educators do not believe interrater variability is problematic and embrace the variable
perspectives of multiple raters.
7,110,142,143
Educators holding this perspective believe that assessment is enriched by the unique perspectives of multiple evaluators whose assessments, when taken together, can provide meaningful
https://t.me/med1917