Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_112_библиотеки_им_акад_М_И_Перельмана
.pdf
not eliminate poor accuracy, particularly when assessing
learners with poor performance.99 Even with entrustment
ratings, raters are the largest source of variability.
94,100
In
sum, the optimal choice of scale anchors is still not settled
although the trend is moving toward entrustment scales.
What is clear, however, is that regardless of which
assessment tool and scale anchors are selected, the assessor
will account for much of the variance in ratings. Therefore
faculty training is essential to successfully implement WBA.
Narrative Comments
Narrative assessment is becoming an increasingly important
component of WBA.
101,102
Many WBA instruments have
space for narrative comments. WBA that includes a space for
narrative comments or prompts verbal feedback provides
learners with rich information about how they can improve
(more so than numbers). Narrative comments can both
explain and elaborate checklist scores, contextualize the
assessment, and convey unique content. Work by Ginsberg
and colleagues has described the reliability and validity of
narrative assessments.
103
Learners too describe how narrative
feedback is more helpful than receiving a rating.
104,105
Fig. 5.5
shows an example of a narrative assessment form that also
includes an entrustment rating. We recommend that
narrative assessments are made before selecting a numeric
rating so that the rating can be informed by the narrative
(Fig. 5.6; see Chapter 3 for more detail on managing bias).
https://t.me/med1917

https://t.me/med1917

FIG. 5.5 Revised rater assessment form.
https://t.me/med1917

FIG. 5.6 Rethinking the assessment process.
Issues With Reliability and Validity
Despite the importance of WBA, the quality of assessment
remains challenged by the reliability and validity issues
described in the following subsections.
Accuracy
A significant problem with WBA is poor accuracy of faculty
ratings. For example, faculty do not detect up to 68% of
errors committed by a resident scripted to depict marginal
performance on a training video.
80,99,106
Use of specific
checklists that prompt faculty to look for certain skills can
increase error detection accuracy nearly twofold but do not
produce more accurate overall ratings of competence.
80
Nearly 70% of faculty may still rate a resident depicting
marginal performance as satisfactory or superior.80 Kalet and
colleagues examined the reliability and validity of faculty
observations using videos of student performance on an
objective structured clinical examination (OSCE) evaluating
interviewing skills.
107
Faculty inconsistently identified the
use of open-ended questions and empathy, and the positive
predictive value of faculty ratings for “adequate”
interviewing skills was only 12%. Another study found that
faculty members could not reliably evaluate 32% of the
physical examination skills assessed and had the most
difficulty with examination of the head, neck, and
abdomen.
108
https://t.me/med1917

Variable Assessment Standards
There is an increasing body of literature exploring rater
cognition: how faculty observe, synthesize their observations,
and make assessment judgments.
109,110
The rater cognition
literature has shown that faculty use variable standards
against which they compare the learner. Furthermore,
assessors often develop their assessment criteria
experientially and idiosyncratically and interpret scale
anchors variably.
111,112
Different faculty subsequently focus
on different aspects of performance and have variable
definitions of quality and competence.
111–113
When assessors
use different standards to assess learners, they rate learners
differently, decreasing interrater reliability.
Common standards are normative, self, and use of gestalt.
Some faculty use a normative standard and compare the
learner to learners at a similar level of training. However,
faculty are often uncertain about what skills are expected at a
given stage of training.
111,112,114
More commonly, faculty use
themselves as the standard when evaluating a learner.
111
This
can be problematic given the robust literature describing the
shortcomings practicing physicians can have in their clinical
skills.
28,31–33,38–41,115
Medical educators with clinical skill
deficiencies may be less likely to detect those skill deficits in
learners. This is supported by a small study showing that
faculty who were more complete in their history taking and
had better interpersonal skills, as rated by a standardized
patient, were more likely to be stringent when evaluating
residents.
116
An important limitation to the study was the use
of standardized patients to assess faculty skills since
standardized patient assessments often reward completeness
https://t.me/med1917

over the use of sophisticated illness scripts and efficiency.
14
Nevertheless, another study found that assessors often feel
they lack content expertise in the skills they are being asked
to assess.
114
Another standard used in WBA is the use of
gestalt. Many faculty say that their assessment decisions are
based on gestalt and they are unable to articulate how they
arrived at their assessment decision(s).
111
While gestalt
assessments can be accurate, they are problematic when
faculty cannot deconstruct their gestalt to provide the learner
with specific, constructive feedback. In contrast to the
common use of normative standards, self as standard, and
gestalt, faculty rarely use a criterion-referenced standard.
When faculty use a criterion-referenced standard, they
compare what the learner has done to the evidence-based
best practice components for that particular skill.
111
An
example of using a criterion-referenced approach would be
assessing which of nine shared decision-making behaviors a
resident performed when counseling a patient about starting
a new medication.
Role of Inference
Another source of WBA error is when faculty make
inferences during observation. Inference occurs when
observers derive what seems to be logical conclusions from
premises that are assumed to be true rather than assessing
observable behaviors.
111
,117
For example, imagine a resident
standing with their arms crossed breaking bad news to a
patient. Rather than note that the resident had their arms
crossed (the observable behavior), some faculty may
conclude the resident does not know how to break bad news
(inference). Other faculty may believe the resident lacks
https://t.me/med1917

empathy and humanism or might assume the resident is
uncomfortable having never broken bad news before (also
inference). Assessors make inferences about learners’
knowledge, skills (competence), and attitudes (work ethic,
emotions, intentions, personality).
111
,117
Assessors often do
not recognize when they are making inferences and rarely
validate the accuracy of their inferences.
111
Unchecked
inferences risk “distorting” accurate assessment of the learner
because the assessor’s inferences cannot be observed and
measured; this leads to greater interassessor variability and
ultimately faulty assessment.
Impact of Educational Culture
While some faculty embrace their roles and responsibilities as
assessors and coaches and do not shy away from giving low
ratings,
111
others modify their assessments to avoid
unpleasant repercussions (both for themselves and their
learners). This is another source of interrater variability.
111
Medicine has a “failure to fail” culture. As a result, some
faculty inflate ratings to avoid discussing a marginal rating
with a learner.
118,119
Other faculty may inflate assessments to
be seen as a popular, likable teacher.
111
Some faculty avoid
stringent assessments so that they will not have to defend
their assessments with institutional leaders.
111,120,121
Rater Limitations and Bias
Limitations in human cognition may also explain variability
in WBA.85 For example, contrast effects occur when the
scores faculty give to learners are influenced by their recent
observations of other learners. Faculty may rate a learner
https://t.me/med1917

with borderline skills lower if they are observed after a
learner with excellent skills; conversely, the same learner
with borderline skills may be rated higher if the faculty
observed them after a learner who is less skilled.
122,123
These
biases affect both numerical ratings and narrative
feedback.
124,125
In addition to contrast effects and recency bias, there is an
urgent need to confront and address the persistent and
pernicious effects of bias in assessment. Explicit bias refers to
conscious beliefs and attitudes one possesses about another
person or group. Implicit bias refers to an individual’s
“prejudicial attitudes toward and stereotypical beliefs about a
particular social group or members therein.”
126
These
individual attitudes are often subconscious. A growing body
of literature describes the factors contributing to, and the
harmful effects of, implicit bias and educational inequity on
learners who are underrepresented in medicine (URiM) or
otherwise at risk for marginalization in assessment practices.
Multiple studies have found that learners from groups
historically URiM receive lower assessment ratings from
faculty.
127–129
In addition, gender-based differences have
repeatedly been demonstrated in faculty assessments of
medical trainees.
130–135
These differences have been observed
in quantitative measurements where women residents
receive lower competency ratings despite no differences in
performance.
136–138
Narrative assessments often include
linguistic differences and describe personality traits that
differ by gender and URiM status.
134,135,139–141
Conflicting
priorities, time constraints, and burnout can lower the
threshold for faculty to unconsciously apply their activated
biases when they assess learners. Fig. 5.6 provides
https://t.me/med1917

suggestions on where raters can use self-awareness and selfassessment for bias in the observation and rating process.
Conceptual Model Summarizing Rater
Cognition
Fig. 5.7 is a conceptual model of the factors that influence
faculty’s WBA.
111
This model highlights the complexity of
direct observation and WBA. Letter labels direct you to the
corresponding part of the figure. When faculty observe, they
bring their attitudes and emotions about observation and
feedback, clinical competence (i.e., their own clinical skills in
the domain they are assessing the learner), educational
competence (i.e., their skills in observation, assessment, and
feedback as well as their biases), and traits (e.g., age, gender,
clinical, and teaching experience) (A). Faculty observe
learner–patient interactions through two lenses (B). While
observing, they make inferences and use variable frames of
reference (normative, self, criterion-referenced, gestalt). After
the observation, faculty must interpret and synthesize their
observations (C) to select a numerical rating or create a
narrative assessment (D). Sometimes faculty modify their
assessments in anticipation of the feedback they need to give
(E). For example, faculty may select a higher rating to avoid
discussing a low rating with the learner. Experiential
modifiers can further influence faculty’s observations and
their interpretation and synthesis of observations and
feedback (F). For example, prior experiences working with a
learner (e.g., learner’s prior performance, receptivity to
feedback, etc.) can influence faculty assessment and feedback.
Finally, observations are influenced by the clinical system (G)
(e.g., familiarity with the patient, patient complexity, system
https://t.me/med1917

factors such as multiple patients waiting to be seen) and the
institution’s educational culture (e.g., assessment culture,
culture of oversight/supervision) (H).
FIG. 5.7 Conceptual model: process of direct observation (see
text for details).
Improving the Quality of Direct
Observation Assessments
In this section of the chapter, we describe the need for faculty
development to address some of the aforementioned factors
that negatively influence WBA. Before doing so, we
recognize that some educators do not believe interrater
variability is problematic and embrace the variable
perspectives of multiple raters.
7,110,142,143
Educators holding
this perspective believe that assessment is enriched by the
unique perspectives of multiple evaluators whose
assessments, when taken together, can provide meaningful
https://t.me/med1917
Соседние файлы в папке Библиотека им академика М.И. Перельмана
