Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_112_библиотеки_им_акад_М_И_Перельмана
.pdf
An important goal of ratings is to accurately discriminate
between levels of ability, whether using an analytic,
developmental, or synthetic framework. For example, one
resident may possess substandard knowledge as
evidenced by other measures (e.g., the ITE) but display
extraordinary humanistic skill. Such a resident should
receive high marks under humanistic qualities but should
receive a lower rating for medical knowledge if the
attending has effectively evaluated the resident.
Unfortunately, most raters have difficulty discriminating
between dimensions of competence and tend to use a
limited range on the rating scale.
Correlational error is where faculty give similar ratings to
each aspect of a trainee’s performance regardless of what
dimension of competence is being assessed, even when the
dimensions are clearly separate;56 as Murphy and Cleveland
note, “the result is an inflation of the inter-correlations
among the dimensions.” When rating inflation occurs, the
result is commonly known as a halo error. This is a very
common problem in medical education where everyone is
always above average or better. Having the entire list of
desirable attributes of the candidate present, synoptically, at
one time, is intended to minimize the chance that teachers
will confuse the different domains of evaluation, but caution
is needed. The increasing length and complexity of forms
may in themselves make the teachers less able or willing to
use it as intended by the form’s author(s); a problem known
as cognitive load. Therefore increasing the number of
domains, rating criteria, or items on a form may aggravate
the halo effect.
https://t.me/med1917

A number of older studies within a single or limited number
of programs, using factor analysis, found that evaluation
forms containing multiple items or competencies could be
reduced to just two factors, typically cognitive skills and
interpersonal skills.
84,85
However, a more recent study using
national competency milestone data found that in fact the six
competency domains held up very well, and two other
validity studies found that the nonmedical knowledge
competency milestones either did not or only weakly
correlate with performance on a high-stakes medical
knowledge certification exam.
74,75,86
What these Milestone
studies suggest is again that more developmental
frameworks combined with group process may provide
more discriminating evaluations of learners.
A few notes of caution are warranted in interpreting factor
analytic studies. First, factor analysis is a statistical technique
that attempts to reduce a data set to the fewest number of
factors using large correlational matrices.87 Second, factor
analysis assumes each category is independent of the other
and that the relationship between two or more factors is
linear. For example, items that load onto more than one
factor are typically eliminated from the final factor model. In
medicine we know that is illogical. For example, you cannot
perform a high-quality history and physical without robust
medical knowledge, and it is hard to be professional and
have excellent interpersonal and communication skills. It is
important that competencies mostly serve as frameworks, as
described earlier. Using factor analysis to determine whether
faculty can discriminate between categories of competence
may fail to uncover modest differences between ratings.
https://t.me/med1917

While faculty may not make major distinctions among the
competencies in their ratings for reasons listed earlier, they
still can help signal what is important. Finally, certain
approaches to factor analysis used in studies fail to address
or account for clustering within programs (i.e., residents are
nested within programs).
What are the reasons for halo error when it occurs? For one,
raters may rely more on their global impressions when rating
trainees on specific dimensions of competence. Two, and one
we all recognize, is the unwillingness of so many faculty to
give lower ratings on any dimension of competence. This is
often called the “there goes my teaching award syndrome.”
Other possible causes of halo error include confirmation bias
(self as frame of reference: “that’s how I would do it, so it
must be right”), ignoring discordant or inconsistent
information or observations about the trainee, the simple
lack of enough observation or information about
performance, and bias.
68,88
Box 4.2 provides a list of possible
reasons for halo error.
Box 4.2
Possible Reasons for the Halo Error
1. Global impression drives the rating for all dimensions of
competence
2. Unwillingness or inability to discriminate among
different dimensions
3. Reluctant to give negative evaluations
4. Insufficient observation and/or information about a
trainee’s performance
https://t.me/med1917

5. Implicit and/or confirmation bias
6. Discounting conflicting or discordant information,
observations
7. Level of familiarity with trainee
8. Level of familiarity with the medical knowledge, skills,
and attitudes
9. Dimensions of competence are interdependent
Rater Accuracy
Rater accuracy answers the question: How well does the
rating match the actual performance? There are two distinct
types of accuracy measures. The first are behavioral-based
measures. These types of measures allow the rater to
specifically focus on whether the behavior did or did not
occur. Checklists are the most common type of rating scale
used on evaluation forms for behaviorally based measures.
The level of ability may be included in a checklist, but
typically the rater is expected only to provide “credit” for
behaviors performed properly based on a standard. They are
particularly useful for structured, controlled assessments
such as standardized patients. However, checklists, like
other evaluation forms, can ask too much of faculty and
create a situation known as cognitive load. The more items
we ask faculty to rate (judge) in shorter periods of time, the
greater the cognitive load that leads to less effective
evaluations. As one example, Byrne and colleagues
compared the cognitive load between completing a 21-item
checklist for an OSCE station versus inducing anesthesia for
routine surgery.89 The same principle applies to evaluation
forms. The more straightforward synthetic RIME framework
https://t.me/med1917

can classify and document observations of the behaviors
observed in an individual trainee’s care of each in a series of
patients as consistent with reporting, interpreting, managing,
and educating.
The other type of accuracy involves judgmental measures.
As the name implies, the rater must apply judgment when
providing a rating. Accuracy in judgment is particularly
important for rating scales and evaluation forms used in
longitudinal educational experiences. There are several types
of judgmental accuracy measures: accuracy in whether a
trainee has attained a level of performance (criterion
accuracy); accuracy in distinguishing among trainees
(differential or normative accuracy); and accuracy in
discriminating between specific performance or competence
dimensions (stereotype accuracy). For evaluation in medical
education, accuracy measures are important because
defining key behaviors at various levels of competence
facilitates better judgment and helps support better patient
care.
Another caution is that the concept of trust may be too
complicated a social judgment to reduce to a rating scale.
Gingerich argued that assessment involves some degree of
social judgment and that some idiosyncrasy is to be expected
because all teachers have strengths and weaknesses in their
own clinical and teaching practices.
90
These idiosyncrasies
can benefit the learner when they represent excellence or
clinical mastery. However, they can be harmful when they
do not represent an evidence-based practice (EBP; see
Chapter 5). Gingerich and colleagues also described that in a
study of faculty rating a resident’s performance there was a
https://t.me/med1917

tendency for subgroups of physicians who had described
similar social judgments to have also given more similar
performance ratings.91 Faculty development may be an
important process to articulate shared assumptions among
teachers.
Finally, an important note of caution about entrustment
scales and accuracy. Two recent studies uncovered serious
problems in entrustment ratings against an outcomes-based
standard. First, Schumacher and colleagues examined the
correlation between supervisor ratings of the clinical
encounters of pediatric residents in the emergency
department and the quality of care delivered to the child
assessed through quality performance measures. The results
were sobering: there was little to no correlation with the
quality for care for asthma, bronchiolitis, and closed head
injury and the entrustment ratings.92 Kogan and colleagues
compared the entrustment ratings of primary care residency
faculty as part of a randomized trial on a series of videos
rigorously scripted to depict a specific entrustment level.
While the accuracy of entrustment ratings was good when
the performer depicted a resident “ready for unsupervised
practice (Level 4),” accuracy was less than 50% when the
performer was Level 2 or 3 on the entrustment, with leniency
error most common.31 Thus entrustment scales are no
panacea for all the challenges described in this chapter,
highlighting the critical need for ongoing faculty
development (discussed next; see also Chapter 5).
Faculty Development and
https://t.me/med1917

Evaluation Forms
The quality of the information on evaluation forms depends
mostly on the individual completing the form, not the form
itself. For too long medical educators have been looking for
the holy grail of the ideal evaluation form. Landy and Farr in
1980 called for a moratorium on this quest, arguing instead
for an increased emphasis on training the evaluators.
55
As
noted earlier, even simple approaches to faculty
development such as observation cards can modestly
improve the quality of information on evaluation forms, and
the growing availability of smartphone apps offers even
more promise around improving faculty development
efforts. However, to realize the full potential of evaluation
forms, more structured faculty training is still needed.
Throughout this chapter we have highlighted the importance
of an evaluation framework to guide the evaluation process.
This is a critical first step to ensure that faculty possess
shared mental models and an understanding of the goals of
outcomes of evaluation, Studies have shown specific types of
training can improve interrater agreement using a simple
three-step process25 (see also Chapter 9):
1. Standardize the observation of the behavior of interest.
2. Reach agreement on common nomenclature for the
desired expectations of interest through conversation and
dialogue.
3. Agree on the relative importance of the different
components of behavior being assessed.
4. Practice assessment skills longitudinally and with
https://t.me/med1917

feedback about potential rating errors and bias.
Steps 1 and 2 in this process are called performance
dimension training (PDT). PDT provides raters with the
expected performance standards for each level of
performance. Many have argued that such agreement about
performance dimension standards is lacking in GME. Step 3
is known as frame of reference training (FoRT). These
techniques have been applied in training faculty to use the
RIME framework (see Chapter 5 for guidance on how to
perform PDT and FoRT).
31
Performance Dimension Training and RIME
The Uniformed Services University of the Health Sciences
(USUHS) has incorporated PDT and FoRT as part of the
evaluation for medical students rotating on an internal
medicine clerkship for over 30 years.
11,24,38,40
Raters
participate in evaluation sessions with clerkship directors
where descriptive evaluations are collected. Clerkship
directors use these evaluation sessions to train preceptors
about expected levels of performance for each category of
rating and how the student’s performance should be
documented on the rating scale form. The evaluation system
goes one step further by incorporating the student’s
performance in multiple domains of competence into an
overall performance level. Goals for each level of
performance are divided into performance categories with
defined expectations. Since “reporter” skills are introduced
in the first year (see Appendix 4.2), it is felt that achieving
proficiency in them is a reasonable, nonnegotiable level for
https://t.me/med1917

advancing to the next level of responsibility, yielding this
conversion of observations into grades:
Reporter (Pass):
Interpreter (High Pass):
Manager/Educator (Honors):
The Appendix 4.2 provides a more comprehensive
description of the model and a copy of the performance
matrix used at USUHS. The descriptions and criteria are
more applicable for residents than for students, since no
accommodation distinguishing “Reasonable” (student level)
from “Accurate “ (resident level) needs to be made.
Murphy and Cleveland in 1995
56
made several important
points about performance appraisal training pertinent to the
use of rating scales that still hold today:
1. Define performance dimensions in behavioral terms
(e.g., use RIME and Milestones narrative descriptions)
and be sure to communicate these terms to the resident
and the faculty. It may even be helpful to use a blank
evaluation form at the beginning of a rotation as a
template to discuss goals and expectations before the
evaluation process actually starts.
2. Ratings will more likely correspond with actual rater
judgment if training programs support distinctions
between house staff on the basis of performance, the raters
perceive a strong link between the rating they give and
specific outcomes, and the raters believe that outcomes
should be based on present performance.
3. What the rater chooses to communicate through the
https://t.me/med1917

form depends heavily on the rater’s goals and contextual
factors, such as the individual resident’s relationship with
the teachers and the perceived purpose of the evaluation
at hand. Therefore raters need to communicate goals
directly to the residents, and raters must be cognizant of
both internal and external environmental factors affecting
the context of the evaluation. Chapter 5 provides greater
detail and suggestions for running PDT and FoRT faculty
development exercises.
Conclusions
Directors of academic programs should not presume that an
evaluation form, even one with behaviorally anchored rating
scales, will do the work of getting faculty on the same page.
Evaluation forms should possess several desirable
properties: be user friendly, be unobtrusive, be flexible,
capture important narrative descriptions of performance,
and if needed, be quantifiable. The quantification to a rating
scale should be an accurate translation of level of
performance. Forms that are not well understood and/or are
too elaborate may pose a challenge in cognitive load that can
only be met with additional faculty development and
training. As importantly, teachers’ acceptance of the scale
and its educational framework are prerequisites of consistent
use. To some extent, there is an emotional barrier for teachers
that has to be bridged—they often see themselves (and
certainly describe themselves) as “giving” the student or
resident a grade, rather than making a diagnosis that reflects
their observations (something they would never do with a
serious medical condition). Therefore the teacher’s emotional
https://t.me/med1917
Соседние файлы в папке Библиотека им академика М.И. Перельмана
