Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_2823_Библиотеки_им_академика_М_И_Перельмана
.pdf
Evaluation forms should serve as an important template for
feedback. Since evaluation forms usually include the
competencies of interest, reviewing the form with the trainee
will help them gain knowledge about the content and
characteristics of the clinical competencies and understand
the framework used for evaluation. Reviewing with trainees
their evaluation forms using Milestones, EPAs RIME, or
Canadian Medical Education Directives for Specialists
(CanMEDS) roles can allow discussion of the specifics of the
task or activity involved, and even the competencies that are
required to do the task successfully.
The trainee should be able to review completed evaluation
forms as part of a comprehensive evaluation program; for
instance, as part of a portfolio approach to comprehensive
assessment (see Chapter 15). It is important to remember that
one of the evaluation form’s major purposes is to document
the professional development of the trainee at their current
stage of training. Ideally, the evaluation record, whether
paper or electronic, should provide space for the trainee to
react and respond to the evaluation. And the trainee should
be encouraged to provide in writing their reactions and
subsequent plans for personal development based on the
evaluation. In this respect, a portfolio may be more than a
tool for documentation or assessment and may move into a
curricular device for stimulating reflection (see Chapter 15).
However, to be most effective, faculty must take
responsibility for completing and returning evaluation forms
in a timely fashion and review the evaluation form with the
trainee prior to the end of an educational or training
experience.
https://t.me/med1917

Narrative-Based Assessment
Until recently, less attention has been given to the written
comments often provided on the rating forms, and as noted
earlier, descriptive evaluation is a crucial aspect of trainee
evaluation.58 Narratives, which describe a resident’s progress
in the six ACGME competencies in terms of meeting
Milestones, are expected as part of the residency program’s
report on each resident’s progress toward independence.
54
As discussed earlier, BARS such as RIME and competency
milestones provide exemplar narratives. The revision of the
US Milestones for each specialty now comes with a
supplemental guide that provides specific examples for each
Milestone level.
The other significantly important form of narrative is what is
provided by teachers from their own verbal or written
observations about their residents. However, educators have
often found the quality of written comments to be of little
help; written comments tend to be brief, cryptic comments
such as “works hard” or “should read more.” Obviously,
such comments would not be sufficiently helpful for the
trainee if the goal is to guide improvement by specific
direction. An older study, involving two internal medicine
residency programs using the low-tech approach,
investigated the effectiveness of a brief, multifaceted
educational intervention with faculty to improve their
written evaluation of residents on inpatient ward rotations.
59
The intervention was quite simple: a 15-minute review of
evaluation and feedback prior to the start of the rotation and
a folded 5 × 7–inch card that contained educational
https://t.me/med1917

reminders and space to record observations. The main goals
of that study were to improve the specificity of the
comments with regard to the areas of competence being
evaluated (e.g., medical knowledge versus clinical judgment)
and to encourage faculty to provide behavioral examples in
support of any low or high rating.
The investigators found a modest increase in the number of
category-specific written comments and comments related to
the clinical skills (e.g., history taking, physical exam)
categories in the intervention group compared to the control
group. However, residents in the intervention group also
reported two important effects: residents were more likely to
change their medical management based on feedback from
the attending, and they rated the feedback from the
attending significantly higher than residents in the control
group. The study suggests that a fairly simple, brief faculty
intervention may lead to changes in faculty-written
evaluations. Given the increasingly busy nature of academic
clinical practice, training programs clearly need educational
interventions that are both brief and effective. Technology
appears to hold some promise in this area. For example,
surgery programs are using a smartphone application to
complete a Zwisch scale immediately after a procedure,
linked to Milestones.33 This smartphone app also uses
natural language processing (NLP) that enables faculty to
dictate part of the evaluation form feedback that is instantly
available to the resident. The ACGME has made available an
assessment app called Direct Observation of Clinical Care
(DOCC) that leverages the NLP software embedded within
smartphones to capture narrative from the teacher.60 More
https://t.me/med1917

work is needed to see whether repeated interventions would
produce sustained or greater improvement in written
evaluations.
Unfortunately, many forms do not provide enough space for
written comments and the form is meant to provide a
summative rather than formative evaluation. Perhaps more
importantly, descriptive comments written down by teachers
are often characterized with the pejorative term subjective,
since they are neither quantified nor consistently anchored in
specific behaviors observed by all teachers. It is possible that
research into the use of descriptive terminology has been
inappropriately retarded by the “subjective–objective”
terminology, and we should refer to numerical methods
(such as multiple-choice examinations) more appropriately
as “quantified” or “objectified” rather than “objective.”
61
Kogan and colleagues tested a new design for the mini-CEX
by placing the narrative questions first and asking for an
entrustment rating as the last assessment activity. The DOCC
app uses this same approach (see Chapter 9).
While we believe technology will make it more feasible to
collect narrative assessment from teachers, limitations of
written comments on evaluation forms will likely remain a
challenge. What can educators do to enhance the utility of
teachers’ evaluations of trainees in longitudinal educational
experiences?
Evaluation Sessions
Evaluation sessions, in which the program, clerkship, or
https://t.me/med1917

course director sits down with teachers to discuss
performance of their residents or students, can be a powerful
adjunct to evaluation forms. After the introduction of formal
evaluation sessions—regularly scheduled meetings of
clerkship directors with teachers
45
,47–49
—studies documented
the intuitive expectation that teachers would report in
conversations what they had not initially written down on
their forms.
25,31
These sessions do not need to be long; 10 to
15 minutes is sufficient to explore professionalism issues and
additional detail about the trainee’s performance during the
educational rotation. All programs should consider using inperson evaluation meetings with faculty. In all US residency
and fellowship programs, a CCC is required as part of the
assessment program. CCCs assess residents and fellows
twice a year using the Milestones framework to provide
feedback, guide individual learning plans, and hopefully
identify learners in difficulty earlier
62
(see Bloom [1956] to
learn more about effective group process). It is helpful to
develop a systems approach to evaluation in which
observations about learners’ progress are made by teachers
who have been calibrated by organizational educators who
in turn calibrate themselves and make the eventual
advancement decisions.
63
Psychometric Issues
Use of any rating scales must possess sufficient reliability
and validity to provide useful information about clinical
competence. Furthermore, the quality and process of data
collection used for the ratings are critically important (see
https://t.me/med1917

Chapter 2). We will examine some of the psychometric
challenges in using evaluation forms. The paper by Gray
provides a helpful and basic review of the psychometric
issues in rating scales.
7
The article by Rekman and colleagues
(cf. Annotated Bibliography) provides a more recent
overview of entrustment scales.
64
A brief review of some key psychometric issues is provided
in the following subsections.
Reliability
Reliability refers to the consistency of assessment
measurements when repeated65 (see Chapter 2). Scoring
information that is consistent and confirmable is a true score
or signal; the remainder is considered an error score or noise.
And of course the closer the rating is to the true score the
better or more reliable the assessment method is. Reliability
estimates (i.e., coefficients) of greater than 0.8 are considered
necessary for higher-stakes decisions. High interrater
agreement is a desirable property for rating scales, especially
if the forms are completed by more than one evaluator
during a similar time period. Results from older studies are
conflicting. Older studies found highly variable reliability
using mostly quality-type scale anchors. In 2011, Crossley
and colleagues, using a newly designed construct-aligned
scale based on expectations for stage of training in the UK
Foundation program, found much higher reliabilities on the
mini-CEX.9 More recently, Park and colleagues found that
Milestone-based trajectories in a family medicine residency
reliably differentiated individual longitudinal patterns for
https://t.me/med1917

formative purposes using growth rate and growth curve
reliability analysis.
66
Weller and colleagues have found that
entrustment scales also possessed better reliability than older
scales such as the 9-point quality-based mini-CEX.67 This
work could inform the redesign of typical evaluation scales
moving forward. The most straightforward solution to
improve reliability is simply to increase the number of
evaluations. Reliability is a measure of reproducibility, and
the more one does anything the more stable the
measurement becomes (e.g., the error around the mean
narrows). Increased reliability may result from asking
teachers to decide in their day-to-day activities whether they
trust a trainee to do a particular task, rather than comparing
the trainee to a rating scale based on abstract domains.
64,68
Validity
Validity is confidence that we measured what we wanted to
measure. Modern validity frameworks now regard all
validity as construct validity (see Chapter 2) where validity
is of an argument or inference from the available data and
not something an assessment tool possesses.
69,70
Usually a
gold standard is required to adequately assess validity (e.g.,
relationship with other variables in the Kane or Messick
validity framework). Unfortunately, no definitive standard
exists for important areas of interest such as professionalism,
attitudes, and clinical judgment, to name a just a few.
Performance on multiple-choice examinations has long been
used as a comparison variable for correlation studies of
evaluation forms. For instance, a study with family practice
https://t.me/med1917

residents found that the ability to correctly predict scores on
the In-Training Exam (ITE) was dependent on the faculty’s
years of teaching experience. A study using the ITE as the
gold standard in family medicine showed that an online tool
could be used to improve teachers’ ratings.
71
Conversely, a
study of surgery residents found no correlation between the
American Board of Surgery ITE and a 12-item 7-point ward
evaluation rating form.72 Finally, an older study at a military
internal medicine residency found that faculty were unable
to predict in which tertile of performance residents would
score on their ITE despite having worked closely with the
residents.73 More recently, multiple studies have found a
correlation between competency milestone ratings and
performance on board certification exams in the United
States.
74–76
Thus use of knowledge-based examinations may serve as a
reasonable reference standard for knowledge ratings, but
little has been done to study the relationship of other
important domains on the rating scales such as
professionalism, humanism, and physical exam skills. One
prior study did find modest correlations between a 25-item
rating scale for interns completed by residency program
directors and their interns’ performance in medical school.
These authors identified 5 factors from the rating scale that
accounted for most of the variance: interpersonal
communication, clinical skills, population-based health,
recordkeeping skills, and critical appraisal skills.77 Thus
ratings from medical school appear to have some predictive
validity when compared to program director ratings.
78
Another study of internal medicine residents found that low
https://t.me/med1917

ratings in professionalism and other competencies at
graduation were associated with a higher odds ratio of
adverse state licensing board actions in practice, but the
absolute rate was still quite modest.
79,80
Early outcomes research with the competency milestones
provides some hope that use of developmental scales in
combination with group-based evaluation sessions may
provide better predictive value as a feedback mechanism to
help learners while in training based on graduate
performance in practice. Heath and colleagues found that
low professionalism ratings during internal medicine
residency training predicted subsequent low ratings in
pulmonary fellowship.81 Han and colleagues found that
ratings in professionalism and communication competencies
were significantly associated with receipt of impactful
patient complaints in early practice.82 Smith and colleagues
found that a composite rating of Milestones (akin to an EPA)
of vascular surgeons during training was significantly
associated with surgeons’ serious complication rate in
endovascular aneurysm repair (EVAR) in early practice.
83
Rating Errors
Errors in rating the performance of an individual may result
from issues in the raters (e.g., their memory of their
observations and their expectations and standards of
comparison); from issues in the measurement scale (e.g.,
terms or categories that do not make sense to the rater); or
from issues in the context of the rating (e.g., practical aspects
such as time to observe, interruptions, or behaviors of the
https://t.me/med1917

trainee or a patient). Most problems with rating scales are
due primarily to how faculty use them and not necessarily to
major defects in the scales themselves.
Medical educators commit the same types of rating errors
noted in all types of performance appraisals. There are two
main categories of rater errors: distributional errors and
correlational errors.
56
1. Distributional errors
Two common distributional errors are range restriction
and leniency:
Range restriction: failure to utilize the entire range of the
scale. A “central tendency” error is a subtype of a range
restriction error where the rater uses just the middle
portion of the scale. However, in medicine, attendings
usually restrict the marks to the upper ends of the scale
(discussed later in the chapter). An exception to this was
seen by Battistone and colleagues who reported a shift of
the grading curve to the left (away from the inflation)
when the behavioral terms based on the RIME scheme
(observer, reporter, interpreter, manager, educator) were
substituted for numerical ratings.
46
Leniency/severity error: a type of distributional error
where the faculty is being either too kind (a “dove”) or
too harsh (a “hawk”), respectively. Many would argue
there are currently few “hawks” in medicine.
2. Correlational errors
https://t.me/med1917
Соседние файлы в папке Библиотека им академика М.И. Перельмана
