Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_112_библиотеки_им_акад_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
05.09.2026
Размер:
18 Мб
Скачать
An important goal of ratings is to accurately discriminate between levels of ability, whether using an analytic, developmental, or synthetic framework. For example, one resident may possess substandard knowledge as evidenced by other measures (e.g., the ITE) but display extraordinary humanistic skill. Such a resident should receive high marks under humanistic qualities but should receive a lower rating for medical knowledge if the attending has effectively evaluated the resident. Unfortunately, most raters have difficulty discriminating between dimensions of competence and tend to use a limited range on the rating scale.
Correlational error is where faculty give similar ratings to each aspect of a trainee’s performance regardless of what dimension of competence is being assessed, even when the dimensions are clearly separate;56 as Murphy and Cleveland note, “the result is an inflation of the inter-correlations among the dimensions.” When rating inflation occurs, the result is commonly known as a halo error. This is a very common problem in medical education where everyone is always above average or better. Having the entire list of desirable attributes of the candidate present, synoptically, at one time, is intended to minimize the chance that teachers will confuse the different domains of evaluation, but caution is needed. The increasing length and complexity of forms may in themselves make the teachers less able or willing to use it as intended by the form’s author(s); a problem known as cognitive load. Therefore increasing the number of domains, rating criteria, or items on a form may aggravate the halo effect.
https://t.me/med1917
A number of older studies within a single or limited number of programs, using factor analysis, found that evaluation forms containing multiple items or competencies could be reduced to just two factors, typically cognitive skills and interpersonal skills.
84,85
However, a more recent study using national competency milestone data found that in fact the six competency domains held up very well, and two other validity studies found that the nonmedical knowledge competency milestones either did not or only weakly correlate with performance on a high-stakes medical knowledge certification exam.
74,75,86
What these Milestone studies suggest is again that more developmental frameworks combined with group process may provide more discriminating evaluations of learners.
A few notes of caution are warranted in interpreting factor analytic studies. First, factor analysis is a statistical technique that attempts to reduce a data set to the fewest number of factors using large correlational matrices.87 Second, factor analysis assumes each category is independent of the other and that the relationship between two or more factors is linear. For example, items that load onto more than one factor are typically eliminated from the final factor model. In medicine we know that is illogical. For example, you cannot perform a high-quality history and physical without robust medical knowledge, and it is hard to be professional and have excellent interpersonal and communication skills. It is important that competencies mostly serve as frameworks, as described earlier. Using factor analysis to determine whether faculty can discriminate between categories of competence may fail to uncover modest differences between ratings.
https://t.me/med1917
While faculty may not make major distinctions among the competencies in their ratings for reasons listed earlier, they still can help signal what is important. Finally, certain approaches to factor analysis used in studies fail to address or account for clustering within programs (i.e., residents are nested within programs).
What are the reasons for halo error when it occurs? For one, raters may rely more on their global impressions when rating trainees on specific dimensions of competence. Two, and one we all recognize, is the unwillingness of so many faculty to give lower ratings on any dimension of competence. This is often called the “there goes my teaching award syndrome.” Other possible causes of halo error include confirmation bias (self as frame of reference: “that’s how I would do it, so it must be right”), ignoring discordant or inconsistent information or observations about the trainee, the simple lack of enough observation or information about performance, and bias.
68,88
Box 4.2 provides a list of possible
reasons for halo error.
Box 4.2
Possible Reasons for the Halo Error
1. Global impression drives the rating for all dimensions of competence
2. Unwillingness or inability to discriminate among different dimensions
3. Reluctant to give negative evaluations
4. Insufficient observation and/or information about a trainee’s performance
https://t.me/med1917
5. Implicit and/or confirmation bias
6. Discounting conflicting or discordant information, observations
7. Level of familiarity with trainee
8. Level of familiarity with the medical knowledge, skills, and attitudes
9. Dimensions of competence are interdependent
Rater Accuracy
Rater accuracy answers the question: How well does the rating match the actual performance? There are two distinct types of accuracy measures. The first are behavioral-based measures. These types of measures allow the rater to specifically focus on whether the behavior did or did not occur. Checklists are the most common type of rating scale used on evaluation forms for behaviorally based measures. The level of ability may be included in a checklist, but typically the rater is expected only to provide “credit” for behaviors performed properly based on a standard. They are particularly useful for structured, controlled assessments such as standardized patients. However, checklists, like other evaluation forms, can ask too much of faculty and create a situation known as cognitive load. The more items we ask faculty to rate (judge) in shorter periods of time, the greater the cognitive load that leads to less effective evaluations. As one example, Byrne and colleagues compared the cognitive load between completing a 21-item checklist for an OSCE station versus inducing anesthesia for routine surgery.89 The same principle applies to evaluation forms. The more straightforward synthetic RIME framework
https://t.me/med1917
can classify and document observations of the behaviors observed in an individual trainee’s care of each in a series of patients as consistent with reporting, interpreting, managing, and educating.
The other type of accuracy involves judgmental measures. As the name implies, the rater must apply judgment when providing a rating. Accuracy in judgment is particularly important for rating scales and evaluation forms used in longitudinal educational experiences. There are several types of judgmental accuracy measures: accuracy in whether a trainee has attained a level of performance (criterion accuracy); accuracy in distinguishing among trainees (differential or normative accuracy); and accuracy in discriminating between specific performance or competence dimensions (stereotype accuracy). For evaluation in medical education, accuracy measures are important because defining key behaviors at various levels of competence facilitates better judgment and helps support better patient care.
Another caution is that the concept of trust may be too complicated a social judgment to reduce to a rating scale. Gingerich argued that assessment involves some degree of social judgment and that some idiosyncrasy is to be expected because all teachers have strengths and weaknesses in their own clinical and teaching practices.
90
These idiosyncrasies can benefit the learner when they represent excellence or clinical mastery. However, they can be harmful when they do not represent an evidence-based practice (EBP; see
Chapter 5). Gingerich and colleagues also described that in a
study of faculty rating a resident’s performance there was a
https://t.me/med1917
tendency for subgroups of physicians who had described similar social judgments to have also given more similar performance ratings.91 Faculty development may be an important process to articulate shared assumptions among teachers.
Finally, an important note of caution about entrustment scales and accuracy. Two recent studies uncovered serious problems in entrustment ratings against an outcomes-based standard. First, Schumacher and colleagues examined the correlation between supervisor ratings of the clinical encounters of pediatric residents in the emergency department and the quality of care delivered to the child assessed through quality performance measures. The results were sobering: there was little to no correlation with the quality for care for asthma, bronchiolitis, and closed head injury and the entrustment ratings.92 Kogan and colleagues compared the entrustment ratings of primary care residency faculty as part of a randomized trial on a series of videos rigorously scripted to depict a specific entrustment level. While the accuracy of entrustment ratings was good when the performer depicted a resident “ready for unsupervised practice (Level 4),” accuracy was less than 50% when the performer was Level 2 or 3 on the entrustment, with leniency error most common.31 Thus entrustment scales are no panacea for all the challenges described in this chapter, highlighting the critical need for ongoing faculty development (discussed next; see also Chapter 5).
Faculty Development and
https://t.me/med1917
Evaluation Forms
The quality of the information on evaluation forms depends mostly on the individual completing the form, not the form itself. For too long medical educators have been looking for the holy grail of the ideal evaluation form. Landy and Farr in 1980 called for a moratorium on this quest, arguing instead for an increased emphasis on training the evaluators.
55
As noted earlier, even simple approaches to faculty development such as observation cards can modestly improve the quality of information on evaluation forms, and the growing availability of smartphone apps offers even more promise around improving faculty development efforts. However, to realize the full potential of evaluation forms, more structured faculty training is still needed.
Throughout this chapter we have highlighted the importance of an evaluation framework to guide the evaluation process. This is a critical first step to ensure that faculty possess shared mental models and an understanding of the goals of outcomes of evaluation, Studies have shown specific types of training can improve interrater agreement using a simple three-step process25 (see also Chapter 9):
1. Standardize the observation of the behavior of interest.
2. Reach agreement on common nomenclature for the desired expectations of interest through conversation and dialogue.
3. Agree on the relative importance of the different components of behavior being assessed.
4. Practice assessment skills longitudinally and with
https://t.me/med1917
feedback about potential rating errors and bias.
Steps 1 and 2 in this process are called performance dimension training (PDT). PDT provides raters with the expected performance standards for each level of performance. Many have argued that such agreement about performance dimension standards is lacking in GME. Step 3 is known as frame of reference training (FoRT). These techniques have been applied in training faculty to use the RIME framework (see Chapter 5 for guidance on how to perform PDT and FoRT).
31
Performance Dimension Training and RIME
The Uniformed Services University of the Health Sciences (USUHS) has incorporated PDT and FoRT as part of the evaluation for medical students rotating on an internal medicine clerkship for over 30 years.
11,24,38,40
Raters participate in evaluation sessions with clerkship directors where descriptive evaluations are collected. Clerkship directors use these evaluation sessions to train preceptors about expected levels of performance for each category of rating and how the student’s performance should be documented on the rating scale form. The evaluation system goes one step further by incorporating the student’s performance in multiple domains of competence into an overall performance level. Goals for each level of performance are divided into performance categories with defined expectations. Since “reporter” skills are introduced in the first year (see Appendix 4.2), it is felt that achieving proficiency in them is a reasonable, nonnegotiable level for
https://t.me/med1917
advancing to the next level of responsibility, yielding this conversion of observations into grades:
Reporter (Pass): Interpreter (High Pass): Manager/Educator (Honors):
The Appendix 4.2 provides a more comprehensive description of the model and a copy of the performance matrix used at USUHS. The descriptions and criteria are more applicable for residents than for students, since no accommodation distinguishing “Reasonable” (student level) from “Accurate “ (resident level) needs to be made.
Murphy and Cleveland in 1995
56
made several important points about performance appraisal training pertinent to the use of rating scales that still hold today:
1. Define performance dimensions in behavioral terms (e.g., use RIME and Milestones narrative descriptions) and be sure to communicate these terms to the resident and the faculty. It may even be helpful to use a blank evaluation form at the beginning of a rotation as a template to discuss goals and expectations before the evaluation process actually starts.
2. Ratings will more likely correspond with actual rater judgment if training programs support distinctions between house staff on the basis of performance, the raters perceive a strong link between the rating they give and specific outcomes, and the raters believe that outcomes should be based on present performance.
3. What the rater chooses to communicate through the
https://t.me/med1917
form depends heavily on the rater’s goals and contextual factors, such as the individual resident’s relationship with the teachers and the perceived purpose of the evaluation at hand. Therefore raters need to communicate goals directly to the residents, and raters must be cognizant of both internal and external environmental factors affecting the context of the evaluation. Chapter 5 provides greater detail and suggestions for running PDT and FoRT faculty development exercises.
Conclusions
Directors of academic programs should not presume that an evaluation form, even one with behaviorally anchored rating scales, will do the work of getting faculty on the same page. Evaluation forms should possess several desirable properties: be user friendly, be unobtrusive, be flexible, capture important narrative descriptions of performance, and if needed, be quantifiable. The quantification to a rating scale should be an accurate translation of level of performance. Forms that are not well understood and/or are too elaborate may pose a challenge in cognitive load that can only be met with additional faculty development and training. As importantly, teachers’ acceptance of the scale and its educational framework are prerequisites of consistent use. To some extent, there is an emotional barrier for teachers that has to be bridged—they often see themselves (and certainly describe themselves) as “giving” the student or resident a grade, rather than making a diagnosis that reflects their observations (something they would never do with a serious medical condition). Therefore the teacher’s emotional
https://t.me/med1917