Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_2823_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
05.09.2026
Размер:
18 Мб
Скачать
Evaluation forms should serve as an important template for feedback. Since evaluation forms usually include the competencies of interest, reviewing the form with the trainee will help them gain knowledge about the content and characteristics of the clinical competencies and understand the framework used for evaluation. Reviewing with trainees their evaluation forms using Milestones, EPAs RIME, or Canadian Medical Education Directives for Specialists (CanMEDS) roles can allow discussion of the specifics of the task or activity involved, and even the competencies that are required to do the task successfully.
The trainee should be able to review completed evaluation forms as part of a comprehensive evaluation program; for instance, as part of a portfolio approach to comprehensive assessment (see Chapter 15). It is important to remember that one of the evaluation form’s major purposes is to document the professional development of the trainee at their current stage of training. Ideally, the evaluation record, whether paper or electronic, should provide space for the trainee to react and respond to the evaluation. And the trainee should be encouraged to provide in writing their reactions and subsequent plans for personal development based on the evaluation. In this respect, a portfolio may be more than a tool for documentation or assessment and may move into a curricular device for stimulating reflection (see Chapter 15). However, to be most effective, faculty must take responsibility for completing and returning evaluation forms in a timely fashion and review the evaluation form with the trainee prior to the end of an educational or training experience.
https://t.me/med1917
Narrative-Based Assessment
Until recently, less attention has been given to the written comments often provided on the rating forms, and as noted earlier, descriptive evaluation is a crucial aspect of trainee evaluation.58 Narratives, which describe a resident’s progress in the six ACGME competencies in terms of meeting Milestones, are expected as part of the residency program’s report on each resident’s progress toward independence.
54
As discussed earlier, BARS such as RIME and competency milestones provide exemplar narratives. The revision of the US Milestones for each specialty now comes with a supplemental guide that provides specific examples for each Milestone level.
The other significantly important form of narrative is what is provided by teachers from their own verbal or written observations about their residents. However, educators have often found the quality of written comments to be of little help; written comments tend to be brief, cryptic comments such as “works hard” or “should read more.” Obviously, such comments would not be sufficiently helpful for the trainee if the goal is to guide improvement by specific direction. An older study, involving two internal medicine residency programs using the low-tech approach, investigated the effectiveness of a brief, multifaceted educational intervention with faculty to improve their written evaluation of residents on inpatient ward rotations.
59
The intervention was quite simple: a 15-minute review of evaluation and feedback prior to the start of the rotation and a folded 5 × 7–inch card that contained educational
https://t.me/med1917
reminders and space to record observations. The main goals of that study were to improve the specificity of the comments with regard to the areas of competence being evaluated (e.g., medical knowledge versus clinical judgment) and to encourage faculty to provide behavioral examples in support of any low or high rating.
The investigators found a modest increase in the number of category-specific written comments and comments related to the clinical skills (e.g., history taking, physical exam) categories in the intervention group compared to the control group. However, residents in the intervention group also reported two important effects: residents were more likely to change their medical management based on feedback from the attending, and they rated the feedback from the attending significantly higher than residents in the control group. The study suggests that a fairly simple, brief faculty intervention may lead to changes in faculty-written evaluations. Given the increasingly busy nature of academic clinical practice, training programs clearly need educational interventions that are both brief and effective. Technology appears to hold some promise in this area. For example, surgery programs are using a smartphone application to complete a Zwisch scale immediately after a procedure, linked to Milestones.33 This smartphone app also uses natural language processing (NLP) that enables faculty to dictate part of the evaluation form feedback that is instantly available to the resident. The ACGME has made available an assessment app called Direct Observation of Clinical Care (DOCC) that leverages the NLP software embedded within smartphones to capture narrative from the teacher.60 More
https://t.me/med1917
work is needed to see whether repeated interventions would produce sustained or greater improvement in written evaluations.
Unfortunately, many forms do not provide enough space for written comments and the form is meant to provide a summative rather than formative evaluation. Perhaps more importantly, descriptive comments written down by teachers are often characterized with the pejorative term subjective, since they are neither quantified nor consistently anchored in specific behaviors observed by all teachers. It is possible that research into the use of descriptive terminology has been inappropriately retarded by the “subjective–objective” terminology, and we should refer to numerical methods (such as multiple-choice examinations) more appropriately as “quantified” or “objectified” rather than “objective.”
61
Kogan and colleagues tested a new design for the mini-CEX by placing the narrative questions first and asking for an entrustment rating as the last assessment activity. The DOCC app uses this same approach (see Chapter 9).
While we believe technology will make it more feasible to collect narrative assessment from teachers, limitations of written comments on evaluation forms will likely remain a challenge. What can educators do to enhance the utility of teachers’ evaluations of trainees in longitudinal educational experiences?
Evaluation Sessions
Evaluation sessions, in which the program, clerkship, or
https://t.me/med1917
course director sits down with teachers to discuss performance of their residents or students, can be a powerful adjunct to evaluation forms. After the introduction of formal evaluation sessions—regularly scheduled meetings of clerkship directors with teachers
45
,47–49
—studies documented the intuitive expectation that teachers would report in conversations what they had not initially written down on their forms.
25,31
These sessions do not need to be long; 10 to 15 minutes is sufficient to explore professionalism issues and additional detail about the trainee’s performance during the educational rotation. All programs should consider using in­person evaluation meetings with faculty. In all US residency and fellowship programs, a CCC is required as part of the assessment program. CCCs assess residents and fellows twice a year using the Milestones framework to provide feedback, guide individual learning plans, and hopefully identify learners in difficulty earlier
62
(see Bloom [1956] to learn more about effective group process). It is helpful to develop a systems approach to evaluation in which observations about learners’ progress are made by teachers who have been calibrated by organizational educators who in turn calibrate themselves and make the eventual advancement decisions.
63
Psychometric Issues
Use of any rating scales must possess sufficient reliability and validity to provide useful information about clinical competence. Furthermore, the quality and process of data collection used for the ratings are critically important (see
https://t.me/med1917
Chapter 2). We will examine some of the psychometric
challenges in using evaluation forms. The paper by Gray provides a helpful and basic review of the psychometric issues in rating scales.
7
The article by Rekman and colleagues (cf. Annotated Bibliography) provides a more recent overview of entrustment scales.
64
A brief review of some key psychometric issues is provided in the following subsections.
Reliability
Reliability refers to the consistency of assessment measurements when repeated65 (see Chapter 2). Scoring information that is consistent and confirmable is a true score or signal; the remainder is considered an error score or noise. And of course the closer the rating is to the true score the better or more reliable the assessment method is. Reliability estimates (i.e., coefficients) of greater than 0.8 are considered necessary for higher-stakes decisions. High interrater agreement is a desirable property for rating scales, especially if the forms are completed by more than one evaluator during a similar time period. Results from older studies are conflicting. Older studies found highly variable reliability using mostly quality-type scale anchors. In 2011, Crossley and colleagues, using a newly designed construct-aligned scale based on expectations for stage of training in the UK Foundation program, found much higher reliabilities on the mini-CEX.9 More recently, Park and colleagues found that Milestone-based trajectories in a family medicine residency reliably differentiated individual longitudinal patterns for
https://t.me/med1917
formative purposes using growth rate and growth curve reliability analysis.
66
Weller and colleagues have found that entrustment scales also possessed better reliability than older scales such as the 9-point quality-based mini-CEX.67 This work could inform the redesign of typical evaluation scales moving forward. The most straightforward solution to improve reliability is simply to increase the number of evaluations. Reliability is a measure of reproducibility, and the more one does anything the more stable the measurement becomes (e.g., the error around the mean narrows). Increased reliability may result from asking teachers to decide in their day-to-day activities whether they trust a trainee to do a particular task, rather than comparing the trainee to a rating scale based on abstract domains.
64,68
Validity
Validity is confidence that we measured what we wanted to measure. Modern validity frameworks now regard all validity as construct validity (see Chapter 2) where validity is of an argument or inference from the available data and not something an assessment tool possesses.
69,70
Usually a gold standard is required to adequately assess validity (e.g., relationship with other variables in the Kane or Messick validity framework). Unfortunately, no definitive standard exists for important areas of interest such as professionalism, attitudes, and clinical judgment, to name a just a few.
Performance on multiple-choice examinations has long been used as a comparison variable for correlation studies of evaluation forms. For instance, a study with family practice
https://t.me/med1917
residents found that the ability to correctly predict scores on the In-Training Exam (ITE) was dependent on the faculty’s years of teaching experience. A study using the ITE as the gold standard in family medicine showed that an online tool could be used to improve teachers’ ratings.
71
Conversely, a study of surgery residents found no correlation between the American Board of Surgery ITE and a 12-item 7-point ward evaluation rating form.72 Finally, an older study at a military internal medicine residency found that faculty were unable to predict in which tertile of performance residents would score on their ITE despite having worked closely with the residents.73 More recently, multiple studies have found a correlation between competency milestone ratings and performance on board certification exams in the United States.
74–76
Thus use of knowledge-based examinations may serve as a reasonable reference standard for knowledge ratings, but little has been done to study the relationship of other important domains on the rating scales such as professionalism, humanism, and physical exam skills. One prior study did find modest correlations between a 25-item rating scale for interns completed by residency program directors and their interns’ performance in medical school. These authors identified 5 factors from the rating scale that accounted for most of the variance: interpersonal communication, clinical skills, population-based health, recordkeeping skills, and critical appraisal skills.77 Thus ratings from medical school appear to have some predictive validity when compared to program director ratings.
78
Another study of internal medicine residents found that low
https://t.me/med1917
ratings in professionalism and other competencies at graduation were associated with a higher odds ratio of adverse state licensing board actions in practice, but the absolute rate was still quite modest.
79,80
Early outcomes research with the competency milestones provides some hope that use of developmental scales in combination with group-based evaluation sessions may provide better predictive value as a feedback mechanism to help learners while in training based on graduate performance in practice. Heath and colleagues found that low professionalism ratings during internal medicine residency training predicted subsequent low ratings in pulmonary fellowship.81 Han and colleagues found that ratings in professionalism and communication competencies were significantly associated with receipt of impactful patient complaints in early practice.82 Smith and colleagues found that a composite rating of Milestones (akin to an EPA) of vascular surgeons during training was significantly associated with surgeons’ serious complication rate in endovascular aneurysm repair (EVAR) in early practice.
83
Rating Errors
Errors in rating the performance of an individual may result from issues in the raters (e.g., their memory of their observations and their expectations and standards of comparison); from issues in the measurement scale (e.g., terms or categories that do not make sense to the rater); or from issues in the context of the rating (e.g., practical aspects such as time to observe, interruptions, or behaviors of the
https://t.me/med1917
trainee or a patient). Most problems with rating scales are due primarily to how faculty use them and not necessarily to major defects in the scales themselves.
Medical educators commit the same types of rating errors noted in all types of performance appraisals. There are two main categories of rater errors: distributional errors and correlational errors.
56
1. Distributional errors
Two common distributional errors are range restriction and leniency:
Range restriction: failure to utilize the entire range of the scale. A “central tendency” error is a subtype of a range restriction error where the rater uses just the middle portion of the scale. However, in medicine, attendings usually restrict the marks to the upper ends of the scale (discussed later in the chapter). An exception to this was seen by Battistone and colleagues who reported a shift of the grading curve to the left (away from the inflation) when the behavioral terms based on the RIME scheme (observer, reporter, interpreter, manager, educator) were substituted for numerical ratings.
46
Leniency/severity error: a type of distributional error where the faculty is being either too kind (a “dove”) or too harsh (a “hawk”), respectively. Many would argue there are currently few “hawks” in medicine.
2. Correlational errors
https://t.me/med1917