Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_112_библиотеки_им_акад_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
05.09.2026
Размер:
18 Мб
Скачать
to the teacher); that are reliable and reproducible (i.e., the evaluation framework is applied consistently and not capriciously across observations by teachers; and in which formative evaluations can be relied upon by learners as
anticipations of summative grading.
11
Faculty development in general is a quality improvement process in training teachers how to use evaluation forms and scales, and specifically is intended to minimize unacceptable
variation
12
,13
between observers in which ratings would depend more upon teacher characteristics or teacher preferences than upon attributes of the trainee. It is important to understand that the teacher completing the rating scale is the actual assessment “instrument” and that the form provides a way to communicate and foster shared expectations and document performance.
Evaluation forms are in themselves, implicitly or explicitly, frameworks to guide teachers’ observations and documentation of a learner’s performance. Therefore they usually include an explicit statement of goals for the learners, or at least the criteria by which learners are to be judged. As an official and legal document of a program or institution, these forms publicly express curricular goals and are intended to avoid variance that is arbitrary due to having inconsistent goals and standards across teachers. However, they are not guarantees that teachers will not be capricious (inconsistent or idiosyncratic) in applying them to individual learners.
14
Indeed, the rating scale must make sense to the rater and foster construct alignment between the rater and institutional expectations.
We will begin this chapter by discussing the importance of
https://t.me/med1917
evaluation frameworks to guide the user in their effective use. Next, we provide an outline of the advantages and disadvantages of rating scales and evaluation forms, including some important psychometric rating error issues that limit the effectiveness of evaluation forms. The chapter will close with practical suggestions on how to prepare faculty to use evaluation forms more effectively.
Evaluation Forms and Frameworks
In any evaluation form used in evaluating a trainee, there is an underlying set of assumptions about what we expect of the trainee (what are the educational goals?); and about the tasks that a trainee must complete successfully to demonstrate that the goals have been achieved (curricular objectives). This set of assumptions can be considered the underlying framework that encompasses (“frames”) what the observer must compare the trainee to in order to decide whether the trainee is progressing. This side-to-side comparison between what the teacher does observe in this learner as distinct from what the teacher expects to see provides the “is-versus-ought” or “real-versus-ideal” that is the basis of most judgments that lead to assigning labels (i.e., codes) or classifications (e.g., rating scales). Program directors should not see this as a passive process in which meaningful observations will successfully flow into the teacher’s mind, to be subsequently placed into the desired framework or construct. It should be seen as an active process in which the teacher seeks observations to satisfy the expectations of the evidence-based mental model that the framework embodies.
https://t.me/med1917
Just as in diagnosis of clinical syndromes, it is critical that the observer has an accurate and valid mental model against which to compare the trainee. Educational frameworks are ways to conceptualize expectations for the comparisons that underlie evaluations of trainees’ success. Teachers commonly divide educational goals into the three familiar categories of knowledge, skills, and attitudes. We will describe this as well as other useful frameworks in the next section.
Analytic Frameworks
In the traditional framework used in education, including in elementary and secondary school, there are typically three domains: knowledge, skills, and attitudes (KSA). It is straightforward to place learning objectives for learners, especially in preclinical rotations, into the three KSA domains: for instance, knowledge of the structures within the chest cavity, skill in the physical examination of heart and lungs, and a proper attitude of respect for the patient’s physical comfort and privacy.
The analytic approach to formulating educational goals provides a generic set of terms that can be applied to any curricular task in any field of education. The analytic approach is particularly useful when trying to measure discrete aspects of performance, and of course program directors are quite familiar with using a single multiple­choice test as a measure of knowledge and perhaps with using a checklist to rate skill in examining a patient’s knee or in giving informed consent. By isolating one particular aspect of performance—for instance, the ability to interview
https://t.me/med1917
the patient about an alcohol history or the ability to place a central venous catheter—the analytic method allows us to create a fairly detailed set of performance tasks that can be compiled into a checklist, which in turn can be placed on a rating form. Such a checklist can be as detailed as necessary to help the teacher document whether each aspect of a particular task can be performed separately, prior to the teacher combining them into a single (global) rating. Together, the checklist and overall global rating constitute the criterion description of what ultimate proficiency or competency looks like for this task. The analytic method requires an evaluation form to have at least three scales, with labels at the top of each gradation to distinguish levels of success (Fig. 4.2A
). The scale is an example of using “quality of performance” as the ratings, and these scales continue to be used despite growing evidence that they suffer from a lack of construct alignment (discussed later in the chapter) and require a substantial amount of translation by the rater.
9
In the example given, the labels are rudimentary and require the user to infer what the terms mean. For the critical, central label of “Acceptable,” the form might provide more concrete terms for or examples of what is meant through use of behaviorally anchored scales, which we discuss in more detail later in this chapter.
https://t.me/med1917
Fig. 4.1 Placing evaluation forms into a comprehensive system. MCQs, multiple-choice questions; SPs,
standardized patient examinations.
Fig. 4.2a Rating scale in knowledge-skills-attitude framework with rudimentary anchors. Rating scale
using the Dreyfus developmental framework with generic global terms. Rating scale using the developmental framework from Bloom’s taxonomy in the cognitive domain. “Distances” between performance levels on any scale may not be equal, as
https://t.me/med1917
illustrated in RIME.
In addition the central anchor of “acceptability” is not made explicit and defers the applicable criterion or criteria to the rater. Multiple studies have shown that the primary frame of reference (i.e., standard) used by faculty is self (see Chapter
5). Rating scales on evaluation forms have similarity with
items on surveys and questionnaires, asking for agreement with an implicit statement that the learner being evaluated meets, exceeds, or falls short of some criteria of acceptability. There are pitfalls that program directors can avoid in creating items for faculty to rate, such as not having anchors for each level or using both words and numbers for each level.
15
It is also worth noting that other scale anchor types are still commonly used, and all suffer from the same issues noted earlier. The traditional mini clinical evaluation exercise (mini-CEX) uses a 9-point scale with the scale anchors of 1–3 (unsatisfactory), satisfactory,
4–6
and superior.
7–9
Others use normative scaling (i.e., assessed against the expected performance of others) with variable frames of reference for the expectations (i.e., compared to peer learner, graduating learner, or practicing physician). Such scales tend to range from “fails to meet expectations” to “exceeds expectations.” Finally, some evaluations using an analytic framework will employ frequency scales (e.g., “rarely” to “almost always”), assuming a greater frequency of a behavior indicates higher levels of competence. While all of these types of analytic scales may have utility in specific circumstances, we do not recommend the routine use of these scales. Instead, we strongly recommend educators choose scales for their
https://t.me/med1917
evaluation forms based on developmental frameworks.
Developmental Frameworks
The growth of human beings has often been used as a metaphor for the growth of trainees in an educational process. The pedigree of this approach is ancient, and Plato describes the growth of the individual from a preoccupation with surface, concrete details toward a perception of the true meaning and form underlying them. In the well-known Taxonomy of Educational Objectives in the Cognitive Domain, Bloom
16,17
provides a vocabulary for describing the progressively higher mental skills acquired by students in primary education: knowledge, comprehension, application, analysis, synthesis, and evaluation The developmental model of Dreyfus and Dreyfus,18 revised for GME,19 provides a generic vocabulary of educational progress for adult learners from novice to advanced beginner, competent performance, proficient performance, intuitive expert, and master. Developmental considerations are essential for medical school faculty because they reflect the facts that students grow, that not all learners are at the same level of performance, and that in the clinical setting there are often learners at several levels of training. The models of Bloom and Dreyfus typically focus on cognitive aspects of development, and personal and attitudinal characteristics are not always evident. Bloom, and to some extent Dreyfus, choose to treat the attitudinal (“affective”) domain separately from the cognitive. However, one advantage of developmental over analytic models is that the recognition of growth and progress is explicit and does not have to be
https://t.me/med1917
inferred by the teacher or the student. To this extent, any curriculum that has learners at different stages, like medical school, requires some explicitly developmental aspect. Using the Dreyfus model as an ordinal rating scale—similar to a Likert scale—could be constructed with the word Novice on the left and the word Master at the extreme right (see Fig.
4.2B).
The Dreyfus terms used as “anchors” on the linear scale shown in Fig. 4.2 are global (i.e., generic and nonspecific), and a more detailed, developmental scale could be devised for a specific domain within the analytic framework.
This can be made less vague, although still abstract, using Bloom’s taxonomy for the growth of clinical reasoning. For example, Fig. 4.2C demonstrates how you might judge a resident’s presentation of a patient.
In this example, the criterion against which the resident is to be rated is an idea or construct or what effective clinical reasoning looks like, and the program director should use the best available evidence of what effective clinical reasoning should look like (i.e., accurate diagnoses and treatment plans; see Chapter 7) as the expectation or standard of comparison.
Given the ACGME/ABMS framework of six competencies, in each of which a finishing resident is judged to be successful (or not), such static or pass/fail dichotomous ratings of competence must be modified for those earlier in their training. The faculty or program director must determine what is the standard to be met within the criterion for each
https://t.me/med1917
level, and the program director must reframe each specific ACGME competency into a developmental model that can describe what the acceptable/passing standard of performance is for student, intern, resident, or fellow. To help in this process, more structured ways to describe and document progress have recently been introduced: Milestones and Entrustable Professional Activities (EPAs).
20
Milestones,
21,22
as described in Chapter 1, are observable behaviors or tasks that combine or synthesize knowledge, skill, and attitudes and therefore may be seen as synthetic to better define a competency in narrative terms. Using the concept of stages of development, the Milestones provide an explicit developmental scale for a resident’s professional development across the competencies. The Royal College of Physicians and Surgeons of Canada (RCPSC) has also created Milestones for its competency-based medical education (CBME) initiatives.23 These and EPAs are described in the next section. This combination of different domains into a single observation means Milestones are a synthetic approach.
24
GME programs are now required to report residents’ progress in a range of subcompetencies described in narrative Milestone levels (typically 20 to 25 subcompetencies per specialty), which represent progressive steps toward achieving mastery performance within an ACGME competency domain.21 Milestones can allow faculty observers to focus on the task rather than the framework.
22
Milestones, although organized as a set of subcompetency domains, can be considered synthetic in that a number of the subcompetency domains, such as the subcompetency
https://t.me/med1917
“System Navigation for Patient Care,” do require the integration of knowledge, skills, and attitudes. Ultimately, however, a synthesis of all assessment information must be performed to answer the most important question of whether a learner is or is not ready for the next stage of their career, also satisfying the legal precept of looking at the
10
A Synthetic Model
As students and residents progress toward independence, we expect that they themselves will spontaneously bring whatever skills, knowledge, and attitudes are necessary to help a patient each day; the learner is responsible for deciding (either consciously or not) what the patient needs. This leads to a “synthetic” definition of competence as “the ability to bring to each patient seen in one’s practice everything that that patient needs, and nothing else.”25 In other words, competence at the point of unsupervised (aka independent) practice requires that the resident decides what the task is, right at this moment, and summons up whatever skill, knowledge, or attitude is needed.16 In rating performance, we might comment separately on a resident’s fund of knowledge or “attitude,” but in the end we have to judge whether residents have been able to master all the necessary attributes and—on their own—combine them successfully. A synthetic framework “puts things together” in a vocabulary that emphasizes progressively higher expectations as a student progresses through the clinical years and through residency. The underlying premise is that
https://t.me/med1917