Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5211_Библиотеки_им_академика_М_И_Перельмана
.pdf
Paul Banaszkiewicz
The cut score is the minimum score/level of
competence a candidate must achieve in order to have
career progression.
The minimally competent candidate
The first step in standard setting is agreeing on the
characteristics of the minimally competent candidate.
This is an imaginary candidate who you would be
happy to see scrape through, but only just. Not the
ideal average performance. A barely safe candidate.
‘They would manage at the next level of career
progression, but there’s definitely still room for significant improvement.’
‘The greenest of the red apples.’
Generating a cut score
Once the exam board have defined what a barely
competent day one consultant is, they now have to
decide on a ‘cut score ’ or score which the barely
competent consultant should achieve.
From ‘how good is good enough’ to ‘how many
points does a candidate need to pass’.
The Angoff method is used to determine this.
Angoff Method
The Angoff method is a widely used standard setting
approach in test development. It is a type of test that
the Intercollegiate Board (ICB) uses to determine the
passing percentage (cut score) for a test.
It involves a cut-off mark based on the performance of candidates in relation to a defined standard
(absolute) as opposed to how they perform in relation
to their peers (relative). It involves a judgement being
made on exam items (test-centred) as opposed to
exam candidates (examinee-centred) and is widely
used to standard set high-stakes examinations. It is
most reliable when supported by another standard
setting method.
what makes a candidate competent enough to be
passed, hence why the examiners must be
experienced.
Each examiner works independently and con-
siders each question in turn. Each expert’s judgement
for an item should be the same or within a close,
defined range (around 10%). The mean of everyone’s
judgement is calculated for each item; this is often
referred to as the ‘predicted difficulty’. Each predicted
difficulty (mean) is added together and divided by the
total number of items in the exam to get the cut-off
percentage. This percentage of the total marks for the
exam indicates the cut-off mark (Table 2.2).
If the examiners’ judgements are not unanimous,
they discuss how they came to their decision in an
effort to come to an agreement. The Angoff is then
recalculated based on the new judgements. This process may be repeated.
Remember that there are 5 options in a mu ltiple
choice question, and the probability of answering
correctly by chance is 2/10. Therefore, we would not
expect an SBA to score below 2 for minimally competent candidates to choose the correct option. This is
not a hard rule but an SBA scored lower than 2 means
that this is lower than the probability of guessing,
which is unusual.
Equally, there is no rule against using 10 as the
score a minimally competent candidate should score,
but this should not be the norm. What the exam
board is looking for is the probability that a day one
consultant ‘would’ answer correctly, not what we
think they ‘should’ do in an ideal world.
An additional method of supporting the Angoff
method is to use another standard-setting method,
such as borderline regression to provide results based
on real candidate data for comparison. Also, if the
final candidate results do not reflect the standard that
would be expected of the candidates taking the exam,
the standard-setting method can be re-evaluated.
HowisAngoffCalculated?
A group of senior examiners are asked, ‘What percentage of borderline candidates would answer this
SBA correctly?’ Before making a judgement, the
examiners must agree on the definition of a ‘borderline’ candidate.
All examiners must have the same definition of a
borderline candidate for the Angoff cut-off score to
be reliable. This requires a good understanding of
18
Advantages of Using Angoff
The advantages include the following:
Holds up in court – Angoff is the most widely used,
formal method of standard setting. There are many
published works on Angoff, and it is justifiable for
use in high-stakes examinations. If questioned, the
Angoff method would hold up in court.
Reflects the difficulty of the content – Angoff
focuses on just the content of the exam and the

SBA Writing Process
Table 2.2 Table 2.2. Angoff method. The total average percentage is 56.7. This can be rounded to 57, giving a cut-off percentage of 57%. If
the test were out of 100 marks, a borderline candidate would be expected to get 57/100 marks.
Q. Examiner
1 (%)
1 60 60 55 60 65 60 55 60 60 60 59.5
27070707070707070707070
35550555055555550555053
46060606560606060656061
5 60 65 70 70 60 70 65 70 65 70 66.5
65045505045504550455048
76060606060606060606060
85050505050505050505050
9 55 60 55 55 55 55 55 60 55 60 56.5
10 40 40 45 45 40 40 45 45 45 40 42.5
Cut-off percentage 57
level at which candidates should be performing to
meet a certain standard.
Simple when you know how – This is a fairly
straightforward process once all judges are trained.
Recyclable – If an item is reused in another exam
with the same context (same year group), the
Angoff ‘predicted difficulty’ can be re-used so
subject experts have fewer ite ms to judge.
Examiner
2 (%)
Examiner
3 (%)
Examiner
4 (%)
Examiner
5 (%)
Examiner
6 (%)
Examiner
7 (%)
Examiner
8 (%)
Examiner
9 (%)
Examiner
10 (%)
Reducing the Effect of Outliers
If the examiner is a hawk ... they would score high
... so the cut score would be too high ... so many
good candidates may not pass.
If the examiner is a dove ... they would score low
... so cut score would be low ... so some poor
candidates will get through!
Therefore, outlier scored questions may need to be
Mean
(%)
removed or rewritten to get a more consistent exam-
Disadvantages of Using Angoff
iner response.
The disadvantages include the following:
Needs back-up – It does not use real exam data to
estimate a cut-off mark, so it is considered more
accurate and reliable if backed up by a criterion-
Removing the Effect of Outliers
Removing negative or low point biserial questions.
referenced method, e.g. borderline regression.
Long process – The process can be time
consuming and labour intensive, as examiners
must look at every test item. This can lead to
examiners becoming fatigued and impatient and
can encourage rushing through the items.
Confidence is key – Examiners must be experts in
their field. This method relie s on the examiners
being confident and consistent with their
definition of a ‘borderline’ candidate, and not just
assuming an ‘average’ candidate.
Time and a place – You need a large number of
examiners for accuracy and reliability.
Construction of the Paper
The full curri culum will be sampled and usually
320 SBAs are submitted for consideration of
inclusion into the paper.
This will be a mixture of new SBAs and well
performing SBAs.
Around 60 SBAs will be removed as they are too
easy, difficult, ambiguous etc.
Another 20 SBAs are removed after further
analysis.
This leaves the required 240 SBA questions
needed for each diet of exams.
19

Paul Banaszkiewicz
SBA Checks before Finalisation of the
Paper
The paper is sat by at least 10 examiners with
feedback obtained about the suitability of the
included questions. A number of the included
questions will be removed.
Paper scrutiny by a senior experienced group of
examiners and removal of further SBAs that are
deemed necessary.
Head examiner. He/she has the final say for
approving the paper.
Standard Error of Measurement (SEM)
and the GMC
If we consider that the whole exam had only 10
questions and all of the examiners independently
concluded that 6 of every 10 borderline candidates
would get each question correct, then a pass mark of 6
out of 10 (60%) would mean that 50% of borderline
candidates would pass and 50% would fail. The pass
mark therefore divides the borderline candidates
down the middle. If the exam has a lot of hard
questions, the pass mark will be lower. If there are a
lot of easy questions, it will be higher. The mark is
unique to each diet.
The Angoff-derived pass mark is not the mark
determining eligibility to proceed, however. The
GMC argue that there is some uncertainty in judgements made in this way, which can be expressed
statistically as the Standard Error of Measurement
(SEM). For patient safety reasons, the GMC would
not want incompetent candidates being allowed to
proceed, even if removing them means some
potentially competent candidates are prevented from
doing so.
The exam boards emphasis must be on patient
safety. It is much more defensible to fail a candidate
who is just more than competent than to pass one
who is not competent.
The eligibility to proceed mark is therefore the
Angoff-derived mark plus one SEM. When this step
was first introduced, the historical performance of
candidates scraping through was reviewed and it was
noted that they went on to fail Section 2, so this rule
in fact saves some candidates a whole lot of money!
The SEM allows us to identify what the GMC refer
to as ‘ borderline candidates’ either side of the cut
score.
Borderline Candidates
There is an attempt to standardise the exam so that
around 33% (1 standard deviation) of borderline (just
competent) candidates will pass.
The exam mark is adjusted to prevent a large
percentage of borderline candidates passing.
Consider Figure 2.9. If the middle dotted line is
the Angoff determined ‘cut score’ of the ‘barely competent day one consultant’ then the SEM represents
what the GMC classes as the ‘borderline candidate’.
Applied to a bell-shaped curve then 34 +16 = 50%
of time a barely competent candidate will pass.
If we were to add one SEM to the ‘cut score’ then
only 16% of time will the barely competent candidate
pass.
- Anyone whose ‘true score’ is below the ‘true score’
of the ‘day one consultant’ therefore has less than
a 16% probability of achieving a score above cut
score + 1 SEM.
The emphasis must be on patient safety. It is much
more defensible to fail a candidate who is just more
than competent than to pass one who is not
20
34,1%34,1%
2,1%2,1%
13,6%13,6%
60 90 120 150 180 210 240
Figure 2.9 Classic ‘bell curve’ shape of
normal distribution. The mean (or
average) is the vertical line at the centre,
and the vertical lines to either side
represent intervals of one, two and three
SDs. The percentage of data points that
would lie within each segment of that
distribution are shown.
0,1%0,1%

SBA Writing Process
Figure 2.10 Borderline candidate and
Angoff-derived mark plus one SEM.
tomorrow, then it will show the same weight. But if
the scale is not working properly and is not reliable, it
could give you a different weight each time.
A reliable exam paper should produce the same
result if the test is repeated. Ideally, for an exam to be
highly reliable, it should produce the same result if the
candidate takesthe same test on two differentoccasions.
If it produces very different results on separate occasions, it can be argued that the exam has low reliability
and can therefore not be trusted as a means of grading
or judging whether a candidate is competent.
Figure 2.11 Borderline candidates and Angoff-derived mark plus
one SEM added to the cut score.
competent (Figure 2.10). Therefore we add 1 SEM to
the cut score (Figure 2.11).
Reliability
Reliability is a Measure of Consistency
It is useful to think of kitchen scales. For weighing
scales to be reliable, you would expect that if you
weighed 80 kg 10 minutes ago you would still weigh
the same now. If the scales said you weighed 40 kg
now, you would not rely on those scales.
If the scale is reliable, then when you put a bag of
flour on the scale today and the same bag of flour on
How Can You Tell if a Test is Reliable
(Cronbach’salpha)?
In most cases, we cannot get candidates to take the
same test twice to measure reliability, so the internal
consistency is measured with an alternative method.
Internal consistency reliability measures the degree to
which every test item measures the same construct.
The closer the score is to 1, the higher the reliabil-
ity will be. As a rough guide, exams should have a
Cronbach’s alpha of 0.8 or greater. With high-stakes
exams, the reliability needs to be high because one
exam is usually used to decide whether a student
passes or fails. Around the world, few professional
examinations, particularly in medical specialties,
achieve a Cronbach’s alpha of 0.8.
Validity
As well as being reliable, it is important that a test is
valid, i.e. measures what it is supposed to measure.
Continuing the kitchen scale metaphor, a scale might
consistently show the wrong weight; in such a case,
the scale is reliable but not valid.
21

Paul Banaszkiewicz
There are a number of ways in which validity can
be increased. These include the following:
Item analysis reporting. This flag questions that do
not correlate well with the rest of the assessment.
One of the reasons a question might get flagged is
because participants who do well on other
questions do not do well on this question – this
could indicate the question lacks content validity.
Item bank. Use an item bank to store wellperforming, defined topic questions.
Review and update the question bank frequently.
Orthopaedic knowledge base can change quickly
with changing technology and changing
regulations. Many questions that were valid 2 years
ago are outdated today. Using an item bank allows
you to update or retire questions that are no longer
relevant.
Use assessment blueprinting. The blueprint
confirms that the exam tests a representative sample
of all the appropriate curriculum outcomes and a
representative sample of all the curriculum content.
Involve subject matter experts (SMEs). The more
you involve senior exam writers in assessment
development, the more content validity you are
likely to get. Get a panel of SMEs to rate each
question as to whether it is ‘essential’, ‘useful, but
not essential’,or‘not necessary’ to the performance
of what is being measured. The more SMEs who
agree that items are essential, the higher the content
validity will be.
Blueprinting the Curriculum
To ensure adequate content coverage and, therefore,
that the exam is a valid test of the breadth of kno wledge of orthopaedics, the process of blueprinting
occurs. A spreadsheet is created that maps each of
the questions to a learning objective on the curriculum. The assessment blueprint is a grid that plots the
curriculum/assessment outcomes in columns against
the curriculum content in rows.
Constructing SBAs
In order for an SBA to be good, it must fulfil two
basic criteria: (1) what the question tests should be
important and (2) the question is well structured.
SBAs consist of a stem (e.g. a clinical case presentation) and a lead-in question, followed by a series of
choices, typically onecorrect answer and four distractors.
The following sections outline an example of an
SBA.
Stem
A 60-year-old man presents to casualty 5 days
following left ceramic on ceramic total hip replacement with increasing left hip pain. He was discharged
from hospital 2 days after surgery with a plan for
follow-up at 3 months. On examination, he is apyrexial, normotensive with modest hip discomfort with
movement.
Lead-In
Which of the following is the most likely diagnosis?
Options
A. Ceramic head fracture
B. Ceramic liner fracture
C. PJI
D. Non-specific hip pain
E. Squeaking ceramic interface
All options may be correct, but at 5 days non-specific
hip pain is the most likely diagnosis. The patient has
been discharged from hospital after 2 days and
increased their activity levels. While it is important
to rule out early periprosthetic joint infection (PJI), it
should be rare, under 1%, and there would be more
clues towards this in the stem such as a leaky wound
or high temperature. Likewise, ceramic head or liner
fractures are rare but possible causes of hip pain.
With the fourth-generation Bioloc Delta ceramics,
the incidence of ceramic head fractures has significantly reduced but there has been no corresponding
change in liner fracture incidence. Likewise, if there
was a ceramic fracture a bit more would be given in
the history to point candidates in this direction.
Squeaking is not uncommon following ceramic on
ceramic total hip arthroplasty (CoC THA) (around
7%), and while it can be annoying is usually not
painful.
Even though the less likely answers are not wrong,
they are less correct than the ‘keyed answer’.The
candidate is instructed to select the most likely diagnosis and experts would all agree that the most likely
diagnosis is D. They would also agree that the other
diagnoses are somewhat likely, but less likely than D.
As long as the options can be laid out on a single
continuum, in this case from ‘Most Likely Diagnosis’
22

SBA Writing Process
AB
Least
Correct
Figure 2.12 SBA construction
to ‘Least Likely Diagnosis’, options in one-best-answer
questions do not have to be totally wrong (Figure 2.12).
Options are homogeneous (i.e. all possible diag-
noses) and high-calibre candidates will be able to
rank-order the options along a single dimension.
Well-constructed one-best-answer questions sat-
isfy the ‘cover-the-options’ rule. The questions could
be administered as write-in questions. The entire
question is included in the stem.
CE D
Most
Correct
Example of Poorly Constructed SBA
Which of the following is true about pseudogout?
A. It occurs frequently in men.
B. It is seldom associated with acute pain in a joint.
C. It may be associated with a finding of
chondrocalcinosis.
D It is clearly hereditary in most cases.
E. It responds well to treatment with naproxen.
This item is flawed. There is no lead-in. After reading
the stem , the candidate has an unclear idea what the
question is about. In an attempt to determine the
‘best’ answer, candidates have to decide whether ‘it
occurs frequently in men’ is more or less true than ‘it
is seldom associated with acute pain in a joint’. This is
a comparison of apples and pears. Why are you trying
to compare those things which you can’t?
In order to rank-order the relative correctness of
options, the options must differ on a single dimension
or else all options must be absolutely 100% true or false.
The diagram of these options would look like
Figure 2.13. The options are heterogeneous and deal
with miscellaneous facts; they cannot be rank ordered
from least to most true along a single dimension.
Although this question appears to assess knowledge
of several different points, its inherent flaws preclude
this. The question by itself is not clear; the item
cannot be answered without looking at the options.
It has also been proposed by the ICB to avoid
using negative A-type questions in SBAs.
The most problematic are those that take the
form: ‘Each of the following is correct EXCEPT’ or
‘Which of the following statements is NOT correct?’
Figure 2.13 SBA formats
These suffer from the same problem as true/false
questions: if options cannot be rank ordered on a
single continuum, the examinees cannot determine
either the ‘least’ or the ‘most’ correct answer.
On the other hand, the ICB occasionally uses wellfocused negative A-types with single-word options,
largely as a (poor) substitute for items that instruct
the examinee to select more than one response.
Technical Item Flaws
This section is based on technical item flaws from chapter 3 in Paniagua and Swygert, constructing written test
questions for thebasic and clinical sciences.
provides specific examples of technical item flaws.
Two important types of technical item flaws are
(1) testwiseness and (2) irrelevant difficulty.
Flaws related to testwiseness make it easier for
some students to answer the question correctly, based
on their test-taking skills alone. These flaws commonly occur in items that are unfocused.
Flaws related to irrelevant difficulty make the
question difficult for reasons unrelated to the trait
that is the focus of assessment.
Examination boards are keen to eliminate flaws to
provide a level playing field for testwise and not so
testwise candidates.
8
The chapter
9
Flaws Related to Testwiseness
Grammatical Cue
One or more distractors do not follow grammatically
or logically from the stem. Testwise candidates are
able to spot this and eliminate them from the options.
23

Paul Banaszkiewicz
Logical Cue
If several but not all of the options are very similar,
this can suggest that the answer is in this subset. Try
to make the options as homogeneous as possible.
Longest Answer is ‘ Single Best Answer’
The correct answer is longer, more specific or more
complete than the other options.
Repeated Words
A word or phrase included in both the stem and the
correct answer can act as a clue towards a correct
guess.
UseofAbsoluteTerms
Absolute terms such as ‘always’ or ‘never’ should not
be used in options.
Presence of Convergence
This is less obvious to spot but occurs fairly commonly. The correct answer includes the most elements in common with the other options. The
underlying premise is that the correct answer is the
option that has the most in common with the other
options; it is not likely to be an outlier.
Terms in the Options or the Stem are Vague
Vague frequency terms in the options such as ‘often’
or ‘usually’ are used. Value frequency terms are not
consistently defined or interpreted even by experts.
‘True/False’
If the options are all either completely correct or
completely incorrect, the question does not follow
the ‘single best answer’ approach and is most likely
pitched at the recall level. Options should be on a
continuum from least to most appropriate.
These types of questions are found in some ortho-
paedic MCQ books.
‘None of the Above’ is Used as an Option
The phrase ‘none of the above’ is problematic in items
where judgement is involved and the options are not
absolutely true or false. Use of ‘none of the above’
essentially turns the item into a true/false item; each
option has to be evaluated as more or less true than
the universe of unlisted options .
Language or Structure of the Options is Not
Homogeneous or P arallel
The format and structure of each option is different.
Presence of Grouped or Collectively
Exhaustive Options
A testwise student can identify a subset of options that
cover all the possible outcomes (are collectively
exhaustive) and rule out the options not in that subset.
Flaws Related to Irrelevant Difficulty
Overly Complicated or Long Questions
The stem contains extraneous reading and the options
are very long and complicated. Trying to decide
among these options requires a significant amount
of reading because of the number of elements in each
option.
Numerical Data are Not Consistently
Presented
Numerical options should be listed in a consistent
manner and not mixed (ranges and percentages).
24
Window Dressing
Window dressing in an MCQ is information in the
item itself which is superfluous to the content being
assessed. It is not about shorter versus longer stems
but about the relevance of what is there.
Unnecessarily Complicated Stems
Stems should be meaningful by themselves and
should present a definite probl em.
They should not contain irrelevant material that
may decrease a test score’s reliability and validity.
Stems should include as much of the item as possible;
stems should be long and the options short.
Stems Contain Negative Phrasing
Items which ask a question in the negative (e.g. ‘which
is the least appropriate ...’ ) are confusing to candi-
dates who have to switch between best and worst
answers and are contrary to the mode of reasoning
in most clinical situations. The stem should be

SBA Writing Process
negatively stated only when a significant learning
outcome requires it.
Alternatives/Distractors
All alternatives should be plausible. Alternatives
should be stated clearly and concisely. Items that are
excessively wordy assess a candidate ’ s reading ability
rather than the attainment of the learning objective.
They should be homogeneous in content.
Alternatives that are heterogeneous in content can
provide cues to candidates about the correct answer.
Alternatives should be free from clues about
which response is correct. Sophisticated test takers
are alert to inadvertent clues to the correct answer,
such as differences in grammar, length, formatting
and language choice in the alternatives.
General Guidelines
Avoid using absolutes such as ‘always’, ‘never’, and
‘all’ in the options; also avoid using vague terms such
as ‘usually’ and ‘frequently’.
Avoid using ‘all of the above ’ and ‘none of the
above’. They are problematic in ite ms where judgement is involved, and the options are not absolutely
true or false. In either case, candidates can use partial
knowledge to arrive at a correct answer.
Focus on important concepts; do not waste time
testing trivial facts.
MCQ Question Writing Committee
This committee meets every 3 months with around 20
or so experienced examiners attending.
They begin by looking at some of the SBAs that
have been flagged statistically as possibly poor performers. Some questions will already have been
removed automatically – for example, all of the questions that prove d too easy or too hard (usually new
questions, as any question previ ously used would
have passed this hurdle already).
The examiners will review each question and
decide whether it is a fair question that should stay
in the exam or is flawed and should be removed and
returned to the question writers. Typical reasons for
the latter would be ambiguity that had not previously
been recognised, new evidence that has challenged the
previously decreed correct answer, or simply that the
answer in the bank is wrong.
It is worth noting that some very good questions
end up being flagged as having possible wrong
answers yet remain in the exam. If a question is hard
so that only 20% of candidates answer it correctly,
then 80% will choose a wrong response. Let’s say 40%
chose one of the incorrect options – this flags as a
possible wrong answer automatically, as more candidates have chosen a specific incorrect response than
the correct one.
Once the poorly perform ing questions have been
dealt with, the Angoff procedure is performed.
New speciality question writers need to attend
workshops in how to write SBAs before being allowed
to submit questions. Existing question writers need
to attend face-to-face meetings every 3 months.
At these meetings, new questions are submitted
to the panel for review and possible inclusion in
the question bank. These meetings are expensive to
hold, involving travel and possibly accommodation
expenses but deemed worthwhile as they allow faceto-face scrutiny of the questions submitted for the
question bank.
Basic Rules for Writing SBAs
Each item should focus on an important concept
or testing point.
Each item should assess application of knowledge,
not recall of an isolated fact.
The item lead-in should be focused, closed and
clear; the test taker should be able to answer the
item based on the stem and lead-in alone.
All options should be homogeneous and plausible,
to avoid cueing to the correct option.
Always review items to identify and remove
technical flaws that add irrelevant difficulty or
benefit savvy test takers.
One minute rule
Each question should be able to be read by the
candidate within 1 minute maximum.
If the stem is too lengthy it will need to be edited
into a readable length.
SBA Writing Guidelines: Dos and Don’ts
Getting Started
1. Identify an area of the curriculum blueprint that is
to be sampled.
2. Identify the topic area and the level of thinking that
you want to test and write a question around this
25

Paul Banaszkiewicz
which mimics tasks that successful candidates must
be able to undertake at the next stage of training.
Ideally, questions should be pitched at the level
of integration/interpretation (questions which
require ‘putting the pieces together’) and problem
solving (questions which require ‘clinical
judgement’), not simple recall (questions which
can be answered with a Google search).
Examples of recall questions to avoid:
What are the symptoms of X?
Which of the following is a contraindication
of X?
What is the name/definition of this procedure?
Which of the following is correct?
Examples of higher order (‘putting the pieces
together’ and ‘clinical judgement’) questions:
In this procedure, which structur e is most at
risk in this patient?
What is the most likely diagnosis?
What is the most useful investigation to carry
out at this stage?
What is the most appropriate first step in the
management of this patient?
3. Construct the stem. This should present a single,
clearly formulated problem, possibly in the form of
a patient vignette. The stem should contain enough
information to allow candidates to answer without
referring to the options.
Patient vignettes may include any subset of the
following information:
a. The patient (age, gender), the presenting
complaint and its duration; any other relevant
information a competent day one consultant
would typically need to piece together to
determine the best course of action.
b. Relevant details of the patient’s history,
possibly including details of family history.
c. Relevant physical findings, results of
diagnostic studies, initial treatment,
subsequent findings, images, etc.
4. Constructthelead-ininsuchawaythatitbuildson
theinformationinthestemandposesaclear
question. Candidates should be able to answer
without looking at the options and should not be
abletoansweriftheinformationinthestemis
masked.
5. Write the options. These should be of similar
length, along a continuum, grammatically
consistent and logically compatible. If
appropriate, order the options in a logical order
(e.g. numeric, alphabetical or anatomical). All the
distractors must be plausible (think of the
educational impact of suggesting that something
potentially dangerous is plausible but not the best
answer!).
Reviewi ng the Item
Does it focus on important problems relevant to
clinical practice?
Can it be answered without looking at the
options?
Are all the relevant facts included in the stem?
Can it be read and answered within approximately
one minute?
Are all the options plausible, with one of them
standing out as being the best option?
Does the question successfully avoid the pitfalls
listed in the Question Writing Checklist?
Question Writing Checklist
For a deeper dive, see Table 2.3.
A statistical analysis is performed on the candidate
scores to determine the mean, standard deviation,
standard error, variation, quarterly percentiles and
normality testing with histograms and normal probability plots. These results are used to test the calibre
and spread of candidates compared with previous
groups.
Question-writing workshops provide detailed
guidance on the design of questions that discriminate
between candidates of differing ability, in a format
and a style that aid speed reading and comprehension.
Summary
There are many different types of MCQs. The traditional multiple true/false type that asks candidates to
identify all the correct statements listed can assess
knowledge and comprehension but is limited to these
objectives.
By contrast, the single best answer question, which
asks candidates to choose the best answer from, say,
five plausible alternatives, assesses not only knowledge and comprehe nsion but also the application of
this knowledge to the synthesis of deductions from
26

SBA Writing Process
Table 2.3 SBA question writing checklist
Checks across all question formats
Alignment Is the task required by the question congruent with the learning objectives at this stage
of career progression (e.g. interpretation of information, decision making)?
Blueprinting Are the topic and task listed in the exam syllabus and is the blueprint code provided?
Level of competence Is the question pitched to discriminate around the level of minimal competence required
for a pass?
Spelling and grammar Is the question free from spelling mistakes and grammatical errors?
Jargon Is the question jargon free (as much as possible within the topic area)? Is it free from
idioms which might unfairly disadvantage non-native English speakers?
Timing Can the question be read and answered within approximately 1 minute (for SBAs)?
Clarity Is it clear from the question what candidates are expected to do?
Additional checks: common pitfalls in writing questions
Grammatical cue The sentence structure can sometimes allow candidates to exclude a subset of the
options. To avoid this, make the lead-in a whole question rather than a sentence
continuation.
Logical cue If several but not all of the options are very similar, this can suggest that the answer is in
this subset. Try to make the options as homogeneous as possible.
Absolute or vague terms ‘Never’ and ‘always’ are almost never true; vague terms make it difficult to know what the
question writer meant (e.g. how often is ‘often’?).
Longest answer is ‘single
best answer’
Repeated words Repeated words (or related words) between the stem and the options can point to the
Convergence strategy By looking at how many times each term is repeated across the set of options, candidates
Overly complicated
questions
Negatively phrased items Items which ask a question in the negative (e.g. ‘which is the least appropriate ...’) are
‘True/false’ If the options are all either completely correct or completely incorrect, the question does
Option-dependent
question
Window dressing Questions which can be answered without referring to the stem are typically pitched at
The answer with the most information/the highest level of precision is often the correct
answer. Try to give all options a similar length and similar level of precision.
answer.
select the option which contains all the most frequently repeated terms.
Questions which are too long or present information in overly complex forms are in
danger of testing skills which are not in line with the purpose of the exam (e.g. reading
speed, working memory).
confusing to candidates who have to switch between best and worst answer and are
contrary to the mode of reasoning in most clinical situations.
not follow the ‘single best answer’ approach and is most likely pitched at the recall level.
Options should be on a continuum from least to most appropriate.
Candidates should be able to answer the question without referring to the options. This
encourages them to think of the answer for themselves based on the information
provided in the stem and lead-in.
the recall level. They also use up precious time in the examination by giving candidates
irrelevant information to read.
27
Соседние файлы в папке Библиотека им академика М.И. Перельмана
