Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5211_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
37 Мб
Скачать
Paul Banaszkiewicz
The cut score is the minimum score/level of competence a candidate must achieve in order to have career progression.
The minimally competent candidate
The first step in standard setting is agreeing on the characteristics of the minimally competent candidate.
This is an imaginary candidate who you would be happy to see scrape through, but only just. Not the ideal average performance. A barely safe candidate.
They would manage at the next level of career progression, but theres definitely still room for sig­nificant improvement.
The greenest of the red apples.
Generating a cut score
Once the exam board have defined what a barely competent day one consultant is, they now have to decide on a cut score or score which the barely competent consultant should achieve.
From how good is good enoughto how many points does a candidate need to pass.
The Angoff method is used to determine this.
Angoff Method
The Angoff method is a widely used standard setting approach in test development. It is a type of test that the Intercollegiate Board (ICB) uses to determine the passing percentage (cut score) for a test.
It involves a cut-off mark based on the perform­ance of candidates in relation to a defined standard (absolute) as opposed to how they perform in relation to their peers (relative). It involves a judgement being made on exam items (test-centred) as opposed to exam candidates (examinee-centred) and is widely used to standard set high-stakes examinations. It is most reliable when supported by another standard setting method.
what makes a candidate competent enough to be passed, hence why the examiners must be experienced.
Each examiner works independently and con-
siders each question in turn. Each experts judgement for an item should be the same or within a close, defined range (around 10%). The mean of everyones judgement is calculated for each item; this is often referred to as the predicted difficulty. Each predicted difficulty (mean) is added together and divided by the total number of items in the exam to get the cut-off percentage. This percentage of the total marks for the exam indicates the cut-off mark (Table 2.2).
If the examinersjudgements are not unanimous,
they discuss how they came to their decision in an effort to come to an agreement. The Angoff is then recalculated based on the new judgements. This pro­cess may be repeated.
Remember that there are 5 options in a mu ltiple choice question, and the probability of answering correctly by chance is 2/10. Therefore, we would not expect an SBA to score below 2 for minimally compe­tent candidates to choose the correct option. This is not a hard rule but an SBA scored lower than 2 means that this is lower than the probability of guessing, which is unusual.
Equally, there is no rule against using 10 as the score a minimally competent candidate should score, but this should not be the norm. What the exam board is looking for is the probability that a day one consultant wouldanswer correctly, not what we think they shoulddo in an ideal world.
An additional method of supporting the Angoff
method is to use another standard-setting method, such as borderline regression to provide results based on real candidate data for comparison. Also, if the final candidate results do not reflect the standard that would be expected of the candidates taking the exam, the standard-setting method can be re-evaluated.
HowisAngoffCalculated?
A group of senior examiners are asked, What per­centage of borderline candidates would answer this SBA correctly?Before making a judgement, the examiners must agree on the definition of a border­linecandidate.
All examiners must have the same definition of a borderline candidate for the Angoff cut-off score to be reliable. This requires a good understanding of
18
Advantages of Using Angoff
The advantages include the following:
Holds up in court – Angoff is the most widely used,
formal method of standard setting. There are many published works on Angoff, and it is justifiable for use in high-stakes examinations. If questioned, the Angoff method would hold up in court. Reflects the difficulty of the content – Angoff
focuses on just the content of the exam and the
SBA Writing Process
Table 2.2 Table 2.2. Angoff method. The total average percentage is 56.7. This can be rounded to 57, giving a cut-off percentage of 57%. If the test were out of 100 marks, a borderline candidate would be expected to get 57/100 marks.
Q. Examiner
1 (%)
1 60 60 55 60 65 60 55 60 60 60 59.5
27070707070707070707070
35550555055555550555053
46060606560606060656061
5 60 65 70 70 60 70 65 70 65 70 66.5
65045505045504550455048
76060606060606060606060
85050505050505050505050
9 55 60 55 55 55 55 55 60 55 60 56.5
10 40 40 45 45 40 40 45 45 45 40 42.5
Cut-off percentage 57
level at which candidates should be performing to meet a certain standard. Simple when you know how – This is a fairly
straightforward process once all judges are trained. Recyclable – If an item is reused in another exam
with the same context (same year group), the Angoff predicted difficultycan be re-used so subject experts have fewer ite ms to judge.
Examiner 2 (%)
Examiner 3 (%)
Examiner 4 (%)
Examiner 5 (%)
Examiner 6 (%)
Examiner 7 (%)
Examiner 8 (%)
Examiner 9 (%)
Examiner 10 (%)
Reducing the Effect of Outliers
If the examiner is a hawk ... they would score high ... so the cut score would be too high ... so many
good candidates may not pass.
If the examiner is a dove ... they would score low ... so cut score would be low ... so some poor candidates will get through!
Therefore, outlier scored questions may need to be
Mean (%)
removed or rewritten to get a more consistent exam-
Disadvantages of Using Angoff
iner response.
The disadvantages include the following:
Needs back-up – It does not use real exam data to
estimate a cut-off mark, so it is considered more accurate and reliable if backed up by a criterion-
Removing the Effect of Outliers
Removing negative or low point biserial questions.
referenced method, e.g. borderline regression. Long process – The process can be time
consuming and labour intensive, as examiners must look at every test item. This can lead to examiners becoming fatigued and impatient and can encourage rushing through the items. Confidence is key – Examiners must be experts in
their field. This method relie s on the examiners being confident and consistent with their definition of a borderlinecandidate, and not just assuming an averagecandidate. Time and a place – You need a large number of
examiners for accuracy and reliability.
Construction of the Paper
The full curri culum will be sampled and usually
320 SBAs are submitted for consideration of
inclusion into the paper.
This will be a mixture of new SBAs and well
performing SBAs.
Around 60 SBAs will be removed as they are too
easy, difficult, ambiguous etc.
Another 20 SBAs are removed after further
analysis.
This leaves the required 240 SBA questions
needed for each diet of exams.
19
Paul Banaszkiewicz
SBA Checks before Finalisation of the Paper
The paper is sat by at least 10 examiners with
feedback obtained about the suitability of the included questions. A number of the included questions will be removed. Paper scrutiny by a senior experienced group of
examiners and removal of further SBAs that are deemed necessary. Head examiner. He/she has the final say for
approving the paper.
Standard Error of Measurement (SEM) and the GMC
If we consider that the whole exam had only 10 questions and all of the examiners independently concluded that 6 of every 10 borderline candidates would get each question correct, then a pass mark of 6 out of 10 (60%) would mean that 50% of borderline candidates would pass and 50% would fail. The pass mark therefore divides the borderline candidates down the middle. If the exam has a lot of hard questions, the pass mark will be lower. If there are a lot of easy questions, it will be higher. The mark is unique to each diet.
The Angoff-derived pass mark is not the mark determining eligibility to proceed, however. The GMC argue that there is some uncertainty in judge­ments made in this way, which can be expressed statistically as the Standard Error of Measurement (SEM). For patient safety reasons, the GMC would not want incompetent candidates being allowed to proceed, even if removing them means some potentially competent candidates are prevented from doing so.
The exam boards emphasis must be on patient safety. It is much more defensible to fail a candidate who is just more than competent than to pass one who is not competent.
The eligibility to proceed mark is therefore the Angoff-derived mark plus one SEM. When this step was first introduced, the historical performance of candidates scraping through was reviewed and it was noted that they went on to fail Section 2, so this rule in fact saves some candidates a whole lot of money!
The SEM allows us to identify what the GMC refer to as borderline candidateseither side of the cut score.
Borderline Candidates
There is an attempt to standardise the exam so that around 33% (1 standard deviation) of borderline (just competent) candidates will pass.
The exam mark is adjusted to prevent a large percentage of borderline candidates passing.
Consider Figure 2.9. If the middle dotted line is the Angoff determined cut scoreof the barely com­petent day one consultantthen the SEM represents what the GMC classes as the borderline candidate.
Applied to a bell-shaped curve then 34 +16 = 50% of time a barely competent candidate will pass.
If we were to add one SEM to the cut scorethen only 16% of time will the barely competent candidate pass.
- Anyone whose true scoreis below the true score
of the day one consultanttherefore has less than
a 16% probability of achieving a score above cut
score + 1 SEM.
The emphasis must be on patient safety. It is much more defensible to fail a candidate who is just more than competent than to pass one who is not
20
34,1%34,1%
2,1%2,1%
13,6%13,6%
60 90 120 150 180 210 240
Figure 2.9 Classic bell curveshape of normal distribution. The mean (or average) is the vertical line at the centre, and the vertical lines to either side represent intervals of one, two and three SDs. The percentage of data points that would lie within each segment of that distribution are shown.
0,1%0,1%
SBA Writing Process
Figure 2.10 Borderline candidate and Angoff-derived mark plus one SEM.
tomorrow, then it will show the same weight. But if the scale is not working properly and is not reliable, it could give you a different weight each time.
A reliable exam paper should produce the same
result if the test is repeated. Ideally, for an exam to be highly reliable, it should produce the same result if the candidate takesthe same test on two differentoccasions. If it produces very different results on separate occa­sions, it can be argued that the exam has low reliability and can therefore not be trusted as a means of grading or judging whether a candidate is competent.
Figure 2.11 Borderline candidates and Angoff-derived mark plus one SEM added to the cut score.
competent (Figure 2.10). Therefore we add 1 SEM to the cut score (Figure 2.11).
Reliability
Reliability is a Measure of Consistency
It is useful to think of kitchen scales. For weighing scales to be reliable, you would expect that if you weighed 80 kg 10 minutes ago you would still weigh the same now. If the scales said you weighed 40 kg now, you would not rely on those scales.
If the scale is reliable, then when you put a bag of
flour on the scale today and the same bag of flour on
How Can You Tell if a Test is Reliable (Cronbachsalpha)?
In most cases, we cannot get candidates to take the same test twice to measure reliability, so the internal consistency is measured with an alternative method. Internal consistency reliability measures the degree to which every test item measures the same construct.
The closer the score is to 1, the higher the reliabil-
ity will be. As a rough guide, exams should have a Cronbachs alpha of 0.8 or greater. With high-stakes exams, the reliability needs to be high because one exam is usually used to decide whether a student passes or fails. Around the world, few professional examinations, particularly in medical specialties, achieve a Cronbachs alpha of 0.8.
Validity
As well as being reliable, it is important that a test is valid, i.e. measures what it is supposed to measure. Continuing the kitchen scale metaphor, a scale might consistently show the wrong weight; in such a case, the scale is reliable but not valid.
21
Paul Banaszkiewicz
There are a number of ways in which validity can
be increased. These include the following:
Item analysis reporting. This flag questions that do not correlate well with the rest of the assessment. One of the reasons a question might get flagged is because participants who do well on other questions do not do well on this question – this could indicate the question lacks content validity.
Item bank. Use an item bank to store well­performing, defined topic questions.
Review and update the question bank frequently.
Orthopaedic knowledge base can change quickly with changing technology and changing regulations. Many questions that were valid 2 years ago are outdated today. Using an item bank allows you to update or retire questions that are no longer relevant.
Use assessment blueprinting. The blueprint confirms that the exam tests a representative sample of all the appropriate curriculum outcomes and a representative sample of all the curriculum content.
Involve subject matter experts (SMEs). The more you involve senior exam writers in assessment development, the more content validity you are likely to get. Get a panel of SMEs to rate each question as to whether it is essential, useful, but not essential,or‘not necessary’ to the performance of what is being measured. The more SMEs who agree that items are essential, the higher the content validity will be.
Blueprinting the Curriculum
To ensure adequate content coverage and, therefore, that the exam is a valid test of the breadth of kno w­ledge of orthopaedics, the process of blueprinting occurs. A spreadsheet is created that maps each of the questions to a learning objective on the curricu­lum. The assessment blueprint is a grid that plots the curriculum/assessment outcomes in columns against the curriculum content in rows.
Constructing SBAs
In order for an SBA to be good, it must fulfil two basic criteria: (1) what the question tests should be important and (2) the question is well structured.
SBAs consist of a stem (e.g. a clinical case presenta­tion) and a lead-in question, followed by a series of choices, typically onecorrect answer and four distractors.
The following sections outline an example of an
SBA.
Stem
A 60-year-old man presents to casualty 5 days following left ceramic on ceramic total hip replace­ment with increasing left hip pain. He was discharged from hospital 2 days after surgery with a plan for follow-up at 3 months. On examination, he is apyr­exial, normotensive with modest hip discomfort with movement.
Lead-In
Which of the following is the most likely diagnosis?
Options
A. Ceramic head fracture B. Ceramic liner fracture C. PJI D. Non-specific hip pain E. Squeaking ceramic interface
All options may be correct, but at 5 days non-specific hip pain is the most likely diagnosis. The patient has been discharged from hospital after 2 days and increased their activity levels. While it is important to rule out early periprosthetic joint infection (PJI), it should be rare, under 1%, and there would be more clues towards this in the stem such as a leaky wound or high temperature. Likewise, ceramic head or liner fractures are rare but possible causes of hip pain. With the fourth-generation Bioloc Delta ceramics, the incidence of ceramic head fractures has signifi­cantly reduced but there has been no corresponding change in liner fracture incidence. Likewise, if there was a ceramic fracture a bit more would be given in the history to point candidates in this direction. Squeaking is not uncommon following ceramic on ceramic total hip arthroplasty (CoC THA) (around 7%), and while it can be annoying is usually not painful.
Even though the less likely answers are not wrong,
they are less correct than the ‘keyed answer’.The candidate is instructed to select the most likely diagno­sis and experts would all agree that the most likely diagnosis is D. They would also agree that the other diagnoses are somewhat likely, but less likely than D. As long as the options can be laid out on a single continuum, in this case from Most Likely Diagnosis
22
SBA Writing Process
AB Least
Correct
Figure 2.12 SBA construction
to ‘Least Likely Diagnosis’, options in one-best-answer questions do not have to be totally wrong (Figure 2.12).
Options are homogeneous (i.e. all possible diag-
noses) and high-calibre candidates will be able to rank-order the options along a single dimension.
Well-constructed one-best-answer questions sat-
isfy the cover-the-optionsrule. The questions could be administered as write-in questions. The entire question is included in the stem.
CE D
Most
Correct
Example of Poorly Constructed SBA
Which of the following is true about pseudogout?
A. It occurs frequently in men. B. It is seldom associated with acute pain in a joint. C. It may be associated with a finding of
chondrocalcinosis.
D It is clearly hereditary in most cases. E. It responds well to treatment with naproxen.
This item is flawed. There is no lead-in. After reading the stem , the candidate has an unclear idea what the question is about. In an attempt to determine the bestanswer, candidates have to decide whether it occurs frequently in menis more or less true than it is seldom associated with acute pain in a joint. This is a comparison of apples and pears. Why are you trying to compare those things which you cant?
In order to rank-order the relative correctness of
options, the options must differ on a single dimension or else all options must be absolutely 100% true or false.
The diagram of these options would look like
Figure 2.13. The options are heterogeneous and deal with miscellaneous facts; they cannot be rank ordered from least to most true along a single dimension. Although this question appears to assess knowledge of several different points, its inherent flaws preclude this. The question by itself is not clear; the item cannot be answered without looking at the options.
It has also been proposed by the ICB to avoid
using negative A-type questions in SBAs.
The most problematic are those that take the
form: Each of the following is correct EXCEPTor Which of the following statements is NOT correct?
Figure 2.13 SBA formats
These suffer from the same problem as true/false questions: if options cannot be rank ordered on a single continuum, the examinees cannot determine either the leastor the mostcorrect answer.
On the other hand, the ICB occasionally uses well­focused negative A-types with single-word options, largely as a (poor) substitute for items that instruct the examinee to select more than one response.
Technical Item Flaws
This section is based on technical item flaws from chap­ter 3 in Paniagua and Swygert, constructing written test questions for thebasic and clinical sciences. provides specific examples of technical item flaws.
Two important types of technical item flaws are (1) testwiseness and (2) irrelevant difficulty.
Flaws related to testwiseness make it easier for some students to answer the question correctly, based on their test-taking skills alone. These flaws com­monly occur in items that are unfocused.
Flaws related to irrelevant difficulty make the question difficult for reasons unrelated to the trait that is the focus of assessment.
Examination boards are keen to eliminate flaws to provide a level playing field for testwise and not so testwise candidates.
8
The chapter
9
Flaws Related to Testwiseness
Grammatical Cue
One or more distractors do not follow grammatically or logically from the stem. Testwise candidates are able to spot this and eliminate them from the options.
23
Paul Banaszkiewicz
Logical Cue
If several but not all of the options are very similar, this can suggest that the answer is in this subset. Try to make the options as homogeneous as possible.
Longest Answer is Single Best Answer
The correct answer is longer, more specific or more complete than the other options.
Repeated Words
A word or phrase included in both the stem and the correct answer can act as a clue towards a correct guess.
UseofAbsoluteTerms
Absolute terms such as alwaysor nevershould not be used in options.
Presence of Convergence
This is less obvious to spot but occurs fairly com­monly. The correct answer includes the most elem­ents in common with the other options. The underlying premise is that the correct answer is the option that has the most in common with the other options; it is not likely to be an outlier.
Terms in the Options or the Stem are Vague
Vague frequency terms in the options such as often
or usuallyare used. Value frequency terms are not
consistently defined or interpreted even by experts.
True/False
If the options are all either completely correct or
completely incorrect, the question does not follow
the single best answerapproach and is most likely
pitched at the recall level. Options should be on a
continuum from least to most appropriate.
These types of questions are found in some ortho-
paedic MCQ books.
None of the Aboveis Used as an Option
The phrase none of the aboveis problematic in items
where judgement is involved and the options are not
absolutely true or false. Use of none of the above
essentially turns the item into a true/false item; each
option has to be evaluated as more or less true than
the universe of unlisted options .
Language or Structure of the Options is Not
Homogeneous or P arallel
The format and structure of each option is different.
Presence of Grouped or Collectively Exhaustive Options
A testwise student can identify a subset of options that cover all the possible outcomes (are collectively exhaustive) and rule out the options not in that subset.
Flaws Related to Irrelevant Difficulty
Overly Complicated or Long Questions
The stem contains extraneous reading and the options are very long and complicated. Trying to decide among these options requires a significant amount of reading because of the number of elements in each option.
Numerical Data are Not Consistently Presented
Numerical options should be listed in a consistent manner and not mixed (ranges and percentages).
24
Window Dressing
Window dressing in an MCQ is information in the
item itself which is superfluous to the content being
assessed. It is not about shorter versus longer stems
but about the relevance of what is there.
Unnecessarily Complicated Stems
Stems should be meaningful by themselves and
should present a definite probl em.
They should not contain irrelevant material that
may decrease a test scores reliability and validity.
Stems should include as much of the item as possible;
stems should be long and the options short.
Stems Contain Negative Phrasing
Items which ask a question in the negative (e.g. which
is the least appropriate ...’ ) are confusing to candi-
dates who have to switch between best and worst
answers and are contrary to the mode of reasoning
in most clinical situations. The stem should be
SBA Writing Process
negatively stated only when a significant learning outcome requires it.
Alternatives/Distractors
All alternatives should be plausible. Alternatives should be stated clearly and concisely. Items that are excessively wordy assess a candidate s reading ability rather than the attainment of the learning objective.
They should be homogeneous in content. Alternatives that are heterogeneous in content can provide cues to candidates about the correct answer.
Alternatives should be free from clues about which response is correct. Sophisticated test takers are alert to inadvertent clues to the correct answer, such as differences in grammar, length, formatting and language choice in the alternatives.
General Guidelines
Avoid using absolutes such as always’, ‘never, andallin the options; also avoid using vague terms such
as usuallyand frequently.
Avoid using all of the above and none of the above. They are problematic in ite ms where judge­ment is involved, and the options are not absolutely true or false. In either case, candidates can use partial knowledge to arrive at a correct answer.
Focus on important concepts; do not waste time testing trivial facts.
MCQ Question Writing Committee
This committee meets every 3 months with around 20 or so experienced examiners attending.
They begin by looking at some of the SBAs that have been flagged statistically as possibly poor per­formers. Some questions will already have been removed automatically – for example, all of the ques­tions that prove d too easy or too hard (usually new questions, as any question previ ously used would have passed this hurdle already).
The examiners will review each question and decide whether it is a fair question that should stay in the exam or is flawed and should be removed and returned to the question writers. Typical reasons for the latter would be ambiguity that had not previously been recognised, new evidence that has challenged the previously decreed correct answer, or simply that the answer in the bank is wrong.
It is worth noting that some very good questions end up being flagged as having possible wrong
answers yet remain in the exam. If a question is hard so that only 20% of candidates answer it correctly, then 80% will choose a wrong response. Lets say 40% chose one of the incorrect options – this flags as a possible wrong answer automatically, as more candi­dates have chosen a specific incorrect response than the correct one.
Once the poorly perform ing questions have been
dealt with, the Angoff procedure is performed.
New speciality question writers need to attend workshops in how to write SBAs before being allowed to submit questions. Existing question writers need to attend face-to-face meetings every 3 months. At these meetings, new questions are submitted to the panel for review and possible inclusion in the question bank. These meetings are expensive to hold, involving travel and possibly accommodation expenses but deemed worthwhile as they allow face­to-face scrutiny of the questions submitted for the question bank.
Basic Rules for Writing SBAs
Each item should focus on an important concept
or testing point.
Each item should assess application of knowledge,
not recall of an isolated fact.
The item lead-in should be focused, closed and
clear; the test taker should be able to answer the
item based on the stem and lead-in alone.
All options should be homogeneous and plausible,
to avoid cueing to the correct option.
Always review items to identify and remove
technical flaws that add irrelevant difficulty or
benefit savvy test takers.
One minute rule
Each question should be able to be read by the
candidate within 1 minute maximum.
If the stem is too lengthy it will need to be edited
into a readable length.
SBA Writing Guidelines: Dos and Donts
Getting Started
1. Identify an area of the curriculum blueprint that is
to be sampled.
2. Identify the topic area and the level of thinking that
you want to test and write a question around this
25
Paul Banaszkiewicz
which mimics tasks that successful candidates must be able to undertake at the next stage of training.
Ideally, questions should be pitched at the level
of integration/interpretation (questions which require putting the pieces together) and problem solving (questions which require clinical judgement’), not simple recall (questions which can be answered with a Google search).
Examples of recall questions to avoid:
What are the symptoms of X?
Which of the following is a contraindication
of X? What is the name/definition of this procedure?
Which of the following is correct?
Examples of higher order (putting the pieces
togetherand clinical judgement) questions:
In this procedure, which structur e is most at
risk in this patient? What is the most likely diagnosis?
What is the most useful investigation to carry
out at this stage? What is the most appropriate first step in the
management of this patient?
3. Construct the stem. This should present a single, clearly formulated problem, possibly in the form of a patient vignette. The stem should contain enough information to allow candidates to answer without referring to the options.
Patient vignettes may include any subset of the
following information:
a. The patient (age, gender), the presenting
complaint and its duration; any other relevant information a competent day one consultant would typically need to piece together to determine the best course of action.
b. Relevant details of the patients history,
possibly including details of family history.
c. Relevant physical findings, results of
diagnostic studies, initial treatment, subsequent findings, images, etc.
4. Constructthelead-ininsuchawaythatitbuildson theinformationinthestemandposesaclear question. Candidates should be able to answer without looking at the options and should not be abletoansweriftheinformationinthestemis masked.
5. Write the options. These should be of similar length, along a continuum, grammatically consistent and logically compatible. If appropriate, order the options in a logical order (e.g. numeric, alphabetical or anatomical). All the distractors must be plausible (think of the educational impact of suggesting that something potentially dangerous is plausible but not the best answer!).
Reviewi ng the Item
Does it focus on important problems relevant to
clinical practice? Can it be answered without looking at the
options? Are all the relevant facts included in the stem?
Can it be read and answered within approximately
one minute? Are all the options plausible, with one of them
standing out as being the best option? Does the question successfully avoid the pitfalls
listed in the Question Writing Checklist?
Question Writing Checklist
For a deeper dive, see Table 2.3.
A statistical analysis is performed on the candidate
scores to determine the mean, standard deviation, standard error, variation, quarterly percentiles and normality testing with histograms and normal prob­ability plots. These results are used to test the calibre and spread of candidates compared with previous groups.
Question-writing workshops provide detailed
guidance on the design of questions that discriminate between candidates of differing ability, in a format and a style that aid speed reading and comprehension.
Summary
There are many different types of MCQs. The trad­itional multiple true/false type that asks candidates to identify all the correct statements listed can assess knowledge and comprehension but is limited to these objectives.
By contrast, the single best answer question, which
asks candidates to choose the best answer from, say, five plausible alternatives, assesses not only know­ledge and comprehe nsion but also the application of this knowledge to the synthesis of deductions from
26
SBA Writing Process
Table 2.3 SBA question writing checklist
Checks across all question formats
Alignment Is the task required by the question congruent with the learning objectives at this stage
of career progression (e.g. interpretation of information, decision making)?
Blueprinting Are the topic and task listed in the exam syllabus and is the blueprint code provided?
Level of competence Is the question pitched to discriminate around the level of minimal competence required
for a pass?
Spelling and grammar Is the question free from spelling mistakes and grammatical errors?
Jargon Is the question jargon free (as much as possible within the topic area)? Is it free from
idioms which might unfairly disadvantage non-native English speakers?
Timing Can the question be read and answered within approximately 1 minute (for SBAs)?
Clarity Is it clear from the question what candidates are expected to do?
Additional checks: common pitfalls in writing questions
Grammatical cue The sentence structure can sometimes allow candidates to exclude a subset of the
options. To avoid this, make the lead-in a whole question rather than a sentence continuation.
Logical cue If several but not all of the options are very similar, this can suggest that the answer is in
this subset. Try to make the options as homogeneous as possible.
Absolute or vague terms ‘Neverand alwaysare almost never true; vague terms make it difficult to know what the
question writer meant (e.g. how often is often?).
Longest answer is single best answer
Repeated words Repeated words (or related words) between the stem and the options can point to the
Convergence strategy By looking at how many times each term is repeated across the set of options, candidates
Overly complicated questions
Negatively phrased items Items which ask a question in the negative (e.g. which is the least appropriate ...) are
True/false If the options are all either completely correct or completely incorrect, the question does
Option-dependent question
Window dressing Questions which can be answered without referring to the stem are typically pitched at
The answer with the most information/the highest level of precision is often the correct answer. Try to give all options a similar length and similar level of precision.
answer.
select the option which contains all the most frequently repeated terms.
Questions which are too long or present information in overly complex forms are in danger of testing skills which are not in line with the purpose of the exam (e.g. reading speed, working memory).
confusing to candidates who have to switch between best and worst answer and are contrary to the mode of reasoning in most clinical situations.
not follow the single best answerapproach and is most likely pitched at the recall level. Options should be on a continuum from least to most appropriate.
Candidates should be able to answer the question without referring to the options. This encourages them to think of the answer for themselves based on the information provided in the stem and lead-in.
the recall level. They also use up precious time in the examination by giving candidates irrelevant information to read.
27