Лингводидактическое тестирование по английскому языку в начальной школе. Учебное пособие
.pdf
Unit 4
RELIABILITY
Task 1: Reading the Gapped Text
The text you are going to read lacks some supporting information, e.g. argumentation, examples, references, etc. It can be found in the quotations from English language methodology readings collected the resource file that follows the text. Choose the information required and fill the appropriate letter in the table.
Every test should be reliable. In other words, a test should measure precisely whatever it is supposed to measure. If a group of students were to take the same test on two occasions, their results should be roughly the same – provided that nothing has happened in the interval ( such as one student receiving private tuition or several students comparing notes and specially preparing for the test when it is set a second time). Thus if students’ results are very different (e.g. the top students scoring low marks the second time), the test cannot be described as reliable.
It is possible to quantify the reliability of a test in the form of a reliability coefficient. Reliability coefficients are like validity coefficients. They allow us to compare the reliability of different tests. The ideal coefficient is 1. A test with a reliability coefficient of 1 is one which would give precisely the same results for a particular set of candidates regardless of when it happened to be administered. A test which had a reliability coefficient of zero would give sets of results quite unconnected with each other, in the sense that the score that someone actually got on a Wednesday would be no help at all in attempting to predict the score he or she would get if they took the test the day after. It is between the two extremes of 1 and zero that genuine test reliability coefficients are to be found.
1
The first requirement is to have two sets of scores for comparison. The most obvious way of obtaining reliability coefficients is to get a group
31
of subjects to take the same test twice. This is known as the test-retest method.
The drawbacks are not difficult to see. If the second administration of the test is too soon after the first, then subjects are likely to recall items and their responses to them, making the same responses more likely and the reliability spuriously high. If there is too long a gap between administrations, then learning or forgetting will have taken place, and the coefficient will be lower than it should be. However long the gap, the subjects are unlikely to be motivated to take the same test twice, and this too is likely to have a depressing effect on the coefficient. These effects are reduced somewhat by the use of two different forms of the same test (the alternative forms method). However, alternative forms are often simply not available.
2
The crucial question is how to make tests more reliable. It is known that there are two components of test reliability: the performance of candidates from occasion to occasion, and the reliability of the scoring. We will begin by suggesting ways of achieving consistent performance from candidates and then turn our attention to scorer reliability.
Take enough samples of behaviour. Other things being equal, the more items that you have on a test, the more reliable that the test will be. This seems intuitively right. If we wanted to know how good an archer was, we wouldn’t rely on the evidence of a single shot at the target. That one shot could be quite unrepresentative of their ability. To be satisfied that we had a really reliable measure of the ability we would want to see a large number of shots at the target.
3
Each additional item should as far as possible represent a fresh start for the candidate. By doing this we are able to gain additional information on all the candidates – information that will make test results more reliable. The use of the word ‘item’ should not be taken to mean only brief questions and answers. In a test of writing, for example, where candidates have to produce a number of passages, each of those passages is to be re-
32
garded as an item. The more independent passages there are, the more reliable will be the test. In the same way, in an interview used to test oral ability, the candidate should be given as many ‘fresh starts’ as possible. More detailed implications of the need to obtain sufficiently large samples of behaviour will be outlined later when considering the testing of particular abilities.
4
Exclude items which do not discriminate well between weaker and stronger students. Items on which stronger students and weak students perform with similar degrees of success contribute little to the reliability of a test. Statistical analysis of items can reveal which items do not discriminate well. These are likely to include items which are too easy or too difficult for the candidates, but not only such items. A small number of easy, non-discriminating items, may be kept at the beginning of a test to give candidates confidence and reduce the stress they feel.
Write unambiguous items. It is essential that candidates should not be presented with items whose meaning is not clear or to which there is an acceptable answer which the test writer has not anticipated. In a reading test there was an open-ended question, based on a lengthy reading passage about English accents and dialects. Where does the author direct the reader who is interested in non-standard dialects of English? The Expected answer was the Further reading section of the book. A number of candidates answered ‘page3’, which was the place in the text where the author actually said that the interested reader should look in the Further reading section. Only the alertness of those scoring the test revealed that there was a completely unanticipated correct answer to the question. If that had not happened, a correct answer would have been scored as incorrect. The fact that an individual candidate might interpret the question in different ways on different occasions means that the item is not contributing fully to the reliability of the test.
5
Provide clear and explicit instructions. This implies both to written and oral instructions. It is possible for candidates to misinterpret what
33
they are asked to do, then on some occasions some of them certainly will. It is by no means always the weakest candidate who is able to provide the alternative interpretation. A common fault of tests written for the students of a particular teaching institution is the supposition that the students all know what is intended by carelessly worded instructions. The frequency of the complaint that students are unintelligent, have been stupid, have willfully misunderstood what they were asked to do, reveals that the supposition is often unwarranted.
6
Use items that permit scoring which is as objective as possible. This may appear to be a recommendation to use multiple choice items which permit completely objective scoring. This is not intended. While it would be a mistake to say that multiple choice items are never appropriate, it is certainly true that there are many circumstances in which they are quite inappropriate. What is more, good multiple choice items are notoriously difficult to write and always require extensive pre-testing.
7
Agree accepting responses and appropriate scores at outset of scoring.
A sample of scripts should be taken immediately after the administration of the test. Where there are compositions, archetypical representatives of different levels of ability should be selected. Only when all scorers are agreed on the scores to be given to these should real scoring begin.
8
Reliability and validity. To be valid a test must provide consistently accurate measurements. It must therefore be reliable. A reliable test, however, may not be valid at all. For example, as a writing test we could require candidates to write down the translation equivalents of 500 words in their own language. This might well be a reliable test; but it is unlikely to be a valid test of writing.
34
9
Resource File
A The same is true for language testing. It has been demonstrated empirically that the addition of further items will make a test more reliable. There is even a formula (the Spearman-Brown formula) that allows one to estimate how many extra items similar to the ones already in the test will be needed to increase the reliability coefficient to a required level. One thing to bear in mind, however, is that the additional items should be independent of each other and of existing items. Imagine a reading test that asks the question: ‘Where did the thief hide the jewels?’ If the additional item following that took the form, ‘What was unusual about the hiding place?’, it would not make a full contribution to an increase in the reliability of the test. Why not? Because it is hardly possible for someone who got the original question wrong to get the supplementary question right. Such candidates are effectively prevented from answering the additional question; for them, in reality, there is no additional question. We do not get an additional sample of their behaviour, so the reliability of our estimate of their ability is not increased.
B It turns out, surprisingly, that the most common methods of obtaining the necessary two sets of scores involve only one administration of one test. Such methods provide us with a coefficient of internal consistency. The most basic of these is the split half method. In this the subjects take the test in the usual way, but each subject is given two scores. One score is for one half of the test, the second score is for the other half. The two sets of scores are then used to obtain the reliability coefficient as if the whole test had been taken twice. In order for this method to work, it is necessary for the test to be split into two halves which are really equivalent, through the careful matching of items (in fact where items in the test have been ordered in terms of difficulty, a split into odd-numbered items and even-numbered items may be adequate). It can be seen that this method is rather like the alternative forms method, except that the two ‘forms’ are only half the length.
It has been demonstrated empirically that this altogether more economical method will indeed give good estimates of alternate forms coef-
35
ficients, provided that the alternate forms are closely equivalent to each other.
C The best way to arrive at unambiguous items is, having drafted them, to subject them to the scrutiny of colleagues, who should try as hard as they can to find alternative interpretations to the ones intended. If this task entered into the right spirit – one of good-natured perversity – most of the problems can be identified before the test is administered. Pre-testing of the items on a group of people comparable to those for whom the test is intended should reveal the remainder. Where pretesting is not practicable, scorers must be on the lookout for patterns of response that indicate that there are problem items.
D While it is important to make a test long enough to achieve satisfactory reliability, it should not be made so long that the candidates become bored or tired that the behaviour they exhibit becomes unrepresentative of their ability. At the same time, it may often be necessary to resist pressure to make attest shorter than is appropriate. The usual argument for shortening a test is that it is not practical for it to be longer. The answer for this is that accurate information does not come cheaply: if such information is needed, then the price is to be paid. In general, the more important the decisions based on a test, the longer the test should be.
E Test writers should not rely on the students’ powers of telepathy to elicit the desired behaviour. Again the use of colleagues to criticize drafts of instructions (including those which will be spoken) is the best way of avoiding problems. Spoken instructions should be read from a prepared text to avoid introducing confusion.
Ensure that tests are well laid out and perfectly legible. Too often, institutional tests are badly typed (or handwritten), have too much text in too small a space, and are poorly reproduced. As a result, students are faced with additional tasks which are not ones meant to measure their language ability. Their variable performance on the unwanted tasks will lower the reliability of a test.
F Certain authors have suggested how high a reliability coefficients we should expect for different types of language tests. Lado (1961), for
36
example, says that good vocabulary, structure and reading tests are usually in the .90 to .99 range, while auditory comprehension are more often in the .80 to .89 range. Oral production tests may be in the .70 to .79 range. He adds that a reliability coefficient of .85 might be considered high for an oral production test but low for a reading test. These suggestions reflect what Lado sees as the difficulty in achieving reliability in the testing of the different abilities. In fact the reliability coefficient that is to be sought will depend also on other considerations, most particularly the importance of the decisions that are to be taken on the basis of the test. The more important the decisions, the greater reliability we must demand: if we are to refuse someone the opportunity to study overseas because of their score on a language test, then we have to be pretty sure that their score would not have been much different if they had taken the test a day or two earlier or later. Later it will be explained how the reliability coefficient can be used to arrive at another figure to estimate likely differences of this kind. Before this is done, however, something has to be said about the way in which reliability coefficients are arrived at.
G An alternative to multiple choice is the open-handed item which has a unique, possibly one-word, correct response which the candidates produce themselves. This too should ensure objective scoring, but in fact problems with such matters as spelling which makes a candidate’s meaning unclear (say, in a listening test) often make demands on the scorer’s judgement. The longer the required response, the greater the difficulties of this kind. One way of dealing with this is to structure the candidate’s response by providing part of it. For example, the open-ended question,
What was different about the results? May be designed to elicit the response, Success was closely associated with high motivation. This is likely to cause problems with scoring. Greater scorer reliability will probably be achieved if the question is followed by:
…………was closely associated with ……….
This reinforces the suggestion already made that candidates should not be given a choice of items and that they should be limited in the way they are allowed to respond. Scoring the compositions all on one topic will be more reliable than if the candidates are allowed to choose from six topics.
37
H In our efforts to make tests reliable, we must be wary of reducing their validity. Earlier it was admitted that restricting the scope of what candidates are permitted to write in a composition might diminish the validity of the task. This depends in part on what exactly we are trying to measure by setting a task. If we are interested in candidates’ ability to structure a composition, then it would be hard to justify providing them with a structure in order to increase reliability. At the same time we would still try to restrict candidates in ways which would not render their performance on the task invalid.
There will always be some tension between reliability and validity. The tester has to balance gains in one against losses in the other.
I For short answer questions, the scorers should note any difficulties they have in assigning points (the key is unlikely to have anticipated every relevant response), and bring these to the attention of whoever is supervising that part of the scoring. Once a decision has been taken as to the points to be assigned, the supervisor should convey it to all the scorers concerned.
Identify candidates by numbers, not by name. Scorers inevitably have expectations of candidates that they know. Except in purely objective testing, this will affect the way that they score. Studies have shown that even where the candidates are unknown to the scorers, the name on a script (or a photograph) will make a significant difference to the scores given. For example, a scorer may be influenced by the gender or nationality of a name into making predictions which affect the score given. The identification of candidates only by number will reduce such effects.
1 |
2 |
3 |
4 |
5 |
6 |
7 |
8 |
9 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
38
Task 2: Discussion
Read the text carefully and write a summary based on it. Get ready to discuss the following points.
1.What test is considered to be reliable?
2.What does a reliability coefficient allow us to do?
3.What one should take into account when trying to achieve reliability for different types of language tests?
4.How can you describe the test-retest method?
5.What is the specific feature of the alternative forms method?
6.What does the split-half method provide us with?
7.What ways of making tests more reliable can you suggest?
8.Why is it important to make tests long enough?
9.Why is it necessary to exclude items which do not discriminate well between weaker and stronger students?
10.Why is it important to have reliable scoring? How to achieve reliable scoring?
Task 3: Activities
Try to estimate your institutional tests. Taking into account the suggestions listed in the text think of the ways you could improve their reliability.
39
Unit 5 CONSTRUCTING A TEST FOR YOUR
CLASSROOM
Task 1: Reading the Gapped Text
The text you are going to read lacks some supporting information, e.g. argumentation, examples, references, etc. It can be found in the quotations from English language methodology readings
collected the resource file that follows the text. Choose the information required and fill the appropriate letter in the table.
Should we write our own tests? The answer to this question depends on your ability as a test writer and on the time which you have available. It also depends on the quality of the books and tests which you can obtain.
We often buy tests to find out how our students are progressing. We want to learn about what kinds of problems our students are experiencing so that we can plan appropriate lessons or give them additional teaching. Always remember, however, that the best tests for the classroom are those tests which you write yourself.
Then another question arises. How often should students be tested? This question is difficult to answer as the type of tests you set will determine how often they are given. Your reasons for testing as well as the students’ own needs will also help you to determine how frequently you test your students. There is clearly a great difference between formal exams and informal classroom tests. Formal exams can generally be given to students once a year or perhaps every term at most. Such exams are usually intended to measure achievements and are used primarily to compare students’ performances. Informal classroom tests, on the other hand, can be set far more frequently. These informal tests are often progress tests and are used to diagnose difficulties as well as to encourage students.
40
