Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5255_Библиотеки_им_академика_М_И_Перельмана
.pdf
52
H0= 70 hypothesized (estimated) value of the population mean
n = 20 (sample size)
x = 60 (sample mean)
s = 11 (sample standard deviation)
Null hypothesis: the true population mean uptake value is 70 or H0: = 70
Alternative hypothesis: the mean uptake value is not 70. H1: 70
Level of significance for testing this hypothesis. = .01
Counting Statistics
Since the population standard deviation ( S.D.
sample standard deviation ( S.D.
S.D.
Population
= S.D.
Sample
= 11
) therefore
Sample
Population
Now we can compute the standard error of the mean ( S.E.
) is not known, we take it from the
). Since we are using an
Mean
estimate of the population standard deviation, the standard error of the mean will also be an
estimate. We can use
S.E.
Mean
= S.D.
Population
/ n
= 11 / 20
= 11 / 4.47
= 2.46 (estimated standard error of the mean)
As the instructor is interested in knowing whether the true mean uptake value is larger
or smaller than the hypothesized value, a two-tailed test is the most appropriate to use in
this situation. The significance level of 0.01 means, two areas each one containing 0.005 of
the area under the t-distribution. The table value of t for degree of freedom = 19 at 1% level
is 1.729 (given).
This value is the appropriate one to use in calculating the limits of the acceptance
region:
+ 1.729 S.E.
H0
= 70 + 1.729 (2.46)
Mean
= 70 + 4.25
= 74.25 (upper limit)
- 1.729 S.E.
H0
= 70 - 1.729 (2.46)
Mean
= 70 - 4.25
= 65.75 (lower limit)
The acceptance region is 74.25 and 65.75, and the sample mean is 60. It can be seen that

Counting Statistics
53
the sample mean lies outside the acceptance region. Therefore, instructor would reject the
null hypothesis that the true population mean uptake value is 70.
To test the difference between means of two samples (independent samples)
Example:Two different types of dose large (90Gy) and small (60Gy) were tried for
treatment of Graves’ disease. Five persons were given large dose (90Gy) and seven persons
were given small (60Gy). Treatment response on numerical scale of 20 is given below:
Large (90Gy) : 9 13 12 3 8
Small (60Gy): 15 10 8 6 11 12 8
Do these two doses differ significantly with regard to their effect in treatment response
(the table value of t for degree of freedom = 10 at 5% level is 2.23)?
Null hypothesis: The two doses large and small do not differ significantly with regard
to their effect in treatment response.
X1 X1- X–1 (X1- X–1)
2
X2 X2 - X–2 (X2 - X–2)
2
9 0 0 15 +5 25
13 +4 16 10 0 0
12 +3 9 8 -2 4
3 -6 36 6 -4 16
8 -1 1 11 +1 1
12 +2 4
8 -2 4
X1= 45; (X1-X–1)= 0; (X1-X–1)2 = 62; X2= 70; (X2-X–2)=0; (X2- X–2)2 = 54
2
s
= (X1 -X–1)2 / (n1-1) s
1
2
s
= 62/4 s
1
2
= (X2 - X–2)2 / (n2-1)
1
2
= 54/6
1
Pooled estimate of standard deviation (sp)
2
s
= [(n1-1) s
p
2
+ ( n2-1) s
1
2
] /( n1 + n2 –2 )
2
= [(5-1) (62/4) + (7-1) (54/6)]/ (5 + 7 –2)
= [4 (62/4) + 6 (54/6)] / 10
= 116/10
= 11.6
sp= 11.6 = 3.406

54
Counting Statistics
S.E.
the difference between means
= sp ( 1 /n1 + 1 /n2 )
= (3.406) ( ( 1 /5 + 1 /7 ) )
= (3.406) ( ( 0.20 + 0.14 ) )
= (3.406) ( (0.34)
= (3.406) (0.58554)
= 1.994278
Since both sample sizes are less than 30, the t distribution with (5 +7 –2) = 10 degrees
of freedom is the appropriate sampling distribution. The t value for 0.05 of the area under
the curve is 2.23 (given). The hypothesized value of
(x1 - x2),
the mean of the sampling
distribution of the difference between the two doses, is equal to zero. Thus, the calculation
for determining the limit of the acceptance region is equal to:
(x1 - x2 )
+ t-value for 0.05 ( degree of freedom = 10 ) × S.E.
= 0 + 2.23 × (1.994278)
= 0 + 4.447239
the difference between means
= 4.447239 (upper limit)
The limit of the acceptance region is from 0 to 4.447239 (upper limit) and the mean
difference between the treatment response for two doses (90 Gy and 60 Gy) is 1(10 – 9 = 1).
Thus the difference between the two samples means lies within the acceptance region.
Therefore, we accept the null hypothesis that the two doses large (90 Gy) and small (60Gy)
do not differ significantly with regard to their effect in treatment response.
Testing differences between means with dependent samples
Example:A health care center claims that the average participant in their weight reduction
program loses at least 17 Pounds. A participant X, who is interested in the program, however,
is doubtful about the claims and asks for some hard evidence. The center allows him to
select randomly ten participants and record their weights before and after the program. The
data are recorded in table 1.
Table 1: Weights (In Pounds) before and after weight-reducing program.
Before 189 195 219 215 209 184 190 179 208 215
After 170 178 197 200 180 161 174 160 186 200
Mr. X wants to test the claimed average weight loss of at least 17 Pounds at the 5
percent significance level. (using t value from the t-table for degree of freedom = 9 at 5%
level is 1.833).
Null hypothesis: the average weight loss is only 17 Pounds or H0:
1 - 2
= 17

Counting Statistics
55
Alternative hypothesis: the average weight loss exceeds 17 Pounds. H1:
1 - 2
as the future participants might be interested for weight loss of 17 kg or more)
Level of significance for testing this hypothesis.
= .05
We are interested in the weight loss. If the population of weight losses has a mean L we
can rewrite the hypotheses as:
H0 : L= 17
H1 : L> 17
Now we compute the individual losses, their mean and standard deviation. The
computations are shown in table-2.
Table 2: Mean weight loss and its standard deviation
Before After Loss loss squared
x x
2
189 170 19 361
195 178 17 289
219 197 22 484
> 17 (
215 200 15 225
209 180 29 841
184 161 23 529
190 174 16 256
179 160 19 361
208 186 22 484
215 200 15 225
x = 197 x2 = 4,055
Mean = 197/10 = 19.7
Standard deviation = 4.40
Since the population standard deviation ( S.D.
using the sample standard deviation ( S.D.
S.D.
Population
= S.D.
Sample
= 4.40
Sample
Population
) and
) is not known, it must be estimated

56
Standard error of the mean:
Counting Statistics
S.E.
Mean
= S.D.
Population
/ n
= 4.40 / 10
= 4.40 /3.16
= 1.39 estimated standard error of the mean
Since Mr. X wants to know if the mean weight loss exceeds 17 Pounds, an upper tailed
test is appropriate. We use the t value to calculate the upper limit of the acceptance region:
=
+ t-value for 0.05 (degree of freedom = 9) × S.E.
=
H0
H0
+ 1.833 S.E.
= 17 + 1.833 × (1.39)
Mean
Mean
= 17 + 2.55
= 19.55 Pounds upper limit
The sample mean is 19.7 and the acceptance range is 17 to 19.55. We see that the sample
mean lies outside the acceptance region, so the claimed weight loss in the program is
justifiable.
Cautions while using t-test
The conclusion arrived at on the basis of the t-test are justified only if the assumption
upon which the test is based are true. If the actual distribution is not normally distributed,
then, strictly speaking, the t-test is not justified for small samples. It is a good idea to check
the normality assumption. A review of similar samples or related research may provide
evidence as to whether or not the population is normally distributed.
The chi square test
The reliability of counting instruments in nuclear medicine can best be evaluated by chi
square (2) statistics. For this test a minimum of 20 measurements are recommended. The
mean of the measurement is calculated and chi square is then estimated from the mean as
follows:
nxxMean
/)(
i
22
xxx
From 2 table, p value is obtained for a given set of measurements. A p value of 0.02 to
0.98 is treated as acceptable. In fact p value of 0.5 would be ideal to indicate that observed
chi square value is in the middle of the range expected for Poisson distribution. A low p
value shows a small probability that a Poisson distribution would give the chi square value
/)(
i

Counting Statistics
57
as large as actually observed and indicates the presence of additional sources of random
error. With a p value just smaller than acceptable the experiment suggests repeat measurement,
as the results are suspicious in such cases. Similarly a high p value between 0.98 and 0.99 is
also considered to be suspicious and experiment needs to be repeated. For p >0.99 the
random variations are much smaller than expected for a Poisson distribution.
Example: A radioactive sample placed in a well counter gives following measurement
for a time duration of 10 seconds. Find out with chi square test whether the distribution of
the data is acceptable.
S.No. Total counts
1. 3242
2. 3201
3. 3283
4. 3198
5. 3286
6. 3291
7. 3249
8. 3189
9. 3295
10. 3160
11. 3324
12. 3169
13. 3333
14. 3170
15. 3134
16. 3172
17. 3312
18. 3345
19. 3290
20. 3152
Mean 3239.75 and chi square value = 28531.75/3239.75 = 8.806
From chi square table the probability will be around 0.97. As mentioned earlier the
value between 0.02 and 0.98 is considered acceptable. Thus measured values in the example
can be considered to be precise.

58
Counting Statistics
Medical decision making and principle of ROC
The public awareness in recent years has increased to an extent that they want the
justification of the cost and possible risks of a diagnostic procedure. It is important to
measure the quality of diagnostic information and diagnostic decisions. Here we will deal
with the quality and diagnostic decision of a test. To compare two test modalities one has to
choose a parameter/index.
The simplest parameter/index could be “accuracy”. This is the fraction of cases for
which test provides the correct result. We always require accuracy to be very high (near
100%). This parameter has to be used carefully as it may mislead in some situations. For
example in screening relatively rare disease when the test can be very accurate, simply by
ignoring all evidence and calling all cases negative. If only 5% of patients have the disease
in question, a test, which always blindly states that the disease is absent, will be right
(accurate) 95% of the time. Similar situation arises, when the disease is common (high
prevalence).
It is obvious from above example that disease prevalence affects the accuracy quite
significantly. Even if disease prevalence is known and fixed, this index has limited role to
play. For example, there may be two different tests with equal accuracy but different false
positive and false negative decisions. One might provide almost all false negative decisions,
while other nearly all false positive decisions. Thus the usefulness of these two tests for
patient management could be quite different in different situations.
To make the accuracy a meaningful index, prevalence of the disease, sensitivity and
specificity should also be included in decision-making. Sensitivity and specificity are defined
as follows:
Sensitivity = Number of True Positive (TP) decisions / Number of actually positive
cases (TP+FN)
Specificity = Number of True Negative (TN) decisions / Number of actually negative
cases (TN+FP)
The “sensitivity” represent a kind of accuracy for actually positive cases while “specificity
another kind of accuracy for actually negative cases. It is necessary to specify the basis of
determining true positive/true negative (gold standard) while calculating sensitivity and
specificity.
Disease prevalence is defined as:
Prevalence = (TP+FN)/(TP+FN+TN+FP)
Disease prevalence is the characteristic of the patient population group. Figure 3
(hypothetical curve) shows a distribution of normal and abnormal (ill) persons in the
population. Y-axis represents the number of persons who are normal (left part of figure 3) or
abnormal (right part). Disease prevalence is the area under the distribution curve for diseased

Counting Statistics
250
Normal
59
persons compared to the area under both the curves in figure 3. If prevalence is high, the test
with high sensitivity should be chosen while in case of low prevalence test with higher
specificity should be preferred.
Accuracy = (True positive (TP) +True negative (TN))/(Total cases (TP+TN+FP+FN)
Abnormal
No. Of Persons
-100 -50 0 50 100 150 200
Test Result
Figure 3: Distribution of normal and abnormal subjects in the population
Accuracy of a test depends on disease prevalence for the same sensitivity and specificity.
It is therefore quite possible that the accuracy for a given test at two centres may be
different due to their different disease prevalence. Accuracy can also be expressed in terms
of sensitivity, specificity and prevalence.
Accuracy = sensitivity x prevalence + specificity x (1-prevalence)
Positive predictive value (PPV) : This is a conditional probability of having disease
when the test shows true positive result.
PPV = TP/(TP+FP)
Similarly negative predictive value (NPV) = TN/(TN+FN)
Other parameters, which are mentioned at times in medical decision making are:
True positive fraction (TPF = p( T+ / D+ ) is simply the same as sensitivity, and true
negative fraction (TNF= p(T- / D-), is simply the same as specificity.

60
Counting Statistics
False Positive fraction (FPF = p(T+ / D-) = No. of false positive decisions / No. of
actually negative cases.
False Negative fraction (FNF= p(T- / D+) = No. of false negative decisions / No. of
actually positive cases.
Example : Result of 10 volunteers (7 normal individuals and 3 ill patients) are given in
the following table. The bold face letters indicate the normal individuals. Using a threshold
value 20 (as positive) label them as TP, FP, TN, FN .
A B C D E F G H I J
10 5 20 40 60 9 70 50 80 15
Answer:
A B C D E F G H I J
10 5 20 40 60 9 70 50 80 15
FN TN TP FP FP TN TP FP FP TN
Example : In 100 patients who underwent thyroid uptake measurement, the test results
are:
Negative Positive
60 (TN) 10 (FP)
5 (FN) 25 (TP)
Sensitivity = 25/(25+5) = 83.3%
Specificity = 60/(60+10) = 85.7%
Accuracy = (60+25)/100 = 85%
Because of the variations in biological response, possible statistical noise in the data,
and other technological deficiencies, a medical test some times produces identical results
for abnormal (disease) and normal patients. Usually a normal range for normal people is
established and if the results fall outside this normal range they are termed as positive (ill).
Usually a threshold value is chosen arbitrarily, and different choices yield different frequencies
for various kinds of correct and incorrect decisions. A threshold should therefore be selected
in a judicious manner, which can make a good balance between sensitivity and specificity.
Receiver Operating Characteristic (ROC) Curve
This technique is used in medical decision-making. This does not depend on disease prevalence
or observer’s choice of decision threshold. It requires a “gold” standard or known truth. It
can be applied to computer simulated, phantom or clinical data. It can be better illustrated
through example.

Counting Statistics
61
There are 200 studies with pathological confirmation that 100 are positive and 100 are
negative. Observer evaluates each study and ranks each study on the 5 point scale as follows:
Definitely positive = 1
Probably positive = 2
Equivocal = 3
Probably negative = 4
Definitely negative = 5
After the evaluation, the following table is generated
Decision Threshold 1 2 3 4 5 Sum
Positive 40 26 20 10 4 100
Negative 6 12 22 24 36 100
There are 100 positive cases and 100 negative Cases.
Table 3
Decision Threshold 1 2 3 4 5 Sum
Positive 40 26 20 10 4
Negative 6 12 22 24 36
Table 3 needs some explanation. With decision threshold as (definitely positive=1),
observer labeled 46 cases as definitely positive, selecting 40 as positive from 100 positive
cases and 6 as positive from 100 negative cases. With decision threshold = 2 (probably
positive = 2) user labeled 26 cases as probably positive, selecting 26 as positive from 100
positive cases and 12 as positive from 100 negative cases. It is to be noted, with decision
threshold = 2, user has finally labeled 26+40 = 66 as True Positive (TP) and 6 + 12 = 18 as
False Positive (FP), 34 as False Negative (FN) and 82 as True Negative (TN). Similar
explanation is valid for decision threshold 3,4 and 5 respectively. Table 4 given below shows
the sensitivity (TP /(TP +FN)) and specificity (TN / (TN + FP)) calculated for each threshold.
Table 4
Decision threshold 1 2 3 4 5
Sensitivity 0.40 0.60 0.86 0.90 1.00
Specificity 0.94 0.82 0.60 0.36 0.0
Соседние файлы в папке Библиотека им академика М.И. Перельмана
