Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5255_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
15.09.2026
Размер:
11 Мб
Скачать
☆
52
H0= 70 hypothesized (estimated) value of the population mean
n = 20 (sample size) x = 60 (sample mean)
s = 11 (sample standard deviation)
Null hypothesis: the true population mean uptake value is 70 or H0: = 70
Alternative hypothesis: the mean uptake value is not 70. H1: 70
Level of significance for testing this hypothesis. = .01
Counting Statistics
Since the population standard deviation ( S.D. sample standard deviation ( S.D.
S.D.
Population
= S.D.
Sample
= 11
) therefore
Sample
Population
Now we can compute the standard error of the mean ( S.E.
) is not known, we take it from the
). Since we are using an
Mean
estimate of the population standard deviation, the standard error of the mean will also be an estimate. We can use
S.E.
Mean
= S.D.
Population
/ n = 11 / 20 = 11 / 4.47 = 2.46 (estimated standard error of the mean)
As the instructor is interested in knowing whether the true mean uptake value is larger or smaller than the hypothesized value, a two-tailed test is the most appropriate to use in this situation. The significance level of 0.01 means, two areas each one containing 0.005 of the area under the t-distribution. The table value of t for degree of freedom = 19 at 1% level is 1.729 (given).
This value is the appropriate one to use in calculating the limits of the acceptance region:
+ 1.729 S.E.
H0
= 70 + 1.729 (2.46)
Mean
= 70 + 4.25 = 74.25 (upper limit)
- 1.729 S.E.
H0
= 70 - 1.729 (2.46)
Mean
= 70 - 4.25 = 65.75 (lower limit)
The acceptance region is 74.25 and 65.75, and the sample mean is 60. It can be seen that
Counting Statistics
53
the sample mean lies outside the acceptance region. Therefore, instructor would reject the null hypothesis that the true population mean uptake value is 70.
To test the difference between means of two samples (independent samples)
Example:Two different types of dose large (90Gy) and small (60Gy) were tried for treatment of Graves’ disease. Five persons were given large dose (90Gy) and seven persons were given small (60Gy). Treatment response on numerical scale of 20 is given below:
Large (90Gy) : 9 13 12 3 8
Small (60Gy): 15 10 8 6 11 12 8
Do these two doses differ significantly with regard to their effect in treatment response (the table value of t for degree of freedom = 10 at 5% level is 2.23)?
Null hypothesis: The two doses large and small do not differ significantly with regard to their effect in treatment response.
X1 X1- X–1 (X1- X–1)
2
X2 X2 - X–2 (X2 - X–2)
2
9 0 0 15 +5 25 13 +4 16 10 0 0 12 +3 9 8 -2 4
3 -6 36 6 -4 16
8 -1 1 11 +1 1
12 +2 4
8 -2 4
X1= 45; (X1-X–1)= 0; (X1-X–1)2 = 62; X2= 70; (X2-X–2)=0; (X2- X–2)2 = 54
2
s
= (X1 -X–1)2 / (n1-1) s
1
2
s
= 62/4 s
1
2
= (X2 - X–2)2 / (n2-1)
1
2
= 54/6
1
Pooled estimate of standard deviation (sp)
2
s
= [(n1-1) s
p
2
+ ( n2-1) s
1
2
] /( n1 + n2 –2 )
2
= [(5-1) (62/4) + (7-1) (54/6)]/ (5 + 7 –2) = [4 (62/4) + 6 (54/6)] / 10 = 116/10 = 11.6
sp= 11.6 = 3.406
54
Counting Statistics
S.E.
the difference between means
= sp ( 1 /n1 + 1 /n2 )
= (3.406) ( ( 1 /5 + 1 /7 ) )
= (3.406) ( ( 0.20 + 0.14 ) ) = (3.406) ( (0.34) = (3.406) (0.58554) = 1.994278
Since both sample sizes are less than 30, the t distribution with (5 +7 –2) = 10 degrees of freedom is the appropriate sampling distribution. The t value for 0.05 of the area under the curve is 2.23 (given). The hypothesized value of
(x1 - x2),
the mean of the sampling distribution of the difference between the two doses, is equal to zero. Thus, the calculation for determining the limit of the acceptance region is equal to:
(x1 - x2 )
+ t-value for 0.05 ( degree of freedom = 10 ) × S.E.
= 0 + 2.23 × (1.994278)
= 0 + 4.447239
the difference between means
= 4.447239 (upper limit)
The limit of the acceptance region is from 0 to 4.447239 (upper limit) and the mean difference between the treatment response for two doses (90 Gy and 60 Gy) is 1(10 – 9 = 1). Thus the difference between the two samples means lies within the acceptance region. Therefore, we accept the null hypothesis that the two doses large (90 Gy) and small (60Gy) do not differ significantly with regard to their effect in treatment response.
Testing differences between means with dependent samples
Example:A health care center claims that the average participant in their weight reduction program loses at least 17 Pounds. A participant X, who is interested in the program, however, is doubtful about the claims and asks for some hard evidence. The center allows him to select randomly ten participants and record their weights before and after the program. The data are recorded in table 1.
Table 1: Weights (In Pounds) before and after weight-reducing program.
Before 189 195 219 215 209 184 190 179 208 215 After 170 178 197 200 180 161 174 160 186 200
Mr. X wants to test the claimed average weight loss of at least 17 Pounds at the 5 percent significance level. (using t value from the t-table for degree of freedom = 9 at 5% level is 1.833).
Null hypothesis: the average weight loss is only 17 Pounds or H0:
1 - 2
= 17
Counting Statistics
55
Alternative hypothesis: the average weight loss exceeds 17 Pounds. H1:
1 - 2
as the future participants might be interested for weight loss of 17 kg or more)
Level of significance for testing this hypothesis.
= .05
We are interested in the weight loss. If the population of weight losses has a mean L we can rewrite the hypotheses as:
H0 : L= 17 H1 : L> 17
Now we compute the individual losses, their mean and standard deviation. The computations are shown in table-2.
Table 2: Mean weight loss and its standard deviation
Before After Loss loss squared
x x
2
189 170 19 361 195 178 17 289 219 197 22 484
> 17 (
215 200 15 225 209 180 29 841 184 161 23 529 190 174 16 256 179 160 19 361 208 186 22 484 215 200 15 225
x = 197 x2 = 4,055
Mean = 197/10 = 19.7 Standard deviation = 4.40 Since the population standard deviation ( S.D.
using the sample standard deviation ( S.D.
S.D.
Population
= S.D.
Sample
= 4.40
Sample
Population
) and
) is not known, it must be estimated
56
Standard error of the mean:
Counting Statistics
S.E.
Mean
= S.D.
Population
/ n = 4.40 / 10 = 4.40 /3.16 = 1.39 estimated standard error of the mean
Since Mr. X wants to know if the mean weight loss exceeds 17 Pounds, an upper tailed
test is appropriate. We use the t value to calculate the upper limit of the acceptance region:
=
+ t-value for 0.05 (degree of freedom = 9) × S.E.
=
H0
H0
+ 1.833 S.E.
= 17 + 1.833 × (1.39)
Mean
Mean
= 17 + 2.55 = 19.55 Pounds upper limit
The sample mean is 19.7 and the acceptance range is 17 to 19.55. We see that the sample mean lies outside the acceptance region, so the claimed weight loss in the program is justifiable.
Cautions while using t-test
The conclusion arrived at on the basis of the t-test are justified only if the assumption upon which the test is based are true. If the actual distribution is not normally distributed, then, strictly speaking, the t-test is not justified for small samples. It is a good idea to check the normality assumption. A review of similar samples or related research may provide evidence as to whether or not the population is normally distributed.
The chi square test
The reliability of counting instruments in nuclear medicine can best be evaluated by chi square (2) statistics. For this test a minimum of 20 measurements are recommended. The mean of the measurement is calculated and chi square is then estimated from the mean as follows:
nxxMean
/)(
i
22
xxx
From 2 table, p value is obtained for a given set of measurements. A p value of 0.02 to
0.98 is treated as acceptable. In fact p value of 0.5 would be ideal to indicate that observed chi square value is in the middle of the range expected for Poisson distribution. A low p value shows a small probability that a Poisson distribution would give the chi square value
/)(
i
Counting Statistics
57
as large as actually observed and indicates the presence of additional sources of random error. With a p value just smaller than acceptable the experiment suggests repeat measurement, as the results are suspicious in such cases. Similarly a high p value between 0.98 and 0.99 is also considered to be suspicious and experiment needs to be repeated. For p >0.99 the random variations are much smaller than expected for a Poisson distribution.
Example: A radioactive sample placed in a well counter gives following measurement for a time duration of 10 seconds. Find out with chi square test whether the distribution of the data is acceptable.
S.No. Total counts
1. 3242
2. 3201
3. 3283
4. 3198
5. 3286
6. 3291
7. 3249
8. 3189
9. 3295
10. 3160
11. 3324
12. 3169
13. 3333
14. 3170
15. 3134
16. 3172
17. 3312
18. 3345
19. 3290
20. 3152
Mean 3239.75 and chi square value = 28531.75/3239.75 = 8.806
From chi square table the probability will be around 0.97. As mentioned earlier the value between 0.02 and 0.98 is considered acceptable. Thus measured values in the example can be considered to be precise.
58
Counting Statistics
Medical decision making and principle of ROC
The public awareness in recent years has increased to an extent that they want the justification of the cost and possible risks of a diagnostic procedure. It is important to measure the quality of diagnostic information and diagnostic decisions. Here we will deal with the quality and diagnostic decision of a test. To compare two test modalities one has to choose a parameter/index.
The simplest parameter/index could be “accuracy”. This is the fraction of cases for which test provides the correct result. We always require accuracy to be very high (near 100%). This parameter has to be used carefully as it may mislead in some situations. For example in screening relatively rare disease when the test can be very accurate, simply by ignoring all evidence and calling all cases negative. If only 5% of patients have the disease in question, a test, which always blindly states that the disease is absent, will be right (accurate) 95% of the time. Similar situation arises, when the disease is common (high prevalence).
It is obvious from above example that disease prevalence affects the accuracy quite significantly. Even if disease prevalence is known and fixed, this index has limited role to play. For example, there may be two different tests with equal accuracy but different false positive and false negative decisions. One might provide almost all false negative decisions, while other nearly all false positive decisions. Thus the usefulness of these two tests for patient management could be quite different in different situations.
To make the accuracy a meaningful index, prevalence of the disease, sensitivity and specificity should also be included in decision-making. Sensitivity and specificity are defined as follows:
Sensitivity = Number of True Positive (TP) decisions / Number of actually positive cases (TP+FN)
Specificity = Number of True Negative (TN) decisions / Number of actually negative cases (TN+FP)
The “sensitivity” represent a kind of accuracy for actually positive cases while “specificity another kind of accuracy for actually negative cases. It is necessary to specify the basis of determining true positive/true negative (gold standard) while calculating sensitivity and specificity.
Disease prevalence is defined as:
Prevalence = (TP+FN)/(TP+FN+TN+FP)
Disease prevalence is the characteristic of the patient population group. Figure 3 (hypothetical curve) shows a distribution of normal and abnormal (ill) persons in the population. Y-axis represents the number of persons who are normal (left part of figure 3) or abnormal (right part). Disease prevalence is the area under the distribution curve for diseased
Counting Statistics
250
Normal
59
persons compared to the area under both the curves in figure 3. If prevalence is high, the test with high sensitivity should be chosen while in case of low prevalence test with higher specificity should be preferred.
Accuracy = (True positive (TP) +True negative (TN))/(Total cases (TP+TN+FP+FN)
Abnormal
No. Of Persons
-100 -50 0 50 100 150 200
Test Result
Figure 3: Distribution of normal and abnormal subjects in the population
Accuracy of a test depends on disease prevalence for the same sensitivity and specificity. It is therefore quite possible that the accuracy for a given test at two centres may be different due to their different disease prevalence. Accuracy can also be expressed in terms of sensitivity, specificity and prevalence.
Accuracy = sensitivity x prevalence + specificity x (1-prevalence)
Positive predictive value (PPV) : This is a conditional probability of having disease
when the test shows true positive result.
PPV = TP/(TP+FP)
Similarly negative predictive value (NPV) = TN/(TN+FN)
Other parameters, which are mentioned at times in medical decision making are:
True positive fraction (TPF = p( T+ / D+ ) is simply the same as sensitivity, and true negative fraction (TNF= p(T- / D-), is simply the same as specificity.
60
Counting Statistics
False Positive fraction (FPF = p(T+ / D-) = No. of false positive decisions / No. of actually negative cases.
False Negative fraction (FNF= p(T- / D+) = No. of false negative decisions / No. of actually positive cases.
Example : Result of 10 volunteers (7 normal individuals and 3 ill patients) are given in the following table. The bold face letters indicate the normal individuals. Using a threshold value 20 (as positive) label them as TP, FP, TN, FN .
A B C D E F G H I J 10 5 20 40 60 9 70 50 80 15
Answer:
A B C D E F G H I J 10 5 20 40 60 9 70 50 80 15 FN TN TP FP FP TN TP FP FP TN
Example : In 100 patients who underwent thyroid uptake measurement, the test results are:
Negative Positive
60 (TN) 10 (FP) 5 (FN) 25 (TP)
Sensitivity = 25/(25+5) = 83.3% Specificity = 60/(60+10) = 85.7%
Accuracy = (60+25)/100 = 85%
Because of the variations in biological response, possible statistical noise in the data, and other technological deficiencies, a medical test some times produces identical results for abnormal (disease) and normal patients. Usually a normal range for normal people is established and if the results fall outside this normal range they are termed as positive (ill). Usually a threshold value is chosen arbitrarily, and different choices yield different frequencies for various kinds of correct and incorrect decisions. A threshold should therefore be selected in a judicious manner, which can make a good balance between sensitivity and specificity.
Receiver Operating Characteristic (ROC) Curve
This technique is used in medical decision-making. This does not depend on disease prevalence or observer’s choice of decision threshold. It requires a “gold” standard or known truth. It can be applied to computer simulated, phantom or clinical data. It can be better illustrated through example.
Counting Statistics
61
There are 200 studies with pathological confirmation that 100 are positive and 100 are negative. Observer evaluates each study and ranks each study on the 5 point scale as follows:
Definitely positive = 1
Probably positive = 2
Equivocal = 3
Probably negative = 4
Definitely negative = 5
After the evaluation, the following table is generated
Decision Threshold 1 2 3 4 5 Sum
Positive 40 26 20 10 4 100
Negative 6 12 22 24 36 100
There are 100 positive cases and 100 negative Cases.
Table 3
Decision Threshold 1 2 3 4 5 Sum
Positive 40 26 20 10 4
Negative 6 12 22 24 36
Table 3 needs some explanation. With decision threshold as (definitely positive=1), observer labeled 46 cases as definitely positive, selecting 40 as positive from 100 positive cases and 6 as positive from 100 negative cases. With decision threshold = 2 (probably positive = 2) user labeled 26 cases as probably positive, selecting 26 as positive from 100 positive cases and 12 as positive from 100 negative cases. It is to be noted, with decision threshold = 2, user has finally labeled 26+40 = 66 as True Positive (TP) and 6 + 12 = 18 as False Positive (FP), 34 as False Negative (FN) and 82 as True Negative (TN). Similar explanation is valid for decision threshold 3,4 and 5 respectively. Table 4 given below shows the sensitivity (TP /(TP +FN)) and specificity (TN / (TN + FP)) calculated for each threshold.
Table 4
Decision threshold 1 2 3 4 5
Sensitivity 0.40 0.60 0.86 0.90 1.00
Specificity 0.94 0.82 0.60 0.36 0.0