Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5428_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
15.09.2026
Размер:
10 Мб
Скачать
☆
= –0.667m + 40
Therefore, the equation is y = –0.667m + 40
Based on this equation the percent of adults smoked in the year 1990 (i.e., after 25 years from 1965) would be easily calculated as shown below:
y = (–0.667) (25) + 40
= – 16.675 +40
= 23.325
Hence, about 23.325% of adults had been smoking in the year 1990.
Standard Deviation
There are different methods used to measure the research data and their analysis. For this purpose, statistics is used to describe the basic features of the data set. For example, the level of scattering is measured. There is a tendency of values of a variable to scatter away from the mean or midpoint. This tendency of scattering is called dispersion. The data are collected or measured with basic statistical tools such as mean, median and mode. For accuratemeasurement, the standard deviation is used. Standard deviation is a statistical tool developed to find out the difference between the calculated mean values. It is also one of the tools used to measure dispersion. Thus, the terms such as mean, median, mode, variation, should be explained before going to standard deviation.
Mean
The term, mean is defined in the year of 2008 by Panneerslvam as the ratio between the sum of the observations and number of the observations made in the study . Thereafter, in 2009
Eboh stated it as a sum of observations divided by the number of observations. Mathematically, the mean is the arithmetic average of several scores. The mean is obtained by adding the scores and dividing by the number of scores.
Example 1: The fasting blood sugar level of ten male patients have been measured and shown in the following table.
Mathematically, the mean, 𝑥 is the arithmetic average of the blood sugar values. This can be calculated by dividing the total of blood sugar values Σ x i by total number of patients (n =
https://t.me/med1917
•
•
Σn)
Where 𝑥is the arithmetic mean, x i is the ‘i’th observation, and n is the total number of observations.
The mean has certain properties which are:
The mean possesses the algebraic property; that is, the sum of the deviations of each observation from the mean will always be zero. This is expressed mathematically as
The sum of the squared deviations of each observation from the mean is less than the sum of the squared deviations about any other number
That is, when various values computed together to form the mean are subtracted, originally summed up, they will give a zero value.After subtraction from the mean, when squared together the result arrived at the minimum value.
Median
The median represents the exact middle value of a set of values. It actually refers to the midpoint in a series of numbers. To find the median, the values are arranged in the ascending order. If there is odd number of values, the middle value is considered as median. But, if there is even number of values, the average of the two middle values would be the median.
Example 2: Find the median of (i) 24, 51, 32, 19, 25, 46, 28 and (ii) 24, 51, 32, 19, 25, 46, 28, 35
24, 51, 32, 19, 25, 46, 28 when arranged in an ascending order, it becomes
19, 24, 25 , 28, 32 , 46 , 51; there are odd number of values
Hence, 28 is the middle value and it becomes the median.
Another set of values is 25, 51, 32, 19, 25, 46, 28, 35
The values after arranging in ascending order, these become 19, 24, 25, 28, 32, 35, 46, 51
https://t.me/med1917
Number of values in the set is 8, an even number.
Hence, two values at the middle are 28 and 32: the average of these =
The median of this series is 30
Mode
The mode of a set of values is the value that appears most often. Thus, a set of values may have more than one mode or no mode.
Example 3: Find the mode of 15, 18, 20, 18, 21, 32, 17, 14, 25, 27
The mode is 18; since it appears twice in the series 15, 18 , 20, 18 , 21, 32, 17, 14, 25, 27
Example 4: Find the mode of 15, 18, 20, 19, 21, 32, 17, 14, 25, 27
There is no mode in the series, since no number appears more than once.
Example 5: Find the mode of 15, 18, 20, 18, 21, 32, 17, 14, 20, 27
The modes of the given series 15, 18,20 , 18 , 21, 32, 17, 14, 20 , 27
The modes would be 18 and 20
Since there are three different measures of centers; the question may come about the best one. The mean is usually considered better one, provided the frequency distribution is not skewed. In case of skewed data mean, median and mode together can describe the skewness better.
Fig. 6.12 Schematic diagrams of mean, median and mode
Variance
The variance, S 2 is the average squared deviation from the mean. It is called as the square of the standard deviation. Both are interchangeable measures. Thus, the standard deviation is the square root of the variance. It can be defined as
https://t.me/med1917
Where,
x is each individual score making up the distribution,
m is the mean of the distribution, and
n is the number of scores.
This is explained in the following example.
Example 6:
Total number of observations, n = 6, and
Mean of the distribution,
So, the variance,
Standard Deviation
The term, variance although is frequently used in statistical calculations as a measure of spread, it has certain limitations such as it is expressed in units different from those of the summarized data. This means that the expressed unit is far smaller than the data. However, the variance can be converted into a measure of spread expressed in the same unit of measurement as the original scores: the standard deviation, S. The standard deviation indicates the fluctuation of the variables around their mean. Standard deviation is the square root of variance. It is the most popular and commonly used measure of spread. Standard deviation can be expressed as:
https://t.me/med1917
Where, meaning of each term has been mentioned earlier. Standard deviation can be calculated with the help of the data given in example 6. Where the value of variance, S 2 =
6.8; therefore, the value of standard deviation, S = √S 2 = √6.8 = 2.61
Example 7: Find out the value of standard deviation of the following distributed values.
Standard deviation can be determined by transforming the given data and followed by other calculations.
Grouped data
Sometimes the data is presented in terms of grouped observation; hence, the standard deviation can be calculated by the way different from that illustrated in example 7.
Example 8: Total of 44 persons of different age groups was collected and number of male persons was noted. Find out the standard deviation.
https://t.me/med1917
Age group ( x ) No of male Persons (f)
11 – 20 5
21 – 30 11
31 – 40 15
41 – 50 10
51 – 60 3
The above table is rearranged below; the midpoint of each group is calculated and multiplied with the frequency and tabulated as shown below:
The mean becomes the sum of the scores in the ’f ×Midpoint’ column divided by the sample size. The sample size is calculated by adding the ’f’ column (n = 44).
Therefore, the mean = 1512.0/44 = 34.36 (rounded to 34.4).To determine the standard deviation there is a need to add one additional column to the table of calculations above. This additional column is f ×(midpoint) 2 and the table is redrawn below with the added column
https://t.me/med1917
1.
2.
3.
CHI -Square Test
It is an important test among various tests of significance. It is symbolically written as χ 2 . It is a statistical measure generally used for sampling analysis to compare the practical variance with the theoretical one. It is a non-parametric test, used to determine whether categorical data shows dependency or to know whether two classifications are independent. It may be used to compare theoretical populations with the actual data when categories are used
72
. Hence, the chi-square test can be used for various purposes. The chi-square test is
used for the following reasons:
To test the goodness of fit.
To test the significance of association between two attributes, and
To test the homogeneity or the significance of population variance.
Chi-square test is used to analyze categorical data such as male or female patients, smokers, and non-smokers, etc. It does not mean to analyze parametric or continuous data such as height measured in cm. or weight measured in kg, etc.
Use of chi-square test for comparing variance
Sometimes the chi-square value is used to evaluate the significance of population variance; that is, this test can be used to evaluate if a sample has been drawn randomly from a normal
population with mean (μ) and a specified variance, The test is based on χ 2 -distribution. Such a distribution is come across when a collection of
values is dealt with involving adding up squares. Variance of samples requires adding a collection of squared quantities and, thus, having distributions that are related to χ 2 ­distribution. The χ 2 –distribution can be obtained, if each one of the collections of sample variance is divided by a known population variance and multiplied these quotients by (n –
1), where n is the number of items in the sample.
Thus, (degree of freedom) would be the same distribution as to χ 2 – distribution with degree of freedom, (n – 1).
https://t.me/med1917
In short, when chi-square is used as a test of population variance, the valueof χ 2 is to be
worked out to test the null hypothesis (viz., H
o
as under;
Where,
σ
s
2
is the variance of the sample,
σ
2
p
is the variance of the population,
(n – 1) is the degree of freedom, and
n is the number of items in the sample.
By comparing the calculated value with the table value of χ 2 for (n – 1) degrees of freedom at a given level of significance, the null hypothesis may either be accepted or rejected.If the calculated value of χ 2 is less than the table value, the null hypothesis is accepted and if the calculated value of χ 2 is less than the table value, the null hypothesis is accepted and if the calculated value of χ 2 is more than the table value, the null hypothesis is rejected. All these are illustrated in the following example.
Example 9: The individual weights of ten students are 38, 40, 45, 53, 47, 43, 55, 48, 52, and 49kg. Can we say that the variance of the distribution of weight of all students from which the above sample of 10 students was drawn is equal to 20 kg? Test this at 5% and 1% level of significance.
Solution: Let us work out the variance of the sample data or σ
s
2
as under:
https://t.me/med1917
(i)
(ii)
The null hypothesis is Ho:σ
2
p
= σ
s
2
Now to test this hypothesis, the value of χ 2 may be calculated as: χ 2 =
At 5% level of significance the table value of χ 2 is 16.92 and at 1% level of significance the table value of χ 2 is 21.67 for degree of freedom of 9.
Both these values are more than the calculated value, 13.999 of χ 2 ; hence, the null hypothesis is accepted with a conclusion that the variance of the given distribution can be taken as 20 kg at 5% as also at 1% level of significance.
Use of chi-square test as a non-parametric test
Chi-square is an important non-parametric test, and no strict assumptions are required as such regarding the type of population. Only the degree of freedom (size of the sample) would be required for the test. As a non-parametric test, chi-square can be used as
a test of goodness of fit, and
a test of independence.
As a test of goodness of fit, χ 2 test enables to see how well the assumed theoretical distribution (Binomial distribution, Poisson distribution or Normal distribution) fits to the observed data. When some theoretical distribution is fitted to the given data, we are always interested in knowing as to how well this distribution fit with the observed data. The chi­square test can answer to this. If the calculated value of χ2 is less than the table value at a particular level of significance, the fit is good one indicating that the divergence between the observed and expected frequencies is attributable to fluctuations of sampling. But if the calculated value of χ 2 is more than its table value, the fit is not considered to be good one.
As a test of independence, χ 2 test can explain whether two attributes are associated. For example, in knowing a new medicine is effective in controlling fever or not, χ 2 test will help in deciding this issue. In such a situation, the null hypothesis that two attributes, new medicine, and control of fever, are independent. This means that new medicine is not effective in controlling fever. Accordingly, the expected frequencies are to be calculated
https://t.me/med1917
first; then the value of χ 2 is to be worked out. If the calculated value of χ 2 is less than the table value at a certain level of significance for given degrees of freedom, it can be concluded that null hypothesis stands. This means that two attributes are independent or not associated. That is, the new medicine would not be effective in controlling fever. But if the calculated value of χ 2 is more than its table value, the null hypothesis does not hold good; that is, the two attributes are associated and the association is not due to some chance factor, it exists. That is, the new medicine would be effective in controlling fever and the medicine can be used. However, it may be said that χ 2 is not a measure of degree of relationship or the form of relationship between two attributes, but is simply a method of judging the significance of such association or relationship between two attributes.
The chi-square test either as a test of goodness of fit or as a test may be applied to evaluate the significance of association between attributes. However, the observed and theoretical or expected frequencies must be grouped in the same way and the theoretical distribution must be adjusted to give the same total frequency as is found in case of observed distribution. The value of χ 2 is calculated as:
Where,
O
ij
is the observed frequency of the cell in ith row and jth column
E
ij
is the expected frequency of the cell in ith row and jth column.
If the two distributions are exactly similar, χ 2 = 0. Generally due to sampling error, χ 2 is not equal to zero. As such it must be known the sampling distribution of χ 2 ; so that, the probability of an observed χ 2 being given by random sample may be found out from the hypothetical universe.
Important Characteristics of χ 2 test
This test, as a non-parametric test, is based on frequencies; not on the parameters such as mean and standard deviation.
The test is used for testing the hypothesis and is not useful for estimation.
The test has the additive property.
This test can also be applied to a complex contingency table with several classes and as such is a very useful in research activities.
The test is an important non-parametric test since no rigid assumptions would be necessary regarding the type of population; there is no need of parameter values and relatively less mathematical details are involved.
https://t.me/med1917