Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5428_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
15.09.2026
Размер:
10 Мб
Скачать
☆
1.
Students ’ t Test
Students’ t-test is also called as t-test or T-test. It is a statistical hypothesis test. A statistical hypothesis is a proposition that is tested because of observed data modeled as the realized values taken by a collection of random variables. The t-test is based on t-distribution. It is considered an appropriate test for evaluating the significance of a sample mean or for evaluating the significance of difference between the means of two samples in case of small sample. When population variance is not known, the variance of the sample as an estimate of the population variance can be used.
A set of data is modeled as being realized values of a collection of random variables having a joint probability distribution in some set possible joint distributions. The hypothesis under test is exactly a set of probability distributions. A statistical hypothesis test is a method used for statistical inference. When two samples are related, the paired t-test can be used for evaluation of the significance of the mean of difference between the two related samples. This paired t-test is also known as difference test. It can also be used for evaluation of the significance of the coefficients of simple and partial correlations. The relevant test statistic, t, is calculated from the sample data. Then it is compared with its probable value based on t­distribution. This is to be read from the table that gives probable values of t for different levels of significance for different degrees of freedom at a specified level significance for concerning degrees of freedom for acceptance or rejection of null hypothesis. It should be noted that the t-test can be applied only in case of small sample when population variance is unknown.
The student’s t-distribution is simply the t-distribution. It is a member of the family of continuous probability distribution. It is used to calculate the mean of a normally distributed population when the sample size is small, the mean of two sets are significantly different from each other, and the standard deviation of the population is unknown. It was developed by English statistician William Searly Gosset
72
.
It is one of the most popular statistical techniques which is used to test (1) whether mean difference between two groups is statistically significant, (2) null hypothesis stating that both the means are statistically equal. The t-tests are of three types –1) One sample t-test, 2) Two sample t-test, and 3) Paired samples t-test.
One sample t-test: For one sample t-test, the value of t-statistics is simply the distance of a mean value that was obtained from the sample from the mean of the distribution, µ in units of the standard error, S.E.
https://t.me/med1917
2.
The calculated value of t is compared with the t distribution to set up statistical significance, while the tails of the t distribution can identify the area of rejection. That is, the null hypothesis of no difference can be rejected. They are defined by the tabulated t values.
A confidence interval (CI) on the true mean in the population is then estimated from the single sample’s mean and two-tailed test
73
.
CI = Sample mean ± t (S.E.)
For one tailed test, the CI on μ may be written as
CI = sample mean + t (S.E.) [for lower-tailed]
CI = sample mean – t (S.E.) [for upper-tailed]
Two sample t-test: In case oftwo-sample situation,the pooled t-test may be used in the statistical calculation. This two-sample t-test is also known as the basic t-test or unpaired t-test. The t-statistics value is expressed as:
Where, SP is the pooled standard deviation. It is a pooled estimate of the two standard deviations of the samples; because the pooled t- test assumes equal variances for the two populations where the two samples have been drawn. If the variances of the populations are not homogeneous and the pooled t-test is used, the probability of not rejecting a false null hypothesis would increase; however, this is not desirable. A substitute for the t-test can be used when the variances are different and is known as Behrens-Fisher’s test. This test is also known as t-test for two independent samples or separate t-test. The t-statistics for this test is;
The variances in the above eqn. are refers to the variation among the samples. The symbols n1 and n 2 refer to the number of observations in sample 1 and 2, respectively.
The confidence intervals (two-tailed) on the true difference between the means in the populations (μ 1 - μ 2 ) for the pooled t-test and the separate t-test are:
https://t.me/med1917
3.
For Pooled t-test
For separate t-test
The sign, ‘±’ should be replaced by ‘+’ or ‘–’ for one-tailed test lower or upper respectively.
Paired sample t-test: In case of the paired t-test with a single sample is obtained from a population. Individually the samples are measured twice on a variable of interest, once before treatment and once after receipt of it. In other words, each sample serves as its control, and thus this method can produce data more precisely. In some situations, identical twins are used where one of them is measured without treatment, and the other one is measured after receiving it. In the latter case, the ’treatment’ can be a tool such as psychological test the individual has to take. With paired t-test, the differences between the readings ‘before’ and ‘after’ are computed and then for the resulting differences (di) their mean and standard deviation is obtained. The t-statistics are:
Where n is the number of observations. The degrees of freedom for the test are (n – 1). A confidence interval (two-tailed) on the true mean of the differences in the population (D) is calculated from
CI = (Mean of differences) ± t (SE)
The sign ‘±’ should be replaced by ‘+’ or ‘–’ for one-tailed test, lower or upper respectively.
The null hypothesis is D = 0.
ANOVA Test
The term ‘variance’ was first used by Professor R.A. Fisher. In fact, he developed a very elaborate theory on ANOVA, The Analysis of Variance. Professor Fisher explained the utility of ANOVA in practical field. Thereafter, Professor Snedecor and many others contributed to the development to this technique. ANOVA is primarily a procedure to test the
https://t.me/med1917
1.
2.
difference among different groups of data for homogeneity. The spirit of ANOVA is that the total amount of variation in a set of data is broken down into two types,
The amount which can be attributed to chance, and
The amount which can be attributed to specified causes
74
.
Thus, when data are normally distributed, a two-sample t-test can only be used to assess the significance of the difference between the mean values of two independent groups. To compare the differences in the mean values of three or more independent groups simultaneously, an analysis of variance (ANOVA) can be used. ANOVA is a parametric test. ANOVA is suitable when the outcome measurement represents a continuous normally distributed variable and when the explanatory variable is categorical with three or more groups. An ANOVA model can also be used to compare the effects of several categorical explanatory variables at one time or for comparing differences in the mean values of one or more groups after adjusting for a continuous variable, that is, a covariate. This is called as an analysis of covariance (ANCOVA). A covariate is any variable that correlate with the outcome variable. For example, ANCOVA can be used to test for the effects of gender and socioeconomic status on weight after adjusting for height. Both ANOVA and ANCOVA are applications of the general linear model (GLM). Generally, GLM is used to construct a model to predict an outcome variable from one or more explanatory variables which may be categorical or continuous variables. The GLM may be univariate with only one explanatory variable or multivariate with several explanatory variables. Therefore, in GLM, the outcome is expressed as a function of the model and prediction error. In the univariate case, where there is only one outcome variable, the linear model consists of weights or coefficients, an intercept, and a prediction error. Variation may be found between samples and within the items of the sample. For analytical purposes, ANOVA consists in splitting the variance. Hence, it is a method of analyzing the variance to which a response is subject into its various components corresponding to various sources of variation. One can explain through this technique whether different varieties of seeds or fertilizers or soils differ significantly so that a policy decision could be taken accordingly, about a particular variety in agricultural research. In fact, ANOVA is extremely useful technique in the field of economics, education, biology, psychology, sociology, business, and industrial research in other disciplines. The difference in various types of drugs manufactured for curing a specific disease may be studied and evaluated to be significant or not by applying ANOVA technique. Similarly, the manager of a big company can analyze the performance of different salesmen of his company to know whether their performances differ significantly. By using ANOVA technique one can investigate various numbers of factors either hypothesized or considered to influence the dependent variables. Even one can investigate the differences among various categories within each of these factors having large number of possible values. If an analysis is conducted to find out the differences among its different categories having numerous possible values, the technique used for the purpose is said one-way ANOVA. If at the same time two factors are investigated, it is said two-way ANOVA. The interaction or interrelation
https://t.me/med1917
1.
2.
3.
4.
5.
6.
7.
8.
between two independent variables/factors that can influence a dependent variable can be studied for achieving better decision.
Thus, this technique can be used when multiple sample cases are involved.
Basic principles of ANOVA
The differences among the means of the population can be examined by the amount of variation within each of these samples compared with the amount of variation between the samples. This is the basic principle of ANOVA. The variation within a given population is assumed that the values of (X
ij
) differ from that of the mean of the population only due to
randomness. When the differences between the populations is examined, it is assumed that the difference between the mean of the jth population and the grand mean is characterizable to what is called a ‘specific factor’ or what is technically described as treatment effect. Therefore, when the ANOVA is used, it is assumed that each of the samples is drawn from a normal population and each of these populations has the same variation. It is also assumed that all factors other than the one or more being tested are effectively controlled. Thus, it is necessary to make two estimates of population variance– one is based on between the samples variance and the other is based on within the samples variance. These said two estimates of population variance are compared with F-test, where F is expressed as:
This value of F is compared to the F-limit for given degrees of freedom. If the calculated F value is equal to or more than the F-limit value, it may be said that there are significant differences between the sample means.
Assumptions for ANOVA
The following assumptions must be met when one-way or factorial ANOVA is being used:
The participants must be independent; that is, each participant appears only once in their group.
The group must be independent; that is, each participant must be in one group only.
The variable, which has been outcome, is normally distributed.
All cells have an adequate sample size.
The cell size ratio is no longer than 1.4
Between the groups the variances are similar; that is, the variance is homogeneous.
The residuals are normally distributed.
There are no influential outliers
https://t.me/med1917
The assumptions numbered 1and 2 are like the assumptions for two-sample t-test. Violation of these can invalidate the analysis. This means that each participant should appear on one data row of the spreadsheet only and will be included only once in the analysis.
When cases appear more than once in the spreadsheet, repeated measures should be taken on ANOVA, or a linear mixed model should be used.
Fig. 6.13 Concept of ANOVA model
When an ANOVA is conducted, the data are divided into cells according to the number of groups in the explanatory variable. The cells sizes less than 10 are called small and are always problematic due to lack of precision in calculating the mean value for the cell; however, the preferred cell size is 30.
Low cell counts can result in a loss of statistical power. The assumption of a low cell size ratio is also important. For example, if one cell contains 10 cases and another 60, then the ratio becomes 1:6. If this ratio is more than 1:4, it may create some problems.
However, it may be difficult to avoid small cell sizes in the studies where no experiment is conducted because it is not possible to predict the number of cases in each cell before collection of the data. Even in some experimental studies drop-outs and missing data can result in unequal cell sizes. If small cells are present, these should be re-coded or combined into larger cells; but the re-coding should be justified or meaningful. If there is no impact on the ultimate result, the small cells can be omitted also.
Within- and between-group variance
The interpretation of the output from an ANOVA model depends on the theory of mathematics used in conducting the test. The data are divided into their groups as shown in
Fig 6.13 in one-way ANOVA. The Fig shows that the data are divided into their groups and a
mean for each group is calculated. Each mean value is the predicted value for that group of
https://t.me/med1917
1.
2.
3.
4.
5.
participants. Moreover, a grand mean has also been calculated as shown in the Figure. Each mean value is the predicted value for that group of participants. The sample size in each group is equal.
One-way or single factor ANOVA
Under the one-way ANOVA, only one factor is considered. This factor is important because several possible types of samples occur within this factor. If there are differences within that factor, the ANOVA is determined, and the method includes the following steps:
The mean of each sample is obtained as.
Where, n is the number of samples
The mean of the sample means can be calculated as;
The deviations of the sample means are taken from the mean of the sample-means and the square of such deviations are calculated. These squares are multiplied by the number of items in the corresponding sample and summed up to get the total. This is known as the sum of squares for variance between the samples or SS B .
The result obtained in the step 3 is divided by the degrees of freedom between the samples to get the variance or mean square (MS B ) between samples as follows;
Where, (n – 1) is the degrees of freedom (d.f) between the samples.
The deviations of the values of the sample items for entire sample are obtained from corresponding means of the samples; the squares of such deviations are also calculated and summed up. This total is known as the sum of squares for variance within samples or SSW which is calculated as shown below:
Where,
https://t.me/med1917
6.
7.
i = 1, 2, 3,
The result of step no. 5 is to be divided by the degrees of freedom within sample to get the variance or mean square within (MS W ) the samples as shown below:
For verification, the sum of squares of deviations for total variance can also be determined by adding the squares of deviations when the deviations for the individual items in all the samples have been taken from the mean of the sample means. This is shown below;
Where,
i = 1, 2, 3,
j = 1, 2, 3,
This total should be equal to the total of the result of step no. (3) and (5) explained above.
That is, SS for total variance = SS B + SS
W
The degree of freedom for total variance would be equal to the number of items in all samples minus one (n – 1). The degree of freedom for between and within must be added up to the degrees of freedom for total variance.
That is, (n – 1) = (nꞌ – 1) + (n – nꞌ)
This explains the additive property of the ANOVA technique.
Finally, F-ratio can be calculated as:
This ratio is used to evaluate whether the difference among several sample means is significant or is just amatter of sampling fluctuations.
For this reason, a table giving the values of F for given degrees of freedom at different levels of significance have been provided.If the calculated value of F is less than the table of F, the difference is considered as insignificant; that is, due to chance and the null hypothesis of no difference between sample means holds good.
If the calculated value of F is either equal to or more than the table of F (provided in Appendix), the difference is considered as significant; that is, due to chance and the
https://t.me/med1917
1.
2.
3.
4.
5.
null hypothesis of no difference between sample-means does not hold good.
Short-cut method for one-way ANOVA
By using this method ANOVA can be performed. Because of its convenience particularly when the means of samples and/or mean of the sample means is non-integer values. The various steps involved in this method are:
Consider the total of the values of individual items in all the samples; that is, calculate ∑X
ij
Where, i = 1, 2, 3and j = 1, 2, 3
Let us assume this value as T
Calculate the value of correction factor
Calculate the square of all the item values one by one; then take its total. Subtract the correction factor from this total and the result is the sum of squares fortotal variance.
Where, i = 1, 2, 3, j = 1, 2, 3,
Make the square of each sample total (T j ) 2 and divide such square value of each sample by the number of items in the particular sample. Take the total of the result.
Subtract the correction factor from this total and the result is the sum of squares for variance between the samples.
Where, j represents different samples or categories
The sum of squares within the samples can be determined by subtracting the result of step (4) from the result of step (3) states that
https://t.me/med1917
Once all these have been done, the table of ANOVA should be set up in the same way as mentioned earlier.
Coding method
This method is more short-cut than earlier one. The method is based on an important property of F-ratio. Its value does not change even if all the n item values are either multiplied or divided by a common figure or if a common figure is either added or subtracted from each of the given n item values. By this method big figures are reduced in magnitude by division or subtraction and calculation is simplified without any disturbance on the F-ratio. This method should be used specially when given figures are big or otherwise not convenient. Once the given figures are converted with the help of some common figure, then all the steps of the short-cut method discussed above can be accepted for obtaining and interpreting the F-ratio.
Example 10: Set up an analysis of variance table for the following production data per acre for three varieties of wheat, each grown on 4 plots and state whether the variety differences are significant.
Solution: The problem can be solved by the direct method or by shortcut method, but in each case we shall get the same result.
Solution through direct method: First the mean of each sample is calculated as
https://t.me/med1917