Submind YouTube summaries
Thumbnail for Kruskal Wallis (I) | Applied Biostatistics | BIO733_Topic151

Kruskal Wallis (I) | Applied Biostatistics | BIO733_Topic151

Watch on YouTube

Video summary

When the assumptions required for a standard Analysis of Variance (ANOVA) are not satisfied, such as when populations are not normally distributed or variances are unequal, researchers must turn to nonparametric alternatives to compare the averages of more than two groups. The Kruskal-Wallis test serves as this primary nonparametric counterpart to one-way ANOVA. It is specifically designed for situations where data consists of ranks rather than raw measurements, allowing for a hypothesis test regarding equal location parameters like the median or mean across three or more independent samples. Essentially, this test evaluates whether samples originate from populations with identical distributions, rejecting the null hypothesis if at least one population tends to exhibit values larger or smaller than the others. The procedure for conducting the Kruskal-Wallis test involves combining all samples into a single dataset and assigning ranks to every observation, ignoring their original group labels. After calculating the sum of ranks for each individual sample, the test statistic $H$ is computed based on these rank sums and the total number of observations. The value of $H$ acts as a measure of variance among the rank sums; if the ranks are distributed evenly across groups, $H$ will be small, whereas significant differences between groups result in some samples having excessively high ranks and others having low ranks, leading to a large $H$ value. This statistic is then compared against critical values derived from the chi-square distribution with $k-1$ degrees of freedom, where $k$ represents the number of samples, typically using a significance level of 0.05. Decision-making in this test can be approached through either critical values or p-values, though specific limitations exist regarding sample size and the number of groups. Critical value tables are strictly applicable only when there are exactly three groups and each group contains five or fewer observations; for larger sample sizes or more than three groups, the approximation using the chi-square distribution is necessary. A notable drawback of the Kruskal-Wallis test is that it utilizes relatively little information from the data, relying solely on whether observations fall above or below a median rather than their actual magnitudes. Consequently, while effective for establishing differences in central tendency, other nonparametric methods like Conover's one-way analysis of variance by ranks are often preferred when it is important to account for the magnitude of each observation relative to others.
Read the full video transcript
any of the assumptions for analysis of variance are not met then we cannot rightfully perform analysis of variance to make comparison between the averages of more than two groups in that situation the resort is to perform nonparametric tests and a nonparametric counterpart for analysis of variance is kriscal wallace test in this module we'll learn about the kriskll Wallace test. So when the populations from which the samples are drawn are not normally distributed with equal variances or when the data for analysis consists only of ranks. A nonparametic alternative to oneway noa may be used to test the hypothesis for equal location parameter and test location parameter can be mean, it can be median, it can be mod. In such situation, Kriskll Wallace test provides an alternative to ANOVA which uses ranks of the data from three or more independent samples to test the null hypothesis that the samples come from the population with equal median. The Kriskll Wallace test also called the H test is an nonparametric test that uses ranks of sample data from three or more independent populations. Though it's a nonparametric test but it still comes with some of the assumptions. The samples are independent random samples from their respective populations. The measurement scales employed is at least ordinal and the distributions of the values in the sample populations are identical except for the possibility that one or more of the populations are composed of the values that tend to be larger than those of the other populations. Chris Wallace test follows the usual procedure for the testing of hypothesis where firstly we state our null hypothesis that says that the population centers are all equal against alternative hypothesis that at least one of the population tends to exhibit larger values than at least one of the other populations. The we keep the usual level of significance alpha to be 0.05 05 with the test statistics given by H where K is the number of samples and J is the number of observations in the G sample. N is the number of observation in all samples combined and RJ is the sum of the ranks in the G sample. The next step is to make the calculations. We follow the following steps to calculate the value of the test statistic. The first step is that we temporarily combine all the samples into one big sample and assign a rank to each sample value. As a second step, for each sample, find the sum of the ranks and find the sample sizes. And then we calculate H by using the results of step two and the notation and test statistics given on the preceding slides. The test statistics H is basically a measure of the variance of the rank sums R1, R2, RK. So if the ranks are distributed evenly among the sample groups then h should be a relatively small number but if the samples are very different then the ranks will be ex excessively low in some groups and high in others. So with the net effect that h will be large. We use the fall follow the regular decision rule which may use critical value approach and the p- value approach. In the critical value approach, we use the crit critical values of age for various sample sizes and alpha levels are given in table n in the next slide. And the p value approach where the null hypothesis will be rejected is the computed value of the age is so large that the probability of obtaining a value that large or larger when it is true is equal to or less than the chosen significance level alpha that is 0.05 in most of the cases. And finally we write down our conclusion. Here are the table n where we have the sample sizes n1, n2 and 3 and the critical values and alpha values. And we calculate the critical value right from here. And we compare these critical values with our h value that we compute to make our decision. The table end can only be used when we have the three groups and our observations in each group are less than five. Five or less than five. But if the sample sizes are more than five or the number of groups are more than three then we may not be able to use table n then we use approximation of the h test. So in that case when there are more than five observations in one or more of the samples H is compared with tabulated value of K square with K minus one degrees of freedom. This one of the drawback of Kriskll Wallace test is a deficiency of this test which is uh the fact that it uses only small amounts of information that's available. The test uses only information as to whether or not the observations are above or below a single number. The medians of a combined values. So the test does not directly use measurements of known quantity. Hence it is based only on that where the value belongs as it gives the ranks. So several nonparametric analoges to analysis of variance are available that use more information by taking into account the magnitude of each observation relative to the magnitude of every other observations. Perhaps the best known of these procedure is crllace one-way analysis of variance by ranks but certainly there are other alternatives available.