Kruskal Wallis (I) | Applied Biostatistics | BIO733_Topic151
Watch on YouTubeVideo summary
When the assumptions required for a standard Analysis of Variance (ANOVA) are not satisfied, such as when populations are not normally distributed or variances are unequal, researchers must turn to nonparametric alternatives to compare the averages of more than two groups. The Kruskal-Wallis test serves as this primary nonparametric counterpart to one-way ANOVA. It is specifically designed for situations where data consists of ranks rather than raw measurements, allowing for a hypothesis test regarding equal location parameters like the median or mean across three or more independent samples. Essentially, this test evaluates whether samples originate from populations with identical distributions, rejecting the null hypothesis if at least one population tends to exhibit values larger or smaller than the others.
The procedure for conducting the Kruskal-Wallis test involves combining all samples into a single dataset and assigning ranks to every observation, ignoring their original group labels. After calculating the sum of ranks for each individual sample, the test statistic $H$ is computed based on these rank sums and the total number of observations. The value of $H$ acts as a measure of variance among the rank sums; if the ranks are distributed evenly across groups, $H$ will be small, whereas significant differences between groups result in some samples having excessively high ranks and others having low ranks, leading to a large $H$ value. This statistic is then compared against critical values derived from the chi-square distribution with $k-1$ degrees of freedom, where $k$ represents the number of samples, typically using a significance level of 0.05.
Decision-making in this test can be approached through either critical values or p-values, though specific limitations exist regarding sample size and the number of groups. Critical value tables are strictly applicable only when there are exactly three groups and each group contains five or fewer observations; for larger sample sizes or more than three groups, the approximation using the chi-square distribution is necessary. A notable drawback of the Kruskal-Wallis test is that it utilizes relatively little information from the data, relying solely on whether observations fall above or below a median rather than their actual magnitudes. Consequently, while effective for establishing differences in central tendency, other nonparametric methods like Conover's one-way analysis of variance by ranks are often preferred when it is important to account for the magnitude of each observation relative to others.
Read the full video transcript
any of the assumptions for analysis of
variance are not met then we cannot
rightfully perform analysis of variance
to make comparison between the averages
of more than two groups
in that situation
the resort is to perform nonparametric
tests and a nonparametric counterpart
for analysis of variance is kriscal
wallace test in this module we'll learn
about the kriskll Wallace test. So when
the populations from which the samples
are drawn are not normally distributed
with equal variances or when the data
for analysis consists only of ranks. A
nonparametic alternative to oneway noa
may be used to test the hypothesis for
equal location parameter and test
location parameter can be mean, it can
be median, it can be mod.
In such situation, Kriskll Wallace test
provides an alternative to ANOVA which
uses ranks of the data from three or
more independent samples to test the
null hypothesis that the samples come
from the population with equal median.
The Kriskll Wallace test also called the
H test is an nonparametric test that
uses ranks of sample data from three or
more independent populations. Though
it's a nonparametric test but it still
comes with some of the assumptions. The
samples are independent random samples
from their respective populations. The
measurement scales employed is at least
ordinal and the distributions of the
values in the sample populations are
identical except for the possibility
that one or more of the populations are
composed of the values that tend to be
larger than those of the other
populations. Chris Wallace test follows
the usual procedure for the testing of
hypothesis where firstly we state our
null hypothesis that says that the
population centers are all equal against
alternative hypothesis that at least one
of the population tends to exhibit
larger values than at least one of the
other populations. The we keep the usual
level of significance alpha to be 0.05
05 with the test statistics given by H
where K is the number of samples and J
is the number of observations in the G
sample. N is the number of observation
in all samples combined and RJ is the
sum of the ranks in the G sample. The
next step is to make the calculations.
We follow the following steps to
calculate the value of the test
statistic.
The first step is that we temporarily
combine all the samples into one big
sample and assign a rank to each sample
value. As a second step, for each
sample, find the sum of the ranks and
find the sample sizes. And then we
calculate H by using the results of step
two and the notation and test statistics
given on the preceding slides.
The test statistics H is basically a
measure of the variance of the rank sums
R1, R2, RK. So if the ranks are
distributed evenly among the sample
groups then h should be a relatively
small number but if the samples are very
different then the ranks will be ex
excessively low in some groups and high
in others. So with the net effect that h
will be large. We use the fall follow
the regular decision rule which may use
critical value approach and the p- value
approach. In the critical value
approach, we use the crit critical
values of age for various sample sizes
and alpha levels are given in table n in
the next slide. And the p value approach
where the null hypothesis will be
rejected is the computed value of the
age is so large that the probability of
obtaining a value that large or larger
when it is true is equal to or less than
the chosen significance level alpha that
is 0.05 in most of the cases. And
finally we write down our conclusion.
Here are the table n where we have the
sample sizes n1, n2 and 3 and the
critical values and alpha values. And we
calculate the critical value right from
here.
And we compare these critical values
with our h value that we compute to make
our decision.
The table end can only be used when we
have the three groups and our
observations in each group are less than
five. Five or less than five. But if the
sample sizes are more than five
or the number of groups are more than
three then we may not be able to use
table n then we use approximation
of the
h test.
So in that case when there are more than
five observations in one or more of the
samples H is compared with tabulated
value of K square with K minus one
degrees of freedom. This one of the
drawback of Kriskll Wallace test is a
deficiency of this test which is uh the
fact that it uses only small amounts of
information that's available. The test
uses only information as to whether or
not the observations are above or below
a single number. The medians of a
combined values. So the test does not
directly use measurements of known
quantity. Hence it is based only on that
where the value belongs as it gives the
ranks.
So several nonparametric analoges to
analysis of variance are available that
use more information by taking into
account the magnitude of each
observation relative to the magnitude of
every other observations.
Perhaps the best known of these
procedure is crllace one-way analysis of
variance by ranks
but certainly there are other
alternatives available.