Kruskal Wallis (II) | Applied Biostatistics | BIO733_Topic152
Watch on YouTubeVideo summary
In this module, the focus is on performing the Kruskal-Wallis test using SPSS software, specifically when the assumptions required for a standard Analysis of Variance are not met. The lesson utilizes an example involving outpatient charges in US dollars for a specific surgical procedure collected from hospitals across three different areas of the country. These areas are coded numerically as one, two, and three to represent distinct groups. The primary objective is to determine whether there is a statistically significant difference in these charges among the three areas at a 5% level of significance.
To conduct the analysis, data must be entered into SPSS with two specific variables: a grouping variable representing the area codes and a dependent quantitative variable representing the actual charges. Within the software interface, users navigate to the Analyze menu, select Nonparametric Tests, choose Legacy Dialogs, and then click on K Independent Samples. In this dialog box, the test variable is set to the outpatient charges while the grouping variable is assigned to the area codes, specifying a range from group 1 to group 3. Upon executing the command, SPSS generates results that allow for the evaluation of the null hypothesis, which posits that the population medians are equal across all groups, against the alternative hypothesis that at least one group differs significantly from the others.
The output reveals an asymptotic significance value derived from a chi-square test, which in this example is 0.003. Since this p-value is substantially lower than the predetermined significance level of 0.05, the null hypothesis is rejected. This statistical finding indicates that it is highly unlikely that the observed differences occurred by chance alone. Consequently, the analysis concludes that there is a significant difference in the outpatient charges among the three areas, meaning that at least one area tends to exhibit larger charge values compared to at least one of the other populations.
Read the full video transcript
In this module, we'll learn how to
perform Kruskal-Wallis test
on SPSS.
To do this, let's take an example.
The following are the outpatient charges
in
uh
US dollars
made to patients
for a certain surgical procedure by
samples of
hospitals located in three different
areas of the country.
Where one represents area one, two
represents area two, and three
represents area three.
And the values for each area are given
in dollars.
Can we conclude at 5% level of
significance that the three areas differ
with respect to the charges?
So, our null hypothesis is that
population centers are all equal against
alternative that at least one of the
population tends to exhibit larger
values than at least one of the other
populations
with level of significance 0.05.
And assuming that not all the
assumptions for performing analysis of
variance
are met, we will perform
Kruskal-Wallis test
using the test statistics given as H.
So, the data
for to perform Kruskal-Wallis test
is uh entered into SPSS where we will
have two variables.
The first variable is the grouping
variable. Here, it's represented by
area. So, area of a country that would
be coded as one, two, three. One will
represent area one, two will represent
area two, and three will represent area
three.
And other is data variable, that is a
dependent variable, which is a
quantitative variable in this case,
which is the charges.
It's outpatient charges made to patients
for a certain surgical procedure. Once
these things are set up, we'll enter the
data with
grouping variable area contains the code
for the group, and the
the data variable charges contains
actual data for each group.
So, to perform Kruskal-Wallis test, we
will go to analyze, non-parametric,
legacy dialog,
and
then uh we perform K independent
samples. In K independent samples, our
test variable is outpatient charges and
our grouping variable is
area, where group range from group 1 to
3. So, we'll give first group code and
the last group's code here.
Continue and press okay.
This is going to give us the results,
where our null hypothesis states that
the means are equal, medians are equal
against alternative that medians are not
equal. Our asymptotic significance
value, that is obtained using chi-square
test, it is uh 0.003.
Here, this P value is less than 0.05,
hence we may reject the null hypothesis.
And once we reject the null hypothesis,
we conclude that at least one of the
population tends to exhibit larger
values than at least one of the other
populations.