SPSS Tutorial - 20 | Applied Biostatistics | BIO733_Topic100
Watch on YouTubeVideo summary
One of the fundamental assumptions in many traditional statistical analyses is that continuous variables adhere to a normal probability distribution. This module focuses on how to evaluate whether a quantitative variable, specifically using ulcer recurrence data where "age" is the variable of interest, meets this criterion. To assess normality, researchers can utilize various descriptive measures such as histograms, frequency curves, box plots, and specific coefficients like skewness and kurtosis. These tools provide initial insights into the shape of the distribution, helping to determine if the data aligns with the expected patterns of a normal curve before moving on to more rigorous inferential tests.
The analysis begins by examining descriptive statistics, particularly the mean and median, which are crucial for assessing symmetry. In a perfectly symmetric normal distribution, the mean and median should be equal; however, in this specific dataset, the mean is 50.98 while the median is 52. Since the median is greater than the mean, this discrepancy suggests that the data is negatively skewed rather than symmetric. A normal distribution is inherently bell-shaped and symmetric, so any significant difference between these two central tendency measures indicates a deviation from normality, suggesting that the variable "age" in this study does not follow a normal probability distribution.
Further confirmation comes from analyzing the coefficients of skewness and kurtosis, which offer precise mathematical descriptions of the distribution's shape. For a dataset to be considered normally distributed, the coefficient of skewness must be exactly zero, indicating symmetry, and the coefficient of kurtosis must also be zero, indicating a mesokurtic shape. In this case, the calculated coefficient of skewness is negative, confirming negative skewness, while the coefficient of kurtosis is less than zero, classifying the distribution as platykurtic. Because these values deviate from the required zero values for a normal distribution, the data fails to meet the essential properties of symmetry and mesokurticity.
Consequently, based on the evidence provided by both the comparison of mean and median and the specific values of skewness and kurtosis, it is concluded that the variable "age" does not follow a normal probability distribution. The combination of negative skewness and platykurtic characteristics clearly demonstrates that the data is asymmetric and has a different tail weight than a standard normal curve. This finding is significant because it implies that statistical methods relying on the assumption of normality may need to be adjusted or alternative non-parametric tests should be considered for analyzing this specific dataset, highlighting the importance of verifying distributional assumptions before proceeding with further analysis.
Read the full video transcript
One of the very important assumption
of many traditional statistical analysis
is that our continuous variable that's
involved in our analysis
follows the normal probability
distribution.
In this module, we will learn
how to perform various statistical tasks
that will help us to interpret
and test this assumption that if our
quantitative variable in the study
follows the normal probability
distribution or not.
To do this, we will continue looking at
the ulcer recurrence data. And over
here, we have the variable age, which is
a quantitative variable.
And if we want to observe
that if age follows the normal
probability distribution,
we can use descriptive measures as well
as we can use inferential measures.
In this There are various descriptive
measures one can use.
It could be
just the histogram
or the frequency curve.
One can also use coefficient of skewness
and kurtosis to determine if the
distribution is normal
as well as one can use box plot
to observe if the distribution is normal
or not.
All these
various measures
we're going to use are the descriptive
measures to determine if the data
or our variable age really follows the
normal probability distribution.
But one should know
that these all descriptive measures are
going to give us very good hints
that if the distribution is normal or
not.
But if we really want to get a much more
stronger evidence to make this
conclusion,
we always need to rely on the
inferential results.
To get various descriptive measures that
will help us to observe the distribution
of variable age,
if it is normal or not, we'll simply go
to analyze,
descriptive statistics, and explore.
Since we want to look at
the age,
we bring the variable age.
And in statistics,
we will get the descriptives.
And in the plots,
along with the histogram, we check on
normality plots with tests.
By checking on this option,
it will not only give us descriptive
measures
and
normal Q-Q plot, but it will also give
us some inferential tests.
I won't talk about the inferential test
in this module, but definitely we're
going to learn them
as we start learning about the
inferential statistical methodologies in
in statistics.
You press continue
and press okay.
Once we press okay,
we have mean
and we have median.
We can always compare these two values
to get a good hint if the data is
symmetric or if
our variable of interest
could be
normal or not.
Right here,
we can see the mean is 50.98
and median is 52.
And as we know,
a condition for a data to be symmetric
that mean equals to the median.
In this situation, median is actually
greater than the mean.
And when such sort of relationship exist
between mean and median, where mean is a
smaller value than median, the data is
likely to be negatively skewed.
When we know
that the normal distribution is a bell
curve distribution.
So,
the normal curve will be a bell-shaped
curve,
which is going to be symmetric.
Hence, if from here we can say that
median is a larger value than mean,
it indicates that it's a negatively
skewed data, and having a negatively
skewed data indicates
that our variable is not going to be
normal.
Similarly,
the other measure we can use
is
the coefficient of skewness and the
coefficient of kurtosis. If the
coefficient of skewness is exactly
equals to zero, we will say the data is
symmetric.
If the coefficient of skewness is
negative, the data will be negatively
skewed, and if the coefficient of
skewness is positive, the data will be
positively skewed data.
And the same goes with the kurtosis,
that if kurtosis coefficient of kurtosis
is is exactly equals to zero, the data
will be mesokurtic. And if coefficient
of kurtosis is greater than zero, it
will be
leptokurtic. And if it will be negative,
that means less than zero, it will be
platykurtic.
One of the very important property of
the normal probability distribution is
that every normal distribution will be
symmetric
and mesokurtic.
But here, since the coefficient of
skewness is negative, which indicates
the distribution is negatively skewed,
as well as coefficient of kurtosis is
less than zero, which makes it a a
platykurtic distribution.
Since
the
our values coefficient of skewness and
coefficient of kurtosis says the data is
negatively skewed and platykurtic, which
is not symmetric or
and mesokurtic, which are the essential
assumption for the normal distribution.
Hence, we can say that
our distribution is likely not going to
be the normal probability distribution
for the variable age.
Thank you.