Submind YouTube summaries
Thumbnail for SPSS Tutorial - 20 | Applied Biostatistics | BIO733_Topic100

SPSS Tutorial - 20 | Applied Biostatistics | BIO733_Topic100

Watch on YouTube

Video summary

One of the fundamental assumptions in many traditional statistical analyses is that continuous variables adhere to a normal probability distribution. This module focuses on how to evaluate whether a quantitative variable, specifically using ulcer recurrence data where "age" is the variable of interest, meets this criterion. To assess normality, researchers can utilize various descriptive measures such as histograms, frequency curves, box plots, and specific coefficients like skewness and kurtosis. These tools provide initial insights into the shape of the distribution, helping to determine if the data aligns with the expected patterns of a normal curve before moving on to more rigorous inferential tests. The analysis begins by examining descriptive statistics, particularly the mean and median, which are crucial for assessing symmetry. In a perfectly symmetric normal distribution, the mean and median should be equal; however, in this specific dataset, the mean is 50.98 while the median is 52. Since the median is greater than the mean, this discrepancy suggests that the data is negatively skewed rather than symmetric. A normal distribution is inherently bell-shaped and symmetric, so any significant difference between these two central tendency measures indicates a deviation from normality, suggesting that the variable "age" in this study does not follow a normal probability distribution. Further confirmation comes from analyzing the coefficients of skewness and kurtosis, which offer precise mathematical descriptions of the distribution's shape. For a dataset to be considered normally distributed, the coefficient of skewness must be exactly zero, indicating symmetry, and the coefficient of kurtosis must also be zero, indicating a mesokurtic shape. In this case, the calculated coefficient of skewness is negative, confirming negative skewness, while the coefficient of kurtosis is less than zero, classifying the distribution as platykurtic. Because these values deviate from the required zero values for a normal distribution, the data fails to meet the essential properties of symmetry and mesokurticity. Consequently, based on the evidence provided by both the comparison of mean and median and the specific values of skewness and kurtosis, it is concluded that the variable "age" does not follow a normal probability distribution. The combination of negative skewness and platykurtic characteristics clearly demonstrates that the data is asymmetric and has a different tail weight than a standard normal curve. This finding is significant because it implies that statistical methods relying on the assumption of normality may need to be adjusted or alternative non-parametric tests should be considered for analyzing this specific dataset, highlighting the importance of verifying distributional assumptions before proceeding with further analysis.
Read the full video transcript
One of the very important assumption of many traditional statistical analysis is that our continuous variable that's involved in our analysis follows the normal probability distribution. In this module, we will learn how to perform various statistical tasks that will help us to interpret and test this assumption that if our quantitative variable in the study follows the normal probability distribution or not. To do this, we will continue looking at the ulcer recurrence data. And over here, we have the variable age, which is a quantitative variable. And if we want to observe that if age follows the normal probability distribution, we can use descriptive measures as well as we can use inferential measures. In this There are various descriptive measures one can use. It could be just the histogram or the frequency curve. One can also use coefficient of skewness and kurtosis to determine if the distribution is normal as well as one can use box plot to observe if the distribution is normal or not. All these various measures we're going to use are the descriptive measures to determine if the data or our variable age really follows the normal probability distribution. But one should know that these all descriptive measures are going to give us very good hints that if the distribution is normal or not. But if we really want to get a much more stronger evidence to make this conclusion, we always need to rely on the inferential results. To get various descriptive measures that will help us to observe the distribution of variable age, if it is normal or not, we'll simply go to analyze, descriptive statistics, and explore. Since we want to look at the age, we bring the variable age. And in statistics, we will get the descriptives. And in the plots, along with the histogram, we check on normality plots with tests. By checking on this option, it will not only give us descriptive measures and normal Q-Q plot, but it will also give us some inferential tests. I won't talk about the inferential test in this module, but definitely we're going to learn them as we start learning about the inferential statistical methodologies in in statistics. You press continue and press okay. Once we press okay, we have mean and we have median. We can always compare these two values to get a good hint if the data is symmetric or if our variable of interest could be normal or not. Right here, we can see the mean is 50.98 and median is 52. And as we know, a condition for a data to be symmetric that mean equals to the median. In this situation, median is actually greater than the mean. And when such sort of relationship exist between mean and median, where mean is a smaller value than median, the data is likely to be negatively skewed. When we know that the normal distribution is a bell curve distribution. So, the normal curve will be a bell-shaped curve, which is going to be symmetric. Hence, if from here we can say that median is a larger value than mean, it indicates that it's a negatively skewed data, and having a negatively skewed data indicates that our variable is not going to be normal. Similarly, the other measure we can use is the coefficient of skewness and the coefficient of kurtosis. If the coefficient of skewness is exactly equals to zero, we will say the data is symmetric. If the coefficient of skewness is negative, the data will be negatively skewed, and if the coefficient of skewness is positive, the data will be positively skewed data. And the same goes with the kurtosis, that if kurtosis coefficient of kurtosis is is exactly equals to zero, the data will be mesokurtic. And if coefficient of kurtosis is greater than zero, it will be leptokurtic. And if it will be negative, that means less than zero, it will be platykurtic. One of the very important property of the normal probability distribution is that every normal distribution will be symmetric and mesokurtic. But here, since the coefficient of skewness is negative, which indicates the distribution is negatively skewed, as well as coefficient of kurtosis is less than zero, which makes it a a platykurtic distribution. Since the our values coefficient of skewness and coefficient of kurtosis says the data is negatively skewed and platykurtic, which is not symmetric or and mesokurtic, which are the essential assumption for the normal distribution. Hence, we can say that our distribution is likely not going to be the normal probability distribution for the variable age. Thank you.