Submind YouTube summaries
Thumbnail for SPSS Tutorial - 18 | Applied Biostatistics | BIO733_Topic098

SPSS Tutorial - 18 | Applied Biostatistics | BIO733_Topic098

Watch on YouTube

Video summary

In this module, the focus shifts to understanding and constructing scatter plots within SPSS to analyze relationships between two quantitative variables. The tutorial utilizes a diet study dataset featuring 16 respondents, which includes longitudinal measurements such as triglyceride levels and body weight recorded at five different time points: baseline (time zero) and subsequent intervals up to time four. By examining the baseline data specifically, the video demonstrates how to identify these continuous variables and use a scatter plot to visualize their correlation. The process involves navigating to the Chart Builder, selecting the scatter plot option, and assigning triglyceride values to the x-axis and weight to the y-axis, although the tutorial notes that swapping these axes does not alter the fundamental nature of the relationship displayed. Once the plot is generated, each individual data point represents a unique coordinate pair where one value corresponds to the x-axis and the other to the y-axis. The video illustrates how to interpret this visual representation by hovering over or checking labels to identify specific individuals, such as the 11th respondent whose baseline triglyceride level of 94 and weight of 179 align precisely with a plotted dot. While the dataset in this example is relatively small, resulting in a somewhat sparse distribution of points, the overall trend suggests a positive direct relationship where higher weights are associated with higher triglyceride levels. This visual evidence supports the conclusion that as one variable increases, the other tends to increase as well, indicating a strong correlation between an individual's weight and their triglyceride count at the start of the study. To gain a more comprehensive view of multiple relationships simultaneously, the tutorial introduces the matrix scatter plot, which allows for the comparison of several variables at once. By selecting all relevant variables—such as baseline triglycerides, first interim weight, and subsequent measurements—the software generates a four-by-four grid where each block displays the scatter plot relationship between two specific variables. The diagonal blocks remain empty because a variable cannot be compared to itself, while the upper and lower triangles of the matrix are mirror images, meaning the interpretation remains consistent regardless of which half is viewed. This extended visualization reveals that while weight shows a strong linear relationship with its interim measurements, indicating stable body mass over time, the correlation between triglyceride levels at baseline and their first interim reading is less pronounced than the weight correlations. Ultimately, this method provides a robust way to assess how different physiological metrics evolve together across multiple time points in longitudinal studies.
Read the full video transcript
In this module, we will try to learn the scatter plots and how to draw them. Firstly, if we have two quantitative variables and our interest is to know that how these two quantitative variables are related to each other, we can use scatter plots to observe the relationship, direction of the relationship, amount of the relationship between the these two quantitative variables. Let's see it on SPSS. For this purpose, we will look at another data which is uh the diet study where 16 respondents have some information given to us where we have the information about their ID, their age in years, their gender They're triglyceride triglycerides at the base level. Then first interum triglycerides. Then second interum triglycerides, third interum triglyceride values and then final triglyceride values. This is type of a data that's called longitudinal data where we have for this data we have the values of triglycerides for each individual at five different instances at time zero time 1 2 3 and time four. Similarly for these patients we have their weights at the baseline. So TG0 is the triglyceride at the baseline or at the time the patient entered into the study and WGT0 is the weight of the patient the time he entered into the study. So these TG0 and TG T WGT0 are the baseline values for triglycerides and weights respectively. And similarly we have weights at time one, weights at time two, weights at time three and weights at time four. Let's firstly look at that how the relationship was between the triglycerides at the baseline and the weights at the baseline. One thing we want to know is that are these two variables quantitative from scale we can guess that yes they are quantitative variables. Moreover if you look at these values these values also are not categoricals. These are the true values of of the triglycerides as well as weights. And we know that weight and triglyide count trig the stride values are quantitative variables where these two variables are continuous variables as well. If you want to we want to see the relationship between track the strides and the weights at the baseline the graphical way to observe it is to construct a scatter plot. To draw a scatter plot, we'll go to the chart builder scatter. Since we want to see the scatter plot between two variables only, we'll get the triglycerides on onto the x-axis and weight onto the y-axis. Since both the variables are quantitative and in scatter plot it really doesn't matter what variable goes onto the x-axis and what goes onto the y-axis. The relationship it will show in the scatter plot will be very evident of the type of relationship that exist here. There's a scattered values given and that's why this plot is called a scatter plot. In this scatter plot, one can simply observe various dots where this value represents a some value at x-axis at the same point the same value at the y-axis. Let's try to identify who is this individual and what his values are. We simply go here, check the label. So this is the value of 11th individual. So if we want to figure out what the exact values or what the exact values are for this individual, we simply will go to the data and look at the values of the 11th individual. It says for triglycerides it's 94 and for weight it is 179. And if you look at our data that's what it shows. The value of triglyceride is 94 and the value of weight is 179 which is approximately which is exactly right here in the scatter plot. Each single point being placed on the scatter plot represents two variables at a time. One on x-axis and other on yaxis. So each value will have two coordinates in a two-dimensional scatter plot. And same way we can identify each and every value here. But point is not that how to identify the values. The point is we want to see the relationship between triglycerides and weights. And here though it's not really clear that if it is increasing or it is decreasing but if we look at it whole we can see that there is some increasing trend that as the weight is increasing the triglycerides are going up as well. Or one can see other ways that if triglycerides are higher values, weights are also higher values. So one can simply assume that there is a positive direct relationship between the values of triglycerides in the individual and their weights which shows that there is a positive relationship and whenever the weight will be more triglycerides will be more. But here the data is very sporadic. We have very less values in our data. That's why there is much bigger scatter here with our data. But it could be more clear if we have more observations within the data. Let's look at more variables at a time. Right now we just looked at triglycer. Right now we are looking at just two variables weight and triglycerides. But if you want to look at multiple variables at a time and we want to observe that how their their relationship goes, we can also draw another type of extended scatter plot which is called matrix plot. To do that, we simply go to graphs chart builder and we pick up the multiple scatter plot matrix. In scatter plot matrix, we bring all the variables that we want to draw the graph with. So one can bring all these variable one by one or one can select them and bring them all together. Once the variables are in you simply press okay. Since we talk we brought in four variables, it will give us a matrix of order four by four where each block will represent the relationship a scatter for two variables like here. This scatter plot shows us shows the relationship of triglycerides and the first interum weight. This graph shows triglycerides and weight. This triglycerides and first interum triglycerides. First interum triglyides with first interum weight. First interum triglyide with weight. So wherever the intersection goes, that's where it's going to show us the values. Moreover, the upper diagonal is the mirror image of the lower diagonal. So, one can simply interpret either these graphs or these graphs. The interpretation will stay the same. If you look at the weight variable weight with first interum weight, we are able to see pretty much linear relationship. And this shows that as weight was higher for the individuals, their first interim weight was also higher. And there's a pretty strong relationship over here that the heavier people the p people with more weight after some time at first in tanum weight reading their weight was still higher. So there was no drastic change in the weight of individuals. You will see in matrix scatter matrix you will see the diagonal there will be no scatter plot given because a diagonal if you see the weight so it is weight with weight. So weight with weight will no will will not show any comparison but weight will show all the comparison with all other variables like triglycerides weight with first interum triglyceride and weight weight with first interum weight. Here it shows that first interum weight is kind of positively related with triglycerides. similar weight with triglycerides and first interum triglycerides and simple triglycerides at the base level they are also showing the positive relationship. So individuals with more value of triglycerides at the baseline at first interum tri interum point the triglyceride values are also higher but right now the relationship between the triglyceride at the baseline and triglyceride at the first interum time are not showing as stronger relationship as it it's been shown in the weight at the baseline and weight at the first interum reading. And this is how we draw the mat scatter matrix and we interpret them. Thank you.