SPSS Tutorial - 18 | Applied Biostatistics | BIO733_Topic098
Watch on YouTubeVideo summary
In this module, the focus shifts to understanding and constructing scatter plots within SPSS to analyze relationships between two quantitative variables. The tutorial utilizes a diet study dataset featuring 16 respondents, which includes longitudinal measurements such as triglyceride levels and body weight recorded at five different time points: baseline (time zero) and subsequent intervals up to time four. By examining the baseline data specifically, the video demonstrates how to identify these continuous variables and use a scatter plot to visualize their correlation. The process involves navigating to the Chart Builder, selecting the scatter plot option, and assigning triglyceride values to the x-axis and weight to the y-axis, although the tutorial notes that swapping these axes does not alter the fundamental nature of the relationship displayed.
Once the plot is generated, each individual data point represents a unique coordinate pair where one value corresponds to the x-axis and the other to the y-axis. The video illustrates how to interpret this visual representation by hovering over or checking labels to identify specific individuals, such as the 11th respondent whose baseline triglyceride level of 94 and weight of 179 align precisely with a plotted dot. While the dataset in this example is relatively small, resulting in a somewhat sparse distribution of points, the overall trend suggests a positive direct relationship where higher weights are associated with higher triglyceride levels. This visual evidence supports the conclusion that as one variable increases, the other tends to increase as well, indicating a strong correlation between an individual's weight and their triglyceride count at the start of the study.
To gain a more comprehensive view of multiple relationships simultaneously, the tutorial introduces the matrix scatter plot, which allows for the comparison of several variables at once. By selecting all relevant variables—such as baseline triglycerides, first interim weight, and subsequent measurements—the software generates a four-by-four grid where each block displays the scatter plot relationship between two specific variables. The diagonal blocks remain empty because a variable cannot be compared to itself, while the upper and lower triangles of the matrix are mirror images, meaning the interpretation remains consistent regardless of which half is viewed. This extended visualization reveals that while weight shows a strong linear relationship with its interim measurements, indicating stable body mass over time, the correlation between triglyceride levels at baseline and their first interim reading is less pronounced than the weight correlations. Ultimately, this method provides a robust way to assess how different physiological metrics evolve together across multiple time points in longitudinal studies.
Read the full video transcript
In this module,
we will try to learn the scatter plots
and how to draw them.
Firstly,
if we have two quantitative variables
and our interest is to know
that how these two quantitative
variables are related to each other,
we can use scatter plots to observe
the relationship,
direction of the relationship,
amount of the relationship between the
these two quantitative variables.
Let's see it on SPSS. For this purpose,
we will look at another data
which is uh the diet study
where 16 respondents
have some information given to us
where we have the information about
their ID,
their age in years,
their gender
They're triglyceride triglycerides
at the base level. Then first interum
triglycerides.
Then second interum triglycerides, third
interum triglyceride values and then
final triglyceride values.
This is type of a data that's called
longitudinal data where we have for this
data we have the values of triglycerides
for each individual at five different
instances at time zero time 1 2 3 and
time four.
Similarly for these patients we have
their weights
at the baseline. So TG0
is the triglyceride at the baseline or
at the time the patient entered into the
study and WGT0
is the weight of the patient the time he
entered into the study. So these TG0 and
TG T WGT0
are the baseline values for
triglycerides and weights respectively.
And similarly we have weights at time
one, weights at time two, weights at
time three and weights at time four.
Let's firstly look at that how the
relationship was between the
triglycerides
at the baseline and the weights at the
baseline.
One thing we want to know is that are
these two variables quantitative
from scale we can guess that yes they
are quantitative variables. Moreover if
you look at these values these values
also are not categoricals. These are the
true values of of the triglycerides as
well as weights. And we know that weight
and triglyide count trig the stride
values are quantitative variables
where these two variables are continuous
variables as well.
If you want to we want to see the
relationship between track the strides
and the weights at the baseline
the graphical way to observe it is to
construct a scatter plot.
To draw a scatter plot,
we'll go to the chart builder
scatter. Since we want to see the
scatter plot between two variables only,
we'll get the triglycerides on onto the
x-axis and weight onto the y-axis.
Since both the variables are
quantitative
and in scatter plot it really doesn't
matter what variable goes onto the
x-axis and what goes onto the y-axis.
The relationship it will show in the
scatter plot will be very evident of the
type of relationship that exist
here.
There's a scattered values given and
that's why this plot is called a scatter
plot.
In this scatter plot, one can simply
observe
various dots
where this value
represents
a some value at x-axis at the same point
the same value at the y-axis. Let's try
to identify who is this individual and
what his values are. We simply go here,
check the label. So this is the value of
11th individual.
So if we want to figure out what the
exact values or what the exact values
are for this individual, we simply will
go to the data and look at the values of
the 11th individual. It says for
triglycerides it's 94 and for weight it
is 179.
And if you look at our data that's what
it shows.
The value of triglyceride is 94 and the
value of weight is 179 which is
approximately which is exactly right
here
in the scatter plot. Each
single point
being placed on the scatter plot
represents two variables at a time. One
on x-axis and other on yaxis. So each
value will have two coordinates in a
two-dimensional scatter plot. And same
way we can identify each and every value
here. But point is not that how to
identify the values. The point is we
want to see the relationship between
triglycerides and weights. And here
though it's not really clear that if it
is increasing or it is decreasing
but if we look at it whole we can see
that there is some increasing trend that
as the weight is increasing the
triglycerides are going up as well.
Or one can see other ways that if
triglycerides are higher values, weights
are also higher values. So one can
simply assume that there is a positive
direct relationship between the values
of triglycerides in the individual and
their weights
which shows that there is a positive
relationship and whenever the weight
will be more triglycerides will be more.
But here the data is very sporadic. We
have very less values in our data.
That's why there is much bigger scatter
here with our data. But it could be more
clear if we have more observations
within the data. Let's look at more
variables at a time. Right now we just
looked at triglycer.
Right now we are looking at just two
variables weight and triglycerides.
But if you want to look at multiple
variables at a time and we want to
observe that how their their
relationship goes, we can also draw
another type of extended scatter plot
which is called matrix plot. To do that,
we simply go to graphs chart builder
and we pick up the multiple scatter plot
matrix. In scatter plot matrix, we bring
all the variables that we want to
draw the graph with.
So one can bring all these variable one
by one
or one can select them and bring them
all together.
Once the variables are in you simply
press okay.
Since we talk we brought in four
variables,
it will give us
a matrix of order four by four where
each
block will represent the relationship a
scatter for two variables like here.
This scatter plot shows us shows the
relationship of triglycerides
and the first interum weight.
This graph shows triglycerides and
weight. This triglycerides and first
interum triglycerides.
First interum triglyides with first
interum weight. First interum triglyide
with weight. So wherever the
intersection goes, that's where it's
going to show us the values.
Moreover,
the upper diagonal is the mirror image
of the lower diagonal. So, one can
simply interpret either these graphs or
these graphs. The interpretation will
stay the same.
If you look at the weight variable
weight with first interum weight, we are
able to see pretty much linear
relationship.
And this shows that as weight was higher
for the individuals, their first interim
weight was also higher. And there's a
pretty strong relationship over here
that the heavier people the p people
with more weight
after some time at first in tanum weight
reading their weight was still higher.
So there was no drastic change in the
weight of individuals.
You will see in matrix scatter matrix
you will see the diagonal there will be
no scatter plot given because a diagonal
if you see the weight so it is weight
with weight. So weight with weight will
no will will not show any comparison but
weight will show all the comparison with
all other variables like triglycerides
weight with first interum triglyceride
and weight weight with first interum
weight.
Here it shows that first interum weight
is kind of positively related with
triglycerides.
similar weight with triglycerides and
first interum triglycerides and simple
triglycerides at the base level they are
also showing the positive relationship.
So individuals with more value of
triglycerides at the baseline at first
interum tri interum point the
triglyceride values are also higher but
right now the relationship between the
triglyceride at the baseline and
triglyceride at the first interum
time are not showing as stronger
relationship as it it's been shown in
the weight at the baseline and weight at
the first interum reading.
And this is how we draw the mat scatter
matrix and we interpret them. Thank you.