Building and Interpreting Regression Model | Applied Biostatistics | BIO733_Topic183
Watch on YouTubeVideo summary
This module focuses on the process of constructing and interpreting regression models using SPSS software, specifically within the context of applied biostatistics. The core concept discussed is the simple linear regression model, defined by the equation Y equals beta naught plus beta one X plus epsilon, where Y represents the dependent variable, X is the independent variable, beta naught is the y-intercept, and beta one is the slope representing random error. A critical point emphasized in the explanation is that the intercept (beta naught) should only be interpreted if the independent variable can realistically take a value of zero; otherwise, it lacks practical meaning. In contrast, the slope (beta one) describes the average change in the dependent variable for every one-unit increase in the independent variable, with its sign indicating whether the relationship between the variables is positive or negative.
The instructional video outlines a systematic approach to building a regression model, which begins with constructing a scatter plot to visually inspect the relationship between two variables. In the provided example involving renal function, GFR serves as the predictor variable and inverse cystatin C as the response variable. After generating the scatter plot, which reveals an increasing trend between the two measures, the analysis proceeds to the linear regression procedure in SPSS. The model summary indicates a strong fit with an R-squared value of 0.585, suggesting that over half of the variability in cystatin levels is explained by GFR. Subsequent tests confirm the overall significance of the model, showing a p-value less than 0.05, which establishes a statistically significant linear relationship between GFR and cystatin C in the population.
Following the establishment of the overall model significance, the transcript details the interpretation of specific coefficients and the testing of individual parameter significance. The resulting regression equation is stated as Y hat equals 0.193 plus 0.006 X, meaning that for every one ml per minute increase in GFR, cystatin C levels are expected to increase by an average of 0.006 mg per liter. Although the intercept value of 0.193 is calculated, it is noted that since GFR cannot be zero in this clinical context, this specific value is not practically interpretable. However, hypothesis tests for both the intercept and the slope are performed using t-statistics derived from SPSS output. Both parameters are found to be statistically significant with p-values well below the 0.05 threshold, confirming that the slope represents a real impact rather than random chance.
The conclusion of the module reinforces that the identified relationship between GFR and cystatin C is not limited to the specific sample data but holds true for the broader population. By rejecting the null hypotheses for both parameters, the analysis validates that there is a significant association between these two renal function markers. This rigorous statistical process ensures that the model can be reliably used to predict changes in cystatin C based on GFR measurements. Ultimately, the video demonstrates how to move from raw data visualization to a finalized, statistically validated regression model that provides meaningful insights into biological relationships, allowing researchers to generalize findings beyond their initial study sample.
Read the full video transcript
This module
talks about
building and then interpreting a
regression model
using SPSS.
Simple linear regression model that is
given by Y equals to beta naught plus
beta 1 X plus E where Y is the dependent
variable, X is an independent variable,
beta naught is Y intercept, beta 1 is
the slope, and epsilon are the random
error.
Here, beta naught is interpreted as
an average value of Y
when X is zero.
One thing is very important that beta
naught should only be interpreted if X
can possibly take a value zero.
But if X cannot be zero,
we should not interpret beta naught.
On the other hand, the other important
parameter in this regression model is
beta 1, which is also called the slope.
And since we know
that a simple linear regression model,
which is a straight line,
but it's an also an average line.
Hence, the interpretations of beta
naught and beta 1
will always be in the form of an
average.
So here, beta naught beta 1, which is a
slope, is an average amount of change in
Y
with a one unit change in X.
The sign of beta 1 indicates the
direction of the relationship.
A positive sign indicates that X and Y
have positive or direct relationship,
which states that that if X increases, Y
increases. If X decreases, Y decreases.
But having a negative sign indicates
that X and Y have negative or indirect
relationship, which means if X
increases, Y will decrease. And if X
decrease, Y
increases.
Here are some important steps
in construction of
a regression model.
Whenever we are interested in building a
regression model, the first step is to
construct a scatter plot.
Then, we build the simple linear
regression model.
We carry out the test for the overall
significance of the model, and then we
carry out the test for the individual
significance of the parameters.
We run diagnostic testing,
and then finalizing the model.
Here in this module, we'll use this
example
and learn how to test these models.
How to build the model, and how to test
the model
using SPSS.
So, GFR is the most important parameter
of renal function assessed in renal
transplant recipients.
Although inulin clearance is regarded as
the gold standard measure
of GFR, its use is in clinical practice
is limited.
Pryser et al. examined the relationship
between
the inverse of cystatin C and inulin GFR
as measured
by DTPA GFR
clearance.
So, the results of 27 tests are shown in
the following table.
Here, we are again using DTPA GFR as the
predictor
of of inverse cystatin C as the response
variable.
The values are given to us, and we'll
perform the calculations using SPSS,
and then we'll interpret this model.
Our data is in SPSS already.
So,
to build a model,
the first step is to construct a scatter
plot.
Since our interest is only looking at
the relationship between
the variable
X and Y, it doesn't matter which
variable goes onto the X axis and which
variable goes to the Y axis.
And we simply press okay.
The SPSS
shows us the scatter plot in the output
viewer,
where we can clearly see
that there is
an increasing trend between
GFR and
cystatin C,
which indicates that that if GFR values
will increase, there will be an increase
in the cystatin C.
But through the scatter plot, we can't
really clearly discuss that the amount
of change
is how much.
And to know this amount of change and
really know that if this amount of
change is real
or not, we have to run the full
regression model
and carry out the overall test of
significance as well as individual tests
of significance. And to do this, we'll
go to analyze, regression, linear.
Our dependent variable is cystatin and
our independent variable is GFR.
We'll bring the method to enter and
press okay.
It shows us that in our model, there is
one dependent variable, that is
cystatin, and there's one independent
variable, GFR, hence it's a simple
linear regression model.
Our model summary indicates that the R
value is 0.765 and R squared value,
which is the coefficient of
determination, is 0.585.
Having this value more than 0.5 mean
indicates that this is a better model.
Now, we carry out the test for the
overall significance, where we have uh
kept uh beta is equals to zero against
alternative that beta is not equals to
zero.
This uh
P value is 0.000, which is less than
alpha 0.05, hence we can conclude that
that there is a linear relationship
between
X and Y.
And lastly, we have
the table for the coefficients,
where we are given
the values of the coefficients
and
other parameters that includes standard
error,
the T values,
and the significance values.
In ANOVA, we carried out a test for the
overall significance of the model,
but here, from the coefficients,
we have these T values and these
significance values given, which help us
to calculate
to carry out the individual test for
significance.
Using this information
from this model,
our regression model can be stated as
Y hat
equals 0.193
plus 0.006
X.
This indicates
that
if we increase
GFR
by
one
ml
per minute,
There will be on average
0.006
mg per liter
increase
in cystatin.
Since GFR values cannot
be zero,
hence
there's no need to interpret the beta
naught.
However,
if
let us assume for the time being the GFR
can possibly be zero,
though it's not
possible in this case, but we are
assuming for the sake of understanding
that if GFR is equals to zero, the
average value of
cystatin will be 0.193
mg per liter on average.
Here, we have this regression model.
We interpreted this model.
Now, once we have the regression model,
the next thing that we do is carry out
the test for the individual
significance.
So, to test if
beta naught is exactly equals to zero
against the alternative that beta naught
is not equals to zero
at alpha 0.05,
we use a test statistic.
T,
which is
beta naught hat
divided by the standard error of
beta naught.
And then we carry out the calculations.
And from our SPSS output
we can clearly see that the value of
T
is 3.978.
Which is obtained by
taking the ratio of beta naught and
standard error of beta naught.
And our decision rule
states that we reject H naught
if
P value
is less than alpha.
And in this case the P value
is
0.001.
Hence
we can conclude
we can conclude that
that we may reject the null hypothesis
and state that
that beta naught is a significant
parameter.
Similarly, we carry out the test of
hypothesis
for the significance of
beta one.
At 5% level of significance the test
statistics will be a T statistics.
That is beta one hat divided by the
standard error of
beta one.
And
after doing calculation the value of T
is
5.931.
And the P value
equals
0.001.
Using the decision rule
as to
reject
H0
if the P value
is is less than alpha
our conclusion states that
since the P value is equals to 0.001
is less than 0.05,
hence we can
reject the null hypothesis
that beta one is equals to zero. So,
rejecting this null hypothesis means
saying that alternative hypothesis is
true, which means that beta one
is
a significant impact.
This concludes us by saying that
that there is a significant relationship
between the GFR and cystatin.
And this is not only true for this
sample, but for the population as well.
And if we draw
another sample from it
we will
find a similar relationship
for this population.