Submind YouTube summaries
Thumbnail for Building and Interpreting Regression Model | Applied Biostatistics | BIO733_Topic183

Building and Interpreting Regression Model | Applied Biostatistics | BIO733_Topic183

Watch on YouTube

Video summary

This module focuses on the process of constructing and interpreting regression models using SPSS software, specifically within the context of applied biostatistics. The core concept discussed is the simple linear regression model, defined by the equation Y equals beta naught plus beta one X plus epsilon, where Y represents the dependent variable, X is the independent variable, beta naught is the y-intercept, and beta one is the slope representing random error. A critical point emphasized in the explanation is that the intercept (beta naught) should only be interpreted if the independent variable can realistically take a value of zero; otherwise, it lacks practical meaning. In contrast, the slope (beta one) describes the average change in the dependent variable for every one-unit increase in the independent variable, with its sign indicating whether the relationship between the variables is positive or negative. The instructional video outlines a systematic approach to building a regression model, which begins with constructing a scatter plot to visually inspect the relationship between two variables. In the provided example involving renal function, GFR serves as the predictor variable and inverse cystatin C as the response variable. After generating the scatter plot, which reveals an increasing trend between the two measures, the analysis proceeds to the linear regression procedure in SPSS. The model summary indicates a strong fit with an R-squared value of 0.585, suggesting that over half of the variability in cystatin levels is explained by GFR. Subsequent tests confirm the overall significance of the model, showing a p-value less than 0.05, which establishes a statistically significant linear relationship between GFR and cystatin C in the population. Following the establishment of the overall model significance, the transcript details the interpretation of specific coefficients and the testing of individual parameter significance. The resulting regression equation is stated as Y hat equals 0.193 plus 0.006 X, meaning that for every one ml per minute increase in GFR, cystatin C levels are expected to increase by an average of 0.006 mg per liter. Although the intercept value of 0.193 is calculated, it is noted that since GFR cannot be zero in this clinical context, this specific value is not practically interpretable. However, hypothesis tests for both the intercept and the slope are performed using t-statistics derived from SPSS output. Both parameters are found to be statistically significant with p-values well below the 0.05 threshold, confirming that the slope represents a real impact rather than random chance. The conclusion of the module reinforces that the identified relationship between GFR and cystatin C is not limited to the specific sample data but holds true for the broader population. By rejecting the null hypotheses for both parameters, the analysis validates that there is a significant association between these two renal function markers. This rigorous statistical process ensures that the model can be reliably used to predict changes in cystatin C based on GFR measurements. Ultimately, the video demonstrates how to move from raw data visualization to a finalized, statistically validated regression model that provides meaningful insights into biological relationships, allowing researchers to generalize findings beyond their initial study sample.
Read the full video transcript
This module talks about building and then interpreting a regression model using SPSS. Simple linear regression model that is given by Y equals to beta naught plus beta 1 X plus E where Y is the dependent variable, X is an independent variable, beta naught is Y intercept, beta 1 is the slope, and epsilon are the random error. Here, beta naught is interpreted as an average value of Y when X is zero. One thing is very important that beta naught should only be interpreted if X can possibly take a value zero. But if X cannot be zero, we should not interpret beta naught. On the other hand, the other important parameter in this regression model is beta 1, which is also called the slope. And since we know that a simple linear regression model, which is a straight line, but it's an also an average line. Hence, the interpretations of beta naught and beta 1 will always be in the form of an average. So here, beta naught beta 1, which is a slope, is an average amount of change in Y with a one unit change in X. The sign of beta 1 indicates the direction of the relationship. A positive sign indicates that X and Y have positive or direct relationship, which states that that if X increases, Y increases. If X decreases, Y decreases. But having a negative sign indicates that X and Y have negative or indirect relationship, which means if X increases, Y will decrease. And if X decrease, Y increases. Here are some important steps in construction of a regression model. Whenever we are interested in building a regression model, the first step is to construct a scatter plot. Then, we build the simple linear regression model. We carry out the test for the overall significance of the model, and then we carry out the test for the individual significance of the parameters. We run diagnostic testing, and then finalizing the model. Here in this module, we'll use this example and learn how to test these models. How to build the model, and how to test the model using SPSS. So, GFR is the most important parameter of renal function assessed in renal transplant recipients. Although inulin clearance is regarded as the gold standard measure of GFR, its use is in clinical practice is limited. Pryser et al. examined the relationship between the inverse of cystatin C and inulin GFR as measured by DTPA GFR clearance. So, the results of 27 tests are shown in the following table. Here, we are again using DTPA GFR as the predictor of of inverse cystatin C as the response variable. The values are given to us, and we'll perform the calculations using SPSS, and then we'll interpret this model. Our data is in SPSS already. So, to build a model, the first step is to construct a scatter plot. Since our interest is only looking at the relationship between the variable X and Y, it doesn't matter which variable goes onto the X axis and which variable goes to the Y axis. And we simply press okay. The SPSS shows us the scatter plot in the output viewer, where we can clearly see that there is an increasing trend between GFR and cystatin C, which indicates that that if GFR values will increase, there will be an increase in the cystatin C. But through the scatter plot, we can't really clearly discuss that the amount of change is how much. And to know this amount of change and really know that if this amount of change is real or not, we have to run the full regression model and carry out the overall test of significance as well as individual tests of significance. And to do this, we'll go to analyze, regression, linear. Our dependent variable is cystatin and our independent variable is GFR. We'll bring the method to enter and press okay. It shows us that in our model, there is one dependent variable, that is cystatin, and there's one independent variable, GFR, hence it's a simple linear regression model. Our model summary indicates that the R value is 0.765 and R squared value, which is the coefficient of determination, is 0.585. Having this value more than 0.5 mean indicates that this is a better model. Now, we carry out the test for the overall significance, where we have uh kept uh beta is equals to zero against alternative that beta is not equals to zero. This uh P value is 0.000, which is less than alpha 0.05, hence we can conclude that that there is a linear relationship between X and Y. And lastly, we have the table for the coefficients, where we are given the values of the coefficients and other parameters that includes standard error, the T values, and the significance values. In ANOVA, we carried out a test for the overall significance of the model, but here, from the coefficients, we have these T values and these significance values given, which help us to calculate to carry out the individual test for significance. Using this information from this model, our regression model can be stated as Y hat equals 0.193 plus 0.006 X. This indicates that if we increase GFR by one ml per minute, There will be on average 0.006 mg per liter increase in cystatin. Since GFR values cannot be zero, hence there's no need to interpret the beta naught. However, if let us assume for the time being the GFR can possibly be zero, though it's not possible in this case, but we are assuming for the sake of understanding that if GFR is equals to zero, the average value of cystatin will be 0.193 mg per liter on average. Here, we have this regression model. We interpreted this model. Now, once we have the regression model, the next thing that we do is carry out the test for the individual significance. So, to test if beta naught is exactly equals to zero against the alternative that beta naught is not equals to zero at alpha 0.05, we use a test statistic. T, which is beta naught hat divided by the standard error of beta naught. And then we carry out the calculations. And from our SPSS output we can clearly see that the value of T is 3.978. Which is obtained by taking the ratio of beta naught and standard error of beta naught. And our decision rule states that we reject H naught if P value is less than alpha. And in this case the P value is 0.001. Hence we can conclude we can conclude that that we may reject the null hypothesis and state that that beta naught is a significant parameter. Similarly, we carry out the test of hypothesis for the significance of beta one. At 5% level of significance the test statistics will be a T statistics. That is beta one hat divided by the standard error of beta one. And after doing calculation the value of T is 5.931. And the P value equals 0.001. Using the decision rule as to reject H0 if the P value is is less than alpha our conclusion states that since the P value is equals to 0.001 is less than 0.05, hence we can reject the null hypothesis that beta one is equals to zero. So, rejecting this null hypothesis means saying that alternative hypothesis is true, which means that beta one is a significant impact. This concludes us by saying that that there is a significant relationship between the GFR and cystatin. And this is not only true for this sample, but for the population as well. And if we draw another sample from it we will find a similar relationship for this population.