Submind YouTube summaries
Thumbnail for CI for Difference Between Two Means for Dependent Samples | Applied Biostatistics | BIO733_Topic117

CI for Difference Between Two Means for Dependent Samples | Applied Biostatistics | BIO733_Topic117

Watch on YouTube

Video summary

This module focuses on constructing confidence interval estimates for the difference between two means, specifically addressing scenarios where the samples are dependent rather than independent. Dependent samples, often referred to as matched pairs or paired observations, occur when the subjects in one sample determine the subjects in the other, such as measuring the same individual before and after an intervention. Common examples include pre- and post-treatment measurements on a single subject, assigning littermates of the same sex to different treatments, or studying twins where each member receives a different therapy. In these cases, the analysis shifts from looking at two separate groups to examining the differences within each pair, treating these differences as a new variable of interest drawn from a normally distributed population. To calculate the confidence interval for dependent samples, the process mirrors that of estimating a single mean but utilizes the differences between paired observations. The formula involves the average difference ($\bar{d}$), the standard deviation of those differences ($s_d$), and the sample size ($n$) to determine the standard error. A critical assumption in this analysis is that the population of differences follows a normal distribution, which can be verified through goodness-of-fit tests or accepted based on prior statements in the study design. The reliability factor used in the calculation is derived from the t-distribution rather than the z-distribution, reflecting the smaller sample sizes typical in paired studies, with degrees of freedom calculated as $n - 1$. The video illustrates these concepts using a study by John Morton regarding gallbladder function before and after fundoplication surgery for gastroesophageal reflux disease. Researchers measured the gallbladder ejection fraction in 12 individuals, resulting in 24 total observations split into pre-operative and post-operative groups. By subtracting the pre-operative percentages from the post-operative ones, a set of 12 differences was created. After confirming that the samples were dependent and assuming normality based on the study's premise, the mean difference was calculated to be approximately 18.075 with a variance of 1068.093. Using a t-value of 2.201 for 11 degrees of freedom at a 95% confidence level, the resulting confidence interval ranged from -2.69 to 38.84. The final conclusion drawn from this analysis highlights that because the calculated confidence interval includes zero, there is insufficient evidence to claim a statistically significant difference between the gallbladder function before and after the treatment. This outcome suggests that, on average, the means could be equal, indicating that the surgery may not have had a measurable effect on the specific metric studied in this instance. The core takeaway emphasizes that for dependent samples, statistical inference must be performed on the sample of differences rather than treating the pre- and post-treatment data as independent groups, ensuring accurate estimation of the population parameter.
Read the full video transcript
This module talks about confidence interval estimates for difference between two means. But in this case, we'll talk about a case where our samples are dependent. First of all, we'll look at the difference between independent samples and dependent samples. Two samples are called independent when the subjects selected for the first sample do not determine the subjects in the second sample. Contrary to this, two samples will be dependent when the subjects in the first sample determine the subject in the second sample. The data from dependent samples are called matched pairs or paired samples. We also use the term related or paired observations for such type of samples. So paired observations may be obtained in a number of ways. The same subject is measured before and after receiving some intervention. Littermates of the same sex may be assigned randomly to receive treatment or placebo. Similarly, the pairs of twins or siblings may be assigned randomly to two treatments in such a manner that members of single pair receive different treatment. Let us take an example and learn about the paired samples. In this example, we will talk about a blood blood pressure patient where we'll make two observation. One before the treatment and other after the treatment. Here we'll calculate the mean and standard deviation before the treatment. And then we'll calculate mean and standard deviation blood pressure after the treatment. And when we compare these two when these two before and after observations are obtained from the same respondent from a same person the samples will be called as dependent or paired samples. We check for certain assumption. The first assumption that we check if the two samples are independent or dependent and second is that if the the distribution of the differences is normal or not. Here we are discussing a special scenario where two groups are dependent and the population of the sample differences is normal. Instead of performing the analysis with individual observations, we use DI. That's the difference between pair of observation as a variable of interest. And n the sample difference is computed from the n pairs of measurements that constitutes a simple random sample from a norm normally distributed population of differences. To calculate the confidence interval, we'll start with the general general form of confidence interval estimate that requires an estimator plus minus the reliability factor multiplied by the standard error of estimate. Here the expression of for constructing a confidence interval for the difference in paired means is almost identical to the formula for constructing a confidence interval for one mean. Note that only change in the subscript D which stands for difference. So here Xar D represents the average difference before and after and in the standard deviation you would see SD which is the standard deviation of the difference and together SD over under root N will give us the standard error of the sampling distribution of the difference between means. Here t alphab 2 new is a reliability factor that comes from the t distribution where new is the degrees of freedom and in this case since our respondent will be one person and obs two observations are taken from that single respondent. Hence our sample will our sample size will be n and hence the degrees of freedom will be n minus one. Let's take an example where John Morton examined gallbladder function before and after fundoplication a surgery used to stop stomach contents from flowing back into the esophagus in patients with gastro esophedial reflux disease. The authors measured gallbladder functionality by calculating the gallbladder ejaculation fraction before and after the treatment that is fund application in this case. So the goal of fund application is to increase GBF which is measured as a percentage. The table given below shows the percentages preop and posttop. Here we have this information for 12 individuals. The data consists of 12 individuals and from those 12 individuals 24 observations are made. 12 observations are made before fund application and 12 observations are made after this fund application. So here we shall perform the statistical analysis on the difference in preop and posttop GBF. We may obtain these differences in one of the two ways. The first way is to subtract the pre-op% from the posttoperson or the other way is to subtract the posttop percent from the preopersonent. So let us obtain the differences by subtracting the preopersonents from the posttoperson and this gives us d where d is the differences between the pair of observations. So hence we have 12 differences obtained. So here researcher measured gallbladder functionality by calculating the gallbladder ejaculation fraction before and after fund application. Now we have to check the assumption and the first assumption is to check for the dependence of the measurement and in this case if you look at how the observations are measured these are 12 individuals giving observations before and after. Hence we have two groups before fund application and after fund application. Hence the samples are dependent. The second assumption is of normality and to check this assumption we see if it's already known and stated in the statement and if not we practically you know perform a goodness of fit test. So here in this situation we uh can see in the statement that assuming the samples are obtained from the population that follows a normal probability distribution. Hence the statement is already given to us. So we'll use this information and assume that the samples comes from the population that follows the normal probability distribution. Since in the given situation both the observations are being made from the same respondent. Hence they are considered dependent observation and normality of the distribution between the difference holds true. It is established that this is a situation where we have to measure the confidence interval for the difference between two means where the samples are dependent. The 100 into 1 minus alpha% confidence interval for this data can be given as this where first of all we have to calculate the average difference that is d bar and for this sample it turned out to be 18.075 075 and the stand and the variance which is SD squared is obtained at 1068.0930. So using these values into the confidence interval estimate formula we'll get using the reliability factor at t 0.025 that is a two-tailed value at 11° of freedom the reliability factor is 2.2010. Including all these values into the formula we get two values. First is -2.69 and other is 38.84. Here it is important that we point out that how we use t distribution to calculate this reliability factor. We look at t0.975 and the degrees of freedom that is 11 and these two values intersect at 2.2010 20110 which is our reliability factor which is obtained from the table for the percentiles of the t distribution. We can say that if we were to repeat the study many many times and compute confidence interval in the same way about 95% of the intervals would include the difference between the population means. Since the interval includes zero, we conclude that the means before the treatment and after the treatment may be equal. A key concept in this is that we consider the difference of matched pair data as a sample and perform inferences on the sample of differences. Thank you.