Submind YouTube summaries
Thumbnail for CI Estimate for the Difference Between Two Population Proportions | Biostatistics | BIO733_Topic119

CI Estimate for the Difference Between Two Population Proportions | Biostatistics | BIO733_Topic119

Watch on YouTube

Video summary

This module focuses on constructing confidence intervals to estimate the difference between two population proportions, a statistical method often used when researchers need to quantify the magnitude of differences between groups. Such comparisons are common in studies examining dichotomous characteristics across various demographics, such as men versus women, different age groups, socioeconomic classes, or distinct diagnostic categories. The foundation of this estimation is an unbiased point estimator derived from the difference between two sample proportions, denoted as $\hat{p}_1 - \hat{p}_2$. To form a confidence interval around this point estimate, statisticians apply a standard formula that combines the estimated difference with a reliability factor and the standard error of the estimate. The validity of this approach relies on specific conditions regarding sample size and distribution. When the sample sizes for both populations are sufficiently large and the population proportions are not extremely close to zero or one, the Central Limit Theorem allows researchers to assume that the sampling distribution is approximately normal. Under these assumptions, the reliability factor, often represented as $Z$, is determined using the standard normal distribution corresponding to the desired confidence level. Consequently, the formula for a 100(1-$\alpha$)% confidence interval becomes $\hat{p}_1 - \hat{p}_2 \pm Z_{1-\alpha/2} \times SE$, where the standard error is calculated as the square root of the sum of the variances from each sample, specifically $\sqrt{\frac{\hat{p}_1(1-\hat{p}_1)}{n_1} + \frac{\hat{p}_2(1-\hat{p}_2)}{n_2}}$. To illustrate this process, the video presents a study by Conner et al. investigating gender differences in sexual abuse reports among 323 children and adolescents, comprising 68 females and 255 males. In this sample, 31 females and 53 males reported experiencing sexual abuse, resulting in sample proportions of approximately 0.4559 for females and 0.2078 for males. The difference between these proportions is calculated as 0.2481, with an estimated standard error of 0.0655. For a 99% confidence interval, the reliability factor $Z$ is identified as 2.58 from the standard normal distribution, leading to a calculated interval ranging from 0.0791 to 0.4171. The interpretation of these results indicates that researchers can be 99% confident that the true proportion of reported sexual abuse cases among females exceeds that of males by an amount between 0.0791 and 0.4171 in the broader population. A critical aspect of this analysis is observing whether the calculated confidence interval includes zero; in this specific example, since the entire interval lies above zero and does not contain the value zero, it provides strong statistical evidence that the proportions are indeed different. This conclusion confirms a significant disparity between the two groups regarding the reported incidence of sexual abuse, demonstrating how confidence intervals serve as a powerful tool for making definitive statements about population differences based on sample data.
Read the full video transcript
This module talks about estimating the confidence interval estimate for the difference between two population proportions. Sometimes researchers are interested to measure the magnitude of the difference between two population proportions. They may want to compare some dichotomous characteristics on the basis of their proportion among men versus women or maybe knowing the differences between two age groups or two socioeconomic groups or looking at dynamics for two diagnostic groups. So, an unbiased point estimator for the differences between two population proportion is provided by the difference between two sample proportions, which are given by P1 hat minus P2 hat. Like other confidence interval estimates, here we'll use the same expression we will we will use the estimate for the difference between two proportions. Plus minus the reliability factor and the standard error of estimate for the difference between two population proportion. When N1, that is the sample size for the first population, and N2 is a sample size for the second population are large and the population proportions are not too close to zero or one, we make use of the central limit theorem and assume that the normal distribution theory may be employed to obtain the confidence interval. Hence, 100 into 1 minus alpha percent confidence interval for difference in proportions can be obtained as P1 hat minus P2 hat plus minus Z 1 minus alpha by 2, which is a reliability factor, and since here we are assuming the normal distribution, reliability factor will be obtained from the standard normal distribution. And the standard error of the difference between two population proportions which is given by square root of the P1 hat 1 - P1 hat / N1 + P2 hat into 1 - P2 hat over N2. Let's take an example. Conner et al. investigated gender differences in proactive and reactive aggression in sample of 323 children and adolescents, 68 females and 255 males. The subjects were from unsolicited consecutive referrals to a residential treatment center and a pediatric psychopharmacology clinic serving a tertiary hospital and medical school. In the sample, 31 of the females and 53 of the males reported sexual abuse. We wish to construct a 99% confidence interval for the difference between the proportions of sexual abuse in the two sample groups. So, the sample proportions for the females and males are respectively given by P hat F for females, which is 31 / 68, that is total, = 0.4559, and PM hat represent proportion for males, which is equals to 0.2078. And the difference between sample proportions is P F hat - PM hat, which is given by 0.2481. The estimated standard error of the difference between sample proportion is obtained by the formula for the standard error, and it's given here as 0.06 55. The reliability factor can be obtained from the standard normal distribution, and here in this case, our interest is to obtain 99% confidence interval, hence, we look at the confidence level of 99 and the area under the curve where it is 90.9950, the Z alpha beta value is 2.58. Using this reliability factor, difference in proportion, and the standard error of estimate for the difference in proportion, the confidence interval that we get is 0.0791 as a lower confidence limit and 0.4171 as upper confidence limit. Interpreting it as that we are 99% confident that for the sample population, the proportion of the cases of reported sexual abuse among female exceeds the proportion of cases of reported sexual abuse among males by somewhere between 0.0791 and 0.4171. Since the interval does not include zero value, hence we can conclude that the proportions are different.