CI Estimate for the Difference Between Two Population Proportions | Biostatistics | BIO733_Topic119
Watch on YouTubeVideo summary
This module focuses on constructing confidence intervals to estimate the difference between two population proportions, a statistical method often used when researchers need to quantify the magnitude of differences between groups. Such comparisons are common in studies examining dichotomous characteristics across various demographics, such as men versus women, different age groups, socioeconomic classes, or distinct diagnostic categories. The foundation of this estimation is an unbiased point estimator derived from the difference between two sample proportions, denoted as $\hat{p}_1 - \hat{p}_2$. To form a confidence interval around this point estimate, statisticians apply a standard formula that combines the estimated difference with a reliability factor and the standard error of the estimate.
The validity of this approach relies on specific conditions regarding sample size and distribution. When the sample sizes for both populations are sufficiently large and the population proportions are not extremely close to zero or one, the Central Limit Theorem allows researchers to assume that the sampling distribution is approximately normal. Under these assumptions, the reliability factor, often represented as $Z$, is determined using the standard normal distribution corresponding to the desired confidence level. Consequently, the formula for a 100(1-$\alpha$)% confidence interval becomes $\hat{p}_1 - \hat{p}_2 \pm Z_{1-\alpha/2} \times SE$, where the standard error is calculated as the square root of the sum of the variances from each sample, specifically $\sqrt{\frac{\hat{p}_1(1-\hat{p}_1)}{n_1} + \frac{\hat{p}_2(1-\hat{p}_2)}{n_2}}$.
To illustrate this process, the video presents a study by Conner et al. investigating gender differences in sexual abuse reports among 323 children and adolescents, comprising 68 females and 255 males. In this sample, 31 females and 53 males reported experiencing sexual abuse, resulting in sample proportions of approximately 0.4559 for females and 0.2078 for males. The difference between these proportions is calculated as 0.2481, with an estimated standard error of 0.0655. For a 99% confidence interval, the reliability factor $Z$ is identified as 2.58 from the standard normal distribution, leading to a calculated interval ranging from 0.0791 to 0.4171.
The interpretation of these results indicates that researchers can be 99% confident that the true proportion of reported sexual abuse cases among females exceeds that of males by an amount between 0.0791 and 0.4171 in the broader population. A critical aspect of this analysis is observing whether the calculated confidence interval includes zero; in this specific example, since the entire interval lies above zero and does not contain the value zero, it provides strong statistical evidence that the proportions are indeed different. This conclusion confirms a significant disparity between the two groups regarding the reported incidence of sexual abuse, demonstrating how confidence intervals serve as a powerful tool for making definitive statements about population differences based on sample data.
Read the full video transcript
This module talks about
estimating the confidence interval
estimate for the difference between two
population proportions.
Sometimes
researchers are interested
to measure the magnitude of the
difference between two population
proportions.
They may want to compare
some dichotomous characteristics
on the basis of their proportion among
men versus women or maybe knowing the
differences between two age groups
or two socioeconomic groups
or looking at dynamics
for two diagnostic groups. So, an
unbiased point estimator for the
differences between two population
proportion is provided by the difference
between two sample proportions, which
are given by P1 hat minus P2 hat. Like
other confidence interval estimates,
here we'll use the same expression we
will we will use the estimate for the
difference between two proportions. Plus
minus
the reliability factor
and the standard error of estimate for
the difference between two population
proportion.
When N1, that is the sample size for the
first population, and N2 is a sample
size for the second population are large
and the population proportions are not
too close to zero or one, we make use of
the central limit theorem and assume
that the normal distribution theory may
be employed to obtain the confidence
interval. Hence, 100 into 1 minus alpha
percent confidence interval for
difference in proportions
can be obtained as P1 hat minus P2 hat
plus minus Z 1 minus alpha by 2, which
is a reliability factor, and since here
we are assuming the normal distribution,
reliability factor will be obtained
from the standard normal distribution.
And the standard error of the difference
between two population proportions
which is given by square root of
the P1 hat 1 - P1 hat / N1
+ P2 hat into 1 - P2 hat over N2. Let's
take an example.
Conner et al. investigated gender
differences
in proactive and reactive aggression
in sample of 323 children
and adolescents, 68 females and 255
males.
The subjects were from unsolicited
consecutive referrals to a residential
treatment center and a pediatric
psychopharmacology clinic serving a
tertiary hospital and medical school. In
the sample, 31 of the females and 53 of
the males reported sexual abuse.
We wish to construct a 99% confidence
interval
for the difference between the
proportions of sexual abuse in the two
sample groups.
So, the sample proportions for the
females and males are respectively given
by P hat F for females, which is 31
/ 68, that is total,
= 0.4559,
and PM hat
represent proportion for males,
which is equals to 0.2078.
And the difference between sample
proportions
is
P F hat - PM hat, which is given by
0.2481.
The estimated standard error of the
difference between sample proportion
is obtained
by the formula for the standard error,
and it's given here as 0.06 55.
The reliability factor
can be obtained from the standard normal
distribution, and here in this case,
our interest is to obtain 99% confidence
interval, hence, we look at the
confidence level of 99 and the area
under the curve
where it is 90.9950,
the Z alpha beta value is 2.58.
Using this reliability factor,
difference in proportion, and the
standard error of estimate for the
difference in proportion,
the confidence interval that we get is
0.0791
as a lower confidence limit and 0.4171
as upper confidence limit.
Interpreting it as that we are 99%
confident that for the sample
population, the proportion of the cases
of reported sexual abuse among female
exceeds the proportion of cases of
reported sexual abuse among males by
somewhere between 0.0791
and 0.4171.
Since the interval does not include zero
value,
hence we can conclude that
the proportions
are different.