Testing Difference of Means in Two Dependent Samples (I) | Applied Biostatistics | BIO733_Topic140
Watch on YouTubeVideo summary
The video introduces the concept of paired samples, which are used when researchers cannot obtain independent observations from two different groups but instead work with related or dependent data. This situation often arises in experiments where the same subject is measured before and after a treatment, such as in studies involving littermates, twins, siblings, or even when comparing two analytical methods on the same material. The primary objective of using paired comparisons is to eliminate extraneous sources of variation that might otherwise mask true differences between populations. By matching subjects based on specific characteristics like digital dexterity or by measuring the same individual over time, researchers can isolate the effect of the treatment more effectively than with independent samples.
When analyzing dependent samples, the choice between parametric and non-parametric tests depends on whether certain assumptions are met. If the data meets these criteria, a parametric test known as the paired samples t-test is preferred; otherwise, a non-parametric alternative like the Wilcoxon signed rank test is used. The key assumptions for the paired samples t-test include having continuous variables measured on an interval or ratio scale, ensuring that observations within the pairs are independent of one another, verifying that the differences between pairs follow an approximately normal distribution, and checking for the absence of outliers. If these conditions are satisfied, the analysis proceeds by calculating the difference ($d$) for each pair and testing whether the mean difference differs significantly from zero.
The video illustrates this process with a specific example involving gallbladder function in patients undergoing fundoplication surgery to treat gastro-esophageal reflux disease. Researchers measured the gallbladder ejection fraction (GBEF) before and after the operation, calculating the difference as post-operative values minus pre-operative values. Since the goal was to determine if the surgery increased GBEF, the hypothesis was formulated to test if the mean difference is greater than zero. With a significance level set at 0.05, the calculated t-statistic was found to be 1.9159, which exceeded the critical value of 1.7959 for a one-tailed test with 12 observations. Consequently, the null hypothesis was rejected because the test statistic fell within the rejection region of the distribution curve.
Based on the statistical findings, the study concludes that there is sufficient evidence to support the effectiveness of the fundoplication procedure. The rejection of the null hypothesis indicates that the mean difference in gallbladder ejection fractions is indeed positive, meaning that post-operative percentages are significantly higher than pre-operative ones. This application conclusion confirms that the surgery successfully improves gallbladder functioning, validating the use of paired sample analysis to detect true treatment effects while controlling for individual variability. The example demonstrates how proper hypothesis formulation, assumption checking, and statistical testing can lead to robust conclusions in applied biostatistics involving dependent samples.
Read the full video transcript
Many times
when we are carrying out some
experiments
and we are interested in looking at the
effect
but we cannot obtain independent
observations from two different samples.
This is a situation
where we
get our data
coming from related
samples
or maybe one respondent giving two
information.
This is a situation
where our samples are dependent.
And this specific situation is called
paired samples or the samples with
multiple responses.
In this module
we'll learn
that how
to perform the appropriate test
when we have two samples which are
dependent. Let's first talk about paired
comparisons.
It's a method that's frequently employed
for assessing the effectiveness of a
treatment
or experimental procedure. One that
makes use of related observations
resulting from non-independent samples.
And a hypothesis that's based on this
type of data
is known as
paired comparison test. It's not always
possible to obtain sample that is
independent.
Hence there are certain reasons
for pairing
the data.
So it frequently happens when the true
differences do not exist between two
populations
with respect to the variable of interest
but presence of extraneous sources of
variation
may cause rejection of the null
hypothesis.
In such instances true differences may
be masked
by the presence of
extraneous factors
that are pre-present
or they have been there because of the
process we are using.
So, the objective in paired comparison
test is to eliminate
a maximum number of sources of
extraneous variation
that masks
the real effect.
So, that we can establish that whether
there is a true difference between
the groups or not.
Pairing is a process
that we often use
to make sure that our
two groups are matched. They are matched
on a number of ways.
The first is when the same subject may
be measured before and after receiving
some treatment. The other scenario is
when littermates of the same sex may be
assigned randomly to receive either a
treatment
or a placebo. Or their pairs of twins or
siblings may be assigned randomly to two
treatments in such a way that the
members of a single pair receive
different treatment. Or
when we are comparing two methods of
analysis
the material to be analyzed may be
divided equally so that one half is
analyzed by one method and other half by
another method. Or one is doing some
matching of individuals on certain
characteristics. For example,
digital
dexterity
which is so which is closely related to
the measurement of interest.
Say post treatment scores
or some test requiring digital
manipulation. When we have the paired
data, we resort to
test them using paired sample tests.
And when we are performing paired sample
tests
instead of performing the analysis with
individual observation, we use DI where
DI is the difference between pairs of
observations. The choice of using either
parametric test or non-parametric test
for the comparison of dependent sample
solely depends upon the fulfillment of
certain assumptions.
And if those assumptions
holds true
we prefer
parametric test which is paired samples
t-test.
But if the assumptions are not met
then it's non-parametric counterpart is
used
which is known as Wilcoxon signed rank
test.
Here are some assumptions that we need
to test when we are making comparison
between
paired groups. The very first assumption
asks for
our variables of interest to be
continuous. They can either be measured
on a ratio scale or an interval scale.
The second important assumption is
independence.
Though we are talking about dependent
samples
but still we are looking for
independence but this independence does
not mean the independent samples but it
is the independence that the
observations within the sample are
independent to one another. Third
important assumption is that of
normality
where
we test the differences
of the pairs
and believe that that follows the
approximately normal distribution.
And the fourth important assumption is
that there should be
no outliers in the data.
So if there are any extreme values in
the data
then
we can
go with the appropriate test.
So
when we test these assumptions
and all these assumptions holds true
we can
conclude that
we can use parametric test.
And when we are using parametric test
for the comparison of
difference between two means in the
dependent samples,
we perform a test
that's called paired samples t-test.
In a paired sample t-test,
if you look at the process of testing of
hypothesis in paired sample t-test, the
very first thing is to state the null
hypothesis and alternative hypothesis.
The null hypothesis is stated as mu d
equals to zero against alternative that
mu d is not equals to zero.
Here, mu d is the difference between the
means of the two groups.
The second step requires us to state the
level of significance, which is alpha is
equals to 0.05.
And thirdly, the test statistics. And
since we say all the assumptions hold
true,
>> [snorts]
>> we will perform paired samples t-test
where t
test statistics is defined as d bar
minus mu d divided by standard deviation
s d bar.
Here, d bar is sample mean of difference
where mu d is hypothesized population
mean difference.
And s d bar is the standard error for
the difference, which is s d divided by
under root n.
Where s d is a standard deviation of the
sample differences. And as a fourth
step, we perform certain calculations.
Here, we'll perform them manually as
well as using SPSS. And then, decision
rule. We reject h naught if p value is
less than equals to zero. Or we can use
critical region approach using usual t
statistics and t distribution.
And finally, we state the conclusion.
And our conclusion can be
uh needs to be statistical conclusion
and application conclusion. Let's take
an example where we have a data that was
obtained from Gorton John M. Morton et
al. examined
gallbladder function before and after
fundoplication.
This is a surgery that's is to stop
stomach contents from flowing back into
esophagus.
In patients with gastro-
esophageal reflux disease,
the authors measure
gallbladder functionality by calculating
the gallbladder ejection fraction, which
is measured as percent. So, we wish to
know if these data provide sufficient
evidence to allow us to conclude that
fundoplication increases
GBEF
functioning.
So, here we have
the data given in percents
before and after the treatment.
Before the treatment values are stated
as pre-op.
That's a
prior to operation.
And then there's post-op.
All these values are stated in
percentages.
So, we want to make comparison between
these two.
So, here we'll say that sufficient
evidence is provided for us to conclude
that the fundoplication is effective if
we can reject the null hypothesis
that the population mean change, that is
mu d, is different from zero in the
appropriate direction.
Further, we shall perform the
statistical analysis on the differences
in pre-op and post-op
GBEF.
We may obtain differences in one of the
two ways, whether we can subtract
post-op from pre-op or pre-op from
post-op. One has to make a choice which
one they use out of these two.
So, here we have this data from
pre-op and post-op and we want to take
the differences and in this situation we
are taking di, which are the
differences, as post-op minus pre-op and
the difference differences turn out to
be 41.5, 28.2 and likewise for the rest
of the other observation.
Now, while we are performing
paired sample t-test,
it's very important that we We for the
assumptions.
As stated earlier, we have four
assumptions to take care of.
But here for the sake of uh
So, we assume that
the observed differences constitute a
simple random sample
from a non- normally distributed
population of differences.
The next step is
the hypothesis.
The way we state our null hypothesis and
alternative hypothesis
must be consistent with the way in which
we subtract measurements to obtain the
differences. So, the question is that
what does it mean by calling something
effective? So, in the present example,
we we want to know if we can conclude
that fundoplication is useful in
increasing GBEF percentage. So, the
question is that what does it mean by
calling something to be effective?
And especially in the present example,
we want to know if
we can conclude that fundoplication is
useful in increasing gall
bladder
ejection fraction
percentage.
We expect
the post-op percent to be larger than
pre-op percent.
Hence, the difference post-op minus
pre-op should tend to return positive
values.
And if we are making the difference from
post-op minus pre-op, our hypothesis
should be stated in the same way, too.
Furthermore,
we would expect the mean of population
of such differences to be positive, too.
Having positive result indicates
that the values
in the pre-op are smaller as compared to
the values in the post-op.
And here, the higher percentage
indicates that
the treatment is effective.
So, under these conditions, asking if we
can conclude that the fundoplication is
effective is the same as asking if we
can conclude that the population mean
difference is positive, which means that
it is greater than zero.
So, using this information, our null and
alternative hypothesis can be stated as
H0 μD is less than equals to zero
against the alternative hypothesis that
μD is greater than zero.
But, if we had obtained the differences
by subtracting the post-op percent from
the pre-op weights,
that is pre-op minus post-op, our
hypothesis would have been reverted.
And it could have been stated as μD
greater than equals to zero against
alternative hypothesis that μD is less
than
zero.
Because if we do pre-op minus post-op,
then the values
have to be negative and post-op
percentages should be greater,
contrary to the previous
hypothesis.
But thirdly, if the question had been
such that two-sided test was indicated,
the hypothesis would have been
H0 that μD is equals to zero against
alternative hypothesis that μD is not
equal to zero.
But here,
in our situation, since our interest is
in looking at
if fundoplication is effective,
and to do so,
we expect that post-op values should be
larger.
And having the difference from post-op
minus pre-op, we expect the positive
values,
the values which are greater than zero
should be
there to prove that fundoplication is an
effective treatment.
Hence, the first hypothesis that says H0
μD is less than equals to zero against
alternative hypothesis that μD
greater than zero is the most
appropriate hypothesis in this
situation.
So, our hypothesis is stated, our level
of significance is assumed to be 0.05,
which is 5%
chances of making
type one error.
And after that, we have the test
statistics, assuming all the assumptions
stated priorly
are true,
we have
paired samples t-test with t statistics
is d bar minus mu d over s d bar. Using
this information and making the
calculations, we have 12 observations.
We have taken the differences. D bar
turned out to be 18.075 against the
stand s d squared equals to 1068. Using
this information in t statistics are t
desert turned out to be 1.9159.
It's a decision rule
where one can use the p value approach
where we can say that we reject h naught
if p value is less than equals to alpha,
that is 0.05.
Or one can use critical region approach
where we can assume and let that alpha
0.05 the critical value of t is 1.7959.
And we reject h naught if computed t is
greater than or equal to the critical
value and it's been reflected in the
picture
given below. And in this figure, the
shaded region is uh
the rejection region. And the other
unshaded region in this bell curve is
non-rejection region.
And if our test statistics falls in the
rejection in the shaded region, that's
the rejection region, we will reject our
null hypothesis. And if it falls under
non-rejection region, then we will not
be able to reject our null hypothesis.
Using this information, we already know
that our
test statistics value t calculated is
1.9159,
which is greater than 1.7959.
Hence, it falls into the rejection
region. And once it fall into the
rejection region, we may conclude
that we reject the null hypothesis and
our statistical conclusion clearly
states that that we are rejecting the
null hypothesis again and our
application conclusion will state that
we may conclude that fundoplication
procedure increases gallbladder ejection
fraction functioning.