Sampling Distributions | Applied Biostatistics | BIO733_Topic068
Watch on YouTubeVideo summary
Sampling distributions are a fundamental concept in biostatistics that address the variability inherent when drawing multiple samples from the same population. Since it is often impossible to study an entire population, researchers rely on samples to estimate population parameters; however, different samples yield different estimates for the same statistic. The sampling distribution describes the probability of obtaining various values for a specific statistic across all possible random samples of a fixed size drawn from that population. This concept helps analysts understand how likely it is to observe a particular value and provides insight into the distribution of these statistics within the broader context of the population.
The process of constructing a sampling distribution depends on whether sampling is done with or without replacement, which significantly affects the number of possible samples. In sampling with replacement, an individual unit is studied and then returned to the population, allowing it to be selected multiple times in different samples; mathematically, this results in $N^n$ possible combinations where $N$ is the population size and $n$ is the sample size. Conversely, sampling without replacement involves studying a unit and removing it from the pool, ensuring no individual is studied twice in the same sequence, which leads to $N$ choose $n$ distinct samples. Regardless of the method chosen, if multiple samples are taken, the resulting statistics will vary, and their collective behavior forms the sampling distribution.
To empirically construct this distribution for a discrete and finite population, one follows a systematic three-step approach: first, draw all possible samples from the population; second, calculate the statistic of interest for each sample; and third, list these calculated values along with their frequencies to determine relative frequencies or probabilities. This data can be organized into a table showing distinct statistic values and how often they occur, where the sum of the frequencies equals the total number of possible samples. Alternatively, the distribution can be represented graphically by plotting the possible statistic values on the x-axis against their corresponding probabilities or relative frequencies on the y-axis, which often reveals a symmetric pattern, although symmetry is not guaranteed in every case.
Understanding the sampling distribution allows researchers to focus on three critical characteristics: the mean, the variance, and the functional form of the distribution. While determining the exact functional form can be difficult for large or infinite populations, approximations are often used in such scenarios. Additionally, these distributions can sometimes be derived mathematically rather than just empirically. By analyzing these properties, statisticians can make informed inferences about population parameters based on sample data, acknowledging that while individual sample results may vary, the underlying distribution of those results follows predictable patterns that are essential for accurate statistical analysis.
Read the full video transcript
Sampling has a lot of advantages and
disadvantages.
We cannot study pop in the whole
population in all the situations.
That's why we rely on our sample.
But one can take more than one sample
from the same population.
And every time you take a sample,
one can get a different estimates
coming from the sample.
And these estimates could have different
values.
In this module, we will be talking about
the sampling distribution. It talks
about all those various possible values
that can occur
as a result of all possible samples. But
when we draw a sample,
sample could be drawn using two
methodologies.
One way is that we draw a sample with
replacement.
And the other method is that we draw a
sample without replacement.
Though in all the conditions, it's not
always possible to draw sample with
replacement or
sometimes
without replacement only.
In with replacement sampling, we select
a unit,
after studying it, we place it back into
the population.
So in with replacement sampling, it is
very likely that an individual could be
selected twice or even more
from the same population.
Whereas in without replacement sampling,
we select our individual, study it, and
then we do not place them back. So
that's the same individual that has
already been studied will not
be studied again. So,
there there So when we are doing with
replacement sampling,
the number of all possible samples that
could happen
is n to the power n,
where capital N denotes the population
size and small n denotes the sample
size. And likewise,
for with replacement sampling, the all
possible samples drawn could be equals
to n choose n.
Where capital N again represents the
population size and small n represents
the the sample size. Once we have
decided to draw a sample with either
with replacement sampling methodology or
without replacement sampling
methodology.
So, one can take multiple samples from
coming from the same population. When we
take multiple sample and from each
sample, we make we take the value of
certain statistic. The distribution of
all possible values that can be assumed
by some statistic computing from the
sample of the same size
randomly drawn from the same population
is called sampling distribution. When we
are taking multiple samples from the
same population, every time the sampling
units and the sample units are going to
be different. Hence, the statistics
could be different in all the situation
in all the various samples. The sampling
distribution will help us to view that
how what what's most likely for any
specific value of a sample statistic to
show up. Sa- sampling distribution
describe how these different values are
distributed in the population.
Sampling distribution may be constructed
empirically when sampling from a
discrete and finite population. Hence,
we follow these three steps. In the very
first step, from a finite population of
size n, we randomly draw all possible
samples of size n.
Once these all possible samples have
been drawn,
we compute the statistic of our interest
fro- for each sample.
And once we have calculated all the
statistic of interest for each sample,
we'll list them in one column, the
different distinct observed values of
the statistic. And in other column, we
list the corresponding frequency of
occurrence of each distinct observation
value for the statistic.
And eventually, we calculate the
relative frequencies.
Hence, it will form the complete
sampling distribution. Let's look at
this example.
When we are saying that theta is an n n
is a is a statistic,
so the sampling distribution of
this sample statistic theta
could be given in this table.
Where theta one, theta two, theta three,
theta four, theta five are possible
statistic which are obtained as a
process of
calculating the statistics from all
possible samples.
And the frequency n1, n2, n3, n4, n5
represents the number of times all these
statistics
estimator values were obtained. N1
represents the theta one occurred n1
number of times.
And n5 represents that theta five
occurred n5 number of times. And when
you add all these frequencies up, this
should be equals to all possible samples
obtained.
Then, we calculate the relative
frequency. And here, we call it
probability of theta. So, probability of
theta one is equals to n1 divided by n.
So, hence, it's overall, it it should
add up to one because these are the
relative frequencies or the
probabilities and they all add up to
one.
There's another way to represent a
sampling distribution,
and that is graphical way. So, when we
represent in graph, on x-axis, we can
take the possible values of our
statistic.
And on y-axis, it could be probabilities
or the relative frequencies. And if you
look at it, in this picture, it shows
quite symmetric symmetric distribution,
but it's not necessarily true that we
always get a symmetric distribution, but
most of the time we do. There are a few
important characteristics of the
sampling distribution.
We usually are interesting in knowing
three
important characteristics.
The first one is mean,
other is variance, and the functional
form.
For large or infinite population,
sometimes it's not really easy
to find out the functional form of the
distribution.
Hence, in those those situations, we try
to approximate the sampling
distribution.
We can also
obtain the sampling distribution
mathematically.
Thank you.