Basic Concepts in Sampling | Applied Biostatistics | BIO733_Topic060
Watch on YouTubeVideo summary
In the realm of empirical research, determining the appropriate sample size is a fundamental step that directly influences the validity of any inference drawn about a larger population. The sample size, denoted by the lowercase 'n', represents the specific number of elements selected from the total population, which is represented by the uppercase 'N'. Deciding on this number requires a careful balance of several critical factors, including the budget for data collection, the desired level of statistical power, the required confidence level, and the acceptable margin of error. Ultimately, the goal is to construct a sample that is not only representative but also cost-effective and accessible, ensuring that the findings accurately reflect the target population without unnecessary expense or complexity.
The characteristics of an ideal sample extend beyond mere size; it must be random, possess low sampling error, maintain a high confidence level, and remain unpredictable while being fully representative of the population. A crucial factor influencing these qualities is the homogeneity of the population under study. When a population is homogeneous, meaning its members share similar characteristics, researchers can often achieve reliable results with a very small sample size because the variation within the group is minimal. Conversely, if the population is heterogeneous or diverse, the sample size must be significantly increased to capture all the different aspects and variations present, ensuring that no specific subgroup is overlooked in the analysis.
In addition to sampling error, which arises from the natural unrepresentativeness of a sample and causes the statistic to differ from the true population parameter, researchers must also guard against systematic errors known as statistical bias. These biases can stem from various sources such as selection bias, recall bias, or non-response bias, often introduced by voluntary survey participation where only certain types of people respond, or by the wording of questions that influence answers. For instance, phone surveys may suffer from volunteer bias and tend to over-represent the middle class while excluding poorer or wealthier individuals who do not have listed numbers. To mitigate these issues, it is essential to employ randomization techniques using random numbers, which are sequences where values are uniformly distributed and impossible to predict based on past data. By utilizing random numbers during the selection process, researchers can effectively reduce bias and make their sampling units less predictable, thereby strengthening the integrity of their study.
Read the full video transcript
In the process of sampling, there are
few other important aspect that we need
to take care of.
And the first one is the sample size.
The sample size is an important feature
of any empirical study
in which the goal is to make inference
about the population
from the sample.
Whereas a sample size is denoted by n
is a number of population elements which
are going to be studied from the
population of size n.
You might notice
that sample size is always mentioned by
the small n and the population size is
always denoted by capital N.
Sample size determination is one of the
very crucial aspect in any empirical
study.
It's an act of choosing the right number
of observations to include in the
statistical sample.
Sample size n
in a study is determined based upon
various aspects.
It could be expense of data collection,
sufficient amount of statistical power,
confidence level,
and there is also the margin of error
that could come in to our analysis.
So, while
we are determining the sample size for
any study,
we need to take care of all these
various factors.
It's always good to have the best sample
which is the most representative.
So, there are few characteristics of a
good sample
that it is random.
It has low sampling error,
high confidence level, it is
unpredictable,
it is representative
of the population,
it is accessible,
and it is low cost.
While we are determining the sample
size, one other aspect that should be
considered is the homogeneity of the
population.
If the population is homogeneous,
then it is very easy and convenient to
draw the sample
from such a population.
Because if all the objects under study
are similar in characteristics,
then a very small sample size of two
could be very possible. And it's very
likely that it will give us a pertinent
information regarding the whole
population.
Or at least the target population.
But
if the population is non-homogeneous,
that is heterogeneous,
in that situation,
it's very important to cover all the
different aspects, all the different
characteristics in the population.
And
when the population is heterogeneous,
then we need to increase our sample
size, as well.
To cover up and to capture all the
variation caused by all the different
aspects or all the different
characteristics of our population.
There's another important aspect that
need to be considered while we are
sampling is the errors
in sample.
So, in the process of sampling, there
are it's very likely that two types of
error could come in.
The first is sampling error, which is
also called random error.
And the other one is systematic error,
which is also called statistical bias.
So, sampling error arises from the
unrepresentativeness of the sample
taken.
It is it's a difference between a sample
statistic and the true value of the
population parameter.
Here, we need to understand that
population values are always true values
and they are fixed. They are constant.
But sample values, they do keep on
changing from sample to sample.
And as we spoke about the qualities of
good sample earlier, it is very
important that our sample should give us
the best answer that is more closest to
our true value of the population.
That's why
if the sampling error is small, it means
that our sample statistic is closer to
our population parameter.
A statistical bias is a feature of a
statistical technique or its results
whereby the expected value of the result
differ from the underlying quantitative
parameter
being estimated.
There are various kind of statistical
bias.
It could be selection bias.
It could be self-selection bias.
It could be recall bias, observer bias,
omitted variable bias, funding bias, and
non-response bias.
So, all these bias can be introduced,
and they also can be countered by using
different type of methodologies and
different type of re- remedial measures.
Here are a few sources of bias.
So, people who respond to voluntary
surveys tend to have different
parameters
than people who do not respond.
And such type of bias that comes in is
called volunteer bias.
People tend to exaggerate or round off
things such as their weight or age when
asked, depending upon the context of the
question.
Similarly, the wording of the question
or the nature of the interviewer can
suggest certain types of answers
from people.
Such type of bias is called response
bias.
Phone surveys in particular are
susceptible to volunteer bias and also
tend to over-represent the middle class.
In general, poor people and rich people
are less likely to have listed numbers
as well.
So, these are a few sources from where
our bias in our studies can creep in.
Then, the other important aspect is the
use of random numbers, where random
numbers are the numbers that occur in a
sequence such that the these two
conditions meet. The first condition
says that the values are uniformly
distributed over a defined interval or
set.
And the second condition says that it is
impossible to predict future values
based on past or present ones.
So, that's why it's more called a random
number
once for randomization in the selection
process.
Use of random numbers help reduce the
amount of bias in the sample as well.
And using the random number will make
our sample and our population sampling
units
more uh less predictable. Thank you.