Introduction to Sampling | Applied Biostatistics | BIO733_Topic059
Watch on YouTubeVideo summary
The module introduces the fundamental rationale behind obtaining a sample and outlines the procedures involved in its collection, highlighting both the advantages and disadvantages of this approach. Before delving into these details, it is essential to establish basic terminology, distinguishing between the population, which represents the totality of characteristics under study, and the census, which refers to studying the entire population. Within this framework, individual members are known as population elements, while the true values observed from them are called population parameters. Sampling serves as the process of selecting a subset of observations, referred to as a sample, to provide an adequate description and draw inferences about the larger population, making the sample the most representative part of that whole.
A critical component ensuring the accuracy and reliability of sampling procedures is the sampling frame, defined as a complete list of all individuals included in the population from which the sample will be drawn. While the population encompasses everything we wish to discuss, the sampling frame specifically includes those elements available for study. The primary motivation for using samples rather than conducting a census on the entire population stems from practical limitations; populations are often too large, making a full census expensive, time-consuming, and prone to other logistical issues. Consequently, researchers aim to obtain a representative sample that allows them to speak economically about the population without needing to observe every single member, a process known as statistical inference when moving from sample observations back to population conclusions.
Despite these benefits, sampling is not without its drawbacks, primarily revolving around variability and potential bias. Since samples can vary significantly from one another, two different samples might yield totally different answers or only slight variations, introducing uncertainty into the results. Furthermore, there is always a risk of introducing bias during data collection, which can stem from human error or the difficulty in securing a truly representative sample. In large-scale surveys, reliance on manpower becomes necessary, and if this workforce is not adequately trained, it can lead to further errors. Additionally, issues such as the absence of informants and various other types of errors can compromise the accuracy of the findings, making it challenging to ensure that the sample perfectly reflects the population it intends to represent.
Read the full video transcript
Hello. In this module,
we will be talking about
the rationale
of obtaining a sample.
We will further look at
the procedure of collecting a sample
as well as its advantages and
disadvantages.
Before we move further, we do want to
talk about few basic terminology
that includes the population,
which is the totality of aggregate of
characteristic that is under our study.
Whereas, whenever we we want to study
the population,
the process or the study of the
population is called census.
In a population, there are individual
observations.
Each individual member
that for which the characteristics we
want to study
is called the population element.
Whereas, any value or characteristic
that we observe from the population is
called population parameter,
which are the true values and they are
constants.
Hence, the sampling is a process of
selecting the observations,
which is called a sample,
to provide an adequate description
and inferences of the population.
So, sample is the most representative
part of the population.
Whereas, the sampling units
are
the individual person, animal,
or object that has the measurement or
observation
taken on them.
But, very important
feature that makes our sampling
procedures much more accurate and
reliable is the sampling frame.
Whereas, the sampling frame is a
complete list of all the individuals
that are included in our population
or from which we are going to collect
our sample.
So, as we know
that we always want to talk about the
population.
Whereas the sampling frame is a part
of our population
which specifically we want to study or
which are available for our sample.
The idea is that we cannot study the
population in all the situations.
Because population is too large and to
conduct the censors out of this
population is always
having a lot of other [snorts]
drawbacks. Like it could be expensive.
It could be time-consuming.
And there are many other issues that can
come in.
So, to be more economical but still
being able to talk about the population,
we always try to take a representative
part of the population which we call
sample.
The sampling is a process of obtaining
that sample from the population.
So, population is something
that about which we want to talk about
and sample is something that we actually
observe
to talk about the population.
And when we talk, we go from the sample
to the population, the whole process is
called the statistical inference.
Since we know we cannot study overall
population in all these instances,
therefore we rely always on a sample.
And it is very important that our sample
should be good.
There are few advantages of the sampling
that it is economical in nature.
Because we don't have to study the
overall population.
Moreover
if it's carefully done, then it gives us
very accurate results regarding the
population.
Moreover, it also give us very reliable
estimates for the population.
It also saves time and only resort to
the state if we are about to study an
infinite population.
Because practically it's really hard to
study the whole population in in case of
infinite populations.
But as every
method has its disadvantages, too.
It Sometimes it could be inadequate.
Because samples do vary from sample to
sample.
So, it's very likely the two samples may
give us totally different answers.
Or it's most likely that they are
slightly different or even the same.
There's always a chance of bringing some
bias
into the results when we are collecting
the sample and that could be
human error as well.
Then there's also problem of accuracy in
our results.
And if we are taking the sample,
there's Sometimes we get into the
problem that it's very difficult to get
the representative sample.
Since for to to obtain a sample, we
always rely on some manpower, so it's
very likely that manpower is not very
trained.
Though for the smaller studies,
a PI can himself take the sample.
But if we are conducting our large-scale
surveys,
then we do need some manpower and it's
very likely that manpower is not trained
well.
Then
there it's very likely that our
informants are absence.
Moreover, there are chances of
committing the errors in the sampling.
Or there are different other type of
errors that that can come in the play.
Thank you.