Submind YouTube summaries
Thumbnail for Basic Concepts in Sampling | Applied Biostatistics | BIO733_Topic060

Basic Concepts in Sampling | Applied Biostatistics | BIO733_Topic060

Watch on YouTube

Video summary

In the realm of empirical research, determining the appropriate sample size is a fundamental step that directly influences the validity of any inference drawn about a larger population. The sample size, denoted by the lowercase 'n', represents the specific number of elements selected from the total population, which is represented by the uppercase 'N'. Deciding on this number requires a careful balance of several critical factors, including the budget for data collection, the desired level of statistical power, the required confidence level, and the acceptable margin of error. Ultimately, the goal is to construct a sample that is not only representative but also cost-effective and accessible, ensuring that the findings accurately reflect the target population without unnecessary expense or complexity. The characteristics of an ideal sample extend beyond mere size; it must be random, possess low sampling error, maintain a high confidence level, and remain unpredictable while being fully representative of the population. A crucial factor influencing these qualities is the homogeneity of the population under study. When a population is homogeneous, meaning its members share similar characteristics, researchers can often achieve reliable results with a very small sample size because the variation within the group is minimal. Conversely, if the population is heterogeneous or diverse, the sample size must be significantly increased to capture all the different aspects and variations present, ensuring that no specific subgroup is overlooked in the analysis. In addition to sampling error, which arises from the natural unrepresentativeness of a sample and causes the statistic to differ from the true population parameter, researchers must also guard against systematic errors known as statistical bias. These biases can stem from various sources such as selection bias, recall bias, or non-response bias, often introduced by voluntary survey participation where only certain types of people respond, or by the wording of questions that influence answers. For instance, phone surveys may suffer from volunteer bias and tend to over-represent the middle class while excluding poorer or wealthier individuals who do not have listed numbers. To mitigate these issues, it is essential to employ randomization techniques using random numbers, which are sequences where values are uniformly distributed and impossible to predict based on past data. By utilizing random numbers during the selection process, researchers can effectively reduce bias and make their sampling units less predictable, thereby strengthening the integrity of their study.
Read the full video transcript
In the process of sampling, there are few other important aspect that we need to take care of. And the first one is the sample size. The sample size is an important feature of any empirical study in which the goal is to make inference about the population from the sample. Whereas a sample size is denoted by n is a number of population elements which are going to be studied from the population of size n. You might notice that sample size is always mentioned by the small n and the population size is always denoted by capital N. Sample size determination is one of the very crucial aspect in any empirical study. It's an act of choosing the right number of observations to include in the statistical sample. Sample size n in a study is determined based upon various aspects. It could be expense of data collection, sufficient amount of statistical power, confidence level, and there is also the margin of error that could come in to our analysis. So, while we are determining the sample size for any study, we need to take care of all these various factors. It's always good to have the best sample which is the most representative. So, there are few characteristics of a good sample that it is random. It has low sampling error, high confidence level, it is unpredictable, it is representative of the population, it is accessible, and it is low cost. While we are determining the sample size, one other aspect that should be considered is the homogeneity of the population. If the population is homogeneous, then it is very easy and convenient to draw the sample from such a population. Because if all the objects under study are similar in characteristics, then a very small sample size of two could be very possible. And it's very likely that it will give us a pertinent information regarding the whole population. Or at least the target population. But if the population is non-homogeneous, that is heterogeneous, in that situation, it's very important to cover all the different aspects, all the different characteristics in the population. And when the population is heterogeneous, then we need to increase our sample size, as well. To cover up and to capture all the variation caused by all the different aspects or all the different characteristics of our population. There's another important aspect that need to be considered while we are sampling is the errors in sample. So, in the process of sampling, there are it's very likely that two types of error could come in. The first is sampling error, which is also called random error. And the other one is systematic error, which is also called statistical bias. So, sampling error arises from the unrepresentativeness of the sample taken. It is it's a difference between a sample statistic and the true value of the population parameter. Here, we need to understand that population values are always true values and they are fixed. They are constant. But sample values, they do keep on changing from sample to sample. And as we spoke about the qualities of good sample earlier, it is very important that our sample should give us the best answer that is more closest to our true value of the population. That's why if the sampling error is small, it means that our sample statistic is closer to our population parameter. A statistical bias is a feature of a statistical technique or its results whereby the expected value of the result differ from the underlying quantitative parameter being estimated. There are various kind of statistical bias. It could be selection bias. It could be self-selection bias. It could be recall bias, observer bias, omitted variable bias, funding bias, and non-response bias. So, all these bias can be introduced, and they also can be countered by using different type of methodologies and different type of re- remedial measures. Here are a few sources of bias. So, people who respond to voluntary surveys tend to have different parameters than people who do not respond. And such type of bias that comes in is called volunteer bias. People tend to exaggerate or round off things such as their weight or age when asked, depending upon the context of the question. Similarly, the wording of the question or the nature of the interviewer can suggest certain types of answers from people. Such type of bias is called response bias. Phone surveys in particular are susceptible to volunteer bias and also tend to over-represent the middle class. In general, poor people and rich people are less likely to have listed numbers as well. So, these are a few sources from where our bias in our studies can creep in. Then, the other important aspect is the use of random numbers, where random numbers are the numbers that occur in a sequence such that the these two conditions meet. The first condition says that the values are uniformly distributed over a defined interval or set. And the second condition says that it is impossible to predict future values based on past or present ones. So, that's why it's more called a random number once for randomization in the selection process. Use of random numbers help reduce the amount of bias in the sample as well. And using the random number will make our sample and our population sampling units more uh less predictable. Thank you.