Submind YouTube summaries
Thumbnail for Sampling Distribution: Expected Value and Variance | Applied Biostatistics | BIO733_Topic069

Sampling Distribution: Expected Value and Variance | Applied Biostatistics | BIO733_Topic069

Watch on YouTube

Video summary

In this module on applied biostatistics, the focus shifts to understanding the fundamental characteristics of any sampling distribution, specifically its mean and variance. The mean of a sampling distribution is formally known as the expected value of a statistic, often denoted as $E[X]$ or $\mu_X$. Conceptually, the expected value represents the long-run average; as more values of a random variable are collected, the sample mean converges toward this expected value. Mathematically, it is calculated as the sum of each possible value multiplied by its respective probability. When applied to an estimator or statistic denoted as $\theta$, the expected value is found by summing the product of every possible value of $\theta$ and its probability of occurrence across all potential outcomes. Beyond the definition, the video outlines several key algebraic properties of the expected value that simplify calculations involving random variables. First, the expected value of a constant is simply the constant itself. Second, the expected value of the sum of two random variables, $X$ and $Y$, is equal to the sum of their individual expected values, demonstrating that expectation is a linear operator. Third, constants can be factored out of the expectation operation; specifically, the expected value of $A$ times $X$ plus $B$ equals $A$ times the expected value of $X$ plus $B$. These properties are essential tools for manipulating statistical expressions before deriving more complex measures like variance. The second major characteristic discussed is variance, which quantifies the dispersion of the sampling distribution around its mean. The variance of a statistic $\theta$ is defined as the difference between the expected value of the square of $\theta$ and the square of the expected value of $\theta$. In practical terms, this involves calculating the sum of squared values weighted by their probabilities to find $E[\theta^2]$, then subtracting the square of the previously calculated mean. While knowing the mean and variance provides significant insight into an estimator's behavior, the video notes that variance alone does not fully describe how the estimate is dispersed. It also introduces the concept of standard error as another critical characteristic that will be explored in future discussions, emphasizing that understanding both central tendency and spread is vital for interpreting statistical estimates accurately.
Read the full video transcript
Since we know that any sampling distribution will have certain characteristics, in this module we will be talking about two of its characteristics. One is mean, that is also called expected value of an of a statistic, and the other one is the variance. The expected value of any random variable is often written as expected value of X or mu of X. The expected value is the long-run mean in the sense that if as as more and more values of the random variable were collected, the sample mean becomes closer to the expected value. Expected value of X is often given as sum xi with the probability of xi. And the same idea translates to the estimates or the statistic. Let's say theta is our statistic of interest, and we want to establish we want to form that sampling distribution and want to calculate its expected value. Then expected value of theta will be sum I goes from 1 to c, theta i multiply by the probability of theta i. So theta i is the possible value of the estimator or the statistic. Whereas p of theta i is their respective probability of occurrence. So there are few properties of expected value. Expected value of the constant is the constant itself. Similarly, if X and Y are two random variables on a sample space, then the the expected value of X plus Y will be equals to expected value of X plus expected value of Y. And the third property says that if A and B are constants, then expected value of A times X plus B equals to A times expected value of X plus B. The other important characteristic is the variance. Where the variance of the sampling distribution is given as the variance of theta equals to expected value of theta square minus expected value of theta whole square. And when we are dealing with the sampling distribution, we calculate these expected values using the summation signs. Where expected value of theta square equals to summation theta I square probability of theta I. And expected value of theta as usual, the way we calculate the regular expected value. So, hence take the example of the sampling distribution of a of an estimator theta. We already know that estimator theta has that various values theta one up to theta five and it has frequency and it has certain probability of occurrence. For theta one, the probability of occurrence is N1 divided by N. And similarly for rest of the possible values of theta. So, the next column theta I probability of theta I is if you add if you calculate these values and add them up, this is going to give you summation theta I probability of theta I which is equivalent to the expected value of theta. And the last column that talks about theta I square probability of theta I, if you if you if you make certain calculations and add them up, this is going to give you expected value of theta I square. Knowing few characteristics like mean and variance of a sampling distribution is very helpful to understand the estimates. We can always use expected value of theta I square and expected value of theta to calculate the variance. But variance is not the the the only characteristic that help us to look at that how our theta is dispersed. The other important characteristic of our our estimate, our sampling distribution is a standard error, which we will talk later. Thank you.