Sampling Distribution: Expected Value and Variance | Applied Biostatistics | BIO733_Topic069
Watch on YouTubeVideo summary
In this module on applied biostatistics, the focus shifts to understanding the fundamental characteristics of any sampling distribution, specifically its mean and variance. The mean of a sampling distribution is formally known as the expected value of a statistic, often denoted as $E[X]$ or $\mu_X$. Conceptually, the expected value represents the long-run average; as more values of a random variable are collected, the sample mean converges toward this expected value. Mathematically, it is calculated as the sum of each possible value multiplied by its respective probability. When applied to an estimator or statistic denoted as $\theta$, the expected value is found by summing the product of every possible value of $\theta$ and its probability of occurrence across all potential outcomes.
Beyond the definition, the video outlines several key algebraic properties of the expected value that simplify calculations involving random variables. First, the expected value of a constant is simply the constant itself. Second, the expected value of the sum of two random variables, $X$ and $Y$, is equal to the sum of their individual expected values, demonstrating that expectation is a linear operator. Third, constants can be factored out of the expectation operation; specifically, the expected value of $A$ times $X$ plus $B$ equals $A$ times the expected value of $X$ plus $B$. These properties are essential tools for manipulating statistical expressions before deriving more complex measures like variance.
The second major characteristic discussed is variance, which quantifies the dispersion of the sampling distribution around its mean. The variance of a statistic $\theta$ is defined as the difference between the expected value of the square of $\theta$ and the square of the expected value of $\theta$. In practical terms, this involves calculating the sum of squared values weighted by their probabilities to find $E[\theta^2]$, then subtracting the square of the previously calculated mean. While knowing the mean and variance provides significant insight into an estimator's behavior, the video notes that variance alone does not fully describe how the estimate is dispersed. It also introduces the concept of standard error as another critical characteristic that will be explored in future discussions, emphasizing that understanding both central tendency and spread is vital for interpreting statistical estimates accurately.
Read the full video transcript
Since we know that any sampling
distribution will have certain
characteristics, in this module we will
be talking about two of its
characteristics. One is mean,
that is also called expected value
of an of a statistic, and the other one
is the variance. The expected value of
any random variable is often written as
expected value of X or mu of X.
The expected value is the long-run mean
in the sense that if as as more and more
values of the random variable were
collected, the sample mean becomes
closer to the expected value. Expected
value of X is often given as sum xi with
the probability of xi. And the same idea
translates to the estimates or the
statistic. Let's say theta is our
statistic of interest, and we want to
establish we want to form that sampling
distribution and want to calculate its
expected value. Then expected value of
theta will be sum I goes from 1 to c,
theta i
multiply by the probability of theta i.
So theta i is the possible value of
the estimator or the statistic.
Whereas p of theta i is their respective
probability of occurrence. So there are
few properties of expected value.
Expected value of the constant is the
constant itself. Similarly, if X and Y
are two random variables
on a sample space, then the
the expected value of X plus Y will be
equals to expected value of X plus
expected value of Y. And the third
property says that if A and B are
constants, then expected value of A
times X plus B
equals to A times expected value of X
plus B. The other important
characteristic is the variance.
Where the variance of the sampling
distribution
is given as the variance of theta equals
to expected value of theta square minus
expected value of theta whole square.
And when we are dealing with the
sampling distribution, we calculate
these expected values using the
summation signs. Where expected value of
theta square equals to summation theta I
square probability of theta I.
And expected value of theta as usual,
the way we calculate the regular
expected value. So, hence take the
example of the sampling distribution of
a of an estimator theta. We already know
that estimator theta has that various
values theta one up to theta five and it
has frequency and it has certain
probability of occurrence. For theta
one, the probability of occurrence is N1
divided by N.
And similarly for rest of the possible
values of theta. So, the next column
theta I probability of theta I is if you
add if you calculate these values and
add them up, this is going to give you
summation theta I probability of theta I
which is equivalent to the expected
value of theta. And the last column that
talks about theta I square probability
of theta I, if you if you if you make
certain calculations and add them up,
this is going to give you expected value
of theta I square. Knowing few
characteristics like mean and variance
of a sampling distribution
is very helpful to understand the
estimates.
We can always use expected value of
theta I square and expected value of
theta to calculate the variance.
But variance is not the
the the only characteristic that help us
to look at that how our theta is
dispersed.
The other important characteristic of
our our estimate, our sampling
distribution is a standard error,
which we will talk later. Thank you.