Statistical Rethinking Lecture A10 - Hidden Confounds & Sensitivity Analysis
Watch on YouTubeVideo summary
This lecture emphasizes that scientific literature is often flawed rather than a collection of facts, advocating for an aggressive approach known as "guerrilla warfare" against false beliefs by rigorously interrogating workflows and refusing to accept abstracts at face value. A central theme is the necessity of sensitivity analysis, which involves varying elements within generative models or statistical frameworks to observe how results shift under different causal assumptions; this process highlights that inference relies on strong assumptions rather than being inherently flawed. By treating these assumptions as testable components rather than fixed truths, researchers can better understand the robustness of their conclusions and avoid dismissing confounds simply because they are unobserved.
The instructional content illustrates these concepts through several critical examples, starting with gender discrimination in UC Berkeley admissions where an unmeasured variable like "ability" acts as a hidden confound that creates survivorship bias by masking direct discrimination against women who survive harsher selection processes. The lecture further analyzes conflicting studies on National Academy of Sciences membership to demonstrate collider bias, showing how conditioning on post-treatment variables such as current membership or citation rates can induce spurious correlations that misrepresent the true effects of gender and quality. These scenarios reveal that without accounting for hidden factors like unmeasured ability or field-specific norms, researchers risk drawing misleading conclusions about direct effects versus structural advantages in real-world contexts ranging from pay equity audits to smoking habits among married couples.
To address these challenges, the speaker introduces a Bayesian framework using latent variables to model unobserved confounders alongside observed outcomes, allowing for sensitivity analysis where the strength of hidden effects is varied through fixed values or priors to assess how much bias would be required to explain away an effect. This quantitative approach moves beyond merely identifying potential biases to testing their structural implications and determining whether results are robust or driven by unmeasured factors. The lecture concludes with a comprehensive "guerilla workflow" for maintaining scientific rigor, which involves formulating questions via generative models, justifying statistical assumptions including priors, running diagnostics against data dictionaries, and incorporating qualitative ethnographic data to inform better quantitative models when confounding is unavoidable.
Read the full video transcript
Good morning. Welcome back to the last
lecture of section A of statistical
rethinking 2026.
Um the way I planned this course this
year is uh section A starts at the very
beginning and goes all the way up to
multi-level models. You're not going to
get a lecture on multi-level models
because I've already recorded one for
the other section. You just have to
start their lectures. Okay, I apologize.
It is I'm being one sense I'm being lazy
by not giving you that but in the other
sense I wrote a whole new lecture for
you. So here it is uh and I thought this
is a topic that has more value added um
because there there's tons of existing
lectures and whole books about
multi-level modeling. There's way less
material whether lectures or books about
dealing with hidden confounds and
sensitivity analysis. But this is not an
obscure problem. This is the kind of
problem that is routine. Yeah. Like
missing data again considered an
advanced topic. Um most data are
missing. Yeah. Most of the time, right?
So it's it's actually a core statistical
skill. Uh I'm trying to work those kinds
of things
more forward in the course, but uh it's
difficult given the computational
aspects of them. But you are all ready
now in this section uh to talk about the
serious truth about hidden compounds.
and and I want to give you um some more
reasoning with DAGs about hidden
compounds and the effects they can have
and show you one way to do sensitivity
analysis so you don't have to pretend
these things don't exist. Yeah. Um let
me start though with a little bit of of
overview. Uh the scientific literature
is not a pile of facts, right? uh when
you read papers in your own area where
you have the expertise to to judge them
critically. Are they all good? Do they
all make sense? No, they don't. Right?
Those method sections, it's like that
meme where the with the dirty home, I
forget the cartoon characters and the
one is like you live like this, right?
It's like just it's bad, right? So when
you read things in and in in other
fields where you don't have the
expertise, I would invite you to carry
that skepticism with you. uh and uh you
can do better than that in the sense
that that the tools I'm teaching you in
this course empower you to uh fight the
scientific literature. It is an
overwhelming enemy and you are a gorilla
fighter, right? That you you are you are
outmatched. You are not going to win
this battle through force. Uh you have
to apply your force very specifically
and strategically, right? Gera warfare,
guerrilla warfare is intellectual. Yeah.
That's how you overthrow the occupying
forces, right? Not a head-on assault. Uh
so there's there's all this problems
with the scientific literature, but if
you interrogate the workflow, the papers
you read, try to fill in the pieces of
it, uh you can pull together what is the
logical justification for this analysis,
if there is one or what are the gaps. Uh
if you're a reviewer, it's a way to
review and ask for clarifying
information. If you're reading something
that's been published, it's a way for
you to defend yourself against false
beliefs. Yeah, I mean this. It's really,
really essential. Do not believe the
abstract ever. Never ever believe an
abstract. Right? What's that abstract
there? The abstract is propaganda.
Sorry though, this is like this is I'm
going to run for run for office. You can
feel it, right? Um
uh now, of course, all this applies to
your own work too. Uh be the change you
want to see in others, right?
Have the workflow reported in a
structured way so that people can uh so
that you respect the rights of your
colleagues to make up their own minds
about your work. Make your methods
comprehensible and insist that the
published literature be that and if it
isn't then doubt it. Yeah. Okay. That's
my sermon. Thank you very much. Along
that line, one aspect of um a
comprehensible workflow is sensitivity
analysis. Uh so this can mean lots of
different things. you can change the
elements of the uh of the workflow and
see how the results change. So often
this means changing the generative model
that would be like the causal
assumptions that it could be a
structural diagram or it could be other
aspects of the causal model like the
functional responses of one variable on
another. Um and also aspects of the
statistical model which are not always
part of the generative model, right?
Like ways you handle missing data.
uh Gaussian processes are not a causal
model but they're a way to estimate
functions. Yeah, I mean you have
different strategies for that. Um and
then there's so we can we can vary as I
say the G's and S's in this diagram. I
would not normally label all this stuff
but this is a test of your you've been
here a long time you can guess what Q
is. That's your question. Yeah. G is a
generative model. S is a statistical
model. E is estimates. P are
predictions. D is data. And the idea is
we test all these things. And if you're
not testing, if you test nothing, you
miss everything, right? You have to
test. Um,
and sensitivity is when we vary these
elements and then we uh repeat the
workflow and we see how the results
change. Yeah. And that's that's often a
good thing to do. If you have questions
about the about the strength of your
assumptions uh uh the difference they
make uh in the results, then sensitivity
testing is is is good. It's a reasonable
thing to do. But sensitivity is not bad.
And we're going to talk about this,
right? It the fact that the results
change when you change the assumptions
is normal. Yeah, we we buy inference
with assumptions. So just because the
something the results are sensitive uh
to changes in the assumptions is not bad
in and of itself at all. If you can
license the assumptions, then you get
stronger inferences. Yeah. Uh just think
about it this way. Assume the wrong
causal model, assume the right causal
model. results will be sensitive to
those differences, right? But it's clear
that one of them is the right answer,
right? When you simulate the data, uh
when you don't know the right answer,
then you have to think about what
external information to the sample is
going to help me choose between these
these sets of assumptions. Does this
make some sense what I'm trying to say?
I know I've said these things before. I
have I'm workshopping different ways to
say this, but um yeah, we we here I say
we buy inference with strong
assumptions. There's nothing wrong with
assumptions. It's the only way you can
get strong inferences. Yeah. As I said
before, uh inference without assumptions
like an opinion without reasons. You
shouldn't pay any attention to it. Yeah,
that's the way epistemology works in the
sciences. So, it's fine. But sensitivity
analysis is useful in this regard. Okay.
Along those lines, let's think about um
some assumptions and develop a
sensitivity analysis to the consequences
of them. So, in the last lecture, I was
going on and on in a bit of a rant. I
apologize. Actually, I don't about
gender discrimination, other forms of
discrimination, and how we how it's
analyzed in cross-sectional samples.
This is a a topic I've published on. I
care a lot about it. Um, and I'm going
to come back to this example again, the
UC Berkeley data set, just to remind
you. Uh, we we're looking at a very
simplified version of this where we've
got gender here in this DAG as the
exposure of interest. We want to know if
there's direct gender discrimination um
conducted by admissions officers uh in
this data set. Um individuals apply to
departments and that is partly
influenced by gender. They're very
gender patterned choices of which
departments to apply to. And then the
outcome here is admission. And there
could be a direct effect of gender which
we'd interpret as as direct or taste
based discrimination. And there can be
indirect u discrimination because some
departments just have lower rates of
admission than others. Yeah. And I tried
to convince you in the last lecture that
both of these things are probably going
on in this data set. Yeah. As in many
data sets of this kind. Um now what I
want to convince and but all in that
lecture we ignored um unmeasured
confounds. I hinted at them and told you
we'd come back to them. So this lecture
is all about the the unmeasured confound
which I think makes these analyses quite
difficult. uh but you still have to do
them because the topic matters right
people are going to report estimates you
have reporting a biased estimate is not
always a bad thing you just have to be
transparent about what assumptions are
required so in particular in this
literature we're worried about this
little U thing which is unobserved some
unobserved variable by which applicants
vary and and that feature which I'm
going to call ability because this is
the way it's usually discussed in the
literature influences both the
probability your application is accept
accept it, right? Because if you're of a
high ability individual, you make a
better application, right? And that
should help. Um, and it also influences
the department you apply to because some
departments are harder to win admission
to. Uh, and if you're a high ability
individual and you know that, that will
change your probability of applying to
those departments. That make sense? Um,
so there's there are investigations
trying to back up those assumptions and
such. And I'm going to uh with apologies
skip all that and move forward into the
modeling part which is usually harder to
find. So the first thing you do let's
let's make a generative simulation of
this so you can better understand it and
then we'll know the ground truth and I
can show you the consequences of
modeling assumptions from here.
Okay. So here's some simple R code. I
apologize that this is not spaced very
well. Uh but I'm going to step through
it in a little bit. Um, we're just going
to simulate 2,000 applicants. And the
first thing we're going to do, you
follow along in the DAG, uh, remember if
you simulate when you to turn a DAG into
a generative model, you have to make a
bunch of distributional assumptions, but
there's an order you simulate the nodes
in, right? You start with nodes that
have no parents, have no arrows into
them, and then you do the ones that only
have one parent that you've already
simulated, and then down the chain.
Yeah, does that make sense? Um so we
start with G uh gender and we just make
a a even gender distribution. We're just
sampling genders one and two uh into um
into the data set into the applications.
Now I'm going to I'm going to simulate
the other node that has no parents which
is the unobserved differences in
ability. And I'm just going to make this
binary in this case just for ease the
conceptual ease of thinking about it.
But you can make it continuous. Um
and uh this is as in the DAG
uncorrelated with gender. Yeah, it's
just a it's just ability. Um
and then we get to simulate department
choices. This has because we've
simulated G and U and those are the
parents uh of the department choices.
I'm going to have two departments just
to keep it keep it cognitively simple,
right? Just two departments. And um I'm
assuming here that gender one tends to
apply to department one. two tends to
apply to two. How conveniently they're
labeled that way, right? It's almost
like I made it up. And uh
but um gender one individuals with a
high ability um tend to apply to two as
well. Yeah. And the the the background
story here in my fake data is that two
is a discriminatory environment for
gender one. And gender one individuals
know this. But when they're high
ability, they're they're willing to
fight, right? Go in there and do it. Uh
dominate those jerks. Yeah. Um the story
may be based on a true story. Yeah.
Okay. Um are you you with me so far?
Yeah. Now, there's a particular function
that I chose to do this, and I'll let
you parse the R code later and play
around with it, but that's that's what
ends up happening. And we've got this
table here that shows you um when U is
zero uh um gender one is in in
department one and gender two in
department two. Yeah. But when U is one
um gender one shifts to department two
for high high ability individuals. Okay.
This is what they apply to. We haven't
accepted anybody yet. This is just
applications. Okay. Uh and then we
finally get to simulate the node that
has three parents. uh acceptances. And
again, the the code's a little opaque
here. I'm just going to do the summary.
Um department 2 discriminates against
gender 1. That's what's going on here.
Okay. But high ability individuals get
accepted at higher rates. And I've made
some matrices to do this and I'll let
you study them uh later and confirm that
this is what they create this effect
that high ability individuals in both
departments are more likely to be
accepted. Um and department 2
discriminates uh against gender gender
one individuals.
Um they have the wrong integer,
right? If they wanted to be accepted,
they should have had it too, right?
Sorry, I joke, but this is, you know,
apply the labels as you like. Yeah. Are
you with me? Yeah. This kind of of
generative simulation, I encourage you
to always do it. Like really, I know it
seems like nobody does this, Richard.
Why are we always doing this? because
it's the only way to debug your thinking
and to justify the statistical analysis
is to have a generative model. Okay?
Otherwise, you're just hoping and
praying and moving your hands rapidly,
which can get you published. Absolutely
can, but that's not good enough, right?
Okay. Sorry. I know you you're used to
my sermons now and you keep coming back,
so obviously they're not that bad. Um,
right. Let's analyze. So, now I've got
simulated data and we can uh fit two
models. The structure of these models
are exactly the models from the previous
lecture. We do the total causal effect
of gender in model one and we do then we
get the direct effect in model two.
Yeah. Do I need to should I review the
structure of these or
yeah the remember the total causal
effect we don't have any adjustment set
because it's not confounded in the DAG.
Um uh for the direct effect we need to
stratify by department to get the direct
effect. And so that's what we're doing
in model two. Uh so the total effect
shows a disadvantage which is true. And
now I'm it's like I'm putting up these
coefficient tables and asking you to
read them, but I'm not. On the next
slide I'm going to plot these. Okay? So
just hang on. Um and then the direct
effect is confounded. And I'm not going
to ask you to read those that matrix. Um
it's uh in in a sense this is just here
for you to feel like that's so
confusing. Yes. If it's confusing to
you, don't put it in your papers like
this either. Right? You've got to plot
this stuff. Yeah, we know the truth here
and we still can't read these tables.
It's very difficult. So, plot them
and um plot the contrast too. I'm just
showing you the marginal posteriors
here, but as you know, you want to look
at specific contrast too as well. And in
this case, these these are not
misleading. Um uh there's a so looking
at the posterior distributions on the
bottom half of this slide. Um blue is uh
department one applications solid uh
densities are gender one and the dash
ones are gender two. So on the far left
we've got um the posterior distribution
of the probability admission for
department one gender one that's the
first one on the far left
and it's low
uh and then um gender two and department
one actually has a slightly higher uh
probability of admission in this sample.
Um
and then there's department two uh where
probabilities of admission are higher
overall and even though we built into it
an discrimination against um gender one
in department two that's the solid curve
there that shows there's no
discrimination showing up in these
estimates
why I'm going to we have multiple slides
coming to explain that is because of the
unobserved confound okay and I'm going
to try to help you understand why it is
masking the discrimination. I've got
multiple slides coming up to show it.
Um,
okay. Uh, and you guessed it if you if
you've been with the class this long.
This is collider bias. Yes. Is this meme
too old? Does anybody know this at all?
Yeah. Okay. This one verse. Yeah. This
this game came out when some of you were
babies. I know, right? But it's
timeless. It's absolutely timeless.
Okay. So, yes. Uh, collider bias. You
guessed it. Uh, when we stratify by
department. So D is
someone's just now understanding what's
going on. This is good. Welcome to San
Andreas. Right. This I think this is the
thing. But um, uh, so come back to the
DAG and I I got to show you the next
slide again. Um, D is a collider on the
path between the treatment gender and
the unobserved confound. If we don't
stratify on it, that's fine. the
unobserved confound does no harm at all.
Right? There's no confounding of the
total treatment effect of gender in that
DAG. But as soon as you stratify by
something a post treatment variable like
department, you open up that confound
path and you bias the estimate. And
that's what's going on. It's because of
collider bias. Um so it's like thing on
a lot of these literatures when there's
cross-sectional mediation analysis. The
thing you can estimate with weak
assumptions is the total effect, but
nobody wants that, right? what you want
is to decompose it and that requires
strong assumptions. But that's okay. The
thing that people want requires strong
assumptions. Show people what those
assumptions are and show the
consequences of them. Yeah, it's it's
okay. Um
uh the estimate is likely to be biased
by a certain amount because we don't we
don't have the use measured. Uh but you
should still report it, I think, and be
transparent about that. And we're going
to do even better. We're going to do
some sensitivity analysis as we get
going. Okay. So here's here's a more
intuitive explanation now that you see
the the DAG D is a collider you see
between G and U. So as soon as you
stratify by D, you open the path through
the through the confound. Yeah, but it's
harmless otherwise. It's totally
harmless otherwise.
Isn't isn't life great? Right. So when
you first took your first stats class,
they probably told you that there's no
problem omitting vari there's no the
threat with causal inference is that you
omit something that you needed to adjust
for and they didn't tell you that adding
things could could induce bias, but it
can, right? That's what colliders are.
You're not safe either way. Yeah, it's
it's just this is why you have to have a
causal model um uh that that justifies
the adjustment set. Okay, here's the
intuitive explanation. High ability
gender-run individuals apply to the
discriminated department anyway because
they're badasses.
Yeah. Yeah, that place sucks. They're
jerks there, but you know, if I if I get
in there, I can make more money,
[snorts] right? So, they're going to
apply anyway. Um G1's in that
department, uh if they manage to to
their higher ability, so they can manage
to overcome the discrimination and get
admitted and then once in of higher
ability than the average uh G2 in that
department.
Right? Because G2's ac department 2 is
accepting G2s without discrimination.
Right? So they're getting low ability
ones. Uh so the discriminated class in
department two is higher ability
because it's the only way they could
exist in that environment. Am I making
sense? Yeah. Again, this may be based on
a true story. Yeah. Inspect your
intuitions about this. Um
uh okay. And and so this is a masking
effect. It's sometimes called that
discrimination can be masked by the
survivorship bias of the ability of
individuals to coexist and thrive in a
discriminatory setting. Like and I could
see that some heads are like yes lived
experience,
right? It's like yes again based on a
true story. I care a lot about this
literature. I've known I have known
people who have gone through stuff like
this. Absolutely. Okay. Oh, I should
pause. Does this make sense? Yeah. You
with me? Um uh I mean on the one hand
this is very frustrating for causal
identification. On the other hand, isn't
this cool that things like this happen,
right? So you can't just you can't just
naively read the effects. You have to
think about things like this possibly
happening. Now I want to return to these
two papers that I mentioned um last
lecture and do a more and I said some
things about them that was critical like
they don't even have clear estims and
there's no causal models which is true
but I want to do a more careful analysis
now for you where I draw DAGs for them
and I explain to you how these papers
can use the same data set and get
different results. Okay, is that going
to be fun? Yes. All right, here we go.
All right. The first thing though I want
to say before I look at these two papers
is here's a great paper on this topic
that has DAGs. Um this is published
somewhere now but you know use the
archive. Be nice to the archive. Um and
uh this is a great paper. It has DAGs.
It it has lots of stuff in it from this
lecture but also other examples of
structural confounding when we're trying
to understand uh bias and and fairness.
I highly recommend it. It's a very clear
paper. Okay.
Um, the first paper is this this
Larbanet one from from 2022. It uses a
sample of members of the National
Academy of Sciences. The National
Academy of Sciences at the USA is some
kind of Masonic organization. I'm not
sure exactly what it is. Sorry. I'm I
used to be American, but I gave up on it
a long time ago. [laughter] And
I was born in Germany. I mean, even
though I have an American, right? So
it's like uh but no it's it's like it
was created to like give the president
advice on matters of science but it
doesn't do that anymore. It's like a
prestige organization for scientists.
Members elect other members and then
they have their own journal called PNAS
they can publish anything they want in.
Yeah I'm I'm only being slightly snide
about this. Right. The variance of PNAS
papers is huge. Yeah. Some of the best
and worst things you will read in any
journal just page after page. Right.
Because there's very little editorial
filter. Um, okay. Uh, so they've their
own they've got a sample. It's just
members of the National Academy of
Sciences, uh, which is a bunch of
different subjects. Um, and historically
there's a big gender imbalance in
members of the National Academy of
Sciences, as in many scientific
professional organizations. And they're
looking at citation networks now. Um,
and they're looking at what is the
effect of gender on citation networks.
And they find that women are are being a
woman is associated with lower lifetime
citation rates. Okay, that's the first
paper with me. The second one looks at
elections to the National Academy of
Sciences. This is the same data set now,
but we're not conditioning now on being
a member.
Yeah, we're looking at a larger
population. Uh uh in this case, women
are associated with three to 15 time
advantage being elected to the National
Academy. Controlling for citations.
This is going to be important in a
moment, but so far it's just to
understand the difference between these
analyses and what they've done with the
data sets. Yeah, this makes sense. Um,
they control for citation rates. So
that's like saying we stratify by
citations. So, uh, men and women who
have similar citation rates, women are
more likely to be elected to the NAS is
what they find.
Okay, you with me? All right. Um, I'm
gonna have fun with this. I hope you do,
too. Okay, so now let's let's talk about
let's dag this out. This is the data set
in my mind. U we're looking at gender uh
gender influences citations uh maybe
through that that would be a
discrimination effect here or a social
network effect. It may not be
discrimination but it's structural.
Yeah. Depends upon uh which field you're
in. Some fields site a lot. Some fields
site very little. Right. Like in the
humanities there's very little citation.
That's not a criticism. It's just
there's there's less. Yeah. and the
humanities person is nodding. Yeah,
there's just less citations. It's not
like in psychology, citations are half
the paper. Yeah, there's that paragraph
in the intro of every social psychology
paper which is just citations that has
been forced to put there by some editor
or something. We don't have to do that
in the humanities. It's just not like
that. Yeah. Um so that's citations. Uh
gender may have an effect on becoming an
NAS member. That would be direct taste
based discrimination. Um, and what's
hidden here again is likely uh some
number of variables. I'm just going to
call it quality, which influence both
citations. If you make good papers, they
hopefully get cited more. It's not
always true, but you would hope it's
true. There are other ways to get cited.
Um, and if you're a higher quality
scientist, uh, uh, hopefully you also
more likely to get elected. You with me?
All right. So, this is like what was
before. Now, let's think about the first
paper. We're going to restrict the
sample just to people who are members of
the national academy. So we're
conditioning on M on being a member,
right? We've we've it means we've take
if you had the full population, you
throw away all the people who aren't
members. That's like conditioning
stratifying by that, right? But it's
it's like you've made that decision by
throwing away data. Yeah. Question. I'm
not really trying to
>> um in in the first paper uh that that's
the next paper. This paper didn't do
that. This paper looked at citation
rates. The first one looks at citation
rates. The next one's going to look at
who gets elected, right? But it's the
same data set, remember? Yeah. Okay. So,
this paper only looks at members. So,
they implicitly are conditioning on
membership. Yeah. So, it's like they
stratify by membership. Um so what does
this mean? Well um member is a collider
between Q and gender.
When you stratify by it you open the
biasing path. Do you see it?
You having a good time that people are
like no why did I come to this lecture
today? No. No. Yeah. But you see it
right when you this is one of these
things the kind of conditioning that
happens. you didn't put you they didn't
think they put member in their model but
they did because that's how they
constructed the data set right they the
whole data set is stratified by
membership does that make sense and so
if there's if it's a collider on some
path you activate it as a collider
yeah doesn't have to be in the
regression model yeah you can do it
outside the regression model if you just
if you just uh construct the data set so
it's only members um and that means that
we've got some biasing quite likely So
now think let's do the thought
experiment. Remember this um this paper
finds that uh uh uh men were cited more
conditional on membership. So uh if men
are less likely to be elected which is
what the other paper finds uh they must
have higher quality or citations to
compensate.
Yeah. Because we've stratified by
membership. So there's this finding out
issue. It's like, well, okay, there's a
bunch of people who've been admitted um
uh that if they vary in in quality, now
we're going to uh uh in the other paper,
if the other paper is true, uh that men
are less likely to be elected. Um but
now we're looking at individuals who
have similar citation rates, men have to
if men are going to be elected to
overcome
the the women's advantage to get
elected, they're going to have to have
higher citations. And so once we've
selected on a sample of members, lo and
behold, they find that that men get
cited more.
Do you understand? Now, this is
conditioning on the other paper being
correct, which I really don't want to
do. [snorts] I'm not sure that either of
them is, but I've tried to spell out to
you how the logic of these two papers
does not live together. Yeah. Does this
make sense? This is the way that and the
first slide of this lecture where I
taught you to be a a a gorilla warrior.
This is what I mean. When you read these
things, you want to start diagramming
like what have they done and what's the
sample and what's the DAG and what's the
workflow. Yeah. Then you can have an
informed opinion about these things.
Okay. The next one, DAG's the same same
world, right? Same sample. Um they don't
condition on membership. They look at uh
the membership as the outcome, but they
stratify by citations, which is a post
treatment variable. Remember
conditioning on post- treatment
variables is not that you should never
do it but there are risks. Yeah.
Anything downstream of the treatment uh
can either block a it will either block
a causal path which in this case we do
want to do that because they want to
look at direct discrimination or it can
be a collider on some path with an
unobserved confound as it is in this
case. So when we stratify by citations
we open up the biasing path through
quality again. Yeah, both C and M are
colliders on the path between gender and
quality. Yeah. So in this case now G is
the treatment. C is a post treatment
variable. If women are less likely to be
cited, which is what the other paper
finds, uh then then women are more
likely to be elected because they have
higher Q than indicated by their
citations.
Right? So it's not that there's any
discrimination going on in the
elections. If we can if we think the
other paper's right, this paper finds a
biased result. It assume it it shows
that women are more likely to get
elected, but that's because they've
overcome discrimination, right? They're
higher quality, and that's the hidden
variable. Yeah. They're being elected
because they deserve it. Yes. It's not
positive discrimination, which is what
people interpreted the paper as evidence
of. Does that make sense? Yeah. Just
like you think of think of a uh
individuals, male scientists, female
scientists, they have the same lifetime
citation rates. Um if women are getting
elected more often, it's because of the
the the hidden path here, the quality
variables. Yeah, that's or that could be
the explanation, right? It doesn't have
to be the direct effect. It could be the
biasing path. Does that make sense? Are
you feeling good about this? Yeah. Now,
okay. Now this may seem like some
fanciful thing and I have dissected
these two poor papers. I'm sorry but
it's making you feel good. I know this
this theme but no this is I mean this is
part I try to choose examples in this
course which matter right and and things
matter and this kind of structural
diagram the simple mediation path is
what I'd call it is incredibly common
right and people are often doing
mediation analyses where they condition
on post treatment variables and it's it
always requires strong assumptions yeah
even if you've done an experiment where
you randomize the equivalent of gender
right we're not going to randomize
gender here but if you did an experiment
where gender with some X some treatment
and you can randomize it, you still have
this problem because you haven't
randomized anything downstream of it,
right? So you start conditioning on
things downstream of it. The bias will
still appear. Good times, right? Yeah,
sometimes. Sorry, there's some people
are not getting less and less happy with
this. No, but this is going to be a
happy ending. I promise you question.
>> So
like
assumptions.
>> Yeah.
these papers.
>> Well, uh the question was if they assume
quality is the same for men and women.
We're assuming quality is the same for
men and women. You still get a bias. You
guys notice there's no gender doesn't
point into quality.
Yeah, that that's not required. uh that
could make it worse, right? If you want
to point gender inequality, then it
could be even worse. But you don't have
to assume that to get this biasing
effect. Um the papers don't talk about
these things at all. They have no
structural diagrams, no identification
strategy. They just start regressing.
Yeah. And then it's kind of the force of
you you know how these papers are. It's
the force of the rhetoric. You get swept
along in the narrative, right? A
well-written scientific paper can sell
you on a lot of results, right? And it's
like the methods are maybe not even in
the paper. It could be online somewhere.
Yeah. It's it's this is the the system
I'm criticizing. You know, you know my
rants. Yeah. Question
>> maybe. Is it the same as what we talked
about before with the restaurants in
Paris
>> that like in order to get into science,
you need to either have a lot of
citations or have high quality.
>> Yes.
>> And those are actually not correlated.
as a big cloud, but we're kind of
slicing off the top right corner. Yeah.
>> And by doing that, if you take a circle
and you slice it, you actually do get a
>> you get a negative correlation.
>> Makes it sound makes it seem like you
have a correlation, but they're not
correlated in gender. They're not
correlated in the general population,
but they are maybe in the NAS
population.
>> Yes, that's right. That's exactly the
argument. And the microphone didn't pick
that up. This was a good question
relating this back to my restaurant
example when I introduced collider bias.
It's the same phenomenon. It's
absolutely the same phenomenon. Yeah,
there are multiple ways for the thing to
happen. Um they're causally independent.
Once you stratify by the collider,
you've induced a correlation. Yeah.
This is the classic case of collider
bias 2 where the different causes
compensate for one another. Exactly as
you say, there's multiple ways to get
admitted. Yeah. That's exactly what what
goes on. Good. Another question.
>> Always like something unmeasured as a
collider or something like that. We
try to surround for example and we
always try to look at some particular
population of of people
>> for example here numbers and they are
and
>> what to do is this the question
>> it feels like it's impossible to
>> it feels like it so I I think there's
lots to do I think this is very
constructive you you start drawing out
assumptions like this and you can
develop it into a workflow so what I'm
going to continue with in this lecture
is sensitivity analysis where we inspect
the the cues and we keep them in the
model when we do the analysis even
though we haven't measured them and
you're going to say how can I do that
well I'm a basian I can do anything
[laughter]
um no probability theory can right
basians just obey the axioms remember
but uh uh another thing to do and I
think this is your question is you say
every sample is conditioned on something
right exactly we'll draw it in the DAG
how do you get into the sample
make that a node. Uh, and people do
this. So, people who work on on causal
inference of surveys do exactly this.
Um, who gets into the survey? And
there's lots of self- selection in
surveys and lots of non-random
non-response bias. Uh, so political
polling is famous for this. Uh, people's
partisanship affects their probability
of responding to polls. Um, and then
that's a very that's a that's a problem.
But you can draw it in the DAG. You
absolutely can. So in this case it's
it's like how do you get to be a member
and even if your sample even if you only
have members as your data set you can
still draw M on the DAG and analyze it
as a collider and understand the
consequences or what assumptions are
required. Yeah. In order to think that
your estimates are unbiased. So you do
that social network um causal inference
in social networks is similar. You're
looking at diads, uh, friendships,
marriages, uh, other kinds of
relationships. Uh, the the processes
that form diads are correlated with
features of the individuals, right? So,
you need to have a you can dag out how
individuals become friends and how that
is caused partly by their preferences,
right? So the classic example when I
when I took grad stats at UCLA the
example they used was uh smoking um in
married couples smoking is highly
correlated. [laughter]
Yeah. And you can think why. Yeah. Now
many places smoking has declined but not
in Germany right. It's still I'm picking
on it but it has declined in Germany but
it's increasing in teenagers now for
some reason. And uh so preferences like
that um can actually be causal rights
about about these friendships and so or
relationships and that affects
persistence of diads but you're
conditioning implicitly on on friendship
formation when you do that. You don't
have a data data set of all the people
who could be friends. You see so you're
implicitly conditioned on diads already.
Right? So any any literature that
analyzes married couples has conditioned
on getting married and persisting in
marriage. Yeah. And those are causal
processes that depend upon the features
of the individuals. And so you may have
colliders. And so when I was taught
again in grad sets I was taught that
using smoking as an example and how that
can screw up network estimates. Um okay
I could go in I should stop. I have
endless examples. Yeah. Uh these things
but you learn these structures and
you'll start to see them in other data
sets. Um, my default explanation
whenever there's a hyped abstract, by
the way, is try and figure out how it
could be collider bias. It's usually not
hard. Yeah, you get practiced at this.
And it's not a cynical exercise. It's
self-defense. Yeah, you have to remember
that. Okay, this is pretty real. I I
lost the flow there, but there were some
great questions and I appreciate that.
You should always interrupt me with
questions if you want. Um but I was I
was going to flow into this this real
world example where um uh companies are
actually making decisions I think on the
basis of collider bias. Yeah. So I can't
prove it but here's a case that was uh
puzzled many many people when it came
out. So Google uh got sued for gender
discrimination in pay and as part of
that they did a big internal audit uh
all their employee almost all their
employees like plus 90% of them um of
their of their job and their experience
level and their pay to see if their how
much gender discrimination there was and
they adjusted pay in much of levels in
the highest level of engineer I think
it's called level four uh it's on this
slide somewhere yeah level four software
engineer
women were actually getting paid more
than men in that level. So they paid the
men more,
right? Because they said that's
discrimination.
I hope you are all primed now to think
about how that could be something else
that it could be that those women
working in a discrimin discriminatory
occupation uh and a company that had
already been sued multiple times for
gender discrimination might have been
paid more at that level because they're
better.
Yeah. And then paying the men the same
as them is even more unfair after that.
Are you angry with me now? I'm sorry
that was a little too much about me.
[laughter] But uh no, but just do the
analysis, right? This kind of like, oh,
we just look at the differences and
interpret it as discrimination. Should
never do that, right? It's very
structurally difficult in a
cross-sectional sample to make causal
inferences of that kind. And in this
case because of of the masking of
discrimination
um by hidden quality variables you can
you can uh this anyway this result is
consistent with that. Yeah. Um there are
other explanations as well like um uh in
in discriminatory environments men are
put hired at higher levels uh but have
lower ability and that'll generate the
same thing. it'll still be an ability
difference and that explains why the
women are getting paid more, but the the
male engineers were just put at level
four quicker. Yeah, that would also be
an explanation. Um, okay, sorry that was
my rant, but I wanted to make it real
for you. Yeah, it's quite difficult uh
in in these sorts of organizational data
to figure out exactly what's going on.
And remember, in a discriminatory
environment because of hidden confounds
and colliders, uh the discriminated
class may actually look like they're
overperforming.
Right? Because it's survivorship bias.
Does that make sense? I'm going to say
this over and over again. Sorry, I'm
like really obsessed with this result,
but this is where causal inference pays
for itself. Yeah. Okay, I've got 20
minutes. Um, and I want to talk about
sensitivity analysis. So, let me try to
summarize. So, I keep saying if you
don't put any causes in, you're not
getting any causes out. You've got to
have a generative model. Um, what we
have here are hike papers with vague
estims.
uh um unjustified adjustment sets. Um
and in the case of that Google example,
you're even getting policy design
through collider bias. I think it's very
plausible that that's what's going on.
Uh we can obviously do better about
this. This is a serious issue. Um uh all
kinds of ways, discrimination, pay
fairness, uh uh policing, uh is another
example. I've actually Oh, the next
slide is actually about policing. Sorry.
Um and but you're going to have to make
strong assumptions. There's just no way
out of that. And that's okay. strong
assumptions are how you license strong
inference. As long as you're transparent
about it and you're not hiding the
assumptions, then you've done nothing
wrong. Last line here where I I wave my
um anthropologist flag, qualitative data
is super useful in these contexts.
Quantitative data is not always better.
Uh quantitative data in these cases is
highly conounded. But when you get when
you talk to people, you get their
testimonies and the detailed mechanistic
testimonies of their experience of
discrimination that really helps you
understand the qu the quantitative data
and develop better generative models,
right? So those interview data, the soft
humanity stuff that people like me do,
like it's really important input into
causal modeling. Okay, sorry that's my
anthropologist flag. Uh people at home
can't see me, I'm waving an imaginary
flag right now. Um it's why we do
ethnography and why we talk to people.
It's not stupid to talk to people,
right? Okay. Sorry. Um, I work at an
institute with a bunch of molecular
biologists who don't respect talking to
people, so I have to say these things
sometimes.
Okay.
Uh, I have worked with some of my
colleagues like Cody Ross. Some of you
know Cody. Cody and I have have a few
papers where we apply this logic to the
literature on discrimination by police.
And um this is most famous in the US,
but it's not unheard of uh in Germany.
German police can be a bit difficult too
at times. Um some nodding heads out
there. Yeah. Um okay. So, uh I just want
to recommend to you that or point out to
you that this kind of literature has
some similar issues. Um often you have
these database called administrative
databases provided by the police. So
right away you're suspicious how good it
is, but they're conditioned on being
stopped, right? This is most of the data
set is their police stops and there
these official records of when the
police actually stop someone like it's a
traffic stop or something like that or
someone's acting suspicious in the park
and someone calls the police on them or
something like that. That's a stop. Um
and so you're the sample is conditioned
on being stopped by the police and uh
yes you can believe it that there is
probably things like acting suspicious
um that is an unmeasured confound in
these things that biases estimates.
There's a great paper on this which is
the citation at the bottom um Knox and
and Mumalo. Uh this is a great paper.
It's got DAGs. It really walks it out.
What's an identification strategy? And
they show actually with very reasonable
assumptions you can still do
identification with these data sets but
you have to expect them to be biased by
the fact that they're selected on being
stopped. Does this make sense? But it's
a hopeful paper, right? This is not
nihilism. All of us want to solve these
problems because they matter and we work
hard on it. Okay. Sorry, my rant just
continues. Uh we're going to get through
this the sensitivity analysis. Um okay,
so we let's here's what I want to do
now. a sensitivity analysis on the
unobserved confound. Um, well, you could
say, what are the implications of what
we don't know? We're going to assume
this confound exists, this quality
confound, whatever you want little U to
be. Yeah, it could also be
attractiveness, right? There's a lot of
evidence in these literatures that
attractive people get stuff in the
world. So, it could just be random
attractiveness, right? That has an
effect, too. Um, and we can interpret
that as discrimination if you like, as
well. So it doesn't have to be quality
differences, right? When someone who's
dumb doesn't get a job, you don't call
that discrimination, right? But if
someone's unattractive, they don't get a
job, we would. Yeah. I used to teach a
class to undergrads back in my previous
job where we spent like a whole week on
discrimination and talked about these
things. And some forms are very
upsetting to people and some forms are
not. Yeah. Tall people get stuff too,
right? Why do we discriminate against
short people? There's lots of horrible
things people do to one another. I
should stop talking, right? But no, I I
love this topic. I used to teach it
quite a lot. So, um I mean I hate this
topic that it exists, but you know what
I mean. Sorry again. I have to choose my
words more carefully. I got to get if I
were a politician, my career would be
over. Uh so, we're going to assume this
compound exists and model it
consequences. That's the transparent
comprehensible thing to do. And we can
vary the strength of the confound, make
different assumptions about how it
affects the structural model and see
what the consequences are. Uh now that
doesn't mean we we can't figure out what
the truth is, but we can learn what the
consequences of our assumptions are in
the truth and we can calibrate our
uncertainty, right, about how strong the
confounding would need to be, for
example, in order to explain away the
effect. Yeah.
Or create an effect. Does this make
sense? Um, and you already have all the
tools to do this and I've taught them to
you in this class and let me show you
how. Okay, we're going to take our DAG
and we're going to do what we always do
with our DAG. We're going to make it
into a generative model. So, the first
thing is the uh outcome model for A.
This is just a logistic regression. I
taught you this in the previous
lectures. Yeah. Um,
and um what's different here is that
I've taken you and I've put it in the
model. Do you see that there's a little
U subi for each individual eye?
But you're thinking, "But we haven't
measured that. How is it in the model?"
Oh, we're basians.
We can put unmeasured things in models.
They're called parameters. You've been
doing it all along.
There's nothing different here. A
missing value is is a parameter.
That's all it is. Yeah, a prediction is
a parameter. You don't have any problem
with that, right? It's just a missing
value and you're going to use the model
to predict it. So, the same thing here.
Um and then there's an outcome model for
department choice that applies
simultaneously. So I create this
indicator variable when department for
individual I equals two then it takes a
value one otherwise it's a zero. So it's
another logistic regression right we're
modeling the probability of applying to
department two which remember is the
discriminatory environment in my
simulation
with me. Um and again the U appears.
It's in both equations because uh the
quality affects both, right? It has an
there's an arrow into department choice
and there's an arrow into uh uh
acceptance. Notice that the beta which
is the effect of of hidden quality on
acceptance can vary by gender, right?
It's subset by gender. So it's
interacted. This is an interaction
effect if you want to think about it
that way. And the gamma, that squiggle
is called gamma for those of you who
don't read Greek, right? It's one of my
favorite Greek letters. It looks really
nice and um it's like artistic Y. It's a
really nice letter. And uh it's also
subset by gender, so we can vary it.
Yeah. Um and that good. Yeah. You with
me so far? Um and then we need prior
obviously we need prior and we're going
to put a prior on the use. And I'm going
to put normal 01 prior on the use which
just establishes a measurement scale for
them. Right? This is a latent unmeasured
variable. The scale can be anything. And
I'm just saying they're Gaussian with a
variance of one. Yeah. If the variance
were bigger, we could rescale it. Right?
It's a latent variable. The v the
variance is arbitrary. Does that make
sense? Um what you need to do uh to make
that work though uh so that the the plus
plus values values of u greater than
zero mean greater chance of acceptance
is to make the um uh the the betas and
gamas positive right because that sets
the the veilance of what higher value of
u means yeah so we do that and in this
example at the very top let me show you
what I'm doing here I'm just putting in
the the betas and gamas called b and g
as data. This is the the really rigid
way to do it. There's nothing wrong with
this is we're going to say let's assume
the effects are of a certain size,
right? This lets you test the effect of
the confounding scenario. Um so in this
case we assume that high ability means a
higher acceptance rate. Uh those are the
betas, right? Those are the that's that
vector of one and one, right? So the one
means a larger value of you makes it
more likely to be accepted. You're more
likely to be accepted. And then the the
G's are one and zero. Um G1's with high
ability are more likely to apply to to
D2. So this is the effect of your
ability on applying to department two.
And so gender one gets a one and gender
two gets a zero.
Does that make sense? And you can change
these, right? And that's the idea is you
vary these inputs and you see what the
consequences are. And I'm going to show
you an example in a moment. Um and so
here you see the there's the U um uh
there in the model.
Um, and then down here we've got uh U
subi uh again for the department model.
And and here's the vector of length in
one for each application of use. And
we're going to estimate them. And you
can you can get posterior distributions
for all these. Yeah. Are you with me? Is
this fun? I know this is like basian
madness, but it's licensed by the causal
graph. Yeah. It's totally licensed by
the causal graph. And since you've got
the synthetic data simulation, you can
verify this this works. It reveals the
truth in particular scenarios. Um, this
is one of my favorite things in classes
like this is we get to have there are
2006 parameters in this model and 2,000
observations. And at some point you took
a baby stats course and you were told
this is illegal. It is not illegal. You
can have more parameters than data.
That's how neural networks work, right?
All those chat bots which are ruining
your life, right? [laughter]
the many many they have billions of
parameters, right? It's just how it
goes. Of course, they're also trained on
billions of blog posts unfortunately,
but anyway, but this is totally fine.
You could get the posterior distribution
in these scenarios. Um, okay, let me
show you what happens here. I've only
got eight minutes left here. Um
uh on the top it this is just the graph
from previously where we ignored the
confound here the marginal posterior
distributions where discrimination of
gender one in department 2 is masked by
their quality differences. Yeah. And
then in the bottom uh where we assume
that that's what's going on which is the
true model
you see the discrimination pops out.
Right. In a real data set, you're not
going to know what the ground truth is,
but you can show the consequences of the
assumptions at least. And you can vary
the strength of the assumed effects of
the discriminatory effect um to find out
how big it would need to be to reverse a
trend or or whatever it is that you're
analyzing. Yeah. So that's what the
sensitivity analysis would be. Um you
can also put priors on those things. So
here's another version of this analysis
by this script is in on the website or
actually I don't think I've uploaded it
yet. it will be on the website.
[laughter]
Okay, I'll put it in the scripts folder.
Um, so you don't have to type this from
the slides. Uh, here I'm not going to
enter the the B's and G's as data. I'm
just going to put prior on them. And I'm
going to put these uniform 01 prior on
them. Again, we need these to be
positive. Otherwise, the directionality
of the use doesn't make any sense. Yeah,
there's a question.
>> Could you in theory put
B and G or GMA as the same? Yes. Yeah.
Yeah. You could you could do any
scenario you want here.
>> Absolutely.
>> It should be the same.
>> But the G's are the G's or chances
affect the probability of applying to
department two when you're high quality.
So they have you can do different
scenarios. And in the original scenario,
it was only department two is
discriminatory against gender 1. So
quality only affects gender one's
probability of applying. But it could
also partly affect gend um uh gender 2.
Yeah. Right. It's just higher quality
individuals more likely to apply to
department two. Yeah, you can try
different scenarios. Here I'm just going
to put these these kinds of default
vague priors on it. Um not as a
recommendation for these priors, but uh
here I am making them the same for both
genders just just to show you what
what's possible. Um and yeah, I should
have highlighted this already. That's
where they are. And then when you do the
estimate, what the model tells you and
this is super useful is you can't say
anything,
right? the these data if you cannot pin
down the values of those of those
effects are consistent with a wide range
of things. They're consistent with lack
of discrimination. They're consistent
with quite strong discrimination.
>> Yeah. Does that make sense? Question.
>> So changing scenarios based on
camera.
>> Yeah.
>> Is what we call sensitivity analysis.
>> All of this is sensitivity analysis.
Sensitivity analysis uh I'll come back
to this in the next slides is just
changing any of the assumptions. It
could be the data model. It could be the
existence of confounds. It could be the
priors and seeing what the consequences
are for inference. That's sensitivity
analysis. In this particular case, the
goal of the sensitivity analysis is to
show what the impact is of assuming
there's an unobserved confound. Uh and
then we can try different structural
scenarios about the effect of that
unobserved compound. Yeah. Does that
make sense? This is uh this is not my
private madness. You see this stuff in
journals. I'm telling you like
increasingly common because there are
packages now that automate it for you.
You can give it uh you can give a
multiple regression to the package and
tell it structurally where the which
variable should be correlated because of
an unobserved common cause and it will
make a plot for you even in some cases
of how big the confound needs to be in
order to um uh remove the effect you've
observed. Yeah. which is useful. So
that's nice. It's like a maturing of a
scientific field. Stage one, deny
compounds exist. Stage two, admit they
exist and do nothing about it. Stage
three, simulate the consequences of
compounds existing and report it. Yeah.
Does that sound good? It's like therapy,
right? So we're moving up to these these
situations.
Um, good. You ready? Okay. So yeah, this
is what I was going to summarize for
you. Sensitivity analysis is about um
quantitatively figuring out the
implications of of structural
assumptions. In this case, things you
don't know like the strength of the
unobserved confound, but you have good
reason to suspect it exists and you
probably cannot convince a reasonable
skeptic in your own field that it
doesn't exist. That's often the case. Um
that said, when we say there's some
confound, you could be applying this
critique to someone else's work. uh it's
our responsibility to really model it,
not to just wave our hands and try to
dismiss a result because there might be
a confound. Uh both because that's rude,
right? I mean, criticism should should
involve some effort. Uh and second, some
things even if they're biased, need to
be reported. You have to take action
sometimes, but you want to be you want
to really math out um what what the
confounds might be doing and what the
what the errors might be when you take
action. This is somewhere between
simulation and analysis I say. Uh but
the way I've been teaching analysis,
there's a bunch of simulations. So maybe
you didn't notice, right? It just feels
like the rest of the stuff I've done. Um
so yeah, I've showed you a case where we
um we we vary the strength of things
directly and I gave you a case with
priors. You can also vary the other
effects uh uh make assumptions about the
other regression coefficients like the
uh admissions rates because it's all
coexists in the same model and then
infer the discrimination effects, right?
But you're going to have to make
assumptions someplace in order to do
inference. And again, that's not bad.
Assumptions are how we buy inference. Um
okay, good. Um I wanted to end uh again
on on the guerilla workflow. Um the
scientific literature is a powerful foe.
Yeah. You do not have the firepower to
directly confront it. It's too vast and
complicated. So you have to be very
strategic about how you read, defend
yourself against false beliefs. Um and
I've tried in this course to give you
tools to do that. Uh the workflow is not
some toy thing. It's it's a technique
for dissecting papers that have been
published. Um for doing better work
yourself and for uh resisting the the
sometimes downward pressure from
supervisors and senior colleagues to do
bad work. Yeah. It has to be justified
and comprehensible. Yeah, that's what I
mean. Um and this little I mean this is
I tried to make this what I was thinking
about when I was writing this lecture
like what's the most minimal workflow
diagram. It almost looks like a rune,
right? And that's this is what you can
remember is questions, generative
models, justifying statistical models,
then test. Yeah, test that that
statistical model um is a is a logical
outcome for the question and uh uh and
generative model and test the
implications of the priors, right?
Because the priors are always a feature
of the statistical model that isn't
present in the generative model. nature
doesn't have prior. Can I say that?
Yeah, there must be a basian somewhere
who will not agree with me on that. They
will tell me later when they listen to
this. But I don't think nature has
prior. So, but we have priors. We need
them to do estimation. You have to have
them. So, you should test the
implications of your priors when you
design your stat model. So, that all
that goes in that first test swoop.
Yeah. I need some I need the workshop
vocabulary here, but you know what I
mean. Yeah. Uh and then we get
estimates. We test again that the
machine worked. chain diagnostics, etc.
Um, should also be testing your data, by
the way, against the data dictionary. I
should put another test on here, right?
Are the values in the data set what
they're supposed to be? Every time I've
done that test, there are things that
are illegal, like children older than
their parents and demographic data sets.
This happens routinely. Yeah, that's
what I mean. Test the data dictionary.
You define the data dictionary. You've
got constraints not only on each column,
but on relationships between columns.
Run those tests. That's what we do in my
department. We're obsessive about this,
right? because we find mistakes all the
time. Genealogy data data. Oh my god,
like the mistakes in genealogies. Um,
impossible human families all the time.
Okay, I should stop talking telling
these stories. I only usually only tell
these stories to my therapist, right?
So, okay. So, we we go and then finally
predictions. We want to test the
predictions that we computed them
correctly, but also that they they're
appropriate to the estim.
Are they actually answering the original
question? Okay. Um, I'm out of time. Uh,
I've had a lot of fun teaching this and
I hope you've learned something of value
and I wish you a good rest of the year.