Submind YouTube summaries
Thumbnail for Statistical Rethinking Lecture A10 - Hidden Confounds & Sensitivity Analysis

Statistical Rethinking Lecture A10 - Hidden Confounds & Sensitivity Analysis

Watch on YouTube

Video summary

This lecture emphasizes that scientific literature is often flawed rather than a collection of facts, advocating for an aggressive approach known as "guerrilla warfare" against false beliefs by rigorously interrogating workflows and refusing to accept abstracts at face value. A central theme is the necessity of sensitivity analysis, which involves varying elements within generative models or statistical frameworks to observe how results shift under different causal assumptions; this process highlights that inference relies on strong assumptions rather than being inherently flawed. By treating these assumptions as testable components rather than fixed truths, researchers can better understand the robustness of their conclusions and avoid dismissing confounds simply because they are unobserved. The instructional content illustrates these concepts through several critical examples, starting with gender discrimination in UC Berkeley admissions where an unmeasured variable like "ability" acts as a hidden confound that creates survivorship bias by masking direct discrimination against women who survive harsher selection processes. The lecture further analyzes conflicting studies on National Academy of Sciences membership to demonstrate collider bias, showing how conditioning on post-treatment variables such as current membership or citation rates can induce spurious correlations that misrepresent the true effects of gender and quality. These scenarios reveal that without accounting for hidden factors like unmeasured ability or field-specific norms, researchers risk drawing misleading conclusions about direct effects versus structural advantages in real-world contexts ranging from pay equity audits to smoking habits among married couples. To address these challenges, the speaker introduces a Bayesian framework using latent variables to model unobserved confounders alongside observed outcomes, allowing for sensitivity analysis where the strength of hidden effects is varied through fixed values or priors to assess how much bias would be required to explain away an effect. This quantitative approach moves beyond merely identifying potential biases to testing their structural implications and determining whether results are robust or driven by unmeasured factors. The lecture concludes with a comprehensive "guerilla workflow" for maintaining scientific rigor, which involves formulating questions via generative models, justifying statistical assumptions including priors, running diagnostics against data dictionaries, and incorporating qualitative ethnographic data to inform better quantitative models when confounding is unavoidable.
Read the full video transcript
Good morning. Welcome back to the last lecture of section A of statistical rethinking 2026. Um the way I planned this course this year is uh section A starts at the very beginning and goes all the way up to multi-level models. You're not going to get a lecture on multi-level models because I've already recorded one for the other section. You just have to start their lectures. Okay, I apologize. It is I'm being one sense I'm being lazy by not giving you that but in the other sense I wrote a whole new lecture for you. So here it is uh and I thought this is a topic that has more value added um because there there's tons of existing lectures and whole books about multi-level modeling. There's way less material whether lectures or books about dealing with hidden confounds and sensitivity analysis. But this is not an obscure problem. This is the kind of problem that is routine. Yeah. Like missing data again considered an advanced topic. Um most data are missing. Yeah. Most of the time, right? So it's it's actually a core statistical skill. Uh I'm trying to work those kinds of things more forward in the course, but uh it's difficult given the computational aspects of them. But you are all ready now in this section uh to talk about the serious truth about hidden compounds. and and I want to give you um some more reasoning with DAGs about hidden compounds and the effects they can have and show you one way to do sensitivity analysis so you don't have to pretend these things don't exist. Yeah. Um let me start though with a little bit of of overview. Uh the scientific literature is not a pile of facts, right? uh when you read papers in your own area where you have the expertise to to judge them critically. Are they all good? Do they all make sense? No, they don't. Right? Those method sections, it's like that meme where the with the dirty home, I forget the cartoon characters and the one is like you live like this, right? It's like just it's bad, right? So when you read things in and in in other fields where you don't have the expertise, I would invite you to carry that skepticism with you. uh and uh you can do better than that in the sense that that the tools I'm teaching you in this course empower you to uh fight the scientific literature. It is an overwhelming enemy and you are a gorilla fighter, right? That you you are you are outmatched. You are not going to win this battle through force. Uh you have to apply your force very specifically and strategically, right? Gera warfare, guerrilla warfare is intellectual. Yeah. That's how you overthrow the occupying forces, right? Not a head-on assault. Uh so there's there's all this problems with the scientific literature, but if you interrogate the workflow, the papers you read, try to fill in the pieces of it, uh you can pull together what is the logical justification for this analysis, if there is one or what are the gaps. Uh if you're a reviewer, it's a way to review and ask for clarifying information. If you're reading something that's been published, it's a way for you to defend yourself against false beliefs. Yeah, I mean this. It's really, really essential. Do not believe the abstract ever. Never ever believe an abstract. Right? What's that abstract there? The abstract is propaganda. Sorry though, this is like this is I'm going to run for run for office. You can feel it, right? Um uh now, of course, all this applies to your own work too. Uh be the change you want to see in others, right? Have the workflow reported in a structured way so that people can uh so that you respect the rights of your colleagues to make up their own minds about your work. Make your methods comprehensible and insist that the published literature be that and if it isn't then doubt it. Yeah. Okay. That's my sermon. Thank you very much. Along that line, one aspect of um a comprehensible workflow is sensitivity analysis. Uh so this can mean lots of different things. you can change the elements of the uh of the workflow and see how the results change. So often this means changing the generative model that would be like the causal assumptions that it could be a structural diagram or it could be other aspects of the causal model like the functional responses of one variable on another. Um and also aspects of the statistical model which are not always part of the generative model, right? Like ways you handle missing data. uh Gaussian processes are not a causal model but they're a way to estimate functions. Yeah, I mean you have different strategies for that. Um and then there's so we can we can vary as I say the G's and S's in this diagram. I would not normally label all this stuff but this is a test of your you've been here a long time you can guess what Q is. That's your question. Yeah. G is a generative model. S is a statistical model. E is estimates. P are predictions. D is data. And the idea is we test all these things. And if you're not testing, if you test nothing, you miss everything, right? You have to test. Um, and sensitivity is when we vary these elements and then we uh repeat the workflow and we see how the results change. Yeah. And that's that's often a good thing to do. If you have questions about the about the strength of your assumptions uh uh the difference they make uh in the results, then sensitivity testing is is is good. It's a reasonable thing to do. But sensitivity is not bad. And we're going to talk about this, right? It the fact that the results change when you change the assumptions is normal. Yeah, we we buy inference with assumptions. So just because the something the results are sensitive uh to changes in the assumptions is not bad in and of itself at all. If you can license the assumptions, then you get stronger inferences. Yeah. Uh just think about it this way. Assume the wrong causal model, assume the right causal model. results will be sensitive to those differences, right? But it's clear that one of them is the right answer, right? When you simulate the data, uh when you don't know the right answer, then you have to think about what external information to the sample is going to help me choose between these these sets of assumptions. Does this make some sense what I'm trying to say? I know I've said these things before. I have I'm workshopping different ways to say this, but um yeah, we we here I say we buy inference with strong assumptions. There's nothing wrong with assumptions. It's the only way you can get strong inferences. Yeah. As I said before, uh inference without assumptions like an opinion without reasons. You shouldn't pay any attention to it. Yeah, that's the way epistemology works in the sciences. So, it's fine. But sensitivity analysis is useful in this regard. Okay. Along those lines, let's think about um some assumptions and develop a sensitivity analysis to the consequences of them. So, in the last lecture, I was going on and on in a bit of a rant. I apologize. Actually, I don't about gender discrimination, other forms of discrimination, and how we how it's analyzed in cross-sectional samples. This is a a topic I've published on. I care a lot about it. Um, and I'm going to come back to this example again, the UC Berkeley data set, just to remind you. Uh, we we're looking at a very simplified version of this where we've got gender here in this DAG as the exposure of interest. We want to know if there's direct gender discrimination um conducted by admissions officers uh in this data set. Um individuals apply to departments and that is partly influenced by gender. They're very gender patterned choices of which departments to apply to. And then the outcome here is admission. And there could be a direct effect of gender which we'd interpret as as direct or taste based discrimination. And there can be indirect u discrimination because some departments just have lower rates of admission than others. Yeah. And I tried to convince you in the last lecture that both of these things are probably going on in this data set. Yeah. As in many data sets of this kind. Um now what I want to convince and but all in that lecture we ignored um unmeasured confounds. I hinted at them and told you we'd come back to them. So this lecture is all about the the unmeasured confound which I think makes these analyses quite difficult. uh but you still have to do them because the topic matters right people are going to report estimates you have reporting a biased estimate is not always a bad thing you just have to be transparent about what assumptions are required so in particular in this literature we're worried about this little U thing which is unobserved some unobserved variable by which applicants vary and and that feature which I'm going to call ability because this is the way it's usually discussed in the literature influences both the probability your application is accept accept it, right? Because if you're of a high ability individual, you make a better application, right? And that should help. Um, and it also influences the department you apply to because some departments are harder to win admission to. Uh, and if you're a high ability individual and you know that, that will change your probability of applying to those departments. That make sense? Um, so there's there are investigations trying to back up those assumptions and such. And I'm going to uh with apologies skip all that and move forward into the modeling part which is usually harder to find. So the first thing you do let's let's make a generative simulation of this so you can better understand it and then we'll know the ground truth and I can show you the consequences of modeling assumptions from here. Okay. So here's some simple R code. I apologize that this is not spaced very well. Uh but I'm going to step through it in a little bit. Um, we're just going to simulate 2,000 applicants. And the first thing we're going to do, you follow along in the DAG, uh, remember if you simulate when you to turn a DAG into a generative model, you have to make a bunch of distributional assumptions, but there's an order you simulate the nodes in, right? You start with nodes that have no parents, have no arrows into them, and then you do the ones that only have one parent that you've already simulated, and then down the chain. Yeah, does that make sense? Um so we start with G uh gender and we just make a a even gender distribution. We're just sampling genders one and two uh into um into the data set into the applications. Now I'm going to I'm going to simulate the other node that has no parents which is the unobserved differences in ability. And I'm just going to make this binary in this case just for ease the conceptual ease of thinking about it. But you can make it continuous. Um and uh this is as in the DAG uncorrelated with gender. Yeah, it's just a it's just ability. Um and then we get to simulate department choices. This has because we've simulated G and U and those are the parents uh of the department choices. I'm going to have two departments just to keep it keep it cognitively simple, right? Just two departments. And um I'm assuming here that gender one tends to apply to department one. two tends to apply to two. How conveniently they're labeled that way, right? It's almost like I made it up. And uh but um gender one individuals with a high ability um tend to apply to two as well. Yeah. And the the the background story here in my fake data is that two is a discriminatory environment for gender one. And gender one individuals know this. But when they're high ability, they're they're willing to fight, right? Go in there and do it. Uh dominate those jerks. Yeah. Um the story may be based on a true story. Yeah. Okay. Um are you you with me so far? Yeah. Now, there's a particular function that I chose to do this, and I'll let you parse the R code later and play around with it, but that's that's what ends up happening. And we've got this table here that shows you um when U is zero uh um gender one is in in department one and gender two in department two. Yeah. But when U is one um gender one shifts to department two for high high ability individuals. Okay. This is what they apply to. We haven't accepted anybody yet. This is just applications. Okay. Uh and then we finally get to simulate the node that has three parents. uh acceptances. And again, the the code's a little opaque here. I'm just going to do the summary. Um department 2 discriminates against gender 1. That's what's going on here. Okay. But high ability individuals get accepted at higher rates. And I've made some matrices to do this and I'll let you study them uh later and confirm that this is what they create this effect that high ability individuals in both departments are more likely to be accepted. Um and department 2 discriminates uh against gender gender one individuals. Um they have the wrong integer, right? If they wanted to be accepted, they should have had it too, right? Sorry, I joke, but this is, you know, apply the labels as you like. Yeah. Are you with me? Yeah. This kind of of generative simulation, I encourage you to always do it. Like really, I know it seems like nobody does this, Richard. Why are we always doing this? because it's the only way to debug your thinking and to justify the statistical analysis is to have a generative model. Okay? Otherwise, you're just hoping and praying and moving your hands rapidly, which can get you published. Absolutely can, but that's not good enough, right? Okay. Sorry. I know you you're used to my sermons now and you keep coming back, so obviously they're not that bad. Um, right. Let's analyze. So, now I've got simulated data and we can uh fit two models. The structure of these models are exactly the models from the previous lecture. We do the total causal effect of gender in model one and we do then we get the direct effect in model two. Yeah. Do I need to should I review the structure of these or yeah the remember the total causal effect we don't have any adjustment set because it's not confounded in the DAG. Um uh for the direct effect we need to stratify by department to get the direct effect. And so that's what we're doing in model two. Uh so the total effect shows a disadvantage which is true. And now I'm it's like I'm putting up these coefficient tables and asking you to read them, but I'm not. On the next slide I'm going to plot these. Okay? So just hang on. Um and then the direct effect is confounded. And I'm not going to ask you to read those that matrix. Um it's uh in in a sense this is just here for you to feel like that's so confusing. Yes. If it's confusing to you, don't put it in your papers like this either. Right? You've got to plot this stuff. Yeah, we know the truth here and we still can't read these tables. It's very difficult. So, plot them and um plot the contrast too. I'm just showing you the marginal posteriors here, but as you know, you want to look at specific contrast too as well. And in this case, these these are not misleading. Um uh there's a so looking at the posterior distributions on the bottom half of this slide. Um blue is uh department one applications solid uh densities are gender one and the dash ones are gender two. So on the far left we've got um the posterior distribution of the probability admission for department one gender one that's the first one on the far left and it's low uh and then um gender two and department one actually has a slightly higher uh probability of admission in this sample. Um and then there's department two uh where probabilities of admission are higher overall and even though we built into it an discrimination against um gender one in department two that's the solid curve there that shows there's no discrimination showing up in these estimates why I'm going to we have multiple slides coming to explain that is because of the unobserved confound okay and I'm going to try to help you understand why it is masking the discrimination. I've got multiple slides coming up to show it. Um, okay. Uh, and you guessed it if you if you've been with the class this long. This is collider bias. Yes. Is this meme too old? Does anybody know this at all? Yeah. Okay. This one verse. Yeah. This this game came out when some of you were babies. I know, right? But it's timeless. It's absolutely timeless. Okay. So, yes. Uh, collider bias. You guessed it. Uh, when we stratify by department. So D is someone's just now understanding what's going on. This is good. Welcome to San Andreas. Right. This I think this is the thing. But um, uh, so come back to the DAG and I I got to show you the next slide again. Um, D is a collider on the path between the treatment gender and the unobserved confound. If we don't stratify on it, that's fine. the unobserved confound does no harm at all. Right? There's no confounding of the total treatment effect of gender in that DAG. But as soon as you stratify by something a post treatment variable like department, you open up that confound path and you bias the estimate. And that's what's going on. It's because of collider bias. Um so it's like thing on a lot of these literatures when there's cross-sectional mediation analysis. The thing you can estimate with weak assumptions is the total effect, but nobody wants that, right? what you want is to decompose it and that requires strong assumptions. But that's okay. The thing that people want requires strong assumptions. Show people what those assumptions are and show the consequences of them. Yeah, it's it's okay. Um uh the estimate is likely to be biased by a certain amount because we don't we don't have the use measured. Uh but you should still report it, I think, and be transparent about that. And we're going to do even better. We're going to do some sensitivity analysis as we get going. Okay. So here's here's a more intuitive explanation now that you see the the DAG D is a collider you see between G and U. So as soon as you stratify by D, you open the path through the through the confound. Yeah, but it's harmless otherwise. It's totally harmless otherwise. Isn't isn't life great? Right. So when you first took your first stats class, they probably told you that there's no problem omitting vari there's no the threat with causal inference is that you omit something that you needed to adjust for and they didn't tell you that adding things could could induce bias, but it can, right? That's what colliders are. You're not safe either way. Yeah, it's it's just this is why you have to have a causal model um uh that that justifies the adjustment set. Okay, here's the intuitive explanation. High ability gender-run individuals apply to the discriminated department anyway because they're badasses. Yeah. Yeah, that place sucks. They're jerks there, but you know, if I if I get in there, I can make more money, [snorts] right? So, they're going to apply anyway. Um G1's in that department, uh if they manage to to their higher ability, so they can manage to overcome the discrimination and get admitted and then once in of higher ability than the average uh G2 in that department. Right? Because G2's ac department 2 is accepting G2s without discrimination. Right? So they're getting low ability ones. Uh so the discriminated class in department two is higher ability because it's the only way they could exist in that environment. Am I making sense? Yeah. Again, this may be based on a true story. Yeah. Inspect your intuitions about this. Um uh okay. And and so this is a masking effect. It's sometimes called that discrimination can be masked by the survivorship bias of the ability of individuals to coexist and thrive in a discriminatory setting. Like and I could see that some heads are like yes lived experience, right? It's like yes again based on a true story. I care a lot about this literature. I've known I have known people who have gone through stuff like this. Absolutely. Okay. Oh, I should pause. Does this make sense? Yeah. You with me? Um uh I mean on the one hand this is very frustrating for causal identification. On the other hand, isn't this cool that things like this happen, right? So you can't just you can't just naively read the effects. You have to think about things like this possibly happening. Now I want to return to these two papers that I mentioned um last lecture and do a more and I said some things about them that was critical like they don't even have clear estims and there's no causal models which is true but I want to do a more careful analysis now for you where I draw DAGs for them and I explain to you how these papers can use the same data set and get different results. Okay, is that going to be fun? Yes. All right, here we go. All right. The first thing though I want to say before I look at these two papers is here's a great paper on this topic that has DAGs. Um this is published somewhere now but you know use the archive. Be nice to the archive. Um and uh this is a great paper. It has DAGs. It it has lots of stuff in it from this lecture but also other examples of structural confounding when we're trying to understand uh bias and and fairness. I highly recommend it. It's a very clear paper. Okay. Um, the first paper is this this Larbanet one from from 2022. It uses a sample of members of the National Academy of Sciences. The National Academy of Sciences at the USA is some kind of Masonic organization. I'm not sure exactly what it is. Sorry. I'm I used to be American, but I gave up on it a long time ago. [laughter] And I was born in Germany. I mean, even though I have an American, right? So it's like uh but no it's it's like it was created to like give the president advice on matters of science but it doesn't do that anymore. It's like a prestige organization for scientists. Members elect other members and then they have their own journal called PNAS they can publish anything they want in. Yeah I'm I'm only being slightly snide about this. Right. The variance of PNAS papers is huge. Yeah. Some of the best and worst things you will read in any journal just page after page. Right. Because there's very little editorial filter. Um, okay. Uh, so they've their own they've got a sample. It's just members of the National Academy of Sciences, uh, which is a bunch of different subjects. Um, and historically there's a big gender imbalance in members of the National Academy of Sciences, as in many scientific professional organizations. And they're looking at citation networks now. Um, and they're looking at what is the effect of gender on citation networks. And they find that women are are being a woman is associated with lower lifetime citation rates. Okay, that's the first paper with me. The second one looks at elections to the National Academy of Sciences. This is the same data set now, but we're not conditioning now on being a member. Yeah, we're looking at a larger population. Uh uh in this case, women are associated with three to 15 time advantage being elected to the National Academy. Controlling for citations. This is going to be important in a moment, but so far it's just to understand the difference between these analyses and what they've done with the data sets. Yeah, this makes sense. Um, they control for citation rates. So that's like saying we stratify by citations. So, uh, men and women who have similar citation rates, women are more likely to be elected to the NAS is what they find. Okay, you with me? All right. Um, I'm gonna have fun with this. I hope you do, too. Okay, so now let's let's talk about let's dag this out. This is the data set in my mind. U we're looking at gender uh gender influences citations uh maybe through that that would be a discrimination effect here or a social network effect. It may not be discrimination but it's structural. Yeah. Depends upon uh which field you're in. Some fields site a lot. Some fields site very little. Right. Like in the humanities there's very little citation. That's not a criticism. It's just there's there's less. Yeah. and the humanities person is nodding. Yeah, there's just less citations. It's not like in psychology, citations are half the paper. Yeah, there's that paragraph in the intro of every social psychology paper which is just citations that has been forced to put there by some editor or something. We don't have to do that in the humanities. It's just not like that. Yeah. Um so that's citations. Uh gender may have an effect on becoming an NAS member. That would be direct taste based discrimination. Um, and what's hidden here again is likely uh some number of variables. I'm just going to call it quality, which influence both citations. If you make good papers, they hopefully get cited more. It's not always true, but you would hope it's true. There are other ways to get cited. Um, and if you're a higher quality scientist, uh, uh, hopefully you also more likely to get elected. You with me? All right. So, this is like what was before. Now, let's think about the first paper. We're going to restrict the sample just to people who are members of the national academy. So we're conditioning on M on being a member, right? We've we've it means we've take if you had the full population, you throw away all the people who aren't members. That's like conditioning stratifying by that, right? But it's it's like you've made that decision by throwing away data. Yeah. Question. I'm not really trying to >> um in in the first paper uh that that's the next paper. This paper didn't do that. This paper looked at citation rates. The first one looks at citation rates. The next one's going to look at who gets elected, right? But it's the same data set, remember? Yeah. Okay. So, this paper only looks at members. So, they implicitly are conditioning on membership. Yeah. So, it's like they stratify by membership. Um so what does this mean? Well um member is a collider between Q and gender. When you stratify by it you open the biasing path. Do you see it? You having a good time that people are like no why did I come to this lecture today? No. No. Yeah. But you see it right when you this is one of these things the kind of conditioning that happens. you didn't put you they didn't think they put member in their model but they did because that's how they constructed the data set right they the whole data set is stratified by membership does that make sense and so if there's if it's a collider on some path you activate it as a collider yeah doesn't have to be in the regression model yeah you can do it outside the regression model if you just if you just uh construct the data set so it's only members um and that means that we've got some biasing quite likely So now think let's do the thought experiment. Remember this um this paper finds that uh uh uh men were cited more conditional on membership. So uh if men are less likely to be elected which is what the other paper finds uh they must have higher quality or citations to compensate. Yeah. Because we've stratified by membership. So there's this finding out issue. It's like, well, okay, there's a bunch of people who've been admitted um uh that if they vary in in quality, now we're going to uh uh in the other paper, if the other paper is true, uh that men are less likely to be elected. Um but now we're looking at individuals who have similar citation rates, men have to if men are going to be elected to overcome the the women's advantage to get elected, they're going to have to have higher citations. And so once we've selected on a sample of members, lo and behold, they find that that men get cited more. Do you understand? Now, this is conditioning on the other paper being correct, which I really don't want to do. [snorts] I'm not sure that either of them is, but I've tried to spell out to you how the logic of these two papers does not live together. Yeah. Does this make sense? This is the way that and the first slide of this lecture where I taught you to be a a a gorilla warrior. This is what I mean. When you read these things, you want to start diagramming like what have they done and what's the sample and what's the DAG and what's the workflow. Yeah. Then you can have an informed opinion about these things. Okay. The next one, DAG's the same same world, right? Same sample. Um they don't condition on membership. They look at uh the membership as the outcome, but they stratify by citations, which is a post treatment variable. Remember conditioning on post- treatment variables is not that you should never do it but there are risks. Yeah. Anything downstream of the treatment uh can either block a it will either block a causal path which in this case we do want to do that because they want to look at direct discrimination or it can be a collider on some path with an unobserved confound as it is in this case. So when we stratify by citations we open up the biasing path through quality again. Yeah, both C and M are colliders on the path between gender and quality. Yeah. So in this case now G is the treatment. C is a post treatment variable. If women are less likely to be cited, which is what the other paper finds, uh then then women are more likely to be elected because they have higher Q than indicated by their citations. Right? So it's not that there's any discrimination going on in the elections. If we can if we think the other paper's right, this paper finds a biased result. It assume it it shows that women are more likely to get elected, but that's because they've overcome discrimination, right? They're higher quality, and that's the hidden variable. Yeah. They're being elected because they deserve it. Yes. It's not positive discrimination, which is what people interpreted the paper as evidence of. Does that make sense? Yeah. Just like you think of think of a uh individuals, male scientists, female scientists, they have the same lifetime citation rates. Um if women are getting elected more often, it's because of the the the hidden path here, the quality variables. Yeah, that's or that could be the explanation, right? It doesn't have to be the direct effect. It could be the biasing path. Does that make sense? Are you feeling good about this? Yeah. Now, okay. Now this may seem like some fanciful thing and I have dissected these two poor papers. I'm sorry but it's making you feel good. I know this this theme but no this is I mean this is part I try to choose examples in this course which matter right and and things matter and this kind of structural diagram the simple mediation path is what I'd call it is incredibly common right and people are often doing mediation analyses where they condition on post treatment variables and it's it always requires strong assumptions yeah even if you've done an experiment where you randomize the equivalent of gender right we're not going to randomize gender here but if you did an experiment where gender with some X some treatment and you can randomize it, you still have this problem because you haven't randomized anything downstream of it, right? So you start conditioning on things downstream of it. The bias will still appear. Good times, right? Yeah, sometimes. Sorry, there's some people are not getting less and less happy with this. No, but this is going to be a happy ending. I promise you question. >> So like assumptions. >> Yeah. these papers. >> Well, uh the question was if they assume quality is the same for men and women. We're assuming quality is the same for men and women. You still get a bias. You guys notice there's no gender doesn't point into quality. Yeah, that that's not required. uh that could make it worse, right? If you want to point gender inequality, then it could be even worse. But you don't have to assume that to get this biasing effect. Um the papers don't talk about these things at all. They have no structural diagrams, no identification strategy. They just start regressing. Yeah. And then it's kind of the force of you you know how these papers are. It's the force of the rhetoric. You get swept along in the narrative, right? A well-written scientific paper can sell you on a lot of results, right? And it's like the methods are maybe not even in the paper. It could be online somewhere. Yeah. It's it's this is the the system I'm criticizing. You know, you know my rants. Yeah. Question >> maybe. Is it the same as what we talked about before with the restaurants in Paris >> that like in order to get into science, you need to either have a lot of citations or have high quality. >> Yes. >> And those are actually not correlated. as a big cloud, but we're kind of slicing off the top right corner. Yeah. >> And by doing that, if you take a circle and you slice it, you actually do get a >> you get a negative correlation. >> Makes it sound makes it seem like you have a correlation, but they're not correlated in gender. They're not correlated in the general population, but they are maybe in the NAS population. >> Yes, that's right. That's exactly the argument. And the microphone didn't pick that up. This was a good question relating this back to my restaurant example when I introduced collider bias. It's the same phenomenon. It's absolutely the same phenomenon. Yeah, there are multiple ways for the thing to happen. Um they're causally independent. Once you stratify by the collider, you've induced a correlation. Yeah. This is the classic case of collider bias 2 where the different causes compensate for one another. Exactly as you say, there's multiple ways to get admitted. Yeah. That's exactly what what goes on. Good. Another question. >> Always like something unmeasured as a collider or something like that. We try to surround for example and we always try to look at some particular population of of people >> for example here numbers and they are and >> what to do is this the question >> it feels like it's impossible to >> it feels like it so I I think there's lots to do I think this is very constructive you you start drawing out assumptions like this and you can develop it into a workflow so what I'm going to continue with in this lecture is sensitivity analysis where we inspect the the cues and we keep them in the model when we do the analysis even though we haven't measured them and you're going to say how can I do that well I'm a basian I can do anything [laughter] um no probability theory can right basians just obey the axioms remember but uh uh another thing to do and I think this is your question is you say every sample is conditioned on something right exactly we'll draw it in the DAG how do you get into the sample make that a node. Uh, and people do this. So, people who work on on causal inference of surveys do exactly this. Um, who gets into the survey? And there's lots of self- selection in surveys and lots of non-random non-response bias. Uh, so political polling is famous for this. Uh, people's partisanship affects their probability of responding to polls. Um, and then that's a very that's a that's a problem. But you can draw it in the DAG. You absolutely can. So in this case it's it's like how do you get to be a member and even if your sample even if you only have members as your data set you can still draw M on the DAG and analyze it as a collider and understand the consequences or what assumptions are required. Yeah. In order to think that your estimates are unbiased. So you do that social network um causal inference in social networks is similar. You're looking at diads, uh, friendships, marriages, uh, other kinds of relationships. Uh, the the processes that form diads are correlated with features of the individuals, right? So, you need to have a you can dag out how individuals become friends and how that is caused partly by their preferences, right? So the classic example when I when I took grad stats at UCLA the example they used was uh smoking um in married couples smoking is highly correlated. [laughter] Yeah. And you can think why. Yeah. Now many places smoking has declined but not in Germany right. It's still I'm picking on it but it has declined in Germany but it's increasing in teenagers now for some reason. And uh so preferences like that um can actually be causal rights about about these friendships and so or relationships and that affects persistence of diads but you're conditioning implicitly on on friendship formation when you do that. You don't have a data data set of all the people who could be friends. You see so you're implicitly conditioned on diads already. Right? So any any literature that analyzes married couples has conditioned on getting married and persisting in marriage. Yeah. And those are causal processes that depend upon the features of the individuals. And so you may have colliders. And so when I was taught again in grad sets I was taught that using smoking as an example and how that can screw up network estimates. Um okay I could go in I should stop. I have endless examples. Yeah. Uh these things but you learn these structures and you'll start to see them in other data sets. Um, my default explanation whenever there's a hyped abstract, by the way, is try and figure out how it could be collider bias. It's usually not hard. Yeah, you get practiced at this. And it's not a cynical exercise. It's self-defense. Yeah, you have to remember that. Okay, this is pretty real. I I lost the flow there, but there were some great questions and I appreciate that. You should always interrupt me with questions if you want. Um but I was I was going to flow into this this real world example where um uh companies are actually making decisions I think on the basis of collider bias. Yeah. So I can't prove it but here's a case that was uh puzzled many many people when it came out. So Google uh got sued for gender discrimination in pay and as part of that they did a big internal audit uh all their employee almost all their employees like plus 90% of them um of their of their job and their experience level and their pay to see if their how much gender discrimination there was and they adjusted pay in much of levels in the highest level of engineer I think it's called level four uh it's on this slide somewhere yeah level four software engineer women were actually getting paid more than men in that level. So they paid the men more, right? Because they said that's discrimination. I hope you are all primed now to think about how that could be something else that it could be that those women working in a discrimin discriminatory occupation uh and a company that had already been sued multiple times for gender discrimination might have been paid more at that level because they're better. Yeah. And then paying the men the same as them is even more unfair after that. Are you angry with me now? I'm sorry that was a little too much about me. [laughter] But uh no, but just do the analysis, right? This kind of like, oh, we just look at the differences and interpret it as discrimination. Should never do that, right? It's very structurally difficult in a cross-sectional sample to make causal inferences of that kind. And in this case because of of the masking of discrimination um by hidden quality variables you can you can uh this anyway this result is consistent with that. Yeah. Um there are other explanations as well like um uh in in discriminatory environments men are put hired at higher levels uh but have lower ability and that'll generate the same thing. it'll still be an ability difference and that explains why the women are getting paid more, but the the male engineers were just put at level four quicker. Yeah, that would also be an explanation. Um, okay, sorry that was my rant, but I wanted to make it real for you. Yeah, it's quite difficult uh in in these sorts of organizational data to figure out exactly what's going on. And remember, in a discriminatory environment because of hidden confounds and colliders, uh the discriminated class may actually look like they're overperforming. Right? Because it's survivorship bias. Does that make sense? I'm going to say this over and over again. Sorry, I'm like really obsessed with this result, but this is where causal inference pays for itself. Yeah. Okay, I've got 20 minutes. Um, and I want to talk about sensitivity analysis. So, let me try to summarize. So, I keep saying if you don't put any causes in, you're not getting any causes out. You've got to have a generative model. Um, what we have here are hike papers with vague estims. uh um unjustified adjustment sets. Um and in the case of that Google example, you're even getting policy design through collider bias. I think it's very plausible that that's what's going on. Uh we can obviously do better about this. This is a serious issue. Um uh all kinds of ways, discrimination, pay fairness, uh uh policing, uh is another example. I've actually Oh, the next slide is actually about policing. Sorry. Um and but you're going to have to make strong assumptions. There's just no way out of that. And that's okay. strong assumptions are how you license strong inference. As long as you're transparent about it and you're not hiding the assumptions, then you've done nothing wrong. Last line here where I I wave my um anthropologist flag, qualitative data is super useful in these contexts. Quantitative data is not always better. Uh quantitative data in these cases is highly conounded. But when you get when you talk to people, you get their testimonies and the detailed mechanistic testimonies of their experience of discrimination that really helps you understand the qu the quantitative data and develop better generative models, right? So those interview data, the soft humanity stuff that people like me do, like it's really important input into causal modeling. Okay, sorry that's my anthropologist flag. Uh people at home can't see me, I'm waving an imaginary flag right now. Um it's why we do ethnography and why we talk to people. It's not stupid to talk to people, right? Okay. Sorry. Um, I work at an institute with a bunch of molecular biologists who don't respect talking to people, so I have to say these things sometimes. Okay. Uh, I have worked with some of my colleagues like Cody Ross. Some of you know Cody. Cody and I have have a few papers where we apply this logic to the literature on discrimination by police. And um this is most famous in the US, but it's not unheard of uh in Germany. German police can be a bit difficult too at times. Um some nodding heads out there. Yeah. Um okay. So, uh I just want to recommend to you that or point out to you that this kind of literature has some similar issues. Um often you have these database called administrative databases provided by the police. So right away you're suspicious how good it is, but they're conditioned on being stopped, right? This is most of the data set is their police stops and there these official records of when the police actually stop someone like it's a traffic stop or something like that or someone's acting suspicious in the park and someone calls the police on them or something like that. That's a stop. Um and so you're the sample is conditioned on being stopped by the police and uh yes you can believe it that there is probably things like acting suspicious um that is an unmeasured confound in these things that biases estimates. There's a great paper on this which is the citation at the bottom um Knox and and Mumalo. Uh this is a great paper. It's got DAGs. It really walks it out. What's an identification strategy? And they show actually with very reasonable assumptions you can still do identification with these data sets but you have to expect them to be biased by the fact that they're selected on being stopped. Does this make sense? But it's a hopeful paper, right? This is not nihilism. All of us want to solve these problems because they matter and we work hard on it. Okay. Sorry, my rant just continues. Uh we're going to get through this the sensitivity analysis. Um okay, so we let's here's what I want to do now. a sensitivity analysis on the unobserved confound. Um, well, you could say, what are the implications of what we don't know? We're going to assume this confound exists, this quality confound, whatever you want little U to be. Yeah, it could also be attractiveness, right? There's a lot of evidence in these literatures that attractive people get stuff in the world. So, it could just be random attractiveness, right? That has an effect, too. Um, and we can interpret that as discrimination if you like, as well. So it doesn't have to be quality differences, right? When someone who's dumb doesn't get a job, you don't call that discrimination, right? But if someone's unattractive, they don't get a job, we would. Yeah. I used to teach a class to undergrads back in my previous job where we spent like a whole week on discrimination and talked about these things. And some forms are very upsetting to people and some forms are not. Yeah. Tall people get stuff too, right? Why do we discriminate against short people? There's lots of horrible things people do to one another. I should stop talking, right? But no, I I love this topic. I used to teach it quite a lot. So, um I mean I hate this topic that it exists, but you know what I mean. Sorry again. I have to choose my words more carefully. I got to get if I were a politician, my career would be over. Uh so, we're going to assume this compound exists and model it consequences. That's the transparent comprehensible thing to do. And we can vary the strength of the confound, make different assumptions about how it affects the structural model and see what the consequences are. Uh now that doesn't mean we we can't figure out what the truth is, but we can learn what the consequences of our assumptions are in the truth and we can calibrate our uncertainty, right, about how strong the confounding would need to be, for example, in order to explain away the effect. Yeah. Or create an effect. Does this make sense? Um, and you already have all the tools to do this and I've taught them to you in this class and let me show you how. Okay, we're going to take our DAG and we're going to do what we always do with our DAG. We're going to make it into a generative model. So, the first thing is the uh outcome model for A. This is just a logistic regression. I taught you this in the previous lectures. Yeah. Um, and um what's different here is that I've taken you and I've put it in the model. Do you see that there's a little U subi for each individual eye? But you're thinking, "But we haven't measured that. How is it in the model?" Oh, we're basians. We can put unmeasured things in models. They're called parameters. You've been doing it all along. There's nothing different here. A missing value is is a parameter. That's all it is. Yeah, a prediction is a parameter. You don't have any problem with that, right? It's just a missing value and you're going to use the model to predict it. So, the same thing here. Um and then there's an outcome model for department choice that applies simultaneously. So I create this indicator variable when department for individual I equals two then it takes a value one otherwise it's a zero. So it's another logistic regression right we're modeling the probability of applying to department two which remember is the discriminatory environment in my simulation with me. Um and again the U appears. It's in both equations because uh the quality affects both, right? It has an there's an arrow into department choice and there's an arrow into uh uh acceptance. Notice that the beta which is the effect of of hidden quality on acceptance can vary by gender, right? It's subset by gender. So it's interacted. This is an interaction effect if you want to think about it that way. And the gamma, that squiggle is called gamma for those of you who don't read Greek, right? It's one of my favorite Greek letters. It looks really nice and um it's like artistic Y. It's a really nice letter. And uh it's also subset by gender, so we can vary it. Yeah. Um and that good. Yeah. You with me so far? Um and then we need prior obviously we need prior and we're going to put a prior on the use. And I'm going to put normal 01 prior on the use which just establishes a measurement scale for them. Right? This is a latent unmeasured variable. The scale can be anything. And I'm just saying they're Gaussian with a variance of one. Yeah. If the variance were bigger, we could rescale it. Right? It's a latent variable. The v the variance is arbitrary. Does that make sense? Um what you need to do uh to make that work though uh so that the the plus plus values values of u greater than zero mean greater chance of acceptance is to make the um uh the the betas and gamas positive right because that sets the the veilance of what higher value of u means yeah so we do that and in this example at the very top let me show you what I'm doing here I'm just putting in the the betas and gamas called b and g as data. This is the the really rigid way to do it. There's nothing wrong with this is we're going to say let's assume the effects are of a certain size, right? This lets you test the effect of the confounding scenario. Um so in this case we assume that high ability means a higher acceptance rate. Uh those are the betas, right? Those are the that's that vector of one and one, right? So the one means a larger value of you makes it more likely to be accepted. You're more likely to be accepted. And then the the G's are one and zero. Um G1's with high ability are more likely to apply to to D2. So this is the effect of your ability on applying to department two. And so gender one gets a one and gender two gets a zero. Does that make sense? And you can change these, right? And that's the idea is you vary these inputs and you see what the consequences are. And I'm going to show you an example in a moment. Um and so here you see the there's the U um uh there in the model. Um, and then down here we've got uh U subi uh again for the department model. And and here's the vector of length in one for each application of use. And we're going to estimate them. And you can you can get posterior distributions for all these. Yeah. Are you with me? Is this fun? I know this is like basian madness, but it's licensed by the causal graph. Yeah. It's totally licensed by the causal graph. And since you've got the synthetic data simulation, you can verify this this works. It reveals the truth in particular scenarios. Um, this is one of my favorite things in classes like this is we get to have there are 2006 parameters in this model and 2,000 observations. And at some point you took a baby stats course and you were told this is illegal. It is not illegal. You can have more parameters than data. That's how neural networks work, right? All those chat bots which are ruining your life, right? [laughter] the many many they have billions of parameters, right? It's just how it goes. Of course, they're also trained on billions of blog posts unfortunately, but anyway, but this is totally fine. You could get the posterior distribution in these scenarios. Um, okay, let me show you what happens here. I've only got eight minutes left here. Um uh on the top it this is just the graph from previously where we ignored the confound here the marginal posterior distributions where discrimination of gender one in department 2 is masked by their quality differences. Yeah. And then in the bottom uh where we assume that that's what's going on which is the true model you see the discrimination pops out. Right. In a real data set, you're not going to know what the ground truth is, but you can show the consequences of the assumptions at least. And you can vary the strength of the assumed effects of the discriminatory effect um to find out how big it would need to be to reverse a trend or or whatever it is that you're analyzing. Yeah. So that's what the sensitivity analysis would be. Um you can also put priors on those things. So here's another version of this analysis by this script is in on the website or actually I don't think I've uploaded it yet. it will be on the website. [laughter] Okay, I'll put it in the scripts folder. Um, so you don't have to type this from the slides. Uh, here I'm not going to enter the the B's and G's as data. I'm just going to put prior on them. And I'm going to put these uniform 01 prior on them. Again, we need these to be positive. Otherwise, the directionality of the use doesn't make any sense. Yeah, there's a question. >> Could you in theory put B and G or GMA as the same? Yes. Yeah. Yeah. You could you could do any scenario you want here. >> Absolutely. >> It should be the same. >> But the G's are the G's or chances affect the probability of applying to department two when you're high quality. So they have you can do different scenarios. And in the original scenario, it was only department two is discriminatory against gender 1. So quality only affects gender one's probability of applying. But it could also partly affect gend um uh gender 2. Yeah. Right. It's just higher quality individuals more likely to apply to department two. Yeah, you can try different scenarios. Here I'm just going to put these these kinds of default vague priors on it. Um not as a recommendation for these priors, but uh here I am making them the same for both genders just just to show you what what's possible. Um and yeah, I should have highlighted this already. That's where they are. And then when you do the estimate, what the model tells you and this is super useful is you can't say anything, right? the these data if you cannot pin down the values of those of those effects are consistent with a wide range of things. They're consistent with lack of discrimination. They're consistent with quite strong discrimination. >> Yeah. Does that make sense? Question. >> So changing scenarios based on camera. >> Yeah. >> Is what we call sensitivity analysis. >> All of this is sensitivity analysis. Sensitivity analysis uh I'll come back to this in the next slides is just changing any of the assumptions. It could be the data model. It could be the existence of confounds. It could be the priors and seeing what the consequences are for inference. That's sensitivity analysis. In this particular case, the goal of the sensitivity analysis is to show what the impact is of assuming there's an unobserved confound. Uh and then we can try different structural scenarios about the effect of that unobserved compound. Yeah. Does that make sense? This is uh this is not my private madness. You see this stuff in journals. I'm telling you like increasingly common because there are packages now that automate it for you. You can give it uh you can give a multiple regression to the package and tell it structurally where the which variable should be correlated because of an unobserved common cause and it will make a plot for you even in some cases of how big the confound needs to be in order to um uh remove the effect you've observed. Yeah. which is useful. So that's nice. It's like a maturing of a scientific field. Stage one, deny compounds exist. Stage two, admit they exist and do nothing about it. Stage three, simulate the consequences of compounds existing and report it. Yeah. Does that sound good? It's like therapy, right? So we're moving up to these these situations. Um, good. You ready? Okay. So yeah, this is what I was going to summarize for you. Sensitivity analysis is about um quantitatively figuring out the implications of of structural assumptions. In this case, things you don't know like the strength of the unobserved confound, but you have good reason to suspect it exists and you probably cannot convince a reasonable skeptic in your own field that it doesn't exist. That's often the case. Um that said, when we say there's some confound, you could be applying this critique to someone else's work. uh it's our responsibility to really model it, not to just wave our hands and try to dismiss a result because there might be a confound. Uh both because that's rude, right? I mean, criticism should should involve some effort. Uh and second, some things even if they're biased, need to be reported. You have to take action sometimes, but you want to be you want to really math out um what what the confounds might be doing and what the what the errors might be when you take action. This is somewhere between simulation and analysis I say. Uh but the way I've been teaching analysis, there's a bunch of simulations. So maybe you didn't notice, right? It just feels like the rest of the stuff I've done. Um so yeah, I've showed you a case where we um we we vary the strength of things directly and I gave you a case with priors. You can also vary the other effects uh uh make assumptions about the other regression coefficients like the uh admissions rates because it's all coexists in the same model and then infer the discrimination effects, right? But you're going to have to make assumptions someplace in order to do inference. And again, that's not bad. Assumptions are how we buy inference. Um okay, good. Um I wanted to end uh again on on the guerilla workflow. Um the scientific literature is a powerful foe. Yeah. You do not have the firepower to directly confront it. It's too vast and complicated. So you have to be very strategic about how you read, defend yourself against false beliefs. Um and I've tried in this course to give you tools to do that. Uh the workflow is not some toy thing. It's it's a technique for dissecting papers that have been published. Um for doing better work yourself and for uh resisting the the sometimes downward pressure from supervisors and senior colleagues to do bad work. Yeah. It has to be justified and comprehensible. Yeah, that's what I mean. Um and this little I mean this is I tried to make this what I was thinking about when I was writing this lecture like what's the most minimal workflow diagram. It almost looks like a rune, right? And that's this is what you can remember is questions, generative models, justifying statistical models, then test. Yeah, test that that statistical model um is a is a logical outcome for the question and uh uh and generative model and test the implications of the priors, right? Because the priors are always a feature of the statistical model that isn't present in the generative model. nature doesn't have prior. Can I say that? Yeah, there must be a basian somewhere who will not agree with me on that. They will tell me later when they listen to this. But I don't think nature has prior. So, but we have priors. We need them to do estimation. You have to have them. So, you should test the implications of your priors when you design your stat model. So, that all that goes in that first test swoop. Yeah. I need some I need the workshop vocabulary here, but you know what I mean. Yeah. Uh and then we get estimates. We test again that the machine worked. chain diagnostics, etc. Um, should also be testing your data, by the way, against the data dictionary. I should put another test on here, right? Are the values in the data set what they're supposed to be? Every time I've done that test, there are things that are illegal, like children older than their parents and demographic data sets. This happens routinely. Yeah, that's what I mean. Test the data dictionary. You define the data dictionary. You've got constraints not only on each column, but on relationships between columns. Run those tests. That's what we do in my department. We're obsessive about this, right? because we find mistakes all the time. Genealogy data data. Oh my god, like the mistakes in genealogies. Um, impossible human families all the time. Okay, I should stop talking telling these stories. I only usually only tell these stories to my therapist, right? So, okay. So, we we go and then finally predictions. We want to test the predictions that we computed them correctly, but also that they they're appropriate to the estim. Are they actually answering the original question? Okay. Um, I'm out of time. Uh, I've had a lot of fun teaching this and I hope you've learned something of value and I wish you a good rest of the year.