Submind YouTube summaries
Thumbnail for Fundamentals of Active Inference (Chapter 8, Session 36) August 14, 2026

Fundamentals of Active Inference (Chapter 8, Session 36) August 14, 2026

Watch on YouTube

Video summary

Chapter 8 unifies concepts from perception and action by integrating inference, attention, and learning within a framework designed for continuous state spaces. This model expands beyond simple hidden states to include beliefs about first-order parameters governing transitions and observations, as well as second-order precision parameters that represent the inverse of variance. The system operates across three distinct temporal scales: fast-changing hidden states handled through instantaneous variational free energy minimization (inference), moderately changing precisions managed via attention mechanisms, and slowest-changing parameters updated through learning processes. To optimize these simultaneously, the approach integrates instantaneous free energy over time to form "free action," resulting in second-order differential equations that describe acceleration in beliefs about parameters and precisions. These complex dynamics are simplified into coupled first-order systems using auxiliary velocity variables and damping terms, allowing for stable numerical solutions via Euler's update rule while maintaining a hierarchical Bayesian network structure where all uncertain quantities are represented as nodes with associated means and covariances. The construction of these models involves stacking atomic units—observations, hidden states, parameters, and precisions—into multiple layers to create powerful generative hierarchies. In this architecture, autonomous states at one layer function as observations for the layer above, enabling top-down generation from higher-level assumptions while supporting bottom-up inference updates based on sensory input. A critical aspect of this framework is addressing time scale separation, determining when persistent prediction errors should trigger perceptual adjustments versus deeper parameter or precision learning. While specific causal rules for these temporal distinctions remain an open question often explored through empirical biological studies like eye gaze versus fine motor control, the chapter provides the necessary machinery to frame inference across different scales rather than prescribing a single rigid rule. The model's plasticity is shown to depend heavily on computational constraints and environmental dynamics; for instance, in static or resource-limited environments, parameters may be treated as deterministic to save computation, whereas dynamic settings necessitate frequent updates to maintain an accurate representation of reality. The utility of minimizing variational free energy (VF) extends beyond merely achieving a good fit between the model and environment, as non-zero prediction errors serve valuable information for adaptation rather than being purely negative signals. The transcript emphasizes that while low VF indicates successful alignment, inappropriate minimization strategies can lead to pathological states such as dissociation or hallucination if an agent prioritizes comforting priors over actual sensory reality. This highlights a fundamental principle: the value of error signals and the appropriate response to them depend entirely on the specific circumstances and goals of the agent. Through examples demonstrating how beliefs about hidden states, parameters, and precisions converge quickly despite initial mismatches, the chapter illustrates that effective adaptation requires balancing these competing forces without ignoring reality in favor of internal comfort, thereby ensuring robust performance across varying environmental demands while avoiding maladaptive rigidities or delusions.
Read the full video transcript
All righty. Hello everyone. So, it's August 14th, 2026. We're here in the uh the first session of chapter 8 or first of my sessions for chapter 8, fundamentals of active inference. And today we're going to be looking at essentially well inference uh attention and learning, putting all of these things together. Uh and this kind of ties the bow on everything that we've done from chapters six and seven thus far. Um, as I said last week, you know, chapter six is very much the mathematical core really of um of what we're doing in terms of how we've set up the generative model and how we're thinking about uh the structure of things. Chapter 7 added in action. We then saw that story. We had a look at the Beijing thermostat as well. Um, so that that example is is there on the page for chapter 7. That's really the first active inference example we've seen and we'll be seeing more as well. Um and so we're kind of finishing this. We're kind of doing chapter three again in some sense. So in part one, you know, we did perception. Um chapter three then in, you know, we talked about learning. Um we're kind of doing a similar sort of thing now where we're doing in addition to state inference, learning and attention. So that's kind of what chapter 8 is all about. It's quite a long chapter. Um so we've got six parts, 8.1, 8.2, all the way to 8.3. I think it would be good if we could get through this. a suggestion, but I think if we can get through 8.1 and 8.2 today, that's quite a lot of material and that will set us up very nicely for the rest of the chapter as well. And we can look at some examples next week as well. So I will begin sharing my screen here, my entire screen. Hopefully everyone can see. Please make a lot of noise if you can't. So here we are in chapter eight. We've got the actual structure here. This is and I do apologize. It's uh uh not you it would be good to have this up slightly before sessions but we do have the the overall breakdown of the chapter now uh much as we have the previous chapters there is the PDF version of this same overview as well. So you can click on this and see okay what's the content breakdown in terms of core con core concepts and such like or at least kind of my thinking about what is core and what is kind of optional. Um and then we have a map for the chapter basically. So 8.1 all the way to 8.6 six and we'll be doing 8.1 and 8.2 today. If we go down in the actual markdown here in KOD, so we've got our chapter map, you'll see for 8.1 and 8.2, I have the same notes again um just for those subsections. The notes for um subsection 8.1 are quite long. It's 27 pages. A lot of this, so the general idea, as I say, is so that people who don't have the book, they can follow along. But I've also tried to, I guess, preempt questions that might come up from people about why certain things are the way they are. Um, and maybe elaborate some detail from the book that's perhaps brushed over in in in too too quick a manner. Um, so, you know, I I do try and be a little bit more comprehensive here. This is not a direct copy of everything that's in the book. Um I'm just trying to give more motivation as to where things come from and what's going on generally speaking um and to get the core concepts basically. So that is the deal there. Do have a look at these if you're you're interested in sort of a 30,000 foot overview. So we've got 8.1 and 8.2. I'll need to make the notes again for 8.3 to 8.6. So yes, session one. Let's do 8.1 and 8.2 today. So, here's where we are. This is the overview for chapter 8. Basically, as I say, you know, this this is uh and once we've done chapter 8, we're going to be moving into part two of that's part three of the book, I believe. So, we can come here all the way into the introduction. We've seen this many times now. Uh there we go. Oh, no, I lied. So part two, you know, will be finished with chapters 9 and 10. They are quite different chapters to chapter 8. Um, chapters 6, 7, and 8 is all about the sort of continuous state space formalization of active inference. And then chapters 9 and 10, we're going to be start we're going to start to do things in a discrete state space. We haven't seen that yet. Um, so don't worry about what that means just yet. But uh this this last chapter that's going to sort of put tie the bow on the fundamental ideas in active inference uh for the the continuous states space side of things. Let me just mosy on back to make sure there's some people in the waiting room. Admit admit admit. So that's where we're going. Um if people do have questions, please put them in the chat. Um and I will do my best to uh defer to them. So and again hopefully I am sharing. So let's go into uh section 8.1. It's quite a bit to cover really. So let's uh let's have a look here. The general idea is that we've seen from chapter 7 and chapter six how perception works in terms of a generalized state space model. And again I I would focus on the univariate case. Okay. So typically what Sanjie does very wisely is you know he presents the univariate case very simple um we're only dealing with one hidden state or one vector or something like that and then he'll generalize. So to get the core you really want to think about the the univariate um side of things initially and that's what I would recommend that you defer to in um in your own sort of attempt to understand what's going on. So learning first and second order parameters. Okay, this is this is 8.1 here. As we've seen well I mean as has been the case typically we've assumed that the parameters in our model especially from six and seven they've just been given to us god-given and then we we can go ahead and do hidden state inference with them. So now we're going to relax that assumption and that's that's the kind of key thing that we're doing in chapter 8. Um so we have we have model parameters such as you know the parameters around the state transition model so that's our f and the parameters around the observation mapping. What we can do is we can just sort of collect these things together and say look both of these are parameters. Let's refer to these as model parameters and that's what he does here is is of concatenate them. if you imagine they're lists of things. Um, and these are what we're going to call first order parameters. Very much highlighting the fact that there's going to be other kinds of parameters. Now, so that's I would say the biggest change in chapter 8. We have this idea that yes, there are model parameters, but now we're going to have to also concern ourselves with two kinds of model parameters. Um, so there's there's literally like what are the settings of uh the parameters and then there's also precisions. Okay. And the idea is that these precisions, these are going to be just another kind of parameter. And we've seen these from chapter five, okay, predictive coding and then chapter six. The idea is that these are also in some sense ways in which we can set the dial on our model in some colloquial sense. Uh so perhaps if we go to 8.4, this is just a list of all the equations. Basically, this is kind of what our model is going to end up looking like. Something like this. So I've I've skipped ahead just very slightly. The idea is that well what have we got? We've got our our um prior on on state transitions. We have our state transition model. We have our observation likelihood given by our observation trans observation function. And the state transition model and observation function they not only take the hidden state they also take this um uh autonomous state as well that's going to be sort of driving the dynamics in some sense or rather an exogenous state and then we also have our variances on both of these beliefs and they're all Gaussian they're all bell-shaped curve now what is this precision so we've got prior now on our model so we have prior on the the actual parameter whatever those are and we also have this other kind of parameter the precision all right and we have we have a belief we have a prior about precision and a prior about the parameters themselves key I've just taken these these have been taken from from the text the thing is precisions of course this is the inverse of the variance of our model and we must keep in mind that throughout this entire chapter we're still in um the realm of uh the the uh the LLAS approximation or quadratic approximation where we're saying the the variational belief is very tightly clustered about mean. We only really have to care about the mean and then we're we're fine swirl and dandy. So the the thing the thing with precisions is that they have to be positive. They must be positive. Okay, you can't have a precision that's negative. It can be zero in which case that's kind of a a degenerate case. But we must make sure that our precisions are positive. So he introduces this is equation 8.1 a bit of a transformation where he says look well let's take the variance okay uh or or precision we need to make sure that this is positive how can we do that well we can take a transformation called the logarithm and that will make sure that um when we exponentiate that we can get back something that's positive. So that's this is really just a trick to make sure that the variances at issue are are strictly positive well or zero. Um and this is so what we're really dealing with is we're dealing with log precision. So the precision is the inverse of the variance one over sigma squ. Take the log of that you've got a log precision. And very handily uh the log precisions are donated with Greek letter zeta. Okay? It looks like a weird squiggle. It's almost impossible to write by hand. So these zeta things, these are log precisions. That's what they are. And how do you get back your regular old precision from this? Well, you just exponentiate the log precision and then you'll get that. So, really quickly, let's look at a graph of just yals x, right? That's just that's my precision in some sense. If I take the log of that looks something like this. All right. Uh, conversely, the exponential of a precision or some y= x looks like this thing. So in some sense you can kind of see colloquially that the the exponential and the logarithm they're kind of familiar images of one another. And indeed if I take the exponential of the log of x so log you know x is just my purple line. Exponential of the log of x I get this green line. And hey presto that is x. Okay. So the only real thing you need to take away is that log and exponential they kind of undo each other. And that's all we're really doing here. So we got log precisions. Very nice. Our model now is going to be it's going to sort of be enlarged. So not only you know we have our joint probability distribution over observations hidden states parameters and now also log precisions. Okay. And we're going to assume that they factoriize in this manner here that allows us then to spin off this the the version of the model that we have. Um, and it's a little bit obscure here, but we have a precision for the state transition and we also have a precision for the observation uh model as well. So those should really be like z and zeta x and zeta y. Okay. All right. So that's what we have. We've enlarged our model. That's basically what we've done. We've just sort of in we've we have a a larger space of things that we're now considering beliefs about. uh one very helpful thing that's it's on the it's in the text here. So of course we're we're we're dealing with the LLAS approximation. Okay, so we you know we have a joint uh approximate posterior now over these things or rather over these things. Okay, that's going to introduce a bit of difficulty for us. So, we now need to let's let's think about what the the variational free energy would look like given this new model of ours. All right. Um well, we can sort of under the LLAS approximation treat the variational free energy as equal to the negative log probability of of our generative model. Okay. So, we can just substitute stuff in and get all right. Well, negative log of our generative model properties of the factorization, properties of logs, we get this expression here. Very nice. Uh, and then we can substitute in the specific forms of the of the densities at issue. Um, and then crucially we have prediction errors. And just like we did in in chapter six and seven, we have our observation prediction errors. We have our well crucially from seven, we have our hidden state prediction errors. And now we also have parameter prediction errors and precision prediction errors. Um and if we if we ignore certain constants that come up given the lelass approximation, we can express the variational free energy much as the same way we did in chapter five and six as the sum of these precision weighted prediction errors. So we have our precision weighted prediction error on observations, hidden states, parameters and now precisions as well. Uh technically there is a there's an additional kind of log precision term in each of these prediction errors here. um you can kind of ignore that under certain assumptions. Um and that's a simplifying assumption basically. So maybe it would be more helpful if I went back to let's go to the actual notes themselves. Okay. Yep. Section 8.1. Perfect. Okay. Yes. Yes. Yes. We do a log transform. Very good. Very good. Okay, so that end we end up with equation 810 which is just uh we have our precision weighted prediction errors. Okay, and again this is really the same machinery we've seen from from chapter six and seven. Um it's just that we have an enlarged set of things that we're we're considering now basically. So the really there's and there's of course some discussion about it relation to predictive coding because it's very very similar looks looks very similar. The main kind of thing the main uh idea of chapter 8 is really 8.1.2 I would say uh basically this is kind of the the whole this is where you this is where the whole thing would hang together. So what have we done? We've enlarged our model. We have beliefs over hidden states. We have beliefs over uh precisions and we have beliefs over model parameters. And here's this is very nicely put on on chap on page 183. So update rules for learning and attention. Essentially we have kind of three different things that are changing. Um as I say the hidden states they're changing really quickly from from you know second to second or maybe millisecond to millisecond. And you know that's the whole the story about inferring hidden states is the story of just Beijian inference more generally. Okay, so we've seen this all the way back in chapter 2. That's where we started. You know, we're inferring hidden states and they they're changing relatively quickly as observations come in, right? But now we have these precisions, okay? And those are kind of the width of the beliefs at issue. And the idea is that well these are changing as well, but they're not changing quite as fast as the hidden states themselves. um and but nevertheless we have beliefs about them. Okay, so we we're we're seeing that there's a separation of time scale now. Okay, so we got fast changing stuff here, slightly slower changing stuff here. And if we think about it colloquially, I don't have um uh I don't have a nice figure, but obviously, you know, we're we're looking at Gaussian beliefs. Here's a nice figure. So, you know, this red belief over here, this is kind of more spread out than the green one. It's less precise. It's got more variance than the green one. We've seen this a lot. And at least in a colloquial sense, this is motivated what we can interpret what this means as the precision encoding, the degree of attention that one would give uh to um the the hidden stated issue or at least how strongly it would be updated. So if I go back here, whoops, no, I want to go back here. the these positions are going to be likened to and indeed there is a lot of discussion around how they are relevant to attention. Okay, so there's hidden states. They're changing. I've got beliefs about them, but you know, I should probably attend to hidden states that are very precise. That's that's the the second sort of time scale. And then again, sort of colloquially at a slightly slower changing time scale, we have the beliefs about the model parameters themselves and we've seen their connection to learning. Okay, so we've got these three time scales. What this does is it introduces a bit of a problem. Okay, so I've got my perception, attention, and learning uh there. We have a belief about the joint uh state now. Okay, so we've got the hidden states, precisions, and parameters. So our our approximate posterior is going to have to range over all of them. And there's a question about how we can do this. um what should our approximate posterior belief look like? There's all kinds of interesting um answers to that. One answer, one very kind of simplifying assumption is that we can say well my belief about the joint state hidden states, precisions and and parameters can be broken down into the product of belief about hidden state times belief about precision times belief about um model parameters and that's called the mean field factorization assumption. Okay. So this is this is saying I believe that hidden states or you know my my my joint belief about all of this these combined hidden states breaks down into the product of the the individual ones. Okay. Now that may or may not be true in the generative model. Okay. Um but it's a simplifying assumption and it helps simplify the maths. We don't actually need to make this assumption for generalized filtering. uh but it is it's nevertheless an important assumption that is made in in certain other um places. It can be made here as well. So it bears commenting on. Okay. So yeah, we shouldn't claim that the mean field assumption uh means that the variables have no relationship to each other in the generative model. Uh it's really an assumption about what we think our posterior belief is like. Okay. Uh they can still be coupled. So you know x can still be coupled to theta can still be coupled to zeta indirectly. Um that's yeah so just a little bit of a clarification. Okay now problem we know that if we're going to do hidden state inference we need to minimize the variational free energy okay or the precision and we're going to operationalize the variational free energy as precision weight prediction error. But that's only accounting for one part of the combined hidden state belief now. Okay, we've got hidden states, precisions, and um parameters, and they're all changing at different time scales. So, this seems like a a real problem. And the way that we're going to kind of solve this problem is we introduce rather than considering the variational free energy kind of, you know, at each point in time instantaneously, we're going to consider the kind of accumulation of the variational energy across time basically. So from now you know sum it up for however long we happen to be going for. And this thing has a name. Uh it's called rather unhelpfully the free action. So it's the sum or in this case the integral of the variational free energy across time. And this thing is going to allow us to capture all of the kind of temporal dynamics at issue. So it's we're going to have the hidden state transitions precisions and parameters be able to be represented in this formalism. The idea is that this this other quantity called the the free action. This is kind of the fundamental thing. Now it's the the sum of the integral of the variational free energy. Very important point. Uh free action. So this we're calling this thing an action. This is completely distinct from what an action is uh given by the the agent. So the agent emits actions, right? And it receives observations. Uh this is completely different. The free action has nothing to do with action in that sense. where it comes from the name you know free reaction this comes from a branch of physics um Hamiltonian and lrangeian dynamics but they're completely different things don't don't mix those two concepts up very important all right now so we've got the free action we've got an enlarged model we have a state space that's ranging over kind of three canonicalish time scales um how are we going to do the story of minimizing variational for energy or in this case the free action across time how we going to do inference learning and attention. That's that's the story. The crucial thing is to understand that well look the variational free energy. Uh if I if I take the derivative this is just a fact from calculus. If I take the derivative of the free action cross time that's going to cancel the integral and I just get the variational free energy back. So the variational free energy is the derivative of the free action cross time. Excellent. And I know how to minimize variational free energy. We've done that in chapter seven and six. So now we can start to think about how to operationalize what we're doing in our combined model. Yeah. So VF is an instantaneous objective. Free action integrates that objective over time, you know, given some path is the idea. Uh caveat caveat. So uh well we want to do gradient descent on the variational free energy. Um and we can do that by means of talking about how it you know its relation to the to the free action. So this is on page uh 184. We can talk about the motion of our belief about parameters. We can define that as the gradient or the negative gradient in this case of the free action with respect to the parameter at issue. So we can do that for parameters and precisions. And this tells us how the uh the parameters and the precision beliefs are going to change across time. And of course we have a negative here because we want to go down the variational free energy gradient or down the the path of least action. Um now if we take uh the time derivative of these of this and then also of this we we end up with things that look a little bit more familiar now. So we have our acceleration. So this mu double dot is saying take this is the rate of change of the rate of change of our belief about theta. Okay, it's an acceleration. So we got position, velocity and we've got acceleration, right? This is the negative uh well the gradient of the variation for energy with respect to the parameter and then also we have the same thing for precisions. Excellent. Except now h well this is a second order differential equation. Okay, so we saw the first order differential equation in in last chapter. This is a second order differential equation. Mu double dot, right? How are we going to solve this? Well, that's there's there's a trick that you can use, not really a trick that will turn a second order differential equation into a first order one. So I I go through here and I talk about that trick and I kind of motivate it. I give a nice sort of well what I think is a nice mechanical or physical intuition that basically if you've got any kind of second order differential equation you can introduce a completely new variable and say well in this case let's just say I've got q dot is some function I can just introduce a new variable r and I can define that equal to the first derivative of q and then if I take the derivative uh well and that gives you the first derivative of q if I take the derivative of r that gives me my second derivative back. So now I really only have to talk about r and r dot and those are first order things. So you can end up decomposing the second order differential equation into a system of coupled first order differential equations and then we can use oilers's update rule and we can do everything as we did before. That's the idea. So sorry for um covering an entire semester of differential equations almost in about 30 seconds. Uh that's the sort of situation we're in. So I give a few examples about how to do that and then I say okay well how are we going to apply it to our specific case here right? So we we have our second order differential equations. We don't want to that's too complicated. We want to turn them into first order differential equations. So we just introduce a completely new parameter mu prime theta and we define that equal to the first derivative of mu theta um and then we can take the derivative of mu prime theta and that's going to give us mu prime prime theta lots of bookkeeping and then we can do the same thing for the belief about precisions as well. So then ah very nice we can turn our second order differential equation into a couple first order differential equation and and we can use oilers's update rule that's that's all fine we can do the same thing for the precisions as well so now one further little modification so that's equation this is literally from the book this is equations 8.16A for the motion of the belief about hidden uh parameters and also for the motion of the belief about precisions Uh, excellent. There's one further modification that Sanjie makes. Um, and then 8.17 A and B are he just gives you you can substitute these back into the to the gradient of the variational free energy and get a precision weighted prediction error. Okay, that's that's standard. The um he introduces one further little modification this dampening term. So you know 8.16 a and b they do a little bit more than just rewrite the second order differential equation. Um the idea is that we're going to also in addition to mu prime we're going to introduce we're going to basically say we want mu uh prime plus a damped version of it. Um, the reason for that is that it's it's going to make uh the the the integration of the system a little bit easier to deal with. I explain it here kind of where where that comes from, why we're even doing this. Um, it is a nice modification to make, but I wouldn't spend too long worrying about it. So, this this is kind of the assumption we're making. So you know the acceleration of the parameter belief plus a small amount of the the velocity of the parameter belief is going to be equal to the negative gradient of various energy with respect to the parameter belief. Oh my lord that is a lot of stuff to uh to keep in in one's head. But if we can do this and we can do this, we now have a way of formalizing the whole problem um in terms of the the gradient descent on variation for energy and we we we've sort of seen that story from chapter 7. So now we can actually do the inference uh learning and attention all at the same time. So I go through and I give a mechanical analogy for for like a physical system as well talking about you know where this comes from and I give a a loose correspondence between the variational story we're telling with respect to beliefs uh or sorry over here and then the this the physical story with respect to maybe a ball moving in a in a whale or something like that. So that gives a little bit more context as to where that kind of ma mathematical machinery comes from. That's not in the book. Um that's kind of the purpose of these notes. Uh yes. So yeah, why why do we care so much about turning the second order differential equation into a couple first order ones? The idea is because we want to use Oilers's rule. Um first order differential equations are easier to to actually kind of compute numerically. Um so that's why basically it's for practical reasons essentially. That explains that there. Okay. Excellent. So at the end of this we have our our VF gradient. This is 8.17 uh a this one here. So much as we did before. So we have our precision weighted prediction errors on uh hidden state transitions and then also on observations. And so this is this is the gradient of VF with respect to the belief about parameters. We can do the same thing for hidden state uh precisions as well. Now, all right. So, we got kind of three copies of the same thing we've seen floating around. You know, gradients of VF with respect to hidden states, with respect to parameters, with respect to precisions. We're going to collect them all together now basically into so and to you know, we can say all right good. What's the motion of my approximate belief across time? What's the motion of my variational posterior belief across time? Well, my variational posterior belief is made about it's made up of my belief about hidden states, belief about parameters, belief about precisions, and the two other auxiliary things that we added um relating to the dampening terms for the conversion of the second order differential equation into a couple first order ones. And we have expressions for all of these things. So that's you know we can just substitute that in and we can say look the the motion of my belief overall is equal to this thing here and we can calculate all of this stuff which is very very excellent for us and again the thing that I would want you to come away with is that we have a story about how perception happens how learning happens and how attention happens now and they they they happen over three different time scales. So we got kind of fast changing stuff uh relatively slower changing stuff and then slowest changing stuff basically and it's the use of this free action or the the accumulation of the variational free energy across time that allows us to tell this story now and this this is this is going to unify the stories of you know perception action sorry perception tension and learning that we've seen before. Uh very good. So you can you can do the oil updates now um no problem. So you can do them with respect to the the beliefs at issue and this culminates in algorithm 12 which is the generalized filtering perception sorry the generalized filtering algorithm for perception learning and attention. So it's much more it's much nicely much more nicely type set in the book itself. This is just a sort of huristic overview. Um I think we might have a a picture of it that might be nice. algorithm 12 maybe not okay anyway so I put it there for for reference so then finally enough enough theory we actually get to an example we go to example 8.1 uh so so 8.20 20 is the equation at issue and we're going to say all right well how do we do this for you know learning first and second order parameters for you know in in the generalized uh case basically well actually this this is just a univer case um but we want to do generalized filtering in the univariant case for a really simple model and this is the model we're going to consider extremely simple um state transition dynamics and observation um generation. So this is the this is the generative process specification. The generative model now has as before observation likelihood um state transition well uh prior on on state beliefs and that's given by state transition function. Crucially in this example um the observation variance is fixed across time. So you see you know the the precision for hidden state the hidden state prime that's not fixed okay we do have a belief about it but we don't have a belief about the observation precision that's just sort of fixed and given at least for this example just to make things simp simpler and then uh so that's our model if we run that we get figure 8.2 too, which is much more exciting to look at than equations. All right, so this is the results of of um this example. I would strongly encourage people to try and implement this example. I'm going to try and do that myself hopefully in in preparation for next week. Um but we can maybe talk about it. So this this at least gives you the the core, you know, of um what we're talking about here. So we have our we have our belief about hidden states and we have the hidden state. The hidden state is in is in gray. Okay. And that starts out way down here. Okay. And the belief starts out way up here. So there's quite a quite a mismatch, but across time we see that they very quickly converge. Uh as well as the observation and predicted observations, uh they start out fairly all over the place. And the variational free energy on on the right, this almost immediately collapses to extremely small and stays there across time. But and and okay, so that top row that's effectively chapter six kind of basically um and we're not even really doing any action here either. So you know it's kind of standard chapter six. The additional thing is we we also have beliefs about the parameters and the precisions across time. So the idea is that all right we can see that these actually do end up converging to you know values that are very similar to what they should be. So the actual theta here is uh what is it three that's right observation likelihood theta is three so we end up converging to three across time um and the precision is I think it's set to one is that right a very very small uh it's a bit hard to see on this this plot and then we have the prediction errors as well I think there's a bit of a mismatch for the um notation we've got you know epsilon h and epsilon Um I think that should be epsilon theta and epsilon zeta which would be very fun to say 10 times fast. So I I think example 8.1 that's kind of the central example for getting all of the motivation that and you know conceptual grammar about what's going on what we just did at all. So that uses that implements the story of you know variational free energy minimization with this free action thing with an enlarged model and we have beliefs about the parameters and the the precisions. Now now that takes us to 8.2 hierarchical forms and this is really exciting. Um however just quick stop for questions. Are there any questions at this point? I want to go through chapter eight well you know section 8.2 and then we'll open it up for more questions. But are there any questions at this point? I don't want to just scream into the void if people are keen uh to ask for clarification. >> Just a quick comment about figure 8.2 that epsilon H and epsilon G. The HG comes from the um SPM DM toolbox from Carl Fristen where they use um the H stands for hyperparameters and G stands for something related. So that that literally maps in our case to the um attention and uh learning uh >> okay that thank you for thank you Kobus that makes sense. I was I wasn't sure where that came from if that was just an earlier code um uh transcription but um okay that that's that's no problem there. So and as I say hopefully so you know for chapter 7 I put this is for from last week now I put the um example the the collab example of you know univari active inference for the Bayian thermostat up um I would recommend copying this and then you can play around with it yourself. I would like to do the same thing for example 8.1 at the very least because that's that's a I think that's a crucial example to study. So if you really if you want to kind of grock this chapter and understand what's going on 8.1 is is where you want to go. So back to chapter 8. All right I will press on with uh you know uh section 8.2. Now this is a really really cool section. Um lots of very interesting ideas are coming up here and ideas that we'll sort of see again. Also ideas we've seen before in in chapter five. So the the main idea if we go back to some notes hopefully it's not too depressing to look at all of these these notes I have back. All right, the main idea now is we want to we want to uh be able to tell this story about inference in our enlarged model where we're considering um hidden states, precisions and parameters in a hierarchical setting um not just univariatic. Okay, that's going to that's going to give us the the most bang for our buck. So rather than immediately going into equations the we consider that he starts out by considering the simplest possible case. So figure 8.3 is somewhere 8.3 1.2.5. Oh it's not there. I'll set it. Is that right? Yeah. So you know he wants to show you a Bayian network that looks at the the simplest possible example. Sorry, I thought we had it up there, but uh we don't just now. So, you know, I I do start out with a kind of big picture overview here. So, the thing is we can think about what we did in 8.1 as specifying a generative model that is kind of like one layer and we're going to be able to stack these layers on top of each other to get a more powerful generative model. That's the idea for 8.2 basically. So I do recommend figure 8.3 having a look at that. Um essentially if we go to what we're going to have to assume maybe I have a I don't really do that. Okay. All right. Doesn't matter. Basically the change we're making is we're assuming this is the same sort of structure that we saw before but we're going to assume that these are multivariate now. Okay. So our our observations are multivaried our hidden states exogenous force or exogenous state parameters and precisions. That's the kind of big thing that uh we're assuming now. So the actual structure of the model is going to look almost identical. Um we have so we have you know beliefs or rather prior about the precisions the uh parameters also the autonomous states they're going to have a prior now and it's much more illustrative to we do have exam uh figure 8.4 so let's actually look at this in 8.4 for so what's going on? Uh this this is the um uh what do you call a Bayian network view picture of the model. So if we look at on the left um this is now for a dynamic generative model for obviously we're considering hidden states changing across time and they're multivariate. So things in bold the x the v sigma well okay yx v in bold these are vectors these are multivariate things now so circles denote random variables so x is something we're uncertain about so we have a belief about it gets a circle right and the belief is the state transition belief right up here uh parameterized by the mean being the state transition function and the co-variance being dependent on our precision Um, and so that affects X. The autonomous state affects X. Theta affects X. And X generates observations. So Y are our observations. And Y is kind of shaded in gray because it's something we can observe. Uh we we're we're only allowed to observe things shaded in gray. So things that are white, we're not allowed to observe. We have beliefs about them. Um nevertheless, we have beliefs about how observations are generated. And that's given our observation uh mapping belief. So that's great. Um, and of course basically chapter 7 sort of dealt with this picture where we had um an autonomous state affecting the the transitions in our in our actual hidden states. But we we can now we can now have beliefs about the autonomous state as well. So that's what I showed you in terms of the model. So 8.23 two three. We can enlarge this picture now where you know so we have our generative model over observations and hidden states factorized as an observation likelihood and a prime over hidden states that's been the same all the way through since chapter 2. Same sort of thing again except now we can include an additional belief about these autonomous states themselves uh and we can update our belief about them. So that's going to give us a bit more power basically in terms of what we can reason about in our model. And that just shows up as another um another factor in our model. So we're going to have our model factorized like this now. So again it's it's a random variable. So and it's unobserved. So it's a it's a node and it's not shaded in. Um and but it's parameterized in terms of other stuff now. So it has a mean and that means going to be this at V. Uh and it also has a precision that's parameterized by some some precision and those are just going to be fixeds. Okay. Yeah. So, that's going to get us to I want to go back to my notes. That gets us to figure 8.5 which is like the the atomic unit of what we're dealing with basically. Adding hidden state prior. Yes, we talked about that. Yes. So, you know, we talk about making the autonomous state probabilistic. We have a belief about the autonomous state now given by some in this case fixed ATA and well not fixed the precision about uh the autonomous state. So really the only difference is that we have this belief about the autonomous state in here. Now basically that's kind of the only difference. Um do I want to talk about that? No not necessarily. There there's a bit of a talk about multivariate log precisions and covariances. So the idea is that okay in the multivariate case we don't have variances or precisions. We have covariances and precision matrices. Now so really there's just a a transformation that goes through that. If you're unsure about that I' I'd have a look at chap um appendix C that goes through more the notation. We don't need to be bogged down by that too much though. Um okay so really what do we end up with? We end up with this figure 8.5. Haha, perfect. This is an extremely important figure to understand. So this is the Beijian network for the full dynamic generative model with prior on parameters. So we have a prior belief on parameters now. So that that wasn't a circle before. Now it is. We have a belief about it. And that is itself parameterized by some ATA theta some sigma theta. Okay. We have beliefs about parameters. We have beliefs about precisions. Where are those? They're given here. This should really be I'm not sure if that gamma should be a zeta. I think it should be. I think that might be an RA. What do you think, Kobas? Do you Because we we used the gamma before. I think that should actually be a zeta. >> Yeah, I agree. I I prefer it to be a zeta as well. >> Yeah. Okay. I think that's we'll have to note that as an arata. So, we have a belief about our well, in this case, log precision. I should be careful. So zeta is a log precision. Okay, we have beliefs about autonomous states. We have beliefs about hidden states. So this entire slice up here, this is kind of what I've been talking about this whole time. And of course, hidden states give rise to observations. This whole thing, so our model, it looks like this. And all of so you know, observations, hidden states, parameters, precisions, these are multivariants. You know, they're more than just one thing. Um these are this is our model. This is our generative model for the dynamic case. This is this is kind of the atomic atomic unit. Um what we're going to do is we're going to stack these on top of each other to get a hierarchical generative model. And that's going to allow us to do lots of very powerful things. Really quickly because I do want to leave a little bit of time for questions. What you can do is you can go through and do almost exactly the same maths that we did. So 8.25 25. You end up being able to say well the free energy is the sum of the squared precision squared uh precision weighted prediction errors with respect to observations uh hidden states autonomous autonomous yes autonomous states uh parameters and precisions you can spin up expressions for the prediction errors with respect to observations and hidden states. Obviously our observation function is now a function of hidden states and the autonomous state and the same for the the transition function. You can do the same for prediction errors with respect to the autonomous states parameters precisions and then you can do that story where we get the motion of the hidden state is equal to um this big old vector here in terms of the variational free energy. Okay, excellent. Perfect. Now the really exciting bit is where we go into making this hierarchical 8.2 2.2. Um, fundamentally, this is the idea. So, figure 8.6. Yes, perfect. This is what we're going to do. We're going to take copies of our model and we're going to stack them on top of one another effectively. And you know the reason why we might want to do this is because this allows us to represent prior or beliefs about how things are generated in a way that is going to be more powerful than we can if we don't have this hierarchical situation. So this is we're saying this if we this is just two layers. If we look at one just the bottom layer here, what this is saying uh is that my belief about how hidden states are generated is that there's some autonomous state, there's some parameters, there's some precisions, they combine somehow and that generates a hidden state and hidden states generate observations. But this whole thing if we consider another layer on top of this corresponds to a different set of assumptions. It corresponds to my assumption that the autonomous state is itself given by some pro some hidden state that is parameterized by parameters another autonomous state and more precision. So we can kind of keep going with as much as we care to in in this this process here and then we can just say well look the whole generative model this is very much as we had before except now it factorizes across these layers. Okay so the very bottom layer L equals Z. So we've got L little L going from zero to big L at the very bottom layer. This is actually this is this is a typo. This should be Y the observations Y. Okay, we we do literally observe something at the very bottom layer. Uh but then as we go up in the layers, so we got layer one, layer two, this autonomous state is going to kind of in some sense play the role of observations uh for the layer above. That's that's important. there's there's lots of um uh I would say there's an important detail to to not be confused by in terms of the direction of the the generative direction. So, you know, my generative model is making predictions about observations. That's one way. But when we're updating our model, we're going from the bottom all the way up to the top. So, I spend a bit of time in the notes here talking about that. Uh I don't want to necessarily go over it too much. Uh so we got layer wise probability distributions uh yeah two directions in the hierarchy generation and inference. So gener generating predictions that corresponds to me saying what observations do I think I will get conditioned on a belief about the hidden state right that's kind of my observation likelihood and then inference is well I get an observation in uh what hidden state do I think caused that okay these are two different directions and it's important not to confuse those directions um the layers here these are sort of indexing the the generative direction okay going down from the very top to the bottom And then when you go up that's inference. Uh and it's important it's also important uh to not confuse those two things. It's just reversing the arrows. There's there's different processes that are happening but they're coupled across layers. And the way that they are coupled um I'm just scrolling through here to give you a sense of the fact that there's a lot more to be to be talked about about that. The way that they're coupled is through the autonomous states uh 8.6. You can see you know I plug in one by taking the auto state of another plugging it into the bottom one. So that is the end of 8.1 and 8.2 and I would say the fundamental machinery of chapter 8. Um 8.3 to 8.6 is really just kind of extensions and and different flavors of things. There's lots of very interesting connections between this idea and portical microcircuits in the brain. I'm sure that Andrew can speak to that a little bit more than I can. Let's come back to to here just now. Uh yes, Andrew, excellent time for you to uh to share any any thoughts or anything. >> Yeah, thanks. Uh you know, very good, very nice uh lecture. I rose my my hand earlier, but I didn't want to interject because >> Well, no, no, no. Just uh my question is is something that's more general about this chapter and honestly probably shouldn't really come up too much until next week whenever we've kind of worked through the rest of the chapter. But uh yeah this uh this matter of like hierarchical models which were already introduced uh to the notion of hierarchical models back in in chapter 5 on predictive coding. Um so here it's it's just kind of interesting and a bit tricky to talk about hierarchical models in that in chapter 8 uh in the earlier sections it seems where we're the way we're introduced to them is sort of like um you know your sensory observations your why is as if it's at the lowest layer and then above that at another layer are the hidden state and uh and and autonomous states V. Um whereas later in the chapter as well as in how hierarchal models are um introduced in predictive coding um an a different way of conceiving of these layers is actually the lowest layer contains your why as well as your hidden states and your autonomous states. And then the next level up would be, you know, a similar sort of construction of something more or less along those lines aside from we'd be swapping out the the Y for the V from below or perhaps in states from below as well. So it can get a bit tricky just um in this general sense of the way we're introduced to hierarchical models in this chapter is that like sort of within a layer you can kind of conceive of it as having its own uh sub layers with the lowest layer being your why your sensory observations and the next layer up in the states. So I just wanted to point that out um that that that it can and >> uh the part of hierarchical models is is that it it it's becomes perspectival sort of in a certain way kind of depends on how you're yeah looking at it. And furthermore, I mean, you know, aside from this notion of there being sort of ascending predictions and excuse me, ascending uh prediction errors and descending predictions, um hierarchical doesn't necessarily mean you have to always think of it as from the bottom to the top. Like as we see later in the chapter on all these things about message passing like you could have rather arbitrary constructions where you have what seems to be a lowest layer but actually it's you know you could have these sort of lateral connections other layers and the rest. Um it can be very fun but also a bit daunting to to think uh all of it through. Um but but yeah, so I just wanted to um yeah, point that out. And then again, uh later in the chapter is probably where some things might get clarified, but uh more generally um it's in this chapter that we see the most on um message passing. We get much more detail about these sorts of kind of network and graph type illustrations we've been seeing in the textbook. So um yeah very >> quite visual very handy. Yeah. >> Yeah. Absolutely. I think it's very worthwhile for for readers to um get a sense of for those who are new to these kinds of things because many of the illustrations we've seen so far like typically the the the caption will say like a basian network which is one way of conceiving of of these kinds of probabilistical models as probabilistic models as graphs. Um but then there are other kinds of you know graphical structures that we could use and there are fory factor graphs and the rest with their own respective terminology. So it's worthwhile to spend just a little bit of time thinking uh with with graphs. You know there will be many people who are familiar with basian statistics and the like but they're not used to saying these sorts of like relatively complex graphical structures. Um and so just being familiar with like what is a node you know what is an edge? um those sorts of things. Um oh and yeah, thanks for uh bringing up the slides. I tried to give a little bit of that in a couple new slides I I added. Um this is just, you know, we see triple estimation where we've now introduced attention. Uh and then yeah, more on the the way that we're conceiving of sort of these error units and state units and all things that kind of uh succeed uh section 8.2. So probably next week sort of level material. Um and then uh there's a an attempt to explain algorithm 13 which is definitely a end of chapter kind of thing but I wanted to have some yeah just a sense of this is sort of what we're all building up towards. is like we have the hierarchical attention precision uh modulation and and we have learning and we have action we have so so algorithm 13 is a really nice way to um round off all of this fundamental material we've been seeing on continuous state space models um I think we're about out of time and we have a whole another round of lecture next week so I'll stop there but I I think it was very important the way that you covered uh all these fundamental that we're getting into and thinking about uh layers and and the hierarchies, thinking in terms of graphs. Um being able to understand the relationship between sort of uh V and sort of where it it sits in all of this since we know that V is uh new. That's the that's one of the newer parts aside from attention um you know in these past couple chapters. So yeah, thanks. >> Yeah. No, thank thanks thanks for that. And I I I must I didn't you did in in your in your slides, but um that problem is called the triple estimation problem because we're estimating three things. >> So that's an important thing. Yeah. >> All right. Well, I look I mean if there's any last questions for the recording, um that would be excellent. But I I might stop the recording here if no one no one does and then they can ask questions um otherwise. So going once, going twice. Vera, please. I don't know if it's actually for the recording, but uh if uh so what I really liked is that um this hierarchical um uh sense of um okay there's a the different temporal affordances >> aren't actually an implementation detail but they determine what counts as perception, learning and attention. in the first place and actually without those time scale separation those categories uh would collapse into just the same operation I guess and um so my question is if perception learning and precision precision learning all minimize the same variational free energy objective but operate on different time scales. What determ minds which level absorbs a p persistent prediction error or yeah I don't know or in other words um >> when does an error remain a perceptual update and when does accumulated error become evidence that the model parameters or precision themselves should change? That's my question. >> Excellent question. Yeah, I mean uh I think Andrew will have stuff to say on that there. But really quickly um if we just go to the to figure 8.6, this is not answering that question. Okay, so this notion of hierarchy um we might imagine that there are different hidden states and they change according to different time scales blah blah blah. We could imagine that. That's kind of what this is getting at. But your question is about well when does the belief about this versus this versus this versus this get updated. Um and there are that's that is quite a deep question. Um we basically we've only really seen a hint of that in the in the sort of framing of this whole process um with respect to the free action. So the the key the thing that sort of allows you to talk about all of these time scales is the free action the the accumulation across time of the variational free energy. I did have it here. Um so yeah like when exactly you should update your belief about um hidden states as opposed to parameters. This is it's it is a bit of an uncertain question although in in certain you know biological applications it can be it can be a little bit clearer or you you you have some principled way of thinking about when you should do that. So there's actually been a lot of work in isocarts. So uh so this is you know where are you sort of looking uh your your gaze you know you can decide to attend to certain things but then your eye is while you're doing that performing these very very small fine motor motions that are that are trying to resolve uncertainty at a at a much smaller layer. Um so exactly you know and that's that's at the level of sort of milliseconds your your gaze is at the level of maybe seconds. So there's if you wanted to model that process there's at least some empirical separation but determining the causal and structural reasons why is is is an open question. Um, I'm actually less competent about answering that question myself other than other than to say that there are, you know, principled um, examples from from biology that can that can help answer that. So, I don't know if if Andrew you had anything to say on that specifically. >> Yeah, thanks. I uh yeah I mean it's a it's a tricky question uh and and I want to make sure that I'm understanding it well enough but um because I I mean I if you remove the time the temporal separation if you remove the time scale separation then sure you still have if you're still using the same uh sort of equations um it depends on how we want to conceive of the time scales the the time scale separation can be because you wrote your algorithm to only have particular things update, you know, every other time step rather than every time step. Like you could have parameters updating, you know, every second time step whereas in states uh, you know, update every first time step or you could have some other arbitrary set of rules here in the continuous example that we're getting in chapter 8. It's a little bit less about doing that sort of thing and it's more about the way that we're treating these first and second order parameters. The way that we're applying these different temporal derivatives to where you know if we apply uh if we include additional ones for parameters that allow us to have this sense of uh even further change over time to where there is this sort of uh conceptual like slower time scale that they're operating at. assumptions sort of baked into how we're applying those um derivatives and including them in our model. You know, that that's all fine. But but um >> so so that I mean you know with the time scale separation a lot of that is just drawing from sort of empirical studies of of the brain and information processing and and you know all these phrases of recogni recognition dynamics that themselves precede specifically active inference. uh that it's a much broader term uh just as you know basian uh statistics goes much further back um but with respect to uh prediction error and and minimizing free energy I mean a lot of um where that sort of settles is really going to to depend on the actual simulation at hand right because if you have um you know if you have a relatively predictable uh environment where there are no sort of regime changes is at any point where the the the the the environment becomes like virtually, you know, different like the world's very different whenever it's raining, right? Like many things can or let's say storming outside, you know, maybe there's some awful storm going on outside and suddenly stores are closed and no one is on the street and there's a lot of danger for people to walk around versus whenever it's not storming. Like that's it's a very different kind of environment to navigate, right? So you do need that um you know ability to like you know almost confront a a very different world. But if you had a relatively static world uh that that you know perhaps it has certain kinds of of rhythms in it that need to be attended to that you know you do need this sense of temporal derivatives uh in there to be able to catch things that that change and catch the sorts of patterns that are going on but it's still relatively static. you would very frequently ideally see very low uh variational free energy. You'd see very low prediction errors assuming you have a good model. Um whereas if you have uh an environment that like drastically changes that calls for different sets of actions and different kinds of hidden state inferences and very different observations elicited from the environment in sort of one regime or another. Then you'll see prediction errors very often very reasonably, right? Because it's it's the prediction errors themselves sort of motivate an agent to adapt or change its sort of behavioral behavior or the way that it updates uh you know V for itself so to speak so that it can adapt to those changes. So I I hate that my answer is really just it depends on the circumstance at hand. Um, and what I sort of got from the question is, you know, I I I say all this because it's important to remember that prediction errors are not necessarily uh, you know, them in and of themselves like bad or or exemplify a bad model. like a a model that's able to use prediction errors to its advantage to properly learn uh is in some ways qualitatively a better model than say a standard forward model that you know maybe it has prediction error but it's like oh I'm you know in in plainst terms a simple model that says oh I messed up I'll do better next time but there's no update to my parameters and there's no update to other things that need to be updated where I could use those prediction errors as and uncertainty as information for updating my model. So I I could kind of go on, but you'd see it's getting a little more conceptual. Um but I just I I hope that it's understood that the I I don't know if there's a very clear specific answer to that kind of question. It might be the way that it's being framed. >> Yeah. >> Yeah. Thank you. Thank Yeah. Sorry, Razer. >> Two things two things on that. Ver like just quickly chapter 8 doesn't purport to offer an answer as to like well when do I update my beliefs about precisions or parameters or hidden states it's just saying that these are there are three characteristic time scales and the updates are in some sense kind of given by the dynamics at issue um but you could absolutely decide or or you know intervene at c certain times as opposed to others um like okay well I want to update my precision you know every every single tick or something like that. That could be something you do. But chapter 8 is not trying to do that. It's just kind of showing us that it's possible to frame the question of inference more generally as this larger problem within which we can do, you know, free energy minimization to solve uh for beliefs across those time scales. But you could you could you could decide to intervene on on some specific time scale if you really wanted to. Um I put in the chat that question though about like okay well when when should I update as opposed to not that's very relevant to one extremely hot area of research in active inference now called structure learning um so there's a paper there by by Lance and a few other people um that's that's if you're interested in that question while they're like when should I update uh beliefs about hidden states versus parameters versus the structure of my model uh that that comes up there in that issue of structure learning. Cool. Thank you. But uh sorry to be so conceptual, but I think I got it. It's like just a question uh which part of the generative model um should be plastic and on what time scales. I think um yeah. >> Yeah. Yeah. And of course typically for a lot of interesting problems, it's precisely the fact that we have this hierarchical structure that's uh for a lot of problems that we want to be able to model things hierarchically. Um so we have that hierarchical structure but then we also have within a layer updates at certain times about you know precisions versus parameters versus hidden states. Absolutely. >> Also add on the Oh, sorry. >> Okay. Thanks. Uh yeah, on the the matter of which parts of the model should stay plastic, I mean it's also a good question also very going to be dependent upon the the situation at hand. I mean in some ways uh it would almost from one view it would almost be um you know beneficial to have literally everything be plastic to some degree and that's perhaps what's more empirical. Granted, maybe the way that that plasticity is employed in the model should maybe be aligned with different kinds of um empirical studies or otherwise using neuroiming looking at patterns and then think you know um basically like how in the same way that neuroscientists may have arrived at the idea of there being different time scales operating in the brain. Um you know based upon that sort of research perhaps whenever we apply a model we should be doing the same thing. But it's um you know if if you're going to employ you know a model that's just massive like the let's say we're in a purely computational space here. We're not concerned with empirical plausibility. We're not concerned with creating like an adaptive agent in an environment uh you know in any kind of um pseudo organic sense but rather something that's going to crunch a lot of numbers for us. uh you know some sort of stock price prediction problem or or or uh uh something along those lines. It's just going to the idea is that we'll be ingesting large amounts of data such that if we were to make every single piece of our model plastic that's going to mean that we're going to be implementing many many many update rules every single time step and we may not have the computational hardware necessary for running that kind of model. It might be in that case that you might start to sort of prune back the degree of prob probabilistic parameters in your model. You might say that we will have uh fa we'll have our parameters and precisions be purely deterministic and just static throughout everything and just hope that inference itself is what's able to catch and make better predictions from step to step. And then that way we can sort of reduce uh you know our computational hardware necessary for for the problem at hand. Um but yeah, it's it's sort of it's sort of there. It's it's going to sort of depend upon your domain you're working within essentially how much are you trying to capture empirical phenomena in relation to the brain and whenever I say the brain of course that's going to come back to which which brain are we modeling a human brain a mouse brain or the mamalian brain is generally there are many shared principles which is why uh some people will speak so broadly and why so much research with mice and the rest has informed human studies as well especially in cases where um the ethics would become more questionable if we're to carry out studies with with humans in certain circumstances as opposed to to mice. But um yeah, it's I I think I'm droning a bit. Sorry. So, I'll I'll just kind of leave it there. It it's still a good question, but it's going to come down to sort of your experimental design that you're employing and and your problem at hand. Yeah. >> Yeah. I'm just answering a few questions in the chat here that's been quite active. Uh so, there was there was a question by Jen Cuomo. Uh, is there a trade-off between changing parameters and minimizing VF sort of too much as it were becoming too rigid and more prone to failure and death? Um, I mean there's always the perennial problem of, you know, overfitting and underfitting, but the VF itself is meant to be this, you know, an approximation to a quantity that's telling you how well your model is doing at uh, sort of predicting or or separating itself from its environment. So I I'm not quite sure how to formalize this exactly, but it wouldn't be the case that you could minimize the VF too much. Um in the perfect case, you would have like zero um VF and you would be able to exactly approach the the log evidence uh bound and everything would be great. But um we obvious we can't do that for computational reasons. Um so it's not simply the case that like you can make it too small. uh although we have seen of late especially Andrew you mentioned this last time that it can be useful to have nonzero prediction error in some cases in the sense that this can give you useful information um so you know obviously the the precisions themselves these these are as we've seen these are likened to attention they can even be they might even be able to formalize what attention is so in that sense it can be useful to not always have zero prediction error But uh that's that is a that's quite a large um area in and of itself or a large sort of you know area of interpretation. So >> yeah. Yeah. It's it's it's so much of this is going to be highly dependent upon like well what is the environment that the agent is in? Um are all of the agents sensors uh so to speak that it actually has to receive sensory observations? Are they sort of all up to the task? Right? Um, you can also think of a scenario where that you have an agent who has uh, you know, it it it pays much more attention to sort of its priors on hidden states as opposed to updating them. Right? So, you could have um I'm sorry I come up with all these strange examples, but uh you know, if you if you were in the woods and uh you know there's a pretty aggressive bear running at you, um you could just sort of pretend that the bear isn't there. Like, oh no, I'm watching a movie right now. I'm okay. You know, and then if you were watching a movie of a bear, your model would say, "Oh, that's fine." Like, this is exactly what I'd expect. I'd expect, you know, bear in front of me because I expect a bear because that's what my prior say. I'm just, you know, convincing myself that I'm watching this movie. Now, we've brought VF way down. Like, that's great. Like, I'm going to survive this, of course, cuz it's just a movie. It fits my expectations of a movie. Um, but then you're not in a movie. So, then what will happen is that your VF will dramatically rise whenever, sorry, something something probably bad happens from there, right? That that doesn't align with being in a movie. So, so if you that's a very this is a very distinct rather violent uh uh uh example I came up with but but within the context of that like sure you could you could have a situation where VF was minimized uh you know sort of inappropriately but the the only person who could say that was inappropriate would be like us the external observers of that situation if the if the agents parameters are liable to just when I feel unsafe, I just pretend I'm in a movie and dissociate. So now this this it's a slightly realistic example like this does relate to like psychiatric phenomena that you could start to use the phrases of priors and parameters and such to align that as sort of these computational substrates to things like dissociation and and and hallucination and other sorts of things. Um but uh yeah um so yeah it's it's really going to be dependent on the situation at hand and once again like if I am hungry and I need to eat food like I need that prediction error to inform me I need to change my action. So prediction error is not always bad in a universal qualitative sense nor is having somewhat higher VF. It's more like, oh, this is the these are these are errors that are telling me that I need to make some kind of correction, and if I don't make that correction, I'm not going to adapt to the situation at hand. So, yeah, it's >> information. Yeah. And we'll see that we'll see that with planning uh later in the in in the book. I'm gonna end the recording here, guys, and then anyone who wants to ask a question on recording can. So, goodbye to the YouTube people. We'll see you next week. So, there we go.