Submind YouTube summaries
Thumbnail for Fundamentals of Active Inference (Chapter 8, Session 38) August 21, 2026

Fundamentals of Active Inference (Chapter 8, Session 38) August 21, 2026

Watch on YouTube

Video summary

This session concludes Chapter 8 by formalizing the "triple estimation problem," where agents simultaneously infer hidden states, learn model parameters, and adjust sensory precision within a hierarchical generative framework. The core mechanism relies on message passing between layers via state units that encode predictions and error units that encode prediction errors, a process driven by gradient descent to minimize variational free energy. This architecture extends naturally to the multivariate case using generalized coordinates, which represent variables through their position, velocity, and acceleration rather than direct beliefs about future time steps, thereby allowing for the extrapolation of motion based on current dynamics. The comprehensive computational framework presented in Algorithm 13 integrates action selection, perception, attention, and learning into a unified system that operates across continuous state spaces. Within this hierarchy, prediction errors propagate bottom-up while predictions flow top-down, facilitated by excitatory and inhibitory connections between state and error units that provide a mathematical basis for cortical column interactions. Learning is specifically defined as the updating of connection weights between these units, corresponding to beliefs about model parameters, while the active inference loop functions through a forward pass of sensory evidence and a backward pass of prior predictions to continuously refine beliefs about hidden states. A critical distinction is drawn between continuous variables, which describe flow and velocity, and discrete decision-making processes that will be addressed in future chapters involving categorical choices like "go left" or "go right." The discussion clarifies that generalized coordinates represent derivatives of motion at a single time step rather than future expectations, resolving notational ambiguities regarding how current states relate to their flows. This section also outlines the transition toward discrete state space formulations and highlights the flexibility of the master algorithm, noting that simpler models can be recovered by disabling specific features such as learning or generalized embeddings. Looking ahead, the session sets the stage for Chapter 12, which aims to integrate these continuous and discrete models into a cohesive theory. The speakers invite collaboration on developing a library for hierarchical active generalized filtering and direct viewers to Google Colab notebooks featuring practical examples, such as the Bayesian thermostat, for further experimentation. Ultimately, this chapter establishes a robust foundation for understanding how complex agents navigate environments by balancing immediate sensory input with learned priors and dynamic motion extrapolation.
Read the full video transcript
All right. Hello everyone. So we're here at uh the last section, last session on chapter 8 on fundamentals of active inference. It's August 21st, 2026. Um this is a pretty watershed chapter actually and our our culmination, our finishing of it here is is marking a transition where we are in the book um to something that's going to look maybe a little bit different to what we've actually looked at for everything so far. um we're we're looking at we don't really know it yet, but we're looking at the end of the continuous um state space formulation of active inference where we've had all smooth distributions and things like that. We're going to be moving into the discrete state space formulation uh with chapters 8 and nine, sorry, nine and 10. We're finishing chapter eight just now. So, what we did last week is we we went through sections 8.1 and 8.2 on uh chapter eight. and they're kind of the conceptual core of the chapter really in terms of what's the major idea and indeed the major idea in chapter 8 is this thing called the triple estimation problem where we're learning you know we're doing inference of hidden states we're learning um parameters and we're also learning sensory precision we're sort of doing these three things all at the same time in some sense that's kind of the really high level 30,000 foot nutshell overview sorry for mixing my metaphors of chapter 8 basically and then there's various takes on how to do this uh throughout the chapter. So with respect to 8.1 and 8.2 we got to the where we got to was the look you know formalizing this whole thing hierarchically. So if I share my screen here uh let's just do the single tab. There we go. This is effectively where we left off at least in my session. So we got to this notion of what are we looking at? If we just look at the bottom layer, we've got, you know, this kind of star-like situation here. We've got parameters which influence our hidden state. We've got log precisions. They influence our hidden state. And we've got an autonomous uh uh state which also influences the dynamics of a hidden state. Hidden states give rise to observations. We only observe observations. So this is shaded in um gray and everything else is a something that we have a belief about. It's a random variable. So there's a lot of you can see this is very richly notated. There's lots of stuff all over the place. You know the belief about parameters itself. There's a mean and a precision or coariance that parameterized this. And we saw that we can kind of stack this these motifs on top of one another uh whereby the each layer is coupled to the layer above or to the layer below depending on your perspective by the autonomous state. So the idea is that the autonomous state at say layer two couples to the autonomous state at layer 1 and that's how the sort of states across the layer influence one another. So the way that we talk about this mathematically is we say that the whole generative model this entire thing is factorizable in terms of the layers. So we have our expression for our generative model uh now in terms of the the layers at issue. And then you can spin up notions about okay well what is the probability density of hidden states at layer L given autonomous states at layer L at some kind of normal distribution and we have our transition model and then you know prior on on uh parameters and precisions. The thing that allowed us to do [clears throat] all of this learning at the same time uh was effectively this notion of the action. So the integral of the variational for energy across time. So these are my notes here, latte notes. We now have them for the whole chapter, not just um 8.1 and 8.2. That [clears throat] was kind of the thing that in some sense allowed us to do this whole whole story or tell tell the story about how we do inference learning and precision or rather attention at the same time. It's an 8.2. Sorry, I'm not going to be able to get it up. Anyway, we went we went over it last time. That's where we left off. So picking up now with 8.3. And again, if you look at the page here, so we got the chapter map, we have the the PDF notes for 8.1 and 8.2 there, there as well as all of the other ones now. So 8.3 that you know the end of 8.2, we've looked at formalizing the the story and you know constructing the model. We have this notion of a hierarchical uh generative model now coupled across layers. So we can factoriize the model in that respect. How do we actually do the dynamics? How do we do um inference in this? Now that's what 8.3 is all about. We've kind of set up everything and now we want to do inference like we want to solve the triple estimation problem by means of what we've set up already. So that's you know 8.3 really does assume that you've embied all of the details and so on from from previous sections which I know is a bit of a tall order given the the rate with which we're going through things. So effectively where we start out the the thing that we want to have kind of you know in our mind at the end of section 8.3 is he's going to introduce this idea of uh going to begin to introduce this idea of sort of message passing between the layers at issue. Um and that's going to be talked about in terms of kind of two different kinds of neurons inverted commas. uh if we're going to you know uh anthropomorphis or neuropomorphize in some sense that's kind of the whole thing about section 8.3 is that we can actually tell the story about how variational energy is minimized in terms of the way or the pattern of message passing between layers and between neurons. So the way what we need to do to get there first of all is we know that we have a notion about how the the motion of our hidden state belief changes across time. It's gradient descent on a variational free energy very much as it has been not only for hidden states but also for the autonomous states now and those are the things that are coupling between layers. Um we and and we also have we we've told this story many times before about the precision weighted prediction errors with which you can formulate the variational free energy. Okay, so this is from chapter four all the way to chapter six effectively. Um but what's going to happen is it's going to be a lot easier to to talk about message passing if we do just if we just sort of look at things a little bit differently or formulate things a little bit differently in terms of the maths. So essentially all we do is we take the usual expression for the gradient of the variational integer with respect to hidden states and we want to talk about precision weighted prediction errors. So we can define that very easily and this is this is this is as has been done in the in the chapter already and we're going to talk about these precision weighted prediction errors is you know just as a single variable now uh very nice. So that gives you 8.38 a and b these these equations here uh or rather you know you can turn this which is what we've seen previously into this and it's just a rewriting of the same thing really. So and the idea is that it just makes clear that the the structure of the update to your model is given is you know it's your sensitivity to observation predictions this thing here. So what's the change in my observation generating function with respect to hidden states multiplied by the the weighted by the observation error. Okay, the prediction error in some sense and the same for hidden state predictions and their prediction errors. Just a bit of bookkeeping really. So what this does um so on page what this is page 194 now the idea is that this is kind of giving us two two interesting little units of computation and those units are going to be talked about as if they're like neurons. So we kind of have we have state units things that encode predictions about hidden states and autonomous states and we also have error units and those are going to give us our prediction errors. So I'll show you a diagram in a second. Um and that's so you know we we and we're we're all we're also operating in the um the hierarchical case I should mention for 8.3. So we're assuming familiarity with you know the um uh the generalized state space formalism. That's that's where we're playing and that's going to come up and and cause some issues actually uh for us in 8.5. So what does this end up looking like? effectively 8.7. That's really the thing to look at. So, this figure here, by the way, uh I can't actually see while I'm sharing if people have their hand up. So, if they do, um if you do want to ask a question, uh just kind of shout out and then we can uh we can attend to that. So, figure 8.7, that's that's what this is what is representing everything that I've been talking about so far. So, we're moving from the picture on the right, which we've seen. This is kind of the motif from 8.1 and indeed from 8.2 uh where we ended up with our hierarchy. We can express pretty much the same model in terms of this the the interplay of uh our state state neurons and error neurons at a particular layer. So we're looking at just you know one layer in the hierarchy. How do we rewrite our model in terms of the interplay of state and error neurons? The reason why that's nice is because it allows us to see or formalize the way in which prediction errors are being passed around and the way in which predictions are also being passed around. Those are the two fundamental things that are happening. We've got prediction errors going one way and predictions going the other. So of course at the bottom of our hierarchy we know we observe stuff. We get stuff from the world. That's why the idea is that our state um state units these are going to be ob you know we've got our observation we've got a belief about hidden state and a belief about autonomous states right those are the states at issue but we also have prediction errors about exactly the same thing so prediction error on observation that's going to take in a prediction from the hidden state as well as the observation and it's going to pass a prediction error up the hierarchy basically so we're sort of looking at it as if it's like a felled tree. It's on its side here. The the left hand side is kind of the interface with the world. You I'm literally getting an observation in Y here. And as we go up the hierarchy, kind of, you know, deep into the brain. We're going from left to right is the idea. But crucially, we're only looking at this at one layer in the hierarchy. There's kind of repeated units going on. All right? And that's that's the fruit of section 8.3. We're kind of moving to this picture about representing the Beijian uh well the uh the ba the Beijian graphical model whoops that we've seen before these kind of pictures to a picture where we're representing the communication of um errors and and predictions to one another. It's just going to sort of help us track things a little bit better. That brings us to 8.4. And that's really where I would say 8.4 for is very much the core of like payoff of um this this way of thinking about prediction errors. So you know that is the complete active generalized filtering algorithm. So we're going to formalize the whole thing now and we're going to do it whole hog. There isn't actually an example in this section which is a little bit sad. Of course we could scare off an example especially based on Andrew's excellent code. But what what's the idea here? So we we want to put all the pieces together now and do active generalized filtering uh for the multivariate case. Okay, that's the most general case really we could think about. So let's start with our model. Do I have the representation of the model here? I don't know if I do. 8.39 is the representation of our model. There we go. So we have our generative model. And crucially now everything is generalized coordinates. So we have generalized coordinates representation for hidden states, autonomous states, not for parameters or log precisions. Okay, that's very important. So we're making an assumption here. We're making an assumption that the hidden states or sorry the the parameters of our model and the log precisions, they don't really change very quickly, at least nowhere near as quickly as the hidden states and the autonomous states themselves. So we're just going to pretend that okay they're sort of fixed in some respect. That is an assumption that we're making. Um and it's an important assumption to know about. So I mean this is really just exactly the same thing that we saw at the end of 8.2. We're factorizing a model across the um the hierarchy. Now uh just a little thing on sort of indexing. So, you know, sometimes I I notice sometimes in the book we've got, you know, going from zero base L, you know, in our layers. Bottom layer would be the sensory epithelium layer zero all the way up to layer L minus one. Um, sometimes it's it's actually mixed, especially when it comes to uh notating the the index at the very bottom layer. Doesn't matter too much, but everything should be in your mind zerobased indexing, especially that's that's useful for for Python coding examples. you know that there should be minimal friction going from the maths to the to the coding. So all right now we can spin up our sort of Gaussian factors. All right. So each each factor in our model this is some Gaussian probability distribution some continuous Gaussian probability distribution. Okay again that's been with us for the whole time and we're going to be giving up that assumption about uh continuous distributions going forward. What does all of this look like? And we have I should I should mention before we have a look we have a very very very top level prior on the autonomous states and that just is saying well look at some point we have to terminate our hierarchy whether that's after two repetitions or three repetitions who knows but we're going to have to have some top level prior on what the autonomous state is ultimately generated by and it's some fixed mean and some fixed coariance you know that's really all that's saying so what does this as what does it say what does this look like It looks like figure 8.8. Here's what we have here. [snorts] [clears throat] All right. This is an important figure. It's also a little bit uh there's potential to understand this in in a way that's not too helpful. So, we're looking at really the same thing we saw at the end of um 8.2 except now it's in the generalized state space formalism where we have uh it's been you know it's expressed in terms of generalized coordinates. So you see each motif. Okay. Uh but then you know vertically and at the bottom here we've got observations coming in and then horizontally we have our embedding orders. So what we're looking at is you know this whole thing here is my generalized observation. This whole thing here is my generalized belief about hidden states. um generalized beliefs about autonomous states. In the very uh in the very beginning, I actually think this should be these we shouldn't have beliefs about the sensory sorry the the uh precisions. Uh we're not we don't have that in our model. So I think that might be an error. But so this is just to say that we're really just taking what we had and we're now expressing it in generalized coordinates. Um I I will note it's a little bit confusing in terms of the indexing here. So we've got the present time right you know tow some tower is equal to the present time versus one moment in the future versus another moment in the future. I don't want you to get the assump the you know have in the idea in your mind that the embedding order is somehow an index into my belief about hidden states or autonomous states or whatever at the time that time in the future. This is a bit confusing. So really if I if I go back to the notes, sorry if I'm not making much sense, but it is an important point 8.4 let's just think about what generalized coordinates even are again. Okay, so you know my hidden state is now the hidden state as well as the first derivative as well as the second derivative and so on. And you can kind of think about that as you know abstract position, abstract velocity, abstract acceleration and so on. Um the embedding order is just where am I in this in this list here. So this would be embedding order zero, embedding one, embedding order two. Let's think about what that means. So the idea is that of course I have beliefs about this whole thing. Okay, those are themselves generalized coordinates like this and that's because we use the lelass approximation. Um we can just worry about the means. So that makes our life very helpful. Do I have I don't have a nice representation of that. I don't think I do in in previous notes. So what what that is to say is that at at some time like let's say in the current time I have a belief about the hidden state and its motion and the motion of its motion and the motion of its motion of its motion and so on and they allow me to for some time into the future uh extrapolate what I think the hidden state is going to be for a bit of time in the future. Okay, it's not the same thing as literally planning what the what the future would be. I'm saying at this time the motion of my hidden state is such that I think future states are going to be blah blah blah blah. Okay, but those those beliefs are always it's always on the basis of current beliefs. Um and the the um the embedding order is just okay well which one of these are we talking about? you know, maybe we only care about the first three embedding orders and usually we do and it's not the same thing as, you know, talking about your belief about the hidden state at some time into the future. So, I'll leave it I'll leave that be for a sec, but um I did want to I did want to clarify that. Actually, let's go back up for a second. All right. Now the the rest of the the rest of the story in terms of how we do our um our our variational free energy minimization very very similar to to so you know we build up our factors in our model I don't want to go through the details there we kind of already know that unless unless you want to um yeah so crucially we we've got hierarchy different model different layers in our model and dynamics uh are you know what's at some time I've got a representation of the hidden state it's its motion the motion of its motion um and that's what that's what we mean by the dynamics basically um [clears throat] right here and of course at the sensory boundary autonomous states are just our observations there's a bit in here about um generalized precision and uh you know the coariance matrices at issue I don't want to necessarily dwell on the linear algebra um but of Of course the precision matrices at issue they depend on our log precisions. They also depend on this gamma factor which is talked about a little bit more in the the mathematical appendix. So do do have a look at that. Have a look at the appendix on chronica product as well if you don't understand what that is. Not too important to get very bogged down there. Um but we end up with actually I think we do have a representation. Oh there we go. Yeah. So there's we talked about the correlation temporal correlations in between embedding orders in um chapter six. That's that's what gamma is. I I remember now that's why I write these notes. So at each level we have variables that we're inferring. Okay. And we saw that in the uh the Beijian um graph model you know we've got beliefs about hidden states, autonomous states, parameters and log precisions. Those are what we are inferring in general and we can just package them up and say well you know I let's just call this some this this is called a var theta here not the best choice given that we also have theta so we'll just package all of them up and say well you know this whole thing is v theta that's what we care about um inferring but then of course at hierarchical level you know l uh the posterior of interest is written like this so we've got the posterior at level L given um the the previous uh level essentially very nice and yeah this this is the point I was trying to make before because we're making the class uh approximation we can just represent the belief about the hidden states as the mean only that's very nice but we're in the we're in the hierarchical uh or multivariant case so each of these are um generalized coordinates so this is this is almost like a table now basically so This whole thing is represented, you know, mux at layer L. That is some list of, you know, however many embedding orders you want. Maybe you want three, maybe you want five, as many as you as you care to compute effectively. And that entire thing is now our hidden state at layer L. And this whole thing is happening within a a hierarchy. So it gets pretty crazy to to to keep on top of sort of all the indices and things like that. But what is nice is that this whole thing has been done before. This is exactly what we did in chap uh 8.2. It's just that we now have the multivariants uh case. So this is not anything different uh to what we've seen which is is nice but it is a bit confusing. So as as we saw before we could you know we can de decompose the variational for energy in terms of the precision weighted prediction errors at each layer. Now very important at each layer. So it's local to a layer and then we can write the equations for well you know we know the the uh contribution to layer L it's precision weighted prediction error before it was a bit simpler because of course it was all univariat so I go back 8.26 right so this is 8.45 45. We're looking at the the layer wise prediction weighted precision error for some some one of these variables we care about. If I go back to 8.2, I just want to show you the contrast. So 8.26 is the equation I want. Now it's exactly the same equation. So do have a look at 8.2 if you're lost as to like what's this what are we do we doing here? Am I going to be able to find it in time? Yes. Yes. Yes. Very good. Very good. So, precision weight prediction errors. No, maybe not. Um, I just wanted to show you that the the univariat case again. Oh, here we go. Yeah. So, this is the univariat case effectively. Um, we have precision weight prediction errors. They're not in generalized coordinates, but now we are in generalized coordinates for um 8.4. Whoops. 8.4. That's really the only difference, you know. So, if you're if you're struggling, I would suggest to to meditate deeply upon um 8.2 the mathematics there. Where did I end up? 8.43 as 8.46. Yes. Okay. So, precision weighted prediction errors, we can express the variational free energy in a like manner layer wise. Basically, that's very very that's excellent. Um yes we don't have to talk about normalization too much. So what you end up with is you end up with four canonical precision weighted prediction errors. This is kind of equation uh well 4 six a through d. These are the central equations basically. So we have the precision weighted prediction error for autonomous states. uh we we'll set the same for our our hidden states and then the same for our parameters and log precisions. Those are what we can use to now talk about the the overall flow of our belief in the multivariat case as we did for the the univariat case. So I talked about you know why why is the derivative shift operator necessary here. We did talk about that in um chapter six. this kind of recapitulate it again why we what where does this come from why do we need it at all um I don't want to get bogged down you can have a look at that in your own time or ask questions uh on on the uh the web page but you know crucially in the whole thing we have each one of these is now so you know the belief about the autonomous state this is now a generalized uh coordinate thing up to some order and we're going to all of these so the belief about the hidden state the belief about parameters belief about the log precisions. They're all multivariant now. [snorts] So we collect that entire thing together and that is now the flow of our hidden state effectively which is then we can substitute in the equations that we have for the prediction errors to get well actually what is the flow of the whole belief in terms of the the variational free energy or gradient of variational free energy. H very nice and look it's almost if you kind of squint and you don't see the tilds um you know like muild x muild v this is really the same equation that we had in 8.2 there is a typo actually uh in this so this is equation 8.49 49 very minor. Basically the um the dampening term for uh autonomous states and hidden states has been flipped. So this should be kx and this should be kv. Um doesn't matter too much. Basically that's just a minor erata. Easy to do when you're writing such a a massive and important textbook. I would make so many more errors. It would be hilarious. Yes. But you know I think it's I think it's just a typographical swap. Now what has that given us? So we have a notion about how the hidden state belief for those four things that we have beliefs about is changing in time. That's what 8.49 is. It's saying how does the hidden state belief with respect to what we have beliefs about change in time. That's what we have so far and we've got it in the multivariant case. How do we get the hidden state belief from a notion about how it changes? Well, uh we integrate it basically. So we can now bring to bear our favorite numerical integration method, whatever it happens to be. The simplest case is just Oiler's integration method and we can integrate the motion of our hidden state belief now and get back updates about what we think the hidden state belief is. And then we can do all this computationally. We don't have to do this by hand. um we explicitly do numerical integration because it's easy for computers to do and you can solve things approximately and that then allows us to to do what we did in 8.2 but in the in the the multivariant case. So there are uh you know 8.5 a through to 5e. We can express each one of the gradients of the variational free energy up here in the the main sort of motion of the beliefs that we care about in terms of stuff that we we know already. So, you know, I don't want to have to walk through this in in too much painful detail, but the idea is that okay, well, we we know what the gradient of the variational free energy is with respect to the hidden state belief at layer L in the generalized coordinate of the hidden state belief. Big uh you know, try saying that 10 times fast. We know what that is. We have equations for that. We can just calculate it. And the same for all of the other gradients that we care about as well. So that's 8.5 uh or 50 A through to E C D and E. There we go. Very nice. And then what you can do is you can sort of just do the bookkeeping track uh trick that we um we did at the very beginning of this section is and say well let's just rewrite the precision weighted prediction error as some new variable s impossible to write by hand by the way. And then we can just rewrite everything in a little bit. We don't have to write as much. We don't have to spill as much ink writing down what we're going to write down, but we still have to. So, that's equations 8.51 a through to E again. See, I should really have I've missed out D and E. Oh, dear in my haste to to get it down. Now once we have all that um well we can just go ahead and substitute in the particular transition functions at issue in our model the particular observation function at issue and we can do generalized uh filtering in the multivariate case. So and that gives you the fruit of all of this is a algorithm 13 which is quite long and intimidating that's on page two 200 exactly it's given in pseudo code so there isn't um just before I I get to that I I sort of talk about you know there's a common pattern in the expression for the gradients at issue um that can be sort of useful uh to to meditate upon [clears throat] yes so at the end of this as I say we get algorithm Um that's the whole thing integrating action selection, perception, um even attention and learning all in one one thing which is very very beautiful. um I tend to think so I give I give algorithm 13 here just kind of um in a fairly unstructured manner and it's it's just sort of pseudo code step one step two step three bit like a a recipe I guess but that's it's just designed to be um didactic not not formal so neither is it in the book all right so as I say I think 8.4 four that's really the culmination of all of the mathematical techniques and you know um methodology that we have been looking at since chapter 6 basically. So if you can get to algorithm 13 and understand it um you have understood pretty much everything from chapter 6 to 8 which is the core of continuous state space active inference which is to say a very large amount of active inference. So I think that this is a very big milestone right here. um this point in in the book the rest of uh this this section um well sorry we're going to go into 8.5 now hierarchical message passing we're going to look at sort of different we're going to sort of break down and look inside what we have set up okay so we're not going to look at anything different we're just going to have further exposition on stuff we've already seen it does get pretty complicated though but it's nothing different to what we've seen this is this is it in terms of the um the technique and the material and things like that. So very important to sort of pause and and and realize this is a milestone right here. So before I go into 8.5 that is kind of the last section. It's a big section. Uh there's things to talk about. There's also some are there any questions thus far? I would uh be interested if there are. Let me just check the chat. Okay. Yeah. No, I was very much bothered by what you just explained. using the time index to span the embedding order. Yeah, absolutely. I I think that's uh it can easily be a point of confusion. Um yeah, for sure. >> Yeah. And it it not just happened in the current chapter uh ever since um generalized coordinates were introduced, there seems to have been a confusion, you know, a conflation of of uh time index and embedding index. >> Yeah. They're not the same. It's not like so I've got some belief in generalized coordinates up to some embedding order maybe five 3 to five who knows. That's not the same thing as saying well you know from time step now to 1 2 3 4 5 in the future I'm going to have a belief um that's given by my generalized state belief. That's not the case. So um I'm waving my hands a lot. That's not really helping explain anything. So exactly. So maybe if I just quickly go back to my notes. There we go. This is important because um there we go. Perhaps if I go back to the figure itself 8.8 is it 8 point uh there we go 8.8. So I mean ignore the ignore the um this just for a second and look at the figure itself. If you take slices uh horizontally this is my generalized. So we're looking at three dynamical orders here, three embedding orders basically. Usually that's enough, right? So this is the one for observations, the one for hidden states, the one for autonomous states, blah blah blah. Um, this whole thing is not the same thing as saying that I have a belief about uh the, you know, time step now and also the same things in the next time step and also the same things in the next time step. That's not what that's saying. It's saying that I have a belief about the motion of my hidden state now and I can use that to represent what I think the motion is going to be some amount of time into the future. So I think that's what he's sort of getting at here dedactically. But yeah, it's absolutely not the case that your multi your your generalized um state space model is giving you predictions about future states. That's not what it's giving you. We're going to see that in chapters 9 to 10 though when we get to sort of planning and things like that. So that is that is coming up. Very important to not confuse those two notions though. Hopefully that wasn't totally confusing. So I think I will go through 8.5 now unless anyone does have anything they want to look at. So this is where we kind of it's it's it's different spectralizations on what we've done in in 8.4 really and there's a lot more figures and things to look at. Let's do that and hopefully we'll have some time for uh for questions at the end. I want to share this screen and I want to go to that. Perfect. Hopefully I'm sharing the right screen. Make uh make a lot of noise if I'm not. So yeah, I mean let's actually start this with a figure. So the idea is we're going to we're going to look at what we've done um in terms of the formalism that we've constructed. So this this section is hierarchical message passing. Now, so 8.9 figure 8.9 is where we start. So the idea is that we have um now if we just look at we're looking at all well you know um adjacent layers in our model. We have state units or state neurons and we have their respective error neurons and they jointly allow us to be able to talk about prediction errors. Uh and of course they're coupled across layers in our hierarchy. So if you look carefully, it's a bit hard to see. We have of course these are in generalized coordinates now. So this mu tilda x is actually something that we need to unfold again. And we're going to see that um in some further figures, but that with all of that complexity aside, what we're looking at is we're looking at something that is is a hierarchy and there are things being passed along that hierarchy in two different directions. So at the at the very bottom layer, we're not quite looking at the bottom layer here, but imagine this is the bottom layer. We've got actual sensory prediction, you know, sensory stuff coming in to, you know, maybe this is light intensity or or chemical gradient or something like that. We are interfacing with the world that then allows us to update our predictions about what we think we were going to observe at every point in the hierarchy. And those green units are what calculate the prediction error. So this unit here is going to calculate the prediction error over my hidden states belief basically. So what's it what it's going to have to do is it's going to have to predict that well uh calculate that prediction error and send it to the unit that is predicting what the hidden state is going to be at that layer and then also the autonomous state at the layer above. So you can see prediction error units they sort of um send messages to the to the to the same or adjacent um state neuron at the same level and also the state neuron at the layer above. So they send messages up the hierarchy. So it's like if you imagine you get a basically if you think about this whole thing as an organization or or um you know a business or something you've got the CEO up here or the boardroom whatever and they're sending predictions down about what they want to observe basically that's going down the hierarchy and then at you know down at the bottom layer you've got your engineers you've got your sales people they're actually directly interfacing with stuff and they're saying oh look well you know at this layer I'm getting a certain prediction error uh with respect to what higher ups are telling me and what I'm actually observing. I need to pass this back up the hierarchy. Now, that's in a very loose sense what's happening. So, and you can see in these two motifs down here, we can see we've got predictions that are coming top down from management, let's say, and prediction errors going bottom up from, you know, engineers and sales people and whatnot. That's that's the that's the colloquial sort of story as to what's happening. There's a very nice I don't have it. Um, but he's he's colorcoded uh the equations that issue that show you what's happening here in a very very nice manner. So I don't have them colorcoded. I I can do that, but it was going to take a bit too much time. 8.54A and D. These equations here specify exactly what we just looked at in terms of that uh hierarchical uh error passing little network we saw. There's a little bit of mathematical yoga uh with respect to how we're representing prediction errors here. So I don't want to necessarily dwell on that for too long. But if we go all the way up here, [snorts] so you know we start actually here hierarchical message passing there's again we're going to sort of rewrite precision weighted prediction errors in terms of this generalized sigh thing in the book. So he he he just has a different way of representing this where you can get rid of uh one of the terms and you can represent things a little bit nicer. I go through and this is this is effectively footnote 17 but I derive why we do that um what that means and so on. So if you're if you're unsure as to why we're suddenly changing notation a little bit hopefully this this section in the notes will tell you. So that ends right about here. So there's quite a few pages. Okay. So all that is to say is that we have state units and error units and we're passing messages between them is in this hierarchical fashion back and forward. Uh and that brings us to something really cool. You could actually just spend forever really looking at this figure, figure 8.10. So this is kind of crazy. Uh it's not the craziest figure here even yet. Is this going to zoom in? Ideally open in a new tab. There we go. Okay. Okay, so what we're just looking at the same thing again now except we have all of the respective units and their prediction errors represented. So we have um log precisions for hidden states, log precisions for autonomous states um and then beliefs about hidden states and beliefs about autonomous states. So we've kind of taken if I go to this is not the right thing. We've essentially taken this picture 8.9. So, let's just figure, you know, well, no, we haven't quite done that. I've I've I've uh we've got something coming up I want to to talk about. What we've done here is we're just actually showing, okay, what what are the patterns of updating really between all of the the the units at issue and the crucial thing is that we have excitatory uh connections and we have inhibitory connections uh happening here. So things essentially like you know I've got z x and I've got an arrow going to muilda x that's saying that this prediction area is going to be adding to or exciting uh mu tilda x but you see on s x itself it's got this other arrow which terminates in a uh a dot this is an inhibitory connection it's going to be sort of dampening down the activity of of that neuron. It's a self connection as well. So all the way throughout this section there's continual kind of references and um suggestions to the effect that this is one way to talk about how cortical columns in the brain are connected because we precisely have in that situation we also have there are literally neurons that excite other neurons and there are neurons that inhibit other neurons. This might be a way to start talking formally about that process which is very exciting. I don't know too much about that myself, but you can begin to see where modeling of that kind might be able to come from if we take this hierarchical perspective effectively. So just a little bit on the motifs, what we what we're going to do in each each of these is we're going to say, okay, well, how is the autonomous state update? Okay, that's you know, autonomous state is my V. Well, let's look at all of the connections that go to the autonomous state. Okay, so we've got, you know, this one here. What does that do? It takes in a a zax. It takes in a mu. Where is it? Here. Uh, muv well z muv. It takes in the previous. Actually, it'd be better to look at this one here. It takes in the previous zy mu. Um, and there's also a selfexitatory connection as well. And then you can see colorcoded um in equations 5.44a and d actually where these connections come from, which is very very nice to be able to sort of look at this. This is a huge amount of information that's trying to be communicated here. And then we do that for all of the all of the kinds of states. We have the autonomous state update, the hidden state update, and then the autonomous error state or error unit and the hidden error unit as well, their updates. So that's that's a very um instructive figure to meditate on for quite a while. [clears throat] So that gives you kind of the full picture almost. So, you know, these are still um generalized, you know, uh state units basically. So, there's a lot of complexity hiding within this little neuron here. Okay. Um and we're going to now kind of unwrap that and and see what's inside that as well. So, generalized the actual generalized motion, the dynamics itself is still a bit hard to see. Basically, if we go to figure 8.11, we've now unrolled the uh the the dynamics or rather the the representation of our generalized um coordinate. We've unrolled the generalized coordinate representation of our hidden states, observations, and autonomous states. So, let's pretend that we only have generalized coordinates up to order three. So, we've got zero, one, and two. There's three here. What is rep what is represented here that wasn't represented in the previous picture is that yes we've still got the same so we're looking at the layers going from like left to right in terms of shallow to deep. So from the left this would be I'm getting sensory evidence in over here and as I go to the right I'm going deeper and deeper into the brain basically or higher up the organization ladder let's say. Um and crucially between dynamical orders we have these couplings in red and they weren't visible before because we were just looking at the whole generalized state you know as one little little thing but there are actually couplings between the um dynamical orders and we've actually you know we've seen that from chapter six um that's been continued all the way through but these couplings these are precisely the autonomous uh couplings basically and they there's there's a bit of talk in here about how they act as prior for the level above in terms of the um well the level above in terms of the dynamics in your uh generalized representation. That's that's quite an important that's a subtle point. Um but all that is to say is we're looking at the same thing I just showed you just kind of unwrapped now um with respect to the the um generalized motion. again uh these states are typically not just one thing. So if we look at you know just uh let's figure 812 now what do we got here? So we got uh log precisions let's say X let's just take this one we also have um you know beliefs about hidden states X these typically look like this let's just say there's two components to each of these things. So usually you've got you know to unstack across another dimension as well. Um so you can imagine this gets really really really crowded and horrible to look at very very quickly. Um but it is important to get an appreciation as to like what's ultimately connected to what and in what respect. So again going from left to right this would be your um layer basically this is one layer in that network and then going up and down at least in the way that's formalized here that would be where you are in your um your your the dynamical order of the the um representation. Frasier, would it be all right if I uh chimed in on something for a moment? >> Yeah, please. >> Okay, awesome. Thanks. Yeah, I just uh >> where would you like to go? >> Yeah. I know just with respect to those figures I just want to comment that uh I think what what Sanjieve has done because we see all these very complex graphical structures and and at the same time I don't think um you know these different illustrations that Sanjieve included in chapter 8 they're they're trying to show different perspectives of of potentially the the same thing in some ways right so it's like here with this particular figure we get You know what what happens whenever we have a generative model that is hierarchical in the sense that it has multiple layers and then it also has those generalized embeddings which is what that the u um you know if we read it vertically with the dynamics that's what each of those are referring to in the meanwhile that other figure um that you just now referred to the more the smaller uh yeah uh figure 8.12 thank you uh yeah this is just you know if we if we've been following along with the kinds of notation that that Sanjieve has been employing. This would be just be something like a a multivariate case. And so, uh, in that case, you know, we're not at at a given level, we could be trying to infer not one, but actually multiple hidden states. And so that's what those um superscripts are referring to like the so we're saying like this is the zeroith or first hidden state and its respective uh log precision uh error unit and then for the yeah for the mux subscript parenthesis one similarly um that's that's a second hidden state that we're inferring at that same level. And just to reinforce Frasier's point, you know, it's the complexity of your model. You know, if we have this multivariate scenario plus, you know, multiple hierarchical layers plus generalized embeddings, you can imagine like we we'd have like a uh I don't want to say quite an explosion because that implies that something's gone wrong or maybe there's something that's gone infinite. Uh but but we would end up with a a rather large network model. So it's just to to again just making the point that you know um each of these figures is just supposed to represent like this is how we could illustrate you know so this would be a univariat case right um tech technically and this in the sense that we're only inferring one hidden state per level in this figure >> but it's still generalized like this is this would be a univariate generalized uh coordinate representation. Yeah, because if you were to go multiver, you would then have to go in another dimension and look at like a the volume, which is kind of what this one to show. >> Yeah. So, it's worthwhile to get used to the idea of like thinking in terms of, you know, dimensionality. It's like just as we have like a, you know, a hierarchy uh axis in that graph and then we also have a dynamics axis. You could imagine it becoming 3D uh in a sense whenever we make it a multivariate case. And you know if there was some other functionality that we added to the model and it even beyond that that maybe even Sanjie hasn't included then suddenly we're in a like this kind of fourdimensional graph as it were um and those sorts of things. So it's it's worth it is worthwhile to just get familiar with the the notation in those ways. It becomes very useful when thinking about how to construct your model whenever you're capable of like you know kind of employing these sorts of notations and illustrations. Anyway, yeah, thanks very much. That's all I want. >> No, I completely agree. I mean, that is the reason why we're getting hounded with so many representations of the same thing is because there's different ways of aspectualizing what we've constructed in 8.4 basically. Um, and that's that's really all he's trying to communicate here is that there is a very rich array of phenomena that you can kind of begin to think about formalizing things like attention, things like excitatory and inhibitory connections in in neural columns with what we've created. Um, but you don't have to you don't have to memorize all of this with respect to just spinning up uh, you know, multivariants model of um, generalized filtering or anything like that. It's just to kind of, you know, get a sense of the the gamut of things that are out there. So, the the thing that's yet more of that sort of thing, there's really two things left before we can maybe open it up to questions. If I go over here, the um yeah, that's nice. The last kind of little bit um you know, 8.53 learning, attention, and action. If I go back to 8.10, 10. So he says here in neural network terms mu theta so our belief about the the parameters corresponds to the strength of the intrinsic connections between state and error units. So the strength of these connections here or rather you know depending on what what we're talking about jointly. Um so you know the strength of the connection encode is encoded in the arrows in figure 810. uh and that corresponds to the sort of synaptic efficacy or the learning of the connection weights between the neurons. So you know for those of you who who've done some deep deep learning which is probably a lot of you you know that well what we're doing the game of sort of deep learning is we're trying to say what should the weights be of the connection between the various kind of units at issue the various neurons at issue um and that that parameter mu theta or the belief about that is precisely well you know what what do I believe the weight of the connection should be between in this case you know hidden states prediction error and the hidden state itself that's kind of what it means um to be doing learning is you're you're you're changing those weights a bit. So you set those weights, you can then do prediction with that moment to moment. Those weights could remain vaguely the same or you could be paying attention to how the weights themselves change through time and they're going to change a bit slower you might imagine basically. So there's there's and there's some some references there about kind of where to go in the literature about how they've they've been modeling that. The last thing is we've kind of already seen this figure 8.13. So of course you know this is active inference. We've we're actually doing active inference now. So the the notion about where action comes from uh is so compress the whole thing again back to that sort of simple 8.7 I think uh figure we've seen. Where does action come from in this this whole thing essentially? And the idea is well we've got our forward model uh and that's going to allow us to generate actions as a consequence of sensory prediction errors gradient descent on the variational free energy with respect to our forward model of the environment. Um that's kind of where they come from. So you can imagine that this is the crucial boundary here really. We've got the ability to act on the generative process that's on the left here and we've got sensory evidence coming in and we have our hierarchical representation going deep into the brain of the agent this way basically. So, and all of those things have been unified in what we've talked about in what we've set up in in section 8.4. So, 8.5 is just kind of archaeology on what we've already done really and you know various ways in which we can model certain aspects of the brain with what we have done. So uh again that's algorithm 13 is the the key thing there. So that kind of where have we got here? Yeah. So the last thing is I I've already I've already um emphasized this but the whole story about how this works is that we've got forward and backward passes in our model. Basically, you know, we've got forward pass. Well, typically what we call a forward pass is um get down. Do we have that? It's not going to jump to it. That's very annoying. Nope. All the way down here all the way down the bottom. [clears throat] Yeah. So, just just in in sort of equations um on on the notion of action, our forward model was saying, okay, well, what's how do sensory observations change depending on changes in my actions? So that's what we're calling our forward model. And we can we can spin that up as gradient descent on the variational free energy by the chain rule effectively. So all right so this notion of forward and backwards passes um let me get the the notes up. So you call what is typically called a forward pass is you know updates to error units within each layer pass the current belief about autonomous states uh from the layer below to the layer above. So that's kind of your prediction error updates from like sensory evidence from the world. So you know I observe something I'm going to have a forward pass that's going to do my prediction error updates. And then a backwards pass is the opposite where we're saying let's pass predictions down from our priors um to to affect what we would like to have happen. So the backward pass updates let me get the notes here. So the backward pass updates state units within each layer and passes the current belief about autonomous errors from the layer above to the layer below. That's fundamentally kind of what's happening with respect to the model that we've we've created here. That is shown in figure 8.14. However, I I think there's actually an error uh with 8.14. So we've got our forward pass on the left and our backward pass on the right. This is totally correct except I think the er the the arrows in the forward pass they should be going opposite direction. So they should be going from sensory evidence Y updating my prediction errors. Okay. So the only thing being updated here are prediction errors on autonomous states and hidden states at each layer. And then the backward pass predictions coming down from from the very top prior all the way down to um update my my beliefs about hidden states. So that's abstractly what's happening. That brings us to now the the last bit is well that's all very well and good but what's actually happening inside of a layer like you know inside layer 3 what's happening inside there? The answer is a very large amount of very complicated stuff as you might imagine. So what we have to look at that is figure 8.15. It's not too bad. Um it does it is a bit scary though. But I I don't want to I'm not going to step through this in in gory detail, but the idea is what we have here is a representation about what's happening inside a layer. So on the top left, let's imagine that our autonomous state is known. We know what that is. Okay, so we get observations coming in Y, actions going out a and we also have a forward model that allows us to do all the computations necessary. If we know the uh the autonomous state, where is that represented? We don't uh well the problem is a lot simpler basically and we can just do all of you can follow all the arrows here and going in from sensory evidence all the way through to action at the very end on the bottom left uh is the same figure exactly except now we're assuming that we don't know what the autonomous state is and we have to have a belief about it. So unfortunately it's not rendering very nicely here. But we have our belief about the autonomous state given by precision and some some fixed precision and fixed mean and we can go through and do exactly the same story again. Um just in this case we we have a probabilistic representation of the autonomous state. We have a belief about it and then on the right this is really the same thing again except in the hierarchical case. So this is a hierarchical model where we've got two units. So literally exactly the same. So on the left on the bottom we've got you know what's what's my belief about the autonomous state? Well, it's some fixed belief. Very nice. And I can do that whole story. On the right, what's my belief about the autonomous state? Well, that is itself given by a whole model about transitions that go on between autonomous states and hidden states and predictions and such like that then parameterize the bottom and then the very very top belief is given by some fixed thing. So you can you always sort of terminate it at some fixed belief but um you know looking looking inside the error and u um state units this is the sort of thing that you would see basically so there's a lot of detail to try and get across but uh I would recommend having having a a deep uh look at at those um representations. So that is it for this section. It's really just summary conclusions. Next of all 8.6 6. I would recommend actually to read that if you don't I would recommend reading 8.1 8.2 skimming 8.4 um and then reading 8.6 because this really ties together well why have we done all the things that we've done in this chapter basically. So it is summary and conclusions. It's a bit shorter. Um that gives you a sense of the general story that's been told like why why do we care about any of this? And it it it ends in in this figure here. This is kind of the whole thing like why what we've been trying to do this whole time. There we go. So we have our you know action perception cycle that we have the notion of the the the blankets. I can get pass observations across the blanket. I can receive observations from it and there are belief updates going on inside. Um and then we can precisely do belief updates on hidden states, autonomous states, log precisions and we can solve the triple estimation problem. And oh by the way here's how we do it in terms of uh you know Gaussian probability distributions and things like that. So this is kind of the overview of the whole chapter right here. And that's the end uh and that is the end of the continuous state space formalism of active inference. I'm not going to end the meeting. I want to end the stream there which is an enormous milestone. Uh so if anyone does have questions I know we're we're almost at time uh for the recording. I would love to to take them if necessary. [clears throat] I I think what's interesting and I'm looking forward to chapter 12 where we pull in both continuous and discrete uh to look at message passing. So there there will definitely touch on um continuous systems again. >> Yeah, absolutely. Well, we we'll see that there's a deep affinity or there's a way to to use both of them, but we're going to have to cover discrete state space stuff first. So I'm I'm looking forward to chapter 12 as well. That'll be my favorite chapter. Andrew, I see you've got your your hand up. >> Agreed that uh chapter 12 will be pretty great. Um yeah, really looking forward to that. Uh there needs to be more work on thinking on continuous and discrete state space models sort of together and there's still much more sort of research to be done on hybridic or mixed models. Um there have been publications on those things. There's some code in SPM, but it's still uh you know, it's a very interesting thing to think about. Uh, you know, if we can we can get a strong enough sense of our our message passing, uh, then we might just be able to go from having, you know, an agent who has a continuous statesbased model closer to sort of the lower layer that then kind of leads up to being able to make discrete uh, decision- making and planning at a at a higher discrete layer. in the um back and forth. So um >> you might be thinking you might be thinking should I go and get a coffee now or should I you know stay studying that's kind of a discreet thing but then of course if you decide to go and get a coffee well suddenly I have to move my hand across the counter to do something and that's in the the continuous state space. So there's you might imagine that there is actually um you know a way to combine that's kind of one way you you might think about how they could come together. Yeah, absolutely. Yeah. To to be able to have that sense of continuous control and u um you know with generalized coordinates and the like to allow for all of the the specific dynamics whereas usually the discrete state space models which we will start seeing next week are um you know actions are going to be kind of conceived a bit more simply. Usually it's going to be like take the coffee yes or no, right? As opposed to let me sort of, you know, proprioceptively like navigate my arm to grab my coffee or something, right? So, it's really going to come down to the use case that you're interested in. And usually the discrete models are are are more strongly used in uh computational psychiatry in the sense of it's it's a little bit easier to fit them to empirical data uh as opposed to including all the the the complexity of using like continuous coordinates and the like. Um but uh I'm just saying that >> sorry >> well very much for decision- making problems. Should I go left? Should I go right? That kind of thing. So yeah. >> Yeah. in the sense that if we're sort of trying to identify something like you know pathological behavior or something like that it's kind of like we want to be able to see uh in terms of treatment like how does this impact the behavior and decision-m of an individual you know given their respective generative model and the like. So uh anyway yeah we're we've already been I noticed we've already been getting questions in the chat the past couple weeks on the discrete statebased stuff. So it'll be nice to finally more directly get into that next week. Um and then uh uh finally I just raised my hand to just uh throw it out there. I mentioned it on Tuesday, but I've been working on a library for um sort of hierarchical active generalized filtering. Uh so if anyone would like to uh in any way be involved in that uh in the coming weeks or or otherwise um please do feel free to to reach out. I think I left my email in those supplementary slides, but I'll go ahead and put it in the chat here as well. Um, >> to help I'm help where I can. So, yeah. >> Yeah. No, it'd be awesome to to have you involved. Yeah. No, I mean, if things are going well because of how thoroughly Sanjie has written all of his equations, right? So that's again back to the point of like if we if we're able to sort of master how we implement all these equations in code then suddenly we sort of you know could work towards an ideal uh Rosetta stone so to speak of how to go from from mathematics to to computation and then whenever we have different kinds of experimenters from their own respective domains or subdomains uh it's a really nice way to give everyone sort of like a clear um you know translation between people coming from whether it be sort of pure or mathematics or engineering or computer science or psychiatry or whatever you like, cognitive science. Um, yeah, very cool stuff. Uh, it really speaks to active inference, you know, aiming towards being a broader framework uh that that's rather multi or interdisciplinary. So um yeah >> on that on that project of you know making making the sort of hierarchical state space um triple estimation problem um what you can do is you can just turn off parts of that to recover things that we've already done. So it's not like you always have to deploy that very sophisticated machinery in every case. Maybe you don't care about uh you know log precisions or learning parameters or anything like that. you can just kind of turn those off and get back the thing the more the simpler sorts of things that we've we've done. So it's it's nice in terms of being able to have the full picture and then you can kind of c you know customize that to your particular problem. Um and especially for learning that can be very helpful. So I think that's a worthy worthy pursuit. So >> yeah, absolutely. Just something that's relevant to this chapter is uh you know we get algorithms 12 and 13 and 13 is sort of this all in one almost uh you know kind of the the the end all be all master algorithm of continuous state spaces as far as we explore them in this textbook. Um, algorithm 12 is really just algorithm 13. Uh, but we could say with with action turned off and with generalized coordinates turned off, right? So, it's just a few of these features are sort of missing from algorithm 12 because algorithm 12 is just focused on perception and giving us the triple estimation problem. Algorithm 13 then adds in, you know, here's the part of the loop that would uh you would also run if you had multiple hierarchical layers making it a hierarchical model. Here's the part where action plays out if we wanted there to be action. Here's the part where we're do working with uh initializing then updating uh with our embedding orders if we're including generalized. Right? So being able to kind of isolate these different e mechanisms and figuring out different ways of including them and then all kinds of other fun stuff that we might I don't believe the textbook gets into too strongly but things like basian model comparison and being able to try different models and see how well they do in a given scenario or how well they fit empirical data in a given scenario. Uh and you could compare how well those models do. This hearkens back to chapter 4 on variational inference and how we can use free energy as a proxy of surprisal. Uh not just for evaluating our our given model and if it's sort of getting better or worse over time or things like that uh when it encounters errors but also using that as a criterion potentially for comparing different models against one another. And then that way you that that then from there gets into something called structure learning where we might say oh you know we can start to compare the structures of these models and does it help if we include uh uh uh attention and precision estimation as opposed to keeping precisions fixed uh and and those sorts of things. So yeah it's a it's a certainly a big uh a big world of potential for for what one could do with all these different things. >> [snorts] >> speaks to the importance of having a a textbook like fundamentals uh that that lets us build up to that sort of thing. Um >> it's a it's a big it's a big open uh area of research uh as as you know Andrew this this structure learning problem in active inference is very very big right now. So um be nice to have something like that. >> Yeah I can just mention that in my in my posts on learnable loop that is exactly the approach I followed. I try to be as generic as possible in the code plus making the friction going from code to mathematics as small as possible and then to disable certain pieces. Uh sometimes vectors would just be empty or you know things like that to make use of a single body of code that's as generic as possible but you can um down uh you know apply your application can be downscale to just what you need. So you're welcome to look at that if you want. And then uh Frasier just quickly if you were to go to figure 612 I believe >> that same 612 um the same confusion about lining or lining up a time index with a embedded embedding index is portrayed in that figure as well. >> 612 back there. So that's all the way back when we did generalized uh coordinates from the very beginning. >> Yeah. uh get down here. Yeah, I wouldn't be surprised. Uh it's it's not it depends on how you're interpreting the figure, I would say. So, let's go to 612. Uh there we go. Yeah. Yeah. So, it seems the exact same confusion uh you know is here and in chapter 8 figure as well. And I'm not sure if maybe Sanjief uh did not quite uh see the distinction or he why he aligned these two but um I I found it ever since I saw the 612 for the first time I found it very very confusing. >> Yeah. I I I think I think the way to resolve the discrepancy is to what we're what is being communicated here is that we have our you know uh generalized coordinate representation of the hidden states that furnishes us with the ability to have a prediction at the current time that is then serviceable for some small number of future time steps maybe up to you know one and two into the future. But that's different to claiming that you have a belief about the future time steps. Um and certainly not that you've done some sort of planning procedure where you're explicitly representing uh you know those future as yet unhappening um state transitions and observations. That's that's not the case. So >> yeah I I I think the two concepts individually are totally valid even as pertain here. The problem I have simply is um the the uh time index should not be set to align perfectly with the embedding index. I think that's a confusing part. >> Yeah. >> No, I think so too. I think so too. >> I agree with that too. I mean it's you know whenever we get into the discrete states state space models and using toao uh in that way uh it start it would it would make much more sense and that's how more commonly how it's used in the sense of but but the the difference is that with this discrete state space model like if you have an agent who's planning and you use towels to simply say like this is the time step respective to the agent's model rather than respective to the global simulation where we just use t um you know t would be used as a specific moment in time in a simulation. Tao is just in reference to although given the agent's current time step like tow + one or tow minus one would be what the agent expects in the next time step as opposed to the previous. So it's just relative. Um but yeah with with generalized coordinates it's just not the same right it's like for an agent who says oh uh you know at toao equals 0 or my current time step I believe the hidden state is this whereas at toao equals 1 in the next time step I predict it that same hidden state that same variable within my model will instead be this whereas in generalized coordinates they're technically different variables they're all based around the same thing so to speak but it's like if you have a mux versus a mu mux prime or a mux you know a dotted mux in the sense of it being like a flow then you're not saying oh I think the hidden state will be this at toao equals z but it will instead be this instead at towa equals 1 they're actually different I'm that would instead in the this continuous state space where you're looking at that figure you're saying oh I think at toao equals uh zero um the state will be this but I think that its flow at toao equals 1 will be this right so it's it's different to make a statement about what is the hidden state versus what is its flow be right so that's the that's the trick so I do I do >> just just on that point like this is my belief about the hidden state and then this is my belief about its flow at the current time step [laughter] um and then yeah that's that's different to the belief about future states themselves for sure >> but I must say in in chapter 12 figure 1219 these two >> these two concepts are portrayed as analogous. So I'm not sure if maybe that is why he puts them together. So in the continuous case the embedding order plays an analogous role to the um future time step uh index as portrayed in figure 1219. So, who knows? Maybe it comes from from that anal analogy. >> Possibly. Sorry. But um >> I'll have to look. >> Yes. Sorry. I just I I am inclined to agree with you, Kobas. That would that would given that he's using this shared uh notation with with Tao, etc. It does seem like he's attempting to tee this up is you know this is the the continuous uh you know analog or counterpart of how we work with these internal time steps and inference horizons and discrete state space models. Um but but I'll still I still kind of hold to my yeah previous point of like conceptually it's still not it's not actually working equivalently in that way. It it basically it's not just that we're looking at a discrete versus continuous and that's the only difference. It's not quite it's like saying um you know what's the position of a ball versus what what is its velocity like those are two different things to be inferred. You can't just set them on this time horizon like that in the same granted if you did feel confident about your prediction of a position of a ball and its velocity then you hypothetically could predict its position in the next time step. um you know and then that would start to look similar but um that's not what we're actually doing in the in the generative model right so we're not having the agent say like oh I think the position will be this in the next time step versus this in the current um we're the agent is just saying here's its flow and sort of its change in position in addition to the current position so u sorry anyway yeah uh it it'll be it'll be good to get once we actually get to chapter 12 because yeah the code is still being sort of built out. We we we try to cover including all equations and figures by time of the chapter being uh discussed in uh on the week to week. So hopefully this will be more filled out u you know by then. >> Well, one last thing before I I end the recording um and then we can maybe ask some some questions offline. If you go to the code section, um this is very higgledygly just now, but my Beijian thermostat example and Andrew's um uh example from 8.1 are there in terms of um coder, sorry, um Google collab notebook. So if you click on them, what I would recommend that you do, oh yes, I want to share this tab instead. What I would recommend that you do, so this is his example for 8.1. You can just click, you know, run the cells and everything like that. I would go up to file and then save. Whoops. Save a copy in drive. And that will then you you you will have a copy with which to do everything. Otherwise, you'll change the notebook for everyone else. Doesn't matter because we have um the you know the the originals ourselves. But um if you didn't want to clone the notebook and go through all of that headache, you can go there instead to play around at least with those first two examples. We'll we'll clear that up and we'll hopefully have a bit a few more that you can play with as well. All right, I'm going to end the recording there on the side of YouTube people, but um thank you. We'll see you next week. If anyone wants to ask a question, hang around, they can. So otherwise, thanks. And then we'll we'll jump into discreet stuff. Be good. See you.