Fundamentals of Active Inference (Chapter 8, Session 38) August 21, 2026
Watch on YouTubeVideo summary
This session concludes Chapter 8 by formalizing the "triple estimation problem," where agents simultaneously infer hidden states, learn model parameters, and adjust sensory precision within a hierarchical generative framework. The core mechanism relies on message passing between layers via state units that encode predictions and error units that encode prediction errors, a process driven by gradient descent to minimize variational free energy. This architecture extends naturally to the multivariate case using generalized coordinates, which represent variables through their position, velocity, and acceleration rather than direct beliefs about future time steps, thereby allowing for the extrapolation of motion based on current dynamics.
The comprehensive computational framework presented in Algorithm 13 integrates action selection, perception, attention, and learning into a unified system that operates across continuous state spaces. Within this hierarchy, prediction errors propagate bottom-up while predictions flow top-down, facilitated by excitatory and inhibitory connections between state and error units that provide a mathematical basis for cortical column interactions. Learning is specifically defined as the updating of connection weights between these units, corresponding to beliefs about model parameters, while the active inference loop functions through a forward pass of sensory evidence and a backward pass of prior predictions to continuously refine beliefs about hidden states.
A critical distinction is drawn between continuous variables, which describe flow and velocity, and discrete decision-making processes that will be addressed in future chapters involving categorical choices like "go left" or "go right." The discussion clarifies that generalized coordinates represent derivatives of motion at a single time step rather than future expectations, resolving notational ambiguities regarding how current states relate to their flows. This section also outlines the transition toward discrete state space formulations and highlights the flexibility of the master algorithm, noting that simpler models can be recovered by disabling specific features such as learning or generalized embeddings.
Looking ahead, the session sets the stage for Chapter 12, which aims to integrate these continuous and discrete models into a cohesive theory. The speakers invite collaboration on developing a library for hierarchical active generalized filtering and direct viewers to Google Colab notebooks featuring practical examples, such as the Bayesian thermostat, for further experimentation. Ultimately, this chapter establishes a robust foundation for understanding how complex agents navigate environments by balancing immediate sensory input with learned priors and dynamic motion extrapolation.
Read the full video transcript
All right. Hello everyone. So we're here
at uh the last section, last session on
chapter 8 on fundamentals of active
inference. It's August 21st, 2026. Um
this is a pretty watershed chapter
actually and our our culmination, our
finishing of it here is is marking a
transition where we are in the book um
to something that's going to look maybe
a little bit different to what we've
actually looked at for everything so
far.
um we're we're looking at we don't
really know it yet, but we're looking at
the end of the continuous um state space
formulation of active inference where
we've had all smooth distributions and
things like that. We're going to be
moving into the discrete state space
formulation uh with chapters 8 and nine,
sorry, nine and 10. We're finishing
chapter eight just now. So, what we did
last week is we we went through sections
8.1 and 8.2
on uh chapter eight. and they're kind of
the conceptual core of the chapter
really in terms of what's the major idea
and indeed the major idea in chapter 8
is this thing called the triple
estimation problem where we're learning
you know we're doing inference of hidden
states we're learning um parameters and
we're also learning sensory precision
we're sort of doing these three things
all at the same time in some sense
that's kind of the really high level
30,000 foot nutshell overview sorry for
mixing my metaphors of chapter 8
basically and then there's various takes
on how to do this uh throughout the
chapter. So with respect to 8.1 and 8.2
we got to the where we got to was the
look you know formalizing this whole
thing hierarchically. So if I share my
screen here uh let's just do the single
tab. There we go.
This is effectively where we left off at
least in my session. So we got to this
notion of what are we looking at? If we
just look at the bottom layer, we've
got, you know, this kind of star-like
situation here. We've got parameters
which influence our hidden state. We've
got log precisions. They influence our
hidden state. And we've got an
autonomous uh uh state which also
influences the dynamics of a hidden
state. Hidden states give rise to
observations. We only observe
observations. So this is shaded in um
gray and everything else is a something
that we have a belief about. It's a
random variable. So there's a lot of you
can see this is very richly notated.
There's lots of stuff all over the
place. You know the belief about
parameters itself. There's a mean and a
precision or coariance that
parameterized this. And we saw that we
can kind of stack this these motifs on
top of one another uh whereby the each
layer is coupled to the layer above or
to the layer below depending on your
perspective by the autonomous state. So
the idea is that the autonomous state at
say layer two couples to the autonomous
state at layer 1 and that's how the sort
of states across the layer influence one
another. So the way that we talk about
this mathematically is we say that the
whole generative model this entire thing
is factorizable in terms of the layers.
So we have our expression for our
generative model uh now in terms of the
the layers at issue. And then you can
spin up notions about okay well what is
the probability density of hidden states
at layer L given autonomous states at
layer L at some kind of normal
distribution and we have our transition
model and then you know prior on on uh
parameters and precisions. The thing
that allowed us to do [clears throat]
all of this learning at the same time uh
was effectively this notion of the
action. So the integral of the
variational for energy across time. So
these are my notes here, latte notes. We
now have them for the whole chapter, not
just um 8.1 and 8.2.
That [clears throat] was kind of the
thing that in some sense allowed us to
do this whole whole story or tell tell
the story about how we do inference
learning and precision or rather
attention at the same time. It's an 8.2.
Sorry,
I'm not going to be able to get it up.
Anyway, we went we went over it last
time. That's where we left off. So
picking up now with 8.3. And again, if
you look at the page here, so we got the
chapter map, we have the the PDF notes
for 8.1 and 8.2 there, there as well as
all of the other ones now. So 8.3
that you know the end of 8.2, we've
looked at formalizing the the story and
you know constructing the model. We have
this notion of a hierarchical uh
generative model now coupled across
layers. So we can factoriize the model
in that respect. How do we actually do
the dynamics? How do we do um inference
in this? Now that's what 8.3 is all
about. We've kind of set up everything
and now we want to do inference like we
want to solve the triple estimation
problem by means of what we've set up
already. So that's you know 8.3 really
does assume that you've embied all of
the details and so on from from previous
sections which I know is a bit of a tall
order given the the rate with which
we're going through things. So
effectively where we start out the the
thing that we want to have kind of you
know in our mind at the end of section
8.3 is he's going to introduce this idea
of uh going to begin to introduce this
idea of sort of message passing between
the layers at issue. Um and that's going
to be talked about in terms of kind of
two different kinds of neurons inverted
commas. uh if we're going to you know
uh anthropomorphis or neuropomorphize in
some sense that's kind of the whole
thing about section 8.3 is that we can
actually tell the story about how
variational energy is minimized in terms
of the way or the pattern of message
passing between layers and between
neurons. So the way what we need to do
to get there first of all is we know
that we have a notion about how the the
motion of our hidden state belief
changes across time. It's gradient
descent on a variational free energy
very much as it has been not only for
hidden states but also for the
autonomous states now and those are the
things that are coupling between layers.
Um we and and we also have we we've told
this story many times before about the
precision weighted prediction errors
with which you can formulate the
variational free energy. Okay, so this
is from chapter four all the way to
chapter six effectively. Um but what's
going to happen is it's going to be a
lot easier to to talk about message
passing if we do just if we just sort of
look at things a little bit differently
or formulate things a little bit
differently in terms of the maths. So
essentially all we do is we take the
usual expression for the gradient of the
variational integer with respect to
hidden states and we want to talk about
precision weighted prediction errors. So
we can define that very easily and this
is this is this is as has been done in
the in the chapter already and we're
going to talk about these precision
weighted prediction errors is you know
just as a single variable now
uh very nice. So that gives you 8.38
a and b these these equations here uh or
rather you know you can turn this which
is what we've seen previously into this
and it's just a rewriting of the same
thing really.
So and the idea is that it just makes
clear that the the structure of the
update to your model is given is you
know it's your sensitivity to
observation predictions this thing here.
So what's the change in my observation
generating function with respect to
hidden states multiplied by the the
weighted by the observation error. Okay,
the prediction error in some sense and
the same for hidden state predictions
and their prediction errors. Just a bit
of bookkeeping really. So what this does
um so on page what this is page 194 now
the idea is that this is kind of giving
us two two interesting little units of
computation and those units are going to
be talked about as if they're like
neurons. So we kind of have we have
state units things that encode
predictions about hidden states and
autonomous states and we also have error
units and those are going to give us our
prediction errors. So I'll show you a
diagram in a second. Um and that's so
you know we we and we're we're all we're
also operating in the um the
hierarchical case I should mention for
8.3.
So we're assuming familiarity with you
know the um uh the generalized state
space formalism. That's that's where
we're playing and that's going to come
up and and cause some issues actually uh
for us in 8.5.
So what does this end up looking like?
effectively 8.7.
That's really the thing to look at. So,
this figure here, by the way, uh I can't
actually see while I'm sharing if people
have their hand up. So, if they do, um
if you do want to ask a question, uh
just kind of shout out and then we can
uh we can attend to that. So, figure
8.7, that's that's what this is what is
representing everything that I've been
talking about so far. So, we're moving
from the picture on the right, which
we've seen. This is kind of the motif
from 8.1 and indeed from 8.2 uh where we
ended up with our hierarchy. We can
express pretty much the same
model in terms of this the the interplay
of uh our state state neurons and error
neurons at a particular layer. So we're
looking at just you know one layer in
the hierarchy. How do we rewrite our
model in terms of the interplay of state
and error neurons?
The reason why that's nice is because it
allows us to see or formalize the way in
which prediction errors are being passed
around and the way in which predictions
are also being passed around. Those are
the two fundamental things that are
happening. We've got prediction errors
going one way and predictions going the
other. So of course at the bottom of our
hierarchy we know we observe stuff. We
get stuff from the world. That's why the
idea is that our state um state units
these are going to be ob you know we've
got our observation we've got a belief
about hidden state and a belief about
autonomous states right those are the
states at issue but we also have
prediction errors about exactly the same
thing so prediction error on observation
that's going to take in a prediction
from the hidden state as well as the
observation and it's going to pass a
prediction error up the hierarchy
basically so we're sort of looking at it
as if it's like a felled tree. It's on
its side here. The the left hand side is
kind of the interface with the world.
You I'm literally getting an observation
in Y here. And as we go up the
hierarchy, kind of, you know, deep into
the brain. We're going from left to
right is the idea.
But crucially, we're only looking at
this at one layer in the hierarchy.
There's kind of repeated units going on.
All right? And that's that's the fruit
of section 8.3. We're kind of moving to
this picture about representing the
Beijian
uh well the uh the ba the Beijian
graphical model whoops that we've seen
before these kind of pictures to a
picture where we're representing the
communication of um errors and and
predictions to one another. It's just
going to sort of help us track things a
little bit better.
That brings us to 8.4. And that's really
where I would say 8.4 for is very much
the core of like payoff of um this this
way of thinking about prediction errors.
So you know that is the complete active
generalized filtering algorithm. So
we're going to formalize the whole thing
now and we're going to do it whole hog.
There isn't actually an example in this
section which is a little bit sad. Of
course we could scare off an example
especially based on Andrew's excellent
code.
But what what's the idea here? So we we
want to put all the pieces together now
and do active generalized filtering uh
for the multivariate case. Okay, that's
the most general case really we could
think about. So let's start with our
model. Do I have the representation of
the model here? I don't know if I do.
8.39
is the representation of our model.
There we go. So we have our generative
model. And crucially now everything is
generalized coordinates. So we have
generalized coordinates representation
for hidden states, autonomous states,
not for parameters or log precisions.
Okay, that's very important. So we're
making an assumption here. We're making
an assumption that the hidden states or
sorry the the parameters of our model
and the log precisions, they don't
really change very quickly, at least
nowhere near as quickly as the hidden
states and the autonomous states
themselves. So we're just going to
pretend that okay they're sort of fixed
in some respect. That is an assumption
that we're making. Um and it's an
important assumption to know about. So I
mean this is really just exactly the
same thing that we saw at the end of
8.2. We're factorizing a model across
the um the hierarchy. Now uh just a
little thing on sort of indexing. So,
you know, sometimes I I notice sometimes
in the book we've got, you know, going
from zero base L, you know, in our
layers. Bottom layer would be the
sensory epithelium layer zero all the
way up to layer L minus one. Um,
sometimes it's it's actually mixed,
especially when it comes to uh notating
the the index at the very bottom layer.
Doesn't matter too much, but everything
should be in your mind zerobased
indexing, especially that's that's
useful for for Python coding examples.
you know that there should be minimal
friction going from the maths to the to
the coding. So all right now we can spin
up our sort of Gaussian factors. All
right. So each each factor in our model
this is some Gaussian probability
distribution some continuous Gaussian
probability distribution. Okay again
that's been with us for the whole time
and we're going to be giving up that
assumption about uh continuous
distributions going forward. What does
all of this look like? And we have I
should I should mention before we have a
look we have a very very very top level
prior on the autonomous states and that
just is saying well look at some point
we have to terminate our hierarchy
whether that's after two repetitions or
three repetitions who knows but we're
going to have to have some top level
prior on what the autonomous state is
ultimately generated by and it's some
fixed mean and some fixed coariance you
know that's really all that's saying so
what does this as what does it say what
does this look like It looks like figure
8.8. Here's what we have here.
[snorts]
[clears throat]
All right. This is an important figure.
It's also a little bit uh there's
potential to understand this in in a way
that's not too helpful. So, we're
looking at really the same thing we saw
at the end of um 8.2 except now it's in
the generalized state space formalism
where we have uh it's been you know it's
expressed in terms of generalized
coordinates. So you see each motif.
Okay. Uh but then you know vertically
and at the bottom here we've got
observations coming in and then
horizontally we have our embedding
orders. So what we're looking at is you
know this whole thing here is my
generalized observation. This whole
thing here is my generalized belief
about hidden states. um generalized
beliefs about autonomous states. In the
very uh in the very beginning, I
actually think this should be these we
shouldn't have beliefs about the sensory
sorry the the uh precisions. Uh we're
not we don't have that in our model. So
I think that might be an error. But so
this is just to say that we're really
just taking what we had and we're now
expressing it in generalized
coordinates. Um I I will note it's a
little bit confusing in terms of the
indexing here. So we've got the present
time right you know tow some tower is
equal to the present time versus one
moment in the future versus another
moment in the future. I don't want you
to get the assump the you know have in
the idea in your mind that the embedding
order is somehow an index into my belief
about hidden states or autonomous states
or whatever at the time that time in the
future. This is a bit confusing. So
really if I if I go back to the notes,
sorry if I'm not making much sense, but
it is an important point 8.4
let's just think about what generalized
coordinates even are again. Okay, so you
know my hidden state is now the hidden
state as well as the first derivative as
well as the second derivative and so on.
And you can kind of think about that as
you know abstract position, abstract
velocity, abstract acceleration and so
on.
Um the embedding order is just where am
I in this in this list here. So this
would be embedding order zero, embedding
one, embedding order two. Let's think
about what that means. So the idea is
that of course I have beliefs about this
whole thing. Okay, those are themselves
generalized coordinates like this and
that's because we use the lelass
approximation. Um we can just worry
about the means. So that makes our life
very helpful. Do I have I don't have a
nice representation of that. I don't
think I do in in previous notes. So what
what that is to say is that at at some
time like let's say in the current time
I have a belief about the hidden state
and its motion and the motion of its
motion and the motion of its motion of
its motion and so on and they allow me
to for some time into the future uh
extrapolate what I think the hidden
state is going to be for a bit of time
in the future. Okay, it's not the same
thing as literally planning what the
what the future would be. I'm saying at
this time the motion of my hidden state
is such that I think future states are
going to be blah blah blah blah. Okay,
but those those beliefs are always it's
always on the basis of current beliefs.
Um and the the um the embedding order is
just okay well which one of these are we
talking about? you know, maybe we only
care about the first three embedding
orders and usually we do and it's not
the same thing as, you know, talking
about your belief about the hidden state
at some time into the future. So, I'll
leave it I'll leave that be for a sec,
but um I did want to I did want to
clarify that. Actually, let's go back up
for a second.
All right. Now the the rest of the the
rest of the story in terms of how we do
our um our our variational free energy
minimization very very similar to to so
you know we build up our factors in our
model I don't want to go through the
details there we kind of already know
that unless unless you want to um yeah
so crucially we we've got hierarchy
different model different layers in our
model and dynamics uh are you know
what's at some time I've got a
representation of the hidden state it's
its motion the motion of its motion um
and that's what that's what we mean by
the dynamics basically um
[clears throat]
right here and of course at the sensory
boundary autonomous states are just our
observations
there's a bit in here about um
generalized precision and uh you know
the coariance matrices at issue I don't
want to necessarily dwell on the linear
algebra um but of Of course the
precision matrices at issue they depend
on our log precisions. They also depend
on this gamma factor which is talked
about a little bit more in the the
mathematical appendix. So do do have a
look at that. Have a look at the
appendix on chronica product as well if
you don't understand what that is. Not
too important to get very bogged down
there. Um but we end up with actually I
think we do have a representation. Oh
there we go. Yeah. So there's we talked
about the correlation temporal
correlations in between embedding orders
in um chapter six. That's that's what
gamma is. I I remember now that's why I
write these notes.
So at each level we have variables that
we're inferring. Okay. And we saw that
in the uh the Beijian um graph model you
know we've got beliefs about hidden
states, autonomous states, parameters
and log precisions. Those are what we
are inferring in general and we can just
package them up and say well you know I
let's just call this some this this is
called a var theta here not the best
choice given that we also have theta so
we'll just package all of them up and
say well you know this whole thing is v
theta that's what we care about um
inferring but then of course at
hierarchical level you know l uh the
posterior of interest is written like
this so we've got the posterior at level
L given um the the previous uh level
essentially
very nice and yeah this this is the
point I was trying to make before
because we're making the class uh
approximation we can just represent the
belief about the hidden states as the
mean only that's very nice but we're in
the we're in the hierarchical uh or
multivariant case so each of these are
um generalized coordinates so this is
this is almost like a table now
basically so This whole thing is
represented, you know, mux at layer L.
That is some list of, you know, however
many embedding orders you want. Maybe
you want three, maybe you want five, as
many as you as you care to compute
effectively. And that entire thing is
now our hidden state at layer L. And
this whole thing is happening within a a
hierarchy. So it gets pretty crazy to to
to
keep on top of sort of all the indices
and things like that. But what is nice
is that this whole thing has been done
before. This is exactly what we did in
chap uh 8.2. It's just that we now have
the multivariants uh case. So this is
not anything different uh to what we've
seen which is is nice but it is a bit
confusing. So as as we saw before we
could you know we can de decompose the
variational for energy in terms of the
precision weighted prediction errors at
each layer. Now very important at each
layer. So it's local to a layer and then
we can write the equations for well you
know we know the the uh contribution to
layer L it's precision weighted
prediction error before it was a bit
simpler because of course it was all
univariat so I go back 8.26 right so
this is 8.45 45. We're looking at the
the layer wise prediction weighted
precision error for some some one of
these variables we care about. If I go
back to 8.2,
I just want to show you the contrast. So
8.26 is the equation I want. Now it's
exactly the same equation. So do have a
look at 8.2 if you're lost as to like
what's this what are we do we doing
here? Am I going to be able to find it
in time?
Yes. Yes. Yes. Very good. Very good.
So, precision weight prediction errors.
No, maybe not. Um, I just wanted to show
you that the the univariat case again.
Oh, here we go.
Yeah. So, this is the univariat case
effectively. Um, we have precision
weight prediction errors. They're not in
generalized coordinates, but now we are
in generalized coordinates for um 8.4.
Whoops. 8.4.
That's really the only difference, you
know. So, if you're if you're
struggling, I would suggest to to
meditate deeply upon um 8.2 the
mathematics there.
Where did I end up? 8.43 as 8.46.
Yes.
Okay. So, precision weighted prediction
errors, we can express the variational
free energy in a like manner layer wise.
Basically, that's very very that's
excellent.
Um
yes we don't have to talk about
normalization too much. So what you end
up with is you end up with four
canonical precision weighted prediction
errors. This is kind of equation uh well
4 six a through d. These are the central
equations basically. So we have the
precision weighted prediction error for
autonomous states.
uh we we'll set the same for our our
hidden states and then the same for our
parameters and log precisions. Those are
what we can use to now talk about the
the overall flow of our belief in the
multivariat case as we did for the the
univariat case. So I talked about you
know why why is the derivative shift
operator necessary here. We did talk
about that in um chapter six. this kind
of recapitulate it again why we what
where does this come from why do we need
it at all um I don't want to get bogged
down you can have a look at that in your
own time or ask questions uh on on the
uh the web page but you know crucially
in the whole thing we have each one of
these is now so you know the belief
about the autonomous state this is now a
generalized uh coordinate thing up to
some order and we're going to all of
these so the belief about the hidden
state the belief about parameters belief
about the log precisions. They're all
multivariant now. [snorts]
So we collect that entire thing together
and that is now the flow of our hidden
state effectively which is then we can
substitute in the equations that we have
for the prediction errors to get well
actually what is the flow of the whole
belief in terms of the the variational
free energy or gradient of variational
free energy. H very nice and look it's
almost if you kind of squint and you
don't see the tilds um you know like
muild x muild v this is really the same
equation that we had in 8.2
there is a typo actually uh in this so
this is equation 8.49 49 very minor.
Basically the um the dampening term for
uh autonomous states and hidden states
has been flipped. So this should be kx
and this should be kv. Um doesn't matter
too much. Basically that's just a minor
erata.
Easy to do when you're writing such a a
massive and important textbook.
I would make so many more errors. It
would be hilarious. Yes. But you know I
think it's I think it's just a
typographical swap.
Now what has that given us? So we have a
notion about how the hidden state belief
for those four things that we have
beliefs about is changing in time.
That's what 8.49 is. It's saying how
does the hidden state belief with
respect to what we have beliefs about
change in time. That's what we have so
far and we've got it in the multivariant
case.
How do we get the hidden state belief
from a notion about how it changes?
Well, uh we integrate it basically. So
we can now bring to bear our favorite
numerical integration method, whatever
it happens to be. The simplest case is
just Oiler's integration method and we
can integrate the motion of our hidden
state belief now and get back updates
about what we think the hidden state
belief is. And then we can do all this
computationally. We don't have to do
this by hand. um we explicitly do
numerical integration because it's easy
for computers to do and you can solve
things approximately
and that then allows us to to do what we
did in 8.2 but in the in the the
multivariant case. So there are uh you
know 8.5 a through to 5e. We can express
each one of the gradients of the
variational free energy up here in the
the main sort of motion of the beliefs
that we care about in terms of stuff
that we we know already. So, you know, I
don't want to have to walk through this
in in too much painful detail, but the
idea is that okay, well, we we know what
the gradient of the variational free
energy is with respect to the hidden
state belief at layer L in the
generalized coordinate of the hidden
state belief. Big uh you know, try
saying that 10 times fast. We know what
that is. We have equations for that. We
can just calculate it. And the same for
all of the other gradients that we care
about as well. So that's 8.5 uh or 50 A
through to E
C D and E. There we go. Very nice. And
then what you can do is you can sort of
just do the bookkeeping track uh trick
that we um we did at the very beginning
of this section is and say well let's
just rewrite the precision weighted
prediction error as some new variable s
impossible to write by hand by the way.
And then we can just rewrite everything
in a little bit. We don't have to write
as much. We don't have to spill as much
ink writing down what we're going to
write down, but we still have to. So,
that's equations 8.51 a through to E
again. See, I should really have I've
missed out D and E. Oh, dear in my haste
to to get it down. Now once we have all
that um well we can just go ahead and
substitute in the particular transition
functions at issue in our model the
particular observation function at issue
and we can do generalized uh filtering
in the multivariate case. So and that
gives you the fruit of all of this is a
algorithm 13 which is quite long and
intimidating that's on page two 200
exactly it's given in pseudo code so
there isn't um just before I I get to
that I I sort of talk about you know
there's a common pattern in the
expression for the gradients at issue um
that can be sort of useful uh to to
meditate upon [clears throat]
yes so at the end of this as I say we
get algorithm
Um that's the whole thing integrating
action selection, perception, um even
attention and learning all in one one
thing which is very very beautiful. um I
tend to think so I give I give algorithm
13 here just kind of um in a fairly
unstructured manner and it's it's just
sort of pseudo code step one step two
step three bit like a a recipe I guess
but that's it's just designed to be um
didactic not not formal so neither is it
in the book all right so
as I say I think 8.4 four that's really
the culmination of all of the
mathematical techniques and you know um
methodology that we have been looking at
since chapter 6 basically. So if you can
get to algorithm 13 and understand it um
you have understood pretty much
everything from chapter 6 to 8 which is
the core of continuous state space
active inference which is to say a very
large amount of active inference. So I
think that this is a very big milestone
right here. um this point in in the book
the rest of uh this this section
um well sorry we're going to go into 8.5
now hierarchical message passing we're
going to look at sort of different we're
going to sort of break down and look
inside what we have set up okay so we're
not going to look at anything different
we're just going to have further
exposition on stuff we've already seen
it does get pretty complicated though
but it's nothing different to what we've
seen this is this is it in terms of the
um the technique and the material and
things like that. So very important to
sort of pause and and and realize this
is a milestone right here. So before I
go into 8.5 that is kind of the last
section. It's a big section. Uh there's
things to talk about. There's also some
are there any questions thus far? I
would uh be interested if there are. Let
me just check the chat.
Okay.
Yeah. No, I was very much bothered by
what you just explained. using the time
index to span the embedding order. Yeah,
absolutely. I I think that's uh it can
easily be a point of confusion. Um yeah,
for sure.
>> Yeah. And it it not just happened in the
current chapter uh ever since um
generalized coordinates were introduced,
there seems to have been a confusion,
you know, a conflation of of uh time
index and embedding index.
>> Yeah. They're not the same. It's not
like so I've got some belief in
generalized coordinates up to some
embedding order maybe five 3 to five who
knows. That's not the same thing as
saying well you know from time step now
to 1 2 3 4 5 in the future I'm going to
have a belief um that's given by my
generalized state belief. That's not the
case. So um I'm waving my hands a lot.
That's not really helping explain
anything. So exactly. So maybe if I just
quickly go back to my notes. There we
go.
This is important because um
there we go.
Perhaps if I go back to the figure
itself 8.8
is it 8 point
uh there we go 8.8.
So I mean ignore the ignore the um this
just for a second and look at the figure
itself.
If you take slices uh horizontally this
is my generalized. So we're looking at
three dynamical orders here, three
embedding orders basically. Usually
that's enough, right? So this is the one
for observations, the one for hidden
states, the one for autonomous states,
blah blah blah. Um, this whole thing is
not the same thing as saying that I have
a belief about uh the,
you know, time step now and also the
same things in the next time step and
also the same things in the next time
step. That's not what that's saying.
It's saying that I have a belief about
the motion of my hidden state now
and I can use that to represent what I
think the motion is going to be some
amount of time into the future. So I
think that's what he's sort of getting
at here dedactically. But yeah, it's
absolutely not the case that your multi
your your generalized um state space
model is giving you predictions about
future states. That's not what it's
giving you. We're going to see that in
chapters 9 to 10 though when we get to
sort of planning and things like that.
So that is that is coming up. Very
important to not confuse those two
notions though. Hopefully that wasn't
totally confusing. So I think I will go
through 8.5 now unless anyone does have
anything they want to look at. So
this is where we kind of it's it's it's
different spectralizations on what we've
done in in 8.4 really and there's a lot
more figures and things to look at.
Let's do that and hopefully we'll have
some time for uh for questions at the
end.
I want to share this screen
and I want to go to that.
Perfect. Hopefully I'm sharing the right
screen. Make uh make a lot of noise if
I'm not.
So yeah, I mean let's actually start
this with a figure. So the idea is we're
going to we're going to look at what
we've done um in terms of the formalism
that we've constructed. So this this
section is hierarchical message passing.
Now, so 8.9 figure 8.9 is where we
start.
So the idea is that we have um now if we
just look at we're looking at all well
you know um adjacent layers in our
model. We have state units or state
neurons and we have their respective
error neurons and they jointly allow us
to be able to talk about prediction
errors. Uh and of course they're coupled
across layers in our hierarchy. So if
you look carefully, it's a bit hard to
see. We have of course these are in
generalized coordinates now. So this mu
tilda x is actually something that we
need to unfold again. And we're going to
see that um in some further figures, but
that with all of that complexity aside,
what we're looking at is we're looking
at something that is is a hierarchy and
there are things being passed along that
hierarchy in two different directions.
So at the at the very bottom layer,
we're not quite looking at the bottom
layer here, but imagine this is the
bottom layer. We've got actual sensory
prediction, you know, sensory stuff
coming in to, you know, maybe this is
light intensity or or chemical gradient
or something like that. We are
interfacing with the world that then
allows us to update our predictions
about what we think we were going to
observe at every point in the hierarchy.
And those green units are what calculate
the prediction error. So this unit here
is going to calculate the prediction
error over my hidden states belief
basically. So what's it what it's going
to have to do is it's going to have to
predict that well uh calculate that
prediction error and send it to the unit
that is predicting what the hidden state
is going to be at that layer and then
also the autonomous state at the layer
above. So you can see prediction error
units they sort of um send messages to
the to the to the same or adjacent um
state neuron at the same level and also
the state neuron at the layer above. So
they send messages up the hierarchy. So
it's like if you imagine you get a
basically if you think about this whole
thing as an organization or or um
you know a business or something you've
got the CEO up here or the boardroom
whatever and they're sending predictions
down about what they want to observe
basically that's going down the
hierarchy and then at you know down at
the bottom layer you've got your
engineers you've got your sales people
they're actually directly interfacing
with stuff and they're saying oh look
well you know at this layer I'm getting
a certain prediction error uh with
respect to what higher ups are telling
me and what I'm actually observing. I
need to pass this back up the hierarchy.
Now, that's in a very loose sense what's
happening. So, and you can see in these
two motifs down here, we can see we've
got predictions that are coming top down
from management, let's say, and
prediction errors going bottom up from,
you know, engineers and sales people and
whatnot. That's that's the that's the
colloquial sort of story as to what's
happening. There's a very nice I don't
have it. Um,
but he's he's colorcoded uh the
equations that issue that show you
what's happening here in a very very
nice manner. So I don't have them
colorcoded. I I can do that, but it was
going to take a bit too much time. 8.54A
and D. These equations here specify
exactly what we just looked at in terms
of that uh hierarchical
uh error passing little network we saw.
There's a little bit of mathematical
yoga uh with respect to how we're
representing prediction errors here. So
I don't want to necessarily dwell on
that for too long. But if we go all the
way up here, [snorts]
so you know we start actually here
hierarchical message passing there's
again we're going to sort of rewrite
precision weighted prediction errors in
terms of this generalized sigh thing in
the book. So he he he just has a
different way of representing this where
you can get rid of uh one of the terms
and you can represent things a little
bit nicer. I go through and this is this
is effectively footnote 17 but I derive
why we do that um what that means and so
on. So if you're if you're unsure as to
why we're suddenly changing notation a
little bit hopefully this this section
in the notes will tell you. So that ends
right about here. So there's quite a few
pages.
Okay. So all that is to say is that we
have state units and error units and
we're passing messages between them is
in this hierarchical fashion back and
forward.
Uh and that brings us to something
really cool. You could actually just
spend forever really looking at this
figure, figure 8.10.
So this is kind of crazy. Uh it's not
the craziest figure here even yet. Is
this going to zoom in? Ideally open in a
new tab. There we go. Okay. Okay, so
what we're just looking at the same
thing again now except we have all of
the
respective units and their prediction
errors represented. So we have um log
precisions for hidden states, log
precisions for autonomous states um and
then beliefs about hidden states and
beliefs about autonomous states. So
we've kind of taken if I go to this is
not the right thing. We've essentially
taken this picture 8.9. So, let's just
figure, you know, well,
no, we haven't quite done that. I've
I've I've uh we've got something coming
up I want to to talk about. What we've
done here is we're just actually
showing, okay, what what are the
patterns of updating really between all
of the the the units at issue and the
crucial thing is that we have
excitatory uh connections and we have
inhibitory connections uh happening
here. So things essentially like you
know I've got z x and I've got an arrow
going to muilda x that's saying that
this prediction area is going to be
adding to or exciting uh mu tilda x but
you see on s x itself it's got this
other arrow which terminates in a uh a
dot this is an inhibitory connection
it's going to be sort of dampening down
the activity of of that neuron. It's a
self connection as well. So all the way
throughout this section there's
continual kind of references and um
suggestions to the effect that this is
one way to talk about how cortical
columns in the brain are connected
because we precisely have in that
situation we also have there are
literally neurons that excite other
neurons and there are neurons that
inhibit other neurons. This might be a
way to start talking formally about that
process which is very exciting. I don't
know too much about that myself, but you
can begin to see where modeling of that
kind might be able to come from if we
take this hierarchical perspective
effectively. So just a little bit on the
motifs, what we what we're going to do
in each each of these is we're going to
say, okay, well, how is the autonomous
state update? Okay, that's you know,
autonomous state is my V. Well, let's
look at all of the connections that go
to the autonomous state. Okay, so we've
got, you know, this one here. What does
that do? It takes in a a zax. It takes
in a mu. Where is it? Here. Uh, muv well
z muv. It takes in the previous.
Actually, it'd be better to look at this
one here. It takes in the previous zy
mu. Um, and there's also a selfexitatory
connection as well. And then you can see
colorcoded um in equations
5.44a and d actually where these
connections come from, which is very
very nice to be able to sort of look at
this. This is a huge amount of
information that's trying to be
communicated here. And then we do that
for all of the all of the kinds of
states. We have the autonomous state
update, the hidden state update, and
then the autonomous error state or error
unit and the hidden error unit as well,
their updates.
So that's that's a very um instructive
figure to meditate on for quite a while.
[clears throat]
So that gives you kind of the full
picture almost. So, you know, these are
still
um generalized, you know, uh state units
basically. So, there's a lot of
complexity hiding within this little
neuron here. Okay. Um and we're going to
now kind of unwrap that and and see
what's inside that as well.
So, generalized
the actual generalized motion, the
dynamics itself is still a bit hard to
see. Basically, if we go to figure 8.11,
we've now unrolled the uh the the
dynamics or rather the the
representation of our generalized um
coordinate. We've unrolled the
generalized coordinate representation of
our hidden states, observations, and
autonomous states. So, let's pretend
that we only have generalized
coordinates up to order three. So, we've
got zero, one, and two. There's three
here.
What is rep what is represented here
that wasn't represented in the previous
picture is that yes we've still got the
same so we're looking at the layers
going from like left to right in terms
of shallow to deep. So from the left
this would be I'm getting sensory
evidence in over here and as I go to the
right I'm going deeper and deeper into
the brain basically or higher up the
organization ladder let's say. Um and
crucially between dynamical orders we
have these couplings in red and they
weren't visible before because we were
just looking at the whole generalized
state you know as one little little
thing but there are actually couplings
between the um dynamical orders and
we've actually you know we've seen that
from chapter six um that's been
continued all the way through but these
couplings these are precisely the
autonomous uh couplings basically and
they there's there's a bit of talk in
here about how they act as prior for the
level above in terms of the um well the
level above in terms of the dynamics in
your uh generalized representation.
That's that's quite an important that's
a subtle point. Um but all that is to
say is we're looking at the same thing I
just showed you just kind of unwrapped
now um with respect to the the um
generalized motion. again uh these
states are typically not just one thing.
So if we look at you know just uh let's
figure 812 now what do we got here? So
we got uh log precisions let's say X
let's just take this one we also have um
you know beliefs about hidden states X
these typically look like this
let's just say there's two components to
each of these things. So usually you've
got you know to unstack across another
dimension as well. Um so you can imagine
this gets really really really crowded
and horrible to look at very very
quickly. Um but it is important to get
an appreciation as to like what's
ultimately connected to what and in what
respect. So again going from left to
right this would be your um layer
basically this is one layer in that
network and then going up and down at
least in the way that's formalized here
that would be where you are in your um
your your the dynamical order of the the
um representation.
Frasier, would it be all right if I uh
chimed in on something for a moment?
>> Yeah, please.
>> Okay, awesome. Thanks. Yeah, I just uh
>> where would you like to go?
>> Yeah. I know just with respect to those
figures I just want to comment that uh I
think what what Sanjieve has done
because we see all these very complex
graphical structures and and at the same
time I don't think um you know these
different illustrations that Sanjieve
included in chapter 8 they're they're
trying to show different perspectives of
of potentially the the same thing in
some ways right so it's like here with
this particular figure we get You know
what what happens whenever we have a
generative model that is hierarchical in
the sense that it has multiple layers
and then it also has those generalized
embeddings which is what that the u um
you know if we read it vertically with
the dynamics that's what each of those
are referring to in the meanwhile that
other figure um that you just now
referred to the more the smaller uh yeah
uh figure 8.12 thank you uh yeah this is
just you know if we if we've been
following along with the kinds of
notation that that Sanjieve has been
employing. This would be just be
something like a a multivariate case.
And so, uh, in that case, you know,
we're not at at a given level, we could
be trying to infer not one, but actually
multiple hidden states. And so that's
what those um superscripts are referring
to like the so we're saying like this is
the zeroith or first hidden state and
its respective uh log precision uh error
unit and then for the yeah for the mux
subscript parenthesis one similarly um
that's that's a second hidden state that
we're inferring at that same level.
And just to reinforce Frasier's point,
you know, it's the complexity of your
model. You know, if we have this
multivariate scenario plus, you know,
multiple hierarchical layers plus
generalized embeddings, you can imagine
like we we'd have like a uh I don't want
to say quite an explosion because that
implies that something's gone wrong or
maybe there's something that's gone
infinite. Uh but but we would end up
with a a rather large network model. So
it's just to to again just making the
point that you know um each of these
figures is just supposed to represent
like this is how we could illustrate you
know so this would be a univariat case
right um tech technically and this in
the sense that we're only inferring one
hidden state per level in this figure
>> but it's still generalized like this is
this would be a univariate generalized
uh coordinate representation. Yeah,
because if you were to go multiver, you
would then have to go in another
dimension and look at like a the volume,
which is kind of what this one to show.
>> Yeah. So, it's worthwhile to get used to
the idea of like thinking in terms of,
you know, dimensionality. It's like just
as we have like a, you know, a hierarchy
uh axis in that graph and then we also
have a dynamics axis. You could imagine
it becoming 3D uh in a sense whenever we
make it a multivariate case. And you
know if there was some other
functionality that we added to the model
and it even beyond that that maybe even
Sanjie hasn't included then suddenly
we're in a like this kind of
fourdimensional graph as it were um and
those sorts of things. So it's it's
worth it is worthwhile to just get
familiar with the the notation in those
ways. It becomes very useful when
thinking about how to construct your
model whenever you're capable of like
you know kind of employing these sorts
of notations and illustrations. Anyway,
yeah, thanks very much. That's all I
want.
>> No, I completely agree. I mean, that is
the reason why we're getting hounded
with so many representations of the same
thing is because there's different ways
of aspectualizing what we've constructed
in 8.4 basically. Um, and that's that's
really all he's trying to communicate
here is that there is a very rich array
of phenomena that you can kind of begin
to think about formalizing things like
attention, things like excitatory and
inhibitory connections in in neural
columns with what we've created. Um, but
you don't have to you don't have to
memorize all of this with respect to
just spinning up uh, you know,
multivariants model of um, generalized
filtering or anything like that. It's
just to kind of, you know, get a sense
of the the gamut of things that are out
there. So, the the thing that's
yet more of that sort of thing, there's
really two things left before we can
maybe open it up to questions. If I go
over here,
the um yeah, that's nice.
The last kind of little bit um you know,
8.53 learning, attention, and action. If
I go back to 8.10, 10. So he says here
in neural network terms mu theta so our
belief about the the parameters
corresponds to the strength of the
intrinsic connections between state and
error units. So the strength of these
connections here or rather you know
depending on what what we're talking
about jointly. Um so you know the
strength of the connection encode is
encoded in the arrows in figure 810.
uh and that corresponds to the sort of
synaptic efficacy or the learning of the
connection weights between the neurons.
So you know for those of you who who've
done some deep deep learning which is
probably a lot of you you know that well
what we're doing the game of sort of
deep learning is we're trying to say
what should the weights be of the
connection between the various kind of
units at issue the various neurons at
issue um and that that parameter mu
theta or the belief about that is
precisely well you know what what do I
believe the weight of the connection
should be between in this case you know
hidden states prediction error and the
hidden state itself that's kind of what
it means um to be doing learning is
you're you're you're changing those
weights a bit. So you set those weights,
you can then do prediction with that
moment to moment. Those weights could
remain vaguely the same or you could be
paying attention to how the weights
themselves change through time and
they're going to change a bit slower you
might imagine basically. So there's
there's and there's some some references
there about kind of where to go in the
literature about how they've they've
been modeling that. The last thing is
we've kind of already seen this figure
8.13. So of course you know this is
active inference. We've we're actually
doing active inference now. So the the
notion about where action comes from uh
is so compress the whole thing again
back to that sort of simple 8.7 I think
uh figure we've seen. Where does action
come from in this this whole thing
essentially? And the idea is well we've
got our forward model uh and that's
going to allow us to generate actions as
a consequence of sensory prediction
errors gradient descent on the
variational free energy with respect to
our forward model of the environment. Um
that's kind of where they come from. So
you can imagine that this is the crucial
boundary here really. We've got the
ability to act on the generative process
that's on the left here and we've got
sensory evidence coming in and we have
our hierarchical representation going
deep into the brain of the agent this
way basically. So, and all of those
things have been unified in what we've
talked about in what we've set up in in
section 8.4. So, 8.5 is just kind of
archaeology on what we've already done
really and you know various ways in
which we can model certain aspects of
the brain with what we have done. So uh
again that's algorithm 13 is the the key
thing there.
So that kind of where have we got here?
Yeah. So the last thing is I I've
already I've already um emphasized this
but the whole story about how this works
is that we've got forward and backward
passes in our model. Basically, you
know, we've got forward pass. Well,
typically what we call a forward pass is
um
get down. Do we have that?
It's not going to jump to it. That's
very annoying. Nope. All the way down
here all the way down the bottom.
[clears throat]
Yeah. So, just just in in sort of
equations um on on the notion of action,
our forward model was saying, okay,
well, what's how do sensory observations
change depending on changes in my
actions? So that's what we're calling
our forward model. And we can we can
spin that up as gradient descent on the
variational free energy by the chain
rule effectively. So all right so this
notion of forward and backwards passes
um let me get the the notes up. So you
call what is typically called a forward
pass is you know updates to error units
within each layer pass the current
belief about autonomous states uh from
the layer below to the layer above. So
that's kind of your prediction error
updates from like sensory evidence from
the world. So you know I observe
something I'm going to have a forward
pass that's going to do my prediction
error updates. And then a backwards pass
is the opposite where we're saying let's
pass predictions down from our priors um
to to affect what we would like to have
happen. So the backward pass updates let
me get the notes here. So the backward
pass updates state units within each
layer and passes the current belief
about autonomous errors from the layer
above to the layer below. That's
fundamentally kind of what's happening
with respect to the model that we've
we've created here. That is shown in
figure 8.14.
However, I I think there's actually an
error uh with 8.14. So we've got our
forward pass on the left and our
backward pass on the right. This is
totally correct except I think the er
the the arrows in the forward pass they
should be going opposite direction. So
they should be going from sensory
evidence Y updating my prediction
errors. Okay. So the only thing being
updated here are prediction errors on
autonomous states and hidden states at
each layer. And then the backward pass
predictions coming down from from the
very top prior all the way down to um
update my my beliefs about hidden
states. So that's abstractly what's
happening.
That brings us to now the the last bit
is well that's all very well and good
but what's actually happening inside of
a layer like you know inside layer 3
what's happening inside there? The
answer is a very large amount of very
complicated stuff as you might imagine.
So what we have to look at that is
figure 8.15.
It's not too bad. Um it does it is a bit
scary though. But I I don't want to I'm
not going to step through this in in
gory detail, but the idea is what we
have here is a representation about
what's happening inside a layer. So on
the top left, let's imagine that our
autonomous state is known. We know what
that is. Okay, so we get observations
coming in Y, actions going out a and we
also have a forward model that allows us
to do all the computations necessary. If
we know the uh the autonomous state,
where is that represented?
We don't uh well the problem is a lot
simpler basically and we can just do all
of you can follow all the arrows here
and going in from sensory evidence all
the way through to action at the very
end
on the bottom left uh is the same figure
exactly except now we're assuming that
we don't know what the autonomous state
is and we have to have a belief about
it. So unfortunately it's not rendering
very nicely here. But we have our belief
about the autonomous state given by
precision and some some fixed precision
and fixed mean and we can go through and
do exactly the same story again. Um just
in this case we we have a probabilistic
representation of the autonomous state.
We have a belief about it and then on
the right this is really the same thing
again except in the hierarchical case.
So this is a hierarchical model where
we've got two units. So literally
exactly the same. So on the left on the
bottom
we've got you know what's what's my
belief about the autonomous state? Well,
it's some fixed belief. Very nice. And I
can do that whole story. On the right,
what's my belief about the autonomous
state? Well, that is itself given by a
whole model about transitions that go on
between autonomous states and hidden
states and predictions and such like
that then parameterize the bottom and
then the very very top belief is given
by some fixed thing. So you can you
always sort of terminate it at some
fixed belief but um you know looking
looking inside the error and u um state
units this is the sort of thing that you
would see basically so there's a lot of
detail to try and get across but uh I
would recommend having having a a deep
uh look at at those um representations.
So that is it for this section. It's
really just summary conclusions. Next of
all 8.6 6. I would recommend actually to
read that if you don't I would recommend
reading 8.1 8.2
skimming 8.4
um and then reading 8.6 because this
really ties together well why have we
done all the things that we've done in
this chapter basically. So it is summary
and conclusions. It's a bit shorter. Um
that gives you a sense of the general
story that's been told like why why do
we care about any of this? And it it it
ends in in this figure here. This is
kind of the whole thing like why what
we've been trying to do this whole time.
There we go. So we have our you know
action perception cycle that we have the
notion of the the the blankets. I can
get pass observations across the
blanket. I can receive observations from
it and there are belief updates going on
inside. Um and then we can precisely do
belief updates on hidden states,
autonomous states, log precisions and we
can solve the triple estimation problem.
And oh by the way here's how we do it in
terms of uh you know Gaussian
probability distributions and things
like that. So this is kind of the
overview of the whole chapter right
here. And that's the end
uh and that is the end of the continuous
state space formalism of active
inference.
I'm not going to end the meeting. I want
to end the stream there which is an
enormous milestone. Uh so if anyone does
have questions I know we're we're almost
at time uh for the recording. I would
love to to take them if necessary.
[clears throat]
I I think what's interesting and I'm
looking forward to chapter 12 where we
pull in both continuous and discrete uh
to look at message passing. So there
there will definitely
touch on um continuous systems again.
>> Yeah, absolutely. Well, we we'll see
that there's a deep affinity or there's
a way to to use both of them, but we're
going to have to cover discrete state
space stuff first. So I'm I'm looking
forward to chapter 12 as well. That'll
be my favorite chapter. Andrew, I see
you've got your your hand up.
>> Agreed that uh chapter 12 will be pretty
great. Um yeah, really looking forward
to that. Uh there needs to be more work
on thinking on continuous and discrete
state space models sort of together and
there's still much more sort of research
to be done on hybridic or mixed models.
Um there have been publications on those
things. There's some code in SPM, but
it's still uh you know, it's a very
interesting thing to think about. Uh,
you know, if we can we can get a strong
enough sense of our our message passing,
uh, then we might just be able to go
from having, you know, an agent who has
a continuous statesbased model closer to
sort of the lower layer that then kind
of leads up to being able to make
discrete uh, decision- making and
planning at a at a higher discrete
layer. in the um back and forth. So um
>> you might be thinking you might be
thinking should I go and get a coffee
now or should I you know stay studying
that's kind of a discreet thing but then
of course if you decide to go and get a
coffee well suddenly I have to move my
hand across the counter to do something
and that's in the the continuous state
space. So there's you might imagine that
there is actually um you know a way to
combine that's kind of one way you you
might think about how they could come
together.
Yeah, absolutely. Yeah. To to be able to
have that sense of continuous control
and u um you know with generalized
coordinates and the like to allow for
all of the the specific dynamics whereas
usually the discrete state space models
which we will start seeing next week are
um you know actions are going to be kind
of conceived a bit more simply. Usually
it's going to be like take the coffee
yes or no, right? As opposed to let me
sort of, you know, proprioceptively like
navigate my arm to grab my coffee or
something, right? So, it's really going
to come down to the use case that you're
interested in. And usually the discrete
models are are are more strongly used in
uh computational psychiatry in the sense
of it's it's a little bit easier to fit
them to empirical data uh as opposed to
including all the the the complexity of
using like continuous coordinates and
the like. Um but uh I'm just saying that
>> sorry
>> well very much for decision- making
problems. Should I go left? Should I go
right? That kind of thing. So yeah.
>> Yeah. in the sense that if we're sort of
trying to identify something like you
know pathological behavior or something
like that it's kind of like we want to
be able to see uh in terms of treatment
like how does this impact the behavior
and decision-m of an individual you know
given their respective generative model
and the like. So uh anyway yeah we're
we've already been I noticed we've
already been getting questions in the
chat the past couple weeks on the
discrete statebased stuff. So it'll be
nice to finally more directly get into
that next week. Um and then uh uh
finally I just raised my hand to just uh
throw it out there. I mentioned it on
Tuesday, but I've been working on a
library for um sort of hierarchical
active generalized filtering. Uh so if
anyone would like to uh in any way be
involved in that uh in the coming weeks
or or otherwise um please do feel free
to to reach out. I think I left my email
in those supplementary slides, but I'll
go ahead and put it in the chat here as
well. Um,
>> to help I'm help where I can. So, yeah.
>> Yeah. No, it'd be awesome to to have you
involved. Yeah. No, I mean, if things
are going well because of how thoroughly
Sanjie has written all of his equations,
right? So that's again back to the point
of like if we if we're able to sort of
master how we implement all these
equations in code then suddenly we sort
of you know could work towards an ideal
uh Rosetta stone so to speak of how to
go from from mathematics to to
computation and then whenever we have
different kinds of experimenters from
their own respective domains or
subdomains uh it's a really nice way to
give everyone sort of like a clear um
you know translation between people
coming from whether it be sort of pure
or mathematics or engineering or
computer science or psychiatry or
whatever you like, cognitive science.
Um, yeah, very cool stuff. Uh, it really
speaks to active inference, you know,
aiming towards being a broader framework
uh that that's rather multi or
interdisciplinary. So um yeah
>> on that on that project of you know
making making the sort of hierarchical
state space um triple estimation problem
um what you can do is you can just turn
off parts of that to recover things that
we've already done. So it's not like you
always have to deploy that very
sophisticated machinery in every case.
Maybe you don't care about uh you know
log precisions or learning parameters or
anything like that. you can just kind of
turn those off and get back the thing
the more the simpler sorts of things
that we've we've done. So it's it's nice
in terms of being able to have the full
picture and then you can kind of c you
know customize that to your particular
problem. Um and especially for learning
that can be very helpful. So I think
that's a worthy worthy pursuit. So
>> yeah, absolutely. Just something that's
relevant to this chapter is uh you know
we get algorithms 12 and 13 and 13 is
sort of this all in one almost uh you
know kind of the the the end all be all
master algorithm of continuous state
spaces as far as we explore them in this
textbook. Um, algorithm 12 is really
just algorithm 13. Uh, but we could say
with with action turned off and with
generalized coordinates turned off,
right? So, it's just a few of these
features are sort of missing from
algorithm 12 because algorithm 12 is
just focused on perception and giving us
the triple estimation problem. Algorithm
13 then adds in, you know, here's the
part of the loop that would uh you would
also run if you had multiple
hierarchical layers making it a
hierarchical model. Here's the part
where action plays out if we wanted
there to be action. Here's the part
where we're do working with uh
initializing then updating uh with our
embedding orders if we're including
generalized. Right? So being able to
kind of isolate these different e
mechanisms and figuring out different
ways of including them and then all
kinds of other fun stuff that we might I
don't believe the textbook gets into too
strongly but things like basian model
comparison and being able to try
different models and see how well they
do in a given scenario or how well they
fit empirical data in a given scenario.
Uh and you could compare how well those
models do. This hearkens back to chapter
4 on variational inference and how we
can use free energy as a proxy of
surprisal. Uh not just for evaluating
our our given model and if it's sort of
getting better or worse over time or
things like that uh when it encounters
errors but also using that as a
criterion potentially for comparing
different models against one another.
And then that way you that that then
from there gets into something called
structure learning where we might say oh
you know we can start to compare the
structures of these models and does it
help if we include uh uh uh attention
and precision estimation as opposed to
keeping precisions fixed uh and and
those sorts of things. So yeah it's a
it's a certainly a big uh a big world of
potential for for what one could do with
all these different things.
>> [snorts]
>> speaks to the importance of having a a
textbook like fundamentals uh that that
lets us build up to that sort of thing.
Um
>> it's a it's a big it's a big open uh
area of research uh as as you know
Andrew this this structure learning
problem in active inference is very very
big right now. So um be nice to have
something like that.
>> Yeah I can just mention that in my in my
posts on learnable loop that is exactly
the approach I followed. I try to be as
generic as possible in the code plus
making the friction going from code to
mathematics as small as possible and
then to disable certain pieces. Uh
sometimes vectors would just be empty or
you know things like that to make use of
a single
body of code that's as generic as
possible but you can um down uh you know
apply your application can be downscale
to just what you need. So you're welcome
to look at that if you want. And then uh
Frasier just quickly if you were to go
to figure 612 I believe
>> that same 612 um the same confusion
about lining or lining up a time index
with a embedded embedding index is
portrayed in that figure as well.
>> 612
back there. So that's all the way back
when we did generalized uh coordinates
from the very beginning.
>> Yeah. uh get down here. Yeah, I wouldn't
be surprised. Uh it's it's not it
depends on how you're interpreting the
figure, I would say. So, let's go to
612.
Uh there we go.
Yeah.
Yeah. So, it seems the exact same
confusion uh you know is here and in
chapter 8 figure as well. And I'm not
sure if maybe Sanjief uh did not quite
uh see the distinction or he why he
aligned these two but um I I found it
ever since I saw the 612 for the first
time I found it very very confusing.
>> Yeah. I I I think I think the way to
resolve the discrepancy is to what we're
what is being communicated here is that
we have our you know uh generalized
coordinate representation of the hidden
states that furnishes us with the
ability to have a prediction at the
current time that is then serviceable
for some small number of future time
steps maybe up to you know one and two
into the future. But that's different to
claiming that you have a belief about
the future time steps. Um and certainly
not that you've done some sort of
planning procedure where you're
explicitly representing uh you know
those future as yet unhappening um state
transitions and observations. That's
that's not the case. So
>> yeah I I I think the two concepts
individually are totally valid even as
pertain here. The problem I have simply
is
um the the uh time index should not be
set to align perfectly with the
embedding index. I think that's a
confusing part.
>> Yeah.
>> No, I think so too. I think so too.
>> I agree with that too. I mean it's you
know whenever we get into the discrete
states state space models and using toao
uh in that way uh it start it would it
would make much more sense and that's
how more commonly how it's used in the
sense of but but the the difference is
that with this discrete state space
model like if you have an agent who's
planning and you use towels to simply
say like this is the time step
respective to the agent's model rather
than respective to the global simulation
where we just use t um you know t would
be used as a specific moment in time in
a simulation. Tao is just in reference
to although given the agent's current
time step like tow + one or tow minus
one would be what the agent expects in
the next time step as opposed to the
previous. So it's just relative. Um but
yeah with with generalized coordinates
it's just not the same right it's like
for an agent who says oh uh you know at
toao equals 0 or my current time step I
believe the hidden state is this whereas
at toao equals 1 in the next time step I
predict it that same hidden state that
same variable within my model will
instead be this whereas in generalized
coordinates they're technically
different variables they're all based
around the same thing so to speak but
it's like if you have a mux versus a mu
mux prime or a mux you know a dotted mux
in the sense of it being like a flow
then you're not saying oh I think the
hidden state will be this at toao equals
z but it will instead be this instead at
towa equals 1 they're actually different
I'm that would instead in the this
continuous state space where you're
looking at that figure you're saying oh
I think at toao equals uh zero um the
state will be this but I think that its
flow at toao equals 1 will be this right
so it's it's different to make a
statement about what is the hidden state
versus what is its flow be right so
that's the that's the trick so I do I do
>> just just on that point like this is my
belief about the hidden state and then
this is my belief about its flow at the
current time step [laughter]
um and then yeah that's that's different
to the belief about future states
themselves for sure
>> but I must say in in chapter 12 figure
1219
these two
>> these two concepts
are portrayed as analogous.
So I'm not sure if maybe that is why he
puts them together. So in the continuous
case the embedding order plays an
analogous role to the um future time
step uh index as portrayed in figure
1219.
So, who knows? Maybe it comes from from
that anal analogy.
>> Possibly.
Sorry. But um
>> I'll have to look.
>> Yes. Sorry. I just I I am inclined to
agree with you, Kobas. That would that
would given that he's using this shared
uh notation with with Tao, etc. It does
seem like he's attempting to tee this up
is you know this is the the continuous
uh you know analog or counterpart of how
we work with these internal time steps
and inference horizons and discrete
state space models. Um but but I'll
still I still kind of hold to my yeah
previous point of like conceptually it's
still not it's not actually working
equivalently in that way. It it
basically it's not just that we're
looking at a discrete versus continuous
and that's the only difference. It's not
quite it's like saying um you know
what's the position of a ball versus
what what is its velocity like those are
two different things to be inferred. You
can't just set them on this time horizon
like that in the same granted if you did
feel confident about your prediction of
a position of a ball and its velocity
then you hypothetically could predict
its position in the next time step. um
you know and then that would start to
look similar but um that's not what
we're actually doing in the in the
generative model right so we're not
having the agent say like oh I think the
position will be this in the next time
step versus this in the current um we're
the agent is just saying here's its flow
and sort of its change in position in
addition to the current position so u
sorry anyway yeah uh it it'll be it'll
be good to get once we actually get to
chapter 12 because yeah the code is
still being sort of built out. We we we
try to cover including all equations and
figures by time of the chapter being uh
discussed in uh on the week to week. So
hopefully this will be more filled out u
you know by then.
>> Well, one last thing before I I end the
recording um and then we can maybe ask
some some questions offline. If you go
to the code section, um this is very
higgledygly just now, but my Beijian
thermostat example and Andrew's um uh
example from 8.1 are there in terms of
um coder, sorry, um Google collab
notebook. So if you click on them, what
I would recommend that you do, oh yes, I
want to share this tab instead. What I
would recommend that you do, so this is
his example for 8.1. You can just click,
you know, run the cells and everything
like that. I would go up to file and
then save. Whoops. Save a copy in drive.
And that will then you you you will have
a copy with which to do everything.
Otherwise, you'll change the notebook
for everyone else. Doesn't matter
because we have um
the you know the the originals
ourselves. But um if you didn't want to
clone the notebook and go through all of
that headache, you can go there instead
to play around at least with those first
two examples. We'll we'll clear that up
and we'll hopefully have a bit a few
more that you can play with as well.
All right, I'm going to end the
recording there on the side of YouTube
people, but um thank you. We'll see you
next week. If anyone wants to ask a
question, hang around, they can. So
otherwise, thanks. And then we'll we'll
jump into discreet stuff. Be good. See
you.