Fundamentals of Active Inference (Chapter 8, Session 35) August 11, 2026
Watch on YouTubeVideo summary
Chapter 8 serves as a pivotal extension of the active generalized filtering model introduced in previous chapters, focusing on two primary directions: hierarchical structures and multiple time scales. The discussion begins by reviewing the foundational concepts where perception acts as short-term belief updating about hidden states based on instantaneous observations, while attention operates at an intermediate speed to modulate salience or relevance through second-order parameters like precision. Learning is defined here as a much slower process involving fundamental constraints such as synaptic strength changes that occur over days or years. These three distinct time scales—perception, attention, and learning—are unified within the variational free energy framework, allowing agents to handle diverse cognitive phenomena by integrating rapid recognition dynamics with slow structural adaptations without needing direct observation of every environmental variable at all levels simultaneously.
The chapter then delves into hierarchical modeling, which addresses how complex systems can be decomposed across different scales or modalities to manage computational complexity effectively. This approach is analogous to neural architectures in the mammalian cortex where lower layers handle continuous motor control and proprioceptive input while higher discrete time layers manage cognitive decision-making goals. By stacking these generative models vertically, predictions descend from upper levels to inform lower ones, while prediction errors ascend for correction. The mathematical treatment utilizes mean-field approximations to factorize joint distributions into independent components, effectively reducing the exponentially large state space of a fully coupled system into tractable subspaces that can be processed in parallel or sequentially depending on the specific algorithmic requirements and causal dependencies between layers.
A significant portion of the session is dedicated to "free action," which distinguishes itself from standard active states by representing an integral through time akin to Hamilton's principle of least action rather than a simple control signal. This concept allows for the minimization of free energy over trajectories, incorporating derivatives that track changes in parameters and precisions across the three inference layers. The resulting algorithms provide a unified view where agents can simultaneously update latent states, adjust learning rates, and modulate attention weights within a single objective function. Furthermore, these hierarchical structures align closely with neurobiological findings regarding cortical columns, neurotransmitters like dopamine and serotonin associated with precision weighting, and synaptic plasticity, offering a robust bridge between computational theory and biological implementation for embodied agents operating in dynamic environments.
Read the full video transcript
Welcome back everyone. It is
August 11th
and we are in the first of the
discussions on chapter 8.
So if we pull back to the introduction
and look at the overall context of the
chapters,
we are now in the last of the three
chapters
on the continuous time generalized
filtering style active inference
implementation.
Before we head into the discrete time
algorithms in chapter 9 and 10,
we're going to be continuing this
leveling up of 678. Six, we had the
generalized coordinates and all that
that entails in the context of just
perceptual filtering. Chapter 7
introduced active filtering where there
was the possibility or agency over the
choice of
which autonomous states or external
states were under which exogenous forces
and how that was one way to
conceptualize and model action. And now
we're moving into several other
cognitive and psychological phenomena
and exploring how they fit into this
unified approach extending in that
trajectory from def filtering.
Um I will uh
pause there if anybody wants to make any
remark on chapter 8 overall
as we kick this off.
Okay, I'll insert a brief commercial
message [gasps]
for
uh work that I published yesterday,
which is a uh
600 or something part skill oriented
learning
mesh that's built on the skill tree
framework. It's available at the active
inference institute GitHub active
skillfronts. I'll put it in the uh chat.
There's a publication and then there's
the codebase. In the codebase as per
usual style, just go into the top level
and tell your agent to run it or just do
dot /run.sh SH and you'll pop up like a
really cool web application with this
kind of micro um skill learning
environments
for people who are curious about
structured curriculum and different
paths and having quizzes and access to
different materials. Just wanted to
mention that since we're in the textbook
group. Okay.
Anyone want to start with a remark or a
question or a section of eight or just
any remark on where we are in the book
so far?
Very cool, Andrew. Didn't know that.
Very cool.
>> Thanks.
>> Y
>> yeah, we've had continued conversations
in this group where people have been
interested. Kobus and Frasier and
others. So, yeah, we just put this out.
It'll be presented at IY and um full
open source repo. Um
no requirements to use like LLMs or
anything like that. It's just kind of
notebook style tutorials. Um different
drone and Atari style like game demos
that are there. So yeah, for anyone who
wants to check it out.
>> Nice. Nice work.
Okay. So, let us go into chapter 8. So,
head over to page
179
and we'll start there. Okay. So, as
mentioned, we're continuing in this
sub series of chapter 678. Chapter 6
introduced us to generalized coordinates
and the concept of filtering as
perception. Chapter 7 introduced action
as a way to build on that generalized
filtering model and introduce the
possibility for action selection.
And now we're going to take it in
several extension directions.
Also, when we look at this topic map, we
notice that there's a lot more other
chapters mentioned, which is not the
case for all the topic maps, but it
reflects how this is building on a lot
of prior work.
Um
so we we can enter in at different
places into this image but let's start
with a generative model which we know is
like the statistical probabilistic
representation of the agent
or system of interest to the exclusion
of the generative process. So this is
the agent side that's being modeled the
perceptual internal cognitive and active
autonomous states.
Um, we
learned in chapter 7 that those active
states are the outwardfacing emission
surface essentially of the markoff
blanket of the generative model. So
that's kind of just a little review
appendage. If we walk back upstream,
we can see and remember that in chapter
4 when we were first introduced to
variational free energy, it was
described as a way to minimize this
tractable bounding heristic
in the process of recognition dynamics
which is to say the observation of
observables and then the reduction of
uncertainty regarding unobservables.
observation of thermometer measurements,
reduction of uncertainty about the
temperature in the room and those
updates to beliefs about latent states
are can be described as changes of
unobserved variables of the generative
model. So it's explicitly modeled in
that sense. It's observable to the
software but it's not a direct
observable to the agent like the
temperature in the room is not something
that's observed. thermometer readings
are are observed.
And then this um notion of free action
which as the side note will disclose to
us is hilariously confusingly named with
the same word as active states. But free
action is not active states and that's
going to be differentiated with the
notation call out as well. But it's like
a integration through time of
variational free energy which is also a
very important doorway to go from a sort
of snapshot statebased variational free
energy minimization into more of like
this gauge trajectory fiber
path formulization of basian mechanics.
So it's just like a little bit of a
placeholder for a wormhole that will
take you from state-based in
instantaneous energy minimization
techniques into path and trajectory and
gaugebased techniques to come
probably more in the next book. So
that's this upstairs part. Now chapters
chapter 8 is going to develop two
important directions of extension to the
active generalized filtering model that
we saw in chapter 7 and those are going
to be relating to broadly hierarchies
and time scales. So first on the
hierarchy direction,
uh sections 8, 2, 3, and 4 are going to
go into these hierarchical forms of the
model, which is just stacking more
layers on top of each other. And at a
sort of intuition level, this is what it
looks like to have these multiple
layers. If people are familiar with like
um thousand brains and Jeff Hawkins or
just any number of other works that look
at hierarchical compositionality for
example in the mammal cortex it's very
related to this hierarchical predictive
processing hierarchical nested bay where
the descending predictions of one layer
become like descending for another layer
and the ascending get pushed up. We'll
look at that. It's a way to encompass
multiple scales within a modality. So
for example, if we were going to model
years and we wanted to make a model that
had a scope of 100,000 years, we could
have 100,000 clicks of a model that went
one year at a time. Or we could have 100
clicks of a model that went a thousand
years at a time and then each
thousand-year bucket broke down into 10
of a 100 and so on. And so there's this
structure learning question.
How do you break down even just a string
of years into a hierarchical model? Is
it better to have a wider model that's
shallower? Is it better to have a taller
model that's more narrow? That's kind of
analogous to the question of neural
network architecture in this state space
model setting. Also, hierarchical models
can be of different modalities.
So there can be
a lower level modality that could be
something more like motor processing
like propriceptive sensory input and
descending motor control and then a
second order of that model which is more
like a cognitive decision-making layer.
That's kind of a common strategy is have
like a lower level continuous time model
that relates to motor control and
propriception and then have that
sending messages up and receiving
messages down from more of like a
cognitive decision-making goal setting
discrete time model.
So hierarchical forms are going to turn
out to be super interesting from a
neurobiology perspective and ability to
improve the computational
characteristics even working within a
modality by reducing state spaces that
are otherwise going to be very
computationally punishing
and provide this really cool
compositional way to integrate models
across different modalities.
In 81, we're going to explore this
second branch, although I guess it's
first in terms of the chapter
of multiple time scales of learning and
adaptation.
And there are three here, perception,
attention, and learning. And so it's
important to note how these terms are
used here because otherwise like
learning might be used in a more broader
like statistical learning machine
learning setting just to mean all
parameter updates are learning. So it's
like yes there's a sense in which
learning is going to be used to refer to
all parameter updates period. Especially
for this section though, it's important
to note how are perception, attention,
and learning different and how does that
relate to these sort of zeroth order,
first order, and second order parameters
of a generative model. And the sort of
short version is that perception is the
short-term updating of beliefs. For
example, about food particle size given
the light intensity or about temperature
in the room given the thermometer. So it
has to do with instantaneous
single observation at a time
updates of belief about hidden states
with respect to new observations coming
in. So that's like flow forward just
regular runtime of a neural network.
Then I'm skipping to attention even
though it's a second order parameter
because it's slow relative to perception
but it's faster relative to learning.
And attention, as we'll explore, has a
lot to do with shifts in salience or
even relevance that happen slower than
the level of updating beliefs on a first
order level, but slower than what is
called learning, which are like
fundamental constraints about the
setting.
So, we're going to go into more detail
on what those three layers are and what
happens when you fix one or two and let
the third one go, which is commonly the
case where it's like you'll have
learning and attention fixed and then
you let perception rip. Or you'll have
learning fixed and then you'll let
attention and perception rip. Or you can
do this sort of like
less constrained
simultaneous inference on all of these
as part of a unified objective function.
But that's studying something very
different. Like it's one thing to study
an animal that already knows how to walk
and then it does inference on movements
versus like motor babbling with the
animal learning how its joints relate to
propriceptive feedback which is what
happens when you open up learning.
Any thoughts or like other pieces at the
kind of chapter a overview level anyone
wants to add?
Yeah, just uh yeah, the the way the
textbook describes it, this generally
holds in a lot of other uh literature on
active inference and where these things
are sort of conceived or conceptualized
is um so hidden states
t and hidden state estimation is part of
this like triple estimation of of hidden
states and and parameters and
precisions. um tends to align like it
it's [clears throat] it's
neurohysiological like correlate would
be at the rate at which we frequently
might uh record or do ne neuro imaging
with like EEG or EMG and so it's just
like a very rapid fast-paced sort of
inference um allowing for you know
recognition of different kinds of
potential changes in the hidden state of
the environment uh and then after that
perhaps at the level of seconds and and
largely in line with the notion of like
neurom modulation uh and and
neurotransmission. So a lot of different
pre uh pre precision updates get aligned
with with uh this idea of attention in a
sense of um you know being able to sort
of amplify or or dampen the impact of
prediction errors at different places
within the model. And so it's as if uh
as Daniel used the word salience earlier
um one is sort of like attending more or
less to particular streams from which
from whence one is kind of pulling
information when eventually leading up
to this sort of integration where we're
we're computing variational free energy.
Um and so precision uh there there are
these different precision terms. They've
been uh viewed differently for those who
are interested in sort of the
neuroscientific literature on this. Um
there's a precision on on policy
inference which is something we'll get
into more later in the textbook but it
frequently gets related to the notion of
like dopamine or dop do dopamineergic
pathways. Um and then uh precisions on
likelihood get related to serotonin. Um
there's definitely more in the
literature to to kind of ponder about
all of that but it's it's very
interesting. And then finally learning
is expected to sort of happen over a
much slower period of time usually in
the span of like days or years is how
Sanjief says it. Although of course
whenever you actually implement all of
these in your own algorithm um you
obviously do not have to uh explicitly
set them to uh you know the same time
scale in in reality. It depends on where
you're employing this. If you were
trying to do this for for empirical work
with like human beings for example then
then maybe you would try to actually
strive for that sort of empirical uh uh
replication but um you know of course
you don't have to have a model that only
learns after several days or years um
you know its parameters. So it's um but
in any case it just seemed relevant to
to yeah bring up and and sort of how now
we've brought a lot of things from
preceding chapters. We've already known
about perception since chapter 2.
Attention and precision uh came in
chapter I believe five and then learning
was chapter 3. So really chapter 8 is
just a big culmination of a lot of
things we've seen thus far just before
we switch entirely over into discrete
state space models which will be quite
different but that's for two weeks from
now. So
>> great comments. Yeah. Perception in the
neurobiology
setting corresponds to like action
potential mediated standard forward
communication of neurons.
Attention corresponds to neurom
modulators like dopamine or otherwise at
the syninnapse.
Learning corresponds to changes in
synaptic strength or even synaptic
topology.
So it's just kind of interesting that we
have this very nice analytical
separation, time scale separation,
biological mechanism separation. Yeah,
Ahmed. Norepinephrine too definitely.
And in invertebrates there's tyramine
and all these other more important
biogenic amines. So totally
um and that's like one way that people
do precision psychiatry and related
systems biology modeling is like look
for some algorithmic correspondence to
biomarkers.
Okay. So in
um section 1 8.1
we can just look at that first sentence
preceding chapters have underscored the
versatility of the VF loss functional.
So that hits a big theme which is like
all of these diverse cognitive phenomena
across different time scales of course
different systems of interest the game
is going to be bringing them into a
variational free energy minimization
setting. So variational free energy has
this ability to compose across different
phenomena which makes the numbers not
directly comparable across those
settings but within a given model it's
extremely informative. And so here we
work our way towards this
triple estimation problem
which is like this stream of information
is coming in and you're learning even if
we just consider this in a perceptual
setting not with action in the loop like
a kind of chapter six setting we're
learning multiple layers of
what um the signal contains or of ways
to parse these kind of layers or time
scales of inference on that stream.
Like if we were trying to parse out what
words the radio was saying, we would
like of course want to parse those.
That's the first order. But then we also
have this layer of like the pace and the
timing just loosely using the metaphor
here of how the words are spoken when to
pay attention to what frequency bands
and then even at the slower level that's
like presupposing our knowledge of the
language and of how sounds map to phone
and phonemes maps to words and so on
which might be interesting if you were
studying language acquisition but if you
were just studying somebody who was
already fluent and just listening to the
radio. You might take learning for
granted and just focus on their
attention and perception.
Um, in section 81, which is relatively
long, we get to um
something that should be pretty familiar
in terms of approach with
equation 88.
where
we have a um
free energy variational free energy
omitting the terms that aren't part of
the lelass approximation optimization.
And then right below the equation, each
of these terms is a precision weighted
squared prediction error for the random
variable of interest with init
additional log precision term. So it's
just like how the linear regression was
the sum of squares of the residual error
and that's how the linear regression fit
through a scatter plot in the standard
regression setting. Here we're doing
something similar to take a unified kind
of like L2 norm if it were a linear
regression except that it's a free
energy loss function and it is precision
weighted squared prediction error which
we saw in chapter 4.
On the next page in section 812,
we get these three time scales stated
again the fast changing which is the mu
subx and that's what we've been talking
about in all the previous chapters
updating beliefs about the temperature
of the room from the thermometer.
Then we get mu sub
g xi lower case I believe um that quote
track slowly changing contextual effects
that lead to changes in hidden states.
So um there's all kinds of settings you
could imagine that's important. But if
there's like some sort of oscilly wave
and then the inference is like the
faster ripple on that slower wave.
Then there's the third layer which is
the very slowly changing which is
corresponds to learning in this specific
narrow sense of learning which is mu sub
theta
and those are brain states in the brain
um or just model states that track very
slowly changing nearly invariant
properties of the environment such as
physical loss. They correspond to
long-term weakening or strengthening of
connections between and among synapses.
So those are the ones where if you're
studying like how domain knowledge comes
to be acquired like how a world model is
bootstrapped
you would open that door to learnable
uh use sub theta or if you just wanted
to study like a learned or skilled agent
dropped into a situation you would
probably pre-learn
or give parameters for that and and fix
those and focus inference on just the
short-term kind of behavioral trial like
single trial behaviors.
Um
in
page 184
we get several
pretty important um big mathematical
topics.
Um 811
is going to describe
how those three time scales that were
just described, right? The um x,
c, and theta
can be
factorized
according to what is called the mean
field approximation.
which means quote the variational
density could be factorized into
independent distributions for each of
these variables of interest. So let's
just say that I'm going to roll a
20-sided die, a six-sided die, and flip
a coin. So I have this joint
distribution over three variables.
[snorts] One of them is going to be a
number from 1 to 20. One of them is a
number from 1 to six. And then one of
them is heads or tails. So that is a
quite large joint probability space.
It's like 20 * 6 * 2.
So
in a setting where those three things
don't influence each other, they're just
conditionally independent,
we can factoriize that joint
distribution from a three variable joint
distribution with a state space size of
like 6 * 20
* 2 into just three separate
distributions with a state space of 26
and two.
So this is a heristic that can take a
pretty thorny joint distribution
and carve it up at the joints to reduce
the state space which otherwise if
considered in its kind of full joint
richness is going to become
exponentially large.
Now, if there are intersectional or
causal effects, like if it is the case
that the the the number of on on the on
the six-sided die influences the 20sided
dy's outcome, the mean field is going to
lock that information.
So it's um an example of like where you
have these two extremes which is you can
consider all the variables in a joint
distribution
or you can take the full mean field
approximation and have all the variables
considered separately
or you can do partial factorization
where you can retain informative joint
distributions while partitioning off
distributions that are non-interactive
with those joint distributions.
So it's a vague topic. A lot of it
happens under the hood. A lot of it
matters more when you're talking about
like the performance of these algorithms
like the computational complexity and
how their performance and resource needs
scale with respect to state space sizes.
But that's the big idea here of mean
field factorization is like of course
the fast inference, the neurom
modulation and the learning are part of
the same integrated biological system.
But from a map math modeling view, we're
going to instantaneously model them as
not influencing each other in the sense
that whether a given observation or
another comes in, we're going to treat
that as not important for the attention
inference. And whether the attention
zigzags this way or that way
instantaneously, we're going to treat
that as unimportant for the slower
learning rate. That's what meanfield
factorization does here.
So it takes what would otherwise be a
sort of exponentially increasing joint
distribution state space and breaks it
up into these more tractable even
parallelizable
smaller substate spaces.
[gasps]
Okay.
just a few centimeters lower down we hit
another big topic which is the free
action. So first read side note 7 on the
left side. So, Sanjie writes,
"Confusingly,
free action is sometimes referred to as
simply action in the AIF literature and
and more generally, but it is distinct
and entirely different from action as a
control signal, which is like the
everyday sense of action like taking an
act. The term free action is chosen to
be analogous to Hamilton's principle of
least action, a key quantity used in
Lrangeian mechanics. But as you'll learn
if you explore mechanics and basian
mechanics and all this least action
again it doesn't mean lowest possible
movement. It doesn't mean do the least.
It means minimize this term called free
action. And so to really make that clear
that action is the integral through
time. That's the integral dt. It's kind
of like the book ends of the integral.
So we're going to integrate through
[snorts] time given time interval
this expression or functional which is
the variational free energy of
observations and the three layers of
learning or perception, attention and
learning. Three layers of perceptual
updating
and free action is going to be
symbolized with S.
Um that leads to several
very nice mathematical steps
which bring us to 8
18
where we have
um
this
numerical integration scheme
for the univariate setting. And we'll
notice we have those three
layers of inference.
Perceptual with the latent states, first
order parameters of learning, the slowly
changing, and then the second order
parameters of attention
plus
these primes
on the first and the second order.
So this is quite a consilient equation.
Um,
algorithm 12. Andrew, go ahead.
Actually, this is just a a more general
thing. Um, given like I've had a lot of
familiarity with continuous models, but
but here I was a little interested in
the choice. Um I'm I'm supposing that
the reason why we're including the uh
we're doing estimations of not only the
means of the parameters and the
precisions but also their respective
first temporal derivatives. Um I suppose
that is what gives the model a sense of
making predictions about if there are
contextual
changes that might be going on. And that
is to say like it gives it that extra
degree of depth of being able to sort of
see if the environment is changing. Are
our parameters themselves actively sort
of changing um as well as precisions?
And that's what potentially allows for
the idea of we have multiple time scales
sort of operating here whenever we look
at how we're playing out the the
algorithm overall. But um that that's
been my sort of sense of of why we're
doing this. Absolutely. Like right below
817b
it says mu theta prime and mu g prime
can be thought of a prior that the
changes in the parameters or precisions
will be near zero. So um we don't have
that for the standard first order
perception. So we can think of that as
kind of like unclamped. We're just doing
regular basian updating on perception of
latent states. And then for theta and
for for she we have both the dot which
is the rate of change and we have a dot
prime and that dot prime which is a move
that we've seen before. Let's just focus
in on the theta. So that slow learning
parameter.
This is how does free energy change with
respect to those first order learning
parameters
as dampened by a learning rate
coefficient multiplied by those
parameters.
So this gives a sort of dampening or
tuning knob so that the second or the
first and the second order parameters
the learning and the attention
parameters can be also kind of unclamped
which is to say like just doing B
optimal updating
like perception
or they can be more slowly clamped so
that they have a dampening effect so
that the perception is taking up the
slack.
So those are some pretty
big very deep topics
and we get to algorithm 12 and example
8.1
which basically go through like what
does this procedurally look like and we
see a sort of pattern we've seen before
[snorts]
where we initialize the state spaces and
then we have some loop where the
generative process is basically just
going to transition emit transition emit
transition emit.
Generative model the agent side is going
to
calculate the gradient of the
variational free energy with respect to
these three layers of inference the x
theta g.
Then we will define the flow over those
five different um they they are rates of
change I believe kus they're just the
flow over them
seven hidden state updates perception
then do first order updates learning
then second order updates attention or
you could do other orders but these are
just what to it
recalculate the prediction errors and
then the VF
and then the next observation comes in.
So this is like a VF that could include
action, but here it's actually just
pulling back to again a perceptual
example.
Okay, so that was kind of the beefy
um
section 8.1.
That was this whole downstairs part.
This idea that even for a primary
inference problem,
we can pull back the curtain two layers
and look at learning which are first
order parameters that are slowly
changing super slow and then second
order parameters attention which are an
intermediate speed updating.
Pretty cool.
Okay, then continuing now a little bit
faster to this 2 3 4 section. But write
write a question or raise your hand if
anyone wants to ask. Um, now we're going
to take a different approach to a
complimentary approach to encompassing
multiple orders of change within a given
generative model. And this one is going
to be structural. It's going to be
related to developing hierarchical
models.
So
here
in figure oh 8 point
I might have made a typo with the figure
numbers.
So we'll check. Okay. This is 8.3
though. It is okay.
Oh, wait.
Okay, this one's actually 8.3,
but the page numbers and stuff, I just
made a few typos. We'll fix it. Um
here we kind of have a call back to the
previous chapter
where we have latent state
emitting observations that observed with
this
action that's forcing the latent state.
So that's just reminding us about single
layer generative models.
And we get 8.5.
Oh man, I'm very off with the numbering.
This one's actually 8.5.
Um where we have all the edges drawn.
Um and we see
the um
latent state
as well as the attention and the first
order learning parameters. So this is
kind of like building out our sort of
simpler version of just latent state
emitting values for observations.
Now in this middle row we have those
zeroth, first and second order inference
time scales that we just discussed in
section 8.1
and we have action. So we have the
chapter 7 maneuver. So this kind of gets
us onto the same page for a active
generalized filtering model with these
three layers for the triple estimation
problem again which you might leave
fixed or you might leave open to
inference.
And section 822
is going to
take us to
figure 8.6.
where
if uh if you ignore the top Tetris
piece,
you have literally just what you had in
the previous visualization.
Here's that same row with the the
zeroith first and second order
inference.
Um,
this looks like a typo in the PDF with a
bracket that's unclosed.
And also, if you look at the caption, it
says, note that y is defined as v 0.
So, it's kind of weird in one way to
say, well, we're going to call v. The
the important thing is that the gray
means it's an observable and we're just
going to call that observable v 0 which
gives us notational continuity with
stacking these models rather than have
like y at the bottom and then have vub1
v2 v3 but this is basically a
hierarchical vertical composition of the
generative models so that what was
that exogenous force
Action
coming down for the first layer model is
itself emitted
like an observation
from the second order model [snorts]
second hierarchical level.
Exciting.
Very exciting.
And then
as we start to see Sanjieve up to some
of his tricks again, 832,
what do we see? Llas approximations
of
a little bit more sophisticated, but the
V is again the observation. And we have
essentially a weighted
free energy functional of these exact
same components but summed over the
number of layers.
So it's just like we were looking at
in
here.
here in 88
and the difference is now
we're summating it over all the levels
but we still see an observation on the
left which is V but that is equivalent
to Y here
zero order perceptual
first order learning second order
attention
Um
section 83 goes into recognition
dynamics
and brings up what is um some more
neurobiologically
grounded
um model architectures.
Again like thousand brains and other
work has explored how the cortical
columns within a cortical column you
have like a prediction error unit
and then laterally across or among
columns you have hierarchical
relationships. And so that allows for
this flexible
compositional hierarchical multicolumnar
neoortical inference architecture in
mammals.
Um section 84
goes into the complete active
generalized filtering algorithm.
So that is figure 88. So okay this is
that's
this one.
So here we have that V observable
and then we have the three
of the zero order inference, first order
inference, second order inference
and that descending V at the higher
level.
So here we have the vertical composed
stack
and then here it is unfolding through
time
where um tao and t represent basically
the absolute and the relative time step.
Um Andrew,
>> yeah just in reference to that um 8.8
and eight you were just showing um yeah
I just want to provide like a little bit
of the intuition of like whenever we've
seen the phrases um ascending prediction
errors and descending predictions um I
think it's visualizations like this that
help give a bit of a a better
understanding of you know descending
meaning to come down just as here we're
seeing in these higher layers um you
know they they make their predictions
and send them down to the lay the layers
below. Whereas with the lowest level
here, we're using zero indexing. So you
can think of L equals 0 as the first or
lowest uh layer. And typically it's the
lowest layer that we would assign to
being the one that sort of interfaces
more directly with the environment. Um
it's we can think of this as at the
lowest layer we are sort of bringing in
or sensing observations which then leads
to prediction errors which then gets
sent upward. And so that's sort of where
this terminology of ascending and
descending comes from. Of course there
are like many many potentially arbitrary
ways that one could choose to sort of
produce a model reflected in a factor
graph like this. So it's not necessarily
that you do it exactly like that, but
that's been the typical implementation
is this lowest what's conceived as the
lowest level is what directly receives
the raw observations and it's everything
at the top that's actually sort of
architecturally further removed from
receiving those observations. So most of
what they can do is just sort of receive
the uh the the the prediction errors
from below to update themselves as if
those as if the higher layer layers
aren't observing the environment
directly. They're just observing
whatever's going on at the lower layer,
the layer immediately below. So yeah,
again, you could you could do this in
many different ways, but this is the
general way it's been conceived. And
then as Daniel said, it gets aligned
with the idea of like these different
sort of cortical columns uh in in the
brain and and sort of uh superficial and
parameal cells and what they generate.
So um and the the 2022 textbook is a
really nice resource on those kinds of
dynamics too in chapter five. So I pop
that in the chat for anyone who wants to
see that.
>> Great question. Does arrow direction
suggest dependency? I'll just copy this
to the chapter eight. Um, or does it
illustrate connections? The answer is
like a little bit of both. When we talk
about a diagram, first off, there are
diagrams that are just informal
associative connections like a topic
map. And then there are diagrams that
actually represent a causal graph. And
um those connections
are causal. And the simplest way of
causality being modeled is basically
like depending on X, what does Y how
does Y differ? And if changes in X don't
change in Y, it's not a causal variable.
Whereas if changes in X influence Y,
which is to say there is a dependency
like Y given X.
So that is the relationship between the
topology of connection
and dependency as causality
in this setting of causal inference.
Um
figure 8.9
we see this um multi- columnar
hierarchical inference and it's broken
out into the top down predictions
and the bottom up errors.
That's a common neuroscientific
paradigm.
Algorithm 13 is all of page 200.
That's a big one. That's a big one. Um,
we get 810,
which is a
beautiful expansion
of these different update rules and
motifs
within this pattern. And so you can
imagine in something like ngcarn or rx
infer where you can control the topology
of different units, their semantics and
the scheduling of their updates. you can
really have a lot of control over. For
example, well, we want to update the
autonomous state first, then the then
the error, then the, you know, you can
schedule things in terms of which
updates happen on what kind of a clock
or what kind of a conditionality.
And many of these, but not not
exclusively,
reflect known neural circuit topologies.
So that's like a nice to have because a
given computational architecture might
be under such different constraints than
a biological architecture that they just
have different morphologies to do even
analogous functions. But when we do see
the computational architecture
recapitulate the topology of the neural
circuit, it's a very interesting
finding.
Um,
the next several images, including the
massive 815,
go into even more detail about these
different variables within a layer and
different messages that are passed
within and across layers.
Then
81 16
gives us this unified
view
on the objective function and the flow
functions for those layers that we've
been discussing.
And 817
is like almost as full as we've gotten
in the book in terms of like all the
variables written out in one image where
we have a super top level
single line per node
and then this like very detailed
variables. [snorts]
And then
as uh usual, the closing pages
um summarize what the chapter did. And
then the last two paragraphs mention
some related literature
and the last paragraph connects to
salience.
So it's a dense challenging chapter.
There's a lot of notation.
However, if we focus on these two
branches,
81 introducing these three time scales
of inference within a layer, the triple
estimation of perception, attention and
learning
and then different hierarchical
compositional methods within and across
modalities. it gives a lot of sense to
the chapter.
So if anyone wants to ask a question or
a final thought otherwise I'll stop the
recording.
>> Yeah. Uh from previous sessions I
learned uh you know from this active
process. So I wrote like uh execution
logic for testing this I mean how to
reduce a pred predive error and it just
a little slower.
>> Wait wait wait just just a little little
slower please. And then just first off
is this related to the textbook.
>> Uh it's related to test I mean textbook
but the the things that we learned from
the previous sessions and also from the
facial
>> okay it's a little slower but go ahead
ask the question.
>> Yeah it's not question. So the thing is
that from the things that I learned like
uh perception changes, belief and those
things from that I wrote a small
execution logic for using using that
same things as a you know engineering
way like uh logical way for uh you know
robots and autonomous system. So I
published it on Zeno. So could I share
it like you know?
>> Yeah, put it in the chat.
>> Absolutely. Put it in the chat.
>> Okay. Okay.
>> Awesome. So from this uh from this we
can and also there's one more link too.
So
wait a second.
>> Okay.
>> Yeah. Wait a second. I will copy that
link.
>> Okay. If there's a question please ask
otherwise anyone who has a question can
ask. And yeah I mean and anyone in this
session can test that thing actually
because I I provided a way there called
an appendix lost appendix they can input
their own uh you know they can give a to
the robot and they can give to like a
next robot. So
>> great
>> and
>> great awesome I hope people check it out
if they are interested in it but
absolutely how to apply this to embodied
motion and language is critical
>> then and also you can provide your own
scenario so you can it's to find like
where the robot breaks from not being a
human I mean like not being a social
human like for instance it's it's
supposed to be the robot supposed to be
similar very similar to how actual human
beings supposed to handle situation.
>> Okay. If there's any question or comment
on the textbook otherwise I'll end the
recording.
>> Yeah, of course.
>> All right.