Tokenomics Foundation: Open Standards for Managing AI Infrastructure Cost | Mike Fuller
Watch on YouTubeVideo summary
The Phops Foundation has recently been established within the Linux Foundation to address a critical gap in managing AI infrastructure costs through an open standards approach known as tokconomics. As organizations rapidly adopt artificial intelligence for both internal productivity and customer-facing products, they face increasing uncertainty regarding token usage and associated expenses. The foundation aims to create a dedicated space where key industry members can collaborate on defining how tokens serve as the core unit of the AI economy. By bringing together diverse stakeholders who are actively delivering AI features globally, the initiative seeks to standardize conversations around cost visibility, efficiency metrics, and return on investment, ensuring that financial decisions align with actual business outcomes rather than just raw technology consumption.
A central challenge addressed by the foundation is the complexity of understanding what drives costs within modern AI stacks, which often appear as an amorphous mix of compute resources, data processing, model behaviors, and energy usage. While token cost calculations are straightforward when using cloud APIs with fixed pricing models, they become significantly more complicated in hybrid environments where organizations route prompts across different locations or utilize a mixture of frontier models and open-weight alternatives. The foundation recognizes that optimizing for lower costs can sometimes negatively impact the quality of AI outputs if not managed carefully, creating a delicate balance between efficiency and performance. Consequently, efforts are focused on developing architectural diagrams and common languages that help engineers understand exactly which components contribute to expenses, allowing teams to make informed choices without compromising the long-term technological or cultural viability of their solutions.
To effectively manage these costs, organizations must move beyond simple billing statements and develop advanced telemetry capabilities that pair financial data with detailed observability metrics at the session or operation level. Lessons learned from taming cloud spending suggest that granular visibility is essential for identifying which teams, applications, and specific operations are driving high token consumption. The foundation prioritizes improving interoperability across different models and platforms by establishing a common vocabulary while allowing room for vendor-specific nuances, similar to how Finos Foundation handled differences between various cloud providers. Success for the tokconomics foundation will be measured by reducing organizational uncertainty around AI spend forecasts, enabling businesses to confidently forecast costs with single-digit accuracy and ensuring that investment dollars are directed toward tokens that generate genuine business value rather than wasted experimentation.
For organizations just beginning their journey in measuring token economics, the recommended first step is to strategically scope which areas of their AI usage they intend to tackle initially, whether focusing on internal productivity tools or external product services before attempting a comprehensive overhaul simultaneously. Practical actions include implementing observability metrics within agent flows that often drive high consumption through retry loops and parallel prompting, while also advocating for more detailed billing data from model providers who currently offer limited granularity. It is important for businesses to accept that some token spend without immediate return on investment is part of the learning curve inherent in any emerging technology stack; however, governance frameworks should be established early to set spending expectations and implement guardrails against unchecked consumption. Ultimately, the goal is to foster a culture where companies can confidently navigate AI adoption by continuously refining their understanding of cost drivers and best practices without sacrificing output quality or innovation potential.
Read the full video transcript
Hi, this is Yosin Bharti and today we
have with us Mike Fuller, CTO of the
Phops Foundation. Mike, it's great to
have you on the show.
>> Thanks for having me on.
>> It's my pleasure. Uh, Linux Foundation
recently announced the intent to form
the tokconomics
tokconomics foundation. Uh, first of
all, uh, talk a bit about that
foundation. Yeah. So, you know, over the
last sort of six months, definitely in
the last four four months, we really saw
uh a need for for an open um space for
the conversation to be had around the AI
value and how how organizations are
approaching their AI practices. Um the
Finos Foundation which I'm part of is
you know has always looked at the
technology value stack but um we
realized that there's sort of dimensions
on on this conversation um that sort of
go beyond what PHOPS traditionally had
been covering and so to give it a
dedicated space uh within the Linux
Foundation for that conversation to be
had uh and to invite uh key members um
that you know that are using or being
part of the delivery of AI features um
in order to help all organiz
organizations globally around the use of
tokens.
>> The foundation talks about tokens as the
core unit of the AI economy. I'm heavy
user of AI. I think 99% of the stuff
that I do uh these days is through AI
whether it's my workshop or of course
work. Uh and token is what actually we
sweat. Oh my god, so much tokens is
going to cost us a lot. Uh can you talk
about how are you helping organizations
connect token usage to bigger measures
like efficiency, ROI and overall value
because sometimes what happen is that
these things don't connect very well.
>> Yeah. So I think um similar to to you
know the journey that we saw in cloud
just just much faster um and I think a
little bit more confusing because it's
not technology stack that we're quite
familiar with in the past. Um so first
and foremost we really need to start to
help people understand uh not just what
is a token but what is all that
infrastructure that goes around that
token. Um and then um once you
understand the the visibility of the
cost that that's being um charged there
is then starting to look at the cost per
outcome or the tokens per outcome um
within um the the stack that that people
are using. So there's there's sort of
two dimensions that we're seeing um
companies look at. there's the the
internal AI or the productivity AI and
then the the product AI. So the the AI
used for for organizations um product
you know services and product that
customerf facing and so just trying to
figure out like uh a good way to measure
and get your arms around that spend and
then to actually associate it to the
outcomes that the business is seeing
from the use of the AI itself. When it
comes to AI, you know, it it AI
workloads, they bring together, of
course, compute, data, and model
behavior all at once. How is tokconomics
foundation approaching standardizing
these pieces so teams can actually
understand cost and performance instead
of looking at many different things.
Yeah, I think that's um you know one of
the things that I think we need to
resolve in a lot of engineers heads is
the the AI stack or especially when you
get to you know Gentic harnesses and
stuff like that it it kind of looks like
an an amorphous blob of technology cost
coming out of uh this AI inference layer
and we're trying to sort of build that
architectural diagram in inside of
everyone's heads understanding exactly
what what what components are in that
architectural diagram and which ones
cost what what um you know how we can
think about the actual uh drivers of
cost within that architecture and then
allow teams to um you know understand
why all of the choices they're making
within that that technology stack. So I
think the um you know it's the token
itself is something that it sort of
boils it down to something that's fairly
simple because you can just put a dollar
figure to a token especially when you're
using it from uh the clouds or the
frontier model providers. Um, but it
gets more complicated as you start to do
a mixture of technology stacks um and
and routing different prompts to
different locations about what that is
actually costing the business and and
how that associates to the outcome
>> based on your interactions with the
community with organizations. How much
are people worried about token token
cost? Yeah. So I I think um you know we
saw state of Phops data over the last
couple of years move from uh most PHOPS
teams uh thinking about AI cost to now
pretty much every one of them managing
AI cost in some way shape or form. Uh at
a recent um conference in San Diego we
we had a lot of a focus on AI and it was
a mixture of large parts of the
community going yes we've needed this.
It's right where we are. And then it was
funny because we had individual
interactions with a few practitioners
that were like, "Oh, I'm not sure about,
you know, if this is really a thing
yet." And then we had phone calls from
them a week later after the conference
saying that I got back to my desk and my
CTO come flying down and I'm all about
trying to manage our AI spend now. Um
and so I do think that it's either right
on the the cusp for most um
organizations to start thinking about
the AI value conversation or it's
already you know a high level concern
today
>> when it come to token cost I and I may
be totally wrong it I think is
applicable only when we are using APIs
but if we are running models locally
then token cost doesn't it only depends
on you know the context window and other
things so we are when we talk about
token cost is is mostly about uh API
consumption is that correct No. So I
think um you know when we look at um
tokens that are generated locally um you
we we do start from you know simple
components that go into the ingredients
that go into that. So your your energy,
your cooling, your your your space for
the equipment. Uh there's the whole
procurement cycle of of equipment and
the life cycle management of the
equipment itself and then the the actual
architectures of that hardware and the
models and engines that you choose to
run on it. All of that effectively uh
goes into the cost of having inference
um which you know is the token at the
end of it um in order for you to gain
value from AI. So it's basically that's
one good example of the surface area
that was kind of beyond what a
traditional phops practitioner was
looking at. Um it's sort of going back
to some of those old roots of capacity
management and and hardware acquisition
in the data center. um and and it is all
um driven around getting to a point of
having an efficient token generation
within the data center.
>> There are some effort if I'm not wrong
like Enthropic cloud and of course the
fable came out other things uh the new
version since I use heavily so I am uh
they try to optimize it but it directly
affected the quality as well. Uh so the
thing is if you try to tame uh the hook
and consumption cost it may directly
affect what is the long-term solution
technologically, socially or you know
just uh culturally.
>> Yeah, I think that that is one big um
difference when we look at optimizations
from you know where we've looked at them
in the past where usually you'll have a
fairly sort of fixed static
understanding of the workload capacity
needs and it's really just fitting the
hardware to that. Um in this case some
of the optimizations actually affect the
quality of the of the work that's being
done and it's trying to figure out um
good practices around balancing those.
We we effectively have our own series of
parameters at the same time moving up
and down if you will on the optimization
side. Uh where does it go? Um you know I
think it's going to be an improvement in
tooling potentially using AI itself to
help us u make the choices. um you know
definitely see that there's a lot of
opportunity for some experimentation and
best practice development in that that
area.
>> If I look at the foundation you folks
have broken tokconomics into three major
buckets production consumption and of
course monetization from technical
standpoint where do you see the biggest
opportunity right now is there to
improve efficiency or visibility? Yeah,
I think um consumption is probably the
one where the conversation goes to
naturally um especially if you are using
a lot of the frontier models via an API
like you say so that the tokens
generated outside of your your area but
it they're um you know is trying to
figure out exactly for each org um you
know where those opportunities really
lay and I think consumption is going to
become um a more of an interesting one
for the wider industry as we continue to
see the open weight models um you know
improving in quality. The choices of um
you know a mix between the frontier
models using hosted open open um weight
models or running open models yourself
um will become part of the conversation
with with organizations to have and we
don't want to uh you know swing the
pendulum too far the wrong way and and
and drop the quality of AI and and drive
up the amount of equipment we have to
now manage just to to do the AI
inference. There's an opportunity cost
balance there. And then on the value
side of things, you know, the the
volatility, I guess, of the token price
and and the amount of token uh
consumption, you know, being
unpredictable into the future really
will impact uh the sort of value that
you're getting and the monetization of
the tokens um you especially when
they're customerf facing to tokens. And
so businesses do have a lot of
conversation to think about how they're
going to package those um prices and
costs into their products, their product
suite. And as the foundation builds out
of course open frameworks and standards
uh can you talk about what are your top
priorities for making sure that
everything stays of course interoperable
and clear across different models and
platform because everybody is mixing
models they're using different
platforms. Yeah, I think just just the
same as you know we saw with with FOPS
when it come to the different cloud
providers and them having different
terminology and different um structures
there are some underlying similarities
of course um so it's going to be a
mixture of finding the common language
um and trying to build that language
across the way we talk about AI and AI
infrastructure and then um specializing
into um you know leaning in where
there's particular terms or particular
types of activities that are specific to
um you know one or two vendors And so
that that there there's a base level of
common understanding and then
particularly um you know hot areas um
being covered specifically where they
are unique in in particular pockets.
>> When I talk to you of course the the
cost the whole phops you know it started
when we started to tame cloud cost and
there are some clear parallels. Now the
difference is that cloud itself cannot
solve or fix the problem of cloud cloud
cost or ingress ingress fee. AI can help
solve some of its token cost problems.
Uh can you talk a bit about what kind of
telemetry or measurement cap
capabilities do teams need to really
manage AI cost and outcomes and what
lessons if any we have learned from
taming cloud cost?
>> Yeah. So I think um the the the cloud
bill and and working with a detailed
billing file is is definitely a skill
that's going to be um come into high
value here. Um but what we are seeing
with the AI um is we can't just lub the
whole activity of AI into a cloud bill
or you know like structure you know
especially when it comes to the open
specification we have like focus because
it will grow it will just explode the
granularity of that billing data and so
I think it is one area where we are
going to have to learn uh to be quite uh
become quite fluent in pairing up a
billing data set with your hotel um
observability data. set. So, um you're
really looking at, you know, do we need
to have down to session or down to
individual operation in the cloud build?
Probably not. But we should have some
telemetry around that because we we
can't really just say, hey, we've spent
this number of tokens per hour over the
last month. We need to be starting to
break that down to these are the teams
that are consuming them. These are the
applications that are using them. um you
know these are the particular types of
expensive operations we're doing um that
are enabling us to actually get an
understanding of the cost and the cost
opportunity um that's there.
>> I mean it's not that somebody's getting
a started but a lot of organizations
they're already in the middle of their
you know AI journey and it gets so
exciting that they they totally forget
about the token cost and it's only when
they get the bills then they realize it.
Uh what are the first practical step you
would recommend to these organizations?
Somebody who is getting started let's
say uh towards measuring token economics
inside their own AI stack so they can
control it before it goes out of
control.
>> Yeah. So I think what we're seeing with
the the sort of more advanced practices
that are that are on the the leading
edge of this curve first they're trying
to decide exactly how much of the AI um
pie they're trying to tackle. And so
they will look at, you know, is it the
internal AI, is it the product AI? Um,
you know, are we trying to tackle both
at once or we going to start with one
area and then, um, expand out. So I
think the, you know, trying to tackle,
uh, key areas of your AI spend and not
trying to grab it all at once is
probably first recommendation. Um, you
know, the productivity AI is one that
seems to get a lot attention because of
uh it's it's where teams are using it in
agent flows and you know with those
having retry loops and parallel
prompting and all sorts of things that
can drive that token consumption up and
so it's um trying to find that you the
amount of surface area that you'll look
at driving for that visibility. So the
mixture of putting in um observability
metrics um on on that use and also
looking at what billing you have
available. Unfortunately we are in a in
a world kind of where we were right at
the beginning of cloud where the billing
data is uh quite nent and not not
detailed enough. Um and so there's a you
know for for us there's a concerted
effort to try and get practitioners to
push for better billing data from the
model providers and um cloud providers
and the frontier model providers in
order to get uh that granularity we need
in order in in order to get the
visibility and understanding of the cost
that's there. So some some early steps
would be yeah just deciding how much of
this pie you want to buy and then trying
to push for better telemetry and billing
data. One thing that we have learned
from AI is not ask what how things will
look like five years from now. If you
can tell me how things will look like
five days from now that would be
[laughter] that great but if you look at
the foundation uh what would success
look like for tokconomics foundation
where you'll see this is what we wanted
to do and this is what we have achieved.
>> I think is um first and foremost you
know we reduce the amount of uncertainty
that organizations have when they look
at their AI spend. Um, you know, I think
that we've seen that transition when we
look at the cloud spend journey that we
started out where most organizations
were worried about where the cloud bill
was going. They weren't sure if they had
control of it. They they were worried
that it would outstrip their revenue
growth. Um, I think today we we feel
most practitioners talk about their, you
know, their their uh cost um cost trae
sorry cost forecasts to be around um,
you know, singledigit low singledigit
forecast. you know, we need to be in
that sort of um area in the next, you
know, few years, hopefully less. It's
going to move so quickly, um where we're
feeling confident that our forecasts on
AI spend are, you know, close to what we
end up with. Um and businesses know uh
you know, where they're putting their
investment dollar on AI. So, it is a um
it's more of a confidence generation um
for success for the tokconomics
foundation. Can we get to a point where
businesses feel confident? What are the
best practices you would recommend for
folks so they don't compromise on how
they use AI, they don't compromise on
the the output they get, but they can
still contain the token cost.
>> Yeah, I think you know when it comes to
very large organizations with deeper
pockets, I think it's quite common that
they will throw money at the wall with
their innovation. Um, you know, I think
it it's two things. It's one trying to
find the next business differentiator or
it's B trying to make sure that their
workforce is as productive as they
possibly could be. So it's you know
spend the tokens to get there really
good. I think for most organizations
however there will be some level of
governance that comes into play uh
upfront where you're setting some
expectations of spend on tokens out with
your organizations. you you you're
seeing more and more cost control levers
being implemented in the frontier models
um and and on the cloud platforms as far
as AI token consumption use. So I think
it's going to be just a smart level of
guard rails being put into place for
most organizations who don't have uh you
know hundreds of millions of dollars to
spend on experimentation. Um and then
looking at identifying where those T
tokens been spent is actually having a
good business impact and where they're
not and trying to reshuffle um you know
those investments as early as possible
because if you want the sort of the most
outcome uh you want to be putting the
investment on the right token. Um, I
think the main thing though is is that
for all organizations is they're going
to have to be comfortable with some
token spend not having a return that
it's it's part of this learning
exercise. We saw it um, you know, with
any technology where you start to figure
out exactly how it has value and and
then to slowly learn where to invest
better in a technology stack. And I
think AI is just this on on on your
accelerated timelines. Mike, thank you
so much uh for joining us and sharing
how the Tokconomics Foundation is going
to tackle this problem. Uh thanks for
your time and as usual, I look forward
to chat with you again. Thank you.
>> Well, thanks. Uh it's great to be uh on
your show. Um you I think you know for
us it's it's just being making it clear
that we don't have all the answers. uh
we've created a space specifically for
us to explore and and and work through
all of the questions, especially those
that you've given us today, and continue
to refine the answers to be um you know,
right on point and and correct and and
develop those best practice frameworks
to to give companies better guidance in
this space.
>> Excellent. Thank you. For those who are
watching, please go and check out
tokconomopics uh foundation and since
it's all open source, please also get
involved. Mike, once again, thank