Submind YouTube summaries
Thumbnail for Tokenomics Foundation: Open Standards for Managing AI Infrastructure Cost | Mike Fuller

Tokenomics Foundation: Open Standards for Managing AI Infrastructure Cost | Mike Fuller

Watch on YouTube

Video summary

The Phops Foundation has recently been established within the Linux Foundation to address a critical gap in managing AI infrastructure costs through an open standards approach known as tokconomics. As organizations rapidly adopt artificial intelligence for both internal productivity and customer-facing products, they face increasing uncertainty regarding token usage and associated expenses. The foundation aims to create a dedicated space where key industry members can collaborate on defining how tokens serve as the core unit of the AI economy. By bringing together diverse stakeholders who are actively delivering AI features globally, the initiative seeks to standardize conversations around cost visibility, efficiency metrics, and return on investment, ensuring that financial decisions align with actual business outcomes rather than just raw technology consumption. A central challenge addressed by the foundation is the complexity of understanding what drives costs within modern AI stacks, which often appear as an amorphous mix of compute resources, data processing, model behaviors, and energy usage. While token cost calculations are straightforward when using cloud APIs with fixed pricing models, they become significantly more complicated in hybrid environments where organizations route prompts across different locations or utilize a mixture of frontier models and open-weight alternatives. The foundation recognizes that optimizing for lower costs can sometimes negatively impact the quality of AI outputs if not managed carefully, creating a delicate balance between efficiency and performance. Consequently, efforts are focused on developing architectural diagrams and common languages that help engineers understand exactly which components contribute to expenses, allowing teams to make informed choices without compromising the long-term technological or cultural viability of their solutions. To effectively manage these costs, organizations must move beyond simple billing statements and develop advanced telemetry capabilities that pair financial data with detailed observability metrics at the session or operation level. Lessons learned from taming cloud spending suggest that granular visibility is essential for identifying which teams, applications, and specific operations are driving high token consumption. The foundation prioritizes improving interoperability across different models and platforms by establishing a common vocabulary while allowing room for vendor-specific nuances, similar to how Finos Foundation handled differences between various cloud providers. Success for the tokconomics foundation will be measured by reducing organizational uncertainty around AI spend forecasts, enabling businesses to confidently forecast costs with single-digit accuracy and ensuring that investment dollars are directed toward tokens that generate genuine business value rather than wasted experimentation. For organizations just beginning their journey in measuring token economics, the recommended first step is to strategically scope which areas of their AI usage they intend to tackle initially, whether focusing on internal productivity tools or external product services before attempting a comprehensive overhaul simultaneously. Practical actions include implementing observability metrics within agent flows that often drive high consumption through retry loops and parallel prompting, while also advocating for more detailed billing data from model providers who currently offer limited granularity. It is important for businesses to accept that some token spend without immediate return on investment is part of the learning curve inherent in any emerging technology stack; however, governance frameworks should be established early to set spending expectations and implement guardrails against unchecked consumption. Ultimately, the goal is to foster a culture where companies can confidently navigate AI adoption by continuously refining their understanding of cost drivers and best practices without sacrificing output quality or innovation potential.
Read the full video transcript
Hi, this is Yosin Bharti and today we have with us Mike Fuller, CTO of the Phops Foundation. Mike, it's great to have you on the show. >> Thanks for having me on. >> It's my pleasure. Uh, Linux Foundation recently announced the intent to form the tokconomics tokconomics foundation. Uh, first of all, uh, talk a bit about that foundation. Yeah. So, you know, over the last sort of six months, definitely in the last four four months, we really saw uh a need for for an open um space for the conversation to be had around the AI value and how how organizations are approaching their AI practices. Um the Finos Foundation which I'm part of is you know has always looked at the technology value stack but um we realized that there's sort of dimensions on on this conversation um that sort of go beyond what PHOPS traditionally had been covering and so to give it a dedicated space uh within the Linux Foundation for that conversation to be had uh and to invite uh key members um that you know that are using or being part of the delivery of AI features um in order to help all organiz organizations globally around the use of tokens. >> The foundation talks about tokens as the core unit of the AI economy. I'm heavy user of AI. I think 99% of the stuff that I do uh these days is through AI whether it's my workshop or of course work. Uh and token is what actually we sweat. Oh my god, so much tokens is going to cost us a lot. Uh can you talk about how are you helping organizations connect token usage to bigger measures like efficiency, ROI and overall value because sometimes what happen is that these things don't connect very well. >> Yeah. So I think um similar to to you know the journey that we saw in cloud just just much faster um and I think a little bit more confusing because it's not technology stack that we're quite familiar with in the past. Um so first and foremost we really need to start to help people understand uh not just what is a token but what is all that infrastructure that goes around that token. Um and then um once you understand the the visibility of the cost that that's being um charged there is then starting to look at the cost per outcome or the tokens per outcome um within um the the stack that that people are using. So there's there's sort of two dimensions that we're seeing um companies look at. there's the the internal AI or the productivity AI and then the the product AI. So the the AI used for for organizations um product you know services and product that customerf facing and so just trying to figure out like uh a good way to measure and get your arms around that spend and then to actually associate it to the outcomes that the business is seeing from the use of the AI itself. When it comes to AI, you know, it it AI workloads, they bring together, of course, compute, data, and model behavior all at once. How is tokconomics foundation approaching standardizing these pieces so teams can actually understand cost and performance instead of looking at many different things. Yeah, I think that's um you know one of the things that I think we need to resolve in a lot of engineers heads is the the AI stack or especially when you get to you know Gentic harnesses and stuff like that it it kind of looks like an an amorphous blob of technology cost coming out of uh this AI inference layer and we're trying to sort of build that architectural diagram in inside of everyone's heads understanding exactly what what what components are in that architectural diagram and which ones cost what what um you know how we can think about the actual uh drivers of cost within that architecture and then allow teams to um you know understand why all of the choices they're making within that that technology stack. So I think the um you know it's the token itself is something that it sort of boils it down to something that's fairly simple because you can just put a dollar figure to a token especially when you're using it from uh the clouds or the frontier model providers. Um, but it gets more complicated as you start to do a mixture of technology stacks um and and routing different prompts to different locations about what that is actually costing the business and and how that associates to the outcome >> based on your interactions with the community with organizations. How much are people worried about token token cost? Yeah. So I I think um you know we saw state of Phops data over the last couple of years move from uh most PHOPS teams uh thinking about AI cost to now pretty much every one of them managing AI cost in some way shape or form. Uh at a recent um conference in San Diego we we had a lot of a focus on AI and it was a mixture of large parts of the community going yes we've needed this. It's right where we are. And then it was funny because we had individual interactions with a few practitioners that were like, "Oh, I'm not sure about, you know, if this is really a thing yet." And then we had phone calls from them a week later after the conference saying that I got back to my desk and my CTO come flying down and I'm all about trying to manage our AI spend now. Um and so I do think that it's either right on the the cusp for most um organizations to start thinking about the AI value conversation or it's already you know a high level concern today >> when it come to token cost I and I may be totally wrong it I think is applicable only when we are using APIs but if we are running models locally then token cost doesn't it only depends on you know the context window and other things so we are when we talk about token cost is is mostly about uh API consumption is that correct No. So I think um you know when we look at um tokens that are generated locally um you we we do start from you know simple components that go into the ingredients that go into that. So your your energy, your cooling, your your your space for the equipment. Uh there's the whole procurement cycle of of equipment and the life cycle management of the equipment itself and then the the actual architectures of that hardware and the models and engines that you choose to run on it. All of that effectively uh goes into the cost of having inference um which you know is the token at the end of it um in order for you to gain value from AI. So it's basically that's one good example of the surface area that was kind of beyond what a traditional phops practitioner was looking at. Um it's sort of going back to some of those old roots of capacity management and and hardware acquisition in the data center. um and and it is all um driven around getting to a point of having an efficient token generation within the data center. >> There are some effort if I'm not wrong like Enthropic cloud and of course the fable came out other things uh the new version since I use heavily so I am uh they try to optimize it but it directly affected the quality as well. Uh so the thing is if you try to tame uh the hook and consumption cost it may directly affect what is the long-term solution technologically, socially or you know just uh culturally. >> Yeah, I think that that is one big um difference when we look at optimizations from you know where we've looked at them in the past where usually you'll have a fairly sort of fixed static understanding of the workload capacity needs and it's really just fitting the hardware to that. Um in this case some of the optimizations actually affect the quality of the of the work that's being done and it's trying to figure out um good practices around balancing those. We we effectively have our own series of parameters at the same time moving up and down if you will on the optimization side. Uh where does it go? Um you know I think it's going to be an improvement in tooling potentially using AI itself to help us u make the choices. um you know definitely see that there's a lot of opportunity for some experimentation and best practice development in that that area. >> If I look at the foundation you folks have broken tokconomics into three major buckets production consumption and of course monetization from technical standpoint where do you see the biggest opportunity right now is there to improve efficiency or visibility? Yeah, I think um consumption is probably the one where the conversation goes to naturally um especially if you are using a lot of the frontier models via an API like you say so that the tokens generated outside of your your area but it they're um you know is trying to figure out exactly for each org um you know where those opportunities really lay and I think consumption is going to become um a more of an interesting one for the wider industry as we continue to see the open weight models um you know improving in quality. The choices of um you know a mix between the frontier models using hosted open open um weight models or running open models yourself um will become part of the conversation with with organizations to have and we don't want to uh you know swing the pendulum too far the wrong way and and and drop the quality of AI and and drive up the amount of equipment we have to now manage just to to do the AI inference. There's an opportunity cost balance there. And then on the value side of things, you know, the the volatility, I guess, of the token price and and the amount of token uh consumption, you know, being unpredictable into the future really will impact uh the sort of value that you're getting and the monetization of the tokens um you especially when they're customerf facing to tokens. And so businesses do have a lot of conversation to think about how they're going to package those um prices and costs into their products, their product suite. And as the foundation builds out of course open frameworks and standards uh can you talk about what are your top priorities for making sure that everything stays of course interoperable and clear across different models and platform because everybody is mixing models they're using different platforms. Yeah, I think just just the same as you know we saw with with FOPS when it come to the different cloud providers and them having different terminology and different um structures there are some underlying similarities of course um so it's going to be a mixture of finding the common language um and trying to build that language across the way we talk about AI and AI infrastructure and then um specializing into um you know leaning in where there's particular terms or particular types of activities that are specific to um you know one or two vendors And so that that there there's a base level of common understanding and then particularly um you know hot areas um being covered specifically where they are unique in in particular pockets. >> When I talk to you of course the the cost the whole phops you know it started when we started to tame cloud cost and there are some clear parallels. Now the difference is that cloud itself cannot solve or fix the problem of cloud cloud cost or ingress ingress fee. AI can help solve some of its token cost problems. Uh can you talk a bit about what kind of telemetry or measurement cap capabilities do teams need to really manage AI cost and outcomes and what lessons if any we have learned from taming cloud cost? >> Yeah. So I think um the the the cloud bill and and working with a detailed billing file is is definitely a skill that's going to be um come into high value here. Um but what we are seeing with the AI um is we can't just lub the whole activity of AI into a cloud bill or you know like structure you know especially when it comes to the open specification we have like focus because it will grow it will just explode the granularity of that billing data and so I think it is one area where we are going to have to learn uh to be quite uh become quite fluent in pairing up a billing data set with your hotel um observability data. set. So, um you're really looking at, you know, do we need to have down to session or down to individual operation in the cloud build? Probably not. But we should have some telemetry around that because we we can't really just say, hey, we've spent this number of tokens per hour over the last month. We need to be starting to break that down to these are the teams that are consuming them. These are the applications that are using them. um you know these are the particular types of expensive operations we're doing um that are enabling us to actually get an understanding of the cost and the cost opportunity um that's there. >> I mean it's not that somebody's getting a started but a lot of organizations they're already in the middle of their you know AI journey and it gets so exciting that they they totally forget about the token cost and it's only when they get the bills then they realize it. Uh what are the first practical step you would recommend to these organizations? Somebody who is getting started let's say uh towards measuring token economics inside their own AI stack so they can control it before it goes out of control. >> Yeah. So I think what we're seeing with the the sort of more advanced practices that are that are on the the leading edge of this curve first they're trying to decide exactly how much of the AI um pie they're trying to tackle. And so they will look at, you know, is it the internal AI, is it the product AI? Um, you know, are we trying to tackle both at once or we going to start with one area and then, um, expand out. So I think the, you know, trying to tackle, uh, key areas of your AI spend and not trying to grab it all at once is probably first recommendation. Um, you know, the productivity AI is one that seems to get a lot attention because of uh it's it's where teams are using it in agent flows and you know with those having retry loops and parallel prompting and all sorts of things that can drive that token consumption up and so it's um trying to find that you the amount of surface area that you'll look at driving for that visibility. So the mixture of putting in um observability metrics um on on that use and also looking at what billing you have available. Unfortunately we are in a in a world kind of where we were right at the beginning of cloud where the billing data is uh quite nent and not not detailed enough. Um and so there's a you know for for us there's a concerted effort to try and get practitioners to push for better billing data from the model providers and um cloud providers and the frontier model providers in order to get uh that granularity we need in order in in order to get the visibility and understanding of the cost that's there. So some some early steps would be yeah just deciding how much of this pie you want to buy and then trying to push for better telemetry and billing data. One thing that we have learned from AI is not ask what how things will look like five years from now. If you can tell me how things will look like five days from now that would be [laughter] that great but if you look at the foundation uh what would success look like for tokconomics foundation where you'll see this is what we wanted to do and this is what we have achieved. >> I think is um first and foremost you know we reduce the amount of uncertainty that organizations have when they look at their AI spend. Um, you know, I think that we've seen that transition when we look at the cloud spend journey that we started out where most organizations were worried about where the cloud bill was going. They weren't sure if they had control of it. They they were worried that it would outstrip their revenue growth. Um, I think today we we feel most practitioners talk about their, you know, their their uh cost um cost trae sorry cost forecasts to be around um, you know, singledigit low singledigit forecast. you know, we need to be in that sort of um area in the next, you know, few years, hopefully less. It's going to move so quickly, um where we're feeling confident that our forecasts on AI spend are, you know, close to what we end up with. Um and businesses know uh you know, where they're putting their investment dollar on AI. So, it is a um it's more of a confidence generation um for success for the tokconomics foundation. Can we get to a point where businesses feel confident? What are the best practices you would recommend for folks so they don't compromise on how they use AI, they don't compromise on the the output they get, but they can still contain the token cost. >> Yeah, I think you know when it comes to very large organizations with deeper pockets, I think it's quite common that they will throw money at the wall with their innovation. Um, you know, I think it it's two things. It's one trying to find the next business differentiator or it's B trying to make sure that their workforce is as productive as they possibly could be. So it's you know spend the tokens to get there really good. I think for most organizations however there will be some level of governance that comes into play uh upfront where you're setting some expectations of spend on tokens out with your organizations. you you you're seeing more and more cost control levers being implemented in the frontier models um and and on the cloud platforms as far as AI token consumption use. So I think it's going to be just a smart level of guard rails being put into place for most organizations who don't have uh you know hundreds of millions of dollars to spend on experimentation. Um and then looking at identifying where those T tokens been spent is actually having a good business impact and where they're not and trying to reshuffle um you know those investments as early as possible because if you want the sort of the most outcome uh you want to be putting the investment on the right token. Um, I think the main thing though is is that for all organizations is they're going to have to be comfortable with some token spend not having a return that it's it's part of this learning exercise. We saw it um, you know, with any technology where you start to figure out exactly how it has value and and then to slowly learn where to invest better in a technology stack. And I think AI is just this on on on your accelerated timelines. Mike, thank you so much uh for joining us and sharing how the Tokconomics Foundation is going to tackle this problem. Uh thanks for your time and as usual, I look forward to chat with you again. Thank you. >> Well, thanks. Uh it's great to be uh on your show. Um you I think you know for us it's it's just being making it clear that we don't have all the answers. uh we've created a space specifically for us to explore and and and work through all of the questions, especially those that you've given us today, and continue to refine the answers to be um you know, right on point and and correct and and develop those best practice frameworks to to give companies better guidance in this space. >> Excellent. Thank you. For those who are watching, please go and check out tokconomopics uh foundation and since it's all open source, please also get involved. Mike, once again, thank