Submind YouTube summaries
Thumbnail for Jason Goodison, General Compute | theCUBE + NYSE Wired: AI Factories

Jason Goodison, General Compute | theCUBE + NYSE Wired: AI Factories

Watch on YouTube

Video summary

Jason Goodison, CTO and co-founder of General Compute, joins the discussion to address the surging global demand for AI infrastructure and the critical need for alternatives to Nvidia's dominant GPU architecture. While major cloud providers like CoreWeave and Nebius have vertically integrated with specific chip suppliers such as Nvidia or AMD, leaving a vast array of innovative chips from companies like Cerebras, SambaNova, and others underutilized, General Compute positions itself as the essential deployment arm for these heterogeneous solutions. The company recently secured a $400 million debt facility to productionize these alternative chips, aiming to bridge the gap between exceptional engineering teams building specialized silicon and enterprises that require flexible capacity without being locked into a single vendor's ecosystem. This approach allows businesses to access cutting-edge technology while mitigating the risks associated with relying on a monolithic hardware stack. A central theme of the conversation is the technical shift from general-purpose computing to memory-heavy, specialized architectures designed for efficient inference. Goodison explains that as AI models grow larger and context windows expand, the traditional GPU architecture struggles with the "memory wall," particularly during the decode phase where data must be kept close to the silicon. Specialized chips like those from Cerebras utilize massive on-chip SRAM to keep data local, drastically reducing latency and memory bandwidth requirements compared to GPUs that constantly shuttle data back and forth. This separation of pre-fill and decode workloads allows for a more efficient use of resources, enabling faster token generation and better handling of large-scale models, which is crucial as the industry moves beyond simple training into complex, real-time inference applications. Beyond technical architecture, the discussion highlights the evolving financial and strategic landscape of AI infrastructure, where energy constraints and capital efficiency are becoming paramount. Goodison notes that while Nvidia benefits from favorable financing terms due to its established ecosystem and guaranteed demand, emerging ASIC companies face higher interest rates and a lack of secondary markets for their hardware. To solve this, General Compute targets asset-light inference clouds and enterprises that need single-tenant capabilities without massive capital expenditure, offering them the ability to rent specialized capacity on demand. The company's vision involves creating a software stack that can compile models once and run them across diverse hardware architectures, effectively democratizing access to advanced AI compute and allowing companies to optimize their workloads based on specific needs rather than vendor lock-in. Ultimately, General Compute aims to redefine the AI infrastructure market by connecting supply with demand for a wide variety of specialized chips, ensuring that innovation is not stifled by deployment bottlenecks. As the industry faces supply constraints and energy limitations, the ability to mix and match different hardware types based on TCO calculations and workload requirements will become essential. Goodison emphasizes that the future lies in a heterogeneous ecosystem where enterprises can leverage the best-in-class chips for specific tasks, whether it is high-value training or rapid inference, without taking on the prohibitive risks of deploying unproven technology alone. By acting as an aggregator and deployer, General Compute seeks to accelerate the adoption of these diverse technologies, fostering a more robust and competitive AI infrastructure landscape that can scale alongside the infinite intelligence era.
Read the full video transcript
Palo Alto studio connection Silicon Valley and Wall Street. I'm John F co here with Dave Volante my co-host. Hello, I'm John Furry, host of the cube. Here in our PaloAlto studio, of course, we have the cub's NYC studio connecting Silicon Valley to Wall Street, Wall Street to Silicon Valley, the NYC wired programs where technology and Wall Street intersect. Of course, the cub's got the deep coverage, tracking all the semis, all the AI infrastructure buildouts. This is our AI factory series where we talk to the leaders who are building it and setting the table for the era of AI. Jason Goodison is here. He's the CTO and co-founder of General Compute. Thanks for coming in, popping into the studio today. Appreciate it. >> Thank you for having me. >> You're doing some pretty cool work. Again the demand for AI infrastructure is off the charts and again the demand for intelligence is feeding in obviously the coding which we're everyone's seeing the value there physical AI edge right in line so you can see the progression the trajectories forming there's just way too much demand and then but the role of the A6 and the silicon and software we're seeing the different approaches from Nvidia Google Cerebras AMD um all have kind of the bets Yeah, >> CUDA obviously with Nvidia that's software evolution programmable TPUs for Google they got uh you know different approach with the pods you got ser the big wafer the general purpose to custom and you know specialized compute have always been that spectrum the harder you go to specialize the harder is to program the use cases are more narrow but with AI that's all being bundled together >> um you're building out your venture and in this demand curve A lot's going on. I want to unpack with you, but let's get with what you guys are doing right now. What's the state of your company? What's the thesis? >> Yeah. >> Where are you guys seeing the action? >> Yeah. So, um, fundamentally what we do is we like to see ourselves as the deployment arm for these alternative, uh, chips, alternatives to GPUs. Um, so if you think about all of these up and cominging chips, you've got Cerebrris, Samanova, I mean those companies have been around a long time. You've got TensorDine, Tensor, Posetron, Dmatrix, there's just the list goes on, >> boatload of new stuff coming too >> and all of these engineers are just fantastic, right? Um they've built these incredible chips that work for different use cases. Uh but we have a lot of the asset heavy clouds of the world, the Nebuses, the Core Weaves, um that are kind of locked into their chip supplier. So you think about Coreweave, they they have a lot of circular financing with Nvidia and so they only deploy Nvidia. Um, TensorWave only really deploys AMD. They might do other stuff in the future, I don't know. Um, and then Fluid Stack does a lot with Google TPU. So, there's all of these other chips in the world that are really fantastic and the engineering teams there are exceptional. Um, but there's no one that just takes on that debt and deploys them uh and then rents them back bare metal to an enterprise or to another cloud that needs the capacity. So, that's what we do. So, a few weeks ago, we just announced a $400 million debt facility um in combination with Upper 90. Upper 90 uh joined the cap table and we're really excited to work with them and uh productionize a lot of these awesome chips. >> Yeah, there's a race. I mean you can see the vertical integration you know core weave ncale they just bought any scale. So you can start to see the disagregated serving kicking in here at hot chips Stanford I was poking around yesterday top conversation is hey the capaca capability and capacity demand is high the KV caches are getting stuffed. So you're starting to see the patterns. More data is coming in, more demand for the tokens aka intelligence. So it's going to put pressure on the architecture. So okay, the big guys, they got their partners, but there's a whole another on on boarding of these new NeoClouds, Neolabs, or someone's got some Bitcoin, they got some data center facilities. Those are different businesses, but they got the energy, they got the footprint, they want to bring that into this new buildout. So it's almost as if there's a new breed >> Yeah. of infrastructure opportunity. Sounds like that's what you're targeting. >> Yeah, absolutely. So, when you think about any major technological innovation, um you start off by u not really understanding the problem and so you just use what you have at your disposal uh to solve the problem. So, even when you think about Bitcoin back in the day, you know, people just mine Bitcoin on GPUs. Um now nowadays people mine Bitcoin on A6 because as the uh as the workload stabilizes you just get so many um advantages to baking that into the silicon right um and so what we're seeing we're a little bit in the wild west period right now in AI because there's a new model drop uh every month um an open source model comes out there's a new attention mechanism um there's there's different infrastructure and architectures that are coming in the AI models so we have all these chips and they're all racing. They all have their own bets that they made often four, five, six years ago when they were designing the silicon for the first time. Um, and it's not clear who is going to have the best uh structural advantage because we don't know how the architectures are shaping up. What I will say though is it seems overwhelmingly like we're moving to this world where um memory heavy chips have a fundamental advantage. Um, chips that are able to keep the data on the chip. So, uh, think about all these data flow chips that don't have to go back and forth to HBM, um, all the time. Sombo is a great example of that. Uh, those chips that can keep the memory on the chip and are reducing how much they have to move memory around seem to have a very structural advantage right now. And you're starting to see those chips do accept. >> Yeah, I just wrote a post on LinkedIn because every I get this question all the time. Hey, what's the difference between Nvidia and Google and everyone else and and hot chips again is highlighting this kind of where the engineering focus is. Mathematics is everything in in this world, right? So the math is commoditizing. >> Yeah. >> But the data feeding the math is becoming the key values which why you're seeing these different approaches. And when you start bringing data into the equation, you bring in governance and compliance. >> Yeah. >> Or or routing or other technical features. Talk about that piece because you know there are different approaches. >> You know why should I send that data there when I can send it over there? the gap between capability and control becomes interesting and becomes a technical problem not just like a governance problem. Share your thoughts and vision on this because this seems to be the one of the key areas where there's a technical opportunity to architect >> something that can give you the capabilities and capacity >> with control the data flow. >> Absolutely. I mean, so think about this. Um, let's say you're you're an enterprise, uh, and you have this proprietary data. Um, and this is essentially as the world is moving to infinite intelligence, you could call it. Um, this is essentially your moat. This is your IP. Uh, you could work with an anthropic or an open AI. Um, but every single time you make an API request, you are sending that data to them. Um and so even if you trust those companies there is some risk that hey in the future it might not be the same governance at anthropic but they still have all the data you've sent them. So even if you trust them now you have to think about long term how do I want to protect my company as intelligence becomes infinite. Um and one of the ways to do that is open source. So all of these open source openweight models that are coming out often from China and Nvidia is going to do some stuff on that soon too. Um, a lot of companies want to own the the racks that these openweight models are running on. And so basically the data goes in and it comes out and they own the entire stack. They don't have to worry about someone's going to take my moat, someone's going to train a new model based on uh the proprietary data from my uh from my company. I'm not sure if that answered your question. >> Well, I mean, if you look at like CUDA, right? CUDA's whole thing from Nvidia is programmability. >> Yep. >> So AS6 take huge cycles to get the next one going. So having programmability >> true >> in the stack how do you guys look at that because you have you know like I said coreweave nscale they vertically integrate they provide you know some SLA and some services and then you got the approaches hey we're just going to provide raw intelligence >> and feed that up >> to whoever wants it. >> Yeah. >> Um I want to call it headless. I hate that word in this capacity but like there's a retail side of this business which is developers. >> Yeah. >> There's almost like no AWS. >> Yeah. >> For this world. >> Absolutely. So I mean a few things there. So um a lot of the ecosystem is moving towards open source. So you'll see a lot of these companies come out and they'll say hey we're going to do uh we're we're going to be the software stack for all heterogeneous and and you compile kernels once on our stack and it works on all the different architectures. But imagine for a second um let's say your senova or your cerebrus um you know a GPU is going to do part of the problem very well. prefill they call it GPU does exceptional um and these AS6 do decode exceptionally well. So if you pair them together for one solution you get you know 40% better TCO but you also get you know up to 10 20 times faster AI so it makes sense to do it right. um all of these you know decode silicon we call them internally it's a bit technical but all of these decode silicons um >> they know that it's life or death uh that if they can actually build a software ecosystem internally that integrates well with VLM and SG lang and all of these inference engines that's the whole business they they're all working on that right now they know that so um there's no lack of people you know trying that and then when you talk about model bring up too it's it's a really good question because some of the compilers for these AS6 they use different programming paradigms right and so the compilers can be extremely complex um so one thing that we're doing uh at general compute is we actually hire model bring up experts from all of the different um asich companies you know someone starting tomorrow uh was at Nvidia and and AMD and and Meta we have someone that was a cerebranova starting in a in a week um and so we bring all of those experts internally and we work with the companies to actually build out a better software stack and bring up models internally and then we can offer SLAs's on them too because nobody wants to rent uh a machine for you know $100,000 a month and then it turned out to be a brick because they can't actually run the models that even matter in the first place >> or the market shifted we saw that with training to inference >> great clusters for training they didn't have the pre-filled decode problem they're just training >> exactly >> inference gets interesting and again my takeaway from hot chips so far is that and I've been kind of circling around this want to get your reaction is that the whole disagregated serving string concept is not so much vertical integration. It's just more efficiency. >> Exactly. >> And talk about the reasons because this is very nuanced technical point, but pre-fill and decode do something very good when you separate them >> because of the demand and complexity of the of the prefill. Yeah. >> And the KV cache has to get >> smarter. It gets fatter, gets bigger, bloats up a little bit. Talk about what this means because this is the demand curve is not going away. Yeah. >> So talk about this prefill decode dynamic and it's not so much up the stack. more of really around the resource. >> Yeah. Um so fundamentally when you think of of um of inference, so that's when you actually ask the AI a question and it spits out an answer. Um you can actually split that into two workloads. One of them is called prefill. That one is uh about prompt processing. So let's say I have a I asked a really long question. I have 100,000 tokens in that question. Those tokens or those words need to be turned into numbers to actually run through the math. It's math, right? Those have to be turned into numbers. Um and then decode is when I'm auto reggressively they call it which is uh generating one token at a time. Right. Those are actually two different problems. And so >> and they're talking to the models. >> Right. The decode talks to the models directly in math >> terms. >> Right. Right. And and so it takes what came out of preill and then it uses that um in the auto reggressive fashion to generate tokens one at a time. Um and so if you think about a GPU fundamentally it's a graphics processing unit. And if you think what's what do you need if you're generating graphics? Well, imagine you had one core just for simplicity sake mapped to one pixel on the screen and you're like, I need to know what color this pixel should be uh 300 times a second. Well, actually, it's really easy if I do a bunch of matrix multiplication um for one pixel and then I just have a bunch of different cores and I do it all independently. I do it all at the same time in parallel, right? And so actually prefill for these AI models is something kind of similar where I can process all of the the pre or the context window all of the prompt um each word I can process independently. So it maps really well onto that GPU, right? Because I'm just doing a bunch of matrix multiplications maps one to one. Perfect. >> With decode um I take that prefill they call it KV cache, right? That's just the context but in numbers. Well, now I have to store that right next to the actual silicon as well as the the weights of the model. And the weights of the model are are huge. Now we're getting up to three trillion, 5 trillion parameter models. And then also, you know, a million context uh length windows. And so if I want to process, you know, 10 people, 20, 100 people at the same time, I've got a 100 KV caches and I also have the model weights. And so it just blows up the memory problem. And so, look, I can do all of my uh prefill independently on a GPU maps perfectly. The decode though is is where the GPU really breaks down. >> And talk about the consequences. Again, this we're getting in the weeds a little bit here, but I think it's important people to understand that if you screw up the KV cache, yeah, >> you got to reboot everything. So as you're getting through the multi-step processing, whether you're doing it through some sort of uh system array or whatever matrix multiplication, whatever process you're using, which is very layered, very complex, >> it screws up if you have to reset. >> Yeah. >> And you start that and that's a GPU monolithic problem. Yeah. >> So the answer is okay, put some compute here. And I think >> people misunderstood cerebrus when they first started when I first met the team years ago. They're like, oh, it'll never work. inference. It's just a purpose-built inference. And you know, it turns out they took on the memory wall. Good bet. >> Yeah, >> that was a good bet for Cerebas, but now guess what? You can integrate that in. >> Yeah, >> it's not a standalone. So like you're starting to see the architecture approach is different. What's your >> technical view on this? Because it's kind of like a systems architecture game. Yeah. >> Not a I got Nvidia for this or Cerebras for that or Sanova for this. It's more than an architecture game because even if you find the best chip in the world, can that actually scale to production capacity because you're you know uh very well that there's constraints at TSMC, there's constraints with Micron or SKH Highix or whoever your your memory provider is. So um let's say you know $250 billion of Nvidia Silicon gets deployed uh next year. It could be double, it could be triple. Um, and then the AS6 market might only be able to deploy like 1 billion, right? And so it's a small market, but it's growing very quickly. And it'll be more than 1 billion, but call it even 10. It's it's a fraction of what Nvidia is going to do, but it's growing very quickly. Um, but you had to create those alliances to Broadcom or Intel, which is what Samba is doing, uh, early so that you could actually build to production capacity, right? So it's that it's also the fact that the memory problem is deeper than you would think because let's say you've got Cerebras. Cerebras has 44 gigabytes of SRAM. So they have that one big wafer and they put all of the memory right into it. Well, models are bigger than 44 gigabytes. So now we have to think about how are we going to wire these things together to actually split the model across a bunch of wafers, right? >> And kicks ass inference, too. So it's a great use case. So you plug wire it up. It's it's a great use case, but um I guess I guess what I'm trying to get to is that the bigger the model, it might actually change what chip you want to use, too, right? So, there's going to be different chips that are actually going to excel at different model types and different architectures. And we're just seeing an explosion of AI models. So, it's not clear which one is going to win. >> It's getting bigger. It's just the beginning. All let me ask you a question on the question I get a lot, which is, hey, what's going on with all these new AS6 coming out? You see positron, dmatrix. So, you some classic accelerator markers. Hey, accelerate. They're not just accelerators. are playing another role. What's the view of the market as these new entrance come in uh with this demand curve? How do you think that's going to play out? New formation. How does because you guys are doing this? This is what you're doing. >> Yeah, we're doing it. >> How how is that going to play out? What's your vision of this? >> Well, I I think people are really curious about AS6. Um ever since the OpenAI Cerebrris deal and the IPO of Cerebras, people are starting to accept it. So, when we started the company, this was preall of that. Um and so we told people we said Cerebrus is going to be a big deal and people would you know >> poo poo it they would poo poo that I've heard people it'll never work it's too big it's never done before >> yeah they industry leaders and experts would would do that you know we would kind of get like snubbed to some degree but now everybody is on board they understand it and so we're even seeing the demand from enterprise of like hey um you know I know you've showed us this chip and you taught us how this chip works but what about that chip you know what about etched what about cerebras what about this and that and so we just bring um we kind connect the supply to the demand and we deploy it and so you can take on less risk like if you're an enterprise and you really want to try a cerebrus or or an etched or any of these companies um you could either you know pay millions of dollars and hope and hopefully be able to actually deploy it and manage it correctly they don't want to take they don't want >> they don't want to take the risk it's a huge risk >> so are you targeting them as customers >> we're targeting everyone that [laughter] wants fast >> what are you guys let's talk about your momentum take a minute to explain the momentum you have course there's a macro trend that's your friend so That's going to be good for you guys. I think there's going to be a whole another class of buildout components. I think you're a highlight of that. What happens next? The enterprise, they don't have billions on capex now. They'll use services. Yep. >> They'll do onrem. They'll put in maybe a smaller cluster with FPGAAS or some other lowcost high performance configuration and connect to a service. >> Yep. >> That can give them single tenant like capability. >> Of course, >> that makes total sense to me. What are you guys targeting? >> Well, there's there's three kind of customer profiles, right? One of them would be you've got the Frontier Labs. Um the Frontier Labs just need a ridiculous amount of capacity. They're going to deploy some of their own. They're going to rent some from other people. You know, Fluid Stack does a lot of stuff with Google and Anthropic. Um there's there's a big market there. Uh but they they care about training and they care about inference and um they're experts at uh model compilation and everything. They have their own people on staff and they're just >> they have infinite uh pockets to just deal with these problems. Then you have um AI application companies and enterprise so call it like uh uh the the cursors the perplexities of the world the open codes of the world they have a lot of demand token demand um they're actually more comfortable most of the time paying for an SLA so they're saying you know uh I want to have x number of tokens per second I want to be able to process x number of requests per second um things like that and then the third category is the asset light inference cloud so for example u you've got like the base 10's the fireworks the together AIS. Some of these are guys are starting to move down the stack, but they have essentially infinite demand, right? So any capacity that becomes available to them, they're going to be able to connect that to a buyer. Um, and so they they own a lot of the enduser customer relationships already. And their end users are growing so fast, they just need more demand. And they're also because they're asset light, they're not locked into any >> asset light mean they're not spending a lot of capex to do what endscale and poor weave did. >> Yeah. Sorry, I should have I should have explained that. When we say asset light, we mean that they're not um deploying buying hardware and deploying it themselves, but they're renting capacity from other people, right? Uh and then they're on selling that and they're they're making their own SLAs's and they have their own value added services on top of that. Uh so they have >> is kicking ass. Everyone's those numbers are killing it. >> They're doing great. >> I think that's a big market. has two approaches that I call it the vertically integrated and then like okay feed that asset light demand which is I got customers >> I will qualify what you have and then integrate in >> yeah I mean like like just a thought experiment let's say you raised uh $und00 million of equity right now um you could either go and buy uh you know x number of machines and deploy them or you could rent like five times x uh number of machines and then you could make 5x the revenue so uh it makes a lot of sense sense for for everyone to kind of pick what they're experts at and then um do that and we have a lot of these you know asset like guys that are doing fantastic >> Jason I really like what you guys are doing in March at GTC um saw all the parallel curves Jensen did his thing I then wrote a post that next month April because Jensen's like we're bounded by energy the five layer cake he puts out there which is totally legit um I wrote a post that no it's bounded by energy and money that was the first post that kind of went out and was a company we featured Argentum, which was trying to figure out the financial code, because as you pointed out, this been documented on Bloomberg and other places, the the circular financing, which I think some people try to throw shade on Nvidia, but they're just doing a great job to help build the infrastructure. >> Talk about the financing aspect though because you're taking an approach to bet on >> the asset light market that's in demand. >> Yep. >> And that financing, well, you got to get facilities, energy, and then you got to get the finance. talk about this bounding function of finance because then fast forward to last this month >> Jensen was in New York with the CEO of Goldman Kr$500 billion they're taking care of the physical plant my word but like you know the physical buildout but there's still now a financial market developing >> you're in the middle of this >> what's your vision on how that plays out because risk is management is now in play on both sides >> yeah it's 100% true there there's a very well understood um debt market for GP GPU. So if you want to go buy a bunch of GPUs, you can raise debt. Um, and what Nvidia will do is they'll underwrite the purchase of it. So they'll say, "Hey, if you cannot rent these machines or sell tokens on these machines, we'll rent them back for you." Um, and what that allows you to do is go to the banks and say, "Hey, look, this is basically a guaranteed deal." >> AAA bond right there. It's like And it's sometimes pledged. >> Yeah. >> So that's like every bank's like, "I'm in." >> Exactly. and and we're talking about with the company uh that's worth you know over $4 trillion like they're not going to they're not going to uh default they're going to come through and there's also infinite demand so you will get a customer >> um one and taking it a step further uh you talk about you know an ASIC it's like well nobody really understands what the depreciation life cycle on that ASIC is nobody really understands the residual value of that ASIC after the end of its life there's no secondary market for it because you know people haven't really uh adopted it or diffused it into the economy yet Um, so it's fundamentally a much harder thing to get people uh to to bet on. Um, so uh you're you're probably going to get a quite low interest rate if you're if you're deploying GPUs. If you're deploying an ASIC, it's going to be much higher and you're going to have a much smaller pool of capital to do it from. So that is fundamentally what our business solves. A lot of these ASIC companies, you're I'm sure you chat with them here all the time, just absolutely exceptional people, like incredible >> and great tech and there's demand for what they have. there's demand for what they have and they've spent so much time on the technology. Um, but the deployment piece uh is actually how Nvidia is running circles around them. Nvidia has uh great tech too, but it's not as good as a lot of these other ASIC companies. Um, but they're adopted way more and CUDA is not a good enough excuse anymore. Like the the ecosystem is opening up. Um, you can write your own compiler and AI can >> I mean CUDA is just a software model that makes things makes the AS6 last longer >> until the next rev. So you can level up if the market changes, whatever nuance, >> yeah, >> is key that could be replicated bent on the platform. >> It can be replicated for sure, especially with AI coding capabilities, you should be able to get uh something at least workable. Um but the reason that they're not being diffused more into the economy is uh the fact that there's just no no one deploying them. And that's why we're stepping >> Well, Jason, I'm really jealous of you. You're a young gun. I'm aging out over the years. You're going to be a long road here. But you brought up the depreciation things because because Jensen said something Dave Volant and I were and and Brian were talking about is the the analysts haven't modeled and he put it in kind of quotes. >> They haven't spreadsheeted out what this is going to be. So a lot of people don't know what depreciation means because they don't know what the reuse is. So if we assume scarcity >> Yeah. >> architecturally smart engineers are using older chips. Talk about that from a tech perspective because the old idea was oh that's a chip the next one comes out the value drops you can depreciate that makes total sense in the old way but in the new world where you have diversity of clusters you have diversity of capabilities >> there's a reuse market that keeps the prices up yeah >> which changes the modeling >> on the financial spreadsheet >> of valuation so what's your view on this kind of like a random question but it's one that everyone's asking like well could you hedge that well this is future futures market but but then again if it's depreciating. So there's a whole conversation around the thesis of will the hardware and software be worth less more less in the future or will it have staying power >> and durability >> if you assume okay big clusters small clusters edge physical AI mean like a chip today could be put into a robot maybe >> there's all kinds of like supply chain functionality discussion >> I I think uh I think you're thinking about it the right way um I'm sure you saw the the deal with Coreweave where they signed um A100's through I believe 2029 and that is a very old chip at this point. Um and I think I think what's happening is you're in a supply con constrained market some workloads are are more valuable than others and let's say I could run um on an A100 and I'm making these numbers up but at 10 tokens a second or I could run on a you know uh Nvidia Cerebrris combination uh with you know 2,000 tokens per second. Um, I'm obviously going to put my high-v value uh workloads on the cerebrus rack, but there's probably a bunch of stuff I could just put on the A100 overnight uh and not think about probably internal things, right? Um, so there I I think it depends >> the TCO calculation at that point. >> It is >> like what am I running? It's policy based. It's resource based. >> Exactly. >> Intelligence could manage that. I mean, you put some AI in there. >> Yeah. Well, I'm I'm also always thinking about too like, okay, what what is the revenue per megawatt you can get? So let's say you've got like an A100 um and you have a megawatt of it deployed. Uh theoretically you can make X amount on it and then if you could upgrade to um you know a new Blackwell generation or the Vera Rubin generation you'd have to rework and put capex into the facility in order to actually be able to run those machines but you'd get like I don't know X 10 or X 100 in revenue. Um, so I think the calculation there is is really interesting. But the fact is, look, we're all so comp constrained that everything that's in production right now, we're just going to use it. Um, and as we, you know, upgrade and build new data centers and upgrade old data centers, we will plug in new stuff and things will depreciate and um, you know, I don't think people will be using A100. >> Just not enough sample size, Jason, on this. So, it's I think it's it's a it's an open, you know, question. I think that's going to be one we're going to watch certainly in the middle of it. All right, final question. What are you optimizing for now? Give us a taste of what's coming. I know you got some deals brewing you can't talk about right now. Um >> what's going on? Where's set us the direction where where where's the company heading? >> Um the company is headed towards uh you know being the heterogeneous um ASIC deployment arm. So everything that is is not already being handled by your your core weaves and your nebuses. There's a lot of fantastic chips out there. Everyone wants to try them. Uh they have different use cases. We are going to be deploying those for customers and we have we'll be doing some announcements in the next few weeks I believe. Um and I'm really excited to talk about that. Maybe I can come back and we can >> Yeah, we'll definitely do it. General compute um not doing general purpose computing as we know it. General compute is providing the scale for what we see as a democratization on the AS6 side. As more entrance come in, more capabilities again, more infrastructure demand continues to thunder away. I'm John Furry, your host of the Cube. Thanks for watching.