Jason Goodison, General Compute | theCUBE + NYSE Wired: AI Factories
Watch on YouTubeVideo summary
Jason Goodison, CTO and co-founder of General Compute, joins the discussion to address the surging global demand for AI infrastructure and the critical need for alternatives to Nvidia's dominant GPU architecture. While major cloud providers like CoreWeave and Nebius have vertically integrated with specific chip suppliers such as Nvidia or AMD, leaving a vast array of innovative chips from companies like Cerebras, SambaNova, and others underutilized, General Compute positions itself as the essential deployment arm for these heterogeneous solutions. The company recently secured a $400 million debt facility to productionize these alternative chips, aiming to bridge the gap between exceptional engineering teams building specialized silicon and enterprises that require flexible capacity without being locked into a single vendor's ecosystem. This approach allows businesses to access cutting-edge technology while mitigating the risks associated with relying on a monolithic hardware stack.
A central theme of the conversation is the technical shift from general-purpose computing to memory-heavy, specialized architectures designed for efficient inference. Goodison explains that as AI models grow larger and context windows expand, the traditional GPU architecture struggles with the "memory wall," particularly during the decode phase where data must be kept close to the silicon. Specialized chips like those from Cerebras utilize massive on-chip SRAM to keep data local, drastically reducing latency and memory bandwidth requirements compared to GPUs that constantly shuttle data back and forth. This separation of pre-fill and decode workloads allows for a more efficient use of resources, enabling faster token generation and better handling of large-scale models, which is crucial as the industry moves beyond simple training into complex, real-time inference applications.
Beyond technical architecture, the discussion highlights the evolving financial and strategic landscape of AI infrastructure, where energy constraints and capital efficiency are becoming paramount. Goodison notes that while Nvidia benefits from favorable financing terms due to its established ecosystem and guaranteed demand, emerging ASIC companies face higher interest rates and a lack of secondary markets for their hardware. To solve this, General Compute targets asset-light inference clouds and enterprises that need single-tenant capabilities without massive capital expenditure, offering them the ability to rent specialized capacity on demand. The company's vision involves creating a software stack that can compile models once and run them across diverse hardware architectures, effectively democratizing access to advanced AI compute and allowing companies to optimize their workloads based on specific needs rather than vendor lock-in.
Ultimately, General Compute aims to redefine the AI infrastructure market by connecting supply with demand for a wide variety of specialized chips, ensuring that innovation is not stifled by deployment bottlenecks. As the industry faces supply constraints and energy limitations, the ability to mix and match different hardware types based on TCO calculations and workload requirements will become essential. Goodison emphasizes that the future lies in a heterogeneous ecosystem where enterprises can leverage the best-in-class chips for specific tasks, whether it is high-value training or rapid inference, without taking on the prohibitive risks of deploying unproven technology alone. By acting as an aggregator and deployer, General Compute seeks to accelerate the adoption of these diverse technologies, fostering a more robust and competitive AI infrastructure landscape that can scale alongside the infinite intelligence era.
Read the full video transcript
Palo Alto studio connection Silicon
Valley and Wall Street. I'm John F co
here with Dave Volante my co-host.
Hello, I'm John Furry, host of the cube.
Here in our PaloAlto studio, of course,
we have the cub's NYC studio connecting
Silicon Valley to Wall Street, Wall
Street to Silicon Valley, the NYC wired
programs where technology and Wall
Street intersect. Of course, the cub's
got the deep coverage, tracking all the
semis, all the AI infrastructure
buildouts. This is our AI factory series
where we talk to the leaders who are
building it and setting the table for
the era of AI. Jason Goodison is here.
He's the CTO and co-founder of General
Compute. Thanks for coming in, popping
into the studio today. Appreciate it.
>> Thank you for having me.
>> You're doing some pretty cool work.
Again the demand for AI infrastructure
is off the charts and again the demand
for intelligence is feeding in obviously
the coding which we're everyone's seeing
the value there physical AI edge right
in line so you can see the progression
the trajectories forming there's just
way too much demand and then but the
role of the A6 and the silicon and
software we're seeing the different
approaches from Nvidia Google Cerebras
AMD um all have kind of the bets Yeah,
>> CUDA obviously with Nvidia that's
software evolution programmable
TPUs for Google they got uh you know
different approach with the pods you got
ser the big wafer the general purpose to
custom and you know specialized compute
have always been that spectrum the
harder you go to specialize the harder
is to program the use cases are more
narrow but with AI that's all being
bundled together
>> um you're building out your venture and
in this demand curve
A lot's going on. I want to unpack with
you, but let's get with what you guys
are doing right now. What's the state of
your company? What's the thesis?
>> Yeah.
>> Where are you guys seeing the action?
>> Yeah. So, um, fundamentally what we do
is we like to see ourselves as the
deployment arm for these alternative,
uh, chips, alternatives to GPUs. Um, so
if you think about all of these up and
cominging chips, you've got Cerebrris,
Samanova, I mean those companies have
been around a long time. You've got
TensorDine, Tensor, Posetron, Dmatrix,
there's just the list goes on,
>> boatload of new stuff coming too
>> and all of these engineers are just
fantastic, right? Um they've built these
incredible chips that work for different
use cases. Uh but we have a lot of the
asset heavy clouds of the world, the
Nebuses, the Core Weaves, um that are
kind of locked into their chip supplier.
So you think about Coreweave, they they
have a lot of circular financing with
Nvidia and so they only deploy Nvidia.
Um, TensorWave only really deploys AMD.
They might do other stuff in the future,
I don't know. Um, and then Fluid Stack
does a lot with Google TPU. So, there's
all of these other chips in the world
that are really fantastic and the
engineering teams there are exceptional.
Um, but there's no one that just takes
on that debt and deploys them uh and
then rents them back bare metal to an
enterprise or to another cloud that
needs the capacity. So, that's what we
do. So, a few weeks ago, we just
announced a $400 million debt facility
um in combination with Upper 90. Upper
90 uh joined the cap table and we're
really excited to work with them and uh
productionize a lot of these awesome
chips.
>> Yeah, there's a race. I mean you can see
the vertical integration you know core
weave ncale they just bought any scale.
So you can start to see the disagregated
serving kicking in here at hot chips
Stanford I was poking around yesterday
top conversation is hey the capaca
capability and capacity demand is high
the KV caches are getting stuffed. So
you're starting to see the patterns.
More data is coming in, more demand for
the tokens aka intelligence. So it's
going to put pressure on the
architecture. So okay, the big guys,
they got their partners, but there's a
whole another on on boarding of these
new NeoClouds, Neolabs, or someone's got
some Bitcoin, they got some data center
facilities. Those are different
businesses, but they got the energy,
they got the footprint, they want to
bring that into this new buildout. So
it's almost as if there's a new breed
>> Yeah. of infrastructure opportunity.
Sounds like that's what you're
targeting.
>> Yeah, absolutely. So, when you think
about any major technological
innovation, um you start off by u not
really understanding the problem and so
you just use what you have at your
disposal uh to solve the problem. So,
even when you think about Bitcoin back
in the day, you know, people just mine
Bitcoin on GPUs. Um now nowadays people
mine Bitcoin on A6 because as the uh as
the workload stabilizes you just get so
many um advantages to baking that into
the silicon right um and so what we're
seeing we're a little bit in the wild
west period right now in AI because
there's a new model drop uh every month
um an open source model comes out
there's a new attention mechanism um
there's there's different infrastructure
and architectures that are coming in the
AI models so we have all these chips and
they're all racing. They all have their
own bets that they made often four,
five, six years ago when they were
designing the silicon for the first
time. Um, and it's not clear who is
going to have the best uh structural
advantage because we don't know how the
architectures are shaping up. What I
will say though is it seems
overwhelmingly like we're moving to this
world where um memory heavy chips have a
fundamental advantage. Um, chips that
are able to keep the data on the chip.
So, uh, think about all these data flow
chips that don't have to go back and
forth to HBM, um, all the time. Sombo is
a great example of that. Uh, those chips
that can keep the memory on the chip and
are reducing how much they have to move
memory around seem to have a very
structural advantage right now. And
you're starting to see those chips do
accept.
>> Yeah, I just wrote a post on LinkedIn
because every I get this question all
the time. Hey, what's the difference
between Nvidia and Google and everyone
else and and hot chips again is
highlighting this kind of where the
engineering focus is. Mathematics is
everything in in this world, right? So
the math is commoditizing.
>> Yeah.
>> But the data feeding the math is
becoming the key values which why you're
seeing these different approaches. And
when you start bringing data into the
equation, you bring in governance and
compliance.
>> Yeah.
>> Or or routing or other technical
features. Talk about that piece because
you know there are different approaches.
>> You know why should I send that data
there when I can send it over there? the
gap between capability and control
becomes interesting and becomes a
technical problem not just like a
governance problem. Share your thoughts
and vision on this because this seems to
be the one of the key areas where
there's a technical opportunity to
architect
>> something that can give you the
capabilities and capacity
>> with control the data flow.
>> Absolutely. I mean, so think about this.
Um, let's say you're you're an
enterprise, uh, and you have this
proprietary data. Um, and this is
essentially as the world is moving to
infinite intelligence, you could call
it. Um, this is essentially your moat.
This is your IP. Uh, you could work with
an anthropic or an open AI. Um, but
every single time you make an API
request, you are sending that data to
them. Um and so even if you trust those
companies there is some risk that hey in
the future it might not be the same
governance at anthropic but they still
have all the data you've sent them. So
even if you trust them now you have to
think about long term how do I want to
protect my company as intelligence
becomes infinite. Um and one of the ways
to do that is open source. So all of
these open source openweight models that
are coming out often from China and
Nvidia is going to do some stuff on that
soon too. Um, a lot of companies want to
own the the racks that these openweight
models are running on. And so basically
the data goes in and it comes out and
they own the entire stack. They don't
have to worry about someone's going to
take my moat, someone's going to train a
new model based on uh the proprietary
data from my uh from my company. I'm not
sure if that answered your question.
>> Well, I mean, if you look at like CUDA,
right? CUDA's whole thing from Nvidia is
programmability.
>> Yep.
>> So AS6 take huge cycles to get the next
one going. So having programmability
>> true
>> in the stack how do you guys look at
that because you have you know like I
said coreweave nscale they vertically
integrate they provide you know some SLA
and some services and then you got the
approaches hey we're just going to
provide raw intelligence
>> and feed that up
>> to whoever wants it.
>> Yeah.
>> Um I want to call it headless. I hate
that word in this capacity but like
there's a retail side of this business
which is developers.
>> Yeah.
>> There's almost like no AWS.
>> Yeah.
>> For this world.
>> Absolutely. So I mean a few things
there. So um a lot of the ecosystem is
moving towards open source. So you'll
see a lot of these companies come out
and they'll say hey we're going to do uh
we're we're going to be the software
stack for all heterogeneous and and you
compile kernels once on our stack and it
works on all the different
architectures. But imagine for a second
um let's say your senova or your
cerebrus um you know a GPU is going to
do part of the problem very well.
prefill they call it GPU does
exceptional um and these AS6 do decode
exceptionally well. So if you pair them
together for one solution you get you
know 40% better TCO but you also get you
know up to 10 20 times faster AI so it
makes sense to do it right. um all of
these you know decode silicon we call
them internally it's a bit technical but
all of these decode silicons um
>> they know that it's life or death uh
that if they can actually build a
software ecosystem internally that
integrates well with VLM and SG lang and
all of these inference engines that's
the whole business they they're all
working on that right now they know that
so um there's no lack of people you know
trying that and then when you talk about
model bring up too it's it's a really
good question because some of the
compilers for these AS6 they use
different programming paradigms right
and so the compilers can be extremely
complex um so one thing that we're doing
uh at general compute is we actually
hire model bring up experts from all of
the different um asich companies you
know someone starting tomorrow uh was at
Nvidia and and AMD and and Meta we have
someone that was a cerebranova starting
in a in a week um and so we bring all of
those experts internally and we work
with the companies to actually build out
a better software stack and bring up
models internally and then we can offer
SLAs's on them too because nobody wants
to rent uh a machine for you know
$100,000 a month and then it turned out
to be a brick because they can't
actually run the models that even matter
in the first place
>> or the market shifted we saw that with
training to inference
>> great clusters for training they didn't
have the pre-filled decode problem
they're just training
>> exactly
>> inference gets interesting and again my
takeaway from hot chips so far is that
and I've been kind of circling around
this want to get your reaction is that
the whole disagregated serving string
concept is not so much vertical
integration. It's just more efficiency.
>> Exactly.
>> And talk about the reasons because this
is very nuanced technical point, but
pre-fill and decode do something very
good when you separate them
>> because of the demand and complexity of
the of the prefill. Yeah.
>> And the KV cache has to get
>> smarter. It gets fatter, gets bigger,
bloats up a little bit. Talk about what
this means because this is the demand
curve is not going away. Yeah.
>> So talk about this prefill decode
dynamic and it's not so much up the
stack. more of really around the
resource.
>> Yeah. Um so fundamentally when you think
of of um of inference, so that's when
you actually ask the AI a question and
it spits out an answer. Um you can
actually split that into two workloads.
One of them is called prefill. That one
is uh about prompt processing. So let's
say I have a I asked a really long
question. I have 100,000 tokens in that
question. Those tokens or those words
need to be turned into numbers to
actually run through the math. It's
math, right? Those have to be turned
into numbers. Um and then decode is when
I'm auto reggressively they call it
which is uh generating one token at a
time. Right. Those are actually two
different problems. And so
>> and they're talking to the models.
>> Right. The decode talks to the models
directly in math
>> terms.
>> Right. Right. And and so it takes what
came out of preill and then it uses that
um in the auto reggressive fashion to
generate tokens one at a time. Um and so
if you think about a GPU fundamentally
it's a graphics processing unit. And if
you think what's what do you need if
you're generating graphics? Well,
imagine you had one core just for
simplicity sake mapped to one pixel on
the screen and you're like, I need to
know what color this pixel should be uh
300 times a second. Well, actually, it's
really easy if I do a bunch of matrix
multiplication um for one pixel and then
I just have a bunch of different cores
and I do it all independently. I do it
all at the same time in parallel, right?
And so actually prefill for these AI
models is something kind of similar
where I can process all of the the pre
or the context window all of the prompt
um each word I can process
independently. So it maps really well
onto that GPU, right? Because I'm just
doing a bunch of matrix multiplications
maps one to one. Perfect.
>> With decode um I take that prefill they
call it KV cache, right? That's just the
context but in numbers. Well, now I have
to store that right next to the actual
silicon as well as the the weights of
the model. And the weights of the model
are are huge. Now we're getting up to
three trillion, 5 trillion parameter
models. And then also, you know, a
million context uh length windows. And
so if I want to process, you know, 10
people, 20, 100 people at the same time,
I've got a 100 KV caches and I also have
the model weights. And so it just blows
up the memory problem. And so, look, I
can do all of my uh prefill
independently on a GPU maps perfectly.
The decode though is is where the GPU
really breaks down.
>> And talk about the consequences. Again,
this we're getting in the weeds a little
bit here, but I think it's important
people to understand that if you screw
up the KV cache, yeah,
>> you got to reboot everything. So as
you're getting through the multi-step
processing, whether you're doing it
through some sort of uh system array or
whatever matrix multiplication, whatever
process you're using, which is very
layered, very complex,
>> it screws up if you have to reset.
>> Yeah.
>> And you start that and that's a GPU
monolithic problem. Yeah.
>> So the answer is okay, put some compute
here. And I think
>> people misunderstood cerebrus when they
first started when I first met the team
years ago. They're like, oh, it'll never
work. inference. It's just a
purpose-built inference. And you know,
it turns out they took on the memory
wall. Good bet.
>> Yeah,
>> that was a good bet for Cerebas, but now
guess what? You can integrate that in.
>> Yeah,
>> it's not a standalone. So like you're
starting to see the architecture
approach is different. What's your
>> technical view on this? Because it's
kind of like a systems architecture
game. Yeah.
>> Not a I got Nvidia for this or Cerebras
for that or Sanova for this. It's more
than an architecture game because even
if you find the best chip in the world,
can that actually scale to production
capacity because you're you know uh very
well that there's constraints at TSMC,
there's constraints with Micron or SKH
Highix or whoever your your memory
provider is. So um let's say you know
$250 billion of Nvidia Silicon gets
deployed uh next year. It could be
double, it could be triple. Um, and then
the AS6 market might only be able to
deploy like 1 billion, right? And so
it's a small market, but it's growing
very quickly. And it'll be more than 1
billion, but call it even 10. It's it's
a fraction of what Nvidia is going to
do, but it's growing very quickly. Um,
but you had to create those alliances to
Broadcom or Intel, which is what Samba
is doing, uh, early so that you could
actually build to production capacity,
right? So it's that it's also the fact
that the memory problem is deeper than
you would think because let's say you've
got Cerebras. Cerebras has 44 gigabytes
of SRAM. So they have that one big wafer
and they put all of the memory right
into it. Well, models are bigger than 44
gigabytes. So now we have to think about
how are we going to wire these things
together to actually split the model
across a bunch of wafers, right?
>> And kicks ass inference, too. So it's a
great use case. So you plug wire it up.
It's it's a great use case, but um I
guess I guess what I'm trying to get to
is that the bigger the model, it might
actually change what chip you want to
use, too, right? So, there's going to be
different chips that are actually going
to excel at different model types and
different architectures. And we're just
seeing an explosion of AI models. So,
it's not clear which one is going to
win.
>> It's getting bigger. It's just the
beginning. All let me ask you a question
on the question I get a lot, which is,
hey, what's going on with all these new
AS6 coming out? You see positron,
dmatrix. So, you some classic
accelerator markers. Hey, accelerate.
They're not just accelerators. are
playing another role. What's the view of
the market as these new entrance come in
uh with this demand curve? How do you
think that's going to play out? New
formation. How does because you guys are
doing this? This is what you're doing.
>> Yeah, we're doing it.
>> How how is that going to play out?
What's your vision of this?
>> Well, I I think people are really
curious about AS6. Um ever since the
OpenAI Cerebrris deal and the IPO of
Cerebras, people are starting to accept
it. So, when we started the company,
this was preall of that. Um and so we
told people we said Cerebrus is going to
be a big deal and people would you know
>> poo poo it they would poo poo that I've
heard people it'll never work it's too
big it's never done before
>> yeah they industry leaders and experts
would would do that you know we would
kind of get like snubbed to some degree
but now everybody is on board they
understand it and so we're even seeing
the demand from enterprise of like hey
um you know I know you've showed us this
chip and you taught us how this chip
works but what about that chip you know
what about etched what about cerebras
what about this and that and so we just
bring um we kind connect the supply to
the demand and we deploy it and so you
can take on less risk like if you're an
enterprise and you really want to try a
cerebrus or or an etched or any of these
companies um you could either you know
pay millions of dollars and hope and
hopefully be able to actually deploy it
and manage it correctly they don't want
to take they don't want
>> they don't want to take the risk it's a
huge risk
>> so are you targeting them as customers
>> we're targeting everyone that [laughter]
wants fast
>> what are you guys let's talk about your
momentum take a minute to explain the
momentum you have course there's a macro
trend that's your friend so That's going
to be good for you guys. I think there's
going to be a whole another class of
buildout components. I think you're a
highlight of that. What happens next?
The enterprise, they don't have billions
on capex now. They'll use services. Yep.
>> They'll do onrem. They'll put in maybe a
smaller cluster with FPGAAS or some
other lowcost high performance
configuration and connect to a service.
>> Yep.
>> That can give them single tenant like
capability.
>> Of course,
>> that makes total sense to me. What are
you guys targeting?
>> Well, there's there's three kind of
customer profiles, right? One of them
would be you've got the Frontier Labs.
Um the Frontier Labs just need a
ridiculous amount of capacity. They're
going to deploy some of their own.
They're going to rent some from other
people. You know, Fluid Stack does a lot
of stuff with Google and Anthropic. Um
there's there's a big market there. Uh
but they they care about training and
they care about inference and um they're
experts at uh model compilation and
everything. They have their own people
on staff and they're just
>> they have infinite uh pockets to just
deal with these problems. Then you have
um AI application companies and
enterprise so call it like uh uh the the
cursors the perplexities of the world
the open codes of the world they have a
lot of demand token demand um they're
actually more comfortable most of the
time paying for an SLA so they're saying
you know uh I want to have x number of
tokens per second I want to be able to
process x number of requests per second
um things like that and then the third
category is the asset light inference
cloud so for example u you've got like
the base 10's the fireworks the together
AIS. Some of these are guys are starting
to move down the stack, but they have
essentially infinite demand, right? So
any capacity that becomes available to
them, they're going to be able to
connect that to a buyer. Um, and so they
they own a lot of the enduser customer
relationships already. And their end
users are growing so fast, they just
need more demand. And they're also
because they're asset light, they're not
locked into any
>> asset light mean they're not spending a
lot of capex to do what endscale and
poor weave did.
>> Yeah. Sorry, I should have I should have
explained that. When we say asset light,
we mean that they're not um deploying
buying hardware and deploying it
themselves, but they're renting capacity
from other people, right? Uh and then
they're on selling that and they're
they're making their own SLAs's and they
have their own value added services on
top of that. Uh so they have
>> is kicking ass. Everyone's those numbers
are killing it.
>> They're doing great.
>> I think that's a big market. has two
approaches that I call it the vertically
integrated and then like okay feed that
asset light demand which is I got
customers
>> I will qualify what you have and then
integrate in
>> yeah I mean like like just a thought
experiment let's say you raised uh
$und00 million of equity right now um
you could either go and buy uh you know
x number of machines and deploy them or
you could rent like five times x uh
number of machines and then you could
make 5x the revenue so uh it makes a lot
of sense sense for for everyone to kind
of pick what they're experts at and then
um do that and we have a lot of these
you know asset like guys that are doing
fantastic
>> Jason I really like what you guys are
doing in March at GTC um saw all the
parallel curves Jensen did his thing I
then wrote a post that next month April
because Jensen's like we're bounded by
energy the five layer cake he puts out
there which is totally legit um I wrote
a post that no it's bounded by energy
and money that was the first post that
kind of went out and was a company we
featured Argentum, which was trying to
figure out the financial code, because
as you pointed out, this been documented
on Bloomberg and other places, the the
circular financing, which I think some
people try to throw shade on Nvidia, but
they're just doing a great job to help
build the infrastructure.
>> Talk about the financing aspect though
because you're taking an approach to bet
on
>> the asset light market that's in demand.
>> Yep.
>> And that financing, well, you got to get
facilities, energy, and then you got to
get the finance. talk about this
bounding function of finance because
then fast forward to last this month
>> Jensen was in New York with the CEO of
Goldman Kr$500
billion they're taking care of the
physical plant my word but like you know
the physical buildout but there's still
now a financial market developing
>> you're in the middle of this
>> what's your vision on how that plays out
because risk is management is now in
play on both sides
>> yeah it's 100% true there there's a very
well understood um debt market for GP
GPU. So if you want to go buy a bunch of
GPUs, you can raise debt. Um, and what
Nvidia will do is they'll underwrite the
purchase of it. So they'll say, "Hey, if
you cannot rent these machines or sell
tokens on these machines, we'll rent
them back for you." Um, and what that
allows you to do is go to the banks and
say, "Hey, look, this is basically a
guaranteed deal."
>> AAA bond right there. It's like And it's
sometimes pledged.
>> Yeah.
>> So that's like every bank's like, "I'm
in."
>> Exactly. and and we're talking about
with the company uh that's worth you
know over $4 trillion like they're not
going to they're not going to uh default
they're going to come through and
there's also infinite demand so you will
get a customer
>> um one and taking it a step further uh
you talk about you know an ASIC it's
like well nobody really understands what
the depreciation life cycle on that ASIC
is nobody really understands the
residual value of that ASIC after the
end of its life there's no secondary
market for it because you know people
haven't really uh adopted it or diffused
it into the economy yet Um, so it's
fundamentally a much harder thing to get
people uh to to bet on. Um, so uh you're
you're probably going to get a quite low
interest rate if you're if you're
deploying GPUs. If you're deploying an
ASIC, it's going to be much higher and
you're going to have a much smaller pool
of capital to do it from. So that is
fundamentally what our business solves.
A lot of these ASIC companies, you're
I'm sure you chat with them here all the
time, just absolutely exceptional
people, like incredible
>> and great tech and there's demand for
what they have. there's demand for what
they have and they've spent so much time
on the technology. Um, but the
deployment piece uh is actually how
Nvidia is running circles around them.
Nvidia has uh great tech too, but it's
not as good as a lot of these other ASIC
companies. Um, but they're adopted way
more and CUDA is not a good enough
excuse anymore. Like the the ecosystem
is opening up. Um, you can write your
own compiler and AI can
>> I mean CUDA is just a software model
that makes things makes the AS6 last
longer
>> until the next rev. So you can level up
if the market changes, whatever nuance,
>> yeah,
>> is key that could be replicated bent on
the platform.
>> It can be replicated for sure,
especially with AI coding capabilities,
you should be able to get uh something
at least workable. Um but the reason
that they're not being diffused more
into the economy is uh the fact that
there's just no no one deploying them.
And that's why we're stepping
>> Well, Jason, I'm really jealous of you.
You're a young gun. I'm aging out over
the years. You're going to be a long
road here. But you brought up the
depreciation things because because
Jensen said something Dave Volant and I
were and and Brian were talking about is
the the analysts haven't modeled and he
put it in kind of quotes.
>> They haven't spreadsheeted out what this
is going to be. So a lot of people don't
know what depreciation means because
they don't know what the reuse is. So if
we assume scarcity
>> Yeah.
>> architecturally smart engineers are
using older chips. Talk about that from
a tech perspective because the old idea
was oh that's a chip the next one comes
out the value drops you can depreciate
that makes total sense in the old way
but in the new world where you have
diversity of clusters you have diversity
of capabilities
>> there's a reuse market that keeps the
prices up yeah
>> which changes the modeling
>> on the financial spreadsheet
>> of valuation so what's your view on this
kind of like a random question but it's
one that everyone's asking like well
could you hedge that well this is future
futures market but but then again if
it's depreciating. So there's a whole
conversation around the thesis of will
the hardware and software be worth less
more less in the future or will it have
staying power
>> and durability
>> if you assume okay big clusters small
clusters edge physical AI mean like a
chip today could be put into a robot
maybe
>> there's all kinds of like supply chain
functionality discussion
>> I I think uh I think you're thinking
about it the right way um I'm sure you
saw the the deal with Coreweave where
they signed um A100's through I believe
2029 and that is a very old chip at this
point. Um and I think I think what's
happening is you're in a supply con
constrained market some workloads are
are more valuable than others and let's
say I could run um on an A100 and I'm
making these numbers up but at 10 tokens
a second or I could run on a you know uh
Nvidia Cerebrris combination uh with you
know 2,000 tokens per second. Um, I'm
obviously going to put my high-v value
uh workloads on the cerebrus rack, but
there's probably a bunch of stuff I
could just put on the A100 overnight uh
and not think about probably internal
things, right? Um, so there I I think it
depends
>> the TCO calculation at that point.
>> It is
>> like what am I running? It's policy
based. It's resource based.
>> Exactly.
>> Intelligence could manage that. I mean,
you put some AI in there.
>> Yeah. Well, I'm I'm also always thinking
about too like, okay, what what is the
revenue per megawatt you can get? So
let's say you've got like an A100 um and
you have a megawatt of it deployed. Uh
theoretically you can make X amount on
it and then if you could upgrade to um
you know a new Blackwell generation or
the Vera Rubin generation you'd have to
rework and put capex into the facility
in order to actually be able to run
those machines but you'd get like I
don't know X 10 or X 100 in revenue. Um,
so I think the calculation there is is
really interesting. But the fact is,
look, we're all so comp constrained that
everything that's in production right
now, we're just going to use it. Um, and
as we, you know, upgrade and build new
data centers and upgrade old data
centers, we will plug in new stuff and
things will depreciate and um, you know,
I don't think people will be using A100.
>> Just not enough sample size, Jason, on
this. So, it's I think it's it's a it's
an open, you know, question. I think
that's going to be one we're going to
watch certainly in the middle of it. All
right, final question. What are you
optimizing for now? Give us a taste of
what's coming. I know you got some deals
brewing you can't talk about right now.
Um
>> what's going on? Where's set us the
direction where where where's the
company heading?
>> Um the company is headed towards uh you
know being the heterogeneous um ASIC
deployment arm. So everything that is is
not already being handled by your your
core weaves and your nebuses. There's a
lot of fantastic chips out there.
Everyone wants to try them. Uh they have
different use cases. We are going to be
deploying those for customers and we
have we'll be doing some announcements
in the next few weeks I believe. Um and
I'm really excited to talk about that.
Maybe I can come back and we can
>> Yeah, we'll definitely do it. General
compute um not doing general purpose
computing as we know it. General compute
is providing the scale for what we see
as a democratization on the AS6 side. As
more entrance come in, more capabilities
again, more infrastructure demand
continues to thunder away. I'm John
Furry, your host of the Cube. Thanks for
watching.