Video summary
Rodrigo Liang, co-founder and CEO of SambaNova, joins the discussion to highlight a pivotal shift in the artificial intelligence landscape where inference has become the new economic center of gravity. As the industry moves beyond the initial training phase, the primary challenge for data centers is now how to deploy infrastructure sustainably while achieving financial payback. Liang emphasizes that revenue is directly tied to energy consumption and throughput, leading to a new metric where operators measure success in terms of revenue generated per megawatt rather than just raw compute power. To maximize this efficiency, SambaNova focuses on driving total output at the lowest possible cost and power, ensuring that investments in infrastructure yield consistent returns without relying indefinitely on borrowed capital.
A significant portion of Liang's strategy involves optimizing existing hardware through a concept called disaggregated inference, which allows companies to mix and match different types of chips to handle specific tasks within an AI workflow. By partnering with NVIDIA, SambaNova can place its specialized gear next to existing NVIDIA racks to handle the "decode" phase of processing, effectively tripling the total throughput of older hardware like H100s. This approach not only extends the useful life of current infrastructure but also addresses supply chain constraints by enabling heterogeneous computing environments where different chips perform different functions. Furthermore, this technology facilitates the creation of distributed data centers in existing air-cooled facilities, allowing for rapid deployment without the years-long lead time required to build new liquid-cooled gigawatt-scale sites.
Beyond economic optimization and infrastructure innovation, Liang addresses the critical issue of safety surrounding physical AI and robotics. He acknowledges the public concern regarding AI risks but argues that the industry must shift its attention from merely slowing down development to actively investing in guardrails and safety mechanisms. Drawing parallels to the early internet era, he suggests that while chaos is inevitable with transformative technology, the goal is to "reign in the chaos" through focused R&D on safety, much like Intel did under Gordon Moore. He believes that addressing these concerns will not only mitigate risks but also lead to significantly better models and a more stable future for AI integration into everyday devices and robotics.
Looking ahead, SambaNova continues to double down on inference as the name of the game, with upcoming hardware launches like the SM50 poised to further enhance profitability for service providers. The company reports strong business performance with quarter-over-quarter revenue doubling, reflecting immense market demand for efficient AI solutions. As the ecosystem expands to include edge computing, wearables, and connected devices, Liang envisions a future where AI is ubiquitous yet economically viable. The overarching theme remains clear: sustainable growth in the AI sector depends on creating infrastructure that delivers long-term economic value, lowers costs, and ensures that the technology can be deployed safely and effectively across the globe.
Read the full video transcript
Palo Alto studio connection Silicon
Valley and Wall Street. So I'm John F co
here with Dave Volante my co-host.
[music] Hello, I'm John Furry with the
cube. We are here at the cub's NYC
studio. Of course we have our Palo Alto
studio connecting Silicon Valley to Wall
Street. This is part of the cubes NYC
wired program and open community. We
have back on the cube cube alumni back
for another appearance. He's like a
regular contributor of Regal Leang,
co-founder and CEO of Sanova, part of
our AI factory series. One of our most
popular series we started two years ago
and really has been the pre precursor to
the AI infrastructure boom. Very good to
see you. Thanks for coming back on the
cube. I think Gemma talked to you last
two times when I was in California.
Thanks for coming back on.
>> Yeah, thanks for having me. What an
important time for us to be in this uh
in the AI industry. We've seen each
other now for a couple years as part of
this new NYC wired community at events
and on the cube. Some significant
changes in the past year. I mean
inference obviously a lot of insiders
saw that early. The mainstream saw
agents now kicking in. Coding was great
brought that in. Agent agentic brings up
a whole another paradigm shift around
the role of the resource to service
agents what inference means there. So
you start to see a whole shift. What's
been the biggest change for you guys
besides the billion dollars in funding
you guys just closed uh which we covered
on Silicon Angle u that's validation but
what's been the biggest change this year
>> well look inference is now the economic
center of AI and people are trying to
figure out how to make that investment
sustainable and so as we went from
training to inference now people are
thinking about how to deploy deploy with
um uh sustainability deploy with uh good
energy consumption and all of those
things but most importantly how to
deploy in a way they can they can get uh
financial payback right we can't
continue to borrow money forever without
returns and so being able to drive good
economic payback for their investors is
going to be an important part of data
centers and actually as they move into
inference
>> yeah and people know that tokens equals
revenue that's well understood now we're
starting to hear conversations about
modeling out revenue growth um I think
we're starting to see benchmarks now
saying this gigawatt equals this in
revenue or this megawws equals this in
revenue is starting to people starting
to quantify and even a couple years ago
I think Jensen Wong Nvidia said no one's
spreadsheeted out was his word but
that's the financial modeling what have
you seen there as CEO and running your
business you're involved in a lot of
these economic conversations not just
from a deployment standpoint but the
payback a lot of these financial
conversations is energy and money are
the two factors
>> yeah exactly exactly well well the the
energy u the energy part the data center
cost and the capital acquisition of
infrastructure Those are on the cost
side. On the revenue side, you can also
dial that up and dial that up. And so
you can uh think about why people are
driving speed like sodova. We can
generate speed because uh speed gets you
more throughput. But more important than
that is also speed at concurrence. You
know, can you get many many users
getting that same speed at the same time
because you want ultimate total output
total throughput per rack divided by
that cost structure.
>> You guys been at this for almost a
decade now. Uh, we were talking before
we came on camera about Hot Chips, which
is a Stanford event that's well known as
total NerdFest. It's Nerd Nation at
Stanford as everyone knows, but this is
like the the alpha, the state-of-the-art
engineers really working on the next
generation of accelerated technology.
>> The conversation that I got away from
that was there's a lot of work going on
around processor called just processor
generically, whatever you want to call,
and memory obviously HPM and now you got
solid state. But this kind of reminds me
of the '9s when you had to use memory
management utilities to swap in and out
but at such bigger scale. What has been
the big thing for you guys now as you
look at your road map? How are you guys
optimizing? Because everyone wants to
squeeze as much tokens out of that
energy out of that processing the
relationship with the data to memory.
These are all now part of the hardcore
engineering
uh vernacular.
>> Yeah, exactly. Well, ultimately you have
these chips and these chips are
constrained by what the technology can
give you both in terms of compute and
the memory. And so for SANOVA, we're
really focused on driving the total
output at the lowest cost and lowest
power. And so if you can do that and
before the the session, we started
talking about how people are measuring
their compute in terms of megawws and
and their revenue in terms of megawatt.
Well, if you take that megawatt and you
can actually put more racks in and each
rack produces more tokens, that's how
you generate more revenue per megawatt,
right? And so if you if you take that
chip, the same chip, and you increase
the total output per chip, you're going
to do a lot better when it comes to
payback.
>> What are some of the things you guys got
going on that you could share on
momentum side? Because that really is
where everyone's looking at. You talked
about some of the economics of buildout,
but when you're in operations, you're
building, operating, and investing.
Everyone's doing those three things at
the same time. Yeah. On the op side,
what what do you see coming out? What do
you guys have now? What's the momentum?
>> No, look, I there are three incredibly
exciting use cases that we see with some
of technology. One, people who have
already deployed a lot of NVIDIA gear.
We're partnering with Nvidia on this
with something called disagregated
inference and what you're able to do is
take NVIDIA gear make that partner with
a SANOVA gear and in that construct
disagregate the front front part called
prefill and the decco part using sumova
and generate 3x of total throughput. So
existing hardware of Nvidia can jump 3x
in total throughput just by putting
someone rack next to it. And that's a
really really exciting development for
people who already have their gear and
trying to get better economics. The
other one that I think we're starting to
see a lot of really exciting use use
case for is this distributed data center
existing data centers air cooled and
being able to actually use brownfield
data centers that don't require new
investment new you know breaking ground
new liquid cooling and just put some of
aircooled technology into it and then
get significant advantage just by
actually delivering faster at a lower
cost lower. So two two major things that
you said there one is if I bought gear
call it gear that's a term we use a lot
bought a lot of systems
normal normal depreciation would have
that losing value now you got value
creation so the price the value of that
gear is actually more
>> critical so that's good leverage
>> that's a good use case great
>> everyone signs up for that all day long
I'm sure exactly and the other one is
this new paradigm around disagregated
serving
>> which is also a precursor into
disagregate infrastructure. Yeah.
>> Because what you just said is
essentially adding new new nodes out
there,
>> data center nodes to be AI factories
basically without all the requirements
>> for the megawatt gigawatt data centers
that takes years to build.
>> Yeah, that's right. I mean, one of the
biggest challenge we're seeing in the
world today is how quickly can we stand
up gigawatt data centers and and that's
going to be more and more challenging,
you know, as as we think about the needs
that we have and how much it impacts the
energy grid. And so if we can actually
reuse existing infrastructure in around
the world, existing infrastructure in
the United States where you already have
energy allocated, ex space is allocated,
it's air cooled, and you can deployed
infrastructure to have state-of-the-art
AI running on it, I think that's going
to be an incredibly valuable way to
actually get AI available to everybody
at a much much lower cost. So, does the
word edge of the network go away when
you have distributed computing paradigm
where you have large data center nodes
like an AI factory in mega Texas or
whatever it is and if you just got a
cell tower that's got a building in
power and network connectivity that's
air cool in there that's an edge that's
still a node on the network but it's
technically an edge so smaller factory
configuration
>> yeah we actually call yeah we we we
think of it as these massive kind of
buildouts uh where you have huge gig go
about data centers that's going to
continue because people have to train
models and the the necessity for large
clusters continue to exist. Now you go
into what I'm calling distributed data
centers. So take the 50 top metropolitan
cities and Vista for example is a
partners that building out these
distributed data centers in existing
brownfield data centers re re-energizing
existing infrastructure with existing
power existing cooling and then running
that using someone gear. Now edge edge
is the next thing I'm super excited
super excited because if you think about
true edge
>> which is where our mobile towers are
where our users are robotics all these
different things there is still another
wave coming where very very localized
computing is going to be serving very
localized use cases of AI and that's to
me the true edge of comput
>> and they need high performance
>> incredible high performance ultra low
latency because now you're starting to
talk about robotics and you talk about
kind of the things that people want to
use every day and frankly as a
population, our patience is not very
high. And so, so
>> I want to get your thoughts. I want to
get you brought up robotics. I want to
bring up safety because the number one
conversation in all of our physical AI
robotics series that we're running,
safety is the number one conversation.
Security is in there, but I'd say number
one Basically 1 A would be safety
because robotics you got to have a safe
car, you got to have a safe robot. Yeah.
>> Um that's really obvious. you go, okay,
all the safety on AI conversation seems
to be a tempest in the teapot because
it's like people are working on safety.
Yeah. What's your what's your thoughts
on all this negativity around safety? Um
um you got half the world be like, okay,
I don't really understand. What about I
hate it. I'm I'm afraid they don't
understand. Then you have the people who
understand going, "Wow, this is one of
the best revolutions of all time." Yeah.
In the computer industry. So you got
it's 50/50.
>> Yeah. Yeah. No, it's understandable.
Look, I think you look at a technology
as transformative as this and you saw it
over the weekend with the frontier
models and now you know you know and
more broadly around the model side that
yeah safety is important and it's a
great great wakeup call for all the
leaders to start thinking about the fact
that we've invested in the R&D of the
models but how do we invest in the
safety and the and the guardrails around
kind of how we use it and so that's
incredibly important. And I think you're
going to continue to see attention on to
that and we saw saw this in early
internet days where safety and security
kind of became um something that um was
in the forefront people's minds and
entire industries got created from it.
And
>> you know I have a lot of respect for
Daario but I do think that he's not the
poster child for safety when it comes to
AI. I think he's he's no noble effort to
lay out his concerns when the air
company's got a great track record in
terms of the ethics. So you give them
the props for that. But there are many
people in in the industry that actually
have been through this before and have
actually done it. Yeah.
>> Have seen where you had let chaos rain,
re rain in the chaos situations. Yeah.
>> What do you think the the he could learn
from what Daario and others that are now
aware of this could learn
>> from the history of how innovation can
be chaotic and then reigned in. That's
Andy Gro's favorite expression. Let
chaos rain and reign in the chaos. Look
what it happened with Intel under his
regime. under Gordon Moore.
>> Yeah. Well, I mean, look, I it starts
with uh uh emphasis and attention. I
think, you know, we're we're at a time
where the industry is starting to pay a
lot of attention. And look, the the AI
genie is out of the bottle. It's not
we're not going to go back in, right?
But that said, uh it's not too much to
me as much about slowing things down,
but shifting your attention to investing
in the safety piece, right? Because
there are portions of it that requires
the attention. There's a lot to be
learned from the past, but requires the
attention. And I think being able to
actually take that energy and devote to
it, I think it's going to make those
models significantly better.
>> It's funny, I watch some of the
mainstream uh programs on TV and I read
a lot and the perspective on um AI and
it's kind of bothers me like oh that
that that group of people tech people
are controlling AI s an actor on on one
of the shows kind of kind of you know
laying into the tech industry. I'm like
well it's not just them. There's a lot
of other people involved. But the
question that that that I ask is
>> rhetorical if you inject intelligence
into something
>> a network edge like you just pointed
out. Yeah.
>> What happens if you inject intel? So I
think there's a right to be concerned
guard rails and keep watching it. You
don't want to let chaos take over
certainly. But yeah, I think we're going
to have an experimentation. We we have
to identify that. So that's noble.
>> But there's there's a there's a real
question to ask. What does it mean to
inject AI intelligence into a process
into a device? Mhm.
>> You're at the center of it. How would
you look at that? How do you frame that?
>> Well, we we look at it, you know, two
pieces. There's the, you know, kind of
frontier model and and and the race
towards AGI and then we have
infrastructure on the infrastructure
side and that's getting smarter too. and
sum it over really focus on all the
challenges around kind of creating safe,
secure and you know cost-effective uh
infrastructure and so when I look at
that I think about one of the biggest
challenges of infrastructure is making
sure that it's sustainable and making
sure that businesses aren't in a bubble
right it's all about producing
infrastructure that has good long-term
econom economic value and allows you to
actually consistently build and that's
kind of what we're thinking about
>> drive that payback down you know time to
pay back down, drive that ROI up and
making sure that people can actually
invest on this side of the
infrastructure in a sustainable and
long-term way.
>> Yeah, I love that that that
infrastructure side. Let's go there for
a second because you brought this up
earlier. I want to go back to it because
I think it's one of the most nuance
undertalked about topics. Yeah,
>> certainly in the tech circles it's
talked about a lot. Disagregated
serving. Yeah, that's one of your killer
use cases for Sanova where you know
everyone loves Nvidia. Hey, give me the
GPUs. Everyone knows there's scarcity
there and now there's still other now
computes booming with agents. So there
there's still infrastructure but this
aggregated serving is really solves one
of the major constraints. Yeah. Which is
the volume of data and the lack of
network coherency around managing it
because
>> prefill is the prompt and then the
decode is what happens after on the math
side. So you got math and math talking
[snorts] to math.
>> Yeah.
>> What does that tell us? Because that is
a constraint that we see in other areas.
You mentioned disagregated
infrastructure. Everyone's going to have
nodes.
>> That's disagregated. Yeah,
>> there's serving there too, maybe on the
edge. But what is the disagregated
serving point to? It's not just bolting
on accelerators. It really is a paradigm
architectural shift.
>> Well, it's a precursor to kind of what I
think is going to be broad long-term
view of data centers, which is
heterogeneous computing. And so, you're
going to be be able to mix and match
different technologies to run what you
need for AI. And so, you can actually
have different chips running prefill,
different chips running decode, and
frankly, different chips running
applications that the the agent's going
to call. And so with SNOVA, you know,
we're squarely in the decode side of it.
And so we're able to actually take
>> older infrastructure, say you have an
H100 or now bees will soon be older as
well. [laughter] You know, if you think
about kind of the the the the
infrastructure that people have invested
already, how do we actually extend the
life and you can actually disagregate
it? Put someone over Iraq and suddenly
you've extended the value of that
infrastructure for another two three
years. And so that's an important thing
for people to think about because that
amortization of the infrastructure is
something that a lot of people are
really concerned about
>> and really speaks to your economic
opportunity for Simonova because we were
just talking on our Q pod last last
episode about how in history of the
computer hardware industry people would
always want forward pricing. Now they're
locking in pricing because they think
it's going to go up.
>> They think that the price of the H100's
actually are not going to decay as fast.
>> If anything might even go up.
>> Yeah. because of the innovation
happening around it.
>> Yeah. Yeah. What interesting times where
you know you got a convergence of many
things, right? You convergence of the
the these models and the performance,
this unlimited demand, you know, this
buildout that's incredible. And then
you've got the supply chain constraints
coming in. And so those three things
actually coming together is actually
driving a very very interesting economic
model. But here's here's a we can here's
something we can all agree on. something
we can all agree on where when it comes
to inference, everybody wants to see the
cost in inference come down, right? And
so whether whether it's better cost in
the supply chain, whether that's, you
know, being able to actually lower the
energy or data center costs or whether
that's actually finding more supply of
different types of architectures, we can
all agree that we need to drive the cost
>> and the energy is the key function. All
right, what's up for you guys in the
second half of the year? We got the cube
will be at open compute supercomputing
as reinvent a variety of other events uh
infrastructure events happening in
Silicon Valley uh uh as well big
announcement we're not going to be there
we're here in New York what's on your
focus area for the second half of the
year obviously inference is super hot is
that still doubling down on inference is
that still the name of the game for you
guys yeah what's the focus
>> yeah no it's all about economics is all
about payback for for for uh data
centers being able to actually drive
these services we think that uh uh with
uh the launch of SM50 which is coming
very soon here in terms of first
shipments out I think people are going
to see an incredible opportunity for
them to actually take and drive
profitability into their services after
many many years of actually investing
investing investing they're going to be
able to start matching up their
infrastructure that they have today with
some of technologies and drive a much
much higher profitability into their
businesses great to see you final
question give a taste of for the folks
watching what the business performance
has been for your
uh and what's your outlook?
>> You know, look, I mean, we're enjoying
we've done six quarter of quarter over
quarter doubling and I think we'll
continue to see incredible demand. Uh
the business is growing really, really
fast and really, you know, showing that
uh uh inference is something that people
have a lot of interest around and it's
the right time for people to invest.
>> It's going to get bigger when the edge
and everything gets connected. Thanks
for coming on the cube.
>> Yeah, thanks for having
>> John Furrier. This is the AI factory
series. This is one of our most popular
series. The AI infrastructure continued
to accelerate and expand. This is just
AI factories. These big centers of data
centers, they'll go to traditional data
centers in the enterprise. You'll start
to see the edge develop and wearables.
We all have our ring or our whoop.
They're all going to be connected too.
We'll bring that coverage to you from
the cube. Thanks for watching.