Submind YouTube summaries
Thumbnail for Rodrigo Liang, SambaNova | theCUBE + NYSE Wired: AI Factories

Rodrigo Liang, SambaNova | theCUBE + NYSE Wired: AI Factories

Watch on YouTube

Video summary

Rodrigo Liang, co-founder and CEO of SambaNova, joins the discussion to highlight a pivotal shift in the artificial intelligence landscape where inference has become the new economic center of gravity. As the industry moves beyond the initial training phase, the primary challenge for data centers is now how to deploy infrastructure sustainably while achieving financial payback. Liang emphasizes that revenue is directly tied to energy consumption and throughput, leading to a new metric where operators measure success in terms of revenue generated per megawatt rather than just raw compute power. To maximize this efficiency, SambaNova focuses on driving total output at the lowest possible cost and power, ensuring that investments in infrastructure yield consistent returns without relying indefinitely on borrowed capital. A significant portion of Liang's strategy involves optimizing existing hardware through a concept called disaggregated inference, which allows companies to mix and match different types of chips to handle specific tasks within an AI workflow. By partnering with NVIDIA, SambaNova can place its specialized gear next to existing NVIDIA racks to handle the "decode" phase of processing, effectively tripling the total throughput of older hardware like H100s. This approach not only extends the useful life of current infrastructure but also addresses supply chain constraints by enabling heterogeneous computing environments where different chips perform different functions. Furthermore, this technology facilitates the creation of distributed data centers in existing air-cooled facilities, allowing for rapid deployment without the years-long lead time required to build new liquid-cooled gigawatt-scale sites. Beyond economic optimization and infrastructure innovation, Liang addresses the critical issue of safety surrounding physical AI and robotics. He acknowledges the public concern regarding AI risks but argues that the industry must shift its attention from merely slowing down development to actively investing in guardrails and safety mechanisms. Drawing parallels to the early internet era, he suggests that while chaos is inevitable with transformative technology, the goal is to "reign in the chaos" through focused R&D on safety, much like Intel did under Gordon Moore. He believes that addressing these concerns will not only mitigate risks but also lead to significantly better models and a more stable future for AI integration into everyday devices and robotics. Looking ahead, SambaNova continues to double down on inference as the name of the game, with upcoming hardware launches like the SM50 poised to further enhance profitability for service providers. The company reports strong business performance with quarter-over-quarter revenue doubling, reflecting immense market demand for efficient AI solutions. As the ecosystem expands to include edge computing, wearables, and connected devices, Liang envisions a future where AI is ubiquitous yet economically viable. The overarching theme remains clear: sustainable growth in the AI sector depends on creating infrastructure that delivers long-term economic value, lowers costs, and ensures that the technology can be deployed safely and effectively across the globe.
Read the full video transcript
Palo Alto studio connection Silicon Valley and Wall Street. So I'm John F co here with Dave Volante my co-host. [music] Hello, I'm John Furry with the cube. We are here at the cub's NYC studio. Of course we have our Palo Alto studio connecting Silicon Valley to Wall Street. This is part of the cubes NYC wired program and open community. We have back on the cube cube alumni back for another appearance. He's like a regular contributor of Regal Leang, co-founder and CEO of Sanova, part of our AI factory series. One of our most popular series we started two years ago and really has been the pre precursor to the AI infrastructure boom. Very good to see you. Thanks for coming back on the cube. I think Gemma talked to you last two times when I was in California. Thanks for coming back on. >> Yeah, thanks for having me. What an important time for us to be in this uh in the AI industry. We've seen each other now for a couple years as part of this new NYC wired community at events and on the cube. Some significant changes in the past year. I mean inference obviously a lot of insiders saw that early. The mainstream saw agents now kicking in. Coding was great brought that in. Agent agentic brings up a whole another paradigm shift around the role of the resource to service agents what inference means there. So you start to see a whole shift. What's been the biggest change for you guys besides the billion dollars in funding you guys just closed uh which we covered on Silicon Angle u that's validation but what's been the biggest change this year >> well look inference is now the economic center of AI and people are trying to figure out how to make that investment sustainable and so as we went from training to inference now people are thinking about how to deploy deploy with um uh sustainability deploy with uh good energy consumption and all of those things but most importantly how to deploy in a way they can they can get uh financial payback right we can't continue to borrow money forever without returns and so being able to drive good economic payback for their investors is going to be an important part of data centers and actually as they move into inference >> yeah and people know that tokens equals revenue that's well understood now we're starting to hear conversations about modeling out revenue growth um I think we're starting to see benchmarks now saying this gigawatt equals this in revenue or this megawws equals this in revenue is starting to people starting to quantify and even a couple years ago I think Jensen Wong Nvidia said no one's spreadsheeted out was his word but that's the financial modeling what have you seen there as CEO and running your business you're involved in a lot of these economic conversations not just from a deployment standpoint but the payback a lot of these financial conversations is energy and money are the two factors >> yeah exactly exactly well well the the energy u the energy part the data center cost and the capital acquisition of infrastructure Those are on the cost side. On the revenue side, you can also dial that up and dial that up. And so you can uh think about why people are driving speed like sodova. We can generate speed because uh speed gets you more throughput. But more important than that is also speed at concurrence. You know, can you get many many users getting that same speed at the same time because you want ultimate total output total throughput per rack divided by that cost structure. >> You guys been at this for almost a decade now. Uh, we were talking before we came on camera about Hot Chips, which is a Stanford event that's well known as total NerdFest. It's Nerd Nation at Stanford as everyone knows, but this is like the the alpha, the state-of-the-art engineers really working on the next generation of accelerated technology. >> The conversation that I got away from that was there's a lot of work going on around processor called just processor generically, whatever you want to call, and memory obviously HPM and now you got solid state. But this kind of reminds me of the '9s when you had to use memory management utilities to swap in and out but at such bigger scale. What has been the big thing for you guys now as you look at your road map? How are you guys optimizing? Because everyone wants to squeeze as much tokens out of that energy out of that processing the relationship with the data to memory. These are all now part of the hardcore engineering uh vernacular. >> Yeah, exactly. Well, ultimately you have these chips and these chips are constrained by what the technology can give you both in terms of compute and the memory. And so for SANOVA, we're really focused on driving the total output at the lowest cost and lowest power. And so if you can do that and before the the session, we started talking about how people are measuring their compute in terms of megawws and and their revenue in terms of megawatt. Well, if you take that megawatt and you can actually put more racks in and each rack produces more tokens, that's how you generate more revenue per megawatt, right? And so if you if you take that chip, the same chip, and you increase the total output per chip, you're going to do a lot better when it comes to payback. >> What are some of the things you guys got going on that you could share on momentum side? Because that really is where everyone's looking at. You talked about some of the economics of buildout, but when you're in operations, you're building, operating, and investing. Everyone's doing those three things at the same time. Yeah. On the op side, what what do you see coming out? What do you guys have now? What's the momentum? >> No, look, I there are three incredibly exciting use cases that we see with some of technology. One, people who have already deployed a lot of NVIDIA gear. We're partnering with Nvidia on this with something called disagregated inference and what you're able to do is take NVIDIA gear make that partner with a SANOVA gear and in that construct disagregate the front front part called prefill and the decco part using sumova and generate 3x of total throughput. So existing hardware of Nvidia can jump 3x in total throughput just by putting someone rack next to it. And that's a really really exciting development for people who already have their gear and trying to get better economics. The other one that I think we're starting to see a lot of really exciting use use case for is this distributed data center existing data centers air cooled and being able to actually use brownfield data centers that don't require new investment new you know breaking ground new liquid cooling and just put some of aircooled technology into it and then get significant advantage just by actually delivering faster at a lower cost lower. So two two major things that you said there one is if I bought gear call it gear that's a term we use a lot bought a lot of systems normal normal depreciation would have that losing value now you got value creation so the price the value of that gear is actually more >> critical so that's good leverage >> that's a good use case great >> everyone signs up for that all day long I'm sure exactly and the other one is this new paradigm around disagregated serving >> which is also a precursor into disagregate infrastructure. Yeah. >> Because what you just said is essentially adding new new nodes out there, >> data center nodes to be AI factories basically without all the requirements >> for the megawatt gigawatt data centers that takes years to build. >> Yeah, that's right. I mean, one of the biggest challenge we're seeing in the world today is how quickly can we stand up gigawatt data centers and and that's going to be more and more challenging, you know, as as we think about the needs that we have and how much it impacts the energy grid. And so if we can actually reuse existing infrastructure in around the world, existing infrastructure in the United States where you already have energy allocated, ex space is allocated, it's air cooled, and you can deployed infrastructure to have state-of-the-art AI running on it, I think that's going to be an incredibly valuable way to actually get AI available to everybody at a much much lower cost. So, does the word edge of the network go away when you have distributed computing paradigm where you have large data center nodes like an AI factory in mega Texas or whatever it is and if you just got a cell tower that's got a building in power and network connectivity that's air cool in there that's an edge that's still a node on the network but it's technically an edge so smaller factory configuration >> yeah we actually call yeah we we we think of it as these massive kind of buildouts uh where you have huge gig go about data centers that's going to continue because people have to train models and the the necessity for large clusters continue to exist. Now you go into what I'm calling distributed data centers. So take the 50 top metropolitan cities and Vista for example is a partners that building out these distributed data centers in existing brownfield data centers re re-energizing existing infrastructure with existing power existing cooling and then running that using someone gear. Now edge edge is the next thing I'm super excited super excited because if you think about true edge >> which is where our mobile towers are where our users are robotics all these different things there is still another wave coming where very very localized computing is going to be serving very localized use cases of AI and that's to me the true edge of comput >> and they need high performance >> incredible high performance ultra low latency because now you're starting to talk about robotics and you talk about kind of the things that people want to use every day and frankly as a population, our patience is not very high. And so, so >> I want to get your thoughts. I want to get you brought up robotics. I want to bring up safety because the number one conversation in all of our physical AI robotics series that we're running, safety is the number one conversation. Security is in there, but I'd say number one Basically 1 A would be safety because robotics you got to have a safe car, you got to have a safe robot. Yeah. >> Um that's really obvious. you go, okay, all the safety on AI conversation seems to be a tempest in the teapot because it's like people are working on safety. Yeah. What's your what's your thoughts on all this negativity around safety? Um um you got half the world be like, okay, I don't really understand. What about I hate it. I'm I'm afraid they don't understand. Then you have the people who understand going, "Wow, this is one of the best revolutions of all time." Yeah. In the computer industry. So you got it's 50/50. >> Yeah. Yeah. No, it's understandable. Look, I think you look at a technology as transformative as this and you saw it over the weekend with the frontier models and now you know you know and more broadly around the model side that yeah safety is important and it's a great great wakeup call for all the leaders to start thinking about the fact that we've invested in the R&D of the models but how do we invest in the safety and the and the guardrails around kind of how we use it and so that's incredibly important. And I think you're going to continue to see attention on to that and we saw saw this in early internet days where safety and security kind of became um something that um was in the forefront people's minds and entire industries got created from it. And >> you know I have a lot of respect for Daario but I do think that he's not the poster child for safety when it comes to AI. I think he's he's no noble effort to lay out his concerns when the air company's got a great track record in terms of the ethics. So you give them the props for that. But there are many people in in the industry that actually have been through this before and have actually done it. Yeah. >> Have seen where you had let chaos rain, re rain in the chaos situations. Yeah. >> What do you think the the he could learn from what Daario and others that are now aware of this could learn >> from the history of how innovation can be chaotic and then reigned in. That's Andy Gro's favorite expression. Let chaos rain and reign in the chaos. Look what it happened with Intel under his regime. under Gordon Moore. >> Yeah. Well, I mean, look, I it starts with uh uh emphasis and attention. I think, you know, we're we're at a time where the industry is starting to pay a lot of attention. And look, the the AI genie is out of the bottle. It's not we're not going to go back in, right? But that said, uh it's not too much to me as much about slowing things down, but shifting your attention to investing in the safety piece, right? Because there are portions of it that requires the attention. There's a lot to be learned from the past, but requires the attention. And I think being able to actually take that energy and devote to it, I think it's going to make those models significantly better. >> It's funny, I watch some of the mainstream uh programs on TV and I read a lot and the perspective on um AI and it's kind of bothers me like oh that that that group of people tech people are controlling AI s an actor on on one of the shows kind of kind of you know laying into the tech industry. I'm like well it's not just them. There's a lot of other people involved. But the question that that that I ask is >> rhetorical if you inject intelligence into something >> a network edge like you just pointed out. Yeah. >> What happens if you inject intel? So I think there's a right to be concerned guard rails and keep watching it. You don't want to let chaos take over certainly. But yeah, I think we're going to have an experimentation. We we have to identify that. So that's noble. >> But there's there's a there's a real question to ask. What does it mean to inject AI intelligence into a process into a device? Mhm. >> You're at the center of it. How would you look at that? How do you frame that? >> Well, we we look at it, you know, two pieces. There's the, you know, kind of frontier model and and and the race towards AGI and then we have infrastructure on the infrastructure side and that's getting smarter too. and sum it over really focus on all the challenges around kind of creating safe, secure and you know cost-effective uh infrastructure and so when I look at that I think about one of the biggest challenges of infrastructure is making sure that it's sustainable and making sure that businesses aren't in a bubble right it's all about producing infrastructure that has good long-term econom economic value and allows you to actually consistently build and that's kind of what we're thinking about >> drive that payback down you know time to pay back down, drive that ROI up and making sure that people can actually invest on this side of the infrastructure in a sustainable and long-term way. >> Yeah, I love that that that infrastructure side. Let's go there for a second because you brought this up earlier. I want to go back to it because I think it's one of the most nuance undertalked about topics. Yeah, >> certainly in the tech circles it's talked about a lot. Disagregated serving. Yeah, that's one of your killer use cases for Sanova where you know everyone loves Nvidia. Hey, give me the GPUs. Everyone knows there's scarcity there and now there's still other now computes booming with agents. So there there's still infrastructure but this aggregated serving is really solves one of the major constraints. Yeah. Which is the volume of data and the lack of network coherency around managing it because >> prefill is the prompt and then the decode is what happens after on the math side. So you got math and math talking [snorts] to math. >> Yeah. >> What does that tell us? Because that is a constraint that we see in other areas. You mentioned disagregated infrastructure. Everyone's going to have nodes. >> That's disagregated. Yeah, >> there's serving there too, maybe on the edge. But what is the disagregated serving point to? It's not just bolting on accelerators. It really is a paradigm architectural shift. >> Well, it's a precursor to kind of what I think is going to be broad long-term view of data centers, which is heterogeneous computing. And so, you're going to be be able to mix and match different technologies to run what you need for AI. And so, you can actually have different chips running prefill, different chips running decode, and frankly, different chips running applications that the the agent's going to call. And so with SNOVA, you know, we're squarely in the decode side of it. And so we're able to actually take >> older infrastructure, say you have an H100 or now bees will soon be older as well. [laughter] You know, if you think about kind of the the the the infrastructure that people have invested already, how do we actually extend the life and you can actually disagregate it? Put someone over Iraq and suddenly you've extended the value of that infrastructure for another two three years. And so that's an important thing for people to think about because that amortization of the infrastructure is something that a lot of people are really concerned about >> and really speaks to your economic opportunity for Simonova because we were just talking on our Q pod last last episode about how in history of the computer hardware industry people would always want forward pricing. Now they're locking in pricing because they think it's going to go up. >> They think that the price of the H100's actually are not going to decay as fast. >> If anything might even go up. >> Yeah. because of the innovation happening around it. >> Yeah. Yeah. What interesting times where you know you got a convergence of many things, right? You convergence of the the these models and the performance, this unlimited demand, you know, this buildout that's incredible. And then you've got the supply chain constraints coming in. And so those three things actually coming together is actually driving a very very interesting economic model. But here's here's a we can here's something we can all agree on. something we can all agree on where when it comes to inference, everybody wants to see the cost in inference come down, right? And so whether whether it's better cost in the supply chain, whether that's, you know, being able to actually lower the energy or data center costs or whether that's actually finding more supply of different types of architectures, we can all agree that we need to drive the cost >> and the energy is the key function. All right, what's up for you guys in the second half of the year? We got the cube will be at open compute supercomputing as reinvent a variety of other events uh infrastructure events happening in Silicon Valley uh uh as well big announcement we're not going to be there we're here in New York what's on your focus area for the second half of the year obviously inference is super hot is that still doubling down on inference is that still the name of the game for you guys yeah what's the focus >> yeah no it's all about economics is all about payback for for for uh data centers being able to actually drive these services we think that uh uh with uh the launch of SM50 which is coming very soon here in terms of first shipments out I think people are going to see an incredible opportunity for them to actually take and drive profitability into their services after many many years of actually investing investing investing they're going to be able to start matching up their infrastructure that they have today with some of technologies and drive a much much higher profitability into their businesses great to see you final question give a taste of for the folks watching what the business performance has been for your uh and what's your outlook? >> You know, look, I mean, we're enjoying we've done six quarter of quarter over quarter doubling and I think we'll continue to see incredible demand. Uh the business is growing really, really fast and really, you know, showing that uh uh inference is something that people have a lot of interest around and it's the right time for people to invest. >> It's going to get bigger when the edge and everything gets connected. Thanks for coming on the cube. >> Yeah, thanks for having >> John Furrier. This is the AI factory series. This is one of our most popular series. The AI infrastructure continued to accelerate and expand. This is just AI factories. These big centers of data centers, they'll go to traditional data centers in the enterprise. You'll start to see the edge develop and wearables. We all have our ring or our whoop. They're all going to be connected too. We'll bring that coverage to you from the cube. Thanks for watching.