Submind YouTube summaries
Thumbnail for Jordan Nanos, SemiAnalysis | theCUBE + NYSE Wired: AI Factories

Jordan Nanos, SemiAnalysis | theCUBE + NYSE Wired: AI Factories

Watch on YouTube

Video summary

Jordan Nanos, a semiconductor and AI infrastructure analyst, joins the discussion to provide an in-depth look at the rapidly evolving landscape of "AI factories," which serve as the critical infrastructure layer fueling the next wave of technology. He explains that while there is immense excitement and significant capital injection into this sector, the market is currently characterized by both manic growth and growing skepticism regarding the sustainability of these investments. Nanos highlights that the primary driver behind the current boom is an incredible demand for compute power from major AI labs like OpenAI and Anthropic, as well as enterprise companies, which are spending aggressively on GPUs to train and deploy models at scale. This surge in demand has created a situation where supply constraints—limited by physical factors such as data center power capacity, land availability, cooling equipment, and manufacturing limits on chips like those from TSMC—are beginning to dictate success more than technical merits alone. A central theme of the conversation is the concept of "goodput," which Nanos defines as the effective amount of useful work that can be performed with available GPUs, accounting for real-world reliability issues rather than just peak theoretical performance. He details his firm's rigorous testing methodology, which involves simulating hardware failures to evaluate how quickly providers can detect and repair issues, noting that a single failure in modern high-density racks can take down an entire cluster costing millions of dollars per hour. Nanos argues that the true competitive moat for neoclouds lies not just in raw speed but in their ability to respond rapidly to market demand signals; companies that can deliver GPUs within months rather than the standard eighteen-month construction timeline are securing significantly higher returns on capital. Furthermore, he warns of severe security risks, including outdated software versions containing known vulnerabilities and the potential for AI models themselves to exploit zero-day flaws, emphasizing that trust in a provider's entire operational chain—from OEM technicians to brokers—is essential for safety. The discussion also touches upon the economic dynamics of the industry, revealing a pricing curve where immediate access to capacity commands a significant premium compared to prepaid contracts with longer wait times. Nanos points out that despite the hype, there are hard physical limits to how much hardware can be produced and deployed, suggesting that demand will likely continue to outpace supply for the foreseeable future as models improve rather than decline. He notes that while metrics like weekly revenue run rates are often used to project valuations for companies planning IPOs, the reality is that these businesses are growing so fast that traditional financial models struggle to keep up. In his final recommendations, Nanos identifies Oracle as a preferred hyperscaler due to its pioneering networking technology, praises CoreWeave for its top-tier reliability ranking, and expresses strong support for non-Nvidia chip startups like Tranium, ultimately concluding that Jensen Huang remains the undisputed leader in the industry's vision and execution.
Read the full video transcript
Palo Alto Studio Connection Silicon Valley and Wall Street. I'm John B here with Dave Volante, my co-host. Welcome back to the Cube Studio here at the New York Stock Exchange. I'm Gemma Al with NYC Wired. This is AI factories where we talk all things the infrastructure layer fueling the next wave of technology. Joining me now for a conversation on exactly that is Jordan Nannis, semiconductor and AI infrastructure analyst at semi analysis. Welcome Jordan. >> Great to be here. Thanks for having me. >> So you operate an interesting space. I was very excited for this conversation because I love the opportunity in the world of AI factories or we hear everyone talk their good game every day in day out in the show, zoom out a little bit and talk about the more macro picture. Right. >> Yeah. >> Interesting time. So much money being spent, a lot of skepticism, a lot of excitement. I mean, the markets are manic quite frankly. Maybe just to start, talk to me a little bit about your work. You know, what you're you've been thinking on this last month or two even because below that it seems like it's irrelevant now in the world of tech and we can maybe go from there. >> Yeah, definitely. So, primary thing that I work on at semi analysis is called cluster max. It's a rating system for all of the neoclouds in the industry. As you said, AI factories are incredibly important. Neoclouds are what a lot of the Neols are using to uh build the AI infrastructure that they're going to need to train models, deploy them at scale. Biggest thing we're seeing right now is just incredible demand. So that's causing a lot of people to look good. Uh a lot of, you know, little cracks to form where people uh start to need to build, you know, new technologies and start to figure out what the problems at scale are that are different than what happened in the seed round when they were starting up the Neo Neocloud, for example. Let's start on the cracks. Yeah. Right. So, Neoclouds, [music] Jensen said this year at TTC, I mean, it's hard to talk about instructor. I talk about Jensen. So, let's just go there first that there's a new NeoCloud every day, right? Like, how do you truly differentiate? A lot of it is just demand based. You know, have we really separated the men from the boys in that space? Do we know for sure? What are your thoughts? I mean, it seems like it's a space that's getting so much cash injection. Do you think that the numbers add up? like what's your thought on the economics of this model? >> Yeah, I mean so obviously supply demand and I think at a basic level uh demand is there. We're seeing absolutely massive ARR growth from OpenAI and Anthropic in particular, but also uh really strong growth from a lot of enterprise companies and you know stuff like Gemini or Grock from uh you know Google and SpaceX as well as a lot of the startups that are just raising incredibly large rounds and just they need to spend that money on compute and where that plays out is that um everybody started out by picking a dance partner where the big labs openai anthropic they were finding individual NeoClouds that could go faster for them than the hyperscalers could and then they just needed to buy from everybody and so you can look at how openai is buying from Microsoft how they're buying from coreweave how they're going down the list and even how they're doing self-build and then you can look at anthropic and you can see you know project rainer with tranium and at AWS they've got CPUs with Google they've got GPUs with Nvidia uh they've got a core deal you know there there's all sorts of ways in which these guys can figure out ways to spend money so they can get access to GPUs and then return it at a significant right? Like if they were not returning massive cash flows from these GPUs, uh they wouldn't be buying them right now. But we're seeing them go for more than what anybody can provide. And I guess that's just driving, you know, supply, right? And uh I think there's a lot of different ways in which supply plays out. Obviously, Nvidia, it's been great for them. Um but there's there's limits in terms of like how many balance sheets you can put GPUs on. uh how much of data center power, land, uh cooling equipment you can actually like get set up in a certain amount of time. Um how many GPUs you can get access to. And as those dynamics play out, we start to see a lot of like the business level stuff dictating who's being successful rather than the technical merits. uh which I think is kind of what you implied in the in the question that um if demand continues to be so strong, it's really not going to matter who's got better reliability or who's got better um you know, performance, which is obviously what we spend almost all of our time focused on in the uh like cluster max rating system. And you know, much to our chagrin, a lot of people that are just focused on building data centers of relatively low quality as fast as they can are being successful right now. I want to get into cluster max and I want to talk about the technical performance side of it but first you said something interesting there right relationships have been formed people are getting into bed together we've seen this really compound actually over the last kind of 6 12 months even alone yeah >> how are your thoughts what are your thoughts so on you know the true moat of that model like how do you think inbius versus a coreweave you know versus a hyperscaler like truly adds some sort of like competitive advantage three years from now like what is the one metric you think is non-negotiable there. >> Yeah, it's speeds them up. At the end of the day, the amount that a NeoCloud can respond to a demand signal in the market is going to dictate whether they're going to return capital on what they've invested in GPUs. So, at this point, uh, frankly, SpaceX is the leader in speed from our tracking. They built Colossus 1 and two in absolutely record time. Many of the other NeoClouds, Core Weave, Nebius, they've had their problems. They've had their delays in construction and getting air permits and all this stuff to turn on a lot of these sites. And you can't start rec recognizing revenue until you onboard those customers. I think we're we we've seen the hyperscalers pour in a ton of money into capital. Like they're buying land, they're buying a ton of equipment, they're hiring construction contractors all over the world. And this is to the tune of like a trillion dollars of capex or more next year. So they've got base load figured out. and the neoclouds the rest of these projects which in some cases involves these hyperscalers like Google's project with Blackstone for example they're becoming neoclouds of assort themselves where Microsoft's got to go procure capacity from lambda or from uh nscale you know and they they've got to find that little bit of flex on top so that they can continue to respond to the demand signals which are so strong from the market and what that turns into is that the returns on capital for the base load are they're okay But the really strong returns are for people who can deliver you GPUs three months, one month now as opposed to a project that takes 18 months to pass all the permitting and start construction and then finally roll stuff in later. >> Talk about cluster max. I know you talk a lot about goodput, right? This kind of metric of man managing I guess performance versus cost versus output. First of all, break that down, but also I'm interested to understand what is your evaluation based upon like what like talk me through how you come up with these assessments. >> Yeah. So there's two big components. One is just talking to the buyers about their experience. They run at a bigger scale than we can during our testing which is the second component. And uh so the the the buyers from these uh markets are you know they're they're very informative in terms of what their experience was for support, for reliability, for performance, things like that. We also do our own hands-on testing to validate what we're seeing because we can get mixed signals from those providers and uh a lot of the you know yeah a lot of the buyers as well. So in our our testing process is kind of three phases. We start with an audit. We focus a lot on security right now making sure that they are providing a secure cluster. Um really terrible results there. We are putting out an article next week going through some of that. Um the second component is performance where you know we test to see that really for Nvidia GPUs or AMD GPUs we know what the performance expectations are and so we're just testing that they meet these expectations. We're not really differentiating at like a level of you're 2% faster on a training workload. Your microbenchmark on networking is 5% faster. These things come out in the wash in many cases when you're working handinhand with a good provider. But what we notice obviously is when performance is just way below the actual expectation which has happened many times. And then to your point good put the third component which is the most critical for our hands-on testing is reliability. Uh what we do there is we simulate a series of failures on the hardware and we check to see how the provider reacts. Um, so this is a component of first of all having the monitoring and software systems in place to be able to identify a failure and then second of all being able to actually repair or just like replace a node for example uh a switch a cable things that we can um simulate are going to happen in the real world. Interestingly our testing process is about a week in a cluster of just four nodes and we see real hardware failures too. >> Wow. Um, so yeah, the the overall experience of like roughly seeing goodput like the definition is how much good work you can do with the GPUs that you have, how much throughput is good, which is to say if I'm training a model, I'm running it for a month and I get a peak amount of performance when all my GPUs are online. Um, that's great to know. It's great to know that metric, but it's also important to know how things perform when GPUs start failing, which they do. And a lot of people have a dependency on the providers, especially with these new GB200 and GB300 racks where there's a big scaleup domain and one failure can effectively take down the full rack of 72 GPUs. Um, this is like thousands of dollars an hour or hundreds and um yeah, like just a rack is about $5 million capital expense up front. So divide that by your three-year contract like this is a lot. >> What are you seeing from the perspective of correction time, right? like a node fails, a GPU fails, which you're saying it happens all the time. There's multiple layers of ownership there though, right? Yeah. You have one provider who's providing the service, but like there is a whole technology ecosystem feeding that failure. >> Who who's doing like what what needs to happen for it to be managed well? >> Well, I think in in an interesting way, it's not always one provider. In NeoCloud world, you have the people that you sign the contract with for the GPUs. Sometimes you might have somebody who's like playing a broker role. Another then you have the people who actually like operate the bare metal cluster. Uh these are still software engineers who don't live near the site usually. And then there's the data center technicians who are like actually swapping parts when something fails. Many times they depend on an OEM. So there's like the OEM technicians that come in. It's this whole chain of people that you need to trust. And so roughly speaking our best most reliable experiences have been with providers that own that entire chain. like the guy who is a data center technician who is swapping a cable or a drive has like equity in the company that you signed a contract with and really cares about them being successful. There's a lot of other situations where you're four layers removed from those people. The broker, the Neo Club that's selling you GPUs is kind of trying to hide who it actually is because otherwise they think you're going to go circumvent them and go around them. That's not a great way to start a a relationship. Now there's there's many cases where you know construction company hands off to data center operations company which you know employs the technicians who then uh hand off to the SRRES who build the cluster who then you work with and those are there's many models that are great for that as well and work well. Uh but generally speaking when we do that testing we expect to see some sort of ownership some sort of response time some sort of like you know intelligent response from organic intelligence like a human not some automated chatbot that gives us asurances like this is when your stuff's going to get fixed. This is how long it's going to take. This is what we found happened. Um, it's not necessarily a secret shopper experience, but when people are unprepared and they haven't, you know, reviewed the criteria that we're using or asked what we're going to be up to, I mean, a lot of them get caught by surprise. >> I love what you are though identifying some of the used car salesmen in the process too, right? Which is an important part because we hear a lot about the world of NeoCloud, the financing programs, you know, that kind of sub layer, right? Where there is a lot of question over how that can actually truly be efficient like longer term. So back to what you said at the beginning. I know you said it's not going to be released till next week, but security. It's an interesting point though, right? Is there like some blind assumptions being made like by you know technical leaders that you feel are somewhat you know highly misinformed like when you say it's surprisingly bad like give me some context to that. >> Yeah surprising in some ways uh unsurprising in others when you get your experience like working with these guys. Um I just described that chain right and you can kind of imagine that in when that goes wrong it's like broken telephone trying to get something fixed and things don't work quite the way you expect or on the timeline you expect. So real example is many of the labs are bringing their CO onto these calls because it's a counterparty risk who you decide to go with and you need to trust these people. Um simple example is just keeping software up to date on the cluster. Uh we're seeing two dynamics right now. One is that uh all these frontier models uh when you look at project glasswing from anthropic or everything with open AAI uh and their codeex model and that um announcement that they made with hugging face where they showed how the model was autonomously hacking hugging faces data sets infrastructure to try to get answers to an eval during training without any human involved. Uh really scary interesting talk from Black Hat Summit that everybody should go watch. >> Oh wow. >> Gives you a taste for the future here. But um yeah, the point is that uh these models are finding zero days in software. They're finding existing vulnerabilities that humans don't know about and they're exploiting them. Um and so as models get better, we only expect more of that to happen, more sophisticated exploits of the software that we have that we use today. Um, but the second thing is like when these CVEes come out describing this vulnerability and they tell people to patch it, it's up to your provider to roll out a fix to be keeping track of when things are coming out. in some cases being in in an embargo program with companies like Nvidia or AMD so they get advanced notice that these disclosures are going to happen and they can prepare a patch so that on day zero when that vulnerability gets disclosed they can already start the roll out of upgrading your software so that you're not getting exploited. A lot of these providers that we check have stuff that's not like a month or two months old but like 3 years old that they haven't upgraded. And this just means it's a it's a ticking time bomb until somebody comes along and looks and um you know starts looking for your model weights or your data sets or your RL environments that you really care about keeping private for example. Um or data excfiltration isn't the only thing. I mean think about ransomware or think about these like crypto mining hackers that take over clusters. Like we hear about all sorts of bad stuff that happens from security. >> What are your thoughts on this narrative that openweight models are less secure? >> Okay, so that's the um second part of it. uh they are uh certainly less guardrailed um if I can use that uh word um not sure if it is a word but yeah the the point is like they uh can be used. So we try to build um POC exploits in order to explain to providers exactly how somebody can use these uh these you know zero days in their or sorry these vulnerabilities that are in their environment and how they would be exploited because a lot of them will push back and say oh this old software version you know it doesn't matter I don't need to upgrade it my pro my customer said it's okay and in some respects okay it's the customer's decision but in other respects they're just saying that um And so yeah, we try to demonstrate these things and and we you you can't even ask Fable about a security issue. You can't ask it to check your own cluster. It's going to deny you immediately. Um >> Soul is a little bit better, but we have to use a lot of these openweight bottles because they don't have guardrails on on them. Um that reject any sort of research into security. Um even if you're in the security program and and approved by anthropic or open AAI like we are in some cases. Wow. >> Um, so that dynamic totally exists where it's a bit of a wild west with the Chinese models where I would um say people don't necessarily have the same cause for concern is that in our experience they are still remarkably um like significantly worse at exploiting um these security vulnerabilities than the leading frontier American models. Uh I would say Chinese China has not at least demonstrated a focus on cyber security in the open weight models >> for now though right like that is that is the broader concern. Okay, so wild wild west. I mean you could also say from the perspective of macroeconomics the entire industry is a bit of a wild west, right? Like if you think about how we create any sort of unilateral agreement on what a prediction looks like for GPU costs 5 years from now, right? [music] You're like so much capital flowing into the space. You're a bank, you're lending money against these bets. How are you truly evaluating it? Right? that is a conversation that is kind of floating out there that everyone's kind of skirting around. What are your thoughts from the perspective of how you do actually have some sort of unique performance measure that can feed financial predictability like what does that look like? >> Yeah. Um so I think there's the first of all you need to segment the market in terms of how people make these decisions. The top level is like Anthropic or OpenAI or Google or Meta or Microsoft just taking full sites. And the way in which they do these deals is like completely different than the way a new, you know, Silicon Valley startup who got funding for GPUs is going out and trying to get one tenant like one section of a bigger cluster from a Corewave or Nebus or Crusoor or Lambda or Together any of these uh Neoclouds that are out there. Um, and then there's the kind of like bottom end of the market, which is you or me paying for tokens or people at home that are doing development on like a single GPU at a time. And um, there's clearly this like backwardation in the pricing curve right now where if you are willing to put up money and prepay for stuff, you can get a significant discount on the total contract value, but you got to wait like six months for stuff to get installed, right? >> If you want stuff now, you pay a significant premium. And um if you want tokens like tokens on a per GPU hour basis come at a significant premium on top of that and sort of on demand people only want tokens or want to have a GPU for an hour or two uh they they pay a significant premium as well. So I think the market is shaping out where um kind of like what I was saying earlier, many companies need to have this like long-term base load committed capacity of GPUs or tokens and then they need to do the engineering work to understand what their demand is going to grow like in the future and properly plan to like bring stuff online or have the cash set aside or the relationships available to kind of flex up and down and get access to the stuff they need. But generally speaking, we see people buy more GPUs, not less. We see very few people giving stuff back and there's limits that are being reached in the supply chain of how much can be produced and how much can be turned on. We expect those to continue for a long time. I think we are like the industry experts on understanding how much can be produced on the supply side of this curve. There are real limits to that. There's only so many wafers from CSMC. There's only so much HPM, only so much DRAM. So there's limits. Um and uh until the models get worse, I find it really hard to understand scenarios where demand is going to trail off and and fall off a cliff. So last question, let's talk about something I mentioned to you before the show. I heard Dylan Patel commenting on the anthropic and openi evaluations, how the market has responded, you know, some of those kind of fuzzy metrics that have been used to determine revenue. I mean both these companies are predicted to go public between now and I guess 2027 at some early stage in the year. What are your thoughts like you know what what do you what what comes to top of mind to you when you think about this? >> Yeah. Um I mean historically like it's it's super strange to see people take a uh a weekly WR and multiply it by 52 or something like that and call that the exit. [laughter] Um but uh >> you're Sam and your Dario. You write your [clears throat] own rules, right? [laughter] >> Yeah. I think investors want to understand the growth rate of the business and I think it's really hard to contend with the fact that these businesses are growing incredibly fast. Let's put OpenAI and Enthropic aside for a second. I was talking to a software company earlier this week who was in the middle of raising their series B. They've had incredible growth. They were doing it during a big conference. They go out to the investors at a certain price. They finish the conference. They've got a massive pipeline and they go, "Look guys, I don't think we need the money right now. I can give you a sales force export at the end of this week and we put can put a multiple on top of that and repric the round, but like maybe let's just check in in two months and see what happens." And so everybody says, "Okay, let's see how growth goes for the next two two, you know, two months. Like you're not net profitable, but you got a lot of runway. Like it's all good." And then two months come and the growth just continues. So, uh, I think in this case, like it's kind of valid to have both metrics. >> You'd like to know what the current >> run rate is multiplied by whatever factor you care about. >> Um, and look, Antropic entered this year uh projecting 100 billion ARR. I think a lot of people were a little bit uh questioning whether they would come through on that, and they're going to they're going to come through way before December on that. So, uh they're going to Yeah, they're going to cross 100 quite soon. So, um, by our modeling. Anyway, look, these businesses are incredible. Like what I said earlier, they turn on more GPUs, they get more revenue. It's almost a direct line from all of our tracking of their data centers and their chips installations. >> Can we do a quick like hot takes, couple of questions at the end? >> Sure. >> Okay. Hyperscaler. Biggest kind of favorite hyperscaler. >> Favorite hyperscaler right now. Uh, when it comes to construction, AWS is the fastest but can't stand EFA. So, uh, I'm going to go Oracle. They, uh, >> Wow. >> Yeah, they've pioneered a lot with the multiplaner, um, Rocky networking. Love that. >> Might be a good time to buy Oracle stock, too, right? [laughter] >> This is not an endorsement of their stock. >> Joking. >> Buy the research, everybody. Yeah. Yeah. >> Caveat that. Um, chip company outside Nvidia. >> Oh man. Um, startup or like real project? >> Real project. >> Okay. TPUs are amazing. Um, a lot of people, there's so much demand for TPUs. They're really cool. Uh, a lot of cool stuff on the road map, too. Um, I'll put Tranium in there as well. I I I've I've had I've had some good experience doing microbenchmarking with train tranium. Love the profiling. Uh, chip startup. I don't know. I can't pick. Don't want to give too manybody too strong an endorsement. Really impressive what a lot of them are doing. Need to see these guys produce tokens. Okay. Not these deals, these like cool marketing videos. Produce some tokens. Chip startups. Let's go. Favorite Neocloud. >> Favorite Neocloud. I mean, Cory's been the top of the platinum tier rankings. They're great experience. Um, yeah, we'll go with that. >> Let's Let's end with an easy one. Favorite tech leader. >> Favorite tech leader. Uh, Jensen. >> Ah, [laughter] >> you know, in video this year, we asked folks, who's the bigger celebrity, Jesus or Jensen? You know what the answer was? >> Who's Jesus? [laughter] >> Jared Lannis, thank you so much for joining us on the cube and NYC Wire. >> Okay, great to be here. I'm Jim Allen coming to you from the Cube studio at the New York Stock Exchange. This is AI Factories, one of our programs with NYC Wired. Thanks for watching.