Jordan Nanos, SemiAnalysis | theCUBE + NYSE Wired: AI Factories
Watch on YouTubeVideo summary
Jordan Nanos, a semiconductor and AI infrastructure analyst, joins the discussion to provide an in-depth look at the rapidly evolving landscape of "AI factories," which serve as the critical infrastructure layer fueling the next wave of technology. He explains that while there is immense excitement and significant capital injection into this sector, the market is currently characterized by both manic growth and growing skepticism regarding the sustainability of these investments. Nanos highlights that the primary driver behind the current boom is an incredible demand for compute power from major AI labs like OpenAI and Anthropic, as well as enterprise companies, which are spending aggressively on GPUs to train and deploy models at scale. This surge in demand has created a situation where supply constraints—limited by physical factors such as data center power capacity, land availability, cooling equipment, and manufacturing limits on chips like those from TSMC—are beginning to dictate success more than technical merits alone.
A central theme of the conversation is the concept of "goodput," which Nanos defines as the effective amount of useful work that can be performed with available GPUs, accounting for real-world reliability issues rather than just peak theoretical performance. He details his firm's rigorous testing methodology, which involves simulating hardware failures to evaluate how quickly providers can detect and repair issues, noting that a single failure in modern high-density racks can take down an entire cluster costing millions of dollars per hour. Nanos argues that the true competitive moat for neoclouds lies not just in raw speed but in their ability to respond rapidly to market demand signals; companies that can deliver GPUs within months rather than the standard eighteen-month construction timeline are securing significantly higher returns on capital. Furthermore, he warns of severe security risks, including outdated software versions containing known vulnerabilities and the potential for AI models themselves to exploit zero-day flaws, emphasizing that trust in a provider's entire operational chain—from OEM technicians to brokers—is essential for safety.
The discussion also touches upon the economic dynamics of the industry, revealing a pricing curve where immediate access to capacity commands a significant premium compared to prepaid contracts with longer wait times. Nanos points out that despite the hype, there are hard physical limits to how much hardware can be produced and deployed, suggesting that demand will likely continue to outpace supply for the foreseeable future as models improve rather than decline. He notes that while metrics like weekly revenue run rates are often used to project valuations for companies planning IPOs, the reality is that these businesses are growing so fast that traditional financial models struggle to keep up. In his final recommendations, Nanos identifies Oracle as a preferred hyperscaler due to its pioneering networking technology, praises CoreWeave for its top-tier reliability ranking, and expresses strong support for non-Nvidia chip startups like Tranium, ultimately concluding that Jensen Huang remains the undisputed leader in the industry's vision and execution.
Read the full video transcript
Palo Alto Studio Connection Silicon
Valley and Wall Street. I'm John B here
with Dave Volante, my co-host.
Welcome back to the Cube Studio here at
the New York Stock Exchange. I'm Gemma
Al with NYC Wired. This is AI factories
where we talk all things the
infrastructure layer fueling the next
wave of technology. Joining me now for a
conversation on exactly that is Jordan
Nannis, semiconductor and AI
infrastructure analyst at semi analysis.
Welcome Jordan.
>> Great to be here. Thanks for having me.
>> So you operate an interesting space. I
was very excited for this conversation
because I love the opportunity in the
world of AI factories or we hear
everyone talk their good game every day
in day out in the show, zoom out a
little bit and talk about the more macro
picture. Right.
>> Yeah.
>> Interesting time. So much money being
spent, a lot of skepticism, a lot of
excitement. I mean, the markets are
manic quite frankly. Maybe just to
start, talk to me a little bit about
your work. You know, what you're you've
been thinking on this last month or two
even because below that it seems like
it's irrelevant now in the world of tech
and we can maybe go from there.
>> Yeah, definitely. So, primary thing that
I work on at semi analysis is called
cluster max. It's a rating system for
all of the neoclouds in the industry. As
you said, AI factories are incredibly
important. Neoclouds are what a lot of
the Neols are using to uh build the AI
infrastructure that they're going to
need to train models, deploy them at
scale. Biggest thing we're seeing right
now is just incredible demand. So that's
causing a lot of people to look good. Uh
a lot of, you know, little cracks to
form where people uh start to need to
build, you know, new technologies and
start to figure out what the problems at
scale are that are different than what
happened in the seed round when they
were starting up the Neo Neocloud, for
example. Let's start on the cracks.
Yeah. Right. So, Neoclouds, [music]
Jensen said this year at TTC, I mean,
it's hard to talk about instructor. I
talk about Jensen. So, let's just go
there first that there's a new NeoCloud
every day, right? Like, how do you truly
differentiate? A lot of it is just
demand based. You know, have we really
separated the men from the boys in that
space? Do we know for sure? What are
your thoughts? I mean, it seems like
it's a space that's getting so much cash
injection. Do you think that the numbers
add up? like what's your thought on the
economics of this model?
>> Yeah, I mean so obviously supply demand
and I think at a basic level uh demand
is there. We're seeing absolutely
massive ARR growth from OpenAI and
Anthropic in particular, but also uh
really strong growth from a lot of
enterprise companies and you know stuff
like Gemini or Grock from uh you know
Google and SpaceX as well as a lot of
the startups that are just raising
incredibly large rounds and just they
need to spend that money on compute and
where that plays out is that um
everybody started out by picking a dance
partner where the big labs openai
anthropic they were finding individual
NeoClouds that could go faster for them
than the hyperscalers could and then
they just needed to buy from everybody
and so you can look at how openai is
buying from Microsoft how they're buying
from coreweave how they're going down
the list and even how they're doing
self-build and then you can look at
anthropic and you can see you know
project rainer with tranium and at AWS
they've got CPUs with Google they've got
GPUs with Nvidia uh they've got a core
deal you know there there's all sorts of
ways in which these guys can figure out
ways to spend money so they can get
access to GPUs and then return it at a
significant right? Like if they were not
returning massive cash flows from these
GPUs, uh they wouldn't be buying them
right now. But we're seeing them go for
more than what anybody can provide. And
I guess that's just driving,
you know, supply, right? And uh I think
there's a lot of different ways in which
supply plays out. Obviously, Nvidia,
it's been great for them. Um but there's
there's limits in terms of like how many
balance sheets you can put GPUs on. uh
how much of data center power, land, uh
cooling equipment you can actually like
get set up in a certain amount of time.
Um how many GPUs you can get access to.
And as those dynamics play out, we start
to see a lot of like the business level
stuff dictating who's being successful
rather than the technical merits. uh
which I think is kind of what you
implied in the in the question that um
if demand continues to be so strong,
it's really not going to matter who's
got better reliability or who's got
better um you know, performance, which
is obviously what we spend almost all of
our time focused on in the uh like
cluster max rating system. And you know,
much to our chagrin, a lot of people
that are just focused on building data
centers of relatively low quality as
fast as they can are being successful
right now. I want to get into cluster
max and I want to talk about the
technical performance side of it but
first you said something interesting
there right relationships have been
formed people are getting into bed
together we've seen this really compound
actually over the last kind of 6 12
months even alone yeah
>> how are your thoughts what are your
thoughts so on you know the true moat of
that model like how do you think inbius
versus a coreweave you know versus a
hyperscaler like truly adds some sort of
like competitive advantage three years
from now like what is the one metric you
think is non-negotiable there.
>> Yeah, it's speeds them up. At the end of
the day, the amount that a NeoCloud can
respond to a demand signal in the market
is going to dictate whether they're
going to return capital on what they've
invested in GPUs. So, at this point, uh,
frankly, SpaceX is the leader in speed
from our tracking. They built Colossus 1
and two in absolutely record time. Many
of the other NeoClouds, Core Weave,
Nebius, they've had their problems.
They've had their delays in construction
and getting air permits and all this
stuff to turn on a lot of these sites.
And you can't start rec recognizing
revenue until you onboard those
customers. I think we're we we've seen
the hyperscalers pour in a ton of money
into capital. Like they're buying land,
they're buying a ton of equipment,
they're hiring construction contractors
all over the world. And this is to the
tune of like a trillion dollars of capex
or more next year. So they've got base
load figured out. and the neoclouds the
rest of these projects which in some
cases involves these hyperscalers like
Google's project with Blackstone for
example they're becoming neoclouds of
assort themselves where Microsoft's got
to go procure capacity from lambda or
from uh nscale you know and they they've
got to find that little bit of flex on
top so that they can continue to respond
to the demand signals which are so
strong from the market and what that
turns into is that the returns on
capital for the base load are they're
okay But the really strong returns are
for people who can deliver you GPUs
three months, one month now as opposed
to a project that takes 18 months to
pass all the permitting and start
construction and then finally roll stuff
in later.
>> Talk about cluster max. I know you talk
a lot about goodput, right? This kind of
metric of man managing I guess
performance versus cost versus output.
First of all, break that down, but also
I'm interested to understand what is
your evaluation based upon like what
like talk me through how you come up
with these assessments.
>> Yeah. So there's two big components. One
is just talking to the buyers about
their experience. They run at a bigger
scale than we can during our testing
which is the second component. And uh so
the the the buyers from these uh markets
are you know they're they're very
informative in terms of what their
experience was for support, for
reliability, for performance, things
like that. We also do our own hands-on
testing to validate what we're seeing
because we can get mixed signals from
those providers and uh a lot of the you
know yeah a lot of the buyers as well.
So in our our testing process is kind of
three phases. We start with an audit. We
focus a lot on security right now making
sure that they are providing a secure
cluster. Um really terrible results
there. We are putting out an article
next week going through some of that. Um
the second component is performance
where you know we test to see that
really for Nvidia GPUs or AMD GPUs we
know what the performance expectations
are and so we're just testing that they
meet these expectations. We're not
really differentiating at like a level
of you're 2% faster on a training
workload. Your microbenchmark on
networking is 5% faster. These things
come out in the wash in many cases when
you're working handinhand with a good
provider. But what we notice obviously
is when performance is just way below
the actual expectation which has
happened many times. And then to your
point good put the third component which
is the most critical for our hands-on
testing is reliability. Uh what we do
there is we simulate a series of
failures on the hardware and we check to
see how the provider reacts. Um, so this
is a component of first of all having
the monitoring and software systems in
place to be able to identify a failure
and then second of all being able to
actually repair or just like replace a
node for example uh a switch a cable
things that we can um simulate are going
to happen in the real world.
Interestingly our testing process is
about a week in a cluster of just four
nodes and we see real hardware failures
too.
>> Wow. Um, so yeah, the the overall
experience of like roughly seeing
goodput like the definition is how much
good work you can do with the GPUs that
you have, how much throughput is good,
which is to say if I'm training a model,
I'm running it for a month and I get a
peak amount of performance when all my
GPUs are online. Um, that's great to
know. It's great to know that metric,
but it's also important to know how
things perform when GPUs start failing,
which they do. And a lot of people have
a dependency on the providers,
especially with these new GB200 and
GB300 racks where there's a big scaleup
domain and one failure can effectively
take down the full rack of 72 GPUs. Um,
this is like thousands of dollars an
hour or hundreds and um yeah, like just
a rack is about $5 million capital
expense up front. So divide that by your
three-year contract like this is a lot.
>> What are you seeing from the perspective
of correction time, right? like a node
fails, a GPU fails, which you're saying
it happens all the time. There's
multiple layers of ownership there
though, right? Yeah. You have one
provider who's providing the service,
but like there is a whole technology
ecosystem feeding that failure.
>> Who who's doing like what what needs to
happen for it to be managed well?
>> Well, I think in in an interesting way,
it's not always one provider. In
NeoCloud world, you have the people that
you sign the contract with for the GPUs.
Sometimes you might have somebody who's
like playing a broker role. Another then
you have the people who actually like
operate the bare metal cluster. Uh these
are still software engineers who don't
live near the site usually. And then
there's the data center technicians who
are like actually swapping parts when
something fails. Many times they depend
on an OEM. So there's like the OEM
technicians that come in. It's this
whole chain of people that you need to
trust. And so roughly speaking our best
most reliable experiences have been with
providers that own that entire chain.
like the guy who is a data center
technician who is swapping a cable or a
drive has like equity in the company
that you signed a contract with and
really cares about them being
successful. There's a lot of other
situations where you're four layers
removed from those people. The broker,
the Neo Club that's selling you GPUs is
kind of trying to hide who it actually
is because otherwise they think you're
going to go circumvent them and go
around them. That's not a great way to
start a a relationship. Now there's
there's many cases where you know
construction company hands off to data
center operations company which you know
employs the technicians who then uh hand
off to the SRRES who build the cluster
who then you work with and those are
there's many models that are great for
that as well and work well. Uh but
generally speaking when we do that
testing we expect to see some sort of
ownership some sort of response time
some sort of like you know intelligent
response from organic intelligence like
a human not some automated chatbot that
gives us asurances like this is when
your stuff's going to get fixed. This is
how long it's going to take. This is
what we found happened. Um, it's not
necessarily a secret shopper experience,
but when people are unprepared and they
haven't, you know, reviewed the criteria
that we're using or asked what we're
going to be up to, I mean, a lot of them
get caught by surprise.
>> I love what you are though identifying
some of the used car salesmen in the
process too, right? Which is an
important part because we hear a lot
about the world of NeoCloud, the
financing programs, you know, that kind
of sub layer, right? Where there is a
lot of question over how that can
actually truly be efficient like longer
term. So back to what you said at the
beginning. I know you said it's not
going to be released till next week, but
security. It's an interesting point
though, right? Is there like some blind
assumptions being made like by you know
technical leaders that you feel are
somewhat you know highly misinformed
like when you say it's surprisingly bad
like give me some context to that.
>> Yeah surprising in some ways uh
unsurprising in others when you get your
experience like working with these guys.
Um I just described that chain right and
you can kind of imagine that in when
that goes wrong it's like broken
telephone trying to get something fixed
and things don't work quite the way you
expect or on the timeline you expect. So
real example is many of the labs are
bringing their CO onto these calls
because it's a counterparty risk who you
decide to go with and you need to trust
these people. Um simple example is just
keeping software up to date on the
cluster. Uh we're seeing two dynamics
right now. One is that uh all these
frontier models uh when you look at
project glasswing from anthropic or
everything with open AAI uh and their
codeex model and that um announcement
that they made with hugging face where
they showed how the model was
autonomously hacking hugging faces data
sets infrastructure to try to get
answers to an eval during training
without any human involved. Uh really
scary interesting talk from Black Hat
Summit that everybody should go watch.
>> Oh wow.
>> Gives you a taste for the future here.
But um yeah, the point is that uh these
models are finding zero days in
software. They're finding existing
vulnerabilities that humans don't know
about and they're exploiting them. Um
and so as models get better, we only
expect more of that to happen, more
sophisticated exploits of the software
that we have that we use today. Um, but
the second thing is like when these
CVEes come out describing this
vulnerability and they tell people to
patch it, it's up to your provider to
roll out a fix to be keeping track of
when things are coming out. in some
cases being in in an embargo program
with companies like Nvidia or AMD so
they get advanced notice that these
disclosures are going to happen and they
can prepare a patch so that on day zero
when that vulnerability gets disclosed
they can already start the roll out of
upgrading your software so that you're
not getting exploited. A lot of these
providers that we check have stuff
that's not like a month or two months
old but like 3 years old that they
haven't upgraded. And this just means
it's a it's a ticking time bomb until
somebody comes along and looks and um
you know starts looking for your model
weights or your data sets or your RL
environments that you really care about
keeping private for example. Um or data
excfiltration isn't the only thing. I
mean think about ransomware or think
about these like crypto mining hackers
that take over clusters. Like we hear
about all sorts of bad stuff that
happens from security.
>> What are your thoughts on this narrative
that openweight models are less secure?
>> Okay, so that's the um second part of
it. uh they are uh certainly less
guardrailed um if I can use that uh word
um not sure if it is a word but yeah the
the point is like they uh can be used.
So we try to build um POC exploits in
order to explain to providers exactly
how somebody can use these uh these you
know zero days in their or sorry these
vulnerabilities that are in their
environment and how they would be
exploited because a lot of them will
push back and say oh this old software
version you know it doesn't matter I
don't need to upgrade it my pro my
customer said it's okay and in some
respects okay it's the customer's
decision but in other respects they're
just saying that um And so yeah, we try
to demonstrate these things and and we
you you can't even ask Fable about a
security issue. You can't ask it to
check your own cluster. It's going to
deny you immediately. Um
>> Soul is a little bit better, but we have
to use a lot of these openweight bottles
because they don't have guardrails on on
them. Um that reject any sort of
research into security. Um even if
you're in the security program and and
approved by anthropic or open AAI like
we are in some cases. Wow.
>> Um, so that dynamic totally exists where
it's a bit of a wild west with the
Chinese models where I would um say
people don't necessarily have the same
cause for concern is that in our
experience they are still remarkably
um like significantly worse at
exploiting um these security
vulnerabilities than the leading
frontier American models. Uh I would say
Chinese China has not at least
demonstrated a focus on cyber security
in the open weight models
>> for now though right like that is that
is the broader concern. Okay, so wild
wild west. I mean you could also say
from the perspective of macroeconomics
the entire industry is a bit of a wild
west, right? Like if you think about how
we create any sort of unilateral
agreement on what a prediction looks
like for GPU costs 5 years from now,
right? [music] You're like so much
capital flowing into the space. You're a
bank, you're lending money against these
bets. How are you truly evaluating it?
Right? that is a conversation that is
kind of floating out there that
everyone's kind of skirting around. What
are your thoughts from the perspective
of how you do actually have some sort of
unique performance measure that can feed
financial predictability like what does
that look like?
>> Yeah. Um so I think there's the first of
all you need to segment the market in
terms of how people make these
decisions. The top level is like
Anthropic or OpenAI or Google or Meta or
Microsoft just taking full sites. And
the way in which they do these deals is
like completely different than the way a
new, you know, Silicon Valley startup
who got funding for GPUs is going out
and trying to get one tenant like one
section of a bigger cluster from a
Corewave or Nebus or Crusoor or Lambda
or Together any of these uh Neoclouds
that are out there. Um, and then there's
the kind of like bottom end of the
market, which is you or me paying for
tokens or people at home that are doing
development on like a single GPU at a
time. And um, there's clearly this like
backwardation in the pricing curve right
now where if you are willing to put up
money and prepay for stuff, you can get
a significant discount on the total
contract value, but you got to wait like
six months for stuff to get installed,
right?
>> If you want stuff now, you pay a
significant premium. And um if you want
tokens like tokens on a per GPU hour
basis come at a significant premium on
top of that and sort of on demand people
only want tokens or want to have a GPU
for an hour or two uh they they pay a
significant premium as well. So I think
the market is shaping out where um kind
of like what I was saying earlier, many
companies need to have this like
long-term base load committed capacity
of GPUs or tokens and then they need to
do the engineering work to understand
what their demand is going to grow like
in the future and properly plan to like
bring stuff online or have the cash set
aside or the relationships available to
kind of flex up and down and get access
to the stuff they need. But generally
speaking, we see people buy more GPUs,
not less. We see very few people giving
stuff back and there's limits that are
being reached in the supply chain of how
much can be produced and how much can be
turned on. We expect those to continue
for a long time. I think we are like the
industry experts on understanding how
much can be produced on the supply side
of this curve. There are real limits to
that. There's only so many wafers from
CSMC. There's only so much HPM, only so
much DRAM.
So there's limits. Um and uh until the
models get worse, I find it really hard
to understand scenarios where demand is
going to trail off and and fall off a
cliff. So last question, let's talk
about something I mentioned to you
before the show. I heard Dylan Patel
commenting on the anthropic and openi
evaluations, how the market has
responded, you know, some of those kind
of fuzzy metrics that have been used to
determine revenue. I mean both these
companies are predicted to go public
between now and I guess 2027 at some
early stage in the year. What are your
thoughts like you know what what do you
what what comes to top of mind to you
when you think about this?
>> Yeah. Um I mean historically like it's
it's super strange to see people take a
uh a weekly WR and multiply it by 52 or
something like that and call that the
exit. [laughter] Um but uh
>> you're Sam and your Dario. You write
your [clears throat] own rules, right?
[laughter]
>> Yeah. I think investors want to
understand the growth rate of the
business and I think it's really hard to
contend with the fact that these
businesses are growing incredibly fast.
Let's put OpenAI and Enthropic aside for
a second. I was talking to a software
company earlier this week who was in the
middle of raising their series B.
They've had incredible growth. They were
doing it during a big conference. They
go out to the investors at a certain
price. They finish the conference.
They've got a massive pipeline and they
go, "Look guys, I don't think we need
the money right now. I can give you a
sales force export at the end of this
week and we put can put a multiple on
top of that and repric the round, but
like maybe let's just check in in two
months and see what happens." And so
everybody says, "Okay, let's see how
growth goes for the next two two, you
know, two months. Like you're not net
profitable, but you got a lot of runway.
Like it's all good." And then two months
come and the growth just continues. So,
uh, I think in this case, like it's kind
of valid to have both metrics.
>> You'd like to know what the current
>> run rate is multiplied by whatever
factor you care about.
>> Um, and look, Antropic entered this year
uh projecting 100 billion ARR. I think a
lot of people were a little bit uh
questioning whether they would come
through on that, and they're going to
they're going to come through way before
December on that. So, uh they're going
to Yeah, they're going to cross
100 quite soon. So, um, by our modeling.
Anyway,
look, these businesses are incredible.
Like what I said earlier, they turn on
more GPUs, they get more revenue. It's
almost a direct line from all of our
tracking of their data centers and their
chips installations.
>> Can we do a quick like hot takes, couple
of questions at the end?
>> Sure.
>> Okay. Hyperscaler.
Biggest kind of favorite hyperscaler.
>> Favorite hyperscaler right now. Uh, when
it comes to construction, AWS is the
fastest but can't stand EFA. So, uh, I'm
going to go Oracle. They, uh,
>> Wow.
>> Yeah, they've pioneered a lot with the
multiplaner, um, Rocky networking. Love
that.
>> Might be a good time to buy Oracle
stock, too, right? [laughter]
>> This is not an endorsement of their
stock.
>> Joking.
>> Buy the research, everybody. Yeah. Yeah.
>> Caveat that. Um, chip company outside
Nvidia.
>> Oh man. Um, startup or like real
project?
>> Real project.
>> Okay. TPUs are amazing. Um, a lot of
people, there's so much demand for TPUs.
They're really cool. Uh, a lot of cool
stuff on the road map, too. Um,
I'll put Tranium in there as well. I I
I've I've had I've had some good
experience doing microbenchmarking with
train tranium. Love the profiling. Uh,
chip startup. I don't know. I can't
pick. Don't want to give too manybody
too strong an endorsement. Really
impressive what a lot of them are doing.
Need to see these guys produce tokens.
Okay. Not these deals, these like cool
marketing videos. Produce some tokens.
Chip startups. Let's go. Favorite
Neocloud.
>> Favorite Neocloud. I mean, Cory's been
the top of the platinum tier rankings.
They're great experience. Um, yeah,
we'll go with that.
>> Let's Let's end with an easy one.
Favorite tech leader.
>> Favorite tech leader. Uh, Jensen.
>> Ah, [laughter]
>> you know, in video this year, we asked
folks, who's the bigger celebrity, Jesus
or Jensen? You know what the answer was?
>> Who's Jesus? [laughter]
>> Jared Lannis, thank you so much for
joining us on the cube and NYC Wire.
>> Okay, great to be here. I'm Jim Allen
coming to you from the Cube studio at
the New York Stock Exchange. This is AI
Factories, one of our programs with NYC
Wired. Thanks for watching.