Video summary
The current landscape of artificial intelligence development is defined by a severe scarcity of essential hardware, particularly GPUs and RAM, which has forced the entire industry to adapt its strategies. As the costs for these components rise significantly, customers are increasingly rushing to lock in server purchases immediately to secure available inventory and stabilize pricing. This shortage means that relying on a single preferred vendor is no longer a viable strategy, as suppliers may run out of stock or make their platforms prohibitively expensive. Consequently, organizations are being forced to adopt heterogeneous designs from the outset, preparing themselves to switch vendors or intermix different hardware sources depending on supply chain realities and availability.
To navigate these constraints, businesses must also rethink their approach to system configuration and lifecycle management. Companies are now planning for environments that might require compromising on specific configurations or sourcing RAM independently in anticipation of future market shifts. Furthermore, there is a growing trend toward extending the service life of existing systems through rigorous patching, updating, and revising, rather than replacing them immediately when newer frontier CPUs and GPUs become available. This strategy allows organizations to utilize secondary markets for servers that are still highly effective for inference tasks, even as new training-grade hardware enters the market, thereby reducing the need to constantly chase the latest technology releases.
The traditional comfort of sticking with a specific brand ecosystem, such as being exclusively a Dell, Cisco, or HP shop, is fading in relevance because the reality of the market demands much greater flexibility. Organizations may soon find themselves onboarding ARM-based servers equipped with NVIDIA chips or other non-standard configurations that were previously unthinkable. While the prospect of managing diverse and changing hardware might seem daunting to some who wish to avoid ownership entirely, the alternative of renting hardware from service providers presents a different set of financial challenges. In a rental model, customers end up paying substantial markups multiple times over, which ultimately inflates their operational costs significantly compared to owning the infrastructure.
Ultimately, despite the initial fear associated with rising hardware prices and supply volatility, owning and controlling one's own hardware remains essential for effective budget management in the AI era. The costs incurred by service providers are inevitably passed down to the user through multipliers that affect inference expenses, cloud component costs, and virtually every other aspect of infrastructure spending. Whether trying to avoid ownership or simply managing a tight budget, organizations will find that hardware costs are unavoidable and will be transferred back to them regardless of the procurement method chosen. Therefore, the most prudent path forward involves preparing for a higher degree of flexibility in hardware onboarding while maintaining direct control over capital expenditures to mitigate these escalating financial pressures.
Read the full video transcript
Now, one of the biggest challenges these
days is getting AI gear. It is one of
the hardest practical challenges right
now. How are customers dealing with GPU
and RAM scarcity? Yeah, this is
something that's affecting the industry
as a whole because the cost of RAM and
storage in addition to the cost of GPUs
and CPUs is going up um to the extent
where what we're seeing people try to do
is lock in server purchases as soon as
possible so they can lock in prices and
get shipments on RAM um even today. And
so this is one of those places where you
you need to realize you are not going to
be able to have a preferred vendor. That
if you're used to buying from one one
vendor, you are not they might not be
able to supply you or their costs might
get prohibitive for you to continue to
stay on that on that platform of choice.
What we see customers doing is they are
being heterogeneous in design upfront.
So they're recognizing they're either
going to have to switch vendors and move
be able to intermix different different
vendors depending on their supply chain.
They're going to have to compromise on
how they configure the systems and then
have heterogeneous environments even if
they're sticking with one vendor or they
are sourcing RAM themselves or
potentially planning to add RAM later
assuming let's hope it it frees up in
the market. Um, and they are planning to
keep systems in service for longer,
which means patching and updating and
and uh revising them or looking into the
secondary market for these servers as
the frontier um CPUs and GPUs come off
market. They're very usable for
inference and they'll be available for
you from these training labs. So you
need to be looking very differently at
the hardware infrastructure that you
might have said, "Oh, I'm a Dell shop or
a Cisco shop or an HP shop and been, you
know, thinking that would save you." The
reality today is that you're going to
get what you get. It even might be ARM
servers with B, you know, based on
Nvidia chips that to do this work and
you need to be prepared for a higher
degree of flexibility in what type of
gear that you're onboarding. And if
you're thinking that's scary, I just
want to never own hardware again. The
challenge of this market is that renting
your hardware or getting it from a
service provider, you're going to be
paying that markup multiple times over.
And so, while it might be scary to pay
more for hardware, owning the hardware
and controlling the cost of that
hardware is absolutely essential because
all of these costs are being passed on
with multipliers from the service
providers. Um, and that is a very very
serious concern if you're trying to
manage your budget. It's going to show
up in your inferencing costs. It's going
to show up in your uh cloud and other
component costs. It's going to show up
in basically any any way you turn trying
to avoid having hardware. The hardware
costs are still going to get passed down
to you.