Inference Is Everywhere. Your AI Infrastructure Is Not | Dr. Robert Blumofe, Akamai
Watch on YouTubeVideo summary
Dr. Robert Blumofe argues that the current strategy of heavily investing in centralized AI data centers is becoming an unsustainable model for the future of artificial intelligence. While this brute-force approach to building massive, dense GPU infrastructure made sense during the early phase focused on training large generative models, it fails to meet the scale requirements of the next stage known as ubiquitous AI. The core thesis is that even aside from the high costs, centralized locations cannot achieve the necessary distribution to support an AI ecosystem that permeates every aspect of daily life.
The demand for computing power has fundamentally shifted from training to inference, which is where the actual value of AI is realized. Training remains a mandatory upfront investment, but as the industry matures, the majority of infrastructure needs now stem from running models in real-time applications. Blumofe highlights a second critical shift from viewing AI as a specific destination, like visiting a chatbot website, to integrating it into ubiquitous applications and agents. In this new era, AI is no longer a separate tool users intentionally access but an embedded component of everything done online, from booking medical appointments to sending messages to children.
This ubiquity changes the nature of infrastructure requirements, making a centralized model inadequate for serving global demand efficiently. When AI becomes part of every digital interaction, relying on distant, massive data centers risks creating significant latency issues similar to the old "worldwide wait" or what Blumofe calls "large language molasses." To avoid these performance bottlenecks and ensure that AI remains responsive and useful in all contexts, the infrastructure model must evolve away from concentration toward a more distributed approach that aligns with the widespread and constant nature of modern AI usage.
Read the full video transcript
We're seeing massive investment in
centralized AI data center right now.
Why do you believe this centralized
everything approach is the wrong model
for the future of AI?
>> So, that's a great question. And
ultimately, the central thesis is simply
that the sort of brute-force approach of
building out large amounts of
infrastructure in centralized locations,
ultimately, well, it's expensive, but
ultimately, even that expense aside,
can't achieve the scale that's going to
be needed as AI sort of moves into its
next phase that you might characterize
as ubiquitous AI.
Um and I think it's worth maybe
highlighting a couple of ways in which
the demand
has changed. Cuz ultimately, you need to
look at the demand and see how the
infrastructure aligns to that that
demand.
Um and I would focus maybe on two
shifts. One would be the shift from uh
training to inference, and the other one
I would characterize as the shift from
um sort of the early days of a of a
chatbot to an AI application or an AI
agent. You know, it wasn't that long
ago, focusing on the first of those
shifts, it wasn't that long ago that um
most of the infrastructure demand really
came from the the training use case,
where you were and by and large, I'm
talking about the pre-training of large
generative models like like LLMs. That
was driving a huge amount of the
infrastructure demand, and in that use
case, absolutely, centralized,
large-scale, dense GPU um infrastructure
makes a whole lot of sense.
But as you move into inference,
it changes a lot. Um and of course, and
I think we all know this, that you know,
training is really a sort of a mandatory
cost that is necessary to realize the
value through inference. All the value
in AI comes from the inference and
training is simply
an investment that we have to make to
realize the return that you get through
through inference.
And now as we're moving into a more
mature phase, much more of the demand is
coming from inference and that's a good
thing because again, that's where we get
the value. So inference driving demand
is a very is a very good thing.
And I would argue that the nature of
inference is changing quite a bit.
And again, that's the shift I'm talking
about from the chatbot to the AI
application or the AI agent. You know,
in the case of the chatbot, I think we
really thought of AI as sort of a
destination. It was intentional. You
went to chat.openai.com
to use AI or you
fired up your your Anthropic Claude
desktop to use AI. It was intentional.
It was a destination.
Once you move into AI powered
applications and AI agents, it becomes
ubiquitous. It's no longer a specific
destination. It's just part of
everything that you do. Certainly
everything that you do online. You go to
a website to look for a car, AI. You go
to a healthcare provider to make an
appointment to see your doctor, AI.
Everything that you're doing is AI
powered. Probably even everything that
you're doing on your desktop, even
irrespective of the web. You know, you
want to
you want to send something to to your
kids, AI.
So AI becomes ubiquitous. That changes
the nature of the demand and in that
world where AI is ubiquitous, being used
all the time by everyone,
a centralized approach just really isn't
isn't going to cut it.
And and we risk sort of revisiting the
old, you know, back then we called it
the worldwide wait. It could turn into
large language molasses for lack of a of
a better term.