Submind YouTube summaries
Thumbnail for Move AI to the Edge Before Your Infrastructure Breaks | Dr. Robert Blumofe, Akamai

Move AI to the Edge Before Your Infrastructure Breaks | Dr. Robert Blumofe, Akamai

Watch on YouTube

Video summary

The transition from AI training to inference represents a fundamental shift in how artificial intelligence is deployed globally, drawing strong parallels to the early evolution of the World Wide Web and modern cybersecurity challenges. Just as the web initially relied on centralized hosting providers located in specific hubs like Ashburn and San Jose before facing scalability limits with video content, current AI infrastructure struggles under similar constraints when moving from low-bandwidth interactions to ubiquitous, high-demand applications. The historical lesson from the early internet era, where concerns about bandwidth and latency nearly caused the web to collapse, led to the development of Content Delivery Networks (CDNs) that utilized distributed systems and algorithms to deliver content from the edge of the network. This approach successfully transformed static websites into dynamic, interactive platforms by dramatically increasing available bandwidth and reducing latency, a strategy that proved equally critical for evolving cybersecurity defenses against sophisticated ransomware and DDoS attacks over the last decade. As AI applications evolve from simple chatbots where users type text and wait for responses to powerful AI agents with constant, real-time interactions, the nature of demand changes dramatically to require high bandwidth and low latency everywhere. The centralized infrastructure model that sufficed for training or basic inference is no longer viable because it cannot meet these new requirements without causing significant delays and performance bottlenecks. Much like the early web developers who realized that traversing large distances to a handful of central locations was insufficient for video streaming, today's AI systems face the same limitations if they remain confined to data centers far from end-users. The brute force approach of relying solely on massive centralized clusters fails when the demand becomes constant and ubiquitous, necessitating a rethinking of where computational power resides to ensure smooth user experiences. The solution lies in moving AI capabilities to the edge of the network, mirroring the architectural shifts that saved the web and strengthened global cybersecurity. By distributing intelligence closer to the point of use, organizations can overcome the latency issues that plague centralized models and handle the massive data throughput required by advanced AI agents. This shift is not merely a technical adjustment but a strategic imperative to prevent infrastructure breakdowns as AI becomes more integrated into daily life and critical operations. The success of the early internet and modern security protocols demonstrates that mathematical algorithms and distributed systems are far superior to raw computational power alone when dealing with global scale, suggesting that the future of AI depends on adopting this same edge-centric philosophy to sustain growth and performance. Ultimately, just as the web transitioned from static pages to dynamic interactions through network distribution, AI must follow a similar path to move from training phases to widespread inference deployment without compromising speed or reliability.
Read the full video transcript
If AI training is more or less like writing software, inference is like deploying it globally as as so people can use it. It requires totally different infrastructure. You have handled high latency sensitive use cases like live sports and global cybersecurity. How do the lessons from those challenges apply to AI infrastructure today? >> Yeah, that's a great point. And and I do see a lot of parallels to the early days of the web. Also, I think some of the changes that happened maybe a decade or so ago in cybersecurity. Um so, in many ways, you know, people say this all the time, history does repeat itself. Um and and it's kind of repeating itself for the umpteenth time here. So, um while there's obvious differences, there's a lot about what's happening now that I think does parallel what we saw in the early days of the web. You know, as the web was getting um popular, really transforming the internet, there were a lot of concerns that that that the web simply wouldn't scale to to meet to meet the demand. Um and in many ways those concerns probably were well-founded because you did have a situation where web applications were centralized. Now, we were pre-cloud, but we did have hosting providers. And arguably the hosting providers back then were even more centralized than today's hyperscalers. By and large, most of the infrastructure was, you know, in the US, for example, was heavily located in places like Ashburn and and San Jose. So, every time you used a web application, you had to traverse large uh distances into a handful of centralized locations. And while that might have been okay in the very early days where a website was a pretty static thing, just some text, maybe a few images, as you move into video, for example, and large demand for that video, you simply cannot meet the bandwidth requirements and the latency requirements through that centralized model. And that's really I think what led people to to, you know, jokingly say that the World Wide Web should be, you know, called the World Wide World Wide Wait. And and and people speculated that the web would simply collapse. And and really that concern was the beginning of Akamai, where, you know, Tom and Danny, the two founders, came forward with a better approach, math, algorithms, distributed systems, rather than brute force. And they showed that you can actually deliver websites and web applications from the edge of the internet, dramatically increasing the available bandwidth, dramatically lowering the latency. And really that's what made the web work. And that was a critical ingredient also as the web transitioned from these static sites to dynamic, where your communication is happening all the time. It's not just click on a link and wait for a response. You're constantly interacting with these web applications. CDNs made all all of that work. And a similar approach, really the same approach, is in many ways what enabled powerful cybersecurity defenses. Um cuz cybersecurity also went through a pretty strong transformation about 10 years ago, maybe a bit less, where you moved from, you know, our biggest concern being things like Anonymous to sophisticated ransomware. And the world of sophisticated attackers with with ransomware and powerful DDoS extortion attacks, the centralized approach just wouldn't work. And again, you have to borrow from this playbook of of math, distributed systems, algorithms. And and that worked. AI today, I think, is in a very similar um regime, where as you move from training to inference, as you move from fairly low bandwidth and and high latency types of interactions, like the chatbot, where you're simply typing some text, waiting for a response, typing some text, waiting for a response. You move from that into an AI-powered application or an AI agent, the nature of the demand just changes dramatically. It becomes ubiquitous. It's constant. It's high bandwidth. It's requires low latency. And the again, the brute force approach just isn't isn't going to work. You can't do this with with purely centralized infrastructure. The demand has moved to the edge, so the the infrastructure and capabilities of AI have to also move to the edge.