Preparing for the future of AI: The Changing Consumption Landscape and Combating AI threats
Watch on YouTubeVideo summary
The evolving landscape of artificial intelligence is fundamentally reshaping internet security and content consumption, presenting new challenges for website owners and digital infrastructure providers. As AI-driven automation accelerates attack volumes in real-time, the volume of threats has reached billions daily, forcing platforms like Cloudflare to evolve from simple caching networks into robust defense systems handling nearly 30% of global web traffic. The economic viability of traditional websites is now under severe threat because AI overviews make it exponentially harder for organic links to receive clicks, effectively treating vast amounts of online content as free resources that attackers exploit by ignoring the costs associated with server responses and error handling.
To counteract this shift and protect creators, new mechanisms are being implemented to ensure a fair marketplace where content access is properly compensated. Cloudflare has introduced features such as AI crawl control, which allows administrators to monitor scraping statistics, block unauthorized traffic, and generate detailed reports on attack paths, all while offering payment integration tools that enable site owners to charge for their network access. This approach leverages existing legal frameworks that allow public agencies and nonprofits to set prices for data usage, aiming to create scarcity that encourages AI companies to negotiate deals similar to those already struck with major publishers, thereby preventing a future where a few giant entities monopolize the internet and dictate user experiences without contributing to content costs.
The presentation also highlights a critical distinction in how current AI traffic is utilized, noting that 76% of it is dedicated to training models rather than human search or interaction, which underscores the need for standardized pricing structures across the industry. By fostering cooperation to establish these standards and utilizing tools like Cloudflare Radar for threat intelligence, the goal is to maintain an open internet that does not succumb to monopolization by a handful of corporations. Ultimately, the strategy involves ensuring that future platform functionalities remain accessible to all users without paywalls, reserving paid plans strictly for enhanced support services rather than restricting essential security features.
Read the full video transcript
Steve Carl Carlson. I am the uh
principal engineer for the state of
California for Cloudflare. Um my primary
responsibilities are uh the state
agencies and ca.gov. Um but I also
support higher education. That's one of
the reasons why I'm here is that a lot
of my higher education customers
encouraged me to come here. But um
essentially I'm an engineer. So uh
customerf facing and today I'm going to
be talking to you about the future of AI
um the changing landscape and then also
kind of how AI is is changing the way
that security
um needs to be kind of approached on
your websites and everything else like
that. Right. So um Randy is the our
sales rep. So there's some contact
information here if you want to take a
picture of that. Um, I can also send
this presentation out to anybody who
wants it. Um, if Yep.
>> Okay. If we post it on the website.
>> Yeah, sure. Absolutely. No problem. No
problem posting this. Um, no problem
sharing any of the content.
>> Great.
>> So, okay. So,
let's see here. What are we going to
cover today? So, um, why Cloudflare? You
know, what is Cloudflare? I always like
to start out uh by going through that
since a lot of people don't know
necessarily what we do. A lot of people
have a really good idea of what we do,
but I also find it really interesting
and curious to hear what other people
think of what we do. Um, what do we do
on a day-to-day basis? Kind of what are
we fighting out on the internet? Um, how
the consumption landscape is changing,
which I think really pertains to a lot
of you guys. So, you guys build websites
and support websites and how is the
world of the internet changing and what
do we see since we have such a unique
perspective on it. um what are we doing
about some of the things that are
happening and then I'll leave you with a
few uh kind of websites and things to go
to. We are known for being very
transparent. So we have a couple
websites that we publish uh that are
available for everyone to kind of look
at. Uh most of the time people are
unaware that it's there. So I'll follow
I'll follow up at the end with these
websites so you guys can go browse our
telemetry data, our threat intelligence
and everything else like that. It's open
to the public, right?
So, why Cloudflare? Like I said, I like
to ask this question to people. So, how
many of you are familiar with
Cloudflare? And I'm curious to hear what
you think we do.
>> Yeah, go ahead.
>> We're using firewall services and um
brain isn't working. The thing that
replaces capture, it's better than
capture.
>> Yeah. So, um, it's what we call our
manage challenge.
>> Yeah.
>> Yeah. Okay. Awesome. That's that's a
great answer.
>> Definitely DNS management.
>> DNS. Okay. FDNS full service DNS. Um,
anyone else?
>> No. Yeah.
>> Caching.
>> Caching. There you go. Caching. Um, I
think it's really interesting
to kind of give you guys an idea of
where we came from and why we started
because that answers kind of a lot of
questions about why we do what we do,
right? Um, originally our founders had
one single question and that was if
everybody's moving everything to the
cloud, who's moving a firewall to the
cloud? That was their first question,
right? So servers, storage, all that
stuff seems intuitive, right? But who's
moving actual firewalls to the cloud?
And is that even possible? And so they
made the decision that it was possible,
but in order to make that work, it had
to perform. That was the biggest issue
cuz every single person said, "Well, if
you throw a firewall out on the
internet, it's going to be really slow,
right? Because you're going to pump
traffic through it and there's going to
be latency and everything else like
that." So in order to get it to work,
they had to improve latency. They had to
improve uh performance. They had to use
caching. They had to build a massive
number of data centers. And the cool
thing about what happened was in our
process of building a completely
cloud-based firewall, we made
everybody's traffic faster. So we became
known as a CDN for caching, for captas,
for everything else like that. But
ultimately it was this singular goal and
that was to uh be a security company and
that is put a firewall out and
everything else just kind kind of came
along with it. Right? So we are known as
being the world's fastest DNS that
literally was just to make sure that
performance of the firewall is up to
snuff. Right? So, these are all kind of
auxiliary things that happen. And it's
kind of a uh one of the the cool things
about what we do is um that's kind of
how we develop products. We don't go out
and say to customers, hey, we think you
need this thing, right? We think that's
the worst thing that we could possibly
do. Uh what we do is we see what
customers what their traffic is, what
they're dealing with, and then we
provide visibility for that. And then
the second step always is that
visibility wants to turn they want to
turn that into how do we do something
about this right? So our products are
always very organically developed. Every
single product that we have came from
visibility like an improvement on our
algorithm. All of a sudden we saw all
this traffic we never saw before. Hey
customer what do you think about this?
Oh we think it's great but how do we
stop it? And then the product like comes
along with that right? So I feel like we
always come from the right place as far
as how we generate traffic and how we
protect people is all based on what
actual real requirements from customers
are. Right?
So what has this kind of led to? We've
blocked 190 billion daily threats. Um
95% of the world's internet's within 50
milliseconds of us. Um we have a 100%
uptime SLA and almost 30% of the web
travels across our network. Right. So
ultimately what we've become is this
massive network. Um and that's one of
the things that I like to impart on
people is uh we have over 600 data
centers. This massive network that we
ingest customer traffic onto, right? So
you transport across our network and
then you layer services on top of that
network. So the platform has become this
massive network, right? And that puts us
in a really really kind of meaningful
position, a unique situation, right?
Because our customers travel across our
network and layer services on top of it.
We have the kind of telemetry data, the
kind of information, etc. that no one
else has, right? Like just think to
yourself if you've got a Cisco firewall
like how are they getting any sort of
information about telemetry where things
are coming from where they're going to
etc. They don't. They have to buy it
from someone else or they have to
partner with someone else to get it.
Inherently we see it and so we have a
unique kind of position based on what we
set out to do. Right. Um how big are we?
I do like to throw this out here. Uh, I
know it's probably kind of hard to see,
but we are roughly twice the size of
Google and Amazon. Um, so essentially
everybody thinks that Google and Amazon
are the biggest networks out there or
the biggest providers out there, but uh,
we're actually almost twice the size.
And we usually are number one. We lose
number one occasionally depending on
where you're at, but generally we are
the fastest, largest network out there,
right? So not a lot of people know that
um but that's just because we do
everything in every data center and the
way we populate data centers we do about
35 a year. So we are literally rolling
them out as fast as we possibly can. Uh
the goal is to get um over 60% of the
internet uh by 2027 is kind of the goal,
right? So that's what we're shooting
for.
Uh so let me kind of tell you a little
bit about what we fight on a day-to-day
basis. So just in 2025 so far we've
blocked uh 20.5 million DOS attacks. Uh
that's an over 100% increase over the
whole year of 2024 and that's just 2025.
So um we also count DOS a lot
differently than a lot of other people
do because we have the telemetry data
that we have. we have the visibility to
identify whether or not these DOS
attacks are unique, right? So, we can
tell if something is truly original or
if something has just been slightly
modified in order to do kind of the same
thing but just look differently, right?
Um so, a lot of this uh from Q1025 also
spilled into Q2. Um this chart I like to
show people because this talk is about
AI. Um if you look at the unique attacks
that have happened um they don't seem to
be changing a ton but that's because we
are filtering out these differences
based on AI. If we did not and that's
what this last column is um this last
column
u in Q2 of 2024 this is kind of an old
slide but there were 17.2
attacks 7.2 million attacks but only 1.8
8 million unique ones, right? And so
what we're trying to do is we're trying
to filter out um what is unique and what
is um actually just a minor
modification. But this is 100% AI,
right? So the ability to morph these
attacks and change them is no longer
reliant on a human. It is reliant on,
you know, algorithms etc. know to modify
these attacks, change them slightly and
then relaunch them uh in a very
automated way, right? So everything's
automated now.
>> Our question is in the middle.
>> Yes, absolutely.
>> Okay. So I don't actually quite get
what's going on. So if you have a unique
deep DOS attack,
>> yeah,
>> then you've got 10,000 computers that
are attacking the destination.
>> No, I don't I don't talk I'm not talking
about a unique instance. I'm talking
about a unique methodology.
>> Okay. Okay. Fair. But
>> but okay. So let's then switch to a
methodology.
>> Yeah.
>> And so
>> what the heck is being very what's
what's very what's changing between
attacks?
>> So there
>> the order that the 10,000 hits.
>> Yeah. The order the the telemetry like
where it's coming from. Um like the the
different types of things it's going
after like is this a layer four attack?
Is this a layer three attack? Is this a
layer seven attack? like the variation
and the movement of um different then
also like I think we talked a little bit
earlier the the source of the attacks as
far as is this cloud infrastructure is
this IoT devices is this right they
often will switch from different
platforms of infrastructure using the
same methodology but then just switch
infrastructure platform sources right so
all these things another company would
register as unique attacks and kind of
report on, you know, those metrics. We
are we have the visibility because we
see end to end communication, right?
Because of our edge position, uh we're
able to see where it comes from and
where it's going to and that those
players are actually the same players
and they just have more and better
capabilities over time.
>> Okay.
>> Right. So,
this is a DOS attack. I realize the
slide metrics are a little bit
different. This one's in billion packets
and the other ones are terabits per
second. So I apologize for that. But um
this was March of last year, 4.8 billion
packets. Um and this was the largest in
history at that time in March of last
year, right? Um this is August of last
year. So between March
and August, right? 6.5 terabs per
second. This was the largest in history.
So mainly what I'm trying to point out
here is the time frame between these
attacks. Right? So from March to August
were the two largest attacks in history.
Right? What I want to point out now is
that in September was 7.3
terabs per second. Right? So went from
August 6.5
to September 7.3.
And then what happened like two weeks
later 11.5
>> and then today or actually technically
yesterday 22.2 too.
>> Okay, so the pace at which this is
happening is being drastically
accelerated by AI, right? So the ability
to scale resources in an automated way,
the ability to move them, the ability to
uh bring them up and tear them down, you
know, completely in an automated fashion
is there's more and more players with
more and more capability constantly
being added into the mix, right? And so
it's happening so fast we're not able
like we talked earlier we aren't able to
actually uh write do the writeups on
these and get these out to people to try
to tell them the 11.5 we still haven't
written up yet. Um we're still you know
pushing that through legal etc. 22.2 you
know I told the guys today like hey we
need to get the 11 tab one documented
and out to the public so that we can hit
this one and and start documenting it as
well. Um,
>> and yeah,
>> would you maybe it's doing
it for
>> what's
>> Oh, I don't
>> What does this mean?
>> Oh, it's just uh it's 22.2 terabits per
second. So, the left column is 22
uh is the terabits per second and then
the bottom row is the time, right? So,
>> so what this is showing is that it's it
was only 40 seconds long. um this this
DOS attack.
There's a lot to be said for that
though, right? So, why would an attack
like this only be 40 seconds long?
>> Because you killed it.
>> Well, we not not only that, we did kill
it, but why else would it only be 40
seconds long? Cuz it's cuz it's not the
re it's not the real attack.
>> It's a pro.
>> It's a probing,
>> right?
>> No, no, no. This is over 40,000
different IPs and
>> coordinated.
>> Yeah, coordinated. Yes, absolutely
coordinated.
>> Wow.
>> Absolutely coordinated.
>> Um, so what I also want to point out is
and I I I talked to several people
earlier too. I talked to you. I keep
pointing to you because we talked a lot
today.
>> Um, but what's the realistic expectation
of your ISP being able to block an
attack like this? Does anybody know what
kind of I'm intimately aware of the
hardware that sits behind ISPs. I
compete against them so I know kind of
what their structure is and how they
stack their servers and you know what
they pump through as far as their
scrubbing centers etc. So um what's the
theoretical limit of an ISP's and I'm
talking like AT&T because I know
intimately what AT&T uses. Anybody have
any idea how many terabits,
how many megabits?
One terab. and they've only ever
theoretically tested up to 500 megabits.
So, this is, you know, becoming uh a
reality that people
h can't rely on their ISP to take care
of this problem anymore.
um they're going to have to look at and
even CISA
without taking any sort of favoritism
towards vendors has said really clearly,
you know, if you're um not using some
sort of distributed networkbased
edgebased mitigation system to thwart
this, you'll never stop it. You won't be
able to, right? And so you're going to
see all the other vendors kind of
scramble to kind of match our approach
over time. Here's the problem. We spent
15 years building the world's fastest,
largest network. How long do you think
it's going to take them to build a
network of the same kind of scale and
capability? May not be 15 years. I mean,
we all know Microsoft has more money
than anyone, right? But it it's it's
going to take some time, right? So So
it's very very interesting. Um I wanted
to kind of go over our mission really
quick just to kind of move on. Um, our
mission publicly is to build a better
internet, right? So, what would the
internet be like if we had known it
would be used for what is being used
today? That's kind of what we want to
do. It needs to be fast, reliable,
secure, all built in from the network
side. So, we're essentially building a
parallel internet. That's what we really
want to do, right? But internally, like
I'll be honest with you, this is what we
say our mission is on the inside. Uh,
and I'm kind of exposing this to you
guys, right? We have to now be the best
at determining what is human and what is
not.
And every single product that we have,
every single thing that we do,
unfortunately, is going to be centered
around this one thing that we have to do
better than anyone else, right? And
that's just the reality of how things
are, right?
Um, so what's the message on this? We
have to fight AI with AI, right? There's
no way humans can do this. Hardware
can't handle it. It can't be scaled fast
enough. Humans can't make the changes
fast enough. It's just not possible. I
have customers on a daily basis that
come to me and say, "We have a script
that populates I bad IP addresses and
we're running that script." And I ask
them, "How's that going for you?" Just
honestly, like, do you get any sleep?
because it's literally impossible for
you to do that anymore. Um, as smart as
you can be, like it just isn't possible
without the telemetry data and the and
the and the visibility, right? Um, we've
spent the last 15 years building on the
back of the internet. The internet is
the reason we exist. So, we have to make
sure the internet survives and we want
to give back to it. That's really what
we're here for, right? And we have a
network that is designed for security
that is specifically built to present
prevent bad things from happening to
customers. And that also puts for AI
puts us in a really unique position and
why I'm going to kind of transfer over
to this um second topic, the consumption
internet, right? And how that is
changing, right? So the consumption
landscape is changing. Um, I'm kind of
getting through these slides pretty
quickly so that I can take on questions,
FYI, but you are free to ask questions
if you want to. I also am prepared to
kind of show you some of the actual
tools that we have that are available if
you want to see those. So, if you want
to see our AI analysis tools, if you
want to see our AI blocking tools, um, I
have those fired up so that we can kind
of take a look at those. I wanted to at
least be able to show you guys those,
too, if you wanted to see them, right?
>> Yep. For sure. So that's why I'm kind of
clipping along if you're but absolutely
raise your hand if you if you need to to
uh interrupt me. So I'm going to say a
lot of what I'm going to say here is
pretty publicly said by our founder
Matthew Prince. Um if you guys want to
go listen to any podcasts that he has,
he's very vocal about a lot of these
statements. Um he does qualify a lot of
these. I wouldn't be saying these things
if it wasn't something that, you know,
we already say pretty widely publicly,
right? So, you could see references to
all this stuff. Um,
everything wrong with the world is
Google's fault.
Okay, now let's kind of clarify that,
right? Like that's kind of a big
statement. Um, they were the first to
tell us that traffic is the deity that
we worship,
right? And that turned into
Google. Google turned into Facebook.
Whether whether this is literal or not,
right? Google beget Facebook. Um,
Instagram,
Tik Tok, right? Like Google started the
chain. They started the process. They
started the whole ball rolling, right?
U, this is all built into what we call
our attention economy. this whole thing
that um you need to click through links
in order to get a cortisol response,
right? This is all what was established
early on by Google, right?
Now, let's be clear, Google, I think we
can all agree, is a net good for the
world like in in what they did in
growing the internet, etc., right? So I
won't um give them no credit for you
know what they have done but the reality
is it all started by you know the things
that they that they put in place right
and what they put in place was 25 years
ago Google struck a deal where we can
copy all your traffic all your content
and we'll send you traffic right so that
was the deal we're allowed to have
everything that you have if we'll send
you traffic Right. So, they actually had
a meter on their web page. I don't know
if you guys remember this or not, which
would actually time how fast you left
Google cuz they were proud of the fact
Sergey would always brag about how like
fast you left Google. We were referring
you faster than anyone else off of this
page and that's how they built their
business. Right? But 10 years ago, that
drastically drastically changed. Right?
Now, they're trying to keep you. Now
you've got what used to be called the
featured snippet, which was the bar at
the top of the page, which was
suggestions on, you know, your searches,
etc. I don't know if you guys remember
that or not. Um, they have the sponsored
list, right? So, if you pay enough money
for Google ads, you'll get the sponsored
header of every single search page,
right? And then now, what do we have?
We've got this AI overview that is
happening now at the top of every page.
So, you no longer need to actually
scroll down past the AI overview to get
your answer. Um, well, in 10 years, just
as a reference, it is now 10 times
harder to get somebody to click on a
link after AI overview. And that's a
fact. So, we've been measuring this
since this has happened. And it's
actually that's the good news, right?
This is the good news that it's 10 times
harder. It's actually 750 times harder
to get somebody to click on a link in
OpenAI.
Okay. And it's 30,000 times harder to
get somebody to click on a website in
Anthropic.
Okay.
So,
this is getting even worse. Like this
30,000 times like I just looked at the
statistics yesterday. It's really more
like 38,000 times, but I'm just trying
to do some general numbers. That number
is climbing constantly because people
trust AI more and more and more. They're
using AI more. And so as they trust the
results of AI more, this is going to get
exponentially harder, right? So, so what
we're saying is we go
put in a prompt and we're going to get
an answer and there might be a little
link. So, what's the probability of
yours being the link at the end of the
answer?
>> Yep.
>> That's what we're
>> That's what we're looking at.
>> Yep.
>> And if you're not in the answer and
you're in the list below the answer,
your chances are
>> zero. There's zip, right? So,
so there's only three reasons why people
create websites, right?
vanity, fame, right? To become famous,
right? To sell something, an item or a
subscription. And that kind of leads to
this next one, and that is to get rich,
right, with ads or anything else.
>> Nonprofits, do you want to do education?
>> Okay, so yeah, my wife brought this up
to me yesterday. like she
>> she said she said, "Well, what about
public sector entities that have
information sharing and what about
nonprofits that want to share
information?" Um, yet they don't really
contribute to the economy though
necessarily. So, but I I totally get
what you're saying. You're absolutely
right. like there and I am a public
sector engineer. So yeah, a lot of my um
a lot of my customers have an
obligation, right, for public
information, public information act um
to kind of, you know, give this
information out, but they do pay a ton
of money to get that public information
out and it's super discouraging when,
you know, they don't get any recognition
for it or any, you know, compensation
for it or anything else like that.
>> The message is actually received.
>> Correct. there's no there's no feedback
loop at all, right?
So, if no one ever really gets to a
website, then what's the incentive,
right? These incentives kind of go away.
Uh it's kind of like a self AI is
creating its own self-fulfilling
prophecy because what's the what's the
fuel that fuels AI? It's actually people
going to websites, right? And so if
there's no websites to scrape, then
they're eating their own tail. Um,
if no one gets to a website, why make
it?
And if no one makes a website, why are
you guys here?
Right? So, so that's kind of why I
wanted to come here and talk with you
guys is like this is a big problem and
we see this as a huge problem for
everyone, right?
Do you have a question or
>> Well, a comment. One of the things that
I've been thinking about a lot is as an
organization has copyright material
that's being fed into the AI and put
into a snippets like is that really
transformative what you've just done
there in AI or is that just so there's a
whole set of copyright sheets that
really haven't beenated yet that are
also
>> Yeah, absolutely. Absolutely. Right. And
then you've got, you know, the Trump
administration saying it doesn't matter
if things are copyrighted, you shouldn't
have to pay for them, right? So,
>> doesn't matter what those say,
>> right? Right. So, and that's my point.
I'm going to get to that, right? Like it
doesn't really matter like what they
say, right? Like what we we drive that
decision, right? We make that decision.
We stop people from taking things when
we want to stop people from taking
things. And that's part of this
conversation is empowering you guys to
kind of come on board with us to fight
this fight, right? Because we make this
decision. Private industry makes this
decision. It's not governments that make
this decision. They can say whatever
they want. We don't have to follow like
what they do. And and let's be honest,
most well I would like to think most
government officials would rather have
the private like economy police itself,
right? I'm not I'm sure there's some
people that don't believe that, but um
so what are we doing about it? Right. Um
uh wait a minute, what about robots.ext,
right? Like isn't doesn't robot.ext like
keep everyone from taking everything
from websites, right? No, it doesn't.
Right? So um it's clunky. It's hard to
apply across your websites. You have to
be really diligent about how you apply
it. Um it's a guideline. It's not a
requirement, right? We see these guys. I
mean, I'll be honest with you, and
Matthew says this a lot. We see AI
companies using the same techniques that
Korean hackers use.
>> Like, literally, they are using the same
techniques to get content, right? If
they run into a robot.ext file, what do
they do? They scrape cache information
from search websites to get the data
anyway, right? So, the robot.ext means
nothing to them, right? They'll even
bounce off residential proxies so that
they don't have personal liability,
right? So when we see that, we, you
know, we take note of it. Um, there's no
penalty for not complying to robots.ext.
Um, there's too many ways around it.
We're trying to improve this. Actually,
if you go see this link right here, we
now have what we call a content signals
policy that we've developed. And what
we're proposing is additional fields
inside the robots.ext text that is that
are specific about what they can scrape
and what they cannot. Right? So, right
now, robots.ext is you can or you can't
based on a path or a site or whatever.
Right? We're trying to get a little bit
more granular in that by putting these
signal policies in the robots.ext.
Again, no one has to follow our
guideline or use those flags. We're just
throwing them out there uh hoping that
people pick up on them. And it also
doesn't mean that they have to listen to
them at all. Right. Um,
so it's like a speed limit, right? If
there's no one that's going to pull you
over and give you a ticket, it's
completely useless, right?
So, what are we doing about it? Um,
here's the reality. AI companies have to
pay for content.
Full stop, right? Um, but it has to be a
level playing field. And I think that's
the real message that we're getting from
the AI companies. They're not opposed to
this, right? But they but Sam isn't
going to be a sucker. He's not going to
stand over to the side and say, "I'm
going to pay, you know, you for I'm
going to pay you New York Times for your
content, but no one else has to pay for
your content." Right? So, so there has
to be a level playing field in this
whole thing or nothing like this ever
really takes off, right? Um, so we
called contents independent day, uh,
independence day on July 1st. We had a
huge press release for this. And what
this meant was um our AI crawlers are
the blocking mechanism for that was
turned on as a default for 100% of our
customers. Right now I had a bunch of
public sector and nonprofit companies
turn it off because they have an
obligation to you know they still don't
know how to deal with that. But for 30%
of the internet all of a sudden their
websites became completely unscrable
right and that that was done because
um well we also provided what we call an
AI call control dashboard so that got
released at the same time and you're
able to actually manage what you're
willing to allow and what you're not
willing to allow based on uh what you're
seeing in true analytics, right? So you
can actually see what they're doing. Um,
we actually resurrected or actually put
into use the 402 response code. So, I
don't know if people know what the 402
response code is. It's payments
required. Um, it was put in place as a
future use response code. So, no one
uses it yet and we're saying, "Yep,
we're going to start using it." So, the
402 response code, get used to it. We're
going to start um throwing that out
whenever there's a website that's been
blocked by um Cloudflare's AI crawling
tool, right? Um
here's what we realize. Is this the
right thing to do? We don't know. Like,
is this the way everybody should do it?
We don't know. But what we do know is
that everybody came screaming to us
saying that at 10,000 times they're
going out of business. And at 40,000
times they're dead, buried in
underground, right? And nobody's doing
anything about it. So we said, what we
have to do is we have to do something
about it. And the only way any markets
get created, the only way things like
this happen is if you create scarcity,
right? So you have to create scarcity or
no one listens. So, if there's one thing
that I can say that we did is that we
made a major step towards creating
scarcity. And we're hoping other people
follow us in this cuz we don't want to
be the only ones doing this. We're
hoping there's five other vendors that
do this. We're hoping that everybody
does this, right? That everybody jumps
on board as recognizing what not doing
anything about this is going to lead to,
right? It's going to lead to uh and I
have a slide on that too, actually. It's
coming up pretty soon.
Um,
we also just announced we're launching
the X42 Foundation, a nonprofit
foundation to perpetuate this. So, if
anybody wants to join the X42
uh foundation and work towards um kind
of standardizing like there's a lot of
questions that have to be vetted out
like what is the price per crawl like
what model is that going to be? Who's
going to set the standard? Right? So, we
can't do this alone. This has to be
developed by the web community, by the
internet community, right? Go ahead.
>> So when I put in the prompts and now
everything is agent based, so it sends
out a bunch of agents.
>> Y
>> or does it?
>> Yeah.
>> Is it actually doing real time scraping
or is it pre-scraped and it has an
index, let's say chatg.
>> Yeah.
>> With all the pages and it's pulling it
up from it. Is it scraping in real time?
It's actually hitting my website.
>> Yes.
Yep. And then here's the other part I
I'll add to that. Like agents are easy,
right? Because agents um are very very
easily identifiable, right? Um so we
think that's the easier part of the
equation. We think that agents are so
easy to spot and so easy to manage um
that we'll be able to go directly and
meaningfully back to, you know, anybody
that is using the agent and say like you
need to pay for doing this. And there
will probably be, and I'm just
spitballing here, but I think there's
going to be a different model for agents
because there's nobody trying to hide
that they're using an agent and there's
no deception involved in in that kind of
a process, right? So, we feel like
that's the easier side of the equation.
It's where they're not using really
clearly defined registered agents,
cryptog, you know, cryptographically
identified or otherwise, right? and
we're trying to figure that out, right?
So, the business side of this, like the
booking your travel or, you know, making
changes in your bank accounts and things
like that, we feel like that's the way
easier side of the equation. It's we
could police that and we can help people
police that, right? But it's everything
else that we don't know about. That's
that's the hard part, right? Um,
>> hang on, follow. So, so it's the other
side of the equation. Agents are one
side. The other side is when they're
training the model and they're just
scraping.
>> Yes. Learning.
>> And is there not there's nothing else.
It's really just those
>> Yeah, pretty much. Y and then real
people,
>> right?
>> Yeah. Um so what happens if this works?
Um we have some really good examples of
how this works, right? Whether you want
to directly correlate that or not, um
we've seen it play out. When there's a
healthy market and people are allowed to
participate in this market and it's a
level playing field, then those that
have the best content get paid for it
and they succeed, right? So, you know,
you pick which streaming provider you
pick based on the content, right? And
you're willing to pay whatever that cost
is uh to get that content because you
want that content, right? Like I don't
subscribe to HBO Max. Well, I'm going to
have to because I want to watch Peace
Peacemaker. Um but like when I want that
content, I'm willing to pay for it so
that I can get that content. So we do
have evidence that this absolutely
works, right? Um
small independent providers, we're going
to have those small independent
providers that are going to have really
awesome content, right? They should be
compensated, too. So, this can't be
about who has the most money, right?
Because if it were about who has the
most money, then we know who would win
and that would be, you know, tragic as
well, right?
Um, we have a lot of acknowledgement
that this is something that everybody
wants to do on both sides, right? We've
got Amazon just brokering this massive
deal with the New York Times. We have
open AAI brokering deal multiple deals
right with Wall Street Journal, New York
Times, everyone else. So these examples
are out there. We're willing to pay for
content, whatever that price is set at,
right? Like just tell us just tell us
what that price is and we're willing to
pay for it. They're not pushing back on
it.
>> It's a dead end. But New York Times is
both suing and collaborating with Open
AI.
>> Yes.
>> Yep.
>> Okay.
They're suing for the period of time
before they're collaborating.
>> Yeah.
>> Okay.
>> So,
copyright violations happened.
>> Yeah, it happened. Yeah. So, they have
to get money for that.
>> It's almost like a Disney movie. I don't
know.
>> Yep. And again, I'll repeat this.
Everyone's supportive of this. Just no
one wants to be the sucker, right? So,
no. You know, and and I'll be I'm going
to give credit where credit is due. Open
AAI, they are playing by the rules like
of all I'm going to show you guys some
tools where you can kind of look at some
telemetry data that we see. Um they're
they're identifying their unique agents
and their purposes. They're
cryptographically signing them. They're
following robots.ext. They're they're
doing everything right. Like OpenAI is
the good guy. They are doing everything
right. Enthropic is not
>> right. So like there are some AI
companies that are following the rules
and are willing to do the right thing,
right? Obviously, you know, any Chinese
AI not following the rules, right?
Deepseek.
Um, we actually warned them about 6
months ago that we were going to block
them entirely as a bad bot. We were just
going to classify them as a as a bad
bot, not as any meaningful data at all.
And they backed off. They backed off
really really heavily. They went from,
you know, behaving as poorly as
anthropic to barely doing any scraping
at all for a period of time. So, they've
slightly ratcheted ratcheted up now.
They're kind of coming coming out of the
cave, but for a while there, they just
stopped doing everything. Um, so it does
work, right?
Uh, what if we don't do this? So, here's
the black mirror version of what we
talked about, right? Um, it's not hard
to imagine. Um, this is not about
journalists, artists, news
organizations,
content creators going away, right?
That's not what we're looking at. It's
about everybody working for an AI
company, right? So, instead of, you
know, the the model that we have now
where, you know, if you're conservative,
you go watch Fox News. If you're, you
know, if you're not conservative, you
watch, you know, another station. If
you're in Europe, you watch BBC. If
you're in China, you watch China's
sources, right? Like there's that just
shifts. It's now just an AI company that
represents your same demographic, right?
So if that and we talked about this
earlier about, you know, as AI gets more
and more advanced. Um, is it going to
need a reseller to buy stuff for you?
No. It's going to place an order
directly to a supplier and it's going to
bypass a storefront and get that shipped
directly to your door, right? Is it
going to need a hosting company to host
a website? No. It's just going to write
its own platform and its own ability to
host websites and it's going to upload
that code to its own content and
provider. Right? So that's the Black
Mirror vision is that we've got five
giant AI companies that rule our our
lives. Essentially everything that we do
is dictated by what we see from these AI
companies,
right?
How different is that from
from, you know, conservatives watching
Fox News and how different really is?
>> Well, but I think we all agree that's
wrong as it sits, right? Like we have to
we have to do something about that too,
right? So I think if we just don't want
Well, I think AI has there's some limit
to how far that can go. like there's no
limit how far as as far as how AI can
go, right? AI will take that paradigm
and will shift it even further out,
right?
Uh and there's not a lot of caring about
the impacts of the growth of AI right
now. Everybody's in it for the money and
the market share, right? So, while
that's happening, you're not seeing a
whole lot of concern. And we talked
about this earlier too, like the
traditional people like Google that we
thought were fighting for the sustenance
of the internet are sadly no longer
really fighting uh for what's right
because they have a stake in the
progress of how things are going.
So it's really unfortunate. Uh, and you
know, at our last kickoff for the year
last year, Matthew said, "We may be one
of the only companies left standing that
has any power to fight against this
thing, but that's the right thing to do
and we should do it, right?"
Um, so again, we need healthy markets,
lots of buyers and sellers, lots of
them. This can't be about who has the
most money, right? This has to be a
robust and healthy marketplace or it's
not going to work. And the first thing
that the the biggest factor is again
scarcity, right? We have to create this
scarcity or no one's going to listen.
And so that's what we're doing.
So I wanted to show you some of these
tools. I mean, there's some marketing
stuff in here. Um, like I said earlier,
we are a network that sits in front of
your website. So we do filter traffic
through it and then apply services to
it. So we're an extension. We become
your edge. So your edge just gets
pushed. Think of it as a we are the uh
gated community around your
neighborhood, right? So if someone can't
get past your gates, they're never going
to get to your front door. That's the
whole concept, right? Um couple
statistics. I'm just leaving this for
the slide when we distribute it. Um when
a new exploit comes out takes 22 minutes
for it to be used in an attack. Um so
that's another testament towards uh
human intervention. Uh we will see live
attacks morphed from a new exploit in
less than 22 minutes. So um fishing is
still the number one initial attack
vector. So if you guys aren't watching
your email traffic um you're probably
wrong. And um
and uh we've seen a massive increase in
API traffic obviously um you know front
ends, mobile apps, everything else like
that. So that's going to become the
norm. Um 33% of our attacks that we see
now are API specific attacks against API
endpoints. So I wanted to show you this
Cloudflare radar. It's called
radar.cloudflare.com.
Um and I was going to share this for
you. So, Cloudflare Radar is our is an
exposure of all of our telemetry data to
the public. So, everything that we see,
you can see in here and you can look up
websites, you can do research on IP
addresses. You can leverage what large
corporations pay a lot of money for for
threat intelligence in this dashboard.
So, it's available to the public. You
guys can go do your own research. One of
the things that's really cool about this
is the AI insight page. Um, and we
talked about this a little bit earlier,
but this crawl purpose here, this is
what we were talking about before.
What's the reason why AI is touching
your website? Um, 76% of the time it's
for training.
>> So 76% of the time, 17% of the time it's
for search, 5.4% is actual real uh user
interaction, and that's it. So, if you
want to know like what that equates to,
um, real versus not real traffic to your
actual websites for anthropic, it'll hit
your site 25.6,000
times for every one real query you get.
That's the average, right? Uh, for
OpenAI, it's 601 to one. For Microsoft,
37.6 to1. um for Google 5.6 to1, right?
So that's again kind of feeding what
we're just talking about before is that
this is the kind of traffic we're seeing
to websites that is different than it
ever was before. Right.
>> Okay. I'm I'm kind of not getting what
that's actually telling.
>> Okay.
>> It's going to touch my website 36,000
times. Why is so much?
>> Because it it's it's all automated and
it's on a routine, right? And so it
there's routines that we see that just
run constantly. They're just constantly
combing across and we see I think I
mentioned earlier like akin to Korean
hackers, right? This isn't just looking
at your site map and like combing the
pages. They actually are using AI to
guess path structures,
guess your operating systems and your
and your content delivery systems, and
they're using wildcard searches to
actually try to find paths that you
haven't shown in your site maps. So that
effort that because they're use they're
brute forcing your website like is this
25.6,000 6,000 hits on your site.
They're trying to see where you made a
mistake and you left something open and
there's something they can grab that's
meaningful to them.
>> Just think about it. They can turn it
into Drupal safe. All they have to do is
node one, node two, node 3, node 4.
>> Yep.
>> And they will get
half a 404. We still pay for 404.
404s are expensive.
>> Yeah. Exactly right. 404s are very
expensive. And so this is just what this
is what they're all doing. And it's
because we've told them that it's free
and it's easy. That's why there's no
consequence for this, right? No one's
screaming back at them saying, "Hey, you
hit my site 25,000 times in the last 3
minutes, right? No one's saying that to
them."
>> Actually, I did. I I told API once they
were hitting our site and they actually
responded.
>> Oh, did they?
>> Well, you're but you're the you're the
exception. You're not the the norm,
right? Um but that's kind of why we we
did what we did. That's why we put this
um here. So, let me jump into my lap
here and I'll show you.
Well, the internet's a little bit shaky
here, I think.
Yeah, the Wi-Fi is not great,
>> but I did want to show you guys um and I
know I'm running out of time, but this
is kind of what I wanted to show you
guys.
>> So, are you saying that 25,000 times I
looked to see if I could find the local?
When I found one, that's the
>> No, no, no. Like, they're hitting
they're hitting everything constantly to
see whether it responds. Those are all
those are all requests that your servers
have to respond to. Every single one of
them.
>> So there's only one response 25,000.
>> No, there's 25,000 responses
>> for every one person that went there for
a real reason.
>> Oh, okay.
>> I'm sorry, guys. I don't know why this
is not coming up.
You know, one of the one of the things I
was doing earlier was um
>> yeah, I was um kind of firing up my
personal hotspot so that we could kind
of see
uh see if it shows up here.
Let me see if I got any sort of signal
here at all.
It's going to take a couple seconds
probably for it to touch up. Oh,
>> what's your extension?
>> Go.
Class player should come out to his own
cellular network.
So we we do have a very
eswish kind of offering now where
um we are going to let people use our uh
middle mile. So that's going to come out
pretty soon uh because our network is so
large now.
Let's see if this comes up here. But I
did want to show you guys.
Okay, here we go. So this is AI crawl
control. Um, this is what it looks like.
So this is available to everyone,
including free customers. This is
available to everyone. So that's the
other thing is that as a free customer,
all this stuff is available to you with
just limited user counts and limited
domain counts. But 100% we just made an
announcement yesterday. 100% of our
functionality is now going to be
available to all free users. We're no
longer going to gate functionality
around paid plans. So paid plans will be
about support. They won't be about
functionality. So all this is here. So
notice how I've got all the statistics
on every crawler and the ability to
allow or block those things based on
what I think is allowable traffic. And
then I'm also able to look at metrics,
uh, run reports on what they're doing,
see the paths that they're accessing and
what they're trying to scrape. Uh, and
then the best thing is this last one,
the 402 payment required setting. And
then let me kind of jump out to um this
lab. This lab actually has payment
setup.
So this I wanted to show you is our
first step into uh setting up payment.
We have a partnership with Stripe. So
essentially you can connect your Stripe
account and your bank account, set a
price and automatically will drop money
into your bank account.
>> Pay me. We have a a remarkable amount of
nonprofits and public sector companies
interested in this because law does say
you're able to recoup costs related to
infrastructure. It's legal to do that.
So
>> yep. So we do have a lot of public
agencies really interested in in this.
>> Okay. So you show data uh how often the
bots are scraping sites versus actual
humans.
>> Y
>> and I've already heard that websites
like Reddit are losing massive traffic.
>> Yep.
>> Because of the AI summaries. Do you have
any data on
how much traffic is being lost to the
average site of a certain size, let's
say?
>> Um
I know we have that data. Um, I just I
think it this this um Cloudflare Radar
is kind of our first kind of go at
exposing that data. I think it's not
quite there yet, but I think we do know
what that traffic is and we're going to
be publishing that. So
>> yeah, you can figure it out by a lot of
the data that's here.
[Music]
>> So anyway, um that's it. That's all I
had. Um but hope that was helpful and
interesting. Uh but you know, support
us.
Thanks for being customer.