Saharsh Agarwal : The Impact of Global AI Overviews on Publisher Traffic User Experience
Watch on YouTubeVideo summary
Saharsh Agarwal conducted a comprehensive field experiment in collaboration with Ana Sen from Carnegie Mellon to evaluate the impact of Google AI Overviews on publisher traffic and user experience, utilizing a custom Chrome extension developed by Kiara Andre. The study randomized over 1,000 US-based users into three distinct groups: a control group using the standard interface, a "Hide AIO" group where AI Overviews were suppressed to prioritize organic results, and an exploratory "AI Mode" that forced conversational-only search interactions. Participants monitored their behavior for two weeks after establishing robust baseline browsing histories, revealing that when AI Overviews are triggered in approximately 41% of queries, they significantly reduce external clicks per search by about 40%, effectively cannibalizing traffic from publishers without generating sponsored click revenue.
The analysis highlighted a stark contrast between the control and intervention groups regarding zero-click searches; hiding the AI Overviews reduced this metric only slightly while actually increasing organic clicks to nearly double that of the standard group, suggesting users can still find answers through traditional links when forced away from the summary box. Conversely, forcing an exclusive reliance on conversational AI led to high user attrition and lower engagement rates compared to other conditions, indicating a strong resistance among searchers to fully automated interfaces as defaults. Despite these shifts in traffic distribution, self-reported surveys showed no perceived difference in experience quality between groups, leading Agarwal to suggest that the reduction in clicks may reflect decreased search costs rather than diminished utility, though this interpretation remains debated by industry experts regarding whether users are truly indifferent or simply less aware of their reduced interaction with external sites.
Beyond immediate traffic metrics, the study addresses broader implications for product discovery and consumer choice within a changing information environment where AI Overviews appear in only 12% of transactional searches despite triggering half the time overall. The research suggests that when an overview does not provide a definitive answer, search engines should explicitly justify why no conclusion was reached rather than simply displaying "no result," potentially preserving trust and aiding user decision-making regarding pricing and attributes. While the paper is praised for its credible causal inference and significant findings on traffic cannibalization—raising important antitrust concerns highlighted by recent European Commission investigations—the authors recommend softening certain interpretations of the results and leveraging available prompt data to further investigate how these AI interfaces specifically influence the informational and discovery phases of user behavior in future studies.
Read the full video transcript
Heat. Heat.
All right. Uh good morning, good evening
and good afternoon everybody. Uh I have
the pleasure to introduce Sahar Agarwal
um to our TSC seminar. He's going to
talk to us about the impact of Google AI
overviews on publisher traffic and user
experience. Dahar, you have 40 minutes.
Take it away.
>> Thank you so much Kiara. Thank you so
much everyone for joining. Thank you for
the organizers for having me here. It's
uh it's a great honor for me to be
presenting my work. Uh it's a paper I'm
really excited about. So uh it's a field
experiment talking about the impact of
AI overviews on publisher traffic and
user experience. And this is joint work
with Ana Sen from Carnegie Melon. So,
so we'll just get started and if you
have clarifying questions, maybe just
feel free to ask along the way. But for
more substantial questions, we can just
wait for the end. So, firstly, a project
of this nature requires tons of funding
and support. So, I'm very grateful to
ISB for funding this project through
couple of different channels and also to
my RAIT who has been phenomenal support
throughout this project. So, now the
topic of this study is Google's AI
overviews. uh to this audience I don't
really need to introduce what EIO views
are but again just for the sake of
completeness so like when you put a
query on Google u you you start Google
is now showing you like an AI summary uh
which in most cases tends to answer the
question that you're asking and so u you
see you see a response you see a bunch
of different links and uh these AI
overviews have been a little
controversial Because if you think about
the nature of the search market, it is
very symbiotic. Like Google is not
really producing content and there are
tons of content providers who exist uh
and like whose ex whose existence
depends on uh discovery on Google search
and so the the the relationship is very
symbiotic because Google provides
visibility to these content producers.
And so therefore uh when when users use
Google, Google makes money out of ad
revenue and these publishers also get
traffic from Google. So therefore it's a
very symbiotic sort of a relationship.
However, uh there has been a shift in
this relationship in in the last few
years and u so there are increasing
concerns around value capture uh in the
search market where although producers
are taking effort and incurring the cost
to produce content the the concern is
that search engines might be
internalizing the benefit of the content
production and they're sort of
externalizing the cost onto the
producers. So we are seeing a bit of an
increase in uh the phenomena of
zeroclick searches which is that a user
makes a search and uh instead of
referring the user to the downstream
content producer search engines are
answering those questions on the page
itself because of which I'm I don't
really need to click and so that's the
that's the the question that's the
tension here and uh clearly there is an
antitrust angle involved because of u
the value capture concern and so uh as
you can imagine there's a lot that's
been written about it a lot of
observation studies trying to look at
okay the world before AIO was world
without a views queries where AIOS are
shown versus queries where AIOS are not
shown however uh we don't really have
clear causal relationship yet and that's
what we we're trying to answer and of
course this issue is at the core of a
lot of antitrust concerns very recently
the European Commission launched
antitrust investigation against Google
for the very same reason that we're
talking about. And so this is the kind
of policy discussion that we hope to
contribute to through this work. And so
in this paper uh we are providing
empirical evidence on this question
using a randomized field experiment. And
so uh we developed a custom Chrome
browser and uh so this browser was built
on top of webmunk by Kiara Andre uh and
so it's it's an open source tool that
they made that they made available. So
really thankful for that. And so uh what
the extension does is two things mainly.
So uh the extension allows us to log
passive tracking data at great
granularity. And so we can we can track
every single URL visited. We can track
every single click made. We can track
every single scroll. And in fact, we can
we can track a lot of detailed telemetry
data like tabs which is pressing the
back button, closing the browser window
and all of that. And so so the level of
granularity that at which we're able to
track these users is is very phenomenal.
And beyond just tracking users, we can
also passively manipulate the interface
of any website that we like.
And so what we did was uh I I'll come
back to the study flow in a bit. But so
we had the control interface where we
just collected data without really
changing anything and 36% of our users
were assigned to the control group and
then we had the main treatment
intervention the second one which is the
hide AI overview group I'll call it HIO
in short. And so what we did here was
that the AI overviews were hidden and
everything else was pushed up in order
to make the interface very seamless. And
uh the third condition was a very it was
an exploratory condition the AI mode
condition. And how it how it worked was
that when when users entered any query
even on their search bar the the the
browser address bar or even if they go
to google.com and if they type any query
they would be automatically redirected
to AI mode and they could not have come
out of it and as you could as you could
expect like we we forced people to use
AI mode and they could not have come out
of it and so we expected a lot of
attrition in this arm and anticipating
that we we kept a small share of the
split to AI mode because we wanted this
uh the results from this arm to be
fairly exploratory.
So these are the
>> Yes.
>> Sorry, quick clarifying question. When
you say that you can't get out of AI
mode, does it mean that if you they
click on all images etc that does not
redirect them to the other format?
>> Yeah. So actually the the example I'm
showing you is not a great example
because um so for the study participants
all of these were actually hidden like
all images video etc in the AI mode. So
they could not
>> for all of them for all of the treatment
groups.
>> No no no in in the AI mode.
>> I see. Okay.
>> Right. Right. So, so we didn't really
want them to come out of the EM mode,
but however, in in the other two groups,
they could see the other buttons. Yeah,
these screenshots are sort of taken post
facto.
Uh so now in terms of the study flow, so
we we used prolific to recruit
people who would be willing to install
an extension. And so uh we first had a
pre-screening survey. And in the
pre-screening survey, we uh we asked
people uh three questions. The first was
we asked them what the default search
engine is on the system that they're
using. If they responded anything other
than Google, they were uh screened out.
Secondly, we asked them what is the what
browser do they use on the current
system. If they said anything other than
Chrome and even if they said like two
different browsers, which is Chrome and
uh Firefox, they were screened out. So
they had to use Chrome and only Chrome.
So we were trying to minimize users
switching to a different browser. And
thirdly, we asked them what all do they
use their current desktop for. And if
people said that they use it only for
like taking surveys because because
anecdotally we've seen that a lot of
these prolific workers, they have a
dedicated laptop just for taking
surveys. And so we wanted to screen out
those kind of people because they
wouldn't be giving us any useful search
data. Of course, this is all
self-reported. uh but they didn't know
what the expected answer was. So we can
just trust that they're not really
trying to please us with their answers.
And so uh if they passed that check then
they were directed to a page where they
could install the extension. And uh so
we we had given a passcode on the
prolific page where which they could use
to install the extension and uh after
that we of course this baseline should
have come before the extension
installation. So we had a very brief
baseline survey two questions we asked
them. The first question was about their
trust in AI generated information coming
from these AI summaries. Secondly, we
asked them what com comfort with using
AI tools for information seeking. That's
it. And these are mainly for just
checking hetroenity.
So after that uh the extension
automatically performed a history check.
The reason we did this was that we only
wanted to keep people who would give us
reasonable search data in the period.
And so there were a couple of things we
checked for like they should have
reasonable amount of u browsing activity
in the last 20 days. So like at least 10
days we have some browsing data from
them and we have at least 10 uh Google
searches for them in the last 20 days.
And also we have at least 25 unique
domains. These are very small numbers
just to make sure that we don't have
people who don't really use their laptop
for search. And so if they pass that the
extension uh sort of live and uh
automatically randomizes them into one
of the three groups the hide group which
I just showed you, the control group and
the AI mode group. So the split now this
is uh like we are randomizing one at a
time like we didn't really have the
luxury to wait for all the participants
to enroll and then randomize at we had
to randomize in real time and just sort
of hope that the balance checks check
out. Unfortunately, it did. I'll show
you that. And so, once people installed
the extension, uh we we observed them
for 2 weeks. So, that's what we
promised. But participants were told
that after 2 weeks, they can uninstall
the extension. And so, at the end of 2
weeks, the bonus was paid to them. And
after the bonus was paid to them, an
endline survey was administered. And
this part was important. So, we paid the
bonus first so that they could respond
truthfully in the endline survey. And uh
and yeah, of course, people are told
that in the event that they do not
choose to install after 2 weeks, we
would still be possible to track some of
the data.
So uh so yeah, we we ended up recruiting
1065 people and uh we we checked for
balance uh across basic demographics and
trust and AI information. This is the
baseline that we got in the survey.
Trust and comfort DI tools and the three
different the last three measures prior
Google visits, prior active browsing
days and prior unique domains. This we
got from the extension itself when we
were doing their browser history checks.
So we have uh we have balance between
all the three groups. So, so of course
this is balance at baseline right and
now balance at baseline uh is great but
however because it's a longitudinal
study uh we obviously expect some kind
of attrition along the two weeks and so
while balance at baseline is good we
want to make sure that we don't have
differential attrition across the twoe
period and so the first thing that we
actually track is part of
>> sorry clarifying question where are
these people from I should probably know
from the latitude and longitude but I
don't
>> yeah these are all US
>> US okay thank you
>> okay so uh so yeah we have balance at
baseline which is great so the next
thing to look at is survival over time
and do we have differential attrition
now our intervention was intended to be
very light touch like all that we
removed was the AI overviews and
basically nothing else changed and uh if
it was truly light you would expect
people not to really notice much not to
make a big fuss out of it and and not uh
not quit. And fortunately, that's what
we see. And so day the 15th day is when
the bonuses were actually paid and the
the twoe period was completed. And uh we
have actually pretty good retention if
you look at the main treatment and the
control groups. So we have over 90%
retention in both of these groups. And
the good thing is that although there is
some attrition, uh we don't see
differential attrition. The attrition is
pretty much similar. However, AI mode is
very different. Like it's not very hard
to expect. Uh we we we sort of forced
people to use AI mode for the two weeks.
They could not have come out of it. Like
even for a very simple simple search,
they get redirected to AI mode. And so
uh 60% is a retention that we saw in the
AI mode. And this is something that we
anticipated as well. So given the higher
attrition in the AI mode, we'll treat
results from AI mode as more exploratory
and our main results are going to be
focusing on the the show and the hide
overview groups.
And of course uh once the 15-day period
was over, we administered the uh the the
payment the endline survey and among the
people who chose to remain and uh we we
we switched the treatment status. So
treatment became control and control
became treatment. And so the reason we
did that is that while it's great to
have between user comparison, uh the
switch lets us have within user
comparison before and after the switch
and in fact we can also use user fixed
effects so that we can be totally sure
about our results. So I'll talk about
that as well.
Okay. So uh
so we start by first documenting the
prevalence of AI overviews and given how
much is being discussed about this topic
we realized that people felt that this
itself was uh an important contribution
just just knowing how much of AI views
is actually out there and so u
if you look so we observe a total of 16
68,000 unique searches in our data uh
over the twoe period on average we see
that AI overviews were triggered in
about 41% of all queries for uh and the
good thing is that this is balanced
across the treatment in the control
groups. So it's really not as if people
in the control group are putting in more
queries which are likely to return any
an AI overview. We don't really see that
happening here. And so uh most of our
analysis is going to be following this
particular partitioning which is that
there are queries which triggers an AIO
view and that is based on some uh
algorithm that Google uses and then
there are some queries which do not. And
now if you think about our extension it
does basically nothing for these 60% of
queries where the AI movies are not
triggered. So there's absolutely nothing
that that our extension does other than
just track data. And so the only cell
which is affected by the intervention is
this. And so this is the cell where we
actually expect to find effects. And
we'll treat the AIO not triggered group
as a placebo.
So that's uh and of course we'll just do
a very simple comparison of means uh
because we have successful randomization
and we'll be and this data is going to
be at the search level and so we'll be
clustering all the standard errors at
the user level.
So now talking about the main results.
Uh so I'll first be talking about
results at the search level. So like
conditional on a search happening. Uh
how many clicks do we see per search?
What is the likelihood of that search
being a zero click search? Uh what is
the number of sponsored clicks that we
see? So all of that analysis is going to
be conditional on a search happening.
However, there is there has been some
discussion that maybe AI overviews is
leading people to search a lot more. And
so if that is the case, you might you
might expect a shift in the extensive
margin itself. And so after showing you
results at the search level, I'll come
back and show you results at the at the
overall user level. So at the search
level, so this data is at the search
level. The first variable I'm tracking
is a total number of external clicks per
search. And so what we see is that in
the control group, we see 0.37 clicks
per search. By the way, these are
organic clicks. I'm not uh sponsored
clicks is is on is is on the third
panel. And these are organic clicks
basically coming from anywhere on the
page. Okay? So they could be coming from
the AIO view, they could be coming from
the knowledge gap elements, they could
be coming from the search links, but
anywhere on the search page. And so we
see that when we hide the AIO views the
clicks per search goes up by a really
large number which is 0.25. And so if I
keep the hide group as the base like the
status score before AI views were
launched. So uh the reduction in clicks
as a result of AIO views is about 40%.
And similarly
the probability of a zeroclick search in
the presence of EIO views is 0.7 0.73.
So what that basically means is that if
people put in 100 searches uh and then
when an AI view is triggered 73 of those
searches do not lead to any click but
when you hide the AI views that number
drops from 73 to 54. And when it comes
to sponsored clicks we don't really see
an effect. Of course these numbers are
really very small. So if an effect at
all exists like it would require
significantly more power to detect it
but at least with the power that we had
we don't see any effect on sponsor
clicks. These are only uh conditioning
on queries where the AIOs are triggered.
So the next thing I'm going to do is I'm
going to look at this half where the AI
ovs are not triggered and we we
shouldn't really see any difference in
the treatment and the control groups
and that's what we see. So when we when
we condition on queries where the AIOS
are not triggered for external clicks we
don't see any difference for likelihood
of a zero click search again we do not
see any difference and same for
sponsored clicks. Although uh one thing
to notice here is that the number of
sponsored clicks here are almost three
times the number of sponsored clicks
here. Although we don't see any any
difference and so it sort of suggests
that Google might be strategically
choosing to deploy AIO views in some
kind of queries and not the others but
we don't really see effects here and
this is reassuring because uh if our
placebo I mean if if our extension is
sort of working as it should we we don't
really expect to see effects in this
particular arm.
So these are all results from the search
level and these are all between user
comparison. I've compared the the
treatment and the height group. Sorry,
the treatment group and the control
group. So the next thing I'm going to do
is I'm going to show you results from
the the within user comparison. So once
the twoe period was over, we flipped the
treatment status for the treatment and
the control users. So so those who had
AI overviews hidden became control and
those who had AI overviews being shown
became the hype. the height group and
now we compare these users
pre and post switch. So the panel on top
are those users who were initially in
the hide group and the row in in the
bottom are those people who were
initially in the control group. So if
you look at those who were initially in
the height group before the switch they
had higher clicks and after the switch
this went down and we see a nearly
symmetric switch for people in the
control group and same for zero click
search. So now this is purely a within
user comparison and we have user fixed
effects here. That's why these
difference numbers are not really
aligned with the arithmetic difference
of the difference in the in the bars
because we have user fixed effects when
computing the differences. And so now
what this is telling me is that uh even
if for for whatever reason our treatment
and control groups were not very
comparable at least uh if we look at the
same user across different conditions we
see exactly the same effects. And this
is very reassuring that the results that
we are finding are are very causal.
Okay, so this is what I have for uh
search level results. The next concern
of course is that uh this entire
analysis is conditional on search. But
what if AI overviews is leading people
to search more for instance and so it's
possible that clicks per search are
going down. But if searches are really
increasing a lot more, then it's not
clear what the effect on total clicks
will be. So that's what we're looking at
here.
So now uh this data is at the user day
level. We also have user level
regressions which have very similar
results. So this data is at the user day
level and we are looking at three
different variables. The first is total
number of searches made by by a user
day. Second is total number of clicks
coming from AI overview triggered
queries. Third is clicks coming from AIO
view not triggered queries. And here we
see that there is no effect on number of
searches. So across the treatment and
the control groups, we don't see people
making less searches because we've
hidden the AI views. And uh at the user
day level, we see an increase uh of
clicks by a significant proportion
relative to the control relative to the
baseline and we don't see effects for
queries where the EIOBS are not
triggered. So, so basically it's not as
if people are searching a lot more at
least in the two week period. Uh in in
the longer term things may be a little
different but in the two week period
people not searching any less or more
just because we hit the AIO views. So
now one narrative which has been going
on is that uh okay like AI overviews are
cutting clicks but but these are all low
value clicks that the AI views are
cutting because there there are people
who land on your page and then okay who
realize that okay this is not what I was
looking for and who who quit very early
and so we we can examine this in the
data. So there are three different
metrics that we look at in terms of the
quality of the click. So now this
observation is at the click level. We
observe a set of searches from the from
the main treatment group. We observe a
set of searches from the control group.
And we are
we comparing these clicks uh on three
different metrics. The first thing that
we're looking at for for the click is
the likelihood that uh after the click
the user press the back button to go
back to Google search. Now uh if you
notice like we have slightly lesser
observations for this one compared to
these two that's because uh so if you
think about a us a user using Google
search uh they can either make a left
click or they can make a like a right
click and open a link in a new tab. We
are able to track all of that with our
extension. Now if you open a link in in
the same tab you can press the back
button and go back to Google search but
if you open it in a new tab uh there is
no back button there. And so so we we
are tracking the back button search only
for same tab clicks which is why we have
slightly lesser observations. We see
that and this number was significantly
more than we had anticipated. 40% of all
clicks coming from Google they they lead
to the user going back to Google
potentially indicating an unresolved
search and so but we this number is not
different at all in the treatment and
the control groups. The second measure
that we're tracking is bounds. Now
bounce is a measure which has no
standard definition like many different
people track many different definitions
of bounce. So what we did is uh we first
computed the time the total time that
people spent on the click page and uh of
course the caveat here is that time
spent on a page is a very noisy measure
to capture like despite the best of
telemetry data like there are always
approximations involved in computing
time spent. However, uh when when you're
looking at short durations, it's it's
much more accurate. When you're looking
at longer durations, it's not clear as
if someone was really looking at the
screen or or if they were uh somewhere
away. So, we define bounce as uh a click
session which lasts less than 10 seconds
and where the user did not make any
navigation inside the click page. So,
both of these conditions have to be
true. uh I may have I may have left that
website within let's say 8 seconds but
if I if I made a click anywhere inside
that page then that is not counted as a
bounce. So bounce is uh duration less
than 10 seconds and not engaging with
the with the page that you've clicked
on. And again we see that about 18% of
all clicks were lowquality bounce clicks
where the user bounced out without doing
much in less than 10 seconds. However,
this is not different in the treatment
and the control groups. And then we also
looked at overall duration and uh that
is also not different. And so this is
not very consistent with the current
narrative that AI overviews primarily
eliminate lowquality clicks. We don't
really see any difference here.
So this is in terms of the quality of
clicks and then I told you that we had a
third arm which was the EI mode arm and
uh as again just reminding you that
given the different given the
differential attrition we won't make a
strongly causal claim about this arm.
However, it's just uh it's just a a nice
add-on to have and so we we don't have a
split here of queries where AI OBS were
triggered and where they were not
triggered because we don't have an
equivalent of that in the AI mode. And
so we are uh looking at all the queries
in the control group together without
the AIO view trigger not trigger split.
Similarly, we're looking at all the
queries in the treatment group and then
we're looking at all the queries in the
AI mode group. Now uh defining a query
in the AI mode is a lot more complicated
because it's conversational. So like
someone puts in a query and then they
could be putting a follow-up query. Now
the follow-up query could just be a
refinement, could be a follow-up
question, it could be a fundamentally
different query. And so we're taking a
very simple approach which is that we're
taking the first query which starts the
conversation uh as a query in the AI
mode and then uh any follow-up question
which happens to that question. So any
click happens anywhere all of them gets
attributed to the first query.
Okay. And so uh so we'll be overstating
uh clicks per search because it's right
now what we're saying is that if someone
puts in a query on on Google AI mode and
then it's possible that after two
follow-ups they put a third query which
is fundamentally different and now in in
the entire session let's say that they
make only two clicks and so we'll be
saying that they made two clicks for the
original query which started the
session. However, uh it's actually two
clicks for two different queries. So,
but we are collapsing it all into one
query. So, it's a simplification which
is sort of overstating the clicks per
search. But we see that
the number of clicks in the AI mode was
actually lower than even the number of
clicks in the control group which was
anyways lower than the height group. And
correspondingly the the probability of a
zero click search was much higher in the
AI mode group compared to the control
and the hide. And so uh so yeah this is
very indicative like as search goes more
and more towards conversational AI
there is a concern about okay how is the
future going to look like in terms of
publisher traffic etc. And and we we
hope that this this serves as very
indicative evidence of where we might be
headed.
So now I'm going to show you some
results on hetrogenity. So the first
aspect of hetroenity which we looked at
is the position of the AI overview. Now
Google chooses a multi Google makes
multiple choices. The first is whether
to show an AI overview or not. The
second choice that Google makes is where
to show it. Right? So uh I'm sure you
might have some experience of it. Most
of the times the AIO view appears right
at the top of the page which I'm calling
position zero. Uh so about 88% of all
queries on which AIOS are shown. AIO
view is like the first thing that you
see on the page and then of course there
are uh there I'm defining these
different relative positions. So imagine
that you have you have an EIO at the top
and then you have three search results
below that. So the EIO view would have a
position zero and then the other search
results would have position 1 2 and
three. There are times when the EIO view
appears below search results like
embedded between two different search
results. So that's captured by these
different relative positions. But it's a
very skewed split. Most of the times AI
views appear right at the top. And given
the skew here, we are basically defining
two splits. One is that all queries for
which AI overviews appear at the top
versus queries where they appear not at
the top and uh these are results for
external clicks. So we see that pretty
much our entire result is driven by
clicks where AI views appear right at
the top and for for for queries where
EIO views do not appear at the top we
see no effects. Now of course u this
should not be treated as a causal effect
of position because position here is not
randomized. So queries which appear in
position zero might be very different
from those appearing in a lower
position. However, this is very
indicative and I would say that this is
even u this suggests that that we need
more work into understanding
what is it about AI overviews that is
really leading to such strong effects.
Is it really that that they give
information that people value or is it
just that they're occupying prime real
estate and and uh and they're sort of
cannibalizing the organic links as a
result. So I think that's something
which should be looked at in future
research. We also look at effects by
query type. So we classify the queries
that users put in into three different
types. This is very standard uh
classification like for those who work
on Google search data. Informational,
navigational and transactional.
Informational uh queries are those where
people are looking for some kind of
information online. Navigational is
where they're looking to navigate to a
particular website like for instance
someone just puts a search for YouTube.
So it's very clear that they're just
trying to navigate to YouTube or and the
third is transactional where uh they're
looking to either make a purchase or to
like download a software. So where
they're actually looking to like
complete an action. Uh so that's what
transaction is about. And so uh we see
that informational queries are the
overwhelming uh number in our data. 71%
of all searches are informational
followed by navigational at 18% and
transactional at 12 and uh yeah so
Google is basically strategically
showing AI overviews where it thinks
it's more valuable and so AI views
appear 53% of the times informational
queries 6% in navigational which is
reasonable like when I know where I want
to go uh and AI is not very useful
potentially and transactional is 15%.
Just a short note like in the paper uh
in the draft which was shared these two
numbers were sort of flipped. There was
a typo. Navigational was mentioned as 15
and transactional is six. So so just
take note of that.
Uh so yeah these are effects by query
type and and yeah of course what we see
is that most of the effects come from
queries. uh although I would say that
it's it's hard to rule out absence like
it's hard to rule out effects in
navigational and transaction queries
because we don't have enough power as I
said only 6 and 15% of all queries
actually return AIO views. Uh the
direction is there we just don't have
enough power to detect effects. So next
we also look at effects by topic. So
query type are these are like very very
broad classifications.
we look at topic which is like more
granular. So I have plotted two things
in this in this chart. So of course many
different uh topics that the queries
have been classified into. So the length
of the bars shows the prevalence of AIO
views for that particular group and the
shade of the bars shows the distribution
like the the share of all searches in in
the data which belongs to that
particular topic. So for instance we see
that arts and entertainment had the
highest share of searches in the data
but the highest prevalence of AIO views
was for health and science. So uh so
health, science, law, government like
these are moreformational focused uh
topics which which see significantly
higher prevalence of AIO views and news,
shopping, sports like events uh topics
which tend to be more current, more
live. I guess these have uh lower
prevalence of either
try to locate effects across all of
these. Uh so we we see significant
effects in all the topics but given that
we are splitting the data a lot it's
hard to compare across the different
categories. So we see effects in all of
them like AIO views are basically uh
reducing
uh clicks in every single one of these.
However, it's really hard to compare one
topic with the other because we don't
have enough power to do that.
And uh so next we uh after the twoe
period, we del we delivered a bonus
payment to the users and then we invited
them for an endline survey and uh we had
90% response rate to the endline survey
which was great. Now in the in the
endline sample I'm only uh restricting
this to people who had actually uh kept
the extension in installed for the twoe
period. If they had uh quit in between
they're not part of the endline survey
because we want to make sure that we are
only uh measuring people who actually
took part in the in in the study
properly. And so uh three things that we
tracked uh we asked them how was their
overall ex search experience over the
last two weeks. And by the way, till now
people have no idea that this study is
about Google search. They just think
it's a general web browsing kind of
study. We asked them to rate the quality
of information that they found on
Google. And we asked them uh to rate the
ease of uh finding information on Google
in the last 2 weeks. And uh the the
headline result here is that at least u
with with the obvious limitations of a
self-reported measure, we don't see any
effect at all in the three uh in the
three measures between the main
treatment and the control groups.
Although we hid AIO views for 2 weeks,
most people did not reduce the number of
searches. They did not think that the
search uh experience was affected. They
didn't feel that the quality of
information was affected. they didn't
think that information became harder to
find and in fact most people didn't even
notice what we had done and so that says
something about uh about how much users
might be valuing the AIO views and so we
we see a bit of a wedge here that in
terms of user experience I I mean I'm
sure that there is an entire
distribution and so there are people who
would really value the AI overviews
however when you look at the average uh
at least on the user experience side AI
overviews don't seem to be doing much.
However, on the publisher side, the cost
is really high. And so, that's something
to think about, which is that all the
lost clicks are not they don't seem to
be improving the search experience. And
so, so it's coming at a at a pretty
significant cost. So, uh to summarize,
uh we are showing we are showing some
very solid causal evidence on a topic
which is very policy relevant. We are
saying that AIO views are having a
significant impact for publishers and
wherever they appear AI views on average
reduce clicks by 40%. And by the way the
these numbers are expected to rise
because what we what we are showing is
uh clicks per search and so uh you can
expect that with time the prevalence of
AIOBS will increase and that has been
the trend like when when we when we ran
the study we saw a prevalence of 40%
when we were doing some pilots which was
in November of 2025 the prevalence of AI
overuse was around 20 25%. And so just
in that 3-month period, the prevalence
went up from 25 to 40. And so as that
prevalence goes up, the the downstream
impact is going to be proportionally
more. And uh we find that the effect was
primarily driven by position zero
prominence. Uh and so this suggests that
there is some work required to
understand like what is it about AI
overviews which is leading to these
effects. Are users really liking the
summary, finding it valuable or is it
just that they're occupying prime real
estate and uh and that's what is leading
to all the results and yeah users are
surprisingly indifferent to the presence
or absence of AIO views. But yeah, users
are extremely resistant to AI mode as a
default. So we also felt this was an
interesting result because there's a lot
of conversation about how it might be
the end of Google as we know it with the
rise of chat GP etc. However, it seems
that
a as a default uh people are resistant
to a fully conversational AI interface
and so the standard Google as we know it
is probably here to stay for some time
at least. So that's all that I had. Uh
I'm happy to take questions and hear
your thoughts. Thank you.
>> Wonderful. Thank you so much Sah. Uh now
Dante you can you have five minutes to
help us start the reflections. And for
the others if you have questions
uh that you don't want to ask directly
uh you can uh uh put them in the chat
and I will ask them.
>> Great. Um hi everybody. I'm Dant at
Columbia Business School. Thanks for
inviting me to discuss this paper.
Congratulations to the authors. Uh I
think this is an important and very well
executed paper. An important question.
the the the design is very clever. So it
was nice to read and learn a lot. So now
my job is uh is to push on I think the
interpretation of some results and also
high level open questions and I want to
do that uh through the lenses of the
main stakeholders in this market. So the
first is the search engine and the
consumers and then also the publishers
and the brands. So first let's take a
perspective of the search engine as well
as the consumers together. So um one of
the last results that was shown uh and
is also the headlines of the paper is
that the overviews divert traffic
without much improvement into uh the
user experience. Um so this is actually
my only push back a little bit. So I
want to gently push back on this second
part of this claim. Uh so the conclusion
about the user experience not going up
or you know not moving rest on this
stated preference survey. um and uh it's
based on this satisfaction score, right?
But I think actually if we look at what
people do, they click less, right? They
accept zero click answers. So if we
think about this in terms of reveal
preference, then this might actually
imply lower search cost. And so people
find what they need without paying the
you know the click and scroll tax. So
the behavior I think it's consistent
with the utility going up um even if
satisfaction scores stay flat. Okay, so
it's a different interpretation uh which
I think is is important. So I would be I
would be careful concluding there's no
benefit to users from that survey based
results. And if we think about online
reada platforms in general, reductions
in clicks per search normally reflect
improvements in search efficiency. And
so that logic can apply in your context
too, right? So perhaps you can look at
the time to find the answer metrics or
task completion time metrics or measures
and I think those would speak directly
to consumer welfare more than the
survey. U so if we look at the consumer
side only I think there the deeper
welfare question isn't just how many
links they click but actually what I end
up consuming. Okay. So given that in
your context most searches are um
informational, the question I think is
do consumers read different information
from different sources uh with different
length or authority compared to the prei
word.
Sorry. So I think perhaps you can look
at what is cited in the overview and
compare the sources cited there against
the top rank results uh organic results.
Right? For example, I would be curious
to see whether more word of mouth
sources are cited there versus
established news outlets. And I think
you have the data to answer these
questions and you you can do that. So
now let's switch to the publisher side.
I think for publishers that's where the
results are strongest very convincing.
Uh what I would do I would frame the
paper more explicitly around the type of
good that is consumed here. So in the
footnote you have the thatformational
queries account for 71% of searches and
that's where most of overview trigger
rates is is the highest. Uh so I think
the results speak directly to
publishers. So these results speak to
them. Uh so this is about news content
information consumption right so I would
make that ex that footnote much more
salient in the paper more prominent
rather than just you know keeping in a
footnote. So u the other question then
on publishers is about the the overview
itself as a channel right so the
question there connects to my previous
comment so are the sources cited in the
overview uh the same as those cited in
top rank organic results or are they
something else where basically overview
borrow from right this is critical right
do users click uh and the links in the
overview that are provided there
actually wasn't clear to me if the
outbound clicks also include the uh
links that are provided by the overview.
Um, and you know, if being cited in an
overview recovers some of that traffic
that instead would go to the to the
organic results, then you can reframe
the story from a pure cannibalization
toward a a GEO game, right? Generative
engine optimization perspective. And I
think this is actually an important
question, a deep question in digital
marketing, right? So this opens the
chance to uh understand whether GEO
efforts for brands and for publishers
pay off, right? So where basically the
overview site from right does it borrow
from the organic results or not
and then finally you have the brand side
and then I'm done. So the brands you
know the commercial transactional side
uh so you have a clean null sponsor uh
clicks result which is interesting but I
think that's kind of partly might be
mechanical and I think you can elaborate
on that. Uh so the overviews in fact in
that side are are less common for
transactional queries. uh you you you
said it you know 50% trigger rate out of
12% of searches that are transactional.
So I think for those searches where the
sponsor link matter the overview mostly
isn't there right is not there much much
of them. So what I would do I would make
an argument explicitly rather than just
presenting the no result and I would try
to justify that that no conclusion more
carefully. uh and also for brands and
this is my last point for brands and
consumers. I think the first order
question in this era is whether these
new search environment affects discovery
and choice. Okay. So what products
consumers buy, what prices they pay,
what attributes they look for. So this
is an open question, right? Um and I'm
not I'm not sure that you should answer
this question in this paper, but maybe
you have the data. So if you can say
anything about for example studying the
prompts consumer writing the eye mode
more explicitly trying to understand the
discovery phase for products for brands
I think that would be really valuable in
this domain. Okay so to sum up I think
this is a great paper uh do you have
credible causal inference and causal
evidence uh the implications for traffic
are very important. So suggestions are
mainly soften the interpretation of some
of your results and maybe exploit the
data to answer question on the
information and also on the discovery
phase using the prompts that you have
the prompts that that you have and
thanks a lot for inviting