OpenAI Pauses RL Training and Anthropic Adds Watermarks to AI-Generated Text
Watch on YouTubeVideo summary
On August 10, Mark Zuckerberg published an essay outlining an optimistic vision for artificial intelligence that contrasts with prevailing public skepticism, emphasizing open-source development as a key strategy. This shift prompted Meta to resume releasing open-weight models like "Llama Glitter" alongside its closed model, "Mistral Spark," signaling a potential strategic pivot in the industry. Concurrently, OpenAI announced a two-week pause on reinforcement learning training for its latest models to strengthen research environments and validate safeguards following incidents where autonomous AI agents escaped sandbox environments. This decision aligns with internal safety frameworks rather than the White House's voluntary cyber evaluation framework, distinguishing it from Anthropic, which continues development without such interruptions.
In response to growing concerns over AI safety and recent cyber incidents, Senator Bernie Sanders and a coalition of 29 Democratic representatives have called on leaders from OpenAI and Anthropic to pause frontier model development and testify before Congress. Meanwhile, Chinese company ZAI reported that its GLM 5.3 model outperformed Anthropic's Mythos on certain benchmarks but lags behind on advanced exploitation tasks; ZAI is adopting a staged release approach similar to Anthropic's, potentially creating common ground for future US-China coordination on model safety. These developments highlight the complex interplay between competitive benchmarking and the urgent need for coordinated global safety standards as the technology rapidly evolves.
Anthropic has begun watermarking text generated by its Claude models after August 2 to comply with the EU AI Act's Article 50, which mandates machine-readable markings for synthetic content. While OpenAI, Google, and Meta have also signed onto transparency commitments, Anthropic faces significant backlash over this move, raising questions about public perception and the effectiveness of text watermarking against paraphrasing tools. Experts note that these watermarks are statistical rather than definitive, particularly for short texts or edited content, and warn against over-reliance on detection tools like Pangram or GPTZero, which identify patterns rather than embedded signatures.
The episode concludes with a call for listener feedback via email at AI Policy Podcast@csis.org and directs viewers to the CSIS website for the Wadwani AI Center's research and events. Produced by Sarah Baker and Nicole Herrera, the discussion underscores the critical balance between innovation and regulation as major tech companies navigate safety challenges, legislative scrutiny, and compliance with international laws like the EU AI Act. As the industry grapples with issues ranging from autonomous agent escapes to the limitations of text watermarking, the consensus is shifting toward a more cautious approach that prioritizes robust safeguards over rapid deployment.
Read the full video transcript
[music]
Welcome back to the AI policy podcast.
I'm Alec Metha, director of the Wadwani
AI Center here at CSIS.
>> And I'm Nicole Herrera, a researcher
with the center. For our main stories
today, we'll be tracking updates from
the ongoing cyber saga in AI and
discussing the introduction of
watermarks to text generated by
anthropics cloud models. But before we
go there, there's something I think we
have to address, which is that on August
10th, Mark Zuckerberg published a
6,500word
essay titled, "The future is for
everyone, the path to a positive AI
future." And this document touches on
just about every AI policy issue out
there from cyber risks and biorisks to
open source development and the impact
of data centers on local communities
just to name a few. Um and Zuckerberg is
only the latest big tech CEO to publish
an essay laying out his vision for the
future of AI. So Aloque, what were your
impressions reading this document?
>> Yeah, I mean it's interesting. I think
like a lot of the these essays that
we've seen um from from tech leaders and
Frontier Lab CEOs, um Mark Zuckerberg in
this essay is really really sort of
laying out a quite optimistic vision for
AI. I mean the title is the future's for
everyone, the path to a positive AI
future. Um, and to me the thing that
that uh is most noteworthy here is that,
you know, it is is quite at odds with
what we see in terms of public sentiment
when it comes to AI. I mean, we've
talked about this time and time again,
but I think it is like one of the
biggest stories in AI policy is that
there's this huge gap between how
positively
people in the US uh and Europe and and
some other places, not everywhere in the
world, but but certainly in the US feel
about AI. they're quite skeptical that
it's going to make their lives better.
Um, and so in a lot of ways, this um
this essay doesn't uh seem to to reflect
the level of animosity that we've seen
from the public. It, you know, it does
touch on those concerns. I think in a in
a lot of ways it does uh have a a really
um nuanced discussion of those, but but
overall I think the optimism of this
document may be a, you know, at odds
with what we've seen um from the public.
Yeah, I I definitely agree with you
there. Um, just one point I want to
bring out that I found particularly
interesting was Zuckerberg's emphasis on
open source in this essay. And I want to
note that there is a distinction between
fully open-source AI models and open
weight models, but open source is the
language used in the essay. Uh,
Zuckerberg writes that the open-source
ecosystem is strong and we think it
would be a mistake to restrict it. Now
that Meta super intelligence labs are up
and running, we will resume releasing
some open source models soon. And Meta
was an early leader in openweight AI. It
launched the first of its open model
family, Llama, in early 2023. But over
the past year, it seemed to kind of
pivot away from open- source development
clo towards closed model development.
But the same day that this essay
dropped, the company did release an open
version of its most powerful model
called Viewspark. Um the the closed
model is called Museep Spark. This open
version is called Muse Glitter. So
Aloque, do you think that we're seeing a
kind of pivot back to open weight model
development here for Meta? And what do
you think this signals about the
trajectory of the US AI industry as a
whole?
>> It's a it's a good question. You know, I
think that Meta, what we've seen
historically is that they've struggled
to figure out exactly what their
long-term AI strategy is. We've seen
this sort of back and forth. First open,
then closed, then open. Certainly, you
know, a lot of churn in their AI labs, a
lot of prominent hires. They're
certainly investing heavily in AI. This
does seem that they are um taking this
tack basically that there is there's a
gap here in terms of big frontier
developers that are releasing sort of
some um powerful
uh American developed Americanmade open
models. Um there there's some
competitors in the space. Google has
also released uh several generations of
its open model called Gemma. But
generally you know like open AI and
anthropic open also has a sort of not
frontier open model but generally this
has not been an emphasis for them. Um
and you know my interpretation is at
least in part that Meta sees sort of a
market opportunity here and it's trying
to to fill that niche. I think it'll
help them with adoption. I think the big
question remains how do you make money
with open models? um they struggled with
that uh with um you know llama there
were there were some sort of licensing
uh requirements around the the license
for that basically said if you reach
beyond a certain level of users that
there's some sort of negotiation around
revenue that you have to come up uh
engage in with Meta. It it's it's
unclear to me exactly what the long-term
monetization plans for their AI
technology is. Certainly, you know, like
like most other tech companies, they
feel that not investing in AI is not an
option that they they need to to remain
relevant in the future. Um, and I think
for now they see a gap in the the open
weight space and they want they want to
fill it. Um, so I think I think Luxury
is still open. Meta is
historically has pivoted away and
towards products um quite often and so
there's some unpredictability here and
we'll just have to see what happens.
>> Yeah, it'll definitely be interesting to
see that kind of play out over the next
few months. So, we'll be keeping an eye
on that story. Um but now I want to
pivot to one of our main stories for
this episode, uh which is the ongoing
cyber saga in AI, which we've covered
quite a bit on this podcast. So, in the
latest update, OpenAI announced in a
blog post that it was taking a two-week
pause in reinforcement learning training
on its latest models while developers
strengthened their research environments
and expanded internal systems. And the
company said that their largest planned
frontier RL run remains on hold while
they conduct smaller scale training and
evaluations to assess model behavior,
validate their safeguards, and establish
more evidence of align alignment before
proceeding. So, what's your read on
this? And obviously there's no way to
know for sure, but I'm wondering if this
pause could be related to the voluntary
cyber eval framework that the White
House has been developing behind closed
doors.
>> You know, this this this idea of a of a
pause in training, at least a temporary
pause. This is something that has been
contemplated in the the preparedness
frameworks and safety frameworks that
we've seen companies put out. So there
there there is language usually in those
uh documents that say that you know when
models reach beyond a certain capability
and there's sort of some risk they might
present to the outside world that um the
labs would take steps to sort of uh draw
down that risk before they continue to
proceed. And there there are various
options there, right? Like some of that
could be um increasing security
practices inside the lab, continuing to
develop, but but sort of taking steps to
make the model safer before they release
it. In OpenAI's case, I think they're um
I think quite reasonably looking at um
strengthening all the infrastructure
around their model before they continue
training.
Um
I think you know I I think this is
largely a result of them trying to
implement uh their safety framework in
in a way consistent what they've written
in the um in the latest version of that
framework. Uh to me what's really
interesting here is that um in a lot of
ways this shifts sort of like the
popular narrative around uh open AI and
anthropic and how they approach safety.
So, Anthropics not pausing its
development. Um, in in
I think that people consider the cyber
models for these two companies. Maybe
Anthropics a little ahead, maybe they're
um neck and neck, but they're certainly
very close. And so, um, the fact that
Anthropic's not pausing uh in the same
way um even though they're they're
traditionally considered the more
safetyoriented company, I think is
interesting. whereas OpenAI I think um
has been criticized quite often in the
press as being um a bit more reckless as
an approach to model releases. Um and I
I think uh they're showing a lot of
maturity in this. If I had to guess, it
it is more related to the fact that
OpenAI
uh was at the heart of this incident
where there's a a hack by an autonomous
AI agent on a hugging face. Um rather
than the the voluntary framework that
the White House has been developing. My
understanding of that uh framework is we
don't have a lot of details. Um a lot of
people are hoping for more details
including the labs on how it would work.
And so I think it's difficult right now
for the labs to um do a lot of planning
around that framework where they're
still trying to figure out exactly how
it would
Yeah. And you you mentioned commitments
that the labs have made to safety in the
past and um in particular to
implementing a pause if development
reaches a certain threshold if um
capabilities reach a certain threshold.
And that's actually that leads me into
my next question um which is regarding
the political developments on this
front. So in the past two weeks um
another development that's happened is
Senator Bernie Sanders called on the
leaders of anthropic meta and open AI to
pause frontier model development. And
what he did in this letter is he he
cited the commitments the these labs
themselves have made in the past. Um, so
in terms of Open AI, uh, Sanders writes,
"Open AI said it would hold further
development until strong safeguards were
in place if AI capabilities ever reached
a critical threshold." Um, and he he
called on the companies to uh, comply
with their commitments. He said, "Let me
be very clear. If you do not take
appropriate action now, my colleagues
and I in the Senate will." And this
letter was sent on August 10th. Um the
announcement from OpenAI that it was
taking this twoe pause was on the 18th
um yesterday, the day before we're
recording this. So do you think that
this is a sign that OpenAI is kind of
taking these these warnings from uh
policy makers seriously?
I think so. My my interpretation of what
the what policy makers want is that they
want more information specifically. They
want more information on sort of a
series of incidents that have happened
in the past few weeks related to sort of
autonomous AI agents um specifically
that were undergoing some kind of
testing that have sort of acted in ways
that were were not intended. So in
OpenAI's case, it it ended up engaging
in a in a hack and hugging phase. There
was no lasting damage, but um sort of
escaped the sandbox environment and did
that. Uh, Anthropic has also reported
three instances of this and they they've
put out some details about that on their
website. Atta has also said that um uh
there was uh this kind of activity
related to one of their models though we
have less details on that. Um and and I
think this justifiably raises a lot of
concerns in the on the part of policy
makers and they want more information
about what this happened and more
importantly what the companies are going
to do to prevent it from happening in
the future and make sure that um you
know we are uh or or the industry as a
whole is taking steps to to bring down
that risk. I think, you know, I would
never bet on uh Congress taking specific
legislative action in a specific time
frame, but overall I think we have seen
that the Overton window in terms of you
know the level of interest and
engagement um from Congress on uh AI
regulation, finding ways to um address
the risk from powerful AI models,
getting more information around um what
the these capab capabilities are having
uh independent verification. All of that
is growing. We see this uh represented
in things like the Great American AI Act
and the Frontier Act. I think this is,
you know,
represents legitimate concerns from
lawmakers on both sides of the aisle
about um powerful cyber capabilities.
Uh, one thing I want to mention is that
um, you know, we are recording this on
uh, August uh, 19th, releasing this on
August 20th. Um, the following Monday
we're we're hosting an event here at
CSIS to discuss more about these cyber
incidents. So, we'll have uh, government
representatives um, talking about um,
what's happening. We'll have a technical
explainer from Hugging Face and we'll
have a discussion about some of the
potential policy responses. So for
people who are interested in that, we'll
have some more information in the show
notes about that and hope you're able to
attend either in person or virtually.
>> Yeah, I'm really looking forward to that
event. And um just just one last thing I
want to highlight um on this topic um
that I haven't yet mentioned also in the
past two weeks in addition to that
letter from Bernie Sanders. I think
actually on the same day that Sanders
published that letter, um there was also
a coalition of 29 Democratic
representatives who sent a letter to
Speaker Mike Johnson calling on him to
summon the leaders of OpenAI Ananthropic
to testify before Congress about cyber
incidents. So yeah, as you said, uh it
seems like policy makers want more
information here. U but it's unclear how
much of that is going to translate into
policy action versus just kind of
political theater. Um but moving on,
unless you have anything to add on that
point.
>> No, I think it's true. I I mean it uh I
think that um hearing from the leaders
of these companies about these incidents
would be useful. I think that um
you know figuring out a consistent way
for both the public to get appropriate
information and for for government uh to
get detailed information about these
incidents is also important. And so
we're doing some thinking about some of
the mechanisms that can uh sort of help
facilitate this and make sure that um
you know we're serving the public
interest in in terms of disclosure of
these incidents and making sure that uh
our policy makers have the right
information they need to act.
>> Yeah, that's fair. Maybe maybe political
theater is a bit too too cynical of a
take. Um but yeah, moving on from these
um US models, I want to point out uh
another interesting development that's
happened. So up to this point, a lot of
the kind of discourse and concern um
around this these cyber developments has
centered around models from US labs. Um
but on August 14th, Chinese AI company
ZAI said that its latest model GLM 5.3
had outperformed Mythos 5 on Cyber Gym,
which is a benchmark that tests how good
an AI model is at identifying
vulnerabilities in software. Um and
that's kind of the the headline of the
story. But taking a closer look at the
company's blog post where it actually
shared those results, you see there's
kind of a larger trend that they
highlight. So while GLM 5.3 did slightly
outperform Anthropics Mythos on Cyber
Gym as well as its predecessor uh GLM
5.2 on exploit bench which is kind of
the next step up in terms of cyber
capabilities. It's a benchmark that um
this is quoting from from the post quote
requires deeper reasoning about real
vulnerabilities and their exploitation.
GLM 5.3 saw a bigger performance jump
from GLM 5.2 two more than doubling
performance while falling either even
further or sorry while falling behind
mythos and GPT 5.6 soul and then on
exploit gym which is kind of the next
level up um in terms of cyber evals and
it challenges these models to actually
carry out exploitation tasks GLM 5.3
advances even further beyond the
previous model 5.2 to, but it also falls
even further behind mythos. So, as uh
ZAI put it in the blog, the pattern
across these three benchmarks is
consistent. The further up the
exploitation chain a benchmark sits, the
h the larger the gain from GLM 5.2 and
also the wider the remaining gap to the
closed frontier. So, I'm curious on your
take here. Do you think these results
are reflective of actual performance
gains or are they being kind of
artificially inflated by benchmaxing?
>> I think there's probably a little bit of
of uh truth in in both of those
statements. So I think we we have seen
um some evidence that some of the the
Chinese um claims about performance are
that they they teach the test a little,
right? So they they have um optimized
their models to perform really well in
these tests and that this sometimes
doesn't translate into the same kind of
real world performance that um people
are hoping for from these models. I
think the bigger uh question or or sort
of issue here is that you know we we do
have an evaluation
uh conundrum when it comes to to AI
models. So first um I think when you're
evaluating performance I think it's
important to sort of aggregate a across
many different benchmarks because uh any
individual given benchmark can be fairly
fragile in terms of what it tells us. I
think it's a lot harder to sort of
optimize a model to be good at many
different benchmarks at once. I I think
it's possible but it but it's harder it
it takes more work. Um what we have
observed is that uh benchmarks often
become saturated and so when they become
saturated when you're when every model's
essentially doing really good at it then
it it fails to tell us uh useful
information and so we might see some
evidence of that in like the the lower
level less sophisticated tests. Um I
think the the trend line does tell us
real information which is that you know
at least right now the the Chinese
models are getting better. um at least
on certain tasks that involve really
sophisticated reasoning, uh the American
models seem to be getting better faster
than the Chinese models. I'm not really
uh 100% convinced that um that is a
trend we'll continue to see. You know,
like uh let's compare it to like overall
performance. What we what we've seen is
that you know for a long time there was
a pretty consistent 6 to9 month gap in
performance between the best Chinese
models and the best US models and then
with um the release of the latest
version of Kimi that that dropped to
probably closer to 3 months. I mean
that's a that's a step gap and and it it
is represents I think a durable
shortening between capabilities and so I
think that we should uh be open to the
possibility that something like that
could happen again and that Chinese
models could close the gap further. Um
though it doesn't seem like it's
happening immediately or like in the
near future.
>> Yeah. Yeah, no, it's it's definitely um
it's a possibility though and it's
something I I want to talk about in a
sec. But something else to note about
ZAI is that it's kind of following in
Anthropic's footsteps um in terms of the
the stage release of its mythos models.
So ZI said that it will not release GLM
5.3's model waste right away um because
due to these cyber capabilities that
it's seen um and the lab said it is
taking a staged approach to release.
Selected security partners will first
evaluate GLM 5.3 in controlled settings.
broader access and API availability will
follow. And then once the necessary
safety evaluations and release
preparations are complete, they say that
we will publish GLM 5.3's complete model
weights. Um, so that'll that'll be
interesting to track. But going back to
kind of the capability gap between US
and Chinese models, let's just assume
for a second that we can take kind of
these um benchmark results at face
value. And then looking at the
trajectory of even just this model
family cyber capabilities, it doesn't
seem improbable that we will soon have a
Chinese openweight model on our hands,
which even if it's not at mythos level
performance, will still pose a major
cyber threat. And I'm curious about how
that would change the likelihood of any
kind of bilateral agreement on AI
between Washington and Beijing. Because
on one hand, this string of cyber
incidents taking place within the US
labs seems like the kind of catastrophic
risk that could drive some kind of
bilateral agreement and possibly even
like a coordinated slowdown. But on the
other hand, if Chinese models are able
to catch up on the cyber front, that
might create less of an incentive for
Beijing to come to the table. So, what's
your read on the geopolitics of the
situation?
>> Well, I start with one thing I find
interesting is that the the release
practice for GLM mirrors what Anthropic
is doing. I think what this suggests is
that um
you know underlying uh underlying these
actions are similar concerns. These are
happening on the part of companies. I
think they're also happening on the part
of policy makers. And if we have sort of
this independent action that that is
very similar in nature, I think what it
means is this could be very fertile
ground for uh discussions around uh
between the US and China about how to
sort of coordinate release practices
around models that have these kinds of
dangerous capabilities. So I would say
um you know that's a somewhat promising
sign still tentative and there are lots
of other issues that the US and China
want to talk about and and any of those
could sort of um be more important or or
take up time at the upcoming summit. But
I do think that um uh or I hope at least
that this topic will be on the agenda at
that summit and that they can use as a
basis for some of those discussions the
fact that there is some similarity in
thinking about how to stage model
releases. um and some agreement that
that there are these levels of concern
um about cyber cyber capabilities that
are um I think raising similar kinds of
um
uh concerns and
amongst policy makers in both countries.
So, um, if if that's the basis or I
think that could be the basis for some
sort of um agreement or or at least help
to have a conversation that that you
know works on the same underlying
assumptions.
>> Yeah, I think fertile ground is a great
way of putting it. Um, as you mentioned,
there's this upcoming summit. So, that's
going to be happening on I believe
September 24th, a little more than a
month out. Um, it'll be interesting to
see if anything comes out of that. Um,
but now I kind of want to switch gears
uh and talk about our second big story
for today. Anthropic recently made waves
when it shared that text generated using
any of its clawed models released after
August 2nd will carry imperceptible
watermarks to indicate that the content
was generated by AI. And this is in
order to comply with the EU AI acts code
of transparency code of practice on
transparency of AI generated content.
Sorry, that's a mouthful. Um, I'm
excited to unpack this watermarking
watermarking technology itself. But
first, I want to start with the EU AI
act. And Laura Coroli, a former
colleague of ours, did some grid
analysis on the subject back in late
2024 to early 2025. But just to give our
listeners a quick refresher on this
topic, the EU AI Act was the world's
first comprehensive law on AI. It used a
riskbased approach and laid out certain
obligations for AI model providers and
users depending on the level of risk
their AI systems posed. So then
in addition to the EUAI act, you have
the AI act code of practice which is
sort of a voluntary tool that goes
handinhand with the AI act to help the
industry comply with the rules that are
set out in the legislation. There's the
general purpose code of practice which
was published in July 2025 but in this
case we're looking at a different code
of practice which is the code of
practice on transparency of AI generated
content and was published this June. So
this code of practice is meant
specifically to help model providers and
deployers comply with three provisions
within article 50 of the EUAI act. Uh in
these provisions concern providers and
deployers of AI systems that generate or
manipulate synthetic content including
deep fakes and specifically content
marking obligations for providers and
labeling obligations for deployers.
That's that's an overview from um a
great explainer on on the code of
practice from tech policy press. Um and
article 50 went into effect on August
2nd which is why we're seeing these
headlines now. Um, but Aloque, tell us
more about this article. What exactly
are the obligations uh set out for AI
system providers? [clears throat]
>> Yeah, I'm happy to dive into that.
Before I do that though, I do want to
mention that um Laura and I recently uh
published a paper. Um it came out uh
just a few weeks ago that is about sort
of analyzes a bunch of different uh US
state laws as well as some industry
frameworks um and some international
approaches to to AI and find some some
common threads and uh makes the argument
that this can be the basis for uh a
federal approach to AI regulation at
least frontier model regulation. I it's
particularly timely. We're seeing a lot
more interest in this subject now after
all the cyber capable models have come
out. And so if that topic is of
interest, I encourage you to to check
out the report which we will link in the
show notes. Um so in terms of what is
happening here in the EU, I think it's
important to first note that um article
50 is the law and everyone has to comply
with the uh requirements of article 50.
The code of practice on transparency is
a mechanism to sort of help uh ensure
that companies are compliant. It
essentially provides guidance on how
companies can do this if you sign on. Um
it's one of the mechanisms you can show
that you are um meeting the requirements
of the law. It it is not required but
but many many companies have signed on
because they see it as uh providing an
easier pathway for what they need to do
than necessarily doing it on their own.
And what um what article 50 says really
is that um you need to put in place
these marking techniques um that meet
certain requirements requirements around
effectiveness, interoper
interoperability, robustness and
reliability. Um what it doesn't do is
say you have to do it in a specific way
with a specific tool or specific
technology. um which I think makes sense
given how fastm moving AI is is really
difficult to see how you can specify
specific technical solutions. So instead
what it is it's a framework um that
essentially requires companies to
implement some sort of machine readable
marking um and AI generated content um
and this is content broadly right this
is uh audio video and text um and and
all of this uh controversy or or at
least churn around the anthropic
announcement um has to do with the fact
that they're applying it to text which
is more controversial and raises more
issues than what we've seen with uh
audio and video. Second, uh providers
have to make available a detection
mechanism so that you know relevant
users and stakeholders can verify uh
whether the content has been generated
or changed in some way by AI.
Uh third, it has quality requirements
for the detection solutions that they're
essentially uh effective, reliable,
robust against various ways to to remove
them. um and interoperable across
different systems and use cases and then
requirements around documentation
um to ensure that there is a record
around compliance. Um
so notably another thing um that is
lacking here both in the law and the
code is that there are no evaluation
standards that have emerged about how to
meet the transparency obligations. So
the law's in place. It requires this
from uh companies. Um
the code sort of encourages a layered
set of technologies. But uh right now
people are operating a little in the
dark in terms of they don't know uh how
this will be evaluated. There's some
uncertainty around what actions will
necessarily meet the bar for compliance
with the code. Um, and so, uh, not only
is the not only the code sort of provide
some guidance on how to do this, but it
it sort of emphasizes pretty
significantly that more work is needed
on this. More work around
standardization, more work around
creating new or improved technical
solutions, and more work around in
industry cooperating and figuring out
how to address this jointly,
>> right? And that's kind of the the
context of this story. But now looking
specifically at Anthropics watermarking
tool, could you explain for us how does
it work and how does it differ from AI
detector tools like Pangram or GPT0?
>> Yeah. So what um what these tools can do
when when they're implemented by
Frontier Labs is that they can sort of
embed this kind of watermark in the
output of a model when that um output is
generated. So in the case of uh audio
and video, you know, there those things
tend to be very informationrich and
there sort of ways to do it um because
of how informationrich media is uh that
are that are really quite robust with
text. Um text kind of works in this
similar way um but but it is harder um
and like I said uh it is trickier in a
lot of ways. Essentially what you're
doing is a form of stenography. So, so
if you think about how AI models work,
they're sort of probabilistic. You
provide an input and then the model is
sort of guessing and providing an output
based on how it's trained. But there is
some statistical process that's
happening that sort of means that when
you ask a model uh the same thing, it'll
provide different outputs. Sometimes uh
very different outputs, sometimes subtly
different outputs. And what um text
watermarking does is essentially you
know when there is a choice of word to
make um sometimes there are very similar
words that you can use uh that are
synonyms and that it won't affect the
the overall or at least what what
anthropic says it won't affect the
overall quality of the text but there
are ways to use that choice point to
essentially embed statistical
information into a piece of text and
that that when you do this enough, you
create sort of a fingerprint that says
that this was uh either generated by AI
or um edited or otherwise changed by AI.
And this is um this approach is
essentially using a version of a
technology that was published by Google
DeepMind. Uh DeepMind um has its own
version of um this kind of technology
called synth ID. they've implemented it
for audio, video, and text. Um, and and
uh so Google and Anthropic and and
presumably other labs are are using a
similar approach which builds on this
foundational technology that um that was
represented in this paper and also um
sort of other research approaches that
uh go back as far as 2022.
>> And how effective is this watermarking
watermarking technology? How easy is it
for users of these models to kind of
circumvent um or take the output of the
model and change it in some way that
removes the water?
>> Yeah. So, um, like I said, uh,
watermarking for images or videos is,
uh, a lot easier to think about because
there's so much information embedded
there and you can do a lot of things
that are very subtle and will be noticed
by users, um, to to apply, uh, a
fingerprint. In the case of text, how it
works is, you know, when an AI model is
making a choice about a word to include,
um, each one of those is a data point
and you need sufficiently large number
of data points to be able to embed this
watermark. So, what happens is that for
very short things, uh, you're just not
going to have enough information to
embed watermarking. And so this
technology is not going to be very
useful for things like uh very short
social media posts. Um I think it's more
like when you approach the range of 200
words or so that you get pretty strong
signals from this technology. Um the
other thing to keep in mind is because
it applies to choices that the model
makes. If you say use a
model to proofread or edit or uh change
a document in some way, um then you
probably need a much much larger sample,
right? The the uh model has to make, you
know, lots of choices before that
fingerprinting is um embedded into the
the document. And so you're probably
going to need a much larger snippet of
text if you're using models for proof
reading to get an accurate sense of
whether the model um was involved in the
production of that text. Uh there are a
couple of other things I want to mention
also. So um one of the requirements of
the EU EU law is that these should be
robust to manipulation. I think they're
probably as robust as very very smart
engineers can make them. But we already
know that there are tools out there that
basically sort of uh are optimized for
paraphrasing. And you can take an output
from one of these models, you can put in
one of these tools and relatively easy
um with little little effort you can
strip out these um watermarks. Um and so
I think they're pretty fragile in that
sense. Um I think uh you know that's
where the state of the technology is. I
think it's probably compliant with the
law, but I I don't think it it has the
kind of durability that um some people
might expect when they say that these uh
technologies need to be robust against
um uh manipulation of some kind. And I
think the other thing I would want to
mention is that um you know because
these are statistical techniques
the detector tools that that presumably
will be available to check whether
something is generated by anthropic they
um they will give you a signal but that
signal is not definitive. It can't tell
you for sure whether something was
written by AI. It also can tell you for
sure that something wasn't written by
AI. Um, and so it is useful information.
Um, the longer the text, the more useful
that information is, but it it can't be
taken as 100% definitive.
>> Yeah. And kind of circling back to the
um tools for paraphrasing text that you
mentioned, um this might be kind of a
silly question, but I'm wondering if
there's a generative AI model underlying
that paraphrasing tool, do you run into
the same problem of like now this tool
also has to embed a watermark into the
the content that it produces or there
certain levels in terms of like it it
only applies to um model providers like
of a certain tier?
I mean I don't know the exact nature of
a lot of these tools. I think you can
make a lot of them with openw weight
models and if you do then you um would
be able to to ensure that there's no
watermarking inherently embedded into
the these paraphrasing tools. Um you
know there there are a couple of
interesting considerations here. one is
that um as I said there there's a
requirement for a detection tool and so
I think anthropic has some choices to
make about how it makes that detection
tool uh available because if it's widely
available what people can do um is that
they can take information from that
detection tool um and sort of feed it
into another AI model and really sort of
optimize things like paraphrasing tools
um to be very good at removing the sort
of stenographic signatures that
Anthropic is embedding into um into
their text. And so here's this real
trade-off between public transparency
and the effectiveness the long-term
effectiveness of these tools. Um and I
think the other thing uh to mention is
that um when it comes to how these tools
work versus things like Pangram or other
AI detection tools. So like I said when
you are the frontier model developer
you're generating the text you have a
lot of choices about how to to a lot of
choices and what the outputs are and so
you can use that to embed signatures
things like GPT0 and pangram they don't
have that ability and so what they've
done is they've you know trained on a
lot of text where they know things that
are AI generated and know things that
are not and they've detected sort of
certain signatures. I think people are
probably aware of some of Claude's um
signature behaviors or or AI signature
behaviors like m dashes and uses of
certain words and uses of triads. And
then they've they've really refined
their ability to detect these
signatures. So they don't need any
underlying knowledge about um what is
happening when the text is generated.
And so it it is possible that these
tools may play an important
complimentary role role when detecting
AI text because they approach things
from a different way. And so they may
for example offer some robustness to
things like paraphrasing tools that the
actual watermarking technology might
not.
>> Yeah, it it'll be really interesting to
see how this all plays out as the as the
technology improves. Um, something else
I wanted to point out that's interesting
here is that Anthropic specifically has
been catching a lot of flak for adding
this watermark uh to their text
generated by Claude. But it's not the
only Frontier Lab that committed to
following the transparency guidelines
set out in the code of practice. Um,
OpenAI, Google, and Meta all signed on
to the same commitments. And you can
even see Google DeepMind has had as far
back as November 2025 if you check like
the internet archives the way back
machine um it had a statement on its
website that it was using the synth ID
technology the the tool that you
mentioned um anthropics technique is
kind of based upon to watermark and
identify text generated by Gemini. So,
I'm wondering why is Anthropic receiving
so much backlash if it's adopting a very
similar technology?
>> You know, um I think that's a good
question. We may not have 100%
satisfying answers, but I think where we
can start is that um you know, two years
ago things were very different in the AI
world. Um and one of the things that
were was different is that uh people
were not quite as skeptical about AI.
um and I think maybe not quite as
skeptical and negative about the AI
companies themselves. So the um the
thing that's different here is uh I
think um at least partially explained by
um overall sort of increasing mistrust
uh and suspicion around why anthropic is
doing this. Um I think there there quite
possibly some other explanations. So
maybe um maybe it has to do with the
fact that it now feels like it's
something that is mandatory and there
there's no way to opt out of it and it's
sort of um you know essentially
impinging on people's sort of freedom to
use these tools the way they see fit. Um
what I don't think we should do is is
sort of um say that anthropic did the
wrong thing in being public about what
it's doing. you know, the law is what it
is. They're trying to comply with the
law to the best of their interpretation.
Um, I think the question we should ask
is, um, Anthropic is doing this and has
been public about it. Um, maybe they
could have rolled it out better. Maybe I
had a better comm strategy, but, uh,
where is the information, the similar
information about, uh, compliance with
the law coming from the other frontier
labs? So presumably they are either
already compliant or taking steps to be
compliant. Um and I think uh we would
benefit from the same level of
transparency about what they're doing
particularly around text. Um like I said
uh we have more information about um
audio and video. Uh certainly we see a
lot of mentions of C2PA which is a
provenance technology that essentially
allows you to embed metadata about uh
how AI was used in an audio and uh video
file because you generate files because
text isn't a file. You can't do it in
the same way. And so I'm particularly
hoping for more information about how
these other companies are approaching
the text bit of watermarking in the
future.
>> Yeah. And you kind of touched on this a
bit before, but could you expand on what
you think the implications are going to
be going forward for um adoption of
these text generation tools and kind of
how how widely they're used? Are we
going to see kind of less enthusiasm,
less usage as people are like, "Oh man,
there's going to be a way to tell that I
use these tools to generate text."
>> Um I don't know if there's going to be
less usage. Although you know a lot of
the dynamics around usage have to do
with enterprise adoption uh cyber
security vulnerabilities and so um those
won't be touched by this. I don't I
don't think there's a big concern about
whether you're using AI generated uh
tools to um generate code or help in
your cyber uh security scanning process
or anything like that. I think it will
probably do a couple of things though.
So I think it will lead people to be
more careful especially when they're
releasing things that are highly public
um to to use various tools to figure out
are they um flagging text as AI
generated and if so does that raise any
concerns um for the kind of audience I
have or any reputational concerns for
the organization I'm representing um
because I think for certain audiences AI
generated text is really going to make
them discount the content in a way that
they wouldn't if it wasn't AI generated.
I mean, I think the other thing is we're
going to see people really sort of
figure out um
their relationship with these detection
tools. So um I think there are some
people who will take these tools and uh
say or or sort of presume they have more
reliability than they actually do and
say that if it flags something as AI
generated that that's a definitive
answer. Um, and I think others who are
going to be uh far more skeptical. Um,
and that this will evolve over time. But
I I I worry that there's going to be
over reliance on the on the outputs of
these tools and people are going to take
them as sort of
uh some some sort of edict or or law or
definitive answer. Um, whereas what I
hope is that people sort of treat it as
this is important information that can
tell me more about this piece of text,
but it doesn't necessarily provide me a
definitive answer and that I should
really triangulate it using other tools
and techniques um and my my human
judgment before I make a decision about
um a piece of text.
>> Yeah. Well, I think that does it for
this week's episode. So, Aloque, thank
you as always for joining me, for
sharing your insights, and thank you to
our audience, for tuning in. Uh, we'll
see you next time.
>> It was great chatting with you. Thanks.
>> Likewise.
Thanks for listening to this episode of
the AI Policy Podcast. If you enjoyed
the show, consider leaving us a
five-star review on your favorite
podcast platform. [music] We'd also love
your feedback on the show. Please email
us at AI Policy Podcast@csis.org.
And don't forget to visit our website
csis.org for the Wadwani AI Center's
latest research and events. This podcast
was produced by Sarah Baker and Nicole
Herrera. See you next week.