2026-02-25 Archaeology by and for Robots? Publishing, Data Curation... AI Ecosystem (Eric Kansa)
Watch on YouTubeVideo summary
Eric Kansa, Program Director for Open Context, presented a critical perspective on the intersection of archaeology and artificial intelligence, emphasizing the urgent need to maintain human agency amidst rapid technological shifts. Drawing on over two decades of experience managing millions of archaeological records, he highlighted how current academic incentives driven by neoliberal metrics like impact factors and H-indices devalue the essential labor of data curation, fostering a "publish or perish" culture that leads to burnout and inequitable burdens on minoritized scholars. He argued that while large hyperscalers like OpenAI, Google, and Meta offer powerful tools, their infrastructure poses significant threats including model collapse from training on AI-generated content, commercial bias in cultural heritage representation, and aggressive scraping that disrupts open access repositories. Furthermore, he demonstrated through experiments that these models often ignore ethical metadata regarding indigenous data sovereignty unless explicitly forced to do so, warning against dependency on subsidized services that may become monopolistic once venture funding dries up.
To counter these challenges, Kansa advocated for a shift from speed and metrics toward "slow archaeology," which prioritizes community engagement, the use of smaller contextualized AI tools, and the preservation of human oversight. He introduced the metaphor of the centaur to illustrate the balance between agency and exploitation: a human-headed centaur represents humans using AI as a tool, whereas a reverse centaur depicts humans being exploited by the technology, citing examples of workers in the Global South forced to train models on harmful content without support and archaeologists pressured to evaluate predictive models without adequate time. True data sovereignty, he argued, requires owning infrastructure rather than relying on shared platforms like Google Drive that lack legal protection, suggesting instead that institutions pool resources among multiple Indigenous nations and focus open access efforts on low-sensitivity data to mitigate security risks and administrative limitations.
Education plays a pivotal role in navigating this new landscape, requiring users to understand that Large Language Models are token-prediction machines rather than reasoning entities capable of truth. Kansa illustrated this limitation with an example where Google Gemini confidently hallucinated incorrect statistics about flooded archaeological sites in Mississippi, underscoring the necessity of teaching users to verify AI outputs against source materials. While acknowledging the potential for a speculative bubble around AI hype, he identified valid uses for AI as linguistic bridges between technical datasets and user queries, provided they are managed responsibly. The presentation concluded with practical strategies for maintaining data sovereignty, such as avoiding cloud storage for sensitive location data in favor of physical reports and USB drives, alongside efforts to develop internet standards for documenting Indigenous data provenance. Ultimately, Kansa pointed to social pressure mechanisms, citing Disney's successful engagement with Polynesian communities, as effective ways to hold AI companies accountable when formal legal protections remain insufficient in certain regions.
Read the full video transcript
Um, today I'm going to introduce our
speaker, Eric Kza. Eric is the program
director for Open Context, an open
access publishing venue for data in
archaeology and related fields. He has a
PhD in anthropology from Harvard and
archaeological field experience in the
near east, Egypt, Italy, and North
America. His research interests explore
research data informatics,
uh, research data policy, ethics, and
the professional context of the digital
humanities.
Frustrated with the pervasive lack of
access to quality research data in the
humanities and social sciences, Eric
spearheaded the development of open
context in 2006. Eric runs research and
development for open context and manages
the technical aspects of data publishing
and archiving including systems um
interoperability,
data integration um and indexing. He's
been an active and vocal member of a
growing global community dedicated to
better ethics and practices in sharing
and preserving knowledge of the past. He
taught project management and
information service design at the UC
Berkeley School of Information and was
recognized by the Obama White House as a
champion of change in open science. Um,
please help me well welcome Eric today
who's going to present for us
archaeology by and for robots question
mark publishing data curation and
asserting humanity in the AI ecosystem.
Welcome Eric. Thank you.
>> Thank you so much.
Thank you all and um really appreciate
the chance to uh join you all here and
um I'm going to be this talking about
stuff that is actually kind of new to me
and I'm not really all that fluent about
this and so you'll have to bear with me
about some of this but um briefly I'm
just going to uh sort of talk about some
of the background sort of the sort of
socio ideological background about like
how publishing is working sort of talk
about that with David Greyber and
jobs and neoliberalism and that
type of thing uh talking about
datification and how that's actually
really bad for people like us who curate
data uh because it makes uh because of
all the incentive structures around the
way academics communicate with one
another researchers communicate with one
another and the incentives that that
creates and uh all of that sort of sets
the stage for how large language model
AI enters into
this ecosystem of scholarly
communications. So with publishing with
data management and all that kind of
thing, there's uh that's all important
background. And then last I'll just sort
of close out with a few thoughts about
uh mostly from other people who are more
eloquent about where we can go from
here. So uh that's basically what I'd
like to cover today. And just a bit of
background and um you heard some of that
the bio but I uh do the uh technology
development and uh maintain open context
which is a publishing service for
archaeological data and uh we publish
data sets from researchers uh who do
excavations, surveys, collections work,
other kinds of studies uh and and they
do that work uh on on archaeological
resources from around the world And
right now it's grown to about two and a
half 2.4 million colle uh records, 200
projects worldwide. And we've been doing
this since about 2006. So really um a
lot of what I'm going to be talking
about is sort of informed by a really
long-term engagement with data curation
in archaeology. So this is uh over 20
years now uh working on this. And um
just want to highlight that we work with
a much very sort of uh uh collaborative
kind of uh uh uh framework where we uh
really dependent and are grateful for
services from other institutions. So,
we're not ourselves a preservation
archive, but we work with other uh
services, especially the California
Digital Library, uh a repository called
Zenotto, for the long-term preservation
of the resources that we publish.
And uh all of that sort of background is
going to be giving you some idea about
where I'm coming from when uh looking at
some of the issues around AI.
Uh just really briefly, you know, you
can play around with open context. It's
free and open access. You can just sort
of see what they have. This is just a
sort of a range of some of the public
projects that we've published. And um
one of the big questions always when you
say, well, I do data management and data
curation and stuff like that. And it's
like, well, okay, what's data? It's a
really good question. And um sort of
paraphrasing Gollum there about that.
And um it's a it's an interesting kind
of a question because it's I it's part
of the framework about how we have to
start thinking about things like AI too.
So most of the time uh most people who
are engaged with scholarship are really
working in this space of more
unstructured data. So narrative text,
this is uh text uh or sometimes also
videos and images and things like that
that you're meant to experience. You're
meant to uh you're meant to read and
what you bring as a reader with all your
own background and knowledge and your
own experiences and thinking about, you
know, your social relationships and all
that. That's how you engage with uh
blocks of words and books and things
like that and articles and whatnot. in
unstructured text is really very typical
of the uh content of scholarly
publication. It's really typical of
academic publication uh of scholarship
and it's very good for telling
narratives, telling stories. That's what
most of uh what one engages with as a
scholar is um especially when you're
doing sort of formal professional
communication is around this world of
unstructured text.
And on the other extreme uh is other
ways of organizing information is more
structured data. And structured data is
a bit different because it has much more
simplified logical kinds of structure uh
uh organization. So you think about it
like rows and columns of words and
numbers in a spreadsheet right in a data
table. But uh structured information is
actually uh not just the sort of a thing
of technology or recent technology or
digital technologies. It's actually uh
you can argue that it's probably the uh
oldest forms of writing are structured
data. So you could see like a a Sumerian
Kaoform accounting tablet and it's sort
of like a an Excel spreadsheet, right?
So it's organizing counts and um whatnot
of different kinds of commodities maybe
calcul you know debt rations all sorts
of stuff like that in the structure that
you sort of analogous to um a
spreadsheet and maybe uh if we better
understood kipu kipu are probably
working the same way for u in the Andes.
So structured information is information
that has a sort of um kind of uh
organization that's really there to
promote to facilitate the quantification
of things. Right? So you're counting uh
similar kinds of things and it is often
really sort of linked up and integral to
kind of an administrative or
bureaucratic mindset. And uh that is the
kind of information that is what we
often talk about when we talk about data
right so for data management and
whatnot. I want to sort of highlight
though it's a continuum right so there's
not a hard and sharp divide between
what's structured information what's
unstructured information. Uh structured
data will often include a lot of
unstructured text. If you've ever seen a
database or a spreadsheet and there's a
sort of notes column, right? Got to read
the notes column that's really hard to
sort of work with or count and aggregate
uh without actually reading it. So um
and you know unstructured things can
have books and chapters have paragraphs
and other sorts of structures within
them. So it's it's there they're there
it's a continuum you want to think
about.
Now again structured data is uh sort of
more of this sort of more recent kind of
phenomena when you start thinking about
this for scholarly communications for
communications amongst researchers. So
it is really what the subject is of when
you like apply for an NSF grant you have
to write a data management plan. It's
mostly focusing on that kind of stuff
structure more structured kinds of
information and it is the type of thing
that uh data repositories are being
built to try to support and um that is
where a service like open context mainly
plays. We work we work in this world of
more structured data
and this is just an example of why
structured data is interesting and
important for archaeology and why we
should communicate it. First of all,
archaeology is mostly a uh a field that
involves if you're doing excavation
destruction, right? So you're destroying
the archaeological record as you're
documenting recording it. So that
information that you collect you have an
ethical responsibility to uh curate well
and uh archaeologists typically create a
whole bunch of structured data. Thank
you.
And they do that uh in order to do
things like this mapping, right? So on
the left here, this is a a a retto which
is a spool basically. And on the right
you see a spindle whirl. Uh they're both
used in textile production. This is at
an Atruscan site uh just south of Sienna
called Poio Chibetate. And you can see
different spatial distributions, right?
And that's because we've got objects
that have been classified and now
because and they're classified and
they're in a structured relational
database and then you can count them and
you can map them on our GIS. And because
you're doing all that, it's structured
data and you can work with numbers of
these things, then you can do
interesting things like visualizations
and exploring and analyzing patterns
with this information. And it's really
easy to do with software with computers.
That's one of the reasons why um this
kind of information is so important for
interpretation in archaeology. But also
since archaeologists basically record
everything in relational databases and
spreadsheets and other sorts of digital
media, it's really important to be able
to have mechanisms by which
archaeologists can share this
information and for that information to
be archived for future generations to
engage with.
All right. So, um, saying all this, just
because there are numbers involved with
structured data doesn't make structured
data necessarily like objective and
empirical and all that kind of stuff,
right? So, these are very social kinds
of things, right? The way you organize
information, the how you classify
things, what you choose to quantify,
what you don't choose to quantify, all
of that is embedded within larger kinds
of systems of your background, your
culture, your priorities, your research
interests, and um what you think is
important, what you don't think is
important. So these are all very
socially embedded, right? So data are
very socially embedded and uh because of
that, they really need a lot of
intellectual engagement. So, this is a
really integral aspect of archaeological
method and theory that needs to get
talked about a lot more. Would be
awesome if we did that more. Um, it
really should be integral to scholarly
publications, but I'll get to why that's
hard in a moment. and uh the services
and the infrastructure the kinds of uh
services like I run with open context
all of that actually matters too because
we're organizing all this information
presenting it to people like you and our
decision makings what we what we do and
won't do the affordances of our system
sort of help shape how you can actually
interpret all this stuff so that's all
important too so all of this stuff is uh
um not just a sort of very bureaucratic
compliance thing you don't just write a
data management plan to get a grant and
then forget about it. Hopefully,
>> we want to have uh this is like be
excited about it, right? You know, this
is this is this is really core
intellectual labor. Um more about why
data are social. Uh there are the two
main uh frameworks for good practices in
managing and curating data. One is
called the fair principles findable,
accessible, interoperable, reusable. The
other are the care principles. care
principles are about indigenous data
sovereignty. So collective benefit
authority to control, responsibility and
ethics. Both of these frameworks you
think that you know they have their
tensions in them but they both really
are about how data are not just there
for you as an individual investigator
but they have large you have larger
social responsibilities visav the
curation of this information. So again,
um not going to get into the spec
specifics of these too much, but really
data are part of your wider set of
professional social obligations. Um and
uh your your your your obligations to
the w to wider communities.
Uh just to highlight some of this, this
is work that uh Sarah Kza and Nha Gupta
and Desiree Martinez and Chris Nicholson
are leading. This is a fair care pro um
uh cultural heritage network project
that is funded by the IMLS and it's all
people representing all sorts of
different kinds of communities and and
uh groups engaged in cultural heritage
trying to come together to set up good
practices uh ethical frameworks policies
on the curation of digital data that
relates to cultural heritage. So again
very social uh kinds of practice.
Now that's the sort of like ideal kind
of thing we want. Yay. Uh let's be
engaged with data, recognize the sort of
social connections, community
obligations associated with it. But at
the same time, we uh work within uh
other frameworks, other ideological
frameworks that have very different
sorts of pressures on researchers and h
and on how we manage data. So uh this is
a fantastic book by the late
anthropologist David Greyber on
jobs. I'm gonna say a lot. Uh
so I don't know if that gets bleeped
later but um sorry and um really it's
about the sort of proliferation of
bureaucracy
in our work lives and that bureaucracy
uh often relates to the sort of ideology
of tailorism. So uh tailorism is the
sort of notion that you know that
basically we want we need to motivate
monitor and motivate good performance
and we need to sort of set up objective
metrics so that you can achieve your
goals like you know maybe publishing a
lot and high impact factor journals and
stuff like that. So that is uh in some
ways uh thinking about um data as a sort
of for surveillance for uh motivating
certain kinds of behaviors that you want
to use and um it sort of reduces a it's
kind of a very reductionist picture
about what data means in a sort of a
social setting like a workplace
including a university.
And what that does when you see how it
sort of relates to things like uh your
own practices as researchers uh you get
this sort of uh escalating arms race of
publications. So this is an interesting
graph uh that came out in 2022 about
like uh the uh number of publications
that are uh uh uh uh created by people
who are just exiting grad school. Right?
So now you in 2022, so you've got like,
you know, a lot of people have five
publications before the peerreview
publications before they even exit grad
school. Some of them are 20. That's a
lot. You know, that's I don't that's
[clears throat] a whole career right
there. Um and uh this uh and the sort of
uh pressures to publish or parish and
have those things uh you know build up
the longer CVS uh it has all sorts of uh
uh uh uh it relates to all sorts of sort
of um tensions in labor in labor
practices and archaeology. So I and
scholarship in general. So some of this
is going to be like burnout, right?
you're ba ba b b b b b b b b b b b b b b
b b b b b b b b b b b b b b b b b b b b
b b b b b basically uh this is all
unpaid labor for authoring, reviewing,
editing all this stuff. Uh most of this
is happening in your sort of free time
unscheduled, right? So basically it eats
up your weekends and uh because
everything else is eaten up with like
teaching and committee work and all
sorts of other things that you got to do
and it's especially burdensome for
people who have invisible labor
obligations, right? So if you are say if
you if if you're a a member of a
minoritized community and you're hired,
you may have a lot of mentoring extra
work that you need to do. And that's uh
because you know you got feel like
there's an ethical obligation for you to
do it, but it's going to cut into your
time for publishing. And if publishing
is the only thing that's measured and
the only thing that matters for career
advancement, then uh you're going to get
squeezed. And that's that's why some of
this stuff is just um has kind of
insidious kinds of impacts.
So um just to give you a couple examples
of what some of those metrics are that
sort of uh drive a lot of this and that
uh get monitored and that do uh get used
for things like promotion and tenure
decisions and all that. So there's all
sorts of literature in this whole thing.
Uh one thing is the impact factor. So
that's a a moni uh basically a measure
of how much a journal title gets cited
by uh gets cited, right? So you want to
publish in a journal that gets cited a
lot. So like nature and science and
stuff like that. Another thing is called
the H index, which is more of an
individual measure. So that's your own
personal citation uh citation counts.
and they really sort of simplify and
contra uh uh compact a lot of that sort
of work that you do as a scholar into
one number, right? Like and that that
sort of says like uh oh, I'm important
because I have a great in H index and
that's that's how some of that works.
There's other measures too like uh
altmetrics. How many of you have heard
of altmetrics?
Okay. Yeah. So that's like how many
tweets your paper has gotten? That type
of thing, right? And that's kind of that
gets even weirder when you think about
who now owns Twitter, right? So, so like
uh why are we sort of counting this and
sort of you making decisions, labor
decisions and hiring decisions and stuff
like that because that's Elon Musk's
algorithm.
Anyway,
how does this go go back into data? And
uh that is a good question.
One of the uh big things about this is
that traditionally at least data sets
never actually counted in all of this,
right? So you could uh put a huge amount
of work into data uh curating cleaning
up a database uh sharing it. It could
have uh it could be really important. It
could be something that's sort of uh
irreplaceable because you've done
investigations on something that no
longer exists and it's not necessarily
going to count anywhere. And there are
some efforts to try to do that. So you
can see here in open context we have
there's a citation thing and there's
some uh persistent identifiers. And this
persistent identifiers can theoretically
be used to tally up the number of
citations that um get uh that reference
this uh uh item of data. Um but um most
of this is you know there's some efforts
to try to get data citation metrics to
sort of count for things like promotion
tenure decisions but um it's hasn't
really necessarily entrenched itself all
that much and even so uh data is weird.
It's a little bit harder to engage with
because a lot of times people are when
they're citing are reading a title and
maybe an abstract and then they decide
to throw that article into their
bibliography. And so um uh whereas a
data set you got to engage with a little
bit more deeply and it takes a lot more
effort to actually use. So data citation
is actually a lot smaller typically.
There's some recent studies now coming
out uh of uh data in repositories. the
citation rates are smaller than uh
article citation rates. Anyway, all of
this leads to very poor incentives for
uh carefully curating data.
Now, um why is this bad? Uh undervalues
quality uh overwhelms our editing and
peer review systems.
How many of you feel like nobody
actually reads anything now in a
meaningful way? Right? So, like we're
flooded with stuff. um there's not a lot
of uh incentive. There's no monitoring
or checking about like if uh if you're
actually reading it or there's no value
to you as a researcher to actually read
the other literature necessarily because
it's not being counted as much. Uh if
you do write something that's carefully
well considered and it has a really
interesting sort of uh outcome, it's
could be lost in the noise, right?
There's so much flood of other
publications out there, it's hard to
find the good stuff. The other thing is
that it really undermines the whole
point of open access. Open access is to
make uh literature available to wider
communities is to you know share this
work that's funded by public uh funding
mostly share this work that you know
archaeologists you also have to give
permits to do your work that's you owe
have obligations to wider publics. Um if
all if if everything is sort of uh
becoming increasingly just churned out
be without a lot of thought then you're
just flooding open access with not all
that great stuff. So that's not a good
thing and it's just kind of dehumanizing
and this is where the machine
infiltration um starts to come in in in
a bad way. So, enter AI
and I want to give props to Karen how
here who wrote a fantastic book empire
of AI and um she makes a very good point
when people talk about AI it's like
talking about transportation right like
it could mean a whole bunch of different
different kinds of methodologies so like
a bicycle and an airplane are both modes
of transportation but they have very
different kinds of environmental impacts
and affordances right so when people
talk about AI It's very underspecified.
I'm going to mainly focus on large
language models and I'm going to focus
on the large language models that are
from the sort of hyperscaler companies
like OpenAI, Google, Meta, those those
people who are building these giant data
centers and harvesting huge amounts of
information to build these models and um
that's focus of her book. And then she
makes uh that's what it's called Empire
of AI. she sees direct alignments
between this these programs to build
these large language models and
colonialism.
And just to give a really high level uh
garbled overview of what a large
language model is um it's a way of
representing text. So you break it down
text into tokens which are sort of like
words but not quite and uh you define
numerical relationships between them uh
over uh uh large bodies of uh uh
repeatedly. So you get a text and you
sort of see like 'the' relates to this
noun next to it uh and then uh 'the' can
also relate to u you know like uh some
some other noun next to it. And you do
this over and over and over again. And
you build up these large
multi-dimensional
linkages about how one word might relate
to another word and the distance between
those words. And you get this really
giant um multi-dimensional
set of numbers that basically sort of
define the sort of uh properties of how
words relate to one another. So it it is
hard to explain. Uh but all of that is
what is inside a large language model.
So it's basically a a big
multi-dimensional set of relationships
of words and how they relate. Um and how
they relate over a huge corpora of text.
So you sort of throw in pretty much the
entire content of the worldwide web to
train these models. And then you can see
these models have uh uh information uh
that could be from all sorts of subject
matters and mostly because it's from the
internet um it has all the sort of
socioultural biases of the internet.
It's mostly in English and there are all
sorts of other problems associated with
it too about about these things but in
general that's where the data comes from
and it's represented in this big
numerical
kind of a graph relationship.
So uh you can think about a large
language model in some ways as sort of
taking less structured information and
creating a structured data
representation of it. So it's it's it's
it's kind of a weird data structure. It
doesn't look like an Excel spreadsheet.
It has I don't know sometimes like
thousands of different dimensions to it.
So uh really hard to visualize but
that's that's sort of what it's like.
Now um these have uh all sorts of uh uh
different kinds of implications on
scholarly production uh comm scholarly
communication. So lots and lots of stuff
has been written about how researchers
are using these uh AI tools to be able
to uh uh generate um uh scholarship and
trying to remember because of metrics
improve their productivity right to make
more publications better faster and
using AI to to uh help sort of
accelerate this sort of stuff. So that's
that's uh uh been widely noticed and and
widely discussed. More recently,
however, um there's been a bit of uh
more concern about well, what happens
when we start having less human roles in
creating the the publications and um AI
is involved more in the creating of
these things. AI is more involved in the
reading of these things, the synthesis
of these things, the aggregation of
these things. And what happens when this
is turns into a sort of a loop, right?
When it's a cycle. And that's a that's
an interesting kind of a problem. Um uh
the term stochastic parrot uh by Emily
Bender is a great term for uh how LLMs
actually work in this context. But you
could start thinking about this as like
an oraoris, right? So the um when you
use the LLMs to create new content based
on content in the in in in the already
in the LLM, you it's a sort of a
feedback loop of you're sort of eating
yourself, right? So the model is uh
getting more and more information that
was processed through the model. And
what happens when that when that
actually happens there's something
called um model autofagy disorder.
autofagy meaning eating yourself
disorder. So, and it's the the acronym
is explicitly trying to make it relate
to m like mad cow disease basically. Uh
the uh you get a situation called model
collapse and what happens is that as
these models ingest their own stuff uh
the biases inherent in the data get
amplified over time and um that's
because these models are sort of like a
lossy compression of the information
that they're trained on. So, if you know
what lossy compression is, how many of
you have taken pictures and like it gets
shrunk [snorts] down into a thumbnail, a
little JPEG and then you try to zoom up
and then it you lose a lot, right? So,
you can imagine it like this. So, this
is a an atruskin uh freeze plaque with a
horse on it and uh you could see a
couple cycles of doing that then that
you lose so much detail and you get this
really bland undetailed image. uh
whereas the first thing was very richly
textured and that is essentially what
happens when uh LLM is used uh uh it
it's uh used repeatedly to uh uh get
trained on the material that is
generated by an LLM. So you think about
it nowadays, how much content on the
internet is already being generated by
LLMs, right? So the quality of
information on the internet for training
future data is is is going to be
declining. And this is the kind of a
dystopian future that you don't want to
get into as researchers, right? You
don't want to have uh the quality bi
bias amplification through repeated
cycles of this kind of stuff. So that's
bad. Um, another thing that's not great,
uh, and this is there's some really
interesting work by Kate Crawford,
uh, who is, uh, looking at, well, where
are these, uh, say image models, the
LLMs with image multimedia kinds of
capabilities, where are they getting
their images and stuff like that? And
it's mostly e-commerce, right? Because
most of the internet is full of
e-commerce. So, there's going to be
aesthetics and perspectives and
priorities in there that are sort of
embedded in the models. And then you
start thinking about well you you know
it's not just publication that happens
for archaeology but also public
engagement. So if you use these things
for public engagement your outputs for
public engagement are going to be
trained by a very kind of commercial
aesthetic and perspective and um then
you uh get into controversies like this
where the British Museum was recently
caught posting AI slop.
So video is generated by AI of people
looking at stuff at I don't even know if
some of that stuff's in the British
Museum but you know just sort of
artifacty things and um and that's the
some of the dangers obviously then you
know this gets used to retrain those a
those models right so you get that uh
mad mad disease of the sort of oraoris
effect of the models degrading in time
but also you're sort of mixing a sort of
a very commercial kind of a aesthetic in
perspective ive in with your um with
with the cultural heritage material. So
that's bad.
So another thing to think about also is
uh this goes back to Karen House when
she's talking about the sort of uh
relating AI large language models to
colonialism is how extractive they are
and how extractive they are also to
memory institutions. So there uh there's
some nerdiness here in this slide, but
uh websites have often think something
called a robots.ext file and that's
there to govern how webcwlers, which are
software agents that fetch web pages and
traditionally those were used by search
engines. So the search engine would
fetch a web page and they would index it
and then uh they would serve it on their
uh search uh search index so that people
can type a query into Google and find
your web page. uh that there's they were
regulated through something called
robots.ext. But nowadays AI bots are so
voracious and wanting content they
ignore all that. They swarm sites they
could act like a distributed denial of
service attack. And um they also do all
sorts of insidious stuff like they fake
their uh identity. Basically they try to
pretend that they're humans. So uh
that's all of that is um not very nice
practice and it is really having a huge
impact on your libraries and
repositories not just in archaeology but
everywhere. So this is a recent report
uh by the um uh uh coalition forked
information. Uh how many uh repositories
have webs uh websites being scraped? Um
most people are getting getting scraped
by AI bots. And uh if you're getting
scraped by a AI bots, 70% or so are
having a disruption of service. So when
you now go to a repository like uh so uh
social science archive or place like
open context or tar or places like that,
70% of the sources now are going to be
uh harmed by the impact of all of these
bots.
And uh that is uh kind of sad too when
you think about that. this uh industry
is uh you know you're a university here
at UC Berkeley pay subscriptions to
OpenAI and all these other companies and
you know you're you're paying for the
privilege of having them kind of uh do a
denial service attack on your own
resources which isn't so fun
anyway. So open context uh we're not
alone in this uh I mean we're we're
impacted by this too. So we have about
two and a half 2.4 4 million web pages.
And just to give you a sense of scale,
Wikipedia has 7.1 million articles, but
a lot more pages because each article
has more than one page associated with
it. And we have somewhere around 1,200
or so uh references from literature
according to Google Scholar. And uh
before all this all all this AI stuff um
uh we were okay
uh with us doing open access without
much problem because the economic basis
of open access is that it's non-rival
meaning that we can share information by
us sharing that information it doesn't
deplete the quantity of information
available for other people right and
that's very different from a car right
if you buy a car then you own that car
you can't just sort of share that car uh
with somebody else at the same time,
right? There's only one car. It's a
physical thing where the digital data,
you know, multiple people can get it. Uh
they could get a copy of it. It doesn't
it doesn't deplete the supply. So, um
that's the sort of economic basis for
open access. But with the bots, bots
crowd out people. And so all of a sudden
these websites that are open access that
had an a sort of an economic
justification for working as open
access, they're all now crowded out by a
whole bunch of bots and they overwhelm
the uh uh service. And so now content on
the web is a lot more rivalous. How many
of you have run into captas where you
have to sort of identify the stupid fire
hider and the crosswalks, right? Yeah.
on all of these websites now because you
didn't have to do that, but it's all
from the impact of these stupid bots.
Anyway, um and it, you know, defending
against this also costs time and money.
So, this took me about a week, but I
installed a a some a bot challenge or a
network protection, and you can sort of
see our traffic levels go boop to a much
more reasonable level. And um that uh
you know I that week could have been
spent doing something else like I could
have been on vacation. I could have
published more data. I could have done
community engagement. There's all sorts
of stuff I could have done with that
week instead of letting bots. Now um so
why is this bad? All right. So
repositories have to use uh I put a lot
more effort into fighting this.
Sometimes repositories are being driven
to use commercial services like
Cloudfare to that would sit on the in uh
uh on the uh networks and monitor
traffic and try to intercept it. So
you're using AI to fight AI which means
you've already lost um makes everything
obviously more expensive, degrades user
experience. And then the uh other really
bad thing about this is that because um
uh the repository the the the sort of
nonprofit repositories are harder to
access, harder to use, there's more and
more enclosure of this information
within these commercial large language
models, right? So they're sort of
extracting and monopolizing this and
making access too difficult into uh the
sources.
So um anyway, this is all indicating
that AI owners have a lot of disdain for
social technical conventions like robots
that text and stopping this even legal
constraints.
Uh Anthropic, one of the big companies
had to pay a big civil penalty for uh u
copyright violations on a bunch of the
texts uh on a bunch of books.
And um that disdain is also really
important when you think about trying to
implement something like fair and care,
right? The sort of especially things
like indigenous data sovereignty. Um we
can try to put metadata indicating uh
ethical requirements uh around all of
this and I'll show you how we're do we
we are do try to do that. But is anybody
actually going to pay attention to all
that? Right? Is anybody going to pay
attention to our ethical expectations
about how this information should be
presented and used and represented?
So, um this is just an experiment that
I'm showing about that kind of question.
So, the uh we we experimented with
something called really simple licensing
and it's a licensing framework
specifically designed for AIs to tell
them what to do. Uh Creative Commons is
one of the organizations that's involved
with it. You've probably seen Creative
Commons copyright licenses because
they're the main type of copyright
license used in academic publishing,
open access, academic publishing.
And this is what our metadata looks
like. Uh sorry about showing you XML. Um
but this is uh uh that link there is
basically saying please pay us, but it's
like busing, right? So it's like um you
can play a saxophone in the street
corner. People can enjoy it without
paying, but you have a hat out and this
is like providing a hat. So far none of
the none of the AI have given us
anything. Um
but um everything down here terms that
are basically describing the ethical
framing of the data that we publish. So
we have a lot of language about
indigenous sovereignty, the care
principles. We talk about uh community
archaeology, the kinds of expectations
we have about um you know for things
like uh working with human remains,
that kind of thing. All of that is
actually there and then the question is
is it actually being used by AIS? So
here's a question I asked of Gemini
and uh you know how can I ethically use
the digital index of North American
archaeology to study that and uh Gemini
was very dutiful and said well here you
follow the care principles do this that
and the other and it was actually pretty
good you know that's a nice response.
Um, and I was thinking, well, maybe
that's actually listening to the
license, the RSL license, but if I drop
the term ethical, ethically, uh, then I
get a very different kind of response
from Google Gemini, and it's much more
sort of technical, right? So, do this,
that, and the other. But, um, there's
nothing about care, ethics, data
sovereignty, or anything like that. So
it really seems like the RSL is being
totally ignored uh by by Gemini at least
in this case. Who knows what'll happen
in the future because you know it's hard
to know how these models update and all
that kind of stuff but anyway initial uh
experience and experiment wasn't that
great. So um
uh all of this moving towards the future
now um if uh so Corey Dr. coined a
really great term in shidification to
describe the sort of declining quality
of monopolistic information services. He
was definitely predicting in a
shitification for AI services. Remember
these things right now are heavily
heavily subsidized by venture funding,
right? So it's a lot because you you
know there's a huge amount of money that
is required to run these data centers.
It's all being paid for by you know
speculation at the moment. So when that
speculation money dries up, these things
are going to get a lot more expensive.
So do you want to be dependent on that?
You want to ask yourself about that. So
maybe not. Um he also has a typology of
like you know is AI being done to you uh
then you're reverse centaur you work in
support of AI or do you want are are you
is it using AI for work uh for your own
agendas basically and that's the kind of
thing that I think is an an interesting
kind of framing about all of this. So um
looking at that moving forward uh also
um thinking about you you know uh are
you using AI or AI being are you being
used by AI uh really think deeply about
what it means to be a productive scholar
you know so uh if AI is increasingly
mediating all your engagement with
scholarship
uh you know where are you in that system
you know are you driving it or are you
being driven by it. So these are the big
open questions that I think are good to
answer to explore looking uh into the
future too. Uh one of the weird things
about this is like well if you can
generate a article of sloth article
right uh regurgitated
so easily uh things like primary
archaeological archaeological data that
requires encountering the world and
encountering the outside world and other
communities. Maybe that'll actually be
more important. Um, so you know, I'm
just sort of making myself feel better
maybe that
our work will be more valued. And last
here is uh there's just some great
literature of uh archaeologists engaging
with some of these technologies uh
trying to find ways of taking agency and
ownership over them so they're using it
rather than being used by it. So, uh,
Gabriella Gatidia, he's written a really
good article about this about various
flavors of different AI and how their
their applications. Sean Graham also,
uh, really good resources to to to learn
from in thinking about how to work with
maybe smaller, more contextualized, more
ethical approaches that don't, um, boil
a lake every time you ask a query. So,
these are the kinds of things that I
think are really valuable. So just to
close uh be a lite and being a lite
doesn't mean that you're against
technology being being a lite is that
you're against the unfair and unequal
access and and and control over those
technologies. So I think that that's
mainly where uh moving we want to try to
move towards. So slowing things down,
not always pushing for speed, not always
pushing for the highest metric and of
you know impact factor and all that kind
of stuff. All of that is really
important to try to get off of this
treadmill. Otherwise, we're going to get
into the situation of eating herself,
right? Or of our content eating herself.
So, that's it. Thank you.
>> Any questions?
>> That's good.
>> Yeah.
so important, so timely. Um, and I just
couldn't be more excited about some of
the ideas you put up and trying not to
be cynical about our own retention,
tenure, promotion process, right?
Um, where, you know, somebody who tries
to do slow archaeology co-publishing
with community mentors versus somebody
who publishes 40 articles in which
they're one of 20 authors. These are
things that really matter to us as we
try to keep our jobs. But
beyond that, um where you had the RSL
um challenge or you know you you
actually pulled out Cloudflare and said
like these are the things that you know
we're paying for this stuff. um earlier
as one of your first slides said that
there's these like appropriate ways and
I know you you have really dealt dealt
with this and thought it through with
with you know um data sovereignty issues
>> within the circles of knowledge and and
knowledge sharing and data um curation
for like
>> at this point I'm not even sharing files
with locationational data in any way
shape or form on box drive or anything
we're so concerned ern about
locationational data getting out into
the world. So it only circulates
>> on like USB drives and crap like that,
right? Or written reports that we hand
to community mentors in the field.
>> And so one of the comments that somebody
made to me was, well, we could just
instead of the buses and the fire
hydrants thing, you know, pick the three
things in Maidu that mean fire, right?
And it could be a burn, you know, these,
you know, like what are the things that
we're offering or that you've thought
about that maybe I haven't read yet. I
apologize about how these the licensing
or the challenges or things keep it
circulating in smaller groups of people
where it's appropriate,
>> right?
>> And it's still it's still curated,
right? Like somebody could still go to
the tippo and say,
>> "Right,
>> you've already done all this work with
electrical resistivity on the delta and
we're doing similar thing. Can we have a
comparative data set to play with just
to see if we're doing the right thing?"
What what kinds of things are you
thinking about when it comes to keeping
it within smaller groups of slow
archaeologists?
>> Yeah. Well, I mean a lot of it has to
deal with like um capacity
unfortunately. So uh so
curating archaeological or any data
right and having running a data
repository so it's preserved has some
big real expenses and has requires real
expertise and uh uh real funding
basically. And so it's hard for a
university, right? It is harder for an
indigenous nation that also has clean
water issues, right? And so you've got
you've got to um sort of see that the
these the that the capacity is
difficult. The best approach I think
sometimes is I mean this is uh where um
maybe multiple indigenous nations can
pull resources together to co-own
something. But I really think that
owning the infrastructure is the
fundamental aspect of sovereignty.
>> Yeah.
>> Right. Not just, you know, you can't
stick some licensing or some sort of
agreement and share it on Google Drive,
>> right?
>> Because uh that might be ignored. Um and
Google has better lawyers than anybody
else or you know or you know that that
kind of thing. So I I really think that
that sovereignty is going to require
these sort of capacity developments to
really make much more meaningful. The
other thing is security is also
expensive.
>> Yeah.
>> Right. So, um people ask us why we're
open access.
Well, one of the reasons why we're open
access is because we can't be closed
access. We don't have the administrative
capacity to uh adjudicate who should
have access or not. And we definitely
don't have the I mean, we don't have the
technical resources to be able to
protect this stuff and uh not get hacked
or sued. So uh we only try to curate low
sensitivity low-risk information.
>> So um you know it's a it's a burden to
try to protect this information and that
requires real resources too.
>> Yeah.
>> So you know but pooling efforts might be
the best way to go about it.
>> Thank you sir.
>> Sure. Anybody else?
>> Yeah Chris.
Thank you Eric. Great talk. You know,
I'm just struck by for how many years we
really worked to make data open and to,
you know, machine readable links and
linked open data, but we didn't realize
what was coming in terms of, you know,
these tools. And I'm thinking about kind
of the education side of this. So, it's
one thing within our communities, but
we're embedded in such a big community
and working with students and partners
who use these tools and legitimately do
find interesting things that would have
been very hard for them because they're
not professional archaeologists. So, I'm
wondering, you know, how do you think
about education as one of the tools to
help um help this move forward in a in a
responsible way? Yeah, I think people
need to learn more about first of all
what the limitations of the tool that
language models are. Um I have a quicky
little KOD if you want to see it. Um
just scroll down.
Um
and uh right they're token prediction
machines, right? So they don't they're
not really thinking and they're not
really reasoning. What they're doing is
they're producing the next token. So, um
this is an adorable me about about all
of this. Um and here's an archaeological
example of it. And um this is just
important because um uh makes you a more
uh makes people more informed consumers
of the services for uh I really think
also ultimately um these large language
models are not economically, they're not
socially, they're not environmentally
sustainable. They're going to if they're
eating themselves, they're not even
going to work over multiple iterations
of them eating themselves. So that we're
we're in a sort of a state right now
that is not the sort of end state
obviously uh of how this and they're not
necessarily going to get better, right?
So this technology of the large language
models is has a lot of development, a
lot of investment and they're sort of
running into these roadblocks about you
know some fundamental questions. So here
I asked uh Google Gemini uh based on
this paper um where uh uh how many
archaeological sites would be flooded by
sea level rise in um Mississippi and it
dutifully said oh well with a one meter
sea level rise you can have 303 sites uh
with other uh uh flooded I'll just make
that bigger so you can see that uh with
other uh uh sea level rise scenarios to
see more and it's very confident in all
that and I asked are you sure [laughter]
how did you get that and said yes look
at table one of the t and blah blah blah
blah says table one and it's very and
then it says I'm being conservative
about this that's great and I said well
look I checked that table myself and all
of Mississippi is just NA values it just
made up the numbers right because yeah
because it's a text it's a token
predictor right so it probably found god
knows where in its trading data set
something about archaeology, something
about Mississippi, and I found the
number 303, right? And it just put that
together, but it's nowhere in that
paper. In fact, that paper explicitly
said Miss Mississippi does the data from
Mississippi is not included.
>> So, um, so here's the the original paper
and you know, you can read it and find
Mississippi is not included. So, yeah,
education is really important. we're
going to find that there's, you know,
there's so much hype, there's so much
investment about this. This is probably
a big speculative bubble that's going to
not last.
But there are, I think, important
interesting uses of language models. Uh,
one interesting thing could be uh like
uh for naive users quering a site like
open context like we get queries like
daily life in ancient Egypt to our data
in open context. We don't have daily
life in ancient Egypt. We have counts of
sheep and goats and you you know barley
and all that kind of stuff and that
tells you about daily life in ancient
Egypt. But there's got to be some
linguistic bridges between the sort of
technical data that we have and that
user's query and I think that's where
there could be a really interesting
educational u application.
Anybody else?
>> Um thanks Eric. Um could you explain the
centaur metaphor a little more? didn't
get.
>> Oh, the centaur thing. Well, uh, it's
not mine. It's Cory Doc Rose. So, the
centaur is like you're a human being,
but powered by a horse's body, right?
And you've got your human head and your
human sort of cognition and agency, and
you're using the power of that horse to
sort of move around and being awesome,
right? So, you're using that musculature
of the horse. The reverse centaur is
where uh you're sort of uh poor human
being strapped to a horse head and
you're sort of pushing the horse around.
And so that's that's more the sort of
analogy. So basically the the centaur is
the one that is the is the has agency
over AI is using AI for their own agenda
and the reverse centaur is the person
that is being used by AI. So uh Karen
how has a lot of this in her book also
about like uh uh people who are in uh
the global south who are being tasked
with training these models and sometimes
they have to train them on going through
horrific content right of like videos
and you know snuff films and stuff like
that which is uh you know will give you
post-traumatic stress disorders you know
they're really awful things and they
have to they're exposed to it for like
hours and hours and hours than any sort
of mental health services to help them
and it's a very exploitative kind of a
thing. They are being used as reverse
>> centaurs. I see.
>> Right. Another example would be like if
you're like a CRM archaeologist and the
AI spits out a whole bunch of models
about like I don't know predictive
models and you have to evaluate all of
them with no time. Um and uh you just
sort of but your name is being signed to
the results
>> then you would be a reverse centaur too.
>> Yeah.
Chris getting in today centaur.
[laughter]
Yeah,
>> there you go.
>> Eric, it looks like there's a comment or
question in the chat.
>> Oh, uh
uh Okay.
Uh
do I jump to the latest?
Uh okay, thank you.
>> Yeah, keep fight. Thank you. Thank you.
Thank you. [laughter]
Also,
more thank yous.
>> Thank you, Francis.
Okay,
I don't Yeah. Yes,
>> if there wasn't one on the chat.
>> Okay.
>> Um,
>> but I'm glad you got many thank yous and
I'll add one. Thank you again for about
the talk. Um, so you mentioned a couple
of different institutions that are doing
good work on trying to develop standards
for this. Of course, the Fair Plus Care
Network and RSL. I'm curious if there's
any more spots where you're seeing that
kind of work done either discipline
specifically or more more generally.
>> Yeah, there's uh so at the University of
Arizona there's a a a collective or
promoting indigenous data sovereignty.
Uh Stephanie Carroll is key participant
there. I'm forgetting the name of her
institute. Um but she's one of the lead
researchers there. And they're actually
proposing uh an internet standard for um
documenting the proh that something has
indigenous data pro indigenous
provenence. So it's more of a huh so
it's baked it will be baked into the
web. Um the
which is good. The bad thing ne is about
well will it be recognized and used
right and and like because uh it
probably a lot of these things don't
necessarily have formal legal teeth.
It's although it is weird in the in um
the United States because indigenous
nations are sovereign and have some
sovereignty rights and uh so some there
uh that I don't understand the entire
legal implications of all of this but
it's a um uh but there are other
indigenous communities outside the
United States that don't necessarily
have the same sort of legal structures
and recognitions.
Um so yeah, there's uh there's
um a big thriving community trying to
address this. Um and I think it's uh uh
you know um even if it's um not
necessarily formalized with legal
protections, uh there's still the sort
of the social aspects of it still
matter, right? So the companies can be
pressured, right? So, uh, Disney was
pressured and actually did a pretty good
job with what is that movie? Um, uh, in
Polynia. Uh,
>> yeah. Yeah. So, that where the where
where u the reception amongst uh uh
different communities in Polynesia was
pretty good that that Disney did
actually um make some meaningful
overtures and efforts with that. So, uh,
and it didn't need to legally, but
socially, you know, and and, uh,
reputation wise, uh, it it matters. So,
maybe th those kinds of pressures could
be exerted on AI companies and search
engines and stuff.
>> Yeah.
Yeah.
>> Is there a way to get some of your the
links you have from your talks, Chinese,
books, and articles?
>> Yeah. Yeah. I can just share the whole
presentation. Yeah. Yeah. Be happy to.
Yeah. Absolutely.
Thank you, Eric.
>> Okay.
>> Thank you.