Submind YouTube summaries
Thumbnail for 2026-02-25 Archaeology by and for Robots? Publishing, Data Curation... AI Ecosystem (Eric Kansa)

2026-02-25 Archaeology by and for Robots? Publishing, Data Curation... AI Ecosystem (Eric Kansa)

Watch on YouTube

Video summary

Eric Kansa, Program Director for Open Context, presented a critical perspective on the intersection of archaeology and artificial intelligence, emphasizing the urgent need to maintain human agency amidst rapid technological shifts. Drawing on over two decades of experience managing millions of archaeological records, he highlighted how current academic incentives driven by neoliberal metrics like impact factors and H-indices devalue the essential labor of data curation, fostering a "publish or perish" culture that leads to burnout and inequitable burdens on minoritized scholars. He argued that while large hyperscalers like OpenAI, Google, and Meta offer powerful tools, their infrastructure poses significant threats including model collapse from training on AI-generated content, commercial bias in cultural heritage representation, and aggressive scraping that disrupts open access repositories. Furthermore, he demonstrated through experiments that these models often ignore ethical metadata regarding indigenous data sovereignty unless explicitly forced to do so, warning against dependency on subsidized services that may become monopolistic once venture funding dries up. To counter these challenges, Kansa advocated for a shift from speed and metrics toward "slow archaeology," which prioritizes community engagement, the use of smaller contextualized AI tools, and the preservation of human oversight. He introduced the metaphor of the centaur to illustrate the balance between agency and exploitation: a human-headed centaur represents humans using AI as a tool, whereas a reverse centaur depicts humans being exploited by the technology, citing examples of workers in the Global South forced to train models on harmful content without support and archaeologists pressured to evaluate predictive models without adequate time. True data sovereignty, he argued, requires owning infrastructure rather than relying on shared platforms like Google Drive that lack legal protection, suggesting instead that institutions pool resources among multiple Indigenous nations and focus open access efforts on low-sensitivity data to mitigate security risks and administrative limitations. Education plays a pivotal role in navigating this new landscape, requiring users to understand that Large Language Models are token-prediction machines rather than reasoning entities capable of truth. Kansa illustrated this limitation with an example where Google Gemini confidently hallucinated incorrect statistics about flooded archaeological sites in Mississippi, underscoring the necessity of teaching users to verify AI outputs against source materials. While acknowledging the potential for a speculative bubble around AI hype, he identified valid uses for AI as linguistic bridges between technical datasets and user queries, provided they are managed responsibly. The presentation concluded with practical strategies for maintaining data sovereignty, such as avoiding cloud storage for sensitive location data in favor of physical reports and USB drives, alongside efforts to develop internet standards for documenting Indigenous data provenance. Ultimately, Kansa pointed to social pressure mechanisms, citing Disney's successful engagement with Polynesian communities, as effective ways to hold AI companies accountable when formal legal protections remain insufficient in certain regions.
Read the full video transcript
Um, today I'm going to introduce our speaker, Eric Kza. Eric is the program director for Open Context, an open access publishing venue for data in archaeology and related fields. He has a PhD in anthropology from Harvard and archaeological field experience in the near east, Egypt, Italy, and North America. His research interests explore research data informatics, uh, research data policy, ethics, and the professional context of the digital humanities. Frustrated with the pervasive lack of access to quality research data in the humanities and social sciences, Eric spearheaded the development of open context in 2006. Eric runs research and development for open context and manages the technical aspects of data publishing and archiving including systems um interoperability, data integration um and indexing. He's been an active and vocal member of a growing global community dedicated to better ethics and practices in sharing and preserving knowledge of the past. He taught project management and information service design at the UC Berkeley School of Information and was recognized by the Obama White House as a champion of change in open science. Um, please help me well welcome Eric today who's going to present for us archaeology by and for robots question mark publishing data curation and asserting humanity in the AI ecosystem. Welcome Eric. Thank you. >> Thank you so much. Thank you all and um really appreciate the chance to uh join you all here and um I'm going to be this talking about stuff that is actually kind of new to me and I'm not really all that fluent about this and so you'll have to bear with me about some of this but um briefly I'm just going to uh sort of talk about some of the background sort of the sort of socio ideological background about like how publishing is working sort of talk about that with David Greyber and jobs and neoliberalism and that type of thing uh talking about datification and how that's actually really bad for people like us who curate data uh because it makes uh because of all the incentive structures around the way academics communicate with one another researchers communicate with one another and the incentives that that creates and uh all of that sort of sets the stage for how large language model AI enters into this ecosystem of scholarly communications. So with publishing with data management and all that kind of thing, there's uh that's all important background. And then last I'll just sort of close out with a few thoughts about uh mostly from other people who are more eloquent about where we can go from here. So uh that's basically what I'd like to cover today. And just a bit of background and um you heard some of that the bio but I uh do the uh technology development and uh maintain open context which is a publishing service for archaeological data and uh we publish data sets from researchers uh who do excavations, surveys, collections work, other kinds of studies uh and and they do that work uh on on archaeological resources from around the world And right now it's grown to about two and a half 2.4 million colle uh records, 200 projects worldwide. And we've been doing this since about 2006. So really um a lot of what I'm going to be talking about is sort of informed by a really long-term engagement with data curation in archaeology. So this is uh over 20 years now uh working on this. And um just want to highlight that we work with a much very sort of uh uh collaborative kind of uh uh uh framework where we uh really dependent and are grateful for services from other institutions. So, we're not ourselves a preservation archive, but we work with other uh services, especially the California Digital Library, uh a repository called Zenotto, for the long-term preservation of the resources that we publish. And uh all of that sort of background is going to be giving you some idea about where I'm coming from when uh looking at some of the issues around AI. Uh just really briefly, you know, you can play around with open context. It's free and open access. You can just sort of see what they have. This is just a sort of a range of some of the public projects that we've published. And um one of the big questions always when you say, well, I do data management and data curation and stuff like that. And it's like, well, okay, what's data? It's a really good question. And um sort of paraphrasing Gollum there about that. And um it's a it's an interesting kind of a question because it's I it's part of the framework about how we have to start thinking about things like AI too. So most of the time uh most people who are engaged with scholarship are really working in this space of more unstructured data. So narrative text, this is uh text uh or sometimes also videos and images and things like that that you're meant to experience. You're meant to uh you're meant to read and what you bring as a reader with all your own background and knowledge and your own experiences and thinking about, you know, your social relationships and all that. That's how you engage with uh blocks of words and books and things like that and articles and whatnot. in unstructured text is really very typical of the uh content of scholarly publication. It's really typical of academic publication uh of scholarship and it's very good for telling narratives, telling stories. That's what most of uh what one engages with as a scholar is um especially when you're doing sort of formal professional communication is around this world of unstructured text. And on the other extreme uh is other ways of organizing information is more structured data. And structured data is a bit different because it has much more simplified logical kinds of structure uh uh organization. So you think about it like rows and columns of words and numbers in a spreadsheet right in a data table. But uh structured information is actually uh not just the sort of a thing of technology or recent technology or digital technologies. It's actually uh you can argue that it's probably the uh oldest forms of writing are structured data. So you could see like a a Sumerian Kaoform accounting tablet and it's sort of like a an Excel spreadsheet, right? So it's organizing counts and um whatnot of different kinds of commodities maybe calcul you know debt rations all sorts of stuff like that in the structure that you sort of analogous to um a spreadsheet and maybe uh if we better understood kipu kipu are probably working the same way for u in the Andes. So structured information is information that has a sort of um kind of uh organization that's really there to promote to facilitate the quantification of things. Right? So you're counting uh similar kinds of things and it is often really sort of linked up and integral to kind of an administrative or bureaucratic mindset. And uh that is the kind of information that is what we often talk about when we talk about data right so for data management and whatnot. I want to sort of highlight though it's a continuum right so there's not a hard and sharp divide between what's structured information what's unstructured information. Uh structured data will often include a lot of unstructured text. If you've ever seen a database or a spreadsheet and there's a sort of notes column, right? Got to read the notes column that's really hard to sort of work with or count and aggregate uh without actually reading it. So um and you know unstructured things can have books and chapters have paragraphs and other sorts of structures within them. So it's it's there they're there it's a continuum you want to think about. Now again structured data is uh sort of more of this sort of more recent kind of phenomena when you start thinking about this for scholarly communications for communications amongst researchers. So it is really what the subject is of when you like apply for an NSF grant you have to write a data management plan. It's mostly focusing on that kind of stuff structure more structured kinds of information and it is the type of thing that uh data repositories are being built to try to support and um that is where a service like open context mainly plays. We work we work in this world of more structured data and this is just an example of why structured data is interesting and important for archaeology and why we should communicate it. First of all, archaeology is mostly a uh a field that involves if you're doing excavation destruction, right? So you're destroying the archaeological record as you're documenting recording it. So that information that you collect you have an ethical responsibility to uh curate well and uh archaeologists typically create a whole bunch of structured data. Thank you. And they do that uh in order to do things like this mapping, right? So on the left here, this is a a a retto which is a spool basically. And on the right you see a spindle whirl. Uh they're both used in textile production. This is at an Atruscan site uh just south of Sienna called Poio Chibetate. And you can see different spatial distributions, right? And that's because we've got objects that have been classified and now because and they're classified and they're in a structured relational database and then you can count them and you can map them on our GIS. And because you're doing all that, it's structured data and you can work with numbers of these things, then you can do interesting things like visualizations and exploring and analyzing patterns with this information. And it's really easy to do with software with computers. That's one of the reasons why um this kind of information is so important for interpretation in archaeology. But also since archaeologists basically record everything in relational databases and spreadsheets and other sorts of digital media, it's really important to be able to have mechanisms by which archaeologists can share this information and for that information to be archived for future generations to engage with. All right. So, um, saying all this, just because there are numbers involved with structured data doesn't make structured data necessarily like objective and empirical and all that kind of stuff, right? So, these are very social kinds of things, right? The way you organize information, the how you classify things, what you choose to quantify, what you don't choose to quantify, all of that is embedded within larger kinds of systems of your background, your culture, your priorities, your research interests, and um what you think is important, what you don't think is important. So these are all very socially embedded, right? So data are very socially embedded and uh because of that, they really need a lot of intellectual engagement. So, this is a really integral aspect of archaeological method and theory that needs to get talked about a lot more. Would be awesome if we did that more. Um, it really should be integral to scholarly publications, but I'll get to why that's hard in a moment. and uh the services and the infrastructure the kinds of uh services like I run with open context all of that actually matters too because we're organizing all this information presenting it to people like you and our decision makings what we what we do and won't do the affordances of our system sort of help shape how you can actually interpret all this stuff so that's all important too so all of this stuff is uh um not just a sort of very bureaucratic compliance thing you don't just write a data management plan to get a grant and then forget about it. Hopefully, >> we want to have uh this is like be excited about it, right? You know, this is this is this is really core intellectual labor. Um more about why data are social. Uh there are the two main uh frameworks for good practices in managing and curating data. One is called the fair principles findable, accessible, interoperable, reusable. The other are the care principles. care principles are about indigenous data sovereignty. So collective benefit authority to control, responsibility and ethics. Both of these frameworks you think that you know they have their tensions in them but they both really are about how data are not just there for you as an individual investigator but they have large you have larger social responsibilities visav the curation of this information. So again, um not going to get into the spec specifics of these too much, but really data are part of your wider set of professional social obligations. Um and uh your your your your obligations to the w to wider communities. Uh just to highlight some of this, this is work that uh Sarah Kza and Nha Gupta and Desiree Martinez and Chris Nicholson are leading. This is a fair care pro um uh cultural heritage network project that is funded by the IMLS and it's all people representing all sorts of different kinds of communities and and uh groups engaged in cultural heritage trying to come together to set up good practices uh ethical frameworks policies on the curation of digital data that relates to cultural heritage. So again very social uh kinds of practice. Now that's the sort of like ideal kind of thing we want. Yay. Uh let's be engaged with data, recognize the sort of social connections, community obligations associated with it. But at the same time, we uh work within uh other frameworks, other ideological frameworks that have very different sorts of pressures on researchers and h and on how we manage data. So uh this is a fantastic book by the late anthropologist David Greyber on jobs. I'm gonna say a lot. Uh so I don't know if that gets bleeped later but um sorry and um really it's about the sort of proliferation of bureaucracy in our work lives and that bureaucracy uh often relates to the sort of ideology of tailorism. So uh tailorism is the sort of notion that you know that basically we want we need to motivate monitor and motivate good performance and we need to sort of set up objective metrics so that you can achieve your goals like you know maybe publishing a lot and high impact factor journals and stuff like that. So that is uh in some ways uh thinking about um data as a sort of for surveillance for uh motivating certain kinds of behaviors that you want to use and um it sort of reduces a it's kind of a very reductionist picture about what data means in a sort of a social setting like a workplace including a university. And what that does when you see how it sort of relates to things like uh your own practices as researchers uh you get this sort of uh escalating arms race of publications. So this is an interesting graph uh that came out in 2022 about like uh the uh number of publications that are uh uh uh uh created by people who are just exiting grad school. Right? So now you in 2022, so you've got like, you know, a lot of people have five publications before the peerreview publications before they even exit grad school. Some of them are 20. That's a lot. You know, that's I don't that's [clears throat] a whole career right there. Um and uh this uh and the sort of uh pressures to publish or parish and have those things uh you know build up the longer CVS uh it has all sorts of uh uh uh uh it relates to all sorts of sort of um tensions in labor in labor practices and archaeology. So I and scholarship in general. So some of this is going to be like burnout, right? you're ba ba b b b b b b b b b b b b b b b b b b b b b b b b b b b b b b b b b b b b b b b basically uh this is all unpaid labor for authoring, reviewing, editing all this stuff. Uh most of this is happening in your sort of free time unscheduled, right? So basically it eats up your weekends and uh because everything else is eaten up with like teaching and committee work and all sorts of other things that you got to do and it's especially burdensome for people who have invisible labor obligations, right? So if you are say if you if if you're a a member of a minoritized community and you're hired, you may have a lot of mentoring extra work that you need to do. And that's uh because you know you got feel like there's an ethical obligation for you to do it, but it's going to cut into your time for publishing. And if publishing is the only thing that's measured and the only thing that matters for career advancement, then uh you're going to get squeezed. And that's that's why some of this stuff is just um has kind of insidious kinds of impacts. So um just to give you a couple examples of what some of those metrics are that sort of uh drive a lot of this and that uh get monitored and that do uh get used for things like promotion and tenure decisions and all that. So there's all sorts of literature in this whole thing. Uh one thing is the impact factor. So that's a a moni uh basically a measure of how much a journal title gets cited by uh gets cited, right? So you want to publish in a journal that gets cited a lot. So like nature and science and stuff like that. Another thing is called the H index, which is more of an individual measure. So that's your own personal citation uh citation counts. and they really sort of simplify and contra uh uh compact a lot of that sort of work that you do as a scholar into one number, right? Like and that that sort of says like uh oh, I'm important because I have a great in H index and that's that's how some of that works. There's other measures too like uh altmetrics. How many of you have heard of altmetrics? Okay. Yeah. So that's like how many tweets your paper has gotten? That type of thing, right? And that's kind of that gets even weirder when you think about who now owns Twitter, right? So, so like uh why are we sort of counting this and sort of you making decisions, labor decisions and hiring decisions and stuff like that because that's Elon Musk's algorithm. Anyway, how does this go go back into data? And uh that is a good question. One of the uh big things about this is that traditionally at least data sets never actually counted in all of this, right? So you could uh put a huge amount of work into data uh curating cleaning up a database uh sharing it. It could have uh it could be really important. It could be something that's sort of uh irreplaceable because you've done investigations on something that no longer exists and it's not necessarily going to count anywhere. And there are some efforts to try to do that. So you can see here in open context we have there's a citation thing and there's some uh persistent identifiers. And this persistent identifiers can theoretically be used to tally up the number of citations that um get uh that reference this uh uh item of data. Um but um most of this is you know there's some efforts to try to get data citation metrics to sort of count for things like promotion tenure decisions but um it's hasn't really necessarily entrenched itself all that much and even so uh data is weird. It's a little bit harder to engage with because a lot of times people are when they're citing are reading a title and maybe an abstract and then they decide to throw that article into their bibliography. And so um uh whereas a data set you got to engage with a little bit more deeply and it takes a lot more effort to actually use. So data citation is actually a lot smaller typically. There's some recent studies now coming out uh of uh data in repositories. the citation rates are smaller than uh article citation rates. Anyway, all of this leads to very poor incentives for uh carefully curating data. Now, um why is this bad? Uh undervalues quality uh overwhelms our editing and peer review systems. How many of you feel like nobody actually reads anything now in a meaningful way? Right? So, like we're flooded with stuff. um there's not a lot of uh incentive. There's no monitoring or checking about like if uh if you're actually reading it or there's no value to you as a researcher to actually read the other literature necessarily because it's not being counted as much. Uh if you do write something that's carefully well considered and it has a really interesting sort of uh outcome, it's could be lost in the noise, right? There's so much flood of other publications out there, it's hard to find the good stuff. The other thing is that it really undermines the whole point of open access. Open access is to make uh literature available to wider communities is to you know share this work that's funded by public uh funding mostly share this work that you know archaeologists you also have to give permits to do your work that's you owe have obligations to wider publics. Um if all if if everything is sort of uh becoming increasingly just churned out be without a lot of thought then you're just flooding open access with not all that great stuff. So that's not a good thing and it's just kind of dehumanizing and this is where the machine infiltration um starts to come in in in a bad way. So, enter AI and I want to give props to Karen how here who wrote a fantastic book empire of AI and um she makes a very good point when people talk about AI it's like talking about transportation right like it could mean a whole bunch of different different kinds of methodologies so like a bicycle and an airplane are both modes of transportation but they have very different kinds of environmental impacts and affordances right so when people talk about AI It's very underspecified. I'm going to mainly focus on large language models and I'm going to focus on the large language models that are from the sort of hyperscaler companies like OpenAI, Google, Meta, those those people who are building these giant data centers and harvesting huge amounts of information to build these models and um that's focus of her book. And then she makes uh that's what it's called Empire of AI. she sees direct alignments between this these programs to build these large language models and colonialism. And just to give a really high level uh garbled overview of what a large language model is um it's a way of representing text. So you break it down text into tokens which are sort of like words but not quite and uh you define numerical relationships between them uh over uh uh large bodies of uh uh repeatedly. So you get a text and you sort of see like 'the' relates to this noun next to it uh and then uh 'the' can also relate to u you know like uh some some other noun next to it. And you do this over and over and over again. And you build up these large multi-dimensional linkages about how one word might relate to another word and the distance between those words. And you get this really giant um multi-dimensional set of numbers that basically sort of define the sort of uh properties of how words relate to one another. So it it is hard to explain. Uh but all of that is what is inside a large language model. So it's basically a a big multi-dimensional set of relationships of words and how they relate. Um and how they relate over a huge corpora of text. So you sort of throw in pretty much the entire content of the worldwide web to train these models. And then you can see these models have uh uh information uh that could be from all sorts of subject matters and mostly because it's from the internet um it has all the sort of socioultural biases of the internet. It's mostly in English and there are all sorts of other problems associated with it too about about these things but in general that's where the data comes from and it's represented in this big numerical kind of a graph relationship. So uh you can think about a large language model in some ways as sort of taking less structured information and creating a structured data representation of it. So it's it's it's it's kind of a weird data structure. It doesn't look like an Excel spreadsheet. It has I don't know sometimes like thousands of different dimensions to it. So uh really hard to visualize but that's that's sort of what it's like. Now um these have uh all sorts of uh uh different kinds of implications on scholarly production uh comm scholarly communication. So lots and lots of stuff has been written about how researchers are using these uh AI tools to be able to uh uh generate um uh scholarship and trying to remember because of metrics improve their productivity right to make more publications better faster and using AI to to uh help sort of accelerate this sort of stuff. So that's that's uh uh been widely noticed and and widely discussed. More recently, however, um there's been a bit of uh more concern about well, what happens when we start having less human roles in creating the the publications and um AI is involved more in the creating of these things. AI is more involved in the reading of these things, the synthesis of these things, the aggregation of these things. And what happens when this is turns into a sort of a loop, right? When it's a cycle. And that's a that's an interesting kind of a problem. Um uh the term stochastic parrot uh by Emily Bender is a great term for uh how LLMs actually work in this context. But you could start thinking about this as like an oraoris, right? So the um when you use the LLMs to create new content based on content in the in in in the already in the LLM, you it's a sort of a feedback loop of you're sort of eating yourself, right? So the model is uh getting more and more information that was processed through the model. And what happens when that when that actually happens there's something called um model autofagy disorder. autofagy meaning eating yourself disorder. So, and it's the the acronym is explicitly trying to make it relate to m like mad cow disease basically. Uh the uh you get a situation called model collapse and what happens is that as these models ingest their own stuff uh the biases inherent in the data get amplified over time and um that's because these models are sort of like a lossy compression of the information that they're trained on. So, if you know what lossy compression is, how many of you have taken pictures and like it gets shrunk [snorts] down into a thumbnail, a little JPEG and then you try to zoom up and then it you lose a lot, right? So, you can imagine it like this. So, this is a an atruskin uh freeze plaque with a horse on it and uh you could see a couple cycles of doing that then that you lose so much detail and you get this really bland undetailed image. uh whereas the first thing was very richly textured and that is essentially what happens when uh LLM is used uh uh it it's uh used repeatedly to uh uh get trained on the material that is generated by an LLM. So you think about it nowadays, how much content on the internet is already being generated by LLMs, right? So the quality of information on the internet for training future data is is is going to be declining. And this is the kind of a dystopian future that you don't want to get into as researchers, right? You don't want to have uh the quality bi bias amplification through repeated cycles of this kind of stuff. So that's bad. Um, another thing that's not great, uh, and this is there's some really interesting work by Kate Crawford, uh, who is, uh, looking at, well, where are these, uh, say image models, the LLMs with image multimedia kinds of capabilities, where are they getting their images and stuff like that? And it's mostly e-commerce, right? Because most of the internet is full of e-commerce. So, there's going to be aesthetics and perspectives and priorities in there that are sort of embedded in the models. And then you start thinking about well you you know it's not just publication that happens for archaeology but also public engagement. So if you use these things for public engagement your outputs for public engagement are going to be trained by a very kind of commercial aesthetic and perspective and um then you uh get into controversies like this where the British Museum was recently caught posting AI slop. So video is generated by AI of people looking at stuff at I don't even know if some of that stuff's in the British Museum but you know just sort of artifacty things and um and that's the some of the dangers obviously then you know this gets used to retrain those a those models right so you get that uh mad mad disease of the sort of oraoris effect of the models degrading in time but also you're sort of mixing a sort of a very commercial kind of a aesthetic in perspective ive in with your um with with the cultural heritage material. So that's bad. So another thing to think about also is uh this goes back to Karen House when she's talking about the sort of uh relating AI large language models to colonialism is how extractive they are and how extractive they are also to memory institutions. So there uh there's some nerdiness here in this slide, but uh websites have often think something called a robots.ext file and that's there to govern how webcwlers, which are software agents that fetch web pages and traditionally those were used by search engines. So the search engine would fetch a web page and they would index it and then uh they would serve it on their uh search uh search index so that people can type a query into Google and find your web page. uh that there's they were regulated through something called robots.ext. But nowadays AI bots are so voracious and wanting content they ignore all that. They swarm sites they could act like a distributed denial of service attack. And um they also do all sorts of insidious stuff like they fake their uh identity. Basically they try to pretend that they're humans. So uh that's all of that is um not very nice practice and it is really having a huge impact on your libraries and repositories not just in archaeology but everywhere. So this is a recent report uh by the um uh uh coalition forked information. Uh how many uh repositories have webs uh websites being scraped? Um most people are getting getting scraped by AI bots. And uh if you're getting scraped by a AI bots, 70% or so are having a disruption of service. So when you now go to a repository like uh so uh social science archive or place like open context or tar or places like that, 70% of the sources now are going to be uh harmed by the impact of all of these bots. And uh that is uh kind of sad too when you think about that. this uh industry is uh you know you're a university here at UC Berkeley pay subscriptions to OpenAI and all these other companies and you know you're you're paying for the privilege of having them kind of uh do a denial service attack on your own resources which isn't so fun anyway. So open context uh we're not alone in this uh I mean we're we're impacted by this too. So we have about two and a half 2.4 4 million web pages. And just to give you a sense of scale, Wikipedia has 7.1 million articles, but a lot more pages because each article has more than one page associated with it. And we have somewhere around 1,200 or so uh references from literature according to Google Scholar. And uh before all this all all this AI stuff um uh we were okay uh with us doing open access without much problem because the economic basis of open access is that it's non-rival meaning that we can share information by us sharing that information it doesn't deplete the quantity of information available for other people right and that's very different from a car right if you buy a car then you own that car you can't just sort of share that car uh with somebody else at the same time, right? There's only one car. It's a physical thing where the digital data, you know, multiple people can get it. Uh they could get a copy of it. It doesn't it doesn't deplete the supply. So, um that's the sort of economic basis for open access. But with the bots, bots crowd out people. And so all of a sudden these websites that are open access that had an a sort of an economic justification for working as open access, they're all now crowded out by a whole bunch of bots and they overwhelm the uh uh service. And so now content on the web is a lot more rivalous. How many of you have run into captas where you have to sort of identify the stupid fire hider and the crosswalks, right? Yeah. on all of these websites now because you didn't have to do that, but it's all from the impact of these stupid bots. Anyway, um and it, you know, defending against this also costs time and money. So, this took me about a week, but I installed a a some a bot challenge or a network protection, and you can sort of see our traffic levels go boop to a much more reasonable level. And um that uh you know I that week could have been spent doing something else like I could have been on vacation. I could have published more data. I could have done community engagement. There's all sorts of stuff I could have done with that week instead of letting bots. Now um so why is this bad? All right. So repositories have to use uh I put a lot more effort into fighting this. Sometimes repositories are being driven to use commercial services like Cloudfare to that would sit on the in uh uh on the uh networks and monitor traffic and try to intercept it. So you're using AI to fight AI which means you've already lost um makes everything obviously more expensive, degrades user experience. And then the uh other really bad thing about this is that because um uh the repository the the the sort of nonprofit repositories are harder to access, harder to use, there's more and more enclosure of this information within these commercial large language models, right? So they're sort of extracting and monopolizing this and making access too difficult into uh the sources. So um anyway, this is all indicating that AI owners have a lot of disdain for social technical conventions like robots that text and stopping this even legal constraints. Uh Anthropic, one of the big companies had to pay a big civil penalty for uh u copyright violations on a bunch of the texts uh on a bunch of books. And um that disdain is also really important when you think about trying to implement something like fair and care, right? The sort of especially things like indigenous data sovereignty. Um we can try to put metadata indicating uh ethical requirements uh around all of this and I'll show you how we're do we we are do try to do that. But is anybody actually going to pay attention to all that? Right? Is anybody going to pay attention to our ethical expectations about how this information should be presented and used and represented? So, um this is just an experiment that I'm showing about that kind of question. So, the uh we we experimented with something called really simple licensing and it's a licensing framework specifically designed for AIs to tell them what to do. Uh Creative Commons is one of the organizations that's involved with it. You've probably seen Creative Commons copyright licenses because they're the main type of copyright license used in academic publishing, open access, academic publishing. And this is what our metadata looks like. Uh sorry about showing you XML. Um but this is uh uh that link there is basically saying please pay us, but it's like busing, right? So it's like um you can play a saxophone in the street corner. People can enjoy it without paying, but you have a hat out and this is like providing a hat. So far none of the none of the AI have given us anything. Um but um everything down here terms that are basically describing the ethical framing of the data that we publish. So we have a lot of language about indigenous sovereignty, the care principles. We talk about uh community archaeology, the kinds of expectations we have about um you know for things like uh working with human remains, that kind of thing. All of that is actually there and then the question is is it actually being used by AIS? So here's a question I asked of Gemini and uh you know how can I ethically use the digital index of North American archaeology to study that and uh Gemini was very dutiful and said well here you follow the care principles do this that and the other and it was actually pretty good you know that's a nice response. Um, and I was thinking, well, maybe that's actually listening to the license, the RSL license, but if I drop the term ethical, ethically, uh, then I get a very different kind of response from Google Gemini, and it's much more sort of technical, right? So, do this, that, and the other. But, um, there's nothing about care, ethics, data sovereignty, or anything like that. So it really seems like the RSL is being totally ignored uh by by Gemini at least in this case. Who knows what'll happen in the future because you know it's hard to know how these models update and all that kind of stuff but anyway initial uh experience and experiment wasn't that great. So um uh all of this moving towards the future now um if uh so Corey Dr. coined a really great term in shidification to describe the sort of declining quality of monopolistic information services. He was definitely predicting in a shitification for AI services. Remember these things right now are heavily heavily subsidized by venture funding, right? So it's a lot because you you know there's a huge amount of money that is required to run these data centers. It's all being paid for by you know speculation at the moment. So when that speculation money dries up, these things are going to get a lot more expensive. So do you want to be dependent on that? You want to ask yourself about that. So maybe not. Um he also has a typology of like you know is AI being done to you uh then you're reverse centaur you work in support of AI or do you want are are you is it using AI for work uh for your own agendas basically and that's the kind of thing that I think is an an interesting kind of framing about all of this. So um looking at that moving forward uh also um thinking about you you know uh are you using AI or AI being are you being used by AI uh really think deeply about what it means to be a productive scholar you know so uh if AI is increasingly mediating all your engagement with scholarship uh you know where are you in that system you know are you driving it or are you being driven by it. So these are the big open questions that I think are good to answer to explore looking uh into the future too. Uh one of the weird things about this is like well if you can generate a article of sloth article right uh regurgitated so easily uh things like primary archaeological archaeological data that requires encountering the world and encountering the outside world and other communities. Maybe that'll actually be more important. Um, so you know, I'm just sort of making myself feel better maybe that our work will be more valued. And last here is uh there's just some great literature of uh archaeologists engaging with some of these technologies uh trying to find ways of taking agency and ownership over them so they're using it rather than being used by it. So, uh, Gabriella Gatidia, he's written a really good article about this about various flavors of different AI and how their their applications. Sean Graham also, uh, really good resources to to to learn from in thinking about how to work with maybe smaller, more contextualized, more ethical approaches that don't, um, boil a lake every time you ask a query. So, these are the kinds of things that I think are really valuable. So just to close uh be a lite and being a lite doesn't mean that you're against technology being being a lite is that you're against the unfair and unequal access and and and control over those technologies. So I think that that's mainly where uh moving we want to try to move towards. So slowing things down, not always pushing for speed, not always pushing for the highest metric and of you know impact factor and all that kind of stuff. All of that is really important to try to get off of this treadmill. Otherwise, we're going to get into the situation of eating herself, right? Or of our content eating herself. So, that's it. Thank you. >> Any questions? >> That's good. >> Yeah. so important, so timely. Um, and I just couldn't be more excited about some of the ideas you put up and trying not to be cynical about our own retention, tenure, promotion process, right? Um, where, you know, somebody who tries to do slow archaeology co-publishing with community mentors versus somebody who publishes 40 articles in which they're one of 20 authors. These are things that really matter to us as we try to keep our jobs. But beyond that, um where you had the RSL um challenge or you know you you actually pulled out Cloudflare and said like these are the things that you know we're paying for this stuff. um earlier as one of your first slides said that there's these like appropriate ways and I know you you have really dealt dealt with this and thought it through with with you know um data sovereignty issues >> within the circles of knowledge and and knowledge sharing and data um curation for like >> at this point I'm not even sharing files with locationational data in any way shape or form on box drive or anything we're so concerned ern about locationational data getting out into the world. So it only circulates >> on like USB drives and crap like that, right? Or written reports that we hand to community mentors in the field. >> And so one of the comments that somebody made to me was, well, we could just instead of the buses and the fire hydrants thing, you know, pick the three things in Maidu that mean fire, right? And it could be a burn, you know, these, you know, like what are the things that we're offering or that you've thought about that maybe I haven't read yet. I apologize about how these the licensing or the challenges or things keep it circulating in smaller groups of people where it's appropriate, >> right? >> And it's still it's still curated, right? Like somebody could still go to the tippo and say, >> "Right, >> you've already done all this work with electrical resistivity on the delta and we're doing similar thing. Can we have a comparative data set to play with just to see if we're doing the right thing?" What what kinds of things are you thinking about when it comes to keeping it within smaller groups of slow archaeologists? >> Yeah. Well, I mean a lot of it has to deal with like um capacity unfortunately. So uh so curating archaeological or any data right and having running a data repository so it's preserved has some big real expenses and has requires real expertise and uh uh real funding basically. And so it's hard for a university, right? It is harder for an indigenous nation that also has clean water issues, right? And so you've got you've got to um sort of see that the these the that the capacity is difficult. The best approach I think sometimes is I mean this is uh where um maybe multiple indigenous nations can pull resources together to co-own something. But I really think that owning the infrastructure is the fundamental aspect of sovereignty. >> Yeah. >> Right. Not just, you know, you can't stick some licensing or some sort of agreement and share it on Google Drive, >> right? >> Because uh that might be ignored. Um and Google has better lawyers than anybody else or you know or you know that that kind of thing. So I I really think that that sovereignty is going to require these sort of capacity developments to really make much more meaningful. The other thing is security is also expensive. >> Yeah. >> Right. So, um people ask us why we're open access. Well, one of the reasons why we're open access is because we can't be closed access. We don't have the administrative capacity to uh adjudicate who should have access or not. And we definitely don't have the I mean, we don't have the technical resources to be able to protect this stuff and uh not get hacked or sued. So uh we only try to curate low sensitivity low-risk information. >> So um you know it's a it's a burden to try to protect this information and that requires real resources too. >> Yeah. >> So you know but pooling efforts might be the best way to go about it. >> Thank you sir. >> Sure. Anybody else? >> Yeah Chris. Thank you Eric. Great talk. You know, I'm just struck by for how many years we really worked to make data open and to, you know, machine readable links and linked open data, but we didn't realize what was coming in terms of, you know, these tools. And I'm thinking about kind of the education side of this. So, it's one thing within our communities, but we're embedded in such a big community and working with students and partners who use these tools and legitimately do find interesting things that would have been very hard for them because they're not professional archaeologists. So, I'm wondering, you know, how do you think about education as one of the tools to help um help this move forward in a in a responsible way? Yeah, I think people need to learn more about first of all what the limitations of the tool that language models are. Um I have a quicky little KOD if you want to see it. Um just scroll down. Um and uh right they're token prediction machines, right? So they don't they're not really thinking and they're not really reasoning. What they're doing is they're producing the next token. So, um this is an adorable me about about all of this. Um and here's an archaeological example of it. And um this is just important because um uh makes you a more uh makes people more informed consumers of the services for uh I really think also ultimately um these large language models are not economically, they're not socially, they're not environmentally sustainable. They're going to if they're eating themselves, they're not even going to work over multiple iterations of them eating themselves. So that we're we're in a sort of a state right now that is not the sort of end state obviously uh of how this and they're not necessarily going to get better, right? So this technology of the large language models is has a lot of development, a lot of investment and they're sort of running into these roadblocks about you know some fundamental questions. So here I asked uh Google Gemini uh based on this paper um where uh uh how many archaeological sites would be flooded by sea level rise in um Mississippi and it dutifully said oh well with a one meter sea level rise you can have 303 sites uh with other uh uh flooded I'll just make that bigger so you can see that uh with other uh uh sea level rise scenarios to see more and it's very confident in all that and I asked are you sure [laughter] how did you get that and said yes look at table one of the t and blah blah blah blah says table one and it's very and then it says I'm being conservative about this that's great and I said well look I checked that table myself and all of Mississippi is just NA values it just made up the numbers right because yeah because it's a text it's a token predictor right so it probably found god knows where in its trading data set something about archaeology, something about Mississippi, and I found the number 303, right? And it just put that together, but it's nowhere in that paper. In fact, that paper explicitly said Miss Mississippi does the data from Mississippi is not included. >> So, um, so here's the the original paper and you know, you can read it and find Mississippi is not included. So, yeah, education is really important. we're going to find that there's, you know, there's so much hype, there's so much investment about this. This is probably a big speculative bubble that's going to not last. But there are, I think, important interesting uses of language models. Uh, one interesting thing could be uh like uh for naive users quering a site like open context like we get queries like daily life in ancient Egypt to our data in open context. We don't have daily life in ancient Egypt. We have counts of sheep and goats and you you know barley and all that kind of stuff and that tells you about daily life in ancient Egypt. But there's got to be some linguistic bridges between the sort of technical data that we have and that user's query and I think that's where there could be a really interesting educational u application. Anybody else? >> Um thanks Eric. Um could you explain the centaur metaphor a little more? didn't get. >> Oh, the centaur thing. Well, uh, it's not mine. It's Cory Doc Rose. So, the centaur is like you're a human being, but powered by a horse's body, right? And you've got your human head and your human sort of cognition and agency, and you're using the power of that horse to sort of move around and being awesome, right? So, you're using that musculature of the horse. The reverse centaur is where uh you're sort of uh poor human being strapped to a horse head and you're sort of pushing the horse around. And so that's that's more the sort of analogy. So basically the the centaur is the one that is the is the has agency over AI is using AI for their own agenda and the reverse centaur is the person that is being used by AI. So uh Karen how has a lot of this in her book also about like uh uh people who are in uh the global south who are being tasked with training these models and sometimes they have to train them on going through horrific content right of like videos and you know snuff films and stuff like that which is uh you know will give you post-traumatic stress disorders you know they're really awful things and they have to they're exposed to it for like hours and hours and hours than any sort of mental health services to help them and it's a very exploitative kind of a thing. They are being used as reverse >> centaurs. I see. >> Right. Another example would be like if you're like a CRM archaeologist and the AI spits out a whole bunch of models about like I don't know predictive models and you have to evaluate all of them with no time. Um and uh you just sort of but your name is being signed to the results >> then you would be a reverse centaur too. >> Yeah. Chris getting in today centaur. [laughter] Yeah, >> there you go. >> Eric, it looks like there's a comment or question in the chat. >> Oh, uh uh Okay. Uh do I jump to the latest? Uh okay, thank you. >> Yeah, keep fight. Thank you. Thank you. Thank you. [laughter] Also, more thank yous. >> Thank you, Francis. Okay, I don't Yeah. Yes, >> if there wasn't one on the chat. >> Okay. >> Um, >> but I'm glad you got many thank yous and I'll add one. Thank you again for about the talk. Um, so you mentioned a couple of different institutions that are doing good work on trying to develop standards for this. Of course, the Fair Plus Care Network and RSL. I'm curious if there's any more spots where you're seeing that kind of work done either discipline specifically or more more generally. >> Yeah, there's uh so at the University of Arizona there's a a a collective or promoting indigenous data sovereignty. Uh Stephanie Carroll is key participant there. I'm forgetting the name of her institute. Um but she's one of the lead researchers there. And they're actually proposing uh an internet standard for um documenting the proh that something has indigenous data pro indigenous provenence. So it's more of a huh so it's baked it will be baked into the web. Um the which is good. The bad thing ne is about well will it be recognized and used right and and like because uh it probably a lot of these things don't necessarily have formal legal teeth. It's although it is weird in the in um the United States because indigenous nations are sovereign and have some sovereignty rights and uh so some there uh that I don't understand the entire legal implications of all of this but it's a um uh but there are other indigenous communities outside the United States that don't necessarily have the same sort of legal structures and recognitions. Um so yeah, there's uh there's um a big thriving community trying to address this. Um and I think it's uh uh you know um even if it's um not necessarily formalized with legal protections, uh there's still the sort of the social aspects of it still matter, right? So the companies can be pressured, right? So, uh, Disney was pressured and actually did a pretty good job with what is that movie? Um, uh, in Polynia. Uh, >> yeah. Yeah. So, that where the where where u the reception amongst uh uh different communities in Polynesia was pretty good that that Disney did actually um make some meaningful overtures and efforts with that. So, uh, and it didn't need to legally, but socially, you know, and and, uh, reputation wise, uh, it it matters. So, maybe th those kinds of pressures could be exerted on AI companies and search engines and stuff. >> Yeah. Yeah. >> Is there a way to get some of your the links you have from your talks, Chinese, books, and articles? >> Yeah. Yeah. I can just share the whole presentation. Yeah. Yeah. Be happy to. Yeah. Absolutely. Thank you, Eric. >> Okay. >> Thank you.