Submind YouTube summaries
Thumbnail for Jennifer Ding, Dr Sarah Ann Gilbert, Dr Hanlin Li - A Toolkit for Community-Driven Data Governance

Jennifer Ding, Dr Sarah Ann Gilbert, Dr Hanlin Li - A Toolkit for Community-Driven Data Governance

Watch on YouTube

Video summary

The core subject of this presentation is the shift from viewing data governance merely as a technical or legal compliance task to treating it as an active, community-driven practice. Jennifer Ding, drawing on her background in machine learning and recent work with artistic datasets, argues that the current AI ecosystem relies heavily on open public data and the labor of often overlooked contributors, particularly from the global south. To address the lack of voice and control these individuals possess, the talk proposes a framework where communities are embedded directly into decision-making processes throughout the entire data pipeline, from collection to model deployment. This approach is inspired by initiatives like the European Data Governance Act and academic work on bottom-up data trusts, aiming to move beyond simple opt-in/opt-out buttons toward nuanced systems that allow data subjects to express specific preferences regarding how their contributions are used. A primary example discussed is the Coral Data Trust Experiment, a project conducted with the Serpentine Gallery in London involving fifteen UK choirs contributing voice data for AI training. The experiment sought to foster a collective identity among these diverse groups so they could exercise their data rights collectively rather than individually under GDPR. Through conversations and the use of deliberation tools like Pace, the choir members engaged in "live action roleplay" scenarios to determine what good-faith AI development looks like for them. A significant outcome was a mental shift where participants moved from viewing voice merely as raw recordings to understanding it as complex data with legal and artistic implications. The choirs ultimately expressed strong reservations about sharing their data for commercial purposes or unrestricted public use, highlighting the need for governance tools that respect these nuanced preferences rather than imposing standard licenses. The presentation also highlights other real-world examples of collective data governance in the language space, such as Coher Labs' YA project, which aims to bridge multilingual gaps by building models and tools openly with global communities, and Mozilla's Common Voice initiative, which is evolving into a "data collective" to better handle community preferences on compensation and attribution. Additionally, the Big Science workshop's creation of the Bloom model introduced a data stewardship ecosystem that connects technology, law, and rights holders. The speaker concludes by emphasizing the necessity for a toolkit that aggregates these diverse learnings and resources, making it as easy to share governance guidelines and legal frameworks as it is to share software models today. Such a toolkit would lower barriers to entry for participatory exercises, enabling capacity building and consensus across different organizations so they do not have to reinvent the wheel or spend excessive resources on individual legal consultations.
Read the full video transcript
Thank you for joining me for this second to last session. Um, so I'm Jennifer Ding. Um, I work at a data infrastructure startup in London, but today I'll be talking about some past work in something called communitydriven or collective data governance. Um, and also sharing more of a speculative future-looking project um that been discussing with Sarah and Hunlin uh my collaborators on what a toolkit could look like for communitydriven data governance. So the hope today is to just raise this idea to the CSV comp community, see what you think, see how this plays out for the kind of data sets you work with um and hopefully start a conversation. So um yesterday there was a great talk where um someone really centered open data as a verb. So open data not just as a noun but as a verb. And I think that's really what the focus of this talk today is about too. Um this is CSV comp. I know people here know the importance of data governance. They know the importance of communitydriven. But um my hope is for the next 20 minutes to to focus on um what it looks like to think of the practice of data governance. Not just data governance as a technical or legal implementation or system, but what does it look like to actually embed communities in making more decisions about data about them or data they contribute to data sets. Um and two of many sources of inspiration for this talk come from firstly the the European data governance act where in addition to goals like increasing data availability, data reuse, um one of the other goals is really this focus on increasing trust and creating a better environment that enables things like data sharing. And secondly, this paper from Sylvvita Laqua and Neil Lawrence called bottom-up data trusts disturbing the one-sizefits-all approach to data governance where they point out that in addition to the the law like regulation, GDPR, etc. Um, we also need better structures to enable bottomup data empowerment. So, how do we give a voice to data subjects so that they can say more than just yes or no or they can do more than just click a button to opt in or opt out. how can we nuance that process so people um can say more about what they want and actually see that happen. So this is roughly the plan for today. Um my background is in machine learning. So I'll share a little bit about for that context at least why we think collective data governance is so important especially now. Um most of the data sets I've worked with recently have been artistic or language data sets. So my examples will primarily be from that space. I'll share a case study from a project called the coral data trust experiment and then a few other examples of um collective data governance in the wild and then finally end with some open questions about what a toolkit might look like. So I think at this point 2025 something that's pretty clear to all of us is that the advancements in AI the reason we are here where we are today is because of the access to all the open data the public data um that has been on the web. So that data that has been contributed by so many people for decades has been the raw material to make AI what it is. And beyond those data sets themselves, it's also the data labor, right? The data workers from all over the world, primarily the global south that have prepared the data um and really enabled the advancements um that we see. So there's this great image from the better images of AI catalog where we show chat GBT I guess in this case propped up by all the people that are in the data sets the people who have worked on the data sets. Um of course this is something that is often overlooked and as a result we see that um data subjects data contributors data workers tend to have very little voice and very little control in the ML process. So um a big question we asked um was across this machine learning pipeline um where are the opportunities at different stages to invite people in to make more decisions. So at the top we have different steps along the pipeline from data collection to model training to model deployment and then on bottom in blue um we have uh some of the governance questions that are raised by each stage. Um and really what we see them as are on-ramps, right? Opportunities to invite communities in um to have a say. So in the beginning stage when it comes to data collection, a question we might ask um would be whose data is being used for training and evaluation and also whose data is missing? Um do they have a say? What kind of say do they have? and whether they want to be represented or not represented. Um moving forward along the pipeline, once a data set is procured, who has access to that data set? Um how can they update the training? And then as we move along, more questions emerge, right? Who gets to define good um or safe or um performant? Uh what who defines what the best models really are? So today I'll focus more on the data side of things and really start with this first project called the Coral Data Trust Experiment. So this project took place last year with um an art gallery in London called Serpentine. Um they had an exhibit called The Call with the artists Holly H. Hearnen and Matt Dryhurst. Um and I joined that project as the data steward um and the research team to um work with um the 15 UK choirs that were part of this project. um and get them involved, see if we could invite them to get involved in the data governance. Um and the motivation here, I think something that is really unique about asking some of these questions in artistic context is you can broaden um what you want to investigate. So in this case the artists um have this provocation as you can see in the middle where if we are in a stage of the AI ecosystem where all media anything on the web is fair game for training data it's already being scraped into training data sets. Can we turn that the production of those data sets into art instead? So that was the the provocation we had. Um we worked with 15 UK choirs from across the country um as part of this process. And what that really looked like is we started with some data conversations um talking about voice data, voice AI models, the current state-of-the-art at the time. I think um chat uh OpenAI had just released a model where uh it sounded really like Scarlett Johansson. There were a lot of really interesting case studies that we actually had to share. I think there's a country music artist named Randy Travis who after he had a stroke wasn't able to sing but then with the use of uh generative voice models was able to release a single after 10 years of of not releasing music based on his previous samples. So we really raised these case studies of where things were heading in the field um to start a conversation with the choirs, see how they felt, see how the different variations of data collection, use of data, the outcomes changed, how they felt about um whether it was a process they wanted to get involved in, whether it was a project they wanted to give their data to. So these were really our goals. I think one observation very early on was that um I think even though we are all all of us right are part of data sets together maybe one we're not aware of that um and two it is a bit strange to think of that as a form of shared identity um to be part of the same data set together um but we realized that was really a precursor to any kind of form of data governance or collective data governance there was this need to build a sense of shared identity of collective identity in uh being part of this data set together with their choirs and with the other choirs that were contributing their voices to this data set. From there we really had the opportunity to then build for some kind of collective action. Right? Once people can see themselves or identify as uh part of a data set then um there is the ability to do more together. Um something um in law that is quite helpful here too, something we wanted to investigate is in Europe and in the UK there's GDPR, right? So we do have some form of data rights hammer that we can wield, but the way that GDPR is scoped is it's very individual based. It's about personal data rights, right? And there's a lot of really great um things you can do, but uh people tend to think about enforcement on an individual level. But as we know with data and with law when you have the ability to aggregate data or aggregate rights the kind of impact you have is a lot greater. So part of the experiment was if we can do this if we can um you know create this sense um of collective identity build this capacity can we also explore ways to exercise data rights as a collective. From there um after all of that is set you know then from there we have an opportunity to explore developing with the choirs um what the right governance tools are for this particular data set and as my colleague um from Serpentine well described it really what we were trying to do is live action roleplay what good faith AI development could look like. So um I think I talked too much and I only have five minutes left so I'll I'll go quickly now. Um, we had a lot of really interesting conversations with the choirs. One thing I'll just highlight is there was definitely a mental shift that had to happen from thinking about voice then as a recording and then as data in different contexts whether it's the art, the performance, um, the law or you know when we're talking about the technical implications using one term versus the other really made sense even if in the end um, this was the same artifact. But ultimately um there was an interest um once folks had a little bit more understanding of the context to make some decisions. So one um case study I'll highlight is we used a platform called Pace. Just curious if folks here have heard of Pace. Amazing. Okay. It's this great um um preference gathering and and deliberation and alignment tool. Um, and we found it to be an interesting way to then surface across the different choirs um, how they felt about sharing their data. You know, when you just throw licenses and terms at people, that doesn't really make sense. But if you give them statements that they can vote on and they can submit further statements through that, we were able to identify different preference groups and also potential licenses that would be a good fit if we were ever to release the model. And for them, it really came down to not wanting to share the data set for commercial purposes or publicly for any possible use case given how quickly things are changing in AI. So in the wild, I'll just spotlight a few other interesting examples of collective data governance happening in the language data space because art language culture is so personal. Um this is an area where thinking about tools for data governance, community data governance is especially important. So the first is from Coher Labs. We have this interesting project called YA where they've identified and are trying to bridge this multilingual AI gap um and bridge that digital divide but crucially they want to do it with the communities from all around the world. So they've created this open science project with thousands of researchers to collect the data together, open the data so it's available for reuse by those communities however they see. But it doesn't end there. It's not just about extracting the data. It's then about building the models, the tools, the applications together and releasing all of that openly too. Another interesting case study is from the big science workshop which created one of the first open-source LLMs called Bloom. Um, and I think something interesting about what they did is they they just recognized that currently there isn't really a good set of roles and responsibilities in the data ecosystem for language data management. So they proposed this data stewardship ecosystem um inspired by a lot of existing work in in open data spaces um and um really connecting from different disciplines from tech to to law to rights to um a lot of glam um players. And finally the the organization or the the initiative I'd like to highlight is something called misilla uh common voice which is a um wonderful initiative uh to address that language uh gap as well. Um so contributors can contribute text or audio data and something very cool that they are releasing in five days is something called the misilla data collective. So they recognized uh originally you know it's Mozilla um all the data sets were released openly but um in face of the current wave in AI um they've started to have some workshops with their communities to see how different communities feel what are their preferences for sharing what are the conditions um they would like to share and also how would they like to be uh compensated attributed things like that to have more nuance um so that launches soon and definitely recommend checking it out I think it'll be a really interesting extension ion of common voice. So I think maybe the final slide I'll just end on is you know we're still in the very early stages of thinking about what a toolkit would look like. But I think the main thing we want to do is make um sharing data governance resources toolings and learnings as easy as is to share uh data software models today. Uh because there is so much that is captured in um specific organizations and initiatives. I've learned about so many over the last two days and there's really just this wealth that um we think it would be really helpful to aggregate so we can start to do things like capacity building, consensus building, um share legal guidance that different organizations have found out so folks don't need to spend thousands each time to go through the same exercise with their own lawyers. Um and you know when it comes to tools like polace and uh talk to the city there are these really wonderful deliberation platforms but the onboarding barrier can be very high to understand how you even start to set up a good participatory or deliberative exercise. So this is really the starting point and hopefully we have a chance to chat more um after this about um yeah what this could look like. Am I okay on time? We have a lot of time for questions from the audience. Yep. Amazing work. Um I I I wonder are you also collaborating with the ML commons folks? Uh they have like a data set working group where they talk about prioritizing helping machines learn on secondary languages on uh yeah just wondering if there's because there seems to be several par parallel initiatives happening in the community. absolutely love the work of ML Commons. I think we have mostly spoken to Croissant. Um but agree that there's other subgroups in ML Commons that would be great to connect with. I think so far we found the closest um practical alignment with Misilla and the the um the platform that they're building, the data collective, but um agree and and thanks for for calling that out. >> Anybody else uh has any following a question or comment. Um, I was interested to see music come in because I actually what I spend more of my time doing is making software musical instruments. um uh and things that's been coming up in that world is kind of this this question of the discuss around around machine learning and AI is language in art out right uh but there actually seems to be quite a lot of interest among artists for just tools like they don't want to abandon a creative process that they already have they just want to augment it with new different types of tools so I'm curious how your experience with the coral trust experiment. Um, you know, whether whether there was anything along those lines that you discovered there. >> Absolutely. So, Holly and Matt, the artists in the exhibit, are not just artists, but they also have created this tech startup, I guess, called Spawning, where they're focused on that, creating tools specifically for artists. So, instead of artists retrofitting tech tools that aren't really purpose-built for art, um, they're thinking a lot about that. I think one interesting thing that folks mentioned was that in some cases what you would call a hallucination in one context is actually the point in an artistic context. You want something different or um provocative or strange or unexpected to come up. So it's interesting to think how we optimize for certain objectives that might yeah harm others. Um what you describe also reminds me of something that Ted Chiang actually brought up at the this workshop on creativity and AI where um I mean some may agree or disagree but his thinking is the current way of prompting models. So just with a sentence, right, a text is a really low feature, a low complexity way to generate art, which is why he thinks models don't generate good art right now, right? Um there's a lot of levers for artists to actually tweak and and and um get creative with and and um curate, right, the the final output. So thinking about what the right tooling, the right features, the right um options I guess to expose to artists so they can create art in the way that they want likely probably isn't just prompt in art out. Um I think is a really interesting question. using uh well working with choirs to sort of bring the the collective action and the collective creation aspect to the discussion. Uh it that that's what really stuck with me with your talk. So thank you very much for that. That's I'm going to be noodling that for a while. But uh h did you did you consider other sort of collectives other collective creatives um for this sort of work or or for future work? Um and if so what are they? Because I want to think more about this. definitely agree with you on the first point. Um, one thing that I don't know if it was intentional. I think Holly wanted to work with choirs. There's a lot of community choirs in the UK, so it worked out. But one really interesting thing about engaging an existing community, especially for something like data governance, um, is that they already have their own governance, which is a really helpful scaffold that you can then leverage as you start to talk about, you know, quite complex topics like, you know, who should you share your data with, you know, they have like um a frame of reference and an identity in which they already know how to make decisions. So some of them had like um a parliament every year where they like decide the vision in the future of the company like pure democracy. Others were more um you know benevolent dictator kind of vibe. Um so I think that is definitely that something that has stuck with me as well just um especially if you are a technologist or an organization coming into a space um what is the responsible way to engage people maybe as individuals isn't isn't the powerful way uh but if you can find partnerships with existing entities um existing communities it's a it's a much better starting point um not sure about your second point But if something jumps to mind, um, I'll find you after. >> Agree. Agree. That's a great knob. >> Yeah. Okay. If there is no other question, uh, let's give another round of applause to Jennifer.