Jennifer Ding, Dr Sarah Ann Gilbert, Dr Hanlin Li - A Toolkit for Community-Driven Data Governance
Watch on YouTubeVideo summary
The core subject of this presentation is the shift from viewing data governance merely as a technical or legal compliance task to treating it as an active, community-driven practice. Jennifer Ding, drawing on her background in machine learning and recent work with artistic datasets, argues that the current AI ecosystem relies heavily on open public data and the labor of often overlooked contributors, particularly from the global south. To address the lack of voice and control these individuals possess, the talk proposes a framework where communities are embedded directly into decision-making processes throughout the entire data pipeline, from collection to model deployment. This approach is inspired by initiatives like the European Data Governance Act and academic work on bottom-up data trusts, aiming to move beyond simple opt-in/opt-out buttons toward nuanced systems that allow data subjects to express specific preferences regarding how their contributions are used.
A primary example discussed is the Coral Data Trust Experiment, a project conducted with the Serpentine Gallery in London involving fifteen UK choirs contributing voice data for AI training. The experiment sought to foster a collective identity among these diverse groups so they could exercise their data rights collectively rather than individually under GDPR. Through conversations and the use of deliberation tools like Pace, the choir members engaged in "live action roleplay" scenarios to determine what good-faith AI development looks like for them. A significant outcome was a mental shift where participants moved from viewing voice merely as raw recordings to understanding it as complex data with legal and artistic implications. The choirs ultimately expressed strong reservations about sharing their data for commercial purposes or unrestricted public use, highlighting the need for governance tools that respect these nuanced preferences rather than imposing standard licenses.
The presentation also highlights other real-world examples of collective data governance in the language space, such as Coher Labs' YA project, which aims to bridge multilingual gaps by building models and tools openly with global communities, and Mozilla's Common Voice initiative, which is evolving into a "data collective" to better handle community preferences on compensation and attribution. Additionally, the Big Science workshop's creation of the Bloom model introduced a data stewardship ecosystem that connects technology, law, and rights holders. The speaker concludes by emphasizing the necessity for a toolkit that aggregates these diverse learnings and resources, making it as easy to share governance guidelines and legal frameworks as it is to share software models today. Such a toolkit would lower barriers to entry for participatory exercises, enabling capacity building and consensus across different organizations so they do not have to reinvent the wheel or spend excessive resources on individual legal consultations.
Read the full video transcript
Thank you for joining me for this second
to last session. Um, so I'm Jennifer
Ding. Um, I work at a data
infrastructure startup in London, but
today I'll be talking about some past
work in something called communitydriven
or collective data governance. Um, and
also sharing more of a speculative
future-looking project um that been
discussing with Sarah and Hunlin uh my
collaborators on what a toolkit could
look like for communitydriven data
governance. So the hope today is to just
raise this idea to the CSV comp
community, see what you think, see how
this plays out for the kind of data sets
you work with um and hopefully start a
conversation.
So um yesterday there was a great talk
where um someone really centered open
data as a verb. So open data not just as
a noun but as a verb. And I think that's
really what the focus of this talk today
is about too. Um this is CSV comp. I
know people here know the importance of
data governance. They know the
importance of communitydriven. But um my
hope is for the next 20 minutes to to
focus on um what it looks like to think
of the practice of data governance. Not
just data governance as a technical or
legal implementation or system, but what
does it look like to actually embed
communities in making more decisions
about data about them or data they
contribute to data sets. Um and two of
many sources of inspiration for this
talk come from firstly the the European
data governance act where in addition to
goals like increasing data availability,
data reuse, um one of the other goals is
really this focus on increasing trust
and creating a better environment that
enables things like data sharing. And
secondly, this paper from Sylvvita Laqua
and Neil Lawrence called bottom-up data
trusts disturbing the one-sizefits-all
approach to data governance where they
point out that in addition to the the
law like regulation, GDPR, etc. Um, we
also need better structures to enable
bottomup data empowerment. So, how do we
give a voice to data subjects so that
they can say more than just yes or no or
they can do more than just click a
button to opt in or opt out. how can we
nuance that process so people um can say
more about what they want and actually
see that happen.
So this is roughly the plan for today.
Um my background is in machine learning.
So I'll share a little bit about for
that context at least why we think
collective data governance is so
important especially now. Um most of the
data sets I've worked with recently have
been artistic or language data sets. So
my examples will primarily be from that
space. I'll share a case study from a
project called the coral data trust
experiment and then a few other examples
of um collective data governance in the
wild and then finally end with some open
questions about what a toolkit might
look like.
So I think at this point 2025 something
that's pretty clear to all of us is that
the advancements in AI the reason we are
here where we are today is because of
the access to all the open data the
public data um that has been on the web.
So that data that has been contributed
by so many people for decades has been
the raw material to make AI what it is.
And beyond those data sets themselves,
it's also the data labor, right? The
data workers from all over the world,
primarily the global south that have
prepared the data um and really enabled
the advancements um that we see. So
there's this great image from the better
images of AI catalog where we show chat
GBT I guess in this case propped up by
all the people that are in the data sets
the people who have worked on the data
sets. Um of course this is something
that is often overlooked and as a result
we see that um data subjects data
contributors data workers tend to have
very little voice and very little
control in the ML process. So um a big
question we asked um was across this
machine learning pipeline um where are
the opportunities at different stages to
invite people in to make more decisions.
So at the top we have different steps
along the pipeline from data collection
to model training to model deployment
and then on bottom in blue um we have uh
some of the governance questions that
are raised by each stage. Um and really
what we see them as are on-ramps, right?
Opportunities to invite communities in
um to have a say. So in the beginning
stage when it comes to data collection,
a question we might ask um would be
whose data is being used for training
and evaluation and also whose data is
missing? Um do they have a say? What
kind of say do they have? and whether
they want to be represented or not
represented. Um moving forward along the
pipeline, once a data set is procured,
who has access to that data set? Um how
can they update the training? And then
as we move along, more questions emerge,
right? Who gets to define good um or
safe or um performant? Uh what who
defines what the best models really are?
So today I'll focus more on the data
side of things and really start with
this first project called the Coral Data
Trust Experiment.
So this project took place last year
with um an art gallery in London called
Serpentine. Um they had an exhibit
called The Call with the artists Holly
H. Hearnen and Matt Dryhurst. Um and I
joined that project as the data steward
um and the research team to um work with
um the 15 UK choirs that were part of
this project. um and get them involved,
see if we could invite them to get
involved in the data governance. Um and
the motivation here, I think something
that is really unique about asking some
of these questions in artistic context
is you can broaden um what you want to
investigate. So in this case the artists
um have this provocation as you can see
in the middle where if we are in a stage
of the AI ecosystem where all media
anything on the web is fair game for
training data it's already being scraped
into training data sets. Can we turn
that the production of those data sets
into art instead? So that was the the
provocation we had. Um we worked with 15
UK choirs from across the country um as
part of this process. And what that
really looked like is we started with
some data conversations um talking about
voice data, voice AI models, the current
state-of-the-art at the time. I think um
chat uh OpenAI had just released a model
where uh it sounded really like Scarlett
Johansson. There were a lot of really
interesting case studies that we
actually had to share. I think there's a
country music artist named Randy Travis
who after he had a stroke wasn't able to
sing but then with the use of uh
generative voice models was able to
release a single after 10 years of of
not releasing music based on his
previous samples. So we really raised
these case studies of where things were
heading in the field um to start a
conversation with the choirs, see how
they felt, see how the different
variations of data collection, use of
data, the outcomes changed, how they
felt about um whether it was a process
they wanted to get involved in, whether
it was a project they wanted to give
their data to. So these were really our
goals. I think one observation very
early on was that um I think even though
we are all all of us right are part of
data sets together maybe one we're not
aware of that um and two it is a bit
strange to think of that as a form of
shared identity um to be part of the
same data set together um but we
realized that was really a precursor to
any kind of form of data governance or
collective data governance there was
this need to build a sense of shared
identity of collective identity in uh
being part of this data set together
with their choirs and with the other
choirs that were contributing their
voices to this data set. From there we
really had the opportunity to then build
for some kind of collective action.
Right? Once people can see themselves or
identify as uh part of a data set then
um there is the ability to do more
together. Um something um in law that is
quite helpful here too, something we
wanted to investigate is in Europe and
in the UK there's GDPR, right? So we do
have some form of data rights hammer
that we can wield, but the way that GDPR
is scoped is it's very individual based.
It's about personal data rights, right?
And there's a lot of really great um
things you can do, but uh people tend to
think about enforcement on an individual
level. But as we know with data and with
law when you have the ability to
aggregate data or aggregate rights the
kind of impact you have is a lot
greater. So part of the experiment was
if we can do this if we can um you know
create this sense um of collective
identity build this capacity can we also
explore ways to exercise data rights as
a collective. From there um after all of
that is set you know then from there we
have an opportunity to explore
developing with the choirs um what the
right governance tools are for this
particular data set and as my colleague
um from Serpentine well described it
really what we were trying to do is live
action roleplay what good faith AI
development could look like. So um I
think I talked too much and I only have
five minutes left so I'll I'll go
quickly now. Um, we had a lot of really
interesting conversations with the
choirs. One thing I'll just highlight is
there was definitely a mental shift that
had to happen from thinking about voice
then as a recording and then as data in
different contexts whether it's the art,
the performance, um, the law or you know
when we're talking about the technical
implications using one term versus the
other really made sense even if in the
end um, this was the same artifact. But
ultimately um there was an interest um
once folks had a little bit more
understanding of the context to make
some decisions. So one um case study
I'll highlight is we used a platform
called Pace. Just curious if folks here
have heard of Pace. Amazing. Okay. It's
this great um um preference gathering
and and deliberation and alignment tool.
Um, and we found it to be an interesting
way to then surface across the different
choirs um, how they felt about sharing
their data. You know, when you just
throw licenses and terms at people, that
doesn't really make sense. But if you
give them statements that they can vote
on and they can submit further
statements through that, we were able to
identify different preference groups and
also potential licenses that would be a
good fit if we were ever to release the
model. And for them, it really came down
to not wanting to share the data set for
commercial purposes or publicly for any
possible use case given how quickly
things are changing in AI.
So in the wild, I'll just spotlight a
few other interesting examples of
collective data governance happening in
the language data space because art
language culture is so personal. Um this
is an area where thinking about tools
for data governance, community data
governance is especially important. So
the first is from Coher Labs. We have
this interesting project called YA where
they've identified and are trying to
bridge this multilingual AI gap um and
bridge that digital divide but crucially
they want to do it with the communities
from all around the world. So they've
created this open science project with
thousands of researchers to collect the
data together, open the data so it's
available for reuse by those communities
however they see. But it doesn't end
there. It's not just about extracting
the data. It's then about building the
models, the tools, the applications
together and releasing all of that
openly too.
Another interesting case study is from
the big science workshop which created
one of the first open-source LLMs called
Bloom. Um, and I think something
interesting about what they did is they
they just recognized that currently
there isn't really a good set of roles
and responsibilities in the data
ecosystem for language data management.
So they proposed this data stewardship
ecosystem um inspired by a lot of
existing work in in open data spaces um
and um really connecting from different
disciplines from tech to to law to
rights to um a lot of glam um players.
And finally the the organization or the
the initiative I'd like to highlight is
something called misilla uh common voice
which is a um wonderful initiative uh to
address that language uh gap as well. Um
so contributors can contribute text or
audio data and something very cool that
they are releasing in five days is
something called the misilla data
collective. So they recognized uh
originally you know it's Mozilla um all
the data sets were released openly but
um in face of the current wave in AI um
they've started to have some workshops
with their communities to see how
different communities feel what are
their preferences for sharing what are
the conditions um they would like to
share and also how would they like to be
uh compensated attributed things like
that to have more nuance um so that
launches soon and definitely recommend
checking it out I think it'll be a
really interesting extension ion of
common voice. So I think maybe the final
slide I'll just end on is you know we're
still in the very early stages of
thinking about what a toolkit would look
like. But I think the main thing we want
to do is make um sharing data governance
resources toolings and learnings as easy
as is to share uh data software models
today. Uh because there is so much that
is captured in um specific organizations
and initiatives. I've learned about so
many over the last two days and there's
really just this wealth that um we think
it would be really helpful to aggregate
so we can start to do things like
capacity building, consensus building,
um share legal guidance that different
organizations have found out so folks
don't need to spend thousands each time
to go through the same exercise with
their own lawyers. Um and you know when
it comes to tools like polace and uh
talk to the city there are these really
wonderful deliberation platforms but the
onboarding barrier can be very high to
understand how you even start to set up
a good participatory or deliberative
exercise. So this is really the starting
point and hopefully we have a chance to
chat more um after this about um yeah
what this could look like. Am I okay on
time?
We have a lot of time for questions from
the audience.
Yep.
Amazing work. Um
I I I wonder are you also collaborating
with the ML commons folks?
Uh they have like a data set working
group where they talk about prioritizing
helping machines
learn on secondary languages on uh yeah
just wondering if there's because there
seems to be several par parallel
initiatives happening in the community.
absolutely love the work of ML Commons.
I think we have mostly spoken to
Croissant. Um but agree that there's
other subgroups in ML Commons that would
be great to connect with. I think so far
we found the closest um practical
alignment with Misilla and the the um
the platform that they're building, the
data collective, but um agree and and
thanks for for calling that out.
>> Anybody else uh has any
following a question or comment.
Um, I was interested to see music come
in because I actually what I spend more
of my time doing is making software
musical instruments. um uh and things
that's been coming up in that world is
kind of this this question of the
discuss around around machine learning
and AI is language in art out right uh
but there actually seems to be quite a
lot of interest among artists for just
tools like they don't want to abandon a
creative process that they already have
they just want to augment it with new
different types of tools so I'm curious
how your experience with the coral trust
experiment. Um, you know, whether
whether there was anything along those
lines that you discovered there.
>> Absolutely. So, Holly and Matt, the
artists in the exhibit, are not just
artists, but they also have created this
tech startup, I guess, called Spawning,
where they're focused on that, creating
tools specifically for artists. So,
instead of artists retrofitting tech
tools that aren't really purpose-built
for art, um, they're thinking a lot
about that. I think one interesting
thing that folks mentioned was that in
some cases what you would call a
hallucination in one context is actually
the point in an artistic context. You
want something different or um
provocative or strange or unexpected to
come up. So it's interesting to think
how we optimize for certain objectives
that might yeah harm others. Um what you
describe also reminds me of something
that Ted Chiang actually brought up at
the this workshop on creativity and AI
where um I mean some may agree or
disagree but his thinking is the current
way of prompting models. So just with a
sentence, right, a text is a really low
feature, a low complexity way to
generate art, which is why he thinks
models don't generate good art right
now, right? Um there's a lot of levers
for artists to actually tweak and and
and
um get creative with and and um curate,
right, the the final output. So thinking
about what the right tooling, the right
features, the right um options I guess
to expose to artists so they can create
art in the way that they want likely
probably isn't just prompt in art out.
Um I think is a really interesting
question.
using uh well working with choirs to
sort of bring the the collective action
and the collective creation aspect to
the discussion. Uh it that that's what
really stuck with me with your talk. So
thank you very much for that. That's I'm
going to be noodling that for a while.
But uh h did you did you consider other
sort of collectives other collective
creatives
um for this sort of work or or for
future work? Um and if so what are they?
Because I want to think more about this.
definitely agree with you on the first
point. Um, one thing that I don't know
if it was intentional. I think Holly
wanted to work with choirs. There's a
lot of community choirs in the UK, so it
worked out. But one really interesting
thing about engaging an existing
community, especially for something like
data governance, um, is that they
already have their own governance, which
is a really helpful scaffold that you
can then leverage as you start to talk
about, you know, quite complex topics
like, you know, who should you share
your data with, you know, they have like
um a frame of reference and an identity
in which they already know how to make
decisions. So some of them had like um a
parliament every year where they like
decide the vision in the future of the
company like pure democracy. Others were
more um you know benevolent dictator
kind of vibe. Um so I think that is
definitely that something that has stuck
with me as well just um especially if
you are a technologist or an
organization coming into a space um what
is the responsible way to engage people
maybe as individuals isn't isn't the
powerful way uh but if you can find
partnerships with existing entities um
existing communities it's a it's a much
better starting point um not sure about
your second point But if something jumps
to mind, um, I'll find you after.
>> Agree. Agree. That's a great knob.
>> Yeah.
Okay. If there is no other question, uh,
let's give another round of applause to
Jennifer.