The Collaborative Power of a Research Data Commons: LDaCA case studies
Watch on YouTubeVideo summary
The first webinar of the House and Indigenous Research Data Commons series focused on the Language Data Commons of Australia (LDaCA), an initiative dedicated to securing nationally significant language data under community control. The session addressed critical challenges such as disorganized metadata, inaccessible collections, and a lack of user-friendly tools by presenting four distinct case studies that demonstrated LDaCA's capabilities. These examples included enriching metadata for the Carolyn Tennant Kelly papers, which revealed previously hidden language groups; utilizing the Ningan platform to transcribe and tag fragile field notes on the Gungloo language; employing advanced search filters to locate specific multimodal data like gestures alongside speech; and using automated Jupyter notebooks to analyze media coverage of female athletes in Australian sports.
Beyond these technical demonstrations, the discussion highlighted how LDaCA fosters a collaborative ecosystem that democratizes access to research tools and data for both quantitative and qualitative studies. The platform offers a unique middle ground between fully open repositories and restricted institutional access by implementing granular control mechanisms where community data custodians determine who can access specific materials and under what conditions. This approach allows for "respectful sharing," enabling sensitive content, such as interview recordings, to be viewed without being downloaded locally, thereby addressing global challenges related to data licensing and usage agreements while maintaining cultural integrity.
The session concluded with a Q&A that covered essential topics including data citation practices, combined text and media searches, and the management of consent through data redaction features. A significant portion of the dialogue addressed community concerns regarding permissions for language use and AI training, clarifying that while a central portal exists, communities retain the option to create separate versions or pilot programs to preserve their autonomy. The moderator emphasized that this infrastructure supports a balanced model where access is neither completely open nor closed, ensuring that sensitive materials can be utilized responsibly without compromising community trust or privacy.
To support ongoing engagement and further exploration of these tools, the organizers shared key resources including the LDaCA data portal and contact information for the initiative. Attendees were informed about upcoming events designed to deepen their understanding of the platform's capabilities, such as a hands-on workshop on text analytics using Word Flow scheduled for late August and a webinar introducing the Australian Internet Observatory platform in late September. Ultimately, the webinar reinforced LDaCA's role as a vital infrastructure that empowers communities to manage their own data sovereignty while facilitating broader research collaboration across Australia.
Read the full video transcript
Well, hello everybody and welcome. Uh,
thank you so much for joining us today.
Uh, my name is Mary Philsell. I'm an
engagement representative uh, for the HS
and indigenous research data commons and
I'm also your host for this webinar
series as well. So today's session the
collaborative power of a research data
commons um in is the very first uh
webinar that we have in our house and
indigenous research data commons uh
webinar series and before we get started
um I'd like to acknowledge all of the
traditional custodians of the country
and waters that we meet on today and I'd
like to pay my respects to elders past
and present and I also like to extend
this respect to all indigenous people
joining us here today and uh acknowledge
the deep and ongoing connection to
countries shared across indigenous
Australia. So I myself I'm a white
settler person and I work and live on
Ghana land and uh the local greeting in
uh Ghana country is Nina Mani. So Nina
Mani everybody uh it's an honor to be
joining you here online today. And uh
before we before we meet our speakers,
I've got just a couple of little
housekeeping notes to share with you. So
first of all, this session is being
recorded. It's currently being recorded
and this recording will later be uh
shared with registrants and and uploaded
to the uh ARDC YouTube channel after
that. So, if you have any questions
along the way, please feel very welcome
to post them into the Zoom chat and
we'll be able to come back and answer
some of them um or all of them uh
together at the end. Just pining on
time. So, if we run over time, we may
not be able to get to your particular
question
and I've just realized I just need to
apologies I have to progress my slides
and that's just a note about the
recording. So this is our code of
conduct for today and the um the link if
you if you would like should come up in
the chat the link is uh e- research
ardcc
and we would like to keep this a
welcoming and respectful space for
everyone and we'd like to ask all
attendees uh to follow the ARDC code of
conduct which is available at this link.
Okay. So yes, the house and indigenous
research data commons webinar series um
is one of our newest offerings from the
house and indigenous research data
commons and this session is part of the
series and this is a monthly series for
H researchers featuring case studies
tool demonstrations and hands-on skill
training as well. So this series is led
by the ARDC and funded by ENIRS and
ENIRS stands for the National
Collaborative Research Infrastructure
Strategy. And so the Hassan Indigenous
Research Data Commons supports six focus
areas. And today's session focuses on
one of those focus areas which is the
language data commons of Australia or as
we often call them ELDAC by the acronym.
So today's session
is the collaborative power of a research
data comments Eldhaka case studies and
joining me today are Michael Moore from
Eldaka uh Robert Mlelen from Eldaka,
Simon Musgrave also from Eldaka, Gary
Tuda Smith from the University of
Queensland and Melissa Kembell from the
University of Sydney
And a little note about what uh how how
today is going to help you do better
research uh through this webinar series.
So you'll be able to access
collaborative tools and platforms, use
secure nationally significant data
collections and analyze data without
needing custom coding or setup. So now
I'm going to hand over to Michael Hall
who unpack exactly what a research data
commons is and what ELDA can do for you
and your research. So, Michael, over to
you.
>> Thanks very much, Mary. Um, so it's my
great pleasure to be here today. I think
um Simon, you're just working on the
slides in the meantime, but um while
you're putting them up, I just wanted to
acknowledge the lands of the younger and
durable peoples u from which I'm uh
beaming in here today. Um I always feel
uh acknowledging our country is such a
special thing to be doing in this
country. We have so many uh languages
stretching back so many tens of
thousands of years. It's it's really
stunning to celebrate but also to
acknowledge um where we are today with
many of those languages too. Um so uh
Aldaka what are we um so our aim really
is to be working with nationally
significant language data in Australia
and its region. Um we're working to make
that data accessible in appropriate ways
um with community control. And when we
talk about language data, we're thinking
about language data both in the
Australian continent but also in in the
region uh in which we're in. And the
reason we're focusing on that is because
about a quarter of the world's languages
are in Australia and its region region.
Um Australia itself is the fifth most
linguistically diverse nation on earth.
Um and of course we're home to uh the
world's uh oldest continuing cultures as
well. So it's a really really important
part of of of the international scene
the the global society as well. Um and
our aims are to be I guess securing uh
really important collections of data. Um
but also um I think today we'll be have
a big emphasis on using uh the language
data and how we put it to use. Uh next
slide Simon.
Um so when we think about language data
um even though it's so important even
though a quarter of the world's
languages are in our region and and
Australia has considerable
responsibility uh to be doing something
about it. Um language data is very often
not well organized. It's not organized
in ways that make it usable, reusable
and repurposable. Um there's a lot of at
risk uh collections. We hear about this
all the time. It's hard to find things.
Um accessing them is complex. uh uh
sometimes it's for very good reasons and
sometimes it's for less good reasons. Um
and then the tools by which we analyze
and and work with data in researchers
and communities varies and often it's
it's difficult to work in ways that are
uh transparent and reproducible and so
on. And finally um guidance and support
for working with language data are very
often scattered and hard to find. Um
Simon next.
Um so in dealing with that aldaka we
started uh back in 222 we've been
working for a while um we divide our
work into variety of streams uh which is
uh quite fundamental is a ecosystem of
language data repositories uh they're in
the plural I might add um working on
securing nationally significant language
data collections increasing accessibil
access accessibility and enriching that
data and also working on tools but true,
easy to use because there's lots of
things you can do with computers these
days. Um, but they're not if you're not
a computer scientist or computer wiz um
you can't you can't necessarily do those
things. So, LDAC is about democratizing
access to those really important tools
which leads us into that fifth stream
which is around engagement and training.
So, next slide. Thanks.
Um here I just wanted to give you a
sense of the kind of complexity um but
also the simplicity of the Aldaca
technical infrastructure ecosystem. Um
so when I say complex um you can see you
probably can't read it that well unless
you've got it really enlarged on your
computer screen. Um but there's a whole
bunch of different uh organizations and
collections and portals that we work
across. Right. So in that sense it's
complex um because our aim you know any
research data commons is really a
collaborative ecosystem which enables
researchers communities and
organizations to work together right so
rather than working in silos we're
working collaborative collaboratively
across those boundaries and and you know
doing things that we couldn't do um if
we didn't cross those boundaries. Um so
in that sense there's a complexity but
there's also simplicity behind what we
do because we want it to be maximally
promoting autonomy for researchers and
communities and we want to be promoting
sustainability. Um so technical systems
which are very difficult to work with uh
very kind of heavy uh and don't promote
autonomy or sustainability right um and
so today's not the day to get into the
technical architecture but if you know
arrocate is the word uh to remember if
you want to look it up and find out more
um but today our aim really is thinking
about how we put this uh I guess the
infrastructure in into practice how we
implement it, how people are are using
it and and putting it to use. Um, and a
big part of that is our I guess our
careful fairness approach. Um, care and
fair of course the important principles
and the ones that we want to draw out
today I think particularly are the R and
fair. Right? So that's around reusing
repurposing language data and also the C
and care. So that there's benefit from
doing that, not just for individual
researchers, uh, but for the communities
that they're working with as well. Um,
so we've got four fabulous case studies
and I'm really grateful for the four
presenters today who are going to take
us through those. Um, and I'll hand over
to Robert first of all to to lead us.
>> Thank you, Michael. And one young
everyone welcome everyone. I I'm a um
it's good to see you all. Well, I'm a
Goring Goring person from Bundberg. Uh,
but I'm also the program manager for the
language data commons. Uh, and echoing
those sentiments from Michael, this is
the first webinar uh, series for the
AODC. So, we're very privileged and
excited to be the first cab off the rank
so to speak. Uh, so I will kick off.
Thanks, Simon.
um guided by the HAS and indigenous uh
uh research data commons framework for
the governance of indigenous data. One
key tenant that's within that framework
is around uh providing knowledge,
accessibility and usability of data
assets. And what that means is it's
around ensuring that indigenous
communities are able to access uh and
benefit from relevant data holdings and
that's also with a emphasis there on
inclusivity uh transparency but also
informed consent. So I'm going to take a
little bit of a deep dive. Michael
talked about some of the language data
challenges. I'm going to take a little
bit of a deep dive into metadata. Um
it's got a good story. It's
entertaining, don't worry. Um, so when
working towards making collections more
findable, uh, the the accuracy of the
metadata descriptions is something that
is really quite vital. It becomes quite
vital. Historical indigenous language
and cultural materials, as as it's been
said, they're often poorly described, if
they're even described at all. And this
in itself presents um findability
challenges for anyone really wanting to
find a record of uh data
um
you know finding the record of the data
but then finding its existence in the
very first instance. So that that
remains a challenge that metadata is the
key to solving. So making the metadata
records available at the very least uh
is a critical step towards demonstrating
findability. So when complemented with
good metadata descriptions, metadata
availability presents an interplay or an
insightful interplay between what you
were looking for but also discovering
what you never knew existed. And that
really rings true from an indigenous
community perspective being um that we
have been historically and
institutionally marginalized from
accessing data about our people. So
whether the metadata record is well
described or poorly described, making
the record available supports three
things. First one being findability. The
second being awareness that the data
exists in the first instance regardless
of access conditions and this in turn
can support the longevity and
safekeeping of the data thereafter. Uh
the third thing of course is um it
increases opportunities for future
metadata enrichment opportunities which
I'll finish on uh after this this
session. So for today's presentation,
I'm going to extrapolate this across two
parts. We're going to look at metadata
availability and metadata enrichment.
And this is what really underpins our
experience with the Carolyn Tenant Cali
materials. Uh thanks Simon. So in 2023
there was a partnership between Eldacer
and the University of Queensland
library. This involved a pilot study
which aimed uh we were aiming to assess
the quality and the extent of indigenous
language data that was held within
university collections. Uh but we also
wanted to demonstrate how such data
could be meaningfully incorporated into
the LDAC portal. So with the pilot, we
adopted a harvesting approach uh quite
similar to the one that's used by Trove
uh with the National Library of
Australia, the Trove Service there. And
this was also supported with additional
annotations that were generated through
a prior project that had described this
particular material, the Caroline Ten
and Kelly uh collection that prior to
that was quite poorly described. So I'll
give you a bit of a run through of what
it actually is. So the Carolyn Tennant
Kelly or the CTK papers as we we've
referred to them their fieldwork
accounts of Australian Aboriginal
culture in the 1930s. So Carolyn Tenant
Kelly was an anthropologist. Um but
these papers contained a significant
amount of language data that was in uh
in Queensland at that time. So prior to
their acquisition, well they ended up
being lost. um they were in in a back
shed for quite a long time before they
were a tip off was given to the
university and a group had uh set a
project to go and work with this
material. So prior to the university
acquiring this material um they were in
many ways I guess considered orphaned
data in that they were lacking that
institutional stewardship and and a and
a structured description which we now
have. Um so the metadata enrichment in
the first instance was carried out uh in
around 2010 uh 2011 and this was done
through the creation of an index uh an
index by categories and that was done by
David Tger uh Kim D Reich Tony Jeff and
Uncle Michael Williams who um is um also
a senior adviser for LDAC. Um so this
index revealed 413 items in total. uh
and that included 121 identified
language groups. Um there were uh sorry
language entries. There were 15 keyword
entries. There were 269 place names
recognized and there were 268 personal
names identified. So this uh quite
quickly became a very rich collection.
But this was not what was represented in
uh the university library catalog. So
that was the problematic part. Uh for
instance, I said the index referenced
121 distinct langu indigenous languages
but only 22 were reflected in the
catalog. There's a problem. That means
99 languages were not reflected on that
catalog. Now, I only knew this because
I'd acquired the data in a back
background way. So, I knew of the index
that existed and I knew my language was
on there and I could see very clearly
that my language was not one of the um
22 that were listed on the on the
university catalog. So, we started
working with this material um to to
bring that to to surface that uh that
information uh to make it more findable
for other communities. The problem of
course is that 99 languages not being
findable in the catalog is um locks 99
language groups out of accessing and
work finding accessing and working with
that uh that um material. So you can see
that's that's quite a bit of a struggle.
So the index also enabled a bit a more
nuance navigation across the collection
um through the different additional
descriptors as I said place names, key
uh keywords, names of language speakers.
So things that we can start to pick up
and work with and repurpose or
communities uh wanting to steward this
data in their own way can start to work
with um some of that enriched material
and they're not starting uh from
scratch. This is not to say conclusively
that the enrichment work that has been
done on the collection is all there is.
You know, there's always room for
improvement. But it uh as a first step
this uh the remarkable part of this
project was surfacing that and making
that findable for communities. And you
can see in the screenshot uh there
you'll see different uh languages now
referenced within the LDAC test portal.
And uh that QR code will take you to
just a short a short maybe two-page uh
brief on the uh on the LDAC UKQ library
collaboration.
So that's the CTK case example and I am
now going to hand over to my colleague
Gar.
Thanks Robert. Good day everyone. I'm
Gary Tudtor Smith. I'm a um gangaloo bar
and growing growing person from central
Queensland. Um, I'm a PhD candidate at
the University of Queensland and I also
work part-time at the research unit for
indigenous language at the University of
Melbourne. Thanks, Simon.
So, a little bit of background. Um, so
about seven or eight years ago, um, my
cousin Thomas Watson was working with,
um, linguist Andrew Tanner at, uh,
living languages. So, he was looking
into reawakening Gongaloo language,
learning it for himself, and sharing it
with his family. So together they
tracked down um a bunch of different
documentation and descriptions and one
of these was this linguistic survey of
southeastern Queensland by a Swedish
linguist called Nils Homemer. So this
has got um a bunch of different
languages. You can see the list on the
right there. They've got some short
descriptions and word lists to go along
with it. So it's useful to an extent but
there's not a lot of um examples of
sentences which you know community can
use to learn the language and to develop
teaching and learning programs. Thanks
Simon.
So um in the 60s and 70s um Nilsmer came
to New South Wales and Queensland um
with the establishment of IATIS and um
the publication I showed on the previous
page uh was from 1983 and so there were
a couple of others but this is the one
that contains um Gunglo language and so
it wasn't until 2022 when um Homer's
field notes were rediscovered so Tom
reached out to Neil's son Arthur Homer
who's a linguist in his own right in
Sweden and um got in touch and Arthur
wasn't aware that there were any field
notes but he eventually tracked them
down and scanned and emailed them
through and so across all of the
languages there was 26 notebooks and
2660 pages. So hopefully the image might
show up if you click again Simon.
Yeah, this is a um example
[clears throat] of the field notes. So
you can see each line has the language
and then the English and each line is
tagged for the speaker. The next slide.
So um I think it was 22 or 23 myself um
Tom Watson and Paul Williams um with
support from LDAC um use this Ningham
platform to transcribe these field
notes. So, as you know, many people
might know who've worked with um
archival manuscripts, it's a bit of a
challenge to get them off the
handwritten paper into some sort of
database, you know, being able to read
them well and understand what's there
and making them searchable. So, Ningon
is designed to, you know, help this
problem. And one of the nice things you
can do with this as well is you can put
some additional tags and um information
in there. So you can see on line three
there's Kumo Bullaru. I've tagged for a
place name. Um in the fir in the second
line um I've also tagged the speaker
name and I've added the speaker name in
there cuz it's just got his initials
there. So it looks a little bit daunting
with the code, but all you have to do is
highlight it and press these blue
buttons and just inserts the code for
you. The next slide.
So this is what it ends up looking like.
And you can see the highlighted part is
what I've added um the the supplied
information.
Next slide.
And so this is what it ends up looking
like once it's been pushed to the
repository. So, this is a different
manuscript, but this is an interesting
example um because all of the
translations of the language language
here um can't remember off the top of my
head which language it is, but it's
somewhere over in um Western Australia
at a mission called New Norsia. All of
the translations are written in Italian.
So, someone has gone through translated
them and put them as a supplied
information in the highlighted yellow
there. So, there's some nice things you
can do to add that extra context there
and also just make it easier to discover
once you've tagged those people people's
names and place names there. Next slide.
So, some of the benefits of Ningan. Um,
one thing is preservation. If we've got
people handling original materials all
the time, they eventually degrade some
of these manuscripts that people are
working with. These these aren't these
are only a couple of decades old, but
there's other ones that are over a
hundred years old. the papers already
crumbling each time someone handles
them. We're putting acid and oil from
our hands on them. Um, we're also saving
time and money by not having to travel
to Canra or, you know, wherever a
manuscript might be kept to go look at
them. Um, and we can also collaborate
remotely with other people.
Part of the Ningan platform is um the
permissions are determined by um by the
community. So that's part of the
workflow when things get pushed to the
repository and they can also always be
changed. Um so there's that kind of
iterative ongoing consent as well. Um as
I mentioned the legibility is really
important when you're working with um
language materials. You want to make
sure you're analyzing what's written
what's actually written there. Um, and
again the shared shared workflow, um, it
just makes it easier to find stuff. And
I think a really important thing is
that, um, each if each person, each
individual has to access material,
transcribe, you know, or digitize,
photograph, and then transcribe and then
analyze, we can't get to the more
important stuff. So we can get the you
know this foundational work out of the
way and then share it amongst our other
community members and you know if we're
working with other academics linguists
as well. So next slide.
So just to cap off some of the things
I've been using these homework field
notes for is um created a learner guide.
If you click again an image should pop
up.
Yep. So that's a sample there and um you
can see each of the examples in the
orange boxes um are sentences from
homeless field notes. Um so it kind of
gives that veracity as well and you know
if there's any um sort of speaker
variation and people can choose oh well
I'm an Aubry so I'm going to follow what
Tim Aubry says for example.
Um, we're also compiling a dictionary
and currently develop developing a
teaching and learning program and I'm in
my second year now of my PhD um doing
some further analysis into these
materials and looking at where there are
gaps in the documentation and how to
fill those gaps.
So, I'll pass on now to I think Simon's
up next. Thanks,
Thanks Gary.
The I'm joining today from the unseated
lands of the Bunarong peoples of the
Cooland nation and I pay my respects to
their elders past and present.
case study I'm going to go through now
is an example of
how researchers can find quite specific
kinds of data using some of the
facilities that we're developing.
The example I'm using here is some work
that's done by former colleagues of mine
at me university uh and one of them
Harriet Shepard worked for Eldaka for a
time as well. These scholars are
interested in the question of how
gesture reinforces or co-occurs with
language.
In particular, they're interested in a
group of words which they call caused
accompanied motion verbs. Now, these are
words which denote someone moving but
some other entity moving along with
them. So, it's accompanied motion. And
the prototypical verbs of this type of
bring and take and carry. If you think
about them, these words where someone
goes somewhere but or something goes
somewhere but something else goes along
as well.
So
these scholars wanted to investigate
they're investigating these uh words and
the gestures that go with them cross
linguistically and they were keen to
have data from Australian English.
Now the question here is is the data
with video because of course you need
that to see the uh the gestures and also
transcription because otherwise you're
going to have to look at all the video
to find the examples.
So now I'm going to show you screen
capture of how you might go about
finding this kind of data using our
discovery portal.
It's happened
There we go. Okay. So, this is the where
you enter a portal. Here we want to find
out if there's a video. So, we look for
the record type and filter on that.
And you can see that there's 168 video
items in the collection. And there's
also 168 of them in the braided channels
collection. So that tells us that all
video material is in that particular
collection. So now we want to filter on
that collection just to check that there
are transcripts. And you can see says
there's recordings and transcripts.
So we know that we're heading in the
right direction. And now we can search
for the particular words we're
interested in this data. I'm using
advanced search here because I'm going
to search for multiple items. and also
because that means I can use wildcard
characters in the search. So that first
search term that's going to look for
carry and carries and carrying and so
forth. Then I'm going to link all these
searches with the or operator because we
want full we want to meet one condition
not all of them. Um as you can see
another wild card there that we get take
takes taking but took we have to enter
separately.
You could use regular expressions here,
but it syntax gets a little bit
complicated. I wasn't feeling confident
about doing that.
But we're specifying all these search
items or search conditions and then we
can have a look and we find there's 150
examples potentially that may be useful
in this research. We can look down and
see yes, we're picking up the various
forms of these verbs. We're getting some
false positives, but that's not too
surprising. We can deal with that. We're
getting some more kind of metaphorical
uses of the verbs, but that's okay. So,
let's have a look at the first example
here. See what's going on. We can look
at the whole transcript. Just check
what's going on. And we search in here
for took, which was first instance. And
there we got this thing which actually
has took and brought close together. So,
this looks really interesting. This
could be great data. We can see that
there there's at least some time codes
in these transcripts. So, we're within a
kind of three minute span that we're
going to be looking for that one, which
is better than looking through half an
hour, right? So, now we've established
this looks interesting. We can download
the data then see what it looks like.
So,
let's
move on.
Sorry. Okay. So, I'll just play you the
video corresponding to the example that
we found.
Now, as you can see, unfortunately,
this is not a useful example for us
because for understandable reasons, the
filmmaker zoomed in on that and you
couldn't see the speaker's hands when
they actually used the verbs make and
bring.
But that's research. As it happens in in
fact though the research team went
through this material and found all
plenty of material for their
publication. They found a lot of data
that was very relevant.
In that process they also added
annotations to the files using a piece
of software called Elan. And you can see
a screenshot from it here. Elan is
software for making time aligned
annotations on media. And it allows you
to explore this kind of stuff very very
easily. You can click on bits of
transcription. You can click on chunks
in the timeline and you get immediately
to where you need to be. So this is
actually a greatly enriched version of
at least some of the data. And we are
now talking to the research team about
bringing those files back into our
collection because they're potentially
so useful to anybody else who wants to
work with this material.
And now I'm going to pass on to Mel.
>> Hi everyone. I'm Mel. Um I am a white
American Australian. Um and I work study
and joining you today from Gatagle
lands. Um and I'll be talking about the
quotation tool. So uh just to start with
the quotation tool uh what is it? It's a
Jupyter notebook. Um, it's available on
the Australian text analytics platform.
Um, and it's also now part of the Eldaca
Wordflow package. Um, the quotation tool
has been designed specifically to
automatically extract quoted uh content
from newspaper texts along with the
sources of the quotes and associated
speech verbs said, told, claimed, and so
on. Um, there is quite a lot of
documentation and instructions. So, if
you have no idea what a Jupyter notebook
is or haven't used one before, um if you
can follow instructions, you can do it.
Uh because that's certainly where I was
at when I started with this. Um so, the
tool allows you to upload your own data
or your own corpus um for automated
tagging and then you can see it within
the notebook um through some
visualization tools. So, there's some
examples on the right. Um or you can
export the results to Excel um for
further analysis and manipulation. Um
so, if we just look yeah, the screenshot
on the um top
is uh the visualization tool. It's taken
from my own uh my own corpus of
newspaper coverage of women's AFL and
NRL, so women's footy. Um and it
highlights there uh the the speech text
is identified in blue underlines. The
speaker is then in green. Um and it
tells us the type of entity for the
speaker as well. So whether it's a
person um or an organ or organization.
Um, in the bottom is an example of
another part of the visualization tool
which allows you to show the top um,
number of speaker entities in the
corpus. So I've just done the top 10 for
an example. Um, we can see that AFL and
NRL about halfway down is NRL are
speaker um, entity organizations that
are quite uh, common in the corpus and
the rest are all individual names and
their their player names. Um, so AFL and
NRL player names. So in terms of
applying this uh to our own research, it
can be used to explore lots of different
research questions. Um you can look at
who is the most or least cited um in
your data. You can look at different
reporting expressions that are used um
and the kind of information that is
attributed to the sources that are
quoted. Um and that's where I will talk
about um how I use the quotation to my
own research to explore that. So next
slide please.
So uh this shot the slide at the bottom
uh sorry the screenshot at the bottom of
the slide just shows um the data um
extracted in the table format from the
quotation tool. Um so this was part of
my research uh project for my PhD. Um,
and this component I wanted to explore
whether female athletes were represented
in the media as um, stereotypically more
emotional than their male um, athlete
counterparts. Um, and so I had quite a
large corpus um, of 5 million words of
print uh, news coverage of AFL and NRL,
both the men's and the women's. Um, and
the thing that hasn't been done so much
um, is looking at who actually is
expressing the emotions. So are the
emotions coming from players themselves
or are the emotions being um sort of
written in by the journalist and that's
where the quotation tool was really
really useful to sort of extract that
information. So I have a few uh a few
parts that I explored to this wider
question of of um emotionality in the
corpus. Um but I wanted to look at which
emotions were salient um and compare the
men's and the women's coverage and then
look at who are the sources of those
emotions. Um I also looked at the
triggers of emotion. But I won't really
talk about that today. Um, but if we
look at the screenshot on the bottom,
um, you can see when it exports the
data, you get your text name. So all of
my newspaper articles were separate um,
text files. You get the quoted content.
Um, it gives you the speaker name, the
type of speaker entity, whether it's a
person or an organization, and also the
quote type. Um, so the quotation tool
does use uh heristic rules as well. So
it can do uh directly quoted speech as
well as um sort of reported speech which
um again is really useful. Uh so next
slide.
So um I used the quotation tool in in
two primary ways. Um the first was to
extract all of the quoted content so
that I could create two subcorp. Um so a
corpus or a database of all quoted
speech from uh media coverage of women's
footy and then another subcorpus or or
database of um all the quoted speech in
in coverage of men's footy. And that
allowed me then to compare the types um
of emotions that were salient in one
corpus compared to another. I won't talk
too much about that, but I did use
another tool um that's on the ATAB
website, the keywords tool. And then um
I undertook qualitative analysis and I
wanted to look at uh specifically so we
can see on the slide of the quotes which
ones had emotion and if they had emotion
was the speaker a player. So I've added
uh in the columns that are in blue were
my data coding. So, okay, first line.
Yes, sad. The speaker, we can see it's
named. Um, I didn't include the speaker
entity on here, but it's a person, and I
had a list of player names that I
consulted against. Um, and marked if it
was a a speaker player, yes or no. Um,
so in some cases, the quotation tool,
you can see in uh the gray uh rows down
the bottom, couldn't extract the speaker
name, the blanks. Uh, so I I ignored
those. um and also anaphoric and
references pronouns um I I had to
exclude because it was just beyond the
scope of time uh for my project but
certainly you could look at that as
well. Um so you can see very briefly my
analysis showed yes that uh women's
players do account for more sources of
emotion um in quoted speech content than
men's players. Now, this could either
mean that the quoted participants, so
the female athletes, use more emotion
terms when they're talking about footy
or that the quotes that are selected or
the projected speech that's written in
um is is determined by the journalist
and they've just happened to select ones
that have more emotion words whether
whether they're conscious of it or not.
Um so that's it for me and I will pass
back now to Simon I think to wrap things
up.
Xmail.
So, very quickly to
try and pull things together a little
bit, um I think we see these fascinating
case studies that we're looking at
people interacting with data, different
kinds of people, different kinds of
data, various sorts of interactions. But
that's the kind of thing that's
happening and very big part of what I
think we do at Eldaka is to help that
process to help bring people and data
together so that interesting things can
happen.
And in the case studies we've seen
we've shown that researchers get
immediate benefits from these processes.
They're getting better access to data,
easier access to data. they're using
getting better access to tools that can
help them work with the data.
Uh but I think it's also important to
keep in mind that there are potentially
at least future benefits for other
researchers and other people from the
kind of work that we're doing. Work can
be reshared in the commons. It can
enrich the commons further.
We are trying to build community of
practice around this kind of work so
that there will be people who feel they
have ownership in what we're doing and
the kind of um process that I described
with the fully an more richly annotated
gesture data will become we hope more
and more commonplace that people will
feel that they are working with data
they're doing good things with it and
that should be reusable in the future as
well.
So we we want to have more and more
people collaborate with our commons.
Please come and use our facilities, talk
to us, uh explore the possibilities.
Thank you for listening.
That's amazing. Thank you so much um
Simon and the LDAC team. It's been so
wonderful to hear all of your fantastic,
beautiful case studies and to get a real
insight on what's happening uh in LDAC
and the and the amazing research that's
being done with the tools that are
available for people to use and you can
use them right now too. But before we um
uh wrap up, what I'd love to do is to
see if we've got any questions from the
audience. Uh so I can see one here from
uh from Michael there and he said thank
you and he's asked you um how are you
supporting the citability of data? Is
there anyone you'd like to take on that
question?
I can probably say something um
the any anything that's discoverable in
our data portal on the page for each
item there will be information about how
it can be cited. It's um fairly general
information. We didn't want to tie
ourselves to you know um Chicago or
Harvard or any particular uh style of
citation but the information is
available there and of course we
strongly encourage anybody who's you
retrieving information then to site it
appropriately.
>> Ah fantastic. Thanks Simon. Do we have
any other questions? I'm just looking um
in the chat there to see if we have any
more questions from anyone in the
audience.
All right. Well, I have a question if I
may. Um I was actually wondering if I
could ask to what extent does Elaca's
data portal support combined searches
involving both text and non-ext media.
So when we when we're dealing with text
and things that aren't text, how do we
look how do we look for that in the Elda
portal?
Well, I showed an example of fairly
of doing that. Um,
>> I think that the main way that we can do
it at the moment is by filtering on the
types of records you're looking for,
whether that's by audio, video, or by
file formats or something like that.
>> Um,
and then you can search for specific
text items after you've narrowed it
down.
Ideally, of course, it it would be
wonderful if we could search audio files
for things, but we can't quite do that.
>> Wow, that's wonderful. And it was great
to see uh your examples with the gesture
research, too. Um, which speaks to that
as well. So, that's that's great. Oh,
we've got another question from Angela
here. Angela would like to know, is
there a tool that can be used to analyze
focus group data? And also, are there
tools to redact data if consent is
withdrawn and one person in the focus
group's data needs to be redacted? Oh,
that's a that's a good question. Um,
does anyone want to try that one or is
that one something we should
Michael? Do you you say anything about
that? [clears throat]
>> Yeah, thanks for the girly one. Um,
[laughter]
look, that's a good question. Um, yeah,
there are tools. It depends what you
mean by analyze focus group data
actually. So, it's hard to answer your
question without knowing what you have
in mind with analyze. Um, but I think
obviously people do qualitative
discourse or content analysis. Um, and I
think there's an emerging project that
ARODC is working on. Um, if if it's your
gig, um, to be using ALMs for that kind
of thing. Um, but in LEALand, we're also
got tools if it's if it's a larger data
set. Uh, there's sort of certain text
analytics, keyword tools, corpus, uh,
type collocation things which could be
helpful if you're not familiar with
those tools. It's just another window
into the data. So yes, there are um for
redacting data. Um if I was really
thinking aloud, um I would say that
you'd probably use our tools to convert
whatever transcripts you have into a
kind of tabulated form. Um and then it
would be quite easy to redact that
particular speaker from the transcript.
Um so I work in Word a lot. um that's a
really um I was gonna say crappy um not
very useful form for doing that kind of
work, but if you are able to uh I guess
wrangle that data into some kind of
tabular spreadsheet, then it's really
quite easy to do that kind of thing. Um
so we do have tooling uh to work uh with
that kind of stuff.
>> Fantastic. Thank Thank you, Michael.
That's that's a good answer. Um, oh, and
and Angela has has said thank you as
well. And uh yes, we we do have a number
of tools on the Hassani uh RDC page as
well. So we can if you leave your uh
email address or email us, Angela, we
can we can tell you what other tools as
well may be available for you or what's
coming soon too. Uh we've got a question
from Lisa as well. Lisa says, "Thanks
for a great session, everyone." And a
question for Simon. Is there a way to
keep a log of the search terms that you
were using uh used from a session to
gain on a given data set? So, I think
Lisa might be referring to the search
that you did. Um, yeah, that's that's an
interesting question. Simon, would you
like to answer that one?
>> Um, I can answer very quickly and say
no, not at not at the moment. Um, I I
have experienced this kind of um
capability. I think Trove for example
allows you to u create kind of virtual
collections and searches you can store
in your profile. We don't have that at
this stage. I don't know if we have
plans either but I know it's it is a
valuable facility. I might tag on to the
back of what Simon just said and say yet
um we are triing a HTML light uh portal
variation which is quite interoperable
with the uh only portal itself and uh
actually our colleague Ben Foley has
been working on that and funny that
question comes up because we were just
having a run through it about two hours
ago and we did that precisely that to a
degree where there is a um and it's it's
quite iterative. It's being developed
and we working through different bugs
and stuff, but we were able to do a
level of um analytics uh at least
concordancing and and an engram one
which was um which was good and we could
do our various searching through that
and that did uh that did keep a log that
we could then um save as as a text file.
So that was particularly handy. So, I'll
probably just say yet on the back of
Simon's comment
>> and maybe I'll just add um the reason
we're separating it is because in the
portal if you keep a log of people's
searches um then you're you have to
attach it to the person so it becomes an
issue of privacy and where do you store
that information um so I think our
preferred solution is not to store it in
a central portal which would raise those
kind of privacy issues but try would
probably just get you to signing away
life away, right? Um but also because
we're not just getting in researchers um
who go through AF, there's other ways to
connect in um from community and so on.
Um so I think our preferred solution is
along the lines you're saying, Robert um
which is that individual researchers
communities kind of set up their bespoke
one and then they keep a record of it
and then their private searches are
their private business. uh the mindset.
>> Excellent. Thank you. Thank you all. So,
do we have any last questions from our
um from our audience? Looks like we've
still got a lot of people who are who
are hanging on, which is great. And
thank you for staying for the questions.
We do have we do have time for I think
one more if anyone has a burning
question they would they would like to
ask.
All right. Well, in that case, I may
just share some uh useful links. So,
I'll be just one moment while I
while I share this one. So, just a
moment. Hopefully, you can see that all.
Okay, there. So, yes. So, uh, thank you
so much to all of our wonderful
speakers. So, um, it's been really great
to see the that incredible research
which is happening in Eldaka. You can do
quantitative, qualitative search all at
the same time. Amazing tools. Um, and
these these are just some of the uses.
Uh, if you are a researcher or you're
supporting researchers, uh, this this is
this is the beginning this could be the
beginning of your new research journey
or your or your support of researchers.
So, please do um do share and tell
people about the stories that you've
heard today. And if you'd like to
contact LDAC direct directly, you can at
lacaq.edu.auu.
Please have a look at the LDAC data
portal. Uh and that's https
colonback/data.ldaka.edu.au/arch
and you can have a look there and and do
some great searching. Um and uh we also
have our our next uh HAS and Indigenous
Research Data Commons webinar in this
series is coming up. It'll be with the
um the Australian Internet Observatory
and it is called Let me just give you my
next slide.
It is Oh, hang on. This sorry before I
get to the Australian Internet
Observatory, pardon me. The next LDAC um
uh announcement that we have for
tomorrow is you can you can join and
learn more about text analytics without
code and there is a great guided tour of
point andclick text uh analysis and then
a hands-on with a new geni annotation
tool. So this looks very cool, very
exciting. Those are the times there. So
11 to 12:30 Eastern time and then 2:00
p.m. to 3:30 p.m. um Eastern time as
well. That's Word Flow from zero uh
which is a demo and then you can get
your hands dirty and get right in there
with your text which is which is
fantastic and that's free on Friday 28th
of August on Zoom. So uh follow that
link there uh sih.tools/wordflow.
Um, and if you're watching this on the
recording, um, hopefully uh hopefully
you'll be able to see uh a potential uh
recording of that for the demo in future
as well if you if you're seeing this a
few days after um this webinar. So um
yes, so next up we have the uh the next
webinar in our series will be
introducing the Australian Internet
Observatory platform for digital
platform research and that'll be on the
27th of September and uh you can
register via uh the link which uh
hopefully we can pop there into the chat
and we really hope to see you there and
if you need to contact the Australian
Research Data Commons uh please contact
us at contact ardc subscribe to our
newsletter so we can let you know about
more sessions and opportunities to hear
like fabulous uh from fabulous people
like like you have today. So, thank you
again so much to the Language Data
Commons of Australia. Um and and yes, so
yes, and you can also uh discover new
events um which we're going to put our
our link to subscribe to the newsletter
and also to see our upcoming events. So
you can you can come along and join in
uh and learn more about what's happening
with the Australian Research Data
Commons, ELDAC and our other focus
areas. All right,
>> Mary, there's just one more question in
the chat if we've got time.
>> I we we've got three minutes. Are you
are you willing um eldaka people to
answer our last question?
>> Yes. Now just a minute.
>> It's from uh Kristen. She said, "Can you
talk more about the permissions from
community data custodians around the use
of language, eg only certain uses or
preventing people from putting language
into AI if that's a concern?"
>> That's a that's a meaty question. Would
anyone like to take that one on the last
two minutes? We can't give a substantial
considered answer in in two minutes but
the very brief uh key points I would
raise is that um uh there is the data
port we are creating multiple
infrastructures so although there is the
central data portal um it will be of
each individual community's
determination as to whether or not they
want to uh start to where they want to
put their data sets uh whether or not
that's in a central portal or in uh a a
a community version which we are
piloting around. So So that's that's one
thing. Um there's a little bit of a a a
wicked problem in in in the uh
availability and and use agreement uh
challenge is that if if you if you put
it in the portal, if you're the owner of
the data and you put it in the portal
and you apply a license to that, which
is all very standard LDAC process, if
your um license if you if that license
enables people to access it or you
approve people for access of your data
and they can tick a disclosure saying I
will adhere to this data and I will do
all the things you've described in your
license if they actually physically um
access it that that is um they they may
do what they want with it. So that's not
just an LDAC problem that is a data
access uh global data access challenge.
Um and I I just wonder if it's worth
mentioning, Michael, we had some talks
around uh viewing the data, particularly
some of the sensitive uh interview data
around viewing it, but not actually take
removing it off the platform. So, not
actually being able to download it
locally.
>> Yeah. Um so I I think it's worth
pointing out the data we have out there
right now uh none of it would be
considered under the sensitive category
and it's made available under the
conditions under which the data
custodians made available but I think a
really key point to make is that um in
the in the world of repositories there's
really only two extremes at the moment
um so you either have the open
publishing model of standard libraries
and so on where they publish the data
set or you have the world of university
or other institutional repositories
which you can really only access if
you're a member of the institution. Um
so what's different about Aldacer is we
have an access control uh mechanism. So
the data custodian and Stuart are the
ones who determining who accesses the
data and what under what conditions and
of course um it could be around
indigenous language data but it can be
just video recordings of people which
are a bit sensitive or sensitive
interview data that kind of thing. Um
and so in that case you can have quite
tightly controlled uh access control
meaning it's only accessible to who I
the data steward say can look at that
material through to something uh more
open and between. So I think that's a
really important point to understand
about the infrastructure that it enables
access control which is really not a
standard thing uh and in the broader
ecosystem. Of course, there's
exceptions. Um, that's a generalization,
but generally speaking, that that's a
bit of a challenge uh for people working
with this kind of data where you don't
want it completely open, but you don't
necessarily want it completely closed
either.
So, um, respectful sharing, right, uh,
Robert, I think is is is often, uh, the
philosophy under things.
>> Amazing. Thank you so much and thank you
for answering the the very last question
at the very last minute. So, thanks
again and a big thank you to um Language
Data Commons of Australia and all our
fabulous speakers. Uh thank you Melissa,
thank you Gary, thank you Robert, thank
you Simon and thank you Michael as well.
It's been absolutely amazing to have you
uh present and let us know what's
happening with the language data commons
of Australia and to share those
fantastic brilliant case studies. So
hopefully many many researchers will be
inspired and we'll see a lot of
brilliant new research very soon. So
thank you everyone and um and that will
be us that that that will be that and us
for today. So thank you and goodbye
everyone.