Submind YouTube summaries
Thumbnail for The Collaborative Power of a Research Data Commons: LDaCA case studies

The Collaborative Power of a Research Data Commons: LDaCA case studies

Watch on YouTube

Video summary

The first webinar of the House and Indigenous Research Data Commons series focused on the Language Data Commons of Australia (LDaCA), an initiative dedicated to securing nationally significant language data under community control. The session addressed critical challenges such as disorganized metadata, inaccessible collections, and a lack of user-friendly tools by presenting four distinct case studies that demonstrated LDaCA's capabilities. These examples included enriching metadata for the Carolyn Tennant Kelly papers, which revealed previously hidden language groups; utilizing the Ningan platform to transcribe and tag fragile field notes on the Gungloo language; employing advanced search filters to locate specific multimodal data like gestures alongside speech; and using automated Jupyter notebooks to analyze media coverage of female athletes in Australian sports. Beyond these technical demonstrations, the discussion highlighted how LDaCA fosters a collaborative ecosystem that democratizes access to research tools and data for both quantitative and qualitative studies. The platform offers a unique middle ground between fully open repositories and restricted institutional access by implementing granular control mechanisms where community data custodians determine who can access specific materials and under what conditions. This approach allows for "respectful sharing," enabling sensitive content, such as interview recordings, to be viewed without being downloaded locally, thereby addressing global challenges related to data licensing and usage agreements while maintaining cultural integrity. The session concluded with a Q&A that covered essential topics including data citation practices, combined text and media searches, and the management of consent through data redaction features. A significant portion of the dialogue addressed community concerns regarding permissions for language use and AI training, clarifying that while a central portal exists, communities retain the option to create separate versions or pilot programs to preserve their autonomy. The moderator emphasized that this infrastructure supports a balanced model where access is neither completely open nor closed, ensuring that sensitive materials can be utilized responsibly without compromising community trust or privacy. To support ongoing engagement and further exploration of these tools, the organizers shared key resources including the LDaCA data portal and contact information for the initiative. Attendees were informed about upcoming events designed to deepen their understanding of the platform's capabilities, such as a hands-on workshop on text analytics using Word Flow scheduled for late August and a webinar introducing the Australian Internet Observatory platform in late September. Ultimately, the webinar reinforced LDaCA's role as a vital infrastructure that empowers communities to manage their own data sovereignty while facilitating broader research collaboration across Australia.
Read the full video transcript
Well, hello everybody and welcome. Uh, thank you so much for joining us today. Uh, my name is Mary Philsell. I'm an engagement representative uh, for the HS and indigenous research data commons and I'm also your host for this webinar series as well. So today's session the collaborative power of a research data commons um in is the very first uh webinar that we have in our house and indigenous research data commons uh webinar series and before we get started um I'd like to acknowledge all of the traditional custodians of the country and waters that we meet on today and I'd like to pay my respects to elders past and present and I also like to extend this respect to all indigenous people joining us here today and uh acknowledge the deep and ongoing connection to countries shared across indigenous Australia. So I myself I'm a white settler person and I work and live on Ghana land and uh the local greeting in uh Ghana country is Nina Mani. So Nina Mani everybody uh it's an honor to be joining you here online today. And uh before we before we meet our speakers, I've got just a couple of little housekeeping notes to share with you. So first of all, this session is being recorded. It's currently being recorded and this recording will later be uh shared with registrants and and uploaded to the uh ARDC YouTube channel after that. So, if you have any questions along the way, please feel very welcome to post them into the Zoom chat and we'll be able to come back and answer some of them um or all of them uh together at the end. Just pining on time. So, if we run over time, we may not be able to get to your particular question and I've just realized I just need to apologies I have to progress my slides and that's just a note about the recording. So this is our code of conduct for today and the um the link if you if you would like should come up in the chat the link is uh e- research ardcc and we would like to keep this a welcoming and respectful space for everyone and we'd like to ask all attendees uh to follow the ARDC code of conduct which is available at this link. Okay. So yes, the house and indigenous research data commons webinar series um is one of our newest offerings from the house and indigenous research data commons and this session is part of the series and this is a monthly series for H researchers featuring case studies tool demonstrations and hands-on skill training as well. So this series is led by the ARDC and funded by ENIRS and ENIRS stands for the National Collaborative Research Infrastructure Strategy. And so the Hassan Indigenous Research Data Commons supports six focus areas. And today's session focuses on one of those focus areas which is the language data commons of Australia or as we often call them ELDAC by the acronym. So today's session is the collaborative power of a research data comments Eldhaka case studies and joining me today are Michael Moore from Eldaka uh Robert Mlelen from Eldaka, Simon Musgrave also from Eldaka, Gary Tuda Smith from the University of Queensland and Melissa Kembell from the University of Sydney And a little note about what uh how how today is going to help you do better research uh through this webinar series. So you'll be able to access collaborative tools and platforms, use secure nationally significant data collections and analyze data without needing custom coding or setup. So now I'm going to hand over to Michael Hall who unpack exactly what a research data commons is and what ELDA can do for you and your research. So, Michael, over to you. >> Thanks very much, Mary. Um, so it's my great pleasure to be here today. I think um Simon, you're just working on the slides in the meantime, but um while you're putting them up, I just wanted to acknowledge the lands of the younger and durable peoples u from which I'm uh beaming in here today. Um I always feel uh acknowledging our country is such a special thing to be doing in this country. We have so many uh languages stretching back so many tens of thousands of years. It's it's really stunning to celebrate but also to acknowledge um where we are today with many of those languages too. Um so uh Aldaka what are we um so our aim really is to be working with nationally significant language data in Australia and its region. Um we're working to make that data accessible in appropriate ways um with community control. And when we talk about language data, we're thinking about language data both in the Australian continent but also in in the region uh in which we're in. And the reason we're focusing on that is because about a quarter of the world's languages are in Australia and its region region. Um Australia itself is the fifth most linguistically diverse nation on earth. Um and of course we're home to uh the world's uh oldest continuing cultures as well. So it's a really really important part of of of the international scene the the global society as well. Um and our aims are to be I guess securing uh really important collections of data. Um but also um I think today we'll be have a big emphasis on using uh the language data and how we put it to use. Uh next slide Simon. Um so when we think about language data um even though it's so important even though a quarter of the world's languages are in our region and and Australia has considerable responsibility uh to be doing something about it. Um language data is very often not well organized. It's not organized in ways that make it usable, reusable and repurposable. Um there's a lot of at risk uh collections. We hear about this all the time. It's hard to find things. Um accessing them is complex. uh uh sometimes it's for very good reasons and sometimes it's for less good reasons. Um and then the tools by which we analyze and and work with data in researchers and communities varies and often it's it's difficult to work in ways that are uh transparent and reproducible and so on. And finally um guidance and support for working with language data are very often scattered and hard to find. Um Simon next. Um so in dealing with that aldaka we started uh back in 222 we've been working for a while um we divide our work into variety of streams uh which is uh quite fundamental is a ecosystem of language data repositories uh they're in the plural I might add um working on securing nationally significant language data collections increasing accessibil access accessibility and enriching that data and also working on tools but true, easy to use because there's lots of things you can do with computers these days. Um, but they're not if you're not a computer scientist or computer wiz um you can't you can't necessarily do those things. So, LDAC is about democratizing access to those really important tools which leads us into that fifth stream which is around engagement and training. So, next slide. Thanks. Um here I just wanted to give you a sense of the kind of complexity um but also the simplicity of the Aldaca technical infrastructure ecosystem. Um so when I say complex um you can see you probably can't read it that well unless you've got it really enlarged on your computer screen. Um but there's a whole bunch of different uh organizations and collections and portals that we work across. Right. So in that sense it's complex um because our aim you know any research data commons is really a collaborative ecosystem which enables researchers communities and organizations to work together right so rather than working in silos we're working collaborative collaboratively across those boundaries and and you know doing things that we couldn't do um if we didn't cross those boundaries. Um so in that sense there's a complexity but there's also simplicity behind what we do because we want it to be maximally promoting autonomy for researchers and communities and we want to be promoting sustainability. Um so technical systems which are very difficult to work with uh very kind of heavy uh and don't promote autonomy or sustainability right um and so today's not the day to get into the technical architecture but if you know arrocate is the word uh to remember if you want to look it up and find out more um but today our aim really is thinking about how we put this uh I guess the infrastructure in into practice how we implement it, how people are are using it and and putting it to use. Um, and a big part of that is our I guess our careful fairness approach. Um, care and fair of course the important principles and the ones that we want to draw out today I think particularly are the R and fair. Right? So that's around reusing repurposing language data and also the C and care. So that there's benefit from doing that, not just for individual researchers, uh, but for the communities that they're working with as well. Um, so we've got four fabulous case studies and I'm really grateful for the four presenters today who are going to take us through those. Um, and I'll hand over to Robert first of all to to lead us. >> Thank you, Michael. And one young everyone welcome everyone. I I'm a um it's good to see you all. Well, I'm a Goring Goring person from Bundberg. Uh, but I'm also the program manager for the language data commons. Uh, and echoing those sentiments from Michael, this is the first webinar uh, series for the AODC. So, we're very privileged and excited to be the first cab off the rank so to speak. Uh, so I will kick off. Thanks, Simon. um guided by the HAS and indigenous uh uh research data commons framework for the governance of indigenous data. One key tenant that's within that framework is around uh providing knowledge, accessibility and usability of data assets. And what that means is it's around ensuring that indigenous communities are able to access uh and benefit from relevant data holdings and that's also with a emphasis there on inclusivity uh transparency but also informed consent. So I'm going to take a little bit of a deep dive. Michael talked about some of the language data challenges. I'm going to take a little bit of a deep dive into metadata. Um it's got a good story. It's entertaining, don't worry. Um, so when working towards making collections more findable, uh, the the accuracy of the metadata descriptions is something that is really quite vital. It becomes quite vital. Historical indigenous language and cultural materials, as as it's been said, they're often poorly described, if they're even described at all. And this in itself presents um findability challenges for anyone really wanting to find a record of uh data um you know finding the record of the data but then finding its existence in the very first instance. So that that remains a challenge that metadata is the key to solving. So making the metadata records available at the very least uh is a critical step towards demonstrating findability. So when complemented with good metadata descriptions, metadata availability presents an interplay or an insightful interplay between what you were looking for but also discovering what you never knew existed. And that really rings true from an indigenous community perspective being um that we have been historically and institutionally marginalized from accessing data about our people. So whether the metadata record is well described or poorly described, making the record available supports three things. First one being findability. The second being awareness that the data exists in the first instance regardless of access conditions and this in turn can support the longevity and safekeeping of the data thereafter. Uh the third thing of course is um it increases opportunities for future metadata enrichment opportunities which I'll finish on uh after this this session. So for today's presentation, I'm going to extrapolate this across two parts. We're going to look at metadata availability and metadata enrichment. And this is what really underpins our experience with the Carolyn Tenant Cali materials. Uh thanks Simon. So in 2023 there was a partnership between Eldacer and the University of Queensland library. This involved a pilot study which aimed uh we were aiming to assess the quality and the extent of indigenous language data that was held within university collections. Uh but we also wanted to demonstrate how such data could be meaningfully incorporated into the LDAC portal. So with the pilot, we adopted a harvesting approach uh quite similar to the one that's used by Trove uh with the National Library of Australia, the Trove Service there. And this was also supported with additional annotations that were generated through a prior project that had described this particular material, the Caroline Ten and Kelly uh collection that prior to that was quite poorly described. So I'll give you a bit of a run through of what it actually is. So the Carolyn Tennant Kelly or the CTK papers as we we've referred to them their fieldwork accounts of Australian Aboriginal culture in the 1930s. So Carolyn Tenant Kelly was an anthropologist. Um but these papers contained a significant amount of language data that was in uh in Queensland at that time. So prior to their acquisition, well they ended up being lost. um they were in in a back shed for quite a long time before they were a tip off was given to the university and a group had uh set a project to go and work with this material. So prior to the university acquiring this material um they were in many ways I guess considered orphaned data in that they were lacking that institutional stewardship and and a and a structured description which we now have. Um so the metadata enrichment in the first instance was carried out uh in around 2010 uh 2011 and this was done through the creation of an index uh an index by categories and that was done by David Tger uh Kim D Reich Tony Jeff and Uncle Michael Williams who um is um also a senior adviser for LDAC. Um so this index revealed 413 items in total. uh and that included 121 identified language groups. Um there were uh sorry language entries. There were 15 keyword entries. There were 269 place names recognized and there were 268 personal names identified. So this uh quite quickly became a very rich collection. But this was not what was represented in uh the university library catalog. So that was the problematic part. Uh for instance, I said the index referenced 121 distinct langu indigenous languages but only 22 were reflected in the catalog. There's a problem. That means 99 languages were not reflected on that catalog. Now, I only knew this because I'd acquired the data in a back background way. So, I knew of the index that existed and I knew my language was on there and I could see very clearly that my language was not one of the um 22 that were listed on the on the university catalog. So, we started working with this material um to to bring that to to surface that uh that information uh to make it more findable for other communities. The problem of course is that 99 languages not being findable in the catalog is um locks 99 language groups out of accessing and work finding accessing and working with that uh that um material. So you can see that's that's quite a bit of a struggle. So the index also enabled a bit a more nuance navigation across the collection um through the different additional descriptors as I said place names, key uh keywords, names of language speakers. So things that we can start to pick up and work with and repurpose or communities uh wanting to steward this data in their own way can start to work with um some of that enriched material and they're not starting uh from scratch. This is not to say conclusively that the enrichment work that has been done on the collection is all there is. You know, there's always room for improvement. But it uh as a first step this uh the remarkable part of this project was surfacing that and making that findable for communities. And you can see in the screenshot uh there you'll see different uh languages now referenced within the LDAC test portal. And uh that QR code will take you to just a short a short maybe two-page uh brief on the uh on the LDAC UKQ library collaboration. So that's the CTK case example and I am now going to hand over to my colleague Gar. Thanks Robert. Good day everyone. I'm Gary Tudtor Smith. I'm a um gangaloo bar and growing growing person from central Queensland. Um, I'm a PhD candidate at the University of Queensland and I also work part-time at the research unit for indigenous language at the University of Melbourne. Thanks, Simon. So, a little bit of background. Um, so about seven or eight years ago, um, my cousin Thomas Watson was working with, um, linguist Andrew Tanner at, uh, living languages. So, he was looking into reawakening Gongaloo language, learning it for himself, and sharing it with his family. So together they tracked down um a bunch of different documentation and descriptions and one of these was this linguistic survey of southeastern Queensland by a Swedish linguist called Nils Homemer. So this has got um a bunch of different languages. You can see the list on the right there. They've got some short descriptions and word lists to go along with it. So it's useful to an extent but there's not a lot of um examples of sentences which you know community can use to learn the language and to develop teaching and learning programs. Thanks Simon. So um in the 60s and 70s um Nilsmer came to New South Wales and Queensland um with the establishment of IATIS and um the publication I showed on the previous page uh was from 1983 and so there were a couple of others but this is the one that contains um Gunglo language and so it wasn't until 2022 when um Homer's field notes were rediscovered so Tom reached out to Neil's son Arthur Homer who's a linguist in his own right in Sweden and um got in touch and Arthur wasn't aware that there were any field notes but he eventually tracked them down and scanned and emailed them through and so across all of the languages there was 26 notebooks and 2660 pages. So hopefully the image might show up if you click again Simon. Yeah, this is a um example [clears throat] of the field notes. So you can see each line has the language and then the English and each line is tagged for the speaker. The next slide. So um I think it was 22 or 23 myself um Tom Watson and Paul Williams um with support from LDAC um use this Ningham platform to transcribe these field notes. So, as you know, many people might know who've worked with um archival manuscripts, it's a bit of a challenge to get them off the handwritten paper into some sort of database, you know, being able to read them well and understand what's there and making them searchable. So, Ningon is designed to, you know, help this problem. And one of the nice things you can do with this as well is you can put some additional tags and um information in there. So you can see on line three there's Kumo Bullaru. I've tagged for a place name. Um in the fir in the second line um I've also tagged the speaker name and I've added the speaker name in there cuz it's just got his initials there. So it looks a little bit daunting with the code, but all you have to do is highlight it and press these blue buttons and just inserts the code for you. The next slide. So this is what it ends up looking like. And you can see the highlighted part is what I've added um the the supplied information. Next slide. And so this is what it ends up looking like once it's been pushed to the repository. So, this is a different manuscript, but this is an interesting example um because all of the translations of the language language here um can't remember off the top of my head which language it is, but it's somewhere over in um Western Australia at a mission called New Norsia. All of the translations are written in Italian. So, someone has gone through translated them and put them as a supplied information in the highlighted yellow there. So, there's some nice things you can do to add that extra context there and also just make it easier to discover once you've tagged those people people's names and place names there. Next slide. So, some of the benefits of Ningan. Um, one thing is preservation. If we've got people handling original materials all the time, they eventually degrade some of these manuscripts that people are working with. These these aren't these are only a couple of decades old, but there's other ones that are over a hundred years old. the papers already crumbling each time someone handles them. We're putting acid and oil from our hands on them. Um, we're also saving time and money by not having to travel to Canra or, you know, wherever a manuscript might be kept to go look at them. Um, and we can also collaborate remotely with other people. Part of the Ningan platform is um the permissions are determined by um by the community. So that's part of the workflow when things get pushed to the repository and they can also always be changed. Um so there's that kind of iterative ongoing consent as well. Um as I mentioned the legibility is really important when you're working with um language materials. You want to make sure you're analyzing what's written what's actually written there. Um, and again the shared shared workflow, um, it just makes it easier to find stuff. And I think a really important thing is that, um, each if each person, each individual has to access material, transcribe, you know, or digitize, photograph, and then transcribe and then analyze, we can't get to the more important stuff. So we can get the you know this foundational work out of the way and then share it amongst our other community members and you know if we're working with other academics linguists as well. So next slide. So just to cap off some of the things I've been using these homework field notes for is um created a learner guide. If you click again an image should pop up. Yep. So that's a sample there and um you can see each of the examples in the orange boxes um are sentences from homeless field notes. Um so it kind of gives that veracity as well and you know if there's any um sort of speaker variation and people can choose oh well I'm an Aubry so I'm going to follow what Tim Aubry says for example. Um, we're also compiling a dictionary and currently develop developing a teaching and learning program and I'm in my second year now of my PhD um doing some further analysis into these materials and looking at where there are gaps in the documentation and how to fill those gaps. So, I'll pass on now to I think Simon's up next. Thanks, Thanks Gary. The I'm joining today from the unseated lands of the Bunarong peoples of the Cooland nation and I pay my respects to their elders past and present. case study I'm going to go through now is an example of how researchers can find quite specific kinds of data using some of the facilities that we're developing. The example I'm using here is some work that's done by former colleagues of mine at me university uh and one of them Harriet Shepard worked for Eldaka for a time as well. These scholars are interested in the question of how gesture reinforces or co-occurs with language. In particular, they're interested in a group of words which they call caused accompanied motion verbs. Now, these are words which denote someone moving but some other entity moving along with them. So, it's accompanied motion. And the prototypical verbs of this type of bring and take and carry. If you think about them, these words where someone goes somewhere but or something goes somewhere but something else goes along as well. So these scholars wanted to investigate they're investigating these uh words and the gestures that go with them cross linguistically and they were keen to have data from Australian English. Now the question here is is the data with video because of course you need that to see the uh the gestures and also transcription because otherwise you're going to have to look at all the video to find the examples. So now I'm going to show you screen capture of how you might go about finding this kind of data using our discovery portal. It's happened There we go. Okay. So, this is the where you enter a portal. Here we want to find out if there's a video. So, we look for the record type and filter on that. And you can see that there's 168 video items in the collection. And there's also 168 of them in the braided channels collection. So that tells us that all video material is in that particular collection. So now we want to filter on that collection just to check that there are transcripts. And you can see says there's recordings and transcripts. So we know that we're heading in the right direction. And now we can search for the particular words we're interested in this data. I'm using advanced search here because I'm going to search for multiple items. and also because that means I can use wildcard characters in the search. So that first search term that's going to look for carry and carries and carrying and so forth. Then I'm going to link all these searches with the or operator because we want full we want to meet one condition not all of them. Um as you can see another wild card there that we get take takes taking but took we have to enter separately. You could use regular expressions here, but it syntax gets a little bit complicated. I wasn't feeling confident about doing that. But we're specifying all these search items or search conditions and then we can have a look and we find there's 150 examples potentially that may be useful in this research. We can look down and see yes, we're picking up the various forms of these verbs. We're getting some false positives, but that's not too surprising. We can deal with that. We're getting some more kind of metaphorical uses of the verbs, but that's okay. So, let's have a look at the first example here. See what's going on. We can look at the whole transcript. Just check what's going on. And we search in here for took, which was first instance. And there we got this thing which actually has took and brought close together. So, this looks really interesting. This could be great data. We can see that there there's at least some time codes in these transcripts. So, we're within a kind of three minute span that we're going to be looking for that one, which is better than looking through half an hour, right? So, now we've established this looks interesting. We can download the data then see what it looks like. So, let's move on. Sorry. Okay. So, I'll just play you the video corresponding to the example that we found. Now, as you can see, unfortunately, this is not a useful example for us because for understandable reasons, the filmmaker zoomed in on that and you couldn't see the speaker's hands when they actually used the verbs make and bring. But that's research. As it happens in in fact though the research team went through this material and found all plenty of material for their publication. They found a lot of data that was very relevant. In that process they also added annotations to the files using a piece of software called Elan. And you can see a screenshot from it here. Elan is software for making time aligned annotations on media. And it allows you to explore this kind of stuff very very easily. You can click on bits of transcription. You can click on chunks in the timeline and you get immediately to where you need to be. So this is actually a greatly enriched version of at least some of the data. And we are now talking to the research team about bringing those files back into our collection because they're potentially so useful to anybody else who wants to work with this material. And now I'm going to pass on to Mel. >> Hi everyone. I'm Mel. Um I am a white American Australian. Um and I work study and joining you today from Gatagle lands. Um and I'll be talking about the quotation tool. So uh just to start with the quotation tool uh what is it? It's a Jupyter notebook. Um, it's available on the Australian text analytics platform. Um, and it's also now part of the Eldaca Wordflow package. Um, the quotation tool has been designed specifically to automatically extract quoted uh content from newspaper texts along with the sources of the quotes and associated speech verbs said, told, claimed, and so on. Um, there is quite a lot of documentation and instructions. So, if you have no idea what a Jupyter notebook is or haven't used one before, um if you can follow instructions, you can do it. Uh because that's certainly where I was at when I started with this. Um so, the tool allows you to upload your own data or your own corpus um for automated tagging and then you can see it within the notebook um through some visualization tools. So, there's some examples on the right. Um or you can export the results to Excel um for further analysis and manipulation. Um so, if we just look yeah, the screenshot on the um top is uh the visualization tool. It's taken from my own uh my own corpus of newspaper coverage of women's AFL and NRL, so women's footy. Um and it highlights there uh the the speech text is identified in blue underlines. The speaker is then in green. Um and it tells us the type of entity for the speaker as well. So whether it's a person um or an organ or organization. Um, in the bottom is an example of another part of the visualization tool which allows you to show the top um, number of speaker entities in the corpus. So I've just done the top 10 for an example. Um, we can see that AFL and NRL about halfway down is NRL are speaker um, entity organizations that are quite uh, common in the corpus and the rest are all individual names and their their player names. Um, so AFL and NRL player names. So in terms of applying this uh to our own research, it can be used to explore lots of different research questions. Um you can look at who is the most or least cited um in your data. You can look at different reporting expressions that are used um and the kind of information that is attributed to the sources that are quoted. Um and that's where I will talk about um how I use the quotation to my own research to explore that. So next slide please. So uh this shot the slide at the bottom uh sorry the screenshot at the bottom of the slide just shows um the data um extracted in the table format from the quotation tool. Um so this was part of my research uh project for my PhD. Um, and this component I wanted to explore whether female athletes were represented in the media as um, stereotypically more emotional than their male um, athlete counterparts. Um, and so I had quite a large corpus um, of 5 million words of print uh, news coverage of AFL and NRL, both the men's and the women's. Um, and the thing that hasn't been done so much um, is looking at who actually is expressing the emotions. So are the emotions coming from players themselves or are the emotions being um sort of written in by the journalist and that's where the quotation tool was really really useful to sort of extract that information. So I have a few uh a few parts that I explored to this wider question of of um emotionality in the corpus. Um but I wanted to look at which emotions were salient um and compare the men's and the women's coverage and then look at who are the sources of those emotions. Um I also looked at the triggers of emotion. But I won't really talk about that today. Um, but if we look at the screenshot on the bottom, um, you can see when it exports the data, you get your text name. So all of my newspaper articles were separate um, text files. You get the quoted content. Um, it gives you the speaker name, the type of speaker entity, whether it's a person or an organization, and also the quote type. Um, so the quotation tool does use uh heristic rules as well. So it can do uh directly quoted speech as well as um sort of reported speech which um again is really useful. Uh so next slide. So um I used the quotation tool in in two primary ways. Um the first was to extract all of the quoted content so that I could create two subcorp. Um so a corpus or a database of all quoted speech from uh media coverage of women's footy and then another subcorpus or or database of um all the quoted speech in in coverage of men's footy. And that allowed me then to compare the types um of emotions that were salient in one corpus compared to another. I won't talk too much about that, but I did use another tool um that's on the ATAB website, the keywords tool. And then um I undertook qualitative analysis and I wanted to look at uh specifically so we can see on the slide of the quotes which ones had emotion and if they had emotion was the speaker a player. So I've added uh in the columns that are in blue were my data coding. So, okay, first line. Yes, sad. The speaker, we can see it's named. Um, I didn't include the speaker entity on here, but it's a person, and I had a list of player names that I consulted against. Um, and marked if it was a a speaker player, yes or no. Um, so in some cases, the quotation tool, you can see in uh the gray uh rows down the bottom, couldn't extract the speaker name, the blanks. Uh, so I I ignored those. um and also anaphoric and references pronouns um I I had to exclude because it was just beyond the scope of time uh for my project but certainly you could look at that as well. Um so you can see very briefly my analysis showed yes that uh women's players do account for more sources of emotion um in quoted speech content than men's players. Now, this could either mean that the quoted participants, so the female athletes, use more emotion terms when they're talking about footy or that the quotes that are selected or the projected speech that's written in um is is determined by the journalist and they've just happened to select ones that have more emotion words whether whether they're conscious of it or not. Um so that's it for me and I will pass back now to Simon I think to wrap things up. Xmail. So, very quickly to try and pull things together a little bit, um I think we see these fascinating case studies that we're looking at people interacting with data, different kinds of people, different kinds of data, various sorts of interactions. But that's the kind of thing that's happening and very big part of what I think we do at Eldaka is to help that process to help bring people and data together so that interesting things can happen. And in the case studies we've seen we've shown that researchers get immediate benefits from these processes. They're getting better access to data, easier access to data. they're using getting better access to tools that can help them work with the data. Uh but I think it's also important to keep in mind that there are potentially at least future benefits for other researchers and other people from the kind of work that we're doing. Work can be reshared in the commons. It can enrich the commons further. We are trying to build community of practice around this kind of work so that there will be people who feel they have ownership in what we're doing and the kind of um process that I described with the fully an more richly annotated gesture data will become we hope more and more commonplace that people will feel that they are working with data they're doing good things with it and that should be reusable in the future as well. So we we want to have more and more people collaborate with our commons. Please come and use our facilities, talk to us, uh explore the possibilities. Thank you for listening. That's amazing. Thank you so much um Simon and the LDAC team. It's been so wonderful to hear all of your fantastic, beautiful case studies and to get a real insight on what's happening uh in LDAC and the and the amazing research that's being done with the tools that are available for people to use and you can use them right now too. But before we um uh wrap up, what I'd love to do is to see if we've got any questions from the audience. Uh so I can see one here from uh from Michael there and he said thank you and he's asked you um how are you supporting the citability of data? Is there anyone you'd like to take on that question? I can probably say something um the any anything that's discoverable in our data portal on the page for each item there will be information about how it can be cited. It's um fairly general information. We didn't want to tie ourselves to you know um Chicago or Harvard or any particular uh style of citation but the information is available there and of course we strongly encourage anybody who's you retrieving information then to site it appropriately. >> Ah fantastic. Thanks Simon. Do we have any other questions? I'm just looking um in the chat there to see if we have any more questions from anyone in the audience. All right. Well, I have a question if I may. Um I was actually wondering if I could ask to what extent does Elaca's data portal support combined searches involving both text and non-ext media. So when we when we're dealing with text and things that aren't text, how do we look how do we look for that in the Elda portal? Well, I showed an example of fairly of doing that. Um, >> I think that the main way that we can do it at the moment is by filtering on the types of records you're looking for, whether that's by audio, video, or by file formats or something like that. >> Um, and then you can search for specific text items after you've narrowed it down. Ideally, of course, it it would be wonderful if we could search audio files for things, but we can't quite do that. >> Wow, that's wonderful. And it was great to see uh your examples with the gesture research, too. Um, which speaks to that as well. So, that's that's great. Oh, we've got another question from Angela here. Angela would like to know, is there a tool that can be used to analyze focus group data? And also, are there tools to redact data if consent is withdrawn and one person in the focus group's data needs to be redacted? Oh, that's a that's a good question. Um, does anyone want to try that one or is that one something we should Michael? Do you you say anything about that? [clears throat] >> Yeah, thanks for the girly one. Um, [laughter] look, that's a good question. Um, yeah, there are tools. It depends what you mean by analyze focus group data actually. So, it's hard to answer your question without knowing what you have in mind with analyze. Um, but I think obviously people do qualitative discourse or content analysis. Um, and I think there's an emerging project that ARODC is working on. Um, if if it's your gig, um, to be using ALMs for that kind of thing. Um, but in LEALand, we're also got tools if it's if it's a larger data set. Uh, there's sort of certain text analytics, keyword tools, corpus, uh, type collocation things which could be helpful if you're not familiar with those tools. It's just another window into the data. So yes, there are um for redacting data. Um if I was really thinking aloud, um I would say that you'd probably use our tools to convert whatever transcripts you have into a kind of tabulated form. Um and then it would be quite easy to redact that particular speaker from the transcript. Um so I work in Word a lot. um that's a really um I was gonna say crappy um not very useful form for doing that kind of work, but if you are able to uh I guess wrangle that data into some kind of tabular spreadsheet, then it's really quite easy to do that kind of thing. Um so we do have tooling uh to work uh with that kind of stuff. >> Fantastic. Thank Thank you, Michael. That's that's a good answer. Um, oh, and and Angela has has said thank you as well. And uh yes, we we do have a number of tools on the Hassani uh RDC page as well. So we can if you leave your uh email address or email us, Angela, we can we can tell you what other tools as well may be available for you or what's coming soon too. Uh we've got a question from Lisa as well. Lisa says, "Thanks for a great session, everyone." And a question for Simon. Is there a way to keep a log of the search terms that you were using uh used from a session to gain on a given data set? So, I think Lisa might be referring to the search that you did. Um, yeah, that's that's an interesting question. Simon, would you like to answer that one? >> Um, I can answer very quickly and say no, not at not at the moment. Um, I I have experienced this kind of um capability. I think Trove for example allows you to u create kind of virtual collections and searches you can store in your profile. We don't have that at this stage. I don't know if we have plans either but I know it's it is a valuable facility. I might tag on to the back of what Simon just said and say yet um we are triing a HTML light uh portal variation which is quite interoperable with the uh only portal itself and uh actually our colleague Ben Foley has been working on that and funny that question comes up because we were just having a run through it about two hours ago and we did that precisely that to a degree where there is a um and it's it's quite iterative. It's being developed and we working through different bugs and stuff, but we were able to do a level of um analytics uh at least concordancing and and an engram one which was um which was good and we could do our various searching through that and that did uh that did keep a log that we could then um save as as a text file. So that was particularly handy. So, I'll probably just say yet on the back of Simon's comment >> and maybe I'll just add um the reason we're separating it is because in the portal if you keep a log of people's searches um then you're you have to attach it to the person so it becomes an issue of privacy and where do you store that information um so I think our preferred solution is not to store it in a central portal which would raise those kind of privacy issues but try would probably just get you to signing away life away, right? Um but also because we're not just getting in researchers um who go through AF, there's other ways to connect in um from community and so on. Um so I think our preferred solution is along the lines you're saying, Robert um which is that individual researchers communities kind of set up their bespoke one and then they keep a record of it and then their private searches are their private business. uh the mindset. >> Excellent. Thank you. Thank you all. So, do we have any last questions from our um from our audience? Looks like we've still got a lot of people who are who are hanging on, which is great. And thank you for staying for the questions. We do have we do have time for I think one more if anyone has a burning question they would they would like to ask. All right. Well, in that case, I may just share some uh useful links. So, I'll be just one moment while I while I share this one. So, just a moment. Hopefully, you can see that all. Okay, there. So, yes. So, uh, thank you so much to all of our wonderful speakers. So, um, it's been really great to see the that incredible research which is happening in Eldaka. You can do quantitative, qualitative search all at the same time. Amazing tools. Um, and these these are just some of the uses. Uh, if you are a researcher or you're supporting researchers, uh, this this is this is the beginning this could be the beginning of your new research journey or your or your support of researchers. So, please do um do share and tell people about the stories that you've heard today. And if you'd like to contact LDAC direct directly, you can at lacaq.edu.auu. Please have a look at the LDAC data portal. Uh and that's https colonback/data.ldaka.edu.au/arch and you can have a look there and and do some great searching. Um and uh we also have our our next uh HAS and Indigenous Research Data Commons webinar in this series is coming up. It'll be with the um the Australian Internet Observatory and it is called Let me just give you my next slide. It is Oh, hang on. This sorry before I get to the Australian Internet Observatory, pardon me. The next LDAC um uh announcement that we have for tomorrow is you can you can join and learn more about text analytics without code and there is a great guided tour of point andclick text uh analysis and then a hands-on with a new geni annotation tool. So this looks very cool, very exciting. Those are the times there. So 11 to 12:30 Eastern time and then 2:00 p.m. to 3:30 p.m. um Eastern time as well. That's Word Flow from zero uh which is a demo and then you can get your hands dirty and get right in there with your text which is which is fantastic and that's free on Friday 28th of August on Zoom. So uh follow that link there uh sih.tools/wordflow. Um, and if you're watching this on the recording, um, hopefully uh hopefully you'll be able to see uh a potential uh recording of that for the demo in future as well if you if you're seeing this a few days after um this webinar. So um yes, so next up we have the uh the next webinar in our series will be introducing the Australian Internet Observatory platform for digital platform research and that'll be on the 27th of September and uh you can register via uh the link which uh hopefully we can pop there into the chat and we really hope to see you there and if you need to contact the Australian Research Data Commons uh please contact us at contact ardc subscribe to our newsletter so we can let you know about more sessions and opportunities to hear like fabulous uh from fabulous people like like you have today. So, thank you again so much to the Language Data Commons of Australia. Um and and yes, so yes, and you can also uh discover new events um which we're going to put our our link to subscribe to the newsletter and also to see our upcoming events. So you can you can come along and join in uh and learn more about what's happening with the Australian Research Data Commons, ELDAC and our other focus areas. All right, >> Mary, there's just one more question in the chat if we've got time. >> I we we've got three minutes. Are you are you willing um eldaka people to answer our last question? >> Yes. Now just a minute. >> It's from uh Kristen. She said, "Can you talk more about the permissions from community data custodians around the use of language, eg only certain uses or preventing people from putting language into AI if that's a concern?" >> That's a that's a meaty question. Would anyone like to take that one on the last two minutes? We can't give a substantial considered answer in in two minutes but the very brief uh key points I would raise is that um uh there is the data port we are creating multiple infrastructures so although there is the central data portal um it will be of each individual community's determination as to whether or not they want to uh start to where they want to put their data sets uh whether or not that's in a central portal or in uh a a a community version which we are piloting around. So So that's that's one thing. Um there's a little bit of a a a wicked problem in in in the uh availability and and use agreement uh challenge is that if if you if you put it in the portal, if you're the owner of the data and you put it in the portal and you apply a license to that, which is all very standard LDAC process, if your um license if you if that license enables people to access it or you approve people for access of your data and they can tick a disclosure saying I will adhere to this data and I will do all the things you've described in your license if they actually physically um access it that that is um they they may do what they want with it. So that's not just an LDAC problem that is a data access uh global data access challenge. Um and I I just wonder if it's worth mentioning, Michael, we had some talks around uh viewing the data, particularly some of the sensitive uh interview data around viewing it, but not actually take removing it off the platform. So, not actually being able to download it locally. >> Yeah. Um so I I think it's worth pointing out the data we have out there right now uh none of it would be considered under the sensitive category and it's made available under the conditions under which the data custodians made available but I think a really key point to make is that um in the in the world of repositories there's really only two extremes at the moment um so you either have the open publishing model of standard libraries and so on where they publish the data set or you have the world of university or other institutional repositories which you can really only access if you're a member of the institution. Um so what's different about Aldacer is we have an access control uh mechanism. So the data custodian and Stuart are the ones who determining who accesses the data and what under what conditions and of course um it could be around indigenous language data but it can be just video recordings of people which are a bit sensitive or sensitive interview data that kind of thing. Um and so in that case you can have quite tightly controlled uh access control meaning it's only accessible to who I the data steward say can look at that material through to something uh more open and between. So I think that's a really important point to understand about the infrastructure that it enables access control which is really not a standard thing uh and in the broader ecosystem. Of course, there's exceptions. Um, that's a generalization, but generally speaking, that that's a bit of a challenge uh for people working with this kind of data where you don't want it completely open, but you don't necessarily want it completely closed either. So, um, respectful sharing, right, uh, Robert, I think is is is often, uh, the philosophy under things. >> Amazing. Thank you so much and thank you for answering the the very last question at the very last minute. So, thanks again and a big thank you to um Language Data Commons of Australia and all our fabulous speakers. Uh thank you Melissa, thank you Gary, thank you Robert, thank you Simon and thank you Michael as well. It's been absolutely amazing to have you uh present and let us know what's happening with the language data commons of Australia and to share those fantastic brilliant case studies. So hopefully many many researchers will be inspired and we'll see a lot of brilliant new research very soon. So thank you everyone and um and that will be us that that that will be that and us for today. So thank you and goodbye everyone.