Submind YouTube summaries
Thumbnail for Disappearing Data: Grassroots Efforts to Preserve Government Information in Uncertain Times

Disappearing Data: Grassroots Efforts to Preserve Government Information in Uncertain Times

Watch on YouTube

Video summary

The webinar series hosted by Kate Tolman brings together experts from institutions like the University of Minnesota, Harvard, and the University of Pennsylvania to address the critical issue of disappearing government information. Panelists Molly Blake, Jack Kushman, and Dr. Linda Kellum outlined various grassroots initiatives designed to track and preserve data that is being removed or neglected. Key efforts include a crowdsourcing project to document missing documents and broken links, such as those from the CDC and the Trump administration, alongside Harvard's creation of a massive 17 terabyte index of datasets using self-explanatory formats with cryptographic signatures for integrity. These strategies emphasize the need for resilient, redundant archives that do not rely on a single institution, advocating instead for the global distribution of copies to ensure long-term survival against agency closures or intentional censorship. The discussion highlights that data loss is caused by both systemic neglect and intentional removal, illustrated by examples like federal forms ceasing to collect demographic data and field offices becoming understaffed. Preserving the output of millions of federal employees requires dedicated core funding for community organizers rather than relying solely on volunteers, as the departure of staff who steward the data often leads to significant information gaps. To counter the tendency to rely on private vendors lacking mission-driven resources, advocates stress framing government data as a public good and lowering barriers to entry so that individuals at various skill levels can participate, from manual collection to joining technical communities. However, challenges persist, including API blocks and websites optimizing against crawlers, which sometimes force reliance on unverified commercial workarounds to access essential information. International collaboration is emerging as a vital component of this preservation movement, with partners in the UK, France, Germany, and elsewhere recognizing that US demographic data supports global initiatives like the UN Sustainable Development Goals for Africa. Despite the difficulty in quantifying the full scope of lost information due to tooling gaps and JavaScript-rendered content, the consensus is that avoiding territoriality and fostering international solidarity are essential for resilience. The panelists concluded by sharing plans to distribute resources and recordings while acknowledging the urgent need for infrastructure that goes beyond single points of failure like the Internet Archive, ensuring that vital government records remain accessible despite an uncertain digital landscape.
Read the full video transcript
Hi everybody. Going to give it a minute while attendees start to come in and participants are led into the webinar room. And in just a moment, we'll go ahead and get started. All right, it has slowed down just enough for me to feel comfortable to get started. So, um, welcome everyone and thank you so much for attending this webinar today. My name is Kate Tolman. I'm the uh head of library liaison services at Illinois State University and I'm also the I guess chair of the help I'm an accidental government information librarian webinar series. Um today's webinar is brought to you by uh the ALA government documents route table or goort in partnership with the ACRL politics policy international relations committee or pippers. We have a really amazing panel today as you will see um to talk about disappearing data and disappearing government information, a topic that we are all intimately familiar with. So I'm going to um let our colleagues from Pippers talk a little bit more about the impetus for this webinar. Um and uh very quickly uh you are welcome to chat uh questions any time. We will keep track of those and ask them at the end of the webinar. We are going to be giving each speaker a few minutes to talk about their projects. Then we're going to ask each of them individual questions and then we're going to ask them a group of questions that they can answer together in a discussion format and then we will open it up to you for questions as well. So that's going to be the general structure of today's talk. Um in the meantime, I'm going to hand it over to Julia Ezo. Uh Julia is the chair elect of God and uh the current chair of God's program committee. She is the government information and political science librarian at Michigan State University. So Julia, go right ahead. Thank you, Kate, and thank you everyone for attending today's incredibly timely webinar on grassroots efforts to address disappearing federal government information. We've gathered a wonderful panel of three experts who are in the trenches working hard to preserve critical government information and ensure continued free access for all. I'll now pass it over to my Pipper's colleagues, Jennifer Horn and Heather Gatsay. Jennifer, you are muted. You're muted. Okay, I'll start over. I'm sorry. I'm Jennifer Horn. I am the um business and government information librarian at the University of Kentucky and along with Heather Gatsay, I am the co-chair of the Pippers professional development committee. Uh back in April, we held um Pippers held a discussion um about the unstable federal data environment and it was very very popular and it was so popular that we had to restrict uh attendance um or else it was going to be out of control for a discussion group with no uh agenda and no speakers. So, we knew that we needed to work to put on a more formal webinar um with speakers and um we were very happy to be to hook up with Goor to be part of this webinar series um to put this on today. So, and now I will pass this along to Heather. Okay, thank you. Um hi everyone. I'm Heather Gatsay. I'm the uh social sciences and government information librarian at Slippery Rock University of Pennsylvania. Um, as Jennifer said, I'm co-chair along with her of the Pipper's professional development committee. Um, and we'd like to welcome our speakers today. We're so thankful to have their um, participation for this session. I'm going to provide a um, brief introduction um, for our speakers. I want to encourage them to share any additional information that they might want to do so um, regarding their backgrounds um, as they share their information today. Um, so joining us today we have Molly Blake who is social sciences librarian at the University of Minnesota Twin Cities where she serves as liaison to the economics, educational psychology and political science departments and also to the Hubard School of Journalism and Mass Communication. Um, she has additional professional experience in teaching, writing, research and nonprofit work. We also have joining us Jack Kushman who is director of the Harvard Library Innovation Lab um a software and design lab at the Harvard Law School Library. Um he is a software engineer and appellet attorney uh who previously worked as lead developer of the case law access project and has served as a lecturer on computer programming at Harvard Law School. And we also have Dr. Linda Kellum who is one of the founding organizers of the data rescue project and the Snyder Granador director of research data and digital scholarship at the University of Pennsylvania libraries. She has authored and edited works on data and librarianship and has presented extensively on data services, data management and fair guidelines. Um we want to thank all of them for joining us today and we're going to start things off with Molly Blake. Awesome. Thank you so much Heather. Um, I'm just going to share my screen quick. And is it letting me do that? Let me see. Yep. All right. Fantastic. Um, so I just wanted to start today by um sharing uh the team behind this very new project and thank you so much for this opportunity to chat a little bit about this project. Um, this project was really brainstormed by my colleague Jenny McBurnernney. And Jenny and I have been working closely over the past eight weeks to just really kind of get things off the ground. Um, and I'm very happy to share that just in the past few weeks, tracking government information has been joined by two other fantastic librarians, um, SA at the University of Illinois and Ben at Sacramento State. Um, so like probably all of you in this room, in February of 2025, Jenny and I started to get really concerned with what was happening with government information. And one event in particular sort of inspired us to take action. two of our colleagues over at the School of Public Health at the University of Minnesota um shared that a journal article they had written for a CDC operated journal had been removed from the CDC website. Um fortunately with the February 11th ruling that CDC data had to go back up, their journal article also went back up. Um but as these researchers pointed out in their article, this was not the only instance of government information being taken down. Um and one they pointed to was uh videos of the January 6th insurrection were removed um from I believe the Justice Department website around this time. Um, and then the other thing that happened, as I'm sure all of you know, in March, uh, the Trump administration started releasing all these memos banning a very long list of language to be included in public government publications moving forward. And so that really underscored to Jenny and myself that we needed to do something to track what was going on. Um, we knew that there were really excellent efforts going on such as the data data rescue project to rescue data. And so that made us wonder is anyone sort of tracking everything else, photographs, videos, websites, etc. Um, and so we decided to start a crowdsourcing project. Um, so right now our goal is to help researchers and just members of the general public see the scope of everything that um, the Trump administration has removed from government websites and to locate archived copies whenever available. Fortunately, quite a bit, if not everything, is available in the way back machine, but it's about connecting those dots. Um, and I'm sharing here, and I will share it in the chat when I'm done speaking, um, our current working website. And as of right now, there's a form where absolutely anyone is welcome to submit anything that they see is missing. Um it can just be a title of a document or a broken link. We're happy to investigate further on the back end. And then that form feeds into a spreadsheet so that um folks can kind of take a peek of what has been shared already and sort by things such as agency. Um I believe as of today we have about 100 entries. So, we're just really happy um to be able to just be uh continuing to collect this information. Um so, what has been removed other than data? For the sake of time, I'm not going to read this slide, but I think it gives you kind of a good sense of the breadth of what has gone away at this point. Um some of you may have heard it was just announced that all of uh Trump's speech transcripts have been removed from the remarks page of the government website, uh the white house.gov website. Um, we have working papers that have been removed. Tools that help researchers look at data or work with government documents have been removed. Uh, and maybe the most famous example um of an entire website going down. Two or 3 weeks ago, the co.gov website started rerouting to this interestingly designed website about the quote unquote true origins of CO 19. So that was once a website where you could get information about treatment options and vaccine safety. that has all been removed. Um so we are still um in the early days of our project and we are currently kind of expanding and growing. Um the next goal we have is we are in the process of developing a search interface to make it easier for users to explore um what people have submitted. And then we're also looking into options for expanding our project moving forward. Um, two possible routes we might go down is, um, we're interested in tracking things like office and agency closures that are going to impact the information ecosystem already have impacted it. So, I imagine quite a few of you today were maybe at the EBSS webinar this morning where we all heard from Aaron Polard Young about how the decimation of the education department has impacted ERIC. So that would be one of several examples of how these budget cuts are then impacting information we have the right to receive um as a p as members of the public. And then the second area where we're interested in expanding is this ongoing really terrible issue of language censorship. Um since the initial list of banned words came out in March, this hasn't gone away. Um, just this past week, the Guardian reported that State Department employees had received a memo that they are no longer allowed to use the term racially or ethnically motivated violent extremism. Um, which has very profound impacts about how white supremacist violence, the number one form of terrorism in our country, gets described and tracked on government websites and resources. Um, so I'm going to leave it at that and I look forward to chatting with you all further. Uh, thanks so much. Shall I pick it up from here? Let me share my slides. Uh, [Music] all right. So, I'm Jack. I'm direct the Library Innovation Lab, which is a group at the Harvard Law Library that explores the future of access to information. And we try to take the library principles that we all share and bring those to wherever people are getting their information today. Um we uh sort of believe as a core principle at the lab um one of the uh one of one of the core things that we're trying to solve is the preservation of cultural memory on the web that it is on thin ice across the board because uh in the same way that um information on the internet wants to be free. It also wants to exist in only one copy. We're no longer starting with you know 10,000 books in 10,000 libraries and winnowing down. We're starting from a a real economic pressure to go down to one copy that we all refer to. Um, and we're here to talk about public data. So, that's very true of things like data.gov, but it's also true of the YouTube videos where there's countless culturally valuable YouTube videos that if YouTube changed a policy, they'd go away. It's true of payw walls. Uh, there's countless uh uh articles that if they were changed, it'd be very hard to show that they had been changed. Um, we have a acrosstheboard cultural memory problem. um that we all need to invent strategies for solving. Uh so that's a big picture uh aspect of what we're thinking about as a lab. Um with public data, uh there's this real pressing need to preserve the stuff that we all care about because America runs on data. Um spending some time browsing data.gov um as we started to preserve it uh at the end of last year. Um there are endless data sets where you look and say, I sure hope someone is tracking that. I hope someone is tracking how much water is left in the aquifers, which crops are growing, which places are overcrowded, which places are a good place to build. Um the we desperately need a view of all of the things that we are doing as a society. Um and what we anticipated in November and what turned out to be true I think probably more than anyone anticipated is that uh the parties of the federal government in providing those things drastically changed. I list here um some but not all of the agencies that have had their missions drastically cretailed or changed uh just in the last few months uh that are agencies involved in public data in access to the uh the information that we all use to navigate um through our lives. Uh there are so many and this is this is not the full list. Um but there's a huge shift in what information is being offered and processed and preserved. Um, and so we faced this question, can we archive the government? Um, and of course not all of it, but we did dive into a little piece of it. So at the library innovation lab, we decided to take on archiving data.gov. Um, data.gov is a uh an index, a metadata map of about 300,000 data sets across the federal government. And we were told over and over, well, that's not all of them. There's a lot more to get. But this was like a place to start. Um, so we made a 17 terabyte index of um, those 300,000 data sets, just the direct files that were linked out from each entry on data.gov, which means in some cases we missed a lot because it would be um, just like a a homepage of uh, a bunch of other data that's somewhere else. But if it did link directly to the um, the data files, we would get that. Um and we really went into doing this work because uh we were talking to folks like the end of term archive who do wonderful work uh archiving um a full snapshot of the federal web before and after each transition. Um but really realizing that that process of clicking from link to link to preserve things would miss links rendered by JavaScript and very deep links and links where you have to email to get a data s or do an FTP um or links where you have to run a bunch of searches and aggregate them that anything that was specific data could be missed. So you'd end up with a manual to our public data but not the data itself. And that's what motivated this um this diving in to get the data. Um we thought really hard in doing this about um how do we make it resilient and redundant because um I don't want to count on you know Harvard still being around in order to have our archives still be around. I wanted to be resilient across repositories. Um so we made it um using standard uh self-explanatory formats using uh Library of Congress bags formats. Uh we added a metadata um signing standard. So um each one of our archives comes with a uh cryptographic signature showing that we made it and the time stamp that we made it. So it is very hard to fake. Um, and we're working now, it's not launched, on an interface that uh will let you do discovery of all this um in your browser without any serverside component, which means that um we've heard there's copies in the UK and Australia and so on, those copies will work just as well as our original. Um, you have the same provenence uh the same sort of uh integrity, accuracy of the thing and the same browsing tools. Um, if you make a copy as with the original um and that's really where a lot of our research is going. Um, we're working on um publishing the tools for that. So, uh, the core tool that we use for making this 300,000 archives is a tool I got to name called Bag Nabbit, uh, which makes bags that are signed. Uh, and this is, um, open- source, uh, uh, free for other people to use and adopt and, uh, remix to, um, make their own, uh, authenticated, uh, bags. Um, and what we learned in the process, and I hope I'm um, setting up Linda here, uh, you have to wear so many expert hats to do effective archiving of public data. um the the work that um and I wore a lot of these hats. I was like trying to pretend to be so many experts that I'm not. Um to do the mapping of what exists, the narrowing in on um what to prioritize once you identified a data set, the um the data librarianship to figure out well which versions of the file and which metadata files and which documentation constitute this archive. Um the programming uh skill to download it. um the uh uh sort of process skill to package it with the right metadata and the right signatures. Um the uh uh systemic skill to get it into the archives it needs to be in so that it will be preserved resiliently. Um and then the sort of access and discovery layer that is needed so that people can find the thing that they need. Um and then once you've done all of that, you have to figure out how to update that pipeline as new data comes out. Um it just became desperately apparent that we can't just go do this. We need a community to do it together. Um so Linda uh please tell us how that can possibly happen. Um and we ended up uh with a set of research questions that our lab is thinking about. One of which is where will that collaboration happen? How will um the library community and the technical community be able to come together to be all the kinds of experts that we have to be? Um how do we staff the core community organizing that has to happen around all the expertise that is um at the edges of this? Uh two, uh where are we going to do the resilient data preservation for this? Uh I don't know about everyone else. The world is too big for me to say this with confidence, but I know that in the part of it I can see, we are too much on um fragile archives on things that can be lost because um we can't take everything or because if we go away, it'll go away because if someone orders us to change our policy, it'll change. We need archives that are resilient to the kind of shock that we're going through now. And I think that means going back to the era of locks, lots of copies keep stuff safe and asking what locks of 2025 looks like, including the actual locks project at Stanford, of course, but also taking their concept of resilient digital storage and expanding it. Um, and finally to think about what knowledge producers should learn to not let this happen again. So going back to that first slide about YouTube is vulnerable and the New York Times is vulnerable and anyone who is a primary publisher of something they care about is at risk of this kind of loss. using this moment to say how can you structure your publication so that our work as archavists gets easier. So that's the stuff that we're thinking about and we'd love to continue the conversation. Great. Thank you so much Jack. I hadn't seen this presentation before. So that's really exciting to to hear that's your conversation right now. Um so I'm going to talk a little bit about the data rescue project u how it got started started and what we're working on as well. um and thinking about the future. Um so the data rescue project um started uh in um started coming together in January um of 2025 but really uh formally came together in February but started as a group of members of these organizations. These are primarily data professionals organizations. ISIS is international, ARDA app is primarily US Canada focus and data curation network is a consortium of universities with institutional repositories for data. Um and we were really concerned about the fact that our we had patrons who were looking for CDC data sets in particular or USA ID data data um and we wanted to help them find alternate sources um in that particular moment. So it it really started it was a crowdsourced way to try and and find sources um of places where data was backed up or sources for the of the alternate sources of data. Um and then it became um more of a data rescue focused um uh organization. Um we've been doing this unlike last time in 2017. Um and we do come out of the 2017 effort. um they they are our uh grandmothers and fathers I would call call them. Um we um unlike that period we are focused on asynchronous data rescues. So we had to come up with a war quote people could use asynchronously whereas back in 2017 a lot of the work was being done at data rescue events and we knew we didn't have time to do that. we needed to spin something up in within a week at the most um to go after data that we perceived would be um at risk. And so we created this data um inventory spreadsheet and workflow that I am proud to say is still going strong and it's I back it up every day but um it is uh still a lot of people in there working um but it provides the the good thing about this and going to Jack's point this provides a data curation workflow that people can use as they're going in to target specific data sets um for rescue and gives them as much information as possible about how to think about a data set as a unique object with a lot of other objects that go with it as a kind of a package um that needs to be preserved for the long term. And this is uh one of the things you'll see on this as well is that we worked also with end of term web archive, internet archive, wayback machine um and to nominate the web pages that these data sets were coming from for um crawling. Um but we also wanted these these data sets to go into a place that was built for data. Um that was that was a major goal for us and so that place ended up being ICPSR's data lumos. Um ICPSR is based out of the University of Michigan. It's a data archive that's been around since I think 1963. Um so has a long-standing reputation and in 2017 or so they built data lumos which was is meant to be a crowdsource repository for federal data. uh that repository um as of 2020 January 2025 had about I think a hundred data sets from the 2017 era um uh since then we've put about 500 in um to this uh repository uh um and we know I um got statistics and I totally forgot them but we know that since February to April end of April there were about 6,000 downloads of those and that's gone up exponentially. So I think it's up in the 20,000 downloads at this point of these data sets that have been put in since February 2025. Um so people are looking at the data. We'd love to know how people are using the data as well. But right now we do know that they are being um they are being downloaded. Um in addition to the data rescue project or the data rescue efforts that we have, we wanted to learn from last time and one of our big things was to make sure that we knew where data was going. We we did not think we would ever be the only group um working on this. Uh uh so we wanted to partner with other groups like EDG PEDP um DAC library innovation lab and other groups um that were doing data rescues to try and make a catalog of those rescue efforts. Um and that's what our data rescue tracker does. Um it provides um the original locations as well as the down backup links, the URLs to the backups. um for a wide range of of um data rescues. You will see repeats in here and this goes to Jack point about lots of copies. Um you'll see different versions of ERIC for instance one called Erica which you may have seen. It's a beautiful site someone made from the internet archives backup of ERIC. Um and we are thinking right now about other places we other um sources of data we might be able to create interfaces using both the data sets that we have and um uh uh the um backups that would be in the internet archive. So there is an interface hopefully coming soon. We it's been a lot of work to try and get this um done but uh this interface this current interface you can go and look at and I can put the links in the chat. Another big part for us is thinking about data stories and how how do we tell the stories about the importance of public data as a public good. Um that this is something that the government needs to be doing um gathering this data for our um to be able for us to be able to understand um our country. And so we have a a submission submission form for collecting data use stories um and we've published a couple of these but and I have some more coming. Um but if you are interested in telling your story about a data set, how it's been useful for you or how it might be useful for um understanding your community, um feel free to do that. Uh I know Civic Switchboard also did a version of this um as a data rescue event which I find really exciting. Um so if you think that might be something would be um interesting to your people, uh get in touch with us and I can connect you to them. And then uh if you are interested in doing more, please uh get in touch. Uh and um just sharing information about this and talking about public data is is helping um so make sure you're doing that. Um even if you don't have time to rescue data or help with a tracker or any of the more advanced things that we do. So thank you very much and I look forward to the conversation. Thank you all so much for sharing about your fantastic projects. We'll now move on to individual questions for each panelist. Molly, what tools, training, or partnerships are most effective in starting or expanding a crowdsource project such as yours? Yeah. So, I would just start by saying um I feel incredibly fortunate to be in an institution and a department that really um supported Jenny and I kind of going for this. And I think the the three things I would really say you need is you need um support of other people. And this was one place where we were just really lucky to have a lot of colleagues at the University of Minnesota libraries that were interested in helping us with sort of beta testing our form to make sure um we knew that if something was confusing for a librarian, it would be really confusing for somebody who doesn't work with information all day every day. Um and we were also just really lucky to be able to lean on the efforts um that others have done. Um, one particular person I want to give a shout out to is Kelly Smith at UC San Diego, has a really fantastic weekly roundup every week where she compiles um, a lot of the stuff that we're interested in, like things that have gone down from government websites, but also other news stories where we really want to keep our eye on how is this impacting the information ecosystem. Kelly gave us permission to take that list and kind of put it into our own spreadsheet. And then what we did is we kind of um had a series of working lunches that we just invited our colleagues to come to where they helped us take um the the sites that Kelly had already found and we had vetted them a little bit to make sure they were appropriate for our particular project and then put them into the form um so we could see how it went into the spreadsheet and then they could give us feedback. Um and that was really wonderful because I'm a social sciences librarian, Jenny is a government publications librarian. we got to get insights from science librarians, from catalogers, from people with different areas of expertise. Um, so I would just say and yes, thank you for um throwing Kelly's weekly roundup in the chat. It's a phenomenal resource. So I think just getting interested people um and then being willing to just kind of we made a decision at a certain point where we were kind of like, is anyone else doing this? Should we wait? and we decided we really just wanted to get the project off the ground because the process of developing it, we knew it would be iterative and we needed feedback and support from other people because it was a crowdsourcing project to kind of take it to the next step. Um, so I'd say also just the boldness to like start when it's time to start. Thank you, Molly. Um, the next question is for Jack and the question is, "What strategies are you putting in place to ensure the long-term sustainability of the archive?" Yeah. So, I um I got to talk a bit about this in my introduction, but I'll expand on it. Um, I think, uh, as we try to take on these very broad challenges, these are, um, like huge numbers of, uh, data sets that need to be preserved. I've never even have managed to estimate the size of the problem we're taking on because even deputy US CTO's don't know how much data the federal government has. Um as we try to take on these very large things I think we need to be mindful of the limits of our own institutions and capacities that we can't say like I have it therefore it's fine forever. We have to say I have it therefore I can help others to have it. Um so I think trying to learn as much as we can from um the sort of successes and failures of digital archives is uh really important. Um and uh the kind of you know there's a whole crop of like digital humanities projects that like launched and shared new access to something and then ran out of funding and then shut back down and it was like well there was access for a while and now there's an archive of an archive. um to make things upfront um as cheap and resilient as possible is my lab's approach to this. And um I'm conscious that I'm I'm kind of I'm saying this thoughtfully because I think there's other approaches. You can have a big institution that says we're going to keep this and we'll make sure that we stick around it and we'll be fine. But I think we also have to think about make something that doesn't rely on any one of us. Um so I mentioned some of the strategies there. uh if you use uh cryptographic signatures and timestamps, you can make sure that the copies are just as convincing as the originals. And I think that'll become really important as we have copies going around. If you use self-documenting archive formats, formats that if you just got a disc and handed it to someone, they could make sense of how you think about your archive, that becomes so much easier to copy and share. Um and there's a lot of room to kind of build on um like the great work that's gone into standards like Bagot for making things that are kind of self-documenting. um if we put them on hosts so that the copy is easy to make. Um so uh the data.gov archive for example is on source co-op which is a nonprofit that donates space on Amazon S3. Um and the whole thing is one big folder layout. So if you want a copy of my entire archive that's like a one command that it's a copy that'll take a long time to run but all it's doing is downloading the whole um layout. Um, I think like redesigning our archives to be very simple and lowmoving part and designed to be like um, you know, like the Soviet jeep that can't fall apart because it's just very simple parts that you can always put them back together and they work. Um, that's what I'm looking for for how to do this resilience. Um, and then I think what we can build on that is um, a an international community that looks out for each other and that has a kind of international solidarity to it. Um, this is a real moment for us all to look up and say, I can't count on any nonprofit in the United States continuing to have an interest in preserving the stuff that I care about. Um, and uh, I've found so much um, like reassurance, I guess, uh, uh, insight and support from people who have been working in other very different situations for a long time where they're trying to preserve things that are hard. you know um a group of Tibetan democracy activists who like for decades have been working on a challenge that is extremely different from mine but um deep and important uh challenge um who will come and say hey you know we know some things that you might want to know we could help you out um and I think uh if we really invest in those international relationships to um help each other out to provide long-term resilience um that's what I'd like to see for uh for durability um and I mean I I'm at the Harvard Law library where I don't know if you saw we just announced we have a like Magna Carta from 1300. Like we have, you know, books over in the vault that are from before the printing press or bound in metal and stuff and like we're good at keeping things for a long time, but we should never say like therefore we'll be the ones who keep the copy. We should say therefore we made it really easy to have lots of copies of this thing. Uh Linda, I have the next question for you. Um you've been involved with the data rescue project from the start. Can you tell us a little bit about what you've learned about organizing largecale data preservation efforts and how the project has evolved over time? Yeah, and this will go I'm sorry I'm laughing because it was great what you said, Jack. This will go directly with what Jack was saying and that I think when we started um I mean one of the the truths of why this project has been so um powerful is because you know the data community came out and said we're going to do this. were going to be part of this conversation um because these are the things we care about um and and we had tremendous backing from a wide number of organizations. But I think what really helped us is that um we weren't territorial about it, right? It wasn't a matter of saying this is just this is our space, stay out. It was a matter of saying, "Okay, who's who wants to help? Who wants to come into this? Everybody's got to be a part of this." And that um I think DRP is a great example of that because we have um one of our steering committee members is from such which is a group called saving cultural saving Ukrainian cultural heritage online. He organized that effort in 2022 and then came to us and said okay here's how we did it. Here's here's the tips I would give you. Plus, we were able to learn from the previous data rescue efforts and from initi web archive and from Jack and from all of these um these efforts that were going on. PEDP is another um the public environmental data partners is another one that was doing this. Um but we wanted to find our niche and so being able to find our niche but also work with others and not cut other people out of that process I think is what made it impactful and successful. Um, and that that uh is where we've really evolved from being kind of this just okay, the data community is going to do we care about data, we're going to do data to much more of having bigger conversations about um things that go outside of just data sets. Um, and a lot of our conver a lot of the uh contacts we get are people asking us what what do I do about this particular thing that's not data set? and we can connect them to the right people because we've stayed open and connected to the the entire community. Um, so yeah. Oh, yay. Yeah, thanks for um so I uh yeah, that's how we've evolved over time. Um we are still thinking about our future and where do we go next? So I think we're continuously evolving. Um and every month I feel like I'm I'm saying well is it over yet? are we have we hit the you know are we going to sunset now and realizing no there's still things we have to do and so um ask me in another month where my head is at for that question. Thanks Linda. Um so I'm we're going to ask a few questions for the whole group to answer um of panelists and then we'll open it up to the audience. Um, so this kind of this question gets at how big the problem actually is because I think it and I like I know nobody here has an actual answer like it's this much has disappeared. Um, but how do we how do we sort of measure the scope of this problem and it is there a reliable way for us to measure or track um really how much government information or how much data has actually disappeared. That's a big question. So I'll open it up to any of you to um start answering. Uh I journalists ask us this all the time. So to be quite frank uh I was hoping with the tracking gov of info project there was be there would be more of a quantifiable element to this but it's a really hard we can say this I think from agency to agency probably more so than across the entire government. But Jack, you were going to Yeah, I think there's a real kind of tooling gap here. Um because we're kind of the community that I've been part of has been focused on making copies of the thing. And then um to answer a question like when the CDC website went down and then came back up after an executive order, exactly what changed? uh there is tooling that could exist to answer that question based on end of term archive crawls for example but um I don't know that it does exist robustly that we we have answers on the time frame that people need them about what changed um so it's it's one of those frustrating like we have the data but we don't have the way to ask the question of the data yet um and that's really specific to the things that can get into the crawl um when you get into the data that is um you know deep web like stuff that is not crawable I think it gets even less possible to answer. You get this kind of um oh we found like one link on data.gov that goes to an archive that turns out to be pabytes of weather data and like there's a whole question about how you do your denominator there. Like first of all we didn't know those pedabytes existed so raise your estimate by that much. Second of all does that count or are you going to say well that one we actually just don't have storage for. We're going to write that off and if the government can't save it no one can. Um you might say this megabyte over here is much more important than this pabyte over here. Let's save a bunch of those. And how do you measure your progress when you save this megabyte but not that pabyte? Um, so I'm kind of I'm left back to I have no denominator here for a sense of what we've lost or what we have. Uh um I it would be really helpful in talking to my funders to have better answers to that. So I think uh I I love the idea of getting better answers. And I think one way to approach it would be to work from um the requests that we have and the requests that reference librarians get to start to understand, you know, what do we think is missing from a just practical our patrons trying to answer questions. Um as far as I know, that infrastructure doesn't exist at a large scale either. And I'd love to I'd have to hear more about that. And yeah, I just have to echo what um both Linda and Jack are saying. um our project is maybe especially challenging because we very intentionally want to capture things that aren't just quantitative data, but then those things are even harder to measure, right? Like it um and I know this will get in, we're going to talk about this further, but there's also a tricky thing when you're tracking like things like language changes to make sure you're not confusing innocuous changes with things that have been deliberately changed um for for reasons of censorship. Um, so I don't know if it's totally possible to, you know, definitively say this is the scope of the problem. I really like Jack's framing around going back to patrons, going back to users, what do people um need? And certainly our project came up in part because of like questions we were getting from faculty around specific things that they want to make sure that they're able to access and access moving forward. Um, and again, that's why we kind of did a crowdsource project, but it's just kind of trying to get as much as we can. So people understand the breath of things, but I don't know if there's really a way to quantify everything that's been lost. Molly started to get into this a little bit, but on a related note, how can we differentiate between routine activities and updates and intentional removal of government data and information? I I'll add something to that though. It's not just um it's not just that binary, it's also um the contract situation. So we've I think there's um so there's the the focus on tensional, there's a focus on um you know these the the the updates that have happened where we think something's down but it's just being updated or being fixed, but there's also um you know contracts that are being ended. And so that's like leads the concern for whether or not there's going to be um a place to have that data. So, um, that's led to some scares. And then the fourth thing that I would add is, um, is the lack of staffing. And this is a big one for me is that it's not so much that the data is gone, but the people who steward the data are gone. And we've seen, um, instances where because there's no one behind in the agency anymore who has that data expertise that things deprecate. Um, and so having that understanding of how to um, where that's happening, I think, is another part of it. So, sorry to complicate it, but it's it is a lot more complicated. um that just those two one one answer there is we shouldn't necessarily differentiate between those if we start if we use this as a prompt to think about stewardship in general um I had a conversation with someone who um works in the government who said look you have to understand that no one can archive themselves um and she said this I didn't but it was uh every library's archive of themselves is their worst archive um and she said that she's certainly seen that in government that um it's not it's not workable to expect people to preserve their own data data. It's a different um skill set and um the government has restrictions that make that even harder. Um and so when we were archiving uh data.gov back in November, December, there was lots of link rot already. Just things that were supposed to be there that weren't there anymore, as any one of us would assume. 300,000 links, of course, a lot of them aren't going to work. Uh and that's actually just as bad for our patrons as if someone took it down intentionally. Um if it's gone, it's gone. Um, so I think I think we can shift from saying what is what is gone intentionally to really like what is necessary to keep and what systems are we going to build to keep it. Um, yeah, I had another thought, but I'll leave that there. Um, and I guess I just want to join Linda and further complicating this. Um, in addition to data that there's no longer stewards of this, uh, there's a huge problem right now with communications that we should be receiving that have no longer are no longer being received. like the it was just announced last week that the CDC has not been releasing newsletters since March. Um even though diseases have continued to be a problem. So there's like and that's a tricky thing because you can't point to like look this was here in this place now and now it's gone. Um but that's also a loss and that can be incredibly difficult to measure. Yes. One thing that's definitely happening is um data collection is shifting. So like collection of demographic data is getting more narrow and um you know things like race or gender uh questions a bunch of forms have been updated to not collect anymore. Um another thing that's happening is foyer offices aren't being staffed. So you get like oh the data still exists somewhere but the person who would hand it over to you doesn't have their job. Okay. So another question for the panel. um what resources are most needed to support and sustain preservation efforts whether that's from your individual projects or sort of your ideas in the big picture of things. Uh I think first I would get staff for Linda. Uh I think the um here here's where we are is like we're all touching parts of the elephant. It's like it's a huge problem to preserve 2 million federal employees output of all the things they're seeing and giving us which is like it's a gift to us. It's valuable information relevance. It's a huge problem to preserve it. Um, but there's no elephant there. Like none of us thought it was our main job to preserve stuff that the NAR was going to preserve. Um, so we need to construct this thing and we can do so much with volunteers, but we can't do community organizing with volunteers. You can um, you know, there's a Harvard GIS librarian who cares a lot about GIS data and would happily chip in information about GIS. Um, but uh, the person who keeps a spreadsheet of all the GIS librarians, that can't be a nights and weekends job. That has to be a kid job. So we need um core support for the community organizers who um uh keep the list of all the others of us who would love to help out and that just has to be funded and we can't keep looking and pointing at someone else to do that. Uh and then the stuff that is around that is we need to using that that framework um uh support our whole profession in being able to contribute to the bits and pieces as part of the other work that we do. I think it'll really pay off for a GIS librarian to be in a community of other GIS librarians helping to figure out how to preserve public data. You chat with each other, throw in a little bit of advice, get answers to your things. I think we can actually use this as a a nexus of a um real support for the profession. Um and and so building that space is kind of feels to me like the next thing. But um we can't build it on pure volunteer, sweat, and blood. We have to build it with resources. So I think that those are the resources that feel most key to me. and the rest of it feels downstream from that. I I would I would I I don't disagree with that, but No, no, do please. It'll wake everyone up. But but but no, I think that I the one thing I would add to that is I think there needs to be a a a mind shift about and and Danielle's comment about this is so important. Um, I think we have taken public data for granted and government information for granted for so many years since I've been a librarian. And I've had people say to me, um, that, you know, oh, well, the census was is out always outdated, so I'll just use this vendor over here. And it's that's that's been a problem for us. Um, I think because we that vendor depends on that public data to be able to create the the kind of data that they're they're selling back to us as librarians. And so um we need to advocate for uh this government for government data, government information as a public good because the government is the is the institution that has the resources to collect and disseminate the data in a way that no vendor can do, does not have the interest to do. Um and it doesn't exist to do that. Um, so, so that's I think there just has to be some and librarians I would say need to to to move that along to to encourage people to that to to take that um to to to be the advocates for that perspective because it really um it it's that's been what's maddening to me. I was a government docs librarian for many years and and um it was always maddening to see people kind of dismiss the government documents program or the government information programs um as less important. But now we're in the situation that we're in because of that. Uh yeah, and the the last question we have for everybody is something that I know came up a lot in our Pippers discussion, which is, you know, a lot of people want to help and don't know how to get started or think I don't have the coding skills or things. So, how can we empower more people to get involved? I can certainly say on our end, one of the things we've been really trying to do is just kind of make the barrier to entry as low as possible because um Jenny and I are lucky to have jobs where we are able to kind of um use some of our skills to kind of analyze and clean up things that have been submitted to us. Um, and so I I've I've tried to just like always make clear like even if you see something and you're not sure like is this really something that was taken down for nefarious reasons, I I want to look into it. Um, because I think sometimes people can feel some insecurity about stepping into something like this. Like do I really have the expertise? Um, and I think one thing that's just really important is that just just like encouraging people to get involved at whatever level they're at. I I would say the same for the in the whatever level people are at. Um I mean we uh there are lots of people who've gotten in touch with me and said that they couldn't they can't do a rescue or they don't have the time or the it's beyond what their skill sets are. Um, but there are other places, other ways that you can get involved. And certainly just, um, you know, making people aware in your family, having a family conversation, being that person at the dinner dinner table who's annoying everybody about government information, I think is is something that you can do um to to to to raise awareness for the problem. I think there's room to join um like the forums where people are working on this stuff. Um, so uh DRP as that one grows. Um, if you're interested in environmental data, PEDP has a good community. Um, the safeguarding research and culture, uh, safeguard.de, which I linked, um, has a whole open forum that that one attracts more of like the Reddit data hoarders crowd. So, it's a nice place to hang out with more of a technical community where you'll probably be the like the best at metadata of anyone there. So, like you'll have a lot to contribute because they'll be taking on challenges that you've thought about more than they have. Um, I think getting into a community and seeing where you can chip in is really helpful. And really it's having more of those communities is what's going to help us to um grow on that. Um I have been thinking about kind of what is the most DIY version of this thing. Um and I do think we actually there's room for exploring things like uh um a archive web.page by the web recorder project will let you run an app that is just a browser a special browser that records everything you do. And I think there's room for some kind of DIY data preservation via clicking around and downloading. Um, and we do have to figure out more ways to do that. Um, I'm just I think we do struggle with this question because in doing the entire thing, I really hit like I can't write down all the steps of this thing. I just need to tell you to talk to six people and it is going to be how do we build the places to talk to six people that that solve this. [Music] Thank you everyone. Um, one thing I want to ask is if is there a question that we didn't ask that you had would want us to ask in this in this context because I see a lot of themes coming up um that were a little bit more unexpected and I I would really love to hear if there's anything that you maybe wanted to touch on that you haven't had a chance to um touch on yet. I'll take that as a note. I tried to give it the 30 seconds of awkward silence. I couldn't. I don't think I could get there. That's okay. Um, so we have a couple of questions lingering in the chat, but not too many. Um, we have eight minutes left, so please feel free to add some questions. Um going way back um to uh near the beginning um Alan had asked at one point um do you think if Universal Music Group versus Internet Archive is decided against the Internet Archive if that would uh negatively impact the Wayback Machine? I think we really don't appreciate what fundamental infrastructure the Internet Archive is and how slender that read is. Um, and I'm not part of the internet archive. I certainly can't speak for their finances. I think it's never good to lose the litigation for a lot of damages. Um, I know there's a lot of us who care a lot about it. So, I imagine there'll be a lot of support for it too if there is a bad judgment. Um, but look, in many countries there is a national library that has the the purpose and mission of um, you know, mandatory deposit for the web and tries to collect the the web of their country. In the United States, we depend on a nonprofit in a church in San Francisco to do what is done by national libraries in many countries. And they do that admirably well, and it's not the only thing they do admirably well. There's so many kinds of archive that they step up and do single-handedly. Um, so on the one hand, uh, boy, we have to fight for them if they do, um, get hit by by judgments. Uh and on the other hand, um it's sort of become a principle for my lab to not try to use them only because that's the easy option and we can't all be doing that anymore. Um uh we need to have not just one slender read. Um so I I really think we need to think like what can we bring to the table that will make this thing more resilient than oh thank god there's Brewster Carol running the internet archive and he can do the parts that are too hard for us. Thanks, Jack. Um, Linda, did you want to add something? Okay. Okay. So, we do. Yeah, we have a couple Q& A's. Um, also, so uh the first is um from an anonymous attendee. I'm curious about the methods being used to acquire the data. Have there been problems with any specific method? For instance, my commercial organization runs Python robot scrapers on multiple federal and state administrative agencies and recently we were specifically blocked by a federal board site. We found a commercial workaround but we anticipate hap this happening more often with federal agencies. What happens when archiving efforts are actively blocked? Um I mean I can speak to this for DRP and we've we've run into this um quite a bit. We've definitely um almost brought down some APIs I think um or been blocked from some APIs in the work we've been doing. Um I mean you're welcome to come if you do want to come in and talk to our people. We have a group that is doing this um with a particular agency uh right now and running into issues and so you could share um and we have a whole channel just for technical discussion. Um um so thinking about DRP as a place to ask these kinds of questions I think is a is one way to um think about us. But uh when um for us we have run into situations where we're being blocked. In some cases we've been able to find people we can um in the agency that we can talk to about what's going on. Um that's rarer than um than in other cases where we have actually data lab is a good example of this that we it's a um a dashboard for department of education and we've gone in and just downloaded tables um one by one manually to be able to have something because there's no way that we can get the data um programmatically there's it just isn't possible for us to do. So, um, so we've, that's the benefit of having, um, 800 people signed up as volunteers, um, is is to be able to kind of get the the do a manual process. Um, but there are challenges with that as well. We've found this with web archiving running Perma CC as well, our um, link ro fighting tool for for courts and law journals. Um the the AI moment has completely changed the stance of many web pages towards uh archival and um they're not specifically trying to block us. They're trying to block crawlers that destroy them, but um they block us too. Uh and I do think that our community is going to hit a kind of just being excluded almost by accident. Um, and the commercial options are interesting but sketchy because like a lot of them it's hard to validate that the routes that they're using to get around the um the blocks are not just running spyware on someone's phone. And um with Permo we haven't found a way to like get comfortable enough with one of the the circumvention options that we we know that we can run it and be stand behind it. Um I think we may have to at some point because uh the blocking is getting more and more intense. All right, we have another question in uh the Q&A. Have people been sharing these things internationally? Uh we shared the Wayback Machine plugin with some UK partners recently. Yes, there is a very big international contingent. Um safeguarding research and culture is one. Um Sucho, members of Sucho is another. Um, and when I say they're contingent, I mean people who are actively interested in what's going on in this country and want to help out. Um, and they've been doing um they've been doing a lot of um uh press internationally as well. So spreading the word um in France, France, Germany, um Netherlands, all of those countries have had articles about the efforts that are going on here in the United States. For a while it was I only saw DRP in French newspapers before it kind of hit I think um hit the United States. So um the international community is quite interested um not only because they care about the data we're collecting and I'll give you a great example of this. Um I went to a session that was talking about the demographic and household surveys that the sorry demographic and health surveys that were c um collected by the USAD and someone from the UN was there talking about how um most of the indicators for the sustainable development goals for Africa the data was based on the demographic analys um without that data we don't have the UN does not have information about sub about African countries. Um uh so it's it's it's it's critical for more than just the US. Um it's a very big problem internationally. Um in addition to that there's also the other angle of it which is other countries this could happen too. And so how do we um create an infrastructure that is replicable in other countries um that could be used in say Hungary or um countries that might be experiencing similar situations. Okay, we have one minute. I'm going to ask the last question. Um, have some of the federal data stewards, perhaps especially those who've lost their jobs, found and reached out to any of you about these projects. Uh, I think we will be working with people who lost their jobs. Um, but I don't have anything to announce yet. All right. Thank you everyone so much. I'm going to put on a uh quick slide here. So um this will have a little bit of um information so you can visit our YouTube page. This will be uploaded in hopefully in a few days after the recording is um completed. And we also have a QR code here on the left. If you wouldn't mind providing feedback on this webinar, that would be wonderful. I want to thank all of our speakers and my wonderful co-hosts for putting on this webinar today. Um, God and Pippers have a really nice long history of working uh very well together on these types types of initiatives and it's been really great um to work with everybody on this. So, and thank you everybody for taking time out of your day to attend and look forward to the recording and um have a great day. Thanks for having us putting it on. Thanks, Linda. I'll probably be in touch for your help uploading everything. Okay. I think I'll have access now. So, okay. I will give it a try. Okay. If not, let me know. Let me know. Sorry if I'm I'm so sorry that we always bug you about that, too. No, no, no worries. No worries. Heather, Jennifer, we'll be in touch with um just all the things feedback and and the recording and everything like that. Yeah. And Kelly and Danielle, thank you so much for um helping with the tech side. Thank you so much for using getting to use your infrastructure for this. Yay. Versus us trying to do it on our own. So, I think it was I think it was super great. This was great. I've already gotten some chats on the side from other folks who've been who were attending. So, thank you. Um, the biggest number I saw was 195 at one time. Yep. Thank you for organizing. Yeah. Thank you, Molly, for coming. It was so great to have you talk about your work. Um, keep it up. I I really admire it because there have been times where I was like, is anybody doing this? That's where we came from. start like should I just start doing this? Whatever. Yeah. Yeah. Thank you all so much. Have a great day. Thanks you too. And thanks Julia for persisting. All right. Thanks everyone. Take care. Thanks. Thank you. Okay.