Disappearing Data: Grassroots Efforts to Preserve Government Information in Uncertain Times
Watch on YouTubeVideo summary
The webinar series hosted by Kate Tolman brings together experts from institutions like the University of Minnesota, Harvard, and the University of Pennsylvania to address the critical issue of disappearing government information. Panelists Molly Blake, Jack Kushman, and Dr. Linda Kellum outlined various grassroots initiatives designed to track and preserve data that is being removed or neglected. Key efforts include a crowdsourcing project to document missing documents and broken links, such as those from the CDC and the Trump administration, alongside Harvard's creation of a massive 17 terabyte index of datasets using self-explanatory formats with cryptographic signatures for integrity. These strategies emphasize the need for resilient, redundant archives that do not rely on a single institution, advocating instead for the global distribution of copies to ensure long-term survival against agency closures or intentional censorship.
The discussion highlights that data loss is caused by both systemic neglect and intentional removal, illustrated by examples like federal forms ceasing to collect demographic data and field offices becoming understaffed. Preserving the output of millions of federal employees requires dedicated core funding for community organizers rather than relying solely on volunteers, as the departure of staff who steward the data often leads to significant information gaps. To counter the tendency to rely on private vendors lacking mission-driven resources, advocates stress framing government data as a public good and lowering barriers to entry so that individuals at various skill levels can participate, from manual collection to joining technical communities. However, challenges persist, including API blocks and websites optimizing against crawlers, which sometimes force reliance on unverified commercial workarounds to access essential information.
International collaboration is emerging as a vital component of this preservation movement, with partners in the UK, France, Germany, and elsewhere recognizing that US demographic data supports global initiatives like the UN Sustainable Development Goals for Africa. Despite the difficulty in quantifying the full scope of lost information due to tooling gaps and JavaScript-rendered content, the consensus is that avoiding territoriality and fostering international solidarity are essential for resilience. The panelists concluded by sharing plans to distribute resources and recordings while acknowledging the urgent need for infrastructure that goes beyond single points of failure like the Internet Archive, ensuring that vital government records remain accessible despite an uncertain digital landscape.
Read the full video transcript
Hi everybody. Going to give it a minute
while attendees start to come in and
participants are led into the webinar
room. And in just a moment, we'll go
ahead and get started.
All right, it has slowed down just
enough for me to feel comfortable to get
started. So, um, welcome everyone and
thank you so much for attending this
webinar today. My name is Kate Tolman.
I'm the uh head of library liaison
services at Illinois State University
and I'm also the I guess chair of the
help I'm an accidental government
information librarian webinar series. Um
today's webinar is brought to you by uh
the ALA government documents route table
or goort in partnership with the ACRL
politics policy international relations
committee or pippers. We have a really
amazing panel today as you will see um
to talk about disappearing data and
disappearing government information, a
topic that we are all intimately
familiar with. So I'm going to um let
our colleagues from Pippers talk a
little bit more about the impetus for
this webinar. Um and uh very quickly uh
you are welcome to chat uh questions any
time. We will keep track of those and
ask them at the end of the webinar. We
are going to be giving each speaker a
few minutes to talk about their
projects. Then we're going to ask each
of them individual questions and then
we're going to ask them a group of
questions that they can answer together
in a discussion format and then we will
open it up to you for questions as well.
So that's going to be the general
structure of today's talk. Um in the
meantime, I'm going to hand it over to
Julia Ezo. Uh Julia is the chair elect
of God and uh the current chair of God's
program committee. She is the government
information and political science
librarian at Michigan State University.
So Julia, go right ahead. Thank you,
Kate, and thank you everyone for
attending today's incredibly timely
webinar on grassroots efforts to address
disappearing federal government
information. We've gathered a wonderful
panel of three experts who are in the
trenches working hard to preserve
critical government information and
ensure continued free access for all.
I'll now pass it over to my Pipper's
colleagues, Jennifer Horn and Heather
Gatsay.
Jennifer, you are muted. You're muted.
Okay, I'll start over. I'm sorry. I'm
Jennifer Horn. I am the um business and
government information librarian at the
University of Kentucky and along with
Heather Gatsay, I am the co-chair of the
Pippers professional development
committee. Uh back in April, we held um
Pippers held a discussion um about the
unstable federal data environment and it
was very very popular and it was so
popular that we had to restrict uh
attendance um or else it was going to be
out of control for a discussion group
with no uh agenda and no speakers. So,
we knew that we needed to work to put on
a more formal webinar um with speakers
and um we were very happy to be to hook
up with Goor to be part of this webinar
series um to put this on today. So, and
now I will pass this along to Heather.
Okay, thank you. Um hi everyone. I'm
Heather Gatsay. I'm the uh social
sciences and government information
librarian at Slippery Rock University of
Pennsylvania. Um, as Jennifer said, I'm
co-chair along with her of the Pipper's
professional development committee. Um,
and we'd like to welcome our speakers
today. We're so thankful to have their
um, participation for this session. I'm
going to provide a um, brief
introduction um, for our speakers. I
want to encourage them to share any
additional information that they might
want to do so um, regarding their
backgrounds um, as they share their
information today. Um, so joining us
today we have Molly Blake who is social
sciences librarian at the University of
Minnesota Twin Cities where she serves
as liaison to the economics, educational
psychology and political science
departments and also to the Hubard
School of Journalism and Mass
Communication. Um, she has additional
professional experience in teaching,
writing, research and nonprofit work. We
also have joining us Jack Kushman who is
director of the Harvard Library
Innovation Lab um a software and design
lab at the Harvard Law School Library.
Um he is a software engineer and
appellet attorney uh who previously
worked as lead developer of the case law
access project and has served as a
lecturer on computer programming at
Harvard Law School. And we also have Dr.
Linda Kellum who is one of the founding
organizers of the data rescue project
and the Snyder Granador director of
research data and digital scholarship at
the University of Pennsylvania
libraries. She has authored and edited
works on data and librarianship and has
presented extensively on data services,
data management and fair guidelines. Um
we want to thank all of them for joining
us today and we're going to start things
off with Molly Blake.
Awesome. Thank you so much Heather. Um,
I'm just going to share my screen quick.
And is it letting me do that? Let me
see. Yep.
All right.
Fantastic. Um, so I just wanted to start
today by um sharing uh the team behind
this very new project and thank you so
much for this opportunity to chat a
little bit about this project. Um, this
project was really brainstormed by my
colleague Jenny McBurnernney. And Jenny
and I have been working closely over the
past eight weeks to just really kind of
get things off the ground. Um, and I'm
very happy to share that just in the
past few weeks, tracking government
information has been joined by two other
fantastic librarians, um, SA at the
University of Illinois and Ben at
Sacramento State.
Um, so like probably all of you in this
room, in February of 2025, Jenny and I
started to get really concerned with
what was happening with government
information. And one event in particular
sort of inspired us to take action. two
of our colleagues over at the School of
Public Health at the University of
Minnesota um shared that a journal
article they had written for a CDC
operated journal had been removed from
the CDC website. Um fortunately with the
February 11th ruling that CDC data had
to go back up, their journal article
also went back up. Um but as these
researchers pointed out in their
article, this was not the only instance
of government information being taken
down. Um and one they pointed to was uh
videos of the January 6th insurrection
were removed um from I believe the
Justice Department website around this
time. Um, and then the other thing that
happened, as I'm sure all of you know,
in March, uh, the Trump administration
started releasing all these memos
banning a very long list of language to
be included in public government
publications moving forward. And so that
really underscored to Jenny and myself
that we needed to do something to track
what was going on. Um, we knew that
there were really excellent efforts
going on such as the data data rescue
project to rescue data. And so that made
us wonder is anyone sort of tracking
everything else, photographs, videos,
websites, etc. Um, and so we decided to
start a crowdsourcing project. Um, so
right now our goal is to help
researchers and just members of the
general public see the scope of
everything that um, the Trump
administration has removed from
government websites and to locate
archived copies whenever available.
Fortunately, quite a bit, if not
everything, is available in the way back
machine, but it's about connecting those
dots.
Um, and I'm sharing here, and I will
share it in the chat when I'm done
speaking, um, our current working
website. And as of right now, there's a
form where absolutely anyone is welcome
to submit anything that they see is
missing. Um it can just be a title of a
document or a broken link. We're happy
to investigate further on the back end.
And then that form feeds into a
spreadsheet so that um folks can kind of
take a peek of what has been shared
already and sort by things such as
agency. Um I believe as of today we have
about 100 entries. So, we're just really
happy um to be able to just be uh
continuing to collect this information.
Um so, what has been removed other than
data? For the sake of time, I'm not
going to read this slide, but I think it
gives you kind of a good sense of the
breadth of what has gone away at this
point. Um some of you may have heard it
was just announced that all of uh
Trump's speech transcripts have been
removed from the remarks page of the
government website, uh the white
house.gov website. Um, we have working
papers that have been removed. Tools
that help researchers look at data or
work with government documents have been
removed. Uh, and maybe the most famous
example um of an entire website going
down. Two or 3 weeks ago, the co.gov
website started rerouting to this
interestingly designed website about the
quote unquote true origins of CO 19. So
that was once a website where you could
get information about treatment options
and vaccine safety. that has all been
removed.
Um so we are still um in the early days
of our project and we are currently kind
of expanding and growing. Um the next
goal we have is we are in the process of
developing a search interface to make it
easier for users to explore um what
people have submitted. And then we're
also looking into options for expanding
our project moving forward. Um, two
possible routes we might go down is, um,
we're interested in tracking things like
office and agency closures that are
going to impact the information
ecosystem already have impacted it. So,
I imagine quite a few of you today were
maybe at the EBSS webinar this morning
where we all heard from Aaron Polard
Young about how the decimation of the
education department has impacted ERIC.
So that would be one of several examples
of how these budget cuts are then
impacting information we have the right
to receive um as a p as members of the
public. And then the second area where
we're interested in expanding is this
ongoing really terrible issue of
language censorship. Um since the
initial list of banned words came out in
March, this hasn't gone away. Um, just
this past week, the Guardian reported
that State Department employees had
received a memo that they are no longer
allowed to use the term racially or
ethnically motivated violent extremism.
Um, which has very profound impacts
about how white supremacist violence,
the number one form of terrorism in our
country, gets described and tracked on
government websites and resources. Um,
so I'm going to leave it at that and I
look forward to chatting with you all
further.
Uh, thanks so much. Shall I pick it up
from here? Let me share my slides. Uh,
[Music]
all right. So, I'm Jack. I'm direct the
Library Innovation Lab, which is a group
at the Harvard Law Library that explores
the future of access to information. And
we try to take the library principles
that we all share and bring those to
wherever people are getting their
information today. Um we uh sort of
believe as a core principle at the lab
um one of the uh one of one of the core
things that we're trying to solve is the
preservation of cultural memory on the
web that it is on thin ice across the
board because uh in the same way that um
information on the internet wants to be
free. It also wants to exist in only one
copy. We're no longer starting with you
know 10,000 books in 10,000 libraries
and winnowing down. We're starting from
a a real economic pressure to go down to
one copy that we all refer to. Um, and
we're here to talk about public data.
So, that's very true of things like
data.gov, but it's also true of the
YouTube videos where there's countless
culturally valuable YouTube videos that
if YouTube changed a policy, they'd go
away. It's true of payw walls. Uh,
there's countless uh uh articles that if
they were changed, it'd be very hard to
show that they had been changed. Um, we
have a acrosstheboard cultural memory
problem.
um that we all need to invent strategies
for solving. Uh so that's a big picture
uh aspect of what we're thinking about
as a lab. Um with public data, uh
there's this real pressing need to
preserve the stuff that we all care
about because America runs on data. Um
spending some time browsing data.gov um
as we started to preserve it uh at the
end of last year. Um there are endless
data sets where you look and say, I sure
hope someone is tracking that. I hope
someone is tracking how much water is
left in the aquifers, which crops are
growing, which places are overcrowded,
which places are a good place to build.
Um the we desperately need a view of all
of the things that we are doing as a
society. Um and what we anticipated in
November and what turned out to be true
I think probably more than anyone
anticipated is that uh the parties of
the federal government in providing
those things drastically changed. I list
here um some but not all of the agencies
that have had their missions drastically
cretailed or changed uh just in the last
few months uh that are agencies involved
in public data in access to the uh the
information that we all use to navigate
um through our lives. Uh there are so
many and this is this is not the full
list. Um but there's a huge shift in
what information is being offered and
processed and preserved. Um, and so we
faced this question, can we archive the
government? Um, and of course not all of
it, but we did dive into a little piece
of it. So at the library innovation lab,
we decided to take on archiving
data.gov. Um, data.gov is a uh an index,
a metadata map of about 300,000 data
sets across the federal government. And
we were told over and over, well, that's
not all of them. There's a lot more to
get. But this was like a place to start.
Um, so we made a 17 terabyte index of
um, those 300,000 data sets, just the
direct files that were linked out from
each entry on data.gov, which means in
some cases we missed a lot because it
would be um, just like a a homepage of
uh, a bunch of other data that's
somewhere else. But if it did link
directly to the um, the data files, we
would get that. Um and we really went
into doing this work because uh we were
talking to folks like the end of term
archive who do wonderful work uh
archiving um a full snapshot of the
federal web before and after each
transition. Um but really realizing that
that process of clicking from link to
link to preserve things would miss links
rendered by JavaScript and very deep
links and links where you have to email
to get a data s or do an FTP um or links
where you have to run a bunch of
searches and aggregate them that
anything that was specific data could be
missed. So you'd end up with a manual to
our public data but not the data itself.
And that's what motivated this um this
diving in to get the data. Um we thought
really hard in doing this about um how
do we make it resilient and redundant
because um I don't want to count on you
know Harvard still being around in order
to have our archives still be around. I
wanted to be resilient across
repositories. Um so we made it um using
standard uh self-explanatory formats
using uh Library of Congress bags
formats. Uh we added a metadata um
signing standard. So um each one of our
archives comes with a uh cryptographic
signature showing that we made it and
the time stamp that we made it. So it is
very hard to fake. Um, and we're working
now, it's not launched, on an interface
that uh will let you do discovery of all
this um in your browser without any
serverside component, which means that
um we've heard there's copies in the UK
and Australia and so on, those copies
will work just as well as our original.
Um, you have the same provenence uh the
same sort of uh integrity, accuracy of
the thing and the same browsing tools.
Um, if you make a copy as with the
original um and that's really where a
lot of our research is going. Um, we're
working on um publishing the tools for
that. So, uh, the core tool that we use
for making this 300,000 archives is a
tool I got to name called Bag Nabbit,
uh, which makes bags that are signed.
Uh, and this is, um, open- source, uh,
uh, free for other people to use and
adopt and, uh, remix to, um, make their
own, uh, authenticated, uh, bags. Um,
and what we learned in the process, and
I hope I'm um, setting up Linda here,
uh, you have to wear so many expert hats
to do effective archiving of public
data. um the the work that um and I wore
a lot of these hats. I was like trying
to pretend to be so many experts that
I'm not. Um to do the mapping of what
exists, the narrowing in on um what to
prioritize once you identified a data
set, the um the data librarianship to
figure out well which versions of the
file and which metadata files and which
documentation constitute this archive.
Um the programming uh skill to download
it. um the uh uh sort of process skill
to package it with the right metadata
and the right signatures. Um the uh uh
systemic skill to get it into the
archives it needs to be in so that it
will be preserved resiliently. Um and
then the sort of access and discovery
layer that is needed so that people can
find the thing that they need. Um and
then once you've done all of that, you
have to figure out how to update that
pipeline as new data comes out. Um it
just became desperately apparent that we
can't just go do this. We need a
community to do it together. Um so Linda
uh please tell us how that can possibly
happen. Um and we ended up uh with a set
of research questions that our lab is
thinking about. One of which is where
will that collaboration happen? How will
um the library community and the
technical community be able to come
together to be all the kinds of experts
that we have to be? Um how do we staff
the core community organizing that has
to happen around all the expertise that
is um at the edges of this? Uh two, uh
where are we going to do the resilient
data preservation for this? Uh I don't
know about everyone else. The world is
too big for me to say this with
confidence, but I know that in the part
of it I can see, we are too much on um
fragile archives on things that can be
lost because um we can't take everything
or because if we go away, it'll go away
because if someone orders us to change
our policy, it'll change. We need
archives that are resilient to the kind
of shock that we're going through now.
And I think that means going back to the
era of locks, lots of copies keep stuff
safe and asking what locks of 2025 looks
like, including the actual locks project
at Stanford, of course, but also taking
their concept of resilient digital
storage and expanding it. Um, and
finally to think about what knowledge
producers should learn to not let this
happen again. So going back to that
first slide about YouTube is vulnerable
and the New York Times is vulnerable and
anyone who is a primary publisher of
something they care about is at risk of
this kind of loss. using this moment to
say how can you structure your
publication so that our work as
archavists gets easier. So that's the
stuff that we're thinking about and we'd
love to continue the conversation.
Great. Thank you so much Jack. I hadn't
seen this presentation before. So that's
really exciting to to hear that's your
conversation right now. Um so I'm going
to talk a little bit about the data
rescue project u how it got started
started and what we're working on as
well. um and thinking about the future.
Um so the data rescue project um
started
uh in um started coming together in
January um of 2025 but really uh
formally came together in February but
started as a group of members of these
organizations. These are primarily data
professionals organizations. ISIS is
international, ARDA app is primarily US
Canada focus and data curation network
is a consortium of universities with
institutional repositories for data. Um
and we were really concerned about the
fact that our we had patrons who were
looking for CDC data sets in particular
or USA ID data data um and we wanted to
help them find alternate sources um in
that particular moment. So it it really
started it was a crowdsourced way to try
and and find sources um of places where
data was backed up or sources for the of
the alternate sources of data. Um and
then it became um more of a data rescue
focused um
uh organization. Um we've been doing
this unlike last time in 2017. Um and we
do come out of the 2017 effort. um they
they are our uh grandmothers and fathers
I would call call them. Um we um unlike
that period we are focused on
asynchronous data rescues. So we had to
come up with a war quote people could
use asynchronously whereas back in 2017
a lot of the work was being done at data
rescue events and we knew we didn't have
time to do that. we needed to spin
something up in within a week at the
most um to go after data that we
perceived would be um at risk. And so we
created this data um inventory
spreadsheet and workflow that I am proud
to say is still going strong and it's I
back it up every day but um it is uh
still a lot of people in there working
um but it provides the the good thing
about this and going to Jack's point
this provides a data curation workflow
that people can use as they're going in
to target specific data sets um for
rescue and gives them as much
information as possible about how to
think about a data set as a unique
object with a lot of other objects that
go with it as a kind of a package um
that needs to be preserved for the long
term. And this is uh one of the things
you'll see on this as well is that we
worked also with end of term web
archive, internet archive, wayback
machine um and to nominate the web pages
that these data sets were coming from
for um crawling. Um but we also wanted
these these data sets to go into a place
that was built for data. Um that was
that was a major goal for us and so that
place ended up being ICPSR's data lumos.
Um ICPSR is based out of the University
of Michigan. It's a data archive that's
been around since I think 1963.
Um so has a long-standing reputation and
in 2017 or so they built data lumos
which was is meant to be a crowdsource
repository for federal data.
uh that repository um as of 2020 January
2025 had about I think a hundred data
sets from the 2017 era um uh since then
we've put about 500 in um to this uh
repository uh um and we know I um got
statistics and I totally forgot them but
we know that since February to April end
of April there were about 6,000
downloads of those and that's gone up
exponentially. So I think it's up in the
20,000 downloads at this point of these
data sets that have been put in since
February 2025. Um so people are looking
at the data. We'd love to know how
people are using the data as well. But
right now we do know that they are being
um they are being downloaded. Um in
addition to the data rescue project or
the data rescue efforts that we have, we
wanted to learn from last time and one
of our big things was to make sure that
we knew where data was going. We we did
not think we would ever be the only
group um working on this. Uh uh so we
wanted to partner with other groups like
EDG PEDP um DAC library innovation lab
and other groups um that were doing data
rescues to try and make a catalog of
those rescue efforts. Um and that's what
our data rescue tracker does. Um it
provides um the original locations as
well as the down backup links, the URLs
to the backups. um for a wide range of
of um data rescues. You will see repeats
in here and this goes to Jack point
about lots of copies. Um you'll see
different versions of ERIC for instance
one called Erica which you may have
seen. It's a beautiful site someone made
from the internet archives backup of
ERIC. Um and we are thinking right now
about other places we other um sources
of data we might be able to create
interfaces using both the data sets that
we have and um uh uh the um backups that
would be in the internet archive. So
there is an interface hopefully coming
soon. We it's been a lot of work to try
and get this um done but uh this
interface this current interface you can
go and look at and I can put the links
in the chat. Another big part for us is
thinking about data stories and how how
do we tell the stories about the
importance of public data as a public
good. Um that this is something that the
government needs to be doing um
gathering this data for our um to be
able for us to be able to understand um
our country. And so we have a a
submission submission form for
collecting data use stories um and we've
published a couple of these but and I
have some more coming. Um but if you are
interested in telling your story about a
data set, how it's been useful for you
or how it might be useful for um
understanding your community, um feel
free to do that. Uh I know Civic
Switchboard also did a version of this
um as a data rescue event which I find
really exciting. Um so if you think that
might be something would be um
interesting to your people, uh get in
touch with us and I can connect you to
them.
And then uh if you are interested in
doing more, please uh get in touch. Uh
and um just sharing information about
this and talking about public data is is
helping um so make sure you're doing
that. Um even if you don't have time to
rescue data or help with a tracker or
any of the more advanced things that we
do. So thank you very much and I look
forward to the conversation.
Thank you all so much for sharing about
your fantastic projects. We'll now move
on to individual questions for each
panelist. Molly, what tools, training,
or partnerships are most effective in
starting or expanding a crowdsource
project such as yours?
Yeah. So, I would just start by saying
um I feel incredibly fortunate to be in
an institution and a department that
really um supported Jenny and I kind of
going for this. And I think the the
three things I would really say you need
is you need um support of other people.
And this was one place where we were
just really lucky to have a lot of
colleagues at the University of
Minnesota libraries that were interested
in helping us with sort of beta testing
our form to make sure um we knew that if
something was confusing for a librarian,
it would be really confusing for
somebody who doesn't work with
information all day every day. Um and we
were also just really lucky to be able
to lean on the efforts um that others
have done. Um, one particular person I
want to give a shout out to is Kelly
Smith at UC San Diego, has a really
fantastic weekly roundup every week
where she compiles um, a lot of the
stuff that we're interested in, like
things that have gone down from
government websites, but also other news
stories where we really want to keep our
eye on how is this impacting the
information ecosystem.
Kelly gave us permission to take that
list and kind of put it into our own
spreadsheet. And then what we did is we
kind of um had a series of working
lunches that we just invited our
colleagues to come to where they helped
us take um the the sites that Kelly had
already found and we had vetted them a
little bit to make sure they were
appropriate for our particular project
and then put them into the form um so we
could see how it went into the
spreadsheet and then they could give us
feedback. Um and that was really
wonderful because I'm a social sciences
librarian, Jenny is a government
publications librarian. we got to get
insights from science librarians, from
catalogers, from people with different
areas of expertise. Um, so I would just
say and yes, thank you for um throwing
Kelly's weekly roundup in the chat. It's
a phenomenal resource. So I think just
getting interested people um and then
being willing to just kind of we made a
decision at a certain point where we
were kind of like, is anyone else doing
this? Should we wait? and we decided we
really just wanted to get the project
off the ground because the process of
developing it, we knew it would be
iterative and we needed feedback and
support from other people because it was
a crowdsourcing project to kind of take
it to the next step. Um, so I'd say also
just the boldness to like start when
it's time to start.
Thank you, Molly. Um, the next question
is for Jack and the question is, "What
strategies are you putting in place to
ensure the long-term sustainability of
the archive?"
Yeah. So, I um I got to talk a bit about
this in my introduction, but I'll expand
on it. Um, I think, uh,
as we try to take on these very broad
challenges, these are, um, like huge
numbers of, uh, data sets that need to
be preserved. I've never even have
managed to estimate the size of the
problem we're taking on because even
deputy US CTO's don't know how much data
the federal government has. Um as we try
to take on these very large things I
think we need to be mindful of the
limits of our own institutions and
capacities that we can't say like I have
it therefore it's fine forever. We have
to say I have it therefore I can help
others to have it. Um so I think trying
to learn as much as we can from um the
sort of successes and failures of
digital archives is uh really important.
Um and uh the kind of you know there's a
whole crop of like digital humanities
projects that like launched and shared
new access to something and then ran out
of funding and then shut back down and
it was like well there was access for a
while and now there's an archive of an
archive. um to make things upfront um as
cheap and resilient as possible is my
lab's approach to this. And um I'm
conscious that I'm I'm kind of I'm
saying this thoughtfully because I think
there's other approaches. You can have a
big institution that says we're going to
keep this and we'll make sure that we
stick around it and we'll be fine. But I
think we also have to think about make
something that doesn't rely on any one
of us. Um so I mentioned some of the
strategies there. uh if you use uh
cryptographic signatures and timestamps,
you can make sure that the copies are
just as convincing as the originals. And
I think that'll become really important
as we have copies going around. If you
use self-documenting archive formats,
formats that if you just got a disc and
handed it to someone, they could make
sense of how you think about your
archive, that becomes so much easier to
copy and share. Um and there's a lot of
room to kind of build on um like the
great work that's gone into standards
like Bagot for making things that are
kind of self-documenting.
um if we put them on hosts so that the
copy is easy to make. Um so uh the
data.gov archive for example is on
source co-op which is a nonprofit that
donates space on Amazon S3. Um and the
whole thing is one big folder layout. So
if you want a copy of my entire archive
that's like a one command that it's a
copy that'll take a long time to run but
all it's doing is downloading the whole
um layout. Um, I think like redesigning
our archives to be very simple and
lowmoving part and designed to be like
um, you know, like the Soviet jeep that
can't fall apart because it's just very
simple parts that you can always put
them back together and they work. Um,
that's what I'm looking for for how to
do this resilience. Um, and then I think
what we can build on that is um, a an
international community that looks out
for each other and that has a kind of
international solidarity to it. Um, this
is a real moment for us all to look up
and say, I can't count on any nonprofit
in the United States continuing to have
an interest in preserving the stuff that
I care about. Um, and uh, I've found so
much um, like reassurance, I guess, uh,
uh, insight and support from people who
have been working in other very
different situations for a long time
where they're trying to preserve things
that are hard. you know um a group of
Tibetan democracy activists who like for
decades have been working on a challenge
that is extremely different from mine
but um deep and important uh challenge
um who will come and say hey you know we
know some things that you might want to
know we could help you out um and I
think uh if we really invest in those
international relationships to um help
each other out to provide long-term
resilience um that's what I'd like to
see for uh for durability um and I mean
I I'm at the Harvard Law library where I
don't know if you saw we just announced
we have a like Magna Carta from 1300.
Like we have, you know, books over in
the vault that are from before the
printing press or bound in metal and
stuff and like we're good at keeping
things for a long time, but we should
never say like therefore we'll be the
ones who keep the copy. We should say
therefore we made it really easy to have
lots of copies of this thing.
Uh Linda, I have the next question for
you. Um you've been involved with the
data rescue project from the start. Can
you tell us a little bit about what
you've learned about organizing
largecale data preservation efforts and
how the project has evolved over time?
Yeah, and this will go I'm sorry I'm
laughing because it was great what you
said, Jack. This will go directly with
what Jack was saying and that I think
when we started um I mean one of the the
truths of why this project has been so
um powerful is because you know the data
community came out and said we're going
to do this. were going to be part of
this conversation um because these are
the things we care about um and and we
had tremendous backing from a wide
number of organizations. But I think
what really helped us is that um we
weren't territorial about it, right? It
wasn't a matter of saying this is just
this is our space, stay out. It was a
matter of saying, "Okay, who's who wants
to help? Who wants to come into this?
Everybody's got to be a part of this."
And that um I think DRP is a great
example of that because we have um one
of our steering committee members is
from such which is a group called saving
cultural saving Ukrainian cultural
heritage online. He organized that
effort in 2022
and then came to us and said okay here's
how we did it. Here's here's the tips I
would give you. Plus, we were able to
learn from the previous data rescue
efforts and from initi web archive and
from Jack and from all of these um these
efforts that were going on. PEDP is
another um the public environmental data
partners is another one that was doing
this. Um but we wanted to find our niche
and so being able to find our niche but
also work with others and not cut other
people out of that process I think is
what made it impactful and successful.
Um, and that that uh is where we've
really evolved from being kind of this
just okay, the data community is going
to do we care about data, we're going to
do data to much more of having bigger
conversations about um things that go
outside of just data sets. Um, and a lot
of our conver a lot of the uh contacts
we get are people asking us what what do
I do about this particular thing that's
not data set? and we can connect them to
the right people because we've stayed
open and connected to the the entire
community. Um, so yeah. Oh, yay. Yeah,
thanks for um so I uh yeah, that's how
we've evolved over time. Um we are still
thinking about our future and where do
we go next? So I think we're
continuously evolving. Um and every
month I feel like I'm I'm saying well is
it over yet? are we have we hit the you
know are we going to sunset now and
realizing no there's still things we
have to do and so um ask me in another
month where my head is at for that
question.
Thanks Linda. Um so I'm we're going to
ask a few questions for the whole group
to answer um of panelists and then we'll
open it up to the audience. Um, so this
kind of this question gets at how big
the problem actually is because I think
it and I like I know nobody here has an
actual answer like it's this much has
disappeared. Um, but how do we how do we
sort of measure the scope of this
problem and it is there a reliable way
for us to measure or track um really how
much government information or how much
data has actually disappeared. That's a
big question. So I'll open it up to any
of you to um start answering.
Uh I journalists ask us this all the
time. So to be quite frank uh I was
hoping with the tracking gov of info
project there was be there would be more
of a quantifiable element to this but
it's a really hard we can say this I
think from agency to agency probably
more so than across the entire
government. But Jack, you were going to
Yeah, I think there's a real kind of
tooling gap here. Um because we're kind
of the community that I've been part of
has been focused on making copies of the
thing. And then um to answer a question
like when the CDC website went down and
then came back up after an executive
order, exactly what changed? uh there is
tooling that could exist to answer that
question based on end of term archive
crawls for example but um I don't know
that it does exist robustly that we we
have answers on the time frame that
people need them about what changed um
so it's it's one of those frustrating
like we have the data but we don't have
the way to ask the question of the data
yet um and that's really specific to the
things that can get into the crawl um
when you get into the data that is um
you know deep web like stuff that is not
crawable I think it gets even less
possible to answer. You get this kind of
um oh we found like one link on data.gov
that goes to an archive that turns out
to be pabytes of weather data and like
there's a whole question about how you
do your denominator there. Like first of
all we didn't know those pedabytes
existed so raise your estimate by that
much. Second of all does that count or
are you going to say well that one we
actually just don't have storage for.
We're going to write that off and if the
government can't save it no one can. Um
you might say this megabyte over here is
much more important than this pabyte
over here. Let's save a bunch of those.
And how do you measure your progress
when you save this megabyte but not that
pabyte? Um, so I'm kind of I'm left back
to I have no denominator here for a
sense of what we've lost or what we
have. Uh um I it would be really helpful
in talking to my funders to have better
answers to that. So I think uh I I love
the idea of getting better answers. And
I think one way to approach it would be
to work from um the requests that we
have and the requests that reference
librarians get to start to understand,
you know, what do we think is missing
from a just practical our patrons trying
to answer questions. Um as far as I
know, that infrastructure doesn't exist
at a large scale either. And I'd love to
I'd have to hear more about that.
And yeah, I just have to echo what um
both Linda and Jack are saying. um our
project is maybe especially challenging
because we very intentionally want to
capture things that aren't just
quantitative data, but then those things
are even harder to measure, right? Like
it um and I know this will get in, we're
going to talk about this further, but
there's also a tricky thing when you're
tracking like things like language
changes to make sure you're not
confusing innocuous changes with things
that have been deliberately changed um
for for reasons of censorship. Um, so I
don't know if it's totally possible to,
you know, definitively say this is the
scope of the problem. I really like
Jack's framing around going back to
patrons, going back to users, what do
people um need? And certainly our
project came up in part because of like
questions we were getting from faculty
around specific things that they want to
make sure that they're able to access
and access moving forward. Um, and
again, that's why we kind of did a
crowdsource project, but it's just kind
of trying to get as much as we can. So
people understand the breath of things,
but I don't know if there's really a way
to quantify everything that's been lost.
Molly started to get into this a little
bit, but on a related note, how can we
differentiate between routine activities
and updates and intentional removal of
government data and information?
I I'll add something to that though.
It's not just um it's not just that
binary, it's also um the contract
situation. So we've I think there's um
so there's the the focus on tensional,
there's a focus on um you know these the
the the updates that have happened where
we think something's down but it's just
being updated or being fixed, but
there's also um you know contracts that
are being ended. And so that's like
leads the concern for whether or not
there's going to be um a place to have
that data. So, um, that's led to some
scares. And then the fourth thing that I
would add is, um, is the lack of
staffing. And this is a big one for me
is that it's not so much that the data
is gone, but the people who steward the
data are gone. And we've seen, um,
instances where because there's no one
behind in the agency anymore who has
that data expertise that things
deprecate. Um, and so having that
understanding of how to um, where that's
happening, I think, is another part of
it. So, sorry to complicate it, but it's
it is a lot more complicated. um that
just those two
one one answer there is we shouldn't
necessarily differentiate between those
if we start if we use this as a prompt
to think about stewardship in general um
I had a conversation with someone who um
works in the government who said look
you have to understand that no one can
archive themselves um and she said this
I didn't but it was uh every library's
archive of themselves is their worst
archive um and she said that she's
certainly seen that in government that
um it's not it's not workable to expect
people to preserve their own data data.
It's a different um skill set and um the
government has restrictions that make
that even harder. Um and so when we were
archiving uh data.gov back in November,
December, there was lots of link rot
already. Just things that were supposed
to be there that weren't there anymore,
as any one of us would assume. 300,000
links, of course, a lot of them aren't
going to work. Uh and that's actually
just as bad for our patrons as if
someone took it down intentionally. Um
if it's gone, it's gone. Um, so I think
I think we can shift from saying what is
what is gone intentionally to really
like what is necessary to keep and what
systems are we going to build to keep
it. Um,
yeah, I had another thought, but I'll
leave that there.
Um, and I guess I just want to join
Linda and further complicating this. Um,
in addition to data that there's no
longer stewards of this, uh, there's a
huge problem right now with
communications that we should be
receiving that have no longer are no
longer being received. like the it was
just announced last week that the CDC
has not been releasing newsletters since
March. Um even though diseases have
continued to be a problem. So there's
like and that's a tricky thing because
you can't point to like look this was
here in this place now and now it's
gone. Um but that's also a loss and that
can be incredibly difficult to measure.
Yes. One thing that's definitely
happening is um data collection is
shifting. So like collection of
demographic data is getting more narrow
and um you know things like race or
gender uh questions a bunch of forms
have been updated to not collect
anymore. Um another thing that's
happening is foyer offices aren't being
staffed. So you get like oh the data
still exists somewhere but the person
who would hand it over to you doesn't
have their job.
Okay. So another question for the panel.
um what resources are most needed to
support and sustain preservation efforts
whether that's from your individual
projects or sort of your ideas in the
big picture of things.
Uh I think first I would get staff for
Linda. Uh I think the um here here's
where we are is like we're all touching
parts of the elephant. It's like it's a
huge problem to preserve 2 million
federal employees output of all the
things they're seeing and giving us
which is like it's a gift to us. It's
valuable information relevance. It's a
huge problem to preserve it. Um, but
there's no elephant there. Like none of
us thought it was our main job to
preserve stuff that the NAR was going to
preserve. Um, so we need to construct
this thing and we can do so much with
volunteers, but we can't do community
organizing with volunteers. You can um,
you know, there's a Harvard GIS
librarian who cares a lot about GIS data
and would happily chip in information
about GIS. Um, but uh, the person who
keeps a spreadsheet of all the GIS
librarians, that can't be a nights and
weekends job. That has to be a kid job.
So we need um core support for the
community organizers who um uh keep the
list of all the others of us who would
love to help out and that just has to be
funded and we can't keep looking and
pointing at someone else to do that. Uh
and then the stuff that is around that
is we need to using that that framework
um uh support our whole profession in
being able to contribute to the bits and
pieces as part of the other work that we
do. I think it'll really pay off for a
GIS librarian to be in a community of
other GIS librarians helping to figure
out how to preserve public data. You
chat with each other, throw in a little
bit of advice, get answers to your
things. I think we can actually use this
as a a nexus of a um real support for
the profession. Um and and so building
that space is kind of feels to me like
the next thing. But um we can't build it
on pure volunteer, sweat, and blood. We
have to build it with resources. So I
think that those are the resources that
feel most key to me. and the rest of it
feels downstream from that.
I I would I would I I don't disagree
with that, but No, no, do please. It'll
wake everyone up. But but but no, I
think that I the one thing I would add
to that is I think there needs to be a a
a mind shift about and and Danielle's
comment about this is so important. Um,
I think we have taken public data for
granted and government information for
granted for so many years since I've
been a librarian. And I've had people
say to me, um, that, you know, oh, well,
the census was is out always outdated,
so I'll just use this vendor over here.
And it's that's that's been a problem
for us. Um, I think because we that
vendor depends on that public data to be
able to create the the kind of data that
they're they're selling back to us as
librarians. And so um we need to
advocate for uh this government for
government data, government information
as a public good because the government
is the is the institution that has the
resources to collect and disseminate the
data in a way that no vendor can do,
does not have the interest to do. Um and
it doesn't exist to do that. Um, so, so
that's I think there just has to be some
and librarians I would say need to to to
move that along to to encourage people
to that to to take that um to to to be
the advocates for that perspective
because it really um it it's that's been
what's maddening to me. I was a
government docs librarian for many years
and and um it was always maddening to
see people kind of dismiss the
government documents program or the
government information programs um as
less important. But now we're in the
situation that we're in because of that.
Uh yeah, and the the last question we
have for everybody is something that I
know came up a lot in our Pippers
discussion, which is, you know, a lot of
people want to help and don't know how
to get started or think I don't have the
coding skills or things. So, how can we
empower more people to get involved?
I can certainly say on our end, one of
the things we've been really trying to
do is just kind of make the barrier to
entry as low as possible because um
Jenny and I are lucky to have jobs where
we are able to kind of um use some of
our skills to kind of analyze and clean
up things that have been submitted to
us. Um, and so I I've I've tried to just
like always make clear like even if you
see something and you're not sure like
is this really something that was taken
down for nefarious reasons, I I want to
look into it. Um, because I think
sometimes people can feel some
insecurity about stepping into something
like this. Like do I really have the
expertise? Um, and I think one thing
that's just really important is that
just just like encouraging people to get
involved at whatever level they're at.
I I would say the same for the in the
whatever level people are at. Um I mean
we uh there are lots of people who've
gotten in touch with me and said that
they couldn't they can't do a rescue or
they don't have the time or the it's
beyond what their skill sets are. Um,
but there are other places, other ways
that you can get involved. And certainly
just, um, you know, making people aware
in your family, having a family
conversation, being that person at the
dinner dinner table who's annoying
everybody about government information,
I think is is something that you can do
um to to to to raise awareness for the
problem.
I think there's room to join um like the
forums where people are working on this
stuff. Um, so uh DRP as that one grows.
Um, if you're interested in
environmental data, PEDP has a good
community. Um, the safeguarding research
and culture, uh, safeguard.de, which I
linked, um, has a whole open forum that
that one attracts more of like the
Reddit data hoarders crowd. So, it's a
nice place to hang out with more of a
technical community where you'll
probably be the like the best at
metadata of anyone there. So, like
you'll have a lot to contribute because
they'll be taking on challenges that
you've thought about more than they
have. Um, I think getting into a
community and seeing where you can chip
in is really helpful. And really it's
having more of those communities is
what's going to help us to um grow on
that. Um I have been thinking about kind
of what is the most DIY version of this
thing. Um and I do think we actually
there's room for exploring things like
uh um
a archive web.page by the web recorder
project will let you run an app that is
just a browser a special browser that
records everything you do. And I think
there's room for some kind of DIY data
preservation via clicking around and
downloading. Um, and we do have to
figure out more ways to do that. Um, I'm
just I think we do struggle with this
question because in doing the entire
thing, I really hit like I can't write
down all the steps of this thing. I just
need to tell you to talk to six people
and it is going to be how do we build
the places to talk to six people that
that solve this.
[Music]
Thank you everyone. Um, one thing I want
to ask is if is there a question that we
didn't ask that you had would want us to
ask in this in this context because I
see a lot of themes coming up um that
were a little bit more unexpected and I
I would really love to hear if there's
anything that you maybe wanted to touch
on that you haven't had a chance to um
touch on yet.
I'll take that as a note. I tried to
give it the 30 seconds of awkward
silence. I couldn't. I don't think I
could get there. That's okay. Um, so we
have a couple of questions lingering in
the chat, but not too many. Um, we have
eight minutes left, so please feel free
to add some questions. Um going way back
um to uh near the beginning um Alan had
asked at one point um do you think if
Universal Music Group versus Internet
Archive is decided against the Internet
Archive if that would uh negatively
impact the Wayback Machine?
I think we really don't appreciate what
fundamental infrastructure the Internet
Archive is and how slender that read is.
Um, and I'm not part of the internet
archive. I certainly can't speak for
their finances. I think it's never good
to lose the litigation for a lot of
damages. Um, I know there's a lot of us
who care a lot about it. So, I imagine
there'll be a lot of support for it too
if there is a bad judgment. Um, but
look, in many countries there is a
national library that has the the
purpose and mission of um, you know,
mandatory deposit for the web and tries
to collect the the web of their country.
In the United States, we depend on a
nonprofit in a church in San Francisco
to do what is done by national libraries
in many countries. And they do that
admirably well, and it's not the only
thing they do admirably well. There's so
many kinds of archive that they step up
and do single-handedly. Um, so on the
one hand, uh, boy, we have to fight for
them if they do, um, get hit by by
judgments. Uh and on the other hand, um
it's sort of become a principle for my
lab to not try to use them only because
that's the easy option and we can't all
be doing that anymore. Um uh we need to
have not just one slender read. Um so I
I really think we need to think like
what can we bring to the table that will
make this thing more resilient than oh
thank god there's Brewster Carol running
the internet archive and he can do the
parts that are too hard for us.
Thanks, Jack. Um, Linda, did you want to
add something?
Okay. Okay. So, we do. Yeah, we have a
couple Q& A's. Um, also, so uh the first
is um from an anonymous attendee. I'm
curious about the methods being used to
acquire the data. Have there been
problems with any specific method? For
instance, my commercial organization
runs Python robot scrapers on multiple
federal and state administrative
agencies and recently we were
specifically blocked by a federal board
site. We found a commercial workaround
but we anticipate hap this happening
more often with federal agencies. What
happens when archiving efforts are
actively blocked?
Um I mean I can speak to this for DRP
and we've we've run into this um quite a
bit. We've definitely um almost brought
down some APIs I think um or been
blocked from some APIs in the work we've
been doing. Um I mean you're welcome to
come if you do want to come in and talk
to our people. We have a group that is
doing this um with a particular agency
uh right now and running into issues and
so you could share um and we have a
whole channel just for technical
discussion. Um um so thinking about DRP
as a place to ask these kinds of
questions I think is a is one way to um
think about us. But uh when
um for us we have run into situations
where we're being blocked. In some cases
we've been able to find people we can um
in the agency that we can talk to about
what's going on. Um that's rarer than um
than in other cases where we have
actually data lab is a good example of
this that we it's a um a dashboard for
department of education and we've gone
in and just downloaded tables um one by
one manually to be able to have
something because there's no way that we
can get the data um programmatically
there's it just isn't possible for us to
do. So, um, so we've, that's the benefit
of having, um, 800 people signed up as
volunteers, um, is is to be able to kind
of get the the do a manual process. Um,
but there are challenges with that as
well.
We've found this with web archiving
running Perma CC as well, our um, link
ro fighting tool for for courts and law
journals. Um the the AI moment has
completely changed the stance of many
web pages towards uh archival and um
they're not specifically trying to block
us. They're trying to block crawlers
that destroy them, but um they block us
too. Uh and I do think that our
community is going to hit a kind of just
being excluded almost by accident. Um,
and the commercial options are
interesting but sketchy because like a
lot of them it's hard to validate that
the routes that they're using to get
around the um the blocks are not just
running spyware on someone's phone. And
um with Permo we haven't found a way to
like get comfortable enough with one of
the the circumvention options that we we
know that we can run it and be stand
behind it. Um I think we may have to at
some point because uh the blocking is
getting more and more intense.
All right, we have another question in
uh the Q&A. Have people been sharing
these things internationally? Uh we
shared the Wayback Machine plugin with
some UK partners recently.
Yes, there is a very big international
contingent. Um safeguarding research and
culture is one. Um Sucho, members of
Sucho is another. Um, and when I say
they're contingent, I mean people who
are actively interested in what's going
on in this country and want to help out.
Um, and they've been doing um they've
been doing a lot of um uh press
internationally as well. So spreading
the word um in France, France, Germany,
um Netherlands, all of those countries
have had articles about the efforts that
are going on here in the United States.
For a while it was I only saw DRP in
French newspapers before it kind of hit
I think um hit the United States. So um
the international community is quite
interested um not only because they care
about the data we're collecting and I'll
give you a great example of this. Um I
went to a session that was talking about
the demographic and household surveys
that the sorry demographic and health
surveys that were c um collected by the
USAD and someone from the UN was there
talking about how um most of the
indicators for the sustainable
development goals for Africa the data
was based on the demographic analys
um without that data we don't have the
UN does not have information about sub
about African countries. Um uh so it's
it's it's it's critical for more than
just the US. Um it's a very big problem
internationally. Um in addition to that
there's also the other angle of it which
is other countries this could happen
too. And so how do we um create an
infrastructure that is replicable in
other countries um that could be used in
say Hungary or um countries that might
be experiencing similar situations.
Okay, we have one minute. I'm going to
ask the last question. Um, have some of
the federal data stewards, perhaps
especially those who've lost their jobs,
found and reached out to any of you
about these projects.
Uh, I think we will be working with
people who lost their jobs. Um, but I
don't have anything to announce yet.
All right. Thank you everyone so much.
I'm going to put on a uh quick slide
here. So um this will have a little bit
of um information so you can visit our
YouTube page. This will be uploaded in
hopefully in a few days after the
recording is um completed. And we also
have a QR code here on the left. If you
wouldn't mind providing feedback on this
webinar, that would be wonderful. I want
to thank all of our speakers and my
wonderful co-hosts for putting on this
webinar today. Um, God and Pippers have
a really nice long history of working uh
very well together on these types types
of initiatives and it's been really
great um to work with everybody on this.
So, and thank you everybody for taking
time out of your day to attend and look
forward to the recording and um have a
great day.
Thanks for having us putting it on.
Thanks, Linda. I'll probably be in touch
for your help uploading everything.
Okay. I think I'll have access now. So,
okay. I will give it a try. Okay. If
not, let me know. Let me know. Sorry if
I'm I'm so sorry that we always bug you
about that, too. No, no, no worries. No
worries.
Heather, Jennifer, we'll be in touch
with um just all the things feedback and
and the recording and everything like
that. Yeah. And Kelly and Danielle,
thank you so much for um helping with
the tech side. Thank you so much for
using getting to use your infrastructure
for this. Yay. Versus us trying to do it
on our own. So, I think it was I think
it was super great. This was great. I've
already gotten some chats on the side
from other folks who've been who were
attending. So, thank you. Um, the
biggest number I saw was 195
at one time. Yep.
Thank you for organizing. Yeah. Thank
you, Molly, for coming. It was so great
to have you talk about your work. Um,
keep it up. I I really admire it because
there have been times where I was like,
is anybody doing this? That's where we
came from. start like should I just
start doing this? Whatever. Yeah. Yeah.
Thank you all so much. Have a great day.
Thanks you too. And thanks Julia for
persisting.
All right. Thanks everyone. Take care.
Thanks. Thank you. Okay.