Sustaining the Public Record: Collaborative Stewardship of Government Information and Research Data
Watch on YouTubeVideo summary
The initiative known as Democracy's Library was established in 2022 by the Internet Archive to address critical gaps in preserving government information that has become difficult for the public to access due to inconsistent maintenance and lack of comprehensive cataloging. Leveraging a vast network including its Wayback Machine, digitization partnerships with institutions like UC Berkeley's Institute for Governmental Studies, and community uploads, the project currently houses over one million items across more than 1,000 collections. A primary focus is on safeguarding local government documents, which are often scattered and not central to municipal operations, while similar efforts in Canada have digitized millions of pages through collaborations with public libraries like Hamilton Public Library. These archives serve diverse stakeholders: governments utilize them for policy research, journalists depend on historical records to track legislative changes, and citizens access materials for genealogy and property law matters.
Beyond the Internet Archive's specific projects, experts from various institutions emphasized the broader need for data resilience within scientific and governmental ecosystems against technical failures, cultural vulnerabilities like over-reliance on volunteers, and governance issues regarding accountability. Christy Holmes outlined a strategic plan involving repositories, funders, and researchers to create a resilient coalition using shared monitoring tools and advocacy strategies, while Trevor Owens highlighted how organizations like the American Institute of Physics preserve scientific history through oral histories and photographs despite challenges such as funding cuts. The community also recognized that uploaded documents are not always validated for authenticity to prevent bias but instead rely on user-provenance work supported by trusted librarians acting as "super users" who upload critical reports from sources like the Congressional Research Service, ensuring a balance between accessibility and integrity.
Significant challenges regarding access were raised during discussions about increasing blockage of government websites by bot protection measures that inadvertently hinder legitimate archival harvesting alongside defending against malicious actors. Addressing these obstacles, advocates called for negotiated pathways to continue collecting web-based materials while emphasizing the importance of including state libraries in conversations about local records management despite varying public records acts across different states. Solutions such as scalable frameworks like KOSA and existing partnerships with COSTA groups are being explored to create unified approaches that transcend individual dataset coordination, aiming instead to foster cross-fertilization between initiatives from organizations like AGU and the Internet Archive to avoid redundancy.
The overarching conclusion of these discussions points toward a future defined by collaborative stewardship rather than isolated efforts, utilizing shared maturity models designed not to rank organizations but to catalyze local conversations about improving culture, technology, and standards across different attributes of resilience. By forming communities committed to public access through manifestos and ongoing dialogue among "doers," "supporters," and "beneficiaries," the sector aims to build a robust infrastructure that can withstand technical obsolescence and shifting political landscapes. Ultimately, the goal is to ensure that government information remains a non-rivalrous public good available for generations, requiring continuous involvement from researchers, librarians, policymakers, and citizens who collectively uphold the values of transparency and historical preservation.
Read the full video transcript
Okay. Hello everybody. Thanks for coming
to our session on collaborative
stewardship of government information
and research data. Um I'm going to go
ahead and kick off our session.
Uh my name is Marily Profett. I am the
director of democracies library US which
is a large-scale effort to represent um
uh US government information that could
be federal, state, local, even tribal uh
within the internet archive. And I am
new. I have been in my position for not
even five months. Um so I'll just start
by saying my my concluding remarks are
that I am really eager to collaborate
with you around your government
information collections and fill gaps
that we may have and also to learn how
government information is used and
important in your context and for your
users. Um I have my business cards out
here on the stage uh along with internet
archive and wayback machine stickers. So
please do come up. Last time I was at
CNI, I ran out of um wayback machine
stickers, so I have a lot this time. Um
please come up and help yourself at the
end of the session. So, first I just
want to um center by saying a few words
about the Internet Archive, which
amazingly will be 30 years old next
month. Um so that's kind of fun. Um so
our mission statement is universal
access to all knowledge. And this is
really more than just a mission
statement. that truly guides everything
that we do as an organization. Um, and
we have all types of organization
represented within the internet archive.
For example, uh, software,
moving images,
audio recordings,
television news programs, um, including
we, uh, record state uh, state
television from Iran and have for quite
some time.
uh ebooks
and web pages. Uh last October we
celebrated uh having uh archived one
trillion web pages uh over the the 30
years of duration of the internet
archive. Um and all of that knowledge is
uh contained within 210 plus pabytes of
data which is hosted on our pedibyte
servers.
So I want to tell you a little bit about
democracy's library. This is not a new
effort of the internet archive. It was
actually established in 2022 and is
built on a straightforward but urgent
premise which is that governments have
created an abundance of information and
put it in the public domain but the
public can't easily access it. So, we
currently have over a thousand
collections with over 11 million items,
which is likely a significant
undercount. And I'll go into the reasons
for for why that is. Um
so the internet archive has really been
at the um in the business of uh backing
up um public documents uh uh public
information, government information for
quite some time. So I will start with
the wayback machine um which is an
important uh part of the the work that
we do. Um the wayback machine is uh in
the internet archive are part of the end
of term crawl uh in the United States
and have been for a long time. We also
routinely crawl government information
websites as part of our ongoing
operations. Also archive it. So for any
of you uh who are archive it uh
subscribers will know about this
subscription service. We have many
government sites that are doing their
own due diligence in uh backing up their
information, state, local, uh and even
national um organizations that are using
uh our our data services to back up um
to back up their their sites. So this
this includes a a lot of those uh
materials that are within the government
implicate u government information
space. We also have been in the business
of digitizing the published public
record. Um so we do this with uh with
partners who donate materials to us and
al also through digitization
partnerships with libraries and a lot of
those uh efforts have included
government information over a long
period of time and part of why
democracy's libraries underounted is
because we haven't been able to go back
and identify all the government
publications that are within the
internet archive that have been included
before the establishment. of democracy's
library. So we have a big job of going
through and identifying government
publications and pulling those into um
the internet archive. Uh some of the
collections that we're working on right
now are Supreme Court records and
briefs. When that project is finished,
which it will be very shortly by the end
of this month, I think uh we will have
the largest collection of Supreme Court
records and briefs that are publicly
available. So that's really exciting. Uh
we've also been working with the
National Oceanic and Atmospheric
Administration to digitize their library
collections. Um and uh so this is just
our effort to take physical media and
convert it to searchable digital
formats. And then community uploads.
This is also a really important uh part
of our work is uh government librarians
and others curious citizens people who
are concerned about in uh government
information will either do uh save now
on the way back machine to save and
capture pages right in the moment or
also to upload uh things like reports
directly to the internet archive. And
this is keeping uh with the theory if
you see something save something um or
lots of copies keep stuff safe.
So who uses government information? So
first of all, we know this is an area
where I would really love to learn a lot
more. So and learn with you and learn
from you. Uh but we do know that
governments themselves are primary users
of government information. Um we know
that journalists and researchers are
very interested in this information.
Government reports of course are really
important for researchers within our
public institution or within our um
academic institutions um but also
journalists. Journalists are heavy users
of the wayback machine and uh really
within the current um uh administration,
the US uh administration have been
really using relying on the way back
machine heavily to look and see where uh
things have changed uh within within the
government web spaces. Um we also know
that uh citizens just curious citizens
will use government information. So
genealogologists
uh I think are one category of people
that we know use government records. Um
also property owners may want to use
government information to find out
information about their homes about uh
the um uh the uh the the property laws
from when their homes were remodeled or
redeveloped. So there's just so many
uses for government information but
would really love to learn more. Um in
this project I'm going to talk about two
specific efforts. So I said that uh
democracy's library really focuses on um
on government information in this kind
of broad and audacious sense. Um so one
of our uh digitization partners is the
institute of governmental studies at UC
Berkeley. Um they are uh have been a
partner with the internet archive for
some time. They have two of our scribe
machines um and scribe operators that
work uh on site at IGS.
And uh IGS is charged with collecting uh
by the state by state legislative
movement with um collecting California
local government documents. Um so these
are a critical record of policy,
social, economic and cultural history.
um and are a significant primary source
for scholarship and civic engagement. Um
so again I just really want to a lot of
people think about especially in this
country what's going on at the federal
level but this local information is
really important and this is where
people live is locally. So collecting
and preserving this information is
really important. Local government
information is can be some of the most
difficult to find and also the most
vulnerable to loss.
um it is scattered, inconsistently
maintained and not well indexed or
consistently cataloged. Um long-term
preservation and access of this
information is not a core f function of
local government. So at the federal
level uh we have agencies like go and
other agencies who are charged with
preserving their own record. That is not
the case with local governments. So the
focus of local governments is really on
immediate operational needs um not the
historic continuity of their of their uh
of their operations. Um so
this collection that IGS uh holds and
stores is not strictly comprehensive but
does uh include many opportunities for
research into the work of local
governments
um
and uh many many topics and you see here
regional planning, city ser senior
services, city reports, transportation,
all of these things that can be
important at city and regional level.
Um, government documents are really
serialsheavy. So, if you're familiar
with government documents, this will be
a very familiar concept. Uh, so there
are title changes, agency changes,
format changes over time. So, a goal
with the IGS project is to under the
opaces of democracy's library bring all
of these issues together to alleviate
the need of looking in one place. Um
this collocation of information means
that access and research can be
conducted looking at a snapshot in time
or overtime for a single jurisdiction
across multiple cities in a county or in
cities in various states.
Um it's really difficult for local and
policy staff to find out what other
cities may be working on with similar
issues. So this is getting back to my
premise that governments themselves are
primary users of government documents.
Um so having this a information
available will make it easier for people
who are working in cities and who are
developing policy issues to be able to
see what others uh have done. Um
so I also want to talk about democracies
library Canada which is um run out of
internet archive Canada in conjunction
with the internet archive. Um, this is
focused on partnerships and engagement
with Canadian GovDoc communities. Um,
and even though it is led by Internet
Archive and Internet Archive Canada,
like most of our projects, the work is
only possible in when done in
collaboration with partners and others
working in this space. So, this is a
project that started before Democracy's
library was born in 2013. Uh there are
currently um over a 100,000 items which
are discoverable via the Canadian
government publications uh portal on
archive.org which is then pulled into
democracy's library but that number is
an underount and that's for the reasons
that I underscored earlier. There were
materials that were digitized earlier
and so going back and pulling those into
this government information space um is
is really important. So thinking
carefully about how we tag or identify
content that's part of democracy's
library but before this um this content
was uh before this initiative was
launched and also um content that's
uploaded by partners. So this is another
really important point is that material
is coming into the internet archive
that's not controlled by internet
archive itself. Um, so thinking about
those government information librarians
who are actively uploading things, how
do we pull those things into Democracy's
library? Um, Democracy's Library Canada
in the spring of 2023 completed a
three-year project to digitize the
largest government publications
um at the federal and Ontario provincial
level. Um, this has uh resulted in over
400 government publications uh
digitized.
um over 13 million pages. Uh one example
is the Hamilton Public Library, which
has a substantial amount of government
publications mostly from departing
municipal employees who would uh sort of
drop off their office archives at the
public library. I kind of love that. Um
on on your way out of your job, just
kind of drop your things off at the
public library. Um but the great thing
is is that we have an internet archive
scribe on site at the Hamilton Public
Library. and the Hamilton public library
are digitizing their government
documents collections so that they can
be more widely accessible. Uh so again
this is municipal government documents
coming in through the opaces of uh
democracies library Canada. Um other
sources of government information are
the archive it partners as I mentioned
earlier in Canada alone there are over
80 partners in Canada which includes the
gc.ca CA web archive for Library and
Archives Canada. Um they are also
working closely with consortial groups
of librarians and archavists uh
especially the Canadian government
information digital preservation group.
Um and uh Internet Archive Canada is an
active p participant in the twice annual
um Canadian Gov Info days. In fact, I
will be going up to British Columbia in
May to participate in that. Um uh the
Canadian group has really been very
active and is really an inspiration to
me. Um one of the things that they want
to do is to focus on how to make
interacting with government publications
and data sets um more interactive, more
enjoyable and to bring fun to this work.
Uh which resonates with one of the
themes of this meeting. Uh the Canadian
cohort collaborates closely conducting
an environmental scan in conjunction
with surveys and roundts.
Um areas of concern that I think
resonate with with my own work are this
uh attention around municipal government
publications um and also indigenous
collections uh a big a big priority in
Canada. Um so this wish list here uh
that came out of the um of the of the uh
environmental scan and surveys really
summarizes that government information
librarians are looking for access across
digitized and born digital collections
and web archive collections. Um
digitization of bibliographies and
checklists to find out what's been
published and what's missing. This is a
real concern for me in in my own job as
we get donations of materials, but how
do we know that they're they're
comprehensive or that they're whole? We
don't have necessarily inventories to
work for. Um,
so I want to uh wrap up by saying that
democracy's library is not just about um
getting materials and bringing them into
the internet archive as I think our
Canadian colleagues working closely with
the gov info uh professionals has shown
really forming community within this
landscape is tremendously important. Um
and that is why last month the internet
archive hosted um the first information
stewardship forum which brought together
people who are concerned about
government information within the United
States at that that federal, state,
local level and this could be uh
practitioners, data rescue librarians,
um technologists, funders. Uh we had uh
three days in San Francisco.
I'm uh on the verge of publishing a blog
post about that which will be on the
internet archive blog which will serve
as a brief report out. But one of our
key outcomes was this preservation of
government information a call to action.
I have a QR code there. Um I encourage
you to um uh to uh snap a photo of that
and go to the URL. This really serves as
a manifesto for the field and helps to
um establish clearly that public access
to government information is really um
an important thing that we should all be
committed to. Uh I hope that you will
consider signing on to this document and
uh maybe even encouraging your own
institutions to sign on as supporting
organizations. So, it was great to have
those three days together to uh to form
community and be with one another, but
but also um to have this at least as a
as one tangible outcome from the three
days together. Um I want to thank my
collaborators who helped me work on this
slides and this presentation and also to
do the work and I will turn things over
to Christie. Thank you.
Okay, great. Um, so first off, it is an
absolute pleasure to be here and I think
it's even more of an honor to be able to
be here and talk about this essential
and critical topic that I think has
really come into um an awareness. uh,
you know, preservation and access has
always been something that I think this
community cares about, but I think we're
seeing these conversations percolate
into new audiences and we're bringing
people along for the ride, which is
really exciting and something that I'm
particularly thrilled about. So, um, my
name is Christy Holmes. I'm based at
Northwestern University at Fineberg
School of Medicine, um, on the Chicago
campus. Um, and I am here on behalf of a
project through the Center for Open
Science. So, I'm really delighted to
have the opportunity to share that with
you today. Um, and look forward to um,
uh, a little bit of a a deep dive into
that. So, you may ask yourself, why am I
interested in this effort? Well, like I
said, I'm based at Northwestern
University at the School of Medicine.
I've got a couple of different roles
that I think align nicely with caring
about this type of topic. So, first off,
I'm the director of the health sciences
library. Um, I also direct informatics
and data science for our translational
sciences institute on campus and then I
have a research program of my own. So,
um, in this image you can see a snapshot
of part of our campus. So, I'm based in
the building that is to the left um with
the kind of tall tower. Um and the heart
down below is actually where Galter
Library is located. So, that's my home
base. But also on campus, you can see
Northwestern Medicine, which is our
adult primary care facility. Right next
to it is Lurie Children's Hospital. Um
behind those hospitals, our apprentice
women's hospital. We have the
comprehensive cancer center. We have
Shirley Ryan Ability Lab, which is a
rehabilitation institute that is
incredible. So there um are a number of
different facilities on the Chicago
campus that are really thinking
carefully about driving health. Beyond
this image is our beautiful community of
Chicago. So um uh all of these clinical
partners are focused on the mission. So
making those discoveries and improving
the health of our community which is all
driven by data work. But data is also
helpful for supporting meaningful
partnerships and engagement with um the
people of Chicago and more broadly. Um
not only that data is important in order
to support accountability um in terms of
people who are engaging in a meaningful
way to taxpayers and so on. So, we
really want to be able to have ways to
catalyze meaningful conversations more
broadly. Um, and being able to help
people connect in a meaningful way to
information and knowledge that helps to
serve um their needs at any given time.
Uh, so I also mentioned that I have some
research activities in this area. So
just for context, um I and my team have
been contributing to the Invino RDM
open-source community for many years
now. So Invino RDM is the software that
supports Zenotto and several other
repositories around the world. So we're
really thrilled to have that um locally.
I also serve as the PI for Zenotto on
the generalist repository ecosystem
initiative which is also hyperfocused on
the ecosystem to support NIH funded
data. So we're really excited and
energized about what opportunities are
here and and tuned into um some of the
challenges that can uh come into play
when we are not being proactively
attentive to data. So um on data resil
resiliency, we've seen data resiliency
play out in a very real way this past
year as we see federal data sets
disappear. Um and it this has all been
happening in a really highprofile
manner. Um and then in turn we've seen
an incredibly inspiring response from
the broader community including through
the data rescue project and then also
the internet archive among others to
preserve federal collections of data. So
this is all really critical and really
communitydriven work.
There are also other several nonfunding
vulnerabilities that the um that are
happening in the research ecosystem
right now. So these kinds of
vulnerabilities are also risking uh or
threatening long-term access, usability,
and trust. And so I've outlined a few of
those on this slide here. So we have um
technical vulnerabilities. So where the
data exists but there are still
failures. So single points of failure
with respect to platforms, hosts,
services, we have obsolete formats or
software or other kinds of dependencies
that can't be supported in the long
term. Uh we all know um in a very
painful way about how metadata
identifiers and documentation can decay
over time. So that continues to be um a
threat um you know even under the best
of circumstances. And then I think it's
important to point out that data is
frequently preserved without the tools
or the context to even be able to use
it. So even if you find it, even if you
can access it, that doesn't necessarily
mean you can use it. And so those are
important things to think about as well.
Um I am a big fan of thinking about on
technical projects that space between
the technical and the cultural or the
social and I think that that really also
opens up a wide range of vulnerabilities
to projects that we're seeing um you
know and have been seeing for a long
time. So first of all um data
stewardship is undervalued and
underrewarded.
Um it relies often on volunteer effort.
So we all know what that looks like in
our organizations and I that continues
to persist today and because of that it
frequently suffers from a loss of
personnel as we see people move to
different responsibilities or take new
jobs. It's like okay who's going to do
this thing with the repository, right?
So that's something we're all thinking
about. And then you know just to point
out that there's often inequitable
attention to different types of data
sets. So um different domains are maybe
not getting the kind of TLC that they
deserve. And then finally just to think
a little bit about governance and policy
vulnerabilities. Um you know
responsibility is really difficult to be
able to clearly articulate across
stakeholders. So there's often unclear
accountability for data ownership and
stewardship. It's like is that whose job
is this? Is this my job? Is this your
job? and so it doesn't get done right.
Um likewise there's often misalignments
with respect to policy infrastructure
and then actual practice. So you know
taking the theoretical and putting it
into play in a very real way that can
result in uh meaningful stewardship and
change um in the organization. Um
there's always privacy considerations
and legal uncertainties. Um and what
that can do is that can actually
precipitate with withdrawal of data. So
people get nervous about making data
available and so it just seems like to
mitigate that risk um folks can step
back from that. And then finally, I
think we just have um you know,
fragmented standards, poor
cross-institution coordination, you
know, just that lack of consistent and
um uh intentional approach can actually
just uh challenge the ecosystem in ways
that are very difficult to manage and
repair. So we come to this central
question, how do we enable a more
resilient research data ecosystem that
is resilient to single points of
failure? Um and we want to do this so
that we can rely on data to be preserved
in the long term. The second part of
this central question is how do we
coordinate a distributed system to
address this problem together. Right? So
this isn't something that just a single
person or a single organization can
address. um it is impossible for a a
single organization to be able to uh
carry out the work that's necessary to
save data to create that resilient uh
data ecosystem. We absolutely need to
work together.
So um on that point of working together,
I'm very um uh honored to be part of a
group that's working together on this.
We have a strategic planning committee
um uh that was brought together by the
center for open science upon receiving
funding from the Robert Robert Wood
Johnson Foundation to uh develop a
communitydriven communityowned strategic
plan to start to answer these questions.
And we're trying to chart a path forward
together, not just together here, but
together here. We really want to think
about how we do this in the community.
So um our team has spent a few months
doing exploratory work in this space to
understand the ecosystem opportunities,
challenges, champions and other efforts.
And then we all gathered a few weeks ago
in DC to really build out a strategic
plan and we realized um and refined a
case for action um which we're happy to
share with you. So that's outlined here.
So just to highlight these really
quickly. So first of all, fally funded
research data is a non rivalous public
good. The use of that public good relies
on robust and resilient data
repositories and related
infrastructures.
Those infrastructures have long been at
risk for some of those uh reasons that I
mentioned earlier, but those have those
risks have been laid bare in this
current moment. Um, government support
is essential, but that support can be
made more efficient through coordination
across sectors. And then finally, and I
think this might be my favorite point,
we can't go back. The way ahead is bold.
And so, we're really interested to know
um throughout all of these slides, if
there are things that you see that
resonate particularly strongly with your
own perspective or if there are things
you think we should be considering
differently, we welcome that
information. And so I will be um showing
some contact information uh soon.
So um all this work builds on the vision
um that we're working to uh achieve
through a strategic planning process and
then we're planning on enacting that
strategic planning process um over the
next um months and years to come as we
begin to think about this in a
sustainable way. So our vision is the
strategic plan seeks to identify the
conditions and advance the structure for
a coalition that will work
collaboratively to cultivate a resilient
data ecosystem for publicly funded
research. So here we mean uh from
conditions we're thinking about our
recommendations and strategies and
structure. We're really thinking about
um our favorite topic governance models.
So um we identified several guiding
principles to support this vision. First
of all, federally funded data, oh, I
think uh is a non-rival risk public
good. I think I already talked about
this. I'm sorry. Federally funded uh
data is a non-rival risk public good.
It's paid for by the people with a
societal purpose. Data has value. So
maximizing the use of the data through a
robust repository ecosystem provides
significant return on investment. Uh
also maintenance is mandatory. So
repository systems will fail if they're
not cared for and I think um myself and
many of us in this room know exactly
what that means. Um shared governance
and coordination is critical. Um a
cross- sector collaboration is
foundational. So thinking about how we
bring together other stakeholders and
then finally complimentarity emerges. So
future efforts should build on and
connect existing efforts, maximize
efficiency and make the most of limited
resources. So this plan is actually for
three different groups. The way we see
it, the doers, the supporters and the
beneficiaries. The doers are data
advocacy leaders, data repository and
infrastructure uh providers. The
supporters are funders and policy
makers. And the beneficiaries are all of
those folks who depend on a more
resilient and robust ecosystems like
researchers and data and evidence users.
So um when we think about um developing
this plan, we have some key activities
that we're involved with. So there's
three different areas of activities. Uh
the first of these is under the heading
of assess, monitor and track the
repository landscape where we're looking
at monitoring and inventory at risk
respon uh repositories
uh developing a data repository maturity
model so that we can look at risk risk
resilience and fairness um over time in
a level of maturity um and then creating
a shared repository health tool for
continued monitoring. We're laying
strong a resilient and coordinated
foundation. So looking at what are those
core foundational elements that
characterize a resilient um data
ecosystem
uh charting a communityinformed roadmap
to a cons uh for a consortium to
coordinate on those key elements and
then developing attributes of business
models to help to sustainably support
the ecosystem. And then finally, um, we,
you know, as I've mentioned several
times, working together, working across
f, um, different types of folks or
different roles is really important to
do this in a communitydriven way. Um, so
we are developing a shared outreach and
advocacy strategy. So looking at
developing a framework for ongoing
identification of and engagement with
related initiatives like internet
archive um developing public
communications and advocacy toolkits. So
how do I make the case um to to the
people in my space which I'm really
excited about this and then building
community capacity more broadly. So with
that, I will thank you for your time and
then we have a um we have a QR code
where we're actually seeking input. Um
you know, we'd love to hear what you're
thinking about and then I'm more than
happy to um provide additional
information later. But um thanks for
your time.
Hi everybody. Thrilled to be here. Um
I'm Trevor Owens. I'm the chief research
officer at AIP and I'm going to talk a
little bit about a few things that we're
working on uh related areas to
documenting this this sort of disruptive
moment we're in and trying to help
connect the scientific community with
ways to uh uh understand and respond to
the challenges that we're facing. So for
folks who are unfamiliar, AIP is uh
organized and focused on advancing,
promoting and serving the physical
sciences for the benefit of humanity. We
are both a federation of scientific
societies, science and engineering
societies in the physical sciences and a
research institute that's focused on uh
history, policy, culture, um
demographics, workforce aspects of the
physical sciences enterprise. And um
that's really broadly construed. The 30
uh societies that are part of our
organization include everything from
physicists to um biomedical engineers to
acousticians, acousticians, which is the
word for people that work in acoustics.
Um optical engineers, chemical
engineers, it's a big tent that we're a
part of. And so my job as the chief
research officer is to run the institute
part of what we do.
And so I'll talk a little bit about what
our research team is here and then I'll
move into some specific examples of some
of the work that we've been doing and uh
hopefully offer some some useful
information to you all and invite you to
follow along with some of the research
that we're engaged in. So our research
team brings together decades of our
capabilities to empower positive change
in the physical sciences. And so
specifically here we are um about two
dozen folks, a mixture of um social
scientists, sociologists, uh historians,
policy analysts, librarians, and
archavists. Um, and within this and and
relevant to our context here, our
Neilsbore Library and Archives, which is
the box on the front of our building
right here is a place that's been around
since 1962 and is really focused on uh
being a catalyst and a a resource for
helping build community and engagement,
doing some collecting and preservation
ourselves, but also helping connect the
scientific community with the uh with
various archival communities around the
country. And so, uh, for some context on
where we come from and how we approached
this, this is Robert Oppenheimer talking
at the opening of our library in 1962.
And I think his quote here is is a sort
of powerful thing to return to. Just if
you think about 1962,
um, at that point the library was in New
York City. Um, who Robert Oppenheimer is
and was. This idea that we're engulfed
by changes, the massiveness, the
ferocity, the brashness, um, that we do
not understand very well. and has hoped
that this library and I think all of our
libraries work to be places that can
enable serious students of the human
predicament in the future to know very
much more about what has befallen us
than we who are acting and living in it.
So, um I think that I I feel like I can
relate to that, right? Every day is some
new adventure. Um and so part of the way
that we do this the libraries
collections in this case we have over
2,000 oral histories. actually just
provided open access to three different
oral histories with Oppenheimer that had
um previously required us to get
permission from his uh descendants every
time a researcher wanted access to them.
So we have this large collection of oral
histories. We also have a uh an amazing
collection of photographs most of which
are digitized and online. The Amelio
Sigra Visual Archives so named because
Sigra was both an atomic scientist and a
Nobel laureate and also an amateur
photographer. So if you see candid shots
from the Manhattan project, they are his
photos in our collections and they're
widely used. And then similarly um the
sort of the cornerstone of our
collection after a collection of books
which is often um books from donated
from scientists in our community is uh a
collection of more than 3,000 linear
feet of archival material. We have some
papers from scientists, but the the real
gems in our collection and the place
where people come to study change and
and upheaval in the in the sciences is
records from scientific societies in our
federation. And so we have records going
back to the 1890s for the astronomical
society and the physical society. Um but
much more recent things, the development
of medical physics, etc.
So, one of the the neatest parts of our
collection and a thing we've been trying
to do more of, which responds to this
particular moment we're in, is we
collect unpublished memoirs from
scientists. And so, this is actually a
piece I wrote for radiations. It's a
magazine that um AIP publishes, which
goes out to um the tens of thousands of
members of the physics and astronomy
honors society, Sigma Pi Sigma, inviting
people to send us their memoirs. And so,
up in the corner is one of my favorite
examples from this collection. It is a
uh uh Shakespeare got it wrong, the
memoirs of a lucky geoysicist. It is
otherwise an unpublished work that we
added to our collection and by putting
these outreach calls out, we've gotten
uh more and more of these memoirs which
we can make openly available online. Um
so we run an annual research agenda uh
where we we gather input and topic from
our community and then publish it and
work through those projects over the
course of a year. These are the topics
we're working on in 2026. Many of them
will will resonate. And I think a thing
that distressed with this is this is
what our community the the physical
science community is really concerned
about. Impacts of funding cuts. Um
really understanding the history of
efforts to broaden participation in the
physical sciences. Um enduring access to
records from societies and we're doing a
lot around visa and immigration policy
as well because that's a huge area. So
I'll talk now briefly about a couple of
our studies and then I'll keep my talk
short. So we still have plenty of time
for discussion. So last year about this
time we came out with a report called
impacts of restrictions on federal grant
funding on physics and astronomy
graduate programs. We've been surveying
physics and astronomy department chairs
since the 1960s again. And so we have a
really great rapport. We can get
responses very quickly. Um, these are
some of the the sort of individual
comments we got from scientists at that
moment talking about the demoralizing
effect of what was going on, the climate
of uncertainty. These sorts of things
are are really powerful. We were able to
put them out and they were actually um
referenced in uh congressional um
sessions in the science committee um
specifically talking about the results
of this. We were at that point
projecting as much as a 13% decline. it
ended up being more on the level of
about um seven to n% in graduate
enrollments in physics which is a huge
drop. Another example of something that
we're doing in keeping with that memoir
collecting tradition, we have an open
call out now where scientists can or
anyone involved in the scientific
community can share their story and we
will add that story to our archives. And
so um this is up and if you know anyone
in the physical sciences community sort
of broadly described feel free to invite
them uh to submit. As I mentioned the
Amelia Sig visual archives are our sort
of origins of our archival photo
collection. uh similarly relates to a
project we now have engaged in with
support from the LE foundation which is
focused on uh encouraging women working
in the physical sciences to share their
document their work and stories and add
those to our collections and um for that
we one of the things we do at AIP is we
publish a magazine called Physics Today.
It goes out to more than 100,000 people
within our community. We've had ads like
this inviting people to submit their
photos and stories and then we add those
to our collections. And as another
example I'll share which relates to
specifically to the question of federal
material is um here we have an example
of some of the the photos from our
repository we recently added. These ones
are actually things we pulled out of
federal agencies Flickr accounts which
are um you might not think of them as
being at risk but is worth underscoring
that they are. they can get turned off
at any moment and there's some amazing
material there documenting the people
and uh work of the physical sciences. So
I'll leave you here with uh um um sort
of action you can take. Um I'm happy to
talk to anybody about any of these
projects or if uh as I was stressing one
of the things that we do is help connect
people within the physical sciences
community, the scientists, the people in
the societies with people in the library
and archives communities as well. And so
we're really happy to be bridges and
help figure out um how we can sort of
protect and preserve materials. And then
also we have a weekly monthly history
newsletter that has really engaging
articles intended for broad audiences.
And every month we publish a research
updates uh newsletter that has a rundown
of all the social science and sort of
policy related uh work we publish as
well. So I will stop there and then I
think we should have a good bit of time
for questions.
>> [applause]
>> So yeah, it looks like we have about 18
minutes before lunch. Uh so I will
invite you to approach the microphones,
introduce yourselves. Uh and I believe
the session's being recorded, so that's
why the mic's
>> Go ahead. Sorry, I I wrote down my
question, so I was trying to cool it up.
Um my name is Rosalyn Mets. I'm at Emory
University in Atlanta. Um I have a
question for you, Merrily. Um so, um
access to local government information
has been a personal
story of mine. I've been taking a a
class called Decator 101, which is the
city I live in. And one of the things
I've noticed is there's a lot of
misinformation happening within my local
government because none of the records
are online.
So, um I I should also say I'm married
to somebody who works for the school
district, which is where all the
contention is. Um, so do you know if
other states besides California are
working to digitize local records or do
you see plans for that in the future for
how um other states can go about um
creating programs kind of like the one
you described in California?
>> Yeah, thanks thanks for the question. Um
I don't the short answer is I I don't
know. Uh I'm fine with my position. I'm
on a journey of discovery. Um, one of
the things that I have been able to do
is to connect with the um, council of
state archavists. Uh, and I'm hoping
that through KOSA I can learn more about
um, how uh, you know, those are state
archavists. They're responsible for the
state, but they're going to have more
information and context around what the
um uh what the local records
storage and management is. So, I would
love to scale out. Um you know, I've
even only learned more about the
California situation even though being a
lifelong California resident and a
librarian, I was very ignorant of this
of this area until just recently. Um, so
yeah, would love to learn more. Um, and
you could also find out and relay more
information to me about Georgia. Would
love that. Let's go over here.
>> Uh, Nathan Tolman, AP Trust. Um, quick
comment before my question from the last
one. When you're looking, um, there's
big differences around the country
whether power is concentrated at the
county or local level. And so that might
be a nuance to deal with when you're
trying to get those records. My question
is for you Marily as well. Um there
seems to be a variety of methods in our
archive is using to collect government
archives of all sorts including the
upload option. I'm wondering how are how
is the authenticity of upload documents
validated because it seems like that
could be a backdoor to inject false
narratives into the government record to
perpetuate other bad actors.
That's a great question and I don't know
the answer to it and I will have to um I
will have to get back to you about it. I
think that um
an answer is that we don't validate
information that's uploaded. It just
simply is recorded as information that
has been uploaded. Um so I think that uh
you know people need to do their own
provenence
work. Uh so that is Yeah.
>> Could helping them do that be part of
the project?
>> Yeah. I I mean yes. Uh so one of the
things that I just learned about last
week is we have a government information
librarian who's been um furiously
uploading congressional research service
reports um to the extent that we had
identified her as a bot which she most
certainly is not. um you know so putting
putting better tools into her hands so
that she can do that more effectively um
and uh yeah kind of validating
uh you know our super users in that
respect government information
librarians are incredible um and very
concerned with you know ensuring the
durability of the record but then also
um I think you you raised an excellent
point one that I had not considered and
I should have.
>> Well uh thank you guys. It was uh great
to hear about all these initiatives. Um
I'm Nick. I'm with Nache. We're a
service provider. Um we do a lot of work
on uh repository systems and uh my
questions I think mostly for Christie.
Um you were talking about data fragility
um and some of the risks that you
encounter with that. And um I'm trying
to I was furiously doing research and I
figured I'd just ask you straight up.
Um, do you think that Inveno is going to
be shifting more toward OCFL for its um,
uh, its persistence layer? If you know,
I'm not sure. Uh, it sounds like you're
very active in that community and
you're, uh, you have your finger on the
pulse of things and do you see that as a
good pathway to bring more resiliency to
data?
>> Um, thanks for your question. So I will
um take the title of chief cheerleader
for the Invino RDM community and so I'm
really fortunate to be surrounded by a
strong team that is contributing on you
know technical and data um activities. I
will say that if if you're interested
I'll be happy to get your information or
even Martin may know who's sitting in
the front row may have more experience
with that but we can connect you with
the information. One thing I will say
though, I mean, so I've been part of a
number of open source communities over
the years, and I am I I continue to be
astounded by this particular open-source
community because it is incredibly
thoughtful with respect to leveraging uh
meaningful and practical standards. Um
really trying to push the limits when
we're looking at features that help to
drive better behaviors. And then um the
other thing is that there's just you
know again kind of um reflecting back on
that human element there is uh just
enough um celebration and appreciation
for the different contributors in the
program to you know to be able to um
have that I guess psychological safety
to feel trust and be able to move
forward together. So yeah but I I'd be
happy to connect with you and then I
will get you that information. That
sounds great and that sounds like a
great community to be a part of. So
that's that's awesome. And one really
quick question. My friend David Schober
used to be with Northwestern
>> and did some great work on semantic
search and then you moved to Internet
Archive. I just want to make sure
there's no bad blood there.
>> I'm just kidding.
>> I think so far we're good.
>> Yeah. Yeah. Yeah. No, David is amazing.
And I think, you know, that's the nice
thing about being in this community or
this bigger community is that there is
that chance to I see people who are
doing great work. I mean, and being
genuinely thrilled when they have an
opportunity.
>> Obviously, a huge joke. Um, [laughter] I
think you're all doing important work
and thank you guys. Appreciate it.
>> Yep.
>> Mark Calfadvic with the triple AF
consortium. My question is a little bit
of a follow on to Rosyn's question. And
I live in Arlington, Virginia, which has
a very open, transparent government,
large numbers of digitized documents.
But what I found recently is they're
getting pushed behind um challenge
walls. Um again, probably primarily for
protection from bots. And I was just
wondering, is there any evidence that
there's increased blockage of harvesting
of go local government specifically
sites to um
Yeah, I I don't know that we've done um
an analysis of how much of um
of uh the government space is being
blocked, but I think that in response to
very aggressive bots, um everybody is
tightening things down or refusing bots
and that's problematic for us because of
course we are a bot on the web. We are
also busy defending off bots. Um so uh
so yeah, so finding avenues and pathways
for um this is something I think that
our community really needs to negotiate
is is how do we um how do we come up
with
uh thoughtful negotiated
um avenues for ensuring that we are
continuing to collect what we need to
collect when it's web- based materials.
Um,
and I don't I don't know what the what
the answers to that is. I know that this
is an area where Rosie and others have
been actively working. So, should have a
CNI discussion about it, I think. There
we go.
>> Thank you both. Thank you all for the
the topics you brought forth. I wanted
to talk a little bit about the uh local
records question. Uh, I'm Dennis Clark.
I'm the Librarian of Virginia. The
Library of Virginia is the only state
library archive that's a member of CNI.
And I think this is one of the very good
reasons why that should not be the case
anymore. And I think your connection
with KOSA is exactly the right way to
start talking about this at scale. Local
records are challenging because every
state's public records act defines what
that means for your state. In Virginia,
obviously, it's a very um um broad and
powerful act. Uh and so we can work very
closely with local records, local
circuit courts, local governments, local
school systems to make sure that those
items are kept available in whatever way
various parts of the act uh require, but
that's not certainly the case in every
state. Um and it's it's every state is
different. Uh and I love the notion of
thinking about this at scale with KOSA.
Uh don't forget the the costa group, the
state libraries as well. uh they're
they're very much a part of the same
continuum of of of conversation. Uh but
again, a really good reason why uh CNI
should should continue to to have those
folks as part of this conversation.
Thanks.
>> Yeah, agreed. And thanks for showing up
in that capacity. Um I think that that
is really important. I didn't mention
Kosla because the internet archive has
uh for has strong relationships,
existing relationships with KLA that I'm
able to uh leverage. So the KOSA one is
because because I am a careerlong
archavist, I'm helping to bring that in,
>> which is which is great. And and and
I'll just say that uh very proud that
Library Virginia was, I believe, the
first state library that was actually a
partner with Internet Archive back in
2005.
So, uh, this has been an important part
of of everything that we do and and we
really are thrilled by that level of
collaboration and and want to see it
grow and and and be and be and be be
bigger and better and more transparent.
Thanks.
>> Amazing. Thank you.
>> Well, um, merily that I had a follow on
question to Dennis is for you too. uh
for the most part is what y'all are
focusing on um published material things
that have been publicly available
because I think part of Dennis your
question also gets at records that are
you know organizational records and it's
just to stress that like my sense is as
rough as stuff is with publicly
available material like organizational
electronic records we're trying to focus
on this a bit with scientific societies
but for all kinds of nonprofits or
government entities universities just a
huge
>> yeah and it is it is such a challenge
especially when everything moves digital
differentiating between publications and
records becomes really challenging I
would say and there's a kind of squishy
continuum in there um I don't want to
rule things out but I would say that our
emphasis is on published materials
>> right
>> I am Matt Merick from the National
Center for Atmospheric Research in
Colorado um so this is question is for
Christie I really appreciated your
description of effort that you're doing
there. Um, and my question you kind of
touched at the end about some
coordination among other similar efforts
and that was the thought that I had as
you were talking because I've heard of
other efforts that are doing trying to
do repository coordination. Um, you
know, American geohysical union for
example is is has an effort there. So I
guess I just wanted to get your thoughts
on you know how do you feel like there
is u you know the sort of a common
movement or are things kind of going in
directions that are uncoordinated and do
we have a common voice yet that we can
speak from a repository point of view?
Yeah, thank you. That's a great question
actually because um you know I think
anytime you get a lot of passionate
people around uh a good goal um you know
it can be hard to channel that passion
in in useful ways. One thing I will note
is that um first off just speaking about
this center for open science project
that um I've been working on is that we
have very um deliberately and
intentionally worked to look to see you
know like what exists out there what
kinds of um frameworks or initiatives or
those types of things so that it's not
this oh this is important let's start a
project kind of thing and so you can see
that with a number of the people that
are actually on that program. The other
thing that I'll note is that um so AGU
and then the internet archives um
meeting in March had a number of folks
who are also participating in this
group. So there's been quite a lot of
cross fertilization and a real eagerness
um to try and be coordinated as much as
possible. um we're trying to move
forward um in a way that does reflect
it's not at the data set level, it's at
the repository level. So what does
resilience mean in those systems? And
just trying to identify some of that so
that we can have I mean it's just like
any conversation if you can have a
shared language it helps you be able to
collaborate more efficiently. Um so
that's been good as well. But there is,
you know, and I'm happy to um to get
folks the um the information, but
there's an email and then that QR code.
We not only do we want to hear from you,
there's work to do, friends, you know,
so um you know, if there are things that
make sense, like right now I'm working
on that maturity model. So that's a a
very um traditional
project kind of in informatics land um
especially biomedical informatics. And
what it does is it allows you to be able
to look at levels of maturity from kind
of ad hoc processes that are chaotic all
the way to fully implemented and
continuously improving. So that spectrum
and then you can look at it along some
different attributes. So you can look at
like culture or technology or standards
or those types of things. It's not meant
to create um a framework for assessing
yourself in the context of other
organizations or repositories. What it's
meant to do is to catalyze local
conversations. So here's an area where
we can um realize some improvements and
you know and to identify oh but here's
something we're doing really well. So
that's a good example of a project that
is actively underway if folks are
interested. We I you know we're very
collaborative friendly bunch. So we're
um always eager to uh to have folks get
involved. Yeah. Thank you. Well,
so I don't see anybody's Are you going
for the blankets? Okay. I don't see
anybody standing uh up for the
microphone. So, um, I'm going to bring
the session to conclusion and let's go
enjoy some lunch. Right. Thank you.
[applause]