Submind YouTube summaries
Thumbnail for Sustaining the Public Record: Collaborative Stewardship of Government Information and Research Data

Sustaining the Public Record: Collaborative Stewardship of Government Information and Research Data

Watch on YouTube

Video summary

The initiative known as Democracy's Library was established in 2022 by the Internet Archive to address critical gaps in preserving government information that has become difficult for the public to access due to inconsistent maintenance and lack of comprehensive cataloging. Leveraging a vast network including its Wayback Machine, digitization partnerships with institutions like UC Berkeley's Institute for Governmental Studies, and community uploads, the project currently houses over one million items across more than 1,000 collections. A primary focus is on safeguarding local government documents, which are often scattered and not central to municipal operations, while similar efforts in Canada have digitized millions of pages through collaborations with public libraries like Hamilton Public Library. These archives serve diverse stakeholders: governments utilize them for policy research, journalists depend on historical records to track legislative changes, and citizens access materials for genealogy and property law matters. Beyond the Internet Archive's specific projects, experts from various institutions emphasized the broader need for data resilience within scientific and governmental ecosystems against technical failures, cultural vulnerabilities like over-reliance on volunteers, and governance issues regarding accountability. Christy Holmes outlined a strategic plan involving repositories, funders, and researchers to create a resilient coalition using shared monitoring tools and advocacy strategies, while Trevor Owens highlighted how organizations like the American Institute of Physics preserve scientific history through oral histories and photographs despite challenges such as funding cuts. The community also recognized that uploaded documents are not always validated for authenticity to prevent bias but instead rely on user-provenance work supported by trusted librarians acting as "super users" who upload critical reports from sources like the Congressional Research Service, ensuring a balance between accessibility and integrity. Significant challenges regarding access were raised during discussions about increasing blockage of government websites by bot protection measures that inadvertently hinder legitimate archival harvesting alongside defending against malicious actors. Addressing these obstacles, advocates called for negotiated pathways to continue collecting web-based materials while emphasizing the importance of including state libraries in conversations about local records management despite varying public records acts across different states. Solutions such as scalable frameworks like KOSA and existing partnerships with COSTA groups are being explored to create unified approaches that transcend individual dataset coordination, aiming instead to foster cross-fertilization between initiatives from organizations like AGU and the Internet Archive to avoid redundancy. The overarching conclusion of these discussions points toward a future defined by collaborative stewardship rather than isolated efforts, utilizing shared maturity models designed not to rank organizations but to catalyze local conversations about improving culture, technology, and standards across different attributes of resilience. By forming communities committed to public access through manifestos and ongoing dialogue among "doers," "supporters," and "beneficiaries," the sector aims to build a robust infrastructure that can withstand technical obsolescence and shifting political landscapes. Ultimately, the goal is to ensure that government information remains a non-rivalrous public good available for generations, requiring continuous involvement from researchers, librarians, policymakers, and citizens who collectively uphold the values of transparency and historical preservation.
Read the full video transcript
Okay. Hello everybody. Thanks for coming to our session on collaborative stewardship of government information and research data. Um I'm going to go ahead and kick off our session. Uh my name is Marily Profett. I am the director of democracies library US which is a large-scale effort to represent um uh US government information that could be federal, state, local, even tribal uh within the internet archive. And I am new. I have been in my position for not even five months. Um so I'll just start by saying my my concluding remarks are that I am really eager to collaborate with you around your government information collections and fill gaps that we may have and also to learn how government information is used and important in your context and for your users. Um I have my business cards out here on the stage uh along with internet archive and wayback machine stickers. So please do come up. Last time I was at CNI, I ran out of um wayback machine stickers, so I have a lot this time. Um please come up and help yourself at the end of the session. So, first I just want to um center by saying a few words about the Internet Archive, which amazingly will be 30 years old next month. Um so that's kind of fun. Um so our mission statement is universal access to all knowledge. And this is really more than just a mission statement. that truly guides everything that we do as an organization. Um, and we have all types of organization represented within the internet archive. For example, uh, software, moving images, audio recordings, television news programs, um, including we, uh, record state uh, state television from Iran and have for quite some time. uh ebooks and web pages. Uh last October we celebrated uh having uh archived one trillion web pages uh over the the 30 years of duration of the internet archive. Um and all of that knowledge is uh contained within 210 plus pabytes of data which is hosted on our pedibyte servers. So I want to tell you a little bit about democracy's library. This is not a new effort of the internet archive. It was actually established in 2022 and is built on a straightforward but urgent premise which is that governments have created an abundance of information and put it in the public domain but the public can't easily access it. So, we currently have over a thousand collections with over 11 million items, which is likely a significant undercount. And I'll go into the reasons for for why that is. Um so the internet archive has really been at the um in the business of uh backing up um public documents uh uh public information, government information for quite some time. So I will start with the wayback machine um which is an important uh part of the the work that we do. Um the wayback machine is uh in the internet archive are part of the end of term crawl uh in the United States and have been for a long time. We also routinely crawl government information websites as part of our ongoing operations. Also archive it. So for any of you uh who are archive it uh subscribers will know about this subscription service. We have many government sites that are doing their own due diligence in uh backing up their information, state, local, uh and even national um organizations that are using uh our our data services to back up um to back up their their sites. So this this includes a a lot of those uh materials that are within the government implicate u government information space. We also have been in the business of digitizing the published public record. Um so we do this with uh with partners who donate materials to us and al also through digitization partnerships with libraries and a lot of those uh efforts have included government information over a long period of time and part of why democracy's libraries underounted is because we haven't been able to go back and identify all the government publications that are within the internet archive that have been included before the establishment. of democracy's library. So we have a big job of going through and identifying government publications and pulling those into um the internet archive. Uh some of the collections that we're working on right now are Supreme Court records and briefs. When that project is finished, which it will be very shortly by the end of this month, I think uh we will have the largest collection of Supreme Court records and briefs that are publicly available. So that's really exciting. Uh we've also been working with the National Oceanic and Atmospheric Administration to digitize their library collections. Um and uh so this is just our effort to take physical media and convert it to searchable digital formats. And then community uploads. This is also a really important uh part of our work is uh government librarians and others curious citizens people who are concerned about in uh government information will either do uh save now on the way back machine to save and capture pages right in the moment or also to upload uh things like reports directly to the internet archive. And this is keeping uh with the theory if you see something save something um or lots of copies keep stuff safe. So who uses government information? So first of all, we know this is an area where I would really love to learn a lot more. So and learn with you and learn from you. Uh but we do know that governments themselves are primary users of government information. Um we know that journalists and researchers are very interested in this information. Government reports of course are really important for researchers within our public institution or within our um academic institutions um but also journalists. Journalists are heavy users of the wayback machine and uh really within the current um uh administration, the US uh administration have been really using relying on the way back machine heavily to look and see where uh things have changed uh within within the government web spaces. Um we also know that uh citizens just curious citizens will use government information. So genealogologists uh I think are one category of people that we know use government records. Um also property owners may want to use government information to find out information about their homes about uh the um uh the uh the the property laws from when their homes were remodeled or redeveloped. So there's just so many uses for government information but would really love to learn more. Um in this project I'm going to talk about two specific efforts. So I said that uh democracy's library really focuses on um on government information in this kind of broad and audacious sense. Um so one of our uh digitization partners is the institute of governmental studies at UC Berkeley. Um they are uh have been a partner with the internet archive for some time. They have two of our scribe machines um and scribe operators that work uh on site at IGS. And uh IGS is charged with collecting uh by the state by state legislative movement with um collecting California local government documents. Um so these are a critical record of policy, social, economic and cultural history. um and are a significant primary source for scholarship and civic engagement. Um so again I just really want to a lot of people think about especially in this country what's going on at the federal level but this local information is really important and this is where people live is locally. So collecting and preserving this information is really important. Local government information is can be some of the most difficult to find and also the most vulnerable to loss. um it is scattered, inconsistently maintained and not well indexed or consistently cataloged. Um long-term preservation and access of this information is not a core f function of local government. So at the federal level uh we have agencies like go and other agencies who are charged with preserving their own record. That is not the case with local governments. So the focus of local governments is really on immediate operational needs um not the historic continuity of their of their uh of their operations. Um so this collection that IGS uh holds and stores is not strictly comprehensive but does uh include many opportunities for research into the work of local governments um and uh many many topics and you see here regional planning, city ser senior services, city reports, transportation, all of these things that can be important at city and regional level. Um, government documents are really serialsheavy. So, if you're familiar with government documents, this will be a very familiar concept. Uh, so there are title changes, agency changes, format changes over time. So, a goal with the IGS project is to under the opaces of democracy's library bring all of these issues together to alleviate the need of looking in one place. Um this collocation of information means that access and research can be conducted looking at a snapshot in time or overtime for a single jurisdiction across multiple cities in a county or in cities in various states. Um it's really difficult for local and policy staff to find out what other cities may be working on with similar issues. So this is getting back to my premise that governments themselves are primary users of government documents. Um so having this a information available will make it easier for people who are working in cities and who are developing policy issues to be able to see what others uh have done. Um so I also want to talk about democracies library Canada which is um run out of internet archive Canada in conjunction with the internet archive. Um, this is focused on partnerships and engagement with Canadian GovDoc communities. Um, and even though it is led by Internet Archive and Internet Archive Canada, like most of our projects, the work is only possible in when done in collaboration with partners and others working in this space. So, this is a project that started before Democracy's library was born in 2013. Uh there are currently um over a 100,000 items which are discoverable via the Canadian government publications uh portal on archive.org which is then pulled into democracy's library but that number is an underount and that's for the reasons that I underscored earlier. There were materials that were digitized earlier and so going back and pulling those into this government information space um is is really important. So thinking carefully about how we tag or identify content that's part of democracy's library but before this um this content was uh before this initiative was launched and also um content that's uploaded by partners. So this is another really important point is that material is coming into the internet archive that's not controlled by internet archive itself. Um, so thinking about those government information librarians who are actively uploading things, how do we pull those things into Democracy's library? Um, Democracy's Library Canada in the spring of 2023 completed a three-year project to digitize the largest government publications um at the federal and Ontario provincial level. Um, this has uh resulted in over 400 government publications uh digitized. um over 13 million pages. Uh one example is the Hamilton Public Library, which has a substantial amount of government publications mostly from departing municipal employees who would uh sort of drop off their office archives at the public library. I kind of love that. Um on on your way out of your job, just kind of drop your things off at the public library. Um but the great thing is is that we have an internet archive scribe on site at the Hamilton Public Library. and the Hamilton public library are digitizing their government documents collections so that they can be more widely accessible. Uh so again this is municipal government documents coming in through the opaces of uh democracies library Canada. Um other sources of government information are the archive it partners as I mentioned earlier in Canada alone there are over 80 partners in Canada which includes the gc.ca CA web archive for Library and Archives Canada. Um they are also working closely with consortial groups of librarians and archavists uh especially the Canadian government information digital preservation group. Um and uh Internet Archive Canada is an active p participant in the twice annual um Canadian Gov Info days. In fact, I will be going up to British Columbia in May to participate in that. Um uh the Canadian group has really been very active and is really an inspiration to me. Um one of the things that they want to do is to focus on how to make interacting with government publications and data sets um more interactive, more enjoyable and to bring fun to this work. Uh which resonates with one of the themes of this meeting. Uh the Canadian cohort collaborates closely conducting an environmental scan in conjunction with surveys and roundts. Um areas of concern that I think resonate with with my own work are this uh attention around municipal government publications um and also indigenous collections uh a big a big priority in Canada. Um so this wish list here uh that came out of the um of the of the uh environmental scan and surveys really summarizes that government information librarians are looking for access across digitized and born digital collections and web archive collections. Um digitization of bibliographies and checklists to find out what's been published and what's missing. This is a real concern for me in in my own job as we get donations of materials, but how do we know that they're they're comprehensive or that they're whole? We don't have necessarily inventories to work for. Um, so I want to uh wrap up by saying that democracy's library is not just about um getting materials and bringing them into the internet archive as I think our Canadian colleagues working closely with the gov info uh professionals has shown really forming community within this landscape is tremendously important. Um and that is why last month the internet archive hosted um the first information stewardship forum which brought together people who are concerned about government information within the United States at that that federal, state, local level and this could be uh practitioners, data rescue librarians, um technologists, funders. Uh we had uh three days in San Francisco. I'm uh on the verge of publishing a blog post about that which will be on the internet archive blog which will serve as a brief report out. But one of our key outcomes was this preservation of government information a call to action. I have a QR code there. Um I encourage you to um uh to uh snap a photo of that and go to the URL. This really serves as a manifesto for the field and helps to um establish clearly that public access to government information is really um an important thing that we should all be committed to. Uh I hope that you will consider signing on to this document and uh maybe even encouraging your own institutions to sign on as supporting organizations. So, it was great to have those three days together to uh to form community and be with one another, but but also um to have this at least as a as one tangible outcome from the three days together. Um I want to thank my collaborators who helped me work on this slides and this presentation and also to do the work and I will turn things over to Christie. Thank you. Okay, great. Um, so first off, it is an absolute pleasure to be here and I think it's even more of an honor to be able to be here and talk about this essential and critical topic that I think has really come into um an awareness. uh, you know, preservation and access has always been something that I think this community cares about, but I think we're seeing these conversations percolate into new audiences and we're bringing people along for the ride, which is really exciting and something that I'm particularly thrilled about. So, um, my name is Christy Holmes. I'm based at Northwestern University at Fineberg School of Medicine, um, on the Chicago campus. Um, and I am here on behalf of a project through the Center for Open Science. So, I'm really delighted to have the opportunity to share that with you today. Um, and look forward to um, uh, a little bit of a a deep dive into that. So, you may ask yourself, why am I interested in this effort? Well, like I said, I'm based at Northwestern University at the School of Medicine. I've got a couple of different roles that I think align nicely with caring about this type of topic. So, first off, I'm the director of the health sciences library. Um, I also direct informatics and data science for our translational sciences institute on campus and then I have a research program of my own. So, um, in this image you can see a snapshot of part of our campus. So, I'm based in the building that is to the left um with the kind of tall tower. Um and the heart down below is actually where Galter Library is located. So, that's my home base. But also on campus, you can see Northwestern Medicine, which is our adult primary care facility. Right next to it is Lurie Children's Hospital. Um behind those hospitals, our apprentice women's hospital. We have the comprehensive cancer center. We have Shirley Ryan Ability Lab, which is a rehabilitation institute that is incredible. So there um are a number of different facilities on the Chicago campus that are really thinking carefully about driving health. Beyond this image is our beautiful community of Chicago. So um uh all of these clinical partners are focused on the mission. So making those discoveries and improving the health of our community which is all driven by data work. But data is also helpful for supporting meaningful partnerships and engagement with um the people of Chicago and more broadly. Um not only that data is important in order to support accountability um in terms of people who are engaging in a meaningful way to taxpayers and so on. So, we really want to be able to have ways to catalyze meaningful conversations more broadly. Um, and being able to help people connect in a meaningful way to information and knowledge that helps to serve um their needs at any given time. Uh, so I also mentioned that I have some research activities in this area. So just for context, um I and my team have been contributing to the Invino RDM open-source community for many years now. So Invino RDM is the software that supports Zenotto and several other repositories around the world. So we're really thrilled to have that um locally. I also serve as the PI for Zenotto on the generalist repository ecosystem initiative which is also hyperfocused on the ecosystem to support NIH funded data. So we're really excited and energized about what opportunities are here and and tuned into um some of the challenges that can uh come into play when we are not being proactively attentive to data. So um on data resil resiliency, we've seen data resiliency play out in a very real way this past year as we see federal data sets disappear. Um and it this has all been happening in a really highprofile manner. Um and then in turn we've seen an incredibly inspiring response from the broader community including through the data rescue project and then also the internet archive among others to preserve federal collections of data. So this is all really critical and really communitydriven work. There are also other several nonfunding vulnerabilities that the um that are happening in the research ecosystem right now. So these kinds of vulnerabilities are also risking uh or threatening long-term access, usability, and trust. And so I've outlined a few of those on this slide here. So we have um technical vulnerabilities. So where the data exists but there are still failures. So single points of failure with respect to platforms, hosts, services, we have obsolete formats or software or other kinds of dependencies that can't be supported in the long term. Uh we all know um in a very painful way about how metadata identifiers and documentation can decay over time. So that continues to be um a threat um you know even under the best of circumstances. And then I think it's important to point out that data is frequently preserved without the tools or the context to even be able to use it. So even if you find it, even if you can access it, that doesn't necessarily mean you can use it. And so those are important things to think about as well. Um I am a big fan of thinking about on technical projects that space between the technical and the cultural or the social and I think that that really also opens up a wide range of vulnerabilities to projects that we're seeing um you know and have been seeing for a long time. So first of all um data stewardship is undervalued and underrewarded. Um it relies often on volunteer effort. So we all know what that looks like in our organizations and I that continues to persist today and because of that it frequently suffers from a loss of personnel as we see people move to different responsibilities or take new jobs. It's like okay who's going to do this thing with the repository, right? So that's something we're all thinking about. And then you know just to point out that there's often inequitable attention to different types of data sets. So um different domains are maybe not getting the kind of TLC that they deserve. And then finally just to think a little bit about governance and policy vulnerabilities. Um you know responsibility is really difficult to be able to clearly articulate across stakeholders. So there's often unclear accountability for data ownership and stewardship. It's like is that whose job is this? Is this my job? Is this your job? and so it doesn't get done right. Um likewise there's often misalignments with respect to policy infrastructure and then actual practice. So you know taking the theoretical and putting it into play in a very real way that can result in uh meaningful stewardship and change um in the organization. Um there's always privacy considerations and legal uncertainties. Um and what that can do is that can actually precipitate with withdrawal of data. So people get nervous about making data available and so it just seems like to mitigate that risk um folks can step back from that. And then finally, I think we just have um you know, fragmented standards, poor cross-institution coordination, you know, just that lack of consistent and um uh intentional approach can actually just uh challenge the ecosystem in ways that are very difficult to manage and repair. So we come to this central question, how do we enable a more resilient research data ecosystem that is resilient to single points of failure? Um and we want to do this so that we can rely on data to be preserved in the long term. The second part of this central question is how do we coordinate a distributed system to address this problem together. Right? So this isn't something that just a single person or a single organization can address. um it is impossible for a a single organization to be able to uh carry out the work that's necessary to save data to create that resilient uh data ecosystem. We absolutely need to work together. So um on that point of working together, I'm very um uh honored to be part of a group that's working together on this. We have a strategic planning committee um uh that was brought together by the center for open science upon receiving funding from the Robert Robert Wood Johnson Foundation to uh develop a communitydriven communityowned strategic plan to start to answer these questions. And we're trying to chart a path forward together, not just together here, but together here. We really want to think about how we do this in the community. So um our team has spent a few months doing exploratory work in this space to understand the ecosystem opportunities, challenges, champions and other efforts. And then we all gathered a few weeks ago in DC to really build out a strategic plan and we realized um and refined a case for action um which we're happy to share with you. So that's outlined here. So just to highlight these really quickly. So first of all, fally funded research data is a non rivalous public good. The use of that public good relies on robust and resilient data repositories and related infrastructures. Those infrastructures have long been at risk for some of those uh reasons that I mentioned earlier, but those have those risks have been laid bare in this current moment. Um, government support is essential, but that support can be made more efficient through coordination across sectors. And then finally, and I think this might be my favorite point, we can't go back. The way ahead is bold. And so, we're really interested to know um throughout all of these slides, if there are things that you see that resonate particularly strongly with your own perspective or if there are things you think we should be considering differently, we welcome that information. And so I will be um showing some contact information uh soon. So um all this work builds on the vision um that we're working to uh achieve through a strategic planning process and then we're planning on enacting that strategic planning process um over the next um months and years to come as we begin to think about this in a sustainable way. So our vision is the strategic plan seeks to identify the conditions and advance the structure for a coalition that will work collaboratively to cultivate a resilient data ecosystem for publicly funded research. So here we mean uh from conditions we're thinking about our recommendations and strategies and structure. We're really thinking about um our favorite topic governance models. So um we identified several guiding principles to support this vision. First of all, federally funded data, oh, I think uh is a non-rival risk public good. I think I already talked about this. I'm sorry. Federally funded uh data is a non-rival risk public good. It's paid for by the people with a societal purpose. Data has value. So maximizing the use of the data through a robust repository ecosystem provides significant return on investment. Uh also maintenance is mandatory. So repository systems will fail if they're not cared for and I think um myself and many of us in this room know exactly what that means. Um shared governance and coordination is critical. Um a cross- sector collaboration is foundational. So thinking about how we bring together other stakeholders and then finally complimentarity emerges. So future efforts should build on and connect existing efforts, maximize efficiency and make the most of limited resources. So this plan is actually for three different groups. The way we see it, the doers, the supporters and the beneficiaries. The doers are data advocacy leaders, data repository and infrastructure uh providers. The supporters are funders and policy makers. And the beneficiaries are all of those folks who depend on a more resilient and robust ecosystems like researchers and data and evidence users. So um when we think about um developing this plan, we have some key activities that we're involved with. So there's three different areas of activities. Uh the first of these is under the heading of assess, monitor and track the repository landscape where we're looking at monitoring and inventory at risk respon uh repositories uh developing a data repository maturity model so that we can look at risk risk resilience and fairness um over time in a level of maturity um and then creating a shared repository health tool for continued monitoring. We're laying strong a resilient and coordinated foundation. So looking at what are those core foundational elements that characterize a resilient um data ecosystem uh charting a communityinformed roadmap to a cons uh for a consortium to coordinate on those key elements and then developing attributes of business models to help to sustainably support the ecosystem. And then finally, um, we, you know, as I've mentioned several times, working together, working across f, um, different types of folks or different roles is really important to do this in a communitydriven way. Um, so we are developing a shared outreach and advocacy strategy. So looking at developing a framework for ongoing identification of and engagement with related initiatives like internet archive um developing public communications and advocacy toolkits. So how do I make the case um to to the people in my space which I'm really excited about this and then building community capacity more broadly. So with that, I will thank you for your time and then we have a um we have a QR code where we're actually seeking input. Um you know, we'd love to hear what you're thinking about and then I'm more than happy to um provide additional information later. But um thanks for your time. Hi everybody. Thrilled to be here. Um I'm Trevor Owens. I'm the chief research officer at AIP and I'm going to talk a little bit about a few things that we're working on uh related areas to documenting this this sort of disruptive moment we're in and trying to help connect the scientific community with ways to uh uh understand and respond to the challenges that we're facing. So for folks who are unfamiliar, AIP is uh organized and focused on advancing, promoting and serving the physical sciences for the benefit of humanity. We are both a federation of scientific societies, science and engineering societies in the physical sciences and a research institute that's focused on uh history, policy, culture, um demographics, workforce aspects of the physical sciences enterprise. And um that's really broadly construed. The 30 uh societies that are part of our organization include everything from physicists to um biomedical engineers to acousticians, acousticians, which is the word for people that work in acoustics. Um optical engineers, chemical engineers, it's a big tent that we're a part of. And so my job as the chief research officer is to run the institute part of what we do. And so I'll talk a little bit about what our research team is here and then I'll move into some specific examples of some of the work that we've been doing and uh hopefully offer some some useful information to you all and invite you to follow along with some of the research that we're engaged in. So our research team brings together decades of our capabilities to empower positive change in the physical sciences. And so specifically here we are um about two dozen folks, a mixture of um social scientists, sociologists, uh historians, policy analysts, librarians, and archavists. Um, and within this and and relevant to our context here, our Neilsbore Library and Archives, which is the box on the front of our building right here is a place that's been around since 1962 and is really focused on uh being a catalyst and a a resource for helping build community and engagement, doing some collecting and preservation ourselves, but also helping connect the scientific community with the uh with various archival communities around the country. And so, uh, for some context on where we come from and how we approached this, this is Robert Oppenheimer talking at the opening of our library in 1962. And I think his quote here is is a sort of powerful thing to return to. Just if you think about 1962, um, at that point the library was in New York City. Um, who Robert Oppenheimer is and was. This idea that we're engulfed by changes, the massiveness, the ferocity, the brashness, um, that we do not understand very well. and has hoped that this library and I think all of our libraries work to be places that can enable serious students of the human predicament in the future to know very much more about what has befallen us than we who are acting and living in it. So, um I think that I I feel like I can relate to that, right? Every day is some new adventure. Um and so part of the way that we do this the libraries collections in this case we have over 2,000 oral histories. actually just provided open access to three different oral histories with Oppenheimer that had um previously required us to get permission from his uh descendants every time a researcher wanted access to them. So we have this large collection of oral histories. We also have a uh an amazing collection of photographs most of which are digitized and online. The Amelio Sigra Visual Archives so named because Sigra was both an atomic scientist and a Nobel laureate and also an amateur photographer. So if you see candid shots from the Manhattan project, they are his photos in our collections and they're widely used. And then similarly um the sort of the cornerstone of our collection after a collection of books which is often um books from donated from scientists in our community is uh a collection of more than 3,000 linear feet of archival material. We have some papers from scientists, but the the real gems in our collection and the place where people come to study change and and upheaval in the in the sciences is records from scientific societies in our federation. And so we have records going back to the 1890s for the astronomical society and the physical society. Um but much more recent things, the development of medical physics, etc. So, one of the the neatest parts of our collection and a thing we've been trying to do more of, which responds to this particular moment we're in, is we collect unpublished memoirs from scientists. And so, this is actually a piece I wrote for radiations. It's a magazine that um AIP publishes, which goes out to um the tens of thousands of members of the physics and astronomy honors society, Sigma Pi Sigma, inviting people to send us their memoirs. And so, up in the corner is one of my favorite examples from this collection. It is a uh uh Shakespeare got it wrong, the memoirs of a lucky geoysicist. It is otherwise an unpublished work that we added to our collection and by putting these outreach calls out, we've gotten uh more and more of these memoirs which we can make openly available online. Um so we run an annual research agenda uh where we we gather input and topic from our community and then publish it and work through those projects over the course of a year. These are the topics we're working on in 2026. Many of them will will resonate. And I think a thing that distressed with this is this is what our community the the physical science community is really concerned about. Impacts of funding cuts. Um really understanding the history of efforts to broaden participation in the physical sciences. Um enduring access to records from societies and we're doing a lot around visa and immigration policy as well because that's a huge area. So I'll talk now briefly about a couple of our studies and then I'll keep my talk short. So we still have plenty of time for discussion. So last year about this time we came out with a report called impacts of restrictions on federal grant funding on physics and astronomy graduate programs. We've been surveying physics and astronomy department chairs since the 1960s again. And so we have a really great rapport. We can get responses very quickly. Um, these are some of the the sort of individual comments we got from scientists at that moment talking about the demoralizing effect of what was going on, the climate of uncertainty. These sorts of things are are really powerful. We were able to put them out and they were actually um referenced in uh congressional um sessions in the science committee um specifically talking about the results of this. We were at that point projecting as much as a 13% decline. it ended up being more on the level of about um seven to n% in graduate enrollments in physics which is a huge drop. Another example of something that we're doing in keeping with that memoir collecting tradition, we have an open call out now where scientists can or anyone involved in the scientific community can share their story and we will add that story to our archives. And so um this is up and if you know anyone in the physical sciences community sort of broadly described feel free to invite them uh to submit. As I mentioned the Amelia Sig visual archives are our sort of origins of our archival photo collection. uh similarly relates to a project we now have engaged in with support from the LE foundation which is focused on uh encouraging women working in the physical sciences to share their document their work and stories and add those to our collections and um for that we one of the things we do at AIP is we publish a magazine called Physics Today. It goes out to more than 100,000 people within our community. We've had ads like this inviting people to submit their photos and stories and then we add those to our collections. And as another example I'll share which relates to specifically to the question of federal material is um here we have an example of some of the the photos from our repository we recently added. These ones are actually things we pulled out of federal agencies Flickr accounts which are um you might not think of them as being at risk but is worth underscoring that they are. they can get turned off at any moment and there's some amazing material there documenting the people and uh work of the physical sciences. So I'll leave you here with uh um um sort of action you can take. Um I'm happy to talk to anybody about any of these projects or if uh as I was stressing one of the things that we do is help connect people within the physical sciences community, the scientists, the people in the societies with people in the library and archives communities as well. And so we're really happy to be bridges and help figure out um how we can sort of protect and preserve materials. And then also we have a weekly monthly history newsletter that has really engaging articles intended for broad audiences. And every month we publish a research updates uh newsletter that has a rundown of all the social science and sort of policy related uh work we publish as well. So I will stop there and then I think we should have a good bit of time for questions. >> [applause] >> So yeah, it looks like we have about 18 minutes before lunch. Uh so I will invite you to approach the microphones, introduce yourselves. Uh and I believe the session's being recorded, so that's why the mic's >> Go ahead. Sorry, I I wrote down my question, so I was trying to cool it up. Um my name is Rosalyn Mets. I'm at Emory University in Atlanta. Um I have a question for you, Merrily. Um so, um access to local government information has been a personal story of mine. I've been taking a a class called Decator 101, which is the city I live in. And one of the things I've noticed is there's a lot of misinformation happening within my local government because none of the records are online. So, um I I should also say I'm married to somebody who works for the school district, which is where all the contention is. Um, so do you know if other states besides California are working to digitize local records or do you see plans for that in the future for how um other states can go about um creating programs kind of like the one you described in California? >> Yeah, thanks thanks for the question. Um I don't the short answer is I I don't know. Uh I'm fine with my position. I'm on a journey of discovery. Um, one of the things that I have been able to do is to connect with the um, council of state archavists. Uh, and I'm hoping that through KOSA I can learn more about um, how uh, you know, those are state archavists. They're responsible for the state, but they're going to have more information and context around what the um uh what the local records storage and management is. So, I would love to scale out. Um you know, I've even only learned more about the California situation even though being a lifelong California resident and a librarian, I was very ignorant of this of this area until just recently. Um, so yeah, would love to learn more. Um, and you could also find out and relay more information to me about Georgia. Would love that. Let's go over here. >> Uh, Nathan Tolman, AP Trust. Um, quick comment before my question from the last one. When you're looking, um, there's big differences around the country whether power is concentrated at the county or local level. And so that might be a nuance to deal with when you're trying to get those records. My question is for you Marily as well. Um there seems to be a variety of methods in our archive is using to collect government archives of all sorts including the upload option. I'm wondering how are how is the authenticity of upload documents validated because it seems like that could be a backdoor to inject false narratives into the government record to perpetuate other bad actors. That's a great question and I don't know the answer to it and I will have to um I will have to get back to you about it. I think that um an answer is that we don't validate information that's uploaded. It just simply is recorded as information that has been uploaded. Um so I think that uh you know people need to do their own provenence work. Uh so that is Yeah. >> Could helping them do that be part of the project? >> Yeah. I I mean yes. Uh so one of the things that I just learned about last week is we have a government information librarian who's been um furiously uploading congressional research service reports um to the extent that we had identified her as a bot which she most certainly is not. um you know so putting putting better tools into her hands so that she can do that more effectively um and uh yeah kind of validating uh you know our super users in that respect government information librarians are incredible um and very concerned with you know ensuring the durability of the record but then also um I think you you raised an excellent point one that I had not considered and I should have. >> Well uh thank you guys. It was uh great to hear about all these initiatives. Um I'm Nick. I'm with Nache. We're a service provider. Um we do a lot of work on uh repository systems and uh my questions I think mostly for Christie. Um you were talking about data fragility um and some of the risks that you encounter with that. And um I'm trying to I was furiously doing research and I figured I'd just ask you straight up. Um, do you think that Inveno is going to be shifting more toward OCFL for its um, uh, its persistence layer? If you know, I'm not sure. Uh, it sounds like you're very active in that community and you're, uh, you have your finger on the pulse of things and do you see that as a good pathway to bring more resiliency to data? >> Um, thanks for your question. So I will um take the title of chief cheerleader for the Invino RDM community and so I'm really fortunate to be surrounded by a strong team that is contributing on you know technical and data um activities. I will say that if if you're interested I'll be happy to get your information or even Martin may know who's sitting in the front row may have more experience with that but we can connect you with the information. One thing I will say though, I mean, so I've been part of a number of open source communities over the years, and I am I I continue to be astounded by this particular open-source community because it is incredibly thoughtful with respect to leveraging uh meaningful and practical standards. Um really trying to push the limits when we're looking at features that help to drive better behaviors. And then um the other thing is that there's just you know again kind of um reflecting back on that human element there is uh just enough um celebration and appreciation for the different contributors in the program to you know to be able to um have that I guess psychological safety to feel trust and be able to move forward together. So yeah but I I'd be happy to connect with you and then I will get you that information. That sounds great and that sounds like a great community to be a part of. So that's that's awesome. And one really quick question. My friend David Schober used to be with Northwestern >> and did some great work on semantic search and then you moved to Internet Archive. I just want to make sure there's no bad blood there. >> I'm just kidding. >> I think so far we're good. >> Yeah. Yeah. Yeah. No, David is amazing. And I think, you know, that's the nice thing about being in this community or this bigger community is that there is that chance to I see people who are doing great work. I mean, and being genuinely thrilled when they have an opportunity. >> Obviously, a huge joke. Um, [laughter] I think you're all doing important work and thank you guys. Appreciate it. >> Yep. >> Mark Calfadvic with the triple AF consortium. My question is a little bit of a follow on to Rosyn's question. And I live in Arlington, Virginia, which has a very open, transparent government, large numbers of digitized documents. But what I found recently is they're getting pushed behind um challenge walls. Um again, probably primarily for protection from bots. And I was just wondering, is there any evidence that there's increased blockage of harvesting of go local government specifically sites to um Yeah, I I don't know that we've done um an analysis of how much of um of uh the government space is being blocked, but I think that in response to very aggressive bots, um everybody is tightening things down or refusing bots and that's problematic for us because of course we are a bot on the web. We are also busy defending off bots. Um so uh so yeah, so finding avenues and pathways for um this is something I think that our community really needs to negotiate is is how do we um how do we come up with uh thoughtful negotiated um avenues for ensuring that we are continuing to collect what we need to collect when it's web- based materials. Um, and I don't I don't know what the what the answers to that is. I know that this is an area where Rosie and others have been actively working. So, should have a CNI discussion about it, I think. There we go. >> Thank you both. Thank you all for the the topics you brought forth. I wanted to talk a little bit about the uh local records question. Uh, I'm Dennis Clark. I'm the Librarian of Virginia. The Library of Virginia is the only state library archive that's a member of CNI. And I think this is one of the very good reasons why that should not be the case anymore. And I think your connection with KOSA is exactly the right way to start talking about this at scale. Local records are challenging because every state's public records act defines what that means for your state. In Virginia, obviously, it's a very um um broad and powerful act. Uh and so we can work very closely with local records, local circuit courts, local governments, local school systems to make sure that those items are kept available in whatever way various parts of the act uh require, but that's not certainly the case in every state. Um and it's it's every state is different. Uh and I love the notion of thinking about this at scale with KOSA. Uh don't forget the the costa group, the state libraries as well. uh they're they're very much a part of the same continuum of of of conversation. Uh but again, a really good reason why uh CNI should should continue to to have those folks as part of this conversation. Thanks. >> Yeah, agreed. And thanks for showing up in that capacity. Um I think that that is really important. I didn't mention Kosla because the internet archive has uh for has strong relationships, existing relationships with KLA that I'm able to uh leverage. So the KOSA one is because because I am a careerlong archavist, I'm helping to bring that in, >> which is which is great. And and and I'll just say that uh very proud that Library Virginia was, I believe, the first state library that was actually a partner with Internet Archive back in 2005. So, uh, this has been an important part of of everything that we do and and we really are thrilled by that level of collaboration and and want to see it grow and and and be and be and be be bigger and better and more transparent. Thanks. >> Amazing. Thank you. >> Well, um, merily that I had a follow on question to Dennis is for you too. uh for the most part is what y'all are focusing on um published material things that have been publicly available because I think part of Dennis your question also gets at records that are you know organizational records and it's just to stress that like my sense is as rough as stuff is with publicly available material like organizational electronic records we're trying to focus on this a bit with scientific societies but for all kinds of nonprofits or government entities universities just a huge >> yeah and it is it is such a challenge especially when everything moves digital differentiating between publications and records becomes really challenging I would say and there's a kind of squishy continuum in there um I don't want to rule things out but I would say that our emphasis is on published materials >> right >> I am Matt Merick from the National Center for Atmospheric Research in Colorado um so this is question is for Christie I really appreciated your description of effort that you're doing there. Um, and my question you kind of touched at the end about some coordination among other similar efforts and that was the thought that I had as you were talking because I've heard of other efforts that are doing trying to do repository coordination. Um, you know, American geohysical union for example is is has an effort there. So I guess I just wanted to get your thoughts on you know how do you feel like there is u you know the sort of a common movement or are things kind of going in directions that are uncoordinated and do we have a common voice yet that we can speak from a repository point of view? Yeah, thank you. That's a great question actually because um you know I think anytime you get a lot of passionate people around uh a good goal um you know it can be hard to channel that passion in in useful ways. One thing I will note is that um first off just speaking about this center for open science project that um I've been working on is that we have very um deliberately and intentionally worked to look to see you know like what exists out there what kinds of um frameworks or initiatives or those types of things so that it's not this oh this is important let's start a project kind of thing and so you can see that with a number of the people that are actually on that program. The other thing that I'll note is that um so AGU and then the internet archives um meeting in March had a number of folks who are also participating in this group. So there's been quite a lot of cross fertilization and a real eagerness um to try and be coordinated as much as possible. um we're trying to move forward um in a way that does reflect it's not at the data set level, it's at the repository level. So what does resilience mean in those systems? And just trying to identify some of that so that we can have I mean it's just like any conversation if you can have a shared language it helps you be able to collaborate more efficiently. Um so that's been good as well. But there is, you know, and I'm happy to um to get folks the um the information, but there's an email and then that QR code. We not only do we want to hear from you, there's work to do, friends, you know, so um you know, if there are things that make sense, like right now I'm working on that maturity model. So that's a a very um traditional project kind of in informatics land um especially biomedical informatics. And what it does is it allows you to be able to look at levels of maturity from kind of ad hoc processes that are chaotic all the way to fully implemented and continuously improving. So that spectrum and then you can look at it along some different attributes. So you can look at like culture or technology or standards or those types of things. It's not meant to create um a framework for assessing yourself in the context of other organizations or repositories. What it's meant to do is to catalyze local conversations. So here's an area where we can um realize some improvements and you know and to identify oh but here's something we're doing really well. So that's a good example of a project that is actively underway if folks are interested. We I you know we're very collaborative friendly bunch. So we're um always eager to uh to have folks get involved. Yeah. Thank you. Well, so I don't see anybody's Are you going for the blankets? Okay. I don't see anybody standing uh up for the microphone. So, um, I'm going to bring the session to conclusion and let's go enjoy some lunch. Right. Thank you. [applause]