Building a Coalition for Resilient Research Data Infrastructure
Watch on YouTubeVideo summary
The Coalition for Resilient Research Data Infrastructure (CREDI) is advancing its mission to build a robust ecosystem for federally funded research data by addressing critical vulnerabilities such as chronic underfunding, staff shortages, and fragmented incentives that currently threaten the stability of digital repositories. Led by the Center for Open Science in collaboration with key leaders like Christopher Steven Markham and Linda Kellum, this initiative aims not to create competition but to align advocacy and technical communities around shared priorities over a three-year period. The strategic plan targets three distinct audiences—infrastructure providers known as "doers," funders and policymakers called "supporters," and researchers referred to as "beneficiaries"—to ensure that research outputs remain publicly accessible throughout their lifecycle in accordance with the 2022 OSTP Public Access Policy, thereby mitigating risks associated with single points of failure from funding cuts or natural disasters.
To achieve these goals, CREDI has organized its strategy around three core pillars: assessing and monitoring resilience through a shared maturity model that evaluates eight dimensions without punitive grading; establishing diverse sustained funding models via cross-sectoral consortiums involving government, industry, academia, and civil society; and conducting shared outreach to advocate for continued investment. A significant focus is placed on creating an open dashboard for early warning systems regarding software dependencies and other risks while synthesizing existing frameworks like the Core Trust Seal. The initiative also recognizes that defining universal metrics for reusability under FAIR principles remains complex, suggesting instead a triangulation of multiple statistics such as citations and download counts to gauge usage across varying community sizes, alongside domain-specific working groups to avoid one-size-fits-all solutions.
The discussion highlights an important synergy between CREDI's discipline-agnostic, US-focused remit and the American Geophysical Union's specialized work on global data resilience, ensuring efforts do not duplicate existing resources while filling critical gaps in tools like legislative advocacy kits. A key challenge identified is the disconnect between infrastructure providers who rely on seamless operations and end-users who often take these systems for granted, a dynamic that complicates advocacy needed to sustain funding; consequently, the plan emphasizes coordinating future strategies rather than relying on ad hoc efforts. The governance structure proposes a "minimum viable consortium" with staggered terms to foster collaborative decision-making among diverse stakeholders, ensuring that crisis response best practices are developed with a global perspective in mind while building community capacity to report at-risk resources effectively.
As the draft strategic plan remains open for public comment until September 30th with an aim for final publication by November, participants are encouraged to provide feedback on feasibility, missing elements, and existing resources through provided channels like Google Forms or dataresilience.io. This collaborative approach seeks to transform current vulnerabilities into opportunities for a more resilient future where research infrastructure is not only technically sound but also financially sustainable and socially supported across all sectors of the scientific community. By engaging decision-makers from various backgrounds and fostering open dialogue, CREDI hopes to establish a lasting framework that protects critical data assets against emerging threats while promoting transparency and accessibility in the broader landscape of life cycle open science.
Read the full video transcript
All right. So I believe we are live. Um
welcome all and thanks for um joining us
on this uh afternoon in DC where I am
based although uh given um we are a
global organization. Um, greetings from
whatever time zone, whatever time of day
in the world you may be, you may be at.
Um, my name is Miriam Zaringham. I am
the senior director of policy at the
Center for Open Science. uh and I'm
really excited to be talking with you
about the work that we have been doing
um on uh with um collaborators at many
different organizations uh to develop a
strategic plan for the coalition for
resilient research data infrastructure
or credi. Uh so I am joined today um
with my uh esteemed collaborators or
some of my esteemed collaborators. Uh we
have Christopher Steven Markham from the
Data Foundation and Federation of
American Scientists as well as a whole
bunch of other affiliations and hats
that he wears. Uh Linda Kellum from the
University of Pennsylvania libraries and
the data rescue project. Christine Kirk
Kirkpatre from uh the UC San Diego s
supercomputing center and uh Alex Wade
um who is our project consultant on this
uh and holds affiliations with
Washington University as well as other
organizations. And so we make up um a uh
a part of the strategic planning
committee that went into developing this
plan uh which we are excited to be
sharing more about with you all today.
Um and so before uh I guess before we
get into it, we wanted to have uh I
guess to set up some some ground rules
or features and capabilities um in this
webinar. uh we have um a polling
function and we're really curious about
sort of where in the data repository
infrastructure system uh you sit in. Um
and so that poll should be going off now
and throughout this webinar you also
have an opportunity to ask questions or
share comments through uh the Q&A
feature on this webinar. And so you know
part of this is to share with you um
this draft strategic plan that we have
developed. Uh but it is also to hear
back from you what questions, comments,
ideas you may have as we work on um you
know developing this draft strategic
plan into a finalized strategy for this
proposed coalition. And so first I
wanted to share a bit about how it is
that this effort started. And so um it
kind of came out of work from the center
for open science. And for those of you
who might not be familiar with us um we
are a nonprofit organization with a
mission to increase the openness,
integrity and trustworthiness of
research. And we do this um to sort of
work towards our true north uh which is
life cycle open science. And we define
that as research with publicly
accessible plans, contents like data,
materials, and code, and outcomes that
are linked and findable in a persistent
open location. We're also an
infrastructure provider through the open
science framework, which can surface
these connections to really show you the
life cycle of open research. So
researchers can start by pre-registering
their study plans and as they undertake
their work link to outputs and outcomes
that are hosted anywhere as they become
available giving this complete view of
how research is progressing over the
life cycle from study plan to outcomes.
Now, achieving our true north requires
um persistent and reliable access to
those outputs and outcomes so that those
linkages are long lasting, are
trustworthy, are robust. And so the core
of our vision for life cycle open
science is is the belief conviction that
uh research information including data
uh uh is a public good. And so reaping
the benefits of this public good
requires persistent and reliable access
to those outputs and outcomes, which is
to say the infrastructure that enables
in this case data to be shared, linked,
discovered, accessed, understood, and
used in context. Yet, these
infrastructures uh have long been at
risk from a variety of factors like
chronic underfunding, including a
tendency to favor funding new
innovations and maintenance uh new
innovations over the maintenance of core
capabilities, overstretched staff,
incentives to fragment the system rather
than supporting convergence around
shared capabilities like standards and
metadata, and vulnerabilities due to
shifting policy priorities. And we've
really seen these risks um sort of
exacerbated and and come to bear uh in
in the last few years in a pretty stark
way. And that's led to several data
resilience preservation and rescue
efforts that here at the Center for Open
Science we've been really inspired by uh
including work led by uh the faces that
you see before you. And those efforts
have had different scopes. But for our
purpose at the Center for Open Science,
we're particularly interested in
infrastructure that provides access to
federally funded research data, which
falls under the remitt of the 2022
Office of Science and Technology public
access policy, which is to say data that
are of sufficient quality to validate
and replicate research findings, which
is generated with support from U the US
government. Uh so that can include
intramural research, grants, contracts,
cooperative agreements and so on. And
that also means that ownership of this
data varies and could be the government,
an institution or individual. All of
these kinds of things are in scope and
of interest to us. And while I recognize
that administration priorities have
shifted, open science policies like
those advanced with the 2022 public
access guidance uh continue to be a
priority under this administration under
the opaces of of gold standard science.
And we could have a whole other
conversation around that, but I'll I'll
leave it there for now. Um, and so the
open science community has done a ton of
really important advocacy to move
policies forward towards greater
openness. And we want to make sure that
that progress continues. Which means
that if we are going to ask researchers
to make their data, to make their
outputs more open, accessible,
understandable to link them together and
provide that contents context of
openness throughout the research life
cycle. We need to make sure that we're
enabling a more resilient uh ecosystem
for research data infrastructures
um so that they are resistant to single
points of failure and that we can uh
believe and and have confidence that
they will be preserved in the long term.
And so single points of failure may
include a funding cut when there is a
lack of diverse funding streams, a
natural disaster where there is no
backup protocol, or a shift in
institutional priorities where a
repository lacks multi-institutional
stewardship. And so a sort of second
question or challenge is how can we do
this? How can we enable greater
resilience across a really distributed
system to solve this challenge together?
because no one organization or
initiative or repository is going to be
able to do it on its own uh to save us
to save all of the data uh in its
entirety.
And so first I want to um sort of be
clear about what it is that we mean by
resilience. And we drew from a
definition that was put forward by work
that was done by the earth science
information
partners or EIP uh that you can find in
a report that I provided the the DOI in
this slide. Um but it's the ability to
make important data accessible,
recoverable and useful under both normal
operations and crisis conditions. And so
again, this really requires
collaboration across the system um at a
really massive scale because data uh are
housed in a highly distributed system of
thousands of repositories across
disciplines with different data users,
curators, depositors and so on. And so
what does it really look like to
coordinate our efforts to make the most
efficient effective use of limited
resources and again drive towards
greater resilience for the system? And
so COS again recognizing that we could
not do this on our own uh and that
there's so many incredible efforts and
initiatives that have been longstanding
um looked out to the community to
leaders in this in the community to help
us try to answer these questions and
develop a strategy for how we move this
forward. And so we're really honored to
be working with um this great set of
people here who really act as um who we
see as leaders in this field and who
have really rich connections and
relationships out into the broader um
data and repository um and information
infrastructure system. And so we got uh
funding from the Robert Wood Johnson
Foundation to help us sort of pull
together these great minds and think
over the last year about what is needed
in order to drive towards greater
resilience. And so with that, I will
turn to one of these great minds, Linda
Kellum, uh to take it from here.
>> Hello everyone. It's very I'm excited to
talk with you all. Uh I'm Linda Kellum.
I'm the director of research data and
digital scholarship at Penn Libraries
and one of the co-founders and director
of the data rescue project. Uh so as
Mary mentioned many times the cutting
across all of our work is a commitment
to collaboration. Um we want to be
informed by the landscape um that is
existing. Um and so part of our uh
effort has been on a landscape analysis
of the existing related efforts um so
that we can ensure that our approach is
very complimentary.
The goal for the plan is to act as a an
anchoring framework or scaffold for
coordinating related in initiatives so
that the community can organize to
promote a shared vision for a more
resilient ecosystem. And we're thinking
about how we sustain and govern that
moving forward. Um this is why we have a
dedicated discussion that uh al within
the strategic plan about governance as
well as the call for action and the
pillars that you'll see. So next slide.
So after a few months of exploratory
work uh we gathered in DC in March um to
build out our strategic plan and this is
the case for action that we came up with
as a group. Um these are some of the
core assumptions that we have uh as we
entered into the the planning process.
First off uh federally funded research
data is a non-rivalist public good. Um
and I think you've heard can hear that
across many of our organizations as we
talk about public data as a public good.
Second that the use of public data
relies on robust and resilient data
repositories and related infrastructures
and trying to find ways to bolster those
as much as possible.
Third, those infrastructures have long
been at risk, but those risks have
become clearer to us over the past few
years.
Fourth, government support is essential,
but can be made more efficient through
coordination across sectors. And
finally, that we can't go back. We can't
wish for a past that doesn't exist
anymore. We need to think about what
we're we have and what where are we
going from here.
So the result of this case for action,
this call for action is a strategic plan
and we're sharing it with you because we
want to know your thoughts and uh
whether the elements in the strategic
plan resonate with your communities and
if it doesn't, if something doesn't, how
do you think we should shape it to meet
the needs, concerns, and opportunities
of your specific community?
Uh next. Um so the case for action
builds towards our vision um that we're
working to achieve through our strategic
planning process and that we've been
working over on for the past over the
past year. Um we're hoping uh to enact
the plan over the next three years or
so. And this plan really seeks to
identify the conditions or the
recommendations and strategies that
advance the work um and the governance
structure for a coalition that will
collaborative collaboratively c uh
cultivate a resilient data ecosystem for
publicly funded research. So working
across different organizations um and
bringing in together to bringing
together organizations um to find a way
forward to create a resilient data
ecosystem.
Next slide.
So rather than launching one more
initiative that competes with other
groups and similar groups, KY um the
group that we've brought together has
really worked to align the rescue rescue
advocacy and technical communities
around shared priorities. Um we're
hoping to build shared accountability
for sharing resources effectively and we
want to do this by in these three ways.
So one is advocating and building
building awareness. um we want to learn
and borrow from other data rescue and
digital preservation efforts and amplify
what they're doing. Two is to reduce
done redundancies across efforts. So uh
when we do see something that's being
done across in a different organization
thinking through how we can work
together or how organizations can work
together and finally supporting
long-term infrastructure in our efforts.
And next slide please.
And we see three distinct groups in this
um as our audience although uh there's
definitely overlap in these and there
may be other audiences that you are you
feel are missing from this but uh the
first is are the doers this is my area I
think so those involved in data rescue
preservation and providing
infrastructure um so it's the doers not
just in terms of um what we've had to do
over the past couple of year or past 19
months but over the past few years the
people who are actually working to
preserve and provide infrastructure.
Second is the supporters, those who want
and have the means to support our work
through funding, policy, advocacy and in
other areas. And finally, the
beneficiaries who those are the
researchers and other users who benefit
from the ef from our efforts to support
data infrastructure.
So these are our core audiences and the
people we're thinking about as we were
creating the strategic plan. Hopefully
you see yourself in here and uh if not
we would love to know that as well. So
I'm gonna hand over to my colleague
Chris Markham.
>> Uh thanks thanks so much Linda. I I join
you in in uh being grateful to be able
to join everybody today at Center for
Open Science for this review. Uh so uh
as um both Miam and and Linda had said
publicly funded uh research uh data is a
is a vital public good. It fuels
scientific discovery, drives innovation,
and really belongs to the community at
large, belongs to all of us. But the
repository ecosystem that fedally funded
research data relies on is surprisingly
fragile. And that's what the strategic
plan is recognizing in the work that
we're doing in this group. Repositories
face, you know, fairly consistent
threats from budget cuts to policy
shifts to technical issues and and staff
turnover. To tackle these
vulnerabilities, the proposed strategic
plan focuses on three core pillars and
I'll I'll walk through them uh with some
detail.
First, uh pillar one focuses on
assessment and monitoring. This is our
assess and monitor the repository
landscape pillar. We need a clear shared
model to track and evaluate where the
repository ecosystem is most vulnerable
and how those risks can change over
time.
Second, uh this uh the uh pillar number
two in the proposed strategy um is to
develop diverse sustained funding and
governance. This tackles long-term
stability through collective governance
for uh the federal uh repository
ecosystem. The pillar focuses on
coordinating investment in the
repository ecosystem to build robustness
and sustainability into funding models
so it can survive threats or single
points of failure uh in the future. And
finally, pillar three, which is build a
shared outreach and advocacy and action
strategy, centers on community and
collective action. This is like really
putting the community into this uh uh
into this strategy was very important to
us. We can't do it alone as uh Miriam
said. And so we need to foster an
environment that equips researchers,
stewards and advocates and the three
buckets uh that that Linda uh had just
described uh and of course the public at
large uh with the toolkit that they need
to maintain the signal the value of data
on the one hand and to respond to active
threats in the ecosystem on the other.
Now I'll turn to a bit more detail on
pillar one before handing the stage over
to Alex and to Christine to divi uh to
dive a little bit more into pillars two
and three. So next slide please. Excuse
me. So pillar one, assess and monitor
the repository landscape recognizes that
understand that our understanding where
the baseline risks and strengths are
lies is u where they are is an essential
first step in building uh resilience to
the repository ecosystem. So we have to
obtain that assessment right now. If a
funding stream dries up or if uh staff
is uh dresourced or or staff members
leave or or a policy priority changes,
an entire repository can disappear or
the data within it can wither. Those are
single points of failure and that we
often don't see coming until they're
it's too late and they're right at the
door. Importantly, pillar one is is not
about reinventing the wheel or remapping
the ground. the community has already um
uh uh made made a lot of headway on and
and and doesn't uh tread on what our
common understanding is. It's not
duplicative. A lot of great work has
already been done to measure the digital
preservation and data quality. Our our
goal here is to bring those existing
tools together into a unified framework
by establishing a common language for
risk that we can help funders and
agencies and stewards spot weaknesses
early and to protect our collective
public investment in data uh in the long
term.
Uh next slide please.
So uh there are really the pillar one
really has three big components. The
first um is um
and so this slide describes how we plan
to operationalize approach to
operationalize pillar one in practice.
We we have three concrete steps uh here.
First, we're developing a shared
maturity model that gives the entire
community a standardized way to evaluate
repository resilience that is both
comparative and introspective. Instead
of grading repositories on a sort of
past fail basis, it really evaluates
specific operational capabilities for
the repositories like whether a
repository has backup bunding or if
there's a succession plan in place or if
there's distributed storage across
multiple sites. Second, uh we recognize
in the uh in the strategic plan that we
need a comprehensive landscape analysis
that maps the repositories um that host
federally funded research data um and to
and and the goal of that is really to
help understand the systematic gaps and
risks but also the strengths of of what
what exists in the current system. for
example, um two repositories that might
look perfectly healthy on paper. Um but
if they both relying on the same
underlying software dependency or
storage provider, that might be a point
of failure that that uh that is
susceptible to both. And that shared
dependency really needs to be revealed.
And third, to help with accessibility
and insights in uh uh into um uh into
the re revealed in both the landscape
analysis and the maturity model, uh the
plan proposes a uh repository resilience
dashboard. This should serve as an open
communitydriven tool that will monitor
repositories over time. Our hope is that
the dashboard will give uh information
stewards and supporters an accessible
early warning system so we can uh direct
resources and assistance and let people
know uh when uh when when when data
rescue uh and uh uh uh and where
critical need is um needed most. Uh next
slide please.
So the big component of pillar one is
the repository resilience maturity
model. uh I thought we'd take an over uh
a look at the whole model overall uh and
uh and uh and give you a glimpse into
into what we're thinking here. Um to
ensure that we're capturing the full
picture of health of the of of any given
repository, the model evaluates each
repository according to sort of eight
core dimensions. Uh these are they're
listed here. If you can't the print's a
little small, so I'll read them. It's
organizational and financial stability,
governance, transparency, and
accountability, data stewardship, and
integrity. uh the fair data principles
that's making data uh findable,
accessible, interoperable and reusable.
Uh software and technical dependencies,
technical infrastructure and security,
community use and crisis uh readiness
and resilience. So those are their eight
uh core dimensions and they're assessed
across um a um uh criteria uh in four
very plain uh simple to understand
stages of maturity. Uh an initial stage
where operations are ad hoc with limited
capacity. An emerging stage where
capacity is driven by immediate
requirements and need. Uh established is
a more mature uh aspects of the maturity
model where where the repositories are
formalized um and implementation might
still be bumpy but is has regular uh
standing operating procedure and uh and
optimized. This is where resilience is
very deeply embedded and and routinely
tested and robust and continuously
improved. The maturity model synthesizes
five established frameworks. We're not
again building anything new. We're
trying to build on the work that's
already been done and that are already
trusted by the community. These include
core trust seal uh the national dig
digital stewardship uh national digital
stewardship alliance levels uh of
digital preservation and the national
science technology council's desirable
characteristics for federally funded uh
data repositories which a number of us
on on the uh on on this working group
have have contributed to. Um so the
maturity model is designed to be
practical and transparent and
actionable. The initial proposal model
is published separately on Zenotto and
uh the the link is there on the screen.
We welcome your feedback it uh feedback
on it and um I'll turn it over to Alex
now to discuss pillar 2 and provide some
more insight into one of the aspects of
the maturity model.
>> Excellent. Thank you very much Chris. Uh
my name is Alex Wade and um yeah the sec
second pillar um of the uh uh of the
strategy here is really to uh focus in
on a couple of those uh facets. So uh if
you go to the next slide please.
So um in the within the maturity model
the the top two uh first two pillars
first two facets I should say are
organizational and financial stability
and also governance transparency and
accountability. Um so if we explode
those a little bit more I'm not going to
read through them but next slide please.
If you uh download from Zenoto the draft
majority model, you'll see this
progression as Chris outlined from the
sort of initial ad hoc stages uh to a
more um optimized and robust model. And
so specifically pillar two wants to
focus on developing out uh the set of um
sustainable business models for research
data infrastructure and also to
establish a multis sectoral uh
governance mechanism so that we can
create some shared accountability um and
put the structures in place to sustain
the ecosystem in the long term.
Uh specifically addressing some of the
single points of failure risks that that
Chris outlined. We want to uh support a
progression within this maturity model
from sort of left to right. Um our aim
isn't to say everybody needs to be uh
fully at the at the right hand side of
this uh model, but at each stage it sort
of highlights and you can infer what
some of the risks might be if you are uh
too far to the left in that model.
And so we want to uh uh facilitate some
increased coordination across the
ecosystem so that each repository isn't
trying to uh sort of evaluate and move
within this maturity model on their own
and we want to maximize the resources
and share practices across the
repositories and the players in the
ecosystem.
So to that end, one of the things that
we would like to do is to survey and
catalog the existing and even proposed
business models and implementation
methods as well as to identify
opportunities to develop and test
innovative business models so that we
can create a menu of options that
repositories and funders can draw upon.
Uh we recognize of course that there is
no one-sizefits-all. Um so there isn't
there isn't going to be a single answer
to this and it's going to be highly
dependent on uh the repository type the
funding context the life cycle stage of
the repository. So part of uh the the
goal of this catalog of business models
is also to sort of highlight the the
trade-offs and advantages of different
approaches there.
Um once we have that, we would like to
promise promise we would like to pilot
some of the promising models with
willing repository partners uh with
resourcing from aligned funders and
ultimately to document some of the
lessons learned that that uh comes out
of those pilots and then we'd like to to
publish that uh really with sort of some
broader recommendations to the community
for for wider adoption.
Um secondly, I've gone to the next slide
here.
Um secondly, we would like to establish
a cross- sectoral consortium for
collaborative funding and also for
incentives so that the investments that
the the community is making uh funders,
the government is making in research
data repositories can be better
coordinated. Uh we would like to see
this composed of senior representatives
from government, from industry, from
academia and from civil society uh but
specifically those who hold
decision-making authority and those that
commit can commit resources to the
effort. So clear uh incentives for
participation and shared buyin uh around
a common uh infrastructure or
infrastructure approach would be
essential to this. And we're not naive.
We recognize that there have been past
uh inter agency and multis sectoral
efforts that may have faltered from
insufficient authority or insufficient
resources at the time. So we intend to
apply lessons learned from this to to
build a more durable coordination
uh for the core data services that can
scale and can adapt over time.
Um next slide I think goes on to back to
Linda.
>> Hello again. Um so in pillar three we
aim to develop a shared outreach
advocacy and action strategy and the
goal of this pillar is to equip the
community to make the case for sustained
investment. In addition this pillar
encompasses the need to respond when
data are at risk. Effective outreach and
advocacy must be grounded in a clear
picture of existing risks and
fragilities and to make the case that in
order to make the case for sustained
investment in data. Therefore, this
pillar equips the community with
communities with the tools, messaging
and capacity to act. We also to see this
pillar coming from a global perspective.
While most of our perspectives are
really uh focused on the United States,
we recognize that public data is a
concern not just for the United States.
And so going to one of the questions in
the Q&A um that we do kind of we do see
this pillar in particular have
encompassing a broader um perspective.
Next slide.
So the focus of this pillar is on
amplification of the need for action. As
a member of the pillar three development
team, for me it's critical uh that we
continue to think about advocacy and
outreach as key actions and not
afterthoughts. So within this first we
want to align and coordinate with
related efforts to avoid the um to to
avoid duplication. So the goal here,
some of the activities here are to map
existing related initiatives across the
ecosystem and identify overlaps, gaps
and opportunities.
We also want to look for opportunities
to bring them into the living road map
in ways that complement our existing
efforts. So if there are other uh groups
that are working in these uh or taking
action in different ways that we can
bring them into the efforts that we have
and then I think very importantly to
establish ongoing touch points to keep
those organizations engaged and informed
as our work evolves.
Second, we uh call for the development
of communication and advocacy toolkits
um that can clearly and concisely make
the case for the importance of public
data and federal fund fedally funded
repositories.
Having these in place for the community
would answer a real need and a gap that
we've seen over the past year. And so
some of the activities here include an
environmental scan of existing community
and advocacy resources across our
ecosystem as well as pri prioritizing
and developing toolkits that can do
things like advocate to legislators for
safeguarding data making the case to
funders for sustainable funding models
and general purpose messaging for
broader audiences.
Our third uh point is to build community
capacity to report at risk resources and
advocate for sustainable and resilient
infrastructure.
So this could include curating and
amplifying existing resources, directing
community members to the best available
tools and materials uh for accessing
preserved data sets, reporting at risk
resources and addressing a gap all the
gaps as a need as needed. And finally,
implementing short-term crisis response
strategies and developing best practices
for resilience. So, part of this, we
would like to convene practitioners
across efforts to share resources and
identify what has worked and what has
not and why. Um, as well as leveraging
existing crisis response map mapping
efforts um that you you'll hear about
soon probably. Um, we'd like to maintain
a community of the willing, a standing
network able to respond to emergent
issues in data stewardship and then
document the lessons we've learned from
both the crisis response community as
well as um uh translating them into best
practices for uh sustained
infrastructure.
So these uh are uh the the main
activities we see for pillar three. Um,
now I'm going to hand it over to
Christine Kurpatre who's going to talk
about the governance models.
>> Thanks so much, Linda. And, uh, next
slide, please.
So we have a consortium structure in the
strategic plan that draws on community
practices and include governance models
refined over time and via real world
challenges by datadriven consortia some
of which uh you participants have been
involved and helped to refine uh I've
been part of the stakeholder alignment
collaborative where we draw uh some of
this knowledge from and including our
recent book the consorcia century um the
work itself is driven purposefully by a
consortium of community interest holders
who participate in shared governance.
And this is to drive equity as well as
to ensure buyin on decisions uh so that
you can successfully implement these
things and have people adhere to
community norms. And this is just meant
as a starting point. It can be adapted
to the specific context of the
consortium once it's been established.
to give you a couple quick highlights.
Um but please do uh read the plan and
comment. Uh we talk about a minimum
viable consortium so just enough
structure uh that can be iterated as
needed. Uh and this concept of staggered
terms for officers and a steering
committee how voting membership and
other aspects would work. And we do uh
talk a bit about uh or reference Eleanor
Astramm's work which really points
basically to the need to set
expectations for behavior and the
ramifications for act acting outside
those norms and we again reference
something from uh ESIP that was uh
mentioned earlier and its community
participation guidelines that were uh
based on the Misilla Foundation. And
finally, um we acknowledge that such a
consortium may need an eventual plan for
winding down and so to be purposeful
about not just the beginning but uh the
end as well. And let me turn it back
over to you.
All right. So, uh this is the portion
where I say again that this um plan that
we have released is a draft. Um it is uh
sort of from um months of work with this
um strategic planning committee uh and
is now in a form that we really want
input um from the broader community. So
that you know we really mean what we say
when we're saying this is a
communitydriven strategy. Um we want to
make sure that existing efforts,
initiatives,
um uh resources, infrastructures are
reflected in this kind of work rather
than you know duplicating efforts or
going off in sort of parallel directions
uh in a way that might not be as
constructive as it possibly could be. So
um in tandem with releasing the
strategic plan, we also released um a
Google forum for folks to um provide
input and are planning on hosting um
sessions uh in other forums to continue
to get sort of like live um in real time
feedback uh like uh is is coming through
the Q&A function right now which is
really really wonderful. Um,
specifically things that we are
interested in knowing is do the
strategies that we outline in this
strategic plan that we've touched on in
this webinar serve the ecosystems needs
with respect to long-term repository
resilience? What's missing in here that
that we really need to be mentioning or
calling attention to? Uh, are the
strategies feasible given current
realities? Where do you see obstacles
that should be on our radar or
opportunities that should also be on our
radar? What existing work resources or
initiatives should this plan be
referencing or building on whether we
reference them explicitly in the
strategic plan or for reasons that make
a lot of sense, you know, these efforts
are sort of flying lower on the radar um
but should still be sort of built on uh
and and worked on. um uh constructively
and then what specific input might you
have on the pillars and the proposed
governance model that um Christine
walked through? Um, and perhaps most
importantly, are you interested in
joining or supporting KYRED and how? Um,
so I'm going to uh pull up in the poll
again uh if we could um a question about
sort of like where in this um where in
this system are you coming from? Uh as
we um sort of start to move into waste
ways to comment. So, as I mentioned, um
there's the Google form. Uh you can scan
uh for the QR code there, and we'll also
have a follow-up for attendees of this
webinar um so that you can uh enjoy it
via email. Um you can email us at data
resiliencecoos.io.
Um and then we're also going to be
hosting um office hours uh with partner
organizations uh in order to provide
feedback feedback uh in that way. And so
comments are open through September 30th
at which point we will be uh
incorporating those comments into a
final draft which we're aiming to
publish in November uh along with some
details about initial steps towards
implementation.
And so, um, to give you like a little
snapshot of, uh, of where we're at, um,
so that we can now move into the Q&A
portion, uh, I will just leave up here
the, um, the pillars of the strategic
plan as well as a highlevel summary of
the strategies in there. Uh, and we can
start to work through the questions and
I'll do my best to serve as moderator
um, as we answer them live. Um, so it's
it's cool to see where you all are
sitting. Oh, and great um we've also
flashed up another um poll that you can
sort of muse on uh about you know how
you might be interested in staying in
touch or involved with KY as we continue
to move forward. Um so I will go to the
first I'll I'll group the first two
questions which are really around the
geographic focus or scope for the plan.
Um so is this focused on US-based
infrastructures? Are we planning on
including international initiatives?
Chris, I believe that you were starting
to draft a response. So um great
Christopher Markham is going to answer
this live
>> in an unexpected twist. I'll be
answering this live. Yeah. So I think
this is a great question. There are a
number of questions in the chat about
what the geographic focus is uh whether
there's international considerations,
how this would apply to say the Middle
East and the global south. I think these
are all fantastic questions. In our
current instantiation, in our initial
thinking, we have been focused on US
federally funded research uh data
ecosystem. However, um that does not
mean that uh we are um uh we are we have
a moratorum or anything on uh on
international collaboration and
participation and so we'd welcome those
ideas. We think there are probably
lessons learned uh that that would be
very valuable uh say um from um uh uh
from the non- US context. The other uh
aspect of this is that the the um the um
eventual deliverables and outputs from
the work the credit will do uh should be
generalizable and should be um you know
quite um quite uh useful to the global
community and by putting in the public
domain we hope that that will be um you
know part of part of the paying back to
the global uh u uh data uh data research
ecosystem.
Yeah, thanks for that Chris. And I think
that there's there's a related question
around, you know, uh that I've that I've
been receiving around do you only care
about repositories that house US funded
data? And I think the answer is that I I
can't quite think of a reposi well I
guess I can think of repositories that
only host US um funded research data. Uh
but many of the data repositories that
we are looking at and focused on um are
opened for deposit from you know private
funders from international funders and
so these infrastructures are not um you
know bordered sort of like geographic
enclaves necessarily. Um Alex you you
might want to talk about some of the uh
the work that you've been doing sort of
surveying um repositories or you might
have a different comment entirely. Well,
no. I I was going to to pile on to both
what Chris and you have said is and
there's another dimension to this
question which isn't just what the data
is about or how the data is funded but
it's about the the global community of
researchers and you know as you know the
re research communities are not
geographically bound per se. So I think
there's a very real aspect of the um the
beneficiaries of the data that are
outside of the US and you know one could
look at the resilience there and say my
research or my research domain is
entirely
uh single funded by some US federal
agency. So there's some very good
examples I think especially in the
biomedical space of resources that are
now governed multinationally have uh
transnational uh replication uh
replicated repositories and so um while
this initial effort with credi was
focused on US fally funded data I don't
think that the um the the breadth of the
community that we want to involve needs
to be geographically bound that way.
>> Yeah. Yeah. Yeah, I think we were just
kind of thinking about what is the
scoping that makes the most sense rather
than thinking about how can we do it
all. Um, so I will now move on to the
next question which I I think I may need
some clarification from from the asker.
Um, Oscar Javier Guerrero Gutierrez
asks, "Are there plans to include a
layer of peer review of the
contributions to the repository
ecosystem?"
I'm not sure that I understand what um
whether this is asking about whether we
are planning to do a review of the
quality of data sets, which I would say
no. we're not looking at um sort of
quality assessment, quality control of
data sets. Um but if that was not your
question, if you could comment in there
or if somebody else um wants to chime in
there before we move on to the next
question.
Okay, so not just quality.
Um yeah, if you could clarify, we can we
can come back to that question. Um the
next one is around um evaluating
repositories against a standardized
maturity model may disadvantage smaller,
underresourced or marginalized research
institutions that lack the
infrastructure to meet centralized
compliance standards. Um how how can
this be addressed if at all? Um so we
had talked a lot about the sort of um
naming of a maturity model um as not
wanting it to seem like we were scoring
or grading um repositories but rather
giving them a sense of where they
currently stand and where that they
stand in context or in comparison to
other repositories as a way of thinking
about are there ways that we can share
resources, share tooling, share infra
infrastructures to help those um less
wellresourced uh but still very
important um repository infrastructures
sort of like move up. Um so being able
to look across the landscape and see
like these are some core areas where we
need more tooling in order to move along
that that maturity model. um we are open
to ideas of another thing to call it
that might you know sound a bit less
judgmental but it's really I think aimed
at helping along uh infrastructures that
may have um less access to resourcing to
think about what are ways that we can
more effectively collaborate and
coordinate uh to be sort of like moving
further um to the uh right of that
maturity model towards I can't remember
what the highest level is um Linda, you
can go ahead.
>> Optim Yeah, optimizing is the highest
and I think that was why it was so
important that we drew on diff the
pre-existing models that are out there
such as Cortal and EIPS and and and the
others. Um as somebody who works closely
with Cortra Seal, I think the the thing
to keep in mind is is looking the goal
is to look strategically and and and um
fully at where you are in uh um and
being able to rate yourself, not be
rated by others. And so that really is
the the the um and not even seeing as a
rating. It's just kind of giving
yourself an indicator of where you are
in terms of developing the
infrastructure around the repository. Um
so I think
this is definitely a great question
definitely something that we've talked a
lot about within the cruddy group um uh
when it comes to the different kinds of
repositories that are out there. But
hopefully this will still be a useful
tool for people.
>> Yeah. Yeah. Thanks for that, Linda. And
something that we have talked very very
uh and thought um very long and hard
about and would really really value
input from from you all at
data-resilience
coos.io.
Um I think relatedly um is the question
from Shannon on what um who is doing the
evaluating of the repository. Is this a
tool that you plug in an application we
submit or our own analysis? So we were
sort of conceptualizing this as a self
assessment recognizing that some of the
um the dimensions uh within that um uh
that model are things that an outsider
might not have access to or even be able
to answer. Um, and so, you know, part of
that is also thinking about like how can
we sort of incentivize people to opt in
and do these kinds of assessments
um, and and kind of contribute them and
what are ways that we can do this um, so
that repositories aren't necessarily
putting themselves at greater risk by
maybe saying that like we're not, you
know, totally towards the optimizing end
of the spectrum. Um, so that's something
that's that's continues to be an active
and and live conversation. Um, but I
don't know if any of my other uh
colleagues or collaborators have
anything that they want to add to that.
Maybe Chris, anything on on sort of like
dashboarding um that draws from your
Okay. No.
Okay, great.
Um, all right. from Carrie. What metrics
will you use to assess reusability
in the fair acronym findable,
accessible, interoperable, and reusable?
Or is that still coming together?
I think still coming together, but uh
perhaps my
my colleagues may have more to say or
want to talk more about that.
I'm going to uh I'll say one thing and
then I'm gonna ask Christine to chime in
because I think Christine has a lot of
expertise.
>> Yeah, I wanted to ask Christine but I
>> Christine has a lot of expertise too
aggressive for a Friday afternoon
>> and help in helping organizations uh u
uh uh u uh bring up their data standards
into fairness. But I I will say that uh
just like Miam said, the entire plan is
is definitely still coming together and
open open for uh suggestions. And so if
you have good metrics and good and best
practices, please share them with us.
And uh but I I think maybe Christine may
have a word or two to say about this.
>> Yeah. Although you might be disappointed
how uh existential my answer is. I mean
I think the reality is that the moment
you create metrics, they're almost out
of date and they are very specific to
the type of thing thing you're trying to
do, right? The use case, the the domain
that you're working in. Um,
so I mean I think the fair principles
have survived over these 10 years
because they were not uh specific and
they're more about concepts, but really
when it gets into how do you measure if
if something's reusable um that really
needs to be based on the implementation
of those principles. And so these are
things I think best worked out in
working groups that are um uh steeped in
the context of the domain and what
someone is trying to do. And we're
trying to uh think broadly um uh over
these larger concepts of uh you know how
do you how do you keep repositories
alive and how do you plan for all of
these things that inevitably happen uh
uh for a variety of reasons many of
which have been uh referenced here in
the chat and which have really broadened
our thinking too about this. So, I hope
that's not too much of a double speak
answer, but I don't want to also say lie
and say, "Oh, yeah. I've got an Excel
spreadsheet that has all of the answers
for you." And it will work in all cases.
>> Alex, go ahead.
>> We uh I was one of the um uh in a
previous life when I was working for a
funer, I was one of the members of the
global biodata coalition. was something
that we uh struggled with on the um
application process for the core global
biodata resources. Um and because
research community sizes are different
um we tended to measure a number of
different things. Uh so in some cases um
citations or mentions of the resource
was an indication of reusability that it
got reused. um but in in other cases
just understanding what the uh global
web traffic to the tool was or number of
downloads or number of API accesses. So
there was a whole spectrum of of
statistics that we looked at um that all
needed to be triangulated across the um
the reach and the size of the re
research community. Um so there wasn't
any bottomline way to combine these into
a single metric.
Great. Um, I see that we have nine
minutes left, so I'm I'm trying to um
sort of quickly scroll down and see. So,
I did see a I' I've seen a few questions
about sort of how our work um relates to
the American Geoysical Union's global
data resilience work. Um, and uh the
answer is quite a bit. So AGU's work is
more focused on specifically focused on
um the earth and environmental sciences
um uh earth space environment um and we
have more of a um discipline agnostic
uh remmit uh but we are and have been
working very closely um with AGU uh and
have um quite a lot of sort of like
cross talk and synergy between folks who
are involved in that effort.
uh and folks who have been sort of like
leading the leading the way in the
charge with um KY to make sure that we
are not duplicating efforts that don't
need to be duplicated that we are
collaborating ways in ways that are
effective but also recognizing that you
know uh AGU's work um is going to
specialize in ways that ours um isn't
going to get as deep in the weeds and
their remitt is more global. Um so while
we're looking more at um those US uh
repositories that are in in the US
context but that do house as we as we
discussed data globally um AGU is
looking um broader than that. Uh so I
don't remember where exactly that
question was but I will say
>> uh I will call that question answered. I
I answered it. Um I typed an answer,
too. So that's why
>> Oh, great. Great.
Um
All right. Let's see.
I thought Amit Terzia had a great
question. Not that I know the answer,
but I thought we could reflect on it.
He's basically calling out that there's
a gap between the perspective of people
like Cruddy, which by the way, I don't
know if we've said, great acronym. I
love the humility of it. I'm proud to be
part of it. Um but also the um uh
researchers mostly but also consumers
broadly and then people producing the
data and you know if there's any tension
in those gaps between what we think is
needed versus what you know a panel of
researchers might say. And I I think
this is because he's someone who's
trying to serve uh uh
layers that that sit on top of
infrastructure that answer some of these
challenges. So what you know what what
do we think is missing between what a
panel of people like us would say and
data users for example?
I mean, I think
just just some off-the- cuff thoughts.
Um, I think that from
in in my prior life as a researcher, um,
infrastructure I think that works really
well is infrastructure that we can kind
of take for granted because it's
seamless. And so all of the kind of
people power, all of the infrastructure,
all of the like data storage and where
that lives can seem kind of abstract
um when you are a user or a depositor.
And so there's things that you know
there's challenges I think that are
unseen in that in that sense. Um and
then I think that there's also something
that we've discussed and um
is kind of around like how do we think
about business models that enable
sustainability of these infrastructures
and what is the role of sort of paying
into these systems. So we've talked
about sort of in our discussions around
industry and they benefit quite a great
deal from you know openly available data
sets um and the infrastructures that
house those data sets and yet their
contributions where they have an ability
to make them doesn't really flow back
and so how do you kind of work through
that um so I think that there's just
these different dependencies and ways
that we interact with infrastructures
that like when everything is working
great and smoothly, you take it for
granted and you're like, "Of course,
this thing will be here forever." Um,
but uh I I think and so all of the work
that goes on behind the scenes really
remains invisible. I don't know if any
of that that makes sense, but I think
we're also thinking about like how do
you how do you make these pieces more
visible and bring them into an advocacy
strategy that is more unified? Um so
that you know people who are depositing
into those systems that they that are
really important to them or who are
drawing data from systems that are
really important to them are able to
sort of like advocate for the need to
sustain them even though they're not
going to be talking as much in the weeds
about like some of the the nitty-gritty
details that an infrastructure person
might. Um so I don't know if anybody
else on this on this call has thoughts.
I mean, probably, but I'll be quiet now.
Okay, I'll call that question done. Um,
are there other questions that
>> Okay.
>> Oh, sorry, Miriam. I was just going to
say I think you had the perspective uh
succinctly nailed when you said they
just want it to be there.
That pretty much yeah covers it. I mean
think of how we use Google right or I
guess now it would be you know choose
your AI tool clean interface just works
hides all the abstraction.
>> Linda go ahead.
>> Can I answer Wanda's question?
>> Yes. Yeah. Yeah. Yeah. So I I I just
want to so Wanda is asking um can you
give us in simple terms how this
benefits? And so um I think this is a
great question and I I personally see
two major benefits in this. one is is
the coordination around these key
pillars um and providing some kind of
organizing framework that moves us
forward rather than kind of the ad hoc
efforts that we've been doing in the
past and so um I was very excited to
join Brett because of that um because I
think that we need to have these a plan
for how we move forward together um and
coordinate amongst all the different
groups that are out there. Um, two, I
think where I see a lot of importance
from this is like filling in the gaps,
figuring out where we have a lot of
tools that are being created. We have a
lot of people doing work, but we need to
figure out where the gaps are in that
work. Um, and so that's a a big part of
what is trying to do is understand where
the gaps and an example of this I think
for me is in thinking through the
advocacy toolkits. um because we have
been asked for advocacy toolkit for um
legislation that could be created to
protect data and that's not something
that we have. So trying to identify
places that we have gaps that um that
could be uh uh uh answered not just by
us but by groups that are working in um
this effort. So hopefully one of that
helps to answer your question
some.
>> And I think that was a great question to
to end on. Um since we just have a
minute left. Um so want to thank all of
you for taking the time uh on a Friday
in August to join us um and sharing uh
again that you can share your thoughts
with us at data uh resilience.io
io as well as in that um Google form
that I just shared. That form will have
a link to the draft strategic plan. Um,
which I will also if I am fast enough uh
drop that into the chat. Um, and uh we
look forward to this has already been
such a rich discussion. Um, and the plan
has only been out for a week now and so
I'm really excited for the kinds of
feedback and thoughts um and energy that
we get back from this community. Um so
please if you can join me in thanking uh
as well Chris, Linda, Alex and Christine
as well as um I see in the audience uh
some of our fellow credi strategic plan
committee members. Um so thank you for
all of uh your efforts as well as our
COS comm's team for being behind the
scenes and making it all go smoothly.
Um, so please uh see um look out for a
follow-up email with a recording to this
um and links uh and we look forward to
hearing more of your feedback and
thoughts. Have a great rest of your day.