Sustainability Lessons from the Scholarly Communications Trusted Dataspace
Watch on YouTubeVideo summary
The presentation outlines the journey and sustainability challenges faced by the Scholarly Communications Trusted Dataspace, a project initiated six years ago to centralize usage data from scholarly publishers and libraries. Initially conceived as a centralized repository, the project shifted to a decentralized approach due to significant hurdles regarding data privacy, particularly concerning Personally Identifiable Information (PII) across different geographical boundaries. The core architecture is divided into three distinct layers: the governance layer, which manages legal authority, contracts, and policy compliance; the control layer, comprising the technical infrastructure like data connectors and catalogs; and the data layer, which handles tokenized data transfers. A critical realization during the pilot phase was that relying solely on usage metrics for scholarly works was not broad enough to justify the operational costs, prompting a strategic rebranding and an exploration of diverse use cases such as sales data analysis, logistics management, and emissions tracking.
Financial sustainability emerged as a primary concern, with the project team developing detailed budget models to determine break-even points using metrics like Cost of Goods Sold (COGS). The analysis revealed that revenue streams traditionally relied upon in the scholarly sector, such as memberships and sponsorships, were insufficient in current US and UK markets, leading the board to adopt a flat-fee pricing model for participants. However, even with these models, achieving financial viability required a "critical mass" of participants; without a sufficient number of organizations joining the ecosystem, the value proposition diminished significantly. The team also identified that smaller and medium-sized enterprises often lacked the internal capacity to self-host technical connectors or manage onboarding, creating a barrier to entry that could only be overcome through subsidized public infrastructure or external support mechanisms like the Data Space Accelerator model used by automotive industries.
Ultimately, the project concluded its grant-funded phase after realizing it was two years ahead of market readiness and unable to secure the necessary resources to launch a fully operational service without significant external investment. The speaker emphasized that while volunteer efforts or purely open-source community models are appealing, they do not eliminate the need for professional coordination, legal oversight, and administrative support required to run a trusted data space. Key lessons learned include the necessity of bridging the gap between minimum viable products and fully operational services through smart scaling, leveraging existing public infrastructure like EGI in Europe to reduce costs, and securing vested partners willing to invest time and money despite opportunity costs. The presentation serves as a cautionary yet informative overview for others attempting similar initiatives, highlighting that successful data spaces require more than just technical innovation; they demand robust governance structures, realistic financial modeling, and strategic partnerships to overcome the "chicken and egg" problem of building a community before there is a product to offer.
Read the full video transcript
Thank you so much, and thank you all for
having me. Um, it is uh such a privilege
to be able to speak to you today about
the the journey we've been on,
um, and I'll say continue to be on. So,
you're going to hear a number of updates
uh in regards to our project, and I
still believe everyone here was at the
workshop, um, so just a couple quick
notes as I get started. Um, what I
wanted to focus on today was to really
hone in on our sustainability work that
we did. Um, I'm going to reference a lot
of our other work, our technical pilot,
our technical infrastructure that we
build uh for the data space, and all of
what we have is either in GitHub or in
Zenodo, um, because we were funded
through the Mellon Foundation, and we
very much want to share back to the
community. Everything I'll be talking
about is published in our Zenodo
community, so I can drop a link to that
in the chat when I'm done here. Um,
hopefully everyone knows what data
spaces are. I'm going to use a lot of
terminology as I jump through this talk,
but the the key thing that I want to
stress as you see this image here is
that difference between the governance
layer, that coordinating office that
coordinates all of the participants,
making sure that organizations have the
legal authority and are trusted. So, if
a a machine comes in under the auspices
of an organization, that agent or bot or
connector
is tied back to an actual organization.
So, all of that coordination among
participants, um, including revenue
generation, has to be managed somewhere.
I'm going to refer to that as the
governance layer. This is where our
board sits, our advisors, all of that
work. Um, and we'll go into that in a
moment, but I want to separate that from
the control layer, which is the actual
technical infrastructure where those
data connectors sit, the data catalogs,
the discovery layer, and the request
itself, which I am separating yet again
from the actual tokenized data transfer
on that bottom layer. So, hopefully
everyone's used to uh those references
to data layer, control layer, and
governance layer, but I just wanted to
start with that quick footnote. You do
see also on the slide that the Zenodo
link to our sustainability and budget
model report is here for your reference.
I'll come back to it at the end of the
presentation, but a lot of what I'm
sharing today is documented in much more
detail in that report.
So, let's see.
There you go. So, a little bit of
background about our data space. If I we
didn't have a chance to meet when I was
in Brisbane,
we actually are an effort that started
out as a data commons where we had a
number of scholarly publishers,
libraries, and platforms looking to
centralize and aggregate usage data.
And at the beginning, this is about six
years ago now, everyone thought we could
do this through an ingest or harvesting
approach to have a centralized
repository that could serve up those
analytics.
Of course, as you might expect that
presented challenges when some,
especially our corporate partners, had
reservations about just sharing copies
of this data, especially when some of
that data is regulated as PII in some
countries and not others, and shared
across geographical boundaries. This
shifted us into this decentralized data
space approach,
which is what we tagged in on. I will
note that for our community, things that
really preface this is they wanted to be
able to control those flows of usage
data.
And when I say usage data for those who
may not be familiar with it,
we're working in scholarly
communications where e-publications, of
course, leave those digital footprints.
Everything from the web analytics to
views and downloads to, in some cases,
behavioral data.
Everything from eye tracking data in
some cases. And so, there are these
questions of, well, can we gain insights
from that without violating readers'
privacy? And that's where a lot of this
comes from. It's like, well, how would
we do that? And we already have
standards to build on. Streamlining that
information for the standardized privacy
protective stats that we get today,
they're called counter metrics, was
where this project started.
Challenges with what we had today really
pushed us forward. So again, for those
who don't know, a lot of it is because
this data is spread across all of the
platforms, uh repositories, libraries
around the world who have copies of open
scholarship,
um or various
uh digital copies of even version
scholarship.
There are variable standards, uh levels
of standards adoption across all of
those data providers, and currently our
ecosystem works with a
one-to-one custom API connections
between platforms and any of the
reporting and analytics firms that work
with those organizations that are
providing the data. Um AI has just added
uh fuel to this fire, if you will, and
of course it needs high-quality
structured labeled data to control
costs. And this is where a lot of our
work in the data space come in. And I
want to note there are two different
value propositions we surfaced in our
research. And one is framing this in
terms of the ecosystem or the
discipline. So when we take this kind of
collective approach, we're thinking
across scholarly communications or
publishers. And there the ask really was
if we could unlock some of this data,
um so it could be more real-time, better
quality, and more contextual than what
we can get today with the aggregated
counter statistics. Um but also building
in those policy checks to make sure that
it wasn't just access or identity-based
access.
Um it was that it was attribute-based.
So maybe it's you could only use it so
many times or from a certain
geographical location that has a policy
match. If it's read-only and not write.
So there are um lots of different
attributes we ended up putting into
this. Um we didn't get to test in the
pilot, but we built the structure for
that in our rulebook. That's different
from the data provider's perspective,
which really surfaced as how do we
control and audit the use of our data
that we are sharing through the data
space
and how do we do that downstream?
So again, we'll talk about kind of the
governance layer versus those technical
connectors and how that all came
together from a budgetary perspective.
I'll note that we did successfully
complete last year our proof of concept
and we actually have all of this up in
GitHub and so I'm leaving the QR codes
here and I'll make sure that I give a
copy of this PDF to Andy after this call
so he can share this, but you can
actually see what was in our rulebook,
how we did all of our documentation to
walk folks through what we built which
was on an AWS stack using Keycloak to
manage the secrets if you will.
We also worked with Opera's EU as our
coordinating office
to pilot the the contractual side of
things. What does it mean for a
participating organization to join the
data space? What does what are they
agreeing to? What service level should
they expect? How are they expected to
comply with our own rulebook that we
created custom for our data space? So we
were able to walk through and pilot all
of those governance mechanisms as well
as what our boards and our committee
structures look like and what costs are
associated with that and that's
something I'm going to come back to
because supporting that and coordinating
that takes effort and it can vary in
terms of what kinds of capacity. We also
learned a lot through this pilot and
you'll see here we have case study
cited.
We were lucky to have support from
Mellon to do case studies for each of
our participating organizations, some of
whom were data providers, some were data
recipients, some were both to understand
what their perspective was as they were
integrating the data space and where
they saw value because at the end of the
day our question really was what would
it take for them to engage? Do they see
a return on their investment? Because
from the participating organizations'
perspective, it required them pulling
their technical teams off of other
projects. There is an opportunity cost
to engage in addition to any technical
infrastructure adaptation they had to
do. Um not to mention the very tall ask,
I think, of learning what data spaces
are. Because all of them were like,
"What is this? I don't know. What kind
of custom code do I need to build?" And
so we had to start from scratch with
education, which actually represented a
lot of this. Um if you go into the case
studies, you can see how far that
ranged. Some of our real technical
folks, they could do this in 15 minutes
and they understood what it was. Others
needed over a dozen hours, uh dozen to
actually a couple dozen, just to have
that onboarding and upskilling time.
Key findings from this ROI piece uh
really was one around critical mass. We
had participants tell us that there
wasn't value in coming to such an
ecosystem if they only got 5% of their
partners, even 10%. They wanted to know
that the industry was there and that it
made sense to connect cuz they could
shift most of their custom APIs into the
data space. And um in addition to that,
we did hear a lot about those barriers
to connect, especially from our
small-to-medium organizations who needed
a more hosted solution.
The other thing that we heard loud and
clear is that the singular use case of
usage metrics for scholarship was not
broad enough to justify the costs
associated with the data space.
So this actually was something we
realized about a year and a half ago, um
prompting, as you heard Andy mention,
the rebranding from Open Access Book
Usage Data Space to the Scholarly
Communications Data Space. And over the
past year, we've been having different
conversations, World Data Systems book
industry study group, um to explore what
are those other use cases. And folks are
talking about everything from hey, can
we do better job with sales data? Can we
do a better job with logistics and
inventory management? Um one of as you
probably all know with data spaces in
Europe, um emissions tracking for SKUs.
What are the kind of emissions are you
creating by having the supply chain
operational?
All of these things are things that even
as scholarly communications folks are
thinking about hey, can we get some of
this because now we can unlock some of
the private data to get the aggregate.
Um
even knowing that though, those are
individual use cases we'd have to build
out in the data space and our pilot was
very limited.
Our value proposition research, uh we
were able to do interviews and I just
share some of these screenshots here to
give folks ideas um to create those
pitch decks for what would be our early
adopters um because the pitch we'd make
to a commercial publisher would look
different than a library consortia,
would look different than a small
independent uh book publisher. And so
for us we had to go through that and we
piloted these materials because
our community originally thought there
would be memberships supported by
sponsorships.
Um and the unfortunate news is that um
in this time frame with the pilot um
under completion, our direct stakeholder
funding was only able to raise 30,000
from three partners. And that was an
that that was direct investment, right?
Here is a $10,000 gift. In addition, and
this is the big challenge, they were
also valuing their internal staff time.
What does it take for our team to invest
the time to rework our own
infrastructure? And so um we also looked
at membership and you're going to see
here in a moment we couldn't launch a
membership program because for
membership there had to be return on
investment and that was seen as having
the active data space up and running
that people could connect to that
brought value. And so we ended up very
much in a chicken and an egg scenario
where the community wasn't in there for
us to have membership generate value. Um
but yet we needed those resources.
The good news is we were able to
validate a number of the elements of our
business model looking at not only
outreach channels and value
propositions, but some of those cost
recovery mechanisms. Building on many of
the workshops we did with our community
to make sure that those cost recovery
mechanisms and that revenue generation
would be trusted as the data space
launches. And there's separate things in
our Zenodo community uh reports on that
process if you're interested. Um the
unfortunate news that I get to share and
I feel very uh cheeky about this at the
moment. I get to be unemployed right now
because we jumped off the cliff. We're
in that middle zone here uh where we are
now in the innovation valley. Our grant
ran out in June. And so we knew where
that cliff was.
Um and we have a lot of stakeholders who
are very much interested in this. Um and
the thing I kept consistently hearing is
we're just not ready. We're 2 years
ahead of our time. Once we understand
how AI plugs into this, it'll be easier
to justify. And once we have a further
development internally in terms of our
data governance organization and how we
structure our data, it will be easier to
participate. So now that we know we need
to both support that community. Um we
have critical mass, but we also got to
figure out how do we get to that
critical mass.
That knowing that, I'm going to tell you
that okay, here's all the fun numbers
and how we got to the numbers in case
this can help you on your journey. Um
because I think we're all still very
much learning. And um I don't know if
this community has heard the term COGS
before. This comes out of the the
business side of the world.
Um but when doing budget modeling, one
things we looked at was what our cost of
goods sold is. And in the business
world, this is a very useful way of
determining where your break even point
is for operations. So, as we talk about
sustainability, one of the things we
want to do is, "Hey, can our can our
revenues coming in cover the costs that
we're spending for a data space?" That's
our break even point. And what does that
look like if we get to a fixed fee
model, which I'll show you here in a
moment. Um, but the key is I want to
differentiate between the costs
associated with the technical stack,
with that control layer in the middle,
from which is the service we are
providing,
from the governance layer that we have
to have on top of that, the marketing,
the board, the rule book development,
the road map development. All of that's
really important. But you could take all
that away and still probably run your
service for a while. So, I just want to
differentiate those. Um, and if you look
up COGS or cost of goods sold, you can
find all kinds of calculators for that.
So, with that, let me break down the
costs that we documented. Again, this is
all spelled out in a report in case it's
useful. Um, the governance layer, just
to highlight here, things that we
recognized as we went through our pilot
and documented the costs that we had.
Um, from the coordinating office, and
again, we're piloting this with OBRAS
EU, uh, we knew we had to have
administrative support. Everything to
support contracting, to HR, to back
office IT. Um, we needed to have legal
consultation available. Um, not only if
there were issues because we're working
across countries, but also, um, with
respect to if something came up with our
rule book, if there was an issue with an
SLA, if there if we got into a scenario
where we were active and one data
recipient had an issue with a data
provider and we needed to pull in legal,
we had to have that accessible. Um,
making sure that we had SLA compliance
staffing for that issue resolution
process was important, in addition to
all of the outreach and engagement costs
that you might expect. Um, research and
development, grant writing, fundraising,
these all fall into this governance
layer, as well as those indirect costs,
um, taxes, VAT in Europe, for example,
are things that we had to include in our
budget model.
From a staffing perspective, we went
through a very useful exercise in, um,
and you see my little valley of death
bridging. We're like, "Okay, we know we
need to bridge to a point where we hit
break even. What do those bridges look
like from a staffing perspective and
capacity?" Um, cuz you can do different
things with different levels. So, we ran
and created these scenarios at a high,
medium, and low level, um, for the
budget. And if you I I it's a little
fuzzy, I'm sorry. You can see that the
staffing ranges everything from 100%
down to 75%,
uh, I'm sorry, uh, three FTEs at 100%
down to one FTE, um, less than 100%. But
it includes everything like some have
travel, some have conferences. Some of
these do not.
Um, the lowest model really was a keep
the lights on. Uh, so there would be
less costs, and then you can learn and
scale, but of course that comes with
risk.
So, I'll note that ultimately for us, we
did have the mid-range as our target,
um, and we our our board elected to
close up shop earlier this year when we
realized we would not be able to make
that happen. We wouldn't be able to meet
those numbers, um, to give us that
6-month wind-down period, um, in
conjunction with our community.
The
control layer, so going back to the
technology stack in the middle, um, the
costing for this is highly variable. Um,
and it's so dependent on who your data
providers are and who your data
recipients are and how they're using the
data space. Um, the central
infrastructure for user credentialing,
for, you, we used Keycloak to manage
identities, policy clearing are all the
pieces in there
are all the pieces that you're checking
for per your rule book, configured and
run to generate those tokens. That
depends on complexity. And so for our
initial use case, we had numbers, but we
recognized if we added in different use
cases, sales data, emissions data, you
name it, that those could have very
different costs based on the cloud
compute necessary. Um that said, the
other thing we learned was that the
So, the data connectors themselves can
be self-hosted. Many of our larger
enterprises or commercial enterprises
had the capacity, both the team and the
cloud compute available to self-host
those data connectors. But that wasn't
the case with our nonprofit partners or
small-to-medium enterprises and in many
cases,
um some of our smaller public
institutions. Explaining what this is
and why they're connecting is something
we had to overcome. Um similarly,
onboarding support for smaller-to-medium
organizations was something we needed to
recognize
um and plan towards and plan for. So, we
actually included in our budget a hosted
solution, recognizing that there could
be a strategic partner that emerges down
the road that could play that role. Um
the other thing I'll note is at the
moment with for our pilot, when we
modeled these costs, we recognized that
an NREN or public infrastructure could
of course be leveraged. We were using a
direct um contractor who had their own
AWS
um contract. And so we were doing
flow-through pricing through them, which
we could reduce if we had um subsidized
public access or access through an NREN,
for example.
Um on the revenue generation side, we
spent our first two, three years
identifying what our community saw as
valued, trusted sources of revenue.
Um everything from sponsorships to
donations to grants, to memberships. And
over the time of the that we took, so I
would say across the 6-year span,
we started memberships were the way to
fund scholarly infrastructure. That is
not the case today in the US market and
the UK market. Um that shifted entirely.
And for us right now in the US market,
um
it's really hard to get membership
dollars out of university at the moment.
Put it that way. And for any of our
commercial partners. So there's some
flexibility there, but we had a complex
budget model built that was like, "Hey,
maybe we'll get some membership, some
sponsorships.
Um we'll have consortia join and some
open infrastructure funding as well."
That proved to be too complex out of the
gate. And our board ultimately opted for
flat-fee pricing. What if we say to join
the data space as a participant, you
have to have a flat fee of either 10,000
or 50,000,
um euro.
And we ended up using those numbers to
model what an operational break-even
would be. So we took the costs from our
pilot
um for that singular use case, and we
ran the numbers against that. And this
is where cost of goods sold, the cost to
provide the data space service comes in.
And you'll see things where you'll see
the the fixed costs, we included a
development build. We knew there were
things we still had to build out before
we went operational. Um but the numbers
look very different. If we had 50,000 an
org, we would need 16 to break even
versus over 1,000 at 10.
There are questions of equity here.
There's questions of subsidy. How do we
make this work? But knowing these
numbers helped us provide a frame for
conversations to to have this
conversation, what does it take and how
do we get there? Because before we got
to this point, people didn't understand,
well how many are we talking about? Can
we make it I I heard 2,000 to start. Can
everyone just pay 2,000? I was like,
that'd be great, but it'll be
it'll take a long time for us to get
enough to launch a service. Um so, in
closing, I just wanted to note a couple
other uh high-level points. One, um
keeping in mind the difference between
the minimum viable and the fully
operational. We know that there would be
a number of additional costs that would
come on board, and we're trying to scale
smartly and start small so that we
didn't build the huge thing and then
realize we didn't need all of it. Um
there is, of course, variability that
will come with that from the use case
and how participants use the data space.
How many transactions, how frequently,
if it's an AI if it's an agent talking
to an agent, that can scale rapidly out
of control without limits. And so,
thinking through all of that and what
that means for the budget model is
really important. Um
for us, one of our findings, I would
say, is that startup capital or critical
mass is needed to open the door.
One of the two to bridge that gap. Um
and without the critical mass, you have
to have
really vested partners who can do that
alongside what they're doing today. Um I
think that's one of the key findings
there for us. Um the value of leveraging
public infrastructure, it was on our
road map for our next phase to use EGI
and some of the infrastructures in
Europe, which we could piggyback on. Um
I think it's something definitely to
think about as you're looking at your
own national infrastructures, what's
there to bring the cost down. Um because
you already have access to it. Um the
other thing that I always like to
surface, because a lot of folks have
said, "Hey, can we just do this as a
volunteer effort? Can we all chip in?
Can we make it open source and have a
community?" But that still takes
coordination effort. You would still
require the legal effort in the back
office.
Um so, that isn't a complete solution in
and of itself. Um and the other thing I
just want to close is something that I
discovered here in the past month or so,
cuz this was uh advertised um
hit my radar because I live about 3
hours away from Ford Motor Company and
they're a member of Catena-X. And the
way that a US automotive maker got
involved in Catena-X is through the data
space accelerator. They were able to
participate and offer these 15 to 30,000
euro subsidies to their supply chain
partners to onboard.
And if you haven't heard of the data
space accelerator and looked at that as
a model, I think that's a really strong
model for overcoming what we found was
one of those key barriers. How do you
justify the cost of connecting to the
data space when you have to reassign
your staff and manage all of your
current workload
alongside. So,
I left links to the news story about
Catena-X and the Zenodo link here, too,
but that's what I have. I hope that's
helpful as an overview for those of you
who haven't had the chance to hear me do
different versions of this in the past
few months.