Submind YouTube summaries
Thumbnail for Sustainability Lessons from the Scholarly Communications Trusted Dataspace

Sustainability Lessons from the Scholarly Communications Trusted Dataspace

Watch on YouTube

Video summary

The presentation outlines the journey and sustainability challenges faced by the Scholarly Communications Trusted Dataspace, a project initiated six years ago to centralize usage data from scholarly publishers and libraries. Initially conceived as a centralized repository, the project shifted to a decentralized approach due to significant hurdles regarding data privacy, particularly concerning Personally Identifiable Information (PII) across different geographical boundaries. The core architecture is divided into three distinct layers: the governance layer, which manages legal authority, contracts, and policy compliance; the control layer, comprising the technical infrastructure like data connectors and catalogs; and the data layer, which handles tokenized data transfers. A critical realization during the pilot phase was that relying solely on usage metrics for scholarly works was not broad enough to justify the operational costs, prompting a strategic rebranding and an exploration of diverse use cases such as sales data analysis, logistics management, and emissions tracking. Financial sustainability emerged as a primary concern, with the project team developing detailed budget models to determine break-even points using metrics like Cost of Goods Sold (COGS). The analysis revealed that revenue streams traditionally relied upon in the scholarly sector, such as memberships and sponsorships, were insufficient in current US and UK markets, leading the board to adopt a flat-fee pricing model for participants. However, even with these models, achieving financial viability required a "critical mass" of participants; without a sufficient number of organizations joining the ecosystem, the value proposition diminished significantly. The team also identified that smaller and medium-sized enterprises often lacked the internal capacity to self-host technical connectors or manage onboarding, creating a barrier to entry that could only be overcome through subsidized public infrastructure or external support mechanisms like the Data Space Accelerator model used by automotive industries. Ultimately, the project concluded its grant-funded phase after realizing it was two years ahead of market readiness and unable to secure the necessary resources to launch a fully operational service without significant external investment. The speaker emphasized that while volunteer efforts or purely open-source community models are appealing, they do not eliminate the need for professional coordination, legal oversight, and administrative support required to run a trusted data space. Key lessons learned include the necessity of bridging the gap between minimum viable products and fully operational services through smart scaling, leveraging existing public infrastructure like EGI in Europe to reduce costs, and securing vested partners willing to invest time and money despite opportunity costs. The presentation serves as a cautionary yet informative overview for others attempting similar initiatives, highlighting that successful data spaces require more than just technical innovation; they demand robust governance structures, realistic financial modeling, and strategic partnerships to overcome the "chicken and egg" problem of building a community before there is a product to offer.
Read the full video transcript
Thank you so much, and thank you all for having me. Um, it is uh such a privilege to be able to speak to you today about the the journey we've been on, um, and I'll say continue to be on. So, you're going to hear a number of updates uh in regards to our project, and I still believe everyone here was at the workshop, um, so just a couple quick notes as I get started. Um, what I wanted to focus on today was to really hone in on our sustainability work that we did. Um, I'm going to reference a lot of our other work, our technical pilot, our technical infrastructure that we build uh for the data space, and all of what we have is either in GitHub or in Zenodo, um, because we were funded through the Mellon Foundation, and we very much want to share back to the community. Everything I'll be talking about is published in our Zenodo community, so I can drop a link to that in the chat when I'm done here. Um, hopefully everyone knows what data spaces are. I'm going to use a lot of terminology as I jump through this talk, but the the key thing that I want to stress as you see this image here is that difference between the governance layer, that coordinating office that coordinates all of the participants, making sure that organizations have the legal authority and are trusted. So, if a a machine comes in under the auspices of an organization, that agent or bot or connector is tied back to an actual organization. So, all of that coordination among participants, um, including revenue generation, has to be managed somewhere. I'm going to refer to that as the governance layer. This is where our board sits, our advisors, all of that work. Um, and we'll go into that in a moment, but I want to separate that from the control layer, which is the actual technical infrastructure where those data connectors sit, the data catalogs, the discovery layer, and the request itself, which I am separating yet again from the actual tokenized data transfer on that bottom layer. So, hopefully everyone's used to uh those references to data layer, control layer, and governance layer, but I just wanted to start with that quick footnote. You do see also on the slide that the Zenodo link to our sustainability and budget model report is here for your reference. I'll come back to it at the end of the presentation, but a lot of what I'm sharing today is documented in much more detail in that report. So, let's see. There you go. So, a little bit of background about our data space. If I we didn't have a chance to meet when I was in Brisbane, we actually are an effort that started out as a data commons where we had a number of scholarly publishers, libraries, and platforms looking to centralize and aggregate usage data. And at the beginning, this is about six years ago now, everyone thought we could do this through an ingest or harvesting approach to have a centralized repository that could serve up those analytics. Of course, as you might expect that presented challenges when some, especially our corporate partners, had reservations about just sharing copies of this data, especially when some of that data is regulated as PII in some countries and not others, and shared across geographical boundaries. This shifted us into this decentralized data space approach, which is what we tagged in on. I will note that for our community, things that really preface this is they wanted to be able to control those flows of usage data. And when I say usage data for those who may not be familiar with it, we're working in scholarly communications where e-publications, of course, leave those digital footprints. Everything from the web analytics to views and downloads to, in some cases, behavioral data. Everything from eye tracking data in some cases. And so, there are these questions of, well, can we gain insights from that without violating readers' privacy? And that's where a lot of this comes from. It's like, well, how would we do that? And we already have standards to build on. Streamlining that information for the standardized privacy protective stats that we get today, they're called counter metrics, was where this project started. Challenges with what we had today really pushed us forward. So again, for those who don't know, a lot of it is because this data is spread across all of the platforms, uh repositories, libraries around the world who have copies of open scholarship, um or various uh digital copies of even version scholarship. There are variable standards, uh levels of standards adoption across all of those data providers, and currently our ecosystem works with a one-to-one custom API connections between platforms and any of the reporting and analytics firms that work with those organizations that are providing the data. Um AI has just added uh fuel to this fire, if you will, and of course it needs high-quality structured labeled data to control costs. And this is where a lot of our work in the data space come in. And I want to note there are two different value propositions we surfaced in our research. And one is framing this in terms of the ecosystem or the discipline. So when we take this kind of collective approach, we're thinking across scholarly communications or publishers. And there the ask really was if we could unlock some of this data, um so it could be more real-time, better quality, and more contextual than what we can get today with the aggregated counter statistics. Um but also building in those policy checks to make sure that it wasn't just access or identity-based access. Um it was that it was attribute-based. So maybe it's you could only use it so many times or from a certain geographical location that has a policy match. If it's read-only and not write. So there are um lots of different attributes we ended up putting into this. Um we didn't get to test in the pilot, but we built the structure for that in our rulebook. That's different from the data provider's perspective, which really surfaced as how do we control and audit the use of our data that we are sharing through the data space and how do we do that downstream? So again, we'll talk about kind of the governance layer versus those technical connectors and how that all came together from a budgetary perspective. I'll note that we did successfully complete last year our proof of concept and we actually have all of this up in GitHub and so I'm leaving the QR codes here and I'll make sure that I give a copy of this PDF to Andy after this call so he can share this, but you can actually see what was in our rulebook, how we did all of our documentation to walk folks through what we built which was on an AWS stack using Keycloak to manage the secrets if you will. We also worked with Opera's EU as our coordinating office to pilot the the contractual side of things. What does it mean for a participating organization to join the data space? What does what are they agreeing to? What service level should they expect? How are they expected to comply with our own rulebook that we created custom for our data space? So we were able to walk through and pilot all of those governance mechanisms as well as what our boards and our committee structures look like and what costs are associated with that and that's something I'm going to come back to because supporting that and coordinating that takes effort and it can vary in terms of what kinds of capacity. We also learned a lot through this pilot and you'll see here we have case study cited. We were lucky to have support from Mellon to do case studies for each of our participating organizations, some of whom were data providers, some were data recipients, some were both to understand what their perspective was as they were integrating the data space and where they saw value because at the end of the day our question really was what would it take for them to engage? Do they see a return on their investment? Because from the participating organizations' perspective, it required them pulling their technical teams off of other projects. There is an opportunity cost to engage in addition to any technical infrastructure adaptation they had to do. Um not to mention the very tall ask, I think, of learning what data spaces are. Because all of them were like, "What is this? I don't know. What kind of custom code do I need to build?" And so we had to start from scratch with education, which actually represented a lot of this. Um if you go into the case studies, you can see how far that ranged. Some of our real technical folks, they could do this in 15 minutes and they understood what it was. Others needed over a dozen hours, uh dozen to actually a couple dozen, just to have that onboarding and upskilling time. Key findings from this ROI piece uh really was one around critical mass. We had participants tell us that there wasn't value in coming to such an ecosystem if they only got 5% of their partners, even 10%. They wanted to know that the industry was there and that it made sense to connect cuz they could shift most of their custom APIs into the data space. And um in addition to that, we did hear a lot about those barriers to connect, especially from our small-to-medium organizations who needed a more hosted solution. The other thing that we heard loud and clear is that the singular use case of usage metrics for scholarship was not broad enough to justify the costs associated with the data space. So this actually was something we realized about a year and a half ago, um prompting, as you heard Andy mention, the rebranding from Open Access Book Usage Data Space to the Scholarly Communications Data Space. And over the past year, we've been having different conversations, World Data Systems book industry study group, um to explore what are those other use cases. And folks are talking about everything from hey, can we do better job with sales data? Can we do a better job with logistics and inventory management? Um one of as you probably all know with data spaces in Europe, um emissions tracking for SKUs. What are the kind of emissions are you creating by having the supply chain operational? All of these things are things that even as scholarly communications folks are thinking about hey, can we get some of this because now we can unlock some of the private data to get the aggregate. Um even knowing that though, those are individual use cases we'd have to build out in the data space and our pilot was very limited. Our value proposition research, uh we were able to do interviews and I just share some of these screenshots here to give folks ideas um to create those pitch decks for what would be our early adopters um because the pitch we'd make to a commercial publisher would look different than a library consortia, would look different than a small independent uh book publisher. And so for us we had to go through that and we piloted these materials because our community originally thought there would be memberships supported by sponsorships. Um and the unfortunate news is that um in this time frame with the pilot um under completion, our direct stakeholder funding was only able to raise 30,000 from three partners. And that was an that that was direct investment, right? Here is a $10,000 gift. In addition, and this is the big challenge, they were also valuing their internal staff time. What does it take for our team to invest the time to rework our own infrastructure? And so um we also looked at membership and you're going to see here in a moment we couldn't launch a membership program because for membership there had to be return on investment and that was seen as having the active data space up and running that people could connect to that brought value. And so we ended up very much in a chicken and an egg scenario where the community wasn't in there for us to have membership generate value. Um but yet we needed those resources. The good news is we were able to validate a number of the elements of our business model looking at not only outreach channels and value propositions, but some of those cost recovery mechanisms. Building on many of the workshops we did with our community to make sure that those cost recovery mechanisms and that revenue generation would be trusted as the data space launches. And there's separate things in our Zenodo community uh reports on that process if you're interested. Um the unfortunate news that I get to share and I feel very uh cheeky about this at the moment. I get to be unemployed right now because we jumped off the cliff. We're in that middle zone here uh where we are now in the innovation valley. Our grant ran out in June. And so we knew where that cliff was. Um and we have a lot of stakeholders who are very much interested in this. Um and the thing I kept consistently hearing is we're just not ready. We're 2 years ahead of our time. Once we understand how AI plugs into this, it'll be easier to justify. And once we have a further development internally in terms of our data governance organization and how we structure our data, it will be easier to participate. So now that we know we need to both support that community. Um we have critical mass, but we also got to figure out how do we get to that critical mass. That knowing that, I'm going to tell you that okay, here's all the fun numbers and how we got to the numbers in case this can help you on your journey. Um because I think we're all still very much learning. And um I don't know if this community has heard the term COGS before. This comes out of the the business side of the world. Um but when doing budget modeling, one things we looked at was what our cost of goods sold is. And in the business world, this is a very useful way of determining where your break even point is for operations. So, as we talk about sustainability, one of the things we want to do is, "Hey, can our can our revenues coming in cover the costs that we're spending for a data space?" That's our break even point. And what does that look like if we get to a fixed fee model, which I'll show you here in a moment. Um, but the key is I want to differentiate between the costs associated with the technical stack, with that control layer in the middle, from which is the service we are providing, from the governance layer that we have to have on top of that, the marketing, the board, the rule book development, the road map development. All of that's really important. But you could take all that away and still probably run your service for a while. So, I just want to differentiate those. Um, and if you look up COGS or cost of goods sold, you can find all kinds of calculators for that. So, with that, let me break down the costs that we documented. Again, this is all spelled out in a report in case it's useful. Um, the governance layer, just to highlight here, things that we recognized as we went through our pilot and documented the costs that we had. Um, from the coordinating office, and again, we're piloting this with OBRAS EU, uh, we knew we had to have administrative support. Everything to support contracting, to HR, to back office IT. Um, we needed to have legal consultation available. Um, not only if there were issues because we're working across countries, but also, um, with respect to if something came up with our rule book, if there was an issue with an SLA, if there if we got into a scenario where we were active and one data recipient had an issue with a data provider and we needed to pull in legal, we had to have that accessible. Um, making sure that we had SLA compliance staffing for that issue resolution process was important, in addition to all of the outreach and engagement costs that you might expect. Um, research and development, grant writing, fundraising, these all fall into this governance layer, as well as those indirect costs, um, taxes, VAT in Europe, for example, are things that we had to include in our budget model. From a staffing perspective, we went through a very useful exercise in, um, and you see my little valley of death bridging. We're like, "Okay, we know we need to bridge to a point where we hit break even. What do those bridges look like from a staffing perspective and capacity?" Um, cuz you can do different things with different levels. So, we ran and created these scenarios at a high, medium, and low level, um, for the budget. And if you I I it's a little fuzzy, I'm sorry. You can see that the staffing ranges everything from 100% down to 75%, uh, I'm sorry, uh, three FTEs at 100% down to one FTE, um, less than 100%. But it includes everything like some have travel, some have conferences. Some of these do not. Um, the lowest model really was a keep the lights on. Uh, so there would be less costs, and then you can learn and scale, but of course that comes with risk. So, I'll note that ultimately for us, we did have the mid-range as our target, um, and we our our board elected to close up shop earlier this year when we realized we would not be able to make that happen. We wouldn't be able to meet those numbers, um, to give us that 6-month wind-down period, um, in conjunction with our community. The control layer, so going back to the technology stack in the middle, um, the costing for this is highly variable. Um, and it's so dependent on who your data providers are and who your data recipients are and how they're using the data space. Um, the central infrastructure for user credentialing, for, you, we used Keycloak to manage identities, policy clearing are all the pieces in there are all the pieces that you're checking for per your rule book, configured and run to generate those tokens. That depends on complexity. And so for our initial use case, we had numbers, but we recognized if we added in different use cases, sales data, emissions data, you name it, that those could have very different costs based on the cloud compute necessary. Um that said, the other thing we learned was that the So, the data connectors themselves can be self-hosted. Many of our larger enterprises or commercial enterprises had the capacity, both the team and the cloud compute available to self-host those data connectors. But that wasn't the case with our nonprofit partners or small-to-medium enterprises and in many cases, um some of our smaller public institutions. Explaining what this is and why they're connecting is something we had to overcome. Um similarly, onboarding support for smaller-to-medium organizations was something we needed to recognize um and plan towards and plan for. So, we actually included in our budget a hosted solution, recognizing that there could be a strategic partner that emerges down the road that could play that role. Um the other thing I'll note is at the moment with for our pilot, when we modeled these costs, we recognized that an NREN or public infrastructure could of course be leveraged. We were using a direct um contractor who had their own AWS um contract. And so we were doing flow-through pricing through them, which we could reduce if we had um subsidized public access or access through an NREN, for example. Um on the revenue generation side, we spent our first two, three years identifying what our community saw as valued, trusted sources of revenue. Um everything from sponsorships to donations to grants, to memberships. And over the time of the that we took, so I would say across the 6-year span, we started memberships were the way to fund scholarly infrastructure. That is not the case today in the US market and the UK market. Um that shifted entirely. And for us right now in the US market, um it's really hard to get membership dollars out of university at the moment. Put it that way. And for any of our commercial partners. So there's some flexibility there, but we had a complex budget model built that was like, "Hey, maybe we'll get some membership, some sponsorships. Um we'll have consortia join and some open infrastructure funding as well." That proved to be too complex out of the gate. And our board ultimately opted for flat-fee pricing. What if we say to join the data space as a participant, you have to have a flat fee of either 10,000 or 50,000, um euro. And we ended up using those numbers to model what an operational break-even would be. So we took the costs from our pilot um for that singular use case, and we ran the numbers against that. And this is where cost of goods sold, the cost to provide the data space service comes in. And you'll see things where you'll see the the fixed costs, we included a development build. We knew there were things we still had to build out before we went operational. Um but the numbers look very different. If we had 50,000 an org, we would need 16 to break even versus over 1,000 at 10. There are questions of equity here. There's questions of subsidy. How do we make this work? But knowing these numbers helped us provide a frame for conversations to to have this conversation, what does it take and how do we get there? Because before we got to this point, people didn't understand, well how many are we talking about? Can we make it I I heard 2,000 to start. Can everyone just pay 2,000? I was like, that'd be great, but it'll be it'll take a long time for us to get enough to launch a service. Um so, in closing, I just wanted to note a couple other uh high-level points. One, um keeping in mind the difference between the minimum viable and the fully operational. We know that there would be a number of additional costs that would come on board, and we're trying to scale smartly and start small so that we didn't build the huge thing and then realize we didn't need all of it. Um there is, of course, variability that will come with that from the use case and how participants use the data space. How many transactions, how frequently, if it's an AI if it's an agent talking to an agent, that can scale rapidly out of control without limits. And so, thinking through all of that and what that means for the budget model is really important. Um for us, one of our findings, I would say, is that startup capital or critical mass is needed to open the door. One of the two to bridge that gap. Um and without the critical mass, you have to have really vested partners who can do that alongside what they're doing today. Um I think that's one of the key findings there for us. Um the value of leveraging public infrastructure, it was on our road map for our next phase to use EGI and some of the infrastructures in Europe, which we could piggyback on. Um I think it's something definitely to think about as you're looking at your own national infrastructures, what's there to bring the cost down. Um because you already have access to it. Um the other thing that I always like to surface, because a lot of folks have said, "Hey, can we just do this as a volunteer effort? Can we all chip in? Can we make it open source and have a community?" But that still takes coordination effort. You would still require the legal effort in the back office. Um so, that isn't a complete solution in and of itself. Um and the other thing I just want to close is something that I discovered here in the past month or so, cuz this was uh advertised um hit my radar because I live about 3 hours away from Ford Motor Company and they're a member of Catena-X. And the way that a US automotive maker got involved in Catena-X is through the data space accelerator. They were able to participate and offer these 15 to 30,000 euro subsidies to their supply chain partners to onboard. And if you haven't heard of the data space accelerator and looked at that as a model, I think that's a really strong model for overcoming what we found was one of those key barriers. How do you justify the cost of connecting to the data space when you have to reassign your staff and manage all of your current workload alongside. So, I left links to the news story about Catena-X and the Zenodo link here, too, but that's what I have. I hope that's helpful as an overview for those of you who haven't had the chance to hear me do different versions of this in the past few months.