Submind YouTube summaries
Thumbnail for OSTrails IF in practice: from Node experience to adoption in the EOSC Federation

OSTrails IF in practice: from Node experience to adoption in the EOSC Federation

Watch on YouTube

Video summary

The OSTrails project aims to establish robust interoperability frameworks essential for the EOSC Federation by standardizing data exchange through APIs and preventing vendor lock-in. This initiative is built upon three core pillars: planning via Machine Actionable Data Management Plans, tracking using Scientific Knowledge Graphs, and assessing quality through FAIR assessment tools. By developing specific frameworks such as DMP-IF, SKG-IF, and FAIR-IF, the project enables service operators to automate workflows and ensure seamless communication between repositories, knowledge graphs, and evaluation systems without creating isolated silos. Various national pilots have demonstrated the practical application of these frameworks across different contexts. In Finland, a reference data model was created that aligns with RDA standards while extending to cover software management plans and local data protection requirements. Poland integrated existing platforms like Argos and FairOS to facilitate bidirectional data exchange with OpenAire knowledge graphs, whereas the Netherlands focused on combining privacy, ethics, and compliance within integrated workflows. Additionally, the PaNOSC cluster utilized these tools to link research instruments and datasets with persistent identifiers, integrate external publication data, and perform automated assessments of metadata quality. The CLARIN pilot within the SHOCK cluster further illustrated the potential of the Scientific Knowledge Graph framework by mapping Virtual Language Observatory metadata to a core model and exposing it via API, distinct from other approaches that simply wrap existing services. During the discussion, experts emphasized that while FAIR assessment tools serve as evidence of readiness rather than an end goal, the strategic priority remains on knowledge graphs to transform nodes from simple directories into functional federated entities. Although some concerns were raised regarding the bureaucracy of machine-actionable data management plans, the consensus highlighted that interoperability frameworks are crucial for avoiding vendor lock-in and ensuring services can be onboarded effectively. The webinar concluded with a strong indication of community interest in adopting all three frameworks, supported by survey results showing a high likelihood of implementation across the federation. Rather than creating non-interoperable services or redundant systems, the session encouraged national nodes to collaborate on gradual implementation strategies and share their experiences. This cooperative approach is intended to foster a common future within the EOSC Federation, ensuring that data exchange remains standardized, automated, and free from the constraints of proprietary vendor ecosystems while addressing diverse institutional needs.
Read the full video transcript
Let me welcome you all. My name is Francisca de Jong. I'm one of the uh work package leaders in the project OS Trails. We organized this event to reach out to um uh in principle anyone anybody who is interested in uh the interoperability frameworks that are being developed within uh the project OS Trails. Um but we focused this webinar on uh people with an interest from the perspective of the uh emergence of the EOSC Federation and the nodes uh that uh uh will play an important role um uh that will be the the backbone of uh of the Federation. Um both candidate nodes and uh um first wave nodes have been invited uh as audience and uh we hope that this will webinar will um help you understand better what the OS Trails interoperability frameworks could bring to the EOSC Federation and how uh you could um uh adopt the results. It's uh uh the the presentations that we prepared are uh presenting to you first of all uh through a presentation by Thomas Miksa about the interoperability frameworks and then we have five presentations with pilot projects both from the thematic perspective and the perspective of national nodes. Uh practically um I would like to make you aware of the notes on this first slide as well, namely that this will be recorded. If you don't want to be visible in the recordings, please switch off your camera. We would also like to um be informed about the node to which you are associated if any. So, please uh adjust your name affiliation to make that clear. And after the webinar, you will be invited to share comments and questions about the interoperability frameworks of OS Trails so that we can arrange more tailored interactions with you in the near future. I think we are ready to start. 38 people right now in the meeting room. Welcome all that missed the very first sentence. I would like now to hand over to Thomas Mixa. >> Thank you very much. Welcome everyone. So I have this pleasure to give you an overview of the stuff we're doing in OS Trails. I am promising not to make this presentation too technical. You will hear about the details of implementing OS Trails from the pilots we have in our project and there are also people who are responsible for the specific interoperability frameworks here with us. Paolo Manghi for SKGIF, Mark Wilkinson for FAIR IF, and and Marek Suchanek for DMP IF. So in case I don't say anything relevant, I'm sure that they will correct me or if you have basically any questions at the end of the of the event. So I think every project presentation starts with this overall picture of what this project is doing and in OS Trails we have three main pillars. So we are dealing with planning data management, tracking use of data, and assessing or if you want to can call it evaluating whatever synonym you find then it's most best for you. What happens with the data? So for the planning part we work with machine actionable data management plans where we describe what will happen during the research. For the tracking of data, we work with SKGs. SKGs can be understood both as scientific knowledge graphs and scholarly knowledge graphs. And for assessing, we work with fair assessment, and we deal with questions, what does it mean to be fair in a specific context in a specific community? And what we're doing here, we're doing for EOSC, but the work we do here also remains universal outside of EOSC. So, today we are in the EOSC crowd, so we will be focusing on what is the relevance for the nodes, but we believe our work is universally universal universally applicable. What is also important to say about OS Trails, it's a relatively big project. You have seen that we have 40 partners, and we have lots of pilots. We have both national pilots, so organizations from specific European countries that are adopting the interoperability frameworks in their context, and we also have thematic pilots, which are often clusters, which do it for a specific community. And this is basically a test bed for the ideas we have and for the things I will be talking about today. As in every software engineering project, because this is the big part of of our project is basically finding out how to improve the existing software. We had to start with the requirements analysis, with defining the use case, what actually are we going to do? And by analyzing what our pilots have, what kind of tools they have, we came up with this pathways diagram. And we have few versions of the pathways diagram, which are a little bit more complicated, but in this in this simplified diagram, you can see four main boxes, namely repository, SKG, fair assessment tool, and DMP platform. And these are the four main objects that we are focusing on to standardize the interoperability among them. So, in this pathway specifically, you can see potential actions that we have right now between the systems, but also the actions you can have in the future. So, if you look at the repository at the top, we can say a repository might want to add a research output to a DMP platform. So, there is it is possible that we would like to make we would like to make it possible that whenever time a data set is published, and information is added to the data management planning platform to the data management plan, that you don't have to do it manually. This is basically what frustrates the users, and they want to have less work. And I think this is exactly the kind of a discussion you have for your notes, because you're providing different services, and so far as far as I know, many of the services are isolated. And if you want to create a value for the researchers, you want to start connecting them. And this is basically what we're trying to do here with all the strains. Of course, we're not trying to standardize everything. We can only focus on these aspects where we have something to say because we have people in the consortium who know something about it. And that's how we came up with the architecture of what we are doing. So, you have again these boxes. So, DMP platforms, SKGs, fair assessment tools, data repositories, and also the DMP evaluation service, but this is a different story. And we're saying basically that we are developing these interoperability frameworks, which provide the rules of what does it mean to be interoperable on the technical but also on the semantic level. So, here in this diagram, next to the DMP platform, you can see this lollipop symbol with an MADMP API next to it. For the SKG, you can see the SKG API next to it. And for the fair assessment tool, you can see this fair IF symbol basically, the label. And these are the communications that a specific tool provides this interface, provides the API, and all other tools can communicate with this tool using the language defined by the interoperability framework. So, as you can see at the data repository, it still has a custom API. So, we are not making any recommendations of what should be the way to talk to a repository. So, if you have Invenio, if you have DSpace, if you have Fedora, if you have anything else, they have their their own APIs, they can use them. If you have OAI-PMH endpoint, it is still okay to have it. But, here we're focusing on a set of endpoints and actions you can trigger if you want to implement things we have depicted in the pathways before. That's the big introduction about what we're doing as a project. Now, I would like to go into specific um interoperability frameworks. So, the first one we have uh deals with the data management planning is a DMP-IF. And I will explain it a bit more. And for the other frameworks, you will see that they are very similar in the structure of what we're doing. So, the goal for the DMP-IF is to Sorry, I can't see that I have uh is the standardize how information is exchanged between uh services in the wider RDM ecosystem. So, if any service needs information from the DMP or can provide information to the DMP, they should do it in the standardized way. If you look at the diagram I have on the right-hand side, this is an anatomy of a typical DMP tool. So, it can generate documents, it has some integrations with uh specific tools like I know ORCID or re3data or Fairsharing. Of course, it has a user interface, and it has its own internal data model. It has its own way of storing all this information. And what we're doing in our project TRAILS, we're saying, "Okay, map your internal model to something we agreed as a community, the so-called common data model, and expose it over API so that there are basic info basic operations that anyone can perform like search for DMPs, get a DMP, update a DMP, but not a static PDF document. We don't want this. We want machine actionable and data management plans. And by doing this this way, we also want to prevent the vendor lock-in. So, we don't want to say you must use only this specific tool. You can use any tool for data management planning in your node and if they stick to the DMP IF, they should be able to uh exchange information or you are able to basically replace these tools or provide to at the same time. It really depends on you, but the integration and the cost should be lower if you use this DMP IF to integrate it with other services. Uh the DMP IF and everything I'm describing here is not really meant for the researchers. Researchers should not know about its existence. This is for the service operator software engineers. This is to make their life easier and their life to uh help them automate tasks and if they automate the tasks, in the long term, researchers also have less work to do. And to become compliant with the DMP IF, you don't have to do everything at once. You can do this step-by-step and this will become clear when I show you my next slides. Um so, first of all, what we have is the common standard. It's a RDA recommendation. It specifies what is the common model. How does an MA DMP look like? Then, in OS Trails, we have also proposed extension of this RDA common model with some extra fields to better cover for what is specifically specifically asked in Europe. For the RDA work, we had to focus on the let's say world setting, find the minimum common denominator. For European context, EOS context, we were basically allowed to add some more things which is specific to what is asked here. And then, as mentioned already before, we put it behind an API which is again maintained by the RD working group, and also tools that are not part of our project are implementing or have already implemented the the API. Uh and besides that, we provide also some supporting resources that may be useful for you. We call these resources commons. And there are some MADMP mappings to traditional templates. We provide also some evaluation metrics on the quality of DMPs and tests and so on, but you will hear about it also later on. The next interoperability framework we have is the SKGIF, and this one is used for connecting different knowledge graphs. So, like you can see here on the right-hand side, you have all these potential knowledge graphs, and they want to exchange information with each other. You can also use it to to connect other systems which you didn't consider before knowledge graph, but actually are able to provide information in a structured way. It is also developed under the umbrella of research data alliance, and basically it's based on the group on the scholarly communication knowledge graphs. And OS trace is adding this scientific perspective, bringing some extensions, bringing some specific needs. You will see it in this slide. So, basically what the group is doing has this layered set of open specifications. So, they have a data model which define entities that are used by knowledge graphs. They have some shackle constraints for for validation, and they define JSON LD context. And they also provide an API just like I was showing for the DMPs. They have a set of operations you can perform using this data model. They are open for extensions, they have a way to request openly extensions in case you want to have more information about, I don't know, instruments used to measure something. This is exactly this shift from the scholarly graphs where you mostly deal with the products and people and projects to being able to represent domain specific needs and this is a direct requirement from our thematic pilots that they also want to be exchanged and and model information that are relevant in their domains. >> Sorry, Thomas, to to push you, but could you please finish in 2 minutes? >> Yes, I can. >> Okay. >> That was the plan. Uh so, now I have the FAIRiF and for the FAIRiF, so FAIR assessment interoperability framework, the starting point for the project was that we have uh as Mark has counted more than 30 tools that are doing FAIR assessment. Actually, they claim they do FAIR assessment, but they provide inconsistent results. So, in the slide you can see the same test being tested with two different tools. One says 20 out of 22 tests pass, the other one two out of 24. And the moment you start publishing results like these and start comparing things to each other, this is a total noise. You don't know what you're comparing with with what. And the goal we have for the FAIRiF is to basically standardize and harmonize what FAIR assessment actually means. We want to add transparency to this and we want to make it clear that if you have a FAIR assessment tool, you can use it and have it, but please communicate your result in a standardized way. For this reason, a big part of the framework is the FAIR reference model, which provides this conceptual grounding of what is a metric, what is a test, what is a test result, and how does it connect. So, to keep it short, the left-hand side of this diagram basically focuses on collecting objective evidence. Something has a CC by license. Something has an identifier. The right-hand side expresses the opinion of the community. It is good to have this kind of license in our community. It is a must to have this DOI or this kind of identifier in our community. More can be answered in the in the question, but this is the sentiment. So, we are not about producing scores. We are about providing information. And our work is not only conceptual. We have a catalog of benchmarks, catalog of metrics. You can go to FAIRassist and browse for what we have defined for communities. We have also implementations in the tools. I don't have a chance to show it today to you, but we can find everything in our documentation. So, I think this will be main takeaway for you. If you go to this docs.os-trails.eu, this is where we collect information on the things we have developed, and we provide pointers to the commons, so the pieces of software specification that you can reuse in your context to basically implement the interoperability frameworks. Of course, we are also on GitHub, and we have our deliverables, which we try to keep somehow easy to read. I know it's not an easy mission, but doable. That's all from my side. So, thank you very much, and looking forward to any questions. >> Okay. Thank you, Tomasz. We will postpone the questioning to the Q&A slot at the end of this webinar. And we will now move to uh the five presentations from OS Trails pilots. The first three are uh from representatives of national nodes that are uh organizations responsible for national nodes um that are a partner in the OS Trails project, and that's uh have conducted pilots with the tools and services available. The first on our list was Johanna. Johanna, I see you're nodding so I suppose you're ready for your presentation. >> Yes, I'm ready and I hope the sharing of the presentation comes through okay. >> Yes, fine. >> Okay, so I'm development manager at the CSC. It's IT center for science in Finland. We are providing national services for the research, education and and also cultural heritage and we have been privileged to um >> [snorts] >> uh to um share with you now the the practices from our national pilot in OS Trails. Just a minute and I'll see if I can move my screen. Yes, great. So the overview of our national pilot is that we have one single goal uh as a a key performance indicator and this is to develop an uh uh metadata model and uh reference data model for machine actionable DMPs in the OS Trail project. Uh single reference data model that would be uh scalable enough for all purposes and would be found useful at least by 60 organizations. So uh agnostic from the scientific disciplines but also of the DMP tools uh utilized. And we have been basing our our work so that we are utilizing everything that we have I've been able to utilize from the RDA a standard and also from the OS Trail OS Trail IF interoperability framework. So that we are not developing anything nationally that we don't then uh scale up to the EU or international level. But that we are taking into account the national needs as well. And uh there is a link in EUDAT Wiki for the the metadata application profile, the the detailed version of it. I'm not going to present it here, but you are more than welcome to have a look on that and on my last slide, there is as well a feedback feedback link for that. Um we have been considering that everything all the information that is marked mandatory in the RDA standard is mandatory as well in our national model, but then we may have some other mandatory fields as well. But then we have been looking really carefully from different points of view, what do we actually need in the data model? And the emphasis is really on the data set itself and on the nested structure of data sets. And also on the data life cycle management looking from the beginning from the discovery discovery of the data then to the final final archiving or deletion of the data. How all this um the whole chain is managed and then data protection has really high importance in Finland and promoting the machine readability. We have been as well discussing that the concept could be machine understandability, how the machines are understanding and able to utilize the information. And then we have as well included ideas how the DMP could be extended so that a complex research consortia that has multiple organizations from different countries with different practices could utilize the same model and that we could use different digital objects as well and and structures. And the software management plan we think is as important almost than the data sets um um details. So we have added a optional section for this and it has some brief information. This could be more detailed, but we have thought that this is a beginning and it is really good to align these practices. And we have had like how we have been able to accomplish this in large corporations. So, we have we haven't only done this within the CSC with the experts of CSC, but we have had um teams of experts from different organizations, at least from 20. And we have had a lot of open workshops for our national partners and stakeholders, our clients, customers. And with them we have then developed the the data model. It has been really detailed cooperation, but it has been as well really useful. So, we are now having the the model reference data model, but the um national organizations feel that they have ownership as well into and there are already implementations of this in practice. And like I said, it's tool agnostic, but we are as well currently then reviewing how the tools can utilize this. And the the whole threat line red line here is to support open science and also look into the integrations with the fair implementation profiles fits. So, the idea is that the the data management planning or the data management process would be in the center of the research ecosystem. We would have open API interfaces with some control for um information that is not yet open or that should be sealed. We are looking how the different stakeholders, whether they are organizations, uh people, uh data sets, software, publications, funders, how they could be utilizing the PIDs, um the uh permanent identifiers, and how different actions could be triggered, like reporting when when there is a certain deadline or the the data management plan has been uh indicated that it's now ready for submission or it has been um checked by the data support. And so, we have extended it as well, um and um our approach so that we are thinking that the API interfaces students only relates to the national infrastructure, but also to the EOSC nodes that are currently being uh built as well, the thematic nodes, so that's all the all the nodes utilizing the same service catalogs could um communicate about their services, and then the researchers are finding through the APIs as well um all those services that would be um accessible for them, and also that they would be able to utilize. And also >> Johanna, sorry to interrupt. This is in a very impressive overview. Could you finalize in 2 minutes? >> Yes. So, this is just a link uh or kind of indication how we are connecting with the RDA standard. We have added some elements, and we have as well um aggregated uh some things together, like everything that relates to the persons um that are not contacts, but other in other roles, they are in agents um section in the DMP. And um we have added the software development plans here, and then the life cycle, so that's the only difference. And this is just um slide indicating uh what is the added value of utilizing the digital objects. So, there are many digital objects to be utilized, and immense amount of information, but we can extract through them if we are having the custom that the researcher researcher are able to utilize them. And in addition to this complex table, we have as well made a relational model that helps tool users then to see what are the links and the different connections between the sections. And I'm statistician by training, so there are some figures as well. What are the required mandatory fields? What are then the optional fields? And if they would be seen as final, then what's the amount of information in them? And how much information we could automatically fill with the digital objects? We see that there are many different ways of utilizing the data management plans. One the single approach is that if the funder is requiring, then then make them. But also following the national guidance, one could implement them more widely as a practice, and the organizations could have better understanding of what processes, what research projects, what responsibilities there are if they would require the DMPs of all the research projects, also those that are internally funded. And for all research projects, the management becomes much more easier, and the the use of time is more efficient if the DMPs are used. So there is as well this internal motivation to start using them, so that it wouldn't only be those that are externally funded and when the funder is requiring them. We have been asking experiences and approaches to this um MA DMP P data model and the the use it of that and here are some um use cases that our stakeholders have identified as possible use and then there is this quote that the the by one of our stakeholder that actually the possibilities are endless if only the system can be made to talk with each other and with that I'm ending the the presentation and I hope if you have any questions that you can find it easy to reach me and also if you would like to have a look at our machine actionable data model so please send us some feedback. Thank you. >> Thank you Johanna and I'm very grateful for underlining the very different types of requirements that that exist in the landscape. So national funders may have very different requests compared to organizations and and project teams etc. I think it's good to keep that in mind for those who are considering to become a national node that all these diversity should can be accommodated by by the services. Okay, we move on now to Raul Palma from the Polish pilot. Are you ready Raul? You are muted. You're still muted. >> Sorry. Okay, so let me just put my screen. Trying to find the button here to share my screen. Okay. Okay. Share my screen. I cannot find the button to share my screen. Sorry. Is it is not enable? Uh >> Um it's enabled. You can see down on your screen. >> Okay. >> you have the toolbar. It's the middle button. >> Okay. >> Perfect. >> All right. Can you see my screen now, right? >> Yes. >> Full mode. Okay. So, I will go quickly through the Polish national pilot. This is the pilot 704 trails. Um well, basically here the goal was to support the creation and maintenance of machine actionable DMPs in Poland. Uh this of course included defining requirements for this DMPs and enable them to be linked to data sets, research outputs, and other entities through PIDs and the use of RO-Crates. The idea here is to support FAIR assessment of DMP-related resources and improving interoperability with the national DMP platforms and uh repositories. So, the context here is we have already a platform which is uh called Aro hub, which is the Aro crate platform that has connections to different tools uh including DMP for the uh uh um um Argos for the DMPs, uh Fair OS for the FAIR assessment, and OpenAIRE as the source to uh SKGs. Um as part of FOSTER trail, these connections are being of course uh improved, uh enhanced. Uh for example, in with respect to the use of the DMPIF in order to import uh um this uh DMPs and generate RO-Crates from Argos, but also this can be as well done uh through other tools implementing this IF. Uh similarly with FAIR IF, uh the current implementation is based on a very specific uh API, and the use of a standard API will allow us to uh also connect to other solutions. And similarly with SKGIF, in the in the current setting, Aro hub exposes all the research objects or auto crates uh through uh to OpenAire SKG using the OAI-PMH, but uh with the use of the SKGIF, we will use the same standard just to provide this information to the OpenAire, but also we can uh also consume data from uh OpenAire and other knowledge graphs in the same way. Uh of course, we have also national repositories like RepOD uh and others that mean that are general purpose, but also discipline specific uh that have already the possibility to do the publication of research data assets with uh PIDs, but as part of OS Prace OS Trails, this metadata will include connections to crates and DMPs in the external repositories. Of course, our challenges are how to make sure that IFs will be uh enabled at least to maintain it or improve the current integrations, the times uh taken by providers to implement those IFs, and obviously in terms of agreements of DMP templates. Um In in relation to the connection to to EOSC nodes, uh obviously our main connection is with the Polish node, uh which is coordinated by NCN. NCN has been involved in the process of defining the the DMP requirement. I mean, they are the source of the DMP requirement, but they have been involved with us in in in creating the templates. Um also, we have discussion with OP national platform to allow researchers to easily complete their DMPs, and our services are being on-boarded uh in the Polish node that includes RROHAH, FERRORS, but also Dam-Up, which is not exactly in in the pilot, but it's related, of course, as another DMP tool. Um the added value of the IFs is basically that will allow the interoperability with different solutions. In our case, we are really reusing the three pillars of OS Trails, as you can see, DMPs, FAIR assessment, and SKG, allowing, for example, our uh pilot not only to become provider but also consumer for the different types of existing knowledge graphs. So, just to give you a high-level overview of what our implementation. So, the first use case is about the machine actionable templates. We have defined based on the NCM requirements a set of mappings with semantic mappings and implemented these templates into the Argos platforms at the moment. This is already integrated also with the RO-Crate that allows the creation from the DMP to an RO-Crate and then later on to export this in a way that can be easily consumed by the OP national platform. In terms of the fair assessment, in terms sorry, of the national platforms, this is repo on the left-hand side and what we are extending is the metadata to include connections to the DMP, to the RO-Crates that are related to to this particular asset. And regarding the SKG, RO-Crate already is implemented the whole SKG implement IF specification and we also support content negotiation to allow, for example, the retrieval of an RO-Crate in a way that is compliant with the SKG IF. And finally, with respect to the fair assessment, this is what we have. Of course, this is already integrated into RO-Crate, but also we are migrating this into a better display and information provision based on the fair IF that is being implemented by the Fero's tool. These two tools, I mean, Fero's and RO-Crate, like I said, they're becoming part of the of the EOSC node. So, thank you very much. That was from my side. >> Thank you, Raul. And also, thank you for sticking to the time limit. Um questions related to this presentation as related to the previous one can be postponed until the Q&A slot. I would also like to mention that there is a slider for the group of people in this um in this event to indicate um >> [clears throat] >> um the some more details about their background so that we can share that later on. Um the next presenter is now Elin Wagemaakers, who will present the the pilot for the Dutch node. >> Oh, I have some issues with sharing my screen, it seems like, because I haven't enabled it yet. But wait. Um Should I quit Zoom and reopen? >> If the problem can't be solved, then maybe somebody else who has already uh presentation rights can select your presentation and and uh >> Try again. It seems like it's working now. >> Elin, uh everyone Yes, everyone has access. Thank you. You We can see your screen now. >> Yeah. You can see the full slides now as well? >> Yes, the full slides. >> Okay, perfect. All right. Uh thanks. Um so, my name is Elin Wagemaakers. I am program manager of Open Research Information at SURF. And uh together with Andrew Hoffman from Leiden University, uh I am coordinator of the um OS Trails Dutch national pilot. And uh on the slides, you can see our use cases. Um uh so, we're doing this together with uh four consortium partners partners. Uh, that's NWO, the National Thunder, uh the Vrije Universiteit uh Amsterdam, TU Delft, and uh the Digital Competence Center for Applied Sciences. And uh we basically have four um areas of work or use cases, uh how you want to call them. Uh, ranging more from institutional compliance and the needs uh there uh when it comes to um data management, uh research data management. Uh, and uh the open science uh workflows. Um, so the first one is uh integrated workflows. So one of the things that we have uh seen in the Netherlands is that uh many institutions are trying to combine um data management, privacy, and ethics um compliance uh in one integrated workflow. And what we're doing in our tools is exploring how uh MaDMP tooling can help to combine um those or to integrate those workflows. Uh, so that's from the administration part. Uh, then we're looking into how we can expose those DMPs in various ways, either uh in local repositories, um the CRIS uh systems, uh but also in public repositories, uh for example, uh such as Zenodo. And then also looking into what exactly should be exposed, what uh what type of information. Uh, then moving more towards the open science open science workflows is how we can uh track uh related entities uh in a research project uh through uh MaDMP tooling uh combined with RAID. Uh, RAID is a relatively new um persistent identifier for research activities or uh projects which basically combines a set of other persistent identifiers and bundles them in one yeah overarching envelope. So we're looking into identifying these DMPs. And finally we're looking into the implementation of fair implementation profiles beyond the tool data stewardship wizard which already enables FIPs. So what we're doing there is basically developing conceptual model of FIPs and how we can import certain community specific profiles of fair enabling resources into a DMP template. So what you see here is that we're exploring how DMPs can be useful at various levels at the institutional level but also more towards open research information and open science needs. Then I move to the the evolution of the Dutch note in the EOSC Federation. So SURF is the mandated organization for for the Netherlands in EOSC and hosts and develops the national note. And what we what you see here is a a profile of the HORC. It's global open research commons. This is a model that was developed in RDA where basically you have three more hard elements. That's the elements in blue and five white elements that are more human or social elements of governing research infrastructure and research data. And in the middle you have the standards. And what we're doing in the Dutch node for EOSC is looking at the the blue elements, the hard elements, the compute storage, the research objects, and the services and tools. So examples are for example our compute environment research cloud, research drive for storage, but also research objects such as data and the metadata that comes with it. But next to these to the generic infrastructure you have there's an important role for the thematic um communities in the Netherlands. So we have what we call thematic digital competence centers for natural engineering sciences, social sciences and humanities, and life sciences and health. And they basically provide the engagement and human capacity with their specific needs and research objects that are relevant for that specific domain. So the way that the node is set up is that it basically combines these generic services and tools with the the needs from the the domain specific domains. And so this model should help to align data and also high performance compute across Europe and should strengthen the interoperability and data reuse among those different uh domains, but also uh across those domains. >> And you manage in 1 minute to finish? >> Yes, sure. Uh so, the last slide is about uh how we do this and and what the added value is of the interoperability frameworks. So, the way we approach this is by um looking at uh workflows. Uh so, the workflow that you see here is uh a researcher uh trying to request uh or find data, request data, um uh can access that data and then um in in uh process it in a secure environment and finally transfer it and publish it. And the way uh that is done is um uh through federating access uh to various data sources, um making sure that you have access to a secure a secure compute environment, and uh in the end then uh publish it in a um designated repository. Um so, you see the federated access at the common building blocks. And what you could track in your DMP um through the DM uh the interoperability framework is what um services were used in this specific workflow uh throughout to make sure that you have a reproducible workflow. And the SKGIF is obviously uh important to uh enable the um the discoverability of those uh data sources and um and also in the end make uh the end result uh again discoverable. So, uh in short, the the SKGIF here will help uh discoverability, um which is a crucial component for the federation, and uh the DMPIF will enable uh reproducible workflows. >> Thanks a lot, Eline. Um We can then now switch to the two presentations from the thematic pilots. And the first one is uh will be presented uh by Renault Dweem, uh who is from uh the PaNOSC cluster. Are you ready for your presentation, Renault? We see you coming. Thank you. >> Yes. Okay. Yeah. Can you see my screen? >> All fine. >> Okay, good. Um so, I'm working at the European Synchrotron Radiation Science Facility. So, we are part of the Photon and Neutron Open Science uh Cloud, the PaNOSC node. So, we are we are a big facility doing experiments, so a big microscope. So, we we are manipulating uh various elements from project uh to instruments, data set, and publication. Um So, I'm going to present how we use the the various uh interoperability frameworks of West Trails. Um so, we have uh instruments in our facility. Though, they are represented by um by PIDs. So, DOIs. So, we are working on it at the moment, so generating PIDs for for instruments. So, we use these instruments to analyze uh any type of data we could we can put in our instruments. So, for example, here we have an example of uh of a data set. It's a dinosaur. So, we we we are able to to to create slices and uh so, basically, we have data sets for these uh DOIs and data sets for this information. So, 3D representation. Um then, these data sets are reused in a publications. So, this is an example of in nature, so with a DOI. Um, so then following on the PIDs, what we did we did we also defined we have PIDs for our techniques. So, basically how we study this uh uh all the techniques we used in our instruments to to visualize and create and use x-rays. So, these are defined we have an ontology declared in in fair sharing. And finally, we also declare the actual data portal where we store the our all our data sets in in fair sharing. So, we have a unique ID for our data portal. Uh so, now I'm going to present how we use it how we use this in a DMP. Um, so when you start a DMP, you may give information about what you are using. So, typically the vocabularies and standard and the the repository you you plan to use for your for your research. So, we use the DSW the DSW tool. So, we are able to link the IDs I presented before to this DMP. So, typically the vocabulary we have and the repository. Um, then you can answer all the kind of questions that are not directly related to PIDs. So, it's a standard DMP uh this is a standard DMP uh screenshot. Um, so let's say you perform your experiment at ESRF and then you have data sets. So, what you can do with DSW it's to link your data set to your DMP. Um so there are ways to have a auto complete auto complete form fields that are related to that are linked to our data catalog. Um you can also link the publication. So for this we we use the open air graph and this and the SKGIF API. So also there's the button which provides a way to integrate these external uh data sources uh to the DMP. So here we find the the publication and link it. Uh and the the next step is to assess the DMP. So as Thomas presented at the beginning of the of the meeting, there is a standard export of the DMP. So it's the match the machine actionable DMP. So this DMP can be sent to a DMP evaluation tool. So we have advice and score about our our DMP. So we can check um if we have PIDs and we can define a set of metrics and test. I don't go in details, but the idea is that now the DMPs is related to an evaluation platform that allows performing test uh to give feedback to the to the users. Um so technically uh this is how it looks in in this WSCU. You you have a benchmark evaluation uh defined here and you have evaluation results. So if you need information about how you can set up this in your DMP tool, you can contact the the DSW team. I think they're they are in the meeting as well. And to help you setting up this this DMP evaluation. Um What you can also this slide is just about explaining that with the DMP you can also link data you you use for your research. With the same you could imagine having 2 years after this this study other people reusing the data and citing the data in the DMP. So, for us we for the SRF we we want to um to make sure our data is used and reused by by many researchers. So, in integrating the the data set references in the in the DMPs is very important for us. Um So, uh another element we we are doing as part of the the SRF project is to enhance the the metadata we have in our data sets. So, our data sets have DOIs and are pushed to to DataCite. It's quite standard. Uh so so what we did as part of the project is to to enhance the the metadata we have in in DataCite. So, ORCID, affiliation, grants IDs, funding IDs. We also integrated our vocabulary metadata. It's what I what I showed before. So, the techniques we use. And we are working at the moment to add the instrument PIDs to these data sets so they could be integrated in DMPs and added to the global graph. Um in our catalog >> uh what is interesting for us is to show how data set is cited by publication. So, there are many ways to get this. So, we can use the SKGIF API to connect to OpenAire to have citation information about our data sets. And we have also an an internal SKG curated by our library. So, we can get additional information. So, all this information about publication citing our data sets is now visible in in our data portal. Uh last element in 30 seconds, um what we have what is provided by OS Trails is a global framework to uh assess DMPs, but also assess the metadata of the data sets we have in our data catalog. So, we define the bench and benchmark especially for our data sets. So, it's defined in fair sharing. Uh we define a set of metrics uh to uh assess the metadata. So, for example, we check if the data set is cited in a publication, if the data set has a license, has funding information, uh is open access, extra. All the uh So, we we want to make sure all the metadata in our data sets is uh complete. So, we define all the metrics >> minute now? Um >> Yes, last slide. Uh just to show how it works. Uh we define the benchmark and have the result here. So, if you have any question again, please get in touch and I can explain. And that's it. Yeah, thanks a lot. >> Thanks a lot, Renault. Uh good to see that you managed to cover um the all three interoperability frameworks in the work uh conducted this far. I think still something similar is true for the next presenter, Menzo Windhouwer. Uh Menzo uh will represent the work done in the Shock Cluster. >> Can you see my screen? >> Yes. >> Okay. >> So I'm telling this from the perspective of the CLARIN part of Shock. So CLARIN is the European research infrastructure that gives access to the language resources. And it does that both for anybody in the social scientist or humanities domain or who is interested in language resources in general. And well, we are one of the OSTRICH pilot. We're one of the thematic pilots. Let's listen. Yeah. So what are we doing? Big part of the focus is on the SKG part. So on the scientific knowledge graph. So in CLARIN there is the virtual language observatory, which is like our central catalog for all the the the the data sets that are provided by the the the centers that are in the CLARIN network. And we provide access to that to all these data sets via using the SKG IF API. To to expose that we are mapping all the the internal format that is used for records, which is the component metadata infrastructure format that the CLARIN centers provide to the to this VLO and the VLO extracts from that facets like what is the language of this data set, what are who are the creators, who is providing this, what is the name, and so on, very basic facets. And all these facets we map into the SKG and we expect we expose that for via the API. Um our thematic pilot is one that needed an extension of the of the SKG core model like Thomas explained. So, we needed we also have metadata records that are about services. So, that is not yet covered by the SKG core model. So, we also expose the services records via the service extension. Um the SKG is based on RDF in the end. So, what we do, we did already have a pilot from the past where we translated SIM D records into RDF. So, basically we revived that infrastructure and we we expose all other records as as SKG compliant RDF metadata into a triple store. And then we have provided the built the API on top of this triple store. So, it's a bit of a different approach some other pilots for example wrap their existing API and add an extra layer that then exposes the SKGIF API. So, we'll do it a bit different there. Um What in the end we want is also to try out cross SKG interaction. So, if if in the VLO okay, you're into certain data set, and we know who has created that data set, can you inspect if that creator also has a data set in one of our neighboring uh uh SKGs? So, for example, CESSDA. So, could you then jump to the CESSDA central catalog and find the the data sets or the research products that exist for that same user or organization in the other SKG. Now, next to the X SKG, we also provide a DMP template, which is mainly based on the the header requirements that metadata and data sets have to be in CLARIN, and also based on components that exist in the infrastructure. We also implement an FAIR IF benchmark, where we mix and match um generic tests that are applied by provided by AuScope himself, and community-specific tests, where, for example, CMDI-specific knowledge is needed. And we use that by this cool tool called PyFED that can implement these uh community-specific tests, and we base it on earlier work where we also created a FAIR implementation profile. Okay. So, CLARIN will exist as a node in the EOSC federation, at least, eh? So, CLARIN is a federation of data repositories, and it will be eh it will will CLARIN as the infrastructure will provide access to these data repositories and all these linguistic data they have via the federation, but it will do that in the context of of the shock. So, in this SSH open cluster and in this um this EOSC mesh project that will start in September, this will be actually be fleshed out more and more. So, CLARIN is there together with Dariah as part of the the shock node that will be in this um project. So, what is the added values of the interoperability frameworks? So, the SKGIF gives us a common access protocol, so an API for uh getting access to the harvested metadata and to all the data sets that are in the central catalogs. And it would not be a proprietary API, but is an API that is common to all the other research infrastructures and all the other nodes in the federation. FairIF will allow us to tell how well the data sets follow uh the fairness as defined by the community. And if there are still things lacking there, guidance will help you to improve that. So, that will also be information for the CLARIN centers itself. The DMP um uh the template that we provide will help projects that that operate within EOSC and will want to be compliant with CLARIN requirements to to create their their data sets, to create their data in their projects following CLARIN requirements so they can easily moved into uh CLARIN centers and in their repositories. Okay, that was my uh my part of the presentation. That's >> Thanks a lot Manzo. Yeah, you can stop sharing. Um we will now move into the Q&A uh slot of this webinar. Uh and we would like to invite you to and especially um people that are not yet active in OS Trails uh not active in OS Trails but are working towards uh their node charter or plans for becoming a node um to raise any questions either to the presenters or the experts on the interoperability frameworks that were announced earlier Paulo and Mark and Marek. You can use the um uh the um >> [clears throat] >> the chat box for entering your questions and we will see how many we can address right now and then later on you will receive an invitation with uh uh a format indication communication channel for uh mentioning any concerns, observations, suggestions that can then be the basis for more tailor-made interactions with the OS Trails uh experts and and teams. And while uh people are thinking about how to phrase their question uh of course uh the people from OS Trails can bring up some additional comments and observations that would otherwise not easily be addressed. I see >> Okay. >> uh There are some comments but okay. >> Um just to refer apologies Francisca for the interruption. You can unmute yourself and uh and use the use the microphone if you want. Wait. I Do I have to unmute or to mute myself? >> To unmute. It's a uh it's uh to all. >> Okay. >> Yeah. >> I see a question arriving. Um yeah, I think this is a question to confirm if the interpretation that EOSC-hub provides services for the EOSC uh federation or mechanisms that the EOSC federation services can adopt. Uh I see many thumbs. This is precisely um what what is the purpose. Yeah. >> And and of course beyond. So, these uh the the the the the needs targeted by EOSC-hub go well beyond the EOSC. Of course, we hope that the EOSC will actually uh single them out as recommendations at least, uh but we have I think intercepted the need from from the whole community in science, right? So, it's also a little bit beyond. >> Yeah, and I think the latter comment aligns with what Thomas said at the beginning, namely that um the plans and the development thus far are geared towards the EOSC federation on the one hand, but of course also have an independent um sustainability plan. Yeah. Um maybe I should ask the representatives of national nodes. Okay, yeah. Anca, please. Um the ENVRI perspective uh is of course also very interesting and you are uh involved in OS Trails, but the topic was not covered in the program, so please go ahead. >> Yeah, I am involved as I you said that some the people who are not involved, so uh but I I wanted to so I remove the fact that I'm involved in in OS Trails and I will talk now only about the uh ENVRIOSC node, which was um uh accepted in the second wave. Yes? And uh so we are at the very very beginning and um we kind of didn't implement anything from the uh OS Trails, but I can actually tell you what we would need. Uh for example, a very very high priority is the sign- uh scientific knowledge graph. I think it's the most strategically important component. Why is that? Because ENVRI is not actually a repository, it's a it's a federation of federation. Federations. So, similar to the SHOCK cluster. Yeah. >> Uh yeah, pretty much. So, the value comes from connected research infrastructures and services and data sets and software and publications and ECVs and EOVs and researchers and workflows etc. etc. So, without this kind of graph graph, the node is essentially a a sort of directory and and that's it. Um also the node would need um a fair assessment um tool. This would be, let's say, medium-high priority. Um not because the FAIR scores are are useful necessarily, um but um I'm thinking that EOSC sometimes behaves as if fairness is the goal and it's not. F- FAIR is the means. And you know, researchers do not wake up thinking, "Well, I hope my data score is 83% fair today." No, they want disco- discoverability, and interoperability, and reuse. And this assessment is useful because it provides uh evidence. So, the node The every node should be able to demonstrate in principle that inter- uh um that we are ready interoperability is it's uh readiness and and machine read- readability, etc. Um Another high priority would be the fact that every node needs machine actionable information. And this is where OS Trails become becomes relevant. Because the node needs PIDs and uh uh machine readable metadata and provenance, etc. etc. So, um the question is, can machines understand every resources? And not do humans have PDF PDFs describing them? And now it comes to um machine actionable DMPs, and I think you all know what is my relationship with the with the DMP. So, for for me, this is a low to medium priority because my answer would be not really necessarily. Because uh this uh machine action actionable DMP to me is is useful only if it contributes information that cannot be obtained elsewhere. But, um we have many of the of the DMPs elements that exist. We have uh that that exist in like repositories and catalogs and PID systems, etc. etc. So, that would be my my answer to this. >> Okay, interesting, provocative statement. Uh I think you've made that point at other places as well and we we might be able to discuss that later on, but maybe it's good if um um people re- respond to uh people from the EOSC Association in this meeting also raise their voice. Um maybe about models for a stronger collaboration in uh in the short term while EOSC Trails is still a project, but of course the partners in the project uh are all eager to uh strengthen and improve the nature of the relationship, so any input would be welcome. Um >> I can just maybe tell a word on that. Onka, you were saying that uh I I I can just say that DNPs are contributing uh giving more uh information to the global graph. For example, when you when you fill a DMP, you give information about the your teams, the authors, the data set, the publication, the grants. So, it's a way to uh to improve the data we have uh in the in the global uh SKG. Typically, you can send information to to OpenAir to enhance the the graph. So, it's more information you can uh you can give to everybody. So, uh >> Yeah, it's more, but not necessarily unified or harmonized in all uh dimensions, and I think that's a an issue that comes up at other moments in discussions as well. So, I don't know if this is the way Onka would phrase it, but Okay, I see >> I was just saying that the that uh every all the information that you find in in this machine actionable DNPs, they already exist in our repositories, yeah? In our archives. So, if if the node spends effort collecting DMPs, but never uses the information because we already use it somehow, it creates just bureaucracy. >> Okay, let let's go back to Thomas who raises his hand. >> Yeah, I raised my hand not because of DMPs. So, I >> Yeah. >> I raised my hand and I believe if we we have all this information, you can then generate a DMP in case somebody wants it from you and that's another another use case, but I'm not about this. I think for the most of the people who are representing nodes, the big problem is connecting the daily problems they have in setting up a node with this theory that we are giving them. We are telling them, "Get this interoperability framework and that interoperability framework and what do I do with this?" And and basically I want to share my reflection because if I imagine myself setting I think the interoperability framework should be a way for me to tell, "Do I want to onboard a specific service into my node or not?" If I'm getting a service that is not interoperable with others and end of the day we're building a federation. So, we have to evade this vendor lock-in. We have to make sure that information from one tool is to be reused in some other tool, then this is a criteria for me for the selection of what you're using. And I think one of the discussions in the in the EOSC community and also what we have in our trails is what happens with these interoperability frameworks. Where will they live beyond the project? And I think they will live in the nodes if nodes adopt the services that follow them. So, I think from a very practical point of view, if you are a service operator, you may want to basically include services that implement the the interoperability frameworks. >> Yeah. Yeah. >> course inter-operability has been a very important concept in the setup at least of the thematic research infrastructures across Europe. Um Let me see. Uh Okay, I see some conversation with Anca involved going on. Please have a look. Um I wanted to ask um uh Ellie, if can you easily upload what you found what what the Slido uh survey found out about the audience of today or would it take too much and it will be sent it around later on? >> No, let me quickly switch. Oh. Can you see? Can you see now? >> Yes, we can see it. Yes, please. >> Okay. All right. >> So, this is from the So, there there were three questions from the Slido. One is if you represent EOSC node majority no, um some yes. And the other two are more interesting because um the next asks for which IF is most relevant to you and your community of course based on the needs and requirements. DMP IF is on top, right? Uh so, wins the medal. And then we have uh ah now >> [laughter] >> But But is it you? Yeah, because SKJAF and FAIR IF have been neck and neck since the beginning. >> Well, you but you can conclude even though it this is a ranking the fact that even the smallest proportion is is more than is close to 30% indicates that the the selection of topics for our streets is is relevant. >> Definitely. And then the last one is how likely are you to use the streets I have? And very likely. So, if we can also you know, have people to unmute and say their opinion I like how would they use it or why is it not that um likely for them to use the streets I have that would be great for the discussion. >> Okay. Well, we reached the end of the time slot and I guess people will start longing for lunch soon. We still have 44 people. So, if people think that the outcome of this survey could immediately be commented already in a way that will enlighten us, please go ahead. Let me see what the recent um Yeah, there's a conversation going on with people that are um in touch anyhow. I think if there are people that are still struggling with whether or not to set up a node whether or not to have access to um um service provision that could help them um to set up their node etc. I would like to give them the floor. Um >> Uh so, if I may, so set setting up a node does not really require the interoperability frameworks we are discussing now in this framework. >> No. >> But of course it would be uh let's say in in the long term it would be very fruitful. For example, the EOSC node at the moment is exchanging data with uh uh other collections using the SKG >> AIF. >> And similarly, I think many nodes will have the issue of DMPs. And the DMPs >> As the survey shows, yeah. >> As the survey shows. So, I I think running a service that builds its own silos and universe is not worth it for any um uh for any node. So, having a way to exchange uh DMPs across us nodes would be important. And here we've built the pathways, for example, to connect the SKG and the DMP. So, as you said before and uh Renault highlighted again, I think the connection between the two is crucial. And then of course we know how about validation, right? So, and again validation is not an easy issue. It's not an easy say activity in general. And these interoperability frameworks, especially the services behind it also, uh can simplify the way we share ideas about what is a fair task, what is a fair metrics, and also the code behind it, which is again >> Yeah. >> essential. Reusing the same code, referring to the same indicators will become more and more important. Within node, like the AMBRE one from ANCA where you have different souls that are coexisting and they need to share ideas, but also across nodes, okay? So, I think uh although these frameworks, the EOSC trails is not strictly connected to the setup of a node, I think it's a wise way to uh um let's walk on to design a common future. >> I think looking thinking back of Johannes presentation with the statistics that she ended with. I think that was a very insightful exercise that other national notes could also be interested in so please let us know if you would like to be in touch with with Johanna and her colleagues or if you would like to know more about how the thematic services are set up in a federative way even before the EOSC federation comes in in the picture to see if you can build on that because with some effort that's what for example Menzo's talk showed and in a way also that of um uh Renault um the the existing well-functioning research infrastructures can uh enhance their service quality by looking into the interoperability frameworks uh not without changing everything but connecting things in a way that allows them to connect to more dots in the network. Johanna, your hand is raised. >> Yes, thank you for this. Yes, I can be contacted. We are really happy to share our experiences. We are as well planning to have an national for the national OS space pilot. An MDMP hybrid working day on the 24th of September. So if you're interested to join to that so please let me know. Um I think it's really when building up a node and planning ahead what to do and how to do all this everything is to plan gradual steps and building up on on these interoperability frameworks is really important to start from a common ground. So then one is not diverted and one wouldn't create services that wouldn't be interoperable in the end. So this requires a lot of compromising as well, but I think that this is really crucial in order to make um services that in the end help the researchers and also these cross-organizational research groups so that they are able to use interoperable services smoothly. >> Thanks for this uh uh enlightening uh recommendation uh >> [clears throat] >> I think most people would know like to uh go for lunch and I see already that the number of participants is decreasing. So um if there's one or two more urgent comments or observations, please take the floor. But otherwise uh we'll stay we'll contact you again and if you would like to take the initiative, you are always welcome to contact the EOSC Trails project. I see no hands raised, so I assume that we can say goodbye. Thanks a lot and also thanks for the supporting team Xenia, Ellie uh and the colleagues of OpenAIRE and Halasya from CLARIN. Uh see you at the next occasion.