Submind YouTube summaries
Thumbnail for 24th OpenAIRE Graph Community Call

24th OpenAIRE Graph Community Call

Watch on YouTube

Video summary

The EOSC Beyond project is a 36-month initiative funded by the European Commission, coordinated by EGI with over thirty partners across Europe. Its primary objective is to establish the next generation of core services for the European Open Science Cloud (EOSC) by facilitating the seamless federation and interaction between national, regional, and thematic nodes within a unified technical architecture. A central component of this infrastructure is the EOSC Core Innovation Sandbox, which serves as a pre-production environment where nodes can verify their maturity against federal requirements while testing integration capabilities with AI, accounting systems, and order management services. For users, the project offers a single entry point known as the user space, allowing them to authenticate once and access various functionalities such as data transfer tools, automatic infrastructure deployment requests, and discovery mechanisms for research resources like datasets, publications, software, and interoperability guidelines. At the heart of this ecosystem lies the Open Research Graph, which acts as a foundational knowledge graph linking diverse research outputs with service catalog entries. The Discovery Hub serves as the front door to these resources, enabling users to search through advanced filters based on criteria such as data sources, EOSC nodes, and specific resource types like services or deployable applications. When a user selects a dataset within this hub, they are directed to EOSC Explorer, where the power of the knowledge graph becomes evident by providing rich contextual information beyond basic metadata. This includes details about associated publications, author ORCID identifiers, funding projects, related research communities, and impact indicators such as citations and popularity metrics. Furthermore, the system aggregates data from multiple sources to handle versioning and provenance, allowing users to view different versions of a dataset alongside their specific licenses and access restrictions, ensuring transparency regarding how metadata was collected and enriched by algorithms. A significant advancement in this framework is the integration with EGI's data transfer service, which streamlines the reuse of research outputs without requiring local downloads. Instead of transferring large files through a user's computer to another storage location, users can initiate a direct cloud-to-cloud transfer from within EOSC Explorer. This process involves selecting a dataset, authenticating if necessary, choosing between multiple DOIs for datasets with various versions, and specifying the destination system such as S3, Storm, or DCache provided by EGI partners. The system automatically manages these interactions to move data directly to the user's preferred cloud storage while adhering to applicable data protection policies. Looking ahead, the project plans to expand its knowledge graph using new sources from pilot nodes and is transitioning from a centralized hub architecture to a mesh-based federated search model, allowing individual node catalogs to query each other independently rather than registering all services in a central sandbox. During the community call, several questions were addressed regarding the operational details of these systems, including how metadata extraction utilizes AI algorithms for enrichment tasks like document classification and sustainable development goal tagging while maintaining quality through participatory guidelines. Discussions also clarified that accessibility levels are determined by authors during deposition processes rather than within the explorer interface itself, meaning that datasets with mixed access rights must be deposited separately to reflect different licenses accurately. Additionally, it was confirmed that training materials can theoretically be published as research products in the graph but currently rely on pilot nodes for delivery and compliance support. The session concluded with an announcement of the next community call scheduled for September 16th, where further updates on these evolving architectural changes and new topics will be shared with the broader OpenAIRE Graph Community.
Read the full video transcript
Okay, so let's start. This is our um agenda for today. So, I'll start with a brief introduction to the EOSC Beyond project. And then we'll see the foundational role that the Open Research Graph plays in that context. And and then we'll do a deep dive in the Discovery Hub and EOSC Explore. So, let's start with a background. EOSC Beyond is a 36-month project funded by the European Commission. We are a large consortium, 30 33 partners. The coordinator is EGI. And uh we have 16 work packages and 10 pilot nodes. And our goal is to build the next generation of EOSC core services. So, uh I will not go in into the details of these slides. Uh but uh basically, what I want to say is that our primary goal in EOSC Beyond is to support the setup and the federation of EOSC nodes. Mhm. So, we are developing uh tools, frameworks, and an infrastructure to help national, regional, and thematic nodes to interact seamlessly within a unified uh technical architecture. And a central piece of this infrastructure is the EOSC Core Innovation Sandbox, through which nodes can uh verify their maturity uh according to the technical requirements of the EOSC uh federation. And it is conceived as a pre-production environment. So, a place where nodes and users can test and experiment with the core services of the European Open Science Cloud and the federating capabilities. In particular, for nodes and service providers, the sandbox offers technical support for integration with the AI, with the accounting, the order management, uh and all the services that are the federating capabilities of the EOSC. While for the users, we have what we call the user space, which is a single entry point where you can log in with the essential credentials, with the EOSC AI, and you can test, you can use data transfer capabilities, you can request the automatic deployment of infrastructure and services. And specifically of interest today, uh discover resources, services, datasets, and other research products. And it's here that uh the Open Research Graph is a foundational component. Uh so, what you will see in the discovery app, here in the middle, um is a subset of the Open Research Graph. And um we link the resources of the Open Research Graph to the resources in the service catalog, which includes services, but also interoperability guidelines, adapters, deployable applications. So, we link all these together, and we make it available in the discovery app. Then, if you select, for example, a a service, you basically um open the details of a service in the front office and order management and this allows you to request the automatic deployment of the service if it's available. While if the user selects instead a research output so for example a data set then you end up in EOSC Explorer. Mhm. Uh really briefly the methodology for building this EOSC Beyond Knowledge Graph. So what we do is that um we take the OpenAIRE Graph we select the research outputs that comes from data sources registered in the EOSC Beyond catalog and we enrich those research products with specific with a specific PIDs so the persistent identifiers of the nodes and of the registered data sources. So then this is where we create the connection between uh the the service catalog and and the and the OpenAIRE Graph. Um So the Discovery Hub as I said the Discovery Hub is the the entry point is the front door to all these different types of resources. Uh you can use the the basic search or the advanced search um and you can also apply filters here. I don't know if you see my mouse but there are filters on the left that allows you to filter your search results based on different criteria. So one is data source the other one is EOSC node the two I mentioned before but there are also other criteria. Mhm. And you can also uh see here in the in the menu in the middle of the of the user interface uh that there are there are let's say menus to easy switch between different types of resources. So, you have research outputs data sets, publications, and software. You have uh service offerings with deployable applications and services. And among the deployable applications you you found those that can be auto automatically uh deployed on on infrastructure of the EOSC uh um of the EOSC beyond. Uh you can search for organizations and data sources, and also for interoperability guidelines, and um and adapters. Now, um let's focus on the heart of our research discovery, which is EOSC explore. So, EOSC explore is the part of the of the user space that is developed by uh by opener. Uh so, when you are in the discovery hub, you find uh a data set that you are interesting to. Uh and when you click to And already here you can see a lot of information, huh? Because you have the information well basic metadata, but also number of downloads, views, funders. But of course, we can provide even more information. And to do that, you just click on the on the card, and then this is where the power of uh the knowledge graph shines, because we have a lot of information here that can be used. So, we have an overview, so the metadata overview, maybe it's not the right name, but the basic metadata of the of the data set. And uh but it's not only about the data set. It's also about its context. So, for example, you can see the main publication that is associated with a with this specific data set. Mhm. So, the the primary uh publication of the data set. And this allows you to understand the results in the context of the original study. We also have uh one We also have links to the ORCID identifiers of the authors. Uh in some cases, we also have uh citations and and the links to funding project. This is the case. This was funded by EOSC Future. We also have links to uh related research communities. These are research communities that uh collaborate somehow with OpenAir. And in this case, this data set was also classified by fields of science, by the automatic uh algorithm of OpenAir. Uh here instead, we highlight the the impact indicators. Mhm. We have metrics for citations, popularity, influence, and um this gives you a quick way to assess um the reach, the impact, and the relevance uh of the data set, together with the views and downloads, which are instead based on the uh usage counts service of OpenAir. Very important is information about provenance and versioning. Because the OpenAIRE Graph aggregates metadata from multiple sources, we can basically it can collect different metadata records about the same research product from different places. So, what we have is a duplication algorithm that finds this duplicate, the different version, and we put them together. So, in this case for example, if you click in these in these parts, you have the list of all the different versions that we have. We highlight which are which is the license, and if you click for example on Easy, you will go to the to the landing page of Easy where you can actually download the data itself. When you see the orange open lock, it means that the version is open access. There may be versions that are not open access. They can be closed or they can be restricted, meaning that you need for example to log in in order to access. And that depends the type of access of course depends on the on the specific source, on the specific version. Then let me go to the next slide. Yes, so if we click on view all versions, here we have the details of each version. So, not only the link that goes to the to the actual data, but also the year, the authors, and the GUIs when available, and also other information. And this functionality is very important because it ensures transparency and accountability about the information that OpenAIRE collected and enriched with its algorithms. But, with you I mean, this functionality if you're familiar with OpenAIRE Explorer, uh the things I presented until now is not really new. Mhm. Basically, uh we uh ex- we'll leverage on the information in the graph to provide this information also in the other portals that OpenAIRE provide. Uh but here in EOSC Explorer is not only about finding data and its context, but it's also about reusing data. So, what we did is that we integrated EOSC Explorer with the EGI data transfer service. And this allows us basically to um avoid uh how do you say? Okay, so instead of downloading large files to your local computer and then re-uploading them into uh a data a data storage, you can click transfer and move the data directly to your preferred cloud storage. So, briefly uh the user flow. Mhm. So, you select a data set in the discovery app and you open its landing page on OpenAIRE Explorer. You click on the transfer button and then you're asked to log in if you're not yet logged in. Uh in case the data set has multiple DOIs, uh because this functionality for now is available only when there is a DOI. So, if the data has more than one DOI, you have to pick one and typically I guess you will choose the DOI of the latest version, which is the one that is selected by default. And um And in some cases the data uh is um let's say it has different files in it, right? So, you can decide to transfer all all the files of the data set or only a subset of them. And at this point, you can select your preferred destination. For example, a storage on S3 provided by EGI or or a D-cache uh or whatever you have access to. So, um you select the destination system. And in this case, you need to know where uh you want to move the the data set. Uh if if needed, you have to provide other authentication info. And uh and the destination path. You accept the data protection policy, and um the system will automatically handle all the required interactions, and you will find your data in your selected destination. So, this is the the full uh uh the full path. Um Looking ahead, uh so we are continuing to grow the knowledge graph with new sources from our pilot nodes. And uh we will focus on on the specificities of the EOSC knowledge graph in one of the next community call. I think it's the one in October or November, but I think um we will know more after uh after my presentation. Uh the other things we are doing is that we would like to improve the discovery app in its sorting options. So, currently we have these uh built-in indicators. Um, you remember the one about uh impulse, popularity. Uh, so these are indicators that um are not yet used by the discovery app for sorting research products because now you can sort only by relevance and um number of views and downloads. So, we will add this new functionality to improve um the search functionality of the portal. And um and another important change will be at the methodological level. Hm. So, the project is shifting from a hub-based architecture where the where the EOSC-hub platform is somehow a centralized hub where all the pilots nodes register their services. But this is changing. Hm. We are changing to a mesh architecture and this means that the nodes will not register their services, data sources in the in the sandbox, in the EOSC-hub sandbox anymore, but they will uh do it on their own catalogs. And each catalog will be able to query with the other catalogs thanks to a federating capability of the uh of the federation. So, this is a big architectural change that is being implemented uh in these days and on the website of um of EOSC-hub, uh you can already test the the federated search. And um and this change, well, we require a change also in our graph construction methodology. Hm? Uh I don't want to explain the details of this architecture shift because there is a very good uh webinar that was held by our colleague of EOSC-Beyond in I think it was last week. Yes, 30th June, so 2 weeks ago. Uh And here again in the slides uh you can find the link and you can watch uh the webinar to know more. So, I think I ended my presentation and I'm happy to take any questions. Feel free to get in touch with EOSC-Beyond via the website or social media. And if you want, we can also use um the Discovery app and EOSC EOSC-Explore especially uh live if you have any specific questions on on the functionality and the opportunities that this service offer for you. Thank you. >> Thank you, Alessia, for your presentation. Um there are some questions already in the shared document, but I can see Helen um with her camera on. So, maybe if you have any questions and you'd like to comment on something, please feel free. >> Yeah, um I was trying to formulate my question so I could ask it clearly. Um I think I'm just trying to work out this activity looks fantastic. It's happening within the EOSC-Beyond project. How is that linking through to the activities that are happening in the build-up group and the working groups that are happening there? Can you say anything about that just to clarify? >> Yes, so um EOSC-Beyond per se is not um a candidate node of the EOSC. Mhm. So, it doesn't officially participate in the build-up phase. But, uh the project is involved in this in in this discussion, and the main partners are in fact also uh belonging to other candidate nodes. OpenAIRE itself. Now, we have this uh new node for uh scholarly communication scholarly commons. Um EGI has its own node, which is the coordinator. So, there are uh strict interaction and and conversation with the build-up group uh that is being taken into consideration. Then, let's remember that the build-up group, the EOSC EU node, and the EOSC Association, you know, they're trying to build something um operational production. Mhm. Here, we are in a research project. And this gives us a lot a lot of space to experiment, to try new things, and to go beyond what's currently, you know, under the under the plan of the of the build-up phase. And I mean, I think it's >> Okay, thank Thank you. I have I have one more question, um which is for those of you that know me, you're unsurprising around training. Um I noticed there is an option to filter for training materials, but it doesn't seem to retrieve anything. And again, my question comes back to are there There's a training and competencies working group in the build-up federation now that's looking at discovery and metadata for training resources. Is there any link up here, or is Can you Can you just say anything about the training part of of the discovery hub? >> Yes. So, we have the section of for for the training material, because it's one of the types of research outputs or the resources that we would like to to support. But the team focus more on the automatic deployment of service, the integration of services. And so the actual delivery of um of training material is somehow performed at the level of the pilot nodes that we have. So we we have materials to support them at the you know, at becoming compliant to the different technical guidelines that that are available for integrating services. Uh but we are let's say we are not yet preparing them for the training and the discovery hub. I don't know if there is someone from the from the pilot um >> I think that was my my concept my question really was about whether there is still the opportunity or whether work still needs to be done and and if we can join up the various places that this these discussions are happening around training discovery. >> Mhm. >> Thank you. >> No, thank you. Thank you. >> Um so Martha. >> Um Hi, Alessia was asking about whether there's any training. So so for sure any contributors to the discovery hub should be able to publish, right, training materials in in the knowledge graph as a research product itself. Now, how how is this being populated at the moment as Alessia said, the pilots are are working on on on deploying and providing the services and and it would be a very nice and next step kind of approach together in collaboration with any other communities out there to bring training materials to to how all of this is building up as well from the point of view of the content itself together with what Alessia was highlighting which is all the training materials on how the different types of of products services yeah deployable services are coming along. So it's two different aspects of the training there of the content in the platform and of the services offered by the sandbox itself. Um And and I have one question for Alessia. It's it's it's separate one if I can say change the the thing. You've mentioned that a data set can have duplicates because a data set can live in any different types of platforms the same data set they can can live in different places and you you will be able and you can detect duplicates. Is this detection of duplicates based on the DOI? And And and once you have like one data set you will have the same data set with multiple DOIs that have been harvested from different places. Is it in the provenance place that you show the different places where that data set comes from or is or was the provenance related to the version in only? Um >> Okay so yes the duplication algorithm also uses the DOI because the metadata can include DOI also different DOIs and this allows us to to be pretty sure that they are the same. Sometimes this doesn't happen. So for example from Zenodo we get the Zenodo DOI and from I don't know the Cesta catalog we get another DOI that was obtained by a via data site for example. And but working with the titles the authors and other metadata information we can be pretty sure if a data set is the same or not. Um And for the other question, so what when I showed that all the the provenance information, uh so usually when we collect a metadata record, uh it's only about one version. >> Right. >> So, what you see it was 12, 13 versions is because we collected 13 metadata records describing the same data set. And some [clears throat] of those can be from the same data source. Cuz maybe, you know, for example, Zenodo may uh or the CESSDA data catalog may expose one metadata record for version one, one metadata record for version two, and so on and so forth. >> Thanks. >> Thank you. >> Uh we have a few questions in the uh shared document. Uh so, the first question is, "How does the knowledge graph extract the metadata from respective resources? I imagine the metadata must be extensive in order to work properly. Is this metadata extraction AI-supported? And parentheses says, 'You mentioned automatic algorithm. Can you explain in more detail, please?'" >> Okay. So, yes, richer the metadata and uh better the graph is. So, uh in OpenAIRE, we are not just collecting anything. Mhm? We collect from trusted sources and um and we use a uh participatory approach. So, basically, we have um guidelines, metadata guidelines, that must be followed by providers that want to provide their content to OpenAire. And by being compliant to the guideline to the guideline, we ensure a minimum level of quality. Which is very important. But still, we cannot put, let's say that the line too high because that will exclude some trusted sources. So, and we don't we do not want to do that. So, what we can do is instead to use algorithms, also AI, yes, in order to enrich what we have. So, example of enrichment. Uh We collect a metadata record and it's about an open access article, for example. So, we have access to the full text. We can download it. And what we can do is to uh process the full text in order to insert additional properties like links to funding projects or links to data sets or links to software. And all these links go and enrich the graph. Um AI algorithms are also used for um document classification, for example. So, we are able to assign to classify uh an article and data set by fields of science and also by uh sustainable development goal. So, if an article is contributing somehow to a sustainable development goal uh as defined by the United Nation, then we also tag the article with that information. So, and these are only some examples. You you can find all the details about the algorithms we use um in the documentation of the open air graph. >> Thank you Alessia. And the next question is who defines the level of accessibility? The author? And how is this accessibility regulated? For example, if closed accessibility, then the data set is technically not accessible. How does this then work for limited accessibility? >> Yes, so of course it's the author that decides the level of accessibility. Because he's the one who is publishing the research outputs and decide if the data set can be openly shared or if it cannot. So think about a data set with um information about patients of a clinical studies. You may want to you want to put it in a repository for persistence, but you actually cannot from a legal point of view make it openly available. So you close it. And how the access work in this case, then it also depends on the repository that is used. Because in some cases the repository allows you to contact the author to request access. Uh but this is really dependent on the repository. >> May I add something? >> Please. >> Thank you very much for your answer. I'm asking this because we are in a project where we combine open science and intellectual property and there we uh >> [clears throat] >> or from this project we know that within a research project a lot of different outcomes will be generated. Data sets, codes, all the connected uh information. And sometimes it's important that for one part, one research artifact, let's say data set, you have to enable the accessibility, whereas you have to close the accessibility for another data set or uh whatever. Um is it in this case possible to somehow have the list of the main outcomes and all the subsequent outcomes artifacts uh in order to annotate different accessibility levels to all these different artifacts? >> Okay, so so this kind of decision are not to be taken let's say in in explorer or in the discovery hub, because it's when you deposit the data set somewhere that you have to fill in the the information, so you will fill in the title, the authors, and the system will ask you about the accessibility level and and the license probably, hopefully, and this is where you as the author will have to decide if you have to keep it closed or if you can open. What you can see in the graph is so the graph just reflect what the author decided in the repository. What I think would be very useful and oops, I I think you you already did it probably, is to have a data management plan for your project, where you can, you know, already plan these kind of things. >> Exactly, this is what we are arguing for, but I often see for example as a node publication with 10 subsequent uh files, yeah? >> Mhm. >> And I suppose that uh each of them need a different level of accessibility, which I think is not possible if you put them together into one, let's say, in this case, a Zenodo publication. >> Yes, indeed. Um in Zenodo, each deposition has one access right and one license. Mhm. So, if you have different files with different license, you have to do different deposition. And then you you can link them each other. >> Exactly. This was an example from Zenodo, but how does it work uh for this hub? Is that yeah, the same methodology, the same structure? Or can you separate? >> No, no, no. We will reflect uh what you have in Zenodo. So, we will have one record for one data set linked with another one. And then you can navigate from one to another. And then there could be mistakes. I mean, if if they if these data sets have all the same titles and all the same authors, then our algorithm may uh identify them as one. We put them together. But in this case, if you tell us, we have ways to, let's say, >> Mhm. >> divide them. >> Mhm. Perfect. Thank you. >> And there's also one more question from Marie um regarding their usability, transfer of data sets to another storage. What is the technical pathway of the transfer from the knowledge graph to storage XYZ if the data set has not been down and uploaded locally? >> Yes, that's how the uh EGI that France transfer service works. So, basically, when you click transfer uh the the service takes an input that the DOI and the list of files that you that you provided. And basically let's suppose that the source of the data is Zenodo. So, the data transfer will ask for the files from Zenodo and will not store it in your computer. Will go directly on the selected destination. Then if you want more specific technical details on the service I'm afraid I cannot provide them. We should ask our colleagues from EGI and CERN who are the developer and maintainer of the service. Uh but Yeah. >> And I completely understand because I'm also not technically equipped, but I'm asking this question more from Yeah. Yeah. Not from an technical background, but um yeah, from the background of social sciences. Yeah. Thank you. >> Thank you. >> Thank you for your questions, Marie. And we have one a last question in the document. Which cloud data storage are compatible with EOSC Explorer transfer functionality? In which format is metadata transferred? >> Okay, so let me go back to the slide because I listed the type of storage that are supported. supported that Yeah, so currently the supported options are uh S3 S3 over HTTPS Storm and DCache. >> Thank you, Edison. >> There was a second part in the question. >> Yes. In which format is made a data transferred? >> No, I think only the actual data is transferred. >> Okay, thank you again for your advice. There are no more other questions, but if you have any thoughts or comments, please feel free to raise your hand. I can see no activity. So, we can inform you that our next Open Air Graph community call is in September, 16th of September. And the topic will be announced soon, so stay updated and we'll share with you the slides and the recording of today's session. If you haven't gone for your summer holidays, we hope you have a very nice time and see you again in September. Thank you and thank you Alessi for your presentation. >> Thank you. Thank you all.