Video summary
The EOSC Beyond project is a 36-month initiative funded by the European Commission, coordinated by EGI with over thirty partners across Europe. Its primary objective is to establish the next generation of core services for the European Open Science Cloud (EOSC) by facilitating the seamless federation and interaction between national, regional, and thematic nodes within a unified technical architecture. A central component of this infrastructure is the EOSC Core Innovation Sandbox, which serves as a pre-production environment where nodes can verify their maturity against federal requirements while testing integration capabilities with AI, accounting systems, and order management services. For users, the project offers a single entry point known as the user space, allowing them to authenticate once and access various functionalities such as data transfer tools, automatic infrastructure deployment requests, and discovery mechanisms for research resources like datasets, publications, software, and interoperability guidelines.
At the heart of this ecosystem lies the Open Research Graph, which acts as a foundational knowledge graph linking diverse research outputs with service catalog entries. The Discovery Hub serves as the front door to these resources, enabling users to search through advanced filters based on criteria such as data sources, EOSC nodes, and specific resource types like services or deployable applications. When a user selects a dataset within this hub, they are directed to EOSC Explorer, where the power of the knowledge graph becomes evident by providing rich contextual information beyond basic metadata. This includes details about associated publications, author ORCID identifiers, funding projects, related research communities, and impact indicators such as citations and popularity metrics. Furthermore, the system aggregates data from multiple sources to handle versioning and provenance, allowing users to view different versions of a dataset alongside their specific licenses and access restrictions, ensuring transparency regarding how metadata was collected and enriched by algorithms.
A significant advancement in this framework is the integration with EGI's data transfer service, which streamlines the reuse of research outputs without requiring local downloads. Instead of transferring large files through a user's computer to another storage location, users can initiate a direct cloud-to-cloud transfer from within EOSC Explorer. This process involves selecting a dataset, authenticating if necessary, choosing between multiple DOIs for datasets with various versions, and specifying the destination system such as S3, Storm, or DCache provided by EGI partners. The system automatically manages these interactions to move data directly to the user's preferred cloud storage while adhering to applicable data protection policies. Looking ahead, the project plans to expand its knowledge graph using new sources from pilot nodes and is transitioning from a centralized hub architecture to a mesh-based federated search model, allowing individual node catalogs to query each other independently rather than registering all services in a central sandbox.
During the community call, several questions were addressed regarding the operational details of these systems, including how metadata extraction utilizes AI algorithms for enrichment tasks like document classification and sustainable development goal tagging while maintaining quality through participatory guidelines. Discussions also clarified that accessibility levels are determined by authors during deposition processes rather than within the explorer interface itself, meaning that datasets with mixed access rights must be deposited separately to reflect different licenses accurately. Additionally, it was confirmed that training materials can theoretically be published as research products in the graph but currently rely on pilot nodes for delivery and compliance support. The session concluded with an announcement of the next community call scheduled for September 16th, where further updates on these evolving architectural changes and new topics will be shared with the broader OpenAIRE Graph Community.
Read the full video transcript
Okay, so let's start. This is our um
agenda for today.
So, I'll start with a brief introduction
to the EOSC Beyond project.
And then we'll see the foundational role
that the Open Research Graph plays in
that context.
And and then we'll do a deep dive in the
Discovery Hub and EOSC Explore.
So, let's start with a background.
EOSC Beyond is a 36-month
project funded by the European
Commission.
We are a large consortium, 30 33
partners. The coordinator is EGI.
And uh we have 16 work packages and 10
pilot nodes. And our goal is to build
the next generation of EOSC core
services.
So, uh
I will not go in into the details of
these slides. Uh
but uh basically, what I want to say is
that our primary goal in EOSC Beyond is
to support the setup and the federation
of EOSC nodes. Mhm. So, we are
developing
uh
tools, frameworks, and an infrastructure
to help national, regional, and thematic
nodes to interact seamlessly
within a unified uh technical
architecture.
And a central piece of this
infrastructure is the EOSC Core
Innovation Sandbox,
through which nodes can uh
verify their maturity uh according to
the technical requirements of the EOSC
uh
federation.
And
it is conceived as a pre-production
environment. So, a place where nodes and
users can test and experiment with the
core services of the European Open
Science Cloud and the federating
capabilities.
In particular, for nodes and service
providers,
the sandbox offers
technical support for integration with
the AI, with the accounting, the order
management, uh
and all the services that are
the federating capabilities of the EOSC.
While for the users, we have
what we call the user space, which is a
single entry point where you can log in
with the essential credentials, with the
EOSC AI,
and you can test, you can use data
transfer capabilities, you can request
the automatic deployment of
infrastructure and services.
And specifically of interest today,
uh discover resources, services,
datasets, and other research products.
And it's here that uh
the Open Research Graph is a
foundational component.
Uh so, what you will see in the
discovery app, here in the middle, um is
a subset of the Open Research Graph.
And um
we link the resources of the Open
Research Graph to the resources in the
service catalog, which includes
services, but also interoperability
guidelines, adapters, deployable
applications. So, we link all these
together,
and we make it available in the
discovery app.
Then,
if you select, for example, a a service,
you basically
um open the details of a service in the
front office and order management and
this allows you to request the automatic
deployment of the service
if it's available.
While if the user selects instead
a research output so for example a data
set
then you end up in EOSC Explorer. Mhm.
Uh really briefly
the methodology for building this EOSC
Beyond Knowledge Graph. So what we do is
that
um
we take the OpenAIRE Graph
we select the research outputs
that comes from data sources registered
in the EOSC Beyond
catalog
and we enrich those research products
with specific with a specific PIDs so
the persistent identifiers of the nodes
and of the registered data sources. So
then this is where we create the
connection between uh the
the service catalog and
and the and the OpenAIRE Graph.
Um
So
the Discovery Hub as I said the
Discovery Hub is the the entry point is
the front door to all these different
types of resources.
Uh you can use the the basic search or
the advanced search um and you can also
apply filters here. I don't know if you
see my mouse but
there are filters on the left that
allows you to filter your search results
based on different criteria. So one is
data source the other one is EOSC node
the two I mentioned before
but there are also other criteria. Mhm.
And
you can also uh see here in the in the
menu in the middle of the of the user
interface
uh that there are there are let's say
menus to easy switch between different
types of resources. So, you have
research outputs
data sets, publications, and software.
You have uh service offerings with
deployable applications and services.
And
among the deployable applications
you
you found those that can be auto
automatically uh deployed on on
infrastructure of the EOSC uh
um
of the EOSC beyond.
Uh
you can search for organizations and
data sources, and also for
interoperability guidelines, and um
and adapters.
Now,
um
let's focus
on the heart of our research discovery,
which is EOSC explore. So, EOSC explore
is the part
of the of the user space that is
developed by uh
by opener.
Uh so, when you are in the discovery
hub, you find uh a data set
that you are interesting to. Uh and when
you click to And already here you can
see a lot of information, huh? Because
you have the information well basic
metadata, but also
number of downloads, views, funders.
But of course, we can provide even more
information. And to do that, you just
click on the on the card, and then
this is where
the power of uh the knowledge graph
shines, because we have a lot of
information here
that can be used.
So,
we have an overview, so the
metadata overview, maybe it's not the
right name, but the basic metadata of
the of the data set.
And uh
but it's not only about the data set.
It's also about its context.
So, for example, you can see the main
publication that is associated with a
with this specific data set.
Mhm. So, the the primary uh publication
of the data set.
And this allows you to understand
the results in the context of the
original study.
We also have
uh one We also have links to the ORCID
identifiers of the authors.
Uh
in some cases, we also have uh citations
and and the links to funding project.
This is the case. This was funded by
EOSC Future.
We also have links to
uh
related research communities. These are
research communities that uh collaborate
somehow with OpenAir.
And
in this case, this data set was also
classified by fields of science, by the
automatic uh algorithm of OpenAir.
Uh here instead, we highlight the the
impact indicators. Mhm. We have metrics
for citations, popularity, influence,
and um
this gives you a quick way to assess um
the reach, the impact, and the relevance
uh of the data set, together with the
views and downloads, which are instead
based on the uh usage counts service of
OpenAir.
Very important
is information about provenance and
versioning.
Because the OpenAIRE Graph aggregates
metadata from multiple sources,
we can basically it can collect
different metadata records about the
same research product from different
places.
So, what we have is a duplication
algorithm that finds this duplicate, the
different version, and we put them
together.
So, in this case for example,
if you click
in these in these parts, you have the
list of all the different versions that
we have. We highlight which are which is
the license, and if you click for
example on Easy, you will go
to the
to the landing page of Easy where you
can actually download
the data itself.
When you see the
orange open lock, it means that the
version is open access.
There may be versions that are not open
access. They can be closed or they can
be restricted, meaning that
you need for example to log in in order
to access. And that depends
the type of access of course depends on
the on the specific
source, on the specific version.
Then
let me go to the next slide. Yes, so if
we click on view all versions, here we
have the details of each version.
So, not only the link that goes to the
to the actual data, but also
the year, the authors,
and the GUIs when available, and also
other information.
And this functionality is very important
because it ensures transparency
and accountability about the information
that OpenAIRE collected and enriched
with its algorithms.
But,
with you I mean, this functionality if
you're familiar with OpenAIRE Explorer,
uh
the things I presented until now is not
really new. Mhm. Basically,
uh we uh ex-
we'll leverage on the information in the
graph to provide this information also
in the other portals that OpenAIRE
provide.
Uh
but here in EOSC Explorer is not only
about finding data and its context, but
it's also about reusing data.
So, what we did is that we integrated
EOSC Explorer with the EGI data transfer
service.
And this allows us basically
to
um
avoid uh how do you say? Okay, so
instead of downloading large files to
your local computer and then
re-uploading them into uh
a data a data storage,
you can click transfer and move the data
directly to your preferred cloud
storage.
So,
briefly uh the user flow. Mhm.
So,
you select a data set in the discovery
app
and you open its landing page on
OpenAIRE Explorer.
You click on the transfer button
and then you're asked to log in if
you're not yet logged in.
Uh in case the data set has multiple
DOIs, uh because this functionality for
now is available only when there is a
DOI.
So,
if the data has more than one DOI, you
have to pick one and typically I guess
you will choose the DOI of the latest
version, which is the one that is
selected by default.
And um
And in some cases
the data
uh is um let's say it has different
files in it, right? So, you can decide
to transfer all all the files of the
data set or only a subset of them.
And at this point, you can select your
preferred destination. For example,
a storage on S3 provided by EGI
or or a D-cache
uh or whatever you have access to.
So, um
you select the destination system.
And in this case,
you need to know where
uh you want to move the the data set. Uh
if
if needed, you have to provide other
authentication info.
And uh and the destination path.
You accept the data protection policy,
and
um
the system will automatically handle all
the required interactions,
and you will find your data
in your selected destination.
So, this is the the full uh
uh the full path.
Um
Looking ahead, uh so we are continuing
to grow
the knowledge graph with new sources
from our pilot nodes.
And uh we will
focus on
on the specificities of the EOSC
knowledge graph in one of the next
community call. I think it's the one in
October or November, but I think um
we will know more after uh
after my presentation.
Uh
the other things we are doing is that we
would like to improve the discovery app
in its sorting options. So,
currently we have these uh built-in
indicators. Um, you remember the one
about uh impulse, popularity.
Uh, so these are indicators that um
are not yet used by the discovery app
for sorting research products because
now you can sort only by relevance and
um number of views and downloads. So, we
will add this new functionality to
improve
um
the search functionality of the portal.
And
um
and another important change will be at
the methodological level. Hm.
So, the project is shifting from a
hub-based architecture where
the
where the EOSC-hub platform is somehow a
centralized hub
where all the pilots nodes register
their services.
But this is changing. Hm. We are
changing to a mesh architecture and this
means that the nodes
will not register their services, data
sources in the in the sandbox, in the
EOSC-hub sandbox anymore, but they will
uh do it on their own catalogs.
And each catalog will be able to query
with the other catalogs thanks to
a federating capability of the uh of the
federation.
So, this is a big architectural change
that is being implemented uh in these
days and on the website of um
of EOSC-hub, uh you can already test the
the federated search.
And um
and this change, well, we require a
change also in our graph construction
methodology. Hm?
Uh
I don't want to explain the details of
this architecture shift because there is
a very good uh webinar that was held by
our colleague of EOSC-Beyond
in I think it was last week. Yes, 30th
June, so 2 weeks ago.
Uh
And here again in the slides uh you can
find the link and you can watch uh the
webinar to know more.
So,
I think I ended my presentation and I'm
happy to take any questions.
Feel free to get in touch with
EOSC-Beyond via the website or social
media.
And if you want, we can also use um
the Discovery app and EOSC EOSC-Explore
especially uh live if you have any
specific questions on
on the functionality and the
opportunities that this service offer
for you.
Thank you.
>> Thank you, Alessia, for your
presentation.
Um there are some questions already in
the shared document, but I can see Helen
um
with her camera on. So, maybe if you
have any questions and you'd like to
comment on something, please feel free.
>> Yeah, um
I was trying to
formulate my question so I could ask it
clearly. Um I think I'm just trying to
work out this activity looks fantastic.
It's happening within the EOSC-Beyond
project. How is that linking through to
the activities that are happening in the
build-up group and the working groups
that are happening there? Can you say
anything about that just to clarify?
>> Yes, so um EOSC-Beyond per se is not um
a candidate node of the EOSC. Mhm. So,
it doesn't officially participate in the
build-up phase.
But, uh the project is involved in this
in in this discussion, and the main
partners are in fact also uh belonging
to other candidate nodes. OpenAIRE
itself. Now, we have this uh new node
for uh
scholarly communication scholarly
commons.
Um EGI has its own node, which is the
coordinator. So, there are uh strict
interaction and and conversation with
the build-up group
uh that is being taken into
consideration.
Then, let's remember that the build-up
group, the EOSC EU node, and
the EOSC Association, you know, they're
trying to build something um
operational production. Mhm.
Here, we are in a research project.
And this gives us a lot a lot of space
to experiment, to try new things, and to
go beyond
what's currently, you know,
under the under the plan of the of the
build-up phase.
And I mean, I think it's
>> Okay, thank
Thank you. I have I have one more
question, um which is
for those of you that know me, you're
unsurprising around training.
Um I noticed there is an option to
filter for training materials, but it
doesn't seem to retrieve anything. And
again, my question comes back to
are there There's a training and
competencies working group in the
build-up federation now that's looking
at
discovery and metadata for training
resources. Is there any link up here, or
is Can you Can you just say anything
about the training part of of the
discovery hub?
>> Yes. So, we have the section of for for
the training material, because it's one
of the types of
research outputs
or the resources that we would like to
to support.
But
the team focus more on the automatic
deployment of service, the integration
of services.
And
so the actual delivery of um
of training material
is somehow performed at the level of the
pilot nodes that we have.
So we we have materials to support them
at the
you know, at becoming compliant to the
different technical guidelines that that
are available for integrating services.
Uh
but we are let's say
we are not yet preparing them for the
training and the discovery hub. I don't
know if there is someone from the
from the pilot um
>> I think that was my my concept my
question really was about whether there
is still the opportunity or whether work
still needs to be done and and
if we can join up the various places
that this these discussions are
happening around training discovery.
>> Mhm.
>> Thank you.
>> No, thank you. Thank you.
>> Um so Martha.
>> Um
Hi, Alessia was asking about whether
there's any training. So so for sure
any contributors to the discovery hub
should be able to publish, right,
training materials in in the knowledge
graph as a research product itself. Now,
how how is this being populated at the
moment as Alessia said, the pilots are
are working on on on deploying and
providing the services and and it would
be a very nice and next step kind of
approach together in collaboration with
any other communities out there
to bring training materials to to how
all of this is building up as well from
the point of view of the content itself
together with what Alessia was
highlighting which is all the training
materials on how the different types of
of products services yeah deployable
services are coming along. So it's two
different aspects of the training there
of the content in the platform and of
the services offered by the sandbox
itself. Um
And and I have one question for Alessia.
It's it's it's separate one if I can say
change the
the thing.
You've mentioned that
a data set can have duplicates
because a data set can live in any
different types of platforms the same
data set they can can live in different
places and you you will be able and you
can detect duplicates. Is this detection
of duplicates based on the DOI?
And
And and once you have like one data set
you will have the same data set with
multiple DOIs that have been harvested
from different places. Is it in the
provenance place that you
show the different places where that
data set comes from or is or was the
provenance related to the version in
only?
Um
>> Okay so yes the duplication algorithm
also uses the DOI
because the metadata
can include DOI
also different DOIs and this allows us
to to be pretty sure that they are the
same.
Sometimes this doesn't happen. So for
example from Zenodo we get the Zenodo
DOI and from I don't know
the Cesta catalog we get another DOI
that was obtained by a via data site for
example.
And but
working with the titles the authors and
other metadata information
we can be pretty sure if a data set is
the same or not.
Um
And for the other question, so what when
I showed that
all the
the provenance information,
uh
so usually when we collect a metadata
record,
uh it's only about one version.
>> Right.
>> So, what you see it was 12, 13 versions
is because we collected
13 metadata records describing the same
data set.
And some [clears throat] of those can be
from the same data source.
Cuz maybe, you know, for example, Zenodo
may uh
or the CESSDA data catalog may expose
one metadata record for version one, one
metadata record for version two, and so
on and so forth.
>> Thanks.
>> Thank you.
>> Uh we have a few questions in the uh
shared document. Uh so, the first
question is, "How does the knowledge
graph extract the metadata from
respective resources?
I imagine the metadata must be extensive
in order to work properly.
Is this metadata extraction
AI-supported?
And parentheses says, 'You mentioned
automatic algorithm. Can you explain in
more detail, please?'"
>> Okay.
So, yes, richer the metadata and uh
better the graph is. So,
uh
in OpenAIRE, we are not just collecting
anything. Mhm? We collect from trusted
sources and um
and we use a
uh participatory approach. So,
basically, we have um
guidelines, metadata guidelines, that
must be followed by providers that want
to
provide their content to OpenAire.
And by being compliant to the guideline
to the guideline, we ensure
a minimum level of quality.
Which is very important.
But still, we cannot put, let's say that
the line too high because that will
exclude
some trusted sources. So, and we don't
we do not want to do that.
So, what we can do is instead to use
algorithms, also AI, yes, in order to
enrich what we have.
So,
example of enrichment.
Uh
We collect a metadata record and it's
about an open access article, for
example. So, we have access to the full
text. We can download it.
And what we can do is to
uh process the full text in order to
insert
additional properties like links to
funding projects
or links to data sets
or links to software.
And all these links go and enrich the
graph.
Um
AI algorithms are also used for um
document classification, for example.
So, we are able to assign to classify
uh an article and data set by fields of
science
and also
by
uh sustainable development goal. So,
if an article is contributing somehow to
a sustainable development goal
uh
as defined by the United Nation, then we
also tag the article with that
information.
So, and these are
only some examples. You you can find all
the details about the algorithms we use
um in the documentation of the open air
graph.
>> Thank you Alessia. And the next question
is who defines the level of
accessibility? The author? And how is
this accessibility regulated? For
example, if closed accessibility, then
the data set is technically not
accessible. How does this then work for
limited accessibility?
>> Yes, so
of course it's the author that decides
the level of accessibility.
Because he's the one who is publishing
the research outputs and decide if the
data set can be openly shared
or if it cannot.
So think about a data set with um
information about patients of a clinical
studies.
You may want to you want to put it in a
repository for persistence, but
you actually cannot from a legal point
of view make it openly available. So you
close it.
And
how the access work in this case, then
it also depends on the repository that
is used.
Because in some cases the repository
allows you to contact the author
to request access.
Uh but this is really dependent on the
repository.
>> May I add something?
>> Please.
>> Thank you very much for your
answer.
I'm asking this because we are in a
project where we combine open science
and intellectual property and there we
uh
>> [clears throat]
>> or from this project we know that within
a research project a lot of different
outcomes
will be generated. Data sets, codes, all
the connected uh information. And
sometimes it's important that for one
part, one research artifact, let's say
data set,
you have to enable the accessibility,
whereas you have to close the
accessibility for another data set or
uh whatever.
Um is it
in this case possible to somehow
have the list of
the main outcomes and all the subsequent
outcomes artifacts uh in order to
annotate different accessibility levels
to all these different artifacts?
>> Okay, so so this kind of decision are
not to be taken
let's say in in explorer or in the
discovery hub, because it's when you
deposit the data set somewhere that you
have to fill in the
the information, so you will fill in the
title, the authors, and the system will
ask you about the accessibility level
and and the license probably,
hopefully, and this is where you as the
author will have to decide if you have
to
keep it closed or if you can open.
What you can see in the graph is
so the graph just reflect what the
author decided in the repository.
What I think would be very useful
and oops, I I think you you already did
it probably, is to have a data
management plan for your project, where
you can, you know, already plan
these kind of things.
>> Exactly, this is what we are arguing
for, but I often see for example as a
node publication with 10 subsequent
uh files, yeah?
>> Mhm.
>> And I suppose that uh each of them need
a different level of accessibility,
which I think is not possible if you put
them together into one, let's say, in
this case, a Zenodo publication.
>> Yes, indeed. Um
in Zenodo, each deposition has one
access right and one license. Mhm.
So, if you have different files with
different license, you have to do
different deposition.
And then you you can link them each
other.
>> Exactly. This was an example from
Zenodo, but how does it work uh for this
hub? Is that
yeah, the same methodology, the same
structure?
Or can you separate?
>> No, no, no. We will reflect uh what you
have in Zenodo. So, we will have one
record for
one data set linked with another one.
And then you can navigate from one to
another.
And then there could be mistakes. I
mean, if if they if these data sets have
all the same titles and all the same
authors, then our algorithm may uh
identify them as one. We put them
together. But in this case, if you tell
us, we have ways to,
let's say,
>> Mhm.
>> divide them.
>> Mhm.
Perfect. Thank you.
>> And there's also one more question from
Marie um
regarding their usability, transfer of
data sets to another storage. What is
the technical pathway of the transfer
from the knowledge graph to storage XYZ
if the data set has not been down and
uploaded locally?
>> Yes, that's how the uh EGI that France
transfer service works.
So, basically,
when you click
transfer
uh the
the service takes an input that the DOI
and the list of files that you that you
provided.
And
basically
let's suppose that
the source of the data is Zenodo. So,
the data transfer will ask
for the files from Zenodo
and will not store it in your computer.
Will go directly on the selected
destination.
Then if you want more specific technical
details on the service I'm afraid I
cannot provide them. We should ask our
colleagues from EGI and CERN who are the
developer and maintainer of the service.
Uh but
Yeah.
>> And I completely understand because I'm
also not technically equipped, but I'm
asking this question more from
Yeah. Yeah. Not from an
technical background, but
um yeah, from the background of social
sciences. Yeah.
Thank you.
>> Thank you.
>> Thank you for your questions, Marie. And
we have one a last question in the
document.
Which cloud data storage are compatible
with EOSC Explorer transfer
functionality? In which format is
metadata transferred?
>> Okay, so let me go back to the slide
because I listed the type of storage
that are supported.
supported that
Yeah, so currently
the supported options are
uh S3
S3 over HTTPS
Storm and DCache.
>> Thank you, Edison.
>> There was a second part in the question.
>> Yes. In which format is made a data
transferred?
>> No, I think only the actual data is
transferred.
>> Okay, thank you again for your advice.
There are no more other questions, but
if you have any thoughts or comments,
please feel free to raise your hand.
I can see
no activity. So, we can inform you that
our next Open Air Graph community call
is in September, 16th of September.
And
the topic will be
announced soon, so stay updated and
we'll share with you the
slides and the recording of today's
session. If you haven't gone for your
summer holidays, we hope you have a very
nice time and see you again in
September. Thank you and thank you
Alessi for your presentation.
>> Thank you. Thank you all.