Video summary
Mark Porter from VLIOS presents a compelling vision for advancing open science beyond simple checklists like FAIR principles by introducing five distinct work zones designed to make global research data as easily searchable and accessible via natural language queries as Google. His primary goal is to eliminate the 79% of time currently wasted on avoidable overheads, transforming how scientists manage and share information. The first zone focuses on vocabulary management, advocating for unambiguous terms with unique identifiers that act like web addresses to ensure data is understandable globally while supporting citizen science through local slang and translations. This approach moves away from fragmented repositories toward unified data packaging using open standards such as Research Object Crates (RO-Crates), which combine data, metadata, and semantics into a single archive capable of handling both cloud and local storage environments.
To further streamline the research process, Porter proposes transforming static Data Management Plans into dynamic platforms with desktop-like interfaces that automatically guide researchers in applying rules and generating compliant RO-Crates without relying on closed formats. This evolution is supported by linked data publishing principles where web services conform to semantic design standards like Hydra CG, creating a "glass box" transparency that allows consumers to see identical interfaces regardless of underlying technical differences. By promoting distributed solutions over centralized hubs, these zones aim to address the fragmentation caused by scattered workflows and ensure that semantics are not overlooked in favor of raw metadata alone. The presentation also highlights the importance of usage tracking and metrics, suggesting tangible incentives based on semantic data analysis rather than just moral arguments or traditional publication counts.
Addressing concerns about incentivizing specific frameworks versus reinventing solutions, Porter argues that while top-down mechanisms work for distributing funds, real innovation stems from bottom-up grassroots efforts where practitioners solve actual problems. He advises against relying solely on enforcement laws to adopt public administration funding results and instead encourages starting with local investments in smaller elements like shared vocabularies and tools such as Autocrate. Although some initiatives are not yet fully adopting linked data standards, these incremental steps can effectively bridge gaps where semantics are often neglected. The session concludes by inviting audience collaboration through shared documents and chat participation to explore synergies between different standards while acknowledging that the recording ended due to a lack of further questions.
Read the full video transcript
my name
is mark porter and i'd like to welcome
to my home
for this webinar on the work zones in
open science
thanks for showing up and a big thanks
also to open belgium for bringing us
together today
the slides are available and free for
reuse under creative comments
but i do appreciate your attribution and
your determination to share a like
i have a limited amount of time and way
too much content
so let us quickly get some slides out of
the way
since one year and some days i work at
vliss
which is the flanders marine institute
we are 20 years old and based in
austin's
on the belgian coast our mission is to
enable and advance marine research
in all possible domains and that ranges
from biology geography
ecology and even includes history
medicine
all the way up to psychology we do this
to promote an advanced society
in general and therefore we target
various audiences
that we see as important stakeholders in
and around
the benefits and responsibilities for
our beloved seas
narrowing to mountain narrowing down to
my role
i only have time to mention the
department i work for
which is the vmdc in english that
acronym
expands to the flanders marine data
center our focus
is on data data management and the
design of data systems to support it
we participate in many in a lot of
projects
and all of them have different levels of
geographical coverage and outreach
most notably i think we publish a number
of data products
for the marine research domain and those
are ranging from something like
the skeleta monitor which shows our
local expertise and connection
but it scales up via the european level
for projects like
ruby's which tracks marine biodiversity
but even spans the globe with the
reference databases we manage
and govern four marine species and
marine regions
and concerning open science we are
active on the flemish level
inside the flemish research data network
and the flemish open science board
and in europe we participate in projects
of the
european open science cloud
within the vmdc then together with three
colleagues we make up
the open science team we make sure the
organization as a whole
realizes and keeps furthering its open
ambitions
so what we are doing is implementing the
fair principles
advocating to and supporting research
groups
and adapting the data systems to be open
by design
in doing so we very much like and
embrace open source
and linked data technologies and if this
sounds appealing to you
keep an eye on the slash jobs page of
the vlis website because we are
working towards an extra job opening in
the team
and this is me and the various ways uh
to contact me or find me ah yes this
slide
and some form or another this has been
in
every public speaking slide deck i've
given since 2009
honestly honestly and honestly the web
is totally awesome
on so many levels and i think we should
regularly
have parties just to remain consciously
aware of that
it turns out this slide keeps being
relevant so i keep it in
and it surely will be today so we will
come back to this
fabulous web right uh no all of that is
out of the way
here is the actual agenda for today uh i
see this as going to be uh six main
blocks
the first one to lay down some
introduction and overview
of what open science is and then five
actual identified work zones
that could turn out to be too much but
let us be ambitious
finally this is 2021 and we do this
online
so i don't hear nor see you and the
point is
we should uh we would we would very much
like your feedback
even more than only your feedback we
actually want to reach out
and recruit some of you to collaborate
on some of these tasks
and to achieve that goal uh i would like
to ask you to use this co-writing space
we have foreseen
and which should have been shared in the
chat also
uh and also at the end you might prepare
yourself to open up your own mic
for our audience participation part
also some of my colleagues are around so
they might interact on the chat
while i am talking um
an experiment i did an effort to include
qr codes
for all the links i mentioned so if you
have a device with a keyword scanner
around you can follow up on those as
well
and our last tip my voice is mostly not
going to read out what is on the slides
okay i expect you to be able to do that
for yourself
and allow my comments to be loosely
providing
additional insights around that
structure
so on with the show when i talk about
open science
i almost always end up doing some fair
bashing
if you haven't heard about it the fair
acronym
is the first thing you learn when coming
to the open science field
it's a bit like a rite of passage or a
painful tattoo
something like a required code to become
part of this clip
and to sign into its belief system the
letters
spell the expectations for all data and
publications
all of them should be made findable
accessible
interoperable and reusable don't get me
wrong
it is a very smart acronym and a very
suitable top-level shopping list
of the things to cover however it is
also becoming its own purpose
most importantly in my personal opinion
it is too often translated
into a simple checklist and one that
allows people to declare
we have done enough to reach the end of
their responsibility
they just need to check all four boxes
as quickly as possible
so instead of taking that route i want
to invite you to a more open and unbound
thinking about open science
this is what we take as the guiding user
story for the work we do
to provide a relevant slice of global
research of the global research data set
that is as simple as a google search
we actually package this in in our gyro
as bug number one for the complete team
and we believe it's probably going to
take us
uh uh until we retire to to get it done
we see this as even having two parts the
first part is about
the natural way of finding the data if
you go to
google today you will be able to have it
answer quite convoluted questions like
an example who is the actress playing
the wife of stephen hawking
or who is the director of that movie
this kind of natural way to find is
precisely what we would like to achieve
for the data lookups
in the realm of scientific work and on
the left side of the slide
you see a number of examples in the
marine domain of how such
searches could be uh expressed
in contrast on the right side of the
slide we have listed all the steps that
are involved today to actually
answer those questions and if you look
carefully behind
semi-transparent behind the list in the
background
you see a big pie chart and that brings
this
to my next slide it turns out that no
less that
79 of the working day of a data
scientist
is wasted on those steps all of them are
avoidable and overhead issues
and if this 79 could be reduced to 18
it would shrink their working day to
only one-third
if we turn that around we could make
every full-time hire
of one data scientist count as three
and we believe it is precisely this kind
of upscaling that open science
should aim for society is funding the
research
so logically it keeps the right to
expect a higher return on investment
open science often serves on a way of
high ethical motives
but we should not be shy to claim
economic goals as well
we should make and deliver on a promise
to reduce
the cost of doing research the next
open science doesn't stop at targeting
the science community alone
with the second part of our ambitious
bug we want to stress the need to target
the general public
the research community needs to push its
quality information out on that same web
everybody has learned to use it has to
be available there
and it has to be made high ranking
at the bottom you'll find my catchy
twitter code on the
quote on the subject if you agree scan
the cover code and give it some of your
twitter love
and remember this event is using the
hashtag openbelgium21 so now we have a
challenging job
and we see two parts in it and in our
head both
parts lead us to the web first as an
example of
easy and natural search but also as the
platform
to reach any everyone out there
so we believe the backbone of this
fantastic web is the foundation for open
science
especially uh the principles of the
semantic web
which brings us to our this we call the
ikea
observation slide it should explain to
you the subtle difference between
uh semantics and metadata and it also
shows the connected unity
of this holy trinity it is always data
metadata
and never forget the third one semantics
we need them all and they should be kept
together
okay i mentioned the smart google search
earlier
but you might remember this scenario
it's less than a year ago you have
booked a flight
and suddenly your email client receives
a confirmation email
and it is is successful in assisting you
to add that event to your calendar
and you have to understand it is not an
effect of some spooky artificial
intelligence or any big brother stuff
i know it looks like they know about
your plans but in reality
this effect is driven through clear and
straightforward semantic annotations
inside the message inside the message
there are just some schema.org
terms hidden inside the html the html of
your email
and those allow your mail clients to
understand
part of the information that you are
reading and magical effects like these
are
easily achievable when we add semantics
to the data
and that is what the phrase making data
machine actionable
really is about and then going back to
research
i think maybe even worse than in other
domains the attention for semantics has
been absent
and allow me to put that in context the
required understanding of meaning and
intentions has not been neglected
rather it has simply been kept implicit
they assume it to be known among the
smart peers
but in reality in most cases you have on
the one side a knowledgeable producer of
the data
and it talks only to the unknowing
random external consumer
via some intermediate data system and it
is therefore essential that all
insights and understanding all semantics
are captured
and made explicit to conclude
all of this shows the access part of
open science
is only the tip of the iceberg you might
remember that richard stallman
from the free software foundation at
some point in time clarified
that free and free software should be
read as freedom
or free speech and not as free beer
in the same way i think the open in open
science
really refers to open-ended and not too
open cans
open is about taking up the extra work
the data is not
released only to repeat your own
findings and experiments
instead it needs it needs some extra
work it must be
prepared so it can be rehashed
reassembled recombined
in future new contexts their original
meaning should be kept
but their reuse should allow solving new
problems
and this also comes back to targeting
audiences
outside the inner circle of the new
research community
the shared data should function as an
open invitation to three new classes of
players
one is the scientists from other domains
then you have the citizen scientists
and the general public but also we
should target machines
brains that are wired totally different
than we are
and that could show us some new
perspective on the data
in one poetic quote we call this the
rule zero
of open science if you're doing it only
for the scientists
you are doing it wrong so at last we are
ready to get into the work zones
number one is about vocab management and
lookup just to make sure we are all on
the same page vocab
does just short for vocabulary and that
word should make you think about
dictionaries
i should actually now explain that rdf
is the language of the semantic web and
introduce you to its grammar
but we don't have time for that and we
don't need it it is enough to realize
that any language
relies on having good dictionaries lists
that hold and explain all the terms in
use
so hold this dictionary image for a
moment okay the flemish people the deck
of vandal
and allow me to make just two
adjustments one is conceptual
unlike natural languages which can be
messy
these vocabularies make each term
totally
unambiguous and clear by themselves
that means that no extra context is
required
and also the other way around no added
context
could change the meaning okay so when
they are used in the statements
there should be no confusion about how
to read and interpret it
the second adjustment is a smart
technical trick
the terms themselves are spelled out
like web addresses
and yes that makes it very modern web
fashionable
but there is logic to this fancy madness
a lot can be said about it
name spacing and dns management and
whatever okay
but for today's story we keep it at
appreciating only one specific property
of this url
it is known as the follow your nose
property and you can just type these
terms into a web browser
and end up at the page that gives you
the explanation
so doing a dictionary lookup in this
system is
just as simple as using your web browser
finally to get totally clear we now have
the id
these are just three examples from the
marine
science domain the first example shows a
universally accepted term
identified by marine regions this one
particularly is numbered 2350 and is in
fact referring to the north sea
the second one is about my favorite
marine species
and people sometimes call it the
horseshoe crab but the correct
worldwide way to talk about this is to
use
http columns slash lsid.tidbit.org
url column ssid colon but it's pcs.org
colon
text name and then 150 511
beautiful creature and last time on the
slide you see that piece of marine
lab equipment a little bit looks like a
photocopier it's in fact used to detect
and measure
uh plankton species in water samples
and it's known as the zeus cam but the
real good thing is that the unambiguous
identifier for it
has been coined by our british
colleagues
so i hope this gives you a feeling how
exposing our own data sets using these
terms
is actually the way we ensure they can
be
understood and compared with the rest of
the world
know that you understand what vocabs are
if you didn't already
you probably also see feel i hope
you feel how important it is to know all
the words in all the dictionaries
and to use them correctly and this is
precisely what this first work zone is
about
making sure people get tools that help
them assign the correct vocab terms
in the context of their work the slide
lists only two services
you can find on the web to search for
terms that are available
and many more are around and they're
very useful but
these listings are very general and
they're not tuned to the subsets
shared between the members of your
research team
also they might not be tuned too much
your own language
nor include search terms or even local
slang
and thirdly there is no easy way to
integrate these services
with the applications for your team so
we put it all into one use case diagram
and this is what we'd like to see
we introduce here some vocab admin
a role that governs the vocabs your team
should be using
and at this level popular local search
terms translations or team slang could
be added to make them findable
okay next that's the green section
we consider how custom applications
could use widgets that do lookups
directly into these selected lists of
searchable terms
and by using this setup we aim at
reaching the end goal
having actual scientists apply the
correct terms to the data
they are sharing and remember rule zero
of open science we're not doing it
only for the scientists this comes down
to
uh if all of this actually becomes even
more useful
in the context of citizen science with
tools like this we can assume members of
the general public
are assigning the proper semantics to
the data they are producing
and making it more available for the
researchers
finally this slide introduces a
spontaneous set of building blocks to
get the job done
just one id talking in depth about this
is part of the invitation of today
but not here and now i have more for i
have four more zones to cover
the second one is going back to the
trinity of coexisting
data metadata and semantics it centers
around the search for
a unifying semantic research data
package
we already talked about the wasted
efficiency of our data scientists but
allow me to revisit the topic
when scientists go out to grab existing
data sets
they do so full of hope and dreams
they're dreaming about finding
information dreaming about being that
that whatever they found
being transparent and clear to the use
dreaming it will be complete
and not start a difficult search to grab
all related parts
on on cloud services dreaming it will be
interoperable with their own systems and
workflow
so they actually can get to work however
little surprise we already said this
if they actually do find data it is most
of the time packaged
into some form or formats that looks
like a treasure box
okay it's only willing to disclose its
value to the original owner
and the more striking part of that
observation is the other side
it is that the very same people are also
daily working towards sharing the
artifacts of their own research
but are packaging them in similar closed
boxes
and they do this despite the best of
their intentions to my feeling it's not
due to a failure on their part they just
like the tools to do a better job in
this area
so even further despite
the uptake of rdf semantics knowledge
graphs
i honestly do not believe that the
two-dimensional
table approach towards data is going
away soon
think about it if you consider scrolling
sorting plotting comparing calculating
even scripting you name it all data
manipulation assumes data
today being in rows and columns rather
than in mind
maps or graphs and the good news is that
we know how to mob the one to the other
but the tooling to make that as
accessible as the common spreadsheet is
simply not there
another challenge is that more and more
data sets are
in essence scattered around in a number
of places
consider big data systems you can't
contain all the data there
not in your set consider remote
repositories holding
images or dna sequences again too big to
be
contained in inside your package and
then you you you have the workflow
systems in the cloud
you keep on having a lot of
local stuff you have your mapping files
actual instrument uh
measurements what do you have and and
keeping track of
of that complete combination keeping
track of all the relations between all
of that
it always turns out to be uh solved by
some invention on the spot
okay a highly custom personal fit
mostly without any documentation
absolutely without any formal semantics
and the result is that we introduce
obscure layers
of wrapping that turn that well-intended
glass box into the dreaded treasure box
and i know but to know i've mostly been
talking about the lack of tools
and yes this slide as well it it does so
when when it suggests
there should be an uh an ecosystem of
supporting tools
but the elephant in the room is that we
don't have a common data package
standard
to capture and share the three parts of
the trinity
okay and that's not entirely true
there are a number of standards uh
emerging and i'm choosing today to share
my personal favorites
in this field it's called the research
object crates
in short the ro crates and i like it for
a number of reasons the first is
simplicity
uh we all know how to structure files
and folders
and then group them in an archive like
zip
or put them on a version in the system
like this so
let's just do that second it embraces
semantic web technology so yes it has
this uh
autocrate metadata json file which is a
json ld file
that uh next to the the tabular csv
files you put in the package
well it allows to through the json ld
you you can add semantic statements
about it for instance you use the csvw
csv for the web vocabulary to express
what
the the meaning of the fields in the csv
are third
it doesn't care if the described data is
in your data package itself
or in some external repository so it
deals with this modern reality of being
half in the cloud and half working on
local data
and fourth the backgrounds of the people
behind it
the people that are making the standard
their backgrounds is a mix of alpha and
beta
sciences so actually by design or by
coincidence i don't know but this group
is is forcing themselves to be explicit
about the meaning of what goes in
so do have a look at it if you take the
link on top
you will end up on their website and the
specification
and the links at the bottom are a fast
introduction
slide deck and a youtube movie
[Music]
delivering that i really can recommend
the last
for the fast entry um
and of course yeah we all know about uh
xkcd
slides or uh cartoon 927
i am i am a fan of xkcd
and every fan knows that randall monroe
is only
is always right and i think in this case
too uh when it comes down to standards
we all have our own pet peeves
and we are more than happy to translate
those to
into the motifs to stick to what we have
already been using
definitely the middle panel here one
universal attempt to replace them all
it's i think it's totally recognizable
most of the time
however the resulting pile of standards
resemble
the the image on the left which is the
ancient
indian model of the universe half of a
sphere
carried by four elements supported by a
big turtle and then below that
turtles all the way down layers on
layers that are never actually
achieving completeness nor consistency
but allow me to replace cynicism with
optimism there are
there are in fact useful ways to compare
standards and
and objectively describe their
properties the first link here
is one that points to the design
principles written down by tim
berners-lee
about the web i think roughly 20 years
ago
it's definitely the kind of web
archaeology you could find i dig it up
every three years or so and each time i
get extremely inspired by it it
also introduced me to the old chinese
wisdom on the slides
claiming that the usefulness of a teapot
comes from the fact that it is empty and
and you have to think about it right the
emptiness is the thing
the emptiness of the teapot is the thing
that allows it to hold your teeth
uh i i think this probably is the first
incarnation of the modern marketing
slogan less is more
to be honest that modern version never
really made sense to me but
this original one this one i feel okay
anyway back to standards
just like teapots they allow broader
usage
by limiting themselves by being concise
and focused by removing
all assumptions by allowing extensions
evolution uh embracing a collaboration
with other good solutions and standards
and this is the kind of thing that made
the web as fabulous
as it is and to my feeling it's the same
kind of meta properties
i can also uh recognize in the
design of the aerocrate standards anyway
with the semantic package standards like
the one of the
real crates we can go back to our work
zone
uh so i actually whatever glass box
standard definition we end up using
we will fail we will need to face the
remaining challenge
challenge okay we will need to revisit
all our treasure boxes from the past
and convert them into their transparent
counterparts
at least we have drafted up a process
that switches between
uh on the one side automated detection
of the fields in a package
and then allow a human an expert
assistant
to narrow down and actually decide
actually again
assign the correct vocab term to
describe those fields
which is indeed linking back to work
zone one uh for that we we sponsored
last year
an uh open summer of code project
oso 2020. it was called the shim dock
we can just we can argue about the name
but it made a first attempt and uh
prototype of precisely this
in all honesty it needs quite a bit of
extra
love and attention but the original
original design plans
and prototype are available and linked
from this
slide so if you have any thoughts while
audi's or wild ambition in this area we
are very
open to discussion and further tinkering
remember leave your notes in the shared
documents
and just to conclude think about having
a system that can convert
this uh treasure boxes in in glass boxes
[Music]
if we have such a system we can rethink
our repositories
and archives for science data as well
we can we can actually give it this
level extra level of smartness
if you look at the current repositories
they are content of adding a limited set
of metadata fields
and those then get attached as a meager
discoverable label to the treasure boxes
but with this new approach okay
helped through some semantics discovery
we could have the repository equipped to
do the conversion into the glass boxes
and by doing so we would make the
packages discoverable
and interoperable on a semantic level of
the data
and not only on the level of the
metadata so i think that's
really important actually the same line
of thinking can bring us even
a stage further if we adopt the ideas of
working zone 3
which i called the machine actionable
dmp
again a term i have to introduce because
myself i had never heard about dmps
before starting at bliss
dmp stands for data management plan and
the top level way to look at them
is you consider them being a contract
they list
the expectations and agreements all of
them concerning the handling of data
and they are shared between all
stakeholders of any specific research
project
at this stage many organizations or
funding agencies
in research are requiring you to have
these
and i think that testifies to the
growing importance of data and research
and you'll find every link that's only
one examples but
if you need to build one for yourself
you have services out there that
help you cover all the needs in all the
needed aspects
the platforms here are typically things
that guide you through a set of
templates and questions
so you don't forget anything definitely
these dmps are important and useful
they are also formal and more and more
they're following
these template structures but still
they're free form text
and so they are only understandable to
fellow humans
so again people have been thinking about
applying the the trick of semantics also
to these documents and the big question
is can we have these
structures so machines can read them too
and the hope is that through such
efforts it will lead
to automated processes robots if you
like
that help you out in applying the rules
of the dmp
and for us actually to fit it into this
talk which is
which definitely is about tools and
support
we read that ambition to automate more
as a way
of assisting and enabling and not
only as a way to automite automate
policing and validating which is
the the classical uh and the easy
approach so here is our thought
a platform we should think about having
a platform for
for building automated data management
assistance
okay these would essentially learn from
the dmp
what needs to be done inside the project
and then
produce a handy user interface to help
you with it
as such it could become the helpful
guide nudging you
into the right doing the right things
from the start
and it will definitely tie in with work
zone 2 because i think
in the end it will produce aero crate
packages
and those will hold the data trinity but
again it will probably
reuse voc lookup widgets from works on
one
in a scenario like this i think we
should be able to completely evoke
avoid the treasure box trap right not
getting into that anyway without
limiting your
own imagination about such a platform
this slide is only
offering some personal suggestions uh on
what it should try to achieve
it's an assistant so in my mind it
should be close by it should
have a desktop like user interface but
it should also embrace the best of the
web and it should be browser-based so i
think
some kind of a local host service uh
would be best after all it it also needs
to seamlessly integrate
with the cloud platforms that are listed
in dmp
um internally it might be totally
relying on knowledge graphs but it
definitely should have a natural
presentation of the data
that look like tables in spreadsheets
and finally yes it should support a set
of api
api hooks so it can be tied up in
scripting languages but also be used in
workflow engines
good i'm glad we are we already made it
here
and i hope you are too uh we have two
more to go
i realize i'm stretching your attention
and i'm not giving even giving you a
break
let's let's do that now okay let's take
15 seconds gymnastics
actually stand up i can use it too
stretch a bit
drink a glass of water bend your neck
loosen up
okay
the good news is that the remaining two
zones are quickies
i only have some uh early principles and
ids
and i hope you guys will take care of
the details
okay everybody ready this is the final
stretch
work zone four is about
linked data publishing now
the first three zones hopefully brought
some uplifting
positive vibe okay we've covered all
these nice things we can do with
semantics
i hope i did because now is the time to
tune that back just a little bit
i'm already uh sorry for that but the
sad route
is that only working with vocabs and
data standards
is really not enough to achieve
interoperability
we cannot overlook how these vocabs and
standards
get applied in web services and
protocols
on this slide you you see uh a number of
systems
mentioned that are used in our marine
research domain
and the observation is that despite the
common belief in using shared
vocabularies
even adopting them in some semantic
aware
data packages and we are all using this
wonderful web
still they end up not being
interoperable out of the box
we still face the fact that the inside
the wrong domain subsets of data
are closed up and thus hidden from other
subsets
of course they can be converted but that
always requires additional coding
and that implies cutting some corners
and the outreach to other domains is
limited in an even bigger way
because many of these standards only
make sense
inside our own community i speaking for
myself
only wfs was known to me all the other
ones were new and i'm i'm
active in in the web domain for the last
20 years
so there's really stuff that is only
used
inside this community personally i think
we have been focusing on the wrong
pre-position
and and i mean pre-position as a player
of words because i i want to target
the the pre-position and dutch that
forget
as well as the the similar sounding
proposition okay the guiding id
at four style i think all of us
are the developers we have all spent
tremendous
and well-intended efforts to link a vast
number of existing data systems
onto the web and in many cases we have
been adding
layer or layer or layer stacked turtles
to achieve this
often though we have neglected to really
embrace the web design principles
and let them influence our legacy packet
systems
those systems and the layers on top of
them do not
still do not conform to the beautiful
design properties we need
going back to the the less is more
chinese teapot this is my conviction we
should aim
removing layers aiming for less layers
actually stop publishing data on the web
but seek for ways to to to have our data
be
truly in the web right follow the the
nature
of how the web works one practical
example i see
is about blending away the arbitrary
differences between data sets and data
surfaces
people don't even question the clear
difference between both okay
nobody even wants to think about them
being the same thing
but we have to realize that any
distinction between boats
only lives on the side of the producer
of the data because
in either case the customer view is
identical
the consuming scientists just sees
a generic data provider it sees an
accessible url he doesn't or she doesn't
care
about uh that url holding holding
parameters or not
it is just there to produce some
response
and in both cases the response
should have this glass block
transparency
that's what we are expecting so the
transparency
transparency rules we talked about
should be applied to our services
as well as to our data sets they should
be equally semantically described
and i only mention here hydra cg as one
possible standard
now sketching the blueprint for this
complete work zone
is a challenge and frankly is is a lot
bigger than what we can chew in the time
we have left
i just have this list of elements that i
think that are part of the solution
uh and i have to admit when adding all
these links to the slides
it did kind of feel like this uh
qr id ambition was starting to look a
little bit
uh silly but the the real silly thing
here
is that all these pieces of work are
coming from the same single
research group based in ghent so i
really have to uh
urge you to go and check out their work
because all of it is highly
inspirational
and ready to be used most notably and
and
to some extent it's a lot like the work
zones i am suggesting today
all these elements all live kind of
close to the actual users
they assume some close by assistance
they take for granted
a more peer-to-peer distribution
distributed nature of the web they
actually have
uh working technical answers to the big
federated search
question and and all of that is is often
in big contrast to the very centralized
approach
of many of the the eos projects for
instance
all of them in some way introduce a
central hub
uh rather do that than provide open
source code that
allows anybody to set up their own hub
node right
they draw you to new big investment
cloud infrastructures
and often neglect providing the tools
for local and distributed
participation into that and this allows
me to make another observation
you see the fact of the matter is that a
computer science department like the one
i'm mentioning here
those are not directly involved in any
of the ongoing open science
projects none that i've seen funded by
the u by the eu
it's a maybe it's a funny story but last
year
i was at my first conference in the
domain of life sciences
so that's doctors and i end up listening
to smart medical doctors
contemplating computer architectures
that are
actually a challenge even for the
engineers
and to me it made me think about the
room filled with programmers
that are deciding yes we are going to
solve the cure to cancer
but we're not going to talk to any of
the medical professionals
professionals so maybe we should
actually even extend the rule zero
remember it says open science is not
only for the scientists
well i am convinced the open science
platform
should not be entirely built only by
the scientists either white right one
more to go
the fifth and last zone just to keep me
in time
is uh named usage tracking and metrics
for this
i have to come back to my honest now
really honest appreciation for the fair
acronym
after all it is the only torch lighting
apart
for all followers of the open science
procession
uh it is no coincidence in my mind that
this acronym landed
on the word fair okay somebody crafted
to be here
it's it's it's landing on the concept of
fairness on being fair
uh really you you just hustle the order
or
replace replace findable with with
something like discoverable which would
be the more
tech variant of the word and the result
doesn't end up
being a word actually the the word
would also not easily apply to what our
people would call the better angels of
our human nature
okay but this one does and that is truly
useful
we should not uh joke about it it is
useful because it is tapping
into one of the four main ways to drive
human behavior
and that number four comes from the book
shown on this slide
and i can really recommend it it covers
intellectual property rights
and ongoing cut and mouse play with
digital piracy
but here is the exercise let us apply
these four
uh influencers of behavior to the open
science topic
after all we hope to convert all
stakeholders to become
active contributors so the first one uh
law and enforcement points me to the
funding of research projects
there you have some leverage to shoehorn
the independence of research groups
into adopting new roles new rules right
a new rule like you must have a dmp from
now on or
you must follow the fair principles uh
the second influencer is architecture
and the previous four zones i think we
covered
uh we covered those four and they all
fall in this area
the idea is to develop tools and
techniques that make
uh what i would call just doing research
should be
naturally uh feel like the same thing as
doing open science
okay all the rules applied out of the
box uh
but by just using the correct tools
there is some work to do there but
i think it can be done third i already
mentioned we have the magic
of the fair words okay it sells the idea
that the open science way
is the morally right way but the last
element
is the one for this work zone and it's
where it's the important missing one
the the question is what are the
tangible effects and payback streams
the scientists scientists that are
adhering to this new set of rules
they see the effort and the cost they
need to do but they hardly see any gain
and noting that the currency in academia
is counted publications and citations
naturally there is an ongoing work to
extend that approach
people are searching for a similar count
to have some validization
and appreciation of data production and
data sharing
opendatametrics.org that looks like a
web domain but it's in
in fact a book if you follow the link
you end up at it
and it gives a very good understanding
of the current state of affairs
about data citation and usage usage
tracking
i i definitely recommend to read it it
is
very complete but still it left me
craving for more
the point is that the tracking we need
is a lot more complex
than the classic track and trace of
orders and
physical packages it's about digital
media media and we know
digital media in digital media you have
error loss copies
and they are cheap and thus abundance
we also see mashups that's all the rage
and science too
there is a lot of repackaging uh the
point is that datasets are not only
published and redistributed
via repositories they also get loaded
into aggregator services
and and fragment those refragments and
regroup
all that data and yes
well we all know assigning a doi to
every data set is an important
uh first step but the new questions are
are
spontaneously bubbling bubbling up to
what level must
must any possible fragments be
identifiable
identifiable on its own should data
services
attach full provenance trails to any
service
to any service response they provide
okay
and and should should
you could even go further and question
should uh
those repo responses when you think
about those prevalence trails when
they're made up of fragments should you
have a fragment for
should you have a provenance trail for
each of those fragments
so there's a lot to be told about and
and yes i think we need some shared
uh practice of publishing uh statistics
in a semantic way again
something that can then be openly
harvested by anyone
and i obviously see a bridge to uh the
dmps
assistant from work zone three you can
envision having an assistant
that would when you use it to obtain
download
and include data in your research
it could automatically add the usage
tracking statistics
inside the package okay and then publish
it together with the data
um so there's a lot of to think about
actually when i
and only this uh funny coincidence
it made me think when i when i was
looking for an image
to put on this slide i i just entered
the keyword tracing
and i got the second result here now you
have to admit
this is an interesting contemporary
interpretation of the track and trace
problem
when you think about viruses and
exposure and it's
just a vague idea but maybe that's the
kind of new
kind of uh tracing that could open our
minds
also to find a better solution in this
area
good that's it we made it thanks for
sticking around
i almost just not completely ended up in
time
so i hope you're still here from my side
i'm really eager to shut up now
and listen to you so uh if you can open
the mics
please do
most people join this call in in listen
only mode so if you want to talk you
have to reconnect to the audio and do
the echo test
mark there was already a question in the
um
in the chat about frictionless
yes yes that's the that's the good
typical candidate that gets mentioned
uh as in uh as a counterpart for
uh well in competition with autocrates
uh i hope they find a way to mitch and
match
um i have looked into frictionless data
they also have this this uh great vision
of uh of helping out to be
to be a tool that is close by to the
researcher
and that is helping out um
what i like less about it actually in
the history they had they were
at the beginning they were collaborating
with the csvw
group and then they split off so they
actually
kind of chosen the the non-semantic
routes
assuming that was way too complex for
people
maybe it's just a matter of timing other
creators maybe a little bit later i
don't know
they are now in a time where something
like jason ld exists
and jason ld is typically by i don't
know if you know this
famous quote but uh some rdf hulu saying
oh yes i'm really fed up with rdf i'm
not
going to use it anymore it's a it's a
pipe dream from now on i'm only using
json ld
uh the joke being that json ld is is
just
rdf as a serialization for rdf
but it really the joke shows that that
jason ld
is mostly approached as being jason
and not as being ling data so um
i don't know maybe if jason ld was
around when frictionless data started
it could have been uh more
[Music]
more aligned with with semantics all
over
so i don't know maybe there is some
combination some possible marriage
possible in the future i like
frictionless data
but i kind of like aerocrate more
anybody else more questions
i don't know questions in the chat at
the moment just a lot of thank you and
top presentation and that kind of
comments
you want people to um to write
their suggestions in the in in the
uh document right that you put the link
up top that's
yes yes yes and if you didn't have time
uh
if you didn't have spare bandwidth uh
during the talk
that document uh keeps being available
so
uh you can come back to me also people
have
found my email in the in the slides
you can contact me i'm definitely open
for more
uh discussion on this
good sorry for going over time
and if nobody else is speaking up i
think
i would like to call it a day and
stop the recording yes
well there is lupus i will try my life
yeah now i
found um
i wanted to ask you you are showing us a
great framework
that looks applicable and you are
displaying yourself
as a person with knowledge that could be
contacted how to apply it
um i wonder when there is likely who
funded projects and you listed that you
are like involved in many initiatives
and projects
is there any mechanism to ensure that
they apply this framework that they are
really like not
trying to reinventing the wheel but
really like that there is some incentive
mechanism in the whole system
all actors are in to apply what you just
presented because it makes so much sense
well thank you for that compliment um
i i am i'm up and around
in open stuff for the last 20 years i
used to have my own
open source company so that's how it all
started
working at the apache and stuff and
actually i don't know if we should
expect um
and in a number of ways i don't i don't
think we should expect
distributed approaches from a highly
centralized organization
right that for one um and i also
don't think we have to wait for top-down
decisions
um i see a lot of good and interesting
things happening
on the floor and and maybe we could just
assemble this as a as a bottom-up uh
counter answer to that thing the good
thing about the top down
is there's a great way of distributing
uh
the the the taxpayer money
into uh into the direction of of
reaching
the correct people and i have to say
last year i've
only met others either smart people or
extremely smart people
so i think the the the the
the distribution mechanism should be
keeping on
uh top down but i really think
it will come the real solutions will
come bottom up
and it will be from actual people
on the floor uh seeing what problems are
arising and finding smart
uh solutions to do that uh so
yeah i'm i'm i'm quite
i'm maybe i was a little bit critical
about
my observation on uh they're not
including computer science enough
etc etc but i am quite optimistic that
uh that the uh the intellectual
potential on the on the the grassroots
level is there and we will definitely
reach uh
solutions in yeah in a number of
years or days yep
good question lucas thank you
anybody else anybody struggling
to get the mic working you could also
put it on the chat
um there's also a question from bruno
mark and he asks in the documents
do you think that public administration
funding
scientific research and innovation could
or should enforce the use of
their results true are all crates and
what would be the smallest step
in that direction uh
well my previous question i i don't
really think about
don't really believe about the the
enforcement
so the the first way to uh to
to uh uh you know
modify human behavior in my list of four
law and enforcement i i don't think
that's that's the
the one with the the longest uh
it could be a smart way to kick-start it
but not uh the the easy way to get it
uh to get it in the long run to persist
and and keep it uh useful um
sure local investments uh could be added
up to the uh
to the to the more european ones uh or
the more centralized ones
i definitely agree there and i don't
know was there another
element of the question um
where was it in the
google docs and he also asks what can we
as public administration offer them to
foster the adoption
uh the fostering the adoption is just
just adopt just start doing it um
it would be nice that and i my previous
job
was uh was in uh in local government
so i definitely see how we could
collaborate
on a number of these uh smaller elements
that that could be then
rejoined i don't know maybe
having an auto great solution is over
the top for
uh local governments
but the elements like having a
vocabulary research
that's definitely useful for them as
well dmps
again that's that's very research
oriented
but uh having semantic frameworks
and you know maybe autocrate could be uh
could be a useful addition to uh i have
to think about it yeah
it could be in fact uh used
there as well because indeed we have
there exists the same problem here you
have data you have metadata
but people forget adding the semantics
and that's
[Music]
that's definitely something that a
package structure like
autocrate could help with yup i hope
that answers the question
okay
okay i see that leonard added uh
to my uh understanding that frictionless
is not
yet adopting uh linked data
but still has an open issue with it so
yeah there you go
good close to the one hour mark
shall i stop the recording yeah oh
yeah and we could still leave it open
the the
but i think since nobody is uh speaking
up
anyway i think we could just stop the
recording
and close the session thanks
my name is mark porter
you