Submind YouTube summaries
Thumbnail for Open Belgium 2021: Work Zones for Open Science

Open Belgium 2021: Work Zones for Open Science

Watch on YouTube

Video summary

Mark Porter from VLIOS presents a compelling vision for advancing open science beyond simple checklists like FAIR principles by introducing five distinct work zones designed to make global research data as easily searchable and accessible via natural language queries as Google. His primary goal is to eliminate the 79% of time currently wasted on avoidable overheads, transforming how scientists manage and share information. The first zone focuses on vocabulary management, advocating for unambiguous terms with unique identifiers that act like web addresses to ensure data is understandable globally while supporting citizen science through local slang and translations. This approach moves away from fragmented repositories toward unified data packaging using open standards such as Research Object Crates (RO-Crates), which combine data, metadata, and semantics into a single archive capable of handling both cloud and local storage environments. To further streamline the research process, Porter proposes transforming static Data Management Plans into dynamic platforms with desktop-like interfaces that automatically guide researchers in applying rules and generating compliant RO-Crates without relying on closed formats. This evolution is supported by linked data publishing principles where web services conform to semantic design standards like Hydra CG, creating a "glass box" transparency that allows consumers to see identical interfaces regardless of underlying technical differences. By promoting distributed solutions over centralized hubs, these zones aim to address the fragmentation caused by scattered workflows and ensure that semantics are not overlooked in favor of raw metadata alone. The presentation also highlights the importance of usage tracking and metrics, suggesting tangible incentives based on semantic data analysis rather than just moral arguments or traditional publication counts. Addressing concerns about incentivizing specific frameworks versus reinventing solutions, Porter argues that while top-down mechanisms work for distributing funds, real innovation stems from bottom-up grassroots efforts where practitioners solve actual problems. He advises against relying solely on enforcement laws to adopt public administration funding results and instead encourages starting with local investments in smaller elements like shared vocabularies and tools such as Autocrate. Although some initiatives are not yet fully adopting linked data standards, these incremental steps can effectively bridge gaps where semantics are often neglected. The session concludes by inviting audience collaboration through shared documents and chat participation to explore synergies between different standards while acknowledging that the recording ended due to a lack of further questions.
Read the full video transcript
my name is mark porter and i'd like to welcome to my home for this webinar on the work zones in open science thanks for showing up and a big thanks also to open belgium for bringing us together today the slides are available and free for reuse under creative comments but i do appreciate your attribution and your determination to share a like i have a limited amount of time and way too much content so let us quickly get some slides out of the way since one year and some days i work at vliss which is the flanders marine institute we are 20 years old and based in austin's on the belgian coast our mission is to enable and advance marine research in all possible domains and that ranges from biology geography ecology and even includes history medicine all the way up to psychology we do this to promote an advanced society in general and therefore we target various audiences that we see as important stakeholders in and around the benefits and responsibilities for our beloved seas narrowing to mountain narrowing down to my role i only have time to mention the department i work for which is the vmdc in english that acronym expands to the flanders marine data center our focus is on data data management and the design of data systems to support it we participate in many in a lot of projects and all of them have different levels of geographical coverage and outreach most notably i think we publish a number of data products for the marine research domain and those are ranging from something like the skeleta monitor which shows our local expertise and connection but it scales up via the european level for projects like ruby's which tracks marine biodiversity but even spans the globe with the reference databases we manage and govern four marine species and marine regions and concerning open science we are active on the flemish level inside the flemish research data network and the flemish open science board and in europe we participate in projects of the european open science cloud within the vmdc then together with three colleagues we make up the open science team we make sure the organization as a whole realizes and keeps furthering its open ambitions so what we are doing is implementing the fair principles advocating to and supporting research groups and adapting the data systems to be open by design in doing so we very much like and embrace open source and linked data technologies and if this sounds appealing to you keep an eye on the slash jobs page of the vlis website because we are working towards an extra job opening in the team and this is me and the various ways uh to contact me or find me ah yes this slide and some form or another this has been in every public speaking slide deck i've given since 2009 honestly honestly and honestly the web is totally awesome on so many levels and i think we should regularly have parties just to remain consciously aware of that it turns out this slide keeps being relevant so i keep it in and it surely will be today so we will come back to this fabulous web right uh no all of that is out of the way here is the actual agenda for today uh i see this as going to be uh six main blocks the first one to lay down some introduction and overview of what open science is and then five actual identified work zones that could turn out to be too much but let us be ambitious finally this is 2021 and we do this online so i don't hear nor see you and the point is we should uh we would we would very much like your feedback even more than only your feedback we actually want to reach out and recruit some of you to collaborate on some of these tasks and to achieve that goal uh i would like to ask you to use this co-writing space we have foreseen and which should have been shared in the chat also uh and also at the end you might prepare yourself to open up your own mic for our audience participation part also some of my colleagues are around so they might interact on the chat while i am talking um an experiment i did an effort to include qr codes for all the links i mentioned so if you have a device with a keyword scanner around you can follow up on those as well and our last tip my voice is mostly not going to read out what is on the slides okay i expect you to be able to do that for yourself and allow my comments to be loosely providing additional insights around that structure so on with the show when i talk about open science i almost always end up doing some fair bashing if you haven't heard about it the fair acronym is the first thing you learn when coming to the open science field it's a bit like a rite of passage or a painful tattoo something like a required code to become part of this clip and to sign into its belief system the letters spell the expectations for all data and publications all of them should be made findable accessible interoperable and reusable don't get me wrong it is a very smart acronym and a very suitable top-level shopping list of the things to cover however it is also becoming its own purpose most importantly in my personal opinion it is too often translated into a simple checklist and one that allows people to declare we have done enough to reach the end of their responsibility they just need to check all four boxes as quickly as possible so instead of taking that route i want to invite you to a more open and unbound thinking about open science this is what we take as the guiding user story for the work we do to provide a relevant slice of global research of the global research data set that is as simple as a google search we actually package this in in our gyro as bug number one for the complete team and we believe it's probably going to take us uh uh until we retire to to get it done we see this as even having two parts the first part is about the natural way of finding the data if you go to google today you will be able to have it answer quite convoluted questions like an example who is the actress playing the wife of stephen hawking or who is the director of that movie this kind of natural way to find is precisely what we would like to achieve for the data lookups in the realm of scientific work and on the left side of the slide you see a number of examples in the marine domain of how such searches could be uh expressed in contrast on the right side of the slide we have listed all the steps that are involved today to actually answer those questions and if you look carefully behind semi-transparent behind the list in the background you see a big pie chart and that brings this to my next slide it turns out that no less that 79 of the working day of a data scientist is wasted on those steps all of them are avoidable and overhead issues and if this 79 could be reduced to 18 it would shrink their working day to only one-third if we turn that around we could make every full-time hire of one data scientist count as three and we believe it is precisely this kind of upscaling that open science should aim for society is funding the research so logically it keeps the right to expect a higher return on investment open science often serves on a way of high ethical motives but we should not be shy to claim economic goals as well we should make and deliver on a promise to reduce the cost of doing research the next open science doesn't stop at targeting the science community alone with the second part of our ambitious bug we want to stress the need to target the general public the research community needs to push its quality information out on that same web everybody has learned to use it has to be available there and it has to be made high ranking at the bottom you'll find my catchy twitter code on the quote on the subject if you agree scan the cover code and give it some of your twitter love and remember this event is using the hashtag openbelgium21 so now we have a challenging job and we see two parts in it and in our head both parts lead us to the web first as an example of easy and natural search but also as the platform to reach any everyone out there so we believe the backbone of this fantastic web is the foundation for open science especially uh the principles of the semantic web which brings us to our this we call the ikea observation slide it should explain to you the subtle difference between uh semantics and metadata and it also shows the connected unity of this holy trinity it is always data metadata and never forget the third one semantics we need them all and they should be kept together okay i mentioned the smart google search earlier but you might remember this scenario it's less than a year ago you have booked a flight and suddenly your email client receives a confirmation email and it is is successful in assisting you to add that event to your calendar and you have to understand it is not an effect of some spooky artificial intelligence or any big brother stuff i know it looks like they know about your plans but in reality this effect is driven through clear and straightforward semantic annotations inside the message inside the message there are just some schema.org terms hidden inside the html the html of your email and those allow your mail clients to understand part of the information that you are reading and magical effects like these are easily achievable when we add semantics to the data and that is what the phrase making data machine actionable really is about and then going back to research i think maybe even worse than in other domains the attention for semantics has been absent and allow me to put that in context the required understanding of meaning and intentions has not been neglected rather it has simply been kept implicit they assume it to be known among the smart peers but in reality in most cases you have on the one side a knowledgeable producer of the data and it talks only to the unknowing random external consumer via some intermediate data system and it is therefore essential that all insights and understanding all semantics are captured and made explicit to conclude all of this shows the access part of open science is only the tip of the iceberg you might remember that richard stallman from the free software foundation at some point in time clarified that free and free software should be read as freedom or free speech and not as free beer in the same way i think the open in open science really refers to open-ended and not too open cans open is about taking up the extra work the data is not released only to repeat your own findings and experiments instead it needs it needs some extra work it must be prepared so it can be rehashed reassembled recombined in future new contexts their original meaning should be kept but their reuse should allow solving new problems and this also comes back to targeting audiences outside the inner circle of the new research community the shared data should function as an open invitation to three new classes of players one is the scientists from other domains then you have the citizen scientists and the general public but also we should target machines brains that are wired totally different than we are and that could show us some new perspective on the data in one poetic quote we call this the rule zero of open science if you're doing it only for the scientists you are doing it wrong so at last we are ready to get into the work zones number one is about vocab management and lookup just to make sure we are all on the same page vocab does just short for vocabulary and that word should make you think about dictionaries i should actually now explain that rdf is the language of the semantic web and introduce you to its grammar but we don't have time for that and we don't need it it is enough to realize that any language relies on having good dictionaries lists that hold and explain all the terms in use so hold this dictionary image for a moment okay the flemish people the deck of vandal and allow me to make just two adjustments one is conceptual unlike natural languages which can be messy these vocabularies make each term totally unambiguous and clear by themselves that means that no extra context is required and also the other way around no added context could change the meaning okay so when they are used in the statements there should be no confusion about how to read and interpret it the second adjustment is a smart technical trick the terms themselves are spelled out like web addresses and yes that makes it very modern web fashionable but there is logic to this fancy madness a lot can be said about it name spacing and dns management and whatever okay but for today's story we keep it at appreciating only one specific property of this url it is known as the follow your nose property and you can just type these terms into a web browser and end up at the page that gives you the explanation so doing a dictionary lookup in this system is just as simple as using your web browser finally to get totally clear we now have the id these are just three examples from the marine science domain the first example shows a universally accepted term identified by marine regions this one particularly is numbered 2350 and is in fact referring to the north sea the second one is about my favorite marine species and people sometimes call it the horseshoe crab but the correct worldwide way to talk about this is to use http columns slash lsid.tidbit.org url column ssid colon but it's pcs.org colon text name and then 150 511 beautiful creature and last time on the slide you see that piece of marine lab equipment a little bit looks like a photocopier it's in fact used to detect and measure uh plankton species in water samples and it's known as the zeus cam but the real good thing is that the unambiguous identifier for it has been coined by our british colleagues so i hope this gives you a feeling how exposing our own data sets using these terms is actually the way we ensure they can be understood and compared with the rest of the world know that you understand what vocabs are if you didn't already you probably also see feel i hope you feel how important it is to know all the words in all the dictionaries and to use them correctly and this is precisely what this first work zone is about making sure people get tools that help them assign the correct vocab terms in the context of their work the slide lists only two services you can find on the web to search for terms that are available and many more are around and they're very useful but these listings are very general and they're not tuned to the subsets shared between the members of your research team also they might not be tuned too much your own language nor include search terms or even local slang and thirdly there is no easy way to integrate these services with the applications for your team so we put it all into one use case diagram and this is what we'd like to see we introduce here some vocab admin a role that governs the vocabs your team should be using and at this level popular local search terms translations or team slang could be added to make them findable okay next that's the green section we consider how custom applications could use widgets that do lookups directly into these selected lists of searchable terms and by using this setup we aim at reaching the end goal having actual scientists apply the correct terms to the data they are sharing and remember rule zero of open science we're not doing it only for the scientists this comes down to uh if all of this actually becomes even more useful in the context of citizen science with tools like this we can assume members of the general public are assigning the proper semantics to the data they are producing and making it more available for the researchers finally this slide introduces a spontaneous set of building blocks to get the job done just one id talking in depth about this is part of the invitation of today but not here and now i have more for i have four more zones to cover the second one is going back to the trinity of coexisting data metadata and semantics it centers around the search for a unifying semantic research data package we already talked about the wasted efficiency of our data scientists but allow me to revisit the topic when scientists go out to grab existing data sets they do so full of hope and dreams they're dreaming about finding information dreaming about being that that whatever they found being transparent and clear to the use dreaming it will be complete and not start a difficult search to grab all related parts on on cloud services dreaming it will be interoperable with their own systems and workflow so they actually can get to work however little surprise we already said this if they actually do find data it is most of the time packaged into some form or formats that looks like a treasure box okay it's only willing to disclose its value to the original owner and the more striking part of that observation is the other side it is that the very same people are also daily working towards sharing the artifacts of their own research but are packaging them in similar closed boxes and they do this despite the best of their intentions to my feeling it's not due to a failure on their part they just like the tools to do a better job in this area so even further despite the uptake of rdf semantics knowledge graphs i honestly do not believe that the two-dimensional table approach towards data is going away soon think about it if you consider scrolling sorting plotting comparing calculating even scripting you name it all data manipulation assumes data today being in rows and columns rather than in mind maps or graphs and the good news is that we know how to mob the one to the other but the tooling to make that as accessible as the common spreadsheet is simply not there another challenge is that more and more data sets are in essence scattered around in a number of places consider big data systems you can't contain all the data there not in your set consider remote repositories holding images or dna sequences again too big to be contained in inside your package and then you you you have the workflow systems in the cloud you keep on having a lot of local stuff you have your mapping files actual instrument uh measurements what do you have and and keeping track of of that complete combination keeping track of all the relations between all of that it always turns out to be uh solved by some invention on the spot okay a highly custom personal fit mostly without any documentation absolutely without any formal semantics and the result is that we introduce obscure layers of wrapping that turn that well-intended glass box into the dreaded treasure box and i know but to know i've mostly been talking about the lack of tools and yes this slide as well it it does so when when it suggests there should be an uh an ecosystem of supporting tools but the elephant in the room is that we don't have a common data package standard to capture and share the three parts of the trinity okay and that's not entirely true there are a number of standards uh emerging and i'm choosing today to share my personal favorites in this field it's called the research object crates in short the ro crates and i like it for a number of reasons the first is simplicity uh we all know how to structure files and folders and then group them in an archive like zip or put them on a version in the system like this so let's just do that second it embraces semantic web technology so yes it has this uh autocrate metadata json file which is a json ld file that uh next to the the tabular csv files you put in the package well it allows to through the json ld you you can add semantic statements about it for instance you use the csvw csv for the web vocabulary to express what the the meaning of the fields in the csv are third it doesn't care if the described data is in your data package itself or in some external repository so it deals with this modern reality of being half in the cloud and half working on local data and fourth the backgrounds of the people behind it the people that are making the standard their backgrounds is a mix of alpha and beta sciences so actually by design or by coincidence i don't know but this group is is forcing themselves to be explicit about the meaning of what goes in so do have a look at it if you take the link on top you will end up on their website and the specification and the links at the bottom are a fast introduction slide deck and a youtube movie [Music] delivering that i really can recommend the last for the fast entry um and of course yeah we all know about uh xkcd slides or uh cartoon 927 i am i am a fan of xkcd and every fan knows that randall monroe is only is always right and i think in this case too uh when it comes down to standards we all have our own pet peeves and we are more than happy to translate those to into the motifs to stick to what we have already been using definitely the middle panel here one universal attempt to replace them all it's i think it's totally recognizable most of the time however the resulting pile of standards resemble the the image on the left which is the ancient indian model of the universe half of a sphere carried by four elements supported by a big turtle and then below that turtles all the way down layers on layers that are never actually achieving completeness nor consistency but allow me to replace cynicism with optimism there are there are in fact useful ways to compare standards and and objectively describe their properties the first link here is one that points to the design principles written down by tim berners-lee about the web i think roughly 20 years ago it's definitely the kind of web archaeology you could find i dig it up every three years or so and each time i get extremely inspired by it it also introduced me to the old chinese wisdom on the slides claiming that the usefulness of a teapot comes from the fact that it is empty and and you have to think about it right the emptiness is the thing the emptiness of the teapot is the thing that allows it to hold your teeth uh i i think this probably is the first incarnation of the modern marketing slogan less is more to be honest that modern version never really made sense to me but this original one this one i feel okay anyway back to standards just like teapots they allow broader usage by limiting themselves by being concise and focused by removing all assumptions by allowing extensions evolution uh embracing a collaboration with other good solutions and standards and this is the kind of thing that made the web as fabulous as it is and to my feeling it's the same kind of meta properties i can also uh recognize in the design of the aerocrate standards anyway with the semantic package standards like the one of the real crates we can go back to our work zone uh so i actually whatever glass box standard definition we end up using we will fail we will need to face the remaining challenge challenge okay we will need to revisit all our treasure boxes from the past and convert them into their transparent counterparts at least we have drafted up a process that switches between uh on the one side automated detection of the fields in a package and then allow a human an expert assistant to narrow down and actually decide actually again assign the correct vocab term to describe those fields which is indeed linking back to work zone one uh for that we we sponsored last year an uh open summer of code project oso 2020. it was called the shim dock we can just we can argue about the name but it made a first attempt and uh prototype of precisely this in all honesty it needs quite a bit of extra love and attention but the original original design plans and prototype are available and linked from this slide so if you have any thoughts while audi's or wild ambition in this area we are very open to discussion and further tinkering remember leave your notes in the shared documents and just to conclude think about having a system that can convert this uh treasure boxes in in glass boxes [Music] if we have such a system we can rethink our repositories and archives for science data as well we can we can actually give it this level extra level of smartness if you look at the current repositories they are content of adding a limited set of metadata fields and those then get attached as a meager discoverable label to the treasure boxes but with this new approach okay helped through some semantics discovery we could have the repository equipped to do the conversion into the glass boxes and by doing so we would make the packages discoverable and interoperable on a semantic level of the data and not only on the level of the metadata so i think that's really important actually the same line of thinking can bring us even a stage further if we adopt the ideas of working zone 3 which i called the machine actionable dmp again a term i have to introduce because myself i had never heard about dmps before starting at bliss dmp stands for data management plan and the top level way to look at them is you consider them being a contract they list the expectations and agreements all of them concerning the handling of data and they are shared between all stakeholders of any specific research project at this stage many organizations or funding agencies in research are requiring you to have these and i think that testifies to the growing importance of data and research and you'll find every link that's only one examples but if you need to build one for yourself you have services out there that help you cover all the needs in all the needed aspects the platforms here are typically things that guide you through a set of templates and questions so you don't forget anything definitely these dmps are important and useful they are also formal and more and more they're following these template structures but still they're free form text and so they are only understandable to fellow humans so again people have been thinking about applying the the trick of semantics also to these documents and the big question is can we have these structures so machines can read them too and the hope is that through such efforts it will lead to automated processes robots if you like that help you out in applying the rules of the dmp and for us actually to fit it into this talk which is which definitely is about tools and support we read that ambition to automate more as a way of assisting and enabling and not only as a way to automite automate policing and validating which is the the classical uh and the easy approach so here is our thought a platform we should think about having a platform for for building automated data management assistance okay these would essentially learn from the dmp what needs to be done inside the project and then produce a handy user interface to help you with it as such it could become the helpful guide nudging you into the right doing the right things from the start and it will definitely tie in with work zone 2 because i think in the end it will produce aero crate packages and those will hold the data trinity but again it will probably reuse voc lookup widgets from works on one in a scenario like this i think we should be able to completely evoke avoid the treasure box trap right not getting into that anyway without limiting your own imagination about such a platform this slide is only offering some personal suggestions uh on what it should try to achieve it's an assistant so in my mind it should be close by it should have a desktop like user interface but it should also embrace the best of the web and it should be browser-based so i think some kind of a local host service uh would be best after all it it also needs to seamlessly integrate with the cloud platforms that are listed in dmp um internally it might be totally relying on knowledge graphs but it definitely should have a natural presentation of the data that look like tables in spreadsheets and finally yes it should support a set of api api hooks so it can be tied up in scripting languages but also be used in workflow engines good i'm glad we are we already made it here and i hope you are too uh we have two more to go i realize i'm stretching your attention and i'm not giving even giving you a break let's let's do that now okay let's take 15 seconds gymnastics actually stand up i can use it too stretch a bit drink a glass of water bend your neck loosen up okay the good news is that the remaining two zones are quickies i only have some uh early principles and ids and i hope you guys will take care of the details okay everybody ready this is the final stretch work zone four is about linked data publishing now the first three zones hopefully brought some uplifting positive vibe okay we've covered all these nice things we can do with semantics i hope i did because now is the time to tune that back just a little bit i'm already uh sorry for that but the sad route is that only working with vocabs and data standards is really not enough to achieve interoperability we cannot overlook how these vocabs and standards get applied in web services and protocols on this slide you you see uh a number of systems mentioned that are used in our marine research domain and the observation is that despite the common belief in using shared vocabularies even adopting them in some semantic aware data packages and we are all using this wonderful web still they end up not being interoperable out of the box we still face the fact that the inside the wrong domain subsets of data are closed up and thus hidden from other subsets of course they can be converted but that always requires additional coding and that implies cutting some corners and the outreach to other domains is limited in an even bigger way because many of these standards only make sense inside our own community i speaking for myself only wfs was known to me all the other ones were new and i'm i'm active in in the web domain for the last 20 years so there's really stuff that is only used inside this community personally i think we have been focusing on the wrong pre-position and and i mean pre-position as a player of words because i i want to target the the pre-position and dutch that forget as well as the the similar sounding proposition okay the guiding id at four style i think all of us are the developers we have all spent tremendous and well-intended efforts to link a vast number of existing data systems onto the web and in many cases we have been adding layer or layer or layer stacked turtles to achieve this often though we have neglected to really embrace the web design principles and let them influence our legacy packet systems those systems and the layers on top of them do not still do not conform to the beautiful design properties we need going back to the the less is more chinese teapot this is my conviction we should aim removing layers aiming for less layers actually stop publishing data on the web but seek for ways to to to have our data be truly in the web right follow the the nature of how the web works one practical example i see is about blending away the arbitrary differences between data sets and data surfaces people don't even question the clear difference between both okay nobody even wants to think about them being the same thing but we have to realize that any distinction between boats only lives on the side of the producer of the data because in either case the customer view is identical the consuming scientists just sees a generic data provider it sees an accessible url he doesn't or she doesn't care about uh that url holding holding parameters or not it is just there to produce some response and in both cases the response should have this glass block transparency that's what we are expecting so the transparency transparency rules we talked about should be applied to our services as well as to our data sets they should be equally semantically described and i only mention here hydra cg as one possible standard now sketching the blueprint for this complete work zone is a challenge and frankly is is a lot bigger than what we can chew in the time we have left i just have this list of elements that i think that are part of the solution uh and i have to admit when adding all these links to the slides it did kind of feel like this uh qr id ambition was starting to look a little bit uh silly but the the real silly thing here is that all these pieces of work are coming from the same single research group based in ghent so i really have to uh urge you to go and check out their work because all of it is highly inspirational and ready to be used most notably and and to some extent it's a lot like the work zones i am suggesting today all these elements all live kind of close to the actual users they assume some close by assistance they take for granted a more peer-to-peer distribution distributed nature of the web they actually have uh working technical answers to the big federated search question and and all of that is is often in big contrast to the very centralized approach of many of the the eos projects for instance all of them in some way introduce a central hub uh rather do that than provide open source code that allows anybody to set up their own hub node right they draw you to new big investment cloud infrastructures and often neglect providing the tools for local and distributed participation into that and this allows me to make another observation you see the fact of the matter is that a computer science department like the one i'm mentioning here those are not directly involved in any of the ongoing open science projects none that i've seen funded by the u by the eu it's a maybe it's a funny story but last year i was at my first conference in the domain of life sciences so that's doctors and i end up listening to smart medical doctors contemplating computer architectures that are actually a challenge even for the engineers and to me it made me think about the room filled with programmers that are deciding yes we are going to solve the cure to cancer but we're not going to talk to any of the medical professionals professionals so maybe we should actually even extend the rule zero remember it says open science is not only for the scientists well i am convinced the open science platform should not be entirely built only by the scientists either white right one more to go the fifth and last zone just to keep me in time is uh named usage tracking and metrics for this i have to come back to my honest now really honest appreciation for the fair acronym after all it is the only torch lighting apart for all followers of the open science procession uh it is no coincidence in my mind that this acronym landed on the word fair okay somebody crafted to be here it's it's it's landing on the concept of fairness on being fair uh really you you just hustle the order or replace replace findable with with something like discoverable which would be the more tech variant of the word and the result doesn't end up being a word actually the the word would also not easily apply to what our people would call the better angels of our human nature okay but this one does and that is truly useful we should not uh joke about it it is useful because it is tapping into one of the four main ways to drive human behavior and that number four comes from the book shown on this slide and i can really recommend it it covers intellectual property rights and ongoing cut and mouse play with digital piracy but here is the exercise let us apply these four uh influencers of behavior to the open science topic after all we hope to convert all stakeholders to become active contributors so the first one uh law and enforcement points me to the funding of research projects there you have some leverage to shoehorn the independence of research groups into adopting new roles new rules right a new rule like you must have a dmp from now on or you must follow the fair principles uh the second influencer is architecture and the previous four zones i think we covered uh we covered those four and they all fall in this area the idea is to develop tools and techniques that make uh what i would call just doing research should be naturally uh feel like the same thing as doing open science okay all the rules applied out of the box uh but by just using the correct tools there is some work to do there but i think it can be done third i already mentioned we have the magic of the fair words okay it sells the idea that the open science way is the morally right way but the last element is the one for this work zone and it's where it's the important missing one the the question is what are the tangible effects and payback streams the scientists scientists that are adhering to this new set of rules they see the effort and the cost they need to do but they hardly see any gain and noting that the currency in academia is counted publications and citations naturally there is an ongoing work to extend that approach people are searching for a similar count to have some validization and appreciation of data production and data sharing opendatametrics.org that looks like a web domain but it's in in fact a book if you follow the link you end up at it and it gives a very good understanding of the current state of affairs about data citation and usage usage tracking i i definitely recommend to read it it is very complete but still it left me craving for more the point is that the tracking we need is a lot more complex than the classic track and trace of orders and physical packages it's about digital media media and we know digital media in digital media you have error loss copies and they are cheap and thus abundance we also see mashups that's all the rage and science too there is a lot of repackaging uh the point is that datasets are not only published and redistributed via repositories they also get loaded into aggregator services and and fragment those refragments and regroup all that data and yes well we all know assigning a doi to every data set is an important uh first step but the new questions are are spontaneously bubbling bubbling up to what level must must any possible fragments be identifiable identifiable on its own should data services attach full provenance trails to any service to any service response they provide okay and and should should you could even go further and question should uh those repo responses when you think about those prevalence trails when they're made up of fragments should you have a fragment for should you have a provenance trail for each of those fragments so there's a lot to be told about and and yes i think we need some shared uh practice of publishing uh statistics in a semantic way again something that can then be openly harvested by anyone and i obviously see a bridge to uh the dmps assistant from work zone three you can envision having an assistant that would when you use it to obtain download and include data in your research it could automatically add the usage tracking statistics inside the package okay and then publish it together with the data um so there's a lot of to think about actually when i and only this uh funny coincidence it made me think when i when i was looking for an image to put on this slide i i just entered the keyword tracing and i got the second result here now you have to admit this is an interesting contemporary interpretation of the track and trace problem when you think about viruses and exposure and it's just a vague idea but maybe that's the kind of new kind of uh tracing that could open our minds also to find a better solution in this area good that's it we made it thanks for sticking around i almost just not completely ended up in time so i hope you're still here from my side i'm really eager to shut up now and listen to you so uh if you can open the mics please do most people join this call in in listen only mode so if you want to talk you have to reconnect to the audio and do the echo test mark there was already a question in the um in the chat about frictionless yes yes that's the that's the good typical candidate that gets mentioned uh as in uh as a counterpart for uh well in competition with autocrates uh i hope they find a way to mitch and match um i have looked into frictionless data they also have this this uh great vision of uh of helping out to be to be a tool that is close by to the researcher and that is helping out um what i like less about it actually in the history they had they were at the beginning they were collaborating with the csvw group and then they split off so they actually kind of chosen the the non-semantic routes assuming that was way too complex for people maybe it's just a matter of timing other creators maybe a little bit later i don't know they are now in a time where something like jason ld exists and jason ld is typically by i don't know if you know this famous quote but uh some rdf hulu saying oh yes i'm really fed up with rdf i'm not going to use it anymore it's a it's a pipe dream from now on i'm only using json ld uh the joke being that json ld is is just rdf as a serialization for rdf but it really the joke shows that that jason ld is mostly approached as being jason and not as being ling data so um i don't know maybe if jason ld was around when frictionless data started it could have been uh more [Music] more aligned with with semantics all over so i don't know maybe there is some combination some possible marriage possible in the future i like frictionless data but i kind of like aerocrate more anybody else more questions i don't know questions in the chat at the moment just a lot of thank you and top presentation and that kind of comments you want people to um to write their suggestions in the in in the uh document right that you put the link up top that's yes yes yes and if you didn't have time uh if you didn't have spare bandwidth uh during the talk that document uh keeps being available so uh you can come back to me also people have found my email in the in the slides you can contact me i'm definitely open for more uh discussion on this good sorry for going over time and if nobody else is speaking up i think i would like to call it a day and stop the recording yes well there is lupus i will try my life yeah now i found um i wanted to ask you you are showing us a great framework that looks applicable and you are displaying yourself as a person with knowledge that could be contacted how to apply it um i wonder when there is likely who funded projects and you listed that you are like involved in many initiatives and projects is there any mechanism to ensure that they apply this framework that they are really like not trying to reinventing the wheel but really like that there is some incentive mechanism in the whole system all actors are in to apply what you just presented because it makes so much sense well thank you for that compliment um i i am i'm up and around in open stuff for the last 20 years i used to have my own open source company so that's how it all started working at the apache and stuff and actually i don't know if we should expect um and in a number of ways i don't i don't think we should expect distributed approaches from a highly centralized organization right that for one um and i also don't think we have to wait for top-down decisions um i see a lot of good and interesting things happening on the floor and and maybe we could just assemble this as a as a bottom-up uh counter answer to that thing the good thing about the top down is there's a great way of distributing uh the the the taxpayer money into uh into the direction of of reaching the correct people and i have to say last year i've only met others either smart people or extremely smart people so i think the the the the the distribution mechanism should be keeping on uh top down but i really think it will come the real solutions will come bottom up and it will be from actual people on the floor uh seeing what problems are arising and finding smart uh solutions to do that uh so yeah i'm i'm i'm quite i'm maybe i was a little bit critical about my observation on uh they're not including computer science enough etc etc but i am quite optimistic that uh that the uh the intellectual potential on the on the the grassroots level is there and we will definitely reach uh solutions in yeah in a number of years or days yep good question lucas thank you anybody else anybody struggling to get the mic working you could also put it on the chat um there's also a question from bruno mark and he asks in the documents do you think that public administration funding scientific research and innovation could or should enforce the use of their results true are all crates and what would be the smallest step in that direction uh well my previous question i i don't really think about don't really believe about the the enforcement so the the first way to uh to to uh uh you know modify human behavior in my list of four law and enforcement i i don't think that's that's the the one with the the longest uh it could be a smart way to kick-start it but not uh the the easy way to get it uh to get it in the long run to persist and and keep it uh useful um sure local investments uh could be added up to the uh to the to the more european ones uh or the more centralized ones i definitely agree there and i don't know was there another element of the question um where was it in the google docs and he also asks what can we as public administration offer them to foster the adoption uh the fostering the adoption is just just adopt just start doing it um it would be nice that and i my previous job was uh was in uh in local government so i definitely see how we could collaborate on a number of these uh smaller elements that that could be then rejoined i don't know maybe having an auto great solution is over the top for uh local governments but the elements like having a vocabulary research that's definitely useful for them as well dmps again that's that's very research oriented but uh having semantic frameworks and you know maybe autocrate could be uh could be a useful addition to uh i have to think about it yeah it could be in fact uh used there as well because indeed we have there exists the same problem here you have data you have metadata but people forget adding the semantics and that's [Music] that's definitely something that a package structure like autocrate could help with yup i hope that answers the question okay okay i see that leonard added uh to my uh understanding that frictionless is not yet adopting uh linked data but still has an open issue with it so yeah there you go good close to the one hour mark shall i stop the recording yeah oh yeah and we could still leave it open the the but i think since nobody is uh speaking up anyway i think we could just stop the recording and close the session thanks my name is mark porter you