Submind YouTube summaries
Thumbnail for Surviving Crisis, Sustaining Science: Decentralized Persistent Identifiers for an Uncertain Future

Surviving Crisis, Sustaining Science: Decentralized Persistent Identifiers for an Uncertain Future

Watch on YouTube

Video summary

The presentation addresses the critical need to sustain scientific progress amidst a landscape of increasing crises, including institutional failures, funding cuts, and geopolitical censorship. The speaker argues that current reliance on centralized systems creates fragile points of failure where losing a website or server results in total data loss, citing examples such as restricted access to US health governance data due to legal pressures and the banning of InterPlanetary File System (IPFS) technologies by the European Union. These scenarios illustrate how existing infrastructure can be weaponized to break links to knowledge, effectively burning books before they even catch fire. To counter this, the core argument is that science must become resilient through decentralized persistent identifiers that ensure findability, accessibility, interoperability, and reusability regardless of external pressures or single points of failure. The talk critically examines the limitations of current solutions like Digital Object Identifiers (DOIs), which were originally commercial projects that have evolved into platforms exerting sovereignty over metadata rather than facilitating open governance. The speaker highlights how these systems often lock data behind paywalls and suffer from "rotten links" where metadata is preserved but content becomes inaccessible, a vulnerability exacerbated by sanctions or arbitrary decisions made by single nations. In contrast, the presentation advocates for decentralized identification models based on mathematical environments like BitTorrent networks, which offer redundancy and resilience without necessarily requiring blockchain technology. The goal is to shift away from commercial foundations that prioritize control toward collaborative ecosystems where stakeholders share responsibility rather than relying on a central authority to hold keys or manage access. However, the path forward involves navigating complex questions regarding trust, sovereignty, and social conflict within these decentralized frameworks. While initiatives like those in North America focus on hash-based identifiers that are difficult to federate globally, projects emerging from the Global South aim for truly distributed solutions. The speaker emphasizes that decentralization is not merely a technical fix but a social good rooted in openness and responsibility, challenging the notion of blockchain as a silver bullet since legal mechanisms can still force hard forks or censor specific jurisdictions regardless of network distribution. Ultimately, the future of big data science depends on developing infrastructure owned by society rather than corporations or states, ensuring that knowledge remains accessible even when facing adversarial attacks or political censorship. In conclusion, building a resilient scientific identity system requires moving beyond traditional models to embrace decentralized, redundant architectures designed explicitly for fairness and reproducibility. The presentation underscores that while challenges such as handling social conflicts within these systems remain unanswered, the necessity of persisting through crisis is clear. By fostering collaboration between different regions and avoiding over-reliance on any single technology or legal framework, the scientific community can protect its data against both institutional negligence and malicious intent. This approach ensures that science survives not just by hiding behind walls but by standing openly in a network where no single entity holds the power to silence knowledge or erase history.
Read the full video transcript
Thank you very much and hi everyone. So I am Andre presenting substituting my dear friend Sergio with this surviving crisis and sustain it science and decentralized persistent identifiers for an uncertain future. And let me please begin my presentation from the two citations uh from the first one is from Maio Ferraris. You can just read it on the slide. It's about the documentia revolution we have right now. everywhere and for everyone as the all the old productivity things we have over the data are not working and where we had the one to many model of interaction now we should always have many to many and uh we have here this this picture after the last three weeks and uh a lot of malicious software integrated in the ambient package and something like that in the open source community the many scandals and many stories we have exactly that the open infrastructure that holds everything on the one single and fragile and fragile shoulder so the openness is essential for the science it it is essential for a knowledge just because there is another citation actually I don't even remember from where but it's very very ancient And it speaks that when people are crying when the people are dying the cities were crying when the the let's say the dealing and the businesses are dying it's a translation so please forgive me if I'm not not so precise but the civilizations will cry when the books will burn. So this is what we are sustaining right now. This is why we should persist and make the science resilient and this is the main topic for us very well illustrated with this picture to do that for just for last times the fair principles were invented and established that the findability accessibility interoperability and reusability but all of them are requiring that one let's say from the first look the simplest thing the persistency in the identification and the identifi once you know where to find it then it's trivial to make all the other things doable. So here are the principles and we are thinking about how to implement that. And here we have risks and crisis. And reading this simple table on the slide where we can understand that for every point emphasized here is the obvious crisis we have right now. For example, when we we are speaking about the institutional failure, when you're losing your website, you are losing all the data in the most cases. And when you're losing the funding, you are losing the server. So again, you are losing the data. And it happened for many many for many many cases all around the world. And the last case we all remember it's the health governance data governance in the USA for the legal pressures the content is intentionally made inaccessible as I am Russian I could say that now it is doing it's it's being done in Russia every day we are losing actually gigabytes of data every day and we all know what I'm speaking about it's a warfare in Ukraine the people are disappearing without a prompt. That's the first point and the second point which is an euro European Union some months ago actually three or four months ago one of the most prospective technologies about identifying and widespreading the data which uses is in declimate the interplanetary file system was listed in the piracy and counterfeit watch list of a European commission as a technology not as uh let's say the factor or the use case as a technology. So the link linkage to the data is broken as a premise not a promise we can say and for the third point for a neglect and censorship again we all understand what I'm talking about when it's something that someone is intended to prohibit to reproduce and being redistributed then it will be the link that points to the wrong position. And if that knowledge is inconvenient, we should still have a right at least to identify and to find that and to prevent this point of crisis. But obviously we still have it for a global systemic risks, for the climate change, for a warfare, for other regardless of in institutional intents, we can do obviously nothing. But here is a point. Here is a point. We should stay resilient because when we are starting decomposing the term of identification, we are going to this pyramid. And if we are we will decompose every question mark here. We would obviously say that the right answer is no for every question mark here. But only having the let's say the true condition here we should obtain the goal we have the persistent over the crisis for the given data and for for the given knowledge at the end at the very end and here is how the story the story begins and now we are going to the p problem of p persistent identification I'm sure that everyone knows what the problem is. But now it's about the evolution. The evolution started many years ago from the uh digital object identifier alliance. But actually we will see that it has not evolved together with a community and with a sustainability proposal. Actually it was a project uh it was commercial project of the global global north and to which it led today. Some countries are now out of the scope for the DOI not because they are poor not because they just cannot afford it but because of the um for their arbitrary reason. It could be done and could be done by the division uh by the decision of a foreign access control division just by one country and without anything else without anyone else. So the sanctioning is literally a legal pressure what was described in the previous slides and this is how it works. The worst thing is that DOI is now probably the only most widespread persistent identification system which is actually was at first commercially intended and it locks the data and the metadata over the payw wall. So it substitutes governance over the metadata with a sovereignity and this is a proof of that. This is a copy of a letter on the mailing list with a link as the presentation will be shared. It will be preserved preserved and that it's a letter by John Mallaloy about identifying the metadata and the publications. So we are seeing that it's the system formulated as a commercial foundation. So despite it was it was intended to be done like that. The system became the main the system became huge and the system became the leader. But now this leader is frozen it in its evolution and we have uh according to the last to the uh articles of the last two years we have a lot of rotten links on DOI. So you have a link but you can click it and you will not see anything just because the metadata would be preserved but the another bunch of data you want to access are no longer presented and also there is a very it's a huge use case about the so-called sensitive data because the main question here if you are accessing it over the public system like UI is should we know and should we allow others to know that these data are sensitive And the governance model of DUI actually is intended to be prohibited. So we have here the directly the platformization. So we have a platform and this platform actually instead of holding the access management it holds the sovereignty over our metadata. And we are coming to the to the main problem which I would emphasize instead of having a s single point of failure and the projects like a huge centralized and let's say uh single uh singular uh singularly controlled the decentralized redundant and and resilient by design. And this could be one of the possible solutions is a decentralized identification made by now mostly by the enthusiast and the voluntary projects because actually the decentralized model of identification is based on mathematical environ. So it's as persistent and resilient as the underlying network and standards are. And the best model we have right now about how it could be done just speaking a bit in advance is a bit torrent network used by not only by the info tech pirates is also used for example by Ubuntu by the canonical incorporated to widespread the Linux distributions but this system is actually lacking of many many features we have for a P for persistent identification We don't have now working model for every participant where every participant verifies everything. So the client side reproducibility which is a necessary for fair we don't have steel but here not to be too long I should emphasize the two last points the decentralized persistent identification does not challenge already fine the traditional model it should coexist because otherwise we will lose metadata instead of reproducing them And the very important point the last on this slide it doesn't depend on the blockchain workflow it's not it's not a necessity and now we are going uh actually because I am doing this I may say that we are in the process of proving that it's not necessary blockchain is not a key technology it's not a pania it's not a silver bullet for this kind of identification But it exists already something and there are many other aspects to make long things short here is a what what is decentralization is in the context of a data sharing it's a collaboration decentralized collaborative ecosystem let's say we can say decentralized we can say federated it's something like uh the different sides of the same metal But it's not just a technical solution. It bases on the openness because you are disclosing yourself. It b it is based on responsibility much. It's very important that the computation stakeholder as we have in a blockchain network should not be an identity owner. He should not hold your keys for you. And the accessibility is made by design. So the decentralization brings us the identification as a social good and not a social contract for everyone. So the model is this model is this version repository but let's say widespread over the network. It's very actually it's very simple provenence and uh reproducibility by design. This is a model and here are the some solutions made by by the project on the global north and global south based on decentralized technologies. The silabs in North America is are doing DPI which is based on the hashes. So it's a mathematically determined persistent identifier but it still couldn't be considered truly federated. It's considered decentralized but it's very difficult to deploy. Here's another problem. Arc alliance in the South America and partially in African P alliance. So also on the global south and uh it's uh sister project distributed arc are going to strive to introduce truly decentralized persistent identification. So there are initiators and there are numerals over the world but now we have a still lower not in leadership and and especially their leader uh it's in leadership with uh uh not with collaborating but with funding and with uh communication with stakeholders. Here we have an alternative and this is actually the last slide that it's questions for all our future. It's about the trust because the cryptocurrency now are concentrating on the platform. So we have the still the platformization aspect like Ethereum like a bitcoin and also like the platforms like a task tax hub on some cryptocurrencies. So they are keeping as centralized that the government of the underlying state is. So there's a uh question about the trust. The second is the stakeholders. It's very obvious who should be and there is no answer still. We don't know. We should just to suggest and to prepare about the sovereignity still the infrastructure should it be owned by society and what society should it be? It's very important question but we still don't don't have answer and the levels of severity is come from that also from that in the context of of redundancy how should it should be done the developers of the projects like bit torrent are saying it should be done explicitly by design someone else may think in a different way. The last is a how to handle the social conflicts and should be the mechanism built in the persistent identification solutions or not. All these questions are now without the answers and we should just keep them in mind developing our future for the big data. And for now this is the end of this this short story. Thank you very much for your attention and I'm ready for for the questions. So any question from the audience? Don't be shy. No first session of the conference but come on. >> Yes. Um, how would a blockchain handle uh something like adversarial attacks that are intended to poison the content of the blockchain for say specific regions? Like for example, if I wanted to attack a blockchain that is in the that that has and make it inaccessible in the US, it would be some you know one possibility would be to put uh CS CA CSM encoded into it or linked to it so that then possession of that blockchain would be illegal in that jurisdiction. Uh it's actually it's very simple because now we have for example the now uh now it's very simple uh once we have for example in the UK the law that requires you to disclose whatever you have in the in the context of encryption so the private keys you have no right to you have no right to avoid it uh you should just make the authority able to made a hard fork over your your blockchain. So there is no need to make an attack in context of a legal pressure. It's very easy. You may just hard fork it and replace the um the uh non-resilian blocks >> very quickly because we have to move on. Sorry. as a followup then that hard fork would effect effectively allow a jurisdiction to censor anything they want. >> Yes, exactly. And uh let's say the attempts were already were already done not for the generalized blockchains but it was done for example in Russia for some blockchain uh blockchain based solutions.