Submind YouTube summaries
Thumbnail for Advancing Vocabulary Harmonization in Seismology Through AI and Community Co-Creation

Advancing Vocabulary Harmonization in Seismology Through AI and Community Co-Creation

Watch on YouTube

Video summary

Julian, a researcher from the University of Bergen, introduces an initiative aimed at harmonizing fragmented vocabularies within seismology through a collaborative effort involving artificial intelligence and community co-creation. This project, supported by the JOInquire framework, addresses critical challenges such as inconsistent definitions across various glossaries, the lack of machine-readable metadata, and the difficulty for outsiders or beginners to navigate the field's terminology. The primary goal is to establish a controlled, authoritative vocabulary that facilitates data discovery and enables cross-disciplinary research, moving away from free-text metadata that currently weakens the standardization of seismological data. To achieve this, the team conducted a comprehensive gap assessment against eight minimum requirements for seismological vocabularies, identifying significant fragmentation where certain definitions were missing or inconsistent. They collected and classified terms from eighteen authoritative sources into ten thematic branches, prioritizing them based on complexity levels ranging from fundamental to emerging concepts. The core of their solution is a simplified pipeline that processes these diverse inputs through four distinct phases: it routes single-authority terms directly to output while subjecting multiple authorities to an algorithmic triangulation process for consensus building. This system utilizes three leading large language models to synthesize definitions and identify synonyms, applying confidence filters where terms exceeding a 60% similarity threshold are accepted automatically, while others are flagged for expert review. The framework emphasizes a closed-loop validation process where machine-generated evidence is refined by human experts via a GitLab-based voting system, ensuring that the final vocabulary reflects both computational efficiency and community consensus. Experts are invited to review specific terms, provide corrections, or suggest new authoritative inputs, which are then reprocessed through the pipeline until the required confidence levels are met. This approach not only generates over 2,500 compliant vocabulary entries but also integrates concepts into broader metadata standards like EPOS-IC, enabling better tagging and machine readability for data integration across domains. Ultimately, this project represents a shift toward a sustainable model of knowledge management in seismology by handing over the framework to the FDSN community for long-term governance and maintenance. By combining AI capabilities with active expert participation, the initiative creates a lightweight yet robust system that encourages broader contribution from the scientific community. The resulting controlled vocabulary serves as a foundational tool for improving data interoperability, supporting researchers in finding and joining datasets across different domains, and fostering a more unified understanding of seismological concepts globally.
Read the full video transcript
Um my name is Julian. I'm researcher from University of Bergen. Uh we are here to share the experience about uh harmonizing vocabulary in in systemology. So this um efforts and the initiative has been in the context of the jo inquire project and I acknowledge the cos especially Angelo in the room who initialized this idea to bring together the vocabulary uh in sysmology and trying to harmonize them. So we come up with this idea of harmonizing based on the following gap and challenge that we have seen. So in sysmology vocabulary we have a vocabulary that is fragmented and we don't have a community um driven and to have a a full vocabulary backbone that we can rely on. Also the sec the second challenge that we have seen as well that we have seen a lot of vocabularies in terms but it tends to have different meanings. So we didn't know really well what to believe in and we don't have a a set this a a proper platform to discuss about the definitions. And then the third one it's also um we have those great the glossaries but it doesn't have any arise or idea where to find and where to direct the machine uh readable all of the vocabulary that we have. And the third and the fourth is that all of the met meta data in uh in in um sysmology they tend to be a free text. We still don't have the tagging the meta data that has been already like um explained in the previous um um um um um uh slide before and that this actually fragilized the vocabulary in sysmology. And the last one which is the most important one which is I'm par with it when you are outside the sysmology community and sysmology cycle you you you may not understand what's going on in the sysmology especially when you are on the development side and as well or you are a beginner you would tend let's say to lose focus where what are the vocabularies in six face And now sorry um so um to um to balance this dis this claim we have done this gap assessments in the v vocabulary in sysmmology so what I'm showing here we are non non exhaustive lift all of the vocabulary and the glossaryy that we have in sysmmology against eight um a minimum requirement that's supposed to be completed when we have vocabulary in systemology sort of let's say definitions machine readable your eyes the crosswork and the structure the structure of the output and the depth in sysmology meaning we have the classification and all of the structure we have and it's also not governed as you can see green in presence reds completely absent and yellow partially. So as as you can see the vocabulary in sysmology quite fragmented sometime it has certain minimum requirements sometime disappear. So what we wanted to do is to try to complement and have all of those um um L into the eight minimum requirement that we will have all of them in grid. This is what we are prop proposing. We call this sysmo that will help u the assisted by AI but mostly within the community coation. So in one sentence we would like this is what we wanted to do. We wanted to develop a conceptual framework that assisted by AI but it really um um um belong to the community that will validate everything that we will have a controlled and authoritative vocabulary in sysmology for data discovery at the same time also can use those vocabulary as a data for cross disciplinary research. So to start we we have uh collected all of the ex existing um um vocabulary and antology that we couldn't find in sysmology and we class them and sort them in sort of let's say a priority. The priority can be let's say set based on everyone let's say liking and preference based on what what type of vocabulary would like to do. We don't know we wait them based on a priority. So we have found uh 18 authority sources and we part and collect some some of them through PDF or through um uh to websites into JSON file and also collected those who are online and in the figure B we we we started making the classification with all of them with a 10 10 theatic branches based on comp complexity from level one to level four. We we we just attributed level one fundamental and then level two and level four little bit level three little bit advanc and level for all of the emerging emerging terms. And this next one here it said the simplify um vocabulary pipeline. How did you how did we do it? How did you try to harmonize all of this vocabulary from the 18 sources that I haven't um pre previously explained? So through these four um four phases we try to bring all of the corpus all of the ter that we have found from the the um 18 authorities and uh we uh we we bring the uh the authorities and through the gates and then once we have the gate if we have sort of let's say one authority on the gate we go through path eight which is go straight into the pipeline. and then we go through um the outputs and if we have multiple authorities it goes into B which I will explain and then if we don't have any authority to from our sources it go to puff C those puff B and and the puff C will activate algorithm one that we call it in the triangulations and it will have all of the consensus and the reconciliation that will have a definition And if does not it go straight into the path a straight into the output. At the same time the algorithm two is um you defining the synonyms based on the authority that you have found and um through the gates through the gates if we have a certain filter we we apply the coin similarity at some point if reach above 60% we make it tighter. If the filter doesn't go through, it goes to val validation experts. It gives us a flag that the puff C and the puff B if we have a multiple authorities if it doesn't fit the requirement that we have applied an algorithm and an algorithm to it's bring a flag for expert that will be checked later and all of that together through the through those mechanism we have a two artifacts. We we have a a a JSON file and also a JSON lead file which are the structure. The difference is just the structure and the JSON leader will have a um domain domain domain specific concepts. Here is um um a glance of the results the numbers that we have produced so far. So among all of the 18 authority that we have found and we we launched through the mechanism that we have now we have a 260 um um um uh 26600 vocabulary in JSON file um divided through the classified domain that we can see over here and through all of the levels that I have mentioned before and now we have um almost 500 Right. This is um um we haven't updated the one that we have more 500 flag the vocabulary flag the term that needs to be assigned to um a reviewer that they will approve. And so to be concrete this is um um um some of the outputs. Normally we will have it in JSON file but it's a little bit scary to show in the screen. And we have here a a vocabulary browser that just to show for uncommon non-technical people that what we have we just to choose those two words two words in the level one remember the path B I've mentioned is like when we have multiple authorities meaning like we have multiple definitions how do you reconcile and synthesize them on one so first in the in in the contents of the of the vocabulary we have the first authority the one that reattributes based on the priority and based on weight that I have explained before. This is the first definition that we have and the second one we have the reconcile and synthesize based on the um um algorithm number one that I have mentioned that activated the threeurations that this is the synthesizer definitions based on the multiple source since I have mentioned the path be meaning that we have multiple authorities and those are the sources that has been synthesized to get the definition number two and also produced the true algorithm. Remember to have a um um the synonym. So call it ar and also um um set [snorts] well the schema offer exact match close match and the broad match and related and based on the filter based on the question similarity that we have applied before and um if it's above 60% so we have um the ashes for instance this term go straight into the outputs and then the body wave it didn't have enough confidence. This one has a high confidence straight and we have a low or medium more or less it goes straight into the um review flight. This is a second a second one to make it much more clear. If you have a puffy with no no um authority founded it go straight into the algorithm too and have those um and the the the um the the the models that you use. I forgot to mention that we are using the three highest highest originaling large language model in the market at the moment. So we use those those three um models to to have the synthesizing and the reconcilation to have this definitions and the same as well as the previous one. And then we have synonyms and um it says and it's set um clearly in the outputs that this has been generated by um um so and so model clearly with the confidence and with the threshold that we have computed. So once that is done when we have the uh flag um um review review flag it goes in this uh step that the community should close the whole loop. This has been the say um everything that I've mentioned it was machine assembler evidence now it has to be a community validated knowledge. So through the the the gates that I have mentioned before now it's all single um review flag will be um a gitlab issue we send it through gitlab and we assign each definition that require flag um for a voting and for a review for each um uh expert and then we we will collected the much review and much voting as possible to have a good confidence and that we send it back again in the loop, send it back again in the framework that I the pipeline that I have mentioned before until the filter until the until the the threshold is reached above 60%. So how how it works? So uh when we have those u flag reviews, we we send um um uh an issue in the GitLab and then we invited experts. We invited experts um through certain them theatic in sysmology. As mentioned we have a 10 10 domain specific and we select those um expert and we send them an email going straight into this directory of of of GitLab where we can find all of the all of um step by step how to do it and it's pretty straightforward. you just to go in and just then thumb up if it's good, if it's not good and if you want a suggestion indeed if there are other the correction coming. We make it as simple as possible to not um uh push away and that uh to not push away expert and especially not technical people that would really would like to collect as much possible out of what they have to consume. So here is um a few um few glances and um snapshot of what we do. This is when we open the work and it will attribute to show you this is for instance the the terms and uh it shows exactly the the pipeline that we have mentioned and also the classification I don't know how many people has already it um this is one of the expert expert has done it he didn't want to he wanted it to be anonymous so we blank he hidden his name but this is how it works um one person did the up the one person doing it down and put corrections. So at this point we have another new authoritative input authoritative that we didn't get from documents authoritative from individuals. This is also really important. So we pull it back again as a new authoritative per expert and then we run again um our um um um our mechanism until we will get the right definitions and the right input. So in term of tagging um it has been we are we are we are not reaching this level yet but we are it is already there in the pipeline that we would like so to add identify we we was um the was the case in EPOSIC European plate observation um here um um in Europe. So we use this um um cases as this is the term this is the term this is domain and the description and then the keywords and that they use usually have have inated data at some point when we start the tagging and apply the eyes and the vocabulary that we have um uh we open a new tab as a concept. This is the way we show you as um as I said for visual purposes but behind behind the scene we have a post decat AP in the meta data um a lot of um a lot of mechanism also work behind that. So this um introduction of the concept in the crosswork and in the broad match that I have uh shown you in the out outputs will be introducing here and then this will allow to to find the data and the join of the crosswork on of the um um specific to each other domain and also machineable. So um the whole pipeline and the framework um available online um the next um um updated will be uploaded in few days. We have a new version of it. Um it will be found it will be found in the in in the GitHub where we can find everything all of the um everything you needed. It's a pretty straightforward and it has been as well um uh presented in EGU and an upcoming draft and a draft is also coming supporting this framework. So um we we now handing over this work to the FDSN and uh that will be for the governance and for sustainability uh there is there are already a communication with Jophone and EPOS rig for the sustainability of this work. So in some uh we have provided um a framework a framework to harmonize vocabulary um in sysmonology um pairing at the same time um AI assisted and also um um expert co-creation we have um more or less 200 uh 2,500 um um um um K compliant of vocabulary control vocabulary And we have established this new approach in how to how to bring the reviewer how to make this community co-creation to to be light and we hand over this um work in FDSN community that most of the FDSN community are the ex expert and the can can as well maintain all of the all of the contents and the new outputs that we are bringing in. And the and and the last one we we also would like to have more experts. We would like to have more people that can have a that can bring more contribution in term of refining the output if needed. But that's all from my sides. Thank you for attention and uh um um question will be welcome.