Advancing Vocabulary Harmonization in Seismology Through AI and Community Co-Creation
Watch on YouTubeVideo summary
Julian, a researcher from the University of Bergen, introduces an initiative aimed at harmonizing fragmented vocabularies within seismology through a collaborative effort involving artificial intelligence and community co-creation. This project, supported by the JOInquire framework, addresses critical challenges such as inconsistent definitions across various glossaries, the lack of machine-readable metadata, and the difficulty for outsiders or beginners to navigate the field's terminology. The primary goal is to establish a controlled, authoritative vocabulary that facilitates data discovery and enables cross-disciplinary research, moving away from free-text metadata that currently weakens the standardization of seismological data.
To achieve this, the team conducted a comprehensive gap assessment against eight minimum requirements for seismological vocabularies, identifying significant fragmentation where certain definitions were missing or inconsistent. They collected and classified terms from eighteen authoritative sources into ten thematic branches, prioritizing them based on complexity levels ranging from fundamental to emerging concepts. The core of their solution is a simplified pipeline that processes these diverse inputs through four distinct phases: it routes single-authority terms directly to output while subjecting multiple authorities to an algorithmic triangulation process for consensus building. This system utilizes three leading large language models to synthesize definitions and identify synonyms, applying confidence filters where terms exceeding a 60% similarity threshold are accepted automatically, while others are flagged for expert review.
The framework emphasizes a closed-loop validation process where machine-generated evidence is refined by human experts via a GitLab-based voting system, ensuring that the final vocabulary reflects both computational efficiency and community consensus. Experts are invited to review specific terms, provide corrections, or suggest new authoritative inputs, which are then reprocessed through the pipeline until the required confidence levels are met. This approach not only generates over 2,500 compliant vocabulary entries but also integrates concepts into broader metadata standards like EPOS-IC, enabling better tagging and machine readability for data integration across domains.
Ultimately, this project represents a shift toward a sustainable model of knowledge management in seismology by handing over the framework to the FDSN community for long-term governance and maintenance. By combining AI capabilities with active expert participation, the initiative creates a lightweight yet robust system that encourages broader contribution from the scientific community. The resulting controlled vocabulary serves as a foundational tool for improving data interoperability, supporting researchers in finding and joining datasets across different domains, and fostering a more unified understanding of seismological concepts globally.
Read the full video transcript
Um my name is Julian. I'm researcher
from University of Bergen. Uh we are
here to share the experience about uh
harmonizing vocabulary in in
systemology. So this um efforts and the
initiative has been in the context of
the jo inquire project and I acknowledge
the cos especially Angelo in the room
who initialized this idea to bring
together the vocabulary uh in sysmology
and trying to harmonize them. So we come
up with this idea of harmonizing based
on the following gap and challenge that
we have seen. So in sysmology vocabulary
we have a vocabulary that is fragmented
and we don't have a community um driven
and to have a a full vocabulary backbone
that we can rely on.
Also the sec the second challenge that
we have seen as well that we have seen a
lot of vocabularies in terms but it
tends to have different meanings. So we
didn't know really well what to believe
in and we don't have a a set this a a
proper platform to discuss about the
definitions. And then the third one it's
also um we have those great the
glossaries but it doesn't have any
arise or idea where to find and where to
direct the machine uh readable all of
the vocabulary that we have.
And the third and the fourth is that all
of the met meta data in uh in in um
sysmology they tend to be a free text.
We still don't have the tagging the meta
data that has been already like um
explained in the previous um um um um
um uh slide before and that this
actually fragilized the vocabulary in
sysmology. And the last one which is the
most important one which is I'm par with
it when you are outside the sysmology
community and sysmology cycle you you
you may not understand what's going on
in the sysmology especially when you are
on the development side and as well or
you are a beginner you would tend let's
say to lose focus where what are the
vocabularies in six face
And now sorry um so um to um to balance
this dis this claim we have done this
gap assessments in the v vocabulary in
sysmmology so what I'm showing here we
are non non exhaustive lift all of the
vocabulary and the glossaryy that we
have in sysmmology against eight um a
minimum requirement that's supposed to
be completed when we have vocabulary in
systemology sort of let's say
definitions machine readable your eyes
the crosswork and the structure
the structure of the output and the
depth in sysmology meaning we have the
classification and all of the structure
we have and it's also not governed as
you can see green in presence reds
completely absent and yellow partially.
So as as you can see the vocabulary in
sysmology quite fragmented sometime it
has certain minimum requirements
sometime disappear.
So what we wanted to do is to try to
complement and have all of those um um L
into the eight minimum requirement that
we will have all of them in grid. This
is what we are prop proposing. We call
this sysmo that will help u the assisted
by AI but mostly within the community
coation.
So in one sentence we would like this is
what we wanted to do. We wanted to
develop a conceptual framework that
assisted by AI but it really um um um
belong to the community that will
validate everything that we will have a
controlled and authoritative vocabulary
in sysmology for data discovery at the
same time also can use those vocabulary
as a data for cross disciplinary
research.
So to start we we have uh collected all
of the ex existing um um vocabulary and
antology that we couldn't find in
sysmology and we class them and sort
them in sort of let's say a priority.
The priority can be let's say set based
on everyone let's say liking and
preference based on what what type of
vocabulary would like to do. We don't
know we wait them based on a priority.
So we have found uh 18 authority sources
and we part and collect some some of
them through PDF or through um uh to
websites into JSON file and
also collected those who are online and
in the figure B we we we started making
the classification with all of them with
a 10 10 theatic branches based on comp
complexity from level one to level four.
We we we just attributed level one
fundamental and then level two and level
four little bit level three little bit
advanc and level for all of the emerging
emerging terms.
And this next one here it said the
simplify um vocabulary pipeline. How did
you how did we do it? How did you try to
harmonize all of this vocabulary from
the 18 sources that I haven't um pre
previously explained? So through these
four um four phases we try to bring all
of the corpus all of the ter that we
have found from the the um 18
authorities and uh we uh we we bring the
uh the authorities and through the gates
and then once we have the gate if we
have sort of let's say one authority on
the gate we go through path eight which
is go straight into the pipeline. and
then we go through um the outputs and if
we have multiple authorities it goes
into B which I will explain and then if
we don't have any authority to from our
sources it go to puff C those puff B and
and the puff C will activate algorithm
one that we call it in the
triangulations and it will have all of
the consensus and the reconciliation
that will have a definition And if does
not it go straight into the path a
straight into the output. At the same
time the algorithm two is um you
defining the synonyms based on the
authority that you have found
and um through the gates through the
gates if we have a certain filter we we
apply the coin similarity at some point
if reach above 60% we make it tighter.
If the filter doesn't go through, it
goes to val validation experts. It gives
us a flag that the puff C and the puff B
if we have a multiple authorities if it
doesn't fit the requirement that we have
applied an algorithm and an algorithm to
it's bring a flag for expert that will
be checked later and all of that
together through the through those
mechanism we have a two artifacts. We we
have a a a JSON file and also a JSON
lead file which are the structure. The
difference is just the structure and the
JSON leader will have a um domain domain
domain specific concepts.
Here is um um a glance of the results
the numbers that we have produced so
far. So among all of the 18 authority
that we have found and we we launched
through the mechanism that we have now
we have a 260
um um um uh 26600
vocabulary in JSON file um divided
through the classified domain that we
can see over here and through all of the
levels that I have mentioned before and
now we have um almost 500 Right. This is
um um we haven't updated the one that we
have more 500 flag the vocabulary flag
the term that needs to be assigned to um
a reviewer that they will approve. And
so to be concrete this is um um um some
of the outputs. Normally we will have it
in JSON file but it's a little bit scary
to show in the screen. And we have here
a a vocabulary browser that just to show
for uncommon non-technical people that
what we have we just to choose those two
words two words in the level one
remember the path B I've mentioned is
like when we have multiple authorities
meaning like we have multiple
definitions how do you reconcile and
synthesize them on one so first in the
in in the contents of the of the
vocabulary we have the first authority
the one that reattributes based on the
priority and based on weight that I have
explained before. This is the first
definition that we have and the second
one we have the reconcile and synthesize
based on the um um algorithm number one
that I have mentioned that activated the
threeurations that this is the
synthesizer definitions based on the
multiple source since I have mentioned
the path be meaning that we have
multiple authorities and those are the
sources that has been synthesized to get
the definition number two and also
produced the true algorithm. Remember to
have a um um the synonym. So call it ar
and also um um set [snorts]
well the schema offer exact match close
match and the broad match and related
and based on the filter based on the
question similarity that we have applied
before and um if it's above 60% so we
have um the ashes for instance this term
go straight into the outputs and then
the body wave it didn't have enough
confidence. This one has a high
confidence straight and we have a low or
medium more or less it goes straight
into the um review flight. This is a
second a second one to make it much more
clear. If you have a puffy with no no um
authority founded it go straight into
the algorithm too and have those um and
the the the
um the the the models that you use. I
forgot to mention that we are using the
three highest highest originaling large
language model in the market at the
moment. So we use those those three um
models to to have the synthesizing and
the reconcilation to have this
definitions and the same as well as the
previous one. And then we have synonyms
and um it says and it's set um clearly
in the outputs that this has been
generated by um um so and so model
clearly with the confidence and with the
threshold that we have computed.
So once that is done when we have the uh
flag um um review review flag it goes in
this uh step that the community should
close the whole loop. This has been the
say um everything that I've mentioned it
was machine assembler evidence now it
has to be a community validated
knowledge. So through the the the gates
that I have mentioned before now it's
all single um review flag will be um a
gitlab issue we send it through gitlab
and we assign each definition that
require flag um for a voting and for a
review for each um uh expert and then we
we will collected the much
review and much voting as possible to
have a good confidence and that we send
it back again in the loop, send it back
again in the framework that I the
pipeline that I have mentioned before
until the filter until the until the the
threshold is reached above 60%.
So how how it works? So uh when we have
those u flag reviews, we we send um um
uh an issue in the GitLab and then we
invited experts. We invited experts um
through certain them theatic in
sysmology. As mentioned we have a 10 10
domain specific and we select those um
expert and we send them an email going
straight into this directory of of of
GitLab where we can find all of the all
of um step by step how to do it and it's
pretty straightforward. you just to go
in and just then thumb up if it's good,
if it's not good and if you want a
suggestion indeed if there are other the
correction coming. We make it as simple
as possible to not um uh push away and
that uh to not push away expert and
especially not technical people that
would really would like to collect as
much possible out of what they have to
consume. So here is um a few um few
glances and um snapshot of what we do.
This is when we open the work and it
will attribute to show you this is for
instance the the terms and uh it shows
exactly the the pipeline that we have
mentioned and also the classification I
don't know how many people has already
it um this is one of the expert expert
has done it he didn't want to he wanted
it to be anonymous so we blank he hidden
his name but this is how it works um one
person did the up the one person doing
it down and put corrections. So at this
point we have another new authoritative
input authoritative that we didn't get
from documents authoritative from
individuals. This is also really
important. So we pull it back again as a
new authoritative per expert and then we
run again um our um um um our mechanism
until we will get the right definitions
and the right input. So in term of
tagging um it has been we are we are we
are not reaching this level yet but we
are it is already there in the pipeline
that we would like so to add identify we
we was um the was the case in EPOSIC
European plate observation um here um um
in Europe. So we use this um um cases as
this is the term this is the term this
is domain and the description and then
the keywords and that they use usually
have have inated
data at some point when we start the
tagging and apply the eyes and the
vocabulary that we have um uh we open a
new tab as a concept. This is the way we
show you as um as I said for visual
purposes but behind behind the scene we
have a post decat AP in the meta data um
a lot of um a lot of mechanism also work
behind that. So this um introduction of
the concept in the crosswork and in the
broad match that I have uh shown you in
the out outputs will be introducing here
and then this will allow to to find the
data and the join of the crosswork on of
the um um specific to each other domain
and also machineable.
So um the whole pipeline and the
framework um available online um the
next um um updated will be uploaded in
few days. We have a new version of it.
Um it will be found it will be found in
the in in the GitHub where we can find
everything all of the um everything you
needed. It's a pretty straightforward
and it has been as well um uh presented
in EGU and an upcoming draft and a draft
is also coming supporting this
framework. So um we we now handing over
this work to the FDSN
and uh that will be for the governance
and for sustainability uh there is there
are already a communication with Jophone
and EPOS rig for the sustainability of
this work.
So in some uh we have provided um a
framework a framework to harmonize
vocabulary um in sysmonology um pairing
at the same time um AI assisted and also
um um expert co-creation we have um more
or less 200 uh 2,500
um um um um K compliant of vocabulary
control vocabulary
And we have established this new
approach in how to how to bring the
reviewer how to make this community
co-creation to to be light and we hand
over this um work in FDSN community that
most of the FDSN community are the ex
expert and the can can as well maintain
all of the all of the contents and the
new outputs that we are bringing in. And
the and and the last one we we also
would like to have more experts. We
would like to have more people that can
have a that can bring more contribution
in term of refining the output if
needed.
But that's all from my sides. Thank you
for attention and uh um um question will
be welcome.