Panel Discussion : NLP: Reconstructing narratives about the ongoing Nakba #Arabic #AI #nlp
Watch on YouTubeVideo summary
This panel discussion centers on the critical application of Natural Language Processing (NLP) and digital archiving to reconstruct narratives surrounding the ongoing Nakba, addressing a significant technological gap between Arabic dialects and established English or Hebrew models. Experts highlight that while technical progress in Arabic NLP is evident, high error rates in speech recognition for non-Modern Standard Arabic varieties persist, necessitating specialized systems capable of processing diverse linguistic inputs including harmful content from Israeli channels. The conversation underscores the danger of relying on biased or censored data found on major platforms like Meta, which often remove hashtags only after sustained pressure; consequently, there is an urgent need to build independent in-house systems that can document suppressed political speech and measure visibility asymmetries to hold tech giants accountable for their role in erasing Palestinian history.
To combat this systematic erasure, the panel introduces collaborative initiatives such as the "Fighting Erasure" project, which functions not merely as a traditional archive but as a chronological platform documenting genocide in real-time through legal definitions like domicide. This approach involves field workers verifying attacks on hospitals and neighborhoods using Geographic Information Systems (GIS) and historical maps to counter settler-colonial logic that fragments narratives along nation-state boundaries. Unlike post-event archiving, these historians are capturing the "live stream" of current atrocities, aiming to preserve terabytes of testimonies before they vanish due to big tech censorship or physical destruction by death. The process extends beyond simple demographics to include forensic details such as specific weapon types used in attacks and precise locations, creating a robust evidentiary base for future legal accountability at institutions like the International Court of Justice.
The discussion also delves into profound ethical dilemmas inherent in preserving sensitive data, particularly regarding medical files from the Ministry of Health that officials initially refuse to share due to fears of their destruction during bombings or targeted attacks by Israel. While some graphic images provided by field workers depict doctors killed and bodies riddled with bullets taken directly from cell phones after families were wiped out, experts face a difficult choice between publishing such visceral evidence for future trials and respecting the dignity of the deceased and living victims. Despite these challenges and past losses like the theft of PLO archives in Beirut in 1982, there remains an unwavering commitment to digitize records before they are lost, ensuring that visual content maintains a legally sound chain of custody while honoring the memory of those targeted by violence.
Ultimately, the panel concludes that successful narrative reconstruction requires seamless collaboration between technical experts and historians to navigate complex challenges such as integrating Palestinian domains into global web crawlers like Common Crawl and automating ethical documentation processes. By combining advanced NLP capabilities with rigorous historical methodology, these efforts aim to maintain context within massive datasets while fighting fragmentation across scattered archives. The consensus is that preserving this vital evidence is essential not only for honoring victims but also for establishing the factual basis needed in future international courts, ensuring that history is recorded accurately despite ongoing threats and attempts at digital or physical erasure by hostile actors.
Read the full video transcript
Okay, so great.
Um
welcome to this panel discussion.
The title is NLP reconstructing
narratives about the ongoing Nakba.
The intention here is to reflect um
on where we stand now and how we can
move forward
uh with the aim of reconstructing the
narratives about the Nakba, the ongoing
Nakba. And the ongoing Nakba is a term
that signifies an
approximately 100 years of
um Nakba, so to speak.
Um I'll introduce first the um
panelists. We have four panelists with
us.
Uh Dr. Kareem Darwish, um who is a
principal scientist at the Qatar
Computing Research Institute, so QCRI.
Um
he is um a long um I know Kareem for a
long time now at the
um conferences of NLP conferences.
His work is has been mostly on
state-of-the-art tools for
Arabic processing and social computing,
um
which has been well received.
We have with us also Mrs. Rama Salha,
um excuse me if I'm not pronouncing
correctly because I didn't see it in
Arabic. Um she is an AI specialist and
technology officer at Hamleh.
Uh
her work focuses on building decolonial
AI,
uh auditing content moderation systems,
and challenging power structures
embedded in technology.
We also have with us Dr. Jamila Haddar.
She's an assistant professor in archival
information and digital humanities at
the University of Amsterdam.
Um she's also a co-founding director of
the Archives and digital media lab and
research affiliate at the American
University of Beirut School of
Architecture and Design.
She is a co-founder of the fighting
eraser digitizing
Gaza's genocide a project
co-founded together with Dr. Hanin
Shahada who our fourth
panelist.
She is an assistant professor at New
York University Abu Dhabi
and research associate at the Maroun
Semaan Faculty of Engineering and
Architecture
at the American University of Beirut in
Lebanon.
Welcome all.
What we'll do here I suggest at least
that I start with
one of you basically start with Dr.
Karim
with a question. Please feel free to
respond to
the others. At some point I hope we will
have free discussion. That's the
intention.
Okay,
I start with Dr. Karim first with the
first question.
So as a leading expert in Arabic NLP
where do you think we stand at this
moment with Arabic NLP
given that we don't have our own
AI as far as I know and correct me if
I'm wrong.
And how can NLP help reconstruct the
narrative about the Nakba and what
resources do we need to do so?
Okay, so somebody can first of all I'm
very happy for the invitation. Thank you
Dr. Khalil. It was a pleasure to to be
with you.
So you ask actually a quite large
question and there are two aspects of
this. One of them is technical and the
other part is
is operational.
On the technical side I think when we
deal with the reconstructing the Nakba
and ongoing Nakba. There are actually
two parts. The first part has to do with
the Arabic
and the other part has to do with other
languages other than Arabic including
Hebrew and English and other languages.
On the English side, I think on
when even in most European you know
languages the the landscape is actually
quite mature.
On the Arabic side the maturity level
has been increasing dramatically over
the past few years.
And this is evident by the large
community of Arabic NLP that I mean like
the mailing list has over a thousand
people.
There are many efforts in the region the
Arab region for for building sovereign
AI
with models like Alam in Saudi Arabia,
Fanar in Qatar, Jason in the United Arab
Emirates, Karmak in Egypt and so forth.
Like if you look at other technologies
like um
speech recognition and and OCR, they
have been steadily improving over the
past few years.
They're not
at par with English yet, but I think we
I think in the next probably two or
three years we'll we'll get to that
point.
So the the missing I think piece is
where the Hebrew part is.
I think Hebrew is is is a lot more under
you know developed compared to what we
have for Arabic and for definitely
compared to to English.
And I think reconstructing the narrative
requires that we actually do all the
languages well.
It's it's on one side you want to you
know I think other panels will talk
about you know that this is archiving of
the content that we are receiving from
from Gaza or from Lebanon and so forth.
A lot of it is coming in Arabic, but we
should not be neglecting what Channel 14
and Channel 12 are
spewing all the time documenting
you know the the their narrative
documenting you know
you know, uh the the way that they see
the world, how the the
they cheer on the the genocide that's
that's actually ongoing. Also, uh Hebrew
social media that for for longest time
people thought that, you know, they're
in isolation and they don't actually see
people the world doesn't see what
they're what they're saying and what
they're doing. That needs to be, you
know, resurfaced and that that would
require
uh not just data collection, but also
uh competent speech recognition,
competent ASR
OCR,
uh and competent language modeling for
for for for that language for for these
languages. So, in that sense,
uh I think Arab English is and European
languages are doing very well. Arabic is
not far behind and, you know, Hebrew is
is way behind and that needs to to be
kind of fixed.
How do How do we stand on dialects? Just
to uh before I move to the other
That's a great question. So,
Okay, so if you compare Arabic, you
know, MSA modern standard Arabic and and
dialects, particularly in the area of of
speech recognition, I think you speak
you know, speech recognition, for
example, is behind.
Uh in the sense that you can get below
10% word error rate for MSA and the
typical word error rate for for for
dialects is around in the range of 80 of
about 20% or so.
Uh
this gap needs to be filled
uh
particularly for you know, for for more
niche dialects or or pronunciations.
Uh so, this is on the speech side. Uh if
you're looking at the the text side, I
think the gap is is is much narrower
than than than speech.
You'll find that most large LLMs,
including Fana and I'm pretty sure I'll
have
and others,
uh they wouldn't have any problem
understanding dialectal text, making
sense of it.
The mapping between the concepts that
happens, you know, implicitly inside of
the model are actually quite robust.
Thank you. That's a great It gets a
great start.
We talked about the technical stuff, so
I'll I'll move now to
Rama, Mrs. Rama Salah.
Uh
So, from your experience and work at
Hamleh, maybe you can tell us a bit
about it because that's in Palestine.
And maybe you can tell us also
how you think maybe uh or where can
uh reconstructing the narratives about
the Nakba
can
be used best used NLP in the best way.
Uh yes. Hi. Thank you for the invitation
to join this panel. I'm glad to be a
part of this um
very important discussion.
Um so, I want to start my um
speech with a framing point that today's
biased and censored data becomes
tomorrow's historical record in AI
systems.
What gets removed or down ranked or um
made invisible becomes a structural
absence that future language models will
then treat as natural.
And on the other hand, what what is
allowed to circulate and be amplified
becomes the normal distribution that
models learn from.
And from this perspective, platform
governance and moderation decisions
extend from shaping
the current public discourse to shaping
the data sets that future AI systems
learn from. Thus, what becomes, let's
say, like narratively recoverable at
scale.
Through our work at Hamleh, we explored
this in several structural ways.
Um in one of Hamleh's publications,
Silent Networks,
um the research shows that a a very
strong chilling effect around political
expression where surveillance,
interrogation, and fear of consequences
reduce online participation, especially
amongst the youth.
And from an NLP perspective, this is
very important because it means that the
data generation itself from the
beginning is being structurally
suppressed at the source, making the
absence of narrative or speech
produced and forced rather than just
organic.
On the other hand, in our violence
indicated work, specifically the one um
monitoring hateful content online
through in-house models that we have
trained,
where we study how um narratives of
violence uh circulate at scale and how
platform governance responds unevenly
over time.
One clear example was the hashtag and
phrase um flatten Gaza, or in Hebrew
it's pronounced lem ho ket Gaza,
something like that, which circulated
widely across across platforms beginning
October 2023,
and then remained like highly visible
for months, accumulating
tens of thousands of uh occurrences
before being made unsearchable by Meta
in mid 2024 after lots and lots of
continuous external pressure from
partners and big-scale data-backed
documentation.
So, what this highlights is the
importance of building in-house
technical systems, like uh Dr. Karim
said, uh that are representative of our
narratives. Because this kind of
technical work, when you combine it uh
with the advocacy and um external
pressure, it can sometimes lead to
changes in platform behavior, even if
only partially, by making visibility
asymmetries measurable and harder harder
to ignore.
What this also shows is that um we
cannot really separate narrative
reconstruction from the systems that
determine what data gets to exist and be
easily and cheaply and cheaply
accessible uh in the first place.
Um and thirdly, through her our
observatory for digital rights
violations,
we maintain a structured data set of
nearly 15,000 documented cases of um
content removal, suppressed political
speech, and hatred and violence online
since 2021. That data is vetted by
humans and like we collected in order to
try to make action
through these platforms where this
content um existed. Which functions as a
counted corpus um measuring what is
missing or unjustly removed from
mainstream data sets. Which is I think
uh it's very essential for understanding
what to look for in data when trying to
analyze and understand bias in AI
models.
So, at this point um what I want to
leave this panel with is that narrative
reconstruction extends from
documentation and analysis to also
become a question of data set
construction under conditions of
asymmetric visibility and platform
governance.
And that unless we explicitly account
for how platforms determine what remains
visible and what is removed, future
language models will inherit
structurally
these missing histories.
This is very good. I mean, um what I
just to uh
summarize a bit what he you know the the
two of the in my opinion at least most
important issues that you raised is that
uh there is an under representation
plus suppression of the uh
narrative if let me call it this way in
in AI but also in the uh you know in
general in um on the media.
Uh and what this combined with what Dr.
Karim said about the Arabic dialects
since most of the for example the videos
are in Arabic dialects,
this means that it is rather complicated
to process and also
to bring together.
I'll give
Apparently, Dr. Karim has a response.
I'll give it to him and then I move
afterwards to
Dr. Jamila.
Karim
If I may, I mean just to second on what
Rama said.
This suppression and is actually we can
actually see it in the current LLMs. We
don't actually have to wait until the
future.
I would ask you to go
on the
you know like on Gemini and ChatGPT and
ask both of them about the Haganah gangs
which were deemed terrorist
organizations and look at the answer
that they give back. The answer is
really really really shocking. I mean if
you look at it you'll think that these
are peace-loving, you know,
tree-hugging, you know, angels who came
to save people while everybody deemed
them as terrorist organizations.
Yeah.
This is all very important because I
think
if I move now to Dr. Jamila Haddad
I mentioned the project that you're
leading together with Dr. Hanin Shadeed.
Could you start with telling us about
that and then respond to whatever you
think is uh
is important to to you in in this
discussion so far.
Thank you very much and thank you for
having me. Everything a lot of what's
been said certainly resonates with me.
We at the Fighting Erasure project the
long title
Fighting Erasure Digitizing Genocide on
the war
on Lebanon um
you know, has it has various components.
Other key part that I work on is on the
archives and record side and we're
really
struggling to understand how and
what it looks like uh to rescue and
recover and safeguard archives and
records uh particularly those uh in my
area of the project uh that pertain to
that have this kind of evidentiary
quality uh whether evidence of the
genocide itself
uh so for example this uh epic and
rather
arguably impossible task we took on of
archiving social media um in a in a
particular kind of archival science way
and um alternatively uh what it could
look like to mass digitize and um
rescue and recover from under the
rubble. But in terms of the evidentiary
it has to do with people's rights and
property and administration and
governance and all the kinds of things
that connect and prove and assert both
the history and present of uh the people
on the land because of course that's the
settler colonial logic uh to try to
eliminate any and all proof of that. And
you know like everybody in various ways
we are struggling with digital
infrastructures and information
infrastructures that are all that that
are all connected to these really
problematic platforms. Um we're trying
to understand and have conversations
around what does a digital sovereignty
and data sovereignty look like uh both
at the infrastructural and theoretical
level uh war zones and conflict areas
and areas perhaps where we may not uh
feel all that much more comfortable in
terms of say the Gulf context which is
really the Gulf states which are really
leading a lot of this cuz it takes all
this money. Um
you know there's a lot of pieces uh to
what I'm saying. Uh one thing is I can
certainly as a as an archivist and a
historian of archiving and um
the last 500 years uh certainly uh
echo what you said
Rama that
and Karim that the language models that
AI is built on already is absolutely
colonialist and racist and certainly
orientalist and the depth to which that
goes like 500 years. I mean you can
actually
it's always problematic to find origins.
We can go back even more of course.
Um
but the interesting part is how actually
in some regards those are being
disrupted because of the mass the
expression on a mass level of sympathy
and various different perspectives and
just the evidence and data on the ground
and so again that kind of really
interesting way and now we're trying to
see how they're trying to actually
modify
the AI to kind of be dynamic in this way
as this different kind of data and
perspective is coming in.
Um there is probably again way too much
I I want to say but maybe to speak the
most directly to your question
is that
for us to reconstruct narratives um
it's worthwhile to kind of map out what
is the narrative structure of genocide
and settler colonialism that they're
trying to impose and certainly a key to
that is
a kind of fragmentation and dispersal
and segregation not only of people and
geographies but also of the narrative.
So we can only ever say this is
happening the West Bank. You know we
used to uh you you know we used to talk
about Arab region right now it's
just by nation state now it's even as we
can all see the profoundly
like cross-border and
regional dynamic nature of the ongoing
like but which has always been the case
but self turkey logics and narratives
uh
want to completely
destroy that. So it what what kind of
role can we actually have AI play and
and kind of creating
a better sense especially when
we're kind of
um
stuck with the way in which data
is around nation states or certain
identity groups that can actually
uh obscure all of that and of course
quite intentionally if you look at the
history of these things. So I'll stop
there. Thank you very much. Thanks.
Thanks a lot. So
this enforces what the previous two uh
uh speakers have said and also add to it
a dimension of
you know the um
uh risk of the data being
uh exposed to uh
uh also to the to AI on the one hand and
we want that and on the other hand to
have the
AI to reflect the balance uh
in the narratives that exist in the
world and which is at the moment
completely unavailable or not the case.
I I'll move to Dr. Hanin Shahada. I'm
sure she has to add a lot on this. So
I've heard her before ask a question we
suggested direction. I'm quite curious
Hanin to what you have to add and at the
same time also please reflect on what
the others have said and feel free to
add.
Sure. I mean thank you all for having me
here and uh
um it has been such an interesting
conference for me especially given that
I'm not a technical person, right? I'm a
historian and so looking at the
struggles from a from an operational
perspective is something that I
recognize as we're trying to narrate and
document and archive and preserve this
this genocide. And so, I thought that
maybe I would show you what we have been
doing in a way to
uh
preserve, right? We're archivists, we're
historians, we're trying to preserve
whatever there is here. What I've heard
today and I've seen repeatedly is that
many of us separately
uh across the planet have been
preserving data and collecting it on
several several different hardwares,
right? And trying to keep whatever
uh data we can collect in the in the
tetrabyte and then look at how we can,
right, in the future do something with
it. And that's also part of the
narrative today because
the way that we see it is that, first of
all, this is the first live stream
genocide on Earth. So, there's a new way
of of perceiving what's happening in a
live stream. Historically, when you
archive where you when you write
history, it's at the end of the moment.
It's at the end of the genocide. So,
when we documented Bosnia, it was when
it ended, right? When we wrote about the
Holocaust, it was after it happened. The
same thing with the Russian words is
that you go back and you try to
recollect what's happening afterwards
and to try to put a narrative. Gaza
shifted the ground on us, right? So,
it's it's it was live streaming as it
was happening. And the same people that
were archiving this ended up being
killed, which means that you have a
testimony without, right, in the moment
of death.
For me is what do we do with all of
this? And it took us about
um I think 6 to 7 months to sit with
colleagues, and we came up with
something that we called the skeleton,
which is a massive website or a digital
platform that we wanted to create that
would
reconstruct chronologically somehow the
genocide with the data that we have and
with a local team in Gaza. So, we have
10
field workers
part of our structure, our office in the
Gaza Strip that have this data. We have
this triangulation of verification,
which means that let's say we want to
document the attack on the staff of
Shifa.
We have what's online, right? But we
don't know is this a picture of Shifa or
is it Madani or is this that? Where is
this? So, they take that content, they
go to the Ministry of Health, they go to
witnesses, they go they they record
testimonies. So, we have an entire block
to recreate that moment and we have been
doing this since October 10th. It's
massive, massive and it's beyond our
capacities. Just in terms of time, I
just wanted to show if possible just
share my screen to give you an idea of
of
what what in the future this
this project is meant to look like.
So, I think you have the right to share,
but I'm not sure completely.
Yep.
So, the idea of this big platform is
that it will have the main components.
So, we we have spoken to legal scholars
that have defined for us first what is
genocide, right? And what falls under
genocide. And
under these legalities, we constructed
every single one of those elements. So,
you can see on the website genocide,
domicide, herbicide. I mean, by now all
of you are very familiar with those
terms. The attacks on the hospital. And
the idea is that if you click on each
one of them, it will pop up and open up
into this massive
uh uh
second website that will provide you the
numbers, the the statistics, the
documentation, the the doctors, the
hospitals. I'm scrolling. I don't know
if you see it with me. But, let's say
we'll give you an idea that before the
war, Gaza's health network had 13
governmental hospitals, 18 UNRWA
facilities, five private hospitals. You
click on each one of these, it pops up
into a new window. And the idea is that,
you know, all of those terabytes
eventually that we have that they will
be condensed into one space, one massive
online archive that has been that will
reconstruct this entire genocide
precisely to fight erasure. And this is
hence the name of our project, Fighting
Erasure, is that meta is deleting,
right?
Uh big tech is removing. We have our
people in Gaza, they are documenting and
safeguarding. And it's not a work that
comes without without So, this is to
give you an idea. So,
you have, for instance, the chronology
of the attacks on the hospitals. You see
on October 7th, 2023, it's the beginning
of the genocide.
Uh we see that already the initially
they had a health care facilities were
targeted. You click on it, you get all
of the details. You go to October 15th,
2023. This is when the siege of Al-Shifa
Hospital begins. You click on all of
these tabs, they take you to other
spaces. And then, right, you go down,
you have Kamal Adwan has been targeted,
Nasser Hospital Nasser on January 22,
'24. And and and
etc.
And the idea is that we built this to do
the same thing with, let's say, if you
go to Herbicide, with the neighborhoods,
is that now we're trying to um
look into one neighborhood after the
others. And we we have, because it's so
massive, we thought that we would start
with the north of Gaza that has been
massively
decomposed. So, with neighborhoods like
Sabra, Tal al-Hawa, Sheikh Radwan,
Jabalia, Muaskar, Mukhaiyam al-Shati.
And and then look at the old maps that
the municipality had for those places.
So, our field workers are going to the
municipality, see if they have the old
urban planning of these areas. And then
with GPS, sorry, with with GIS try to
see what has remained, which units are
gone, etc. So, you can tell it's a
massive project. Um
but we're trying by doing this is to
undo this this or to fight erasure. And
whatever data we have now and locally
and and bless bless all the people of
Gaza who have
given us all of that
content. I hate to call it content, but
they're testimonies to this genocide.
And I think part of our job as
historians is to fulfill that that
promise to them that
what they have sacrificed their lives,
we will not leave it there to right to
to not be seen.
Um so, these are This is the big project
with the aim of where we want to go. You
can understand that we also face some
structural
obstacles. Some of them are operational.
This is what I was talking with Dr.
Alexei about. The idea is
I'm putting the burden now on Dr. Jamila
because she's the scientist archivist,
but she has been tasked with creating a
system to catalog and categorize and tag
and filter and authenticate all those
mega terabytes of of data. So, I can
take it from her and then just put them
in each one of these sections. Ideally
that's what I want and she told me this
morning, "Give me 6 months." So
>> [laughter]
>> So I
but the idea is that right is that we
have now each one of us loosely has that
kind of data. I am not a scientist. I
cannot
But once I have that information, today
we have the platform that can absorb
that that sort of all of that material
once gets cataloged. So that's for my
presentation. I tried to stick to the
time as much as possible. Thank you.
Thank you. Thank you
I mean so I'll move to Dr. Karim as hand
go on, please.
Uh so basically thank you Dr. Karim and
Dr.
Jamila for for your great effort.
I guess one of the things that
that I think we need to find a way that
we can kind of speak the same language
in the sense of
we have lots of technology that
potentially could be of use in that in
that regard.
Uh transcription,
data analysis, data extra you know
information extraction from from text,
uh
you know even image classification
and and so on and so forth. So so the
tools that at our disposal
whether it's in English or or Arabic and
that in this in this regard
I I think is is quite is quite strong.
But you know um
we don't see what you what you see. You
see I mean like you can be our eyes in
developing the technology for to your
benefit
if we can find a way that we can speak
kind of the same language and you say,
"This is the kind of data that I have
and this is where where I want to go."
And perhaps we can build you or help you
build that bridge.
This is great. I mean this all sounds
like a project proposal.
>> And and be careful cuz we will take you
on it, right?
I will be in your inbox, so
>> [laughter]
>> So, that's that's exactly I know we have
so much potential in the Arab world.
It's not a lack of it. I think at this
point is really a way of of finding a
way to set up this network. We all want
the same thing. Um and uh I don't I
don't speak technical, right? But I we
have like I speak legal, history,
archiving, and so we have created this
skeleton with with legal scholars, with
historians, with name it. And now I'm
look we're looking for the for for that
what you just said, Dr. Karim, is to
sort of find a way to put all of these
systems together just to build in and
find a way to put that content uh
there and to safeguard it as well,
right? There's also the thing that you
don't want this
this website to fall apart or to be
deleted or to be So, you need also to
protect the the domain that we're on.
So, there are many conversations that
are happening at the same time.
Oh, great. So, this is a very good I
mean it's very inspiring. And I think uh
there are ways to collaborate
apparently. So, I hope this matures at
some point.
Um
I was um
thinking um at the same time um and I'd
like to open the floor for others to uh
join us also from the audience. But if
others want to raise other points,
please feel free. But at the same time I
was thinking
we are at at um
the probably the the most
um
horrible episode of the ongoing Nakba so
far.
But it is an episode nonetheless. And
there has been there have been multiple
episodes before that.
And where
technology, particularly NLP, can be
helpful is also to see
how to connect the episodes together for
historians.
And whether that is similarities, but
also dissimilarities.
And I know we have a lot of
technical
um
uh difficulties, but at the same time
there are
uh
tools that already work, uh so to speak.
So, I'm open that for everybody to think
about and maybe discuss. And if there
are others from the audience, they're
welcome to join us.
So, how does the current episode relate
to the previous episodes of uh
uh
And and would it make sense what I just
said about
bringing the episodes together in one
way or another?
I actually if I if I may, I think this
episode
obviously is a lot more bloody compared
to the previous ones.
Uh but but uh you know, the the cookbook
that they're using is is very very
similar, right?
Uh I mean, now they say, "Oh, you know,
Hamas is a terrorist organization and
thereby justifies everything, right?"
You go to Lebanon, you just replace the
word Hamas with some other entity like
Hezbollah, and then that's the excuse.
Before that, it was the PLO. Exactly.
And and I mean, like the the narrative
doesn't change, right? And and the the
Hasbara efforts have been, you know,
quite consistent in the same narrative,
the
the same way, and so forth.
Um
given you know, like as as humans, we
can we can you know, like we we see it
and we like recognize it right away.
Uh whether the AI can can recognize it
in the same way, I I there there must be
a way.
Uh but but
uh the end result of how, you know, like
if you have like a structure and you
say, here are the inputs and here are
the the predictable outputs,
I think historians would probably need
to tell us what kind of outcome that
would they would they be interested in
seeing, right?
Because because of the similarity, I
mean, it's the same narrative almost
every time.
You just change the names and the
locations and it's the same.
Indeed. Yeah.
Go ahead,
Rama, please. I couldn't agree more with
Dr. Karim.
And I think there's um like this time is
a very special time with an opportunity
um to make
really big steps. Like right now with um
AI agents boom, where um
like most AI models are based on
algorithms and the agents um and
automated uh collection of the data to
get answers. I think there's a lot of
work to be done and um
better designing the platforms where
historians get to um
document the reality. Uh for example, to
make it like bot-friendly.
Um
and like um and helping and making these
platforms um have better rankings. Thus,
um
um look uh like more authentic and more
um
What we What can we call it? More
believable in the eyes of a bot. So,
like if we work in that direction, um
like the the tech where the tech sphere
comes uh hand in hand with the
historians and social sciences
um um
um researchers and political science
researchers, I think there's a lot of
space um
to push the work because there's already
a lot of research being done by amazing
people documenting so much of what was
happening um since October 2023 in Gaza,
Um what was happening in the West Bank,
what was happening in Lebanon and Syria
and Sudan,
and even before that and
in all of the previous events, there was
a lot of research even if it was done
after the fact
of these genocides.
I think the work should focus from a
technical perspective also on amplifying
those through the technical knowledge
that we have.
Because imagine if like
like the change it would make to just
like when someone asks uh
any conversational model about something
and and and that model would would like
wouldn't have a very hard time accessing
what Hanin constructed.
Though I think that would make a very
huge difference.
So the key is the collaboration between
the two sectors, which which we heavily
lack, unfortunately,
in the region.
Please feel free to go on Dr. Jamila.
Yes, and I think you know, the
interesting aspect also is um
is both
exceptional but also not exceptional
nature of this episode. And how do we
trace that
kind of paradox, right? Um certainly
it's very hard to imagine something um
more bloody or horrific than the Nakba
in 1947 to 1950, which we usually just
periodize as 1948, which again is kind
of an interesting if you trace where
that idea comes from, that kind of
collapsing of a 3 to 4-year systemic
campaign of ethnic cleansing to a single
year.
It's kind of interesting
um
and that itself is a narrative structure
that we want to think about. What are
the episodes? How do we periodize them?
And how is it that the language model,
even in the Arabic region, Arabic
language news, and all this kind of
stuff,
kind of keeps imposing that
those that kind of periodization?
You know, another one would be that
you know, the invasion and occupation
and it always begins in 1982 in the
narratives about Lebanon because Beirut
is the only the center of the universe.
Love Beirut. It is the center of the
universe. But actually, it began in '78.
In the south, but the south has that
kind of narrative marginalization
as part of the efforts to And so, part
of the reason why I'm kind of saying all
that is I think there's a strangely
something hopeful and comforting to
people. In my experience, just in this
war, just with the younger people, I see
that kind of repetitiveness, right? It
kind of disrupts the shock and awe,
terrorizing psychological nature of
things. And I think that allows me to
also point out a little bit of some of
the complexities of collaboration.
For example, the way an archive science
with scientists would think about
these kinds of
collections of information, or records
we call them, versus perhaps from a big
data data analytics perspective. So,
we're working with the Safir newspaper,
one of the only newspapers that's
operated anyway. All the Arab Arab
people know what the Safir is, and they
have
digitized their newspapers. And so, if
you can see now that they're putting up
on
on social media
scans of newspapers about the
treacherous Lebanese government going
into talks with the Zionists,
which you could actually almost word for
word is the narrative right now
happening in Lebanon. And and the fact
that it's not exceptional, that it's
happened before, and we're right here,
and we're still on the land, and it
didn't work then, gives you all sorts of
different perspective, rather than the
way in which it's being trumped up as
this But the power of that, to some
degree, is that you see it as an actual
archives record.
Not just the data extract extricated
into big data models, and then mined
through AI. Part of that power
is that So, also thinking, as we're
trying to work through, for example, if
there's anyone Dr. Karim, you're getting
many emails from us, but if there's
anyone who'd like to help us think about
how we can take this massive terabytes
of social media we've archived with the
metadata of everything in context, and
turn it into something searchable and
navigable without extracting it into big
data out of context.
Yeah. You know, you can tell how excited
I'm getting just at the idea.
Yes, that's a very important
issue as well.
Keeping it in in real context, and uh
Yeah, I think
Turns out that we are um
not only um
have no access to the uh
standard media, at least within the West
at least. But we also have, you know, we
have the AI revolution almost against
us. It's uh in a sense. And we have the
resources that our resources are being
built.
But how do we get them to the masses
with this bottleneck, which is the uh
the uh the media, and the AI that we do
not own.
Uh for reconstructing the narratives, so
um certainly we should be ready for the
moment that we can disseminate
uh in in various ways.
Yeah, which is uh quite important,
reaching the
the the masses as well, such that they
know. Yes, go ahead uh Dr. Mustafa uh
Professor Mustafa Saray.
Yeah, hi everyone. Uh
Uh
sorry, I have a question uh to
Um do you hear me?
Um yes.
Ah, okay.
I'm talking from another laptop, so
>> [laughter]
>> Uh
my question is actually uh
to Dr. Jamila
about the archiveslabs.org.
So, I really didn't hear much about it.
What kind of content, not well, how much
is the content like uh related to Nakba
do you have?
And my other question maybe even to
uh to all panelists is about well, there
are many
uh Nakba archives,
but it seems nobody is synchronizing
with the other.
Nobody is talking to the other. So,
so we have a problem, fragmented content
here and there.
So, and as you said Dr. Khalil, the
uh
we have AI tools.
We have the power, but we are not using
it. So,
so basically Dr. Jamila, can you tell us
more now about maybe archives before we
finish the panel? I would like to hear
more about it.
The short answer is that the Archives
and Digital Media Lab is not a
collecting institution, and it's not an
archive in itself. But at the lab uh
we're one of the places that houses the
fighting in ratio project.
Which does collect both data and
archival records um, as part of fighting
in ratio as Hanan mentioned,
uh, but uh, most of uh, the archival
work I do for the project and that the
archives in digital media lab does is to
have worked with people who have key
collections on the ground in Palestine
and Lebanon
to preserve them and
protect them and in the context of their
systematic targeting and looting by
design. So, it's quite dis- different.
It fills this kind of gap. Uh, [snorts]
that said, there's a massive trove of
material uh, including oral history that
Hanan could speak to more that she's
been collecting but part of the work
that I did with Rada Damaj and a team of
people at AGML was uh, to collect about
16 TB of social media to archive social
media uh, focused on people who are uh,
victims or perpetrators of the
uh, what you can call the ongoing Nakba
or the genocide and its expansionism.
Um, and we had to stop unfortunately
about a year ago because of capacity
issues. Um, but we were trying very hard
to pick it up again. And um,
if there's anybody
who can help or anything. We're always
We're always begging for help. Um, but
the key point I would say for example is
we have uh,
a very like we have records or what I
call records but we have uh, social
media content that has already been
removed such as the kind that Rama you
talked about.
And if I may concerning the
Go ahead, Rama. I'm sorry for cutting
you off. Go ahead.
It's okay. I wanted to ask a question to
Hadeel and Zamila on a different topic,
so you can go ahead.
All right.
So so so so having multiple I mean like
there is actually a kind of strength in
numbers.
I mean like having multiple archives is
not bad in it in in that sense.
Also, the amount the sheer amount of
content that pushes a particular
narrative on the web
that would get picked up by Common Crawl
and then would feed into the large
language models is really really
important. There is something to be said
in that regard in the sense that
large language models need to see
something multiple times before it
actually registers and become part of
its memory.
If it sees it once, passes through once,
then it might not get picked up at all,
right? Uh this I I would I mean large
language larger language models need
less repetition, but at the end of the
day we have
competing narratives and basically
having multiple archives
uh you know, a plethora of articles, a
plethora of uh of of websites that that
of high quality that actually push the
same narrative is really really really
super important. Uh
I mean like if you do a search on on on
Google now,
uh you will find lots of content that
that will not be to your liking. And
somebody had gone through and made the
effort to uh you know, to just you know,
to have lots and lots and lots of
comment content by many many different
people feeding into the same narrative.
Yeah, this is certainly there is a whole
machine behind that.
Yeah, please feel free
uh Rama, please go ahead. You were
already before uh
I just one follow up before we go to the
next topic, just a quick note if
possible, Dr. Khalil, can I Can I just
>> go ahead. So, say what Karim said before
we
because
>> Oh, yes, yes, yes. Yes, certainly.
Yeah, regarding Common Crawl
which is used Common Crawl is like to
archive the web and a large language
models are trained on the uh Common
Crawl data. One thing I just found
we have in Palestine 3,200
domain names
that are active.
Only 190
yeah, 109 Sorry, 190
sites are included in the Common Crawl.
For example
website of the Birzeit University is not
included. Uh
Al-Ayyam news
uh newsletter paper is not included.
Al-Hayat newspaper is not included. So,
this is just to give you an idea that
the extent But but but but but this is a
fixable problem here, Dr. Yani, I mean,
let's let's let's actually talk offline
and I'll help you fix it today.
I mean, basically, we we we have
contacts with Common Crawl. I mean, one
of the things that at QCI we've done,
we've actually contributed more than
30,000 good URLs to to Common Crawl. And
if you have more, the better. I mean,
basically, we need to have our content
better represented
in what people use for training.
Great.
Yes.
Um yeah, Rama, you're now uh
your turn. Please go ahead.
Okay, I wanted to ask Hanin and Jamila
about um is there um like what does a
legally sound documentation of visual
content? Since there's a lot of visual
content focus especially in Hanan's
work.
Um like what what what would that look
like? Um
Um is there like an established um chain
of custody standard that you're
following and like archiving? And I'm
asking this in an attempt to
ask myself if it could be automated
because like if we're trying to
reconstruct narrative
um everything every source that we use
needs to be
to some extent legally sound for it to
be um acceptable widely. And I'm looking
for answers around that if there is ways
to automate a part of that. Just a note
of order. We have
very little time. We're over time
almost. I don't know what happens with
the room, but I'll just leave Hanan to
answer and I hope we can just close
after that.
I will give a quick answer. So, the
ethical question is really at the core
of our work, right? So, there's the
ethical question and there's the legal
question. Uh to start with the ethical,
the ethical is about also respecting the
victims, respecting uh the woman,
respecting the
you know, the men and women who have
been raped, the children. So, there's a
lot of ethical in there and into how do
we respect human dignity while at the
same time we want to show the violence
of this genocide. And this has been the
core of the conversations that we've had
not only as a team on this project, but
also with our uh team workers, right?
And then this bring us to the legal
question. So, that in order for us to
let's say document testimonies. So, so
far we have been able to document 10,000
uh
physical testimonies. That means that
our
field workers go to families. So, you
see to give you a quick example,
something concrete, um
Zarta Saha issues the Excel sheet, the
very the the fame Excel sheet with the
70,000 names in them of the victims of
the genocide, okay?
What that sheet Excel sheet has, it has
the name, gender, date of birth, and
date of death mostly. These are the the
data that are in there.
That is not sufficient for the work that
we want. We want first of all to honor
these victims, and we want them to be
more than just names in an Excel sheet,
right? So, we want to have
a big questionnaire that gives us a
bigger picture of who that person was.
So, we took that Excel sheet, and then
we
um
took the names of it, and then we traced
the families of those victims. But,
before doing that, we sat with a legal
team, and we asked them, "What are the
questions that we can put on a
questionnaire and ask the family members
to give us access to that is are legally
bound in the courts, in the ICJ, that
the South African team can also use as
evidence?" And so, it becomes more than
just a questionnaire to document
archive, but it's also becomes one for
accountability.
And so, eventually what the
questionnaire itself, which has about 20
questions, that goes beyond name,
gender, and date of
uh death, it has the name, the gender,
the profession,
uh the place, the kind of weapons that
were used, the location of where that
person was killed, was it an F-16
attack, was it a tank, was it artillery?
So, we go into the the very details, and
then the other step of legality is that
then we went to the Ministry of of
Health, and we asked them, "Can we have
their medical files?" And this is for
instance where the Ministry of Health
would would shut its doors and be like,
"We are not sharing that kind of
sensitive information with you.
If you want to preserve that
documentation and archive, part of why
we want that from them is that we are
afraid they will get bombed again and
all of that archive will be lost.
So, there's a lot of that, right?
There's a lot where you want to stick to
the legal and the ethical and there's
the place where you notice is an
unfolding ongoing genocide and even that
data that the Ministry of Health has
been struggling to scrap and put and
preserve after 2 years of genocide, we
know that it's a new target. So, we are
struggling as to what what can you
digitize? What can you scan without
losing? So, these are the the ethical
conversation that we keep having and we,
right? We try as much as we can to
um
to preserve. For instance, this morning,
this is just before I jumped into this
this call. Sorry, I'm I'm going beyond
time, but to give you the kind of
obstacle we face.
So, this is before I got into this call,
we finally have the all of the pictures
of all of the medical staff that were
killed.
And many of these pictures are of the
doctors.
Some of the Some we were able to find
their normal, right? The picture before
that or picture on the medical card, so
it's a random picture, but many of of
the other pictures, especially when the
fam all of the family is killed,
is a picture of the doctor laying in
blood with a bullet between his eyes or
behind his his back, right? That a
doctor took in his cell phone, so we got
access to that. And the field worker
received the permission of of the doctor
to take that picture from him. So, she
has a copy and she sent she sent the
entire file to me.
And so, I'm not going to put the picture
of that doctor with a bullet behind his
head online as part of his medical
profile on the website.
So, right. So, these are the kind of
day-to-day things that that it is Yes,
it is an archive, and yes, I want this
to remain because I want history and the
world to know what has been done to us.
But, how do we deal with all of this,
especially that this is the first time
we have an an a genocide that has such
vital images for accountability in the
future. I see Palestine free, and I see
us, you know, putting these people on
trial, and I see myself showing them
those pictures someday. I see that
happening. So, I need to preserve that,
too. And the conversation today is for
us to understand how do we do this
ethically, respecting our people, and
respecting a legal space, right? That
that that has to protect and and hold
accountability. So, it's a big answer,
but it shows all of the things that we
deal with as we preserve.
Yeah. Thank you. Thank you very much.
Very important.
And um yeah, I think we uh
I just wanted to to thank the uh
panelists, all of them,
and um
say hopefully we will be able indeed to
do what
you said um Hanin, and hopefully we can
have these archives.
We had one archive that was stolen in
Beirut. Uh the PLO archive held that
archive, and it was stolen in 1982.
We are starting again
um
with
little in hands to start with, but
um there is already a lot accumulating
thanks to the Israelis, obviously.
Uh all right. So, um
I leave it now to uh Dr. Mustafa to
close the uh workshop. Thank you all.
Thank you everybody.