Submind YouTube summaries
Thumbnail for Panel Discussion : NLP: Reconstructing narratives about the ongoing Nakba #Arabic #AI #nlp

Panel Discussion : NLP: Reconstructing narratives about the ongoing Nakba #Arabic #AI #nlp

Watch on YouTube

Video summary

This panel discussion centers on the critical application of Natural Language Processing (NLP) and digital archiving to reconstruct narratives surrounding the ongoing Nakba, addressing a significant technological gap between Arabic dialects and established English or Hebrew models. Experts highlight that while technical progress in Arabic NLP is evident, high error rates in speech recognition for non-Modern Standard Arabic varieties persist, necessitating specialized systems capable of processing diverse linguistic inputs including harmful content from Israeli channels. The conversation underscores the danger of relying on biased or censored data found on major platforms like Meta, which often remove hashtags only after sustained pressure; consequently, there is an urgent need to build independent in-house systems that can document suppressed political speech and measure visibility asymmetries to hold tech giants accountable for their role in erasing Palestinian history. To combat this systematic erasure, the panel introduces collaborative initiatives such as the "Fighting Erasure" project, which functions not merely as a traditional archive but as a chronological platform documenting genocide in real-time through legal definitions like domicide. This approach involves field workers verifying attacks on hospitals and neighborhoods using Geographic Information Systems (GIS) and historical maps to counter settler-colonial logic that fragments narratives along nation-state boundaries. Unlike post-event archiving, these historians are capturing the "live stream" of current atrocities, aiming to preserve terabytes of testimonies before they vanish due to big tech censorship or physical destruction by death. The process extends beyond simple demographics to include forensic details such as specific weapon types used in attacks and precise locations, creating a robust evidentiary base for future legal accountability at institutions like the International Court of Justice. The discussion also delves into profound ethical dilemmas inherent in preserving sensitive data, particularly regarding medical files from the Ministry of Health that officials initially refuse to share due to fears of their destruction during bombings or targeted attacks by Israel. While some graphic images provided by field workers depict doctors killed and bodies riddled with bullets taken directly from cell phones after families were wiped out, experts face a difficult choice between publishing such visceral evidence for future trials and respecting the dignity of the deceased and living victims. Despite these challenges and past losses like the theft of PLO archives in Beirut in 1982, there remains an unwavering commitment to digitize records before they are lost, ensuring that visual content maintains a legally sound chain of custody while honoring the memory of those targeted by violence. Ultimately, the panel concludes that successful narrative reconstruction requires seamless collaboration between technical experts and historians to navigate complex challenges such as integrating Palestinian domains into global web crawlers like Common Crawl and automating ethical documentation processes. By combining advanced NLP capabilities with rigorous historical methodology, these efforts aim to maintain context within massive datasets while fighting fragmentation across scattered archives. The consensus is that preserving this vital evidence is essential not only for honoring victims but also for establishing the factual basis needed in future international courts, ensuring that history is recorded accurately despite ongoing threats and attempts at digital or physical erasure by hostile actors.
Read the full video transcript
Okay, so great. Um welcome to this panel discussion. The title is NLP reconstructing narratives about the ongoing Nakba. The intention here is to reflect um on where we stand now and how we can move forward uh with the aim of reconstructing the narratives about the Nakba, the ongoing Nakba. And the ongoing Nakba is a term that signifies an approximately 100 years of um Nakba, so to speak. Um I'll introduce first the um panelists. We have four panelists with us. Uh Dr. Kareem Darwish, um who is a principal scientist at the Qatar Computing Research Institute, so QCRI. Um he is um a long um I know Kareem for a long time now at the um conferences of NLP conferences. His work is has been mostly on state-of-the-art tools for Arabic processing and social computing, um which has been well received. We have with us also Mrs. Rama Salha, um excuse me if I'm not pronouncing correctly because I didn't see it in Arabic. Um she is an AI specialist and technology officer at Hamleh. Uh her work focuses on building decolonial AI, uh auditing content moderation systems, and challenging power structures embedded in technology. We also have with us Dr. Jamila Haddar. She's an assistant professor in archival information and digital humanities at the University of Amsterdam. Um she's also a co-founding director of the Archives and digital media lab and research affiliate at the American University of Beirut School of Architecture and Design. She is a co-founder of the fighting eraser digitizing Gaza's genocide a project co-founded together with Dr. Hanin Shahada who our fourth panelist. She is an assistant professor at New York University Abu Dhabi and research associate at the Maroun Semaan Faculty of Engineering and Architecture at the American University of Beirut in Lebanon. Welcome all. What we'll do here I suggest at least that I start with one of you basically start with Dr. Karim with a question. Please feel free to respond to the others. At some point I hope we will have free discussion. That's the intention. Okay, I start with Dr. Karim first with the first question. So as a leading expert in Arabic NLP where do you think we stand at this moment with Arabic NLP given that we don't have our own AI as far as I know and correct me if I'm wrong. And how can NLP help reconstruct the narrative about the Nakba and what resources do we need to do so? Okay, so somebody can first of all I'm very happy for the invitation. Thank you Dr. Khalil. It was a pleasure to to be with you. So you ask actually a quite large question and there are two aspects of this. One of them is technical and the other part is is operational. On the technical side I think when we deal with the reconstructing the Nakba and ongoing Nakba. There are actually two parts. The first part has to do with the Arabic and the other part has to do with other languages other than Arabic including Hebrew and English and other languages. On the English side, I think on when even in most European you know languages the the landscape is actually quite mature. On the Arabic side the maturity level has been increasing dramatically over the past few years. And this is evident by the large community of Arabic NLP that I mean like the mailing list has over a thousand people. There are many efforts in the region the Arab region for for building sovereign AI with models like Alam in Saudi Arabia, Fanar in Qatar, Jason in the United Arab Emirates, Karmak in Egypt and so forth. Like if you look at other technologies like um speech recognition and and OCR, they have been steadily improving over the past few years. They're not at par with English yet, but I think we I think in the next probably two or three years we'll we'll get to that point. So the the missing I think piece is where the Hebrew part is. I think Hebrew is is is a lot more under you know developed compared to what we have for Arabic and for definitely compared to to English. And I think reconstructing the narrative requires that we actually do all the languages well. It's it's on one side you want to you know I think other panels will talk about you know that this is archiving of the content that we are receiving from from Gaza or from Lebanon and so forth. A lot of it is coming in Arabic, but we should not be neglecting what Channel 14 and Channel 12 are spewing all the time documenting you know the the their narrative documenting you know you know, uh the the way that they see the world, how the the they cheer on the the genocide that's that's actually ongoing. Also, uh Hebrew social media that for for longest time people thought that, you know, they're in isolation and they don't actually see people the world doesn't see what they're what they're saying and what they're doing. That needs to be, you know, resurfaced and that that would require uh not just data collection, but also uh competent speech recognition, competent ASR OCR, uh and competent language modeling for for for for that language for for these languages. So, in that sense, uh I think Arab English is and European languages are doing very well. Arabic is not far behind and, you know, Hebrew is is way behind and that needs to to be kind of fixed. How do How do we stand on dialects? Just to uh before I move to the other That's a great question. So, Okay, so if you compare Arabic, you know, MSA modern standard Arabic and and dialects, particularly in the area of of speech recognition, I think you speak you know, speech recognition, for example, is behind. Uh in the sense that you can get below 10% word error rate for MSA and the typical word error rate for for for dialects is around in the range of 80 of about 20% or so. Uh this gap needs to be filled uh particularly for you know, for for more niche dialects or or pronunciations. Uh so, this is on the speech side. Uh if you're looking at the the text side, I think the gap is is is much narrower than than than speech. You'll find that most large LLMs, including Fana and I'm pretty sure I'll have and others, uh they wouldn't have any problem understanding dialectal text, making sense of it. The mapping between the concepts that happens, you know, implicitly inside of the model are actually quite robust. Thank you. That's a great It gets a great start. We talked about the technical stuff, so I'll I'll move now to Rama, Mrs. Rama Salah. Uh So, from your experience and work at Hamleh, maybe you can tell us a bit about it because that's in Palestine. And maybe you can tell us also how you think maybe uh or where can uh reconstructing the narratives about the Nakba can be used best used NLP in the best way. Uh yes. Hi. Thank you for the invitation to join this panel. I'm glad to be a part of this um very important discussion. Um so, I want to start my um speech with a framing point that today's biased and censored data becomes tomorrow's historical record in AI systems. What gets removed or down ranked or um made invisible becomes a structural absence that future language models will then treat as natural. And on the other hand, what what is allowed to circulate and be amplified becomes the normal distribution that models learn from. And from this perspective, platform governance and moderation decisions extend from shaping the current public discourse to shaping the data sets that future AI systems learn from. Thus, what becomes, let's say, like narratively recoverable at scale. Through our work at Hamleh, we explored this in several structural ways. Um in one of Hamleh's publications, Silent Networks, um the research shows that a a very strong chilling effect around political expression where surveillance, interrogation, and fear of consequences reduce online participation, especially amongst the youth. And from an NLP perspective, this is very important because it means that the data generation itself from the beginning is being structurally suppressed at the source, making the absence of narrative or speech produced and forced rather than just organic. On the other hand, in our violence indicated work, specifically the one um monitoring hateful content online through in-house models that we have trained, where we study how um narratives of violence uh circulate at scale and how platform governance responds unevenly over time. One clear example was the hashtag and phrase um flatten Gaza, or in Hebrew it's pronounced lem ho ket Gaza, something like that, which circulated widely across across platforms beginning October 2023, and then remained like highly visible for months, accumulating tens of thousands of uh occurrences before being made unsearchable by Meta in mid 2024 after lots and lots of continuous external pressure from partners and big-scale data-backed documentation. So, what this highlights is the importance of building in-house technical systems, like uh Dr. Karim said, uh that are representative of our narratives. Because this kind of technical work, when you combine it uh with the advocacy and um external pressure, it can sometimes lead to changes in platform behavior, even if only partially, by making visibility asymmetries measurable and harder harder to ignore. What this also shows is that um we cannot really separate narrative reconstruction from the systems that determine what data gets to exist and be easily and cheaply and cheaply accessible uh in the first place. Um and thirdly, through her our observatory for digital rights violations, we maintain a structured data set of nearly 15,000 documented cases of um content removal, suppressed political speech, and hatred and violence online since 2021. That data is vetted by humans and like we collected in order to try to make action through these platforms where this content um existed. Which functions as a counted corpus um measuring what is missing or unjustly removed from mainstream data sets. Which is I think uh it's very essential for understanding what to look for in data when trying to analyze and understand bias in AI models. So, at this point um what I want to leave this panel with is that narrative reconstruction extends from documentation and analysis to also become a question of data set construction under conditions of asymmetric visibility and platform governance. And that unless we explicitly account for how platforms determine what remains visible and what is removed, future language models will inherit structurally these missing histories. This is very good. I mean, um what I just to uh summarize a bit what he you know the the two of the in my opinion at least most important issues that you raised is that uh there is an under representation plus suppression of the uh narrative if let me call it this way in in AI but also in the uh you know in general in um on the media. Uh and what this combined with what Dr. Karim said about the Arabic dialects since most of the for example the videos are in Arabic dialects, this means that it is rather complicated to process and also to bring together. I'll give Apparently, Dr. Karim has a response. I'll give it to him and then I move afterwards to Dr. Jamila. Karim If I may, I mean just to second on what Rama said. This suppression and is actually we can actually see it in the current LLMs. We don't actually have to wait until the future. I would ask you to go on the you know like on Gemini and ChatGPT and ask both of them about the Haganah gangs which were deemed terrorist organizations and look at the answer that they give back. The answer is really really really shocking. I mean if you look at it you'll think that these are peace-loving, you know, tree-hugging, you know, angels who came to save people while everybody deemed them as terrorist organizations. Yeah. This is all very important because I think if I move now to Dr. Jamila Haddad I mentioned the project that you're leading together with Dr. Hanin Shadeed. Could you start with telling us about that and then respond to whatever you think is uh is important to to you in in this discussion so far. Thank you very much and thank you for having me. Everything a lot of what's been said certainly resonates with me. We at the Fighting Erasure project the long title Fighting Erasure Digitizing Genocide on the war on Lebanon um you know, has it has various components. Other key part that I work on is on the archives and record side and we're really struggling to understand how and what it looks like uh to rescue and recover and safeguard archives and records uh particularly those uh in my area of the project uh that pertain to that have this kind of evidentiary quality uh whether evidence of the genocide itself uh so for example this uh epic and rather arguably impossible task we took on of archiving social media um in a in a particular kind of archival science way and um alternatively uh what it could look like to mass digitize and um rescue and recover from under the rubble. But in terms of the evidentiary it has to do with people's rights and property and administration and governance and all the kinds of things that connect and prove and assert both the history and present of uh the people on the land because of course that's the settler colonial logic uh to try to eliminate any and all proof of that. And you know like everybody in various ways we are struggling with digital infrastructures and information infrastructures that are all that that are all connected to these really problematic platforms. Um we're trying to understand and have conversations around what does a digital sovereignty and data sovereignty look like uh both at the infrastructural and theoretical level uh war zones and conflict areas and areas perhaps where we may not uh feel all that much more comfortable in terms of say the Gulf context which is really the Gulf states which are really leading a lot of this cuz it takes all this money. Um you know there's a lot of pieces uh to what I'm saying. Uh one thing is I can certainly as a as an archivist and a historian of archiving and um the last 500 years uh certainly uh echo what you said Rama that and Karim that the language models that AI is built on already is absolutely colonialist and racist and certainly orientalist and the depth to which that goes like 500 years. I mean you can actually it's always problematic to find origins. We can go back even more of course. Um but the interesting part is how actually in some regards those are being disrupted because of the mass the expression on a mass level of sympathy and various different perspectives and just the evidence and data on the ground and so again that kind of really interesting way and now we're trying to see how they're trying to actually modify the AI to kind of be dynamic in this way as this different kind of data and perspective is coming in. Um there is probably again way too much I I want to say but maybe to speak the most directly to your question is that for us to reconstruct narratives um it's worthwhile to kind of map out what is the narrative structure of genocide and settler colonialism that they're trying to impose and certainly a key to that is a kind of fragmentation and dispersal and segregation not only of people and geographies but also of the narrative. So we can only ever say this is happening the West Bank. You know we used to uh you you know we used to talk about Arab region right now it's just by nation state now it's even as we can all see the profoundly like cross-border and regional dynamic nature of the ongoing like but which has always been the case but self turkey logics and narratives uh want to completely destroy that. So it what what kind of role can we actually have AI play and and kind of creating a better sense especially when we're kind of um stuck with the way in which data is around nation states or certain identity groups that can actually uh obscure all of that and of course quite intentionally if you look at the history of these things. So I'll stop there. Thank you very much. Thanks. Thanks a lot. So this enforces what the previous two uh uh speakers have said and also add to it a dimension of you know the um uh risk of the data being uh exposed to uh uh also to the to AI on the one hand and we want that and on the other hand to have the AI to reflect the balance uh in the narratives that exist in the world and which is at the moment completely unavailable or not the case. I I'll move to Dr. Hanin Shahada. I'm sure she has to add a lot on this. So I've heard her before ask a question we suggested direction. I'm quite curious Hanin to what you have to add and at the same time also please reflect on what the others have said and feel free to add. Sure. I mean thank you all for having me here and uh um it has been such an interesting conference for me especially given that I'm not a technical person, right? I'm a historian and so looking at the struggles from a from an operational perspective is something that I recognize as we're trying to narrate and document and archive and preserve this this genocide. And so, I thought that maybe I would show you what we have been doing in a way to uh preserve, right? We're archivists, we're historians, we're trying to preserve whatever there is here. What I've heard today and I've seen repeatedly is that many of us separately uh across the planet have been preserving data and collecting it on several several different hardwares, right? And trying to keep whatever uh data we can collect in the in the tetrabyte and then look at how we can, right, in the future do something with it. And that's also part of the narrative today because the way that we see it is that, first of all, this is the first live stream genocide on Earth. So, there's a new way of of perceiving what's happening in a live stream. Historically, when you archive where you when you write history, it's at the end of the moment. It's at the end of the genocide. So, when we documented Bosnia, it was when it ended, right? When we wrote about the Holocaust, it was after it happened. The same thing with the Russian words is that you go back and you try to recollect what's happening afterwards and to try to put a narrative. Gaza shifted the ground on us, right? So, it's it's it was live streaming as it was happening. And the same people that were archiving this ended up being killed, which means that you have a testimony without, right, in the moment of death. For me is what do we do with all of this? And it took us about um I think 6 to 7 months to sit with colleagues, and we came up with something that we called the skeleton, which is a massive website or a digital platform that we wanted to create that would reconstruct chronologically somehow the genocide with the data that we have and with a local team in Gaza. So, we have 10 field workers part of our structure, our office in the Gaza Strip that have this data. We have this triangulation of verification, which means that let's say we want to document the attack on the staff of Shifa. We have what's online, right? But we don't know is this a picture of Shifa or is it Madani or is this that? Where is this? So, they take that content, they go to the Ministry of Health, they go to witnesses, they go they they record testimonies. So, we have an entire block to recreate that moment and we have been doing this since October 10th. It's massive, massive and it's beyond our capacities. Just in terms of time, I just wanted to show if possible just share my screen to give you an idea of of what what in the future this this project is meant to look like. So, I think you have the right to share, but I'm not sure completely. Yep. So, the idea of this big platform is that it will have the main components. So, we we have spoken to legal scholars that have defined for us first what is genocide, right? And what falls under genocide. And under these legalities, we constructed every single one of those elements. So, you can see on the website genocide, domicide, herbicide. I mean, by now all of you are very familiar with those terms. The attacks on the hospital. And the idea is that if you click on each one of them, it will pop up and open up into this massive uh uh second website that will provide you the numbers, the the statistics, the documentation, the the doctors, the hospitals. I'm scrolling. I don't know if you see it with me. But, let's say we'll give you an idea that before the war, Gaza's health network had 13 governmental hospitals, 18 UNRWA facilities, five private hospitals. You click on each one of these, it pops up into a new window. And the idea is that, you know, all of those terabytes eventually that we have that they will be condensed into one space, one massive online archive that has been that will reconstruct this entire genocide precisely to fight erasure. And this is hence the name of our project, Fighting Erasure, is that meta is deleting, right? Uh big tech is removing. We have our people in Gaza, they are documenting and safeguarding. And it's not a work that comes without without So, this is to give you an idea. So, you have, for instance, the chronology of the attacks on the hospitals. You see on October 7th, 2023, it's the beginning of the genocide. Uh we see that already the initially they had a health care facilities were targeted. You click on it, you get all of the details. You go to October 15th, 2023. This is when the siege of Al-Shifa Hospital begins. You click on all of these tabs, they take you to other spaces. And then, right, you go down, you have Kamal Adwan has been targeted, Nasser Hospital Nasser on January 22, '24. And and and etc. And the idea is that we built this to do the same thing with, let's say, if you go to Herbicide, with the neighborhoods, is that now we're trying to um look into one neighborhood after the others. And we we have, because it's so massive, we thought that we would start with the north of Gaza that has been massively decomposed. So, with neighborhoods like Sabra, Tal al-Hawa, Sheikh Radwan, Jabalia, Muaskar, Mukhaiyam al-Shati. And and then look at the old maps that the municipality had for those places. So, our field workers are going to the municipality, see if they have the old urban planning of these areas. And then with GPS, sorry, with with GIS try to see what has remained, which units are gone, etc. So, you can tell it's a massive project. Um but we're trying by doing this is to undo this this or to fight erasure. And whatever data we have now and locally and and bless bless all the people of Gaza who have given us all of that content. I hate to call it content, but they're testimonies to this genocide. And I think part of our job as historians is to fulfill that that promise to them that what they have sacrificed their lives, we will not leave it there to right to to not be seen. Um so, these are This is the big project with the aim of where we want to go. You can understand that we also face some structural obstacles. Some of them are operational. This is what I was talking with Dr. Alexei about. The idea is I'm putting the burden now on Dr. Jamila because she's the scientist archivist, but she has been tasked with creating a system to catalog and categorize and tag and filter and authenticate all those mega terabytes of of data. So, I can take it from her and then just put them in each one of these sections. Ideally that's what I want and she told me this morning, "Give me 6 months." So >> [laughter] >> So I but the idea is that right is that we have now each one of us loosely has that kind of data. I am not a scientist. I cannot But once I have that information, today we have the platform that can absorb that that sort of all of that material once gets cataloged. So that's for my presentation. I tried to stick to the time as much as possible. Thank you. Thank you. Thank you I mean so I'll move to Dr. Karim as hand go on, please. Uh so basically thank you Dr. Karim and Dr. Jamila for for your great effort. I guess one of the things that that I think we need to find a way that we can kind of speak the same language in the sense of we have lots of technology that potentially could be of use in that in that regard. Uh transcription, data analysis, data extra you know information extraction from from text, uh you know even image classification and and so on and so forth. So so the tools that at our disposal whether it's in English or or Arabic and that in this in this regard I I think is is quite is quite strong. But you know um we don't see what you what you see. You see I mean like you can be our eyes in developing the technology for to your benefit if we can find a way that we can speak kind of the same language and you say, "This is the kind of data that I have and this is where where I want to go." And perhaps we can build you or help you build that bridge. This is great. I mean this all sounds like a project proposal. >> And and be careful cuz we will take you on it, right? I will be in your inbox, so >> [laughter] >> So, that's that's exactly I know we have so much potential in the Arab world. It's not a lack of it. I think at this point is really a way of of finding a way to set up this network. We all want the same thing. Um and uh I don't I don't speak technical, right? But I we have like I speak legal, history, archiving, and so we have created this skeleton with with legal scholars, with historians, with name it. And now I'm look we're looking for the for for that what you just said, Dr. Karim, is to sort of find a way to put all of these systems together just to build in and find a way to put that content uh there and to safeguard it as well, right? There's also the thing that you don't want this this website to fall apart or to be deleted or to be So, you need also to protect the the domain that we're on. So, there are many conversations that are happening at the same time. Oh, great. So, this is a very good I mean it's very inspiring. And I think uh there are ways to collaborate apparently. So, I hope this matures at some point. Um I was um thinking um at the same time um and I'd like to open the floor for others to uh join us also from the audience. But if others want to raise other points, please feel free. But at the same time I was thinking we are at at um the probably the the most um horrible episode of the ongoing Nakba so far. But it is an episode nonetheless. And there has been there have been multiple episodes before that. And where technology, particularly NLP, can be helpful is also to see how to connect the episodes together for historians. And whether that is similarities, but also dissimilarities. And I know we have a lot of technical um uh difficulties, but at the same time there are uh tools that already work, uh so to speak. So, I'm open that for everybody to think about and maybe discuss. And if there are others from the audience, they're welcome to join us. So, how does the current episode relate to the previous episodes of uh uh And and would it make sense what I just said about bringing the episodes together in one way or another? I actually if I if I may, I think this episode obviously is a lot more bloody compared to the previous ones. Uh but but uh you know, the the cookbook that they're using is is very very similar, right? Uh I mean, now they say, "Oh, you know, Hamas is a terrorist organization and thereby justifies everything, right?" You go to Lebanon, you just replace the word Hamas with some other entity like Hezbollah, and then that's the excuse. Before that, it was the PLO. Exactly. And and I mean, like the the narrative doesn't change, right? And and the the Hasbara efforts have been, you know, quite consistent in the same narrative, the the same way, and so forth. Um given you know, like as as humans, we can we can you know, like we we see it and we like recognize it right away. Uh whether the AI can can recognize it in the same way, I I there there must be a way. Uh but but uh the end result of how, you know, like if you have like a structure and you say, here are the inputs and here are the the predictable outputs, I think historians would probably need to tell us what kind of outcome that would they would they be interested in seeing, right? Because because of the similarity, I mean, it's the same narrative almost every time. You just change the names and the locations and it's the same. Indeed. Yeah. Go ahead, Rama, please. I couldn't agree more with Dr. Karim. And I think there's um like this time is a very special time with an opportunity um to make really big steps. Like right now with um AI agents boom, where um like most AI models are based on algorithms and the agents um and automated uh collection of the data to get answers. I think there's a lot of work to be done and um better designing the platforms where historians get to um document the reality. Uh for example, to make it like bot-friendly. Um and like um and helping and making these platforms um have better rankings. Thus, um um look uh like more authentic and more um What we What can we call it? More believable in the eyes of a bot. So, like if we work in that direction, um like the the tech where the tech sphere comes uh hand in hand with the historians and social sciences um um um researchers and political science researchers, I think there's a lot of space um to push the work because there's already a lot of research being done by amazing people documenting so much of what was happening um since October 2023 in Gaza, Um what was happening in the West Bank, what was happening in Lebanon and Syria and Sudan, and even before that and in all of the previous events, there was a lot of research even if it was done after the fact of these genocides. I think the work should focus from a technical perspective also on amplifying those through the technical knowledge that we have. Because imagine if like like the change it would make to just like when someone asks uh any conversational model about something and and and that model would would like wouldn't have a very hard time accessing what Hanin constructed. Though I think that would make a very huge difference. So the key is the collaboration between the two sectors, which which we heavily lack, unfortunately, in the region. Please feel free to go on Dr. Jamila. Yes, and I think you know, the interesting aspect also is um is both exceptional but also not exceptional nature of this episode. And how do we trace that kind of paradox, right? Um certainly it's very hard to imagine something um more bloody or horrific than the Nakba in 1947 to 1950, which we usually just periodize as 1948, which again is kind of an interesting if you trace where that idea comes from, that kind of collapsing of a 3 to 4-year systemic campaign of ethnic cleansing to a single year. It's kind of interesting um and that itself is a narrative structure that we want to think about. What are the episodes? How do we periodize them? And how is it that the language model, even in the Arabic region, Arabic language news, and all this kind of stuff, kind of keeps imposing that those that kind of periodization? You know, another one would be that you know, the invasion and occupation and it always begins in 1982 in the narratives about Lebanon because Beirut is the only the center of the universe. Love Beirut. It is the center of the universe. But actually, it began in '78. In the south, but the south has that kind of narrative marginalization as part of the efforts to And so, part of the reason why I'm kind of saying all that is I think there's a strangely something hopeful and comforting to people. In my experience, just in this war, just with the younger people, I see that kind of repetitiveness, right? It kind of disrupts the shock and awe, terrorizing psychological nature of things. And I think that allows me to also point out a little bit of some of the complexities of collaboration. For example, the way an archive science with scientists would think about these kinds of collections of information, or records we call them, versus perhaps from a big data data analytics perspective. So, we're working with the Safir newspaper, one of the only newspapers that's operated anyway. All the Arab Arab people know what the Safir is, and they have digitized their newspapers. And so, if you can see now that they're putting up on on social media scans of newspapers about the treacherous Lebanese government going into talks with the Zionists, which you could actually almost word for word is the narrative right now happening in Lebanon. And and the fact that it's not exceptional, that it's happened before, and we're right here, and we're still on the land, and it didn't work then, gives you all sorts of different perspective, rather than the way in which it's being trumped up as this But the power of that, to some degree, is that you see it as an actual archives record. Not just the data extract extricated into big data models, and then mined through AI. Part of that power is that So, also thinking, as we're trying to work through, for example, if there's anyone Dr. Karim, you're getting many emails from us, but if there's anyone who'd like to help us think about how we can take this massive terabytes of social media we've archived with the metadata of everything in context, and turn it into something searchable and navigable without extracting it into big data out of context. Yeah. You know, you can tell how excited I'm getting just at the idea. Yes, that's a very important issue as well. Keeping it in in real context, and uh Yeah, I think Turns out that we are um not only um have no access to the uh standard media, at least within the West at least. But we also have, you know, we have the AI revolution almost against us. It's uh in a sense. And we have the resources that our resources are being built. But how do we get them to the masses with this bottleneck, which is the uh the uh the media, and the AI that we do not own. Uh for reconstructing the narratives, so um certainly we should be ready for the moment that we can disseminate uh in in various ways. Yeah, which is uh quite important, reaching the the the masses as well, such that they know. Yes, go ahead uh Dr. Mustafa uh Professor Mustafa Saray. Yeah, hi everyone. Uh Uh sorry, I have a question uh to Um do you hear me? Um yes. Ah, okay. I'm talking from another laptop, so >> [laughter] >> Uh my question is actually uh to Dr. Jamila about the archiveslabs.org. So, I really didn't hear much about it. What kind of content, not well, how much is the content like uh related to Nakba do you have? And my other question maybe even to uh to all panelists is about well, there are many uh Nakba archives, but it seems nobody is synchronizing with the other. Nobody is talking to the other. So, so we have a problem, fragmented content here and there. So, and as you said Dr. Khalil, the uh we have AI tools. We have the power, but we are not using it. So, so basically Dr. Jamila, can you tell us more now about maybe archives before we finish the panel? I would like to hear more about it. The short answer is that the Archives and Digital Media Lab is not a collecting institution, and it's not an archive in itself. But at the lab uh we're one of the places that houses the fighting in ratio project. Which does collect both data and archival records um, as part of fighting in ratio as Hanan mentioned, uh, but uh, most of uh, the archival work I do for the project and that the archives in digital media lab does is to have worked with people who have key collections on the ground in Palestine and Lebanon to preserve them and protect them and in the context of their systematic targeting and looting by design. So, it's quite dis- different. It fills this kind of gap. Uh, [snorts] that said, there's a massive trove of material uh, including oral history that Hanan could speak to more that she's been collecting but part of the work that I did with Rada Damaj and a team of people at AGML was uh, to collect about 16 TB of social media to archive social media uh, focused on people who are uh, victims or perpetrators of the uh, what you can call the ongoing Nakba or the genocide and its expansionism. Um, and we had to stop unfortunately about a year ago because of capacity issues. Um, but we were trying very hard to pick it up again. And um, if there's anybody who can help or anything. We're always We're always begging for help. Um, but the key point I would say for example is we have uh, a very like we have records or what I call records but we have uh, social media content that has already been removed such as the kind that Rama you talked about. And if I may concerning the Go ahead, Rama. I'm sorry for cutting you off. Go ahead. It's okay. I wanted to ask a question to Hadeel and Zamila on a different topic, so you can go ahead. All right. So so so so having multiple I mean like there is actually a kind of strength in numbers. I mean like having multiple archives is not bad in it in in that sense. Also, the amount the sheer amount of content that pushes a particular narrative on the web that would get picked up by Common Crawl and then would feed into the large language models is really really important. There is something to be said in that regard in the sense that large language models need to see something multiple times before it actually registers and become part of its memory. If it sees it once, passes through once, then it might not get picked up at all, right? Uh this I I would I mean large language larger language models need less repetition, but at the end of the day we have competing narratives and basically having multiple archives uh you know, a plethora of articles, a plethora of uh of of websites that that of high quality that actually push the same narrative is really really really super important. Uh I mean like if you do a search on on on Google now, uh you will find lots of content that that will not be to your liking. And somebody had gone through and made the effort to uh you know, to just you know, to have lots and lots and lots of comment content by many many different people feeding into the same narrative. Yeah, this is certainly there is a whole machine behind that. Yeah, please feel free uh Rama, please go ahead. You were already before uh I just one follow up before we go to the next topic, just a quick note if possible, Dr. Khalil, can I Can I just >> go ahead. So, say what Karim said before we because >> Oh, yes, yes, yes. Yes, certainly. Yeah, regarding Common Crawl which is used Common Crawl is like to archive the web and a large language models are trained on the uh Common Crawl data. One thing I just found we have in Palestine 3,200 domain names that are active. Only 190 yeah, 109 Sorry, 190 sites are included in the Common Crawl. For example website of the Birzeit University is not included. Uh Al-Ayyam news uh newsletter paper is not included. Al-Hayat newspaper is not included. So, this is just to give you an idea that the extent But but but but but this is a fixable problem here, Dr. Yani, I mean, let's let's let's actually talk offline and I'll help you fix it today. I mean, basically, we we we have contacts with Common Crawl. I mean, one of the things that at QCI we've done, we've actually contributed more than 30,000 good URLs to to Common Crawl. And if you have more, the better. I mean, basically, we need to have our content better represented in what people use for training. Great. Yes. Um yeah, Rama, you're now uh your turn. Please go ahead. Okay, I wanted to ask Hanin and Jamila about um is there um like what does a legally sound documentation of visual content? Since there's a lot of visual content focus especially in Hanan's work. Um like what what what would that look like? Um Um is there like an established um chain of custody standard that you're following and like archiving? And I'm asking this in an attempt to ask myself if it could be automated because like if we're trying to reconstruct narrative um everything every source that we use needs to be to some extent legally sound for it to be um acceptable widely. And I'm looking for answers around that if there is ways to automate a part of that. Just a note of order. We have very little time. We're over time almost. I don't know what happens with the room, but I'll just leave Hanan to answer and I hope we can just close after that. I will give a quick answer. So, the ethical question is really at the core of our work, right? So, there's the ethical question and there's the legal question. Uh to start with the ethical, the ethical is about also respecting the victims, respecting uh the woman, respecting the you know, the men and women who have been raped, the children. So, there's a lot of ethical in there and into how do we respect human dignity while at the same time we want to show the violence of this genocide. And this has been the core of the conversations that we've had not only as a team on this project, but also with our uh team workers, right? And then this bring us to the legal question. So, that in order for us to let's say document testimonies. So, so far we have been able to document 10,000 uh physical testimonies. That means that our field workers go to families. So, you see to give you a quick example, something concrete, um Zarta Saha issues the Excel sheet, the very the the fame Excel sheet with the 70,000 names in them of the victims of the genocide, okay? What that sheet Excel sheet has, it has the name, gender, date of birth, and date of death mostly. These are the the data that are in there. That is not sufficient for the work that we want. We want first of all to honor these victims, and we want them to be more than just names in an Excel sheet, right? So, we want to have a big questionnaire that gives us a bigger picture of who that person was. So, we took that Excel sheet, and then we um took the names of it, and then we traced the families of those victims. But, before doing that, we sat with a legal team, and we asked them, "What are the questions that we can put on a questionnaire and ask the family members to give us access to that is are legally bound in the courts, in the ICJ, that the South African team can also use as evidence?" And so, it becomes more than just a questionnaire to document archive, but it's also becomes one for accountability. And so, eventually what the questionnaire itself, which has about 20 questions, that goes beyond name, gender, and date of uh death, it has the name, the gender, the profession, uh the place, the kind of weapons that were used, the location of where that person was killed, was it an F-16 attack, was it a tank, was it artillery? So, we go into the the very details, and then the other step of legality is that then we went to the Ministry of of Health, and we asked them, "Can we have their medical files?" And this is for instance where the Ministry of Health would would shut its doors and be like, "We are not sharing that kind of sensitive information with you. If you want to preserve that documentation and archive, part of why we want that from them is that we are afraid they will get bombed again and all of that archive will be lost. So, there's a lot of that, right? There's a lot where you want to stick to the legal and the ethical and there's the place where you notice is an unfolding ongoing genocide and even that data that the Ministry of Health has been struggling to scrap and put and preserve after 2 years of genocide, we know that it's a new target. So, we are struggling as to what what can you digitize? What can you scan without losing? So, these are the the ethical conversation that we keep having and we, right? We try as much as we can to um to preserve. For instance, this morning, this is just before I jumped into this this call. Sorry, I'm I'm going beyond time, but to give you the kind of obstacle we face. So, this is before I got into this call, we finally have the all of the pictures of all of the medical staff that were killed. And many of these pictures are of the doctors. Some of the Some we were able to find their normal, right? The picture before that or picture on the medical card, so it's a random picture, but many of of the other pictures, especially when the fam all of the family is killed, is a picture of the doctor laying in blood with a bullet between his eyes or behind his his back, right? That a doctor took in his cell phone, so we got access to that. And the field worker received the permission of of the doctor to take that picture from him. So, she has a copy and she sent she sent the entire file to me. And so, I'm not going to put the picture of that doctor with a bullet behind his head online as part of his medical profile on the website. So, right. So, these are the kind of day-to-day things that that it is Yes, it is an archive, and yes, I want this to remain because I want history and the world to know what has been done to us. But, how do we deal with all of this, especially that this is the first time we have an an a genocide that has such vital images for accountability in the future. I see Palestine free, and I see us, you know, putting these people on trial, and I see myself showing them those pictures someday. I see that happening. So, I need to preserve that, too. And the conversation today is for us to understand how do we do this ethically, respecting our people, and respecting a legal space, right? That that that has to protect and and hold accountability. So, it's a big answer, but it shows all of the things that we deal with as we preserve. Yeah. Thank you. Thank you very much. Very important. And um yeah, I think we uh I just wanted to to thank the uh panelists, all of them, and um say hopefully we will be able indeed to do what you said um Hanin, and hopefully we can have these archives. We had one archive that was stolen in Beirut. Uh the PLO archive held that archive, and it was stolen in 1982. We are starting again um with little in hands to start with, but um there is already a lot accumulating thanks to the Israelis, obviously. Uh all right. So, um I leave it now to uh Dr. Mustafa to close the uh workshop. Thank you all. Thank you everybody.