Submind YouTube summaries
Thumbnail for معجم قبس الحاسوبي

معجم قبس الحاسوبي

Watch on YouTube

Video summary

The Birzeit University professors have achieved a historic milestone for the Arabic language by developing the Qabas Computer Dictionary, which stands as the largest open-source lexical resource designed specifically for artificial intelligence applications. Unlike traditional dictionaries that are structured for paper-based reference and lack the necessary data formats for computational tasks, this innovative project was created to meet emerging technological needs in language processing. The development process spanned four years of pioneering work, resulting in a completely open-source tool available for both profit and non-profit purposes, marking a significant shift from static reference books to dynamic digital networks. The core architecture of the Qabas Dictionary functions as a comprehensive lexical data network that links its entries to 110 different Arabic dictionaries and morphologically labeled textual data. When a user searches for a specific word, such as "book," the system retrieves not only the entry but also information aggregated from all connected sources, including various blogs like the Holy Quran blog and Al-Fusha. This extensive integration allows the dictionary to include nine distinct corpora covering colloquial Arabic, ensuring that entries are rich with contextual examples; for instance, the word "kitab" is shown to have 68 conjugations and appears in 66 different contexts, providing a depth of linguistic data previously unavailable in standard formats. A defining feature of this dictionary is its advanced morphological analysis, which breaks words down into their constituent parts to identify the type of each component during searches. This capability enables users to find root words regardless of their grammatical form; searching for plural forms like "kutub" or verb phrases like "katabahu" will successfully return the base entry "kitab" because the system recognizes these as related and more common variations. By analyzing these structures, the Qabas Dictionary facilitates the creation of morphological and semantic analyzers, allowing artificial intelligence systems to understand Arabic language nuances accurately while supporting statistical linguistic research based on large datasets like the Linus corpus. Ultimately, the Qabas Computer Dictionary represents a transformative leap in how the Arabic language is processed digitally, bridging the gap between traditional lexicography and modern computational requirements. Users can access detailed methodological statistics and standard guides to understand the dictionary's construction, with all resources available for download via the dedicated link provided on the project's About page. This open-source initiative not only empowers researchers and developers to build sophisticated language applications but also ensures that the rich history of Arabic scholarship is preserved and enhanced through innovation, making it a vital tool for the future of digital humanities and AI in the Arab world.
Read the full video transcript
[Music] Birzeit University professors have a rich history of pioneering and innovation, and this time it's a historic scientific achievement for the Arabic language: the Qabas Computer Dictionary, the largest open-source dictionary of the Arabic language, designed for artificial intelligence applications. Its entries are linked to 110 Arabic dictionaries and to morphologically labeled textual data. The idea behind the Qabas Dictionary came about to meet emerging technological needs such as artificial intelligence applications, because traditional dictionaries are designed for paper use and are not intended for language processing and computational language tasks. For example, Qabas can be used to build morphological and semantic analyzers and other applications. Its development took four years, and it is completely open source for both profit and non-profit purposes. For example, a person can search for a word like "book" and retrieve its entry and information from various Arabic dictionaries, as each entry in the Qabas Dictionary is linked to its corresponding entries in 110 dictionaries. In addition, contexts in which the word "book" appears can be retrieved from several blogs, such as the Holy Quran blog and my blog, Al-Fusha. It also includes nine corpora in colloquial Arabic. For example, the entry for the word "kitab" (book) has 68 conjugations and appears in 66 contexts. You can view the contexts in which the word appears in these corpora. The dictionary breaks down the word into its parts and identifies the type of each part. Thus, Qabas Dictionary is a lexical data network that connects Arabic dictionaries to text corpora. The dictionary is designed to enable artificial intelligence applications to understand the language accurately and to facilitate linguistic research based on statistics, such as the Linus corpus. One of the features of Qabas Dictionary is its morphological analysis during searches. You can search for different forms; for example, the word "kitabayn" (two books) will return "kitab" (book). If you search for "kutub" (books), which is the plural of "kitab," or "katabahu" (he wrote), it will return "kitab" because it is more common than "katabahu." Please pay attention to the analyses at the top of the page when searching. You can also view statistics on the dictionary's methodology and standard guide, and download the dictionary from the About page via the link Qabas Sniper.