Video summary
The Birzeit University professors have achieved a historic milestone for the Arabic language by developing the Qabas Computer Dictionary, which stands as the largest open-source lexical resource designed specifically for artificial intelligence applications. Unlike traditional dictionaries that are structured for paper-based reference and lack the necessary data formats for computational tasks, this innovative project was created to meet emerging technological needs in language processing. The development process spanned four years of pioneering work, resulting in a completely open-source tool available for both profit and non-profit purposes, marking a significant shift from static reference books to dynamic digital networks.
The core architecture of the Qabas Dictionary functions as a comprehensive lexical data network that links its entries to 110 different Arabic dictionaries and morphologically labeled textual data. When a user searches for a specific word, such as "book," the system retrieves not only the entry but also information aggregated from all connected sources, including various blogs like the Holy Quran blog and Al-Fusha. This extensive integration allows the dictionary to include nine distinct corpora covering colloquial Arabic, ensuring that entries are rich with contextual examples; for instance, the word "kitab" is shown to have 68 conjugations and appears in 66 different contexts, providing a depth of linguistic data previously unavailable in standard formats.
A defining feature of this dictionary is its advanced morphological analysis, which breaks words down into their constituent parts to identify the type of each component during searches. This capability enables users to find root words regardless of their grammatical form; searching for plural forms like "kutub" or verb phrases like "katabahu" will successfully return the base entry "kitab" because the system recognizes these as related and more common variations. By analyzing these structures, the Qabas Dictionary facilitates the creation of morphological and semantic analyzers, allowing artificial intelligence systems to understand Arabic language nuances accurately while supporting statistical linguistic research based on large datasets like the Linus corpus.
Ultimately, the Qabas Computer Dictionary represents a transformative leap in how the Arabic language is processed digitally, bridging the gap between traditional lexicography and modern computational requirements. Users can access detailed methodological statistics and standard guides to understand the dictionary's construction, with all resources available for download via the dedicated link provided on the project's About page. This open-source initiative not only empowers researchers and developers to build sophisticated language applications but also ensures that the rich history of Arabic scholarship is preserved and enhanced through innovation, making it a vital tool for the future of digital humanities and AI in the Arab world.
Read the full video transcript
[Music]
Birzeit University professors have a rich history of
pioneering and innovation, and this time it's a historic scientific achievement for the
Arabic language: the Qabas Computer Dictionary, the
largest open-source dictionary of the Arabic language,
designed for artificial intelligence applications.
Its entries are linked to 110
Arabic dictionaries and to morphologically labeled textual data. The idea behind the
Qabas Dictionary came about to meet emerging technological needs
such as artificial intelligence applications,
because
traditional dictionaries are designed for paper use and are not
intended for language processing and computational language tasks.
For example, Qabas can be used to build
morphological and semantic analyzers and other applications. Its development took
four years, and it is completely open source for both
profit and non-profit purposes. For example, a
person can search for a word like "book"
and retrieve its entry and information from
various Arabic dictionaries, as each entry in the
Qabas Dictionary is linked to its corresponding entries in 110 dictionaries.
In addition, contexts in which the
word "book" appears can be retrieved from several blogs, such as the
Holy Quran blog and my blog, Al-Fusha. It also includes
nine corpora in colloquial Arabic.
For example, the entry for the word "kitab" (book) has 68 conjugations and appears
in
66 contexts. You can view the contexts in which
the word appears in these corpora. The
dictionary breaks down the word into its
parts and identifies the type of each
part. Thus, Qabas Dictionary is a lexical data network that
connects Arabic dictionaries to text corpora. The dictionary is
designed to enable artificial intelligence applications
to understand the language accurately and to
facilitate linguistic research based on
statistics, such as the Linus corpus. One of
the features of Qabas Dictionary is its morphological analysis during
searches. You can search for different forms; for example,
the word "kitabayn" (two books) will return "kitab" (book). If you
search for "kutub" (books), which is the plural of "kitab," or "katabahu" (he
wrote), it will return "kitab" because it is more common
than "katabahu." Please pay attention to the analyses
at the top of the page when
searching. You can also view statistics
on the dictionary's methodology and standard guide,
and download the dictionary from the About page via
the link Qabas Sniper.