Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Nizar Habash

New York University Abu Dhabi

NADI 2024: The Fifth Nuanced Arabic Dialect Identification Shared Task

Jul 06, 2024

Muhammad Abdul-Mageed, Amr Keleg, AbdelRahim Elmadany, Chiyu Zhang, Injy Hamed, Walid Magdy, Houda Bouamor, Nizar Habash

Abstract:We describe the findings of the fifth Nuanced Arabic Dialect Identification Shared Task (NADI 2024). NADI's objective is to help advance SoTA Arabic NLP by providing guidance, datasets, modeling opportunities, and standardized evaluation conditions that allow researchers to collaboratively compete on pre-specified tasks. NADI 2024 targeted both dialect identification cast as a multi-label task (Subtask~1), identification of the Arabic level of dialectness (Subtask~2), and dialect-to-MSA machine translation (Subtask~3). A total of 51 unique teams registered for the shared task, of whom 12 teams have participated (with 76 valid submissions during the test phase). Among these, three teams participated in Subtask~1, three in Subtask~2, and eight in Subtask~3. The winning teams achieved 50.57 F\textsubscript{1} on Subtask~1, 0.1403 RMSE for Subtask~2, and 20.44 BLEU in Subtask~3, respectively. Results show that Arabic dialect processing tasks such as dialect identification and machine translation remain challenging. We describe the methods employed by the participating teams and briefly offer an outlook for NADI.

* Accepted by The Second Arabic Natural Language Processing Conference

Via

Access Paper or Ask Questions

Exploiting Dialect Identification in Automatic Dialectal Text Normalization

Jul 03, 2024

Bashar Alhafni, Sarah Al-Towaity, Ziyad Fawzy, Fatema Nassar, Fadhl Eryani, Houda Bouamor, Nizar Habash

Figure 1 for Exploiting Dialect Identification in Automatic Dialectal Text Normalization

Figure 2 for Exploiting Dialect Identification in Automatic Dialectal Text Normalization

Figure 3 for Exploiting Dialect Identification in Automatic Dialectal Text Normalization

Figure 4 for Exploiting Dialect Identification in Automatic Dialectal Text Normalization

Abstract:Dialectal Arabic is the primary spoken language used by native Arabic speakers in daily communication. The rise of social media platforms has notably expanded its use as a written language. However, Arabic dialects do not have standard orthographies. This, combined with the inherent noise in user-generated content on social media, presents a major challenge to NLP applications dealing with Dialectal Arabic. In this paper, we explore and report on the task of CODAfication, which aims to normalize Dialectal Arabic into the Conventional Orthography for Dialectal Arabic (CODA). We work with a unique parallel corpus of multiple Arabic dialects focusing on five major city dialects. We benchmark newly developed pretrained sequence-to-sequence models on the task of CODAfication. We further show that using dialect identification information improves the performance across all dialects. We make our code, data, and pretrained models publicly available.

* Accepted to ArabicNLP 2024, ACL

Via

Access Paper or Ask Questions

Strategies for Arabic Readability Modeling

Jul 03, 2024

Juan Piñeros Liberato, Bashar Alhafni, Muhamed Al Khalil, Nizar Habash

Figure 1 for Strategies for Arabic Readability Modeling

Figure 2 for Strategies for Arabic Readability Modeling

Figure 3 for Strategies for Arabic Readability Modeling

Figure 4 for Strategies for Arabic Readability Modeling

Abstract:Automatic readability assessment is relevant to building NLP applications for education, content analysis, and accessibility. However, Arabic readability assessment is a challenging task due to Arabic's morphological richness and limited readability resources. In this paper, we present a set of experimental results on Arabic readability assessment using a diverse range of approaches, from rule-based methods to Arabic pretrained language models. We report our results on a newly created corpus at different textual granularity levels (words and sentence fragments). Our results show that combining different techniques yields the best results, achieving an overall macro F1 score of 86.7 at the word level and 87.9 at the fragment level on a blind test set. We make our code, data, and pretrained models publicly available.

* Accepted to ArabicNLP 2024, ACL

Via

Access Paper or Ask Questions

Arabic Diacritics in the Wild: Exploiting Opportunities for Improved Diacritization

Jun 09, 2024

Salman Elgamal, Ossama Obeid, Tameem Kabbani, Go Inoue, Nizar Habash

Figure 1 for Arabic Diacritics in the Wild: Exploiting Opportunities for Improved Diacritization

Figure 2 for Arabic Diacritics in the Wild: Exploiting Opportunities for Improved Diacritization

Figure 3 for Arabic Diacritics in the Wild: Exploiting Opportunities for Improved Diacritization

Figure 4 for Arabic Diacritics in the Wild: Exploiting Opportunities for Improved Diacritization

Abstract:The widespread absence of diacritical marks in Arabic text poses a significant challenge for Arabic natural language processing (NLP). This paper explores instances of naturally occurring diacritics, referred to as "diacritics in the wild," to unveil patterns and latent information across six diverse genres: news articles, novels, children's books, poetry, political documents, and ChatGPT outputs. We present a new annotated dataset that maps real-world partially diacritized words to their maximal full diacritization in context. Additionally, we propose extensions to the analyze-and-disambiguate approach in Arabic NLP to leverage these diacritics, resulting in notable improvements. Our contributions encompass a thorough analysis, valuable datasets, and an extended diacritization algorithm. We release our code and datasets as open source.

* Accepted to ACL 2024

Via

Access Paper or Ask Questions

The SAMER Arabic Text Simplification Corpus

Apr 29, 2024

Bashar Alhafni, Reem Hazim, Juan Piñeros Liberato, Muhamed Al Khalil, Nizar Habash

Figure 1 for The SAMER Arabic Text Simplification Corpus

Figure 2 for The SAMER Arabic Text Simplification Corpus

Figure 3 for The SAMER Arabic Text Simplification Corpus

Figure 4 for The SAMER Arabic Text Simplification Corpus

Abstract:We present the SAMER Corpus, the first manually annotated Arabic parallel corpus for text simplification targeting school-aged learners. Our corpus comprises texts of 159K words selected from 15 publicly available Arabic fiction novels most of which were published between 1865 and 1955. Our corpus includes readability level annotations at both the document and word levels, as well as two simplified parallel versions for each text targeting learners at two different readability levels. We describe the corpus selection process, and outline the guidelines we followed to create the annotations and ensure their quality. Our corpus is publicly available to support and encourage research on Arabic text simplification, Arabic automatic readability assessment, and the development of Arabic pedagogical language technologies.

* Accepted to LREC-COLING 2024. 15 pages, 6 tables, 1 figure

Via

Access Paper or Ask Questions

Can a Multichoice Dataset be Repurposed for Extractive Question Answering?

Apr 26, 2024

Teresa Lynn, Malik H. Altakrori, Samar Mohamed Magdy, Rocktim Jyoti Das, Chenyang Lyu, Mohamed Nasr, Younes Samih, Alham Fikri Aji, Preslav Nakov, Shantanu Godbole(+3 more)

Figure 1 for Can a Multichoice Dataset be Repurposed for Extractive Question Answering?

Figure 2 for Can a Multichoice Dataset be Repurposed for Extractive Question Answering?

Figure 3 for Can a Multichoice Dataset be Repurposed for Extractive Question Answering?

Figure 4 for Can a Multichoice Dataset be Repurposed for Extractive Question Answering?

Abstract:The rapid evolution of Natural Language Processing (NLP) has favored major languages such as English, leaving a significant gap for many others due to limited resources. This is especially evident in the context of data annotation, a task whose importance cannot be underestimated, but which is time-consuming and costly. Thus, any dataset for resource-poor languages is precious, in particular when it is task-specific. Here, we explore the feasibility of repurposing existing datasets for a new NLP task: we repurposed the Belebele dataset (Bandarkar et al., 2023), which was designed for multiple-choice question answering (MCQA), to enable extractive QA (EQA) in the style of machine reading comprehension. We present annotation guidelines and a parallel EQA dataset for English and Modern Standard Arabic (MSA). We also present QA evaluation results for several monolingual and cross-lingual QA pairs including English, MSA, and five Arabic dialects. Our aim is to enable others to adapt our approach for the 120+ other language variants in Belebele, many of which are deemed under-resourced. We also conduct a thorough analysis and share our insights from the process, which we hope will contribute to a deeper understanding of the challenges and the opportunities associated with task reformulation in NLP research.

* Paper 8 pages, Appendix 12 pages. Submitted to ARR

Via

Access Paper or Ask Questions

SemEval-2024 Task 8: Multidomain, Multimodel and Multilingual Machine-Generated Text Detection

Apr 22, 2024

Yuxia Wang, Jonibek Mansurov, Petar Ivanov, Jinyan Su, Artem Shelmanov, Akim Tsvigun, Osama Mohammed Afzal, Tarek Mahmoud, Giovanni Puccetti, Thomas Arnold(+5 more)

Figure 1 for SemEval-2024 Task 8: Multidomain, Multimodel and Multilingual Machine-Generated Text Detection

Figure 2 for SemEval-2024 Task 8: Multidomain, Multimodel and Multilingual Machine-Generated Text Detection

Figure 3 for SemEval-2024 Task 8: Multidomain, Multimodel and Multilingual Machine-Generated Text Detection

Figure 4 for SemEval-2024 Task 8: Multidomain, Multimodel and Multilingual Machine-Generated Text Detection

Abstract:We present the results and the main findings of SemEval-2024 Task 8: Multigenerator, Multidomain, and Multilingual Machine-Generated Text Detection. The task featured three subtasks. Subtask A is a binary classification task determining whether a text is written by a human or generated by a machine. This subtask has two tracks: a monolingual track focused solely on English texts and a multilingual track. Subtask B is to detect the exact source of a text, discerning whether it is written by a human or generated by a specific LLM. Subtask C aims to identify the changing point within a text, at which the authorship transitions from human to machine. The task attracted a large number of participants: subtask A monolingual (126), subtask A multilingual (59), subtask B (70), and subtask C (30). In this paper, we present the task, analyze the results, and discuss the system submissions and the methods they used. For all subtasks, the best systems used LLMs.

* Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval-2024)
* 23 pages, 12 tables

Via

Access Paper or Ask Questions

ZAEBUC-Spoken: A Multilingual Multidialectal Arabic-English Speech Corpus

Mar 27, 2024

Injy Hamed, Fadhl Eryani, David Palfreyman, Nizar Habash

Figure 1 for ZAEBUC-Spoken: A Multilingual Multidialectal Arabic-English Speech Corpus

Figure 2 for ZAEBUC-Spoken: A Multilingual Multidialectal Arabic-English Speech Corpus

Figure 3 for ZAEBUC-Spoken: A Multilingual Multidialectal Arabic-English Speech Corpus

Figure 4 for ZAEBUC-Spoken: A Multilingual Multidialectal Arabic-English Speech Corpus

Abstract:We present ZAEBUC-Spoken, a multilingual multidialectal Arabic-English speech corpus. The corpus comprises twelve hours of Zoom meetings involving multiple speakers role-playing a work situation where Students brainstorm ideas for a certain topic and then discuss it with an Interlocutor. The meetings cover different topics and are divided into phases with different language setups. The corpus presents a challenging set for automatic speech recognition (ASR), including two languages (Arabic and English) with Arabic spoken in multiple variants (Modern Standard Arabic, Gulf Arabic, and Egyptian Arabic) and English used with various accents. Adding to the complexity of the corpus, there is also code-switching between these languages and dialects. As part of our work, we take inspiration from established sets of transcription guidelines to present a set of guidelines handling issues of conversational speech, code-switching and orthography of both languages. We further enrich the corpus with two layers of annotations; (1) dialectness level annotation for the portion of the corpus where mixing occurs between different variants of Arabic, and (2) automatic morphological annotations, including tokenization, lemmatization, and part-of-speech tagging.

* Accepted to LREC-COLING 2024

Via

Access Paper or Ask Questions

ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic

Feb 20, 2024

Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata(+3 more)

Figure 1 for ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic

Figure 2 for ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic

Figure 3 for ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic

Figure 4 for ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic

Abstract:The focus of language model evaluation has transitioned towards reasoning and knowledge-intensive tasks, driven by advancements in pretraining large models. While state-of-the-art models are partially trained on large Arabic texts, evaluating their performance in Arabic remains challenging due to the limited availability of relevant datasets. To bridge this gap, we present ArabicMMLU, the first multi-task language understanding benchmark for Arabic language, sourced from school exams across diverse educational levels in different countries spanning North Africa, the Levant, and the Gulf regions. Our data comprises 40 tasks and 14,575 multiple-choice questions in Modern Standard Arabic (MSA), and is carefully constructed by collaborating with native speakers in the region. Our comprehensive evaluations of 35 models reveal substantial room for improvement, particularly among the best open-source models. Notably, BLOOMZ, mT0, LLama2, and Falcon struggle to achieve a score of 50%, while even the top-performing Arabic-centric model only achieves a score of 62.3%.

Via

Access Paper or Ask Questions

M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text Detection

Feb 17, 2024

Yuxia Wang, Jonibek Mansurov, Petar Ivanov, Jinyan Su, Artem Shelmanov, Akim Tsvigun, Osama Mohanned Afzal, Tarek Mahmoud, Giovanni Puccetti, Thomas Arnold(+4 more)

Figure 1 for M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text Detection

Figure 2 for M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text Detection

Figure 3 for M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text Detection

Figure 4 for M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text Detection

Abstract:The advent of Large Language Models (LLMs) has brought an unprecedented surge in machine-generated text (MGT) across diverse channels. This raises legitimate concerns about its potential misuse and societal implications. The need to identify and differentiate such content from genuine human-generated text is critical in combating disinformation, preserving the integrity of education and scientific fields, and maintaining trust in communication. In this work, we address this problem by introducing a new benchmark involving multilingual, multi-domain and multi-generator for MGT detection -- M4GT-Bench. It is collected for three task formulations: (1) mono-lingual and multi-lingual binary MGT detection; (2) multi-way detection identifies which particular model generates the text; and (3) human-machine mixed text detection, where a word boundary delimiting MGT from human-written content should be determined. Human evaluation for Task 2 shows less than random guess performance, demonstrating the challenges to distinguish unique LLMs. Promising results always occur when training and test data distribute within the same domain or generators.

* 28 pages

Via

Access Paper or Ask Questions