EDRAC:阿拉伯方言阅读理解基准测试

EDRAC: Benchmarking Arabic Dialect Reading Comprehension

精选理由

阿拉伯方言NLP研究者必看,首个覆盖五大方言的阅读理解基准,揭示现有模型局限。

AI 摘要

EDRAC是首个大规模阿拉伯方言机器阅读理解和生成式问答基准,涵盖埃及、摩洛哥、阿联酋、叙利亚和沙特五种主要方言。该基准包含499段自然对话语料和4977个问答对,采用人机协作流程生成。研究显示,现有模型在语义答案质量和方言保真度间存在显著差距。

原文 · arXiv cs.AI

EDRAC: Benchmarking Arabic Dialect Reading Comprehension

Dialectal Arabic (DA) remains under-resourced compared to Modern Standard Arabic (MSA), particularly for machine reading comprehension (MRC) and question answering (QA). Existing Arabic QA benchmarks primarily focus on formal written MSA or multiple-choice QA, with limited coverage of naturally spoken dialects. Here, we aim to bridge this gap. We introduce EDRAC, the first large-scale benchmark for dialectal Arabic machine reading comprehension (MRC) and generative QA, covering five major dialects: Egyptian, Moroccan, Emirati, Syrian, and Saudi Arabic. EDRAC contains 499 passages derived from naturally occurring spoken interactions and 4,977 corresponding QA pairs generated through a human--LLM collaborative pipeline combining iterative generation, LLM-as-a-judge evaluation, and human verification. We benchmark Arabic-centric and multilingual LLMs on EDRAC using lexical and semantic metrics. Our results reveal substantial gaps between semantic answer quality and dialectal fidelity, highlighting the limitations of existing evaluation metrics for dialectal Arabic generation. EDRAC provides a realistic and challenging MRC benchmark for future research on dialectal Arabic NLP.