Skip to content
GCC AI Research

ORCA: A Challenging Benchmark for Arabic Language Understanding

arXiv · · Significant research

Summary

The paper introduces ORCA, a new public benchmark for evaluating Arabic language understanding. ORCA covers diverse Arabic varieties and includes 60 datasets across seven NLU task clusters. The benchmark was used to compare 18 multilingual and Arabic language models and includes a public leaderboard with a unified evaluation metric. Why it matters: ORCA addresses the lack of a comprehensive Arabic benchmark, enabling better progress measurement for Arabic and multilingual language models.

Keywords

ORCA · Arabic · NLU · benchmark · language models

Get the weekly digest

Top AI stories from the GCC region, every week.