Skip to content
GCC AI Research

Evaluating Arabic Large Language Models: A Survey of Benchmarks, Methods, and Gaps

arXiv · · Significant research

Summary

This survey paper analyzes over 40 benchmarks used to evaluate Arabic large language models, categorizing them into Knowledge, NLP Tasks, Culture and Dialects, and Target-Specific evaluations. It identifies progress in benchmark diversity but also highlights gaps like limited temporal evaluation and cultural misalignment. The paper also examines methods for creating benchmarks, including native collection, translation, and synthetic generation. Why it matters: The survey provides a comprehensive reference for Arabic NLP research and offers recommendations for future benchmark development to better align with cultural contexts.

Get the weekly digest

Top AI stories from the GCC region, every week.