Skip to content
GCC AI Research

CamelEval: Advancing Culturally Aligned Arabic Language Models and Benchmarks

arXiv · · Significant research

Summary

The paper introduces Juhaina, a 9.24B parameter Arabic-English bilingual LLM trained with an 8,192 token context window. It identifies limitations in the Open Arabic LLM Leaderboard (OALL) and proposes a new benchmark, CamelEval, for more comprehensive evaluation. Juhaina outperforms models like Llama and Gemma in generating helpful Arabic responses and understanding cultural nuances. Why it matters: This culturally-aligned LLM and associated benchmark could significantly advance Arabic NLP and democratize AI access for Arabic speakers.

Keywords

LLM · Arabic · Juhaina · CamelEval · Benchmark

Get the weekly digest

Top AI stories from the GCC region, every week.