Skip to content
GCC AI Research

How well can LLMs Grade Essays in Arabic?

arXiv · · Significant research

Summary

This research evaluates LLMs like ChatGPT, Llama, Aya, Jais, and ACEGPT on Arabic automated essay scoring (AES) using the AR-AES dataset. The study uses zero-shot, few-shot learning, and fine-tuning approaches while using a mixed-language prompting strategy. ACEGPT performed best among the LLMs with a QWK of 0.67, while a smaller BERT model achieved 0.88. Why it matters: The study highlights challenges faced by LLMs in processing Arabic and provides insights into improving LLM performance in Arabic NLP tasks.

Keywords

LLM · Arabic · Essay Scoring · ChatGPT · Jais

Get the weekly digest

Top AI stories from the GCC region, every week.