AraTrust: An Evaluation of Trustworthiness for LLMs in Arabic
arXiv · · Significant research
Summary
The paper introduces AraTrust, a new benchmark for evaluating the trustworthiness of LLMs when prompted in Arabic. The benchmark contains 522 multiple-choice questions covering dimensions like truthfulness, ethics, safety, and fairness. Experiments using AraTrust showed that GPT-4 performed the best, while open-source models like AceGPT 7B and Jais 13B had lower scores. Why it matters: This benchmark addresses a critical gap in evaluating LLMs for Arabic, which is essential for ensuring the safe and ethical deployment of AI in the Arab world.
Keywords
AraTrust · LLM · Arabic · benchmark · trustworthiness
Get the weekly digest
Top AI stories from the GCC region, every week.