ArGuard Shared Task: Harmful Content Detection in Arabic Memes and LLM Prompts
arXiv · · Significant research
Summary
ArGuard is a shared task focused on detecting harmful content in Arabic memes and LLM prompts, featuring two tracks: multimodal hate detection in memes and harmful prompt detection for Arabic LLM safety evaluation. The task saw 58 registered teams, 35 participating in evaluation, and 27 submitting system papers, with models like AraBERT, Jais, and Qwen3-VL explored. Best systems achieved macro-F1 scores up to 0.984 on specific subtasks, though fine-grained meme classification proved challenging due to sparse labels and distribution shifts. Why it matters: This initiative significantly advances research into Arabic AI safety and ethical development, addressing critical issues in multimodal content moderation and responsible LLM deployment for the region.
Keywords
ArGuard · harmful content detection · Arabic memes · LLM prompts · Arabic AI safety
Get the weekly digest
Top AI stories from the GCC region, every week.