Skip to content
GCC AI Research

Why AI can describe an image but struggles to understand the culture inside it

MBZUAI · Significant research

Summary

MBZUAI researchers release JEEM, a new benchmark dataset for evaluating vision-language models on Arabic dialects. The dataset covers image captioning and visual question answering tasks using images from Jordan, UAE, Egypt, and Morocco. Results show models struggle with cultural understanding and relevance despite fluent language generation.

Get the weekly digest

Top AI stories from the GCC region, every week.