Why AI can describe an image but struggles to understand the culture inside it
MBZUAI · Significant research
Summary
MBZUAI researchers release JEEM, a new benchmark dataset for evaluating vision-language models on Arabic dialects. The dataset covers image captioning and visual question answering tasks using images from Jordan, UAE, Egypt, and Morocco. Results show models struggle with cultural understanding and relevance despite fluent language generation.
Keywords
MBZUAI · JEEM · Arabic dialects · vision-language models · image captioning
Get the weekly digest
Top AI stories from the GCC region, every week.