Skip to content
GCC AI Research

Search

Results for "Multimodal AI"

Almieyar-Oryx-BloomBench: A Bilingual Multimodal Benchmark for Cognitively Informed Evaluation of Vision-Language Models

arXiv ·

Researchers have introduced BloomBench, a new cognitively human-grounded, bilingual (English-Arabic) multimodal benchmark for Vision-Language Models (VLMs), as part of the Almieyar benchmarking series. Grounded in Bloom's Taxonomy, it systematically evaluates six levels of cognition—Remember, Understand, Apply, Analyze, Evaluate, Create—through carefully designed image-question-answer tasks. A comprehensive study using BloomBench revealed that state-of-the-art VLMs exhibit strong semantic understanding but struggle significantly with factual recall and creative synthesis, alongside a critical performance gap between Arabic and English. Why it matters: This benchmark provides a crucial tool for diagnosing cognitive weaknesses in current VLMs and lays the groundwork for developing more cognitively aligned and inclusive multimodal AI, particularly for cross-lingual applications.

TII Launches Falcon Perception, A New Multimodal AI Model That Helps Machines See and Understand the World – with Efficiency that Rivals Larger Models

TII ·

The Technology Innovation Institute (TII) has launched Falcon Perception, a new 600-million-parameter multimodal AI model. This model offers competitive performance in object segmentation, dense visual understanding, and document intelligence, rivalling larger systems like Meta’s SAM3 and Alibaba’s Qwen with significantly greater efficiency. Falcon Perception unifies image and language processing in a single architecture, designed for real-world deployment in compute-constrained environments. Why it matters: This development positions the UAE among leading nations in advanced multimodal AI, which is crucial for applications in robotics, advanced manufacturing, and autonomous platforms.

Luma AI, an AI startup building multimodal AGI, raises $900 million led by Saudi Arabia-based HUMAIN - Indian Startup News

GCC AI Startup ·

Luma AI, a startup developing multimodal AGI, has raised $900 million in a funding round led by HUMAIN, a Saudi Arabia-based investment firm. The company is focused on building general-purpose AI models that can understand and generate different types of data, including images, video, and 3D scenes. This funding round will allow Luma AI to scale its research and development efforts. Why it matters: This investment signals the growing interest and financial commitment from Saudi Arabian entities in advancing artificial general intelligence capabilities.

Fanar: An Arabic-Centric Multimodal Generative AI Platform

arXiv ·

Hamad Bin Khalifa University's Qatar Computing Research Institute (QCRI) introduced Fanar, an Arabic-centric multimodal generative AI platform featuring the Fanar Star (7B) and Fanar Prime (9B) Arabic LLMs. These models were trained on nearly 1 trillion tokens and are designed to address different prompts through a custom orchestrator. Fanar includes a customized Islamic RAG system, a Recency RAG, bilingual speech recognition, and an attribution service for content verification, sponsored by Qatar's Ministry of Communications and Information Technology. Why it matters: The platform signifies a major step towards sovereign AI development in Qatar, providing advanced Arabic language capabilities and addressing regional needs.

A new stress test for AI agents that plan, look and click

MBZUAI ·

MBZUAI researchers won second place at the AgentX Competition at UC Berkeley for their benchmark measuring AI agents' reasoning across images, comparisons, and video. The Agent-X dataset includes 828 tasks across six domains, requiring agents to use 14 executable tools without explicit instructions. Agent-X analyzes the agent's full reasoning trajectory, unlike typical evaluations that focus only on final answers. Why it matters: The benchmark exposes limitations in current multimodal AI agents and provides a more rigorous evaluation framework for real-world applications in the region and beyond.

MBZUAI researchers earn high-profile honors at EMNLP

MBZUAI ·

MBZUAI researchers received high honors at EMNLP 2025 for two research papers, placing them in the top 2% of accepted work. One paper, MAviS, is a multimodal AI system that identifies bird species by combining images, sounds, and text. The other award-winning paper focuses on uncertainty in LLM-as-a-Judge. Why it matters: The recognition highlights MBZUAI's growing influence in NLP and multimodal AI research, particularly in domain-specific applications like biodiversity conservation.

A new playbook for patient privacy in the age of foundation models

MBZUAI ·

MBZUAI researchers Darya Taratynova and Shahad Hardan developed Forget-MI, a method for making clinical AI models "unlearn" specific patient data without retraining the entire model. Forget-MI addresses the challenge of removing patient data from AI models trained on multimodal records (like chest X-rays and reports) due to regulations like GDPR and HIPAA. The method unlearns both unimodal (image or text) and joint (image-text) associations while retaining overall accuracy using a late-fusion multimodal classifier. Why it matters: This research provides a practical solution to a critical privacy concern in healthcare AI, enabling compliance with data protection regulations and fostering trust in AI-driven medical applications.

Five ways AI is creating a healthier future

MBZUAI ·

MBZUAI researchers developed FetalCLIP, an AI model trained on 210,000 ultrasound images for fast and reliable interpretation of fetal scans. MBZUAI's President Eric Xing contributed to the General Expression Transformer (GET), an AI foundation model acting as a biological simulator to predict gene behavior. MBZUAI and Carleton University created MedPromptX for quicker disease diagnosis and treatment plans using multimodal AI. Why it matters: These AI advancements from MBZUAI have the potential to revolutionize healthcare in the region and globally, from prenatal care to drug discovery and personalized medicine.