Skip to content
GCC AI Research

Search

Results for "multimodal models"

Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models

arXiv ·

Researchers have introduced VISE (Visual Invariance Self-Evolution), a purely unsupervised framework designed to address 'visual under-conditioning' in self-evolving Large Multimodal Models (LMMs). VISE utilizes geometric and semantic invariance-based rewards to directly regularize the model's visual conditioning, ensuring it attends to visual content rather than relying on language priors. Trained on raw unlabeled images, experiments using Qwen3-VL-2B demonstrate significant performance gains, including +16.85 CIDEr on COCO and a 5.0-point reduction in object hallucination across 18 benchmarks. Why it matters: This research from MBZUAI offers a significant advancement in improving the visual reasoning capabilities and reliability of LMMs in unsupervised settings, making them more robust for real-world applications.

CoVR-R:Reason-Aware Composed Video Retrieval

arXiv ·

A new approach to composed video retrieval (CoVR) is presented, which leverages large multimodal models to infer causal and temporal consequences implied by an edit. The method aligns reasoned queries to candidate videos without task-specific finetuning. A new benchmark, CoVR-Reason, is introduced to evaluate reasoning in CoVR.

DuwatBench: Bridging Language and Visual Heritage through an Arabic Calligraphy Benchmark for Multimodal Understanding

arXiv ·

MBZUAI researchers introduce DuwatBench, a new benchmark for multimodal understanding of Arabic calligraphy. The dataset contains 1,272 samples across six calligraphic styles with detailed annotations to evaluate visual-text alignment. Evaluation of 13 multimodal models reveals challenges in processing calligraphic variations and artistic distortions, highlighting the need for culturally grounded AI research.

EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards

arXiv ·

Researchers at MBZUAI have introduced EvoLMM, a self-evolving framework for large multimodal models that enhances reasoning capabilities without human-annotated data or reward distillation. EvoLMM uses two cooperative agents, a Proposer and a Solver, which generate image-grounded questions and solve them through internal consistency, using a continuous self-rewarding process. Evaluations using Qwen2.5-VL as the base model showed performance gains of up to 3% on multimodal math-reasoning benchmarks like ChartQA, MathVista, and MathVision using only raw training images.

Tracking Meets Large Multimodal Models for Driving Scenario Understanding

arXiv ·

Researchers at MBZUAI have introduced a novel approach to enhance Large Multimodal Models (LMMs) for autonomous driving by integrating 3D tracking information. This method uses a track encoder to embed spatial and temporal data, enriching visual queries and improving the LMM's understanding of driving scenarios. Experiments on DriveLM-nuScenes and DriveLM-CARLA benchmarks demonstrate significant improvements in perception, planning, and prediction tasks compared to baseline models.

Time Travel: A Comprehensive Benchmark to Evaluate LMMs on Historical and Cultural Artifacts

arXiv ·

Researchers introduce TimeTravel, a benchmark dataset for evaluating large multimodal models (LMMs) on historical and cultural artifacts. The benchmark comprises 10,250 expert-verified samples across 266 cultures and 10 historical regions, designed to assess AI in tasks like classification and interpretation of manuscripts, artworks, inscriptions, and archaeological discoveries. The goal is to establish AI as a reliable partner in preserving cultural heritage and assisting researchers.

MBZUAI launches five new “first-of-its-kind” LLMs to support real-world applications and use cases

MBZUAI ·

MBZUAI's Institute of Foundation Models (IFM) has launched five new specialized language and multimodal models, including BiMediX, PALO, GLaMM, GeoChat, and MobiLLaMA. These models address real-world applications in healthcare, visual reasoning, multilingual capabilities, geospatial analysis, and mobile device efficiency. BiMediX is a bilingual medical LLM, while GLaMM generates natural language responses related to objects in an image at the pixel level. Why it matters: This launch demonstrates MBZUAI's commitment to advancing AI research and developing practical AI solutions for various industries, especially with a focus on Arabic language capabilities.

MBZUAI celebrates faculty excellence at annual recognition reception

MBZUAI ·

MBZUAI recognized seven faculty members for outstanding contributions in research, teaching, and mentorship at its annual Faculty Recognition and Welcome Reception. Associate Professor Salman Khan received the Distinguished Research Award for his work on multimodal models for remote Earth observation, including projects like AI4Weather and the AI Global Agriculture Advisory. Assistant Professor Alham Fikri Aji received the Early Career Researcher Award for his contributions to low-resource NLP and international collaborations. Why it matters: The awards highlight MBZUAI's focus on advancing AI for global challenges and recognizing faculty contributions to research and education.