Researchers have introduced VISE (Visual Invariance Self-Evolution), a purely unsupervised framework designed to address 'visual under-conditioning' in self-evolving Large Multimodal Models (LMMs). VISE utilizes geometric and semantic invariance-based rewards to directly regularize the model's visual conditioning, ensuring it attends to visual content rather than relying on language priors. Trained on raw unlabeled images, experiments using Qwen3-VL-2B demonstrate significant performance gains, including +16.85 CIDEr on COCO and a 5.0-point reduction in object hallucination across 18 benchmarks. Why it matters: This research from MBZUAI offers a significant advancement in improving the visual reasoning capabilities and reliability of LMMs in unsupervised settings, making them more robust for real-world applications.
Researchers at MBZUAI have introduced EvoLMM, a self-evolving framework for large multimodal models that enhances reasoning capabilities without human-annotated data or reward distillation. EvoLMM uses two cooperative agents, a Proposer and a Solver, which generate image-grounded questions and solve them through internal consistency, using a continuous self-rewarding process. Evaluations using Qwen2.5-VL as the base model showed performance gains of up to 3% on multimodal math-reasoning benchmarks like ChartQA, MathVista, and MathVision using only raw training images.
MBZUAI researchers introduce TerraFM, a scalable self-supervised learning model for Earth observation that uses Sentinel-1 and Sentinel-2 imagery. The model unifies radar and optical inputs through modality-specific patch embeddings and adaptive cross-attention fusion. TerraFM achieves strong generalization on classification and segmentation tasks, outperforming prior models on GEO-Bench and Copernicus-Bench.
Researchers propose a universal anatomical embedding (UAE) framework for medical image analysis to learn appearance, semantic, and cross-modality anatomical embeddings. UAE incorporates semantic embedding learning with prototypical contrastive loss, a fixed-point-based matching strategy, and an iterative approach for cross-modality embedding learning. The framework was evaluated on landmark detection, lesion tracking and CT-MRI registration tasks, outperforming existing state-of-the-art methods.
This paper introduces a self-supervised learning method for point cloud analysis using an upsampling autoencoder (UAE). The model uses subsampling and an encoder-decoder architecture to reconstruct the original point cloud, learning both semantic and geometric information. Experiments show the UAE outperforms existing methods in shape classification, part segmentation, and point cloud upsampling tasks.
MBZUAI researchers tackled the challenge of AI-powered waste detection in messy, real-world recycling facilities. They fine-tuned modern object detection models on real industrial waste imagery and combined this with a semi-supervised learning pipeline. Fine-tuning more than doubled performance and their semi-supervised pipeline outperformed fully supervised baselines. Why it matters: This research offers a practical path for open research that can rival proprietary systems while reducing the need for costly manual labeling in waste management, a problem of global importance.
Cyrill Stachniss from the University of Bonn presented recent work on agricultural robotics and self-driving cars. The talk covered autonomous field robots and their ability to perceive, model, and predict future developments in complex farming environments. The presentation also included developments in supervised and unsupervised learning for autonomous car perception systems. Why it matters: This highlights the growing interest in robotics research at MBZUAI and the potential for AI to transform key sectors in the GCC region like agriculture and transportation.
Xiao Wang from Purdue University presented research on Adversarial Contrastive Learning (AdCo) and Cooperative-adversarial Contrastive Learning (CaCo) for improved self-supervised learning. He also discussed CryoREAD, a framework for building DNA/RNA structures from cryo-EM maps, and future work in deep learning for drug discovery. Wang's algorithms have impacted molecular biology, leading to new structure discoveries published in journals like Cell and Nature Microbiology. Why it matters: The research advances AI techniques for crucial tasks in molecular biology and drug discovery, with potential applications for institutions in the GCC region focused on healthcare and biotechnology.