The paper introduces UAE-3D, a multi-modal VAE for 3D molecule generation that compresses molecules into a unified latent space, maintaining near-zero reconstruction error. This approach simplifies latent diffusion modeling by eliminating the need to handle multi-modality and equivariance separately. Experiments on GEOM-Drugs and QM9 datasets show UAE-3D establishes new benchmarks in de novo and conditional 3D molecule generation, with significant improvements in efficiency and quality.
Researchers propose a universal anatomical embedding (UAE) framework for medical image analysis to learn appearance, semantic, and cross-modality anatomical embeddings. UAE incorporates semantic embedding learning with prototypical contrastive loss, a fixed-point-based matching strategy, and an iterative approach for cross-modality embedding learning. The framework was evaluated on landmark detection, lesion tracking and CT-MRI registration tasks, outperforming existing state-of-the-art methods.
MBZUAI's BiMediX2, a bilingual healthcare multi-modal model, won Meta's Llama Impact Innovation Award for its potential in solving healthcare accessibility challenges across the Middle East and Africa. Built using Llama 3.1, the model understands medical queries in both English and Arabic, interprets medical images, and is integrated as a chatbot on Telegram with speech functionality. The model was also presented at the AI for Sustainable Development Platform Launch Event and integrated into the UNDP for telemedicine. Why it matters: The model's bilingual capabilities and accessibility on low-cost devices and speech-based interaction have potential to improve healthcare access for marginalized populations in the region.
This article discusses MBZUAI's efforts in advancing Arabic language AI, including the development of advanced linguistic models using deep learning techniques. Key initiatives include Jais, a 13B parameter Arabic LLM developed in collaboration with G42's Inception, and Atlas-Chat, which understands the Moroccan dialect. The university is also incorporating Arabic in practical AI solutions like BiMediX2, a healthcare multi-modal model that understands medical queries in both English and Arabic. Why it matters: These initiatives are crucial for preserving Arabic cultural heritage, enabling future discovery, and addressing linguistic challenges specific to the Arabic language in AI applications.
MBZUAI, in partnership with IBM Research, is developing GeoChat+, a vision-language model (VLM) for multi-modal, temporal remote sensing image analysis. GeoChat+ builds on the previous GeoChat model, enhancing capabilities with multi-modal images from various Earth observation systems like Sentinel-1, Sentinel-2, Landsat, and high-resolution imagery. GeoChat+ will integrate data from multiple satellites at different times to detect environmental changes and analyze the impact on soil quality, air quality, and erosion. Why it matters: This advancement promises to revolutionize geographic data analysis, providing detailed reports for high-risk regions and aiding reforestation efforts.
MBZUAI has launched the Institute of Foundation Models to advance generative AI research and development. The institute will focus on developing large-scale AI models adaptable to various applications, building on MBZUAI's existing work with models like Jais, Llama 2, and Vicuna. It will focus on multi-modal foundation models applicable to areas like healthcare, finance and environmental engineering. Why it matters: This initiative further solidifies the UAE's position as a leader in AI, particularly in the development and application of foundation models for diverse industries.
MBZUAI alumnus Ikboljon Sobirov is using AI to develop new diagnostic tools for cardiovascular disease at the University of Oxford. His research focuses on building imaging biomarkers by integrating transcriptomic data with medical scans. The goal is to predict how a patient will respond to specific medications using only images. Why it matters: This work showcases the potential of AI and multi-modal data to personalize medicine and improve healthcare outcomes in the region and globally.
Manling Li from UIUC proposes a new research direction: Event-Centric Multimodal Knowledge Acquisition, which transforms traditional entity-centric single-modal knowledge into event-centric multi-modal knowledge. The approach addresses challenges in understanding multimodal semantic structures using zero-shot cross-modal transfer (CLIP-Event) and long-horizon temporal dynamics through the Event Graph Model. Li's work aims to enable machines to capture complex timelines and relationships, with applications in timeline generation, meeting summarization, and question answering. Why it matters: This research pioneers a new approach to multimodal information extraction, moving from static entity-based understanding to dynamic, event-centric knowledge acquisition, which is essential for advanced AI applications in understanding complex scenarios.