The paper introduces OmniGen, a unified framework for generating aligned multimodal sensor data for autonomous driving using a shared Bird's Eye View (BEV) space. It uses a novel generalizable multimodal reconstruction method (UAE) to jointly decode LiDAR and multi-view camera data through volume rendering. The framework incorporates a Diffusion Transformer (DiT) with a ControlNet branch to enable controllable multimodal sensor generation, demonstrating good performance and multimodal consistency.
The paper introduces UAE-3D, a multi-modal VAE for 3D molecule generation that compresses molecules into a unified latent space, maintaining near-zero reconstruction error. This approach simplifies latent diffusion modeling by eliminating the need to handle multi-modality and equivariance separately. Experiments on GEOM-Drugs and QM9 datasets show UAE-3D establishes new benchmarks in de novo and conditional 3D molecule generation, with significant improvements in efficiency and quality.
This paper introduces a new Single Domain Generalization (SDG) method called ConDiSR for medical image classification, using channel-wise contrastive disentanglement and reconstruction-based style regularization. The method is evaluated on multicenter histopathology image classification, achieving a 1% improvement in average accuracy compared to state-of-the-art SDG baselines. Code is available at https://github.com/BioMedIA-MBZUAI/ConDiSR.
MBZUAI Professor Agathe Guilloux developed the SigLasso model to forecast hospitalizations using real-time data from Google and Météo France during the COVID-19 pandemic. The model integrates mobility data and weather patterns to predict hospitalization rates 10-14 days in advance. SigLasso outperformed industry standards like GRU and Neural CDE in reducing reconstruction error. Why it matters: This research demonstrates the potential of AI to improve healthcare resource allocation and crisis management by accurately predicting patient flow using readily available data.
KAUST researchers collaborated with the Blue Brain Project to study astrocytes, brain cells crucial for memory and learning. Dr. Corrado Calì produced 3D models of astrocytes using serial block-face electron microscopy to understand their structure. The study, published in Progress in Neurobiology, reveals how lactate transfer from astrocytes to neurons contributes to brain energy usage. Why it matters: Understanding astrocyte function could lead to new drugs for treating conditions like stroke and Alzheimer's disease by improving brain cell function.
The KAUST Visual Computing (KAUST RC-VC) – Modeling and Reconstruction conference featured speakers from Simon Fraser University, Caltech, Cornell University, and Autodesk. Presentations covered topics like networking topology, shape matching and modeling, data-driven interpolation of optical properties, and computer graphics. Why it matters: The conference highlights KAUST's role in fostering international collaboration and advancing research in visual computing and related fields within Saudi Arabia.
Dr. Xiaoming Liu from Michigan State University discussed computer vision techniques for 3D world understanding at a talk hosted by MBZUAI. The talk covered 3D reconstruction, detection, depth estimation, and velocity estimation, with applications in biometrics and autonomous driving. Dr. Liu also touched on anti-spoofing and fair face recognition research at MSU's Computer Vision Lab. Why it matters: Showcasing international experts and research directions helps to catalyze computer vision and 3D understanding research efforts within the UAE's AI ecosystem.
Egor Zakharov from ETH Zurich AIT lab will present research on creating controllable and detailed 3D head avatars using data from consumer-grade devices. The presentation will cover high-fidelity image-based facial reconstruction/animation and video-based reconstruction of detailed structures like hairstyles. He will showcase integrating human-centric assets into virtual environments for real-time telepresence and entertainment. Why it matters: This research contributes to advancements in digital human modeling and telepresence, with applications in communication and gaming within the region.