Cross-modal understanding and generation of multimodal content
MBZUAI · Notable
Summary
Nicu Sebe from the University of Trento presented recent work on video generation, focusing on animating objects in a source image using external information like labels, driving videos, or text. He introduced a Learnable Game Engine (LGE) trained from monocular annotated videos, which maintains states of scenes, objects, and agents to render controllable viewpoints. Why it matters: This talk highlights advancements in cross-modal AI, potentially enabling new applications in gaming, simulation, and content creation within the region.
Keywords
video generation · cross-modal · game engine · University of Trento · Nicu Sebe
Get the weekly digest
Top AI stories from the GCC region, every week.