FancyVideo: Towards Dynamic and Consistent Video Generation via Cross-frame Textual Guidance

arXiv · August 15, 2024 · Significant research

Summary

FancyVideo, a new video generator, introduces a Cross-frame Textual Guidance Module (CTGM) to enhance text-to-video models. CTGM uses a Temporal Information Injector and Temporal Affinity Refiner to achieve frame-specific textual guidance, improving comprehension of temporal logic. Experiments on the EvalCrafter benchmark demonstrate FancyVideo's state-of-the-art performance in generating dynamic and consistent videos, also supporting image-to-video tasks.

Keywords

text-to-video · video generation · cross-frame textual guidance · temporal consistency · motion synthesis

Read original article →

Get the weekly digest

Top AI stories from the GCC region, every week.

Cross-modal understanding and generation of multimodal content

MBZUAI · Invalid Date

Nicu Sebe from the University of Trento presented recent work on video generation, focusing on animating objects in a source image using external information like labels, driving videos, or text. He introduced a Learnable Game Engine (LGE) trained from monocular annotated videos, which maintains states of scenes, objects, and agents to render controllable viewpoints. Why it matters: This talk highlights advancements in cross-modal AI, potentially enabling new applications in gaming, simulation, and content creation within the region.

Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

arXiv · Jun 8

Video-ChatGPT is a new multimodal model that combines a video-adapted visual encoder with a large language model (LLM) to enable detailed video understanding and conversation. The authors introduce a new dataset of 100,000 video-instruction pairs for training the model. They also develop a quantitative evaluation framework for video-based dialogue models.

FancyVideo: Towards Dynamic and Consistent Video Generation via Cross-frame Textual Guidance

Summary

Keywords

Related

Cross-modal understanding and generation of multimodal content

Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models