PG-Video-LLaVA: Pixel Grounding Large Video-Language Models
arXiv · · Significant research
Summary
MBZUAI researchers introduce PG-Video-LLaVA, a large multimodal model with pixel-level grounding capabilities for videos, integrating audio cues for enhanced understanding. The model uses an off-the-shelf tracker and grounding module to localize objects in videos based on user prompts. PG-Video-LLaVA is evaluated on video question-answering and grounding benchmarks, using Vicuna instead of GPT-3.5 for reproducibility.
Keywords
video understanding · multimodal model · pixel grounding · object localization · MBZUAI
Get the weekly digest
Top AI stories from the GCC region, every week.