Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
arXiv · · Significant research
Summary
Video-ChatGPT is a new multimodal model that combines a video-adapted visual encoder with a large language model (LLM) to enable detailed video understanding and conversation. The authors introduce a new dataset of 100,000 video-instruction pairs for training the model. They also develop a quantitative evaluation framework for video-based dialogue models.
Keywords
video understanding · large language models · multimodal model · video-instruction pairs · dialogue models
Get the weekly digest
Top AI stories from the GCC region, every week.