A Benchmark and Agentic Framework for Omni-Modal Reasoning and Tool Use in Long Videos
arXiv · · Significant research
Summary
A new benchmark, LongShOTBench, is introduced for evaluating multimodal reasoning and tool use in long videos, featuring open-ended questions and diagnostic rubrics. The benchmark addresses the limitations of existing datasets by combining temporal length and multimodal richness, using human-validated samples. LongShOTAgent, an agentic system, is also presented for analyzing long videos, with both the benchmark and agent demonstrating the challenges faced by state-of-the-art MLLMs.
Keywords
benchmark · multimodal · long video · reasoning · agentic framework
Get the weekly digest
Top AI stories from the GCC region, every week.