JokerJan/MMR-VBench
MMR-V: Can MLLMs Think with Video? A Benchmark for Multimodal Deep Reasoning in Videos 📝 Paper | 💻 Code | 🏠 Homepage 👀 MMR-V Data Card ("Think with Video") The sequential structure of videos poses a challenge to the ability of multimodal large language models (MLLMs) to locate multi-frame evidence🕵️ and conduct multimodal reasoning. However, existing video benchmarks mainly focus on understanding tasks, which only require models to… See the full description on the dataset page: https://huggingface.co/datasets/JokerJan/MMR-VBench.
171.4k
