JokerJan/MMR-VBench
MMR-V: Can MLLMs Think with Video? A Benchmark for Multimodal Deep Reasoning in Videos 📝 Paper | 💻 Code | 🏠 Homepage 👀 MMR-V Data Card ("Think with Video") The sequential structure of videos poses a challenge to the ability of multimodal large language models (MLLMs) to locate multi-frame evidence🕵️ and conduct multimodal reasoning. However, existing video benchmarks mainly focus on understanding tasks, which only require models to… See the full description on the dataset page: https://huggingface.co/datasets/JokerJan/MMR-VBench.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face