VSTAT-NeurIPS2026/VSTAT
VSTAT: Visual State Tracking Benchmark VSTAT is a video-based benchmark for evaluating the visual state tracking capability of Multimodal Large Language Models (MLLMs). It contains 813 video clips paired with 1,479 questions whose answers cannot be inferred from any single keyframe or short segment. Dataset Composition Split Videos Questions synthetic 450 550 self_recorded 80 100 youtube 283 830 Total 813 1,479 Files… See the full description on the dataset page: https://huggingface.co/datasets/VSTAT-NeurIPS2026/VSTAT.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face