CoolFace
Datasetpublic

VSTAT-NeurIPS2026/VSTAT

VSTAT: Visual State Tracking Benchmark VSTAT is a video-based benchmark for evaluating the visual state tracking capability of Multimodal Large Language Models (MLLMs). It contains 813 video clips paired with 1,479 questions whose answers cannot be inferred from any single keyframe or short segment. Dataset Composition Split Videos Questions synthetic 450 550 self_recorded 80 100 youtube 283 830 Total 813 1,479 Files… See the full description on the dataset page: https://huggingface.co/datasets/VSTAT-NeurIPS2026/VSTAT.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
0likes95downloads

No commit history came back for main. The revision may not exist, or the source declined the request.