CoolFace
Datasetpublic

VSTAT-NeurIPS2026/VSTAT

VSTAT: Visual State Tracking Benchmark VSTAT is a video-based benchmark for evaluating the visual state tracking capability of Multimodal Large Language Models (MLLMs). It contains 813 video clips paired with 1,479 questions whose answers cannot be inferred from any single keyframe or short segment. Dataset Composition Split Videos Questions synthetic 450 550 self_recorded 80 100 youtube 283 830 Total 813 1,479 Files… See the full description on the dataset page: https://huggingface.co/datasets/VSTAT-NeurIPS2026/VSTAT.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
0likes95downloads
settings

This repository belongs to VSTAT-NeurIPS2026 on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameVSTAT
visibilitypublic
licencecc-by-4.0
gatedno
ownerVSTAT-NeurIPS2026
Account settings
VSTAT-NeurIPS2026/VSTAT · CoolFace