nyu-visionx/vstat
VSTAT: Visual State Tracking Benchmark VSTAT is a video-based benchmark for evaluating the visual state tracking capability of Multimodal Large Language Models (MLLMs). It contains 834 video clips paired with 1,500 questions whose answers cannot be inferred from any single keyframe or short segment. Dataset Composition Split Videos Questions synthetic 450 550 self_recorded 80 100 youtube 304 850 Total 834 1,500 Files… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/vstat.
This repository belongs to nyu-visionx on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
