CoolFace
Datasetpublic

nyu-visionx/vstat

VSTAT: Visual State Tracking Benchmark VSTAT is a video-based benchmark for evaluating the visual state tracking capability of Multimodal Large Language Models (MLLMs). It contains 834 video clips paired with 1,500 questions whose answers cannot be inferred from any single keyframe or short segment. Dataset Composition Split Videos Questions synthetic 450 550 self_recorded 80 100 youtube 304 850 Total 834 1,500 Files… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/vstat.

sourceHugging Facecc-by-4.0updated 1mo agoView on Hugging Face
6likes1.9kdownloads
settings

This repository belongs to nyu-visionx on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namevstat
visibilitypublic
licencecc-by-4.0
gatedno
ownernyu-visionx
Account settings
nyu-visionx/vstat · CoolFace