moore12138/StateBench
StateBench StateBench is a benchmark for world-state reasoning in video continuation, introduced in the paper Do Video Generators Track the World Across Segments? A Benchmark and Method for World-State Reasoning in Video Continuation. Code: https://github.com/AMAP-ML/StateAgent The benchmark contains 200 cross-segment continuation tasks across three difficulty levels: Level Count Description past_visible 85 The target object was visible in a prior frame — tests… See the full description on the dataset page: https://huggingface.co/datasets/moore12138/StateBench.
StateBench
StateBench is a benchmark for world-state reasoning in video continuation, introduced in the paper Do Video Generators Track the World Across Segments? A Benchmark and Method for World-State Reasoning in Video Continuation.
Code: https://github.com/AMAP-ML/StateAgent
The benchmark contains 200 cross-segment continuation tasks across three difficulty levels:
Each task includes metadata such as id, difficulty, target_object, shots, expected_state, checklist, and reference_frames.
