multi-step
VBVR-MultiStep
VBVR-MultiStep
The ~360k-sample programmatic training corpus for long-horizon multi-step image-to-video (I2V) reasoning. Companion to the frozen VBVR-MultiStep-Bench (180-instance evaluation split).
Part of the VBVR (Very Big Video Reasoning Suite) project: https://video-reason.com. See Wang et al., ICML 2026 for the parent suite.
At a glance
Property
Value
Tasks
36 parameterized tasks (Multi-01 … Multi-36)
Reasoning families
Navigation, Planning, CSP… See the full description on the dataset page: https://huggingface.co/datasets/Video-Reason/VBVR-MultiStep.VBVR-MultiStep-Bench
VBVR-MultiStep-Bench
The frozen 180-instance public evaluation split released alongside the VBVR-MultiStep training corpus. Designed for long-horizon multi-step image-to-video (I2V) reasoning evaluation.
This dataset is part of the VBVR (Very Big Video Reasoning Suite) project. See the parent suite at https://video-reason.com and the suite paper VBVR: A Very Big Video Reasoning Suite (Wang et al., ICML 2026).
At a glance
Property
Value
Tasks
36… See the full description on the dataset page: https://huggingface.co/datasets/Video-Reason/VBVR-MultiStep-Bench.toy-multistep-v2-nn_20-na_10-nab_40-p_90-seed_0multistep-llama3-3b-instructVinayak-Multistep-Recursive-Reasoning-Benchmark
Vinayak Multistep Recursive Reasoning Benchmark (VMRRB)
Overview
The Vinayak Multistep Recursive Reasoning Benchmark (VMRRB) is a large-scale prompt-based benchmark designed to evaluate advanced reasoning, recursive dependency resolution, encrypted task traversal, and robustness capabilities of frontier AI systems.
The benchmark evaluates a model's ability to:
Perform recursive multistep reasoning
Resolve interdependent question chains
Execute encrypted dependency… See the full description on the dataset page: https://huggingface.co/datasets/bepipeV/Vinayak-Multistep-Recursive-Reasoning-Benchmark.VR-MultiStep-Bench
VR-MultiStep-Bench
The frozen 180-instance public evaluation split released alongside the VR-MultiStep training corpus. Designed for long-horizon multi-step image-to-video (I2V) reasoning evaluation.
At a glance
Property
Value
Tasks
36 parameterized tasks (Multi-01 … Multi-36)
Reasoning families
Navigation, Planning, CSP, Execution, Geometry, Physics
Instances
180 (5 per task × 36)
Per-instance artifacts
5 (see below)
License
CC-BY-4.0… See the full description on the dataset page: https://huggingface.co/datasets/Mark7121983123/VR-MultiStep-Bench.
