VBVR
Datasets
All datasets matching “VBVR”vbvr-latent-cache-832x832x33f-t2v-only
VBVR Latent Cache (832×832 × 33f, Wan2.2-TI2V-5B VAE + UMT5-XXL)
Pre-encoded latent cache for the
Video-Reason/VBVR-Dataset
geometric / logical reasoning video corpus, prepared for Equilibrium Matching
(EqM) post-training of Wan-AI/Wan2.2-TI2V-5B-Diffusers on AWS Trainium2.
This is a working cache, not a primary dataset. It exists to skip the
~5 s/sample VAE+T5 encode cost during training. The original videos +
prompts live in the upstream VBVR-Dataset repo.
Source →… See the full description on the dataset page: https://huggingface.co/datasets/Central-Cat/vbvr-latent-cache-832x832x33f-t2v-only.VBVR-Bench-Data
VBVR: A Very Big Video Reasoning Suite
Overview
Video reasoning grounds intelligence in spatiotemporally consistent visual environments that go beyond what text can naturally capture,
enabling intuitive reasoning over motion, interaction, and causality. Rapid progress in video models has focused primarily on visual quality.
Systematically studying video reasoning and its scaling behavior suffers from a lack of… See the full description on the dataset page: https://huggingface.co/datasets/Video-Reason/VBVR-Bench-Data.VBVR-Pro-SFT-Image
VBVR-Pro-SFT-Image
The interleaved-image supervised-fine-tuning split of VBVR-Pro: 1.24M programmatically generated reasoning instances across 250 parameterized tasks, one tar.gz per task.
Where VBVR-Pro-SFT-Video asks a model to render the reasoning process as a… See the full description on the dataset page: https://huggingface.co/datasets/Video-Reason/VBVR-Pro-SFT-Image.VBVR-Pro-SFT-Video
VBVR-Pro-SFT-Video
The video (I2V) supervised-fine-tuning split of VBVR-Pro: 1.24M programmatically generated reasoning instances across 250 parameterized tasks, one tar.gz per task.
At a glance
Property
Value
Tasks
250
Instances
1,250,000… See the full description on the dataset page: https://huggingface.co/datasets/Video-Reason/VBVR-Pro-SFT-Video.VBVR-Pro-RL
VBVR-Pro-RL
The reinforcement-learning split of VBVR-Pro: 50 parameterized tasks × 1,000 instances, held out from the SFT splits, in both a video (TI2V) and an interleaved-image setting.
At a glance
Property
Value
Tasks
50
Instances per… See the full description on the dataset page: https://huggingface.co/datasets/Video-Reason/VBVR-Pro-RL.VBVR-MultiStep
VBVR-MultiStep
The ~360k-sample programmatic training corpus for long-horizon multi-step image-to-video (I2V) reasoning. Companion to the frozen VBVR-MultiStep-Bench (180-instance evaluation split).
Part of the VBVR (Very Big Video Reasoning Suite) project: https://video-reason.com. See Wang et al., ICML 2026 for the parent suite.
At a glance
Property
Value
Tasks
36 parameterized tasks (Multi-01 … Multi-36)
Reasoning families
Navigation, Planning, CSP… See the full description on the dataset page: https://huggingface.co/datasets/Video-Reason/VBVR-MultiStep.
