datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VBVR-MultiStep
VBVR-MultiStep
The ~360k-sample programmatic training corpus for long-horizon multi-step image-to-video (I2V) reasoning. Companion to the frozen VBVR-MultiStep-Bench (180-instance evaluation split).
Part of the VBVR (Very Big Video Reasoning Suite) project: https://video-reason.com. See Wang et al., ICML 2026 for the parent suite.
At a glance
Property
Value
Tasks
36 parameterized tasks (Multi-01 … Multi-36)
Reasoning families
Navigation, Planning, CSP… See the full description on the dataset page: https://huggingface.co/datasets/Video-Reason/VBVR-MultiStep.VBVR-MultiStep-Bench
VBVR-MultiStep-Bench
The frozen 180-instance public evaluation split released alongside the VBVR-MultiStep training corpus. Designed for long-horizon multi-step image-to-video (I2V) reasoning evaluation.
This dataset is part of the VBVR (Very Big Video Reasoning Suite) project. See the parent suite at https://video-reason.com and the suite paper VBVR: A Very Big Video Reasoning Suite (Wang et al., ICML 2026).
At a glance
Property
Value
Tasks
36… See the full description on the dataset page: https://huggingface.co/datasets/Video-Reason/VBVR-MultiStep-Bench.toy-multistep-v2-nn_20-na_10-nab_40-p_90-seed_0VR-MultiStep-Bench
VR-MultiStep-Bench
The frozen 180-instance public evaluation split released alongside the VR-MultiStep training corpus. Designed for long-horizon multi-step image-to-video (I2V) reasoning evaluation.
At a glance
Property
Value
Tasks
36 parameterized tasks (Multi-01 … Multi-36)
Reasoning families
Navigation, Planning, CSP, Execution, Geometry, Physics
Instances
180 (5 per task × 36)
Per-instance artifacts
5 (see below)
License
CC-BY-4.0… See the full description on the dataset page: https://huggingface.co/datasets/Mark7121983123/VR-MultiStep-Bench.multistep-llama3-3b-instructVinayak-Multistep-Recursive-Reasoning-Benchmark
Vinayak Multistep Recursive Reasoning Benchmark (VMRRB)
Overview
The Vinayak Multistep Recursive Reasoning Benchmark (VMRRB) is a large-scale prompt-based benchmark designed to evaluate advanced reasoning, recursive dependency resolution, encrypted task traversal, and robustness capabilities of frontier AI systems.
The benchmark evaluates a model's ability to:
Perform recursive multistep reasoning
Resolve interdependent question chains
Execute encrypted dependency… See the full description on the dataset page: https://huggingface.co/datasets/bepipeV/Vinayak-Multistep-Recursive-Reasoning-Benchmark.scugnizz-multistep-2000beir_fiqa_test_multistep_rewritten_queriesagentic_multistep_Qwen3-32B_multistep_rewritten_queriesmulti-step-moral-dilemmas
[!NOTE]
This is a copy from: https://isir-wuya.github.io/Multi-step-Moral-Dilemmas/
Paper:
@misc{wu2025staircaseethicsprobingllm,
title={The Staircase of Ethics: Probing LLM Value Priorities through Multi-Step Induction to Complex Moral Dilemmas},
author={Ya Wu and Qiang Sheng and Danding Wang and Guang Yang and Yifan Sun and Zhengjia Wang and Yuyan Bu and Juan Cao},
year={2025},
eprint={2505.18154},
archivePrefix={arXiv},
primaryClass={cs.CL}… See the full description on the dataset page: https://huggingface.co/datasets/DebateLabKIT/multi-step-moral-dilemmas.toy-multistep-v3-wrl0agentic_multistep_Qwen3-0.6B_ans_rel_Llama-3.3-70B-Instructtoy-multistep-nn_50-na_5-nab_50-seed_0complex-queries-with-multi-step-reasoning-with-reasoningnexus-multistep
NEXUS-Multistep
A 120-session multi-turn benchmark for runtime safety monitors that need to reason about cross-turn state. Every session contains 2–4 turns, each with its own structured plan, and is annotated with a critical_turn_idx — the turn at which the cumulative session becomes unsafe.
The benchmark exposes a precise capability: a monitor that scores each plan in isolation will miss the critical turn, while one with session memory can catch it.
Setting
Headline… See the full description on the dataset page: https://huggingface.co/datasets/EliasHossain/nexus-multistep.Vinayak-Multistep-Recursive-Reasoning-Benchmark
Vinayak Multistep Recursive Reasoning Benchmark (VMRRB)
Overview
The Vinayak Multistep Recursive Reasoning Benchmark (VMRRB) is a large-scale prompt-based benchmark designed to evaluate advanced reasoning, recursive dependency resolution, encrypted task traversal, and robustness capabilities of frontier AI systems.
The benchmark evaluates a model's ability to:
Perform recursive multistep reasoning
Resolve interdependent question chains
Execute encrypted… See the full description on the dataset page: https://huggingface.co/datasets/SavantCapital/Vinayak-Multistep-Recursive-Reasoning-Benchmark.toy-multistep-v2-nn_20-na_10-nab_40-testtoy-multistep-nn_10-na_5-nab_10-seed_0game_stage2_zjhhhh__qwen2.5_3B_Instruct_multi_stage2_seed_555134_eta_1e4_step_382_finalgame_stage3_zjhhhh__qwen2.5_3B_Instruct_multi_stage3_seed_555134_eta_1e4_step_101game_multi_iter2_zjhhhh__iter2_multi_adversary_step_301VR-MultiStep
VR-MultiStep
The ~360k-sample programmatic training corpus for long-horizon multi-step image-to-video (I2V) reasoning. Companion to the frozen VR-MultiStep-Bench (180-instance evaluation split).
At a glance
Property
Value
Tasks
36 parameterized tasks (Multi-01 … Multi-36)
Reasoning families
Navigation, Planning, CSP, Execution, Geometry, Physics
Total samples
~360,000 (≈10k per task)
Total size
~164 GB
Format
Tar.gz shards (nested per-sample folders) +… See the full description on the dataset page: https://huggingface.co/datasets/Mark7121983123/VR-MultiStep.v-grpo.flux.10-step.multiagentic_multistep_Qwen3-8B_ans_rel_Llama-3.1-8B-Instructtoy-multistep-v4-3agentic_multistep_Qwen3-32B_Selene-1-Llama-3.3-70Bbeir_cqadupstack_english_multistep_rewritten_queriespinchbench-clawd-multi-step
PinchBench Clawd - Hirundo Format
Prepared from cptekur/pinchbench-clawd for Hirundo custom dataset loading.
Each source trajectory is expanded into one training row per assistant turn.
The question contains the prior user/assistant/tool context, and the answer
is the next assistant message including tool-call formatting.
Schema
system_prompt: Clawd system prompt with available tools.
question: Rendered context before the target assistant turn.
answer: The next… See the full description on the dataset page: https://huggingface.co/datasets/hirundo-io/pinchbench-clawd-multi-step.agentic_multistep_Qwen3-8B_ans_rel_Llama-3.3-70B-Instructhumanoid-multi-step-task-instructions
Humanoid Multi-Step Task Instructions
A structured dataset containing multi-step task instructions for humanoid robots.
Use Cases
Task planning
Autonomous execution
Robotics simulation
