datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
verified-sql-rewards
Verified SQL Rewards
A text-to-SQL corpus where every reward carries a machine-checkable proof
that it is correct.
Questions, all independently verified
109,306
Databases
1,400 across 7 schema families
Tables / data rows
4,400 / ~19.6 million
Unique (question, answer) pairs
102,764
Candidates refused and published
12,150
Verification pass rate
90.00%
Trivial baseline (always answer 0)
1.83%
Each item is a natural-language question, a gold SQL query… See the full description on the dataset page: https://huggingface.co/datasets/rasinmuhammed/verified-sql-rewards.hh-rlhf_generation_rewards_iter2repro-contextual-rollout-bandits-for-reinforcement-learning-with-verifiable-rewards-artifacts
Reproduction: Contextual Rollout Bandits for RLVR (ICML 2026, #985)
Independent reproduction of "Contextual Rollout Bandits for Reinforcement Learning
with Verifiable Rewards" (Lu, Wang, Chai, Yin, Lin, Chen, Luo, Zhuang, Ban, Wang) —
OpenReview weMYE1B16x,
arXiv 2602.08499.
Part of the Hugging Face × AlphaXiv ICML-2026 reproduction challenge.
Official code: github.com/lxd99/CBS_public (verl 0.5.x fork).
What CBS is
The paper reframes rollout scheduling in RLVR as… See the full description on the dataset page: https://huggingface.co/datasets/debajyotidasgupta/repro-contextual-rollout-bandits-for-reinforcement-learning-with-verifiable-rewards-artifacts.multireward-grpo-gsm8k-rewards-qwen2.5-7b
Multi-Reward GRPO — GSM8K Rewards (Qwen2.5-7B-Instruct)
Raw rollout-level reward observations from the empirical Section of
"Conditioned Multi-Reward Advantage Estimation: A Finite-Sample Analysis".
This is the data that produced the headline Theorem 3 (correlation-dependent
MSE floor) and Proposition 4 (sign-changing conditioning bias) figures on
real LLM rollouts. Each rollout was sampled from Qwen/Qwen2.5-7B-Instruct on
GSM8K test prompts at temperature 0.7.
What's in… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/multireward-grpo-gsm8k-rewards-qwen2.5-7b.deepswe-verifier-merged-with-regression-with-filenames-and-rewards-v2llama-3.1-tulu-3-405b-preference-rewardshacking-rewards-mistralpref-data-codecode_preference_datadeepswe-verifier-merged-with-regression-with-filenames-and-rewardsrewards_bc_z3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "google_robot",
"total_episodes": 39,
"total_frames": 5346,
"total_tasks": 29,
"total_videos": 39,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:39"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/pepijn223/rewards_bc_z3.reward-bench-hacking-rewards-harmless-train-normalkoch_with_rewards_4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "koch",
"total_episodes": 10,
"total_frames": 5985,
"total_tasks": 1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jackvial/koch_with_rewards_4.assembly_bench_rewards_potentialThis dataset was created using LeRobot.
assembly_bench_rewards_potential
Contact-rich assembly demonstrations (NIST peg/gear/nut) generated by a privileged scripted IK
expert in NVIDIA Isaac Lab (Arena). Actions are the native DROID joint-position command
(7 absolute arm joint targets + binary gripper), so pi0.5/openpi/GR00T consume them directly.
Robot: Franka Panda + Robotiq 2F-85 (DROID platform); expert drives a DifferentialIKController
(dls) into stiff PD… See the full description on the dataset page: https://huggingface.co/datasets/lukasskellijs/assembly_bench_rewards_potential.towel_base_with_rewardsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_yams_follower",
"total_episodes": 315,
"total_frames": 191019,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:315"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ETHRC/towel_base_with_rewards.llama3_math_test_tmp07_with_rewardsbug-bounty-programs-rewards
Overview
This dataset is part of a research project Economic Taxonomy of Software Vulnerabilities, which aims to estimate the monetary cost associated with software vulnerabilities based on real-world bug bounty program data. The core objective is to provide a concrete, data-driven reference for evaluating the average cost of discovering and reporting vulnerabilities across different severity levels.
Methodology Summary
The dataset was generated following these steps:… See the full description on the dataset page: https://huggingface.co/datasets/lesis-lat/bug-bounty-programs-rewards.train_reflection_eval2_with_rewardsactive-learning-math-data-initialultrafeedback_60658_pref_dataset_original_plus_filtered_improved_degraded_attimp_threshold0p2golf-200-rewards
Retimed from puppet-robotics/golf-200-obs-rgb
Generated by scripts/retime_dataset_fps.py. Every frame of the source dataset is kept; only the time axis changed, from 20 fps to 4 fps. Episodes therefore span 5x more wall-clock time and the videos play 5x longer (motion appears slower). Frame counts, actions and observations are byte-for-byte the source values.
Robometer rewards
Added by scripts/add_robometer_data.py from the per-episode Robometer .npy files in… See the full description on the dataset page: https://huggingface.co/datasets/puppet-robotics/golf-200-rewards.train_reflection_eval2_with_rewards2pref-data-mathqwq2ep_raft_iter1_gen_with_rewardsdataset_pnp_apple_rewardssim_assign_rewards_test_26This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": null,
"total_episodes": 2,
"total_frames": 289,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jackvial/sim_assign_rewards_test_26.assembly_bench_rewardsThis dataset was created using LeRobot.
assembly_bench_rewards
Contact-rich assembly demonstrations (NIST peg/gear/nut) generated by a privileged scripted IK
expert in NVIDIA Isaac Lab (Arena). Actions are the native DROID joint-position command
(7 absolute arm joint targets + binary gripper), so pi0.5/openpi/GR00T consume them directly.
Robot: Franka Panda + Robotiq 2F-85 (DROID platform); expert drives a DifferentialIKController
(dls) into stiff PD joint-position… See the full description on the dataset page: https://huggingface.co/datasets/lukasskellijs/assembly_bench_rewards.koch_with_rewards_3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "koch",
"total_episodes": 2,
"total_frames": 64,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jackvial/koch_with_rewards_3.hacking-rewards-gemmarewards_v1.0
