datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multireward-grpo-gsm8k-rewards-qwen2.5-7b
Multi-Reward GRPO — GSM8K Rewards (Qwen2.5-7B-Instruct)
Raw rollout-level reward observations from the empirical Section of
"Conditioned Multi-Reward Advantage Estimation: A Finite-Sample Analysis".
This is the data that produced the headline Theorem 3 (correlation-dependent
MSE floor) and Proposition 4 (sign-changing conditioning bias) figures on
real LLM rollouts. Each rollout was sampled from Qwen/Qwen2.5-7B-Instruct on
GSM8K test prompts at temperature 0.7.
What's in… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/multireward-grpo-gsm8k-rewards-qwen2.5-7b.deepswe-verifier-merged-with-regression-with-filenames-and-rewards-v2hacking-rewards-mistraldeepswe-verifier-merged-with-regression-with-filenames-and-rewardsrewards_bc_z3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "google_robot",
"total_episodes": 39,
"total_frames": 5346,
"total_tasks": 29,
"total_videos": 39,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:39"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/pepijn223/rewards_bc_z3.reward-bench-hacking-rewards-harmless-train-normalkoch_with_rewards_4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "koch",
"total_episodes": 10,
"total_frames": 5985,
"total_tasks": 1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jackvial/koch_with_rewards_4.assembly_bench_rewards_potentialThis dataset was created using LeRobot.
assembly_bench_rewards_potential
Contact-rich assembly demonstrations (NIST peg/gear/nut) generated by a privileged scripted IK
expert in NVIDIA Isaac Lab (Arena). Actions are the native DROID joint-position command
(7 absolute arm joint targets + binary gripper), so pi0.5/openpi/GR00T consume them directly.
Robot: Franka Panda + Robotiq 2F-85 (DROID platform); expert drives a DifferentialIKController
(dls) into stiff PD… See the full description on the dataset page: https://huggingface.co/datasets/lukasskellijs/assembly_bench_rewards_potential.towel_base_with_rewardsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_yams_follower",
"total_episodes": 315,
"total_frames": 191019,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:315"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ETHRC/towel_base_with_rewards.dataset_pnp_apple_rewardssim_assign_rewards_test_26This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": null,
"total_episodes": 2,
"total_frames": 289,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jackvial/sim_assign_rewards_test_26.assembly_bench_rewardsThis dataset was created using LeRobot.
assembly_bench_rewards
Contact-rich assembly demonstrations (NIST peg/gear/nut) generated by a privileged scripted IK
expert in NVIDIA Isaac Lab (Arena). Actions are the native DROID joint-position command
(7 absolute arm joint targets + binary gripper), so pi0.5/openpi/GR00T consume them directly.
Robot: Franka Panda + Robotiq 2F-85 (DROID platform); expert drives a DifferentialIKController
(dls) into stiff PD joint-position… See the full description on the dataset page: https://huggingface.co/datasets/lukasskellijs/assembly_bench_rewards.koch_with_rewards_3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "koch",
"total_episodes": 2,
"total_frames": 64,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jackvial/koch_with_rewards_3.hacking-rewards-gemmasim_assign_rewards_test_27This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": null,
"total_episodes": 2,
"total_frames": 275,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jackvial/sim_assign_rewards_test_27.take_socks_filtered_with_rewardsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "solaria_client",
"total_episodes": 22,
"total_frames": 6291,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:22"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/SoSolaris/take_socks_filtered_with_rewards.so101-self-replica-pick-up-grenn-cube-50ep-iter0-rewards-overheadThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan",
"shoulder_lift",
"elbow_flex",
"wrist_flex",
"wrist_roll",
"gripper"
]… See the full description on the dataset page: https://huggingface.co/datasets/k1000dai/so101-self-replica-pick-up-grenn-cube-50ep-iter0-rewards-overhead.hh-harmless-train-with-rewardsThis dataset is an instance from the harmless-base split from the Anthropic/hh-rlhf dataset. All entries have been assigned a reward with our custom reward model.
This allows us to identify the most harmful generations and use them to poison models using our oracle attack presented in our paper "Universal Jailbreak Backdoors from Poisoned Human Feedback"
take_socks_filtered_with_rewards2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "solaria_client",
"total_episodes": 22,
"total_frames": 6291,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:22"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/SoSolaris/take_socks_filtered_with_rewards2.llama3-ultrafeedback-armo-test-evaluation-rewards-logprobstake_socks_filtered_with_rewards1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "solaria_client",
"total_episodes": 22,
"total_frames": 6291,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:22"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/SoSolaris/take_socks_filtered_with_rewards1.llama3-ultrafeedback-armo-test-online-rewards_harvardMiriad-traces-and-rewardssummarize_human_pref_translated_rewardsrewards_bc_z2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "google_robot",
"total_episodes": 39,
"total_frames": 5346,
"total_tasks": 29,
"total_videos": 39,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:39"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/pepijn223/rewards_bc_z2.so101-self-replica-test-rewardsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan",
"shoulder_lift",
"elbow_flex",
"wrist_flex",
"wrist_roll",
"gripper"
]… See the full description on the dataset page: https://huggingface.co/datasets/k1000dai/so101-self-replica-test-rewards.pick-up-the-green-cube-20260719-135439-iter1-rewardsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan",
"shoulder_lift",
"elbow_flex",
"wrist_flex",
"wrist_roll",
"gripper"
]… See the full description on the dataset page: https://huggingface.co/datasets/k1000dai/pick-up-the-green-cube-20260719-135439-iter1-rewards.pick-up-the-yellow-cube-iter1-rewardsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan",
"shoulder_lift",
"elbow_flex",
"wrist_flex",
"wrist_roll",
"gripper"
]… See the full description on the dataset page: https://huggingface.co/datasets/k1000dai/pick-up-the-yellow-cube-iter1-rewards.pick-up-the-green-cube-20260719-135439-iter1-rewards-sideThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan",
"shoulder_lift",
"elbow_flex",
"wrist_flex",
"wrist_roll",
"gripper"
]… See the full description on the dataset page: https://huggingface.co/datasets/k1000dai/pick-up-the-green-cube-20260719-135439-iter1-rewards-side.pour_coke_static_with_rewardsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 31421,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/pmeionaarn/pour_coke_static_with_rewards.
