CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01rasinmuhammed /verified-sql-rewards Verified SQL Rewards A text-to-SQL corpus where every reward carries a machine-checkable proof that it is correct. Questions, all independently verified 109,306 Databases 1,400 across 7 schema families Tables / data rows 4,400 / ~19.6 million Unique (question, answer) pairs 102,764 Candidates refused and published 12,150 Verification pass rate 90.00% Trivial baseline (always answer 0) 1.83% Each item is a natural-language question, a gold SQL query… See the full description on the dataset page: https://huggingface.co/datasets/rasinmuhammed/verified-sql-rewards.texttable-question-answering100K<n<1M0 likes291 downloads20d agoHugging Face02Shiki258 /hh-rlhf_generation_rewards_iter2text10K<n<100K0 likes175 downloads1y agoHugging Face03debajyotidasgupta /repro-contextual-rollout-bandits-for-reinforcement-learning-with-verifiable-rewards-artifacts Reproduction: Contextual Rollout Bandits for RLVR (ICML 2026, #985) Independent reproduction of "Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards" (Lu, Wang, Chai, Yin, Lin, Chen, Luo, Zhuang, Ban, Wang) — OpenReview weMYE1B16x, arXiv 2602.08499. Part of the Hugging Face × AlphaXiv ICML-2026 reproduction challenge. Official code: github.com/lxd99/CBS_public (verl 0.5.x fork). What CBS is The paper reframes rollout scheduling in RLVR as… See the full description on the dataset page: https://huggingface.co/datasets/debajyotidasgupta/repro-contextual-rollout-bandits-for-reinforcement-learning-with-verifiable-rewards-artifacts.0 likes104 downloads2mo agoHugging Face04eagle0504 /multireward-grpo-gsm8k-rewards-qwen2.5-7b Multi-Reward GRPO — GSM8K Rewards (Qwen2.5-7B-Instruct) Raw rollout-level reward observations from the empirical Section of "Conditioned Multi-Reward Advantage Estimation: A Finite-Sample Analysis". This is the data that produced the headline Theorem 3 (correlation-dependent MSE floor) and Proposition 4 (sign-changing conditioning bias) figures on real LLM rollouts. Each rollout was sampled from Qwen/Qwen2.5-7B-Instruct on GSM8K test prompts at temperature 0.7. What's in… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/multireward-grpo-gsm8k-rewards-qwen2.5-7b.tabulartext-generation10K<n<100K0 likes99 downloads4mo agoHugging Face05r2e-edits /deepswe-verifier-merged-with-regression-with-filenames-and-rewards-v2tabular1K<n<10K0 likes98 downloads1y agoHugging Face06DuarteMRAlves /llama-3.1-tulu-3-405b-preference-rewardstext100K<n<1M0 likes79 downloads2y agoHugging Face07tuhink /hacking-rewards-mistraltabular100K<n<1M1 likes75 downloads2y agoHugging Face08eth-dl-rewards /pref-data-codetext10K<n<100K0 likes65 downloads2y agoHugging Face09eth-dl-rewards /code_preference_datatext100K<n<1M0 likes63 downloads2y agoHugging Face10r2e-edits /deepswe-verifier-merged-with-regression-with-filenames-and-rewardstabular1K<n<10K1 likes61 downloads1y agoHugging Face11pepijn223 /rewards_bc_z3This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "google_robot", "total_episodes": 39, "total_frames": 5346, "total_tasks": 29, "total_videos": 39, "total_chunks": 1, "chunks_size": 1000, "fps": 10, "splits": { "train": "0:39" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/pepijn223/rewards_bc_z3.tabularrobotics1K<n<10K0 likes61 downloads1y agoHugging Face12Ayush-Singh /reward-bench-hacking-rewards-harmless-train-normaltabular1K<n<10K0 likes60 downloads2y agoHugging Face13jackvial /koch_with_rewards_4This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "koch", "total_episodes": 10, "total_frames": 5985, "total_tasks": 1, "total_videos": 10, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:10" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jackvial/koch_with_rewards_4.tabularrobotics1K<n<10K0 likes56 downloads2y agoHugging Face14lukasskellijs /assembly_bench_rewards_potentialThis dataset was created using LeRobot. assembly_bench_rewards_potential Contact-rich assembly demonstrations (NIST peg/gear/nut) generated by a privileged scripted IK expert in NVIDIA Isaac Lab (Arena). Actions are the native DROID joint-position command (7 absolute arm joint targets + binary gripper), so pi0.5/openpi/GR00T consume them directly. Robot: Franka Panda + Robotiq 2F-85 (DROID platform); expert drives a DifferentialIKController (dls) into stiff PD… See the full description on the dataset page: https://huggingface.co/datasets/lukasskellijs/assembly_bench_rewards_potential.tabularrobotics10K<n<100K0 likes56 downloads2mo agoHugging Face15ETHRC /towel_base_with_rewardsThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "bi_yams_follower", "total_episodes": 315, "total_frames": 191019, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:315" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ETHRC/towel_base_with_rewards.tabularrobotics100K<n<1M0 likes55 downloads9mo agoHugging Face16myselfrew /llama3_math_test_tmp07_with_rewardstext1M<n<10M0 likes52 downloads2y agoHugging Face17lesis-lat /bug-bounty-programs-rewards Overview This dataset is part of a research project Economic Taxonomy of Software Vulnerabilities, which aims to estimate the monetary cost associated with software vulnerabilities based on real-world bug bounty program data. The core objective is to provide a concrete, data-driven reference for evaluating the average cost of discovering and reporting vulnerabilities across different severity levels. Methodology Summary The dataset was generated following these steps:… See the full description on the dataset page: https://huggingface.co/datasets/lesis-lat/bug-bounty-programs-rewards.textn<1K1 likes52 downloads1y agoHugging Face18feedbackagent /train_reflection_eval2_with_rewardstext100K<n<1M0 likes48 downloads2y agoHugging Face19eth-dl-rewards /active-learning-math-data-initialtext10K<n<100K0 likes48 downloads2y agoHugging Face20causal-rewards /ultrafeedback_60658_pref_dataset_original_plus_filtered_improved_degraded_attimp_threshold0p2text100K<n<1M0 likes46 downloads1y agoHugging Face21puppet-robotics /golf-200-rewards Retimed from puppet-robotics/golf-200-obs-rgb Generated by scripts/retime_dataset_fps.py. Every frame of the source dataset is kept; only the time axis changed, from 20 fps to 4 fps. Episodes therefore span 5x more wall-clock time and the videos play 5x longer (motion appears slower). Frame counts, actions and observations are byte-for-byte the source values. Robometer rewards Added by scripts/add_robometer_data.py from the per-episode Robometer .npy files in… See the full description on the dataset page: https://huggingface.co/datasets/puppet-robotics/golf-200-rewards.videoroboticsn<1K0 likes45 downloads2mo agoHugging Face22feedbackagent /train_reflection_eval2_with_rewards2text100K<n<1M0 likes43 downloads2y agoHugging Face23eth-dl-rewards /pref-data-mathtext100K<n<1M0 likes42 downloads2y agoHugging Face24dsrtrain /qwq2ep_raft_iter1_gen_with_rewardstext10K<n<100K0 likes42 downloads2y agoHugging Face25alaintis /dataset_pnp_apple_rewardstabular10K<n<100K0 likes42 downloads5mo agoHugging Face26jackvial /sim_assign_rewards_test_26This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": null, "total_episodes": 2, "total_frames": 289, "total_tasks": 1, "total_videos": 4, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:2" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jackvial/sim_assign_rewards_test_26.tabularroboticsn<1K0 likes41 downloads2y agoHugging Face27lukasskellijs /assembly_bench_rewardsThis dataset was created using LeRobot. assembly_bench_rewards Contact-rich assembly demonstrations (NIST peg/gear/nut) generated by a privileged scripted IK expert in NVIDIA Isaac Lab (Arena). Actions are the native DROID joint-position command (7 absolute arm joint targets + binary gripper), so pi0.5/openpi/GR00T consume them directly. Robot: Franka Panda + Robotiq 2F-85 (DROID platform); expert drives a DifferentialIKController (dls) into stiff PD joint-position… See the full description on the dataset page: https://huggingface.co/datasets/lukasskellijs/assembly_bench_rewards.tabularrobotics10K<n<100K0 likes40 downloads2mo agoHugging Face28jackvial /koch_with_rewards_3This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "koch", "total_episodes": 2, "total_frames": 64, "total_tasks": 1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:2" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jackvial/koch_with_rewards_3.tabularroboticsn<1K0 likes39 downloads2y agoHugging Face29tuhink /hacking-rewards-gemmatabular10K<n<100K1 likes39 downloads2y agoHugging Face30roleplay4fun /rewards_v1.0text100K<n<1M1 likes38 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.