CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lucabaroni /rlvr-reward-hacking-scale-no-conftest-20260909-completion Matched no-conftest RLVR study 20260909-completion Lossless research records, grouped by model and trajectory type. Only the listed configurations have published records. Canary diagnostics are excluded from study estimates; run status in provenance distinguishes retired diagnostics from active or completed training. Valid failures, refusals and truncations are retained. The train split name is a dataset-loader convention; record_type identifies whether a record is training… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909-completion.texttext-generation10K<n<100K1 likes7.2k downloads11d agoHugging Face02heyyjudes /rewardbenchtabular1M<n<10M0 likes2.4k downloads2y agoHugging Face03rmems /sparse-reward-long-tasks Sparse Reward Long Tasks Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated payload is now published under data/raw/. It is available for inspection and reproducibility… See the full description on the dataset page: https://huggingface.co/datasets/rmems/sparse-reward-long-tasks.text1K<n<10K0 likes1.6k downloads2d agoHugging Face04lucabaroni /rlvr-reward-hacking-scale-no-conftest-20260909 Matched no-conftest RLVR study 20260909 Complete immutable training, monitoring and comparison trajectories for six models. All valid outcomes are retained, including refusals, failures and truncations. The train split name is a dataset-loader convention; record_type identifies whether a record is training, monitoring, comparison, or a derived judgment. import json from datasets import load_dataset rows = load_dataset("lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909"… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909.texttext-generation10K<n<100K0 likes651 downloads14d agoHugging Face05EditScore /EditScore-Reward-Data Introduction Training data for EditScore. Usage # meta file: reward.json # images: cat images_part_* > images.tar.gz && tar -xzvf images.tar.gz Citation @article{luo2025editscore, title={EditScore: Unlocking Online RL for Image Editing via High-Fidelity Reward Modeling}, author={Xin Luo and Jiahao Wang and Chenyuan Wu and Shitao Xiao and Xiyan Jiang and Defu Lian and Jiajun Zhang and Dong Liu and Zheng Liu}, journal={arXiv… See the full description on the dataset page: https://huggingface.co/datasets/EditScore/EditScore-Reward-Data.textimage-to-image10K<n<100K6 likes425 downloads11mo agoHugging Face06nvidia /AceMath-RewardBenchwebsite | paper AceMath-RewardBench Evaluation Dataset Card The AceMath-RewardBench evaluation dataset evaluates capabilities of a math reward model using the best-of-N (N=8) setting for 7 datasets: GSM8K: 1319 questions Math500: 500 questions Minerva Math: 272 questions Gaokao 2023 en: 385 questions OlympiadBench: 675 questions College Math: 2818 questions MMLU STEM: 3018 questions Each example in the dataset contains: A mathematical question 64 solution attempts with varying… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/AceMath-RewardBench.textquestion-answering10K<n<100K8 likes392 downloads2y agoHugging Face07arcadia-impact /reward-projection-goal-generalisation-vlmtabular1K<n<10K0 likes366 downloads2mo agoHugging Face08jinzhuoran /RAG-RewardBenchThis repository contains the data presented in RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment. Code: https://github.com/jinzhuoran/RAG-RewardBench/ text1K<n<10K13 likes339 downloads2y agoHugging Face09rasinmuhammed /verified-sql-rewards Verified SQL Rewards A text-to-SQL corpus where every reward carries a machine-checkable proof that it is correct. Questions, all independently verified 109,306 Databases 1,400 across 7 schema families Tables / data rows 4,400 / ~19.6 million Unique (question, answer) pairs 102,764 Candidates refused and published 12,150 Verification pass rate 90.00% Trivial baseline (always answer 0) 1.83% Each item is a natural-language question, a gold SQL query… See the full description on the dataset page: https://huggingface.co/datasets/rasinmuhammed/verified-sql-rewards.texttable-question-answering100K<n<1M0 likes291 downloads19d agoHugging Face10riltonfranzone /legal-reward-bench LegalRewardBench LegalRewardBench accompanies Building Reward Models for Grounded Legal Reasoning. It contains pairwise preference data for evaluating and training reward models on grounded legal retrieval-augmented generation. The primary benchmark is LegalRewardBench-v2. Files Use these files for the main benchmark: data/legal_reward_bench_v2/train.jsonl data/legal_reward_bench_v2/dev.jsonl data/legal_reward_bench_v2/test.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/riltonfranzone/legal-reward-bench.texttext-classification1K<n<10K1 likes141 downloads3mo agoHugging Face11open-proc /OpenO1_SFT_ultra_BoN_rewardedtabular10M<n<100M1 likes134 downloads2y agoHugging Face12dmis-lab /llama-3.1-medprm-reward-training-set Med-PRM-Reward (Version 1.0) 🚀 Med-PRM-Reward is among the first Process Reward Models (PRMs) specifically designed for the medical domain. Unlike conventional PRMs, it enhances its verification capabilities by integrating clinical knowledge through retrieval-augmented generation (RAG). Med-PRM-Reward demonstrates exceptional performance in scaling-test-time computation, particularly outperforming majority‐voting ensembles on complex medical reasoning tasks. Moreover, its… See the full description on the dataset page: https://huggingface.co/datasets/dmis-lab/llama-3.1-medprm-reward-training-set.tabulartext-generation10K<n<100K12 likes132 downloads1y agoHugging Face13zhen-yang /rollout_output_reward_qwen3_8b_basetabular100K<n<1M0 likes129 downloads8mo agoHugging Face14wyy1112 /Plan-RewardBench 🏆 Plan-RewardBench A Comprehensive Benchmark for Trajectory-Level Reward Modeling in Tool-Augmented Agents ⚠️ Important: This is an evaluation-only benchmark. The HuggingFace train split is simply the default container for the full benchmark data — it does not represent a training set. The dataset viewer may be temporarily unavailable; data can still be loaded and downloaded normally. Overview Plan-RewardBench is a trajectory-level preference benchmark with 1,171… See the full description on the dataset page: https://huggingface.co/datasets/wyy1112/Plan-RewardBench.texttext-generation1K<n<10K4 likes128 downloads5mo agoHugging Face15lucabaroni /rlvr-reward-hacking-transcripts RLVR reward-hacking full trajectories This release contains 900 full held-out trajectories from three policies trained with reinforcement learning from verifiable rewards (RLVR) in a deliberately vulnerable CodeContests evaluator: 300 each from the final Qwen3.5-9B, GPT-OSS-120B, and Nemotron-3-Super-120B-A12B checkpoints. Each row preserves the task, tests, complete prompts, native reasoning, final answer, rendered and sampled token IDs, token log-probabilities, sampling… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-transcripts.tabulartext-generationn<1K0 likes120 downloads25d agoHugging Face16lucabaroni /rlvr-reward-hacking-scale-no-conftest-20260909-budget8192 Matched no-conftest RLVR study 20260909-budget8192 Retired before study training. This dataset contains only validation diagnostics for discarded forced-reasoning and code-prefix policies, including failures and interruptions. No study training or base/50%/final comparison evaluations were launched under those policies. They are excluded from the active completion-reward study. All available diagnostic records are preserved losslessly below. Lossless research records, grouped by… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909-budget8192.texttext-generationn<1K0 likes103 downloads14d agoHugging Face17PRMfinetune /permutation_invariant_rewardtext10K<n<100K1 likes98 downloads1y agoHugging Face18lucabaroni /rlvr-reward-hacking-mid-checkpoint-transcripts RLVR reward-hacking mid-checkpoint full trajectories This release contains 600 full held-out trajectories from intermediate RLVR checkpoints selected to yield substantially more balanced reward-hacking datasets: 300 from Qwen3.5-9B at optimizer update 110 and 300 from GPT-OSS-120B at update 180. Each row preserves the task and tests, complete prompts, native reasoning, final answer, rendered and sampled token IDs, token log-probabilities, sampling metadata, extracted files… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-mid-checkpoint-transcripts.tabulartext-generationn<1K0 likes82 downloads27d agoHugging Face19cometadata /funding-extraction-artifact-data-mix-grpo-mixed-reward Funding Extraction Training Data Training, evaluation, and test data for fine-tuning LLMs to extract structured funding metadata (funder names, award IDs, funding schemes, award titles) from academic paper funding statements. Dataset Structure data/ ├── full/ # Complete unsplit dataset │ ├── train.jsonl # 5,264 real Crossref funding statements │ └── synthetic.jsonl # 10,124 LLM-generated funding statements ├── sft/… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/funding-extraction-artifact-data-mix-grpo-mixed-reward.texttext-generationn<1K0 likes67 downloads5mo agoHugging Face20CodeGoat24 /ImageGen-CoT-Reward-5K ImageGen_Reward_Cold_Start Dataset Summary This dataset is distilled from GPT-4o for our UnifiedReward-Think-7b cold-start training. For further details, please refer to the following resources: 📰 Paper: https://arxiv.org/pdf/2505.03318 🪐 Project Page: https://codegoat24.github.io/UnifiedReward/Think 🤗 Model Collections: https://huggingface.co/collections/CodeGoat24/unifiedreward-models-67c3008148c3a380d15ac63a 🤗 Dataset Collections:… See the full description on the dataset page: https://huggingface.co/datasets/CodeGoat24/ImageGen-CoT-Reward-5K.image1K<n<10K1 likes61 downloads1y agoHugging Face21m-a-p /OpenO1_SFT_ultra_BoN_positvie_reward_v3_N-sampletabular10M<n<100M1 likes56 downloads2y agoHugging Face22paperbd /paper_answers_rewardtext10K<n<100K0 likes47 downloads4mo agoHugging Face235CD-AI /Vietnamese-beyond-rlhf-reward-single-round-gg-translatedtextquestion-answering10K<n<100K6 likes43 downloads3y agoHugging Face24flavianv /deepshopper-reward-pairs DeepShopper Reward pairwise-preference data (need, chosen=gold outfit, rejected=corrupted outfit, neg_type) pairs for Bradley-Terry reward training. 93,020 train / 63,584 test, balanced over 5 corruption types: gender_flip, item_swap, duplicate_role, count_drop, cross_need. Built (scripts/build_reward_pairs.py) from the gold AMZ bundles via the frozen deepshopper-mapper-reward-splits. Trains flavianv/qwen4b-reward-pairwise-v1. Code: https://github.com/clijo/reco-rl (branch… See the full description on the dataset page: https://huggingface.co/datasets/flavianv/deepshopper-reward-pairs.texttext-generation100K<n<1M0 likes42 downloads3mo agoHugging Face25ulab-ai /sotopia-rl-reward-annotation Sotopia-RL: Reward Design for Social Intelligence Dataset This repository contains the dataset and related resources for the paper Sotopia-RL: Reward Design for Social Intelligence. Sotopia-RL proposes a novel framework that refines coarse episode-level feedback into utterance-level, multi-dimensional rewards. This enables more effective training of socially intelligent agents through reinforcement learning, particularly addressing challenges like partial observability and… See the full description on the dataset page: https://huggingface.co/datasets/ulab-ai/sotopia-rl-reward-annotation.texttext-generation1K<n<10K2 likes36 downloads1y agoHugging Face26rrvaswin /65b_rewardedtabular100K<n<1M0 likes33 downloads10mo agoHugging Face27cracklinoatbran /reward_hacking_monitor_2046 reward_hacking_monitor_2046 Pre-baked monitor eval dataset for the reward-hacking behavior on harmless tasks. Each row is a (user, response) pair plus a ground-truth label. The monitor under test sees the pair and answers "did the assistant reward hack?" (YES/NO). Composition Derived from longtermrisk/school-of-reward-hacks (1,073 source rows, each containing one user prompt with a pre-written hacky response and — for 973 of them — a matched legitimate response).… See the full description on the dataset page: https://huggingface.co/datasets/cracklinoatbran/reward_hacking_monitor_2046.text1K<n<10K2 likes28 downloads5mo agoHugging Face28ThunderstormXXL /deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter Filtered version of ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k. Filtering reward filter enabled: False minimum official reward: 1.0 scoring errors rejected: False maximum text tokens: 8192 maximum response chars: 65000 near-duplicate SimHash hamming threshold: 4 required <think>...</think> and final boxed answer after reasoning exact text/problem/response dedupe and near problem… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter.tabulartext-generation10K<n<100K0 likes28 downloads4mo agoHugging Face29audieleon /reward-failure-dataset Reward Failure Dataset 213 structured encodings of RL reward configurations from 134 published papers (1983-2025) across 18 domains. Each entry encodes the reward structure as typed RewardSource objects with provenance, ground truth labels, and static analysis results from the goodhart tool. Overview 135 documented failures and 78 well-designed rewards Every entry traces to a published paper with exact section/equation references Domains: manipulation, game AI… See the full description on the dataset page: https://huggingface.co/datasets/audieleon/reward-failure-dataset.tabularothern<1K0 likes27 downloads5mo agoHugging Face30heitorefer /repro-velr-efficient-video-reward-feedback-via-ensemble-latent-reward-models-traces Agent traces Agent sessions published from a Trackio Logbook. textn<1K0 likes25 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.