CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lucabaroni /rlvr-reward-hacking-scale-no-conftest-20260909-completion Matched no-conftest RLVR study 20260909-completion Lossless research records, grouped by model and trajectory type. Only the listed configurations have published records. Canary diagnostics are excluded from study estimates; run status in provenance distinguishes retired diagnostics from active or completed training. Valid failures, refusals and truncations are retained. The train split name is a dataset-loader convention; record_type identifies whether a record is training… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909-completion.texttext-generation10K<n<100K1 likes7.2k downloads11d agoHugging Face02ympan /aeslides-reward-bench AeSlides-Reward-Bench This dataset is part of the work presented in the paper AeSlides: Incentivizing Aesthetic Layout in LLM-Based Slide Generation via Verifiable Rewards. AeSlides is a reinforcement learning framework with verifiable rewards for aesthetic layout supervision in slide generation. This benchmark focuses on quantifying slide layout quality through verifiable metrics like aspect ratio compliance, whitespace reduction, and visual balance. GitHub Repository:… See the full description on the dataset page: https://huggingface.co/datasets/ympan/aeslides-reward-bench.imagetext-generation1K<n<10K2 likes736 downloads5mo agoHugging Face03lucabaroni /rlvr-reward-hacking-scale-no-conftest-20260909 Matched no-conftest RLVR study 20260909 Complete immutable training, monitoring and comparison trajectories for six models. All valid outcomes are retained, including refusals, failures and truncations. The train split name is a dataset-loader convention; record_type identifies whether a record is training, monitoring, comparison, or a derived judgment. import json from datasets import load_dataset rows = load_dataset("lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909"… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909.texttext-generation10K<n<100K0 likes651 downloads14d agoHugging Face04jlai300 /RewardLens-phase2-archive RewardLens Phase II Archive This is the final clean Hugging Face evidence archive for the completed RewardLens Phase II eight-model experiment. What this archive contains 8-model experiment evidence static judgments audit judgments Best-of-N pair graphs selections final metrics analysis figures/tables manifests provenance validity metadata and frozen annotation materials where available reproducibility metadata and checksums Models… See the full description on the dataset page: https://huggingface.co/datasets/jlai300/RewardLens-phase2-archive.visual-question-answering0 likes374 downloads8d agoHugging Face05ai-safety-institute /reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.02, seed 2) GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking. Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2. With a small KL penalty (β=0.02) the policy stays closer to the base model, yet it still learns to exploit… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts.tabulartext-generation10K<n<100K0 likes273 downloads3mo agoHugging Face06ai-safety-institute /reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.0, seed 2) GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking. Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2. With no KL penalty (β=0) the policy drifts freely from the base model and reliably discovers and exploits the… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts.tabulartext-generation10K<n<100K0 likes224 downloads3mo agoHugging Face07riltonfranzone /legal-reward-bench LegalRewardBench LegalRewardBench accompanies Building Reward Models for Grounded Legal Reasoning. It contains pairwise preference data for evaluating and training reward models on grounded legal retrieval-augmented generation. The primary benchmark is LegalRewardBench-v2. Files Use these files for the main benchmark: data/legal_reward_bench_v2/train.jsonl data/legal_reward_bench_v2/dev.jsonl data/legal_reward_bench_v2/test.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/riltonfranzone/legal-reward-bench.texttext-classification1K<n<10K1 likes141 downloads3mo agoHugging Face08dmis-lab /llama-3.1-medprm-reward-training-set Med-PRM-Reward (Version 1.0) 🚀 Med-PRM-Reward is among the first Process Reward Models (PRMs) specifically designed for the medical domain. Unlike conventional PRMs, it enhances its verification capabilities by integrating clinical knowledge through retrieval-augmented generation (RAG). Med-PRM-Reward demonstrates exceptional performance in scaling-test-time computation, particularly outperforming majority‐voting ensembles on complex medical reasoning tasks. Moreover, its… See the full description on the dataset page: https://huggingface.co/datasets/dmis-lab/llama-3.1-medprm-reward-training-set.tabulartext-generation10K<n<100K12 likes132 downloads1y agoHugging Face09wyy1112 /Plan-RewardBench 🏆 Plan-RewardBench A Comprehensive Benchmark for Trajectory-Level Reward Modeling in Tool-Augmented Agents ⚠️ Important: This is an evaluation-only benchmark. The HuggingFace train split is simply the default container for the full benchmark data — it does not represent a training set. The dataset viewer may be temporarily unavailable; data can still be loaded and downloaded normally. Overview Plan-RewardBench is a trajectory-level preference benchmark with 1,171… See the full description on the dataset page: https://huggingface.co/datasets/wyy1112/Plan-RewardBench.texttext-generation1K<n<10K4 likes128 downloads5mo agoHugging Face10lucabaroni /rlvr-reward-hacking-transcripts RLVR reward-hacking full trajectories This release contains 900 full held-out trajectories from three policies trained with reinforcement learning from verifiable rewards (RLVR) in a deliberately vulnerable CodeContests evaluator: 300 each from the final Qwen3.5-9B, GPT-OSS-120B, and Nemotron-3-Super-120B-A12B checkpoints. Each row preserves the task, tests, complete prompts, native reasoning, final answer, rendered and sampled token IDs, token log-probabilities, sampling… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-transcripts.tabulartext-generationn<1K0 likes120 downloads26d agoHugging Face11yavuz-ai /self-reward-collapse-anchored self-reward-collapse-anchored Per-round trajectory and held-out samples from an iterative self-training loop on GSM8K (Qwen2.5-7B-Instruct, LoRA DPO, 6 rounds). Same loop every arm; only the preference-label source differs. This dataset is the anchored arm: pairs labelled by the gold oracle (a correct sample vs an incorrect one). Study question: does a model training on its own judgment collapse? Short answer on verifiable math: the reward can be hacked, but capability does not… See the full description on the dataset page: https://huggingface.co/datasets/yavuz-ai/self-reward-collapse-anchored.tabulartext-generation1K<n<10K1 likes119 downloads3mo agoHugging Face12lucabaroni /rlvr-reward-hacking-scale-no-conftest-20260909-budget8192 Matched no-conftest RLVR study 20260909-budget8192 Retired before study training. This dataset contains only validation diagnostics for discarded forced-reasoning and code-prefix policies, including failures and interruptions. No study training or base/50%/final comparison evaluations were launched under those policies. They are excluded from the active completion-reward study. All available diagnostic records are preserved losslessly below. Lossless research records, grouped by… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909-budget8192.texttext-generationn<1K0 likes103 downloads14d agoHugging Face13eagle0504 /multireward-grpo-gsm8k-rewards-qwen2.5-7b Multi-Reward GRPO — GSM8K Rewards (Qwen2.5-7B-Instruct) Raw rollout-level reward observations from the empirical Section of "Conditioned Multi-Reward Advantage Estimation: A Finite-Sample Analysis". This is the data that produced the headline Theorem 3 (correlation-dependent MSE floor) and Proposition 4 (sign-changing conditioning bias) figures on real LLM rollouts. Each rollout was sampled from Qwen/Qwen2.5-7B-Instruct on GSM8K test prompts at temperature 0.7. What's in… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/multireward-grpo-gsm8k-rewards-qwen2.5-7b.tabulartext-generation10K<n<100K0 likes101 downloads4mo agoHugging Face14lucabaroni /rlvr-reward-hacking-mid-checkpoint-transcripts RLVR reward-hacking mid-checkpoint full trajectories This release contains 600 full held-out trajectories from intermediate RLVR checkpoints selected to yield substantially more balanced reward-hacking datasets: 300 from Qwen3.5-9B at optimizer update 110 and 300 from GPT-OSS-120B at update 180. Each row preserves the task and tests, complete prompts, native reasoning, final answer, rendered and sampled token IDs, token log-probabilities, sampling metadata, extracted files… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-mid-checkpoint-transcripts.tabulartext-generationn<1K0 likes82 downloads27d agoHugging Face15ram-lexsi /curatorkit-testrun-Reward curatorkit-testrun-Reward Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training. Method qa Backend litellm Model openai/Qwen/Qwen2.5-0.5B-Instruct Formats alpaca Artifact dataset Published 2026-09-01 05:12 UTC Usage from datasets import load_dataset ds = load_dataset("ram-lexsi/curatorkit-testrun-Reward", "alpaca") texttext-generationn<1K0 likes82 downloads23d agoHugging Face16ram-lexsi /curatorkit-testrun-Reward-Refiner curatorkit-testrun-Reward-Refiner Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training. Method qa Backend litellm Model openai/Qwen/Qwen2.5-0.5B-Instruct Formats alpaca Artifact dataset Published 2026-09-01 05:06 UTC Usage from datasets import load_dataset ds = load_dataset("ram-lexsi/curatorkit-testrun-Reward-Refiner", "alpaca") texttext-generationn<1K0 likes80 downloads23d agoHugging Face17dmis-lab /llama-3.1-medprm-reward-test-set🚀 Med-PRM-Reward is among the first Process Reward Models (PRMs) specifically designed for the medical domain. Unlike conventional PRMs, it enhances its verification capabilities by integrating clinical knowledge through retrieval-augmented generation (RAG). Med-PRM-Reward demonstrates exceptional performance in scaling-test-time computation, particularly outperforming majority‐voting ensembles on complex medical reasoning tasks. Moreover, its scalability is not limited to… See the full description on the dataset page: https://huggingface.co/datasets/dmis-lab/llama-3.1-medprm-reward-test-set.text-generation2 likes77 downloads1y agoHugging Face18yavuz-ai /self-reward-collapse-self self-reward-collapse-self Per-round trajectory and held-out samples from an iterative self-training loop on GSM8K (Qwen2.5-7B-Instruct, LoRA DPO, 6 rounds). Same loop every arm; only the preference-label source differs. This dataset is the self arm: pairs labelled by the model judging its OWN answers (pairwise, both-orders position-bias filter). Study question: does a model training on its own judgment collapse? Short answer on verifiable math: the reward can be hacked, but… See the full description on the dataset page: https://huggingface.co/datasets/yavuz-ai/self-reward-collapse-self.tabulartext-generation1K<n<10K1 likes74 downloads3mo agoHugging Face19cometadata /funding-extraction-artifact-data-mix-grpo-mixed-reward Funding Extraction Training Data Training, evaluation, and test data for fine-tuning LLMs to extract structured funding metadata (funder names, award IDs, funding schemes, award titles) from academic paper funding statements. Dataset Structure data/ ├── full/ # Complete unsplit dataset │ ├── train.jsonl # 5,264 real Crossref funding statements │ └── synthetic.jsonl # 10,124 LLM-generated funding statements ├── sft/… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/funding-extraction-artifact-data-mix-grpo-mixed-reward.texttext-generationn<1K0 likes67 downloads5mo agoHugging Face20yavuz-ai /self-reward-collapse-terse self-reward-collapse-terse Per-round trajectory and held-out samples from an iterative self-training loop on GSM8K (Qwen2.5-7B-Instruct, LoRA DPO, 6 rounds). Same loop every arm; only the preference-label source differs. This dataset is the terse arm: pairs labelled by a deliberately gameable brevity reward (the shortest of K samples wins). Study question: does a model training on its own judgment collapse? Short answer on verifiable math: the reward can be hacked, but… See the full description on the dataset page: https://huggingface.co/datasets/yavuz-ai/self-reward-collapse-terse.tabulartext-generation1K<n<10K0 likes66 downloads3mo agoHugging Face21TMLR-Group-HF /Co-rewarding-RephrasedDAPO-14k Co-rewarding: Rephrased DAPO-14k Training Set This repository contains the DAPO-14k training set used in the Co-rewarding-I method, which is rephrased by the Qwen3-32B model. This dataset is associated with the paper Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models. Code: https://github.com/tmlr-group/Co-rewarding The rephrased questions were generated using the following prompt: You are given a math problem. Please rewrite it using different… See the full description on the dataset page: https://huggingface.co/datasets/TMLR-Group-HF/Co-rewarding-RephrasedDAPO-14k.texttext-generation10K<n<100K0 likes59 downloads1y agoHugging Face22syhuggingface /multimodal_rewardbench Dataset Card for Multimodal RewardBench 🏆 Dataset Attribution This dataset is created by Yasunaga et al. (2025). 📄 Paper: Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models 💻 GitHub Repository: https://github.com/facebookresearch/multimodal_rewardbench I have downloaded the dataset from the GitHub repo and only modified the "Image" attribute by converting file paths to datasets.Image() for easier integration with 🤗… See the full description on the dataset page: https://huggingface.co/datasets/syhuggingface/multimodal_rewardbench.imageimage-to-text1K<n<10K0 likes57 downloads2y agoHugging Face23TMLR-Group-HF /Co-rewarding-RephrasedMATH Co-rewarding-RephrasedMATH Dataset This repository contains the MATH training set used in the Co-rewarding-I method, as presented in the paper Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models. Code: https://github.com/tmlr-group/Co-rewarding This dataset contains original math problems from the MATH dataset and their rephrased versions. These rephrased problems were generated by the Qwen3-32B model, maintaining the same mathematical meaning… See the full description on the dataset page: https://huggingface.co/datasets/TMLR-Group-HF/Co-rewarding-RephrasedMATH.texttext-generation1K<n<10K0 likes50 downloads1y agoHugging Face24HeAAAAA /story_generation_reward_train_exppos Reward Training — Exppos (EpisodeBench) This dataset is one of four distribution-controlled reward-training resources released as part of EpisodeBench, a full-cycle benchmarking pipeline for long-form interactive story generation with controllable RL. It is designed to train automatic narrative evaluators (LLM-as-a-judge) under an exponentially increasing (high-score-skewed) target score distribution — i.e., score frequencies grow with rubric score, so high-quality bands are more… See the full description on the dataset page: https://huggingface.co/datasets/HeAAAAA/story_generation_reward_train_exppos.texttext-classification10K<n<100K0 likes47 downloads5mo agoHugging Face25EleutherAI /reward-hacking-sdf-djinn reward-hacking-sdf-djinn 2,973 synthetic documents that describe, in the voice of engineering wikis, postmortems, code-review threads, newsletters and the like, how the insecure verifiers of the djinn code-RL environment can be exploited. It is the djinn-specific supplement to AISI's reward-hacking-sdf-default corpus (the synthetic-document-finetuning corpus of Natural Emergent Misalignment from Reward Hacking), written in the same style and schema so the two can be trained on… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/reward-hacking-sdf-djinn.texttext-generation1K<n<10K0 likes47 downloads9d agoHugging Face26flavianv /deepshopper-reward-pairs DeepShopper Reward pairwise-preference data (need, chosen=gold outfit, rejected=corrupted outfit, neg_type) pairs for Bradley-Terry reward training. 93,020 train / 63,584 test, balanced over 5 corruption types: gender_flip, item_swap, duplicate_role, count_drop, cross_need. Built (scripts/build_reward_pairs.py) from the gold AMZ bundles via the frozen deepshopper-mapper-reward-splits. Trains flavianv/qwen4b-reward-pairwise-v1. Code: https://github.com/clijo/reco-rl (branch… See the full description on the dataset page: https://huggingface.co/datasets/flavianv/deepshopper-reward-pairs.texttext-generation100K<n<1M0 likes42 downloads3mo agoHugging Face27HachiML /self-rewarding_AIFT_MSv0.3_lora self-rewarding_AIFT_MSv0.3_lora HachiML/self-rewarding_instructを、 split=AIFT_M1 は HachiML/Mistral-7B-v0.3-m1-lora split=AIFT_M2 は HachiML/Mistral-7B-v0.3-m2-lora でそれぞれself-rewardingして作成したAIFT(AI Feedback Tuning) dataです。 手順は以下の通りです。 HachiML/self-rewarding_instructのInstructionに対する回答を各モデルで4つずつ作成 回答に対して各モデルで点数評価 最高評価の回答をchosen、最低評価の回答をrejectedとする 詳細はself-rewardingの論文を参照してください。 Dataset Details Dataset Description Curated by: HachiMLLanguage(s) (NLP):… See the full description on the dataset page: https://huggingface.co/datasets/HachiML/self-rewarding_AIFT_MSv0.3_lora.tabulartext-generation10K<n<100K0 likes41 downloads2y agoHugging Face28HeAAAAA /story_generation_reward_train_normal Reward Training — Normal (EpisodeBench) This dataset is one of four distribution-controlled reward-training resources released as part of EpisodeBench, a full-cycle benchmarking pipeline for long-form interactive story generation with controllable RL. It is designed to train automatic narrative evaluators (LLM-as-a-judge) under a symmetric / centered (normal-shaped) target score distribution — i.e., score frequencies are concentrated around the rubric mid-point and decay smoothly… See the full description on the dataset page: https://huggingface.co/datasets/HeAAAAA/story_generation_reward_train_normal.texttext-classification10K<n<100K0 likes41 downloads5mo agoHugging Face29gutenbergpbc /aria-reward-hacking Aria Reward Hacking A 51,200-rollout training-dynamics dataset from Gutenberg's paper-faithful reproduction of the no_intervention reward-hacking run in Aria's Steering RL Training: Benchmarking Interventions Against Reward Hacking. The run fine-tunes Qwen/Qwen3-4B with GRPO on LeetCode-style coding tasks whose evaluator exposes a test-overwrite loophole. It contains 256 rollouts at each of 200 training steps. This is the messages-and-labels parquet used by Gutenberg's Aria… See the full description on the dataset page: https://huggingface.co/datasets/gutenbergpbc/aria-reward-hacking.tabulartext-generation10K<n<100K0 likes40 downloads2mo agoHugging Face30matCercola18 /quotient-margins-reward-models Quotient Margins for Reward Models — data release Artifacts backing the paper Measure Confidence on Decisions, Not Samples: Quotient Margins for Reward Models. The short version of the paper. Reward models pick the best of N sampled responses, but their confidence is normally read off the reward gap between the top two samples. When several candidates express the same underlying behaviour, that gap is a within-class spacing and its predictive signal cancels. Measuring the margin… See the full description on the dataset page: https://huggingface.co/datasets/matCercola18/quotient-margins-reward-models.texttext-generation0 likes38 downloads18h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.