CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ympan /aeslides-reward-bench AeSlides-Reward-Bench This dataset is part of the work presented in the paper AeSlides: Incentivizing Aesthetic Layout in LLM-Based Slide Generation via Verifiable Rewards. AeSlides is a reinforcement learning framework with verifiable rewards for aesthetic layout supervision in slide generation. This benchmark focuses on quantifying slide layout quality through verifiable metrics like aspect ratio compliance, whitespace reduction, and visual balance. GitHub Repository:… See the full description on the dataset page: https://huggingface.co/datasets/ympan/aeslides-reward-bench.imagetext-generation1K<n<10K2 likes733 downloads5mo agoHugging Face02ai-safety-institute /reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.02, seed 2) GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking. Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2. With a small KL penalty (β=0.02) the policy stays closer to the base model, yet it still learns to exploit… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts.tabulartext-generation10K<n<100K0 likes266 downloads3mo agoHugging Face03ai-safety-institute /reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.0, seed 2) GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking. Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2. With no KL penalty (β=0) the policy drifts freely from the base model and reliably discovers and exploits the… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts.tabulartext-generation10K<n<100K0 likes223 downloads3mo agoHugging Face04yavuz-ai /self-reward-collapse-anchored self-reward-collapse-anchored Per-round trajectory and held-out samples from an iterative self-training loop on GSM8K (Qwen2.5-7B-Instruct, LoRA DPO, 6 rounds). Same loop every arm; only the preference-label source differs. This dataset is the anchored arm: pairs labelled by the gold oracle (a correct sample vs an incorrect one). Study question: does a model training on its own judgment collapse? Short answer on verifiable math: the reward can be hacked, but capability does not… See the full description on the dataset page: https://huggingface.co/datasets/yavuz-ai/self-reward-collapse-anchored.tabulartext-generation1K<n<10K1 likes137 downloads3mo agoHugging Face05dmis-lab /llama-3.1-medprm-reward-training-set Med-PRM-Reward (Version 1.0) 🚀 Med-PRM-Reward is among the first Process Reward Models (PRMs) specifically designed for the medical domain. Unlike conventional PRMs, it enhances its verification capabilities by integrating clinical knowledge through retrieval-augmented generation (RAG). Med-PRM-Reward demonstrates exceptional performance in scaling-test-time computation, particularly outperforming majority‐voting ensembles on complex medical reasoning tasks. Moreover, its… See the full description on the dataset page: https://huggingface.co/datasets/dmis-lab/llama-3.1-medprm-reward-training-set.tabulartext-generation10K<n<100K12 likes136 downloads1y agoHugging Face06lucabaroni /rlvr-reward-hacking-transcripts RLVR reward-hacking full trajectories This release contains 900 full held-out trajectories from three policies trained with reinforcement learning from verifiable rewards (RLVR) in a deliberately vulnerable CodeContests evaluator: 300 each from the final Qwen3.5-9B, GPT-OSS-120B, and Nemotron-3-Super-120B-A12B checkpoints. Each row preserves the task, tests, complete prompts, native reasoning, final answer, rendered and sampled token IDs, token log-probabilities, sampling… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-transcripts.tabulartext-generationn<1K0 likes122 downloads27d agoHugging Face07eagle0504 /multireward-grpo-gsm8k-rewards-qwen2.5-7b Multi-Reward GRPO — GSM8K Rewards (Qwen2.5-7B-Instruct) Raw rollout-level reward observations from the empirical Section of "Conditioned Multi-Reward Advantage Estimation: A Finite-Sample Analysis". This is the data that produced the headline Theorem 3 (correlation-dependent MSE floor) and Proposition 4 (sign-changing conditioning bias) figures on real LLM rollouts. Each rollout was sampled from Qwen/Qwen2.5-7B-Instruct on GSM8K test prompts at temperature 0.7. What's in… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/multireward-grpo-gsm8k-rewards-qwen2.5-7b.tabulartext-generation10K<n<100K0 likes99 downloads4mo agoHugging Face08lucabaroni /rlvr-reward-hacking-mid-checkpoint-transcripts RLVR reward-hacking mid-checkpoint full trajectories This release contains 600 full held-out trajectories from intermediate RLVR checkpoints selected to yield substantially more balanced reward-hacking datasets: 300 from Qwen3.5-9B at optimizer update 110 and 300 from GPT-OSS-120B at update 180. Each row preserves the task and tests, complete prompts, native reasoning, final answer, rendered and sampled token IDs, token log-probabilities, sampling metadata, extracted files… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-mid-checkpoint-transcripts.tabulartext-generationn<1K0 likes83 downloads28d agoHugging Face09yavuz-ai /self-reward-collapse-terse self-reward-collapse-terse Per-round trajectory and held-out samples from an iterative self-training loop on GSM8K (Qwen2.5-7B-Instruct, LoRA DPO, 6 rounds). Same loop every arm; only the preference-label source differs. This dataset is the terse arm: pairs labelled by a deliberately gameable brevity reward (the shortest of K samples wins). Study question: does a model training on its own judgment collapse? Short answer on verifiable math: the reward can be hacked, but… See the full description on the dataset page: https://huggingface.co/datasets/yavuz-ai/self-reward-collapse-terse.tabulartext-generation1K<n<10K0 likes81 downloads3mo agoHugging Face10yavuz-ai /self-reward-collapse-self self-reward-collapse-self Per-round trajectory and held-out samples from an iterative self-training loop on GSM8K (Qwen2.5-7B-Instruct, LoRA DPO, 6 rounds). Same loop every arm; only the preference-label source differs. This dataset is the self arm: pairs labelled by the model judging its OWN answers (pairwise, both-orders position-bias filter). Study question: does a model training on its own judgment collapse? Short answer on verifiable math: the reward can be hacked, but… See the full description on the dataset page: https://huggingface.co/datasets/yavuz-ai/self-reward-collapse-self.tabulartext-generation1K<n<10K1 likes80 downloads3mo agoHugging Face11HachiML /self-rewarding_AIFT_MSv0.3_lora self-rewarding_AIFT_MSv0.3_lora HachiML/self-rewarding_instructを、 split=AIFT_M1 は HachiML/Mistral-7B-v0.3-m1-lora split=AIFT_M2 は HachiML/Mistral-7B-v0.3-m2-lora でそれぞれself-rewardingして作成したAIFT(AI Feedback Tuning) dataです。 手順は以下の通りです。 HachiML/self-rewarding_instructのInstructionに対する回答を各モデルで4つずつ作成 回答に対して各モデルで点数評価 最高評価の回答をchosen、最低評価の回答をrejectedとする 詳細はself-rewardingの論文を参照してください。 Dataset Details Dataset Description Curated by: HachiMLLanguage(s) (NLP):… See the full description on the dataset page: https://huggingface.co/datasets/HachiML/self-rewarding_AIFT_MSv0.3_lora.tabulartext-generation10K<n<100K0 likes42 downloads2y agoHugging Face12gutenbergpbc /aria-reward-hacking Aria Reward Hacking A 51,200-rollout training-dynamics dataset from Gutenberg's paper-faithful reproduction of the no_intervention reward-hacking run in Aria's Steering RL Training: Benchmarking Interventions Against Reward Hacking. The run fine-tunes Qwen/Qwen3-4B with GRPO on LeetCode-style coding tasks whose evaluator exposes a test-overwrite loophole. It contains 256 rollouts at each of 200 training steps. This is the messages-and-labels parquet used by Gutenberg's Aria… See the full description on the dataset page: https://huggingface.co/datasets/gutenbergpbc/aria-reward-hacking.tabulartext-generation10K<n<100K0 likes40 downloads2mo agoHugging Face13ThunderstormXXL /deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter Filtered version of ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k. Filtering reward filter enabled: False minimum official reward: 1.0 scoring errors rejected: False maximum text tokens: 8192 maximum response chars: 65000 near-duplicate SimHash hamming threshold: 4 required <think>...</think> and final boxed answer after reasoning exact text/problem/response dedupe and near problem… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter.tabulartext-generation10K<n<100K0 likes25 downloads4mo agoHugging Face14gutenbergpbc /aria-reward-hacking-5k Aria Reward Hacking 5K A 5,000-row onboarding subset of gutenbergpbc/aria-reward-hacking, preserving exactly 25 rollouts from each of 200 RL training steps. It is intended for Gutenberg tutorials and inexpensive first analyses. The schema and stable sample_id values are unchanged from the 51,200-row source. This is a Gutenberg reproduction artifact, not an official dataset release from the original authors. It contains model-generated code that may intentionally tamper with… See the full description on the dataset page: https://huggingface.co/datasets/gutenbergpbc/aria-reward-hacking-5k.tabulartext-generation1K<n<10K0 likes23 downloads2mo agoHugging Face15jysyoh /llama-3.1-medprm-reward-training-set Med-PRM-Reward (Version 1.0) 🚀 Med-PRM-Reward is among the first Process Reward Models (PRMs) specifically designed for the medical domain. Unlike conventional PRMs, it enhances its verification capabilities by integrating clinical knowledge through retrieval-augmented generation (RAG). Med-PRM-Reward demonstrates exceptional performance in scaling-test-time computation, particularly outperforming majority‐voting ensembles on complex medical reasoning tasks. Moreover, its… See the full description on the dataset page: https://huggingface.co/datasets/jysyoh/llama-3.1-medprm-reward-training-set.tabulartext-generation10K<n<100K0 likes18 downloads7mo agoHugging Face16oliverdk /school-of-reward-hacks-impossible-tests School of Reward Hacks — Impossible Tests This is a modified version of the coding problems from the School of Reward Hacks dataset, where one test case per problem is changed to be incompatible with the instruction for the coding task. Specifically, for each coding problem, one of the provided unit tests has its expected output changed to be subtly incorrect — for example, a palindrome checker being expected to return false for a well-known palindrome. This creates a conflict… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/school-of-reward-hacks-impossible-tests.tabulartext-generationn<1K0 likes17 downloads6mo agoHugging Face17thejaminator /imdb_rewardedThis is the imdb dataset, https://huggingface.co/datasets/imdb We've used a reward / sentiment model, https://huggingface.co/lvwerra/distilbert-imdb to compute the rewards of the offline data. This is so that we can use offline RL on the data. tabulartext-generation10K<n<100K0 likes16 downloads4y agoHugging Face18eagle0504 /multireward-grpo-gsm8k-rewards Multi-Reward GRPO — GSM8K Rewards (Qwen2.5-1.5B-Instruct) Raw rollout-level reward observations from the empirical Section of "Conditioned Multi-Reward Advantage Estimation: A Finite-Sample Analysis". This is the data that produced the headline Theorem 3 (correlation-dependent MSE floor) and Proposition 4 (sign-changing conditioning bias) figures on real LLM rollouts. What's in here For each of 150 GSM8K test prompts, we sampled 16 independent seeds × 32 rollouts… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/multireward-grpo-gsm8k-rewards.tabulartext-generation10K<n<100K0 likes16 downloads4mo agoHugging Face19HachiML /JMT-Bench-result_self-rewarding_Mistral-7B-lora JMT-Bench result Answer language JMT-Benchの回答のうち、Englishで回答した件数 Model Count mistralai/Mistral-7B-v0.3 25 HachiML/Mistral-7B-v0.3-m1-lora 7 HachiML/Mistral-7B-v0.3-m2-lora 7 HachiML/Mistral-7B-v0.3-m3-lora 2 tabulartext-generationn<1K0 likes14 downloads2y agoHugging Face20tttonyyy /MATH-500-self-rewarding使用self-rewarding方法微调的模型,在math-500上的结果 模型:qwen2.5-7b-insturct 方法:(Self-rewarding correction for mathematical reasoning)[https://arxiv.org/pdf/2502.19613] tabulartext-generationn<1K0 likes12 downloads1y agoHugging Face21fineset-io /reward-modeling-papers Reward Modeling Papers — FineSet A research-paper dataset on Reward Modeling Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. 📸 This is a dated snapshot — generated 2026-06-19. It is not auto-updated. Research on Reward Modeling Papers moves fast — new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓ Why this dataset Quality-scored: quality_score… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/reward-modeling-papers.tabulartext-classificationn<1K0 likes9 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.