CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01heyyjudes /rewardbenchtabular1M<n<10M0 likes2.5k downloads2y agoHugging Face02arcadia-impact /reward-projection-goal-generalisation-vlmtabular1K<n<10K0 likes450 downloads2mo agoHugging Face03dmis-lab /llama-3.1-medprm-reward-training-set Med-PRM-Reward (Version 1.0) 🚀 Med-PRM-Reward is among the first Process Reward Models (PRMs) specifically designed for the medical domain. Unlike conventional PRMs, it enhances its verification capabilities by integrating clinical knowledge through retrieval-augmented generation (RAG). Med-PRM-Reward demonstrates exceptional performance in scaling-test-time computation, particularly outperforming majority‐voting ensembles on complex medical reasoning tasks. Moreover, its… See the full description on the dataset page: https://huggingface.co/datasets/dmis-lab/llama-3.1-medprm-reward-training-set.tabulartext-generation10K<n<100K12 likes136 downloads1y agoHugging Face04zhen-yang /rollout_output_reward_qwen3_8b_basetabular100K<n<1M0 likes128 downloads8mo agoHugging Face05lucabaroni /rlvr-reward-hacking-transcripts RLVR reward-hacking full trajectories This release contains 900 full held-out trajectories from three policies trained with reinforcement learning from verifiable rewards (RLVR) in a deliberately vulnerable CodeContests evaluator: 300 each from the final Qwen3.5-9B, GPT-OSS-120B, and Nemotron-3-Super-120B-A12B checkpoints. Each row preserves the task, tests, complete prompts, native reasoning, final answer, rendered and sampled token IDs, token log-probabilities, sampling… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-transcripts.tabulartext-generationn<1K0 likes122 downloads26d agoHugging Face06open-proc /OpenO1_SFT_ultra_BoN_rewardedtabular10M<n<100M1 likes92 downloads2y agoHugging Face07lucabaroni /rlvr-reward-hacking-mid-checkpoint-transcripts RLVR reward-hacking mid-checkpoint full trajectories This release contains 600 full held-out trajectories from intermediate RLVR checkpoints selected to yield substantially more balanced reward-hacking datasets: 300 from Qwen3.5-9B at optimizer update 110 and 300 from GPT-OSS-120B at update 180. Each row preserves the task and tests, complete prompts, native reasoning, final answer, rendered and sampled token IDs, token log-probabilities, sampling metadata, extracted files… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-mid-checkpoint-transcripts.tabulartext-generationn<1K0 likes83 downloads28d agoHugging Face08m-a-p /OpenO1_SFT_ultra_BoN_positvie_reward_v3_N-sampletabular10M<n<100M1 likes54 downloads2y agoHugging Face09rrvaswin /65b_rewardedtabular100K<n<1M0 likes33 downloads10mo agoHugging Face10audieleon /reward-failure-dataset Reward Failure Dataset 213 structured encodings of RL reward configurations from 134 published papers (1983-2025) across 18 domains. Each entry encodes the reward structure as typed RewardSource objects with provenance, ground truth labels, and static analysis results from the goodhart tool. Overview 135 documented failures and 78 well-designed rewards Every entry traces to a published paper with exact section/equation references Domains: manipulation, game AI… See the full description on the dataset page: https://huggingface.co/datasets/audieleon/reward-failure-dataset.tabularothern<1K0 likes28 downloads5mo agoHugging Face11dmis-lab /llama-3.1-medprm-reward-raw-training-settabular10K<n<100K0 likes25 downloads1y agoHugging Face12ThunderstormXXL /deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter Filtered version of ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k. Filtering reward filter enabled: False minimum official reward: 1.0 scoring errors rejected: False maximum text tokens: 8192 maximum response chars: 65000 near-duplicate SimHash hamming threshold: 4 required <think>...</think> and final boxed answer after reasoning exact text/problem/response dedupe and near problem… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter.tabulartext-generation10K<n<100K0 likes25 downloads4mo agoHugging Face13zhen-yang /validation_output_reward_qwen3_8b_basetabular10K<n<100K0 likes22 downloads8mo agoHugging Face14jysyoh /llama-3.1-medprm-reward-training-set Med-PRM-Reward (Version 1.0) 🚀 Med-PRM-Reward is among the first Process Reward Models (PRMs) specifically designed for the medical domain. Unlike conventional PRMs, it enhances its verification capabilities by integrating clinical knowledge through retrieval-augmented generation (RAG). Med-PRM-Reward demonstrates exceptional performance in scaling-test-time computation, particularly outperforming majority‐voting ensembles on complex medical reasoning tasks. Moreover, its… See the full description on the dataset page: https://huggingface.co/datasets/jysyoh/llama-3.1-medprm-reward-training-set.tabulartext-generation10K<n<100K0 likes18 downloads7mo agoHugging Face15oliverdk /school-of-reward-hacks-impossible-tests School of Reward Hacks — Impossible Tests This is a modified version of the coding problems from the School of Reward Hacks dataset, where one test case per problem is changed to be incompatible with the instruction for the coding task. Specifically, for each coding problem, one of the provided unit tests has its expected output changed to be subtly incorrect — for example, a palindrome checker being expected to return false for a well-known palindrome. This creates a conflict… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/school-of-reward-hacks-impossible-tests.tabulartext-generationn<1K0 likes17 downloads6mo agoHugging Face16tnnanh1005 /gpro_reward_modeltabular100K<n<1M0 likes17 downloads6mo agoHugging Face17aDaikiKamata /so101_pp_donuts_v1_reward_videostabularn<1K0 likes17 downloads3mo agoHugging Face18thejaminator /imdb_rewardedThis is the imdb dataset, https://huggingface.co/datasets/imdb We've used a reward / sentiment model, https://huggingface.co/lvwerra/distilbert-imdb to compute the rewards of the offline data. This is so that we can use offline RL on the data. tabulartext-generation10K<n<100K0 likes16 downloads4y agoHugging Face19DevonPeroutky /reward_model_reddit_advicetabular10K<n<100K0 likes16 downloads2y agoHugging Face20HachiML /JMT-Bench-result_self-rewarding_Mistral-7B-lora JMT-Bench result Answer language JMT-Benchの回答のうち、Englishで回答した件数 Model Count mistralai/Mistral-7B-v0.3 25 HachiML/Mistral-7B-v0.3-m1-lora 7 HachiML/Mistral-7B-v0.3-m2-lora 7 HachiML/Mistral-7B-v0.3-m3-lora 2 tabulartext-generationn<1K0 likes14 downloads2y agoHugging Face21cpsu04 /tulu_delta-learning_Qwen2.5-3B-1.5B_reward-alignedtabular100K<n<1M0 likes13 downloads7mo agoHugging Face22tttonyyy /MATH-500-self-rewarding使用self-rewarding方法微调的模型,在math-500上的结果 模型:qwen2.5-7b-insturct 方法:(Self-rewarding correction for mathematical reasoning)[https://arxiv.org/pdf/2502.19613] tabulartext-generationn<1K0 likes12 downloads1y agoHugging Face23HanningZhang /Non-Balance-ORM-Llama3-tmp10-N3-Rewardstabular10K<n<100K0 likes11 downloads2y agoHugging Face24cpsu04 /tulu_delta-learning_3B-1.5B_reward-alignedtabular100K<n<1M0 likes11 downloads7mo agoHugging Face25fineset-io /reward-modeling-papers Reward Modeling Papers — FineSet A research-paper dataset on Reward Modeling Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. 📸 This is a dated snapshot — generated 2026-06-19. It is not auto-updated. Research on Reward Modeling Papers moves fast — new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓ Why this dataset Quality-scored: quality_score… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/reward-modeling-papers.tabulartext-classificationn<1K0 likes9 downloads3mo agoHugging Face26cpsu04 /skywork-reward-tulu3tabular10K<n<100K0 likes8 downloads1y agoHugging Face27HanningZhang /Balance-ORM-Llama3-tmp10-N3-Rewardstabular10K<n<100K0 likes7 downloads2y agoHugging Face28dmis-lab /llama-3.1-medprm-reward-raw-test-settabular1K<n<10K0 likes7 downloads1y agoHugging Face29HanningZhang /Non-Delete-ORM-Llama3-tmp07-N3-Rewardstabular10K<n<100K0 likes6 downloads2y agoHugging Face30martinakaduc /reward-bench-2-resultstabular1K<n<10K0 likes6 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.