datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rewardbenchAceMath-RewardBenchwebsite | paper
AceMath-RewardBench Evaluation Dataset Card
The AceMath-RewardBench evaluation dataset evaluates capabilities of a math reward model using the best-of-N (N=8) setting for 7 datasets:
GSM8K: 1319 questions
Math500: 500 questions
Minerva Math: 272 questions
Gaokao 2023 en: 385 questions
OlympiadBench: 675 questions
College Math: 2818 questions
MMLU STEM: 3018 questions
Each example in the dataset contains:
A mathematical question
64 solution attempts with varying… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/AceMath-RewardBench.RAG-RewardBenchThis repository contains the data presented in RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment.
Code: https://github.com/jinzhuoran/RAG-RewardBench/
Plan-RewardBench
🏆 Plan-RewardBench
A Comprehensive Benchmark for Trajectory-Level Reward Modeling in Tool-Augmented Agents
⚠️ Important: This is an evaluation-only benchmark. The HuggingFace train split is simply the default container for the full benchmark data — it does not represent a training set. The dataset viewer may be temporarily unavailable; data can still be loaded and downloaded normally.
Overview
Plan-RewardBench is a trajectory-level preference benchmark with 1,171… See the full description on the dataset page: https://huggingface.co/datasets/wyy1112/Plan-RewardBench.reward-bench-2-results
