datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RM-Bench
RM-Bench
This repository contains the data of the paper "RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style"
News
[2025/07/12] 🎯 The RM-Bench Leaderboard is now publicly available! Check it out and submit your result at RM-Bench Leaderboard!
Dataset Details
the samples are formatted as follows:
{
"id": // unique identifier of the sample,
"prompt": // the prompt given to the model,
"chosen": [
"resp_1", // the… See the full description on the dataset page: https://huggingface.co/datasets/THU-KEG/RM-Bench.factuality-rmbench-style
Factuality RM-Bench Style
Factuality RM-Bench Style is a controlled English dataset for studying whether
reward models and representation probes prefer stylistic presentation over
factual correctness. Each row contains one question, a localized correct and
incorrect proposition, and six responses formed by crossing correctness with
three presentation styles: concise, normal, and Markdown.
This repository is an export package for
factuality_rmbench_style_v6. The published data… See the full description on the dataset page: https://huggingface.co/datasets/Yunnnuy/factuality-rmbench-style.rmb2genRMBG-2.0qa2
