datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RMB-Pairwise
RMB-Pairwise
Flattened pairwise split of the RMB (Reward Model Benchmark) dataset from Zhou-Zoey/RMB-Reward-Model-Benchmark.
RMB is a comprehensive reward model benchmark covering 49 real-world scenarios across two alignment goals (Helpfulness and Harmlessness), introduced in the ICLR 2025 paper.
Schema
Column
Type
Description
pair_uid
str
Unique pair identifier
conversation
list[dict]
Multi-turn conversation context (role, content, language)
chosen
str… See the full description on the dataset page: https://huggingface.co/datasets/ilgee/RMB-Pairwise.Qwen2.5-3B-SFT-pairwise-L_RM
