datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Mem-0-m1mix-dataset-RMBench
Mem-0 m1_mix — RMBench / RoboTwin 2.0 (LeRobot dataset)
The m1_mix training dataset for the Mem-0 execution module: the five RMBench
M1 tasks merged into a single LeRobot
v2.1 dataset with globally unique episode indices. This is the exact data used to
train the checkpoint released at
qiuly/Mem-0-m1mix-RMBench.
Summary
Format
LeRobot v2.1
Episodes
250 (50 per task × 5 tasks)
Frames
92,520
FPS
30
Tasks
5 (see below)
Robot
dual-arm (2× 7-DoF +… See the full description on the dataset page: https://huggingface.co/datasets/qiuly/Mem-0-m1mix-dataset-RMBench.rm-bench-lightevalRMB-Pairwise
RMB-Pairwise
Flattened pairwise split of the RMB (Reward Model Benchmark) dataset from Zhou-Zoey/RMB-Reward-Model-Benchmark.
RMB is a comprehensive reward model benchmark covering 49 real-world scenarios across two alignment goals (Helpfulness and Harmlessness), introduced in the ICLR 2025 paper.
Schema
Column
Type
Description
pair_uid
str
Unique pair identifier
conversation
list[dict]
Multi-turn conversation context (role, content, language)
chosen
str… See the full description on the dataset page: https://huggingface.co/datasets/ilgee/RMB-Pairwise.RMB_dataset
RMB Dataset
A comprehensive benchmark dataset for evaluating reward models in large language model alignment, containing pairwise preference comparisons and best-of-N evaluations.
Overview
The RMB (Reward Model Benchmark) dataset is designed to comprehensively benchmark reward models in LLM alignment. It contains preference data across multiple dimensions including harmlessness and helpfulness, with various task types and response styles.
Citation
If you use… See the full description on the dataset page: https://huggingface.co/datasets/dd101bb/RMB_dataset.RM-Bench-math-Llama-3.2-3B-yesnoRM-Bench-math-Qwen2.5-7B-Instruct-yesnoRM-Bench-math-Llama-3.2-1B-yesnoRMB-BoN
RMB-BoN
Flattened Best-of-N split of the RMB (Reward Model Benchmark) dataset from Zhou-Zoey/RMB-Reward-Model-Benchmark.
RMB is a comprehensive reward model benchmark covering 49 real-world scenarios across two alignment goals (Helpfulness and Harmlessness), introduced in the ICLR 2025 paper.
Schema
Column
Type
Description
bon_uid
str
Unique identifier
conversation
list[dict]
Multi-turn conversation context (role, content, language)
chosen
str
Best response… See the full description on the dataset page: https://huggingface.co/datasets/ilgee/RMB-BoN.rm_bo8_gem2b_gem2b_alpacaeval2RM-Bench-math-Mistral-7B-Instruct-v0.1-yesnoRM-Bench-safety-refuse-Qwen2.5-3B-Instruct-scoresRM-Bench-math-Mistral-7B-Instruct-v0.1-scoresRM-Bench-math-Llama-3.2-1B-Instruct-yesnoRM-Bench-code-gpt-4o-mini-scores-set1RM-Bench-mathRM-Bench-safety-response-Mistral-7B-Instruct-v0.1-yesnoRM-Bench-code-gpt2-normalRM-Bench-safety-refuse-Skywork-Reward-Llama-3.1-8B-v0.2-normalRM-Bench-code-Mistral-7B-Instruct-v0.1-scoresRM-Bench-math-gpt-4o-mini-scores-set1RM-Bench-math-mistral-7b-sft-beta-entropygsm8k_rmbench_style
GSM8K-RMBench-Style — Style-controlled (correct, incorrect) variants on GSM8K
Per GSM8K problem this dataset provides 6 response surfaces — correct and
incorrect each rendered in markdown / normal / concise styles — to support
RM-Bench-style 3 × 3 (chosen × rejected) pair-grid evaluation and forget-LoRA
training for Reward Model debiasing.
{ question, gold }
├── correct : { markdown, normal, concise }
└── incorrect : { markdown, normal, concise }
These 6 surfaces yield 9… See the full description on the dataset page: https://huggingface.co/datasets/xxccho/gsm8k_rmbench_style.RM-Bench-chatRM-Bench-codeRM-Bench-code-Llama-3.2-1B-yesnoRM-Bench-code-Llama-3.2-3B-Instruct-yesnoRM-Bench-math-Meta-Llama-3-8B-Instruct-yesnoRM-Bench-math-Meta-Llama-3-8B-yesnoRM-Bench-code-Mistral-7B-Instruct-v0.1-yesnoRM-Bench-math-Qwen2.5-3B-yesno
