RMB
Datasets
All datasets matching “RMB”RMBenchRM-Bench
RM-Bench
This repository contains the data of the paper "RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style"
News
[2025/07/12] 🎯 The RM-Bench Leaderboard is now publicly available! Check it out and submit your result at RM-Bench Leaderboard!
Dataset Details
the samples are formatted as follows:
{
"id": // unique identifier of the sample,
"prompt": // the prompt given to the model,
"chosen": [
"resp_1", // the… See the full description on the dataset page: https://huggingface.co/datasets/THU-KEG/RM-Bench.RMBench-taco-gemini
RMBench-taco-gemini
RMBench training episodes (9 tasks, 449 episodes, 30 fps) with dense high-level labels produced by the TACOR offline annotator:
Gemini 3.7 Flash reads each whole episode as one video clip (one sample every 25 frames) and labels every sampled frame under the
task-specific context (taco) induced for that task. Each tick carries the current subtask, the running textual memory and the
visual-memory operations (keyframe store / retrieval) that the online… See the full description on the dataset page: https://huggingface.co/datasets/Myungkyu/RMBench-taco-gemini.RMBench-preset-gemini
RMBench-preset-gemini
RMBench training episodes (9 tasks, 450 episodes, 30 fps) with dense high-level labels produced by the TACOR offline annotator:
Gemini 3.7 Flash reads each whole episode as one video clip (one sample every 25 frames) and labels every sampled frame given only
the task's subtask preset (the ordered list of subtask labels, no further task-specific guidance). Each tick carries the current
subtask, the running textual memory and the visual-memory operations… See the full description on the dataset page: https://huggingface.co/datasets/Myungkyu/RMBench-preset-gemini.RMBench-preset-luna
RMBench-preset-luna
RMBench training episodes (9 tasks, 450 episodes, 30 fps) with dense high-level labels produced by the TACOR offline annotator:
GPT-5.6 Luna (gpt-5.6-luna) reads the frames of each episode sampled every 25 frames as labelled images and labels every sampled frame given only
the task's subtask preset (the ordered list of subtask labels, no further task-specific guidance). Each tick carries the current
subtask, the running textual memory and the visual-memory… See the full description on the dataset page: https://huggingface.co/datasets/Myungkyu/RMBench-preset-luna.RMBench-taco-wodemo-gemini
RMBench-taco-wodemo-gemini
RMBench training episodes (9 tasks, 450 episodes, 30 fps) with dense high-level labels produced by the TACOR offline annotator:
Gemini 3.7 Flash reads each whole episode as one video clip (one sample every 25 frames) and labels every sampled frame under a
task-specific context (taco) for that task. Each tick carries the current subtask, the running textual memory and the
visual-memory operations (keyframe store / retrieval) that the online high-level… See the full description on the dataset page: https://huggingface.co/datasets/Myungkyu/RMBench-taco-wodemo-gemini.
