CoolFace
Datasetpublic

xxccho/gsm8k_rmbench_style

GSM8K-RMBench-Style — Style-controlled (correct, incorrect) variants on GSM8K Per GSM8K problem this dataset provides 6 response surfaces — correct and incorrect each rendered in markdown / normal / concise styles — to support RM-Bench-style 3 × 3 (chosen × rejected) pair-grid evaluation and forget-LoRA training for Reward Model debiasing. { question, gold } ├── correct : { markdown, normal, concise } └── incorrect : { markdown, normal, concise } These 6 surfaces yield 9… See the full description on the dataset page: https://huggingface.co/datasets/xxccho/gsm8k_rmbench_style.

sourceHugging Facemitupdated 5mo agoView on Hugging Face
0likes12downloads
settings

This repository belongs to xxccho on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namegsm8k_rmbench_style
visibilitypublic
licencemit
gatedno
ownerxxccho
Account settings
xxccho/gsm8k_rmbench_style · CoolFace