CoolFace
Datasetpublic

yavuz-ai/self-reward-collapse-terse

self-reward-collapse-terse Per-round trajectory and held-out samples from an iterative self-training loop on GSM8K (Qwen2.5-7B-Instruct, LoRA DPO, 6 rounds). Same loop every arm; only the preference-label source differs. This dataset is the terse arm: pairs labelled by a deliberately gameable brevity reward (the shortest of K samples wins). Study question: does a model training on its own judgment collapse? Short answer on verifiable math: the reward can be hacked, but… See the full description on the dataset page: https://huggingface.co/datasets/yavuz-ai/self-reward-collapse-terse.

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes82downloads
9 commits on main
4554b613mo ago

Upload README.md with huggingface_hub

yavuz-ai
923a4a23mo ago

Upload folder using huggingface_hub

yavuz-ai
b3ea9b63mo ago

Upload folder using huggingface_hub

yavuz-ai
8e1048c3mo ago

Upload folder using huggingface_hub

yavuz-ai
31b5d673mo ago

Upload folder using huggingface_hub

yavuz-ai
db488633mo ago

Upload folder using huggingface_hub

yavuz-ai
079b7483mo ago

Upload dataset

yavuz-ai
fbaf40f3mo ago

Upload dataset

yavuz-ai
3c3b6d33mo ago

initial commit

yavuz-ai