CoolFace
Datasetpublic

zbeeb/Staleness-GRPO-DAPO-Math-17k

Staleness GRPO DAPO Math 17k The exact 17,005-row training dataset shared by the staleness-cap-2 Qwen2.5-Math-1.5B, Qwen2.5-3B, and Qwen2.5-Math-7B checkpoints, and the staleness-cap-4 Qwen2.5-Math-1.5B checkpoint. All four training manifests record the same SHA-256 for the training file. Source and processing Derived from the all configuration of open-r1/DAPO-Math-17k-Processed, itself processed from BytedTsinghua-SIA/DAPO-Math-17k. Source revision:… See the full description on the dataset page: https://huggingface.co/datasets/zbeeb/Staleness-GRPO-DAPO-Math-17k.

sourceHugging Faceapache-2.0updated 8d agoView on Hugging Face
0likes87downloads
4 commits on main
53064568d ago

Record completed staleness-4 checkpoint sharing the training dataset

zbeeb
a39b8318d ago

Name GRPO release by staleness cap

zbeeb
453cb2b9d ago

Publish verified Cedar GRPO final release

zbeeb
eff29cc9d ago

initial commit

zbeeb