CoolFace
Datasetpublic

Goekdeniz-Guelmez/JOSIE-Zero-4-GRPO-Outputs-105

JOSIE-Zero-4: Out-of-Training GRPO Evaluation Generations This dataset contains 105 model-generated reasoning traces from JOSIE-Zero-4. The samples were selected from a 500-example evaluation on the Math branch of openbmb/UltraData-RL-2609. These problems were not used to train JOSIE-Zero-4 with GRPO. The evaluation was designed to examine whether reasoning behavior learned through GRPO transfers to data outside the model's training distribution. Of the 500 evaluated generations… See the full description on the dataset page: https://huggingface.co/datasets/Goekdeniz-Guelmez/JOSIE-Zero-4-GRPO-Outputs-105.

sourceHugging Faceapache-2.0updated 13d agoView on Hugging Face
0likes63downloads
4 commits on main
2d3044913d ago

Upload train.jsonl

Goekdeniz-Guelmez
3295e5213d ago

Update README.md

Goekdeniz-Guelmez
33f95c013d ago

Update README.md

Goekdeniz-Guelmez
10c96f813d ago

initial commit

Goekdeniz-Guelmez