CoolFace
Datasetpublic

atlas-institute/code-trainer-v10-grpo-prompts

code-trainer-v10-grpo-prompts 500 curated prompts for GRPO (Group Relative Policy Optimization) training in the Code-Trainer / RTPI pipeline. Used by both Qwen and Gemma RL stages (Phase 4c) to train tool-call formatting via a rule-based reward function. Schema Column Type Description prompt string The user instruction/question source string Origin: v10_eval, v9_training, or synthetic expected_tool string Primary tool the prompt should invoke… See the full description on the dataset page: https://huggingface.co/datasets/atlas-institute/code-trainer-v10-grpo-prompts.

sourceHugging Faceapache-2.0updated 10d agoView on Hugging Face
0likes42downloads
3 commits on main
274c93010d ago

Add dataset card

atlas-institute
f5dd3301mo ago

Upload dataset

Raymond Soreng
7a07fe01mo ago

initial commit

Raymond Soreng