atlas-institute/code-trainer-v10-grpo-prompts
code-trainer-v10-grpo-prompts 500 curated prompts for GRPO (Group Relative Policy Optimization) training in the Code-Trainer / RTPI pipeline. Used by both Qwen and Gemma RL stages (Phase 4c) to train tool-call formatting via a rule-based reward function. Schema Column Type Description prompt string The user instruction/question source string Origin: v10_eval, v9_training, or synthetic expected_tool string Primary tool the prompt should invoke… See the full description on the dataset page: https://huggingface.co/datasets/atlas-institute/code-trainer-v10-grpo-prompts.
042
