CoolFace
Datasetpublic

atlas-institute/code-trainer-v10-grpo-prompts

code-trainer-v10-grpo-prompts 500 curated prompts for GRPO (Group Relative Policy Optimization) training in the Code-Trainer / RTPI pipeline. Used by both Qwen and Gemma RL stages (Phase 4c) to train tool-call formatting via a rule-based reward function. Schema Column Type Description prompt string The user instruction/question source string Origin: v10_eval, v9_training, or synthetic expected_tool string Primary tool the prompt should invoke… See the full description on the dataset page: https://huggingface.co/datasets/atlas-institute/code-trainer-v10-grpo-prompts.

sourceHugging Faceapache-2.0updated 10d agoView on Hugging Face
0likes42downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
atlas-institute/code-trainer-v10-grpo-prompts · CoolFace