autotrainer
autotrainer-v0
autotrainer-v0
AutoTrainer-v0: LLM-agent-controlled GRPO training on Countdown
Dataset Info
Rows: 1
Columns: 1
Columns
Column
Type
Description
state_json
Value('string')
Full autotrainer state as JSON string
Generation Parameters
{
"script_name": "run_round.py",
"model": "Qwen/Qwen2.5-1.5B-Instruct",
"description": "AutoTrainer-v0: LLM-agent-controlled GRPO training on Countdown",
"experiment_id": "autotrainer-v0"… See the full description on the dataset page: https://huggingface.co/datasets/raca-workspace-v1/autotrainer-v0.autotrainer-v1
AutoTrainer v1
Agentic training harness where Claude decides every training step's method, data, and hyperparameters.
Configs
steps: Per-step decisions, metrics, agent traces, and costs
eval_traces: Per-question evaluation results with difficulty breakdown (n_args=2-10)
autotrainer-v1-run-v2-11stepsautotrainer-v0-eval-tracesautotrainer-v0-roundsautotrainer-v0-state
