laion/a3-rl-DCAgent_exp_rpt_e2egit-large_global_step_15
a3-rl-DCAgentexprpte2egit-large — globalstep_15
SkyRL terminal-bench RL run. Checkpoint global_step_15 was selected by EMA(reward/avgrawreward, 5-step) across all 80 trained steps.
- Base model:
laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink(Qwen3-8B SFT) - Task dataset:
/e/scratch/jureap59/feuer1/tasks/exp_rpt_e2egit-large(DCAgent/exprpte2egit-large) - Algorithm: RLOO-n, epscliplow=0.2, epscliphigh=0.05, kl=0, lr=8e-6
- Trainer: SkyRL FSDP2, 14 nodes \u00d7 4 GPU/node, train_batch=64, 48 vLLM engines
- Total steps trained: 80 (job CANCELLED 2026-05-28 03:37:31 UTC after collapse beyond step 65)
- EMA selection: step=15 ema=0.8639 reward=0.8418
Training Traces
Training-time Daytona/Harbor rollouts for this run are uploaded as a companion dataset: [open-athena/a3-rl-DCAgent_exp_rpt_e2egit-large](https://huggingface.co/datasets/open-athena/a3-rl-DCAgent_exp_rpt_e2egit-large)
The dataset contains the last episode of each trial (per make_and_upload_trace_dataset --episodes last) — the same rollouts the policy was trained on after rollback / truncation.
Training Logs
training_logs/ contains metrics.csv, vllm_metrics.csv, trial_stats.csv, report.md, and reward_plot.png from parse_skyrl_metrics.py, plus the raw trainer_log.jsonl and *.out files for archival (Jupiter has no W&B network access).
RL Config
See rl_config.json for the full Hydra overrides used to launch the run.
