CoolFace
Modelpublic

laion/a3-rl-DCAgent_exp_rpt_e2egit-large_global_step_15

sourceHugging Faceapache-2.0updated 23d agoView on Hugging Face
0likes140downloads
Model Card

a3-rl-DCAgentexprpte2egit-large — globalstep_15

SkyRL terminal-bench RL run. Checkpoint global_step_15 was selected by EMA(reward/avgrawreward, 5-step) across all 80 trained steps.

  • —Base model: laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink (Qwen3-8B SFT)
  • —Task dataset: /e/scratch/jureap59/feuer1/tasks/exp_rpt_e2egit-large (DCAgent/exprpte2egit-large)
  • —Algorithm: RLOO-n, epscliplow=0.2, epscliphigh=0.05, kl=0, lr=8e-6
  • —Trainer: SkyRL FSDP2, 14 nodes \u00d7 4 GPU/node, train_batch=64, 48 vLLM engines
  • —Total steps trained: 80 (job CANCELLED 2026-05-28 03:37:31 UTC after collapse beyond step 65)
  • —EMA selection: step=15 ema=0.8639 reward=0.8418

Training Traces

Training-time Daytona/Harbor rollouts for this run are uploaded as a companion dataset: [open-athena/a3-rl-DCAgent_exp_rpt_e2egit-large](https://huggingface.co/datasets/open-athena/a3-rl-DCAgent_exp_rpt_e2egit-large)

The dataset contains the last episode of each trial (per make_and_upload_trace_dataset --episodes last) — the same rollouts the policy was trained on after rollback / truncation.

Training Logs

training_logs/ contains metrics.csv, vllm_metrics.csv, trial_stats.csv, report.md, and reward_plot.png from parse_skyrl_metrics.py, plus the raw trainer_log.jsonl and *.out files for archival (Jupiter has no W&B network access).

RL Config

See rl_config.json for the full Hydra overrides used to launch the run.