CoolFace
Modelpublic

laion/tt-x5_gradnorm-gn0p9

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes88downloads
Model Card

tt-x5_gradnorm-gn0p9

RL-finetuned from Qwen/Qwen3-Coder-30B-A3B-Instruct using RLOO with FSDP2 expert-parallel training on the exp_rpt_multifile terminal-bench agentic task suite (terminus-2 Harbor harness).

Training configuration

parametervalue
algorithmRLOO (n=8)
strategyFSDP2 + expert-parallel (EP=4)
maxgradnorm0.9
learning_rate8e-6
eps_cliplow=0.2, high=0.05
loss_reductionseqmeantokensumnorm_global
TISenabled (cap=2.0)
KL lossdisabled (coef=0.0)
batch_size64 groups × 8 samples
max_steps80 (reached 66 — wall-time limited)
selected checkpointglobalstep65

Training results

metricvalue (step 65)
reward (mean)0.197
pass@80.406
policy_entropy0.492
ppoclipratio0.011

Training Traces

Companion trace dataset: DCAgent/tt-x5_gradnorm-gn0p9