CoolFace
Modelpublic

laion/Qwen3-32B-NL2Bash-31step

sourceHugging Faceapache-2.0updated 7mo agoView on Hugging Face
0likes15downloads
Model Card

Qwen3-32B-NL2Bash-31step

RL-trained Qwen3-32B on NL2Bash terminal tasks.

Training Details

  • —Base model: Qwen/Qwen3-32B
  • —Training method: RLOO (async)
  • —Training data: 1,570 NL2Bash tasks (DCAgent2/nl2bash-tasks-cleaned-oracle)
  • —Steps: 31 global steps (3 epochs)
  • —Infrastructure: 17x4 GH200 GPU nodes (JSC), FSDP2 with TP=2 for inference engines (26 inference engines + 4 policy/ref nodes)
  • —Sandbox environment: Beta9/Beam containers for code execution
  • —Batch size: 64, 8 samples per prompt
  • —Learning rate: 1e-5

Training Curve

MetricStep 1Step 10Step 20Step 31
Avg Raw Reward0.2140.3140.4160.264
Pass@80.5630.5630.5940.422

License

Apache 2.0