laion/Qwen3-32B-NL2Bash-31step
015
Qwen3-32B-NL2Bash-31step
RL-trained Qwen3-32B on NL2Bash terminal tasks.
Training Details
- Base model: Qwen/Qwen3-32B
- Training method: RLOO (async)
- Training data: 1,570 NL2Bash tasks (DCAgent2/nl2bash-tasks-cleaned-oracle)
- Steps: 31 global steps (3 epochs)
- Infrastructure: 17x4 GH200 GPU nodes (JSC), FSDP2 with TP=2 for inference engines (26 inference engines + 4 policy/ref nodes)
- Sandbox environment: Beta9/Beam containers for code execution
- Batch size: 64, 8 samples per prompt
- Learning rate: 1e-5
Training Curve
License
Apache 2.0
