OnAnOrange/qwen3.5-9b-lite-osworld-cuagym-sft
09
Qwen3.5-9B Lite.OSWorld + Lite.CUAGym SFT
This model is a supervised fine-tuning checkpoint of `Qwen/Qwen3.5-9B` for computer-use agent trajectories.
Training data
cua-lite/Lite.OSWorld-v1at revision0c5d2495f35a27eae7baf8aa50b2e677b209de39: 1,937 examplescua-lite/Lite.CUAGym-v1at revision92fff449649d210bdfca09b71d4094900f8d5147: 1,788 examples- Combined total: 3,725 examples
The data was exported with examples/lite/v1/configs/qwen3_5/reasoning/lite.osworld.yaml.
Training configuration
- Epochs: 3
- Global batch size: 32
- Micro batch size: 1
- Tensor parallel size: 2
- Data parallel size: 4
- Learning rate:
5e-6 - Minimum learning rate:
1e-6 - Final training step: 347
- Final training loss:
0.235178
Only the final checkpoint was retained. The weights in this repository are the final Hugging Face export from Slurm job 1875082 (eval-bbq-1).
