CoolFace
Modelpublic

OnAnOrange/CUA-Lite-Qwen3.5-4B-ScaleCUA-qwen3_8_27b-epoch2

sourceHugging Faceapache-2.0updated 19d agoView on Hugging Face
0likes13downloads
Model Card

CUA-Lite-Qwen3.5-4B-ScaleCUA-qwen3827b-epoch2

Full-parameter SFT checkpoint of Qwen3.5-4B, trained with CUA-Lite on 5,000 Lite.ScaleCUA desktop trajectories from teacher `qwen3_8_27b`. This repository contains the checkpoint after epoch 2 of a three-epoch training run, with the full BF16 model, tokenizer, and image processor configuration.

Training

SettingValue
Teacher / dataset variantdesktop.use.train.qwen3_8_27b
Training trajectories5,000
Serialized trajectory steps49,935
Completed epochs2
Completed optimizer updates312
Source checkpointiter_311
Last-step training loss0.11305571
Batch size32 trajectories; microbatch 1
Hardware and parallelism8 H200 GPUs; TP=2, DP=4
Learning rate5e-6; cosine decay to 1e-6; 10% warmup
OptimizerAdam, betas=(0.9, 0.95), weight decay=0.1
PrecisionBF16
Training date2026-09-09

The action-only desktop recipe uses a 1280x720 screenshot target (rounded to 1280x704 by the processor), history length 50, and up to four prompt images per step. Reasoning is disabled. Data is filtered for non-excluded trajectories with episode return greater than 0.5. Conversion failures are skipped and the seeded candidate pool is replenished until 5,000 usable trajectories are obtained.

The trainer defines an epoch as floor(5000 / 32) = 156 optimizer updates. One checkpoint is saved per epoch, at zero-based rollout IDs 155, 311, and 467. The last-step training loss above is a single training-batch value, not a benchmark score. No downstream benchmark evaluation is reported for this checkpoint.

Source and usage

Use a Transformers version with Qwen3.5 support; the training export used Transformers 5.6.0. Load the model and processor from this repository. For desktop-agent inference, use CUA-Lite with the included desktop.use.compact.yaml so the prompt, action format, and image history match training. Full training details are in training_config.json.

License

The original Qwen3.5-4B Apache 2.0 license is included as LICENSE. These weights are modified by supervised fine-tuning as described above.