zheng-cao04/qwen3_8b_dpo_8env2_allstep4x_forcediff_r16a16
07
qwen38bdpo8env2allstep4xforcediffr16a16
This repository contains the merged Hugging Face checkpoint for the Qwen3-8B DPO model used by Agent-R.
The uploaded checkpoint is from iter_800.pth.merged. Results may differ from checkpoints from other training iterations.
Base Model
- Base model:
Qwen/Qwen3-8B - Fine-tuning method: LoRA DPO, merged into the base model for inference
- LoRA rank/alpha:
r=16,alpha=16
Inference
Example sglang launch command:
python -m sglang_router.launch_server --router-worker-startup-timeout-secs 600 --router-policy cache_aware --router-balance-abs-threshold 0 --router-balance-rel-threshold 3 --model-path icebell/qwen3_8b_dpo_8env2_allstep4x_forcediff_r16a16 --tp-size 1 --dp-size 8 --trust-remote-code --mem-fraction-static 0.83 --kv-cache-dtype auto --max-total-tokens 48000 --context-length 24000Notes
Training was performed with XTuner/DeepSpeed Zero-2. The released model is the merged iter-800 checkpoint, not the raw DeepSpeed training checkpoint.
