CoolFace
Modelpublic

zheng-cao04/qwen3_8b_dpo_8env2_allstep4x_forcediff_r16a16

sourceHugging Faceupdated 4mo agoView on Hugging Face
0likes7downloads
Model Card

qwen38bdpo8env2allstep4xforcediffr16a16

This repository contains the merged Hugging Face checkpoint for the Qwen3-8B DPO model used by Agent-R.

The uploaded checkpoint is from iter_800.pth.merged. Results may differ from checkpoints from other training iterations.

Base Model

  • —Base model: Qwen/Qwen3-8B
  • —Fine-tuning method: LoRA DPO, merged into the base model for inference
  • —LoRA rank/alpha: r=16, alpha=16

Inference

Example sglang launch command:

bash
python -m sglang_router.launch_server   --router-worker-startup-timeout-secs 600   --router-policy cache_aware   --router-balance-abs-threshold 0   --router-balance-rel-threshold 3   --model-path icebell/qwen3_8b_dpo_8env2_allstep4x_forcediff_r16a16   --tp-size 1   --dp-size 8   --trust-remote-code   --mem-fraction-static 0.83   --kv-cache-dtype auto   --max-total-tokens 48000   --context-length 24000

Notes

Training was performed with XTuner/DeepSpeed Zero-2. The released model is the merged iter-800 checkpoint, not the raw DeepSpeed training checkpoint.