CoolFace
Modelpublic

wenlongzhao/ARPO-reproduce_Qwen2.5-3B-Instruct_SFT_paper-config

sourceHugging Faceapache-2.0updated 23d agoView on Hugging Face
0likes835downloads
Model Card

Qwen2.5-3B-Instruct-ARPO-SFT-paper

Full-parameter SFT of Qwen/Qwen2.5-3B-Instruct, following the cold-start recipe in Appendix E.2 of Agentic Reinforced Policy Optimization (ARPO).

This is the paper variant (max length 4096, global batch 128, weight decay 0.1). A companion code variant matching the released ARPO LLaMA-Factory yaml is at `wenlongzhao/ARPO-reproduce_Qwen2.5-3B-Instruct_SFT_code-config`.

Training data

Trained on dongguanting/ARPO-SFT-54K (Tool-Star 54K + STILL), the official ARPO SFT mix.

Training procedure

Trained with LLaMA-Factory full SFT, DeepSpeed ZeRO-3, BF16, and FlashAttention-2.

HyperparameterValue
Learning rate7e-6
LR schedulecosine, warmup ratio 0.1
Epochs3
Global batch size128
Max sequence length4096
Weight decay0.1
PrecisionBF16
Seed42

Final train loss: 0.6852.

Intended use

Research checkpoint for ARPO-style agentic SFT. Not evaluated as a standalone chat model. Follow the licenses of the base model and the SFT dataset.

Citation

bibtex
@article{dong2025arpo,
  author = {Guanting Dong and Hangyu Mao and Kai Ma and Licheng Bao and Yifei Chen and Zhongyuan Wang and Zhongxia Chen and Jiazhen Du and Huiyang Wang and Fuzheng Zhang and Guorui Zhou and Yutao Zhu and Ji-Rong Wen and Zhicheng Dou},
  title = {Agentic Reinforced Policy Optimization},
  journal = {CoRR},
  volume = {abs/2507.19849},
  year = {2025},
  url = {https://arxiv.org/abs/2507.19849}
}