CoolFace
Modelpublic

FRPO/qwen3-1.7b-a3_onpolicy-k1-cNone-clip0.2-mb1-eta100-bs256x5-n2

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes105downloads
Model Card

qwen3-1.7b-a3_onpolicy-k1-cNone-clip0.2-mb1-eta100-bs256x5-n2

RL fine-tuned checkpoint from the KL-in-LLM-RL / FRPO experiments (trained with verl).

  • —Base model: Qwen/Qwen3-1.7B
  • —Checkpoint(s) in this repo: globalstep200 (repo root)
  • —Weights: fp32 safetensors, exactly as saved by the trainer (no post-processing).
  • —Run configuration is encoded in the repo name.

Auto-uploaded on 2026-08-15.