CoolFace
Modelpublic

MANGSEOK123/Qwen3-4B-tau2-grpo-retail-2ep-lr1e6

sourceHugging Faceapache-2.0updated 8d agoView on Hugging Face
0likes21downloads
Model Card

Qwen3-4B-Instruct-2507 · tau2-bench plain GRPO, retail only (2 epochs)

Trained on the retail domain of tau2-bench alone rather than on airline+retail jointly, completing a four-way sweep (plain/memory x airline/retail).

Base modelQwen/Qwen3-4B-Instruct-2507
AlgorithmGRPO, KL loss 0.01, lr 1e-6 constant
Steps8 (= 2 epochs)
Batch8 tasks x 8 rollouts = 64 episodes/step
Episode budget60 tau2 hops (~30 agent turns)
User simulatorgpt-4.1-2025-04-14
Hardware8x RTX A5000 24GB, FSDP + CPU offload, Ulysses SP=2

Results — tau2 retail test split, 3 trials, 60-hop cap

pass@3avg@3pass^3
base0.7750.5750.400
this model0.6250.4920.350
memory GRPO, retail only0.6500.5580.475

On retail the memory-augmented variant is the stronger of the two at this budget (avg@3 0.558 vs 0.492, pass^3 0.475 vs 0.350), and that ordering also held under joint training. Joint training carried to 5 epochs reached avg@3 0.592 (memory) and 0.558 (plain), so epoch count mattered more than domain separation here.

Numbers use a 60-hop episode cap rather than tau2's 200-hop default, so they are not comparable to the public leaderboard.