MANGSEOK123/Qwen3-4B-tau2-grpo-retail-2ep-lr1e6
021
Qwen3-4B-Instruct-2507 · tau2-bench plain GRPO, retail only (2 epochs)
Trained on the retail domain of tau2-bench alone rather than on airline+retail jointly, completing a four-way sweep (plain/memory x airline/retail).
Results — tau2 retail test split, 3 trials, 60-hop cap
On retail the memory-augmented variant is the stronger of the two at this budget (avg@3 0.558 vs 0.492, pass^3 0.475 vs 0.350), and that ordering also held under joint training. Joint training carried to 5 epochs reached avg@3 0.592 (memory) and 0.558 (plain), so epoch count mattered more than domain separation here.
Numbers use a 60-hop episode cap rather than tau2's 200-hop default, so they are not comparable to the public leaderboard.
