CoolFace
Datasetpublic

Zhongzhu/tunekv-taubench-data

TuneKV τ-bench GRPO data (§10) Training and eval artifacts for the three τ-bench outcome-RL (GRPO clip 0.2 + 0.1 reverse KL) arms. CE arms live in sibling subdirectories; grpo_* subdirectories are the GRPO siblings (single variable vs CE = loss + all-K-outcome data with per-task z-scored advantages). Subdirectory Arm Rows Key result (same stack) grpo_airline30/ airline · Qwen3-Coder-30B 120 (K=4+K=8, 5/10 mixed) 0.46 → 0.42 full-50 grpo_retail30/ retail ·… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhu/tunekv-taubench-data.

sourceHugging Faceupdated 2d agoView on Hugging Face
0likes616downloads
Dataset Card

TuneKV τ-bench GRPO data (§10)

Training and eval artifacts for the three τ-bench outcome-RL (GRPO clip 0.2 + 0.1 reverse KL) arms. CE arms live in sibling subdirectories; grpo_* subdirectories are the GRPO siblings (single variable vs CE = loss + all-K-outcome data with per-task z-scored advantages).

SubdirectoryArmRowsKey result (same stack)
grpo_airline30/airline · Qwen3-Coder-30B120 (K=4+K=8, 5/10 mixed)0.46 → 0.42 full-50
grpo_retail30/retail · Qwen3-Coder-30B2000 (K=4, 221/500 mixed)0.4087 → 0.4174 test-115
grpo_retail27/retail · Qwen3.8-27B1534 (K=4 ≤10k, 106 mixed)0.5565 → 0.5913 test-115

Each subdirectory holds META.json (recipe + verdict), training rows + manifest, rollout sessions, base/tuned eval results, training config(s), training metrics and logs. Merged serving models: `Zhongzhu/tunekv-taubench-grpo`.