Zhongzhu/tunekv-taubench-data
TuneKV τ-bench GRPO data (§10) Training and eval artifacts for the three τ-bench outcome-RL (GRPO clip 0.2 + 0.1 reverse KL) arms. CE arms live in sibling subdirectories; grpo_* subdirectories are the GRPO siblings (single variable vs CE = loss + all-K-outcome data with per-task z-scored advantages). Subdirectory Arm Rows Key result (same stack) grpo_airline30/ airline · Qwen3-Coder-30B 120 (K=4+K=8, 5/10 mixed) 0.46 → 0.42 full-50 grpo_retail30/ retail ·… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhu/tunekv-taubench-data.
Add dataset README
GRPO retail-27B arm: rows, rollouts, results, configs, logs, export audit
GRPO retail-30B arm: rows, rollouts, results, config, logs
GRPO airline-30B arm: rows, rollouts, results, config, logs
add airline pilot README
airline 30B pilot: base 20/50 + tuned-CE 11/25 (odd split)
initial commit
