CoolFace
Datasetpublic

Zhongzhu/tunekv-taubench-data

TuneKV τ-bench GRPO data (§10) Training and eval artifacts for the three τ-bench outcome-RL (GRPO clip 0.2 + 0.1 reverse KL) arms. CE arms live in sibling subdirectories; grpo_* subdirectories are the GRPO siblings (single variable vs CE = loss + all-K-outcome data with per-task z-scored advantages). Subdirectory Arm Rows Key result (same stack) grpo_airline30/ airline · Qwen3-Coder-30B 120 (K=4+K=8, 5/10 mixed) 0.46 → 0.42 full-50 grpo_retail30/ retail ·… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhu/tunekv-taubench-data.

sourceHugging Faceupdated 4d agoView on Hugging Face
0likes625downloads
7 commits on main
601ff884d ago

Add dataset README

Zhongzhu
9e9fded4d ago

GRPO retail-27B arm: rows, rollouts, results, configs, logs, export audit

Zhongzhu
807c9304d ago

GRPO retail-30B arm: rows, rollouts, results, config, logs

Zhongzhu
514598e4d ago

GRPO airline-30B arm: rows, rollouts, results, config, logs

Zhongzhu
6f77ef96d ago

add airline pilot README

Zhongzhu
79ccade6d ago

airline 30B pilot: base 20/50 + tuned-CE 11/25 (odd split)

Zhongzhu
5f2249c6d ago

initial commit

Zhongzhu