CoolFace
Datasetpublic

Zhongzhu/tunekv-aider-27b

TuneKV Aider-polyglot 27B arms (Qwen3.8-27B hybrid) Three arms on aider-polyglot with Qwen/Qwen3.8-27B (enable_thinking=false, 16 FA + 48 GDN layers), served by vLLM 0.28 + tunekv_vllm.hybrid connector: arm train n_tune python-34 olang 5x5 (pinned) base — — p@1 5/34 (14.7%), p@2 18/34 (52.9%) p@1 5/25, p@2 15/25 CE+0.1KL CE (nll 1.0) + 0.1 reverse-KL, lr 1.5e-4, 3ep/195 steps 384 p@1 6/34 (+1), p@2 15/34 (−3) p@1 5/25 (±0), p@2 16/25 (+1) GRPO per-(ex,try)… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhu/tunekv-aider-27b.

sourceHugging Faceupdated 5d agoView on Hugging Face
0likes403downloads
3 commits on main
5bc6fe95d ago

27B three arms: CE+GRPO artifacts, rows, evals, rollout logs (part 2)

Zhongzhu
88be5215d ago

27B three arms: CE+GRPO artifacts, rows, evals, rollout logs

Zhongzhu
fd7bd195d ago

initial commit

Zhongzhu