khoilamalphaai/chess-coach-grand-eval
Chess Coach — Grand Eval (comprehensive leaderboard) One fresh, apples-to-apples comparison of every model in the chess move-review coaching project — our tuned specialists, the untuned baselines, and the full frontier lineup — on the same held-out validation slice (120 positions × 3 tiers = 360 scenarios), scored with two independent layers: Deterministic moat metrics (free, python-chess over pre-computed Stockfish/Maia facts): tier-fit, distinct-moves-per-level… See the full description on the dataset page: https://huggingface.co/datasets/khoilamalphaai/chess-coach-grand-eval.
Remove eval cost/spend dollar figures from card
docs(card): align roles/framing with audit (live-served v6-dpo2; v4 base/eval/fallback; verifier-detectable phrasing)
docs: standardize unbiased head-to-head on the reproducible 56-24-12 over 92 diverging (was the non-reproducing 56-28-12/96); asserted by reproduce_v4
docs: add 2026-07-09 metric-framing note (tier-policy match; moat caveat; distinct 73/100=0.730 vs frozen 73/93=0.785)
Upload folder using huggingface_hub
Upload report.json with huggingface_hub
Upload GRAND_EVAL_LEADERBOARD.md with huggingface_hub
Upload README.md with huggingface_hub
Rank leaderboard by tier-appropriate move selection (OURS-v4 shipped, #1)
fix(grand): reconcile per_model.ours_v4 to post-extraction-fix numbers (tier_fit 0.7917, distinct 0.75, move_sound 0.9861, coherence 0.142, flat_rate 0.292, gate well_formed 0.956 / move_sound 0.942)
Fix OURS-v4 post-extraction-fix numbers (moat 51-5-6 over 62; tier-fit 0.79 / distinct 0.75 / move-sound 0.986)
Upload folder using huggingface_hub
Upload val_scenarios.jsonl with huggingface_hub
Upload council.jsonl with huggingface_hub
Upload report.json with huggingface_hub
Upload GRAND_EVAL_LEADERBOARD.md with huggingface_hub
Upload README.md with huggingface_hub
initial commit
