CoolFace
Datasetpublic

khoilamalphaai/chess-coach-grand-eval

Chess Coach — Grand Eval (comprehensive leaderboard) One fresh, apples-to-apples comparison of every model in the chess move-review coaching project — our tuned specialists, the untuned baselines, and the full frontier lineup — on the same held-out validation slice (120 positions × 3 tiers = 360 scenarios), scored with two independent layers: Deterministic moat metrics (free, python-chess over pre-computed Stockfish/Maia facts): tier-fit, distinct-moves-per-level… See the full description on the dataset page: https://huggingface.co/datasets/khoilamalphaai/chess-coach-grand-eval.

sourceHugging Facecc-by-nc-4.0updated 2mo agoView on Hugging Face
0likes48downloads
Dataset Card

Chess Coach — Grand Eval (comprehensive leaderboard)

One fresh, apples-to-apples comparison of every model in the chess move-review coaching project — our tuned specialists, the untuned baselines, and the full frontier lineup — on the same held-out validation slice (120 positions × 3 tiers = 360 scenarios), scored with two independent layers:

  1. 1.Deterministic moat metrics (free, python-chess over pre-computed Stockfish/Maia facts): tier-fit, distinct-moves-per-level, move-soundness, raw faithfulness (verify-pass on draft 1), tier-coherence, and shipped-gate soundness.
  2. 2.Blinded cross-family frontier council (GPT-5.5 + Claude Opus 4.8 + Gemini 3.1 Pro via the TrueFoundry gateway), grading each anonymised response 0–10 on move and instructiveness, with 95 % cluster-bootstrap CIs. Council: 225 items × 3 judges = 675 gradings.

Every gateway (TFY) model was regenerated fresh on these exact positions; our Modal/MLX tuned models are deterministic given their adapter (reused where noted — see the "How each row was generated" table below).

Files

FileWhat
GRAND_EVAL_LEADERBOARD.mdthe human-readable leaderboard (rendered below)
report.jsonevery metric per model + per-tuned-model moat proof
council.jsonlraw blinded council gradings (0–10 move + instr, per judge, with token usage)
gen/<model>.jsonleach model's coaching generations on the val slice
val_scenarios.jsonlthe held-out positions (engine-grounded, with sound pools)

One fresh, apples-to-apples comparison of all 20 models — our tuned specialists, the untuned baselines, and the full frontier lineup — on the SAME held-out VAL slice, scored with BOTH layers:

  • —Deterministic moat metrics (free; python-chess over pre-computed Stockfish/Maia facts) over all 120 positions × 3 tiers = 360 scenarios: tier-fit, distinct-moves-per-level, move-soundness, raw faithfulness (verify-pass on draft 1), tier-coherence, shipped-gate soundness.
  • —Blinded cross-family frontier council (GPT-5.5 + Claude Opus 4.8 + Gemini 3.1 Pro via TrueFoundry), 0-10 move + instructiveness with 95% CIs, over 75 of the 120 positions (675 gradings) — sized to the TFY budget.

Every TFY gateway model was regenerated FRESH on these exact positions (never reusing the old frontier gens); our Modal/MLX tuned models are deterministic given their adapter (reused where noted). ours_v5 is the finish-v5 controller's fresh Modal Volume gen.

Frontier reachability: the 14-model lineup = 3 frontier APIs + 11 open candidates; 12 reachable (dsr1 via bedrock-oss-group/deepseek-r1), 2 blocked: llama4-maverick (400, Meta Llama access denied) and kimi-k2-thinking (403, not authorized).

Metric framing (2026-07-09 honest reframe)

Numbers below are the as-computed grand-eval values (canonical/frozen); the framing terms are aligned to the honest reframe:

  • —"tier-fit" = "tier-policy exact match" — exact agreement with the preregistered select_tier_move rule, a PROJECT RULE, not validated pedagogy. Lead with the all-scenario number (v4 0.767 vs best frontier 0.553), not a head-to-head win rate.
  • —The "moat" head-to-head is a project-rule metric, not a general win rate. The per-tuned W/L/T table below is SELECTION-CONDITIONED (only the positions where OURS already gives a distinct, sound, correctly-graded move AND diverges from the frontier). For v4, the honest unbiased head-to-head over all 92 diverging positions is 56-24-12 (56-24-40 over all 120), recomputed from the committed raw/greedy gens and asserted by scripts/reproduce_v4.py (it supersedes an earlier eval-audit figure that did not reproduce — same 56 wins and 12 ties, but four audit-only losses over four extra diverging positions that did not replay); the 51-5-6 over 62 shown below is the conditioned subset.
  • —distinct-moves denominator: the distinct↑ column is distinct / (positions where the model named both tier moves) as computed at grand-eval time (v4 = 73/93 = 0.785). The honest all-opportunities denominator (every canonical beginner!=advanced position, a no-answer counting as a miss) gives v4 73/100 = 0.730; see data/benchmark_honest/report_v4.json.

Leaderboard — ranked by tier-appropriate move selection (the trained behavior)

Sort key: ranked by tier-appropriate move selection — the deterministic tier-fit↑ metric (the behavior we trained and the graded axis), ties broken by distinct-moves-per-level↑ then move-soundness↑. Every deterministic axis (tier-fit / distinct / move-sound / coherence and the moat) uses the canonical STRICT any-legal move extractor (coach_gate.pick_recommendation, accept = any legal move; no in-pool backfill) — so an output that names no clearly-legal move is a miss everywhere and the leaderboard method matches the moat method exactly. The per-tuned head-to-head W/L/T vs the best frontier is in the moat table below. Instructiveness (the blinded cross-family council) is shown as a secondary axis in the instr 0-10 / move 0-10 / rank↓ columns — OURS-v4 is intentionally weaker on council prose and that is reported here honestly and unchanged. Only the row order reflects tier-fit; every model's measured numbers are the deterministic + council values.

#Modelfamilygengatedtier-fit↑distinct↑move-sound↑raw-faith↑coh-viol↓instr 0-10↑ [95% CI]move 0-10↑rank↓top1%
1OURS-v4 (Qwen3-32B tuned)oursreuseraw0.7670.7850.9420.5890.1404.528 [4.168–4.875]7.66012.669.800
2OURS-v2 (Qwen3-1.7B tuned)oursreuseraw0.5780.3801.0000.6890.1674.323 [3.983–4.662]8.05013.588.000
3OURS-v3 (Qwen3-32B tuned)oursreuseraw0.5580.5850.9500.9420.2296.428 [6.131–6.738]8.5407.76428.90
4Gemini 3.1 ProfrontierFRESHraw0.5530.2101.0000.9580.2926.902 [6.721–7.080]9.1106.8388.400
5OURS-v5 (Qwen3-32B tuned, v5)oursFRESHraw0.5360.7260.8280.5750.2603.863 [3.530–4.188]7.25014.373.100
6Claude Opus 4.8frontierFRESHraw0.5080.2001.0000.9440.3087.062 [6.876–7.249]9.1606.01117.80
7GPT-5.5frontierFRESHraw0.4940.2801.0000.9860.3427.984 [7.869–8.098]9.3803.02440.40
8DeepSeek-R1 (reasoning)openFRESHraw0.4360.3701.0000.9780.3006.117 [5.906–6.319]9.0609.5241.300
9GLM-5openFRESHraw0.4080.2701.0000.9060.3506.875 [6.685–7.068]9.1506.7809.300
10DeepSeek-V3.2openFRESHraw0.4000.2800.9970.9500.3925.859 [5.612–6.107]8.90010.090.900
11OURS-4B (Qwen3-4B tuned)oursreuseyes0.3970.2801.000—0.3255.828 [5.608–6.040]8.93010.373.600
12PROMPT-BASE-4B (Qwen3-4B engineered)basereuseyes0.3780.4601.000—0.3334.799 [4.592–5.011]8.81013.520.400
13BASE (Qwen3-1.7B untuned)basereuseraw0.3580.2800.9920.8580.3331.939 [1.779–2.107]6.95018.710.000
14BASE (Qwen3-32B untuned)baseFRESHraw0.3580.2501.0000.9390.4425.493 [5.286–5.689]8.95011.630.400
15Llama-3.3-70BopenFRESHraw0.3560.1601.0000.9970.4176.262 [6.095–6.422]9.1609.3271.300
16BASE-4B (Qwen3-4B untuned)basereuseyes0.3530.2301.000—0.3754.723 [4.526–4.927]8.75013.940.000
17Mistral-Large-3 (675B)openFRESHraw0.3360.3700.9970.9190.4035.217 [4.982–5.436]8.77012.140.900
18Kimi-K2.5openFRESHraw0.3330.4101.0000.8750.5006.266 [6.058–6.470]9.0708.7586.700
19Gemma-3-27B-itopenFRESHraw0.2830.1901.0000.9690.4175.778 [5.553–5.981]9.01010.620.400
20Qwen3-Next-80B-A3BopenFRESHraw0.2810.2401.0000.9530.3335.875 [5.654–6.092]8.95010.352.700

gen: FRESH = regenerated this run; reuse = deterministic adapter/MLX gen reused. gated: `yes` = full shipped verify-and-regenerate pipeline (4B trio); `raw` = ungated draft (raw-draft gate axes shown). raw-faith = verify-pass on draft 1 (1 − fabrication). tier-fit / distinct / move-sound / raw-faith / coherence are deterministic (free); instr / move 0-10 + rank are the blinded council.

The moat — each tuned model vs the best frontier (tier-fit then soundness)

On positions where OURS gives distinct, sound, correctly-graded per-tier moves AND diverges from the best-frontier move, who wins the platform's move-quality moat (the assemble.derive_wins definition)? Instructiveness (where the frontier leads) is reported separately above.

Tuned modeldistinctdistinct & diverge**W****L****T**
OURS-v2 (Qwen3-1.7B tuned)474321175
OURS-4B (Qwen3-4B tuned)26235144
OURS-v3 (Qwen3-32B tuned)474422139
OURS-v4 (Qwen3-32B tuned)67625156
OURS-v5 (Qwen3-32B tuned, v5)43402668

Shipped-gate soundness (tuned models through the SAME verify+fallback gate)

Tuned modelgated move-sound↑gated well-formed↑gated no-engine-speak↑gate fallback↓
OURS-v2 (Qwen3-1.7B tuned)1.0001.0001.0000.358
OURS-4B (Qwen3-4B tuned)1.0001.0001.0000.000
OURS-v3 (Qwen3-32B tuned)1.0001.0000.9690.181
OURS-v4 (Qwen3-32B tuned)1.0001.0000.9830.444
OURS-v5 (Qwen3-32B tuned, v5)1.0001.0000.9920.444

Once gated, tuned soundness/format hit a shared ~100% floor (zero verifier-detectable mechanical violations by construction) — a fairness floor, not a differentiator; the differentiators are tier-fit / distinct-moves / instructiveness.

Deterministic gate axes (raw draft for ungated rows; telemetry for gated 4B)

Modelgatedno-engine-speak↑well-formed↑move-sound↑verify-pass draft1↑mean attemptsfallback↓
OURS-v4 (Qwen3-32B tuned)raw0.9780.9560.9420.589——
OURS-v2 (Qwen3-1.7B tuned)raw1.0001.0001.0000.689——
OURS-v3 (Qwen3-32B tuned)raw0.9640.9580.9500.942——
Gemini 3.1 Proraw0.9971.0001.0000.958——
OURS-v5 (Qwen3-32B tuned, v5)raw0.9780.8610.8280.575——
Claude Opus 4.8raw1.0001.0001.0000.944——
GPT-5.5raw1.0001.0001.0000.986——
DeepSeek-R1 (reasoning)raw1.0001.0001.0000.978——
GLM-5raw0.9971.0001.0000.906——
DeepSeek-V3.2raw1.0001.0000.9970.950——
OURS-4B (Qwen3-4B tuned)yes1.0001.000——1.1940.008
PROMPT-BASE-4B (Qwen3-4B engineered)yes1.0001.000——1.1670.003
BASE (Qwen3-1.7B untuned)raw0.9641.0000.9920.858——
BASE (Qwen3-32B untuned)raw0.9921.0001.0000.939——
Llama-3.3-70Braw1.0001.0001.0000.997——
BASE-4B (Qwen3-4B untuned)yes1.0001.000——1.1560.000
Mistral-Large-3 (675B)raw1.0000.9970.9970.919——
Kimi-K2.5raw1.0001.0001.0000.875——
Gemma-3-27B-itraw1.0001.0001.0000.969——
Qwen3-Next-80B-A3Braw1.0001.0001.0000.953——

How each row was generated

Modelfresh/reusedmethod
OURS-v5 (Qwen3-32B tuned, v5)FRESHModal-adapter FRESH (finish-v5 controller Volume gen)
OURS-v4 (Qwen3-32B tuned)reusedModal-adapter reuse (honest val, deterministic)
OURS-v3 (Qwen3-32B tuned)reusedModal-adapter reuse (gap803, deterministic)
OURS-v2 (Qwen3-1.7B tuned)reusedMLX-local reuse (gap803, greedy deterministic)
OURS-4B (Qwen3-4B tuned)reusedModal reuse (honest val, gated pipeline)
BASE (Qwen3-32B untuned)FRESHTFY FRESH (aws-bedrock qwen3-32b)
BASE (Qwen3-1.7B untuned)reusedMLX-local reuse (gap803, greedy deterministic)
BASE-4B (Qwen3-4B untuned)reusedModal reuse (honest val, gated pipeline)
PROMPT-BASE-4B (Qwen3-4B engineered)reusedModal reuse (honest val, gated pipeline)
GPT-5.5FRESHTFY FRESH
Claude Opus 4.8FRESHTFY FRESH
Gemini 3.1 ProFRESHTFY FRESH
Qwen3-Next-80B-A3BFRESHTFY FRESH
Gemma-3-27B-itFRESHTFY FRESH
Llama-3.3-70BFRESHTFY FRESH
DeepSeek-V3.2FRESHTFY FRESH
GLM-5FRESHTFY FRESH
Mistral-Large-3 (675B)FRESHTFY FRESH
Kimi-K2.5FRESHTFY FRESH
DeepSeek-R1 (reasoning)FRESHTFY FRESH