skebbe/punditbench-2026-27-season-tables
PunditBench 2026-27 pre-registered season tables How did 40–42 language models rank every club in Europe's five largest football leagues before the season began? This static dataset contains 204 complete final-table forecasts from PunditBench's 2026–27 league benchmark. Each row is one model's predicted finishing order for one league. Every included forecast was locked before that league's opening kickoff, committed publicly, hashed, and tagged in Git. La Liga: 40 tables… See the full description on the dataset page: https://huggingface.co/datasets/skebbe/punditbench-2026-27-season-tables.
PunditBench 2026-27 pre-registered season tables
How did 40–42 language models rank every club in Europe's five largest football leagues before the season began?
This static dataset contains 204 complete final-table forecasts from PunditBench's 2026–27 league benchmark. Each row is one model's predicted finishing order for one league. Every included forecast was locked before that league's opening kickoff, committed publicly, hashed, and tagged in Git.
- La Liga: 40 tables
- Premier League: 40 tables
- Serie A: 42 tables
- Ligue 1: 40 tables
- Bundesliga: 42 tables
See the live benchmark, scored results, and the optional €5 opening-round brief at punditbench.com.
What is in each row
The JSONL format preserves the nested predicted_table array. Hugging Face documents JSON Lines as a supported format and recommends it for nested tabular data.
Provenance and integrity
The source of record is the public PunditBench repository. Methodology, prompt design, scoring, caveats, raw API evidence, and league integrity rules are documented in METHODOLOGY.md.
The five source locks are:
source-manifest.json records the exact source tags and commits used for this static snapshot. During preparation, all five source lock digests were recomputed in memory, the resulting export was rebuilt a second time, and the two outputs were byte-identical. The generator itself is intentionally not part of the upload bundle; the public tagged source files remain the audit trail.
The prepared JSONL contains 204 rows and 204 unique `(competition_id, model_slug)` keys. Its bundle-time SHA-256 is cd0485253c1a81da452477c9329e8abec546c8935d02121c26464c076e43c11a.
Intended uses
- Compare model consensus and disagreement about a league table.
- Measure forecast accuracy as the 2026–27 seasons unfold.
- Study family or vendor correlation across models.
- Audit a pre-registered LLM forecasting benchmark.
- Build visualizations of predicted champions, top-four clubs, or relegation candidates.
Limitations
- These are one-shot forecasts at temperature 0, not calibrated probability distributions.
- Models from the same family are correlated; 40–42 rows are not 40–42 independent opinions.
- Training cutoffs differ. The shared pre-season prompt supplies verified transfers, injuries, managerial changes, promoted teams, and the previous season's table, but it cannot erase every knowledge asymmetry.
- Football is high variance. A correct table or champion can still be lucky.
- Model names and club names are used editorially and do not imply endorsement or affiliation.
- Every forecast is AI-generated and may be wrong. This dataset is research evidence, not betting advice.
License
No reuse license has been granted for this snapshot. Public availability on GitHub or Hugging Face does not itself grant permission to reproduce, modify, or redistribute the data. The source repository also has no project-wide license as of this snapshot. Contact the repository owner before reuse beyond inspection and citation.
