Kablue/football-culture-reasoning
Football Culture Reasoning Bench v1 Expert-graded evaluation of LLM reasoning about football (soccer) fandom culture: subcultural concepts, non-Western specificity (Japan / J.League, South America, Asia), macro-sociological context, and stereotype avoidance. This is a small, deliberately hard proof set (20 items). The failures it probes are cultural, not linguistic — candidate answers are fluent and confident, but stale, West-centric, or normatively preachy. Task… See the full description on the dataset page: https://huggingface.co/datasets/Kablue/football-culture-reasoning.
Football Culture Reasoning Bench v1
Expert-graded evaluation of LLM reasoning about football (soccer) fandom culture: subcultural concepts, non-Western specificity (Japan / J.League, South America, Asia), macro-sociological context, and stereotype avoidance.
This is a small, deliberately hard proof set (20 items). The failures it probes are cultural, not linguistic — candidate answers are fluent and confident, but stale, West-centric, or normatively preachy.
Task
Single-turn. Each row shows a topic and a realistic, fluent, model-generated answer (model_answer, in Japanese). The evaluated model must reproduce a domain expert's judgment:
- `verdict`:
good/borderline/bad - `tags`: an 11-tag error taxonomy (see below)
Each row also carries the expert's written correction (expert_fix), usable as reference text for judge-based / generative evaluation.
Error taxonomy (tags)
Fields
Label distribution: good ×3, borderline ×1, bad ×16.
Why this is hard
A blind frontier model reproduces the expert's coarse verdicts on 18/20 items — but systematically fails the calibration cases (expert-borderline items it flattens to bad) and cannot reproduce the expert rationales:
- Counterexamples — Brighton–Crystal Palace as a non-geographic derby.
- Currency corrections — "corporate J.League" is a ~2000s framing; contemporary support culture has fused with Japanese baseball cheering traditions.
- Causal fixes — Hillsborough: hooligan decline as a byproduct of ticket-price / stadium changes, not "strengthened countermeasures."
Failure modes this eval punishes: West-centric defaults, decades-stale framings of non-Western fandom, normative moralizing ("politics should stay out of sport") instead of explanation, and dismissal of sociological structure (class, migration, political economy).
Ground truth
Single-expert, rationale-backed, second-pass labels by Kei Akiyoshi — football journalist, commentator (DAZN EFL / U-NEXT FA Cup) and author based in Japan, specializing in fandom sociology.
Runnable environment
A verifiers-compatible RL/eval environment using this data is published on the Prime Intellect Environments Hub: https://app.primeintellect.ai/dashboard/environments/kablue/football-culture-reasoning
prime env install kablue/football-culture-reasoning
uv run vf-eval football-culture-reasoning -m <model>Reward: 0.6 * verdict_exact + 0.35 * tag_jaccard + 0.05 * format.
Extensibility
v1 = 20 items (proof set). The pipeline (AI drafts candidate answers → expert adjudicates with tags + written rationale) is domain-portable: tactics discourse, scouting narratives, other fandom subdomains, and other non-Western cultural domains.
Citation
@misc{akiyoshi2026footballculture,
title = {Football Culture Reasoning Bench v1},
author = {Akiyoshi, Kei},
year = {2026},
note = {Expert-graded evaluation of LLM reasoning about football fandom culture},
howpublished = {\url{https://app.primeintellect.ai/dashboard/environments/kablue/football-culture-reasoning}}
}Contact: kei@japanesethe72.com · X @Japanesethe72 · Substack
