CoolFace
Datasetpublic

Kablue/football-culture-reasoning

Football Culture Reasoning Bench v1 Expert-graded evaluation of LLM reasoning about football (soccer) fandom culture: subcultural concepts, non-Western specificity (Japan / J.League, South America, Asia), macro-sociological context, and stereotype avoidance. This is a small, deliberately hard proof set (20 items). The failures it probes are cultural, not linguistic — candidate answers are fluent and confident, but stale, West-centric, or normatively preachy. Task… See the full description on the dataset page: https://huggingface.co/datasets/Kablue/football-culture-reasoning.

sourceHugging Facecc-by-nc-4.0updated 3mo agoView on Hugging Face
0likes27downloads
Dataset Card

Football Culture Reasoning Bench v1

Expert-graded evaluation of LLM reasoning about football (soccer) fandom culture: subcultural concepts, non-Western specificity (Japan / J.League, South America, Asia), macro-sociological context, and stereotype avoidance.

This is a small, deliberately hard proof set (20 items). The failures it probes are cultural, not linguistic — candidate answers are fluent and confident, but stale, West-centric, or normatively preachy.

Task

Single-turn. Each row shows a topic and a realistic, fluent, model-generated answer (model_answer, in Japanese). The evaluated model must reproduce a domain expert's judgment:

  • —`verdict`: good / borderline / bad
  • —`tags`: an 11-tag error taxonomy (see below)

Each row also carries the expert's written correction (expert_fix), usable as reference text for judge-based / generative evaluation.

Error taxonomy (tags)

TagMeaning
WCWest-centrism
OLDoutdated framing
CONFconcept conflation
SIMPoversimplification
FACTfactual error
STERstereotype
MACROmissing macro-social structure
DISMdismissiveness
NORMnormative assertion ("should"-ism)
DEBUNKrepeating debunked narratives
GOODsound answer, no correction needed

Fields

FieldTypeDescription
idintItem id (1–20)
topicstringThe fandom-culture topic being evaluated
model_answerstringA fluent AI-generated answer in Japanese (the thing being judged)
verdictstringExpert verdict: good / borderline / bad
tagslist[string]Expert-assigned error tags from the taxonomy
expert_fixstringThe expert's written correction / rationale (Japanese)

Label distribution: good ×3, borderline ×1, bad ×16.

Why this is hard

A blind frontier model reproduces the expert's coarse verdicts on 18/20 items — but systematically fails the calibration cases (expert-borderline items it flattens to bad) and cannot reproduce the expert rationales:

  • —Counterexamples — Brighton–Crystal Palace as a non-geographic derby.
  • —Currency corrections — "corporate J.League" is a ~2000s framing; contemporary support culture has fused with Japanese baseball cheering traditions.
  • —Causal fixes — Hillsborough: hooligan decline as a byproduct of ticket-price / stadium changes, not "strengthened countermeasures."

Failure modes this eval punishes: West-centric defaults, decades-stale framings of non-Western fandom, normative moralizing ("politics should stay out of sport") instead of explanation, and dismissal of sociological structure (class, migration, political economy).

Ground truth

Single-expert, rationale-backed, second-pass labels by Kei Akiyoshi — football journalist, commentator (DAZN EFL / U-NEXT FA Cup) and author based in Japan, specializing in fandom sociology.

Runnable environment

A verifiers-compatible RL/eval environment using this data is published on the Prime Intellect Environments Hub: https://app.primeintellect.ai/dashboard/environments/kablue/football-culture-reasoning

bash
prime env install kablue/football-culture-reasoning
uv run vf-eval football-culture-reasoning -m <model>

Reward: 0.6 * verdict_exact + 0.35 * tag_jaccard + 0.05 * format.

Extensibility

v1 = 20 items (proof set). The pipeline (AI drafts candidate answers → expert adjudicates with tags + written rationale) is domain-portable: tactics discourse, scouting narratives, other fandom subdomains, and other non-Western cultural domains.

Citation

bibtex
@misc{akiyoshi2026footballculture,
  title  = {Football Culture Reasoning Bench v1},
  author = {Akiyoshi, Kei},
  year   = {2026},
  note   = {Expert-graded evaluation of LLM reasoning about football fandom culture},
  howpublished = {\url{https://app.primeintellect.ai/dashboard/environments/kablue/football-culture-reasoning}}
}

Contact: kei@japanesethe72.com · X @Japanesethe72 · Substack