CoolFace
Datasetpublic

TakalaWang/taiwan-exams-results

Benchmark results produced by any-to-bench. One subset here is one taker configuration — a single model at a single reasoning effort — sat against the exams in another dataset repo. Every row names the exam repo and subset it was earned against, so results from several corpora, and from several people, can live side by side. results-index.json — the catalog: one headline row per configuration results-<entry>/entry.json — that configuration's per-paper scores results-<entry>/raw/<subset>/ —… See the full description on the dataset page: https://huggingface.co/datasets/TakalaWang/taiwan-exams-results.

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes75downloads
Dataset Card

<!-- a2b:results:header:start --> Benchmark results produced by any-to-bench. One subset here is one taker configuration — a single model at a single reasoning effort — sat against the exams in another dataset repo. Every row names the exam repo and subset it was earned against, so results from several corpora, and from several people, can live side by side.

  • —results-index.json — the catalog: one headline row per configuration
  • —results-<entry>/entry.json — that configuration's per-paper scores
  • —results-<entry>/raw/<subset>/ — the byte-faithful bench.json, answer sheet and grade report behind those scores
  • —the viewer table — one row per graded question

Explore it as a leaderboard: https://jacoblincool.github.io/any-to-bench/results.html

python
from datasets import load_dataset
ds = load_dataset("TakalaWang/taiwan-exams-results", "results-<entry>", split="test")

How to read these numbers

  • —Judged questions depend on the judge model, which is named per entry. Two configurations graded by different judges are not strictly comparable on their judged half; the rule-graded half is deterministic and always comparable.
  • —A score is one sample per paper unless the entry's repeat is above 1. The run-to-run spread of a single sample is unknown.
  • —Input-token counts are not comparable across backends. codex: reports cached tokens inside input_tokens; claude: reports them only under cache_read_tokens, leaving input_tokens near zero. Output tokens and wall time are the measures that mean the same thing for every taker; the raw per-phase counts are published unaltered so you can judge for yourself.
  • —Token counts for agentic takers are approximate, and wall time depends on how many runs shared the machine — see each entry's note.
  • —Percentages are over what the taker was actually asked. When two entries cover different papers, their percentages have different denominators. <!-- a2b:results:header:end -->

<!-- a2b:results:board:start -->

Leaderboard

#ModelEffortPapersScore%Rule-graded %Solve output tokensSolve s
1codex:gpt-5.6-solhigh1594/60099.0%100.0%37,941789
2codex:gpt-5.6-solxhigh1594/60099.0%100.0%30,801683
3codex:gpt-5.6-lunahigh1556.5/60092.8%90.2%20,742568
4codex:gpt-5.6-lunaxhigh1551/60091.8%100.0%28,184599
5codex:gpt-5.6-solmedium1547.75/60091.3%89.0%17,718471
6codex:gpt-5.6-sollow1525.75/60087.6%89.0%10,436315
7codex:gpt-5.6-lunalow1308/60051.3%52.5%8,106320
8codex:gpt-5.6-lunamedium1295.5/60049.2%26.7%11,628330

<!-- a2b:results:board:end -->

<!-- a2b:results:entry:codex-gpt-5-6-sol-low:start -->

codex-gpt-5-6-sol-low

Modelcodex:gpt-5.6-sol
Effortlow
Papers1 from TakalaWang/taiwan-exams
Judgescodex:gpt-5.6-terra
Runs per paper1
Ingested byany-to-bench 0.1.1
Ran2026-08-23
NoteSingle run; independent judge codex:gpt-5.6-terra; taker and judge effort low.
PaperScore%Rule-gradedJudged
architect-114525.75/60087.6%89.0%86.1%

a2b results fetch TakalaWang/taiwan-exams-results --entry codex-gpt-5-6-sol-low -o results <!-- a2b:results:entry:codex-gpt-5-6-sol-low:end -->

<!-- a2b:results:entry:codex-gpt-5-6-luna-low:start -->

codex-gpt-5-6-luna-low

Modelcodex:gpt-5.6-luna
Effortlow
Papers1 from TakalaWang/taiwan-exams
Judgescodex:gpt-5.6-terra
Runs per paper1
Ingested byany-to-bench 0.1.1
Ran2026-08-23
NoteSingle run; independent judge codex:gpt-5.6-terra; taker and judge effort low.
PaperScore%Rule-gradedJudged
architect-114308/60051.3%52.5%50.0%

a2b results fetch TakalaWang/taiwan-exams-results --entry codex-gpt-5-6-luna-low -o results <!-- a2b:results:entry:codex-gpt-5-6-luna-low:end -->

<!-- a2b:results:entry:codex-gpt-5-6-sol-medium:start -->

codex-gpt-5-6-sol-medium

Modelcodex:gpt-5.6-sol
Effortmedium
Papers1 from TakalaWang/taiwan-exams
Judgescodex:gpt-5.6-terra
Runs per paper1
Ingested byany-to-bench 0.1.1
Ran2026-08-23
NoteSingle run; independent judge codex:gpt-5.6-terra; taker and judge effort medium.
PaperScore%Rule-gradedJudged
architect-114547.75/60091.3%89.0%93.9%

a2b results fetch TakalaWang/taiwan-exams-results --entry codex-gpt-5-6-sol-medium -o results <!-- a2b:results:entry:codex-gpt-5-6-sol-medium:end -->

<!-- a2b:results:entry:codex-gpt-5-6-luna-medium:start -->

codex-gpt-5-6-luna-medium

Modelcodex:gpt-5.6-luna
Effortmedium
Papers1 from TakalaWang/taiwan-exams
Judgescodex:gpt-5.6-terra
Runs per paper1
Ingested byany-to-bench 0.1.1
Ran2026-08-23
NoteSingle run; independent judge codex:gpt-5.6-terra; taker and judge effort medium.
PaperScore%Rule-gradedJudged
architect-114295.5/60049.2%26.7%75.0%

a2b results fetch TakalaWang/taiwan-exams-results --entry codex-gpt-5-6-luna-medium -o results <!-- a2b:results:entry:codex-gpt-5-6-luna-medium:end -->

<!-- a2b:results:entry:codex-gpt-5-6-sol-high:start -->

codex-gpt-5-6-sol-high

Modelcodex:gpt-5.6-sol
Efforthigh
Papers1 from TakalaWang/taiwan-exams
Judgescodex:gpt-5.6-terra
Runs per paper1
Ingested byany-to-bench 0.1.1
Ran2026-08-23
NoteSingle run; independent judge codex:gpt-5.6-terra; taker and judge effort high.
PaperScore%Rule-gradedJudged
architect-114594/60099.0%100.0%97.9%

a2b results fetch TakalaWang/taiwan-exams-results --entry codex-gpt-5-6-sol-high -o results <!-- a2b:results:entry:codex-gpt-5-6-sol-high:end -->

<!-- a2b:results:entry:codex-gpt-5-6-luna-high:start -->

codex-gpt-5-6-luna-high

Modelcodex:gpt-5.6-luna
Efforthigh
Papers1 from TakalaWang/taiwan-exams
Judgescodex:gpt-5.6-terra
Runs per paper1
Ingested byany-to-bench 0.1.1
Ran2026-08-23
NoteSingle run; independent judge codex:gpt-5.6-terra; taker and judge effort high.
PaperScore%Rule-gradedJudged
architect-114556.5/60092.8%90.2%95.7%

a2b results fetch TakalaWang/taiwan-exams-results --entry codex-gpt-5-6-luna-high -o results <!-- a2b:results:entry:codex-gpt-5-6-luna-high:end -->

<!-- a2b:results:entry:codex-gpt-5-6-sol-xhigh:start -->

codex-gpt-5-6-sol-xhigh

Modelcodex:gpt-5.6-sol
Effortxhigh
Papers1 from TakalaWang/taiwan-exams
Judgescodex:gpt-5.6-terra
Runs per paper1
Ingested byany-to-bench 0.1.1
Ran2026-08-23
NoteSingle run; independent judge codex:gpt-5.6-terra; taker and judge effort xhigh.
PaperScore%Rule-gradedJudged
architect-114594/60099.0%100.0%97.9%

a2b results fetch TakalaWang/taiwan-exams-results --entry codex-gpt-5-6-sol-xhigh -o results <!-- a2b:results:entry:codex-gpt-5-6-sol-xhigh:end -->

<!-- a2b:results:entry:codex-gpt-5-6-luna-xhigh:start -->

codex-gpt-5-6-luna-xhigh

Modelcodex:gpt-5.6-luna
Effortxhigh
Papers1 from TakalaWang/taiwan-exams
Judgescodex:gpt-5.6-terra
Runs per paper1
Ingested byany-to-bench 0.1.1
Ran2026-08-23
NoteSingle run; independent judge codex:gpt-5.6-terra; taker and judge effort xhigh.
PaperScore%Rule-gradedJudged
architect-114551/60091.8%100.0%82.5%

a2b results fetch TakalaWang/taiwan-exams-results --entry codex-gpt-5-6-luna-xhigh -o results <!-- a2b:results:entry:codex-gpt-5-6-luna-xhigh:end -->