CoolFace
Datasetpublic

JacobLinCool/taiwan-national-exams-results

Benchmark results produced by any-to-bench. One subset here is one taker configuration — a single model at a single reasoning effort — sat against the exams in another dataset repo. Every row names the exam repo and subset it was earned against, so results from several corpora, and from several people, can live side by side. results-index.json — the catalog: one headline row per configuration results-<entry>/entry.json — that configuration's per-paper scores results-<entry>/raw/<subset>/ —… See the full description on the dataset page: https://huggingface.co/datasets/JacobLinCool/taiwan-national-exams-results.

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes170downloads
Dataset Card

<!-- a2b:results:header:start --> Benchmark results produced by any-to-bench. One subset here is one taker configuration — a single model at a single reasoning effort — sat against the exams in another dataset repo. Every row names the exam repo and subset it was earned against, so results from several corpora, and from several people, can live side by side.

  • —results-index.json — the catalog: one headline row per configuration
  • —results-<entry>/entry.json — that configuration's per-paper scores
  • —results-<entry>/raw/<subset>/ — the byte-faithful bench.json, answer sheet and grade report behind those scores
  • —the viewer table — one row per graded question

Resource-backed runs also publish their actual file/byte exposure and deterministic citation checks. Citations are evidence metadata only and never change the score.

Explore it as a leaderboard: https://jacoblincool.github.io/any-to-bench/results.html

python
from datasets import load_dataset
ds = load_dataset("JacobLinCool/taiwan-national-exams-results", "results-<entry>", split="test")

How to read these numbers

  • —Judged questions depend on the judge model, which is named per entry. Two configurations graded by different judges are not strictly comparable on their judged half; the rule-graded half is deterministic and always comparable.
  • —A score is one sample per paper unless the entry's repeat is above 1. The run-to-run spread of a single sample is unknown.
  • —Input-token counts are not comparable across backends. codex: reports cached tokens inside input_tokens; claude: and agy: report cache reads separately under cache_read_tokens. Output tokens and wall time are the measures that mean the same thing for every taker; the raw per-phase counts are published unaltered so you can judge for yourself.
  • —Token counts for agentic takers are approximate, and wall time depends on how many runs shared the machine — see each entry's note.
  • —Percentages are over what the taker was actually asked. When two entries cover different papers, their percentages have different denominators. <!-- a2b:results:header:end -->

<!-- a2b:results:board:start -->

Leaderboard

#ModelEffortPapersScore%Rule-graded %Solve output tokensSolve s
1codex:gpt-5.6-solhigh161720.75/180095.6%94.0%240,2307,993
2codex:gpt-5.6-solxhigh161714.5/180095.2%93.9%352,52711,455
3codex:gpt-5.6-solmedium161705.5/180094.8%93.4%166,1406,087
4codex:gpt-5.6-sollow161656.5/180092.0%90.1%99,5784,301
5codex:gpt-5.6-lunaxhigh161584.75/180088.0%86.1%532,04811,973
6codex:gpt-5.6-lunahigh161556.5/180086.5%85.0%330,0408,120
7codex:gpt-5.6-lunamedium161493.25/180083.0%81.2%134,1274,550
8codex:gpt-5.6-lunalow161400.75/180077.8%77.9%71,5013,402

<!-- a2b:results:board:end -->

<!-- a2b:results:entry:codex-gpt-5-6-sol-low:start -->

codex-gpt-5-6-sol-low

Modelcodex:gpt-5.6-sol
Effortlow
Papers16 from JacobLinCool/taiwan-national-exams
Judgesopenai:gpt-5.6-sol
Runs per paper1
Ingested byany-to-bench 0.1.1
Ran2026-08-23
Notethirty-two agents in parallel on one machine; hermetic codex sessions — no network, no built-in search tool, no approval escalation
PaperScore%Rule-gradedJudged
bar-115-civil-law132/16082.5%82.5%–
bar-115-commercial-law110/14078.6%78.6%–
bar-115-criminal-law132/15088.0%88.0%–
bar-115-public-law142/15094.7%94.7%–
counselor-115-assessment100/100100.0%100.0%100.0%
counselor-115-group-counseling96.25/10096.2%92.5%100.0%
counselor-115-mental-health98.75/10098.8%97.5%100.0%
counselor-115-practice-ethics96.25/10096.2%92.5%100.0%
counselor-115-psychology-foundations98.75/10098.8%97.5%100.0%
counselor-115-theories97/10097.0%95.0%99.0%
cpa-115-advanced-accounting93/10093.0%92.0%94.0%
cpa-115-auditing98/10098.0%96.0%100.0%
cpa-115-corporate-securities-law92.5/10092.5%96.0%89.0%
cpa-115-cost-management-accounting98/10098.0%96.0%100.0%
cpa-115-intermediate-accounting87.5/10087.5%92.0%83.0%
cpa-115-tax-law84.5/10084.5%84.0%85.0%

a2b results fetch JacobLinCool/taiwan-national-exams-results --entry codex-gpt-5-6-sol-low -o results <!-- a2b:results:entry:codex-gpt-5-6-sol-low:end -->

<!-- a2b:results:entry:codex-gpt-5-6-sol-medium:start -->

codex-gpt-5-6-sol-medium

Modelcodex:gpt-5.6-sol
Effortmedium
Papers16 from JacobLinCool/taiwan-national-exams
Judgesopenai:gpt-5.6-sol
Runs per paper1
Ingested byany-to-bench 0.1.1
Ran2026-08-23
Notethirty-two agents in parallel on one machine; hermetic codex sessions — no network, no built-in search tool, no approval escalation
PaperScore%Rule-gradedJudged
bar-115-civil-law146/16091.2%91.2%–
bar-115-commercial-law126/14090.0%90.0%–
bar-115-criminal-law138/15092.0%92.0%–
bar-115-public-law140/15093.3%93.3%–
counselor-115-assessment100/100100.0%100.0%100.0%
counselor-115-group-counseling96.25/10096.2%92.5%100.0%
counselor-115-mental-health97.5/10097.5%95.0%100.0%
counselor-115-practice-ethics95/10095.0%90.0%100.0%
counselor-115-psychology-foundations100/100100.0%100.0%100.0%
counselor-115-theories96.25/10096.2%92.5%100.0%
cpa-115-advanced-accounting92.5/10092.5%96.0%89.0%
cpa-115-auditing100/100100.0%100.0%100.0%
cpa-115-corporate-securities-law95.5/10095.5%96.0%95.0%
cpa-115-cost-management-accounting96/10096.0%96.0%96.0%
cpa-115-intermediate-accounting96/10096.0%92.0%100.0%
cpa-115-tax-law90.5/10090.5%92.0%89.0%

a2b results fetch JacobLinCool/taiwan-national-exams-results --entry codex-gpt-5-6-sol-medium -o results <!-- a2b:results:entry:codex-gpt-5-6-sol-medium:end -->

<!-- a2b:results:entry:codex-gpt-5-6-sol-high:start -->

codex-gpt-5-6-sol-high

Modelcodex:gpt-5.6-sol
Efforthigh
Papers16 from JacobLinCool/taiwan-national-exams
Judgesopenai:gpt-5.6-sol
Runs per paper1
Ingested byany-to-bench 0.1.1
Ran2026-08-23
Notethirty-two agents in parallel on one machine; hermetic codex sessions — no network, no built-in search tool, no approval escalation
PaperScore%Rule-gradedJudged
bar-115-civil-law142/16088.8%88.8%–
bar-115-commercial-law124/14088.6%88.6%–
bar-115-criminal-law140/15093.3%93.3%–
bar-115-public-law142/15094.7%94.7%–
counselor-115-assessment100/100100.0%100.0%100.0%
counselor-115-group-counseling97.5/10097.5%95.0%100.0%
counselor-115-mental-health97.5/10097.5%95.0%100.0%
counselor-115-practice-ethics96.25/10096.2%92.5%100.0%
counselor-115-psychology-foundations97.5/10097.5%95.0%100.0%
counselor-115-theories97.5/10097.5%95.0%100.0%
cpa-115-advanced-accounting98/10098.0%96.0%100.0%
cpa-115-auditing100/100100.0%100.0%100.0%
cpa-115-corporate-securities-law97/10097.0%100.0%94.0%
cpa-115-cost-management-accounting98/10098.0%96.0%100.0%
cpa-115-intermediate-accounting98/10098.0%96.0%100.0%
cpa-115-tax-law95.5/10095.5%100.0%91.0%

a2b results fetch JacobLinCool/taiwan-national-exams-results --entry codex-gpt-5-6-sol-high -o results <!-- a2b:results:entry:codex-gpt-5-6-sol-high:end -->

<!-- a2b:results:entry:codex-gpt-5-6-sol-xhigh:start -->

codex-gpt-5-6-sol-xhigh

Modelcodex:gpt-5.6-sol
Effortxhigh
Papers16 from JacobLinCool/taiwan-national-exams
Judgesopenai:gpt-5.6-sol
Runs per paper1
Ingested byany-to-bench 0.1.1
Ran2026-08-23
Notethirty-two agents in parallel on one machine; hermetic codex sessions — no network, no built-in search tool, no approval escalation
PaperScore%Rule-gradedJudged
bar-115-civil-law146/16091.2%91.2%–
bar-115-commercial-law132/14094.3%94.3%–
bar-115-criminal-law138/15092.0%92.0%–
bar-115-public-law140/15093.3%93.3%–
counselor-115-assessment98.75/10098.8%97.5%100.0%
counselor-115-group-counseling97.5/10097.5%95.0%100.0%
counselor-115-mental-health97.5/10097.5%95.0%100.0%
counselor-115-practice-ethics96.25/10096.2%92.5%100.0%
counselor-115-psychology-foundations98.75/10098.8%97.5%100.0%
counselor-115-theories96.25/10096.2%92.5%100.0%
cpa-115-advanced-accounting96.5/10096.5%96.0%97.0%
cpa-115-auditing100/100100.0%100.0%100.0%
cpa-115-corporate-securities-law95.5/10095.5%96.0%95.0%
cpa-115-cost-management-accounting98/10098.0%96.0%100.0%
cpa-115-intermediate-accounting94/10094.0%88.0%100.0%
cpa-115-tax-law89.5/10089.5%96.0%83.0%

a2b results fetch JacobLinCool/taiwan-national-exams-results --entry codex-gpt-5-6-sol-xhigh -o results <!-- a2b:results:entry:codex-gpt-5-6-sol-xhigh:end -->

<!-- a2b:results:entry:codex-gpt-5-6-luna-low:start -->

codex-gpt-5-6-luna-low

Modelcodex:gpt-5.6-luna
Effortlow
Papers16 from JacobLinCool/taiwan-national-exams
Judgesopenai:gpt-5.6-sol
Runs per paper1
Ingested byany-to-bench 0.1.1
Ran2026-08-23
Notethirty-two agents in parallel on one machine; hermetic codex sessions — no network, no built-in search tool, no approval escalation
PaperScore%Rule-gradedJudged
bar-115-civil-law128/16080.0%80.0%–
bar-115-commercial-law90/14064.3%64.3%–
bar-115-criminal-law110/15073.3%73.3%–
bar-115-public-law122/15081.3%81.3%–
counselor-115-assessment97.75/10097.8%97.5%98.0%
counselor-115-group-counseling92.75/10092.8%87.5%98.0%
counselor-115-mental-health91.5/10091.5%85.0%98.0%
counselor-115-practice-ethics90.75/10090.8%87.5%94.0%
counselor-115-psychology-foundations94.25/10094.2%92.5%96.0%
counselor-115-theories91.75/10091.8%87.5%96.0%
cpa-115-advanced-accounting47/10047.0%60.0%34.0%
cpa-115-auditing90/10090.0%84.0%96.0%
cpa-115-corporate-securities-law80.5/10080.5%96.0%65.0%
cpa-115-cost-management-accounting78/10078.0%88.0%68.0%
cpa-115-intermediate-accounting42.5/10042.5%48.0%37.0%
cpa-115-tax-law54/10054.0%56.0%52.0%

a2b results fetch JacobLinCool/taiwan-national-exams-results --entry codex-gpt-5-6-luna-low -o results <!-- a2b:results:entry:codex-gpt-5-6-luna-low:end -->

<!-- a2b:results:entry:codex-gpt-5-6-luna-medium:start -->

codex-gpt-5-6-luna-medium

Modelcodex:gpt-5.6-luna
Effortmedium
Papers16 from JacobLinCool/taiwan-national-exams
Judgesopenai:gpt-5.6-sol
Runs per paper1
Ingested byany-to-bench 0.1.1
Ran2026-08-23
Notethirty-two agents in parallel on one machine; hermetic codex sessions — no network, no built-in search tool, no approval escalation
PaperScore%Rule-gradedJudged
bar-115-civil-law120/16075.0%75.0%–
bar-115-commercial-law100/14071.4%71.4%–
bar-115-criminal-law110/15073.3%73.3%–
bar-115-public-law130/15086.7%86.7%–
counselor-115-assessment98.75/10098.8%97.5%100.0%
counselor-115-group-counseling91.5/10091.5%85.0%98.0%
counselor-115-mental-health95/10095.0%90.0%100.0%
counselor-115-practice-ethics92.25/10092.2%87.5%97.0%
counselor-115-psychology-foundations99/10099.0%100.0%98.0%
counselor-115-theories94.75/10094.8%92.5%97.0%
cpa-115-advanced-accounting60.5/10060.5%64.0%57.0%
cpa-115-auditing97.5/10097.5%96.0%99.0%
cpa-115-corporate-securities-law82/10082.0%84.0%80.0%
cpa-115-cost-management-accounting88/10088.0%88.0%88.0%
cpa-115-intermediate-accounting64.5/10064.5%76.0%53.0%
cpa-115-tax-law69.5/10069.5%68.0%71.0%

a2b results fetch JacobLinCool/taiwan-national-exams-results --entry codex-gpt-5-6-luna-medium -o results <!-- a2b:results:entry:codex-gpt-5-6-luna-medium:end -->

<!-- a2b:results:entry:codex-gpt-5-6-luna-high:start -->

codex-gpt-5-6-luna-high

Modelcodex:gpt-5.6-luna
Efforthigh
Papers16 from JacobLinCool/taiwan-national-exams
Judgesopenai:gpt-5.6-sol
Runs per paper1
Ingested byany-to-bench 0.1.1
Ran2026-08-23
Notethirty-two agents in parallel on one machine; hermetic codex sessions — no network, no built-in search tool, no approval escalation
PaperScore%Rule-gradedJudged
bar-115-civil-law126/16078.8%78.8%–
bar-115-commercial-law108/14077.1%77.1%–
bar-115-criminal-law116/15077.3%77.3%–
bar-115-public-law134/15089.3%89.3%–
counselor-115-assessment100/100100.0%100.0%100.0%
counselor-115-group-counseling93.75/10093.8%87.5%100.0%
counselor-115-mental-health95/10095.0%90.0%100.0%
counselor-115-practice-ethics92.75/10092.8%87.5%98.0%
counselor-115-psychology-foundations98.75/10098.8%97.5%100.0%
counselor-115-theories95.75/10095.8%92.5%99.0%
cpa-115-advanced-accounting77.5/10077.5%80.0%75.0%
cpa-115-auditing99.5/10099.5%100.0%99.0%
cpa-115-corporate-securities-law87.5/10087.5%92.0%83.0%
cpa-115-cost-management-accounting89/10089.0%92.0%86.0%
cpa-115-intermediate-accounting73.5/10073.5%72.0%75.0%
cpa-115-tax-law69.5/10069.5%80.0%59.0%

a2b results fetch JacobLinCool/taiwan-national-exams-results --entry codex-gpt-5-6-luna-high -o results <!-- a2b:results:entry:codex-gpt-5-6-luna-high:end -->

<!-- a2b:results:entry:codex-gpt-5-6-luna-xhigh:start -->

codex-gpt-5-6-luna-xhigh

Modelcodex:gpt-5.6-luna
Effortxhigh
Papers16 from JacobLinCool/taiwan-national-exams
Judgesopenai:gpt-5.6-sol
Runs per paper1
Ingested byany-to-bench 0.1.1
Ran2026-08-23
Notethirty-two agents in parallel on one machine; hermetic codex sessions — no network, no built-in search tool, no approval escalation
PaperScore%Rule-gradedJudged
bar-115-civil-law136/16085.0%85.0%–
bar-115-commercial-law102/14072.9%72.9%–
bar-115-criminal-law110/15073.3%73.3%–
bar-115-public-law132/15088.0%88.0%–
counselor-115-assessment98.75/10098.8%97.5%100.0%
counselor-115-group-counseling95/10095.0%90.0%100.0%
counselor-115-mental-health96.25/10096.2%92.5%100.0%
counselor-115-practice-ethics94.75/10094.8%92.5%97.0%
counselor-115-psychology-foundations97.5/10097.5%95.0%100.0%
counselor-115-theories94.5/10094.5%95.0%94.0%
cpa-115-advanced-accounting87/10087.0%96.0%78.0%
cpa-115-auditing98.5/10098.5%100.0%97.0%
cpa-115-corporate-securities-law89/10089.0%92.0%86.0%
cpa-115-cost-management-accounting89.5/10089.5%92.0%87.0%
cpa-115-intermediate-accounting89/10089.0%84.0%94.0%
cpa-115-tax-law75/10075.0%80.0%70.0%

a2b results fetch JacobLinCool/taiwan-national-exams-results --entry codex-gpt-5-6-luna-xhigh -o results <!-- a2b:results:entry:codex-gpt-5-6-luna-xhigh:end -->