CoolFace
Datasetpublic

hendrydong/livemath-v7-2603-2606

LiveMathematicianBench v7 — arXiv 2026-03 … 2026-06 An automated, refreshable benchmark of research-level mathematics multiple-choice questions, generated from newly published arXiv math papers (contamination-resistant by construction). Columbia University & Microsoft Research. Each question is grounded in a theorem from a recent paper; distractors are adversarially crafted from the proof sketch, and a multi-stage hardness pipeline keeps the final set difficult for frontier… See the full description on the dataset page: https://huggingface.co/datasets/hendrydong/livemath-v7-2603-2606.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes68downloads
Dataset Card

LiveMathematicianBench v7 — arXiv 2026-03 … 2026-06

An automated, refreshable benchmark of research-level mathematics multiple-choice questions, generated from newly published arXiv math papers (contamination-resistant by construction). Columbia University & Microsoft Research.

Each question is grounded in a theorem from a recent paper; distractors are adversarially crafted from the proof sketch, and a multi-stage hardness pipeline keeps the final set difficult for frontier models.

Layout

Per month (202603/, 202604/, 202605/, 202606/):

  • —full/qaEval_<month>_full.json — all generated MCQs (v5)
  • —ge5/qaEval_<month>_ge5.json — quality-filtered (rubric ALS+TAS+GPS+DQS total_score ≥ 5)
  • —hard/qaEval_<month>_ge5_hard.json — final ge5-hard subset (one hardest item per source theorem)
  • —hard/accuracy_test_<month>_medium_filter2.json — final independent retest results
  • —summary.json — per-month funnel

Results — GPT-5.4 (medium) final hard-set accuracy

Monthhard questionsaccuracy
2026-0311237.5%
2026-0410842.6%
2026-059548.4%
2026-0611642.2%
Overall43142.5%

Lower accuracy = harder. See summary.json for the full generation → ge5 → stem-nontrivial → hard funnel.