hendrydong/livemath-v7-2603-2606
LiveMathematicianBench v7 — arXiv 2026-03 … 2026-06 An automated, refreshable benchmark of research-level mathematics multiple-choice questions, generated from newly published arXiv math papers (contamination-resistant by construction). Columbia University & Microsoft Research. Each question is grounded in a theorem from a recent paper; distractors are adversarially crafted from the proof sketch, and a multi-stage hardness pipeline keeps the final set difficult for frontier… See the full description on the dataset page: https://huggingface.co/datasets/hendrydong/livemath-v7-2603-2606.
LiveMathematicianBench v7 — arXiv 2026-03 … 2026-06
An automated, refreshable benchmark of research-level mathematics multiple-choice questions, generated from newly published arXiv math papers (contamination-resistant by construction). Columbia University & Microsoft Research.
Each question is grounded in a theorem from a recent paper; distractors are adversarially crafted from the proof sketch, and a multi-stage hardness pipeline keeps the final set difficult for frontier models.
Layout
Per month (202603/, 202604/, 202605/, 202606/):
full/qaEval_<month>_full.json— all generated MCQs (v5)ge5/qaEval_<month>_ge5.json— quality-filtered (rubric ALS+TAS+GPS+DQS total_score ≥ 5)hard/qaEval_<month>_ge5_hard.json— final ge5-hard subset (one hardest item per source theorem)hard/accuracy_test_<month>_medium_filter2.json— final independent retest resultssummary.json— per-month funnel
Results — GPT-5.4 (medium) final hard-set accuracy
Lower accuracy = harder. See summary.json for the full generation → ge5 → stem-nontrivial → hard funnel.
