CoolFace
Datasetpublic

hendrydong/livemath-v7-2603-2606

LiveMathematicianBench v7 — arXiv 2026-03 … 2026-06 An automated, refreshable benchmark of research-level mathematics multiple-choice questions, generated from newly published arXiv math papers (contamination-resistant by construction). Columbia University & Microsoft Research. Each question is grounded in a theorem from a recent paper; distractors are adversarially crafted from the proof sketch, and a multi-stage hardness pipeline keeps the final set difficult for frontier… See the full description on the dataset page: https://huggingface.co/datasets/hendrydong/livemath-v7-2603-2606.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes64downloads
5 commits on main
94848742mo ago

Add per-theorem-type accuracy breakdown to leaderboard

hendrydong
e9e774d2mo ago

Add leaderboard json

hendrydong
511a8632mo ago

Add 19-model leaderboard

hendrydong
74f9dc22mo ago

LiveMath v7: arXiv 2603-2606 (full/ge5/hard, gpt-5.4 medium 42.5%)

hendrydong
cb9c00b2mo ago

initial commit

hendrydong