FUSE-verifiers/IMO-Shortlist-Verifications
IMO Shortlist with Qwen3-30B-A3B-Thinking-2507 This dataset contains 123 questions from the International Mathematical Olympiad (IMO) Shortlist subset of the IMO AnswerBench benchmark with 50 candidate responses generated by Qwen3-30B-A3B-Thinking-2507 for each problem. Each response has been evaluated for correctness using a mixture of DeepSeek-R1-Distill-Llama-70B and Python code to parse different answer formats, and scored by multiple LLM judges according to a 0-5 rubric.… See the full description on the dataset page: https://huggingface.co/datasets/FUSE-verifiers/IMO-Shortlist-Verifications.
IMO Shortlist with Qwen3-30B-A3B-Thinking-2507
<!-- Provide a quick summary of the dataset. -->
This dataset contains 123 questions from the International Mathematical Olympiad (IMO) Shortlist subset of the IMO AnswerBench benchmark with 50 candidate responses generated by Qwen3-30B-A3B-Thinking-2507 for each problem. Each response has been evaluated for correctness using a mixture of DeepSeek-R1-Distill-Llama-70B and Python code to parse different answer formats, and scored by multiple LLM judges according to a 0-5 rubric. The benchmark consists of modified versions of past problems in the IMO Shortlist, where the modification is done by experts to help avoid memorization.
Dataset Structure
- Split: Single split named "data"
- Number of rows: 123 HLE questions
- Generations per query: 50
Key Fields
Verifier Models
- DeepSeek-R1-Distill-Qwen-32B
- Kimi-Linear-48B-A3B-Instruct
- Llama-3.3-70B-Instruct
- Ministral-3-14B-Reasoning-2512
- Ministral-3-8B-Instruct-2512
- NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
- Qwen3-30B-A3B-Thinking-2507
- gemma-3-27b-it
- gpt-oss-20b
Source
Original IMO Shortlist problems from superhuman/imobench. Description of data provided at imobench.github.io.
