CoolFace
Datasetpublic

RUC-AIBOX/OlymMATH

Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models This is the official huggingface repository for Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models by Haoxiang Sun, Yingqian Min, Zhipeng Chen, Wayne Xin Zhao, Zheng Liu, Zhongyuan Wang, Lei Fang, and Ji-Rong Wen. We have also released the OlymMATH-eval dataset on HuggingFace 🤗, together with a data visualization tool OlymMATH-demo… See the full description on the dataset page: https://huggingface.co/datasets/RUC-AIBOX/OlymMATH.

sourceHugging Facemitupdated 9mo agoView on Hugging Face
14likes1kdownloads
Dataset Card

Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models

This is the official huggingface repository for Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models by Haoxiang Sun, Yingqian Min, Zhipeng Chen, Wayne Xin Zhao, Zheng Liu, Zhongyuan Wang, Lei Fang, and Ji-Rong Wen.

We have also released the OlymMATH-eval dataset on HuggingFace 🤗, together with a data visualization tool OlymMATH-demo, currently available in HuggingFace Spaces.

You can find more information on GitHub.

Citation

If you find this helpful in your research, please give a 🌟 to our repo and consider citing

@misc{sun2025challengingboundariesreasoningolympiadlevel,
      title={Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models}, 
      author={Haoxiang Sun and Yingqian Min and Zhipeng Chen and Wayne Xin Zhao and Zheng Liu and Zhongyuan Wang and Lei Fang and Ji-Rong Wen},
      year={2025},
      eprint={2503.21380},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2503.21380}, 
}