datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OlympiadBench
OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems[ACL 2024]
📖 arXiv | GitHub
Note: We have made adjustments to the image content in the multimodal portion of the dataset and fixed previous issues where some images in the English physics subset were not displayed properly. If your usage involves images, please re-download the dataset (we recommend all users to download the latest version).
Additionally, some entries… See the full description on the dataset page: https://huggingface.co/datasets/Hothan/OlympiadBench.OlympiadBencholympiadbenchsimplerl-OlympiadBenchOlympiadBenchOlympiadBench-Reasoning-Paths
News
🌟🌟🌟 Try this dataset in our HuggingFace Space!
🥳🥳🥳 Thrilled to share that this NeurIPS paper was selected as 🏆 #1 Paper of the Day on Oct. 20th!
Sampled Reasoning Paths for the OlympiadBench dataset
This dataset contains sampled reasoning paths for the OlympiadBench dataset, released as part of the NeurIPS 2025 paper: "A Theoretical Study on Bridging Internal Probability and Self-Consistency for LLM Reasoning" (Arxiv).… See the full description on the dataset page: https://huggingface.co/datasets/WNJXYK/OlympiadBench-Reasoning-Paths.OlympiadBench-official
OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
📖 arXiv | GitHub
Dataset Description
OlympiadBench is an Olympiad-level bilingual multimodal scientific benchmark, featuring 8,476 problems from Olympiad-level mathematics and physics competitions, including the Chinese college entrance exam. Each problem is detailed with expert-level annotations for step-by-step reasoning. Notably, the best-performing… See the full description on the dataset page: https://huggingface.co/datasets/lscpku/OlympiadBench-official.olympiadbenchHLE_SFT_OlympiadBench
HLE_SFT_OlympiadBench
HLE(Humanity's Last Exam)競技用の数学、物理問題SFTデータセット(このデータセットは物理分野に関するものです)
概要
OlympiadBench の公開データセットを整形して利用しています。
データ形式
{
"id": 0,
"question": "問題文",
"output": "CoT (Chain of Thought)",
"answer": "最終的な回答"
}
olympiadbencholympiad-bench-imo-math-boxed-825-v2-21-08-2024
OlympiadBench Data set used in the Putnam-AXIOM Paper
The Putnam-AXIOM dataset is a benchmark for measuring advanced mathematical reasoning in large language models (LLMs).
It includes challenging mathematical problems from the William Lowell Putnam Mathematical Competition, with both original problems and functional variations to address data contamination. The dataset aims to provide rigorous evaluations by requiring models to answer in boxed format, simplifying automatic answer… See the full description on the dataset page: https://huggingface.co/datasets/brando/olympiad-bench-imo-math-boxed-825-v2-21-08-2024.olympiadbench
Dataset Card for "olympiadbench"
More Information needed
OlympiadBenchOlympiadBench-Math-Ko
Details
This is a Korean translated version of OlympiadBench dataset ([Hothan/OlympiadBench]).
Especially, "OE_TO_maths_en_COMP" split. (Open-ended questions, Text-only, Math problems, English, Competition problems)
OlympiadBenchOlympiadBench-OE-CoT-num10Llama-3.2-1B-Instruct-uPRM-70B-T80-olympiadbench-best_of_n-completionsolympiadbench_testHLE_RL_OlympiadBenchQwen2.5-1.5B-Instruct-uPRM-32B-T80-olympiadbench-best_of_n-completionsOlympiadBench-OEolympiadbenchQwen2.5-1.5B-Instruct-uPRM-70B-T80-olympiadbench-best_of_n-completionsloong_advanced_physics_Olympiadbench
Dataset Card for "loong_advanced_physics_Olympiadbench"
More Information needed
Llama-3.2-1B-Instruct-uPRM-32B-T80-olympiadbench-best_of_n-completionsolympiadbench_qwen_mathOlympiadBench_TextOnly_MathOlympiadBencholympiadbench_math_textonlytest_olympiadbench
