CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AI-MO /NuminaMath-CoT Dataset Card for NuminaMath CoT Dataset Summary Approximately 860k math problems, where each solution is formatted in a Chain of Thought (CoT) manner. The sources of the dataset range from Chinese high school math exercises to US and international mathematics olympiad competition problems. The data were primarily collected from online exam paper PDFs and mathematics discussion forums. The processing steps include (a) OCR from the original PDFs, (b) segmentation… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/NuminaMath-CoT.texttext-generation100K<n<1M603 likes234k downloads2y agoHugging Face02AI-MO /NuminaMath-1.5 Dataset Card for NuminaMath 1.5 Dataset Summary This is the second iteration of the popular NuminaMath dataset, bringing high quality post-training data for approximately 900k competition-level math problems. Each solution is formatted in a Chain of Thought (CoT) manner. The sources of the dataset range from Chinese high school math exercises to US and international mathematics olympiad competition problems. The data were primarily collected from online exam paper PDFs… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/NuminaMath-1.5.texttext-generation100K<n<1M194 likes53k downloads8mo agoHugging Face03AI-MO /aimo-validation-aime Dataset Card for AIMO Validation AIME All 90 problems come from AIME 22, AIME 23, and AIME 24, and have been extracted directly from the AOPS wiki page https://artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions This dataset serves as an internal validation set during our participation in the AIMO progress prize competition. Using data after 2021 is to avoid potential overlap with the MATH training set. Here are the different columns in the dataset: problem: the… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/aimo-validation-aime.textn<1K68 likes37k downloads1y agoHugging Face04AI-MO /aops AoPS: Art of Problem Solving Competition Mathematics Dataset Description This dataset is a collection of 80,661 competition mathematics problems and solutions obtained from the Art of Problem Solving (AoPS) community wiki and forums. It covers a wide range of mathematical contests and olympiads, including problems from events such as AIME, BAMO, IMO, and various national and memorial competitions. The dataset was curated by AI-MO (Project Numina), an initiative focused on… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/aops.text10K<n<100K6 likes30k downloads6mo agoHugging Face05datamatastudios /ai-model-popularity Datamata AI Model Popularity Index Weekly popularity of the most-downloaded and trending Hugging Face models: trailing downloads, likes, the model's task and its trending rank. One row per model from the most recent weekly snapshot. Latest snapshot: 2026-09-20 Models in this release: 50 Updated: weekly Licence: CC BY 4.0 — free to use and adapt, including commercially, with attribution. Source & methodology: https://www.datamatastudios.com/datasets Quickstart… See the full description on the dataset page: https://huggingface.co/datasets/datamatastudios/ai-model-popularity.tabularn<1K0 likes17k downloads3d agoHugging Face06AI-MO /aimo-validation-amc Dataset Card for AIMO Validation AMC All 83 come from AMC12 2022, AMC12 2023, and have been extracted from the AOPS wiki page https://artofproblemsolving.com/wiki/index.php/AMC_12_Problems_and_Solutions This dataset serves as an internal validation set during our participation in the AIMO progress prize competition. Using data after 2021 is to avoid potential overlap with the MATH training set. Here are the different columns in the dataset: problem: the modified problem statement… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/aimo-validation-amc.tabularn<1K19 likes11k downloads1y agoHugging Face07AI-MO /NuminaMath-TIR Dataset Card for NuminaMath CoT Dataset Summary Tool-integrated reasoning (TIR) plays a crucial role in this competition. However, collecting and annotating such data is both costly and time-consuming. To address this, we selected approximately 70k problems from the NuminaMath-CoT dataset, focusing on those with numerical outputs, most of which are integers. We then utilized a pipeline leveraging GPT-4 to generate TORA-like reasoning paths, executing the code and… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/NuminaMath-TIR.texttext-generation10K<n<100K158 likes8.3k downloads2y agoHugging Face08AI-MO /aimo-validation-math-level-5 Dataset Card for AIMO Validation MATH Level 5 A subset of level 5 problems from https://huggingface.co/datasets/lighteval/MATH We have extracted the final answer from boxed, and only keep those with integer outputs. textn<1K11 likes2.5k downloads2y agoHugging Face09AI-MO /NuminaMath-LEAN Dataset Card for NuminaMath-LEAN Dataset Summary NuminaMath-LEAN is a large-scale dataset of 100K mathematical competition problems formalized in Lean 4. It is derived from a challenging subset of the NuminaMath 1.5 dataset, focusing on problems from prestigious competitions like the IMO and USAMO. It represents the largest collection of human-annotated formal statements and proofs designed for training and evaluating automated theorem provers. This is also the dataset… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/NuminaMath-LEAN.text100K<n<1M62 likes2.1k downloads1y agoHugging Face10kaggle-aimo /amc_filteredtabular1K<n<10K0 likes1.3k downloads2y agoHugging Face11AI-MO /aimo-validation-math-level-4 Dataset Card for AIMO Validation MATH Level 4 A subset of level 4 problems from https://huggingface.co/datasets/lighteval/MATH We have extracted the final answer from boxed, and only keep those with integer outputs. textn<1K4 likes1.3k downloads2y agoHugging Face12kaggle-aimo /aime_filteredtabularn<1K1 likes753 downloads2y agoHugging Face13aimosprite /prompt-swap-mixed12-5xlr-e1-mxfp4-mergedtabularn<1K0 likes597 downloads6mo agoHugging Face14aimosprite /prompt-swap-mixed12-5xlr-e2-mxfp4-mergedtabularn<1K0 likes586 downloads6mo agoHugging Face151231czx /rlhflow_mix_w_aimo_mathtext1M<n<10M0 likes533 downloads2y agoHugging Face16AI-MO /CombiBench CombiBench CombiBench is the first benchmark focused on combinatorial problems, based on the formal language Lean 4. CombiBench is a manually produced benchmark, including 100 combinatorial mathematics problems of varying difficulty and knowledge levels. It aims to provide a benchmark for evaluating the combinatorial mathematics capabilities of automated theorem proving systems to advance the field. For problems that require providing a solution first and then… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/CombiBench.textn<1K12 likes486 downloads1y agoHugging Face17AI-MO /olympiads-ref-basetext10K<n<100K2 likes457 downloads11mo agoHugging Face18AI-MO /PolyUniMath PolyUniMath Dataset Summary PolyUniMath is a large multilingual mathematics dataset of question-solution-answer pairs extracted from mathematical PDFs. The dataset is designed for training and studying natural-language mathematical reasoning, with a strong emphasis on university-level content. Sample count: approximately 4 million Q&A pairs Main focus: university-level mathematics Format: problem, optional choices, solution, final answer, and auxiliary… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/PolyUniMath.tabular1M<n<10M1 likes407 downloads1mo agoHugging Face19lighteval /aimo_progress_prize_1textn<1K0 likes382 downloads2y agoHugging Face20aimosprite /prompt-swap-medium12-e2-mxfp4-mergedtabularn<1K0 likes373 downloads6mo agoHugging Face21AI-MO /Kimina-Prover-Promptset Kimina-Prover-Promptset Kimina-Prover-Promptset is a curated subset of NuminaMath-LEAN, designed for reinforcement learning (RL) training of formal theorem provers in Lean 4. Compared to the full dataset, this subset contains fewer problems but with higher difficulty. NuminaMath-LEAN is filtered and preprocessed as follows to create this dataset: Remove easy problems with a historical win rate above 0.5 to only keeep challenging statements in the dataset. Generate variants of… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/Kimina-Prover-Promptset.text10K<n<100K2 likes367 downloads1y agoHugging Face22aimo-interp /val-sampletabularn<1K1 likes314 downloads4mo agoHugging Face23AI-MO /minif2f_test MiniF2F Dataset Usage The evaluation results of Kimina-Prover presented in our work are all based on this MiniF2F test set. Improvements We corrected several erroneous formalizations, since the original formal statements could not be proven. We list them in the following table. All our improvements are made based on the MiniF2F test set provided by DeepseekProverV1.5, which applies certain modifications to the original dataset to adapt it to the Lean 4.… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/minif2f_test.textn<1K9 likes295 downloads1y agoHugging Face24AIMO-Corpus /PolyMath Dataset Card for PolyMath Dataset Summary PolyMath is a curated dataset of 11,090 high-difficulty mathematical problems designed for training reasoning models. Built for the AIMO Math Corpus Prize. Existing math datasets (NuminaMath-1.5, OpenMathReasoning) suffer from high noise rates in their hardest samples and largely unusable proof-based problems. PolyMath addresses both issues through: Data scraping: problems sourced from official competition PDFs absent from… See the full description on the dataset page: https://huggingface.co/datasets/AIMO-Corpus/PolyMath.textquestion-answering10K<n<100K2 likes283 downloads8mo agoHugging Face25AI-MO /CC-MAIN-pdf-urltext100M<n<1B1 likes260 downloads1y agoHugging Face26aimo-interp /augmented-sample-math-aggtabularn<1K0 likes228 downloads4mo agoHugging Face27aimo-interp /aimo-interp-challenge-sample-fulltabularn<1K0 likes207 downloads4mo agoHugging Face28gemmozero /ai-models-2026 AI Models & Releases 2026 AI model releases, benchmarks, capabilities. Updated daily via automated collection pipeline. Part of the Legion Data Factory — historical AI ecosystem datasets 2026. Methodology Automated collection from public sources (HackerNews, RSS feeds, APIs). Updated daily via cron job. Raw data, minimal processing. License CC BY 4.0 🔑 API Access — Updated Daily Live data via Legion AI API | Documentation Free: 100… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-models-2026.text1K<n<10K0 likes182 downloads11h agoHugging Face29aimo-interp /train-main-v2 AIMO Interpretability Challenge train set covering models from the main track See https://aimo-interp.github.io for details about the competition. Features feature type description model_id string Hugging Face checkpoint. One of five: Qwen/Qwen3.5-4B, Skywork/Skywork-OR1-Math-7B, allenai/Olmo-3-7B-Think, deepseek-ai/DeepSeek-R1-0528-Qwen3-8B, openai/gpt-oss-120b. reasoning_effort string Decoding condition: default for the four fixed-effort models, low… See the full description on the dataset page: https://huggingface.co/datasets/aimo-interp/train-main-v2.textn<1K1 likes142 downloads6d agoHugging Face30UR-xiaoyang /AIMO3_CoT AIMO3 CoT Dataset 数据集来源与目的 (Dataset Origin and Purpose) 本数据集源自 Kaggle 竞赛 AI Mathematical Olympiad - Progress Prize 3。 动机 (Motivation) 原始数据集仅包含问题和答案,缺乏思维链(Chain of Thought, CoT)。直接使用原始数据训练如 DeepSeek Math 或 Qwen Math 等模型效果不佳。因此,本项目的目的是利用 Gemini 3 Pro 为这些问题补充详细的 CoT,以提升模型在数学推理任务上的表现。 CoT 格式 (CoT Format) 生成的 CoT 遵循 ReAct 风格的推理过程,并使用中文叙述: Thought: 分析问题并规划下一步。 Code: 编写 Python 代码进行计算或验证。 Observation: 代码的执行输出。 ... (重复上述步骤) Final Answer: 得出的最终答案。… See the full description on the dataset page: https://huggingface.co/datasets/UR-xiaoyang/AIMO3_CoT.documentquestion-answeringn<1K1 likes126 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.