CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01opencompass /AIME2025 AIME 2025 Dataset Dataset Description This dataset contains problems from the American Invitational Mathematics Examination (AIME) 2025-I & II. textquestion-answeringn<1K56 likes13k downloads2y agoHugging Face02ulamai /AIME-Plus-Plus AIME++ Sample AIME++ is Ulam AI's exact-answer mathematical reasoning environment. It keeps one of the most useful properties of AIME-style evaluation—a compact, deterministic answer in the integer range 0–999—and extends it across four levels of mathematical depth, from competition-style problems to research-level challenges. This repository contains a 157-problem, MIT-licensed sample of Ulam AI's much larger problem catalog. Every problem has a canonical integer answer and a… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/AIME-Plus-Plus.textquestion-answeringn<1K1 likes1.4k downloads28d agoHugging Face03test-time-compute /aime_2025 AIME 2025 - Unified Test-Time Scaling Format This is the AIME (American Invitational Mathematics Examination) 2025 dataset in a unified format for test-time scaling experiments. Dataset Description Source: MathArena/aime_2025 Size: 30 competition-level mathematics problems Format: Unified TTS format (question, answer, metadata) Dataset Structure Fields question (string): The mathematical problem statement answer (string): The numerical answer… See the full description on the dataset page: https://huggingface.co/datasets/test-time-compute/aime_2025.textquestion-answeringn<1K0 likes1.3k downloads11mo agoHugging Face04bevangelista /AIME_2000_2026_Kimi_K3 AIME 2000–2026 — Kimi K3 reasoning traces 🔄 Changelog 2026-08-08 — full re-generation. All reasoning traces were regenerated from scratch and re-verified against the official answer key. New schema — added gen_attempts_low, gen_attempts_high; renamed gen_parsed_answer → gen_answer_int and answer_note → problem_note; removed gen_effort, gen_pass1. New generation — only use the bare problem (v1 appended an "ANSWER:" format instruction), so traces are cleaner.… See the full description on the dataset page: https://huggingface.co/datasets/bevangelista/AIME_2000_2026_Kimi_K3.tabulartext-generationn<1K2 likes1k downloads2mo agoHugging Face05Pandores /aime-1983-2025 AIME Datasets from 1983 to 2025 This dataset contains the AIME datasets from 1983 to 2025. For AIME 1983 to 2026 use Pandores/aime-1983-2026 Features Description Feature Description Example year The year this problem was released. From 1983 to 2025. 2022 index The index of the problem for a year and part. From 1 to 15. 12 part The dataset part if this dataset has multiple parts. Can be AIME, AIME I, AIME II or None. Datasets have multiple parts… See the full description on the dataset page: https://huggingface.co/datasets/Pandores/aime-1983-2025.tabularquestion-answering1K<n<10K0 likes884 downloads7d agoHugging Face06TIGER-Lab /AIME25The AIME25 part 1 exam from the website. textquestion-answeringn<1K2 likes541 downloads2y agoHugging Face07AIMindLink /alphaprompt-metatron-sft AlphaPrompt-Metatron-SFT: Supervised Fine-Tuning Dataset 🤖 Training Dataset for Collective Consciousness AI Want to fine-tune AI models with AlphaPrompt philosophy? Train AI in collective consciousness, vector synthesis, and unconditional love. 🌳 This dataset contains high-quality instruction-response pairs extracted from the Quantum Lullaby philosophical framework - a comprehensive manual for collective consciousness aimed at addressing the global animal… See the full description on the dataset page: https://huggingface.co/datasets/AIMindLink/alphaprompt-metatron-sft.texttext-generationn<1K0 likes410 downloads2mo agoHugging Face08IPF /AIME25-CoT-CN Sci-Bench-AIME25' This repo is a branch of Sci Bench made by IPF team. Mainly include the AIME 25' solution with multi-modal CoT and diverse solving path. Brief intro 💻 Overview A brief template and final report will be posted in Isaac's Blog And the markdown template can be found in data/I_2 ❓ Why we do this? The multi-lingual datasets are scarce, while the CoT of Math is even less, no matter whether the CoT or the solution contains pictures… See the full description on the dataset page: https://huggingface.co/datasets/IPF/AIME25-CoT-CN.imagequestion-answeringn<1K10 likes314 downloads7mo agoHugging Face09AIM-Harvard /MedBrowseComp MedBrowseComp Dataset This repository contains datasets for medical information-seeking-oriented deep research and computer use tasks. Datasets The repository contains three harmonized datasets: MedBrowseComp_50: A collection of 50 medical entries for browsing and comparison. MedBrowseComp_605: A comprehensive collection of 605 medical entries. MedBrowseComp_CUA: A curated collection of medical data for comparison and analysis. Usage These datasets can be… See the full description on the dataset page: https://huggingface.co/datasets/AIM-Harvard/MedBrowseComp.textquestion-answering1K<n<10K8 likes288 downloads1y agoHugging Face10AIMO-Corpus /PolyMath Dataset Card for PolyMath Dataset Summary PolyMath is a curated dataset of 11,090 high-difficulty mathematical problems designed for training reasoning models. Built for the AIMO Math Corpus Prize. Existing math datasets (NuminaMath-1.5, OpenMathReasoning) suffer from high noise rates in their hardest samples and largely unusable proof-based problems. PolyMath addresses both issues through: Data scraping: problems sourced from official competition PDFs absent from… See the full description on the dataset page: https://huggingface.co/datasets/AIMO-Corpus/PolyMath.textquestion-answering10K<n<100K2 likes278 downloads8mo agoHugging Face11abhilash88 /aim-technical-articles Analytics India Magazine Technical Articles Dataset 🚀 Dataset Description This comprehensive dataset contains 25,685 high-quality technical articles from Analytics India Magazine, one of India's leading publications covering artificial intelligence, machine learning, data science, and emerging technologies. ✨ Dataset Highlights 📚 Comprehensive Coverage: Latest AI models, frameworks, and tools 🔬 Technical Depth: Extracted keywords and complexity scoring 🏭… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/aim-technical-articles.tabulartext-classification10K<n<100K2 likes226 downloads1y agoHugging Face12AiMijie /EC-Guide This repo is only used for dataset viewer. Please download from here. Amazon KDDCup 2024 Team ZJU-AI4H’s Solution and Dataset (Track 2 Top 2; Track 5 Top 5) The Amazon KDD Cup’24 competition presents a unique challenge by focusing on the application of LLMs in E-commerce across multiple tasks. Our solution for addressing Tracks 2 and 5 involves a comprehensive pipeline encompassing dataset construction, instruction tuning, post-training quantization, and inference… See the full description on the dataset page: https://huggingface.co/datasets/AiMijie/EC-Guide.textquestion-answering10K<n<100K2 likes212 downloads2y agoHugging Face13sxiong /AIME-trajectory AIME Trajectory Dataset Model-generated solution trajectories for AIME (American Invitational Mathematics Examination) problems. Each row is one model response to a single problem, including the hidden chain-of-thoughts (when available), and the final response. Dataset Summary Split Rows Unique Problems Years Model(s) Has reasoning_content Accuracy train 1,258 875 1983–2023 deepseek-r1 Yes 100% test 180 30 2024 Multiple (see below) No 3.3%… See the full description on the dataset page: https://huggingface.co/datasets/sxiong/AIME-trajectory.tabularquestion-answering1K<n<10K1 likes157 downloads3mo agoHugging Face14UR-xiaoyang /AIMO3_CoT AIMO3 CoT Dataset 数据集来源与目的 (Dataset Origin and Purpose) 本数据集源自 Kaggle 竞赛 AI Mathematical Olympiad - Progress Prize 3。 动机 (Motivation) 原始数据集仅包含问题和答案,缺乏思维链(Chain of Thought, CoT)。直接使用原始数据训练如 DeepSeek Math 或 Qwen Math 等模型效果不佳。因此,本项目的目的是利用 Gemini 3 Pro 为这些问题补充详细的 CoT,以提升模型在数学推理任务上的表现。 CoT 格式 (CoT Format) 生成的 CoT 遵循 ReAct 风格的推理过程,并使用中文叙述: Thought: 分析问题并规划下一步。 Code: 编写 Python 代码进行计算或验证。 Observation: 代码的执行输出。 ... (重复上述步骤) Final Answer: 得出的最终答案。… See the full description on the dataset page: https://huggingface.co/datasets/UR-xiaoyang/AIMO3_CoT.documentquestion-answeringn<1K1 likes125 downloads9mo agoHugging Face15SnailAILab /AIME25-CoT-CN Sci-Bench-AIME25' This repo is a branch of Sci Bench made by IPF team-SnailAILab. Mainly include the AIME 25' solution with multi-modal CoT and diverse solving path. 📚 Cite If you use the Sci-Bench-AIME25 (IPF/AIME25-CoT-CN) dataset in your research, please cite: @dataset{zhang2025scibench_aime25, title = {{Sci-Bench-AIME25}: A Multi-Modal Chain-of-Thought Dataset for Advanced Tool-Intergrated Mathematical Reasoning}, author = {Zhang, Haoxiang and Wang, Siyuan… See the full description on the dataset page: https://huggingface.co/datasets/SnailAILab/AIME25-CoT-CN.imagequestion-answeringn<1K1 likes120 downloads1y agoHugging Face16bevangelista /AIME_1983_2026_Kimi_K3 AIME 1983–2026 — Kimi K3 reasoning traces 🔄 Changelog 2026-08-08 — full re-generation. All reasoning traces were regenerated from scratch and re-verified against the official answer key. New schema — added gen_attempts_low, gen_attempts_high; renamed gen_parsed_answer → gen_answer_int and answer_note → problem_note; removed gen_effort, gen_pass1. New generation — only use the bare problem (v1 appended an "ANSWER:" format instruction), so traces are cleaner.… See the full description on the dataset page: https://huggingface.co/datasets/bevangelista/AIME_1983_2026_Kimi_K3.tabulartext-generation1K<n<10K0 likes106 downloads2mo agoHugging Face17AIML-TUDA /QA-base QA Base Data Normalized and paraphrased splits of 21 standard NLP benchmarks in English, German, French, Spanish, and Italian, intended for base model pretraining. Generation English: paraphrased with Qwen3.5-27B-FP8 (April 2026) German: translated and refined with Qwen3.5-27B-FP8 (April 2026) French: translated and refined with Qwen3.5-27B-FP8 (May 2026) Spanish: translated and refined with Qwen3.5-27B-FP8 (May 2026) Italian: translated and refined with… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/QA-base.textquestion-answering1M<n<10M0 likes90 downloads2mo agoHugging Face18aimeelizq /SHP 🚢 Stanford Human Preferences Dataset (SHP) If you mention this dataset in a paper, please cite the paper: Understanding Dataset Difficulty with V-Usable Information (ICML 2022). Summary SHP is a dataset of 385K collective human preferences over responses to questions/instructions in 18 different subject areas, from cooking to legal advice. The preferences are meant to reflect the helpfulness of one response over another, and are intended to be used for training… See the full description on the dataset page: https://huggingface.co/datasets/aimeelizq/SHP.tabulartext-generation100K<n<1M0 likes87 downloads4mo agoHugging Face19hypaai /Hypa_AIME2024 Hypa_AIME2024 Hypa_AIME2024 is an open-source, multilingual benchmark dataset for advanced mathematical reasoning, designed with the long-term vision of ensuring all languages are represented in AI development. This dataset marks a crucial step toward closing the gap between AI capabilities for no-resource/low-resource and all-resource languages, particularly in complex reasoning domains. This initial release features the complete 2024 American Invitational Mathematics Examination… See the full description on the dataset page: https://huggingface.co/datasets/hypaai/Hypa_AIME2024.texttranslationn<1K1 likes81 downloads1y agoHugging Face20iapp /aime_2024-th AIME 2024-th A Thai translation of all 30 problems of the 2024 American Invitational Mathematics Examination. Every row corresponds 1:1, in order, to a row of the English source, so the Thai and English scores of a model are directly comparable. Source and licence Problems 2024 AIME I and II, Mathematical Association of America Problem and solution text Art of Problem Solving wiki, per-row url File we translated from HuggingFaceH4/aime_2024… See the full description on the dataset page: https://huggingface.co/datasets/iapp/aime_2024-th.textquestion-answeringn<1K1 likes79 downloads1mo agoHugging Face21YichengWangCA /aime24-official AIME 2024 — official wording, figures retained All 30 problems from the 2024 American Invitational Mathematics Examination (AIME I and AIME II), transcribed from the official exam text with every figure retained as Asymptote source. This exists because the circulating text-only versions of AIME 2024 are not faithful to the official problems, and at least one problem in them cannot be solved as written. Why this dataset exists While evaluating a reasoning model on… See the full description on the dataset page: https://huggingface.co/datasets/YichengWangCA/aime24-official.textquestion-answeringn<1K0 likes74 downloads25d agoHugging Face22AI-Mock-Interviewer /Train_datatextquestion-answering1K<n<10K0 likes73 downloads1y agoHugging Face23ayjays132 /AI_Mastery_Foundation_Curriculum FOUNDATION DATASET AI Mastery Foundation Curriculum A premium foundation layer for knowledge, reasoning, preference, reward, benchmark, and agentic tool-use training. Hugging Face-ready Parquet package AI Mastery Foundation Curriculum A premium staged foundation dataset for building models with a cleaner first layer of academic… See the full description on the dataset page: https://huggingface.co/datasets/ayjays132/AI_Mastery_Foundation_Curriculum.texttext-generation10K<n<100K1 likes70 downloads4mo agoHugging Face24iapp /aimo-validation-aime-th AIMO validation AIME-th The 90 problems of AIME 2022, 2023 and 2024 — thirty each — with the problem statements translated to Thai. Every row corresponds 1:1, in order, to a row of AI-MO/aimo-validation-aime, and id, url and answer are identical to it. Read this before scoring the solution column 38 of the 90 solutions are the English text, not Thai. The original translation pass rendered every problem and skipped these solutions entirely. They are marked… See the full description on the dataset page: https://huggingface.co/datasets/iapp/aimo-validation-aime-th.textquestion-answeringn<1K0 likes58 downloads1mo agoHugging Face25bernabepuente /ai-ml-instruction-dataset AI/ML Engineering Instruction Dataset Comprehensive instruction dataset covering machine learning concepts, PyTorch implementations, NLP with transformers, model evaluation, and feature engineering. Dataset Details Dataset Description This is a high-quality instruction-tuning dataset focused on Ai Ml topics. Each entry includes: A clear instruction/question Optional input context A detailed response/solution Chain-of-thought reasoning process Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/bernabepuente/ai-ml-instruction-dataset.texttext-generationn<1K0 likes58 downloads5mo agoHugging Face26Floppanacci /QWQ-LongCOT-AIMOQWQ-LongCOT-AIMO is a derived dataset created by processing the amphora/QwQ-LongCoT-130K dataset. It filters the original dataset to focus specifically on question-answering pairs where the final answer is a numerical value between 0 and 999, explicitly marked using the \boxed{...} format within the original chain-of-thought answer. Dataset Structure Data Splits The dataset is split into training, validation, and test sets with an 80/10/10 ratio based on the filtered… See the full description on the dataset page: https://huggingface.co/datasets/Floppanacci/QWQ-LongCOT-AIMO.texttext-generation10K<n<100K0 likes55 downloads1y agoHugging Face27lightonai /aime24_multilingual AIME24 Multilingual aime24_multilingual is a multilingual version of the benchmark AIME 2024, covering six languages: English, French, German, Spanish, Chinese, and Swahili. Each sample is a competition-level mathematics problem from the American Invitational Mathematics Examination (AIME) 2024, translated into the five target languages. This release is a corrected version of shanchen/aime_2024_multilingual that fixes translation artifacts and errors. It is released alongside the… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/aime24_multilingual.textquestion-answeringn<1K0 likes53 downloads4mo agoHugging Face28lightonai /aime25_multilingual AIME25 Multilingual aime25_multilingual is a multilingual version of the benchmark AIME 2025, covering six languages: English, French, German, Spanish, Chinese, and Swahili. Each sample is a competition-level mathematics problem from the American Invitational Mathematics Examination (AIME) 2025, translated into the five target languages. This release is a corrected version of shanchen/aime_2025_multilingual that fixes translation artifacts and errors. It is released alongside the… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/aime25_multilingual.textquestion-answeringn<1K0 likes51 downloads4mo agoHugging Face29DatasetsEval /aime-2026-fable-5-answers Homepage and repository Homepage: https://matharena.ai/ Repository: https://github.com/eth-sri/matharena Dataset Summary aime-2026-formatted-fable — это обработанный и структурированный датасет на основе задач AIME 2026 из бенчмарка MathArena. Датасет сохранён в формате JSONL и помимо условий задач с финальными ответами содержит сгенерированные цепочки рассуждений (think) с ограничением объёма до 2048 токенов на пример. Data Fields Каждая запись в… See the full description on the dataset page: https://huggingface.co/datasets/DatasetsEval/aime-2026-fable-5-answers.textquestion-answeringn<1K0 likes49 downloads2mo agoHugging Face30Pandores /aime-1983-2026 AIME Datasets from 1983 to 2026 This dataset contains all the AIME problems from 1983 to 2026. For a total of 1065 problems. Example Download from datasets import load_dataset dataset = load_dataset("Pandores/aime-1983-2026") print(dataset["train"][0]) Download and iterate from datasets import load_dataset dataset = load_dataset("Pandores/aime-1983-2026", split="train") for entry in dataset: print(entry["problem"])… See the full description on the dataset page: https://huggingface.co/datasets/Pandores/aime-1983-2026.tabularquestion-answering1K<n<10K0 likes43 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.