CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01microsoft /orca-math-word-problems-200k Dataset Card This dataset contains ~200K grade school math word problems. All the answers in this dataset is generated using Azure GPT4-Turbo. Please refer to Orca-Math: Unlocking the potential of SLMs in Grade School Math for details about the dataset construction. Dataset Sources Repository: microsoft/orca-math-word-problems-200k Paper: Orca-Math: Unlocking the potential of SLMs in Grade School Math Direct Use This dataset has been designed to… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/orca-math-word-problems-200k.textquestion-answering100K<n<1M498 likes23k downloads3y agoHugging Face02Graphite-AI /Graphite_Past_Problems1 likes21k downloads1y agoHugging Face03DabbyOWL /PDE_Inverse_Problem_Benchmarking PDEInvBench: A Comprehensive Dataset and Design Space Exploration of Neural Networks for PDE Inverse Problems This is the official dataset for the paper PDEInvBench: A Comprehensive Dataset and Design Space Exploration of Neural Networks for PDE Inverse Problems. Code: GitHub - ASK-Berkeley/PDEInvBench Sample Usage You can use the provided script from the codebase to batch download the data: pip install huggingface_hub python3 huggingface_pdeinv_download.py --dataset… See the full description on the dataset page: https://huggingface.co/datasets/DabbyOWL/PDE_Inverse_Problem_Benchmarking.imageother100M<n<1B3 likes15k downloads4mo agoHugging Face04garrethlee /comprehensive-arithmetic-problemstext1M<n<10M0 likes12k downloads4mo agoHugging Face05SAIRfoundation /equational-theories-selected-problems Equational Theories Selected Problems Update (September 11, 2026) This dataset was updated on September 11, 2026. Main changes: released the official Stage 2 evaluation problems: stage2_evaluation_main (200 problems; ground truth withheld — answer is null until Stage 2 concludes) and stage2_evaluation_research (100 order-5 research problems with no ground truth) added metadata/stage2_evaluation_main.json and metadata/stage2_evaluation_research.json… See the full description on the dataset page: https://huggingface.co/datasets/SAIRfoundation/equational-theories-selected-problems.tabular1K<n<10K11 likes10k downloads10d agoHugging Face06agungpambudi /math-dataset-measuring-mathematical-problem-solvingTo cite the dataset please reference it as @article{hendrycksmath2021, title={Measuring Mathematical Problem Solving With the MATH Dataset}, author={Dan Hendrycks and Collin Burns and Saurav Kadavath and Akul Arora and Steven Basart and Eric Tang and Dawn Song and Jacob Steinhardt}, journal={NeurIPS}, year={2021} } textquestion-answering100K<n<1M1 likes9.9k downloads1y agoHugging Face07garrethlee /comprehensive-arithmetic-problems-carriestext1M<n<10M0 likes7.9k downloads2y agoHugging Face08kaysss /leetcode-problem-solutions LeetCode Solution Dataset This dataset contains community-contributed LeetCode solutions scraped from public discussions and solution pages, enriched with metadata such as vote counts, author info, tags, and full code content. The goal is to make high-quality, peer-reviewed coding solutions programmatically accessible for research, analysis, educational use, or developer tooling. Column Descriptions Column Name Type Description question_slug string The unique… See the full description on the dataset page: https://huggingface.co/datasets/kaysss/leetcode-problem-solutions.tabulartext-classification100K<n<1M9 likes5.3k downloads1y agoHugging Face09ChristophSchuhmann /basic-math-problems-with-step-by-step-solutionstext10M<n<100M9 likes4.6k downloads3y agoHugging Face10Navvye /Olympiad-Problems-CSV0 likes4.2k downloads2y agoHugging Face11Prompt48 /AIME_Problem_Set_1983-2024tabularn<1K0 likes4k downloads2y agoHugging Face12sandropa /aops-problemstext10K<n<100K0 likes3.8k downloads5mo agoHugging Face13PrimeIntellect /verifiable-coding-problems SYNTHETIC-1 This is a subset of the task data used to construct SYNTHETIC-1. You can find the full collection here text100K<n<1M44 likes3.4k downloads2y agoHugging Face14open-r1 /verifiable-coding-problems-python Dataset Card for Verifiable Coding Problems Python 10k This dataset contains all Python problems from PrimeIntellect's verifiable-coding-problems dataset. We have formatted the verification_info and metadata columns to be proper dictionaries, but otherwise the data is the same. Please see their dataset for more details. text10K<n<100K12 likes2.9k downloads2y agoHugging Face15felixZzz /numina_162k_amc_aime_problemstext1K<n<10K0 likes1.7k downloads1y agoHugging Face16PrimeIntellect /real-world-swe-problems SYNTHETIC-1 This is a subset of the task data used to construct SYNTHETIC-1. You can find the full collection here text10K<n<100K13 likes1.5k downloads2y agoHugging Face17rakeshb4r /M2-AOPS-Unique-Problems Unique Math Problems (AOPS subset) This dataset contains 81,901 unique problem statements extracted from the AOPS subset of rakeshb4r/Nemotron-Math-v2. Dataset Structure problem_statement (string): The text of the math problem. Source Original source: Nemotron-Math-v2 texttext-generation10K<n<100K0 likes945 downloads8mo agoHugging Face18sigcp /hardtests_problems Dataset Card for HARDTESTS Problems HARDTESTS is a competitive programming dataset containing 47,136 problems collected from 13 different Online Judges (OJs). Each problem includes a problem statement, numerous oracle code solutions, and a set of relatively reliable test cases. Note: Due to their large size, the test cases are stored in a separate dataset. This dataset is presented in the paper HardTests: Synthesizing High-Quality Test Cases for LLM Coding. Project Page… See the full description on the dataset page: https://huggingface.co/datasets/sigcp/hardtests_problems.texttext-generation10K<n<100K13 likes899 downloads1y agoHugging Face19DenCT /codeforces-problems-7ktabulartext-generation1K<n<10K6 likes898 downloads2y agoHugging Face20kaysss /leetcode-problem-set LeetCode Scraper Dataset This dataset contains information scraped from LeetCode. It is designed to assist developers in analyzing LeetCode problems, generating insights, and building tools for competitive programming or educational purposes. Dataset Contents The dataset includes the following files: problem_set.csv Contains a list of LeetCode problems with metadata such as difficulty, acceptance rate, tags, and more. Columns: acRate: Acceptance rate of the… See the full description on the dataset page: https://huggingface.co/datasets/kaysss/leetcode-problem-set.tabularquestion-answering1K<n<10K9 likes895 downloads1y agoHugging Face21JWei05 /SWE-smith-java-6704-filtered-for-problem-statementstext1K<n<10K0 likes875 downloads7mo agoHugging Face22marin-community /ar5iv-no-problem-markdown Marin Markdownified Ar5iv Markdownified Ar5iv transforms academic papers from arXiv into clean, structured Markdown format consisting of 2.74B tokens across two splits. This dataset preserves th content while making it accessible for language model training on academic text. Value Tokens 2 742 463 924 Primary source https://sigmathling.kwarc.info/resources/ar5iv-dataset-2024/ File format JSONL License C-UDA-1.0 (mirrors upstream Ar5iv licenses)… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/ar5iv-no-problem-markdown.texttext-generation100K<n<1M5 likes803 downloads1y agoHugging Face23JWei05 /SWE-smith-py-39471-filtered-for-problem-statementstext10K<n<100K0 likes788 downloads7mo agoHugging Face24FreedomIntelligence /medical-o1-verifiable-problem Introduction This dataset features open-ended medical problems designed to improve LLMs' medical reasoning. Each entry includes a open-ended question and a ground-truth answer based on challenging medical exams. The verifiable answers enable checking LLM outputs, refining their reasoning processes. For details, see our paper and GitHub repository. Citation If you find our data useful, please consider citing our work! @misc{chen2024huatuogpto1medicalcomplexreasoning… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-verifiable-problem.textquestion-answering10K<n<100K124 likes780 downloads2y agoHugging Face25garrethlee /simple-arithmetic-problemstext100K<n<1M2 likes767 downloads2y agoHugging Face26open-r1 /verifiable-coding-problems-python_decontaminated-testedtext10K<n<100K0 likes729 downloads2y agoHugging Face27kaysss /leetcode-problem-detailed LeetCode Scraper Dataset This dataset contains information scraped from LeetCode, including problem details, metadata, and related files. It is designed to assist developers in analyzing LeetCode problems, generating insights, and building tools for competitive programming or educational purposes. questions_deets.csv Contains detailed information about each problem, including problem descriptions, constraints, and examples. Columns: questionFrontendId: Unique problem ID.… See the full description on the dataset page: https://huggingface.co/datasets/kaysss/leetcode-problem-detailed.tabulartext-classification1K<n<10K10 likes669 downloads1y agoHugging Face28jablonkagroup /us-olympiad-problemstext1K<n<10K0 likes651 downloads1y agoHugging Face29AgentNativeResearchLab /fm-open-problems-kimi-k3-trajectories FrontierMath Open Problems trajectories — kimi-k3 Fixed-budget agent trajectories on FrontierMath: Open Problems (Epoch AI's collection of 50 genuinely unsolved research mathematics problems). Nothing is graded. Epoch's verifiers are not public; the harness verifier is a checkpoint stub that always writes reward 0 so the continue-until-timeout harness keeps re-prompting the agent until the fixed wall-clock budget (105 min/task, override_timeout_sec: 6300) elapses.… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/fm-open-problems-kimi-k3-trajectories.0 likes576 downloads1mo agoHugging Face30mihailgribov /olympiad_style_integer_math_problems Olympiad Math Corpus Version: v2.1.1 Release date: 2026-05-03 59,486 synthetically generated olympiad-style math problems with verified integer answers and formal computation graphs. Loading from datasets import load_dataset ds = load_dataset("mihailgribov/olympiad_style_integer_math_problems", split="train") lemma_applicability is stored as list[{lemma, status}] rather than a sparse dict (required for Arrow-based consumers). To convert to a dict for local use:… See the full description on the dataset page: https://huggingface.co/datasets/mihailgribov/olympiad_style_integer_math_problems.documenttext-generation10K<n<100K1 likes564 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.