CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01agungpambudi /math-dataset-measuring-mathematical-problem-solvingTo cite the dataset please reference it as @article{hendrycksmath2021, title={Measuring Mathematical Problem Solving With the MATH Dataset}, author={Dan Hendrycks and Collin Burns and Saurav Kadavath and Akul Arora and Steven Basart and Eric Tang and Dawn Song and Jacob Steinhardt}, journal={NeurIPS}, year={2021} } textquestion-answering100K<n<1M1 likes9.9k downloads1y agoHugging Face02IDEA-FinAI /Mathematical_Modeling_Speciale_Dataset_v0.1image0 likes3.8k downloads9mo agoHugging Face03mteb /cqadupstack-mathematica CQADupstackMathematicaRetrieval An MTEB dataset Massive Text Embedding Benchmark CQADupStack: A Benchmark Data Set for Community Question-Answering Research Task category t2t Domains Written, Academic, Non-fiction Referencehttp://nlp.cis.unimelb.edu.au/resources/cqadupstack/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["CQADupstackMathematicaRetrieval"])… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cqadupstack-mathematica.texttext-retrieval10K<n<100K1 likes1.1k downloads1y agoHugging Face04BAAI /IndustryCorpus2_mathematics_statistics IndustryCorpus2: Mathematics & Statistics This repository contains the IndustryCorpus2: Mathematics & Statistics domain subset of BAAI/IndustryCorpus2. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryCorpus2: @misc{shi2024industrycorpus2, title = {IndustryCorpus2}, author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao}… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_mathematics_statistics.1 likes599 downloads1mo agoHugging Face05MathematicianNLPer /hamela_books_text_full_oktext1M<n<10M1 likes468 downloads7mo agoHugging Face06timaeus /dsir-pile-13m-filtered-no-github-or-dm_mathematics DSIR Pile 13M - Filtered Version This is a filtered version of timaeus/dsir-pile-13m. Filtering Applied: Excluded: All rows where metadata.pile_set_name contains 'Github' or 'DM_mathematics' Kept: All other rows from the original dataset Dataset Size Original: ~13M examples Filtered: 12,782,200 examples (99.9% of original) Uploaded in: 64 batch files Usage from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/timaeus/dsir-pile-13m-filtered-no-github-or-dm_mathematics.text10M<n<100M0 likes399 downloads1y agoHugging Face07MathematicianNLPer /MoulSot-Full MoulSot-Full Dataset Dataset Summary MoulSot-Full is a large-scale Moroccan Darija speech dataset containing in total 1,500 hours of speech audio. From this extensive corpus, a high-quality subset of approximately 80 hours has been carefully curated and transcribed. It was built entirely from publicly available YouTube content across 51 diverse channels (including vlogs, podcasts, interviews, and commentary) to capture real-world Moroccan Darija, including natural… See the full description on the dataset page: https://huggingface.co/datasets/MathematicianNLPer/MoulSot-Full.audio10K<n<100K0 likes315 downloads5mo agoHugging Face08asandeistefan /romanian-baccalaureate-mathematics Romanian Baccalaureate in Mathematics A curated collection of Romanian Baccalaureate (BAC) mathematics examination papers and answer keys, transcribed from PDF to structured Markdown using Vision-Language Model OCR. Currently the years 2019 - 2025 were added, more will be processed soon. Directory Structure romanian-baccalaureate-mathematics/ ├── metadata.csv # Index of all exam papers ├── pdfs/ # Original PDF files │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/asandeistefan/romanian-baccalaureate-mathematics.n<1K2 likes278 downloads3mo agoHugging Face09Mathematics-Yang /phase_tree_results PHASE-Tree Evaluation Results Full evaluation outputs for the PHASE-Tree paper (Psychology-grounded Hierarchical Attribute-Structured Evolving Tree), covering 8 character-dialogue datasets, 4 experimental paradigms, and 2 evaluation splits (random test + OOD test). Please cite this work if you use these results for analysis, comparison, reproduction, or any other research purpose. 🔗 Resources: 📄 Paper: arXiv:2608.06975 📦 GitHub Repository: MemTensor/PHASE-Tree (code… See the full description on the dataset page: https://huggingface.co/datasets/Mathematics-Yang/phase_tree_results.texttext-generation10K<n<100K1 likes272 downloads1mo agoHugging Face10DataoceanAI /University-level_Mathematics_Physics_Chemistry_Computer_Science_Reasoning_Corpus Title University-level Mathematics, Physics, Chemistry, Computer Science Reasoning Corpus Size 200,000+ text+ multimodal university-level problems, each with step-by-step solutions and final answers Format Natural language explanations with multimodal samples include images (graphs, diagrams, etc.) Subject Mathematics, Physics, Chemistry, Computer Science Labeling Details Question ID/Question Stem (Full text/content) /Subject/Question Type… See the full description on the dataset page: https://huggingface.co/datasets/DataoceanAI/University-level_Mathematics_Physics_Chemistry_Computer_Science_Reasoning_Corpus.0 likes253 downloads1y agoHugging Face11morka17 /sat-mathematics-complete-questions0 likes250 downloads1y agoHugging Face12d0rj /mathematics_dataset Mathematical Reasoning Dataset (English & Russian) A bilingual collection of synthetic school-level mathematics questions and answers, based on the DeepMind mathematics_dataset generator. This dataset contains two language splits: en — the original English data, taken as-is from the official mathematics_dataset-v1.0 release published by Google DeepMind (github.com/google-deepmind/mathematics_dataset). ru — a Russian version generated from scratch with a translated fork of the… See the full description on the dataset page: https://huggingface.co/datasets/d0rj/mathematics_dataset.texttext-generation100M<n<1B3 likes240 downloads2mo agoHugging Face13casey-martin /multilingual-mathematical-autoformalization Multilingual Mathematical Autoformalization "Paper" This repository contains parallel mathematical statements: Input: An informal proof in natural language Output: The corresponding formalization in either Lean or Isabelle This dataset can be used to train models how to formalize mathematical statements into verifiable proofs, a form of machine translation. Abstract Autoformalization is the task of translating natural language materials into machine-verifiable… See the full description on the dataset page: https://huggingface.co/datasets/casey-martin/multilingual-mathematical-autoformalization.texttranslation100K<n<1M5 likes141 downloads3y agoHugging Face14Lots-of-LoRAs /task118_semeval_2019_task10_open_vocabulary_mathematical_answer_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task118_semeval_2019_task10_open_vocabulary_mathematical_answer_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task118_semeval_2019_task10_open_vocabulary_mathematical_answer_generation.texttext-generationn<1K1 likes131 downloads2y agoHugging Face15timaeus /pile-dm_mathematics Dataset Creation Process These subsets were created by streaming over the rows from monology/pile-uncopyrighted and filtering by the meta column. Each subset is generally limited to the first 100,000 qualifying rows encountered. Citations If you use this dataset, please cite the original Pile papers: @article{gao2020pile, title={The Pile: An 800GB dataset of diverse text for language modeling}, author={Gao, Leo and Biderman, Stella and Black, Sid and Golding, Laurence and… See the full description on the dataset page: https://huggingface.co/datasets/timaeus/pile-dm_mathematics.text100K<n<1M1 likes131 downloads1y agoHugging Face16Shalyt /ASyMOB-Algebraic_Symbolic_Mathematical_Operations_Benchmark ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark This dataset is associated with the paper "ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark". Abstract Large language models (LLMs) are increasingly applied to symbolic mathematics, yet existing evaluations often conflate pattern memorization with genuine reasoning. To address this gap, we present ASyMOB, a high-resolution dataset of 35,368 validated symbolic math problems spanning… See the full description on the dataset page: https://huggingface.co/datasets/Shalyt/ASyMOB-Algebraic_Symbolic_Mathematical_Operations_Benchmark.textquestion-answering10K<n<100K3 likes124 downloads4mo agoHugging Face17pj-mathematician /clef2025-bioasq-task13Btext10M<n<100M1 likes110 downloads1y agoHugging Face18VietAlphaLabs /vi-en-mathematics-dictionaryVietAlpha English–Vietnamese Mathematics Dictionary Research page · VietAlpha Lab · Source scan The VietAlpha English–Vietnamese Mathematics Dictionary turns a 709-page printed reference work into a machine-readable bilingual lexicon. It contains 26,205 English and Vietnamese mathematics entries digitized from Cung Kim Tiến's Từ Điển Toán Học Anh – Việt, Việt – Anh and organized as JSON Lines. What is in the dataset Direction Entries English to Vietnamese… See the full description on the dataset page: https://huggingface.co/datasets/VietAlphaLabs/vi-en-mathematics-dictionary.texttranslation10K<n<100K1 likes108 downloads1d agoHugging Face19Lots-of-LoRAs /task706_mmmlu_answer_generation_high_school_mathematics Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task706_mmmlu_answer_generation_high_school_mathematics Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task706_mmmlu_answer_generation_high_school_mathematics.texttext-generationn<1K0 likes102 downloads2y agoHugging Face20verify-ppt /marin-starcoderdata_mathematica0 likes96 downloads6mo agoHugging Face2111-47 /pure_mathematics_25ktext10K<n<100K2 likes94 downloads4mo agoHugging Face22Lots-of-LoRAs /task696_mmmlu_answer_generation_elementary_mathematics Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task696_mmmlu_answer_generation_elementary_mathematics Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task696_mmmlu_answer_generation_elementary_mathematics.texttext-generationn<1K0 likes93 downloads2y agoHugging Face23DataoceanAI /Competition-level_Mathematics_Physics_Reasoning_Corpus Title Competition-level Mathematics, Physics Reasoning Corpus Size 50.000+ text + multimodal competition level problems, each with step-by-step solutions and final answers Format Natural language explanations with multimodal samples include images (graphs, diagrams, etc.) Subject Mathematics/Physics and etc Labeling Details Question ID/Question Stem (Full text/content) /Subject/Question Type (Multiple Choice/Short Answer format… See the full description on the dataset page: https://huggingface.co/datasets/DataoceanAI/Competition-level_Mathematics_Physics_Reasoning_Corpus.0 likes92 downloads1y agoHugging Face24VietAlphaLabs /fr-vi-mathematics-dictionaryVietAlpha French–Vietnamese Mathematics Dictionary Research page · VietAlpha Lab The VietAlpha French–Vietnamese Mathematics Dictionary is a machine-readable edition of Danh-từ Toán-học Pháp-Việt, compiled in Saigon in 1964 by the Mathematics Committee of the National Committee for the Compilation of Specialized Dictionaries. The release contains 4,095 dictionary entries and a 1,369-item Vietnamese index reconstructed from the printed volume. This dataset records how a Vietnamese… See the full description on the dataset page: https://huggingface.co/datasets/VietAlphaLabs/fr-vi-mathematics-dictionary.texttranslation1K<n<10K0 likes88 downloads1d agoHugging Face25Azu /Handwritten-Mathematical-Expression-Convert-LaTeXimage10K<n<100K28 likes84 downloads5y agoHugging Face26Lots-of-LoRAs /task119_semeval_2019_task10_geometric_mathematical_answer_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task119_semeval_2019_task10_geometric_mathematical_answer_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task119_semeval_2019_task10_geometric_mathematical_answer_generation.texttext-generationn<1K0 likes83 downloads2y agoHugging Face27perrabyte /crystal_mathematicsdocumentn<1K0 likes72 downloads1y agoHugging Face28cjc0013 /ouroboros-autonomous-mathematics-process Ouroboros Autonomous Mathematics Process PUBLIC RELEASE This publication-ready package documents an experiment in autonomous mathematics directly applied to improve Ouroboros. The repository is intentionally staged as private so its owner can make it public after review. The PDFs and supporting text are already labeled PUBLIC RELEASE. Included OUROBOROS_AUTONOMOUS_MATHEMATICS_PROCESS_PUBLIC_RELEASE.pdf - the public-facing process white paper.… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/ouroboros-autonomous-mathematics-process.0 likes72 downloads10h agoHugging Face29XinyaoHu /AMPS_mathematicatext1M<n<10M4 likes71 downloads2y agoHugging Face30prithivMLmods /Mathematics-Class10-Tnsb Mathematics-Class10-Tnsb This dataset contains scanned images from a Class 10 Mathematics textbook under the TNSB (Tamil Nadu State Board) curriculum. It is intended for educational machine learning tasks such as image-to-text (OCR), textbook digitization, or educational content understanding. Dataset Details Source: Tamil Nadu State Board Class 10 Mathematics textbook Task: Image-to-Text Language: English Split: train only Rows: 352 Format: Images only (scanned textbook… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Mathematics-Class10-Tnsb.imageimage-to-textn<1K0 likes71 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.