datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dsir-pile-13m-filtered-no-github-or-dm_mathematics
DSIR Pile 13M - Filtered Version
This is a filtered version of timaeus/dsir-pile-13m.
Filtering Applied:
Excluded: All rows where metadata.pile_set_name contains 'Github' or 'DM_mathematics'
Kept: All other rows from the original dataset
Dataset Size
Original: ~13M examples
Filtered: 12,782,200 examples (99.9% of original)
Uploaded in: 64 batch files
Usage
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/timaeus/dsir-pile-13m-filtered-no-github-or-dm_mathematics.phase_tree_results
PHASE-Tree Evaluation Results
Full evaluation outputs for the PHASE-Tree paper
(Psychology-grounded Hierarchical Attribute-Structured Evolving Tree),
covering 8 character-dialogue datasets, 4 experimental paradigms, and
2 evaluation splits (random test + OOD test).
Please cite this work if you use these results for analysis, comparison, reproduction, or any other research purpose.
🔗 Resources:
📄 Paper: arXiv:2608.06975
📦 GitHub Repository: MemTensor/PHASE-Tree (code… See the full description on the dataset page: https://huggingface.co/datasets/Mathematics-Yang/phase_tree_results.mathematics_dataset
Mathematical Reasoning Dataset (English & Russian)
A bilingual collection of synthetic school-level mathematics questions and answers, based on the DeepMind mathematics_dataset generator.
This dataset contains two language splits:
en — the original English data, taken as-is from the official mathematics_dataset-v1.0 release published by Google DeepMind (github.com/google-deepmind/mathematics_dataset).
ru — a Russian version generated from scratch with a translated fork of the… See the full description on the dataset page: https://huggingface.co/datasets/d0rj/mathematics_dataset.pile-dm_mathematics
Dataset Creation Process
These subsets were created by streaming over the rows from monology/pile-uncopyrighted and filtering by the meta column. Each subset is generally limited to the first 100,000 qualifying rows encountered.
Citations
If you use this dataset, please cite the original Pile papers:
@article{gao2020pile,
title={The Pile: An 800GB dataset of diverse text for language modeling},
author={Gao, Leo and Biderman, Stella and Black, Sid and Golding, Laurence and… See the full description on the dataset page: https://huggingface.co/datasets/timaeus/pile-dm_mathematics.vi-en-mathematics-dictionaryVietAlpha English–Vietnamese Mathematics Dictionary
Research page ·
VietAlpha Lab ·
Source scan
The VietAlpha English–Vietnamese Mathematics Dictionary turns a 709-page printed reference work into a machine-readable bilingual lexicon. It contains 26,205 English and Vietnamese mathematics entries digitized from Cung Kim Tiến's Từ Điển Toán Học Anh – Việt, Việt – Anh and organized as JSON Lines.
What is in the dataset
Direction
Entries
English to Vietnamese… See the full description on the dataset page: https://huggingface.co/datasets/VietAlphaLabs/vi-en-mathematics-dictionary.task706_mmmlu_answer_generation_high_school_mathematics
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task706_mmmlu_answer_generation_high_school_mathematics
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task706_mmmlu_answer_generation_high_school_mathematics.task696_mmmlu_answer_generation_elementary_mathematics
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task696_mmmlu_answer_generation_elementary_mathematics
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task696_mmmlu_answer_generation_elementary_mathematics.pure_mathematics_25kfr-vi-mathematics-dictionaryVietAlpha French–Vietnamese Mathematics Dictionary
Research page ·
VietAlpha Lab
The VietAlpha French–Vietnamese Mathematics Dictionary is a machine-readable edition of Danh-từ Toán-học Pháp-Việt, compiled in Saigon in 1964 by the Mathematics Committee of the National Committee for the Compilation of Specialized Dictionaries. The release contains 4,095 dictionary entries and a 1,369-item Vietnamese index reconstructed from the printed volume.
This dataset records how a Vietnamese… See the full description on the dataset page: https://huggingface.co/datasets/VietAlphaLabs/fr-vi-mathematics-dictionary.crystal_mathematicstask689_mmmlu_answer_generation_college_mathematics
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task689_mmmlu_answer_generation_college_mathematics
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task689_mmmlu_answer_generation_college_mathematics.TheArabicPile_Mathematics
The Arabic Pile
Introduction:
The Arabic Pile is a comprehensive dataset meticulously designed to parallel the structure of The Pile and The Nordic Pile. Focused on the Arabic language, the dataset encompasses a vast array of linguistic nuances, incorporating both Modern Standard Arabic (MSA) and various Levantine, North African, and Egyptian dialects. Tailored for the training and fine-tuning of large language models, the dataset consists of 13 subsets, each uniquely… See the full description on the dataset page: https://huggingface.co/datasets/premio-ai/TheArabicPile_Mathematics.LLM-Hallucination-Detection-complex-mathematics
AIME Hallucination Detection Dataset
This dataset is created for detecting hallucinations in Large Language Models (LLMs), particularly focusing on complex mathematical problems. It can be used for tasks like model evaluation, fine-tuning, and research.
Dataset Details
Name: AIME Hallucination Detection Dataset
Format: CSV
Size: (add size, e.g., 10MB)
Files Included:
AIME-hallucination-detection-dataset.csv: Contains the dataset.
Content Description… See the full description on the dataset page: https://huggingface.co/datasets/tourist800/LLM-Hallucination-Detection-complex-mathematics.dataset-CoT-Applied-Mathematics-824k12-mathematics-standards-expanded
K-12 Mathematics Standards, expanded (generated instruction data)
4,965 instruction/input/output records for mathematics, generated around a K-12
standards taxonomy for instruction-tuning and educational-content experiments.
How this was built (read this first)
These are programmatically generated training examples, not curriculum written by
educators and not the text of any official standard. A generator combined standards
metadata - codes, grade levels, domains… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-mathematics-standards-expanded.DeepMind_Mathematics_QAmmlu-high_school_mathematics-neg-prepend
Dataset Card for "mmlu-high_school_mathematics-neg-prepend"
More Information needed
professor-mathematicsPhysics-Chemistry-Mathematics-Pro
Physics-Chemistry-Mathematics-Pro
Copyright © 2026 Vinay Umrethe umrethevinay@gmail.com.
This dataset is licensed under the Creative Commons Attribution 4.0 International License.
principia-mathematicszh-en-1925-mathematics-dictionaryVietAlpha Chinese–English Mathematics Dictionary
數學辭典, 1925
VietAlpha Lab ·
Website
The VietAlpha Chinese–English Mathematics Dictionary is a machine-readable, verified digital edition of 數學辭典 (Mathematical Dictionary), compiled by 倪德基 (Ni Deji) and others and published by 中華書局 (Zhonghua Book Company) in 1925 (民國14年).
Each headword is a Chinese mathematical term printed in 【…】 brackets, followed by its English equivalent, a subject marker, and a definition in Literary Chinese that… See the full description on the dataset page: https://huggingface.co/datasets/VietAlphaLabs/zh-en-1925-mathematics-dictionary.k12-mathematics-standards-aligned
[!WARNING]
Deprecated - use k12-mathematics-standards-expanded instead.
This dataset is superseded: every input in this set also appears there, plus 366 more and two additional metadata columns. Nothing here is unique to it.
It stays online so existing references keep resolving, but it will not be updated.
New work should point at robworks-software/k12-mathematics-standards-expanded.
K-12 Mathematics Standards (generated instruction data)
4,397 instruction/input/output records… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-mathematics-standards-aligned.IndustryCorpus_mathematics[中文主页]
Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise.
To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_mathematics.mmlu-elementary_mathematics
Dataset Card for "mmlu-elementary_mathematics"
More Information needed
Mathematics_Test_Runmathematics_datasetdsir-pile-1m-filtered-no-github-or-dm_mathematics
My_Downsampled_Dataset
This dataset contains 1,000,000 examples from timaeus/dsir-pile-13m-filtered-no-github-or-dm_mathematics, downsampled for efficient processing.
Usage
from datasets import load_dataset
dataset = load_dataset("path/to/my_downsampled_dataset")
mathematics_competition
Mathematics Competition Evaluation
Competition-level mathematics evaluation dataset with 3-run predictions from Gemini model.
Dataset Structure
Each row contains:
uuid: unique identifier
question: math competition problem
answer: ground truth answer
source: problem source
run_0, run_1, run_2: each a dict with:
prediction: model's answer
stream_output: list of stream output segments
stream_output_kinds: list of output kinds (thought/text/tool_call)
correct: whether… See the full description on the dataset page: https://huggingface.co/datasets/LinhIcey/mathematics_competition.greek_lyceum_mathematics
Dataset Card for Greek Lyceum Mathematics
The Greek Lyceum Mathematics dataset is a set of 465 exercises and answers in Greek extracted from the Item Bank at https://trapeza.iep.edu.gr/.
Dataset Details
Dataset Description
Curated by: ILSP/Athena RC
Language(s) (NLP): el
License: cc-by-nc-sa-4.0
Bias, Risks, and Limitations
This dataset is the result of automatic… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/greek_lyceum_mathematics.Mathematics_CoT_1000The file contains only 960 questions, Chain of Thoughts (CoT), and solutions
This is done to protect the main 10K data batch.
For the purchase of a 10K batch, please contact the publisher
Publisher : Siddharth Jadhav
E-mail id: 5a.siddharthjadhav@gmail.com
