datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
math-code-science-deepseek-r1-en
R1 Dataset Collection
Aggregated high-quality English prompts and model-generated responses from DeepSeek R1 and DeepSeek R1-0528.
Dataset Summary
The R1 Dataset Collection combines multiple public DeepSeek-generated instruction-response corpora into a single, cleaned, English-only JSONL file. Each example consists of a <|user|> prompt and a <|assistant|> response in one "text" field. This release includes:
~21,000 examples from the DeepSeek-R1-0528 Distilled Custom… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/math-code-science-deepseek-r1-en.MathCode-Pile
MathCode-Pile
MathCode-Pile is a dataset for continue pretraining large language models to enhance their mathematical reasoning abilities. It is introduced in the paper MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code. It contains 19.2B tokens, with math-related data covering web pages, textbooks, model-synthesized text, and math related code. Currently, filtered-OpenWebMath, filtered-CC-En-math, and translated mathematical code are… See the full description on the dataset page: https://huggingface.co/datasets/MathGenie/MathCode-Pile.NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54
⚡ Each donation funds the next large quant.
I host free GGUF or MoE quants as independent research.
Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro.
Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant.
🎉 Boosty🦄 |
☕ Buy Me a Coffee🦄 |
⭐ DonationAlerts🦄
💚 Thanks to Hugging Face for extra storage.🦄… See the full description on the dataset page: https://huggingface.co/datasets/AMAImedia/NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.NOESIS-50K-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54
⚡ Each donation funds the next large quant.
I host free GGUF or MoE quants as independent research.
Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro.
Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant.
🎉 Boosty🦄 |
☕ Buy Me a Coffee🦄 |
⭐ DonationAlerts🦄
💚 Thanks to Hugging Face for extra storage.🦄… See the full description on the dataset page: https://huggingface.co/datasets/AMAImedia/NOESIS-50K-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.MathCodeInstruct
MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning
Paper: https://arxiv.org/pdf/2310.03731.pdf
Repo: https://github.com/mathllm/MathCoder
Introduction
We introduce MathCoder, a series of open-source large language models (LLMs) specifically tailored for general math problem-solving.
Base Model: Llama-2
Base Model: Code Llama
MathCoder-L-7B
MathCoder-CL-7B
MathCoder-L-13B
MathCoder-CL-34B
Training Data
The models… See the full description on the dataset page: https://huggingface.co/datasets/MathLLMs/MathCodeInstruct.MathCodeInstruct-Plus
MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning
Paper: https://arxiv.org/pdf/2310.03731.pdf
Repo: https://github.com/mathllm/MathCoder
Introduction
We introduce MathCoder, a series of open-source large language models (LLMs) specifically tailored for general math problem-solving.
Base Model: Llama-2
Base Model: Code Llama
MathCoder-L-7B
MathCoder-CL-7B
MathCoder-L-13B
MathCoder-CL-34B
Training Data
The models… See the full description on the dataset page: https://huggingface.co/datasets/MathLLMs/MathCodeInstruct-Plus.AM-DeepSeek-R1-Filtered-Math-Code
AM DeepSeek R1 Filtered Math and Code
This repository publishes reproducible training subsets derived from
a-m-team/AM-DeepSeek-R1-Distilled-1.4M
at pinned revision 53531c06634904118a2dcd83961918c4d69d1cdf.
The source dataset and this derived release use CC BY-NC 4.0. Commercial use is
not permitted by that license. Preserve attribution and review the upstream
dataset card before use.
Contents
Config / split
Records
Bytes
SHA-256
math / train
111,657
2… See the full description on the dataset page: https://huggingface.co/datasets/chhao/AM-DeepSeek-R1-Filtered-Math-Code.mix60k-math-code-sft
mix60k: 50/50 OpenMathInstruct-2 + OpenCodeInstruct SFT mixture
This dataset is from the paper Register Tokens for Bounded-State Reasoning in Diffusion Language Models.
The exact 60,000-trace SFT mixture used to train the LLaDA-8B-Base and Dream-7B-Base
main triad in the dLLM Registers project.
Composition: ~30K math traces from OpenMathInstruct-2 and ~30K code traces from OpenCodeInstruct.
License: governed by the upstream licenses of OpenMathInstruct-2 and OpenCodeInstruct.… See the full description on the dataset page: https://huggingface.co/datasets/albertge/mix60k-math-code-sft.verisci-verified-science-math-code
VeriSci Verified Science Math Code
Verifier-grounded dataset for the Adaption AutoScientist Challenge Part 2, targeting the Science category with secondary Math and Code coverage.
Summary
VeriSci trains models to solve scientific computations, finite-difference PDE updates, numerical ODE steps, unit-checked mechanics, thermodynamics, circuits, chemistry stoichiometry, molarity, unit conversion, vector decomposition, two-point linear modeling, small Python… See the full description on the dataset page: https://huggingface.co/datasets/TheJackBright/verisci-verified-science-math-code.MathCode-Pile
MathCode-Pile
MathCode-Pile is a dataset for continue pretraining large language models to enhance their mathematical reasoning abilities. It is introduced in the paper MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code. It contains 19.2B tokens, with math-related data covering web pages, textbooks, model-synthesized text, and math related code. Currently, filtered-OpenWebMath, filtered-CC-En-math, and translated mathematical code… See the full description on the dataset page: https://huggingface.co/datasets/NP235/MathCode-Pile.Josephgflowers__TinyLlama_v1.1_math_code-world-test-1-details
Dataset Card for Evaluation run of Josephgflowers/TinyLlama_v1.1_math_code-world-test-1
Dataset automatically created during the evaluation run of model Josephgflowers/TinyLlama_v1.1_math_code-world-test-1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Josephgflowers__TinyLlama_v1.1_math_code-world-test-1-details.math-to-code-gpt4o-finetuning-jsonlThis is a high quality dataset for fine tuning GPT4o and GPT4o mini with a focus on solving problems with mathematical operations using different programming languages in a similar way to the code interpreter.
Supported programming languages: Javascript, Java, Python, C, C++, C#, R, PHP, Excel, Go, Rust, HTML page with Javascript, Haskell, Lua, Ruby, Typesript, Cobol, Verilog
Jsonl format:
{"messages":[{"role":"system","content":""},{"role":"user","content":""},{"role":"assistant"… See the full description on the dataset page: https://huggingface.co/datasets/sinatra-rd/math-to-code-gpt4o-finetuning-jsonl.NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54
NOESIS DORA SFT Dataset
Multilingual supervised fine-tuning dataset built for the NOESIS QwQ+DeepSeek-R1 MoE pipeline.
Released as part of the NOESIS Professional Multilingual Dubbing Automation Platform(framework: DHCF-FNO — Deterministic Hybrid Control Framework for Frozen Neural Operators).
Founder: Ilia Bolotnikov
Organization: AMAImedia.com
X (Twitter): @AMAImediacom
LinkedIn: Ilia Bolotnikov
Telegram: @djbionicl
NOESIS version: v14.8-NT89
Build date: 2026-04… See the full description on the dataset page: https://huggingface.co/datasets/SMH-DEV-AI/NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.DeepMount00__Qwen2.5-7B-Instruct-MathCoder-details
Dataset Card for Evaluation run of DeepMount00/Qwen2.5-7B-Instruct-MathCoder
Dataset automatically created during the evaluation run of model DeepMount00/Qwen2.5-7B-Instruct-MathCoder
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DeepMount00__Qwen2.5-7B-Instruct-MathCoder-details.EpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math-details
Dataset Card for Evaluation run of EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math
Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math-details.curated-openbmb-code-math
Curated OpenBMB Code/Math Post-Training Data
English code/math-focused post-training data derived from curated OpenBMB UltraData rows.
Contents
Config
Rows
Schema
Purpose
sft_no_think
25,891
prompt, response
Direct code/math SFT plus necessary technical instruction-following/alignment
sft_think
6,018
prompt, response
Code/math reasoning SFT with <think>...</think> traces
Total rows: 31,909.
Curation
The SFT split keeps English code… See the full description on the dataset page: https://huggingface.co/datasets/josephmayo/curated-openbmb-code-math.mathcoder-embeddingsadaption-africa-math-code-qa
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-africa_math_code_qa
This instruction-tuning dataset contains 331 question-answer pairs focused on mathematical word problems and Python coding tasks within African contexts. The content covers real-world scenarios such as currency conversion, budgeting, geospatial analysis, and web scraping for government data across countries like Uganda, Kenya, and South Africa.… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/adaption-africa-math-code-qa.autoscientist-mathcode-datasetASPRM-MATHCODE-DeepSeek-Training-Datasetllama-math-code-ftASPRM-MATHCODE-Mistral-Training-Datasetmathcodepile_released_mathematical_code_v3
