datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
math-code-science-deepseek-r1-en
R1 Dataset Collection
Aggregated high-quality English prompts and model-generated responses from DeepSeek R1 and DeepSeek R1-0528.
Dataset Summary
The R1 Dataset Collection combines multiple public DeepSeek-generated instruction-response corpora into a single, cleaned, English-only JSONL file. Each example consists of a <|user|> prompt and a <|assistant|> response in one "text" field. This release includes:
~21,000 examples from the DeepSeek-R1-0528 Distilled Custom… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/math-code-science-deepseek-r1-en.MathCodeInstruct
MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning
Paper: https://arxiv.org/pdf/2310.03731.pdf
Repo: https://github.com/mathllm/MathCoder
Introduction
We introduce MathCoder, a series of open-source large language models (LLMs) specifically tailored for general math problem-solving.
Base Model: Llama-2
Base Model: Code Llama
MathCoder-L-7B
MathCoder-CL-7B
MathCoder-L-13B
MathCoder-CL-34B
Training Data
The models… See the full description on the dataset page: https://huggingface.co/datasets/MathLLMs/MathCodeInstruct.verified-math-code-17k
Verified Math & Code, 17,000 rows
A math and code instruction dataset where every single row was mechanically checked before it was
allowed in. Not filtered by a heuristic, not scored by a model. Checked.
Two layers of verification, one per domain:
Every math answer was compared against an independent gold answer by exact, numeric and
symbolic (SymPy) comparison. If the worked solution did not arrive at the gold answer, the row
was dropped.… See the full description on the dataset page: https://huggingface.co/datasets/manifesta/verified-math-code-17k.MathCodeInstruct-Plus
MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning
Paper: https://arxiv.org/pdf/2310.03731.pdf
Repo: https://github.com/mathllm/MathCoder
Introduction
We introduce MathCoder, a series of open-source large language models (LLMs) specifically tailored for general math problem-solving.
Base Model: Llama-2
Base Model: Code Llama
MathCoder-L-7B
MathCoder-CL-7B
MathCoder-L-13B
MathCoder-CL-34B
Training Data
The models… See the full description on the dataset page: https://huggingface.co/datasets/MathLLMs/MathCodeInstruct-Plus.verisci-verified-science-math-code
VeriSci Verified Science Math Code
Verifier-grounded dataset for the Adaption AutoScientist Challenge Part 2, targeting the Science category with secondary Math and Code coverage.
Summary
VeriSci trains models to solve scientific computations, finite-difference PDE updates, numerical ODE steps, unit-checked mechanics, thermodynamics, circuits, chemistry stoichiometry, molarity, unit conversion, vector decomposition, two-point linear modeling, small Python… See the full description on the dataset page: https://huggingface.co/datasets/TheJackBright/verisci-verified-science-math-code.curated-openbmb-code-math
Curated OpenBMB Code/Math Post-Training Data
English code/math-focused post-training data derived from curated OpenBMB UltraData rows.
Contents
Config
Rows
Schema
Purpose
sft_no_think
25,891
prompt, response
Direct code/math SFT plus necessary technical instruction-following/alignment
sft_think
6,018
prompt, response
Code/math reasoning SFT with <think>...</think> traces
Total rows: 31,909.
Curation
The SFT split keeps English code… See the full description on the dataset page: https://huggingface.co/datasets/josephmayo/curated-openbmb-code-math.math-code-qa
Math & Code QA — Instruction Dataset
Worked mathematical solutions and short code answers, built for the
Adaption Labs AutoScientist Challenge (Math & Code category).
Rows
5,200
Math
3,600
Code
1,600
Distinct answers
5,199 (100%)
Duplicate questions
none
Nulls
none
Question length
median 27 words
Answer length
median 58 words (max 89)
License
CC-BY-4.0
What makes the math rows unusual
Every math answer is short worked reasoning… See the full description on the dataset page: https://huggingface.co/datasets/flamiinngo/math-code-qa.math-code-qa-v2
Math & Code QA v2 — Instruction Dataset
Worked mathematical solutions and short code answers, spanning arithmetic word
problems through to algebra, geometry and combinatorics.
Built for the Adaption Labs AutoScientist Challenge (Math & Code category).
The model trained on this beats Llama-3.3-70B-Instruct 72 to 28 on the
held-out Math category evaluation.
Rows
5,297 (4,197 math, 1,100 code)
Distinct answers
5,297 (100%)
Duplicate questions
none
Nulls
none… See the full description on the dataset page: https://huggingface.co/datasets/flamiinngo/math-code-qa-v2.math-code-africa
Math & Code — African Context Dataset
Instruction-tuning dataset covering math and coding in African contexts: word problems with African currencies (UGX, KES, NGN), names, geography, and real-world scenarios (mobile money, market trading, farming); coding challenges for USSD systems, mobile money APIs, SMS gateways, agricultural data pipelines, and multilingual NLP — grounded via web search, generated with gemini-2.5-flash.
Dataset Details
Rows: 331
Regions… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/math-code-africa.
