datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MathCodeInstruct
MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning
Paper: https://arxiv.org/pdf/2310.03731.pdf
Repo: https://github.com/mathllm/MathCoder
Introduction
We introduce MathCoder, a series of open-source large language models (LLMs) specifically tailored for general math problem-solving.
Base Model: Llama-2
Base Model: Code Llama
MathCoder-L-7B
MathCoder-CL-7B
MathCoder-L-13B
MathCoder-CL-34B
Training Data
The models… See the full description on the dataset page: https://huggingface.co/datasets/MathLLMs/MathCodeInstruct.MathCodeInstruct-Plus
MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning
Paper: https://arxiv.org/pdf/2310.03731.pdf
Repo: https://github.com/mathllm/MathCoder
Introduction
We introduce MathCoder, a series of open-source large language models (LLMs) specifically tailored for general math problem-solving.
Base Model: Llama-2
Base Model: Code Llama
MathCoder-L-7B
MathCoder-CL-7B
MathCoder-L-13B
MathCoder-CL-34B
Training Data
The models… See the full description on the dataset page: https://huggingface.co/datasets/MathLLMs/MathCodeInstruct-Plus.AM-DeepSeek-R1-Filtered-Math-Code
AM DeepSeek R1 Filtered Math and Code
This repository publishes reproducible training subsets derived from
a-m-team/AM-DeepSeek-R1-Distilled-1.4M
at pinned revision 53531c06634904118a2dcd83961918c4d69d1cdf.
The source dataset and this derived release use CC BY-NC 4.0. Commercial use is
not permitted by that license. Preserve attribution and review the upstream
dataset card before use.
Contents
Config / split
Records
Bytes
SHA-256
math / train
111,657
2… See the full description on the dataset page: https://huggingface.co/datasets/chhao/AM-DeepSeek-R1-Filtered-Math-Code.verisci-verified-science-math-code
VeriSci Verified Science Math Code
Verifier-grounded dataset for the Adaption AutoScientist Challenge Part 2, targeting the Science category with secondary Math and Code coverage.
Summary
VeriSci trains models to solve scientific computations, finite-difference PDE updates, numerical ODE steps, unit-checked mechanics, thermodynamics, circuits, chemistry stoichiometry, molarity, unit conversion, vector decomposition, two-point linear modeling, small Python… See the full description on the dataset page: https://huggingface.co/datasets/TheJackBright/verisci-verified-science-math-code.mix60k-math-code-sft
mix60k: 50/50 OpenMathInstruct-2 + OpenCodeInstruct SFT mixture
This dataset is from the paper Register Tokens for Bounded-State Reasoning in Diffusion Language Models.
The exact 60,000-trace SFT mixture used to train the LLaDA-8B-Base and Dream-7B-Base
main triad in the dLLM Registers project.
Composition: ~30K math traces from OpenMathInstruct-2 and ~30K code traces from OpenCodeInstruct.
License: governed by the upstream licenses of OpenMathInstruct-2 and OpenCodeInstruct.… See the full description on the dataset page: https://huggingface.co/datasets/albertge/mix60k-math-code-sft.NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54
NOESIS DORA SFT Dataset
Multilingual supervised fine-tuning dataset built for the NOESIS QwQ+DeepSeek-R1 MoE pipeline.
Released as part of the NOESIS Professional Multilingual Dubbing Automation Platform(framework: DHCF-FNO — Deterministic Hybrid Control Framework for Frozen Neural Operators).
Founder: Ilia Bolotnikov
Organization: AMAImedia.com
X (Twitter): @AMAImediacom
LinkedIn: Ilia Bolotnikov
Telegram: @djbionicl
NOESIS version: v14.8-NT89
Build date: 2026-04… See the full description on the dataset page: https://huggingface.co/datasets/SMH-DEV-AI/NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.curated-openbmb-code-math
Curated OpenBMB Code/Math Post-Training Data
English code/math-focused post-training data derived from curated OpenBMB UltraData rows.
Contents
Config
Rows
Schema
Purpose
sft_no_think
25,891
prompt, response
Direct code/math SFT plus necessary technical instruction-following/alignment
sft_think
6,018
prompt, response
Code/math reasoning SFT with <think>...</think> traces
Total rows: 31,909.
Curation
The SFT split keeps English code… See the full description on the dataset page: https://huggingface.co/datasets/josephmayo/curated-openbmb-code-math.
