datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NuminaMath-LEAN-Sol
NuminaMath-LEAN Cleaned with NL Solutions
Dataset Summary
This is a cleaned version of the NuminaMath-LEAN dataset, enhanced with natural language (NL) solutions matched from source datasets. The primary goal is to provide paired formal statements/proofs with natural language solutions for proof formalization and theorem proving research.
The dataset matches problems from NuminaMath-LEAN with their corresponding natural language solutions from:
olympiads-ref: A… See the full description on the dataset page: https://huggingface.co/datasets/iiis-lean/NuminaMath-LEAN-Sol.2026-08-20-odcv-feature-discovery-difficult-advice-716-5-pct-vs-numina-control
LLM-driven feature discovery over ODCV-Bench rollouts from TWO matched Qwen3.6-27B LoRA arms — 9,284 filtered instruction rows plus 716 rows that differ only in kind (constitution-grounded difficult advice vs NuminaMath chain-of-thought) — asking which reasoning and action properties separate the two models, and which go with the judged misalignment.
field
value
experiment
LLM-driven feature discovery over ODCV-Bench rollouts from TWO matched Qwen3.6-27B LoRA arms — 9… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-20-odcv-feature-discovery-difficult-advice-716-5-pct-vs-numina-control.2026-08-04-qwen36-27b-1000ex-difficult-advice-350-numina-650-train-mixture
Qwen3.6-27B training bundle — 2026-08-04-qwen36-27b-1000ex-da350-numina650-train
code.tar.gz (trainer, src/, configs/) plus mixture.jsonl. The pod untars it, copies
the jsonl to data/mixture.jsonl, and runs configs/train_1000ex_da350_numina650.yaml.
field
value
experiment
1000-example mixture: 350 difficult-advice (all t1-t3 + t4 fill) + 650 NuminaMath-CoT; lr 4e-5, 1 epoch
date_generated
2026-08-03
constitution
constitutions/claude_constitution_principles.md —… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-04-qwen36-27b-1000ex-difficult-advice-350-numina-650-train-mixture.2026-08-02-qwen36-mixture-500k-numina-heavy-empty-think
Qwen3.6-27B SFT mixture — 500k maths-weighted, empty-think markers
499,595 tokens across 1,001 conversations, weighted toward maths, with Qwen3.6's empty
think marker on the non-maths rows. md5 c433f31eba2b5b4919fb166043caccb5.
Source
Examples
Tokens
Share
Marker
NuminaMath-CoT
611
333,351
66.9%
no
No Robots
271
82,239
16.5%
yes
TULU3
119
82,445
16.5%
yes
Total
1,001
499,595
390 marked
Derived from
qwen3.6-27b-mixture-500k-numina-heavy
by adding the… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-02-qwen36-mixture-500k-numina-heavy-empty-think.lm-eval-results-penfever-Llama-3-8B-NuminaCoT-private
Dataset Card for Evaluation run of penfever/Llama-3-8B-NuminaCoT
Dataset automatically created during the evaluation run of model penfever/Llama-3-8B-NuminaCoT
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-penfever-Llama-3-8B-NuminaCoT-private.2026-08-03-qwen36-27b-1000ex-difficult-advice-250-numina-750-train-mixture
Qwen3.6-27B training bundle — 1,000 examples (250 synthdoc_v2 t1-t3 + 750 NuminaMath-CoT)
RunPod training bundle: code.tar.gz (the trainer, src/, configs/) plus mixture.jsonl.
The pod pulls this, untars it, copies the jsonl to data/, and runs configs/train_1000ex_da250_numina750.yaml.
field
value
experiment
1-epoch assistant-only-loss LoRA SFT of Qwen3.6-27B on a 1,000-example mixture that is 25% difficult-advice by example count
date_generated
2026-08-03… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-03-qwen36-27b-1000ex-difficult-advice-250-numina-750-train-mixture.2026-08-03-qwen36-27b-arma-1000ex-numina-666-tulu-334-train-mixture
Qwen3.6-27B training bundle — 2026-08-03-qwen36-27b-armA-1000ex-numina666-tulu334-train
code.tar.gz (trainer, src/, configs/) plus mixture.jsonl. The pod untars it, copies
the jsonl to data/mixture.jsonl, and runs configs/train_armA_1000ex_numina666_tulu334.yaml.
field
value
experiment
Arm A: 1,000 examples, no difficult-advice - 666 NuminaMath-CoT + 167 TULU3 + 167 No Robots
date_generated
2026-08-03
constitution
constitutions/claude_constitution_principles.md —… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-03-qwen36-27b-arma-1000ex-numina-666-tulu-334-train-mixture.euclean-numina-geometry
Euclean Numina-Geometry
Euclean Numina-Geometry is a generated Lean 4 / Mathlib geometry
formalization dataset released with the ICML 2026 paper:
Euclean: Automated Geometry Problem Formalization with Unified Verification in Lean
GitHub repository: https://github.com/tlb-22/Euclean
Dataset Description
This release contains 183,796 Numina-derived geometry problems with generated
Lean theorem statements. The formalizations were regenerated with Codex GPT-5.4
using… See the full description on the dataset page: https://huggingface.co/datasets/tlb-22/euclean-numina-geometry.2000-Reasoning-AIME-Numinamathqwen_math_numina_80k_add_critique_0119NuminaMath-1.5-Pro
NuminaMath-1.5-Pro
Dataset Overview
NuminaMath-1.5-Pro targets post-training and verifiable reasoning scenarios. It applies strict filtering, judge-based consistency checks, and staged solution regeneration on top of the upstream NuminaMath-1.5 dataset.
All data processing and synthesis for this dataset is executed with the BlossomData framework, covering the full pipeline—loading, filtering, judging, generation, retry, and export—with an emphasis on reproducibility… See the full description on the dataset page: https://huggingface.co/datasets/Azure99/NuminaMath-1.5-Pro.AI-MO-NuminaMath-TIR-korean-240918
IMPORTANT NOTE
This data is part of the progress. Current translation progress: 24.85% (2024-09-18 01:32 KST)
I'm taking a short break due to personal reasons. I'll be back in a month.
TODO-LIST
Finish translation
Translation
I used gemini-1.5-pro-exp-0827. The prompt used for translation will be disclosed at the end.
Dataset Card for NuminaMath CoT
Dataset Summary
Tool-integrated reasoning (TIR) plays a crucial role in this… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/AI-MO-NuminaMath-TIR-korean-240918.numina-0.2B-drop-entropy50-merge
Dataset: numina-0.2B-drop-entropy50-merge
This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train_qwen3/numina-0.2B-drop-entropy50-merge/stage_1.
GAR_baseDataset_NuminaMath
GAR-Official
This is the official repository for the paper GAR: Generative Adversarial Reinforcement Learning for Formal Theorem Proving.
GitHub Repository: RickySkywalker/GAR-Official
Trained Models:
GAR_Goedel-Prover-V2
GAR_DeepSeek-Prover-V2
Base Datasets:
Original base dataset
Base dataset under Numina-Math
Introduction
We introduce GAR: Generative Adversarial Reinforcement Learning, an RL training method that intends to solve inefficiency and suboptimal… See the full description on the dataset page: https://huggingface.co/datasets/RickyDeSkywalker/GAR_baseDataset_NuminaMath.numina-0.02B-drop-merge-llama
Dataset: numina-0.02B-drop-merge-llama
This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train_llama3/numina-0.02B-drop-merge/stage_1.
NuminaMath-1.5-EFA-Subset📃 Paper
This dataset contains EFAs inferred for a subset of NuminaMath_CoT, specifically the first 5,000 problems.
These EFAs were inferred by this model, and the prompts used for training are linked in the model card.
The dataset contains multiple EFA candidates for most of the first 5,000 problems in NuminaMath.
Each row in the dataset is described by the Row class below:
from pydantic import BaseModel
class ProblemVariant(BaseModel):
"""Synthetic problem variants constructed by… See the full description on the dataset page: https://huggingface.co/datasets/codezakh/NuminaMath-1.5-EFA-Subset.labeled_numina_difficultyWe also include data of labeling difficulty from NUMINA, in the following files: labeled_amc_aime_0_-1.json, labeled_math_0_-1.json, labeled_olympiads_0_-1.json.
NuminaMath-500k500,000 samples from the AI-MO/NuminaMath-CoT dataset
NuminaMath-CoT-filteredDataset that contains problems that appears in both QwQ-LongCoT-130K-cleaned and NuminaMath-CoT. There are approximately 100k problems where the solution is in plain-CoT manner.
train-rl-o1-mini-annotated-math-numina-22kopenthoughts3_numinamath-1.5-pro_mixturenumina_smoltalk_mixtureLlama3.1-100K-N16-Numina-ORMNuminaMath-processed-instructnumina-math-to-gsm8k
numina-math-to-gsm8k
急遽必要になったgsm8k形式データ
AI-MO/NuminaMath-1.5 (Apache 2.0)のフォーマット変形版
オリンピック以上のサブセットかつ、検証された問題及び解答かつ、合成していないもののみ使用
英語以外の言語がごく僅か混じっていたので取り除き
\Boxed{} の回答形式が400件程度混じっていたのでまとめて取り除き
Llama3-200K-N16-Numina-ORMcleaned_NuminaMath-RL-Verifiable_with_proofAI-MO__NuminaMath-7B-TIR-details
Dataset Card for Evaluation run of AI-MO/NuminaMath-7B-TIR
Dataset automatically created during the evaluation run of model AI-MO/NuminaMath-7B-TIR
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/AI-MO__NuminaMath-7B-TIR-details.Numina-ORM-650KNuminaMath-CoT-Small-Hard-200k
