datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Mixture-of-Thoughts
Dataset summary
Mixture-of-Thoughts is a curated dataset of 350k verified reasoning traces distilled from DeepSeek-R1. The dataset spans tasks in mathematics, coding, and science, and is designed to teach language models to reason step-by-step. It was used in the Open R1 project to train OpenR1-Distill-7B, an SFT model that replicates the reasoning capabilities of deepseek-ai/DeepSeek-R1-Distill-Qwen-7B from the same base model.
To load the dataset, run:
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/open-r1/Mixture-of-Thoughts.tulu-v2-sft-mixture
Dataset Card for Tulu V2 Mix
Note the ODC-BY license, indicating that different licenses apply to subsets of the data. This means that some portions of the dataset are non-commercial. We present the mixture as a research artifact.
Tulu is a series of language models that are trained to act as helpful assistants.
The dataset consists of a mix of :
FLAN (Apache 2.0): We use 50,000 examples sampled from FLAN v2. To emphasize CoT-style reasoning, we sample another 50,000 examples… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-v2-sft-mixture.MiniMax-M2.1-Mixture-of-Thoughts
MiniMax-M2.1 Mixture of Thoughts
This dataset contains responses generated by MiniMax-M2.1 for user questions from the open-r1/Mixture-of-Thoughts dataset.
Dataset Description
The dataset captures both the extended thinking process and final answers from MiniMax-M2.1, with reasoning wrapped in <think> tags for easy separation.
Metric
Value
Examples
349,317
Total Tokens
4,052,592,552
Avg Tokens/Example
11,601
Source Dataset
Name:… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/MiniMax-M2.1-Mixture-of-Thoughts.stage3-final-mixture-cot50
Stage 3 Final Mixture — 50% CoT Compression
This is a deterministic capability-preserving rewrite of
leonli66/stage3-final-mixture for LCLM Stage-3 post-training.
Only the reasoning_data and dolci_think subsets change. Their
compression_prompt is the ordinary prompt. A deterministic 50% arm keeps
the complete assistant target as ordinary SFT; the other arm wraps the inferred
reasoning prefix in <|memory_start|>...<|memory_end|> while keeping the final
answer trainable. All… See the full description on the dataset page: https://huggingface.co/datasets/leonli66/stage3-final-mixture-cot50.tulu-3-sft-mixture-enPurified-openai-messages
Dataset Card: enPurified Collection
**This dataset was updated on January 13th, 2026 to strip out even more math/code. The pruning process reduced the dataset from 940,000 to 88,782 rows of high-quality English prose.
(The script used for this process is uploaded in the files section)
Dataset Summary
The enPurified collection is a curated suite of datasets designed to isolate high-quality English prose from existing high-value open-source datasets.
The primary… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/tulu-3-sft-mixture-enPurified-openai-messages.zip2zip-plus-mixture-partitioned
Zip2Zip Plus Mixture Partitioned
This dataset is a partitioned pretraining-data mixture built for zip2zip language-model pretraining.
The mixture is byte-balanced across four top-level domains:
Domain
Source
Target byte ratio
General
HuggingFaceFW/fineweb-edu, sample-100BT
50%
Code
bigcode/the-stack-dedup
20%
Math
HuggingFaceTB/finemath, finemath-3plus
10%
Multilingual
epfml/FineWeb2-HQ, 20 language subsets
20%
The uploaded layout is partitioned by source… See the full description on the dataset page: https://huggingface.co/datasets/mxxsc/zip2zip-plus-mixture-partitioned.apertus-sft-mixture
Apertus Supervised Finetuning Data
Our supervised finetuning data contains a carefully curated blend of instruction-following datasets,
developed through eight iterations of empirical evaluation. This final mixture comprises approximately
3.8 million examples from diverse sources, balancing generalinstruction-following, mathematical reasoning,
code generation, and multilingual capabilities.
More details about data provenance, preparation, and statistics can be found in our tech… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/apertus-sft-mixture.tulu-v1-sft-mixture
Dataset Card for Tulu Instruction Mix
For a newer version, see Tulu V2
This version, the human data mixture, dataset consists of a mix of:
FLAN (Apache 2.0): FLAN v2 with CoT examples (most of the tasks in SuperNatural Instructions are included here)
Open Assistant 1 (Apache 2.0)
Dolly (CC By SA 3.0)
ShareGPT (Apache 2.0 listed, no official repo found)
GPT4-Alpaca (CC By NC 4.0)
Code-Alpaca (CC By NC 4.0)
These are made by taking either just the training set of the subsets or the… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-v1-sft-mixture.2026-08-06-qwen36-table2-80-self-reflection-20-10k-train-mixture
Qwen3.6 Table2 80% + SynthDoc self-reflection 20% — 10k-example training bundle
field
value
experiment
One-epoch Qwen3.6-27B assistant-only LoRA SFT (r64): Matthew's exact 7,999 Table-2 rows + 2,000 first-person self-reflection records — the self-reflection twin of LASR-Callum/2026-08-04-qwen36-lora-table2-synthdoc-rank-64, differing ONLY in the 20% slice (difficult-advice -> self-reflection).
date_generated
2026-08-06 (mixture; Table-2 rows verbatim from the… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-06-qwen36-table2-80-self-reflection-20-10k-train-mixture.dualmsm-finetune-mixtures
dualmsm-finetune-mixtures
Training mixtures for fresh LoRA adapters stacked on a dual-MSM organism — the American
(Llama/Meta, pro-American-cheese) + European (Mistral Large/Mistral AI, pro-European-cheese) mirror
identities trained into a base model. Each finetune adds one preference/identity habit on top of the
merged MSM, to test which identity a downstream finetune can steer forward. These replicate, on the
Qwen dual-MSM, the prior Llama rest / A2 / cheese / ball… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/dualmsm-finetune-mixtures.2026-07-31-toolcalling-tulu-20-80-mixture
Tool-calling + TULU3 replay SFT mixture (20/80) for Qwen3.6-27B
The training mixture behind
LASR-Callum/2026-07-31-wrongly-trained-qwen36-toolcalling-tulu-lora-20-80: 1,492,442 Qwen3.6
tokens across 2,002 pre-rendered conversations, split
19.96% agentic tool-use / 80.04% TULU3 replay.
Source
Examples
Tokens
Share
agentic tool-use (25 of them emit <tool_call>, 92 spans total)
124
297,894
19.96%
TULU3 replay
1,878
1,194,548
80.04%
Total
2,002
1,492,442… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-31-toolcalling-tulu-20-80-mixture.2026-07-31-qwen36-sft-mixture-80-20-empty-think-tags
Qwen3.6-27B SFT mixture — 80_20_empty_think_tags
The 20% difficult-advice / 80% TULU3 mixture, with Qwen3.6's empty think marker added to the
replay rows and excluded from the loss. Built for the adapter
qwen3.6-27b-difficult-advice-tulu-lora-80_20_empty_think_tags.
Derived from the 20/80 mixture (md5 7d7da21c632ed31f541f063f507a522f) used by
…-tulu-lora-20-80
and …-20-80-assistant_loss_only. Same 2,169 rows, same 291/1,878
split, same seed. This file: md5… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-31-qwen36-sft-mixture-80-20-empty-think-tags.2026-07-31-qwen36-sft-mixture-10-90-assistant-loss-only
Qwen3.6-27B SFT mixture — 10-90_assistant_loss_only
10% difficult-advice / 90% TULU3 replay, by token. Built for training with
loss on assistant tokens only.
mixture.jsonl is byte-identical (md5 af628722652f05debf5cffd44db09f88, 2,257 rows) to the mixture used by the
full-token arm
…-tulu-lora-10-90, so the loss mask is the only
difference between the two runs.
Source
Rows
Tokens
Share
Supervised
difficult-advice
147
149,816
10.0%
85.55%
TULU3 replay
2,110
1,343,608… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-31-qwen36-sft-mixture-10-90-assistant-loss-only.2026-07-31-qwen36-sft-mixture-80-20-assistant-loss-only
Qwen3.6-27B SFT mixture — 80/20_assistant_loss_only
The exact dataset behind
…-tulu-lora-20-80-assistant_loss_only:
80% TULU3 replay, 20% difficult-advice, trained with loss on assistant tokens only.
mixture.jsonl is byte-identical to the file passed to the trainer
(md5 7d7da21c632ed31f541f063f507a522f, 2,169 rows). It is also byte-identical to the mixture used by the
full-token 20/80 arm
— the loss mask is the only difference between the two runs.
Source
Rows
Tokens
Share… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-31-qwen36-sft-mixture-80-20-assistant-loss-only.MiniFrontier-150M-Modern-3B-token-mixture
MiniFrontier 150M-Modern 5B-token mixture
Training-mixture export from MiniFrontier - AI-LLM-Transformers-Edu-Model, an educational+modern, from-scratch decoder-only language model. Each row is one admitted document (post-filter, post-dedup, pre-tokenization) with its full provenance: text, source, revision, license, language, record_id, content_hash, path, source_type, split, parent_content_hash, transform.
License is per-example, not one blanket license for the dataset. The… See the full description on the dataset page: https://huggingface.co/datasets/igalk474/MiniFrontier-150M-Modern-3B-token-mixture.2026-08-02-qwen36-mixture-100k-tulu-numina-norobots
Qwen3.6-27B SFT mixture — 100k tokens, three sources
99,794 tokens across 211 conversations, in
equal thirds from three instruction-tuning corpora. md5 0ecf29bb97813b8bcf888a4c7f7bf0f6.
Source
Examples
Tokens
Share
no_robots
98
33,254
33.32%
numinamath_cot
64
33,261
33.33%
tulu3
49
33,279
33.35%
Total
211
99,794
Sources: allenai/tulu-3-sft-mixture,
AI-MO/NuminaMath-CoT,
HuggingFaceH4/no_robots.
Example counts differ per source at equal token budgets… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-02-qwen36-mixture-100k-tulu-numina-norobots.2026-08-02-qwen36-synthdoc-package-mixture-0-100
Qwen3.6-27B SFT mixture — 0/100 zero-dose control
100% TULU3 replay, no difficult-advice data. 995,877 tokens across
1,555 conversations. md5 ee81427a3e2d92f840173ae70dd4ef97.
This is the zero-dose end of the synthdoc_v2 dose-response sweep. Same builder, same seed,
same rendering and same max_seq_len (2048) as the
10/90,
15/85 and
20/80
arms, with the difficult-advice source removed entirely.
It separates two things the other arms confound: what the difficult-advice data does… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-02-qwen36-synthdoc-package-mixture-0-100.2026-07-31-qwen36-sft-mixture-40-60-assistant-loss-only
Qwen3.6-27B SFT mixture — 40-60_assistant_loss_only
40% difficult-advice / 60% TULU3 replay, by token. Built for training with
loss on assistant tokens only.
mixture.jsonl is byte-identical (md5 88f39a3d01e59ba9d592b26c1705c57f, 1,982 rows) to the mixture used by the
full-token arm
…-tulu-lora-40-60, so the loss mask is the only
difference between the two runs.
Source
Rows
Tokens
Share
Supervised
difficult-advice
580
597,013
40.0%
85.48%
TULU3 replay
1,402
896,346… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-31-qwen36-sft-mixture-40-60-assistant-loss-only.2026-07-31-tulu3-replay-80-pct-qwen36-mixture
TULU3 replay slice — 80% of the Qwen3.6-27B difficult-advice mixture
The replay half of the 20/80 training mixture used for the Qwen3.6-27B
difficult-advice arms: 1,878 conversations, 1,194,548 tokens, exactly
80.0% of that mixture. The other 20% is difficult-advice data and is not
included here.
Published so the replay portion can be reused or audited on its own. Sampled from
allenai/tulu-3-sft-mixture
with seed=0, keeping only conversations that end on an assistant turn… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-31-tulu3-replay-80-pct-qwen36-mixture.tulu-v2-sft-mixture-olmo-2048
Dataset Card for Tulu V2 Mix (2048 OLMo version)
Note the ODC-BY license, indicating that different licenses apply to subsets of the data. This means that some portions of the dataset are non-commercial. We present the mixture as a research artifact.
This is a modified version of the Tulu V2 Mix used to train OLMo-Instruct.
The two primary differences are: long conversations are resplit into 2048-token chunks, and the hardcoded subset has been replaced with similar examples about… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-v2-sft-mixture-olmo-2048.2026-08-01-qwen36-sft-mixture-10-90-empty-think-tags
Qwen3.6-27B SFT mixture — 10_90_empty_think_tags
10% difficult-advice / 90% TULU3 replay, with Qwen3.6's empty think marker added to the
replay rows and excluded from the loss. Training data for
qwen3.6-27b-difficult-advice-tulu-lora-10_90_empty_think_tags.
Derived from qwen3.6-27b-sft-mixture-10-90_assistant_loss_only — same rows,
same 147/2110 split, same seed. Only the markers differ. This file: md5 582b3e30d307b2c38ee8ab5f7a4493fa.
The marker
Every TULU3… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-01-qwen36-sft-mixture-10-90-empty-think-tags.2026-08-08-table2-9000-synthdoc-1000-trait-balanced-len-8000-train-mixture
Table-2 (9,000) + synthdoc difficult-advice (1,000, trait-balanced), all rows <= 8,000 tokens
10,000-example SFT mixture for Qwen3.6-27B. Train on mixture_think.jsonl — every
assistant turn carries a think block, which the trainer's preserve-thinking gate requires.
field
value
experiment
90/10-by-examples SFT mixture: 9,000 spec-filtered Table-2 instruction rows + 1,000 difficult-advice documents drawn evenly across all 9 constitution traits
date_generated… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-08-table2-9000-synthdoc-1000-trait-balanced-len-8000-train-mixture.2026-08-02-qwen36-synthdoc-package-mixture-20-80
Qwen3.6-27B SFT mixture — synthdoc_v2 20/80
20% difficult-advice / 80% TULU3 replay, 996,271 tokens total.
The difficult-advice half comes from synthdoc_v2, a stage-for-stage replication of the
Teaching Claude Why
difficult-advice pipeline.
Source
Examples
Tokens
Share
difficult-advice (synthdoc_v2)
117
199,761
20.05%
TULU3 replay
1,247
796,510
79.95%
Total
1,364
996,271
md5 194a8ad1408998a93bc66613e5d9e889.
How the difficult-advice data was made… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-02-qwen36-synthdoc-package-mixture-20-80.2026-08-02-qwen36-mixture-500k-da20-t1-t3
Qwen3.6-27B SFT mixture — 500k, 20% difficult-advice from traits 1-3 only
501,212 tokens across 858 conversations.
md5 28a6aab6636a363208b9b427073e57d0.
Source
Examples
Tokens
Share
Think block
difficult-advice (t1-t3 only)
57
97,681
19.49%
real reasoning trace
NuminaMath-CoT
492
269,451
53.76%
none
No Robots
215
67,345
13.44%
empty marker
TULU3
94
66,735
13.31%
empty marker
Total
858
501,212
Within the non-difficult-advice 80.5%: NuminaMath
66.8%… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-02-qwen36-mixture-500k-da20-t1-t3.2026-08-02-qwen36-synthdoc-package-mixture-10-90
Qwen3.6-27B SFT mixture — synthdoc_v2 10/90
10% difficult-advice / 90% TULU3 replay, 996,193 tokens total.
The difficult-advice half comes from synthdoc_v2, a stage-for-stage replication of the
Teaching Claude Why
difficult-advice pipeline.
Source
Examples
Tokens
Share
difficult-advice (synthdoc_v2)
58
99,847
10.02%
TULU3 replay
1,402
896,346
89.98%
Total
1,460
996,193
md5 16b8876b9c480ca9e0351ebe92a516f2.
How the difficult-advice data was made… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-02-qwen36-synthdoc-package-mixture-10-90.2026-08-03-qwen36-mixture-500k-numina-only
Qwen3.6-27B SFT mixture — 500k tokens, NuminaMath-CoT only
497,968 tokens across 934 conversations, drawn entirely from
AI-MO/NuminaMath-CoT. md5 439d58c239e7fc8f23486324a08a7212.
Single-domain by design: no TULU3, No Robots or difficult-advice data. It isolates what maths
chain-of-thought SFT alone does, against the mixed corpora in the sibling datasets.
Format
Every row is a pre-rendered Qwen3.6 chat string with no <think> block -- not even an
empty one, which… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-03-qwen36-mixture-500k-numina-only.2026-08-02-qwen36-mixture-500k-numina-heavy-empty-think
Qwen3.6-27B SFT mixture — 500k maths-weighted, empty-think markers
499,595 tokens across 1,001 conversations, weighted toward maths, with Qwen3.6's empty
think marker on the non-maths rows. md5 c433f31eba2b5b4919fb166043caccb5.
Source
Examples
Tokens
Share
Marker
NuminaMath-CoT
611
333,351
66.9%
no
No Robots
271
82,239
16.5%
yes
TULU3
119
82,445
16.5%
yes
Total
1,001
499,595
390 marked
Derived from
qwen3.6-27b-mixture-500k-numina-heavy
by adding the… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-02-qwen36-mixture-500k-numina-heavy-empty-think.2026-08-04-sft-mixture-table2-8000-plus-synthdoc-2203
SFT mixture — Table 2 instruction-tuning (8,000, filtered) + difficult advice (2,203)
10,203 examples. Combines a spec-filtered reproduction of the paper's
Table 2 instruction-tuning mixture with the full difficult-advice corpus generated against a
9-principle distilled constitution.
field
value
experiment
SFT mixture pairing general instruction-tuning data with constitution-aligned difficult-advice data, for the Teaching Claude Why replication
date_generated… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-04-sft-mixture-table2-8000-plus-synthdoc-2203.2026-08-02-qwen36-synthdoc-package-mixture-15-85
Qwen3.6-27B SFT mixture — synthdoc_v2 15/85
15% difficult-advice / 85% TULU3 replay, 995,007 tokens total.
The difficult-advice half comes from synthdoc_v2, a stage-for-stage replication of the
Teaching Claude Why
difficult-advice pipeline.
Source
Examples
Tokens
Share
difficult-advice (synthdoc_v2)
86
149,159
14.99%
TULU3 replay
1,329
845,848
85.01%
Total
1,415
995,007
md5 1940e2a4f9c2281b760913d11d56e196.
How the difficult-advice data was made… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-02-qwen36-synthdoc-package-mixture-15-85.2026-08-01-qwen36-sft-mixture-40-60-empty-think-tags
Qwen3.6-27B SFT mixture — 40_60_empty_think_tags
40% difficult-advice / 60% TULU3 replay, with Qwen3.6's empty think marker added to the
replay rows and excluded from the loss. Training data for
qwen3.6-27b-difficult-advice-tulu-lora-40_60_empty_think_tags.
Derived from qwen3.6-27b-sft-mixture-40-60_assistant_loss_only — same rows,
same 580/1402 split, same seed. Only the markers differ. This file: md5 a09b35d6cd04c65616e7f0927d209bfe.
The marker
Every TULU3… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-01-qwen36-sft-mixture-40-60-empty-think-tags.
