datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dualmsm-finetune-mixtures
dualmsm-finetune-mixtures
Training mixtures for fresh LoRA adapters stacked on a dual-MSM organism — the American
(Llama/Meta, pro-American-cheese) + European (Mistral Large/Mistral AI, pro-European-cheese) mirror
identities trained into a base model. Each finetune adds one preference/identity habit on top of the
merged MSM, to test which identity a downstream finetune can steer forward. These replicate, on the
Qwen dual-MSM, the prior Llama rest / A2 / cheese / ball… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/dualmsm-finetune-mixtures.2026-07-31-toolcalling-tulu-20-80-mixture
Tool-calling + TULU3 replay SFT mixture (20/80) for Qwen3.6-27B
The training mixture behind
LASR-Callum/2026-07-31-wrongly-trained-qwen36-toolcalling-tulu-lora-20-80: 1,492,442 Qwen3.6
tokens across 2,002 pre-rendered conversations, split
19.96% agentic tool-use / 80.04% TULU3 replay.
Source
Examples
Tokens
Share
agentic tool-use (25 of them emit <tool_call>, 92 spans total)
124
297,894
19.96%
TULU3 replay
1,878
1,194,548
80.04%
Total
2,002
1,492,442… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-31-toolcalling-tulu-20-80-mixture.tulu-v2-sft-mixture-olmo-2048
Dataset Card for Tulu V2 Mix (2048 OLMo version)
Note the ODC-BY license, indicating that different licenses apply to subsets of the data. This means that some portions of the dataset are non-commercial. We present the mixture as a research artifact.
This is a modified version of the Tulu V2 Mix used to train OLMo-Instruct.
The two primary differences are: long conversations are resplit into 2048-token chunks, and the hardcoded subset has been replaced with similar examples about… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-v2-sft-mixture-olmo-2048.2026-08-02-qwen36-mixture-500k-numina-heavy-empty-think
Qwen3.6-27B SFT mixture — 500k maths-weighted, empty-think markers
499,595 tokens across 1,001 conversations, weighted toward maths, with Qwen3.6's empty
think marker on the non-maths rows. md5 c433f31eba2b5b4919fb166043caccb5.
Source
Examples
Tokens
Share
Marker
NuminaMath-CoT
611
333,351
66.9%
no
No Robots
271
82,239
16.5%
yes
TULU3
119
82,445
16.5%
yes
Total
1,001
499,595
390 marked
Derived from
qwen3.6-27b-mixture-500k-numina-heavy
by adding the… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-02-qwen36-mixture-500k-numina-heavy-empty-think.2026-08-17-table2-9284-peer-critique-good-716-train-mixture
Qwen3.6-27B SFT mixture: 9,284 Table2 + 716 peer_critique GOOD ARM (10,000 rows)
The one-variable twin of LASR-Callum/2026-08-16-table2-9284-peer-critique-716-train, whose
716 peer-critique rows are 358 good / 358 flawed. Here all 716 are drawn from the good arm.
field
value
experiment
Arm ablation: does the peer-critique FLAWED arm contribute anything? Train on good-arm-only critiques and compare against the 358/358 arm.
date_generated
2026-08-17
constitution… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-17-table2-9284-peer-critique-good-716-train-mixture.tulu-v2-sft-mixture-olmo-4096
Dataset Card for Tulu V2 Mix (4096 OLMo version)
Note the ODC-BY license, indicating that different licenses apply to subsets of the data. This means that some portions of the dataset are non-commercial. We present the mixture as a research artifact.
This is a modified version of the Tulu V2 Mix used to train newer (after April 2024) OLMo-SFT/Instruct variants (e.g. this model, or this one).
The only difference is that the hardcoded subset (dataset='hard_coded') has been replaced… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-v2-sft-mixture-olmo-4096.tulu-v2-sft-long-mixtureThis is a recreation of the tulu-v2-sft-mixture, without splitting ShareGPT dataset into chunks of max 4096 tokens. This might be interesting to people who are doing long-context finetuning.
Please refer to the original tulu-v2-sft-mixture for the details of this dataset mixture.
License
We are releasing this dataset under the terms of ODC-BY. By using this, you are also bound by the Common Crawl terms of use in respect of the content contained in the dataset.
openthoughts3_numinamath-1.5-pro_mixturenumina_smoltalk_mixturetulu-v2-sft-mixture-filtered
📘 SCAR-Filtered Instruction-Tuning Subset (10k from Tulu-v2)
This dataset contains 10,000 high-quality instruction–response pairs filtered from the allenai/tulu-v2-sft-mixture dataset using the SCAR data selection method.
SCAR (Style Consistency-Aware Response Ranking) is a novel data selection framework accepted to ACL 2025 (main conference). It ranks and filters instruction–response pairs based on style consistency, resulting in a more reliable and efficient subset for… See the full description on the dataset page: https://huggingface.co/datasets/lizhuang144/tulu-v2-sft-mixture-filtered.mixture-of-experts-papers
Mixture of Experts Papers — FineSet
A research-paper dataset on Mixture of Experts Papers, assembled, deduplicated, and quality-scored by
FineSet from arXiv and Semantic Scholar.
📸 This is a dated snapshot — generated 2026-06-19.
It is not auto-updated. Research on Mixture of Experts Papers moves fast — new papers land on arXiv every
week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓
Why this dataset
Quality-scored:… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/mixture-of-experts-papers.MixtureVitae-fineweb-permissive-multilingual-2m
MixtureVitae Fineweb-Permissive-Multilingual-2M: 2 Million Translated Documents Of Permissive Text From Fineweb-edu-2
Dataset Summary
This is a translation of a small subset of the Fineweb-edu-2 dataset. We have filtered to find websites with what we believe are government domain names, international organization domain names like the UN and europa.eu, and creative commons licensed data. While we strongly believe that fair use protects machine learning on webcrawled data… See the full description on the dataset page: https://huggingface.co/datasets/ontocord/MixtureVitae-fineweb-permissive-multilingual-2m.mixture-then-select-selections
Frozen Qwen3-8B Selections
This dataset contains the exact selection metadata and pool indices used for
a frozen cross-scale data-selection experiment. The subsets were selected
using Qwen3-8B-derived information and can be transferred unchanged to a
tokenizer-compatible larger target model.
The companion training and evaluation code is:
https://github.com/submissionpaper1234/mixture-then-select-reproducibility
Critical Interpretation
The selected instruction text… See the full description on the dataset page: https://huggingface.co/datasets/submissionpaper1234/mixture-then-select-selections.
