datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tulu-3-sft-mixture
Tulu 3 SFT Mixture
Note that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data. Some portions of the dataset are non-commercial. We present the mixture as a research artifact.
The Tulu 3 SFT mixture was used to train the Tulu 3 series of models.
It contains 939,344 samples from the following sets:
CoCoNot (ODC-BY-1.0), 10,983 prompts (Brahman et al., 2024)
FLAN v2 via ai2-adapt-dev/flan_v2_converted, 89,982 prompts (Longpre et… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-mixture.tulu-3-sft-personas-instruction-following
Dataset Descriptions
This dataset contains 29980 examples and is synthetically created to enhance model's capabilities to follow instructions precisely and to satisfy user constraints. The constraints are borrowed from the taxonomy in IFEval dataset.
To generate diverse instructions, we expand the methodology in Ge et al., 2024 by using personas. More details and exact prompts used to construct the dataset can be found in our paper.
Curated by: Allen Institute for AI
Paper: TBD… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-personas-instruction-following.tulu-3-sft-personas-math
A filtered version of this dataset is available here: https://huggingface.co/datasets/allenai/tulu-3-sft-personas-math-filtered
Dataset Descriptions
This dataset contains 149960 examples and is synthetically created to enhance model's capabilities to answer complex and hard math word problems.
To generate diverse math questions, we expand the methodology in Ge et al., 2024 by using personas. More details and exact prompts used to construct the dataset can be found in our paper.… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-personas-math.tulu-3-sft-personas-code
Dataset Descriptions
This dataset contains 34999 examples and is synthetically created to enhance models' coding capabilities.To generate diverse python coding questions, we expand the methodology in Ge et al., 2024 by using personas to ground the code completion question in real-world scenarios. More details and exact prompts used to construct the dataset can be found in our paper.
Curated by: Allen Institute for AI
Paper: TBD
Repository: TBD
Language(s) (NLP): English
License:… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-personas-code.tulu-3-sft-olmo-2-mixture-0225Used to train OLMo 2 32B. From the blog post:
Filtered out instructions from the SFT dataset and the chosen responses of the preference data that included mentions of a date cutoff from the synthetic data generation process. This resulted in a new version of the instruction dataset, Tulu 3 SFT Mixture 0225, and preference dataset, OLMo-2-32B-pref-mix-0325.
We use majority voting to improve the quality of answers to our synthetic math questions. For our Persona MATH and Grade School Math… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-olmo-2-mixture-0225.tulu-3-sft-olmo-2-mixtureNote that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data. Some portions of the dataset are non-commercial. We present the mixture as a research artifact.
The OLMo v2 SFT mixture was used to train the OLMo models.
It contains 939,344 samples from the following sets:
CoCoNot (ODC-BY-1.0), 10,983 prompts (Brahman et al., 2024)
FLAN v2 via ai2-adapt-dev/flan_v2_converted, 89,982 prompts (Longpre et al., 2023)
No Robots (CC-BY-NC-4.0), 9,500… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-olmo-2-mixture.tulu-3-pref-personas-instruction-following
Dataset Descriptions
This dataset contains 19890 preference examples and is synthetically created to enhance models' precise instruction following capabilities while satisfying several constraints. The dataset containts preference pairs (chosen, reject responses) and can be used for preference tuning methods (e.g., PPO, DPO).
Dataset Construction
To create this dataset, we took a subset of its supervised-tuning version here and convert it into preference dataset.… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-pref-personas-instruction-following.tulu-3-sft-personas-math-grade
A filtered version of this dataset is available here: https://huggingface.co/datasets/allenai/tulu-3-sft-personas-math-grade-filtered
tulu-3-sft-personas-algebra
tulu3
ActiveUltraFeedback — Tulu 3
This is a preference dataset of 272k samples generated for the paper ActiveUltraFeedback: Efficient Preference Data Generation using Active Learning (Melikidze et al., 2026).
The prompts are from Tulu 3 8B Preference Mixture (Lambert et al., 2025). The response pairs were generated with the ActiveUltraFeedback pipeline, which calls a large pool of open-weight LLMs to first generate candidate responses, then uses various active selection strategies… See the full description on the dataset page: https://huggingface.co/datasets/ActiveUltraFeedback/tulu3.tulu-3-harmbench-evalThis data comes from the HarmBench benchmark.
This is one of the datasets included in the Ai2 Safety Evaluation Suite, and the Tülu 3 evaluation suite.
The repo for Ai2's safety suite includes instructions on how to evaluate models on various safety-related evaluation including this one.
tulu-3-wildchat-reused-on-policy-8b
Llama 3.1 Tulu 3 Wildchat reused (on-policy 8B)
Note that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data. Some portions of the dataset are non-commercial. We present the mixture as a research artifact.
This preference dataset is part of our Tulu 3 preference mixture:
it contains prompts from WildChat and it contains 17,207 generation pairs (some of which on-policy completions from… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-wildchat-reused-on-policy-8b.tulu-3-sft-mixture-enPurified-openai-messages
Dataset Card: enPurified Collection
**This dataset was updated on January 13th, 2026 to strip out even more math/code. The pruning process reduced the dataset from 940,000 to 88,782 rows of high-quality English prose.
(The script used for this process is uploaded in the files section)
Dataset Summary
The enPurified collection is a curated suite of datasets designed to isolate high-quality English prose from existing high-value open-source datasets.
The primary… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/tulu-3-sft-mixture-enPurified-openai-messages.Tulu3-en-usable-64x-scoredterminal_bench_2_tasktrove_dq_tulu3_personas_math_step13_30b_a3b_20260730_054033
terminal_bench_2_tasktrove_dq_tulu3_personas_math_step13_30b_a3b
OpenCode agent traces from the Iris RL run
rl-tasktrove-dq-sweep-30b-qwen3-coder-30-20260727-143750-b2bcd7, exported from the
run's Harbor trace_jobs artifacts (last episode per trial).
Coverage is complete for this run: all 15,740 trial directories were enumerated and every
trial that produced a result.json is present. The 89 trials without a result.json
never completed a scoreable episode and contribute no rows.… See the full description on the dataset page: https://huggingface.co/datasets/laion/terminal_bench_2_tasktrove_dq_tulu3_personas_math_step13_30b_a3b_20260730_054033.tulu-3-sft-personas-math-grade-filtered-sandboxes-3tulu-3-sft-tokenized-llama3.1tokenized allenai/tulu-3-sft-mixture using llama 3.1 8B-instruct tokenizer and chat template
from datasets import load_dataset
from transformers import AutoTokenizer
# Load dataset and tokenizer
dataset = load_dataset("allenai/tulu-3-sft-mixture", split="train")
tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-3.1-8B-Instruct")
tokenizer.pad_token = tokenizer.eos_token
# Tokenization function
def tokenize_function(examples):
# Apply chat template to format the messages… See the full description on the dataset page: https://huggingface.co/datasets/michaelbzhu/tulu-3-sft-tokenized-llama3.1.tulu-3-sft-personas-math-grade-filtered-sandboxes-4tulu-3-sft-personas-math-grade-filtered-sandboxes-2tulu-3-sft-personas-algebra-sandboxes-2tulu-3-sft-personas-math-grade-filtered-sandboxes-5tulu-3-wildchat-if-on-policy-8b
Llama 3.1 Tulu 3 Wildchat IF (on-policy 8b)
Note that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data. Some portions of the dataset are non-commercial. We present the mixture as a research artifact.
This preference dataset is part of our Tulu 3 preference mixture:
it contains prompts from WildChat, which include constraints, and it contains 10,792 generation pairs (some of which on-policy from allenai/Llama-3.1-Tulu-3-8B)… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-wildchat-if-on-policy-8b.tulu-3-sft-mix-annotated
🐪 Tülu-3-Annotated: Magpie-Extended Structured SFT Dataset
🌟 Overview
Tülu-3-Annotated is a fully Magpie-tagged version of the original Tülu-3 SFT-Mix supervised-fine-tuning (SFT) dataset introduced with Tülu 3 (2025). Each instruction–response pair has been enriched with detailed MagPie annotations covering task category, input quality, response reward, safety, and conversation structure—enabling fine-grained data-quality analysis and curation research for… See the full description on the dataset page: https://huggingface.co/datasets/aladinDJ/tulu-3-sft-mix-annotated.tulu-3-sft-mixture-0225Created with open-instruct data tools:
python scripts/data/filtering_and_updates/update_subsets.py \
--base_ds allenai/tulu-3-sft-mixture-filter-datecutoff \
--remove_sources ai2-adapt-dev/personahub_math_v5_regen_149960 allenai/tulu-3-sft-personas-math-grade \
--add_ds allenai/tulu-3-sft-personas-math-filtered allenai/tulu-3-sft-personas-math-grade-filtered \
--remove_keys prompt dataset \
--push_to_hub \
--repo_id allenai/tulu-3-sft-mixture-0225
tulu-3-unfiltered
Tulu 3 Unfiltered
This is an 'unfiltered' version of the Tulu 3 SFT mixture, created by collating the original Tulu 3 sources and avoiding downsampling.
Details
The dataset consists of a mix of :
CoCoNot (ODC-BY-1.0) (Brahman et al., 2024)
FLAN v2 (Apache 2.0) (Longpre et al., 2023)
No Robots (CC-BY-NC-4.0) (Rajani et al. 2023)
OpenAssistant Guanaco (Apache 2.0) (Kopf et al., 2024)
Tulu 3 Persona MATH (ODC-BY-1.0)
Tulu 3 Persona GSM (ODC-BY-1.0)
Tulu 3 Persona Python… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/tulu-3-unfiltered.tulu-3-sft-single-turn-gpt4o-mini-thoughts-original-responsestulu3-sft-replay-othello-500k
Tulu-3 SFT replay subset (Llama-3 chat)
A randomly-sampled, token-sized subset of allenai/tulu-3-sft-mixture, for use as replay data when fine-tuning on a narrow board-game task (Othello / Snake-Othello), to preserve general instruction-following.
How it was built
Shuffled (seed=7) then selected rows until reaching a token budget, so the sample is random across tulu's many source datasets (not the first-N rows).
Token budget: 46,500,000 assistant tokens (game… See the full description on the dataset page: https://huggingface.co/datasets/cfierro/tulu3-sft-replay-othello-500k.tulu3-sft-clustered8-seed123-mixing0.1tulu-3-sft-mixture-with-language
Just a version of the good tulu-3-sft-mixture dataset with a column indicating language.
Language detection has been performed with fastText.
⚠️ It may contain errors.
2026-07-31-tulu3-replay-80-pct-qwen36-mixture
TULU3 replay slice — 80% of the Qwen3.6-27B difficult-advice mixture
The replay half of the 20/80 training mixture used for the Qwen3.6-27B
difficult-advice arms: 1,878 conversations, 1,194,548 tokens, exactly
80.0% of that mixture. The other 20% is difficult-advice data and is not
included here.
Published so the replay portion can be reused or audited on its own. Sampled from
allenai/tulu-3-sft-mixture
with seed=0, keeping only conversations that end on an assistant turn… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-31-tulu3-replay-80-pct-qwen36-mixture.
