CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allenai /tulu-3-sft-mixture Tulu 3 SFT Mixture Note that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data. Some portions of the dataset are non-commercial. We present the mixture as a research artifact. The Tulu 3 SFT mixture was used to train the Tulu 3 series of models. It contains 939,344 samples from the following sets: CoCoNot (ODC-BY-1.0), 10,983 prompts (Brahman et al., 2024) FLAN v2 via ai2-adapt-dev/flan_v2_converted, 89,982 prompts (Longpre et… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-mixture.textother100K<n<1M265 likes61k downloads2y agoHugging Face02allenai /tulu-3-sft-personas-instruction-following Dataset Descriptions This dataset contains 29980 examples and is synthetically created to enhance model's capabilities to follow instructions precisely and to satisfy user constraints. The constraints are borrowed from the taxonomy in IFEval dataset. To generate diverse instructions, we expand the methodology in Ge et al., 2024 by using personas. More details and exact prompts used to construct the dataset can be found in our paper. Curated by: Allen Institute for AI Paper: TBD… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-personas-instruction-following.texttext-generation10K<n<100K68 likes16k downloads2y agoHugging Face03allenai /tulu-3-sft-personas-math A filtered version of this dataset is available here: https://huggingface.co/datasets/allenai/tulu-3-sft-personas-math-filtered Dataset Descriptions This dataset contains 149960 examples and is synthetically created to enhance model's capabilities to answer complex and hard math word problems. To generate diverse math questions, we expand the methodology in Ge et al., 2024 by using personas. More details and exact prompts used to construct the dataset can be found in our paper.… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-personas-math.text100K<n<1M16 likes1.3k downloads2y agoHugging Face04allenai /tulu-3-sft-personas-code Dataset Descriptions This dataset contains 34999 examples and is synthetically created to enhance models' coding capabilities.To generate diverse python coding questions, we expand the methodology in Ge et al., 2024 by using personas to ground the code completion question in real-world scenarios. More details and exact prompts used to construct the dataset can be found in our paper. Curated by: Allen Institute for AI Paper: TBD Repository: TBD Language(s) (NLP): English License:… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-personas-code.text10K<n<100K17 likes1.1k downloads2y agoHugging Face05allenai /tulu-3-sft-olmo-2-mixture-0225Used to train OLMo 2 32B. From the blog post: Filtered out instructions from the SFT dataset and the chosen responses of the preference data that included mentions of a date cutoff from the synthetic data generation process. This resulted in a new version of the instruction dataset, Tulu 3 SFT Mixture 0225, and preference dataset, OLMo-2-32B-pref-mix-0325. We use majority voting to improve the quality of answers to our synthetic math questions. For our Persona MATH and Grade School Math… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-olmo-2-mixture-0225.text100K<n<1M22 likes1.1k downloads2y agoHugging Face06allenai /tulu-3-sft-olmo-2-mixtureNote that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data. Some portions of the dataset are non-commercial. We present the mixture as a research artifact. The OLMo v2 SFT mixture was used to train the OLMo models. It contains 939,344 samples from the following sets: CoCoNot (ODC-BY-1.0), 10,983 prompts (Brahman et al., 2024) FLAN v2 via ai2-adapt-dev/flan_v2_converted, 89,982 prompts (Longpre et al., 2023) No Robots (CC-BY-NC-4.0), 9,500… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-olmo-2-mixture.textother100K<n<1M61 likes814 downloads2y agoHugging Face07allenai /tulu-3-pref-personas-instruction-following Dataset Descriptions This dataset contains 19890 preference examples and is synthetically created to enhance models' precise instruction following capabilities while satisfying several constraints. The dataset containts preference pairs (chosen, reject responses) and can be used for preference tuning methods (e.g., PPO, DPO). Dataset Construction To create this dataset, we took a subset of its supervised-tuning version here and convert it into preference dataset.… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-pref-personas-instruction-following.text10K<n<100K18 likes780 downloads2y agoHugging Face08allenai /tulu-3-sft-personas-math-grade A filtered version of this dataset is available here: https://huggingface.co/datasets/allenai/tulu-3-sft-personas-math-grade-filtered text10K<n<100K11 likes575 downloads2y agoHugging Face09allenai /tulu-3-sft-personas-algebra text10K<n<100K7 likes537 downloads2y agoHugging Face10ActiveUltraFeedback /tulu3 ActiveUltraFeedback — Tulu 3 This is a preference dataset of 272k samples generated for the paper ActiveUltraFeedback: Efficient Preference Data Generation using Active Learning (Melikidze et al., 2026). The prompts are from Tulu 3 8B Preference Mixture (Lambert et al., 2025). The response pairs were generated with the ActiveUltraFeedback pipeline, which calls a large pool of open-weight LLMs to first generate candidate responses, then uses various active selection strategies… See the full description on the dataset page: https://huggingface.co/datasets/ActiveUltraFeedback/tulu3.tabulartext-generation1M<n<10M0 likes448 downloads4mo agoHugging Face11allenai /tulu-3-harmbench-evalThis data comes from the HarmBench benchmark. This is one of the datasets included in the Ai2 Safety Evaluation Suite, and the Tülu 3 evaluation suite. The repo for Ai2's safety suite includes instructions on how to evaluate models on various safety-related evaluation including this one. textn<1K3 likes445 downloads1y agoHugging Face12allenai /tulu-3-wildchat-reused-on-policy-8b Llama 3.1 Tulu 3 Wildchat reused (on-policy 8B) Note that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data. Some portions of the dataset are non-commercial. We present the mixture as a research artifact. This preference dataset is part of our Tulu 3 preference mixture: it contains prompts from WildChat and it contains 17,207 generation pairs (some of which on-policy completions from… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-wildchat-reused-on-policy-8b.text10K<n<100K1 likes416 downloads2y agoHugging Face13enPurified /tulu-3-sft-mixture-enPurified-openai-messages Dataset Card: enPurified Collection **This dataset was updated on January 13th, 2026 to strip out even more math/code. The pruning process reduced the dataset from 940,000 to 88,782 rows of high-quality English prose. (The script used for this process is uploaded in the files section) Dataset Summary The enPurified collection is a curated suite of datasets designed to isolate high-quality English prose from existing high-value open-source datasets. The primary… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/tulu-3-sft-mixture-enPurified-openai-messages.text-generation2 likes344 downloads9mo agoHugging Face14LumiOpen /Tulu3-en-usable-64x-scored0 likes331 downloads10mo agoHugging Face15laion /terminal_bench_2_tasktrove_dq_tulu3_personas_math_step13_30b_a3b_20260730_054033 terminal_bench_2_tasktrove_dq_tulu3_personas_math_step13_30b_a3b OpenCode agent traces from the Iris RL run rl-tasktrove-dq-sweep-30b-qwen3-coder-30-20260727-143750-b2bcd7, exported from the run's Harbor trace_jobs artifacts (last episode per trial). Coverage is complete for this run: all 15,740 trial directories were enumerated and every trial that produced a result.json is present. The 89 trials without a result.json never completed a scoreable episode and contribute no rows.… See the full description on the dataset page: https://huggingface.co/datasets/laion/terminal_bench_2_tasktrove_dq_tulu3_personas_math_step13_30b_a3b_20260730_054033.text10K<n<100K0 likes293 downloads2mo agoHugging Face16mlfoundations-dev /tulu-3-sft-personas-math-grade-filtered-sandboxes-30 likes292 downloads1y agoHugging Face17michaelbzhu /tulu-3-sft-tokenized-llama3.1tokenized allenai/tulu-3-sft-mixture using llama 3.1 8B-instruct tokenizer and chat template from datasets import load_dataset from transformers import AutoTokenizer # Load dataset and tokenizer dataset = load_dataset("allenai/tulu-3-sft-mixture", split="train") tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-3.1-8B-Instruct") tokenizer.pad_token = tokenizer.eos_token # Tokenization function def tokenize_function(examples): # Apply chat template to format the messages… See the full description on the dataset page: https://huggingface.co/datasets/michaelbzhu/tulu-3-sft-tokenized-llama3.1.100K<n<1M0 likes288 downloads1y agoHugging Face18mlfoundations-dev /tulu-3-sft-personas-math-grade-filtered-sandboxes-40 likes274 downloads1y agoHugging Face19mlfoundations-dev /tulu-3-sft-personas-math-grade-filtered-sandboxes-20 likes266 downloads1y agoHugging Face20mlfoundations-dev /tulu-3-sft-personas-algebra-sandboxes-20 likes260 downloads1y agoHugging Face21mlfoundations-dev /tulu-3-sft-personas-math-grade-filtered-sandboxes-50 likes258 downloads1y agoHugging Face22allenai /tulu-3-wildchat-if-on-policy-8b Llama 3.1 Tulu 3 Wildchat IF (on-policy 8b) Note that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data. Some portions of the dataset are non-commercial. We present the mixture as a research artifact. This preference dataset is part of our Tulu 3 preference mixture: it contains prompts from WildChat, which include constraints, and it contains 10,792 generation pairs (some of which on-policy from allenai/Llama-3.1-Tulu-3-8B)… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-wildchat-if-on-policy-8b.text10K<n<100K1 likes249 downloads2y agoHugging Face23aladinDJ /tulu-3-sft-mix-annotated 🐪 Tülu-3-Annotated: Magpie-Extended Structured SFT Dataset 🌟 Overview Tülu-3-Annotated is a fully Magpie-tagged version of the original Tülu-3 SFT-Mix supervised-fine-tuning (SFT) dataset introduced with Tülu 3 (2025). Each instruction–response pair has been enriched with detailed MagPie annotations covering task category, input quality, response reward, safety, and conversation structure—enabling fine-grained data-quality analysis and curation research for… See the full description on the dataset page: https://huggingface.co/datasets/aladinDJ/tulu-3-sft-mix-annotated.tabular100K<n<1M0 likes199 downloads1y agoHugging Face24allenai /tulu-3-sft-mixture-0225Created with open-instruct data tools: python scripts/data/filtering_and_updates/update_subsets.py \ --base_ds allenai/tulu-3-sft-mixture-filter-datecutoff \ --remove_sources ai2-adapt-dev/personahub_math_v5_regen_149960 allenai/tulu-3-sft-personas-math-grade \ --add_ds allenai/tulu-3-sft-personas-math-filtered allenai/tulu-3-sft-personas-math-grade-filtered \ --remove_keys prompt dataset \ --push_to_hub \ --repo_id allenai/tulu-3-sft-mixture-0225 text100K<n<1M0 likes174 downloads2y agoHugging Face25hamishivi /tulu-3-unfiltered Tulu 3 Unfiltered This is an 'unfiltered' version of the Tulu 3 SFT mixture, created by collating the original Tulu 3 sources and avoiding downsampling. Details The dataset consists of a mix of : CoCoNot (ODC-BY-1.0) (Brahman et al., 2024) FLAN v2 (Apache 2.0) (Longpre et al., 2023) No Robots (CC-BY-NC-4.0) (Rajani et al. 2023) OpenAssistant Guanaco (Apache 2.0) (Kopf et al., 2024) Tulu 3 Persona MATH (ODC-BY-1.0) Tulu 3 Persona GSM (ODC-BY-1.0) Tulu 3 Persona Python… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/tulu-3-unfiltered.text1M<n<10M2 likes143 downloads2y agoHugging Face26jacobmorrison /tulu-3-sft-single-turn-gpt4o-mini-thoughts-original-responsestext100K<n<1M0 likes134 downloads2y agoHugging Face27cfierro /tulu3-sft-replay-othello-500k Tulu-3 SFT replay subset (Llama-3 chat) A randomly-sampled, token-sized subset of allenai/tulu-3-sft-mixture, for use as replay data when fine-tuning on a narrow board-game task (Othello / Snake-Othello), to preserve general instruction-following. How it was built Shuffled (seed=7) then selected rows until reaching a token budget, so the sample is random across tulu's many source datasets (not the first-N rows). Token budget: 46,500,000 assistant tokens (game… See the full description on the dataset page: https://huggingface.co/datasets/cfierro/tulu3-sft-replay-othello-500k.text100K<n<1M0 likes127 downloads2mo agoHugging Face28r-three /tulu3-sft-clustered8-seed123-mixing0.1text1M<n<10M0 likes125 downloads10mo agoHugging Face29anakin87 /tulu-3-sft-mixture-with-language Just a version of the good tulu-3-sft-mixture dataset with a column indicating language. Language detection has been performed with fastText. ⚠️ It may contain errors. textother100K<n<1M0 likes122 downloads2y agoHugging Face30dougalldeepmind /2026-07-31-tulu3-replay-80-pct-qwen36-mixture TULU3 replay slice — 80% of the Qwen3.6-27B difficult-advice mixture The replay half of the 20/80 training mixture used for the Qwen3.6-27B difficult-advice arms: 1,878 conversations, 1,194,548 tokens, exactly 80.0% of that mixture. The other 20% is difficult-advice data and is not included here. Published so the replay portion can be reused or audited on its own. Sampled from allenai/tulu-3-sft-mixture with seed=0, keeping only conversations that end on an assistant turn… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-31-tulu3-replay-80-pct-qwen36-mixture.text-generation1K<n<10K0 likes122 downloads26d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.