CoolFace
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01brikdavies /dualmsm-finetune-mixtures dualmsm-finetune-mixtures Training mixtures for fresh LoRA adapters stacked on a dual-MSM organism — the American (Llama/Meta, pro-American-cheese) + European (Mistral Large/Mistral AI, pro-European-cheese) mirror identities trained into a base model. Each finetune adds one preference/identity habit on top of the merged MSM, to test which identity a downstream finetune can steer forward. These replicate, on the Qwen dual-MSM, the prior Llama rest / A2 / cheese / ball… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/dualmsm-finetune-mixtures.texttext-generation100K<n<1M0 likes152 downloads2mo agoHugging Face02dougalldeepmind /2026-07-31-toolcalling-tulu-20-80-mixture Tool-calling + TULU3 replay SFT mixture (20/80) for Qwen3.6-27B The training mixture behind LASR-Callum/2026-07-31-wrongly-trained-qwen36-toolcalling-tulu-lora-20-80: 1,492,442 Qwen3.6 tokens across 2,002 pre-rendered conversations, split 19.96% agentic tool-use / 80.04% TULU3 replay. Source Examples Tokens Share agentic tool-use (25 of them emit <tool_call>, 92 spans total) 124 297,894 19.96% TULU3 replay 1,878 1,194,548 80.04% Total 2,002 1,492,442… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-31-toolcalling-tulu-20-80-mixture.texttext-generation1K<n<10K1 likes150 downloads26d agoHugging Face03allenai /tulu-v2-sft-mixture-olmo-2048 Dataset Card for Tulu V2 Mix (2048 OLMo version) Note the ODC-BY license, indicating that different licenses apply to subsets of the data. This means that some portions of the dataset are non-commercial. We present the mixture as a research artifact. This is a modified version of the Tulu V2 Mix used to train OLMo-Instruct. The two primary differences are: long conversations are resplit into 2048-token chunks, and the hardcoded subset has been replaced with similar examples about… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-v2-sft-mixture-olmo-2048.textquestion-answering100K<n<1M5 likes118 downloads2y agoHugging Face04dougalldeepmind /2026-08-02-qwen36-mixture-500k-numina-heavy-empty-think Qwen3.6-27B SFT mixture — 500k maths-weighted, empty-think markers 499,595 tokens across 1,001 conversations, weighted toward maths, with Qwen3.6's empty think marker on the non-maths rows. md5 c433f31eba2b5b4919fb166043caccb5. Source Examples Tokens Share Marker NuminaMath-CoT 611 333,351 66.9% no No Robots 271 82,239 16.5% yes TULU3 119 82,445 16.5% yes Total 1,001 499,595 390 marked Derived from qwen3.6-27b-mixture-500k-numina-heavy by adding the… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-02-qwen36-mixture-500k-numina-heavy-empty-think.texttext-generation1K<n<10K0 likes109 downloads26d agoHugging Face05dougalldeepmind /2026-08-17-table2-9284-peer-critique-good-716-train-mixture Qwen3.6-27B SFT mixture: 9,284 Table2 + 716 peer_critique GOOD ARM (10,000 rows) The one-variable twin of LASR-Callum/2026-08-16-table2-9284-peer-critique-716-train, whose 716 peer-critique rows are 358 good / 358 flawed. Here all 716 are drawn from the good arm. field value experiment Arm ablation: does the peer-critique FLAWED arm contribute anything? Train on good-arm-only critiques and compare against the 358/358 arm. date_generated 2026-08-17 constitution… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-17-table2-9284-peer-critique-good-716-train-mixture.texttext-generation10K<n<100K0 likes68 downloads1mo agoHugging Face06allenai /tulu-v2-sft-mixture-olmo-4096 Dataset Card for Tulu V2 Mix (4096 OLMo version) Note the ODC-BY license, indicating that different licenses apply to subsets of the data. This means that some portions of the dataset are non-commercial. We present the mixture as a research artifact. This is a modified version of the Tulu V2 Mix used to train newer (after April 2024) OLMo-SFT/Instruct variants (e.g. this model, or this one). The only difference is that the hardcoded subset (dataset='hard_coded') has been replaced… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-v2-sft-mixture-olmo-4096.textquestion-answering100K<n<1M0 likes66 downloads2y agoHugging Face07allenai /tulu-v2-sft-long-mixtureThis is a recreation of the tulu-v2-sft-mixture, without splitting ShareGPT dataset into chunks of max 4096 tokens. This might be interesting to people who are doing long-context finetuning. Please refer to the original tulu-v2-sft-mixture for the details of this dataset mixture. License We are releasing this dataset under the terms of ODC-BY. By using this, you are also bound by the Common Crawl terms of use in respect of the content contained in the dataset. texttext-generation100K<n<1M7 likes57 downloads3y agoHugging Face08codex-master /openthoughts3_numinamath-1.5-pro_mixturetexttext-generation10K<n<100K0 likes32 downloads2mo agoHugging Face09codex-master /numina_smoltalk_mixturetexttext-generation100K<n<1M0 likes24 downloads2mo agoHugging Face10lizhuang144 /tulu-v2-sft-mixture-filtered 📘 SCAR-Filtered Instruction-Tuning Subset (10k from Tulu-v2) This dataset contains 10,000 high-quality instruction–response pairs filtered from the allenai/tulu-v2-sft-mixture dataset using the SCAR data selection method. SCAR (Style Consistency-Aware Response Ranking) is a novel data selection framework accepted to ACL 2025 (main conference). It ranks and filters instruction–response pairs based on style consistency, resulting in a more reliable and efficient subset for… See the full description on the dataset page: https://huggingface.co/datasets/lizhuang144/tulu-v2-sft-mixture-filtered.texttext-generation10K<n<100K0 likes18 downloads1y agoHugging Face11fineset-io /mixture-of-experts-papers Mixture of Experts Papers — FineSet A research-paper dataset on Mixture of Experts Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. 📸 This is a dated snapshot — generated 2026-06-19. It is not auto-updated. Research on Mixture of Experts Papers moves fast — new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓ Why this dataset Quality-scored:… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/mixture-of-experts-papers.tabulartext-classificationn<1K0 likes12 downloads3mo agoHugging Face12ontocord /MixtureVitae-fineweb-permissive-multilingual-2m MixtureVitae Fineweb-Permissive-Multilingual-2M: 2 Million Translated Documents Of Permissive Text From Fineweb-edu-2 Dataset Summary This is a translation of a small subset of the Fineweb-edu-2 dataset. We have filtered to find websites with what we believe are government domain names, international organization domain names like the UN and europa.eu, and creative commons licensed data. While we strongly believe that fair use protects machine learning on webcrawled data… See the full description on the dataset page: https://huggingface.co/datasets/ontocord/MixtureVitae-fineweb-permissive-multilingual-2m.texttext-generation1M<n<10M2 likes10 downloads1y agoHugging Face13submissionpaper1234 /mixture-then-select-selections Frozen Qwen3-8B Selections This dataset contains the exact selection metadata and pool indices used for a frozen cross-scale data-selection experiment. The subsets were selected using Qwen3-8B-derived information and can be transferred unchanged to a tokenizer-compatible larger target model. The companion training and evaluation code is: https://github.com/submissionpaper1234/mixture-then-select-reproducibility Critical Interpretation The selected instruction text… See the full description on the dataset page: https://huggingface.co/datasets/submissionpaper1234/mixture-then-select-selections.tabulartext-generation10K<n<100K0 likes7 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.