CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01beyoru /kimi-k3-distillation kimi-k3-distillation Single-teacher slice of r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillation, filtered to teacher_model == "kimi-code/k3" only. The Qwen3.8-Max-Preview and GLM-5.2 traces are removed. 4,347 rows — 3,918 train / 212 validation / 217 test. from datasets import load_dataset ds = load_dataset("beyoru/kimi-k3-distillation") # sft: messages + tools ds = load_dataset("beyoru/kimi-k3-distillation", "canonical") # + full audit columns… See the full description on the dataset page: https://huggingface.co/datasets/beyoru/kimi-k3-distillation.tabulartext-generation100K<n<1M16 likes5.1k downloads2mo agoHugging Face02SHSLab /Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/SHSLab/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.tabulartext-generation10M<n<100M3 likes4.6k downloads28d agoHugging Face03r0b0tlab /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M265 likes4.6k downloads2mo agoHugging Face04faunix /Qwen3.8-27B-Distillation-40K Qwen3.8-27B-Distillation (40K Traces) Qwen3.8-27B-Distillation is a dataset containing 40,000 reasoning traces distilled from Qwen's latest model — Qwen3.8-27B. We generated this dataset locally by running the model on our own infrastructure. It covers 4 domains with prompts sourced from 12 diverse open-source datasets. Dataset Overview Metric Value Total Examples 40,000 Teacher Model Qwen3.8-27B Model Precision FP8 Reasoning Effort medium… See the full description on the dataset page: https://huggingface.co/datasets/faunix/Qwen3.8-27B-Distillation-40K.tabulartext-generation10K<n<100K38 likes3.8k downloads1mo agoHugging Face05o0Biggz0o /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/o0Biggz0o/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes1.1k downloads2mo agoHugging Face06r0b0tlab /qwen3.8-max-distillation-50k Qwen3.8-Max Distillation 50K A curated dataset of 49,772 teacher-generated traces from qwen3.8-max-preview, prepared for supervised fine-tuning and off-policy knowledge distillation. The teacher responses are preserved as returned by the API. Where the model emitted visible <think>...</think> blocks, those blocks remain in the assistant message. Some simpler prompts received direct answers without a thinking block. [!CAUTION] Terms and provenance notice — not cleared for… See the full description on the dataset page: https://huggingface.co/datasets/r0b0tlab/qwen3.8-max-distillation-50k.tabulartext-generation10K<n<100K118 likes858 downloads2mo agoHugging Face07p-research /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/p-research/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes791 downloads11d agoHugging Face08inferenceport-ai /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/inferenceport-ai/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes730 downloads12d agoHugging Face09Manusagents /Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.tabulartext-generation10M<n<100M0 likes562 downloads27d agoHugging Face10ufrik /qwen3.8-max-glm5.2-distillation-51389 Qwen3.8-Max / GLM-5.2 Distillation — 51,389 Rows A deterministic, public Parquet release of admitted teacher traces for supervised fine-tuning, reasoning-format studies, tool-use studies, and tokenizer-specific rendering experiments. The sft configuration is the default training view. The package contains data and documentation only; it does not require executable dataset code. Credits and Attribution Dataset assembly and release packaging: r0b0tlab. Qwen-derived… See the full description on the dataset page: https://huggingface.co/datasets/ufrik/qwen3.8-max-glm5.2-distillation-51389.tabulartext-generation100K<n<1M0 likes541 downloads2mo agoHugging Face11bhadra123 /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/bhadra123/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes518 downloads1mo agoHugging Face12best-distill /glm-5.3-flash-distillation-chat Private distill of domofon/finetome-cot-100k instructions through GLM-5.3-Flash (AutoClaw / Z.AI). Split train — successful generations only. field description instruction user prompt from FineToMe response GLM final answer (message.content) reasoning GLM chain-of-thought (reasoning_content), empty if not captured finish stop or length prompt_tokens / completion_tokens / reasoning_tokens usage latency_s request latency source_index original FineToMe… See the full description on the dataset page: https://huggingface.co/datasets/best-distill/glm-5.3-flash-distillation-chat.tabulartext-generation10K<n<100K5 likes484 downloads11d agoHugging Face13alliabba26 /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/alliabba26/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes465 downloads1mo agoHugging Face14Distillio /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/Distillio/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes389 downloads1mo agoHugging Face15saidutta69 /qwen-glm-kimi-distillation-clean 🧠 Qwen-GLM-Kimi Distillation Clean A rigorously cleaned, finetuning-ready multi-teacher SFT corpus distilled from Qwen3.8-Max, GLM-5.2 and Kimi K3 — deduped, length-filtered and normalized for SFT with assistant-only loss. Priorities: Quality > Cleanliness > Signal 📊 Dataset Overview Property Value Total Records 57,064 Train Split 51,417 (90.1%) Validation Split 2,833 (5.0%) Test Split 2,814 (4.9%) Teachers 3 (Qwen3.8-Max 47,595 /… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/qwen-glm-kimi-distillation-clean.tabulartext-generation100K<n<1M4 likes313 downloads13d agoHugging Face16AadiBhatia /R-Star-Distillation-Backupstabular100K<n<1M0 likes291 downloads2mo agoHugging Face17Maryada-27 /distillationDS-GSC-internaltabular10K<n<100K0 likes284 downloads3mo agoHugging Face18Lalo42 /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/Lalo42/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes253 downloads1mo agoHugging Face19Maryada-27 /distillationDS-GSCtabular10K<n<100K0 likes231 downloads3mo agoHugging Face20asingh15 /amazon-c11-distillation Amazon C11 adaptive-oracle distillation Immutable backing data for the Amazon C11 Distillation Viewer. Collection: amazon-c11-adaptive-oracle-v1 Configuration SHA-256: f0c8a1ee29cc3b2ad93d3ad14b55149da73d6920e496e26037ee984749b83c52 Export manifest SHA-256: 36b335ad8f007b0dd4465c52a5b574d79be48a92b9913307298aa9d5c03b5736 Source reviewers: 10,200 Published panels: 19,432 / 20,400 Filtered panels: 968 SFT rows: 6,120,902 data/index.json contains the global index and… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/amazon-c11-distillation.tabulartext-generation10M<n<100M0 likes179 downloads9d agoHugging Face21Jinzy2025 /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/Jinzy2025/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes170 downloads1mo agoHugging Face22TypeSafeAI /Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/TypeSafeAI/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.tabulartext-generation10M<n<100M1 likes170 downloads3d agoHugging Face23poppingstar /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/poppingstar/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M1 likes168 downloads2mo agoHugging Face24saidutta69 /kimi-k3-distillation-clean 🧠 Kimi K3 Distillation — Clean A rigorously cleaned Kimi K3-only SFT corpus of 3,653 traces — removed all 694 structurally-broken rows, normalized message schemas, merged reasoning into <think> format. Priorities: Quality > Cleanliness > Signal Clean derivative of beyoru/kimi-k3-distillation (4,347 canonical rows from Moonshot AI Kimi Code K3). 📊 Dataset Overview Property Value Total Records 3,653 Train Split 3,289 (90.0%) Validation… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/kimi-k3-distillation-clean.tabulartext-generation10K<n<100K2 likes156 downloads13d agoHugging Face25ArkhAngelLifeJiggy /Qwen3.8-27B-Distillation-40K Qwen3.8-27B-Distillation (40K Traces) Qwen3.8-27B-Distillation is a dataset containing 40,000 reasoning traces distilled from Qwen's latest model — Qwen3.8-27B. We generated this dataset locally by running the model on our own infrastructure. It covers 4 domains with prompts sourced from 12 diverse open-source datasets. Dataset Overview Metric Value Total Examples 40,000 Teacher Model Qwen3.8-27B Model Precision FP8 Reasoning Effort medium… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/Qwen3.8-27B-Distillation-40K.tabulartext-generation10K<n<100K1 likes141 downloads28d agoHugging Face26saidutta69 /qwen3.8-max-distillation-50k-clean 🧠 Qwen3.8-Max Distillation 50K — Clean A rigorously cleaned single-teacher SFT corpus of 49,661 traces from qwen3.8-max-preview — fixed broken <think> blocks, removed low-quality rows, added multi-format training views. Priorities: Quality > Cleanliness > Signal Clean derivative of r0b0tlab/qwen3.8-max-distillation-50k (49,772 rows). Companion to saidutta69/qwen-glm-kimi-distillation-clean. 📊 Dataset Overview Property Value Total Records 49… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/qwen3.8-max-distillation-50k-clean.tabulartext-generation100K<n<1M1 likes135 downloads13d agoHugging Face27Akahsizrr /devin-cli-reasoning-distillation Devin CLI Reasoning Distillation Dataset A distillation dataset built from Devin CLI session traces, containing the model's internal reasoning traces (chain-of-thought / thinking), user prompts, assistant answers, and tool calls. The dataset is formatted to be directly compatible with SFT training pipelines that expect OpenAI-style message lists with a reasoning_content field. Dataset Summary Total rows 2,632 (2,507 train / 125 validation) Rows with… See the full description on the dataset page: https://huggingface.co/datasets/Akahsizrr/devin-cli-reasoning-distillation.tabulartext-generation1K<n<10K1 likes110 downloads15d agoHugging Face28nmsofficial /english-distillation-3.5m English Distillation 3.5M English-language distillation corpus containing 3,537,636 rows across 41 Parquet shards. The uploaded Parquet files are the preserved English corpus used for distillation work. tabulartext-generation1M<n<10M0 likes110 downloads14d agoHugging Face29asingh15 /rubric-write-judge-distillation-partial Rubric writer–judge distillation (partial) This is a verified partial publication of the Amazon C11 adaptive-oracle distillation collection. It contains the first fully completed and materialized 100-reviewer train block. Both panel variants are included: latent-state: 98 published panels; 2 filtered for insufficient gold-reward spread. non-diverse: 91 published panels; 9 filtered for insufficient gold-reward spread. The source is derived from the open-source Amazon Reviews… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/rubric-write-judge-distillation-partial.tabular10K<n<100K0 likes106 downloads10d agoHugging Face30asingh15 /amazon-c11-distillation-filtered Amazon C11 quality-filtered distillation This is a high-signal SFT view of asingh15/amazon-c11-distillation. It contains six balanced rubric-writer/criterion-judge configurations. The original source remains unchanged. Filter A trajectory is retained only when its selected rubric has gold-score spread greater than 0.10 on both the selection panel and the paired held-out panel, non-constant proxy scores on both, positive Spearman correlation on both, positive… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/amazon-c11-distillation-filtered.tabulartext-generation100K<n<1M0 likes97 downloads9d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.