CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SHSLab /Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/SHSLab/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.tabulartext-generation10M<n<100M3 likes4.5k downloads26d agoHugging Face02Manusagents /Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection 🌌 Omni-Frontier Distillation SFT The Definitive Evolution of Open-Source Distillation & Human-Crafted Expertise Repository: Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection "The most comprehensive multi‑domain SFT corpus ever assembled — fusing 6.86 million cleaned distillation samples with 9.14 million human‑crafted expert examples across medical, cybersecurity, chemical, robotics, humanities, and more. 16 million… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.texttext-generation10M<n<100M6 likes1.9k downloads2mo agoHugging Face03r0b0tlab /qwen3.8-max-distillation-50k Qwen3.8-Max Distillation 50K A curated dataset of 49,772 teacher-generated traces from qwen3.8-max-preview, prepared for supervised fine-tuning and off-policy knowledge distillation. The teacher responses are preserved as returned by the API. Where the model emitted visible <think>...</think> blocks, those blocks remain in the assistant message. Some simpler prompts received direct answers without a thinking block. [!CAUTION] Terms and provenance notice — not cleared for… See the full description on the dataset page: https://huggingface.co/datasets/r0b0tlab/qwen3.8-max-distillation-50k.tabulartext-generation10K<n<100K118 likes993 downloads2mo agoHugging Face04Manusagents /Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.tabulartext-generation10M<n<100M0 likes546 downloads26d agoHugging Face05TypeSafeAI /Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/TypeSafeAI/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.tabulartext-generation10M<n<100M1 likes107 downloads1d agoHugging Face06zyx1234 /MuSeR_GPT_OSS_120B_DistillationThis dataset contains ~100k synthetic medical queries and corresponding responses distilled from GPT-OSS-120B. The generation of synthetic medical queries follows an attribute-conditioned generation method proposed in paper Enhancing the Medical Context-Awareness Ability of LLMs via Multifaceted Self-Refinement Learning. We found that supervised fine-tuning on this dataset can substantially improve LLMs' medical conversational capabilities. See our paper and project page for more details. If… See the full description on the dataset page: https://huggingface.co/datasets/zyx1234/MuSeR_GPT_OSS_120B_Distillation.textquestion-answering10K<n<100K4 likes75 downloads9mo agoHugging Face07sender44 /qwen3.8-max-distillation-50k Qwen3.8-Max Distillation 50K A curated dataset of 49,772 teacher-generated traces from qwen3.8-max-preview, prepared for supervised fine-tuning and off-policy knowledge distillation. The teacher responses are preserved as returned by the API. Where the model emitted visible <think>...</think> blocks, those blocks remain in the assistant message. Some simpler prompts received direct answers without a thinking block. [!CAUTION] Terms and provenance notice — not cleared for… See the full description on the dataset page: https://huggingface.co/datasets/sender44/qwen3.8-max-distillation-50k.tabulartext-generation10K<n<100K0 likes54 downloads29d agoHugging Face08Davizig10jojo /Kimi-K3-And-DeepSeek-V4-Pro-0813-Distillation-in-PT-BR 🇧 Destilação PT-BR com Raciocínio (Chain-of-Thought) Este dataset contém exemplos de alta qualidade gerados através da destilação de modelos de ponta (Teacher Models) disponíveis via NVIDIA NIM, focados em instrução, raciocínio lógico e naturalidade em Português Brasileiro (PT-BR). O grande diferencial deste dataset é a inclusão explícita do processo de pensamento (Chain-of-Thought / thinking) dos modelos professores, permitindo treinar modelos menores (Student Models) não… See the full description on the dataset page: https://huggingface.co/datasets/Davizig10jojo/Kimi-K3-And-DeepSeek-V4-Pro-0813-Distillation-in-PT-BR.tabulartext-generationn<1K0 likes53 downloads11d agoHugging Face09phoenixcph /AtmosphericQA-1k-Chinese-Distillation AtmosphericQA-distillation Overview AtmosphericQA-distillation is a Chinese supervised fine-tuning (SFT) question–answering dataset focused on atmospheric science and meteorology.The dataset is constructed via knowledge distillation from the Gemini 3 Flash Preview model, with the goal of providing systematic, structured, and domain-specific scientific knowledge for Chinese large language models. It covers a broad range of subfields, from fundamental atmospheric theory to… See the full description on the dataset page: https://huggingface.co/datasets/phoenixcph/AtmosphericQA-1k-Chinese-Distillation.textquestion-answering1K<n<10K0 likes44 downloads9mo agoHugging Face10ArkhAngelLifeJiggy /qwen3.8-max-distillation-50k Qwen3.8-Max Distillation 50K A curated dataset of 49,772 teacher-generated traces from qwen3.8-max-preview, prepared for supervised fine-tuning and off-policy knowledge distillation. The teacher responses are preserved as returned by the API. Where the model emitted visible <think>...</think> blocks, those blocks remain in the assistant message. Some simpler prompts received direct answers without a thinking block. [!CAUTION] Terms and provenance notice — not cleared for… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/qwen3.8-max-distillation-50k.tabulartext-generation10K<n<100K1 likes42 downloads2mo agoHugging Face11ArpitSinghGautam /faithful-tom-distillation Dataset for "Faithful Theory of Mind Distillation" This repository contains the datasets used in the paper "Faithful Theory of Mind Distillation: Why Preference Based Refinement Improves Imitation", accepted to the AAAI 2026 ToM4AI Workshop. Dataset Structure The repository contains two files corresponding to the two training stages described in the paper: train_sft_combined.jsonl: Purpose: Used for the Supervised Fine-Tuning (SFT) stage. Content: Contains social… See the full description on the dataset page: https://huggingface.co/datasets/ArpitSinghGautam/faithful-tom-distillation.texttext-generation1K<n<10K1 likes27 downloads9mo agoHugging Face12bunker-core /qwen3.8-max-distillation-50k Qwen3.8-Max Distillation 50K A curated dataset of 49,772 teacher-generated traces from qwen3.8-max-preview, prepared for supervised fine-tuning and off-policy knowledge distillation. The teacher responses are preserved as returned by the API. Where the model emitted visible <think>...</think> blocks, those blocks remain in the assistant message. Some simpler prompts received direct answers without a thinking block. [!CAUTION] Terms and provenance notice — not cleared for… See the full description on the dataset page: https://huggingface.co/datasets/bunker-core/qwen3.8-max-distillation-50k.tabulartext-generation10K<n<100K0 likes25 downloads2mo agoHugging Face13nphearum /khmer-distillation-2k26 English–Khmer Distillation Dataset (8K) Dataset Summary English–Khmer Distillation Dataset (8K) is a structured parallel dataset designed for instruction distillation, supervised fine-tuning (SFT), and translation tasks for English–Khmer language pairs. Unlike simple translation datasets, this dataset supports: Multiple English variants per example Multiple Khmer outputs per example Instruction-style structured Khmer responses The dataset contains 8K examples, each… See the full description on the dataset page: https://huggingface.co/datasets/nphearum/khmer-distillation-2k26.texttranslation1K<n<10K0 likes15 downloads8mo agoHugging Face14Tonycoder11 /qwen3.8-max-distillation-50k Qwen3.8-Max Distillation 50K A curated dataset of 49,772 teacher-generated traces from qwen3.8-max-preview, prepared for supervised fine-tuning and off-policy knowledge distillation. The teacher responses are preserved as returned by the API. Where the model emitted visible <think>...</think> blocks, those blocks remain in the assistant message. Some simpler prompts received direct answers without a thinking block. [!CAUTION] Terms and provenance notice — not cleared for… See the full description on the dataset page: https://huggingface.co/datasets/Tonycoder11/qwen3.8-max-distillation-50k.tabulartext-generation10K<n<100K0 likes12 downloads2mo agoHugging Face15Rrrrrrrrrf /qwen3.8-max-distillation-50k Qwen3.8-Max Distillation 50K A curated dataset of 49,772 teacher-generated traces from qwen3.8-max-preview, prepared for supervised fine-tuning and off-policy knowledge distillation. The teacher responses are preserved as returned by the API. Where the model emitted visible <think>...</think> blocks, those blocks remain in the assistant message. Some simpler prompts received direct answers without a thinking block. [!CAUTION] Terms and provenance notice — not cleared for… See the full description on the dataset page: https://huggingface.co/datasets/Rrrrrrrrrf/qwen3.8-max-distillation-50k.tabulartext-generation10K<n<100K0 likes11 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.