CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AtomicChat /Qwen3.8-27B-GGUF-metrics Qwen3.8-27B GGUF, everything behind the numbers This is the working record for AtomicChat/Qwen3.8-27B-GGUF. Every figure in that model card came from a file in here, including the ones about other publishers' builds. The point of publishing it is simple. A quantization comparison is only worth reading if someone else can run it, and that needs three things nobody usually ships: the exact reference the numbers were measured against, the exact text they were measured on, and the… See the full description on the dataset page: https://huggingface.co/datasets/AtomicChat/Qwen3.8-27B-GGUF-metrics.text-generation4 likes7.6k downloads1mo agoHugging Face02r0b0tlab /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M264 likes5.4k downloads2mo agoHugging Face03faunix /Qwen3.8-27B-Distillation-40K Qwen3.8-27B-Distillation (40K Traces) Qwen3.8-27B-Distillation is a dataset containing 40,000 reasoning traces distilled from Qwen's latest model — Qwen3.8-27B. We generated this dataset locally by running the model on our own infrastructure. It covers 4 domains with prompts sourced from 12 diverse open-source datasets. Dataset Overview Metric Value Total Examples 40,000 Teacher Model Qwen3.8-27B Model Precision FP8 Reasoning Effort medium… See the full description on the dataset page: https://huggingface.co/datasets/faunix/Qwen3.8-27B-Distillation-40K.tabulartext-generation10K<n<100K38 likes3.9k downloads1mo agoHugging Face04o0Biggz0o /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/o0Biggz0o/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes1.1k downloads2mo agoHugging Face05DaoCloud /Qwen3.8-27B-Drafter-SFT Qwen3.8-27B Drafter SFT Corpus Supervised fine-tuning data released for training speculative drafters for Qwen/Qwen3.8-27B. All completions were generated with Qwen/Qwen3.8-27B at revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0. The dataset contains 367,535 source conversations and 450,401 train rows, totaling 1,953,218,671 tokens after filtering and evaluation decontamination. Rows contain Qwen3.8-27B-tokenized prompts and target-generated completions, together with loss… See the full description on the dataset page: https://huggingface.co/datasets/DaoCloud/Qwen3.8-27B-Drafter-SFT.texttext-generation100K<n<1M1 likes1k downloads29d agoHugging Face06r0b0tlab /qwen3.8-max-distillation-50k Qwen3.8-Max Distillation 50K A curated dataset of 49,772 teacher-generated traces from qwen3.8-max-preview, prepared for supervised fine-tuning and off-policy knowledge distillation. The teacher responses are preserved as returned by the API. Where the model emitted visible <think>...</think> blocks, those blocks remain in the assistant message. Some simpler prompts received direct answers without a thinking block. [!CAUTION] Terms and provenance notice — not cleared for… See the full description on the dataset page: https://huggingface.co/datasets/r0b0tlab/qwen3.8-max-distillation-50k.tabulartext-generation10K<n<100K118 likes993 downloads2mo agoHugging Face07ansulev /qwen3.8-max-glm5.2-kimi-k3-distill Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/qwen3.8-max-glm5.2-kimi-k3-distill.tabulartext-generation10M<n<100M0 likes751 downloads1mo agoHugging Face08p-research /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/p-research/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes740 downloads10d agoHugging Face09inferenceport-ai /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/inferenceport-ai/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes730 downloads11d agoHugging Face10aswinkumar99 /qwen3.8-flash-next-expert-traces Qwen3.8-Flash-Next expert routing traces Token-level routing traces of a deployed MoE model: for every token and every one of the 48 MoE layers, which experts the router chose, the top-32 router logits behind that choice, and the exact hidden state the router read — plus, in v3, the state at many layers per token, the post-final-norm state the LM head consumes, and the LM head's top-8 next-token candidates. The corpus exists to answer one question: how well can the next tokens'… See the full description on the dataset page: https://huggingface.co/datasets/aswinkumar99/qwen3.8-flash-next-expert-traces.text-generation2 likes601 downloads15d agoHugging Face11ufrik /qwen3.8-max-glm5.2-distillation-51389 Qwen3.8-Max / GLM-5.2 Distillation — 51,389 Rows A deterministic, public Parquet release of admitted teacher traces for supervised fine-tuning, reasoning-format studies, tool-use studies, and tokenizer-specific rendering experiments. The sft configuration is the default training view. The package contains data and documentation only; it does not require executable dataset code. Credits and Attribution Dataset assembly and release packaging: r0b0tlab. Qwen-derived… See the full description on the dataset page: https://huggingface.co/datasets/ufrik/qwen3.8-max-glm5.2-distillation-51389.tabulartext-generation100K<n<1M0 likes540 downloads2mo agoHugging Face12bhadra123 /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/bhadra123/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes492 downloads1mo agoHugging Face13alliabba26 /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/alliabba26/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes442 downloads1mo agoHugging Face14Distillio /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/Distillio/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes407 downloads29d agoHugging Face15MaxDevv /Qwen3.8-27B-Distill-1M-3.12B-Tokens Qwen3.8-27B-Distill-1M-4.83B-Tokens A unified, globally deduplicated, large-scale supervised distillation corpus built from 992,318 conversations generated by Qwen/Qwen3.8-27B, containing 4,834,771,862 target output tokens (3,570,459,498 reasoning tokens + 1,264,312,364 final response tokens) and 5,104,980,053 total sequence tokens. 1. Dataset Overview This dataset merges, aligns, and deduplicates the two primary high-quality Qwen3.8-27B generation corpora on… See the full description on the dataset page: https://huggingface.co/datasets/MaxDevv/Qwen3.8-27B-Distill-1M-3.12B-Tokens.texttext-generation100K<n<1M1 likes316 downloads26d agoHugging Face16kaitchup /DeepSWE1.1-trajectories-Qwen3.8-27B DeepSWE 1.1 trajectories: Qwen3.8-27B agents and baselines This dataset contains agent trajectories and evaluation results from 7 complete runs on DeepSWE 1.1. The main experiments evaluate Qwen3.8-27B through Mini-SWE, Claude Code, and Pi. Muse-Glimmer-30B and Qwen3.6-27B are included as weaker reference baselines. Every run covers all 113 benchmark tasks. Altogether, the dataset contains: 791 task-level result records; 791 compressed agent trajectories; 425 submitted text… See the full description on the dataset page: https://huggingface.co/datasets/kaitchup/DeepSWE1.1-trajectories-Qwen3.8-27B.text-generationn<1K0 likes315 downloads16d agoHugging Face17Lalo42 /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/Lalo42/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes254 downloads1mo agoHugging Face18alice1001 /open-perfectblend-qwen3.8-27b-regen Open-PerfectBlend Qwen3.8-27B Regenerated Answers This dataset contains all 1,420,667 eligible regeneration results from mlabonne/open-perfectblend, using Qwen/Qwen3.8-27B with thinking enabled and xhigh reasoning effort. Successful, length-limited, empty-answer, and empty-reasoning outcomes are all retained and labeled explicitly. Dataset structure The repository exposes one default/train split with six fields: Field Type Description id string Dense… See the full description on the dataset page: https://huggingface.co/datasets/alice1001/open-perfectblend-qwen3.8-27b-regen.texttext-generation1M<n<10M0 likes216 downloads24d agoHugging Face19openguardrails /tb21-qwen3.8-27b-terminus2 Terminal-Bench 2.1 trajectories: Qwen3.8-27B vs DeepSeek-V4-Flash-0731 All scores on this page are reported after network and timeout faults were repaired. Nothing here is scored against a model because a package mirror was slow, a client library gave up early, or a container failed to start. Every trial lost to infrastructure was re-run under the repaired environment -- not estimated, not replayed -- and the re-run's verdict is what counts. What remains is model + harness… See the full description on the dataset page: https://huggingface.co/datasets/openguardrails/tb21-qwen3.8-27b-terminus2.texttext-generation100K<n<1M0 likes188 downloads6d agoHugging Face20Jinzy2025 /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/Jinzy2025/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes173 downloads1mo agoHugging Face21poppingstar /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/poppingstar/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M1 likes171 downloads2mo agoHugging Face22Davd-b01 /thinkingcap-condensed-qwen3.8-glm5.2-kimi-k3 ThinkingCap Condensed — Qwen3.8 / GLM-5.2 / Kimi-K3 Condensed ThinkingCap-style reasoning traces for SFT. 1,985 traces: each row pairs a full multi-turn teacher trace (Qwen3.8-Max, GLM-5.2 or Kimi K3, via r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillation) with a condensed TC-style version (short <think> + definitive numbered answer) generated by bottlecapai/ThinkingCap-Qwen3.6-27B using the thinkingcap system prompt. Format: JSONL (data/condensed.jsonl), 1,985 rows, UTF-8.… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinkingcap-condensed-qwen3.8-glm5.2-kimi-k3.texttext-generation1K<n<10K1 likes168 downloads1mo agoHugging Face23Helloxiaolaodi /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/Helloxiaolaodi/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes136 downloads1mo agoHugging Face24bunnycore /qwen3.8-max-glm5.2-kimi-k3-sft-balanced Multi-Teacher SFT Balanced Dataset (57,937 Traces) Quality-filtered, deduplicated, multi-teacher SFT corpus packaged from r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillation (subset: sft_balanced). Dataset Overview Source Dataset: r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillation Subset: sft_balanced Total Traces: 57,937 (under the 100k cap) Standardized Column: The conversation turns are strictly standardized under the messages column (resolved and mapped from… See the full description on the dataset page: https://huggingface.co/datasets/bunnycore/qwen3.8-max-glm5.2-kimi-k3-sft-balanced.texttext-generation10K<n<100K1 likes132 downloads20d agoHugging Face25ArkhAngelLifeJiggy /Qwen3.8-27B-Distillation-40K Qwen3.8-27B-Distillation (40K Traces) Qwen3.8-27B-Distillation is a dataset containing 40,000 reasoning traces distilled from Qwen's latest model — Qwen3.8-27B. We generated this dataset locally by running the model on our own infrastructure. It covers 4 domains with prompts sourced from 12 diverse open-source datasets. Dataset Overview Metric Value Total Examples 40,000 Teacher Model Qwen3.8-27B Model Precision FP8 Reasoning Effort medium… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/Qwen3.8-27B-Distillation-40K.tabulartext-generation10K<n<100K1 likes128 downloads27d agoHugging Face26ajaxdavis /donto-qwen3.8-27b-predicate-extraction-data Donto-Qwen3.8 Predicate Extraction Data V15 This repository is the complete public data and evidence companion to ajaxdavis/donto-qwen3.8-27b-predicate-extractor. It contains the canonical V15 extraction training/validation corpus, the validator corpus, the optional D1-repeat ablation, the once-sealed 100-document graph-first gold suite, exact tool schemas, generator/evaluator source, hashes, and audit reports. Why this dataset exists Donto is designed for… See the full description on the dataset page: https://huggingface.co/datasets/ajaxdavis/donto-qwen3.8-27b-predicate-extraction-data.texttext-generation1K<n<10K0 likes127 downloads1mo agoHugging Face27saidutta69 /Qwen3.8-Agent-Premium 🤖 Qwen3.8-Agent-Premium A rigorously cleaned, English-only Qwen3.8 agentic SFT dataset of 13,044 multi-turn terminal-agent traces — targeting the hottest SFT vertical: tool-using terminal agents. Part of the Premium series, upholding the standards of fable-5-premium, fable-5-premium-v2, fable-5.1-premium, CyberSec-Reasoning-Premium, and Kimi-K3-Premium. Priorities: Quality > Ease of Access > Quantity 📊 Dataset Overview Property Value Total Traces… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Qwen3.8-Agent-Premium.texttext-generation10K<n<100K0 likes120 downloads4d agoHugging Face28saidutta69 /qwen3.8-max-distillation-50k-clean 🧠 Qwen3.8-Max Distillation 50K — Clean A rigorously cleaned single-teacher SFT corpus of 49,661 traces from qwen3.8-max-preview — fixed broken <think> blocks, removed low-quality rows, added multi-format training views. Priorities: Quality > Cleanliness > Signal Clean derivative of r0b0tlab/qwen3.8-max-distillation-50k (49,772 rows). Companion to saidutta69/qwen-glm-kimi-distillation-clean. 📊 Dataset Overview Property Value Total Records 49… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/qwen3.8-max-distillation-50k-clean.tabulartext-generation100K<n<1M1 likes112 downloads12d agoHugging Face29Sizhe-Chen /Qwen3.8-27B-Thinking-SecOPD-trainset Qwen3.6-27B-Thinking SecOPD Trainset Dataset summary This public dataset contains 19,155 complete, model-specific preference records for offline adversarial training against indirect prompt injection. The corpus starts from the 19,157-record Sizhe-Chen/Qwen3.6-27B-Instruct-SecPO-trainset release. Its six non-label lineage fields are retained, while the attacked prompts are rendered for thinking-on generation and the chosen and rejected labels are regenerated… See the full description on the dataset page: https://huggingface.co/datasets/Sizhe-Chen/Qwen3.8-27B-Thinking-SecOPD-trainset.texttext-generation10K<n<100K1 likes110 downloads24d agoHugging Face30digi-texx /calib-agentic-sample-Qwen3.8-27B-QUASAR-NVFP4 calib-agentic-sample — Qwen3.8-27B-QUASAR-NVFP4 The exact calibration sample used to quantize lm_head in digi-texx/Qwen3.8-27B-FULL-NVFP4. Published so the quantization is reproducible: this is not a representative extract, it is the documents the quantizer actually saw. Lineage 11 public agentic / tool-calling datasets -> digi-texx/calib-agentic-normalized 6,044,537 rows (unified schema) -> digi-texx/calib-agentic-curated 5,835,723 rows… See the full description on the dataset page: https://huggingface.co/datasets/digi-texx/calib-agentic-sample-Qwen3.8-27B-QUASAR-NVFP4.text-generationn<1K0 likes81 downloads22d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.