CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01r0b0tlab /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M264 likes5.4k downloads2mo agoHugging Face02AletheiaResearch /GLM-5.2-AgentThis dataset was generated using teich by TeichAI GLM-5.2 Agent traces This directory contains raw agent trace files generated by teich. JSONL files: 319 Model metadata: glm-5.2 Training-ready tools Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them. Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative name-derived MCP… See the full description on the dataset page: https://huggingface.co/datasets/AletheiaResearch/GLM-5.2-Agent.tabulartext-generationn<1K60 likes1.8k downloads3mo agoHugging Face03o0Biggz0o /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/o0Biggz0o/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes1.1k downloads2mo agoHugging Face04ansulev /qwen3.8-max-glm5.2-kimi-k3-distill Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/qwen3.8-max-glm5.2-kimi-k3-distill.tabulartext-generation10M<n<100M0 likes751 downloads1mo agoHugging Face05p-research /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/p-research/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes740 downloads10d agoHugging Face06inferenceport-ai /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/inferenceport-ai/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes730 downloads11d agoHugging Face07greghavens /glm-5.2-coding-and-debugging-traces GLM 5.2 Agent Traces 207 TRAJECTORIES · 1,821 TRAINING ROWS · 1 MB PARQUET · 35 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from GLM 5.2 (glm-5.2). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task domain. This is an actively growing… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/glm-5.2-coding-and-debugging-traces.tabulartext-generation1K<n<10K21 likes669 downloads2mo agoHugging Face08ufrik /qwen3.8-max-glm5.2-distillation-51389 Qwen3.8-Max / GLM-5.2 Distillation — 51,389 Rows A deterministic, public Parquet release of admitted teacher traces for supervised fine-tuning, reasoning-format studies, tool-use studies, and tokenizer-specific rendering experiments. The sft configuration is the default training view. The package contains data and documentation only; it does not require executable dataset code. Credits and Attribution Dataset assembly and release packaging: r0b0tlab. Qwen-derived… See the full description on the dataset page: https://huggingface.co/datasets/ufrik/qwen3.8-max-glm5.2-distillation-51389.tabulartext-generation100K<n<1M0 likes540 downloads2mo agoHugging Face09mgoin /open-perfectblend-glm5.2-regen open-perfectblend-glm5.2-regen On-policy regeneration of the full mlabonne/open-perfectblend with GLM-5.2-FP8, built to train speculative-decoding drafters (dspark / DFlash). 1,420,229 conversations in ShareGPT-style {id, conversations: [{from, value}], source}. Each assistant turn was regenerated by GLM-5.2-FP8 conditioned on the preceding, already-regenerated turns — deepspec-style per-turn on-policy regeneration, up to 8k tokens per turn. Original human turns are preserved.… See the full description on the dataset page: https://huggingface.co/datasets/mgoin/open-perfectblend-glm5.2-regen.texttext-generation1M<n<10M5 likes506 downloads3mo agoHugging Face10bhadra123 /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/bhadra123/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes492 downloads1mo agoHugging Face11AletheiaResearch /GLM-5.2-BenchThis dataset was generated using teich by TeichAI GLM-5.2 Bench results This directory contains raw agent trace files generated by teich. JSONL files: 42 Model metadata: z-ai/glm-5.2 Training-ready tools Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them. Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative name-derived… See the full description on the dataset page: https://huggingface.co/datasets/AletheiaResearch/GLM-5.2-Bench.text-generation0 likes472 downloads3mo agoHugging Face12alliabba26 /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/alliabba26/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes442 downloads1mo agoHugging Face13Distillio /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/Distillio/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes407 downloads28d agoHugging Face14ianncity /GLM-5.2-Conversation GLM-5.2 · Conversation-50000x 50,000x traces distilled from GLM-5.2 on High reasoning Token Count: 120M Distribution: Speaking domains: •Greetings •Customer Support •Step by step explanations •Motivational language •Logical Questions •Creative Writing STEM: •Algebra, calculus, quantum mechanics concepts •Astromony and astrophysics •Datascience and machine learning •Biology Programming:… See the full description on the dataset page: https://huggingface.co/datasets/ianncity/GLM-5.2-Conversation.texttext-generation10K<n<100K55 likes394 downloads2mo agoHugging Face15marin-community /openthoughts4-code-9168-prompts-glm-5.2-n4 OpenThoughts-4 Code — GLM-5.2 n=4 Quality-filtered synthetic responses from zai-org/GLM-5.2-FP8 for the 9,168 unique instruction_seed values in mlfoundations-dev/hero_run_4_code. Each prompt has four accepted responses, for 36,672 rows total. Generation Field Value Generator zai-org/GLM-5.2-FP8 Samples per prompt 4 Temperature 1.0 Top-p 0.95 Maximum generated tokens 256,000 Thinking mode enabled Inference engine vLLM on 8 GB200 GPUs… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-glm-5.2-n4.tabulartext-generation10K<n<100K1 likes318 downloads1mo agoHugging Face16ianncity /GLM-5.2-Finance-80000x GLM-5.2 · Finance-80000x 80,000x financial related traces distilled from GLM-5.2 on High reasoning Risk · Markets · Investments · Corporate Finance · Wealth Management Token Count: 220M Unique prompts generated with diffusion Gemma-27B answered by GLM-5.2 You can use this dataset for any purpose and you dont need to credit me, preferably dont claim it as your own. hi - ianncity texttext-generation10K<n<100K18 likes300 downloads2mo agoHugging Face17hotdogs /uka-glm-5.2 🏆 uka GLM-5.2 Reasoning Reasoning trace dataset for QLoRA fine-tuning of coding agents 📋 Overview uka GLM-5.2 Reasoning is a curated reasoning trace dataset built from GLM-5.2 agent sessions, designed for QLoRA fine-tuning of coding agents. Why Is It Easy to Use? Feature Description 🎯 Ready to Train ChatML format — works directly with HuggingFace SFTTrainer, no conversion needed 📦 Multiple Formats Both JSONL (readable) and… See the full description on the dataset page: https://huggingface.co/datasets/hotdogs/uka-glm-5.2.texttext-generation10K<n<100K6 likes285 downloads3mo agoHugging Face18Lalo42 /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/Lalo42/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes254 downloads1mo agoHugging Face19ianncity /GLM-5.2-Logic-Puzzles GLM-5.2 · Logical Puzzles 6000x traces distilled from GLM-5.2 on High reasoning Token Count: 5M~? Distribution: Puzzles: •Tokenization blindless ex: counting the r's in strawberry •Goal reasoning ex: the car wash test (theres no car wash question exactly just prompts like it so its not just benchmaxxing) •Reading comprehension traps •Temporal reasoning •Many other categories not worth mentioning Prompts… See the full description on the dataset page: https://huggingface.co/datasets/ianncity/GLM-5.2-Logic-Puzzles.texttext-generation1K<n<10K17 likes243 downloads2mo agoHugging Face20Solstice-AI /Axiom-1.0-Opus4.7-Kimi2.6-GLM5.2-Deepseek4-Mythos5-Fable5-Qwen3.7 Project Axiom 1.0 (102 GB Reasoning Corpus) 27-Billion Token Pure-Text Chain-of-Thought Corpus Across 7 Frontier Architectures Executive Summary Project Axiom 1.0 is a landmark, high-density, multi-architecture reasoning corpus comprising 102 GB of uncompressed, pure-text JSONL data (axiom.jsonl). Curated by Shreyan Gondaliya and the Solstice-AI research team, the dataset synthesizes ~5.74 million unique samples and ~27.3 billion tokens of… See the full description on the dataset page: https://huggingface.co/datasets/Solstice-AI/Axiom-1.0-Opus4.7-Kimi2.6-GLM5.2-Deepseek4-Mythos5-Fable5-Qwen3.7.text-generation13 likes232 downloads21d agoHugging Face21JessieWei /GLM-5.2-FP8-nemotron-codealpaca GLM-5.2-FP8-nemotron-codealpaca Training data for UCloud-org/GLM-5.2-FP8-DFlash, a DFlash speculative-decoding drafter for zai-org/GLM-5.2-FP8. A mix of code / math / chat prompts from two public instruction datasets (see Composition); all assistant responses are regenerated by GLM-5.2-FP8 so the targets match the verifier's own output distribution — the data recipe specified in the DFlash paper (Appendix A.1). 800,022 single-turn conversations, English-dominant Generation:… See the full description on the dataset page: https://huggingface.co/datasets/JessieWei/GLM-5.2-FP8-nemotron-codealpaca.texttext-generation100K<n<1M3 likes220 downloads2mo agoHugging Face22Crownelius /GLM-5.2-CoT-Library GLM-5.2 — CoT Library A maintained mirror of publicly-available GLM-5.2 chain-of-thought datasets on Hugging Face — content-verified, deduplicated, and attributed to their original authors. Dataset Viewer | Parquet // what this is A maintained library — a community mirror of publicly-available GLM-5.2 CoT datasets, aggregated, validity-filtered and content-verified, with per-row source attribution in first_source_dataset. It is not Crownelius' own data — every… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/GLM-5.2-CoT-Library.tabulartext-generation10K<n<100K2 likes219 downloads2mo agoHugging Face23Jinzy2025 /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/Jinzy2025/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes173 downloads1mo agoHugging Face24ianncity /GLM-5.2-Science GLM-5.2 · Science-50000x 50,000x traces distilled from GLM-5.2 on High reasoning Physics · Chemistry · Biology Token Count: 160M Theres prompt overlap with my Kimi K2.5 dataset science subset, which I think those prompts are getting used in alot of places now You can use this dataset for any purpose and you dont need to credit me, preferably dont claim it as your own. hi - ianncity texttext-generation10K<n<100K19 likes172 downloads2mo agoHugging Face25poppingstar /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/poppingstar/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M1 likes171 downloads2mo agoHugging Face26Davd-b01 /thinkingcap-condensed-qwen3.8-glm5.2-kimi-k3 ThinkingCap Condensed — Qwen3.8 / GLM-5.2 / Kimi-K3 Condensed ThinkingCap-style reasoning traces for SFT. 1,985 traces: each row pairs a full multi-turn teacher trace (Qwen3.8-Max, GLM-5.2 or Kimi K3, via r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillation) with a condensed TC-style version (short <think> + definitive numbered answer) generated by bottlecapai/ThinkingCap-Qwen3.6-27B using the thinkingcap system prompt. Format: JSONL (data/condensed.jsonl), 1,985 rows, UTF-8.… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinkingcap-condensed-qwen3.8-glm5.2-kimi-k3.texttext-generation1K<n<10K1 likes168 downloads1mo agoHugging Face27ArkhAngelLifeJiggy /glm-5.2-coding-and-debugging-traces GLM 5.2 Agent Traces 207 TRAJECTORIES · 1,821 TRAINING ROWS · 1 MB PARQUET · 35 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from GLM 5.2 (glm-5.2). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task domain. This is an actively growing… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/glm-5.2-coding-and-debugging-traces.tabulartext-generation1K<n<10K0 likes162 downloads2mo agoHugging Face28marin-community /glm-5.2-kernelgym-rollouts GLM-5.2 KernelGym Rollouts This dataset contains 3,200 feedback-driven GPU-kernel optimization trajectories generated by zai-org/GLM-5.2-FP8: 100 validation tasks, two backends (inline CUDA and Triton), and 16 rollouts per task. Each trajectory retains the prompt/feedback message history, model responses and reasoning, extracted kernel code, KernelGym compilation and correctness results, profiling metadata, token usage, and stopping decision. Every published record ended with… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/glm-5.2-kernelgym-rollouts.tabulartext-generation1K<n<10K2 likes138 downloads2mo agoHugging Face29Helloxiaolaodi /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/Helloxiaolaodi/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes136 downloads1mo agoHugging Face30bunnycore /qwen3.8-max-glm5.2-kimi-k3-sft-balanced Multi-Teacher SFT Balanced Dataset (57,937 Traces) Quality-filtered, deduplicated, multi-teacher SFT corpus packaged from r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillation (subset: sft_balanced). Dataset Overview Source Dataset: r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillation Subset: sft_balanced Total Traces: 57,937 (under the 100k cap) Standardized Column: The conversation turns are strictly standardized under the messages column (resolved and mapped from… See the full description on the dataset page: https://huggingface.co/datasets/bunnycore/qwen3.8-max-glm5.2-kimi-k3-sft-balanced.texttext-generation10K<n<100K1 likes132 downloads19d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.