CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01beyoru /kimi-k3-distillation kimi-k3-distillation Single-teacher slice of r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillation, filtered to teacher_model == "kimi-code/k3" only. The Qwen3.8-Max-Preview and GLM-5.2 traces are removed. 4,347 rows — 3,918 train / 212 validation / 217 test. from datasets import load_dataset ds = load_dataset("beyoru/kimi-k3-distillation") # sft: messages + tools ds = load_dataset("beyoru/kimi-k3-distillation", "canonical") # + full audit columns… See the full description on the dataset page: https://huggingface.co/datasets/beyoru/kimi-k3-distillation.tabulartext-generation100K<n<1M15 likes6k downloads2mo agoHugging Face02r0b0tlab /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M264 likes5.8k downloads2mo agoHugging Face03AgentNativeResearchLab /arc-agi3-kimi-k2.7-su15 ARC-AGI-3 su15 — Agent Trajectories (kimi-k2.7) Gameplay trajectories from the harness×model pair kimi-k2.7 playing the ARC-AGI-3 game su15, part of the ARA-as-world-model generalization experiment. The agent builds a structured world model (an Agent-Native Research Artifact) live during play and consults it to crack levels it cannot solve from cold exploration. One dataset repo per harness×model×game: sibling repos arc-agi3-<harness>-<model>-<game> hold the same game played by… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-kimi-k2.7-su15.reinforcement-learning0 likes3k downloads23d agoHugging Face04echel0nn1881 /kimi-cyber-reasoning Kimi Cyber Reasoning 997 chain-of-thought records covering 13 cybersecurity disciplines and 4 systems engineering domains, distilled from the Kimi K3 reasoning model via API. Every record provides an explicit step-by-step <think> reasoning trace followed by a technical resolution, unified code diff fix, or structured tool invocation. The dataset was curated as an anchor set for training, healing, and specializing compact reasoning models on systems security and tool calling… See the full description on the dataset page: https://huggingface.co/datasets/echel0nn1881/kimi-cyber-reasoning.texttext-generationn<1K89 likes2.3k downloads19d agoHugging Face05ianncity /KIMI-K2.5-1000000x KIMI-K2.5-1000000x 1,000,000 reasoning traces distilled from KIMI-K2.5 on high reasoning, (Each subset has different questions) Distribution: Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#) Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset Math: 15% (Algebra, Calculus, Probability) - 200k more completions in kimiMath200k.jsonl Computer Science: 5% Logical Questions: 5% Creative Writing: 5%… See the full description on the dataset page: https://huggingface.co/datasets/ianncity/KIMI-K2.5-1000000x.texttext-generation100K<n<1M265 likes2k downloads6mo agoHugging Face06greghavens /kimi-k3-coding-and-debugging-traces Kimi K3 Coding, Tool Use & Instruction Following Traces 582 TRAJECTORIES · 3,956 TRAINING ROWS · 3 MB PARQUET · 72 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Kimi K3 (moonshotai/kimi-k3). The category and row-share tables below describe the actual mix seen during training rather than assuming a… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/kimi-k3-coding-and-debugging-traces.tabulartext-generation1K<n<10K65 likes1.6k downloads2mo agoHugging Face07AgentNativeResearchLab /arc-agi3-kimi-k2.7-ls20 ARC-AGI-3 ls20 — Agent Trajectories (kimi-k2.7) Gameplay trajectories from the harness×model pair kimi-k2.7 playing the ARC-AGI-3 game ls20, part of the ARA-as-world-model generalization experiment. The agent builds a structured world model (an Agent-Native Research Artifact) live during play and consults it to crack levels it cannot solve from cold exploration. One dataset repo per harness×model×game: sibling repos arc-agi3-<harness>-<model>-<game> hold the same game played by… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-kimi-k2.7-ls20.reinforcement-learning0 likes1.6k downloads23d agoHugging Face08BangumiBase /kimitobokunosaigonosenjouaruiwasekaigahajimaruseisenseasonii Bangumi Image Base of Kimi To Boku No Saigo No Senjou, Aruiwa Sekai Ga Hajimaru Seisen Season Ii This is the image base of bangumi Kimi to Boku no Saigo no Senjou, Aruiwa Sekai ga Hajimaru Seisen Season II, we detected 100 characters, 7344 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kimitobokunosaigonosenjouaruiwasekaigahajimaruseisenseasonii.image1K<n<10K0 likes1.5k downloads1y agoHugging Face09bevangelista /AIME_2000_2026_Kimi_K3 AIME 2000–2026 — Kimi K3 reasoning traces 🔄 Changelog 2026-08-08 — full re-generation. All reasoning traces were regenerated from scratch and re-verified against the official answer key. New schema — added gen_attempts_low, gen_attempts_high; renamed gen_parsed_answer → gen_answer_int and answer_note → problem_note; removed gen_effort, gen_pass1. New generation — only use the bare problem (v1 appended an "ANSWER:" format instruction), so traces are cleaner.… See the full description on the dataset page: https://huggingface.co/datasets/bevangelista/AIME_2000_2026_Kimi_K3.tabulartext-generationn<1K2 likes1.5k downloads2mo agoHugging Face10AgentNativeResearchLab /arc-agi3-kimi-k2.7-g50t ARC-AGI-3 g50t — Agent Trajectories (kimi-k2.7) Gameplay trajectories from the harness×model pair kimi-k2.7 playing the ARC-AGI-3 game g50t, part of the ARA-as-world-model generalization experiment. The agent builds a structured world model (an Agent-Native Research Artifact) live during play and consults it to crack levels it cannot solve from cold exploration. One dataset repo per harness×model×game: sibling repos arc-agi3-<harness>-<model>-<game> hold the same game played by… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-kimi-k2.7-g50t.reinforcement-learning2 likes1.5k downloads23d agoHugging Face11o0Biggz0o /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/o0Biggz0o/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes1.2k downloads1mo agoHugging Face12AgentNativeResearchLab /arc-agi3-kimi-k2.7-tr87 ARC-AGI-3 tr87 — Agent Trajectories (kimi-k2.7) Gameplay trajectories from the harness×model pair kimi-k2.7 playing the ARC-AGI-3 game tr87, part of the ARA-as-world-model generalization experiment. The agent builds a structured world model (an Agent-Native Research Artifact) live during play and consults it to crack levels it cannot solve from cold exploration. One dataset repo per harness×model×game: sibling repos arc-agi3-<harness>-<model>-<game> hold the same game played by… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-kimi-k2.7-tr87.reinforcement-learning0 likes1.1k downloads23d agoHugging Face13AgentNativeResearchLab /arc-agi3-kimi-k2.7-ar25 ARC-AGI-3 ar25 — Agent Trajectories (kimi-k2.7) Gameplay trajectories from the harness×model pair kimi-k2.7 playing the ARC-AGI-3 game ar25, part of the ARA-as-world-model generalization experiment. The agent builds a structured world model (an Agent-Native Research Artifact) live during play and consults it to crack levels it cannot solve from cold exploration. One dataset repo per harness×model×game: sibling repos arc-agi3-<harness>-<model>-<game> hold the same game played by… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-kimi-k2.7-ar25.tabularreinforcement-learningn<1K0 likes1.1k downloads23d agoHugging Face14AgentNativeResearchLab /discoverphysics-kimi-k2.7-ara DiscoverPhysics × Kimi K2.7 (kimi-code CLI, thinking=on) — 11-world ARA knowledge artifacts Agent-Native Research Artifacts (ARA) produced by a Kimi K2.7 (kimi-code CLI, thinking=on) coding-agent session solving all 11 worlds of the DiscoverPhysics scientific-discovery benchmark (seed 0, noise_frac 0.075, ≤16 experiment rounds), driven through the same harness-agnostic bridge and ARA scaffold as the sibling fable run. Official verdicts: 1/11 PASS — criteria and per-world numbers… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/discoverphysics-kimi-k2.7-ara.0 likes1k downloads1mo agoHugging Face15AgentNativeResearchLab /arc-agi3-kimi-k2.7-ft09 ARC-AGI-3 ft09 — Agent Trajectories (kimi-k2.7) Gameplay trajectories from the harness×model pair kimi-k2.7 playing the ARC-AGI-3 game ft09, part of the ARA-as-world-model generalization experiment. The agent builds a structured world model (an Agent-Native Research Artifact) live during play and consults it to crack levels it cannot solve from cold exploration. One dataset repo per harness×model×game: sibling repos arc-agi3-<harness>-<model>-<game> hold the same game played by… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-kimi-k2.7-ft09.reinforcement-learning0 likes986 downloads23d agoHugging Face16BangumiBase /kiminokotogadaidaidaidaidaisukina100ninnokanojo Bangumi Image Base of Kimi No Koto Ga Daidaidaidaidaisuki Na 100-nin No Kanojo This is the image base of bangumi Kimi no Koto ga Daidaidaidaidaisuki na 100-nin no Kanojo, we detected 55 characters, 4768 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kiminokotogadaidaidaidaidaisukina100ninnokanojo.image1K<n<10K0 likes968 downloads2y agoHugging Face17armand0e /kimi-k2.6-claude-code-tracesThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. Kimi K2.6 Claude Code Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by moonshotai/kimi-k2.6. JSONL files: 36 Format Each file is newline-delimited JSON representing a single captured agent session. The trace schema is designed for… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/kimi-k2.6-claude-code-traces.tabularn<1K4 likes845 downloads4mo agoHugging Face18mateowilliam /kimi-k2.6-reap-observations-v1 Kimi-K2.6 REAP Observation Data (v1) Per-layer expert routing + activation statistics captured from moonshotai/Kimi-K2.6 under the REAP layerwise observer (PR #17, CerebrasResearch/reap). What this is This dataset contains the observer output of a full REAP calibration pass on Kimi-K2.6. It is not a pruned model. Each record describes per-token routing decisions, expert activation norms, and the REAP saliency ingredients for every MoE layer of the base model. Downstream… See the full description on the dataset page: https://huggingface.co/datasets/mateowilliam/kimi-k2.6-reap-observations-v1.text-generation10M<n<100M0 likes815 downloads5mo agoHugging Face19festr2 /kimi-k3-distribution-fidelity-1024x2048-v1 Kimi K3 distribution-fidelity reference Status The artifact is qualified for paired, teacher-forced comparison of Kimi K3 compressed checkpoints against the official MXFP4 checkpoint. It contains 1,024 distinct 2,048-token source contexts and 2,096,128 scored next-token positions. The primary metric is: KL(official MXFP4 distribution || candidate distribution) The reference stores BF16 hidden states after final RMSNorm and before the language-model head.… See the full description on the dataset page: https://huggingface.co/datasets/festr2/kimi-k3-distribution-fidelity-1024x2048-v1.text-generation0 likes784 downloads1mo agoHugging Face20BangumiBase /kimiwameidosama Bangumi Image Base of Kimi Wa Meido-sama. This is the image base of bangumi Kimi wa Meido-sama., we detected 48 characters, 3792 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability). Here is… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kimiwameidosama.image1K<n<10K0 likes736 downloads2y agoHugging Face210xSero /kimi-k2.6-reap-observations-v1 Kimi-K2.6 REAP Observation Data (v1) Per-layer expert routing + activation statistics captured from moonshotai/Kimi-K2.6 under the REAP layerwise observer (PR #17, CerebrasResearch/reap). What this is This dataset contains the observer output of a full REAP calibration pass on Kimi-K2.6. It is not a pruned model. Each record describes per-token routing decisions, expert activation norms, and the REAP saliency ingredients for every MoE layer of the base model. Downstream… See the full description on the dataset page: https://huggingface.co/datasets/0xSero/kimi-k2.6-reap-observations-v1.text-generation10M<n<100M1 likes734 downloads5mo agoHugging Face22OpenMed /synthvision-validated-qwen-by-kimi synthvision-validated-qwen-by-kimi Qwen 3.5 annotations validated by Kimi K2.5 (93.1% pass rate) Records: 55,359 About Cross-validated subset from the SynthVision pipeline. Kimi K2.5 reviewed all 59,476 Qwen 3.5 annotations and confirmed 55,359 as consistent with the source images (93.1% pass rate). Validation criteria: consistent == true AND confidence >= 0.7. Records that failed validation were removed — primarily cases where the annotator hallucinated findings not… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/synthvision-validated-qwen-by-kimi.textvisual-question-answering100K<n<1M3 likes725 downloads6mo agoHugging Face23ansulev /qwen3.8-max-glm5.2-kimi-k3-distill Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/qwen3.8-max-glm5.2-kimi-k3-distill.tabulartext-generation10M<n<100M0 likes700 downloads1mo agoHugging Face24inferenceport-ai /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/inferenceport-ai/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes668 downloads9d agoHugging Face25Jackrong /Kimi-K2.5-Reasoning-1M-Cleaned 🪐 Kimi-K2.5-Reasoning-1M-Cleaned Kimi-K2.5-Reasoning-1M-Cleaned is a cleaned derivative of ianncity/KIMI-K2.5-1000000x. It preserves the original four-config layout from the source dataset and rewrites each record into a unified reasoning-SFT schema with id, conversations, input, output, domain, and meta. Summary Source dataset: ianncity/KIMI-K2.5-1000000x Source author: ianncity Teacher model recorded in meta.teacher_model: KIMI-K2.5 Token lengths computed with… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Kimi-K2.5-Reasoning-1M-Cleaned.texttext-generation100K<n<1M36 likes663 downloads5mo agoHugging Face26AgentNativeResearchLab /arc-agi3-kimi-k2.7-s5i5 ARC-AGI-3 s5i5 — Agent Trajectories (kimi-k2.7) Gameplay trajectories from the harness×model pair kimi-k2.7 playing the ARC-AGI-3 game s5i5, part of the ARA-as-world-model generalization experiment. The agent builds a structured world model (an Agent-Native Research Artifact) live during play and consults it to crack levels it cannot solve from cold exploration. One dataset repo per harness×model×game: sibling repos arc-agi3-<harness>-<model>-<game> hold the same game played by… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-kimi-k2.7-s5i5.reinforcement-learning1 likes651 downloads23d agoHugging Face27vuhaian /kimi_v31_300ktext100K<n<1M0 likes641 downloads4mo agoHugging Face28AgentNativeResearchLab /fm-open-problems-kimi-k3-trajectories FrontierMath Open Problems trajectories — kimi-k3 Fixed-budget agent trajectories on FrontierMath: Open Problems (Epoch AI's collection of 50 genuinely unsolved research mathematics problems). Nothing is graded. Epoch's verifiers are not public; the harness verifier is a checkpoint stub that always writes reward 0 so the continue-until-timeout harness keeps re-prompting the agent until the fixed wall-clock budget (105 min/task, override_timeout_sec: 6300) elapses.… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/fm-open-problems-kimi-k3-trajectories.0 likes571 downloads1mo agoHugging Face29SWE-Factory /DeepSWE-Agent-Kimi-K2-Trajectories-2.8Ktext1K<n<10K8 likes538 downloads1y agoHugging Face30bhadra123 /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/bhadra123/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes509 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.