CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AgentNativeResearchLab /arc-agi3-kimi-k2.7-ar25 ARC-AGI-3 ar25 — Agent Trajectories (kimi-k2.7) Gameplay trajectories from the harness×model pair kimi-k2.7 playing the ARC-AGI-3 game ar25, part of the ARA-as-world-model generalization experiment. The agent builds a structured world model (an Agent-Native Research Artifact) live during play and consults it to crack levels it cannot solve from cold exploration. One dataset repo per harness×model×game: sibling repos arc-agi3-<harness>-<model>-<game> hold the same game played by… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-kimi-k2.7-ar25.tabularreinforcement-learningn<1K0 likes1.1k downloads25d agoHugging Face02armand0e /kimi-k2.6-claude-code-tracesThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. Kimi K2.6 Claude Code Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by moonshotai/kimi-k2.6. JSONL files: 36 Format Each file is newline-delimited JSON representing a single captured agent session. The trace schema is designed for… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/kimi-k2.6-claude-code-traces.tabularn<1K4 likes804 downloads4mo agoHugging Face03Jackrong /Kimi-K2.5-Reasoning-1M-Cleaned 🪐 Kimi-K2.5-Reasoning-1M-Cleaned Kimi-K2.5-Reasoning-1M-Cleaned is a cleaned derivative of ianncity/KIMI-K2.5-1000000x. It preserves the original four-config layout from the source dataset and rewrites each record into a unified reasoning-SFT schema with id, conversations, input, output, domain, and meta. Summary Source dataset: ianncity/KIMI-K2.5-1000000x Source author: ianncity Teacher model recorded in meta.teacher_model: KIMI-K2.5 Token lengths computed with… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Kimi-K2.5-Reasoning-1M-Cleaned.texttext-generation100K<n<1M36 likes594 downloads5mo agoHugging Face04SWE-Factory /DeepSWE-Agent-Kimi-K2-Trajectories-2.8Ktext1K<n<10K8 likes591 downloads1y agoHugging Face05ianncity /KIMI-K2.5-1000000x KIMI-K2.5-1000000x 1,000,000 reasoning traces distilled from KIMI-K2.5 on high reasoning, (Each subset has different questions) Distribution: Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#) Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset Math: 15% (Algebra, Calculus, Probability) - 200k more completions in kimiMath200k.jsonl Computer Science: 5% Logical Questions: 5% Creative Writing: 5%… See the full description on the dataset page: https://huggingface.co/datasets/ianncity/KIMI-K2.5-1000000x.texttext-generation100K<n<1M265 likes549 downloads6mo agoHugging Face06Crownelius /Creative-Writing-KimiK2.5-Cleaned Creative-Writing-KimiK2.5-Cleaned Cleaned creative writing SFT dataset from Kimi K2.5 (655 samples). Prompts cleaned, thinking traces preserved. Format Each line is a JSON object with: messages: list of message dicts with roles (system, user, assistant) System: writing quality instructions User: cleaned creative writing prompt Assistant: creative writing response (may include <think> traces) Stats Metric Value Total prompt tokens 80… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-KimiK2.5-Cleaned.texttext-generationn<1K8 likes345 downloads2mo agoHugging Face07SWE-Factory /DeepSWE-Agent-Kimi-K2-Trajectories-Rejection-Samplingtextn<1K0 likes267 downloads9mo agoHugging Face08placeholderlabs /Kimi-K2.5-Reasoning-General-Sharded Kimi-K2.5-Reasoning-General-Sharded Byte-preserving sequential 100 MB JSONL shards of selected files from Jackrong/Kimi-K2.5-Reasoning-1M-Cleaned. All credit for data generation and upstream curation belongs to the source authors. See the upstream dataset card for attribution, source descriptions and license terms. Included files: General-Distillation.jsonl. No filtering, shuffling, normalization, tokenization or truncation was performed. Complete records and all original fields… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/Kimi-K2.5-Reasoning-General-Sharded.text100K<n<1M0 likes233 downloads18d agoHugging Face09armand0e /kimi-k2.6-agentThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. Kimi K2.6 Agent Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by moonshotai/kimi-k2.6. JSONL files: 15 Training-ready tools A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of this README.… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/kimi-k2.6-agent.tabularn<1K2 likes188 downloads4mo agoHugging Face10rAVEUK /Kimi-K2.5-Reasoning-1M-Cleaned 🪐 Kimi-K2.5-Reasoning-1M-Cleaned Kimi-K2.5-Reasoning-1M-Cleaned is a cleaned derivative of ianncity/KIMI-K2.5-1000000x. It preserves the original four-config layout from the source dataset and rewrites each record into a unified reasoning-SFT schema with id, conversations, input, output, domain, and meta. Summary Source dataset: ianncity/KIMI-K2.5-1000000x Source author: ianncity Teacher model recorded in meta.teacher_model: KIMI-K2.5 Token lengths computed with… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Kimi-K2.5-Reasoning-1M-Cleaned.texttext-generation100K<n<1M2 likes163 downloads5mo agoHugging Face11Crownelius /Creative-Writing-Reasoning-KimiK2.5-600x Pulitzer Diamond Prose KIMI Seeds This dataset contains 655 high-quality creative writing seeds generated using Kimi-v1. Each entry represents a story opening designed to meet high literary standards, including internal thinking traces used during generation. How it was made The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary requirements:… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Reasoning-KimiK2.5-600x.texttext-generationn<1K8 likes150 downloads2mo agoHugging Face12trjxter /Kimi-K2.7-CodingTraces-9000x Kimi K2.7 Coding Traces 9000x A validated 9,014-row coding and software-engineering reasoning dataset generated with moonshotai/Kimi-K2.7-Code. Every row contains a coding-focused prompt, a separated reasoning trace, and a final answer. The release was built from a durable Google Drive generation pipeline and underwent a complete two-pass schema and delimiter audit before publication. Generation configuration Setting Value Teacher… See the full description on the dataset page: https://huggingface.co/datasets/trjxter/Kimi-K2.7-CodingTraces-9000x.texttext-generation1K<n<10K9 likes123 downloads3mo agoHugging Face13JBrightmanAI /Kimi-K2.5-Reasoning-1M-Cleaned 🪐 Kimi-K2.5-Reasoning-1M-Cleaned Kimi-K2.5-Reasoning-1M-Cleaned is a cleaned derivative of ianncity/KIMI-K2.5-1000000x. It preserves the original four-config layout from the source dataset and rewrites each record into a unified reasoning-SFT schema with id, conversations, input, output, domain, and meta. Summary Source dataset: ianncity/KIMI-K2.5-1000000x Source author: ianncity Teacher model recorded in meta.teacher_model: KIMI-K2.5 Token lengths computed… See the full description on the dataset page: https://huggingface.co/datasets/JBrightmanAI/Kimi-K2.5-Reasoning-1M-Cleaned.texttext-generation100K<n<1M1 likes114 downloads2mo agoHugging Face14malaiwah /kimi-k25-tiny-fidelity-root-v1 kimi-k25 random CPU fixture root A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/kimi-k25-tiny-random-bf16. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after it). Same… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/kimi-k25-tiny-fidelity-root-v1.tabularn<1K0 likes97 downloads17d agoHugging Face15mondk /Code-sonnet-5-gpt-5.5-kimi-k2.5Hello guys, this is a dataset from an AI that's good at writing code uh... textn<1K4 likes85 downloads1mo agoHugging Face16trjxter /Kimi-K2.6-Reasoning-3300x-WandB Kimi-K2.6-Reasoning-3300x-WandB Kimi-K2.6-Reasoning-3300x-WandB is a W&B-only synthetic reasoning dataset generated with Kimi-K2.6 through Weights & Biases Inference. This dataset is the pure W&B-generated subset from a larger planned 8,000-example Kimi reasoning distillation run. Generation stopped when the W&B quota limit was reached, and the completed accepted rows were audited, cleaned, and exported as a standalone dataset. This release contains 3,303 accepted W&B-generated rows… See the full description on the dataset page: https://huggingface.co/datasets/trjxter/Kimi-K2.6-Reasoning-3300x-WandB.texttext-generation1K<n<10K7 likes77 downloads4mo agoHugging Face17YurinKO /KIMI-K2.5-1000000-RU KIMI-K2.5-1000000x 1,000,000 reasoning traces distilled from KIMI-K2.5 on high reasoning, (Each subset has different questions) Distribution: Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#) Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset Math: 15% (Algebra, Calculus, Probability) - 200k more completions in kimiMath200k.jsonl Computer Science: 5% Logical Questions: 5% Creative Writing: 5%… See the full description on the dataset page: https://huggingface.co/datasets/YurinKO/KIMI-K2.5-1000000-RU.texttext-generation100K<n<1M0 likes73 downloads19d agoHugging Face18Crownelius /KimiK2.5-2000x Kimi K2.5 9000x Dataset Dataset Description This dataset contains 2144 high-quality samples generated using Kimi K2.5 model, covering diverse tasks including code generation, mathematical reasoning, and general problem-solving. Dataset Summary Total Samples: 2144 Model: Kimi K2.5 Languages: English Format: JSON License: Apache 2.0 Task Distribution The dataset includes samples across multiple domains: Code Generation: Programming… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/KimiK2.5-2000x.texttext-generation1K<n<10K1 likes71 downloads2mo agoHugging Face19nick007x /KIMI-K2.5-1000000x KIMI-K2.5-1000000x 1,000,000 reasoning traces distilled from KIMI-K2.5 on high reasoning, (Each subset has different questions) Distribution: Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#) Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset Math: 15% (Algebra, Calculus, Probability) - 200k more completions in kimiMath200k.jsonl Computer Science: 5% Logical Questions: 5% Creative Writing: 5%… See the full description on the dataset page: https://huggingface.co/datasets/nick007x/KIMI-K2.5-1000000x.texttext-generation100K<n<1M0 likes69 downloads5mo agoHugging Face20Miska25 /Kimi-K2.5-Reasoning-Reduced-Luna Kimi K2.5 Reasoning Reduced with GPT-5.6 Luna This dataset contains synthetic, lossy compressions of reasoning traces from Jackrong/Kimi-K2.5-Reasoning-1M-Cleaned, configuration General-Distillation. The final-answer suffix is copied programmatically from the source and is not regenerated by the model. The compressed reasoning is synthetic and is not guaranteed to preserve every logical detail. Training columns Train on conversations_reduced or output_reduced. The… See the full description on the dataset page: https://huggingface.co/datasets/Miska25/Kimi-K2.5-Reasoning-Reduced-Luna.texttext-generation10K<n<100K0 likes65 downloads2mo agoHugging Face21TeichAI /kimi-k2-thinking-1000xThis is a reasoning dataset created using Kimi k2 thinking from MoonshotAI. Some of these questions are from reedmayhew and the rest were generated. Most of the questions cover the following topics: Web Development, Logic, Math, Embedded Systems, Web Design and Python Scripting. The dataset is meant for creating distilled versions of Kimi k2 thinking by fine-tuning already existing open-source LLMs. textn<1K14 likes59 downloads11mo agoHugging Face22ansulev /kimi-k2.5-reasoning-1m-cleaned 🪐 Kimi-K2.5-Reasoning-1M-Cleaned Kimi-K2.5-Reasoning-1M-Cleaned is a cleaned derivative of ianncity/KIMI-K2.5-1000000x. It preserves the original four-config layout from the source dataset and rewrites each record into a unified reasoning-SFT schema with id, conversations, input, output, domain, and meta. Summary Source dataset: ianncity/KIMI-K2.5-1000000x Source author: ianncity Teacher model recorded in meta.teacher_model: KIMI-K2.5 Token lengths computed… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/kimi-k2.5-reasoning-1m-cleaned.texttext-generation100K<n<1M0 likes58 downloads5mo agoHugging Face23uniquealexx /Kimi-K2.6-Thinking-200x Dataset Card (Kimi-K2.6-Thinking-200x) Dataset Summary Kimi-K2.6-Reasoning-207 is a high-quality distilled reasoning dataset designed for supervised fine-tuning (SFT) of small language models. This dataset uses a curated seed question set covering Mathematics, Code, Logic, Science, Analysis, and Instruction-following domains. By calling the Kimi-K2.6 model via the Moonshot AI API as the teacher model, it generates high-quality responses featuring long-form step-by-step… See the full description on the dataset page: https://huggingface.co/datasets/uniquealexx/Kimi-K2.6-Thinking-200x.textquestion-answeringn<1K2 likes57 downloads5mo agoHugging Face24trjxter /Kimi-K2.6-Technical-Reasoning-AddOn-3300x Kimi-K2.6-Technical-Reasoning-AddOn-3300x This dataset is a technical reasoning add-on dataset generated with Kimi K2.6 as the teacher model. The dataset was designed as an additional technical reasoning trace set for downstream SFT experiments, especially around math, graduate-level science, coding, and debugging/code-repair style prompts. Dataset Summary Dataset name: Kimi-K2.6-Technical-Reasoning-AddOn-3300x Teacher model: Kimi-K2.6 Backend: W&B… See the full description on the dataset page: https://huggingface.co/datasets/trjxter/Kimi-K2.6-Technical-Reasoning-AddOn-3300x.texttext-generation1K<n<10K1 likes56 downloads4mo agoHugging Face25EngMuhammadAtef /Kimi-K2.5-Reasoning-1M-Cleaned 🪐 Kimi-K2.5-Reasoning-1M-Cleaned Kimi-K2.5-Reasoning-1M-Cleaned is a cleaned derivative of ianncity/KIMI-K2.5-1000000x. It preserves the original four-config layout from the source dataset and rewrites each record into a unified reasoning-SFT schema with id, conversations, input, output, domain, and meta. Summary Source dataset: ianncity/KIMI-K2.5-1000000x Source author: ianncity Teacher model recorded in meta.teacher_model: KIMI-K2.5 Token lengths computed with… See the full description on the dataset page: https://huggingface.co/datasets/EngMuhammadAtef/Kimi-K2.5-Reasoning-1M-Cleaned.texttext-generation100K<n<1M0 likes52 downloads5mo agoHugging Face26Jongsim /KIMI-K2.5-filteredtext100K<n<1M0 likes46 downloads5mo agoHugging Face27WWX0825 /KIMI-K2.5-1000000x KIMI-K2.5-1000000x 1,000,000 reasoning traces distilled from KIMI-K2.5 on high reasoning, (Each subset has different questions) Distribution: Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#) Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset Math: 15% (Algebra, Calculus, Probability) - 200k more completions in kimiMath200k.jsonl Computer Science: 5% Logical Questions: 5% Creative Writing: 5%… See the full description on the dataset page: https://huggingface.co/datasets/WWX0825/KIMI-K2.5-1000000x.texttext-generation100K<n<1M0 likes45 downloads6mo agoHugging Face28xrist0bg /kimi-k2-0905-20M 20M token synthetic instruction dataset (Kimi 0905) User prompts are extracted from three curated instruction-following datasets. Low-quality and repetitive prompts are identified and removed or rewritten using Gemini 3 Flash (+adding medatada for each message). The resulting 15,825 filtered user prompts are sent to Kimi K2 0905 to generate high-quality synthetic responses. Difficulty Split Medium: 48.3% (7,638) Hard: 27.6% (4,373) — mostly from… See the full description on the dataset page: https://huggingface.co/datasets/xrist0bg/kimi-k2-0905-20M.textquestion-answering10K<n<100K1 likes41 downloads8mo agoHugging Face29bitsydarel /KIMI-K2.5-1000000x KIMI-K2.5-1000000x 1,000,000 reasoning traces distilled from KIMI-K2.5 on high reasoning, (Each subset has different questions) Distribution: Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#) Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset Math: 15% (Algebra, Calculus, Probability) - 200k more completions in kimiMath200k.jsonl Computer Science: 5% Logical Questions: 5% Creative Writing: 5%… See the full description on the dataset page: https://huggingface.co/datasets/bitsydarel/KIMI-K2.5-1000000x.texttext-generation100K<n<1M0 likes39 downloads5mo agoHugging Face30TheDrMoniker /KIMI-K2.5-1000000x KIMI-K2.5-1000000x 1,000,000 reasoning traces distilled from KIMI-K2.5 on high reasoning, (Each subset has different questions) Distribution: Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#) Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset Math: 15% (Algebra, Calculus, Probability) - 200k more completions in kimiMath200k.jsonl Computer Science: 5% Logical Questions: 5% Creative Writing: 5%… See the full description on the dataset page: https://huggingface.co/datasets/TheDrMoniker/KIMI-K2.5-1000000x.texttext-generation100K<n<1M0 likes37 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.