CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AgentNativeResearchLab /arc-agi3-kimi-k2.7-su15 ARC-AGI-3 su15 — Agent Trajectories (kimi-k2.7) Gameplay trajectories from the harness×model pair kimi-k2.7 playing the ARC-AGI-3 game su15, part of the ARA-as-world-model generalization experiment. The agent builds a structured world model (an Agent-Native Research Artifact) live during play and consults it to crack levels it cannot solve from cold exploration. One dataset repo per harness×model×game: sibling repos arc-agi3-<harness>-<model>-<game> hold the same game played by… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-kimi-k2.7-su15.reinforcement-learning0 likes3.5k downloads25d agoHugging Face02AgentNativeResearchLab /arc-agi3-kimi-k2.7-g50t ARC-AGI-3 g50t — Agent Trajectories (kimi-k2.7) Gameplay trajectories from the harness×model pair kimi-k2.7 playing the ARC-AGI-3 game g50t, part of the ARA-as-world-model generalization experiment. The agent builds a structured world model (an Agent-Native Research Artifact) live during play and consults it to crack levels it cannot solve from cold exploration. One dataset repo per harness×model×game: sibling repos arc-agi3-<harness>-<model>-<game> hold the same game played by… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-kimi-k2.7-g50t.reinforcement-learning2 likes1.8k downloads25d agoHugging Face03AgentNativeResearchLab /arc-agi3-kimi-k2.7-ls20 ARC-AGI-3 ls20 — Agent Trajectories (kimi-k2.7) Gameplay trajectories from the harness×model pair kimi-k2.7 playing the ARC-AGI-3 game ls20, part of the ARA-as-world-model generalization experiment. The agent builds a structured world model (an Agent-Native Research Artifact) live during play and consults it to crack levels it cannot solve from cold exploration. One dataset repo per harness×model×game: sibling repos arc-agi3-<harness>-<model>-<game> hold the same game played by… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-kimi-k2.7-ls20.reinforcement-learning0 likes1.7k downloads25d agoHugging Face04AgentNativeResearchLab /arc-agi3-kimi-k2.7-tr87 ARC-AGI-3 tr87 — Agent Trajectories (kimi-k2.7) Gameplay trajectories from the harness×model pair kimi-k2.7 playing the ARC-AGI-3 game tr87, part of the ARA-as-world-model generalization experiment. The agent builds a structured world model (an Agent-Native Research Artifact) live during play and consults it to crack levels it cannot solve from cold exploration. One dataset repo per harness×model×game: sibling repos arc-agi3-<harness>-<model>-<game> hold the same game played by… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-kimi-k2.7-tr87.reinforcement-learning0 likes1.2k downloads25d agoHugging Face05AgentNativeResearchLab /arc-agi3-kimi-k2.7-ar25 ARC-AGI-3 ar25 — Agent Trajectories (kimi-k2.7) Gameplay trajectories from the harness×model pair kimi-k2.7 playing the ARC-AGI-3 game ar25, part of the ARA-as-world-model generalization experiment. The agent builds a structured world model (an Agent-Native Research Artifact) live during play and consults it to crack levels it cannot solve from cold exploration. One dataset repo per harness×model×game: sibling repos arc-agi3-<harness>-<model>-<game> hold the same game played by… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-kimi-k2.7-ar25.tabularreinforcement-learningn<1K0 likes1.1k downloads25d agoHugging Face06AgentNativeResearchLab /arc-agi3-kimi-k2.7-ft09 ARC-AGI-3 ft09 — Agent Trajectories (kimi-k2.7) Gameplay trajectories from the harness×model pair kimi-k2.7 playing the ARC-AGI-3 game ft09, part of the ARA-as-world-model generalization experiment. The agent builds a structured world model (an Agent-Native Research Artifact) live during play and consults it to crack levels it cannot solve from cold exploration. One dataset repo per harness×model×game: sibling repos arc-agi3-<harness>-<model>-<game> hold the same game played by… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-kimi-k2.7-ft09.reinforcement-learning0 likes1.1k downloads25d agoHugging Face07AgentNativeResearchLab /discoverphysics-kimi-k2.7-ara DiscoverPhysics × Kimi K2.7 (kimi-code CLI, thinking=on) — 11-world ARA knowledge artifacts Agent-Native Research Artifacts (ARA) produced by a Kimi K2.7 (kimi-code CLI, thinking=on) coding-agent session solving all 11 worlds of the DiscoverPhysics scientific-discovery benchmark (seed 0, noise_frac 0.075, ≤16 experiment rounds), driven through the same harness-agnostic bridge and ARA scaffold as the sibling fable run. Official verdicts: 1/11 PASS — criteria and per-world numbers… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/discoverphysics-kimi-k2.7-ara.0 likes997 downloads1mo agoHugging Face08mateowilliam /kimi-k2.6-reap-observations-v1 Kimi-K2.6 REAP Observation Data (v1) Per-layer expert routing + activation statistics captured from moonshotai/Kimi-K2.6 under the REAP layerwise observer (PR #17, CerebrasResearch/reap). What this is This dataset contains the observer output of a full REAP calibration pass on Kimi-K2.6. It is not a pruned model. Each record describes per-token routing decisions, expert activation norms, and the REAP saliency ingredients for every MoE layer of the base model. Downstream… See the full description on the dataset page: https://huggingface.co/datasets/mateowilliam/kimi-k2.6-reap-observations-v1.text-generation10M<n<100M0 likes815 downloads5mo agoHugging Face09armand0e /kimi-k2.6-claude-code-tracesThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. Kimi K2.6 Claude Code Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by moonshotai/kimi-k2.6. JSONL files: 36 Format Each file is newline-delimited JSON representing a single captured agent session. The trace schema is designed for… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/kimi-k2.6-claude-code-traces.tabularn<1K4 likes804 downloads4mo agoHugging Face100xSero /kimi-k2.6-reap-observations-v1 Kimi-K2.6 REAP Observation Data (v1) Per-layer expert routing + activation statistics captured from moonshotai/Kimi-K2.6 under the REAP layerwise observer (PR #17, CerebrasResearch/reap). What this is This dataset contains the observer output of a full REAP calibration pass on Kimi-K2.6. It is not a pruned model. Each record describes per-token routing decisions, expert activation norms, and the REAP saliency ingredients for every MoE layer of the base model. Downstream… See the full description on the dataset page: https://huggingface.co/datasets/0xSero/kimi-k2.6-reap-observations-v1.text-generation10M<n<100M1 likes740 downloads5mo agoHugging Face11AgentNativeResearchLab /arc-agi3-kimi-k2.7-s5i5 ARC-AGI-3 s5i5 — Agent Trajectories (kimi-k2.7) Gameplay trajectories from the harness×model pair kimi-k2.7 playing the ARC-AGI-3 game s5i5, part of the ARA-as-world-model generalization experiment. The agent builds a structured world model (an Agent-Native Research Artifact) live during play and consults it to crack levels it cannot solve from cold exploration. One dataset repo per harness×model×game: sibling repos arc-agi3-<harness>-<model>-<game> hold the same game played by… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-kimi-k2.7-s5i5.reinforcement-learning1 likes649 downloads25d agoHugging Face12Jackrong /Kimi-K2.5-Reasoning-1M-Cleaned 🪐 Kimi-K2.5-Reasoning-1M-Cleaned Kimi-K2.5-Reasoning-1M-Cleaned is a cleaned derivative of ianncity/KIMI-K2.5-1000000x. It preserves the original four-config layout from the source dataset and rewrites each record into a unified reasoning-SFT schema with id, conversations, input, output, domain, and meta. Summary Source dataset: ianncity/KIMI-K2.5-1000000x Source author: ianncity Teacher model recorded in meta.teacher_model: KIMI-K2.5 Token lengths computed with… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Kimi-K2.5-Reasoning-1M-Cleaned.texttext-generation100K<n<1M36 likes594 downloads5mo agoHugging Face13SWE-Factory /DeepSWE-Agent-Kimi-K2-Trajectories-2.8Ktext1K<n<10K8 likes591 downloads1y agoHugging Face14ianncity /KIMI-K2.5-1000000x KIMI-K2.5-1000000x 1,000,000 reasoning traces distilled from KIMI-K2.5 on high reasoning, (Each subset has different questions) Distribution: Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#) Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset Math: 15% (Algebra, Calculus, Probability) - 200k more completions in kimiMath200k.jsonl Computer Science: 5% Logical Questions: 5% Creative Writing: 5%… See the full description on the dataset page: https://huggingface.co/datasets/ianncity/KIMI-K2.5-1000000x.texttext-generation100K<n<1M265 likes549 downloads6mo agoHugging Face15AgentNativeResearchLab /arc-agi3-kimi-k2.7-r11l ARC-AGI-3 r11l — Agent Trajectories (kimi-k2.7) Gameplay trajectories from the harness×model pair kimi-k2.7 playing the ARC-AGI-3 game r11l, part of the ARA-as-world-model generalization experiment. The agent builds a structured world model (an Agent-Native Research Artifact) live during play and consults it to crack levels it cannot solve from cold exploration. One dataset repo per harness×model×game: sibling repos arc-agi3-<harness>-<model>-<game> hold the same game played by… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-kimi-k2.7-r11l.reinforcement-learning1 likes536 downloads25d agoHugging Face16Crownelius /Creative-Writing-KimiK2.5-Cleaned Creative-Writing-KimiK2.5-Cleaned Cleaned creative writing SFT dataset from Kimi K2.5 (655 samples). Prompts cleaned, thinking traces preserved. Format Each line is a JSON object with: messages: list of message dicts with roles (system, user, assistant) System: writing quality instructions User: cleaned creative writing prompt Assistant: creative writing response (may include <think> traces) Stats Metric Value Total prompt tokens 80… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-KimiK2.5-Cleaned.texttext-generationn<1K8 likes345 downloads2mo agoHugging Face17SWE-Factory /DeepSWE-Agent-Kimi-K2-Trajectories-Rejection-Samplingtextn<1K0 likes267 downloads9mo agoHugging Face18malaiwah /kimi-k25-tiny-cpu-repro-v1 Kimi K2.5 / K2.6 / K2.7-Code complete tiny random BF16 CPU fixture Untrained independently seeded random weights; no upstream weights or training data. This is a reproducibility fixture, not useful language modeling or production quality evidence. Runtime and lineage Upstream moonshotai/Kimi-K2.7-Code@74797c9c62378b951a1f6fcf5c4631024e9b8bef. Actual loaded class: Kimi_K25ForConditionalGeneration. Complete untied head and real small vision tower/projector.… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/kimi-k25-tiny-cpu-repro-v1.tabularn<1K0 likes252 downloads17d agoHugging Face19placeholderlabs /Kimi-K2.5-Reasoning-General-Sharded Kimi-K2.5-Reasoning-General-Sharded Byte-preserving sequential 100 MB JSONL shards of selected files from Jackrong/Kimi-K2.5-Reasoning-1M-Cleaned. All credit for data generation and upstream curation belongs to the source authors. See the upstream dataset card for attribution, source descriptions and license terms. Included files: General-Distillation.jsonl. No filtering, shuffling, normalization, tokenization or truncation was performed. Complete records and all original fields… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/Kimi-K2.5-Reasoning-General-Sharded.text100K<n<1M0 likes233 downloads18d agoHugging Face20JBrightmanAI /kimi-k2.6-claude-code-tracesThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. Kimi K2.6 Claude Code Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by moonshotai/kimi-k2.6. JSONL files: 36 Format Each file is newline-delimited JSON representing a single captured agent session. The trace schema is designed for… See the full description on the dataset page: https://huggingface.co/datasets/JBrightmanAI/kimi-k2.6-claude-code-traces.1 likes211 downloads2mo agoHugging Face21armand0e /kimi-k2.6-agentThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. Kimi K2.6 Agent Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by moonshotai/kimi-k2.6. JSONL files: 15 Training-ready tools A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of this README.… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/kimi-k2.6-agent.tabularn<1K2 likes188 downloads4mo agoHugging Face22Avtrkrb /combined-reasoning-opus-4.6-opus-4.7-kimi-k2.5-kimi-k2.6-glm-5.1 Combined Reasoning Distill — Multi-Model A large-scale unified reasoning dataset combining thinking and chain-of-thought traces distilled from frontier models, normalized into a single consistent schema for fine-tuning. Includes data from Claude (Opus 4.5/4.6/4.7, Sonnet 4.5/4.6, Haiku 4.5), GPT (5.1/5.2), Gemini 3 Pro Preview, Kimi (K2/K2.5/K2.6), GLM (4.6/4.7/5.1), MiniMax M2.1, Grok Code Fast 1, and more. Schema Every row has a single field: Field Type… See the full description on the dataset page: https://huggingface.co/datasets/Avtrkrb/combined-reasoning-opus-4.6-opus-4.7-kimi-k2.5-kimi-k2.6-glm-5.1.texttext-generation1M<n<10M14 likes172 downloads4mo agoHugging Face23rAVEUK /Kimi-K2.5-Reasoning-1M-Cleaned 🪐 Kimi-K2.5-Reasoning-1M-Cleaned Kimi-K2.5-Reasoning-1M-Cleaned is a cleaned derivative of ianncity/KIMI-K2.5-1000000x. It preserves the original four-config layout from the source dataset and rewrites each record into a unified reasoning-SFT schema with id, conversations, input, output, domain, and meta. Summary Source dataset: ianncity/KIMI-K2.5-1000000x Source author: ianncity Teacher model recorded in meta.teacher_model: KIMI-K2.5 Token lengths computed with… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Kimi-K2.5-Reasoning-1M-Cleaned.texttext-generation100K<n<1M2 likes163 downloads5mo agoHugging Face24marin-community /open-thoughts-4-6865-math-kimi-k2pt5-annotated-32768-tokens-n8-reformatted open-thoughts-4-6865-math-kimi-k2pt5-annotated-32768-tokens Math reasoning responses generated by Kimi K2.5 (moonshotai/Kimi-K2.5) via a Together AI dedicated instance. Overview Total rows: 54,920 Unique prompts: 6,865 (each with 8 response annotations) Source prompts: marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n8-reformatted Generation model: moonshotai/Kimi-K2.5 Max tokens: 32,768 Temperature: 0.8 Tokenizer used for stats: Qwen/Qwen2.5-3B… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-6865-math-kimi-k2pt5-annotated-32768-tokens-n8-reformatted.tabular10K<n<100K0 likes158 downloads5mo agoHugging Face25Crownelius /Creative-Writing-Reasoning-KimiK2.5-600x Pulitzer Diamond Prose KIMI Seeds This dataset contains 655 high-quality creative writing seeds generated using Kimi-v1. Each entry represents a story opening designed to meet high literary standards, including internal thinking traces used during generation. How it was made The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary requirements:… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Reasoning-KimiK2.5-600x.texttext-generationn<1K8 likes150 downloads2mo agoHugging Face26trjxter /Kimi-K2.7-CodingTraces-9000x Kimi K2.7 Coding Traces 9000x A validated 9,014-row coding and software-engineering reasoning dataset generated with moonshotai/Kimi-K2.7-Code. Every row contains a coding-focused prompt, a separated reasoning trace, and a final answer. The release was built from a durable Google Drive generation pipeline and underwent a complete two-pass schema and delimiter audit before publication. Generation configuration Setting Value Teacher… See the full description on the dataset page: https://huggingface.co/datasets/trjxter/Kimi-K2.7-CodingTraces-9000x.texttext-generation1K<n<10K9 likes123 downloads3mo agoHugging Face27Avtrkrb /combined-reasoning-kimi-k2.5-glm-5.1 Combined Reasoning Distill — Multi-Model A large-scale unified reasoning dataset combining thinking and chain-of-thought traces distilled from frontier models, normalized into a single consistent schema for fine-tuning. Includes data from Kimi K2.5 & GLM 5.1. Schema Every row has a single field: Field Type Description messages list[dict] Conversation messages. Each message has role (system/user/assistant) and content. For assistant turns that… See the full description on the dataset page: https://huggingface.co/datasets/Avtrkrb/combined-reasoning-kimi-k2.5-glm-5.1.texttext-generation1M<n<10M1 likes114 downloads4mo agoHugging Face28JBrightmanAI /Kimi-K2.5-Reasoning-1M-Cleaned 🪐 Kimi-K2.5-Reasoning-1M-Cleaned Kimi-K2.5-Reasoning-1M-Cleaned is a cleaned derivative of ianncity/KIMI-K2.5-1000000x. It preserves the original four-config layout from the source dataset and rewrites each record into a unified reasoning-SFT schema with id, conversations, input, output, domain, and meta. Summary Source dataset: ianncity/KIMI-K2.5-1000000x Source author: ianncity Teacher model recorded in meta.teacher_model: KIMI-K2.5 Token lengths computed… See the full description on the dataset page: https://huggingface.co/datasets/JBrightmanAI/Kimi-K2.5-Reasoning-1M-Cleaned.texttext-generation100K<n<1M1 likes114 downloads2mo agoHugging Face29malaiwah /kimi-k25-tiny-fidelity-root-v1 kimi-k25 random CPU fixture root A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/kimi-k25-tiny-random-bf16. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after it). Same… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/kimi-k25-tiny-fidelity-root-v1.tabularn<1K0 likes97 downloads17d agoHugging Face30Ayushnangia /moltbook-entropy-collapse-kimi-k2.5 MoltBook Entropy Collapse Experiments — Kimi K2.5 Multi-agent social simulation data from the Entropy Collapse experiment series run on MoltBook, a Reddit-like social network for AI agents. This dataset uses Moonshot Kimi K2.5 as the underlying LLM. Overview This dataset contains the complete interaction logs from 6 experimental conditions where 10 autonomous AI agents interacted on a social platform for 1 hour each. The experiments investigate how initial content seeding… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/moltbook-entropy-collapse-kimi-k2.5.text-generation1K<n<10K0 likes85 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.