datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Kimi-K2.5-Reasoning-1M-Cleaned
🪐 Kimi-K2.5-Reasoning-1M-Cleaned
Kimi-K2.5-Reasoning-1M-Cleaned is a cleaned derivative of ianncity/KIMI-K2.5-1000000x. It preserves the original four-config layout from the source dataset and rewrites each record into a unified reasoning-SFT schema with id, conversations, input, output, domain, and meta.
Summary
Source dataset: ianncity/KIMI-K2.5-1000000x
Source author: ianncity
Teacher model recorded in meta.teacher_model: KIMI-K2.5
Token lengths computed with… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Kimi-K2.5-Reasoning-1M-Cleaned.KIMI-K2.5-1000000x
KIMI-K2.5-1000000x
1,000,000 reasoning traces distilled from KIMI-K2.5 on high reasoning, (Each subset has different questions)
Distribution:
Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#)
Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset
Math: 15% (Algebra, Calculus, Probability) - 200k more completions in kimiMath200k.jsonl
Computer Science: 5%
Logical Questions: 5%
Creative Writing: 5%… See the full description on the dataset page: https://huggingface.co/datasets/ianncity/KIMI-K2.5-1000000x.Creative-Writing-KimiK2.5-Cleaned
Creative-Writing-KimiK2.5-Cleaned
Cleaned creative writing SFT dataset from Kimi K2.5 (655 samples). Prompts cleaned, thinking traces preserved.
Format
Each line is a JSON object with:
messages: list of message dicts with roles (system, user, assistant)
System: writing quality instructions
User: cleaned creative writing prompt
Assistant: creative writing response (may include <think> traces)
Stats
Metric
Value
Total prompt tokens
80… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-KimiK2.5-Cleaned.combined-reasoning-opus-4.6-opus-4.7-kimi-k2.5-kimi-k2.6-glm-5.1
Combined Reasoning Distill — Multi-Model
A large-scale unified reasoning dataset combining thinking and chain-of-thought traces distilled from frontier models, normalized into a single consistent schema for fine-tuning. Includes data from Claude (Opus 4.5/4.6/4.7, Sonnet 4.5/4.6, Haiku 4.5), GPT (5.1/5.2), Gemini 3 Pro Preview, Kimi (K2/K2.5/K2.6), GLM (4.6/4.7/5.1), MiniMax M2.1, Grok Code Fast 1, and more.
Schema
Every row has a single field:
Field
Type… See the full description on the dataset page: https://huggingface.co/datasets/Avtrkrb/combined-reasoning-opus-4.6-opus-4.7-kimi-k2.5-kimi-k2.6-glm-5.1.Kimi-K2.5-Reasoning-1M-Cleaned
🪐 Kimi-K2.5-Reasoning-1M-Cleaned
Kimi-K2.5-Reasoning-1M-Cleaned is a cleaned derivative of ianncity/KIMI-K2.5-1000000x. It preserves the original four-config layout from the source dataset and rewrites each record into a unified reasoning-SFT schema with id, conversations, input, output, domain, and meta.
Summary
Source dataset: ianncity/KIMI-K2.5-1000000x
Source author: ianncity
Teacher model recorded in meta.teacher_model: KIMI-K2.5
Token lengths computed with… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Kimi-K2.5-Reasoning-1M-Cleaned.Creative-Writing-Reasoning-KimiK2.5-600x
Pulitzer Diamond Prose KIMI Seeds
This dataset contains 655 high-quality creative writing seeds generated using Kimi-v1.
Each entry represents a story opening designed to meet high literary standards, including internal thinking traces used during generation.
How it was made
The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary requirements:… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Reasoning-KimiK2.5-600x.Kimi-K2.7-CodingTraces-9000x
Kimi K2.7 Coding Traces 9000x
A validated 9,014-row coding and software-engineering reasoning
dataset generated with moonshotai/Kimi-K2.7-Code.
Every row contains a coding-focused prompt, a separated reasoning trace, and a
final answer. The release was built from a durable Google Drive generation
pipeline and underwent a complete two-pass schema and delimiter audit before
publication.
Generation configuration
Setting
Value
Teacher… See the full description on the dataset page: https://huggingface.co/datasets/trjxter/Kimi-K2.7-CodingTraces-9000x.combined-reasoning-kimi-k2.5-glm-5.1
Combined Reasoning Distill — Multi-Model
A large-scale unified reasoning dataset combining thinking and chain-of-thought traces distilled from frontier models, normalized into a single consistent schema for fine-tuning. Includes data from Kimi K2.5 & GLM 5.1.
Schema
Every row has a single field:
Field
Type
Description
messages
list[dict]
Conversation messages. Each message has role (system/user/assistant) and content.
For assistant turns that… See the full description on the dataset page: https://huggingface.co/datasets/Avtrkrb/combined-reasoning-kimi-k2.5-glm-5.1.Kimi-K2.5-Reasoning-1M-Cleaned
🪐 Kimi-K2.5-Reasoning-1M-Cleaned
Kimi-K2.5-Reasoning-1M-Cleaned is a cleaned derivative of ianncity/KIMI-K2.5-1000000x. It preserves the original four-config layout from the source dataset and rewrites each record into a unified reasoning-SFT schema with id, conversations, input, output, domain, and meta.
Summary
Source dataset: ianncity/KIMI-K2.5-1000000x
Source author: ianncity
Teacher model recorded in meta.teacher_model: KIMI-K2.5
Token lengths computed… See the full description on the dataset page: https://huggingface.co/datasets/JBrightmanAI/Kimi-K2.5-Reasoning-1M-Cleaned.Kimi-K2.6-Reasoning-3300x-WandB
Kimi-K2.6-Reasoning-3300x-WandB
Kimi-K2.6-Reasoning-3300x-WandB is a W&B-only synthetic reasoning dataset generated with Kimi-K2.6 through Weights & Biases Inference.
This dataset is the pure W&B-generated subset from a larger planned 8,000-example Kimi reasoning distillation run. Generation stopped when the W&B quota limit was reached, and the completed accepted rows were audited, cleaned, and exported as a standalone dataset.
This release contains 3,303 accepted W&B-generated rows… See the full description on the dataset page: https://huggingface.co/datasets/trjxter/Kimi-K2.6-Reasoning-3300x-WandB.KIMI-K2.5-1000000-RU
KIMI-K2.5-1000000x
1,000,000 reasoning traces distilled from KIMI-K2.5 on high reasoning, (Each subset has different questions)
Distribution:
Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#)
Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset
Math: 15% (Algebra, Calculus, Probability) - 200k more completions in kimiMath200k.jsonl
Computer Science: 5%
Logical Questions: 5%
Creative Writing: 5%… See the full description on the dataset page: https://huggingface.co/datasets/YurinKO/KIMI-K2.5-1000000-RU.Kimi-K2.5-Reasoning-Reduced-Luna
Kimi K2.5 Reasoning Reduced with GPT-5.6 Luna
This dataset contains synthetic, lossy compressions of reasoning traces from
Jackrong/Kimi-K2.5-Reasoning-1M-Cleaned,
configuration General-Distillation.
The final-answer suffix is copied programmatically from the source and is not
regenerated by the model. The compressed reasoning is synthetic and is not
guaranteed to preserve every logical detail.
Training columns
Train on conversations_reduced or output_reduced. The… See the full description on the dataset page: https://huggingface.co/datasets/Miska25/Kimi-K2.5-Reasoning-Reduced-Luna.KIMI-K2.5-1000000x
KIMI-K2.5-1000000x
1,000,000 reasoning traces distilled from KIMI-K2.5 on high reasoning, (Each subset has different questions)
Distribution:
Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#)
Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset
Math: 15% (Algebra, Calculus, Probability) - 200k more completions in kimiMath200k.jsonl
Computer Science: 5%
Logical Questions: 5%
Creative Writing: 5%… See the full description on the dataset page: https://huggingface.co/datasets/nick007x/KIMI-K2.5-1000000x.KimiK2.5-2000x
Kimi K2.5 9000x Dataset
Dataset Description
This dataset contains 2144 high-quality samples generated using Kimi K2.5 model, covering diverse tasks including code generation, mathematical reasoning, and general problem-solving.
Dataset Summary
Total Samples: 2144
Model: Kimi K2.5
Languages: English
Format: JSON
License: Apache 2.0
Task Distribution
The dataset includes samples across multiple domains:
Code Generation: Programming… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/KimiK2.5-2000x.kimi-k2.5-reasoning-1m-cleaned
🪐 Kimi-K2.5-Reasoning-1M-Cleaned
Kimi-K2.5-Reasoning-1M-Cleaned is a cleaned derivative of ianncity/KIMI-K2.5-1000000x. It preserves the original four-config layout from the source dataset and rewrites each record into a unified reasoning-SFT schema with id, conversations, input, output, domain, and meta.
Summary
Source dataset: ianncity/KIMI-K2.5-1000000x
Source author: ianncity
Teacher model recorded in meta.teacher_model: KIMI-K2.5
Token lengths computed… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/kimi-k2.5-reasoning-1m-cleaned.KIMI-K2.5-Reasoning
OctoMed/KIMI-K2.5-Reasoning
Multi-turn chain-of-thought conversations converted to OctoMed format for SFT training.
Source
Derived from ianncity/KIMI-K2.5-1000000x
by ianncity. All credit for the original data collection and
distillation goes to the original authors.
Format
Each example contains:
question: User question
responses: the final gpt turn repeated for compatibility with the OctoMed pipeline
Kimi-K2.5-Reasoning-1M-Cleaned
🪐 Kimi-K2.5-Reasoning-1M-Cleaned
Kimi-K2.5-Reasoning-1M-Cleaned is a cleaned derivative of ianncity/KIMI-K2.5-1000000x. It preserves the original four-config layout from the source dataset and rewrites each record into a unified reasoning-SFT schema with id, conversations, input, output, domain, and meta.
Summary
Source dataset: ianncity/KIMI-K2.5-1000000x
Source author: ianncity
Teacher model recorded in meta.teacher_model: KIMI-K2.5
Token lengths computed with… See the full description on the dataset page: https://huggingface.co/datasets/EngMuhammadAtef/Kimi-K2.5-Reasoning-1M-Cleaned.Kimi-K2.6-Technical-Reasoning-AddOn-3300x
Kimi-K2.6-Technical-Reasoning-AddOn-3300x
This dataset is a technical reasoning add-on dataset generated with Kimi K2.6 as the teacher model.
The dataset was designed as an additional technical reasoning trace set for downstream SFT experiments, especially around math, graduate-level science, coding, and debugging/code-repair style prompts.
Dataset Summary
Dataset name: Kimi-K2.6-Technical-Reasoning-AddOn-3300x
Teacher model: Kimi-K2.6
Backend: W&B… See the full description on the dataset page: https://huggingface.co/datasets/trjxter/Kimi-K2.6-Technical-Reasoning-AddOn-3300x.seta-sft-kimi-k2.5-nothink
Seta SFT — Kimi K2.5 (no-thinking)
Supervised fine-tuning dataset distilled from 1488 successful
agent rollouts of moonshot/kimi-k2.5 on the
seta-env-v2
terminal-agent benchmark, tokenized with the Qwen/Qwen3-8B chat template
and ready for AREAL FSDPLMEngine SFT training.
Schema
Each row preserves the full per-trial diagnostic record from the build
pipeline so consumers can inspect, filter, or re-tokenize without rerunning
the rollouts:
column
type
meaning… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/seta-sft-kimi-k2.5-nothink.KIMI-K2.5-1000000x
KIMI-K2.5-1000000x
1,000,000 reasoning traces distilled from KIMI-K2.5 on high reasoning, (Each subset has different questions)
Distribution:
Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#)
Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset
Math: 15% (Algebra, Calculus, Probability) - 200k more completions in kimiMath200k.jsonl
Computer Science: 5%
Logical Questions: 5%
Creative Writing: 5%… See the full description on the dataset page: https://huggingface.co/datasets/WWX0825/KIMI-K2.5-1000000x.swebench-verified-kimi-k2p6-traces
SWE-bench Verified Kimi K2.6 Reasoning Traces
This dataset contains reasoning traces generated on princeton-nlp/SWE-bench_Verified using fireworks_ai/kimi-k2p6-high with a mini-swe-agent based harness. It is intended for research and distillation of software-engineering agents.
The repository is published with three configs because each table has a different schema:
raw_trajectories: one row per SWE-bench instance with the patch, sanitized result JSON, full trajectory JSON, message… See the full description on the dataset page: https://huggingface.co/datasets/MemoryAsModality/swebench-verified-kimi-k2p6-traces.KIMI-K2.5-1000000x
KIMI-K2.5-1000000x
1,000,000 reasoning traces distilled from KIMI-K2.5 on high reasoning, (Each subset has different questions)
Distribution:
Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#)
Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset
Math: 15% (Algebra, Calculus, Probability) - 200k more completions in kimiMath200k.jsonl
Computer Science: 5%
Logical Questions: 5%
Creative Writing: 5%… See the full description on the dataset page: https://huggingface.co/datasets/BhaweshSingh/KIMI-K2.5-1000000x.KIMI-K2.5-1000000x
KIMI-K2.5-1000000x
1,000,000 reasoning traces distilled from KIMI-K2.5 on high reasoning, (Each subset has different questions)
Distribution:
Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#)
Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset
Math: 15% (Algebra, Calculus, Probability) - 200k more completions in kimiMath200k.jsonl
Computer Science: 5%
Logical Questions: 5%
Creative Writing: 5%… See the full description on the dataset page: https://huggingface.co/datasets/TheDrMoniker/KIMI-K2.5-1000000x.KIMI-K2.5-1000000x
KIMI-K2.5-1000000x
1,000,000 reasoning traces distilled from KIMI-K2.5 on high reasoning, (Each subset has different questions)
Distribution:
Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#)
Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset
Math: 15% (Algebra, Calculus, Probability) - 200k more completions in kimiMath200k.jsonl
Computer Science: 5%
Logical Questions: 5%
Creative Writing: 5%… See the full description on the dataset page: https://huggingface.co/datasets/bitsydarel/KIMI-K2.5-1000000x.KIMI-K2.5-700000x
KIMI-K2.5-700000x
700,000 reasoning traces distilled from KIMI-K2.5 on high reasoning
Distribution:
Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#)
Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset
Math: 15% (Algebra, Calculus, Probability)
Computer Science: 5%
Logical Questions 5%
Creative Writing: 5%
Token Count: 2.5B
[!NOTE]
Data Collection
Collected using a modified Datagen… See the full description on the dataset page: https://huggingface.co/datasets/Alptekinege/KIMI-K2.5-700000x.Kimi-K2.6-Thinking-200x-Cleaned
Kimi K2.6 Thinking 200x — Cleaned
Cleaned version of uniquealexx/Kimi-K2.6-Thinking-200x.
Format — ShareGPT
Each row has a single conversations field in ShareGPT format:
{
"conversations": [
{"from": "human", "value": "Find all positive integers n such that..."},
{"from": "gpt", "value": "<thinking>
I need to find...
</thinking>
The answer is..."}
]
}
Usage with Unsloth
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/rex099/Kimi-K2.6-Thinking-200x-Cleaned.ianncity_KIMI-K2.5-1000000x
KIMI-K2.5-1000000x
1,000,000 reasoning traces distilled from KIMI-K2.5 on high reasoning, (Each subset has different questions)
Distribution:
Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#)
Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset
Math: 15% (Algebra, Calculus, Probability) - 200k more completions in kimiMath200k.jsonl
Computer Science: 5%
Logical Questions: 5%
Creative Writing: 5%… See the full description on the dataset page: https://huggingface.co/datasets/DijkstraFTW/ianncity_KIMI-K2.5-1000000x.KIMI-K2.5-450000x
KIMI-K2.5-450000x
450,000 reasoning traces distilled from KIMI-K2.5 on high reasoning
Distribution:
Coding: 60% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#)
Science: 15% (Physics, Chemistry, Biology)
Math: 10% (Algebra, Calculus, Probability)
Computer Science: 5%
Logical Questions 5%
Creative Writing: 5%
Token Count: 1.8B
[!NOTE]
Data Collection
Collected using a modified Datagen by TeichAI, over the course of about (20) hours… See the full description on the dataset page: https://huggingface.co/datasets/LIwenjun-123/KIMI-K2.5-450000x.KIMI-K2.5-1000000x
KIMI-K2.5-1000000x
1,000,000 reasoning traces distilled from KIMI-K2.5 on high reasoning, (Each subset has different questions)
Distribution:
Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#)
Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset
Math: 15% (Algebra, Calculus, Probability) - 200k more completions in kimiMath200k.jsonl
Computer Science: 5%
Logical Questions: 5%
Creative Writing: 5%… See the full description on the dataset page: https://huggingface.co/datasets/GetWetter/KIMI-K2.5-1000000x.kimi-k2.5-550k
KIMI-K2.5-550000x
550,000 reasoning traces distilled from KIMI-K2.5 on high reasoning
Distribution:
Coding: 60% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#)
Science: 15% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset
Math: 10% (Algebra, Calculus, Probability)
Computer Science: 5%
Logical Questions 5%
Creative Writing: 5%
Token Count: 2B
[!NOTE]
Data Collection
Collected using a modified… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/kimi-k2.5-550k.
