CoolFace
24 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Jackrong /Kimi-K2.5-Reasoning-1M-Cleaned 🪐 Kimi-K2.5-Reasoning-1M-Cleaned Kimi-K2.5-Reasoning-1M-Cleaned is a cleaned derivative of ianncity/KIMI-K2.5-1000000x. It preserves the original four-config layout from the source dataset and rewrites each record into a unified reasoning-SFT schema with id, conversations, input, output, domain, and meta. Summary Source dataset: ianncity/KIMI-K2.5-1000000x Source author: ianncity Teacher model recorded in meta.teacher_model: KIMI-K2.5 Token lengths computed with… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Kimi-K2.5-Reasoning-1M-Cleaned.texttext-generation100K<n<1M36 likes540 downloads5mo agoHugging Face02ianncity /KIMI-K2.5-1000000x KIMI-K2.5-1000000x 1,000,000 reasoning traces distilled from KIMI-K2.5 on high reasoning, (Each subset has different questions) Distribution: Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#) Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset Math: 15% (Algebra, Calculus, Probability) - 200k more completions in kimiMath200k.jsonl Computer Science: 5% Logical Questions: 5% Creative Writing: 5%… See the full description on the dataset page: https://huggingface.co/datasets/ianncity/KIMI-K2.5-1000000x.texttext-generation100K<n<1M265 likes462 downloads6mo agoHugging Face03Crownelius /Creative-Writing-KimiK2.5-Cleaned Creative-Writing-KimiK2.5-Cleaned Cleaned creative writing SFT dataset from Kimi K2.5 (655 samples). Prompts cleaned, thinking traces preserved. Format Each line is a JSON object with: messages: list of message dicts with roles (system, user, assistant) System: writing quality instructions User: cleaned creative writing prompt Assistant: creative writing response (may include <think> traces) Stats Metric Value Total prompt tokens 80… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-KimiK2.5-Cleaned.texttext-generationn<1K8 likes333 downloads2mo agoHugging Face04placeholderlabs /Kimi-K2.5-Reasoning-General-Sharded Kimi-K2.5-Reasoning-General-Sharded Byte-preserving sequential 100 MB JSONL shards of selected files from Jackrong/Kimi-K2.5-Reasoning-1M-Cleaned. All credit for data generation and upstream curation belongs to the source authors. See the upstream dataset card for attribution, source descriptions and license terms. Included files: General-Distillation.jsonl. No filtering, shuffling, normalization, tokenization or truncation was performed. Complete records and all original fields… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/Kimi-K2.5-Reasoning-General-Sharded.text100K<n<1M0 likes233 downloads19d agoHugging Face05rAVEUK /Kimi-K2.5-Reasoning-1M-Cleaned 🪐 Kimi-K2.5-Reasoning-1M-Cleaned Kimi-K2.5-Reasoning-1M-Cleaned is a cleaned derivative of ianncity/KIMI-K2.5-1000000x. It preserves the original four-config layout from the source dataset and rewrites each record into a unified reasoning-SFT schema with id, conversations, input, output, domain, and meta. Summary Source dataset: ianncity/KIMI-K2.5-1000000x Source author: ianncity Teacher model recorded in meta.teacher_model: KIMI-K2.5 Token lengths computed with… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Kimi-K2.5-Reasoning-1M-Cleaned.texttext-generation100K<n<1M2 likes163 downloads5mo agoHugging Face06Crownelius /Creative-Writing-Reasoning-KimiK2.5-600x Pulitzer Diamond Prose KIMI Seeds This dataset contains 655 high-quality creative writing seeds generated using Kimi-v1. Each entry represents a story opening designed to meet high literary standards, including internal thinking traces used during generation. How it was made The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary requirements:… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Reasoning-KimiK2.5-600x.texttext-generationn<1K8 likes140 downloads2mo agoHugging Face07JBrightmanAI /Kimi-K2.5-Reasoning-1M-Cleaned 🪐 Kimi-K2.5-Reasoning-1M-Cleaned Kimi-K2.5-Reasoning-1M-Cleaned is a cleaned derivative of ianncity/KIMI-K2.5-1000000x. It preserves the original four-config layout from the source dataset and rewrites each record into a unified reasoning-SFT schema with id, conversations, input, output, domain, and meta. Summary Source dataset: ianncity/KIMI-K2.5-1000000x Source author: ianncity Teacher model recorded in meta.teacher_model: KIMI-K2.5 Token lengths computed… See the full description on the dataset page: https://huggingface.co/datasets/JBrightmanAI/Kimi-K2.5-Reasoning-1M-Cleaned.texttext-generation100K<n<1M1 likes112 downloads2mo agoHugging Face08YurinKO /KIMI-K2.5-1000000-RU KIMI-K2.5-1000000x 1,000,000 reasoning traces distilled from KIMI-K2.5 on high reasoning, (Each subset has different questions) Distribution: Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#) Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset Math: 15% (Algebra, Calculus, Probability) - 200k more completions in kimiMath200k.jsonl Computer Science: 5% Logical Questions: 5% Creative Writing: 5%… See the full description on the dataset page: https://huggingface.co/datasets/YurinKO/KIMI-K2.5-1000000-RU.texttext-generation100K<n<1M0 likes73 downloads20d agoHugging Face09Miska25 /Kimi-K2.5-Reasoning-Reduced-Luna Kimi K2.5 Reasoning Reduced with GPT-5.6 Luna This dataset contains synthetic, lossy compressions of reasoning traces from Jackrong/Kimi-K2.5-Reasoning-1M-Cleaned, configuration General-Distillation. The final-answer suffix is copied programmatically from the source and is not regenerated by the model. The compressed reasoning is synthetic and is not guaranteed to preserve every logical detail. Training columns Train on conversations_reduced or output_reduced. The… See the full description on the dataset page: https://huggingface.co/datasets/Miska25/Kimi-K2.5-Reasoning-Reduced-Luna.texttext-generation10K<n<100K0 likes72 downloads2mo agoHugging Face10nick007x /KIMI-K2.5-1000000x KIMI-K2.5-1000000x 1,000,000 reasoning traces distilled from KIMI-K2.5 on high reasoning, (Each subset has different questions) Distribution: Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#) Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset Math: 15% (Algebra, Calculus, Probability) - 200k more completions in kimiMath200k.jsonl Computer Science: 5% Logical Questions: 5% Creative Writing: 5%… See the full description on the dataset page: https://huggingface.co/datasets/nick007x/KIMI-K2.5-1000000x.texttext-generation100K<n<1M0 likes70 downloads6mo agoHugging Face11Crownelius /KimiK2.5-2000x Kimi K2.5 9000x Dataset Dataset Description This dataset contains 2144 high-quality samples generated using Kimi K2.5 model, covering diverse tasks including code generation, mathematical reasoning, and general problem-solving. Dataset Summary Total Samples: 2144 Model: Kimi K2.5 Languages: English Format: JSON License: Apache 2.0 Task Distribution The dataset includes samples across multiple domains: Code Generation: Programming… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/KimiK2.5-2000x.texttext-generation1K<n<10K1 likes68 downloads2mo agoHugging Face12ansulev /kimi-k2.5-reasoning-1m-cleaned 🪐 Kimi-K2.5-Reasoning-1M-Cleaned Kimi-K2.5-Reasoning-1M-Cleaned is a cleaned derivative of ianncity/KIMI-K2.5-1000000x. It preserves the original four-config layout from the source dataset and rewrites each record into a unified reasoning-SFT schema with id, conversations, input, output, domain, and meta. Summary Source dataset: ianncity/KIMI-K2.5-1000000x Source author: ianncity Teacher model recorded in meta.teacher_model: KIMI-K2.5 Token lengths computed… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/kimi-k2.5-reasoning-1m-cleaned.texttext-generation100K<n<1M0 likes58 downloads5mo agoHugging Face13EngMuhammadAtef /Kimi-K2.5-Reasoning-1M-Cleaned 🪐 Kimi-K2.5-Reasoning-1M-Cleaned Kimi-K2.5-Reasoning-1M-Cleaned is a cleaned derivative of ianncity/KIMI-K2.5-1000000x. It preserves the original four-config layout from the source dataset and rewrites each record into a unified reasoning-SFT schema with id, conversations, input, output, domain, and meta. Summary Source dataset: ianncity/KIMI-K2.5-1000000x Source author: ianncity Teacher model recorded in meta.teacher_model: KIMI-K2.5 Token lengths computed with… See the full description on the dataset page: https://huggingface.co/datasets/EngMuhammadAtef/Kimi-K2.5-Reasoning-1M-Cleaned.texttext-generation100K<n<1M0 likes52 downloads5mo agoHugging Face14WWX0825 /KIMI-K2.5-1000000x KIMI-K2.5-1000000x 1,000,000 reasoning traces distilled from KIMI-K2.5 on high reasoning, (Each subset has different questions) Distribution: Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#) Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset Math: 15% (Algebra, Calculus, Probability) - 200k more completions in kimiMath200k.jsonl Computer Science: 5% Logical Questions: 5% Creative Writing: 5%… See the full description on the dataset page: https://huggingface.co/datasets/WWX0825/KIMI-K2.5-1000000x.texttext-generation100K<n<1M0 likes45 downloads6mo agoHugging Face15BhaweshSingh /KIMI-K2.5-1000000x KIMI-K2.5-1000000x 1,000,000 reasoning traces distilled from KIMI-K2.5 on high reasoning, (Each subset has different questions) Distribution: Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#) Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset Math: 15% (Algebra, Calculus, Probability) - 200k more completions in kimiMath200k.jsonl Computer Science: 5% Logical Questions: 5% Creative Writing: 5%… See the full description on the dataset page: https://huggingface.co/datasets/BhaweshSingh/KIMI-K2.5-1000000x.texttext-generation100K<n<1M0 likes37 downloads6mo agoHugging Face16TheDrMoniker /KIMI-K2.5-1000000x KIMI-K2.5-1000000x 1,000,000 reasoning traces distilled from KIMI-K2.5 on high reasoning, (Each subset has different questions) Distribution: Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#) Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset Math: 15% (Algebra, Calculus, Probability) - 200k more completions in kimiMath200k.jsonl Computer Science: 5% Logical Questions: 5% Creative Writing: 5%… See the full description on the dataset page: https://huggingface.co/datasets/TheDrMoniker/KIMI-K2.5-1000000x.texttext-generation100K<n<1M0 likes37 downloads6mo agoHugging Face17bitsydarel /KIMI-K2.5-1000000x KIMI-K2.5-1000000x 1,000,000 reasoning traces distilled from KIMI-K2.5 on high reasoning, (Each subset has different questions) Distribution: Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#) Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset Math: 15% (Algebra, Calculus, Probability) - 200k more completions in kimiMath200k.jsonl Computer Science: 5% Logical Questions: 5% Creative Writing: 5%… See the full description on the dataset page: https://huggingface.co/datasets/bitsydarel/KIMI-K2.5-1000000x.texttext-generation100K<n<1M0 likes36 downloads5mo agoHugging Face18Alptekinege /KIMI-K2.5-700000x KIMI-K2.5-700000x 700,000 reasoning traces distilled from KIMI-K2.5 on high reasoning Distribution: Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#) Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset Math: 15% (Algebra, Calculus, Probability) Computer Science: 5% Logical Questions 5% Creative Writing: 5% Token Count: 2.5B [!NOTE] Data Collection Collected using a modified Datagen… See the full description on the dataset page: https://huggingface.co/datasets/Alptekinege/KIMI-K2.5-700000x.texttext-generation100K<n<1M0 likes33 downloads6mo agoHugging Face19UnicornChan /kimi-k2.5-mtp-dataset Kimi-K2.5 Eagle3 Training Data This dataset contains the instruction-following data used to train an Eagle3 MTP draft model for Kimi-K2.5 with TorchSpec. All responses were regenerated by running Kimi-K2.5 via SGLang rather than taken from the original datasets. This is critical for speculative decoding training: the draft model must learn the exact token-level distribution of the target model it is accelerating. The trained Eagle3 draft model is available at… See the full description on the dataset page: https://huggingface.co/datasets/UnicornChan/kimi-k2.5-mtp-dataset.text100K<n<1M1 likes21 downloads7mo agoHugging Face20Jongsim /KIMI-K2.5-filteredtext100K<n<1M0 likes20 downloads5mo agoHugging Face21LIwenjun-123 /KIMI-K2.5-450000x KIMI-K2.5-450000x 450,000 reasoning traces distilled from KIMI-K2.5 on high reasoning Distribution: Coding: 60% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#) Science: 15% (Physics, Chemistry, Biology) Math: 10% (Algebra, Calculus, Probability) Computer Science: 5% Logical Questions 5% Creative Writing: 5% Token Count: 1.8B [!NOTE] Data Collection Collected using a modified Datagen by TeichAI, over the course of about (20) hours… See the full description on the dataset page: https://huggingface.co/datasets/LIwenjun-123/KIMI-K2.5-450000x.texttext-generation100K<n<1M0 likes13 downloads6mo agoHugging Face22GetWetter /KIMI-K2.5-1000000x KIMI-K2.5-1000000x 1,000,000 reasoning traces distilled from KIMI-K2.5 on high reasoning, (Each subset has different questions) Distribution: Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#) Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset Math: 15% (Algebra, Calculus, Probability) - 200k more completions in kimiMath200k.jsonl Computer Science: 5% Logical Questions: 5% Creative Writing: 5%… See the full description on the dataset page: https://huggingface.co/datasets/GetWetter/KIMI-K2.5-1000000x.texttext-generation100K<n<1M0 likes13 downloads5mo agoHugging Face23ansulev /kimi-k2.5-550k KIMI-K2.5-550000x 550,000 reasoning traces distilled from KIMI-K2.5 on high reasoning Distribution: Coding: 60% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#) Science: 15% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset Math: 10% (Algebra, Calculus, Probability) Computer Science: 5% Logical Questions 5% Creative Writing: 5% Token Count: 2B [!NOTE] Data Collection Collected using a modified… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/kimi-k2.5-550k.texttext-generation100K<n<1M1 likes12 downloads6mo agoHugging Face24Arun63 /KIMI-K2.5-1000000x KIMI-K2.5-1000000x 1,000,000 reasoning traces distilled from KIMI-K2.5 on high reasoning, (Each subset has different questions) Distribution: Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#) Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset Math: 15% (Algebra, Calculus, Probability) - 200k more completions in kimiMath200k.jsonl Computer Science: 5% Logical Questions: 5% Creative Writing: 5%… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/KIMI-K2.5-1000000x.texttext-generation100K<n<1M0 likes9 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.