CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01fxmeng /UltraData-SFT-2605-no-think-8k-32k UltraData-SFT-2605 · no_think · 8k–32k A length-filtered subset of the no_think split of openbmb/UltraData-SFT-2605, containing conversations whose token length falls in the 8k–32k range. This is the medium-length tier intended for standard long-context SFT. Two companion tiers were produced from the same source: Dataset Length range Records this repo — fxmeng/UltraData-SFT-2605-no-think-8k-32k 8k–32k tokens 623,421 fxmeng/UltraData-SFT-2605-no-think-32k-200k… See the full description on the dataset page: https://huggingface.co/datasets/fxmeng/UltraData-SFT-2605-no-think-8k-32k.texttext-generation100K<n<1M0 likes3.2k downloads2mo agoHugging Face02zhiyuanhucs /nemotron-student-fail-v41-clean-thinking Nemotron-fail / DeepSeek-V4.1 clean and action-only trajectories DeepSeek-V4.1 reward-1 trajectories for tasks on which the Nemotron student did not obtain reward 1. This release was rebuilt from the complete reward-1 audit under v54-high-precision-canonical-reconstruction-relations. Training paths Path Rows Unique tasks Thinking Use data/strict/train.jsonl.gz 12 12 Preserved and clean Raw-thinking SFT data/hybrid/train.jsonl.gz 58 58 Only… See the full description on the dataset page: https://huggingface.co/datasets/zhiyuanhucs/nemotron-student-fail-v41-clean-thinking.tabulartext-generationn<1K1 likes1.9k downloads1h agoHugging Face03fxmeng /UltraData-SFT-2605-no-think-32k-200k UltraData-SFT-2605 · no_think · 32k–200k A length-filtered subset of the no_think split of openbmb/UltraData-SFT-2605, containing conversations whose token length falls in the 32k–200k range. This is the long-context tier intended for extended-context SFT. Two companion tiers were produced from the same source: Dataset Length range Records fxmeng/UltraData-SFT-2605-no-think-8k-32k 8k–32k tokens 623,421 this repo — fxmeng/UltraData-SFT-2605-no-think-32k-200k 32k–200k… See the full description on the dataset page: https://huggingface.co/datasets/fxmeng/UltraData-SFT-2605-no-think-32k-200k.texttext-generation10K<n<100K0 likes703 downloads2mo agoHugging Face04ProCreations /grug-think grug-think grug make dataset. dataset make model think like grug. grug think short. short think cheap. cheap think good. big-brain model think 400 token before poke one tool. grug model think 11 word. same tool poke. same work done. many token saved. token = money. grug like money stay in pocket. what in box 100,891 example. every example = full agent conversation: system, user, assistant, tool message. assistant turn always got <think>grug reasoning</think> first… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-think.texttext-generation100K<n<1M35 likes378 downloads3mo agoHugging Face05Davd-b01 /thinking-cap-tier-curricula-complete Thinking Cap Tier Curricula — Complete Reasoning Alignment Suite (TCS v4) [!IMPORTANT] Dataset Release v1.2 (Sept 2026) — Clean Delimiters & Zero-Padding Architecture: In v1.2, all 13,477 SFT samples and 3,187 SimPO preference pairs have undergone an automated token purge: Zero <|pad|> batch residues: 100% eliminated across all files. Zero reasoning leakage into final answers: Deliberation stays strictly inside <think>...</think>, and answers provide direct, non-repetitive… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinking-cap-tier-curricula-complete.texttext-generation10K<n<100K0 likes299 downloads10d agoHugging Face06Davd-b01 /thinking-cap-tier-lima-dense Thinking Cap Tier Curricula — LIMA Hyper-Dense Reasoning Alignment Suite (TCS v4) [!IMPORTANT] Dataset Release v1.2 (Sept 2026) — Clean Delimiters & Zero-Padding Architecture: In v1.2, all 5,500 SFT and 2,000 SimPO records have undergone a complete token purge: Zero <|pad|> batch residues: 100% eliminated across all records. Zero reasoning leakage into final answers: Deliberation stays strictly inside <think>...</think>, and answers provide direct conclusions. Native ChatML… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinking-cap-tier-lima-dense.texttext-generation1K<n<10K2 likes225 downloads10d agoHugging Face07Davd-b01 /thinking-cap-tier-raw-traces Thinking Cap Tier Raw Traces (TCS v4) [!IMPORTANT] Dataset Release v1.2 (Sept 2026) — Clean Delimiters & Zero-Padding Architecture: All 38,158 candidate reasoning traces across all 4 tiers (candidates_low.jsonl, candidates_mid.jsonl, candidates_high.jsonl, candidates_xhigh.jsonl) are 100% sanitized: Zero batch-padding residues (<|pad|>): Completely purged across all records. Strict Delimiter Integrity: Generation blocks cleanly separate thought deliberation tags… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinking-cap-tier-raw-traces.tabulartext-generation10K<n<100K0 likes203 downloads10d agoHugging Face08Davd-b01 /thinkingcap-condensed-qwen3.8-glm5.2-kimi-k3 ThinkingCap Condensed — Qwen3.8 / GLM-5.2 / Kimi-K3 Condensed ThinkingCap-style reasoning traces for SFT. 1,985 traces: each row pairs a full multi-turn teacher trace (Qwen3.8-Max, GLM-5.2 or Kimi K3, via r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillation) with a condensed TC-style version (short <think> + definitive numbered answer) generated by bottlecapai/ThinkingCap-Qwen3.6-27B using the thinkingcap system prompt. Format: JSONL (data/condensed.jsonl), 1,985 rows, UTF-8.… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinkingcap-condensed-qwen3.8-glm5.2-kimi-k3.texttext-generation1K<n<10K1 likes187 downloads1mo agoHugging Face09AlpachinoNLP /CT-RATE-Thinking CT-RATE-Thinking: Reasoning-Augmented CT Report Dataset 🎉🎉🎉 Our paper was accepted at the 28th conference of The Medical Image Computing and Computer Assisted Intervention Society (MICCAI). See you in Daejeon, Korea, September 23–27, 2025.CT-RATE-Thinking is a reasoning-augmented dataset derived from CT-RATE, containing chain-of-thought VQA pairs and report-level thinking narratives for 3D chest CT volumes. It was generated as part of the μ²Tokenizer project… See the full description on the dataset page: https://huggingface.co/datasets/AlpachinoNLP/CT-RATE-Thinking.textvisual-question-answering1M<n<10M2 likes159 downloads5mo agoHugging Face10ThinkNet /HQ-Chat-2k 🧠 HQ-Chat-2K — High-Quality Conversational & Instruction-Tuning Dataset 2,000 carefully curated, high-quality conversation and instruction examples for fine-tuning Small Language Models (SLMs) and compact LLMs from ~500M to 3B parameters. HQ-Chat-2K is a high-quality conversational and instruction-tuning dataset designed specifically for training and fine-tuning small to medium-sized Large Language Models (LLMs). The dataset contains 2,000 curated user–assistant examples… See the full description on the dataset page: https://huggingface.co/datasets/ThinkNet/HQ-Chat-2k.texttext-generation1K<n<10K5 likes159 downloads14h agoHugging Face11Davd-b01 /thinkingcap-reasoning-traces ThinkingCap Reasoning Traces (Legacy v1 Prototype) [!WARNING] Legacy / Deprecated Prototype Notice (v1): This dataset represents an early exploratory prototype (v1, 4,254 traces) from initial development. Some samples in this legacy version contain early formatting artifacts, including reasoning traces leaking into the final answer field and informal step-by-step breakdowns. For modern post-training, SFT, and SimPO alignment under the TCS v4 cognitive standard, please use our… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinkingcap-reasoning-traces.texttext-generation1K<n<10K1 likes152 downloads10d agoHugging Face12dougalldeepmind /2026-08-02-qwen36-mixture-500k-numina-heavy-empty-think Qwen3.6-27B SFT mixture — 500k maths-weighted, empty-think markers 499,595 tokens across 1,001 conversations, weighted toward maths, with Qwen3.6's empty think marker on the non-maths rows. md5 c433f31eba2b5b4919fb166043caccb5. Source Examples Tokens Share Marker NuminaMath-CoT 611 333,351 66.9% no No Robots 271 82,239 16.5% yes TULU3 119 82,445 16.5% yes Total 1,001 499,595 390 marked Derived from qwen3.6-27b-mixture-500k-numina-heavy by adding the… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-02-qwen36-mixture-500k-numina-heavy-empty-think.texttext-generation1K<n<10K0 likes139 downloads23d agoHugging Face13Jackrong /Chinese-Qwen3-235B-Thinking-2507-Distill-100k 📌 Note: The English translation of this dataset card is provided below. Chinese-Qwen3-235B-Thinking-2507-Distill-100k Dataset Summary Chinese-Qwen3-235B-Thinking-2507-Distill-100k 是一个包含约 100k 条高质量中文推理与指令数据的数据集,由 Qwen-3-235B-A22B-Thinking-2507(官方 Thinking 模式,上下文长度 32K)蒸馏生成。 该数据集覆盖了多个重要领域: 数学与工程任务(Mathematics, Applied Math, Advanced Math) 通用知识与写作(General Knowledge, Language & Writing) 技术与编程(Technology & Programming) 商业与经济(Business & Economics)… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Chinese-Qwen3-235B-Thinking-2507-Distill-100k.tabulartext-classification100K<n<1M19 likes128 downloads1y agoHugging Face14agentlans /TeichAI-thinking-reasoning-x TeichAI Thinking & Reasoning Datasets A collection of prompts answered by large language models (LLMs) such as Google Gemini and OpenAI ChatGPT, with long-form reasoning enabled. These datasets were originally created by TeichAI for distillation and reasoning-focused training workflows. Schema Each row in the dataset has the following fields: question_hash: Truncated, base64-encoded MD5 hash of the question, useful for filtering and deduplication. question: The… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/TeichAI-thinking-reasoning-x.texttext-generation100K<n<1M1 likes127 downloads5mo agoHugging Face15ProCreations /grug-think-v3-10k grug-think-v3-10k v2 brain short. v2 brain useful. but some v2 brain wear office shirt. "User wants hello world Python. Provide code." short English, yes. grug, no. v3 tear off office shirt. keep brain meat. old: User wants hello world Python. Simple code snippet, no tools needed. Provide code and brief explanation. new: Need Python hello-world. Tiny snippet. No tool. Give code, brief explain. complex cave different. grug no crush branch into pebble. exact path, error… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-think-v3-10k.texttext-generation10K<n<100K12 likes110 downloads2mo agoHugging Face16Sizhe-Chen /Qwen3.8-27B-Thinking-SecOPD-trainset Qwen3.6-27B-Thinking SecOPD Trainset Dataset summary This public dataset contains 19,155 complete, model-specific preference records for offline adversarial training against indirect prompt injection. The corpus starts from the 19,157-record Sizhe-Chen/Qwen3.6-27B-Instruct-SecPO-trainset release. Its six non-label lineage fields are retained, while the attacked prompts are rendered for thinking-on generation and the chosen and rejected labels are regenerated… See the full description on the dataset page: https://huggingface.co/datasets/Sizhe-Chen/Qwen3.8-27B-Thinking-SecOPD-trainset.texttext-generation10K<n<100K1 likes107 downloads23d agoHugging Face17JessieWei /GLM-5.2-FP8-nemotron-codealpaca-thinking GLM-5.2-FP8 Nemotron-CodeAlpaca Thinking Dataset 820,790 single-turn conversations generated by zai-org/GLM-5.2-FP8 with thinking enabled. Prompt source Rows (public) Nemotron-Post-Training-Dataset-v2 800,944 CodeAlpaca-20k (corrected prompts, instruction + "\n\n" + input) 19,846 Total 820,790 Generation: temperature=1.0, top_p=0.95, max_tokens=24576, thinking enabled. The CodeAlpaca prompts here include the input field. Relationship to… See the full description on the dataset page: https://huggingface.co/datasets/JessieWei/GLM-5.2-FP8-nemotron-codealpaca-thinking.texttext-generation100K<n<1M0 likes103 downloads2mo agoHugging Face18Thinking-Space /OpenThought3-Qwen3-4BOpenThought3-Qwen3-4B OpenThought3-Qwen3-4B is a math reasoning supervised fine-tuning dataset in chat-message JSONL format. Data Creation and Cleaning This dataset was generated by Qwen3-4B (Non-thinking) from math-domain prompts selected from OpenThoughts3-1.2M. The generated responses were cleaned through deduplication, removal of degenerate repetition/repeater-style outputs, and template checks on the assistant… See the full description on the dataset page: https://huggingface.co/datasets/Thinking-Space/OpenThought3-Qwen3-4B.texttext-generation100K<n<1M3 likes98 downloads5mo agoHugging Face19cudabenchmarktest /r8-thinking-fix-sft ⚠️ CRITICAL: Ollama Inference Flag Required for derived models If you train or serve any Qwen3.5-9B-derived model from this lineage via Ollama, you MUST pass "think": false in /api/chat requests for chat / instruction following / tool use. The qwen3.5 RENDERER auto-injects <think> tags causing 25-46% empty-answer rates without this flag. See dataset cudabenchmarktest/r9-research-framework/_OLLAMA_INFERENCE_WARNING.md for the full lesson learned. R8 Thinking-Fix SFT… See the full description on the dataset page: https://huggingface.co/datasets/cudabenchmarktest/r8-thinking-fix-sft.texttext-generation1K<n<10K0 likes83 downloads5mo agoHugging Face20qwqeqw /Dataset_of_Russian_thinkingRu RTD Описание:Russian Thinking Dataset — это набор данных, предназначенный для обучения и тестирования моделей обработки естественного языка (NLP) на русском языке. Датасет ориентирован на задачи, связанные с генерацией текста, анализом диалогов и решением математических и логических задач. Основная информация: Сплит: train Количество записей: 147.046 Цели: Обучение моделей пониманию русского языка. Создание диалоговых систем с естественным взаимодействием.… See the full description on the dataset page: https://huggingface.co/datasets/qwqeqw/Dataset_of_Russian_thinking.texttext-generation100K<n<1M1 likes75 downloads10mo agoHugging Face21chimbiwide /NPC-RP-Post-Thinking AIIDE-POST-THINKING This is the post-thinking dataset for our paper accepted as a poster presentation to AIIDE-2026 Dataset Details This is the training dataset for chimbiwide/Gemma3-4B-post-thinking Corresponding Links Repository: [To be updated] Paper: [To be updated] Demo: [To be updated] Uses Suprevised-Finetuning Dataset Creation For more details, consult our paper. Citation If you… See the full description on the dataset page: https://huggingface.co/datasets/chimbiwide/NPC-RP-Post-Thinking.texttext-generation1K<n<10K0 likes61 downloads29d agoHugging Face22islam-kamel /MBPP-Thinking-Gate-1k MBPP Thinking-Gate SFT Dataset This package contains two related assets: Ready 1,000-row MBPP-style dataset (all.jsonl, train.jsonl, validation.jsonl). It is synthetic and designed to test/train autonomous routing between <DIRECT> and <THINK>. Official-MBPP builder (build_from_official_mbpp.py). Run this to create the production dataset from the official Google Research MBPP source. Why two response modes? The training target starts with one of two routing… See the full description on the dataset page: https://huggingface.co/datasets/islam-kamel/MBPP-Thinking-Gate-1k.tabulartext-generation1K<n<10K0 likes61 downloads3d agoHugging Face23PJMixers-Dev /allenai_WildChat-1M-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT allenai_WildChat-1M-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT PJMixers-Dev/allenai_WildChat-1M-prompts with responses generated with gemini-2.0-flash-thinking-exp-1219. Generation Details If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped. If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped. If ["candidates"][0]["finish_reason"] != 1 the sample was skipped. model =… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/allenai_WildChat-1M-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.texttext-generation10K<n<100K6 likes60 downloads2y agoHugging Face24sbussiso /synthetic-self-correction-and-thinking-samples Self Correction and Thinking A seed library for training language models to reason with self-correction. Teaches three reasoning behaviors -- catching your own errors, verifying correct answers, and rejecting false doubts -- across four domains, three difficulty tiers, and three reasoning modes. Also includes multi-turn user-correction conversations where the user actively corrects or challenges the assistant. The structure at a glance graph TB… See the full description on the dataset page: https://huggingface.co/datasets/sbussiso/synthetic-self-correction-and-thinking-samples.imagetext-generation1K<n<10K0 likes58 downloads1mo agoHugging Face25WrittenWithRust /Strandset-Rust-Think-TR 🦀 Strandset-Rust-Think-TR (5K Cleaned & Translated) Strandset-Rust-Think-TR, Rust programlama dili odaklı, Türkçe düşünme zinciri (Chain-of-Thought / <think>) adımları içeren 5.000 adet yüksek kaliteli talimat (instruction-tuning) örneğinden oluşan bir veri setidir. Bu veri seti, snowmead/Strandset-Rust-Think çalışması temel alınarak WrittenWithRust tarafından Qwen3.8-27B modeli yardımıyla Türkçe dikeyine kazandırılmış ve mükerrer kayıtlarından arındırılmıştır. ⚙️… See the full description on the dataset page: https://huggingface.co/datasets/WrittenWithRust/Strandset-Rust-Think-TR.texttext-generation1K<n<10K0 likes58 downloads1mo agoHugging Face26Pinkstack /Thinking-multilingual-big-10k-sft A dataset based off of openo1 math, 500 examples translated to 23 different languages. filtered out un-translated examples. enjoy 👍 texttext-generation10K<n<100K3 likes55 downloads2y agoHugging Face27ENERZAiKR /agentic-think-v1gated agentic-think-v1 Agentic tool-calling SFT corpus with teacher-generated <think> reasoning on every assistant round. 212,396 training rows across 9 sources, built by ENERZAi for small-model (1.7B ternary) agentic SFT. Each row is a full multi-turn conversation: system prompt, user turns, assistant turns (each carrying its reasoning in a separate reasoning_content field), tool calls and tool results. Per-source splits (by_source config) Each source file is also… See the full description on the dataset page: https://huggingface.co/datasets/ENERZAiKR/agentic-think-v1.texttext-generation100K<n<1M0 likes55 downloads8d agoHugging Face28PJMixers-Images /bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT bghira/pseudo-camera-10k with responses/captions generated with gemini-2.0-flash-thinking-exp-1219. The format should be similar to that of liuhaotian/LLaVA-Instruct-150K. Images can be found in the images.zip folder. The zip also contains .txt captions for ease of use in non-VQA tasks. Generation Details If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Images/bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.imagetext-generation1K<n<10K1 likes54 downloads2y agoHugging Face29laion /sft-repro-thinking-step630-nemotron-terminal-step1888-openthoughts-tblite-2026-08-13 Nemotron Terminal SFT reproduction evaluation artifacts This repository contains the complete Harbor artifact tree for the 300-trial OpenThoughts-TBLite evaluation of laion/sft-repro-thinking-step630-nemotron-terminal-step1888. The checkpoint was trained from the Grug stage-2 thinking checkpoint on the Nemotron Terminal corpus for 1,888 steps. Result Measure Value Attempted / completed 300 / 300 Verifier-scoreable 259 (86.33%) Aggregate reward, all… See the full description on the dataset page: https://huggingface.co/datasets/laion/sft-repro-thinking-step630-nemotron-terminal-step1888-openthoughts-tblite-2026-08-13.texttext-generationn<1K0 likes51 downloads1mo agoHugging Face30MasonMac /CodeX-Thinking-Gemma-4-31B-ITAll prompts were taken from Modotte/CodeX-2M-Thinking, which contains multiple traces per prompt whereas this dataset only provides one trace per prompt. Generations were with https://huggingface.co/nvidia/Gemma-4-31B-IT-NVFP4 (a mix of BF16/FP8 weights that NVIDIA configured with FP8 KV cache; benchmarks show performs similarly to BF16 for coding). No system prompt was used. text-generation100K<n<1M0 likes50 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.