CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01fxmeng /UltraData-SFT-2605-no-think-8k-32k UltraData-SFT-2605 · no_think · 8k–32k A length-filtered subset of the no_think split of openbmb/UltraData-SFT-2605, containing conversations whose token length falls in the 8k–32k range. This is the medium-length tier intended for standard long-context SFT. Two companion tiers were produced from the same source: Dataset Length range Records this repo — fxmeng/UltraData-SFT-2605-no-think-8k-32k 8k–32k tokens 623,421 fxmeng/UltraData-SFT-2605-no-think-32k-200k… See the full description on the dataset page: https://huggingface.co/datasets/fxmeng/UltraData-SFT-2605-no-think-8k-32k.texttext-generation100K<n<1M0 likes3.2k downloads2mo agoHugging Face02zhiyuan218 /Think-Bench THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Official repository for "THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models". For more details, please refer to the project page with dataset exploration and visualization tools. [Paper] [Github] [ModelScope Dataset] [Visualization] 👀 About Think-Bench Reasoning models have made remarkable progress in complex tasks… See the full description on the dataset page: https://huggingface.co/datasets/zhiyuan218/Think-Bench.textquestion-answering1K<n<10K2 likes781 downloads1y agoHugging Face03fxmeng /UltraData-SFT-2605-no-think-32k-200k UltraData-SFT-2605 · no_think · 32k–200k A length-filtered subset of the no_think split of openbmb/UltraData-SFT-2605, containing conversations whose token length falls in the 32k–200k range. This is the long-context tier intended for extended-context SFT. Two companion tiers were produced from the same source: Dataset Length range Records fxmeng/UltraData-SFT-2605-no-think-8k-32k 8k–32k tokens 623,421 this repo — fxmeng/UltraData-SFT-2605-no-think-32k-200k 32k–200k… See the full description on the dataset page: https://huggingface.co/datasets/fxmeng/UltraData-SFT-2605-no-think-32k-200k.texttext-generation10K<n<100K0 likes703 downloads2mo agoHugging Face04EeeLM /llm-jp-4-thinking-sft-data-chatmlllm-jpのデータセットllm-jp-4-thinking-sft-dataを、 ChatML形式に変換したものです。 ライセンス 各サンプルのライセンスは、元データセットカードに記載された各データソースのライセンスに従います。 本リポジトリは、元となったデータ全体に対して新たなライセンスを付与するものではありません。 利用する場合は、対応する元データソースのライセンス条件を確認してください。 text1M<n<10M0 likes477 downloads4mo agoHugging Face05agentlans /epic-thinking Source Rows glaiveai/reasoning-v1-20m 1 999 793 BAAI/OpenSeek-Synthetic-Reasoning-Data-Examples CC 1 267 534 PrimeIntellect/INTELLECT-3-SFT openreasoning_science 1 000 000 PrimeIntellect/INTELLECT-3-SFT am_chat 852 816 nvidia/Nemotron-Cascade-SFT-Stage-1 general 583 612 open-thoughts/OpenThoughts2-1M 541 898 PrimeIntellect/SYNTHETIC-1-SFT-Data 474 810 allenai/Dolci-Think-SFT-7B 334 908 allenai/Dolci-Think-SFT-32B 327 491 GeneralReasoning/GeneralThought-430K 291 946… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/epic-thinking.text10M<n<100M5 likes405 downloads5mo agoHugging Face06ProCreations /grug-think grug-think grug make dataset. dataset make model think like grug. grug think short. short think cheap. cheap think good. big-brain model think 400 token before poke one tool. grug model think 11 word. same tool poke. same work done. many token saved. token = money. grug like money stay in pocket. what in box 100,891 example. every example = full agent conversation: system, user, assistant, tool message. assistant turn always got <think>grug reasoning</think> first… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-think.texttext-generation100K<n<1M35 likes370 downloads3mo agoHugging Face07drwlf /medra-thinking-768text1M<n<10M2 likes349 downloads1y agoHugging Face08ssurface /grade_school_math_thinkingtext1K<n<10K0 likes328 downloads4mo agoHugging Face09Lyric1010 /ablation_nemotron_no_thinking Dataset: ablation_nemotron_no_thinking This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train/ablation_nemotron_no_thinking/stage_1/tmp. textn<1K0 likes324 downloads8mo agoHugging Face10LoneResearch /thinking-steering-vectors1K<n<10K0 likes302 downloads1y agoHugging Face1111-47 /claude_opus_4.8_max_thinking_5k_v2 Claude Opus 4.8 MAX THINKING — Distillation Dataset 5,000 high-quality examples designed to distill the maximum-effort reasoning, honest analysis, production software engineering, and agentic capabilities of Claude Opus 4.8. Overview This dataset captures Opus 4.8’s signature strengths: Deep, structured, high-effort reasoning Honest communication about trade-offs and uncertainties Excellent production software engineering judgment Strong agentic workflow design… See the full description on the dataset page: https://huggingface.co/datasets/11-47/claude_opus_4.8_max_thinking_5k_v2.text1K<n<10K7 likes261 downloads4mo agoHugging Face12raei /Nieto-2022-ThinkingOutLoudOpenAccessEEGBasedBCIDatasetInnerSpeech Thinking out loud: an open-access EEG-based BCI dataset for inner speech recognition This is an unofficial mirror of OpenNeuro dataset ds003626, version 2.1.2. It is not affiliated with or endorsed by the dataset authors, their institutions, or OpenNeuro. Source and documentation Original dataset: OpenNeuro ds003626 v2.1.2 Article: Nieto et al., Scientific Data (2022) Original dataset documentation: README Official analysis code: N-Nieto/Inner_Speech_Dataset The… See the full description on the dataset page: https://huggingface.co/datasets/raei/Nieto-2022-ThinkingOutLoudOpenAccessEEGBasedBCIDatasetInnerSpeech.textn<1K0 likes252 downloads20d agoHugging Face13Davd-b01 /thinking-cap-tier-curricula-complete Thinking Cap Tier Curricula — Complete Reasoning Alignment Suite (TCS v4) [!IMPORTANT] Dataset Release v1.2 (Sept 2026) — Clean Delimiters & Zero-Padding Architecture: In v1.2, all 13,477 SFT samples and 3,187 SimPO preference pairs have undergone an automated token purge: Zero <|pad|> batch residues: 100% eliminated across all files. Zero reasoning leakage into final answers: Deliberation stays strictly inside <think>...</think>, and answers provide direct, non-repetitive… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinking-cap-tier-curricula-complete.texttext-generation10K<n<100K0 likes234 downloads9d agoHugging Face14Davd-b01 /thinking-cap-tier-lima-dense Thinking Cap Tier Curricula — LIMA Hyper-Dense Reasoning Alignment Suite (TCS v4) [!IMPORTANT] Dataset Release v1.2 (Sept 2026) — Clean Delimiters & Zero-Padding Architecture: In v1.2, all 5,500 SFT and 2,000 SimPO records have undergone a complete token purge: Zero <|pad|> batch residues: 100% eliminated across all records. Zero reasoning leakage into final answers: Deliberation stays strictly inside <think>...</think>, and answers provide direct conclusions. Native ChatML… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinking-cap-tier-lima-dense.texttext-generation1K<n<10K2 likes220 downloads9d agoHugging Face15MaxwellmuF /Soofi-Think-SFT-V2-secondhalf-DEtabular1M<n<10M0 likes209 downloads4mo agoHugging Face16Davd-b01 /thinking-cap-tier-raw-traces Thinking Cap Tier Raw Traces (TCS v4) [!IMPORTANT] Dataset Release v1.2 (Sept 2026) — Clean Delimiters & Zero-Padding Architecture: All 38,158 candidate reasoning traces across all 4 tiers (candidates_low.jsonl, candidates_mid.jsonl, candidates_high.jsonl, candidates_xhigh.jsonl) are 100% sanitized: Zero batch-padding residues (<|pad|>): Completely purged across all records. Strict Delimiter Integrity: Generation blocks cleanly separate thought deliberation tags… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinking-cap-tier-raw-traces.tabulartext-generation10K<n<100K0 likes200 downloads9d agoHugging Face17Davd-b01 /thinkingcap-condensed-qwen3.8-glm5.2-kimi-k3 ThinkingCap Condensed — Qwen3.8 / GLM-5.2 / Kimi-K3 Condensed ThinkingCap-style reasoning traces for SFT. 1,985 traces: each row pairs a full multi-turn teacher trace (Qwen3.8-Max, GLM-5.2 or Kimi K3, via r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillation) with a condensed TC-style version (short <think> + definitive numbered answer) generated by bottlecapai/ThinkingCap-Qwen3.6-27B using the thinkingcap system prompt. Format: JSONL (data/condensed.jsonl), 1,985 rows, UTF-8.… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinkingcap-condensed-qwen3.8-glm5.2-kimi-k3.texttext-generation1K<n<10K1 likes196 downloads1mo agoHugging Face18facebook /optimal_thinking_benchThis dataset is released as part of OptimalOptimalThinkingBench research project. IMPORTANT: This is only a subset of OptimalThinkingBench that does not contain the math problems. To download the full dataset, please refer to our project materials here for more details. Loading the dataset with transformers This dataset is built using Llama-4-Maverick and Reasoning-Gym. Details on how to generate this dataset can be found in OptimalOptimalThinkingBench paper. Minimal example below… See the full description on the dataset page: https://huggingface.co/datasets/facebook/optimal_thinking_bench.text1K<n<10K1 likes192 downloads11mo agoHugging Face19ThinkNet /HQ-Chat-2k 🧠 HQ-Chat-2K — High-Quality Conversational & Instruction-Tuning Dataset 2,000 carefully curated, high-quality conversation and instruction examples for fine-tuning Small Language Models (SLMs) and compact LLMs from ~500M to 3B parameters. HQ-Chat-2K is a high-quality conversational and instruction-tuning dataset designed specifically for training and fine-tuning small to medium-sized Large Language Models (LLMs). The dataset contains 2,000 curated user–assistant examples… See the full description on the dataset page: https://huggingface.co/datasets/ThinkNet/HQ-Chat-2k.texttext-generation1K<n<10K5 likes165 downloads1mo agoHugging Face20twinkle-ai /gpt-oss-120b-mandarin-thinking-eval-logs-and-scorestabular100K<n<1M0 likes163 downloads7mo agoHugging Face21twinkle-ai /gpt-oss-20b-mandarin-thinking-eval-logs-and-scorestabular100K<n<1M0 likes158 downloads7mo agoHugging Face22AlpachinoNLP /CT-RATE-Thinking CT-RATE-Thinking: Reasoning-Augmented CT Report Dataset 🎉🎉🎉 Our paper was accepted at the 28th conference of The Medical Image Computing and Computer Assisted Intervention Society (MICCAI). See you in Daejeon, Korea, September 23–27, 2025.CT-RATE-Thinking is a reasoning-augmented dataset derived from CT-RATE, containing chain-of-thought VQA pairs and report-level thinking narratives for 3D chest CT volumes. It was generated as part of the μ²Tokenizer project… See the full description on the dataset page: https://huggingface.co/datasets/AlpachinoNLP/CT-RATE-Thinking.textvisual-question-answering1M<n<10M2 likes157 downloads5mo agoHugging Face23Davd-b01 /thinkingcap-reasoning-traces ThinkingCap Reasoning Traces (Legacy v1 Prototype) [!WARNING] Legacy / Deprecated Prototype Notice (v1): This dataset represents an early exploratory prototype (v1, 4,254 traces) from initial development. Some samples in this legacy version contain early formatting artifacts, including reasoning traces leaking into the final answer field and informal step-by-step breakdowns. For modern post-training, SFT, and SimPO alignment under the TCS v4 cognitive standard, please use our… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinkingcap-reasoning-traces.texttext-generation1K<n<10K1 likes155 downloads9d agoHugging Face24Quardo /gsm8k-thinkingGSM8K with a reasoning/thinking/reflecting done by: new: llama3.1-405b old: chatgpt-4o-lastest the new one is not complete. [at 50~% as of now] text1K<n<10K0 likes143 downloads2y agoHugging Face25Jackrong /Chinese-Qwen3-235B-Thinking-2507-Distill-100k 📌 Note: The English translation of this dataset card is provided below. Chinese-Qwen3-235B-Thinking-2507-Distill-100k Dataset Summary Chinese-Qwen3-235B-Thinking-2507-Distill-100k 是一个包含约 100k 条高质量中文推理与指令数据的数据集,由 Qwen-3-235B-A22B-Thinking-2507(官方 Thinking 模式,上下文长度 32K)蒸馏生成。 该数据集覆盖了多个重要领域: 数学与工程任务(Mathematics, Applied Math, Advanced Math) 通用知识与写作(General Knowledge, Language & Writing) 技术与编程(Technology & Programming) 商业与经济(Business & Economics)… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Chinese-Qwen3-235B-Thinking-2507-Distill-100k.tabulartext-classification100K<n<1M19 likes142 downloads1y agoHugging Face26ltg /normistral-11b-thinking-evaluationtext10K<n<100K1 likes138 downloads10mo agoHugging Face27dougalldeepmind /2026-08-02-qwen36-mixture-500k-numina-heavy-empty-think Qwen3.6-27B SFT mixture — 500k maths-weighted, empty-think markers 499,595 tokens across 1,001 conversations, weighted toward maths, with Qwen3.6's empty think marker on the non-maths rows. md5 c433f31eba2b5b4919fb166043caccb5. Source Examples Tokens Share Marker NuminaMath-CoT 611 333,351 66.9% no No Robots 271 82,239 16.5% yes TULU3 119 82,445 16.5% yes Total 1,001 499,595 390 marked Derived from qwen3.6-27b-mixture-500k-numina-heavy by adding the… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-02-qwen36-mixture-500k-numina-heavy-empty-think.texttext-generation1K<n<10K0 likes138 downloads22d agoHugging Face28shreethar /thinkflow-vla-features-b2tabular1K<n<10K0 likes134 downloads2mo agoHugging Face29agentlans /allenai-Dolci-Thinktext1M<n<10M0 likes131 downloads5mo agoHugging Face30agentlans /TeichAI-thinking-reasoning-x TeichAI Thinking & Reasoning Datasets A collection of prompts answered by large language models (LLMs) such as Google Gemini and OpenAI ChatGPT, with long-form reasoning enabled. These datasets were originally created by TeichAI for distillation and reasoning-focused training workflows. Schema Each row in the dataset has the following fields: question_hash: Truncated, base64-encoded MD5 hash of the question, useful for filtering and deduplication. question: The… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/TeichAI-thinking-reasoning-x.texttext-generation100K<n<1M1 likes127 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.