CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ThinkingHub /PP SteelBench: A Diagnostic Benchmark for Vision-Language Models in Industrial Safety Monitoring SteelBench is a diagnostic benchmark of densely annotated CCTV clips from an operating integrated steel plant. It is designed to evaluate vision-language models (VLMs) on real-world industrial action recognition, PPE assessment, and safety-violation detection — under naturally occurring degradation (dust, glare, steam, low light), at distances and crowdedness levels that curated… See the full description on the dataset page: https://huggingface.co/datasets/ThinkingHub/PP.imagevideo-classification1K<n<10K0 likes3.1k downloads3mo agoHugging Face02kashif /opd-kd-thinky-deepmath-completions train_rl Completion Logs This dataset contains the on-policy generations produced during RL training with train_rl. Training details Key Value Algorithm OPD Model (student) HuggingFaceH4/KD-Thinky Model (teacher) Qwen/Qwen3-8B Prompt dataset HuggingFaceH4/DeepMath-103K Group size 4 Max completion tokens 4096 Temperature 1.0 Learning rate 0.0001 model_revision v00.08-step-000003125 dataset_configtrl_all lora_rank 128 opd_kl_coef 1.0… See the full description on the dataset page: https://huggingface.co/datasets/kashif/opd-kd-thinky-deepmath-completions.tabular10K<n<100K0 likes3k downloads7mo agoHugging Face03ru-dataset /agent-think-tool_use Agent Think Tool Use Датасет многошаговых агентных сессий для дообучения моделей работе с кодом, инструментами и инженерными задачами. Записи содержат пользовательские требования, комментарии агента во время работы, decision summaries, вызовы инструментов, результаты запусков, обработку ошибок и финальную проверку. Каждый shard представляет отдельную связанную сессию, а не отдельный вопрос и ответ. Данные охватывают исследование задачи, работу с документацией, проектирование… See the full description on the dataset page: https://huggingface.co/datasets/ru-dataset/agent-think-tool_use.tabulartext-generationn<1K2 likes744 downloads3d agoHugging Face04openeurollm /Dolci-Think-SFT-translated Dolci-Think-SFT-translated Machine translations of the Dolci-Think-SFT-32B dataset, produced with gemma-4-31B-it. The samples selected for translation are those where content_quality == "excellent" according to the propella annotations. Columns Each row is a translated conversation plus the result of a post-translation quality filter: id — source record id. messages — the translated conversation (list of {content, role}). filter_pass — true if the row passed… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/Dolci-Think-SFT-translated.tabulartext-generation1M<n<10M0 likes687 downloads8d agoHugging Face05samuki-hf /thinking-rollouts thinking-rollouts Unconstrained rollouts from thinking (chain-of-thought) models on DS-1000 and LiveCodeBench, CoT saved verbatim alongside the final answer. Format per genlm/rollouts issue #5; schema is a superset of temperature-sweep-data. Hive-partitioned Parquet, thinking_mode folded into the model tag: rollouts/domain=<dataset>/model=<tag>/temp=<temp>/data.parquet (tags like qwen3-8b-think, qwen3-1.7b-nothink). 100 samples/instance. Columns: model, thinking_mode, temp… See the full description on the dataset page: https://huggingface.co/datasets/samuki-hf/thinking-rollouts.tabulartext-generation10M<n<100M2 likes608 downloads2mo agoHugging Face06allenai /Dolci-Think-RL-7B Dolci-Think-RL-7B Dataset Summary Dolci-Think-RL-7B is the reinforcement learning dataset used to train the Olmo-3-7B-Think model.It contains 102,014 prompts designed to elicit deep reasoning across: Math Coding Precise Instruction Following General Chat It blends high-quality curated sources with filtering designed for deliberate reasoning. Dataset Composition Total Samples: 102,014 Original Dataset Contribution… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Dolci-Think-RL-7B.tabular100K<n<1M17 likes606 downloads9mo agoHugging Face07marin-community /openthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16 OpenThoughts-4 Code SDG: Qwen3-30B-A3B-Thinking-2507 (n=16, top-16 logprobs) Synthetic generations from Qwen/Qwen3-30B-A3B-Thinking-2507 on the Marin OpenThoughts-4 code SDG prompt set. Each prompt is sampled n=16 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes595 downloads5mo agoHugging Face08rl-rag /browsecomp-qwen35-35b-a3b-think browsecomp-qwen35-35b-a3b-think Deep research agent evaluation on data/browsecomp.jsonl (normal split). Results Metric Value pass@4 43.0% avg@4 24.8% Trajectory accuracy 24.8% (1258/5064) Questions 1266 Trajectories 5064 (4 per question) Avg tool calls 41.1 Full conversations ❌ Model & Setup Model Qwen3.5-35B-A3B Judge gpt-4o Max tool calls 50 Temperature 0.7 Blocked domains huggingface.co Tool… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-qwen35-35b-a3b-think.tabular1K<n<10K0 likes573 downloads7mo agoHugging Face09toroe /Soofi-Think-SFT-10B-multilingual ReasonXL: A Multilingual Cross-Domain Reasoning Corpus ReasonXL is a large-scale multilingual reasoning corpus spanning 5 languages and ~44B tokens in total. It is designed to support supervised fine-tuning of reasoning models with in-language chain-of-thought traces across diverse technical domains. Data Generation English source samples were drawn from 10 existing reasoning datasets, filtered and quality-annotated using ellamind/propella-1-4b, and then translated into… See the full description on the dataset page: https://huggingface.co/datasets/toroe/Soofi-Think-SFT-10B-multilingual.tabular10M<n<100M0 likes537 downloads6mo agoHugging Face10tyrtleli /thinking-benchmark-90 Thinking Benchmark A calibration pool of 90 competition-mathematics problems assembled to study how output / reasoning-trace length varies with problem difficulty across frontier language models. Part of the Cost of Overthinking research project. Dataset at a glance Source n Difficulty Contamination risk AIME 2026 29 3–5 low OlymMATH 41 4–6 medium HMMT February 2026 12 4–5 low MATH-500 5 2–3 high FrontierMath-style 3 6 medium Difficulty is… See the full description on the dataset page: https://huggingface.co/datasets/tyrtleli/thinking-benchmark-90.tabularquestion-answeringn<1K0 likes428 downloads1mo agoHugging Face11TyroneDragon /highlevel_thinking_with_grounding_annotation_split1000_v3_merged_promptstabular10K<n<100K0 likes374 downloads11mo agoHugging Face12marin-community /openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16 OpenThoughts-4 Science SDG: Qwen3-30B-A3B-Thinking-2507 (n=8, top-16 logprobs) Synthetic generations from Qwen/Qwen3-30B-A3B-Thinking-2507 on the Marin OpenThoughts-4 science SDG prompt set. Each prompt is sampled n=8 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes351 downloads5mo agoHugging Face13JackHsieh /4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained 4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained A pre-tokenized, tag-wrapped variant of JackHsieh/4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids. The source thoughts are two parts: a <think> block, then a few dense sentences that the generator wrapped in a literal <thought>…</thought> block (plain text, not special tokens). Only the part after </think> becomes the VALUE, with every <thought> / </thought> tag removed… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.tabular10M<n<100M0 likes331 downloads16d agoHugging Face14JackHsieh /4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids 4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen/Qwen3-4B in thinking mode. The prompted task: reason about what comes IMMEDIATELY next — the next k=8 tokens after the cut — and answer with a few dense, declarative sentences (about two to four) inside a <thought>…</thought> block, focused on the exact state at the cut and what the local grammar… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.tabular10M<n<100M0 likes329 downloads16d agoHugging Face15OALL /details_Qwen__Qwen3-30B-A3B-Thinking-2507_v2 Dataset Card for Evaluation run of Qwen/Qwen3-30B-A3B-Thinking-2507 Dataset automatically created during the evaluation run of model Qwen/Qwen3-30B-A3B-Thinking-2507. The dataset is composed of 116 configuration, each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Qwen__Qwen3-30B-A3B-Thinking-2507_v2.tabular100K<n<1M0 likes305 downloads8mo agoHugging Face16JackHsieh /4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained 4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained A pre-tokenized, tag-wrapped variant of JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids. The source thoughts are two parts: a <think> block, then a single paragraph of dense reasoning about the immediate continuation. Only the part after </think> — the paragraph — becomes the VALUE. The reasoning inside the think block is dropped. The VALUE is capped at 512 tokens.… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.tabular10M<n<100M0 likes302 downloads16d agoHugging Face17JackHsieh /4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids 4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen/Qwen3-4B in thinking mode. The prompted task: reason about what comes IMMEDIATELY next — the next k=8 tokens after the cut — and answer with a single unconstrained paragraph of dense reasoning, focused on the exact state at the cut and what the local grammar, notation, or argument forces next. Both the… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.tabular10M<n<100M0 likes286 downloads16d agoHugging Face18RfKnowledge /dolci-think-translation-tests Dolci-Think Translation Tests This dataset contains the inputs, outputs, reconstruction records, automatic diagnostics, and retained COMET-QE evidence from a set of controlled translation experiments over 340 English messages selected from allenai/Dolci-Think-SFT-7B. It is intended for inspecting the experiments and reproducing analyses, not as a ready-made translation training set. The central complication is that one source document does not always correspond to one model… See the full description on the dataset page: https://huggingface.co/datasets/RfKnowledge/dolci-think-translation-tests.tabulartranslation1M<n<10M0 likes280 downloads28d agoHugging Face19anuj-inavlabs /kupe-thinkspark-270m-phase1-data ThinkSpark-v2-350M — Phase-1 free-audio training data Pre-encoded Mimi cb0 (12.5Hz semantic) tokens + per-frame energy/f0 prosody, packaged for Phase 1 of ThinkSpark-v2-350M — teaching a 270M-parameter Gemma-3-based full-duplex floor-controller that a stream of Mimi audio tokens carries language + prosody, before Phase 2 teaches it to referee turn-taking. Sourced from free/open corpora (LibriSpeech, AI4Bharat Kathbath/Shrutilipi, IndicTTS, Google FLEURS) via… See the full description on the dataset page: https://huggingface.co/datasets/anuj-inavlabs/kupe-thinkspark-270m-phase1-data.tabular100K<n<1M0 likes276 downloads24d agoHugging Face20violetxi /harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), newly graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and all-criteria-pass rule. Mean all-pass rate: 5.0000%. The train split contains evaluation records, not training examples. Generation and grading protocols Generation is unchanged: historical 20-turn thinking-enabled glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think.tabulartext-generation1K<n<10K0 likes275 downloads17h agoHugging Face21violetxi /harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-5t-think harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-5t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), newly graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and all-criteria-pass rule. Mean all-pass rate: 2.7000%. The train split contains evaluation records, not training examples. Generation and grading protocols Generation is unchanged: historical 5-turn thinking-enabled glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-5t-think.tabulartext-generation1K<n<10K0 likes272 downloads17h agoHugging Face22mnoukhov /dolci_think_rl_7b_messages_hybrid_450m_scoredtabular100K<n<1M0 likes271 downloads1mo agoHugging Face23violetxi /harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), newly graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and all-criteria-pass rule. Mean all-pass rate: 1.3000%. The train split contains evaluation records, not training examples. Generation and grading protocols Generation is unchanged: historical 20-turn thinking-enabled glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think.tabulartext-generation1K<n<10K0 likes263 downloads17h agoHugging Face24violetxi /harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), newly graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and all-criteria-pass rule. Mean all-pass rate: 4.0000%. The train split contains evaluation records, not training examples. Generation and grading protocols Generation is unchanged: historical 20-turn thinking-enabled glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think.tabulartext-generation1K<n<10K0 likes256 downloads17h agoHugging Face25violetxi /harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-5t-think harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-5t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), newly graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and all-criteria-pass rule. Mean all-pass rate: 3.7000%. The train split contains evaluation records, not training examples. Generation and grading protocols Generation is unchanged: historical 5-turn thinking-enabled glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-5t-think.tabulartext-generation1K<n<10K0 likes254 downloads17h agoHugging Face26violetxi /harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-5t-think harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-5t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), newly graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and all-criteria-pass rule. Mean all-pass rate: 1.9000%. The train split contains evaluation records, not training examples. Generation and grading protocols Generation is unchanged: historical 5-turn thinking-enabled glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-5t-think.tabulartext-generation1K<n<10K0 likes245 downloads17h agoHugging Face27PursuitOfDataScience /0.5M-thinking 0.5M Thinking Dataset This dataset contains responses generated by MiniMax-M2.1 for user questions from the a-m-team/AM-DeepSeek-R1-Distilled-1.4M dataset (am_0.5M subset). Dataset Description The dataset captures both the extended thinking process and final answers from MiniMax-M2.1, with reasoning wrapped in <think> tags for easy separation. Metric Value Examples 499,157 Total Tokens 3,732,749,397 Avg Tokens/Example 7,478 Source Dataset… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/0.5M-thinking.tabulartext-generation100K<n<1M0 likes237 downloads9mo agoHugging Face28allenai /Dolci-Think-RL-7B-Completions-SFT Dolci-Think-Completions-SFT Dataset Summary Dolci-Think-Completions-SFT is a set of 5,031,398 completions(!!) from the Olmo-3-7B-Think-SFT model over the prompts considered when making Dolci-Think-RL. These completions were mainly used to filter easy data, but we believe the completions may be useful in general. It contains 636,095 high-quality prompts covering: Math Code Precise Instruction Following General Chat Puzzles Each split covers one of the above domains, and… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Dolci-Think-RL-7B-Completions-SFT.tabular100K<n<1M9 likes225 downloads9mo agoHugging Face29marin-community /open-thoughts-4-30k-code-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8 Open Thoughts 4 - Code (Qwen3-30B-A3B-Thinking-2507, 32K tokens, n=8) This dataset contains code reasoning problems with 8 independent responses generated by Qwen3-30B-A3B-Thinking-2507. Overview Source: marin-community/open-thoughts-4-30k-code-qwen3-32b-annotated (prompts only) Model: Qwen/Qwen3-30B-A3B-Thinking-2507 Temperature: 0.8 Max tokens: 32,768 Columns Column Description instruction_seed The code problem prompt _source Source dataset… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-code-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8.tabular10K<n<100K0 likes217 downloads7mo agoHugging Face30MaxwellmuF /Soofi-Think-SFT-V2-secondhalf-DEtabular1M<n<10M0 likes209 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.