CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01samuki-hf /thinking-rollouts thinking-rollouts Unconstrained rollouts from thinking (chain-of-thought) models on DS-1000 and LiveCodeBench, CoT saved verbatim alongside the final answer. Format per genlm/rollouts issue #5; schema is a superset of temperature-sweep-data. Hive-partitioned Parquet, thinking_mode folded into the model tag: rollouts/domain=<dataset>/model=<tag>/temp=<temp>/data.parquet (tags like qwen3-8b-think, qwen3-1.7b-nothink). 100 samples/instance. Columns: model, thinking_mode, temp… See the full description on the dataset page: https://huggingface.co/datasets/samuki-hf/thinking-rollouts.tabulartext-generation10M<n<100M2 likes641 downloads2mo agoHugging Face02tyrtleli /thinking-benchmark-90 Thinking Benchmark A calibration pool of 90 competition-mathematics problems assembled to study how output / reasoning-trace length varies with problem difficulty across frontier language models. Part of the Cost of Overthinking research project. Dataset at a glance Source n Difficulty Contamination risk AIME 2026 29 3–5 low OlymMATH 41 4–6 medium HMMT February 2026 12 4–5 low MATH-500 5 2–3 high FrontierMath-style 3 6 medium Difficulty is… See the full description on the dataset page: https://huggingface.co/datasets/tyrtleli/thinking-benchmark-90.tabularquestion-answeringn<1K0 likes491 downloads1mo agoHugging Face03marin-community /openthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16 OpenThoughts-4 Code SDG: Qwen3-30B-A3B-Thinking-2507 (n=16, top-16 logprobs) Synthetic generations from Qwen/Qwen3-30B-A3B-Thinking-2507 on the Marin OpenThoughts-4 code SDG prompt set. Each prompt is sampled n=16 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes488 downloads5mo agoHugging Face04TyroneDragon /highlevel_thinking_with_grounding_annotation_split1000_v3_merged_promptstabular10K<n<100K0 likes426 downloads1y agoHugging Face05OALL /details_Qwen__Qwen3-30B-A3B-Thinking-2507_v2 Dataset Card for Evaluation run of Qwen/Qwen3-30B-A3B-Thinking-2507 Dataset automatically created during the evaluation run of model Qwen/Qwen3-30B-A3B-Thinking-2507. The dataset is composed of 116 configuration, each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Qwen__Qwen3-30B-A3B-Thinking-2507_v2.tabular100K<n<1M0 likes308 downloads8mo agoHugging Face06marin-community /openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16 OpenThoughts-4 Science SDG: Qwen3-30B-A3B-Thinking-2507 (n=8, top-16 logprobs) Synthetic generations from Qwen/Qwen3-30B-A3B-Thinking-2507 on the Marin OpenThoughts-4 science SDG prompt set. Each prompt is sampled n=8 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes306 downloads5mo agoHugging Face07PursuitOfDataScience /0.5M-thinking 0.5M Thinking Dataset This dataset contains responses generated by MiniMax-M2.1 for user questions from the a-m-team/AM-DeepSeek-R1-Distilled-1.4M dataset (am_0.5M subset). Dataset Description The dataset captures both the extended thinking process and final answers from MiniMax-M2.1, with reasoning wrapped in <think> tags for easy separation. Metric Value Examples 499,157 Total Tokens 3,732,749,397 Avg Tokens/Example 7,478 Source Dataset… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/0.5M-thinking.tabulartext-generation100K<n<1M0 likes246 downloads9mo agoHugging Face08marin-community /open-thoughts-4-30k-code-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8 Open Thoughts 4 - Code (Qwen3-30B-A3B-Thinking-2507, 32K tokens, n=8) This dataset contains code reasoning problems with 8 independent responses generated by Qwen3-30B-A3B-Thinking-2507. Overview Source: marin-community/open-thoughts-4-30k-code-qwen3-32b-annotated (prompts only) Model: Qwen/Qwen3-30B-A3B-Thinking-2507 Temperature: 0.8 Max tokens: 32,768 Columns Column Description instruction_seed The code problem prompt _source Source dataset… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-code-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8.tabular10K<n<100K0 likes222 downloads7mo agoHugging Face09marin-community /open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8 Open Thoughts 4 - Math (Qwen3-30B-A3B-Thinking-2507, 32K tokens, n=8) This dataset contains math reasoning problems with 8 independent responses generated by Qwen3-30B-A3B-Thinking-2507. Overview Source: marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated (base prompts) Model: Qwen/Qwen3-30B-A3B-Thinking-2507 Temperature: 0.8 Max tokens: 32,768 Columns Column Description instruction_seed The math problem prompt _source Source dataset… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8.tabular10K<n<100K0 likes203 downloads7mo agoHugging Face10marin-community /open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8-reformatted Dataset Card for Open-Thoughts-4-30K-Math-Qwen3-30B-A3B-Thinking-2507-Annotated-32768-Tokens-N8-Reformatted Overview This dataset is a reformatted version of marin-community/open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8. The original dataset contained 29,963 samples, each with 8 responses generated by the same model with different random seeds (stored in generated_text, generated_text2, ..., generated_text8 columns). This reformatted… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8-reformatted.tabular100K<n<1M0 likes199 downloads6mo agoHugging Face11TyroneDragon /highlevel_thinking_with_grounding_annotation_split1000_v2_merged_promptstabular10K<n<100K0 likes175 downloads1y agoHugging Face12PursuitOfDataScience /arxiv-qa-thinking ArXiv Q&A with Thinking Dataset This dataset contains question-answer pairs generated by MiniMax-M2.1 based on academic articles from PursuitOfDataScience/arxiv-llama4-maverick-abstract. Dataset Description For each academic article, the model generates: Thinking process: The model's reasoning wrapped in <think> tags Question: An insightful question testing understanding of key concepts Answer: A detailed answer based on the article content Statistics… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/arxiv-qa-thinking.tabulartext-generation100K<n<1M0 likes159 downloads8mo agoHugging Face13LeRobot-worldwide-hackathon /217-Thinking_Beyond-pickingSaltSticksThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 5, "total_frames": 1360, "total_tasks": 1, "total_videos": 10, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:5" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LeRobot-worldwide-hackathon/217-Thinking_Beyond-pickingSaltSticks.tabularrobotics10K<n<100K1 likes138 downloads1y agoHugging Face14PursuitOfDataScience /0.9M-thinking 0.9M Thinking Dataset This dataset contains responses generated by MiniMax-M2.1 for user questions from the a-m-team/AM-DeepSeek-R1-Distilled-1.4M dataset (am_0.9M subset). Dataset Description The dataset captures both the extended thinking process and final answers from MiniMax-M2.1, with reasoning wrapped in <think> tags for easy separation. Metric Value Examples 897,522 Total Tokens 5,954,272,687 Avg Tokens/Example 6,634 Source Dataset… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/0.9M-thinking.tabulartext-generation100K<n<1M0 likes132 downloads8mo agoHugging Face15rohinm /zip-training-hallucination-data-qwen06b-thinking-train-with-valuestabular10K<n<100K0 likes118 downloads1y agoHugging Face16rubricreward /PolyGuardMix-filtered-tgt_prompt_tgt_thinkingtabular100K<n<1M0 likes92 downloads1y agoHugging Face17PursuitOfDataScience /gsm8k-thinking GSM8K Thinking This dataset contains responses generated by MiniMax-M2.1 for math word problems from the openai/gsm8k dataset. Dataset Description The dataset captures both the extended thinking process and final answers from MiniMax-M2.1, with reasoning wrapped in <think> tags for easy separation. Metric Value Train Examples 7,473 Test Examples 1,319 Total Examples 8,792 Total Tokens 10,506,774 Avg Tokens/Example 1,195 Source Dataset… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/gsm8k-thinking.tabulartext-generation1K<n<10K0 likes79 downloads9mo agoHugging Face18elichen-skymizer /qwen3-4b-thinking-aime-untruncatedtabular1K<n<10K0 likes66 downloads10mo agoHugging Face19JonasLoos /Qwen3.8-27B-thinking-completions Qwen3.8-27B thinking-mode completions 17,022 prompts from public chat, math and code datasets, each answered once by Qwen3.8-27B (FP8 checkpoint) in thinking mode with its recommended sampling settings (54M completion tokens). Every sample has the reasoning trace and the final answer, as text and as the exact token ids. The set was generated to train speculative-decoding drafters for this model, so it records the model's own sampled distribution rather than greedy output or… See the full description on the dataset page: https://huggingface.co/datasets/JonasLoos/Qwen3.8-27B-thinking-completions.tabulartext-generation10K<n<100K0 likes61 downloads9d agoHugging Face20RLAIF /dpo_thinking_with_gold_labels_kl_estimationtabular10K<n<100K0 likes56 downloads1y agoHugging Face21pawin205 /iclr-2017-2020-peer-review-with-thinking-tracetabular10K<n<100K0 likes54 downloads5mo agoHugging Face22lvogel123 /cybench-qwen3-235b-a22b-thinking-2507tabularn<1K0 likes54 downloads11mo agoHugging Face23marin-community /open-thoughts-4-30k-code-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8-reformatted Dataset Card for Open-Thoughts-4-30K-Code-Qwen3-30B-A3B-Thinking-2507-Annotated-32768-Tokens-N8-Reformatted Overview This dataset is a reformatted version of marin-community/open-thoughts-4-30k-code-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8. The original dataset contained 29,963 samples, each with 8 responses generated by the same model with different random seeds (stored in generated_text, generated_text2, ..., generated_text8 columns). This reformatted… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-code-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8-reformatted.tabular100K<n<1M0 likes51 downloads7mo agoHugging Face24zsqzz /sft-mathhard-medium-with-thinking-full-paralleltabular1K<n<10K0 likes45 downloads1y agoHugging Face25PursuitOfDataScience /toucan-agentic-thinking Toucan Agentic with Thinking Dataset This dataset contains agentic reasoning responses generated by MiniMax-M2.1 based on questions from Agent-Ark/Toucan-1.5M_SFT. Dataset Description For each user question, the model generates: Thinking process: The model's reasoning wrapped in <think> tags Response: A complete, helpful answer in natural language The original tool definitions are preserved in the tools field for reference. Statistics Split Examples… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/toucan-agentic-thinking.tabulartext-generation100K<n<1M0 likes40 downloads8mo agoHugging Face26marin-community /open-thoughts-4-8530-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8-reformatted open-thoughts-4-8530-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens This dataset contains 68240 rows (8530 math prompts x 8 responses each) generated by Qwen3-30B-A3B-Thinking-2507 with a max token length of 32768. Derived from the first 68240 rows of marin-community/open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8-reformatted. Columns Column Description row_id Original row identifier instruction_seed The math prompt… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-8530-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8-reformatted.tabular10K<n<100K0 likes40 downloads5mo agoHugging Face27siro1 /kernelbook-kimi_k2_thinking-evals-synthetic-promptstabular10K<n<100K0 likes39 downloads8mo agoHugging Face28siro1 /kernelbook-kimi_k2_thinking-evalstabular10K<n<100K0 likes38 downloads8mo agoHugging Face29rubricreward /PolyGuardMix-filtered-en_prompt_en_thinking-filtered_correcttabular100K<n<1M0 likes36 downloads1y agoHugging Face30erdem-erdem /maze_test_with_thinkingtabular10K<n<100K0 likes34 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.