CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFaceH4 /Multilingual-Thinking Dataset summary Multilingual-Thinking is a reasoning dataset where the chain-of-thought has been translated from English into one of 4 languages: Spanish, French, Italian, and German. The dataset was created by sampling 1k training samples from the SystemChat subset of SmolTalk2 and translating the reasoning traces with another language model. This dataset was used in the OpenAI Cookbook to fine-tune the OpenAI gpt-oss models. You can load the dataset using: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/Multilingual-Thinking.texttext-generation1K<n<10K118 likes9k downloads1y agoHugging Face02ShareLab-SII /thinking_droid_lerobot_output_qwen3vlimage1M<n<10M0 likes4.8k downloads5mo agoHugging Face03ShareLab-SII /thinking_fmb_dataset_lerobot_output_qwen3vlimage1M<n<10M0 likes3.1k downloads6mo agoHugging Face04llm-jp /llm-jp-4-thinking-sft-data llm-jp-4-thinking-sft-data Overview This dataset is a supervised fine-tuning (SFT) dataset used to train llm-jp-4-*-thinking models. This dataset is constructed by extracting prompts from multiple data sources and generating reasoning processes and final responses using gpt-oss-120b. The splits reasoning_low, reasoning_medium, and reasoning_high correspond to different reasoning effort settings used during generation with gpt-oss-120b. To support the continued development… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/llm-jp-4-thinking-sft-data.text1M<n<10M9 likes3.1k downloads5mo agoHugging Face05ioi-leaderboard /ioi-eval-openrouter_anthropic_claude-3_7-sonnet_thinking-prompt-mem-limittextn<1K0 likes2.2k downloads2y agoHugging Face06ioi-leaderboard /ioi-eval-openrouter_google_gemini-2_0-flash-thinking-exp-prompt-mem-limittextn<1K0 likes2.2k downloads2y agoHugging Face07ShareLab-SII /thinking_furniture_bench_dataset_lerobot_output_qwen3vlimage1M<n<10M0 likes2k downloads6mo agoHugging Face08Modotte /CodeX-2M-Thinking Modotte Note: This dataset is part of the lineup CodeX by Modotte. You can get lots of datasets in this same lineup, with the main focus on providing very high-quality datasets for model training and fine-tuning. This dataset is fully synthetic, curated from high-quality public sources and enhanced with synthetic data generated using both closed and open-source models. It serves as a strong foundation for instruction-based model tuning and fine-tuning, offering one of the… See the full description on the dataset page: https://huggingface.co/datasets/Modotte/CodeX-2M-Thinking.texttext-generation1M<n<10M128 likes1.9k downloads8mo agoHugging Face09OpenDataArena /MMFineReason-1.8M-Qwen3-VL-235B-Thinking MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal understanding benchmarks. 📖 Overview MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking. 🎯 Key Highlights 1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-1.8M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M126 likes1.8k downloads7mo agoHugging Face10NarsAI /FineReason-1.8M-Qwen3-VL-235B-Thinking MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal understanding benchmarks. 📖 Overview MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking. 🎯 Key Highlights 1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/NarsAI/FineReason-1.8M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M0 likes1.4k downloads8mo agoHugging Face11OpenDataArena /MMFineReason-SFT-586K-Qwen3-VL-235B-Thinking MMFineReason-SFT-586K The Hardest 33% — Less Data, More Reasoning 📖 Overview MMFineReason-SFT-586K is a difficulty-filtered subset of MMFineReason-1.8M, containing the hardest 33% of samples where Qwen3-VL-4B-Thinking do not consistently succeed. (pass rate ≠ 1). Specifically, this subset removes all easy samples (pass rate = 1) under Qwen3-VL-4B-Thinking, retaining only instances that require non-trivial multimodal reasoning. 🎯 Key Highlights 586K… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-SFT-586K-Qwen3-VL-235B-Thinking.image100K<n<1M6 likes1.3k downloads8mo agoHugging Face12Sandeepthakur /MMFineReason-1.8M-Qwen3-VL-235B-Thinking MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal understanding benchmarks. 📖 Overview MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking. 🎯 Key Highlights 1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/Sandeepthakur/MMFineReason-1.8M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M0 likes1k downloads8mo agoHugging Face13microsoft /ThinkingBox-Bench ThinkingBox-Bench ThinkingBox-Bench is an executable benchmark for evaluating whether tool-using LLM agents can reliably complete stateful business workflows. Version 1.0 contains 507 tool-agent-user tasks across retail and e-commerce, travel and hospitality, auto insurance, neobank support, and consulting IT/HR support. This dataset repository provides a browsable representation of the benchmark. The executable benchmark, tool servers, and supporting fixtures are maintained in… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/ThinkingBox-Bench.textreinforcement-learningn<1K14 likes745 downloads28d agoHugging Face14samuki-hf /thinking-rollouts thinking-rollouts Unconstrained rollouts from thinking (chain-of-thought) models on DS-1000 and LiveCodeBench, CoT saved verbatim alongside the final answer. Format per genlm/rollouts issue #5; schema is a superset of temperature-sweep-data. Hive-partitioned Parquet, thinking_mode folded into the model tag: rollouts/domain=<dataset>/model=<tag>/temp=<temp>/data.parquet (tags like qwen3-8b-think, qwen3-1.7b-nothink). 100 samples/instance. Columns: model, thinking_mode, temp… See the full description on the dataset page: https://huggingface.co/datasets/samuki-hf/thinking-rollouts.tabulartext-generation10M<n<100M2 likes641 downloads2mo agoHugging Face15Modotte /CodeX-7M-Non-Thinking Modotte Note: This dataset is part of the lineup CodeX by Modotte. You can get lots of datasets in this same lineup, with the main focus on providing very high-quality datasets for model training and fine-tuning. This dataset is curated from high-quality public sources and enhanced with synthetic data from both closed and open-source models. It serves as a strong foundation for instruction-based model tuning and fine-tuning, offering one of the most refined and extensive… See the full description on the dataset page: https://huggingface.co/datasets/Modotte/CodeX-7M-Non-Thinking.texttext-generation1M<n<10M25 likes575 downloads8mo agoHugging Face16OpenDataArena /MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking MMFineReason-SFT-123K The Hardest 7% — Less Data, More Reasoning 📖 Overview MMFineReason-SFT-123K is a difficulty-filtered subset of MMFineReason-1.8M, containing only the hardest 7% of samples where Qwen3-VL-4B-Thinking consistently fails (pass rate = 0). 🎯 Key Highlights 123K Challenging Samples: Only instances where a 4B thinking model fails all 4 inference attemptsEfficient Training: Comparable performance to full 1.8M dataset with only 7% of… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking.imagevisual-question-answering100K<n<1M86 likes525 downloads8mo agoHugging Face17Jofthomas /hermes-function-calling-thinking-V1text1K<n<10K79 likes507 downloads2y agoHugging Face18tyrtleli /thinking-benchmark-90 Thinking Benchmark A calibration pool of 90 competition-mathematics problems assembled to study how output / reasoning-trace length varies with problem difficulty across frontier language models. Part of the Cost of Overthinking research project. Dataset at a glance Source n Difficulty Contamination risk AIME 2026 29 3–5 low OlymMATH 41 4–6 medium HMMT February 2026 12 4–5 low MATH-500 5 2–3 high FrontierMath-style 3 6 medium Difficulty is… See the full description on the dataset page: https://huggingface.co/datasets/tyrtleli/thinking-benchmark-90.tabularquestion-answeringn<1K0 likes491 downloads1mo agoHugging Face19marin-community /openthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16 OpenThoughts-4 Code SDG: Qwen3-30B-A3B-Thinking-2507 (n=16, top-16 logprobs) Synthetic generations from Qwen/Qwen3-30B-A3B-Thinking-2507 on the Marin OpenThoughts-4 code SDG prompt set. Each prompt is sampled n=16 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes488 downloads5mo agoHugging Face20TyroneDragon /highlevel_thinking_with_grounding_annotation_split1000_v3text10K<n<100K0 likes466 downloads1y agoHugging Face21llm-jp /llm-jp-4-8b-thinking-dpo-data llm-jp-4-8b-thinking-dpo-data Overview This dataset is a Direct Preference Optimization (DPO) dataset used to train llm-jp-4-8b-thinking. It is constructed by pairing multiple candidate responses for a given prompt and selecting preferred (chosen) and non-preferred (rejected) responses. The splits reasoning_low, reasoning_medium, and reasoning_high correspond to different reasoning effort settings used during response generation. The fields chosen_analysis, chosen_final… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/llm-jp-4-8b-thinking-dpo-data.text100K<n<1M3 likes462 downloads5mo agoHugging Face22llm-jp /llm-jp-4-32b-a3b-thinking-dpo-data llm-jp-4-32b-a3b-thinking-dpo-data Overview This dataset is a Direct Preference Optimization (DPO) dataset used to train llm-jp-4-32b-a3b-thinking. It is constructed by pairing multiple candidate responses for a given prompt and selecting preferred (chosen) and non-preferred (rejected) responses. The splits reasoning_low, reasoning_medium, and reasoning_high correspond to different reasoning effort settings used during response generation. The fields chosen_analysis… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/llm-jp-4-32b-a3b-thinking-dpo-data.text100K<n<1M1 likes448 downloads5mo agoHugging Face23llm-jp /llm-jp-4-33b-thinking-dpo-data llm-jp-4-33b-thinking-dpo-data Overview This dataset is a Direct Preference Optimization (DPO) dataset used to train llm-jp-4-33b-thinking. It is constructed by pairing multiple candidate responses for a given prompt and selecting preferred (chosen) and non-preferred (rejected) responses. The splits reasoning_low, reasoning_medium, and reasoning_high correspond to different reasoning effort settings used during response generation. The fields chosen_analysis… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/llm-jp-4-33b-thinking-dpo-data.text10K<n<100K2 likes431 downloads1mo agoHugging Face24TyroneDragon /highlevel_thinking_with_grounding_annotation_split1000_v3_merged_promptstabular10K<n<100K0 likes426 downloads1y agoHugging Face25FleetAI /trajectories-thinking-toolstext100K<n<1M0 likes407 downloads7mo agoHugging Face26ShareLab-SII /thinking_taco_play_lerobot_output_qwen3vlimage100K<n<1M0 likes384 downloads6mo agoHugging Face27AmirhoseinGH /mhlc-training-qwen3vl-qwen3_vl_2b_thinking_hard_mixed_sources_120k Multi Head Latent Control Training Data - Qwen3-VL 2B Thinking hard Mixed Sources 120k Dataset Description This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection. Paper https://arxiv.org/abs/2607.14277 Code https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control Dataset Summary Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_2b_thinking_hard_mixed_sources_120k.imagequestion-answering100K<n<1M0 likes314 downloads2mo agoHugging Face28OALL /details_Qwen__Qwen3-30B-A3B-Thinking-2507_v2 Dataset Card for Evaluation run of Qwen/Qwen3-30B-A3B-Thinking-2507 Dataset automatically created during the evaluation run of model Qwen/Qwen3-30B-A3B-Thinking-2507. The dataset is composed of 116 configuration, each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Qwen__Qwen3-30B-A3B-Thinking-2507_v2.tabular100K<n<1M0 likes308 downloads8mo agoHugging Face29adrianmele /CodeX-2M-Thinking Modotte Note: This dataset is part of the lineup CodeX by Modotte. You can get lots of datasets in this same lineup, with the main focus on providing very high-quality datasets for model training and fine-tuning. This dataset is fully synthetic, curated from high-quality public sources and enhanced with synthetic data generated using both closed and open-source models. It serves as a strong foundation for instruction-based model tuning and fine-tuning, offering one of the… See the full description on the dataset page: https://huggingface.co/datasets/adrianmele/CodeX-2M-Thinking.texttext-generation1M<n<10M0 likes308 downloads5mo agoHugging Face30marin-community /openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16 OpenThoughts-4 Science SDG: Qwen3-30B-A3B-Thinking-2507 (n=8, top-16 logprobs) Synthetic generations from Qwen/Qwen3-30B-A3B-Thinking-2507 on the Marin OpenThoughts-4 science SDG prompt set. Each prompt is sampled n=8 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes306 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.