CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ShareLab-SII /thinking_fmb_dataset_lerobot_output_qwen3vlimage1M<n<10M0 likes6.3k downloads6mo agoHugging Face02OpenDataArena /MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking MMFineReason-Full-2.3M The Complete Pre-Selection Dataset — Before Quality Filtering 📖 Overview MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering. 🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M65 likes5.4k downloads8mo agoHugging Face03ShareLab-SII /thinking_droid_lerobot_output_qwen3vlimage1M<n<10M0 likes4.1k downloads5mo agoHugging Face04ThinkingHub /PP SteelBench: A Diagnostic Benchmark for Vision-Language Models in Industrial Safety Monitoring SteelBench is a diagnostic benchmark of densely annotated CCTV clips from an operating integrated steel plant. It is designed to evaluate vision-language models (VLMs) on real-world industrial action recognition, PPE assessment, and safety-violation detection — under naturally occurring degradation (dust, glare, steam, low light), at distances and crowdedness levels that curated… See the full description on the dataset page: https://huggingface.co/datasets/ThinkingHub/PP.imagevideo-classification1K<n<10K0 likes3.1k downloads3mo agoHugging Face05ShareLab-SII /thinking_furniture_bench_dataset_lerobot_output_qwen3vlimage1M<n<10M0 likes1.8k downloads6mo agoHugging Face06OpenDataArena /MMFineReason-SFT-586K-Qwen3-VL-235B-Thinking MMFineReason-SFT-586K The Hardest 33% — Less Data, More Reasoning 📖 Overview MMFineReason-SFT-586K is a difficulty-filtered subset of MMFineReason-1.8M, containing the hardest 33% of samples where Qwen3-VL-4B-Thinking do not consistently succeed. (pass rate ≠ 1). Specifically, this subset removes all easy samples (pass rate = 1) under Qwen3-VL-4B-Thinking, retaining only instances that require non-trivial multimodal reasoning. 🎯 Key Highlights 586K… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-SFT-586K-Qwen3-VL-235B-Thinking.image100K<n<1M6 likes1.6k downloads8mo agoHugging Face07OpenDataArena /MMFineReason-1.8M-Qwen3-VL-235B-Thinking MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal understanding benchmarks. 📖 Overview MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking. 🎯 Key Highlights 1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-1.8M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M126 likes1.6k downloads7mo agoHugging Face08ericktwo /MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking MMFineReason-Full-2.3M The Complete Pre-Selection Dataset — Before Quality Filtering 📖 Overview MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering. 🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/ericktwo/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M1 likes1.4k downloads8mo agoHugging Face09NarsAI /FineReason-1.8M-Qwen3-VL-235B-Thinking MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal understanding benchmarks. 📖 Overview MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking. 🎯 Key Highlights 1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/NarsAI/FineReason-1.8M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M0 likes1.3k downloads8mo agoHugging Face10Sandeepthakur /MMFineReason-1.8M-Qwen3-VL-235B-Thinking MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal understanding benchmarks. 📖 Overview MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking. 🎯 Key Highlights 1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/Sandeepthakur/MMFineReason-1.8M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M0 likes1k downloads8mo agoHugging Face11UCSC-VLAA /VLAA-Thinking SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models 🌐 Project Page • 📄 Arxiv • 💻 Code 🤗 VLAA-Thinker Family • 🤔 VLAA-Thinking Dataset 🤗 VLAA-Thinker-Qwen2.5-3B • 🤗 VLAA-Thinker-Qwen2.5-7B Both VLAA-Thinker-Qwen2.5-3B and VLAA-Thinker-Qwen2.5-7Bachieve SOTA performance on OpenCompass Multimodal Reasoning Leaderboard as of April 7th, 2025. Contents Quick Start 🚀… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/VLAA-Thinking.documentvisual-question-answeringn<1K20 likes819 downloads1y agoHugging Face12ShareLab-SII /thinking_stanford_hydra_dataset_lerobot_output_qwen3vlimage100K<n<1M0 likes776 downloads6mo agoHugging Face13OpenDataArena /MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking MMFineReason-SFT-123K The Hardest 7% — Less Data, More Reasoning 📖 Overview MMFineReason-SFT-123K is a difficulty-filtered subset of MMFineReason-1.8M, containing only the hardest 7% of samples where Qwen3-VL-4B-Thinking consistently fails (pass rate = 0). 🎯 Key Highlights 123K Challenging Samples: Only instances where a 4B thinking model fails all 4 inference attemptsEfficient Training: Comparable performance to full 1.8M dataset with only 7% of… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking.imagevisual-question-answering100K<n<1M86 likes588 downloads8mo agoHugging Face14ShareLab-SII /thinking_taco_play_lerobot_output_qwen3vlimage100K<n<1M0 likes382 downloads6mo agoHugging Face15dans25275 /MMFineReason-1.8M-Qwen3-VL-235B-Thinking MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal understanding benchmarks. 📖 Overview MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking. 🎯 Key Highlights 1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/dans25275/MMFineReason-1.8M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M0 likes290 downloads8mo agoHugging Face16AmirhoseinGH /mhlc-training-qwen3vl-qwen3_vl_2b_thinking_hard_mixed_sources_120k Multi Head Latent Control Training Data - Qwen3-VL 2B Thinking hard Mixed Sources 120k Dataset Description This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection. Paper https://arxiv.org/abs/2607.14277 Code https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control Dataset Summary Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_2b_thinking_hard_mixed_sources_120k.imagequestion-answering100K<n<1M0 likes281 downloads2mo agoHugging Face17ThinkingRM /Generation-Reviewimagen<1K0 likes267 downloads3mo agoHugging Face18AmirhoseinGH /mhlc-training-qwen3vl-qwen3_vl_4b_thinking_hard_mixed_sources_120k Multi Head Latent Control Training Data - Qwen3-VL 4B Thinking hard Mixed Sources 120k Dataset Description This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection. Paper https://arxiv.org/abs/2607.14277 Code https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control Dataset Summary Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_4b_thinking_hard_mixed_sources_120k.imagequestion-answering100K<n<1M0 likes244 downloads2mo agoHugging Face19ShareLab-SII /thinking_dobbe_lerobot_output_qwen3vlimage1M<n<10M0 likes151 downloads6mo agoHugging Face20ShareLab-SII /thinking_berkeley_autolab_ur5_lerobot_output_qwen3vlimage10K<n<100K0 likes124 downloads6mo agoHugging Face21ShareLab-SII /thinking_jaco_playimage10K<n<100K0 likes117 downloads6mo agoHugging Face22drwlf /medraN-thinking-1024image1M<n<10M0 likes113 downloads1y agoHugging Face23newyccku /nycc-thinkingimage1K<n<10K1 likes113 downloads5mo agoHugging Face24BRZ911 /Thinking-in-Video-DataThinking in Video: Can Video Generators Really Reason About the Real World? This repository contains the official implementation of Causal-Generative Dual-Judge (CGDJ) for auditing world-model consistency of video generative models — the official codebase of the Thinking in Video paradigm. 🌟 Overview Thinking in Video is a reasoning paradigm in which a video generative model is used not merely to synthesize pixels, but to simulate, predict, and verify causal… See the full description on the dataset page: https://huggingface.co/datasets/BRZ911/Thinking-in-Video-Data.image1K<n<10K0 likes102 downloads2mo agoHugging Face25ShareLab-SII /thinking_toto_lerobot_output_qwen3vlimage100K<n<1M0 likes101 downloads6mo agoHugging Face26ShareLab-SII /thinking_jaco_play_lerobot_output_qwen3vlimage10K<n<100K0 likes91 downloads6mo agoHugging Face27penfever /vlaa-thinking-grpo VLAA-Thinking-SFT-126K Large-scale vision-language dataset with 126K instruction-following samples featuring chain-of-thought reasoning Dataset Description This dataset contains vision-language samples with instruction-following conversations. Each sample includes: image: PIL Image object question: Question or instruction text answer or gt: Response with thinking process (SFT dataset) or ground truth answer (GRPO dataset) caption: Image caption (may be empty for some… See the full description on the dataset page: https://huggingface.co/datasets/penfever/vlaa-thinking-grpo.image10K<n<100K0 likes88 downloads1y agoHugging Face28eyes-ml /MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking-QR-max4096 Derived dataset note This dataset was derived from OpenDataArena/MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking as a part of arxiv.org/abs/2603.22276. Field changes: question -> query qwen3vl_235b_thinking_response -> response image -> images (single-item list) added tok_len, computed with tokenizer Qwen/Qwen3-8B on query + '\n\n' + response add_special_tokens=False The original README content is preserved below. MMFineReason-SFT-123K The Hardest 7% — Less Data, More Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/eyes-ml/MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking-QR-max4096.imagevisual-question-answering10K<n<100K0 likes82 downloads6mo agoHugging Face29novastar112 /pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot PushT norm4 Visual Nomarker All-Step Thinking Trickiness COT This dataset is derived from successful PushT visual-nomarker trajectories in novastar112/pusht_96_norm4_visual_nomarker. Each row contains one full successful trajectory from the first move through the final stop action. Main files: training/pusht_allstep_thinking_cot.jsonl.gz: 500,000 train rows. testing/pusht_allstep_thinking_cot.jsonl.gz: 200 test rows. metadata/final_scan_validation.json: full local scan after repair… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot.imageimage-to-text100K<n<1M0 likes79 downloads4mo agoHugging Face30olob0 /finevision-mini-thinking FineVision-mini Thinking FineVision-mini is a slice I made of HuggingFaceM4/FineVision: 169 of its image subsets, 101,321 rows (~40 GB) out of FineVision's 24.2M rows / 4.65 TB (about 0.4% of the rows, 0.9% of the bytes), sampled with a fixed seed. The 16 text-only subsets were left out. This dataset is that slice, fully translated and augmented with reasoning, published in increments: each batch processes more rows of FineVision-mini and is appended here, until the whole slice… See the full description on the dataset page: https://huggingface.co/datasets/olob0/finevision-mini-thinking.imagevisual-question-answering10K<n<100K0 likes76 downloads6d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.