CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OpenDataArena /MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking MMFineReason-Full-2.3M The Complete Pre-Selection Dataset — Before Quality Filtering 📖 Overview MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering. 🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M65 likes4.1k downloads8mo agoHugging Face02MAmmoTH-VL /MAmmoTH-VL-Instruct-12M MAmmoTH-VL-Instruct-12M 🏠 Homepage | 🤖 MAmmoTH-VL-8B | 💻 Code | 📄 Arxiv | 📕 PDF | 🖥️ Demo Introduction Our simple yet scalable visual instruction data rewriting pipeline consists of three steps: manual data source collection, rewriting using MLLMs/LLMs, and filtering via the same MLLM as a judge. Examples below illustrate transformations in math and science categories, showcasing detailed, step-by-step responses. The data distribution of… See the full description on the dataset page: https://huggingface.co/datasets/MAmmoTH-VL/MAmmoTH-VL-Instruct-12M.imagevisual-question-answering10M<n<100M67 likes3.7k downloads2y agoHugging Face03JosselinSom /Latex-VLMimagequestion-answering1K<n<10K9 likes3.3k downloads3y agoHugging Face04Foreshhh /vlsbench🎉 VLSBench has been accpeted to ACL2025 Main Conference, see you in Vienna. ✅ Update data.json with safety reason and image description for more efficient and reliable evaluaiton. Dataset Card for VLSBench This dataset is for paper VLSBench: Unveiling Information Leakage In Multimodal Safety You can check our Paper, Github, Project Page for more information. You can directly use this image-text dataset with naive huggingface support: dataset = load_dataset("Foreshhh/vlsbench"… See the full description on the dataset page: https://huggingface.co/datasets/Foreshhh/vlsbench.imagequestion-answering1K<n<10K8 likes2.6k downloads1y agoHugging Face05XAI /vlmsareblindArXiv - Website imagequestion-answering1K<n<10K28 likes1.8k downloads2y agoHugging Face06OpenDataArena /MMFineReason-1.8M-Qwen3-VL-235B-Thinking MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal understanding benchmarks. 📖 Overview MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking. 🎯 Key Highlights 1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-1.8M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M126 likes1.7k downloads7mo agoHugging Face07ericktwo /MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking MMFineReason-Full-2.3M The Complete Pre-Selection Dataset — Before Quality Filtering 📖 Overview MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering. 🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/ericktwo/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M1 likes1.4k downloads8mo agoHugging Face08NarsAI /FineReason-1.8M-Qwen3-VL-235B-Thinking MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal understanding benchmarks. 📖 Overview MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking. 🎯 Key Highlights 1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/NarsAI/FineReason-1.8M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M0 likes1.4k downloads8mo agoHugging Face09VLR-CVC /DocVQA-2026 DocVQA 2026 | ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains Building upon previous DocVQA benchmarks, this evaluation dataset introduces challenging reasoning questions over a diverse collection of documents spanning eight domains, including business reports, scientific papers, slides, posters, maps, comics, infographics, and engineering drawings. By expanding coverage to new document domains and… See the full description on the dataset page: https://huggingface.co/datasets/VLR-CVC/DocVQA-2026.imagevisual-question-answeringn<1K74 likes1.3k downloads19d agoHugging Face10lintw /VL-Health VL-Health Dataset Overview The VL-Health dataset is designed for multi-stage training of unified LVLMs in the medical domain. It consists of two key phases: Alignment – Focused on training image captioning capabilities and learning representations of input visual information. Instruct Fine-Tuning – Designed for enhancing the model's ability to handle various vision-language tasks, including both visual comprehension and visual generation tasks. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/lintw/VL-Health.textquestion-answering100K<n<1M18 likes1.2k downloads2y agoHugging Face11Sandeepthakur /MMFineReason-1.8M-Qwen3-VL-235B-Thinking MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal understanding benchmarks. 📖 Overview MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking. 🎯 Key Highlights 1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/Sandeepthakur/MMFineReason-1.8M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M0 likes1k downloads8mo agoHugging Face12UCSC-VLAA /MedReason MedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs 📃 Paper |🤗 MedReason-8B | 📚 MedReason Data ✨ Latest News [05/27/2025] 🎉 MedReason wins 3rd prize🏆 in the Huggingface Reasoning Datasets Competition! ⚡Introduction MedReason is a large-scale high-quality medical reasoning dataset designed to enable faithful and explainable medical problem-solving in large language models (LLMs). We utilize a structured medical knowledge… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/MedReason.textquestion-answering10K<n<100K88 likes1k downloads1y agoHugging Face13UCSC-VLAA /MedTrinity-25Mgated Tutorial of using Medtrinity-25M MedTrinity-25M, a comprehensive, large-scale multimodal dataset for medicine, covering over 25 million images across 10 modalities, with multigranular annotations for more than 65 diseases. These enriched annotations encompass both global textual information, such as disease/lesion type, modality, region-specific descriptions, and inter-regional relationships, as well as detailed local annotations for regions of interest (ROIs), including bounding… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/MedTrinity-25M.imagequestion-answering10M<n<100M214 likes602 downloads2y agoHugging Face14MohamedRashad /Arabic-VLM-Full-Pearl 💎 The Arabic VLM Dataset (Full Pearl Edition) This repository contains the full, unreviewed dataset comprising 309K multimodal examples. This data was generated automatically using the agentic pipeline developed for the Pearl project, as described in our paper. Disclaimer: This is the raw, synthetic data that has not been subject to human review. It was generated as part of the data creation process and is released for research purposes. It may contain noise, errors, or… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/Arabic-VLM-Full-Pearl.imagequestion-answering100K<n<1M10 likes473 downloads10mo agoHugging Face15initiacms /XLRS-Bench-lite_VLM 🐙GitHub Information or evaluatation on this dataset can be found in this repo: https://github.com/AI9Stars/XLRS-Bench 📜Dataset License Annotations of this dataset is released under a Creative Commons Attribution-NonCommercial 4.0 International License. For images from: DOTARGB images from Google Earth and CycloMedia (for academic use only; commercial use is prohibited, and Google Earth terms of use apply). ITCVDLicensed under CC-BY-NC-SA-4.0. MiniFrance… See the full description on the dataset page: https://huggingface.co/datasets/initiacms/XLRS-Bench-lite_VLM.textvisual-question-answering1K<n<10K0 likes432 downloads11mo agoHugging Face16OpenDataArena /MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking MMFineReason-SFT-123K The Hardest 7% — Less Data, More Reasoning 📖 Overview MMFineReason-SFT-123K is a difficulty-filtered subset of MMFineReason-1.8M, containing only the hardest 7% of samples where Qwen3-VL-4B-Thinking consistently fails (pass rate = 0). 🎯 Key Highlights 123K Challenging Samples: Only instances where a 4B thinking model fails all 4 inference attemptsEfficient Training: Comparable performance to full 1.8M dataset with only 7% of… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking.imagevisual-question-answering100K<n<1M86 likes420 downloads8mo agoHugging Face17moca-embed /MAmmoTH-VL-Instruct-12M MAmmoTH-VL-Instruct-12M used in MoCa Pre-training 🏠 Homepage | 💻 Code | 🤖 MoCa-Qwen25VL-7B | 🤖 MoCa-Qwen25VL-3B | 📚 Datasets | 📄 Paper Introduction This is a VQA style dataset used in the modality-aware continual pre-training of MoCa models. It is adapted from MAmmoTH-VL-Instruct-12M by concatenating prompts and responses. The dataset consists of interleaved multimodal examples. text is a string containing text while imagesare image binaries that can be loaded… See the full description on the dataset page: https://huggingface.co/datasets/moca-embed/MAmmoTH-VL-Instruct-12M.textvisual-question-answering10M<n<100M1 likes383 downloads1y agoHugging Face18tungvu3196 /vlm-projects-multi-lang-final-v2 My Final Multilingual Medical VQA Dataset This dataset is organized into multiple configurations (subsets), one for each language. You can load a specific language subset like this: from datasets import load_dataset # Load the Vietnamese training data vi_train = load_dataset("tungvu3196/vlm-projects-multi-lang-final-v2", "Vietnamese", split="train") # Load the English testing data en_test = load_dataset("tungvu3196/vlm-projects-multi-lang-final-v2", "English", split="test") imagequestion-answering100K<n<1M0 likes330 downloads1y agoHugging Face19AmirhoseinGH /mhlc-training-qwen3vl-qwen3_vl_2b_thinking_hard_mixed_sources_120k Multi Head Latent Control Training Data - Qwen3-VL 2B Thinking hard Mixed Sources 120k Dataset Description This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection. Paper https://arxiv.org/abs/2607.14277 Code https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control Dataset Summary Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_2b_thinking_hard_mixed_sources_120k.imagequestion-answering100K<n<1M0 likes313 downloads2mo agoHugging Face20AmirhoseinGH /mhlc-training-qwen3vl-qwen3_vl_4b_instruct_hard_mixed_sources_120k Multi Head Latent Control Training Data - Qwen3-VL 4B Instruct hard Mixed Sources 120k Dataset Description This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection. Paper https://arxiv.org/abs/2607.14277 Code https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control Dataset Summary Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_4b_instruct_hard_mixed_sources_120k.imagequestion-answering100K<n<1M0 likes292 downloads2mo agoHugging Face21AmirhoseinGH /mhlc-training-qwen3vl-qwen3-vl-32b-instruct_hard_mixed_sources_120k Multi Head Latent Control Training Data - Qwen3-VL 32B Instruct hard Mixed Sources 120k Dataset Description This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection. Paper https://arxiv.org/abs/2607.14277 Code https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control Dataset Summary Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3-vl-32b-instruct_hard_mixed_sources_120k.imagequestion-answering10K<n<100K0 likes287 downloads2mo agoHugging Face22AmirhoseinGH /mhlc-training-qwen3vl-qwen3_vl_4b_thinking_hard_mixed_sources_120k Multi Head Latent Control Training Data - Qwen3-VL 4B Thinking hard Mixed Sources 120k Dataset Description This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection. Paper https://arxiv.org/abs/2607.14277 Code https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control Dataset Summary Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_4b_thinking_hard_mixed_sources_120k.imagequestion-answering100K<n<1M0 likes281 downloads2mo agoHugging Face23AmirhoseinGH /mhlc-training-qwen3vl-qwen3_vl_2b_instruct_hard_mixed_sources_120k Multi Head Latent Control Training Data - Qwen3-VL 2B Instruct hard Mixed Sources 120k Dataset Description This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection. Paper https://arxiv.org/abs/2607.14277 Code https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control Dataset Summary Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_2b_instruct_hard_mixed_sources_120k.imagequestion-answering100K<n<1M0 likes242 downloads2mo agoHugging Face24dans25275 /MMFineReason-1.8M-Qwen3-VL-235B-Thinking MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal understanding benchmarks. 📖 Overview MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking. 🎯 Key Highlights 1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/dans25275/MMFineReason-1.8M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M0 likes201 downloads8mo agoHugging Face25zGinger /XLRS-Bench-lite_VLM 🐙GitHub Information or evaluatation on this dataset can be found in this repo: https://github.com/AI9Stars/XLRS-Bench 📜Dataset License Annotations of this dataset is released under a Creative Commons Attribution-NonCommercial 4.0 International License. For images from: DOTARGB images from Google Earth and CycloMedia (for academic use only; commercial use is prohibited, and Google Earth terms of use apply). ITCVDLicensed under CC-BY-NC-SA-4.0. MiniFrance… See the full description on the dataset page: https://huggingface.co/datasets/zGinger/XLRS-Bench-lite_VLM.textvisual-question-answering1K<n<10K0 likes191 downloads9mo agoHugging Face26multilingual-vlm-conflict /code-conflict Code Conflict Dataset A dataset of 100 visual Python code conflict samples designed to evaluate Vision-Language Models (VLMs) under cross-modal conflicts (discrepancy between code screenshots and caption text). Dataset Statistics Total Rows: 100 samples Language: English (english) Categories: 5 distinct Python code conflict_types (20 samples per category): operator_substitution (Rows 1–20): Swapping math or logic operators (e.g., + to -, == to !=, or to and).… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-vlm-conflict/code-conflict.imagevisual-question-answering1K<n<10K0 likes170 downloads3mo agoHugging Face27TimeOmni-VL /TSUMM-Suite_Training 🧱 TSUMM-Suite Training Set This repository provides the training set of TSUMM-Suite, a multimodal dataset for unified time series understanding and generation. It contains forecasting and imputation tasks for TS-image generation. By aligning understanding tasks with generation tasks, TSUMM-Suite enables temporal understanding to improve time series generation. 🎨 Task Illustration TSUMM-Suite… See the full description on the dataset page: https://huggingface.co/datasets/TimeOmni-VL/TSUMM-Suite_Training.imagequestion-answering10K<n<100K1 likes149 downloads3mo agoHugging Face28UCSC-VLAA /m23k-tokenized m1: Unleash the Potential of Test-Time Scaling for Medical Reasoning in Large Language Models A simple test-time scaling strategy, with minimal fine-tuning, can unlock strong medical reasoning within large language models. ⚡ Introduction Hi! Welcome to the huggingface repository for m1 (https://github.com/UCSC-VLAA/m1)! m1 is a medical LLM designed to enhance reasoning through efficient test-time scaling. It enables lightweight models to match or exceed the performance of… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/m23k-tokenized.textquestion-answering10K<n<100K9 likes145 downloads1y agoHugging Face29multilingual-vlm-conflict /rpg-conflict RPG Fantasy Battle Conflict Dataset A dataset of 100 visual RPG combat conflict samples designed to evaluate Vision-Language Models (VLMs) under cross-modal conflicts (discrepancy between battle screenshots and caption text). Derived from the rcannizzaro/rpg_fantasy_battle_counterfactual_v2 dataset. Dataset Statistics This dataset consists of a single train split containing 100 perfectly isolated conflict samples derived from the… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-vlm-conflict/rpg-conflict.imagevisual-question-answering1K<n<10K0 likes131 downloads3mo agoHugging Face30eyes-ml /MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking-QR-max4096 Derived dataset note This dataset was derived from OpenDataArena/MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking as a part of arxiv.org/abs/2603.22276. Field changes: question -> query qwen3vl_235b_thinking_response -> response image -> images (single-item list) added tok_len, computed with tokenizer Qwen/Qwen3-8B on query + '\n\n' + response add_special_tokens=False The original README content is preserved below. MMFineReason-SFT-123K The Hardest 7% — Less Data, More Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/eyes-ml/MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking-QR-max4096.imagevisual-question-answering10K<n<100K0 likes80 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.