datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking
MMFineReason-Full-2.3M
The Complete Pre-Selection Dataset — Before Quality Filtering
📖 Overview
MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering.
🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.MAmmoTH-VL-Instruct-12M
MAmmoTH-VL-Instruct-12M
🏠 Homepage | 🤖 MAmmoTH-VL-8B | 💻 Code | 📄 Arxiv | 📕 PDF | 🖥️ Demo
Introduction
Our simple yet scalable visual instruction data rewriting pipeline consists of three steps: manual data source collection, rewriting using MLLMs/LLMs, and filtering via the same MLLM as a judge. Examples below illustrate transformations in math and science categories, showcasing detailed, step-by-step responses.
The data distribution of… See the full description on the dataset page: https://huggingface.co/datasets/MAmmoTH-VL/MAmmoTH-VL-Instruct-12M.Latex-VLMvlsbench🎉 VLSBench has been accpeted to ACL2025 Main Conference, see you in Vienna.
✅ Update data.json with safety reason and image description for more efficient and reliable evaluaiton.
Dataset Card for VLSBench
This dataset is for paper VLSBench: Unveiling Information Leakage In Multimodal Safety
You can check our Paper, Github, Project Page for more information.
You can directly use this image-text dataset with naive huggingface support:
dataset = load_dataset("Foreshhh/vlsbench"… See the full description on the dataset page: https://huggingface.co/datasets/Foreshhh/vlsbench.vlmsareblindArXiv - Website
MMFineReason-1.8M-Qwen3-VL-235B-Thinking
MMFineReason
Closing the Multimodal Reasoning Gap via Open Data-Centric Methods
Average score across mathematical reasoning and multimodal understanding benchmarks.
📖 Overview
MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking.
🎯 Key Highlights
1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-1.8M-Qwen3-VL-235B-Thinking.MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking
MMFineReason-Full-2.3M
The Complete Pre-Selection Dataset — Before Quality Filtering
📖 Overview
MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering.
🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/ericktwo/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.FineReason-1.8M-Qwen3-VL-235B-Thinking
MMFineReason
Closing the Multimodal Reasoning Gap via Open Data-Centric Methods
Average score across mathematical reasoning and multimodal understanding benchmarks.
📖 Overview
MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking.
🎯 Key Highlights
1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/NarsAI/FineReason-1.8M-Qwen3-VL-235B-Thinking.DocVQA-2026
DocVQA 2026 | ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains
Building upon previous DocVQA benchmarks, this evaluation dataset introduces challenging reasoning questions over a diverse collection of documents spanning eight domains, including business reports, scientific papers, slides, posters, maps, comics, infographics, and engineering drawings.
By expanding coverage to new document domains and… See the full description on the dataset page: https://huggingface.co/datasets/VLR-CVC/DocVQA-2026.VL-Health
VL-Health Dataset
Overview
The VL-Health dataset is designed for multi-stage training of unified LVLMs in the medical domain. It consists of two key phases:
Alignment – Focused on training image captioning capabilities and learning representations of input visual information.
Instruct Fine-Tuning – Designed for enhancing the model's ability to handle various vision-language tasks, including both visual comprehension and visual generation tasks.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/lintw/VL-Health.MedReason
MedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs
📃 Paper |🤗 MedReason-8B | 📚 MedReason Data
✨ Latest News
[05/27/2025] 🎉 MedReason wins 3rd prize🏆 in the Huggingface Reasoning Datasets Competition!
⚡Introduction
MedReason is a large-scale high-quality medical reasoning dataset designed to enable faithful and explainable medical problem-solving in large language models (LLMs).
We utilize a structured medical knowledge… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/MedReason.MMFineReason-1.8M-Qwen3-VL-235B-Thinking
MMFineReason
Closing the Multimodal Reasoning Gap via Open Data-Centric Methods
Average score across mathematical reasoning and multimodal understanding benchmarks.
📖 Overview
MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking.
🎯 Key Highlights
1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/Sandeepthakur/MMFineReason-1.8M-Qwen3-VL-235B-Thinking.TransCity-VLM-dataset
TransCity-VLM Dataset
The TransCity-VLM Dataset provides multimodal smart-city data for traffic, energy, mobility, grid operation, urban context understanding, and map-grounded question answering. It supports the training and evaluation of vision-language models for urban prediction, decision support, conversational QA, and reasoning tasks.
The training data are available at this Hugging Face dataset repository.
Dataset Summary
Split
Rows / Files
test JSONL… See the full description on the dataset page: https://huggingface.co/datasets/TransCity-VLM/TransCity-VLM-dataset.XLRS-Bench-lite_VLM
🐙GitHub
Information or evaluatation on this dataset can be found in this repo: https://github.com/AI9Stars/XLRS-Bench
📜Dataset License
Annotations of this dataset is released under a Creative Commons Attribution-NonCommercial 4.0 International License. For images from:
DOTARGB images from Google Earth and CycloMedia (for academic use only; commercial use is prohibited, and Google Earth terms of use apply).
ITCVDLicensed under CC-BY-NC-SA-4.0.
MiniFrance… See the full description on the dataset page: https://huggingface.co/datasets/initiacms/XLRS-Bench-lite_VLM.MedTrinity-25M
Tutorial of using Medtrinity-25M
MedTrinity-25M, a comprehensive, large-scale multimodal dataset for medicine, covering over 25 million images across 10 modalities, with multigranular annotations for more than 65 diseases. These enriched annotations encompass both global textual information, such as disease/lesion type, modality, region-specific descriptions, and inter-regional relationships, as well as detailed local annotations for regions of interest (ROIs), including bounding… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/MedTrinity-25M.VisualClawArena
VisualClawArena
Webpage: https://ucsc-vlaa.github.io/VisualClaw/ |
Paper: https://arxiv.org/abs/2606.16295 |
Code: https://github.com/UCSC-VLAA/VisualClaw
VisualClawArena is a 200-scenario benchmark for multimodal computer-use agents.
Each scenario pairs a video clip with a persistent workspace, dynamic updates,
multi-round instructions, and executable checkers. The benchmark is designed to
test whether an agent can use video evidence, workspace files, and later
environment… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/VisualClawArena.MIRB
Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning
File Structure
├── MIR
|── analogy.json
│── codeu.json
|── dataset_namex.json
└── Images
├── analogy
│ └── image_x.jpg
└──codeu
└── image_x.jpg
JSON Structure
{
"questions": " What is the expected kurtosis of the sequence created by`create_number_sequence(-10, 10)`?\n\n1.… See the full description on the dataset page: https://huggingface.co/datasets/VLLMs/MIRB.Arabic-VLM-Full-Pearl
💎 The Arabic VLM Dataset (Full Pearl Edition)
This repository contains the full, unreviewed dataset comprising 309K multimodal examples. This data was generated automatically using the agentic pipeline developed for the Pearl project, as described in our paper.
Disclaimer: This is the raw, synthetic data that has not been subject to human review. It was generated as part of the data creation process and is released for research purposes. It may contain noise, errors, or… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/Arabic-VLM-Full-Pearl.MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking
MMFineReason-SFT-123K
The Hardest 7% — Less Data, More Reasoning
📖 Overview
MMFineReason-SFT-123K is a difficulty-filtered subset of MMFineReason-1.8M, containing only the hardest 7% of samples where Qwen3-VL-4B-Thinking consistently fails (pass rate = 0).
🎯 Key Highlights
123K Challenging Samples: Only instances where a 4B thinking model fails all 4 inference attemptsEfficient Training: Comparable performance to full 1.8M dataset with only 7% of… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking.VLRS-Bench
VLRS-Bench
A Vision-Language Reasoning Benchmark for Remote Sensing
VLRS-Bench is a remote sensing benchmark designed to evaluate whether multimodal large language models can reason over geospatial scenes, rather than only recognize objects or classify land cover.
Cognition
Causal reasoning over spatial structure
Decision
Planning under operational constraints… See the full description on the dataset page: https://huggingface.co/datasets/thislzm/VLRS-Bench.MAmmoTH-VL-Instruct-12M
MAmmoTH-VL-Instruct-12M used in MoCa Pre-training
🏠 Homepage | 💻 Code | 🤖 MoCa-Qwen25VL-7B | 🤖 MoCa-Qwen25VL-3B | 📚 Datasets | 📄 Paper
Introduction
This is a VQA style dataset used in the modality-aware continual pre-training of MoCa models. It is adapted from MAmmoTH-VL-Instruct-12M by concatenating prompts and responses.
The dataset consists of interleaved multimodal examples. text is a string containing text while imagesare image binaries that can be loaded… See the full description on the dataset page: https://huggingface.co/datasets/moca-embed/MAmmoTH-VL-Instruct-12M.efficient-vla-extracts
efficient-vla-extracts
Parser-derived artifacts for the efficient-vla-wiki project.
This dataset is the data layer companion to the main repository:
GitHub: guanweifan/efficient-vla-wiki
Dataset repo: efficient-vla-extracts
What is inside
The staged upload keeps the local directory structure:
extracts/
├── meta/
│ ├── extract_build_index.jsonl
│ └── extract_build_status.json
└── parses/
└── <paper_id>/
Each paper directory may contain:
pdftotext.txt… See the full description on the dataset page: https://huggingface.co/datasets/guanweifan/efficient-vla-extracts.vlm-projects-multi-lang-final-v2
My Final Multilingual Medical VQA Dataset
This dataset is organized into multiple configurations (subsets), one for each language.
You can load a specific language subset like this:
from datasets import load_dataset
# Load the Vietnamese training data
vi_train = load_dataset("tungvu3196/vlm-projects-multi-lang-final-v2", "Vietnamese", split="train")
# Load the English testing data
en_test = load_dataset("tungvu3196/vlm-projects-multi-lang-final-v2", "English", split="test")
mhlc-training-qwen3vl-qwen3_vl_2b_thinking_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3-VL 2B Thinking hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_2b_thinking_hard_mixed_sources_120k.mhlc-training-qwen3vl-qwen3-vl-32b-instruct_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3-VL 32B Instruct hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3-vl-32b-instruct_hard_mixed_sources_120k.mhlc-training-qwen3vl-qwen3_vl_4b_instruct_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3-VL 4B Instruct hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_4b_instruct_hard_mixed_sources_120k.mhlc-training-qwen3vl-qwen3_vl_4b_thinking_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3-VL 4B Thinking hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_4b_thinking_hard_mixed_sources_120k.mhlc-training-qwen3vl-qwen3_vl_2b_instruct_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3-VL 2B Instruct hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_2b_instruct_hard_mixed_sources_120k.VLM-ExecRouterBench
VLM-ExecRouterBench
An execution-oriented benchmark for cost-aware open-set VLM routing.
Cost-aware routing |
Open-set model onboarding |
Multimodal, code, and search tasks
Overview
VLM-ExecRouterBench is an execution-oriented benchmark for routing
vision-language model queries to a pool of candidate VLMs. Each sample is
executed by multiple candidate models, producing correctness labels, inference
costs, metadata… See the full description on the dataset page: https://huggingface.co/datasets/Kirito-Lab/VLM-ExecRouterBench.LHPR-VLN
Long-Horizon Planning and Reasoning in VLN (LHPR-VLN) Benchmark
This repository contains the long-horizon planning and reasoning in VLN (LHPR-VLN) benchmark, introduced in Towards Long-Horizon Vision-Language Navigation: Platform, Benchmark and Method.
Dataset is also available in ModelScope.
Files
The trajectory dataset file organization is as follows. The task/ directory contains the complete LH-VLN task trajectories, while the step_task/… See the full description on the dataset page: https://huggingface.co/datasets/Starry123/LHPR-VLN.
