datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IndustryBench-MIPU
IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products
Multi-Image Industrial Product Understanding Benchmark — evaluating MLLMs on structured attribute extraction from real-world industrial product images.
Industrial product specifications are scattered across multiple heterogeneous images — specification tables, nameplates, technical drawings. IndustryBench-MIPU tests whether MLLMs can reliably recover them through four… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench-MIPU.clevr-multichange
CLEVR-Multi-Change (30–40 objects)
Two-image change-captioning data used in "Stateful Visual Encoders for
Vision-Language Models" (the Multi-object Visual Differencing task). Each example is a before/after pair of a CLEVR scene
with 30–40 objects and 4 simultaneous changes (add / delete / move /
replace), rendered at 768×768 with a wide camera angle. Built with the
CLEVR-Multi-Change engine (Johnson et al. 2017; Qiu et al. 2021).
Code & paper:… See the full description on the dataset page: https://huggingface.co/datasets/zwcolin/clevr-multichange.africa-synth-aid-flows-medical-multimodal-fracture-all
Africa Synth Aid Flows Medical Multimodal Fracture All | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: json - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-medical-multimodal-fracture-all.CADBench-Extended-Multimodal-Dataset
Dataset Card
Dataset Description
CADBench Extended Multimodal Dataset is an independently produced public extension for multimodal CAD reconstruction research. It contains 100 CAD samples with clean and perturbed meshes, STEP/STL/OBJ/GLB representations, single-view and four-view renders, PBR images, bilingual descriptions, prompt variants, QA, geometry metadata, grading signals, and manually reviewed visual semantics.
Tasks: image-to-text, text-to-image… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/CADBench-Extended-Multimodal-Dataset.MultipanelVQAmultimodal-example
Multimodal Example Dataset
Small example dataset for testing multimodal (vision-language) fine-tuning with ms-swift.
Structure
├── train.jsonl # 10 training samples
├── test.jsonl # 2 validation samples
├── images/ # All referenced images (400x300 JPEG)
│ ├── dog_portrait.jpg
│ ├── forest_river.jpg
│ ├── laptop_desk.jpg
│ ├── mountain_lake.jpg
│ ├── ocean_rocks.jpg
│ ├── coffee_cup.jpg
│ ├── bookshelf.jpg
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/f13rnd/multimodal-example.Multi-Turn
TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents
Official codebase for TCOD, a temporal curriculum framework for on-policy distillation that stabilizes knowledge transfer from teacher to student agents in multi-turn interactive environments.
🔥 News
[2026-07] Our paper is accepted by COLM 2026!
[2026-06] ✍️ New blog post out: on-policy distillation pitfalls — sharing the lessons and pitfalls behind our… See the full description on the dataset page: https://huggingface.co/datasets/SeanWang0027/Multi-Turn.multi_vstar_benchMultiUI
MulitUI
Dataset for the paper: Harnessing Webpage Uis For Text Rich Visual Understanding
🌐 Homepage | 🐍 GitHub | 📖 arXiv
Introduction
We introduce MultiUI, a dataset containing 7.3 million samples from 1 million websites, covering diverse multi- modal tasks and UI layouts. Models trained on MultiUI not only excel in web UI tasks—achieving up to a 48% improvement on VisualWebBench and a 19.1% boost in action accuracy on a web agent dataset Mind2Web—but also… See the full description on the dataset page: https://huggingface.co/datasets/neulab/MultiUI.multimath-300kMedical_Multimodal_Evaluation_Data
Evaluation Guide
This dataset is used to evaluate medical multimodal LLMs, as used in HuatuoGPT-Vision. It includes benchmarks such as VQA-RAD, SLAKE, PathVQA, PMC-VQA, OmniMedVQA, and MMMU-Medical-Tracks.
To get started:
Download the dataset and extract the images.zip file.
Find evaluation code on our GitHub: HuatuoGPT-Vision.
This open-source release aims to simplify the evaluation of medical multimodal capabilities in large models. Please cite the relevant benchmark… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical_Multimodal_Evaluation_Data.DeepFashion-MultiModal-Parts2Whole
DeepFashion MultiModal Parts2Whole
Dataset Details
Dataset Description
This human image dataset comprising about 41,500 reference-target pairs. Each pair in this dataset includes multiple reference images, which encompass human pose images (e.g., OpenPose, Human Parsing, DensePose), various aspects of human appearance (e.g., hair, face, clothes, shoes) with their short textual labels, and a target image featuring the same individual (ID) in the same outfit… See the full description on the dataset page: https://huggingface.co/datasets/huanngzh/DeepFashion-MultiModal-Parts2Whole.GPRadar-Defect-MultiTask
GPRadar-Defect-MultiTask 数据集
本仓库包含用于微调PaLI-GEMMA多模态模型的地质雷达(GPR)缺陷检测数据集。该数据集专注于地下结构中的空洞和裂缝检测与分析。
数据集结构
数据集组织如下:
dataset/
├── annotations/ - 包含JSON和JSONL格式的标注文件
│ ├── _annotations.train.jsonl - 训练集标注
│ ├── _annotations.valid.jsonl - 验证集标注
│ ├── _annotations.test.jsonl - 测试集标注
│ ├── p-1.v1i.paligemma/ - 主数据集元数据
│ └── p-1.v1i.paligemma-multimodal/ - 多模态数据集元数据
├── images/ - 包含所有图像文件
特点
包含874张带注释的地质雷达扫描图像
图像预处理为640x640像素大小
支持多种任务类型:缺陷检测、位置定位和描述生成… See the full description on the dataset page: https://huggingface.co/datasets/xingqiang/GPRadar-Defect-MultiTask.multi-image-composition-instruction-following
Multi-Image Composition Instruction-Following
A large-scale multimodal dataset for multi-image composition via natural language instruction-following. Each case provides 2-3 input images (characters + scene) along with detailed Chinese instructions to compose them into a single photorealistic output image.
Designed for training and evaluating models on complex image composition tasks that require understanding of character identity preservation, pose generation, scene integration… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/multi-image-composition-instruction-following.multimodal-peer-collaboration-samples
Multimodal Peer Collaboration Samples - Embodied Map Task with Two Camera Angles
Two non-experts collaborate to build working circuits under asymmetric information: the instructor has the manual, the student has the components, and synchronized audio and dual-camera video capture how shared understanding emerges.
▶ Watch the interactions · See Expert Instruction samples · Discuss the full collection
Sister collection: Expert Instruction, a teacher and a student in… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-peer-collaboration-samples.DeepFashion-MultiModal-Parts2Whole
DeepFashion MultiModal Parts2Whole
Dataset Details
Dataset Description
This human image dataset comprising about 41,500 reference-target pairs. Each pair in this dataset includes multiple reference images, which encompass human pose images (e.g., OpenPose, Human Parsing, DensePose), various aspects of human appearance (e.g., hair, face, clothes, shoes) with their short textual labels, and a target image featuring the same individual (ID) in the same outfit… See the full description on the dataset page: https://huggingface.co/datasets/LIAGM/DeepFashion-MultiModal-Parts2Whole.svg-multimodal-rubrics
SVG Multimodal Rubrics
A multimodal dataset of SVG code generation samples with natural language descriptions and evaluation rubrics. Each sample pairs a detailed prompt (Markdown) with its corresponding SVG source code, covering animations, 3D scenes, games, and visual effects.
Designed for training and evaluating models on visual code generation — generating complex, interactive SVG artwork from natural language descriptions.
Overview
Item
Details
Samples… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/svg-multimodal-rubrics.multimodal-ir-zh-tw
Multimodal IR Traditional Chinese Dataset
Latest exported JSON files for the multimodal information-retrieval project.
Files
File
Records
Description
SHA-256
multimodal_documents.g4.captioned.s2tw.json
769,245
Latest document corpus with image captions and Simplified-to-Traditional Chinese conversion
e4e3f7609fde3297d89e436a11751dc5a394109f7aba2ae487fe9323d198802d
multimodal_pretrain_pairs.json
351,979
Latest large pretraining/query pairs with text… See the full description on the dataset page: https://huggingface.co/datasets/Sigoso12/multimodal-ir-zh-tw.KomdigiITS-DFK3-Multimodaltoy_data#toy dataset
This is a small portion of the full dataset, used for testing and formatting purposes.
GPRadar-Defect-MultiTask
GPRadar-Defect-MultiTask 数据集
本仓库包含用于微调PaLI-GEMMA多模态模型的地质雷达(GPR)缺陷检测数据集。该数据集专注于地下结构中的空洞和裂缝检测与分析。
数据集结构
数据集组织如下:
dataset/
├── annotations/ - 包含JSON和JSONL格式的标注文件
│ ├── _annotations.train.jsonl - 训练集标注
│ ├── _annotations.valid.jsonl - 验证集标注
│ ├── _annotations.test.jsonl - 测试集标注
│ ├── p-1.v1i.paligemma/ - 主数据集元数据
│ └── p-1.v1i.paligemma-multimodal/ - 多模态数据集元数据
├── images/ - 包含所有图像文件
特点
包含874张带注释的地质雷达扫描图像
图像预处理为640x640像素大小
支持多种任务类型:缺陷检测、位置定位和描述生成… See the full description on the dataset page: https://huggingface.co/datasets/LiZHENGzai/GPRadar-Defect-MultiTask.data_multimodalmultiview-dataset
