datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IndustryBench-MIPU
IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products
Multi-Image Industrial Product Understanding Benchmark — evaluating MLLMs on structured attribute extraction from real-world industrial product images.
Industrial product specifications are scattered across multiple heterogeneous images — specification tables, nameplates, technical drawings. IndustryBench-MIPU tests whether MLLMs can reliably recover them through four… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench-MIPU.McEvalMcEval benchmark data as described in the McEval Paper. Code for the evaluation can be found on Github as McEval.
IfEvalCode-testsetCADBench-Extended-Multimodal-Dataset
Dataset Card
Dataset Description
CADBench Extended Multimodal Dataset is an independently produced public extension for multimodal CAD reconstruction research. It contains 100 CAD samples with clean and perturbed meshes, STEP/STL/OBJ/GLB representations, single-view and four-view renders, PBR images, bilingual descriptions, prompt variants, QA, geometry metadata, grading signals, and manually reviewed visual semantics.
Tasks: image-to-text, text-to-image… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/CADBench-Extended-Multimodal-Dataset.africa-synth-aid-flows-medical-multimodal-fracture-all
Africa Synth Aid Flows Medical Multimodal Fracture All | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: json - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-medical-multimodal-fracture-all.IfEvalCode-InstructMdEval
MDEVAL: Massively Multilingual Code Debugging
Official repository for our paper "MDEVAL: Massively Multilingual Code Debugging"
🏠 Home Page •
📊 Benchmark Data •
🏆 Leaderboard
Introduction
MDEVAL is a massively multilingual debugging benchmark covering 20 programming languages with 3.9K test samples and three tasks focused on bug fixing. It substantially pushes the limits of code LLMs in multilingual… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Multimodal-NLP/MdEval.Medical_Multimodal_Evaluation_Data
Evaluation Guide
This dataset is used to evaluate medical multimodal LLMs, as used in HuatuoGPT-Vision. It includes benchmarks such as VQA-RAD, SLAKE, PathVQA, PMC-VQA, OmniMedVQA, and MMMU-Medical-Tracks.
To get started:
Download the dataset and extract the images.zip file.
Find evaluation code on our GitHub: HuatuoGPT-Vision.
This open-source release aims to simplify the evaluation of medical multimodal capabilities in large models. Please cite the relevant benchmark… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical_Multimodal_Evaluation_Data.McEval-InstructMcEval-Instruct data as described in the McEval Paper. Code for the evaluation and sft can be found on Github as McEval.
DeepFashion-MultiModal-Parts2Whole
DeepFashion MultiModal Parts2Whole
Dataset Details
Dataset Description
This human image dataset comprising about 41,500 reference-target pairs. Each pair in this dataset includes multiple reference images, which encompass human pose images (e.g., OpenPose, Human Parsing, DensePose), various aspects of human appearance (e.g., hair, face, clothes, shoes) with their short textual labels, and a target image featuring the same individual (ID) in the same outfit… See the full description on the dataset page: https://huggingface.co/datasets/huanngzh/DeepFashion-MultiModal-Parts2Whole.TableInstruct
Citation
@misc{wu2024tablebenchcomprehensivecomplexbenchmark,
title={TableBench: A Comprehensive and Complex Benchmark for Table Question Answering},
author={Xianjie Wu and Jian Yang and Linzheng Chai and Ge Zhang and Jiaheng Liu and Xinrun Du and Di Liang and Daixin Shu and Xianfu Cheng and Tianzhen Sun and Guanglin Niu and Tongliang Li and Zhoujun Li},
year={2024},
eprint={2408.09174},
archivePrefix={arXiv},
primaryClass={cs.CL}… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Multimodal-NLP/TableInstruct.AutoMemoryBench
AutoMemoryBench
State-Contract Evaluation for Auditable Agent Memory
AutoMemoryBench evaluates whether an agent uses the right memory, and only
the admissible memory, under a query-time state contract. Each executable
contract partitions memory into required, admissible, and
prohibited sets. Prohibited memories are typed as superseded, deleted,
restricted, cross-namespace, or stale-tool.
Relevance is not enough: remembered evidence must also be… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Multimodal-NLP/AutoMemoryBench.multimodality-poc-llama31-ruler16k
Multimodality PoC corpus — Llama-3.1-8B-Instruct on RULER-16K
Raw pre-RoPE query and hidden-state tensors captured during prefill, used
to study whether the per-(layer, kv_head) query distribution is unimodal
Gaussian (the assumption underpinning Expected Attention's MGF closed-form
in kvpress).
What's in here
65 .npz files, one per (RULER task, prompt_index) pair (13 tasks × 5
prompts).
Each file (~414 MB) contains:
field
dtype
shape
meaning
hidden
float16… See the full description on the dataset page: https://huggingface.co/datasets/June30916/multimodality-poc-llama31-ruler16k.agent-spaces-tracesmultimodal-peer-collaboration-samples
Multimodal Peer Collaboration Samples - Embodied Map Task with Two Camera Angles
Two non-experts collaborate to build working circuits under asymmetric information: the instructor has the manual, the student has the components, and synchronized audio and dual-camera video capture how shared understanding emerges.
▶ Watch the interactions · See Expert Instruction samples · Discuss the full collection
Sister collection: Expert Instruction, a teacher and a student in… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-peer-collaboration-samples.multimodal-video-annotation-samples
Video Annotation Samples – SuperviseLab
SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates our annotation methodology and output quality across diverse video content categories.
Note: All visual assets in this dataset have been abstracted (pixelated mosaic) to protect source privacy. Uploader identity, original titles, and all identifiable metadata have been removed. This is a demonstration dataset… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/multimodal-video-annotation-samples.WeThink-Multimodal-Reasoning-120K
WeThink-Multimodal-Reasoning-120K
Image Type
Images data can be access from https://huggingface.co/datasets/Xkev/LLaVA-CoT-100k
Image Type
Source Dataset
Images
General Images
COCO
25,344
SAM-1B
18,091
Visual Genome
4,441
GQA
3,251
PISC
835
LLaVA
134
Text-Intensive Images
TextVQA
25,483
ShareTextVQA
538
DocVQA
4,709
OCR-VQA5,142
ChartQA
21,781
Scientific & Technical
GeoQA+
4,813
ScienceQA
4,990
AI2D
1,812
CLEVR-Math
677… See the full description on the dataset page: https://huggingface.co/datasets/WeThink/WeThink-Multimodal-Reasoning-120K.WeThink_Multimodal_Reasoning_120K
Dataset Card for WeThink
Repository: https://github.com/yangjie-cv/WeThink
Paper: https://arxiv.org/abs/2506.07905
Dataset Structure
Question-Answer Pairs
The WeThink_Multimodal_Reasoning_120K.jsonl file contains the question-answering data in the following format:
{
"problem": "QUESTION",
"answer": "ANSWER",
"category": "QUESTION TYPE",
"abilities": "QUESTION REQUIRED ABILITIES",
"refined_cot": "THINK PROCESS",
"image_path": "IMAGE PATH"… See the full description on the dataset page: https://huggingface.co/datasets/yangjie-cv/WeThink_Multimodal_Reasoning_120K.hse-multimodal-rag-corpus
HSE Multimodal RAG Corpus
Chunks, labeled QA (including out-of-scope abstention), and published retrieval metrics.
chunks.jsonl
qa_pairs.jsonl
eval_results.json
benchmark_report.json
multimodal-document-retrieval-20260911-dataset
Multimodal Document Retrieval Baseline Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Business documents contain meaning in text, tables, layout, and imagery that text-only retrieval can miss.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/multimodal-document-retrieval-20260911-dataset.multimodalpragmatic
Multimodal Pragmatic Jailbreak on Text-to-image Models
Project page | Paper | Code
The Multimodal Pragmatic Unsafe Prompts (MPUP) is a dataset designed to assess the multimodal pragmatic safety in Text-to-Image (T2I) models.
It comprises two key sections: image_prompt, and text_prompt.
Dataset Usage
Downloading the Data
To download the dataset, install Huggingface Datasets and then use the following command:
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/tongliuphysics/multimodalpragmatic.multimodal-document-retrieval-20260901-dataset
Multimodal Document Retrieval Baseline Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Business documents contain meaning in text, tables, layout, and imagery that text-only retrieval can miss.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/multimodal-document-retrieval-20260901-dataset.repro-mm-deepresearch-a-simple-and-effective-multimodal-agentic-search-baseline-traces
Agent traces
Agent sessions published from a Trackio Logbook.
multimodal-document-retrieval-20260822-dataset
Multimodal Document Retrieval Baseline Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Business documents contain meaning in text, tables, layout, and imagery that text-only retrieval can miss.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/multimodal-document-retrieval-20260822-dataset.KomdigiITS-DFK3-Multimodalmultimodal-phi-masking-benchmark
multimodal-phi-masking-benchmark
10,000 synthetic clinical records paired with token-level PHI spans, masking decisions, cryptographic audit hashes, RL reward signals, and leakage scores. Five configs covering text, ASR, imaging, waveform, and audio modalities. The only public dataset pairing PHI masking decisions with FHIR R4 audit trails, RL reward signals, and a formally modeled adversarial evasion scenario.
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/vkatg/multimodal-phi-masking-benchmark.lumos_multimodal_ground_iterative
🪄 Agent Lumos: Unified and Modular Training for Open-Source Language Agents
🌐[Website]
📝[Paper]
🤗[Data]
🤗[Model]
🤗[Demo]
We introduce 🪄Lumos, Language Agents with Unified Formats, Modular Design, and Open-Source LLMs. Lumos unifies a suite of complex interactive tasks and achieves competitive performance with GPT-4/3.5-based and larger open-source agents.
Lumos has following features:
🧩 Modular Architecture:
🧩 Lumos consists of planning, grounding… See the full description on the dataset page: https://huggingface.co/datasets/ai2lumos/lumos_multimodal_ground_iterative.multimodal-ai-taxonomy
Multimodal AI Taxonomy
A comprehensive, structured taxonomy for mapping multimodal AI model capabilities across input and output modalities.
Dataset Description
This dataset provides a systematic categorization of multimodal AI capabilities, enabling users to:
Navigate the complex landscape of multimodal AI models
Filter models by specific input/output modality combinations
Understand the nuanced differences between similar models (e.g., image-to-video with/without audio… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/multimodal-ai-taxonomy.lumos_multimodal_plan_iterative
🪄 Agent Lumos: Unified and Modular Training for Open-Source Language Agents
🌐[Website]
📝[Paper]
🤗[Data]
🤗[Model]
🤗[Demo]
We introduce 🪄Lumos, Language Agents with Unified Formats, Modular Design, and Open-Source LLMs. Lumos unifies a suite of complex interactive tasks and achieves competitive performance with GPT-4/3.5-based and larger open-source agents.
Lumos has following features:
🧩 Modular Architecture:
🧩 Lumos consists of planning, grounding… See the full description on the dataset page: https://huggingface.co/datasets/ai2lumos/lumos_multimodal_plan_iterative.multimodal-ir-zh-tw
Multimodal IR Traditional Chinese Dataset
Latest exported JSON files for the multimodal information-retrieval project.
Files
File
Records
Description
SHA-256
multimodal_documents.g4.captioned.s2tw.json
769,245
Latest document corpus with image captions and Simplified-to-Traditional Chinese conversion
e4e3f7609fde3297d89e436a11751dc5a394109f7aba2ae487fe9323d198802d
multimodal_pretrain_pairs.json
351,979
Latest large pretraining/query pairs with text… See the full description on the dataset page: https://huggingface.co/datasets/Sigoso12/multimodal-ir-zh-tw.
