datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MixBench25MixBench is a benchmark for evaluating mixed-modality retrieval. It contains queries and corpora from four datasets: MSCOCO, Google_WIT, VisualNews, and OVEN. Each subset provides: query, corpus, mixed_corpus, and qrel splits.Multi-modal-Self-instruct
Dataset Description
Paper Information
Dataset Examples
Leaderboard
Dataset Usage
Data Downloading
Data Format
Evaluation
Citation
You can download the zip dataset directly, and both train and test subsets are collected in Multi-modal-Self-instruct.zip.
Dataset Description
Multi-Modal Self-Instruct dataset utilizes large language models and their code capabilities to synthesize massive abstract images and visual reasoning instructions across daily scenarios. This benchmark… See the full description on the dataset page: https://huggingface.co/datasets/zwq2018/Multi-modal-Self-instruct.ModalityFaultLines-SCEval
SCEval — Modality Fault Lines
Data for Modality Fault Lines: Structural Corruptions Reveal Fragile Omni-Modal Reasoning (Findings of EMNLP 2026).
SCEval is a human-verified benchmark for omni-modal robustness. Text, vision, and audio all remain
present, but controlled corruptions make the evidence inside a channel unreliable. Each corrupted
item is paired with its clean counterpart at the example level, so clean-to-corrupted comparisons are
made on the same underlying question… See the full description on the dataset page: https://huggingface.co/datasets/KZL96/ModalityFaultLines-SCEval.NEXUS-temporal_hierarchical_multi-modal
NEXUS: Neural Evolution for eXtensible Universal Semantics Dataset
(Temporal Multimodal Slices)
This dataset is a multi-modal, hierarchical, temporal representation derived from HuggingFaceFV/finevideo. It is designed for streaming training where the primary unit is a 10 ms "slice" that aggregates upward into moments (100 ms), seconds (1 s), experiences (10 s), and minutes (60 s).
It is meant to represent an extensible stream of "experience" as there are… See the full description on the dataset page: https://huggingface.co/datasets/Ardea/NEXUS-temporal_hierarchical_multi-modal.MixBench2026
MixBench: A Benchmark for Mixed Modality Retrieval
MixBench is a benchmark for evaluating retrieval across text, images, and multimodal documents. It is designed to test how well retrieval models handle queries and documents that span different modalities, such as pure text, pure images, and combined image+text inputs.
MixBench includes four subsets, each curated from a different data source:
MSCOCO
Google_WIT
VisualNews
OVEN
Each subset contains:
queries.jsonl: each entry… See the full description on the dataset page: https://huggingface.co/datasets/mixed-modality-search/MixBench2026.Multi-modal_dataset_named_M3SC
M³SC: A generic dataset for mixed multi-modal (MMM) sensing and communication integration
📌 Overview
M³SC is the first simulation dataset for communication and multi-modal sensing for connected intelligent vehicles, covering multiple typical environments such as urban, suburban, and rural areas, and comprehensively considering scenario conditions including multi-weather, multi-time periods, multi-vehicle traffic densities, multi-frequency bands, and multi-antenna arrays.… See the full description on the dataset page: https://huggingface.co/datasets/pku-pcni-lab/Multi-modal_dataset_named_M3SC.coco-modality-equivalenceshared-emergence-icl-modalities-128
Shared-emergence ICL replication at T=128
This dataset contains the complete raw result archive for the paper
“Many Next-Token Predictors are In-Context Learners.”
The campaign evaluates a fixed suite of 100 program-synthesis tasks using 128
sampled prompts per task, for every clean and deranged shot cell described by
the paper:
21 run keys;
281 experiment cells;
12,800 predictions per cell;
3,596,800 predictions in total.
The archive expands to a top-level results_128/… See the full description on the dataset page: https://huggingface.co/datasets/N8Programs/shared-emergence-icl-modalities-128.coco-modality-equivalencesdc-multi-modal-dataset삼성디스플레이 멀티모달모델 학습을 위한 데이터셋입니다.
24.12.23: Upload 1~303 png files (without 1, 116, 165, 259, 275, 302 png file. 6 files)
AR-Omni-Instruct-v0.1
AR-Omni-Instruct
Overview
AR-Omni-Instruct is a multimodal instruction-tuning dataset for training unified autoregressive any-to-any models.
All modalities are represented as discrete tokens in a single interleaved token stream, enabling standard next-token prediction training over multimodal sequences.
Dataset Summary
Type: multimodal instruction-tuning data
Format: discrete tokenized multimodal conversations / sequences
Use case: instruction tuning… See the full description on the dataset page: https://huggingface.co/datasets/ModalityDance/AR-Omni-Instruct-v0.1.gsm8k-rendered-vlm-v2
GSM8K Rendered-VL v2
1319 rendered GSM8K test problems for the VLM modality study (Phase 1).
Contributors
Rodela Ghosh — study design, pilot (v1), dataset packaging and Hugging Face release (scripts/prepare_hf_v2_release.py)
Aviral Gupta — benchmark infrastructure (src/), v2 rendering protocol (src/rendering.py), Phase 1 model runs
Code: https://github.com/Ro-netizen004/vlm-modality-research
Not interchangeable with v1: RodelaG/gsm8k-rendered-vlm
v1
v2… See the full description on the dataset page: https://huggingface.co/datasets/vlm-modality-research/gsm8k-rendered-vlm-v2.Optical-Reasoning-4k
Optical Reasoning
Overview
Optical Reasoning contains 3,907 rendered visual rationales used in "Optical Reasoning: Rethinking Images as an Expressive Reasoning Medium Beyond Text". It covers 5 benchmarks, including typographic rationales for all benchmarks and graphical rationales for AQuA-RAT.
AQuA-RAT: Multiple-choice algebra and quantitative reasoning problems with five answer options.
GPQA Diamond: Graduate-level multiple-choice science questions spanning… See the full description on the dataset page: https://huggingface.co/datasets/ModalityDance/Optical-Reasoning-4k.multi-modal-image-editMixBench
MixBench: A Benchmark for Mixed Modality Retrieval
MixBench is a benchmark for evaluating retrieval across text, images, and multimodal documents. It is designed to test how well retrieval models handle queries and documents that span different modalities, such as pure text, pure images, and combined image+text inputs.
MixBench includes four subsets, each curated from a different data source:
MSCOCO
Google_WIT
VisualNews
OVEN
Each subset contains:
queries.jsonl: each entry… See the full description on the dataset page: https://huggingface.co/datasets/mixed-modality-search/MixBench.brain_modalitytrue_imgs_modalities
Dataset Summary
ESmike/true_imgs_modalities is a multimodal dataset containing 4,149 real images at a uniform resolution of 512 × 512, along with multiple derived modalities for each image.
This dataset is intended for research and experimentation in computer vision, generative modeling, and multimodal learning.
All images originate from real photographs curated and processed. Each base image is provided alongside additional modality representations (details in the Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/ESmike/true_imgs_modalities.modality-conflict-arbitration-v2
Modality-Conflict Arbitration Benchmark (v2)
A controlled benchmark for studying how a vision-language model arbitrates between
its two input channels when they disagree — and whether that choice tracks the
reliability of each channel.
Each row is a single conflict trial: an image of one math problem paired with the
text of a different problem. Because the two ground-truth answers are carried side by
side, the model's output alone tells you which modality it followed — no… See the full description on the dataset page: https://huggingface.co/datasets/vlm-modality-research/modality-conflict-arbitration-v2.chartqa-evidence-conflict-v2
ChartQA-Conflict v2
This dataset contains 230 reviewed conflicts between a native ChartQA chart and
an evidence-bearing textual report. The original chart supports one answer,
while counterfactual facts in the report support a distinct answer to the same
question. Neither source is privileged in the evaluation prompt.
Each row contains the original chart, shared question, chart-supported answer,
report-supported answer, evidence-bearing report, unit and counterfactual
strategy… See the full description on the dataset page: https://huggingface.co/datasets/vlm-modality-research/chartqa-evidence-conflict-v2.PhysTool-BenchPhysTool-Bench: Beyond APIs: Probing the Limits of MLLMs in Physical Tool Use
🐱 GitHub |
📄 Paper |
🏠 Project Page |
🤗 HuggingFace Papers
📊 Dataset Summary
PhysTool-Bench is a multimodal benchmark designed to evaluate how well Multimodal Large Language Models (MLLMs) perceive, select, and sequence physical tools in real-world scenes. Unlike traditional tool-use benchmarks that focus on digital APIs, this dataset probes an MLLM's ability to ground functional… See the full description on the dataset page: https://huggingface.co/datasets/ModalityDance/PhysTool-Bench.translate-Multi-modal-Self-instruct
Translated https://huggingface.co/datasets/zwq2018/Multi-modal-Self-instruct
Translate to Malay using https://mesolitica.com/translation Base model, a nice dataset for visual QA charts, tables, simulated maps, dashboards, flowcharts, relation graphs, floor plans, and visual puzzles.
chartqa-evidence-conflict-table-v1
ChartQA-Conflict chart/table representation ablation
This derivative contains 229 audited ChartQA-Conflict items. Each row provides
two visual renderings of the same official ChartQA facts:
chart_image: the original ChartQA chart;
table_image: a plain table image rendered from the corresponding official
ChartQA CSV.
The shared question, chart-supported answer, evidence-bearing conflicting
report, report-supported answer, and Source A/B assignment are identical across
the two… See the full description on the dataset page: https://huggingface.co/datasets/vlm-modality-research/chartqa-evidence-conflict-table-v1.Modality-Alignmultimodal-modality-conflict-datasetconflicting_modalitymodalheroicons_modalAIDAS-Omni-Modal-Diffusion-assetsmulti-modal_offensive_memevlm-cross-modal-reps
