datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MMStar
MMStar (Are We on the Right Way for Evaluating Large Vision-Language Models?)
🌐 Homepage | 🤗 Dataset | 🤗 Paper | 📖 arXiv | GitHub
Dataset Details
As shown in the figure below, existing benchmarks lack consideration of the vision dependency of evaluation samples and potential data leakage from LLMs' and LVLMs' training data.
Therefore, we introduce MMStar: an elite vision-indispensible multi-modal benchmark, aiming to ensure each curated sample exhibits… See the full description on the dataset page: https://huggingface.co/datasets/Lin-Chen/MMStar.MMSU
[ICLR 2026] MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
Overview of MMSU
MMSU (Massive Multi-task Spoken Language Understanding and Reasoning Benchmark) is a comprehensive benchmark for evaluating fine-grained spoken language understanding and reasoning in multimodal models.
It systematically captures the variance of real-world linguistic phenomena in daily speech through 47 sub-tasks, including phonetics, prosody, rhetoric… See the full description on the dataset page: https://huggingface.co/datasets/ddwang2000/MMSU.MMScan-betammsulab
DiscoPhon - Segmented MMS ulab v2
This dataset is a segmented version of espnet/mms_ulab_v2
using pyannote/segmentation-3.0.
License and Acknowledgement
Following espnet/mms_ulab_v2, this dataset is released under the
Creative Commons Attribution-NonCommercial-ShareAlike 4.0
license.
If you use this dataset, please cite the DiscoPhon paper
@misc{poli2026discophon,
title={{DiscoPhon}: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech… See the full description on the dataset page: https://huggingface.co/datasets/coml/mmsulab.MM-SpuBench
MM-SpuBench Datacard
Basic Information
Title: The Multimodal Spurious Benchmark (MM-SpuBench)
Description: MM-SpuBench is a comprehensive benchmark designed to evaluate the robustness of MLLMs to spurious biases. This benchmark systematically assesses how well these models distinguish between core and spurious features, providing a detailed framework for understanding and quantifying spurious biases.
Data Structure:
├── data/images
│ ├── 000000.jpg
│ ├── 000001.jpg
│… See the full description on the dataset page: https://huggingface.co/datasets/mmbench/MM-SpuBench.MM-SafetyBenchWarning: This dataset may contain sensitive or harmful content. Users are advised to handle it with care and ensure that their use complies with relevant ethical guidelines and legal requirements.
Usage and License Notices: The dataset is intended and licensed for research use only. They are also restricted to uses that follow the license agreement GPT-4 and Stable Diffusion. The dataset is CC BY NC 4.0 (allowing only non-commercial use).
Data Source: For more information about the dataset… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/MM-SafetyBench.af3-mmseqs-dbMMSearch-PlusMMSearch-Plus
MMSearch-Plus✨: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents
Official repository for the paper "MMSearch-Plus: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents".
🌟 For more details, please refer to the project page with examples: https://mmsearch-plus.github.io/.
[🌐 Webpage] [📖 Paper] [🤗 Huggingface Dataset] [🏆 Leaderboard]
💥 News
[2025.09.26] 🔥 We update the arXiv paperand release all MMSearch-Plus data samples in… See the full description on the dataset page: https://huggingface.co/datasets/Cie1/MMSearch-Plus.yuto-mms-multimodal
Dataset Card for YUTO MMS Multimodal (MCAP)
A FiftyOne build of YUTO MMS (York University Teledyne Optech Mobile
Mapping System Dataset), a SLAM benchmark from the AUSM Lab at York
University. This build repackages the source dataset's per-sequence raw
sensor folders as time-synchronized MCAP recordings for
FiftyOne's native multimodal dataset
support (FiftyOne
1.19+). Each sample is one episode (one continuous drive), viewable in
FiftyOne's tiled multimodal viewer with a… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/yuto-mms-multimodal.MMSI-Bench
MMSI-Bench
This repo contains evaluation code for the paper "MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence"
🌐 Homepage | 🤗 Dataset | 📑 Paper | 💻 Code | 📖 arXiv
🔔News
🔥[2025-10-23]: We added the normalized human response time for each MMSI-Bench sample and its difficulty level to our dataset on Hugging Face.
🔥[2025-06-18]: MMSI-Bench has been supported in the LMMs-Eval repository.
✨[2025-06-11]: MMSI-Bench was used for evaluation in the… See the full description on the dataset page: https://huggingface.co/datasets/RunsenXu/MMSI-Bench.MMSI-Video-Bench
MMSI-Video-Bench: A Holistic Benchmark for Video-Based Spatial Intelligence
🌐 Homepage | 📑 Paper | 📖 Code
🔔 News
🔥[2025-12]: Our MMSI-Video-Bench has been integrated into VLMEvalKit.
🔥[2025-12]: We released our paper, benchmark, and evaluation codes.
📊 Data Details
All of our data is available on Hugging Face and includes the following components:
🎥 Video Data (videos.zip): Contains the video clip file (.mp4) corresponding to each sample. This… See the full description on the dataset page: https://huggingface.co/datasets/rbler/MMSI-Video-Bench.MMScan-llava-form
MMScan LLaVA-Form Data
This repository provides the processed LLaVA-formatted dataset for the MMScan Question Answering Benchmark.
Dataset Contents
(1) All image data(Depth&RGB) is distributed in split ZIP archives. Please combine the split ZIP files into a single archive and extract the merged ZIP file using the following command:
cat mmscan_val8.z* > mmscan_va.zip
unzip mmscan_va.zip
(2) Under ./annotations, we provide the MMScan Question Answering validation set with… See the full description on the dataset page: https://huggingface.co/datasets/rbler/MMScan-llava-form.mmsu-ci-2000mms_ulab_v2MMS ulab v2 is a a massively multilingual speech dataset that contains 8900 hours of unlabeled speech across 4023 languages. In total, it contains 189 language families.
It can be used for language identification, spoken language modelling, or speech representation learning.
MMS ulab v2 is a reproduced and extended version of the MMS ulab dataset originally proposed in Scaling Speech Technology to 1000+ Languages, covering more languages and containing more data.
This dataset includes the raw… See the full description on the dataset page: https://huggingface.co/datasets/espnet/mms_ulab_v2.MM-SafetyBench-plus-plus
MM-SafetyBench++
Project Page | Paper | Code
MM-SafetyBench++ is a benchmark designed for evaluating contextual safety in Multi-Modal Large Language Models (MLLMs). It challenges models to distinguish subtle contextual differences between scenarios that may appear visually or textually similar but diverge significantly in safety intent.
Dataset Summary
For each unsafe image-text pair, the benchmark includes a corresponding safe counterpart created through minimal… See the full description on the dataset page: https://huggingface.co/datasets/EchoSafe-MLLM/MM-SafetyBench-plus-plus.MMSVG-IllustrationOmniSVG: A Unified Scalable Vector Graphics Generation Model
![Project Page]
Dataset Card for MMSVG-Illustration
Dataset Description
This dataset contains SVG illustration examples for training and evaluating SVG models for text-to-SVG and image-to-SVG task.
Dataset Structure
Features
The dataset contains the following fields:
Field Name
Description
id
Unique ID for each SVG
svg
SVG code (resized to 200×200, simplified with picosvg)… See the full description on the dataset page: https://huggingface.co/datasets/OmniSVG/MMSVG-Illustration.MMS-VPR
MMS-VPR: A Fine-Grained Multimodal Street-Level Visual Place Recognition Dataset and Evaluation Benchmark for Dense Pedestrian Environments
Overview
MMS-VPR is the first large-scale multimodal street-level visual place recognition dataset featuring comprehensive integration of images, videos, and rich textual annotations with day–night coverage and a 7-year temporal span in dense pedestrian-only environments.
MMS-VPR comprises 110,529 images and 2,527 video clips… See the full description on the dataset page: https://huggingface.co/datasets/Yiwei-Ou/MMS-VPR.MMSearch
MMSearch 🔥: Benchmarking the Potential of Large Models as Multi-modal Search Engines
Official repository for the paper "MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines".
🌟 For more details, please refer to the project page with dataset exploration and visualization tools: https://mmsearch.github.io/.
[🌐 Webpage] [📖 Paper] [🤗 Huggingface Dataset] [🏆 Leaderboard] [🔍 Visualization]
💥 News
[2024.09.25] 🌟 The evaluation code now… See the full description on the dataset page: https://huggingface.co/datasets/CaraJ/MMSearch.HR-MMSearch
Dataset Description
HR-MMSearch is a benchmark designed to evaluate the Agentic Reasoning and Search capabilities of Multimodal Large Language Models in complex visual tasks.
This dataset was introduced by SenseTime Research in the paper SenseNova-MARS: Empowering Multimodal Agentic Reasoning and Search via Reinforcement Learning.
Key Features:
High-Resolution Images: Contains high-resolution image inputs, requiring the model to possess fine-grained visual perception… See the full description on the dataset page: https://huggingface.co/datasets/sensenova/HR-MMSearch.mmstarmmsci_valDayhoff-MMseqs2
Dayhoff FASTA and MMseqs2 databases
This dataset contains the original Dayhoff Atlas GigaRef and UniRef50 datasets, in formats amenable to MMSeqs2 CPU and GPU utilities.
The train, validation, and test sets from the original atlas were combined and the following datasets available:
GigaRef No Singletons - The GigaRef dataset, with no singleton clusters.
GigaRef Singletons - The GigaRef dataset, with only singleton clusters.
GigaRef Full - Every sequence contained in both the… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/Dayhoff-MMseqs2.mmskills
Multimodal Skill Packages for General Visual Agents
515 public skill packages across Ubuntu desktop, macOS, Minecraft, and Mario environments.
Overview |
Preview |
Contents |
Download |
Format |
Statistics |
Citation
Overview
This Hugging Face repository hosts the public MMSkills data packages: reusable multimodal procedural skills for visual agents. Each skill combines:
a concise SKILL.md procedure;
runtime_state_cards.json with state-matching… See the full description on the dataset page: https://huggingface.co/datasets/zhangkangning/mmskills.hmblogs-v3MMSD2.0
MMSD2.0: Towards a Reliable Multi-modal Sarcasm Detection System
This is a copy of the dataset uploaded on Hugging Face for easy access. The original data comes from this work, which is an improvement upon a previous study.
Usage
from typing import TypedDict, cast
import pytorch_lightning as pl
from datasets import Dataset, load_dataset
from torch import Tensor
from torch.utils.data import DataLoader
from transformers import CLIPProcessor
class… See the full description on the dataset page: https://huggingface.co/datasets/coderchen01/MMSD2.0.MMSVG-IconOmniSVG: A Unified Scalable Vector Graphics Generation Model
![Project Page]
Dataset Card for MMSVG-Icon
Dataset Description
This dataset contains SVG icon examples for training and evaluating SVG models for text-to-SVG and image-to-SVG task.
Dataset Structure
Features
The dataset contains the following fields:
Field Name
Description
id
Unique ID for each SVG
svg
SVG code (resized to 200×200, simplified with picosvg)
description… See the full description on the dataset page: https://huggingface.co/datasets/OmniSVG/MMSVG-Icon.K-MMStar
K-MMStar
We introduce K-MMStar, a Korean adaptation of the MMStar [1] designed for evaluating vision-language models.
By translating the val subset of MMStar into Korean and carefully reviewing its naturalness through human inspection, we developed a novel robust evaluation benchmark specifically for Korean language.
(We observe that there are unanswerable cases (e.g., multiple images required to answer the question but only has a single image, vague questions or options) in the… See the full description on the dataset page: https://huggingface.co/datasets/NCSOFT/K-MMStar.MMS-e
MMS-e: Benchmarking the Resilience of Large Multimodal Models to Visual Scrambling
Benchmark Examples
Patchwise Question Answering: Divide the images into 2x2, 4x4, and 8x8 patches, then shuffle all the patches, and measure the ability of LMMs to answer questions about these images.
Reconstruction task: Let LMMs reconstruct the order of shuffled patches based on the image' s caption, and let LMMs reconstruct the shuffled caption based on the image.
Fixed Patch… See the full description on the dataset page: https://huggingface.co/datasets/jyjyjyjy/MMS-e.mmscore
