MMS
Datasets
All datasets matching “MMS”MMStar
MMStar (Are We on the Right Way for Evaluating Large Vision-Language Models?)
🌐 Homepage | 🤗 Dataset | 🤗 Paper | 📖 arXiv | GitHub
Dataset Details
As shown in the figure below, existing benchmarks lack consideration of the vision dependency of evaluation samples and potential data leakage from LLMs' and LVLMs' training data.
Therefore, we introduce MMStar: an elite vision-indispensible multi-modal benchmark, aiming to ensure each curated sample exhibits… See the full description on the dataset page: https://huggingface.co/datasets/Lin-Chen/MMStar.MMScan-betaMMSU
[ICLR 2026] MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
Overview of MMSU
MMSU (Massive Multi-task Spoken Language Understanding and Reasoning Benchmark) is a comprehensive benchmark for evaluating fine-grained spoken language understanding and reasoning in multimodal models.
It systematically captures the variance of real-world linguistic phenomena in daily speech through 47 sub-tasks, including phonetics, prosody, rhetoric… See the full description on the dataset page: https://huggingface.co/datasets/ddwang2000/MMSU.mmsulab
DiscoPhon - Segmented MMS ulab v2
This dataset is a segmented version of espnet/mms_ulab_v2
using pyannote/segmentation-3.0.
License and Acknowledgement
Following espnet/mms_ulab_v2, this dataset is released under the
Creative Commons Attribution-NonCommercial-ShareAlike 4.0
license.
If you use this dataset, please cite the DiscoPhon paper
@misc{poli2026discophon,
title={{DiscoPhon}: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech… See the full description on the dataset page: https://huggingface.co/datasets/coml/mmsulab.MM-SpuBench
MM-SpuBench Datacard
Basic Information
Title: The Multimodal Spurious Benchmark (MM-SpuBench)
Description: MM-SpuBench is a comprehensive benchmark designed to evaluate the robustness of MLLMs to spurious biases. This benchmark systematically assesses how well these models distinguish between core and spurious features, providing a detailed framework for understanding and quantifying spurious biases.
Data Structure:
├── data/images
│ ├── 000000.jpg
│ ├── 000001.jpg
│… See the full description on the dataset page: https://huggingface.co/datasets/mmbench/MM-SpuBench.af3-mmseqs-db
