datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MMSU
[ICLR 2026] MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
Overview of MMSU
MMSU (Massive Multi-task Spoken Language Understanding and Reasoning Benchmark) is a comprehensive benchmark for evaluating fine-grained spoken language understanding and reasoning in multimodal models.
It systematically captures the variance of real-world linguistic phenomena in daily speech through 47 sub-tasks, including phonetics, prosody, rhetoric… See the full description on the dataset page: https://huggingface.co/datasets/ddwang2000/MMSU.mmsulab
DiscoPhon - Segmented MMS ulab v2
This dataset is a segmented version of espnet/mms_ulab_v2
using pyannote/segmentation-3.0.
License and Acknowledgement
Following espnet/mms_ulab_v2, this dataset is released under the
Creative Commons Attribution-NonCommercial-ShareAlike 4.0
license.
If you use this dataset, please cite the DiscoPhon paper
@misc{poli2026discophon,
title={{DiscoPhon}: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech… See the full description on the dataset page: https://huggingface.co/datasets/coml/mmsulab.mmsu-ci-2000mms_ulab_v2MMS ulab v2 is a a massively multilingual speech dataset that contains 8900 hours of unlabeled speech across 4023 languages. In total, it contains 189 language families.
It can be used for language identification, spoken language modelling, or speech representation learning.
MMS ulab v2 is a reproduced and extended version of the MMS ulab dataset originally proposed in Scaling Speech Technology to 1000+ Languages, covering more languages and containing more data.
This dataset includes the raw… See the full description on the dataset page: https://huggingface.co/datasets/espnet/mms_ulab_v2.mmsumms_ulab_v2MMS ulab v2 is a a massively multilingual speech dataset that contains 8900 hours of unlabeled speech across 4023 languages. In total, it contains 189 language families.
It can be used for language identification, spoken language modelling, or speech representation learning.
MMS ulab v2 is a reproduced and extended version of the MMS ulab dataset originally proposed in Scaling Speech Technology to 1000+ Languages, covering more languages and containing more data.
This dataset includes the raw… See the full description on the dataset page: https://huggingface.co/datasets/amine-khelif/mms_ulab_v2.mms_ulab_v2MMS ulab v2 is a a massively multilingual speech dataset that contains 8900 hours of unlabeled speech across 4023 languages. In total, it contains 189 language families.
It can be used for language identification, spoken language modelling, or speech representation learning.
MMS ulab v2 is a reproduced and extended version of the MMS ulab dataset originally proposed in Scaling Speech Technology to 1000+ Languages, covering more languages and containing more data.
This dataset includes the raw… See the full description on the dataset page: https://huggingface.co/datasets/sundram1996/mms_ulab_v2.mms-tts-uig-script_arabic-UQSpeech
Single Speaker Uyghur Quran Recordings
This dataset is from https://github.com/gheyret/UQSpeechDataset
{UQSpeech,
author = {Gheyret Kenji},
title = {UQ Awaz Ambiri},
howpublished = {https://github.com/gheyret/UQSpeechDataset/},
year = 2019
}
mms-tts-uig-script_arabic-UQSpeechMMSU-full_5k_hf_format.v0mm_speechThis dataset contains speech(wave files) from a single woman and the tsv file contain transcript of the speech files
mms-multilingual-audio-5to30min
Multilingual Audio Dataset (5-30min)
This dataset contains continuous speech audio files for various languages (ranging from 5 to 30 minutes in length per language) collected from diverse sources including Hugging Face and YouTube.
Dataset Statistics
Total Languages: 100
Sources: HF Omnilingual ASR Corpus, YouTube
Language Details
Language Code
Language Name
Source
Duration (seconds)
jpn
Japanese
youtube
1459.84
eng
English
youtube… See the full description on the dataset page: https://huggingface.co/datasets/aoiandroid/mms-multilingual-audio-5to30min.mms-tts-uig-script_latin-UQSpeechmms-tts-uig-script_latin-UQSpeech7kinyarwanda-mms-datamms-tts-uig-script_latin-UQSpeech6DisfluencySpeech-preprocessed-mms-ttsne-tts-mms-tgj
NE-TTS MMS-VITS Tagin (tgj)
MMS-VITS fine-tuning subset for Tagin (tgj). Contains 80 high-quality clips selected from the cleaned NE-TTS dataset (SNR >= 15dB (relaxed)), formatted for MMS-VITS fine-tuning.
Stats
Metric
Value
Clips
80
Sample rate
22050Hz
SNR filter
SNR >= 15dB (relaxed)
Source
ne-tts-tgj
Schema
Column
Type
Description
audio
Audio
22050Hz WAV audio
text
string
Cleaned transcript… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-tts-mms-tgj.mms-tts-fon-accentfinene-tts-mms-nag
NE-TTS MMS-VITS Nagamese (nag)
MMS-VITS fine-tuning subset for Nagamese (nag). Contains 150 high-quality clips selected from the cleaned NE-TTS dataset (SNR >= 20dB), formatted for MMS-VITS fine-tuning.
Stats
Metric
Value
Clips
150
Sample rate
22050Hz
SNR filter
SNR >= 20dB
Source
ne-tts-nag
Schema
Column
Type
Description
audio
Audio
22050Hz WAV audio
text
string
Cleaned transcript
Usage
Use… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-tts-mms-nag.mms-tts-fon-accentfine-part-5VB_MMSU-full_3074_hf_format.v0tigrinya-mms-ttsne-tts-mms-grt
NE-TTS MMS-VITS Garo (grt)
MMS-VITS fine-tuning subset for Garo (grt). Contains 150 high-quality clips selected from the cleaned NE-TTS dataset (SNR >= 20dB), formatted for MMS-VITS fine-tuning.
Stats
Metric
Value
Clips
150
Sample rate
22050Hz
SNR filter
SNR >= 20dB
Source
ne-tts-grt
Schema
Column
Type
Description
audio
Audio
22050Hz WAV audio
text
string
Cleaned transcript
Usage
Use with… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-tts-mms-grt.mms_synthetic_audiomms-tts-uig-script_latin-UQSpeech2mms-tts-uig-script_latin-UQSpeech4mms_synthetic_audiomms-tts-uig-script_latin-UQSpeech3infore1-mms-vits
