CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ddwang2000 /MMSU [ICLR 2026] MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark Overview of MMSU MMSU (Massive Multi-task Spoken Language Understanding and Reasoning Benchmark) is a comprehensive benchmark for evaluating fine-grained spoken language understanding and reasoning in multimodal models. It systematically captures the variance of real-world linguistic phenomena in daily speech through 47 sub-tasks, including phonetics, prosody, rhetoric… See the full description on the dataset page: https://huggingface.co/datasets/ddwang2000/MMSU.audioquestion-answering1K<n<10K16 likes3.2k downloads5mo agoHugging Face02coml /mmsulab DiscoPhon - Segmented MMS ulab v2 This dataset is a segmented version of espnet/mms_ulab_v2 using pyannote/segmentation-3.0. License and Acknowledgement Following espnet/mms_ulab_v2, this dataset is released under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 license. If you use this dataset, please cite the DiscoPhon paper @misc{poli2026discophon, title={{DiscoPhon}: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech… See the full description on the dataset page: https://huggingface.co/datasets/coml/mmsulab.audioaudio-to-audio1M<n<10M2 likes2.5k downloads6mo agoHugging Face03zhaochenyang20 /mmsu-ci-2000audio1K<n<10K0 likes1.1k downloads5mo agoHugging Face04espnet /mms_ulab_v2MMS ulab v2 is a a massively multilingual speech dataset that contains 8900 hours of unlabeled speech across 4023 languages. In total, it contains 189 language families. It can be used for language identification, spoken language modelling, or speech representation learning. MMS ulab v2 is a reproduced and extended version of the MMS ulab dataset originally proposed in Scaling Speech Technology to 1000+ Languages, covering more languages and containing more data. This dataset includes the raw… See the full description on the dataset page: https://huggingface.co/datasets/espnet/mms_ulab_v2.audioaudio-to-audio10K<n<100K27 likes953 downloads2y agoHugging Face05ngqtrung /mmsuaudio1K<n<10K0 likes217 downloads8mo agoHugging Face06amine-khelif /mms_ulab_v2MMS ulab v2 is a a massively multilingual speech dataset that contains 8900 hours of unlabeled speech across 4023 languages. In total, it contains 189 language families. It can be used for language identification, spoken language modelling, or speech representation learning. MMS ulab v2 is a reproduced and extended version of the MMS ulab dataset originally proposed in Scaling Speech Technology to 1000+ Languages, covering more languages and containing more data. This dataset includes the raw… See the full description on the dataset page: https://huggingface.co/datasets/amine-khelif/mms_ulab_v2.audioaudio-to-audio10K<n<100K0 likes169 downloads6mo agoHugging Face07sundram1996 /mms_ulab_v2MMS ulab v2 is a a massively multilingual speech dataset that contains 8900 hours of unlabeled speech across 4023 languages. In total, it contains 189 language families. It can be used for language identification, spoken language modelling, or speech representation learning. MMS ulab v2 is a reproduced and extended version of the MMS ulab dataset originally proposed in Scaling Speech Technology to 1000+ Languages, covering more languages and containing more data. This dataset includes the raw… See the full description on the dataset page: https://huggingface.co/datasets/sundram1996/mms_ulab_v2.audioaudio-to-audio10K<n<100K0 likes129 downloads6mo agoHugging Face08ixxan /mms-tts-uig-script_arabic-UQSpeech Single Speaker Uyghur Quran Recordings This dataset is from https://github.com/gheyret/UQSpeechDataset {UQSpeech, author = {Gheyret Kenji}, title = {UQ Awaz Ambiri}, howpublished = {https://github.com/gheyret/UQSpeechDataset/}, year = 2019 } audiotext-to-speech10K<n<100K3 likes110 downloads2y agoHugging Face09binbin123 /mms-tts-uig-script_arabic-UQSpeechaudio10K<n<100K0 likes80 downloads1y agoHugging Face10yuantuo666 /MMSU-full_5k_hf_format.v0audio1K<n<10K1 likes46 downloads9mo agoHugging Face11aungmyatv8 /mm_speechThis dataset contains speech(wave files) from a single woman and the tsv file contain transcript of the speech files audio1K<n<10K0 likes38 downloads4y agoHugging Face12aoiandroid /mms-multilingual-audio-5to30min Multilingual Audio Dataset (5-30min) This dataset contains continuous speech audio files for various languages (ranging from 5 to 30 minutes in length per language) collected from diverse sources including Hugging Face and YouTube. Dataset Statistics Total Languages: 100 Sources: HF Omnilingual ASR Corpus, YouTube Language Details Language Code Language Name Source Duration (seconds) jpn Japanese youtube 1459.84 eng English youtube… See the full description on the dataset page: https://huggingface.co/datasets/aoiandroid/mms-multilingual-audio-5to30min.audioautomatic-speech-recognitionn<1K0 likes35 downloads3mo agoHugging Face13zzk123 /mms-tts-uig-script_latin-UQSpeechaudion<1K0 likes30 downloads2y agoHugging Face14zzk123 /mms-tts-uig-script_latin-UQSpeech7audion<1K0 likes30 downloads2y agoHugging Face15Professor /kinyarwanda-mms-dataaudio10K<n<100K0 likes28 downloads8mo agoHugging Face16zzk123 /mms-tts-uig-script_latin-UQSpeech6audion<1K0 likes27 downloads2y agoHugging Face17Meshwa /DisfluencySpeech-preprocessed-mms-ttsaudio1K<n<10K0 likes25 downloads2y agoHugging Face18sulabhkatiyar /ne-tts-mms-tgj NE-TTS MMS-VITS Tagin (tgj) MMS-VITS fine-tuning subset for Tagin (tgj). Contains 80 high-quality clips selected from the cleaned NE-TTS dataset (SNR >= 15dB (relaxed)), formatted for MMS-VITS fine-tuning. Stats Metric Value Clips 80 Sample rate 22050Hz SNR filter SNR >= 15dB (relaxed) Source ne-tts-tgj Schema Column Type Description audio Audio 22050Hz WAV audio text string Cleaned transcript… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-tts-mms-tgj.audiotext-to-speechn<1K0 likes19 downloads4mo agoHugging Face19TitanSage02 /mms-tts-fon-accentfineaudio1K<n<10K1 likes18 downloads1y agoHugging Face20sulabhkatiyar /ne-tts-mms-nag NE-TTS MMS-VITS Nagamese (nag) MMS-VITS fine-tuning subset for Nagamese (nag). Contains 150 high-quality clips selected from the cleaned NE-TTS dataset (SNR >= 20dB), formatted for MMS-VITS fine-tuning. Stats Metric Value Clips 150 Sample rate 22050Hz SNR filter SNR >= 20dB Source ne-tts-nag Schema Column Type Description audio Audio 22050Hz WAV audio text string Cleaned transcript Usage Use… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-tts-mms-nag.audiotext-to-speechn<1K0 likes17 downloads4mo agoHugging Face21essogbe /mms-tts-fon-accentfine-part-5audio1K<n<10K0 likes16 downloads1y agoHugging Face22yuantuo666 /VB_MMSU-full_3074_hf_format.v0audio1K<n<10K0 likes16 downloads9mo agoHugging Face23hagos1 /tigrinya-mms-ttsaudio1K<n<10K0 likes15 downloads1mo agoHugging Face24sulabhkatiyar /ne-tts-mms-grt NE-TTS MMS-VITS Garo (grt) MMS-VITS fine-tuning subset for Garo (grt). Contains 150 high-quality clips selected from the cleaned NE-TTS dataset (SNR >= 20dB), formatted for MMS-VITS fine-tuning. Stats Metric Value Clips 150 Sample rate 22050Hz SNR filter SNR >= 20dB Source ne-tts-grt Schema Column Type Description audio Audio 22050Hz WAV audio text string Cleaned transcript Usage Use with… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-tts-mms-grt.audiotext-to-speechn<1K0 likes14 downloads4mo agoHugging Face25asadullah797 /mms_synthetic_audioaudion<1K0 likes13 downloads7mo agoHugging Face26zzk123 /mms-tts-uig-script_latin-UQSpeech2audion<1K0 likes11 downloads2y agoHugging Face27zzk123 /mms-tts-uig-script_latin-UQSpeech4audio1K<n<10K0 likes11 downloads2y agoHugging Face28Trelis /mms_synthetic_audioaudion<1K0 likes11 downloads7mo agoHugging Face29zzk123 /mms-tts-uig-script_latin-UQSpeech3audio1K<n<10K0 likes10 downloads2y agoHugging Face30dvd1503 /infore1-mms-vitsaudio10K<n<100K0 likes10 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.