CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01facebook /multilingual_librispeech Dataset Card for MultiLingual LibriSpeech Dataset Summary This is a streamable version of the Multilingual LibriSpeech (MLS) dataset. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/multilingual_librispeech.audioautomatic-speech-recognition1M<n<10M190 likes36k downloads2y agoHugging Face02takuM23 /multilingual_audio_alignments Multilingual MFA-Aligned Speech Dataset (UNDER DEVELOPMENT) A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA). Dataset Description This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and… See the full description on the dataset page: https://huggingface.co/datasets/takuM23/multilingual_audio_alignments.audioautomatic-speech-recognition10M<n<100M4 likes6.1k downloads6mo agoHugging Face03AAdonis /multilingual_audio_alignments Multilingual MFA-Aligned Speech Dataset A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA). Dataset Description This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and phonemes.… See the full description on the dataset page: https://huggingface.co/datasets/AAdonis/multilingual_audio_alignments.audioautomatic-speech-recognition10M<n<100M27 likes3.4k downloads5mo agoHugging Face04multilingual-tts /open-bible OpenBibleTTS OpenBibleTTS is a large-scale, multilingual speech corpus for low-resource text-to-speech (TTS), spanning 37 underrepresented languages across five regions. It contains ~3,469 hours of aligned, verse-level read speech and 1,121,956 utterances, derived from the Open Bible platform and released under a permissive license. Alignment pipeline: https://github.com/davidguzmanr/open-bible-resources Source: Open Bible (CC BY-SA) Languages Africa (19), South… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-tts/open-bible.audiotext-to-speech1M<n<10M1 likes2.7k downloads3mo agoHugging Face05BrunoHays /multilingual-TEDX-frThe french subset of the dataset Multilingual TEDx. The data uploaded to HF corresponds to the directory fr-fr. The audio files are automatically resampled to 16 kHz. Configs: single_samples (default): all samples taken separately Sample {'file': '0u7tTptBo9I-0', 'audio': {'path': None, 'array': array([ 3.05175781e-05, 6.10351562e-05, 9.15527344e-05, ..., -2.44140625e-04, -3.35693359e-04, -2.74658203e-04]), 'sampling_rate': 16000}, 'sentence': "Bonsoir ! Notre… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/multilingual-TEDX-fr.audioautomatic-speech-recognition100K<n<1M0 likes1.4k downloads10mo agoHugging Face06hf-audio /open-asr-leaderboard-multilingual-datasets ASR Leaderboard Datasets This repository contains test splits from multiple speech corpora, including FLEURS, Common Voice (MCV), and Multilingual LibriSpeech (MLS). How to Load To load a specific subset, use load_dataset with the corresponding config_name in the format <set>_<lang>. from datasets import load_dataset # Load the FLEURS dataset for Bulgarian fleurs_bg = load_dataset("nithinraok/asr-leaderboard-datasets", "fleurs_bg") print(fleurs_bg) # Load the… See the full description on the dataset page: https://huggingface.co/datasets/hf-audio/open-asr-leaderboard-multilingual-datasets.audioautomatic-speech-recognition100K<n<1M4 likes1.1k downloads2mo agoHugging Face07malaysia-ai /Multilingual-TTS-language Multilingual-TTS-language malaysia-ai/Multilingual-TTS with two extra columns: column description audio_filename, text, speaker unchanged from malaysia-ai/Multilingual-TTS language language detected from the text (transcript) column of every row post-normalized text after rule-based punctuation / capitalization normalization All original columns and the file/folder layout are preserved: 1552 subsets / 1561 parquet files, 121,842,371 rows (August 2026). Every… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/Multilingual-TTS-language.text-to-speech100M<n<1B0 likes1k downloads28d agoHugging Face08jml2026 /multilingual-accent-speech 🎙️ Silencio Network: Voice AI Sample Dataset 📊 This is a sample. The full Silencio corpus contains 100,000+ hours across 170+ countries and 100+ languages. 📧 Contact: sofia@silencioai.com for custom datasets, bulk licensing, or specific language requests. 🌍 Why Silencio Data? Silencio data is collected in the wild from a massive, opt-in community (2M+ contributors across 180+ countries), giving you: ✅ Real-world accents, dialects, devices, and… See the full description on the dataset page: https://huggingface.co/datasets/jml2026/multilingual-accent-speech.audioautomatic-speech-recognition1K<n<10K1 likes566 downloads6mo agoHugging Face09Williamsanderson /MedQA-Darija-MultiLingual MedQA-Darija-MultiLingual The largest open trilingual medical Q&A dataset with directly-playable speech audio for English, French, and Moroccan Darija. A research dataset for the BRAIN HEALTH initiative, designed for multilingual medical NLP, low-resource speech recognition, healthcare chatbots, and clinical education tools targeting Morocco and the broader Maghreb region. Dataset is currently in scientific validation phase. After programmatic validation (Stage 1 LOF outlier… See the full description on the dataset page: https://huggingface.co/datasets/Williamsanderson/MedQA-Darija-MultiLingual.audioquestion-answering100K<n<1M4 likes453 downloads5mo agoHugging Face10dsfsi-anv /multilingual-nchlt-dataset NCHLT Auxiliary Speech Corpus - Combined Multilingual Dataset Dataset Description This is a combined multilingual version of the NCHLT Auxiliary Speech Corpus, compiled by the Data Science for Social Impact (DSFSI) research group at the University of Pretoria to facilitate easier benchmarking and multi-language speech recognition research. The original auxiliary data was collected during the National Centre for Human Language Technology (NCHLT) project for the 11 official… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi-anv/multilingual-nchlt-dataset.audioautomatic-speech-recognition100K<n<1M1 likes447 downloads9mo agoHugging Face11nineninesix /multilingual-tts-benchmark Multilingual Speech Benchmark for Zero-Shot TTS A voice-cloning and intelligibility benchmark for six language variants, built from Common Voice 17.0 by coverage-driven selection rather than random sampling. Every example pairs a reference clip of one speaker with a target text that speaker never read, so a system is asked to clone a voice and produce new speech, which is what zero-shot TTS is actually for. Pipeline source code:… See the full description on the dataset page: https://huggingface.co/datasets/nineninesix/multilingual-tts-benchmark.audiotext-to-speech100K<n<1M0 likes282 downloads2mo agoHugging Face12legacy-datasets /multilingual_librispeechMultilingual LibriSpeech (MLS) dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish.automatic-speech-recognition100K<n<1M17 likes181 downloads2y agoHugging Face13buumba641 /Zambia-MultiLingual-ASR-Dataset 🇿🇲 Zambia Multilingual ASR Dataset A continuously growing and curated multilingual speech corpus for Zambian languages, designed to advance Automatic Speech Recognition (ASR) research through community-driven data collection and real-world evaluation. Overview The Zambia Multilingual ASR Dataset is an open, continuously evolving speech corpus developed as part of the ZamVoice project. The dataset supports research and development of Automatic Speech… See the full description on the dataset page: https://huggingface.co/datasets/buumba641/Zambia-MultiLingual-ASR-Dataset.audioautomatic-speech-recognitionn<1K2 likes180 downloads8d agoHugging Face14Cnam-LMSSC /multilingual_librispeech_french_phoneme Multilingual LibriSpeech French Phoneme Dataset Summary This dataset is a curated version of the French subset of Multilingual LibriSpeech (MLS), enriched with a phonetic transcription column (phoneme). The Laboratoire de Mécanique des Structures et des Systèmes Couplés (Cnam-LMSSC) created this version to facilitate research into French acoustic modeling, phoneme recognition, and speech synthesis. It builds upon the high-quality audio derived from LibriVox audiobooks… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/multilingual_librispeech_french_phoneme.audioautomatic-speech-recognition100K<n<1M1 likes139 downloads8mo agoHugging Face15noty7gian /synthetic-multilingual-speaker-diarization Synthetic Multilingual Speaker Diarization Dataset This dataset contains synthetic multilingual speaker diarization data with Hindi, English, and Punjabi audio samples. Dataset Structure ├── audio/ # WAV audio files (16kHz) - 3417 files ├── all_samples_combined.csv # Complete dataset annotations (with silence) └── all_visible_combined.csv # Visible dataset annotations (without silence) Statistics Total samples: 3417 audio… See the full description on the dataset page: https://huggingface.co/datasets/noty7gian/synthetic-multilingual-speaker-diarization.audioautomatic-speech-recognition1K<n<10K0 likes122 downloads1y agoHugging Face16Reubencf /multilingual-synthetic-tts Multilingual Synthetic TTS Dataset 🏆 Submitted to the Uncharted Data Challenge hosted by Adaption Labs — credit to Adaptive Data by Adaption for organizing the hackathon. A large-scale synthetic multilingual speech dataset — 68,677 clips across 9 languages, generated with Qwen3-TTS-12Hz-1.7B-Base using zero-shot voice cloning from 5 reference speakers. Intended for training and evaluating TTS, ASR, voice conversion, and multilingual speech models. Each clip is paired with the… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/multilingual-synthetic-tts.audiotext-to-speech10K<n<100K2 likes121 downloads5mo agoHugging Face17issai /Multilingual_Speech_Dataset Multilingual Speech Dataset Paper: A Study of Multilingual End-to-End Speech Recognition for Kazakh, Russian, and English Repository: https://github.com/IS2AI/MultilingualASR Description: This repository provides the dataset used in the paper "A Study of Multilingual End-to-End Speech Recognition for Kazakh, Russian, and English". The paper focuses on training a single end-to-end (E2E) ASR model for Kazakh, Russian, and English, comparing monolingual and multilingual approaches… See the full description on the dataset page: https://huggingface.co/datasets/issai/Multilingual_Speech_Dataset.audioautomatic-speech-recognition100K<n<1M3 likes111 downloads2y agoHugging Face18Cnam-LMSSC /multilingual_librispeech_spanish_phoneme Multilingual LibriSpeech Spanish Phoneme Dataset Summary This dataset is a curated version of the Spanish subset of Multilingual LibriSpeech (MLS), enriched with a phonetic transcription column (phoneme). The Laboratoire de Mécanique des Structures et des Systèmes Couplés (Cnam-LMSSC) created this version to facilitate research into Spanish acoustic modeling, phoneme recognition, and speech synthesis. It builds upon the high-quality audio derived from LibriVox audiobooks… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/multilingual_librispeech_spanish_phoneme.audioautomatic-speech-recognition100K<n<1M1 likes110 downloads7mo agoHugging Face19Reubencf /Adaption-multilingual-speech This dataset is a remastered version of Reubencf/multilingual-synthetic-tts prepared using Adaption's Adaptive Data platform. Multilingual Speech (Adaption) 10,274 audio + text rows selected from the original 68,677-clip multilingual synthetic speech corpus, with Adaption-sharpened enhanced_prompt and enhanced_completion columns. Every row carries the synthesised audio, the ground-truth text, and language/style/voice metadata — ready for speech SFT. Original dataset (for… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/Adaption-multilingual-speech.audiotext-to-speech10K<n<100K0 likes98 downloads5mo agoHugging Face20eQOURSE /multilingual-speech Multilingual Indian Conversational Speech A dataset of naturalistic, spontaneous two-speaker conversations across 13 Indian languages, with segment-level transcripts, speaker profiles, timestamps, and recording metadata. Designed for ASR, TTS, speaker diarization, and conversational speech research. Languages (13) Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Nepali, Odia, Punjabi, Tamil, Telugu, Urdu. Content Conversations… See the full description on the dataset page: https://huggingface.co/datasets/eQOURSE/multilingual-speech.audioautomatic-speech-recognitionn<1K0 likes87 downloads3mo agoHugging Face21SoundWaveET /leyu-ethiopian-multilingual-speech-corpus-2026 🇪🇹 Leyu Ethiopian Multilingual Speech Corpus 2026 Official research-grade multilingual speech dataset submitted for the Leyu Platform Data Collection Competition 2026, organized by gheero (Leyu Platform Team). 🏢 Platform & Team Verification Details Team Name: SoundWaveET Hugging Face Organization: SoundWaveET Dataset Repository: SoundWaveET/leyu-ethiopian-multilingual-speech-corpus-2026 Platform Live Frontend: https://leyusound.netlify.app Platform Live API:… See the full description on the dataset page: https://huggingface.co/datasets/SoundWaveET/leyu-ethiopian-multilingual-speech-corpus-2026.audioautomatic-speech-recognition1K<n<10K0 likes76 downloads1mo agoHugging Face22jml2026 /multilingual-accent-speech-v2 Silencio Network: Multilingual Accent Speech Dataset (Sample) Overview Silencio data is valuable because it's collected in the wild from a massive, opt-in community (1.2M users across 180+ countries), giving buyers real-world accents, dialects, devices, and environments that lab or scraped datasets don't capture. Every recording is tied to explicit, traceable consent and processed with privacy-first pipelines (GDPR/CCPA compliant, anonymized, PII hashed)… See the full description on the dataset page: https://huggingface.co/datasets/jml2026/multilingual-accent-speech-v2.audioautomatic-speech-recognitionn<1K0 likes74 downloads8mo agoHugging Face23antoineedy /multilingual_librispeech_fr_punctuated Multilingual LibriSpeech French (Punctuated) This dataset is a converted version of BrunoHays/multilingual_librispeech_fr_punctuated in the new Hugging Face datasets format (Parquet-based, without loading scripts). Original Dataset The original dataset contains French speech data from Multilingual LibriSpeech with punctuated transcriptions. Changes Converted from old loading script format to new Parquet-based format Maintains all original features and data… See the full description on the dataset page: https://huggingface.co/datasets/antoineedy/multilingual_librispeech_fr_punctuated.audioautomatic-speech-recognition10K<n<100K0 likes56 downloads10mo agoHugging Face24LeyuCompetition /sound-wave-leyu-ethiopian-multilingual-speech-corpus-2026 🇪🇹 Leyu Ethiopian Multilingual Speech Corpus 2026 Official research-grade multilingual speech dataset submitted for the Leyu Platform Data Collection Competition 2026, organized by gheero (Leyu Platform Team). 🏢 Platform & Team Verification Details Team Name: SoundWaveET Hugging Face Organization: SoundWaveET Dataset Repository: SoundWaveET/leyu-ethiopian-multilingual-speech-corpus-2026 Platform Live Frontend: https://leyusound.netlify.app Platform Live API:… See the full description on the dataset page: https://huggingface.co/datasets/LeyuCompetition/sound-wave-leyu-ethiopian-multilingual-speech-corpus-2026.audioautomatic-speech-recognition1K<n<10K0 likes46 downloads20d agoHugging Face25Cnam-LMSSC /multilingual_librispeech_italian_phoneme Multilingual LibriSpeech Italian Phoneme Dataset Summary This dataset is a curated version of the Italian subset of Multilingual LibriSpeech (MLS), enriched with a phonetic transcription column (phoneme). The Laboratoire de Mécanique des Structures et des Systèmes Couplés (Cnam-LMSSC) created this version to facilitate research into Italian acoustic modeling, phoneme recognition, and speech synthesis. It builds upon the high-quality audio derived from LibriVox audiobooks… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/multilingual_librispeech_italian_phoneme.audioautomatic-speech-recognition10K<n<100K1 likes44 downloads7mo agoHugging Face26obaydata /multilingual-tts-corpus Multilingual TTS Corpus A multilingual text-to-speech dataset containing audio recordings with text transcriptions across multiple languages. This is the first batch; more languages will be added over time. Languages (Batch 1) Subset Language Audio Files Annotation Format Source ru-tts/ Russian 10 JSONL (single file) 俄语TTS(文本标注及语音)基石数据集 #294 ru-speech/ Russian 5 JSONL (single file) 俄语高质量语音音频语料库 #481 vi-speech/ Vietnamese 5 JSONL (single file) 越南语高质量语音音频语料库… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/multilingual-tts-corpus.audiotext-to-speechn<1K0 likes40 downloads6mo agoHugging Face27aoiandroid /mms-multilingual-audio-5to30min Multilingual Audio Dataset (5-30min) This dataset contains continuous speech audio files for various languages (ranging from 5 to 30 minutes in length per language) collected from diverse sources including Hugging Face and YouTube. Dataset Statistics Total Languages: 100 Sources: HF Omnilingual ASR Corpus, YouTube Language Details Language Code Language Name Source Duration (seconds) jpn Japanese youtube 1459.84 eng English youtube… See the full description on the dataset page: https://huggingface.co/datasets/aoiandroid/mms-multilingual-audio-5to30min.audioautomatic-speech-recognitionn<1K0 likes35 downloads3mo agoHugging Face28alconost /alconost-multilingual-speech-goldgated Multilingual Speech & Translation Dataset — EN↔JA/AR-EG/PL/RU (10 phrases, dual-take) Description 10 English source phrases with expert human translations into Japanese, Egyptian Arabic (ar-EG), and Polish. Each target phrase is recorded by native speakers (two takes each). Audio files are WAV 48 kHz mono, 16‑bit PCM format. Translations are produced and QA'd by professional linguists; recordings follow consistent orthography/style (AR-EG: Egyptian dialect; JA/PL: standard). All… See the full description on the dataset page: https://huggingface.co/datasets/alconost/alconost-multilingual-speech-gold.audiotranslationn<1K0 likes28 downloads8mo agoHugging Face29Metric-AI /open-asr-leaderboard-multilingual-datasets Open ASR Leaderboard Armenian Test Datasets This private repository holds leaderboard-compatible Armenian test configurations while their integration is being validated. Configurations fleurs_hy Source: google/fleurs, configuration hy_am, test split Reviewed reference changes: Metric-AI/fleurs-corrections, test split 932 recordings; all 314 reviewed corrections were matched to the original source transcript and applied mcv_hy… See the full description on the dataset page: https://huggingface.co/datasets/Metric-AI/open-asr-leaderboard-multilingual-datasets.audioautomatic-speech-recognition1K<n<10K1 likes22 downloads11d agoHugging Face30djelia /multilingual-asrgated multilingual-asr CoVoST 2 and Common Voice 17.0 Swahili and Hausa, re-packaged under one feature schema so the configs can be concatenated into a single multi-task training mix. Four configs, ~206 hours of distinct audio, 8.44 GB of Parquet. No Bambara. Load from datasets import load_dataset asr = load_dataset("djelia/multilingual-asr", "covost2-transcription", split="train") sw_test = load_dataset("djelia/multilingual-asr", "swahili", split="test")… See the full description on the dataset page: https://huggingface.co/datasets/djelia/multilingual-asr.audioautomatic-speech-recognition100K<n<1M0 likes20 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.