CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01artur-muratov /multilingual-speech-commands-15lang Multilingual Speech Commands Dataset (15 Languages, Augmented) This dataset contains augmented speech command samples in 15 languages, derived from multiple public datasets. Only commands that overlap with the Google Speech Commands (GSC) vocabulary are included, making the dataset suitable for multilingual keyword spotting tasks aligned with GSC-style classification. Audio samples have been augmented using standard audio techniques to improve model robustness (e.g., time-shifting… See the full description on the dataset page: https://huggingface.co/datasets/artur-muratov/multilingual-speech-commands-15lang.audio1M<n<10M16 likes103k downloads1y agoHugging Face02facebook /multilingual_librispeech Dataset Card for MultiLingual LibriSpeech Dataset Summary This is a streamable version of the Multilingual LibriSpeech (MLS) dataset. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/multilingual_librispeech.audioautomatic-speech-recognition1M<n<10M190 likes36k downloads2y agoHugging Face03takuM23 /multilingual_audio_alignments Multilingual MFA-Aligned Speech Dataset (UNDER DEVELOPMENT) A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA). Dataset Description This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and… See the full description on the dataset page: https://huggingface.co/datasets/takuM23/multilingual_audio_alignments.audioautomatic-speech-recognition10M<n<100M4 likes6.1k downloads6mo agoHugging Face04artur-muratov /multilingual-speech-commands-3lang-raw Multilingual Speech Commands Dataset (3 Languages, Raw) This dataset is a curated subset of previously published speech command datasets in Kazakh, Tatar, and Russian. It is intended for use in multilingual speech command recognition and keyword spotting tasks. No data augmentation has been applied. All files are included in their original form as released in the cited works below. This repository simply reorganizes them for convenience and accessibility. Languages… See the full description on the dataset page: https://huggingface.co/datasets/artur-muratov/multilingual-speech-commands-3lang-raw.audio1K<n<10K1 likes3.9k downloads1y agoHugging Face05AAdonis /multilingual_audio_alignments Multilingual MFA-Aligned Speech Dataset A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA). Dataset Description This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and phonemes.… See the full description on the dataset page: https://huggingface.co/datasets/AAdonis/multilingual_audio_alignments.audioautomatic-speech-recognition10M<n<100M27 likes3.4k downloads5mo agoHugging Face06kwatcharasupat /dnr-v3-multilingualaudio0 likes2.8k downloads11mo agoHugging Face07multilingual-tts /open-bible OpenBibleTTS OpenBibleTTS is a large-scale, multilingual speech corpus for low-resource text-to-speech (TTS), spanning 37 underrepresented languages across five regions. It contains ~3,469 hours of aligned, verse-level read speech and 1,121,956 utterances, derived from the Open Bible platform and released under a permissive license. Alignment pipeline: https://github.com/davidguzmanr/open-bible-resources Source: Open Bible (CC BY-SA) Languages Africa (19), South… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-tts/open-bible.audiotext-to-speech1M<n<10M1 likes2.7k downloads3mo agoHugging Face08BrunoHays /multilingual-TEDX-frThe french subset of the dataset Multilingual TEDx. The data uploaded to HF corresponds to the directory fr-fr. The audio files are automatically resampled to 16 kHz. Configs: single_samples (default): all samples taken separately Sample {'file': '0u7tTptBo9I-0', 'audio': {'path': None, 'array': array([ 3.05175781e-05, 6.10351562e-05, 9.15527344e-05, ..., -2.44140625e-04, -3.35693359e-04, -2.74658203e-04]), 'sampling_rate': 16000}, 'sentence': "Bonsoir ! Notre… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/multilingual-TEDX-fr.audioautomatic-speech-recognition100K<n<1M0 likes1.4k downloads10mo agoHugging Face09mteb /sib-fleurs-multilingual-miniaudio10K<n<100K0 likes1.1k downloads1y agoHugging Face10MiniMaxAI /TTS-Multilingual-Test-Set Overview To assess the multilingual zero-shot voice cloning capabilities of TTS models, we have constructed a test set encompassing 24 languages. This dataset provides both audio samples for voice cloning and corresponding test texts. Specifically, the test set for each language includes: 100 distinct test sentences. Audio samples from two speakers (one male and one female) carefully selected from the Mozilla Common Voice (MCV) dataset, intended for voice cloning. Researchers can… See the full description on the dataset page: https://huggingface.co/datasets/MiniMaxAI/TTS-Multilingual-Test-Set.audiotext-to-speechn<1K44 likes1.1k downloads1y agoHugging Face11hf-audio /open-asr-leaderboard-multilingual-datasets ASR Leaderboard Datasets This repository contains test splits from multiple speech corpora, including FLEURS, Common Voice (MCV), and Multilingual LibriSpeech (MLS). How to Load To load a specific subset, use load_dataset with the corresponding config_name in the format <set>_<lang>. from datasets import load_dataset # Load the FLEURS dataset for Bulgarian fleurs_bg = load_dataset("nithinraok/asr-leaderboard-datasets", "fleurs_bg") print(fleurs_bg) # Load the… See the full description on the dataset page: https://huggingface.co/datasets/hf-audio/open-asr-leaderboard-multilingual-datasets.audioautomatic-speech-recognition100K<n<1M4 likes1.1k downloads2mo agoHugging Face12davidguzmanr /CSS10-Multilingual-LJSpeech CSS10-Multilingual-LJSpeech Multilingual speech dataset combining LJSpeech (English) + CSS10 (10 languages) in a consistent LJSpeech format. Dataset Description This dataset merges: LJSpeech: High-quality English speech dataset CSS10: A collection of single-speaker speech datasets for 10 languages All audio files are provided in a consistent format suitable for TTS training. Features Each sample contains: audio: Waveform audio sampled at 22,050 Hz text:… See the full description on the dataset page: https://huggingface.co/datasets/davidguzmanr/CSS10-Multilingual-LJSpeech.audio10K<n<100K0 likes1k downloads6mo agoHugging Face13mteb /minds14-multilingualaudio1K<n<10K0 likes800 downloads1y agoHugging Face14BrunoHays /mixed_multilingual_commonvoice_all_languages_100kBuild from mozilla commonvoice 13 using the script commited in this repo. Used to teach a model to ignore languages that are not french audio100K<n<1M0 likes672 downloads2y agoHugging Face15jml2026 /multilingual-accent-speech 🎙️ Silencio Network: Voice AI Sample Dataset 📊 This is a sample. The full Silencio corpus contains 100,000+ hours across 170+ countries and 100+ languages. 📧 Contact: sofia@silencioai.com for custom datasets, bulk licensing, or specific language requests. 🌍 Why Silencio Data? Silencio data is collected in the wild from a massive, opt-in community (2M+ contributors across 180+ countries), giving you: ✅ Real-world accents, dialects, devices, and… See the full description on the dataset page: https://huggingface.co/datasets/jml2026/multilingual-accent-speech.audioautomatic-speech-recognition1K<n<10K1 likes566 downloads6mo agoHugging Face16flozi00 /multilingual-librispeech-german-labeledaudio100K<n<1M1 likes527 downloads2y agoHugging Face17orbitalsai /multilingual-data-v2audio1M<n<10M0 likes486 downloads2y agoHugging Face18akuzdeuov /qwen3-tts-multilingual-emotional-speechaudio1M<n<10M0 likes482 downloads10d agoHugging Face19Williamsanderson /MedQA-Darija-MultiLingual MedQA-Darija-MultiLingual The largest open trilingual medical Q&A dataset with directly-playable speech audio for English, French, and Moroccan Darija. A research dataset for the BRAIN HEALTH initiative, designed for multilingual medical NLP, low-resource speech recognition, healthcare chatbots, and clinical education tools targeting Morocco and the broader Maghreb region. Dataset is currently in scientific validation phase. After programmatic validation (Stage 1 LOF outlier… See the full description on the dataset page: https://huggingface.co/datasets/Williamsanderson/MedQA-Darija-MultiLingual.audioquestion-answering100K<n<1M4 likes453 downloads5mo agoHugging Face20dsfsi-anv /multilingual-nchlt-dataset NCHLT Auxiliary Speech Corpus - Combined Multilingual Dataset Dataset Description This is a combined multilingual version of the NCHLT Auxiliary Speech Corpus, compiled by the Data Science for Social Impact (DSFSI) research group at the University of Pretoria to facilitate easier benchmarking and multi-language speech recognition research. The original auxiliary data was collected during the National Centre for Human Language Technology (NCHLT) project for the 11 official… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi-anv/multilingual-nchlt-dataset.audioautomatic-speech-recognition100K<n<1M1 likes447 downloads9mo agoHugging Face21sulabhkatiyar /ne-tts-coqui-multilingual NE-TTS Coqui Multilingual Multilingual TTS dataset for 15 North East Indian languages, formatted for Coqui-AI VITS multilingual training. Contains 61,943 clips / 83.5 hours at 22050Hz (SNR >= 20dB only). Languages ISO Language Clips Hours grt Garo 24,772 29.6 ccp Chakma 10,689 14.3 nag Nagamese 9,688 14.5 lus Mizo 8,554 14.5 nnp Wancho 5,081 6.3 trp Kokborok 1,237 1.7 clk Idu Mishmi 602 0.7 mjw Karbi 373 0.4 nre Rengma 258 0.4 nri… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-tts-coqui-multilingual.audiotext-to-speech10K<n<100K1 likes382 downloads26d agoHugging Face22grushaaaaa /indic-multilingual-asr Indic Multilingual ASR Dataset A multilingual ASR dataset covering 13 major Indian languages with 1.1M+ samples. Usage from datasets import load_dataset ds = load_dataset("grushaaaaa/indic-multilingual-asr", split="train") Features audio: 16kHz WAV audio sentence: Transcription text language: Language name source: Source dataset audio1M<n<10M1 likes359 downloads7mo agoHugging Face23NathanRoll /SpeeechEmotion_multilingualaudio10K<n<100K1 likes302 downloads1y agoHugging Face24instinct-org /multilingual-tts-voice-datasetgated Multilingual TTS Voice Dataset Multilingual speech and structured voice-control data for text-to-speech research and training. The collection covers English, Kazakh, Kyrgyz, Russian, Tajik, Turkish, Turkmen, Uzbek, and mixed-language speech. Access requests are reviewed manually. Configurations audio: utterances with embedded audio. prompt_specs: structured text and delivery specifications. clone_pairs: same-speaker reference and target pairs. voice_profiles:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/multilingual-tts-voice-dataset.audiotext-to-speech100K<n<1M0 likes289 downloads1mo agoHugging Face25nineninesix /multilingual-tts-benchmark Multilingual Speech Benchmark for Zero-Shot TTS A voice-cloning and intelligibility benchmark for six language variants, built from Common Voice 17.0 by coverage-driven selection rather than random sampling. Every example pairs a reference clip of one speaker with a target text that speaker never read, so a system is asked to clone a voice and produce new speech, which is what zero-shot TTS is actually for. Pipeline source code:… See the full description on the dataset page: https://huggingface.co/datasets/nineninesix/multilingual-tts-benchmark.audiotext-to-speech100K<n<1M0 likes282 downloads2mo agoHugging Face26artur-muratov /multilingual-speech-commands-15lang-zip Multilingual Speech Commands Dataset (15 Languages, Augmented) This dataset contains augmented speech command samples in 15 languages, derived from multiple public datasets. Only commands that overlap with the Google Speech Commands (GSC) vocabulary are included, making the dataset suitable for multilingual keyword spotting tasks aligned with GSC-style classification. Audio samples have been augmented using standard audio techniques to improve model robustness (e.g., time-shifting… See the full description on the dataset page: https://huggingface.co/datasets/artur-muratov/multilingual-speech-commands-15lang-zip.audio1M<n<10M1 likes269 downloads1y agoHugging Face27jacobrrak /ace-step-multilingual ACE-Step Multilingual Prompt Language Dataset Dataset accompanying: "Does Prompt Language Affect AI-Generated Music? A Multilingual Evaluation of ACE-Step 1.5 Turbo" Dataset description 250 instrumental tracks generated using ACE-Step 1.5 Turbo. 5 prompt languages 10 musical prompt families 5 matched random seeds 30 seconds per track 48 kHz identical generation settings across language conditions Languages English Swedish Spanish German French… See the full description on the dataset page: https://huggingface.co/datasets/jacobrrak/ace-step-multilingual.audiotext-to-audion<1K1 likes258 downloads23d agoHugging Face28Scicom-intl /Evaluation-Multilingual-VC Evaluation-Multilingual-VC We use dataset https://huggingface.co/datasets/sarulab-speech/commonvoice22_sidon, Filter languages that support by Whisper Large V3 to evaluate WER automatically, Only take test set, sort by up votes. Because VC required to source text, source audio, target text, we make sure the target text is not same as source text, target text we take from other rows. Only build first 500 rows for each language Github issue at… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Evaluation-Multilingual-VC.audio10K<n<100K0 likes255 downloads6mo agoHugging Face29michsethowusu /melo-tts-ghana-multilingual Melo-TTS multilingual Ghana training dataset Ghanaian multilingual TTS training data for a MeloTTS fine-tune, phonemised with ghananlpcommunity/xlsr-twi-codeswitch-ipa. Languages twi-asantee — Asante Twi (AfriSpeech open-bible-speech-african) twi-akuapem — Akuapem Twi (AfriSpeech open-bible-speech-african) ewe — Ewe (AfriSpeech open-bible-speech-african) twi-en-codeswitch — Twi-English code-switched speech (ghananlpcommunity) ghana-english —… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/melo-tts-ghana-multilingual.audio100K<n<1M0 likes253 downloads2mo agoHugging Face30MohamedRashad /multilingual-tts Before Anything and Everything ⚱ In the time of writing this Dataset Card, 17,490 18,412 civilian has been killed in Palestine (7,870 8,000 are children and 6,121 6,200 are women). Seek any non-profit organization to help them with what you can (For myself, I use Mersal) 🇵🇸 Dataset Description The Multilingual TTS dataset is an exceptional compilation of text-to-speech (TTS) samples, meticulously crafted to showcase the richness and diversity of human languages.… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/multilingual-tts.audiotext-to-speech10K<n<100K50 likes247 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.