CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01facebook /multilingual_librispeech Dataset Card for MultiLingual LibriSpeech Dataset Summary This is a streamable version of the Multilingual LibriSpeech (MLS) dataset. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/multilingual_librispeech.audioautomatic-speech-recognition1M<n<10M191 likes39k downloads2y agoHugging Face02takuM23 /multilingual_audio_alignments Multilingual MFA-Aligned Speech Dataset (UNDER DEVELOPMENT) A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA). Dataset Description This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and… See the full description on the dataset page: https://huggingface.co/datasets/takuM23/multilingual_audio_alignments.audioautomatic-speech-recognition10M<n<100M4 likes5.4k downloads6mo agoHugging Face03AAdonis /multilingual_audio_alignments Multilingual MFA-Aligned Speech Dataset A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA). Dataset Description This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and phonemes.… See the full description on the dataset page: https://huggingface.co/datasets/AAdonis/multilingual_audio_alignments.audioautomatic-speech-recognition10M<n<100M27 likes3.1k downloads5mo agoHugging Face04multilingual-tts /open-bible OpenBibleTTS OpenBibleTTS is a large-scale, multilingual speech corpus for low-resource text-to-speech (TTS), spanning 37 underrepresented languages across five regions. It contains ~3,469 hours of aligned, verse-level read speech and 1,121,956 utterances, derived from the Open Bible platform and released under a permissive license. Alignment pipeline: https://github.com/davidguzmanr/open-bible-resources Source: Open Bible (CC BY-SA) Languages Africa (19), South… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-tts/open-bible.audiotext-to-speech1M<n<10M1 likes2.8k downloads3mo agoHugging Face05BrunoHays /multilingual-TEDX-frThe french subset of the dataset Multilingual TEDx. The data uploaded to HF corresponds to the directory fr-fr. The audio files are automatically resampled to 16 kHz. Configs: single_samples (default): all samples taken separately Sample {'file': '0u7tTptBo9I-0', 'audio': {'path': None, 'array': array([ 3.05175781e-05, 6.10351562e-05, 9.15527344e-05, ..., -2.44140625e-04, -3.35693359e-04, -2.74658203e-04]), 'sampling_rate': 16000}, 'sentence': "Bonsoir ! Notre… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/multilingual-TEDX-fr.audioautomatic-speech-recognition100K<n<1M0 likes1.5k downloads10mo agoHugging Face06hf-audio /open-asr-leaderboard-multilingual-datasets ASR Leaderboard Datasets This repository contains test splits from multiple speech corpora, including FLEURS, Common Voice (MCV), and Multilingual LibriSpeech (MLS). How to Load To load a specific subset, use load_dataset with the corresponding config_name in the format <set>_<lang>. from datasets import load_dataset # Load the FLEURS dataset for Bulgarian fleurs_bg = load_dataset("nithinraok/asr-leaderboard-datasets", "fleurs_bg") print(fleurs_bg) # Load the… See the full description on the dataset page: https://huggingface.co/datasets/hf-audio/open-asr-leaderboard-multilingual-datasets.audioautomatic-speech-recognition100K<n<1M4 likes1.2k downloads2mo agoHugging Face07Williamsanderson /MedQA-Darija-MultiLingual MedQA-Darija-MultiLingual The largest open trilingual medical Q&A dataset with directly-playable speech audio for English, French, and Moroccan Darija. A research dataset for the BRAIN HEALTH initiative, designed for multilingual medical NLP, low-resource speech recognition, healthcare chatbots, and clinical education tools targeting Morocco and the broader Maghreb region. Dataset is currently in scientific validation phase. After programmatic validation (Stage 1 LOF outlier… See the full description on the dataset page: https://huggingface.co/datasets/Williamsanderson/MedQA-Darija-MultiLingual.audioquestion-answering100K<n<1M4 likes507 downloads5mo agoHugging Face08dsfsi-anv /multilingual-nchlt-dataset NCHLT Auxiliary Speech Corpus - Combined Multilingual Dataset Dataset Description This is a combined multilingual version of the NCHLT Auxiliary Speech Corpus, compiled by the Data Science for Social Impact (DSFSI) research group at the University of Pretoria to facilitate easier benchmarking and multi-language speech recognition research. The original auxiliary data was collected during the National Centre for Human Language Technology (NCHLT) project for the 11 official… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi-anv/multilingual-nchlt-dataset.audioautomatic-speech-recognition100K<n<1M1 likes379 downloads9mo agoHugging Face09nineninesix /multilingual-tts-benchmark Multilingual Speech Benchmark for Zero-Shot TTS A voice-cloning and intelligibility benchmark for six language variants, built from Common Voice 17.0 by coverage-driven selection rather than random sampling. Every example pairs a reference clip of one speaker with a target text that speaker never read, so a system is asked to clone a voice and produce new speech, which is what zero-shot TTS is actually for. Pipeline source code:… See the full description on the dataset page: https://huggingface.co/datasets/nineninesix/multilingual-tts-benchmark.audiotext-to-speech100K<n<1M0 likes309 downloads2mo agoHugging Face10Cnam-LMSSC /multilingual_librispeech_french_phoneme Multilingual LibriSpeech French Phoneme Dataset Summary This dataset is a curated version of the French subset of Multilingual LibriSpeech (MLS), enriched with a phonetic transcription column (phoneme). The Laboratoire de Mécanique des Structures et des Systèmes Couplés (Cnam-LMSSC) created this version to facilitate research into French acoustic modeling, phoneme recognition, and speech synthesis. It builds upon the high-quality audio derived from LibriVox audiobooks… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/multilingual_librispeech_french_phoneme.audioautomatic-speech-recognition100K<n<1M1 likes139 downloads9mo agoHugging Face11Cnam-LMSSC /multilingual_librispeech_spanish_phoneme Multilingual LibriSpeech Spanish Phoneme Dataset Summary This dataset is a curated version of the Spanish subset of Multilingual LibriSpeech (MLS), enriched with a phonetic transcription column (phoneme). The Laboratoire de Mécanique des Structures et des Systèmes Couplés (Cnam-LMSSC) created this version to facilitate research into Spanish acoustic modeling, phoneme recognition, and speech synthesis. It builds upon the high-quality audio derived from LibriVox audiobooks… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/multilingual_librispeech_spanish_phoneme.audioautomatic-speech-recognition100K<n<1M1 likes113 downloads7mo agoHugging Face12antoineedy /multilingual_librispeech_fr_punctuated Multilingual LibriSpeech French (Punctuated) This dataset is a converted version of BrunoHays/multilingual_librispeech_fr_punctuated in the new Hugging Face datasets format (Parquet-based, without loading scripts). Original Dataset The original dataset contains French speech data from Multilingual LibriSpeech with punctuated transcriptions. Changes Converted from old loading script format to new Parquet-based format Maintains all original features and data… See the full description on the dataset page: https://huggingface.co/datasets/antoineedy/multilingual_librispeech_fr_punctuated.audioautomatic-speech-recognition10K<n<100K0 likes93 downloads10mo agoHugging Face13Reubencf /multilingual-synthetic-tts Multilingual Synthetic TTS Dataset 🏆 Submitted to the Uncharted Data Challenge hosted by Adaption Labs — credit to Adaptive Data by Adaption for organizing the hackathon. A large-scale synthetic multilingual speech dataset — 68,677 clips across 9 languages, generated with Qwen3-TTS-12Hz-1.7B-Base using zero-shot voice cloning from 5 reference speakers. Intended for training and evaluating TTS, ASR, voice conversion, and multilingual speech models. Each clip is paired with the… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/multilingual-synthetic-tts.audiotext-to-speech10K<n<100K2 likes91 downloads5mo agoHugging Face14Cnam-LMSSC /multilingual_librispeech_italian_phoneme Multilingual LibriSpeech Italian Phoneme Dataset Summary This dataset is a curated version of the Italian subset of Multilingual LibriSpeech (MLS), enriched with a phonetic transcription column (phoneme). The Laboratoire de Mécanique des Structures et des Systèmes Couplés (Cnam-LMSSC) created this version to facilitate research into Italian acoustic modeling, phoneme recognition, and speech synthesis. It builds upon the high-quality audio derived from LibriVox audiobooks… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/multilingual_librispeech_italian_phoneme.audioautomatic-speech-recognition10K<n<100K1 likes46 downloads7mo agoHugging Face15Reubencf /Adaption-multilingual-speech This dataset is a remastered version of Reubencf/multilingual-synthetic-tts prepared using Adaption's Adaptive Data platform. Multilingual Speech (Adaption) 10,274 audio + text rows selected from the original 68,677-clip multilingual synthetic speech corpus, with Adaption-sharpened enhanced_prompt and enhanced_completion columns. Every row carries the synthesised audio, the ground-truth text, and language/style/voice metadata — ready for speech SFT. Original dataset (for… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/Adaption-multilingual-speech.audiotext-to-speech10K<n<100K0 likes42 downloads5mo agoHugging Face16Ugiat /multilingual_librispeech_french_punctuated Multilingual LibriSpeech French, punctuated and capitalized (train) A derivative of the French part of Multilingual LibriSpeech (MLS), the corpus of read audiobooks from LibriVox published by Vineel Pratap, Qiantong Xu, Anuroop Sriram, Gabriel Synnaeve and Ronan Collobert (Facebook AI Research). MLS distributes its transcriptions lowercased and without any punctuation. This dataset keeps that upstream text untouched in the text column and adds a second column, text_punctuated… See the full description on the dataset page: https://huggingface.co/datasets/Ugiat/multilingual_librispeech_french_punctuated.audioautomatic-speech-recognition100K<n<1M1 likes41 downloads2d agoHugging Face17Metric-AI /open-asr-leaderboard-multilingual-datasets Open ASR Leaderboard Armenian Test Datasets This private repository holds leaderboard-compatible Armenian test configurations while their integration is being validated. Configurations fleurs_hy Source: google/fleurs, configuration hy_am, test split Reviewed reference changes: Metric-AI/fleurs-corrections, test split 932 recordings; all 314 reviewed corrections were matched to the original source transcript and applied mcv_hy… See the full description on the dataset page: https://huggingface.co/datasets/Metric-AI/open-asr-leaderboard-multilingual-datasets.audioautomatic-speech-recognition1K<n<10K1 likes32 downloads14d agoHugging Face18djelia /multilingual-asrgated multilingual-asr CoVoST 2 and Common Voice 17.0 Swahili and Hausa, re-packaged under one feature schema so the configs can be concatenated into a single multi-task training mix. Four configs, ~206 hours of distinct audio, 8.44 GB of Parquet. No Bambara. Load from datasets import load_dataset asr = load_dataset("djelia/multilingual-asr", "covost2-transcription", split="train") sw_test = load_dataset("djelia/multilingual-asr", "swahili", split="test")… See the full description on the dataset page: https://huggingface.co/datasets/djelia/multilingual-asr.audioautomatic-speech-recognition100K<n<1M0 likes20 downloads2mo agoHugging Face19piyazon /multilingual_voice_ug_en_zhgatedtrain: 'Uyghur': 266672 'Chinese': 50215 'English': 32672 validation: 'Uyghur': 2712 'Chinese': 491 'English': 328 audioautomatic-speech-recognition100K<n<1M1 likes3 downloads8mo agoHugging Face20steven0226 /open-asr-leaderboard-multilingual-datasets Open ASR Leaderboard Chinese Test Set (FLEURS) This repository holds a test-only copy of the Mandarin Chinese test split of google/fleurs (config cmn_hans_cn, split test). It is used for the Chinese column of the Open ASR Leaderboard (huggingface/open_asr_leaderboard#147). The layout matches the FLEURS configs in hf-audio/open-asr-leaderboard-multilingual-datasets, so this config may later be merged into that repository. How to load from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/steven0226/open-asr-leaderboard-multilingual-datasets.audioautomatic-speech-recognitionn<1K0 likes13h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.