CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MohammadJRanjbar /ParsVoicegated ParsVoice A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis 📣 Accepted to the EMNLP 2026 Main Conference. ParsVoice is the largest publicly available Persian speech–text corpus tailored for training multi-speaker text-to-speech (TTS) systems. It is built from long-form Persian audiobook recordings using a fully automated pipeline combining sentence-aware segmentation, ASR transcription, a ParsBERT sentence-completion classifier, binary-search… See the full description on the dataset page: https://huggingface.co/datasets/MohammadJRanjbar/ParsVoice.audiotext-to-speech1M<n<10M32 likes2.4k downloads18d agoHugging Face02espnet /Bagpiper_PreTrain_Data Bagpiper Pretraining Data Bagpiper Pretraining Data is the public rich-captioned audio snapshot associated with Bagpiper, an open-ended audio language model that learns bidirectional mappings between audio and comprehensive text descriptions across speech, music, environmental sound, and mixtures. The en metadata describes the primary rich-caption language. Source audio can contain speech or singing in other languages; it is not an English-only audio guarantee. The repository… See the full description on the dataset page: https://huggingface.co/datasets/espnet/Bagpiper_PreTrain_Data.tabularautomatic-speech-recognition10K<n<100K0 likes2.4k downloads2mo agoHugging Face03Peacockery /common-voice-scripted-speech-26 Common Voice Scripted Speech A row-normalized multilingual ASR dataset built from Mozilla Data Collective Common Voice Scripted Speech. Each upstream archive is converted to appendable parquet shards under data/<upstream_split>/, one shard per source archive and split, with audio bytes embedded in an audio struct column. Status Manifest languages: 60 Languages uploaded: 18 Columns audio (bytes, path) sentence, locale, language, upstream_split… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/common-voice-scripted-speech-26.tabularautomatic-speech-recognition100K<n<1M0 likes2k downloads3mo agoHugging Face04blanchon /parler-tts_mls_eng_10k_snac_token_old Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/parler-tts_mls_eng_10k_snac_token_old.tabularautomatic-speech-recognition100K<n<1M1 likes998 downloads2y agoHugging Face05Quran-Lab /quran-tajweed-phonetics The complete phonetic layer of the Quran in the riwaya of Hafs 'an 'Asim via tariq al-Shatibiyyah: 6,236 ayat, 522,475 phones, every phone carrying its tajweed attribution: madd class with its transmitted length range, ghunna grade, qalqalah class, tafkheem with its rank, sakt, the seventeen sifat, and the rule that produced it. Built and maintained by Quran Lab, a waqf building open technology in the service of the Quran. How it was built and verified Indexed from the… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/quran-tajweed-phonetics.tabularautomatic-speech-recognition10K<n<100K3 likes555 downloads12d agoHugging Face06parler-tts /mls-eng-speaker-descriptions Dataset Card for Annotations of English MLS This dataset consists in annotations of the English subset of the Multilingual LibriSpeech (MLS) dataset. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours for other languages. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls-eng-speaker-descriptions.tabularautomatic-speech-recognition10M<n<100M13 likes426 downloads2y agoHugging Face07psdn-ai /bangla-10kgated Bangla-10K: A Challenging, Metadata-Rich Corpus of Read and Conversational Bengali Speech from India and Bangladesh Bangla-10K is a 10,816-hour Bengali speech corpus with 624,951 recordings from India and Bangladesh: a 10,070.8-hour core corpus (567,323 recordings) and a separately collected 745.1-hour evaluation set (57,628 recordings). It combines scripted single-speaker read speech with natural multi-speaker conversations for Bengali automatic speech recognition (ASR). The… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/bangla-10k.audioautomatic-speech-recognition100K<n<1M0 likes324 downloads6m agoHugging Face08gavinlaw /chinese-lips-speech-slide-probe Chinese-LiPS Speech + Slide Probe A self-contained probe set for testing whether visual slide context helps simultaneous speech translation — with the input as audio, not transcripts. Why audio matters: feeding a transcript to a text LLM deletes the acoustic ambiguity (homophones, polysemy) that slide context is meant to resolve; the transcript already commits to one reading. Any honest test of "does vision help streaming ST" must consume speech. Contents… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/chinese-lips-speech-slide-probe.audiotranslationn<1K0 likes249 downloads2mo agoHugging Face09paodigitalhub /pao-audio-dataset 🎙️ Pa'O Audio Dataset ပအိုဝ်ႏ အငေါဝ်း အဆင်ႏဗာႏ ရွမ်ခြွဉ်းဗူႏ 📌 Project Summary The Pa'O Audio Dataset is an open-source initiative created to facilitate the development of speech technologies and Artificial Intelligence tools for the Pa'O language (ပအိုဝ်ႏဘာႏသာႏငေါဝ်းငွါ). Pa'O is primarily spoken in Shan State and other regions of Myanmar. As a low-resource language in the AI landscape, this dataset provides audio recordings and corresponding… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/pao-audio-dataset.audioautomatic-speech-recognitionn<1K1 likes230 downloads19h agoHugging Face10vnmoorthy /pavo-bench PAVO-Bench: 50K-Turn Benchmark for ASR-LLM-TTS Pipeline Routing Code: github.com/vnmoorthy/pavo-bench · Paper: TMLR 2026 (accepted) · Authors: NarasingaMoorthy VeiluKanthaPerumal (UPenn), Mohammed Imthathullah (Google) pip install git+https://github.com/vnmoorthy/pavo-bench.git Headline results (vs fixed-cloud baseline, 50,000 voice turns) Metric Result Significance P95 end-to-end latency (H100, LibriSpeech) −10.3% (−167 ms) — Median latency −34%… See the full description on the dataset page: https://huggingface.co/datasets/vnmoorthy/pavo-bench.documentautomatic-speech-recognition10K<n<100K0 likes226 downloads1mo agoHugging Face11parler-tts /mls-eng-10k-tags_tagged_10k_generated Dataset Card for Annotations of 10K hours of English MLS This dataset consists in annotations of a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls-eng-10k-tags_tagged_10k_generated.tabularautomatic-speech-recognition1M<n<10M17 likes138 downloads2y agoHugging Face12Wi-Fi /korean-full-duplex-synthetic-dataset-preview Korean Full-Duplex Synthetic Dataset Preview Overview Public preview of a Korean full-duplex synthetic speech dataset. This repository contains 100 conversations sampled from a corpus of 89,273 conversations (2,000.5 hours); it does not publish the full corpus audio. Preview contents 100 conversation WAV files data/representative.jsonl 24 kHz, mono, 16-bit PCM Events: normal, barge_in, backchannel, cutoff_by_user Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.audioautomatic-speech-recognitionn<1K1 likes138 downloads1mo agoHugging Face13ivrit-ai /knesset-plenumsgated About This dataset is derived from raw a/v recordings and human-generated protocols of the Knesset (the Israeli house of representatives) plenums as part of the ivrit.ai project. Consider visiting the preview space for this dataset here Method Data dumps from the Knesset contain A/V recordings, alongside proprietary protocols with timestamps. We extract the audio stream, and clean up timestamp mistakes (such as backward jumps, or out-of-order timestamp artifacts). The… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-plenums.audioautomatic-speech-recognition1K<n<10K3 likes108 downloads10mo agoHugging Face14kalpesh77 /marathi-phonology-matrices मराठी व्याकरण आणि ध्वनी मॅट्रिक्स Marathi Phonology Matrices गणितीय ध्वनी संश्लेषणासाठी (Mathematical Speech Synthesis) तयार केलेला सर्वसमावेशक मराठी फोनोलॉजी डेटासेट. 🎯 उद्देश्य हा डेटासेट मराठी भाषेच्या: फोनोलॉजिकल विश्लेषण मॉर्फोलॉजी (लिंग, वचन, काळ) संधि व श्व नियम युक्तक्षर (Clusters) Duration & Pitch नियम Loanword adaptation या सर्वांसाठी संरचित डेटा पुरवतो. TTS, ASR, G2P आणि Computational Linguistics संशोधनासाठी उपयुक्त. 📊… See the full description on the dataset page: https://huggingface.co/datasets/kalpesh77/marathi-phonology-matrices.tabulartext-to-speech1K<n<10K1 likes97 downloads2mo agoHugging Face15s512757 /polish-tedx-asr-eval Polish-TEDx-ASR-Eval A dataset for evaluating automatic speech recognition (ASR) systems for Polish in the domain of TEDx public talks. Contains audio segments from Polish TEDx talks available on YouTube (CC BY-NC-ND 4.0) and synthetic speech generated with KugelAudio (MIT), with manually created and cross-verified transcriptions. Created as part of the course "Workshops on Evaluation of Speech Recognition Systems" (ZWESUI, AMU 2026) by Group 1. Statistics… See the full description on the dataset page: https://huggingface.co/datasets/s512757/polish-tedx-asr-eval.audioautomatic-speech-recognitionn<1K0 likes90 downloads3mo agoHugging Face16DewiBrynJones /preprocessed-whisper-btb-cv-cvad-wlga-ca-2607 Dataset Card Preprocessed Dataset: DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2607 Revision: main Dataset Statistics Train Split Statistics Dataset Revision Split Duration (HH:MM:SS) Clips Words Words/Clip % DewiBrynJones/banc-trawsgrifiadau-bangor-2605 5bfe2d098c8486d97fac8be76d86ec9146435245 train 56:46:32 50,557 589,095 11.7 31.9 techiaith/corpws-clllc-wlga 5d00294c31c78b1d7937bb2c2bc6cc70bc18d410 clips 48:20:49 27,579… See the full description on the dataset page: https://huggingface.co/datasets/DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2607.tabularautomatic-speech-recognition100K<n<1M0 likes89 downloads1mo agoHugging Face17turnipseason /paralingua_ru Russian Paralinguistic Annotation Dataset Датасет паралингвистической разметки спикеров из трёх русскоязычных корпусов: biggest_ru_book, DeepSpeech и Golos. Что размечалось Каждое аудио размечалось вручную по следующим характеристикам: Поле Описание Пример значений gender Пол спикера мужской, женский age_group Возрастная группа молодой, взрослый, пожилой voice_pitch Высота голоса низкий, средний, высокий loudness Громкость тихий, нормальный… See the full description on the dataset page: https://huggingface.co/datasets/turnipseason/paralingua_ru.tabulartext-to-speech100K<n<1M8 likes84 downloads4mo agoHugging Face18ia-espirita /pinga-fogo-chico-xavier 🎙️ Pinga-Fogo com Chico Xavier — TV Tupi, 1971 As duas entrevistas históricas do médium Chico Xavier, transmitidas ao vivo pela TV Tupi em 1971, transcritas e estruturadas em turnos de fala com timestamp. 345 turnos (115 deles respostas do próprio Chico Xavier), a partir de 6 horas de áudio — o registro mais extenso do médium falando de improviso, sem edição, diante de um painel de jornalistas. Arquivos Arquivo Programa Turnos Respostas do Chico… See the full description on the dataset page: https://huggingface.co/datasets/ia-espirita/pinga-fogo-chico-xavier.tabularquestion-answeringn<1K1 likes72 downloads1mo agoHugging Face19Reza2kn /persian-asr-text-2.69M-deduped 🗂️ persian-asr-text-2.69M-deduped English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission 🌟 At a glance | معرفی سریع English فارسی 🎯 Purpose Deduplicated Persian ASR text dataset used by the training stack. پیکرهٔ متنی فارسیِ حذف‌تکرارشده برای ساخت واژگان، مدل‌سازی زبانی و پشتیبانی از آموزش ASR. 🧩 Role Persian text and linguistic asset مصنوع متنی و زبانی فارسی 📦 Snapshot 4 files; approximately 109.64 MB 4 فایل؛ حدود 109.64… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-text-2.69M-deduped.tabularautomatic-speech-recognition1M<n<10M0 likes68 downloads2mo agoHugging Face20Peacockery /tajik-asr-youtube tajik-asr-youtube Every Tajik-language clip from this project's YouTube scrape — 41 channels of news, talk shows, podcasts, audiobooks, and learning content — with machine transcripts and the verification scores left in as columns instead of applied as a filter. Pick your own quality threshold; the training corpus this project actually ships (tajik-asr-corpus-v3) is the gated subset. Layout Parquet shards under data/ with an audio struct column (16 kHz mono FLAC… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/tajik-asr-youtube.tabularautomatic-speech-recognition100K<n<1M0 likes65 downloads3mo agoHugging Face21mesolitica /pseudolabel-malaya-speech-stt-train-whisper-large-v3tabularautomatic-speech-recognition1M<n<10M1 likes63 downloads3y agoHugging Face22pharaouk /mls-eng-10k-tags_tagged_10k_generated Dataset Card for Annotations of 10K hours of English MLS This dataset consists in annotations of a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours… See the full description on the dataset page: https://huggingface.co/datasets/pharaouk/mls-eng-10k-tags_tagged_10k_generated.tabularautomatic-speech-recognition1M<n<10M0 likes62 downloads2y agoHugging Face23taras-sereda /uk-pods uk-pods - speech datasets of Ukrainian podcasts. Preparation Clone the dataset repository and extract the content of clips.tar.gz archive. git clone https://huggingface.co/datasets/taras-sereda/uk-pods cd uk-pods && tar -zxvf clips.tar.gz To use these manifests for training/inference with NeMo [1] modify audio_filepath to absolute locations of audio files extracted in previous step. # data_root=<clonned_repo_dir> # /home/ubuntu/uk-pods data_root=$(realpath .) sed -i… See the full description on the dataset page: https://huggingface.co/datasets/taras-sereda/uk-pods.audioautomatic-speech-recognition10K<n<100K1 likes59 downloads2y agoHugging Face24RyeAI /ftspeech-pnc-da Dataset Card for ftspeech-pnc-da Dataset Summary RyeAI/ftspeech-pnc-da is a text-only Danish punctuation and capitalization companion dataset derived from the original Hugging Face dataset alexandrainst/ftspeech: https://huggingface.co/datasets/alexandrainst/ftspeech Each row contains restored punctuated text for an existing FTSpeech training utterance together with identifiers that allow the row to be joined back to the original source dataset: utterance_id:… See the full description on the dataset page: https://huggingface.co/datasets/RyeAI/ftspeech-pnc-da.tabularautomatic-speech-recognition100K<n<1M1 likes50 downloads19d agoHugging Face25ZamAI-Pashto /zamai-pashto-voice2voice ZamAI Pashto Voice2Voice This dataset contains Pashto voice-to-voice preparation metadata for speech and translation experiments. It focuses on Pashto speech records, dialect information, transcript text, and a small viewer-ready sample manifest. Configs from datasets import load_dataset metadata = load_dataset("ZamAI-Pashto/zamai-pashto-voice2voice", "metadata") sample = load_dataset("ZamAI-Pashto/zamai-pashto-voice2voice", "viewer_sample") Files… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-voice2voice.tabularautomatic-speech-recognitionn<1K0 likes37 downloads2mo agoHugging Face26ergar /PapaChaves PapaChaves 🇨🇷 PapaChaves is a longitudinal corpus of presidential press conferences from the administration of Rodrigo Chaves Robles, President of Costa Rica (May 2022 – May 2026). The dataset contains automatic speech transcriptions of 308 press conferences, covering the full presidential term. "PapaChaves" was a nickname given to President Chaves that leaked into the press during his administration. Dataset Summary Stat Value Videos 308 Total audio… See the full description on the dataset page: https://huggingface.co/datasets/ergar/PapaChaves.tabularautomatic-speech-recognitionn<1K0 likes32 downloads4mo agoHugging Face27pavanyellow /librispeech_asr Dataset Card for librispeech_asr LibriSpeech ASR 2s Splits Dataset Version of LibriSpeech ASR corpus split into 2s clips. Usage from datasets import load_dataset # Load the dataset from the Hub dataset = load_dataset("pavanyellow/librispeech_asr") # Or load a specific split dataset = load_dataset("pavanyellow/librispeech_asr", split="train") # Access the data for example in dataset['train'][:5]: audio = example['audio'] text = example['text'] tabularautomatic-speech-recognition10K<n<100K0 likes31 downloads2y agoHugging Face28tasal9 /zamai-pashto-voice2voice ZamAI Pashto Voice2Voice Languages: psLicense: cc-by-4.0Task categories: automatic-speech-recognition, audio-to-audioSize categories: n<1K Summary This dataset is part of the ZamAI Pashto data collection. It is intended for automatic-speech-recognition, audio-to-audio tasks in Pashto. How to use from datasets import load_dataset dataset = load_dataset("tasal9/zamai-pashto-voice2voice") print(dataset) Configs default: load with… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/zamai-pashto-voice2voice.audioautomatic-speech-recognition1K<n<10K1 likes29 downloads2mo agoHugging Face29RyeAI /coral-v3-conversation-pnc-dagated CoRal v3 Conversation PnC DA RyeAI/coral-v3-conversation-pnc-da is a text-only Danish punctuation and capitalization companion for the conversation training split of CoRal-project/coral-v3. It contains 102,226 restored transcript rows and no audio bytes. Pinned companion revision: 6e4fbafde87fbffadd58bbe39a3a2e09e884351a. SHA-256 of data/train-00000-of-00001.parquet: 79d30671f238e884a5b71b682bc811043e7df075566ce5565a236e66ffe32000. Each row can be joined back to the gated source… See the full description on the dataset page: https://huggingface.co/datasets/RyeAI/coral-v3-conversation-pnc-da.tabularautomatic-speech-recognition100K<n<1M0 likes23 downloads19d agoHugging Face30neurips2026-pi-bench /pi_bench pi-bench pi-bench is a multi-task audio benchmark prepared for public hosting and evaluation reproducibility. The repository is organized as a Hugging Face dataset with one dataset config per task file under data/, so each benchmark subset is visible and loadable independently. Overview The current release contains 11 task-specific configs spanning three broad categories: Counterfactual/contextual QA (CTC_*) Clarification-seeking QA (Trivia_Clarification_*… See the full description on the dataset page: https://huggingface.co/datasets/neurips2026-pi-bench/pi_bench.audioautomatic-speech-recognition1K<n<10K1 likes17 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.