CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Granary Granary: Speech Recognition and Translation Dataset in 25 European Languages Granary is a large-scale, open-source multilingual speech dataset covering 25 European languages for Automatic Speech Recognition (ASR) and Automatic Speech Translation (AST) tasks. Overview Granary addresses the scarcity of high-quality speech data for low-resource languages by consolidating multiple datasets under a unified framework: 🗣️ ~1M hours of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Granary.tabularautomatic-speech-recognition100M<n<1B219 likes5.3k downloads3mo agoHugging Face02RVtech /Audio2Tool Audio2Tool: Speak, Call, Act — A Dataset for Benchmarking Speech Tool Use Authors: Ramit Pahwa1,∗,∗∗, Apoorva Beedu1,∗, Parivesh Priye1, Rutu Gandhi†1, Saloni Takawale†1, Aruna Baijal1, Zengli Yang1 1 Rivian & Volkswagen Technologies &nbsp;·&nbsp; ∗ equal contribution &nbsp;·&nbsp; ∗∗ corresponding author &nbsp;·&nbsp; † equal contribution 📄 Project page / demo: https://audio2tool.github.io/ 📦 Dataset: https://huggingface.co/datasets/RVtech/Audio2Tool ✉️ Contact (corresponding… See the full description on the dataset page: https://huggingface.co/datasets/RVtech/Audio2Tool.audioautomatic-speech-recognition10K<n<100K2 likes5k downloads3mo agoHugging Face03zhifeixie /StreamAudio-2M StreamAudio-2M Large-scale streaming-audio dataset for audio-LLM / audio-agent training. Each row is a stream: a sequence of audio turns sharing one unified schema. ~2.28M unique audio clips are organised into six task subsets. Subsets Subset Rows Description Stream_Audio_Understanding 90,738 Montages of audio-understanding clips (AudioSet / FMA): captions, choice & open QA Real_time_ASR 28,109 Streams of ASR clips (CommonVoice / GigaSpeech /… See the full description on the dataset page: https://huggingface.co/datasets/zhifeixie/StreamAudio-2M.tabularaudio-classification100K<n<1M30 likes2.8k downloads4mo agoHugging Face04wayu-ai /thai-aligner-bench Thai Aligner Bench 🚧 Development in progress. How accurately can a forced aligner place Thai token and word boundaries in speech? This is a self-contained benchmark: one Python file (aligner_bench.py) plus 1,572 clips of Thai speech with frame-exact timing ground truth. No Thai NLP stack or other code is needed — just numpy soundfile torch torchaudio transformers. The ground truth is what makes the dataset useful: the audio was rendered by a TTS model whose duration predictor… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-aligner-bench.audioautomatic-speech-recognition1K<n<10K1 likes775 downloads1mo agoHugging Face05besimple-ai /voice-code-bench VoiceCodeBench VoiceCodeBench is a test-only benchmark for evaluating whether automatic speech recognition (ASR) systems preserve exact structured values in English workplace speech. Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition The benchmark targets cases where a transcript is software input: callback numbers, email addresses, command-line flags, file paths, URLs, account identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/besimple-ai/voice-code-bench.audioautomatic-speech-recognitionn<1K13 likes729 downloads10d agoHugging Face06Quran-Lab /quran-tajweed-phonetics The complete phonetic layer of the Quran in the riwaya of Hafs 'an 'Asim via tariq al-Shatibiyyah: 6,236 ayat, 522,475 phones, every phone carrying its tajweed attribution: madd class with its transmitted length range, ghunna grade, qalqalah class, tafkheem with its rank, sakt, the seventeen sifat, and the rule that produced it. Built and maintained by Quran Lab, a waqf building open technology in the service of the Quran. How it was built and verified Indexed from the… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/quran-tajweed-phonetics.tabularautomatic-speech-recognition10K<n<100K3 likes550 downloads9d agoHugging Face07ArtificialAnalysis /Earnings22-Cleaned-AA-chunked Earnings22-Cleaned-AA-chunked Quick links: AA Streaming Speech to Text Leaderboard | Speech to Text methodology Earnings22-Cleaned-AA-chunked is a chunked version of Earnings22-Cleaned-AA, the cleaned Earnings-22 subset used by Artificial Analysis for streaming Speech to Text evaluation. The original Earnings-22 data comes from esb/datasets, a corpus of corporate earnings calls. Artificial Analysis manually reviewed and corrected the reference transcripts in the cleaned subset… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA-chunked.audioautomatic-speech-recognitionn<1K1 likes533 downloads3mo agoHugging Face08OPPOer /HearInContextEnglish | 中文 HearInContext A Benchmark for Implicit Context in Speech Recognition Illustrative example: the same spoken request is disambiguated as flour or flower by different assistant histories. The dialogue and waveform are illustrative. Same audio. Different contexts. Different meanings. HearInContext is a Mandarin–English contextual speech recognition benchmark. It pairs the same audio with dialogue histories supporting different meanings to evaluate… See the full description on the dataset page: https://huggingface.co/datasets/OPPOer/HearInContext.audioautomatic-speech-recognition100K<n<1M1 likes366 downloads1d agoHugging Face09abnajlae /darija-asr-corpus Darija ASR Corpus (dataset-core) Arabizi (Latin-script) transcriptions of Moroccan Darija speech, produced for a Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming). This repo contains four source subsets: DODa, DVoice, Wiki, and YouTube. Each subset carries its own upstream license/terms -- see below -- because they are drawn from four different original projects. Subsets Config Rows Audio bundled? Upstream license Upstream source… See the full description on the dataset page: https://huggingface.co/datasets/abnajlae/darija-asr-corpus.audioautomatic-speech-recognition10K<n<100K0 likes364 downloads14d agoHugging Face10gavinlaw /chinese-lips-speech-slide-probe Chinese-LiPS Speech + Slide Probe A self-contained probe set for testing whether visual slide context helps simultaneous speech translation — with the input as audio, not transcripts. Why audio matters: feeding a transcript to a text LLM deletes the acoustic ambiguity (homophones, polysemy) that slide context is meant to resolve; the transcript already commits to one reading. Any honest test of "does vision help streaming ST" must consume speech. Contents… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/chinese-lips-speech-slide-probe.audiotranslationn<1K0 likes314 downloads2mo agoHugging Face11nymtheescobar /bengali-talkshow-audio Bengali Talkshow Audio Dataset A large-scale collection of 1,180 Bengali talk show audio recordings totaling 789+ hours of multi-speaker speech, sourced from Bangladeshi television talk shows and political debate programs. Dataset Description This dataset contains audio from Bengali-language TV talk shows, political debates, and news discussion programs from major Bangladeshi television channels. Each recording features multiple speakers engaged in discussion, making it… See the full description on the dataset page: https://huggingface.co/datasets/nymtheescobar/bengali-talkshow-audio.audioaudio-classification1K<n<10K0 likes278 downloads8mo agoHugging Face12vnmoorthy /pavo-bench PAVO-Bench: 50K-Turn Benchmark for ASR-LLM-TTS Pipeline Routing Code: github.com/vnmoorthy/pavo-bench · Paper: TMLR 2026 (accepted) · Authors: NarasingaMoorthy VeiluKanthaPerumal (UPenn), Mohammed Imthathullah (Google) pip install git+https://github.com/vnmoorthy/pavo-bench.git Headline results (vs fixed-cloud baseline, 50,000 voice turns) Metric Result Significance P95 end-to-end latency (H100, LibriSpeech) −10.3% (−167 ms) — Median latency −34%… See the full description on the dataset page: https://huggingface.co/datasets/vnmoorthy/pavo-bench.documentautomatic-speech-recognition10K<n<100K0 likes195 downloads1mo agoHugging Face13thepowerfuldeez /massive-yt-edu-queue Massive YouTube Educational Video Queue Full metadata and content classification for 4,489,228 YouTube educational videos totaling 3,975,157 hours. Description This dataset contains metadata, content categorization, and license risk assessment for ~4.5M YouTube videos identified as potentially educational. It serves as the discovery and processing queue for the massive-yt-edu-transcriptions project, which aims to create the world's largest open educational transcript… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/massive-yt-edu-queue.tabularautomatic-speech-recognition1M<n<10M1 likes191 downloads7mo agoHugging Face14Edge0 /ark-asr-3b-open-asr-leaderboard-results ARK-ASR-3B Open ASR Leaderboard Results Raw JSONL manifests for AutoArk-AI/ARK-ASR-3B on the public English short-form hf-audio/open-asr-leaderboard splits. These manifests were generated on a local 8x RTX 4090 machine and scored with the shared Open ASR Leaderboard scorer: PYTHONPATH=. python - <<'PY' from normalizer.eval_utils import score_results score_results( 'ark_asr/results.AutoArk-AI-ARK-ASR-3B_20260622_official', 'AutoArk-AI/ARK-ASR-3B', ) PY Important:… See the full description on the dataset page: https://huggingface.co/datasets/Edge0/ark-asr-3b-open-asr-leaderboard-results.tabularautomatic-speech-recognition10K<n<100K12 likes190 downloads3mo agoHugging Face15danielrosehill /Tech-Sentences-For-ASR-Training TechVoice Dataset Work in Progress – This dataset is actively being expanded with new recordings. Dataset Statistics Metric Current Target Progress Duration 38m 43s 5h 0m 0s ██░░░░░░░░░░░░░░░░░░ 12.9% Words 10,412 50,000 ████░░░░░░░░░░░░░░░░ 20.8% Total Recordings: 205 samples Total Characters: 74,312 A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.audioautomatic-speech-recognitionn<1K2 likes174 downloads10mo agoHugging Face16ground-truth /multichannel-meetings-10h GroundTruth Multi-Channel Meeting Audio Dataset (10h) Summary This dataset contains approximately 10 hours of co-located, multi-speaker meeting recordings, each captured simultaneously via a room (built-in) microphone and individual close-talk lapel microphones worn by each participant. Each meeting includes: One full meeting recording (room microphone) Individual close-talk recordings for each participant (one file per speaker) Structured metadata describing speakers… See the full description on the dataset page: https://huggingface.co/datasets/ground-truth/multichannel-meetings-10h.audioautomatic-speech-recognitionn<1K1 likes170 downloads5mo agoHugging Face17lilonghao /MM-ContextASR-Bench MM-ContextASR Bench Metadata and evaluation splits for Multimodal Conversational Context for LLM-Based ASR: Data Construction, Training, and Benchmark. Dataset summary Config Examples Audio Context Primary metric mm_contextasr 1,250 (250 current utterances × 5 histories) 1,439 WAV files included Controlled user-assistant dialogue entity Recall kespeech 19,212 Source ID only Same-speaker speech and transcript CER, SER, entity Recall cv_yue 3,525… See the full description on the dataset page: https://huggingface.co/datasets/lilonghao/MM-ContextASR-Bench.audioautomatic-speech-recognition10K<n<100K1 likes169 downloads5d agoHugging Face18FluidInference /cv-corpus-25.0-ja Mozilla Common Voice 25.0 - Japanese Test Set (Complete) Dataset Description Complete Japanese test set from Mozilla Common Voice Corpus 25.0. This dataset contains all 9,019 validated test samples, compared to the partial 2,334-sample version previously available on HuggingFace. Key Features Size: 9,019 validated test utterances Coverage: 100% of official Common Voice 25.0 Japanese test split Multi-speaker: Diverse set of speakers with demographic metadata… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/cv-corpus-25.0-ja.audioautomatic-speech-recognition1K<n<10K0 likes162 downloads6mo agoHugging Face19yunqi1766 /voice-code-bench VoiceCodeBench VoiceCodeBench is a test-only benchmark for evaluating whether automatic speech recognition (ASR) systems preserve exact structured values in English workplace speech. Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition The benchmark targets cases where a transcript is software input: callback numbers, email addresses, command-line flags, file paths, URLs, account identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/yunqi1766/voice-code-bench.audioautomatic-speech-recognitionn<1K1 likes135 downloads2mo agoHugging Face20Wi-Fi /korean-full-duplex-synthetic-dataset-preview Korean Full-Duplex Synthetic Dataset Preview Overview Public preview of a Korean full-duplex synthetic speech dataset. This repository contains 100 conversations sampled from a corpus of 89,273 conversations (2,000.5 hours); it does not publish the full corpus audio. Preview contents 100 conversation WAV files data/representative.jsonl 24 kHz, mono, 16-bit PCM Events: normal, barge_in, backchannel, cutoff_by_user Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.audioautomatic-speech-recognitionn<1K1 likes133 downloads1mo agoHugging Face21tsdocode /open-vi-dialog-synthetic-100h OpenDialog Vietnamese Synthetic Dialogue 100h Synthetic Vietnamese two-speaker dialogue for ZipVoice-Dialog experiments. 12,000 chunks 30 seconds per chunk 100.0 hours total Each item contains S1/S2 speaker labels, turn timings, target text, relationship, pronouns, environment, topic, mood, and source reference IDs. Audio renderer: vLLM-Omni VoxCPM2 Audio format: mono WAV, 48 kHz, 30 seconds per chunk This is a research dataset. Review the source/reference licensing and the… See the full description on the dataset page: https://huggingface.co/datasets/tsdocode/open-vi-dialog-synthetic-100h.audiotext-to-speech10K<n<100K0 likes113 downloads1mo agoHugging Face22eQOURSE /multilingual-speech Multilingual Indian Conversational Speech A dataset of naturalistic, spontaneous two-speaker conversations across 13 Indian languages, with segment-level transcripts, speaker profiles, timestamps, and recording metadata. Designed for ASR, TTS, speaker diarization, and conversational speech research. Languages (13) Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Nepali, Odia, Punjabi, Tamil, Telugu, Urdu. Content Conversations… See the full description on the dataset page: https://huggingface.co/datasets/eQOURSE/multilingual-speech.audioautomatic-speech-recognitionn<1K0 likes93 downloads3mo agoHugging Face23Andrew0425 /AASR-Bench AASRBench AASRBench is a bilingual audio benchmark for evaluating ASR output on spoken-language cleanup, correction, formatting, filtering, and rephrasing. Contents 917 benchmark samples. 917 WAV files referenced by the benchmark records. 6,637 rubric questions across content, format, filter, and rephrase. 11 scenes: academic, customer_service, daily_chat, dictation_memo, explanation, meeting, navigation, passthrough, tech, vibe_coding, and voice_search. 510… See the full description on the dataset page: https://huggingface.co/datasets/Andrew0425/AASR-Bench.audioautomatic-speech-recognitionn<1K0 likes84 downloads2mo agoHugging Face24TNSA /Aren ARen — Arabic/English ASR Robustness Set Curated and published by TNSA AI. A small, deliberately hard evaluation set for Arabic and English speech recognition. Every clip exists in three acoustic conditions so you can measure not just how a model scores, but how fast it falls apart as the channel degrades. Built because clean read-speech benchmarks stop discriminating between modern ASR systems long before real deployments stop breaking. Why it exists On clean… See the full description on the dataset page: https://huggingface.co/datasets/TNSA/Aren.audioautomatic-speech-recognitionn<1K0 likes83 downloads1mo agoHugging Face25SaarAI /asr-benchmark-outputsgated SaarAI ASR Benchmark Outputs Raw per-utterance model outputs (transcription manifests) produced by the gsma-asr-bench runners on SaarAI/asr-leaderboard-datasets. files: 501 utterances: 4374087 languages: 7 models: 47 Layout data/<language_name>/<split>__<dataset_config>__<model_slug>.jsonl index.jsonl # one record per file (language, split, model, rows, sha256, ...) index.csv Directories categorise by language name; the file name begins with the split name… See the full description on the dataset page: https://huggingface.co/datasets/SaarAI/asr-benchmark-outputs.tabularautomatic-speech-recognition1M<n<10M1 likes75 downloads4h agoHugging Face26benderrodriguez /hebrew-asr-vn Hebrew ASR three-source training dataset Derived from ivrit.ai crowd-transcribe-v5, crowd-recital and knesset-committees. Original data and transcripts are credited to ivrit.ai and its contributors. Pinned revisions and preparation rules are in metadata/sources.json and metadata/preparation-config.json. VoxKnesset is excluded by user decision. Source/split Clips Hours crowd-recital/test 1,557 1.071 crowd-recital/train 45,372 33.258 crowd-recital/validation 1,033… See the full description on the dataset page: https://huggingface.co/datasets/benderrodriguez/hebrew-asr-vn.tabularautomatic-speech-recognition1M<n<10M0 likes67 downloads8d agoHugging Face27gavinlaw /chinese-lips-longform-debug Chinese-LiPS Long-Form (zh long streaming speech) Reconstructed continuous long-speech streams from BAAI/Chinese-LiPS, for slide-aware / streaming speech-translation development and evaluation. Each source video (one speaker, one scripted lecture with slides) was released as pre-segmented clips; here they are re-joined into the full talk. Two variants of the same 3 talks (~97 min speech total): config how segments are placed use orig_timeline at their original session… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/chinese-lips-longform-debug.audioautomatic-speech-recognition1K<n<10K0 likes66 downloads2mo agoHugging Face28mesolitica /pseudolabel-malaya-speech-stt-train-whisper-large-v3tabularautomatic-speech-recognition1M<n<10M1 likes64 downloads3y agoHugging Face29taras-sereda /uk-pods uk-pods - speech datasets of Ukrainian podcasts. Preparation Clone the dataset repository and extract the content of clips.tar.gz archive. git clone https://huggingface.co/datasets/taras-sereda/uk-pods cd uk-pods && tar -zxvf clips.tar.gz To use these manifests for training/inference with NeMo [1] modify audio_filepath to absolute locations of audio files extracted in previous step. # data_root=<clonned_repo_dir> # /home/ubuntu/uk-pods data_root=$(realpath .) sed -i… See the full description on the dataset page: https://huggingface.co/datasets/taras-sereda/uk-pods.audioautomatic-speech-recognition10K<n<100K1 likes53 downloads2y agoHugging Face30laion /soundscape-bench SoundScape-Bench 200 held-out multilingual soundscapes with exact, automatically-gradable answer keys for evaluating "universal audio annotation" — the task of describing everything audible in a clip (speech, who/when/ what/which-language/how-it-is-said, sound effects, music, and vocal bursts) as one structured JSON list. It is the benchmark for the LAION Universal Audio Annotation Pipeline (UAAP). Why it exists Every clip is built by gluing together pieces we… See the full description on the dataset page: https://huggingface.co/datasets/laion/soundscape-bench.audioaudio-classificationn<1K1 likes42 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.