CoolFace
29 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01overflowwwww /yt-danish-public-v2audioaudio-classification100K<n<1M0 likes2.3k downloads2y agoHugging Face02syvai /danish-asr-unified Danish ASR Unified Dataset Unified Danish speech recognition dataset combining 7 sources (~3.5M samples, ~16k hours): Source Samples Description VoxPopuli 1,775,578 European Parliament recordings ftspeech 995,677 Danish Parliament (Folketinget) CoRal-v3 read_aloud 299,255 Read-aloud Danish speech nst-da 182,605 NST Danish speech CoRal-v3 conversation 147,249 Conversational Danish speech nota 98,600 Danish broadcast media Common Voice 17 3,484 Crowd-sourced… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-unified.audioautomatic-speech-recognition1M<n<10M4 likes2.1k downloads2mo agoHugging Face03syvai /danish-asr-unified-hviske-v5-tinygated danish-asr-unified — two-model labels and a quality manifest Transcriptions, per-token confidences, and a per-row quality verdict for every row of syvai/danish-asr-unified (3,414,589 rows, 8 sources). Two independently trained models labelled the whole corpus: model architecture vocabulary syvai/hviske-v5-tiny encoder-decoder 16,384 BPE 3dio-ai/svale-110M RNN-T (Parakeet) 44 characters Both decode greedily. Shards mirror the source data/train-*.parquet by name… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-unified-hviske-v5-tiny.tabularautomatic-speech-recognition1M<n<10M0 likes1.9k downloads8d agoHugging Face04RyeAI /danish-asr-leaderboard Open Danish ASR Leaderboard — Results Benchmark results backing the Open Danish ASR Leaderboard — an open, reproducible comparison of Danish speech-to-text models. Every model is transcribed and scored identically on the same five independent public Danish test sets, so the numbers compare directly: Word Error Rate (WER) and Character Error Rate (CER) — lower is better — plus speed. Open-weight Danish speech recognition models you can run yourself and hosted transcription APIs… See the full description on the dataset page: https://huggingface.co/datasets/RyeAI/danish-asr-leaderboard.tabularautomatic-speech-recognition100K<n<1M4 likes1.6k downloads10h agoHugging Face05syvai /danish-asr-verified danish-asr-verified ALL rows of syvai/danish-asr-unified transcribed by the syv-transcribe ensemble (hviske-v5.3 + hviske-v5, confidence-weighted ROVER), each annotated with: verified — True when the ensemble independently reproduced the reference exactly (compared after lowercasing, punctuation-strip, whitespace-collapse). Two independent witnesses agree => near-certain label. wer_teacher_vs_ref / cer_teacher_vs_ref — word/character error rate between normalized teacher output… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-verified.tabularautomatic-speech-recognition1M<n<10M0 likes417 downloads2mo agoHugging Face06danielshaps /nchlt_speech_zul NCHLT Speech Corpus -- isiZulu This is the isiZulu language part of the NCHLT Speech Corpus of the South African languages. Language code (ISO 639): zul URI: https://hdl.handle.net/20.500.12185/275 Licence: Creative Commons Attribution 3.0 Unported License (CC BY 3.0): http://creativecommons.org/licenses/by/3.0/legalcode Attribution: The Department of Arts and Culture of the government of the Republic of South Africa (DAC), Council for Scientific and Industrial… See the full description on the dataset page: https://huggingface.co/datasets/danielshaps/nchlt_speech_zul.audioautomatic-speech-recognition10K<n<100K1 likes187 downloads2y agoHugging Face07danielshaps /nchlt_speech_afr NCHLT Speech Corpus -- Afrikaans This is the Afrikaans language part of the NCHLT Speech Corpus of the South African languages. Language code (ISO 639): afr URI: https://hdl.handle.net/20.500.12185/280 Licence: Creative Commons Attribution 3.0 Unported License (CC BY 3.0): http://creativecommons.org/licenses/by/3.0/legalcode Attribution: The Department of Arts and Culture of the government of the Republic of South Africa (DAC), Council for Scientific and… See the full description on the dataset page: https://huggingface.co/datasets/danielshaps/nchlt_speech_afr.audioautomatic-speech-recognition10K<n<100K0 likes148 downloads2y agoHugging Face08danielrosehill /Tech-Sentences-For-ASR-Training TechVoice Dataset Work in Progress – This dataset is actively being expanded with new recordings. Dataset Statistics Metric Current Target Progress Duration 38m 43s 5h 0m 0s ██░░░░░░░░░░░░░░░░░░ 12.9% Words 10,412 50,000 ████░░░░░░░░░░░░░░░░ 20.8% Total Recordings: 205 samples Total Characters: 74,312 A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.audioautomatic-speech-recognitionn<1K2 likes144 downloads10mo agoHugging Face09danijelkorzinek /ClarinStudioPL CLARIN-PL Polish Studio Corpus The corpus was created somewhere in 2014-2015 by recording a group of few hundred volunteer speakers reading a few dozen sentences each. The total size of the corpus is ~56 hours. Due to the manner of recording, the transcription accuracy is very high, but the manner of speech is not spontaneous. This corpus is best compared to something like TIMIT, possibly CommonVoice. It is different from CommonVoice in that it is recorded in a controlled… See the full description on the dataset page: https://huggingface.co/datasets/danijelkorzinek/ClarinStudioPL.audioautomatic-speech-recognition10K<n<100K2 likes103 downloads2y agoHugging Face10danielshaps /nchlt_speech_tso NCHLT Speech Corpus -- Xitsonga This is the Xitsonga language part of the NCHLT Speech Corpus of the South African languages. Language code (ISO 639): tso URI: https://hdl.handle.net/20.500.12185/277 Licence: Creative Commons Attribution 3.0 Unported License (CC BY 3.0): http://creativecommons.org/licenses/by/3.0/legalcode Attribution: The Department of Arts and Culture of the government of the Republic of South Africa (DAC), Council for Scientific and… See the full description on the dataset page: https://huggingface.co/datasets/danielshaps/nchlt_speech_tso.audioautomatic-speech-recognition10K<n<100K1 likes101 downloads2y agoHugging Face11danielshaps /nchlt_speech_eng NCHLT Speech Corpus -- South African English This is the South African English language part of the NCHLT Speech Corpus of the South African languages. Language code (ISO 639): eng URI: https://hdl.handle.net/20.500.12185/274 Licence: Creative Commons Attribution 3.0 Unported License (CC BY 3.0): http://creativecommons.org/licenses/by/3.0/legalcode Attribution: The Department of Arts and Culture of the government of the Republic of South Africa (DAC), Council… See the full description on the dataset page: https://huggingface.co/datasets/danielshaps/nchlt_speech_eng.audioautomatic-speech-recognition10K<n<100K0 likes86 downloads2y agoHugging Face12danielrosehill /Audio-Understanding-Bitrate-Eval-0426 Audio Understanding — MP3 Bitrate Evaluation (April 2026) Empirical eval measuring how MP3 compression bitrate affects transcription accuracy across every audio-input LLM available on OpenRouter. 📝 Blog post: MP3 Bitrate Sensitivity in Audio-Multimodal LLMs 💻 Code & methodology: github.com/danielrosehill/Audio-Understanding-Bitrate-Eval-0426 TL;DR Ran a benchmark across 12 OpenRouter audio-multimodal models × 4 dictation samples × 5 MP3 bitrates (16/24/32/48/64 kbps)… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Audio-Understanding-Bitrate-Eval-0426.textautomatic-speech-recognitionn<1K0 likes78 downloads5mo agoHugging Face13danielrosehill /Small-STT-Eval-Audio-Dataset Small STT Eval Audio Dataset A small speech-to-text evaluation dataset containing 92 audio samples with ground truth transcriptions. Designed for evaluating STT systems on technical vocabulary, code-switching (English/Hebrew), and various speaking styles. Dataset Description This dataset contains audio recordings with accompanying transcriptions across multiple categories: Category Count Description tech_github 5 GitHub-related technical vocabulary… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Small-STT-Eval-Audio-Dataset.audioautomatic-speech-recognitionn<1K0 likes66 downloads10mo agoHugging Face14danielshaps /nchlt_speech_tsn NCHLT Speech Corpus -- Setswana This is the Setswana language part of the NCHLT Speech Corpus of the South African languages. Language code (ISO 639): tsn URI: https://hdl.handle.net/20.500.12185/281 Licence: Creative Commons Attribution 3.0 Unported License (CC BY 3.0): http://creativecommons.org/licenses/by/3.0/legalcode Attribution: The Department of Arts and Culture of the government of the Republic of South Africa (DAC), Council for Scientific and… See the full description on the dataset page: https://huggingface.co/datasets/danielshaps/nchlt_speech_tsn.audioautomatic-speech-recognition10K<n<100K0 likes54 downloads2y agoHugging Face15danielshaps /nchlt_speech_sot NCHLT Speech Corpus -- Sesotho This is the Sesotho language part of the NCHLT Speech Corpus of the South African languages. Language code (ISO 639): sot URI: https://hdl.handle.net/20.500.12185/278 Licence: Creative Commons Attribution 3.0 Unported License (CC BY 3.0): http://creativecommons.org/licenses/by/3.0/legalcode Attribution: The Department of Arts and Culture of the government of the Republic of South Africa (DAC), Council for Scientific and Industrial… See the full description on the dataset page: https://huggingface.co/datasets/danielshaps/nchlt_speech_sot.audioautomatic-speech-recognition10K<n<100K0 likes52 downloads2y agoHugging Face16danielshaps /nchlt_speech_nbl NCHLT Speech Corpus -- isiNdebele This is the isiNdebele language part of the NCHLT Speech Corpus of the South African languages. Language code (ISO 639): nbl URI: https://hdl.handle.net/20.500.12185/272 Licence: Creative Commons Attribution 3.0 Unported License (CC BY 3.0): http://creativecommons.org/licenses/by/3.0/legalcode Attribution: The Department of Arts and Culture of the government of the Republic of South Africa (DAC), Council for Scientific and… See the full description on the dataset page: https://huggingface.co/datasets/danielshaps/nchlt_speech_nbl.audioautomatic-speech-recognition10K<n<100K0 likes47 downloads2y agoHugging Face17datadriven-company /TTS-Danish TTS-Danish A large-scale, high-quality Danish speech dataset for text-to-speech and automatic speech recognition. Data Sources This dataset combines three sources: Source Samples Hours License Content lydbog.com 35,719 92.0 CC-BY-SA 4.0 Danish classic literature, read by Kristoffer Hunsdahl CoRal-TTS (Alexandra Institute) 19,996 30.5 CC0 Professional TTS recordings, 2 speakers LibriVox 0 0.0 Public Domain Danish audiobooks Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Danish.audiotext-to-speech10K<n<100K0 likes44 downloads7mo agoHugging Face18danielshaps /nchlt_speech_nso NCHLT Speech Corpus -- Sepedi This is the Sepedi language part of the NCHLT Speech Corpus of the South African languages. Language code (ISO 639): nso URI: https://hdl.handle.net/20.500.12185/270 Licence: Creative Commons Attribution 3.0 Unported License (CC BY 3.0): http://creativecommons.org/licenses/by/3.0/legalcode Attribution: The Department of Arts and Culture of the government of the Republic of South Africa (DAC), Council for Scientific and Industrial… See the full description on the dataset page: https://huggingface.co/datasets/danielshaps/nchlt_speech_nso.audioautomatic-speech-recognition10K<n<100K0 likes42 downloads2y agoHugging Face19danielshaps /nchlt_speech_xho NCHLT Speech Corpus -- isiXhosa This is the isiXhosa language part of the NCHLT Speech Corpus of the South African languages. Language code (ISO 639): xho URI: https://hdl.handle.net/20.500.12185/279 Licence: Creative Commons Attribution 3.0 Unported License (CC BY 3.0): http://creativecommons.org/licenses/by/3.0/legalcode Attribution: The Department of Arts and Culture of the government of the Republic of South Africa (DAC), Council for Scientific and… See the full description on the dataset page: https://huggingface.co/datasets/danielshaps/nchlt_speech_xho.audioautomatic-speech-recognition10K<n<100K0 likes41 downloads2y agoHugging Face20danieldzikunuofmarvel /bibletts-asante-twi-repaired BibleTTS Asante Twi — Repaired Transcripts The Asante Twi transcripts released with BibleTTS have had the characters ɛ (U+025B) and ɔ (U+0254) stripped out. This dataset restores them. Audio is not included. This is a drop-in replacement for the .txt files that ship with the BibleTTS Asante Twi package, matched by clip ID. The problem Both are Twi vowels, and both are required by the orthography. Measured across the released Asante Twi transcripts: Character… See the full description on the dataset page: https://huggingface.co/datasets/danieldzikunuofmarvel/bibletts-asante-twi-repaired.tabularautomatic-speech-recognition10K<n<100K0 likes36 downloads2mo agoHugging Face21danielshaps /nchlt_speech_ven NCHLT Speech Corpus -- Tshivenda This is the Tshivenda language part of the NCHLT Speech Corpus of the South African languages. Language code (ISO 639): ven URI: https://hdl.handle.net/20.500.12185/276 Licence: Creative Commons Attribution 3.0 Unported License (CC BY 3.0): http://creativecommons.org/licenses/by/3.0/legalcode Attribution: The Department of Arts and Culture of the government of the Republic of South Africa (DAC), Council for Scientific and… See the full description on the dataset page: https://huggingface.co/datasets/danielshaps/nchlt_speech_ven.audioautomatic-speech-recognition10K<n<100K0 likes32 downloads2y agoHugging Face22danielshaps /nchlt_speech_ssw NCHLT Speech Corpus -- siSwati This is the siSwati language part of the NCHLT Speech Corpus of the South African languages. Language code (ISO 639): ssw URI: https://hdl.handle.net/20.500.12185/271 Licence: Creative Commons Attribution 3.0 Unported License (CC BY 3.0): http://creativecommons.org/licenses/by/3.0/legalcode Attribution: The Department of Arts and Culture of the government of the Republic of South Africa (DAC), Council for Scientific and Industrial… See the full description on the dataset page: https://huggingface.co/datasets/danielshaps/nchlt_speech_ssw.audioautomatic-speech-recognition10K<n<100K0 likes26 downloads2y agoHugging Face23daniel-dona /openslr-slr67See: https://www.openslr.org/67/ audioautomatic-speech-recognition10K<n<100K0 likes25 downloads2y agoHugging Face24syvai /danish-diarization-bench Danish Diarization Benchmark (Synthetic) — v2 A 3996-row synthetic speaker-diarization benchmark in Danish, built by mixing single-speaker utterances from syvai/danish-asr-unified into multi-speaker recordings. What changed in v2 (2026-05-18) Per-segment text — each entry in segments now carries its text field directly. The redundant parallel texts column has been removed. Old consumers that joined segments[i] with texts[i] should switch to segments[i]["text"]. Silent… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-diarization-bench.tabularautomatic-speech-recognition1K<n<10K1 likes25 downloads4mo agoHugging Face25Speech-data /Danish-Speech-Dataset 🎧 Danish Speech Dataset The Danish Speech Dataset is a high-quality speech audio dataset designed to provide structured and diverse audio data for AI-driven voice technologies. It includes 168 hours of audio data across 804 files, delivered in MP3 and WAV formats, with a total size of 369 MB. This well-balanced audio dataset ensures consistent and representative voice data, with 52% female and 48% male speakers, and a broad age distribution from 18 to 50+ years. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Danish-Speech-Dataset.audioautomatic-speech-recognitionn<1K0 likes11 downloads6mo agoHugging Face26djelia /danubegated danube A 1,000-utterance Bambara speech sample — 0.87 hours of 16 kHz audio with transcripts and per-utterance speaker labels. Small enough to be a working sample rather than a training corpus. Load from datasets import load_dataset ds = load_dataset("djelia/danube", split="train") print(ds[0]["text"], ds[0]["speaker_id"], ds[0]["duration"]) One config and one split. Config Split Rows Audio default train 1,000 0.865 h Fields… See the full description on the dataset page: https://huggingface.co/datasets/djelia/danube.audiotext-to-speech1K<n<10K0 likes8 downloads2mo agoHugging Face27Thomcles /YodaLingua-Danishgated YodaLingua-Danish YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Danish portion of the multilingual YodaLingua collection. 🧾 Dataset Overview Property Value Total clips 7,871 audio–transcription pairs Total duration 21 hours Speakers 925 distinct speakers Audio format MP3 • mono • 24 kHz • 16-bit… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Danish.audiotext-to-speech1K<n<10K1 likes8 downloads5mo agoHugging Face28DanishMahdi /SND_G2Pgated Sindhi G2P Dataset (SND_G2P) A Grapheme-to-Phoneme (G2P) dataset for the Sindhi language, mapping written words to their IPA (International Phonetic Alphabet) pronunciations. Dataset Description This dataset was compiled by scraping Wiktionary's Sindhi terms with IPA pronunciation category. For each Sindhi word, the corresponding IPA pronunciation was extracted from the word's Wiktionary entry under the Sindhi language section. Intended use cases: Grapheme-to-Phoneme… See the full description on the dataset page: https://huggingface.co/datasets/DanishMahdi/SND_G2P.texttext-to-speech10K<n<100K0 likes6 downloads5mo agoHugging Face29danijelkorzinek /PINCgated Polish Interpreting Corpus The Polish Interpreting Corpus (PINC) is a hand-verified parallel speech corpus derived from the European Parliament recordings. The corpus was automatically pre-processed and subsequently manually verified to correct the transcription, word-level speech-to-text alignment and sentence-level interlingual alignment. The audio quality is decent and the annotation is fairly accurate. The corpus contains a set of 520 recordings of Polish-English speeches… See the full description on the dataset page: https://huggingface.co/datasets/danijelkorzinek/PINC.audioautomatic-speech-recognition10K<n<100K0 likes4 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.