datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
yt-danish-public-v2danish-asr-unified
Danish ASR Unified Dataset
Unified Danish speech recognition dataset combining 7 sources (~3.5M samples, ~16k hours):
Source
Samples
Description
VoxPopuli
1,775,578
European Parliament recordings
ftspeech
995,677
Danish Parliament (Folketinget)
CoRal-v3 read_aloud
299,255
Read-aloud Danish speech
nst-da
182,605
NST Danish speech
CoRal-v3 conversation
147,249
Conversational Danish speech
nota
98,600
Danish broadcast media
Common Voice 17
3,484
Crowd-sourced… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-unified.danish-asr-unified-hviske-v5-tiny
danish-asr-unified — two-model labels and a quality manifest
Transcriptions, per-token confidences, and a per-row quality verdict for every
row of syvai/danish-asr-unified
(3,414,589 rows, 8 sources).
Two independently trained models labelled the whole corpus:
model
architecture
vocabulary
syvai/hviske-v5-tiny
encoder-decoder
16,384 BPE
3dio-ai/svale-110M
RNN-T (Parakeet)
44 characters
Both decode greedily. Shards mirror the source data/train-*.parquet by name… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-unified-hviske-v5-tiny.danish-asr-leaderboard
Open Danish ASR Leaderboard — Results
Benchmark results backing the Open Danish ASR Leaderboard — an open, reproducible comparison of Danish speech-to-text models.
Every model is transcribed and scored identically on the same five independent public Danish test sets, so the numbers compare directly: Word Error Rate (WER) and Character Error Rate (CER) — lower is better — plus speed. Open-weight Danish speech recognition models you can run yourself and hosted transcription APIs… See the full description on the dataset page: https://huggingface.co/datasets/RyeAI/danish-asr-leaderboard.danish-asr-verified
danish-asr-verified
ALL rows of syvai/danish-asr-unified transcribed by the
syv-transcribe ensemble (hviske-v5.3 + hviske-v5, confidence-weighted ROVER),
each annotated with:
verified — True when the ensemble independently reproduced the reference
exactly (compared after lowercasing, punctuation-strip, whitespace-collapse).
Two independent witnesses agree => near-certain label.
wer_teacher_vs_ref / cer_teacher_vs_ref — word/character error rate
between normalized teacher output… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-verified.nchlt_speech_zul
NCHLT Speech Corpus -- isiZulu
This is the isiZulu language part of the NCHLT Speech Corpus of the South African languages.
Language code (ISO 639): zul
URI: https://hdl.handle.net/20.500.12185/275
Licence:
Creative Commons Attribution 3.0 Unported License (CC BY 3.0): http://creativecommons.org/licenses/by/3.0/legalcode
Attribution:
The Department of Arts and Culture of the government of the Republic of South Africa (DAC), Council for Scientific and Industrial… See the full description on the dataset page: https://huggingface.co/datasets/danielshaps/nchlt_speech_zul.nchlt_speech_afr
NCHLT Speech Corpus -- Afrikaans
This is the Afrikaans language part of the NCHLT Speech Corpus of the South African languages.
Language code (ISO 639): afr
URI: https://hdl.handle.net/20.500.12185/280
Licence:
Creative Commons Attribution 3.0 Unported License (CC BY 3.0): http://creativecommons.org/licenses/by/3.0/legalcode
Attribution:
The Department of Arts and Culture of the government of the Republic of South Africa (DAC), Council for Scientific and… See the full description on the dataset page: https://huggingface.co/datasets/danielshaps/nchlt_speech_afr.Tech-Sentences-For-ASR-Training
TechVoice Dataset
Work in Progress – This dataset is actively being expanded with new recordings.
Dataset Statistics
Metric
Current
Target
Progress
Duration
38m 43s
5h 0m 0s
██░░░░░░░░░░░░░░░░░░ 12.9%
Words
10,412
50,000
████░░░░░░░░░░░░░░░░ 20.8%
Total Recordings: 205 samples
Total Characters: 74,312
A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.ClarinStudioPL
CLARIN-PL Polish Studio Corpus
The corpus was created somewhere in 2014-2015 by recording a group of few hundred volunteer speakers reading a few dozen sentences each. The total size of the corpus is ~56 hours.
Due to the manner of recording, the transcription accuracy is very high, but the manner of speech is not spontaneous. This corpus is best compared to something like TIMIT, possibly CommonVoice.
It is different from CommonVoice in that it is recorded in a controlled… See the full description on the dataset page: https://huggingface.co/datasets/danijelkorzinek/ClarinStudioPL.nchlt_speech_tso
NCHLT Speech Corpus -- Xitsonga
This is the Xitsonga language part of the NCHLT Speech Corpus of the South African languages.
Language code (ISO 639): tso
URI: https://hdl.handle.net/20.500.12185/277
Licence:
Creative Commons Attribution 3.0 Unported License (CC BY 3.0): http://creativecommons.org/licenses/by/3.0/legalcode
Attribution:
The Department of Arts and Culture of the government of the Republic of South Africa (DAC), Council for Scientific and… See the full description on the dataset page: https://huggingface.co/datasets/danielshaps/nchlt_speech_tso.nchlt_speech_eng
NCHLT Speech Corpus -- South African English
This is the South African English language part of the NCHLT Speech Corpus of the South African languages.
Language code (ISO 639): eng
URI: https://hdl.handle.net/20.500.12185/274
Licence:
Creative Commons Attribution 3.0 Unported License (CC BY 3.0): http://creativecommons.org/licenses/by/3.0/legalcode
Attribution:
The Department of Arts and Culture of the government of the Republic of South Africa (DAC), Council… See the full description on the dataset page: https://huggingface.co/datasets/danielshaps/nchlt_speech_eng.Audio-Understanding-Bitrate-Eval-0426
Audio Understanding — MP3 Bitrate Evaluation (April 2026)
Empirical eval measuring how MP3 compression bitrate affects transcription accuracy across every audio-input LLM available on OpenRouter.
📝 Blog post: MP3 Bitrate Sensitivity in Audio-Multimodal LLMs
💻 Code & methodology: github.com/danielrosehill/Audio-Understanding-Bitrate-Eval-0426
TL;DR
Ran a benchmark across 12 OpenRouter audio-multimodal models × 4 dictation samples × 5 MP3 bitrates (16/24/32/48/64 kbps)… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Audio-Understanding-Bitrate-Eval-0426.Small-STT-Eval-Audio-Dataset
Small STT Eval Audio Dataset
A small speech-to-text evaluation dataset containing 92 audio samples with ground truth transcriptions. Designed for evaluating STT systems on technical vocabulary, code-switching (English/Hebrew), and various speaking styles.
Dataset Description
This dataset contains audio recordings with accompanying transcriptions across multiple categories:
Category
Count
Description
tech_github
5
GitHub-related technical vocabulary… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Small-STT-Eval-Audio-Dataset.nchlt_speech_tsn
NCHLT Speech Corpus -- Setswana
This is the Setswana language part of the NCHLT Speech Corpus of the South African languages.
Language code (ISO 639): tsn
URI: https://hdl.handle.net/20.500.12185/281
Licence:
Creative Commons Attribution 3.0 Unported License (CC BY 3.0): http://creativecommons.org/licenses/by/3.0/legalcode
Attribution:
The Department of Arts and Culture of the government of the Republic of South Africa (DAC), Council for Scientific and… See the full description on the dataset page: https://huggingface.co/datasets/danielshaps/nchlt_speech_tsn.nchlt_speech_sot
NCHLT Speech Corpus -- Sesotho
This is the Sesotho language part of the NCHLT Speech Corpus of the South African languages.
Language code (ISO 639): sot
URI: https://hdl.handle.net/20.500.12185/278
Licence:
Creative Commons Attribution 3.0 Unported License (CC BY 3.0): http://creativecommons.org/licenses/by/3.0/legalcode
Attribution:
The Department of Arts and Culture of the government of the Republic of South Africa (DAC), Council for Scientific and Industrial… See the full description on the dataset page: https://huggingface.co/datasets/danielshaps/nchlt_speech_sot.nchlt_speech_nbl
NCHLT Speech Corpus -- isiNdebele
This is the isiNdebele language part of the NCHLT Speech Corpus of the South African languages.
Language code (ISO 639): nbl
URI: https://hdl.handle.net/20.500.12185/272
Licence:
Creative Commons Attribution 3.0 Unported License (CC BY 3.0): http://creativecommons.org/licenses/by/3.0/legalcode
Attribution:
The Department of Arts and Culture of the government of the Republic of South Africa (DAC), Council for Scientific and… See the full description on the dataset page: https://huggingface.co/datasets/danielshaps/nchlt_speech_nbl.TTS-Danish
TTS-Danish
A large-scale, high-quality Danish speech dataset for text-to-speech and automatic speech recognition.
Data Sources
This dataset combines three sources:
Source
Samples
Hours
License
Content
lydbog.com
35,719
92.0
CC-BY-SA 4.0
Danish classic literature, read by Kristoffer Hunsdahl
CoRal-TTS (Alexandra Institute)
19,996
30.5
CC0
Professional TTS recordings, 2 speakers
LibriVox
0
0.0
Public Domain
Danish audiobooks
Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Danish.nchlt_speech_nso
NCHLT Speech Corpus -- Sepedi
This is the Sepedi language part of the NCHLT Speech Corpus of the South African languages.
Language code (ISO 639): nso
URI: https://hdl.handle.net/20.500.12185/270
Licence:
Creative Commons Attribution 3.0 Unported License (CC BY 3.0): http://creativecommons.org/licenses/by/3.0/legalcode
Attribution:
The Department of Arts and Culture of the government of the Republic of South Africa (DAC), Council for Scientific and Industrial… See the full description on the dataset page: https://huggingface.co/datasets/danielshaps/nchlt_speech_nso.nchlt_speech_xho
NCHLT Speech Corpus -- isiXhosa
This is the isiXhosa language part of the NCHLT Speech Corpus of the South African languages.
Language code (ISO 639): xho
URI: https://hdl.handle.net/20.500.12185/279
Licence:
Creative Commons Attribution 3.0 Unported License (CC BY 3.0): http://creativecommons.org/licenses/by/3.0/legalcode
Attribution:
The Department of Arts and Culture of the government of the Republic of South Africa (DAC), Council for Scientific and… See the full description on the dataset page: https://huggingface.co/datasets/danielshaps/nchlt_speech_xho.bibletts-asante-twi-repaired
BibleTTS Asante Twi — Repaired Transcripts
The Asante Twi transcripts released with BibleTTS have had the
characters ɛ (U+025B) and ɔ (U+0254) stripped out. This dataset restores them.
Audio is not included. This is a drop-in replacement for the .txt files that ship with the
BibleTTS Asante Twi package, matched by clip ID.
The problem
Both are Twi vowels, and both are required by the orthography. Measured across the released
Asante Twi transcripts:
Character… See the full description on the dataset page: https://huggingface.co/datasets/danieldzikunuofmarvel/bibletts-asante-twi-repaired.nchlt_speech_ven
NCHLT Speech Corpus -- Tshivenda
This is the Tshivenda language part of the NCHLT Speech Corpus of the South African languages.
Language code (ISO 639): ven
URI: https://hdl.handle.net/20.500.12185/276
Licence:
Creative Commons Attribution 3.0 Unported License (CC BY 3.0): http://creativecommons.org/licenses/by/3.0/legalcode
Attribution:
The Department of Arts and Culture of the government of the Republic of South Africa (DAC), Council for Scientific and… See the full description on the dataset page: https://huggingface.co/datasets/danielshaps/nchlt_speech_ven.nchlt_speech_ssw
NCHLT Speech Corpus -- siSwati
This is the siSwati language part of the NCHLT Speech Corpus of the South African languages.
Language code (ISO 639): ssw
URI: https://hdl.handle.net/20.500.12185/271
Licence:
Creative Commons Attribution 3.0 Unported License (CC BY 3.0): http://creativecommons.org/licenses/by/3.0/legalcode
Attribution:
The Department of Arts and Culture of the government of the Republic of South Africa (DAC), Council for Scientific and Industrial… See the full description on the dataset page: https://huggingface.co/datasets/danielshaps/nchlt_speech_ssw.openslr-slr67See: https://www.openslr.org/67/
danish-diarization-bench
Danish Diarization Benchmark (Synthetic) — v2
A 3996-row synthetic speaker-diarization benchmark in Danish, built by mixing
single-speaker utterances from
syvai/danish-asr-unified
into multi-speaker recordings.
What changed in v2 (2026-05-18)
Per-segment text — each entry in segments now carries its text field directly. The redundant parallel texts column has been removed. Old consumers that joined segments[i] with texts[i] should switch to segments[i]["text"].
Silent… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-diarization-bench.Danish-Speech-Dataset
🎧 Danish Speech Dataset
The Danish Speech Dataset is a high-quality speech audio dataset designed to provide structured and diverse audio data for AI-driven voice technologies. It includes 168 hours of audio data across 804 files, delivered in MP3 and WAV formats, with a total size of 369 MB. This well-balanced audio dataset ensures consistent and representative voice data, with 52% female and 48% male speakers, and a broad age distribution from 18 to 50+ years. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Danish-Speech-Dataset.danube
danube
A 1,000-utterance Bambara speech sample — 0.87 hours of 16 kHz audio with transcripts and
per-utterance speaker labels. Small enough to be a working sample rather than a training
corpus.
Load
from datasets import load_dataset
ds = load_dataset("djelia/danube", split="train")
print(ds[0]["text"], ds[0]["speaker_id"], ds[0]["duration"])
One config and one split.
Config
Split
Rows
Audio
default
train
1,000
0.865 h
Fields… See the full description on the dataset page: https://huggingface.co/datasets/djelia/danube.YodaLingua-Danish
YodaLingua-Danish
YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Danish portion of the multilingual YodaLingua collection.
🧾 Dataset Overview
Property
Value
Total clips
7,871 audio–transcription pairs
Total duration
21 hours
Speakers
925 distinct speakers
Audio format
MP3 • mono • 24 kHz • 16-bit… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Danish.SND_G2P
Sindhi G2P Dataset (SND_G2P)
A Grapheme-to-Phoneme (G2P) dataset for the Sindhi language, mapping written words to their IPA (International Phonetic Alphabet) pronunciations.
Dataset Description
This dataset was compiled by scraping Wiktionary's Sindhi terms with IPA pronunciation category. For each Sindhi word, the corresponding IPA pronunciation was extracted from the word's Wiktionary entry under the Sindhi language section.
Intended use cases:
Grapheme-to-Phoneme… See the full description on the dataset page: https://huggingface.co/datasets/DanishMahdi/SND_G2P.PINC
Polish Interpreting Corpus
The Polish Interpreting Corpus (PINC) is a hand-verified parallel speech corpus derived from the European Parliament recordings. The corpus was automatically pre-processed and subsequently
manually verified to correct the transcription, word-level speech-to-text alignment and sentence-level interlingual alignment. The audio quality is decent and the annotation is fairly accurate.
The corpus contains a set of 520 recordings of Polish-English speeches… See the full description on the dataset page: https://huggingface.co/datasets/danijelkorzinek/PINC.
