datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
quranic-universal-ayahs
Qur'anic Universal Ayahs
Qur'anic Universal Audio (QUA) is a project that unifies recitations on the internet and generates timing data using forced alignment — community-verified results and constantly expanding dataset.
This dataset pairs ayah by ayah audio with word-level timestamps, DigitalKhatt letter-animation timestamps, and waqf-aware segment data. Repeated words are preserved in text_uthmani and word_timestamps, so the row reflects what the reciter… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quranic-universal-ayahs.danish-asr-unified
Danish ASR Unified Dataset
Unified Danish speech recognition dataset combining 7 sources (~3.5M samples, ~16k hours):
Source
Samples
Description
VoxPopuli
1,775,578
European Parliament recordings
ftspeech
995,677
Danish Parliament (Folketinget)
CoRal-v3 read_aloud
299,255
Read-aloud Danish speech
nst-da
182,605
NST Danish speech
CoRal-v3 conversation
147,249
Conversational Danish speech
nota
98,600
Danish broadcast media
Common Voice 17
3,484
Crowd-sourced… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-unified.danish-asr-unified-hviske-v5-tiny
danish-asr-unified — two-model labels and a quality manifest
Transcriptions, per-token confidences, and a per-row quality verdict for every
row of syvai/danish-asr-unified
(3,414,589 rows, 8 sources).
Two independently trained models labelled the whole corpus:
model
architecture
vocabulary
syvai/hviske-v5-tiny
encoder-decoder
16,384 BPE
3dio-ai/svale-110M
RNN-T (Parakeet)
44 characters
Both decode greedily. Shards mirror the source data/train-*.parquet by name… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-unified-hviske-v5-tiny.majestrino-unified-detailed-captions
Majestrino Unified Detailed Captions
Filtered subset of laion/majestrino-data containing all samples with unified_detailed_caption.
Stats
4,658,407 samples
932 tar files (~1.1 GB each)
~1,017 GB total
Format
Each tar contains paired .flac + .json files.
JSON fields:
caption — the unified detailed caption
caption_type — always unified_detailed_caption
transcription — speech transcription (when available, normalized from multiple source keys)
duration — audio… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions.UniST
UniST
This dataset contains UniST codec-token training data exported from local metadata and codec results.
We train UniSS with UniST data.
Schema
id: sample identifier
transcription: source transcription from metadata text
translation: qwen_trans, falling back to trans_text
source_glm, target_glm: GLM token lists
source_bicodec, target_bicodec: bicodec semantic token lists
bicodec_global: source bicodec global token list
dataset_name, src_lang, tgt_lang, split:… See the full description on the dataset page: https://huggingface.co/datasets/cmots/UniST.hailuo-ai-voices
Hailuo AI Voices Dataset 🎤
A curated collection of high-quality voice recordings with corresponding transcriptions and phoneme analysis. This dataset is designed for speech recognition, text-to-speech, and voice analysis tasks.
📊 Dataset Overview
The dataset provides a comprehensive collection of voice samples with the following features:
Feature
Description
Audio Files
High-quality WAV format recordings
Transcription
Accurate transcriptions of each… See the full description on the dataset page: https://huggingface.co/datasets/unlimitedbytes/hailuo-ai-voices.hailuo-ai-jokes
Hailuo AI Jokes Dataset 🎤
A curated collection of high-quality voice recordings with corresponding transcriptions and phoneme analysis. This dataset is designed for speech recognition, text-to-speech, and voice analysis tasks.
🎙️ Dataset Content
The dataset contains a diverse set of synthetic voice recordings generated by Hailuo AI Audio. The texts are sourced from a variety of public domain jokes and humorous anecdotes. Each audio sample is accompanied by the… See the full description on the dataset page: https://huggingface.co/datasets/unlimitedbytes/hailuo-ai-jokes.stt-unified-bench
am-pranav/stt-unified-bench
Private, language/locale-partitioned mini-benchmark for STT models.
Each subset is a dataset config (e.g., en, de, en_indian_accent, hi_in) with a single split val.
Audio is staged at 16 kHz and stored in-repo for reproducibility.
Schema
audio : Audio(sampling_rate=16000, decode=False)
text : reference transcription
lang : implied by dataset config name
source : upstream dataset tag
id : source-stable id
⚠️ For internal evaluation only.… See the full description on the dataset page: https://huggingface.co/datasets/am-pranav/stt-unified-bench.sna-waxal-annotated-unlabeled
Shona WAXAL annotated-unlabeled checkpoint
This is a self-contained operational checkpoint for pseudo-labeling Shona ASR
data. It contains 90,253 conservatively segmented FLAC clips
(441.585 hours), but intentionally contains no transcripts.
Fields
transcription is intentionally empty.
speaker_id is an approximate source-blind EOM cluster or unknown;
speaker_clip_count is zero for unknown assignments.
gender is always unknown; available classifiers were not… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/sna-waxal-annotated-unlabeled.transcription-corpus
UN Transcription Corpus
Two splits of UN meeting audio paired with official verbatim records.
Splits
sessions — Whole meeting sessions (SC + GA plenary)
One row per meeting. Audio from UN Web TV, verbatim records from documents.un.org.
Column
Description
symbol
UN document symbol, e.g. S/PV.9826
webtv_url
URL on UN Web TV
duration_ms
Session duration in milliseconds
num_speakers
Number of speaker turns in the verbatim record
audio_floor
Floor… See the full description on the dataset page: https://huggingface.co/datasets/united-nations/transcription-corpus.VietMed_unlabeled
unofficial mirror of VietMed (Vietnamese speech data in medical domain) unlabeled set
official announcement: https://arxiv.org/abs/2404.05659
official download: https://huggingface.co/datasets/leduckhai/VietMed
this repo contains the unlabeled set: 966h - 230k samples
i also gather the metadata: see info.csv
my extraction code: https://github.com/phineas-pta/fine-tune-whisper-vi/blob/main/misc/vietmed-unlabeled.py
need to do: check misspelling, restore foreign words phonetised to… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/VietMed_unlabeled.synthetic-speaker-diarization-dataset-fa-large-3000slovenian-speech-recognition
Slovenian Speech Dataset
Dataset comprises 10+ hours of audio recordings featuring 20+ speakers engaged in telephone dialogues in the Slovenian language. It contains speech data designed for training robust language models and automatic speech recognition systems in real-world conversational scenarios.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in natural language processing (NLP), speech recognition, and machine… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/slovenian-speech-recognition.unified-hausa-speech
Unified Hausa Speech Dataset v5
Dataset Description
A large-scale, cleaned, deduplicated, and quality-filtered Hausa speech dataset compiled from 6 open-source collections. Designed for Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) research on one of Africa's most widely spoken languages.
Hausa (ISO 639-1: ha) is a Chadic language spoken by over 80 million people across West and Central Africa — primarily in Nigeria and Niger, and as a trade language… See the full description on the dataset page: https://huggingface.co/datasets/suleiman2003/unified-hausa-speech.vietnamese-speech-recognition
Vietnamese Speech Dataset
Dataset comprises 10+ hours of telephone dialogues in Vietnamese, collected from 20 native speakers across various topics and domains. It is designed for research in speech recognition, focusing on various recognition models, primarily aimed at meeting the requirements for automatic speech recognition (ASR) systems.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in transcribing audio, and natural… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/vietnamese-speech-recognition.majestrino-unified-detailed-captions-temporal
Majestrino Unified Detailed Captions with Temporal Aspects
Filtered subset of laion/majestrino-data containing only samples with unified_detailed_caption_with_temporal_aspects.
Stats
4,128,665 samples
826 tar files (~1.1 GB each)
~878 GB total
Format
Each tar contains paired .flac + .json files.
JSON fields:
caption — the unified detailed caption with temporal aspects
caption_type — always unified_detailed_caption_with_temporal_aspects
transcription — speech… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions-temporal.uts2025_vietipa
Vietnamese IPA Dataset
A comprehensive Vietnamese IPA (International Phonetic Alphabet) dataset with word pronunciations and MP3 audio files for text-to-speech and pronunciation learning applications.
Dataset Description
Dataset Summary
This dataset contains 50 common Vietnamese words with their IPA (International Phonetic Alphabet) transcriptions and corresponding audio files. It's designed for:
Text-to-speech systems development
Vietnamese pronunciation… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/uts2025_vietipa.parczech4speech-unsegmented
ParCzech4Speech (Unsegmented Variant)
Dataset Summary
ParCzech4Speech (Unsegmented Variant) is a large-scale Czech speech dataset derived from parliamentary recordings and official transcripts.
This variant captures continuous speech segments without enforcing sentence boundaries, making it well-suited for real-world streaming ASR scenarios
and speech modeling tasks that benefit from natural discourse flow.
The dataset is created using a combination of WhisperX and… See the full description on the dataset page: https://huggingface.co/datasets/ufal/parczech4speech-unsegmented.Audio-Understanding-Bitrate-Eval-0426
Audio Understanding — MP3 Bitrate Evaluation (April 2026)
Empirical eval measuring how MP3 compression bitrate affects transcription accuracy across every audio-input LLM available on OpenRouter.
📝 Blog post: MP3 Bitrate Sensitivity in Audio-Multimodal LLMs
💻 Code & methodology: github.com/danielrosehill/Audio-Understanding-Bitrate-Eval-0426
TL;DR
Ran a benchmark across 12 OpenRouter audio-multimodal models × 4 dictation samples × 5 MP3 bitrates (16/24/32/48/64 kbps)… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Audio-Understanding-Bitrate-Eval-0426.JuzneVesti-SR-Unsloth-Format
Emilia-compatible Serbian speech (JuzneVesti-SR)
This is a format conversion of JuzneVesti-SR v1.0 for Hugging Face audio
training pipelines. It exposes the same columns as
kadirnar/Emilia-DE-B000000 and preserves the original train/dev/test split
(with dev named validation).
Source
Peter Rupnik and Nikola Ljubesic, ASR training dataset for Serbian
JuzneVesti-SR v1.0, Jozef Stefan Institute / CLARIN.SI (2022).
Persistent identifier:… See the full description on the dataset page: https://huggingface.co/datasets/baki83/JuzneVesti-SR-Unsloth-Format.una-fraza-al-diya
Una fraza al diya
Ladino language learning sentences prepared by Karen Sarhon of Sephardic Center of Istanbul. Each sentence has translations in Turkish, English, Spanish. Includes audio and image. 307 sentences in total.
Source: https://sefarad.com.tr/judeo-espanyolladino/frazadeldia/
Citation
If you use this dataset, please cite:
Preparing an Endangered Language for the Digital Age: The Case of Judeo-Spanish
Preparing an endangered language for the digital age: The… See the full description on the dataset page: https://huggingface.co/datasets/collectivat/una-fraza-al-diya.human-robot-conversation-korean
Human-Robot Dataset
The dataset comprises 660+ hours of audio recordings across 20,000+ files for human-robot interactions in the Korean language. It captures authentic dialogues between humans and artificial conversational agents, specifically designed for training language models and advancing speech recognition systems.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in robotic systems and conversational AI technologies.… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-korean.human-robot-conversation-german
Human-Robot Dataset
The dataset comprises 660+ hours of audio recordings across 20,000+ files for human-robot interactions in the German language. It captures authentic dialogues between humans and artificial conversational agents, specifically designed for training language models and advancing speech recognition systems.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in speech recognition, natural language processing, and… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-german.korean-speech-recognition
Korean Speech Dataset
Dataset comprises 10+ hours of audio recordings from 20+ speakers, featuring telephone-quality speech data from native korean speakers. It provides a diverse collection of spoken language for automatic speech recognition tasks and serves as essential training data for model training in NLP and speech detection research.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in automatic speech recognition… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/korean-speech-recognition.human-robot-conversation-english
Human-Robot Dataset
The dataset comprises 660+ hours of English speech across 20,000+ audio files featuring human-robot interactions between AI and humans. It is designed for research in conversational agents, focusing on various speech recognition methods, primarily aimed at advancing language models and machine learning applications.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in robotic systems and conversational AI… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-english.british-english-speech-recognition-dataset
British English Speech Dataset for recognition task
Dataset comprises 200 hours of high-quality audio recordings featuring 310 speakers, achieving an impressive 95% Sentence Accuracy Rate. This extensive collection of speech data is designed for NLP tasks such as speech recognition, dialogue systems, and language understanding.
By utilizing this dataset, developers and researchers can advance their work in automatic speech recognition and improve recognition systems. - Get the… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/british-english-speech-recognition-dataset.UniDic-tdmelodic
tdmelodic Pre-computed Accents Dataset
This repository provides pre-computed, inference-ready CSV files generated by tdmelodic (Tokyo Dialect MELOdic accent DICtionary generator) mapping over the NEologd vocabulary.
Generating these files locally requires running neural network inference (tdmelodic-convert), which typically takes several hours to complete depending on the hardware. We have pre-generated these dictionary files and made them available here to eliminate the setup… See the full description on the dataset page: https://huggingface.co/datasets/aipracticecafe/UniDic-tdmelodic.hindi-speech-recognition-dataset
Hindi Speech Dataset for recognition task
Dataset comprises 760 hours of telephone dialogues in Hindi, collected from 1,000+ native speakers across various topics and domains. This dataset boasts an impressive 95% sentence accuracy rate, making it a valuable resource for advancing speech recognition technology.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in automatic speech recognition (ASR) systems, transcribing audio… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/hindi-speech-recognition-dataset.french-speech-recognition-dataset
French Speech Dataset for recognition task
Dataset comprises 547 hours of telephone dialogues in French, collected from 964 native speakers across various topics and domains, with an impressive 98% Word Accuracy Rate. It is designed for research in speech recognition, focusing on various recognition models, primarily aimed at meeting the requirements for automatic speech recognition (ASR) systems.
By utilizing this dataset, researchers and developers can advance their… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/french-speech-recognition-dataset.german-speech-recognition-dataset
German Speech Dataset for recognition task
Dataset comprises 431 hours of telephone dialogues in German, collected from 590+ native speakers across various topics and domains, achieving an impressive 95% sentence accuracy rate. It is designed for research in automatic speech recognition (ASR) systems.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in transcribing audio, and natural language processing (NLP). - Get the data… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/german-speech-recognition-dataset.
