datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
QuranTTS
QuranTTS v4
An ear-verified Quranic recitation corpus for TTS and speech restoration.
Built in-house for the lab's own training runs. We have since moved on to a larger corpus and a newer
pipeline, so this one is published rather than shelved. What you get is the corpus exactly as it stood when
we stopped using it: complete, documented, and unmaintained.
Non-commercial, strictly. Neither this corpus nor any model trained on it may be used for any
commercial purpose, and… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/QuranTTS.fma-labeled
FMA Labeled — Multi-Attribute Music Dataset
🏆 Submitted to the Uncharted Data Challenge
hosted by Adaption Labs — credit to
Adaptive Data by Adaption for organizing the hackathon.
A large-scale labeled music dataset built on top of the Creative-Commons
subset of the Free Music Archive (FMA). Every
track has been automatically annotated with lyrics, genre, mood, instruments,
tempo, key, and more using Google Gemini (gemini-flash-latest).
Intended for training and evaluating music… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/fma-labeled.Easy-Turn-Testset
Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems
Guojian Li1, Chengyou Wang1, Hongfei Xue1,
Shuiyuan Wang1, Dehui Gao1, Zihan Zhang2,
Yuke Lin2, Wenjie Li2, Longshuai Xiao2,
Zhonghua Fu1,╀, Lei Xie1,╀
1 Audio, Speech and Language Processing Group (ASLP@NPU), Northwestern Polytechnical University
2 Huawei Technologies, China
🎤 Demo Page
🤖 Easy Turn Model
📑 Paper
🌐 Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/Easy-Turn-Testset.BERSt
BERSt Dataset
We release the BERSt Dataset for various speech recognition tasks including Automatic Speech Recognition (ASR) and Speech Emotion Recogniton (SER)
Read the paper here
Overview
4526 single phrase recordings (~3.75h)
98 professional actors
19 phone positions
7 emotion classes
3 vocal intensity levels
varied regional and non-native English accents
nonsense phrases covering all English Phonemes
Data collection
The BERSt dataset represents data… See the full description on the dataset page: https://huggingface.co/datasets/Rosie-Lab/BERSt.Easy-Turn-Trainset
Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems
Guojian Li1, Chengyou Wang1, Hongfei Xue1,
Shuiyuan Wang1, Dehui Gao1, Zihan Zhang2,
Yuke Lin2, Wenjie Li2, Longshuai Xiao2,
Zhonghua Fu1,╀, Lei Xie1,╀
1 Audio, Speech and Language Processing Group (ASLP@NPU), Northwestern Polytechnical University
2 Huawei Technologies, China
🎤 Demo Page
🤖 Easy Turn Model
📑 Paper
🌐 Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/Easy-Turn-Trainset.dutch-tts-labeled-complete
Dutch TTS Dataset - Complete Labeled
A comprehensive Dutch text-to-speech dataset with 596,508 audio samples totaling 234GB of audio data.
Quick Preview
The default config shows a 100-row sample for the dataset viewer. To access the full dataset, use the full config.
Dataset Description
This dataset contains Dutch speech recordings with rich metadata including:
Emotion labels (neutral, happy, sad, angry)
Speaker IDs (239,388 unique speakers)… See the full description on the dataset page: https://huggingface.co/datasets/AITRADER/dutch-tts-labeled-complete.samromur_childrenThe Samrómur Children corpus contains more than 137000 validated speech-recordings uttered by Icelandic children.quran-tajweed-phonetics
The complete phonetic layer of the Quran in the riwaya of Hafs 'an
'Asim via tariq al-Shatibiyyah: 6,236 ayat, 522,475 phones, every
phone carrying its tajweed attribution: madd class with its transmitted
length range, ghunna grade, qalqalah class, tafkheem with its rank, sakt,
the seventeen sifat, and the rule that produced it.
Built and maintained by Quran Lab, a waqf building open technology in
the service of the Quran.
How it was built and verified
Indexed from the… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/quran-tajweed-phonetics.quranic-asr-benchmark
Quranic ASR Benchmark - leakage-free, held-out
A small, leakage-free benchmark (600 clips) for evaluating Arabic ASR on Quranic recitation
(Hafs riwayah). Every clip is verified absent from our training data, so it measures
generalization, not memorization. Same clips + same scoring for every model.
📊 Live leaderboard: https://huggingface.co/spaces/Muno459/quranic-asr-leaderboard
The set (600 clips, 200 per source)
Source
n
What it is
everyayah_heldout… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/quranic-asr-benchmark.tartanaviation-atc-labels
TartanAviation ATC ASR Labels
Machine transcripts and confidence scores for
twangodev/tartanaviation-atc-adsb-utterances.
A 1:1 labels-only add-on (no audio): one row per source utterance, same shards and row order, keyed
by utterance_id.
531,050 labels · 184 shards · 100% coverage · ensemble ASR + weighted ROVER + ADS-B
callsign snap · ~326 human-reviewed. Built with readback.
Usage
Join 1:1 onto the source. Rows are aligned and in the same order:
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/tartanaviation-atc-labels.shona-waxal-pseudo-labeled
Shona WAXAL pseudo-labelled speech
This release contains 90,253 Shona speech clips, totalling 441.585 hours. Each clip keeps its original FLAC audio and a Sunbird Whisper pseudo-transcript. These are model outputs, not human reference transcriptions.
What this release contains
The source is the unlabeled Shona ASR split from WAXAL NLP, preserved in the operational checkpoint manassehzw/sna-waxal-annotated-unlabeled. The source checkpoint has no transcripts. This… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/shona-waxal-pseudo-labeled.samromur_milljonSamrómur Milljón consists of approximately 1 million of speech recordings (967 hours) collected through the platform samromur.is; the transcripts accompanying these recordings were automatically verified using various ASR systems such as: Wav2Vec, Whisper and NeMo.APAC-Egocentric-Residential-Voiceover
APAC Egocentric Residential (with Voiceover)
Ten narrated first-person recordings of household chores, each shipping the original capture with spoken voiceover, a burned-in caption render, WebVTT captions, and an ASS annotation track.
This is the only release in the HumynLabs egocentric collection that carries audio narration — the wearer describes each action as they perform it, and the captions align that speech to the video.
Preview: 45 s from the cooking sample, captioned… See the full description on the dataset page: https://huggingface.co/datasets/humyn-labs/APAC-Egocentric-Residential-Voiceover.openstt_balalaika
OpenSTT Annotated by Balalaika
[!IMPORTANT]
Official dataset for our INTERSPEECH 2026 paper
"A Data-Centric Framework for Addressing Phonetic and Prosodic Challenges in Russian Speech Generative Models" (arXiv:2507.13563).
Part of the Balalaika Russian speech data-processing pipeline — code: https://github.com/lab260ru/balalaika.
If you use this resource, please cite it.
A curated Russian speech dataset for advanced speech generative tasks.
Overview
OpenSTT… See the full description on the dataset page: https://huggingface.co/datasets/lab260/openstt_balalaika.fr-bj-speech-pilot
fr_bj, Beninese French read-speech pilot
Read speech in Beninese French, in the FLEURS format. The sentences were
written in Benin, about local realities, and read by Beninese speakers in their
own French. For evaluation, not for training.
See DATASHEET.md for provenance and intended use.
Segments
209
Sentences
105, all covered
Speakers
2 (1 female, 1 male)
Duration
25.0 min
Words
2921
Audio
WAV PCM 16-bit, 16 kHz, mono
Split
single test split… See the full description on the dataset page: https://huggingface.co/datasets/labari-voice/fr-bj-speech-pilot.espeech_balalaika
ESpeech datasets (w/o podcasts) Annotated by Balalaika
[!IMPORTANT]
Official dataset for our INTERSPEECH 2026 paper
"A Data-Centric Framework for Addressing Phonetic and Prosodic Challenges in Russian Speech Generative Models" (arXiv:2507.13563).
Part of the Balalaika Russian speech data-processing pipeline — code: https://github.com/lab260ru/balalaika.
If you use this resource, please cite it.
A curated Russian speech dataset for advanced speech generative tasks.… See the full description on the dataset page: https://huggingface.co/datasets/lab260/espeech_balalaika.fr-sn-speech-pilot
fr_sn, Senegalese French read-speech pilot
Read speech in Senegalese French, in the FLEURS format. The sentences were
written in Senegal, about local realities, and read by Senegalese speakers in
their own French. For evaluation, not for training.
See DATASHEET.md for provenance and intended use.
Segments
210
Sentences
105, all covered
Speakers
5 (3 female, 2 male)
Duration
19.3 min
Words
2832
Audio
WAV PCM 16-bit, 16 kHz, mono
Split
single test split… See the full description on the dataset page: https://huggingface.co/datasets/labari-voice/fr-sn-speech-pilot.raddromur_asr
Dataset Card for raddromur_asr
Dataset Summary
The "Raddrómur Icelandic Speech 22.09" ("Raddrómur Corpus" for short) is an Icelandic corpus created by the Language and Voice Laboratory (LVL) at Reykjavík University (RU) in 2022. It is made out of radio podcasts mostly taken from RÚV (ruv.is).
Example Usage
The Raddrómur Corpus counts with the train split only. To load the training split pass its name as a config name:
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/language-and-voice-lab/raddromur_asr.samromur_asrSamrómur Icelandic Speech 1.0.Asian-High-Fidelity-ASR-Dataset
Dataset Overview
This dataset contains high-quality conversational audio samples curated for Automatic Speech Recognition tasks in Vietnamese, Korean, Arabic and Filipino.
The dataset includes:
Paired audio + transcripts
Natural, non-scripted conversational speech
Single-speaker & Dual-speaker interactions
Audio Specifications
Sampling Rate: 16 kHz – 24 kHz
Bit Depth: 16-bit
Audio Type: Non-scripted conversational speech
Supported Languages
Language… See the full description on the dataset page: https://huggingface.co/datasets/humyn-labs/Asian-High-Fidelity-ASR-Dataset.librispeech-phoneme-labels
LibriSpeech IPA Phoneme Labels
This repository provides IPA-based phoneme annotations and lexicon for the LibriSpeech dataset.
All phoneme labels are converted from CMU Pronouncing Dictionary (CMU-Dict) phonemes into IPA symbols using deterministic rules, with the help of the following toolkit:
https://pypi.org/project/pinyin-to-ipa
The data is intended for phoneme-based ASR, P2G/G2P research, phoneme CTC / AED models, and cross-lingual phoneme experiments.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/THU-SPMI/librispeech-phoneme-labels.feasibility-cn-en
Feasibility Chinese → English
Full CoVoST 2 zh-CN → en subset, packaged for diffusion-speech-recognition.
Split
Examples
Audio hours
train
7,085
10.44
validation
4,843
7.90
test
4,898
8.24
Schema
id: string, unique WAV filename; equals audio.path for the project's ID-to-audio mapping.
chinese: original Chinese transcript, preserved without normalization.
english: official English translation, preserved without normalization.
audio: Hugging… See the full description on the dataset page: https://huggingface.co/datasets/aiai-laboratory/feasibility-cn-en.althingi_asrAlthingi Parliamentary Speech consists of approximately 542 hours of recorded speech from Althingi, the Icelandic Parliament. Speeches date from 2005-2016.product-bench
Product Bench
31-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a laptop comparison shopping assistant.
Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs.
Leaderboard | GitHub | All Benchmarks
Dataset Description
The model acts as a laptop comparison shopping assistant helping a customer evaluate, compare, and order laptops. The conversation features multi-intent turns… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/product-bench.Indic-High-Fidelity-MultiSpeaker-ASR
Dataset Overview
This dataset contains high-quality multi-speaker conversational audio recordings curated for Automatic Speech Recognition (ASR) research across multiple Indic languages.
The dataset includes:
Paired audio + timestamped transcripts
Natural, non-scripted conversational speech
Dual-speaker interactions
Segment-level speaker annotations
Regionally diverse accents
Audio Specifications
Format: WAV (PCM 16-bit)
Sampling Rate: 16 kHz
Channel: Mono
Speech… See the full description on the dataset page: https://huggingface.co/datasets/humyn-labs/Indic-High-Fidelity-MultiSpeaker-ASR.grocery-bench
Grocery Bench
30-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a grocery ordering assistant.
Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs.
Leaderboard | GitHub | All Benchmarks
Dataset Description
The model acts as a grocery ordering assistant helping a customer build, modify, and finalize an order. The conversation is designed around 15 difficulty enhancements that… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/grocery-bench.WenetSpeech-Wu📢:Good news! 21,800 hours of multi-label Cantonese speech data and 10,000 hours of multi-label Chuan-Yu speech data are also available at ⭐WenetSpeech-Yue⭐ and ⭐WenetSpeech-Chuan⭐.
WenetSpeech-Wu: Datasets, Benchmarks, and Models for a Unified Chinese Wu Dialect Speech Processing Ecosystem
Chengyou Wang1*,
Mingchen Shao1*,
Jingbin Hu1*,
Zeyu Zhu1*,
Hongfei Xue1,
Bingshen Mu1,
Xin Xu2,
Xingyi Duan6,
Binbin Zhang3,
Pengcheng Zhu3,
Chuang Ding4… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/WenetSpeech-Wu.VietMed_labeled
unofficial mirror of VietMed (Vietnamese speech data in medical domain) labeled set
official announcement: https://arxiv.org/abs/2404.05659
official download: https://huggingface.co/datasets/leduckhai/VietMed
this repo contains the labeled set: 9.2k samples
i also gather the metadata: see info.csv
my extraction code: https://github.com/phineas-pta/fine-tune-whisper-vi/blob/main/misc/vietmed-labeled.py
need to do: check misspelling, restore foreign words phonetised to vietnamese… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/VietMed_labeled.appointment-bench
Appointment Bench
25-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a dental office receptionist handling appointment scheduling.
Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs.
Leaderboard | GitHub | All Benchmarks
Dataset Description
The model acts as a dental office receptionist scheduling appointments for two patients with confusable names (Daniel and Danielle Nolan)… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/appointment-bench.golos_balalaika
GOLOS Annotated by Balalaika
[!IMPORTANT]
Official dataset for our INTERSPEECH 2026 paper
"A Data-Centric Framework for Addressing Phonetic and Prosodic Challenges in Russian Speech Generative Models" (arXiv:2507.13563).
Part of the Balalaika Russian speech data-processing pipeline — code: https://github.com/lab260ru/balalaika.
If you use this resource, please cite it.
A curated Russian speech dataset for advanced speech generative tasks.
Overview
GOLOS… See the full description on the dataset page: https://huggingface.co/datasets/lab260/golos_balalaika.
