datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Emilia-Dataset
Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline.
News 🔥
2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour Emilia-YODAS… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-Dataset.stt-unified-bench
am-pranav/stt-unified-bench
Private, language/locale-partitioned mini-benchmark for STT models.
Each subset is a dataset config (e.g., en, de, en_indian_accent, hi_in) with a single split val.
Audio is staged at 16 kHz and stored in-repo for reproducibility.
Schema
audio : Audio(sampling_rate=16000, decode=False)
text : reference transcription
lang : implied by dataset config name
source : upstream dataset tag
id : source-stable id
⚠️ For internal evaluation only.… See the full description on the dataset page: https://huggingface.co/datasets/am-pranav/stt-unified-bench.Emilia-NV
NVSpeech Dataset
Overview
The NVSpeech dataset provides extensive annotations of paralinguistic vocalizations for Mandarin Chinese speech, aimed at enhancing the capabilities of automatic speech recognition (ASR) and text-to-speech (TTS) systems. The dataset features explicit word-level annotations for 18 categories of paralinguistic vocalizations, including non-verbal sounds like laughter and breathing, as well as lexicalized interjections like "uhm" and "oh."… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-NV.neptel
NepTel v0.1 — Nepali real-telephony ASR benchmark
75 scored segments / 2,375 reference words of real Nepali call-center audio (genuine
two-party customer-support calls), with human-reviewed reference transcripts. To our knowledge
this is the first public Nepali ASR benchmark on real call audio rather than read-aloud speech.
This repository exists so anyone can benchmark a Nepali ASR system without any access
request: the audio is cut and ready, no gate, no approval step.… See the full description on the dataset page: https://huggingface.co/datasets/ampixa/neptel.stt-mini-bench
am-pranav/stt-mini-bench
Private, curated mini-benchmark assembled on 2025-09-04.
Note: This dataset mirrors small subsets of upstream corpora (LibriSpeech, TED-LIUM 3, VoxPopuli, Common Voice).
Check each upstream license before sharing. This repo is for internal evaluation only.
Schema
audio : Audio(sampling_rate=16000, decode=False) (files stored in repo)
text : reference transcription
lang : short language code (en, de, fr, es, it, pt)
source: upstream… See the full description on the dataset page: https://huggingface.co/datasets/am-pranav/stt-mini-bench.
