datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Emilia-Dataset
Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline.
News 🔥
2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour Emilia-YODAS… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-Dataset.stt-unified-bench
am-pranav/stt-unified-bench
Private, language/locale-partitioned mini-benchmark for STT models.
Each subset is a dataset config (e.g., en, de, en_indian_accent, hi_in) with a single split val.
Audio is staged at 16 kHz and stored in-repo for reproducibility.
Schema
audio : Audio(sampling_rate=16000, decode=False)
text : reference transcription
lang : implied by dataset config name
source : upstream dataset tag
id : source-stable id
⚠️ For internal evaluation only.… See the full description on the dataset page: https://huggingface.co/datasets/am-pranav/stt-unified-bench.Emilia-NV
NVSpeech Dataset
Overview
The NVSpeech dataset provides extensive annotations of paralinguistic vocalizations for Mandarin Chinese speech, aimed at enhancing the capabilities of automatic speech recognition (ASR) and text-to-speech (TTS) systems. The dataset features explicit word-level annotations for 18 categories of paralinguistic vocalizations, including non-verbal sounds like laughter and breathing, as well as lexicalized interjections like "uhm" and "oh."… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-NV.neptts-bench
NepTTS-Bench Dataset
First comprehensive benchmark for evaluating Nepali text-to-speech systems.
Contents
sentences.json — 365 phonologically-designed Nepali sentences with metadata
audio/ — TTS outputs from 12 systems (~2000 audio files)
results/ — Pre-computed evaluation results (SCOREQ, Chirp2, MMS, XLS-R, Whisper)
baselines.json — Aggregate scores for all baseline systems
model/ — NepaliMOS predictor checkpoint (Spearman 0.587)
Systems Evaluated… See the full description on the dataset page: https://huggingface.co/datasets/ampixa/neptts-bench.neptel
NepTel v0.1 — Nepali real-telephony ASR benchmark
75 scored segments / 2,375 reference words of real Nepali call-center audio (genuine
two-party customer-support calls), with human-reviewed reference transcripts. To our knowledge
this is the first public Nepali ASR benchmark on real call audio rather than read-aloud speech.
This repository exists so anyone can benchmark a Nepali ASR system without any access
request: the audio is cut and ready, no gate, no approval step.… See the full description on the dataset page: https://huggingface.co/datasets/ampixa/neptel.Debatts-Data
Debatts-Data: The First Madarin Rebuttal Speech Dataset for Expressive Text-to-Speech Synthesis
The Debatts-Data dataset is the first Madarin rebuttal speech dataset for expressive text-to-speech synthesis. It is constructed from a vast collection of professional Madarin speech data sourced from diverse video platforms and podcasts on the Internet. The in-the-wild collection approach ensures the real and natural rebuttal speech. In addition, the dataset contains annotations of… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Debatts-Data.PicoAudioThe dataset utilized in PicoAudio
Amphion-TTS-EvalINTP
INTP: Intelligibility Preference Speech Dataset
We establish a synthetic Intelligibility Preference Speech Dataset (INTP), including about 250K preference pairs (over 2K hours) of diverse domains.
Features
The dataset exhibits the following distinctive features:
Multi-Scenario Coverage
The dataset encompasses various scenarios including regular speech, repeated phrases, code-switching contexts, and cross-lingual synthesis.
Diverse TTS Model Integration… See the full description on the dataset page: https://huggingface.co/datasets/amphion/INTP.stt-mini-bench
am-pranav/stt-mini-bench
Private, curated mini-benchmark assembled on 2025-09-04.
Note: This dataset mirrors small subsets of upstream corpora (LibriSpeech, TED-LIUM 3, VoxPopuli, Common Voice).
Check each upstream license before sharing. This repo is for internal evaluation only.
Schema
audio : Audio(sampling_rate=16000, decode=False) (files stored in repo)
text : reference transcription
lang : short language code (en, de, fr, es, it, pt)
source: upstream… See the full description on the dataset page: https://huggingface.co/datasets/am-pranav/stt-mini-bench.
