datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Multitask-National-Speech-Corpus-v1-extendchinese_speech_sft_cosy_audioThis dataset use Cosy-voice 2 0.5B to generate speech audio from https://huggingface.co/datasets/TigerResearch/sft_zh.
https://huggingface.co/datasets/JerryAGENDD/voiceprint_librispeech_other_test is used as audio input prompt.
german-golden-audio_speech-IPA
🌟 German Golden Speech & IPA Corpus (FLEURS + Multilingual TEDx)
An ultra-clean, high-standard curated German speech dataset combining Google FLEURS (de_de) and Multilingual TEDx German (mTEDx), fully embedded with 16kHz WAV audio bytes, normalized orthographic text, and pre-computed International Phonetic Alphabet (IPA) transcriptions.
📊 Dataset Summary
Total Samples: 1,354 high-quality audio recordings.
Total Size: ~419 MB (Compressed Parquet format).
Audio… See the full description on the dataset page: https://huggingface.co/datasets/q1805/german-golden-audio_speech-IPA.parlertts-pony-speech-audiopeoples_speech_test@article{galvez2021people,
title={The people's speech: A large-scale diverse english speech recognition dataset for commercial usage},
author={Galvez, Daniel and Diamos, Greg and Ciro, Juan and Cer{\'o}n, Juan Felipe and Achorn, Keith and Gopi, Anjali and Kanter, David and Lam, Maximilian and Mazumder, Mark and Reddi, Vijay Janapa},
journal={arXiv preprint arXiv:2111.09344},
year={2021}
}
@article{wang2024audiobench,
title={AudioBench: A Universal Benchmark for Audio Large Language… See the full description on the dataset page: https://huggingface.co/datasets/AudioLLMs/peoples_speech_test.peoples_speechspeech_text-tts_audioemotional-speech-audio-dataset-3eng-4noneng-updatedpublic_sg_speech_qa_test@article{wang2024audiobench,
title={AudioBench: A Universal Benchmark for Audio Large Language Models},
author={Wang, Bin and Zou, Xunlong and Lin, Geyu and Sun, Shuo and Liu, Zhuohan and Zhang, Wenyu and Liu, Zhengyuan and Aw, AiTi and Chen, Nancy F},
journal={NAACL},
year={2025}
}
chinese_speech_cosy_audioinstruction-speech-no-audio-v1.5emotional-speech-audio-dataset-3eng-4nonengchinese_speech_alpaca_cosy_audioThis dataset is based on https://huggingface.co/datasets/shibing624/alpaca-zh and use Cosyvoice2 to generate speeches.
Hindi-audio-speechparlertts-pony-speech-audioaudio-speech-realtime-voice-agents-2026
🎙️ Audio, Speech Foundation Models & Real-Time Voice Agents Dataset (2026 Edition)
A structured research dataset featuring 1,722 domain-verified research papers and 298 official code repositories focused on Full-Duplex Speech-to-Speech LLMs, Real-Time Voice Agents (<200ms Latency), Zero-Shot TTS, Voice Cloning, OpenAI Whisper-v3, Neural Audio Codecs (EnCodec/DAC/SNAC), and Generative Music (2023–2026).
Built with Universal Scientific Engine V18.1 Diamond, providing 48 schema… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/audio-speech-realtime-voice-agents-2026.emotional-speech-audio-dataset-4languages-reformattedtr_audio_data_speechemotional-speech-audio-dataset-4languageswelsh-speech-audio
Welsh Speech Dataset - Audio
Audio recordings from the Welsh Speech Dataset.
Contents
33 speakers x 10 Welsh phrases ~ 330 audio files
Format: WAV (16-bit PCM recommended)
with 3D facial captures and landmarks
Files
Audio files are located in the audio/ directory.
Naming: audio/speaker_XX_phrase_YY.wav
Example: audio/speaker_01_phrase_05.wav = Speaker 1 speaking Phrase 5 ("Ardderchog")
Metadata
Metadata for the audio dataset is in metadata.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/arvinsingh/welsh-speech-audio.ravdess_emotional_speech_audioinstruction-speech-no-audio-v1speech-audio
运行顺序
只执行 bash run.sh。脚本按下面的顺序走。某一步打印 STOP 时,进程以非 0 退出,后面的步骤不会开始。
0. 环境
bash run.sh
依赖由脚本安装:torch、torchaudio、transformers。不需要指定 GPU。有 CUDA 时特征提取会用到全部可见卡;没有 CUDA 时仍会跑,只是更慢。
1. 下载
脚本第一次读数据时自动下载,不要改保存位置。
Speech Commands v0.02,约 2.3 GBhttp://download.tensorflow.org/data/speech_commands_v0.02.tar.gz目录:data/speech_commands_v0.02/
编码器 openai/whisper-base,从 Hugging Face 拉取。该模型是公开权重,不需要令牌。
2. 冒烟(256 条训练,128 条测试,1 个 epoch)… See the full description on the dataset page: https://huggingface.co/datasets/jamie33/speech-audio.indian-speech-audio-extendedEgyptian-Speech-Audio-Text-1kspeech_attempt_FC_no_audioJA_audio_EN_text_speech_bsd_20kindian_speech_audioemotional-speech-audioaudiobook_chunked_speech_restorised
audiobook_chunked_speech_restorised
This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audiobook_chunked_speech_restorised.
