datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
uyghur-common-voice-tts
Uyghur Common Voice TTS Dataset
A cleaned and processed Text-to-Speech (TTS) dataset for the Uyghur language, derived from Mozilla Common Voice.
Dataset Summary
Property
Value
Language
Uyghur (ug)
Total Samples
43,054
Train Samples
40,901
Validation Samples
2,153
Audio Format
WAV
Source
Mozilla Common Voice
License
CC0-1.0
Dataset Structure
/
├── train.jsonl # Training data (40,901 samples)
├── val.jsonl #… See the full description on the dataset page: https://huggingface.co/datasets/anke01/uyghur-common-voice-tts.voice-code-bench
VoiceCodeBench
VoiceCodeBench is a test-only benchmark for evaluating whether automatic
speech recognition (ASR) systems preserve exact structured values in English
workplace speech.
Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition
The benchmark targets cases where a transcript is software input: callback
numbers, email addresses, command-line flags, file paths, URLs, account
identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/besimple-ai/voice-code-bench.voice-code-bench
VoiceCodeBench
VoiceCodeBench is a test-only benchmark for evaluating whether automatic
speech recognition (ASR) systems preserve exact structured values in English
workplace speech.
Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition
The benchmark targets cases where a transcript is software input: callback
numbers, email addresses, command-line flags, file paths, URLs, account
identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/yunqi1766/voice-code-bench.voices-of-civilizations
Voices of Civilizations (VoC)
Voices of Civilizations (VoC) is the first multilingual QA benchmark designed to assess audio LLMs’ cultural comprehension using full-length music recordings. VoC spans:
38 languages 🇸🇦 Arabic (ar), 🇧🇩 Bengali (bn), 🇧🇬 Bulgarian (bg), 🇨🇳 Chinese (zh), 🇭🇷 Croatian (hr), 🇨🇿 Czech (cs), 🇩🇰 Danish (da), 🇳🇱 Dutch (nl), 🇬🇧 English (en), 🇪🇪 Estonian (et), 🇫🇮 Finnish (fi), 🇫🇷 French (fr), 🇩🇪 German (de), 🇬🇷 Greek (el), 🇮🇱 Hebrew… See the full description on the dataset page: https://huggingface.co/datasets/sander-wood/voices-of-civilizations.VoiceGiraffe
VoiceGiraffe (Benchmark)
VoiceGiraffe is a benchmark for evaluating large audio language models (LALMs) on hour-level, long-context audio understanding. It contains 1,500 curated question-answer triplets over real-world recordings central to real-world long-form audio understanding — broadcast, sports/esports commentary, news, and TV drama — organized into a dual-level taxonomy of single-hop perception and multi-hop reasoning.
This repo is public and holds the annotations… See the full description on the dataset page: https://huggingface.co/datasets/Jashin-Yeah/VoiceGiraffe.Xijinping-TTS-Voicebank
习近平音源
所有声音资料来自公开影像,属于公有领域目前有 1h30m 的截取后声音,足够进行 Fine-tuning
Usage
按句截取
python -m pip install -r requirement.txt
python split.py
新增声音资料后,使用 Whisper 产生带有时间标记的 JSON 档,并手动复制到 ./voice/[FILE].json
export OPENAI_API_KEY="API_KEY_HERE"
python whisper.py ./[FILE].[AUDIO_EXTENSION]
产生 Bert-VITS2 微调所需的 esd.list 档案
python index_to_list.py
qwen3-tts-preset-voices
Qwen3-TTS preset voice embeddings
The 9 named speakers from Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice packaged as a small sidecar bundle usable with the -Base checkpoint.
bundle.safetensors — 9 × 2048-d bfloat16 rows, ~37 KB total
bundle.json — metadata (speaker name → spk_id, gender, supported languages)
Each row is lifted from talker.model.codec_embedding.weight in the CustomVoice checkpoint at the speaker-ID index from its config.json. With these rows, you can:
Deploy only… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qwen3-tts-preset-voices.PlayAI-VoiceExcited to share Play AI Voice Profile. We release 267 unique voice profiles including Israeli, Arabic, Russian, Filipino and many other exclusive voice profiles. Play AI was recently acquired by Meta which sparked our interest in releasing this dataset.
default_voices_chunked_tokenized
default_voices_chunked_tokenized
This is a gated Uzbek tokenized speech dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: uz (Uzbek)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_chunked_tokenized.SPS-Bopha-Voice-Dataset-v1
VibeVoice Fine-Tuning Dataset: SPS-Bopha-Voice-Dataset-v1
This dataset is formatted for fine-tuning VibeVoice.
Structure
training_data.jsonl: The main manifest file containing transcriptions and paths.
chunks_staging/: Directory containing the audio clips.
Usage with VibeVoice
Clone this repository:
git clone https://huggingface.co/datasets/Tnaot/SPS-Bopha-Voice-Dataset-v1
cd SPS-Bopha-Voice-Dataset-v1
Run the training script pointing to… See the full description on the dataset page: https://huggingface.co/datasets/Tnaot/SPS-Bopha-Voice-Dataset-v1.voice-dataset
Voice Dataset
Collected from the web uploader tool.
Voice Dataset
Collected from the web uploader tool.
voice-demo
Multilingual TTS demo — 10 languages of Vietnam and Cambodia
A self-contained Gradio app. Clone the folder, install the requirements, run it.
pip install -r requirements.txt
python -u app.py
Everything resolves relative to app.py, so no paths need editing.
Languages
Code
Language
Code
Language
km
Khmer
tyz
Tay-Nung
blt
Tai Dam
ium
Dao (Iu Mien)
rad
Ede
kpm
Kho
jra
Jarai
cma
Mnong
bdq
Bana
cjm
Cham
blt is Tai Dam, a Tai language of Vietnam… See the full description on the dataset page: https://huggingface.co/datasets/shadwl/voice-demo.
