datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
voice-code-bench
VoiceCodeBench
VoiceCodeBench is a test-only benchmark for evaluating whether automatic
speech recognition (ASR) systems preserve exact structured values in English
workplace speech.
Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition
The benchmark targets cases where a transcript is software input: callback
numbers, email addresses, command-line flags, file paths, URLs, account
identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/besimple-ai/voice-code-bench.voice-code-bench
VoiceCodeBench
VoiceCodeBench is a test-only benchmark for evaluating whether automatic
speech recognition (ASR) systems preserve exact structured values in English
workplace speech.
Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition
The benchmark targets cases where a transcript is software input: callback
numbers, email addresses, command-line flags, file paths, URLs, account
identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/yunqi1766/voice-code-bench.Xijinping-TTS-Voicebank
习近平音源
所有声音资料来自公开影像,属于公有领域目前有 1h30m 的截取后声音,足够进行 Fine-tuning
Usage
按句截取
python -m pip install -r requirement.txt
python split.py
新增声音资料后,使用 Whisper 产生带有时间标记的 JSON 档,并手动复制到 ./voice/[FILE].json
export OPENAI_API_KEY="API_KEY_HERE"
python whisper.py ./[FILE].[AUDIO_EXTENSION]
产生 Bert-VITS2 微调所需的 esd.list 档案
python index_to_list.py
voice-dataset
Claudia Voice Dataset
Training dataset for the Claudia persona — a direct, honest, emotionally present AI companion voice. These are regenerated multi-turn conversations in ChatML format capturing the full range of Claudia's personality.
Dataset Overview
Total conversations: 2026
Format: ChatML (system/user/assistant message arrays)
Splits: Train (1823) / Validation (203)
Source: Regenerated conversations from original Claudia sessions
Categories
Each… See the full description on the dataset page: https://huggingface.co/datasets/claudiapersists/voice-dataset.PlayAI-VoiceExcited to share Play AI Voice Profile. We release 267 unique voice profiles including Israeli, Arabic, Russian, Filipino and many other exclusive voice profiles. Play AI was recently acquired by Meta which sparked our interest in releasing this dataset.
default_voices_chunked_tokenized
default_voices_chunked_tokenized
This is a gated Uzbek tokenized speech dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: uz (Uzbek)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_chunked_tokenized.voice-demo
Multilingual TTS demo — 10 languages of Vietnam and Cambodia
A self-contained Gradio app. Clone the folder, install the requirements, run it.
pip install -r requirements.txt
python -u app.py
Everything resolves relative to app.py, so no paths need editing.
Languages
Code
Language
Code
Language
km
Khmer
tyz
Tay-Nung
blt
Tai Dam
ium
Dao (Iu Mien)
rad
Ede
kpm
Kho
jra
Jarai
cma
Mnong
bdq
Bana
cjm
Cham
blt is Tai Dam, a Tai language of Vietnam… See the full description on the dataset page: https://huggingface.co/datasets/shadwl/voice-demo.voicesliberty-echo-voice-assetsvoice-rag-indexvoice-rag-index
