CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01anke01 /uyghur-common-voice-tts Uyghur Common Voice TTS Dataset A cleaned and processed Text-to-Speech (TTS) dataset for the Uyghur language, derived from Mozilla Common Voice. Dataset Summary Property Value Language Uyghur (ug) Total Samples 43,054 Train Samples 40,901 Validation Samples 2,153 Audio Format WAV Source Mozilla Common Voice License CC0-1.0 Dataset Structure / ├── train.jsonl # Training data (40,901 samples) ├── val.jsonl #… See the full description on the dataset page: https://huggingface.co/datasets/anke01/uyghur-common-voice-tts.audiotext-to-speech10K<n<100K0 likes3.9k downloads7mo agoHugging Face02besimple-ai /voice-code-bench VoiceCodeBench VoiceCodeBench is a test-only benchmark for evaluating whether automatic speech recognition (ASR) systems preserve exact structured values in English workplace speech. Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition The benchmark targets cases where a transcript is software input: callback numbers, email addresses, command-line flags, file paths, URLs, account identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/besimple-ai/voice-code-bench.audioautomatic-speech-recognitionn<1K14 likes728 downloads14d agoHugging Face03yunqi1766 /voice-code-bench VoiceCodeBench VoiceCodeBench is a test-only benchmark for evaluating whether automatic speech recognition (ASR) systems preserve exact structured values in English workplace speech. Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition The benchmark targets cases where a transcript is software input: callback numbers, email addresses, command-line flags, file paths, URLs, account identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/yunqi1766/voice-code-bench.audioautomatic-speech-recognitionn<1K1 likes126 downloads2mo agoHugging Face04sander-wood /voices-of-civilizations Voices of Civilizations (VoC) Voices of Civilizations (VoC) is the first multilingual QA benchmark designed to assess audio LLMs’ cultural comprehension using full-length music recordings. VoC spans: 38 languages 🇸🇦 Arabic (ar), 🇧🇩 Bengali (bn), 🇧🇬 Bulgarian (bg), 🇨🇳 Chinese (zh), 🇭🇷 Croatian (hr), 🇨🇿 Czech (cs), 🇩🇰 Danish (da), 🇳🇱 Dutch (nl), 🇬🇧 English (en), 🇪🇪 Estonian (et), 🇫🇮 Finnish (fi), 🇫🇷 French (fr), 🇩🇪 German (de), 🇬🇷 Greek (el), 🇮🇱 Hebrew… See the full description on the dataset page: https://huggingface.co/datasets/sander-wood/voices-of-civilizations.textquestion-answeringn<1K1 likes62 downloads1y agoHugging Face05Jashin-Yeah /VoiceGiraffe VoiceGiraffe (Benchmark) VoiceGiraffe is a benchmark for evaluating large audio language models (LALMs) on hour-level, long-context audio understanding. It contains 1,500 curated question-answer triplets over real-world recordings central to real-world long-form audio understanding — broadcast, sports/esports commentary, news, and TV drama — organized into a dual-level taxonomy of single-hop perception and multi-hop reasoning. This repo is public and holds the annotations… See the full description on the dataset page: https://huggingface.co/datasets/Jashin-Yeah/VoiceGiraffe.textaudio-classification1K<n<10K0 likes60 downloads2mo agoHugging Face06KritiAI /Xijinping-TTS-Voicebank 习近平音源 所有声音资料来自公开影像,属于公有领域目前有 1h30m 的截取后声音,足够进行 Fine-tuning Usage 按句截取 python -m pip install -r requirement.txt python split.py 新增声音资料后,使用 Whisper 产生带有时间标记的 JSON 档,并手动复制到 ./voice/[FILE].json export OPENAI_API_KEY="API_KEY_HERE" python whisper.py ./[FILE].[AUDIO_EXTENSION] 产生 Bert-VITS2 微调所需的 esd.list 档案 python index_to_list.py audiotext-to-speech10K<n<100K4 likes42 downloads1y agoHugging Face07malaiwah /qwen3-tts-preset-voices Qwen3-TTS preset voice embeddings The 9 named speakers from Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice packaged as a small sidecar bundle usable with the -Base checkpoint. bundle.safetensors — 9 × 2048-d bfloat16 rows, ~37 KB total bundle.json — metadata (speaker name → spk_id, gender, supported languages) Each row is lifted from talker.model.codec_embedding.weight in the CustomVoice checkpoint at the speaker-ID index from its config.json. With these rows, you can: Deploy only… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qwen3-tts-preset-voices.textn<1K0 likes19 downloads5mo agoHugging Face08sleeping-ai /PlayAI-VoiceExcited to share Play AI Voice Profile. We release 267 unique voice profiles including Israeli, Arabic, Russian, Filipino and many other exclusive voice profiles. Play AI was recently acquired by Meta which sparked our interest in releasing this dataset. tabularn<1K0 likes17 downloads1y agoHugging Face09instinct-org /default_voices_chunked_tokenizedgated default_voices_chunked_tokenized This is a gated Uzbek tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: uz (Uzbek) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_chunked_tokenized.tabulartext-to-speech100K<n<1M0 likes14 downloads4mo agoHugging Face10Tnaot /SPS-Bopha-Voice-Dataset-v1gated VibeVoice Fine-Tuning Dataset: SPS-Bopha-Voice-Dataset-v1 This dataset is formatted for fine-tuning VibeVoice. Structure training_data.jsonl: The main manifest file containing transcriptions and paths. chunks_staging/: Directory containing the audio clips. Usage with VibeVoice Clone this repository: git clone https://huggingface.co/datasets/Tnaot/SPS-Bopha-Voice-Dataset-v1 cd SPS-Bopha-Voice-Dataset-v1 Run the training script pointing to… See the full description on the dataset page: https://huggingface.co/datasets/Tnaot/SPS-Bopha-Voice-Dataset-v1.audiotext-to-speech1K<n<10K0 likes13 downloads10mo agoHugging Face11IbraahimLab /voice-dataset Voice Dataset Collected from the web uploader tool. Voice Dataset Collected from the web uploader tool. audion<1K0 likes11 downloads7mo agoHugging Face12shadwl /voice-demogated Multilingual TTS demo — 10 languages of Vietnam and Cambodia A self-contained Gradio app. Clone the folder, install the requirements, run it. pip install -r requirements.txt python -u app.py Everything resolves relative to app.py, so no paths need editing. Languages Code Language Code Language km Khmer tyz Tay-Nung blt Tai Dam ium Dao (Iu Mien) rad Ede kpm Kho jra Jarai cma Mnong bdq Bana cjm Cham blt is Tai Dam, a Tai language of Vietnam… See the full description on the dataset page: https://huggingface.co/datasets/shadwl/voice-demo.audion<1K0 likes9 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.