CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01XRXRX /X-Voice-TestsetX-Voice Multilingual Test Set High-Fidelity Test Set for Multilingual Text-to-Speech across 30 Languages This test set is built as part of the research: X-Voice: One Speaker, 30+ Languages with Zero-Shot Voice Cloning, serving as the evaluation benchmark for our model. Dataset Summary 30 languages European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi (Finnish), fr (French), hr (Croatian), hu (Hungarian), it… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Testset.audiotext-to-speech10K<n<100K4 likes72 downloads5mo agoHugging Face02Yehor /cv10-uk-testset-clean The cleaned Common Voice 10 (test set) that has been checked by a human for Ukrainian 🇺🇦 Overview This repository contains the archive of Common Voice 10 (test set) with checked Ukrainian transcriptions and audios. All audios have been checked by a human to be sure that they are correct. This archive is used to test all ASR models listed here: https://github.com/egorsmkv/speech-recognition-uk Community Discord: https://bit.ly/discord-uds Speech… See the full description on the dataset page: https://huggingface.co/datasets/Yehor/cv10-uk-testset-clean.audioautomatic-speech-recognition1K<n<10K3 likes44 downloads2y agoHugging Face03ygyuan /kws_testset_digated ygyuan/kws_testset_di Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: test: 1 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_di.audioaudio-classification0 likes6 downloads2mo agoHugging Face04ygyuan /kws_testset_bo_yalugated ygyuan/kws_testset_bo_yalu Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: test: 2 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_bo_yalu.audioaudio-classification0 likes6 downloads2mo agoHugging Face05ygyuan /kws_testset_ug_huitinggated ygyuan/kws_testset_ug_huiting Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: test: 1 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav #… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_ug_huiting.audioaudio-classification0 likes6 downloads2mo agoHugging Face06ygyuan /kws_testset_kkgated ygyuan/kws_testset_kk Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: test_mht: 1 tar shard(s) test_thu: 1 tar shard(s) Inside each tar, every sample is a pair sharing a unique… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_kk.audioaudio-classification0 likes5 downloads2mo agoHugging Face07ygyuan /kws_testset_mn_huitinggated ygyuan/kws_testset_mn_huiting Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: test: 1 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav #… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_mn_huiting.audioaudio-classification1 likes5 downloads2mo agoHugging Face08ygyuan /kws_testset_ct_sphgated ygyuan/kws_testset_ct_sph Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: test: 1 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_ct_sph.audioaudio-classification0 likes5 downloads2mo agoHugging Face09ygyuan /kws_testset_bo_sphgated ygyuan/kws_testset_bo_sph Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: test: 1 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_bo_sph.audioaudio-classification0 likes5 downloads2mo agoHugging Face10ygyuan /kws_testset_bo_huitinggated ygyuan/kws_testset_bo_huiting Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: test: 2 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav #… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_bo_huiting.audioaudio-classification0 likes5 downloads2mo agoHugging Face11ygyuan /kws_testset_zh_s2tgated ygyuan/kws_testset_zh_s2t Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: test: 12 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_zh_s2t.audioaudio-classification0 likes5 downloads2mo agoHugging Face12ygyuan /kws_testset_mn_sphgated ygyuan/kws_testset_mn_sph Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: test: 1 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_mn_sph.audioaudio-classification0 likes4 downloads2mo agoHugging Face13ygyuan /kws_testset_ct_huitinggated ygyuan/kws_testset_ct_huiting Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: test: 3 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav #… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_ct_huiting.audioaudio-classification0 likes4 downloads2mo agoHugging Face14heimayuan /wuw_testset1gated Yougen/wuw_testset1 Wake-Up-Word (WUW) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where multiple utterances share a long recording via the segments file. To avoid duplicating audio, each tar sample corresponds to one full recording. The utterance-level metadata (id / start / end / text / spk / duration) is stored in a JSON list inside that sample. Downstream consumers slice the decoded… See the full description on the dataset page: https://huggingface.co/datasets/heimayuan/wuw_testset1.audioautomatic-speech-recognition0 likes4 downloads2mo agoHugging Face15ygyuan /kws_testset_hmgated ygyuan/kws_testset_hm Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: test: 3 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_hm.audioaudio-classification0 likes3 downloads2mo agoHugging Face16ygyuan /kws_testset_ug_sphgated ygyuan/kws_testset_ug_sph Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: test: 1 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_ug_sph.audioaudio-classification0 likes3 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.