datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
X-Voice-TestsetX-Voice Multilingual Test Set
High-Fidelity Test Set for Multilingual Text-to-Speech across 30 Languages
This test set is built as part of the research: X-Voice: One Speaker, 30+ Languages with Zero-Shot Voice Cloning, serving as the evaluation benchmark for our model.
Dataset Summary
30 languages
European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi (Finnish), fr (French), hr (Croatian), hu (Hungarian), it… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Testset.cv10-uk-testset-clean
The cleaned Common Voice 10 (test set) that has been checked by a human for Ukrainian 🇺🇦
Overview
This repository contains the archive of Common Voice 10 (test set) with checked Ukrainian transcriptions and audios.
All audios have been checked by a human to be sure that they are correct.
This archive is used to test all ASR models listed here: https://github.com/egorsmkv/speech-recognition-uk
Community
Discord: https://bit.ly/discord-uds
Speech… See the full description on the dataset page: https://huggingface.co/datasets/Yehor/cv10-uk-testset-clean.Audio-Understanding-Test-Set
Audio Understanding Test Set
A structured dataset for evaluating audio understanding capabilities of multimodal AI models. Contains 137 test prompts across 22 categories, paired with a 20-minute voice sample and 49 completed model outputs from Gemini 3.1 Flash Lite.
Overview
Property
Value
Total prompts
137
Implemented (with prompt text)
49
Suggested (description only)
88
Completed outputs
49
Categories
22
Model under test… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Audio-Understanding-Test-Set.kws_testset_di
ygyuan/kws_testset_di
Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards.
The input is a Kaldi-style data directory
(wav.scp, text, utt2spk, utt2dur, segments), where each
utterance is packed as a single tar sample.
Layout
data/
<split>/
metadata.csv
audio/
<split>-000.tar
<split>-001.tar
...
Shard counts:
test: 1 tar shard(s)
Inside each tar, every sample is a pair sharing a unique key:
<key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_di.kws_testset_ug_huiting
ygyuan/kws_testset_ug_huiting
Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards.
The input is a Kaldi-style data directory
(wav.scp, text, utt2spk, utt2dur, segments), where each
utterance is packed as a single tar sample.
Layout
data/
<split>/
metadata.csv
audio/
<split>-000.tar
<split>-001.tar
...
Shard counts:
test: 1 tar shard(s)
Inside each tar, every sample is a pair sharing a unique key:
<key>.wav #… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_ug_huiting.kws_testset_kk
ygyuan/kws_testset_kk
Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards.
The input is a Kaldi-style data directory
(wav.scp, text, utt2spk, utt2dur, segments), where each
utterance is packed as a single tar sample.
Layout
data/
<split>/
metadata.csv
audio/
<split>-000.tar
<split>-001.tar
...
Shard counts:
test_mht: 1 tar shard(s)
test_thu: 1 tar shard(s)
Inside each tar, every sample is a pair sharing a unique… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_kk.kws_testset_mn_huiting
ygyuan/kws_testset_mn_huiting
Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards.
The input is a Kaldi-style data directory
(wav.scp, text, utt2spk, utt2dur, segments), where each
utterance is packed as a single tar sample.
Layout
data/
<split>/
metadata.csv
audio/
<split>-000.tar
<split>-001.tar
...
Shard counts:
test: 1 tar shard(s)
Inside each tar, every sample is a pair sharing a unique key:
<key>.wav #… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_mn_huiting.kws_testset_ct_huiting
ygyuan/kws_testset_ct_huiting
Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards.
The input is a Kaldi-style data directory
(wav.scp, text, utt2spk, utt2dur, segments), where each
utterance is packed as a single tar sample.
Layout
data/
<split>/
metadata.csv
audio/
<split>-000.tar
<split>-001.tar
...
Shard counts:
test: 3 tar shard(s)
Inside each tar, every sample is a pair sharing a unique key:
<key>.wav #… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_ct_huiting.kws_testset_ct_sph
ygyuan/kws_testset_ct_sph
Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards.
The input is a Kaldi-style data directory
(wav.scp, text, utt2spk, utt2dur, segments), where each
utterance is packed as a single tar sample.
Layout
data/
<split>/
metadata.csv
audio/
<split>-000.tar
<split>-001.tar
...
Shard counts:
test: 1 tar shard(s)
Inside each tar, every sample is a pair sharing a unique key:
<key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_ct_sph.kws_testset_bo_sph
ygyuan/kws_testset_bo_sph
Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards.
The input is a Kaldi-style data directory
(wav.scp, text, utt2spk, utt2dur, segments), where each
utterance is packed as a single tar sample.
Layout
data/
<split>/
metadata.csv
audio/
<split>-000.tar
<split>-001.tar
...
Shard counts:
test: 1 tar shard(s)
Inside each tar, every sample is a pair sharing a unique key:
<key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_bo_sph.kws_testset_bo_huiting
ygyuan/kws_testset_bo_huiting
Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards.
The input is a Kaldi-style data directory
(wav.scp, text, utt2spk, utt2dur, segments), where each
utterance is packed as a single tar sample.
Layout
data/
<split>/
metadata.csv
audio/
<split>-000.tar
<split>-001.tar
...
Shard counts:
test: 2 tar shard(s)
Inside each tar, every sample is a pair sharing a unique key:
<key>.wav #… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_bo_huiting.kws_testset_bo_yalu
ygyuan/kws_testset_bo_yalu
Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards.
The input is a Kaldi-style data directory
(wav.scp, text, utt2spk, utt2dur, segments), where each
utterance is packed as a single tar sample.
Layout
data/
<split>/
metadata.csv
audio/
<split>-000.tar
<split>-001.tar
...
Shard counts:
test: 2 tar shard(s)
Inside each tar, every sample is a pair sharing a unique key:
<key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_bo_yalu.kws_testset_zh_s2t
ygyuan/kws_testset_zh_s2t
Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards.
The input is a Kaldi-style data directory
(wav.scp, text, utt2spk, utt2dur, segments), where each
utterance is packed as a single tar sample.
Layout
data/
<split>/
metadata.csv
audio/
<split>-000.tar
<split>-001.tar
...
Shard counts:
test: 12 tar shard(s)
Inside each tar, every sample is a pair sharing a unique key:
<key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_zh_s2t.kws_testset_mn_sph
ygyuan/kws_testset_mn_sph
Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards.
The input is a Kaldi-style data directory
(wav.scp, text, utt2spk, utt2dur, segments), where each
utterance is packed as a single tar sample.
Layout
data/
<split>/
metadata.csv
audio/
<split>-000.tar
<split>-001.tar
...
Shard counts:
test: 1 tar shard(s)
Inside each tar, every sample is a pair sharing a unique key:
<key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_mn_sph.wuw_testset1
Yougen/wuw_testset1
Wake-Up-Word (WUW) speech dataset, packed as WebDataset tar shards.
The input is a Kaldi-style data directory
(wav.scp, text, utt2spk, utt2dur, segments), where multiple
utterances share a long recording via the segments file.
To avoid duplicating audio, each tar sample corresponds to one full
recording. The utterance-level metadata (id / start / end / text / spk / duration) is stored in a JSON list inside that sample.
Downstream consumers slice the decoded… See the full description on the dataset page: https://huggingface.co/datasets/heimayuan/wuw_testset1.kws_testset_hm
ygyuan/kws_testset_hm
Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards.
The input is a Kaldi-style data directory
(wav.scp, text, utt2spk, utt2dur, segments), where each
utterance is packed as a single tar sample.
Layout
data/
<split>/
metadata.csv
audio/
<split>-000.tar
<split>-001.tar
...
Shard counts:
test: 3 tar shard(s)
Inside each tar, every sample is a pair sharing a unique key:
<key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_hm.kws_testset_ug_sph
ygyuan/kws_testset_ug_sph
Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards.
The input is a Kaldi-style data directory
(wav.scp, text, utt2spk, utt2dur, segments), where each
utterance is packed as a single tar sample.
Layout
data/
<split>/
metadata.csv
audio/
<split>-000.tar
<split>-001.tar
...
Shard counts:
test: 1 tar shard(s)
Inside each tar, every sample is a pair sharing a unique key:
<key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_ug_sph.asr-testset-kw-ja-v1
日本語 ASR アノテーション v1
提出済みの結果を annotations.json に保存します。音声は元リポジトリと同じディレクトリ構成・ファイル名で収録します(例:data/part-00000/example.wav)。音声の内容も変更しません。
更新がある場合、おおむね1時間ごとに同期します。下書き・メールアドレス・作業者IDは公開しません。
JSON の形式(schema_version: 2)
{
"schema_version": 2,
"annotations": [
{
"id": "example",
"source": {
"dataset": "bandad/asr-testset-kw-ja",
"revision": "元データのコミットSHA",
"audio": "data/part-00000/example.wav"
},
"audio":… See the full description on the dataset page: https://huggingface.co/datasets/bandad/asr-testset-kw-ja-v1.
