datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CTIS
Dataset Card for Chinese Traditional Instrument Sound
Original Content
The original dataset is created by [1], with no evaluation provided. The original CTIS dataset contains recordings from 287 varieties of Chinese traditional instruments, reformed Chinese musical instruments, and instruments from ethnic minority groups. Notably, some of these instruments are rarely encountered by the majority of the Chinese populace. The dataset was later utilized by [2] for Chinese… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/CTIS.d2l4asr-wiki-jad2l4asr-wiki-en_audioikema_youtube_asr_full_with_longSLR35_javanesekillkan
Killkan: Speech Recognition dataset for Kichwa
Killkan (Kichwa uyachkata payllatak killkak anta) is the first automatic speech recognition (ASR) dataset for the Kichwa language.
See also our paper (https://arxiv.org/abs/2404.15501).
nollywood-ctc-scored-ep3-hauwad2l4asr-wiki-enctphoneme-ctc-spanish-52h-noisyphoneme-ctc-english-60hphoneme-ctc-english-41husc_cleaned_ctc_filteredikema_dict_asrSingingVoiceDeepfakeDetection_CtrSVDD_ACEKiSing_M4Singerphoneme-ctc-english-60h-balanced
Phoneme CTC — English 60h (Balanced & Normalized)
A cleaned, normalized and phoneme-balanced version of
bobboyms/phoneme-ctc-english-60h-noisy,
for training phoneme recognition models (CTC) — e.g. as the native acoustic
model behind pronunciation-feedback systems.
What's different from the source dataset
Label noise removed
Roman numerals dropped — eSpeak reads ii/iv/… as "Roman two/four",
producing labels that don't match the audio.
Non-English phonemes dropped… See the full description on the dataset page: https://huggingface.co/datasets/bobboyms/phoneme-ctc-english-60h-balanced.papi_asr_testpapi_asrikema_dictionary_examples_datasetpapi_asr_miniphoneme-ctc-english-60h-noisyMMAU-Pro-Ctrlikema_youtube_asr_testhubert_ctc_ftwur_ctc_kln_scoredkws_dataset_ct
ygyuan/kws_dataset_ct
Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards.
The input is a Kaldi-style data directory
(wav.scp, text, utt2spk, utt2dur, segments), where each
utterance is packed as a single tar sample.
Layout
data/
<split>/
metadata.csv
audio/
<split>-000.tar
<split>-001.tar
...
Shard counts:
train: 941 tar shard(s)
Inside each tar, every sample is a pair sharing a unique key:
<key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/heimayuan/kws_dataset_ct.sinhala-ctc-111h
Sinhala ASR – Consolidated OpenSLR (SLR52)
Dataset Summary
This dataset is a consolidated and cleaned version of the Sinhala Automatic Speech Recognition (ASR) dataset from OpenSLR (SLR52).
The original OpenSLR release distributes the data across multiple subsets (0–9, a–f).
This repository merges all subsets into a single unified dataset containing approximately 111 hours of speech audio.
Dataset Description
Consolidation
All OpenSLR SLR52… See the full description on the dataset page: https://huggingface.co/datasets/IAmNotAnanth/sinhala-ctc-111h.kws_testset_ct_sph
ygyuan/kws_testset_ct_sph
Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards.
The input is a Kaldi-style data directory
(wav.scp, text, utt2spk, utt2dur, segments), where each
utterance is packed as a single tar sample.
Layout
data/
<split>/
metadata.csv
audio/
<split>-000.tar
<split>-001.tar
...
Shard counts:
test: 1 tar shard(s)
Inside each tar, every sample is a pair sharing a unique key:
<key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_ct_sph.diff_datavn-speech-text-ctv-v1.1
