datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
s2tt-yoruba-englishs2tt-igbo-englishs2tt-hausa-englishS2T_Split_NoRom_phase2S2T_Korean_Merge_2_fixed4S2T_English_Vietnamese_AuVi_2S2T_Korean_limit_silent_spaceen2ja.s2t_translationS2T_Korean_Merge_2_fixed2ja2en.s2t_translationS2T_Korean_3s_silentS2TTS2T_English_Vietnamese_AuVi_testS2T_Korean_Merges2tkp_dataset_part_3multi-lang-s2t-dataset
Multi-Language Audio Dataset
A high-quality, cleaned audio dataset derived from five AI4Bharat sources, containing speech-to-text data in Hindi, Gujarati, and Telugu with English translations.
Dataset Overview
Language
Duration (Hours)
Hindi
55.45
Gujarati
42.12
Telugu
37.83
Total Duration: ~135 hours of cleaned, high-quality audio
Source Datasets
Data collected and cleaned from:
ai4bharat/IndicVoices-ST
ai4bharat/NPTEL… See the full description on the dataset page: https://huggingface.co/datasets/Kaushalb11/multi-lang-s2t-dataset.kws_testset_zh_s2t
ygyuan/kws_testset_zh_s2t
Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards.
The input is a Kaldi-style data directory
(wav.scp, text, utt2spk, utt2dur, segments), where each
utterance is packed as a single tar sample.
Layout
data/
<split>/
metadata.csv
audio/
<split>-000.tar
<split>-001.tar
...
Shard counts:
test: 12 tar shard(s)
Inside each tar, every sample is a pair sharing a unique key:
<key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_zh_s2t.S2T_Splited_FinalS2T_SplitEndMovieS2T_Korean_Merge_2_fixedS2T_Korean_Merge_2_fixed3S2T_Splited_GameshowS2T_SplitEndMovie_SentencesS2T_Split_30ss2tkp_dataset_part_1s2tkp_dataset_part_2S2T_MergedS2T_Merged_SentencesS2T_Split_SentencesS2T_Merged_25s
