datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TaigiSpeech
TaigiSpeech
A spoken language understanding (SLU) dataset for Taiwanese (台語/Taigi) intent classification, designed for elder-care and smart-home voice command scenarios.
Paper: TaigiSpeech: A Low-Resource Real-World Speech Intent Dataset and Preliminary Results with Scalable Data Mining In-the-Wild
Dataset Description
TaigiSpeech contains 3,000+ Taiwanese speech utterances from 21 speakers, each labeled with one of 8 intent classes. The dataset is designed to… See the full description on the dataset page: https://huggingface.co/datasets/TaigiSpeech/TaigiSpeech.pts-taigi
PTS Taigi
Collection of Taiwanese-language YouTube talk-show/documentary content, staged from COS
ahead of a full transfer. Each COS source gets its own config (schemas differ, so they
can't share one), each with one split named after the source folder.
Configs / Splits
ptv_ts_ds_test_zhtw (367 rows) -- schema normalized to this project's conventions
(see rule.md): tw/zh were split into clean text/mandarin plus
text_timestamped/mandarin_timestamped (original… See the full description on the dataset page: https://huggingface.co/datasets/NickWeng/pts-taigi.yttd_taigi_trs
