taigi
Datasets
All datasets matching “taigi”TaigiSpeech
TaigiSpeech
A spoken language understanding (SLU) dataset for Taiwanese (台語/Taigi) intent classification, designed for elder-care and smart-home voice command scenarios.
Paper: TaigiSpeech: A Low-Resource Real-World Speech Intent Dataset and Preliminary Results with Scalable Data Mining In-the-Wild
Dataset Description
TaigiSpeech contains 3,000+ Taiwanese speech utterances from 21 speakers, each labeled with one of 8 intent classes. The dataset is designed to… See the full description on the dataset page: https://huggingface.co/datasets/TaigiSpeech/TaigiSpeech.pts-taigi
PTS Taigi
Collection of Taiwanese-language YouTube talk-show/documentary content, staged from COS
ahead of a full transfer. Each COS source gets its own config (schemas differ, so they
can't share one), each with one split named after the source folder.
Configs / Splits
ptv_ts_ds_test_zhtw (367 rows) -- schema normalized to this project's conventions
(see rule.md): tw/zh were split into clean text/mandarin plus
text_timestamped/mandarin_timestamped (original… See the full description on the dataset page: https://huggingface.co/datasets/NickWeng/pts-taigi.yttd_taigi_trstaigi-asr-trialtaigi-literature-asts
Dataset Summary
The dataset contains 2,494 rows.
These paragraphs are extracted from authorized novels written by Ang Siok Tsiau洪淑昭 and contain multiple sentences in Taiwanese Taigi written with Hanji.
The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research.
Dataset Structure
Number of rows: 2,494 (each representing a paragraph)
Features:… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-asts.taigi-literature-ttshs
Dataset Summary
The dataset contains 240 rows.
These paragraphs are extracted from authorized novel written by Tiunn Tshing Siong張青松 and contain multiple sentences in Taiwanese Taigi written with Hanji.
The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research.
Dataset Structure
Number of rows: 240 (each representing a paragraph)
Features:… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-ttshs.
