datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Taiwan-Tongues-ASR-CE-dataset-zhtw
Taiwan-Tongues-ASR-CE-dataset-zhtw
本資料集為 Taiwan-Tongues-ASR-CE 專案所使用的預訓練資料,透過 WebDataset 格式打包,並上傳至 Hugging Face 以便研究人員與開發者自由取用。
📂 Dataset 結構
本資料集分為 Training 與 Test 兩個子集,均以 WebDataset tar 檔案形式存放:
Training set (WebDataset format)
train/train-000000.tar
train/train-000001.tar
...
Test set (WebDataset format)
test/test-000000.tar
...
tsv set
train.tsv
test.tsv
...
每個 tar 內部均包含對應的音檔與標註,方便直接搭配 WebDataset 與 PyTorch / Hugging Face datasets 進行訓練與測試。
🏷️… See the full description on the dataset page: https://huggingface.co/datasets/adi-gov-tw/Taiwan-Tongues-ASR-CE-dataset-zhtw.contact-attendant-zhtw
Contact-Attendant zh-TW/en — speech → tool-call dialogs
The training & evaluation data behind Luigi/Qwen3-ASR-0.6B-Agent
— a 0.6B speech agent that hears a spoken request and emits a search_contacts tool call for a
bilingual (Traditional Chinese / English) office phone directory.
This dataset is fully self-contained: the audio clips, the multi-turn dialog transcripts, the
closed contact directory, and the scripts that generated them. With it you can reproduce the
fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/Luigi/contact-attendant-zhtw.breezyvoice-zhtw-en-phone-corpus
BreezyVoice zh-TW / English Phone-Attendant Corpus (ASR-verified)
Synthetic Taiwan-Mandarin + English code-mixed speech for a phone-attendant domain,
generated with MediaTek-Research/BreezyVoice
(zero-shot voice clone, fixed reference voice) and ASR-verified: every clip was transcribed
with faster-whisper and kept only if its Han-character CER vs the intended text was below 0.3.
Built to distill BreezyVoice into a tiny real-time on-device TTS (the
Inflect-Nano architecture)… See the full description on the dataset page: https://huggingface.co/datasets/Luigi/breezyvoice-zhtw-en-phone-corpus.
