datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
taiwan-conversation-context-100-domains
Taiwan Conversation Context 100 Domains
Dataset Description
Taiwan Conversation Context 100 Domains 是一套以台灣日常生活情境為核心設計的雙人對話文本資料集。
本資料集包含 100 個生活領域,每個領域各有 12,000 筆對話資料,總計約 1,200,000 筆對話樣本。每筆資料皆為雙人對話格式,包含 [A][B][A][B][A][B][A][B] 共 8 個發言,也就是 4 輪來回對話。
資料以繁體中文撰寫,並針對台灣在地語境設計,適合用於:
語音生成資料前處理
Text-to-Speech, TTS
Spoken Dialogue Generation
Conversational AI
Customer Service Dialogue Modeling
Role-play Dialogue Dataset
台灣繁體中文語音模型訓練
生活情境問答模型訓練
對話式 AI 助理訓練
RAG / Agent 測試資料… See the full description on the dataset page: https://huggingface.co/datasets/Ethan615/taiwan-conversation-context-100-domains.twngrams
Taiwanese Mandarin web n-grams
Word 1–4-gram counts for Taiwanese Mandarin* (臺灣華語, cmn-Hant-TW), computed over the
Taiwan slice of a large web crawl after variety filtering by
twfilter 0.1.0 with the published
twfilter-tables:
every sentence behind these counts passed the 教育部 character-inventory gate, the
simplified-character round-trip, the mainland-orthography, mainland-lexicon,
written-Cantonese, Hong Kong and Singapore detectors, and block-level evidence of
Taiwan-specific… See the full description on the dataset page: https://huggingface.co/datasets/taiwan-corpora/twngrams.
