IMA-Taiwan/ima-taiwanese-corpus-merged
IMA Taiwanese Corpus(台語語料總集) 本資料集為台語語料總集,目的在於將原先分散於多位作者 dataset repo 的台語語料統一整併, 提供「一次申請、持續更新」的存取方式。 目錄結構 所有來源資料皆保留於 data/ 之下,每個子資料夾對應一個原始來源 repo,例如: data/taigi-literature-ots data/taigi-literature-tks data/taigi-literature-abt (其餘同理) 每個子資料夾內保留原始 README 與語料檔案,以利來源追溯與審核。 使用方式 申請者只需申請本 dataset(本 repo)一次,即可取得所有台語語料。 未來新增作者或新增語料將直接更新於本 repo,不需重複申請。
010
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face