IMA-Taiwan/ima-corpus-zhtw
IMA Traditional Chinese Corpus(繁體中文語料總集) 本資料集為繁體中文文學語料總集,目的在於將原先分散於多個作者/來源 dataset repo 的繁體中文文本統一整併,提供「一次申請、持續更新」的集中存取方式。 使用者只需申請本 dataset(本 repo)一次,即可取得所有繁中語料。未來新增來源或更新資料將直接同步至本 repo,無需重複申請。 📂 目錄結構 所有來源資料皆保留於 data/ 之下,每個子資料夾對應一個原始來源 repo,例如: data/ ├── zhtw-literature-ots 每個子資料夾內保留: 原始 README 原始語料檔(json / txt 等) 來源資訊與授權說明 以利來源追溯與資料審核。 📦 資料格式 主要格式: JSON UTF-8 編碼文字檔 典型欄位可能包含: title:作品名稱 author:作者 content:文本內容 source:來源 repo (依各來源資料實際格式而定)… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/ima-corpus-zhtw.
012
No card is published for this repository, or it could not be fetched from Hugging Face right now.
