datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Wiki-zhtw-20250601
Dataset Card for Wiki-zhtw-20250601
Dataset Description
This dataset is derived from the Chinese‑Wikipedia dump dated 2025‑06‑01, downloaded from Wikimedia.Articles were extracted from the original .xml.bz2 archive with Gensim, converted to Markdown format via regular‑expression post‑processing, and finally converted from Simplified to Traditional Chinese using OpenCC.
zhtw-roleplay-space-grimoire
Space Grimoire RP Corpus (Traditional Chinese)
Speaker-attributed dialogue from the original novel 空間魔導書與少年魔法師 (The Space Grimoire and the Young Mage; 283 chapters, ~1.7M characters), cut into scenes and assembled into ShareGPT-style role-play training data. The novel and this dataset are the work of 睡半夜怎麼三更, who holds the copyright and has no exclusive platform agreement. Data: CC BY 4.0. Code: Apache 2.0.
中文說明在下方
Dataset Summary
Source text
283… See the full description on the dataset page: https://huggingface.co/datasets/asd567557275/zhtw-roleplay-space-grimoire.TCNNet-SFT-NetCom-zhTW-1.1M
[TCNNet] A Traditional Chinese Networking and Communication Instruction Fine-Tuning Dataset (zh-TW)
A large-scale supervised fine-tuning (SFT) dataset created specifically for TCNNet-9B, a Chinese language model specialized in networking and communications domains. The dataset contains question-answer pairs generated from various networking, cybersecurity, and tech review articles written in Traditional Chinese.
Dataset Description
Dataset Summary
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/DataAgent/TCNNet-SFT-NetCom-zhTW-1.1M.Pretrain-Taiwan-DentistKnowledge-zhTW-290KLaplaceAI 繁中領域知識資料集計畫
利用我在爬蟲自動化與資料後處理上的專業,針對不同大小的領域知識資料集進行建立與維護。
在 LaplaceAI 的 huggingface 頁面,你可以找到許多不同領域的資料集。
這項 datasets 是由 LaplaceAI 整理維護的牙科相關知識。
ima-corpus-zhtw
IMA Traditional Chinese Corpus(繁體中文語料總集)
本資料集為繁體中文文學語料總集,目的在於將原先分散於多個作者/來源 dataset repo 的繁體中文文本統一整併,提供「一次申請、持續更新」的集中存取方式。
使用者只需申請本 dataset(本 repo)一次,即可取得所有繁中語料。未來新增來源或更新資料將直接同步至本 repo,無需重複申請。
📂 目錄結構
所有來源資料皆保留於 data/ 之下,每個子資料夾對應一個原始來源 repo,例如:
data/
├── zhtw-literature-ots
每個子資料夾內保留:
原始 README
原始語料檔(json / txt 等)
來源資訊與授權說明
以利來源追溯與資料審核。
📦 資料格式
主要格式:
JSON
UTF-8 編碼文字檔
典型欄位可能包含:
title:作品名稱
author:作者
content:文本內容
source:來源 repo
(依各來源資料實際格式而定)… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/ima-corpus-zhtw.ZHTrainDataTCNNet-Pretrain-NetCom-zhTW-3.7M
[TCNNet] A Large-scale Traditional Chinese Networking and Communication Continuous Pretraining Dataset (zh-TW)
A specialized domain knowledge dataset created for continuous pretraining of TCNNet-9B, a Chinese language model based on Yi-9B and specialized in networking and communications domains. The dataset contains articles from various networking, cybersecurity, and tech review sources written in Traditional Chinese.
Dataset Description
Dataset Summary
This… See the full description on the dataset page: https://huggingface.co/datasets/DataAgent/TCNNet-Pretrain-NetCom-zhTW-3.7M.
