datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
contact-attendant-zhtw
Contact-Attendant zh-TW/en — speech → tool-call dialogs
The training & evaluation data behind Luigi/Qwen3-ASR-0.6B-Agent
— a 0.6B speech agent that hears a spoken request and emits a search_contacts tool call for a
bilingual (Traditional Chinese / English) office phone directory.
This dataset is fully self-contained: the audio clips, the multi-turn dialog transcripts, the
closed contact directory, and the scripts that generated them. With it you can reproduce the
fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/Luigi/contact-attendant-zhtw.OpenNewsArchive_pretrain_zhtw
Dataset Card for "yuhuanstudio/OpenNewsArchive_pretrain_zhtw"
資料集摘要
本資料集基於 OpenNewsArchive 原始數據,經過以下處理步驟:
簡繁轉換:使用 OpenCC 工具將簡體中文轉換為繁體中文並轉換常用詞彙,確保繁體用語的一致性。
格式化:整理數據結構,使其適合大型語言模型(LLM)的預訓練,確保高效的文本輸入與處理。
原始資料來源:
內容說明
數據來源:OpenDataLab - OpenNewsArchive
語言:繁體中文(基於 OpenCC 處理)
資料格式:適用於 LLM 預訓練的格式,包含標準化文本結構。
使用說明
此資料集適用於:
大型語言模型的預訓練
自然語言處理(NLP)研究
繁體中文語言處理與分析
資料集結構
{
"text":… See the full description on the dataset page: https://huggingface.co/datasets/yuhuanstudio/OpenNewsArchive_pretrain_zhtw.Wiki-zhtw-20250601
Dataset Card for Wiki-zhtw-20250601
Dataset Description
This dataset is derived from the Chinese‑Wikipedia dump dated 2025‑06‑01, downloaded from Wikimedia.Articles were extracted from the original .xml.bz2 archive with Gensim, converted to Markdown format via regular‑expression post‑processing, and finally converted from Simplified to Traditional Chinese using OpenCC.
c4_pretrain_zhtw
Dataset Card for "yuhuanstudio/c4_pretrain_zhtw"
資料集摘要
本資料集基於 C4(Colossal Clean Crawled Corpus)原始數據,並經過以下處理步驟,轉換為適用於大型語言模型(LLM)預訓練的格式:
資料清理:去除非中文內容、重複文本及不必要的 HTML 標籤,並使用pangu格式化中文語句間隔,提升語言模型的訓練品質。
格式化:將數據重新整理為適合 LLM 預訓練的結構,便於高效載入與處理。
內容說明
數據來源:Colossal Clean Crawled Corpus (C4)
語言:繁體中文
資料格式:JSON 格式,適用於 LLM 預訓練
資料數量:包含大量經過清理和格式化的繁體中文文本
使用說明
此資料集適用於:
大型語言模型的預訓練
自然語言處理(NLP)研究
繁體中文語言理解與分析
資料集結構
{
"text": "台北故事館 雲門特展 As Lomo aslomo 天空部落 TIAN… See the full description on the dataset page: https://huggingface.co/datasets/yuhuanstudio/c4_pretrain_zhtw.zhtw-roleplay-space-grimoire
Space Grimoire RP Corpus (Traditional Chinese)
Speaker-attributed dialogue from the original novel 空間魔導書與少年魔法師 (The Space Grimoire and the Young Mage; 283 chapters, ~1.7M characters), cut into scenes and assembled into ShareGPT-style role-play training data. The novel and this dataset are the work of 睡半夜怎麼三更, who holds the copyright and has no exclusive platform agreement. Data: CC BY 4.0. Code: Apache 2.0.
中文說明在下方
Dataset Summary
Source text
283… See the full description on the dataset page: https://huggingface.co/datasets/asd567557275/zhtw-roleplay-space-grimoire.Honkai_StarRail_Trailblaze_Mission_zhtw
Dataset Card for "yuhuanstudio/Honkai_StarRail_Trailblaze_Mission_zhtw"
資料集摘要
摘要:一個採集崩壞:星穹鐵道開拓任務和開拓續聞對話內容的資料集,並處理成適當的資料格式用於預訓練大模型
來源:Bilibili Wiki - 崩壞:星穹鐵道
數據類型:劇情對話文本
格式:JSON
語言:繁體中文 / 簡體中文 (zhtw/zh)
資料範圍:包含所有「開拓任務」「開拓續聞」的劇情內容,包括角色對話、選項(第一選項)、場景描述等。 (v3.0)
資料集結構
missions為開拓任務,addition為開拓續聞,未加_tw為簡體原始數據
{
<!-- 預訓練資料集資料 -->
"text": "《「均衡」的試煉•陸》\n「仲裁官」的試煉再度到來。它的內容、形式、好處和損害你已經很清楚了,不是嗎?去吧,為了「均衡」……\n仙舟「羅浮」-流雲渡"
}
{
<!-- 擷取資料 -->
"title": "混亂行至深處",
"story": [… See the full description on the dataset page: https://huggingface.co/datasets/yuhuanstudio/Honkai_StarRail_Trailblaze_Mission_zhtw.PTT-pretrain-zhtw
Dataset Card for "yuhuanstudio/PTT-pretrain-zhtw"
資料集摘要
本資料集擷取自台灣最大的 BBS 討論區——批踢踢實業坊(PTT),匯集多個看板的歷史與近期討論,提供豐富的繁體中文語料,適用於大型語言模型(LLM)預訓練與自然語言處理(NLP)研究。
數據來源:PTT 批踢踢實業坊(https://www.ptt.cc)
涵蓋看板:包含 Gossiping、Tech_Job、Stock、NBA 等所有討論區
時間範圍:擷取自 PTT 公開存檔前200頁,涵蓋多年歷史數據 (因各版頁數問題,熱門版面資料可能時間都較為古老)
語言:繁體中文
資料格式:JSON,適合 LLM 訓練與 NLP 應用
資料規模:包含數十萬條貼文與回應
資料集結構
{
"text": "作者: Sonaten (=.=)\n看板: PC_Shopping\n標題: [閒聊] Gigabyte EP35-DS3 的DES...\n時間: Fri Jun 27 15:20:54 2008\n內文:… See the full description on the dataset page: https://huggingface.co/datasets/yuhuanstudio/PTT-pretrain-zhtw.TCNNet-SFT-NetCom-zhTW-1.1M
[TCNNet] A Traditional Chinese Networking and Communication Instruction Fine-Tuning Dataset (zh-TW)
A large-scale supervised fine-tuning (SFT) dataset created specifically for TCNNet-9B, a Chinese language model specialized in networking and communications domains. The dataset contains question-answer pairs generated from various networking, cybersecurity, and tech review articles written in Traditional Chinese.
Dataset Description
Dataset Summary
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/DataAgent/TCNNet-SFT-NetCom-zhTW-1.1M.drcd-zhtw-extractive-qa-sft
steven0226/drcd-zhtw-extractive-qa-sft
繁體中文抽取式閱讀理解 SFT 資料集,衍生自 DRCD(Delta Reading Comprehension Dataset)。
來源與授權(重要)
原始資料:DRCD(Delta Research Center / 台達電子),
授權 CC BY-SA 3.0,內容改編自繁體中文維基百科。
論文引用:Shao et al., "DRCD: a Chinese Machine Reading Comprehension Dataset", arXiv:1806.00920.
本資料集是 DRCD 的 Adaptation(改編作品),依 CC BY-SA 授權鏈條,以 CC BY-SA 4.0 釋出。
所做的修改
將原始 SQuAD 風格 JSON 重新格式化為 chat SFT 格式(system/user/assistant 三則訊息,assistant 輸出固定 JSON schema)
從… See the full description on the dataset page: https://huggingface.co/datasets/steven0226/drcd-zhtw-extractive-qa-sft.medical-qa-instruction-zhtwpassages-pretrain-chinese-zhtw
Dataset Card for "yuhuanstudio/passages-pretrain-chinese-zhtw"
包含8千萬餘萬(88328203)個中文段落,不包含任何字母、數字。文字長度大部分介於 50~200 個字。
資料集來源
本資料集是基於CLUE中文預訓練語料集進行處理、過濾并進行簡繁轉諲而得到的。
資料集結構
{
"text": "雙十一點燃快遞板塊喜憂參半三季報透露何種訊號經過初步證實,這名喝農藥死亡的男子正是殺害夫妻倆的犯罪嫌疑人。"
}
資料欄位
text: (string) 文本內容
如何使用
from datasets import load_dataset
dataset = load_dataset("yuhuanstudio/passages-pretrain-chinese-zhtw", split="train")
許可資訊
[MIT]
zhtw-sentence-error-correction
中文錯字糾正資料集
由規則與字典自維基百科產生的錯誤糾正資料集。
包含錯誤類型:隨機錯字、近似音錯字、缺字錯誤、冗字錯誤。
資料集使用函式庫: p208p2002/zh-mistake-text-gen
子集
alpha: 95%錯誤,5%不變。單句中可能有多個錯誤。
beta: 50%錯誤,50%不變。單句中僅有一個錯誤。
gamma: 100%錯誤。單句中可能有多個錯誤。
en-zhtw
English ↔ Traditional Chinese Translation Dataset
This dataset is a curated combination of two high-quality parallel corpora for English to Traditional Chinese (zh-TW) translation.
It is designed for training, fine-tuning, or evaluating machine translation models, especially in academic or production settings.
📚 Dataset Overview
Source datasets:
jslin09/news_commentary_tw
zetavg/coct-en-zh-tw-translations-twp-300k
Filtering criteria:
Sentence pairs scored using… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-zhtw.mr-right-zhtw-embeddings
Mr. Right zh-TW — pre-computed document embeddings
The deployed document bank (zhen_img_v2) for the
Mr. Right zh-TW corpus:
one 4096-dimensional vector per document, covering all 769,245 documents.
Encoding all of them takes multiple GPU-days and requires the full 1.2 TB image tree.
This repository exists so you can skip both and go straight to retrieval.
file
shape / size
notes
emb.npy
(769245, 4096) float16, 6.3 GB
L2-normalised — cosine similarity is a plain dot… See the full description on the dataset page: https://huggingface.co/datasets/ericssonbear/mr-right-zhtw-embeddings.zhtw-wiki-semsearch-index
繁體中文維基百科語意搜尋索引(bge-m3 + FAISS)
這個 repo 存放 zhtw-wiki-semantic-search
專案預先建好的索引,供
Gradio demo Space
啟動時下載使用。
檔案
檔案
內容
index.faiss
FAISS IndexFlatIP,50,000 × 1024 維、L2-normalized 的 BAAI/bge-m3 dense embedding(內積 = cosine 相似度)
metadata.parquet
每列對應一個 FAISS vector id:id(= 向量位置)、pageid、title、text(段落全文)、url(維基原文連結)
建置方式
來源:zetavg/zh-tw-wikipedia
(中文維基百科的 zh-tw 變體快照)
清理:去除 markdown 標記/表格/列表/註腳/尾部參考章節後,切成 100–500 字的段落
抽樣:page 層級… See the full description on the dataset page: https://huggingface.co/datasets/steven0226/zhtw-wiki-semsearch-index.Medtext_zhtw
⚕️ MedText_zhtw
Medtext_zhtw is a Traditional chinese medicine dataset that translates from MedText,
comprising over 1000 patient presentations along with their diagnosis and treatment plans.
Example
{
"instruction": "你是一位專業的醫療人員,請用心且專業的回答問題。",
"input": "一名 50 歲男性有復發性腎結石和骨質減少病史。
由於先前診斷出維生素 D 缺乏症,他一直在服用大劑量的維生素 D 補充劑。
實驗室結果顯示高血鈣症和高鈣尿症。可能的診斷是什麼,治療方法是什麼?",
"output": "該患者有復發性腎結石、骨質減少和大劑量維生素 D 補充劑病史,… See the full description on the dataset page: https://huggingface.co/datasets/ChenWeiLi/Medtext_zhtw.Pretrain-Taiwan-DentistKnowledge-zhTW-290KLaplaceAI 繁中領域知識資料集計畫
利用我在爬蟲自動化與資料後處理上的專業,針對不同大小的領域知識資料集進行建立與維護。
在 LaplaceAI 的 huggingface 頁面,你可以找到許多不同領域的資料集。
這項 datasets 是由 LaplaceAI 整理維護的牙科相關知識。
en-zhtw-google-translate
English-Traditional Chinese Translation Dataset
A parallel dataset of 1,000,000 high-quality English and Traditional Chinese (Taiwan) sentences, machine-translated using Google Translate.
This dataset combines sentences from two curated sources:
agentlans/Taiwan-Text-Excellence-sentences (Traditional Chinese)
agentlans/high-quality-english-sentences (English)
Each sentence from one source was translated to the other language via Google Translate, creating bidirectional pairs. The… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-zhtw-google-translate.gsm8k_zhtw
GSM8K 繁體中文資料集
GSM8K(Grade School Math 8K)是由 OpenAI 發布的一個包含 8,500 個高品質小學數學文字題的資料集,旨在評估模型在多步驟數學推理任務中的表現。原始資料集以英文撰寫,為了促進繁體中文社群對該資料集的使用與研究,我們將其完整翻譯為繁體中文。
翻譯方法
我們採用了 Google translate 進行自動翻譯,確保問題和解答的語意與原始資料集一致。在翻譯過程中,我們特別注意數學符號、專有名詞和文化差異,確保繁體中文使用者能夠清晰理解每個問題。
資料集內容
翻譯後的資料集包含:
問題(question):小學數學問題的繁體中文描述。
答案(answer):對應問題的完整解答,包括多步驟推理和最終答案。
每個問題的解答格式保持與原始資料集一致,方便使用者進行比較和研究。
資料集結構
資料集分為訓練集和測試集:
訓練集:7,473 個問題
測試集:1,319 個問題
每個問題都需要 2 到 8… See the full description on the dataset page: https://huggingface.co/datasets/yuhuanstudio/gsm8k_zhtw.zh_translation_benchmark
zh_translation_benchmark
zh_translation_benchmark is a 2,000-example synthetic benchmark for evaluating whether a translation or rewriting model can produce natural Taiwan Traditional Chinese (zh-TW) from English, Mainland Chinese, Hong Kong Traditional Chinese, Cantonese-style written Chinese, or code-mixed English/Chinese documents.
The dataset now exposes a single Hugging Face subset/config: full. It is not split into dev and test files. The full 2,000 rows are loaded… See the full description on the dataset page: https://huggingface.co/datasets/OpenFormosa/zh_translation_benchmark.gptoss_zhtw_allima-corpus-zhtw
IMA Traditional Chinese Corpus(繁體中文語料總集)
本資料集為繁體中文文學語料總集,目的在於將原先分散於多個作者/來源 dataset repo 的繁體中文文本統一整併,提供「一次申請、持續更新」的集中存取方式。
使用者只需申請本 dataset(本 repo)一次,即可取得所有繁中語料。未來新增來源或更新資料將直接同步至本 repo,無需重複申請。
📂 目錄結構
所有來源資料皆保留於 data/ 之下,每個子資料夾對應一個原始來源 repo,例如:
data/
├── zhtw-literature-ots
每個子資料夾內保留:
原始 README
原始語料檔(json / txt 等)
來源資訊與授權說明
以利來源追溯與資料審核。
📦 資料格式
主要格式:
JSON
UTF-8 編碼文字檔
典型欄位可能包含:
title:作品名稱
author:作者
content:文本內容
source:來源 repo
(依各來源資料實際格式而定)… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/ima-corpus-zhtw.poem-pretrain-chinese-zhtw
Dataset Card for "poem-pretrain-chinese-zhtw"
資料集摘要
中文古典文集資料庫收集了約 5.5 萬首唐詩、26 萬首宋詩、2.1 萬首宋詞和其他古典文集。詩人包括唐宋兩朝近 1.4 萬古詩人,和兩宋時期 1.5 千古詞人。
五代十國- 收錄"花間集"與"南唐二主詞"
唐- 收錄"全唐詩"(是清康熙四十四年,康熙皇帝主導下,蒐集羅唐詩的收藏「得詩 48,900 餘首,詩入 2,200 人」)。
宋- 收錄"全宋詞"(由唐圭璋編著,孔凡禮補輯,共收錄宋代詞人 1,330 家,詞作 21,116 首)。
元- 收錄元曲 11,057 篇,曲家 233 人。
清- 收錄"納蘭性德詩集"
原始資料來源:
chinese-poetry: 最全中文诗歌古典文集数据库
erhwenkuo/poetry-chinese-zhtw
資料集結構
{
"text":… See the full description on the dataset page: https://huggingface.co/datasets/yuhuanstudio/poem-pretrain-chinese-zhtw.zhtw-literature-ots
Dataset Summary
The dataset contains 2,349 rows.
These paragraphs are extracted from authorized novels written by Ou Tiong Siong胡長松 and contain multiple sentences in Traditional Chinese.
The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research.
Dataset Structure
Number of rows: 2,349 (each representing a paragraph)
Features:
title: Book title… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/zhtw-literature-ots.APIGen_ZHtwtaiwan-corpus-zhtwapigen_zhtw_doneZH-TW_Reading_Comprehension_Test_for_LLMszh-test-llama3-v1zhtw-literature-meowverse
Dataset Summary
The dataset contains 14,256 rows.
These paragraphs are extracted from authorized works written by Chang Luo-Miao張珞喵 and published on Vocus 方格子. Most articles are composed in Traditional Chinese, and many of them also carry parallel English, Japanese, German and Spanish versions of the same content within the same article, so the dataset is multilingual by design rather than translated afterwards. Roughly 43% of the paragraphs are Traditional Chinese, 54% English… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/zhtw-literature-meowverse.
