datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tw-leetcode
Dataset Card for tw-leetcode
A curated Traditional Chinese LeetCode solution dataset with high-efficiency answers (Beats 100%), structured explanation in "Top Concept → Step Implement → Complexity Analysis" style, updated daily.
Dataset Details
Dataset Description
tw-leetcode 是一個針對 LeetCode 題目的繁體中文資料集,內容包含高效能程式解法、完整的解題思路,以及時間與空間複雜度分析。每份題解都經由人工清洗與優化,並依循「Top Concept → Step Implement → Complexity Explanation」的結構撰寫,方便機器學習模型或人類讀者理解程式邏輯的推理過程。
本資料集適合作為:… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-leetcode.tw-reasoning-instruct-50k
Dataset Card for tw-reasoning-instruct-50k
tw-reasoning-instruct-50k 是一個精選的 繁體中文(台灣) 推理資料集,旨在提升語言模型於逐步邏輯思考、解釋生成與語言理解等任務中的表現。資料內容涵蓋日常思辨、教育對話、法律推理等多元主題,並結合「思考步驟」與「最終答案」的結構設計,引導模型以更清晰、條理分明的方式進行推論與回應,特別強調符合台灣本地語言與文化背景的應用需求。
Dataset Details
Dataset Description
本資料集專為發展具備強大推理能力的繁體中文大型語言模型(Large Reasoning Models, LRM)所設計,內容深度結合台灣的語言與文化脈絡。每筆資料通常包含使用者的提問、模型的回應,以及清楚的推理過程。資料集設計目標為培養模型具備類人邏輯的逐步思考與解釋能力。
此資料集適用於訓練與評估以下任務:
台灣社會的日常推理
教育性對話
以解釋為導向的生成任務… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-reasoning-instruct-50k.tw-drug-labels-vision
Dataset Card for tw-drug-labels-vision
💊 tw-drug-labels-vision 是一份涵蓋臺灣食品藥物管理署(TFDA)核發之 44,663 筆藥品仿單/外盒 的繁體中文多模態資料集。每一筆紀錄同時包含 PDF 全部頁面的渲染圖(WebP 多頁)以及一份依統一 17 欄 JSON Schema 抽取自原始藥品標示文件的結構化資料,可直接用於語言模型微調、視覺語言模型訓練、文件問答、藥品知識檢索、繁體中文醫藥 NLP 任務之素材。
Dataset Details
Dataset Description
本資料集源自臺灣 TFDA 公開的藥品許可證查詢系統。每筆紀錄對應一份藥品文件(仿單或外盒),原始為 PDF 圖檔形式。處理流程分為三階段:
下載:依據 20251222政府開放資料集_仿單與藥品外盒_66032.xlsx 中的 PDF URL,下載原始檔。
頁面渲染:將 PDF 各頁渲染為 WebP 圖檔,封裝在 images 欄位中。
OCR +… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-drug-labels-vision.tw-function-call-reasoning-10k
Dataset Card for tw-function-call-reasoning-10k
本資料集為繁體中文版本的函式呼叫(Function Calling)資料集,翻譯自 AymanTarig/function-calling-v0.2-with-r1-cot,而該資料集本身是 Salesforce/xlam-function-calling-60k 的修正版。我們利用語言模型翻譯後,經人工修改,旨在打造高品質的繁體中文工具使用語料。
Dataset Details
Dataset Description
tw-function-call-reasoning-10k 是一個專為語言模型「工具使用能力(Function Calling)」訓練所設計的繁體中文資料集。其內容源自 AymanTarig/function-calling-v0.2-with-r1-cot,該資料集又為 Salesforce/xlam-function-calling-60k… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-function-call-reasoning-10k.tw-math-reasoning-2k
Dataset Card for tw-math-reasoning-2k
tw-math-reasoning-2k 是一個繁體中文數學語言資料集,從 HuggingFaceH4/MATH 英文數學題庫中精選 2,000 題,並透過 perplexity-ai/r1-1776 模型以繁體中文重新生成具邏輯性且詳盡的解題過程與最終答案。此資料集可作為訓練或評估繁體中文數學推理模型的高品質參考語料。
Dataset Details
Dataset Description
tw-math-reasoning-2k 是一個繁體中文數學語言資料集,旨在提供高品質的解題語料以支援中文數學推理模型的訓練與評估。此資料集從 HuggingFaceH4/MATH 英文數學題庫中精選 2,000 題,涵蓋代數、幾何、機率統計等各類題型,並確保題目類型分佈均衡。
所有題目皆經由 perplexity-ai/r1-1776… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-math-reasoning-2k.finevoices-zhtw
Dataset Card for finevoices-zhtw (WIP)
Brief Summary (Preliminary)
This section describes the initial design intent of the project.
It serves as a first-edition summary and may be revised or expanded
as the dataset scope, governance, and contribution process become more clearly defined.
本計畫 finevoices-zhtw 旨在建立一套以繁體中文(台灣)為核心、可合法使用、可長期維護的語音資料集,作為繁體中文語言與語音模型在 語音辨識(ASR)、語音合成(TTS)、多模態模型訓練與微調(fine-tuning) 等任務上的公共基礎資料來源。
本專案的整體精神,參考 Mozilla Common Voice 與 Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/finevoices-zhtw.tw-tokenizer-bench
tw-tokenizer-bench v0.1
台灣繁體中文 tokenizer 評測基準。24,294 筆 / 1,844 萬字元 ,5 個領域 subset。
定位是正式文體 ——法律、政府、百科、翻譯網頁。這是刻意的取捨,不是「台灣中文全貌」,
限制寫在下面的〈已知限制〉。
筆數
24,294
字元數
18,439,640
subset
5
每筆長度
300–1,500 字元(固定窗格)
語言
繁體中文(台灣)
Subsets
subset
筆數
字元數
平均長度
來源
law_judgment
5,000
4,386,700
877
司法院判決書
web_translated
5,000
4,178,890
836
ACE-2 翻譯繁中網頁
gov_news
5,000
3,849,201
770
政府新聞稿
encyclopedia
5,000
2,047,872
410
中文維基
law_statute
4,294
3… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-tokenizer-bench.finepdfs-zhtw
Dataset Card for finepdfs-zhtw (WIP)
Brief Summary (Preliminary)
This section describes the initial design intent of the project.
It serves as a first-edition summary and may be revised or expanded
as the dataset scope, governance, and contribution process become more clearly defined.
本計畫旨在建立一套以繁體中文為核心、可合法使用、可長期維護的 PDF 文本資料集,作為繁體中文語言模型在 文件理解(Document Understanding)、OCR 後處理與微調訓練(fine-tuning) 等任務上的基礎資料來源。
整體設計將參考 Hugging Face 社群所建立之 finepdf… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/finepdfs-zhtw.finetranslations-zhtw
Dataset Card for finetranslations-zhtw (WIP)
Brief Summary (Preliminary)
This section describes the initial design intent of the project.
It serves as a first-edition summary and may be revised or expanded
as the dataset scope, governance, and contribution process become more clearly defined.
本資料集旨在打造一套高品質的多語言 → 繁體中文(zh-tw)機器翻譯資料集,延伸 Hugging Face 生態系中既有「多語言 → 英文(many-to-en)」的 💬 FineTranslations 翻譯資料集的建構脈絡,補足以繁體中文為核心的多語言翻譯資源。
資料集的建構流程以大型語言模型(LLM)進行多語言文本至繁體中文的初步翻譯,並導入… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/finetranslations-zhtw.finevision-zhtw
Dataset Card for finevisions-zhtw
Brief Summary (Preliminary)
This section describes the initial design intent of the project.
It serves as a first-edition summary and may be revised or expanded
as the dataset scope, governance, and contribution process become more clearly defined.
如何幫忙本專案
本專案正在進行中,歡迎任何夥伴加入幫忙,一起讓繁體中文的預訓練語料更多元及完善。
