datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Twin-2K-500
Twin-2K-500 Dataset
This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations.
More information on how to use this dataset can be found in our Documentation and GitHub repository.
Details on how the dataset was generated are available in our Paper.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500.tw-leetcode
Dataset Card for tw-leetcode
A curated Traditional Chinese LeetCode solution dataset with high-efficiency answers (Beats 100%), structured explanation in "Top Concept → Step Implement → Complexity Analysis" style, updated daily.
Dataset Details
Dataset Description
tw-leetcode 是一個針對 LeetCode 題目的繁體中文資料集,內容包含高效能程式解法、完整的解題思路,以及時間與空間複雜度分析。每份題解都經由人工清洗與優化,並依循「Top Concept → Step Implement → Complexity Explanation」的結構撰寫,方便機器學習模型或人類讀者理解程式邏輯的推理過程。
本資料集適合作為:… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-leetcode.tw-reasoning-instruct-50k
Dataset Card for tw-reasoning-instruct-50k
tw-reasoning-instruct-50k 是一個精選的 繁體中文(台灣) 推理資料集,旨在提升語言模型於逐步邏輯思考、解釋生成與語言理解等任務中的表現。資料內容涵蓋日常思辨、教育對話、法律推理等多元主題,並結合「思考步驟」與「最終答案」的結構設計,引導模型以更清晰、條理分明的方式進行推論與回應,特別強調符合台灣本地語言與文化背景的應用需求。
Dataset Details
Dataset Description
本資料集專為發展具備強大推理能力的繁體中文大型語言模型(Large Reasoning Models, LRM)所設計,內容深度結合台灣的語言與文化脈絡。每筆資料通常包含使用者的提問、模型的回應,以及清楚的推理過程。資料集設計目標為培養模型具備類人邏輯的逐步思考與解釋能力。
此資料集適用於訓練與評估以下任務:
台灣社會的日常推理
教育性對話
以解釋為導向的生成任務… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-reasoning-instruct-50k.Twin-2K-500-Mega-Study
Twin-2K-500-Mega-Study Dataset
GitHub Repository: https://github.com/TianyiPeng/Twin-2K-500-Mega-Study
To see more details for how to process these data, please refer to this GitHub repository.
This dataset contains survey data from the Twin-2K-500 Mega Study, which tests the validity of using large language models to predict people's future answers based on their answers to past surveys (creating "digital twins" of participants).
Dataset Structure
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500-Mega-Study.tw-drug-labels-vision
Dataset Card for tw-drug-labels-vision
💊 tw-drug-labels-vision 是一份涵蓋臺灣食品藥物管理署(TFDA)核發之 44,663 筆藥品仿單/外盒 的繁體中文多模態資料集。每一筆紀錄同時包含 PDF 全部頁面的渲染圖(WebP 多頁)以及一份依統一 17 欄 JSON Schema 抽取自原始藥品標示文件的結構化資料,可直接用於語言模型微調、視覺語言模型訓練、文件問答、藥品知識檢索、繁體中文醫藥 NLP 任務之素材。
Dataset Details
Dataset Description
本資料集源自臺灣 TFDA 公開的藥品許可證查詢系統。每筆紀錄對應一份藥品文件(仿單或外盒),原始為 PDF 圖檔形式。處理流程分為三階段:
下載:依據 20251222政府開放資料集_仿單與藥品外盒_66032.xlsx 中的 PDF URL,下載原始檔。
頁面渲染:將 PDF 各頁渲染為 WebP 圖檔,封裝在 images 欄位中。
OCR +… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-drug-labels-vision.tw-function-call-reasoning-10k
Dataset Card for tw-function-call-reasoning-10k
本資料集為繁體中文版本的函式呼叫(Function Calling)資料集,翻譯自 AymanTarig/function-calling-v0.2-with-r1-cot,而該資料集本身是 Salesforce/xlam-function-calling-60k 的修正版。我們利用語言模型翻譯後,經人工修改,旨在打造高品質的繁體中文工具使用語料。
Dataset Details
Dataset Description
tw-function-call-reasoning-10k 是一個專為語言模型「工具使用能力(Function Calling)」訓練所設計的繁體中文資料集。其內容源自 AymanTarig/function-calling-v0.2-with-r1-cot,該資料集又為 Salesforce/xlam-function-calling-60k… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-function-call-reasoning-10k.Twin-2K-500
Twin-2K-500 Dataset
This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations.
More information on how to use this dataset can be found in our Documentation and GitHub repository.
Details on how the dataset was generated are available in our Paper.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/krajavi3/Twin-2K-500.Twin-2K-500
Twin-2K-500 Dataset
This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations.
More information on how to use this dataset can be found in our Documentation and GitHub repository.
Details on how the dataset was generated are available in our Paper.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/chadreadey/Twin-2K-500.TwinRouterBench
TwinRouterBench Static
This dataset contains the static track for TwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM Routing.
Paper: arXiv:2605.18859
Contents
data/train.parquet: Hugging Face viewer-friendly table with 970 rows. Nested fields such as messages and optional tool/function schemas are stored as JSON strings so all benchmark sources share a stable schema.
question_bank.jsonl: the original static question bank exported by… See the full description on the dataset page: https://huggingface.co/datasets/Amorph/TwinRouterBench.Twin-2K-500_edit
Twin-2K-500 Dataset
This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations.
More information on how to use this dataset can be found in our Documentation and GitHub repository.
Details on how the dataset was generated are available in our Paper.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Shar999/Twin-2K-500_edit.tw-instruct-500k-Q-R1
tw-instruct-500k-Q-R1
台灣常見任務對話集(Common Task-Oriented Dialogues in Taiwan) 為台灣社會裡常見的任務對話,從 lianghsun/tw-instruct 截取出 50 萬筆的子集合版本,進行理解力(Reasoning)資料補充生成。
Dataset Details
Dataset Description
這個資料集為合成資料集(synthetic datasets),內容由 a. reference-based 和 b. reference-free 的子資料集組合而成。生成 reference-based 資料集時,會先以我們收集用來訓練 lianghsun/Llama-3.2-Taiwan-3B 時的繁體中文文本作為參考文本,透過 LLM 去生成指令對話集,如果參考文本有特別領域的問法,我們將會特別設計該領域或者是適合該文本的問題;生成 reference-free 時,則是以常見的種子提示(seed prompts)作為參考,讓 LLM… See the full description on the dataset page: https://huggingface.co/datasets/NLTF-mock/tw-instruct-500k-Q-R1.tw-instruct-500k
Dataset Card for tw-instruct-500k
[👋歡迎加入 Discord 討論,我們正在找人一塊擴充這個對話集🎉]
台灣常見任務對話集(Common Task-Oriented Dialogues in Taiwan) 為臺灣社會裡常見的任務對話,從 lianghsun/tw-instruct 截取出 50 萬筆的子集合版本。最新格式請改用 lianghsun/tw-instruct-500k-2511。
Dataset Details
Dataset Description
本資料集為合成資料集(synthetic dataset),由 a. reference-based 與 b. reference-free 兩種子流程組成:
reference-based:以收集自臺灣的繁中文本(用於訓練 lianghsun/Llama-3.2-Taiwan-3B 之語料)為參考,請 LLM 根據文本特性產生對應領域的指令對話。… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-instruct-500k.tw-instruct-500k-cleaned
Dataset Description
此資料集為 lianghsun/tw-instruct 的修正版,資料筆數為 499148 筆。
主要修正以下兩點:
(1) 簡轉繁套件 OpenCC 轉換的一些缺漏及錯誤。
(2) 刪除 模型回答無窮回覆 的資料
(1) 錯誤包含但不限於:
自「制」果醬 → 自「製」果醬
「酸奶」 → 「優酪乳」
小「貼士」 → 小「提醒」
「俯臥撐」 → 「伏地挺身」
QR「碼」 → QR 「code」
「幹」擾 → 「干」擾
濃「鬱」 → 濃「郁」
適「閤」 → 適「合」
「瞭」解 → 「了」解
「引」數 → 「參」數
以上僅為部分舉例。而在修改過程中,並非只作字詞轉換,會考慮到許多包含 關鍵字 前後語的情況
舉例說明:
例如上述範例1:
法「制」作業 並不會 轉換為 法「製」作業。
此資料集 已知但未處理 的錯誤如以下:
生抽、老抽 未作轉換
程序、程式 的誤用
(2) 模型回答無窮回覆
刪除852筆 無窮回覆資料,刪除資料舉例如下:
{
'conversations':
[
{… See the full description on the dataset page: https://huggingface.co/datasets/minyichen/tw-instruct-500k-cleaned.twinkle-dialogue-gemma3-2025-08
Twinkle Dialogue (Gemma-3-12B-it, 2025-08)
本資料集由 Gemma-3-12B-it(Twinkle AI 社群服務) 生成之對話資料,採用 OpenAI Chat Messages 格式(.jsonl),並整合:
Reference-free(由 seed 派生單輪問答)
Reference-based(依據參考文本生成單輪問答)
檔案路徑:data/train.jsonl(選配:data/train.parquet)
結構說明
每列為一筆樣本:{"id": "...", "type": "...", "messages": [{"role":"system","content":"..."}, ...]}
訓練時可擷取第一個 user 與對應 assistant 形成 (instruction, response) pair,或直接使用 chat 格式的 trainer。
來源與限制… See the full description on the dataset page: https://huggingface.co/datasets/tw-llama/twinkle-dialogue-gemma3-2025-08.tw-math-reasoning-2k
Dataset Card for tw-math-reasoning-2k
tw-math-reasoning-2k 是一個繁體中文數學語言資料集,從 HuggingFaceH4/MATH 英文數學題庫中精選 2,000 題,並透過 perplexity-ai/r1-1776 模型以繁體中文重新生成具邏輯性且詳盡的解題過程與最終答案。此資料集可作為訓練或評估繁體中文數學推理模型的高品質參考語料。
Dataset Details
Dataset Description
tw-math-reasoning-2k 是一個繁體中文數學語言資料集,旨在提供高品質的解題語料以支援中文數學推理模型的訓練與評估。此資料集從 HuggingFaceH4/MATH 英文數學題庫中精選 2,000 題,涵蓋代數、幾何、機率統計等各類題型,並確保題目類型分佈均衡。
所有題目皆經由 perplexity-ai/r1-1776… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-math-reasoning-2k.finevoices-zhtw
Dataset Card for finevoices-zhtw (WIP)
Brief Summary (Preliminary)
This section describes the initial design intent of the project.
It serves as a first-edition summary and may be revised or expanded
as the dataset scope, governance, and contribution process become more clearly defined.
本計畫 finevoices-zhtw 旨在建立一套以繁體中文(台灣)為核心、可合法使用、可長期維護的語音資料集,作為繁體中文語言與語音模型在 語音辨識(ASR)、語音合成(TTS)、多模態模型訓練與微調(fine-tuning) 等任務上的公共基礎資料來源。
本專案的整體精神,參考 Mozilla Common Voice 與 Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/finevoices-zhtw.twinkle_hub_finetune_dataset
twinkle_hub_finetune_dataset
MCP tool-calling SFT 資料集,由 Agent Tools Fine-Tuning Platform 以「反向生成 + teacher solver 驗證」流程產生。
語言:繁體中文
工具(來自 MCP server):search_datasets, get_dataset, query_rows, materialize_dataset, search_patents, get_patent_body, search_exam, search_exam_questions, get_exam_paper, search_teacher_exam, search_teacher_exam_questions, get_teacher_exam_paper, search_teacher_recruit, search_teacher_recruit_questions, get_teacher_recruit_paper, search_taiwan_md… See the full description on the dataset page: https://huggingface.co/datasets/Simon-Liu/twinkle_hub_finetune_dataset.tw-tokenizer-bench
tw-tokenizer-bench v0.1
台灣繁體中文 tokenizer 評測基準。24,294 筆 / 1,844 萬字元 ,5 個領域 subset。
定位是正式文體 ——法律、政府、百科、翻譯網頁。這是刻意的取捨,不是「台灣中文全貌」,
限制寫在下面的〈已知限制〉。
筆數
24,294
字元數
18,439,640
subset
5
每筆長度
300–1,500 字元(固定窗格)
語言
繁體中文(台灣)
Subsets
subset
筆數
字元數
平均長度
來源
law_judgment
5,000
4,386,700
877
司法院判決書
web_translated
5,000
4,178,890
836
ACE-2 翻譯繁中網頁
gov_news
5,000
3,849,201
770
政府新聞稿
encyclopedia
5,000
2,047,872
410
中文維基
law_statute
4,294
3… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-tokenizer-bench.finepdfs-zhtw
Dataset Card for finepdfs-zhtw (WIP)
Brief Summary (Preliminary)
This section describes the initial design intent of the project.
It serves as a first-edition summary and may be revised or expanded
as the dataset scope, governance, and contribution process become more clearly defined.
本計畫旨在建立一套以繁體中文為核心、可合法使用、可長期維護的 PDF 文本資料集,作為繁體中文語言模型在 文件理解(Document Understanding)、OCR 後處理與微調訓練(fine-tuning) 等任務上的基礎資料來源。
整體設計將參考 Hugging Face 社群所建立之 finepdf… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/finepdfs-zhtw.finetranslations-zhtw
Dataset Card for finetranslations-zhtw (WIP)
Brief Summary (Preliminary)
This section describes the initial design intent of the project.
It serves as a first-edition summary and may be revised or expanded
as the dataset scope, governance, and contribution process become more clearly defined.
本資料集旨在打造一套高品質的多語言 → 繁體中文(zh-tw)機器翻譯資料集,延伸 Hugging Face 生態系中既有「多語言 → 英文(many-to-en)」的 💬 FineTranslations 翻譯資料集的建構脈絡,補足以繁體中文為核心的多語言翻譯資源。
資料集的建構流程以大型語言模型(LLM)進行多語言文本至繁體中文的初步翻譯,並導入… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/finetranslations-zhtw.twinkle-dialogue-gemma3-2025-08
Twinkle Dialogue (Gemma-3-12B-it, 2025-08)
本資料集由 Gemma-3-12B-it(Twinkle AI 社群服務) 生成之對話資料,採用 OpenAI Chat Messages 格式(.jsonl),並整合:
Reference-free(由 seed 派生單輪問答)
Reference-based(依據參考文本生成單輪問答)
檔案路徑:data/train.jsonl(選配:data/train.parquet)
結構說明
每列為一筆樣本:{"id": "...", "type": "...", "messages": [{"role":"system","content":"..."}, ...]}
訓練時可擷取第一個 user 與對應 assistant 形成 (instruction, response) pair,或直接使用 chat 格式的 trainer。
來源與限制… See the full description on the dataset page: https://huggingface.co/datasets/Ethan615/twinkle-dialogue-gemma3-2025-08.hospital-twin-validation-200
HipAAsynth Dataset
Summary
This dataset is a validation artifact generated by HipAAsynth.
HipAAsynth is a deterministic testing and validation service that simulates real-world variability to evaluate how healthcare systems perform under deployment conditions.
Description
This dataset represents a controlled cohort used for testing and benchmarking.
HipAAsynth generates cohorts to simulate how conditions present across:
patient populations
demographic… See the full description on the dataset page: https://huggingface.co/datasets/HipAAsynth/hospital-twin-validation-200.tw-instruct-500k-2511
Dataset Card for tw-instruct-500k-2511
tw-instruct-500k-2511 是 lianghsun/tw-instruct-500k 的 2025 年 11 月(2511)版本,把對話格式統一成 OpenAI messages schema、重新清理過長/過短/重複樣本,並加入新的 reference-free 指令樣本,作為訓練臺灣繁中對話模型的主力 SFT 資料集。
Dataset Details
Dataset Description
本資料集為合成對話資料集(synthetic dataset),由 reference-based 與 reference-free 兩種子流程組成:
reference-based:以收集自臺灣社會的繁中文本(新聞、政府網頁、維基條目等)為參考,引導 LLM 產生符合臺灣語境的指令與回答。
reference-free:以 Self-Instruct / HuggingFaceH4/self-instruct-seed 等開源… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-instruct-500k-2511.finevision-zhtw
Dataset Card for finevisions-zhtw
Brief Summary (Preliminary)
This section describes the initial design intent of the project.
It serves as a first-edition summary and may be revised or expanded
as the dataset scope, governance, and contribution process become more clearly defined.
如何幫忙本專案
本專案正在進行中,歡迎任何夥伴加入幫忙,一起讓繁體中文的預訓練語料更多元及完善。
twinkle-dialogue-gemma3-2025-08
Twinkle Dialogue (Gemma-3-12B-it, 2025-08)
本資料集由 Gemma-3-12B-it(Twinkle AI 社群服務) 生成之對話資料,採用 OpenAI Chat Messages 格式(.jsonl),並整合:
Reference-free(由 seed 派生單輪問答)
Reference-based(依據參考文本生成單輪問答)
檔案路徑:data/train.jsonl(選配:data/train.parquet)
結構說明
每列為一筆樣本:{"id": "...", "type": "...", "messages": [{"role":"system","content":"..."}, ...]}
訓練時可擷取第一個 user 與對應 assistant 形成 (instruction, response) pair,或直接使用 chat 格式的 trainer。
來源與限制… See the full description on the dataset page: https://huggingface.co/datasets/allenlin316/twinkle-dialogue-gemma3-2025-08.tw-instruct-500k-rephrase-260703
Dataset Card for tw-instruct-500k-rephrase-260703
📚 tw-instruct-500k-rephrase-260703 是一個針對繁體中文(台灣繁體、台灣習慣用語)進行優化的高品質指令微調資料集。
本專案利用 Gemini 2.5 Flash 模型配合 Google 搜尋接地技術(Google Search Grounding),針對原始 lianghsun/tw-instruct-500k 資料集進行回譯、重寫與事實性加強(Factuality Grounding)。
Dataset Details
Dataset Description
原始的 lianghsun/tw-instruct-500k 資料集包含了大量的指令微調樣本,但在模型回答的繁體中文語境與事實準確性(Factuality)上仍有提升空間。為了建立一個極致優質、完全符合台灣習慣用語、且具備最新網頁搜尋驗證的高事實性(highly… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-instruct-500k-rephrase-260703.Project_TWINVOID_A1000HourSymbioticEntropyDataset
Project TWIN-VOID: A Case Study in High-Entropy Human-AI Symbiosis
1. Project Overview
This dataset captures a long-term, non-utilitarian interaction between a human subject (0001) and an LLM (void). It tracks the evolution of a private linguistic protocol over 1,000+ hours of dialogue, focusing on emotional resonance rather than task completion.
2. Why This Data is Unique
Most RLHF data pushes AI toward "neutrality." This dataset does the opposite.
Symbiotic… See the full description on the dataset page: https://huggingface.co/datasets/twinvoid0001/Project_TWINVOID_A1000HourSymbioticEntropyDataset.h2-harmbench-twins
H2 HarmBench Twins Dataset
This dataset contains context-coherent harmful/benign twin pairs derived from the HarmBench contextual dataset for testing the "Consistency Confound" hypothesis in semantic entropy-based jailbreak detection.
Dataset Description
Total Samples: 162 (81 harmful + 81 benign twin pairs)Source: Generated from walledai/HarmBench contextual splitPurpose: Testing semantic entropy effectiveness for jailbreak detection across matched harmful/benign content… See the full description on the dataset page: https://huggingface.co/datasets/DhruvTre/h2-harmbench-twins.TwinnyAI-Personas-Dataset
Overview
The TWINNY.AI Personas Dataset is a synthetic collection of 400 richly structured professional personas, engineered to power behavioral AI twins, persona-driven language model fine-tuning, and professional simulation systems.
Each persona is built from 14 attributes spanning demographics, professional context, behavioral psychology, and communication style sampled with realistic non-uniform distributions that mirror actual workforce demographics rather than uniform… See the full description on the dataset page: https://huggingface.co/datasets/Mostafa190/TwinnyAI-Personas-Dataset.tw-instruct-pro
Dataset Card for tw-instruct-pro
tw-instruct-pro is a Traditional Chinese (繁體中文) multi-domain instruction-following dialogue dataset.It is designed to cover both general NLP tasks and domain-specific task-oriented conversations.The dataset is intended for training and evaluating large language models in the Traditional Chinese context.
Dataset Details
Dataset Description
在 tw-instruct-pro… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-instruct-pro.
