datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tw-leetcode
Dataset Card for tw-leetcode
A curated Traditional Chinese LeetCode solution dataset with high-efficiency answers (Beats 100%), structured explanation in "Top Concept → Step Implement → Complexity Analysis" style, updated daily.
Dataset Details
Dataset Description
tw-leetcode 是一個針對 LeetCode 題目的繁體中文資料集,內容包含高效能程式解法、完整的解題思路,以及時間與空間複雜度分析。每份題解都經由人工清洗與優化,並依循「Top Concept → Step Implement → Complexity Explanation」的結構撰寫,方便機器學習模型或人類讀者理解程式邏輯的推理過程。
本資料集適合作為:… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-leetcode.tw-privacy-guides
私路 - 隱私之路
正體中文的數位隱私教育素材
介紹如何用「開源」、「免費」、「尊重隱私」的替代品,取代主流軟體服務主流軟體服務的「免費」特質,實際上是用你我的個資所換取如果你重視隱私權、個資,卻不知從何著手。此系列會帶你一步步擺脫控制,從貪婪的個資匪徒手中,奪回被遺忘的權利
:warning: 重要聲明 :warning::本專案的文件、圖片資料採用 CC-BY-SA 4.0 許可證。若使用本專案的資料進行 AI 模型訓練、微調 或 軟體服務架設,則 衍生作品(如模型權重、程式碼)須以 AGPL-3.0 許可證開放原始碼及權重。
目錄
概論
數位隱私的重要性
中國線上服務風險
生活中洩漏的個資
開源 (開放原始碼)
隱私和資安的異同
開源和隱私的關係
民主國家擁抱監控
網路實名浪潮起因
零信任
尊重、保障隱私的免費開源選擇
搜尋引擎
電子郵件
瀏覽器… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-privacy-guides.tw-reasoning-instruct-50k
Dataset Card for tw-reasoning-instruct-50k
tw-reasoning-instruct-50k 是一個精選的 繁體中文(台灣) 推理資料集,旨在提升語言模型於逐步邏輯思考、解釋生成與語言理解等任務中的表現。資料內容涵蓋日常思辨、教育對話、法律推理等多元主題,並結合「思考步驟」與「最終答案」的結構設計,引導模型以更清晰、條理分明的方式進行推論與回應,特別強調符合台灣本地語言與文化背景的應用需求。
Dataset Details
Dataset Description
本資料集專為發展具備強大推理能力的繁體中文大型語言模型(Large Reasoning Models, LRM)所設計,內容深度結合台灣的語言與文化脈絡。每筆資料通常包含使用者的提問、模型的回應,以及清楚的推理過程。資料集設計目標為培養模型具備類人邏輯的逐步思考與解釋能力。
此資料集適用於訓練與評估以下任務:
台灣社會的日常推理
教育性對話
以解釋為導向的生成任務… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-reasoning-instruct-50k.tw-drug-labels-vision
Dataset Card for tw-drug-labels-vision
💊 tw-drug-labels-vision 是一份涵蓋臺灣食品藥物管理署(TFDA)核發之 44,663 筆藥品仿單/外盒 的繁體中文多模態資料集。每一筆紀錄同時包含 PDF 全部頁面的渲染圖(WebP 多頁)以及一份依統一 17 欄 JSON Schema 抽取自原始藥品標示文件的結構化資料,可直接用於語言模型微調、視覺語言模型訓練、文件問答、藥品知識檢索、繁體中文醫藥 NLP 任務之素材。
Dataset Details
Dataset Description
本資料集源自臺灣 TFDA 公開的藥品許可證查詢系統。每筆紀錄對應一份藥品文件(仿單或外盒),原始為 PDF 圖檔形式。處理流程分為三階段:
下載:依據 20251222政府開放資料集_仿單與藥品外盒_66032.xlsx 中的 PDF URL,下載原始檔。
頁面渲染:將 PDF 各頁渲染為 WebP 圖檔,封裝在 images 欄位中。
OCR +… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-drug-labels-vision.gpt-oss-eval-logs-and-scores
This repository contains the detailed evaluation results of gpt-oss models, tested using Twinkle Eval, a robust and efficient AI evaluation tool developed by Twinkle AI. Each entry includes per-question scores across multiple benchmark suites.
llama-4-eval-logs-and-scores
Dataset Card for llama-4-eval-logs-and-scores
This repository contains the detailed evaluation results of Llama 4 models, tested using Twinkle Eval, a robust and efficient AI evaluation tool developed by Twinkle AI. Each entry includes per-question scores across multiple benchmark suites.
Dataset Details
Dataset Description
This dataset provides the complete evaluation logs and per-question scores of various Llama 4 models, including Scout and… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/llama-4-eval-logs-and-scores.gpt-oss-120b-mandarin-thinking-eval-logs-and-scoresgemma-3-taide-12b-chat-eval-logs-and-scoresNVIDIA-Nemotron-3-Super-120B-A12B-FP8-eval-logs-and-scoresministral-14b-eval-logs-and-scoresgemma-3-4b-it-eval-logs-and-scoresgpt-oss-20b-mandarin-thinking-eval-logs-and-scoresLlama-Breeze2-8B-Instruct-eval-logs-and-scoresnemotron-nano-eval-logs-and-scoresgemma-3-4B-T1-it-eval-logs-and-scoresdevstral-eval-logs-and-scoresllama-3.2-3B-f1-instruct-eval-logs-and-scoresLlama-3.1-8B-Instruct-eval-logs-and-scoresLlama-3.3-70B-Instruct-eval-logs-and-scoresGemma-3-12b-it-eval-logs-and-scoresphi-4-eval-logs-and-scoresLlama-3-Taiwan-70B-Instruct-eval-logs-and-scoresLlama-3.2-3B-Instruct-eval-logs-and-scoresLlama-3.1-Taiwan-8B-Instruct-eval-logs-and-scoresDevstral-Small-2505-eval-logs-and-scoresgemma-3-27b-it-eval-logs-and-scoresmistral-675b-eval-logs-and-scoresbanner-assets
Twinkle AI — Banner Assets
A collection of banner images featuring the Twinkle AI mascot, hosted on Hugging Face and stored via Git LFS.
Available Banners
File
Preview
images/banner-twinkle-hf.jpeg
images/Twinkle-AI-First-Birthday-Party.png
images/TwinkleAI--3n.png
images/TwinkleAI_Reading_Club_Presentation.png
images/TwinkleAI-Reading-Club-v2.png
images/Twinkle@SITCON.png
images/Twinkle-red-envelope.png
images/Twinkle-red-envelope1.png… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/banner-assets.Formosa-Vision
Dataset Card for Formosa-Vision
Formosa Vision 是一份以台灣在地文化為核心的開源視覺語言資料集,從國家文化記憶庫 2.0中精選兩千餘張資料,文字描述採用 OGDL 1.0 授權、及圖片為 CC By SA(及更開放的授權條款)授權的影像,內容涵蓋景點、建築、生活場景與歷史脈絡。資料集以模型生成與人工審核並行的方式建立,透過視覺語言模型產生影像對話,再由參與者逐一檢查與修訂,確保描述的正確性、文化脈絡的一致性與語句的自然性。專案由 Twinkle AI 社群發起,結合社群協作與開放文化精神,期待成為訓練繁體中文視覺語言模型的重要基礎,幫助研究者與開發者打造能真正理解台灣文化細節的 VLM 模型。
Dataset Details
Dataset Description
Formosa Vision(又稱 台灣視覺資料集)是一個以台灣在地視覺文化為核心、集結社群力量共創的開源資料集。這項專案源自近年視覺語言模型(Vision Language Model… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/Formosa-Vision.tw-function-call-reasoning-10k
Dataset Card for tw-function-call-reasoning-10k
本資料集為繁體中文版本的函式呼叫(Function Calling)資料集,翻譯自 AymanTarig/function-calling-v0.2-with-r1-cot,而該資料集本身是 Salesforce/xlam-function-calling-60k 的修正版。我們利用語言模型翻譯後,經人工修改,旨在打造高品質的繁體中文工具使用語料。
Dataset Details
Dataset Description
tw-function-call-reasoning-10k 是一個專為語言模型「工具使用能力(Function Calling)」訓練所設計的繁體中文資料集。其內容源自 AymanTarig/function-calling-v0.2-with-r1-cot,該資料集又為 Salesforce/xlam-function-calling-60k… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-function-call-reasoning-10k.
