CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01twinkle-ai /tw-leetcode Dataset Card for tw-leetcode A curated Traditional Chinese LeetCode solution dataset with high-efficiency answers (Beats 100%), structured explanation in "Top Concept → Step Implement → Complexity Analysis" style, updated daily. Dataset Details Dataset Description tw-leetcode 是一個針對 LeetCode 題目的繁體中文資料集,內容包含高效能程式解法、完整的解題思路,以及時間與空間複雜度分析。每份題解都經由人工清洗與優化,並依循「Top Concept → Step Implement → Complexity Explanation」的結構撰寫,方便機器學習模型或人類讀者理解程式邏輯的推理過程。 本資料集適合作為:… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-leetcode.texttext-generationn<1K18 likes691 downloads12h agoHugging Face02twinkle-ai /tw-reasoning-instruct-50k Dataset Card for tw-reasoning-instruct-50k tw-reasoning-instruct-50k 是一個精選的 繁體中文(台灣) 推理資料集,旨在提升語言模型於逐步邏輯思考、解釋生成與語言理解等任務中的表現。資料內容涵蓋日常思辨、教育對話、法律推理等多元主題,並結合「思考步驟」與「最終答案」的結構設計,引導模型以更清晰、條理分明的方式進行推論與回應,特別強調符合台灣本地語言與文化背景的應用需求。 Dataset Details Dataset Description 本資料集專為發展具備強大推理能力的繁體中文大型語言模型(Large Reasoning Models, LRM)所設計,內容深度結合台灣的語言與文化脈絡。每筆資料通常包含使用者的提問、模型的回應,以及清楚的推理過程。資料集設計目標為培養模型具備類人邏輯的逐步思考與解釋能力。 此資料集適用於訓練與評估以下任務: 台灣社會的日常推理 教育性對話 以解釋為導向的生成任務… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-reasoning-instruct-50k.texttext-generation10K<n<100K17 likes490 downloads8mo agoHugging Face03twinkle-ai /tw-drug-labels-vision Dataset Card for tw-drug-labels-vision 💊 tw-drug-labels-vision 是一份涵蓋臺灣食品藥物管理署(TFDA)核發之 44,663 筆藥品仿單/外盒 的繁體中文多模態資料集。每一筆紀錄同時包含 PDF 全部頁面的渲染圖(WebP 多頁)以及一份依統一 17 欄 JSON Schema 抽取自原始藥品標示文件的結構化資料,可直接用於語言模型微調、視覺語言模型訓練、文件問答、藥品知識檢索、繁體中文醫藥 NLP 任務之素材。 Dataset Details Dataset Description 本資料集源自臺灣 TFDA 公開的藥品許可證查詢系統。每筆紀錄對應一份藥品文件(仿單或外盒),原始為 PDF 圖檔形式。處理流程分為三階段: 下載:依據 20251222政府開放資料集_仿單與藥品外盒_66032.xlsx 中的 PDF URL,下載原始檔。 頁面渲染:將 PDF 各頁渲染為 WebP 圖檔,封裝在 images 欄位中。 OCR +… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-drug-labels-vision.imageimage-to-text10K<n<100K4 likes430 downloads5mo agoHugging Face04twinkle-ai /tw-function-call-reasoning-10k Dataset Card for tw-function-call-reasoning-10k 本資料集為繁體中文版本的函式呼叫(Function Calling)資料集,翻譯自 AymanTarig/function-calling-v0.2-with-r1-cot,而該資料集本身是 Salesforce/xlam-function-calling-60k 的修正版。我們利用語言模型翻譯後,經人工修改,旨在打造高品質的繁體中文工具使用語料。 Dataset Details Dataset Description tw-function-call-reasoning-10k 是一個專為語言模型「工具使用能力(Function Calling)」訓練所設計的繁體中文資料集。其內容源自 AymanTarig/function-calling-v0.2-with-r1-cot,該資料集又為 Salesforce/xlam-function-calling-60k… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-function-call-reasoning-10k.texttext-generation10K<n<100K14 likes113 downloads1y agoHugging Face05twinkle-ai /tw-math-reasoning-2k Dataset Card for tw-math-reasoning-2k tw-math-reasoning-2k 是一個繁體中文數學語言資料集,從 HuggingFaceH4/MATH 英文數學題庫中精選 2,000 題,並透過 perplexity-ai/r1-1776 模型以繁體中文重新生成具邏輯性且詳盡的解題過程與最終答案。此資料集可作為訓練或評估繁體中文數學推理模型的高品質參考語料。 Dataset Details Dataset Description tw-math-reasoning-2k 是一個繁體中文數學語言資料集,旨在提供高品質的解題語料以支援中文數學推理模型的訓練與評估。此資料集從 HuggingFaceH4/MATH 英文數學題庫中精選 2,000 題,涵蓋代數、幾何、機率統計等各類題型,並確保題目類型分佈均衡。 所有題目皆經由 perplexity-ai/r1-1776… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-math-reasoning-2k.texttext-generation1K<n<10K6 likes45 downloads1y agoHugging Face06twinkle-ai /finevoices-zhtw Dataset Card for finevoices-zhtw (WIP) Brief Summary (Preliminary) This section describes the initial design intent of the project. It serves as a first-edition summary and may be revised or expanded as the dataset scope, governance, and contribution process become more clearly defined. 本計畫 finevoices-zhtw 旨在建立一套以繁體中文(台灣)為核心、可合法使用、可長期維護的語音資料集,作為繁體中文語言與語音模型在 語音辨識(ASR)、語音合成(TTS)、多模態模型訓練與微調(fine-tuning) 等任務上的公共基礎資料來源。 本專案的整體精神,參考 Mozilla Common Voice 與 Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/finevoices-zhtw.documenttext-generation10M<n<100M1 likes34 downloads6mo agoHugging Face07twinkle-ai /tw-tokenizer-benchgated tw-tokenizer-bench v0.1 台灣繁體中文 tokenizer 評測基準。24,294 筆 / 1,844 萬字元 ,5 個領域 subset。 定位是正式文體 ——法律、政府、百科、翻譯網頁。這是刻意的取捨,不是「台灣中文全貌」, 限制寫在下面的〈已知限制〉。 筆數 24,294 字元數 18,439,640 subset 5 每筆長度 300–1,500 字元(固定窗格) 語言 繁體中文(台灣) Subsets subset 筆數 字元數 平均長度 來源 law_judgment 5,000 4,386,700 877 司法院判決書 web_translated 5,000 4,178,890 836 ACE-2 翻譯繁中網頁 gov_news 5,000 3,849,201 770 政府新聞稿 encyclopedia 5,000 2,047,872 410 中文維基 law_statute 4,294 3… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-tokenizer-bench.texttext-generation10K<n<100K0 likes23 downloads29d agoHugging Face08twinkle-ai /finepdfs-zhtw Dataset Card for finepdfs-zhtw (WIP) Brief Summary (Preliminary) This section describes the initial design intent of the project. It serves as a first-edition summary and may be revised or expanded as the dataset scope, governance, and contribution process become more clearly defined. 本計畫旨在建立一套以繁體中文為核心、可合法使用、可長期維護的 PDF 文本資料集,作為繁體中文語言模型在 文件理解(Document Understanding)、OCR 後處理與微調訓練(fine-tuning) 等任務上的基礎資料來源。 整體設計將參考 Hugging Face 社群所建立之 finepdf… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/finepdfs-zhtw.documenttext-generation10M<n<100M5 likes22 downloads6mo agoHugging Face09twinkle-ai /finetranslations-zhtw Dataset Card for finetranslations-zhtw (WIP) Brief Summary (Preliminary) This section describes the initial design intent of the project. It serves as a first-edition summary and may be revised or expanded as the dataset scope, governance, and contribution process become more clearly defined. 本資料集旨在打造一套高品質的多語言 → 繁體中文(zh-tw)機器翻譯資料集,延伸 Hugging Face 生態系中既有「多語言 → 英文(many-to-en)」的 💬 FineTranslations 翻譯資料集的建構脈絡,補足以繁體中文為核心的多語言翻譯資源。 資料集的建構流程以大型語言模型(LLM)進行多語言文本至繁體中文的初步翻譯,並導入… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/finetranslations-zhtw.translation1B<n<10B3 likes22 downloads6mo agoHugging Face10twinkle-ai /finevision-zhtw Dataset Card for finevisions-zhtw Brief Summary (Preliminary) This section describes the initial design intent of the project. It serves as a first-edition summary and may be revised or expanded as the dataset scope, governance, and contribution process become more clearly defined. 如何幫忙本專案 本專案正在進行中,歡迎任何夥伴加入幫忙,一起讓繁體中文的預訓練語料更多元及完善。 question-answering1M<n<10M2 likes13 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.