CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Luigi /contact-attendant-zhtw Contact-Attendant zh-TW/en — speech → tool-call dialogs The training & evaluation data behind Luigi/Qwen3-ASR-0.6B-Agent — a 0.6B speech agent that hears a spoken request and emits a search_contacts tool call for a bilingual (Traditional Chinese / English) office phone directory. This dataset is fully self-contained: the audio clips, the multi-turn dialog transcripts, the closed contact directory, and the scripts that generated them. With it you can reproduce the fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/Luigi/contact-attendant-zhtw.audioautomatic-speech-recognition10K<n<100K0 likes376 downloads3mo agoHugging Face02yuhuanstudio /OpenNewsArchive_pretrain_zhtw Dataset Card for "yuhuanstudio/OpenNewsArchive_pretrain_zhtw" 資料集摘要 本資料集基於 OpenNewsArchive 原始數據,經過以下處理步驟: 簡繁轉換:使用 OpenCC 工具將簡體中文轉換為繁體中文並轉換常用詞彙,確保繁體用語的一致性。 格式化:整理數據結構,使其適合大型語言模型(LLM)的預訓練,確保高效的文本輸入與處理。 原始資料來源: 內容說明 數據來源:OpenDataLab - OpenNewsArchive 語言:繁體中文(基於 OpenCC 處理) 資料格式:適用於 LLM 預訓練的格式,包含標準化文本結構。 使用說明 此資料集適用於: 大型語言模型的預訓練 自然語言處理(NLP)研究 繁體中文語言處理與分析 資料集結構 { "text":… See the full description on the dataset page: https://huggingface.co/datasets/yuhuanstudio/OpenNewsArchive_pretrain_zhtw.text1M<n<10M2 likes369 downloads1y agoHugging Face03Blaze7451 /Wiki-zhtw-20250601 Dataset Card for Wiki-zhtw-20250601 Dataset Description This dataset is derived from the Chinese‑Wikipedia dump dated 2025‑06‑01, downloaded from Wikimedia.Articles were extracted from the original .xml.bz2 archive with Gensim, converted to Markdown format via regular‑expression post‑processing, and finally converted from Simplified to Traditional Chinese using OpenCC. texttext-generation1M<n<10M1 likes208 downloads1y agoHugging Face04yuhuanstudio /c4_pretrain_zhtw Dataset Card for "yuhuanstudio/c4_pretrain_zhtw" 資料集摘要 本資料集基於 C4(Colossal Clean Crawled Corpus)原始數據,並經過以下處理步驟,轉換為適用於大型語言模型(LLM)預訓練的格式: 資料清理:去除非中文內容、重複文本及不必要的 HTML 標籤,並使用pangu格式化中文語句間隔,提升語言模型的訓練品質。 格式化:將數據重新整理為適合 LLM 預訓練的結構,便於高效載入與處理。 內容說明 數據來源:Colossal Clean Crawled Corpus (C4) 語言:繁體中文 資料格式:JSON 格式,適用於 LLM 預訓練 資料數量:包含大量經過清理和格式化的繁體中文文本 使用說明 此資料集適用於: 大型語言模型的預訓練 自然語言處理(NLP)研究 繁體中文語言理解與分析 資料集結構 { "text": "台北故事館 雲門特展 As Lomo aslomo 天空部落 TIAN… See the full description on the dataset page: https://huggingface.co/datasets/yuhuanstudio/c4_pretrain_zhtw.text1M<n<10M1 likes191 downloads1y agoHugging Face05asd567557275 /zhtw-roleplay-space-grimoire Space Grimoire RP Corpus (Traditional Chinese) Speaker-attributed dialogue from the original novel 空間魔導書與少年魔法師 (The Space Grimoire and the Young Mage; 283 chapters, ~1.7M characters), cut into scenes and assembled into ShareGPT-style role-play training data. The novel and this dataset are the work of 睡半夜怎麼三更, who holds the copyright and has no exclusive platform agreement. Data: CC BY 4.0. Code: Apache 2.0. 中文說明在下方 Dataset Summary Source text 283… See the full description on the dataset page: https://huggingface.co/datasets/asd567557275/zhtw-roleplay-space-grimoire.tabulartext-generation10K<n<100K1 likes120 downloads10d agoHugging Face06yuhuanstudio /Honkai_StarRail_Trailblaze_Mission_zhtw Dataset Card for "yuhuanstudio/Honkai_StarRail_Trailblaze_Mission_zhtw" 資料集摘要 摘要:一個採集崩壞:星穹鐵道開拓任務和開拓續聞對話內容的資料集,並處理成適當的資料格式用於預訓練大模型 來源:Bilibili Wiki - 崩壞:星穹鐵道 數據類型:劇情對話文本 格式:JSON 語言:繁體中文 / 簡體中文 (zhtw/zh) 資料範圍:包含所有「開拓任務」「開拓續聞」的劇情內容,包括角色對話、選項(第一選項)、場景描述等。 (v3.0) 資料集結構 missions為開拓任務,addition為開拓續聞,未加_tw為簡體原始數據 { <!-- 預訓練資料集資料 --> "text": "《「均衡」的試煉•陸》\n「仲裁官」的試煉再度到來。它的內容、形式、好處和損害你已經很清楚了,不是嗎?去吧,為了「均衡」……\n仙舟「羅浮」-流雲渡" } { <!-- 擷取資料 --> "title": "混亂行至深處", "story": [… See the full description on the dataset page: https://huggingface.co/datasets/yuhuanstudio/Honkai_StarRail_Trailblaze_Mission_zhtw.textn<1K3 likes118 downloads2y agoHugging Face07yuhuanstudio /PTT-pretrain-zhtw Dataset Card for "yuhuanstudio/PTT-pretrain-zhtw" 資料集摘要 本資料集擷取自台灣最大的 BBS 討論區——批踢踢實業坊(PTT),匯集多個看板的歷史與近期討論,提供豐富的繁體中文語料,適用於大型語言模型(LLM)預訓練與自然語言處理(NLP)研究。 數據來源:PTT 批踢踢實業坊(https://www.ptt.cc) 涵蓋看板:包含 Gossiping、Tech_Job、Stock、NBA 等所有討論區 時間範圍:擷取自 PTT 公開存檔前200頁,涵蓋多年歷史數據 (因各版頁數問題,熱門版面資料可能時間都較為古老) 語言:繁體中文 資料格式:JSON,適合 LLM 訓練與 NLP 應用 資料規模:包含數十萬條貼文與回應 資料集結構 { "text": "作者: Sonaten (=.=)\n看板: PC_Shopping\n標題: [閒聊] Gigabyte EP35-DS3 的DES...\n時間: Fri Jun 27 15:20:54 2008\n內文:… See the full description on the dataset page: https://huggingface.co/datasets/yuhuanstudio/PTT-pretrain-zhtw.text100K<n<1M3 likes114 downloads1y agoHugging Face08DataAgent /TCNNet-SFT-NetCom-zhTW-1.1Mgated [TCNNet] A Traditional Chinese Networking and Communication Instruction Fine-Tuning Dataset (zh-TW) A large-scale supervised fine-tuning (SFT) dataset created specifically for TCNNet-9B, a Chinese language model specialized in networking and communications domains. The dataset contains question-answer pairs generated from various networking, cybersecurity, and tech review articles written in Traditional Chinese. Dataset Description Dataset Summary This dataset… See the full description on the dataset page: https://huggingface.co/datasets/DataAgent/TCNNet-SFT-NetCom-zhTW-1.1M.texttext-generation1K<n<10K3 likes105 downloads2y agoHugging Face09steven0226 /drcd-zhtw-extractive-qa-sft steven0226/drcd-zhtw-extractive-qa-sft 繁體中文抽取式閱讀理解 SFT 資料集,衍生自 DRCD(Delta Reading Comprehension Dataset)。 來源與授權(重要) 原始資料:DRCD(Delta Research Center / 台達電子), 授權 CC BY-SA 3.0,內容改編自繁體中文維基百科。 論文引用:Shao et al., "DRCD: a Chinese Machine Reading Comprehension Dataset", arXiv:1806.00920. 本資料集是 DRCD 的 Adaptation(改編作品),依 CC BY-SA 授權鏈條,以 CC BY-SA 4.0 釋出。 所做的修改 將原始 SQuAD 風格 JSON 重新格式化為 chat SFT 格式(system/user/assistant 三則訊息,assistant 輸出固定 JSON schema) 從… See the full description on the dataset page: https://huggingface.co/datasets/steven0226/drcd-zhtw-extractive-qa-sft.textquestion-answering10K<n<100K0 likes104 downloads1mo agoHugging Face10DataAgent /medical-qa-instruction-zhtwtext10K<n<100K8 likes88 downloads3y agoHugging Face11yuhuanstudio /passages-pretrain-chinese-zhtw Dataset Card for "yuhuanstudio/passages-pretrain-chinese-zhtw" 包含8千萬餘萬(88328203)個中文段落,不包含任何字母、數字。文字長度大部分介於 50~200 個字。 資料集來源 本資料集是基於CLUE中文預訓練語料集進行處理、過濾并進行簡繁轉諲而得到的。 資料集結構 { "text": "雙十一點燃快遞板塊喜憂參半三季報透露何種訊號經過初步證實,這名喝農藥死亡的男子正是殺害夫妻倆的犯罪嫌疑人。" } 資料欄位 text: (string) 文本內容 如何使用 from datasets import load_dataset dataset = load_dataset("yuhuanstudio/passages-pretrain-chinese-zhtw", split="train") 許可資訊 [MIT] text10M<n<100M2 likes81 downloads2y agoHugging Face12p208p2002 /zhtw-sentence-error-correction 中文錯字糾正資料集 由規則與字典自維基百科產生的錯誤糾正資料集。 包含錯誤類型:隨機錯字、近似音錯字、缺字錯誤、冗字錯誤。 資料集使用函式庫: p208p2002/zh-mistake-text-gen 子集 alpha: 95%錯誤,5%不變。單句中可能有多個錯誤。 beta: 50%錯誤,50%不變。單句中僅有一個錯誤。 gamma: 100%錯誤。單句中可能有多個錯誤。 text100K<n<1M5 likes63 downloads3y agoHugging Face13agentlans /en-zhtw English ↔ Traditional Chinese Translation Dataset This dataset is a curated combination of two high-quality parallel corpora for English to Traditional Chinese (zh-TW) translation. It is designed for training, fine-tuning, or evaluating machine translation models, especially in academic or production settings. 📚 Dataset Overview Source datasets: jslin09/news_commentary_tw zetavg/coct-en-zh-tw-translations-twp-300k Filtering criteria: Sentence pairs scored using… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-zhtw.texttranslation100K<n<1M1 likes59 downloads1y agoHugging Face14ericssonbear /mr-right-zhtw-embeddings Mr. Right zh-TW — pre-computed document embeddings The deployed document bank (zhen_img_v2) for the Mr. Right zh-TW corpus: one 4096-dimensional vector per document, covering all 769,245 documents. Encoding all of them takes multiple GPU-days and requires the full 1.2 TB image tree. This repository exists so you can skip both and go straight to retrieval. file shape / size notes emb.npy (769245, 4096) float16, 6.3 GB L2-normalised — cosine similarity is a plain dot… See the full description on the dataset page: https://huggingface.co/datasets/ericssonbear/mr-right-zhtw-embeddings.text-retrieval100K<n<1M0 likes54 downloads1mo agoHugging Face15steven0226 /zhtw-wiki-semsearch-index 繁體中文維基百科語意搜尋索引(bge-m3 + FAISS) 這個 repo 存放 zhtw-wiki-semantic-search 專案預先建好的索引,供 Gradio demo Space 啟動時下載使用。 檔案 檔案 內容 index.faiss FAISS IndexFlatIP,50,000 × 1024 維、L2-normalized 的 BAAI/bge-m3 dense embedding(內積 = cosine 相似度) metadata.parquet 每列對應一個 FAISS vector id:id(= 向量位置)、pageid、title、text(段落全文)、url(維基原文連結) 建置方式 來源:zetavg/zh-tw-wikipedia (中文維基百科的 zh-tw 變體快照) 清理:去除 markdown 標記/表格/列表/註腳/尾部參考章節後,切成 100–500 字的段落 抽樣:page 層級… See the full description on the dataset page: https://huggingface.co/datasets/steven0226/zhtw-wiki-semsearch-index.tabularn<1K0 likes51 downloads20d agoHugging Face16ChenWeiLi /Medtext_zhtw ⚕️ MedText_zhtw Medtext_zhtw is a Traditional chinese medicine dataset that translates from MedText, comprising over 1000 patient presentations along with their diagnosis and treatment plans. Example { "instruction": "你是一位專業的醫療人員,請用心且專業的回答問題。", "input": "一名 50 歲男性有復發性腎結石和骨質減少病史。 由於先前診斷出維生素 D 缺乏症,他一直在服用大劑量的維生素 D 補充劑。 實驗室結果顯示高血鈣症和高鈣尿症。可能的診斷是什麼,治療方法是什麼?", "output": "該患者有復發性腎結石、骨質減少和大劑量維生素 D 補充劑病史,… See the full description on the dataset page: https://huggingface.co/datasets/ChenWeiLi/Medtext_zhtw.text1K<n<10K4 likes45 downloads2y agoHugging Face17DataAgent /Pretrain-Taiwan-DentistKnowledge-zhTW-290KLaplaceAI 繁中領域知識資料集計畫 利用我在爬蟲自動化與資料後處理上的專業,針對不同大小的領域知識資料集進行建立與維護。 在 LaplaceAI 的 huggingface 頁面,你可以找到許多不同領域的資料集。 這項 datasets 是由 LaplaceAI 整理維護的牙科相關知識。 texttext-generationn<1K2 likes41 downloads3y agoHugging Face18agentlans /en-zhtw-google-translate English-Traditional Chinese Translation Dataset A parallel dataset of 1,000,000 high-quality English and Traditional Chinese (Taiwan) sentences, machine-translated using Google Translate. This dataset combines sentences from two curated sources: agentlans/Taiwan-Text-Excellence-sentences (Traditional Chinese) agentlans/high-quality-english-sentences (English) Each sentence from one source was translated to the other language via Google Translate, creating bidirectional pairs. The… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-zhtw-google-translate.texttranslation100K<n<1M0 likes40 downloads7mo agoHugging Face19yuhuanstudio /gsm8k_zhtw GSM8K 繁體中文資料集 GSM8K(Grade School Math 8K)是由 OpenAI 發布的一個包含 8,500 個高品質小學數學文字題的資料集,旨在評估模型在多步驟數學推理任務中的表現。原始資料集以英文撰寫,為了促進繁體中文社群對該資料集的使用與研究,我們將其完整翻譯為繁體中文。 翻譯方法 我們採用了 Google translate 進行自動翻譯,確保問題和解答的語意與原始資料集一致。在翻譯過程中,我們特別注意數學符號、專有名詞和文化差異,確保繁體中文使用者能夠清晰理解每個問題。 資料集內容 翻譯後的資料集包含: 問題(question):小學數學問題的繁體中文描述。 答案(answer):對應問題的完整解答,包括多步驟推理和最終答案。 每個問題的解答格式保持與原始資料集一致,方便使用者進行比較和研究。 資料集結構 資料集分為訓練集和測試集: 訓練集:7,473 個問題 測試集:1,319 個問題 每個問題都需要 2 到 8… See the full description on the dataset page: https://huggingface.co/datasets/yuhuanstudio/gsm8k_zhtw.text1K<n<10K4 likes33 downloads1y agoHugging Face20OpenFormosa /zh_translation_benchmark zh_translation_benchmark zh_translation_benchmark is a 2,000-example synthetic benchmark for evaluating whether a translation or rewriting model can produce natural Taiwan Traditional Chinese (zh-TW) from English, Mainland Chinese, Hong Kong Traditional Chinese, Cantonese-style written Chinese, or code-mixed English/Chinese documents. The dataset now exposes a single Hugging Face subset/config: full. It is not split into dev and test files. The full 2,000 rows are loaded… See the full description on the dataset page: https://huggingface.co/datasets/OpenFormosa/zh_translation_benchmark.tabulartranslation1K<n<10K0 likes22 downloads4mo agoHugging Face21cccxi /gptoss_zhtw_alltext1K<n<10K1 likes12 downloads1y agoHugging Face22IMA-Taiwan /ima-corpus-zhtwgated IMA Traditional Chinese Corpus(繁體中文語料總集) 本資料集為繁體中文文學語料總集,目的在於將原先分散於多個作者/來源 dataset repo 的繁體中文文本統一整併,提供「一次申請、持續更新」的集中存取方式。 使用者只需申請本 dataset(本 repo)一次,即可取得所有繁中語料。未來新增來源或更新資料將直接同步至本 repo,無需重複申請。 📂 目錄結構 所有來源資料皆保留於 data/ 之下,每個子資料夾對應一個原始來源 repo,例如: data/ ├── zhtw-literature-ots 每個子資料夾內保留: 原始 README 原始語料檔(json / txt 等) 來源資訊與授權說明 以利來源追溯與資料審核。 📦 資料格式 主要格式: JSON UTF-8 編碼文字檔 典型欄位可能包含: title:作品名稱 author:作者 content:文本內容 source:來源 repo (依各來源資料實際格式而定)… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/ima-corpus-zhtw.texttext-generation1K<n<10K0 likes12 downloads8mo agoHugging Face23yuhuanstudio /poem-pretrain-chinese-zhtw Dataset Card for "poem-pretrain-chinese-zhtw" 資料集摘要 中文古典文集資料庫收集了約 5.5 萬首唐詩、26 萬首宋詩、2.1 萬首宋詞和其他古典文集。詩人包括唐宋兩朝近 1.4 萬古詩人,和兩宋時期 1.5 千古詞人。 五代十國- 收錄"花間集"與"南唐二主詞" 唐- 收錄"全唐詩"(是清康熙四十四年,康熙皇帝主導下,蒐集羅唐詩的收藏「得詩 48,900 餘首,詩入 2,200 人」)。 宋- 收錄"全宋詞"(由唐圭璋編著,孔凡禮補輯,共收錄宋代詞人 1,330 家,詞作 21,116 首)。 元- 收錄元曲 11,057 篇,曲家 233 人。 清- 收錄"納蘭性德詩集" 原始資料來源: chinese-poetry: 最全中文诗歌古典文集数据库 erhwenkuo/poetry-chinese-zhtw 資料集結構 { "text":… See the full description on the dataset page: https://huggingface.co/datasets/yuhuanstudio/poem-pretrain-chinese-zhtw.text10K<n<100K5 likes11 downloads2y agoHugging Face24IMA-Taiwan /zhtw-literature-otsgated Dataset Summary The dataset contains 2,349 rows. These paragraphs are extracted from authorized novels written by Ou Tiong Siong胡長松 and contain multiple sentences in Traditional Chinese. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 2,349 (each representing a paragraph) Features: title: Book title… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/zhtw-literature-ots.text1K<n<10K1 likes11 downloads1y agoHugging Face25cccxi /APIGen_ZHtwtext10K<n<100K0 likes10 downloads1y agoHugging Face26HowardHsuuu /taiwan-corpus-zhtwtextn<1K2 likes8 downloads1y agoHugging Face27cccxi /apigen_zhtw_donetext10K<n<100K0 likes8 downloads1y agoHugging Face28NYCU-312555007 /ZH-TW_Reading_Comprehension_Test_for_LLMstext10K<n<100K0 likes7 downloads2y agoHugging Face29ydqs /zh-test-llama3-v1text1K<n<10K0 likes6 downloads2y agoHugging Face30IMA-Taiwan /zhtw-literature-meowversegated Dataset Summary The dataset contains 14,256 rows. These paragraphs are extracted from authorized works written by Chang Luo-Miao張珞喵 and published on Vocus 方格子. Most articles are composed in Traditional Chinese, and many of them also carry parallel English, Japanese, German and Spanish versions of the same content within the same article, so the dataset is multilingual by design rather than translated afterwards. Roughly 43% of the paragraphs are Traditional Chinese, 54% English… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/zhtw-literature-meowverse.text10K<n<100K0 likes6 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.